This final capstone project analyzes transaction data to identify patterns associated with fraudulent activity — a foundational, rule-based approach to fraud analytics before moving into machine learning.
C15.2 Business Problem
A payments company wants to understand what typically distinguishes fraudulent transactions from normal ones — by amount and by time of day — to build simple early-warning flagging rules.
C15.3 Dataset Description
Column
Description
transaction_id
Unique transaction identifier
amount
Transaction amount (Rs.)
hour
Hour of day the transaction occurred (0-23)
label
Normal or Fraud (based on historical review)
Sample scope: 500 transactions, including 20 confirmed fraud cases.
C15.4 Objectives
Compare transaction amounts between normal and fraudulent transactions
Analyze whether fraud rate varies by hour of day
Identify simple, explainable rules that flag likely fraud
Recommend a practical monitoring approach for the business
C15.5 Step-by-Step Solution
Step 1: Load the transaction dataset into Pandas
Step 2: Compare amount distributions for Normal vs Fraud using a box plot
Step 3: Calculate fraud rate by hour of day
Step 4: Define a simple rule-based flag (e.g. amount > threshold)
Step 5: Test the rule's accuracy against known fraud labels
Step 6: Summarize a practical fraud-monitoring recommendation
C15.6 Complete Python Code
► fraud_detection.py
import pandas as pd
import matplotlib.pyplot as plt
df = pd.read_csv("transactions.csv")
avg_by_label = df.groupby('label')['amount'].mean()
print("Average Amount - Normal vs Fraud:\n", avg_by_label.round(2))
fraud_rate_by_hour = df.groupby('hour')['label'].apply(lambda x: (x=='Fraud').mean()*100)
print("Fraud Rate % by Hour:\n", fraud_rate_by_hour.round(1))
# Simple rule-based flag
threshold = df[df['label']=='Fraud']['amount'].quantile(0.25)
df['flagged'] = df['amount'] > threshold
accuracy = (df['flagged'] == (df['label']=='Fraud')).mean() * 100
print(f"Simple Rule Accuracy: {accuracy:.1f}%")
C15.7 Expected Output
Output
Average Amount - Normal vs Fraud: Normal 1,480 Fraud 8,650
C15.8 Visualizations
Figure C15.1 — Transaction amount comparison between normal and fraudulent transactions.Figure C15.2 — Fraud rate (%) by hour of day.
C15.9 Key Insights
Fraudulent transactions average Rs.8,650 — nearly 6x higher than the average normal transaction (Rs.1,480), making amount a strong initial flagging signal.
Fraud shows no single obvious time-of-day concentration in this sample, suggesting amount is a more reliable signal than timing alone for this dataset.
A simple amount-threshold rule alone can catch a meaningful share of fraud cases, showing that even basic rule-based analytics provide real protective value before adding ML models.
C15.10 Business Recommendations
Implement an immediate manual-review flag for any transaction above the identified amount threshold, as a simple first line of defense.
Continue collecting labeled fraud data to eventually train a proper machine learning classification model, using this rule-based analysis as the baseline to beat.
Combine amount-based rules with other signals (device, location, transaction frequency) in future iterations for a more robust, layered fraud-detection system.
✓ Instructor Tip
This project is an excellent bridge to any future Machine Learning module the institute may add — frame it as 'fraud detection before ML' to set that expectation with students.