This module brings together everything learned so far — NumPy, Pandas, Matplotlib, and Seaborn — into complete, real, end-to-end analytics workflows, exactly as they're used in the industry.
| Stage | Tools Used | Goal |
|---|---|---|
| 1. Load Data | Pandas (read_csv/read_excel) | Bring raw data into Python |
| 2. Explore | df.info(), describe(), head() | Understand structure & quality |
| 3. Clean | Pandas (dropna, fillna, dedupe) | Fix missing values, duplicates |
| 4. Transform | Pandas, NumPy | Create new columns, aggregate, reshape |
| 5. Visualize | Matplotlib, Seaborn | Reveal patterns and trends |
| 6. Report Insights | Written summary + charts | Communicate to stakeholders |
Business Question: Which region and product category should get more marketing budget next quarter?
import pandas as pd
df = pd.read_csv("retail_sales.csv")
df = df.dropna(subset=['revenue'])
region_perf = df.groupby('region')['revenue'].sum().sort_values(ascending=False)
category_perf = df.groupby('category')['revenue'].sum().sort_values(ascending=False)
print("Top Region:", region_perf.index[0])
print("Top Category:", category_perf.index[0])
growth = df.groupby('month')['revenue'].sum().pct_change() * 100
print("Month-over-Month Growth %:\n", growth.round(2))Business Question: Which department has the highest employee attrition, and why?
import pandas as pd
df = pd.read_csv("hr_data.csv")
attrition_rate = df.groupby('department')['attrition'].mean() * 100
print("Attrition Rate by Department (%):\n", attrition_rate.round(2))
avg_tenure = df.groupby('department')['years_at_company'].mean()
print("Average Tenure by Department:\n", avg_tenure.round(1))
correlation = df[['satisfaction_score','attrition']].corr()
print(correlation)Business Question: Which marketing channel gives the best return on ad spend?
import pandas as pd
df = pd.read_csv("campaign_data.csv") # columns: channel, spend, revenue
df['roi'] = (df['revenue'] - df['spend']) / df['spend'] * 100
channel_roi = df.groupby('channel')['roi'].mean().sort_values(ascending=False)
print("Average ROI % by Channel:\n", channel_roi.round(2))
best_channel = channel_roi.idxmax()
print("Best performing channel:", best_channel)Professional analysts wrap repeated logic into functions, making analysis faster and less error-prone across projects.
import pandas as pd
def quick_summary(df, group_col, value_col):
"""Returns total, average, and count grouped by a category."""
summary = df.groupby(group_col)[value_col].agg(["sum","mean","count"])
summary.columns = ['Total', 'Average', 'Count']
return summary.sort_values("Total", ascending=False)
result = quick_summary(df, "category", "revenue")
print(result)result.to_csv("category_summary.csv")
result.to_excel("category_summary.xlsx", sheet_name="Summary")
import matplotlib.pyplot as plt
result["Total"].plot(kind="bar", title="Revenue by Category")
plt.savefig("category_chart.png", dpi=150, bbox_inches="tight")Basic (5 Questions)
1. Load any dataset and print a full summary (shape, info, describe).
2. Group your dataset by one categorical column and calculate the sum of a numeric column.
3. Create one bar chart summarizing your groupby result.
4. Export a Pandas DataFrame to both a CSV and Excel file.
5. Write a one-function summary that takes a DataFrame and returns basic stats.
Intermediate (5 Questions)
1. Recreate the Retail Sales Analysis example using your own sample dataset.
2. Calculate month-over-month growth % for a revenue column using pct_change().
3. Build a reusable function quick_summary() and test it on 2 different datasets.
4. Create and export a chart alongside your summary CSV, matching filenames.
5. Write 3 business insights based on your own groupby() analysis.
Advanced (5 Questions)
1. Build a complete end-to-end workflow (load → clean → transform → visualize → export) on a dataset of your choice, following all 6 workflow stages.
2. Recreate the HR Attrition or Marketing ROI case study using your own synthetic dataset.
3. Build a toolkit of 3 reusable analytics functions (summary, outlier check, chart export) and use them together on one dataset.
4. Create a 1-page 'Executive Summary' combining 2 charts and 3 written business recommendations.
5. Compare 2 different time periods in a dataset (e.g. Q1 vs Q2) and write a comparative analysis.
Q1. What is typically the FIRST stage of any analytics workflow?
Q2. Which function calculates period-over-period growth %?
Q3. Why build reusable functions for analysis?
Q4. Which format is best for sharing results with non-technical managers?
Q5. What should every analytics workflow end with?
Assignment: Complete Business Analytics Workflow
Expected Output: A complete, professional analytics project script from raw data to a stakeholder-ready summary, following the industry-standard workflow.
Use this space to write down key points, doubts, and your own examples from today's session.