Exploratory Data Analysis (EDA) is the process of investigating a dataset before drawing conclusions or building models — understanding its structure, spotting problems, and finding initial patterns. This is where every real data analytics project actually begins.
import pandas as pd
df = pd.read_csv("ecommerce_orders.csv")
print(df.shape) # rows, columns
print(df.head()) # first 5 rows
print(df.info()) # data types, non-null counts
print(df.describe()) # statistical summarydf.columns = df.columns.str.strip().str.lower() # clean column names
df = df.drop_duplicates() # remove duplicate rows
df['order_date'] = pd.to_datetime(df['order_date']) # fix data typesprint(df.isnull().sum())
df['discount'] = df['discount'].fillna(0)
df = df.dropna(subset=['customer_id']) # drop rows missing critical infoOutliers are unusually extreme values that can distort averages and mislead analysis. Box plots are the standard tool for spotting them visually.
import matplotlib.pyplot as plt
plt.boxplot(df['order_value'], vert=False)
plt.title("Order Value Distribution - Outlier Detection")
plt.show()Q1 = df['order_value'].quantile(0.25)
Q3 = df['order_value'].quantile(0.75)
IQR = Q3 - Q1
lower = Q1 - 1.5*IQR
upper = Q3 + 1.5*IQR
df_clean = df[(df['order_value'] >= lower) & (df['order_value'] <= upper)]Before analyzing, classify each column as numeric or categorical, and understand what each represents in the business context.
| Column | Type | Business Meaning |
|---|---|---|
| order_value | Numeric | Total amount of the order in Rs. |
| category | Categorical | Product category (Electronics, Fashion, etc.) |
| order_date | Datetime | When the order was placed |
| customer_id | Categorical (ID) | Unique identifier for each customer |
| discount | Numeric | Discount applied to the order |
print(df[['order_value','discount','quantity']].corr())A correlation value close to +1 or -1 indicates a strong relationship; values near 0 suggest little to no linear relationship between the two columns.
revenue_by_category = df.groupby('category')['order_value'].sum().sort_values(ascending=False)
revenue_by_category.plot(kind="bar", color="navy")
plt.title("Revenue by Product Category")
plt.show()Business Problem: An online store wants to understand what's driving sales and where to focus marketing spend.
import pandas as pd
import matplotlib.pyplot as plt
df = pd.read_csv("ecommerce_orders.csv")
# 1. Understand
print(df.shape, df.info())
# 2. Clean
df = df.drop_duplicates()
df['discount'] = df['discount'].fillna(0)
# 3. Handle Outliers (IQR method)
Q1, Q3 = df['order_value'].quantile([0.25, 0.75])
IQR = Q3 - Q1
df = df[df['order_value'] <= Q3 + 1.5*IQR]
# 4. Analyze
top_category = df.groupby('category')['order_value'].sum().idxmax()
avg_order = df['order_value'].mean()
print("Top Category:", top_category)
print("Average Order Value:", round(avg_order, 2))Basic (5 Questions)
1. Load any CSV dataset and print its shape, info(), and describe().
2. Check for missing values in your dataset and count them per column.
3. Remove duplicate rows from a sample dataset.
4. Create a box plot for one numeric column to check for outliers.
5. List the data type (numeric/categorical) of each column in your dataset.
Intermediate (5 Questions)
1. Use the IQR method to detect and remove outliers from a numeric column.
2. Create a correlation matrix for 3+ numeric columns in your dataset.
3. Group your dataset by a categorical column and visualize the totals with a bar chart.
4. Fill missing values in a numeric column using its median instead of 0.
5. Write 3 business insights based on a groupby() analysis of your own dataset.
Advanced (5 Questions)
1. Perform a complete EDA (understand, clean, handle missing values, detect outliers, visualize) on a real public dataset (e.g. Titanic or a sample sales CSV).
2. Build a correlation heatmap and identify the two most strongly correlated variables.
3. Detect and separately analyze outlier records — are they errors, or legitimate rare cases?
4. Create a complete written EDA summary report (1 page) with at least 5 business insights.
5. Compare two subgroups in your dataset (e.g. two regions or two time periods) and explain any meaningful differences you find.
Q1. What is the main goal of EDA?
Q2. Which method is commonly used to detect outliers?
Q3. What does a correlation value near 0 indicate?
Q4. Which function shows column data types and missing value counts?
Q5. Why should you always end EDA with business insights?
Assignment: Complete EDA on a Retail Sales Dataset
Expected Output: A complete, well-documented EDA notebook or script covering every stage from raw data to clear business recommendations.
Use this space to write down key points, doubts, and your own examples from today's session.