SAMANTUS Python for Data Analytics — Complete Training Manual MODULE 16
← Back to Course Index

Exploratory Data Analysis (EDA)


On This Page

Module 16: Exploratory Data Analysis (EDA)


16.1  What is EDA?

Exploratory Data Analysis (EDA) is the process of investigating a dataset before drawing conclusions or building models — understanding its structure, spotting problems, and finding initial patterns. This is where every real data analytics project actually begins.

✓ Instructor Tip
Tell students: 'You should never trust a dataset until you've explored it yourself — EDA is how you build that trust.'

16.2  Understanding the Dataset

► First Steps on Any New Dataset
import pandas as pd

df = pd.read_csv("ecommerce_orders.csv")

print(df.shape)         # rows, columns
print(df.head())        # first 5 rows
print(df.info())        # data types, non-null counts
print(df.describe())    # statistical summary

16.3  Cleaning Data

► Common Cleaning Steps
df.columns = df.columns.str.strip().str.lower()   # clean column names
df = df.drop_duplicates()                          # remove duplicate rows
df['order_date'] = pd.to_datetime(df['order_date'])  # fix data types

16.4  Handling Missing Values

► Investigating & Handling Missing Data
print(df.isnull().sum())

df['discount'] = df['discount'].fillna(0)
df = df.dropna(subset=['customer_id'])   # drop rows missing critical info
✗ Common Mistake
Filling every missing value with 0 without thinking can badly distort averages. Decide column by column: does 0 make business sense, or is the column mean/median more appropriate?

16.5  Detecting Outliers

Outliers are unusually extreme values that can distort averages and mislead analysis. Box plots are the standard tool for spotting them visually.

► Detecting Outliers with a Box Plot
import matplotlib.pyplot as plt

plt.boxplot(df['order_value'], vert=False)
plt.title("Order Value Distribution - Outlier Detection")
plt.show()
Figure 16.1 — Order values cluster between Rs.200-2000, with 3 extreme outliers above Rs.8500.
Figure 16.1 — Order values cluster between Rs.200-2000, with 3 extreme outliers above Rs.8500.
ℹ Interpretation
Three orders above Rs.8,500 stand far apart from the rest — these could be bulk B2B orders, data-entry errors, or fraud attempts, and should always be investigated individually before deciding whether to keep or remove them.
► Removing Outliers using IQR Method
Q1 = df['order_value'].quantile(0.25)
Q3 = df['order_value'].quantile(0.75)
IQR = Q3 - Q1
lower = Q1 - 1.5*IQR
upper = Q3 + 1.5*IQR

df_clean = df[(df['order_value'] >= lower) & (df['order_value'] <= upper)]

16.6  Feature Understanding

Before analyzing, classify each column as numeric or categorical, and understand what each represents in the business context.

ColumnTypeBusiness Meaning
order_valueNumericTotal amount of the order in Rs.
categoryCategoricalProduct category (Electronics, Fashion, etc.)
order_dateDatetimeWhen the order was placed
customer_idCategorical (ID)Unique identifier for each customer
discountNumericDiscount applied to the order

16.7  Correlation Analysis

► Checking Correlation Between Numeric Columns
print(df[['order_value','discount','quantity']].corr())

A correlation value close to +1 or -1 indicates a strong relationship; values near 0 suggest little to no linear relationship between the two columns.

16.8  Visualization

► Visualizing Category Performance
revenue_by_category = df.groupby('category')['order_value'].sum().sort_values(ascending=False)
revenue_by_category.plot(kind="bar", color="navy")
plt.title("Revenue by Product Category")
plt.show()
Figure 16.2 — Total revenue generated by each product category.
Figure 16.2 — Total revenue generated by each product category.

16.9  Complete Case Study — E-Commerce Sales EDA

Business Problem: An online store wants to understand what's driving sales and where to focus marketing spend.

► ecommerce_eda.py
import pandas as pd
import matplotlib.pyplot as plt

df = pd.read_csv("ecommerce_orders.csv")

# 1. Understand
print(df.shape, df.info())

# 2. Clean
df = df.drop_duplicates()
df['discount'] = df['discount'].fillna(0)

# 3. Handle Outliers (IQR method)
Q1, Q3 = df['order_value'].quantile([0.25, 0.75])
IQR = Q3 - Q1
df = df[df['order_value'] <= Q3 + 1.5*IQR]

# 4. Analyze
top_category = df.groupby('category')['order_value'].sum().idxmax()
avg_order = df['order_value'].mean()

print("Top Category:", top_category)
print("Average Order Value:", round(avg_order, 2))
Output
Top Category: Electronics
Average Order Value: 1198.45

16.10  Deriving Business Insights

✓ Instructor Tip
Always end an EDA with 3-5 plain-English business insights — this is what separates a data analyst from someone who just runs code.

16.11  Practical Exercises

Basic (5 Questions)

1. Load any CSV dataset and print its shape, info(), and describe().

2. Check for missing values in your dataset and count them per column.

3. Remove duplicate rows from a sample dataset.

4. Create a box plot for one numeric column to check for outliers.

5. List the data type (numeric/categorical) of each column in your dataset.

Intermediate (5 Questions)

1. Use the IQR method to detect and remove outliers from a numeric column.

2. Create a correlation matrix for 3+ numeric columns in your dataset.

3. Group your dataset by a categorical column and visualize the totals with a bar chart.

4. Fill missing values in a numeric column using its median instead of 0.

5. Write 3 business insights based on a groupby() analysis of your own dataset.

Advanced (5 Questions)

1. Perform a complete EDA (understand, clean, handle missing values, detect outliers, visualize) on a real public dataset (e.g. Titanic or a sample sales CSV).

2. Build a correlation heatmap and identify the two most strongly correlated variables.

3. Detect and separately analyze outlier records — are they errors, or legitimate rare cases?

4. Create a complete written EDA summary report (1 page) with at least 5 business insights.

5. Compare two subgroups in your dataset (e.g. two regions or two time periods) and explain any meaningful differences you find.

16.12  Module Quiz (MCQs)

Q1. What is the main goal of EDA?

Q2. Which method is commonly used to detect outliers?

Q3. What does a correlation value near 0 indicate?

Q4. Which function shows column data types and missing value counts?

Q5. Why should you always end EDA with business insights?

ℹ Answer Key
1-b, 2-b, 3-b, 4-b, 5-b

16.13  Module Assignment

Assignment: Complete EDA on a Retail Sales Dataset

Expected Output: A complete, well-documented EDA notebook or script covering every stage from raw data to clear business recommendations.

16.14  Interview Questions — Module 16


16.15  Student Notes Page

Use this space to write down key points, doubts, and your own examples from today's session.