SAMANTUS Python for Data Analytics — Complete Training Manual CAPSTONE PROJECT 07
← Back to Course Index

HR Analytics


On This Page

Capstone Project 7: HR Analytics


C7.1  Project Overview

This project analyzes company-wide workforce data to understand headcount distribution and the relationship between job satisfaction and attrition risk.

C7.2  Business Problem

The HR leadership team wants a clear picture of where employees are concentrated across departments, and whether job satisfaction meaningfully predicts who is likely to leave.

C7.3  Dataset Description

ColumnDescription
employee_idUnique employee identifier
departmentDepartment name
satisfaction_scoreJob satisfaction score (1-10)
attritionWhether the employee left (1) or stayed (0)
salary_bandSalary band (Low, Medium, High)

Sample scope: 200 employees across 5 departments.

C7.4  Objectives


C7.5  Step-by-Step Solution

C7.6  Complete Python Code

► hr_analytics.py
import pandas as pd
import matplotlib.pyplot as plt

df = pd.read_csv("hr_workforce.csv")
df = df.dropna(subset=['satisfaction_score'])

headcount = df['department'].value_counts()
print("Headcount by Department:\n", headcount)

avg_satisfaction = df.groupby('attrition')['satisfaction_score'].mean()
print("Avg Satisfaction (0=Stayed, 1=Left):\n", avg_satisfaction.round(2))

plt.scatter(df["satisfaction_score"], df["attrition"], alpha=0.5)
plt.title("Satisfaction vs Attrition")
plt.savefig("satisfaction_scatter.png")

C7.7  Expected Output

Output
Avg Satisfaction (0=Stayed, 1=Left):
0   7.4
1   4.8

C7.8  Visualizations

Figure C7.1 — Headcount distribution across departments.
Figure C7.1 — Headcount distribution across departments.
Figure C7.2 — Relationship between job satisfaction and attrition.
Figure C7.2 — Relationship between job satisfaction and attrition.

C7.9  Key Insights

C7.10  Business Recommendations