CASE STUDY: EMPLOYEE’S DATA ANALYSIS
Step 1: Import and Load Dataset
import pandas as pd
df=pd.read_csv("F:\\DESKTOP\\ACSE\\DHV MINOR\\Case study\\[Link]")
Meaning:
Loads the 5000-employee dataset into a DataFrame.
Step 2: Initial Inspection
[Link]()
[Link]
[Link]()
[Link]()
We check:
• Number of rows (should be 5000)
• Number of columns
• Data types
• Summary statistics
MODULE 2: BASIC HR ANALYTICS
2.1 Total Employees
[Link][0]
Gives total employee count.
2.2 Department-wise Employee Count
df["Department"].value_counts()
Meaning:
Shows distribution of employees across departments.
2.3 Average Salary
df["Salary"].mean()
Meaning:
Calculates company-wide average salary.
2.4 Highest and Lowest Salary
df["Salary"].max()
df["Salary"].min()
2.5 Total Payroll Cost
df["Salary"].sum()
Meaning:
Total monthly salary expenditure.
MODULE 3: ADVANCED ANALYSIS
3.1 Salary by Department
[Link]("Department")["Salary"].mean()
Insight:
Identifies highest paying department.
3.2 Experience vs Salary Relationship
df[["Salary","Experience"]].corr()
If correlation is positive → Salary increases with experience.
3.3 Top 10 Highest Paid Employees
[Link](10, "Salary")
3.4 Bottom 10 Lowest Paid Employees
[Link](10, "Salary")
MODULE 4: CATEGORIZATION
4.1 Salary Bands
df["Salary_Band"] = [Link](
df["Salary"],
bins=[0,50000,80000,120000],
labels=["Low","Medium","High"]
)
Meaning:
Classifies employees into salary categories.
4.2 Senior vs Junior Employees
import numpy as np
df["Level"] = [Link](df["Experience"] >= 10, "Senior","Junior")
MODULE 5: STATISTICAL ANALYSIS
5.1 Standard Deviation of Salary
df["Salary"].std()
Measures salary variation.
5.2 Skewness
df["Salary"].skew()
Checks whether salary distribution is skewed.
5.3 Quartiles
df["Salary"].quantile([0.25,0.50,0.75])
Shows salary percentiles.
MODULE 6: WINDOW ANALYSIS
6.1 Rolling Average Salary (5 Employees)
df["Salary"].rolling(5).mean()
6.2 Cumulative Payroll
df["Salary"].cumsum()
MODULE 7: CITY ANALYSIS
7.1 Employees per City
df["City"].value_counts()
7.2 Average Salary per City
[Link]("City")["Salary"].mean()
MODULE 8: DATA CLEANING PRACTICE
Check Missing Values
[Link]().sum()
Remove Duplicates
df.drop_duplicates()
MODULE 9: BUSINESS INSIGHTS
From analysis, management can determine:
• Highest salary department
• Salary distribution spread
• Experience impact
• Payroll burden
• Regional distribution