0% found this document useful (0 votes)
4 views4 pages

Python Code Employees

This case study outlines the analysis of a dataset containing information on 5000 employees, focusing on various HR analytics such as total employee count, department-wise distribution, average salary, and payroll costs. It includes advanced analyses like salary by department, experience vs salary correlation, and categorization into salary bands. Additionally, it emphasizes data cleaning practices and provides business insights for management decisions.

Uploaded by

asifasshu6
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views4 pages

Python Code Employees

This case study outlines the analysis of a dataset containing information on 5000 employees, focusing on various HR analytics such as total employee count, department-wise distribution, average salary, and payroll costs. It includes advanced analyses like salary by department, experience vs salary correlation, and categorization into salary bands. Additionally, it emphasizes data cleaning practices and provides business insights for management decisions.

Uploaded by

asifasshu6
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

CASE STUDY: EMPLOYEE’S DATA ANALYSIS

Step 1: Import and Load Dataset


import pandas as pd

df=pd.read_csv("F:\\DESKTOP\\ACSE\\DHV MINOR\\Case study\\[Link]")


Meaning:
Loads the 5000-employee dataset into a DataFrame.

Step 2: Initial Inspection


[Link]()
[Link]
[Link]()
[Link]()
We check:
• Number of rows (should be 5000)
• Number of columns
• Data types
• Summary statistics

MODULE 2: BASIC HR ANALYTICS


2.1 Total Employees
[Link][0]
Gives total employee count.

2.2 Department-wise Employee Count


df["Department"].value_counts()
Meaning:
Shows distribution of employees across departments.

2.3 Average Salary


df["Salary"].mean()
Meaning:
Calculates company-wide average salary.
2.4 Highest and Lowest Salary
df["Salary"].max()
df["Salary"].min()

2.5 Total Payroll Cost


df["Salary"].sum()
Meaning:
Total monthly salary expenditure.

MODULE 3: ADVANCED ANALYSIS


3.1 Salary by Department
[Link]("Department")["Salary"].mean()
Insight:
Identifies highest paying department.

3.2 Experience vs Salary Relationship


df[["Salary","Experience"]].corr()
If correlation is positive → Salary increases with experience.

3.3 Top 10 Highest Paid Employees


[Link](10, "Salary")

3.4 Bottom 10 Lowest Paid Employees


[Link](10, "Salary")

MODULE 4: CATEGORIZATION
4.1 Salary Bands
df["Salary_Band"] = [Link](
df["Salary"],
bins=[0,50000,80000,120000],
labels=["Low","Medium","High"]
)
Meaning:
Classifies employees into salary categories.

4.2 Senior vs Junior Employees


import numpy as np

df["Level"] = [Link](df["Experience"] >= 10, "Senior","Junior")

MODULE 5: STATISTICAL ANALYSIS


5.1 Standard Deviation of Salary
df["Salary"].std()
Measures salary variation.

5.2 Skewness
df["Salary"].skew()
Checks whether salary distribution is skewed.

5.3 Quartiles
df["Salary"].quantile([0.25,0.50,0.75])
Shows salary percentiles.

MODULE 6: WINDOW ANALYSIS


6.1 Rolling Average Salary (5 Employees)
df["Salary"].rolling(5).mean()

6.2 Cumulative Payroll


df["Salary"].cumsum()

MODULE 7: CITY ANALYSIS


7.1 Employees per City
df["City"].value_counts()
7.2 Average Salary per City
[Link]("City")["Salary"].mean()

MODULE 8: DATA CLEANING PRACTICE


Check Missing Values
[Link]().sum()
Remove Duplicates
df.drop_duplicates()

MODULE 9: BUSINESS INSIGHTS


From analysis, management can determine:
• Highest salary department
• Salary distribution spread
• Experience impact
• Payroll burden
• Regional distribution

You might also like