Hindu College Of Engineering
Sonipat (Haryana)
Data Analytics With Python
Assignment
Submitted To: Submitted By:
Prof. Neeraj Goyal Nikita
23011001061
ASSIGNMENT NO:1
1) Differentiate between the following with the help of examples:-
1. Qualitative (nominal and ordinal) and Quantitative (interval and range)
Ans)
Qualitative Quantitative
• Qualitative data tells the features of the • Quantitative data is the type of data that
data in the statistics. represents the numerical value of the data.
• Qualitative Data is also called Categorical • They are also called Numerical Data.
Data.
• Qualitative data is further categorized into • Quantitative data is further classified into
two categories: two categories:
• Nominal Data • Interval Data
• Ordinal Data • Range Data
• Nominal Data: is a type of data that • Interval Data: A scale used to label
consists of categories or names that cannot variables that have a natural order
be ordered or ranked. • no “true zero” value.
• Examples: Gender (Male or female), • Examples: Credit Scores: Measured from
• Blood type (A, B, AB, O). 300 to 850
• SAT Scores: Measured from 400 to 1,600
• Ordinal Data: is a type of data that • Range Data: tells you how spread out the
consists of categories that can be ordered data is.
or ranked. • Range = Maximum value − Minimum
• Examples: Education level (Elementary, value
Middle, High School, College). • Example: If the ages are 10, 12, 15, and
18, the range = 18 − 10 = 8.
2. Various central of tendencies
Ans)
Mean Median Mode
• Arithmetic average of all • Middle value of ordered • Most frequent value
values data
• Highly affected by • Not affected by extreme • Not affected by extreme
extreme values values values
• Not suitable for skewed • Suitable for skewed data • May not exist or may be
data more than one
• Example: • Example: • Example:
• Data: 5, 10, 10, 15, 20 • Data: 5, 10, 10, 15, 20 • Data: 5, 10, 10, 15, 20
=(5+10+10+15+20)÷ 5 = Middle value = 10 = Most repeated value =
= 60 ÷ 5 = 12 10
3. Skewness and Kurtosis
Ans)
Skewness Kurtosis
• Measures asymmetry of a distribution • Measures peakedness and tail heaviness
• Shows left or right tilt • Shows sharpness of peak
• Indicates direction of skew • Indicates presence of extreme values
• Zero means symmetrical distribution • Zero means normal distribution
• Positive Skewness means right-skewed • Positive Kurtosis →Reflect a sharper
peak with heavier tails and more extreme
values
• Negative Skewness means left-skewed • Negative Kurtosis → Reflect a flatter
peak with lighter tails and fewer extreme
values
• Describes shape horizontally • Describes shape vertically and in the tails
4. Pearson’s Coefficient and Spearsman coefficient
Ans)
Pearson’s Coefficient Spearsman coefficient
• Measures linear relationships between • Measures monotonic relationships, where
variables variables move consistently in one
direction
• Works with continuous , interval or • Suitable for ordinal, ranked, interval, or
ratio data. ratio data.
• Sensitive to outliers, which can skew the • Resistant to outliers since it uses ranks
correlation value. instead of raw data.
• Based on covariance and standard • Based on ranking the data points and
deviations of raw values. calculating the difference in ranks.
• Ranges from -1 to 1 (negative, positive, or • Ranges from -1 to 1 (negative, positive, or
no linear correlation) no monotonic correlation).
5. Population and Sample
Ans)
Population Sample
• The population includes all members of a • A sample is a subset of the population
specified group
• It includes all members • It includes only selected members
• Usually very large in size • Relatively small in size
• Data collection is difficult • Data collection is easier
• Results are exact and accurate • Results are approximate
• Example: All households in a city • Example: 500 households from the city
QUES 2) The following are the figures of profits earned by 1400 companies during 2003-
04
Profits (Rs. Lakhs) Number of Companies
200-400 500
400-600 300
600-800 280
800-1000 120
1000-1200 100
1200-1400 80
1400-1600 20
a. Calculate the average profits for all the companies.
Ans)
Profits (Rs. Lakhs) Number of Companies(f) Midpoint(x) Total Profit(f*x)
200-400 500 300 150000
400-600 300 500 150000
600-800 280 700 196000
800-1000 120 900 108000
1000-1200 100 1100 110000
1200-1400 80 1300 104000
1400-1600 20 1500 30000
Total=1400 Total=868000
b. Calculate the median profit.
Profits (Rs. Lakhs) Number of Cumulative
Companies Frequency(cf)
200-400 500 500
400-600 300 500+300=800
600-800 280 800+280=1080
800-1000 120 1080+120=1200
1000-1200 100 1200+100=1300
1200-1400 80 1300+80=1380
1400-1600 20 1380+20=1400
N/2 = 1400/2 => 700
Median class = 400–600
Where:
• l = 400
• cf = 500
• f = 300
• h = 200
c. Write a python script to draw the for the above [Link] the horizontal line
indicating the mean, median and mode values in above graph
import numpy as np
import [Link] as plt
Profits=[(200,400),(400,600),(600,800),(800,1000),(1000,1200),(1200,1400),(1400,1600)]
x=[Link]([300,500,700,900,1100,1300,1500])
f=[Link]([500,300,280,120,100,80,20])
#--Histogram--
data=[Link](x,f)
[Link]()
bins=[200,400,600,800,1000,1200,1400,1600]
[Link](data,bins=bins)
[Link](mean, linestyle='--',label='Mean')
[Link](median, linestyle='-', label='Median')
[Link](mode,linestyle=':',label='Mode')
[Link]("Profit(Lakhs)")
[Link]("Frequency")
[Link]("Histogram showing ,Mean, Median, Mode")
[Link]()
[Link]()
print("Mode ", mode)
print("Mean Profit =",mean)
print("Median =",median)
d. Calculate the second quartile for the above data.
Ans) Second Quartile means median
import numpy as np
Profits=[(200,400),(400,600),(600,800),(800,1000),(1000,1200),(1200,1400),(1400,1600)]
x=[Link]([300,500,700,900,1100,1300,1500])
f=[Link]([500,300,280,120,100,80,20])
#--Median--
cf=[Link](f)
i=[Link](cf,[Link](f)/2)
L=Profits[i][0]
h=Profits[i][1] - Profits[i][0]
median= L+(([Link](f)/2 - (cf[i-1] if i>0 else 0)) / f[i])*h
print("Median =",median)
OUTPUT
Ques 3) Based on the frequency distribution given below, compute the following statistical
measures to characterise the distribution.
i) Coefficient of Variation
ii) Standard Deviation
iii) Modal Values
Annual Tax Paid (Rs. Thousand) No. of Managers
5-10 18
10-15 30
15-20 46
20-25 28
25-30 20
30-35 12
35-40 6
With the help of python script determine if the graph for above data is normal or not.
Ans)
Annual Tax Paid No. of Midpoint(x) fx x²f
(Rs. Thousand) Managers(f)
5-10 18 7.5 135 1012.5
10-15 30 12.5 375 4687.5
15-20 46 17.5 805 14087.5
20-25 28 22.5 630 14175
25-30 20 27.5 550 15125
30-35 12 32.5 390 12675
35-40 6 37.5 225 8437.5
N=160 ∑fx=3110 ∑fx2=701...
N=18+30+46+28+20+12+6=160
(i) Mean
ii) Standard Deviation
iii) Modal Value
Modal class = 15–20 (highest frequency = 46)
(iv) Coefficient of Variation (C.V.)
With the help of python script determine if the graph for above data is normal or not.
Ans)
import [Link] as plt
import numpy as np
classes = [(5,10),(10,15),(15,20),(20,25),(25,30),(30,35),(35,40)]
frequencies = [18,30,46,28,20,12,6]
midpoints = [(a+b)/2 for a,b in classes]
[Link](midpoints, frequencies, width=4)
[Link]("Annual Tax Paid (Rs. Thousand)")
[Link]("Number of Managers")
[Link]("Distribution of Annual Tax Paid")
[Link]()
Output) The distribution is not normal; it is positively skewed (right-skewed).
Ques 4) Ten competitors in beauty contents are ranked by three judges in the following
order:
1st Judge 1 6 5 10 3 2 4 9 7
2nd Judge 3 5 8 4 7 10 2 1 6
3rd Judge 6 4 9 8 1 2 3 10 5
Use the rank correlation coefficient to determine which pair of judges has the nearest approach
to common tastes in beauty.
Ans)
Formula (Spearman’s Rank Correlation)
1. Correlation between Judge 1 and Judge 2
2. Correlation between Judge 1 and Judge 3
3. Correlation
between Judge
2 and Judge 3
Judge 2 and Judge 3 have the highest positive rank correlation, indicating the closest similarity in
judging beauty.