0% found this document useful (0 votes)
4 views67 pages

Descriptive Statistics: Central Tendency

Uploaded by

kunaldeb866
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as KEY, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views67 pages

Descriptive Statistics: Central Tendency

Uploaded by

kunaldeb866
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as KEY, PDF, TXT or read online on Scribd

Module 3: Descriptive Statistics 9

Hours

Arithmetic mean
Averages of position:
Median
Quartiles
Deciles
Percentiles
Mode
Standard deviation
co-efficient of variation.

[Link].a
[Link]
Arithmetic Mean
Arithmetic mean

[Link].a
[Link]
Introduction

Central Tendency
Central tendency is a descriptive summary of a dataset through a
single value that reflects the center of the data distribution. Along
with the variability (dispersion) of a dataset, central tendency is a
branch of descriptive statistics.

[Link].a
[Link]
Measures of Central Tendency
Generally, the central tendency of a dataset can be described using the
following measures:

Mean (Average): Represents the sum of all values in a dataset divided


by the total number of the values.

Median: The middle value in a dataset that is arranged in ascending order


(from the smallest value to the largest value). If a dataset contains an even
number of values, the median of the dataset is the mean of the two middle
values.

Mode: Defines the most frequently occurring value in a dataset.

[Link].a
[Link]
Arithmetic Mean

Averages are useful because they:


Summarise a large amount of data into a single value; and
indicate that there is some variability around this single value within
the original data
In general language arithmetic mean is same as the average of data. It
is the representative value of the group of data. The arithmetic mean
between two numbers is defined to be the sum of numbers divided by
the quantity of numbers.

[Link].a
[Link]
Formula for Arithmetic Mean

[Link].a
[Link]
Formula for Arithmetic Mean - Grouped data

[Link].a
[Link]
Ungrouped data – problems
1. Ramesh has been working on programing and updating a Web site
for his company for the past 15 months. The following numbers
represent the number of hours Ramesh has worked on this Web site for
each of the past 7 months: 24, 25, 31, 50, 53, 66, 78
What is the mean (average) number of hours that Stephen worked on
this Web site each month?

[Link].a
[Link]
Solution
Step 1: Add the numbers to determine the total number of hours he
worked.
24 + 25 + 33 + 50 + 53 + 66 + 78 = 329

Step 2: Divide the total by the number of months.


329/7 = 47

The mean number of hours that Rakesh worked each month was 47.

[Link].a
[Link]
Ungrouped data – problems
2. From 3rd August onwards, for 14 continuous days, Mark had an
average of 24.5 hits on his Web site per day. In the first two days that
Technology Titans was open for business, that is on 1st and 2nd August
and the Web site received 42 and 53 hits respectively. Determine the
new average for hits on the Web site for 16 days altogether.

[Link].a
[Link]
Solution

Step 1: Multiply the given average by 14 to determine the total number of hits
on Mark's Web site.
24.5 x 14 = 343
Step 2: Add the hits for the first two days his business was open.
343 + 42 + 53 = 438
Step 3: Divide this new total by 16 to determine the new average.
Mean = 438/16 = 27.375
The average number of hits Mark's Web site has received per day
since it was launched is 27.375.

[Link].a
[Link]
Arithmetic mean for discrete
series

[Link].a
[Link]
Numericals
From the following information on the number of
defective components in 1000 boxes:

Calculate the arithmetic mean of defective


components for the whole of the production line.

[Link].a
[Link]
Numericals
A company is planning to improve plant safety. For
this, accident data for the last 50 weeks was
compiled. These data are grouped into the
frequency distribution as shown below. Calculate the
Arithmetic mean of the number of accidents per
week. Use direct method.

[Link].a
[Link]
Numericals
168 handloom factories have the following
distribution of average number of workers in various
income groups:

Find the mean salary paid to the workers.

[Link].a
[Link]
Numericals
The pass result of 50 students who took a
class test is given below. If the average
marks for all the students were 51.6, find
out the average marks of the students who
failed.

[Link].a
[Link]
Median

Median may be defined as the middle value such that half of the
observations are smaller and other half are larger than this value,
provided the data set are arranged in a sequential order of
magnitude.

[Link].a
[Link]
Median of ungrouped data

If N is odd
Median =

If N is even

[Link].a
[Link]
Median of grouped data

Median = L+ ( )h

N= Number of values or total frequency


L= lower limit of the median class
f = frequency of the median class
cf = cumulative frequency of the preceding the
median class
h = class size of the median class = UL-LL

[Link].a
[Link]
Median of ungrouped data A top running athlete in a typical 200-
metre training session runs in the following times: 26.1, 25.6, 25.7,
25.2 , 25.0 [Link] would you calculate his median time?

[Link].a
[Link]
Median of ungrouped data A top running athlete in a typical 200-
metre training session runs in the following times: 26.1, 25.6, 25.7,
25.2 et 25.0 [Link] would you calculate his median time?
First, the values are put in ascending order: 25.0, 25.2, 25.6, 25.7, 26.1.
Then, using the following formula, figure out which value is the middle
value. Remember that n represents the number of values in the data
set.
Median = {(n + 1) ÷ 2}th value = (5 + 1) ÷ 2= 3
The third value in the data set will be the median. Since 25.6 is the third
value, 25.6 seconds would be the median time.
Median= 25.6 seconds

[Link].a
[Link]
Median of grouped data The following frequency distribution
gives the monthly consumption of electricity of 68 consumers of
a locality. Find the median.

[Link].a
[Link]
Solution:
Step 1: Find Cumulative Frequency
Step 2: Find Median Class
Step 3: Apply formula to find Median
Applying the formula,
Median = 125 + (34 -22)/20 * 20
= 137

[Link].a
[Link]
Numericals
A survey was conducted to determine the age (in years) of 120
automobiles. The result of such a survey is as follows:

What is the median age for the autos?

[Link].a
[Link]
CALCULATION OF QUARTILES, DECILES & PERCENTILES

[Link].a
[Link]
Quartiles
The values of the variate which divide the total frequency into four equal parts, are called quartiles. That value of the variate which
divides the total frequency into two equal parts is called median. The lower quartile or first quartile denoted by Q1 divides the
frequency between the lowest value and the median into two equal parts and similarly the upper quartile (or third quartile) denoted
by Q3 divides the frequency between the median and the greatest value into two equal parts. The formulas for computation of
quartiles are given by

n= Number of values or total frequency


l= lower limit of the median class
f = frequency of the median class
cf = cumulative frequency of the preceding the median class
h = class size of the median class = UL-LL

[Link].a
[Link]
Deciles

The values of the variate which divide the total frequency into ten equal
parts are called deciles. The formulas for computation are given by

where l, cf , n, f, i. have the same meaning as in the formula for median.

[Link].a
[Link]
Percentiles

[Link].a
[Link]
Some important observations :

[Link].a
[Link]
Quartiles, Deciles and Percentiles for grouped
data
Examples1. Calculate Quartile-3, Deciles-7, Percentiles-20 from the following
grouped data

[Link].a
[Link]
Solution:

Solution

[Link].a
[Link]
Solution
Here, n=10Q3 class :Class with (3n/4)th value of the
observation in cf column=(3⋅10/4)th value of the
observation in cf column=(7.5)th value of the
observation in cf columnand it lies in the class 6-8.
∴Q3 class : 6-8The lower boundary point of 6-8 is 6.
∴L=6

[Link].a
[Link]
D7 class :Class with (7n/10)th value of the
observation in cf column
=(7⋅10/10)th value of the observation
in cf column=(7)th value of the
observation in cf columnand it lies in the
class 4-6.∴D7 class : 4-6The lower
boundary point of 4-6 is 4.∴L=4

[Link].a
[Link]
P20 class :Class with (20n/100)th value of the
observation in cf column=(20⋅10/100)th value
of the observation in cf column=(2)th value of
the observation in cf columnand it lies in the
class 2-4.∴P20 class : 2-4The lower boundary
point of 2-4 is 2.∴L=2

[Link].a
[Link]
Numericals
The following distribution gives the pattern of overtime work per week
done by 100 employees of a company.
Calculate median, first quartile, seventh decile and sixtieth percentile.

[Link].a
[Link]
Numericals
Find the median wage of a daily worker from the following data. Also
calculate the third quartile, fourth decile and 79th percentile.
The total number of daily workers employed is 6000. Out of which 5%
earn less than Rs.350 per day, 1160 earn from Rs.351 to Rs.400 per day,
30% earn from Rs.401 to Rs. 450 per day, 1000 earn from Rs.451 to
Rs.500 per day, 20% earn from Rs.501 to Rs.550 per day and the rest
earn Rs.551 or more per day.

[Link].a
[Link]
Numericals
A firm houses 4500 employees on the basis of daily wages, 5% earn less
than Rs.250 per day, 870 earn from Rs.251 to Rs.300 per day, 30% earn
from Rs.301 to Rs. 350 per day, 750 earn from Rs.351 to Rs.400 per day,
20% earn from Rs.401 to Rs.450 per day and the rest earn Rs.451 or
more per day.
What is the median wage? Also find the first quartile, fourth decile and
18th percentile

[Link].a
[Link]
Numerical
s
In a factory employing 3000 persons, 5% earn less than Rs.150 per day,
580 earn from Rs.151 to Rs.200 per day, 30% earn from Rs.201 to Rs.
250 per day, 500 earn from Rs.251 to Rs.300 per day, 20% earn from
Rs.301 to Rs.350 per day and the rest earn Rs.351 or more per day.
What is the median wage? Also find the third quartile, eighth decile and
43rd percentile.

[Link].a
[Link]
Numerical
sThe Construction workers Federation of India reported the minimum
wage declared by various state governments in 2023. Find the upper
and lower quartile, 7th decile, 5th decile, 50th percentile.

[Link].a
[Link]
Mode
Mode is that value of the observation which occurs maximum
number of times.
Examples1. Calculate Mode from the following data
3,13,11,15,5,4,2,3,2Mode :In the given data, the
observation 2,3 occurs maximum number of times (2)∴Z=2,3

[Link].a
[Link]
Mode formula for
grouped data

[Link].a
[Link]
Illustrative Example:Compute for the mode using grouped data.

[Link].a
[Link]
Dispersion in
Statistics
In statistics, dispersion is the extent to which a distribution is
stretched or squeezed.
Dispersion is also called as variability, scatter, spread etc.,
Common examples of measures of statistical dispersion are the
variance, standard deviation, interquartile range etc.,

[Link].a
[Link]
Range

[Link].a
[Link]
Formulae

[Link].a
[Link]
Numericals
Find the range and co-efficient of range for the
following data.

[Link].a
[Link]
Series 1: 10, 10, 10, 10, 10

Series 2: 6, 8, 10, 12, 14

Series 3: 1, 2, 8, 19, 20

[Link].a
[Link]
Mean absolute deviation

The average value of the absolute deviations from each


of the observations from the mean value is said to be
mean absolute deviation.

[Link].a
[Link]
Formulae

[Link].a
[Link]
Numericals
Find the interquartile range, quartile deviation and co-
efficient of quartile deviation for the following data.

[Link].a
[Link]
Numericals
Find the interquartile range, quartile deviation and co-
efficient of quartile deviation for the following data.

[Link].a
[Link]
Formulae

[Link].a
[Link]
Numericals
Find the mean absolute deviation and co-efficient of
MAD for the following distribution of sales (in
thousands of rupees) in a co-operative store.

[Link].a
[Link]
Numericals
Find the mean absolute deviation and co-efficient of
MAD for the following data.

[Link].a
[Link]
Consider the following Data Sets and compute Range:
Data Set A: 25 37 35 37 38 39 39 44
Data Set B: 46 58 59 59 70 71 84 90

Range for Data Set A = Max. – Min. = 44 – 25 = 19


Range for Data Set B = Max. – Min. = 90 – 46 = 44
Conclusion: Range is a crude measure and it has not affected by
other numbers in the Data set, as it is based on only Minimum and
maximum Values in the data set.

[Link].a
[Link]
[Link].a
[Link]
Inter Quartile Range (IQR):
Inter Quartile Range (IQR): As we are aware that Mean
Absolute Deviation is a Measure of Dispersion in which deviations
are measured from the Average – Mean, or Median or Mode. But
practically, in some situations data may not be symmetrical for which
Mean is not a right measure.
In such cases, Median is a good measure of Central Tendency and a
Measure of Dispersion with reference to Quartiles places an
important role, as Median is the 2nd Quartile or Middle Quartile.
So, there is a need to check the variation among the Quartiles, which
otherwise can be called as Inter Quartile Range (IQR), a Deviation
with the Quartiles.

[Link].a
[Link]
Quartile Deviation (QD):

[Link].a
[Link]
Quartile Deviation (QD):

[Link].a
[Link]
Example: Consider the following Data Sets and
compute QD:
Data Set A: 16 12 14 19 11 13 15
Data Set B: 13 10 15 12 19 14 15

[Link].a
[Link]
Solution
Data Set A:
n = 7, Data in order
Data Set A: 1112 13 14 15
16 19

Q1 is located at the next Integer


after 1.75 = 2th term, Q1 = 12
(from the ordered data)

[Link].a
[Link]
Formula for Quartile Deviation (QD):For Grouped
Data:

[Link].a
[Link]
[Link].a
[Link]
[Link].a
[Link]
Solution

[Link].a
[Link]
Standard Deviation

[Link].a
[Link]
[Link].a
[Link]

You might also like