LESSON III.
FREQUENCY DISTRIBUTION
Regardless of whether an ordered array or a stem –and-leaf display is
selected for organizing data, as the number of observations obtained gets large,
the data need to be further condensed into summary of table in order to properly
present, analyze, and interpret the findings. This data can be arranged into
class groupings according to conveniently established divisions of the range of
the observations. This arrangement of data in tabular form is called frequency
distribution.
When the observations are grouped or condensed into a frequency
distribution, the process of data analysis and interpretation becomes much
more manageable and meaningful. The major data characteristics can be
approximated, which compensates for the fact that when the data are grouped,
the initial information pertaining to individual observations that was previously
available is lost through the grouping process.
In constructing frequency distribution, attention must be given to
selecting the appropriate number of class groupings for the table, obtaining a
suitable class interval, or width of each class grouping, and establishing the
boundaries of each class groupings to avoid overlapping.
At the end of the chapter, you should be able to:
1. identify the different steps in constructing frequency distribution; and
2. construct a cumulative frequency distribution.
FREQUENCY DISTRIBUTION
Table 3.3 gives the weekly earnings of 100 employees of a large
company. The first column lists the classes, which represent the variable weekly
earnings. For quantitative data, an interval that includes all the values that fall
within two numbers—the lower and upper limits—is called a class. Note that the
classes always represent a variable. As we can observe, the classes are
nonoverlapping; that is, each value on earnings belongs to one and only one
class. The second column in the table lists the number of employees who have
earnings within each class. For example, 9 employees of this company
earn Php5,001 to Php7,000 per week. The numbers listed in the second column
are called the frequencies, which give the number of values that belong to
different classes. The frequencies are denoted by f. Frequency distribution
for quantitative data lists all the classes and the number of values that belong
to each class. Data presented in the form of a frequency distribution are called
grouped data.
TABLE 3.3 Weekly Earnings of 100 Employees of a Company
Number of Frequency
Variable Weekly Earnings Employees column
f
5001 to 7,000 9
7,001 to 9,000 22 Frequency of
Third Class 9,001 to 11,000 39 the third
11,001 to 13,000 15 class
13,001 to 15,000 9
15,001 to 17,000 6
Lower limit of
the sixth class Upper limit of
the sixth class
The frequency of a class represents the number of values in the data set
that fall in that class. Table 3.3 contains of six classes. Each class has a lower
limit and an upper limit. The values 5001, 7001, 9001, 11001, 13001, and 15001
give the lower limits, and the values 7000, 9000, 11000, 13000, 15000, and
17000 are the upper limits of the six classes, respectively. The data presented
in Table 3.3 are an illustration of a frequency distribution table. Whereas the
data that list individual values are called ungrouped data, the data presented
in a frequency distribution table are called grouped data.
Constructing Frequency Distribution
When constructing a frequency distribution table, we need to follow the
following steps.
Example:
Assumed the data in table 3.4, test scores of 50 students in Statistics.
TABLE 3.4 Test scores of 50 Students in Statistics
63 67 73 80 76
88 71 60 56 52
65 85 63 51 76
62 46 90 40 55
88 60 63 78 86
63 42 79 77 60
72 70 83 54 43
76 87 62 48 72
83 78 62 47 52
85 75 55 90 40
Step 1: Determine the range (R) of the distribution
The range refers to the difference between the highest and the lowest
scores.
Range = Highest Score – Lowest Score
R=H–L
R = 90 – 40
R = 50
Step 2: Determine the class size (𝑖) by dividing the by the described number of
class intervals. The number of classes for a frequency distribution table
varies from 5 to 20, depending mainly on the number of observations in
the data set. It is preferable to have more classes as the size of a data
increases. The decision about the number of classes is arbitrarily made
by the data organizer.
Let us use 10.
Class size = Range ÷ 10
𝑖 = 50 ÷ 10
𝑖=5
If the obtained 𝑖 is not whole number, round it off to the nearest whole
number.
Step 3: When the class size is 5, all the lower class limit must be multiple of
1. The lower class interval should include the lowest score while highest class
interval must contain the highest score. Any convenient number that is equal
to or less than the smallest value in the data set can be used as the lower limit
of the first class.
Step 4: Tally the frequencies for each interval and sum them.
Step 5: Find the class marks or midpoint of the class intervals. It is the point
halfway between the boundaries of each class and is representative of
the data within that class.
For example, the class mark or midpoint of 50 – 54.
50 + 54 = 104 ÷ 2 = 52.
TABLE 3.5 Completed Frequency Distribution
Class Interval Tally Marks Frequency (f) Class Marks (x)
40 – 44 IIII 4 42
45 – 49 III 3 47
50 – 54 IIII 4 52
55 – 59 III 3 57
60 – 64 IIII - IIII 10 62
65 – 69 II 2 67
70 – 74 IIII 5 72
75 – 79 IIII - III 8 77
80 – 84 III 3 82
85 – 89 IIII - I 6 87
90 – 94 II 2 92
N = 50
CUMULATIVE FREQUENCY DISTRIBUTION
A table of cumulative frequency distribution is another useful technique
in tabulating data. It provides information about sets of data that cannot be
obtained from the frequency distribution itself. Cumulative frequency
distribution is the sum of the class and all classes below it in a frequency
distribution. All that means is you are adding up a value and all the values that
came before it.
Table 3.6 demonstrate how cumulative frequency distribution is
constructed by using the example in table 3.4.
TABLE 3.6 Cumulative Frequency Distribution
Cumulative
Class Interval Frequency (f) Class Marks (x)
Frequency (<cf)
40 – 44 4 42 4
45 – 49 3 47 7
50 – 54 4 52 11
55 – 59 3 57 14
60 – 64 10 62 24
65 – 69 2 67 26
70 – 74 5 72 31
75 – 79 8 77 39
80 – 84 3 82 42
85 – 89 6 87 48
90 – 94 2 92 50
N = 50
Descriptive statistics is the term given to the analysis of data that helps
describe, show, or summarize data in a meaningful way. Descriptive statistics
do not, however, allow us to make conclusions beyond the data we have
analyzed or reach conclusions regarding any hypotheses we might have made.
They are simply a way to describe the data. It is very important because if we
simply presented our raw data it would be hard to visualize what the data was
showing, especially if there was a lot of it.
In Chapter 3 we discussed how to summarize data using different
methods and to display data using graphs. Graphs are one important
component of statistics; however, it is also important to numerically describe
the main characteristics of a data set. The numerical summary measures, such
as the ones that identify the center and spread of a distribution, identify many
important features of a distribution. For example, the techniques learned in
Chapter 2 can help us graph data on family incomes. However, if we want to
know the income of a “typical” family (given by the center of the distribution),
the spread of the distribution of incomes, or the relative position of a family with
a particular income, the numerical summary measures can provide more
detailed information. The measures that we discuss in this chapter include
measures of (1) central tendency, (2) dispersion (or spread), and (3) position.
General Objectives:
At the end of the chapter, you should be able to:
1. recognize the different types of descriptive statistics using grouped
and ungrouped data;
2. describe the measures of central tendencies, measures of position,
and measures of dispersion;
59
3. identify the steps in computing descriptive statistics using MS Excel;
and
4. appreciate the ease of computing problems with the aid of computer.
LESSON I. MEASURES OF CENTRAL TENDENCY
A measure of central tendency is a summary statistic that represents the
center point or typical value of a dataset. These measures indicate where most
values in a distribution fall and are also referred to as the central location of a
distribution. You can think of it as the tendency of data to cluster around a
middle value. In statistics, the three most common measures of central
tendency are the mean, median, and mode. Each of these measures calculates
the location of the central point using a different method.
Choosing the best measure of central tendency depends on the type of
data you have. In this lesson, we explore these measures of central tendency,
show you how to calculate them, and how to determine which one is best for
your data.
At the end of the lesson, you should be able to:
1. recognize the mean, median, and the mode using grouped and
ungrouped data; and
2. describe the measures of central tendencies.
MEASURES OF CENTRAL TENDENCY USING UNGROUPED DATA
We often represent a data set by numerical summary measures, usually
called the typical values. A measure of central tendency gives the center of a
frequency distribution curve. This section discusses three different measures of
central tendency: the mean, the median, and the mode. We will learn how to
calculate each of these measures for ungrouped data. Recall from Chapter 3
that the data that give information on each member of the population or sample
individually are called ungrouped data, whereas grouped data are presented in
the form of a frequency distribution table.
60
The Mean
The mean, also called the arithmetic mean, is the most frequently used
measure of central tendency. This module will use the words mean and average
synonymously. For ungrouped data, the mean is obtained by dividing the sum
of all values by the number of values in the data set:
𝑆𝑢𝑚 𝑜𝑓 𝑎𝑙𝑙 𝑣𝑎𝑙𝑢𝑒𝑠
𝑀𝑒𝑎𝑛 =
𝑁𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑣𝑎𝑙𝑢𝑒𝑠
You can denote the observations by X1, X2,… Xn as the first, second…
and the last observation, where the mean calculated for sample data is denoted
by 𝑥 (read as “x bar”), and the mean calculated for population data is denoted
by µ (Greek letter mu).
Calculating Mean for Ungrouped Data The mean for ungrouped data is
obtained by dividing the sum of all values by the number of values in the data
set. Thus,
Mean for population data: ∑𝑥
µ=
𝑁
Mean for sample data:
Where the ∑𝑥 is the sum of all values, N is the population size, n is the sample
size, µ is the population mean, and is the sample mean
𝑥
Example: (Population Mean)
The number of students in 5 sections of the grade 10 of ABC National
High School are: Amethyst = 50, Diamond = 45, Jade = 46, Pearl = 45, and
Sapphire = 40. Treating the data as a population, find the population mean of
the 5 sections of the grade 10 of ABC National High School.
Solution:
∑𝑥 50 + 45 + 46 + 45 + 40 226
µ= = = = 𝟒𝟓. 𝟐𝟎
𝑁 5 5
61
Example: (Sample Mean)
A sample of 12 high school students was asked on how many hours had
they spent on watching the television last week. The responses are listed
below.
5 12 15 20 15 24 0 24 30 10 16 18
Solution
∑𝑥 5 + 12 + 15 + 20 + 15 + 24 + 0 + 24 + 30 + 10 + 16 + 18 189
𝑥= = = = 𝟏𝟓. 𝟕𝟓
𝑛 12 12
The Median
The median is the value of the middle term in a data set that has been
ranked in decreasing or decreasing order. As is obvious from the definition of
the median, it divides a ranked data set into two equal parts. The calculation of
the median consists of the following two steps:
1. Rank the data set in increasing or decreasing order.
2. Find the middle term. The value of this term is the median.
Note that if the number of observations in a data set is odd, then the
median is given by the value of the middle term in the ranked data. However, if
the number of observations is even, then the median is given by the average of
the values of the two middle terms. To illustrates this, consider the following
values:
Examples:
1. When the number of observations is odd, say n= 9.
Find the median of: 13 43 23 20 51 64 49 80 55
Solution
Arrange the data in ascending or descending order.
13 20 23 43 49 51 55 64 80
The median is the middle score of the data; therefore, 49 is the median.
2. When the number of observations is even, say n=10.
Find the median of: 34 56 89 42 26 14 28 98 56 78
62
Solution
Arrange the data in ascending or descending order.
14 26 28 34 42 56 56 78 89 98
The median is the average of the two middle scores; therefore, getting
the average of the 5th and 6th values, 49 is the median.
The Mode
The mode is a French word that means fashion – an item that is most
popular or common. In statistics, the mode is the value that occurs with the
highest frequency in a data set. If there is no common score, the said data has
no mode. A distribution with only one mode is said to be unimodal while a
distribution with two or more modes is described as multi-modal.
Examples:
1. Find the mode of the following scores:
0 1 3 5 3 3 8 9 3 4
Solution:
Simply get the value of the most frequent appearing value. The mode of
the given data is 3.
2. Find the mode of:
0 0 0 0 2 4 3 4 4 4 9 7 8 5
Solution:
There are two values that appeared four times; therefore, the modes of
the given data are 0 and 4.
MEASURES OF CENTRAL TENDENCY USING GROUPED DATA
Data which are arranged in a frequency distribution are called Grouped
Data. When the number or items is too large, it is best to compute for the
measures of central Tendency and variability using the frequency distribution.
The Mean
There are two methods we can use to compute for the mean from
grouped data: the long method and the coded deviation method.
63
Calculating Mean for Grouped Data
∑ 𝑓𝑥
𝑥=
𝑁
Where f is the frequency, x is the class mark or the midpoint, and N is the
total observation or frequency.
Example:
Using the table 3.5 in chapter 3.
Class Interval f x fx
40 – 44 4 42 168
45 – 49 3 47 141
50 – 54 4 52 208
55 – 59 3 57 171
60 – 64 10 62 620
65 – 69 2 67 134
70 – 74 5 72 360
75 – 79 8 77 616
80 – 84 3 82 246
85 – 89 6 87 522
90 – 94 2 92 184
N = 50 ∑fx=3,370
Solution:
∑𝑓𝑥 3370
𝑥= = = 𝟔𝟕. 𝟓𝟎
𝑁 50
One will notice that the accuracy of the result from grouped data has
been affected by the loss of the identity of the original numbers. However, we
can still say that is good approximation of the true mean.
64
The Median
It is the value of the middle in an ordered arrangement of data. In an
ordered distribution, half of the terms are located above the median and half
are below the median. The median is a positional measure; hence, the values
of the individual item in a distribution do not affect the median. It is not
influenced by extreme value. The highest value in a distribution does not enter
computation of the median. This is an advantage when there are some terms
of extreme value in the distribution that would pull up the average of the whole
distribution, thus presenting a biased average or a value not “typical” of the
terms in the distribution.
To compute the median from grouped data we also must determine the
“less than” cumulative frequency. The median is the sum of the lower limit of
the median class and a fractional part of the class interval size.
Calculating Median for Grouped Data
𝑁
[ 2 − 𝐹]𝑖 Where:
𝑚𝑑 = 𝐿𝑚 + md = Median
𝑓 𝐿𝑚 = Lower Limit
N = total frequency
F = Less than cumulative frequency
f = frequency of the median
i = class interval
Example:
Using the table 3.5 in chapter 3.
Solution:
𝑁 50 𝑁
[ 2 −𝐹]𝑖 [25−24]5
= = 25 𝑚𝑑 = 𝐿𝑚 + = 64.5 + = 67
2 2 𝑓 2
65
The Mode
The mode in a frequency distribution is within the class interval with the
highest frequency. The class interval with the frequency is known as the modal
class. A crude mode may be determined by taking the class mark with the
highest frequency. However, this rough approximation may be improved by
considering the frequencies adjoining the modal class.
Calculating Mode for Grouped Data
⬚
∆1
𝑚𝑜𝑑𝑒 = 𝐿𝑚 + ൨𝑖
∆1 + ∆2
Where:
Lm is the lower limit of the modal class (this is the class interval with
the highest frequency)
∆1 is the difference between the highest frequency and the
frequency above it.
∆2 is the difference between the highest frequency and the
frequency below it.
𝑖 is the class interval
Example:
Using the table 3.5 in chapter 3.
Class Interval f
40 – 44 4
45 – 49 3
50 – 54 4 ∆1
55 – 59 3
60 – 64 10
65 – 69 2
70 – 74 5 Modal Class
75 – 79 8
∆2
80 – 84 3
85 – 89 6
90 – 94 2
66
Solution:
∆1
𝑚𝑜𝑑𝑒 = 𝐿𝑚𝑜 + ൨𝑖
∆1 + ∆2
(10 − 3)
= 59.5 + ൨5
(10 − 3) + (10 − 2)
= 𝟔𝟏. 𝟖𝟑
EXERCISES
CONCEPTS AND PROCEDURES
1.1. Explain how the value of the median is determined for a data set
contains an odd number of observations and for a data set that
contains an even number of observations.
1.2. Is it possible for a quantitative data set to have no mean, no
median, or no mode? Give an example of a data set for which this
summary does not exist.
APPLICATIONS
1.3. The following data give the numbers of car thefts that occurred in a
Region during the past 12 days.
6 3 7 11 4 3 8 7 2 6 9 15
Find the mean, median, and mode.
1.4. A brochure from the department of public safety recommends that
motorists should carry 12 items (flashlights, blankets, and so forth)
in their vehicles for emergency use while driving in winter. The
following data give the number of items out of these 12 that were
carried in their vehicles by 15 randomly selected motorists.
5 3 7 8 0 1 0 5 1 21 07 6 7
1 19
Find the mean, median, and mode for these data. Are the values of
these summary measures population parameters or sample
statistics? Explain.
67
1.5. Consider the following scores as the results of the College Entrance
Examination of 30 students who have taken the admission test.
85 81 94 100 77 98 96 98 67 88
74 79 80 96 83 93 87 77 66 83
66 72 86 91 70 83 94 90 87 61
Prepare a frequency distribution table with a class interval of 5. Find
for the mean, median, mode.
1.6. Given the frequency distribution of the daily commuting times (in
minutes) from home to work for all 25 employees of a company.
Daily Commuting Time
Number of Employees
(minutes)
0 to less than 10 4
10 to less than 20 9
20 to less than 30 6
30 to less than 40 4
40 to less than 50 2
Calculate the mean, median and mode of the daily times
1.7. Given the frequency distribution of the number of orders received
each day during the past 50 days at the office of a mail-order
company.
Number of Orders Number of Days
0 to less than 10 4
10 to less than 20 12
20 to less than 30 20
30 to less than 40 14
Calculate the mean, median and mode.
68
LESSON II. MEASURES OF POSITION
Measures of position give us a way to see where a certain data point or
value falls in a sample or distribution. A measure can tell us whether a value is
about the average, or whether it is unusually high or low. Measures of position
are used for quantitative data that falls on some numerical scale. Sometimes,
measures can be applied to ordinal variables— those variables that have an
order, like first, second…fiftieth.
The idea of the median can be extended to the discussion of quantiles.
Quantiles are values which divide the distribution into a given number of equal
parts. The median divides the distribution into two equal parts. Other types of
quantiles which we will discuss in this lesson are percentiles, quartiles, and
deciles.
At the end of the lesson, you should be able to:
1. recognize the quartiles, percentiles, and the deciles using grouped
and ungrouped data; and
2. describe the measures of position.
THE QUARTILES
The quartiles divide the distribution into four equal parts, each
comprising 25% of the observation. The quartiles are Q1 (first quartile), Q2
(second quartile or the median), Q3 (third quartile), and Q4 (fourth quartile).
The following are the formula of quartiles in ungrouped and grouped
data. Please take note that we get from these formula is not the exact values
but only the position of the given data.
Calculating the Quartiles for Ungrouped Data
𝑘(𝑛 + 1) Example:
(𝑛 + 1) 3(𝑛 + 1)
Position of Qk= Q1 = Q3=
4 4 4
69
Calculating the Quartiles for Grouped Data
𝑘𝑛
− < 𝑐𝑓
𝑄𝑘 = 𝐿𝑚 + 4 𝑖
𝑓
Where:
<cf is the cumulative frequency less than of the class immediately
preceding the desired quartile
Lm is the exact lower limit of the class where desired quartile is
located
i is the class interval
f is the frequency of the desired quartile
kn
is the conventional formula to guide in locating the appropriate
4
desired quartile class
THE DECILES
The deciles divide the distribution into ten equal parts, each comprising
10% of the observation. The median is the 5th decile.
The following are the formula of quartiles in ungrouped and grouped
data. Please take note that we get from these formula is not the exact values
but only the position of the given data.
Calculating the Deciles
For Ungrouped Data
𝑘(𝑛 + 1) Example:
Position of Dk = 5(𝑛 + 1) 8(𝑛 + 1)
D5 = D8 =
10 10 10
For Grouped Data
𝑘𝑛
−< 𝑐𝑓
𝐷𝑘 = 𝐿𝑚 + 10 𝑖
𝑓
Where:
<cf is the cumulative frequency less than of the class
immediately preceding the desired decile
Lm is the exact lower limit of the class where desired decile is
located
i is the class interval
f is the frequency of the desired decile
𝑘𝑛
is the conventional formula to guide in locating the
10
appropriate desired decile class
70
THE PERCENTILES
The percentiles divide the distribution into 100 equal parts, each
comprising 1% of the observation. The median describes the 50th percentile.
The following are the formula of quartiles in ungrouped and grouped
data.
Calculating the Percentiles
For Ungrouped Data
𝑘(𝑛 + 1) Example:
Position of Pk = 30(𝑛 + 1) 70(𝑛 + 1)
100 P30 = P70 =
100 100
For Grouped Data
𝑘𝑛−< 𝑐𝑓
𝑃𝑘 = 𝐿𝑚 + ൨𝑖
𝑓
Where:
Lm is the the exact lower limit of the class interval containing
desired percentile
n is the number of cases or the total frequency
k is the desired percentile
i is the class interval
f is the frequency of the desired percentile
< cf is the cumulative frequency immediately below
class interval containing desired percentile
Example for Ungrouped Data
Find the P20, Q2, and D8 of the following data.
15 13 9 4 8 25 28 30 18
71
Solution:
Step 1: Arrange the given data according to magnitude or size
(descending or ascending order).
30
28
P20
25
18
15
Q2
13
9
8
4
D8
Step 2: Compute the position of the given percentile, quartile, or decile
in the distribution using its corresponding formula.
P20 = 20(𝑛 + 1) = 20(9 + 1) = 200 = 𝟐
100 100 100
Q2 = 2(𝑛 + 1) = 2(9 + 1) = 20 = 𝟓
4 4 4
8(𝑛 + 1) 8(9 + 1) 80
D8 = = = =𝟖
10 10 10
Step 2: Starting from the lowest score, locate the score corresponding
to the obtained position in the distribution.
P20 = 8 Q2 = 15 D8 = 28
Interpolate to get the score if the obtained position in step 2 is not exact.
Using the same example, follow the steps below:
Step 1: Solve for the position of Q3.
3(𝑛 + 1) 3(9 + 1) 30
Q3 = = = = 𝟕. 𝟓
4 4 4
72
Step 2: Since 7.4 is between the 7th and 8th score from the lowest, we
must interpolate to find the answer. Take the 7th score from the
lowest which is 25 and the 8th score which is 28.
Step 3: Get the difference of the two scores.
8th score – 7th score
28 – 25 = 3
Step 4: Then, multiply the difference by the decimal part obtained in step
1.
3 x .5 = 1.5
Step 5: Finally, add the product obtained in step 4 to the lower score in
the identified location in step 2.
Thus, Q3 = 25 + 1.5 = 26.5
Example Grouped Data
Using the same frequency distribution in chapter 3, table 3.5. Find the
P40, Q3, and D2.
Class
F x <cf
Interval
40 – 44 4 42 4
45 – 49 3 47 7
50 – 54 4 52 11 Desired 2nd
Decile
55 – 59 3 57 14
60 – 64 10 62 24 Desired 20th
Percentile
65 – 69 2 67 26
70 – 74 5 72 31
75 – 79 8 77 39 Desired 3rd
Quartile
80 – 84 3 82 42
85 – 89 6 87 48
90 – 94 2 92 50
N = 50
73
Solution:
𝑘𝑛 −< 𝑐𝑓 3𝑁 2𝑁
= 𝐿𝑚 + ൨𝑖 −< 𝑐𝑓 −< 𝑐𝑓
P40 𝑓 Q3= 𝐿𝑚 + 4 𝑖 D2 = 𝐿𝑚 + 10 𝑖
𝑓 𝑓
np = 50 x 40% = 20 3𝑁 3(50) 2𝑁 2(50)
= = 𝟑𝟕. 𝟓 = = 𝟏𝟎
4 4 10 10
20 − 14 10 − 7
P40 = 59.5 + ൨5 37.5 − 31 D2 = 49.5 + ൨5
10 Q3 = 74.5 + ൨5 4
8
6 3
P40 = 59.5 + ൨ 5 6.5 D2 = 49.5 + ൨ 5
10 Q3 = 74.5 + ൨5 4
8
P40 = 59.5 + 3 Q3 = 74.5 + 4.06 D2= 49.5 + 3.75
3
P40 = 𝟔𝟐. 𝟓𝟎 Q3 = 𝟕𝟖. 𝟓𝟔 D2 = 49.5 + ൨ 5
4
LESSON III. MEASURES OF DISPERSION
The range, interquartile range, standard deviation, and the variance
provide us the distances of scores from the measures of central tendency. It
can also be used to establish the actual similarities or the differences of the
distribution. In general, these are employed to further characterize the
distribution of test scores.
These measures simply approximate the central value of the distribution.
However, such descriptions are not enough to be able to adequately describe
the characteristics of a set of data. Hence, there is a need to consider how the
values are scattered on either side of values in the distribution.
At the end of the lesson, you should be able to:
1. recognize the range, interquartile range, standard deviation, and
variance; and
2. describe the measures of dispersion.
74
THE RANGE
It is the simplest and easiest measures of variability. It simply measures
how far the highest score is to the highest score. It does not tell anything about
the scores between these two extreme scores. Thus, it is considered as the
least satisfactory measure of variability.
Calculating the Range
For Ungrouped Data
Where:
𝑅 =H−L R = Range
H = Highest Score
L = Lowest Score
For Grouped Data
Where:
𝑅 = HUL – LLM R = Range
HUL = Upper Limit of the Highest class Interval
LLM = Lower Limit of the Lowest class Interval
75
THE INTERQUARTILE RANGE
The interquartile range is a measure of where the “middle fifty” is in a
data set. Where a range is a measure of where the beginning and end are in a
set, an interquartile range is a measure of where the bulk of the values lie. This
value is obtained by getting the difference of the 3rd and 1st quartiles in a set of
data.
Calculating the Interquartile Range
Where:
IQR= Q3 – Q1 IQR = Interquartile Range
Q3 = 3rd Quartile
Q1 = 1st Quartile
THE STANDARD DEVIATION AND THE VARIANCE
It is the measure of variability that involves all scores in the distribution
rather than through extreme scores. It may be referred to as the root-mean
square of the deviation from the mean. It is considered the most important
measure of variability. In general, a lower value of the standard deviation for a
data set indicates that the value of the data sets is spread over a relatively
smaller range around the mean.
Calculating the Standard Deviation and the Variance
For Standard Deviation
Ungrouped Data Grouped Data
∑(𝑥 − 𝑥 )2 ∑𝑓(𝑥 − 𝑥 )2
𝑠=ඨ 𝑠=ඨ
𝑛−1 𝑛−1
For Variance
Ungrouped Data Grouped Data
∑(𝑥 −𝑥 )2 ∑𝑓(𝑥 −𝑥 )2
2
𝑠 = 𝑠2 =
𝑛−1 𝑛−1
76
Example:
For Ungrouped Data
A sample of 5 high school students was asked on how many hours had
they spent on watching the television last week. Using the samples below, find
for the (a) range, (b) interquartile range, (c) standard deviation and (d) variance.
18 19 12 11 10
x ẋ (x - ẋ) (x - ẋ)2
18 14 4 16
19 14 5 25
12 14 -2 4
11 14 -3 9
10 14 -4 16
∑X=70 ∑(x - ẋ)2 =70
Solution:
Solve for the mean (ẋ)
∑𝑥 70
𝑥= = = 𝟏𝟒
𝑛 5
a. Range
R=H–L
= 19 – 10
R=9
b. Interquartile Range
(𝑛 + 1)
IQR = Q3 – Q1 Q1 =
4
Q3 = 3(𝑛 + 1)
4
= 18.50 − 10.50
(5 + 1) 3(5 + 1)
=𝟖 = =
4 4
Q1 = 1.5 Q3 = 4.5
11 – 10 = 1 19 – 18 = 1
1 x .5 = 0.5 1 x .5 = 0.5
Q1 = 10 + 0.5 Q3 = 18 + 0.5
= 𝟏𝟎. 𝟓𝟎 = 𝟏𝟖. 𝟓𝟎
77
c. Standard Deviation
∑(𝑥 − 𝑥 )2 70
𝑠=ඨ = ඨ = 𝟒. 𝟏𝟖
𝑛−1 4
d. Variance
∑(𝑥 − 𝑥 )2 70
𝑠2 = = = 𝟏𝟕. 𝟓𝟎
𝑛−1 4
Example:
For Grouped Data
Using the same frequency distribution in chapter 3, table 3.5. Find the
find for the (a) range, (b) interquartile range, (c) standard deviation and (d)
variance.
Class
f x fx <cf x-ẋ (x - ẋ)2 f(x - ẋ)2
Interval
40 – 44 4 42 168 4 -25.4 645.16 2580.64
45 – 49 3 47 141 7 -20.4 416.16 1248.48
50 – 54 4 52 208 11 -15.4 237.16 948.64
55 – 59 3 57 171 14 -10.4 108.16 324.48
60 – 64 10 62 620 24 -5.4 29.16 291.6
65 – 69 2 67 134 26 -0.4 0.16 0.32
70 – 74 5 72 360 31 4.6 21.16 105.8
75 – 79 8 77 616 39 9.6 92.16 737.28
80 – 84 3 82 246 42 14.6 213.16 639.48
85 – 89 6 87 522 48 19.6 384.16 2304.96
90 – 94 2 92 184 50 24.6 605.16 1210.32
N = 50 ∑fx=3,370 ∑f(x - ẋ)2 = 10,392
Solution:
a. Range b. Interquartile range
R = HUL – LLM IQR = Q3 – Q1
= 94.5 – 39.5 = 78.56 – 57
R = 55 IQR = 21.56
78
3𝑁 𝑁
−< 𝑐𝑓 −< 𝑐𝑓
Q3 = 𝐿𝑚 + 4 𝑖 Q1 = 𝐿𝑚 + 4 𝑖
𝑓 𝑓
3𝑁 3(50) 𝑁 50
= = 𝟑𝟕. 𝟓 = = 12.5
4 4 4 4
Q3 = 𝟕𝟖. 𝟓𝟔 Q3 = 𝟓𝟕
b. Standard deviation
∑𝑓(𝑥 − 𝑥 )2
𝑠=ඨ
𝑛−1 𝑥 =
∑𝑓𝑥 3370
𝑥= = = 𝟔𝟕. 𝟓𝟎
10,392 𝑁 50
𝑠=ඨ
50 − 1
𝒔 = 𝟏𝟒. 𝟓𝟔
c. Variance
∑𝑓(𝑥 −𝑥 )2
𝑠2 =
𝑛−1
79
10,392
𝑠2 =
49
𝒔𝟐 = 𝟐𝟏𝟐. 𝟎𝟖
80
LESSON IV. COMPUTER APPLICATION USING MS EXCEL
Statistics is a mathematical science involving the collection,
interpretation, measurement, enumerations or estimation analysis, and
presentation of natural or social phenomena, through application of various
tools and technique the raw data becomes meaningful and generates the
information’s for decision making purpose. It is the systematic arrangement of
data and information exhibits their inner relation between the things. Now
statistics holds a central position in almost every field of research like Industry,
Commerce, Trade, Physics, Chemistry, Economics, Mathematics, Biology,
Botany, Psychology, Astronomy, management of decision making etc. in this
lesson, we will discuss on the statistical tools with the use of computer
application specifically the Microsoft Excel which can help the calculation and
interpretation of data in a very efficient and effective manner.
At the end of the lesson, you should be able to:
1. identify the steps in computing descriptive statistics using MS Excel;
and
2. appreciate the ease of computing problems with the aid of computer.
STEPS ON HOW TO COMPUTE THE DESCRIPTIVE STATISTICS USING
MS EXCEL ANALYSIS TOOLPAK:
Step 1: Type your data into Excel, in a single column. For example, if
you have ten items in your data set, type them into cells A1
through A10.
Step 2: Click the “Data” tab and then click “Data Analysis” in the
Analysis group.
Step 3: Highlight “Descriptive Statistics” in the pop-up Data Analysis
window and click “OK”.
81
Step 4: Type an input range into the “Input Range” text box. For this
example, type “A1:A10” into the box.
Step 5: Check the “Labels in first row” check box if you have titled the
column in row 1, otherwise leave the box unchecked.
Step 6: Type a cell location into the “Output Range” box. For example,
type “C1.” Make sure that two adjacent columns do not have
data in them.
Step 7: Click the “Summary Statistics” check box and then click “OK”
to display Excel descriptive statistics. A list of descriptive
statistics will be returned in the column you selected as the
Output Range.
Example:
The following examples will illustrate how MS Excel will help you resolve
exercises in a very convenient way.
Consider the scores obtained by 20 students in their college entrance
test.
Student Score Student Score
A 81 K 90
B 90 L 90
C 78 M 90
D 65 N 88
E 85 O 76
F 79 P 94
G 80 Q 97
H 81 R 73
I 74 S 72
J 75 T 69
82
Solution:
Step 1. Input the Given Data
to the Excel
Step 2 & 3
Click Descriptive Statistic Click Data Analysis
Click Data
83
Step 7 & 8
Click “OK”
Type a cell into the
input range
“B2:B21”
Check the column
Check the Type a cell into the Check the or rows
“Labels in first row” output range “Summary Statistics”
“D4”
84
TABLE 4.1 Results of the Statistical Function
Statistical Function Results
Mean 81.35
Standard Error 1.973875429
Median 80.5
Mode 90
Standard Deviation 8.827439278
Sample Variance 77.92368421
Kurtosis -0.88771227
Skewness 0.02342738
Range 32
Minimum 65
Maximum 97
Sum 1627
Count 20
You can use other statistical functions to solve other problems like z-test,
t-test, correlation, regression, and many more.
Page 85 of 33