Data Classification and Frequency Distribution
Data Classification and Frequency Distribution
ORGANIZATION/CLASSIFICATION OF DATA
Measurements or counting gives rise to raw data. Raw data are collected data that have not been
organized numerically. Raw data are difficult to comprehend because it lacks organization,
summarization, which renders it meaningless. Thus, the raw data has to be put in some order
through classification and tabulation so as to reduce its volume and heterogeneity.
• Comprehensiveness: Classification should cover all the items of the data. In other words, it
should be so comprehensive that it classifies all items in some group or class.
• Clarity: There should be no confusion of the placement of any data item in a group or class.
That is, classification should be absolutely clear.
• Homogeneity: The items within a specific group or class should be similar to each other.
• Elastic: As the purpose of classification changes, one should be able to change the basis of
classification.
Frequency: Is the number of times a certain value or class of values of the data occurs. The
sum of the frequencies equal to the total observations(n) or sample size n.
Frequency Distribution
A tabular arrangement of data by their classes together with the corresponding class
frequencies is called frequency distribution or frequency table.
The techniques use to organize data depend on the type of variable (Quantitative(numerical) or
Qualitative(categorical)) associated with such data.
Two types of frequency distributions that are most often used are the:
Page 1 of 12
Categorical Frequency Distribution:
A categorical frequency distribution is a table used to organize data that can be placed in specific
categories, such as nominal- or ordinal-level data. The categorical frequency distribution is
done by tallying responses by categories and place the results in tables. This can be
done to construct a summary table to organize the data for a single categorical variable
or construct a contingency table to organize the data from two or more categorical
variables.
a. Summary Table:
A summary table is usually constructed for a single categorical variable. The table present the
tallied responses as frequencies or percentages for each category. The table helps you to see the
differences among the categories by displaying the frequency(amount) or percentage of items in a
set of categories in a separate column.
For example, the data below represents the blood groups of 40 students in a Biostatistics class.
Construct a frequency distribution for the data.
Page 2 of 12
B. Quantitative variable
Page 3 of 12
ii. Grouped Frequency Distribution: The data are organized into groups or
intervals with their corresponding frequencies.
Table 2.0: Height of 100 students in STA 131 class
c. Class boundaries:
Class boundaries are defined to eliminate any gaps between the classes (it has one more decimal
place than the data.). Class boundaries are those limits which are determined mathematically to
make an interval of a continuous variable continuous in both directions, and no gap exists between
classes. It’s the actual or real limits of a class interval.
a. The lower extreme point is called lower class boundary
b. The Upper extreme point is called Upper class boundary
Class boundaries are obtained as follows:
Page 4 of 12
1
Lower Class boundary= lower class limit− 2 𝛼
1
Upper Class boundary= upper class limit+ 𝛼
2
Where 𝛼 is the difference between the upper-class limit of any class interval and lower-class
limit of the next class interval.
d. Class mark (xc) or Mid-point of an interval:
1. The class mark is the midpoint of the class interval and is obtained by adding the lower-
class limit and upper-class limit and dividing by 2.
2. The class mark is also called the class midpoint.
3. It is used as representative value of the class interval for the calculation of mean, standard
deviation and other measures.
4. Class mark is the value representing the class interval. It is calculated as:
Table: Class limit, Class boundary, Class mark, Width, Relative frequency and Percentage
Relative frequency
Class Class Class limits Class Class Class Relative %
Interval Frequency Boundaries Mark Width Freq. Relative
Lower Upper Freq.
Class Frequency: The number of observations falling within a class is called its class frequency.
Total Frequency: The sum of all the frequencies is called total frequency.
Relative frequency: It is ratio of the frequency of the class to the total frequency. It’s used to
compare two or more frequency distributions or two or more items in the same frequency
distribution. The relative frequency is not expressed as percentage and its defined as:
Page 5 of 12
𝐹𝑟𝑒𝑞𝑢𝑒𝑛𝑐𝑦 𝑜𝑓 𝑡ℎ𝑒 𝐶𝑙𝑎𝑠𝑠
𝐑𝐞𝐥𝐚𝐭𝐢𝐯𝐞 𝐅𝐫𝐞𝐪𝐮𝐞𝐧𝐜𝐲 =
𝑇𝑜𝑡𝑎𝑙 𝑓𝑟𝑒𝑞𝑢𝑒𝑛𝑐𝑦
Percentage Relative Frequency: This is the ratio of the frequency of a class to the total
frequency expressed as percentage. Its defined as:
𝐹𝑟𝑒𝑞𝑢𝑒𝑛𝑐𝑦 𝑜𝑓 𝑡ℎ𝑒 𝐶𝑙𝑎𝑠𝑠
𝐏𝐞𝐫𝐜𝐞𝐧𝐭𝐚𝐠𝐞 (%) 𝐑𝐞𝐥𝐚𝐭𝐢𝐯𝐞 𝐅𝐫𝐞𝐪𝐮𝐞𝐧𝐜𝐲 = × 100
𝑇𝑜𝑡𝑎𝑙 𝑓𝑟𝑒𝑞𝑢𝑒𝑛𝑐𝑦
Guidelines for Number of Classes
i. There should be between 5 and 20 classes.
ii. The class width should be an odd number. This will guarantee that the class midpoints are
integers instead of decimal.
iii. The classes must be mutually exclusive. This means that no data value can fall into two
different classes.
iv. The classes must be all inclusive or exhaustive. This means that all data values must be
included.
v. The classes must be continuous. There are no gaps in a frequency distribution.
vi. The classes must be equal in width. The exception here is the first or last class. It is possible
to have an "below or " and above" class. This is often used with ages.
Creating a Grouped Frequency Distribution
l. Find the largest and smallest values
2. Compute the Range(R): Maximum - Minimum
3. Select the number of classes(K) desired. This is usually between 5 and 20.
4. Find the class width by dividing the range by the number of classes and rounding up.
Size(Width) = R/ K
a. You must round up, not off. Normally 3.2 would round to be 3, but in rounding up, it
becomes 4. If the range divided by the number of classes gives an integer value (no
remainder), then you can either add one to the number of classes or add one to the class
width.
Sometimes you are instructed to use a certain number of classes.
b. Pick a suitable value less than or equal to minimum value.
c. . Your starting value is the lower limit of the first class. Continue to add the class width to
this lower limit to get the rest of the lower limits.
d. To find the upper limit of the first class, subtract one from the lower limit of the second
class. Then continue to add the class width to this upper limit to find the rest of the upper
limits.
Page 6 of 12
5. Tally the data.
6. Find the frequencies.
7. Find the boundaries by subtracting 0.5 units from the lower limits and adding 0.5
units from the upper limits (if the data are recorded without decimal).
ii. More than cumulative frequency: The number of observations “greater than” a value is
called more than cumulative frequency. It’s the number of observations that are greater
than the lower limit of a class interval.
1. It’s used to find out the number of observations less than or more than any given value.
2. It’s used to find out the number of observations falling between any two specified values of
the variable.
100
Page 7 of 12
CONSTRUCTION OF FREQUENCY DISTRIBUTION
(1) Find the range of the data: The range is the difference between the largest and the
smallest values.
(2) Decide the approximate number of classes in which the data are to be grouped. Where
the number of classes to used is not given, the number of classes can be estimated using
H.A. Sturge’s formula given as:
K = 1 + 3.322log N
Where K= Number of classes and N = total number of observations.
(3) Determine the approximate class size: The size (width)of class interval is obtained by
dividing the range of data by the number of classes and is denoted by W (class interval
width(size))
Size/ width= Range/ K
In the case of fractional results, the next higher whole number is taken as the size of the
class interval.
(4) Decide the starting point: The lower-class limits or class boundary should cover the
smallest value in the raw data. Usually class intervals of multiple of 5 are commonly used.
(5) Determine the remaining class limits (boundary): When the lowest class boundary has
been decided, you can compute the upper-class boundary by adding the class interval size
to the lower-class boundary. The remaining lower- and upper-class limits may be
determined by adding the class interval size repeatedly till the largest value of the data is
observed in the class.
(6) Distribute the data into respective classes: All the observations are divided into
respective classes by using the tally bar (tally mark) method, which is suitable for
tabulating the observations into respective classes. The number of tally bars is counted to
get the frequency against each class. The frequency of all the classes is noted to get the
grouped data or frequency distribution of the data. The total of the frequency columns
must be equal to the number of observations.
Page 8 of 12
2. Number of Classes = 1+3.322logN
Number of Classes = 1+3.322log57
Number of Classes = 1+3.322(1.75587) = 6.833
Approximately 7 class intervals
Effect of grouping:
As a result of grouping, it is possible to detect a pattern in the figures but grouping results in the
loss of information i.e. calculations made from a grouped frequency distribution can never be
exact, and consequently excessive accuracy can only result in spurious accuracy.
Page 9 of 12
Construction of Frequency Distribution, Relative frequency and Cumulative Relative
Frequency
Example 2:
The following data represents the percent change in tuition levels at public, four-year colleges
(inflation adjusted) from 2008 to 2013 (Weismann, 2013). Create a frequency distribution,
histogram, and ogive for the data.
19.5 40.8 57.0 15.1 17.4 5.2 13.0 15.6 51.5 15.6 14.5 22.4 19.5 31.3 21.7 27.0
13.1 26.8 24.3 38.0 21.1 9.3 46.7 14.5 78.4 67.3 21.1 22.4 5.3 17.3 17.5 36.6
72.0 63.2 15.1 2.2 17.5 36.7 2.8 16.2 20.5 17.8 30.1 63.6 17.8 23.2 25.3 21.4
28.5 9.4;
Solution:
2. Pick the number of classes: Since there are 50 data points, Let’s use 8.
Every value in the data should fall into exactly one of the classes. No data values should fall right
on the boundary of two classes.
Page 10 of 12
Table 2.2. Frequency Distribution for Tuition Levels at Public, Four-Year Colleges
Page 11 of 12
Tutorial Questions
1. Construct a frequency distribution with the suitable class interval size for
marks obtained by a class of 50 students as given below:
23, 50, 38, 42, 63, 75, 12, 33, 26, 39, 35, 47, 43, 52, 56, 59, 64, 77, 15, 21, 51, 54,
72, 68, 36, 65, 52, 60, 27, 34, 47, 48, 55, 58, 59, 62, 51, 48, 50, 41, 57, 65, 54, 43,
56, 44, 30, 46, 67, 53
b. Create the column of class boundaries, Class marks, %Relative frequency and
Cumulative frequency.
2. The following is the distribution of ages of new employees at a factory
a. Obtain the class boundaries and class marks of the class intervals
b. What is the upper-class limit of the class 30-39?
c. What is the lower-class limit of the class 50-59?
d. What is the class mark of the class 40-49?
e. What is the class width of the class 40-49?
f. What is the lower-class boundary of the class 30-39?
3. A medical research team studied the ages of patients who had strokes caused by stress. The
ages of 34 patients who suffered stress strokes were as follows.
29 30 36 41 45 50 57 61 28 50 36 58
60 38 36 47 40 32 58 46 61 40 55 32
61 56 45 46 62 36 38 40 50 27
Page 12 of 12
3.0 GRAPHICAL PRESENTATION OF DATA
. The representation of quantitative data suitably through charts and diagrams is known as
Graphical Representation of Data (Statistical Information). Graphs are used for presenting
statistical data in an attractive way. They enable us to visualize the whole meaning of
complex data at a glance. The main object of diagrammatic representation is to emphasis
the relative position of different subdivisions and not simply to record details.
Advantages of Graphical Representation:
i. It is easily understood by all.
ii. The data can be presented in a more attractive form.
iii. It shows the trend and tendency of values of the variable.
iv. Diagrammatic representations are useful to detect mistakes at the time of data
computations.
v. It shows relationship between two or more sets of figures.
vi. It has the universal applicability.
vii. It is helpful in assimilating the data readily and quickly.
Disadvantages of Graphical Representation:
i. It does not show details or all the facts.
ii. Graphical representation can reveal only the approximate position.
iii. It takes a lot of time to prepare of graph.
The Different Types of Diagrammatic Representation:
There are various types of graphs in the form of charts and diagrams. Some of them are:
1. Line Diagram or Graph.
2. Bar Chart
3. Pie Chart
4. Stem plot
5. Histogram.
6. Frequency polygon.
7. Ogives (cumulative frequency polygon).
3.1 Modes of Graphical Representation of Data:
The data in the form of raw scores is known as ungrouped data and when It is organized into
frequency distribution then it is referred to as grouped data.
Separate modes and methods are used to represent these two types of data grouped.
Page 1 of 21
PIE CHART
A pie chart is an angular representation of a statistical data with several sub-divisions in a
circle. It is a graph that shows the relative frequency distribution of a nominal variable. Each
component is expressed as a percentage of total value. The size of the angle shows their
relative frequency.
This type of graph can be a good choice when you want to emphasize that one variable is
especially frequent or infrequent, or you want to present the overall composition of a variable.
A disadvantage of pie charts is that it’s difficult to see small differences between frequencies.
As a result, it’s also not a good option if you want to compare the frequencies of different
values.
Example 1:
A campus press polled a sample of 280 undergraduate students in order to study student attitude
toward a proposed change in the dormitory regulations. The following tables shows the response
Oppose 51 51 51
× 3600 = 660 × 100 = 18%
280 280
Total 280 3600 100%
Page 2 of 21
Figure 1.0-: Pie Chart Showing the distributions of Students’ responses
BAR CHART
A bar chart is a graph that shows the frequency or relative frequency distribution of a
categorical variable (nominal or ordinal).
The y-axis of the bars shows the frequencies or relative frequencies, and the x-axis shows the
values. Each value is represented by either horizontal or vertical bars, and the length or height
of the bar shows the frequency of the value.
The following types of Bar chart are commonly used:
1. Simple Bar Chart
2. Multiple/Clustered Bar Chart and
3. Component/Stacked Bar Chart
A bar chart is a good choice when you want to compare the frequencies of different values or
groups. Some bar graphs present bars clustered in groups of more than one (Multiple bar
graphs), and others show the bars divided into subparts to show cumulative effect (stacked or
component bar graphs). Bar graphs are especially useful when categorical data is being used.
It’s much easier to compare the heights of bars than the angles of pie chart slices.
SIMPLE BAR GRAPH
This graph is used to show a single characteristic of group, categories or years.
It consists of a number of equally spaced vertical bars of uniform width originating from a
horizontal axis and is shaded. The bars are usually arranged according to relative magnitude
of bars. The limitation of simple bar chart is that only one variable can be represented on it.
Example 2:
In a study of 100 women, the numbers show below indicate the major reason why each
surveyed worked outside her home. Construct a bar graph for the data.
Page 3 of 21
Table 3.2: Frequency distribution of why women work outside the home.
Reason No. of Women
A: To support self/family 62
B: For extra money 18
C: For something different to do 12
D: others 8
Frequency distribution of why women work outside the home.
60
50
40
No of Women
30
20
10
0
A B C D
Reasons
Example 3: Create simple bar Chart of the data below represent the revenue(dollars)
collected by a State for the months of April to August:
Amounts(million$) 7 12 28 3 41
20
10
0
Months
Figure 1.1: A Bar Chart of Why Women Work Outside the Home
CLUSTERED AND STACKED BAR CHART
CLUSTERED / MULTIPLE BAR CHART: Clustered Bar Chart (also known as Grouped
bar chart or Multiple bar chart) is great for displaying and comparing multiple sets of data
over the same categories. In this case numerical values of major categories are arranged in
ascending or descending order so that categories can be readily distinguished.
Page 4 of 21
For example:
1. Sales revenue of various departments of the company over several years
2. Prices of different commodities at various quarter of a given year.
Example 3:
1. The table below shows the distribution of students offering STA131 by Departments and
gender.
Table 3.3: Distribution of students offering STA131 Course
Gender Departments
STA Maths Engineering Tele Info Total
Com. Tech
Male 40 35 110 40 75 300
Female 30 20 60 35 55 200
Total 0 55 170 75 130 500
Male
100
Female
80
No of Students
60
40
20
0
Departments
Page 5 of 21
i. Each bar in the component bar is subdivided into several component parts.
ii. A single bar represents the aggregate value whereas the component parts represent
the component values of the aggregate value.
iii. It shows the relationship between the different components and difference in main
bar.
iv. Different shades or colours are used to distinguish the various components.
Example 4: Using the data in table 3.3, draw a stacked or component bar chart.
Gender Departments
STA Maths Engineering Tele Info Total
Com. Tech
Male 40 35 110 40 75 300
Female 30 20 60 35 55 200
Total 70 55 170 75 130 500
Distribution of students offering ST A131 course
100 120 140
No of Students
80
60
40
20
0
Departments
Male
Female
120
100
No of Students
80
60
40
20
0
Departments
Page 6 of 21
Class work:
Example 5:
A researcher asked her class to pick who would win in a battle of superhero. Below is a
frequency table and charts of the results:
Figure 1.4: Bar chart showing the proportion of votes for Superhero Battle Contestants
Page 7 of 21
Figure 1.5: Pie chart showing the proportion of votes for Superhero Battle Contestants.
Example 5: The data below represent the scores for the first exam were as follows (smallest
to largest):
33; 42; 49; 49; 53; 55; 55; 61; 63; 67; 68; 68; 69; 69; 72; 73; 74; 78; 80; 83; 88; 88; 88; 90;
92; 94; 94; 94; 94; 96; 100
The stem-plot shows that most scores fell in the 60s, 70s, 80s, and 90s. Eight out of the 31
scores or approximately 26% (8/31) were in the 90s or 100, a fairly high number of As.
Page 8 of 21
Class Work
1. The data are the distances (in kilometres) from a home to local supermarkets. 1.1; 2.3;
3.3; 4.5; 4.7; 2.7; 4.5; 4.8; 6.7; 12.3;1.5; 2.5; 3.2; 4.2; 5.5; 3.5; 3.8; 6.5; 3.3; 4.0; 5.6;
Box Plot: Box plot (Box and Whisker diagram or plot) is a convenient way of graphically
depicting groups of numerical data through their five number summaries: minimum, lower
quartile(Q1), median(Q2), Upper quartile(Q3) and maximum observation with outliers plotted
individually.
Characteristics of Box Plot:
i. A central box spans the quartile
ii. A line in the box mark the median
iii. Observations more than 1.5 X IQR (inter quartile range) outside the central box
plotted individually as possible outlier.
iv. Lines called Whiskers, extend from the box and to the smallest and largest
observations that are not outlier.
v. Box plots can be drawn either horizontally or vertically.
vi. For symmetric distribution, box plots are symmetric but the symmetric box plot
does not imply symmetric distribution.
vii. The upper part of the box i.e. differences between Q3 and Median is bigger than
the lower part (difference between Median and Q1) in case of positively skewed
distribution and smaller in case of negatively skewed.
Uses of Box Plot:
1. The box plots are most useful for comparing distributions.
2. It is a quick way of examining one or more sets of data graphically.
3. The spacings between the different parts of the box indicate the degrees of dispersions
and skewness in the data and identify outliers.
Inter Quartile Range (IQR):
It is a measure of the spread based on quartiles. The range from third quartile(Q3) to first
quartile(Q1) is called interquartile range.
IQR= Q3 – Q1
Semi-Interquartile Range or Quartile Deviation
The half of the difference between the third quartile(Q3) and first quartile(Q1) is called
Semi-Interquartile Range or Quartile Deviation
OUTLIERS:
An outlier is defined as an observation which falls more than 1.5 * IQR (called step)
above Q3 or below Q1
Page 9 of 21
a. If an observation falls more than 3* IQR above Q3 or below Q1, then it is known as
extreme outlier.
b. The observation falling between 1.5 * IQR and 3* IQR above Q3 or below Q1, it is
known as suspect outlier.
c. Inner-fences and Outer fences
i. The value [Q1- 1.5*IQR, Q3 + 1.5* IQR] are known inner-fences
ii. The value [Q1- 3*IQR, Q3 + 3* IQR] are known outer-fences
The centre of distribution A is the lowest of the three distributions (median is 0.11). The
distribution is positively skewed, because the whisker and half-box are longer on the right
side of the median than on the left side.
Distribution B is approximately symmetric, because both half-boxes are almost the same
length (0.11 on the left side and 0.10 on the right side). It’s the most concentrated distribution
because the interquartile range is 0.21, compared to 0.30 for distribution A and 0.26 for
distribution C.
Page 10 of 21
The centre of distribution C is the highest of the three distributions (median is 0.88). The
distribution C is negatively skewed because the whisker and half-box are longer on the left
side of the median than on the right side.
All three distributions include potential outliers. Let’s take distribution A, for example. The
interquartile range is Q3 - Q1 = 0.32 – 0.02 = 0.30. According to the definition used by the
function in R software, all values higher than Q3 + 1.5 x (Q3 - Q1) = 0.32 + 1.5 x 0.30 = 0.77
are outside the right whisker and indicated by a circle. There are two potential outliers in
distribution A.
Building A Box and Whisker Plot
Example 1:
Minimum 82
Lower quartile 94
Median 95
Upper quartile 102
Maximum 110
Example2:
5 7 1 9 11 22 15
Page 11 of 21
Example 1: a simple box and whisker plot
Suppose you have the math test results for a class of 15 students. Here are the results:
91 95 54 69 80 85 88 73 71 70 66 90 86 84 73
This is an odd set of data – you have 15 data points. It means the middle point is 80 as there
are 7 data points above it and 7 numbers below.
Step 3: Find the middle points of the two halves divided by the median (find the upper and
lower quartiles).
Page 12 of 21
So, we can determine that the five-number summary for the class of students is:
54, 70, 80, 88, 95.
Now we are absolutely ready to draw our box and whisker plot.
As you see, the plot is divided into four groups: a lower whisker, a lower box half, an upper
box half, and an upper whisker. Each of those groups shows 25% of the data because we have
an equal amount of data in each group.
LINE GRAPH
Another type of graph that is useful for specific data values is a line graph. In the particular
line graph shown in fig. 1.5, the x-axis (horizontal axis) consists of data values and the y-axis
(vertical axis) consists of frequency points. The frequency points are connected using line
segments. Line graph are appropriate for ungrouped (quantitative) frequency distribution.
Example: In a survey, 40 mothers were asked how many numbers of birth each had. The
results are shown in Table. Construct a line graph.
No of births Frequency
0 2
1 5 Table 3.5: A frequency distribution
2 8 table showing the distribution of births
3 14 by 40 randomly selected mothers in
4 7 Oyun, LGA of Kwara State.
5 4
Total 40
Number of Births
Fig. 1.5: A line graph showing the number of Births for 40 randomly selected mothers in
Oyun, LGA of Kwara State.
Page 13 of 21
Tutorial Question:
In a survey, 40 people were asked how many times per year they had their car in the
mechanic shop for repairs. Construct a line graph of the data.
Number of times in Mechanic Shop Frequency
0 7
1 10
2 14
3 9
Step 3: Draw rectangles with bases as class boundary and corresponding frequencies as
heights.
Step 4 The height of each rectangle is proportional to the corresponding class frequency if the
intervals are equal. The area of every individual rectangle is proportional to the corresponding
class frequency if the intervals are unequal.
.
Page 14 of 21
Example of Histogram
Although bar charts and histograms are similar, there are important differences:
S/No Differences Bar chart Histogram
1 Type of Categorical Quantitative
variable
2 Value Ungrouped (values) Grouped (interval classes)
grouping
3 Bar spacing Space can be between bars Space is not allow between
bars
4 Bar order Can be in any order Can only be ordered from
lowest to highest
FREQUENCY POLYGON
Page 15 of 21
Solution:
To plot a frequency polygon with a histogram, we need to follow these steps to construct a
histogram:
• The cost of living index is represented on the x-axis.
• The number of weeks is represented on the y-axis.
• Now rectangular bars of widths equal to the class- size and the length of the bars
corresponding to a frequency of the class interval are drawn.
To calculate the midpoint, we use the formula:
Class mark = (Upper Limit + Lower Limit) / 2 So,
Class mark = (150 + 140)/2 = 145, (160 + 150)/2 = 155 and so on
While plotting the graph, we also mark the before and after class mark as well. In this case, the
before is 135 and the after is 205. ABCDEFGH represents the given data graphically in form of
frequency polygon i.e. those are the midpoints. Hence, the frequency polygons graph is given below:
Page 16 of 21
Difference between Frequency Polygons and Histogram:
Even though a frequency polygon graph is similar to a histogram and can be plotted with or
without a histogram, the two graphs are yet different from each other. The two graphs have
their own unique properties that show the difference visually. The differences are:
Page 17 of 21
Step 3: Against each upper-class boundary, plot the cumulative frequencies.
Step 4: Connect all the points from step 3 with straight lines
Step 5: Make sure you include the point with the lowest class boundary and the 0 cumulative
frequency.
Class Work:
Example: Draw an ogive for the following frequency distribution of test marks for a group of
50 students by less than method.
Marks 20-24 25-29 30-34 35-39 40-44 45-49 50-54 55-59 60-64 65-69
Frequency 1 2 4 8 11 9 7 4 3 1
Solution:
Page 18 of 21
Fig. 1.6: Cumulative frequency polygon (Ogive) for test scores
Uses of Ogive
1. It is used to find the median, quartiles, decile and percentiles.
2. It is also useful in finding the cumulative frequency corresponding to a given
value of the variable.
3. To find out the number of observations which are expected to lie between two
given values.
General Rules for Graphical Representation of Data
There are certain rules to effectively present the information in the graphical representation.
They are:
1- Suitable Title: Make sure that the appropriate title is given to the graph which indicates
the subject of the presentation.
2- Measurement Unit: Mention the measurement unit in the graph.
3- Proper Scale: To represent the data in an accurate manner, choose a proper scale.
4- Index: Index the appropriate colors, shades, lines, and design in the graphs for better
understanding.
5- Data Sources: Include the source of information wherever it is necessary at the bottom
of the graph.
6- Keep it Simple: Construct a graph in an easy way that everyone can understand.
7- Neat: Choose the correct size, fonts, colors. etc in such a way that the graph should be a
visual aid for the presentation of information
Tutorial Questions:
Question 1
The data below were the ages of 30 randomly selected staff in a particular university:
32; 54; 32; 42; 33; 34; 50; 38; 40; 42; 53; 56; 43; 44; 51; 52; 46; 47; 60; 47; 48; 52;
57; 48; 50; 52; 48; 49; 57; 61
Page 19 of 21
Question 2
2. The table below shows the frequency distribution of test marks for 120
students.
Marks
1-10 11-20 21-30 31-40 41-50 51-60 61-70 71-80 81-90 91-100
(%)
Number
of 1 6 8 15 17 24 22 15 9 3
students
Draw the following graph:
i. Frequency polygon
ii. Histogram
iii. Less than and more than cumulative frequency polygon (ogive)
Question 4
c. What are the median and the semi-interquartile range of the following
sample?
Page 20 of 21
Question 5
The following shows the distribution of scores obtained by SS2 students in mathematics
mock exam.
Class 21-30 31-40 41-50 51-60 61-70 71-80 81-90 91-100
interval
Frequency 2 6 7 10 11 9 4 1
Question 6
1. The table contains data from a study of daily study time for 40 students from
Statistics 101. Construct an ogive from the data.
Minutes on
Homework Number ofstudents Relativefrequency Cumulativefrequency
0-15 2 0.05 2
16-30 4 0.10 6
31-45 8 0.20 14
46-60 18 0.45 32
61-75 4 0.10 36
76-90 4 0.10 40
2. An annual survey sent to retail store managers contained the question "Did your store suffer
any losses due to employee theft?" The responses are summarized in the table for two years,
2000 and 2005. Construct a multiple bar graph of the data, then describe any trends.
Page 21 of 21
MEASURES OF LOCATION:
Mean or Average: The arithmetic mean is the sum of all observations divided by the number of
observations. The mean of a set of n observations 𝑥1 , 𝑥2 , … . . , 𝑥𝑛 is defined as
𝒏
𝟏 𝟏
̅ = (𝒙𝟏 + 𝒙𝟐 + ⋯ … . . +𝒙𝒏 ) = ∑ 𝒙𝒊
𝑿
𝒏 𝒏
𝒊=𝟏
(𝒇𝟏 𝒙𝟏 + 𝒇𝟐 𝒙𝟐 + ⋯ … . . +𝒇𝒏 𝒙𝒏 ) ∑ 𝒇𝒊 𝒙𝒊
̅=
𝑿 =
𝑓1 + 𝑓2 , + ⋯ … . +𝑓𝑛 ∑ 𝒇𝒊
Merits
1. The mean is simple and easily understood , most popular measures of central tendency.
2. Every observation in the data is used to compute the mean.
3. The mean is unique,there is only one mean for any given data sets.
4. The mean is used for further statistical analysis.
Demerits:
Example 1: The following are weights in pounds of 10 randomly selected children at a day-care
center: 68 ,63, 42, 27, 30, 36 ,28, 32 ,79 ,27. Compute the sample mean
Solution: n=10
1 1
𝑋̅=𝑛 (𝑥1 + 𝑥2 + ⋯ … . . +𝑥𝑛 ) = 𝑛 ∑𝑛𝑖=1 𝑥𝑖
1
𝑋̅ = (68 + 63 + 42 + 27 + ⋯ … . . +27)
10
𝑛
1
𝑋̅ = ∑ 432 = 43.2
10
𝑖=1
The average or mean weigths of the 10 children is 43.2 pounds
1
B. Arithmetic mean for ungrouped/ grouped frequency data, the mean is defined as:
(𝒇𝟏 𝒙𝟏 + 𝒇𝟐 𝒙𝟐 + ⋯ … . . +𝒇𝒏 𝒙𝒏 ) ∑ 𝒇𝒊 𝒙𝒊 ∑ 𝒇𝒙
̅=
𝑿 = = ;
𝑓1 + 𝑓2 , + ⋯ … . +𝑓𝑛 ∑ 𝒇𝒊 𝑵
where 𝑁 = ∑ 𝒇𝒊 𝑎𝑛𝑑 𝑥𝑖 are the mid-points or class marks for grouped frequency
STEPS:
Marks 30 40 50 60 70 80 90
No. of 15 20 10 15 20 15 5
students
Solution: this is an Ungrouped frequency data
Marks(x) No. of 𝒇𝒙
students(f)
30 15 450
40 20 800
50 10 500
60 15 900
70 20 1400
80 15 1200
90 5 450
Total ∑ 𝒇𝒊 = 𝟏𝟎𝟎 ∑ 𝒇𝒙 =5,700
∑ 𝒇𝒊 𝒙𝒊 𝟓𝟕𝟎𝟎
̅ ==
𝑿 = = 𝟓𝟕
∑ 𝒇𝒊 𝟏𝟎𝟎
2
Example 3: Compute the mean of the grouped frequency data given below:
Solution:
Class mark(x) is the mid –point of each interval defined as the average of the lower and upper
limits of a given class interval
∑ 𝑓𝑖 𝑥𝑖 = 2086.5 and ∑ 𝒇𝒊 = 𝟓𝟕
∑ 𝒇𝒊 𝒙𝒊 𝟐𝟎𝟖𝟔. 𝟓
𝑋̅ = = = 𝟑𝟔. 𝟔𝟏𝒍𝒃
∑ 𝒇𝒊 𝟓𝟕
Step;
1. Assumed mean is taken to be the middle value of x (or class mark(x)) which correspond
to the middle value of the frequency distribution.
3
a. In case of n observations data, the mean is computed as:
𝒏
𝟏
̅ = 𝑨 + ∑ 𝒅𝒊
𝑿
𝒏
𝒊=𝟏
𝑤ℎ𝑒𝑟𝑒 𝑑𝑖 = 𝑋𝑖 − 𝐴 and A= assumed Mean
b. For frequency distribution, the mean is computed as:
Step:
i. Obtain the class mark(x) or mid-point(x)
ii. Create a column for 𝑑𝑖 = 𝑋𝑖 − 𝐴; i.e. deviation of each value of x from the
assumed mean
iii. Multiply the respective class frequency by the corresponding deviation(𝑑𝑖 ) and
sum up the column to get ∑𝒏𝒊=𝟏 𝒇𝒅𝒊
iv. The Arithmetic mean by Assumed mean or shot cut method is given by:
𝒏
𝟏
̅ = 𝑨 + ∑ 𝒇𝒅𝒊
𝑿
𝒏
𝒊=𝟏
Example 4: Find the mean weight of the following students by assumed mean method. The
weights are in kg: 67, 69, 66, 68, 63, 76, 72, 74, 70, 65
Value (x) 𝑑𝑖 = 𝑋𝑖 − 68
63 -5
65 -3
66 -2
67 -1
68 0
69 1
70 2
72 4
74 6
76 8
𝒏
∑ 𝒅𝒊 = 𝟏𝟎
𝒊=𝟏
4
Example 3: Compute the mean of the grouped frequency data given below using an assumed
mean of 44.5
5
GEOMETRIC MEAN
The geometric mean of a series containing n observations is the nth root of the product of the
values. The geometric mean(GM) of a set of n observations 𝑥1 , 𝑥2 , … . . , 𝑥𝑛 is defined as
𝑮𝒎 = 𝒏√𝑥1 × 𝑥2 × … . .× 𝑥𝑛
𝟏
𝑮𝑴 = (𝑥1 × 𝑥2 × … . .× 𝑥𝑛 )𝒏
Taking the log of both sides
1 1
𝑙𝑜𝑔𝐺𝑀 = 𝑙𝑜𝑔(𝑥1 + 𝑥2 + ⋯ . . +𝑥𝑛 ) = (𝑙𝑜𝑔𝑥1 + 𝑙𝑜𝑔𝑥2 + ⋯ . . +𝑙𝑜𝑔𝑥𝑛 )
𝑛 𝑛
∑ 𝑙𝑜𝑔𝑥𝑖
𝑙𝑜𝑔𝐺𝑀 =
𝑛
Therefore,
∑ 𝑙𝑜𝑔𝑥𝑖
𝑮𝑴 = 𝑨𝒏𝒕𝒊𝒍𝒐𝒈 [ ]
𝑛
Example 1: The data below represent the scores of nine(9) students in a test. Find
the geometric mean score. 45, 32, 37, 46, 39, 36, 41, 48, 36.
Solution:
Scores(𝑥𝑖 ) 𝑙𝑜𝑔𝑥𝑖
45 1.6532
32 1.5052
37 1.5682
46 1.6628
39 1.5911
36 1.5563
41 1.6128
48 1.6212
36 1.5563
∑ 𝑙𝑜𝑔𝑥𝑖 = 14.387
∑ 𝑙𝑜𝑔𝑥𝑖 14.387
𝑮𝑴 = 𝑨𝒏𝒕𝒊𝒍𝒐𝒈 [ ] = 𝒂𝒏𝒕𝒊 − 𝒍𝒐𝒈 [ ] = 𝒂𝒏𝒕𝒊𝒍𝒐𝒈(𝟏. 𝟓𝟗𝟖𝟓𝟔)
𝑛 9
𝑮𝑴 = 𝟑𝟗. 𝟔𝟗
6
Example 1: The following frequency distribution of weights of 60 apples.
Find the geometric mean for the grouped data.
marks 65-84 85-104 105-124 125-144 145-164 165-184 185-204 Total
No of 9 10 17 10 5 4 5 60
students
Solution:
∑ 𝑓𝑖 𝑙𝑜𝑔𝑥𝑖 93.4559
𝐺𝑀 = 𝐴𝑛𝑡𝑖𝑙𝑜𝑔 [ ∑ 𝑓𝑖
] = 𝐴𝑛𝑡𝑖𝑙𝑜𝑔 [ ] = 𝐴𝑛𝑡𝑖𝑙𝑜𝑔(1.5576)
60
𝐺𝑀 = 𝐴𝑛𝑡𝑖𝑙𝑜𝑔(1.5576) = 36.11
7
HARMONIC MEAN
Harmonic Mean is the reciprocal of the arithmetic mean of the reciprocals of a finite set of
numbers. It is calculated by dividing the number of observations by the sum of reciprocals of
each value of the data set.
𝑛
𝐻𝑀 = [ ] 𝑎𝑛𝑑
1
∑𝑖=1 ( )
𝑥𝑖
𝑛
𝐻𝑀 = [ ]
1
∑𝑖=1 𝑓𝑖 ( )
𝑥𝑖
Example1: The following are weights in pounds of 10 randomly selected children at a day-care
center: 68 ,63, 42, 27, 30, 36 ,28, 32 ,79 ,27. Compute the Harmonic mean
Solution
Values(𝑥𝑖 ) 1
𝑥𝑖
68 0.0147
63 0.0158
42 0.0238
27 0.03704
30 0.03333
36 0.02778
28 0.0357
32 0.03125
79 0.01266
27 0.03704
Total 1
∑ ( ) = 0.2691
𝑥𝑖
𝑖=1
8
𝑛 10
𝐻𝑀 = [ ]= = 37.16
1 0.2691
∑𝑖=1 ( )
𝑥𝑖
Class mark(𝑥𝑖 ) f 1 1
𝑓( )
𝑥𝑖 𝑥𝑖
(65+84)/2= 74.5 9 0.0134 0.1206
94.5 10 0.0106 0.1058
114.5 17 0.00873 0.1484
134.5 5 0.00744 0.0372
154.5 4 0.00647 0.02588
174.5 5 0.00573 0.02865
Total ∑ 𝑓𝑖 =60 0.46653
𝑛 60
𝐻𝑀 = [ ]= = 128.61
1 0.46653
∑𝑖=1 𝑓𝑖 ( )
𝑥𝑖
Merits of H.M
1. It is rigidly defined.
4. It is the most suitable average when it is desired to give greater weight to smaller observations
and less weight to the larger ones.
Demerits of H.M
9
4. It gives greater importance to small items and is therefore, useful only when small items
have to be given greater weightage.
5. It is rarely used in grouped data
TUTORIAL QUESTION
Find
i. Arithmetic Mean
ii. Geometric Mean and
iii. Harmonic Mean.
2. Find the Arithmetic mean, Harmonic mean and Geometric Mean of each set of
observations below.
(i) 3, 4, 7, 3, 5, 2, 6, 10
(ii) 8, 10, 12, 14, 7, 16, 5, 7, 9, 11
(iii) 17, 18, 16, 17, 17, 14, 22, 15, 16, 17, 14, 12
3. A survey of 100 households on the number of cars own in each household resulted in the
data given below.
No of cars 0 1 2 3 4 Total
No of 5 70 21 3 1 100
students
Calculate the
i. Arithmetic mean number of cars per household
ii. Geometric mean number of cars per household
iii. Harmonic mean number of cars per household
4. A manager keeps a record of the number of calls she makes each day on her mobile
phone.
No of calls 0 1 2 3 4 5 6 7 8 Total
per day
frequency 3 4 7 8 12 10 14 3 1 62
Calculate the Arithmetric mean, Harmonic and Geometric mean number of calls per day
Frequency(f) 6 16 21 8 51
10
MEDIAN:
Another useful measure of location is the median. If the observations in the data set are
arranged in increasing or decreasing order, the median is the middle observation, which divides
the data set into equal halves.
Procedures:
11
iii. The median is calculated by
𝑛+1 𝑡ℎ
𝑀=( 2
) ; 𝑛 = ∑𝑓
iv. The first value with a cumulative frequency greater than depth of the median is the
median. However, if the depth of the median is exactly 0.5 more than the
cumulative of the previous class, then the median is the mid-point between the two
classes.
Example1; The data below represent the monthly allowance of 100 students
Solution: Let arrange the allowance in ascending order and obtain the cumulative frequency
Allowance in No of Cumulative
ascending Students(f) frequency (Cf)
order
80 12 12
100 20 32
150 22 54
180 26 80
200 16 96
260 4 100
Total 100
Since n=100(even)
𝑀𝑒𝑑𝑖𝑎𝑛 = 150
12
C. For Grouped Frequency data:
Procedures:
i. Here data is given in the form of a frequency table with class interval.
ii. Obtain the cumulative frequencies column.
𝑡ℎ
iii. The median class is determined by (∑2𝑓) 𝑡𝑒𝑟𝑚
iv. The median is calculated by the formular:
∑𝑓 𝑤𝑚
Median=𝐿𝑚 + [ − 𝐶𝐹𝑏𝑚 ] ×
2 𝑓𝑚
∑ 𝑓 = Total frequency
Advantages of Median
Disadvantages of Median
i. The median unlike the mean does not use the whole data set.
ii. The median does not use for any further statistical analysis.
Example 1: The following data gives the marks obtained by students in STA 131.
Find the median.
marks 11-20 21-30 31-40 41-50 51-60 61-70 71-80 Total
No of 42 38 125 84 45 36 30 400
students
13
Solution:
[200−80] 120
Median=30.5 + × 10 = 30.5 + 205 × 10
205
14
Weight frequency Cumulative
interval frequency
(CF)
10- 19 5 5
20- 29 19 24
30- 39 10 34
40- 49 13 47
50- 59 4 51
60- 69 4 55
70- 79 2 57
Total 57
Using the CF column, the 29th observation falls in class interval 30-39 and the boundaries of
that class is 29.5-39.5
Tutorial Questions
15
MODE
Mode of a set of data is defined as the value that occurs most frequently among the value of
the variable. In frequency distribution any value with the highest frequency is the mode.
For ungrouped frequency
By inspection:
Procedure: In case of n observation or ungrouped frequency data, mode can be obtained by
examine the data value with the highest frequency.
Example1: Find the mode from the data given below:
1. 1, 1, 3, 3, 1, 5, 3, 2, 2, 3, 5, 3, 3, 6, 7, 2, 3, 4, 4,6
Solution:
Value(x) No of times (f)
1 3
2 3
3 7
4 2
5 2
6 2
7 1
Total ∑ 𝑓 = 𝑁 =20
The table indicate that the value “3” has the maximum frequency.
Hence, mode is 3
The mode is especially useful for describing qualitative variables or quantitative variables that
take on a small number of possible values. If two values occur more than others but equally
frequently, we say the data are bimodal, or more generally multimodal. The term bimodal is
also used to describe distribution in which there are two peaks, not necessarily of the same
height.
For grouped frequency data, the Mode is defined as:
Procedure:
i. The modal class is the class with the highest frequency
ii. Find the mode using the formular below:
𝑜 𝑓 −𝑓𝑏
Mode=𝐿𝑜 + [2𝑓 −𝑓 ] × 𝑤𝑜
𝑜 𝑏 −𝑓𝑎
𝐿𝑜 = lower class boundary of the modal class
𝑓𝑜 - Frequency of the modal class
𝑓𝑏 - Frequency before the modal class
𝑓𝑎 - Frequency after the modal class
𝑤𝑜 − Class size of the modal class
Advantages of Mode
16
i. Mode is the easiest average to compute.
ii. Mode is not generally influenced by extreme values.
iii. Mode can be determined for categorical variables
Disadvantages of Mode
i. The mode may not exist.
ii. The mode may not be unique.
iii. The Mode does not lend itself to further statistical Analysis.
Example: Compute the Mode and Median of the grouped data below:
Frequency(f) 5 19 10 13 4 4 2 57
Solution:
𝑜 𝑓 −𝑓𝑏
1. Mode=𝐿𝑜 + [2𝑓 −𝑓 ] × 𝑤𝑜
𝑜 𝑏 −𝑓𝑎
Therefore,
14
Mode=19.5 + [23] × 10 = 19.5 + 0.6086 𝑋10 = 19.5 + 6.086
Mode= 25 .59
17
2.1 MEASURE OF PARTITION
The partition values are the measures used to divide the total number of observations from a
distribution into a certain number of equal parts. The measure of partition are;
I. Measure of quartile
II. Measure of decile
III. Measure of percentile
𝑁 + 1 𝑡ℎ
𝑄𝑖 = 𝑖 × ( ) 𝑖𝑡𝑒𝑚 ; 𝑤ℎ𝑒𝑟𝑒 𝑖 = 1, 2, 3
4
2. For grouped frequency data, the formula is then given by:
𝑖×𝑁
( 4 − 𝐶𝑓𝑏 )
𝑄𝑖 = 𝑙1 + [ ]×𝑐
𝑓𝑖
8+1 𝑡ℎ 9 𝑡ℎ
𝑄1 = 1 × ( ) 𝑖𝑡𝑒𝑚 = (4) 𝑖𝑡𝑒𝑚
4
𝑄1 = (2.25)𝑡ℎ 𝑖𝑡𝑒𝑚
𝑄1 = (2)𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 + 0.25(3𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 − 2𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 )
𝑁+1 𝑡ℎ
𝑄𝑖 = 𝑖 × ( ) 𝑖𝑡𝑒𝑚
4
8+1 𝑡ℎ 27 𝑡ℎ
𝑄3 = 3 × ( ) 𝑖𝑡𝑒𝑚 = ( 4 ) 𝑖𝑡𝑒𝑚
4
𝑄3 = (6.75)𝑡ℎ 𝑖𝑡𝑒𝑚
𝑄3 = (6)𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 + 0.75(7𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 − 6𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 )
𝑄3 = 52 + 0.75(57 − 52) = 52 + 3.75 = 55.75
QUARTILE FOR UNGROUPED FREQUENCY DATA:
1. Find cumulative frequencies
2. Find
𝑁 + 1 𝑡ℎ
𝑄𝑖 = 𝑖 × ( ) 𝑖𝑡𝑒𝑚
4
𝑁+1 𝑡ℎ
3. Check the cumulative frequencies, the value just greater than 𝑖 × ( ) 𝑖𝑡𝑒𝑚 , the
4
corresponding value of x is Qi
Example 2.1.2: For ungrouped frequency data
Compute Q1 and Q3 for the data relating to age in years of 543 members in a village
Age 20 30 40 50 60 70 80
No of 3 61 132 153 140 51 3
member
Solution:
X F CF
20 3 3
30 61 64
40 132 196
50 153 349
60 140 489
70 51 540
80 3 543
Total ∑ 𝑓 = 543
𝑁 + 1 𝑡ℎ
𝑄𝑖 = 𝑖 × ( ) 𝑖𝑡𝑒𝑚
4
First Quartile (𝑸𝟏 ):
𝑁 + 1 𝑡ℎ
𝑄1 = 𝑖 × ( ) 𝑖𝑡𝑒𝑚
4
543 + 1 𝑡ℎ 544 𝑡ℎ
𝑄1 = 1 × ( ) 𝑖𝑡𝑒𝑚 = ( ) 𝑖𝑡𝑒𝑚
4 4
𝑄1 = (136)𝑡ℎ 𝑖𝑡𝑒𝑚
From the cumulative column, the corresponding value of x for which the cumulative frequency
of 136 is found is the ith quartile value. So, after cumulative frequency of 64, value of 40 has 65
to 196 observations in which 136 is within the range. Therefore,
𝑄1 = 40 𝑦𝑒𝑎𝑟𝑠
𝑁+1 𝑡ℎ
𝑄3 = 3 × ( ) 𝑖𝑡𝑒𝑚
4
543 + 1 𝑡ℎ
𝑄1 = 3 × ( ) 𝑖𝑡𝑒𝑚
4
𝑄1 = (3 × 136)𝑡ℎ 𝑖𝑡𝑒𝑚 = (408)𝑡ℎ 𝑖𝑡𝑒𝑚
𝑄3 = 60 𝑦𝑒𝑎𝑟𝑠
FOR GROUPED FREQUENCY DATA, THE FORMULA IS THEN GIVEN BY:
1. Find the cumulative frequency
2. Find the ith quartile using:
𝑖 × 𝑁 𝑡ℎ
𝑄𝑖 = ( ) 𝑖𝑡𝑒𝑚
4
3. 𝑄𝑖 Class is the class interval corresponding to the value of the cumulative frequency just
greater than (i × n)/4.
Where i = 1, 2, 3, 4
N = total of all frequency value
𝑙𝑖 = lower class limit of the ith quartile class
𝑓𝑖 = frequency of the ith quartile class
𝑐 = class interval
𝑓𝐶𝑏 = cumulative frequency preceding the quartile class
Solution:
wages Frequency CF
30-32 12 12
𝑄1= 32-34 18 30
34-36 16 46
36-38 14 60
𝑸𝟑= 38- 12 72
40
40-42 8 80
42-44 6 86
Total ∑ 𝑓 = 86
(21.5 − 12)
𝑄1 = 32 + [ ]×2
18
(19)
𝑄1 = 32 + [ ]
18
𝑄1 = 32 + 1.06 = 𝟑𝟑. 𝟎𝟔
(64.5 − 60)
𝑄3 = 38 + [ ]×2
12
(9)
𝑄3 = 38 + [ ]
12
𝑄3 = 38 + 0.75 = 𝟑𝟖. 𝟕𝟓
𝑄3−𝑄1
𝑆𝑒𝑚𝑖 𝑖𝑛𝑡𝑒𝑟 − 𝑞𝑢𝑎𝑟𝑡𝑖𝑙𝑒 𝑟𝑎𝑛𝑔𝑒(𝑆𝐼𝑄𝑅) = 2
𝑄3−𝑄1
2. 𝑆𝑒𝑚𝑖 𝑖𝑛𝑡𝑒𝑟𝑞𝑢𝑎𝑟𝑡𝑖𝑙𝑒 𝑟𝑎𝑛𝑔𝑒 = 2
55.75 − 29.7
𝑆𝐼𝑄𝑅 = = 14.88
2
From example 2.1.3;
The first and third quartile are:
𝑄3−𝑄1
2. 𝑆𝑒𝑚𝑖 𝑖𝑛𝑡𝑒𝑟𝑞𝑢𝑎𝑟𝑡𝑖𝑙𝑒 𝑟𝑎𝑛𝑔𝑒 = 2
𝑁 + 1 𝑡ℎ
𝐷𝑖 = 𝑖 ( ) 𝑖𝑡𝑒𝑚 ; 𝑤ℎ𝑒𝑟𝑒 𝑖 = 1, 2, 3, 4, … . . ,9
10
B. For grouped frequency data, the formula is then given by:
𝑖×𝑁
( 10 − 𝐶𝑓𝑏 )
𝐷𝑖 = 𝑙1 + [ ]×𝐶
𝑓𝑖
Examples 2.2.1
1. Find the D6 for the following data 11, 25, 20, 15, 24, 28, 19, 21.
Solution:
Arrange in an ascending order 11, 15, 19, 20, 21, 24, 25, 28
N=8
𝑁 + 1 𝑡ℎ
𝐷𝑖 = 𝑖 × ( ) 𝑖𝑡𝑒𝑚
10
8 + 1 𝑡ℎ 9 𝑡ℎ
𝐷6 = 6 × ( ) 𝑖𝑡𝑒𝑚 = 6 × ( ) 𝑖𝑡𝑒𝑚
10 10
54 𝑡ℎ
𝐷6 = ( ) 𝑖𝑡𝑒𝑚 = (5.4)𝑡ℎ 𝑖𝑡𝑒𝑚
10
𝑡ℎ
𝐷6 = (5) 𝑣𝑎𝑙𝑢𝑒 + 0.4(6𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 − 5𝑡ℎ 𝑣𝑎𝑙𝑢𝑒)
𝐷6 = 21 + 0.4(24 − 21) = 21 + 0.4(3)
𝐷6 = 21 + 1.2 = 22.2
3. 𝐷𝑖 Class is the class interval corresponding to the value of the cumulative frequency just
greater than (i × N)/10.
Example 2.2.2: Calculate D5 for the frequency distribution of monthly income of workers in a factory.
Income (in 0-4 4-8 8-12 12-16 16-20 20-24 24-28 28-32
thousand)
No of 10 12 8 7 5 8 4 6
persons
Solution:
Class F CF
0-4 10 10
4-8 12 22
8-12 8 30
12-16 7 37
16-20 5 42
20-24 8 50
24-28 4 54
28-32 6 60
Total ∑ 𝑓 = 60
𝑖 × 𝑁 𝑡ℎ
𝐷𝑖 = ( ) 𝑖𝑡𝑒𝑚
10
5 × 60 𝑡ℎ
𝐷5 = ( ) 𝑖𝑡𝑒𝑚 = (30)𝑡ℎ 𝑖𝑡𝑒𝑚
10
Checking the cumulative frequency column, 30th observation fall in 3rd class, hence
The 5th decile lies in the group 8 – 12
𝑖×𝑁
( 10 − 𝐶𝑓𝑏 )
𝐷𝑖 = 𝑙1 + [ ]×𝑐
𝑓𝑖
(30 − 22)
𝐷5 = 8 + [ ]×4
8
8
𝐷5 = 8 + [ ] × 4 = 8 + 4
8
𝐷5 = 12
𝑁 + 1 𝑡ℎ
𝑃𝑖 = 𝑖 × ( ) 𝑖𝑡𝑒𝑚 ; 𝑤ℎ𝑒𝑟𝑒 𝑖 = 1, 2, 3, 4, … . . ,99
100
B. For grouped frequency data, the formula is then given by:
𝑖×𝑁
( 100 − 𝐶𝑓𝑏 )
𝑃𝑖 = 𝑙1 + [ ]×𝐶
𝑓𝑖
𝑤ℎ𝑒𝑟𝑒 𝑖 = 1, 2, 3, 4, … . . ,99
N=∑ 𝑓 = total of all frequency value
𝑙𝑖 = lower class limit of the percentile class
𝑓𝑖 = frequency of the percentile class
𝑐 = class interval
𝐶𝑓𝑚 = cumulative frequency preceding the percentile class
Example 2.3.1: The following is the monthly income (in thousand) of 8 persons working in a
factory. Find P30 income value. 10, 14, 36, 25, 15, 21, 29, 17.
Solution:
Arrange the data in an ascending order. (n = 8)
10, 14, 15, 17, 21, 25, 29, 36
N=8
𝑁 + 1 𝑡ℎ
𝑃𝑖 = 𝑖 × ( ) 𝑖𝑡𝑒𝑚
100
270 𝑡ℎ
𝑃30 =( ) 𝑖𝑡𝑒𝑚 = (2.7)𝑡ℎ 𝑖𝑡𝑒𝑚
100
𝑃30 = (2)𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 + 0.7(3𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 − 2𝑡ℎ 𝑣𝑎𝑙𝑢𝑒)
𝑃30 = 14 + 0.7(15 − 14) = 14 + 0.7(1)
𝑃30 = 14 + 0.7 = 14.7
Example 2.3.2:
Calculate P61 for the following data relating to the height of the plants in a garden.
Height 0-5 5-10 10-15 15-20 20-25 25-30
(cm)
No of 18 20 36 40 26 16
plant
Solution:
Class F CF
0–5 18 18
5 - 10 20 38
10 - 15 36 74
15 - 20 40 114
20 - 25 26 140
25 - 30 16 156
𝑖 × 𝑁 𝑡ℎ
𝑃𝑖 = ( ) 𝑖𝑡𝑒𝑚
100
61 × 156 𝑡ℎ
𝑃61 = ( ) 𝑖𝑡𝑒𝑚 = (95.16)𝑡ℎ 𝑖𝑡𝑒𝑚
100
Checking the cumulative frequency column, 95.16th observation fall in 4th class, hence
The 61th percentile lies in the group 15 – 20
156
(61 × 100) − 74
𝑃61 = 15 + ( )×5
40
95.16 − 74
𝑃61 = 15 + ( )×5
40
105.8
𝑃61 = 15 + ( ) = 15 + 2.645
40
𝑃61 = 17.645
NOTE that:
If the sample data collected is deliberately biased and includes data points with specific
characteristics, this can cause the distribution to be Asymmetrical. In an asymmetrical
distribution the two sides will not be mirror images of each other.
For this purpose, we use other two statistical measures that compare the shape to the normal
curve called Skewness and Kurtosis. Skewness and Kurtosis are the two important
characteristics of distribution that are studied in descriptive statistics
a). SKEWNESS:
Skewness can be defined as a statistical measure that describes the lack of symmetry or
asymmetry in the probability distribution of a dataset. It quantifies the degree to which the
data deviates from a perfectly symmetrical distribution, such as a normal (bell-shaped)
distribution. Skewness is a valuable statistical term because it provides insight into the shape
and nature of a dataset’s distribution. For example, understanding whether a dataset is
positively or negatively skewed can be important in various fields, including finance,
economics, and data analysis, as it can impact the interpretation of data and the choice of
statistical techniques.
STA 131: MEASURES OF SHAPE: SKEWNESS AND KURTOSIS
2
BY DR. R. B. AFOLAYAN
Types of Skewness
Positive skewness and negative skewness are two different ways that a dataset’s distribution
can deviate from perfect symmetry (a normal distribution). They describe the direction of the
skew or asymmetry in the data.
In a positively skewed distribution, the tail on the right side (the larger values) is longer than
the tail on the left side (the smaller values). This means that the majority of data points are
concentrated on the left side of the distribution, and there are some extreme values on the
right side. In the case of a positively skewed dataset, Mean > Median > Mode
Examples of positively skewed data include income distribution (where most people earn a
moderate income, but a few earn extremely high incomes), exam scores (where most students
score in a certain range, but a few score exceptionally high), and stock market returns (where
most days have modest returns, but a few days may have very high returns).
In a negatively skewed distribution, the tail on the left side (the smaller values) is longer than
the tail on the right side (the larger values). This implies that most of the data points are
concentrated on the right side of the distribution, with a few extreme values on the left side.
In the case of a negatively skewed dataset, Mean < Median < Mode
For example, the symmetrical and skewed distributions are shown by curves as:
Measurement of Skewness
2. If mode is not defined for a distribution, we cannot find Sk. But empirical
relation between mean, median and mode states that, for a moderately symmetrical
distribution, we have Mean - Mode ≈ 3 (Mean - Median). Hence Karl Pearson's
coefficient of skewness is defined in terms of median as:
3(𝑚𝑒𝑎𝑛 − 𝑚𝑒𝑑𝑖𝑎𝑛)
𝑆𝑘 =
𝑠𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛(𝜎)
1
𝑆=√ ∑(𝑋 − 𝑋̅)2
𝑛−1
𝑆 = √26.79 = 5.18
Step 4. Calculate the mode
It is clear from the data set that 100 is the most frequently occurring value in
the data. Hence, mode of given data is 100.
Pearson’s skewness coefficient
1. With respect to Mean and Median:
3(𝑚𝑒𝑎𝑛 − 𝑚𝑒𝑑𝑖𝑎𝑛) 3[95.3 − 97]
𝑆𝑘 = = = −0.9846
𝑠𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛(𝜎) 5.18
𝑄1 + 𝑄3 − 2𝑄2
𝐵𝑠 =
𝑄3 − 𝑄1
Coefficient of Bowley’s Measure
a) If 𝐵𝑠 = 0, the distribution is perfectly symmetric about the mean (no
skewness).
b) If 𝐵𝑠 < 0, the distribution is negatively skewed (left-skewed), meaning the
tail on the left side of the distribution is longer or heavier.
c) If 𝐵𝑠 > 0, the distribution is positively skewed (right-skewed), indicating
that the tail on the right side of the distribution is longer or heavier.
Solution:
Since N= 9;
Step 1: Calculate the median (Q2)
Arrange the data in ascending order
20, 24, 28, 32, 35, 40, 42, 45, 50.
𝑀𝑒𝑑𝑖𝑎𝑛 = 𝑄2 = 𝑋𝑛+1 = 𝑋5
2
(𝑛 + 1) 10
𝑄1 = = = 2.5𝑡ℎ 𝑣𝑎𝑙𝑢𝑒
4 4
𝑄1 = 2𝑛𝑑 + 0.5(3𝑟𝑑 − 2𝑛𝑑 ) = 24 + 0.5(28 − 24)
𝑄1 = 24 + 2 = 26
𝑄1 = 26
To find Q3, consider the values to the right of the median: 40, 42, 45, 50.
3(𝑛 + 1) 30
𝑄3 = = = 7.5𝑡ℎ 𝑣𝑎𝑙𝑢𝑒
4 4
𝑄3 = 7𝑡ℎ + 0.5(8𝑡ℎ − 7𝑡ℎ) = 42 + 0.5(45 − 42)
𝑄3 = 42 + 1.5 = 43.5
𝑄3 = 43.5
1
𝜎 = √ ∑(𝑋 − 𝑋̅)2
𝑛
𝑤ℎ𝑒𝑟𝑒 𝑁 = ∑ 𝑓
1
𝜎 = √ ∑ 𝑓(𝑋 − 𝑋̅)2
𝑛
̅ )3
∑ 𝑓(𝑋 − 𝑋
𝑀3 =
∑𝑓
𝑤ℎ𝑒𝑟𝑒 𝑡ℎ𝑒 𝑀2 𝑎𝑛𝑑 𝑀3 𝑎𝑟𝑒 𝑠𝑒𝑐𝑜𝑛𝑑 𝑎𝑛𝑑 𝑡ℎ𝑖𝑟𝑑 𝑐𝑒𝑛𝑡𝑟𝑎𝑙 𝑚𝑜𝑚𝑒𝑛𝑡 𝑎𝑏𝑜𝑢𝑡 𝑡ℎ𝑒 𝑚𝑒𝑎𝑛.
Interpretation of Skewness:
• For skewness values between -0.5 and 0.5, the data exhibit approximate
symmetry.
• Skewness values within the range of -1 and -0.5 (negative skewed) or 0.5 and
1(positive skewed) indicate slightly skewed data distributions.
• Data with skewness values less than -1 (negative skewed) or greater than 1
(positive skewed) are considered highly skewed.
KURTOSIS:
Kurtosis is a statistical measure that quantifies the shape of a probability distribution. It
provides information about the tails and peakedness of the distribution compared to a normal
distribution. The measure of kurtosis is very helpful in the selection of an appropriate
average. For example, for normal distribution, mean is most appropriate; for a leptokurtic
distribution, median is most appropriate; and for platykurtic distribution, the quartile range is
most appropriate.
∑ 𝑓(𝑋 − 𝑋̅ )2
𝑠𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛 = 𝜎 = √
∑𝑓
∑ 𝑓(𝑋 − 𝑋̅)4
𝑀4 =
∑𝑓
Karl Pearson classified curves into three types on the basis of the shape of their peaks. These
are Mesokurtic, leptokurtic and platykurtic. These three types of curves are shown in
figure below:
1. Mesokurtic (Kurtosis = 3)
Mesokurtic is the same as the normal distribution, which means kurtosis is near 0. The value
1- The characteristic of a frequency distribution that ascertains its symmetry about the mean is
called skewness. On the other hand, Kurtosis means the relative pointedness of the standard
bell curve, defined by the frequency distribution.
2- Skewness is a measure of the degree of lop-sidedness in the frequency distribution.
Conversely, kurtosis is a measure of degree of tailedness in the frequency distribution.
3- Skewness is an indicator of lack of symmetry, i.e. both left and right sides of the curve are
unequal, with respect to the central point. As against this, kurtosis is a measure of data, that
is either peaked or flat, with respect to the probability distribution.
4- Skewness shows how much and in which direction, the values deviate from the mean? In
contrast, kurtosis explain how tall and sharp the central peak is.
Frequency 5 7 15 28 8 63
Solution:
class Mid-Point 𝐟𝐗 (𝐗 (𝐗 − 𝟐𝟗. 𝟐𝟗)𝟐
interval Frequency(f) (𝑿) − 𝟐𝟗. 𝟐𝟗)
0 - 10 5 5 25 -24.29 590.0041
10 - 20 7 15 105 -14.29 204.2041
20 - 30 15 25 375 -4.29 18.4041
30 - 40 28 35 980 5.71 32.6041
40 - 50 8 45 360 15.71 246.8041
Total ∑ 𝑓𝑋 =63 ------- ∑ 𝑓𝑋 -------- ∑(𝑋 − 𝑋̅)2
= 1845;
= 1092.0205
Calculate the Mean
∑ 𝑓𝑋 1845
𝑀𝑒𝑎𝑛(𝑋̅) = = = 29.29
∑𝑓 63
𝑀3 ∑(𝑋 − 𝑋̅)3
𝑆𝑘𝑒𝑤𝑛𝑒𝑠𝑠 = 3 =
𝜎 (𝑛 − 1) × 𝑆 3
𝑀4 ∑(𝑋 − 𝑋̅)4
𝐾𝑢𝑟𝑡𝑜𝑠𝑖𝑠 = 4 =
𝜎 (𝑛 − 1) × 𝑆 4
Where S is the sample standard deviation and its defined as:
∑(𝑋 − 𝑋̅)2
𝑆= √
𝑛−1