0% found this document useful (0 votes)
16 views74 pages

Data Classification and Frequency Distribution

The document discusses the organization and classification of raw data, emphasizing the importance of structuring data for clarity and comprehension. It outlines characteristics of good classification, types of frequency distributions, and methods for constructing frequency tables, including cumulative frequency distributions. Additionally, it provides guidelines for determining the number of classes and constructing grouped frequency distributions.

Uploaded by

adebayohakee
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
16 views74 pages

Data Classification and Frequency Distribution

The document discusses the organization and classification of raw data, emphasizing the importance of structuring data for clarity and comprehension. It outlines characteristics of good classification, types of frequency distributions, and methods for constructing frequency tables, including cumulative frequency distributions. Additionally, it provides guidelines for determining the number of classes and constructing grouped frequency distributions.

Uploaded by

adebayohakee
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

EDWARD CARES

ORGANIZATION/CLASSIFICATION OF DATA
Measurements or counting gives rise to raw data. Raw data are collected data that have not been
organized numerically. Raw data are difficult to comprehend because it lacks organization,
summarization, which renders it meaningless. Thus, the raw data has to be put in some order
through classification and tabulation so as to reduce its volume and heterogeneity.

Characteristics of a Good Classification

• Comprehensiveness: Classification should cover all the items of the data. In other words, it
should be so comprehensive that it classifies all items in some group or class.

• Clarity: There should be no confusion of the placement of any data item in a group or class.
That is, classification should be absolutely clear.

• Homogeneity: The items within a specific group or class should be similar to each other.

• Suitability: The attribute or characteristic according to which classification is done should


agree with the purpose of classification.

• Stability: A particular kind of investigation should be effective on the same set of


classification.

• Elastic: As the purpose of classification changes, one should be able to change the basis of
classification.

Construction of frequency distribution


When summarizing large masses of raw data, it is often useful to distribute the data into classes,
or categories, and determine the number of individuals belonging to each class, called the class
frequency.

Frequency: Is the number of times a certain value or class of values of the data occurs. The
sum of the frequencies equal to the total observations(n) or sample size n.
Frequency Distribution
A tabular arrangement of data by their classes together with the corresponding class
frequencies is called frequency distribution or frequency table.
The techniques use to organize data depend on the type of variable (Quantitative(numerical) or
Qualitative(categorical)) associated with such data.

TYPES OF FREQUENCY DISTRIBUTION

Two types of frequency distributions that are most often used are the:

i. Categorical frequency distribution


ii. Quantitative frequency distribution

Page 1 of 12
Categorical Frequency Distribution:

A categorical frequency distribution is a table used to organize data that can be placed in specific
categories, such as nominal- or ordinal-level data. The categorical frequency distribution is
done by tallying responses by categories and place the results in tables. This can be
done to construct a summary table to organize the data for a single categorical variable
or construct a contingency table to organize the data from two or more categorical
variables.

a. Summary Table:
A summary table is usually constructed for a single categorical variable. The table present the
tallied responses as frequencies or percentages for each category. The table helps you to see the
differences among the categories by displaying the frequency(amount) or percentage of items in a
set of categories in a separate column.

For example, the data below represents the blood groups of 40 students in a Biostatistics class.
Construct a frequency distribution for the data.

Page 2 of 12
B. Quantitative variable

i. Ungrouped/ discrete frequency distribution


The frequency is constructed for a data based on a single data value for each class. This is used
when each distinct data occurs a number of times.
Example: Given below, are the wing length measurements (to the nearest whole millimeter) of 50
laughing doves.

Page 3 of 12
ii. Grouped Frequency Distribution: The data are organized into groups or
intervals with their corresponding frequencies.
Table 2.0: Height of 100 students in STA 131 class

TERMS ASSOCIATED WITH FREQUENCY DISTRIBUTION

i. Class intervals and class limits


Class Interval: class interval is a range of values into which data is grouped for the
purpose organizing large data set. A class interval is defined by two values:
a. Lower Class Limits
This is the smallest value that belong to a class interval. It is inclusive.
b. UPPER Class Limit
This is the largest value that belong to a class interval.
The end numbers 10-19 are called class limits; the smaller number (10) is the lower limit, and the
larger number (19) is the upper-class limit. A class interval with no upper- or lower-class limit
indicated is called an open class interval. For example, referring to age groups of individuals, the
class interval “ 65 years and above” is an open class interval.

c. Class boundaries:
Class boundaries are defined to eliminate any gaps between the classes (it has one more decimal
place than the data.). Class boundaries are those limits which are determined mathematically to
make an interval of a continuous variable continuous in both directions, and no gap exists between
classes. It’s the actual or real limits of a class interval.
a. The lower extreme point is called lower class boundary
b. The Upper extreme point is called Upper class boundary
Class boundaries are obtained as follows:

Page 4 of 12
1
Lower Class boundary= lower class limit− 2 𝛼
1
Upper Class boundary= upper class limit+ 𝛼
2
Where 𝛼 is the difference between the upper-class limit of any class interval and lower-class
limit of the next class interval.
d. Class mark (xc) or Mid-point of an interval:

1. The class mark is the midpoint of the class interval and is obtained by adding the lower-
class limit and upper-class limit and dividing by 2.
2. The class mark is also called the class midpoint.
3. It is used as representative value of the class interval for the calculation of mean, standard
deviation and other measures.
4. Class mark is the value representing the class interval. It is calculated as:

𝐿𝑜𝑤𝑒𝑟 𝑐𝑙𝑎𝑠𝑠 𝑙𝑖𝑚𝑖𝑡 − 𝑈𝑝𝑝𝑒𝑟 𝑐𝑙𝑎𝑠𝑠 𝑙𝑖𝑚𝑖𝑡


𝑋𝑐 =
2

e. The size, or width, of a class interval


The size, or width, of a class interval is the difference between the lower- and upper-class
boundaries and is also referred to as the class width, class size, or class length. If all class intervals
of a frequency distribution have equal widths, this common width is denoted by c. In such case c is
equal to the difference between two successive lower-class limits or two successive upper-class
limits. For example, the boundaries of the class interval 10-19 is 9.5 - 19.5, the size = 19.5-9.5= 10

Table: Class limit, Class boundary, Class mark, Width, Relative frequency and Percentage
Relative frequency
Class Class Class limits Class Class Class Relative %
Interval Frequency Boundaries Mark Width Freq. Relative
Lower Upper Freq.

15- 19 18 15 19 14.5 19.5 17 5 0.18 18%


20- 24 34 20 24 19.5 24.5 22 5 0.34 34%
25- 29 21 25 29 24.5 29.5 27 5 0.21 21%
30- 34 12 30 34 29.5 34.5 32 5 0.12 12%
35- 39 9 35 39 34.5 39.5 37 5 0.09 9%
40-44 6 40 44 40.5 44.5 42 5 0.06 6%
100 1.00 100%

Class Frequency: The number of observations falling within a class is called its class frequency.
Total Frequency: The sum of all the frequencies is called total frequency.

Relative frequency: It is ratio of the frequency of the class to the total frequency. It’s used to
compare two or more frequency distributions or two or more items in the same frequency
distribution. The relative frequency is not expressed as percentage and its defined as:

Page 5 of 12
𝐹𝑟𝑒𝑞𝑢𝑒𝑛𝑐𝑦 𝑜𝑓 𝑡ℎ𝑒 𝐶𝑙𝑎𝑠𝑠
𝐑𝐞𝐥𝐚𝐭𝐢𝐯𝐞 𝐅𝐫𝐞𝐪𝐮𝐞𝐧𝐜𝐲 =
𝑇𝑜𝑡𝑎𝑙 𝑓𝑟𝑒𝑞𝑢𝑒𝑛𝑐𝑦

Percentage Relative Frequency: This is the ratio of the frequency of a class to the total
frequency expressed as percentage. Its defined as:
𝐹𝑟𝑒𝑞𝑢𝑒𝑛𝑐𝑦 𝑜𝑓 𝑡ℎ𝑒 𝐶𝑙𝑎𝑠𝑠
𝐏𝐞𝐫𝐜𝐞𝐧𝐭𝐚𝐠𝐞 (%) 𝐑𝐞𝐥𝐚𝐭𝐢𝐯𝐞 𝐅𝐫𝐞𝐪𝐮𝐞𝐧𝐜𝐲 = × 100
𝑇𝑜𝑡𝑎𝑙 𝑓𝑟𝑒𝑞𝑢𝑒𝑛𝑐𝑦
Guidelines for Number of Classes
i. There should be between 5 and 20 classes.
ii. The class width should be an odd number. This will guarantee that the class midpoints are
integers instead of decimal.
iii. The classes must be mutually exclusive. This means that no data value can fall into two
different classes.
iv. The classes must be all inclusive or exhaustive. This means that all data values must be
included.
v. The classes must be continuous. There are no gaps in a frequency distribution.
vi. The classes must be equal in width. The exception here is the first or last class. It is possible
to have an "below or " and above" class. This is often used with ages.
Creating a Grouped Frequency Distribution
l. Find the largest and smallest values
2. Compute the Range(R): Maximum - Minimum
3. Select the number of classes(K) desired. This is usually between 5 and 20.
4. Find the class width by dividing the range by the number of classes and rounding up.
Size(Width) = R/ K
a. You must round up, not off. Normally 3.2 would round to be 3, but in rounding up, it
becomes 4. If the range divided by the number of classes gives an integer value (no
remainder), then you can either add one to the number of classes or add one to the class
width.
Sometimes you are instructed to use a certain number of classes.
b. Pick a suitable value less than or equal to minimum value.
c. . Your starting value is the lower limit of the first class. Continue to add the class width to
this lower limit to get the rest of the lower limits.
d. To find the upper limit of the first class, subtract one from the lower limit of the second
class. Then continue to add the class width to this upper limit to find the rest of the upper
limits.

Page 6 of 12
5. Tally the data.
6. Find the frequencies.
7. Find the boundaries by subtracting 0.5 units from the lower limits and adding 0.5
units from the upper limits (if the data are recorded without decimal).

Cumulative Frequency Distribution:


Cumulative frequency corresponding to a class is the sum of all the frequencies up to and including
that class. It is obtained by adding the frequency of that class to all the frequencies of the previous
classes. Cumulative frequencies are of two types:
i. Less than cumulative frequency: The number of observations up to a given value is
called less than Cumulative frequency. Its obtained by adding all the frequencies of the
class that are less than the upper limit of the class.

ii. More than cumulative frequency: The number of observations “greater than” a value is
called more than cumulative frequency. It’s the number of observations that are greater
than the lower limit of a class interval.

Uses of Cumulative Frequency

1. It’s used to find out the number of observations less than or more than any given value.

2. It’s used to find out the number of observations falling between any two specified values of
the variable.

3. It’s used to find median, quartiles and percentiles.

Table: Showing Less than and More than Cumulative frequency

Class Class Cumulative Frequency Class Class

Interval Frequency Less than More than Mark Width

15- 19 18 < 19 18 More than 15 100 17 5

20- 24 34 < 24 52 More than 20 82 22 5

25- 29 21 < 29 73 More than 25 48 27 5

30- 34 12 < 34 85 More than 30 27 32 5

35- 39 9 < 39 94 More than 35 15 37 5

40-44 6 < 44 100 More than 40 6 42 5

100

Page 7 of 12
CONSTRUCTION OF FREQUENCY DISTRIBUTION

The following steps are involved in the construction of a frequency distribution.

(1) Find the range of the data: The range is the difference between the largest and the
smallest values.
(2) Decide the approximate number of classes in which the data are to be grouped. Where
the number of classes to used is not given, the number of classes can be estimated using
H.A. Sturge’s formula given as:
K = 1 + 3.322log N
Where K= Number of classes and N = total number of observations.
(3) Determine the approximate class size: The size (width)of class interval is obtained by
dividing the range of data by the number of classes and is denoted by W (class interval
width(size))
Size/ width= Range/ K
In the case of fractional results, the next higher whole number is taken as the size of the
class interval.

(4) Decide the starting point: The lower-class limits or class boundary should cover the
smallest value in the raw data. Usually class intervals of multiple of 5 are commonly used.
(5) Determine the remaining class limits (boundary): When the lowest class boundary has
been decided, you can compute the upper-class boundary by adding the class interval size
to the lower-class boundary. The remaining lower- and upper-class limits may be
determined by adding the class interval size repeatedly till the largest value of the data is
observed in the class.
(6) Distribute the data into respective classes: All the observations are divided into
respective classes by using the tally bar (tally mark) method, which is suitable for
tabulating the observations into respective classes. The number of tally bars is counted to
get the frequency against each class. The frequency of all the classes is noted to get the
grouped data or frequency distribution of the data. The total of the frequency columns
must be equal to the number of observations.

Page 8 of 12
2. Number of Classes = 1+3.322logN
Number of Classes = 1+3.322log57
Number of Classes = 1+3.322(1.75587) = 6.833
Approximately 7 class intervals

3. Class Interval Size (W) = Range /No. of Classes = 67 / 7 = 9.57 or 10

Effect of grouping:
As a result of grouping, it is possible to detect a pattern in the figures but grouping results in the
loss of information i.e. calculations made from a grouped frequency distribution can never be
exact, and consequently excessive accuracy can only result in spurious accuracy.

The reasons for constructing a frequency distribution are:


1. To organize the data in a meaningful, intelligible way.
2. To enable the reader to determine the nature and shape of the distribution.
3. To facilitate computational procedures for measures of average and spread.
4. To enable the researcher to draw charts and graphs for the presentation of data.
5. To enable the reader to make comparisons among different data sets.

Page 9 of 12
Construction of Frequency Distribution, Relative frequency and Cumulative Relative
Frequency

Example 2:

The following data represents the percent change in tuition levels at public, four-year colleges
(inflation adjusted) from 2008 to 2013 (Weismann, 2013). Create a frequency distribution,
histogram, and ogive for the data.

19.5 40.8 57.0 15.1 17.4 5.2 13.0 15.6 51.5 15.6 14.5 22.4 19.5 31.3 21.7 27.0

13.1 26.8 24.3 38.0 21.1 9.3 46.7 14.5 78.4 67.3 21.1 22.4 5.3 17.3 17.5 36.6

72.0 63.2 15.1 2.2 17.5 36.7 2.8 16.2 20.5 17.8 30.1 63.6 17.8 23.2 25.3 21.4

28.5 9.4;

Solution:

1. Find the range:


largest value - smallest value = 78.4 −2.2 =76.2

2. Pick the number of classes: Since there are 50 data points, Let’s use 8.

3. Find the class width:


width = range/ 8= 76.2/8≈9.525
Since the data has one decimal place, then the class width should round to one decimal
place. Make sure you round up.
width =9.6

4. Find the class limits:


2.2+9.6=11.8; 11.8+9.6=21.4; 21.4+9.6=31.0;

5. Find the class boundaries:


Since the data has one decimal place, the class boundaries should have two decimal places,
so subtract 0.05 from the lower-class limit to get the class boundaries. Add 0.05 to the
upper-class limit for the last class’s boundary.
2.2−0.05=2.15;11.8−0.05=11.75;21.4−0.05=21.35

Every value in the data should fall into exactly one of the classes. No data values should fall right
on the boundary of two classes.

6. Find the class midpoints:


midpoint = lower limit + upper limit /2
(2.2+11.7)/2=6.95; (11.8+21.3)/2=16.55

7. Tally and find the frequency of the data:

Page 10 of 12
Table 2.2. Frequency Distribution for Tuition Levels at Public, Four-Year Colleges

Class Class Class


Tally F RF CF
Limits Boundaries Midpoint

2.2- 11.7 2.15- 11.75 6.95 |||||||||| 6 0.12 6

11.8- 21.3 11.75- 21.35 16.55 |||||||||||||||||||||||||||||||| 20 0.40 26

21.4- 30.9 21.35- 30.95 26.15 |||||||||||||||||| 11 0.22 37

31.0- 45.0 30.95- 40.55 35.75 |||||||| 4 0.08 41

40.6- 50.1 40.55- 50.15 45.35 |||| 2 0.04 43

50.2- 59.7 50.15- 59.75 54.95 |||| 2 0.04 45

59.8- 69.3 59.75- 69.35 64.55 |||||| 3 0.06 48

69.4- 78.9 69.35- 78.95 74.15 |||| 2 0.04 50

RF= Relative Frequency and CF=Cumulative Frequency

Page 11 of 12
Tutorial Questions

1. Construct a frequency distribution with the suitable class interval size for
marks obtained by a class of 50 students as given below:

23, 50, 38, 42, 63, 75, 12, 33, 26, 39, 35, 47, 43, 52, 56, 59, 64, 77, 15, 21, 51, 54,
72, 68, 36, 65, 52, 60, 27, 34, 47, 48, 55, 58, 59, 62, 51, 48, 50, 41, 57, 65, 54, 43,
56, 44, 30, 46, 67, 53

b. Create the column of class boundaries, Class marks, %Relative frequency and
Cumulative frequency.
2. The following is the distribution of ages of new employees at a factory

a. Obtain the class boundaries and class marks of the class intervals
b. What is the upper-class limit of the class 30-39?
c. What is the lower-class limit of the class 50-59?
d. What is the class mark of the class 40-49?
e. What is the class width of the class 40-49?
f. What is the lower-class boundary of the class 30-39?
3. A medical research team studied the ages of patients who had strokes caused by stress. The
ages of 34 patients who suffered stress strokes were as follows.
29 30 36 41 45 50 57 61 28 50 36 58

60 38 36 47 40 32 58 46 61 40 55 32

61 56 45 46 62 36 38 40 50 27

Use 8 classes beginning with a lower-class limit of 25

i. Construct a frequency distribution for these ages.


ii. Create a column for Class boundaries, class mark and cumulative frequency, relative
frequency and cumulative relative frequency.
iii. Draw cumulative frequency curve
iv. Draw the histogram and frequency polygon

Page 12 of 12
3.0 GRAPHICAL PRESENTATION OF DATA

. The representation of quantitative data suitably through charts and diagrams is known as
Graphical Representation of Data (Statistical Information). Graphs are used for presenting
statistical data in an attractive way. They enable us to visualize the whole meaning of
complex data at a glance. The main object of diagrammatic representation is to emphasis
the relative position of different subdivisions and not simply to record details.
Advantages of Graphical Representation:
i. It is easily understood by all.
ii. The data can be presented in a more attractive form.
iii. It shows the trend and tendency of values of the variable.
iv. Diagrammatic representations are useful to detect mistakes at the time of data
computations.
v. It shows relationship between two or more sets of figures.
vi. It has the universal applicability.
vii. It is helpful in assimilating the data readily and quickly.
Disadvantages of Graphical Representation:
i. It does not show details or all the facts.
ii. Graphical representation can reveal only the approximate position.
iii. It takes a lot of time to prepare of graph.
The Different Types of Diagrammatic Representation:
There are various types of graphs in the form of charts and diagrams. Some of them are:
1. Line Diagram or Graph.
2. Bar Chart
3. Pie Chart
4. Stem plot
5. Histogram.
6. Frequency polygon.
7. Ogives (cumulative frequency polygon).
3.1 Modes of Graphical Representation of Data:
The data in the form of raw scores is known as ungrouped data and when It is organized into
frequency distribution then it is referred to as grouped data.

Separate modes and methods are used to represent these two types of data grouped.

A. Graphical Representation of Ungrouped/Categorical Data:


For the ungrouped data, the following graphical presentations are used:
1. Line graphs
2. Bar Chart
3. Pie Chart
4. Pictograms
5. Box Plot

Page 1 of 21
PIE CHART
A pie chart is an angular representation of a statistical data with several sub-divisions in a
circle. It is a graph that shows the relative frequency distribution of a nominal variable. Each
component is expressed as a percentage of total value. The size of the angle shows their
relative frequency.
This type of graph can be a good choice when you want to emphasize that one variable is
especially frequent or infrequent, or you want to present the overall composition of a variable.
A disadvantage of pie charts is that it’s difficult to see small differences between frequencies.
As a result, it’s also not a good option if you want to compare the frequencies of different
values.
Example 1:

A campus press polled a sample of 280 undergraduate students in order to study student attitude
toward a proposed change in the dormitory regulations. The following tables shows the response

Table 3.0 : Frequency distribution of responses to dormitor regulations

Draw a Pie chart to show the student responses.


Solution:
Table 3.1: Calculation for Degrees and Percentages
Group Frequency Degrees Percentage
Support 152 152 152
× 3600 = 1950 × 100 = 54%
280 280
Neutral 77 77 77
× 3600 = 990 × 100 = 28%
280 280

Oppose 51 51 51
× 3600 = 660 × 100 = 18%
280 280
Total 280 3600 100%

Page 2 of 21
Figure 1.0-: Pie Chart Showing the distributions of Students’ responses
BAR CHART
A bar chart is a graph that shows the frequency or relative frequency distribution of a
categorical variable (nominal or ordinal).
The y-axis of the bars shows the frequencies or relative frequencies, and the x-axis shows the
values. Each value is represented by either horizontal or vertical bars, and the length or height
of the bar shows the frequency of the value.
The following types of Bar chart are commonly used:
1. Simple Bar Chart
2. Multiple/Clustered Bar Chart and
3. Component/Stacked Bar Chart
A bar chart is a good choice when you want to compare the frequencies of different values or
groups. Some bar graphs present bars clustered in groups of more than one (Multiple bar
graphs), and others show the bars divided into subparts to show cumulative effect (stacked or
component bar graphs). Bar graphs are especially useful when categorical data is being used.
It’s much easier to compare the heights of bars than the angles of pie chart slices.
SIMPLE BAR GRAPH
This graph is used to show a single characteristic of group, categories or years.
It consists of a number of equally spaced vertical bars of uniform width originating from a
horizontal axis and is shaded. The bars are usually arranged according to relative magnitude
of bars. The limitation of simple bar chart is that only one variable can be represented on it.
Example 2:
In a study of 100 women, the numbers show below indicate the major reason why each
surveyed worked outside her home. Construct a bar graph for the data.

Page 3 of 21
Table 3.2: Frequency distribution of why women work outside the home.
Reason No. of Women
A: To support self/family 62
B: For extra money 18
C: For something different to do 12
D: others 8
Frequency distribution of why women work outside the home.
60
50
40
No of Women

30
20
10
0

A B C D

Reasons

Example 3: Create simple bar Chart of the data below represent the revenue(dollars)
collected by a State for the months of April to August:

Amounts(million$) 7 12 28 3 41

Months March April May June July

Bar Chart Showing Distribution of Federal Allocation to a State


40
30
Allocation(dollars)

20
10
0

March April May June July

Months

Figure 1.1: A Bar Chart of Why Women Work Outside the Home
CLUSTERED AND STACKED BAR CHART
CLUSTERED / MULTIPLE BAR CHART: Clustered Bar Chart (also known as Grouped
bar chart or Multiple bar chart) is great for displaying and comparing multiple sets of data
over the same categories. In this case numerical values of major categories are arranged in
ascending or descending order so that categories can be readily distinguished.

Page 4 of 21
For example:
1. Sales revenue of various departments of the company over several years
2. Prices of different commodities at various quarter of a given year.
Example 3:
1. The table below shows the distribution of students offering STA131 by Departments and
gender.
Table 3.3: Distribution of students offering STA131 Course
Gender Departments
STA Maths Engineering Tele Info Total
Com. Tech
Male 40 35 110 40 75 300
Female 30 20 60 35 55 200
Total 0 55 170 75 130 500

Distribution of students offering ST A131 course

Male
100

Female
80
No of Students

60
40
20
0

STAT MATHS ENG. TElCOM [Link]

Departments

Figure 1.2: Multiple or Clustered Bar Chart


Figure 1.2 shows the distributions of Male and Female(gender) in each department. The
height of each bar is the frequency of Male and female in each department. This type of graph
is used for two categorical variables. One categorical variable is presented on the x-axis why
the second category represent the bars and the corresponding cell values represent the height
of the bar.
A STACKED / COMPONENT BAR CHART:
Is a type of bar graph that represents the proportional contribution of individual data points in
comparison to a total. The height or length of each bar represents how much each group
contributes to the total. Here, the height of the bar represents a particular variable group total
and cell values for the second variable are used to show how each group of the second
variable contributes to the total.

Page 5 of 21
i. Each bar in the component bar is subdivided into several component parts.
ii. A single bar represents the aggregate value whereas the component parts represent
the component values of the aggregate value.
iii. It shows the relationship between the different components and difference in main
bar.
iv. Different shades or colours are used to distinguish the various components.
Example 4: Using the data in table 3.3, draw a stacked or component bar chart.
Gender Departments
STA Maths Engineering Tele Info Total
Com. Tech
Male 40 35 110 40 75 300
Female 30 20 60 35 55 200
Total 70 55 170 75 130 500
Distribution of students offering ST A131 course
100 120 140
No of Students

80
60
40
20
0

STAT MATHS ENG. TElCOM [Link]

Departments

Distribution of students offering STA131 course


140

Male
Female
120
100
No of Students

80
60
40
20
0

STAT MATHS ENG. TElCOM [Link]

Departments

Figure 1.3: Stacked or Component Bar Chart


From figure 1.3, the height of the bar equals to the total students in each department. Each bar
is divided into gender components using the corresponding cell values. For instance, Statistics
as a total of seventy (70) students classified as male (40) and Female (30).

Page 6 of 21
Class work:
Example 5:
A researcher asked her class to pick who would win in a battle of superhero. Below is a
frequency table and charts of the results:

Table 3.4: Calculation for Degrees and Percentages


Group Frequency Degrees Percentage
Batman 52 52 52
× 3600 = 146.30 × 100% = 41%
128 128
Captain 25 25 25
× 3600 = 70.30 × 100% = 19%
128 128
American
34
Iron Man 34 × 3600 = 95.60 34/128 × 100% = 27%
128
Superman 17 17 17
× 3600 = 47.80 × 100% = 13%
128 128
Total 128 3600 100%
A pie chart and bar chart of these results are shown below:

Figure 1.4: Bar chart showing the proportion of votes for Superhero Battle Contestants

Page 7 of 21
Figure 1.5: Pie chart showing the proportion of votes for Superhero Battle Contestants.

STEM- AND- LEAVE PLOT


One simple graph, the stem-and-leaf graph or stem plot, comes from the field of exploratory
data analysis. It is a good choice when the data sets are small. To create the plot, divide each
observation of data into a stem and a leaf. The leaf consists of a final significant digit. For
example, 23 has stem two and leaf three. The number 432 has stem 43 and leaf two.
Likewise, the number 5,432 has stem 543 and leaf two. The decimal 9.3 has stem nine and
leaf three. Write the stems in a vertical line from smallest to largest. Draw a vertical line to
the right of the stems. Then write the leaves in increasing order next to their corresponding
stem. The advantage in a stem-and-leaf plot is that all values are listed, unlike a histogram,
which gives classes of data values

Example 5: The data below represent the scores for the first exam were as follows (smallest
to largest):

33; 42; 49; 49; 53; 55; 55; 61; 63; 67; 68; 68; 69; 69; 72; 73; 74; 78; 80; 83; 88; 88; 88; 90;
92; 94; 94; 94; 94; 96; 100

Table 3.5: Stem and Leaf Plot


Stem Leaf
3 3
4 299
5 355
6 1378899
7 2348
8 03888
9 0244446
10 0

The stem-plot shows that most scores fell in the 60s, 70s, 80s, and 90s. Eight out of the 31
scores or approximately 26% (8/31) were in the 90s or 100, a fairly high number of As.

Page 8 of 21
Class Work

1. The data are the distances (in kilometres) from a home to local supermarkets. 1.1; 2.3;
3.3; 4.5; 4.7; 2.7; 4.5; 4.8; 6.7; 12.3;1.5; 2.5; 3.2; 4.2; 5.5; 3.5; 3.8; 6.5; 3.3; 4.0; 5.6;

a. Create a stem-plot using the data and identify any outliers.


b. Do the data seem to have any concentration of values?

Box and Whisker Plot

Box Plot: Box plot (Box and Whisker diagram or plot) is a convenient way of graphically
depicting groups of numerical data through their five number summaries: minimum, lower
quartile(Q1), median(Q2), Upper quartile(Q3) and maximum observation with outliers plotted
individually.
Characteristics of Box Plot:
i. A central box spans the quartile
ii. A line in the box mark the median
iii. Observations more than 1.5 X IQR (inter quartile range) outside the central box
plotted individually as possible outlier.
iv. Lines called Whiskers, extend from the box and to the smallest and largest
observations that are not outlier.
v. Box plots can be drawn either horizontally or vertically.
vi. For symmetric distribution, box plots are symmetric but the symmetric box plot
does not imply symmetric distribution.
vii. The upper part of the box i.e. differences between Q3 and Median is bigger than
the lower part (difference between Median and Q1) in case of positively skewed
distribution and smaller in case of negatively skewed.
Uses of Box Plot:
1. The box plots are most useful for comparing distributions.
2. It is a quick way of examining one or more sets of data graphically.
3. The spacings between the different parts of the box indicate the degrees of dispersions
and skewness in the data and identify outliers.
Inter Quartile Range (IQR):
It is a measure of the spread based on quartiles. The range from third quartile(Q3) to first
quartile(Q1) is called interquartile range.
IQR= Q3 – Q1
Semi-Interquartile Range or Quartile Deviation
The half of the difference between the third quartile(Q3) and first quartile(Q1) is called
Semi-Interquartile Range or Quartile Deviation
OUTLIERS:
An outlier is defined as an observation which falls more than 1.5 * IQR (called step)
above Q3 or below Q1

Page 9 of 21
a. If an observation falls more than 3* IQR above Q3 or below Q1, then it is known as
extreme outlier.
b. The observation falling between 1.5 * IQR and 3* IQR above Q3 or below Q1, it is
known as suspect outlier.
c. Inner-fences and Outer fences
i. The value [Q1- 1.5*IQR, Q3 + 1.5* IQR] are known inner-fences
ii. The value [Q1- 3*IQR, Q3 + 3* IQR] are known outer-fences

The centre of distribution A is the lowest of the three distributions (median is 0.11). The
distribution is positively skewed, because the whisker and half-box are longer on the right
side of the median than on the left side.
Distribution B is approximately symmetric, because both half-boxes are almost the same
length (0.11 on the left side and 0.10 on the right side). It’s the most concentrated distribution
because the interquartile range is 0.21, compared to 0.30 for distribution A and 0.26 for
distribution C.

Page 10 of 21
The centre of distribution C is the highest of the three distributions (median is 0.88). The
distribution C is negatively skewed because the whisker and half-box are longer on the left
side of the median than on the right side.
All three distributions include potential outliers. Let’s take distribution A, for example. The
interquartile range is Q3 - Q1 = 0.32 – 0.02 = 0.30. According to the definition used by the
function in R software, all values higher than Q3 + 1.5 x (Q3 - Q1) = 0.32 + 1.5 x 0.30 = 0.77
are outside the right whisker and indicated by a circle. There are two potential outliers in
distribution A.
Building A Box and Whisker Plot

Example 1:

Given the information below, draw a box and whisker plot.

Minimum 82
Lower quartile 94
Median 95
Upper quartile 102
Maximum 110

Example2:

Draw a box and whisker plot for this sample:

5 7 1 9 11 22 15

Page 11 of 21
Example 1: a simple box and whisker plot
Suppose you have the math test results for a class of 15 students. Here are the results:

91 95 54 69 80 85 88 73 71 70 66 90 86 84 73

Step 1: Order the data points from least to greatest.


54 66 69 70 71 73 73 80 84 85 86 88 90 91 95

Step 2: Find the median of the data:

This is an odd set of data – you have 15 data points. It means the middle point is 80 as there
are 7 data points above it and 7 numbers below.

Step 3: Find the middle points of the two halves divided by the median (find the upper and
lower quartiles).

Step 4: Find the extreme values.


This is the easiest part. You need to find the largest and smallest data values.
Extreme values = 54 and 95.

Page 12 of 21
So, we can determine that the five-number summary for the class of students is:
54, 70, 80, 88, 95.
Now we are absolutely ready to draw our box and whisker plot.

As you see, the plot is divided into four groups: a lower whisker, a lower box half, an upper
box half, and an upper whisker. Each of those groups shows 25% of the data because we have
an equal amount of data in each group.

LINE GRAPH

Another type of graph that is useful for specific data values is a line graph. In the particular
line graph shown in fig. 1.5, the x-axis (horizontal axis) consists of data values and the y-axis
(vertical axis) consists of frequency points. The frequency points are connected using line
segments. Line graph are appropriate for ungrouped (quantitative) frequency distribution.

Example: In a survey, 40 mothers were asked how many numbers of birth each had. The
results are shown in Table. Construct a line graph.

No of births Frequency
0 2
1 5 Table 3.5: A frequency distribution
2 8 table showing the distribution of births
3 14 by 40 randomly selected mothers in
4 7 Oyun, LGA of Kwara State.
5 4
Total 40

Number of Births
Fig. 1.5: A line graph showing the number of Births for 40 randomly selected mothers in
Oyun, LGA of Kwara State.

Page 13 of 21
Tutorial Question:
In a survey, 40 people were asked how many times per year they had their car in the
mechanic shop for repairs. Construct a line graph of the data.
Number of times in Mechanic Shop Frequency
0 7
1 10
2 14
3 9

B. GRAPHICAL REPRESENTATION OF GROUPED DATA


For the grouped data, the following graphical presentations are used:
1. Histogram
2. Frequency Polygon
3. Cumulative Frequency Curve (ogive)
HISTOGRAM
A histogram is an effective visual summary of several important characteristics of a variable.
It is a graph that shows the frequency or relative frequency distribution of a quantitative
variable. It looks similar to a bar chart. The continuous variable is grouped into interval
classes, just like a grouped frequency table. At a glance, you can see a variable’s central
tendency and variability, as well as what probability distribution it appears to follow, such
as a normal, Poisson, or uniform distribution.
Steps to Construct a Histogram
Step 1: Use the frequency distribution table and class boundaries.
Step 2: Label the x-axis with the class boundaries of each class. Label the y-axis with
frequencies.

Step 3: Draw rectangles with bases as class boundary and corresponding frequencies as
heights.

Step 4 The height of each rectangle is proportional to the corresponding class frequency if the
intervals are equal. The area of every individual rectangle is proportional to the corresponding
class frequency if the intervals are unequal.
.

Page 14 of 21
Example of Histogram

Although bar charts and histograms are similar, there are important differences:
S/No Differences Bar chart Histogram
1 Type of Categorical Quantitative
variable
2 Value Ungrouped (values) Grouped (interval classes)
grouping
3 Bar spacing Space can be between bars Space is not allow between
bars
4 Bar order Can be in any order Can only be ordered from
lowest to highest

FREQUENCY POLYGON

Frequency Polygon: A frequency polygon is a straight-line graph used to represent the


frequency of various data. It is an area diagram represented in the form of curve obtained by
joining the middle points of the tops of the rectangles in a histogram or joining the mid-points of
class intervals at the height of frequencies by straight lines.

Steps to Construct a Frequency Polygon


Step 1: Use the frequency distribution table and midpoints(class Mark) values.
Step 2: Label the x-axis with the midpoints of each class. Label the y-axis with frequencies.
Step 3: Place a point corresponding to each class and its frequency.
Step 4: Connect all the points from step 3 with straight lines. Make sure to include one class
interval above and below your highest and lowest classes so that the frequency polygon touches
the x-axis on both sides.
Class Work:
Example: In a city, the weekly observations made in a study on the cost of a living index are given in
the following table: Draw a frequency polygon for the data below with a histogram.

Page 15 of 21
Solution:
To plot a frequency polygon with a histogram, we need to follow these steps to construct a
histogram:
• The cost of living index is represented on the x-axis.
• The number of weeks is represented on the y-axis.
• Now rectangular bars of widths equal to the class- size and the length of the bars
corresponding to a frequency of the class interval are drawn.
To calculate the midpoint, we use the formula:
Class mark = (Upper Limit + Lower Limit) / 2 So,
Class mark = (150 + 140)/2 = 145, (160 + 150)/2 = 155 and so on

While plotting the graph, we also mark the before and after class mark as well. In this case, the
before is 135 and the after is 205. ABCDEFGH represents the given data graphically in form of
frequency polygon i.e. those are the midpoints. Hence, the frequency polygons graph is given below:

Page 16 of 21
Difference between Frequency Polygons and Histogram:
Even though a frequency polygon graph is similar to a histogram and can be plotted with or
without a histogram, the two graphs are yet different from each other. The two graphs have
their own unique properties that show the difference visually. The differences are:

CUMULATIVE FREQUENCY CURVE (OGIVE)


The graph for cumulative frequency is called an ogive (o-jive). The cumulative frequencies
are plotted against the corresponding upper-class boundaries and the successive points are
joined by straight lines, the diagram or curve obtained is known as Ogive or Cumulative
frequency polygon.
To create an ogive, follow the steps below:
Steps to Construct a Cumulative frequency Polygon or Ogive
Step 1: Use the cumulative frequency distribution table and Class boundary.
Step 2: Label the x-axis with the upper-class boundary of each class. Label the y-axis with
Cumulative frequencies.

Page 17 of 21
Step 3: Against each upper-class boundary, plot the cumulative frequencies.
Step 4: Connect all the points from step 3 with straight lines
Step 5: Make sure you include the point with the lowest class boundary and the 0 cumulative
frequency.
Class Work:

Example: Draw an ogive for the following frequency distribution of test marks for a group of
50 students by less than method.

Marks 20-24 25-29 30-34 35-39 40-44 45-49 50-54 55-59 60-64 65-69
Frequency 1 2 4 8 11 9 7 4 3 1

Solution:

Class frequency Less Than Cumulative More than Cumulative


interval frequency frequency

20-24 1 Less than 24.5 1 19.5 or more 50

25-29 2 Less than 29.5 3 24.5 or More 49

30-34 4 Less than 34.5 7 29.5 or More 47

35-39 8 Less than 39.5 15 34.5 or More 43

40-44 11 Less than 44.5 26 39.5 or More 35

45-49 9 Less than 49.5 35 44.5 or More 24

50-54 7 Less than 54.5 42 49.5 or More 15

55-59 4 Less than 59.5 46 54.5 or More 8

60-64 3 Less than 64.5 49 59.5 or More 4

65-69 1 Less than 69.5 50 64.5 or More 1

Page 18 of 21
Fig. 1.6: Cumulative frequency polygon (Ogive) for test scores

Uses of Ogive
1. It is used to find the median, quartiles, decile and percentiles.
2. It is also useful in finding the cumulative frequency corresponding to a given
value of the variable.
3. To find out the number of observations which are expected to lie between two
given values.
General Rules for Graphical Representation of Data
There are certain rules to effectively present the information in the graphical representation.
They are:
1- Suitable Title: Make sure that the appropriate title is given to the graph which indicates
the subject of the presentation.
2- Measurement Unit: Mention the measurement unit in the graph.
3- Proper Scale: To represent the data in an accurate manner, choose a proper scale.
4- Index: Index the appropriate colors, shades, lines, and design in the graphs for better
understanding.
5- Data Sources: Include the source of information wherever it is necessary at the bottom
of the graph.
6- Keep it Simple: Construct a graph in an easy way that everyone can understand.
7- Neat: Choose the correct size, fonts, colors. etc in such a way that the graph should be a
visual aid for the presentation of information

Tutorial Questions:
Question 1

The data below were the ages of 30 randomly selected staff in a particular university:

32; 54; 32; 42; 33; 34; 50; 38; 40; 42; 53; 56; 43; 44; 51; 52; 46; 47; 60; 47; 48; 52;
57; 48; 50; 52; 48; 49; 57; 61

a. Construct a stem plot for the data

Page 19 of 21
Question 2

Construct both Pie and bar graph to represent the data


Question 3

2. The table below shows the frequency distribution of test marks for 120
students.
Marks
1-10 11-20 21-30 31-40 41-50 51-60 61-70 71-80 81-90 91-100
(%)

Number
of 1 6 8 15 17 24 22 15 9 3
students
Draw the following graph:
i. Frequency polygon
ii. Histogram
iii. Less than and more than cumulative frequency polygon (ogive)

Tutorial Questions on Box Plot

Question 4

a. Given the information below, draw a box and whisker plot.

Min.=10, Q1=14, Q2=16, Q3=20 and Max. =29

b. Draw a box and whisker plot for these samples:


a. 17 22 18 33 14 36 39 41 25 31 18 19 16 21 21
b. 22 12 10 15 13 25 12 24 31 28

c. What are the median and the semi-interquartile range of the following
sample?

Page 20 of 21
Question 5

The following shows the distribution of scores obtained by SS2 students in mathematics
mock exam.
Class 21-30 31-40 41-50 51-60 61-70 71-80 81-90 91-100
interval
Frequency 2 6 7 10 11 9 4 1

1. Construct a cumulative frequency table for the distribution.


2. Draw a cumulative frequency curve for the distribution.
3. From your graph, estimate the:
4. median
5. interquartile range
6. pass mark, if 70% of the students passed.

Question 6

1. The table contains data from a study of daily study time for 40 students from
Statistics 101. Construct an ogive from the data.
Minutes on
Homework Number ofstudents Relativefrequency Cumulativefrequency
0-15 2 0.05 2
16-30 4 0.10 6
31-45 8 0.20 14
46-60 18 0.45 32
61-75 4 0.10 36
76-90 4 0.10 40

2. An annual survey sent to retail store managers contained the question "Did your store suffer
any losses due to employee theft?" The responses are summarized in the table for two years,
2000 and 2005. Construct a multiple bar graph of the data, then describe any trends.

Page 21 of 21
MEASURES OF LOCATION:
Mean or Average: The arithmetic mean is the sum of all observations divided by the number of
observations. The mean of a set of n observations 𝑥1 , 𝑥2 , … . . , 𝑥𝑛 is defined as
𝒏
𝟏 𝟏
̅ = (𝒙𝟏 + 𝒙𝟐 + ⋯ … . . +𝒙𝒏 ) = ∑ 𝒙𝒊
𝑿
𝒏 𝒏
𝒊=𝟏

For grouped frequency data, the mean is defined as:

(𝒇𝟏 𝒙𝟏 + 𝒇𝟐 𝒙𝟐 + ⋯ … . . +𝒇𝒏 𝒙𝒏 ) ∑ 𝒇𝒊 𝒙𝒊
̅=
𝑿 =
𝑓1 + 𝑓2 , + ⋯ … . +𝑓𝑛 ∑ 𝒇𝒊

where 𝑥𝑖 are the mid-points or class marks

Merits and Demerits

Merits

1. The mean is simple and easily understood , most popular measures of central tendency.
2. Every observation in the data is used to compute the mean.
3. The mean is unique,there is only one mean for any given data sets.
4. The mean is used for further statistical analysis.

Demerits:

1. The mean is not defined for categorical data.


2. The mean is affected by some few extreme values(outliers).

Example 1: The following are weights in pounds of 10 randomly selected children at a day-care
center: 68 ,63, 42, 27, 30, 36 ,28, 32 ,79 ,27. Compute the sample mean

Solution: n=10
1 1
𝑋̅=𝑛 (𝑥1 + 𝑥2 + ⋯ … . . +𝑥𝑛 ) = 𝑛 ∑𝑛𝑖=1 𝑥𝑖
1
𝑋̅ = (68 + 63 + 42 + 27 + ⋯ … . . +27)
10
𝑛
1
𝑋̅ = ∑ 432 = 43.2
10
𝑖=1
The average or mean weigths of the 10 children is 43.2 pounds

1
B. Arithmetic mean for ungrouped/ grouped frequency data, the mean is defined as:

Let 𝒙𝟏 , 𝒙𝟐 … … . . , 𝒙𝒏 be n observations with frequencies 𝑓1 , 𝑓2 , … … . , 𝑓𝑛 . Then the arithmetic


mean is defined as:

(𝒇𝟏 𝒙𝟏 + 𝒇𝟐 𝒙𝟐 + ⋯ … . . +𝒇𝒏 𝒙𝒏 ) ∑ 𝒇𝒊 𝒙𝒊 ∑ 𝒇𝒙
̅=
𝑿 = = ;
𝑓1 + 𝑓2 , + ⋯ … . +𝑓𝑛 ∑ 𝒇𝒊 𝑵

where 𝑁 = ∑ 𝒇𝒊 𝑎𝑛𝑑 𝑥𝑖 are the mid-points or class marks for grouped frequency

STEPS:

1. Obtain the class mark(x) for grouped frequency data


2. Multiply each value(x) or class mark(x) with its corresponding frequencies for ungrouped
or grouped frequency data respectively and sum them up to get ∑ 𝒇𝒙
3. Divide the total value ∑ 𝒇𝒙 by total frequency ∑ 𝒇𝒊

Example 2: Find the Arithmetic mean from the frquency table.

Marks 30 40 50 60 70 80 90
No. of 15 20 10 15 20 15 5
students
Solution: this is an Ungrouped frequency data

Marks(x) No. of 𝒇𝒙
students(f)
30 15 450
40 20 800
50 10 500
60 15 900
70 20 1400
80 15 1200
90 5 450
Total ∑ 𝒇𝒊 = 𝟏𝟎𝟎 ∑ 𝒇𝒙 =5,700

∑ 𝒇𝒊 𝒙𝒊 𝟓𝟕𝟎𝟎
̅ ==
𝑿 = = 𝟓𝟕
∑ 𝒇𝒊 𝟏𝟎𝟎

2
Example 3: Compute the mean of the grouped frequency data given below:

Weight 10-19 20-29 30-39 40-49 50-59 60-69 70-79 Total


interval
Frequency(f) 5 19 10 13 4 4 2 57
Solution:

Weight Interval Frequency(f) Mark 𝒇𝒙


class(x)
10-19 5 14.5 72.5
20-29 19 24.5 465.5
30-39 10 34.5 345.0
40-49 13 44.5 578.5
50-59 4 54.5 218.0
60-69 4 64.5 258.0
70-79 2 74.5 149.0
Total ∑ 𝒇𝒊 =57 ∑ 𝒇𝒙 = 𝟐𝟎𝟖𝟔. 𝟓

Solution:
Class mark(x) is the mid –point of each interval defined as the average of the lower and upper
limits of a given class interval

For interval : 10-19, the mid-point= (10 +19)/2= 14.5

For grouped frequency data, the mean is defined as;


(𝑓 𝑥 +𝑓 𝑥 +⋯…..+𝑓𝑛 𝑥𝑛 ) ∑ 𝑓𝑖 𝑥𝑖
𝑋̅ = = 1 1 2 2 =∑
𝑓1 + 𝑓2 ,+ ⋯….+𝑓𝑛 𝑓𝑖

∑ 𝑓𝑖 𝑥𝑖 = 2086.5 and ∑ 𝒇𝒊 = 𝟓𝟕

∑ 𝒇𝒊 𝒙𝒊 𝟐𝟎𝟖𝟔. 𝟓
𝑋̅ = = = 𝟑𝟔. 𝟔𝟏𝒍𝒃
∑ 𝒇𝒊 𝟓𝟕

C. Arithmetic Mean Using Assumed Mean Method


The method is applied when the frquency or value of the variables are quite large. The assumed
mean or short cut methods is applied to ease the rigors of find the product of frequency and the
value of the variable.

Step;

1. Assumed mean is taken to be the middle value of x (or class mark(x)) which correspond
to the middle value of the frequency distribution.

3
a. In case of n observations data, the mean is computed as:
𝒏
𝟏
̅ = 𝑨 + ∑ 𝒅𝒊
𝑿
𝒏
𝒊=𝟏
𝑤ℎ𝑒𝑟𝑒 𝑑𝑖 = 𝑋𝑖 − 𝐴 and A= assumed Mean
b. For frequency distribution, the mean is computed as:
Step:
i. Obtain the class mark(x) or mid-point(x)
ii. Create a column for 𝑑𝑖 = 𝑋𝑖 − 𝐴; i.e. deviation of each value of x from the
assumed mean
iii. Multiply the respective class frequency by the corresponding deviation(𝑑𝑖 ) and
sum up the column to get ∑𝒏𝒊=𝟏 𝒇𝒅𝒊
iv. The Arithmetic mean by Assumed mean or shot cut method is given by:
𝒏
𝟏
̅ = 𝑨 + ∑ 𝒇𝒅𝒊
𝑿
𝒏
𝒊=𝟏

Example 4: Find the mean weight of the following students by assumed mean method. The
weights are in kg: 67, 69, 66, 68, 63, 76, 72, 74, 70, 65

Solution:let us take A= 68 as assumed mean

Value (x) 𝑑𝑖 = 𝑋𝑖 − 68
63 -5
65 -3
66 -2
67 -1
68 0
69 1
70 2
72 4
74 6
76 8
𝒏

∑ 𝒅𝒊 = 𝟏𝟎
𝒊=𝟏

Therefore, the arithmetic Mean is given as:


𝒏
𝟏 𝟏𝟎
̅ = 𝑨 + ∑ 𝒅𝒊 = 𝟔𝟖 +
𝑿
𝒏 𝟏𝟎
𝒊=𝟏
̅ = 𝟔𝟖 + 𝟏 = 𝟔𝟗𝒌𝒈
𝑿

4
Example 3: Compute the mean of the grouped frequency data given below using an assumed
mean of 44.5

Weight 10-19 20-29 30-39 40-49 50-59 60-69 70-79 Total


interval
Frequency(f) 5 19 10 13 4 4 2 57
Solution:

Weight Interval Frequency(f) Mark 𝑑𝑖 = 𝑋𝑖 − 44.5 𝒇𝒅𝒊


class(x)
10-19 5 14.5 -30 -150
20-29 19 24.5 -20 -380
30-39 10 34.5 -10 -100
40-49 13 44.5 0 0
50-59 4 54.5 10 40
60-69 4 64.5 20 80
70-79 2 74.5 30 60
𝒏
Total ∑ 𝒇𝒊 =57
∑ 𝒇𝒅𝒊 = −𝟒𝟓𝟎
𝒊=𝟏
The Arithmetic Mean is given as:
𝒏
𝟏
̅ = 𝑨 + ∑ 𝒇𝒅𝒊
𝑿
𝒏
𝒊=𝟏
−𝟒𝟓𝟎
̅ = 𝟒𝟒. 𝟓 +
𝑿 = 𝟒𝟒. 𝟓 − 𝟕. 𝟗 = 𝟑𝟔. 𝟔𝒍𝒃
𝟓𝟕

5
GEOMETRIC MEAN
The geometric mean of a series containing n observations is the nth root of the product of the
values. The geometric mean(GM) of a set of n observations 𝑥1 , 𝑥2 , … . . , 𝑥𝑛 is defined as
𝑮𝒎 = 𝒏√𝑥1 × 𝑥2 × … . .× 𝑥𝑛
𝟏
𝑮𝑴 = (𝑥1 × 𝑥2 × … . .× 𝑥𝑛 )𝒏
Taking the log of both sides
1 1
𝑙𝑜𝑔𝐺𝑀 = 𝑙𝑜𝑔(𝑥1 + 𝑥2 + ⋯ . . +𝑥𝑛 ) = (𝑙𝑜𝑔𝑥1 + 𝑙𝑜𝑔𝑥2 + ⋯ . . +𝑙𝑜𝑔𝑥𝑛 )
𝑛 𝑛
∑ 𝑙𝑜𝑔𝑥𝑖
𝑙𝑜𝑔𝐺𝑀 =
𝑛
Therefore,
∑ 𝑙𝑜𝑔𝑥𝑖
𝑮𝑴 = 𝑨𝒏𝒕𝒊𝒍𝒐𝒈 [ ]
𝑛
Example 1: The data below represent the scores of nine(9) students in a test. Find
the geometric mean score. 45, 32, 37, 46, 39, 36, 41, 48, 36.
Solution:
Scores(𝑥𝑖 ) 𝑙𝑜𝑔𝑥𝑖
45 1.6532
32 1.5052
37 1.5682
46 1.6628
39 1.5911
36 1.5563
41 1.6128
48 1.6212
36 1.5563
∑ 𝑙𝑜𝑔𝑥𝑖 = 14.387

∑ 𝑙𝑜𝑔𝑥𝑖 14.387
𝑮𝑴 = 𝑨𝒏𝒕𝒊𝒍𝒐𝒈 [ ] = 𝒂𝒏𝒕𝒊 − 𝒍𝒐𝒈 [ ] = 𝒂𝒏𝒕𝒊𝒍𝒐𝒈(𝟏. 𝟓𝟗𝟖𝟓𝟔)
𝑛 9
𝑮𝑴 = 𝟑𝟗. 𝟔𝟗

For grouped frequency data: The geometric mean is given as


∑ 𝑓 𝑙𝑜𝑔𝑥𝑖
𝐺𝑀 = 𝐴𝑛𝑡𝑖𝑙𝑜𝑔 [ 𝑖 ∑ 𝑓𝑖
]; 𝑛 = ∑ 𝑓𝑖

6
Example 1: The following frequency distribution of weights of 60 apples.
Find the geometric mean for the grouped data.
marks 65-84 85-104 105-124 125-144 145-164 165-184 185-204 Total
No of 9 10 17 10 5 4 5 60
students
Solution:

Class mark(x) f Log x f logx


(65+84)/2= 74.5 9 1.8722 16.8498
94.5 10 1.9754 19.754
114.5 17 2.0588 34.9996
134.5 5 2.1287 10.6435
154.5 4 2.1889 8.7556
174.5 5 2.2418 11.209
Total ∑ 𝑓𝑖 =60 ∑ 𝑓𝑖 𝑙𝑜𝑔𝑥𝑖 =93.4559

∑ 𝑓𝑖 𝑙𝑜𝑔𝑥𝑖 93.4559
𝐺𝑀 = 𝐴𝑛𝑡𝑖𝑙𝑜𝑔 [ ∑ 𝑓𝑖
] = 𝐴𝑛𝑡𝑖𝑙𝑜𝑔 [ ] = 𝐴𝑛𝑡𝑖𝑙𝑜𝑔(1.5576)
60

𝐺𝑀 = 𝐴𝑛𝑡𝑖𝑙𝑜𝑔(1.5576) = 36.11

7
HARMONIC MEAN

Harmonic Mean is the reciprocal of the arithmetic mean of the reciprocals of a finite set of
numbers. It is calculated by dividing the number of observations by the sum of reciprocals of
each value of the data set.

Harmonic Mean(HM) for n observations 𝒙𝟏 , 𝒙𝟐 … … . . , 𝒙𝒏 is given as:

𝑛
𝐻𝑀 = [ ] 𝑎𝑛𝑑
1
∑𝑖=1 ( )
𝑥𝑖

For Frequency Distribution data

𝑛
𝐻𝑀 = [ ]
1
∑𝑖=1 𝑓𝑖 ( )
𝑥𝑖

Example1: The following are weights in pounds of 10 randomly selected children at a day-care
center: 68 ,63, 42, 27, 30, 36 ,28, 32 ,79 ,27. Compute the Harmonic mean

Solution

Values(𝑥𝑖 ) 1
𝑥𝑖
68 0.0147
63 0.0158
42 0.0238
27 0.03704
30 0.03333
36 0.02778
28 0.0357
32 0.03125
79 0.01266
27 0.03704
Total 1
∑ ( ) = 0.2691
𝑥𝑖
𝑖=1

8
𝑛 10
𝐻𝑀 = [ ]= = 37.16
1 0.2691
∑𝑖=1 ( )
𝑥𝑖

Example 2: The following frequency distribution of weights of 60 apples.


Find the Harmonic mean for the grouped data.
marks 65-84 85-104 105-124 125-144 145-164 165-184 185-204 Total
No of 9 10 17 10 5 4 5 60
students
Solution:

Class mark(𝑥𝑖 ) f 1 1
𝑓( )
𝑥𝑖 𝑥𝑖
(65+84)/2= 74.5 9 0.0134 0.1206
94.5 10 0.0106 0.1058
114.5 17 0.00873 0.1484
134.5 5 0.00744 0.0372
154.5 4 0.00647 0.02588
174.5 5 0.00573 0.02865
Total ∑ 𝑓𝑖 =60 0.46653

𝑛 60
𝐻𝑀 = [ ]= = 128.61
1 0.46653
∑𝑖=1 𝑓𝑖 ( )
𝑥𝑖

Merits of H.M

1. It is rigidly defined.

2. It is defined on all observations.

3. It is amenable to further algebraic treatment.

4. It is the most suitable average when it is desired to give greater weight to smaller observations
and less weight to the larger ones.

Demerits of H.M

1. It is not easily understood.


2. It is difficult to compute.
3. It is only a summary figure and may not be the actual item in the serie.

9
4. It gives greater importance to small items and is therefore, useful only when small items
have to be given greater weightage.
5. It is rarely used in grouped data

TUTORIAL QUESTION

1. In a survey of 10 households, the number of children was found to be


4, 1, 5, 4, 3, 7, 2, 3, 4, 1

Find

i. Arithmetic Mean
ii. Geometric Mean and
iii. Harmonic Mean.
2. Find the Arithmetic mean, Harmonic mean and Geometric Mean of each set of
observations below.
(i) 3, 4, 7, 3, 5, 2, 6, 10
(ii) 8, 10, 12, 14, 7, 16, 5, 7, 9, 11
(iii) 17, 18, 16, 17, 17, 14, 22, 15, 16, 17, 14, 12
3. A survey of 100 households on the number of cars own in each household resulted in the
data given below.

No of cars 0 1 2 3 4 Total
No of 5 70 21 3 1 100
students
Calculate the
i. Arithmetic mean number of cars per household
ii. Geometric mean number of cars per household
iii. Harmonic mean number of cars per household
4. A manager keeps a record of the number of calls she makes each day on her mobile
phone.

No of calls 0 1 2 3 4 5 6 7 8 Total
per day
frequency 3 4 7 8 12 10 14 3 1 62
Calculate the Arithmetric mean, Harmonic and Geometric mean number of calls per day

5. The table below gives data on the heights, in cm, of 51 children.

Weight interval 140- 149 150-159 160-169 170-179 Total

Frequency(f) 6 16 21 8 51

Calculate the Arithmetric mean, Harmonic and Geometric mean height.

10
MEDIAN:
Another useful measure of location is the median. If the observations in the data set are
arranged in increasing or decreasing order, the median is the middle observation, which divides
the data set into equal halves.

A. Calculation for n observations:

Procedures:

1. Arrange the data in (either ascending or descending) order of magnitude.


𝑡ℎ
2. If the number of observations n is odd there will be a unique median, then (𝑛+1
2
)
observation from either end in the ordered sequence is the median.
𝑡ℎ 𝑡ℎ
3. If n is even, in this case there are two middle terms: (𝑛2) and (𝑛2+1) .
4. The Median is average of the two middle terms:
𝑡ℎ 𝑡ℎ
(𝑛
2
) + (𝑛
2
+1)
𝑀= 2

Example: Find the median of the following observations

a. 21, 12, 49, 37, 88, 40, 55, 74, 63


b. 88, 72, 33, 29, 70, 86, 54, 91, 61, 57
Solution: first arrange the data in order
a. 12, 21, 37, 40, [49], 55, 63, 74, 88
Since n=9 (odd)
𝑛+1 𝑡ℎ 9+1 𝑡ℎ
𝑀=( 2
) =( 2
) = 5𝑡ℎ 𝑜𝑏𝑠𝑒𝑟𝑣𝑎𝑡𝑖𝑜𝑛𝑠
The median= 49
b. Arrange the data in order
29, 33, 54, 57, [61], [70], 72, 86, 88, 91
Here, n=10(even)
Therefore, the median
𝑛 𝑡ℎ 𝑛 𝑡ℎ
𝑀 = 𝑎𝑣𝑒𝑟𝑎𝑔𝑒 𝑜𝑓 ( 2) 𝑡𝑒𝑟𝑚 𝑎𝑛𝑑 ( 2+1) 𝑡𝑒𝑟𝑚𝑠
𝑀 = 𝑎𝑣𝑒𝑟𝑎𝑔𝑒 𝑜𝑓 5𝑡ℎ 𝑡𝑒𝑟𝑚 𝑎𝑛𝑑 6𝑡ℎ 𝑡𝑒𝑟𝑚𝑠
61 + 70
𝑀= = 65.2
2

B. For ungrouped frequency data:


i. Arrange the data in (either ascending or descending) order of magnitude.
ii. Prepare a table showing the corresponding frequencies and cumulative frequencies.

11
iii. The median is calculated by
𝑛+1 𝑡ℎ
𝑀=( 2
) ; 𝑛 = ∑𝑓
iv. The first value with a cumulative frequency greater than depth of the median is the
median. However, if the depth of the median is exactly 0.5 more than the
cumulative of the previous class, then the median is the mid-point between the two
classes.

Example1; The data below represent the monthly allowance of 100 students

Allowances($) 100 150 80 200 260 180 Total


No of 20 22 12 16 4 26 100
students

Solution: Let arrange the allowance in ascending order and obtain the cumulative frequency

Allowance in No of Cumulative
ascending Students(f) frequency (Cf)
order
80 12 12
100 20 32
150 22 54
180 26 80
200 16 96
260 4 100
Total 100
Since n=100(even)

Therefore, the median


𝑛 𝑡ℎ 𝑛 𝑡ℎ
𝑀 = 𝑎𝑣𝑒𝑟𝑎𝑔𝑒 𝑜𝑓 ( 2) 𝑡𝑒𝑟𝑚 𝑎𝑛𝑑 ( 2+1) 𝑡𝑒𝑟𝑚𝑠
100 𝑡ℎ 100 𝑡ℎ
𝑀=( 2
) 𝑡𝑒𝑟𝑚 𝑎𝑛𝑑 ( 2
+1)

𝑀 = 𝑎𝑣𝑒𝑟𝑎𝑔𝑒 𝑜𝑓 50𝑡ℎ 𝑡𝑒𝑟𝑚 𝑎𝑛𝑑 51𝑡ℎ 𝑡𝑒𝑟𝑚𝑠


The median is the average of observations at 50th and 51th of the cumulative
frequency.
150+150
𝑀= 2
=150

𝑀𝑒𝑑𝑖𝑎𝑛 = 150

12
C. For Grouped Frequency data:
Procedures:
i. Here data is given in the form of a frequency table with class interval.
ii. Obtain the cumulative frequencies column.
𝑡ℎ
iii. The median class is determined by (∑2𝑓) 𝑡𝑒𝑟𝑚
iv. The median is calculated by the formular:
∑𝑓 𝑤𝑚
Median=𝐿𝑚 + [ − 𝐶𝐹𝑏𝑚 ] ×
2 𝑓𝑚

𝐿𝑚 = lower class boundary of the median class

𝐶𝐹𝑏𝑚 - Cumulative frequency before the median class

𝑓𝑚 - Frequency of the median class

𝑤𝑚 - Class size of the median class

∑ 𝑓 = Total frequency

Advantages of Median

(i) It is easy to define and understood


(ii) It is not influence by extreme values
(iii) The median is used to find the average of an open-ended distribution.
(iv) Its unique and exist for any set of numerical data.
(v) It is the best measure for qualitative data.

Disadvantages of Median

i. The median unlike the mean does not use the whole data set.
ii. The median does not use for any further statistical analysis.

Example 1: The following data gives the marks obtained by students in STA 131.
Find the median.
marks 11-20 21-30 31-40 41-50 51-60 61-70 71-80 Total
No of 42 38 125 84 45 36 30 400
students

13
Solution:

Class interval No of Class Cumulative


students boundary frequency(cf)
(f)
11-20 42 10.5- 20.5 42
21-30 38 20.5-30.5 80
31-40 125 30.5- 40.5 205
41-50 84 40.5-50.5 289
51-60 45 50.5-60.5 334
61-70 36 60.5-70.5 370
71-80 30 70.5-80.5 400
𝑡ℎ
Median class is the class that contain the (∑2𝑓) 𝑜𝑏𝑠𝑒𝑟𝑣𝑎𝑡𝑖𝑜𝑛
∑ 𝑓 𝑁 400
= = = 200
2 2 2
It is more than the cumulative frequency 80, but less than c.f. 205.
Hence the median class is 31-40
𝑁
= 200, 𝐿𝑚 = 30.5; 𝑓𝑚 = 205; 𝑐𝑓𝑚 = 80 𝑎𝑛𝑑 𝑤𝑚 = 10
2
∑𝑓
−𝐶𝐹𝑏𝑚
Median=𝐿𝑚 + [ 2 ] × 𝑤𝑚
𝑓𝑚

[200−80] 120
Median=30.5 + × 10 = 30.5 + 205 × 10
205

𝑀𝑒𝑑𝑖𝑎𝑛 = 30.5 + 9.6 = 40.1

Example 2: Compute the Median of the grouped data given below:

Weight 10- 20- 30- 40- 50- 60- 70- Total


interval 19 29 39 49 59 69 79
Frequency(f) 5 19 10 13 4 4 2 57
Total frequency =57 which is odd number

Median class is the class with ½(57+1)=58/2= 29th observation;

14
Weight frequency Cumulative
interval frequency
(CF)
10- 19 5 5
20- 29 19 24
30- 39 10 34
40- 49 13 47
50- 59 4 51
60- 69 4 55
70- 79 2 57
Total 57
Using the CF column, the 29th observation falls in class interval 30-39 and the boundaries of
that class is 29.5-39.5

Therefore, 𝐿𝑚 = 29.5 ; 𝐶𝐹𝑏𝑚 =24 ; 𝑓𝑚 =10; 𝑊𝑚 =10 and ∑ 𝑓 =57


∑𝑓 𝑤𝑚 57 10
Median=𝐿𝑚 + [ − 𝐶𝐹𝑏𝑚 ] × = 29.5 + [ 2 − 24] × 10
2 𝑓𝑚

Median= 29.5 + [28.5 − 24] × 1= 29.5 + [4.5]= 34

Tutorial Questions

15
MODE
Mode of a set of data is defined as the value that occurs most frequently among the value of
the variable. In frequency distribution any value with the highest frequency is the mode.
For ungrouped frequency
By inspection:
Procedure: In case of n observation or ungrouped frequency data, mode can be obtained by
examine the data value with the highest frequency.
Example1: Find the mode from the data given below:
1. 1, 1, 3, 3, 1, 5, 3, 2, 2, 3, 5, 3, 3, 6, 7, 2, 3, 4, 4,6
Solution:
Value(x) No of times (f)
1 3
2 3
3 7
4 2
5 2
6 2
7 1
Total ∑ 𝑓 = 𝑁 =20
The table indicate that the value “3” has the maximum frequency.
Hence, mode is 3

The mode is especially useful for describing qualitative variables or quantitative variables that
take on a small number of possible values. If two values occur more than others but equally
frequently, we say the data are bimodal, or more generally multimodal. The term bimodal is
also used to describe distribution in which there are two peaks, not necessarily of the same
height.
For grouped frequency data, the Mode is defined as:
Procedure:
i. The modal class is the class with the highest frequency
ii. Find the mode using the formular below:
𝑜 𝑓 −𝑓𝑏
Mode=𝐿𝑜 + [2𝑓 −𝑓 ] × 𝑤𝑜
𝑜 𝑏 −𝑓𝑎
𝐿𝑜 = lower class boundary of the modal class
𝑓𝑜 - Frequency of the modal class
𝑓𝑏 - Frequency before the modal class
𝑓𝑎 - Frequency after the modal class
𝑤𝑜 − Class size of the modal class
Advantages of Mode

16
i. Mode is the easiest average to compute.
ii. Mode is not generally influenced by extreme values.
iii. Mode can be determined for categorical variables
Disadvantages of Mode
i. The mode may not exist.
ii. The mode may not be unique.
iii. The Mode does not lend itself to further statistical Analysis.
Example: Compute the Mode and Median of the grouped data below:

Weight 10-19 20-29 30-39 40-49 50-59 60-69 70-79 Total


interval

Frequency(f) 5 19 10 13 4 4 2 57

Solution:

𝑜 𝑓 −𝑓𝑏
1. Mode=𝐿𝑜 + [2𝑓 −𝑓 ] × 𝑤𝑜
𝑜 𝑏 −𝑓𝑎

Mode is the observation with highest frequency

Step 1: determine the modal class and its boundaries

Since the highest frequency is 19 and it falls in class interval 20-29

Modal class Boundaries= 19.5- 29.5,

Therefore,

𝐿𝑜 =19.5; 𝑓𝑜 = 19; 𝑓𝑏 = 5; 𝑓𝑎 = 10 and 𝑤𝑜 =10


19−5
Mode=19.5 + [2(19)−5−10] × 10

14
Mode=19.5 + [23] × 10 = 19.5 + 0.6086 𝑋10 = 19.5 + 6.086

Mode= 25 .59

17
2.1 MEASURE OF PARTITION
The partition values are the measures used to divide the total number of observations from a
distribution into a certain number of equal parts. The measure of partition are;
I. Measure of quartile
II. Measure of decile
III. Measure of percentile

2.1.1 MEASURE OF QUARTILE:


The measure of quartile is the measure of partition of the value of a distribution at each part when
the whole distribution is being divided into four parts. The formula for quartile is given by:
1. For n observations and ungrouped frequency data
The formula for quartile is given by:

𝑁 + 1 𝑡ℎ
𝑄𝑖 = 𝑖 × ( ) 𝑖𝑡𝑒𝑚 ; 𝑤ℎ𝑒𝑟𝑒 𝑖 = 1, 2, 3
4
2. For grouped frequency data, the formula is then given by:
𝑖×𝑁
( 4 − 𝐶𝑓𝑏 )
𝑄𝑖 = 𝑙1 + [ ]×𝑐
𝑓𝑖

Example 2.1.1: For n observations data


Compute Q1 and Q3 for the data relating to the marks of 8 students in an examination given below
25, 48,32,52,21,64,29,57.
Solution:
N=8
Arrange the values in ascending order, 21, 25, 29, 32, 48, 52, 57,64
𝑁+1 𝑡ℎ
𝑄𝑖 = 𝑖 × ( ) 𝑖𝑡𝑒𝑚
4

8+1 𝑡ℎ 9 𝑡ℎ
𝑄1 = 1 × ( ) 𝑖𝑡𝑒𝑚 = (4) 𝑖𝑡𝑒𝑚
4

𝑄1 = (2.25)𝑡ℎ 𝑖𝑡𝑒𝑚
𝑄1 = (2)𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 + 0.25(3𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 − 2𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 )

STA 131: INTRDUCTION TO STATISTICAL INFERENCE BY DR. R. B. AFOLAYAN 1


𝑄1 = 25 + 0.25(29 − 25 ) = 32 + 0.25(4) = 25 + 1 = 26

𝑁+1 𝑡ℎ
𝑄𝑖 = 𝑖 × ( ) 𝑖𝑡𝑒𝑚
4

8+1 𝑡ℎ 27 𝑡ℎ
𝑄3 = 3 × ( ) 𝑖𝑡𝑒𝑚 = ( 4 ) 𝑖𝑡𝑒𝑚
4

𝑄3 = (6.75)𝑡ℎ 𝑖𝑡𝑒𝑚
𝑄3 = (6)𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 + 0.75(7𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 − 6𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 )
𝑄3 = 52 + 0.75(57 − 52) = 52 + 3.75 = 55.75
QUARTILE FOR UNGROUPED FREQUENCY DATA:
1. Find cumulative frequencies
2. Find
𝑁 + 1 𝑡ℎ
𝑄𝑖 = 𝑖 × ( ) 𝑖𝑡𝑒𝑚
4
𝑁+1 𝑡ℎ
3. Check the cumulative frequencies, the value just greater than 𝑖 × ( ) 𝑖𝑡𝑒𝑚 , the
4
corresponding value of x is Qi
Example 2.1.2: For ungrouped frequency data
Compute Q1 and Q3 for the data relating to age in years of 543 members in a village

Age 20 30 40 50 60 70 80
No of 3 61 132 153 140 51 3
member

Solution:

X F CF
20 3 3
30 61 64
40 132 196
50 153 349
60 140 489
70 51 540
80 3 543
Total ∑ 𝑓 = 543

STA 131: INTRDUCTION TO STATISTICAL INFERENCE BY DR. R. B. AFOLAYAN 2


Where 𝑁 = ∑ 𝑓

𝑁 + 1 𝑡ℎ
𝑄𝑖 = 𝑖 × ( ) 𝑖𝑡𝑒𝑚
4
First Quartile (𝑸𝟏 ):

𝑁 + 1 𝑡ℎ
𝑄1 = 𝑖 × ( ) 𝑖𝑡𝑒𝑚
4
543 + 1 𝑡ℎ 544 𝑡ℎ
𝑄1 = 1 × ( ) 𝑖𝑡𝑒𝑚 = ( ) 𝑖𝑡𝑒𝑚
4 4
𝑄1 = (136)𝑡ℎ 𝑖𝑡𝑒𝑚
From the cumulative column, the corresponding value of x for which the cumulative frequency
of 136 is found is the ith quartile value. So, after cumulative frequency of 64, value of 40 has 65
to 196 observations in which 136 is within the range. Therefore,
𝑄1 = 40 𝑦𝑒𝑎𝑟𝑠

Third Quartile (𝑸𝟑 ):

𝑁+1 𝑡ℎ
𝑄3 = 3 × ( ) 𝑖𝑡𝑒𝑚
4

543 + 1 𝑡ℎ
𝑄1 = 3 × ( ) 𝑖𝑡𝑒𝑚
4
𝑄1 = (3 × 136)𝑡ℎ 𝑖𝑡𝑒𝑚 = (408)𝑡ℎ 𝑖𝑡𝑒𝑚
𝑄3 = 60 𝑦𝑒𝑎𝑟𝑠
FOR GROUPED FREQUENCY DATA, THE FORMULA IS THEN GIVEN BY:
1. Find the cumulative frequency
2. Find the ith quartile using:
𝑖 × 𝑁 𝑡ℎ
𝑄𝑖 = ( ) 𝑖𝑡𝑒𝑚
4

3. 𝑄𝑖 Class is the class interval corresponding to the value of the cumulative frequency just
greater than (i × n)/4.

4. The ith quartile is calculated by the formula below:

STA 131: INTRDUCTION TO STATISTICAL INFERENCE BY DR. R. B. AFOLAYAN 3


𝑖×𝑁
( 4 − 𝐶𝑓𝑏 )
𝑄𝑖 = 𝑙1 + [ ]×𝑐
𝑓𝑖

Where i = 1, 2, 3, 4
N = total of all frequency value
𝑙𝑖 = lower class limit of the ith quartile class
𝑓𝑖 = frequency of the ith quartile class
𝑐 = class interval
𝑓𝐶𝑏 = cumulative frequency preceding the quartile class

Example 2.1.3: For Grouped frequency data


Calculate the Q1 and Q3 for the wages of the laborers given below
Wages 30-32 32-34 34-36 36-38 38-40 40-42 42-44
(#000)
laborers 12 18 16 14 12 8 6

Solution:

wages Frequency CF
30-32 12 12
𝑄1= 32-34 18 30
34-36 16 46
36-38 14 60
𝑸𝟑= 38- 12 72
40
40-42 8 80
42-44 6 86
Total ∑ 𝑓 = 86

1st Quartile position/ Class:


𝑖 × 𝑁 𝑡ℎ
𝑄𝑖 = ( ) 𝑖𝑡𝑒𝑚
4
1 × 86 𝑡ℎ
𝑄1 = ( ) 𝑖𝑡𝑒𝑚 = (21.5)𝑡ℎ 𝑖𝑡𝑒𝑚
4
Checking the cumulative frequency column 21.5th observation fall in 2nd class, hence
The first quartile lies in the group 32 – 34
𝑖×𝑁
( 4 − 𝐶𝑓𝑏 )
𝑄𝑖 = 𝑙𝑖 + [ ]×𝑐
𝑓𝑖

STA 131: INTRDUCTION TO STATISTICAL INFERENCE BY DR. R. B. AFOLAYAN 4


(21.5 − 𝐶𝑓𝑏 )
𝑄1 = 𝑙1 + [ ]×𝑐
𝑓𝑖
𝑙1 = 32; 𝐶𝑓𝑏 = 12; 𝑓1 = 18; 𝑐 = 2

(21.5 − 12)
𝑄1 = 32 + [ ]×2
18
(19)
𝑄1 = 32 + [ ]
18
𝑄1 = 32 + 1.06 = 𝟑𝟑. 𝟎𝟔

3rd Quartile position/ Class:


𝑖 × 𝑁 𝑡ℎ
𝑄𝑖 = ( ) 𝑖𝑡𝑒𝑚
4
3 × 86 𝑡ℎ
𝑄3 = ( ) 𝑖𝑡𝑒𝑚 = (64.5)𝑡ℎ 𝑖𝑡𝑒𝑚
4
Checking the cumulative frequency column 64.5th observation fall in 4th class, hence
The first quartile lies in the group 38 – 40
𝑖×𝑁
( 4 − 𝐶𝑓𝑏 )
𝑄3 = 𝑙3 + [ ]×𝑐
𝑓𝑖
(21.5 − 𝐶𝑓𝑏 )
𝑄1 = 𝑙3 + [ ]×𝑐
𝑓3
𝑙1 = 38; 𝐶𝑓𝑏 = 60; 𝑓1 = 12; 𝑐 = 2

(64.5 − 60)
𝑄3 = 38 + [ ]×2
12
(9)
𝑄3 = 38 + [ ]
12
𝑄3 = 38 + 0.75 = 𝟑𝟖. 𝟕𝟓

Inter-Quartile Range and Semi Inter-Quartile Range


Inter-quartile range is the difference between the first quartile (Q1) and the third quartile
(Q3) while semi inter-quartile range is the average of difference between Q1 and Q3. It is
therefore given by the formula:
𝐼𝑛𝑡𝑒𝑟 − 𝑞𝑢𝑎𝑟𝑡𝑖𝑙𝑒 𝑟𝑎𝑛𝑔𝑒(𝐼𝑄𝑅) = 𝑄3 − 𝑄1

𝑄3−𝑄1
𝑆𝑒𝑚𝑖 𝑖𝑛𝑡𝑒𝑟 − 𝑞𝑢𝑎𝑟𝑡𝑖𝑙𝑒 𝑟𝑎𝑛𝑔𝑒(𝑆𝐼𝑄𝑅) = 2

STA 131: INTRDUCTION TO STATISTICAL INFERENCE BY DR. R. B. AFOLAYAN 5


Example2.1.4
From the example 2.1.1 given above, calculate the Inter-quartile and semi inter-quartile
range.
Solution:
From example 1,
The first and third quartile are:
Q1= 26 and Q3= 55.75
1. 𝐼𝑛𝑡𝑒𝑟 − 𝑞𝑢𝑎𝑟𝑡𝑖𝑙𝑒 𝑟𝑎𝑛𝑔𝑒(𝐼𝑄𝑅) = 𝑄3 − 𝑄1 = 55.75 − 26 = 29.75

𝑄3−𝑄1
2. 𝑆𝑒𝑚𝑖 𝑖𝑛𝑡𝑒𝑟𝑞𝑢𝑎𝑟𝑡𝑖𝑙𝑒 𝑟𝑎𝑛𝑔𝑒 = 2

55.75 − 29.7
𝑆𝐼𝑄𝑅 = = 14.88
2
From example 2.1.3;
The first and third quartile are:

Q1 = 33.06 and Q3 = 38.75


1. 𝐼𝑛𝑡𝑒𝑟 − 𝑞𝑢𝑎𝑟𝑡𝑖𝑙𝑒 𝑟𝑎𝑛𝑔𝑒(𝐼𝑄𝑅) = 𝑄3 − 𝑄1 = 38.75 − 33.06 = 5.69

𝑄3−𝑄1
2. 𝑆𝑒𝑚𝑖 𝑖𝑛𝑡𝑒𝑟𝑞𝑢𝑎𝑟𝑡𝑖𝑙𝑒 𝑟𝑎𝑛𝑔𝑒 = 2

38.75 − 33.06 5.69


𝑆𝐼𝑄𝑅 = = = 2.845
2 2

2.2.1 MEASURE OF DECILE:


This is the measure of partition of the value of a distribution part when the whole distribution is
being divided into ten parts. The decile is given by:
A. For n observations and ungrouped frequency data
The formula for quartile is given by:

𝑁 + 1 𝑡ℎ
𝐷𝑖 = 𝑖 ( ) 𝑖𝑡𝑒𝑚 ; 𝑤ℎ𝑒𝑟𝑒 𝑖 = 1, 2, 3, 4, … . . ,9
10
B. For grouped frequency data, the formula is then given by:
𝑖×𝑁
( 10 − 𝐶𝑓𝑏 )
𝐷𝑖 = 𝑙1 + [ ]×𝐶
𝑓𝑖

STA 131: INTRDUCTION TO STATISTICAL INFERENCE BY DR. R. B. AFOLAYAN 6


𝑊ℎ𝑒𝑟𝑒 𝑖 = 1, 2, 3, … , 9
N= total of all frequency value
𝑙𝑖 = lower class limit of the ith decile class
𝑓𝑖 = frequency of the ith decile class
𝑐 = class interval
𝐶𝑓𝑏 = cumulative frequency preceding the decile class

Examples 2.2.1
1. Find the D6 for the following data 11, 25, 20, 15, 24, 28, 19, 21.

Solution:
Arrange in an ascending order 11, 15, 19, 20, 21, 24, 25, 28

N=8

𝑁 + 1 𝑡ℎ
𝐷𝑖 = 𝑖 × ( ) 𝑖𝑡𝑒𝑚
10
8 + 1 𝑡ℎ 9 𝑡ℎ
𝐷6 = 6 × ( ) 𝑖𝑡𝑒𝑚 = 6 × ( ) 𝑖𝑡𝑒𝑚
10 10

54 𝑡ℎ
𝐷6 = ( ) 𝑖𝑡𝑒𝑚 = (5.4)𝑡ℎ 𝑖𝑡𝑒𝑚
10
𝑡ℎ
𝐷6 = (5) 𝑣𝑎𝑙𝑢𝑒 + 0.4(6𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 − 5𝑡ℎ 𝑣𝑎𝑙𝑢𝑒)
𝐷6 = 21 + 0.4(24 − 21) = 21 + 0.4(3)
𝐷6 = 21 + 1.2 = 22.2

FOR GROUPED FREQUENCY DATA, THE FORMULA IS THEN GIVEN BY:


1. Find the cumulative frequency
2. Find the ith Decile using:
𝑖 × 𝑁 𝑡ℎ
𝐷𝑖 = ( ) 𝑖𝑡𝑒𝑚
10

3. 𝐷𝑖 Class is the class interval corresponding to the value of the cumulative frequency just
greater than (i × N)/10.

4. The ith decile is calculated by the formula below:

STA 131: INTRDUCTION TO STATISTICAL INFERENCE BY DR. R. B. AFOLAYAN 7


𝑖×𝑁
( 10 − 𝐶𝑓𝑏 )
𝐷𝑖 = 𝑙1 + [ ]×𝑐
𝑓𝑖

Example 2.2.2: Calculate D5 for the frequency distribution of monthly income of workers in a factory.

Income (in 0-4 4-8 8-12 12-16 16-20 20-24 24-28 28-32
thousand)
No of 10 12 8 7 5 8 4 6
persons

Solution:

Class F CF
0-4 10 10
4-8 12 22
8-12 8 30
12-16 7 37
16-20 5 42
20-24 8 50
24-28 4 54
28-32 6 60
Total ∑ 𝑓 = 60

1st Decile position/ Class:

𝑖 × 𝑁 𝑡ℎ
𝐷𝑖 = ( ) 𝑖𝑡𝑒𝑚
10
5 × 60 𝑡ℎ
𝐷5 = ( ) 𝑖𝑡𝑒𝑚 = (30)𝑡ℎ 𝑖𝑡𝑒𝑚
10

Checking the cumulative frequency column, 30th observation fall in 3rd class, hence
The 5th decile lies in the group 8 – 12

The 5th decile lies in the group 8-12

𝑖×𝑁
( 10 − 𝐶𝑓𝑏 )
𝐷𝑖 = 𝑙1 + [ ]×𝑐
𝑓𝑖

STA 131: INTRDUCTION TO STATISTICAL INFERENCE BY DR. R. B. AFOLAYAN 8


𝑙1 = 8; 𝐶𝑓𝑏 = 22; 𝑓1 = 8; 𝑐 = 4

(30 − 22)
𝐷5 = 8 + [ ]×4
8
8
𝐷5 = 8 + [ ] × 4 = 8 + 4
8
𝐷5 = 12

2.3.1 MEASURE OF PERCENTILE:


Percentile is a measure of location of the value of a distribution at each part when the whole
distribution is being divided into 100 equal parts. The percentile formula is given:
A. For n observations and ungrouped frequency data
The formula for quartile is given by:

𝑁 + 1 𝑡ℎ
𝑃𝑖 = 𝑖 × ( ) 𝑖𝑡𝑒𝑚 ; 𝑤ℎ𝑒𝑟𝑒 𝑖 = 1, 2, 3, 4, … . . ,99
100
B. For grouped frequency data, the formula is then given by:
𝑖×𝑁
( 100 − 𝐶𝑓𝑏 )
𝑃𝑖 = 𝑙1 + [ ]×𝐶
𝑓𝑖

𝑤ℎ𝑒𝑟𝑒 𝑖 = 1, 2, 3, 4, … . . ,99
N=∑ 𝑓 = total of all frequency value
𝑙𝑖 = lower class limit of the percentile class
𝑓𝑖 = frequency of the percentile class
𝑐 = class interval
𝐶𝑓𝑚 = cumulative frequency preceding the percentile class
Example 2.3.1: The following is the monthly income (in thousand) of 8 persons working in a
factory. Find P30 income value. 10, 14, 36, 25, 15, 21, 29, 17.
Solution:
Arrange the data in an ascending order. (n = 8)
10, 14, 15, 17, 21, 25, 29, 36
N=8

𝑁 + 1 𝑡ℎ
𝑃𝑖 = 𝑖 × ( ) 𝑖𝑡𝑒𝑚
100

STA 131: INTRDUCTION TO STATISTICAL INFERENCE BY DR. R. B. AFOLAYAN 9


8 + 1 𝑡ℎ 9 𝑡ℎ
𝑃30 = 30 × ( ) 𝑖𝑡𝑒𝑚 = 30 × ( ) 𝑖𝑡𝑒𝑚
100 100

270 𝑡ℎ
𝑃30 =( ) 𝑖𝑡𝑒𝑚 = (2.7)𝑡ℎ 𝑖𝑡𝑒𝑚
100
𝑃30 = (2)𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 + 0.7(3𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 − 2𝑡ℎ 𝑣𝑎𝑙𝑢𝑒)
𝑃30 = 14 + 0.7(15 − 14) = 14 + 0.7(1)
𝑃30 = 14 + 0.7 = 14.7

Example 2.3.2:
Calculate P61 for the following data relating to the height of the plants in a garden.
Height 0-5 5-10 10-15 15-20 20-25 25-30
(cm)
No of 18 20 36 40 26 16
plant

Solution:

Class F CF
0–5 18 18
5 - 10 20 38
10 - 15 36 74
15 - 20 40 114
20 - 25 26 140
25 - 30 16 156

1st Percentile position/ Class:

𝑖 × 𝑁 𝑡ℎ
𝑃𝑖 = ( ) 𝑖𝑡𝑒𝑚
100
61 × 156 𝑡ℎ
𝑃61 = ( ) 𝑖𝑡𝑒𝑚 = (95.16)𝑡ℎ 𝑖𝑡𝑒𝑚
100

Checking the cumulative frequency column, 95.16th observation fall in 4th class, hence
The 61th percentile lies in the group 15 – 20

The 61th percentile lies in the group 15-20

STA 131: INTRDUCTION TO STATISTICAL INFERENCE BY DR. R. B. AFOLAYAN 10


𝑁
(61 × 100) − 𝐶𝑓𝑚
𝑃61 = 𝑙1 + × 𝐶
𝑓𝑖

𝑙1 = 15; 𝐶𝑓𝑏 = 74; 𝑓1 = 40; 𝑐 = 5; 𝑁 = 156

156
(61 × 100) − 74
𝑃61 = 15 + ( )×5
40
95.16 − 74
𝑃61 = 15 + ( )×5
40

105.8
𝑃61 = 15 + ( ) = 15 + 2.645
40

𝑃61 = 17.645
NOTE that:

25th percentile P25 = Q1 = First quartile


50th percentile P50 = median = Q2= Second quartile
75th percentile P75 = Q3= Third quartile
10th percentile P10 = D1 =First decile
20th percentile P20 = D2= Second decile, etc.,
Tutorial Questions

1) The middle value of an ordered series is called


a)2nd quartile b) 5th decile
c) 50th percentile d) all the above
Ans: all the above
2) For a set of values, the model value can be
a) Unimodal b) bimodal
c) Trimodal d) All of these
d) Ans: all the above
3) Mode is suitable for qualitative data. True or False?
Ans: True
4) Decile divides the group in to ten equal parts. True or False?
Ans: True
5) Mean is affected by extreme values. True or False?
Ans: True
6) Geometric mean can be calculated for negative values. True or False?
Ans: False

STA 131: INTRDUCTION TO STATISTICAL INFERENCE BY DR. R. B. AFOLAYAN 11


Measures of Shape: Skewness and Kurtosis
The measure of central tendency and measure of dispersion can describe the distribution but
they are not sufficient to describe the nature of the distribution. Measures of Shape as a
descriptive statistic can help us to determine how numbers of data points in a data set are
distributed. It helps us to find the patterns that might be hidden and are easily discernible
once the data is displayed on graphs.
The form of data can be understood by looking at the distribution of data elements across the
space. The distribution can be classified as Symmetrical distribution (e.g., Normal
Distribution) and Asymmetrical Distribution (Skewed Distribution).

Symmetrical Distribution: The distribution of data set is symmetric if it looks the


same to the left and right of the centre point. A normal distribution is a true symmetric
distribution of observed values. When a histogram is constructed on values that are normally
distributed, the shape of columns forms a symmetrical bell shape. This is why this
distribution is also known as a 'normal curve' or 'bell curve'.

The following graph is an example of a normal distribution.

NORMAL DISTRIBUTION CURVE

STA 131: MEASURES OF SHAPE: SKEWNESS AND KURTOSIS


1
BY DR. R. B. AFOLAYAN
Key features of the normal distribution:

i. symmetrical shape or Bell Shape


ii. mode, median and mean are the same and are together in the centre of the
curve. i.e. mode= median=mean
iii. there can only be one mode (i.e. there is only one value which is most
frequently observed)
iv. most of the data are clustered around the centre, while the more extreme
values on either side of the centre become less rare as the distance from the
centre increases (i.e. About 68% of values lie within one standard deviation (σ)
away from the mean; about 95% of the values lie within two standard
deviations; and about 99.7% are within three standard deviations. This is
known as the empirical rule or the 3-sigma rule.
Asymmetric Distribution

If the sample data collected is deliberately biased and includes data points with specific
characteristics, this can cause the distribution to be Asymmetrical. In an asymmetrical
distribution the two sides will not be mirror images of each other.
For this purpose, we use other two statistical measures that compare the shape to the normal
curve called Skewness and Kurtosis. Skewness and Kurtosis are the two important
characteristics of distribution that are studied in descriptive statistics
a). SKEWNESS:
Skewness can be defined as a statistical measure that describes the lack of symmetry or
asymmetry in the probability distribution of a dataset. It quantifies the degree to which the
data deviates from a perfectly symmetrical distribution, such as a normal (bell-shaped)
distribution. Skewness is a valuable statistical term because it provides insight into the shape
and nature of a dataset’s distribution. For example, understanding whether a dataset is
positively or negatively skewed can be important in various fields, including finance,
economics, and data analysis, as it can impact the interpretation of data and the choice of
statistical techniques.
STA 131: MEASURES OF SHAPE: SKEWNESS AND KURTOSIS
2
BY DR. R. B. AFOLAYAN
Types of Skewness
Positive skewness and negative skewness are two different ways that a dataset’s distribution
can deviate from perfect symmetry (a normal distribution). They describe the direction of the
skew or asymmetry in the data.

1. Positive Skewness (Right Skew)

In a positively skewed distribution, the tail on the right side (the larger values) is longer than
the tail on the left side (the smaller values). This means that the majority of data points are
concentrated on the left side of the distribution, and there are some extreme values on the
right side. In the case of a positively skewed dataset, Mean > Median > Mode
Examples of positively skewed data include income distribution (where most people earn a
moderate income, but a few earn extremely high incomes), exam scores (where most students
score in a certain range, but a few score exceptionally high), and stock market returns (where
most days have modest returns, but a few days may have very high returns).

2. Negative Skewness (Left Skew)

In a negatively skewed distribution, the tail on the left side (the smaller values) is longer than
the tail on the right side (the larger values). This implies that most of the data points are
concentrated on the right side of the distribution, with a few extreme values on the left side.
In the case of a negatively skewed dataset, Mean < Median < Mode
For example, the symmetrical and skewed distributions are shown by curves as:

Measurement of Skewness

I. Karl Pearson’s Measure


Karl Pearson’s Measure of Skewness uses the mean, median, and standard deviation of the
given data set to quantify the asymmetry or lack of symmetry in the distribution. It is a

STA 131: MEASURES OF SHAPE: SKEWNESS AND KURTOSIS


3
BY DR. R. B. AFOLAYAN
dimensionless number that provides valuable insights into the shape of a dataset’s
distribution.

1. Karl Pearson’s Coefficient of Skewness

1. With respect to Mean and Mode:


(𝑚𝑒𝑎𝑛 − 𝑚𝑜𝑑𝑒)
𝑆𝑘 =
𝑠𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛(𝜎)

2. If mode is not defined for a distribution, we cannot find Sk. But empirical
relation between mean, median and mode states that, for a moderately symmetrical
distribution, we have Mean - Mode ≈ 3 (Mean - Median). Hence Karl Pearson's
coefficient of skewness is defined in terms of median as:
3(𝑚𝑒𝑎𝑛 − 𝑚𝑒𝑑𝑖𝑎𝑛)
𝑆𝑘 =
𝑠𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛(𝜎)

Coefficient of Karl Pearson’s Measure


a) if 𝑆𝑘 = 0, it indicates a perfectly symmetric distribution (normal
distribution) where the data is evenly balanced on both sides of the
mean.
b) If 𝑆𝑘 > 0, it suggests a positively skewed distribution where the tail on
the right side is longer or fatter, and the majority of data points are
concentrated on the left side of the mean.
c) If 𝑆𝑘 < 0, it indicates a negatively skewed distribution where the tail on
the left side is longer or fatter, and the majority of data points are
concentrated on the right side of the mean.
Example of Karl Pearson’s Measure:
Calculate Pearson’s skewness coefficient for exam scores of 10 random
sample of students in STA 131 class:
85, 88, 92, 94, 96, 98, 100, 100, 100, 100.
Solution:
Step 1: Calculate the mean

STA 131: MEASURES OF SHAPE: SKEWNESS AND KURTOSIS


4
BY DR. R. B. AFOLAYAN
∑ 𝑥 85 + 88 + 92 + 94 + 96 + 98 + 100 + 100 + 100 + 100
𝑀𝑒𝑎𝑛(𝑋̅ ) = =
𝑛 10
∑ 𝑥 953
𝑀𝑒𝑎𝑛(𝑋̅ ) = = = 95.3
𝑛 10
Step 2: Calculate the Median
Since N= 10 ;
Arrange the data in ascending order
85, 88, 92, 94, 96, 98, 100, 100, 100, 100
𝑋𝑛 + 𝑋𝑛+1 𝑋10 + 𝑋10+1 𝑋5 + 𝑋6
2 2 2 2
𝑀𝑒𝑑𝑖𝑎𝑛 = = =
2 2 2
The median is the average of the 5th and 6th values when sorted in
ascending order:
96 + 98 194
𝑀𝑒𝑑𝑖𝑎𝑛 = = = 97
2 2
Step 3: Calculate the sample standard deviation;
The sample variance is given by
1
𝑆2 = ∑(𝑋 − 𝑋̅)2
𝑛−1
And the standard deviation

1
𝑆=√ ∑(𝑋 − 𝑋̅)2
𝑛−1

1 (85 − 97)2 + ⋯ + (100 − 97)2 268.1


2
𝑠 = ̅ 2
∑(𝑋 − 𝑋) = = = 26.79
𝑛−1 10 − 1 9

𝑆 = √26.79 = 5.18
Step 4. Calculate the mode
It is clear from the data set that 100 is the most frequently occurring value in
the data. Hence, mode of given data is 100.
Pearson’s skewness coefficient
1. With respect to Mean and Median:
3(𝑚𝑒𝑎𝑛 − 𝑚𝑒𝑑𝑖𝑎𝑛) 3[95.3 − 97]
𝑆𝑘 = = = −0.9846
𝑠𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛(𝜎) 5.18

STA 131: MEASURES OF SHAPE: SKEWNESS AND KURTOSIS


5
BY DR. R. B. AFOLAYAN
2. With respect to Mean and Mode:
𝑚𝑒𝑎𝑛 − 𝑚𝑜𝑑𝑒 95.3 − 100
𝑆𝑘 = = = −0.9073
𝑠𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛(𝜎) 5.18

2. Bowley’s Measure of Skewness


Bowley’s Skewness Coefficient, named after the British economist Arthur Lyon Bowley, is a
statistical measure used to assess the skewness or asymmetry in a probability distribution.
Unlike some other skewness measures that rely on moments or deviations from the mean,
Bowley’s Skewness Coefficient is based on quartiles. This coefficient provides a simple and
intuitive way to understand the direction and magnitude of skewness in a dataset. Bowley’s
Skewness Coefficient is especially useful when dealing with data that may not follow a
normal distribution or when a robust measure of skewness is required.

𝑄1 + 𝑄3 − 2𝑄2
𝐵𝑠 =
𝑄3 − 𝑄1
Coefficient of Bowley’s Measure
a) If 𝐵𝑠 = 0, the distribution is perfectly symmetric about the mean (no
skewness).
b) If 𝐵𝑠 < 0, the distribution is negatively skewed (left-skewed), meaning the
tail on the left side of the distribution is longer or heavier.
c) If 𝐵𝑠 > 0, the distribution is positively skewed (right-skewed), indicating
that the tail on the right side of the distribution is longer or heavier.

Example of Bowley’s Measure:


Calculate Bowley’s Measure of Skewness for the following dataset
representing the ages of a group of people in a sample:
20, 24, 28, 32, 35, 40, 42, 45, 50.

Solution:
Since N= 9;
Step 1: Calculate the median (Q2)
Arrange the data in ascending order
20, 24, 28, 32, 35, 40, 42, 45, 50.
𝑀𝑒𝑑𝑖𝑎𝑛 = 𝑄2 = 𝑋𝑛+1 = 𝑋5
2

The median is the 5th value when sorted in ascending order:


𝑀𝑒𝑑𝑖𝑎𝑛 = 𝑄2 = 35

STA 131: MEASURES OF SHAPE: SKEWNESS AND KURTOSIS


6
BY DR. R. B. AFOLAYAN
Step 2: Calculate the first quartile (Q1)

(𝑛 + 1) 10
𝑄1 = = = 2.5𝑡ℎ 𝑣𝑎𝑙𝑢𝑒
4 4
𝑄1 = 2𝑛𝑑 + 0.5(3𝑟𝑑 − 2𝑛𝑑 ) = 24 + 0.5(28 − 24)
𝑄1 = 24 + 2 = 26
𝑄1 = 26

Step 3: Calculate the third quartile (Q3)

To find Q3, consider the values to the right of the median: 40, 42, 45, 50.

3(𝑛 + 1) 30
𝑄3 = = = 7.5𝑡ℎ 𝑣𝑎𝑙𝑢𝑒
4 4
𝑄3 = 7𝑡ℎ + 0.5(8𝑡ℎ − 7𝑡ℎ) = 42 + 0.5(45 − 42)
𝑄3 = 42 + 1.5 = 43.5
𝑄3 = 43.5

Step 4: Substitute the above values in the formula

𝑄1 + 𝑄3 − 2𝑄2 26 + 43.5 − 2(35)


𝐵𝑠 = = = −0.02857
𝑄3 − 𝑄1 43.5 − 26

Interpretation: Since B is negative (B < 0), the distribution is negatively


skewed (left-skewed). This means that the tail of the distribution is longer on
the left side, indicating that there may be outliers or high values on the right
side of the data.
3. MEASURE OF SKEWNESS BY METHOD OF MOMENT
The skewness of n observations 𝑥1 , 𝑥2 , … . . , 𝑥𝑛 defined as:
̅) 3
𝑀3 ∑(𝑋 − 𝑋 𝑀3
𝑆𝑘 = 3 = 3
= 3
𝜎 𝑁×𝜎 (√𝑀2 )
Population standard deviation

1
𝜎 = √ ∑(𝑋 − 𝑋̅)2
𝑛

Third Central Moment

STA 131: MEASURES OF SHAPE: SKEWNESS AND KURTOSIS


7
BY DR. R. B. AFOLAYAN
1 ̅ )3
𝑀3 = ∑(𝑋 − 𝑋
𝑛

And the skewness for frequency distribution data is given as:


̅ )3
∑ 𝑓 (𝑋 − 𝑋 𝑀3 𝑀3
𝑆𝑘 = = 3 = 3 ;
𝑁 × 𝜎3 (√𝜎 2 ) (√𝑀2 )

𝑤ℎ𝑒𝑟𝑒 𝑁 = ∑ 𝑓

1
𝜎 = √ ∑ 𝑓(𝑋 − 𝑋̅)2
𝑛

̅ )3
∑ 𝑓(𝑋 − 𝑋
𝑀3 =
∑𝑓
𝑤ℎ𝑒𝑟𝑒 𝑡ℎ𝑒 𝑀2 𝑎𝑛𝑑 𝑀3 𝑎𝑟𝑒 𝑠𝑒𝑐𝑜𝑛𝑑 𝑎𝑛𝑑 𝑡ℎ𝑖𝑟𝑑 𝑐𝑒𝑛𝑡𝑟𝑎𝑙 𝑚𝑜𝑚𝑒𝑛𝑡 𝑎𝑏𝑜𝑢𝑡 𝑡ℎ𝑒 𝑚𝑒𝑎𝑛.
Interpretation of Skewness:
• For skewness values between -0.5 and 0.5, the data exhibit approximate
symmetry.
• Skewness values within the range of -1 and -0.5 (negative skewed) or 0.5 and
1(positive skewed) indicate slightly skewed data distributions.
• Data with skewness values less than -1 (negative skewed) or greater than 1
(positive skewed) are considered highly skewed.

KURTOSIS:
Kurtosis is a statistical measure that quantifies the shape of a probability distribution. It
provides information about the tails and peakedness of the distribution compared to a normal
distribution. The measure of kurtosis is very helpful in the selection of an appropriate
average. For example, for normal distribution, mean is most appropriate; for a leptokurtic
distribution, median is most appropriate; and for platykurtic distribution, the quartile range is
most appropriate.

MEASURE OF KURTOSIS BY METHOD OF MOMENT


The Kurtosis of n observations 𝑥1 , 𝑥2 , … . . , 𝑥𝑛 defined as:
𝑀4 ∑(𝑋 − 𝑋̅)4
𝐾𝑢𝑟𝑡𝑜𝑠𝑖𝑠 = 4 =
𝜎 𝑁 × 𝜎4

STA 131: MEASURES OF SHAPE: SKEWNESS AND KURTOSIS


8
BY DR. R. B. AFOLAYAN
1
𝑠𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛 = 𝜎 = √ ∑(𝑋 − 𝑋̅)2
𝑁
1
𝑀4 = ∑(𝑋 − 𝑋̅ )4
𝑁
And the Kurtosis for frequency distribution data is given as:
𝑀4 𝑀4 ∑ 𝑓(𝑋 − 𝑋̅)4
𝐾𝑢𝑟𝑡𝑜𝑠𝑖𝑠 = 4 = 4 =
𝜎 (√𝑀2 ) 𝑁 × 𝜎4

∑ 𝑓(𝑋 − 𝑋̅ )2
𝑠𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛 = 𝜎 = √
∑𝑓

∑ 𝑓(𝑋 − 𝑋̅)4
𝑀4 =
∑𝑓
Karl Pearson classified curves into three types on the basis of the shape of their peaks. These
are Mesokurtic, leptokurtic and platykurtic. These three types of curves are shown in
figure below:

1. Mesokurtic (Kurtosis = 3)

Mesokurtic is the same as the normal distribution, which means kurtosis is near 0. The value

of Kurtosis is equal or approximately 3. In Mesokurtic, distributions are moderate in breadth,

and curves are a medium peaked height.

2. Leptokurtic (Kurtosis > 3)


Leptokurtic has very long and thick tails, which means there are more chances of outliers.
Positive values of kurtosis indicate that distribution is peaked and possesses thick tails.
STA 131: MEASURES OF SHAPE: SKEWNESS AND KURTOSIS
9
BY DR. R. B. AFOLAYAN
Extremely positive kurtosis indicates a distribution where more numbers are located in the
tails of the distribution instead of around the mean.
3. Platykurtic (Kurtosis < 3)
Platykurtic having a thin tail and stretched around the centre means. Most data points are
present in high proximity to the mean. A platykurtic distribution is flatter (less peaked) when
compared with the normal distribution.
Excess Kurtosis
In statistics and probability theory, researchers use excess kurtosis to compare the kurtosis
coefficient with that of a normal distribution. Excess kurtosis can be positive (Leptokurtic
distribution), negative (Platykurtic distribution), or near zero (Mesokurtic distribution). Since
normal distributions have a kurtosis of 3, excess kurtosis is calculated by subtracting 3 from
the estimate of kurtosis.
Excess kurtosis = Kurtosis – 3
Types of Excess Kurtosis
1. Leptokurtic or heavy-tailed distribution (kurtosis more than normal distribution).
2. Mesokurtic (kurtosis same as the normal distribution).
3. Platykurtic or short-tailed distribution (kurtosis less than normal distribution).

Key Differences Between Skewness and Kurtosis


This is the fundamental differences between skewness and kurtosis:

1- The characteristic of a frequency distribution that ascertains its symmetry about the mean is
called skewness. On the other hand, Kurtosis means the relative pointedness of the standard
bell curve, defined by the frequency distribution.
2- Skewness is a measure of the degree of lop-sidedness in the frequency distribution.
Conversely, kurtosis is a measure of degree of tailedness in the frequency distribution.
3- Skewness is an indicator of lack of symmetry, i.e. both left and right sides of the curve are
unequal, with respect to the central point. As against this, kurtosis is a measure of data, that
is either peaked or flat, with respect to the probability distribution.
4- Skewness shows how much and in which direction, the values deviate from the mean? In
contrast, kurtosis explain how tall and sharp the central peak is.

STA 131: MEASURES OF SHAPE: SKEWNESS AND KURTOSIS


10
BY DR. R. B. AFOLAYAN
Examples: Calculate Skewness and Kurtosis from the following grouped data.
Class Interval 0 – 10 10 – 20 20 – 30 30 – 40 40 - 50 Total

Frequency 5 7 15 28 8 63

Solution:
class Mid-Point 𝐟𝐗 (𝐗 (𝐗 − 𝟐𝟗. 𝟐𝟗)𝟐
interval Frequency(f) (𝑿) − 𝟐𝟗. 𝟐𝟗)
0 - 10 5 5 25 -24.29 590.0041
10 - 20 7 15 105 -14.29 204.2041
20 - 30 15 25 375 -4.29 18.4041
30 - 40 28 35 980 5.71 32.6041
40 - 50 8 45 360 15.71 246.8041
Total ∑ 𝑓𝑋 =63 ------- ∑ 𝑓𝑋 -------- ∑(𝑋 − 𝑋̅)2
= 1845;
= 1092.0205
Calculate the Mean
∑ 𝑓𝑋 1845
𝑀𝑒𝑎𝑛(𝑋̅) = = = 29.29
∑𝑓 63

𝑓(𝑋 − 𝑋̅)2 𝑓(𝑋 − 𝑋̅)3 𝑓(𝑋 − 𝑋̅)4


2950.021 -71655.998 1740524
1429.429 -20426.536 291895.2
276.0615 -1184.3038 5080.663
912.9148 5212.7435 29764.77
1974.433 31018.339 487298.1

∑ 𝑓(𝑋 − 𝑋̅)2 ∑ 𝑓(𝑋 − 𝑋̅)3 ∑ 𝑓(𝑋 − 𝑋̅)4


= −57035.755;
= 7542.858 = 2554563

STA 131: MEASURES OF SHAPE: SKEWNESS AND KURTOSIS


11
BY DR. R. B. AFOLAYAN
Calculate the Population Standard deviation

∑ 𝑓(𝑋 − 𝑋̅)2 7542.858


𝑠𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛 = 𝜎 = √ =√
∑𝑓 63

𝑠𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛 = √119.7279 = 10.94


Calculate the third moment, 𝑀3
∑ 𝑓(𝑋 − 𝑋̅)3 −57035.755
𝑀3 = = = −905.33
∑𝑓 63
Calculate the forth moment, 𝑀4
∑ 𝑓(𝑋 − 𝑋̅)4 2554563
𝑀4 = = = 40,548.62
∑𝑓 63
Now, 𝜎 = 10.94 ; 𝑀3 = −905.33 𝑎𝑛𝑑 𝑀4 = 40,548.62
Calculate Skewness
𝑀3 −905.33 −905.33
𝑆𝑘 = = = = −0.69
𝜎 3 (10.94)3 1309.3386
The data is slightly negatively skewed
Calculate Kurtosis
𝑀4 40,548.62 40,548.62
𝐾𝑢𝑟𝑡𝑜𝑠𝑖𝑠 = = = = 2.83
𝜎4 (10.94)4 14324.1641
Since Kurtosis is less than 3, the distribution is platykurtic with thin tails and less peaked.
The skewness and Kurtosis for sample data is defined as:

𝑀3 ∑(𝑋 − 𝑋̅)3
𝑆𝑘𝑒𝑤𝑛𝑒𝑠𝑠 = 3 =
𝜎 (𝑛 − 1) × 𝑆 3

𝑀4 ∑(𝑋 − 𝑋̅)4
𝐾𝑢𝑟𝑡𝑜𝑠𝑖𝑠 = 4 =
𝜎 (𝑛 − 1) × 𝑆 4
Where S is the sample standard deviation and its defined as:

∑(𝑋 − 𝑋̅)2
𝑆= √
𝑛−1

STA 131: MEASURES OF SHAPE: SKEWNESS AND KURTOSIS


12
BY DR. R. B. AFOLAYAN
TUTORIAL QUESTIONS

1. Given that 𝑥1 , 𝑥2 , … . . , 𝑥𝑛 is a random sample of size n is summarized as


follows:
𝟖 𝟖 𝟖
𝟐 ̅ )𝟐 = 𝟒𝟑. 𝟖𝟕𝟓
∑ 𝑿 = 𝟐𝟎𝟓 ; ∑ 𝑿 = 𝟓𝟐𝟗𝟕; ∑(𝑿 − 𝑿
𝒊=𝟏 𝒊=𝟏 𝒊=𝟏
𝟖 𝟖
̅ )𝟑 = 𝟓𝟖. 𝟏𝟐𝟑𝟏 ; ∑(𝑿 − 𝑿
∑(𝑿 − 𝑿 ̅ )𝟒 = 𝟓𝟎𝟐. 𝟖𝟐𝟓
𝒊=𝟏 𝒊=𝟏

2. The following measures were computed for a frequency distribution:


Mean = 50, coefficient of Variation = 35% and Karl Pearson's Coefficient of
Skewness = - 0.25. Compute Standard Deviation, Mode and Median of the
distribution.
3. . The data below represent the summaries of distributions of students mark in an
exam. Compute the moment coefficient of skewness and kurtosis and interpret the
results.
∑ 𝑓 = 50 ; ∑ 𝑓𝑥 = 2010; ∑ 𝑓(𝑋 − 𝑋̅)2 = 5898

∑ 𝑓(𝑋 − 𝑋̅)3 = 21460.8 ; ∑ 𝑓(𝑋 − 𝑋̅)4 = 2491416


4. In an experiment, the height of 100 seedlings were measured to the nearest centimetre and
the results were recorded as below:
Heights(cm) 20 – 24 25- 29 30 - 34 35 - 39 40 - 44 45 - 49
No of 3 19 25 20 18 15
Plants
Compute the Karl Pearson’s coefficient of Skewness and interpret the result.
5. Compute Bowley’s coefficient of Skewness of the following data set
47, 53, 62, 71, 83, 21, 43, 47, 41.
6. Compute Karl Pearson’s coefficient of Skewness of the following data set
36, 44, 86, 31, 37, 44, 86, 35, 60, 51
7. . The rainfall recorded in various places of five districts in a week are given below.

Compute the Bowley’s coefficient of Skewness for rainfall distribution data.


8. The first three moment of a distribution about the value 3 of a variable are 2,10 and 30
respectively. Obtain Mean, 𝑀2 , 𝑀3 and Skewness. Comment upon the nature of
skewness.

STA 131: MEASURES OF SHAPE: SKEWNESS AND KURTOSIS


13
BY DR. R. B. AFOLAYAN

You might also like