STATISTICS: MATH 353
By
Dr. Samuel Asante Gyamerah
Department of Statistics and Actuarial Science
KNUST
January 24, 2023
1 / 58
Descriptive Data Analysis
The observations (data) obtained from a survey or
experiment may represent a sample selected from a
population.
These observations are usually too many to gain an insight
into the nature of the acquired information.
Generally making it impossible for one to convey much
information about the characteristics of the population
under study.
This session basically gives the various descriptive methods for
summarizing survey or experimental data.
2 / 58
INTRODUCTION
Raw Data
The data below represent the highest temperature recorded on
a particular day by a remote sensor in 50 countries.
112, 100, 127, 120, 134, 118, 105, 110, 109, 112, 110, 118, 117,
116, 118, 112, 114, 114, 105, 109, 107, 112, 114, 115, 118, 117,
118, 122, 106, 110, 116, 108, 110, 121, 113, 120, 119, 111, 104,
111, 120, 113, 120, 117, 105, 110, 118, 112, 114, 114
What can you say about this data?
3 / 58
MEASURES OF CENTRAL TENDENCY
Measures of Central Tendency measures, often referred to as
averages, describe the centre of any given data set. They are
useful measures to summarize a given data set.
Examples of measures of central tendency
The Mean
Arithmetic mean
Weighted mean
Trimmed mean
Geometric mean
Harmonic mean
The Median
The Mode
4 / 58
MEASURES OF CENTRAL TENDENCY
5 / 58
MEASURES OF CENTRAL TENDENCY
6 / 58
MEASURES OF CENTRAL TENDENCY
7 / 58
MEASURES OF CENTRAL TENDENCY
8 / 58
MEASURES OF CENTRAL TENDENCY
9 / 58
MEASURES OF CENTRAL TENDENCY
10 / 58
MEASURES OF CENTRAL TENDENCY
11 / 58
MEASURES OF CENTRAL TENDENCY
12 / 58
MEASURES OF CENTRAL TENDENCY
13 / 58
MEASURES OF CENTRAL TENDENCY
14 / 58
MEASURES OF CENTRAL TENDENCY
15 / 58
MEASURES OF CENTRAL TENDENCY
16 / 58
MEASURES OF CENTRAL TENDENCY
The Median - Example
17 / 58
MEASURES OF CENTRAL TENDENCY
The Median - Example
18 / 58
MEASURES OF CENTRAL TENDENCY
19 / 58
MEASURES OF CENTRAL TENDENCY
The Median - Example
20 / 58
MEASURES OF CENTRAL TENDENCY
The Median - Example
21 / 58
MEASURES OF CENTRAL TENDENCY
22 / 58
MEASURES OF CENTRAL TENDENCY
23 / 58
MEASURES OF CENTRAL TENDENCY
Measures of Dispersion
Measures of Dispersion describe the spread or variability in a
data set.
Measures of dispersion include the following: range, mean
deviation, variance, standard deviation and coefficient of
variation. 24 / 58
MEASURES OF DISPERSION
The Range
Range = Largest value Smallest value
The following represents the minimum temperature recorded by
a remote sensor for 25 countries in Europe.
-8.1 3.2 5.9 8.1 12.3
-5.1 4.1 6.3 9.2 13.3
-3.1 4.6 7.9 9.5 14.0
-1.4 4.8 7.9 9.7 15.0
1.2 5.7 8.0 10.3 22.1
Highest value: 22.1 Lowest value: -8.1
Range = Highest valuelowest value
= 22.1 − (−8.1) = 30.2 25 / 58
MEASURES OF DISPERSION
26 / 58
MEASURES OF DISPERSION
The Mean Deviation - Example
27 / 58
MEASURES OF DISPERSION
28 / 58
MEASURES OF DISPERSION
√
σ= σ2
29 / 58
MEASURES OF DISPERSION
Variance and Standard Deviation: Example
30 / 58
MEASURES OF DISPERSION
31 / 58
MEASURES OF DISPERSION
The Variance and Standard Deviation: Example
32 / 58
MEASURES OF DISPERSION
The Coefficient of Variation
33 / 58
MEASURES OF DISPERSION
The Coefficient of Variation
34 / 58
MEASURES OF POSITION
Measures of Position describe the location (position) of a particu-
lar value in a given distribution of data. The position is described
by quartiles, deciles and percentiles.
Quartiles split the ranked data into 4 segments with an
equal number of values per segment
The first quartile, Q1 , is the value for which 25% of the
observations are smaller and 75% are larger
Q2 is the same as the median (50% are smaller, 50% are
larger)
Only 25% of the observations are greater than the third
quartile
35 / 58
MEASURES OF POSITION
The Quartiles
The quartiles are found by determining the value in the
appropriate position in the ranked data, where
First quartile position: Q1 = (n + 1)/4
Second quartile position: Q2 = (n + 1)/2 (the median position)
Third quartile position: Q3 = 3(n + 1)/4
Interquartile range: IQR = (Q3 − Q1 )
Semi Interquartile range: Semi-IQR = (Q3 − Q1 )/2
where n is the number of observed values
Percentiles
Percentiles divide the data into 100 equal parts.
25th percentile= Q1
50th percentile= Q2
75th percentile= Q3
36 / 58
MEASURES OF POSITION
The Quartiles: Example
37 / 58
MEASURES OF POSITION
The Quartiles: Example
Find the other quartiles
The Decile
Deciles are the values that divide the set of data into ten equal
[Link] fifth decile is the median 38 / 58
MEASURES OF POSITION
Measures of position for grouped data
Measures of position for grouped data can be determined by:
Interpolation method
The formula for the kth percentile for a grouped data is given by
ck
Pk = lk + (nk − Fk )
fk
where
K= Percentile (0.1,0.2,)
lk = lower class boundary of the class in which the k th
percentile lies
Ck = the class width of the k th percentile class boundary
Fk = the cumulative frequency just before the k th
percentile class boundary
fk = the frequency of the k th percentile class boundary
39 / 58
MEASURES OF POSITION
Measures of position for grouped data: Example
Find the 10th , 45th and 90th percentile of the data
Length(mm) Frequency Cum. Frequency
118-126 3 3
127-135 5 8
136-144 9 17
145-153 12 29
154-162 5 34
163-171 4 38
172-180 2 40
Total 40 -
40 / 58
MEASURES OF POSITION
Measures of position for grouped data: Example
Solution
The k th percentile can be calculated by
ck
Pk = lk + (nk − Fk )
fk
C0.1
P0.1 = l0.1 + (0.1 × 40 − F0.1 ) =
f0.1
126.5 + 95 (0.1 × 40 − 3) = 128.3
C0.45
P0.45 = l0.45 + (0.45 × 40 − F0.45 ) =
f0.45
9
144.5 + 12 (0.45 × 40 − 17) = 144.5
C0.9
P0.9 = l0.9 + (0.9 × 40 − F0.9 ) =
f0.9
9
162.5 + 4 (0.9 × 40 − 34) = 167
41 / 58
MEASURES OF SHAPE
Measures of shape determine whether the distribution of data
exhibits a symmetric pattern or stretch out in a particular direc-
tion. Two of such measures of shape are the
Skewness
Kurtosis/Peakness
NOTE: − 3 ≤ γ ≥ 3
42 / 58
MEASURES OF SHAPE
43 / 58
MEASURES OF SHAPE
44 / 58
MEASURES OF SHAPE
45 / 58
MEASURES OF SHAPE
46 / 58
MEASURES OF SHAPE
Skewness: Example
The lengths of stay on the cancer floor of a hospital were
organised into a frequency distribution. The mean length of
stay was 28 days, the median, 25 days and mode, 23 days. The
standard deviation was computed to be 4.2 days. Compute the
skewness.
3(mean − median) 3(28 − 25)
Sk = = = 2.14
standard deviation 4.2
Kurtosis/Peakness
The degree of peakness or kurtosis of a distribution is described
by the coefficient of kurtosis expressed as
1
− Q1 )
2 (Q3
k=
P90 − P10
47 / 58
MEASURES OF SHAPE
Kurtosis/Peakness
If k=3 The distribution is said to be symmetrical or
normal
If k<3 The distribution flattens at the centre than the
normal distribution
If k>3 the distribution is more peaked than the normal
distribution
Graphs of Distributions indicating their Peakness
48 / 58
ASSIGNMENT
The data below show the age distribution of cases of malaria
reported during year at a hospital.
34 17 25 37 19 19 27 19 44 24 24 22 32 12 13 16 18 14 12 16 14
17 10 16 22 20 15 15 10 10 14 17 20 18 13 32 13 13 18 30 24 34
44 31 43 40 28 31 15 22 15 31 18 27 35 35 20 32 38 32
Organize the data into a frequency distribution table
Calculate the coefficient of skewness and kurtosis and
interpret your results.
49 / 58
Exploratory Data Analysis (EDA)
The purpose of exploratory data analysis is to examine
data to find out what information can be discovered about
the data such as the center and the spread.
The measure of central tendency used in EDA is the
median. The measure of variation used in EDA is the
interquartile range Q3 − Q1 .
In EDA the data are represented graphically using a
boxplot (sometimes called a box-and-whisker plot).
50 / 58
The Five-Number Summary and Boxplots
A boxplot can be used to graphically represent the data set.
These plots involve five specific values:
1 The lowest value of the data set (i.e., minimum)
2 Q1
3 The median
4 Q3
5 The highest value of the data set (i.e., maximum)
Boxplot
A boxplot is a graph of a data set obtained by drawing a
horizontal line from the minimum data value to Q1 , drawing a
horizontal line from Q3 to the maximum data value, and
drawing a box whose vertical sides pass through Q1 and Q3 with
a vertical line inside the box passing through the median or Q2 .
51 / 58
Box-plot
52 / 58
Box-plot
53 / 58
Procedure for constructing a boxplot
Find the five-number summary for the data values, that is,
the maximum and minimum data values, Q1 and Q3, and
the median.
Draw a horizontal axis with a scale such that it includes
the maximum and minimum data values.
Draw a box whose vertical sides go through Q1 and Q3,
and draw a vertical line though the median.
Draw a line from the minimum data value to the left side
of the box and a line from the maximum data value to the
right side of the box.
54 / 58
Box-plot
55 / 58
Information Obtained from a Boxplot
If the median is near the center of the box, the distribution
is approximately symmetric.
If the median falls to the left of the center of the box, the
distribution is positively skewed.
If the median falls to the right of the center, the
distribution is negatively skewed.
If the lines are about the same length, the distribution is
approximately symmetric.
If the right line is larger than the left line, the distribution
is positively skewed.
If the left line is larger than the right line, the distribution
is negatively skewed.
56 / 58
Box-plot
57 / 58
Thank You.
58 / 58