0% found this document useful (0 votes)
4 views58 pages

Unit 2

The document provides an overview of descriptive data analysis, focusing on methods for summarizing survey or experimental data. It discusses measures of central tendency, dispersion, position, and shape, along with examples and applications. Additionally, it covers exploratory data analysis and the construction of boxplots to visually represent data distributions.

Uploaded by

mikeofosu280
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views58 pages

Unit 2

The document provides an overview of descriptive data analysis, focusing on methods for summarizing survey or experimental data. It discusses measures of central tendency, dispersion, position, and shape, along with examples and applications. Additionally, it covers exploratory data analysis and the construction of boxplots to visually represent data distributions.

Uploaded by

mikeofosu280
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

STATISTICS: MATH 353

By
Dr. Samuel Asante Gyamerah

Department of Statistics and Actuarial Science


KNUST

January 24, 2023

1 / 58
Descriptive Data Analysis

The observations (data) obtained from a survey or


experiment may represent a sample selected from a
population.
These observations are usually too many to gain an insight
into the nature of the acquired information.
Generally making it impossible for one to convey much
information about the characteristics of the population
under study.
This session basically gives the various descriptive methods for
summarizing survey or experimental data.

2 / 58
INTRODUCTION

Raw Data
The data below represent the highest temperature recorded on
a particular day by a remote sensor in 50 countries.

112, 100, 127, 120, 134, 118, 105, 110, 109, 112, 110, 118, 117,
116, 118, 112, 114, 114, 105, 109, 107, 112, 114, 115, 118, 117,
118, 122, 106, 110, 116, 108, 110, 121, 113, 120, 119, 111, 104,
111, 120, 113, 120, 117, 105, 110, 118, 112, 114, 114

What can you say about this data?

3 / 58
MEASURES OF CENTRAL TENDENCY

Measures of Central Tendency measures, often referred to as


averages, describe the centre of any given data set. They are
useful measures to summarize a given data set.
Examples of measures of central tendency
The Mean
Arithmetic mean
Weighted mean
Trimmed mean
Geometric mean
Harmonic mean
The Median
The Mode

4 / 58
MEASURES OF CENTRAL TENDENCY

5 / 58
MEASURES OF CENTRAL TENDENCY

6 / 58
MEASURES OF CENTRAL TENDENCY

7 / 58
MEASURES OF CENTRAL TENDENCY

8 / 58
MEASURES OF CENTRAL TENDENCY

9 / 58
MEASURES OF CENTRAL TENDENCY

10 / 58
MEASURES OF CENTRAL TENDENCY

11 / 58
MEASURES OF CENTRAL TENDENCY

12 / 58
MEASURES OF CENTRAL TENDENCY

13 / 58
MEASURES OF CENTRAL TENDENCY

14 / 58
MEASURES OF CENTRAL TENDENCY

15 / 58
MEASURES OF CENTRAL TENDENCY

16 / 58
MEASURES OF CENTRAL TENDENCY

The Median - Example

17 / 58
MEASURES OF CENTRAL TENDENCY
The Median - Example

18 / 58
MEASURES OF CENTRAL TENDENCY

19 / 58
MEASURES OF CENTRAL TENDENCY

The Median - Example

20 / 58
MEASURES OF CENTRAL TENDENCY

The Median - Example

21 / 58
MEASURES OF CENTRAL TENDENCY

22 / 58
MEASURES OF CENTRAL TENDENCY

23 / 58
MEASURES OF CENTRAL TENDENCY

Measures of Dispersion
Measures of Dispersion describe the spread or variability in a
data set.

Measures of dispersion include the following: range, mean


deviation, variance, standard deviation and coefficient of
variation. 24 / 58
MEASURES OF DISPERSION
The Range
Range = Largest value Smallest value

The following represents the minimum temperature recorded by


a remote sensor for 25 countries in Europe.

-8.1 3.2 5.9 8.1 12.3


-5.1 4.1 6.3 9.2 13.3
-3.1 4.6 7.9 9.5 14.0
-1.4 4.8 7.9 9.7 15.0
1.2 5.7 8.0 10.3 22.1

Highest value: 22.1 Lowest value: -8.1

Range = Highest valuelowest value


= 22.1 − (−8.1) = 30.2 25 / 58
MEASURES OF DISPERSION

26 / 58
MEASURES OF DISPERSION

The Mean Deviation - Example

27 / 58
MEASURES OF DISPERSION

28 / 58
MEASURES OF DISPERSION


σ= σ2
29 / 58
MEASURES OF DISPERSION
Variance and Standard Deviation: Example

30 / 58
MEASURES OF DISPERSION

31 / 58
MEASURES OF DISPERSION
The Variance and Standard Deviation: Example

32 / 58
MEASURES OF DISPERSION

The Coefficient of Variation

33 / 58
MEASURES OF DISPERSION

The Coefficient of Variation

34 / 58
MEASURES OF POSITION
Measures of Position describe the location (position) of a particu-
lar value in a given distribution of data. The position is described
by quartiles, deciles and percentiles.
Quartiles split the ranked data into 4 segments with an
equal number of values per segment

The first quartile, Q1 , is the value for which 25% of the


observations are smaller and 75% are larger
Q2 is the same as the median (50% are smaller, 50% are
larger)
Only 25% of the observations are greater than the third
quartile
35 / 58
MEASURES OF POSITION
The Quartiles
The quartiles are found by determining the value in the
appropriate position in the ranked data, where
First quartile position: Q1 = (n + 1)/4
Second quartile position: Q2 = (n + 1)/2 (the median position)
Third quartile position: Q3 = 3(n + 1)/4
Interquartile range: IQR = (Q3 − Q1 )
Semi Interquartile range: Semi-IQR = (Q3 − Q1 )/2
where n is the number of observed values

Percentiles
Percentiles divide the data into 100 equal parts.
25th percentile= Q1
50th percentile= Q2
75th percentile= Q3
36 / 58
MEASURES OF POSITION

The Quartiles: Example

37 / 58
MEASURES OF POSITION
The Quartiles: Example
Find the other quartiles

The Decile
Deciles are the values that divide the set of data into ten equal
[Link] fifth decile is the median 38 / 58
MEASURES OF POSITION

Measures of position for grouped data


Measures of position for grouped data can be determined by:
Interpolation method
The formula for the kth percentile for a grouped data is given by
ck
Pk = lk + (nk − Fk )
fk
where
K= Percentile (0.1,0.2,)
lk = lower class boundary of the class in which the k th
percentile lies
Ck = the class width of the k th percentile class boundary
Fk = the cumulative frequency just before the k th
percentile class boundary
fk = the frequency of the k th percentile class boundary
39 / 58
MEASURES OF POSITION

Measures of position for grouped data: Example


Find the 10th , 45th and 90th percentile of the data

Length(mm) Frequency Cum. Frequency


118-126 3 3
127-135 5 8
136-144 9 17
145-153 12 29
154-162 5 34
163-171 4 38
172-180 2 40
Total 40 -

40 / 58
MEASURES OF POSITION

Measures of position for grouped data: Example


Solution
The k th percentile can be calculated by
ck
Pk = lk + (nk − Fk )
fk
C0.1
P0.1 = l0.1 + (0.1 × 40 − F0.1 ) =
f0.1
126.5 + 95 (0.1 × 40 − 3) = 128.3
C0.45
P0.45 = l0.45 + (0.45 × 40 − F0.45 ) =
f0.45
9
144.5 + 12 (0.45 × 40 − 17) = 144.5
C0.9
P0.9 = l0.9 + (0.9 × 40 − F0.9 ) =
f0.9
9
162.5 + 4 (0.9 × 40 − 34) = 167

41 / 58
MEASURES OF SHAPE
Measures of shape determine whether the distribution of data
exhibits a symmetric pattern or stretch out in a particular direc-
tion. Two of such measures of shape are the
Skewness
Kurtosis/Peakness

NOTE: − 3 ≤ γ ≥ 3
42 / 58
MEASURES OF SHAPE

43 / 58
MEASURES OF SHAPE

44 / 58
MEASURES OF SHAPE

45 / 58
MEASURES OF SHAPE

46 / 58
MEASURES OF SHAPE

Skewness: Example
The lengths of stay on the cancer floor of a hospital were
organised into a frequency distribution. The mean length of
stay was 28 days, the median, 25 days and mode, 23 days. The
standard deviation was computed to be 4.2 days. Compute the
skewness.
3(mean − median) 3(28 − 25)
Sk = = = 2.14
standard deviation 4.2

Kurtosis/Peakness
The degree of peakness or kurtosis of a distribution is described
by the coefficient of kurtosis expressed as
1
− Q1 )
2 (Q3
k=
P90 − P10
47 / 58
MEASURES OF SHAPE
Kurtosis/Peakness
If k=3 The distribution is said to be symmetrical or
normal
If k<3 The distribution flattens at the centre than the
normal distribution
If k>3 the distribution is more peaked than the normal
distribution

Graphs of Distributions indicating their Peakness

48 / 58
ASSIGNMENT

The data below show the age distribution of cases of malaria


reported during year at a hospital.

34 17 25 37 19 19 27 19 44 24 24 22 32 12 13 16 18 14 12 16 14
17 10 16 22 20 15 15 10 10 14 17 20 18 13 32 13 13 18 30 24 34
44 31 43 40 28 31 15 22 15 31 18 27 35 35 20 32 38 32
Organize the data into a frequency distribution table
Calculate the coefficient of skewness and kurtosis and
interpret your results.

49 / 58
Exploratory Data Analysis (EDA)

The purpose of exploratory data analysis is to examine


data to find out what information can be discovered about
the data such as the center and the spread.

The measure of central tendency used in EDA is the


median. The measure of variation used in EDA is the
interquartile range Q3 − Q1 .

In EDA the data are represented graphically using a


boxplot (sometimes called a box-and-whisker plot).

50 / 58
The Five-Number Summary and Boxplots

A boxplot can be used to graphically represent the data set.


These plots involve five specific values:
1 The lowest value of the data set (i.e., minimum)
2 Q1
3 The median
4 Q3
5 The highest value of the data set (i.e., maximum)

Boxplot
A boxplot is a graph of a data set obtained by drawing a
horizontal line from the minimum data value to Q1 , drawing a
horizontal line from Q3 to the maximum data value, and
drawing a box whose vertical sides pass through Q1 and Q3 with
a vertical line inside the box passing through the median or Q2 .

51 / 58
Box-plot

52 / 58
Box-plot

53 / 58
Procedure for constructing a boxplot

Find the five-number summary for the data values, that is,
the maximum and minimum data values, Q1 and Q3, and
the median.

Draw a horizontal axis with a scale such that it includes


the maximum and minimum data values.

Draw a box whose vertical sides go through Q1 and Q3,


and draw a vertical line though the median.

Draw a line from the minimum data value to the left side
of the box and a line from the maximum data value to the
right side of the box.

54 / 58
Box-plot

55 / 58
Information Obtained from a Boxplot

If the median is near the center of the box, the distribution


is approximately symmetric.

If the median falls to the left of the center of the box, the
distribution is positively skewed.

If the median falls to the right of the center, the


distribution is negatively skewed.

If the lines are about the same length, the distribution is


approximately symmetric.

If the right line is larger than the left line, the distribution
is positively skewed.

If the left line is larger than the right line, the distribution
is negatively skewed.

56 / 58
Box-plot

57 / 58
Thank You.

58 / 58

You might also like