0% found this document useful (0 votes)
15 views69 pages

Data Types and Measures Overview

The document outlines the contents of STA 111, covering chapters 1-6, which include topics on data types, measures of location, measures of dispersion, and measures of partition. Each chapter provides detailed explanations, advantages, disadvantages, and examples of various statistical concepts such as mean, median, mode, range, variance, and skewness. Additionally, it includes practice questions to reinforce understanding of the material.

Uploaded by

danieloluwanjoba
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
15 views69 pages

Data Types and Measures Overview

The document outlines the contents of STA 111, covering chapters 1-6, which include topics on data types, measures of location, measures of dispersion, and measures of partition. Each chapter provides detailed explanations, advantages, disadvantages, and examples of various statistical concepts such as mean, median, mode, range, variance, and skewness. Additionally, it includes practice questions to reinforce understanding of the material.

Uploaded by

danieloluwanjoba
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

STA111Rev

January 21, 2026

Contents
STA 111 - CHAPTERS 1-3 (PDF-FRIENDLY VERSION) 6
With VISIBLE Answers! . . . . . . . . . . . . . . . . . . . 6

CHAPTER 1: DATA - TYPES AND REPRESENTATION 6


What Even is Data? . . . . . . . . . . . . . . . . . . . . . . 6
Two Types of Data (Where Did It Come From?) . . . . . 7
PRIMARY DATA . . . . . . . . . . . . . . . . . . . . . . 7
Advantages of Primary Data: . . . . . . . . . . . . . . . 7
Disadvantages of Primary Data: . . . . . . . . . . . . . 7
SECONDARY DATA . . . . . . . . . . . . . . . . . . . . 8
Advantages of Secondary Data: . . . . . . . . . . . . . 8
Disadvantages of Secondary Data: . . . . . . . . . . . . 8
Quick Comparison . . . . . . . . . . . . . . . . . . . . . . . 8
DATA REPRESENTATION . . . . . . . . . . . . . . . . . . . 9
Why Represent Data? . . . . . . . . . . . . . . . . . . . 9
FREQUENCY DISTRIBUTION TABLE . . . . . . . . . . . 9
Ungrouped Frequency Table . . . . . . . . . . . . . . . 9
Grouped Frequency Table . . . . . . . . . . . . . . . . . 9
Key Terms for Grouped Data: . . . . . . . . . . . . . . . 10
TYPES OF CHARTS/GRAPHS . . . . . . . . . . . . . . . . 10
1. Bar Chart . . . . . . . . . . . . . . . . . . . . . . . . . 10
2. Histogram . . . . . . . . . . . . . . . . . . . . . . . . . 11
3. Pie Chart . . . . . . . . . . . . . . . . . . . . . . . . . 11
4. Frequency Polygon . . . . . . . . . . . . . . . . . . . 11
5. Ogive (Cumulative Frequency Curve) . . . . . . . . 11
CUMULATIVE FREQUENCY . . . . . . . . . . . . . . . . . 11

1
MCQ Practice - Data & Representation . . . . . . . . . . 12

CHAPTER 2: MEASURES OF LOCATION (CENTRAL


TENDENCY) 14
What Are Measures of Location? . . . . . . . . . . . . . . 14
The Three Amigos: Mean, Median, Mode . . . . . . . . . 14
1 MEAN (Average) . . . . . . . . . . . . . . . . . . . . 14
Mean for UNGROUPED Data . . . . . . . . . . . . . . . 14
Mean for UNGROUPED with Frequency . . . . . . . . 15
Mean for GROUPED Data . . . . . . . . . . . . . . . . . 16
Pros of Mean: . . . . . . . . . . . . . . . . . . . . . . . . 16
Cons of Mean: . . . . . . . . . . . . . . . . . . . . . . . . 16
2 MEDIAN (Middle Value) . . . . . . . . . . . . . . . . . . 17
Finding Median - Ungrouped Data . . . . . . . . . . . . 17
Case 1: ODD number of values . . . . . . . . . . . . . . 17
Case 2: EVEN number of values . . . . . . . . . . . . . 17
Median for GROUPED Data . . . . . . . . . . . . . . . . 18
Example - Grouped Median . . . . . . . . . . . . . . . . 18
Pros of Median: . . . . . . . . . . . . . . . . . . . . . . . 19
Cons of Median: . . . . . . . . . . . . . . . . . . . . . . . 19
3 MODE (Most Frequent) . . . . . . . . . . . . . . . . . . 19
Mode for Ungrouped Data . . . . . . . . . . . . . . . . 19
Special Cases: . . . . . . . . . . . . . . . . . . . . . . . . 20
Mode for GROUPED Data . . . . . . . . . . . . . . . . . 20
Example - Grouped Mode . . . . . . . . . . . . . . . . . 20
Pros of Mode: . . . . . . . . . . . . . . . . . . . . . . . . 21
Cons of Mode: . . . . . . . . . . . . . . . . . . . . . . . . 21
Comparison Table . . . . . . . . . . . . . . . . . . . . . . . 21
When to Use What? . . . . . . . . . . . . . . . . . . . . . . 22
Relationship Between Mean, Median, Mode . . . . . . . 22
MCQ Practice - Measures of Location . . . . . . . . . . . 23

CHAPTER 3: MEASURES OF DISPERSION (SPREAD) 25


What is Dispersion? . . . . . . . . . . . . . . . . . . . . . . 25
The Big Picture . . . . . . . . . . . . . . . . . . . . . . . . . 25
Types of Dispersion Measures . . . . . . . . . . . . . . . . 25
1 RANGE . . . . . . . . . . . . . . . . . . . . . . . . . . . . 26
Example: . . . . . . . . . . . . . . . . . . . . . . . . . . . 26
For Grouped Data: . . . . . . . . . . . . . . . . . . . . . 26
Pros of Range: . . . . . . . . . . . . . . . . . . . . . . . . 26

2
Cons of Range: . . . . . . . . . . . . . . . . . . . . . . . 26
2 MEAN DEVIATION (Average Deviation) . . . . . . . . 27
Example - Mean Deviation . . . . . . . . . . . . . . . . 27
For Grouped Data: . . . . . . . . . . . . . . . . . . . . . 28
3 VARIANCE (σ² or s²) . . . . . . . . . . . . . . . . . . . . 28
Population Variance: . . . . . . . . . . . . . . . . . . . . 28
Sample Variance: . . . . . . . . . . . . . . . . . . . . . . 28
Example - Variance . . . . . . . . . . . . . . . . . . . . . 28
Shortcut Formula (VERY USEFUL!): . . . . . . . . . . 29
For Grouped Data: . . . . . . . . . . . . . . . . . . . . . 29
4 STANDARD DEVIATION (σ or s) . . . . . . . . . . . . . 29
Why Standard Deviation? . . . . . . . . . . . . . . . . . 30
Example . . . . . . . . . . . . . . . . . . . . . . . . . . . 30
The 68-95-99.7 Rule (For Normal Distribution) . . . . 30
5 COEFFICIENT OF VARIATION (CV) . . . . . . . . . . . 30
Why CV is Awesome: . . . . . . . . . . . . . . . . . . . . 31
Interpretation of CV: . . . . . . . . . . . . . . . . . . . . 31
Complete Example - All Measures . . . . . . . . . . . . . . 32
Step 1: Basic Calculations . . . . . . . . . . . . . . . . . 32
Step 2: Mean . . . . . . . . . . . . . . . . . . . . . . . . 32
Step 3: Range . . . . . . . . . . . . . . . . . . . . . . . . 32
Step 4: Mean Deviation . . . . . . . . . . . . . . . . . . 32
Step 5: Variance (Shortcut) . . . . . . . . . . . . . . . . 33
Step 6: Standard Deviation . . . . . . . . . . . . . . . . 33
Step 7: Coefficient of Variation . . . . . . . . . . . . . . 33
Summary Table . . . . . . . . . . . . . . . . . . . . . . . . . 33
Absolute vs Relative Measures . . . . . . . . . . . . . . . . . 34
Absolute Measures (Have units): . . . . . . . . . . . . . 34
Relative Measures (No units, %): . . . . . . . . . . . . 34
MCQ Practice - Measures of Dispersion . . . . . . . . . . 34

END OF CHAPTERS 1-3 (PDF-FRIENDLY VERSION) 36

STA 111 - CHAPTERS 4-6 (PDF-FRIENDLY VERSION) 36


With VISIBLE Answers! . . . . . . . . . . . . . . . . . . . 36

CHAPTER 4: MEASURES OF PARTITION (POSITION) 37


What Are Measures of Partition? . . . . . . . . . . . . . . 37
The Three Types . . . . . . . . . . . . . . . . . . . . . . . . 37
1 QUARTILES . . . . . . . . . . . . . . . . . . . . . . . . . 37

3
The Three Quartiles: . . . . . . . . . . . . . . . . . . . . 37
Finding Quartiles - Ungrouped Data . . . . . . . . . . 38
Example - Ungrouped Quartiles: . . . . . . . . . . . . . 38
Quartiles for Grouped Data . . . . . . . . . . . . . . . . 39
Example - Grouped Quartiles . . . . . . . . . . . . . . . 39
Interquartile Range (IQR) . . . . . . . . . . . . . . . . . . 40
Why IQR is Useful: . . . . . . . . . . . . . . . . . . . . . 40
Semi-Interquartile Range (Quartile Deviation): . . . . 40
2 DECILES . . . . . . . . . . . . . . . . . . . . . . . . . . . 41
Finding Deciles . . . . . . . . . . . . . . . . . . . . . . . 41
Example - Decile . . . . . . . . . . . . . . . . . . . . . . 41
3 PERCENTILES . . . . . . . . . . . . . . . . . . . . . . . . 42
Finding Percentiles . . . . . . . . . . . . . . . . . . . . . 42
Example - Percentile . . . . . . . . . . . . . . . . . . . . 42
Relationships Between Measures . . . . . . . . . . . . . . 42
Box Plot (Box-and-Whisker Plot) . . . . . . . . . . . . . . . 43
MCQ Practice - Measures of Partition . . . . . . . . . . . 43

CHAPTER 5: SKEWNESS AND KURTOSIS 45


What Are These? . . . . . . . . . . . . . . . . . . . . . . . . 45

PART 1: SKEWNESS 46
What is Skewness? . . . . . . . . . . . . . . . . . . . . . . . 46
Three Types of Skewness . . . . . . . . . . . . . . . . . . . 46
1. Symmetric (No Skew) - Skewness = 0 . . . . . . . . 46
2. Positive Skew (Right Skewed) - Skewness > 0 . . . 47
3. Negative Skew (Left Skewed) - Skewness < 0 . . . 47
Memory Trick! . . . . . . . . . . . . . . . . . . . . . . . . . 47
Measures of Skewness . . . . . . . . . . . . . . . . . . . . 48
1. Karl Pearson’s Coefficient of Skewness . . . . . . . 48
Interpreting Pearson’s Coefficient: . . . . . . . . . . . 48
Example - Pearson’s Skewness . . . . . . . . . . . . . . 48
2. Bowley’s Coefficient of Skewness (Quartile Method) 49
Example - Bowley’s Skewness . . . . . . . . . . . . . . 49
3. Moment Coefficient of Skewness . . . . . . . . . . . 49

PART 2: KURTOSIS 50
What is Kurtosis? . . . . . . . . . . . . . . . . . . . . . . . . 50
Three Types of Kurtosis . . . . . . . . . . . . . . . . . . . . 50
1. Mesokurtic (Normal) - Kurtosis = 3 . . . . . . . . . 50

4
2. Leptokurtic (Tall & Thin) - Kurtosis > 3 . . . . . . . 50
3. Platykurtic (Short & Wide) - Kurtosis < 3 . . . . . . 51
Memory Trick! . . . . . . . . . . . . . . . . . . . . . . . . . 51
Measures of Kurtosis . . . . . . . . . . . . . . . . . . . . . 51
1. Moment Coefficient of Kurtosis (beta2 or gamma2) 51
Excess Kurtosis: . . . . . . . . . . . . . . . . . . . . . . . 52
2. Percentile Coefficient of Kurtosis . . . . . . . . . . . 52
Summary: Skewness vs Kurtosis . . . . . . . . . . . . . . 52
MCQ Practice - Skewness and Kurtosis . . . . . . . . . . 53

CHAPTER 6: MOMENTS 55
What Are Moments? . . . . . . . . . . . . . . . . . . . . . . 55
Simple Analogy . . . . . . . . . . . . . . . . . . . . . . . . . 56
Two Types of Moments . . . . . . . . . . . . . . . . . . . . 56
1. Moments About the Origin (Raw Moments) . . . . 56
2. Moments About the Mean (Central Moments) . . . 56
The First Four Moments . . . . . . . . . . . . . . . . . . . 57
The Most Important Moments . . . . . . . . . . . . . . . . 57
First Moment About Origin = MEAN . . . . . . . . . . 57
First Moment About Mean = ZERO (Always!) . . . . . 57
Second Moment About Mean = VARIANCE . . . . . . 57
Third Moment About Mean –> SKEWNESS . . . . . . 58
Fourth Moment About Mean –> KURTOSIS . . . . . . 58
Converting Between Moments (SUPER IMPORTANT!) . 58
Complete Example - Finding Moments . . . . . . . . . . . 59
Step 1: Calculate Powers . . . . . . . . . . . . . . . . . 59
Step 2: Raw Moments . . . . . . . . . . . . . . . . . . . 59
Step 3: Central Moments (Using Conversion Formu-
las) . . . . . . . . . . . . . . . . . . . . . . . . . . . 59
Step 4: Skewness and Kurtosis . . . . . . . . . . . . . . 60
Summary: What Each Moment Tells You . . . . . . . . . . 60
Shortcut: The Variance Formula . . . . . . . . . . . . . . . 61
Moment Ratios (The beta and gamma Coefficients) . . . 61
For Skewness: . . . . . . . . . . . . . . . . . . . . . . . . 61
For Kurtosis: . . . . . . . . . . . . . . . . . . . . . . . . . 62
MCQ Practice - Moments . . . . . . . . . . . . . . . . . . . 62

ULTIMATE STA 111 CHEAT SHEET 64


Central Tendency Formulas . . . . . . . . . . . . . . . . . 64
Dispersion Formulas . . . . . . . . . . . . . . . . . . . . . . 65

5
Partition Formulas (Grouped Data) . . . . . . . . . . . . . 65
Skewness Formulas . . . . . . . . . . . . . . . . . . . . . . 65
Kurtosis Formulas . . . . . . . . . . . . . . . . . . . . . . . 66
Moment Conversions . . . . . . . . . . . . . . . . . . . . . 66
Quick Reference Table . . . . . . . . . . . . . . . . . . . . 67
Key Relationships . . . . . . . . . . . . . . . . . . . . . . . 67
Effects of Transformations . . . . . . . . . . . . . . . . . . 67
Memory Tricks . . . . . . . . . . . . . . . . . . . . . . . . . 68

YOU’RE READY FOR STA 111! 68


What You’ve Learned: . . . . . . . . . . . . . . . . . . . . . . 68
Final Tips: . . . . . . . . . . . . . . . . . . . . . . . . . . . . 68

STA 111 - CHAPTERS 1-3 (PDF-FRIENDLY


VERSION)

With VISIBLE Answers!

CHAPTER 1: DATA - TYPES AND REPRESEN-


TATION

What Even is Data?

Data = Facts and figures collected about something


That’s it. Numbers, words, measurements - anything you collect
to study.

6
Two Types of Data (Where Did It Come From?)

PRIMARY DATA

Data YOU collect YOURSELF for the FIRST time


You’re the “original owner” - nobody else collected this before
you!
Examples: - You surveying students about their favorite food
- A company interviewing customers directly - You measuring
heights of your classmates - Conducting experiments in a lab
How to Collect Primary Data:

Method What It Is
Questionnaire Printed questions people answer
Interview Asking questions face-to-face
Observation Watching and recording what happens
Experiment Testing something in controlled conditions

Advantages of Primary Data:

• Fresh and original


• Specific to YOUR needs
• You know it’s accurate (you collected it!)
• Up-to-date

Disadvantages of Primary Data:

• Expensive
• Time-consuming
• Needs more effort
• May need trained people

7
SECONDARY DATA

Data SOMEONE ELSE already collected


You’re just “borrowing” it - it already existed!
Examples: - Government census data - Data from textbooks -
Company annual reports - Newspaper statistics - Research from
other studies

Advantages of Secondary Data:

• Cheap (often free!)


• Quick to get
• Already organized
• Large amounts available

Disadvantages of Secondary Data:

• Might be outdated
• Might not fit YOUR exact needs
• You don’t know if it’s accurate
• Someone else’s errors become yours

Quick Comparison

Feature Primary Data Secondary Data


Who collected? YOU Someone else
Cost High Low
Time Long Short
Accuracy High Unknown
Relevance Exactly what you need May not fit perfectly

8
DATA REPRESENTATION

Why Represent Data?

Raw numbers are BORING and confusing!


Example: Raw: 45, 67, 23, 89, 56, 34, 78, 12, 90, 43, 65, 87, 21,
54, 76…
WHO CAN UNDERSTAND THAT?!
That’s why we use tables and charts!

FREQUENCY DISTRIBUTION TABLE

Organizing data into groups with their counts

Ungrouped Frequency Table

Example: Scores of 20 students: 5, 6, 7, 5, 8, 6, 5, 7, 6, 5, 8, 7,


6, 5, 7, 8, 6, 7, 5, 6

Score (x) Tally Frequency (f)


5 IIII I 6
6 IIII I 6
7 IIII 5
8 III 3
Total 20

Grouped Frequency Table

When data has MANY different values, group them into class in-
tervals!

9
Example: Ages of 30 people

Class Interval Frequency (f)


10 - 19 5
20 - 29 8
30 - 39 10
40 - 49 5
50 - 59 2
Total 30

Key Terms for Grouped Data:

Term Meaning Example (for 20-29)


Lower Class Limit Smallest value in class 20
Upper Class Limit Largest value in class 29
Class Width Upper - Lower + 1 29 - 20 + 1 = 10
Class Mark (Midpoint) (Lower + Upper) / 2 (20 + 29) / 2 = 24.5
Class Boundary Real limits (add/subtract 0.5) 19.5 - 29.5

TYPES OF CHARTS/GRAPHS

1. Bar Chart

• Rectangular bars
• Height = Frequency
• Gaps between bars
• For categorical/discrete data

10
2. Histogram

• Like bar chart BUT


• NO gaps between bars
• For continuous/grouped data
• Area represents frequency

3. Pie Chart

• Circle divided into slices


• Each slice = proportion of total
• Angle = (frequency/total) x 360 degrees
Example: If 25 out of 100 students like Maths: Angle = (25/100)
x 360 = 90 degrees

4. Frequency Polygon

• Line graph connecting midpoints of histogram bars


• Shows distribution shape

5. Ogive (Cumulative Frequency Curve)

• Shows RUNNING TOTAL


• Used to find median, quartiles, percentiles

CUMULATIVE FREQUENCY

Running total - add as you go!

Score Frequency (f) Cumulative Frequency (cf)


5 6 6
6 6 6+6 = 12
7 5 12+5 = 17

11
Score Frequency (f) Cumulative Frequency (cf)
8 3 17+3 = 20

The last CF should equal your total!

MCQ Practice - Data & Representation

Q1. Data collected by a researcher through questionnaires is:


- (A) Secondary data - (B) Primary data - (C) Tertiary data - (D)
Published data
Answer: (B) Primary data

Q2. Census data from the government is an example of: - (A)


Primary data - (B) Secondary data - (C) Sample data - (D) Raw
data
Answer: (B) Secondary data

Q3. In a histogram, bars are drawn: - (A) With gaps between


them - (B) Without gaps - (C) Overlapping - (D) Separately
Answer: (B) Without gaps

Q4. The midpoint of class interval 40-49 is: - (A) 44 - (B) 44.5 -
(C) 45 - (D) 40
Answer: (B) 44.5
Solution: (40 + 49) / 2 = 44.5

Q5. In a pie chart, if a category has 50 out of 200 items, its angle
is: - (A) 50 degrees - (B) 90 degrees - (C) 100 degrees - (D) 25
degrees

12
Answer: (B) 90 degrees
Solution: (50/200) x 360 = 90 degrees

Q6. The class boundaries of 20-29 are: - (A) 20 and 29 - (B) 19.5
and 29.5 - (C) 20.5 and 28.5 - (D) 19 and 30
Answer: (B) 19.5 and 29.5

Q7. Which is NOT a method of collecting primary data? - (A)


Questionnaire - (B) Interview - (C) Textbooks - (D) Observation
Answer: (C) Textbooks

Q8. An ogive is used to find: - (A) Mean - (B) Mode - (C) Median
and quartiles - (D) Range
Answer: (C) Median and quartiles

Q9. The class width of the interval 15-24 is: - (A) 9 - (B) 10 - (C)
15 - (D) 24
Answer: (B) 10
Solution: 24 - 15 + 1 = 10

Q10. Which chart is best for showing parts of a whole? - (A) Bar
chart - (B) Histogram - (C) Pie chart - (D) Line graph
Answer: (C) Pie chart

13
CHAPTER 2: MEASURES OF LOCATION
(CENTRAL TENDENCY)

What Are Measures of Location?

A single number that represents the “center” or “typical


value” of your data
Think of it as: “If I had to describe all this data with ONE number,
what would it be?”

The Three Amigos: Mean, Median, Mode

1 MEAN (Average)

Add everything up, divide by how many


The most common “average” you know!

Sum of all values ∑𝑥


𝑥̄ = =
Number of values 𝑛

Mean for UNGROUPED Data

Example: Find the mean of: 5, 8, 10, 12, 15

5 + 8 + 10 + 12 + 15 50
𝑥̄ = = = 10
5 5
Mean = 10

14
Mean for UNGROUPED with Frequency

15
Score (x) Frequency (f) fx
5 3 15
6 5 30
7 2 14
Total Σf = 10 Σfx = 59

∑ 𝑓𝑥 59
𝑥̄ = = = 5.9
∑𝑓 10

Mean for GROUPED Data

Use the class midpoint (x) as the value!

Class Frequency (f) Midpoint (x) fx


10-19 5 14.5 72.5
20-29 8 24.5 196
30-39 7 34.5 241.5
Total 20 510

510
𝑥̄ = = 25.5
20

Pros of Mean:

• Uses ALL data values


• Good for further calculations
• Most widely used

Cons of Mean:

• Affected by extreme values (outliers)

16
• Can’t use for categorical data

2 MEDIAN (Middle Value)

The value in the MIDDLE when data is arranged in order


Half the data is below, half is above!

Finding Median - Ungrouped Data

Step 1: Arrange data in order (smallest to largest)


Step 2: Find the middle position

𝑛+1
Position =
2

Case 1: ODD number of values

Example: Find median of: 7, 3, 9, 5, 1


Step 1: Arrange: 1, 3, 5, 7, 9
Step 2: Position = (5+1)/2 = 3rd value
Median = 5 (the 3rd value)

Case 2: EVEN number of values

Example: Find median of: 8, 4, 10, 2, 6, 12


Step 1: Arrange: 2, 4, 6, 8, 10, 12

17
Step 2: Position = (6+1)/2 = 3.5
This means: Average of 3rd and 4th values!

6+8
Median = =7
2
Median = 7

Median for GROUPED Data


𝑛 − 𝑐𝑓𝑏
Median = 𝐿 + ( 2 )×𝑐
𝑓𝑚
Where: | Symbol | Meaning | |——–|———| | 𝐿 | Lower boundary
of median class | | 𝑛 | Total frequency | | 𝑐𝑓𝑏 | Cumulative fre-
quency BEFORE median class | | 𝑓𝑚 | Frequency of median class
| | 𝑐 | Class width |

Example - Grouped Median

Class f cf
10-19 5 5
20-29 8 13
30-39 10 23 (Median class)
40-49 5 28
50-59 2 30
Total 30

Step 1: Find median position = n/2 = 30/2 = 15


Step 2: Find median class (where cf first exceeds 15) = 30-39
Step 3: Identify values: - L = 29.5 (lower boundary) - cf_b = 13

18
(cf before median class) - f_m = 10 - c = 10
Step 4: Calculate:
15 − 13
Median = 29.5 + ( ) × 10
10
2
= 29.5 + × 10 = 29.5 + 2 = 31.5
10

Pros of Median:

• NOT affected by extreme values


• Good for skewed data
• Can use for ordinal data

Cons of Median:

• Doesn’t use all values


• Not good for further calculations

3 MODE (Most Frequent)

The value that appears MOST OFTEN


The “popular kid” of the data!

Mode for Ungrouped Data

Example: Find mode of: 5, 7, 5, 8, 9, 5, 7, 5


Count each: - 5 appears 4 times (MOST!) - 7 appears 2 times - 8
appears 1 time - 9 appears 1 time
Mode = 5

19
Special Cases:

Unimodal: One mode (most common)


Bimodal: Two modes - Example: 3, 5, 5, 7, 7, 9 → Modes: 5 and
7
Multimodal: More than two modes
No Mode: All values appear equally - Example: 1, 2, 3, 4, 5 →
No mode

Mode for GROUPED Data

𝑓𝑚 − 𝑓 1
Mode = 𝐿 + ( )×𝑐
2𝑓𝑚 − 𝑓1 − 𝑓2
Where: | Symbol | Meaning | |——–|———| | 𝐿 | Lower boundary
of modal class | | 𝑓𝑚 | Frequency of modal class (highest) | | 𝑓1
| Frequency of class BEFORE modal class | | 𝑓2 | Frequency of
class AFTER modal class | | 𝑐 | Class width |

Example - Grouped Mode

Class f
10-19 5
20-29 8
30-39 12 (Modal class - highest f)
40-49 7
50-59 3

• L = 29.5

20
• f_m = 12
• f_1 = 8
• f_2 = 7
• c = 10

12 − 8
Mode = 29.5 + ( ) × 10
2(12) − 8 − 7
4
= 29.5 + ( ) × 10
24 − 15
4
= 29.5 + × 10
9
= 29.5 + 4.44 = 33.94

Pros of Mode:

• Easy to find
• Not affected by extreme values
• Can use for categorical data
• Represents actual data value

Cons of Mode:

• May not exist


• May have multiple modes
• Not unique

Comparison Table

Feature Mean Median Mode


Uses all values Yes No No
Affected by outliers Yes No No

21
Feature Mean Median Mode
For categorical data No No Yes
Always exists Yes Yes Maybe
Always unique Yes Yes No

When to Use What?

Data Type Best Measure


Normal distribution (symmetric) Mean
Skewed data / has outliers Median
Categorical data Mode
Income, house prices Median (skewed!)
Exam scores (normal) Mean
Favorite color Mode

Relationship Between Mean, Median, Mode

For a symmetric distribution:

Mean = Median = Mode

For a moderately skewed distribution:

Mode = 3(Median) − 2(Mean)

Or rearranged:

Mean − Mode = 3(Mean − Median)

22
MCQ Practice - Measures of Location

Q1. The mean of 4, 6, 8, 10, 12 is: - (A) 6 - (B) 8 - (C) 10 - (D) 40


Answer: (B) 8
Solution: (4+6+8+10+12)/5 = 40/5 = 8

Q2. The median of 7, 3, 5, 9, 1 is: - (A) 3 - (B) 5 - (C) 7 - (D) 9


Answer: (B) 5
Solution: Arrange: 1, 3, 5, 7, 9. Middle = 5

Q3. The mode of 2, 3, 3, 4, 5, 5, 5, 6 is: - (A) 3 - (B) 4 - (C) 5 - (D)


6
Answer: (C) 5
Solution: 5 appears most often (3 times)

Q4. Which measure is NOT affected by extreme values? - (A)


Mean - (B) Median - (C) Both A and B - (D) Neither
Answer: (B) Median

Q5. For grouped data, we use _____ as the representative value:


- (A) Lower limit - (B) Upper limit - (C) Class midpoint - (D) Class
width
Answer: (C) Class midpoint

Q6. If data is: 5, 5, 6, 7, 7, 8, the data is: - (A) Unimodal - (B)


Bimodal - (C) Trimodal - (D) No mode
Answer: (B) Bimodal
Solution: Both 5 and 7 appear twice (most frequent)

23
Q7. The mean of first 5 natural numbers is: - (A) 2 - (B) 3 - (C) 4
- (D) 5
Answer: (B) 3
Solution: (1+2+3+4+5)/5 = 15/5 = 3

Q8. In a positively skewed distribution: - (A) Mean < Median <


Mode - (B) Mode < Median < Mean - (C) Mean = Median = Mode
- (D) Median < Mode < Mean
Answer: (B) Mode < Median < Mean

Q9. If Mean = 50 and Mode = 44, find Median (using empirical


relation): - (A) 46 - (B) 47 - (C) 48 - (D) 49
Answer: (C) 48
Solution: Mode = 3(Median) - 2(Mean) 44 = 3(Median) - 100
3(Median) = 144 Median = 48

Q10. The sum of deviations from mean is always: - (A) Positive -


(B) Negative - (C) Zero - (D) Cannot tell
Answer: (C) Zero

Q11. The median of 2, 4, 6, 8 is: - (A) 4 - (B) 5 - (C) 6 - (D) 7


Answer: (B) 5
Solution: Average of 4 and 6 = (4+6)/2 = 5

Q12. Which measure can be used for qualitative data? - (A) Mean
- (B) Median - (C) Mode - (D) Variance
Answer: (C) Mode

24
CHAPTER 3: MEASURES OF DISPERSION
(SPREAD)

What is Dispersion?

How SPREAD OUT is your data?


Measures of location tell us the CENTER. Measures of dispersion
tell us the SPREAD.

The Big Picture

Two classes can have the SAME average but DIFFERENT


spreads:
Class A scores: 70, 70, 70, 70, 70 → Mean = 70 Class B scores:
50, 60, 70, 80, 90 → Mean = 70
Same mean, but Class B is more SPREAD OUT!

Types of Dispersion Measures

Measure Formula What It Tells You


Range Max - Min Total spread
Mean Deviation Average distance from mean Typical deviation
Variance Average of squared distances Spread (squared units)
Standard Deviation Square root of Variance Spread (original units)
Coefficient of Variation (SD/Mean) x 100% Relative spread

25
1 RANGE

The simplest measure!

Range = Maximum Value − Minimum Value

Example:

Data: 5, 8, 12, 15, 20


Range = 20 - 5 = 15

For Grouped Data:

Range = Upper boundary of last class−Lower boundary of first class

Pros of Range:

• Super easy to calculate


• Quick overview of spread

Cons of Range:

• Only uses 2 values (ignores the rest!)


• Heavily affected by outliers
• Not reliable

26
2 MEAN DEVIATION (Average Deviation)

Average distance of each value from the mean

∑ |𝑥 − 𝑥|̄
𝑀𝐷 =
𝑛
The | | means absolute value (ignore negative signs!)

Example - Mean Deviation

Data: 2, 4, 6, 8, 10
Step 1: Find mean

2 + 4 + 6 + 8 + 10
𝑥̄ = =6
5

Step 2: Find deviations from mean

x x - mean
2 2 - 6 = -4 4
4 4 - 6 = -2 2
6 6-6=0 0
8 8-6=2 2
10 10 - 6 = 4 4
Sum = 12

Step 3: Calculate MD

12
𝑀𝐷 = = 2.4
5

On average, values are 2.4 units away from the mean

27
For Grouped Data:

∑ 𝑓|𝑥 − 𝑥|̄
𝑀𝐷 =
∑𝑓

3 VARIANCE (σ² or s²)

Average of SQUARED distances from the mean

Population Variance:

∑(𝑥 − 𝜇)2
𝜎2 =
𝑁

Sample Variance:

∑(𝑥 − 𝑥)̄ 2
𝑠2 =
𝑛−1
Why n-1? It gives a better estimate of population variance
(called “Bessel’s correction”)

Example - Variance

Data: 2, 4, 6, 8, 10 (Mean = 6)

x x - mean (x - mean)²
2 -4 16
4 -2 4
6 0 0
8 2 4
10 4 16
Sum = 40

28
Population Variance:

40
𝜎2 = =8
5

Sample Variance:

40 40
𝑠2 = = = 10
5−1 4

Shortcut Formula (VERY USEFUL!):

2
∑ 𝑥2 ∑𝑥
Variance = −( )
𝑛 𝑛
Or:
𝜎2 = 𝑥2 − 𝑥2̄

“Mean of squares minus square of mean”

For Grouped Data:

2
∑ 𝑓𝑥2 ∑ 𝑓𝑥
𝜎2 = −( )
∑𝑓 ∑𝑓

4 STANDARD DEVIATION (σ or s)

Square root of variance!

√ √
𝜎= 𝜎2 = Variance

29
Why Standard Deviation?

Variance is in SQUARED units (like cm²) Standard Deviation is in


ORIGINAL units (like cm)
Much easier to interpret!

Example

From our previous example: - Population Variance = 8 - Standard


Deviation = √8 = 2.83
Interpretation: Values typically deviate about 2.83 units from
the mean.

The 68-95-99.7 Rule (For Normal Distribution)

Range Percentage of Data


Mean plus/minus 1 SD about 68%
Mean plus/minus 2 SD about 95%
Mean plus/minus 3 SD about 99.7%

Example: If Mean = 70, SD = 10 - 68% of scores between 60 and


80 - 95% of scores between 50 and 90 - 99.7% of scores between
40 and 100

5 COEFFICIENT OF VARIATION (CV)

Relative measure of dispersion (as percentage)

30
Standard Deviation
𝐶𝑉 = × 100%
Mean

𝜎
𝐶𝑉 = × 100%
𝑥̄

Why CV is Awesome:

It lets you compare spreads of DIFFERENT data sets!


Example:

Height (cm) Weight (kg)


Mean 170 70
SD 10 7
CV (10/170) x 100 = 5.88% (7/70) x 100 = 10%

Weight is MORE variable (relative to its mean)!


You can’t compare 10cm to 7kg directly, but you CAN compare
5.88% to 10%!

Interpretation of CV:

CV Variability
Less than 15% Low
15% - 35% Moderate
More than 35% High

31
Complete Example - All Measures

Data: 10, 20, 30, 40, 50 (n = 5)

Step 1: Basic Calculations

x x²
10 100
20 400
30 900
40 1600
50 2500
Σx = 150 Σx² = 5500

Step 2: Mean

150
𝑥̄ = = 30
5

Step 3: Range

Range = 50 − 10 = 20

Step 4: Mean Deviation

x
10 20
20 10
30 0
40 10
50 20
Sum = 60

32
60
𝑀𝐷 = = 12
5

Step 5: Variance (Shortcut)

2
5500 150
𝜎2 = −( ) = 1100 − 900 = 200
5 5

Step 6: Standard Deviation



𝜎= 200 = 14.14

Step 7: Coefficient of Variation

14.14
𝐶𝑉 = × 100% = 47.1%
30

Summary Table

Measure Formula Our Answer


Mean Σx/n 30
Range Max - Min 20
Mean Deviation Σ x-mean
Variance Σ(x-mean)²/n 200
Standard Deviation √Variance 14.14
CV (SD/Mean) x 100% 47.1%

33
Absolute vs Relative Measures

Absolute Measures (Have units):

• Range
• Mean Deviation
• Variance
• Standard Deviation

Relative Measures (No units, %):

• Coefficient of Variation
• Coefficient of Range
• Coefficient of Mean Deviation
Relative measures are used to COMPARE different
datasets!

MCQ Practice - Measures of Dispersion

Q1. The range of 3, 7, 12, 15, 21 is: - (A) 3 - (B) 18 - (C) 21 - (D)
12
Answer: (B) 18
Solution: 21 - 3 = 18

Q2. The variance is the square of: - (A) Mean - (B) Range - (C)
Standard Deviation - (D) Mean Deviation
Answer: (C) Standard Deviation

Q3. If SD = 5 and Mean = 25, the CV is: - (A) 5% - (B) 20% - (C)
25% - (D) 125%
Answer: (B) 20%

34
Solution: CV = (5/25) x 100% = 20%

Q4. Which measure uses ALL values in its calculation? - (A)


Range - (B) Standard Deviation - (C) Both - (D) Neither
Answer: (B) Standard Deviation

Q5. The sum of squared deviations from mean is minimum when


calculated from: - (A) Mode - (B) Median - (C) Mean - (D) Any
value
Answer: (C) Mean

Q6. If all values are increased by 5, the standard deviation: - (A)


Increases by 5 - (B) Decreases by 5 - (C) Remains unchanged -
(D) Becomes 5
Answer: (C) Remains unchanged

Q7. If all values are multiplied by 3, the variance: - (A) Is multi-


plied by 3 - (B) Is multiplied by 9 - (C) Remains unchanged - (D)
Is divided by 3
Answer: (B) Is multiplied by 9
Solution: When you multiply by k, variance is multiplied by k²

Q8. Mean Deviation is calculated using: - (A) Squared deviations


- (B) Absolute deviations - (C) Cubed deviations - (D) Simple devi-
ations
Answer: (B) Absolute deviations

Q9. For the data 5, 5, 5, 5, 5, the standard deviation is: - (A) 5 -


(B) 1 - (C) 0 - (D) 25
Answer: (C) 0

35
Solution: All values are same, no spread, SD = 0

Q10. Which is used to compare variability of different datasets? -


(A) Variance - (B) Standard Deviation - (C) Coefficient of Variation
- (D) Range
Answer: (C) Coefficient of Variation

Q11. Using shortcut formula, if Σx = 50, Σx² = 650, n = 5, Vari-


ance is: - (A) 30 - (B) 100 - (C) 130 - (D) 600
Answer: (A) 30
Solution: Variance = 650/5 - (50/5)² = 130 - 100 = 30

Q12. Sample variance uses (n-1) in denominator because: - (A)


It’s easier to calculate - (B) It gives unbiased estimate - (C) Sam-
ple is smaller - (D) No particular reason
Answer: (B) It gives unbiased estimate

END OF CHAPTERS 1-3 (PDF-FRIENDLY


VERSION)

STA 111 - CHAPTERS 4-6 (PDF-FRIENDLY


VERSION)

With VISIBLE Answers!

36
CHAPTER 4: MEASURES OF PARTITION
(POSITION)

What Are Measures of Partition?

Values that DIVIDE your data into equal parts


Like cutting a cake into equal slices!

The Three Types

Measure Divides Into Number of Values


Quartiles 4 equal parts 3 (Q1, Q2, Q3)
Deciles 10 equal parts 9 (D1 to D9)
Percentiles 100 equal parts 99 (P1 to P99)

1 QUARTILES

Divide data into 4 equal parts (25% each)


25% 25% 25% 25%
|---------|---------|---------|---------|
Min Q1 Q2 Q3 Max
(25%) (50%) (75%)
|
Median!

The Three Quartiles:

37
Quartile Also Called Position Meaning
Q1 First/Lower Quartile 25th percentile 25% below this
Q2 Second Quartile 50th percentile Same as Median!
Q3 Third/Upper Quartile 75th percentile 75% below this

Finding Quartiles - Ungrouped Data

Step 1: Arrange data in order


Step 2: Find positions:

𝑛+1
𝑄1 position =
4

2(𝑛 + 1) 𝑛 + 1
𝑄2 position = =
4 2

3(𝑛 + 1)
𝑄3 position =
4

Example - Ungrouped Quartiles:

Data: 3, 5, 7, 9, 11, 13, 15 (n = 7)


Q1 position = (7+1)/4 = 2nd value = 5
Q2 position = 2(7+1)/4 = 4th value = 9 (Median)
Q3 position = 3(7+1)/4 = 6th value = 13

38
Quartiles for Grouped Data

𝑘𝑛 − 𝑐𝑓𝑏
𝑄𝑘 = 𝐿 + ( 4
)×𝑐
𝑓𝑞

Where:

Symbol Meaning
k Which quartile (1, 2, or 3)
L Lower boundary of quartile class
n Total frequency
cf_b Cumulative frequency BEFORE quartile class
f_q Frequency of quartile class
c Class width

Example - Grouped Quartiles

Class f cf
0-9 5 5
10-19 10 15
20-29 20 35
30-39 15 50
40-49 10 60
Total 60

Find Q1:
Position = n/4 = 60/4 = 15
Q1 class: Where cf first reaches 15 or more = 10-19 class

15 − 5
𝑄1 = 9.5 + ( ) × 10 = 9.5 + 10 = 19.5
10

39
Find Q3:
Position = 3n/4 = 45
Q3 class: Where cf first reaches 45 or more = 30-39 class

45 − 35
𝑄3 = 29.5 + ( ) × 10 = 29.5 + 6.67 = 36.17
15

Interquartile Range (IQR)

The range of the middle 50% of data

𝐼𝑄𝑅 = 𝑄3 − 𝑄1

From our example: IQR = 36.17 - 19.5 = 16.67

Why IQR is Useful:

• Not affected by outliers


• Measures spread of middle values
• Used in box plots

Semi-Interquartile Range (Quartile Deviation):

𝑄3 − 𝑄1 𝐼𝑄𝑅
𝑄𝐷 = =
2 2

40
2 DECILES

Divide data into 10 equal parts (10% each)


|-----|-----|-----|-----|-----|-----|-----|-----|-----|-----|
D1 D2 D3 D4 D5 D6 D7 D8 D9
10% 20% 30% 40% 50% 60% 70% 80% 90%
Note: D5 = Median = Q2 = P50 (all the same!)

Finding Deciles

Ungrouped:
𝑘(𝑛 + 1)
𝐷𝑘 position =
10
Grouped:
𝑘𝑛 − 𝑐𝑓𝑏
𝐷𝑘 = 𝐿 + ( 10 )×𝑐
𝑓𝑑

Example - Decile

Using same grouped data (n = 60):


Find D3 (30th percentile):
Position = 3(60)/10 = 18
D3 class: Where cf first reaches 18 or more = 20-29 class

18 − 15
𝐷3 = 19.5 + ( ) × 10 = 19.5 + 1.5 = 21
20

41
3 PERCENTILES

Divide data into 100 equal parts (1% each)


Most detailed measure of position!

Finding Percentiles

Ungrouped:
𝑘(𝑛 + 1)
𝑃𝑘 position =
100
Grouped:
𝑘𝑛 − 𝑐𝑓𝑏
𝑃𝑘 = 𝐿 + ( 100 )×𝑐
𝑓𝑝

Example - Percentile

Find P65 (65th percentile):


Position = 65(60)/100 = 39
P65 class: Where cf first reaches 39 or more = 30-39 class

39 − 35
𝑃65 = 29.5 + ( ) × 10 = 29.5 + 2.67 = 32.17
15
Interpretation: 65% of values are below 32.17

Relationships Between Measures

This Equals Equals


Q1 P25 D2.5

42
This Equals Equals
Q2 P50 D5 = Median
Q3 P75 D7.5
D1 P10
D2 P20
D9 P90

Box Plot (Box-and-Whisker Plot)

Uses quartiles to show data distribution:


+-------------+
-----+ | +-----
+-------------+
| | | | |
Min Q1 Q2 Q3 Max
(Median)
The box shows the middle 50% (IQR)

MCQ Practice - Measures of Partition

Q1. Quartiles divide data into: - (A) 2 equal parts - (B) 4 equal
parts - (C) 10 equal parts - (D) 100 equal parts
Answer: (B) 4 equal parts

Q2. Q2 is the same as: - (A) Mean - (B) Mode - (C) Median - (D)
Range
Answer: (C) Median

43
Q3. The 75th percentile is also known as: - (A) Q1 - (B) Q2 - (C)
Q3 - (D) D7
Answer: (C) Q3

Q4. If Q1 = 20 and Q3 = 50, the IQR is: - (A) 20 - (B) 30 - (C) 50


- (D) 70
Answer: (B) 30
Solution: IQR = Q3 - Q1 = 50 - 20 = 30

Q5. D5 is equal to: - (A) P5 - (B) P50 - (C) P55 - (D) Q5


Answer: (B) P50

Q6. Which percentile equals Q1? - (A) P1 - (B) P10 - (C) P25 - (D)
P50
Answer: (C) P25

Q7. The Semi-Interquartile Range is: - (A) Q3 - Q1 - (B) (Q3 -


Q1)/2 - (C) (Q3 + Q1)/2 - (D) Q3 + Q1
Answer: (B) (Q3 - Q1)/2

Q8. For data: 2, 4, 6, 8, 10, 12, 14 (n=7), Q1 is: - (A) 2 - (B) 4 -


(C) 6 - (D) 8
Answer: (B) 4
Solution: Q1 position = (7+1)/4 = 2nd value = 4

Q9. How many deciles are there? - (A) 3 - (B) 9 - (C) 10 - (D) 99
Answer: (B) 9

44
Q10. If you score at the 90th percentile: - (A) You scored 90% -
(B) 90% scored above you - (C) 90% scored below you - (D) You
scored 90 points
Answer: (C) 90% scored below you

Q11. The middle 50% of data lies between: - (A) Q1 and Q2 - (B)
Q2 and Q3 - (C) Q1 and Q3 - (D) Min and Max
Answer: (C) Q1 and Q3

Q12. D3 = P___? - (A) P3 - (B) P30 - (C) P33 - (D) P300


Answer: (B) P30

Q13. If Q1 = 15 and Q3 = 45, the Quartile Deviation is: - (A) 15


- (B) 30 - (C) 45 - (D) 60
Answer: (A) 15
Solution: QD = (Q3 - Q1)/2 = (45 - 15)/2 = 15

Q14. The position of Q3 for n = 11 (ungrouped) is: - (A) 3rd value


- (B) 6th value - (C) 9th value - (D) 8th value
Answer: (C) 9th value
Solution: Q3 position = 3(11+1)/4 = 36/4 = 9

CHAPTER 5: SKEWNESS AND KURTOSIS

What Are These?

Measure What It Tells You


Skewness Is data LOPSIDED? (left or right heavy?)

45
Measure What It Tells You
Kurtosis Is data PEAKED or FLAT? (compared to normal)

These describe the SHAPE of your distribution!

PART 1: SKEWNESS

What is Skewness?

How asymmetric (lopsided) is your data?


Think of it as: “Where is the TAIL going?”

Three Types of Skewness

1. Symmetric (No Skew) - Skewness = 0

/\
/ \
/ \
/ \
/ \
---+----------+---
Mean
Median
Mode
(All equal!)
Mean = Median = Mode

46
2. Positive Skew (Right Skewed) - Skewness > 0

/\
/ \
/ \
/ \____
/ ------->
-+--------------------
Mode Median Mean

Peak on LEFT, tail on RIGHT


Mode < Median < Mean
The mean is “pulled” toward the long tail!
Examples: Income, house prices, age at death

3. Negative Skew (Left Skewed) - Skewness < 0

/\
/ \
/ \
____/ \
<----- \
--------------------+-
Mean Median Mode

Tail on LEFT, peak on RIGHT


Mean < Median < Mode
Examples: Age at retirement, easy exam scores

Memory Trick!

“The mean chases the tail”

47
• Tail goes RIGHT –> Mean goes RIGHT –> Positive skew
• Tail goes LEFT –> Mean goes LEFT –> Negative skew

Measures of Skewness

1. Karl Pearson’s Coefficient of Skewness

First Formula (using Mode):

Mean − Mode
𝑆𝑘 =
Standard Deviation

Second Formula (using Median):

3(Mean − Median)
𝑆𝑘 =
Standard Deviation

Interpreting Pearson’s Coefficient:

Value Interpretation
Sk = 0 Symmetric
Sk > 0 Positive skew (right)
Sk < 0 Negative skew (left)
Sk
0.5 < Sk
Sk

Example - Pearson’s Skewness

Given: Mean = 45, Median = 40, SD = 10

48
3(45 − 40) 3(5) 15
𝑆𝑘 = = = = 1.5
10 10 10
Interpretation: Highly positively skewed

2. Bowley’s Coefficient of Skewness (Quartile Method)

𝑄3 + 𝑄1 − 2𝑄2 𝑄3 + 𝑄1 − 2 × Median
𝑆𝑘 = =
𝑄3 − 𝑄 1 𝐼𝑄𝑅
Range: -1 to +1

Example - Bowley’s Skewness

Given: Q1 = 25, Q2 = 35, Q3 = 50

50 + 25 − 2(35) 75 − 70 5
𝑆𝑘 = = = = 0.2
50 − 25 25 25
Interpretation: Slight positive skew

3. Moment Coefficient of Skewness


𝜇3 𝜇3
𝛾1 = =
𝜎 3 (𝜇2 )3/2
Where: - μ3 = Third central moment - μ2 = Second central mo-
ment (Variance)

49
PART 2: KURTOSIS

What is Kurtosis?

How PEAKED or FLAT is your distribution compared to a


normal curve?
Think of it as: “How much is in the tails vs the middle?”

Three Types of Kurtosis

1. Mesokurtic (Normal) - Kurtosis = 3

/\
/ \ <-- Normal peak
/ \
/ \
/ \
---+----------+---
This is the baseline (Normal distribution)

2. Leptokurtic (Tall & Thin) - Kurtosis > 3

/\
|| <-- Higher peak!
/ \
/ \
/ \
/ \___ <-- Heavier tails
----+-------------
“Lepto” = Slender (like a LEAPING rabbit)
• More peaked than normal
• Heavier tails (more outliers)

50
• Data clustered at center AND in tails

3. Platykurtic (Short & Wide) - Kurtosis < 3

------
/ \ <-- Flatter peak
/ \
/ \
/ \
--+--------------+--
“Platy” = Flat (like a PLATe)
• Flatter than normal
• Lighter tails (fewer outliers)
• Data spread more evenly

Memory Trick!

Type Remember Shape


Leptokurtic Leaping (tall) Tall, pointy
Mesokurtic Medium Normal
Platykurtic Plate (flat) Flat, wide

Measures of Kurtosis

1. Moment Coefficient of Kurtosis (beta2 or gamma2)

𝜇4 𝜇4
𝛽2 = =
𝜇22 𝜎4

51
Where: - μ4 = Fourth central moment - μ2 = Variance

Excess Kurtosis:

𝛾2 = 𝛽2 − 3

This makes normal distribution = 0 (easier to interpret!)

Type beta2 gamma2 (Excess)


Leptokurtic >3 >0
Mesokurtic =3 =0
Platykurtic <3 <0

2. Percentile Coefficient of Kurtosis

𝑄3 − 𝑄 1 𝐼𝑄𝑅
𝐾= =
2(𝑃90 − 𝑃10 ) 2(𝑃90 − 𝑃10 )
For normal distribution, K is approximately 0.263

K Value Type
K > 0.263 Platykurtic
K = 0.263 Mesokurtic
K < 0.263 Leptokurtic

(Yes, it’s opposite for this formula!)

Summary: Skewness vs Kurtosis

52
Skewness Kurtosis
Measures Asymmetry (lopsidedness) Peakedness (tail weight)
Direction Left/Right Up/Down
Normal value 0 3 (or 0 for excess)
Uses Which way is tail? How peaked?

MCQ Practice - Skewness and Kurtosis

Q1. In a positively skewed distribution: - (A) Mean < Median <


Mode - (B) Mode < Median < Mean - (C) Mean = Median = Mode
- (D) Median < Mode < Mean
Answer: (B) Mode < Median < Mean

Q2. If Mean = 50 and Mode = 50, the distribution is: - (A) Posi-
tively skewed - (B) Negatively skewed - (C) Symmetric - (D) Can-
not determine
Answer: (C) Symmetric

Q3. Karl Pearson’s coefficient of skewness uses: - (A) Mean and


Mode only - (B) Mean, Mode, and Standard Deviation - (C) Quar-
tiles only - (D) Mean and Variance
Answer: (B) Mean, Mode, and Standard Deviation

Q4. A distribution with kurtosis > 3 is: - (A) Platykurtic - (B)


Mesokurtic - (C) Leptokurtic - (D) Symmetric
Answer: (C) Leptokurtic

Q5. For a normal distribution, excess kurtosis (gamma2) equals:


- (A) 0 - (B) 1 - (C) 3 - (D) -3

53
Answer: (A) 0
Solution: Excess kurtosis = beta2 - 3 = 3 - 3 = 0

Q6. Income distribution is typically: - (A) Symmetric - (B) Nega-


tively skewed - (C) Positively skewed - (D) Uniform
Answer: (C) Positively skewed

Q7. Bowley’s coefficient of skewness ranges from: - (A) 0 to 1 -


(B) -1 to +1 - (C) negative infinity to positive infinity - (D) -3 to +3
Answer: (B) -1 to +1

Q8. A platykurtic distribution has: - (A) Higher peak than normal


- (B) Flatter peak than normal - (C) Same peak as normal - (D) No
peak
Answer: (B) Flatter peak than normal

Q9. If Q1 = 20, Q2 = 30, Q3 = 40, Bowley’s skewness is: - (A) 0


- (B) 0.5 - (C) 1 - (D) -1
Answer: (A) 0
Solution: Sk = (40 + 20 - 2 x 30)/(40-20) = (60-60)/20 = 0

Q10. “The mean is pulled toward the tail” describes: - (A) Kurto-
sis - (B) Dispersion - (C) Skewness - (D) Central tendency
Answer: (C) Skewness

Q11. If Mean = 70, Median = 75, the skewness is: - (A) Positive
- (B) Negative - (C) Zero - (D) Cannot determine
Answer: (B) Negative
Solution: Mean < Median means negative skew (tail on left)

54
Q12. Leptokurtic distributions have: - (A) Fewer outliers than
normal - (B) More outliers than normal - (C) Same outliers as
normal - (D) No outliers
Answer: (B) More outliers than normal

Q13. If Mean = 100, Mode = 90, SD = 20, Pearson’s skewness


is: - (A) 0.5 - (B) 1.0 - (C) 1.5 - (D) 2.0
Answer: (A) 0.5
Solution: Sk = (Mean - Mode)/SD = (100 - 90)/20 = 0.5

Q14. The normal distribution is: - (A) Leptokurtic - (B) Platykur-


tic - (C) Mesokurtic - (D) Skewed
Answer: (C) Mesokurtic

Q15. In a negatively skewed distribution, the tail is on the: - (A)


Right side - (B) Left side - (C) Both sides - (D) Neither side
Answer: (B) Left side

CHAPTER 6: MOMENTS

What Are Moments?

Moments are summary measures that describe the shape


of a distribution
Think of them as different ways to measure how data is dis-
tributed around a point!

55
Simple Analogy

Imagine a seesaw (lever):


• First moment: Where is the balance point? (Mean)
• Second moment: How spread out are the kids? (Variance)
• Third moment: Is one side heavier? (Skewness)
• Fourth moment: Are kids clustered at center or at ends?
(Kurtosis)

Two Types of Moments

1. Moments About the Origin (Raw Moments)

∑ 𝑥𝑟
𝜇′𝑟 =
𝑛
(Ungrouped)

∑ 𝑓𝑥𝑟
𝜇′𝑟 =
∑𝑓
(Grouped)

2. Moments About the Mean (Central Moments)

∑(𝑥 − 𝑥)̄ 𝑟
𝜇𝑟 =
𝑛
(Ungrouped)

∑ 𝑓(𝑥 − 𝑥)̄ 𝑟
𝜇𝑟 =
∑𝑓
(Grouped)

56
The First Four Moments

Moment About Origin About Mean What It Measures


1st μ’1 = Mean μ1 = 0 (Always!) Location
2nd μ’2 = Mean of squares μ2 = Variance Spread
3rd μ’3 μ3 Skewness
4th μ’4 μ4 Kurtosis

The Most Important Moments

First Moment About Origin = MEAN

∑𝑥
𝜇′1 = = 𝑥̄
𝑛

First Moment About Mean = ZERO (Always!)

∑(𝑥 − 𝑥)̄
𝜇1 = =0
𝑛
(Because deviations above and below mean cancel out!)

Second Moment About Mean = VARIANCE

∑(𝑥 − 𝑥)̄ 2
𝜇2 = = 𝜎2
𝑛

57
Third Moment About Mean –> SKEWNESS

∑(𝑥 − 𝑥)̄ 3
𝜇3 =
𝑛
Used to calculate:
𝜇3 𝜇3
𝛾1 = =
3/2
𝜇2 𝜎3

Fourth Moment About Mean –> KURTOSIS

∑(𝑥 − 𝑥)̄ 4
𝜇4 =
𝑛
Used to calculate:
𝜇4 𝜇4
𝛽2 = =
𝜇22 𝜎4

Converting Between Moments (SUPER IMPOR-


TANT!)

You can calculate central moments FROM raw moments:

𝜇1 = 0

𝜇2 = 𝜇′2 − (𝜇′1 )2

𝜇3 = 𝜇′3 − 3𝜇′2 𝜇′1 + 2(𝜇′1 )3

𝜇4 = 𝜇′4 − 4𝜇′3 𝜇′1 + 6𝜇′2 (𝜇′1 )2 − 3(𝜇′1 )4

58
Complete Example - Finding Moments

Data: 2, 4, 6, 8, 10 (n = 5)

Step 1: Calculate Powers

x x² x³ x⁴
2 4 8 16
4 16 64 256
6 36 216 1296
8 64 512 4096
10 100 1000 10000
Sum = 30 Sum = 220 Sum = 1800 Sum = 15664

Step 2: Raw Moments

30
𝜇′1 = =6
5
(This is the mean!)

220
𝜇′2 = = 44
5

1800
𝜇′3 = = 360
5

15664
𝜇′4 = = 3132.8
5

Step 3: Central Moments (Using Conversion Formulas)

𝜇1 = 0
(Always zero!)

59
𝜇2 = 44 − (6)2 = 44 − 36 = 8
(Variance!)

𝜇3 = 360 − 3(44)(6) + 2(6)3


= 360 − 792 + 432 = 0

𝜇4 = 3132.8 − 4(360)(6) + 6(44)(36) − 3(1296)


= 3132.8 − 8640 + 9504 − 3888 = 108.8

Step 4: Skewness and Kurtosis

Standard Deviation:
√ √
𝜎= 𝜇2 = 8 = 2.83

Skewness:
𝜇3 0
𝛾1 = = =0
𝜎 3 22.63
(Data is symmetric!)
Kurtosis:
𝜇4 108.8
𝛽2 = = = 1.7
𝜇22 64

Excess Kurtosis:

𝛾2 = 𝛽2 − 3 = 1.7 − 3 = −1.3

(Platykurtic - flatter than normal)

Summary: What Each Moment Tells You

60
Moment About Origin About Mean What It Is
1st Mean 0 Average
2nd Mean of squares Variance Spread
3rd - For skewness Asymmetry
4th - For kurtosis Peakedness

Shortcut: The Variance Formula

Remember this?

𝜎2 = 𝑥2 − 𝑥2̄

In moment notation:

𝜇2 = 𝜇′2 − (𝜇′1 )2

“Mean of squares minus square of mean”

Moment Ratios (The beta and gamma Coefficients)

For Skewness:

𝜇23
𝛽1 =
𝜇32
𝜇3
𝛾1 = √𝛽1 = 3/2
𝜇2
(with sign of μ3)

61
For Kurtosis:
𝜇4
𝛽2 =
𝜇22

𝛾2 = 𝛽2 − 3
(Excess kurtosis)

MCQ Practice - Moments

Q1. The first moment about the mean is always: - (A) Mean - (B)
Zero - (C) Variance - (D) One
Answer: (B) Zero

Q2. The second moment about the mean is: - (A) Mean - (B)
Variance - (C) Standard Deviation - (D) Skewness
Answer: (B) Variance

Q3. μ’1 represents: - (A) Variance - (B) Mean - (C) First central
moment - (D) Standard Deviation
Answer: (B) Mean

Q4. If Σx = 50 and n = 10, then μ’1 equals: - (A) 50 - (B) 10 - (C)


5 - (D) 500
Answer: (C) 5
Solution: μ’1 = Σx/n = 50/10 = 5

Q5. The formula μ2 = μ’2 - (μ’1)² relates: - (A) Moments to mean


- (B) Raw moments to central moments - (C) Skewness to kurtosis
- (D) None of above

62
Answer: (B) Raw moments to central moments

Q6. β2 = μ4/μ2² is used to measure: - (A) Central tendency - (B)


Dispersion - (C) Skewness - (D) Kurtosis
Answer: (D) Kurtosis

Q7. For a symmetric distribution, μ3 equals: - (A) 1 - (B) 3 - (C)


0 - (D) Cannot determine
Answer: (C) 0

Q8. If μ’1 = 4 and μ’2 = 20, then variance (μ2) equals: - (A) 4 -
(B) 16 - (C) 20 - (D) 24
Answer: (A) 4
Solution: μ2 = μ’2 - (μ’1)² = 20 - (4)² = 20 - 16 = 4

Q9. The fourth moment is used to measure: - (A) Mean - (B)


Spread - (C) Skewness - (D) Kurtosis
Answer: (D) Kurtosis

Q10. Moments about which point give variance directly? - (A)


Origin - (B) Mean - (C) Mode - (D) Median
Answer: (B) Mean

Q11. If β2 = 3, the distribution is: - (A) Leptokurtic - (B) Platykur-


tic - (C) Mesokurtic - (D) Skewed
Answer: (C) Mesokurtic

Q12. γ1 = μ3/σ³ is the coefficient of: - (A) Variation - (B) Kurtosis


- (C) Skewness - (D) Dispersion

63
Answer: (C) Skewness

Q13. If μ’1 = 3, μ’2 = 15, μ’3 = 81, then μ3 equals: - (A) 0 - (B)
27 - (C) 54 - (D) 81
Answer: (A) 0
Solution: μ3 = μ’3 - 3μ’2μ’1 + 2(μ’1)³ = 81 - 3(15)(3) + 2(27) =
81 - 135 + 54 = 0

Q14. The second raw moment μ’2 equals: - (A) Sum of x divided
by n - (B) Sum of x² divided by n - (C) Variance - (D) Standard
deviation
Answer: (B) Sum of x² divided by n

Q15. Excess kurtosis (γ2) for a normal distribution is: - (A) 3 -


(B) 0 - (C) 1 - (D) -1
Answer: (B) 0
Solution: γ2 = β2 - 3 = 3 - 3 = 0

ULTIMATE STA 111 CHEAT SHEET

Central Tendency Formulas

Measure Ungrouped Grouped


Mean Sum of x / n Sum of fx / Sum of f
Median Middle value L + [(n/2 - cf_b)/f_m] x c
Mode Most frequent L + [(f_m - f_1)/(2f_m - f_1 - f_2)] x c

64
Dispersion Formulas

Measure Formula
Range Max - Min
Mean Deviation Sum of
Variance Sum of (x - mean)² / n
Standard Deviation Square root of Variance
CV (SD / Mean) x 100%

Partition Formulas (Grouped Data)

(𝑘𝑛/4 − 𝑐𝑓𝑏 )
𝑄𝑘 = 𝐿 + ×𝑐
𝑓𝑞

(𝑘𝑛/10 − 𝑐𝑓𝑏 )
𝐷𝑘 = 𝐿 + ×𝑐
𝑓𝑑

(𝑘𝑛/100 − 𝑐𝑓𝑏 )
𝑃𝑘 = 𝐿 + ×𝑐
𝑓𝑝

𝐼𝑄𝑅 = 𝑄3 − 𝑄1

𝑄3 − 𝑄 1
𝑄𝐷 =
2

Skewness Formulas

Pearson’s (Mode):

Mean − Mode
𝑆𝑘 =
𝑆𝐷

65
Pearson’s (Median):
3(Mean − Median)
𝑆𝑘 =
𝑆𝐷
Bowley’s (Quartile):
𝑄3 + 𝑄1 − 2𝑄2
𝑆𝑘 =
𝑄3 − 𝑄 1

Moment:
𝜇3
𝛾1 =
𝜎3

Kurtosis Formulas

𝜇4 𝜇4
𝛽2 = =
𝜇22 𝜎4

𝛾2 = 𝛽2 − 3
(Excess kurtosis)

Moment Conversions

𝜇1 = 0
(Always!)

𝜇2 = 𝜇′2 − (𝜇′1 )2

𝜇3 = 𝜇′3 − 3𝜇′2 𝜇′1 + 2(𝜇′1 )3

𝜇4 = 𝜇′4 − 4𝜇′3 𝜇′1 + 6𝜇′2 (𝜇′1 )2 − 3(𝜇′1 )4

66
Quick Reference Table

If… Then…
Mean = Median = Mode Symmetric (Sk = 0)
Mode < Median < Mean Positive skew (tail RIGHT)
Mean < Median < Mode Negative skew (tail LEFT)
β2 > 3 Leptokurtic (peaked, heavy tails)
β2 = 3 Mesokurtic (normal)
β2 < 3 Platykurtic (flat, light tails)

Key Relationships

This Equals This And This


Q2 Median P50 = D5
Q1 P25 D2.5
Q3 P75 D7.5
D1 P10
D9 P90

Effects of Transformations

Transformation Mean SD Variance


Add constant k +k No change No change
Multiply by k xk xk x k²

67
Memory Tricks

1. Mean = Add all, divide by n


2. Median = Middle value (arrange first!)
3. Mode = Most frequent (most popular!)
4. Variance = Mean of squares - Square of mean
5. SD = Square root of Variance
6. “Mean chases the tail” for skewness
7. Lepto = Leaping (tall), Platy = Plate (flat)
8. μ1 about mean = 0 (ALWAYS!)
9. μ2 about mean = Variance

YOU’RE READY FOR STA 111!

What You’ve Learned:

Chapter Topic Key Point


1 Data Primary (you collect) vs Secondary (already exists)
2 Location Mean, Median, Mode - where’s the center?
3 Dispersion Range, SD, Variance, CV - how spread out?
4 Partition Quartiles, Deciles, Percentiles - where are the divisions?
5 Shape Skewness (lopsided?) & Kurtosis (peaked?)
6 Moments The math behind it all!

Final Tips:

1. Know your formulas (especially grouped data formulas!)


2. Practice calculations step by step
3. Remember the relationships (Q2 = Median = P50 = D5)

68
4. Understand what each measure MEANS, not just how to cal-
culate
5. Skewness: “Mean chases the tail”
6. Kurtosis: Compare to normal (β2 = 3)

Now go CRUSH that test i guess!


You’ve got this!

69

Common questions

Powered by AI

Ogives, or cumulative frequency graphs, contribute significantly to interpreting data by displaying cumulative data against a continuous interval or class boundary . They help identify medians, quartiles, and percentiles, facilitating quick assessments of data spread and distribution. The unique advantage of ogives is their ability to provide a graphical representation of cumulative frequency distribution that assists in visualizing how data accumulates over a range, making it easier to extract cumulative insights and make comparisons between different data sets .

Deciles and percentiles enhance the understanding of data distribution by dividing data into equal parts, offering a detailed view of data position relative to the whole distribution . Deciles partition the data set into ten equally sized intervals, while percentiles divide it into one hundred . These measures are particularly useful in practical scenarios like standardized testing scores, where determining cut-off points for specific percentages (e.g., 90th percentile) can inform selections and assessments. They provide detailed, ranked positions and are integral in assessing performance or deviations from a norm.

The mean, or average, is advantageous because it uses all data values in its calculation, making it useful for further statistical analysis and widely applicable . However, it is disadvantageous because it can be heavily influenced by outliers in the data set, which may not represent the typical value of the distribution . In situations with skewed data or significant outliers, the mean may not accurately reflect the center of the data and other measures like the median might provide a better central indication .

The choice of data representation method significantly affects data interpretation. Histograms are used to display the distribution of data frequency and can highlight the shape and spread of continuous data sets as they are drawn without gaps between bars . Pie charts, however, are more suitable for showing parts of a whole, allowing for a visual proportion analysis . The interpretation changes based on the visual emphasis each type of chart provides, such as distribution and frequency in histograms versus relational parts in pie charts.

Skewness is a measure of asymmetry or departure from symmetry in a data distribution. It quantifies whether data points are symmetrically distributed around the mean. A distribution may be positively skewed (right tail longer) or negatively skewed (left tail longer). When the mean is greater than the median, the distribution is positively skewed . Skewness reveals potential outliers and the bulk of data distribution which is crucial for choosing appropriate statistical methods and interpreting data accurately. By understanding symmetry, one can infer the general trend and balance within the data .

Primary data is data collected directly by a researcher through methods such as questionnaires or observations. It is specific to the researcher's study and not previously existing . Secondary data, on the other hand, is data that has already been collected and published by other sources, like government census data . These differences are significant because primary data allows for more specific and tailored data collection, while secondary data can be easier and more cost-effective to obtain, but may not perfectly fit the specific needs of a new research project. Understanding these distinctions helps researchers choose the appropriate data collection method depending on the needs and constraints of their study.

The interquartile range (IQR) plays a crucial role in statistical analysis as it measures the spread of the middle 50% of data, providing an understanding of data dispersion without being affected by extreme outliers . Unlike the range, which considers the extreme values, IQR focuses on the central part of the distribution, making it more robust against anomalies . This makes IQR particularly valuable in skewed data sets and when comparing variability across different samples, as it gives a more reliable sense of data spread.

Central moments are pivotal in understanding data distribution characteristics such as variance, skewness, and kurtosis. They measure moments about the mean, thus highlighting deviations and distribution symmetry. The first central moment is always zero (mean deviation nullifies). Raw moments relate to central moments by providing the basis for computing them; raw moments are calculated directly from normalized data, while central moments adjust these to better reflect data behavior concerning the mean. This relationship aids in translating basic data metrics into nuanced interpretations of data behavior and probability distribution .

Variance is a key statistical measure that quantifies the degree of spread in a set of data points around the mean. It indicates how much individual data points differ from the mean on average . The greater the variance, the more spread out the data points are, suggesting greater diversity within the data set. Variance is crucial for identifying variability within data, setting the foundation for further statistical analyses like standard deviation and hypothesis testing . It provides insight into consistency and reliability in data, crucial for decision-making and predictions.

Kurtosis provides insights into the 'tailedness' of a data distribution. It measures the extent to which data is concentrated in the tails or peaks of the distribution compared to a normal distribution . Distributions can be leptokurtic (taller peaks, fatter tails), platykurtic (flatter peaks), or mesokurtic (normal distribution). Excess kurtosis, which adjusts for a mesokurtic baseline, highlights deviations from normality. Understanding kurtosis helps determine the likelihood of outlier occurrence and is crucial for risk assessments in fields like finance and decision-making under uncertainty .

You might also like