STATISTICS
Complete Study Guide
10 Essential Topics for Mastery
1. Definition & Types of Statistics
2. Data / Variable & Its Types
3. Sample vs. Population
4. Parameter vs. Statistic
5. Types of Data Collection
6. Tabular & Graphical Data Representation
7. Measures of Central Tendency, Position & Dispersion
8. Exploratory Data Analysis & Box Plot
9. Absolute & Relative Measures of Dispersion
10. Qualitative-wise Quantitative Analysis
Designed for clarity · Packed with examples · Ready to use
Statistics Study Guide Comprehensive Reference
1 Definition and Types of Statistics
What is Statistics?
Statistics is the science of collecting, organizing, analyzing, interpreting, and presenting data to
make informed decisions. It bridges raw numbers and meaningful knowledge.
KEY INSIGHT
Statistics turns raw data into meaningful knowledge, helping us make sense of the world through numbers.
Two Major Branches
Branch Focus Goal Example
Descriptive Summarizing & Describe what Mean exam score
Statistics organizing data is in the data of a class is 75
Inferential Drawing conclusions Make predictions Predict election
Statistics from samples about a population results from a poll
Further Sub-Types
Type Description
Applied Statistics
Uses statistical methods to solve real-world problems (economics, medicine, engineering).
Mathematical Statistics Theoretical foundation — probability theory, derivations, proofs.
Business Statistics Applied in business for decision-making, quality control, forecasting.
Biostatistics Applied in biology and public health to analyze health-related data.
REMEMBER
Descriptive = DESCRIBE what you have. Inferential = INFER beyond what you have.
Statistics Study Guide Page 2 © 2024 Study Notes
Statistics Study Guide Comprehensive Reference
2 Data / Variable and Its Types
What is Data?
Data are raw facts and figures collected for analysis. A Variable is any characteristic that can take
different values across individuals or observations.
Classification of Data
DATA / VARIABLE TYPES
QUALITATIVE (Categorical) QUANTITATIVE (Numerical)
Data that represents categories or groups. Data that represents counts or measurements.
Cannot be measured numerically. Can be ordered and computed upon.
NOMINAL ORDINAL
No natural order Has natural order
Ex: Colors, Gender, Ex: Rating (1-5),
Blood Group Education Level
DISCRETE CONTINUOUS
Counted, whole numbers Measured, any value in range
Ex: No. of students, Ex: Height, Weight,
No. of cars Temperature
Scale of Measurement (Stevens' Scales)
Scale Properties Operations Example
Nominal Categories only Count, Mode Gender, Religion
Ordinal Ordered categories Median, Rank Satisfaction Rating
Interval Equal intervals, no true zero Mean, SD Temperature (°C)
Ratio Equal intervals + true zero All statistics Height, Weight, Age
MEMORY TIP
NOIR → Nominal, Ordinal, Interval, Ratio (each level has MORE information than the previous one).
Statistics Study Guide Page 3 © 2024 Study Notes
Statistics Study Guide Comprehensive Reference
3 Difference Between Sample and Population
Core Concepts
POPULATION SAMPLE
The entire group of individuals A subset of the population
Definition
or items of interest selected for study
Size Usually very large (N) Smaller, manageable (n)
Fixed but often unknown Estimated from data
Parameters
(µ, σ, P) (x■, s, p■)
Feasibility Often impractical to study Practical and cost-effective
Symbol (mean) µ (mu) x■ (x-bar)
Symbol (SD) σ (sigma) s
Survey of 500 students
Example All students in Pakistan
in Lahore
Visual Analogy
Think of a bag of 10,000 marbles. You cannot count every marble, so you draw a handful of 50. The bag
is the population and your handful is the sample. The proportion of red marbles in your handful estimates
the proportion in the whole bag.
KEY RULE
A good sample must be REPRESENTATIVE — it should reflect the characteristics of the population accurately.
Types of Sampling Methods
Method How It Works Best For
Simple Random Every member has equal chance Homogeneous populations
Systematic Every kth member selected Large, ordered lists
Stratified Population divided into strata; random from each Diverse populations
Cluster Groups selected; all members in group included
Geographically spread populations
Convenience Easiest to reach members selected Exploratory research (less reliable)
Statistics Study Guide Page 4 © 2024 Study Notes
Statistics Study Guide Comprehensive Reference
4 Difference Between Parameter and Statistic
Definitions
Parameter: A numerical value that describes a characteristic of a population. It is fixed (though usually
unknown).
Statistic: A numerical value that describes a characteristic of a sample. It varies from sample to sample.
Feature PARAMETER STATISTIC
Describes Population Sample
Known / Unknown Usually unknown Calculated from data
Fixed or Variable Fixed value Varies by sample
Mean µ (mu) x■ (x-bar)
Standard Deviation σ (sigma) s
Proportion P (or π) p■ (p-hat)
Variance σ² s²
Size N (capital) n (lowercase)
QUICK TRICK
"P" for Parameter → Population (both start with P). "S" for Statistic → Sample (both start with S).
Example: The average height of ALL students in a university (N=5000) is µ = 165 cm — this is a
PARAMETER. A sample of 100 students has x■ = 163 cm — this is a STATISTIC.
Statistics Study Guide Page 5 © 2024 Study Notes
Statistics Study Guide Comprehensive Reference
5 Types of Data Collection
Primary vs Secondary Data
PRIMARY DATA SECONDARY DATA
Collected firsthand by the Already collected by someone else;
Definition
researcher for specific purpose used for a different purpose
Direct observation, surveys, Published reports, census data,
Source
experiments, interviews government records, databases
Cost Expensive & time-consuming Cheaper & faster
Accuracy More accurate & specific May not perfectly fit your need
A researcher surveys 200 people Using WHO health statistics
Example
about their diet for a study on obesity
Methods of Primary Data Collection
Method Description Advantages Limitations
Researcher asks questions
Interview Detailed, flexible Expensive, interviewer bias
directly to respondent
Written/online survey
Questionnaire Large reach, cheap Low response rate
with structured questions
Watch and record
Observation Natural data, no recall bias Observer effect
behavior or events
Controlled environment;
Experiment Establishes causality Artificial setting
manipulate variables
Group discussion
Focus Groups Rich qualitative insights Not generalizable
on a specific topic
In-depth study of
Case Study Detailed understanding Not representative
one unit
REMEMBER
Before collecting data, always define your RESEARCH QUESTION clearly — it determines which method is best.
Statistics Study Guide Page 6 © 2024 Study Notes
Statistics Study Guide Comprehensive Reference
6 Tabular & Graphical Representation of Data
A. Frequency Distribution Table
A Frequency Distribution organizes raw data into classes/categories with their frequencies.
Class Interval Tally Frequency (f) Relative Freq. Cumulative Freq.
10 – 20 |||| 4 4/30 = 0.13 4
20 – 30 |||| |||| 9 9/30 = 0.30 13
30 – 40 |||| |||| || 12 12/30 = 0.40 25
40 – 50 |||| 5 5/30 = 0.17 30
TOTAL 30 1.00 —
KEY TERMS
Class Width = Upper Limit - Lower Limit | Midpoint = (UL + LL)/2 | Relative Freq = f/n | Cumulative Freq = running
B. Bar Charts
Type Description Use Case
Simple Bar Chart One bar per category; height = frequency
Comparing one variable across groups
Multiple (Grouped)
Bars side-by-side for each sub-group
Comparing two or more groups side by side
Bar Chart
Component (Stacked)
Bars stacked; each segment = sub-group
Showing part-to-whole within each group
Bar Chart
Simple Bar Chart Histogram Style15 Pie Chart
18
A
14 12
11 B
10 9 C
8
8 6 D
E
A B C D Jan Feb Mar Apr May Jun
C. Histogram
A Histogram is like a bar chart but for continuous data. Bars are adjacent (no gaps) because class
intervals are continuous. The area of each bar represents frequency (not just height).
Statistics Study Guide Page 7 © 2024 Study Notes
Statistics Study Guide Comprehensive Reference
Histogram
10-20 20-30 30-40 40-50 50-60 60-70
Class Intervals
BAR vs HISTOGRAM
Bar Chart: Qualitative/Discrete data | Gaps between bars | Height = frequency Histogram: Continuous/Quantitative
data | NO gaps | Area = frequency
Statistics Study Guide Page 8 © 2024 Study Notes
Statistics Study Guide Comprehensive Reference
7 Measures of Central Tendency, Position & Dispersion
A. Measures of Central Tendency
These measures identify the center or typical value of a dataset.
Measure Formula When to Use Resistant to Outliers?
x■ = Σx / n
Mean (Arithmetic
Symmetric data, interval/ratio scale
NO — sensitive to outliers
Average)
For grouped: x■ = Σ(f·m) / Σf
Middle value when sorted
Median Skewed data, ordinal/ratioYES
scale
— not affected by extremes
Even n: avg of two middle values
Most frequently occurring value
Mode Nominal/ordinal data; finding most
YEScommon
— unaffected by outliers
(can be multiple or none)
SKEWNESS RULE
Symmetric: Mean = Median = Mode Right (Positive) Skew: Mode < Median < Mean Left (Negative) Skew: Mean <
Median < Mode
B. Measures of Position (Quantiles)
These divide ordered data into equal parts to locate a value's position in the distribution.
Measure Divides Into Key Values Formula (ungrouped)
Quartiles (Q) 4 equal parts Q1 (25%), Q2=Median (50%), Q3 (75%)
Qi = value at i(n+1)/4 th position
Deciles (D) 10 equal parts D1–D9; D5 = Median Di = value at i(n+1)/10 th position
Percentiles (P) 100 equal parts P25=Q1, P50=Median, P75=Q3 Pi = value at i(n+1)/100 th position
C. Measures of Dispersion
These measure the spread or variability of data around the center.
Measure Formula Interpretation
Range R = Max – Min Simplest measure; very sensitive to outliers
Mean Deviation (MD) MD = Σ|x – x■| / n Average absolute distance from the mean
s² = Σ(x–x■)² / (n–1)
Variance (s²) Average squared deviation; in squared units
[Population: σ² = Σ(x–µ)² / N]
Standard Deviation s = √s² (or σ = √σ²) Most used; same units as data; interpretable
Statistics Study Guide Page 9 © 2024 Study Notes
Statistics Study Guide Comprehensive Reference
IQR IQR = Q3 – Q1 Range of middle 50%; resistant to outliers
WHEN TO USE WHICH
Use Range for quick rough idea | Use IQR for skewed data | Use Variance/SD for further calculations | Use SD to
Statistics Study Guide Page 10 © 2024 Study Notes
Statistics Study Guide Comprehensive Reference
8 Exploratory Data Analysis (EDA) and Box Plot
What is EDA?
Exploratory Data Analysis (EDA) is an approach to analyzing datasets to summarize their main
characteristics, often visually, before applying formal modeling or hypothesis testing. Introduced by John
Tukey in 1977.
EDA Goal How Achieved
Understand data structure Shape of distribution, types of variables
Detect outliers & anomalies Box plots, scatter plots, z-scores
Find patterns & relationships Correlation, heatmaps, pair plots
Check assumptions Normality tests, residual plots
Guide model selection Based on shape, spread, skewness
The Box Plot (Box-and-Whisker Plot)
A Box Plot is a standardized visual summary of a dataset's distribution based on five key values:
Minimum, Q1, Median, Q3, and Maximum.
Box Plot (Box-and-Whisker Plot)
IQR = Q3 - Q1
Outlier
Min Q1 Median Q3 Max
Five-Number Summary
Component What It Is How to Calculate
Minimum Smallest non-outlier value Smallest value ≥ Q1 − 1.5×IQR
Q1 (25th Pctile) Lower quartile; 25% data below Median of lower half
Median (Q2) Middle value; 50% data below Middle of sorted data
Q3 (75th Pctile) Upper quartile; 75% data below Median of upper half
Maximum Largest non-outlier value Largest value ≤ Q3 + 1.5×IQR
Outlier Detection
Statistics Study Guide Page 11 © 2024 Study Notes
Statistics Study Guide Comprehensive Reference
Using the 1.5 × IQR Rule:
• Lower Fence = Q1 − 1.5 × IQR
• Upper Fence = Q3 + 1.5 × IQR
• Any value beyond these fences is considered an outlier.
EDA TOOLS SUMMARY
Histograms (shape) | Box Plots (spread & outliers) | Scatter Plots (relationships) | Bar Charts (categories) | Heat
Statistics Study Guide Page 12 © 2024 Study Notes
Statistics Study Guide Comprehensive Reference
9 Absolute and Relative Measures of Dispersion
Two Categories of Dispersion Measures
Absolute measures express dispersion in the same units as the data. Relative measures express
dispersion as a ratio or percentage — allowing comparison between different datasets.
ABSOLUTE MEASURES RELATIVE MEASURES
Definition Dispersion in original data units Unit-free ratio (dimensionless)
Use Case Describing spread within ONE dataset
COMPARING spread of TWO or more datasets
Affected by units? YES NO
Examples Range, MD, Variance, SD, IQR CV, Coefficient of Range, Coefficient of MD
Absolute Measures — Formulas
Measure Formula Notes
Range R = X_max – X_min Simplest; highly sensitive to outliers
Mean Deviation
MD = Σ|xi – x■| / n Uses absolute values; ignores sign
(about Mean)
Mean Deviation
MD = Σ|xi – Median| / n Minimum value of MD is about median
(about Median)
Variance s² = Σ(xi–x■)² / (n–1) Penalizes large deviations more (squared)
Standard Deviation s = √[Σ(xi–x■)²/(n–1)] Most important; same units as data
IQR IQR = Q3 – Q1 Range of middle 50%; robust to outliers
Relative Measures — Formulas
Measure Formula Interpretation
Coefficient of
CV = (s / x■) × 100% Most widely used; lower CV = more consistent data
Variation (CV)
Coefficient of
CR = (Max–Min) / (Max+Min) Ranges from 0 to 1; 0 = no spread
Range
Coefficient of
CMD = MD / x■ (or MD / Median)Relative dispersion around mean or median
Mean Deviation
Coefficient of
CQD = (Q3–Q1) / (Q3+Q1) IQR-based; robust to outliers
Quartile Deviation
Statistics Study Guide Page 13 © 2024 Study Notes
Statistics Study Guide Comprehensive Reference
CLASSIC EXAMPLE
Dataset A: Ages (mean=30, SD=5) | Dataset B: Incomes in PKR (mean=50000, SD=8000) CV_A = 5/30×100 =
16.7% | CV_B = 8000/50000×100 = 16% Interpretation: Incomes are SLIGHTLY more consistent than ages (lower
CV).
Statistics Study Guide Page 14 © 2024 Study Notes
Statistics Study Guide Comprehensive Reference
10 Qualitative-wise Quantitative Analysis
What is it?
This approach involves computing quantitative (numerical) measures separately for each category
of a qualitative variable. It allows us to compare distributions of a numeric variable across different
groups or categories.
Example: Comparing exam scores (quantitative) separately for Male and Female (qualitative) students.
Steps in Qualitative-wise Quantitative Analysis
Step Action Example
1 Identify qualitative variable (grouping factor) Gender: Male / Female
2 Identify quantitative variable to be analyzed Exam Score (0–100)
3 Split data by each category Group A: Male scores | Group B: Female scores
4 Compute measures for each group separately Mean, Median, SD, Min, Max for each group
5 Compare and interpret the groups Which group scored higher? More variable?
Example: Exam Scores by Department
Statistic Engineering (n=30) Medicine (n=28) Arts (n=25)
Mean (x■) 72.4 80.1 65.3
Median 74 81 66
Mode 75 82 63
Standard Dev (s) 10.2 7.8 12.5
Min 48 61 38
Max 95 98 90
CV (%) 14.1% 9.7% 19.1%
INTERPRETATION
Medicine students have the highest mean (80.1) and lowest CV (9.7%) — they score highest AND most
consistently. Arts students have the most variability (CV=19.1%) — scores are most spread out in this group.
Common Tools for This Analysis
Tool / Technique What It Does
Statistics Study Guide Page 15 © 2024 Study Notes
Statistics Study Guide Comprehensive Reference
Group-wise Descriptive Stats Mean, Median, SD computed per group (as shown above)
Side-by-side Box Plots Visually compares distributions of groups; shows medians, spreads & outliers
Grouped Bar Charts Compares frequencies or means across categories
Contingency Tables Cross-tabulates two categorical variables to find associations
One-way ANOVA Tests if means differ significantly across 3+ groups (inferential)
t-Test (2 groups) Tests if means of exactly two groups are significantly different
WHY IS THIS IMPORTANT?
Many real-world questions require comparing groups: Do men and women earn differently? Does treatment A work
QUICK REFERENCE SUMMARY
Topic Key Takeaway
1. Statistics Descriptive (describe data) vs Inferential (draw conclusions)
2. Data Types Qualitative (Nominal/Ordinal) | Quantitative (Discrete/Continuous) | NOIR scales
3. Sample vs Population Population=all (N,µ,σ) | Sample=subset (n,x■,s) | Must be representative
4. Parameter vs Statistic Parameter=population (fixed, unknown) | Statistic=sample (calculated, varies)
5. Data Collection Primary (firsthand) vs Secondary (existing) | 6 main primary methods
6. Representation FDT | Bar (Simple/Multiple/Component) | Pie | Histogram (no gaps!)
7. Central Tendency Mean (sensitive) | Median (robust) | Mode (nominal) | Skewness rule
8. EDA & Box Plot 5-number summary: Min,Q1,Median,Q3,Max | Outliers: 1.5×IQR rule
9. Dispersion Absolute (same units) | Relative (unit-free, use CV to compare datasets)
10. QW Quant Analysis Compute stats per qualitative group → Compare → Interpret differences
Statistics Study Guide Page 16 © 2024 Study Notes