0% found this document useful (0 votes)
4 views16 pages

Statistics Study Guide

This document is a comprehensive study guide on statistics, covering 10 essential topics including definitions, types of data, sampling methods, and measures of central tendency and dispersion. It provides detailed explanations, examples, and visual representations to aid understanding. The guide is designed for clarity and practical application in various fields such as business and health.

Uploaded by

abdtahir7676
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views16 pages

Statistics Study Guide

This document is a comprehensive study guide on statistics, covering 10 essential topics including definitions, types of data, sampling methods, and measures of central tendency and dispersion. It provides detailed explanations, examples, and visual representations to aid understanding. The guide is designed for clarity and practical application in various fields such as business and health.

Uploaded by

abdtahir7676
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

STATISTICS

Complete Study Guide

10 Essential Topics for Mastery

1. Definition & Types of Statistics

2. Data / Variable & Its Types

3. Sample vs. Population

4. Parameter vs. Statistic

5. Types of Data Collection

6. Tabular & Graphical Data Representation

7. Measures of Central Tendency, Position & Dispersion

8. Exploratory Data Analysis & Box Plot

9. Absolute & Relative Measures of Dispersion

10. Qualitative-wise Quantitative Analysis

Designed for clarity · Packed with examples · Ready to use


Statistics Study Guide Comprehensive Reference

1 Definition and Types of Statistics

What is Statistics?
Statistics is the science of collecting, organizing, analyzing, interpreting, and presenting data to
make informed decisions. It bridges raw numbers and meaningful knowledge.

KEY INSIGHT
Statistics turns raw data into meaningful knowledge, helping us make sense of the world through numbers.

Two Major Branches


Branch Focus Goal Example

Descriptive Summarizing & Describe what Mean exam score


Statistics organizing data is in the data of a class is 75

Inferential Drawing conclusions Make predictions Predict election


Statistics from samples about a population results from a poll

Further Sub-Types

Type Description

Applied Statistics
Uses statistical methods to solve real-world problems (economics, medicine, engineering).

Mathematical Statistics Theoretical foundation — probability theory, derivations, proofs.

Business Statistics Applied in business for decision-making, quality control, forecasting.

Biostatistics Applied in biology and public health to analyze health-related data.

REMEMBER
Descriptive = DESCRIBE what you have. Inferential = INFER beyond what you have.

Statistics Study Guide Page 2 © 2024 Study Notes


Statistics Study Guide Comprehensive Reference

2 Data / Variable and Its Types

What is Data?
Data are raw facts and figures collected for analysis. A Variable is any characteristic that can take
different values across individuals or observations.

Classification of Data

DATA / VARIABLE TYPES

QUALITATIVE (Categorical) QUANTITATIVE (Numerical)

Data that represents categories or groups. Data that represents counts or measurements.
Cannot be measured numerically. Can be ordered and computed upon.

NOMINAL ORDINAL
No natural order Has natural order
Ex: Colors, Gender, Ex: Rating (1-5),
Blood Group Education Level

DISCRETE CONTINUOUS
Counted, whole numbers Measured, any value in range
Ex: No. of students, Ex: Height, Weight,
No. of cars Temperature

Scale of Measurement (Stevens' Scales)


Scale Properties Operations Example

Nominal Categories only Count, Mode Gender, Religion

Ordinal Ordered categories Median, Rank Satisfaction Rating

Interval Equal intervals, no true zero Mean, SD Temperature (°C)

Ratio Equal intervals + true zero All statistics Height, Weight, Age

MEMORY TIP
NOIR → Nominal, Ordinal, Interval, Ratio (each level has MORE information than the previous one).

Statistics Study Guide Page 3 © 2024 Study Notes


Statistics Study Guide Comprehensive Reference

3 Difference Between Sample and Population

Core Concepts
POPULATION SAMPLE

The entire group of individuals A subset of the population


Definition
or items of interest selected for study

Size Usually very large (N) Smaller, manageable (n)

Fixed but often unknown Estimated from data


Parameters
(µ, σ, P) (x■, s, p■)

Feasibility Often impractical to study Practical and cost-effective

Symbol (mean) µ (mu) x■ (x-bar)

Symbol (SD) σ (sigma) s

Survey of 500 students


Example All students in Pakistan
in Lahore

Visual Analogy
Think of a bag of 10,000 marbles. You cannot count every marble, so you draw a handful of 50. The bag
is the population and your handful is the sample. The proportion of red marbles in your handful estimates
the proportion in the whole bag.

KEY RULE
A good sample must be REPRESENTATIVE — it should reflect the characteristics of the population accurately.

Types of Sampling Methods


Method How It Works Best For

Simple Random Every member has equal chance Homogeneous populations

Systematic Every kth member selected Large, ordered lists

Stratified Population divided into strata; random from each Diverse populations

Cluster Groups selected; all members in group included


Geographically spread populations

Convenience Easiest to reach members selected Exploratory research (less reliable)

Statistics Study Guide Page 4 © 2024 Study Notes


Statistics Study Guide Comprehensive Reference

4 Difference Between Parameter and Statistic

Definitions
Parameter: A numerical value that describes a characteristic of a population. It is fixed (though usually
unknown).
Statistic: A numerical value that describes a characteristic of a sample. It varies from sample to sample.

Feature PARAMETER STATISTIC

Describes Population Sample

Known / Unknown Usually unknown Calculated from data

Fixed or Variable Fixed value Varies by sample

Mean µ (mu) x■ (x-bar)

Standard Deviation σ (sigma) s

Proportion P (or π) p■ (p-hat)

Variance σ² s²

Size N (capital) n (lowercase)

QUICK TRICK
"P" for Parameter → Population (both start with P). "S" for Statistic → Sample (both start with S).

Example: The average height of ALL students in a university (N=5000) is µ = 165 cm — this is a
PARAMETER. A sample of 100 students has x■ = 163 cm — this is a STATISTIC.

Statistics Study Guide Page 5 © 2024 Study Notes


Statistics Study Guide Comprehensive Reference

5 Types of Data Collection

Primary vs Secondary Data


PRIMARY DATA SECONDARY DATA

Collected firsthand by the Already collected by someone else;


Definition
researcher for specific purpose used for a different purpose

Direct observation, surveys, Published reports, census data,


Source
experiments, interviews government records, databases

Cost Expensive & time-consuming Cheaper & faster

Accuracy More accurate & specific May not perfectly fit your need

A researcher surveys 200 people Using WHO health statistics


Example
about their diet for a study on obesity

Methods of Primary Data Collection


Method Description Advantages Limitations

Researcher asks questions


Interview Detailed, flexible Expensive, interviewer bias
directly to respondent

Written/online survey
Questionnaire Large reach, cheap Low response rate
with structured questions

Watch and record


Observation Natural data, no recall bias Observer effect
behavior or events

Controlled environment;
Experiment Establishes causality Artificial setting
manipulate variables

Group discussion
Focus Groups Rich qualitative insights Not generalizable
on a specific topic

In-depth study of
Case Study Detailed understanding Not representative
one unit

REMEMBER
Before collecting data, always define your RESEARCH QUESTION clearly — it determines which method is best.

Statistics Study Guide Page 6 © 2024 Study Notes


Statistics Study Guide Comprehensive Reference

6 Tabular & Graphical Representation of Data

A. Frequency Distribution Table


A Frequency Distribution organizes raw data into classes/categories with their frequencies.

Class Interval Tally Frequency (f) Relative Freq. Cumulative Freq.

10 – 20 |||| 4 4/30 = 0.13 4

20 – 30 |||| |||| 9 9/30 = 0.30 13

30 – 40 |||| |||| || 12 12/30 = 0.40 25

40 – 50 |||| 5 5/30 = 0.17 30

TOTAL 30 1.00 —

KEY TERMS
Class Width = Upper Limit - Lower Limit | Midpoint = (UL + LL)/2 | Relative Freq = f/n | Cumulative Freq = running

B. Bar Charts

Type Description Use Case

Simple Bar Chart One bar per category; height = frequency


Comparing one variable across groups

Multiple (Grouped)
Bars side-by-side for each sub-group
Comparing two or more groups side by side
Bar Chart

Component (Stacked)
Bars stacked; each segment = sub-group
Showing part-to-whole within each group
Bar Chart

Simple Bar Chart Histogram Style15 Pie Chart


18
A
14 12
11 B
10 9 C
8
8 6 D
E

A B C D Jan Feb Mar Apr May Jun

C. Histogram
A Histogram is like a bar chart but for continuous data. Bars are adjacent (no gaps) because class
intervals are continuous. The area of each bar represents frequency (not just height).

Statistics Study Guide Page 7 © 2024 Study Notes


Statistics Study Guide Comprehensive Reference

Histogram

10-20 20-30 30-40 40-50 50-60 60-70


Class Intervals

BAR vs HISTOGRAM
Bar Chart: Qualitative/Discrete data | Gaps between bars | Height = frequency Histogram: Continuous/Quantitative
data | NO gaps | Area = frequency

Statistics Study Guide Page 8 © 2024 Study Notes


Statistics Study Guide Comprehensive Reference

7 Measures of Central Tendency, Position & Dispersion

A. Measures of Central Tendency


These measures identify the center or typical value of a dataset.

Measure Formula When to Use Resistant to Outliers?

x■ = Σx / n
Mean (Arithmetic
Symmetric data, interval/ratio scale
NO — sensitive to outliers
Average)
For grouped: x■ = Σ(f·m) / Σf

Middle value when sorted


Median Skewed data, ordinal/ratioYES
scale
— not affected by extremes
Even n: avg of two middle values

Most frequently occurring value


Mode Nominal/ordinal data; finding most
YEScommon
— unaffected by outliers
(can be multiple or none)

SKEWNESS RULE
Symmetric: Mean = Median = Mode Right (Positive) Skew: Mode < Median < Mean Left (Negative) Skew: Mean <
Median < Mode

B. Measures of Position (Quantiles)


These divide ordered data into equal parts to locate a value's position in the distribution.

Measure Divides Into Key Values Formula (ungrouped)

Quartiles (Q) 4 equal parts Q1 (25%), Q2=Median (50%), Q3 (75%)


Qi = value at i(n+1)/4 th position

Deciles (D) 10 equal parts D1–D9; D5 = Median Di = value at i(n+1)/10 th position

Percentiles (P) 100 equal parts P25=Q1, P50=Median, P75=Q3 Pi = value at i(n+1)/100 th position

C. Measures of Dispersion
These measure the spread or variability of data around the center.

Measure Formula Interpretation

Range R = Max – Min Simplest measure; very sensitive to outliers

Mean Deviation (MD) MD = Σ|x – x■| / n Average absolute distance from the mean

s² = Σ(x–x■)² / (n–1)
Variance (s²) Average squared deviation; in squared units
[Population: σ² = Σ(x–µ)² / N]

Standard Deviation s = √s² (or σ = √σ²) Most used; same units as data; interpretable

Statistics Study Guide Page 9 © 2024 Study Notes


Statistics Study Guide Comprehensive Reference

IQR IQR = Q3 – Q1 Range of middle 50%; resistant to outliers

WHEN TO USE WHICH


Use Range for quick rough idea | Use IQR for skewed data | Use Variance/SD for further calculations | Use SD to

Statistics Study Guide Page 10 © 2024 Study Notes


Statistics Study Guide Comprehensive Reference

8 Exploratory Data Analysis (EDA) and Box Plot

What is EDA?
Exploratory Data Analysis (EDA) is an approach to analyzing datasets to summarize their main
characteristics, often visually, before applying formal modeling or hypothesis testing. Introduced by John
Tukey in 1977.

EDA Goal How Achieved

Understand data structure Shape of distribution, types of variables

Detect outliers & anomalies Box plots, scatter plots, z-scores

Find patterns & relationships Correlation, heatmaps, pair plots

Check assumptions Normality tests, residual plots

Guide model selection Based on shape, spread, skewness

The Box Plot (Box-and-Whisker Plot)


A Box Plot is a standardized visual summary of a dataset's distribution based on five key values:
Minimum, Q1, Median, Q3, and Maximum.

Box Plot (Box-and-Whisker Plot)

IQR = Q3 - Q1

Outlier

Min Q1 Median Q3 Max

Five-Number Summary
Component What It Is How to Calculate

Minimum Smallest non-outlier value Smallest value ≥ Q1 − 1.5×IQR

Q1 (25th Pctile) Lower quartile; 25% data below Median of lower half

Median (Q2) Middle value; 50% data below Middle of sorted data

Q3 (75th Pctile) Upper quartile; 75% data below Median of upper half

Maximum Largest non-outlier value Largest value ≤ Q3 + 1.5×IQR

Outlier Detection

Statistics Study Guide Page 11 © 2024 Study Notes


Statistics Study Guide Comprehensive Reference

Using the 1.5 × IQR Rule:


• Lower Fence = Q1 − 1.5 × IQR
• Upper Fence = Q3 + 1.5 × IQR
• Any value beyond these fences is considered an outlier.

EDA TOOLS SUMMARY


Histograms (shape) | Box Plots (spread & outliers) | Scatter Plots (relationships) | Bar Charts (categories) | Heat

Statistics Study Guide Page 12 © 2024 Study Notes


Statistics Study Guide Comprehensive Reference

9 Absolute and Relative Measures of Dispersion

Two Categories of Dispersion Measures


Absolute measures express dispersion in the same units as the data. Relative measures express
dispersion as a ratio or percentage — allowing comparison between different datasets.

ABSOLUTE MEASURES RELATIVE MEASURES

Definition Dispersion in original data units Unit-free ratio (dimensionless)

Use Case Describing spread within ONE dataset


COMPARING spread of TWO or more datasets

Affected by units? YES NO

Examples Range, MD, Variance, SD, IQR CV, Coefficient of Range, Coefficient of MD

Absolute Measures — Formulas

Measure Formula Notes

Range R = X_max – X_min Simplest; highly sensitive to outliers

Mean Deviation
MD = Σ|xi – x■| / n Uses absolute values; ignores sign
(about Mean)

Mean Deviation
MD = Σ|xi – Median| / n Minimum value of MD is about median
(about Median)

Variance s² = Σ(xi–x■)² / (n–1) Penalizes large deviations more (squared)

Standard Deviation s = √[Σ(xi–x■)²/(n–1)] Most important; same units as data

IQR IQR = Q3 – Q1 Range of middle 50%; robust to outliers

Relative Measures — Formulas


Measure Formula Interpretation

Coefficient of
CV = (s / x■) × 100% Most widely used; lower CV = more consistent data
Variation (CV)

Coefficient of
CR = (Max–Min) / (Max+Min) Ranges from 0 to 1; 0 = no spread
Range

Coefficient of
CMD = MD / x■ (or MD / Median)Relative dispersion around mean or median
Mean Deviation

Coefficient of
CQD = (Q3–Q1) / (Q3+Q1) IQR-based; robust to outliers
Quartile Deviation

Statistics Study Guide Page 13 © 2024 Study Notes


Statistics Study Guide Comprehensive Reference

CLASSIC EXAMPLE
Dataset A: Ages (mean=30, SD=5) | Dataset B: Incomes in PKR (mean=50000, SD=8000) CV_A = 5/30×100 =
16.7% | CV_B = 8000/50000×100 = 16% Interpretation: Incomes are SLIGHTLY more consistent than ages (lower
CV).

Statistics Study Guide Page 14 © 2024 Study Notes


Statistics Study Guide Comprehensive Reference

10 Qualitative-wise Quantitative Analysis

What is it?
This approach involves computing quantitative (numerical) measures separately for each category
of a qualitative variable. It allows us to compare distributions of a numeric variable across different
groups or categories.

Example: Comparing exam scores (quantitative) separately for Male and Female (qualitative) students.

Steps in Qualitative-wise Quantitative Analysis


Step Action Example

1 Identify qualitative variable (grouping factor) Gender: Male / Female

2 Identify quantitative variable to be analyzed Exam Score (0–100)

3 Split data by each category Group A: Male scores | Group B: Female scores

4 Compute measures for each group separately Mean, Median, SD, Min, Max for each group

5 Compare and interpret the groups Which group scored higher? More variable?

Example: Exam Scores by Department


Statistic Engineering (n=30) Medicine (n=28) Arts (n=25)

Mean (x■) 72.4 80.1 65.3

Median 74 81 66

Mode 75 82 63

Standard Dev (s) 10.2 7.8 12.5

Min 48 61 38

Max 95 98 90

CV (%) 14.1% 9.7% 19.1%

INTERPRETATION
Medicine students have the highest mean (80.1) and lowest CV (9.7%) — they score highest AND most
consistently. Arts students have the most variability (CV=19.1%) — scores are most spread out in this group.

Common Tools for This Analysis


Tool / Technique What It Does

Statistics Study Guide Page 15 © 2024 Study Notes


Statistics Study Guide Comprehensive Reference

Group-wise Descriptive Stats Mean, Median, SD computed per group (as shown above)

Side-by-side Box Plots Visually compares distributions of groups; shows medians, spreads & outliers

Grouped Bar Charts Compares frequencies or means across categories

Contingency Tables Cross-tabulates two categorical variables to find associations

One-way ANOVA Tests if means differ significantly across 3+ groups (inferential)

t-Test (2 groups) Tests if means of exactly two groups are significantly different

WHY IS THIS IMPORTANT?


Many real-world questions require comparing groups: Do men and women earn differently? Does treatment A work

QUICK REFERENCE SUMMARY


Topic Key Takeaway

1. Statistics Descriptive (describe data) vs Inferential (draw conclusions)

2. Data Types Qualitative (Nominal/Ordinal) | Quantitative (Discrete/Continuous) | NOIR scales

3. Sample vs Population Population=all (N,µ,σ) | Sample=subset (n,x■,s) | Must be representative

4. Parameter vs Statistic Parameter=population (fixed, unknown) | Statistic=sample (calculated, varies)

5. Data Collection Primary (firsthand) vs Secondary (existing) | 6 main primary methods

6. Representation FDT | Bar (Simple/Multiple/Component) | Pie | Histogram (no gaps!)

7. Central Tendency Mean (sensitive) | Median (robust) | Mode (nominal) | Skewness rule

8. EDA & Box Plot 5-number summary: Min,Q1,Median,Q3,Max | Outliers: 1.5×IQR rule

9. Dispersion Absolute (same units) | Relative (unit-free, use CV to compare datasets)

10. QW Quant Analysis Compute stats per qualitative group → Compare → Interpret differences

Statistics Study Guide Page 16 © 2024 Study Notes

You might also like