LECTURE 5
SAMPLING DISTRIBUTION
NSNA | 2510 | CDS6224 : Statistical Data Analysis
Important Notation
• A numerical value associated with a population is called a
“parameter”: (FIXED)
• Population mean: 𝜇 (mu)
• Population standard deviation: 𝜎 (sigma)
• Population proportion: 𝑝
• A numerical value computed from a sample is called “statistic”:
(VARIABLE)
• Sample mean: 𝑥ҧ (x-bar)
• Sample standard deviation: 𝑠
• Sample proportion: 𝑝ෝ (p-hat)
NSNA | 2510 | CDS6224 : Statistical Data Analysis
The random nature of samples
• Let’s assume that we take many random samples of same size
from the same population
• If we calculated the sample means from each of these samples,
what do you expect these sample means to be?
• In other words, how is the behavior of sample means calculated
from the many samples of same size taken from the same
population?
• What is the distribution of these sample means?
• Distribution of a sample statistic is called sampling distribution
NSNA | 2510 | CDS6224 : Statistical Data Analysis
Sampling Distribution of a Sample Mean
• The distribution of all possible values taken by the statistics when
all possible samples of a fixed size n are taken from the
population.
• The probability distribution of that statistic
NSNA | 2510 | CDS6224 : Statistical Data Analysis
Sampling Distribution of the Sample Mean
• When we repeatedly take random samples of size 𝑛 from a
population with a known mean (𝜇) and standard deviation (𝜎), the
mean of each sample will vary
• Some sample means will be higher than the population mean, and
others lower – these values form what we call the sampling
distribution of the sample mean ( 𝑥)ҧ
• For any population with mean (𝜇) and standard deviation (𝜎)
• The mean, or center of the sampling distribution of 𝑥,ҧ is equal to the
population mean 𝜇: 𝜇𝑥ҧ = 𝜇
• The standard deviation of the sampling distribution is 𝜎/ 𝑛, where n is the
sample size: 𝜎𝑥ҧ = 𝜎/ 𝑛
NSNA | 2510 | CDS6224 : Statistical Data Analysis
Sampling Distribution of the Sample Mean
• Mean of a sampling distribution of 𝑥ҧ
• There is no tendency for a sample mean to fall systematically above or
below 𝜇, even if the distribution of the raw data is skewed
• Mean of the sampling distribution is an unbiased estimate of the
population mean 𝜇 – it will be “correct on average” in many samples
• Standard deviation of a sampling distribution of 𝑥ҧ
• The standard deviation of the sampling distribution measures how much
the sample statistic varies from sample to sample
• It is smaller than the standard deviation of the population by a factor of
𝑛
NSNA | 2510 | CDS6224 : Statistical Data Analysis
Normally Distributed Populations
• When a variable in a population is
normally distributed, the sampling
distribution of 𝑥ҧ for all possible
samples of size 𝑛 is also normally
distributed
NSNA | 2510 | CDS6224 : Statistical Data Analysis
Example 1
• In a large population of adults, the
mean IQ is 112 with standard
deviation 20. Suppose 200 adults are
randomly selected for a market
research campaign.
• Population distribution: 𝑁(𝜇 =
112; 𝜎 = 20)
• Sampling distribution for n = 200
is 𝑁(𝜇 = 112; 𝜎/ 𝑛 = 1.414)
NSNA | 2510 | CDS6224 : Statistical Data Analysis
Example 2
• Hypokalemia is diagnosed when a patient’s blood potassium level
drops below 3.5mEq/dl. Let’s assume we know a patient whose
measured potassium levels vary daily according to a normal
distribution N(𝜇 = 3.8, 𝜎 = 0.2)
• Scenario 1: One-Time Measurement
• Probability that this patient will be misdiagnosed with Hypokalemia
𝑥−𝜇 3.5−3.8
• 𝑧= = = −1.5, 𝑃 𝑧 < −1.5 = 0.0668 ≈ 7%
𝜎 0.2
• Scenario 2: Average of 4 Daily Readings
• Probability that this patient will be misdiagnosed with Hypokalemia
ҧ
𝑥−𝜇 3.5−3.8
• 𝑧= = = −3, 𝑃 𝑧 < −3 = 0.0013 ≈ 0.1%
𝜎/ 𝑛 0.2/ 4
NSNA | 2510 | CDS6224 : Statistical Data Analysis
Example 2 – cont.
NSNA | 2510 | CDS6224 : Statistical Data Analysis
The Central Limit Theorem
• When randomly sampling from any population with mean 𝜇 and
standard deviation 𝜎, when n is large enough, the sampling
𝜎
distribution of 𝑥ҧ is approximately normal: ~𝑁(𝜇, )
𝑛
Sampling
Population with
distribution of
strongly skewed
𝑥ҧ for 𝑛 = 2
distribution
observations
Sampling
Sampling
distribution of
distribution of
𝑥ҧ for 𝑛 = 25
𝑥ҧ for 𝑛 = 10
observations
observations
NSNA | 2510 | CDS6224 : Statistical Data Analysis
• Large samples are not always attainable
• Sometimes the cost, difficulty, or preciousness of what is studies
drastically limits any possible sample size
• Blood samples/biopsies: No more than a handful of repetitions are
acceptable. Oftentimes, we even make do with just one
• Opinion polls have a limited sample size due to time and cost of
operation. During election times, though, sample sizes are increased for
better accuracy
• Not all variables are normally distributed
• Income, for example, is typically strongly skewed
NSNA | 2510 | CDS6224 : Statistical Data Analysis
Example 3
• Let’s consider the very large database of
individual incomes from the Bureau of Labor
Statistics as our population. It is strongly
right skewed.
• We take 1000 simple random sample of
100 incomes, calculate the sample
mean for each, and make a histogram
of these 1000 means
• We also take 1000 simple random
sample of 25 incomes, calculate the
sample mean for each, and make a
histogram of these 1000 means.
• Which histogram corresponds to samples of
size 100? 25?
NSNA | 2510 | CDS6224 : Statistical Data Analysis
How large a sample size?
• It depends on the population distribution. More observations are
required if the population distribution is far from normal.
• A sample size of 25 is generally enough to obtain a normal sampling
distribution from a strong skewness or even mild outliers
• A sample size of 40 will typically be good enough to overcome extreme
skewness and outliers
NSNA | 2510 | CDS6224 : Statistical Data Analysis