Applied Statistics for
Data Analytics
Module 3: Confidence intervals
Confidence intervals
Module 3 introduction
Module 3 outline
Confidence intervals Confidence intervals LLMs for
Inferential statistics
for means for proportions confidence intervals
Define inferential Construct intervals Interpret results
Construct intervals
statistics
Interpret intervals Interpret intervals Create simulation
interfaces
Manage precision Perform
of the estimate calculations
Create
visualizations
Diamond prices
Confidence intervals
Inferential statistics
Scenario Figure out the rate of employee satisfaction:
Satisfaction is likely high
100 employees 82.0% satisfied
500 employees 91.0% satisfied
10,000
employees
500 employees 91.0% satisfied
500 employees 92.1% satisfied
You
Data Analyst
500 employees 88.7% satisfied
Sean Barnes
Larger samples provide more reliable
Inferential statistics ●
estimates than smaller samples
500 100
Population Sample data Inference
● Different samples show variability, even
10,000 when drawn from the same population
Employees Employee survey 91.0% 92.1% 88.7%
satisfied satisfied satisfied
● Allows you to quantify your level of
confidence in your estimate
Sean Barnes
Descriptive statistics Inferential statistics
State facts about sample Use sample data to draw conclusions
about population
● “In a sample of 100 employees, 82 ● Based on a sample of 100
said they were satisfied.” employees, employee satisfaction
among all 10,000 employees is likely
● “A survey of 160 parents found that
between 88% and 91%.
the parents with newborns got 6.1
hours of sleep while the parents with ● A survey of parents concluded that
older children got 8.2 hours.” parents of older children sleep 2
hours more per night compared with
parents of newborns.
Describe the characteristics of a sample Use characteristics of sample to make
conclusions about population
Sean Barnes
Inferential statistics
● Use probability to draw conclusions about
population based on sample statistics
○ Taking into account:
Size of sample
Variability of the sample
Sample
Other factors
Sean Barnes
Inferential vs. descriptive statistics
Low-stakes
Descriptive statistics decision
✅ “82 out of 100 employees
are satisfied”
⛔ Can’t generalize that to all
employees
10,000
employees
Inferential statistics
✅ “The true proportion of
satisfaction is 88-91%”
✅ Population parameter is
likely to fall within those High-stakes
two values decision
Sean Barnes
Confidence intervals
Point & interval estimates
Point estimates
x̄ s p̂
Sample mean Sample standard Sample proportion
deviation
Represent “best guess” about
population parameters:
µ σ p
Population mean Population Population proportion
standard deviation
Sean Barnes
Point estimates Interval estimates
❌ Don’t contain information about the ✅ Contain information about how
confidence of estimate confident you can be
Example: Example:
Random sample of 25 movies from 2013
Arriving within 10 to 15 mins
x̄ of durations = 121 minutes
How certain can you be that true Arriving in exactly 10 mins
population mean is exactly 121?
Sean Barnes
Visualizing intervals
Employee Satisfaction by Department
Error bars: visual representation Range of values for
of an interval estimate true proportion
84.8
80
Proportion of Satisfied Employees
● Inference about true 72.7
population proportion based
on sample 60 60.3
Least precise Most precise
● Wider range ● Narrowest
🔁 If repeated many times, 40
expect true population value
to fall within this range
20
0
Engineering Design Sales
Sean Barnes
Confidence intervals
Sampling distributions &
the central limit theorem
Scenario
💬 Task: Estimate average score on a
professional certification exam
🏆 Possible scores: 0 to 100
Sample 1 Sample 2
📝 Ask 50 random people 📝 Ask 50 random people
their scores their scores
You
Data Analyst
x̄ Same value? x̄
of scores of scores
Sean Barnes
Central limit theorem
n > 30 Increase your Estimate becomes
If you take sufficiently large samples from any sample size more precise
distribution and calculate their means:
● Sample means will be normally distributed Standard deviation of your sample mean
(standard error of the mean)
● Mean of distribution will equal μ
True population standard deviation
SE =
Square root of your sample size
Central tendency of the
sample mean is around
the population mean As n gets larger, √n
Normally gets larger, but at a
distributed slower rate
Sean Barnes
Sample size and precision
Diminishing returns for
precision
√n grows
quickly √n levels off
Large gains in
precision from adding
a few more values
As n gets larger As n gets larger and larger
Sean Barnes
Central limit theorem
Central limit theorem applies to:
Even if sample data is not normal:
● Sample mean x̄
● Sample means will be normal as long as n is
sufficiently large ● Sample proportion p̂
In some cases:
Distribution of ● Sample variance s2
sample means
● Sample standard deviation s
Sean Barnes
Confidence intervals
Demo: confidence intervals
in action
Confidence intervals
Confidence intervals
Scenario Task: Figure out how long it takes to deliver
the pastries
● Before 7am
Sample data:
🗓 Monitor the delivery truck for 30 days
⏰ Record time it takes to get from bakery
to zoo
x̄ = 43 minutes s = 11 minutes
Maybe due to chance 🐇 Unusually fast
You your deliveries were: 🐢 Unusually slow
Data Analyst
Sean Barnes
Confidence interval x̄ = 43 minutes s = 11 minutes
Best guess estimate for
Sampling Distribution of Delivery Time Means
mean delivery time (μ)
43 Standard error (SE) s 11
≈ = ≈2
x̄ - 2×SE x̄ + 2×SE √n √30
= 43 - 2×2 = 43 + 2×2
= 39 = 47 Best guess: s
Variability Sample size Confidence
With 95% confidence, the true 95% of
mean delivery time is between 39 possible means
and 47 minutes.
1 SE 1 SE 1 SE 1 SE
Sean Barnes
Confidence interval
Mean delivery time (μ) ● Range of values used to estimate a
population parameter
39 47
● Quantifies the uncertainty of an estimate
Confidence by the relative width of the range
interval
Broader range Narrow range
High uncertainty More precision
Understanding the possible average
delivery time can help with precise
scheduling
Sean Barnes
Scenario
● Before 7am Your sample is one of
thousands of possible
samples
50% chance of 1σ 1σ 50% chance of
below average being longer
time than average
2σ 2σ
3σ 3σ
You
Data Analyst
50% below the mean 50% above the mean
Sean Barnes
Probability that sample mean 43 is
Scenario within 2 standard deviations of μ?
● Before 7am
Sample mean x̄
approximates the true
population mean Sampling distribution
is normally distributed
2σ 2σ
You
Data Analyst Based on the two sigma rule: 95% chance
Sean Barnes
Confidence intervals
Mechanisms of confidence
intervals
Example With 95% confidence, the true average delivery
time is between 39 and 47 minutes
● 95% confidence level reflects
how reliable your estimation
method is
● A method that's designed to be
right about 95% of the time
● Population parameter is:
● Fixed, but unknown to you
● In this particular interval or
it’s not
36 38 40 42 44 46 48 50
Delivery times (minutes)
Sean Barnes
Confidence interval for the population mean
z-score
● Controls how “confident” you
Where: are that interval contains the
population mean
● x̄ - sample mean
● s - sample standard deviation ● Number of standard
deviations from mean in
● n - sample size standard normal distribution
● z - z-score value from the
standard normal distribution
Sean Barnes
Confidence interval for the population mean
To calculate a 95% confidence interval:
z-score: 1.96
● The two sigma rule is just an
estimate
● z = 2 is associated with slightly 2σ 2σ
higher confidence than 95%
● Use z = 1.96 to be more precise
Estimate 95% confidence interval
Sean Barnes
Confidence interval for the population mean
To calculate a 95% confidence interval: Confidence interval z score
90% 1.645
95% 1.96
99% 2.576
Margin of error
⬆ Higher level of confidence
● Constructs the confidence interval
⬆ Use a higher z score
● Helps gauge the precision of the
estimate ⬆ Generate a wider confidence interval
Sean Barnes
Choosing a confidence level
90% confidence level 95% confidence level 99% confidence level
Used when missing true value is Balances confidence with Used when you want to
less important potential for error minimize the risk of error
Estimate avg. user rating Estimate avg. delivery time Estimate pollution in a river
Users on average rate a The average delivery time Estimating pollution
new feature 7.2 out of 10 is between 39 and 47 in a river
minutes
✅ A less precise estimate is ✅ Don’t need to schedule too ✅ Can help reduce risk of
acceptable to help make initial much buffer time harmful impact
decisions
Sean Barnes
Confidence intervals
Understanding
margin of error
Margin of error
Determines how wide your
confidence interval is
✅ A narrower confidence interval is more desirable
because it means you have a more precise estimate
Depends on three factors:
● Desired confidence level 36 38 40 42 44 46 48 50
i.e. the amount Delivery times (minutes)
● Standard deviation of variability in
your data
● Sample size
🗓 Precision allows for better
scheduling
Sean Barnes
Desired confidence level = = =
90% confidence level z = 1.645
Confidence Intervals for Delivery Times
z x 1 = 1.645 x 1 = 1.645
95% confidence level z = 1.96
Probability Density
z x 1 = 1.96 x 1 = 1.96
3.29 mins
99% confidence level z = 2.576
3.92 mins
z x 1 = 2.576 x 1 = 2.576
5.152 mins
36 38 40 42 44 46 48 50
Delivery Time (minutes)
Sean Barnes
Width vs. Confidence level = = =
90% confidence level
Confidence Intervals for Delivery Times
95% confidence level
Probability Density
● 19% wider
● Gain 5% in relative confidence
3.29 mins
99% confidence level
3.92 mins
● 31% wider
5.152 mins
● Gain 4% in relative confidence
36 38 40 42 44 46 48 50
Delivery Time (minutes)
Sean Barnes
95% confidence interval
Variability
= =
s = 10 Confidence Intervals for Delivery Times
3.92 mins
7.84 mins
s = 20
Probability Density
11.76 mins
s = 30
If data has a lot of variability, expect a
less precise estimate. 36 38 40 42 44 46 48 50
Delivery Time (minutes)
Sean Barnes
95% confidence interval
Variability
= =
s = 10
Confidence Intervals for Delivery Times
3.29 mins
7.84 mins
s = 20
Probability Density
11.76 mins
s = 30
City bus system Rural bus system
✅ Smaller variability makes it easier to
If data has a lot of variability, expect a
estimate the true average arrival time
less precise estimate. 36 38 40 42 44 46 48 50
Delivery Time (minutes)
Sean Barnes
Sample size
n = 100 Confidence Intervals for Delivery Times
3.92 mins
2.78 mins
Probability Density
n = 200
2.62 mins
n = 300
36 38 40 42 44 46 48 50
Delivery Time (minutes)
Sean Barnes
Sample size
A larger sample:
Confidence Intervals for Delivery Times
● Allows you to construct a narrower
confidence interval 3.92 mins
● Reduces the size of the margin of 2.78 mins
error
Probability Density
2.62 mins
● Produces diminishing returns
n = 100
n = 200 → Reduced interval by 29%
n = 300 → Reduced interval by 6%
36 38 40 42 44 46 48 50
Delivery Time (minutes)
Sean Barnes
Sample size
● This relationship is a basic
fact of statistics
● Often need a relatively
Negative
relationship small sample for
Decreases very large population
margin of error
Slope flattens out
the higher n goes ● Adding more and more data
will narrow interval
● In very high sample sizes, you
will observe diminishing impact
Larger sample size
Sean Barnes
Summary
You can achieve a narrower confidence interval in three ways:
● By lowering your confidence level, which can increase your
chances of missing the true value
● By working with data that has less variability, which is
generally out of your control
● By increasing your sample size, though this approach has
diminishing returns
Sean Barnes
Confidence intervals
Demo: confidence intervals
for means
Confidence intervals
Confidence intervals
for proportions
Scenario 💬 Problem: “What proportion of deliveries are on time?”
Deliveries that made p
it to the zoo by 7am
True proportion of
on-time deliveries
🚚 Collect a sample of 30 deliveries
📝 Record if they were: ✅ On time ❌ Not on time
You
Data Analyst p̂
Estimate for the true
proportion of p
Sean Barnes
Scenario p̂ = 0.6 18 out of 30
on-time deliveries
Sampling distribution of p̂
Normally
distributed
2σ 2σ
You
Data Analyst
p
95% chance
Sean Barnes
Low variability
Standard error
The standard error for the proportion is:
Square root
Use p̂ to estimate p̂ = 0.9 p̂ (1 - p̂ ) = 0.09
Sample size
Higher variability
● p(1-p) represents the variability
● Higher variability when p̂ is closer to 0.5
● Take the square root so number is at the p̂ = 0.6 p̂ (1 - p̂ ) = 0.24
original scale
More even mix of success and failures
Sean Barnes
Confidence interval
The interval is defined as:
✅ p̂ – estimate of true proportion of p
✅ – standard error
Margin of error
Confidence interval for the mean:
represents the uncertainty
in your estimate
Sean Barnes
Confidence interval
The interval is defined as:
✅ p̂ – estimate of true proportion of p
✅ – standard error
Confidence interval for the mean:
Sean Barnes
Confidence interval
18 out of 30
p̂ = 0.6 on-time deliveries
With 95% confidence, the true
proportion of on time deliveries
is somewhere within this range
Wide interval is reflected in:
● High variability of the data
● Relatively small sample size
0.2 0.4 0.6 0.8 1.0
Sean Barnes
Confidence intervals
Demo: confidence
intervals for proportions
Confidence intervals
Interpretation with LLMs
Confidence intervals
Simulating random
sampling with LLMs
Confidence intervals
Inference and visualization
with LLMs