ISOM 2500 -
BUSINESS
STATISTICS
PROF. LANCELOT JAMES
CONFIDENCE INTERVALS
READINGS: CHAPTER 15
OVERVIEW
Discrete
Random
Variables
Descriptive Random
Probability
Statistics Variables
Continuous
Random
Variables
Sampling
Distribution
Simple Linear Hypothesis
Estimation
Regression Testing
Confidence
Intervals
GOALS FOR THIS TOPIC
Know what is Confidence Interval and how to interpret it
Know how to estimate sample size needed to achieve a specific
margin of error
WHY A POINT ESTIMATE IS NOT ENOUGH?
̂
Three drawbacks of point estimators, 𝑥 and 𝑝
It is almost certain that the estimate will be wrong based on a single sample
What if we want to know how close this estimator is to the parameter?
Intuitively, larger samples will produce more accurate results, but point
estimators alone does not fully reflect the effect of large sample size
INTERVAL ESTIMATE
Imagine you try to capture a butterfly in the dark. Is it better to use
A dart?
Or a large net?
Instead of using a point estimate (using a single value to estimate the
population parameter), we can use an interval estimate, where we use a range
of values to estimate the population parameter (we say that the population
parameter is within our interval estimate)
INTERVAL ESTIMATE
An interval estimate draws inferences about a population by
estimating the value of an unknown parameter using an interval
Based on a sample from the population, a confidence interval is a
range of plausible values for a parameter
𝜇
MOTIVATING EXAMPLE:
LAUNCH A NEW CREDIT CARD
To launch an affinity credit card, the contemplated launch process proposes
sending pre-approved application to 𝑁= 100,000 alumni of a large university
(population)
Two parameters of the population determine whether the card will be
profitable:
𝑝, the proportion who will return the application
𝜇, the average monthly balance carried by those who accept the card
To estimate the parameters, the credit card issuer sent pre-
approved application to a sample of 1,000 alumni. Of these,
140 accepted the offer and received a card
Variable Statistic
Number of offers 1000
Number accepted 140
Proportion who accepted ̂
𝑝 = 0.14
Average balance 𝑥 = 1990.5
SD of balance 𝑠 = 2833.33
PART I: CONFIDENCE INTERVAL FOR THE
PROPORTION
Recall the sampling variability issue:
Each time we take a random sample from a population, we are likely to get a different set of
individuals and calculate a different statistic. This is called sampling variability.
If we take a lot of random samples of the same size from a given population, the variation from sample
to sample - the sampling distribution - will follow a predictable pattern.
SAMPLING DISTRIBUTION OF THE SAMPLE
PROPORTION
̂ count of successes in the sample 𝑋
𝑋 ∼ Bin(𝑛, 𝑝), and 𝑝 = 𝑛
=
𝑛
̂
For instance, in an SRS of 50 students from an undergrad class, 10 are exchange student. 𝑝 =
10/50 = 0.2 (proportion of exchange student in the sample)
𝑝(1−𝑝)
𝜇 ̂ = 𝑝, 𝜎 ̂ =
𝑝 𝑝 𝑛
Because the mean is 𝑝, we say that the sample proportion is an unbiased estimator of the
population proportion 𝑝
SAMPLING DISTRIBUTION OF THE SAMPLE
PROPORTION
Normal Approximation to Binomial
distribution:
Under certain conditions (check slides
3b), Bin(𝑛, 𝑝) can be approximated by
𝑁(𝑛𝑝, 𝑛𝑝(1 − 𝑝))
̂
Then the sampling distribution of 𝑝 is
𝑝(1−𝑝)
approximately 𝑁(𝑝, )
𝑛
HOW TO SET THE LENGTH L
̂
𝑝−𝑝
Step 1: based on the sampling distribution, we know that 𝑃(−1.96 ≤ ≤ 1.96) =
𝑝(1−𝑝)
𝑛
0.95
𝑝(1−𝑝) ̂ 𝑝(1−𝑝)
Step 2: moving the terms 𝑃(𝑝 − 1.96 ≤ 𝑝 ≤ 𝑝 + 1.96 ) = 0.95
𝑛 𝑛
̂ ̂ 𝑝(1−𝑝) ̂ 𝑝(1−𝑝)
Step 3: switch the role of 𝑝 and 𝑝𝑃(𝑝 − 1.96 ≤ 𝑝 ≤ 𝑝 + 1.96 ) = 0.95
𝑛 𝑛
̂ ̂ ̂
𝑝−𝐿 𝑝 𝑝+𝐿
̂ 𝑝(1−𝑝) ̂ 𝑝(1−𝑝)
𝑃(𝑝 − 1.96 ≤ 𝑝 ≤ 𝑝 + 1.96 ) = 0.95
𝑛 𝑛
𝑝(1−𝑝)
So 𝐿 = 1.96
𝑛
And the 95% confidence interval for 𝑝 is
̂ 𝑝(1 − 𝑝) ̂ 𝑝(1 − 𝑝)
[𝑝 − 1.96 , 𝑝 + 1.96 ]
𝑛 𝑛
In general, the 100(1 − 𝛼)%confidence interval for 𝑝 is
̂ 𝑝(1−𝑝) ̂ 𝑝(1−𝑝)
[𝑝 − 𝑧𝛼Τ2 ,𝑝 + 𝑧𝛼Τ2 ]
𝑛 𝑛
1 − 𝛼 : confidence level
𝑧𝛼Τ2 : critical value. For 𝛼 between 0 and 1, 𝑧𝛼Τ2 is the value such that there is an 𝛼 Τ2 probability of being above
that value in a normal distribution, i.e., the upper-percentile.
𝑝(1−𝑝)
𝑧𝛼Τ2 : margin of error. It is also the L we defined
𝑛
̂
With unknown 𝑝, we plug in 𝑝:
̂ ̂ ̂ ̂
̂ 𝑝(1 − 𝑝) ̂ 𝑝(1 − 𝑝)
[𝑝 − 𝑧𝛼Τ2 , 𝑝 + 𝑧𝛼Τ2 ]
𝑛 𝑛
CHECK THE CONDITIONS
Simple Random Sample (SRS):the observed sample is a
simple random sample from the relevant population
̂ ̂
Sample size condition: Both 𝑛𝑝 and 𝑛(1 − 𝑝) are larger
than or equal to 5
̂ ̂
̂ 𝑝(1−𝑝)
In the data, the estimated standard error is 𝑆𝐸(𝑝) = =
𝑛
0.14(1−0.14)
≈ 0.011
1000
The 95% CI for 𝑝 is [0.14 − 1.96 ∗ 0.011,0.14 + 1.96 ∗ 0.011] =
[11.84%, 16.16%]
Interpretation to non-expert: we are 95% confident that the population
proportion that will accept this offer is between about 12% and 16%
Is it correct to say 𝑃(11.84% ≤ 𝑝 ≤ 16.16%) = 95%?
STATISTICAL INTERPRETATION
𝑝
We say “we are 95% confident that…”
But NEVER say “the probability that the true proportion
𝑝 is between 12% and 16% is 0.95”
For a realized confidence interval, the true mean 𝑝 is
either in there or not
Statistically, it really means: if you line up the 95%
confidence intervals from many, many samples, 95% of
these intervals would cover the population parameter 𝑝
PART II: CONFIDENCE INTERVAL FOR THE MEAN
𝜎2
Recall from sampling variability that 𝑋 ∼ 𝑁(𝜇, )
𝑛
𝑋−𝜇
Therefore 𝑃(−𝑧𝛼Τ2 ≤ ≤ 𝑧𝛼Τ2 ) = 1 − 𝛼 , suppose we use the upper-percentile notation.
𝜎Τ 𝑛
𝜎 𝜎
Rearranging the terms: 𝑃(𝑋 − 𝑧𝛼Τ2 ≤𝜇≤𝑋+ 𝑧𝛼Τ2 ) =1−𝛼
𝑛 𝑛
𝜎 𝜎
Thus, with known 𝜎 , the 100(1 − 𝛼)%CI for 𝜇 is [𝑋 − 𝑧𝛼Τ2 , 𝑋 + 𝑧𝛼Τ2 ]
𝑛 𝑛
𝑋 —— point estimate; 𝑧𝛼Τ2 —— critical value;
𝜎 𝜎
—— standard error; 𝑧𝛼Τ2 —— margin of error
𝑛 𝑛
SUBSTITUTE 𝜎 BY 𝑠
𝜎 𝑠
𝑆𝐸(𝑋) = is replaced by 𝑆𝐸(𝑋) =
𝑛 𝑛
𝑋−𝜇
Recall the sampling distribution ∼ 𝑡𝑛−1 , for 𝑛 < 30
𝑠Τ 𝑛
Therefore the 100(1 − 𝛼)%CI should be adjusted as
𝑠 𝑠
[𝑋 − 𝑡𝛼/2,(𝑛−1) , 𝑋 + 𝑡𝛼/2,(𝑛−1) ]
𝑛 𝑛
What is 𝑡𝛼/2,(𝑛−1) ?
𝑡𝛼/2,𝑛−1 is the critical point of a t-
distribution with df n-1.
The area to the right of 𝑡𝛼/2,𝑛−1 under
a 𝑡𝑛−1 curve is 𝛼 Τ2.
Example: for the 95% CI, 𝛼 = 0.05
and suppose 𝑛 = 20, what is the
critical value?
BACK TO THE CREDIT CARD EXAMPLE
The monthly balance for the 𝑛 = 140 customers: 𝑥 = 1990.5 and 𝑠 =
2833.33
𝑋−𝜇
Since n>30, we may treat ∼ 𝑁(0,1)
𝑠Τ 𝑛
The 95% CI for the balance is [1990.5 − 1.96 ∗ (2833.33/ 140), 1990.5 + 1.96 ∗
(2833.33/ 140)] = [1523.553,2457.447]
COMMON CONFUSIONS
Three examples of incorrect interpretation of the confidence interval
[1523.553, 2457.447] in the credit card case
“95% of all customers keep a balance between $1523.553 and $2457.447”
“The mean balance of 95% of sample of this size will fall between
$1523.553 and $2457.447”
“The mean balance 𝜇 is between $1523.553 and $2457.447”
GENERAL FORMULA
Point Estimate ± (Critical Value)x(Standard Error)
𝜎
Known 𝜎; normal population or large sample (𝑛 ≥ 30) 𝑋ത 𝑍𝛼ൗ
2 𝑛
𝑠
𝑡𝛼ൗ
Unknown 𝜎; normal population and small sample 𝑋ത 2,𝑛−1 𝑛
𝑠
𝑍𝛼ൗ
Unknown 𝜎; large sample 𝑋ത 2 𝑛
𝑍𝛼ൗ 𝑝(1
ො − 𝑝)ො
Unknown p; 𝑛𝑝 ≥ 5 and 𝑛(1 − 𝑝) ≥ 5 𝑝ො 2
𝑛
For all the other cases, we cannot apply the z or t intervals
MANIPULATING CONFIDENCE INTERVALS
Obtaining Ranges for Related Quantities
If [L to U] is a 100(1 – α)% confidence interval for µ,
then [c x L to c x U] is a 100 (1 – α)% confidence interval for c x µ and [c + L to c + U] is a
100(1 – α)% confidence interval for c + µ.
Copyright © 2011 Pearson Education, Inc.
MANIPULATING CONFIDENCE INTERVALS
Changing the Problem
▪ Creating a new variable is preferable to combining confidence intervals.
▪ Consider the following:
Let Y = profit earned from each customer. A customer who does not accept the card costs the bank $8. Each customer who accepts the
card costs the bank $58 but the bank earns 10% on the revolving credit card balance.
Copyright © 2011 Pearson Education, Inc.
MANIPULATING CONFIDENCE INTERVALS
Creating a New Variable
▪ Therefore, profit ($) earned from a customer is
yi = -8 if offer is not accepted
0.10 (Balance) – 58 if offer is accepted.
Copyright © 2011 Pearson Education, Inc.
MANIPULATING CONFIDENCE INTERVALS
Summary Statistics for Profit Earned
For 100,000 offers, the 95% confidence interval for total profit is from $557,000 to $2,017,000.
Copyright © 2011 Pearson Education, Inc.
MARGIN OF ERROR
̂ ̂
𝜎 𝑠 𝑝(1−𝑝)
𝐿= 𝑧𝛼Τ2 , or 𝑡𝛼/2,(𝑛−1) , or 𝑧𝛼Τ2 ,…
𝑛 𝑛 𝑛
Three factors affects the margin of error
Confidence level: 90%, 95%
Sample size: n
The population variation or the sample variation
SAMPLE SIZE REQUIREMENT
To control the Margin Error to be less than some L, there are some
requirements on sample sizes.
For proportions
̂ ̂ 2 ̂ ̂
𝑝(1−𝑝) 𝑧𝛼Τ2 𝑝(1−𝑝)
with unknown 𝑝: 𝐿 = 𝑧𝛼Τ2 →𝑛≥
𝑛 𝐿2
2
𝑧𝛼 Τ2 ×0.25
Sometimes, since 𝑎(1 − 𝑎) ≤ 1/2 for any 𝑎, we can simply use 𝑛 ≥ , which is greater
𝐿2
2 ̂ ̂ 2
𝑧𝛼Τ2 𝑝(1−𝑝) 𝑧𝛼 Τ2 𝑝(1−𝑝)
than or
𝐿2 𝐿2
EXAMPLE - SAMPLE SIZE FOR 𝑝
What sample size is needed to estimate the true proportion of defective be bulbs
within ±5% with 90% confidence? Out of a population of 1000 bulbs, we randomly
select 100 of which, and 30 among them are defective
̂
𝑝 = 0.3, 𝑛 = 100
2 ̂ ̂
𝑧𝛼Τ2 𝑝(1−𝑝) 1.6452 ×0.3×0.7
𝑛≥ = = 227.3
𝐿2 0.052
So the minimum sample size requirement is 228
SAMPLE SIZE REQUIREMENT
To control the Margin Error to be less than some L, there are
some requirements on sample sizes.
For mean 𝜇
2 2
𝜎 𝑧𝛼 Τ2 𝜎
with known 𝜎: 𝐿 = 𝑧𝛼Τ2 →𝑛≥
𝑛 𝐿2
2 2
𝑧𝛼 Τ2 𝑠
with unknown 𝜎: 𝑛 ≥
𝐿2
EXAMPLE - SAMPLE SIZE FOR 𝜇
A lumber company has acquired the rights to a large tract of land containing thousands of
trees
A lumber company needs to be able to estimate the amount of trees it can harvest in a
tract of land to determine whether the effort will be profitable
To do so, it must estimate the mean diameter of the trees
It decides to estimate that parameter to within ±1 inch with 95% confidence
A forester familiar with the territory guesses that the diameters are normally distributed
with a standard deviation 6 inches
Question 1: determine the number of trees he should sample
Question 2: After the sample is taken, the forester discovered
that the sample mean is 25 and the sample standard
deviation is 12. He decided to use the sample standard
deviation instead of his old guess. What is the 90% confidence
interval then?
Q & AS
For the same sample, the width of a confidence interval will
be
A. Narrower for 99% confidence than 95% confidence
B. Wider for a sample size of 100 than for a sample size of
50
C. Narrower for 90% confidence than 95% confidence
Q & AS
With the same confidence level, as standard deviation
increases, samples size need to _______ to achieve a
specified margin of error
A. Increase
B. Decrease
C. Remains the same
CONFIDENCE INTERVAL
Provides range of values based on observations from one sample
Gives information about closeness to unknown population
parameter
Stated in terms of confident
Examples: 90% CI for the 𝜇 , 95% CI for the 𝑝
MORE EXAMPLES - TEXTBOOK EXPENSE
A random sample of 59 students spent an average of $273.20 on Spring 2009
textbooks. The sample standard deviation was $94.40
Determine a 99% confidence interval for the true mean textbook expense of the
population
94.4
273.20 ± 2.574 ∗ ( ) = 273.20 ± 31.58
59
Interpretation: we can be 99% confident that the average amount spent by all
students was between $241.62 and $304.78
MORE EXAMPLES - CALL CENTER
A software distributor has opened a new customer call center to assist customers with
the installation and use of their software. The manager is assuming that the service
time, X, is a normal random variable. To get some info. on the population service
time, the manager has selected a random sample of 81 calls and found the average
service time is 24m, with s=4.2m
What is the 90% CI for the population mean service time?
The manager would like to take another sample and determine a 90% confidence
interval that has a margin of error of 1 minutes. What sample size would be needed to
achieve this goal?
TAKE AWAY FOR TOPIC 4B
Confidence Interval for population mean or Confidence Interval for population proportion
Calculation
Interpretation
Things that affect the width of the confidence interval
Sample variation, population variation
Sample size
Confidence level
Sample size requirement calculation
WHERE DO WE GO FROM HERE?
Discrete
Random
Variables
Descriptive Random
Probability
Statistics Variables
Continuous
Random
Variables
Sampling
Distribution
Simple Linear Hypothesis
Estimation
Regression Testing
Confidence
Intervals