0% found this document useful (0 votes)
7 views29 pages

Understanding Surveys and Statistics

The document provides an overview of survey methodology, including key terminology such as population, parameter, sample, and statistic. It discusses the importance of unbiased sampling methods, the concept of statistical inference, and the Central Limit Theorem for sample proportions. Additionally, it highlights potential biases in survey methods and the criteria for evaluating survey quality.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views29 pages

Understanding Surveys and Statistics

The document provides an overview of survey methodology, including key terminology such as population, parameter, sample, and statistic. It discusses the importance of unbiased sampling methods, the concept of statistical inference, and the Central Limit Theorem for sample proportions. Additionally, it highlights potential biases in survey methods and the criteria for evaluating survey quality.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

7.

1 Learning about the world through surveys

Survey Terminology
A survey is one of the most often encountered applications of statistics. Most news programs,
newspapers, magazines, websites ask and report on surveys or polls daily. If a survey is done
correctly, we can learn a lot.

Population: a group of individuals we wish to study

• Recall from Chapter 1: Population is “all ____________”

Parameter: a numerical value that characterizes some aspect of the population

• Two types of parameters: 1. Population proportion (p)


2. Population mean (𝝁)

Example: Population: All students at West Valley College

Parameter:
1. The population proportion (p) of all students at West Valley College who work
part time

2. The population mean weight (𝜇) of all students at West Valley College

Census: a survey where every member of the population is measured

Note: If the population is relatively small, we can find the exact value of the population
parameter by conducting a census. But for most populations, they are too large for a census (i.e.
it will be too expensive or too time-consuming to conduct a census).

Sample: a collection of individuals taken from the population of interest

• Recall from Chapter 1: Sample is “### ____________”

Statistic: a numerical characteristic of a sample

• Recall from Chapter 1: When studying categorical data, the statistics was a proportion.
• Recall from Chapter 3: When studying numerical data, the statistics was a mean.
• Note: Another word for “statistic” is “estimator” since we use statistics to estimate
parameters.

Example: Sample: 100 students at West Valley College

Statistic:
1. The proportion of the 100 students at West Valley College who work part time
is 80%

2. The mean weight of the 100 students at West Valley College is 123 pounds.
Example 1: Identify each of the following as either a parameter or a statistic, and whether the
survey which gathered the information was a census or sample. Circle the correct answer.

a) Following the 2010 national midterm election, 12% of the governors of the 50 United
States were female.

The proportion 12% is a: parameter or statistics

Because “the governors of the 50 United States” is a: population or sample

Therefore, the survey was a: census or sample

b) In a national survey of 1300 high school students (grades 9 to 12), 32% of the
respondents reported that someone had bullied them at school.

The proportion 32% is a: parameter or statistics

Because “1300 high school students” is a: population or sample

Therefore, the survey was a: census or sample

c) A study of 6076 adults in public rest rooms (in Atlanta, Chicago, New York City, and San
Francisco) found that 23% did not wash their hands before exciting.

The proportion 23% is a: parameter or statistics

Because “6076 adults” is a: population or sample

Therefore, the survey was a: census or sample

d) Of the 12 men that have walked on the moon, the average age at time of their moonwalk
was 39 years, 11 months, 15 days.

The mean age of 39 years, 11 months, 15 days is a: parameter or statistics

Because “the 12 men that have walked on the moon” is a: population or sample

Therefore, the survey was a: census or sample


Statistical Inference: drawing conclusions about a population on the basis of observing
only a small subset of that population

Idea: Observe small sample → Infer → Generalize whole population

Warning: If the population is completely known, then there is no need to make an


inference since we can simply study the entire population

Note: Statistical inference always involves uncertainty, so an important component


of this science is measure our uncertainty. At any time we can collect data and find
the value of a statistic. In contrast, the value of a parameter is almost always
unknown.

Example 2: A survey of 750 random Americans found that 24% believed in reincarnation.
Identify the following:

a) Population:

b) Sample:

c) Parameter:

d) Statistic:

e) Statistical inference:

Notation
Statistics (based on data) Parameters (typically unknown)
Sample mean = 𝑥̅ Population mean = 𝜇
Sample standard deviation = 𝑠 Population standard deviation = 𝜎
Sample proportion = 𝑝̂ Population proportion = 𝑝
Bias in Survey Methods

Biased method: if the survey method has a tendency to produce an untrue value

Two types of bias in surveys:


1. Sampling Bias: occurs when the method produces a sample that is not
representative of the population. Types of sampling bias:
a) Voluntary-response bias: people tend to respond only if they have
strong feelings about the result
b) Nonresponse bias: those selected for the survey refuse to answer
c) Convenience bias: a sample made up of people who are easy to reach
2. Measurement bias: occurs when the survey questions do not produce true
answers (i.e. confusing wording, misleading questions)

Example 4: Identify the possible bias if the population is all Americans. Circle all that
apply.
a) A student asked all 250 of her Facebook friends if they preferred Facebook to
Twitter.

Bias: Voluntary response Nonresponse Convenience Measurement

b) A researcher asked 500 randomly selected people, “Are you in favor of the unfair tax
burden that the hard working successful business people have so that the lazy
unemployed can receive a paycheck without working?”

Bias: Voluntary response Nonresponse Convenience Measurement

c) On July 4, CNN posted on their website a question asking if they supported the
current US military operations. 18,943 people responded.

Bias: Voluntary response Nonresponse Convenience Measurement

d) 100 randomly selected Americans were asked by a researcher, “Do you currently
have a sexually transmitted disease?”

Bias: Voluntary response Nonresponse Convenience Measurement

e) A researcher stood outside a grocery store and asked 250 shoppers, “Do you eat out
at a restaurant at least three times per week?”

Bias: Voluntary response Nonresponse Convenience Measurement

f) Gallop randomly selected 1000 phone numbers from the yellow pages and then
called to ask if they supported government funding of high speed rail
Bias: Voluntary response Nonresponse Convenience Measurement
Critiquing Surveys in the Media
When presented with results of a survey, ask yourself the following questions:
1. What percentage of people who were asked to participate actually did so?
2. Did the researchers choose people to participate in the survey or did the people
themselves choose to participate?
3. Did the researcher leave out whole segments of the population who are likely to
answer the question differently from the rest of the population?

Random Samples
How do we collect a sample that has as little bias as possible and is representative of the
population? Take a random sample.

Random sample: taken in such a way that every individual in the population is equally
likely to be chosen

Idea: Random sample → Minimize bias

To find a random sample:


1. Assign a number to each and every member of the population
2. Use a random number generator (like a calculator) to select our sample

Example 5: A coach must select two players to serve as captains at the beginning of a
soccer match. He has 12 players on his team and, to be fair, wants to randomly select the
two captains.
1. List all players’ names and assign each a number 1 to 12
2. Go to the calculator and push MATH  PRB then choose option 5: randInt(
3. Enter randInt(1, 12, 2) to randomly select 2 players out of those labeled 1 to 12
4. Identify the two captains randomly selected
7.2: Measuring the Quality of a Survey

In order to judge a survey, statisticians instead evaluate the method used for the survey,
and not the outcome. The method must be unbiased/highly accurate (center) and
highly precise/low variation (spread).

Accuracy and Bias


o Bias is a measure of the accuracy (that is, bias measures the distance between the
center of the sampling distribution and the population parameter).
o The more unbiased the sample, the more accurate the measure of the estimate.
o Random selection from the population ensure unbiased results.

Example 1: Suppose basketball players from the NBA are measured to estimate the
proportion of all Americans who are taller than 6 feet.
a) Will there be bias in this measure?
b) Will the measure overestimate or underestimate the true proportion?

Precision and Standard Error


o Standard error (SE) is a measure of precision (where precision is reflected in the
spread of the sampling distribution and is measured by using the standard deviation
of the sampling distribution, called the standard error)
o The smaller the standard error (i.e. the less variation), the more precise the
estimate will be from sample to sample

Example 2: In order to estimate the proportion of tall people in the US, we use 3 randomly
selected Americans.
a) Is this measure unbiased? Will this estimate be close to or far from the true
proportion of tall people in the US?
b) Will the standard error for this estimate be large or small?
Introduction to a Sampling Distribution

Recall: The symbol for a sample proportion is

Flipping a fair coin and landing on a tail

The theoretical probability a fair coin will land on a tail is P(tail) =

Sample size: n =

Trial Proportion of tails from that trial


Trial 1

Trial 2

Trial 3

Trial 4

Trial 5

Rolling a fair six-sided dice and landing on a 5

The theoretical probability a fair six-sided dice will land on a 5 is P(land on 5) =

Sample size: n =

Trial Proportion of 5s from that trial


Trial 1

Trial 2

Trial 3

Trial 4

Trial 5
Sampling Distribution

o The previous example illustrates that the sample proportion 𝒑


̂ is a random
variable.

o This probability distribution of 𝒑


̂ is called sampling distribution.

o The sampling distribution will have a shape, center, and spread just like other
numerical distributions

Key Points

o The mean of all sample proportions 𝒑 ̂ always equals the population proportion
(That is, the results are accurate and there is no bias). This is represented by the
mathematical formula: 𝝁𝒑̂ = 𝒑

o The standard error (SE) will be smaller for larger sample sizes, therefore
improving precision. This is represented by the mathematical formula:
𝒑(𝟏 − 𝒑)
𝝈𝒑̂ = 𝑺𝑬 = √
𝒏

o The size of the population has no effect on the distribution of all sample
proportions, as long as the population size is at least 10 times larger than the
sample size.

Example 3: Only 65% of insured women get annual Pap tests. Find the mean and standard
error for the sampling distribution of the sample proportion of women who get annual Pap
tests with a sample size of 500.
a) What value should we expect for our sample proportion?

b) What is the standard error?

c) We expect __________% of women to get annual Pap tests, give or take ___________%.
7.3: The Central Limit Theorem for Sample Proportions

Central Limit Theorem


o The Central Limit Theorem is central to the study of statistical inference. Without it,
none of the inferential methods in this class would work.
o There are two versions of the CLT
1. one for sample proportions (Section 7.3)
2. one for sample means (Section 9.2)

Idea of the Central Limit Theorem for Sample Proportions


o The CLT for sample proportions allows us to approximate a sampling distribution
without having to do simulations
o It tells us that, if some basic conditions are met, the sampling distribution of the
sample proportion is close to the normal distribution

Conditions for the CLT for Sample Proportions


The following conditions must all be met for the CLT for sample proportions to be valid:

1. Random: The sample is collected randomly

2. Large sample: The sample has:


i. at least 10 “successes”
➢ Formula to verify if # of success not given: 𝒏𝒑 ≥ 𝟏𝟎
AND
ii. at least 10 “failures”
➢ Formula to verify if # of failures not given: 𝒏(𝟏 − 𝒑) ≥ 𝟏𝟎

3. Large population: The population size is at least 10 times the sample size
➢ Formula to verify: 𝑵 ≥ 𝟏𝟎𝒏

Central Limit Theorem for Sample Proportions


If we take a large random sample from a population, and if the population size is much
larger than the sample size, then the sampling distribution of 𝑝̂ is approximately

• Shape is normal (symmetric, unimodal, bell-shaped)

• Center (i.e. the mean of 𝑝̂ ) is the population proportion


o Center is 𝝁𝒑̂ = 𝒑

• Spread (i.e. the standard deviation of 𝑝̂ ) is the standard error


𝒑(𝟏−𝒑)
o Spread is 𝝈𝒑̂ = 𝑺𝑬 = √ 𝒏

𝒑(𝟏−𝒑)
• Notation: 𝑵 (𝒑, √ 𝒏
)
Note: If you don’t know the value of p, then you can substitute the value of 𝑝̂ to calculate
the standard error.
o Center is 𝝁𝒑̂ = 𝒑
̂
̂(𝟏−𝒑
𝒑 ̂)
o Spread is 𝝈𝒑̂ = √ 𝒏

Notes about the CLT Conditions


1. Random: Since random sampling is usually impossible to do, other sampling
techniques are often used instead. There is no way to check this just by looking at
the data. You will just have to trust the researcher’s report on how the data was
collated. However, if the sample is collected randomly, we know the sample
proportion 𝑝̂ equals the population proportion p (that is, there is no bias)
2. Large sample: A large sample size is absolutely necessary
3. Large population: Typically the population of interest is very large, but we should
still be aware of this requirement. This is especially true when samples are collected
without replacement

Finding the number of success and number of failures from the total and percentage

Example A: A recent study of 400 students found that 65% carry calculators. Is this
sample large enough to satisfy the Central Limit Theorem condition of “Large Sample”
where both the number of success and number of failures are greater than or equal to 10?
Total: n = Proportion: 𝑝̂ =

Success:

x = # of successes = (total) * (percentage in decimal form) = 𝑛𝑝̂ =

Failure:

# of failures = (total) – (# of successes) = n – x

Is the “Large Sample” conditions satisfied?

Example B: A hospital employs 346 nurses and 35% of them are male. Is this sample large
enough to satisfy the Central Limit Theorem condition of “Large Sample” where both the
number of success and number of failures are greater than or equal to 10?
Total: n = Proportion: 𝑝̂ =

Success:

x = # of successes = (total) * (percentage in decimal form) = 𝑛𝑝̂ =

Failure:

# of failures = (total) – (# of successes) = n – x

Is the “Large Sample” conditions satisfied?


Note: Please write your notes for Examples 1 – 3 on a separate piece of binder paper,
as each example is very long.

Example 1: 200 randomly selected American drivers were asked if they text while driving.
48 admitted they did. Apply the Central Limit Theorem.
a) Check conditions to see if CLT can be applied.
1. Randomly selected?

2. Large Sample:
i. Success =

x = # of success =

ii. Failure =

# of failures = n – x

3. Large population:
There are definitely more than _____________ ______________________________.
(10n) (population)

b) Identify the shape, center, and spread of the sampling distribution of the sample
proportion.
Shape =

Center =

Spread =

c) Find the probability that more than 30% text while driving.

normalcdf(

The probability ___________________________________________ of ________ randomly selected


(restriction) (n)

____________________________________ will _____________________________________ is ____________.


(population) (characteristic) (probability)
Example 2: According to the CDC at the end of 2011, a proportion of .004 people have HIV.
You want to sample 1000 random people. Check conditions for the CLT.

1. Randomly selected?

2. Large Sample:
i. Success =

Note: Since the number of successes is not given, then use the
sample size (n) and proportion (p) to find the number of successes
and number of failures.

x = # of success = np =

ii. Failure =

# of failures = n – x =

3. Large population:
There are definitely more than _____________ ______________________________.
(10n) (population)
Example 3: Time magazine reported (June 17, 2002) that 80% of all brides take the last
name of their new husband. A random sample of 100 brides is obtained.
a) Find the probability that more than 90% of brides take their husband’s last name?

Check conditions to see if CLT can be applied.

1. Randomly selected?

2. Large Sample:
i. Success =

x = # of success =

ii. Failure =

# of failures = n – x

3. Large population:
There are definitely more than _____________ ______________________________.
(10n) (population)

Since all three conditions of the Central Limit Theorem for proportions are met, then
we can identify the shape, center, and spread of the sampling distribution of the
sample proportion.

Shape =

Center =

Spread =

normalcdf(

The probability ___________________________________________ of ________ randomly selected


(restriction) (n)

_______________ will ____________________________________________________________ is ____________.


(population) (characteristic) (probability)
b) What is the probability that, in a random sample of 100 brides, 60 or fewer take
their husband’s last name?

normalcdf(

The probability ___________________________________________ of ________ randomly selected


(restriction) (n)

_______________ will ____________________________________________________________ is ____________.


(population) (characteristic) (probability)
7.4: Estimating the Population Proportion with Confidence Intervals

Statistical Inference

o There are two main types of statistical inference we study in this course:
1. Estimation of parameters through confidence intervals (CI)
2. Decision-making about parameter values through hypothesis test (HT)

o In this section 7.4 we focus on estimating a population proportion

Point Estimate vs. Interval Estimate

o A point estimate is a single number that is our “best initial guess” for the
parameter.

o An interval estimate is an interval of number within which the parameter value


is believed to fall.

Confidence Interval vs. Confidence Level

o A confidence interval is an interval containing the most believable values for a


parameter.

o The probability that this method produces an interval that contains the
parameter is called the confidence level.
o This is a number chosen to be close to 1, most commonly 0.95

o Warning: The confidence level is not the probability that the individual interval
contains the parameter
o Instead, the confidence level tells us how often the estimation method is
successful
Finding a Confidence Interval
o Confidence interval estimates are always of the form:
(𝒑𝒐𝒊𝒏𝒕 𝒆𝒔𝒕𝒊𝒎𝒂𝒕𝒆) ± (𝒎𝒂𝒓𝒈𝒊𝒏 𝒐𝒇 𝒆𝒓𝒓𝒐𝒓) = 𝒑 ̂±𝒎
o The margin of error measures how accurate the point estimate is (that is, tells how
far from the population value the estimate can be), and has the following structure:
𝒎𝒂𝒓𝒈𝒊𝒏 𝒐𝒇 𝒆𝒓𝒓𝒐𝒓 = 𝒎 = z* ∙ 𝑺𝑬
o z* is a critical number telling us how many standard errors to include in the
margin of error

Margin of Error
o The z* in the margin of error formula comes from the Normal distribution, so we
must always check that the Central Limit Theorem for proportion applies
before using this formula. Once verified, the margin of error is found as follows:

Confidence Interval for a Proportion


o A confidence interval for a population proportion has the following structure:
̂ ± z* ∙ 𝑺𝑬
𝒑
o Finding the SE requires we know the value of p, but in real life we don’t know this
(that’s why we’re estimating it!)
o So we substitute our sample proportion in for p and use the estimated standard
error instead
o Thus, a confidence interval for a proportion is given by:
̂(𝟏−𝒑
𝒑 ̂)
̂ ± z* ∙ √
𝒑 𝒏

Example 1: A recent survey of 367 randomly selected college students showed that 312
have Facebook accounts.
a) Point estimate =

b) Standard error =

c) Margin of error for a 95% CI =

d) Confidence interval:

e) Interpretation: We are _____% confident that between _____% and _____% of all

___________________________________ have ________________________________.


Interpreting Confidence Intervals

o Individual Interval: We are _____% confident that between ______% and ______% of all
C-level Lower Upper
_______________ are __________________.
Population Characteristic

o Prediction: The confidence interval gives a set of plausible values for the
population proportion. If a value is not in the confidence interval, we conclude
that is it implausible (not impossible, but very unlikely).

o Method: For every random sample, there corresponds a #% confidence interval.


#% of these confidence intervals will successfully contain the population
proportion and (𝟏 − #)% will not.

Example 2: Each student in a class of 30 students was assigned one random line of 10
random digits (i.e. 0 – 9). Each student then counted the number of even digits in their 10-
digit line.
a) On average, in the list of 10 digits, how many even-numbered digits would each
student find?

b) If each student found an 80% confidence interval for the percentage of even-
numbered digits, how many intervals (out of 30) would you…
i. Expect to capture 50% even digits (i.e. 4 7 3 2 5 6 1 0 4 3)?

ii. Expect not to capture 50% (i.e. 5 7 2 1 0 8 3 5 2 9)?

TI-83/84 Instructions for Confidence Intervals


Finding a Confidence Interval for One Proportion:
1. Verify the conditions of the Central Limit Theorem for Proportions for a valid
interval
a) Random sample
b) Large sample: at least 10 success AND at least 10 failures
c) Large population: population size is at least 10 times the sample size
2. Press STAT arrow over to TESTS.
3. Select A:1-PropZInt… from the list provided, and press ENTER.
4. Enter the number of “successes”(x), the sample size(n), and the confidence level
(C-level) then Calculate.
5. The output should be the Confidence Interval, but also the sample proportion and
the sample size which should match with the rest of your data.
6. Write a sentence interpreting the interval within the context of the problem
Note: Please write your notes for Examples 3 – 5 on a separate piece of binder paper,
as each example is very long.

Example 3: In 2000, the GSS asked: “Are you willing to pay much higher prices in order to
protect the environment?” Of 1154 randomly selected respondents, 518 were willing to.
a) Find and interpret a 95% confidence interval for the population proportion of adult
Americans willing to do so at the time of the survey.

Check conditions to see if CLT can be applied.

1. Randomly selected?

2. Large Sample:
i. Success =

x = # of success =

ii. Failure =

# of failures = n – x

3. Large population:
There are definitely more than _____________ ______________________________.
(10n) (population)

Calculate the Confidence Interval

1-PropZInt(
x, n, C-level

Confidence Interval:

Interpretation: We are _______% confident that between _______% and _______%


(confidence level) (lower) (upper)

of all ___________________________________ are _________________________________________.


(population) (characteristic)

______________________________________________________________________________________.

b) Are 50% of adult Americans willing to pay higher prices in order to protect the
environment?

Since the entire confidence interval _______________________________________________________,

then 50% of adult Americans are ___________________ to pay higher prices to protect the
environment.
Example 4: During the 2006 election, an exit poll of 2705 randomly selected California
voters found that 56.52% had voted to re-elect Arnold Schwartzenegger (AS) for governor.
Based on this data, can we predict that he would win re-election?

Check conditions to see if CLT can be applied.

1. Randomly selected?

2. Large Sample:
i. Success =

x = # of success =

ii. Failure =

# of failures = n – x

3. Large population:
There are definitely more than _____________ ______________________________.
(10n) (population)

Idea: If a confidence level is not given, assume a 95% confidence level.

Calculate the Confidence Interval:

1-PropZInt(
x, n, C-level

Confidence Interval:

Interpretation: We are _________% confident that between ___________% and ___________%


(confidence level) (lower) (upper)

of all ___________________________________ will ______________________________________________________.


(population) (characteristic)

Idea: To win re-election, Schwartzenegger would need over 50% of California votes.

Since the entire confidence interval _______________________________________________________________,

then we predict that Schwartzenegger would ____________ reelection.


Example 5: Suppose you have been hired by a political consulting firm. Your task is to use
data to predict whether a ballot proposition will pass. In order to pass, the proposition
needs to win more than 50% of the votes cast. A random sample of 1000 likely voters
surveyed one week before the election found that 515 were in favor the proposition. Based
on the statistical analysis given below, would you predict the ballot proposition would
pass? Why or why not?
1-PropZInt
(.48402, .54598)
pˆ = .515
n = 1000

We would ___________________________ the ballot proposed would pass since the entire

confidence interval is ____________________________________________________________________.

Key points:
• Our confidence is in the process that produces confidence intervals, not in any
particular interval. It is incorrect to say that a particular confidence interval has a
___% chance of including the true population parameter. Instead, we say that the
process that produces intervals capture the true population parameter with a ___%
probability.
• The true population parameter never changes. Either it is always within the
confidence interval or it is never within the confidence interval.

Example 6: What effect does changing the confidence level have on the width of a
confidence interval and the margin of error?

Confidence Confidence Interval Width of the Margin of Error


Level (Lower, Upper) Confidence Interval 𝑊𝑖𝑑𝑡ℎ 𝑈𝑝𝑝𝑒𝑟 − 𝐿𝑜𝑤𝑒𝑟
= =
= Upper – Lower 2 2

Idea: The bigger the confidence level, the bigger the confidence interval.

Why? Confidence level represents how confident you are that the process will contain the
parameter. So the larger the interval, the higher the chance the interval contains the
parameter, the more confident you feel.
EXTRA: Find the Sample Size given a Margin of Error (Proportions)

Fixing Margin of Error


o Quite often, researchers choose the size of the margin of error they wish to report
for a survey, and then collect a large enough sample to achieve it
o For example, most surveys and polls reported on the news use a 95% confidence
level and a margin of error of 3%
o In this section, we introduce a formula that allows us to determine the sample size
needed to estimate a proportion for a required margin of error

Sample Size for Proportions


o To find the sample size n, needed to obtain a margin of error of size m, we use the
following formula:
𝑧∗ 𝟐 𝟏
𝒏=( ) ( )
𝒎 𝟒
o Note that z* is our critical value multiplier from the margin of error formula
C-level 𝑧 ∗ : critical value
99% 2.58
95% 1.96 → close to 2
90% 1.645
80% 1.28
o Important: Always round up to the next whole number.

Example 1:
a) Find the sample size of voters you would need to predict the results of a local
election with a confidence level of 95% and margin of error of 3%.

To have ____________ confidence and a margin of error of ______________,

the minimum sample size to survey is ___________________________________________________.

b) Find the sample size of voters you would need to predict the results of a local
election with a confidence level of 99% and margin of error of 2%.

To have ____________ confidence and a margin of error of ______________,

the minimum sample size to survey is ___________________________________________________.

Observation: As the confidence level increases and margin of error decreases, then the
sample size increases.
7.5: Comparing Two Population Proportions with Confidence

Comparing Two Populations

o Many research questions require we compare two groups


o Men vs. women
o Treatment group vs. control group
o Old data vs. new data
o We now expand the confidence interval procedure for one proportion to allow
comparison of two proportions
o Note, however, we will now have two of everything:
o Two population proportions to compare: 𝒑𝟏 𝐚𝐧𝐝 𝒑𝟐
o Two sample sizes: 𝒏𝟏 𝐚𝐧𝐝 𝒏𝟐
o Two sample proportions: 𝒑̂ 𝟏 𝐚𝐧𝐝 𝒑̂𝟐
o Where the subscript 1 corresponds to population 1
o Where the subscript 2 corresponds to population 2

Estimating the Difference

o When asked about the difference between two numbers, we want to know how far
apart those two numbers are
o Much of our analysis in comparing two samples is based on subtraction
o Thus our comparison of two proportions will be based on the statistic:
𝒑̂
𝟏 − 𝒑 ̂𝟐
o This statistic will be used to estimate the difference between two population
proportions: 𝒑𝟏 − 𝒑𝟐

Significant Differences

o It might seem strange to go to so much trouble to see if two numbers are different
o After all, can’t we just look at two numbers and see if they are different?
o Even when two proportions are equal in the population, their sample
proportions can be different.
o This is due to the fact that we only look at a sample, not the entire population
o Thus, confidence intervals are one method for determining whether different
sample proportions reflect significant (or real) differences in the population

General Confidence Interval Procedure for Two Proportions

o First we find a confidence interval, at the significance level we think best, for the
difference in proportions 𝒑𝟏 − 𝒑𝟐 .
o Then we check to see whether that interval includes 0
o If it includes 0, then this suggests that the two population proportions are
equal.
o If it does not include 0, then this suggests that one of the population
proportions is significantly greater than the other
Conditions for a Valid Interval

1. Both samples are random samples


➢ Assume this is true even if not given

2. Samples are independent of one another


➢ Selection of one sample doesn’t affect the selection of the other

3. Both samples are large:


➢ Formulas to verify:
i. 𝒏𝟏 𝒑 ̂𝟏 ≥ 𝟏𝟎
ii. 𝒏𝟏 (𝟏 − 𝒑 ̂)
𝟏 ≥ 𝟏𝟎
iii. 𝒏𝟐 𝒑 ̂𝟐 ≥ 𝟏𝟎
iv. 𝒏𝟐 (𝟏 − 𝒑 ̂)
𝟐 ≥ 𝟏𝟎
➢ In other words, the number of successes and failures for each sample needs
to be at least 10

4. Both populations are large:


➢ Formulas to verify:
i. 𝑵𝟏 ≥ 𝟏𝟎𝒏𝟏
ii. 𝑵𝟐 ≥ 𝟏𝟎𝒏𝟐
➢ In other words, each population size needs to be at least 10 times the sample
size

Computing the CI on a Calculator


To compute a confidence interval for the difference in two proportions on the TI-83/84
calculator:
1. Go to the calculator: STAT → TESTS
2. Choose option B: 2-PropZInt
3. Enter x1, n1, x2, n2, then enter the C-Level
4. Highlight Calculate and press ENTER
5. Report the interval your calculator gives you

Interpreting the Interval

1. If the entire confidence interval for p1 − p2 is positive, then 𝒑𝟏 is significantly


larger than 𝒑𝟐 .
➢ That is, if 𝒑𝟏 − 𝒑𝟐 > 𝟎, then 𝒑𝟏 > 𝒑𝟐 .
➢ The bounds of the confidence interval tell us by how much

2. If the entire confidence interval for p1 − p2 is negative, then 𝒑𝟐 is significantly


larger than 𝒑𝟏 .
➢ That is, if 𝒑𝟏 − 𝒑𝟐 < 𝟎, then 𝒑𝟏 < 𝒑𝟐 .
➢ The bounds of the confidence interval tell us by how much

3. If the confidence interval contains 0 (i.e. the signs change from negative to positive)
then there is no significant difference between the two proportions
➢ That is, if 𝒑𝟏 − 𝒑𝟐 = 𝟎, then 𝒑𝟏 = 𝒑𝟐 .
Interpretation Templates
o If the entire confidence interval for p1 − p2 is positive:
▪ We are _____% confident that the proportion of ___________________________________
(C-level) (Characteristic)
is between _____% and _____% larger for _____________ than it is for ______________.
(Population 1) (Population 2)

o If the entire confidence interval for p1 − p2 is negative:


▪ We are _____% confident that the proportion of ___________________________________
(C-level) (Characteristic)
is between _____% and _____% larger for _____________ than it is for ______________.
(Population 2) (Population 1)

o If the confidence interval contains zero:


▪ We are _____% confident that there is no significant difference in the
(C-level)
proportion of ________________ and the proportion of ____________________.
(Population 1) (Population 2)

Note: Please write your notes for Examples 1- 3 on a separate piece of binder paper.

Example 1: According to the Pew Research Center, 47% of 3000 randomly selected
respondents to a poll in April 2012 reported that they strongly favored gay marriage. Back
in 2004, only 31% of 3000 randomly selected respondents said the same thing.

a) Can we conclude based on these two percentages alone, that a greater percentage of
people favored gay marriage in 2012 than in 2004?

__________, even though a greater proportion of the SAMPLE of respondents favor gay
marriage in 2012 (47%) compared to the SAMPLE in 2004 (31%), these

________________________________________________________________________________________________

________________________________________________________________________________________________;
that is, the POPULATION of ALL respondents might have a:
o Greater proportion favoring gay marriage in 2012
o Greater proportion favoring gay marriage in 2004
o The exact same proportion favoring gay marriage in both 2012 and 2004

b) Check the conditions for a valid confidence interval are met.

Label the two groups. When comparing two different time periods, generally, we make
Group 1 the more recent time period.

Group 1:

Group 2:
Check the four conditions for a valid interval

1. Are BOTH samples randomly selected?

2. Are the two groups independent of each other?

3. Check that both samples are large enough:

Success =

Group 1:

𝑛1 = 𝑝
̂1 =

i. # success from group 1 = x1 = 𝑛1 ̂


𝑝1

ii. # failure from group 1 = 𝑛1 − x1

Group 2:

𝑛2 = 𝑝
̂2 =

iii. # success from group 2 = x2 = 𝑛2 𝑝


̂2

iv. # failure from group 2 = 𝑛2 − x2

4. Check that both populations are large enough:

There are definitely over _____________________ ___________________________________


10n1 Group 1

There are definitely over _____________________ ___________________________________


10n2 Group 2

c) Find a 95% Confidence Interval for the difference in proportions supporting gay
marriage and interpret.

2-PropZInt(
x1, n1, x2, n2, C-level

Confidence Interval:
Sidework: Since the entire interval is ___________________, then p1 – p2 is _____________________.

Interpretation:

We are _____% confident that the proportion of _____________________________________


(C-level) (Characteristic)

________________________________________________________ is between ________% and __________%


(Characteristic)

larger in ______________ than it was in ______________.


Example 2: The Harris Poll conducted a survey in they asked, “How many tattoos do you
currently have on your body?” Of the 1205 randomly selected males, 181 responded that
they had at least one tattoo. Of the 1097 randomly selected females, 143 responded that
they had at least one tattoo.
a) What are the two populations of interest?

Population 1:

Population 2:

b) Researchers want to compare the proportion with at least one tattoo between these
two groups. Check the conditions for a valid confidence interval.

1. Are BOTH samples randomly selected?

2. Are the two groups independent of each other?

3. Check that both samples are large enough:

Success = having at least one tattoo

Population 1: i. # success from group 1 =

ii. # failure from group 1 =

Population 2: iii. # success from group 2 =

iv. # failure from group 2 =

4. Check that both populations are large enough:

Population 1: There are definitely over 12,050 males.

Population 2: There are definitely over 10,970 females.

c) Use your calculator to estimate pM – pF with a 95% confidence interval.

2-PropZInt(
x1, n1, x2, n2, C-level

Confidence Interval:
d) Is this interval positive, negative, or does it include zero? What does this mean?

Sidework: Since the entire interval is _______________, then p 1 – p2 is ___________________.

e) Interpret your interval with a complete sentence.

Interpretation:

We are _____% confident that there is ______________________________________________________


(C-level)

in the proportion of ____________ and _____________ having __________________________________


(Characteristic)
Example 3: The San Mateo County Clerk wishes to improve voter registration. One
method under consideration is to send reminders in the mail to all citizens in the county
who are eligible to register. As part of a pilot study 1250 random potential voters were
selected and divided into two groups, Group 1: 625 voters received no reminder and 295
registered, Group 2: 625 voters were sent a reminder and 350 registered to vote. First
check the Central Limit Theorem, and then estimate the 95% confidence Interval and
interpret.

Population 1:

Population 2:

Assume the conditions for a valid confidence interval hold.

Find the 95% confidence interval.

2-PropZInt(
x1, n1, x2, n2, C-level

Confidence Interval:

Sidework: Since the entire interval is _______________, then p 1 – p2 is ___________________.

Interpretation:

We are _____% confident that the proportion of _____________________________________


(C-level) (Characteristic)

________________________________________________________ is between ________% and __________%


(Characteristic)

larger __________________________________________________________________________________________

than if __________________________________________________________________________________________.

You might also like