0% found this document useful (0 votes)
15 views5 pages

6 2lesson

This document covers a lesson on sampling distribution for a proportion, focusing on the development of an 'E-Nose' for detecting Parkinson's disease through sebum analysis. It includes guided notes, discussion questions, and practice problems related to statistical concepts such as sample size, probability, and the effectiveness of the E-Nose compared to random guessing. The lesson is designed to engage students in understanding statistical inference through real-world applications.

Uploaded by

awesomeknight26
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
15 views5 pages

6 2lesson

This document covers a lesson on sampling distribution for a proportion, focusing on the development of an 'E-Nose' for detecting Parkinson's disease through sebum analysis. It includes guided notes, discussion questions, and practice problems related to statistical concepts such as sample size, probability, and the effectiveness of the E-Nose compared to random guessing. The lesson is designed to engage students in understanding statistical inference through real-world applications.

Uploaded by

awesomeknight26
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Name: _________________

AP Statistics Handout: Lesson 6.2 (Day 2)


Topics: sampling distribution for a proportion, spread and sample size
Note: This activity was inspired by Doug Tyson’s amazing lesson.

Lesson 6.2 (Day 2) Guided Notes1

Making an “E-Nose”
Joy’s ability to detect Parkinson’s with scent is truly

Image from study*


amazing. But, there’s a problem: She’s only one
person. For patients across the world to benefit from
early detection of Parkinson’s, we have to develop a
detection tool that can be deployed in many places at
once.

Researchers at Zhejiang University (in China) have attempted to do just that. Inspired by Joy’s story, they
created an “E-Nose” that detects compounds produced in the sebum (an oily substance emitted in the
skin). Their system attempts to detect the specific mix of compounds in the sebum of Parkinson’s
patients. They published a study* on their results, which we’ll analyze below.

*Study: Wei Fu et al. Artificial Intelligent Olfactory System for the Diagnosis of Parkinson’s
Disease. ACS Omega, 2022; 7 (5): 4001 DOI: 10.1021/acsomega.1c05060

Sampling Distribution for a Proportion

The researchers tested the E-Nose on 24 people (12 with Parkinson’s, 12 without Parkinson’s).

1. In a world where the E-Nose doesn’t work at all (i.e. it’s only as good as random guessing)...
a) How many of its guesses would you expect to be correct? Why?

b) What proportion of its guesses would you expect to be correct? Why?

2. In their first set of tests, the E-Nose correctly guessed (Parkinson’s or no Parkinson’s) 17 out of the 24
people (70.8% correct). In your opinion, is this convincing evidence that the E-Nose is more effective
than randomly guessing? Why or why not?

1
Notes are meant to accompany this Desmos activity:
[Link]

Material adapted from the Skew The Script curriculum ([Link])


2

3. Your class will now return to the Desmos activity2 to simulate the proportion of correct guesses (𝑝̂ )
from a random machine (over many trials). Create a copy of the distribution of simulations below,
labeling each simulated proportion as 𝑝̂ .

0.30 0.40 0.50 0.60 0.70

4. Recall: The E-Nose made correct guesses 17 out of 24 times (𝑝̂ = 0.708). In a world where the
machine was randomly guessing, would this result be surprising? Justify using the distribution you
sketched above.

5. Use the Desmos activity to add even more trials. Describe the shape of the resulting distribution. Does
it look like a distribution you’re already familiar with? If so, which one?

It turns out, under certain conditions, we can model the sampling distribution for a proportion using a
normal curve! Here is the formula:
𝑝(1−𝑝)
𝑝̂ ~ Normal(𝜇𝑝̂ = 𝑝, 𝜎𝑝̂ = √ )
𝑛

6. In this situation, we’re assuming the true proportion of correct guesses will be p = 0.5. We are
performing this study on a sample of 24 patients (n = 24). Use these numbers to calculate the
parameters of the normal distribution (using the formula above). Then, sketch this normal curve on top
of the distribution you made in #3. Does the curve fit?

7. Using the normal curve, calculate the probability that the E-Nose gets 70.8% of its guesses correct (or
more) by chance alone. Is this similar to the probability you estimated in Question #4? Are you
convinced that the E-Nose was more effective than random guessing? Why or why not?

2
Direct URL to Desmos activity: [Link]

Material adapted from the Skew The Script curriculum ([Link])


3

Lesson 6.2 (Day 2) Discussion


Discussion Question: Imagine that the researchers sampled n = 240 patients for their study (again, half
with Parkinson’s and half without). Imagine that among those, the E-Nose made correct guesses for 170
170
patients (𝑝̂ = 240 = 70.8%).

a) Compared to before, does this evidence leave you more convinced, less convinced, or equally
convinced that the E-Nose does better than random guessing? Why or why not?

b) With this higher sample size, calculate the new probability of getting 70.8% of guesses correct
by chance alone. Is this probability higher or lower than what you found in Question #7? Why
did the probability increase/decrease?

Lesson 6.2 (Day 2) Practice

1) In 2023, 60% of students who took the AP Statistics exam passed with a score of 3 or higher. A
random sample of 300 students who took that same exam were selected, and it was found that 67%
passed with a score of 3 or higher.
a) In this scenario, which value represents the parameter 𝑝 (the population proportion of exam
passers) and what value represents the statistic 𝑝̂ (the sample proportion of exam passers)?

Material adapted from the Skew The Script curriculum ([Link])


4

b) Find the mean and standard deviation of the sampling distribution for 𝑝̂ (assume all inference
conditions are met).

c) Interpret the mean and standard deviation of the sampling distribution of 𝑝̂ in context.

2) Your teacher gives students an opportunity to blindly draw a marble out of a well-mixed bag, in hopes
of picking a green marble. A green marble means you get a homework pass. The teacher claims that 20%
of the marbles in the bag are green. Across all of the teachers’ classes, there are 80 students, and each
student gets to draw 1 marble. Each student puts their marble back in the bag after each draw. You
recorded the proportion of draws that found a green marble and found that only 6 students got a green
marble. You start to wonder if the teacher is lying in her claim that 20% of the bag’s marbles are green.
a) What is the parameter p in this scenario?

b) Assuming the teacher is telling the truth, find mean and standard deviation of the sampling
distribution for 𝑝̂ (the sample proportion of 80 draws that got a green marble). Interpret these
values in context. Assume the relevant inference conditions are met.

c) Assuming the bag is composed of 20% green marbles, find the probability of students getting 6
or fewer green marbles out of 80 draws. Assume all conditions for inference are met.

d) Does the previous probability convince you that the teacher is lying? Why or why not?’

Material adapted from the Skew The Script curriculum ([Link])


5

3) A lake has been stocked with a variety of fish, 35% of which are small-mouth bass. A fisherman from
another town has come to fish at the lake and doesn’t know what proportion of the fish in the lake are
bass. The fisherman plans to catch fish at various randomly chosen locations throughout the lake. Would
the fisherman be more likely to precisely estimate the true proportion of bass in the lake by catching 10
fish or 50 fish? Explain your choice.

Further Practice

Teachers: provide exercises from your AP Stats textbook (or other resources) about the content covered
in this lesson. This lesson matches with the following sections of the most widely-used AP Stats
textbooks:

• The Practice of Statistics (AP Edition), 4th-6th editions: section 7.2


• Stats: Modeling the World (AP Edition), 4th/5th editions: ch 17, 3rd edition: ch 18
• Statistics: Learning from Data (AP Edition), 2nd edition: section 8.2
• Advanced High School Statistics, section 4.1

Material adapted from the Skew The Script curriculum ([Link])

You might also like