0% found this document useful (0 votes)
10 views2 pages

Coffee Experiment and Sample Size Issues

The document discusses the Central Limit Theorem for proportions, explaining that the sampling distribution of a proportion will follow a normal distribution with mean p and standard deviation p(1-p)/n. It includes an example calculating the probability that the average time spent texting by a sample of students exceeds 84 minutes. Additionally, it highlights the importance of sample size through an anecdote about King Gustav III's coffee experiment, which suffered from a very small sample size leading to unreliable conclusions.

Uploaded by

Dev Chan
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
10 views2 pages

Coffee Experiment and Sample Size Issues

The document discusses the Central Limit Theorem for proportions, explaining that the sampling distribution of a proportion will follow a normal distribution with mean p and standard deviation p(1-p)/n. It includes an example calculating the probability that the average time spent texting by a sample of students exceeds 84 minutes. Additionally, it highlights the importance of sample size through an anecdote about King Gustav III's coffee experiment, which suffered from a very small sample size leading to unreliable conclusions.

Uploaded by

Dev Chan
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

​ ​ ​ ​ ​ ​ ​ ​ ​ Foundations of Data Science

​ ​ ​ ​ ​ ​ ​ ​ ​ Prof. Dinesh Kumar U


​ Module: 3

Central Limit Theorem for Proportions

The central limit theorem for proportion is stated as follows:

If we have a population in which a characteristic has a proportion of p, then the


sampling distribution of the proportion (that is 𝑝 calculated from several samples
of size n) will follow a normal distribution with mean p and standard deviation
𝑝(1 − 𝑝)/𝑛.

The Central Limit Theorem for proportions can be stated as follow:

If X1, X2, …, Xn are counts from a Bernoulli trial with probability of success p,
E(Xi) = p and Var(Xi) = p*(1-p), then the sampling distribution of the probability of
success (say 𝑝) follows an approximate normal distribution with mean p and
𝑝−𝑝
standard error 𝑝(1 − 𝑝)/𝑛, where n is the sample size. The variable
𝑝(1−𝑝)/𝑛
converges to a standard normal distribution.
X-bar - mu
Example ----------------

It is believed that college students in Bangalore spend, on average, 80 minutes a


day texting on their mobile phones, with a corresponding standard deviation of 25
minutes. Data from a sample of 100 students has been collected to calculate the
amount of time spent texting. Calculate the probability that the average time spent
by this sample of students will exceed 84 minutes.

Solution: Using the central limit theorem, the mean of the sampling distribution is
from Population
80 and the corresponding standard deviation is 25/ 100 = 2. 5. Mean and Std. Deviation,
calculate
Mean and Std Deviation for
The probability that the sample average is greater than 84 minutes is: Sampling Distribution
RV of Sampling Distribution
​ ​ 𝑃 𝑍>( 84−80
2.5 ) = 𝑃(𝑍 > 1. 6) = 0. 05479

Importance of Sample Size – King Gustav III’s Coffee Experiment (Afshari,


2017) King Gustav III of Sweden (1746 – 1792) was convinced that coffee posed

© AllRightsReserved. This document has been authored by Professor Dinesh Kumar U and is permitted for use only within the course
"Fundamentals of Data Sciences" delivered in the online course format by IIM Bangalore. No part ofthis document, including any logo,
data, illustrations, pictures, scripts, may be reproduced, or stored in a retrieval system or transmitted in any form or by any means –
electronic, mechanical, photocopying, recording or otherwise – without the prior permission of the author.
​ ​ ​ ​ ​ ​ ​ ​ ​ Foundations of Data Science
​ ​ ​ ​ ​ ​ ​ ​ ​ Prof. Dinesh Kumar U
​ Module: 3

several health hazards and wanted to educate Swedes about the same. Coffee
drinking was becoming popular in Sweden during 18th century, despite a ban.
King Gustav III wanted to demonstrate the health hazards of coffee using an
experiment. Luckily for him, a pair of identical twins convicted for a murder
were scheduled for an execution at the time. Gustav wanted to use them for his
“coffee experiment”. Gustav III commutated their death sentence to life
imprisonment under the condition that one of them must drink 3 pots of coffee
every day, and the other 3 pots of tea. Gustav III was convinced that they would
fall sick and die quickly. He also appointed 2 doctors to monitor the health of the
twins. In 1972, King Gustav III was assassinated and did not live to see the result
of the experiment that he started. Both doctors he appointed to monitor the twins
died few years later, and the twin who was drinking tea died at the age of 83. No
one recorded the age at which the coffee drinker died. Well, coffee was the
winner .

One of the major problems with this experiment was the sample size; it is very
difficult to arrive at a statistical inference based on just 2 observations.

© AllRightsReserved. This document has been authored by Professor Dinesh Kumar U and is permitted for use only within the course
"Fundamentals of Data Sciences" delivered in the online course format by IIM Bangalore. No part ofthis document, including any logo,
data, illustrations, pictures, scripts, may be reproduced, or stored in a retrieval system or transmitted in any form or by any means –
electronic, mechanical, photocopying, recording or otherwise – without the prior permission of the author.

Common questions

Powered by AI

In the context of the central limit theorem, the standard deviation of the sampling distribution, often called the standard error, indicates the dispersion of sample means around the population mean. It is calculated as the population standard deviation divided by the square root of the sample size. A smaller standard deviation in the sampling distribution implies that the sample means are tightly clustered around the population mean, enabling more precise estimates of the population parameter .

King Gustav III's experiment illustrates several challenges typical in longitudinal health studies, such as maintaining participant involvement over long periods and accounting for the natural increase in variables over time. Additionally, unforeseen events, like the death of subjects, can disrupt data collection. This unpredictability and reliance on long-term outcomes make it hard to draw causal inferences, especially without large sample sizes or control over extraneous variables .

The potential biases and limitations of King Gustav III's coffee experiment include its extremely limited sample size of just two subjects, which makes it difficult to generalize findings. Furthermore, without a control group or randomization, confounding variables could affect the results. The lack of detailed records of health outcomes and ages at death further complicates any form of reliable analysis or conclusions about coffee's health impacts .

Considering the extremely small sample size and the singular nature of the cases tested, King Gustav III's method was inadequate for drawing scientific conclusions about the health effects of coffee. The lack of repeatability and the susceptibility to variances not controlled for in the experiment's design reduce the validity of any findings derived from it. Additionally, without statistical tools like randomization or control groups, confounds such as genetic predispositions of the twins, environmental influences, or other lifestyle factors remain unaddressed .

Modern scientists can learn several statistical lessons from King Gustav III's experiment, notably the impact of sample size on the reliability of results, the importance of control groups and randomization in reducing bias, and the need for a well-structured methodology to permit valid inferences. The experiment also highlights the significance of clear and comprehensive documentation to provide context and details, enabling reproducibility and further analysis by others .

Standard error affects probability calculations by determining how much sampling variability to expect. In the context of the central limit theorem, it is the standard deviation of the sampling distribution. A smaller standard error indicates the sample mean is likely to be close to the population mean, affecting the calculated probabilities of observed outcomes. For example, in the case where we calculate the probability of a sample mean exceeding a threshold, the standard error scales the Z-score, affecting the resulting probability .

Ethical considerations in experiments similar to King Gustav III's include informed consent, the right to withdraw, risk of harm, and the use of human subjects without their full autonomy. In the 18th-century context, the twins were convicts coerced into participation, lacking real choice or awareness of potential risks. Today, such an experiment would be deemed unethical because it fails to adhere to principles of autonomy, beneficence, and justice, essential to human subject research .

The central limit theorem for proportions states that if we have a population with a characteristic proportion p, then the sampling distribution of the proportion calculated from several samples will approximate a normal distribution with mean p and standard deviation sqrt[p(1-p)/n]. This holds true when samples are large enough, even if the original data is not normally distributed .

The importance of sample size is highlighted in King Gustav III's coffee experiment, which involved only two participants (twins). Such a small sample size makes it difficult to draw reliable statistical inferences as it does not adequately represent the population. This undermines the experiment’s ability to conclusively demonstrate the health hazards of coffee .

In the students' texting example, the central limit theorem helps calculate the probability of observing an empirical outcome by approximating the distribution of sample means to a normal distribution. Given a sample mean, the population mean, and the standard deviation, we calculate the Z-score to find the probability that the average time spent texting exceeds a specific value. This method creatively uses probability to predict empirical outcomes in real-world contexts by assessing how typical or atypical the observed sample mean is within the expected distribution .

You might also like