0% found this document useful (0 votes)
2 views7 pages

Null Hypothesis in Tea Experiment Study

This document discusses an experimental study involving a 'tea lady' who can distinguish between cups of tea based on the order of milk addition. It emphasizes the importance of the null hypothesis, randomization, sample size, effect size, and statistical power in experimental design, referencing Fisher's original work. Additionally, it highlights the need for careful planning and adherence to experimental designs to avoid bias and ensure credible results.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views7 pages

Null Hypothesis in Tea Experiment Study

This document discusses an experimental study involving a 'tea lady' who can distinguish between cups of tea based on the order of milk addition. It emphasizes the importance of the null hypothesis, randomization, sample size, effect size, and statistical power in experimental design, referencing Fisher's original work. Additionally, it highlights the need for careful planning and adherence to experimental designs to avoid bias and ensure credible results.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Part C: Tea lady (experimental study)

In this section, students will discuss the questions below within their groups. Each group will be
assigned a question by the TA and asked to present their answer to the whole class.

1.​ Why is the null hypothesis a valuable tool in experimental design? What is the null
hypothesis that will be tested in this experiment? Referring to your answers to the
questions at home, was the hypothesis you chose different than the null hypothesis
given by Fisher in his essay? Explain your answer.

N0 - The order in which the tea components are added does not have an effect on the ladies
ability to discriminate between the two types of cups.

It gives a baseline framework that can be either disproven or proven. Most of science is built
upon data and quantitative values, so being able to “disprove” a null hypothesis to accept the
alternative hypothesis gives a study more credit. It also acts as a reference for a study, no
significance would mean no difference, and is a reference point for data that would have a
difference. The hypothesis chosen was different from the null hypothesis given by Fisher, in the
sense that Fisher’s null hypothesis talked about frequency distributions. In contrast, the null
hypothesis we created was more specific, stating that there would be no significance in how the
tea lady was able to discern the tea leaves.

2.​ What reason did you give for why it is important to offer the model more than just two
cups (one with the milk added first and one with the milk added second)? Was your
answer the same as Fisher’s answer? Based on the essay, please describe Fisher’s
answer to this question.

It is important to offer the model more than just two cups (one with the milk added first and one
with the milk added second) because one cup can be guessed correctly by chance, but more
cups allow statistical assessment. Yes, our answer is similar to the Fisher's answer, since they
recommended 8 cups (4 of each type) to allow meaningful probability. To explain, there are 70
ways of choosing groups of 4 objects out of 8 (8x7x6x5 = 1680 / 4x3x2x1 = 24). If the cups are
indistinguishable, there is a 1 in 70 chance she can guess it all correctly, which is the probability
that represents the number of expected frequency of perfect success (based on the null
hypothesis - no ability to discriminate). Therefore, increasing the number of cups would make it
even harder to guess, strengthening the test, whereas fewer amounts of cups could cause a
perfect guess by chance, making the study statistically weak. In summary, if you have two cups
its a 50/50 chance, probability to high, Need to have more than one cup because its less likely
to happen by chance.

3.​ How many cups did you say the lady should taste (before reading the article)? How
many cups did Fisher say that the “tea lady” in the story should taste? Please describe
fully Fisher’s answer to this question, including any mathematical considerations. Was
your answer the same as Fisher’s answer? If not, how is it different? When thinking back
to your test of significance, what factors might impact your faith in your conclusions?
What factors limit or constrain experimental design?

I said about 4 cups before reading the article. The Fisher said that the “tea lady” in the story
should taste 8 cups (4 each), and the combination is of 4 out of 8, which results in 70 possible
outcomes (8x7x6x5/4x3x2x1), where a perfect score is 1/70 by pure chance. Some factors that
might impact our faith in the conclusions, or factors that limit/constraint experimental design, are
sample size, chance variation, and uncontrolled differences.

4.​ Describe exactly how the cups should be prepared. Does every cup need to be exactly
the same in every way except the order of the addition of milk and tea? Can you actually
make every cup identical? Before reading Fisher’s essay, did you think that every cup
needed to be exactly the same in every way except the order of the addition of milk and
coffee? Does Fisher believe that every cup should be prepared identically? Describe
Fisher’s explanation for how to deal with uncontrollable variation among cups.

Each amount of tea and milk should be measured with the same measuring device for every
cup in order to account for as much accuracy as possible. It would be impossible to make each
cup identical because of human error. I did believe that every cup needed to be prepared
identically in order for this study to be accurate.

Fisher stated that it was almost impossible to create the perfect conditions in which every cup is
prepared the same. Hard to control variation - variation is inherent in all systems.

Distributes variation since randomization handles the remaining differences

5.​ Why is randomization important in experimental design? In what order should the cups
be presented? What method or decision rules might you use to decide which cup is to be
offered first, second, etc.? Before reading Fisher’s essay, what method did you
recommend for choosing the order that the cups should be presented? Was your answer
different than what Fisher recommends in his essay? Please describe Fisher’s
explanation for how to choose the cup order.

-​ It removes bias and improves the validity of experiments. It offers a more realistic
approach to science because the majority of the time, what they're testing for, occurs
randomly in time in the environment. As well, non-random experiments are easily
affected by bias because the research may perpetuate certain underlying intentions by
providing only certain items at certain times. First, 5 out of 10 of the cups would be filled
with milk first then 5 the other way. All cups would be covered and then randomly
organized using a random number generator.
-​ The method we recommended for choosing the order that the cups would be presented
would be to line them up in a line and put the number of cups (i.e. 10 samples / 10 cups)
into a random number generator. Then the cups would be selected randomly in any
order out of the sample size until all cups are finished. The lady would mark which cup
she considered to be milk infused or not. In the case of Fisher, he did a very similar
method but used less cups, I assume to reduce the tasting fatigue, and did a 4 by 4 split
of whether the milk was added first or later. Then he randomly selected cups from the 8
total cups and asked the lady to organize the cups into separate groups.

Trying the equalize the variation

Note: Random number generators are considered pseudo-random bc they are built on Al
algorithm so they are not truly random
Note: if you coin flip to decide what order each groups will be presented to the lady and they are
all the same, repeat the coin flip
-​ Introducing randomization again can also introduce bias (experimenter should NOT
correct “bad” case of randomization)

6.​ Would it be acceptable to add two more cups to the original 8 cups? Why or why
not? What is the value of deciding the experimental design before you begin an
experiment and not changing it in the middle of the experiment?

I think it would be acceptable to add 2 more cups as long as the total amount of cups in each of
the 2 groups remains consistent. The more cups the lady tests, the more likely that her results
are not due to random chance Adding 2 more groups, 1 in each group would mean there's 9 in
each group which is an odd number which Fisher wanted to avoid

Choose even number of cups - reduces bias and skewed results


(bc symmetrical dataset, less likely to see outliers bc 2 extreme opposing outliers can cancel
themselves out within a symmetric dataset)

It is often better to create and stick to an experimental design chosen before the experiment
begins. This helps to reduce bias that can occur when scientists or researchers may have an
inclination to change the experimental design as they realize they are not getting the results that
they want or significant results. Sticking to the original experimental design will often provide a
more reliable answer to whether there is a true cause and effect relationship. It will also make
the experiment more credible and reproducible → more accurate and reliable results.

7.​ Rumor has it that the real lady who had tea with Fisher on an afternoon in the 1920s was
able to accurately tell whether the milk was added first or second every single time she
was offered a new cup. Can you conclude from her success that most people can tell the
difference between milk-first and tea-first cups? Design an experiment that would test
this broader question.
I think it would be inaccurate to assume that most people can tell the difference between
milk-first and tea-first cups because each individual’s taste palette is so unique and different
from others. A way to test this would be to have a group of 20 individuals of different
backgrounds and experiences with food and tasting to go through this same experiment with 8
cups 4 with milk-first and 4 with tea-first.

Only tested on one individual - you would need to increase the sample size to generalize results

PART D: The importance of effect sizes and statistical power


Refer to the article by Sullivan and Feinn (2012). As a class, you will answer the following
questions in your own words:

1. What is the difference between effect size and statistical significance? Why is it
important to report both?
-​ Effect size is the magnitude or strength of the effect, and statistical significance only tells
you whether the observed effect is due to chance or is a real effect. It's important to
report both because statistically significant results may be meaningless and have a tiny
effect, while statistically insignificant results may have a large effect, but due to a small
sample size, it’s considered insignificant.
-​ Correlation coefficient - Effect size
-​ P test - statistical significance
-​ Effect size gives you the limit of too many trials (eg. too many trials)

2. What is statistical power and why is it important to calculate it? What does it depend
on?
-​ Statistical power is the probability that a study will reject the null hypothesis correctly. Its
important because it helps prevent a type II error which is when you think a relationship
doesn't exist when it actually does. Depends on sample size and effect size.

The probability that your test is going to find statistical significance between the interventions
-​ Large effect size →to see the effect you need a small sample size
-​ Small effect size → need a large sample size

3. How can you increase the power of your study?


-​ Larger sample size

4. A small sample size might cause a researcher to wrongly accept the Ho. Thus, a general
principle is that the larger the sample size, the greater the likelihood of obtaining significant
results when they indeed exist. This is because larger sample sizes give more accurate
estimates of the actual population than do small sample sizes. However, this does not mean
that researchers should always use huge sample sizes. Besides potential cost, what is another
potential issue of including sample sizes that are too large in a study? Hint: refer to the example
of the aspirin study in the article).?
-​ Regardless of whether or not there is a difference, having a large sample size (too large)
will wrongly create a statistically significant result.
-​ It will almost always depict a statistically significant result if the sample size is large
enough
-​ If you can not have a smaller sample size - Choose a low alpha (p value of 0.001)
-​ Having too large a sample size or too many trials can be deleterious because it can
make the results seem significant when it’s really just a large effect size?

5. Taking all into consideration, briefly discuss different factors that may contribute to obtain
non-significant results in a study
-​ Confounding variables can cause results to be misconstrued / unaccounted for and can
result in non-significant results
-​ There may be little difference between the groups being chosen
-​ Study was poorly designed

Tutorial Notes​
Tutorial 7 is a small snippet of what we will be doing in test 2

●​ Good experiment design: consider how to manipulate independent variable, how


subjects will be assigned, how to precisely measure dependent variables, and how to
control confounding variables
●​ Experiment should declare the minimum effect size → ensure the experiment can detect
significant results that are biologically relevant
Effect size: deductible difference of biological difference
-​ How big is the effect between two groups

●​ Odds ratio (OR, ratio of odds fo an outcome occurring in the exposed group compared
to odds of outcome occurring in the unexposed group) hazard ratio (ratio of hazards b.w
3 groups), relative risk of risk ratio (RR, ratio of the risk occurring in exposed vs
unexposed group), and correlation coefficient r (quantifies the degree of correlation
between independent and dependent variables)

●​ Null hypothesis - no effect


●​ Ha- some association or effect

Ear Length and Age in Men


●​ Type of study: observational, cross sectional (weaker compared to longitudinal), clinical
(human subjects even though no intervention is applied).
●​ Ho- no association/relationship b/w age and ear size
●​ Ha - there is an association b/w age and ear size
●​ Independent variable: age & dependent variable: ear size
●​ There’s no p value, but they came to conclusion by “linear regression equation and 95%
CI of 0.17 to 0.27)
○​ As long as CI does NOT contain the null hypothesis value (assumed to be 0 if it's
not given) → results are significant! (reject Ho, fail to reject Ha)
●​ Control group: none
●​ External validity: does not apply to women because they only studied men, there is no
correlation coefficient (value from 0-1, closer to 1 = stronger correlation)
●​ Better study design: longitudinal study (easier to detect causation), attrition bias (people
dropping out of study), experimental study (instead of observational)
○​ Obs has no intervention or control (exp has both)
○​
●​
Mathematics of a Lady Testing TEa
-​ A lady claims that she can discriminate whether the mild or the tea infusion was fist
added to the cup

Quiz answers

●​ Any deviation
1.​ Define scientific bias and where it can result from?
Any deviation of results or inferences from the truth, or processes leading to such deviation.
Bias can result from several sources: one sided or systemic variations in measurement from the
true value *systemic error), flaws in study design, deviation of inferences, interpretations, or
analysis based on flawed data or data collection etc. There is not sense of prejudice or
subjectivity in the assessment of bias under these conditions

2. Conflict of interest results in a retraction when


a)​ Any conflict of interest exits
b)​ Conflict of interest is not disclosed
c)​ Conflict of interest is disclosed
●​ Any deviation

3. Which of the following are a reason for retraction


a)​ Civil proceedings
b)​ concerns/issues about results
c)​ Copyright claims
d)​ All of the above
4. What is the role of the editor in the publication process?
The editor makes the final decision about whether a research paper will be published in a
journal. The editor uses the comments of the referee when making this decision.

5. What is meant by the term "peer review”?


a.​ You classmate has reviewed your article
b.​ The masters/PHD student in your lab has reviewed your article
c.​ Experts in the field review your article prior to publication
d.​ Your supervisor has reviewed your article

6. Please identify two areas (stages of conducting an experiment) where scientific bias can
occur
-​ Data collection, data analysis, interpretation, publication
*areas = stages of an experiment

Test:
-​ Null hypothesis: no association between … and …
-​

You might also like