Research Methods
Procedures for Hypothesis Testing
Dr. Sven Magg, Tayfun Alpay, Dr. Annika Peters
[Link]
Plan for today!
• Spurious effects in experiments
Hypothesis Testing
• What is a good hypothesis?
• Two procedures to test a hypothesis
• What is a one- or two-tailed test?
• How to define a good test?
• The dilemma of interpreting the results!
Dr. Sven Magg Research Methods - Hypothesis Testing 2
Effects to watch out for!
Always first: Run pilot study (small groups, dry run) to
• test your design, not the hypothesis (“debug” the protocol)
• calibrate measurements and parameters (e.g. number of subjects, number of trials,
etc.)
• test your measurements (reliability & validity) and analysis
Sometimes study is well designed, but still goes wrong
Common things to look out for:
• Boundary Effects (Ceiling and floor effects)
• Regression effects
• Order effects
• Sampling Biases
Dr. Sven Magg Research Methods - Hypothesis Testing 3
Ceiling & Floor effects
Group 1 2 3 4 5 6 7 Ø
Test 9 10 9 8 10 10 9 9,29
Control 9 9 8 9 10 9 9 9,0
Scores on 7 tests on a scale 0-10
Scores between test and control
are almost equal
Both are close to maximum value!
Maybe the tests were too simple?
Watch for boundary effects when results are near a
possible maximum/minimum value that can be reached
Dr. Sven Magg Research Methods - Hypothesis Testing 4
Ceiling & Floor effects
How to detect boundary effects?
1. Run a pilot study
2. Estimate best/worst bounds for recorded values
3. If both test & control are near this boundary
Boundary effect may be lurking
Just because your are far away from the absolute boundary of your scale
does not mean there is no boundary effect!
Practical boundary not necessarily the max/min of the recording scale!!
Dr. Sven Magg Research Methods - Hypothesis Testing 5
Other boundary effects
Cut-Off points often create
statistical “artifacts”
Be careful with random numbers
and limits (think about deviation!)
Better to “reflect” values at limits
Dr. Sven Magg Research Methods - Hypothesis Testing 6
Regression Effects
Test: 1 2 3 4 5 6 7 Ø
Alg 1 0 2 4 5 8 9 10 5.4
Alg 2 3 7 5 6 5.25
Algorithm 1 is tested on 7 problems
Due to time issues, only test improved algorithm on
problems where Alg1 performed below average
Claim: Algorithm 2 is an improvement!
Regression towards the mean!
We expect values to be better if the result depends on a
chance component!
For a second test, always choose a representative sample
Dr. Sven Magg Research Methods - Hypothesis Testing 7
Order Effects
Often there are sequences in the experiment procedure
• Robot has to complete a series of tasks
• Humans are presented a sequence of stimuli
What kind of effects can the order of the sequence have?
• The order has effects on the performance
• Two groups receive same sequence, but one is more sensitive to a specific order
than the other
Often subtle!
• garbage collection in Java at different points in time
• Learning curve different depending on sequence in training
Dr. Sven Magg Research Methods - Hypothesis Testing 8
Order Effects
How to detect order effects?
Counterbalancing
• Run problem on all permutations of a sequence
• Expensive! Sequence: Alg1: Alg2:
• Maybe just run a few to test a,b 10 15
b,a 12 20
Often impossible to run all permutations
Sometimes the sequence is already part of the study
If you are not sure: Include counterbalancing in pilot study!
Dr. Sven Magg Research Methods - Hypothesis Testing 9
Sampling Bias
When the collected sample is disproportionally biased towards a result
• Statistics on computer usage per day taken from a questionnaire distributed at the
Informatikum
• 2 samples taken at different days in front of Audimax
• Internet poll on political opinion on the Fox News Website
• Internet polls in general
Can be very subtle
• Assumtion: gender and #siblings are independent
• But, parents often have children until they have a boy
• i.e. gender and #siblings are NOT independent
Dr. Sven Magg Research Methods - Hypothesis Testing 10
Sampling Bias
How to detect sampling bias?
If we have a well-designed experiment with a random sample: mean of y is
different for different conditions of x
BUT: shape of distribution of y the same
If the shape changes for different levels of x, this hints at another factor
influencing membership in the sample!
Dr. Sven Magg Research Methods - Hypothesis Testing 11
What have we learned?
3. Run a pilot study
• test the protocol, procedures AND analysis (“debugging”)
• check for validity and reliability of measures
• look for spurious effects and sampling bias
4. Discuss and interpret results
• Do they show what you have expected?
• Do they actually answer your question?
• Did you address all competing hypotheses?
5. Run and repeat the experiment
• Repetition to disclose effects of interaction between independent and noise
variables
Dr. Sven Magg Research Methods - Experiment Design 12
Statistical Inference
If we have a sample drawn from a population, we can ask two kind of
questions:
1. How “good” is an estimate for a parameter of the population, drawn from this
sample?
How confident are we that the estimate is close to the real parameter value?
Example: Guessing the average number of blonde students from a snapshot
count in the Mensa
2. When answering a yes/no question about the population using the sample, how
likely is it that we are wrong?
Example: My software A is more accurate in guessing the weather than
software B
Dr. Sven Magg Research Methods - Hypothesis Testing 13
Hypotheses
1. a suggested explanation for a group of facts or phenomena, either
accepted as a basis for further verification (working hypothesis) or
accepted as likely to be true […]
3. (Philosophy / Logic) an unproved theory; a conjecture
[Collins English Dictionary]
has to be testable and falsifiable
follows from observation, exploratory study, or just idea
Big question: Is my hypothesis correct?
• What does correct mean? How well can I prove it?
• When do I consider it verified?
Dr. Sven Magg Research Methods - Hypothesis Testing 14
Hypotheses
Falsifiability and testability
• “All students are female”
Falsifiable by a single male student
• “When green aliens land in Hamburg, they always step of their spaceship with their
middle foot first”
Falsifiable in principle, but not in practice
• “A god-like being exists”
“Albert Einstein was the best physicist in the world!”
“What is the sun?”
Usually not scientifically falsifiable/testable by experiment
It has to be possible to think of a hypothesis stating the opposite and you can
both test them in practice
Dr. Sven Magg Research Methods - Hypothesis Testing 15
Let’s gamble first…
You watch a gambler throwing a die three times and always scoring a 6
You want to accuse him of using a manipulated die!
What are the chances of you being right?
Your assumption is that the die is fair and under this assumption you think it’s
unlikely to score three 6s
Two competing Hypothesis:
• The die is fair (𝐻0 ) (and the result is down to chance)
• The die is manipulated (𝐻1 )
Dr. Sven Magg Research Methods - Hypothesis Testing 16
Gambler’s Thinking Process
We do not know how the die was manipulated…..
If the die is fair, we know what the probabilities are:
• 1/6 to get a specific number in one go
• 1/(6 ∗ 6 ∗ 6) = 1/216 = 0.0046 to get three 6s
You are therefore 99.54% certain that the die was not fair?
You reject 𝐻0 with a chance of 𝑝 = 0.0046 to be wrong
What does this say about the chance of the die being manipulated?
Dr. Sven Magg Research Methods - Hypothesis Testing 17
What did we do?
We have…
1. …stated a null hypothesis 𝐻0 (Die is fair)
2. …thought about a formula to calculate chances, if 𝐻0 is true (binomial
distribution)
3. …observed a result (gathered a sample (6,6,6))
4. …calculated the probability 𝑝 of event to happen if 𝐻0 is true
5. …used 𝑝 as strength of evidence against 𝐻0
Science is easy!
Is 0.46% low enough to accuse the 2m professional boxer of cheating?
Dr. Sven Magg Research Methods - Hypothesis Testing 18
Rejecting is better
Why is my hypothesis the “alternative” hypothesis 𝐻1 ?
You can’t prove a hypothesis with statistics on a sample
But we can estimate the likelihood that a sample was drawn from a given
population!
𝐻0 : Sample from this (known) population
𝐻1 : Sample from a different population
I can statistically evaluate the likelihood that my sample 𝑋 came from a given
population and, if low, reject 𝐻0
Rejecting 𝐻0 Evidence for 𝐻1
Dr. Sven Magg Research Methods - Hypothesis Testing 19
Another example
Hypothesis: There are more male than female students in computer science!
Step 1: State a Null-Hypothesis
Dr. Sven Magg Research Methods - Hypothesis Testing 20
Group Task! 2 5
What is the difference between these
two hypotheses:
1. There are more male than female
students in computer science
2. The ratio of male and female
students is not equal in CS
When do you reject the hypothesis?
Dr. Sven Magg Research Methods - Hypothesis Testing 21
Another example
Hypothesis: The ratio of male and female students is not equal in CS
Step 1: State a Null-Hypothesis
• 𝐻0 : 𝑃 𝑚𝑎𝑙𝑒 = 0.5
Step 2: Find a sampling distribution for 𝑯𝟎
• From 𝐻0 : 𝑃(𝑚𝑎𝑙𝑒) = 𝑃(𝑓𝑒𝑚𝑎𝑙𝑒) = 0.5
• Binomial Distribution
Step 3: Gather a sample statistic
• 30 students in the Mensa:
x = 20 (male students)
Step 4: Calculate P(x=20): 0.028
Step 5: Decide?
Dr. Sven Magg Research Methods - Hypothesis Testing 22
p-values for regions
Can we reject 𝐻0 because having 20 males in a sample of 30 is unlikely?
How about 21? Or 28?
We would reject if we
see 20 or more!
If 𝐻0 would be rejected for several results, the probability of the combined
result is the sum of individual values:
𝑃𝑜𝑛𝑒𝑇𝑎𝑖𝑙𝑒𝑑 = 𝑃(20) + ⋯ + 𝑃(30) = 0.049!
Dr. Sven Magg Research Methods - Hypothesis Testing 23
Rejection Regions
1-Tailed Test
• for all values greater (smaller)
than a given sample statistic
• Used for directional hypotheses
(e.g. “greater than”)
2-Tailed Test
• Reject 𝐻0 if observed value is greater or lower than one of
two “cut-off” points
• 𝑃𝑡𝑤𝑜𝑇𝑎𝑖𝑙𝑒𝑑 = 0.099
• Typical use:
Reject 𝐻0 when the observed
value differs more than a given
maximum from the mean
Dr. Sven Magg Research Methods - Hypothesis Testing 24
Group Task! 2 5
Thoughts on p-values:
1. What can we use them for?
2. What do they mean for my
Hypothesis 𝐻1 ?
3. What would be good values for
rejection of 𝐻0 ?
Dr. Sven Magg Research Methods - Hypothesis Testing 25
History excursion
“Early” Ronald Fisher [1925]
• Inductive inference: Use direct probability P(Data| 𝐻0 )
• Only use a Null-Hypothesis 𝐻0 ≙ “happened by chance”, “no effect”
• Use known distribution of a test statistic T, assuming 𝐻0
• Set significance level (.05/.01/.001) following a convention
• Calculate p-value to check whether there is a significant (= backed by statistics)
divergence
• Significance value is a genuine feature of the test
“Late” Ronald Fisher [1956]
• Calculate exact p-value from the data
• Significance level is a feature of the data themselves
• No use of an arbitrary convention
Dr. Sven Magg Research Methods - Hypothesis Testing 26
History excursion
Ronald Fisher combined
• Use known distribution of a test statistic T, assuming 𝐻0
• Determine density of values that exceed observed value
• Use p value as strength p-value Strength of evidence
of evidence against H0 0.100 Borderline (or weak)
• p-Value is sample-based 0.050 Moderate
measure of evidence 0.025 Substantial
against null hypothesis 0.010 Strong
0.005 Very Strong
• We report exact p-value,
0.001 Overwhelming
NOT a decision
“P-Values and "significance" measure the probability
of data given the hypothesis, not the probability
of the hypothesis given the data.” - Morris DeGroot
Dr. Sven Magg Research Methods - Hypothesis Testing 27
Neyman-Pearson
Neyman-Pearson: [1928]
Define 𝐻0 and alternative Hypothesis 𝐻1
There are two errors you can make:
• Type I: False rejection (probability )
• Type II: False acceptance (probability )
𝑯𝟏 is true 𝑯𝟎 is true
Correct Outcome Wrong Outcome
Power (1-) Type I (-)Error
Reject 𝑯𝟎
True Positive (TP) False Positive (FP)
Significance Level
Wrong Outcome Correct Outcome
Accept 𝑯𝟎 Type II (-)Error Specificity (1-)
False Negative (FN) True Negative (TN)
Dr. Sven Magg Research Methods - Hypothesis Testing 28
Group Task! 2 5
Colour the regions for , ,power
and specificity
Dr. Sven Magg Research Methods - Hypothesis Testing 29
Group Task! 2 5
Colour the regions for , ,power
and specificity
Dr. Sven Magg Research Methods - Hypothesis Testing 30
and
Error probabilities:
𝑃 𝑇𝑦𝑝𝑒 𝐼 𝐸𝑟𝑟𝑜𝑟 =
𝑃 𝑅𝑒𝑗𝑒𝑐𝑡 𝐻0 𝐻0 𝑡𝑟𝑢𝑒 =
𝑃 𝑇𝑦𝑝𝑒 𝐼𝐼 𝐸𝑟𝑟𝑜𝑟 =
𝑃 𝐴𝑐𝑐𝑒𝑝𝑡 𝐻0 𝐻1 𝑡𝑟𝑢𝑒 =
The power of the test (= 𝑃 𝐴𝑐𝑐𝑒𝑝𝑡 𝐻1 𝐻1 𝑡𝑟𝑢𝑒 ) is dependent on:
• , which is set a priori
• the degree of separation between the two distributions given by δ = 𝜇1 − 𝜇0 (=
Effect size)
• the variance(s) of the population(s)
• the sample size N
Dr. Sven Magg Research Methods - Hypothesis Testing 31
Power of a test
Analogy
• You are searching for an item in your room
• Power: “What are your chances that you would find the item”
• Depends on:
How long you are searching Sample Size
The size of the item Size of effect, i.e. degree of separation
The messiness of the room Standard deviation
• There is a high chance to find a large item in a clean room if you spend a long time
searching!
• If you can’t find it, you can be confident it wasn’t there
• Power of an experiment: If there really is an effect, how high are the chances that
the experiment would find it?
Dr. Sven Magg Research Methods - Hypothesis Testing 32
Power of a test
The power of the test (= 𝑃 𝐴𝑐𝑐𝑒𝑝𝑡 𝐻1 𝐻1 𝑡𝑟𝑢𝑒 ) is dependent on:
• , which is set a priori
• the degree of separation between the two distributions given by δ = 𝜇1 − 𝜇0
= Effect Size
• the variance(s) of the
population(s)
• the sample size N
2 of those can be changed…
2 can be estimated AFTER the
experiment
Dr. Sven Magg Research Methods - Confidence Intervals 33
How to get the power?
Power often shown as power curves for one of the four parameters , δ,
variances (σ1 and σ2 ), and N
Using Monte-Carlo:
1. Fix all parameters but one
2. Draw samples from both
populations
3. Run the test and see whether
you would reject 𝐻0
4. Repeat 2.& 3. k times and calculate #rejections/k
5. Repeat 2.- 4. for all needed values of the free parameter
Dr. Sven Magg Research Methods - Confidence Intervals 34
Sample Size
You can increase the power of the test by increasing N
But should you?
We usually do not increase confidence in our test by increasing N!
• If we want to estimate a parameter, sample size should be as large as you can
afford (CI: 𝜇 = 𝑥ҧ ± 𝑡𝜎ො𝑥ҧ )
• If we want to test a hypothesis, samples should be no larger than required to show
the effect
If we test a hypothesis and are confident to reject 𝐻0 with sample size N, we
don’t gain much by increasing N
On the contrary…..
Dr. Sven Magg Research Methods - Confidence Intervals 35
Sample Size
Can samples be too big?
By increasing N you can
• boost any real effect to statistical significance
• boost any meaningless effect to significance
Do not fish for (statistical) significance!
What we want to know: How much predictive power does our result have?
In other words: Does knowing which population the sample came from give us
the power to predict it’s value?
Dr. Sven Magg Research Methods - Confidence Intervals 36
Sample size
Sample ഥ
𝒙 s N
Example from Cohen A 147.95 11.10 1000
2-sample test: 𝑡 = 2.468 B 146.77 10.16 1000
with 1998 degrees of freedom A&B 147.36 2000
Statistical significant difference (𝑝 ≤ .05)
If I hand you one sample from A and let you guess whether it is above the
combined mean A&B, how well would you do?
517 values from A exceed the combined mean
464 from B do as well
You would guess correctly for 51.7% of samples!
If you don’t know the origin of the sample: 50%
Dr. Sven Magg Research Methods - Confidence Intervals 37
Sample Size
We can roughly estimate predictive power
If we know (or can estimate):
• the population variances σ𝐴 2 and σ𝐵 2 of populations A & B
• The variance σ𝑃 2 of the combined population
then predictive power means reduction in variance from knowing the population
Relative reduction by knowing sample is from A:
2
σ𝑃 2 −σ𝐴
𝜔2 =
σ𝑃 2
2
If 𝜔2 2
= 0, then σ𝑃 = σ𝐴 and no prediction is possible
𝜔2 = 1 means no variance in A, i.e. perfect prediction
Dr. Sven Magg Research Methods - Confidence Intervals 38
Sample Size
We can only estimate 𝜔2 since we try to find the population parameters!
If both variances are equal for both populations, a rough estimate is:
2
2
𝑡 −1
𝜔ෝ = 2
𝑡 + 𝑁1 + 𝑁2 − 1
In the example: 𝜔 ෝ 2 =0.0025
If we would have drawn only 100 samples each: 𝜔 ෝ 2 =0.027
Increasing N decreases the predictive power and therefore the meaningfulness
of significant findings
Dr. Sven Magg Research Methods - Confidence Intervals 39
Neyman-Pearson
Neyman-Pearson: [1928]
Define 𝐻0 and alternative Hypothesis 𝐻1
There are two errors you can make:
• Type I: False rejection (probability )
• Type II: False acceptance (probability )
𝑯𝟏 is true 𝑯𝟎 is true
Correct Outcome Wrong Outcome
Power (1-) Type I (-)Error
Reject 𝑯𝟎
True Positive (TP) False Positive (FP)
Significance Level
Wrong Outcome Correct Outcome
Accept 𝑯𝟎 Type II (-)Error Specificity (1-)
False Negative (FN) True Negative (TN)
Dr. Sven Magg Research Methods - Hypothesis Testing 40
Neyman-Pearson- procedure
Step 1+2 as before (including definition of 𝐻1 )
Step 3: Set 𝜶
• Decide on a maximum acceptable probability 𝛼 of incorrectly rejecting 𝐻0
Step 4: Find cut-off points
• Use sampling distribution 𝑁ℎ to find critical values 𝑐 + and 𝑐 −
• Set 𝑐 + and 𝑐 − such that P 𝑁ℎ ≥ 𝑐 + + 𝑃 𝑁ℎ ≤ 𝑐 − ≤ 𝛼
Step 5: Gather the sample statistic (run experiment)
Step 6: Decide using X and cut-off points
• If 𝑥 ≥ 𝑐 + 𝑜𝑟 𝑥 ≤ 𝑐 − , reject 𝐻0
In the example: Set 𝛼 = 0.05 𝑐 + = 21 and 𝑐 − = 9
• 𝑃 𝑁ℎ ≤ 9 + 𝑃 𝑁ℎ ≥ 21 = 0.043 ≤ 𝛼
Dr. Sven Magg Research Methods - Hypothesis Testing 41
Neyman-Pearson
First define and before running the test
and are probabilities of making errors of type I / II in the long run and
therefore features of the test
Set and not by a convention but after a detailed cost-benefit analysis of the
consequences
You have to think prior to the experiment about meaningful values for and
• What are the consequences of making a type I or II error?
• Values often used: =.05 and =.2 (= power of 80%)
• Sometimes you want a power of almost 100% and can accept a high probability of
errors of type I
Dr. Sven Magg Research Methods - Hypothesis Testing 42
Neyman-Pearson
Rules of inductive behaviour
Rule gives you decision (reject/accept 𝐻0 ) without final statement whether we
believe 𝐻0 is true/false
An optimal statistical test minimises while keeping at a set bound.
Nowadays often mixed forms of Fisher/Neyman-Pearson can be found that
would satisfy neither:
• Report significance level (usually 1-3 stars or “ns”) and exact p-value
• Define and use it to reject hypotheses after calculating p
• Use a conventional and of 0.05 and .2
• etc….
Dr. Sven Magg Research Methods - Hypothesis Testing 43
What should you do?
Decide for one side of the debate!
Either:
1. Think about and set and before the test and report findings as significant or
not, stating the significance level
“There was a significant effect (=.05)” or “.. (𝑝 ≤ .01)”, etc.
2. Calculate p-Value for the sample after the test and report exact p-Value without
reporting a decision about 𝐻0
In the first case, think about the consequences of your errors.
If you can’t think of any: Use 2.
Dr. Sven Magg Research Methods - Hypothesis Testing 44
What have we learned?
1. Good hypotheses have to be falsifiable in practice
2. We define 𝐻0 (and maybe 𝐻1 ) and try to reject 𝐻0
3. When we reject, there is a chance that we are wrong
4. P-Values are a measure of the probability to falsely reject 𝐻0
5. Fisher’s testing strategy using sample-based p-values
6. Neyman-Pearson’s strategy is to set and
7. We can either use a calculated P-Value directly as the strength of evidence
against 𝐻0 or set bounds and reject 𝐻0 if the sample statistic is within the
rejection regions
8. We have to find sampling distributions to make decisions!
Dr. Sven Magg Research Methods - Hypothesis Testing 45