0% found this document useful (0 votes)
5 views45 pages

Hypothesis Testing Research Methods Guide

The document discusses research methods for hypothesis testing, emphasizing the importance of pilot studies to identify spurious effects, boundary effects, regression effects, order effects, and sampling biases. It outlines procedures for formulating and testing hypotheses, including the distinction between null and alternative hypotheses, and the significance of p-values in statistical inference. Additionally, it covers the historical context of hypothesis testing and the types of errors that can occur during the process.

Uploaded by

vincent.rolfs
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views45 pages

Hypothesis Testing Research Methods Guide

The document discusses research methods for hypothesis testing, emphasizing the importance of pilot studies to identify spurious effects, boundary effects, regression effects, order effects, and sampling biases. It outlines procedures for formulating and testing hypotheses, including the distinction between null and alternative hypotheses, and the significance of p-values in statistical inference. Additionally, it covers the historical context of hypothesis testing and the types of errors that can occur during the process.

Uploaded by

vincent.rolfs
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Research Methods

Procedures for Hypothesis Testing

Dr. Sven Magg, Tayfun Alpay, Dr. Annika Peters

[Link]
Plan for today!

• Spurious effects in experiments


Hypothesis Testing
• What is a good hypothesis?
• Two procedures to test a hypothesis
• What is a one- or two-tailed test?
• How to define a good test?
• The dilemma of interpreting the results!

Dr. Sven Magg Research Methods - Hypothesis Testing 2


Effects to watch out for!

 Always first: Run pilot study (small groups, dry run) to


• test your design, not the hypothesis (“debug” the protocol)
• calibrate measurements and parameters (e.g. number of subjects, number of trials,
etc.)
• test your measurements (reliability & validity) and analysis

 Sometimes study is well designed, but still goes wrong


 Common things to look out for:
• Boundary Effects (Ceiling and floor effects)
• Regression effects
• Order effects
• Sampling Biases
Dr. Sven Magg Research Methods - Hypothesis Testing 3
Ceiling & Floor effects
Group 1 2 3 4 5 6 7 Ø
Test 9 10 9 8 10 10 9 9,29
Control 9 9 8 9 10 9 9 9,0
Scores on 7 tests on a scale 0-10

 Scores between test and control


are almost equal
 Both are close to maximum value!

 Maybe the tests were too simple?


 Watch for boundary effects when results are near a
possible maximum/minimum value that can be reached

Dr. Sven Magg Research Methods - Hypothesis Testing 4


Ceiling & Floor effects

 How to detect boundary effects?


1. Run a pilot study
2. Estimate best/worst bounds for recorded values
3. If both test & control are near this boundary
 Boundary effect may be lurking

 Just because your are far away from the absolute boundary of your scale
does not mean there is no boundary effect!
 Practical boundary not necessarily the max/min of the recording scale!!

Dr. Sven Magg Research Methods - Hypothesis Testing 5


Other boundary effects

 Cut-Off points often create


statistical “artifacts”
 Be careful with random numbers
and limits (think about deviation!)
 Better to “reflect” values at limits

Dr. Sven Magg Research Methods - Hypothesis Testing 6


Regression Effects
Test: 1 2 3 4 5 6 7 Ø
Alg 1 0 2 4 5 8 9 10 5.4
Alg 2 3 7 5 6 5.25

 Algorithm 1 is tested on 7 problems


 Due to time issues, only test improved algorithm on
problems where Alg1 performed below average
 Claim: Algorithm 2 is an improvement!

 Regression towards the mean!


 We expect values to be better if the result depends on a
chance component!
 For a second test, always choose a representative sample
Dr. Sven Magg Research Methods - Hypothesis Testing 7
Order Effects

 Often there are sequences in the experiment procedure


• Robot has to complete a series of tasks
• Humans are presented a sequence of stimuli
 What kind of effects can the order of the sequence have?
• The order has effects on the performance
• Two groups receive same sequence, but one is more sensitive to a specific order
than the other

 Often subtle!
• garbage collection in Java at different points in time
• Learning curve different depending on sequence in training

Dr. Sven Magg Research Methods - Hypothesis Testing 8


Order Effects

 How to detect order effects?

 Counterbalancing
• Run problem on all permutations of a sequence
• Expensive! Sequence: Alg1: Alg2:
• Maybe just run a few to test a,b 10 15
b,a 12 20

 Often impossible to run all permutations


 Sometimes the sequence is already part of the study
 If you are not sure: Include counterbalancing in pilot study!
Dr. Sven Magg Research Methods - Hypothesis Testing 9
Sampling Bias

 When the collected sample is disproportionally biased towards a result


• Statistics on computer usage per day taken from a questionnaire distributed at the
Informatikum
• 2 samples taken at different days in front of Audimax
• Internet poll on political opinion on the Fox News Website
• Internet polls in general
 Can be very subtle
• Assumtion: gender and #siblings are independent
• But, parents often have children until they have a boy
• i.e. gender and #siblings are NOT independent

Dr. Sven Magg Research Methods - Hypothesis Testing 10


Sampling Bias

 How to detect sampling bias?


 If we have a well-designed experiment with a random sample: mean of y is
different for different conditions of x
 BUT: shape of distribution of y the same

 If the shape changes for different levels of x, this hints at another factor
influencing membership in the sample!

Dr. Sven Magg Research Methods - Hypothesis Testing 11


What have we learned?

3. Run a pilot study


• test the protocol, procedures AND analysis (“debugging”)
• check for validity and reliability of measures
• look for spurious effects and sampling bias
4. Discuss and interpret results
• Do they show what you have expected?
• Do they actually answer your question?
• Did you address all competing hypotheses?
5. Run and repeat the experiment
• Repetition to disclose effects of interaction between independent and noise
variables

Dr. Sven Magg Research Methods - Experiment Design 12


Statistical Inference

 If we have a sample drawn from a population, we can ask two kind of


questions:

1. How “good” is an estimate for a parameter of the population, drawn from this
sample?
How confident are we that the estimate is close to the real parameter value?
Example: Guessing the average number of blonde students from a snapshot
count in the Mensa

2. When answering a yes/no question about the population using the sample, how
likely is it that we are wrong?
Example: My software A is more accurate in guessing the weather than
software B
Dr. Sven Magg Research Methods - Hypothesis Testing 13
Hypotheses
1. a suggested explanation for a group of facts or phenomena, either
accepted as a basis for further verification (working hypothesis) or
accepted as likely to be true […]
3. (Philosophy / Logic) an unproved theory; a conjecture
[Collins English Dictionary]

 has to be testable and falsifiable


 follows from observation, exploratory study, or just idea

 Big question: Is my hypothesis correct?


• What does correct mean? How well can I prove it?
• When do I consider it verified?

Dr. Sven Magg Research Methods - Hypothesis Testing 14


Hypotheses

 Falsifiability and testability


• “All students are female”
 Falsifiable by a single male student
• “When green aliens land in Hamburg, they always step of their spaceship with their
middle foot first”
 Falsifiable in principle, but not in practice
• “A god-like being exists”
“Albert Einstein was the best physicist in the world!”
“What is the sun?”
 Usually not scientifically falsifiable/testable by experiment
 It has to be possible to think of a hypothesis stating the opposite and you can
both test them in practice

Dr. Sven Magg Research Methods - Hypothesis Testing 15


Let’s gamble first…

 You watch a gambler throwing a die three times and always scoring a 6
 You want to accuse him of using a manipulated die!

 What are the chances of you being right?


 Your assumption is that the die is fair and under this assumption you think it’s
unlikely to score three 6s

 Two competing Hypothesis:


• The die is fair (𝐻0 ) (and the result is down to chance)
• The die is manipulated (𝐻1 )

Dr. Sven Magg Research Methods - Hypothesis Testing 16


Gambler’s Thinking Process

 We do not know how the die was manipulated…..


 If the die is fair, we know what the probabilities are:
• 1/6 to get a specific number in one go
• 1/(6 ∗ 6 ∗ 6) = 1/216 = 0.0046 to get three 6s

 You are therefore 99.54% certain that the die was not fair?
 You reject 𝐻0 with a chance of 𝑝 = 0.0046 to be wrong

 What does this say about the chance of the die being manipulated?

Dr. Sven Magg Research Methods - Hypothesis Testing 17


What did we do?

 We have…
1. …stated a null hypothesis 𝐻0 (Die is fair)
2. …thought about a formula to calculate chances, if 𝐻0 is true (binomial
distribution)
3. …observed a result (gathered a sample (6,6,6))
4. …calculated the probability 𝑝 of event to happen if 𝐻0 is true
5. …used 𝑝 as strength of evidence against 𝐻0

 Science is easy! 
 Is 0.46% low enough to accuse the 2m professional boxer of cheating?

Dr. Sven Magg Research Methods - Hypothesis Testing 18


Rejecting is better

 Why is my hypothesis the “alternative” hypothesis 𝐻1 ?


 You can’t prove a hypothesis with statistics on a sample
 But we can estimate the likelihood that a sample was drawn from a given
population!

 𝐻0 : Sample from this (known) population


 𝐻1 : Sample from a different population

 I can statistically evaluate the likelihood that my sample 𝑋 came from a given
population and, if low, reject 𝐻0
 Rejecting 𝐻0  Evidence for 𝐻1
Dr. Sven Magg Research Methods - Hypothesis Testing 19
Another example

 Hypothesis: There are more male than female students in computer science!
 Step 1: State a Null-Hypothesis

Dr. Sven Magg Research Methods - Hypothesis Testing 20


Group Task! 2 5

What is the difference between these


two hypotheses:
1. There are more male than female
students in computer science
2. The ratio of male and female
students is not equal in CS
When do you reject the hypothesis?
Dr. Sven Magg Research Methods - Hypothesis Testing 21
Another example

 Hypothesis: The ratio of male and female students is not equal in CS


 Step 1: State a Null-Hypothesis
• 𝐻0 : 𝑃 𝑚𝑎𝑙𝑒 = 0.5
 Step 2: Find a sampling distribution for 𝑯𝟎
• From 𝐻0 : 𝑃(𝑚𝑎𝑙𝑒) = 𝑃(𝑓𝑒𝑚𝑎𝑙𝑒) = 0.5
• Binomial Distribution
 Step 3: Gather a sample statistic
• 30 students in the Mensa:
x = 20 (male students)
 Step 4: Calculate P(x=20): 0.028
 Step 5: Decide?

Dr. Sven Magg Research Methods - Hypothesis Testing 22


p-values for regions

 Can we reject 𝐻0 because having 20 males in a sample of 30 is unlikely?


 How about 21? Or 28?
 We would reject if we
see 20 or more!

 If 𝐻0 would be rejected for several results, the probability of the combined


result is the sum of individual values:
 𝑃𝑜𝑛𝑒𝑇𝑎𝑖𝑙𝑒𝑑 = 𝑃(20) + ⋯ + 𝑃(30) = 0.049!

Dr. Sven Magg Research Methods - Hypothesis Testing 23


Rejection Regions

 1-Tailed Test
• for all values greater (smaller)
than a given sample statistic
• Used for directional hypotheses
(e.g. “greater than”)
 2-Tailed Test
• Reject 𝐻0 if observed value is greater or lower than one of
two “cut-off” points
• 𝑃𝑡𝑤𝑜𝑇𝑎𝑖𝑙𝑒𝑑 = 0.099
• Typical use:
Reject 𝐻0 when the observed
value differs more than a given
maximum from the mean
Dr. Sven Magg Research Methods - Hypothesis Testing 24
Group Task! 2 5

Thoughts on p-values:
1. What can we use them for?
2. What do they mean for my
Hypothesis 𝐻1 ?
3. What would be good values for
rejection of 𝐻0 ?

Dr. Sven Magg Research Methods - Hypothesis Testing 25


History excursion

 “Early” Ronald Fisher [1925]


• Inductive inference: Use direct probability P(Data| 𝐻0 )
• Only use a Null-Hypothesis 𝐻0 ≙ “happened by chance”, “no effect”
• Use known distribution of a test statistic T, assuming 𝐻0
• Set significance level (.05/.01/.001) following a convention
• Calculate p-value to check whether there is a significant (= backed by statistics)
divergence
• Significance value is a genuine feature of the test
 “Late” Ronald Fisher [1956]
• Calculate exact p-value from the data
• Significance level is a feature of the data themselves
• No use of an arbitrary convention
Dr. Sven Magg Research Methods - Hypothesis Testing 26
History excursion

 Ronald Fisher combined


• Use known distribution of a test statistic T, assuming 𝐻0
• Determine density of values that exceed observed value
• Use p value as strength p-value Strength of evidence
of evidence against H0 0.100 Borderline (or weak)

• p-Value is sample-based 0.050 Moderate

measure of evidence 0.025 Substantial

against null hypothesis 0.010 Strong


0.005 Very Strong
• We report exact p-value,
0.001 Overwhelming
NOT a decision
“P-Values and "significance" measure the probability
of data given the hypothesis, not the probability
of the hypothesis given the data.” - Morris DeGroot
Dr. Sven Magg Research Methods - Hypothesis Testing 27
Neyman-Pearson

 Neyman-Pearson: [1928]
 Define 𝐻0 and alternative Hypothesis 𝐻1
 There are two errors you can make:
• Type I: False rejection (probability )
• Type II: False acceptance (probability )
𝑯𝟏 is true 𝑯𝟎 is true
Correct Outcome Wrong Outcome
Power (1-) Type I (-)Error
Reject 𝑯𝟎
True Positive (TP) False Positive (FP)
Significance Level
Wrong Outcome Correct Outcome
Accept 𝑯𝟎 Type II (-)Error Specificity (1-)
False Negative (FN) True Negative (TN)

Dr. Sven Magg Research Methods - Hypothesis Testing 28


Group Task! 2 5

Colour the regions for , ,power


and specificity
Dr. Sven Magg Research Methods - Hypothesis Testing 29
Group Task! 2 5

Colour the regions for , ,power


and specificity
Dr. Sven Magg Research Methods - Hypothesis Testing 30
 and 

 Error probabilities:
𝑃 𝑇𝑦𝑝𝑒 𝐼 𝐸𝑟𝑟𝑜𝑟 =
𝑃 𝑅𝑒𝑗𝑒𝑐𝑡 𝐻0 𝐻0 𝑡𝑟𝑢𝑒 = 
𝑃 𝑇𝑦𝑝𝑒 𝐼𝐼 𝐸𝑟𝑟𝑜𝑟 =
𝑃 𝐴𝑐𝑐𝑒𝑝𝑡 𝐻0 𝐻1 𝑡𝑟𝑢𝑒 = 
 The power of the test (= 𝑃 𝐴𝑐𝑐𝑒𝑝𝑡 𝐻1 𝐻1 𝑡𝑟𝑢𝑒 ) is dependent on:
• , which is set a priori
• the degree of separation between the two distributions given by δ = 𝜇1 − 𝜇0 (=
Effect size)
• the variance(s) of the population(s)
• the sample size N

Dr. Sven Magg Research Methods - Hypothesis Testing 31


Power of a test

 Analogy
• You are searching for an item in your room
• Power: “What are your chances that you would find the item”
• Depends on:
 How long you are searching Sample Size
 The size of the item Size of effect, i.e. degree of separation
 The messiness of the room Standard deviation
• There is a high chance to find a large item in a clean room if you spend a long time
searching!
• If you can’t find it, you can be confident it wasn’t there
• Power of an experiment: If there really is an effect, how high are the chances that
the experiment would find it?

Dr. Sven Magg Research Methods - Hypothesis Testing 32


Power of a test

 The power of the test (= 𝑃 𝐴𝑐𝑐𝑒𝑝𝑡 𝐻1 𝐻1 𝑡𝑟𝑢𝑒 ) is dependent on:


• , which is set a priori
• the degree of separation between the two distributions given by δ = 𝜇1 − 𝜇0
= Effect Size
• the variance(s) of the
population(s)
• the sample size N

 2 of those can be changed…


 2 can be estimated AFTER the
experiment

Dr. Sven Magg Research Methods - Confidence Intervals 33


How to get the power?

 Power often shown as power curves for one of the four parameters , δ,
variances (σ1 and σ2 ), and N

 Using Monte-Carlo:
1. Fix all parameters but one
2. Draw samples from both
populations
3. Run the test and see whether
you would reject 𝐻0
4. Repeat 2.& 3. k times and calculate #rejections/k
5. Repeat 2.- 4. for all needed values of the free parameter

Dr. Sven Magg Research Methods - Confidence Intervals 34


Sample Size

 You can increase the power of the test by increasing N


 But should you?

 We usually do not increase confidence in our test by increasing N!


• If we want to estimate a parameter, sample size should be as large as you can
afford (CI: 𝜇 = 𝑥ҧ ± 𝑡𝜎ො𝑥ҧ )
• If we want to test a hypothesis, samples should be no larger than required to show
the effect
 If we test a hypothesis and are confident to reject 𝐻0 with sample size N, we
don’t gain much by increasing N
 On the contrary…..

Dr. Sven Magg Research Methods - Confidence Intervals 35


Sample Size

 Can samples be too big?


 By increasing N you can
• boost any real effect to statistical significance
• boost any meaningless effect to significance
 Do not fish for (statistical) significance!

 What we want to know: How much predictive power does our result have?
 In other words: Does knowing which population the sample came from give us
the power to predict it’s value?

Dr. Sven Magg Research Methods - Confidence Intervals 36


Sample size
Sample ഥ
𝒙 s N
 Example from Cohen A 147.95 11.10 1000
 2-sample test: 𝑡 = 2.468 B 146.77 10.16 1000
with 1998 degrees of freedom A&B 147.36 2000
 Statistical significant difference (𝑝 ≤ .05)
 If I hand you one sample from A and let you guess whether it is above the
combined mean A&B, how well would you do?
 517 values from A exceed the combined mean
 464 from B do as well
 You would guess correctly for 51.7% of samples!
 If you don’t know the origin of the sample: 50%

Dr. Sven Magg Research Methods - Confidence Intervals 37


Sample Size

 We can roughly estimate predictive power


 If we know (or can estimate):
• the population variances σ𝐴 2 and σ𝐵 2 of populations A & B
• The variance σ𝑃 2 of the combined population
 then predictive power means reduction in variance from knowing the population
 Relative reduction by knowing sample is from A:
2
σ𝑃 2 −σ𝐴
𝜔2 =
σ𝑃 2
2
 If 𝜔2 2
= 0, then σ𝑃 = σ𝐴 and no prediction is possible
 𝜔2 = 1 means no variance in A, i.e. perfect prediction

Dr. Sven Magg Research Methods - Confidence Intervals 38


Sample Size

 We can only estimate 𝜔2 since we try to find the population parameters!


 If both variances are equal for both populations, a rough estimate is:
2
2
𝑡 −1
𝜔ෝ = 2
𝑡 + 𝑁1 + 𝑁2 − 1
 In the example: 𝜔 ෝ 2 =0.0025
 If we would have drawn only 100 samples each: 𝜔 ෝ 2 =0.027

 Increasing N decreases the predictive power and therefore the meaningfulness


of significant findings

Dr. Sven Magg Research Methods - Confidence Intervals 39


Neyman-Pearson

 Neyman-Pearson: [1928]
 Define 𝐻0 and alternative Hypothesis 𝐻1
 There are two errors you can make:
• Type I: False rejection (probability )
• Type II: False acceptance (probability )
𝑯𝟏 is true 𝑯𝟎 is true
Correct Outcome Wrong Outcome
Power (1-) Type I (-)Error
Reject 𝑯𝟎
True Positive (TP) False Positive (FP)
Significance Level
Wrong Outcome Correct Outcome
Accept 𝑯𝟎 Type II (-)Error Specificity (1-)
False Negative (FN) True Negative (TN)

Dr. Sven Magg Research Methods - Hypothesis Testing 40


Neyman-Pearson- procedure

 Step 1+2 as before (including definition of 𝐻1 )


 Step 3: Set 𝜶
• Decide on a maximum acceptable probability 𝛼 of incorrectly rejecting 𝐻0
 Step 4: Find cut-off points
• Use sampling distribution 𝑁ℎ to find critical values 𝑐 + and 𝑐 −
• Set 𝑐 + and 𝑐 − such that P 𝑁ℎ ≥ 𝑐 + + 𝑃 𝑁ℎ ≤ 𝑐 − ≤ 𝛼
 Step 5: Gather the sample statistic (run experiment)
 Step 6: Decide using X and cut-off points
• If 𝑥 ≥ 𝑐 + 𝑜𝑟 𝑥 ≤ 𝑐 − , reject 𝐻0
 In the example: Set 𝛼 = 0.05  𝑐 + = 21 and 𝑐 − = 9
• 𝑃 𝑁ℎ ≤ 9 + 𝑃 𝑁ℎ ≥ 21 = 0.043 ≤ 𝛼

Dr. Sven Magg Research Methods - Hypothesis Testing 41


Neyman-Pearson

 First define  and  before running the test


  and  are probabilities of making errors of type I / II in the long run and
therefore features of the test
 Set  and  not by a convention but after a detailed cost-benefit analysis of the
consequences

 You have to think prior to the experiment about meaningful values for  and 
• What are the consequences of making a type I or II error?
• Values often used: =.05 and =.2 (= power of 80%)
• Sometimes you want a power of almost 100% and can accept a high probability of
errors of type I

Dr. Sven Magg Research Methods - Hypothesis Testing 42


Neyman-Pearson

 Rules of inductive behaviour


 Rule gives you decision (reject/accept 𝐻0 ) without final statement whether we
believe 𝐻0 is true/false
 An optimal statistical test minimises  while keeping  at a set bound.
 Nowadays often mixed forms of Fisher/Neyman-Pearson can be found that
would satisfy neither:
• Report significance level (usually 1-3 stars or “ns”) and exact p-value
• Define  and use it to reject hypotheses after calculating p
• Use a conventional  and  of 0.05 and .2
• etc….

Dr. Sven Magg Research Methods - Hypothesis Testing 43


What should you do?

 Decide for one side of the debate!


 Either:
1. Think about and set  and  before the test and report findings as significant or
not, stating the significance level
“There was a significant effect (=.05)” or “.. (𝑝 ≤ .01)”, etc.
2. Calculate p-Value for the sample after the test and report exact p-Value without
reporting a decision about 𝐻0

 In the first case, think about the consequences of your errors.


 If you can’t think of any: Use 2.

Dr. Sven Magg Research Methods - Hypothesis Testing 44


What have we learned?

1. Good hypotheses have to be falsifiable in practice


2. We define 𝐻0 (and maybe 𝐻1 ) and try to reject 𝐻0
3. When we reject, there is a chance that we are wrong
4. P-Values are a measure of the probability to falsely reject 𝐻0
5. Fisher’s testing strategy using sample-based p-values
6. Neyman-Pearson’s strategy is to set  and 
7. We can either use a calculated P-Value directly as the strength of evidence
against 𝐻0 or set bounds and reject 𝐻0 if the sample statistic is within the
rejection regions
8. We have to find sampling distributions to make decisions!
Dr. Sven Magg Research Methods - Hypothesis Testing 45

You might also like