0% found this document useful (0 votes)
16 views60 pages

Power and A/B Testing in Data Science

This lecture focuses on the principles of statistical power in hypothesis testing, particularly in A/B testing and bootstrap sampling. It explains the importance of effect size, sample size, and variance in determining statistical power, which is the probability of correctly rejecting a false null hypothesis. Additionally, the document discusses the methodology of permutation tests to analyze differences in sample distributions.

Uploaded by

salaarmasood321
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
16 views60 pages

Power and A/B Testing in Data Science

This lecture focuses on the principles of statistical power in hypothesis testing, particularly in A/B testing and bootstrap sampling. It explains the importance of effect size, sample size, and variance in determining statistical power, which is the probability of correctly rejecting a false null hypothesis. Additionally, the document discusses the methodology of permutation tests to analyze differences in sample distributions.

Uploaded by

salaarmasood321
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

CS334: Principles and

Techniques of Data Science


Lecture 16
Mobin Javed

Slides adapted from Deborah Nolan, Ani Adhikari, Rachel Glennerster, and Ihsan Ayyub Qazi
Outline
● Power of a Test
● A/B Testing
● Estimation: Bootstrap Sampling
Hypothesis Testing:
Can the Conclusion be Wrong?
Yes.

Null is true Alternative is


true
Test rejects the
null ❌ ✅
Test doesn’t
reject the null ✅ ❌
Power: Effect size= 1SE
Sample size = 4,000
0.5

0.45

0.4

H0 H
True effect=H0
0.35
True effect=H
0.3
Significance

0.25

0.2

0.15

0.1

0.05

0
-4 -3 -2 -1 0 1 2 3 4 5 6

J - PAL | SAMPLING AND SAMPLE SIZE 58


The power to avoid Type II errors
● Statistical power is the probability that, if the true effect
is of a given size, our proposed experiment will be able
to distinguish the estimated effect from zero
● Power is the probability of avoiding Type II errors
○ Type II errors: Failing to reject the null hypothesis when it is
false
● Traditionally, we aim for 80% power (some aim for 90%)
● Low power means we may not find a significant effect
even though an effect exists
Questions
● What influences power?
● How do we calculate power in practice?
Power: main ingredients
1. Effect size
Effect
Effect Size: 1*SE Size: 1 * SE
• Hypothesized effect
0.5
size determines distance between
1 Standard
means 0.45 Error

0.4
True effect=H0

H0 0.35

0.3
H True effect=H

0.25

0.2

0.15

0.1

0.05

0
-4 -3 -2 -1 0 1 2 3 4 5 6

J - PAL | SAMPLING AND SAMPLE SIZE 45


Effect Size: 3*SE
Effect Size: 3 * SE
0.5

0.45 3*SE

0.4
True effect=H0

H0 0.35

0.3
H True effect=H

0.25

0.2

0.15

0.1

0.05

0
-4 -3 -2 -1 0 1 2 3 4 5 6

Bigger hypothesized
Bigger hypothesizedeffect
effect size distributions
size àdistributions farther
farther apart
apart
Effect size
Effect size 3*SE:3Power=
* SE:91%
Power= 91%
0.5
True effect=H0
0.45
True effect=H

0.4 Power

H0 0.35

0.3
H
0.25

0.2

0.15

0.1

0.05

0
-4 -3 -2 -1 0 1 2 3 4 5 6

Bigger effect
Bigger size
Effect means more
size means or less power?
more power
Power: main ingredients
1. Effect Size
2. Sample Size
Power: Effect size = 1 SE
Sample
Power: Effectsize = 1,000
size = 1 SE, Sample size = 1,000
0.5

0.45

H0 H
True effect=H0
0.4

True effect=H
0.35
Significance
0.3

0.25

0.2

0.15

0.1

0.05

0
-4 -3 -2 -1 0 1 2 3 4 5 6

J - PAL | SAMPLING AND SAMPLE SIZE 57


Power: Effect size= 1SE
Power:Sample
Effectsize
size = 1 SE,
= 4,000 Sample size = 4,000
0.5

0.45

0.4

H0 H
True effect=H0
0.35
True effect=H
0.3
Significance

0.25

0.2

0.15

0.1

0.05

0
-4 -3 -2 -1 0 1 2 3 4 5 6

J - PAL | SAMPLING AND SAMPLE SIZE 58


Power: 64%
Power: 65%
0.5

0.45

0.4

H0 H
True effect=H0
0.35
True effect=H
0.3
Power

0.25

0.2

0.15

0.1

0.05

0
-4 -3 -2 -1 0 1 2 3 4 5 6

J - PAL | SAMPLING AND SAMPLE SIZE 59


Power: Sample size
Power: Sample size = 9,000
= 9,000
0.5

0.45

0.4

H0 H
True effect=H0
0.35
True effect=H
0.3
Significance

0.25

0.2

0.15

0.1

0.05

0
-4 -3 -2 -1 0 1 2 3 4 5 6

J - PAL | SAMPLING AND SAMPLE SIZE 60


Power: 91% Power: 91%
0.5

0.45

0.4

H0 H
True effect=H0
0.35
True effect=H
0.3
Power

0.25

0.2

0.15

0.1

0.05

0
-4 -3 -2 -1 0 1 2 3 4 5 6

J - PAL | SAMPLING AND SAMPLE SIZE 61


Power: main ingredients
1. Effect Size
2. Sample Size
3. Variance (of the underlying population)
Low variance
Low variance sample sample
0.5

0.45

0.4

H0 H
True effect=H0
0.35
True effect=H
0.3
Significance
0.25

0.2

0.15

0.1

0.05

0
-4 -3 -2 -1 0 1 2 3 4 5 6

Estimates willwillbe
Estimates bemore tightly
more tightly clustered
clustered
Low variance
Low variance sample sample
0.5

0.45

0.4

H0 H
True effect=H0
0.35
True effect=H
0.3
Power
0.25

0.2

0.15

0.1

0.05

0
-4 -3 -2 -1 0 1 2 3 4 5 6

Estimates
Estimates will be will
more be more
tightly tightly
clustered clustered
Higher power
Higher variance sample
Higher variance sample 0.5

0.45

True effect=H0
0.4
True effect=H
0.35
Significance

H0 0.3

0.25
H
0.2

0.15

0.1

0.05

0
-4 -3 -2 -1 0 1 2 3 4 5 6

Estimates will be more dispersed


Estimates will be more dispersed
Power: main ingredients
1. Effect Size
2. Sample Size
3. Variance
4. Proportion of sample in T vs. C
Sample
Sample split:split:
50% C,50%
50% T C, 50% T
0.5

0.45

H0 H
0.4

0.35 True effect=H0

0.3 True effect=H

Significance
0.25

0.2

0.15

0.1

0.05

0
-4 -3 -2 -1 0 1 2 3 4 5 6

Equal splitEqual
gives distributions that are the same “fatness”
split gives distributions that are the same “fatness”
Power: 91% Power: 91%
0.5

0.45

H0 0.4

0.35
H
0.3 True effect=H0

True effect=H
0.25
Power
0.2

0.15

0.1

0.05

0
-4 -3 -2 -1 0 1 2 3 4 5 6

J - PAL | SAMPLING AND SAMPLE SIZE 73


If it’s not 50-50 split?
● What happens to the relative fatness if the split is not
50-50
● Say 25-75?
Sample
Sample split:
split: 25% 25%
C, 75% T C, 75% T
0.5

0.45

0.4

H0 H
0.35

0.3

True effect=H0
0.25

True effect=H
0.2
Significance
0.15

0.1

0.05

0
-4 -3 -2 -1 0 1 2 3 4 5 6

Uneven distributions, not efficient, i.e., less power


Uneven distributions, not efficient, i.e. less power
Power: 83% Power: 83%
0.5

0.45

H0 H
0.4

0.35

0.3 True effect=H0

True effect=H
0.25
Power
0.2

0.15

0.1

0.05

0
-4 -3 -2 -1 0 1 2 3 4 5 6

J - PAL | SAMPLING AND SAMPLE SIZE 76


Allocation ratio
● Definition: the fraction of the total sample allocated to
the treatment group is the allocation ratio
● Usually, for a given sample size, power is maximized
when half sample allocated to treatment, half to control
● Diminishing marginal benefit to precision from adding
sample, so best to add equally
PowerMDE
Power equation: Equation

Significance Variance
Level
Effect Size Power
2
1
EffectSize t1 t * *
P1 P N
Proportion in
Treatment Sample
Size
79
J - PAL | SAMPLING AND SAMPLE SIZE
Summary
● Statistical power is the probability that the study will
find a significant impact when there is one
● Power calculations have to be done to understand the
required sample size to detect impact
○ It is very common for studies to be underpowered
A/B Testing
Is the distribution of samples in A different from
samples in B?
The Groups and the Question
● Random sample of mothers of newborns. Compare:
○ (A) Birth weights of babies of mothers who smoked during
pregnancy
○ (B) Birth weights of babies of mothers who didn’t smoke
● 715 non-smokers
● 459 smokers
Distribution of Baby Weights

Question: Could the difference be due to chance alone?


Hypotheses
● Null:
○ In the population, the distributions of the birth weights of the
babies in the two groups are the same (i.e., they are different in
the sample just due to chance selection of babies and mothers)
● Alternative:
○ In the population, the babies of the mothers who smoked
weighed less, on average, than the babies of the non-smokers
Test Statistic
● Group A: smokers
● Group B: non-smokers

● Statistic: Difference between average weights


Group B average - Group A average

● Large values of this statistic favor the alternative


Simulating Under the Null

...

Non-smoker Non-smoker Smoker Non-smoker Smoker

120 oz 113 oz 128 oz 136 oz 108 oz


Simulating Under the Null

...

Smoker Non-smoker Non-smoker Smoker Non-smoker

120 oz 113 oz 128 oz 136 oz 108 oz


Simulating Under the Null
● If the null hypothesis is true, all rearrangements of the
birth weights among the two groups should be equally
likely (this is known as a permutation test)
● Plan:
○ Shuffle all the birth weights
○ Assign some to “Group A” and the rest to “Group B”,
maintaining the two sample sizes
○ Find the difference between the averages of the two shuffled
groups
○ Repeat
Prediction Under the Null Hypothesis

The observed difference in the original


sample is about −9.27 ounces, which
doesn't even appear on the horizontal
scale of the histogram

The conclusion of the test is that the


data favor the alternative over the null
Permutation Tests
● Tests based on random permutations of the data are
called permutation tests
● A permutation test is a type of a non-parametric test
that allows us to make inferences without making
statistical assumptions that underlie parametric tests
Other Tests
● Parametric: t-test, Z-test
○ Assume Normality

● Non-parametric: Wilcoxon, Mann-Whitney U


Randomized Controlled Experiment
● Sample A: Control group
● Sample B: Treatment group
● If the treatment and control groups are selected at
random, then you can make causal conclusions
● Any difference in outcomes between the two groups
could be due to
○ chance
○ the treatment
Bootstrap Sampling
Inference: Estimation
● How big is an unknown parameter?

● If we have a census (that is, the whole population):


○ We just calculate the parameter and we’re done!

● If we don’t have a census:


○ We take a random sample from the population
○ Use a statistic as an estimate of the parameter
Variability of the Estimate
● One sample ➜ One estimate
o But the random sample could have come out differently
o And so the estimate could have been different
● Main question:
○ How different could the estimate have been?
● The variability of the estimate tells us something about
how accurate the estimate is:
estimate = parameter + error
Where to Get Another Sample?
● One sample ➜ One estimate
● To get many values of the estimate, we need many
random samples
● Can’t go back and sample again from the population:
○ No time, no money
● Stuck?
The Bootstrap
● A technique for simulating repeated random sampling

● All that we have is the original sample


○ … which is large and random
○ Therefore, it probably resembles the population

● So we sample at random from the original sample!


Why the Bootstrap Works

population sample resamples

All of these look pretty similar, most likely!


Why We Need the Bootstrap

population sample resamples

What we wish What we


we could get really get
The Bootstrap Principle
● The bootstrap principle:
○ Bootstrap-world sampling ≈ Real-world sampling
● Not always true!
○ … but reasonable if sample is large enough
● We hope that:
a. Variability of bootstrap estimate
b. Distribution of bootstrap errors
...are similar to what they are in the real world
Key to Resampling
● From the original sample,
○ draw at random
○ with replacement
○ as many values as the original sample contained

● The size of the new sample has to be the same as the


original one, so that the two estimates are comparable
95% Confidence Interval
● Interval of estimates of a parameter
● Based on random sampling
● 95% is called the confidence level
○ Could be any percent between 0 and 100
○ Higher level means wider intervals
● The confidence is in the process that generated the
interval:
○ It generates a “good” interval about 95% of the time
Interpreting Confidence Intervals
● Suppose we would like to estimate the population median
● We draw 5000 bootstrap samples each of size 500
● For each sample, we find the median

Population Median: $110,305.79


Interpreting Confidence Intervals
● The two ends of the "middle
95%" interval of resampled
medians: $102,285 and $115,557
● The "middle 95%" interval of
estimates captured the parameter
in our example. But was that a
fluke?
● To see how frequently the interval
contains the parameter, we have
to run the entire process over and
over again
Interpreting Confidence Intervals
● Specifically, we will repeat the ● The statistical theory of the
following process 100 times: bootstrap says that the count
○ Draw an original sample of size should be around 95 (i.e., the
500 from the population
number of intervals
○ Carry out 5,000 replications of the
bootstrap process and generate containing the true median)
the "middle 95%" interval of
resampled medians
● We will end up with 100
intervals, and count how many
of them contain the population
median
Confidence Interval Summary
● Suppose we would like to estimate the 95% confidence interval for the
population mean
● Steps
○ Draw a large random sample from the population
○ Bootstrap your random sample and get an estimate from the new random sample
○ Repeat the above step thousands of times, and get thousands of estimates
○ Pick off the "middle 95%" interval of all the estimates
○ That gives you one interval of estimates
● Now if you repeat the entire process 100 times, ending up with 100 intervals,
then about 95 of those 100 intervals will contain the population parameter
● In other words, this process of estimation captures the parameter about 95%
of the time
Can You Use a CI Like This?
Suppose an approximate 95% confidence interval for the
average age of the mothers in the population is (26.9,
27.6) years
True or False:
● About 95% of the mothers in the population were
between 26.9 years and 27.6 years old.
Answer: False! We’re estimating that their average age is
in this interval
CI and Bootstrap Sampling
● Bootstrap sampling is often used for constructing CIs
● There are scenarios when it may not be appropriate to
rely on bootstrap sampling
o If you’re trying to estimate very high or very low percentiles, or
min and max
o If you’re trying to estimate any parameter that’s greatly
affected by rare elements of the population
o If the original sample is very small
Using a CI for Hypothesis Testing
● Null hypothesis: Population average = x
● Alternative hypothesis: Population average ≠ x
● Cutoff for P-value: p%
● Method:
○ Construct a (100-p)% confidence interval for the population
average
○ If x is not in the interval, reject the null
○ If x is in the interval, can’t reject the null
Thank you!

You might also like