Clinical Review & Education
JAMA Guide to Statistics and Methods
Sample Size Calculation for a Hypothesis Test
Lynne Stokes, PhD
In this issue of JAMA, Koegelenberg et al1 report the results of a
Figure. Power for Detecting Difference and Sample Size
randomized clinical trial (RCT) that investigated whether treat-
ment with a nicotine patch in addition to varenicline produced 100
MDD = 14%
higher rates of smoking absti-
nence than varenicline alone. 80 MDD = 12%
Related article page 155 The primary results were posi-
tive; that is, patients receiving 60
Power, %
the combination therapy were more likely to achieve continuous
abstinence at 12 weeks than patients receiving varenicline alone.
40
The absolute difference in the abstinence rate was estimated to
be approximately 14%, which was statistically significant at level
20
α = .05.
These findings differed from the results reported in 2 previous
0
studies2,3 of the same question, which detected no difference in 0 50 100 150 200 250 300 350
treatments. What explains this difference? One explanation of- Sample Size for Each Group, No.
fered by the authors is that the previous studies “…may have been
inadequately powered,” which means the sample size in those stud- For a baseline rate of 45% and a minimum detectable difference (MDD) of 14%,
the target sample size of 398 (199 in each group) will produce a power of 80%
ies may have been too small to identify a difference between the
when α is set to .05. When the MDD is 12%, the resulting sample size is 542
treatments tested. (2 × 271) to achieve a power of 80%.
Use of the Method
Why Is Power Analysis Used? minimum detectable difference, or MDD, of 14%), when signifi-
The sample size in a research investigation should be large enough cance level α is set to .05. For this scenario, the authors’ target
that differences occurring by chance are rare but should not be larger sample size of 398 (199 in each group) will produce a power of
than necessary, to avoid waste of resources and to prevent expo- 80%. All these values (45%, 14%, .05, 80%) must be selected at
sure of research participants to risk associated with the interven- the planning stage of the study to carry out this calculation. The
tions. With any study, but especially if the study sample size is very significance level and power are “rule-of-thumb” choices and are
small, any difference in observed rates can happen by chance and typically not based on the specifics of the study. If the researcher
thus cannot be considered statistically significant. wants to reduce the probability of making a type I error (α = .05)
In developing the methods for a study, investigators conduct a or to increase the probability of detecting the specified difference
power analysis to calculate sample size. The power of a hypothesis (power = 80%), then these values can be changed. Either change
test is the probability of obtaining a statistically significant result will require a larger sample size.
when there is a true difference in treatments. For example, sup- Selecting the baseline rate and MDD requires the expertise of
pose, as Koegelenberg et al1 did, that the smoking abstinence rate the researcher. The baseline rate is typically available from the lit-
were 45% for varenicline alone and 14% larger, or 59%, for the com- erature, because this rate is often based on a therapy that has
bination regimen. Power is the probability that, under these condi- been studied. The MDD choice is more subjective. It should be a
tions, the trial would detect a difference in rates large enough to be clinically meaningful rate difference, or a scientifically important
statistically significant at a certain level α (ie, α is the probability of a rate difference, or both, that is also feasible to detect. For
type I error, which occurs by rejecting a null hypothesis that is ac- example, if the combination therapy of varenicline and nicotine
tually true). patch increased abstinence by 0.1%, this difference would not be
Power can also be thought of as the probability of the comple- of practical benefit, would require an extremely large sample size,
ment of a type II error. If we accept a 20% type II error for a differ- and would thus be too small a setting for the MDD. If the MDD
ence in rates of size d, we are saying that there is a 20% chance that were specified as 50%, the new therapy would have to be 95%
we do not detect the difference between groups when the differ- effective (45% + 50%) before there would be a high chance of
ence in their rates is d. The complement of this, 0.8 = 1 − 0.2, or the detecting any difference, so would be too large for the MDD. The
statistical power, means that when a difference of d exists, there is authors based their choice of MDD = 14% on a compromise
an 80% chance that our statistical test will detect it. between their judgment of a clinically important difference, 12%,
The Figure illustrates the relationship between sample size and the scientifically meaningful value of 16%. The 16% rate was
and power for the test described. The orange line shows the the observed difference in a study that compared varenicline
power for the parameter settings above (baseline rate of 45% and alone and together with nicotine gum.4 Thus, the ability to con-
180 JAMA July 9, 2014 Volume 312, Number 2 [Link]
Copyright 2014 American Medical Association. All rights reserved.
Downloaded From: [Link] by a Karolinska Institutet University Library User on 12/12/2022
JAMA Guide to Statistics and Methods Clinical Review & Education
firm a difference that is slightly smaller for a related treatment How Should This Method’s Findings Be Interpreted
was considered scientifically important. in This Particular Study?
A power analysis can help with the interpretation of study findings
What Are the Limitations of Power Analysis? when statistically significant effects are not found. However, be-
Calculation of sample size requires predictions of baseline rates and cause the findings in the study by Koegelenberg et al1 were statis-
MDD, which may not be readily available, before the study begins. tically significant, interpretation of a lack of significance was unnec-
The sample size is especially sensitive to the MDD. This is illus- essary. If no statistically significant difference in abstinence rates had
trated by the blue line in the Figure, which shows the sample size been found, the authors could have noted that, “The study was suf-
needed in this study if the MDD were set to 12%. The resulting sample ficiently powered to have a high chance of detecting a difference of
size is 542 (2 × 271) to achieve a power of 80%. 14% in abstinence rates. Thus, any undetected difference is likely
This method of conducting a power analysis might also pro- to be of little clinical benefit.”
duce the incorrect sample size if the data analysis conducted dif-
fers from that planned. For example, if abstinence were affected Caveats to Consider When Looking at Results Based
by other covariates, such as age, and the groups were unbalanced on Power Analysis
on this variable, other analyses might be used, such as logistic Sample size calculation based on any power analysis requires input
regression models accounting for covariate differences. The from the researcher prior to the study. Some of these are assump-
sample size that would be appropriate for one analysis may be tions and predictions of fact (such as the baseline rate), which may
too large or small to achieve the same power with another ana- be incorrect. Others reflect the clinical judgment of the researcher
lytic procedure. (eg, MDD), with which the reader may disagree. If a statistically sig-
nificant effect is not found, the reader should assess whether either
Why Did the Authors Use Power Analysis in This Particular Study? of these are concerns.
The number of research participants available for any study is lim- The reader should also not interpret a lack of significance for an
ited by resources. However, the authors were aware that previous outcome other than the one on which the power analysis was based
studies comparing these treatments had found no significant dif- as confirmation that no difference exists, because the analysis is spe-
ference in abstinence rates. This can occur even if a difference cific to the parameter settings. For example, no significant differ-
exists if the sample size is too small. The authors wanted to ence was found in this study for most adverse events rates, al-
ensure that their sample size was adequate to detect even a small though the power analysis does not apply to these rates. Thus, the
but clinically important difference, so they carefully evaluated sample size may not be adequate to interpret that finding to con-
sample size. firm that no meaningful difference in these outcomes exists.
ARTICLE INFORMATION Disclosure of Potential Conflicts of Interest and effective in helping smokers quit than varenicline
Author Affiliation: Department of Statistical none were reported. alone? a randomised controlled trial. BMC Med.
Science, Southern Methodist University, Dallas, 2013;11:140.
Texas. REFERENCES 3. Ebbert JO, Burke MV, Hays JT, Hurt RD.
Corresponding Author: Lynne Stokes, PhD, 1. Koegelenberg CFN, Noor F, Bateman ED, et al. Combination treatment with varenicline and
Department of Statistical Science, Southern Efficacy of varenicline combined with nicotine nicotine replacement therapy. Nicotine Tob Res.
Methodist University, PO Box 750100, Dallas, TX replacement therapy vs varenicline alone for 2009;11(5):572-576.
75275 (slstokes@[Link]). smoking cessation: a randomized clinical trial. JAMA. 4. Besada NA, Guerrero AC, Fernandez MI, Ulibarri
doi:10.1001/jama.2014.7195. MM, Jiménez-Ruiz CA. Clinical experience from a
Conflict of Interest Disclosures: The author has
completed and submitted the ICMJE Form for 2. Hajek P, Smith KM, Dhanji AR, McRobbie H. Is a smokers clinic combining varenicline and nicotine
combination of varenicline and nicotine patch more gum. Eur Respir J. 2010;36(suppl 54):462s.
[Link] JAMA July 9, 2014 Volume 312, Number 2 181
Copyright 2014 American Medical Association. All rights reserved.
Downloaded From: [Link] by a Karolinska Institutet University Library User on 12/12/2022