0% found this document useful (0 votes)
17 views52 pages

Sample Size Calculations in Research

The document is the third edition of 'Sample Size Calculations in Clinical Research' edited by Shein-Chung Chow, Jun Shao, Hansheng Wang, and Yuliya Lokhnygina, published by CRC Press in 2018. It covers essential topics such as regulatory requirements, study design, and statistical methods for sample size calculations in clinical trials. The book includes various methodologies and practical considerations for researchers in the field of clinical medicine and drug development.

Uploaded by

lythienphuc2014
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
17 views52 pages

Sample Size Calculations in Research

The document is the third edition of 'Sample Size Calculations in Clinical Research' edited by Shein-Chung Chow, Jun Shao, Hansheng Wang, and Yuliya Lokhnygina, published by CRC Press in 2018. It covers essential topics such as regulatory requirements, study design, and statistical methods for sample size calculations in clinical trials. The book includes various methodologies and practical considerations for researchers in the field of clinical medicine and drug development.

Uploaded by

lythienphuc2014
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Sample Size Calculations

in Clinical Research
Third Edition
Published Titles

1. Design and Analysis of Animal Studies in Pharmaceutical Development, Shein-Chung


Chow and Jen-pei Liu
2. Basic Statistics and Pharmaceutical Statistical Applications, James E. De Muth
3. Design and Analysis of Bioavailability and Bioequivalence Studies, Second Edition,
Revised and Expanded, Shein-Chung Chow and Jen-pei Liu
4. Meta-Analysis in Medicine and Health Policy, Dalene K. Stangl and Donald A. Berry
5. Generalized Linear Models: A Bayesian Perspective, Dipak K. Dey, Sujit K. Ghosh, and
Bani K. Mallick
6. Difference Equations with Public Health Applications, Lemuel A. Moyé and Asha Seth
Kapadia
7. Medical Biostatistics, Abhaya Indrayan and Sanjeev B. Sarmukaddam
8. Statistical Methods for Clinical Trials, Mark X. Norleans
9. Causal Analysis in Biomedicine and Epidemiology: Based on Minimal Sufficient Causation,
Mikel Aickin
10. Statistics in Drug Research: Methodologies and Recent Developments, Shein-Chung
Chow and Jun Shao
11. Sample Size Calculations in Clinical Research, Shein-Chung Chow, Jun Shao, and
Hansheng Wang
12. Applied Statistical Design for the Researcher, Daryl S. Paulson
13. Advances in Clinical Trial Biostatistics, Nancy L. Geller
14. Statistics in the Pharmaceutical Industry, 3rd Edition, Ralph Buncher and Jia-Yeong
Tsay
15. DNA Microarrays and Related Genomics Techniques: Design, Analysis, and Interpretation
of Experiments, David B. Allsion, Grier P. Page, T. Mark Beasley, and Jode W.
Edwards
16. Basic Statistics and Pharmaceutical Statistical Applications, Second Edition, James E. De
Muth
17. Adaptive Design Methods in Clinical Trials, Shein-Chung Chow and Mark Chang
18. Handbook of Regression and Modeling: Applications for the Clinical and Pharmaceutical
Industries, Daryl S. Paulson
19. Statistical Design and Analysis of Stability Studies, Shein-Chung Chow
20. Sample Size Calculations in Clinical Research, Second Edition, Shein-Chung Chow, Jun
Shao, and Hansheng Wang
Sample Size Calculations
in Clinical Research
Third Edition

Shein-Chung Chow, Jun Shao,


Hansheng Wang, and Yuliya Lokhnygina
CRC Press
Taylor & Francis Group
6000 Broken Sound Parkway NW, Suite 300
Boca Raton, FL 33487-2742

© 2018 by Taylor & Francis Group, LLC


CRC Press is an imprint of Taylor & Francis Group, an Informa business

No claim to original U.S. Government works

Printed on acid-free paper

International Standard Book Number-13: 978-1-138-74098-3 (Hardback)

This book contains information obtained from authentic and highly regarded sources. Reasonable efforts have been
made to publish reliable data and information, but the author and publisher cannot assume responsibility for the
validity of all materials or the consequences of their use. The authors and publishers have attempted to trace the
copyright holders of all material reproduced in this publication and apologize to copyright holders if permission to
publish in this form has not been obtained. If any copyright material has not been acknowledged please write and let
us know so we may rectify in any future reprint.

Except as permitted under U.S. Copyright Law, no part of this book may be reprinted, reproduced, transmitted, or
utilized in any form by any electronic, mechanical, or other means, now known or hereafter invented, including pho-
tocopying, microfilming, and recording, or in any information storage or retrieval system, without written permission
from the publishers.

For permission to photocopy or use material electronically from this work, please access [Link] (http://
[Link]/) or contact the Copyright Clearance Center, Inc. (CCC), 222 Rosewood Drive, Danvers, MA
01923, 978-750-8400. CCC is a not-for-profit organization that provides licenses and registration for a variety of users.
For organizations that have been granted a photocopy license by the CCC, a separate system of payment has been
arranged.

Trademark Notice: Product or corporate names may be trademarks or registered trademarks, and are used only for
identification and explanation without intent to infringe.

Library of Congress Cataloging-in-Publication Data

Names: Chow, Shein-Chung, 1955- editor. | Shao, Jun (Statistician) editor. | Wang, Hansheng,
1977- editor. | Lokhnygina, Yuliya, editor.
Title: Sample size calculations in clinical research / [edited by] Shein-Chung Chow, Jun Shao,
Hansheng Wang, Yuliya Lokhnygina.
Description: Third edition. | Boca Raton : Taylor & Francis, 2017. | Series:
Chapman & Hall/CRC biostatistics series | “A CRC title, part of the Taylor & Francis imprint,
a member of the Taylor & Francis Group, the academic division of T&F Informa plc.” | Includes
bibliographical references.
Identifiers: LCCN 2017011239 | ISBN 9781138740983 (hardback)
Subjects: LCSH: Clinical medicine--Research--Statistical methods. | Drug development--Statistical
methods. | Sampling (Statistics)
Classification: LCC R853.S7 S33 2017 | DDC 610.72/7--dc23 LC record available at
[Link]

Visit the Taylor & Francis Web site at


[Link]
and the CRC Press Web site at
[Link]
Contents

Preface........................................................................................................................................... xvii

1. Introduction..............................................................................................................................1
1.1 Regulatory Requirement...............................................................................................2
1.1.1 Adequate and Well-Controlled Clinical Trials............................................. 2
1.1.2 Substantial Evidence........................................................................................ 3
1.1.3 Why At Least Two Studies?.............................................................................3
1.1.4 Substantial Evidence with a Single Trial....................................................... 4
1.1.5 Sample Size........................................................................................................5
1.2 Basic Considerations......................................................................................................5
1.2.1 Study Objectives...............................................................................................6
1.2.2 Study Design.....................................................................................................6
1.2.3 Hypotheses........................................................................................................ 7
[Link] Test for Equality.................................................................................7
[Link] Test for Noninferiority......................................................................8
[Link] Test for Superiority............................................................................8
[Link] Test for Equivalence.......................................................................... 8
[Link] Relationship among Noninferiority,
Superiority, and Equivalence........................................................... 9
1.2.4 Primary Study Endpoint..................................................................................9
1.2.5 Clinically Meaningful Difference................................................................ 10
1.3 Procedures for Sample Size Calculation................................................................... 11
1.3.1 Type I and Type II Errors............................................................................... 11
1.3.2 Precision Analysis.......................................................................................... 12
1.3.3 Power Analysis................................................................................................ 13
1.3.4 Probability Assessment.................................................................................. 15
1.3.5 Reproducibility Probability........................................................................... 16
1.3.6 Sample Size Reestimation without Unblinding......................................... 18
1.4 Aims and Structure of this Book............................................................................... 18
1.4.1 Aim of this Book............................................................................................. 18
1.4.2 Structure of this Book.................................................................................... 19

2. Considerations Prior to Sample Size Calculation........................................................... 21


2.1 Confounding and Interaction..................................................................................... 21
2.1.1 Confounding................................................................................................... 21
2.1.2 Interaction........................................................................................................22
2.1.3 Remark.............................................................................................................22
2.2 One-Sided Test versus Two-Sided Test..................................................................... 23
2.2.1 Remark............................................................................................................. 24
2.3 Crossover Design versus Parallel Design................................................................. 24
2.3.1 Intersubject and Intrasubject Variabilities.................................................. 25
2.3.2 Crossover Design............................................................................................ 25
2.3.3 Parallel Design................................................................................................ 26
2.3.4 Remark............................................................................................................. 26

v
vi Contents

2.4 Subgroup/Interim Analyses....................................................................................... 26


2.4.1 Group Sequential Boundaries....................................................................... 27
2.4.2 Alpha Spending Function............................................................................. 28
2.5 Data Transformation.................................................................................................... 29
2.5.1 Remark............................................................................................................. 31
2.6 Practical Issues............................................................................................................. 31
2.6.1 Unequal Treatment Allocation...................................................................... 31
2.6.2 Adjustment for Dropouts or Covariates...................................................... 32
2.6.3 Mixed-Up Randomization Schedules.......................................................... 33
2.6.4 Treatment or Center Imbalance.................................................................... 35
2.6.5 Multiplicity...................................................................................................... 37
2.6.6 Multiple-Stage Design for Early Stopping.................................................. 37
2.6.7 Rare Incidence Rate........................................................................................ 38

3. Comparing Means................................................................................................................. 39
3.1 One-Sample Design..................................................................................................... 39
3.1.1 Test for Equality.............................................................................................. 40
3.1.2 Test for Noninferiority/Superiority.............................................................42
3.1.3 Test for Equivalence........................................................................................44
3.1.4 An Example..................................................................................................... 45
[Link] Test for Equality............................................................................... 45
[Link] Test for Noninferiority.................................................................... 46
[Link] Test for Equivalence........................................................................ 46
3.2 Two-Sample Parallel Design....................................................................................... 47
3.2.1 Test for Equality.............................................................................................. 47
3.2.2 Test for Noninferority/Superiority............................................................... 50
3.2.3 Test for Equivalence........................................................................................ 51
3.2.4 An Example..................................................................................................... 52
[Link] Test for Equality............................................................................... 53
[Link] Test for Noninferiority.................................................................... 53
[Link] Test for Equivalence........................................................................54
3.2.5 Remarks............................................................................................................54
3.3 Two-Sample Crossover Design..................................................................................54
3.3.1 Test for Equality.............................................................................................. 55
3.3.2 Test for Noninferiority/Superiority............................................................. 56
3.3.3 Test for Equivalence........................................................................................ 57
3.3.4 An Example..................................................................................................... 58
[Link] Therapeutic Equivalence................................................................ 58
[Link] Noninferiority.................................................................................. 58
3.3.5 Remarks............................................................................................................ 59
3.4 Multiple-Sample One-Way ANOVA.......................................................................... 59
3.4.1 Pairwise Comparison..................................................................................... 60
3.4.2 Simultaneous Comparison............................................................................ 61
3.4.3 An Example..................................................................................................... 61
3.4.4 Remarks............................................................................................................ 62
3.5 Multiple-Sample Williams Design............................................................................63
3.5.1 Test for Equality..............................................................................................64
3.5.2 Test for Noninferiority/Superiority.............................................................65
Contents vii

3.5.3 Test for Equivalence........................................................................................65


3.5.4 An Example..................................................................................................... 66
3.6 Practical Issues............................................................................................................. 66
3.6.1 One-Sided versus Two-Sided Test................................................................ 67
3.6.2 Parallel Design versus Crossover Design.................................................... 67
3.6.3 Sensitivity Analysis........................................................................................ 68

4. Large Sample Tests for Proportions................................................................................... 71


4.1 One-Sample Design..................................................................................................... 71
4.1.1 Test for Equality.............................................................................................. 72
4.1.2 Test for Noninferiority/Superiority............................................................. 73
4.1.3 Test for Equivalence........................................................................................ 74
4.1.4 An Example..................................................................................................... 74
[Link] Test for Equality............................................................................... 75
[Link] Test for Noninferiority.................................................................... 75
[Link] Test for Equivalence........................................................................ 75
4.1.5 Remarks............................................................................................................ 75
4.2 Two-Sample Parallel Design....................................................................................... 76
4.2.1 Test for Equality..............................................................................................77
4.2.2 Test for Noninferiority/Superiority.............................................................77
4.2.3 Test for Equivalence........................................................................................ 78
4.2.4 An Example..................................................................................................... 79
[Link] Test for Equality...............................................................................80
[Link] Test for Noninferiority....................................................................80
[Link] Test for Superiority..........................................................................80
[Link] Test for Equivalence........................................................................ 81
4.2.5 Remarks............................................................................................................ 81
4.3 Two-Sample Crossover Design.................................................................................. 82
4.3.1 Test for Equality..............................................................................................83
4.3.2 Test for Noninferiority/Superiority.............................................................84
4.3.3 Test for Equivalence........................................................................................84
4.3.4 An Example..................................................................................................... 85
[Link] Test for Equality............................................................................... 85
[Link] Test for Noninferiority.................................................................... 86
[Link] Test for Equivalence........................................................................ 86
4.3.5 Remarks............................................................................................................ 86
4.4 One-Way Analysis of Variance.................................................................................. 86
4.4.1 Pairwise Comparison..................................................................................... 87
4.4.2 An Example..................................................................................................... 87
4.4.3 Remarks............................................................................................................ 88
4.5 Williams Design........................................................................................................... 88
4.5.1 Test for Equality.............................................................................................. 89
4.5.2 Test for Noninferiority/Superiority............................................................. 89
4.5.3 Test for Equivalence........................................................................................ 90
4.5.4 An Example..................................................................................................... 91
[Link] Test for Equality............................................................................... 91
[Link] Test for Superiority.......................................................................... 92
[Link] Test for Equivalence........................................................................ 92
viii Contents

4.6 Relative Risk—Parallel Design................................................................................... 92


4.6.1 Test for Equality.............................................................................................. 93
4.6.2 Test for Noninferiority/Superiority............................................................. 94
4.6.3 Test for Equivalence........................................................................................ 94
4.6.4 An Example..................................................................................................... 95
[Link] Test for Equality............................................................................... 95
[Link] Test for Superiority.......................................................................... 96
[Link] Test for Equivalence........................................................................ 96
4.7 Relative Risk—Crossover Design.............................................................................. 96
4.7.1 Test for Equality.............................................................................................. 97
4.7.2 Test for Noninferiority/Superiority............................................................. 97
4.7.3 Test for Equivalence........................................................................................ 98
4.8 Practical Issues............................................................................................................. 99
4.8.1 Exact and Asymptotic Tests.......................................................................... 99
4.8.2 Variance Estimates.......................................................................................... 99
4.8.3 Stratified Analysis......................................................................................... 101
4.8.4 Equivalence Test for More Than Two Proportions.................................. 102
5. Exact Tests for Proportions................................................................................................ 103
5.1 Binomial Test.............................................................................................................. 103
5.1.1 The Procedure............................................................................................... 103
5.1.2 Remarks.......................................................................................................... 104
5.1.3 An Example................................................................................................... 104
5.2 Negative Binomial...................................................................................................... 105
5.2.1 Negative Binomial Distribution.................................................................. 106
5.2.2 Sample Size Requirement............................................................................ 107
5.3 Fisher’s Exact Test....................................................................................................... 108
5.3.1 The Procedure............................................................................................... 109
5.3.2 Remarks.......................................................................................................... 109
5.3.3 An Example................................................................................................... 109
5.4 Optimal Multiple-Stage Designs for Single-Arm Trials....................................... 111
5.4.1 Optimal Two-Stage Designs........................................................................ 111
5.4.2 Flexible Two-Stage Designs......................................................................... 113
5.4.3 Optimal Three-Stage Designs..................................................................... 114
5.5 Flexible Designs for Multiple-Arm Trials............................................................... 122
5.6 Remarks....................................................................................................................... 129
6. Tests for Goodness-of-Fit and Contingency Tables...................................................... 131
6.1 Tests for Goodness-of-Fit.......................................................................................... 131
6.1.1 Pearson’s Test................................................................................................. 131
6.1.2 An Example................................................................................................... 132
6.2 Test for Independence: Single Stratum................................................................... 133
6.2.1 Pearson’s Test................................................................................................. 134
6.2.2 Likelihood Ratio Test................................................................................... 135
6.2.3 An Example................................................................................................... 136
6.3 Test for Independence: Multiple Strata................................................................... 136
6.3.1 Cochran–Mantel–Haenszel Test................................................................. 137
6.3.2 An Example................................................................................................... 138
6.4 Test for Categorical Shift........................................................................................... 138
6.4.1 McNemar’s Test............................................................................................. 139
Contents ix

6.4.2 Stuart–Maxwell Test..................................................................................... 141


6.4.3 Examples........................................................................................................ 142
[Link] McNemar’s Test............................................................................. 142
[Link] Stuart–Maxwell Test..................................................................... 142
6.5 Carryover Effect Test................................................................................................. 143
6.5.1 Test Procedure............................................................................................... 143
6.5.2 An Example................................................................................................... 145
6.6 Practical Issues........................................................................................................... 145
6.6.1 Local Alternative versus Fixed Alternative.............................................. 145
6.6.2 Random versus Fixed Marginal Total....................................................... 146
6.6.3 r × c versus p × r × c.................................................................................... 146

7. Comparing Time-to-Event Data....................................................................................... 147


7.1 Basic Concepts............................................................................................................ 147
7.1.1 Survival Function......................................................................................... 148
7.1.2 Median Survival Time................................................................................. 148
7.1.3 Hazard Function........................................................................................... 148
7.1.4 An Example................................................................................................... 149
7.2 Exponential Model..................................................................................................... 150
7.2.1 Test for Equality............................................................................................ 152
7.2.2 Test for Noninferiority/Superiority........................................................... 153
7.2.3 Test for Equivalence...................................................................................... 154
7.2.4 An Example................................................................................................... 155
[Link] Test for Equality............................................................................. 155
[Link] Test for Superiority........................................................................ 156
[Link] Test for Equivalence...................................................................... 156
7.2.5 Remarks.......................................................................................................... 156
[Link] Unconditional versus Conditional.............................................. 156
[Link] Losses to Follow-Up, Dropout, and Noncompliance............... 157
7.3 Cox’s Proportional Hazards Model......................................................................... 158
7.3.1 Test for Equality............................................................................................ 159
7.3.2 Test for Noninferiority/Superiority........................................................... 161
7.3.3 Test for Equivalence...................................................................................... 162
7.3.4 An Example................................................................................................... 162
[Link] Test for Equality............................................................................. 162
[Link] Test for Superiority........................................................................ 163
[Link] Test for Equivalence...................................................................... 163
7.4 Weighted Log-Rank Test........................................................................................... 163
7.4.1 Tarone–Ware Test.......................................................................................... 163
7.4.2 An Example................................................................................................... 165
7.5 Practical Issues........................................................................................................... 168
7.5.1 Binomial versus Time to Event................................................................... 168
7.5.2 Local Alternative versus Fixed Alternative.............................................. 168
7.5.3 One-Sample versus Historical Control...................................................... 168

8. Group Sequential Methods............................................................................................... 169


8.1 Pocock’s Test............................................................................................................... 169
8.1.1 The Procedure............................................................................................... 169
8.1.2 An Example................................................................................................... 171
x Contents

8.2 O’Brien and Fleming’s Test....................................................................................... 172


8.2.1 The Procedure............................................................................................... 173
8.2.2 An Example................................................................................................... 174
8.3 Wang and Tsiatis’ Test............................................................................................... 175
8.3.1 The Procedure............................................................................................... 175
8.3.2 An Example................................................................................................... 175
8.4 Inner Wedge Test........................................................................................................ 177
8.4.1 The Procedure............................................................................................... 177
8.4.2 An Example................................................................................................... 178
8.5 Binary Variables......................................................................................................... 180
8.5.1 The Procedure............................................................................................... 180
8.5.2 An Example................................................................................................... 180
8.6 Time-to-Event Data.................................................................................................... 181
8.6.1 The Procedure............................................................................................... 181
8.6.2 An Example................................................................................................... 182
8.7 Alpha-Spending Function........................................................................................ 183
8.8 Sample Size Reestimation......................................................................................... 185
8.8.1 The Procedure............................................................................................... 185
8.8.2 An Example................................................................................................... 186
8.9 Conditional Power..................................................................................................... 187
8.9.1 Comparing Means........................................................................................ 187
8.9.2 Comparing Proportions............................................................................... 188
8.10 Practical Issues........................................................................................................... 189

9. Comparing Variabilities..................................................................................................... 191


9.1 Comparing Intrasubject Variabilities...................................................................... 191
9.1.1 Parallel Design with Replicates.................................................................. 192
[Link] Test for Equality............................................................................. 192
[Link] Test for Noninferiority/Superiority............................................ 193
[Link] Test for Similarity.......................................................................... 194
[Link] An Example.................................................................................... 195
9.1.2 Replicated Crossover Design...................................................................... 195
[Link] Test for Equality............................................................................. 197
[Link] Test for Noninferiority/Superiority............................................ 198
[Link] Test for Similarity.......................................................................... 198
[Link] An Example.................................................................................... 199
9.2 Comparing Intrasubject CVs.................................................................................... 200
9.2.1 Simple Random Effects Model.................................................................... 200
[Link] Test for Equality............................................................................. 201
[Link] Test for Noninferiority/Superiority............................................ 202
[Link] Test for Similarity.......................................................................... 203
[Link] An Example.................................................................................... 203
9.2.2 Conditional Random Effects Model........................................................... 204
[Link] Test for Equality............................................................................. 206
[Link] Test for Noninferiority/Superiority............................................ 206
[Link] Test for Similarity.......................................................................... 207
[Link] An Example.................................................................................... 208
[Link] Remarks.......................................................................................... 208
Contents xi

9.3 Comparing Intersubject Variabilities...................................................................... 209


9.3.1 Parallel Design with Replicates.................................................................. 209
[Link] Test for Equality............................................................................. 210
[Link] Test for Noninferiority/Superiority............................................ 211
[Link] An Example.................................................................................... 212
9.3.2 Replicated Crossover Design...................................................................... 213
[Link] Test for Equality............................................................................. 213
[Link] Test for Noninferiority/Superiority............................................ 215
[Link] An Example.................................................................................... 216
9.4 Comparing Total Variabilities.................................................................................. 217
9.4.1 Parallel Designs without Replicates........................................................... 217
[Link] Test for Equality............................................................................. 218
[Link] Test for Noninferiority/Superiority............................................ 219
[Link] Test for Similarity.......................................................................... 219
[Link] An Example.................................................................................... 220
9.4.2 Parallel Design with Replicates.................................................................. 221
[Link] Test for Equality............................................................................. 221
[Link] Test for Noninferiority/Superiority............................................222
[Link] An Example....................................................................................223
9.4.3 The Standard 2 × 2 Crossover Design....................................................... 224
[Link] Test for Equality............................................................................. 224
[Link] Test for Noninferiority/Superiority............................................ 226
[Link] An Example.................................................................................... 227
9.4.4 Replicated 2 × 2m Crossover Design......................................................... 227
[Link] Test for Equality............................................................................. 227
[Link] Test for Noninferiority/Superiority............................................ 229
[Link] An Example.................................................................................... 230
9.5 Practical Issues........................................................................................................... 231

10. Bioequivalence Testing...................................................................................................... 233


10.1 Bioequivalence Criteria.............................................................................................234
10.2 Average Bioequivalence............................................................................................ 235
10.2.1 An Example................................................................................................... 237
10.3 Population Bioequivalence....................................................................................... 238
10.3.1 An Example................................................................................................... 240
10.4 Individual Bioequivalence........................................................................................ 242
10.4.1 An Example................................................................................................... 246
10.5 In Vitro Bioequivalence.............................................................................................. 248
10.5.1 An Example................................................................................................... 252
10.6 Sample Size Requirement for Analytical Similarity Assessment of
Biosimilar Products................................................................................................... 253
10.6.1 FDA’s Tiered Approach................................................................................ 253
10.6.2 Sample Size Requirement............................................................................ 253

11. Dose–Response Studies..................................................................................................... 257


11.1 Continuous Response................................................................................................ 257
11.1.1 Linear Contrast Test..................................................................................... 258
11.2 Binary Response......................................................................................................... 261
xii Contents

11.3 Time-to-Event Endpoint............................................................................................ 262


11.4 Williams’ Test for Minimum Effective Dose......................................................... 264
11.5 Cochran–Armitage’s Test for Trend........................................................................ 268
11.6 Dose Escalation Trials............................................................................................... 271
11.6.1 A + B Escalation Design without Dose De-Escalation............................ 272
11.6.2 A + B Escalation Design with Dose De-Escalation.................................. 274
11.7 Concluding Remarks................................................................................................. 276

12. Microarray Studies.............................................................................................................. 277


12.1 Literature Review....................................................................................................... 277
12.2 FDR Control................................................................................................................ 278
12.2.1 Model and Assumptions............................................................................. 278
12.2.2 Sample Size Calculation............................................................................... 280
12.3 FWER Control............................................................................................................. 288
12.3.1 Multiple Testing Procedures....................................................................... 288
12.3.2 Sample Size Calculation............................................................................... 290
12.3.3 Leukemia Example....................................................................................... 293
12.4 Concluding Remarks................................................................................................. 295

13. Bayesian Sample Size Calculation................................................................................... 297


13.1 Posterior Credible Interval Approach..................................................................... 298
13.1.1 Three Selection Criteria............................................................................... 298
[Link] Average Coverage Criterion......................................................... 299
[Link] Average Length Criterion............................................................ 299
[Link] Worst Outcome Criterion.............................................................300
13.1.2 One Sample....................................................................................................300
[Link] Known Precision...........................................................................300
[Link] Unknown Precision...................................................................... 301
[Link] Mixed Bayesian-Likelihood......................................................... 302
13.1.3 Two-Sample with Common Precision.......................................................305
[Link] Known Common Precision..........................................................306
[Link] Unknown Common Precision..................................................... 307
13.1.4 Two-Sample with Unequal Precisions.......................................................308
[Link] Known Precision........................................................................... 310
[Link] Unknown Precisions..................................................................... 311
13.2 Posterior Error Approach.......................................................................................... 312
13.2.1 Posterior Error Rate...................................................................................... 312
13.2.2 Comparing Means........................................................................................ 314
13.3 Bootstrap-Median Approach.................................................................................... 316
13.3.1 Background.................................................................................................... 317
13.3.2 Bootstrap-Median Approach...................................................................... 318
13.4 Concluding Remarks................................................................................................. 319

14. Nonparametrics.................................................................................................................... 321


14.1 Violation of Assumptions......................................................................................... 321
14.2 One-Sample Location Problem................................................................................ 323
14.2.1 Remark........................................................................................................... 326
14.2.2 An Example................................................................................................... 326
Contents xiii

14.3 Two-Sample Location Problem................................................................................ 327


14.3.1 Remark........................................................................................................... 329
14.3.2 An Example................................................................................................... 330
14.4 Test for Independence............................................................................................... 330
14.4.1 An Example................................................................................................... 333
14.5 Practical Issues...........................................................................................................334
14.5.1 Bootstrapping................................................................................................334
14.5.2 Comparing Variabilities...............................................................................334
14.5.3 Multiple-Sample Location Problem............................................................334
14.5.4 Testing Scale Parameters............................................................................. 335

15. Sample Size Calculations for Cluster Randomized Trials......................................... 337


15.1 Unmatched Trials....................................................................................................... 338
15.1.1 Comparison of Means.................................................................................. 339
15.1.2 Comparison of Proportions.........................................................................340
15.1.3 Comparison of Incidence Rates.................................................................. 341
15.1.4 Further Remarks...........................................................................................342
15.2 Matched Trials............................................................................................................342
15.2.1 Comparison of Means..................................................................................343
15.2.2 Comparison of Proportions.........................................................................344
15.2.3 Comparison of Incidence Rates..................................................................345
15.3 Stratified Trials........................................................................................................... 347

16. Test for Homogeneity of Two Zero-Inflated Poisson Population.............................. 349


16.1 Zero-Inflated Poisson Distribution.......................................................................... 350
16.2 Testing Differences between Treatment Groups................................................... 351
16.2.1 Testing the Difference in Both Groups of Zeros and Nonzeros............ 352
16.2.2 Testing the Difference in the Groups of Zeros......................................... 354
16.2.3 Testing the Difference of the Groups of Nonzeros.................................. 355
16.3 Sample Size Calculation............................................................................................ 356
16.3.1 Testing the Difference in the Groups of Both Zeros and Nonzeros...... 356
16.3.2 Testing the Difference in the Groups of Zeros between Treatments...... 357
16.3.3 Testing the Difference in the Groups of Nonzeros between
Treatments..................................................................................................... 357
16.3.4 An Example................................................................................................... 361
16.4 Multivariate ZIP......................................................................................................... 363
16.4.1 Bivariate ZIP.................................................................................................. 363
16.4.2 Comparing the Effects of Control and Test Treatment............................ 366
16.4.3 Sample Size Calculation............................................................................... 367
16.4.4 An Example................................................................................................... 368
16.5 Concluding Remarks................................................................................................. 370
Appendix............................................................................................................................... 371

17. Sample Size for Clinical Trials with Extremely Low Incidence Rate....................... 373
17.1 Clinical Studies with Extremely Low Incidence Rate.......................................... 374
17.2 Classical Methods for Sample Size Determination............................................... 374
17.2.1 Power Analysis.............................................................................................. 374
17.2.2 Precision Analysis........................................................................................ 375
17.2.3 Remarks.......................................................................................................... 376
xiv Contents

17.3 Chow and Chiu’s Procedure for Sample Size Estimation.................................... 378
17.3.1 Basic Idea of Chow and Chiu’s Procedure................................................ 378
17.3.2 Sensitivity Analysis...................................................................................... 379
17.3.3 An Example................................................................................................... 380
17.4 Data Safety Monitoring Procedure.......................................................................... 380
17.5 Concluding Remarks................................................................................................. 383

18. Sample Size Calculation for Two-Stage Adaptive Trial Design................................ 389
18.1 Types of Two-Stage Adaptive Designs.................................................................... 390
18.2 Analysis and Sample Size for Category SS Adaptive Designs............................ 391
18.2.1 Theoretical Framework................................................................................ 392
18.2.2 Two-Stage Design......................................................................................... 394
18.2.3 Conditional Power........................................................................................ 397
18.3 Analysis and Sample Size for Category II SD Adaptive Designs....................... 398
18.3.1 Continuous Endpoints................................................................................. 398
18.3.2 Binary Responses.......................................................................................... 402
18.3.3 Time-to-Event Endpoints.............................................................................405
18.4 Analysis and Sample Size for Category III DS and IV DD Two-Stage
Adaptive Designs....................................................................................................... 414
18.4.1 Nonadaptive Version.................................................................................... 415
18.4.2 Adaptive Version........................................................................................... 416
18.4.3 A Case Study of Hepatitis C Virus Infection............................................ 417
18.5 Concluding Remarks................................................................................................. 419

19. Simulation-Based Sample Size and Power Analysis................................................... 421


19.1 Example: Survival Study with Nonconstant Treatment Effect...........................422
19.2 Example: Cluster Randomized Study with Stepped Wedge Design.................. 424

20. Sample Size Calculation in Other Areas........................................................................ 427


20.1 QT/QTc Studies with Time-Dependent Replicates............................................... 427
20.1.1 Study Designs and Models.......................................................................... 428
20.1.2 Power and Sample Size Calculation........................................................... 429
20.1.3 Extension........................................................................................................ 433
20.1.4 Remarks..........................................................................................................434
20.2 Propensity Analysis in Nonrandomized Studies................................................. 435
20.2.1 Weighted Mantel–Haenszel Test................................................................ 435
20.2.2 Power and Sample Size................................................................................ 436
20.2.3 Simulations.................................................................................................... 439
20.2.4 Concluding Remarks....................................................................................440
20.3 ANOVA with Repeated Measures...........................................................................440
20.3.1 Statistical Model............................................................................................440
20.3.2 Hypotheses Testing......................................................................................442
20.3.3 Sample Size Calculation...............................................................................443
20.3.4 An Example...................................................................................................443
20.4 Quality of Life............................................................................................................445
20.4.1 Time Series Model........................................................................................446
20.4.2 Sample Size Calculation...............................................................................448
20.4.3 An Example...................................................................................................448
Contents xv

20.5 Bridging Studies......................................................................................................... 449


20.5.1 Sensitivity Index........................................................................................... 449
20.5.2 Assessment of Similarity............................................................................. 452
20.5.3 Remarks.......................................................................................................... 457
20.6 Vaccine Clinical Trials............................................................................................... 457
20.6.1 Reduction in Disease Incidence.................................................................. 458
20.6.2 Evaluation of Vaccine Efficacy with Extremely Low
Disease Incidence.......................................................................................... 459
20.6.3 Relative Vaccine Efficacy.............................................................................. 461
20.6.4 Composite Efficacy Measure....................................................................... 461
20.6.5 Remarks..........................................................................................................463

Bibliography................................................................................................................................. 465
Index.............................................................................................................................................. 481
Preface

This third edition is expanded to 20 chapters with five completely new chapters and
numerous updates and new sections of the existing chapters. These new chapters include
cluster randomized trial design (Chapter 15), zero-inflated Poisson distribution (Chapter
16), sample size estimation based on clinical trial simulation (Chapter 17), two-stage seam-
less adaptive trial design (Chapter 18), and sample size estimation based on clinical trial
simulation (Chapter 19). New sections in the existing chapters include negative binomial
regression in Chapter 5 and sample size requirement for analytical similarity assessment
in Chapter 10. Numerous updates have been included in several chapters especially the
introduction chapter. Similar to the first two editions, this third edition also concentrates
on concepts and implementation of methodology rather than technical details. We keep
the mathematics and statistics covered in this edition as fundamental as possible and illus-
trate the concepts and implementation through real working examples whenever possible.
The first and second editions of this book have been well received by pharmaceuti-
cal and clinical scientists/researchers and biostatisticians. The first two editions are
widely used as a reference source and a graduate textbook in pharmaceutical and clinical
research and development. Following the tradition of the first two editions, the purpose of
this third edition is also to provide a comprehensive and unified presentation of the prin-
ciples and methodologies for sample size calculation under various designs and hypoth-
eses for various clinical trials across different therapeutic areas. This revision is to give a
well-balanced summary of current regulatory requirements and recently developed sta-
tistical methods in this area. It is our continuing goal to provide a complete, comprehen-
sive, and updated reference and textbook in the area of sample size calculation in clinical
research. In the past decade, tremendous progress has been made in statistical methodol-
ogy to utilize innovative design and analysis in pharmaceutical and clinical research to
improve the probability of success in pharmaceutical and clinical development. These
methods include the use of adaptive design methods in clinical trials and others such as
negative binomial regression for recurrent incidence rates, zero-inflated Poisson for the
number of lesions, randomized cluster randomized trials, clinical trials with extremely
low ­incidence rates, and sample size requirement for analytical similarity assessment for
biosimilar products.
We have received much positive feedback and constructive suggestions from scientists
and researchers in the biopharmaceutical industry, regulatory agencies such as the U.S.
FDA, and academia. This third edition provides a comprehensive coverage of the current
state-of-the-art methodology for sample size estimation of clinical trials. Therefore, we
strongly believe that this new and expanded third edition is not only an extremely useful
reference book for scientists and researchers, regulatory reviewers, clinicians, and biostat-
isticians in biopharmaceutical industry, academia, and regulatory agencies, but may also
serve as a textbook for graduate students in clinical trials and biostatistics-related courses.
In addition, this third edition can serve as a bridge among the biopharmaceutical industry,
regulatory agencies, and academia.
We would like to thank David Grubbs at Taylor & Francis for his outstanding admin-
istrative assistance and support. We also express our sincere appreciation to the many
clinicians, scientists, researchers, and biostatisticians from the FDA, academia, and

xvii
xviii Preface

pharmaceutical industry for their valuable feedback, support, and encouragement.


Finally, the views expressed are those of the authors and not necessarily those of Duke
University School of Medicine, Durham, North Carolina, USA; University of Wisconsin,
Madison, Wisconsin, USA; and Peking University, Beijing, China. We are solely respon-
sible for the contents and errors of this edition. Any comments and suggestions will be
very much appreciated.
1
Introduction

In clinical research, during the planning stage of a clinical study, the following questions
are of particular interest to the investigators: (i) How many subjects are needed to have
a desired power for detecting a clinically meaningful difference (e.g., an 80% chance of
correctly detecting a clinically meaningful difference)? (ii) What is the trade off between
cost-effectiveness and power if only a small number of subjects are available for the
study due to limited budget and/or certain medical considerations. To address these
questions, a statistical evaluation for sample size calculation is often performed based
on certain statistical inference (e.g., power or confidence interval) of the primary study
endpoint with certain assurance. In clinical research, sample size calculation plays an
important role for assuring validity, accuracy, reliability, and integrity of the intended
clinical study.
For a given study, sample size calculation is usually performed based on some statisti-
cal criteria controlling type I error (e.g., a desired confidence level) and/or type II error
(i.e., a desired power). For example, we may choose sample size in such a way that there
is a desired precision at a fixed confidence level (i.e., fixed type I error). This approach is
referred to as precision analysis for sample size calculation. The method of precision ­analysis
is simple and easy to perform and yet it may have a small chance of correctly detecting
a true difference. As an alternative, the method of prestudy power analysis is usually con-
ducted to estimate sample size. The concept of the prestudy power analysis is to select the
required sample size for achieving the desired power for detecting a clinically or scientifi-
cally meaningful difference at a fixed type I error rate. In clinical research, the prestudy
power analysis is probably the most commonly used method for sample size calculation.
In this book, we will focus on sample size calculation based on power analysis for various
situations in clinical research.
In clinical research, to provide an accurate and reliable sample size calculation, an
appropriate statistical test for the hypotheses of interest is necessarily derived under the
study design. The hypotheses should be established to reflect the study objectives under
the study design. In practice, it is not uncommon to observe discrepancies among study
­objectives (hypotheses), study design, statistical analysis (test statistic), and sample size
calculation. These discrepancies can certainly distort the validity and integrity of the
intended clinical trial.
In Section 1.1, regulatory requirement regarding the role of sample size calculation
in clinical research is discussed. In Section 1.2, we provide some basic considerations
for sample size calculation. These basic considerations include study objectives, design,
hypotheses, primary study endpoint, and clinically meaningful difference. The con-
cepts of type I and type II errors and procedures for sample size calculation based
on precision analysis, power analysis, probability assessment, and reproducibility
­probability are given in Section 1.3. The aim and structure of this book is given in
Section 1.4.2.

1
2 Sample Size Calculations in Clinical Research

1.1 Regulatory Requirement


As indicated in Chow and Liu (1998, 2003, 2013), the process of drug research and devel-
opment is a lengthy and costly process. This lengthy and costly process is necessary not
only to demonstrate the efficacy and safety of the drug product under investigation, but
also to ensure that the study drug product possesses good drug characteristics such as
­identity, strength, quality, purity, and stability after it is approved by the regulatory author-
ity. This lengthy process includes drug discovery, formulation, animal study, ­laboratory
development, clinical development, and regulatory submission. As a result, clinical
­
­development plays an important role in the process of drug research and development
because all of the tests are conducted on humans. For approval of a drug p ­ roduct under
investigation, the United States Food and Drug Administration (FDA) requires that at least
two adequate and well-controlled clinical studies be conducted for providing substan-
tial evidence regarding the efficacy and safety of the drug product (FDA, 1988). However,
the following scientific/statistical questions are raised: (i) What is the definition of an
­adequate and well-controlled clinical study? (ii) What evidence is considered ­substantial?
(iii) Why do we need at least two studies? (iv) Will a single large trial be ­sufficient to
­provide substantial evidence for approval? (v) If a single large trial can provide substantial
evidence for approval, how large is considered large? In what follows, we will address
these questions.

1.1.1 Adequate and Well-Controlled Clinical Trials


Section 314.126 of 21 CFR (Code of Federal Regulation) provides the definition of an
­adequate and well-controlled study, which is summarized in Table 1.1.
As can be seen from Table 1.1, an adequate and well-controlled study is judged by
eight characteristics specified in the CFR. These characteristics include study objectives,
methods of analysis, design, selection of subjects, assignment of subjects, participants of
studies, assessment of responses, and assessment of the effect. For study objectives, it is
required that the study objectives be clearly stated in the study protocol such that they
can be formulated into statistical hypotheses. Under the hypotheses, appropriate statisti-
cal methods should be described in the study protocol. A clinical study is not considered
adequate and well controlled if the employed study design is not valid. A valid study
design allows a quantitative assessment of drug effect with a valid comparison with a

TABLE 1.1
Characteristics of an Adequate and Well-Controlled Study
Criteria Characteristics
Objectives Clear statement of investigation’s purpose
Methods of analysis Summary of proposed or actual methods of analysis
Design Valid comparison with a control to provide a quantitative assessment of drug effect
Selection of subjects Adequate assurance of the disease or conditions under study
Assignment of subjects Minimization of bias and assurance of comparability of groups
Participants of studies Minimization of bias on the part of subjects, observers, and analysis
Assessment of responses Well defined and reliable
Assessment of the effect Requirement of appropriate statistical methods
Introduction 3

control. The selection of a sufficient number of subjects with the disease or conditions
under study is one of the keys to the integrity of an adequate and well-controlled study.
In an adequate and well-controlled clinical study, subjects should be randomly assigned
to treatment groups to minimize potential bias by ensuring comparability between treat-
ment groups with respect to demographic variables such as age, gender, race, height and
weight, and other patient characteristics or prognostic factors such as medical history
and disease severity. An adequate and well-controlled study requires that the primary
study endpoint or response variable be well defined and assessed with a certain degree
of accuracy and reliability. To achieve this goal, statistical inferences on the drug effect
should be obtained based on the responses of the primary study endpoint observed from
the ­sufficient ­number of subjects using appropriate statistical methods derived under the
study design and objectives.

1.1.2 Substantial Evidence


The substantial evidence as required in the Kefauver–Harris amendments to the Food
and Drug and Cosmetics Act in 1962 is defined as the evidence consisting of adequate and
well-controlled investigations, including clinical investigations, by experts qualified
by ­scientific training and experience to evaluate the effectiveness of the drug involved,
on the basis of which it could fairly and responsibly be concluded by such experts that
the drug will have the effect it purports to have under the conditions of use prescribed,
­recommended, or suggested in the labeling or proposed labeling thereof. Based on this
amendment, the FDA requests that reports of adequate and well-controlled investiga-
tions provide the ­primary basis for determining whether there is substantial evidence to
­support the claims of new drugs and antibiotics.

1.1.3 Why At Least Two Studies?


As indicated earlier, the FDA requires that at least two adequate and well-controlled clini-
cal trials be conducted for providing substantial evidence regarding the effectiveness
and safety of the test drug under investigation for regulatory review and approval. In
practice, it is prudent to plan for more than one trial in the phase III study because of
any or combination of the following reasons: (i) lack of pharmacological rationale, (ii) a
new pharmacological principle, (iii) phase I and phase II data are limited or unconvinc-
ing, (iv) a therapeutic area with a history of failed studies or failures to confirm seem-
ingly convincing results, (v) a need to demonstrate efficacy and/or tolerability in different
subpopulations, with different comedication or other interventions, relative to different
competitors, and (vi) any other needs to address additional questions in the phase III
program.
Shao and Chow (2002) and Chow, Shao, and Hu (2002) pointed out that the purpose
of requiring at least two clinical studies is not only to assure the reproducibility but also
to provide valuable information regarding generalizability. Reproducibility is referred to
as whether the clinical results are reproducible from location (e.g., study site) to location
within the same region or from region to region, while generalizability is referred to as
whether the clinical results can be generalized to other similar patient populations within
the same region or from region to region. When the sponsor of a newly developed or
approved drug product is interested in getting the drug product into the marketplace from
one region (e.g., where the drug product is developed and approved) to another region,
it is a concern that differences in ethnic factors could alter the efficacy and safety of the
4 Sample Size Calculations in Clinical Research

drug product in the new region. As a result, it is recommended that a bridging study be
conducted to generate a limited amount of clinical data in the new region to extrapolate
the clinical data between the two regions (ICH, 1998a).
In practice, it is often of interest to determine whether a clinical trial that produced posi-
tive clinical results provides substantial evidence to assure reproducibility and generaliz-
ability of the clinical results. In this chapter, the reproducibility of a positive clinical result
is studied by evaluating the probability of observing a positive result in a future clinical
study with the same study protocol, given that a positive clinical result has been observed.
The generalizability of clinical results observed from a clinical trial will be evaluated by
means of a sensitivity analysis with respect to changes in mean and standard deviation of
the primary clinical endpoints of the study.

1.1.4 Substantial Evidence with a Single Trial


Although the FDA requires that at least two adequate and well-controlled clinical trials
be conducted for providing substantial evidence regarding the effectiveness of the drug
product under investigation, a single trial may be accepted for regulatory approval under
certain circumstances. In 1997, FDA published the Modernization Act (FDAMA), which
includes a provision (Section 115 of FDAMA) to allow data from one adequate and well-
controlled clinical trial investigation and confirmatory evidence to establish effectiveness
for risk/benefit assessment of drug and biological candidates for approval under c­ ertain
circumstances. This provision essentially codified an FDA policy that had existed for
­several years but whose application had been limited to some biological products approved
by the Center for Biologic Evaluation and Research (CBER) of the FDA and a few pharma-
ceuticals, especially orphan drugs such as zidovudine and lamotrigine. As can be seen
from Table 1.2, a relatively strong significant result observed from a single clinical trial
(say, p-value is less than 0.001) would have about 90% chance of reproducing the result in
future clinical trials.
Consequently, a single clinical trial is sufficient to provide substantial evidence for
demonstration of efficacy and safety of the medication under study. However, in 1998,
FDA published a guidance that shed light on this approach despite the FDA having
recognized that advances in sciences and practice of drug development may permit
an expanded role for the single controlled trial in contemporary clinical development
(FDA, 1998).

TABLE 1.2
Estimated Reproducibility Probability Based on
Results from a Single Trial
t-Statistic p-Value Reproducibility
1.96 0.050 0.500
2.05 0.040 0.536
2.17 0.030 0.583
2.33 0.020 0.644
2.58 0.010 0.732
2.81 0.005 0.802
3.30 0.001 0.901
Introduction 5

1.1.5 Sample Size


As the primary objective of most clinical trials is to demonstrate the effectiveness and
safety of drug products under investigation, sample size calculation plays an impor-
tant role in the planning stage to ensure that there are sufficient subjects for providing
­accurate and reliable assessment of the drug products with certain statistical assurance.
In practice, hypotheses regarding medical or scientific questions of the study drug are
usually formulated based on the primary study objectives. The hypotheses are then evalu-
ated using appropriate statistical tests under a valid study design to ensure that the test
results are accurate and reliable with certain statistical assurance. It should be noted that
a valid s­ ample size calculation can only be done based on appropriate statistical tests for
the hypotheses that can reflect the study objectives under a valid study design. It is then
suggested that the hypotheses be clearly stated when performing a sample size calcula-
tion. Each of the above hypotheses has different requirements for sample size to achieve a
desired statistical assurance (e.g., 80% power or 95% assurance in precision).
Basically, sample size calculation can be classified into sample size estimation/
determination, sample size justification, sample size adjustment, and sample size reesti-
mation. Sample size estimation/determination is referred to the calculation of the required
sample size for achieving some desired statistical assurance of accuracy and reliability
such as 80% power, while sample size justification is to provide statistical justification for
a selected sample size, which is often a small number due to budget constraints and/or
some medical considerations. In most clinical trials, sample size is necessarily adjusted
for some factors such as dropouts or covariates to yield sufficient number of evaluable
subjects for a valid statistical assessment of the study medicine. This type of sample size
calculation is known as sample size adjustment. In many clinical trials, it may be desir-
able to conduct interim analyses (planned or unplanned) during the conduct of the trial.
For clinical trials with planned or unplanned interim analyses, it is suggested that sample
size be adjusted for controlling an overall type I error rate at the nominal significance
level (e.g., 5%). In addition, when performing interim analyses, it is also desirable to per-
form sample size reestimation based on cumulative information observed up to a specific
time point to determine whether the selected sample size is sufficient to achieve a desired
power at the end of the study. Sample size reestimation may be performed in a blinded or
unblinded fashion depending upon whether the process of sample size reestimation will
introduce bias to clinical evaluation of subjects beyond the time point at which the interim
analysis or sample size reestimation is performed. In this book, however, our emphasis
will be placed on sample size estimation/determination. The concept can be easily applied
to (i) sample size justification for a selected sample size, (ii) sample size adjustment with
respect to some factors such as dropouts or covariates, and (iii) sample size reestimation in
clinical trials with planned or unplanned interim analyses.

1.2 Basic Considerations


In clinical research, sample size calculation may be performed based on precision analysis,
power analysis, probability assessment, or other statistical inferences. To provide an accu-
rate and reliable sample size calculation, it is suggested that an appropriate statistical test
for the hypotheses of interest be derived under the study design. The hypotheses should be
6 Sample Size Calculations in Clinical Research

established to reflect the study objectives and should be able to address statistical/­medical
questions of interest under the study design. As a result, a typical procedure for sample
size calculation is to determine or estimate sample size based on an appropriate statistical
method or test, which is derived under the hypotheses and the study design, for testing the
hypotheses to achieve a certain degree of statistical inference (e.g., 95% assurance or 80%
power) on the effect of the test drug under investigation. As indicated earlier, in ­practice,
it is not uncommon to observe discrepancies among study objectives (hypotheses), study
design, statistical analysis (test statistic), and sample size calculation. These discrepancies
certainly have an impact on sample size calculation in clinical research. Therefore, it is
suggested that the following be carefully considered when performing sample size cal-
culation: (i) the study objectives or the hypotheses of interest be clearly stated, (ii) a valid
design with appropriate statistical tests be used, (iii) sample size be determined based on
the test for the hypotheses of interest, and (iv) sample size be determined based on the pri-
mary study endpoint, and (v) sample size be determined based on the clinically meaning-
ful difference of the primary study endpoint that the clinical study is intended to detect.

1.2.1 Study Objectives


In clinical research, it is important to clearly state the study objectives of the intended
­clinical trials. The objectives of clinical studies may include one or more of the ­following
four objectives: (i) demonstrate/confirm efficacy, (ii) establish a safety profile, (iii) pro-
vide an adequate basis for assessing the benefit/risk relationship to support labeling,
and (iv) establish the dose–response relationship (ICH, 1998b). Since most clinical studies
are c­ onducted for the clinical evaluation of efficacy and safety of drug products under
­investigation, it is suggested that the following study objectives related to efficacy and
safety be clarified before choosing an appropriate design strategy for the intended trial.

Safety
Equivalence Noninferiority Superiority
Equivalence E/E E/N E/S
Efficacy Noninferiority N/E N/N N/S
Superiority S/E S/N S/S

For example, if the intent of the planned clinical study is to develop an alternative ther-
apy to the standard therapy that is quite toxic, then we may consider the strategy of E/S,
which is to show that the test drug has equal efficacy but less toxicity (superior safety).
The study objectives will certainly have an impact on the sample size calculation. Sample
size calculation provides the required sample size for achieving the study objectives.

1.2.2 Study Design


In clinical trials, different designs may be employed to achieve the study objectives.
A valid study design is necessarily chosen to collect relevant clinical information for
achieving the study objectives by addressing some statistical/medical hypotheses of
­interest, which are formulated to reflect the study objectives.
In clinical research, commonly employed study designs include parallel-group design,
crossover design, enrichment design, and titration design (see, e.g., Chow and Liu,
1998). The design strategy can certainly affect sample size calculation because statistical
Introduction 7

methods or tests are usually derived under the hypotheses and study design. As an exam-
ple, Fleming (1990) discussed the following design strategies that are commonly used in
­clinical therapeutic equivalence/noninferiority and superiority trials.

Design Description
Classical STD + TEST versus STD
Active control TEST versus STD
Dual purpose TEST versus STD versus STD + TEST

The classical design is to compare the combination of a test drug (TEST) and a standard
therapy (STD) (i.e., STD + TEST) against STD to determine whether STD + TEST yields
superior efficacy. When the intent is to determine whether a test drug could be used as
an alternative to a standard therapy, one may consider an active control design involving
direct randomization to either TEST or STD. This occurs frequently when STD is quite toxic
and the intent is to develop an alternative therapy that is less toxic, yet equally ­efficacious.
To achieve both objectives, a dual-purpose design strategy is useful.
Note that in practice, a more complicated study design, which may consist of a
­combination of the above designs, may be chosen to address more complicated statistical/­
medical questions regarding the study drug. In this case, standard procedure for sample
size ­calculation may not be directly applicable and a modification will be necessary.

1.2.3 Hypotheses
In most clinical trials, the primary study objective is usually related to the evaluation of
the effectiveness and safety of a drug product. For example, it may be of interest to show
that the study drug is effective and safe as compared to a placebo for some intended indi-
cations. In some cases, it may be of interest to show that the study drug is as effective as,
superior to, or equivalent to an active control agent or a standard therapy. In practice,
hypotheses regarding medical or scientific questions of the study drug are usually for-
mulated based on the primary study objectives. The hypotheses are then evaluated using
appropriate statistical tests under a valid study design.
In clinical trials, a hypothesis is usually referred to as a postulation, assumption, or
statement that is made about the population regarding the effectiveness and safety of a
drug under investigation. For example, the statement that there is a direct drug effect
is a hypothesis regarding the treatment effect. For testing the hypotheses of interest, a
random sample is usually drawn from the targeted population to evaluate hypotheses
about the drug product. A statistical test is then performed to determine whether the null
hypothesis would be rejected at a prespecified significance level. Based on the test result,
conclusion(s) regarding the hypotheses can be drawn. The selection of the hypothesis
depends upon the study objectives. In clinical research, commonly considered hypotheses
include point hypotheses for testing equality and interval hypothesis for testing equiva-
lence/noninferiority and superiority, which are described below. A typical approach for
demonstration of the efficacy and safety of a test drug under investigation is to test the
following hypotheses.

[Link] Test for Equality


H 0 : µT = µP versus H a : µT ≠ µP , (1.1)
8 Sample Size Calculations in Clinical Research

where µT and µP are the mean response of the outcome variable for the test drug and
the placebo, respectively. We first show that there is a statistically significant difference
between the test drug and the placebo by rejecting the null hypothesis, and then demon-
strate that there is a high chance of correctly detecting a clinically meaningful difference
if such difference truly exists.

[Link] Test for Noninferiority


In clinical trials, one may wish to show that the test drug is as effective as an active agent
or a standard therapy. In this case, Blackwelder (1982) suggested testing the following
hypotheses:

H 0 : µS − µT ≥ δ versus H a : µS − µT < δ , (1.2)

where µS is the mean for a standard therapy and δ is a difference of clinical importance.
The concept is to reject the null hypothesis and conclude that the difference between
the test drug and the standard therapy is less than a clinically meaningful difference δ
and hence the test drug is as effective as the standard therapy. This study objective is not
uncommon in clinical trials, especially when the test drug is considered to be less toxic,
easier to administer, or less expensive than the established standard therapy.

[Link] Test for Superiority


To show superiority of a test drug over an active control agent or a standard therapy, we
may consider the following hypotheses:

H 0 : µT − µS ≤ δ versus H a : µT − µS > δ. (1.3)

The rejection of the above null hypothesis suggests that the difference between the test
drug and the standard therapy is greater than a clinically meaningful difference. Therefore,
we may conclude that the test drug is superior to the standard therapy by rejecting the null
hypothesis of Equation 1.3. Note that the above hypotheses are also known as hypotheses
for testing clinical superiority. When δ = 0, the above hypotheses are usually referred to as
hypotheses for testing statistical superiority.

[Link] Test for Equivalence


In practice, unless there is some prior knowledge regarding the test drug, usually we do
not know the performance of the test drug as compared to the standard therapy. Therefore,
hypotheses (1.2) and (1.3) are not preferred because they have predetermined the perfor-
mance of the test drug as compared to the standard therapy. As an alternative, the follow-
ing hypotheses for therapeutic equivalence are usually considered:

H 0 :|µT − µS |≥ δ versus H a : µT − µS < δ. (1.4)

We then conclude that the difference between the test drug and the standard therapy is
of no clinical importance if the null hypothesis of Equation 1.4 is rejected.
Introduction 9

µs – M µs µs + M

Inferiority Equivalence Superiority

Noninferiority
Nonsuperiority
One-sided equivalence

FIGURE 1.1
Relationship among noninferiority, superiority, and equivalence.

[Link] Relationship among Noninferiority, Superiority, and Equivalence


To study the relationship among noninferiority, equivalence, and superiority, we first
assume that the noninferiority margin, equivalence limit, and superiority margin are
the same. Let M denote the noninferiority margin (also equivalence limit and superior-
ity m ­ argin). Also, let µT and µS be the mean responses of the test treatment and standard
therapy (active control agent), respectively. If we assume that an observed mean response
on the right-hand side of µS is an indication of improvement, then the relationship among
noninferiority, equivalence, and superiority is illustrated in Figure 1.1.
It should be noted that a valid sample size calculation can only be done based on
­appropriate statistical tests for the hypotheses that can reflect the study objectives under
a valid study design. It is then suggested that the hypotheses be clearly stated when
­performing a sample size calculation. Each of the above hypotheses has different require-
ment for sample size to achieve a desired power or precision of the corresponding tests.

1.2.4 Primary Study Endpoint


A study objective (hypotheses) will define what study variable is to be considered as
the primary clinical endpoint and what comparisons or investigations are deemed most
clinically relevant. The primary clinical endpoints depend upon therapeutic areas and
the indications that the test drugs sought for. For example, for coronary artery disease/
angina in cardiovascular system, patient mortality is the most important clinical endpoint
in clinical trials assessing the beneficial effects of drugs on coronary artery disease. For
congestive heart failure, patient mortality, exercise tolerance, the number of hospitaliza-
tions, and cardiovascular morbidity are common endpoints in trials assessing the effects
of drugs in congestive heart failure, while mean change from baseline in systolic and dia-
stolic blood pressure and cardiovascular mortality and morbidity are commonly used in
hypertension trials. Other examples include change in forced expiratory volume in one
second (FEV1) for asthma in respiratory system, cognitive and functional scales specially
designed to assess Alzheimer’s disease and Parkinson’s disease in central nervous system,
tender joints and pain-function endpoints (e.g., Western Ontario and McMaster University
Osteoarithritis Index) for osteoarthritis in musculoskeletal system, and the incidence of
bone fracture for osteoporosis in the endocrine system.
It can be seen that the efficacy of a test drug in the treatment of a certain disease may
be characterized through multiple clinical endpoints. Capizzi and Zhang (1996) clas-
sify the clinical endpoints into primary, secondary, and tertiary endpoints. Endpoints
that satisfy the following criteria are considered primary endpoints: (i) should be of
10 Sample Size Calculations in Clinical Research

biological and/or clinical importance, (ii) should form the basis of the objectives of the
trial, (iii) should not be highly correlated, (iv) should have sufficient power for the statis-
tical hypotheses formulated from the objectives of the trial, and (v) should be relatively
few (e.g., at most 4). Sample size calculation based on detecting a difference in some or
all primary clinical endpoints may result in a high chance of false-positive and false-
negative results for the evaluation of the test drug. Thus, it is suggested that sample
size calculation should be performed based on a single primary study endpoint under
certain assumption of the single primary endpoint. More discussion regarding the issue
of false-positive and false-negative rates caused by multiple primary endpoints can be
found in Chow and Liu (1998).

1.2.5 Clinically Meaningful Difference


In clinical research, the determination of a clinically meaningful difference, denoted by δ,
is critical in clinical trials such as equivalence/noninferiority trials. In therapeutic equiva-
lence trials, δ is known as the equivalence limit, while δ is referred to as the noninferiority
margin in noninferiority trials. The noninferiority margin reflects the degree of inferiority
of the test drug under investigation as compared to the standard therapy that the trials
attempt to exclude.
A different choice of δ may affect the sample size calculation and may alter the conclu-
sion of clinical results. Thus, the choice of δ is critical at the planning stage of a ­clinical
study. In practice, there is no golden rule for the determination of δ in clinical trials.
As indicated in the ICH E10 Draft Guideline, the noninferiority margin cannot be chosen
to be greater than the smallest effect size that the active drug would be reliably expected to
have ­compared with placebo in the setting of the planned trial, but may be smaller based
on clinical judgment (ICH, 1999). The ICH E10 Guideline suggests that the noninferiority
margin be identified based on past experience in placebo-controlled trials of adequate
design under conditions similar to those planned for the new trial. In addition, the ICH
E10 Guideline emphasizes that the determination of δ should be based on both statisti-
cal reasoning and clinical judgment, which should not only reflect uncertainties in the
­evidence on which the choice is based, but also be suitably conservative.
In some cases, regulatory agencies do provide clear guidelines for the selection of an
appropriate δ for clinical trials. For example, as indicated by Huque and Dubey (1990), the
FDA proposed some noninferiority margins for some clinical endpoints (binary responses)
such as cure rate for anti-infective drug products (e.g., topical antifungals or vaginal
­antifungals). These limits are given in Table 1.3. For example, if the cure rate is between
80% and 90%, it is suggested that the noninferiority margin or a clinically ­meaningful dif-
ference be chosen as δ = 15%.

TABLE 1.3
Noninferiority Margins for Binary Responses
δ (%) Response Rate for the Active Control (%)
20 50–80
15 80–90
10 90–95
5 >95
Source: FDA Anti-Infectives Drug Guideline (FDA, 1977).
Introduction 11

On the other hand, for bioequivalence trials with healthy volunteers, the margin of
δ = log(1.25) for mean difference on log-transformed data such as area under the blood
or plasma concentration­–time curve (AUC) or maximum concentration Cmax is considered
(FDA, 2001).
Most recently, FDA published a draft guidance on Non-Inferiority Clinical Trials, which
recommends that two noninferiority margins, namely, M1 and M2, should be considered
(FDA, 2010b). The 2010 FDA draft guidance indicated that M1 is based on (i) the treatment
effect estimated from the historical experience with the active control drug, (ii) assessment
of the likelihood that the current effect of the active control is similar to the past effect
(the constancy assumption), and (iii) assessment of the quality of the noninferiority trial,
particularly looking for defects that could reduce a difference between the active control
and the new drug. Thus, M1 is defined as the entire effect of the active control assumed
to be present in the noninferiority study. On the other hand, FDA indicates that M2 is
selected based on a clinical judgment, which is never greater than M1 even if for active
control drugs with small effects. It should be noted that a clinical judgment might argue
that a larger difference is not clinically important. Ruling out that a difference between
the active control and test treatment that is larger than M1 is a critical finding that supports
the conclusion of effectiveness.
In clinical trials, the choice of δ may depend upon absolute change, percent change, or
effect size of the primary study endpoint. In practice, a standard effect size (i.e., effect size
adjusted for standard deviation) between 0.25 and 0.5 is usually chosen as δ, if no prior
knowledge regarding clinical performance of the test drug is available. This recommenda-
tion is made based on the fact that the standard effect size of clinical importance observed
from most clinical trials is within the range of 0.25 and 0.5.

1.3 Procedures for Sample Size Calculation


In practice, sample size may be determined based on either precision analysis or power
analysis. Precision analysis and power analysis for sample size determination are usually
performed by controlling type I error (or confidence level) and type II error (or power),
respectively. In what follows, we will first introduce the concepts of type I and type II
errors.

1.3.1 Type I and Type II Errors


In practice, two kinds of errors occur when testing hypotheses. If the null hypothesis is
rejected when it is true, then a type I error has occurred. If the null hypothesis is not
rejected when it is false, then a type II error has been made. The probabilities of making
type I and type II errors, denoted by α and β, respectively, are given below:

α = P{type I error}
= P{reject H 0 when H 0 is true},
β = P{type II error}
= P{fail to reject H 0 when H 0 is false}.
12 Sample Size Calculations in Clinical Research

An upper bound for α is a significance level of the test procedure. Power of the test is
defined as the probability of correctly rejecting the null hypothesis when the null hypoth-
esis is false, that is,

Power = 1 − β
= P{reject H 0 when H 0 is false}.

As an example, suppose one wishes to test the following hypotheses:

H 0 : The drug is ineffective versus H a : The drug is effective.

Then, a type I error occurs if we conclude that the drug is effective when in fact it is not.
On the other hand, a type II error occurs if we claim that the drug is ineffective when in
fact it is effective. In clinical trials, none of these errors is desirable. With a fixed sample
size, a typical approach is to avoid a type I error but at the same time to decrease a type
II error so that there is a high chance of correctly detecting a drug effect when the drug is
indeed effective. Typically, when the sample size is fixed, α decreases as β increases and
α increases as β decreases. The only approach to decrease both α and β is to increase the
sample size. Sample size is usually determined by controlling both type I error (or confi-
dence level) and type II error (or power).
In what follows, we will introduce the concepts of precision analysis and power analysis
for sample size determination based on type I error and type II error, respectively.

1.3.2 Precision Analysis


In practice, the maximum probability of committing a type I error that one can tolerate is
usually considered as the level of significance. The confidence level, 1 − α, then reflects
the probability or confidence of not rejecting the true null hypothesis. Since the confidence
interval approach is equivalent to the method of hypotheses testing, we may determine
sample size required based on type I error rate using the confidence interval approach.
For a (1 − α)100% confidence interval, the precision of the interval depends on its width.
The narrower the interval is, the more precise the inference. Therefore, the precision analy-
sis for sample size determination is to consider the maximum half width of the (1 − α)100%
confidence interval of the unknown parameter that one is willing to accept. Note that the
maximum half width of the confidence interval is usually referred to as the maximum error
of an estimate of the unknown parameter. For example, let y1, y2, …, yn be independent and
identically distributed normal random variables with mean µ and variance σ2. When σ2 is
known, a (1 − α)100% confidence interval for µ can be obtained as

σ
y ± zα/2 ,
n

where zα/ 2 is the upper (α/2)th quantile of the standard normal distribution. The maxi-
mum error, denoted by E, in estimating the value of µ that one is willing to accept is then
defined as

σ
E =| y − µ|= zα/2 .
n
Introduction 13

Thus, the sample size required can be chosen as

zα2 /2σ 2
n= . (1.5)
E2

Note that the maximum error approach for choosing n is to attain a specified precision
while estimating µ, which is derived based only on the interest of type I error. A nonpara-
metric approach can be obtained by using the following Chebyshev’s inequality

σ2
P{| y − µ|≤ E} ≥ 1 − ,
nE2

and hence

σ2
n= . (1.6)
αE2

Note that the precision analysis for sample size determination is very easy to apply either
based on Equation 1.5 or 1.6. For example, suppose we wish to have a 95% assurance that
the error in the estimated mean is less than 10% of the standard deviation (i.e., 0.1σ). Thus,

σ
zα/2 = 0.1σ.
n

Hence,

zα2 /2σ 2 (1.96)2 σ 2


n= = = 384.2 ≈ 385.
E2 (0.1σ )2

The above concept can be applied to binary data (or proportions). In addition, it can be
easily implemented for sample size determination when comparing two treatments.

1.3.3 Power Analysis


Since a type I error is usually considered to be a more important and/or serious error
that one would like to avoid, a typical approach in hypothesis testing is to control α at an
­acceptable level and try to minimize β by choosing an appropriate sample size. In other
words, the null hypothesis can be tested at predetermined level (or nominal level) of
­significance with a desired power. This concept for the determination of sample size is
usually referred to as power analysis for sample size determination.
For the determination of sample size based on power analysis, the investigator is required
to specify the following information. First of all, select a significance level at which the
chance of wrongly concluding that a difference exists when in fact there is no real differ-
ence (type I error) one is willing to tolerate. Typically, a 5% level of significance is chosen to
reflect a 95% confidence regarding the unknown parameter. Second, select a desired power
at which the chance of correctly detecting a difference when the difference one wishes to
achieve truly exists. A conventional choice of power is either 90% or 80%. Third, specify
14 Sample Size Calculations in Clinical Research

a clinically meaningful difference. In most clinical trials, the objective is to demonstrate


effectiveness and safety of a drug under study as compared to a placebo. Therefore, it is
important to specify what difference, in terms of the primary endpoint, is considered to
be of clinical or scientific importance. Denote such a difference by Δ. If the investigator
will settle for detecting only a large difference, then fewer subjects will be needed. If the
difference is relatively small, a larger study group (i.e., a larger number of subjects) will be
needed. Finally, the knowledge regarding the standard deviation (i.e., σ) of the primary end-
point considered in the study is also required for sample size determination. A very precise
method of measurement (i.e., a small σ) will permit detection of any given difference with a
much smaller sample size than would be required with a less precise measurement.
Suppose there are two group of observations, namely, xi, i = 1, …, n1 (treatment) and
yi, i = 1, …, n2 (control). Assume that xi and yi are independent and normally distributed
with means µ1 and µ2 and variances σ12 and σ 22 , respectively. Suppose the hypotheses of
interest are

H 0 : µ1 = µ2 versus H1 : µ1 ≠ µ2 .

For simplicity and illustration purpose, we assume (i) σ12 and σ 22 are known, which
may be estimated from pilot studies or historical data, and (ii) n1 = n2 = n. Under these
­assumptions, a Z-statistic can be used to test the mean difference. The Z-test is given by

x−y
Z= .
 σ12   σ22 
  +  
 n   n 

Under the null hypothesis of no treatment difference, Z is distributed as N(0, 1). Hence,
we reject the null hypothesis when

|Z|> zα/2 .

Under the alternative hypothesis that µ1 = µ2 + δ (without loss of generality, we assume


δ > 0), a clinically meaningful difference, Z is distributed as N(µ*, 1), where

δ
µ* = > 0.
 σ   σ22 
2
  +  
1
 n   n 

The corresponding power is then given by

P{|N (µ*, 1)|> zα/2 } ≈ P{N (µ*, 1) > zα/2 }


= P{N (0, 1) > zα/2 − µ*}.

To achieve the desired power of (1 − β)100%, we set

zα/2 − µ* = −zβ .
Introduction 15

This leads to

n=
(σ12 + σ22 )(zα/2 + zβ )2 . (1.7)
δ2

To apply the above formula for sample size calculation, consider a double-blind, pla-
cebo-controlled clinical trial. Suppose the objective of the study is to compare a test drug
with a control and the standard deviation for the treatment group is 1 (i.e., σ1 = 1) and the
standard deviation of the control group is 2 (i.e., σ2 = 2). Then, by choosing α = 5% and
β = 10%, we have

n=
(σ12 + σ22 )(zα/2 + zβ )2 = (12 + 22 )(1.96 + 1.28)2 ≈ 53.
δ2 12

Thus, a total of 106 subjects is required for achieving a 90% power for the detection of a
clinically meaningful difference of δ = 1 at the 5% level of significance.

1.3.4 Probability Assessment


In practice, sample size calculation based on power analysis for detecting a small differ-
ence in the incidence rate of rare events (e.g., 3 per 10,000) may not be appropriate. In this
case, a very large sample size is required to observe a single event, which is not practical.
In addition, small difference in the incidence rate (e.g., 2 per 10,000 versus 3 per 10,000)
may not be of practical/clinical interest. Alternatively, it may be of interest to justify the
sample size based on a probability statement, for example, there is a certain assurance (say
(1 − ε)100%) that the mean incidence rate of the treatment group is less than that of the
­control group with probability (1 − α)100%.
Suppose there are two groups of observations, namely, xi, i = 1, …, n (treatment) and yi,
i = 1, …, n (control). Assume that xi and yi are independent and identically distributed as
Bernoulli random variables with mean p1 and p2, that is, B(1, p1) and B(1, p2), respectively,
and that

P(x ≤ y |n) + P( x > y |n) = 1

for large n. Then

P(x < y |n) = p.

The hypotheses of interest are

H 0 : p ∉ (ε , 1) versus H1 : p ∈ (ε , 1),

for some ε > 0, where p = P(x = y|n) for some n. A test for the hypothesis that p ∈ (ε,1) is

φ( x , y) = I (x < y ).
16 Sample Size Calculations in Clinical Research

We then reject the null hypothesis if φ(x, y) = 1. Then, given p1 < p2, the power is given by

Power = P( x < y )
 
 
 x − y − ( p1 − p2 ) p2 − p1 
= P  < 
 p1(1 − p1 ) + p2 (1 − p2 ) p1(1 − p1 ) + p2 (1 − p2 ) 
 
 n n 
 
 
 p2 − p1 
≈ Φ  .
 p1(1 − p1 ) + p2 (1 − p2 ) 
 
 n 

Therefore, for a given power 1 − β, the sample size, n, can be estimated by letting

( p2 − p1 )
= zβ ,
p1(1 − p1 ) + p2 (1 − p2 )
n

which gives

zβ2 [ p1(1 − p1 ) + p2 (1 − p2 )]
n= .
( p2 − p1 )2

To illustrate the above procedure, consider a double-blind, active-control trial. The objec-
tive of this trial is to compare a test drug with a reference drug (active control). Suppose the
event rate of the reference drug is 0.075 and the event rate of the test drug is 0.030. Then,
with β = 10%, we have

1.28 2 (0.075×(1 − 0.075) + 0.030 ×(1 − 0.030))


n= ≈ 80.
(0.075 − 0.030)2

Thus, a total of 160 subjects is needed to achieve a 90% power for observing less accident
rate in test drug group.

1.3.5 Reproducibility Probability


As indicated, current regulation for approval of a test drug under investigation requires
at least two adequate and well-controlled clinical trials be conducted for proving substan-
tial evidence regarding the effectiveness and safety of the drug product. Shao and Chow
(2002) investigated the probability of reproducibility of the second trial and developed
an ­estimated power approach. As a result, sample size calculation of the second clini-
cal trial can be performed based on the concept of reproducibility probability. Suppose
these are two groups of observations obtained in the first trial, namely, x1i, i = 1, …, n1
Introduction 17

(treatment) and x2i, i = 1, …, n2 (control). Assume that x1i and x2i are independent and
normally d­ istributed with means µ1 and µ2 and common variances σ2, respectively. The
hypotheses of interest are

H 0 : µ1 = µ2 versus H1 : µ1 ≠ µ2 .

When σ2 is known, we reject H0 at the 5% level of significance if and only if |T| > tn−2,
where tn−2 is the (1 − α/2)th percentile of the t-distribution with n − 2 degrees of freedom,
n = n1 + n2,

x1 − x2
T= ,
(n1 − 1)s + (n2 − 1)s22
2
1 1 1
+
n−2 n1 n2

and xi and s12 are the sample means and variances calculated based on data from the ith
group, respectively. Thus, the power of T is given by

p(θ) = P(|T ( x)|> tn−2 )


(1.8)
= 1 − Tn−2 (tn−2 |θ) + Tn−2 (−tn−2 |θ),

where

µ1 − µ2
θ=
1 1
σ +
n1 n2

and Tn−2(⋅|θ) denotes the distribution function of the t-distribution with n − 2 degrees of
freedom and the noncentrality parameter θ. Let x be the observed data from the first trial
and T(x) be the value of T based on x. Replacing θ in the power in Equation 1.8 by its
­estimate T(x), the estimated power

Pˆ = p(T ( x)) = 1 − Tn−2 (Tn−2 |T ( x)) + Tn−2 (−tn−2 |T ( x)),

is defined by Shao and Chow (2002) as a reproducibility probability for the second trial.
Based on this concept, sample size calculation for the second trial can be obtained as

(T * /ΔT )2
n* = ,
1 1
+
4n1 4n2

where T* is the value obtained such that a desired reproducibility probability is attained
and Δ is given by

1+ ∈ /(µ1 − µ2 )
Δ= ,
C
18 Sample Size Calculations in Clinical Research

where ε and C reflect the population mean and variance changes in the second trial.
In other words, in the second trial, it is assumed that the population mean difference is
changed from µ1 − µ2 to µ1 − µ2 + ε and the population variance is changed from σ2 to
C2σ2, where C > 0.

1.3.6 Sample Size Reestimation without Unblinding


In clinical trials, it is desirable to perform a sample size reestimation based on clinical
data accumulated up to the time point. If the reestimated sample size is bigger than the
originally planned sample size, then it is necessary to increase the sample size to achieve
the desired power at the end of the trial. On the other hand, if the reestimated sample size
is smaller than the originally planned sample size, a sample size reduction is justifiable.
Basically, sample size reestimation involves either unblinding or without unblinding of
the treatment codes. In practice, it is undesirable to perform a sample size reestimation
with unblinding of the treatment codes as even the significance level will be adjusted
for potential statistical penalty for the unblinding. Thus, sample size reestimation with-
out unblinding the treatment codes has become very attractive. Shih (1993) and Shih and
Zhao (1997) proposed some procedures without unblinding for sample size reestimation
within interim data for double-blinded clinical trials with binary outcomes. The detailed
­procedure for sample size reestimation with unblinding will be given in Chapter 8.
In practice, it is suggested that a procedure for sample size reestimation be specified in
the study protocol and should be performed by an external statistician who is independent
of the project team. It is also recommended that a data monitoring committee (DMC) be
considered to maintain the scientific validity and integrity of the clinical trial when per-
forming sample size reestimation at the interim stage of the trial. More details regarding
sample size reestimation are provided in Chapter 8.

1.4 Aims and Structure of this Book


1.4.1 Aim of this Book
As indicated earlier, sample size calculation plays an important role in clinical research.
Sample size calculation is usually performed using an appropriate statistical test for the
hypotheses of interest to achieve a desired power for the detection of a clinically m
­ eaningful
difference. The hypotheses should be established to reflect the study objectives for clinical
investigation under the study design. In practice, however, it is not uncommon to observe
discrepancies among study objectives (or hypotheses), study design, statistical analysis
(or test statistic), and sample size calculation. These inconsistencies often result in (i) wrong
test for right hypotheses, (ii) right test for wrong hypotheses, (iii) wrong test for wrong
hypotheses, or (iv) right test for right hypotheses with insufficient power. Therefore, the
aim of this book is to provide a comprehensive and unified presentation of statistical con-
cepts and methods for sample size calculation in various situations in clinical research.
Moreover, this book will focus on the interactions between clinicians and biostatisti-
cians that often occur during various phases of clinical research and development. This
book is also intended to provide a well-balanced summarization of current and emerging
clinical issues and recently developed statistical methodologies in the area of sample size
Introduction 19

c­ alculation in clinical research. Although this book is written from the v


­ iewpoint of clini-
cal research and development, the principles and concepts presented in this book can also
be applied to a nonclinical setting.

1.4.2 Structure of this Book


It is our goal to provide a comprehensive reference book for clinical researchers,
­pharmaceutical scientists, clinical or medical research associates, clinical programmers
or data coordinators, and biostatisticians in the areas of clinical research and develop-
ment, regulatory agencies, and academia. The scope of this book covers sample size cal-
culation for studies that may be conducted during various phases of clinical research and
­development. Basically, this book consists of 20 chapters, which are outlined below.
Chapter 1 provides a brief introduction and a review of regulatory requirement regard-
ing sample size calculation in clinical research for drug development. Also included in
this chapter are statistical procedures for sample size calculation based on precision analy-
sis, power analysis, probability assessment, and reproducibility probability. Chapter 2 cov-
ers some statistical considerations such as the concept of confounding and interaction, a
one-sided test versus or a two-sided test in clinical research, a crossover design versus a
parallel design, subgroup/interim analysis, and data transformation. Also included in this
chapter are unequal treatment allocation, adjustment for dropouts or covariates, the effect
of mixed-up treatment codes, treatment or center imbalance, multiplicity, multiple-stage
design for early stopping, and sample size calculation based on rare incidence rate.
Chapter 3 focuses on sample size calculation for comparing means with one sample, two
samples, and multiple samples. Formulas are derived under different hypotheses testing
for equality, superiority, noninferiority, and equivalence with equal or unequal treatment
allocation. In addition, sample size calculation based on Bayesian approach is also consid-
ered in this chapter. Chapter 4 deals with sample size calculation for comparing propor-
tions based on large sample tests. Formulas for sample size calculation are derived under
different hypotheses testing for equality, superiority, noninferiority, and equivalence with
equal or unequal treatment allocation. In addition, issues in sample size calculation based
on the confidence interval of the relative risk and/or odds ratio between treatments are
also examined.
Chapter 5 considers sample size calculation for binary responses based on exact tests
such as the binomial test and Fisher’s exact test. Also included in this chapter are optimal
and flexible multiple-stage designs that are commonly employed in phase II cancer trials.
The emphasis of Chapter 6 is placed on tests for contingency tables such as the goodness-
of-fit test and test for independence. Procedures for sample size calculation are derived
under different hypotheses for testing equality, superiority, noninferiority, and equiva-
lence with equal or unequal treatment allocation.
Chapter 7 provides sample size calculation for comparing time-to-event data using Cox’s
proportional hazards model and weighted log-rank test. Formulas are derived under dif-
ferent hypotheses testing for equality, superiority, noninferiority, and equivalence with
equal or unequal treatment allocation. Chapter 8 considers the problems of sample size
estimation and reestimation in group sequential trials with various alpha spending func-
tions. Also included in this chapter are the study of conditional power for the assessment
of futility and a proposed procedure for sample size reestimation without unblinding.
Chapter 9 discusses statistical methods and the corresponding sample size calcula-
tion for comparing intrasubject variabilities, intrasubject coefficient of variations (CV),
intersubject variabilities, and total variabilities under replicated crossover designs and
20 Sample Size Calculations in Clinical Research

­ arallel-group designs with replicates. Chapter 10 summarizes sample size calculation


p
for the assessment of population bioequivalence, individual bioequivalence, and in vitro
bioequivalence under replicated crossover designs as suggested in the FDA 2001 guidance
(FDA, 2001).
Chapter 11 summarizes sample size calculation for dose-ranging studies, includ-
ing the determination of minimum effective dose (MED) and maximum tolerable dose
(MTD). Chapter 12 considers sample size calculation for microarray studies controlling
false ­discovery rate (FDR) and family-wise error rate (FWER). Bayesian sample size cal-
culation is discussed in Chapter 13. Sample size calculation based on nonparametrics for
­comparing means with one or two samples is discussed in Chapter 14. Chapter 15 dis-
cusses statistical methods for data analysis under a cluster randomized trial. Sample size
requirements for analysis at the cluster level and at individual subject level (within the
cluster) are examined.
Sample size requirements for comparing clinical data (e.g., number of lesions) that ­follow
zero-inflated Poisson distribution are studied in Chapter 16. Also included in this chapter are
formulas for sample size calculation for testing noninferiority/superiority and equivalence.
Chapter 17 provides a compromised approach (compromise between precision analysis and
power analysis in conjunction with probability assessment) for sample size calculation for
clinical studies with extremely low incidence rate. Chapter 18 provides a comprehensive
summarization of sample size requirements for two-stage seamless adaptive trial designs
with/without different study objectives and study endpoints at ­different stages.
Chapter 19 provides an alternative method for sample size calculation using clinical
trial simulation when there exists no closed form for test statistics and/or formulas or
procedures for sample size calculation for complex clinical studies. Chapter 20 includes
sample size calculations in other areas of clinical research such as QT/QTc studies, where
the QT interval is a measure of the time between the start of the Q wave and the end of
T wave in the heart’s electrical cycle and QTc is corrected QT interval in cardiology, the
use of propensity score analysis in nonrandomized or observational studies, analysis of
variance with repeated measurements, quality-of-life assessment, bridging studies, and
vaccine clinical trials.
For each chapter, whenever possible, real examples concerning clinical studies of
­various therapeutic areas are included to demonstrate the clinical and statistical concepts,
­interpretations, and their relationships and interactions. Comparisons regarding the rela-
tive merits and disadvantages of statistical methods for sample size calculation in various
therapeutic areas are discussed whenever deemed appropriate. In addition, if applicable,
topics for future research development are provided.
2
Considerations Prior to Sample Size Calculation

As indicated in Chapter 1, sample size calculation should be performed using appropriate sta-
tistical methods or tests for hypotheses, which can reflect the study objectives under the study
design based on the primary study endpoint of the intended trial. As a result, some information
including study design, hypotheses, mean response and the associated variability of the pri-
mary study endpoint, and the desired power at a specified α level of significance are required
when performing sample size calculation. For good statistics practice, some statistical consid-
erations such as stratification with respect to possible confounding/interaction factors, the use
of a one-sided test or a two-sided test, the choice of a parallel design or a crossover design,
subgroup/interim analyses, and data transformation are important for performing an accu-
rate and reliable sample size calculation. In addition, some practical issues that are commonly
encountered in clinical trials, which may have an impact on sample size calculation, should also
be taken into consideration when performing sample size calculation. These practical issues
include unequal treatment allocation, adjustment for dropouts or covariates, mixed-up treat-
ment codes, treatment study center imbalance, multiplicity, multiple-stage design for early
stopping, and sample size calculation based on rare or extremely low incidence rate.
In Section 2.1, we introduce the concepts of confounding and interaction effects in clinical
trials. Section 2.2 discusses the controversial issues between the use of a one-sided test and
a two-sided test in clinical research. In Section 2.3, we summarize the difference in sample
size calculation between a crossover design and a parallel design. The concepts of group
sequential boundaries and alpha spending function in subgroup/interim analyses in clinical
trials are discussed in Section 2.4. Section 2.5 clarifies some issues that are commonly seen
in sample size calculation based on transformed data such as log-transformed data under
a parallel design or a crossover design. Section 2.6 provides a discussion regarding some
practical issues that have impact on sample size calculation in clinical trials. These issues
include unequal treatment allocation in randomization, sample size adjustment for dropouts
or covariates, the effect of mixed-up treatment codes during the conduct of clinical trials, the
loss in power for treatment and/or center imbalance, the issue of multiplicity in multiple pri-
mary endpoints and/or multiple comparisons, multiple-stage design for early stopping, and
sample size calculation based on rare or extremely low incidence rate in safety assessment.

2.1 Confounding and Interaction


2.1.1 Confounding
Confounding effects are defined as effects contributed by various factors that cannot be
separated by the design under study (Chow and Liu, 1998, 2003, 2013). Confounding is an
important concept in clinical research. When confounding effects are observed in a clini-
cal trial, the treatment effect cannot be assessed because it is contaminated by other effects
contributed by various factors.

21
22 Sample Size Calculations in Clinical Research

In clinical trials, there are many sources of variation that have an impact on the primary
clinical endpoints for clinical evaluation of a test drug under investigation. If some of
these variations are not identified and properly controlled, they can become mixed with
the treatment effect that the trial is designed to demonstrate, in which case the treatment
effect is said to be confounded by effects due to these variations. In clinical trials, there are
many subtle, unrecognizable, and seemingly innocent confounding factors that can cause
ruinous results of clinical trials. Moses (1992) gives the example of the devastating result
in the confounder being the personal choice of a patient. The example concerns a polio-
vaccine trial that was conducted on 2 million children worldwide to investigate the effect
of Salk poliomyelitis vaccine. This trial reported that the incidence rate of polio was lower
in children whose parents refused injection than in those who received placebo after their
parents gave permission (Meier, 1989). After an exhaustive examination of the data, it was
found that susceptibility to poliomyelitis was related to the differences between families
who gave the permission and those who did not. Therefore, it is not clear whether the
effect of the incidence rate is due to the effect of Salk poliomyelitis vaccine or due to the
difference between families giving permission.

2.1.2 Interaction
An interaction effect between factors is defined as a joint effect with one or more contribut-
ing factors (Chow and Liu, 1998, 2003, 2013). Interaction is also an important concept in clin-
ical research. The objective of a statistical interaction investigation is to conclude whether
the joint contribution of two or more factors is the same as the sum of the contributions
from each factor when considered alone. When interactions among factors are observed,
an overall assessment on the treatment effect is not appropriate. In this case, it is suggested
that the treatment must be carefully evaluated for those effects contributed by the factors.
In clinical research, almost all adequate and well-controlled clinical trials are multicenter
trials. For multicenter trials, the FDA requires that the treatment-by-center interaction be
examined to evaluate whether the treatment effect is consistent across all centers. As a
result, it is suggested that statistical tests for homogeneity across centers (i.e., for detect-
ing treatment-by-center interaction) be provided. The significant level used to declare the
significance of a given test for a treatment-by-center interaction should be considered in
light of the sample size involved. Gail and Simon (1985) classify the nature of interaction as
either quantitative or qualitative. A quantitative interaction between treatment and center
indicates that the treatment differences are in the same direction across centers but the
magnitude differs from center to center, while a qualitative interaction reveals that sub-
stantial treatment differences occur in different directions in different centers. More dis-
cussion regarding treatment-by-center interaction can be found in Chow and Shao (2002).

2.1.3 Remark
In clinical trials, a stratified randomization is usually employed with respect to some prog-
nostic factors or covariates, which may have confounding and interaction effects on the
evaluation of the test drug under investigation. Confounding or interaction effects may
alter the conclusion of the evaluation of the test drug under investigation. Thus, a strati-
fied randomization is desirable if the presence of the confounding and/or interaction
effects of some factors is doubtful. In practice, although sample size calculation according
to some stratification factors can be similarly performed within each combination of the
stratification factors, it is not desirable to have too many stratification factors. Therefore, it
Considerations Prior to Sample Size Calculation 23

is suggested that the possible confounding and/or interaction effects of the stratification
factors should be carefully evaluated before a sample size calculation is performed and a
stratified randomization is carried out.

2.2 One-Sided Test versus Two-Sided Test


In clinical research, there has been a long discussion on whether a one-sided test or a two-
sided test should be used for clinical evaluation of a test drug under investigation. Sample
size calculations based on a one-sided test and a two-sided test are different at a fixed α
level of significance. As it can be seen from Equation 1.3, the sample size for comparing the
two means can be obtained as

(σ12 + σ 22 )( zα/2 + zβ )2
n= .
δ2

When σ12 = σ 22 = σ 2 , the above formula reduces to

2σ 2 ( zα/2 + zβ )2
n= .
δ2

If δ = cσ, then the sample size formula can be rewritten as

2( zα/2 + zβ )2
n= .
c2

Table 2.1 provides a comparison for sample sizes obtained based on a one-sided test or
a two-sided test at the α level of significance. The results indicate that sample size may be
reduced by about 21% when switching from a two-sided test to a one-sided test for testing at
the 5% level of significance with an 80% power for detection of a difference of 0.5 standard
deviation.
The pharmaceutical industry prefers a one-sided test for demonstration of clinical
superiority based on the argument that they will not run a study if the test drug would

TABLE 2.1
Sample Sizes Based on One-Sided Test and Two-Sided
Test at α Level of Significance
One-Sided Two-Sided
Test Test

α δ* 80% 90% 80% 90%


0.05 0.25σ 198 275 252 337
0.50σ 50 69 63 85
1.00σ 13 18 16 22
0.01 0.25σ 322 417 374 477
0.50σ 81 105 94 120
1.00σ 21 27 24 30
24 Sample Size Calculations in Clinical Research

TABLE 2.2
Comparison between One-Sided Test and Two-Sided Test at α Level
of Significance
Characteristic One-Sided Test Two-Sided Test
Hypotheses Noninferiority/superiority Equality/equivalence
One trial 1/20 1/40
Two trials 1/400 1/1600

be worse. In practice, however, many drug products such as drug products in the central
nervous system may show a superior placebo effect as compared to the drug effect. This
certainly argues against the use of a one-sided test. Besides, a one-sided test allows more
bad drug products to be approved because of chances as compared to a two-sided test.
As indicated earlier, the FDA requires that at least two adequate and well-controlled
clinical studies be conducted to provide substantial evidence regarding the effectiveness
and safety of a test drug under investigation. For each of the two adequate and well-con-
trolled clinical trials, suppose the test drug is evaluated at the 5% level of significance.
Table 2.2 provides a summary of comparison between one-sided test and two-sided test in
clinical research. For the one-sided test procedure, the false-positive rate is one out of 400
trials (i.e., 0.25%) for the two trials, while the false-positive rate is one out of 1600 trials (i.e.,
0.0625%) for two trials when applying a two-sided test.
Some researchers from the academia and the pharmaceutical industry consider this
false-positive rate as acceptable and the evidence provided by the two clinical trials using
the one-sided test procedure as rather substantial. Hence, the one-sided test procedure
should be recommended. However, in practice, a two-sided test may be preferred because
placebo effect may be substantial in many drug products such as drug products regarding
diseases in the central nervous system.

2.2.1 Remark
Dubey (1991) indicated that the FDA prefers a two-sided test over a one-sided test proce-
dure in clinical research and development of drug products. In situations where (i) there
is concern with outcomes in only one tail and (ii) it is completely inconceivable that results
could go in the opposite direction, one-sided test procedure may be appropriate (Dubey,
1991). Dubey (1991) provided situations where one-sided test procedure may be justified.
These situations include (i) toxicity studies, (ii) safety evaluation, (iii) the analysis of occur-
rences of adverse drug reaction data, (iv) risk evaluation, and (v) laboratory data.

2.3 Crossover Design versus Parallel Design


As indicated earlier, an adequate and well-controlled clinical trial requires that a valid
study design be employed for a valid assessment of the effect of the test drug under inves-
tigation. As indicated in Chow and Liu (1998), commonly used study designs in clinical
research include parallel design, crossover design, enrichment design, titration design, or
a combination of these designs. Among these designs, crossover and parallel designs are
probably the two most commonly employed study designs.
Considerations Prior to Sample Size Calculation 25

2.3.1 Intersubject and Intrasubject Variabilities


Chow and Liu (1998, 2003, 2013) suggested that relative merits and disadvantages of candi-
date designs should be carefully evaluated before an appropriate design is chosen for the
intended trial. The clarification of the intrasubject and intersubject variabilities is essential
for sample size calculation in clinical research when a crossover design or a parallel design
is employed.
Intrasubject variability is the variability that could be observed by repeating experi-
ments on the same subject under the same experimental condition. The source of intrasu-
bject variability could be multifold. One important source is biological variability. Exactly
the same results may not be obtained even if they are from the same subject under the same
experimental condition. Another important source is measurement or calculation error.
For example, in a bioequivalence study with healthy subjects, it could be (i) the error when
measuring the blood or plasma concentration–time curve, (ii) the error when calculating
AUC (area under the curve), and/or (iii) the error of rounding after log-transformation.
Intrasubject variability could be eliminated if we could repeat the experiment infinitely
many times (in practice, this just means a large number of times) on the same subject under
the same experimental condition and then take the average. The reason is that intrasubject
variability tends to cancel each other on average in a large scale. If we repeat the experi-
ment on different subjects infinitely many times, it is possible that we may still see that the
averages of the responses from different subjects are different from each other even if the
experiments are carried out under exactly the same conditions. Then, what causes this dif-
ference or variation? It is not due to intrasubject variability, which has been eliminated by
averaging infinitely repeated experiments; it is not due to experimental conditions, which
are exactly the same for different subjects. Therefore, this difference or variation can only
be due to the unexplained difference between the two subjects.
It should be pointed out that sometimes people may call the variation observed from
different subjects under the same experimental condition intersubject variability, which is
different from the intersubject variability defined here. The reason is that the variability
observed from different subjects under the same experimental condition could be due
to unexplained difference among subjects (pure intersubject variability); it also could be
due to the biological variability, or measurement error associated with different experi-
ments on different subjects (intrasubject variability). Therefore, it is clear that the observed
variability from different subjects incorporates two components. They are, namely, pure
intersubject variability and intrasubject variability. We refer to it as the total intersubject
variability. For simplicity, it is also called total variability, which is the variability one
would observe from a parallel design.
In practice, no experiment can be carried out infinitely many times. It is also not always
true that the experiment can be repeatedly carried out on the same subject under the same
experimental condition. But, we can still assess these two variability components (intra-
and inter-) under certain statistical models, for example, a mixed effects model.

2.3.2 Crossover Design


A crossover design is a modified randomized block design in which each block receives
more than one treatment at different dosing periods. In a crossover design, subjects are
randomly assigned to receive a sequence of treatments, which contains all the treatments
in the study. For example, for a standard two-sequence, two-period 2 × 2 crossover design,
subjects are randomly assigned to receive one of the two sequences of treatments (say, RT
26 Sample Size Calculations in Clinical Research

and TR), where T and R represent the test drug and the reference drug, respectively. For
subjects who are randomly assigned to the sequence of RT, they receive the reference drug
first and then crossover to receive the test drug after a sufficient length of washout. The
major advantage of a crossover design is that it allows a within-subject (or intrasubject)
comparison between treatments (each subject serves as its own control) by removing the
between-subject (or intersubject) variability from the comparison. Let µT and µR be the
mean of the responses of the study endpoint of interest. Also, let σS2 and σ e2 be the intersub-
ject variance and intrasubject variance, respectively. Define θ = (µT − µR)/µR and assume
that the equivalence limit is δ = 0.2µR. Then, under a two-sequence, two-period crossover
design, the formula for sample size calculation is given by (see, also Chow and Wang, 2001)

CV 2 (tα ,2 n−2 + tβ /2 ,2 n−2 )2


n≥ ,
(0.2 −|θ|)2

where CV = σe/µR.

2.3.3 Parallel Design


A parallel design is a complete randomized design in which each subject receives one and
only one treatment in a random fashion. The parallel design does not provide independent
estimates for the intrasubject variability for each treatment. As a result, the assessment of
treatment effect is made based on the total variability, which includes the intersubject vari-
ability and the intrasubject variability.
Under a parallel design, assuming that the equivalence limit δ = 0.2µR, the following
formula is useful for sample size calculation (Chow and Wang, 2001):

2CV 2 (tα , 2 n−2 + tβ/2 , 2 n−2 )2


n≥ ,
(0.2−|θ |)2

where CV = σ/µR and σ 2 = σS2 + σ e2 .

2.3.4 Remark
In summary, in a parallel design, the comparison is made based on the intersubject varia-
tion, while in a crossover design the comparison is made based on the intrasubject variation.
As a result, sample size calculation under a parallel design or a crossover design is similar
and yet different. Note that the above formulas for sample size calculation are obtained
based on raw data. Sample size formulas based on log-transformation data under a parallel
design or a crossover design can be similarly obtained (Chow and Wang, 2001). More dis-
cussion regarding data transformation such as a log-transformation is given in Section 2.5.

2.4 Subgroup/Interim Analyses


In clinical research, subgroup analyses are commonly performed in clinical trials.
Subgroup analyses may be performed with respect to subject prognostic or confounding
Considerations Prior to Sample Size Calculation 27

factors such as demographics or subject characteristics at baseline. The purpose of this


type of subgroup analysis is to isolate the variability due to the prognostic or confound-
ing factors for an unbiased and reliable assessment of the efficacy and safety of the test
drug under investigation. In addition, many clinical trial protocols may call for an interim
analysis or a number of interim analyses during the conduct of the trials for the purpose
of establishing early efficacy and/or safety monitoring. The rationale for interim analyses
of accumulating data in clinical trials has been well established in the literature. See, for
example, Armitage et al. (1969), Haybittle (1971), Peto et al. (1976), Pocock (1977), O’Brien
and Fleming (1979), Lan and DeMets (1983), PMA (1993), and DeMets and Lan (1994).

2.4.1 Group Sequential Boundaries


For interim analyses in clinical trials, it is suggested that the number of planned interim
analyses should be specified in the study protocol. Let N be the total planned sample size
with equal allocation to the two treatments. Suppose that K interim analyses is planned
with equal increment of accumulating data. Then we can divide the duration of the clini-
cal trial into K intervals. Within each stage, the data of n = N/K patients are accumulated.
At the end of each interval, an interim analysis can be performed using the Z-statistic,
denoted by Zi, with the data accumulated up to that point. Two decisions will be made
based on the result of each interim analysis. First, the trial will continue if

|Zi|≤ zi , i = 1,, K − 1, (2.1)

where the zi are some critical values that are known as the group sequential boundaries. We
fail to reject the null hypothesis if

|Zi |≤ zi , for all i = 1, ... , K . (2.2)

Note that we may terminate the trial if the null hypothesis is rejected at any of the K
interim analyses (|Zi| > zi, i = 1, … , K). For example, at the end of the first interval, an
interim analysis is carried out with data from n subjects. If we fail to reject the null hypoth-
esis, we continue the trial to the second planned interim analysis. Otherwise, we reject the
null hypothesis and may stop the trial. The trial may be terminated at the final analysis if
we fail to reject the null hypothesis at the final analysis. Then we declare that the data from
the trial provide sufficient evidence to doubt the validity of the null hypothesis. Otherwise,
the null hypothesis is rejected and we conclude that there is statistically significant differ-
ence in change from baseline between the test drug and the control.
In contrast to the fixed sample where only one final analysis is performed, K analyses are
carried out for the K-stage group sequential procedure. Suppose that the nominal signifi-
cance level for each of the K interim analyses is still 5%. Then, because of repeated testing
based on the accumulated data, the overall significance level is inflated. In other words, the
probability of declaring at least one significance result increases due to K interim analyses.
Various methods have been proposed to maintain the overall significance level at the pre-
specified nominal level. One of the early methods was proposed by Haybittle (1971) and
Peto et al. (1976). They proposed to use 3.0 as group sequential boundaries for all interim
analyses except for the final analysis for which they suggested 1.96. In other words,

3.0, if i = 1,…, K − 1,
zi = 
1.96, if i = K .
28 Sample Size Calculations in Clinical Research

Therefore, their method can be summarized as follows:

Step 1: At each of the K interim analyses, compute Zi, i = 1, … , K − 1.


Step 2: If the absolute value of Zi crosses 3.0, then reject the null hypothesis and rec-
ommend a possible early termination of the trial; otherwise, continue the trial to
the next planned interim analysis and repeat Steps 1 and 2.
Step 3: For the final analysis, use 1.96 for the boundary. Trial stops here regardless of
whether the null hypothesis is rejected.

Haybittle and Peto’s method is very simple. However, it is a procedure with ad hoc
boundaries that are independent of the number of planned interim analyses and stage of
interim analyses. Pocock (1977) proposed different group sequential boundaries, which
depend upon the number of planned interim analyses. However, his boundaries are con-
stant at each stage of interim analyses. Since limited information is included in the early
stages of interim analyses, O’Brien and Fleming (1979) suggested posing conservative
boundaries for interim analyses scheduled to be carried out during an early phase of the
trial. Their boundaries not only depend upon the number of interim analyses but are also
a function of stages of interim analysis. As a result, the O’Brien–Fleming boundaries can
be calculated as follows:

ck k
zik = , 1 ≤ i ≤ k ≤ K, (2.3)
i

where ck is the critical value for a total of k planned interim analyses. As an example, sup-
pose that five planned interim analyses are scheduled. Then, c5 = 2.04 and boundaries for
each stage of these five interim analyses are given as

2.04 5
zi5 = , 1 ≤ i ≤ 5.
i

Thus, O’Brien–Fleming boundary for the first interim analysis is equal to


(2.04)( 5 ) = 4.561. The O’Brien–Fleming boundaries for the other four interim analyses
can be similarly computed as 3.225, 2.633, 2.280, and 2.040, respectively. The O’Brien–
Fleming boundaries are very conservative so that the early trial results must be extreme
for any prudent and justified decision-making in recommendation of a possible early ter-
mination when very limited information is available. On the other hand, for the late phase
of the trial when the accumulated information approaches the required maximum infor-
mation, their boundaries also become quite close to the critical value when no interim
analysis had been planned. As a result, the O’Brien–Fleming method does not require a
significant increase in the sample size for what has already been planned. Therefore, the
O’Brien–Fleming group sequential boundaries have become one of the most popular pro-
cedures for the planned interim analyses of clinical trials.

2.4.2 Alpha Spending Function


The idea of the alpha spending function proposed by Lan and DeMets (1983) is to spend (i.e.,
distribute) the total probability of false-positive risk as a continuous function of the infor-
mation time. The implementation of the alpha spending function requires the selection and
Considerations Prior to Sample Size Calculation 29

specification of the spending function in advance in the protocol. One cannot change and
choose another spending function in the middle of the trial. Geller (1994) suggested that
the spending function should be convex and have the property that the same value of a test
statistic is more compelling as the sample sizes increase. Because of its flexibility and lack of
requirement for total information and equal increment of information, there is a potential
to abuse the alpha spending function by increasing the frequency of interim analyses as the
results approach the boundary. However, DeMets and Lan (1994) reported that alteration of
the frequency of interim analyses has very little impact on the overall significance level if
the O’Brien–Fleming-type or Pocock-type continuous spending function is used.
Pawitan and Hallstrom (1990) studied the alpha spending function with the use of the
permutation test. The permutation test is conceptually simple and provides an exact test
for small sample sizes. In addition, it is valid for complicated stratified analysis in which
the exact sampling distribution is, in general, unknown and large-sample approximation
may not be adequate. Consider the one-sided alternative. For the kth interim analyses,
under the assumption of no treatment effect, the null joint permutation distribution of test
statistics (Z1, … , ZK) can be obtained by random permutation of treatment assignments on
the actual data. Let (Z1*b ,, ZKb
*
), b = 1, … , B, be the statistics computed from B treatment
assignments and B be the total number of possible permutations. Given α(s1), α(s2) − α(s1),
… , α(sK) − α(sK−1), the probabilities of type I error allowed to spend at successive interim
analyses, the one-sided boundaries z1, … , zK can be determined by

Number of (Z1* > z1 )


= α(s1 ),
B
and

Number of (Z1* > z1 or Z2* > z2 ,  , or Zk* > zk )


= α(sk ) − α(sk −1 ),
B

k = 1, … , K. If B is very large, then the above method can be executed with a random
sample with replacement of size B. The α spending function for an overall significance
level of 2.5% for one-sided alternative is given by

α
 s, if s < 1,
α(s) =  2

 α , if s = 1.

In the interest of controlling the overall type I error rate at the α level of significance,
sample size is necessarily adjusted according to the α spending function to account for
the planned interim analyses. In some cases, sample size reestimation without unblinding
may be performed according to the procedure described in Section 1.3 of Chapter 1. More
details can be found in Chapter 8.

2.5 Data Transformation


In clinical research, data transformation on clinical response of the primary study end-
point may be necessarily performed before statistical analysis for a more accurate and
30 Sample Size Calculations in Clinical Research

reliable assessment of the treatment effect. For example, for bioavailability and bioequiva-
lence studies with healthy human subjects, the FDA requires that a log-transformation be
performed before data analysis. Two drug products are claimed bioequivalent in terms of
drug absorption if the 90% confidence interval of the ratio of means of the primary phar-
macokinetic (PK) parameters, such as area under the blood or plasma concentration–time
curve (AUC) and maximum concentration (Cmax), is entirely within the bioequivalence lim-
its of (80%, 125%). Let µT and µR be the population means of the test drug and the reference
drug, respectively. Also, let X and Y be the PK responses for the test drug and the reference
drug. After log-transformation, we assume that log X and log Y follow normal distribu-
tions with means µX∗ and µY∗ and variance σ2. Then,
* 2 * 2
µT = E(X ) = e µX +(σ /2) and µR = E(Y ) = e µY +(σ /2) ,

which implies

µ 
log  T  = log (e µ X +µY ) = µX∗ − µY∗ .
* *

 µR 

Under both the crossover and parallel design, an exact (1 − α)100% confidence interval
for µX∗ − µY∗ can be obtained based on the log-transformed data. Hence, an exact (1 − α)100%
confidence interval for µT/µR can be obtained after the back transformation.
Chow and Wang (2001) provided sample size formulas under a parallel design or a cross-
over design with and without log-transformation. These formulas are different but very
similar. In practice, scientists often confuse them with one another. The following discus-
sion may be helpful for clarification.
We note that the sample size derivation is based on normality assumption for the raw
data and log-normality assumption for the transformed data. Thus, it is of interest to study
the distribution of log X when X is normally distributed with mean µ and variance σ2.
Note that

 X − µ  σ 2
Var  = = CV 2 .
 µ  µ2

If CV is sufficiently small, (X − µ)/µ is close to 0. As a result, by Taylor’s expansion,

 X − µ  X − µ
log X − log µ = log 1 + ≈ .
 µ  µ

Then,

X −µ
log X ≈ log µ + ~ N (log µ , CV 2 ).
µ

This indicates that when CV is small, log X is still approximately normally distributed,
even if X is from a normal population. Therefore, the procedure based on log-transformed
Considerations Prior to Sample Size Calculation 31

TABLE 2.3
Posterior Power Evaluation under a Crossover Design
Data Type Power

 0.2 
Raw data 1 − 2Φ  − tα ,n1+n2−2 
 CV (1/n1 ) + (1/n2 ) 

 0.223 
Log-transformed data 1 − 2Φ  − tα ,n1+n2−2 
 σ e (1/n1 ) + (1/n2 ) 

data is robust in some sense. In addition, the CV observed from the raw data is very simi-
lar to the variance obtained from the log-transformed data.
Traditionally, for the example regarding bioavailability and bioequivalence (BE) with
raw data, BE can be established if the 90% confidence interval for µT − µR is entirely within
the interval of (−0.2µR, 0.2µR) (Chow and Liu, 1992). This is the reason why 0.2 appears in
the formula for raw data. However, both the 1992 FDA and the 2000 FDA guidances rec-
ommended that a log-transformation be performed before bioequivalence assessment is
made. For log-transformed data, the BE can be established if the 90% confidence interval
for µT/µR is entirely located in the interval (80%, 125%). That is why log 1.25 appears in the
formula for log-transformed data. It should be noted that log 1.25 = −log 0.8 = 0.2231. In
other words, the BE limit for the raw data is symmetric about 0 (i.e., ±0.2µR), while the BE
limit for the log-transformed data is also symmetric about 0 after log-transformation.

2.5.1 Remark
For the crossover design, since each subject serves as its own control, the intersubject vari-
ation is removed from comparison. As a result, the formula for sample size calculation
derived under a crossover design only involves the intrasubject variability. On the other
hand, for the parallel design, formula for sample size calculation under a parallel design
includes both the inter- and intrasubject variabilities. In practice, it is easy to get confused
with the sample size calculation and/or evaluation of posterior power based on either raw
data or log-transformed data under either a crossover design or a parallel design (Chow
and Wang, 2001). As an example, posterior powers based on raw data and log-transformed
data under a crossover design when the true mean difference is 0 are given in Table 2.3.

2.6 Practical Issues


2.6.1 Unequal Treatment Allocation
In a parallel design or a crossover design comparing two or more than two treatments,
sample sizes in each treatment group (for parallel design) or in each sequence of treat-
ments (for crossover design) may not be the same. For example, when conducting a
placebo-controlled clinical trial with very ill patients or patients with severe or life-threat-
ening diseases, it may not be ethical to put too many patients in the placebo arm. In this
case, the investigator may prefer to put fewer patients in the placebo (if the placebo arm
is considered necessary to demonstrate the effectiveness and safety of the drug under
32 Sample Size Calculations in Clinical Research

investigation). A typical ratio of patient allocation for situations of this kind is 1:2, that is,
each patient will have a one-third chance to be assigned to the placebo group and two-
third chance to receive the active drug. For different ratios of patient allocation, the sample
size formulas discussed can be directly applied with appropriate modification of the cor-
responding degrees of freedom in the formulas.
When there is unequal treatment allocation, say κ to 1 ratio, sample size for comparing
two means can be obtained as

(σ12 /κ + σ 22 )( zα/2 + zβ )2
n= .
δ2

When κ = 1, the above formula reduces to Equation 1.3. When σ12 = σ 22 = σ 2 , we have

(κ + 1)σ 2 ( zα/2 + zβ )2
n= .
κδ 2

Note that unequal treatment allocation will have an impact on randomization in clinical
trials, especially in multicenter trials. To maintain the integrity of blinding of an intended
trial, a blocking size of 2 or 4 in randomization is usually employed. A blocking size of 2
guarantees that one of the subjects in the block will be randomly assigned to the treatment
group and the other one will be randomly assigned to the control group. In a multicenter
trial comparing two treatments, if we consider a 2 to 1 allocation, the size of each block has
to be a multiple of 3, that is, 3, 6, or 9. In the treatment of having a minimum of two blocks
in each center, each center is required to enroll a minimum of six subjects. As a result, this
may have an impact on the selection of the number of centers. As indicated in Chow and
Liu (1998), as a rule of thumb, it is not desirable to have the number of subjects in each cen-
ter to be less than the number of centers. As a result, it is suggested that the use of a κ to
1 treatment allocation in multicenter trials should take into consideration the blocking size
in randomization and the number of centers selected.

2.6.2 Adjustment for Dropouts or Covariates


At the planning stage of a clinical study, sample size calculation provides the number
of evaluable subjects required for achieving a desired statistical assurance (e.g., an 80%
power). In practice, we may have to enroll more subjects to account for potential dropouts.
For example, if the sample size required for an intended clinical trial is n and the potential
dropout rate is p, then we need to enroll n/(1 − p) subjects to obtain n evaluable subjects at
the completion of the trial. It should also be noted that the investigator may have to screen
more patients to obtain n/(1 − p) qualified subjects at the entry of the study based on inclu-
sion/exclusion criteria of the trial.
Fleiss (1986) pointed out that a required sample size may be reduced if the response vari-
able can be described by a covariate. Let n be the required sample size per group when the
design does not call for the experimental control of a prognostic factor. Also, let n* be the
required sample size for the study with the factor controlled. The relative efficiency (RE)
between the two designs is defined as

n
RE = .
n*
Considerations Prior to Sample Size Calculation 33

As indicated by Fleiss (1986), if the correlation between the prognostic factor (covariate)
and the response variable is r, then RE can be expressed as

100
RE = .
1− r 2

Hence, we have

n∗ = n(1− r 2 ).

As a result, the required sample size per group can be reduced if the correlation exists.
For example, a correlation of r = 0.32 could result in a 10% reduction in the sample size.

2.6.3 Mixed-Up Randomization Schedules


Randomization plays an important role in the conduct of clinical trials. Randomization not
only generates comparable groups of patients who constitute representative samples from
the intended patient population, but also enables valid statistical tests for clinical evalu-
ation of the study drug. Randomization in clinical trials involves random recruitment of
patients from the targeted patient population and random assignment of patients to the
treatments. Under randomization, statistical inference can be drawn under some prob-
ability distribution assumption of the intended patient population. The probability dis-
tribution assumption depends on the method of randomization under a randomization
model. A study without randomization results in the violation of the probability distri-
bution assumption and consequently no accurate and reliable statistical inference on the
evaluation of the safety and efficacy of the study drug can be drawn.
A problem commonly encountered during the conduct of a clinical trial is that a propor-
tion of treatment codes are mixed up in randomization schedules. Mixing up treatment
codes can distort the statistical analysis based on the population or randomization model.
Chow and Shao (2002) quantitatively studied the effect of mixed-up treatment codes on
the analysis based on the intention-to-treat (ITT) population, which are described below.
Consider a two-group parallel design for comparing a test drug and a control (placebo),
where n1 patients are randomly assigned to the treatment group and n2 patients are ran-
domly assigned to the control group. When randomization is properly applied, the popu-
lation model holds and responses from patients are normally distributed. Consider first
the simplest case where two patient populations (treatment and control) have the same
variance σ2 and σ2. Let µ1 and µ2 be the population means for the treatment and the control,
respectively. The null hypothesis that µ1 = µ2 (i.e., there is no treatment effect) is rejected
at the α level of significance if

|x1 − x2 |
> zα/2 , (2.4)
σ (1/n1 ) + (1/n2 )

where x1 is the sample mean of responses from patients in the treatment group, x2 is the
sample mean of responses from patients in the control group, and zα/2 is the upper (α/2)
th percentile of the standard normal distribution. Intuitively, mixing up treatment codes
does not affect the significance level of the test.

You might also like