Understanding Randomized Controlled Trials
Understanding Randomized Controlled Trials
This chapter describes methodological and design considerations central to the scientific
evaluation of clinical treatment methods via randomized clinical trials (RCTs). Matters of
design, procedure, measurement, data analysis, and reporting are each considered in
turn. Specifically, the authors examine different types of controlled comparisons, random
assignment, the evaluation of treatment response across time, participant selection, study
setting, properly defining and checking the integrity of the independent variable (i.e.,
treatment condition), dealing with participant attrition and missing data, evaluating
clinical significance and mechanisms of change, and consolidated standards for
communicating study findings to the scientific community. After addressing
considerations related to the design and implementation of the traditional RCT, the
authors turn their attention to important extensions and variations of the RCT. These
treatment study designs include equivalency designs, sequenced treatment designs,
prescriptive designs, adaptive designs, and preferential treatment designs. Examples
from the recent clinical psychology literature are provided, and guidelines are suggested
for conducting treatment evaluations that maximize both scientific rigor and clinical
relevance.
Keywords: Randomized clinical trial, RCT, normative comparisons, random assignment, treatment integrity,
equivalency designs, sequenced treatment designs
The randomized controlled trial (RCT)—a group comparison design in which participants
are randomly assigned to treatment conditions—constitutes the most rigorous and
objective methodological design for evaluating therapeutic outcomes. In this chapter we
focus on RCT research strategies that maximize both scientific rigor and clinical
relevance (for consideration of single-case, multiple-baseline, and small pilot trial
designs, see Chapter 3 in this volume). We organize the present chapter around (a) RCT
design considerations, (b) RCT procedural considerations, (c) RCT measurement
Page 1 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
considerations, (d) RCT data analysis, and (e) RCT reporting. We then turn our attention
to extensions and variations of the traditional RCT, which offer various adjustments for
clinical generalizability, while at the same time sacrificing important elements of internal
validity. Although all of the methodological and design ideals presented may not always
be achieved within a single RCT, our discussions provide exemplars of the RCT.
Design Considerations
To adequately assess the causal impact of a therapeutic intervention, clinical researchers
must use control procedures derived from experimental science. In the RCT, the
intervention applied constitutes the experimental manipulation, and thus to have
confidence that an intervention is responsible for observed changes, extraneous factors
must be experimentally “controlled.” The objective is to distinguish intervention effects
from any changes that (p. 41) result from other factors, such as the passage of time,
patient expectancies of change, therapist attention, repeated assessments, and simple
regression to the mean. To maximize internal validity, the clinical researcher must
carefully select control/comparison condition(s), randomly assign participants across
treatment conditions, and systematically evaluate treatment response across time. We
now consider each of these RCT research strategies in turn.
Page 2 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
Importantly, not all control conditions are “created equal.” Deciding which form of control
condition to select for a particular study (e.g., no treatment, waitlist, attention-placebo,
standard treatment as usual) requires careful deliberation (see Table 4.1 for recent
examples from the literature). In a no-treatment control condition, comparison
participants are evaluated in repeated assessments, separated by an interval of time
equal in duration to the treatment provided to those in the experimental treatment
condition. Any changes seen in the treated participants are compared to changes seen in
the nontreated participants. When, relative to nontreated participants, the treated
participants show significantly greater improvements, the experimental treatment may be
credited with producing the observed changes. Several important rival hypotheses are
eliminated in a no-treatment design, including effects due to the passage of time,
maturation, (p. 42) spontaneous remission, and regression to the mean. Importantly,
however, other potentially important confounding factors not specific to the experimental
treatment—such as patient expectancies to get better, or meeting with a caring and
attentive clinician—are not ruled out in a no-treatment control design. Accordingly, no-
treatment control conditions may be useful in earlier stages of treatment development,
but to establish broad empirical support for an intervention, more informative control
procedures are preferred.
Page 3 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
Page 4 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
A more revealing variant of the no-treatment condition is the waitlist condition. Here,
participants in the waitlist condition expect that after a certain period of time they will
receive treatment, and accordingly may anticipate upcoming changes (which may in turn
affect their symptoms). Changes are evaluated at uniform intervals across the waitlist and
experimental conditions, and if we assume the participants in the waitlist and treatment
conditions are comparable (e.g., comparable baseline symptom severity and gender, age,
and ethnicity distributions), we can then infer that changes in the treated participants
relative to waitlist participants are likely due to the intervention rather than to
expectations of impending change. However, as with no-treatment conditions, waitlist
conditions are of limited value for evaluating treatments that have already been examined
relative to “inactive” conditions.
Page 5 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
equated (Kazdin, 2003). When the experimental treatment and the standard care
intervention share comparable durations and participant and therapist expectancies, the
researcher can evaluate the relative efficacy of the interventions.
Random Assignment
Although random assignment does not ensure participant comparability across conditions
on all measures, randomization procedures do rigorously maximize the likelihood of
comparability. An alternative procedure, randomized blocks assignment, or assignment by
stratified blocks, involves matching (arranging) prospective participants in subgroups
that contain participants that are highly comparable on key dimensions (e.g.,
socioeconomic status indicators) and contain the same number of participants as the
number of conditions. For example, if the study requires three conditions—a standard
Page 7 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
Page 8 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
Page 9 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
Page 10 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
high). Stratified blocking offers a viable option to ensure that all treatments are
conducted by comparable therapists. It is wise to gather data on therapist variables (e.g.,
expertise, experience, allegiance) and examine their relationships to participant
outcomes.
For proper evaluation, intervention procedures across treatments must be equated for
key variables such as (a) duration; (b) length, intensity, and frequency of contacts with
participants; (c) credibility of treatment rationale; (d) treatment setting; and (e) degree of
involvement of persons significant to the participant. These factors may be the basis for
two alternative therapies (e.g., conjoint vs. individual marital therapy). In such cases, the
nonequated feature constitutes an experimentally manipulated variable rather than a
factor to control.
What is the best method of measuring change when two alternative treatments are being
compared? Importantly, measures should cover the range of symptoms and functioning
targeted for change, tap costs and potential negative side effects, and be unbiased with
respect to the alternate interventions. Assessments should not be differentially sensitive
to one treatment over another. Treatment comparisons will be misleading if measures are
not equally sensitive to the types of changes that most likely result from each intervention
type.
Procedural Considerations
We now address key RCT procedural considerations, including (a) sample selection, (b)
study (p. 46) setting, (c) defining the independent variable, and (d) checking the integrity
of the independent variable.
Page 11 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
Sample Selection
Selecting a sample to best represent the clinical population of interest requires careful
deliberation. A selected sample refers to a sample of participants who may require
treatment but who may otherwise only approximate clinically disordered persons. By
contrast, RCTs optimize external validity when treatments are applied and evaluated with
actual treatment-seeking patients. Consider a study investigating the effects of a
treatment on social anxiety disorder. The researcher could use (a) a sample of patients
diagnosed with social anxiety disorder via structured diagnostic interviews (genuine
clinical sample), (b) a sample consisting of a group of individuals who self-report shyness
(analogue sample), or (c) a sample of socially anxious persons after excluding cases with
depressed mood and/or substance use (highly select sample). This last sample may meet
full diagnostic criteria for social anxiety disorder but are nevertheless highly selected.
From a feasibility standpoint, clinical researchers may find it easier to recruit analogue
samples relative to genuine clinical samples, and such samples may afford a greater
ability to control various conditions and minimize threats to internal validity. At the same
time, analogue and select samples compromise external validity—these individuals are
not necessarily comparable to patients seen in typical clinical practice (and may not
qualify as an RCT). With respect to social anxiety disorder, for instance, one could
question whether social anxiety disorder in genuine clinical populations compares
meaningfully to self-reported shyness (see Heiser, Turner, Beidel, & Roberson-Nay, 2009).
When deciding whether to use clinical, analogue, or select samples, the researcher needs
to consider how the study results will be interpreted and generalized. Regrettably,
nationally representative data show that standard exclusion criteria set for clinical
treatment studies exclude up to 75 percent of affected individuals in the general
population who have major depression (Blanco, Olfson, Goodwin, et al., 2008).
Researchers must consider patient diversity when deciding which samples to study.
Research supporting the efficacy of psychological treatments has historically been
conducted with predominantly European-American samples, although this is rapidly
changing (see Huey & Polo, 2008). Although racially and ethnically diverse samples may
be similar in many ways to single-ethnicity samples, one can question the extent to which
efficacy findings from predominantly European-American samples can be generalized to
ethnic-minority samples (Bernal, Bonilla, & Bellido, 1995; Bernal & Scharron-Del-Rio,
2001; Hall, 2001; Olfson, Cherry, & Lewis-Fernandez, 2009; Sue, 1998). Investigations
have also addressed the potential for bias in diagnoses and in the provision of mental
health services to ethnic-minority patients (e.g., Flaherty & Meaer, 1980; Homma-True,
Green, Lopez, & Trimble, 1993; Lopez, 1989; Snowden, 2003).
A simple rule is that the research sample should reflect the broad population to which the
study results are to be generalized. To generalize to a single-ethnicity group, one must
study a single-ethnicity sample. To generalize to a diverse population, one must study a
diverse sample, as most RCTs strive to accomplish. Barriers to care must be reduced and
Page 12 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
outreach efforts employed to inform minorities of available services (see Sweeney, Robins,
Ruberu, & Jones, 2005; Yeh, McCabe, Hough, Dupuis, & Hazen, 2003) and include them
in the research. Walders and Drotar (2000) provide guidelines for recruiting and working
with ethnically diverse samples.
After the fact, appropriate statistical analyses can examine potential differential
outcomes (see Arnold et al., 2003; Treadwell, Flannery-Schroeder, & Kendall, 1994).
Although grouping and analyzing research participants by racial or ethnic status is a
common analytic approach, this approach is simplistic because it fails to address
variations in each patient's degree of ethnic identity. It is often the degree to which an
individual identifies with an ethnocultural group or community, and not simply his or her
ethnicity itself, that may moderate response to treatment. For further consideration of
this important issue, the reader is referred to Chapter 21 in this volume.
Study Setting
Some have questioned whether outcomes found at select research centers can transport
to clinical practice settings, and thus the question of whether an intervention can be
transported to other service settings requires independent evaluation (Southam-Gerow,
Ringeisen, & Sherrill, 2006). It is not sufficient to demonstrate treatment efficacy within
a narrowly defined sample in a highly selective setting. One should study, rather than
assume, that a treatment found to be efficacious within a research (p. 47) clinical setting
will be efficacious in a clinical service setting (see Hoagwood, 2002; Silverman, Kurtines,
& Hoagwood, 2004; Southam-Gerow et al., 2006; Weisz, Donenberg, Han, & Weiss, 1995;
Weisz, Weiss, & Donenberg, 1992). Closing the gap between RCTs and clinical practice
requires transporting effective treatments (getting “what works” into practice) and
identifying additional research into those factors that may be involved in successful
transportation (e.g., patient, therapist, researcher, service delivery setting; see Kendall &
Southam-Gerow, 1995; Silverman et al., 2004). Methodological issues relevant to the
conduct of research evaluating the transportability of treatments to “real-world” settings
can be found in Chapter 5 in this volume.
Page 13 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
training and contribute meaningfully to replication (Dobson & Hamilton, 2002; Dobson &
Shaw, 1988).
The merits of manual-based treatments are not universally agreed upon. Debate has
ensued regarding the appropriate use of manual-based treatments versus a more variable
approach typically found in clinical practice (see Addis, Cardemil, Duncan, & Miller, 2006;
Addis & Krasnow, 2000; Westen, Novotny, & Thompson-Brenner, 2004). Some have
argued that treatment manuals limit therapist creativity and place restrictions on the
individualization that the clinicians use (see also Waltz, Addis, Koerner, & Jacobson, 1993;
Wilson, 1995). Indeed, some therapy manuals may appear “cookbook-ish,” and some lack
attention to the clinical sensitivities needed for implementation and individualization, but
our experience and data suggest that this is not the norm. An empirical evaluation from
our laboratory found that the use of a manual-based treatment for child anxiety disorders
(Kendall & Hedtke, 2006) did not restrict therapist flexibility (Kendall & Chu, 1999).
Although it is not the goal of manual-based treatments to have clinicians perform
treatment in a rigid manner, this misperception has restricted some clinicians’ openness
to manual-based interventions (Addis & Krasnow, 2000).
Several modern treatment manuals allow the therapist to attend to each patient's specific
circumstances, clinical needs, concerns, and comorbid diagnoses without deviating from
the core treatment strategies detailed in the manual. The goal is to include provisions for
standardized implementation of therapy while using a personalized case formulation
(e.g., see Suveg, Comer, Furr, & Kendall, 2006). Importantly, use of manual-based
treatments does not eliminate the potential for differential therapist effects. Researchers
examine therapist variables within the context of manual-based treatments (e.g.,
therapeutic relationship-building behaviors, flexibility, warmth) that may relate to
treatment outcome (Creed & Kendall, 2005; Karver et al., 2008; Shirk et al., 2008; see
also Chapter 9 in this volume for a full consideration of designing, conducting, and
evaluating therapy process research).
Page 14 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
To help ensure that the treatments are indeed implemented as intended, it is wise to
require that a treatment plan be followed, that therapists are (p. 48) carefully trained,
and that sufficient supervision is available throughout. The researcher is wise to conduct
an independent check on the manipulation. For example, treatment sessions are recorded
so that an independent rater can listen to and/or watch the recordings and provide
quantifiable judgments regarding key characteristics of the treatment. Such a
manipulation check provides the necessary assurance that the described treatment was
indeed provided as intended. Digital audio and video recordings are inexpensive, can be
used for subsequent training, and can be analyzed to answer key research questions.
Therapy session recordings evaluated in RCTs not only provide a check on the treatment
within each separate study but also allow for a check on the comparability of treatments
provided across studies. That is, the therapy provided as CBT in one researcher's RCT
could be checked to assess its comparability to other teams’ CBT.
A recently completed clinical trial from our research program comparing two active-
treatment conditions for childhood anxiety disorders against an active attention control
condition (Kendall et al., 2008) illustrates a procedural plan for integrity checks. First, we
developed a checklist of the strategies and content called for in each session by the
respective treatment manuals. A panel of expert clinicians served as independent raters
who used the checklists to rate randomly selected video segments from randomly
selected cases. The panel of raters was trained on nonstudy cases until they reached an
interrater reliability of Cohen's κ ≥ .85. After ensuring reliability, the panel used the
checklists to assess whether the appropriate content was covered for randomly selected
segments that were representative of all sessions, conditions, and therapists. For each
coded session, we computed an integrity ratio corresponding to the number of checklist
items covered by the therapist divided by the total number of items that should have been
included. Integrity check results indicated that across the conditions, 85 to 92 percent of
intended content was in fact covered.
It is also wise for the RCT researcher to evaluate the quality of treatment provided. A
therapist may strictly adhere to a treatment manual and yet fail to administer the
treatment in an otherwise competent manner, or he or she may administer therapy while
significantly deviating from the manual. In both cases, the operational definition of the
independent variable (i.e., the treatment manual) has been violated, treatment integrity
impaired, and replication rendered impossible (Dobson & Shaw, 1988). When a treatment
fails to demonstrate expected gains, one can examine the adequacy with which the
treatment was implemented (see Hollon, Garber, & Shelton, 2005). It is also of interest to
investigate potential variations in treatment outcome that may be associated with
differences in the quality of the treatment provided (Garfield, 1998; Kendall & Hollon,
1983). Expert judges are needed to make determinations of differential quality prior to
Page 15 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
Page 16 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
Measurement Considerations
Page 17 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
Data Analysis
Data analysis is an active process through which we extract useful information from the
data we have collected in ways that allow us to make statistical inferences about the
larger population that a given sample was selected to represent. Data do not “speak” for
themselves. Although a comprehensive statistical discussion about RCT data analysis is
beyond the present scope (the reader is referred to Jaccard & Guilamo-Ramos, 2002a,
2002b; Kraemer & Kupfer, 2006; Kraemer, Wilson, Fairburn, & Agras, 2002; and Chapters
14 and 16 in this volume) in this section, we discuss three areas that merit consideration
in the context of RCT data analysis: (a) addressing missing data and attrition, (b)
assessing clinical significance, and (c) evaluating mechanisms of change (i.e., mediators
and moderators).
Page 18 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
are large numbers of noncompleters or when attrition varies across conditions (Leon et
al., 2006; Molenberghs et al., 2004).
Regardless of how diligently researchers work to prevent attrition, data will likely be lost.
Although attrition rates vary across RCTs and treated clinical populations, Mason (1999)
estimated that most researchers can expect roughly 20 percent of their sample to
withdraw or be removed from a study prior to completion. To address this matter,
researchers can conduct and report two sets of analyses: (a) analyses of outcomes for
treatment completers and (b) analyses of outcomes for all participants who were included
at the time of randomization (i.e., the intent-to-treat sample). Treatment-completer
analyses involve the evaluation of only those who actually completed treatment and
examine what the effects of treatment are when someone completes a full treatment
course. Treatment refusers, treatment dropouts, and participants who fail to adhere to
treatment schedules are not included in such analyses. Reports of such treatment
outcomes may be somewhat elevated because they represent (p. 50) the results for only
those who adhered to and completed the treatment. A more conservative approach to
addressing missing data, intent-to-treat analysis, entails the evaluation of outcomes for all
participants involved at the point of randomization. As proponents of intent-to-treatment
analyses we say, “once randomized, always analyzed.”
Page 19 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
the nonidentical datasets, the results are pooled and the resulting variability addresses
the uncertainty of the true value of the missing data.
Data produced by RCTs are submitted to statistical tests of significance. Mean scores for
participants in each condition are compared, within-group and between-group variability
is considered, and the analysis produces a numerical figure, which is then checked
against critical values. Statistical significance is achieved when the magnitude of the
mean difference is beyond that which could have resulted by chance alone
(conventionally defined as p 〈 .05). Tests of statistical significance are essential as they
inform us that the degree of change was likely not due to chance.
Importantly, statistical tests alone do not provide evidence of clinical significance. Sole
reliance on statistical significance can lead to perceiving treatment gains as potent when
in fact they may be clinically insignificant. For example, imagine that the results of a
treatment outcome study demonstrate that mean Beck Depression Inventory (BDI) scores
are significantly lower at posttreatment than pretreatment. An examination of the means,
however, reveals only a small but reliable shift from a mean of 29 to a mean of 26. With
larger sample sizes, this difference may well achieve statistical significance at the
conventional p 〈 .05 level (i.e., over 95 percent chance that the finding is not due to
chance alone), yet perhaps be of limited practical significance. Both before and after
treatment, the scores are within the range considered indicative of clinical levels of
depressive distress (Kendall, Hollon, Beck, Hammen, & Ingram, 1987), and such a small
magnitude of change may have little effect on a person's (p. 51) life impairment (Gladis,
Gosch, Dishuk, & Crits-Christoph, 1999). Conversely, statistically meager results may
Page 20 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
Normative comparisons (Kendall & Grove, 1988; Kendall, Marrs-Garcia, Nath, &
Sheldrick, 1999) are conducted in several steps. First, the researcher selects a normative
group for posttreatment comparison. Given that several well-established measures
provide normative data (e.g., the BDI, the Child Behavior Checklist), investigators may
choose to rely on these preexisting normative samples. However, when normative data do
not exist, or when the treatment sample is qualitatively different on key factors (e.g.,
socioeconomic status indicators, age), it may be necessary to collect one's own normative
data. In a typical RCT, when using statistical tests to compare groups, the investigator
assumes equivalency across groups (null hypothesis) and aims to find that they are not
(alternate hypothesis). However, when the goal is to show that treated individuals are
equivalent to “normal” individuals on some factor (i.e., are indistinguishable from
normative comparisons), traditional hypothesis-testing methods are inadequate. One uses
an equivalency testing method to circumvent this problem (Kendall, Marrs-Garcia, et al.,
1999) that examines whether the difference between the treatment and normative groups
is within some predetermined range. When used in conjunction with traditional
hypothesis testing, this approach allows conclusions to be drawn about the equivalency of
groups (see, e.g., Jarrett, Vittengl, Doyle, & Clark, 2007; Pelham et al., 2000; Westbrook &
Kirk, 2007, for examples of normative comparisons), thus testing that posttreatment data
are within a normative range on the measure of interest. For example, Weisz and
colleagues (1997) utilized normative comparisons in a trial in which elementary school
children with mild to moderate symptoms of depression were randomly assigned either to
a Primary and Secondary Control Enhancement Training (PASCET) program or to a no-
treatment control group. Normative comparisons were used to determine whether
participants’ scores on two depression measures, the Children's Depression Inventory
and the Revised Children's Depression Rating Scale, fell within one standard deviation
above elementary school norm groups at pretreatment, posttreatment, and 9-month
follow-up time points. Utilizing normative comparisons allowed the authors to conclude
that children who had received the treatment intervention were more likely to fall within
the normal range on depression measures than children in the no-treatment control
condition.
Page 21 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
The Reliable Change Index (RCI; Jacobson, Follette, & Revenstorf, 1984; Jacobson &
Traux, 1991) is another popular method to examine clinically significant change. The RCI
entails calculating the number of participants moving from a dysfunctional to a normative
range. Specifically, the research calculates a difference score (posttreatment minus
pretreatment) divided by the standard error of measurement (calculated based on the
reliability of the measure). The RCI is influenced by the magnitude of change and the
reliability of the measure. The RCI has been used in RCT research, although its
originators point out that it has at times been misapplied (Jacobson, Roberts, Berns, &
McGlinchey, 1999). When used in conjunction with reliable measures and appropriate
cutoff scores, it can be a valuable tool for assessing clinical significance.
The RCT researcher is often interested in identifying (a) the conditions that dictate when
a treatment is more or less effective and (b) the processes through which a treatment
produces change. Addressing such issues necessitates the specification of moderator and
mediator variables (Baron & Kenny, 1986; Holmbeck, 1997; Kraemer et al., 2002). A
moderator is a variable that delineates the conditions under which a given treatment is
related to an outcome. Conceptually, moderators identify on whom and under what
circumstances treatments have different effects (Kraemer et al., 2002). A moderator is
functionally a variable that influences either the strength or direction of a relationship
between an independent variable (treatment) and a dependent variable (outcome). For
example, if in an RCT the experimental treatment was found (p. 52) to be more effective
with men than with women, but this gender effect was not found in response to the
control treatment, then gender would be considered a moderator of the association
between treatment and outcome. Treatment moderators help clarify for consumers of the
treatment outcome literature which patients might be most responsive to which
treatments, and for which patients alternative treatments might be sought. Importantly,
when a variable broadly predicts outcome across all treatment conditions in an RCT,
conceptually that variable is simply a predictor, and not a moderator (see Kraemer et al.,
2002).
On the other hand, a mediator is a variable that serves to explain the process by which a
treatment affects an outcome. Conceptually, mediators identify how and why treatments
take effect (Kraemer et al., 2002). The mediator effect reveals the mechanism through
which the independent variable (e.g., treatment) is related to outcome (e.g., treatment-
related changes). Accordingly, mediational models are inherently causal models, and in
the context of an RCT, significant meditational pathways inform us about causal
relationships. If an effective treatment for child externalizing problems was found to have
an impact on parenting behavior, which in turn was found to have a significant influence
on child externalizing behavior, then parent behavior would be considered to mediate the
treatment-to-outcome relationship (provided certain statistical criteria were met; see
Page 22 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
Holmbeck, 1997). Specific statistical methods used to evaluate the presence of treatment
moderation and mediation can be found elsewhere (see Chapter 15 in this volume).
Next, the researcher must decide where to submit the report. We recommend that
researchers consider submitting RCT findings to peer-reviewed journals only. Publishing
RCT outcomes in a refereed journal (i.e., one that employs the peer-review process)
signals that the work has been accepted and approved for publication by a panel of
impartial and qualified reviewers (i.e., independent researchers knowledgeable in the
area but not involved with the RCT). Consumers should be highly cautious of RCTs
published in journals that do not place manuscript submissions through a rigorous peer-
review process. Although the peer-review process slows down the speed with which one
is able to communicate RCT results, much to the chagrin of the excited researcher who
just completed an investigation, it is nonetheless one of the indispensable safeguards that
we have to ensure that our collective knowledge base is drawn from studies meeting
acceptable standards. Typically, the review process is “blind,” meaning that the authors
Page 23 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
of the article do not know the identities of the peer reviewers who are considering their
manuscript. Many journals now employ a double-blind peer-review process in which the
identities of study authors are also not known to the peer reviewers.
Equivalency Designs
Page 24 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
Barlow and colleagues, for example, are currently testing the efficacy of a transdiagnostic
treatment (Unified Protocol for Emotional Disorders; Barlow, Farchione, Fairholme,
Ellard, Boisseau, et al., 2010) for anxiety disorders. The proposed analyses include a
rigorous comparison of the Unified Protocol (UP) against single-diagnosis psychological
treatment protocols (SDPs). Statistical equivalence will be used to test the hypothesis
that the UP is statistically equivalent to SDPs. An a priori confidence interval around
change in the clinical severity rating (CSR) will be utilized to evaluate statistical
equivalence among treatments. The potential finding that the UP is indeed equivalent to
SDPs in the treatment of anxiety disorders, regardless of specific diagnosis, would have
important implications for treatment dissemination and transportability.
Benchmarking equivalency designs allow for meaningful comparison groups with which
to gauge the progress of treated participants in a clinical trial. The comparison data are
typically readily available, given that they may include samples that have been used to
obtain normative data for specific measures, or research participants whose outcome
data are included in reported results in published studies. In addition, as noted earlier,
equivalency tests can be conducted to determine the clinical significance of treatment
Page 25 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
When the aim of a research study is to determine the most effective sequence of
treatments for an identified patient population, a sequenced treatment design may be
utilized. This design involves the assignment of study participants to a particular
sequence of treatment and control/comparison conditions. The order in which conditions
are assigned may be random, as in a randomized sequence design. In other sequenced
treatment designs, factors such as participant characteristics, individual treatment
outcomes, or participant preferences may influence the sequence of administered
treatments. These variations on sequenced treatment designs—prescriptive, adaptive,
and preferential treatment designs, respectively—are outlined in further detail below.
The prescriptive treatment design recognizes that individual patient characteristics play a
key role in treatment outcomes and assigns treatment condition based on these patient
characteristics. The basis of this treatment design aims to improve upon nomothetic data
models by incorporating idiographic data to treatment assignments (see Barlow & Nock,
2009). Study participants who are matched to treatment conditions based on individual
characteristics (e.g., psychiatric comorbidity, levels of distress and impairment, readiness
to change, etc.) may experience greater gains than those who are not matched to
interventions based on patient characteristics (Beutler & Harwood, 2000). In a
prescriptive (p. 55) treatment design, the clinical researcher studies the effectiveness of
a treatment decision-making algorithm as opposed to a set treatment protocol.
Participants do not have an equal chance of receiving study treatments, as is the case in
the standard RCT. Instead, what remains consistent across participants is the application
of the same decision-making algorithm, which can lead to a variety of sequenced
treatment courses.
Page 26 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
One of the primary goals of prescriptive treatment research designs is to examine the
effectiveness of treatments tailored to individuals for whom those treatments are thought
to work best. Key patient dimensions that have been found in nomothetic evaluations to
be effective mediators or moderators of treatment outcome can lay the groundwork for
decision rules to assign participants to particular interventions, or alternatively, can lead
to the administration or omission of a specific module within a larger intervention.
Prescriptive treatment designs offer opportunities to better develop and tailor efficacious
treatments to patients with varied characteristics.
Illustrating the utility of the adaptive treatment design, the Sequenced Treatment
Alternatives to Relieve Depression (STAR*D) study assessed the effectiveness of
depression treatments in patients with major depressive disorder (Rush et al., 2004).
Participants advanced through four levels of treatment and were assigned a particular
course of treatment depending on their response to treatment up until that point. In level
1, all participants were given citalopram for 12 to 14 weeks. Those who became
symptom-free after this period could continue on citalopram for a 12-month follow-up
period, while those who did not become symptom-free moved on to level 2. Levels 2 and 3
allowed participants to choose another medication or cognitive therapy (switch) or
augment their current medication with another medication or cognitive therapy (add-on).
In level 4, participants who were not yet symptom-free were taken off their medications
and randomly assigned to either a monoamine oxidase inhibitor (MAOI) or the
Page 27 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
combination of venlafaxine extended release with mirtazapine. At each stage of the study,
participants were assigned to treatments based on their previous treatment responses.
Roughly half of the participants became symptom-free after two treatment levels, and
roughly 70 percent of study completers became symptom-free over the course of all four
treatment levels.
The preferential treatment design allows study participants to choose the treatment
condition(s) to which they are assigned. This approach considers patient preferences,
which emulates the process that typically occurs in clinical practice. Taking into account
patient preferences in a treatment study can result in a better understanding of which
individuals will fare best when administered specific interventions under circumstances
that incorporate their preferences in determining treatment selection. Proponents often
argue that assigning treatments based on patient preference may increase other factors
known to positively affect treatment outcomes, including patient motivation, attitudes
toward treatment, and expectations of treatment success.
Lin and colleagues (2005) utilized a preferential treatment design to explore the effects of
matching patient preferences and interventions in a population of adults with major
depression. Participants were offered antidepressant medication and/or counseling based
on patient preference, where appropriate. Participants who were matched to their
treatment preference exhibited more positive treatment outcomes at 3- and 9-month
follow-up evaluations than participants who were not matched to their preferred
treatment condition.
Page 28 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
Multiple-groups crossover designs are best suited for the evaluation of interventions that
would not expectedly retain effects once they are removed, as is the case in the
evaluation of a therapeutic medication with a very short half-life. These designs are more
difficult to implement in the evaluation of psychosocial interventions, which often
produce effects that are somewhat irreversible (e.g., the learning of a skill, or the
acquisition of important knowledge). How can the clinical researcher evaluate separate
treatment phases when it is not possible to completely remove the intervention? In such
situations, crossover designs are misguided.
Proponents of sequential designs argue that designs that are informed by patient
characteristics, outcomes, and preferences provide patients with uniquely individualized
care within a clinical trial. The argument suggests that an appropriate match (p. 57)
between patient characteristics and treatment type will optimize success in producing
significant treatment effects and lead to a heightened understanding of interventions that
are best suited to a variety of patients and circumstances in clinical practice (Luborsky et
al., 2002). In this way, systematic evaluation is extended to the very decision-making
Page 29 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
algorithms that occur in real-world clinical practice, an important element not afforded
by the standard RCT. However, whereas these approaches increase clinical relevance and
may enhance the ability to generalize findings from research to clinical practice, they also
decrease scientific rigor by eliminating the uniformity of randomization to experimental
conditions.
Conclusion
The RCT offers the most rigorous method of examining the causal impact of therapeutic
interventions. After reviewing the essential considerations relevant for matters of design,
procedure, measurement, and data analysis, one recognizes that no one single clinical
trial, even with optimal design and procedures, can alone answer the relevant questions
about the efficacy and effectiveness of therapy. Rather, a series and collection of studies,
with varying designs and approaches, is necessary. The extensions and variations of the
RCT addressed in this chapter address important features relevant to clinical practice
that are not informed by the standard RCT; however, each modification to the standard
RCT decreases the maximized internal validity achieved in the standard RCT. Criteria for
determining evidence-based practice have been proposed (American Psychological
Association, 2005, Chambless & Hollon, 1998), and the quest to identify such treatments
continues. The goal is for the research to be rigorous, with the end goal being to optimize
clinical decision making and practice for those affected by emotional and behavioral
disorders and problems.
Controlled trials play a vital role in facilitating a dialogue between academic clinical
psychology and the public and private sector (e.g., insurance payers, Department of
Health and Human Services, policymakers). The results of controlled clinical evaluations
are increasingly being examined by both professional associations and managed care
organizations with the intent of formulating clinical practice guidelines for cost-effective
care that provides optimized service to those in need. In the absence of compelling data,
there is the risk that psychological practice will be co-opted and exploited in the service
of only profitability and/or cost containment. Clinical trials must retain scientific rigor to
enhance the ability of practitioners to deliver effective treatment procedures to
individuals in need.
References
Addis, M., Cardemil, E. V., Duncan, B., & Miller, S. (2006). Does manualization improve
therapy outcomes? In J. C. Norcross, L. E. Beutler, & R. F. Levant (Eds.), Evidence-based
practices in mental health (pp. 131–160). Washington, DC: American Psychological
Association.
Page 30 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
Addis, M., & Krasnow, A. (2000). A national survey of practicing psychologists’ attitudes
toward psychotherapy treatment manuals. Journal of Consulting and Clinical Psychology,
68, 331–339.
Altman, D. G., & Bland, J. M. (1995). Absence of evidence is not evidence of absence.
British Medical Journal, 311 (7003), 485.
Anderson, E. M., & Lambert, M. J. (2001). A survival analysis of clinical significant change
in outpatient psychotherapy. Journal of Clinical Psychology, 57, 875–888.
Arnold, L. E., Elliott, M., Sachs, L., Bird, H., Kraemer, H. C., Wells, K. C., et al. (2003).
Effects of ethnicity on treatment attendance, stimulant response/dose, and 14-month
outcome in ADHD. Journal of Consulting and Clinical Psychology, 71, 713–727.
Barlow, D. H., Farchione, T. J., Fairholme, C. P., Ellard, K. K., Boisseau, C. L., Allen, L. B.,
& Ehrenreich May, J. T. (2010). Unified protocol for transdiagnostic treatment of
emotional disorders: Therapist guide. New York: Oxford University Press.
Barlow, D. H., & Nock, M. K. (2009). Why can't we be more idiographic in our research?
Perspectives on Psychological Science, 4(1), 19–21.
Baron, R. M., & Kenny, D. A. (1986). The mediator-moderator variable distinction in social
psychological research: Conceptual, strategic, and statistical consideration. Journal of
Personality and Social Psychology, 51, 1173–1182.
Begg, C. B., Cho, M. K., Eastwood, S., Horton, R., Moher, D., Olkin, I., et al. (1996).
Improving the quality of reporting of randomized controlled trials: The CONSORT
statement. Journal of the American Medical Association, 276, 637–639.
Beidel, D. C., Turner, S. M., Sallee, F. R., Ammerman, R. T., Crosby, L. A., & Pathak, S.
(2007). SET-C vs. Fluoxetine in the treatment of childhood social phobia. Journal of the
American Academy of Child and Adolescent Psychiatry, 46, 1622–1632.
Bernal, G., Bonilla, J., & Bellido, C. (1995). Ecological validity and cultural sensitivity for
outcome research: Issues for the cultural adaptation and development of psychosocial
treatments with Hispanics. Journal of Abnormal Child Psychology, 23, 67–82.
Bernal, G., & Scharron-Del-Rio, M. R. (2001). Are empirically supported treatments valid
for ethnic minorities? Toward an alternative approach for treatment research. Cultural
Diversity and Ethnic Minority Psychology, 7, 328–342.
Page 31 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
Blanco, C., Olfson, M., Goodwin, R. D., Ogburn, E., Liebowitz, M. R., Nunes, E. V., &
Hasin, D. S. (2008). Generalizability of clinical trial results for major depression to
community samples: Results from the National Epidemiologic Survey on Alcohol and
Related Conditions. Journal of Clinical Psychiatry, 69, 1276–1280.
Brown, R. A., Evans, M., Miller, I., Burgess, E., & Mueller, T. (1997). Cognitive-behavioral
treatment for depression in alcoholism. Journal of Consulting and Clinical Psychology, 65,
715–726.
Curry, J., Silva, S., Rohde, P., Ginsburg, G., Kratochvil, C., Simons, A., et al. (2011).
Recovery and recurrence following treatment for adolescent major depression. Archives
of General Psychiatry, 68(3), 263–269.
Dawson, R., & Lavori, P. W. (2004). Placebo-free designs for evaluating new mental health
treatments: The use of adaptive treatment strategies. Statistics in Medicine, 23, 3249–
3262.
De Los Reyes, A., & Kazdin, A. E. (2005). Informant discrepancies in the assessment of
childhood psychopathology: A critical review, theoretical framework, and
recommendations for further study. Psychological Bulletin, 131, 483–509.
Dobson, K. S., & Hamilton, K. E. (2002). The stage model for psychotherapy manual
development: A valuable tool for promoting evidence-based practice. Clinical Psychology:
Science and Practice, 9, 407–409.
Page 32 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
Dobson, K. S., Hollon, S. D., Dimidjian, S., Schmaling, K. B., Kohlenberg, R. J., Gallop, R.
J., et al. (2008). Randomized trial of behavioral activation, cognitive therapy, and
antidepressant medication in the prevention of relapse and recurrence in major
depression. Journal of Consulting and Clinical Psychology, 76, 468–477.
Dobson, K. S., & Shaw, B. (1988). The use of treatment manuals in cognitive therapy.
Experience and issues. Journal of Consulting and Clinical Psychology, 56, 673–682.
Flaherty, J. A., & Meaer, R. (1980). Measuring racial bias in inpatient treatment. American
Journal of Psychiatry, 137, 679–682.
Gladis, M. M., Gosch, E. A., Dishuk, N. M., & Crits-Cristoph, P. (1999). Quality of life:
Expanding the scope of clinical significance. Journal of Consulting and Clinical
Psychology, 67, 320–331.
Hedeker, D., & Gibbons, R. D. (1994). A random-effects ordinal regression model for
multilevel analysis. Biometrics, 50, 933–944.
Heiser, N. A., Turner, S. M., Beidel, D. C., & Roberson-Nay, R. (2009). Differentiating
social phobia from shyness. Journal of Anxiety Disorders, 23, 469–476.
Hoagwood, K. (2002). Making the translation from research to its application: The je ne
sais pas of evidence-based practices. Clinical Psychology: Science and Practice, 9, 210–
213.
Hollon, S. D., Garber, J., & Shelton, R. C. (2005). Treatment of depression in adolescents
with cognitive behavior therapy and medications: A commentary on the TADS project.
Cognitive and Behavioral Practice, 12, 149–155.
Page 33 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
psychology literatures. Journal of Consulting and Clinical and Clinical Psychology, 65,
599–610.
Homma-True, R., Greene, B., Lopez, S. R., & Trimble, J. E. (1993). Ethnocultural diversity
in clinical psychology. Clinical Psychologist, 46, 50–63.
Hood, S., D., Potokar, J. P., Davies, S. J., Hince, D. A., Morris, K., Seddon, K. M., et al.
(2010). Dopaminergic challenges in social anxiety disorder: Evidence for dopamine D3
desensitization following successful treatment with serotonergic antidepressants. Journal
of Psychopharmacology, 24(5), 709–716.
Huey, S. J., & Polo, A. J. (2008). Evidence-based psychosocial treatments for ethnic
minority youth. Journal of Clinical Child and Adolescent Psychology, 37, 262–301.
Jaccard, J., & Guilamo-Ramos, V. (2002a). Analysis of variance frameworks in clinical child
and adolescent psychology: Issues and recommendations. Journal of Clinical Child and
Adolescent Psychology, 31, 130–146.
Jaccard, J., & Guilamo-Ramos, V. (2002b). Analysis of variance frameworks in clinical child
and adolescent psychology: Advanced issues and recommendations. Journal of Clinical
Child and Adolescent Psychology, 31, 278–294.
Jacobson, N. S., Follette, W. C., & Revenstorf, D. (1984). Psychotherapy outcome research:
Methods for reporting variability and evaluating clinical significance. Behavior Therapy,
15, 336–352.
Jacobson, N. S., Roberts, L. J., Berns, S. B., & McGlinchey, J. B. (1999). Methods for
defining and determining the clinical significance of treatment effects. Description,
application, and alternatives. Journal of Consulting and Clinical Psychology, 67, 300–307.
Jacobson, N. S., & Traux, P. (1991). Clinical significance: A statistic approach to defining
meaningful change in psychotherapy research. Journal of Consulting and Clinical
Psychology, 59, 12–19.
Jarrett, R. B., Vittengl, J. R., Doyle, K., & Clark, L. A. (2007). Changes in cognitive content
during and following cognitive therapy for recurrent depression: Substantial and
enduring, but not predictive of change in depressive symptoms. Journal of Consulting and
Clinical Psychology, 75, 432–446.
Page 34 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
Jones, B., Jarvis, P., Lewis, J. A., & Ebbutt, A. F. (1996). Trials to assess equivalence: the
importance of rigorous methods. British Medical Journal, 313 (7048), 36–39.
Karver, M., Shirk, S., Handelsman, J. B., Fields, S., Crisp, H., Gudmundsen, G., &
McMakin, D. (2008). Relationship processes in youth psychotherapy: Measuring alliance,
alliance-building behaviors, and client involvement. Journal of Emotional and Behavioral
Disorders, 16, 15–28.
Kazdin, A. E. (2003). Research design in clinical psychology (4th ed.). Boston, MA: Allyn
and Bacon.
Kazdin, A. E., & Bass, D. (1989). Power to detect differences between alternative
treatments in comparative psychotherapy outcome research. Journal of Consulting and
Clinical Psychology, 57, 138–147.
Kendall, P. C., & Beidas, R. S. (2007). Smoothing the trail for dissemination of evidence-
based practices for youth: Flexibility within fidelity. Professional Psychology: Research
and Practice, 38, 13–20.
Kendall, P. C., & Grove, W. (1988). Normative comparisons in therapy outcome. Behavioral
Assessment, 10, 147–158.
Kendall, P. C., & Hedtke, K. A. (2006). Cognitive-behavioral therapy for anxious children
(3rd ed.). Ardmore, PA: Workbook Publishing.
Kendall, P. C., & Hollon, S. D. (1983). Calibrating therapy: Collaborative archiving of tape
samples from therapy outcome trials. Cognitive Therapy and Research, 7, 199–204.
Kendall, P. C., Hollon, S., Beck, A. T., Hammen, C., & Ingram, R. (1987). Issues and
recommendations regarding use of the Beck Depression Inventory. Cognitive Therapy and
Research, 11, 289–299.
Kendall, P. C., Hudson, J. L., Gosch, E., Flannery-Schroeder, E., & Suveg, C. (2008).
Cognitive-behavioral therapy for anxiety disordered youth: A randomized clinical trial
evaluating child and family modalities. Journal of Consulting and Clinical Psychology, 76,
282–297.
Page 35 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
Kendall, P. C., Marrs-Garcia, A., Nath, S. R., & Sheldrick, R. C. (1999). Normative
comparisons for the evaluation of clinical significance. Journal of Consulting and Clinical
Psychology, 67, 285–299.
Kendall, P. C., Safford, S., Flannery-Schroeder, E., & Webb, A. (2004). Child anxiety
treatment: Outcomes in adolescence and impact on substance use and depression at 7.4-
year follow-up. Journal of the Consulting and Clinical Psychology, 72, 276–287.
Kendall, P. C., & Sugarman, A. (1997). Attrition in the treatment of childhood anxiety
disorders. Journal of Consulting and Clinical Psychology, 65, 883–888.
Kendall, P. C., & Suveg, C. (2008). Treatment outcome studies with children: Principles of
proper practice. Ethics and Behavior, 18, 215–233.
Kraemer, H. C., & Kupfer, D. J. (2006). Size of treatment effects and their importance to
clinical research and practice. Biological Psychiatry, 59, 990–996.
Kraemer, H. C., Wilson, G. T., Fairburn, C. G., & Agras, W. S. (2002). Mediators and
moderators of treatment effects in randomized clinical trials. Archives of General
Psychiatry, 59, 877–883.
Laird, N. M., & Ware, J. H. (1982). Random-effects models for longitudinal data.
Biometrics, 38, 963–974.
Leon, A. C., Mallinckrodt, C. H., Chuang-Stein, C., Archibald, D. G., Archer, G. E., &
Chartier, K. (2006). Attrition in randomized controlled clinical trials: Methodological
issues in psychopharmacology. Biological Psychiatry, 59, 1001–1005.
Lin, P., Campbell, D. G., Chaney, E. F., Liu, C., Heagerty, P., Felker, B. L., et al. (2005). The
influence of patient preference on depression treatment in primary care. Annals of
Behavioral Medicine, 30(2), 164–173.
Little, R. J. A., & Rubin, D. (2002). Statistical analysis with missing data (2nd ed.). New
York: Wiley.
Lopez, S. R. (1989). Patient variable biases in clinical judgment: Conceptual overview and
methodological considerations. Psychological Bulletin, 106, 184–204.
Page 36 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
Luborsky, L., Rosenthal, R., Diguer, L., Andrusyna, T. P., Berman, J. S., Levitt, J. T., et al.
(2002). The dodo bird verdict is alive and well—mostly. Clinical Psychology: Science and
Practice, 9(1), 2–12.
Marcus, S. M., Gorman, J., Shea, M. K., Lewin, D., Martinez, J., Ray, S., et al. (2007). A
comparison of medication side effect reports by panic disorder patients with and without
concomitant cognitive behavior therapy. American Journal of Psychiatry, 164, 273–275.
Mason, M. J. (1999). A review of procedural and statistical methods for handling attrition
and missing data. Measurement and Evaluation in Counseling and Development, 32, 111–
118.
Moher, D., Schulz, K. F., & Altman, D. (2001). The CONSORT Statement: Revised
(p. 60)
Molenberghs, G., Thijs, H., Jansen, I., Beunckens, C., Kenward, M. G., Mallinckrodt, C., &
Carroll, R. (2004). Analyzing incomplete longitudinal clinical trial data. Biostatistics, 5,
445–464.
Mufson, L., Dorta, K. P., Wickramaratne, P., Nomura, Y., Olfson, M., & Weissman, M. M.
(2004). A randomized effectiveness trial of interpersonal psychotherapy for depressed
adolescents. Archives of General Psychiatry, 61, 577–584.
Neuner, F., Onyut, P. L., Ertl, V., Odenwald, M., Schauer, E., & Elbert, T. (2008). Treatment
of posttraumatic stress disorder by trained lay counselors in an African refugee
settlement: A randomized controlled trial. Journal of Consulting and Clinical Psychology,
76, 686–694.
Olfson, M., Cherry, D., & Lewis-Fernandez, R. (2009). Racial differences in visit duration
of outpatient psychiatric visits. Archives of General Psychiatry, 66, 214–221.
Page 37 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
Pelham, W. E., Jr., Gnagy, E. M., Greiner, A. R., Hoza, B., Hinshaw, S. P., Swanson, J. M., et
al. (2000). Behavioral versus behavioral and psychopharmacological treatment in ADHD
children attending a summer treatment program. Journal of Abnormal Child Psychology,
28, 507–525.
Perepletchikova, F., & Kazdin, A. E. (2005). Treatment integrity and therapeutic change:
Issues and research recommendations. Clinical Psychology: Science and Practice, 12,
365–383.
Reis, B. F., & Brown, L. G. (2006). Preventing therapy dropout in the real world: The
clinical utility of videotape preparation and client estimate of treatment duration.
Professional Psychology: Research and Practice, 37, 311–316.
Rush, A. J., Fava, M., Wisniewski, S. R., Lavori, P. W., Trivedi, M. H., Sackeim, H. A., et al.
(2004). Sequenced treatment alternatives to relieve depression (STAR*D): rationale and
design. Controlled Clinical Trials, 25(1), 119–142.
Sachs, G. S., Thase, M. E., Otto, M. W., Bauer, M., Miklowitz, D., Wisniewski, S. R., et al.
(2003). Rationale, design, and methods of the systematic treatment enhancement
program for bipolar disorder (STEP-BD). Biological Psychiatry, 53(11), 1028–1042.
Shirk, S. R., Gudmundsen, G., Kaplinski, H., & McMakin, D. L. (2008). Alliance and
outcome in cognitive-behavioral therapy for adolescent depression. Journal of Clinical
Child and Adolescent Psychology, 37, 631–639.
Snowden, L. R. (2003). Bias in mental health assessment and intervention: Theory and
evidence. American Journal of Public Health, 93, 239–243.
Suveg, C., Comer, J. S., Furr, J. M., & Kendall, P. C. (2006). Adapting manualized CBT for a
cognitively delayed child with multiple anxiety disorders. Clinical Case Studies, 5, 488–
510.
Sweeney, M., Robins, M., Ruberu, M., & Jones, J. (2005). African-American and Latino
families in TADS: Recruitment and treatment considerations. Cognitive and Behavioral
Practice, 12, 221–229.
Page 38 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
Treadwell, K., Flannery-Schroeder, E. C., & Kendall, P. C. (1994). Ethnicity and gender in
a sample of clinic-referred anxious children: Adaptive functioning, diagnostic status, and
treatment outcome. Journal of Anxiety Disorders, 9, 373–384.
Treatment for Adolescents with Depression Study (TADS) Team. (2004). Fluoxetine,
cognitive-behavioral therapy, and their combination for adolescents with depression:
Treatment for Adolescents with Depression Study (TADS) randomized controlled trial.
Journal of the American Medical Association, 292, 807–820.
Vanable, P. A., Carey, M. P., Carey, K. B., & Maisto, S. A. (2002). Predictors of participation
and attrition in a health promotion study involving psychiatric outpatients. Journal of
Consulting and Clinical Psychology, 70, 362–368.
Walders, N., & Drotar, D. (2000). Understanding cultural and ethnic influences in
research with child clinical and pediatric psychology populations. In D. Drotar (Ed.),
Handbook of research in pediatric and clinical child psychology (pp. 165–188). New York:
Springer.
Walkup, J. T., Albano, A. M., Piacentini, J., Birmaher, B., Compton, S. N., et al. (2008)
Cognitive behavioral therapy, sertraline, or a combination in childhood anxiety. New
England Journal of Medicine, 359, 1–14.
Waltz, J., Addis, M. E., Koerner, K., & Jacobson, N. S. (1993). Testing the integrity of a
psychotherapy protocol: Assessment of adherence and competence. Journal of Consulting
and Clinical Psychology, 61, 620–630.
Weersing, R. V., & Weisz, J. R. (2002). Community clinic treatment of depressed youth:
Benchmarking usual care against CBT clinical trials. Journal of Consulting and Clinical
Psychology, 70(2), 299–310.
Weisz, J., Donenberg, G. R., Han, S. S., & Weiss, B. (1995). Bridging the gap between
laboratory and clinic in child and adolescent psychotherapy. Journal of Consulting and
Clinical Psychology, 63, 688–701.
Weisz, J. R., Thurber, C. A., Sweeney, L., Proffitt, V. D., & LeGagnoux, G. L. (1997). Brief
treatment of mild-to-moderate child depression using primary and secondary control
enhancement training. Journal of Consulting and Clinical Psychology, 65(4), 703–707.
Weisz, J. R., Weiss, B., & Donenberg, G. R. (1992). The lab versus the clinic: Effects of
child and adolescent psychotherapy. American Psychologist, 47, 1578–1585.
Westbrook, D., & Kirk, J. (2007). The clinical effectiveness of cognitive behaviour
(p. 61)
therapy: Outcome for a large sample of adults treated in routine practice. Behaviour
Research and Therapy, 43, 1243–1261.
Page 39 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).
Westen, D., Novotny, C., & Thompson-Brenner, H. (2004). The empirical status of
empirically supported psychotherapies: Assumptions, findings, and reporting in
controlled clinical trials. Psychological Bulletin, 130, 631–663.
Yeh, M., McCabe, K., Hough, R. L., Dupuis, D., & Hazen, A. (2003). Racial and ethnic
differences in parental endorsement of barriers to mental health services in youth. Mental
Health Services Research, 5, 65–77.
Philip C. Kendall
Jonathan S. Comer
Candice Chow
Page 40 of 40
PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).