0% found this document useful (0 votes)
8 views40 pages

Understanding Randomized Controlled Trials

This chapter discusses the methodological and design considerations essential for conducting randomized controlled trials (RCTs) in clinical psychology. It covers various aspects such as design, procedure, measurement, data analysis, and reporting, while also exploring different types of control conditions and their implications for internal validity. The authors emphasize the importance of rigorous evaluation methods to ensure both scientific rigor and clinical relevance in treatment evaluations.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
8 views40 pages

Understanding Randomized Controlled Trials

This chapter discusses the methodological and design considerations essential for conducting randomized controlled trials (RCTs) in clinical psychology. It covers various aspects such as design, procedure, measurement, data analysis, and reporting, while also exploring different types of control conditions and their implications for internal validity. The authors emphasize the importance of rigorous evaluation methods to ensure both scientific rigor and clinical relevance in treatment evaluations.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

The Randomized Controlled Trial: Basics and Beyond

Oxford Handbooks Online


The Randomized Controlled Trial: Basics and Beyond
Philip C. Kendall, Jonathan S. Comer, and Candice Chow
The Oxford Handbook of Research Strategies for Clinical Psychology
Edited by Jonathan S. Comer and Philip C. Kendall

Print Publication Date: Apr 2013


Subject: Psychology, Clinical Psychology, Psychological Methods and Measurement
Online Publication Date: Aug 2013 DOI: 10.1093/oxfordhb/9780199793549.013.0004

Abstract and Keywords

This chapter describes methodological and design considerations central to the scientific
evaluation of clinical treatment methods via randomized clinical trials (RCTs). Matters of
design, procedure, measurement, data analysis, and reporting are each considered in
turn. Specifically, the authors examine different types of controlled comparisons, random
assignment, the evaluation of treatment response across time, participant selection, study
setting, properly defining and checking the integrity of the independent variable (i.e.,
treatment condition), dealing with participant attrition and missing data, evaluating
clinical significance and mechanisms of change, and consolidated standards for
communicating study findings to the scientific community. After addressing
considerations related to the design and implementation of the traditional RCT, the
authors turn their attention to important extensions and variations of the RCT. These
treatment study designs include equivalency designs, sequenced treatment designs,
prescriptive designs, adaptive designs, and preferential treatment designs. Examples
from the recent clinical psychology literature are provided, and guidelines are suggested
for conducting treatment evaluations that maximize both scientific rigor and clinical
relevance.

Keywords: Randomized clinical trial, RCT, normative comparisons, random assignment, treatment integrity,
equivalency designs, sequenced treatment designs

The randomized controlled trial (RCT)—a group comparison design in which participants
are randomly assigned to treatment conditions—constitutes the most rigorous and
objective methodological design for evaluating therapeutic outcomes. In this chapter we
focus on RCT research strategies that maximize both scientific rigor and clinical
relevance (for consideration of single-case, multiple-baseline, and small pilot trial
designs, see Chapter 3 in this volume). We organize the present chapter around (a) RCT
design considerations, (b) RCT procedural considerations, (c) RCT measurement

Page 1 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

considerations, (d) RCT data analysis, and (e) RCT reporting. We then turn our attention
to extensions and variations of the traditional RCT, which offer various adjustments for
clinical generalizability, while at the same time sacrificing important elements of internal
validity. Although all of the methodological and design ideals presented may not always
be achieved within a single RCT, our discussions provide exemplars of the RCT.

Design Considerations
To adequately assess the causal impact of a therapeutic intervention, clinical researchers
must use control procedures derived from experimental science. In the RCT, the
intervention applied constitutes the experimental manipulation, and thus to have
confidence that an intervention is responsible for observed changes, extraneous factors
must be experimentally “controlled.” The objective is to distinguish intervention effects
from any changes that (p. 41) result from other factors, such as the passage of time,
patient expectancies of change, therapist attention, repeated assessments, and simple
regression to the mean. To maximize internal validity, the clinical researcher must
carefully select control/comparison condition(s), randomly assign participants across
treatment conditions, and systematically evaluate treatment response across time. We
now consider each of these RCT research strategies in turn.

Page 2 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

Selecting Control Condition(s)

Comparisons of participants randomly assigned to different treatment conditions are


essential to control for factors other than treatment. In a “controlled” treatment
evaluation, comparable participants are randomly placed into either the experimental
condition, composed of those who receive the intervention, or a control condition,
composed of those who do not receive the intervention. The efficacy of treatment over
and above the outcome produced by extraneous factors (e.g., the passage of time) can be
determined by comparing prospective changes shown by participants across conditions.

Importantly, not all control conditions are “created equal.” Deciding which form of control
condition to select for a particular study (e.g., no treatment, waitlist, attention-placebo,
standard treatment as usual) requires careful deliberation (see Table 4.1 for recent
examples from the literature). In a no-treatment control condition, comparison
participants are evaluated in repeated assessments, separated by an interval of time
equal in duration to the treatment provided to those in the experimental treatment
condition. Any changes seen in the treated participants are compared to changes seen in
the nontreated participants. When, relative to nontreated participants, the treated
participants show significantly greater improvements, the experimental treatment may be
credited with producing the observed changes. Several important rival hypotheses are
eliminated in a no-treatment design, including effects due to the passage of time,
maturation, (p. 42) spontaneous remission, and regression to the mean. Importantly,
however, other potentially important confounding factors not specific to the experimental
treatment—such as patient expectancies to get better, or meeting with a caring and
attentive clinician—are not ruled out in a no-treatment control design. Accordingly, no-
treatment control conditions may be useful in earlier stages of treatment development,
but to establish broad empirical support for an intervention, more informative control
procedures are preferred.

Page 3 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

Table 4.1 Types of Control Conditions in Treatment Outcome Research

Recent Example in Literature

Control Definition Description Reference


Condition

No- Control participants Adults with anxiety Varley et


treatment are administered symptoms were randomly al. (2011)
control assessments on assigned to a standard self-
repeated occasions, help condition, an
separated by an augmented self-help
interval of time equal condition, or a control
to the length of condition in which they did
treatment. not receive any
intervention.

Waitlist Control participants Adolescents with anxiety Spence et


control are assessed before disorders were randomly al. (2011)
and after a designated assigned to Internet-
duration of time but delivered CBT, face- to-face
receive the treatment CBT, or to a waitlist control
following the waiting group.
period. They may
anticipate change due
to therapy.

Attention- Control participants School-age children with Miller et al.


placebo/ receive a treatment anxiety symptoms were (2011)
nonspecific that involves randomly assigned to
control nonspecific factors either a cognitive-
(e.g., attention, contact behavioral group
with a therapist). intervention or an attention
control in which students
were read to in small
groups.

Page 4 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

Standard Control participants Depressed veterans were Mohr et al.


treatment/ receive an intervention randomly assigned to (2011)
routine care that is the current either telephone-
control practice for treatment administered cognitive-
of the problem under behavioral therapy or
study. standard care through
community-based
outpatient clinics.

A more revealing variant of the no-treatment condition is the waitlist condition. Here,
participants in the waitlist condition expect that after a certain period of time they will
receive treatment, and accordingly may anticipate upcoming changes (which may in turn
affect their symptoms). Changes are evaluated at uniform intervals across the waitlist and
experimental conditions, and if we assume the participants in the waitlist and treatment
conditions are comparable (e.g., comparable baseline symptom severity and gender, age,
and ethnicity distributions), we can then infer that changes in the treated participants
relative to waitlist participants are likely due to the intervention rather than to
expectations of impending change. However, as with no-treatment conditions, waitlist
conditions are of limited value for evaluating treatments that have already been examined
relative to “inactive” conditions.

No-treatment and waitlist conditions in study designs introduce important ethical


considerations, particularly with vulnerable populations (see Kendall & Suveg, 2008). For
ethical purposes, the functioning of waitlist participants must be carefully monitored to
ensure that they are safely able to tolerate the treatment delay. If a waitlist participant
experiences a clinical emergency requiring immediate professional attention during the
waitlist interval, the provision of emergency professional services undoubtedly
compromises the integrity of the waitlist condition. In addition, to maximize internal
validity, the duration of the control condition should be equal to the duration of the
experimental treatment condition to ensure that differences in response across conditions
cannot be attributed simply to differential passages of time. Now suppose a 24-session
treatment takes 6 months to provide—is it ethical to withhold treatment for such a long
wait period (see Bersoff & Bersoff, 1999)? The ethical response to this question varies
across clinical conditions. It may be ethical to incorporate a waitlist design when
evaluating an experimental treatment for obesity, but a waitlist design may be unethical
when evaluating an experimental treatment for suicidal participants. Moreover, with
increasing waitlist durations, the problem of differential attrition arises, which
compromises study interpretation. If attrition rates are higher in a waitlist condition, the
sample in the control condition may be different from the sample in the treatment
condition, and no longer representative of the larger group. For appropriate
interpretation of study results, it is important to recognize that the smaller waitlist group
at the end of the study now represents only patients who could tolerate and withstand a
prolonged waitlist period.

Page 5 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

An alternative to waitlist control condition is the attention-placebo control condition (or


nonspecific treatment condition), which accounts for key effects that might be due simply
to regularly meeting with and getting the attention of a warm and knowledgeable
therapist. For example, in a recent RCT, Kendall and colleagues (2008) randomly assigned
children with anxiety disorders to receive one of two forms of cognitive-behavioral
treatment (CBT; either individual or family CBT) or to a manual-based family education,
support, and attention (FESA) condition. Individual and family-based CBT showed
superiority over FESA in reducing children's principal anxiety diagnosis. Given the
attentive and supportive nature of FESA, it could be inferred that gains associated with
CBT were not likely attributable to “common therapy factors” such as learning about
emotions, receiving support from an understanding therapist, and having opportunities to
discuss the child's difficulties.

Developing and implementing a successful attention-placebo control condition requires


careful deliberation. Attention placebos must credibly instill positive expectations in
participants and provide comparable professional contact, while at the same time they
must be devoid of specific therapeutic techniques hypothesized to be effective. For ethical
purposes, participants must be fully informed of and willing to take a chance on receiving
a psychosocial placebo condition. Even then, a credible attention-placebo condition may
be difficult for therapists to accomplish, particularly if they do not believe that the
treatment will offer any benefit to the participant. Methodologically, it is difficult to
ensure that study therapists share comparable positive expectancies for their attention-
placebo participants as they do for their participants who are receiving more active
treatment (O'Leary & Borkovec, 1978). “Demand characteristics” suggest that when
study therapists predict a favorable treatment response, participants will tend to improve
accordingly (Kazdin, 2003), (p. 43) which in turn affects the interpretability of study
findings. Similarly, whereas participants in an attention-placebo condition may have high
baseline expectations, they may grow disenchanted when no meaningful changes are
emerging. The clinical researcher is wise to assess participant expectations for change
across conditions so that if an experimental treatment outperforms an attention-placebo
control condition, the impact of differential participant expectations across conditions can
be evaluated.

Inclusion of an attention-placebo control condition, when carefully designed, offers


advantages from an internal validity standpoint. Treatment components across conditions
are carefully specified and the clinical researcher maintains tight control over the
differential experiences of participants across conditions. At the same time, such designs
typically compare an experimental treatment to a treatment condition that has been
developed for the purposes of the study and that does not exist in actual clinical practice.
The use of a standard treatment comparison condition (or treatment-as-usual condition)
affords evaluation of an experimental treatment relative to the intervention that is
currently available and being applied. Including standard treatment as the control
condition offers advantages over attention-placebo, waitlist, and no-treatment controls.
Ethical concerns about no-treatment conditions are quelled, and, as all participants
receive care, attrition is likely to be minimized, and nonspecific factors are likely to be
Page 6 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

equated (Kazdin, 2003). When the experimental treatment and the standard care
intervention share comparable durations and participant and therapist expectancies, the
researcher can evaluate the relative efficacy of the interventions.

In a recent example, Mufson and colleagues (2004) randomly assigned depressed


adolescents to interpersonal psychotherapy (IPT-A) or to “treatment as usual” in school-
based mental health clinics. Adolescents treated with IPT-A relative to treatment as usual
showed greater symptom reduction and improved overall functioning. Given this design,
the researchers were able to infer that IPT-A outperformed the existing standard of care
for depressed adolescents in the school settings. Importantly, in standard treatment
comparisons, it is critical that both the experimental treatment and the standard (routine)
treatment are implemented in a high-quality fashion (Kendall & Hollon, 1983).

Random Assignment

To achieve baseline comparability between study conditions, random assignment is


essential. Random assignment in the context of an RCT ensures that every participant has
an equal chance of being assigned to the active treatment condition or to the control
condition(s). Random assignment, however, does not guarantee comparability across
conditions—simply as a result of chance, one resultant group may be different on some
variables (e.g., household income, occupational impairment, comorbidity). Appropriate
statistical tests can be used to evaluate the comparability of participants across treatment
conditions.

Problems arise when random assignment is not incorporated into a group-comparison


design of treatment response. Consider a situation in which participants do not have an
equal chance of being assigned to the experimental and control conditions. For example,
suppose a researcher were to allow depressed participants to elect for themselves
whether to participate in the active treatment or in a waitlist treatment condition. If
participants in the active treatment condition subsequently showed greater symptom
reductions than waitlist participants, the research cannot rule out the possibility that
posttreatment symptom differences could have resulted from prestudy differences
between the participants (e.g., selection bias). Participants who choose not to receive
treatment immediately may not be ready to work on their depression and may be
meaningfully different from those depressed participants who are immediately ready to
work on their symptoms.

Although random assignment does not ensure participant comparability across conditions
on all measures, randomization procedures do rigorously maximize the likelihood of
comparability. An alternative procedure, randomized blocks assignment, or assignment by
stratified blocks, involves matching (arranging) prospective participants in subgroups
that contain participants that are highly comparable on key dimensions (e.g.,
socioeconomic status indicators) and contain the same number of participants as the
number of conditions. For example, if the study requires three conditions—a standard

Page 7 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

treatment, an experimental treatment, and a waitlist condition—participants can be


arranged in matching groups of three so that each trio is highly comparable on
preselected features. Members in each trio are then randomly assigned to one of the
three conditions, in turn increasing the probability that each condition will contain
comparable participants while at the same time retaining a critical randomization
procedure.

Evaluating Treatment Response Across Time

In the RCT, it is essential to evaluate participant functioning on the dependent variables


(p. 44) (e.g., presenting symptoms) prior to treatment initiation. Such pretreatment (or

“baseline”) assessments provide critical data to evaluate between-groups comparability at


treatment outset, as well as within-groups treatment response. Posttreatment
assessments of participants are essential to examine the comparative efficacy of
treatment versus control conditions. Importantly, evidence of acute treatment efficacy
(i.e., improvement immediately upon therapy completion) may not be indicative of long-
term success (maintenance). At posttreatment, treatment effects may be appreciable but
fail to exhibit maintenance at a follow-up assessment. Accordingly, we recommend that
treatment outcome studies systematically include a follow-up assessment. Follow-up
assessments (e.g., 6 months, 9 months, 18 months) are essential to demonstrations of
treatment efficacy and are a signpost of methodological rigor. Maintenance is
demonstrated when a treatment produces results at the follow-up assessment that are
comparable to those found at posttreatment (i.e., improvements from pretreatment and
an absence of detrimental change from posttreatment to follow-up).

Follow-up evaluations can help to identify differential treatment effects of considerable


clinical utility. For example, two treatments may produce comparable effects at the end of
treatment, but one may be more effective in the prevention of relapse (see Anderson &
Lambert, 2001, for demonstration of survival analysis in clinical psychology). When two
treatments show comparable response at posttreatment, yet one is associated with a
higher relapse rate over time, follow-up evaluations provide critical data to support
selection of one treatment over the other. For example, Brown and colleagues (1997)
compared CBT and relaxation training for depression in alcoholism. When considering
the average (mean) days abstinent and drinks per day as dependent variables, measured
at pretreatment and at 3 and 6 months posttreatment, the authors found that although
both treatments produced comparable acute gains, CBT was superior to relaxation
training in maintaining the gains.

Follow-up evaluations can also be used to detect continued improvement—the benefits of


some interventions may accumulate over time and possibly expand to other domains of
functioning. Policymakers and researchers are increasingly interested in expanding
intervention research to consider potential indirect effects on the prevention of secondary
problems. In a long-term (7.4 years) follow-up of individuals treated with CBT for
childhood anxiety (Kendall, Safford, Flannery-Schroeder & Webb, 2004), it was found that

Page 8 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

positive responders relative to less-positive responders had fewer problems with


substance use at the long-term follow-up (see also Kendall & Kessler, 2002). In another
example, participants in the Treatment for Adolescents with Depression Study (TADS)
were followed for 5 years after study entry (Curry et al., 2011). TADS evaluated the
relative effectiveness of fluoxetine, CBT, and their combination in the treatment of
adolescents with major depressive disorder (see Treatment for Adolescents with
Depression Study [TADS] Team, 2004). The Survey of Outcomes Following Treatment for
Adolescent Depression (SOFTAD) was an open, 3.5-year follow-up period extending
beyond the TADS 1-year follow-up period. Initial acute outcomes (measured directly after
treatment) found combination treatment to be associated with significantly greater
outcomes relative to fluoxetine or CBT alone, and CBT showed no incremental response
over pill placebo immediately following treatment (TADS Team, 2004). However, by the 5-
year follow-up, 96 percent of participants, regardless of treatment condition, experienced
remission of their major depressive episode, and 88 percent recovered by 2 years (Curry
et al., 2011). Importantly, gains identified at long-term follow-ups are fully attributable to
the initial treatment only if one determines that participants did not seek or receive
additional treatments during the follow-up interval. As an example, during the TADS 5-
year follow-up interim, 42 percent of participants received psychotherapy and 44 percent
received antidepressant medication (Curry et al., 2011). Appropriate statistical tests are
needed to account for differences across conditions when services have been rendered
during the long-term follow-up interval.

Page 9 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

Multiple Treatment Comparisons

To evaluate the comparative (or relative) efficacy of therapeutic interventions,


researchers use between-groups designs with more than one active treatment condition.
Such between-groups designs offer direct comparisons of one treatment with one or more
alternative active treatments. Importantly, whereas larger effect sizes can be reasonably
expected in evaluations comparing an active treatment to an inactive condition, smaller
differences are to be expected when distinguishing among multiple active treatments.
Accordingly, sample size considerations are influenced by whether the comparison is
between a treatment and a control condition or (p. 45) one treatment versus another
known-to-be-effective treatment (see Kazdin & Bass, 1989; see Chapter 12 in this volume
for a full consideration of statistical power). Research aiming to identify reliable
differences in response between two active treatments will need to evaluate a larger
sample of participants than research comparing an active condition to an inactive
treatment.

In a recent example utilizing multiple active treatment comparisons, Walkup and


colleagues (2008) examined the efficacy of CBT, sertraline, and their combination in a
placebo-controlled trial with children diagnosed with separation anxiety disorder,
generalized anxiety disorder, and/or social phobia. Participants were assigned to CBT,
sertraline, their combination, or a placebo pill for 12 weeks. Patients receiving the three
active treatments all fared significantly better than those in the placebo group, with the
combination of sertraline and CBT yielding the most favorable treatment outcomes.
Specifically, analyses revealed a significant clinical response in roughly 81 percent of
youth treated with a combination of CBT and sertraline, 60 percent of youth treated with
CBT alone, 55 percent of youth treated with sertraline alone, and 24 percent receiving
placebo alone.

As in the above-mentioned designs, it is wise to check the participant comparability


across conditions on important variables (e.g., baseline functioning, prior therapy
experience, socioeconomic indicators, treatment preferences/expectancies) before
continuing with statistical evaluation of the intervention effects. Multiple treatment
comparisons are optimal when each participant is randomly assigned to receive one and
only one treatment. As previously noted, a randomized block procedure, with participants
blocked on preselected variable(s) (e.g., baseline severity), can be used. Comparability
across therapists who are administering the different treatments is also essential.
Therapists conducting each type of treatment should be equivalent in training,
experience, intervention expertise, treatment allegiance, and expectation that the
intervention will be effective. To control for therapist variables, one method has each
study therapist conduct each type of intervention in the study. This method is optimized
when cases are randomly assigned to therapists who are equally expert and favorably
disposed toward each treatment. For example, an intervention test would have reduced
validity if a group of psychodynamic therapists were asked to conduct both a CBT (in
which their expertise is low) and a psychodynamic therapy (in which their expertise is

Page 10 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

high). Stratified blocking offers a viable option to ensure that all treatments are
conducted by comparable therapists. It is wise to gather data on therapist variables (e.g.,
expertise, experience, allegiance) and examine their relationships to participant
outcomes.

For proper evaluation, intervention procedures across treatments must be equated for
key variables such as (a) duration; (b) length, intensity, and frequency of contacts with
participants; (c) credibility of treatment rationale; (d) treatment setting; and (e) degree of
involvement of persons significant to the participant. These factors may be the basis for
two alternative therapies (e.g., conjoint vs. individual marital therapy). In such cases, the
nonequated feature constitutes an experimentally manipulated variable rather than a
factor to control.

What is the best method of measuring change when two alternative treatments are being
compared? Importantly, measures should cover the range of symptoms and functioning
targeted for change, tap costs and potential negative side effects, and be unbiased with
respect to the alternate interventions. Assessments should not be differentially sensitive
to one treatment over another. Treatment comparisons will be misleading if measures are
not equally sensitive to the types of changes that most likely result from each intervention
type.

Special issues are presented in comparisons of psychological and psychopharmacological


treatments (e.g., Beidel et al., 2007; Dobson et al., 2008; Marcus et al., 2007; MTA
Cooperative Group, 1999; Pediatric OCD Treatment Study Team, 2004; Walkup et al,
2008). For example, when and how should placebo medications be used in comparison to
or with psychological treatment? How should expectancy effects be addressed? How
should differential attrition be handled statistically and/or conceptually? How should
inherent differences in professional contact across psychological and pharmacological
interventions be addressed? Follow-up evaluations become particularly important after
acute treatment phases are discontinued. Psychological treatment effects may persist
after treatment, whereas the effects of medications may not persist upon medication
discontinuation. (Interested readers are referred to Hollon, 1996; Hollon & DeRubeis,
1981; Jacobson & Hollon, 1996a, 1996b, for thorough consideration of these issues.)

Procedural Considerations
We now address key RCT procedural considerations, including (a) sample selection, (b)
study (p. 46) setting, (c) defining the independent variable, and (d) checking the integrity
of the independent variable.

Page 11 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

Sample Selection

Selecting a sample to best represent the clinical population of interest requires careful
deliberation. A selected sample refers to a sample of participants who may require
treatment but who may otherwise only approximate clinically disordered persons. By
contrast, RCTs optimize external validity when treatments are applied and evaluated with
actual treatment-seeking patients. Consider a study investigating the effects of a
treatment on social anxiety disorder. The researcher could use (a) a sample of patients
diagnosed with social anxiety disorder via structured diagnostic interviews (genuine
clinical sample), (b) a sample consisting of a group of individuals who self-report shyness
(analogue sample), or (c) a sample of socially anxious persons after excluding cases with
depressed mood and/or substance use (highly select sample). This last sample may meet
full diagnostic criteria for social anxiety disorder but are nevertheless highly selected.

From a feasibility standpoint, clinical researchers may find it easier to recruit analogue
samples relative to genuine clinical samples, and such samples may afford a greater
ability to control various conditions and minimize threats to internal validity. At the same
time, analogue and select samples compromise external validity—these individuals are
not necessarily comparable to patients seen in typical clinical practice (and may not
qualify as an RCT). With respect to social anxiety disorder, for instance, one could
question whether social anxiety disorder in genuine clinical populations compares
meaningfully to self-reported shyness (see Heiser, Turner, Beidel, & Roberson-Nay, 2009).
When deciding whether to use clinical, analogue, or select samples, the researcher needs
to consider how the study results will be interpreted and generalized. Regrettably,
nationally representative data show that standard exclusion criteria set for clinical
treatment studies exclude up to 75 percent of affected individuals in the general
population who have major depression (Blanco, Olfson, Goodwin, et al., 2008).

Researchers must consider patient diversity when deciding which samples to study.
Research supporting the efficacy of psychological treatments has historically been
conducted with predominantly European-American samples, although this is rapidly
changing (see Huey & Polo, 2008). Although racially and ethnically diverse samples may
be similar in many ways to single-ethnicity samples, one can question the extent to which
efficacy findings from predominantly European-American samples can be generalized to
ethnic-minority samples (Bernal, Bonilla, & Bellido, 1995; Bernal & Scharron-Del-Rio,
2001; Hall, 2001; Olfson, Cherry, & Lewis-Fernandez, 2009; Sue, 1998). Investigations
have also addressed the potential for bias in diagnoses and in the provision of mental
health services to ethnic-minority patients (e.g., Flaherty & Meaer, 1980; Homma-True,
Green, Lopez, & Trimble, 1993; Lopez, 1989; Snowden, 2003).

A simple rule is that the research sample should reflect the broad population to which the
study results are to be generalized. To generalize to a single-ethnicity group, one must
study a single-ethnicity sample. To generalize to a diverse population, one must study a
diverse sample, as most RCTs strive to accomplish. Barriers to care must be reduced and

Page 12 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

outreach efforts employed to inform minorities of available services (see Sweeney, Robins,
Ruberu, & Jones, 2005; Yeh, McCabe, Hough, Dupuis, & Hazen, 2003) and include them
in the research. Walders and Drotar (2000) provide guidelines for recruiting and working
with ethnically diverse samples.

After the fact, appropriate statistical analyses can examine potential differential
outcomes (see Arnold et al., 2003; Treadwell, Flannery-Schroeder, & Kendall, 1994).
Although grouping and analyzing research participants by racial or ethnic status is a
common analytic approach, this approach is simplistic because it fails to address
variations in each patient's degree of ethnic identity. It is often the degree to which an
individual identifies with an ethnocultural group or community, and not simply his or her
ethnicity itself, that may moderate response to treatment. For further consideration of
this important issue, the reader is referred to Chapter 21 in this volume.

Study Setting

Some have questioned whether outcomes found at select research centers can transport
to clinical practice settings, and thus the question of whether an intervention can be
transported to other service settings requires independent evaluation (Southam-Gerow,
Ringeisen, & Sherrill, 2006). It is not sufficient to demonstrate treatment efficacy within
a narrowly defined sample in a highly selective setting. One should study, rather than
assume, that a treatment found to be efficacious within a research (p. 47) clinical setting
will be efficacious in a clinical service setting (see Hoagwood, 2002; Silverman, Kurtines,
& Hoagwood, 2004; Southam-Gerow et al., 2006; Weisz, Donenberg, Han, & Weiss, 1995;
Weisz, Weiss, & Donenberg, 1992). Closing the gap between RCTs and clinical practice
requires transporting effective treatments (getting “what works” into practice) and
identifying additional research into those factors that may be involved in successful
transportation (e.g., patient, therapist, researcher, service delivery setting; see Kendall &
Southam-Gerow, 1995; Silverman et al., 2004). Methodological issues relevant to the
conduct of research evaluating the transportability of treatments to “real-world” settings
can be found in Chapter 5 in this volume.

Defining the Independent Variable

Proper treatment evaluation necessitates that the treatment must be adequately


described and detailed in order to replicate the evaluation in another setting, or to be
able to show and teach others how to conduct the treatment. Treatment manuals achieve
the required description and detail of the treatment. Treatment manuals enhance internal
validity and treatment integrity and allow for comparison of treatments across formats
and contexts, while at the same time reducing potential confounds (e.g., differences in
the amount of clinical contact, type and amount of training). Therapist manuals facilitate

Page 13 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

training and contribute meaningfully to replication (Dobson & Hamilton, 2002; Dobson &
Shaw, 1988).

The merits of manual-based treatments are not universally agreed upon. Debate has
ensued regarding the appropriate use of manual-based treatments versus a more variable
approach typically found in clinical practice (see Addis, Cardemil, Duncan, & Miller, 2006;
Addis & Krasnow, 2000; Westen, Novotny, & Thompson-Brenner, 2004). Some have
argued that treatment manuals limit therapist creativity and place restrictions on the
individualization that the clinicians use (see also Waltz, Addis, Koerner, & Jacobson, 1993;
Wilson, 1995). Indeed, some therapy manuals may appear “cookbook-ish,” and some lack
attention to the clinical sensitivities needed for implementation and individualization, but
our experience and data suggest that this is not the norm. An empirical evaluation from
our laboratory found that the use of a manual-based treatment for child anxiety disorders
(Kendall & Hedtke, 2006) did not restrict therapist flexibility (Kendall & Chu, 1999).
Although it is not the goal of manual-based treatments to have clinicians perform
treatment in a rigid manner, this misperception has restricted some clinicians’ openness
to manual-based interventions (Addis & Krasnow, 2000).

Effective use of manual-based treatments must be preceded by adequate training


(Barlow, 1989). Clinical professionals cannot become proficient in the administration of
therapy simply by reading a manual. Interactive training, flexible application, and
ongoing clinical supervision are essential to ensure proper conduct of manual-based
therapy: The goal has been referred to as “flexibility within fidelity” (Kendall & Beidas,
2007).

Several modern treatment manuals allow the therapist to attend to each patient's specific
circumstances, clinical needs, concerns, and comorbid diagnoses without deviating from
the core treatment strategies detailed in the manual. The goal is to include provisions for
standardized implementation of therapy while using a personalized case formulation
(e.g., see Suveg, Comer, Furr, & Kendall, 2006). Importantly, use of manual-based
treatments does not eliminate the potential for differential therapist effects. Researchers
examine therapist variables within the context of manual-based treatments (e.g.,
therapeutic relationship-building behaviors, flexibility, warmth) that may relate to
treatment outcome (Creed & Kendall, 2005; Karver et al., 2008; Shirk et al., 2008; see
also Chapter 9 in this volume for a full consideration of designing, conducting, and
evaluating therapy process research).

Checking the Integrity of the Independent Variable

Careful checking of the manipulated variable is required in any rigorous experimental


research. In the RCT, the manipulated variable is typically treatment or a key
characteristic of treatment. By experimental design, all participants are not treated the
same. However, just because the study has been so designed does not guarantee that the
independent variable (treatment) has been implemented as intended. In the course of a

Page 14 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

study—whether due to insufficient therapist training, therapist variables, lack of manual


specification, inadequate therapist monitoring, participant demand characteristics, or
simple error variance—the treatment that was assigned may not in fact be the treatment
that was provided (see also Perepletchikova & Kazdin, 2005).

To help ensure that the treatments are indeed implemented as intended, it is wise to
require that a treatment plan be followed, that therapists are (p. 48) carefully trained,
and that sufficient supervision is available throughout. The researcher is wise to conduct
an independent check on the manipulation. For example, treatment sessions are recorded
so that an independent rater can listen to and/or watch the recordings and provide
quantifiable judgments regarding key characteristics of the treatment. Such a
manipulation check provides the necessary assurance that the described treatment was
indeed provided as intended. Digital audio and video recordings are inexpensive, can be
used for subsequent training, and can be analyzed to answer key research questions.
Therapy session recordings evaluated in RCTs not only provide a check on the treatment
within each separate study but also allow for a check on the comparability of treatments
provided across studies. That is, the therapy provided as CBT in one researcher's RCT
could be checked to assess its comparability to other teams’ CBT.

A recently completed clinical trial from our research program comparing two active-
treatment conditions for childhood anxiety disorders against an active attention control
condition (Kendall et al., 2008) illustrates a procedural plan for integrity checks. First, we
developed a checklist of the strategies and content called for in each session by the
respective treatment manuals. A panel of expert clinicians served as independent raters
who used the checklists to rate randomly selected video segments from randomly
selected cases. The panel of raters was trained on nonstudy cases until they reached an
interrater reliability of Cohen's κ ≥ .85. After ensuring reliability, the panel used the
checklists to assess whether the appropriate content was covered for randomly selected
segments that were representative of all sessions, conditions, and therapists. For each
coded session, we computed an integrity ratio corresponding to the number of checklist
items covered by the therapist divided by the total number of items that should have been
included. Integrity check results indicated that across the conditions, 85 to 92 percent of
intended content was in fact covered.

It is also wise for the RCT researcher to evaluate the quality of treatment provided. A
therapist may strictly adhere to a treatment manual and yet fail to administer the
treatment in an otherwise competent manner, or he or she may administer therapy while
significantly deviating from the manual. In both cases, the operational definition of the
independent variable (i.e., the treatment manual) has been violated, treatment integrity
impaired, and replication rendered impossible (Dobson & Shaw, 1988). When a treatment
fails to demonstrate expected gains, one can examine the adequacy with which the
treatment was implemented (see Hollon, Garber, & Shelton, 2005). It is also of interest to
investigate potential variations in treatment outcome that may be associated with
differences in the quality of the treatment provided (Garfield, 1998; Kendall & Hollon,
1983). Expert judges are needed to make determinations of differential quality prior to

Page 15 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

the examination of differential outcomes for high- versus low-quality therapy


implementation (see Waltz et al., 1993). McLeod and colleagues (in press) provide a
description of procedural issues in the conduct of quality assurance and treatment
integrity checks.

Page 16 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

Measurement Considerations

Assessing the Dependent Variable(s)

No single measure can serve as the sole indicator of participants’ treatment-related


gains. Rather, a variety of methods, measures, data sources, and sampling domains (e.g.,
symptoms, distress, functional impairment, quality of life) are used to assess outcomes. A
rigorous treatment RCT will consider using assessments of participant self-report;
participant test/task performance; therapist judgments and ratings; archival or
documentary records (e.g., health care visits and costs, work and school records);
observations by trained, unbiased, blinded observers; rating by significant people in the
participant's life; and independent judgments by professionals. Outcomes are more
compelling when observed by independent (blind) evaluators than when based solely on
the therapist's opinion or the participant's self-reports.

Collecting data on variables of interest from multiple reporters (e.g., treatment


participant, family members, peers) can be particularly important when assessing
children and adolescents. Such a multi-informant strategy is critical as features of
cognitive development may compromise youth self-reports, and children may simply offer
what they believe to be the desired responses. And so in RCTs with youth, collecting
additional data from important adults in children's lives who observe them across
different settings (e.g., parents, teachers) is essential. Importantly, however, because
emotions and mood are partially internal phenomena, some symptoms may be less known
to parents and teachers, and some observable symptoms may occur in situations outside
the home or school. Accordingly, an inherent dilemma with a multi-informant assessment
strategy is that discrepancies among informants are to be expected (Comer & Kendall,
2004). Research shows low to moderate concordance rates (p. 49) among informants in
the assessment of youth (De Los Reyes & Kazdin, 2005), with particularly low agreement
among child internalizing symptoms (Comer & Kendall, 2004).

A multimodal strategy relies on multiple inquiries to evaluate an underlying construct of


interest. For example, assessing family functioning may include family members
completing self-report forms on their perceptions of relationships in the family, as well as
conducting structured behavioral observations of family members interacting to be coded
by independent raters. Statistical packages can integrate data obtained from multimodal
assessment strategies (see Chapter 16 in this volume). The increasing availability of
handheld communication devices and personal digital assistants allows researchers to
incorporate experience sampling methodology (ESM), in which people report on their
emotions and behavior in real-world situations (in situ). ESM data provide naturalistic
information on patterns in day-to-day functioning (see Chapter 11 in this volume).

Page 17 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

In a well-designed RCT, multiple targets are assessed to determine treatment evaluation.


For example, one can measure the presence of a diagnosis, overall well-being,
interpersonal skills, self-reported mood, family functioning, occupational impairment, and
health-related quality of life. No one target captures all, and using multiple targets
facilitates an examination of therapeutic changes when changes occur, and the absence of
change when interventions are less beneficial. However, inherent in a multiple-domain
assessment strategy is the fact that it is rare that a treatment produces uniform effects
across assessed domains. Suppose a treatment, relative to a control condition, improves
participants’ severity of anxiety, but not their overall quality of life. In an RCT designed to
evaluate improved anxiety symptoms and quality of life, should the treatment be deemed
efficacious if only one of two measures showed gains? The Range of Possible Changes
model (De Los Reyes & Kazdin, 2006) calls for a multidimensional conceptualization of
intervention change. In this spirit, we recommend that RCT researchers be explicit about
the domains of functioning expected to change and the relative magnitude of such
expected changes. We also caution consumers of the treatment outcome literature
against simplistic dichotomous appraisals of treatments as efficacious or not.

Data Analysis
Data analysis is an active process through which we extract useful information from the
data we have collected in ways that allow us to make statistical inferences about the
larger population that a given sample was selected to represent. Data do not “speak” for
themselves. Although a comprehensive statistical discussion about RCT data analysis is
beyond the present scope (the reader is referred to Jaccard & Guilamo-Ramos, 2002a,
2002b; Kraemer & Kupfer, 2006; Kraemer, Wilson, Fairburn, & Agras, 2002; and Chapters
14 and 16 in this volume) in this section, we discuss three areas that merit consideration
in the context of RCT data analysis: (a) addressing missing data and attrition, (b)
assessing clinical significance, and (c) evaluating mechanisms of change (i.e., mediators
and moderators).

Addressing Missing Data and Attrition

Not every participant assigned to treatment actually completes participation in an RCT. A


loss of research participants (attrition) may occur just after randomization, during
treatment, prior to posttreatment evaluation, or during the follow-up interval.
Increasingly, researchers are evaluating predictors and correlates of attrition to elucidate
the nature of treatment dropout, to understand treatment tolerability, and to enhance the
sustainability of mental health services in the community (Kendall & Sugarman, 1997;
Reis & Brown, 2006; Vanable, Carey, Carey, & Maisto., 2002). However, from a research
methods standpoint, attrition can be problematic for data analysis, such as when there

Page 18 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

are large numbers of noncompleters or when attrition varies across conditions (Leon et
al., 2006; Molenberghs et al., 2004).

Regardless of how diligently researchers work to prevent attrition, data will likely be lost.
Although attrition rates vary across RCTs and treated clinical populations, Mason (1999)
estimated that most researchers can expect roughly 20 percent of their sample to
withdraw or be removed from a study prior to completion. To address this matter,
researchers can conduct and report two sets of analyses: (a) analyses of outcomes for
treatment completers and (b) analyses of outcomes for all participants who were included
at the time of randomization (i.e., the intent-to-treat sample). Treatment-completer
analyses involve the evaluation of only those who actually completed treatment and
examine what the effects of treatment are when someone completes a full treatment
course. Treatment refusers, treatment dropouts, and participants who fail to adhere to
treatment schedules are not included in such analyses. Reports of such treatment
outcomes may be somewhat elevated because they represent (p. 50) the results for only
those who adhered to and completed the treatment. A more conservative approach to
addressing missing data, intent-to-treat analysis, entails the evaluation of outcomes for all
participants involved at the point of randomization. As proponents of intent-to-treatment
analyses we say, “once randomized, always analyzed.”

Careful consideration is required when selecting an appropriate analytic method to


handle missing endpoint data because different methods can produce different outcomes
(see Chapter 19 in this volume). Researchers address missing endpoint data via one of
several ways: (a) last observation carried forward (LOCF), (b) substituting pretreatment
scores for posttreatment scores, (c) multiple imputation methods, and (d) mixed-effects
models. LOCF analysis assumes that participants who drop out remain constant on the
outcome variable from their last assessed point through the posttreatment evaluation. For
example, if a participant drops out at week 6, the data from the week 5 assessment would
be substituted for his or her missing posttreatment assessment data. The LOCF approach
can be problematic however, as the last data collected may not be representative of the
dropout participant's ultimate progress or lack of progress at posttreatment, given that
participants may change after dropping out of treatment. The use of pretreatment data as
posttreatment data (a conservative and not recommended method) simply inserts
pretreatment scores for cases of attrition as posttreatment scores, assuming that
participants who drop out make no change from their initial baseline state. Critics of
pretreatment substitution and LOCF argue that these crude methods introduce
systematic bias and fail to take into account the uncertainty of posttreatment functioning
(see Leon et al., 2006). More current missing data imputation methods are grounded in
statistical theory and incorporate the uncertainty regarding the true value of the missing
data. Multiple imputation methods impute a range of values for the missing data,
incorporating the uncertainty of the true values of missing data and generating a number
of nonidentical datasets (Little & Rubin, 2002). After the researcher conducts analyses on

Page 19 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

the nonidentical datasets, the results are pooled and the resulting variability addresses
the uncertainty of the true value of the missing data.

Mixed-effects modeling, which relies on linear and/or logistic regression to address


missing data in the context of random (e.g., participant) and fixed effects (e.g., treatment,
age, sex) (see Hedeker & Gibbons, 1994, 1997; Laird & Ware, 1982), can be used (see
Neuner et al., 2008, for an example). Mixed-effects modeling may be particularly useful in
addressing missing data if numerous assessments are collected across a treatment trial
(e.g., weekly symptom reports are collected).

Despite sophisticated data analytic approaches to accounting for missing data, we


recommend that researchers attempt to contact noncompleting participants and re-
evaluate them at the time when the treatment would have ended. This method accounts
for the passage of time, as both dropouts and treatment completers are evaluated over
time periods of the same duration, and minimizes any potential error introduced by
statistical imputation and modeling approaches to missing data. If this method is used,
however, it is important to determine whether dropouts sought and/or received
alternative treatments in the interim.

Assessing the Persuasiveness of Therapeutic Outcomes

Data produced by RCTs are submitted to statistical tests of significance. Mean scores for
participants in each condition are compared, within-group and between-group variability
is considered, and the analysis produces a numerical figure, which is then checked
against critical values. Statistical significance is achieved when the magnitude of the
mean difference is beyond that which could have resulted by chance alone
(conventionally defined as p 〈 .05). Tests of statistical significance are essential as they
inform us that the degree of change was likely not due to chance.

Importantly, statistical tests alone do not provide evidence of clinical significance. Sole
reliance on statistical significance can lead to perceiving treatment gains as potent when
in fact they may be clinically insignificant. For example, imagine that the results of a
treatment outcome study demonstrate that mean Beck Depression Inventory (BDI) scores
are significantly lower at posttreatment than pretreatment. An examination of the means,
however, reveals only a small but reliable shift from a mean of 29 to a mean of 26. With
larger sample sizes, this difference may well achieve statistical significance at the
conventional p 〈 .05 level (i.e., over 95 percent chance that the finding is not due to
chance alone), yet perhaps be of limited practical significance. Both before and after
treatment, the scores are within the range considered indicative of clinical levels of
depressive distress (Kendall, Hollon, Beck, Hammen, & Ingram, 1987), and such a small
magnitude of change may have little effect on a person's (p. 51) life impairment (Gladis,
Gosch, Dishuk, & Crits-Christoph, 1999). Conversely, statistically meager results may

Page 20 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

disguise meaningful changes in participant functioning. As Kazdin (1999) put it,


sometimes a little can mean a lot, and vice versa.

Clinical significance refers to the persuasiveness or meaningfulness of the magnitude of


change (Kendall, 1999). Whereas statistical significance tests address the question,
“Were there treatment-related changes?” tests of clinical significance address the
question, “Were treatment-related changes meaningful and convincing?” Specifically, this
can be made operational as changes on a measure of the presenting problem (e.g.,
anxiety symptoms) that result in the participants being returned to within normal limits
on that same measure. Several approaches for measuring clinically significant change
have been developed, two of which are normative sample comparison and reliable change
index.

Normative comparisons (Kendall & Grove, 1988; Kendall, Marrs-Garcia, Nath, &
Sheldrick, 1999) are conducted in several steps. First, the researcher selects a normative
group for posttreatment comparison. Given that several well-established measures
provide normative data (e.g., the BDI, the Child Behavior Checklist), investigators may
choose to rely on these preexisting normative samples. However, when normative data do
not exist, or when the treatment sample is qualitatively different on key factors (e.g.,
socioeconomic status indicators, age), it may be necessary to collect one's own normative
data. In a typical RCT, when using statistical tests to compare groups, the investigator
assumes equivalency across groups (null hypothesis) and aims to find that they are not
(alternate hypothesis). However, when the goal is to show that treated individuals are
equivalent to “normal” individuals on some factor (i.e., are indistinguishable from
normative comparisons), traditional hypothesis-testing methods are inadequate. One uses
an equivalency testing method to circumvent this problem (Kendall, Marrs-Garcia, et al.,
1999) that examines whether the difference between the treatment and normative groups
is within some predetermined range. When used in conjunction with traditional
hypothesis testing, this approach allows conclusions to be drawn about the equivalency of
groups (see, e.g., Jarrett, Vittengl, Doyle, & Clark, 2007; Pelham et al., 2000; Westbrook &
Kirk, 2007, for examples of normative comparisons), thus testing that posttreatment data
are within a normative range on the measure of interest. For example, Weisz and
colleagues (1997) utilized normative comparisons in a trial in which elementary school
children with mild to moderate symptoms of depression were randomly assigned either to
a Primary and Secondary Control Enhancement Training (PASCET) program or to a no-
treatment control group. Normative comparisons were used to determine whether
participants’ scores on two depression measures, the Children's Depression Inventory
and the Revised Children's Depression Rating Scale, fell within one standard deviation
above elementary school norm groups at pretreatment, posttreatment, and 9-month
follow-up time points. Utilizing normative comparisons allowed the authors to conclude
that children who had received the treatment intervention were more likely to fall within
the normal range on depression measures than children in the no-treatment control
condition.

Page 21 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

The Reliable Change Index (RCI; Jacobson, Follette, & Revenstorf, 1984; Jacobson &
Traux, 1991) is another popular method to examine clinically significant change. The RCI
entails calculating the number of participants moving from a dysfunctional to a normative
range. Specifically, the research calculates a difference score (posttreatment minus
pretreatment) divided by the standard error of measurement (calculated based on the
reliability of the measure). The RCI is influenced by the magnitude of change and the
reliability of the measure. The RCI has been used in RCT research, although its
originators point out that it has at times been misapplied (Jacobson, Roberts, Berns, &
McGlinchey, 1999). When used in conjunction with reliable measures and appropriate
cutoff scores, it can be a valuable tool for assessing clinical significance.

Evaluating Change Mechanisms

The RCT researcher is often interested in identifying (a) the conditions that dictate when
a treatment is more or less effective and (b) the processes through which a treatment
produces change. Addressing such issues necessitates the specification of moderator and
mediator variables (Baron & Kenny, 1986; Holmbeck, 1997; Kraemer et al., 2002). A
moderator is a variable that delineates the conditions under which a given treatment is
related to an outcome. Conceptually, moderators identify on whom and under what
circumstances treatments have different effects (Kraemer et al., 2002). A moderator is
functionally a variable that influences either the strength or direction of a relationship
between an independent variable (treatment) and a dependent variable (outcome). For
example, if in an RCT the experimental treatment was found (p. 52) to be more effective
with men than with women, but this gender effect was not found in response to the
control treatment, then gender would be considered a moderator of the association
between treatment and outcome. Treatment moderators help clarify for consumers of the
treatment outcome literature which patients might be most responsive to which
treatments, and for which patients alternative treatments might be sought. Importantly,
when a variable broadly predicts outcome across all treatment conditions in an RCT,
conceptually that variable is simply a predictor, and not a moderator (see Kraemer et al.,
2002).

On the other hand, a mediator is a variable that serves to explain the process by which a
treatment affects an outcome. Conceptually, mediators identify how and why treatments
take effect (Kraemer et al., 2002). The mediator effect reveals the mechanism through
which the independent variable (e.g., treatment) is related to outcome (e.g., treatment-
related changes). Accordingly, mediational models are inherently causal models, and in
the context of an RCT, significant meditational pathways inform us about causal
relationships. If an effective treatment for child externalizing problems was found to have
an impact on parenting behavior, which in turn was found to have a significant influence
on child externalizing behavior, then parent behavior would be considered to mediate the
treatment-to-outcome relationship (provided certain statistical criteria were met; see

Page 22 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

Holmbeck, 1997). Specific statistical methods used to evaluate the presence of treatment
moderation and mediation can be found elsewhere (see Chapter 15 in this volume).

Reporting the Results


Communicating study findings to the scientific community constitutes the final stage of
conducting an RCT. A quality report will present outcomes in the context of previous
related work (e.g., discussing how the findings build on and support previous work;
discussing the ways in which findings are discrepant from previous work and why this
may be the case), as well as consider shortcomings and limitations that can direct future
empirical efforts and theory in the area. To prepare a well-constructed report, the
researcher must provide all of the relevant information for the reader to critically
appraise, interpret, and/or replicate study findings. It has been suggested that there have
been inadequacies in the reporting of RCTs (see Westen et al., 2004). Inadequacies in the
reporting of RCTs can result in bias in estimating treatment effectiveness (Moher, Schulz,
& Altman, 2001). An international group of epidemiologists, statisticians, and journal
editors developed a set of consolidated standards for reporting trials (i.e., CONSORT; see
Begg et al., 1996) in order to maximize transparency in RCT reporting. CONSORT
guidelines consist of a 22-item checklist of study features that can bias estimates of
treatment effects, or that are critical to judging the reliability or relevance of RCT
findings, and consequently should be included in a comprehensive research report. A
quality report will address each of these 22 items. Importantly, participant flow should be
characterized at each research stage. The researcher reports the specific numbers of
participants who were randomly assigned to each treatment condition, who received
treatments as assigned, who participated in posttreatment evaluations, and who
participated in follow-up evaluations. It has become standard practice for scientific
journals to require a CONSORT flow diagram. See Figure 4.1 for an example of a flow
diagram used in reporting to depict participant flow at each stage of an RCT.

Next, the researcher must decide where to submit the report. We recommend that
researchers consider submitting RCT findings to peer-reviewed journals only. Publishing
RCT outcomes in a refereed journal (i.e., one that employs the peer-review process)
signals that the work has been accepted and approved for publication by a panel of
impartial and qualified reviewers (i.e., independent researchers knowledgeable in the
area but not involved with the RCT). Consumers should be highly cautious of RCTs
published in journals that do not place manuscript submissions through a rigorous peer-
review process. Although the peer-review process slows down the speed with which one
is able to communicate RCT results, much to the chagrin of the excited researcher who
just completed an investigation, it is nonetheless one of the indispensable safeguards that
we have to ensure that our collective knowledge base is drawn from studies meeting
acceptable standards. Typically, the review process is “blind,” meaning that the authors

Page 23 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

of the article do not know the identities of the peer reviewers who are considering their
manuscript. Many journals now employ a double-blind peer-review process in which the
identities of study authors are also not known to the peer reviewers.

Extensions and Variations of the RCT


Thus far, we have
addressed considerations
related to the design and
implementation of the
standard RCT. We now
turn our attention to
important extensions and
variations of the RCT.
These (p. 53) treatment
study designs—which
Click to view larger include equivalency
Figure 4.1 Example of flow diagram used in designs and sequenced
reporting to depict participant flow at each study treatment designs—
stage.
address key questions that
cannot be adequately
addressed with the traditional RCT. We discuss each of these designs in turn and note
some of their strengths and limitations.

Equivalency Designs

As varied therapeutic treatment interventions are becoming readily available, research is


needed to determine their relative efficacy. Sometimes, the researcher is not interested in
evaluating the superiority of one treatment over another, but rather that a treatment
produces comparable results to another treatment that differs in key ways. For example,
a researcher may be interested in determining whether an individual treatment protocol
can yield comparable results when administered in a group format. The researcher may
not hold a hypothesis that the group format would produce superior outcomes, but if it
could be demonstrated that the two treatments produce equivalent outcomes, the group
treatment may nonetheless be preferred due to the efficiency of treating multiple patients
in the same amount of time. In another example, a researcher may be interested in
comparing a cross-diagnostic treatment (e.g., one that flexibly addresses any of the
common child anxiety disorders—separation anxiety disorder, social anxiety disorder, or
generalized anxiety disorders) relative to single-disorder treatment protocols for those
specific disorders. The researcher may not hold a hypothesis that the cross-diagnostic

Page 24 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

protocol produces superior outcomes over single-disorder treatment protocols, but if it


could be demonstrated that it produces equivalent outcomes, parsimony would suggest
that the cross-diagnostic protocol would be the most efficient to broadly disseminate.

In equivalency research designs, significance tests are utilized to determine the


equivalence of treatment outcomes observed across multiple active treatments. While a
standard significance test would be used in a comparative trial, such a test could not
conclude equivalency between treatments because a nonsignificant difference does not
necessarily signify equivalence (Altman & Bland, 1995). In an equivalency design, a
confidence interval is established to define a range of points within which (p. 54)
treatments may be deemed essentially equivalent (Jones, Jarvis, Lewis, & Ebbutt, 1996).
To minimize bias, this confidence interval must be determined prior to data collection.

Barlow and colleagues, for example, are currently testing the efficacy of a transdiagnostic
treatment (Unified Protocol for Emotional Disorders; Barlow, Farchione, Fairholme,
Ellard, Boisseau, et al., 2010) for anxiety disorders. The proposed analyses include a
rigorous comparison of the Unified Protocol (UP) against single-diagnosis psychological
treatment protocols (SDPs). Statistical equivalence will be used to test the hypothesis
that the UP is statistically equivalent to SDPs. An a priori confidence interval around
change in the clinical severity rating (CSR) will be utilized to evaluate statistical
equivalence among treatments. The potential finding that the UP is indeed equivalent to
SDPs in the treatment of anxiety disorders, regardless of specific diagnosis, would have
important implications for treatment dissemination and transportability.

A variation of the equivalency research design is the benchmarking design, which


involves a quantitative comparison between treatment outcomes collected in a current
study and results from similar treatment outcome studies. Demonstrating equivalence in
such a study design allows the researcher to determine whether results from a current
treatment evaluation are equivalent to findings reported elsewhere in the literature.
Results of a trial are evaluated, or benchmarked, against the findings from other
comparable trials. Weersing and Weisz (2002) used a benchmarking design to assess
differences in the effectiveness of community psychotherapy for depressed youth versus
evidence-based CBT provided in RCTs. The authors aggregated data from all available
clinical trials evaluating the effects of best-practice treatment, determined the pooled
effect sizes associated with depressed youth treated in these clinical trials, and
benchmarked these data with outcomes of depressed youth treated in community mental
health clinics. They found that outcomes of youth treated in community care settings
were more similar to youth in control conditions than to youth treated with CBT.

Benchmarking equivalency designs allow for meaningful comparison groups with which
to gauge the progress of treated participants in a clinical trial. The comparison data are
typically readily available, given that they may include samples that have been used to
obtain normative data for specific measures, or research participants whose outcome
data are included in reported results in published studies. In addition, as noted earlier,
equivalency tests can be conducted to determine the clinical significance of treatment

Page 25 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

outcomes—that is, the extent to which posttreatment functioning identified in a treated


group is comparable to functioning among normative comparisons (Kendall, Marrs-
Garcia, et al., 1999).

Sequenced Treatment Designs

When interventions are applied, a treated participant's symptoms may improve


(treatment response), may get worse (deterioration), may neither improve nor deteriorate
(treatment nonresponse), or may improve somewhat but not to a satisfactory extent
(partial response). In clinical practice, over the course of treatment important clinical
decisions must be made regarding when to escalate treatment, augment treatment with
another intervention, or switch to another supported intervention. The standard RCT
design does not provide sufficient data with which to inform the optimal sequence of
treatment for cases of nonresponse, partial response, or deterioration.

When the aim of a research study is to determine the most effective sequence of
treatments for an identified patient population, a sequenced treatment design may be
utilized. This design involves the assignment of study participants to a particular
sequence of treatment and control/comparison conditions. The order in which conditions
are assigned may be random, as in a randomized sequence design. In other sequenced
treatment designs, factors such as participant characteristics, individual treatment
outcomes, or participant preferences may influence the sequence of administered
treatments. These variations on sequenced treatment designs—prescriptive, adaptive,
and preferential treatment designs, respectively—are outlined in further detail below.

The prescriptive treatment design recognizes that individual patient characteristics play a
key role in treatment outcomes and assigns treatment condition based on these patient
characteristics. The basis of this treatment design aims to improve upon nomothetic data
models by incorporating idiographic data to treatment assignments (see Barlow & Nock,
2009). Study participants who are matched to treatment conditions based on individual
characteristics (e.g., psychiatric comorbidity, levels of distress and impairment, readiness
to change, etc.) may experience greater gains than those who are not matched to
interventions based on patient characteristics (Beutler & Harwood, 2000). In a
prescriptive (p. 55) treatment design, the clinical researcher studies the effectiveness of
a treatment decision-making algorithm as opposed to a set treatment protocol.
Participants do not have an equal chance of receiving study treatments, as is the case in
the standard RCT. Instead, what remains consistent across participants is the application
of the same decision-making algorithm, which can lead to a variety of sequenced
treatment courses.

Although a prescriptive treatment design may enhance clinical generalizability—as


practitioners will typically incorporate patient characteristics into treatment planning—
this design introduces serious threats to internal validity. In a variation, the randomized
prescriptive treatment design randomizes participants to either a blind randomization

Page 26 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

algorithm or an experimental treatment algorithm. For example, algorithm A may


randomly assign participants to one of three treatment conditions (i.e., the blind
randomization algorithm), and algorithm B may match participants to each of the three
treatment conditions based on baseline data hypothesized to inform the optimal
treatment assignment (i.e., the experimental treatment algorithm). Here, the researcher
is interested in which algorithm condition is superior, rather than what is the absolute
effect of a specific treatment protocol.

One of the primary goals of prescriptive treatment research designs is to examine the
effectiveness of treatments tailored to individuals for whom those treatments are thought
to work best. Key patient dimensions that have been found in nomothetic evaluations to
be effective mediators or moderators of treatment outcome can lay the groundwork for
decision rules to assign participants to particular interventions, or alternatively, can lead
to the administration or omission of a specific module within a larger intervention.
Prescriptive treatment designs offer opportunities to better develop and tailor efficacious
treatments to patients with varied characteristics.

In the adaptive treatment design, a participant's course of treatment is determined by his


or her clinical response across the trial. In the traditional RCT, a comparison is typically
made between an innovative treatment and some sort of placebo or accepted standard.
Some argue that a more clinically relevant design involves a comparison between an
innovative treatment and an adaptive strategy in which a participant's treatment
condition is switched based on treatment outcome to date (Dawson & Lavori, 2004). With
the adaptive treatment study design, clinical researchers can also switch participants
from one experimental group to another if a particular intervention is lacking in
effectiveness for an individual patient (see Chapter 5 in this volume). After a participant
reaches a predetermined deterioration threshold, or if he or she fails to meet a response
threshold before a given point during a trial, the participant may be switched from the
innovative treatment to the accepted standard, or vice versa. In this way, the adaptive
treatment option allows the clinical researcher to determine the relative efficacy of the
innovative treatment if the adaptive strategy produces significantly better outcomes than
the standard treatment (Dawson & Lavori, 2004).

Illustrating the utility of the adaptive treatment design, the Sequenced Treatment
Alternatives to Relieve Depression (STAR*D) study assessed the effectiveness of
depression treatments in patients with major depressive disorder (Rush et al., 2004).
Participants advanced through four levels of treatment and were assigned a particular
course of treatment depending on their response to treatment up until that point. In level
1, all participants were given citalopram for 12 to 14 weeks. Those who became
symptom-free after this period could continue on citalopram for a 12-month follow-up
period, while those who did not become symptom-free moved on to level 2. Levels 2 and 3
allowed participants to choose another medication or cognitive therapy (switch) or
augment their current medication with another medication or cognitive therapy (add-on).
In level 4, participants who were not yet symptom-free were taken off their medications
and randomly assigned to either a monoamine oxidase inhibitor (MAOI) or the

Page 27 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

combination of venlafaxine extended release with mirtazapine. At each stage of the study,
participants were assigned to treatments based on their previous treatment responses.
Roughly half of the participants became symptom-free after two treatment levels, and
roughly 70 percent of study completers became symptom-free over the course of all four
treatment levels.

The Systematic Treatment Enhancement Program for Bipolar Disorder (STEP-BD), a


national, longitudinal public health initiative, was implemented in an effort to gauge the
effectiveness of treatments for adults with bipolar disorder (Sachs et al., 2003). All
participants were initially invited to enter Standard Care Pathways (SCPs), which
involved clinical care delivered by a STEP-BD clinician. At any point during their
participation in SCP treatment, participants could become eligible for one of the
Randomized Care Pathways (RCPs) for acute depression, refractory depression, or
relapse (p. 56) prevention. Upon meeting nonresponse criteria (i.e., failure to respond to
treatment within the first 12 weeks, or failure to respond to two or more antidepressants
in the current depressive episode), participants were randomly assigned to one of the
RCP treatment arms. After participating in one of these treatment arms, participants
could return to SCP or opt to participate in another RCP. Some of the treatment arms also
allowed the treating clinician to exclude a participant from an RCP based on his or her
particular presentation. The treatment was therefore adaptive in nature, allowing for
flexibility and some element of decision making to occur within the trial. Importantly,
although such flexibility may enhance generalizability to clinical settings and when used
appropriately can guide clinical practice, this flexibility introduces serious threats to
internal validity. Accordingly, this design does not allow inferences to be made about the
absolute benefit of various interventions.

The preferential treatment design allows study participants to choose the treatment
condition(s) to which they are assigned. This approach considers patient preferences,
which emulates the process that typically occurs in clinical practice. Taking into account
patient preferences in a treatment study can result in a better understanding of which
individuals will fare best when administered specific interventions under circumstances
that incorporate their preferences in determining treatment selection. Proponents often
argue that assigning treatments based on patient preference may increase other factors
known to positively affect treatment outcomes, including patient motivation, attitudes
toward treatment, and expectations of treatment success.

Lin and colleagues (2005) utilized a preferential treatment design to explore the effects of
matching patient preferences and interventions in a population of adults with major
depression. Participants were offered antidepressant medication and/or counseling based
on patient preference, where appropriate. Participants who were matched to their
treatment preference exhibited more positive treatment outcomes at 3- and 9-month
follow-up evaluations than participants who were not matched to their preferred
treatment condition.

Page 28 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

Importantly, outcomes identified in preferential treatment designs are intertwined with


the confound of patient preferences. Accordingly, clinical researchers are wise to use
preferential treatment designs only after treatment efficacy has first been established for
the various treatment arms in a randomized design.

In a multiple-groups crossover design, participants are randomly assigned to receive a


sequence of at least two treatments, one of which may be a control condition. In this
design, participants act as their own controls, as at some point during the trial, they
receive each of the experimental and control/comparison conditions. Because each
participant is his or her own control, the risk of having comparison groups that are
dissimilar on variables such as demographic characteristics, severity of presenting
symptoms, and comorbidities is eliminated. Precautions should be taken to ensure that
the effects of one treatment intervention have receded before starting participants on the
next treatment intervention.

Illustrations of multiple-groups crossover designs can often be found in clinical trials


testing the efficacy of various medications. Hood and colleagues (2010) utilized a double-
blind crossover design in a study with untreated and selective serotonin reuptake
inhibitor (SSRI)-remitted patients with social anxiety disorder. Participants were
administered a single dose of either pramipexole (a dopamine agonist) or sulpiride (a
dopamine antagonist). One week later, participants received a single dose of the
medication they had not received the previous week. Following each medication
administration, participants were asked to rate their anxiety and mood, and they were
invited to engage in anxiety-provoking tasks. The authors concluded that untreated
participants experienced significant increases in anxiety symptoms following anxiety-
provoking tasks after both medications. In contrast, SSRI-remitted participants
experienced elevated anxiety under the effects of sulpiride and decreased anxiety levels
under the effects of pramipexole.

Multiple-groups crossover designs are best suited for the evaluation of interventions that
would not expectedly retain effects once they are removed, as is the case in the
evaluation of a therapeutic medication with a very short half-life. These designs are more
difficult to implement in the evaluation of psychosocial interventions, which often
produce effects that are somewhat irreversible (e.g., the learning of a skill, or the
acquisition of important knowledge). How can the clinical researcher evaluate separate
treatment phases when it is not possible to completely remove the intervention? In such
situations, crossover designs are misguided.

Proponents of sequential designs argue that designs that are informed by patient
characteristics, outcomes, and preferences provide patients with uniquely individualized
care within a clinical trial. The argument suggests that an appropriate match (p. 57)
between patient characteristics and treatment type will optimize success in producing
significant treatment effects and lead to a heightened understanding of interventions that
are best suited to a variety of patients and circumstances in clinical practice (Luborsky et
al., 2002). In this way, systematic evaluation is extended to the very decision-making

Page 29 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

algorithms that occur in real-world clinical practice, an important element not afforded
by the standard RCT. However, whereas these approaches increase clinical relevance and
may enhance the ability to generalize findings from research to clinical practice, they also
decrease scientific rigor by eliminating the uniformity of randomization to experimental
conditions.

Conclusion
The RCT offers the most rigorous method of examining the causal impact of therapeutic
interventions. After reviewing the essential considerations relevant for matters of design,
procedure, measurement, and data analysis, one recognizes that no one single clinical
trial, even with optimal design and procedures, can alone answer the relevant questions
about the efficacy and effectiveness of therapy. Rather, a series and collection of studies,
with varying designs and approaches, is necessary. The extensions and variations of the
RCT addressed in this chapter address important features relevant to clinical practice
that are not informed by the standard RCT; however, each modification to the standard
RCT decreases the maximized internal validity achieved in the standard RCT. Criteria for
determining evidence-based practice have been proposed (American Psychological
Association, 2005, Chambless & Hollon, 1998), and the quest to identify such treatments
continues. The goal is for the research to be rigorous, with the end goal being to optimize
clinical decision making and practice for those affected by emotional and behavioral
disorders and problems.

Controlled trials play a vital role in facilitating a dialogue between academic clinical
psychology and the public and private sector (e.g., insurance payers, Department of
Health and Human Services, policymakers). The results of controlled clinical evaluations
are increasingly being examined by both professional associations and managed care
organizations with the intent of formulating clinical practice guidelines for cost-effective
care that provides optimized service to those in need. In the absence of compelling data,
there is the risk that psychological practice will be co-opted and exploited in the service
of only profitability and/or cost containment. Clinical trials must retain scientific rigor to
enhance the ability of practitioners to deliver effective treatment procedures to
individuals in need.

References
Addis, M., Cardemil, E. V., Duncan, B., & Miller, S. (2006). Does manualization improve
therapy outcomes? In J. C. Norcross, L. E. Beutler, & R. F. Levant (Eds.), Evidence-based
practices in mental health (pp. 131–160). Washington, DC: American Psychological
Association.

Page 30 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

Addis, M., & Krasnow, A. (2000). A national survey of practicing psychologists’ attitudes
toward psychotherapy treatment manuals. Journal of Consulting and Clinical Psychology,
68, 331–339.

Altman, D. G., & Bland, J. M. (1995). Absence of evidence is not evidence of absence.
British Medical Journal, 311 (7003), 485.

American Psychological Association. (2005). Policy statement on evidence-based practice


in psychology. Retrieved August 27, 2011, from [Link]
evidence/[Link].

Anderson, E. M., & Lambert, M. J. (2001). A survival analysis of clinical significant change
in outpatient psychotherapy. Journal of Clinical Psychology, 57, 875–888.

Arnold, L. E., Elliott, M., Sachs, L., Bird, H., Kraemer, H. C., Wells, K. C., et al. (2003).
Effects of ethnicity on treatment attendance, stimulant response/dose, and 14-month
outcome in ADHD. Journal of Consulting and Clinical Psychology, 71, 713–727.

Barlow, D. H. (1989). Treatment outcome evaluation methodology with anxiety disorders:


Strengths and key issues. Advances in Behavior Research and Therapy, 11, 121–132.

Barlow, D. H., Farchione, T. J., Fairholme, C. P., Ellard, K. K., Boisseau, C. L., Allen, L. B.,
& Ehrenreich May, J. T. (2010). Unified protocol for transdiagnostic treatment of
emotional disorders: Therapist guide. New York: Oxford University Press.

Barlow, D. H., & Nock, M. K. (2009). Why can't we be more idiographic in our research?
Perspectives on Psychological Science, 4(1), 19–21.

Baron, R. M., & Kenny, D. A. (1986). The mediator-moderator variable distinction in social
psychological research: Conceptual, strategic, and statistical consideration. Journal of
Personality and Social Psychology, 51, 1173–1182.

Begg, C. B., Cho, M. K., Eastwood, S., Horton, R., Moher, D., Olkin, I., et al. (1996).
Improving the quality of reporting of randomized controlled trials: The CONSORT
statement. Journal of the American Medical Association, 276, 637–639.

Beidel, D. C., Turner, S. M., Sallee, F. R., Ammerman, R. T., Crosby, L. A., & Pathak, S.
(2007). SET-C vs. Fluoxetine in the treatment of childhood social phobia. Journal of the
American Academy of Child and Adolescent Psychiatry, 46, 1622–1632.

Bernal, G., Bonilla, J., & Bellido, C. (1995). Ecological validity and cultural sensitivity for
outcome research: Issues for the cultural adaptation and development of psychosocial
treatments with Hispanics. Journal of Abnormal Child Psychology, 23, 67–82.

Bernal, G., & Scharron-Del-Rio, M. R. (2001). Are empirically supported treatments valid
for ethnic minorities? Toward an alternative approach for treatment research. Cultural
Diversity and Ethnic Minority Psychology, 7, 328–342.

Page 31 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

Bersoff, D. M., & Bersoff, D. N. (1999). Ethical perspectives in clinical research. In


(p. 58)

P. C. Kendall, J. Butcher, & G. Holmbeck (Eds.), Handbook of research methods in clinical


psychology (pp. 31–55). New York, NY: Wiley.

Beutler, L. E., & Harwood, M. T. (2000). Prescriptive therapy: A practical guide to


systematic treatment selection. New York: Oxford University Press.

Blanco, C., Olfson, M., Goodwin, R. D., Ogburn, E., Liebowitz, M. R., Nunes, E. V., &
Hasin, D. S. (2008). Generalizability of clinical trial results for major depression to
community samples: Results from the National Epidemiologic Survey on Alcohol and
Related Conditions. Journal of Clinical Psychiatry, 69, 1276–1280.

Brown, R. A., Evans, M., Miller, I., Burgess, E., & Mueller, T. (1997). Cognitive-behavioral
treatment for depression in alcoholism. Journal of Consulting and Clinical Psychology, 65,
715–726.

Chambless, D. L., & Hollon, S. D. (1998). Defining empirically supported therapies.


Journal of Consulting and Clinical Psychology, 66, 7–18.

Comer, J. S., & Kendall, P. C. (2004). A symptom-level examination of parent-child


agreement in the diagnosis of anxious youths. Journal of the American Academy of Child
and Adolescent Psychiatry, 43, 878–886.

Creed, T. A., & Kendall, P. C. (2005). Therapist alliance-building behavior within a


cognitive-behavioral treatment for anxiety in youth. Journal of Consulting and Clinical
Psychology, 73, 498–505.

Curry, J., Silva, S., Rohde, P., Ginsburg, G., Kratochvil, C., Simons, A., et al. (2011).
Recovery and recurrence following treatment for adolescent major depression. Archives
of General Psychiatry, 68(3), 263–269.

Dawson, R., & Lavori, P. W. (2004). Placebo-free designs for evaluating new mental health
treatments: The use of adaptive treatment strategies. Statistics in Medicine, 23, 3249–
3262.

De Los Reyes, A., & Kazdin, A. E. (2005). Informant discrepancies in the assessment of
childhood psychopathology: A critical review, theoretical framework, and
recommendations for further study. Psychological Bulletin, 131, 483–509.

De Los Reyes, A., & Kazdin, A. E. (2006). Conceptualizing changes in behavior in


intervention research: The range of possible changes model. Psychological Review, 113,
554–583.

Dobson, K. S., & Hamilton, K. E. (2002). The stage model for psychotherapy manual
development: A valuable tool for promoting evidence-based practice. Clinical Psychology:
Science and Practice, 9, 407–409.

Page 32 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

Dobson, K. S., Hollon, S. D., Dimidjian, S., Schmaling, K. B., Kohlenberg, R. J., Gallop, R.
J., et al. (2008). Randomized trial of behavioral activation, cognitive therapy, and
antidepressant medication in the prevention of relapse and recurrence in major
depression. Journal of Consulting and Clinical Psychology, 76, 468–477.

Dobson, K. S., & Shaw, B. (1988). The use of treatment manuals in cognitive therapy.
Experience and issues. Journal of Consulting and Clinical Psychology, 56, 673–682.

Flaherty, J. A., & Meaer, R. (1980). Measuring racial bias in inpatient treatment. American
Journal of Psychiatry, 137, 679–682.

Garfield, S. (1998). Some comments on empirically supported psychological treatments.


Journal of Consulting and Clinical Psychology, 66, 121–125.

Gladis, M. M., Gosch, E. A., Dishuk, N. M., & Crits-Cristoph, P. (1999). Quality of life:
Expanding the scope of clinical significance. Journal of Consulting and Clinical
Psychology, 67, 320–331.

Hall, G. C. N. (2001). Psychotherapy research with ethnic minorities: Empirical, ethnical,


and conceptual issues. Journal of Consulting and Clinical Psychology, 69, 502–510.

Hedeker, D., & Gibbons, R. D. (1994). A random-effects ordinal regression model for
multilevel analysis. Biometrics, 50, 933–944.

Hedeker, D., & Gibbons, R. D. (1997). Application of random-effects pattern-mixture


models for missing data in longitudinal studies. Psychological Methods, 2, 64–78.

Heiser, N. A., Turner, S. M., Beidel, D. C., & Roberson-Nay, R. (2009). Differentiating
social phobia from shyness. Journal of Anxiety Disorders, 23, 469–476.

Hoagwood, K. (2002). Making the translation from research to its application: The je ne
sais pas of evidence-based practices. Clinical Psychology: Science and Practice, 9, 210–
213.

Hollon, S. D. (1996). The efficacy and effectiveness of psychotherapy relative to


medications. American Psychologist, 51, 1025–1030.

Hollon, S. D., & DeRubeis, R. J. (1981). Placebo-psychotherapy combinations:


Inappropriate representation of psychotherapy in drug-psychotherapy comparative trials.
Psychological Bulletin, 90, 467–477.

Hollon, S. D., Garber, J., & Shelton, R. C. (2005). Treatment of depression in adolescents
with cognitive behavior therapy and medications: A commentary on the TADS project.
Cognitive and Behavioral Practice, 12, 149–155.

Holmbeck, G. N. (1997). Toward terminological, conceptual, and statistical clarity in the


study of mediators and moderators: Examples from the child-clinical and pediatric

Page 33 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

psychology literatures. Journal of Consulting and Clinical and Clinical Psychology, 65,
599–610.

Homma-True, R., Greene, B., Lopez, S. R., & Trimble, J. E. (1993). Ethnocultural diversity
in clinical psychology. Clinical Psychologist, 46, 50–63.

Hood, S., D., Potokar, J. P., Davies, S. J., Hince, D. A., Morris, K., Seddon, K. M., et al.
(2010). Dopaminergic challenges in social anxiety disorder: Evidence for dopamine D3
desensitization following successful treatment with serotonergic antidepressants. Journal
of Psychopharmacology, 24(5), 709–716.

Huey, S. J., & Polo, A. J. (2008). Evidence-based psychosocial treatments for ethnic
minority youth. Journal of Clinical Child and Adolescent Psychology, 37, 262–301.

Jaccard, J., & Guilamo-Ramos, V. (2002a). Analysis of variance frameworks in clinical child
and adolescent psychology: Issues and recommendations. Journal of Clinical Child and
Adolescent Psychology, 31, 130–146.

Jaccard, J., & Guilamo-Ramos, V. (2002b). Analysis of variance frameworks in clinical child
and adolescent psychology: Advanced issues and recommendations. Journal of Clinical
Child and Adolescent Psychology, 31, 278–294.

Jacobson, N. S., Follette, W. C., & Revenstorf, D. (1984). Psychotherapy outcome research:
Methods for reporting variability and evaluating clinical significance. Behavior Therapy,
15, 336–352.

Jacobson, N. S., & Hollon, S. D. (1996a). Cognitive-behavior therapy versus


pharmacotherapy: Now that the jury's returned its verdict, it's time to present the rest of
the evidence. Journal of Consulting and Clinical Psychology, 74, 74–80.

Jacobson, N. S., & Hollon, S. D. (1996b). Prospects for future comparisons


(p. 59)

between drugs and psychotherapy: Lessons from the CBT-versus-pharmacotherapy


exchange. Journal of Consulting and Clinical Psychology, 64, 104–108.

Jacobson, N. S., Roberts, L. J., Berns, S. B., & McGlinchey, J. B. (1999). Methods for
defining and determining the clinical significance of treatment effects. Description,
application, and alternatives. Journal of Consulting and Clinical Psychology, 67, 300–307.

Jacobson, N. S., & Traux, P. (1991). Clinical significance: A statistic approach to defining
meaningful change in psychotherapy research. Journal of Consulting and Clinical
Psychology, 59, 12–19.

Jarrett, R. B., Vittengl, J. R., Doyle, K., & Clark, L. A. (2007). Changes in cognitive content
during and following cognitive therapy for recurrent depression: Substantial and
enduring, but not predictive of change in depressive symptoms. Journal of Consulting and
Clinical Psychology, 75, 432–446.

Page 34 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

Jones, B., Jarvis, P., Lewis, J. A., & Ebbutt, A. F. (1996). Trials to assess equivalence: the
importance of rigorous methods. British Medical Journal, 313 (7048), 36–39.

Karver, M., Shirk, S., Handelsman, J. B., Fields, S., Crisp, H., Gudmundsen, G., &
McMakin, D. (2008). Relationship processes in youth psychotherapy: Measuring alliance,
alliance-building behaviors, and client involvement. Journal of Emotional and Behavioral
Disorders, 16, 15–28.

Kazdin, A. E. (1999). The meanings and measurement of clinical significance. Journal of


Consulting and Clinical Psychology, 67, 332–339.

Kazdin, A. E. (2003). Research design in clinical psychology (4th ed.). Boston, MA: Allyn
and Bacon.

Kazdin, A. E., & Bass, D. (1989). Power to detect differences between alternative
treatments in comparative psychotherapy outcome research. Journal of Consulting and
Clinical Psychology, 57, 138–147.

Kendall, P. C. (1999). Introduction to the special section: Clinical Significance. Journal of


Consulting and Clinical Psychology, 67, 283–284.

Kendall, P. C., & Beidas, R. S. (2007). Smoothing the trail for dissemination of evidence-
based practices for youth: Flexibility within fidelity. Professional Psychology: Research
and Practice, 38, 13–20.

Kendall, P. C., & Chu, B. (1999). Retrospective self-reports of therapist flexibility in a


manual-based treatment for youths with anxiety disorders. Journal of Clinical Child
Psychology, 29, 209–220.

Kendall, P. C., & Grove, W. (1988). Normative comparisons in therapy outcome. Behavioral
Assessment, 10, 147–158.

Kendall, P. C., & Hedtke, K. A. (2006). Cognitive-behavioral therapy for anxious children
(3rd ed.). Ardmore, PA: Workbook Publishing.

Kendall, P. C., & Hollon, S. D. (1983). Calibrating therapy: Collaborative archiving of tape
samples from therapy outcome trials. Cognitive Therapy and Research, 7, 199–204.

Kendall, P. C., Hollon, S., Beck, A. T., Hammen, C., & Ingram, R. (1987). Issues and
recommendations regarding use of the Beck Depression Inventory. Cognitive Therapy and
Research, 11, 289–299.

Kendall, P. C., Hudson, J. L., Gosch, E., Flannery-Schroeder, E., & Suveg, C. (2008).
Cognitive-behavioral therapy for anxiety disordered youth: A randomized clinical trial
evaluating child and family modalities. Journal of Consulting and Clinical Psychology, 76,
282–297.

Page 35 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

Kendall, P. C., & Kessler, R. C. (2002). The impact of childhood psychopathology


interventions on subsequent substance abuse: Policy implications, comments, and
recommendations. Journal of Consulting and Clinical Psychology, 70, 1303–1306.

Kendall, P. C., Marrs-Garcia, A., Nath, S. R., & Sheldrick, R. C. (1999). Normative
comparisons for the evaluation of clinical significance. Journal of Consulting and Clinical
Psychology, 67, 285–299.

Kendall, P. C., Safford, S., Flannery-Schroeder, E., & Webb, A. (2004). Child anxiety
treatment: Outcomes in adolescence and impact on substance use and depression at 7.4-
year follow-up. Journal of the Consulting and Clinical Psychology, 72, 276–287.

Kendall, P. C., & Southam-Gerow, M. A. (1995). Issues in the transportability of treatment:


The case of anxiety disorders in youth. Journal of Consulting and Clinical Psychology, 63,
702–708.

Kendall, P. C., & Sugarman, A. (1997). Attrition in the treatment of childhood anxiety
disorders. Journal of Consulting and Clinical Psychology, 65, 883–888.

Kendall, P. C., & Suveg, C. (2008). Treatment outcome studies with children: Principles of
proper practice. Ethics and Behavior, 18, 215–233.

Kraemer, H. C., & Kupfer, D. J. (2006). Size of treatment effects and their importance to
clinical research and practice. Biological Psychiatry, 59, 990–996.

Kraemer, H. C., Wilson, G. T., Fairburn, C. G., & Agras, W. S. (2002). Mediators and
moderators of treatment effects in randomized clinical trials. Archives of General
Psychiatry, 59, 877–883.

Laird, N. M., & Ware, J. H. (1982). Random-effects models for longitudinal data.
Biometrics, 38, 963–974.

Leon, A. C., Mallinckrodt, C. H., Chuang-Stein, C., Archibald, D. G., Archer, G. E., &
Chartier, K. (2006). Attrition in randomized controlled clinical trials: Methodological
issues in psychopharmacology. Biological Psychiatry, 59, 1001–1005.

Lin, P., Campbell, D. G., Chaney, E. F., Liu, C., Heagerty, P., Felker, B. L., et al. (2005). The
influence of patient preference on depression treatment in primary care. Annals of
Behavioral Medicine, 30(2), 164–173.

Little, R. J. A., & Rubin, D. (2002). Statistical analysis with missing data (2nd ed.). New
York: Wiley.

Lopez, S. R. (1989). Patient variable biases in clinical judgment: Conceptual overview and
methodological considerations. Psychological Bulletin, 106, 184–204.

Page 36 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

Luborsky, L., Rosenthal, R., Diguer, L., Andrusyna, T. P., Berman, J. S., Levitt, J. T., et al.
(2002). The dodo bird verdict is alive and well—mostly. Clinical Psychology: Science and
Practice, 9(1), 2–12.

Marcus, S. M., Gorman, J., Shea, M. K., Lewin, D., Martinez, J., Ray, S., et al. (2007). A
comparison of medication side effect reports by panic disorder patients with and without
concomitant cognitive behavior therapy. American Journal of Psychiatry, 164, 273–275.

Mason, M. J. (1999). A review of procedural and statistical methods for handling attrition
and missing data. Measurement and Evaluation in Counseling and Development, 32, 111–
118.

Moher, D., Schulz, K. F., & Altman, D. (2001). The CONSORT Statement: Revised
(p. 60)

recommendations for improving the quality of reports of parallel-group randomized trials.


Journal of the American Medical Association, 285, 1987–1991.

Molenberghs, G., Thijs, H., Jansen, I., Beunckens, C., Kenward, M. G., Mallinckrodt, C., &
Carroll, R. (2004). Analyzing incomplete longitudinal clinical trial data. Biostatistics, 5,
445–464.

MTA Cooperative Group. (1999). A 14-month randomized clinical trial of treatment


strategies for attention-deficit/hyperactivity disorder. Archives of General Psychiatry, 56,
1088–1096.

Mufson, L., Dorta, K. P., Wickramaratne, P., Nomura, Y., Olfson, M., & Weissman, M. M.
(2004). A randomized effectiveness trial of interpersonal psychotherapy for depressed
adolescents. Archives of General Psychiatry, 61, 577–584.

Neuner, F., Onyut, P. L., Ertl, V., Odenwald, M., Schauer, E., & Elbert, T. (2008). Treatment
of posttraumatic stress disorder by trained lay counselors in an African refugee
settlement: A randomized controlled trial. Journal of Consulting and Clinical Psychology,
76, 686–694.

O'Leary, K. D., & Borkovec, T. D. (1978). Conceptual, methodological, and ethical


problems of placebo groups in psychotherapy research. American Psychologist, 33, 821–
830.

Olfson, M., Cherry, D., & Lewis-Fernandez, R. (2009). Racial differences in visit duration
of outpatient psychiatric visits. Archives of General Psychiatry, 66, 214–221.

Pediatric OCD Treatment Study (POTS) Team. (2004). Cognitive-behavior therapy,


sertraline, and their combination for children and adolescents with obsessive-compulsive
disorder: The Pediatric OCD Treatment Study (POTS) randomized controlled trial. Journal
of the American Medical Association, 292, 1969–1976.

Page 37 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

Pelham, W. E., Jr., Gnagy, E. M., Greiner, A. R., Hoza, B., Hinshaw, S. P., Swanson, J. M., et
al. (2000). Behavioral versus behavioral and psychopharmacological treatment in ADHD
children attending a summer treatment program. Journal of Abnormal Child Psychology,
28, 507–525.

Perepletchikova, F., & Kazdin, A. E. (2005). Treatment integrity and therapeutic change:
Issues and research recommendations. Clinical Psychology: Science and Practice, 12,
365–383.

Reis, B. F., & Brown, L. G. (2006). Preventing therapy dropout in the real world: The
clinical utility of videotape preparation and client estimate of treatment duration.
Professional Psychology: Research and Practice, 37, 311–316.

Rush, A. J., Fava, M., Wisniewski, S. R., Lavori, P. W., Trivedi, M. H., Sackeim, H. A., et al.
(2004). Sequenced treatment alternatives to relieve depression (STAR*D): rationale and
design. Controlled Clinical Trials, 25(1), 119–142.

Sachs, G. S., Thase, M. E., Otto, M. W., Bauer, M., Miklowitz, D., Wisniewski, S. R., et al.
(2003). Rationale, design, and methods of the systematic treatment enhancement
program for bipolar disorder (STEP-BD). Biological Psychiatry, 53(11), 1028–1042.

Shirk, S. R., Gudmundsen, G., Kaplinski, H., & McMakin, D. L. (2008). Alliance and
outcome in cognitive-behavioral therapy for adolescent depression. Journal of Clinical
Child and Adolescent Psychology, 37, 631–639.

Silverman, W. K., Kurtines, W. M., & Hoagwood, K. (2004). Research progress on


effectiveness, transportability, and dissemination of empirically supported treatments:
Integrating theory and research. Clinical Psychology: Science and Practice, 11, 295–299.

Snowden, L. R. (2003). Bias in mental health assessment and intervention: Theory and
evidence. American Journal of Public Health, 93, 239–243.

Southam-Gerow, M. A., Ringeisen, H. L., & Sherrill, J. T. (2006). Integrating interventions


and services research: Progress and prospects. Clinical Psychology: Science and Practice,
13, 1–8.

Sue, S. (1998). In search of cultural competence in psychotherapy and counseling.


American Psychologist, 53, 440–448.

Suveg, C., Comer, J. S., Furr, J. M., & Kendall, P. C. (2006). Adapting manualized CBT for a
cognitively delayed child with multiple anxiety disorders. Clinical Case Studies, 5, 488–
510.

Sweeney, M., Robins, M., Ruberu, M., & Jones, J. (2005). African-American and Latino
families in TADS: Recruitment and treatment considerations. Cognitive and Behavioral
Practice, 12, 221–229.

Page 38 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

Treadwell, K., Flannery-Schroeder, E. C., & Kendall, P. C. (1994). Ethnicity and gender in
a sample of clinic-referred anxious children: Adaptive functioning, diagnostic status, and
treatment outcome. Journal of Anxiety Disorders, 9, 373–384.

Treatment for Adolescents with Depression Study (TADS) Team. (2004). Fluoxetine,
cognitive-behavioral therapy, and their combination for adolescents with depression:
Treatment for Adolescents with Depression Study (TADS) randomized controlled trial.
Journal of the American Medical Association, 292, 807–820.

Vanable, P. A., Carey, M. P., Carey, K. B., & Maisto, S. A. (2002). Predictors of participation
and attrition in a health promotion study involving psychiatric outpatients. Journal of
Consulting and Clinical Psychology, 70, 362–368.

Walders, N., & Drotar, D. (2000). Understanding cultural and ethnic influences in
research with child clinical and pediatric psychology populations. In D. Drotar (Ed.),
Handbook of research in pediatric and clinical child psychology (pp. 165–188). New York:
Springer.

Walkup, J. T., Albano, A. M., Piacentini, J., Birmaher, B., Compton, S. N., et al. (2008)
Cognitive behavioral therapy, sertraline, or a combination in childhood anxiety. New
England Journal of Medicine, 359, 1–14.

Waltz, J., Addis, M. E., Koerner, K., & Jacobson, N. S. (1993). Testing the integrity of a
psychotherapy protocol: Assessment of adherence and competence. Journal of Consulting
and Clinical Psychology, 61, 620–630.

Weersing, R. V., & Weisz, J. R. (2002). Community clinic treatment of depressed youth:
Benchmarking usual care against CBT clinical trials. Journal of Consulting and Clinical
Psychology, 70(2), 299–310.

Weisz, J., Donenberg, G. R., Han, S. S., & Weiss, B. (1995). Bridging the gap between
laboratory and clinic in child and adolescent psychotherapy. Journal of Consulting and
Clinical Psychology, 63, 688–701.

Weisz, J. R., Thurber, C. A., Sweeney, L., Proffitt, V. D., & LeGagnoux, G. L. (1997). Brief
treatment of mild-to-moderate child depression using primary and secondary control
enhancement training. Journal of Consulting and Clinical Psychology, 65(4), 703–707.

Weisz, J. R., Weiss, B., & Donenberg, G. R. (1992). The lab versus the clinic: Effects of
child and adolescent psychotherapy. American Psychologist, 47, 1578–1585.

Westbrook, D., & Kirk, J. (2007). The clinical effectiveness of cognitive behaviour
(p. 61)

therapy: Outcome for a large sample of adults treated in routine practice. Behaviour
Research and Therapy, 43, 1243–1261.

Page 39 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017


The Randomized Controlled Trial: Basics and Beyond

Westen, D., Novotny, C., & Thompson-Brenner, H. (2004). The empirical status of
empirically supported psychotherapies: Assumptions, findings, and reporting in
controlled clinical trials. Psychological Bulletin, 130, 631–663.

Wilson, G. T. (1995). Empirically validated treatments as a basis for clinical practice:


Problems and prospects. In S. C. Hayes, V. M. Follette, R. D. Dawes & K. Grady (Eds.),
Scientific standards of psychological practice: Issues and recommendations (pp. 163–
196). Reno, NV: Context Press.

Yeh, M., McCabe, K., Hough, R. L., Dupuis, D., & Hazen, A. (2003). Racial and ethnic
differences in parental endorsement of barriers to mental health services in youth. Mental
Health Services Research, 5, 65–77.

Philip C. Kendall

Philip C. Kendall, Department of Psychology, Temple University.

Jonathan S. Comer

Jonathan S. Comer, Florida International University

Candice Chow

Candice Chow, Department of Psychology, Wellesley College

Page 40 of 40

PRINTED FROM OXFORD HANDBOOKS ONLINE ([Link]). (c) Oxford University Press, 2015. All Rights
Reserved. Under the terms of the licence agreement, an individual user may print out a PDF of a single chapter of a title in
Oxford Handbooks Online for personal use (for details see Privacy Policy).

Subscriber: Universita di Milano Bicocca; date: 08 May 2017

You might also like