0% found this document useful (0 votes)
6 views9 pages

Experimental Design in the 3Rs Framework

The document discusses the application of the 3Rs—replacement, reduction, and refinement—in animal experimentation, emphasizing the importance of good experimental design and analysis to enhance animal welfare while achieving research objectives. It highlights public support for animal research under ethical considerations and outlines the principles of the 3Rs as a framework for conducting humane experiments. The authors also address the tension between using fewer animals and ensuring high-quality evidence, advocating for careful sample size calculations and appropriate experimental designs to optimize research outcomes.

Uploaded by

ruwaidirham2024
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
6 views9 pages

Experimental Design in the 3Rs Framework

The document discusses the application of the 3Rs—replacement, reduction, and refinement—in animal experimentation, emphasizing the importance of good experimental design and analysis to enhance animal welfare while achieving research objectives. It highlights public support for animal research under ethical considerations and outlines the principles of the 3Rs as a framework for conducting humane experiments. The authors also address the tension between using fewer animals and ensuring high-quality evidence, advocating for careful sample size calculations and appropriate experimental designs to optimize research outcomes.

Uploaded by

ruwaidirham2024
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

The Place of Experimental Design and Statistics in the 3Rs

Richard M. A. Parker and William J. Browne

Abstract disease, gauging the efficacy of potentially therapeutic


interventions, or testing product toxicity. Public surveys
The 3Rs—replacement, reduction, and refinement—can be largely indicate majority support for research using animals.
applied to any animal experiment by researchers and other In the United States, for example, 57% of respondents to

Downloaded from [Link] by guest on 11 July 2025


bodies seeking to conduct those studies in as humane a man- Gallup’s Values and Beliefs Survey in 2014 believed that
ner as possible. Key to the success of this endeavor is an ap- medical testing on animals was morally acceptable (Gallup
preciation of the principles of good experimental design and 2014). Similarly, the majority of the surveyed British public
analysis; these need to be considered in concert before any accept the use of animal experimentation, but conditional on
data is collected. Indeed, many of the principles central to factors such as there being no unnecessary suffering (66%
helping achieve the objectives of the 3Rs—such as conduct- agree), and there being no alternative to their use in medical
ing valid, reliable, and efficient experiments; clearly and research (63% agree) (Ipsos MORI 2012). As such, there is
transparently reporting findings; and ensuring that an appre- a commonly held tension between the perceived utility of
ciation and understanding of animal welfare plays a central animal experimentation and concerns regarding the cost to
role in laboratory practice—are to the betterment of research animals. This tension is reflected in legislation too: for
per se. instance, both the UK’s Animals (Scientific Procedures)
Act 1986 (ASPA) and the European Union’s Directive
Key Words: 3Rs; animal welfare; experimental design; lab- 2010/63/EU on the Protection of Animals Used for Scientific
oratory animals; reduction; refinement; replacement; statistics Purposes explicitly apply the 3Rs, of replacement, reduc-
tion, and refinement, as indeed does the legislation of a
Introduction variety of other jurisdictions, albeit implicitly.

any millions of nonhuman animals (hereafter “ani-

M mals”) are used every year as subjects in laboratory


experiments. Whilst estimates of the number of an-
imals used annually worldwide are very imprecise due to dif-
The 3Rs

Proposed by Russell and Burch in The Principles of Humane


ferences in how different nations collect and report statistics, Experimental Technique (1959), the 3Rs are akin to an ethical
estimates nevertheless extend into the tens and hundreds of algorithm that can be applied to any animal experiment by re-
millions. For example, in a review of the topic, Taylor and searchers and other bodies seeking to conduct those studies in
colleagues (2008) noted that prior published estimates ranged as humane a manner as possible.
from 50 to 200 million, before deriving their own estimate of Russell and Burch demarcated the first of the 3Rs, replace-
115 million, a figure they thought likely to be a conserva- ment, into relative and absolute. Absolute replacement, the op-
tive one. timal outcome, refers to changing an experimental protocol that
The uses to which these animals are put are manifold, from uses animals in some capacity—be it as whole-animal study
more fundamental biological research to applied studies con- subjects, sources of tissue, etc.—into one that does not. For ex-
cerning human health: be it the further understanding of ample, replacing the study of biological processes conducted in
vivo, or from animal-derived tissue in vitro, to one conducted
by the construction of virtual models on a computer (in silico).
Richard M. A. Parker, PhD, is a research associate in the Animal Behaviour
and Welfare Research Group, School of Veterinary Sciences, and the Centre
Otherwise, replacements may be relative, for example, using an
for Multilevel Modelling, University of Bristol, UK. William J. Browne is established cell line that was originally derived from animals,
professor of biostatistics in the Animal Behaviour and Welfare Research or replacing a given animal with one thought less likely to suf-
Group, School of Veterinary Sciences, and Director of the Centre for fer as a result of the proposed experimental conditions (in the
Multilevel Modelling, University of Bristol, UK. terminology of the UK’s legislation, this would be replacing a
Address correspondence and reprint requests to Dr. Richard M. A. Parker,
Animal Behaviour and Welfare Research Group, University of Bristol, Lang-
protected with an unprotected animal).
ford House, Langford, Bristol, BS40 5DU, UK or email [Link]@ The second R, reduction, refers to reducing the numbers of
[Link]. animals used without meaningfully diminishing the amount

ILAR Journal, Volume 55, Number 3, doi: 10.1093/ilar/ilu044


© The Author 2014. Published by Oxford University Press on behalf of the Institute for Laboratory Animal Research. All rights reserved.
For permissions, please email: [Link]@[Link] 477
and quality of information gleaned from experiments. This key biological characteristics thought sufficiently analogous
may involve reducing the number of animals used in any giv- to those of humans to allow inferences to be fruitfully applied
en experiment but, perhaps counterintuitively, could also in- across species.
volve increasing the number if this results in a net reduction of Experiments are often, of course, designed to test hypo-
numbers across a ( potentially) longer series of experiments. theses: for example, if we hypothesize that a new intervention
Finally, the third R of refinement refers to “any approach (A) has greater efficacy than an alternative (B) in treating a
which avoids or minimises the actual or potential pain, dis- particular medical condition, we can gauge empirical support
tress and other adverse effects experienced at any time during for this hypothesis by testing specific predictions generated
the life of the animals involved, and which enhances their from it. For instance, we may test a prediction that animals
wellbeing” (Buchanan-Smith et al. 2005)—such concerns, undergoing intervention A will have lower levels of a disease
either implicitly or explicitly, ultimately relate to the subjec- marker—a proxy measure of the underlying disease state
tive experience of the animal (e.g., Dawkins 1990). of interest—than animals undergoing intervention B. Alterna-
As an ethical algorithm, the 3Rs can be applied to a range tively, rather than testing predictions, the focus of experiments
of scenarios, such as field research, or the use of animals in

Downloaded from [Link] by guest on 11 July 2025


may be to estimate relationships such as the shape of a dose-
veterinary or biological education, and of course most perti- response curve. Naturally, different experimental designs and
nently here, laboratory experiments. At present, however, the analyses are best suited to answer the different sorts of ques-
extent to which the 3Rs are successfully applied in laboratory tions researchers can pose. One principle is common through-
research is a moot point. For example, a number of concerns out, however, namely that experimental design and analysis go
have been raised recently about the manner in which animal hand in hand, and this observation is central to fruitfully apply-
experiments are designed, analyzed, or reported (i.e., whether ing the 3Rs, be it in developing, and validating, alternatives to
the most efficient use is being made of the animals that are animal use; reducing the number of animals that are used with
employed) (e.g., Button et al. 2013; Kilkenny et al. 2009; little or no cost to inferential gain; or refining procedures so that
Sena et al. 2010), and the extent to which the experiments they avoid, or minimize, pain, distress, suffering, and lasting
can be generalized to the target population about which infer- harm or otherwise improve animal welfare (NC3Rs 2014).
ences are to be made (e.g., Hackam and Redelmeier 2006; The relationship between experimental design and analysis
Perel et al. 2007; van der Worp et al. 2010). More generally, in the 3Rs is a large topic, in part because any improvement in
recent technological developments, such as methods of ge- a study’s design and analysis can be construed as an applica-
netic modification, have resulted in increases in the numbers tion of the 3Rs: ensuring animals are used as efficiently as
of animals used (e.g., after a downturn, the number of animals possible in experiments that are valid and reliable may con-
used in regulated procedures in Great Britain, at least, is in- tribute to a net reduction in the number of subjects used. In
creasing, chiefly due to the breeding of genetically modified what follows, therefore, we offer an overview of some of
and harmful mutants) (Home Office 2013). Similarly, mem- the main points to consider, and direct the reader to other re-
ber states contributing to a recent report on the number of an- sources for a deeper treatment of specific topics. We focus
imals used for experimental and other scientific purposes in mostly on reduction, although naturally many of the princi-
the European Union reported an increase in research using ples that contribute to helping achieve this objective can be
transgenic mice (European Commission 2013). applied to choosing appropriate alternatives (replacement),
or titrating the effect of refinement against any apparent costs
of a modification in experimental protocol.
Experimental Design and Statistics

Experiments are tools we use to empirically infer something Reduction


about how the world works: to estimate quantities and test
( potentially causal) effects we cannot directly observe or Many research programmes need to carefully manage the
measure. So, for example, if a research team is interested in number of observational units they study, be it for reasons
assessing the efficacy of a certain intervention in treating a of funding, logistics, time available, and so on. Laboratory-
particular medical condition, it’s not possible, nor even nec- based research on animals shares these constraints, of course,
essary, to test the whole population of interest. Rather, a sam- but Russell and Burch’s principle of reducing the number of
ple is studied, and inferences are made about the larger animals used in such settings is naturally informed by some-
population from which that sample is drawn. Here, because thing else, namely an ethical consideration of the costs, to
that population has not been directly observed, inferences those animals, of their use as subjects.
are made with uncertainty: they may not be what we would Reducing the number of animals used in an experiment can
regard as a true statement if we had access to data from the result from a change in experimental design and analysis such
full population. In many animal experiments, of course, an that the same amount of information is obtained from the
additional inferential leap is made, across species: a research- study despite a reduction in the animal-level sample size
er may be unable to study a sample of the population of ulti- (i.e., a researcher is able to provide as good, or better, an
mate interest (e.g., humans with a particular medical answer to his or her research question even though fewer
condition) and so studies a sample from a population with animals are being used in an experiment). Depending on

478 ILAR Journal


the nature of the experiment, reuse of animals may be a pos- animals. However, the same observed difference based on a
sibility, but since this has the potential to further compromise sample average of, say, thirty subjects would be clearer evi-
the welfare of any animal reused, it may run contrary to one dence of a real difference. More formally, each group mean
of the other objectives of the 3Rs, namely that of refinement has uncertainty that we measure
pffiffiffi through the standard error
(e.g., Fenwick et al. 2009). Alternatively, a researcher might of the mean (SEM) = s= n, where s is our estimate of the
consider using more animals than initially anticipated in a standard deviation of the population (for which we use the
given study, if that means that that experiment alone provides standard deviation of the sample), and n is the sample size.
a considerably better test of their research question of interest. Here as n increases we decrease the SEM and thus estimate
This might then lead to a net reduction in the total number of the population mean with greater precision. Therefore we
animals used when considering the experiment as part of a are better able to distinguish the outcome of one intervention
series (i.e., compared with conducting a larger number of group from the other.
smaller experiments, each a poorer test) (Button et al. Similarly, if we find that the sample mean for animals in
2013). Otherwise, careful consideration of the prospects of group A is considerably lower than that for animals in group

Downloaded from [Link] by guest on 11 July 2025


an experiment may result in the conclusion that it is so unlike- B, and the outcome measure varies little between the animals
ly to answer the question that the use of animals is not justi- in each group (for example, only one animal in group A is
fied, and the experiment is dropped. above the sample mean of group B), then we will have
Thus, in applying the principle of reduction to a given ex- more evidence of a significant difference than if the outcome
periment, there is a tension between using as few animals as measures from the animals exhibited large variability, with
possible and ensuring the quality of evidence one can glean considerable overlap between the two groups. Relating this
from that experiment remains sufficiently high, whilst re- to the formula above, if s is increased for fixed n, then the
specting the other objectives of the 3Rs. This involves select- SEM increases, and thus the population mean is estimated
ing an experimental design, and subsequent analysis, that can with less precision.
most help to reliably and validly answer the research question So the number of animals to be sampled (the sample size),
of interest, but efficiently so, using no more animals than nec- the anticipated difference between the group means (the un-
essary. Here, appropriate sample-size calculations play a key standardized effect size), and the variability in the two sam-
role (e.g., Festing et al. 2002). ples are central to how much weight we give any evidence as
being in support, or otherwise, of our hypothesis of interest.
Here we have been implicitly comparing our experimental,
Judging Appropriate Sample Sizes or alternative, hypothesis (H1) with the null hypothesis (H0)
that it makes no difference at all to the outcome measure
Ideally, if there were no costs of doing so—ethical, financial, which intervention the subjects undergo. Rather than infor-
or otherwise—researchers wishing to uncover the nature of mally speculating to what extent our experimental evidence
causal effects and relationships in the population of interest is in keeping with our alternative hypothesis, however, we
would test the whole population. Clearly, in the vast majority naturally need to formalize this above a hunch (i.e., we
of cases, this is neither possible nor desirable, and so re- need a rule/criterion for deciding between these two hypo-
searchers must judge how small a sample size they can ob- theses). In this case, a natural rule would be to consider the
serve whilst still giving themselves sufficient probability of value of the absolute difference in sample means and then
uncovering the truth about their research question of interest. reject the null hypothesis if this difference is greater than
As such, there are a variety of methods and formulae that re- some chosen value that depends upon the variability in the
searchers can use to help guide their choice of sample size samples and their respective sizes. The choice of value, or
( plus see suggested references at the end of this section). threshold, is a balance between making two types of error.
Whilst the relevant procedures can be technical, it’s important The larger we make the threshold, the more often we will
to note that thinking about what aspects of the study influence reject the null hypothesis both if it is false (true positive)
the desired sample size, and how they interrelate, is rather but also if it is true (false positive), with the latter known as
intuitive. a type I error. Conversely the smaller we make the threshold,
For example, if we were to test the hypothesis that a new the more often we fail to reject the null hypothesis both if it is
intervention (A) significantly decreases the level of a disease true (true negative) but also if it is false (false negative), with
marker compared to an alternative intervention (B), then we the latter this time known as a type II error. We can fix the
might collect an outcome measure for a group of animals that probability of making a type I error (α) by choosing a specific
are exposed to intervention A and calculate the mean value threshold; this is known as the significance level of the test.
for the group. We would then compare this to the average val- As there is only one threshold but two types of error that are
ue from animals in another group that are exposed to interven- inversely related, having fixed α, the only way to reduce the
tion B. When considering group sizes, clearly an observed probability of a type II error (β) is therefore to increase the pre-
difference based on sample averages derived from just two cision of our estimates by increasing the sample size (aside
animals in each group might simply be due to the selected from using a better design that might reduce variability).
animals. We may have gotten a very different average had We generally refer to the probability of rejecting the null
we repeated the procedure and sampled a different group of hypothesis when it is false (i.e., 1- β), as the power of the

Volume 55, Number 3, doi: 10.1093/ilar/ilu044 2014 479


test. Since increasing sample size will increase power, sample- critique), for example, in the case of a completely novel inter-
size calculations are also referred to as power calculations (e.g., vention where there is no pilot data available.
see Krzywinski and Altman 2013b for a succinct primer, This may seem an unsatisfactory state of affairs, but a
including a supplementary interactive worksheet allowing the researcher can nevertheless profitably use his or her best judg-
reader to graphically investigate the relationship between pow- ment, having digested all relevant prior information, to esti-
er, sample size, effect size, etc.). mate the appropriate sample size given a variety of possible
So we have identified five pieces of information—the scenarios. For example, for a given variability (as measured
sample size, effect size, variability, significance level, and by a variance), what would be a reasonable sample size to
power—that determine whether we formally reject (or not) test the hypothesis of interest if the effect size was the smallest
the null hypothesis. If we know four of these five pieces of it could be whilst remaining of scientific interest? Very small
information, we can, typically, derive the fifth via an appro- effect sizes might be real, but uncovering them might make
priate calculation. For example, to estimate an advisable sam- only a trivial addition to our understanding of the biological
ple size we need to decide on a tolerable probability of world or the benefits of a new treatment being not practically

Downloaded from [Link] by guest on 11 July 2025


yielding a false positive result (the significance level, by con- important. If we think the effect size is actually likely to be
vention p = 0.05), an acceptable desired probability of reject- larger than this, though, we naturally also need to take appro-
ing the null hypothesis when it is false (e.g., a power of 0.8, priate calculations based on this best estimate into account
0.9, etc.), the effect size (how big a biological effect we wish when making our decision, as otherwise we risk unnecessar-
to detect), and the variability within the data. ily using too many animals. Similarly, sample size calcula-
Although the significance level and power are generally set tions could be made for a selection of possible variances.
by convention, the effect size and, perhaps especially, the var- As such, the researcher can attempt to map estimated sample
iability, are more challenging to estimate. Prior (e.g., pub- sizes for a range of possible scenarios (e.g., Lenth 2001), and
lished) studies can offer useful guidance, but often, of use these to help make an informed judgment along with a
course, a given experimental protocol is novel (or, if it is not, consideration of the costs and benefits of the research (e.g.,
it may be novel to a particular laboratory, and there may be the costs to the subjects’ welfare of the research, financial,
unforeseen between-laboratory differences). In addition, the logistical, and other personnel costs, and the likely benefits
estimates of interest (in particular for the variability) might that will stem from having performed a good test of the
not actually be reported (e.g., Kilkenny et al. 2009), or, if hypothesis of interest).
they apparently are, there may nevertheless be some ambiguity The manner in which this is done (i.e., the actual sample
as to whether they refer to the specific parameter of interest, for size calculation) depends on the proposed experimental de-
instance a component of the variance that is not of primary sign, including the type of variables (continuous, categorical,
concern to you (e.g., Lenth 2001). In addition, if prior studies etc.) to be collected, structure of dependency (for example,
have low power (are underpowered)—a pervasive problem in any nesting), probability of missing data, the inferences of
many research disciplines—any true effect sizes that are found key interest, and so on. For example, if the main concern is
are more likely to be overinflated. This has been referred to as to test whether there is a significant difference between treat-
the “winner’s curse”: since underpowered experiments cannot ments on outcome variable y, then this requires a sample size
detect small effect sizes, then in such cases only samples that, calculation to be conducted specific to that inference. If, in the
by the nature of random variation, find larger effect sizes (than same experiment, the researcher is also interested in whether
the underlying true small effect sizes) will be statistically sig- another explanatory variable significantly predicts y, or in es-
nificant and reported, resulting in true effect size magnitudes timating the between-animal variance, etc., then these will not
being overestimated (Button et al. 2013). be estimated with the same power as the treatment-related
Preliminary pilot studies may help researchers derive these difference (i.e., each of these research questions, despite the
estimates, under conditions comparable to those in which the fact they all relate to data from the same experiment, requires
proposed study will be undertaken, although of course there its own unique sample size calculation). The more complex
will be considerable uncertainty (given the small sample size) the experiment and analysis required, the more unknown var-
associated with any estimates of effect size and variance iabilities that need estimating and the more complex the sam-
gleaned from them (e.g., Kraemer et al. 2006; Thabane ple size calculation.
et al. 2010). Note that one possible solution to deriving cer- There are a considerable range of useful pedagogic re-
tain estimates is to consider the effect size relative to the var- sources to aid the researcher in calculating samples sizes
iability, and set this to a particular value based on whether the and effect sizes (e.g., including but not confined to Cohen
effect is hypothesized to be “small,” “medium,” or “large” 1992; Lenth 2001; Nakagawa and Cuthill 2007; Krzywinski
( following Cohen’s [1988] original demarcation). If one and Altman 2013b; Howell 2013 [e.g., Chapter 8]). With
uses a standardized effect size such as this, it is then not nec- regard to software, there are a variety of useful programs
essary to estimate the variability directly as it will cancel out available, for instance, InVivoStat (Clark et al. 2012),
in the power calculations, making it analytically attractive. G*Power (Faul et al. 2007, 2009), PiFace (Lenth 2006-9),
However, note that some, including Cohen himself, recom- PS (Dupont and Plummer 1990), and nQuery Advisor +
mend this approach be used with a degree of caution and as nTerim (Statistical Solutions 2014), to name but a few.
a last resort (e.g., Ellis 2010; see also Lenth 2001 for a robust Amongst them is our own MLPowSim (Browne et al.

480 ILAR Journal


2009), designed in the main for more complex experiments. higher-level units unnecessarily dispenses with information
MLPowSim calculates power by simulating data (“virtual that can be more efficiently analyzed in the same analysis. Hi-
studies”) based on parameters set by the user. Whilst designed erarchical designs such as these pose the question of what is
for research scenarios over which the researcher typically has the optimum level of replication at each level of the design to
considerably less control than in a laboratory environment, achieve a certain power (Bate and Clark 2014 [e.g., note the
this does mean it is a useful resource for any researcher wish- example on p.105]; Browne et al. 2009; RJ Tempelman, un-
ing to estimate sample sizes for certain more complex de- published observations).
signs, including multilevel models with missing data and a Increasing the (standardized) effect size might in certain
lack of balance. experiments be achieved by choosing a more extreme treat-
ment (e.g., Krzywinski and Altman 2013b). Such aspects
of the design may be naturally determined by what is clinical-
Increasing Power and the Use of Efficient ly relevant, or constrained by ethical considerations if a more
Designs extreme treatment is likely to induce an increasingly negative

Downloaded from [Link] by guest on 11 July 2025


outcome for the animal. More generally, the choice of out-
To increase the power of a proposed experiment, then, the re- come measure will have a very important bearing on sample-
searcher can either (a) increase the sample size, (b) decrease size calculations. It naturally needs to be a sensitive proxy of
the variation in the current experiment, and/or (c) investigate the underlying construct of interest, and thus the choice will
alternative measures that will answer the same research hy- reflect the performance of outcome measures in prior work;
pothesis but result in an increased (standardized) effect size. scientific considerations of mechanism; as well as how inva-
As mentioned earlier, there may be occasions in which a sive (with respect to the 3Rs’ objective of refinement), and
sample-size calculation implies that a greater number of sub- also logistically and financially feasible, taking the measure
jects should be used than a researcher may have assumed be- will be.
forehand. On face value, this is contrary to the 3Rs’ objective With regard to variation, known (or possible) sources of
of reduction, but not necessarily so if it results in an apprecia- variance can be controlled for (and/or investigated) in the ex-
bly better test of the hypothesis. For example, by increasing perimental design by appropriate use of blocking, inclusion
power, one increases the chance of rejecting the null hypo- of factors measuring them, choice of subjects, etc. (e.g.,
thesis when it is false and increases the likelihood that any Bate and Clark, 2014; Festing et al. 2002). Preferentially tar-
statistically significant result found actually relates to a true geting (i.e., controlling for) the sources of greatest residual
effect. Indeed, given the chronic problem of underpowered variance is naturally a sensible strategy. Note that, in a nested
studies, the need to reconsider how resources are best de- design, the overall variance is made up of components at each
ployed is pressing (see Button et al. 2013 for a discussion). level of the hierarchy, so if a researcher takes more than one
Alternatively, if a sample size calculation implies that a great- measure from each animal, and the within-animal variance is
er sample size than originally envisaged might be necessary greater than the between-animal variance, then this would
to achieve a certain power, this may imply—having weighed suggest he or she is best advised to first address sources of
the costs and benefits—that the experiment is simply not within-animal variation (e.g.; Bate and Clark 2014; RJ Tem-
worth conducting. pelman, unpublished observations). A change in design, such
Whilst underpowered, rather than overpowered, experi- as employing combinations of factors rather than testing one
ments tend to characterize laboratory-based animal studies, factor at a time, can allow more information to be gleaned
it’s nevertheless important to note that the benefits, in terms from the same number of animals, and can be designed
of the uncertainty with which any subsequent inferences are with a relatively small number of animals per combination
drawn, will diminish, per subject, as sample
pffiffiffi size is increased. of factors by reaping the benefits of hidden replication (e.g.,
Since the SEM is proportional to 1= n for known standard Bate and Clark 2014; Festing et al. 2002; Shaw et al. 2002).
deviation, precision will increase at a slower rate than data
collection (Krzywinski and Altman 2013a), and so the matter
is one of careful judgment (Bacchetti et al. 2005). The sample Other Ways to Achieve Net Reduction
size can also be manipulated without increasing the number in Animal Usage
of animals, for example, by taking more samples from those
animals. Any design in which observations are clustered in As well as ensuring sufficient power, the researcher can use
this manner needs to be modeled using appropriate statistical the tools of experimental design and statistics to address
methods; all else being equal, measurements taken from the bias stemming from other sources. Naturally, if a treatment ef-
same animal are likely to be more correlated than measure- fect on an outcome variable is found in a given experiment,
ments taken from different animals, and standard errors will then this is only of interest if it doesn’t represent the effect of
be underestimated if this lack of independence is not taken unacknowledged factors (i.e., it is internally valid and not in-
into account. Hierarchical structures can be encountered in fluenced by bias unwittingly introduced in the experimental
other scenarios too: for example, animals within cages, or design or analysis). This is obviously a considerable concern
even replicates across labs. In such instances, aggregating when considering how we might reduce the overall numbers
(i.e., taking a summary of an outcome measure) across of animals in experiments, as the better each test of a

Volume 55, Number 3, doi: 10.1093/ilar/ilu044 2014 481


hypothesis is, the fewer the “false leads” (with associated an- study of mice applied to humans). A study that has poor ex-
imal usage) generated, the less corrective replication of previ- ternal validity is a study in which animals were used unnec-
ously compromised experimental work is needed, and so on. essarily and, reflecting the ethical implications (with regard to
There is a large amount of valuable literature detailing both animals and humans), there is an important and lively
methods to ensure bias is minimized, including resources spe- debate concerning the success, or otherwise, of this venture
cific to laboratory animal studies (e.g., Bate and Clark 2014; (e.g., Hackam and Redelmeier 2006; Perel et al. 2007; van
Festing and Altman 2002; Festing et al. 2002). In essence, der Worp et al. 2010).
much of this concerns protecting us against ourselves (i.e., Finally, accurate and clear reporting of experimental proto-
ensuring our own cognitive and behavioural biases do not cols, analytical methods, and subsequent results ensures ex-
jeopardize the integrity of our experimental results). For ex- periments are not reproduced unnecessarily (i.e., erroneously
ample, the importance of randomization, blinding (not only conducted under the assumption they have not been per-
of researchers, but of technical and veterinary staff too, e.g., formed before), but also facilitates the faithful reproduction
Bate and Clark 2014), and identifying and controlling for the of experimental protocols when it is deemed desirable and
a more informed critique of the study’s findings. As such,

Downloaded from [Link] by guest on 11 July 2025


influence of known confounding variables via the use of
blocking is rightly stressed. there is a growing trend toward providing access, alongside
The researcher may have further sway over the inferences a summarizing article, to the materials necessary to reproduce
reported: so-called “researcher degrees of freedom” the analysis reported, such as an annotated dataset and com-
(Simmons et al. 2011), or “p-hacking” may be employed to mand syntax used to analyze it (e.g., Diggle and Zeger 2010;
render a result more significant, for example, than alternative Groves and Godlee 2012). Furthermore, in some disciplines
analytical strategies might have found. This may be in the there is a movement to publish protocols prior to data collec-
choice of variables reported, model selection, the choice of tion and analysis, helping guard against any undue sway attrib-
a statistical test, the treatment of outliers, choice of transfor- utable to researcher “degrees of freedom,” and guaranteeing
mations, and so on. Some of the terminology used to describe the publication of results regardless of whether they are null
such practices could be construed as implying a degree of or not (e.g., Chambers 2013; Munafo and Strain 2014).
underhandedness on the part of a researcher, fishing for Transparency is also important in facilitating a critical ap-
significance, but whilst that may occasionally be the case, praisal of the experiment, as is the choice of which statistics
many such analytical choices are generally-accepted and and charts to present. The latter can, for example, help render
appear reasonable (Gelman and Loken 2013). Given suffi- any future meta-analysis considerably more tractable, thus
cient interrogation using commonplace methods of analysis, ensuring the experiment can be put to further use as a data
many datasets will yield p values under 0.05 (e.g., Bennett point in a larger retrospective systematic appraisal (e.g.,
et al. 2009; Wagenmakers et al. 2011). Kilkenny et al. 2010; Nakagawa and Cuthill 2007). Further-
As such, phenomena such as the chronic problem of under- more, some intuitively reasonable urges, such as conducting
powered experiments (Button et al. 2013), a disproportionate and reporting post hoc power analyses, on closer inspection
prevalence of published p values that are marginally under prove to offer nothing more than the p values already reported
0.05 (Masicampo and Lalande 2012), results not in keeping (e.g., Bate and Clark 2014; Hoenig and Heisey 2001;
with the experimental hypothesis residing in metaphorical file Nakagawa and Foster 2004).
drawers rather than the pages of journals (e.g., Dwan et al. To end on a pragmatic note, it is worth noting that whilst
2013; Rosenthal 1979; ter Riet et al. 2012) likely reflect a the very valuable body of work advocating, and advising
variety of investigator-level and more systemic factors. For on, better standards of experimental design, analysis, and re-
instance journals’ publication policies valuing positive results porting continues to grow, translating this into better practice
and novelty above null findings and study replication, the is a considerable challenge. It is commonplace to hear re-
bearing of high-impact publications on the trajectory of scien- searchers’ disquiet at being advised, sometimes apocalypti-
tists’ careers, and also perhaps an investigator’s own satisfac- cally, that the methods they have dutifully and carefully
tion that their efforts have indeed found evidence in keeping learned might actually retard scientific progress, and that
with their pet hypothesis, or the simple urge to construct a the larger community in which they work has gross systemic
coherent narrative, rather than presenting a more confusing failures (e.g., Button et al. 2013; Ioannidis 2005; Nakagawa
picture harder to explain (Anon 2013) all shape what is pub- and Cuthill 2007). It is naturally important to alert the com-
lished. The trend for increasing transparency (see below) munity to any fundamental problems that exist in the conduct
seeks to, at least partly, remedy this. of science, but also, of course, to remain mindful that im-
External validity, on the other hand, concerns the extent to provements in experimental design, analysis, and reporting
which the results from a given experiment can be applied to (and any positive impact on the 3Rs that ensues) can only
the population about which the investigators ultimately wish be made if those conducting the work are carried along
inferences to be made. Sometimes this target population will with that movement (e.g., Sharpe 2013). To this end, efforts
be a more obviously close match to the sample studied, as in to incentivize transparency and study replication, and to pro-
some ethological studies, the assessment of certain veterinary mote education in experimental design and statistical analy-
interventions, etc. In many other instances, generalizations sis, are increasingly being recognized as important drivers
will be sought across species (e.g., inferences drawn from a for change (Munafo et al. 2014).

482 ILAR Journal


Replacement likely to involve suffering. As such, and pending the
possibility of any future replacement, there may be scope to
To the extent that tests involving the absolute replacement of at least alleviate the severity, or shorten the duration, of the
animal subjects are as reliable and valid as the animal-based induced disorder (e.g., Littin et al. 2008; Wolfensohn et al.
experiments they replaced, this is the ultimate manifestation 2013; Workman et al. 2010).
of a humane corrective to current procedure. If they are not Otherwise, the reported success in training animals in a
as valid or reliable, however, then how humane the proposed manner that facilitates collecting physiological data, or con-
replacement is becomes less certain, particularly if the exper- ducting general veterinary checks—likely reducing any stress
iment is designed as a model for medical or toxicity testing. associated with capture and restraint—has led to some advo-
With regard to relative replacement with subjects thought cating such techniques be more widely adopted (e.g., Perlman
less likely to suffer as a result of the proposed experimental et al. 2012). Even relatively modest changes in the manner in
conditions, establishing evidence for grades of subjective suf- which mice are handled, for example, can reduce behavior as-
fering across animals can be—philosophically (e.g., Nagel sociated with anxiety and fear (Gouveia and Hurst 2013;

Downloaded from [Link] by guest on 11 July 2025


1974) and scientifically (e.g., Sherwin 2001)—a challenging Hurst and West 2010).
endeavor, but a necessary one, and so by-and-large consensus Stress may introduce noise into an experiment, and may
opinions are drawn based on a consideration of factors such as even interact with treatments of interest; therefore, in addition
apparent behavioral and cognitive complexity and phylogeny. to its obvious intrinsic value, refinement can result in an ex-
So, for example, certain legislation, such as the Animals periment being a better test of a hypothesis, insofar as vari-
Scientific Procedures Act (ASPA) in Great Britain, draws ability and/or bias can be reduced. Trying to understand the
boundaries both between species, and between developmen- laboratory world from the animals’ point of view can not
tal stages within species, which divides those protected by the only further the aim of refinement from a welfare perspective
Act from the remainder. but can also help the researcher detect potential confounders
The development of alternative methods in which protect- and sources of noise, be it auditory frequencies that we are not
ed animals are replaced as subjects is a rapidly developing privy to but rodents are (e.g., Burman et al. 2007); fluorescent
area, somewhat dictated by the speed of technological ad- light flicker that, unlike captive birds, we cannot perceive
vance (Balls 2007; Langley et al. 2007). Validation protocols (e.g., Greenwood et al. 2004); and so on. Indeed, the subtle
assessing alternative methods, such as those of the Organisa- sensitivity of laboratory animals’ responses to the environ-
tion for Economic Co-operation and Development (OECD ment around them can often surprise and fascinate (e.g.,
2005), assess the relevance and reliability of a method Sorge et al. 2014).
with regard to a defined purpose. Here relevance is akin to
predictive validity (i.e., the extent to which the test predicts
the effect of interest in the target population), whereas reli- Conclusion
ability refers to a test’s reproducibility (across laboratories
A carefully considered appreciation of experimental design
and across time). These definitions are substrate neutral.
and statistics prior to data collection is key when seeking to
That is, they do not fall foul of the high fidelity fallacy (Balls
successfully apply the 3Rs, and indeed many of the principles
2007; Russell and Burch 1959): the assumption that (e.g.)
central to helping realize their objectives—for instance, con-
placental mammalian animal models will necessarily be bet-
ducting valid, reliable, and efficient experiments; clearly and
ter models of within-human effects as a consequence of their
transparently reporting findings; and ensuring that an appre-
greater apparent similarity (compared to alternatives) to us.
ciation and understanding of animal welfare plays a central
All else being equal, the substrate of an experiment—be it
role in laboratory practice—are to the betterment of research
in silico, in vitro, etc.—is of less importance than the ability
per se.
of the experiment to reliably predict the effect of interest (i.e.,
“what is essential is not that a model system ‘looks like’
the system of interest, but that it behaves like it” Acknowledgments
[Richmond 2010]).
We thank two anonymous referees for helpful and construc-
tive comments.
Refinement
With regard to refinement, Russell and Burch (1959) distin- References
guished the direct inhumanity that may result from the ex-
perimental procedures themselves, from the contingent Anon. 2013. Should scientists tell stories? Nat Methods 10:1037–1037.
inhumanity that may result from “the infliction of distress Bacchetti P, Wolf LE, Segal MR, McCulloch CE. 2005. Ethics and sample
size. Am J Epidemiol 161:105–110.
as an incidental and inadvertent by-product of the use of
Balls M. 2007. Alternatives to animal experiments: Time to focus on replace-
the procedure, which is not necessary for its success.” ment. AATEX 12:145–154.
With regard to the first of these, animal models designed Bate ST, Clark RA. 2014. The Design and Statistical Analysis of Animal
to simulate human-based diseases are, by their very nature, Experiments. Cambridge, UK: Cambridge University Press.

Volume 55, Number 3, doi: 10.1093/ilar/ilu044 2014 483


Bennett CM, Baird AA, Miller MB, Wolford GL. 2009. Neural correlates of rescent lighting affect the welfare of captive European starlings? Appl
interspecies perspective taking in the post-mortem Atlantic Salmon: An Anim Behav Sci 86:145–159.
argument for multiple comparisons correction. Poster Presented at the Groves T, Godlee F. 2012. Open science and reproducible research. Br Med J
15th Annual Meeting of the Organization for Human Brain Mapping, 344:e4383.
San Francisco, USA. Hackam DG, Redelmeier DA. 2006. Translation of research evidence from
Browne WJ, Golalizadeh Lahi M, Parker RMA. 2009. A Guide to Sample animals to humans. JAMA 296:1731–1732.
Size Calculations for Random Effect Models via Simulation and Hoenig JM, Heisey DM. 2001. The abuse of power: The pervasive fallacy of
the MLPowSim Software Package. Bristol, UK: Centre for Multilevel power calculations for data analysis. Am Stat 55:19–24.
Modelling, University of Bristol. Home Office. 2013. Annual Statistics of Scientific Procedures on Living
Buchanan-Smith HM, Rennie AE, Vitale A, Pollo S, Prescott MJ, Animals, Great Britain 2012. London: The Stationery Office.
Morton DB. 2005. Harmonising the definition of refinement. Anim Howell DC. 2013. Statistical Methods for Psychology. Canada: Wadsworth
Welfare 14:379–384. Cengage Learning.
Burman OHP, Ilyat A, Jones G, Mendl M. 2007. Ultrasonic vocalizations as Hurst JL, West RS. 2010. Taming anxiety in laboratory mice. Nat Methods
indicators of welfare for laboratory rats (Rattus norvegicus). Appl Anim 7:825–U1516.
Behav Sci 104:116–129. Ioannidis JPA. 2005. Why most published research findings are false. Plos
Button KS, Ioannidis JPA, Mokrysz C, Nosek BA, Flint J, Robinson ESJ, Med 2:696–701.

Downloaded from [Link] by guest on 11 July 2025


Munafo MR. 2013. Power failure: Why small sample size undermines Ipsos MORI. 2012. Views On the Use of Animals in Scientific Research.
the reliability of neuroscience. Nature Rev Neurosci 14:365–376. London: Ipsos MORI.
Chambers CD. 2013. Registered Reports: A new publishing initiative at Kilkenny C, Browne WJ, Cuthill IC, Emerson M, Altman DG. 2010. Improv-
Cortex. Cortex 49:609–610. ing bioscience research reporting: The ARRIVE guidelines for reporting
Clark RA, Shoaib M, Hewitt KN, Stanford SC, Bate ST. 2012. A comparison animal research. Plos Biol 8:e1000412.
of InVivoStat with other statistical software packages for analysis of data Kilkenny C, Parsons N, Kadyszewski E, Festing MFW, Cuthill IC, Fry D,
generated from animal experiments. J Psychopharmacol 26:1136–1142. Hutton J, Altman DG. 2009. Survey of the quality of experimental de-
Cohen J. 1988. Statistical Power Analysis for the Behavioral Sciences. Hill- sign, statistical analysis and reporting of research using animals. Plos
sdale, NJ: Erlbaum. One 4:e7824.
Cohen J. 1992. A power primer. Psychol Bull 112:155–159. Kraemer HC, Mintz J, Noda A, Tinklenberg J, Yesavage JA. 2006. Caution
Dawkins MS. 1990. From an animal’s point of view - motivation, fitness, and regarding the use of pilot studies to guide power calculations for study
animal welfare. Behav Brain Sci 13:1–9. proposals. Arch Gen Psychiatry 63:484–489.
Diggle PJ, Zeger SL. 2010. Editorial. Biostatistics 11:375–375. Krzywinski M, Altman N. 2013a. Importance of being uncertain.
Dupont WD, Plummer WD. 1990. Power and sample size calculations: A re- Nat Methods 10:809–810.
view and computer program. Control Clin Trials 11:116–128. Krzywinski M, Altman N. 2013b. Power and sample size. Nat Methods
Dwan K, Gamble C, Williamson PR, Kirkham JJ; Reporting Bias Group. 10:1139–1140.
2013. Systematic review of the empirical evidence of study publication Langley G, Evans T, Holgate ST, Jones A. 2007. Replacing animal experi-
bias and outcome reporting bias - an updated review. Plos One 8:e66844. ments: choices, chances and challenges. BioEssays 29:918–926.
Ellis PD. 2010. The Essential Guide to Effect Sizes: Statistical Power, Lenth RV. 2001. Some practical guidelines for effective sample size determi-
Meta-Analyses and the Interpretation of Research Results. Cambridge, nation. Am Stat 55:187–193.
UK: Cambridge University Press. Lenth RV. 2006–9. Java Applets for Power and Sample Size [Internet]. Avail-
European Commission. 2013. Report from the Commissin to the Council and able from: [Link]
the European Parliament: Seventh Report on the Statistics on the Number Littin K, Acevedo A, Browne W, Edgar J, Mendl M, Owen D, Sherwin C,
of Animals used for Experimental and other Scientific Purposes in the Wurbel H, Nicol C. 2008. Towards humane end points: behavioural
Member States of the European Union. Brussels: European Commission. changes precede clinical signs of disease in a Huntington’s disease
Faul F, Erdfelder E, Buchner A, Lang A-G. 2009. Statistical power analyses model. Proc R Soc Lond [Biol] 275:1865–1874.
using G*Power 3.1: Tests for correlation and regression analyses. Behav Masicampo EJ, Lalande DR. 2012. A peculiar prevalence of p values
Res Meth 41:1149–1160. just below .05. Q J Exp Psychol 65:2271–2279.
Faul F, Erdfelder E, Lang A-G, Buchner A. 2007. G*Power 3: A flexible stat- Munafo M, Noble S, Browne WJ, Brunner D, Button K, Ferreira J,
istical power analysis program for the social, behavioural, and biomedical Holmans P, Langbehn D, Lewis G, Lindquist M, Tilling K,
sciences. Behav Res Meth 39:175–191. Wagenmakers E-J, Blumenstein R. 2014. Scientific rigor and the art of
Fenwick N, Griffin G, Gauthier C. 2009. The welfare of animals used in sci- motorcycle maintenance. Nat Biotech 32:871–873.
ence: How the “Three Rs” ethic guides improvements. Can Vet J Rev Munafo MR, Strain E. 2014. Registered Reports: A new submission format at
50:523–529. drug and alcohol dependence. Drug and Alcohol Depend 137:1–2.
Festing MFW, Altman DG. 2002. Guidelines for the design and statistical Nagel T. 1974. What is it like to be a bat. Philosoph Rev 83:435–450.
analysis of experiments using laboratory animals. ILAR J 43:244–258. Nakagawa S, Cuthill IC. 2007. Effect size, confidence interval and statistical
Festing MFW, Overend P, Gaines Das R, Cortina Borja M, Berdoy M. 2002. significance: a practical guide for biologists. Biol Rev 82:591–605.
The Design of Animal Experiments: Reducing the Use of Animals in Nakagawa S, Foster TM. 2004. The case against retrospective statistical pow-
Research Through Better Experimental Design. London: Royal Society er analyses with an introduction to power analysis. Acta Ethologica
of Medicine Press. 7:103–108.
Gallup. 2014. Gallup News Service - Gallup Poll Social Series: Values and NC3Rs. 2014. The 3Rs [Internet]. Available from: [Link]
Beliefs - Final Topline. Washington, DC: Gallup. the-3rs
Gelman A, Loken E. 2013. The garden of forking paths: why multiple com- OECD. 2005. Guidance Document no. 34 on the Validation and International
parisons can be a problem, even when there is no “fishing expedition” or Acceptance of New or Updated Test Methods for Hazard Assessment.
“p-hacking” and the research hyptothesis was posited ahead of time. [In- Paris: OECD.
ternet]. Available from: [Link] Perel P, Roberts I, Sena E, Wheble P, Briscoe C, Sandercock P, Macleod M,
unpublished/p_hacking.pdf Mignini LE, Jayaram P, Khan KS. 2007. Comparison of treatment effects
Gouveia K, Hurst JL. 2013. Reducing mouse anxiety during handling: Effect between animal experiments and clinical trials: systematic review. Br
of experience with handling tunnels. Plos One 8:e66401. Med J 334:197–200.
Greenwood VJ, Smith EL, Goldsmith AR, Cuthill IC, Crisp LH, Perlman JE, Bloomsmith MA, Whittaker MA, McMillan JL, Minier DE,
Walter-Swan MB, Bennett ATD. 2004. Does the flicker frequency of fluo- McCowan B. 2012. Implementing positive reinforcement animal

484 ILAR Journal


training programs at primate laboratories. Appl Anim Behav Sci Statistical Solutions. 2014. nQuery Advisor + nTerim. Cork, Ireland: Statisti-
137:114–126. cal Solutions.
Richmond J. 2010. The Three Rs. In: Hubrecht R, Kirkwood J, editors. The Taylor K, Gordon N, Langley G, Higgins W. 2008. Estimates for
UFAW Handbook on the Care and Management of Laboratory and Other worldwide laboratory animal use in 2005. Altern Lab Anim: ATLA
Research Animals. 8th ed. p. 5–22. 36:327–342.
Rosenthal R. 1979. The file drawer problem and tolerance for null results. ter Riet G, Korevaar DA, Leenaars M, Sterk PJ, Van Noorden CJF,
Psychol Bull 86:638–641. Bouter LM, Lutter R, Elferink RPO, Hooft L. 2012. Publication bias
Russell WMS, Burch RL. 1959. The Principles of Humane Experimental in laboratory animal research: A survey on magnitude, drivers, conse-
Technique. London: Methuen & Co. Ltd. quences and potential solutions. Plos One 7:e43404.
Sena ES, van der Worp HB, Bath PMW, Howells DW, Macleod MR. 2010. Thabane L, Ma J, Chu R, Cheng J, Ismaila A, Rios LP, Robson R,
Publication bias in reports of animal stroke studies leads to major over- Thabane M, Giangregorio L, Goldsmith CH. 2010. A tutorial on pilot
statement of efficacy. Plos Biol 8:e1000344. studies: the what, why and how. BMC Med Res Methodol 10:1.
Sharpe D. 2013. Why the resistance to statistical innovations? bridging the van der Worp HB, Howells DW, Sena ES, Porritt MJ, Rewell S, O’Collins V,
communication gap. Psychol Meth 18:572–582. Macleod MR. 2010. Can animal models of disease reliably inform human
Shaw R, Festing MFW, Peers I, Furlong L. 2002. Use of factorial designs to studies? Plos Med 7:e1000245.
optimize animals experiments and reduce animal use. ILAR J 43:223–232. Wagenmakers EJ, Wetzels R, Borsboom D, van der Maas HLJ. 2011. Why

Downloaded from [Link] by guest on 11 July 2025


Sherwin CM. 2001. Can invertebrates suffer? Or, how robust is psychologists must change the way they analyze their data: the case of
argument-by-analogy? Anim Welfare 10:S103–S118. psi: comment on Bem (2011). J Pers Soc Psychol 100:426–432.
Simmons JP, Nelson LD, Simonsohn U. 2011. False-positive psychology: Wolfensohn S, Hawkins P, Lilley E, Anthony D, Chambers C, Lane S,
undisclosed flexibility in data collection and analysis allows presenting Lawton M, Robinson S, Voipio H-M, Woodhall G. 2013. Reducing suf-
anything as significant. Psychol Sci 22:1359–1366. fering in animal models and procedures involving seizures, convulsions
Sorge RE, Martin LJ, Isbester KA, Sotocinal SG, Rosen S, Tuttle AH, and epilepsy. J Pharmacol Toxicol Meth 67:9–15.
Wieskopf JS, Acland EL, Dokova A, Kadoura B, Leger P, Workman P, Aboagye EO, Balkwill F, Balmain A, Bruder G, Chaplin DJ,
Mapplebeck JC, McPhail M, Delaney A, Wigerblad G, Schumann AP, Double JA, Everitt J, Farningham DAH, Glennie MJ, Kelland LR,
Quinn T, Frasnelli J, Svensson CI, Sternberg WF, Mogil JS. 2014. Olfac- Robinson V, Stratford IJ, Tozer GM, Watson S, Wedge SR, Eccles SA,
tory exposure to males, including men, causes stress and related analgesia Navaratnam V, Ryder S. 2010. Guidelines for the welfare and use of an-
in rodents. Nat Methods 11:629–632. imals in cancer research. Br J Cancer 102:1555–1577.

Volume 55, Number 3, doi: 10.1093/ilar/ilu044 2014 485

You might also like