Methods
Examining Test-Retest
Reliability
An Intra-Class
Correlation Approach
Miaofen Yen ▼ Li-Hua Lo
Background: There are limitations of inter-class correlation (usually, the Pearson’s relations. First, the purpose of the
product-moment correlation) in evaluating reliability of an instrument. An intra-class Pearson’s product-moment correlation
correlation approach would be appropriate to examine the test-retest reliability of an is to determine the relationship be-
instrument. tween two variables. Theoretically, it is
Objectives: To introduce an intra-class correlation approach to examine the test-retest not appropriate to apply the correla-
reliability of an instrument. tion to a case of two measures from the
same variable, such as the test and the
Method: The method of constructing an intra-class correlation coefficient was presented.
retest scores for a concept. Second, it is
Results: The instrument with visual analog scales to measure the perceived competence difficult to determine the test-to-test
to practice breast self-examination and perceived barriers to breast self-examination variation when multiple tests are
was tested. The intra-class correlation coefficient was 0.640. administered. For example, if a con-
Conclusion: The intra-class correlation is an alternative to test the reliability of an instru- cept is measured three times repeatedly,
ment and it is more sensitive to the detection of systematic error. three scores are obtained. Tradition-
Key Words: instrumentation • intra-class correlation • test-retest reliability ally, three correlation coefficients for
every two of these three scores would
be calculated and examined. However,
the correlation coefficient of all three
R eliability testing is important in
the evaluation of instruments
that are used to measure concepts in
intra-class correlation (ICC) that pro-
vides a solution to limitations of evalu-
ating test-retest reliability. Intra-class
scores cannot be generated at the same
time. Intra-class correlations resolve
this problem when three scores are
nursing research. “Reliability can be correlation, also known as generaliz- examined simultaneously. Third, the
defined as the degree of consistency ability coefficient (Lo, 1996), originated Pearson’s product-moment correlation
between two measures of the same from the generalizability theory and has cannot detect the existence of a sys-
thing” (Mehrens & Lehmann, 1991, been thoroughly discussed by Burns tematic error. In Figure 1 the test scores
p. 249). Statistically, the reliability of (1998). This approach uses analysis of represent the knowledge about breast
a measure is defined as the proportion variance and allows the calculation of self-examination (BSE). Both studies
of the observed score variance due to error variances from each source. For show perfect correlations between the
the true scores, that is, the ratio of example, this approach can be used to test and the retest scores. Nevertheless,
true score variance to total score vari- investigate within person variation (i.e., the differences between the test and the
ance. A common practice of test-retest individual differences), and test-to-test retest scores for each subject are 2 in
reliability in nursing research is to cal- variation (i.e., the difference between study A and 30 in study B. Even
culate the Pearson’s product-moment the test and retest scores, separately).
correlation between the test and the
retest scores. However, Safrit (1976) Limitations of Pearson’s
identified at least three limitations to Miaofen Yen, PhD, RN, is Associate Pro-
Product-Moment fessor, National Cheng Kung University,
this practice, which are discussed in Tainan, Taiwan, ROC.
Correlation
detail in the next section. Li-Hua Lo, PhD, RN, is Associate Profes-
The purpose of this presentation is There are at least three limitations in sor, National Cheng Kung University,
to describe a method to calculate the using Pearson’s product-moment cor- Tainan, Taiwan, ROC.
Nursing Research January/February 2002 Vol 51, No 1 [Link] 59
60 Recruitment and Retention Nursing Research January/February 2002 Vol 51, No 1 [Link]
correlation coefficient leaves systematic
errors undetected. The ICC coefficients
for study A and study B are 0.991 and
0.024, respectively. Both studies, A and
B, have systematic errors. The Pearson’s
product-moment correlation coeffi-
cients do not detect the systematic
errors. On the other hand, the ICC coef-
ficients reflect the magnitude of the
errors. The Pearson’s product-moment
correlation would not reflect the relia-
bility of the test as well as ICC.
Intra-Class Correlation
Approach
According to Shrout and Fleiss (1979)
three issues should be considered
when using ICC approach. First, the
design should be a reliability study
rather than a correlational study. For
example, in a test-retest reliability
study detecting the effect of various
sources of errors rather than the cor-
relation between two measures is the
purpose. Factors that may generate
the measurement errors, such as time,
rater, or sex, should be taken into con-
sideration.
Second, it is necessary to select an
appropriate statistical model for a
reliability study. Selection of a one-
way or a two-way random model
would depend on the sources of vari-
ance taken into account. When the
sources of errors cannot be separated
and are pooled, a one-way model
should be applied. If the error vari-
ances come from various sources and
can be separated, a two-way random
model should be considered. For
example, when evaluating an instru-
ment for the knowledge of BSE, a
source of error, educational prepara-
tion, might influence the stability of
the instrument. In this situation, a
two-way random model to evaluate
the instrument is preferred because
the error variances may come from
educational preparation and other
FIGURE 1. Two studies of test and retest measures. unknown sources. When one cannot
separate the error variance of educa-
tional preparation from others, then,
a one-way random model should be
though the test and the retest scores are retest scores for study A and study B are selected.
perfectly correlated, that is, the Pear- perfectly correlated. A larger correlation Additionally, the error variances
son’s correlation coefficients are both coefficient between the test and the from each possible source would be
equal to 1.00, a systematic error exists retest scores means the tested instru- influenced by the characteristics of
in study B. ment is more reliable than one that variables (random or fixed) according
Based on the Pearson’s product- results in a smaller correlation coeffi- to the study design. A random variable
moment correlation, the test and the cient. However, in the example, a large can be defined as “a variable that
Nursing Research January/February 2002 Vol 51, No 1 [Link] Recruitment and Retention 61
The test-retest reliability of the
TABLE 1. Data for Intra-Class Correlation Coefficients (N = 10) scale was completed using the ICC
approach. The variable of interest was
Subjects Test Scores Retest Scores measured twice for each subject
(Table 1). Using the analysis of vari-
1 129.9 113.4
ance to calculate the intra-class corre-
2 132.7 128.4 lation coefficient, two assumptions
3 118.9 122.4 were considered. The normality and
4 111.4 91.6 equal variance of the sample were
5 145.3 149.8 evaluated through descriptive statis-
6 148.3 140.6
tics and Levene’s tests. The normality
test was performed using Fisher’s
7 107.8 122.7 skewness coefficient (Pett, 1997, p.
8 141.9 134.2 36), which is expressed as equation 1.
9 141.9 121.3
Fisher Skewness Coefficient =
10 126.8 142.6
Skewness
Standard Error Skewness
assumes its values according to a set of measures and the number of measures The Fisher skewness coefficient for
probabilities” (Crocker & Algina, ranges from 1 to k. The one-way ran- the test scores (first measure) equals
1986, p. 107). For example, the dom model is discussed below. ⫺0.584 (⫺0.401/0.687) and the retest
knowledge score of BSE ranges from 0 scores (second measure) equals ⫺1.147
to 100. The possible scores for each (⫺0.788/0.687). These two coefficients
Example Demonstration
subject have an equal opportunity are both between ⫹1.96 and ⫺1.96
ranging from 0 to 100; therefore, the The example is from a study to evalu- indicating that the distributions of
knowledge score is a random variable. ate the reliability and validity of an scores are not significantly different
If the sources of variances are gener- instrument to measure the perceived from a normal distribution. Next, the
ated by random variables, a random competence to practice BSE and per- assumption of equal variances was
model should be applied. A variable ceived barriers to BSE. This instru- tested using SPSS® for Windows. One-
(i.e., sex), is fixed if only the conditions ment was adapted from a 20-item way analysis of variance with Levene’s
(i.e., male and female), of the variable five-point Likert-type format scale. test gave the value of Levene’s statistic
are of interest for a decision (Roe- The major modification from the orig- for the test and retest scores (0.049, p
broeck, Harlaar, & Lankhorst, 1993). inal scale was using visual analog ⫽ .828) indicating that the two mea-
Third, the number of measures to scale to replace the Likert-type format sures had equal variances.
be conducted in the study is a consid- in order to get more information The Pearson’s product-moment
eration. For example, the notation of about the perceived competence and correlation coefficient between the
ICC can be expressed as ICC(1, k), barriers to BSE. A group of nurses (N test and retest scores was calculated (r
where “1” denotes the one-way ⫽ 10) completed the instrument twice ⫽ 0.643, p ⫽ .045). The ICC coeffi-
model; “k” denotes the number of in 2 weeks. cient was calculated through SPSS(r)
Table 2. Output From SPSS® for Windows for Intra-Class Correlation Analysis
ICC 95% C.I.
(Test Value = .000) Lower Upper F DF Sig.
Analysis A: Comparing test
scores and retest scores
Single Measure 0.640 0.093 0.895 4.557 (9,10.0) 0.013
Average Measure 0.781 0.171 0.945 4.557 (9,10.0) 0.013
Analysis B: Comparing test
scores and retest scores*
Single Measure 0.554 ⫺0.040 0.865 3.488 (9,10.0) 0.032
Average Measure 0.713 ⫺0.084 0.928 3.488 (9,10.0) 0.032
ICC = intra-class correlation; C. I. = confidence interval; F = F values; DF = degrees of freedom; Sig. = significance
*A systematic error of 12 points introduced.
62 Recruitment and Retention Nursing Research January/February 2002 Vol 51, No 1 [Link]
for Windows (Table 2). Two ICC correlation is the correlation between assess dependability. Research in Nurs-
two different variables and it is inap- ing & Health, 21, 83-90.
coefficients appear in the output: one Crocker, L., & Algina, J. (1986). Introduc-
is a single measure intra-class correla- propriate to be used in the reliability
analysis. Furthermore, the Pearson’s tion to classical and modern test theory.
tion and the other is an average mea- New York, NY: Holt, Rinehart, & Win-
sure intra-class correlation. In this product-moment correlation coefficient
ston.
example, the interest is in the single cannot detect the existence of system- Lo, S. K. (1996). An empirical discussion
measure intra-class correlation coeffi- atic errors. Therefore, the ICC coeffi- of commonly encountered statistical
cient [ICC(1,1) ⫽ 0.640] because in cient would be the appropriate ap- problems in nursing research. Nursing
reality the instrument would only be proach to evaluate the test and retest Research, 4(3), 298-306. (Chinese).
administered once to a subject at one reliability. ▼ Mehrens, W. A., & Lehmann, I. J. (1991).
Measurement and evaluation in educa-
period of time. If the instrument will
tion and psychology. Philadelphia, PA:
be administered twice or more times Holt, Rinehart and Winston.
Accepted for publication July 12, 2001.
in each time period, then, the average Pett, M. A. (1997). Nonparametric statis-
The authors thank Sing Kai Lo, PhD, Profes-
measure ICC should be reported. tics for health care research. Thousand
sor, Department of Rehabilitation Sciences,
Although consensus among re- The Hong Kong Polytechnic University, for Oaks, CA: Sage.
searchers has not been reached on the suggestions and comments in the process of Roebroeck, M. E., Harlaar, J., & Lank-
significance of the ICC, the higher the developing this article, and Jia-Jer Liou, PhD, horst, G. J. (1993). The application of
estimation, and the better the reliabil- Acting Director, Department of Products & generalizability theory to reliability
Process Development, Taiwan Sugar assessment: An illustration using iso-
ity. Research Institute, for his editing assistance.
In this example, the Pearson’s prod- metric force measurements. Physical
Corresponding author: Li-Hua Lo, PhD, RN, Therapy, 73(6), 386-395.
uct-moment correlation coefficient (r ⫽ Department of Nursing, College of Medicine, Safrit, M. J. (1976). Reliability theory.
0.643) and the ICC coefficient (ICC(1,1) National Cheng Kung University, 1 Univer-
Washington, DC: American Alliance for
⫽ 0.640) were similar for the analysis sity Road, Tainan 70101, Taiwan, ROC (e-
Health, Physical Education, and Recre-
mail: lhlo@[Link]).
A. However, when the systematic error ation.
of 12 points was introduced, the two Shrout, P. E., & Fleiss, J. L. (1979). Intra-
coefficients (r ⫽ 0.643 and ICC(1,1) ⫽ References class correlations: Uses in assessing
0.554) were slightly different. Theoret- Burns, K. J. (1998). Beyond classical relia- rater reliability. Psychological Bulletin,
ically, the Pearson’s product-moment bility: Using generalizability theory to 86(2), 420-422.