Unit III
Using & Interpreting the Coefficient of Reliability
Purpose of the Reliability Coefficient
The nature of the test
Homogeneity vs heterogeneity of test items
Dynamic vs Static characteristics
Restriction vs inflation of range When range of either variable is restricted, r tends to be lower
Odd-even does not work for speed tests
Speed test vs Power test
Split test on the basis of time (t1+t4 and t2+t3)
Criterion-referenced tests Traditional r may not be appropriate as the focus is
(versus Norm-referenced) on mastery, not on individual differences
Alternatives to the True Score Model
Domain sampling theory: seeks to estimate the extent to which specific sources
of variation under defined conditions are contributing to the test score
Generalizability theory by Lee J. Cronbach: universe score replaces that of
a true score
A generalizability study examines how generalizable scores from a
particular test are if the test is administered in different situations.
In the decision study, developers examine the usefulness of test scores in helping the test user
make decisions
Classical Test Theory
3 imp considerations:
X=T+e
Two sources of variance:
Constant error – affects validity not reliability
Variable error – affects both
Item response theory: a way to model the probability that a person with X ability
will be able to perform at a level of Y.
A synonym for IRT in the academic literature is latent-trait
theory.
Two important parameters: Item Difficulty and Item Discrimination
Different ‘weights’ are assigned to different items, not plain 1 or zero as in
classical theory
Reliability and Individual Scores
The Standard Error of SEM or SEM
Measurement
The standard error of measurement is the tool used to estimate or infer
the extent to which an observed score deviates from a true score
We may define the standard error of measurement as the standard deviation of a
theoretically normal distribution of test scores obtained by one person on
equivalent tests.
A test is administered to a person on e.g. 100 occasions. The person will get scores
like: 118, 120, 123, 114, 117, 125…..and so on.
Get a normal distribution of these scores. The SD of this distribution is the SEM
116 120 124 When SEM = 4
The Standard Error of the Difference
between Two Scores
True differences in score due to intervention (e.g. psychotherapy)
Comparisons between scores are made using the standard error of the difference, a
statistical measure that can aid a test user in determining how large a difference
should be before it is considered statistically significant