Module 4: Reliability
NOTES PROPER ADDITIONAL INFO
RELIABILITY COEFFICIENT OF
EQUIVALENCE
● consistency of scores produced by a measurement instrument. - Alternate
● E is random, and therefore uncorrelated with either T or X. Forms
● Crocker & Algina (1986) - Measure of
○ reliability as a measure of instrument consistency in the scores that it yields. how these
● A true score on a psychological variable can be roughly thought of as the person’s tests are
average score on that assessment if that assessment was taken an infinite number of similar
times. - Correlate two
○ True score = the person’s average score if they took the test an infinite number tests
of times.
● the proportion of observed score variance accounted for by the true score. TEMPORAL
○ High reliability = scores mostly reflect the true construct. STABILITY
○ Low reliability = scores influenced mostly by error. - Aka
● reliability is concerned with how scores vary across replications of a measurement given coefficient of
to the same individual over time (Haertel, 2006). stability
● Confidence Intervals - Test-Retest
○ Always report confidence intervals for reliability.
INTERNAL
CONSISTENCY
- Cronbach’s
Alpha
ESTIMATION OF RELIABILITY
1. Alternate Forms
● should be equivalent in terms of content and difficulty of the items.
● Two different but equivalent test forms that measure the same construct.
● Assuming that this is the case, both forms are given to a single sample of individuals either
simultaneously, or within a very narrow window of time, and the correlation between scores on the two
forms serves as the reliability estimate.
● Counterbalancing
○ half of the sample receives Form A followed by Form B, and half of the sample receives Form B
followed by Form A.
● Forms must be parallel assessment
○ Same true score and the same error variance
2. Test Retest
● involves giving a single sample of persons the same assessment at two points in time.
● scores obtained at each of these time points can then be correlated with one another
● Test and retest can balance each other’s limitations
● Use one test and use it to retest in a different occasion
● An acceptable time lag for use in test-retest will vary depending upon the nature of the trait being
assessed, and the extent to which it is temporally stable
● Assumptions
○ Correct time interval
■ Not too short (avoid memory effects)
■ Not too long (trait might change)
○ Equal error variance at both times
■ Testing conditions should be identical.
○ Trait stability must be reasonable.
● Low test-retest correlation may reflect:
○ Trait instability, Memory effect, testing conditions (not just poor reliability.)
3. Alternate Forms & Test-Retest Reliability Estimates
● can be used in conjunction with one another so that researchers obtain a single measure of both stability
over time and equivalence across test forms.
● serves as an estimate of both temporal stability, and equivalence of the forms.
4. Split-Half Reliability & The Spearman-Brown Formula
● Splitting a test into 2
● the halves should essentially be parallel forms
● SPLIT HALF
○ it focuses on larger units of analysis, two halves of the same test or form, rather than on
individual items.
○ The entire set of items is then divided into two equally sized halves, with the intent that the halves
are equivalent in terms of item difficulty, item content, and location in the instrument itself
○ For this reason, we will need to adjust the correlation between the halves in order to reflect the
reliability of the full instrument with the current sample.
○ Reliability estimates are directly impacted by the length of the scale such as shorter scales yield
lower reliability
○ Half of one test is correlated (hinanati) and the other half of another test
○ CON: if binabawasan mo ung item, bumababa ang reliability
● SPREARMAN-BROWN FORMULA
○ Corrects split half
○ Used to adjust the split-half correlation to estimate reliability of the full test.
○ obtain an estimate of reliability for the scores
● Rulon (1939) suggested a method for estimating split-halves reliability that does not require the use of
Spearman-Brown prophecy formula
1. finding this reliability estimate is to calculate the difference between scores on the two halves for
each individual in the sample.
a. Perhaps the most widely used method for creation of the halves is to divide the
instrument into odd and even components.
2. A second approach to creating the two halves requires the researcher to first order the items
based on their difficulty values, and then assign them alternately into one of two halves, working
down the list of item difficulty values.
3. A third method for creating the two halves is to randomly assign items to one half or the other
without regard to their position in the test, their item difficulty values, or any other item
characteristic.
5. Reliability Estimate Based on Item Covariances
● Covariance
○ is the numerator of the correlation coefficient and is itself a (unstandardized) measure of the
relationship between two variables.
○ A primary advantage of reliability estimation methods relying on item covariances is that they do
not require multiple administrations of the assessment, nor do they require that we use an
arbitrary method for dividing the assessment into two halves. In addition, several of these
item-covariance based approaches have methods for calculating confidence intervals.
6. Cronbach’s Alpha & Similar Reliability Estimates
● Most common method for estimating reliability
● is probably the most widely used and reported estimate of reliability throughout the social science
● KR-20 is just a special case of Cronbach’s α. Therefore, when researchers apply standard statistical
software packages to dichotomous item data in order to obtain α, they are in fact calculating the KR-20
value.
● Guttman (1945) described six such coefficients, which he referred to as Lambda 1 through Lambda 6.
7. Omega
● An alternative manner in which reliability estimation might be constructed is through the prism of factor
analysis.
● McDonald (1999) described how the true score model or idea can be reconceptualized through the factor
analysis model.
● With these ideas in mind, we know that the factor analysis model can provide estimates of error
variance for the items through the unique variance associated with each.
● McDonald (1999) showed that in the context of congeneric tests, the square of the sum of the factor
loadings is an estimate of the true score variance
● Procedure
○ Fit a factor model
○ Extract:
■ Factor loadings
■ Error variances
○ Compute ω:
■ Numerator: (sum of loadings)²
■ Denominator: (same numerator + sum of error variances)
● Requirements
○ Good model fit
○ Model must be identified (fix factor variance to 1)
● Advantages
○ More accurate than alpha
○ Useful for congeneric measures
○ Can apply to bifactor models
8. Stratified Alpha
● Used for tests with subscales
● Used when estimating reliability for a scale consisting of several subscales
● Total scale composed of several distinct but related domains.
● perhaps the optimal method for estimating reliability of a composite scale.
● Works Best When
○ Subscale items have similar factor loadings
○ Subscales are conceptually strong
QUIZ ANSWERS 😛
1. The correction for this reliability approach is known as the Spearman-Brown prophecy formula- Split Half
2. The most commonly used methods for estimating reliability are based upon relationships among individuals items on the
assessment - Cronbach's a
3. This approach to estimate reliability is particularly useful when it is important for an instrument to exhibit temporal stability -
Test-Retest
4. The correlation between the halves is thus an estimate of the reliability of these shortened parallel forms of the measure -
Split-Half
5. It is concerned with how scores vary across replications of a measurement given to the same individual overtime -Reliability
6. The value which is commonly referred to as reliability can be interpreted as the proportion of observed score variance
accounted for by the true score. - Reliability
7. The larger value of pxx suggests that the observed score on an assessment is by and large a function of the level of the ____
that is being measured on that given instrument. - latent construct
8. The correlation between the forms can then be treated as both an estimate of reliability and a measure of equivalence
between the test scores. - alternate forms
9. They describe reliability as a measure of instrument consistency in the score that it yields. - Crocker & Algina
10. A primary advantage of reliability estimation methods relying on item ____ is that they do not require multiple administrations
of the assessment, nor do they require that we use arbitrary methods for dividing the assessment into two halves. -
Covariances