HISTORY AND MEANING
“Reliability tells us how much of a score is real—and how much is just noise.”
Why Important for
Psychologist Contribution
Reliability
Measurement of individual Started scientific
Francis Galton
differences measurement
Charles Made consistency
Introduced correlation
Spearman measurable
Edward Applied measurement
Promoted testing in education
Thorndike practically
Measured internal
Lee Cronbach Developed Cronbach’s alpha
consistency
1. Logical Meaning of Reliability (Simple Understanding)
Reliability means consistency of measurement.
Ask:
“If I measure the same thing again, will I get a similar result?”
Reliability refers to the consistency or stability of a test. Logically, it means that
repeated measurements under similar conditions yield similar results. In
Classical Test Theory, reliability is defined as the proportion of variance in
observed scores that is attributable to true scores. It is expressed as the ratio of
true score variance to observed score variance. Higher reliability indicates lower
measurement error and greater trustworthiness of scores.
Example
You step on a weighing machine:
o 60 kg, 60.2 kg, 59.8 kg → ✅ Reliable
o 60 kg, 65 kg, 55 kg → ❌ Not reliable
In testing:
“If a student takes the same test again, will the score be similar?”
YES → Reliable test
NO → Unreliable test
Logical Core Idea
“A reliable test gives stable and consistent results, with minimal random error.”
2. Deeper Logical Insight
Link it to CTT:
“Since every score = true ability + error,
a reliable test is one where error is small and true ability dominates.”
One-line clarity:
“Reliability tells us how much we can trust a score.”
3. Technical Meaning of Reliability (Formal Definition)
Now shift to exam language:
Reliability is the proportion of variance in observed scores that is due to true
scores.
Formula
r_{xx} = \frac{\sigma_T^2}{\sigma_X^2}
Where:
r_{xx} = reliability
\sigma_T^2 = variance of true scores
\sigma_X^2 = variance of observed scores
4. Interpret the Formula (Most Important Part)
Say:
“Out of the total differences we see in scores,
how much is real and how much is error?”
Example:
If:
Reliability = 0.80
👉 Meaning:
80% of score variation = real differences
20% = error
5. Link with Error (Very Important)
\text{Error variance} = 1 - \text{Reliability}
So:
High reliability → low error
Low reliability → high error
6. Practical Meaning
High Reliability Test:
Stable scores
Less influenced by mood, guessing
More trustworthy
Low Reliability Test:
Scores fluctuate a lot
Less dependable
CLASSICAL TEST THEORY
Core Philosophy of CTT
Classical Test Theory assumes that an observed score is composed of a true
score and random error (X = T + E). It focuses on estimating reliability, which
reflects the proportion of true variance in observed scores. Key concepts include
standard error of measurement and item analysis (difficulty and discrimination).
The theory assumes random, uncorrelated errors and stable true scores. Despite
limitations like sample dependency, it remains widely used due to its simplicity.
CTT is based on a very intuitive psychological idea:
Any score we observe is imperfect.
When a student scores 75 on a test, that is not their exact ability, but a
combination of:
their true ability
plus measurement error
2. Fundamental Equation
X=T+E
X (Observed Score) → What we actually measure
T (True Score) → The person’s actual, stable ability
E (Error) → Random influences (fatigue, mood, guessing, distractions)
Understanding “True Score”
True score is:
The average score a person would obtain over infinite repeated testing under
ideal conditions
Example:
If a student takes the same test 100 times:
Scores may vary: 72, 78, 75, 74…
The average (~75) = True score
3. Nature of Error
CTT emphasizes random error, not systematic bias.
Types of Errors:
1. Random Error
Mood
Guessing
Temporary distractions
Cancels out over time
2. Systematic Error (CTT limitation)
Poorly designed test
Biased questions
➡️CTT does not handle this well
4. Assumptions of CTT (Very Important for Exams)
1. Mean of error = 0
→ Errors cancel out in the long run
2. Error is uncorrelated with true score
→ Smart students are not systematically “more lucky/unlucky”
3. Errors across tests are uncorrelated
→ Error in Test 1 ≠ Error in Test 2
4. True score is stable
→ Trait does not change during measurement
5. Reliability: Heart of CTT
Reliability answers:
“How much of the observed score is actually true?”
r_{xx} = \frac{\sigma_T^2}{\sigma_X^2}
High reliability → less error
Low reliability → more noise
Interpretation
Reliability Meaning
0.90+ Excellent
0.70–0.80 Acceptable
<0.60 Poor
Methods of Estimating Reliability
1. Test–Retest
Same test, different time
Measures stability
2. Parallel Forms
Two equivalent tests
Measures equivalence
3. Split-Half
Divide test into two halves
Correlate scores
4. Cronbach’s Alpha
Internal consistency
Most widely used
6. Standard Error of Measurement (SEM)
SEM = SD \sqrt{1 - r_{xx}}
SEM tells us:
“How much a score might fluctuate due to error”
Practical Meaning
If:
Score = 80
SEM = 3
Then:
Likely true score range = 77 to 83
[Link] Analysis (Very Practical for Test Construction)
A. Item Difficulty (p-value)
p = \frac{\text{Number correct}}{\text{Total}}
p = 0.90 → very easy
p = 0.30 → difficult
Ideal range: 0.30–0.70
🔹 B. Item Discrimination
Ability of item to differentiate between high and low performers
High scorers should answer correctly
Low scorers should answer incorrectly
Measured using:
Point-biserial correlation
8. Parallel Forms & True Score Model
CTT also says:
If two tests are parallel:
Same mean
Same variance
Same reliability
Then:
A person’s true score remains the same across both tests
9. Limitations of CTT (Critical Thinking)
1. Sample-dependent
o Item difficulty changes with group
2. Test-dependent
o Scores depend on test used
3. Equal error assumption
o Assumes same error for all ability levels (not realistic)
4. Cannot analyze individual item behavior deeply
10. Why CTT Still Matters
Despite limitations:
Easy to understand and apply
Requires small samples
Widely used in:
o School exams
o Psychological tests
o Surveys
Significance of Reliability
Ensures consistency and stability of measurement
Increases confidence in test scores
Forms the basis for further evaluation of tests
Essential for meaningful interpretation of results
Conclusion
In conclusion, reliability refers to the consistency and dependability of test
scores. Classical Test Theory explains reliability by distinguishing between true
scores and observed scores, emphasizing that every measurement contains some
variation. A reliable test is one in which observed scores closely reflect true
ability and remain stable across repeated measurements.
TYPES OF RELIABILITY
Test-Retest Reliability
Introduction
Test–retest reliability is a method of estimating reliability that assesses the
stability of test scores over time. It is based on the principle that if a test is
reliable, it should yield consistent results when administered to the same
individuals under similar conditions at different points in time.
Meaning and Concept
Test–retest reliability refers to the degree of correlation between scores obtained
by the same individuals on two separate administrations of the same test.
Procedure
1. A test is administered to a group of individuals (Time 1).
2. After a specific time interval, the same test is administered again (Time
2).
3. The scores from both administrations are correlated.
A high correlation coefficient indicates high reliability
A low correlation indicates poor reliability
Example
Suppose a group of students takes an aptitude test:
First test → Scores: 70, 75, 80
After two weeks → Scores: 72, 74, 79
Since the scores are very similar, the test demonstrates high test–retest
reliability.
However, if the scores change drastically, it indicates low stability and poor
reliability.
Assumptions
Test–retest reliability is based on the following assumptions:
1. Stability of Trait
The characteristic being measured (e.g., intelligence) remains constant
over time.
2. No Practice Effect
Individuals do not remember or improve due to repeated exposure to the
test.
3. Similar Testing Conditions
The environment and administration conditions remain consistent.
4. No Significant External Influence
No major changes (learning, fatigue, emotional shifts) occur between the
two tests.
Advantages
1. Simple and Direct Method
Easy to understand and apply.
2. Measures Stability
Provides a clear estimate of consistency over time.
3. Useful for Stable Traits
Particularly suitable for traits like intelligence, aptitude, and personality.
4. Widely Used
Commonly used in educational and psychological testing.
Limitations
1. Practice Effect
Participants may remember answers, artificially increasing reliability.
2. Time Interval Problem
o Too short → memory effect
o Too long → real change in ability
3. Not Suitable for Changing Traits
Traits like mood or anxiety fluctuate over time.
4. Testing Conditions May Vary
Environmental differences can affect scores.
5. Dropout of Participants
Same group may not be available for retesting.
Conclusion
Test–retest reliability is an important method for assessing the temporal stability
of a test. While it provides a clear indication of consistency over time, it is
influenced by factors such as practice effects and time interval. Therefore, it is
most appropriate for measuring stable characteristics and should be used
carefully while controlling external variables.
Measurement of Test–Retest Reliability
Introduction
Reliability in psychological measurement refers to the consistency and stability
of test scores. One of the most important methods of measuring reliability is
test–retest reliability, which evaluates the stability of scores over time. This
method is particularly useful for determining whether a test produces consistent
results when administered on different occasions.
Meaning of Test–Retest Reliability
Test–retest reliability refers to the degree of correlation between scores obtained
by the same individuals on the same test administered at two different points in
time.
If the test is reliable, individuals should obtain similar scores across both
administrations.
Measurement Approach
The measurement of test–retest reliability involves the use of correlation
techniques.
Steps:
1. A test is administered to a group of individuals (Time 1).
2. After a specified time interval, the same test is administered again (Time
2).
3. Scores from both administrations are recorded.
4. The two sets of scores are correlated using a correlation coefficient
(usually Pearson’s r).
Formula for Measurement
r = \text{Correlation between scores at Time 1 and Time 2}
Where:
r = reliability coefficient
Interpretation of Coefficient
r ≈ 1.00 → Very high reliability (high stability)
r ≈ 0.80+ → Good reliability
r < 0.60 → Poor reliability
A higher correlation indicates that the test scores are stable over time.
Example
Suppose a group of students takes an aptitude test:
Student Time 1 Score Time 2 Score
A 70 72
B 65 66
C 80 79
Since the scores are very similar across both occasions, the correlation will be
high, indicating high test–retest reliability.
However, if the scores differ widely, it would indicate low stability and poor
reliability.
Assumptions of Measurement
Test–retest reliability is based on several assumptions:
1. Stability of Trait
The characteristic being measured should remain constant over time.
2. No Practice Effect
Participants should not remember answers or improve due to prior exposure.
3. Appropriate Time Interval
Too short → memory effect
Too long → real change in ability
4. Similar Testing Conditions
Testing conditions should remain consistent across both administrations.
Advantages
1. Direct Measure of Stability
It provides a clear estimate of how stable test scores are over time.
2. Simple to Compute
Only requires correlation between two sets of scores.
3. Useful for Stable Traits
Suitable for measuring intelligence, aptitude, and personality traits.
4. Widely Accepted
Commonly used in educational and psychological research.
Limitations
1. Practice Effect
Participants may perform better in the second test due to familiarity.
2. Time Interval Issues
Choosing an appropriate interval is difficult.
3. Not Suitable for Changing Traits
Traits like mood, emotions, or fatigue may vary over time.
4. Attrition Problem
Same participants may not be available for retesting.
5. Environmental Differences
Changes in testing conditions can affect results.
Factors Affecting Measurement
Length of time interval
Nature of the trait measured
Testing environment
Participant motivation
Significance
Measurement of test–retest reliability is important because:
It ensures temporal stability of test scores
It helps in evaluating the consistency of measurement tools
It increases confidence in longitudinal assessments
It is essential in standardized testing
Conclusion
In conclusion, test–retest reliability is an important method for measuring the
stability of test scores over time. By correlating scores from two
administrations, it provides a quantitative estimate of reliability. Although it is
affected by factors such as practice effects and time interval, it remains a
valuable tool for assessing consistency in psychological measurement,
especially for stable traits.
Parallel (Alternate) Forms Reliability
Introduction
Reliability is a crucial aspect of psychological and educational measurement,
referring to the consistency and stability of test scores. Among the various
methods of estimating reliability, Parallel Forms Reliability, also known as
Alternate Forms Reliability, plays an important role in ensuring that different
versions of a test produce equivalent results. This method is especially useful in
situations where repeated testing is required and the use of the same test may
lead to memory or practice effects.
Meaning and Concept
Parallel forms reliability refers to the degree of consistency between two
equivalent forms of a test administered to the same group of individuals. The
fundamental idea is that if two tests are truly parallel, they should measure the
same construct, have the same level of difficulty, and yield similar results for
the same individuals.
In simple terms, this method answers the question:
Thus, reliability is established when scores on both forms show a high degree of
correlation.
Characteristics of Parallel Forms
For two tests to be considered parallel, they must satisfy certain conditions:
Both forms should measure the same construct or ability
The level of difficulty should be equal
The content coverage should be similar
The variance and mean scores of both tests should be approximately
equal
These characteristics ensure that both forms are interchangeable.
Procedure
The procedure for determining parallel forms reliability involves the following
steps:
1. Two equivalent forms of a test (Form A and Form B) are constructed.
2. Both forms are administered to the same group of individuals. This can
be done:
o At the same time (one after another), or
o With a short time interval
3. Scores obtained on both forms are recorded.
4. The scores are then correlated using a correlation coefficient.
A high correlation coefficient indicates that both forms are reliable and
equivalent.
Example
Consider a situation where a teacher prepares two sets of a mathematics test:
Form A includes algebra and arithmetic problems
Form B includes similar algebra and arithmetic problems of equal
difficulty
A group of students takes both forms:
Student 1: 75 (Form A), 77 (Form B)
Student 2: 68 (Form A), 70 (Form B)
Student 3: 82 (Form A), 80 (Form B)
Since the scores are very similar across both forms, the correlation between
them will be high, indicating high parallel forms reliability.
However, if Form B were significantly more difficult, the scores would differ
widely, resulting in low reliability.
Assumptions of Parallel Forms Reliability
The method is based on several important assumptions:
1. Equivalence of Test Forms
Both forms must be truly parallel in terms of difficulty, content, and structure.
2. Measurement of Same Trait
Both forms should measure the same psychological construct or ability.
3. Stability of Trait
The trait being measured should remain stable during the testing period.
4. No Practice Effect
Exposure to one form should not influence performance on the other.
5. Similar Testing Conditions
The testing environment and administration procedures should remain constant.
Advantages
Parallel forms reliability offers several advantages:
1. Eliminates Memory Effect
Since different forms are used, participants cannot rely on memory of previous
answers.
2. Useful for Repeated Testing
It is highly suitable in situations where testing is conducted multiple times, such
as competitive exams.
3. Ensures Fairness
Different versions of a test can be used without compromising comparability.
4. Measures Equivalence
It directly assesses whether different forms of a test are equally reliable.
Limitations
Despite its usefulness, this method has certain limitations:
1. Difficulty in Construction
Creating two truly equivalent forms is challenging and requires expertise.
2. Time-Consuming
Developing, administering, and analyzing two tests requires more time and
effort.
3. Possibility of Content Variation
Even slight differences in content or difficulty can affect results.
4. Fatigue Effect
If both forms are administered in one session, participants may become tired,
affecting performance.
5. Costly
Developing multiple test forms increases the cost of test construction.
Comparison with Other Methods
Unlike test–retest reliability, it avoids memory effects.
Unlike split-half reliability, it uses two full tests instead of dividing one
test.
It is more rigorous but also more demanding than other methods.
Significance
Parallel forms reliability is particularly important in:
Standardized testing
Entrance examinations
Large-scale assessments
Situations requiring repeated measurements
It ensures that different test versions are equally valid and reliable, maintaining
fairness and consistency.
Conclusion
In conclusion, parallel forms reliability is a valuable method for assessing the
equivalence of different test forms. It ensures that scores remain consistent even
when the test content changes. Although it requires careful construction and
may be time-consuming, it provides a robust estimate of reliability, especially in
contexts where repeated testing is necessary. Therefore, it plays a significant
role in maintaining the integrity and fairness of psychological and educational
assessments.
Split-Half Reliability
Introduction
Reliability is an essential characteristic of any psychological or educational test,
referring to the consistency and dependability of measurement. Among the
various methods of estimating reliability, split-half reliability is one of the most
widely used methods for assessing internal consistency. It evaluates how
consistently different parts of a test measure the same construct
Meaning and Concept
Split-half reliability refers to the degree of consistency between two halves of
the same test. The underlying assumption is that if all items in a test measure the
same ability or trait, then different parts of the test should yield similar results.
Procedure
The procedure for determining split-half reliability involves the following steps:
1. A test is administered to a group of individuals.
2. The test is divided into two equivalent halves. Common methods include:
o Odd-even split (most common)
o First half vs second half
3. Scores for each half are calculated separately.
4. The scores of the two halves are correlated using a correlation coefficient.
5. Since the reliability is calculated for half the test, the Spearman-Brown
prophecy formula is applied to estimate the reliability of the full test.
Spearman-Brown Formula
The Spearman-Brown formula is used to adjust the reliability coefficient
obtained from half the test:
r_{tt} = \frac{2r}{1 + r}
Where:
r = correlation between two halves
r_{tt} = reliability of the full test
This formula corrects the underestimation of reliability due to halving the test
length.
Example
Consider a 20-item test administered to students. The test is split into:
Odd-numbered items
Even-numbered items
A student scores:
18 out of 20 on odd items
17 out of 20 on even items
If most students show similar consistency across both halves, the correlation
between the two halves will be high, indicating good split-half reliability.
However, if scores vary significantly between halves, the test may lack internal
consistency.
Assumptions of Split-Half Reliability
Split-half reliability is based on several assumptions:
1. Homogeneity of Items
All items in the test measure the same construct or ability.
2. Equal Difficulty of Halves
Both halves of the test should be comparable in terms of difficulty and content.
3. Independence of Errors
Errors in one half of the test are not related to errors in the other half.
4. Single Administration
The test is administered only once, and conditions remain constant.
Advantages
Split-half reliability offers several benefits:
1. Requires Only One Administration
Unlike test–retest reliability, there is no need to administer the test twice.
2. Time and Cost Efficient
It saves time and resources since only one test session is required.
3. Useful for Internal Consistency
It provides a quick estimate of how well the items in a test work together.
4. Eliminates Memory and Practice Effects
Since the test is taken only once, there is no risk of recall or practice influencing
results.
Limitations
Despite its usefulness, split-half reliability has certain limitations:
1. Dependence on Method of Splitting
Different ways of dividing the test can produce different reliability coefficients.
2. Not Fully Representative
The reliability estimate is based on only half the test, which may not reflect the
entire test accurately.
3. Assumption of Equal Halves
It assumes that both halves are equivalent, which may not always be true.
4. Limited Use for Small Tests
In tests with few items, splitting may not produce meaningful results.
Comparison with Other Methods
Compared to test–retest reliability, split-half avoids time-related issues.
Compared to parallel forms reliability, it does not require constructing
two tests.
However, it is less comprehensive than methods like Cronbach’s alpha,
which consider all possible item combinations.
Significance
Split-half reliability is particularly useful in:
Educational testing
Personality assessment
Questionnaire construction
It helps in identifying whether test items are internally consistent and measuring
the same construct.
Conclusion
In conclusion, split-half reliability is an important method for assessing the
internal consistency of a test. By comparing two halves of the same test, it
provides an estimate of how well the items function together. Although it has
limitations related to test splitting and assumptions of equivalence, it remains a
practical and efficient method for evaluating reliability. Its simplicity and
effectiveness make it a widely used technique in psychological and educational
measurement.
Measurement of Internal Consistency
Introduction
Internal consistency is an important method of estimating reliability in
psychological and educational measurement. It refers to the degree to which
items within a test are consistent with each other and measure the same
construct. The measurement of internal consistency is essential in determining
whether a test is homogeneous and coherent, ensuring that all items contribute
meaningfully to the overall score.
Meaning of Internal Consistency
Internal consistency reflects the inter-relatedness among test items. If a test is
designed to measure a single trait, such as intelligence, anxiety, or aptitude, then
all items should be positively correlated with one another.
In simple terms, it answers:
A test with high internal consistency ensures that individuals who perform well
on some items are likely to perform well on others.
Need for Measuring Internal Consistency
The measurement of internal consistency is important because:
It ensures homogeneity of items
It helps in identifying irrelevant or inconsistent items
It improves the quality of test construction
It provides a reliable estimate without requiring repeated testing
Methods of Measuring Internal Consistency
Internal consistency is measured using several statistical methods, the most
important being:
1. Split-Half Reliability
Measurement Approach
The test is divided into two halves (e.g., odd-even items).
Scores for both halves are calculated.
The correlation between the two halves is computed.
The result is adjusted using the Spearman-Brown prophecy formula.
Formula:
r_{tt} = \frac{2r}{1 + r}
Where:
r = correlation between halves
r_{tt} = reliability of full test
Interpretation
High correlation → high internal consistency
Low correlation → items are inconsisten
Example
In a 20-item test:
Odd items score = 16
Even items score = 15
Since both halves produce similar results, the test shows good internal
consistency.
2. Cronbach’s Alpha
Measurement Approach
Cronbach’s alpha is the most widely used method for measuring internal
consistency. It calculates the average correlation among all items in a test.
Formula:
\alpha = \frac{k}{k-1}\left(1 - \frac{\sum \sigma_i^2}{\sigma_X^2}\right)
Where:
k = number of items
\sigma_i^2 = variance of individual items
\sigma_X^2 = total test variance
Interpretation
0.90 and above → Excellent
0.80 – 0.89 → Good
0.70 – 0.79 → Acceptable
Higher alpha indicates that items are closely related and consistent.
Example
In a personality test, if all questions related to anxiety show similar response
patterns, Cronbach’s alpha will be high, indicating strong internal consistency.
3. Kuder–Richardson Methods (KR-20 and KR-21)
Measurement Approach
These methods are used for tests with dichotomous items (right/wrong).
KR-20 Formula:
KR_{20} = \frac{k}{k-1}\left(1 - \frac{\sum pq}{\sigma_X^2}\right)
Where:
p = proportion of correct responses
q=1-p
KR-21 Formula:
Simplified version using mean and variance
Less accurate
Interpretation
High KR value → high internal consistency
Low KR value → poor item consistency
Example
In a multiple-choice test, if students who answer one question correctly also
answer others correctly, the KR value will be high.
Assumptions of Internal Consistency Measurement
1. Unidimensionality
All items measure a single construct.
2. Homogeneity of Items
Items are similar in content and purpose.
3. Independence of Errors
Errors associated with items are not correlated.
4. Consistency Across Items
Each item contributes equally to the construct.
Advantages
1. Single Administration
No need for repeated testing.
2. Efficient and Economical
Saves time and resources.
3. Comprehensive Estimate
Especially with Cronbach’s alpha.
4. Widely Applicable
Useful in questionnaires, scales, and tests.
Limitations
1. Not Suitable for Multidimensional Tests
If a test measures multiple traits, internal consistency may be misleading.
2. High Value Does Not Ensure Validity
Items may be consistent but not measure the intended construct.
3. Affected by Test Length
Longer tests tend to show higher reliability.
4. Assumption of Equal Contribution
Not all items contribute equally in reality.
Significance
Measurement of internal consistency is crucial in:
Psychological test construction
Educational assessments
Personality and attitude scales
It ensures that a test is coherent, reliable, and scientifically sound.
Conclusion
In conclusion, the measurement of internal consistency plays a vital role in
determining the reliability of a test. Methods such as split-half reliability,
Cronbach’s alpha, and Kuder–Richardson formulas provide statistical estimates
of how well test items function together. Among these, Cronbach’s alpha is the
most widely used due to its comprehensive nature. Despite certain limitations,
internal consistency remains an essential tool for ensuring that tests are reliable
and meaningful.
Additional for exam purpose
Kuder–Richardson Formulas
The Kuder–Richardson (KR) formulas are statistical techniques used to measure
the internal consistency reliability of a test. These formulas are particularly
applicable to tests consisting of dichotomous items, where responses are scored
in two categories such as right/wrong, yes/no, or true/false. They help determine
how consistently the items of a test measure a single underlying construct.
Concept of Internal Consistency
Internal consistency refers to the degree of homogeneity among test items, that
is, the extent to which all items in a test measure the same ability or trait. If the
items are consistent, individuals who perform well on one item are likely to
perform well on other items also. The Kuder–Richardson formulas provide a
quantitative estimate of this consistency.
Types of Kuder–Richardson Formulas
There are two main types:
1. KR-20 Formula
KR-20 is the most widely used and accurate measure of internal consistency for
dichotomous items.
KR_{20} = \frac{k}{k-1}\left(1 - \frac{\sum pq}{\sigma_X^2}\right)
Where:
k = total number of items
p = proportion of correct responses to an item
q = 1 - p = proportion of incorrect responses
\sum pq = sum of the product of p and q for all items
\sigma_X^2 = variance of total test scores
Explanation:
The term pq represents the variance of each item, and the formula compares the
sum of item variances with the total test variance. If items are highly consistent,
the total variance will mainly reflect true score variance, resulting in a higher
reliability coefficient.
Characteristics:
Takes into account item difficulty
Provides accurate estimation
Preferred in most testing situations
2. KR-21 Formula
KR-21 is a simplified version of KR-20 used when detailed item-level data is
unavailable.
KR_{21} = \frac{k}{k-1}\left(1 - \frac{M(k - M)}{k\sigma_X^2}\right)
Where:
k = number of items
M = mean score
\sigma_X^2 = variance of total scores
Explanation:
KR-21 assumes that all items have equal difficulty, which simplifies
calculations but reduces accuracy.
Characteristics:
Easier to compute
Requires less data
Less precise than KR-20
Interpretation of KR Coefficient
The value of KR reliability ranges from 0 to 1:
0.90 and above → Excellent reliability
0.80 – 0.89 → Good reliability
0.70 – 0.79 → Acceptable
Below 0.60 → Poor reliability
Higher values indicate that the test items are highly consistent and measure the
same construct effectively.
Relation to Cronbach’s Alpha
The Kuder–Richardson formula, especially KR-20, is considered a special case
of Cronbach’s Alpha, used specifically when items are dichotomous. While
Cronbach’s alpha can be applied to items with multiple scoring categories, KR
formulas are limited to binary responses.
Significance and Uses
Widely used in objective-type tests such as multiple-choice exams
Helps in test construction and evaluation
Assists in identifying inconsistent or weak items
Ensures that the test provides reliable and stable scores
Limitations
Applicable only to dichotomous items
KR-21 may produce biased estimates due to its assumptions
High reliability does not guarantee validity
Does not account for multidimensional constructs
Conclusion
In conclusion, the Kuder–Richardson formulas are essential tools in
psychometrics for assessing the internal consistency of tests with dichotomous
items. KR-20 is preferred for its accuracy, while KR-21 serves as a simpler
alternative when limited data is available. These formulas play a vital role in
ensuring that a test is reliable, consistent, and suitable for measuring a specific
construct.