0% found this document useful (0 votes)
2 views33 pages

Reliability

Uploaded by

Medhavi Gugnani
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views33 pages

Reliability

Uploaded by

Medhavi Gugnani
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

HISTORY AND MEANING

“Reliability tells us how much of a score is real—and how much is just noise.”

Why Important for


Psychologist Contribution
Reliability
Measurement of individual Started scientific
Francis Galton
differences measurement
Charles Made consistency
Introduced correlation
Spearman measurable
Edward Applied measurement
Promoted testing in education
Thorndike practically
Measured internal
Lee Cronbach Developed Cronbach’s alpha
consistency

1. Logical Meaning of Reliability (Simple Understanding)

Reliability means consistency of measurement.

Ask:

“If I measure the same thing again, will I get a similar result?”

Reliability refers to the consistency or stability of a test. Logically, it means that


repeated measurements under similar conditions yield similar results. In
Classical Test Theory, reliability is defined as the proportion of variance in
observed scores that is attributable to true scores. It is expressed as the ratio of
true score variance to observed score variance. Higher reliability indicates lower
measurement error and greater trustworthiness of scores.

Example

 You step on a weighing machine:


o 60 kg, 60.2 kg, 59.8 kg → ✅ Reliable
o 60 kg, 65 kg, 55 kg → ❌ Not reliable

In testing:

“If a student takes the same test again, will the score be similar?”
 YES → Reliable test
 NO → Unreliable test

Logical Core Idea

“A reliable test gives stable and consistent results, with minimal random error.”

2. Deeper Logical Insight

Link it to CTT:

“Since every score = true ability + error,


a reliable test is one where error is small and true ability dominates.”

One-line clarity:

“Reliability tells us how much we can trust a score.”

3. Technical Meaning of Reliability (Formal Definition)

Now shift to exam language:

Reliability is the proportion of variance in observed scores that is due to true


scores.

Formula

r_{xx} = \frac{\sigma_T^2}{\sigma_X^2}

Where:

 r_{xx} = reliability
 \sigma_T^2 = variance of true scores
 \sigma_X^2 = variance of observed scores

4. Interpret the Formula (Most Important Part)

Say:

“Out of the total differences we see in scores,


how much is real and how much is error?”
Example:

If:

 Reliability = 0.80

👉 Meaning:

 80% of score variation = real differences


 20% = error

5. Link with Error (Very Important)

\text{Error variance} = 1 - \text{Reliability}

So:

 High reliability → low error


 Low reliability → high error

6. Practical Meaning

High Reliability Test:

 Stable scores
 Less influenced by mood, guessing
 More trustworthy

Low Reliability Test:

 Scores fluctuate a lot


 Less dependable

CLASSICAL TEST THEORY

Core Philosophy of CTT

Classical Test Theory assumes that an observed score is composed of a true


score and random error (X = T + E). It focuses on estimating reliability, which
reflects the proportion of true variance in observed scores. Key concepts include
standard error of measurement and item analysis (difficulty and discrimination).
The theory assumes random, uncorrelated errors and stable true scores. Despite
limitations like sample dependency, it remains widely used due to its simplicity.

CTT is based on a very intuitive psychological idea:

Any score we observe is imperfect.

When a student scores 75 on a test, that is not their exact ability, but a
combination of:

 their true ability


 plus measurement error

2. Fundamental Equation

X=T+E

 X (Observed Score) → What we actually measure


 T (True Score) → The person’s actual, stable ability
 E (Error) → Random influences (fatigue, mood, guessing, distractions)

Understanding “True Score”

True score is:

The average score a person would obtain over infinite repeated testing under
ideal conditions

Example:

If a student takes the same test 100 times:

 Scores may vary: 72, 78, 75, 74…


 The average (~75) = True score

3. Nature of Error

CTT emphasizes random error, not systematic bias.

Types of Errors:
1. Random Error

 Mood
 Guessing
 Temporary distractions

Cancels out over time

2. Systematic Error (CTT limitation)

 Poorly designed test


 Biased questions

➡️CTT does not handle this well

4. Assumptions of CTT (Very Important for Exams)

1. Mean of error = 0

→ Errors cancel out in the long run

2. Error is uncorrelated with true score

→ Smart students are not systematically “more lucky/unlucky”

3. Errors across tests are uncorrelated

→ Error in Test 1 ≠ Error in Test 2

4. True score is stable

→ Trait does not change during measurement

5. Reliability: Heart of CTT

Reliability answers:

“How much of the observed score is actually true?”

r_{xx} = \frac{\sigma_T^2}{\sigma_X^2}

 High reliability → less error


 Low reliability → more noise
Interpretation

Reliability Meaning
0.90+ Excellent
0.70–0.80 Acceptable
<0.60 Poor

Methods of Estimating Reliability

1. Test–Retest

 Same test, different time


 Measures stability

2. Parallel Forms

 Two equivalent tests


 Measures equivalence

3. Split-Half

 Divide test into two halves


 Correlate scores

4. Cronbach’s Alpha

 Internal consistency
 Most widely used

6. Standard Error of Measurement (SEM)

SEM = SD \sqrt{1 - r_{xx}}

SEM tells us:

“How much a score might fluctuate due to error”

Practical Meaning
If:

 Score = 80
 SEM = 3

Then:

 Likely true score range = 77 to 83

[Link] Analysis (Very Practical for Test Construction)

A. Item Difficulty (p-value)

p = \frac{\text{Number correct}}{\text{Total}}

 p = 0.90 → very easy


 p = 0.30 → difficult

Ideal range: 0.30–0.70

🔹 B. Item Discrimination

Ability of item to differentiate between high and low performers

 High scorers should answer correctly


 Low scorers should answer incorrectly

Measured using:

 Point-biserial correlation

8. Parallel Forms & True Score Model

CTT also says:

If two tests are parallel:

 Same mean
 Same variance
 Same reliability

Then:

A person’s true score remains the same across both tests


9. Limitations of CTT (Critical Thinking)

1. Sample-dependent
o Item difficulty changes with group
2. Test-dependent
o Scores depend on test used
3. Equal error assumption
o Assumes same error for all ability levels (not realistic)
4. Cannot analyze individual item behavior deeply

10. Why CTT Still Matters

Despite limitations:

 Easy to understand and apply


 Requires small samples
 Widely used in:
o School exams
o Psychological tests
o Surveys

Significance of Reliability

 Ensures consistency and stability of measurement


 Increases confidence in test scores
 Forms the basis for further evaluation of tests
 Essential for meaningful interpretation of results

Conclusion

In conclusion, reliability refers to the consistency and dependability of test


scores. Classical Test Theory explains reliability by distinguishing between true
scores and observed scores, emphasizing that every measurement contains some
variation. A reliable test is one in which observed scores closely reflect true
ability and remain stable across repeated measurements.

TYPES OF RELIABILITY

Test-Retest Reliability
Introduction

Test–retest reliability is a method of estimating reliability that assesses the


stability of test scores over time. It is based on the principle that if a test is
reliable, it should yield consistent results when administered to the same
individuals under similar conditions at different points in time.

Meaning and Concept

Test–retest reliability refers to the degree of correlation between scores obtained


by the same individuals on two separate administrations of the same test.

Procedure

1. A test is administered to a group of individuals (Time 1).


2. After a specific time interval, the same test is administered again (Time
2).
3. The scores from both administrations are correlated.

 A high correlation coefficient indicates high reliability


 A low correlation indicates poor reliability

Example

Suppose a group of students takes an aptitude test:

 First test → Scores: 70, 75, 80


 After two weeks → Scores: 72, 74, 79

Since the scores are very similar, the test demonstrates high test–retest
reliability.

However, if the scores change drastically, it indicates low stability and poor
reliability.

Assumptions

Test–retest reliability is based on the following assumptions:

1. Stability of Trait
The characteristic being measured (e.g., intelligence) remains constant
over time.

2. No Practice Effect

Individuals do not remember or improve due to repeated exposure to the


test.

3. Similar Testing Conditions

The environment and administration conditions remain consistent.

4. No Significant External Influence

No major changes (learning, fatigue, emotional shifts) occur between the


two tests.

Advantages

1. Simple and Direct Method

Easy to understand and apply.

2. Measures Stability

Provides a clear estimate of consistency over time.

3. Useful for Stable Traits

Particularly suitable for traits like intelligence, aptitude, and personality.

4. Widely Used

Commonly used in educational and psychological testing.

Limitations

1. Practice Effect

Participants may remember answers, artificially increasing reliability.

2. Time Interval Problem


o Too short → memory effect
o Too long → real change in ability
3. Not Suitable for Changing Traits
Traits like mood or anxiety fluctuate over time.

4. Testing Conditions May Vary

Environmental differences can affect scores.

5. Dropout of Participants

Same group may not be available for retesting.

Conclusion

Test–retest reliability is an important method for assessing the temporal stability


of a test. While it provides a clear indication of consistency over time, it is
influenced by factors such as practice effects and time interval. Therefore, it is
most appropriate for measuring stable characteristics and should be used
carefully while controlling external variables.

Measurement of Test–Retest Reliability

Introduction

Reliability in psychological measurement refers to the consistency and stability


of test scores. One of the most important methods of measuring reliability is
test–retest reliability, which evaluates the stability of scores over time. This
method is particularly useful for determining whether a test produces consistent
results when administered on different occasions.

Meaning of Test–Retest Reliability

Test–retest reliability refers to the degree of correlation between scores obtained


by the same individuals on the same test administered at two different points in
time.

If the test is reliable, individuals should obtain similar scores across both
administrations.

Measurement Approach
The measurement of test–retest reliability involves the use of correlation
techniques.

Steps:

1. A test is administered to a group of individuals (Time 1).


2. After a specified time interval, the same test is administered again (Time
2).
3. Scores from both administrations are recorded.
4. The two sets of scores are correlated using a correlation coefficient
(usually Pearson’s r).

Formula for Measurement

r = \text{Correlation between scores at Time 1 and Time 2}

Where:

 r = reliability coefficient

Interpretation of Coefficient

 r ≈ 1.00 → Very high reliability (high stability)


 r ≈ 0.80+ → Good reliability
 r < 0.60 → Poor reliability

A higher correlation indicates that the test scores are stable over time.

Example

Suppose a group of students takes an aptitude test:


Student Time 1 Score Time 2 Score
A 70 72
B 65 66
C 80 79

Since the scores are very similar across both occasions, the correlation will be
high, indicating high test–retest reliability.

However, if the scores differ widely, it would indicate low stability and poor
reliability.

Assumptions of Measurement

Test–retest reliability is based on several assumptions:

1. Stability of Trait

The characteristic being measured should remain constant over time.

2. No Practice Effect

Participants should not remember answers or improve due to prior exposure.

3. Appropriate Time Interval

 Too short → memory effect


 Too long → real change in ability
4. Similar Testing Conditions

Testing conditions should remain consistent across both administrations.

Advantages

1. Direct Measure of Stability

It provides a clear estimate of how stable test scores are over time.

2. Simple to Compute

Only requires correlation between two sets of scores.

3. Useful for Stable Traits

Suitable for measuring intelligence, aptitude, and personality traits.

4. Widely Accepted

Commonly used in educational and psychological research.

Limitations
1. Practice Effect

Participants may perform better in the second test due to familiarity.

2. Time Interval Issues

Choosing an appropriate interval is difficult.

3. Not Suitable for Changing Traits

Traits like mood, emotions, or fatigue may vary over time.

4. Attrition Problem

Same participants may not be available for retesting.

5. Environmental Differences

Changes in testing conditions can affect results.

Factors Affecting Measurement

 Length of time interval


 Nature of the trait measured
 Testing environment
 Participant motivation

Significance

Measurement of test–retest reliability is important because:

 It ensures temporal stability of test scores


 It helps in evaluating the consistency of measurement tools
 It increases confidence in longitudinal assessments
 It is essential in standardized testing

Conclusion

In conclusion, test–retest reliability is an important method for measuring the


stability of test scores over time. By correlating scores from two
administrations, it provides a quantitative estimate of reliability. Although it is
affected by factors such as practice effects and time interval, it remains a
valuable tool for assessing consistency in psychological measurement,
especially for stable traits.
Parallel (Alternate) Forms Reliability

Introduction

Reliability is a crucial aspect of psychological and educational measurement,


referring to the consistency and stability of test scores. Among the various
methods of estimating reliability, Parallel Forms Reliability, also known as
Alternate Forms Reliability, plays an important role in ensuring that different
versions of a test produce equivalent results. This method is especially useful in
situations where repeated testing is required and the use of the same test may
lead to memory or practice effects.

Meaning and Concept

Parallel forms reliability refers to the degree of consistency between two


equivalent forms of a test administered to the same group of individuals. The
fundamental idea is that if two tests are truly parallel, they should measure the
same construct, have the same level of difficulty, and yield similar results for
the same individuals.

In simple terms, this method answers the question:

Thus, reliability is established when scores on both forms show a high degree of
correlation.

Characteristics of Parallel Forms

For two tests to be considered parallel, they must satisfy certain conditions:

 Both forms should measure the same construct or ability


 The level of difficulty should be equal
 The content coverage should be similar
 The variance and mean scores of both tests should be approximately
equal

These characteristics ensure that both forms are interchangeable.

Procedure

The procedure for determining parallel forms reliability involves the following
steps:
1. Two equivalent forms of a test (Form A and Form B) are constructed.
2. Both forms are administered to the same group of individuals. This can
be done:
o At the same time (one after another), or
o With a short time interval
3. Scores obtained on both forms are recorded.
4. The scores are then correlated using a correlation coefficient.

A high correlation coefficient indicates that both forms are reliable and
equivalent.

Example

Consider a situation where a teacher prepares two sets of a mathematics test:

 Form A includes algebra and arithmetic problems


 Form B includes similar algebra and arithmetic problems of equal
difficulty

A group of students takes both forms:

 Student 1: 75 (Form A), 77 (Form B)


 Student 2: 68 (Form A), 70 (Form B)
 Student 3: 82 (Form A), 80 (Form B)

Since the scores are very similar across both forms, the correlation between
them will be high, indicating high parallel forms reliability.

However, if Form B were significantly more difficult, the scores would differ
widely, resulting in low reliability.

Assumptions of Parallel Forms Reliability

The method is based on several important assumptions:

1. Equivalence of Test Forms

Both forms must be truly parallel in terms of difficulty, content, and structure.

2. Measurement of Same Trait

Both forms should measure the same psychological construct or ability.

3. Stability of Trait
The trait being measured should remain stable during the testing period.

4. No Practice Effect

Exposure to one form should not influence performance on the other.

5. Similar Testing Conditions

The testing environment and administration procedures should remain constant.

Advantages

Parallel forms reliability offers several advantages:

1. Eliminates Memory Effect

Since different forms are used, participants cannot rely on memory of previous
answers.

2. Useful for Repeated Testing

It is highly suitable in situations where testing is conducted multiple times, such


as competitive exams.

3. Ensures Fairness

Different versions of a test can be used without compromising comparability.

4. Measures Equivalence

It directly assesses whether different forms of a test are equally reliable.

Limitations

Despite its usefulness, this method has certain limitations:

1. Difficulty in Construction

Creating two truly equivalent forms is challenging and requires expertise.

2. Time-Consuming

Developing, administering, and analyzing two tests requires more time and
effort.
3. Possibility of Content Variation

Even slight differences in content or difficulty can affect results.

4. Fatigue Effect

If both forms are administered in one session, participants may become tired,
affecting performance.

5. Costly

Developing multiple test forms increases the cost of test construction.

Comparison with Other Methods

 Unlike test–retest reliability, it avoids memory effects.


 Unlike split-half reliability, it uses two full tests instead of dividing one
test.
 It is more rigorous but also more demanding than other methods.

Significance

Parallel forms reliability is particularly important in:

 Standardized testing
 Entrance examinations
 Large-scale assessments
 Situations requiring repeated measurements

It ensures that different test versions are equally valid and reliable, maintaining
fairness and consistency.

Conclusion

In conclusion, parallel forms reliability is a valuable method for assessing the


equivalence of different test forms. It ensures that scores remain consistent even
when the test content changes. Although it requires careful construction and
may be time-consuming, it provides a robust estimate of reliability, especially in
contexts where repeated testing is necessary. Therefore, it plays a significant
role in maintaining the integrity and fairness of psychological and educational
assessments.
Split-Half Reliability

Introduction

Reliability is an essential characteristic of any psychological or educational test,


referring to the consistency and dependability of measurement. Among the
various methods of estimating reliability, split-half reliability is one of the most
widely used methods for assessing internal consistency. It evaluates how
consistently different parts of a test measure the same construct

Meaning and Concept

Split-half reliability refers to the degree of consistency between two halves of


the same test. The underlying assumption is that if all items in a test measure the
same ability or trait, then different parts of the test should yield similar results.

Procedure

The procedure for determining split-half reliability involves the following steps:

1. A test is administered to a group of individuals.


2. The test is divided into two equivalent halves. Common methods include:
o Odd-even split (most common)
o First half vs second half
3. Scores for each half are calculated separately.
4. The scores of the two halves are correlated using a correlation coefficient.
5. Since the reliability is calculated for half the test, the Spearman-Brown
prophecy formula is applied to estimate the reliability of the full test.

Spearman-Brown Formula

The Spearman-Brown formula is used to adjust the reliability coefficient


obtained from half the test:

r_{tt} = \frac{2r}{1 + r}

Where:

 r = correlation between two halves


 r_{tt} = reliability of the full test
This formula corrects the underestimation of reliability due to halving the test
length.

Example

Consider a 20-item test administered to students. The test is split into:

 Odd-numbered items
 Even-numbered items

A student scores:

 18 out of 20 on odd items


 17 out of 20 on even items

If most students show similar consistency across both halves, the correlation
between the two halves will be high, indicating good split-half reliability.

However, if scores vary significantly between halves, the test may lack internal
consistency.

Assumptions of Split-Half Reliability

Split-half reliability is based on several assumptions:

1. Homogeneity of Items

All items in the test measure the same construct or ability.

2. Equal Difficulty of Halves

Both halves of the test should be comparable in terms of difficulty and content.

3. Independence of Errors

Errors in one half of the test are not related to errors in the other half.

4. Single Administration

The test is administered only once, and conditions remain constant.

Advantages
Split-half reliability offers several benefits:

1. Requires Only One Administration

Unlike test–retest reliability, there is no need to administer the test twice.

2. Time and Cost Efficient

It saves time and resources since only one test session is required.

3. Useful for Internal Consistency

It provides a quick estimate of how well the items in a test work together.

4. Eliminates Memory and Practice Effects

Since the test is taken only once, there is no risk of recall or practice influencing
results.

Limitations

Despite its usefulness, split-half reliability has certain limitations:

1. Dependence on Method of Splitting

Different ways of dividing the test can produce different reliability coefficients.

2. Not Fully Representative

The reliability estimate is based on only half the test, which may not reflect the
entire test accurately.

3. Assumption of Equal Halves

It assumes that both halves are equivalent, which may not always be true.

4. Limited Use for Small Tests

In tests with few items, splitting may not produce meaningful results.

Comparison with Other Methods

 Compared to test–retest reliability, split-half avoids time-related issues.


 Compared to parallel forms reliability, it does not require constructing
two tests.
 However, it is less comprehensive than methods like Cronbach’s alpha,
which consider all possible item combinations.

Significance

Split-half reliability is particularly useful in:

 Educational testing
 Personality assessment
 Questionnaire construction

It helps in identifying whether test items are internally consistent and measuring
the same construct.

Conclusion

In conclusion, split-half reliability is an important method for assessing the


internal consistency of a test. By comparing two halves of the same test, it
provides an estimate of how well the items function together. Although it has
limitations related to test splitting and assumptions of equivalence, it remains a
practical and efficient method for evaluating reliability. Its simplicity and
effectiveness make it a widely used technique in psychological and educational
measurement.

Measurement of Internal Consistency

Introduction

Internal consistency is an important method of estimating reliability in


psychological and educational measurement. It refers to the degree to which
items within a test are consistent with each other and measure the same
construct. The measurement of internal consistency is essential in determining
whether a test is homogeneous and coherent, ensuring that all items contribute
meaningfully to the overall score.

Meaning of Internal Consistency

Internal consistency reflects the inter-relatedness among test items. If a test is


designed to measure a single trait, such as intelligence, anxiety, or aptitude, then
all items should be positively correlated with one another.

In simple terms, it answers:


A test with high internal consistency ensures that individuals who perform well
on some items are likely to perform well on others.

Need for Measuring Internal Consistency

The measurement of internal consistency is important because:

 It ensures homogeneity of items


 It helps in identifying irrelevant or inconsistent items
 It improves the quality of test construction
 It provides a reliable estimate without requiring repeated testing

Methods of Measuring Internal Consistency

Internal consistency is measured using several statistical methods, the most


important being:

1. Split-Half Reliability

Measurement Approach

 The test is divided into two halves (e.g., odd-even items).


 Scores for both halves are calculated.
 The correlation between the two halves is computed.
 The result is adjusted using the Spearman-Brown prophecy formula.

Formula:

r_{tt} = \frac{2r}{1 + r}

Where:

 r = correlation between halves


 r_{tt} = reliability of full test

Interpretation

 High correlation → high internal consistency


 Low correlation → items are inconsisten
Example

In a 20-item test:

 Odd items score = 16


 Even items score = 15

Since both halves produce similar results, the test shows good internal
consistency.

2. Cronbach’s Alpha

Measurement Approach

Cronbach’s alpha is the most widely used method for measuring internal
consistency. It calculates the average correlation among all items in a test.

Formula:

\alpha = \frac{k}{k-1}\left(1 - \frac{\sum \sigma_i^2}{\sigma_X^2}\right)

Where:

 k = number of items
 \sigma_i^2 = variance of individual items
 \sigma_X^2 = total test variance

Interpretation

 0.90 and above → Excellent


 0.80 – 0.89 → Good
 0.70 – 0.79 → Acceptable

Higher alpha indicates that items are closely related and consistent.

Example

In a personality test, if all questions related to anxiety show similar response


patterns, Cronbach’s alpha will be high, indicating strong internal consistency.

3. Kuder–Richardson Methods (KR-20 and KR-21)


Measurement Approach

These methods are used for tests with dichotomous items (right/wrong).

KR-20 Formula:

KR_{20} = \frac{k}{k-1}\left(1 - \frac{\sum pq}{\sigma_X^2}\right)

Where:

 p = proportion of correct responses


 q=1-p

KR-21 Formula:

 Simplified version using mean and variance


 Less accurate

Interpretation

 High KR value → high internal consistency


 Low KR value → poor item consistency

Example

In a multiple-choice test, if students who answer one question correctly also


answer others correctly, the KR value will be high.
Assumptions of Internal Consistency Measurement

1. Unidimensionality

All items measure a single construct.

2. Homogeneity of Items

Items are similar in content and purpose.

3. Independence of Errors

Errors associated with items are not correlated.

4. Consistency Across Items

Each item contributes equally to the construct.

Advantages

1. Single Administration

No need for repeated testing.

2. Efficient and Economical

Saves time and resources.

3. Comprehensive Estimate

Especially with Cronbach’s alpha.

4. Widely Applicable

Useful in questionnaires, scales, and tests.

Limitations

1. Not Suitable for Multidimensional Tests


If a test measures multiple traits, internal consistency may be misleading.

2. High Value Does Not Ensure Validity

Items may be consistent but not measure the intended construct.

3. Affected by Test Length

Longer tests tend to show higher reliability.

4. Assumption of Equal Contribution

Not all items contribute equally in reality.

Significance

Measurement of internal consistency is crucial in:

 Psychological test construction


 Educational assessments
 Personality and attitude scales

It ensures that a test is coherent, reliable, and scientifically sound.

Conclusion

In conclusion, the measurement of internal consistency plays a vital role in


determining the reliability of a test. Methods such as split-half reliability,
Cronbach’s alpha, and Kuder–Richardson formulas provide statistical estimates
of how well test items function together. Among these, Cronbach’s alpha is the
most widely used due to its comprehensive nature. Despite certain limitations,
internal consistency remains an essential tool for ensuring that tests are reliable
and meaningful.
Additional for exam purpose

Kuder–Richardson Formulas

The Kuder–Richardson (KR) formulas are statistical techniques used to measure


the internal consistency reliability of a test. These formulas are particularly
applicable to tests consisting of dichotomous items, where responses are scored
in two categories such as right/wrong, yes/no, or true/false. They help determine
how consistently the items of a test measure a single underlying construct.

Concept of Internal Consistency

Internal consistency refers to the degree of homogeneity among test items, that
is, the extent to which all items in a test measure the same ability or trait. If the
items are consistent, individuals who perform well on one item are likely to
perform well on other items also. The Kuder–Richardson formulas provide a
quantitative estimate of this consistency.

Types of Kuder–Richardson Formulas

There are two main types:

1. KR-20 Formula

KR-20 is the most widely used and accurate measure of internal consistency for
dichotomous items.

KR_{20} = \frac{k}{k-1}\left(1 - \frac{\sum pq}{\sigma_X^2}\right)


Where:

 k = total number of items


 p = proportion of correct responses to an item
 q = 1 - p = proportion of incorrect responses
 \sum pq = sum of the product of p and q for all items
 \sigma_X^2 = variance of total test scores

Explanation:

The term pq represents the variance of each item, and the formula compares the
sum of item variances with the total test variance. If items are highly consistent,
the total variance will mainly reflect true score variance, resulting in a higher
reliability coefficient.

Characteristics:

 Takes into account item difficulty


 Provides accurate estimation
 Preferred in most testing situations

2. KR-21 Formula

KR-21 is a simplified version of KR-20 used when detailed item-level data is


unavailable.

KR_{21} = \frac{k}{k-1}\left(1 - \frac{M(k - M)}{k\sigma_X^2}\right)

Where:
 k = number of items
 M = mean score
 \sigma_X^2 = variance of total scores

Explanation:

KR-21 assumes that all items have equal difficulty, which simplifies
calculations but reduces accuracy.

Characteristics:

 Easier to compute
 Requires less data
 Less precise than KR-20

Interpretation of KR Coefficient

The value of KR reliability ranges from 0 to 1:

 0.90 and above → Excellent reliability


 0.80 – 0.89 → Good reliability
 0.70 – 0.79 → Acceptable
 Below 0.60 → Poor reliability

Higher values indicate that the test items are highly consistent and measure the
same construct effectively.

Relation to Cronbach’s Alpha


The Kuder–Richardson formula, especially KR-20, is considered a special case
of Cronbach’s Alpha, used specifically when items are dichotomous. While
Cronbach’s alpha can be applied to items with multiple scoring categories, KR
formulas are limited to binary responses.

Significance and Uses

 Widely used in objective-type tests such as multiple-choice exams


 Helps in test construction and evaluation
 Assists in identifying inconsistent or weak items
 Ensures that the test provides reliable and stable scores

Limitations

 Applicable only to dichotomous items


 KR-21 may produce biased estimates due to its assumptions
 High reliability does not guarantee validity
 Does not account for multidimensional constructs

Conclusion

In conclusion, the Kuder–Richardson formulas are essential tools in


psychometrics for assessing the internal consistency of tests with dichotomous
items. KR-20 is preferred for its accuracy, while KR-21 serves as a simpler
alternative when limited data is available. These formulas play a vital role in
ensuring that a test is reliable, consistent, and suitable for measuring a specific
construct.

You might also like