UNIT–5
RELIABILITY OF THE ASSESSMENT
TOOLS
Written by:
Dr. Muhammad Tanveer Afzal
Reviewed by:
Prof. Dr. Rehana Masrur
CONTENT
Sr. No Topic Page No
Introduction ...............................................................................................................103
Objectives .................................................................................................................103
5.1 Reliability ....................................................................................................104
5.2 Types of Reliability......................................................................................104
5.3 Factor Affecting Rehability .........................................................................109
5.4 Usability of Assessment Tools.....................................................................111
5.5 Summary ......................................................................................................112
5.6 Self Assessment Questions ..........................................................................112
5.7 References/Suggested Readings .................................................................114
INTRODUCTION
Assessment is an integral part of teaching-learning process which allows teachers to evaluate their
student’s achievement during an educational course. Many teachers feel deficiency in preparing and
grading exams, and most students are fearful of taking them. Yet test is a significant educational tool.
Therefore, this tool must be reliable and valid in such a way that everyone has credibility on its results.
Every classroom assessment measure must be appropriately reliable and valid, whether, it is the routine
classroom achievement test, attitudinal measure, or performance assessment. A measure must first be
reliable before it can be valid.
Teachers have been designing achievement tests since decades. But before preparing a test a teacher or
external exam designer must be aware of the qualities of an achievement test. A measure that ignores the
basic principles of developing a test may produce such results that may be unacceptable for the students,
and will not be measuring the actual performance.
Therefore this particular unit is meant for the prospective teachers addressing the concept and meaning of
the reliability, its types, factors affecting reliability of the tests and the usability of the tests.
OBJECTIVES
After studying this unit, prospective teachers will be able to:
define reliability in their own words.
apply the different methods of assuring reliability on the tests.
identify the factors affecting reliability.
construct a test and check how much reliable it is.
identify measures for reducing the problems in conducting the tests.
5.1. Reliability
What does the term reliability mean? Reliability means Trustworthy. A test score is called reliable when
we have reasons for believing the test score to be stable and objective. For example if the same test is
given to two classes and is marked by different teachers even then it produced the similar results, it may
be considered as reliable. Stability and trustworthiness depends upon the degree to which score is free of
chance error. We must first build a conceptual bridge between the question asked by the individual (i.e.
are my scores reliable?) and how reliability is measured scientifically. This bridge is not as simple as it
may first appear. When a person thinks of reliability, many things may come into his mind – my friend is
very reliable, my car is very reliable, my internet bill-paying process is very reliable, my client’s
performance is very reliable, and so on. The characteristics being addressed are the concepts such as
consistency, dependability, predictability, variability etc. Note that implicit, reliability statements, is the
behaviour, machine performance, data processes, and work performance may sometimes not reliable.
The question is “how much the scores of tests vary over different observations?”
5.1.1 Some Definitions of Reliability:
According to Merriam Webster Dictionary:
“Reliability is the extent to which an experiment, test, or measuring procedure yields the same results on
repeated trials.”
According to Hopkins & Antes (2000):
“Reliability is the consistency of observations yielded over repeated recordings either for one subject or a
set of subjects.”
Joppe (2000) defines reliability as:
“…The extent to which results are consistent over time and an accurate representation of the total
population under study is referred to as reliability and if the results of a study can be reproduced under a
similar methodology, then the research instrument is considered to be reliable.” (p. 1)
The more general definition of the reliability is: The degree to which a score is stable and consistent when
measured at different times (test-retest reliability), in different ways (parallel-forms and alternate-forms),
or with different items within the same scale (internal consistency).
5.2 Types of Reliability
Reliability is one of the most important elements of test quality. It has to do with the consistency, or
reproducibility, of an examinee's performance in the test. It's not possible to calculate reliability exactly.
Instead, we have to estimate reliability, and this is always an imperfect attempt. Here, we introduce the
major reliability estimators and talk about their strengths and weaknesses.
There are six general classes of reliability estimates, each of which estimates reliability in a different
way. They are:
i) Inter-Rater or Inter-Observer Reliability
To assess the degree to which different raters/observers give consistent estimates of the same
phenomenon. That is if two teachers mark same test and the results are similar, so it indicates the inter-
rater or inter-observer reliability.
ii) Test-Retest Reliability:
To assess the consistency of a measure from one time to another, when a same test is administered twice
and the results of both administrations are similar, this constitutes the test-retest reliability. Students may
remember and may be mature after the first administration creates a problem for test-retest reliability.
iii) Parallel-Form Reliability:
To assess the consistency of the results of two tests constructed in the same way from the same content
domain. Here the test designer tries to develop two tests of the similar kinds and after administration the
results are similar then it will indicate the parallel form reliability.
iv) Internal Consistency Reliability:
To assess the consistency of results across items within a test, it is correlation of the individual items
score with the entire test.
v) Split half Reliability:
To assess the consistency of results comparing two halves of single test, these halves may be even odd
items on the single test.
vi) Kuder-Richardson Reliability:
To assess the consistency of the results using all the possible split halves of a test.
Let's discuss each of these in turn.
5.2.1. Inter-Rater or Inter-Observer Reliability
Whenever we observe or activities of humans, we have to think about the procedure for reliable and
consistent results. For this two or more than two observers are assigned to observe the students or
teachers. So how do we determine whether two observers are being consistent in their observations? We
probably should establish inter-rater reliability by considering the similarity of the scores awarded by the
two observers. After all, if we use data to establish reliability, and we find that reliability is low. We
should have to focus upon the criteria established for the observation. And if it is tried first in the actual
situation then it may help to develop the reasonable criteria for the observation, and may be more
objective.
There are two major ways to actually estimate inter-rater reliability. If your measurement consists of
categories -- the raters are checking off which category each observation falls in -- you can calculate the
percent of agreement between the raters. For instance, let's say you had 100 observations that were being
rated by two raters. For each observation, the rater could check one of three categories. Imagine that on
86 of the 100 observations, the raters checked the same category. In this case, the percent of agreement
would be 86%. OK, it's a crude measure, but it does give an idea of how much agreement exists, and it
works no matter how many categories are used for each observation.
The other major way to estimate inter-rater reliability is appropriate when the measure is a continuous
one. There, all you need to do is calculate the correlation between the ratings of the two observers. For
instance, they might be rating the overall level of activity in a classroom on a 1-to-7 scale. You could
have them give their rating at regular time intervals (e.g., every 30 seconds). The correlation between
these ratings would give you an estimate of the reliability or consistency between the raters.
One might think of this type of reliability as "calibrating" the observers. There are other things one could
do to encourage reliability between observers, even without estimating it. For instance, in a psychiatric
unit where every morning a nurse had to do a ten-item rating of each patient on the unit. Of course, it’s
difficult to count on the same nurse being present every day, so there is a need to find a way to assure that
any of the nurses would give comparable ratings. The way we did, it was to hold weekly "calibration"
meetings where we would have all of the nurses ratings for several patients and discuss why they chose
the specific values they did. If there were disagreements, the nurses would discuss them and attempt to
come up with rules for deciding when they would give a "3" or a "4" for a rating on a specific item.
Although this was not an estimate of reliability, it probably went a long way towards improving the
reliability between raters.
Activity 5.1: Develop an essay type test for any class, administer it, get it marked from two raters and
then compare the marks given by the two raters for each question.
5.2.2. Test-Retest Reliability
Test-retest is a statistical method used to determine a test's reliability. The test is performed twice; in the
case of a questionnaire, this would mean giving a group of participants the same questionnaire on two
different occasions.
This form of reliability is used to judge the consistency of results across items on the same test.
Essentially, you are comparing test items that measure the same construct to determine the tests internal
consistency. When you see a question that seems very similar to another test question, it may indicate that
the two questions are being used to gauge reliability. Because the two questions are similar and designed
to measure the same thing, the test taker should answer both questions the same, which would indicate
that the test has internal consistency.
We estimate test-retest reliability when we administer the same test to the same sample on two different
occasions. This approach assumes that there is no substantial change in the construct being measured
between the two occasions. The amount of time allowed between measures is critical. We know that if we
measure the same thing twice that the correlation between the two observations will depend in part by
how much time elapses between the two measurement occasions. The shorter the time gap, the higher the
correlation; the longer the time gap, the lower the correlation. This is because the two observations are
related over time -- the closer in time we get the more similar the factors that contribute to error. Since
this correlation is the test-retest estimate of reliability, you can obtain considerably different estimates
depending on the interval.
Activity 5.2: Develop a test of English for sixth grade students, administer it twice with a gap of six
weeks, find the relationship between the scores of students between 1st and 2nd
administration.
5.2.3. Split-Half Reliability
Suppose you have to develop a test of 30 items and you want to know that how reliable the test is? What
you have to do is to administer the test, mark it and divide it in to two parts, in such a way that place all
the even numbered items (2,4,6…………) in one half and the odd numbered items (1,3,5…………..) in
the second. Calculate the reliability by using the Spearman-Brown prophecy formula given below.
Actually in split-half reliability we randomly divide all items that claim to measure the same contents into
two sets. We administer the entire instrument to a sample of students and calculate the total score for each
randomly divided half. The split-half reliability estimate is simply the correlation between these two total
scores.
Normally a single test is used to make two shorter alternate forms. This method has the advantage that
only one test administration is required, and therefore memory and the practice and maturation effects are
not involved. Furthermore, it does not require two tests. So it has many advantages over parallel form and
test-retest methods, therefore it is the most frequently used method of finding internal consistency of the
classroom tests. The formula used for the reliability of the full test is Spearman-Brown prophecy formula
as given below.
2(reliability of the half test)
Reliability of the Full Test = ______________________
1+ (reliability of the half test)
5.2.4 Parallel-Form Reliability
In parallel form reliability we have to create two different tests from the same contents to measure the
same learning outcomes. The easiest way to accomplish this is to write a large set of questions that
address the same contents and then randomly divide the questions into two sets. Now it’s time to
administer both instruments to the same students at the same time. The correlation between the two
parallel forms is the estimate of reliability. One major problem with this approach is that you have to be
able to write lots of items that reflect the same contents. This is often no easy to do job. Furthermore, this
approach makes the assumption that the randomly divided halves are parallel or equivalent. Even by
chance, this will sometimes not be the case. The parallel forms approach is very similar to the split-half
reliability described earlier. The major difference is that parallel forms are constructed so that the two
forms can be used independent of each other and considered equivalent measures. For instance, we might
be concerned about a testing threat to internal validity. If we use Form A for the pretest and Form B for
the posttest, we minimize that problem. It would even be better if we randomly assign individuals to
receive Form A or B on the pretest and then switch them on the posttest. With split-half reliability we
have an instrument that we wish to use as a single measurement instrument and only develop randomly
split halves for purposes of estimating reliability.
Activity 5.3: Make two tests of mathematics and compare its reliability through Parallel-Forms
Reliability method.
5.2.5. Internal Consistency Reliability
In internal consistency reliability estimation, we use our single test. The test is administered to a group of
students on one occasion to estimate reliability. In effect we judge the reliability of the instrument by
estimating how well the items that reflect the same content give similar results. We are looking at how
consistent the results are for different items for the same construct within the measure. There are a wide
variety of internal consistency measures that can be used.
5.2.6. Kuder Richardson Reliability
The estimates of internal consistency of the test are commonly calculated by using Kuder-Richardson
methods. These measures to extent to which items within one form of the test have as much in common
with one another as do the items in that one form with corresponding items in an equivalent form. The
strength of this estimate of reliability depends upon the context to which the entire test represents a single,
fairly consistent measure of a concept.
Normally these estimates are lower than the split halves but estimates higher than the test-retest and
parallel form estimates. These techniques are also called item total correlations. There are different
techniques to estimate the internal consistency of the test using K-R procedures, but two of them are more
frequently used by the measurement experts. The first KR-20 is difficult to calculate as it is based on the
information of the percentages of the students passing each item on the test. However, it gives more
accurate results (Kubiszyn and Borich, 2003). The KR-20 formula is given below.
KR20 Formula
Where “pq” provides a test score error variance for an "average" person, we know that the sampled
people vary, i.e., the variance of their raw scores is greater than zero. Persons with high or low scores
have less score error variance than those with scores near fifty percent correct where the score error
variance is maximum. Since the "average" person variance used in the KR20 formula is always larger
than the lower score error variance of persons with extreme scores, it must always overestimate their
score error variances.
The second formula, which is easier to calculate but slightly less accurate is called KR21. It requires only
the information about the number of items, the mean of the test score and the standard deviation. The
formula KR21 is as under.
n 2 mn m
r1
2 n 1
Studies indicated that this formula provide good results even when the item difficulties are not consistent.
5.3 Factors Affecting Reliability
Reliability of the test is an important characteristic as we use the test results for the future decisions about
the students’ educational advances and for the job selection and many more. The methods to assure the
reliability of the tests have been discussed. Many examples have been provided in order to in-depth
understanding of the concepts. Here we shall focus upon the different factors that may affect the
reliability of the test. The degree of the affect of each factor varies from the situation to situation.
Controlling the factor may improve the reliability and otherwise it may lower the consistency of
production of scores. Some of the factors that directly or indirectly affect the test reliability are given as
under.
5.3.1. Test Length
As a rule, adding more homogeneous questions to a test will increase the test's reliability. The more
observations there are of a specific trait, the more accurate the measure is likely to be. Adding more
questions to a psychological test is similar to adding finer distinctions on a measuring tape.
5.3.2. Method Used to Estimate Reliability
The reliability coefficient is an estimate that can change depending on the method used to calculate it. The
method chosen to estimate the reliability should fit the way in which the test will be used.
5.3.3 Heterogeneity of Scores
Heterogeneity is referred as the differences among the scores obtained from class. You may say that there
are some students who got high scores and some students who got low scores or intelligent students who
got high scores and other one got low scores or the difference could be due to any reason may be income
level, intelligence of the students, parents qualification etc. Whichever is the reason for the variability of
the scores the greater the variability (range) of test scores, the higher the reliability. Increasing the
heterogeneity of the examinee sample increases variability (individual differences) thus reliability
increases.
5.3.4 Difficulty
A test that is too difficult or too easy reduces the reliability (e.g., fewer test-takers get the answers
correctly or vice-versa). A moderate level of difficulty increases test reliability.
5.3.5 Errors that Can Increase or Decrease Individual Scores:
There might be some errors committed by the test developers that also affect the reliability of the tests
developed by teachers. These errors initially affect the students’ scores, mean deviate the scores from the
true ability of the students, and therefore affect the reliability. A careful consideration of these factors
may help to measure the true ability of the students.
The test itself: the overall look of the test may affect the students score. Normally a test is written
in well readable font size and style, the language of the test should be simple and understandable.
The test administration: After the development of the test, the test developer may have to prepare
the manual of the test administration, the time, environment, invigilation, and the anxiety also
affects students’ performance while attempting the test. Therefore the uniform administration of
the test leads to the increased reliability.
The test scoring: Marking of the test is another factor towards the variation in the scores of the
students. Normally there are many raters to rate the students’ responses/answers on the test.
Objective type test items and the marking rubric for essay type/ supply type test items help to get
the consistent scores.
Ensuring the Reliability of Test:
The most straightforward ways to improve a test’s reliability are`
First, calculate the item-test correlations and rewrite or reject any that are too low. Any item that does
not correlate with the total test at least (point-biserial) r = .25, should be reconsidered.
Second, look at the items that did correlate well and write more like them. The longer the test, the higher
the reliability will be.
5.4 Usability of Assessment Tools
Another important feature of a good assessment tool (Classroom test) is its usability. Classroom teachers
are well familiar with issues related to the usability and practicality of the tests, but they need to think of
how practical matters relate to testing. Usability refers to the extent to which a test can be used by
students and teachers to achieve specified goals in an effective and efficient manner. It also refers to
facilities available to test developers regarding both administration and scoring procedures of a test. As
far as administration is concerned, test developers should be attentive to the possibilities of giving a test
under reasonably acceptable conditions. For example, suppose a team of experts decide on giving a
listening comprehension test to large groups of examinees. In this case, test developers should make sure
those facilities such as audio equipments and/or suitable acoustic rooms are available. Otherwise, no
matter how reliable and valid the test may be, it will not be practical.
Regarding the scoring procedures of a test, one should pay attention to the problem of ease of scoring as
well as ease of interpretation of scores. For instance, assume that composition tests are excellent
indicators of language ability. Would it be possible to use it in large scale administrations? How would
the compositions be scored? How long would it take to score them? All these questions relate to the
usability of the test in terms of scoring. Therefore, test developers should be very careful in selecting and
administering a test. The test should be practical, i.e., it should be easy to administer, easy to score, and
easy to interpret the scores in other words easy to use.
A good classroom test should be “teacher-friendly”. A teacher should be able to develop, administer and
mark it within the available time and with available resources. Classroom tests are only valuable to
students when they are returned promptly and when the feedback from assessment is understood by the
student. In this way, students can benefit from the test-taking process. The issues regarding usability of
the test include cost of test development and maintenance, time (for development and test length),
resources (everything from computer access, copying facilities, AV equipment to storage space), ease of
marking, availability of suitable/trained markers and administrative logistics.
The following are two very important aspects that contribute towards the usability of the test.
Transparency
In simple words transparency is a process which requires from teachers to maintain objectivity and the
honesty for developing, administering, marking and reporting the test results. Transparency refers to the
availability of clear, accurate information to students about testing. Such information should include
outcomes to be evaluated, formats used, weighting of items and sections, time allowed to complete the
test, and grading criteria. Transparency makes students part of the testing process. No one could doubt
any aspect of the testing process. It also requires setting rules and keeping record of the testing process.
Security
Most teachers feel that security is an issue only in large-scale, high-stakes testing. However, security is
part of both reliability and validity. If a teacher invests time and energy in developing good tests that
accurately reflect the course outcomes, then it is desirable to be able to recycle the tests or similar
materials. This is especially important if analyses show that the items, distracters and test sections are
valid and discriminating. In some parts of the world, cultural attitudes towards “collaborative test-taking”
are a threat to test security and thus to reliability and validity. As a result, there is a trade-off between
letting tests into the public domain and giving students adequate information about tests.
5.5 Summary
This unit dealt with the reliability and usability of a good test. First, the concepts were defined, and then
the methods of estimating and assuring reliability and the factors affecting was discussed in detail.
Finally, the concept of practicality was explained.
The procedures for test construction may seem tedious. However, regardless of the complexity of the
tasks in determining the reliability and usability of a test, these concepts are essential parts of test
construction. It means that in order to have an acceptable and applicable test, upon which reasonably
sound decisions can be made, test developers should go through planning, preparing, reviewing, and
pretesting processes.
Without determining these parameters, nobody is ethically allowed to use a test for practical purposes.
Otherwise, the test users are bound to make inexcusable mistakes, unreasonable decisions and unrealistic
appraisals.
5.6 Self Assessment Questions
5.6.1 Essay Type
1. Define the term reliability and elaborate the importance and scope of reliability of a test.
2. State different types of reliability and explain each type with examples.
3. Give the limitations of test retest, split half and parallel form reliability methods.
4. Identify different factors affecting reliability a test also suggest measures to control the impact of
these factors.
5. Discuss the problems encountered by teachers and students while using the tests.
5.6.2 Objective Type
I Mark the following statements as true or false.
Assessment is an integral part of teaching learning process.
If a test measures for what it is designed to measure then it is a reliable test.
If the scores of the two administration of a test are consistent then it is called the test-
retest reliability of the test.
Administering two different forms of the test at a time is method of split half reliability.
If item does not correlate with the total test scores it should be reconsidered.
5.7 Reference/ Suggested Readings:
Anastasi, A. (1982). Psychological Testing. New York: Macmillan.
Babour, R. S. (1998). Mixing Qualitative Methods: Quality Assurance or Qualitative quagmire?
Qualitative Health Research, 8(3), 352-361.
Bazovsky, I. (1961). Reliability Theory and Practice. Prentice-Hall Report.
Bogdan, R. C. & Biklen, S. K. (1998). Qualitative Research in Education: An Introduction to Theory and
Methods (3rd ed.). Needham Heights, MA: Allyn & Bacon.
Cohen, R. J., Swerdlik, M. E., & Phillips, S. M. (1996). Psychological Testing and Measurement: An
Introduction to Tests and Measurement. Mountain View, CA: Mayfield Publishing Company.
Crooks, T. J. (1988). The Impact of Classroom Evaluation Practices on Students. Review of Educational
Research, 58(4): 438-481.
Hopkins, C.D. & Antes, R.L. (2000). Classroom Measurement and Evaluation, (3rd Ed). F.E. Peacock
Publishers, Int. ITASCA, ILLIONS.
Kubiszyn, T. & Borich, G. (2003). Educational Testing and Measurement: Classroom Application and
Practice. New York, Johan Wiley and Sons, Inc.
Joppe, M. (2000). The Research Process. Retrieved December 16, 2006, from
[Link]