INTRODUCTION
Meaning of Psychological Test:
A psychological test is a standardised procedure to measure quantitatively or qualitatively one or more
than one aspect of a trait by means of a sample of verbal or non-verbal behaviour. The purpose of a
psychological test is twofold. First, it attempts to compare the same individual on two or more than
two aspects of a trait; and second, two are more than two persons may be compared on the same trait.
Such a measurement may be either quantitative or qualitative. In the words of Bean (1953), a test is
“an organised succession of the stimuli designed to measure quantitatively or to evaluate qualitatively
some mental process, trait or characteristics.” Likewise, Anastasi and Urbina (1997) have defined a
psychological test as “essentially an objective and standardised measure of sample of behaviour.”
Classification of Tests:
1. On the basis of the criterion of administrative conditions (individual test, group test)
2. On the basis of the criterion of scoring (objective test, subjective test)
3. On the basis of the criterion of time limit in producing the response (speed test, power test)
4. On the basis of the criterion of the nature or contents of items (verbal test, non-verbal test,
performance test, non-language test)
5. On the basis of the criterion of purpose or objective (intelligence test, aptitude test, personality test,
achievement test)
6. On the basis of the criterion of standardisation (instructions, scoring, reliability, validity etc.)
Characteristics of a Good Test:
• Objectivity: A test must have the trait of objectivity, i.e., it must be free from the subjective
element so that there is complete interpersonal agreement among experts regarding the
meaning of the items and scoring of the test. Obviously, objectivity here relates to two aspects
of the test – objectivity of the items and objectivity of the scoring system. By objectivity of
items is meant that the items should be phrased in such a manner that they are interpreted in
exactly the same way by all those who take the test. For ensuring objectivity of items, items
must have uniformity of order of presentation (That is either ascending or descending order).
By objectivity of scoring is meant that the scoring method of the test should be a standard one
so that complete uniformity can be maintained when the test is scored by different experts at
different times.
• Reliability: A test must also be reliable. Reliability here refers to self-correlation of the test. It
shows the extent to which the results obtained are consistent when the test is administered
once or more than once on the same sample with a reasonable time gap. Consistency in results
obtained in a single administration is an index of internal consistency of the test and
consistency in results obtained upon testing and retesting is an index of temporal consistency.
Reliability, thus, includes both internal consistency as well as temporal consistency. For a test
to be called sound, it must be reliable because reliability indicates the extent to which the
scores obtained in the test are free from such internal defects of standardisation which are
likely to produce errors of measurement.
• Validity: Validity is another prerequisite for a test to be sound. Validity indicates the extent to
which the test measures what it intends to measure, when compared with some outside
independent criterion. In other words, it is the correlation of the taste with some outside
criterion. The criterion should be an independent one and should be regarded as the best index
of trait or ability being measured by the test. Generally, validity of the test is dependent upon
the reliability because a test which yields inconsistent results (poor reliability) is ordinarily
not expected to correlate with some outside independent criterion.
• Norms: A test must also be guided by certain norms. Norms refer to the average performance
of a representative centre on a given test. There are four common types of norms – age norms,
grade norms, percentile norms and standards score norms. Depending upon the purpose and
use, a test constructor prepares any of these norms for his test. Norms help in interpretation of
the scores. In the absence of norms, no meaning can be added to the score obtained on the
taste.
• Practicability / Usability: A test must also be practicable/ usable from the point of view of
the time taken in its completion, length, scoring, etc. In other words, the test should not be
lengthy and the scoring method must not be difficult nor one which can only be done by
highly specialised persons. In addition, the test should be economical from the point of view
of money also.
Concept and Definition of Reliability:
Reliability is one of the important characteristics of any test. A well-made scientific instrument should
give consistent results. Reliability refers to this consistency of scores or measurement which is
reflected in reproducibility of the scores. A test is said to be consistent over a given period of time
when all the examinees retain their same relative ranks of two separate testing with the same test.
According to Anastasi and Urbina (1997), reliability refers to “the consistency of scores obtained by
the same individuals when re-examined the test on different occasions or with different sets of
equivalent items, or under variable examining conditions. The correlation coefficient indicating
temporal stability is known as the coefficient of stability and the correlation coefficient indicating
internal consistency is known as the coefficient of internal consistency or the alpha coefficient.
Methods (or Types) of Reliability:
There are four most common methods of estimating the reliability coefficient of test scores. These
methods are
[Link]-retest reliability: Repetition of a test is the simplest method of determining agreement
between two sets of scores: the test is given and repeated on the same group, and the correlation
computed between the first and second set of scores. Although test-retest is sometimes the only
available procedure, the method is open to several serious objections. If the test is repeated
immediately, many subjects will recall their first answers and spend their time on new material, thus
tending to increase their scores – sometimes by a good deal. Besides immediate memory effects,
practice and the confidence induced by familiarity with the material will almost certainly affect scores
when the test is taken for a second time. Moreover, transfer effects are likely to be different from
person to person. The test-retest method will estimate less accurately the reliability of a test which
contains novel features and is highly susceptible to practice than it will estimate the reliability of the
test scores which involve familiar and well-learned operations little affected by practice. Owing to
difficulties in controlling conditions which influence scores on retest, the test-retest method is
generally less useful than are the other methods.
[Link] Consistency Reliability or Split-Half Method: Internal consistency reliability indicates
the homogeneity of the test. If all the items of the test measure to same function or trait, the test is said
to be a homogeneous one and its internal consistency reliability would be pretty high. The most
common method of estimating internal consistency reliability is the split-half method in which the test
is divided into two equal or nearly equal halves.
3. Parallel-forms reliability or Alternate-forms reliability or Equivalent-form reliability or
Comparable-forms reliability: When alternative to parallel forms of a test can be constructed, the
correlation between Form A, for example, and Form B may be taken as a measure of the self-
correlation of the test. Under these conditions, the reliability coefficient becomes an index of the
equivalence of the two forms of the test. Parallel forms are usually available for standard
psychological and educational achievement tests. The alternate forms method is satisfactory when
sufficient time has intervened between the administration of the two forms to weaken or eliminate
memory and practise effects. When Form B of a test follows Form A closely, scores on the second
form of the test will often be increased because of familiarity. In drawing up alternate test forms, care
must be exercised to match test materials for content, difficulty and form; and precautions must be
taken not to have the items in the two forms too similar. When alternate forms are virtually identical,
reliability is too high; whereas when parallel forms are not sufficiently alike, reliability will be too
low.
4. Rational Equivalence: The method of rational equivalence represents an attempt to get an estimate
of the reliability of a test, free from the objections raised against the methods outlined above. Two
forms of a test are defined as ‘equivalent’ when corresponding items, a, A, b, B, etc., are
interchangeable; and when the inter-item correlations are the same for both forms. The method of
rational equivalence stresses the Inter-correlations of the items in the test and the correlations of the
items with the test as a whole. Kuder and Richardson (1973) did a series of researchers to remove
some of the difficulties of the split of method of estimating reliability. The main requirements for the
use of K-R formulas are:
i. All items of the test should be homogeneous, that is, each item should measure the same
factor or factors in the same proportion.
ii. Items should be scored either as +1 or 0, that is, all correct answers should be scored as +1
and all incorrect answers should be scored as 0.
Factors Influencing Reliability of Test Scores:
The reliability of test scores is influenced by a large number of factors and all these factors can be
categorised under two heads: extrinsic and intrinsic.
a) Extrinsic Factors:
Important extrinsic factors affecting reliability of a test may be enumerated as follows:
1. Group Variability: When the group of examinees being tested is homogeneous in ability, the
reliability of the test scores is likely to be lowered. But when the examinees vary widely in
their range of ability, that is, the group of examinees is a heterogeneous one; the reliability of
the test scores is likely to be high.
2. Guessing by the examinees: Guessing in a test is an important source of unreliability. In two
alternative response options there is a 50% chance of answering the items correctly on the
basis of the guess. In multiple choice items the chances of getting the answer correctly purely
by guessing are reduced. Guessing has two important effects upon the total scores. First, it
tests to raise the total score and thereby makes the reliability coefficient spuriously high.
Second, guessing contributes to the measurement error since the examinees differ in
exercising their luck over guessing the correct answer.
3. Environmental conditions: As far as possible, the testing environment should be uniform.
Arrangement should be such that light, sound and other comforts are equal and uniform to all
the examinees, otherwise it will tend to lower the reliability of the test scores.
4. Momentary fluctuations in the examinee: Momentary fluctuations influence the test scores
sometimes by raising the score and sometimes by lowering it. Accordingly, they tend to affect
the reliability. A broken pencil, momentary distraction by the sudden sound of an aeroplane
flying above, anxiety regarding non-completion of home-work, mistake in giving answer and
knowing no way to change it, are some of the afctors which explain momentary fluctuations
in the examinee.
b) Intrinsic Factors:
The main intrinsic factors affecting the reliability of a test are as follows:
1. Length of the test: A longer test tends to yield a higher reliability coefficient than a
shorter test. Lengthening the test or averaging total test scores obtained from several
repetitions of the same test tends to increase the reliability. It has been demonstrated that
averaging the test scores of several applications essentially gives the same result as
increasing the length of the test. For example, suppose the test scores are being averaged
after its three repeated applications and the test is lengthened three times the present
length and administered once, then statistically the result of both averaging and
lengthening will be same. Care has to be taken to see that added items should have the
same variance and the same inter-item correlation as items of the original test. When the
test has been increased or decreased by any multiple, the Spearman-Brown formula may
be used to estimate the reliability of the test.
2. Range of the test scores: If the obtained total scores on the test are very close to each
other, that is, if there is lesser variability among them, the reliability of the test is lowered.
On the other hand, if the total scores on the test vary widely, the reliability of the test is
increased.
3. Homogeneity of items: Homogeneity of the items is an important factor in reliability.
The concept of homogeneity of items includes two things – item reliability (or inter-item
correlation) and the homogeneity of function or trait measured from one item to another.
When the items measure different functions and the inter-correlations of items are zero or
near it (that is, when the test is heterogeneous one). The reliability is zero or very low.
When all the items measure the same function or trait and when the inter-item correlation
is high, the reliability of the test is also high.
4. Discrimination value: When the test is composed of discriminating items, the item-total
test correlation is likely to be high and then, the reliability is also likely to be high. But
when the items do not discriminate well between superior and inferior, that is, when the
items have poor discrimination values, the item-total correlation is affected, which
ultimately attenuates the reliability of the test.
5. Scorer Reliability: Scorer reliability (also known as reader reliability) is also an
important factor which affects the reliability of the test. By scorer reliability is meant how
closely two or more scores agree in scoring or rating the same set of responses. If they do
not agree, the reliability is likely to be lowered.
Reliability of Speed Test:
The distinction between speed and power test is difficult to be drawn in actual practice. In fact, this
distinction is one of degree and most tests depend upon both power and speed in varying proportion.
Single trial reliability coefficients like those found by Odd-even or Kuder-Richardson techniques are
inapplicable to speed tests. In speed test the individual differences in test scores are dependent upon
the speed of performances and as such, the reliability coefficient found by these methods will be
spuriously inflated. Now question may arise that what alternatives are available for estimating the
reliability of a speed test. Some suggestions can be made like this: Test-retest method, if applicable,
can be applied. Equivalent-form reliability can also be properly employed. Split-half techniques can
also be applied, provided the split is made in terms of time rather than in terms of items (Anastasi &
Urbina, 2002).
How to improve Reliability of test scores?
Reliability of test scores can be improved by controlling those factors which adversely affect the
reliability of the test. The following suggestions are useful for improving the reliability.
1. The group of examinees should be heterogeneous, that is, the examinees should vary widely
in their ability or trait being measured.
2. Items should be homogeneous.
3. Test should be preferably be a longer one.
4. As far as possible, items should be of moderate difficulty values; in other words, the indexes
of item difficulty should have the range of 0.40- 0.50-0.60.
5. Items should be discriminatory ones.
A part from these suggestions, there are two common approaches for improving the reliability of the
test. One approach emphasizes upon the length of the test and another approach emphasizes upon
throwing out items that pull down the reliability.
The approach emphasizing upon increasing the length of the test assumes that if new items similar to
the original set of items are added, the reliability of the test would tend to increase. Following
domain-sampling model, each item in the test is an independent sample of the trait or ability being
measured. The larger the sample, the more likely the fact that the test will represent the true
characteristics. According to this model, the reliability of a test increases as the number of items
increase.
Index of Reliability:
Index of reliability is statistically defined as the correlation coefficient between the obtained scores
and their true counterparts. This statistic indicates the extent to which we can depend upon obtained
scores as a measure of true scores. Index of reliability, thus, gives the maximum correlation which the
test is capable of yielding in its present form. Index of reliability is statistically equal to the square
root of the reliability coefficient of the test. Hence the formula is
r1∞ = √ rtt
where, r1∞ = the index of reliability and rtt = the reliability coefficient of the test