Assessment of
Learning
OUTCOME BASED
MODULE
ED 9
Module 6:
Item Analysis and
Validation
Student Signature: Date
Returned:
ED 9
Learning Objectives:
At the end of this module, learners would be able to:
a. Explain the meaning of item analysis, item validity, reliability, item difficulty,
discrimination index.
b. Determine the validity and reliability of the given test items.
c. Determine the quality of the test item by its difficulty index, discrimination index and
plausibility of options (for a selected – response test.
tem Analysis: Difficulty index and Quality index
Read:
The teacher normally prepares a draft of the test. Such a draft is subjected to
item analysis and validation in order to ensure that the final version of the test would be
useful and functional/ First, the teacher tries out the daft test to a group of students of
similar characteristics as the intended test takers (try out phase). From the try-out
group, each item will be analysed in terms of its ability to discriminate between those
who know and those who does not know and also its level of difficulty (item analysis
phase). The item analysis will provide information that will allow the teacher to decide
whether to revise or replace an item (item revision phase). Then finally, the final draft of
the test is subjected to validation if the intent is to make use of the test as a standard
test for particular unit or grading period. We shall be concerned with these concepts in
this Chapter.
Two important characteristics of an item that will be of interest of a teacher are:
Item Difficulty
Discrimination Index
Item Difficulty is defined as the number of students who are able to answer the item
correctly, divided by the total number of students. It is usually presented in percentage.
Thus:
Item Difficulty = Number of Students with Correct Answers/ total number of
students
ED 9
Example 1: What is the difficulty index of an item if 25 students are unable to answer it
correctly while 75 answered it correctly.
Solution:
Item Difficulty = 75 / (25 + 75)
= 75 / 100
= 75%
The denominator is the total number of students who took the test, hence the
25+75 on the denominator of the first mathematical sentence.
Example 2: 25 students answered the items correctly while 75 students did not. The
total number of students is 100 so the difficulty index is:
Solution:
Item difficulty = 25 / 100
= 25%
It is a more difficult test item than that one with a difficulty index of 75. A high
percentage indicates that the test item is easy while a low percentage indicates a
difficult item.
The problem with this difficulty index is that it might not actually indicate that the
item is difficult or easy. For example, a student unfamiliar with the subject will never be
able to answer the given item. The following rule of thumb is being used when deciding
whether the test item is too easy or too difficult.
Range of Difficulty Index Interpretation Action
0-0.25 difficult Revise or discard
0.26-0.75 right difficulty Retain
0.76-above easy Revise or discard
Difficult items tend to discriminate between those who know the answer and
those who do not know the answer. Conversely, easy items cannot discriminate
between these two groups of students. We are therefore interested in deriving a
measure that will tell us whether an item can discriminate between these two groups of
students. Such measure is called Index of Discrimination.
An easy way to derive such measure is to measure how difficult an item is with
respect to those in the upper 25% of the class and how difficult it is with respect to the
lower 25% of the class. If the 25% of the class found the item easy yet the lower 25%
found it difficult, then the item can discriminate properly between these two groups.
Thus:
ED 9
Index of discrimination = DU – DL
Where: U = Upper Group, L = Lower Group
Example 3: Obtain the index of discrimination of the item if the upper 25% of the class
had a difficulty index of 0.60 (i.e., 60% of the upper 25% of the class got the correct
answer) while the lower 25% of the class had a difficulty index of 0.20.
Solution: Here, DU = 0.60 while DL = 0.20. Thus, the index of discrimination is:
Index of Discrimination = 0.60-0.20 = 0.40.
Discrimination index is the difference between the proportion of the top scorers
who got an item correct and the proportion of the lower scorers who got the item right.
The discrimination index range is between -1 and +1. The closer the discrimination
index is to +1, the more effectively the item can discriminate or distinguish between the
two groups of students. A negative discrimination index means more from the lower
group got the item correctly. The last item is not good and so must be discarded.
Theoretically, the index of discrimination can range from -1.0 (when DU = 0 and
DL = 1) to 1.0 (when DU = 1 and DL = 0). When the index of discrimination is equal to -
1, then this means that all of the lower 25% of the students got the correct answers
while all of the upper 25% got the wrong answer. In a sense, such an index
discriminates correctly between the two groups but the item itself is highly questionable.
Why should the bright ones get the wrong answer and the poor ones get the right
answer? On the other hand, if the index of discrimination is 1.0, then this means that all
the lower 25% failed to get the correct answer while all of the upper 25% got the correct
answer. This is a perfectly discriminating item and is the ideal item should be included in
the test. From these discussions, let us agree to discard or revise the items that have
negative discrimination index for although they discriminate correctly between the upper
and lower 25% of the class the content of the item itself may be highly dubious or
doubtful. As in the case of the index of difficulty, we have the following rule of thumb:
Index Range Interpretation Action
-1.0 to -0.50 Can discriminate but the Discard
item is questionable
-0.55 to 0.45 Non discriminating Revise
0.46 to 1.0 Discriminating item Include
Example 4: Consider a multiple choice type of test of which the following data were
obtained.
Item Options
1. A. B* C D
0 40 20 20 Total
ED 9
0 15 5 0 Upper 25%
0 5 10 5 Lower 25%
The correct response is B. Let us compute the difficulty index and the index of
discrimination.
Difficulty Index = No. of students getting the correct response / total
= 40/100 = 40%, within range of “good item”
The discrimination index can similarly be computed:
DU = no. of student in the upper 25% with correct response / no. of student in the upper
25%
= 15/20 = 0.75 = 75%
DL = no. of student in the lower 25% with correct response / no. of student in the lower
25%
= 5/20 = 0.25 = 25%
Discrimination Index = DU-DL = 0.75 – 0.25 = 0.50 or 50%
Thus, the item also has a “good discriminating power”.
It is also instructive to note that the distracter “A” is not a good distracter since it
was never selected by the students. It is an implausible distracter. Distracter C and D
appears to have a good appeal as distracters. They are plausible distracters.
Index of difficulty:
Ru+ RL
P= T x100
Where:
Ru = The number in the upper group who answered the item correctly.
RL = The number in the lower group who answered the item correctly.
T = The total number who tried the item.
Index of item Discriminating Power
Ru+ RL
D= 1
T
2
Where:
ED 9
P = percentage who answered the item correctly.
R = number who answered the item correctly.
T = total number who tried the item.
on 2: Validation and Validity
After performing the item analysis and revising the items which needs revision,
the next step is to validate the instrument. The purpose of the validation is to determine
the characteristics of the whole test itself, namely, the validity and reliability of the test.
Validation is the process of collecting and analysing evidence to support the
meaningfulness and usefulness of the test.
Validity – validity is the extent to which a test measures what it purports to measure or
as referring to the appropriateness, correctness, meaningfulness and usefulness of the
specific decisions a teacher makes based on the test results.
A test is valid when it is aligned with the learning outcome.
A teacher who conducts test validation might want to gather different kinds of
evidence. There are essentially three main types of evidence that may be collected
namely:
Content-related evidence of validity: refers to the content and format of the
instrument. How appropriate is the content? How comprehensive? Does it logically get
at the intended variable? How adequately does the sample of items or questions
represents the contents to be assessed?
Criterion-related evidence of validity: refers to the relationship between scores
obtained using the instrument and scores obtained using one or more other tests (often
called criterion). How strong is this relationship? How well do scores estimate present or
predict future performance of a certain type?
Construct-related evidence of validity: refers to the nature of the physiological
construct or characteristic being measured by the test. How well does a measure of the
construct explain differences in the behaviour of the individuals or their performance on
a certain task?
The usual procedure for determining the content validity may be describe as
follows: The teacher writes out the objectives of the test based on the table of
specifications and then gives these together with the test to at least two (2) experts
along with the description of the intended test takers. The experts look at the objectives,
read over the items in the test items in the test and then place a check mark in front of
each question or item that they feel does not measure one or more objectives. They
ED 9
also place a check mark in front of each objective not assessed by any item in the test.
The teacher then rewrites any item checked and resubmit them to the experts and/or
writes new items to cover those objectives not covered by the existing test. This
continues until the experts approve of all the items and also until the experts agree that
all of the objectives are sufficiently covered by the test.
Lesson 2: Reliability
Reliability refers to the consistency of the scores obtained – how consistent they are
for each individual from one administration of an instrument to another and from one set
of items to another. We already gave the formula for computing the reliability of a test:
for internal consistency; for instance, we could use the split-half method or the Kuder –
Richardson formulae (KR – 20 or KR – 21).
Reliability and Validity are related concepts. If an instrument is unreliable, it
cannot get valid outcomes. As reliability improves, validity may improve (or it may not).
However, if an instrument is shown scientifically to be valid then it is almost certain that
it is also reliable.
The following table is a standard followed almost universally in educational test
measurement.
Reliability Interpretation
0.90 and above Excellent reliability; at the level of the best standardized tests.
0.80 to 0.90 Very good for a classroom test.
0.70 to 0.80 Good for a classroom test; in the range of most. There are
probably a few items which could be improved.
0.60 to 0.70 Somewhat low. This test needs to be supplemented by other
measures (e.g., more tests) to determine grades. There are
probably some items which could be improved.
0.50 to 0.60 Suggests the need for revision of test, unless it is quite short (ten
or fewer items). The test definitely needs to be supplemented by
other measures (e.g., more tests) for grading.
0.50 or below Questionable reliability. This test should not contribute heavily to
the course grade and it needs revision.
Let’s check your understanding!!!
A. Item analysis – Discrimination and Difficulty Index
1. Solve for the difficulty index of each test item:
Item No. 1 2 3 4 5
No. of Correct Responses 2 10 20 30 15
ED 9
No. of Students 50 30 30 30 40
Difficulty Index
2. Which is the most difficult? Most easy?
3. Which needs revision? Which should be discarded? Why?
B. Solve for the discrimination index of the following test items:
Item No. 1 2 3 4 5
UG LG UG LG UG LG UG LG UG LG
No. of Correct 12 20 10 20 20 10 10 24 20 5
responses
No. of 25 25 25 25 25 25 25 25 25 25
Students
Discriminatio
n Index
Based on the computed discrimination index, which are good test items? Which
are not good test items?
C. Enumerate the three types of validity evidence. Which of the following
evidence is the most difficult to measure and why?
ED 9
References:
Published:
Assessment in Learning 1, 4th Edition, R.L. Navarro, LPT, PhD
R.G. Santos, LPT, PhD
B.B. Corpuz, LPT, PhD
K to 12 English Curriculum Guide, May 2016 ([Link]
Congratulations for Finishing Module 5.
Keep up the Good Work!
ED 9