0% found this document useful (0 votes)
7 views14 pages

Item Analysis and Validation Techniques

This document discusses item analysis and validation in educational assessments, focusing on item difficulty, discrimination index, validity, and reliability. It outlines the importance of analyzing test items to ensure they effectively measure student knowledge and the need for validation to confirm the test's meaningfulness. The document also explains the types of validity evidence and the relationship between reliability and validity in testing.

Uploaded by

camercado
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views14 pages

Item Analysis and Validation Techniques

This document discusses item analysis and validation in educational assessments, focusing on item difficulty, discrimination index, validity, and reliability. It outlines the importance of analyzing test items to ensure they effectively measure student knowledge and the need for validation to confirm the test's meaningfulness. The document also explains the types of validity evidence and the relationship between reliability and validity in testing.

Uploaded by

camercado
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

ASSESSMENT

IN LEARNING
1
Module 6

ITEM ANALYSIS AND


VALIDATION
Overview

In this chapter, the following concepts need to be discussed


to know the effectiveness of a test and to make it useful and
functional: Item Analysis including the two important
characteristics of an item that will be of interest of a teacher, the
item difficulty and discrimination index, Validation and Reliability.

Learning Objectives

At the end of this lesson, you should be able to:

• Define item analysis, item validity and reliability;

• Compare item difficulty and discrimination index;

• Understand the basic item analysis statistics;


• Cite the benefits of item analysis;

• Calculate the index of difficulty in every item of test; and


• Explain the 3 main types of validity evidence.
Keywords/Concepts

Item Analysis
a. item difficulty Basic Item Analysis
b. discrimination Statistics Validity
index

Lesson Proper

6.1 Item Analysis and Validation

The teacher normally prepares to revise or replace an item (item


a draft of the test. Such a draft is revision phase). Then, finally, the
subjected to item analysis and final draft of the test is subjected to
validation in order to ensure that validation if the intent is to make
the final version of the test would be use of the test as a standard test for
useful and functional. First, the the particular unit or grading
teacher tries out the draft test to a period. We shall be concerned with
group of students of similar these concepts in this chapter.
characteristics as the intended test
takers (try-out phase). From the try- I. Item Analysis
out group, each item will be There are two important
analyzed in terms of its ability to characteristics of an item that will
discriminate between those who be of interest to the teacher: (a) item
know and those who do not know difficulty (b) discrimination index
and also its level of difficulty (item We shall learn how to measure these
analysis phase). The item analysis characteristics and apply our
will provide information that will knowledge in making a decision
allow the teacher to decide whether about the item in question.
The difficulty of an item or item difficulty is defined as the number
of students who are able to answer the item correctly divided by the total
number of students. Thus:
The item difficulty is usually expressed in percentage.
Example: What is the item difficulty index of an item if 25 students are
unable to answer it correctly while 75 answered it correctly?
Here, the total number of students is 100, hence, the item difficulty index
is 75/100 or 75%.
One problem with this type of difficulty index is that it may not actually
indicate that the item is difficult (or easy). A student who does not know
the subject matter will naturally be unable to answer the item correctly
even if the question is easy. How do we decide the basis of this index
whether the item is too difficult or too easy? The following arbitrary rule is
often used in the literature:

Difficult items tend to discriminate between those who know and


those who do not know the answer. Conversely, easy items cannot
discriminate between these two groups of students. We are therefore
interested in deriving a measure that will tell us whether an item can
discriminate between these two groups of students. Such a measure is
called an index of discrimination.
An easy way to derive such a measure is to measure how difficult an
item is with respect to those in the upper 25% of the class and how difficult
it is with respect to those in the lower 25% of the class. If the upper 25% of
the class found the item easy yet the lower 25% found it difficult, then the
item can discriminate properly between these two groups. Thus:
Index of discrimination = DU – DL
Example: Obtain the index of discrimination of an item if the upper 25% of
the class had a difficulty index of 0.60% (i.e. 60% of the upper 25% got the
correct answer) while the lower 25% of the class had a difficulty index of
0.20.
Here, DU = 0.60 while DL = 0.20, thus index of discrimination = .60 - .20 =
.40.
Theoretically, the index of discrimination can range from -1.0 (when
DU = 0 and DL = l) to 1.0 (when DU = 1 and DL = 0). When the index of
discrimination is equal to -1, then this means that all of the lower 25% of
the students got the correct answer while all of the upper 25% got the
wrong answer. In a sense, such an index discriminates correctly between
the two groups but the item itself is highly questionable. Why should the
bright ones get the wrong answer and the poor ones get the right answer?
On the other hand, if the index of discrimination is 1.0, then this means
that all of the lower 25% failed to get the correct answer while all of the
upper 25% got the correct answer. This is a perfectly discriminating item
and is the ideal item that should be included in the test. From these
discussions, let us agree to discard or revise all items that have negative
discrimination index for although they discriminate correctly between the
upper and lower 25% of the class, the content of the item itself may be
highly dubious. As in the case of the index of difficulty, we have the
following rule of thumb:
The correct response is B. Let us compute the difficulty index and index of
discrimination: Difficulty index = no. of students getting correct
response/total = 40/100 = 40%, within range of “good item”
The discrimination index can similarly be computed: DU = no. of students
in upper 25% with correct response/no. of students in the upper 25% =
15/20 = .75 or 75%
DL = no. of students in lower 25% with correct response/no. of students in
the lower 25% = 5/20 = .25 or 25% Discrimination index = DU – DL = .75 -
.25 = .50 or 50%. Thus, the item also has a "good discriminating power",
It is also instructive to note that the distracter A is not an effective
distracter since this was never selected by the students. Distracters C and
D appear to have appeal as distracters.
Basic Item Analysis Statistics
The Michigan State University Measurement and Evaluation
Department reports a number of item statistics which aid in evaluating the
effectiveness of an item. The first of these is the index of difficulty which
MSU (http//[Link]/dept/) defines as the proportion of the total
group who got the item wrong. “Thus a high index indicates a difficult item
and a low index indicates an easy item. Some item analysts prefer an index
of difficulty which is the proportion of the total group who got an item
right. This index may be obtained by marking the PROPORTION RIGHT
option on the item analysis header sheet. Whichever index is selected is
shown as the INDEX OF DIFFICULTY on the item analysis print-out. For
classroom achievement tests, most test constructors desire items with
indices of difficulty no lower than 20 nor higher than 80, with an average
index of difficulty from 30 or 40 to a maximum of 60.
The INDEX OF DISCRIMINATION is the difference between the
proportion of the upper group who got an item right and the proportion of
the lower group who got the item right. This index is dependent upon the
difficulty of an item. It may reach a maximum value of 100 for an item with
an index of difficulty of 50, that is, when 100% of the upper group and none
of the lower group answer the item correctly. For items of less than or
greater than 50 difficulty, the index of discrimination has a maximum value
of less than 100. Interpreting the Index of Discrimination document
contains a more detailed discussion of the index of discrimination.”
(http//[Link]/dept).
More Sophisticated Discrimination Index
Item discrimination refers to the ability of an item to differentiate
among students on the basis of how well they know the material being
tested. Various hand calculation procedures have traditionally been used
to compare item responses to total test scores using high and low scoring
groups of students. Computerized analyses provide more accurate
assessment of the discrimination power of items because they take into
account responses of all students rather than just high and low scoring
groups.
The item discrimination index provided by ScorePak® is a Pearson
Product Moment correlation between student responses to a particular
item and total scores on all other items on the test. This index is the
equivalent of a point-biserial coefficient in this application. It provides an
estimate of the degree to which an individual item is measuring the same
thing as the rest of the items.
Because the discrimination index reflects the degree to which an item
and the test as a whole are measuring a unitary ability or attribute, values
of the coefficient will tend to be lower for tests measuring a wide range of
content areas than for more homogeneous tests. Item discrimination
indices must always be interpreted in the context of the type of test which
is being analyzed. Items with low discrimination indices are often
ambiguously worded and should be examined. Items with negative indices
should be examined to determine why a negative value was obtained. For
example, a negative value may indicate that the item was mis-keyed, so that
students who knew the material tended to choose an unkeyed, but correct,
response option.
Tests with high internal consistency consist of items with mostly
positive relationships with total test score. In practice, values of the
discrimination index will seldom exceed .50 because of the differing shapes
of item and total score distributions. ScorePak® classifies item
discrimination as “good” if the index is above .30; “fair” if it is between .10
and.30; and “poor” if it is below .10.
A good item is one that has good discriminating ability and has
sufficient level of difficult (not too difficult nor too easy). In the two tables
presented for the levels of difficulty and discrimination there is a little area
of intersection where the two indices will coincide (between 0.56 to 0.67)
which represent the good items in a test. (Source: Office of Educational
Assessment, Washington DC, USA
[Link]
item_analysis. html)
At the end of the Item Analysis report, test items are listed according
to their degrees of difficulty (easy, medium, hard) and discrimination
(good, fair, poor). These distributions provide a quick overview of the test,
and can be used to identify items which are not performing well and which
can perhaps be improved or discarded.
6.2 Validation

After performing the item analysis and revising the items which need
revision, the next step is to validate the instrument. The purpose of
validation is to determine the characteristics of the whole test itself,
namely, the validity and reliability of the test. Validation is the process of
collecting and analyzing evidence to support the meaningfulness and
usefulness of the test.
Validity. Validity is the extent to which a test measures what it
purports to measure or as referring to the appropriateness, correctness,
meaningfulness and usefulness of the specific decisions a teacher makes
based on the test results. These two definitions of validity differ in the
sense that the first definition refers to the test itself while the second refers
to the decisions made by the teacher based on the test. A test is valid when
it is aligned to the learning outcome.
A teacher who conducts test validation might want to gather different
kinds of evidence. There are essentially three main types of evidence that
may be collected: content-related evidence of validity, criterion-related
evidence of validity and construct-related evidence of validity. Content-
related evidence of validity refers to the content and format of the
instrument. How appropriate is the content? How comprehensive? Does it
logically get at the intended variable? How adequately does the sample of
items or questions represent the content to be assessed?
Criterion-related evidence of validity refers to the relationship
between scores obtained using the instrument and scores obtained using
one or more other tests (often called criterion). How strong is this
relationship? How well do such scores estimate present or predict future
performance of a certain type?
Construct-related evidence of validity refers to the nature of the
psychological construct or characteristic being measured by the test. How
well does a measure of the construct explain differences in the behavior of
the individuals or their performance on a certain task?
The usual procedure for determining content validity may be
described as follows: The teacher writes out the objectives of the test based
on the table of specifications and then gives these together with the test to
at least two (2) experts along with a description of the intended test takers.
The experts look at the objectives, read over the items in the test and place
a check mark in front of each question or item that they feel does not
measure one or more objectives. They also place a check mark in front of
each objective not assessed by any item in the test. The teacher then
rewrites any item so checked and resubmits to the experts and/or writes
new items to cover those objectives not heretofore covered by the existing
test. This continues until the experts approve of all items and also until the
experts agree that all of the objectives are sufficiently covered by the test.
In order to obtain evidence of criterion-related validity, the teacher
usually compares scores on the test in question with the scores on some
other independent criterion test which presumably has already high
validity. For example, if a test is designed to measure mathematics ability
of students and it correlates highly with a standardized mathematics
achievement test (external criterion), then we say we have high criterion-
related evidence of validity. In particular, this type of criterion-related
validity is called its concurrent validity. Another type of criterion-related
validity is called predictive validity wherein the test scores in the
instrument are correlated with scores on a later performance (criterion
measure) of the students. For example, the mathematics ability test
constructed by the teacher may be correlated with their later performance
in a division wide mathematics achievement test.
Apart from the use of correlation coefficient in measuring criterion-
related validity, Gronlund suggested using the so-called expectancy table.
This table is easy to construct and consists of the test (predictor) categories
listed on the left-hand side and the criterion categories listed horizontally
along the top of the chart. For example, suppose that a mathematics
achievement test is constructed and the scores are categorized as high,
average, and low. The criterion measure used is the final average grades of
the students in high school: Very Good, Good, and Needs Improvement.
The two-way table lists down the number of students falling under each of
the possible pairs (test, grade) as shown below:
The expectancy table shows that there were 20 students getting high
test scores and subsequently rated excellent in terms of their final grades;
25 students got average scores and subsequently rated good in their finals;
and finally, 14 students obtained low test scores and were later graded as
needing improvement. The evidence for particular test tends to indicate
that students getting high scores on it would be graded excellent; average
scores on it would be rated good later; and students getting low scores on
the test would be graded as needing improvement later. We will not be able
to discuss the measurement of construct-related validity in this book since
the method to be used require sophisticated statistical techniques falling
in the category of factor analysis.

6.3 Reliability

Reliability refers to the consistency of the scores obtained — how


consistent they are for each individual from one administration of an
instrument to another and from one set of items to another. We already
gave the formula for computing the reliability of a test: for internal
consistency; for instance, we could use the split-half method or the Kuder-
Richardson formulae (KR-20 or KR-21).
Reliability and validity are related concepts. If an instrument is
unreliable, it cannot yet valid outcomes. As reliability improves, validity
may improve (or it may not). However, if an instrument is shown
scientifically to be valid then it is almost certain that it is also reliable.
The following table is a standard followed almost universally in
educational tests and measurement.
ASSESSMENT

Performance Task

Answer the following.


a. Give at least three (3) benefits of Item Analysis.
b. In your own words explain in not more than two (2) sentences each
1. Validity
2. 2. Criterion-related evidence of validity
3. 3. Construct-related evidence of validity
c. How reliability and validity related? Explain in not more than three (3)
sentences.
CRITERIA FOR EVALUATION: Your output will be assessed based on the
following: Content (15 points) and Grammar and Spelling (5 points).
Assessment Task

A. Answer the following question.


• What is the purpose of item analysis?
• What are the steps in item analysis? Explain each step.
• How item analysis can increase teaching efficiency and assessment
accuracy?
B. Find the index of difficulty and index of item discriminating power.
1. Given: Ru=4 RL= 8 T=80
2. Given: Ru= 2 RL=6 T=60
3. Given: Ru= 3 RL= 6 T= 50
4. Given: Ru= 5 RL= 9 T= 80
5. Given: Ru= 4 RL= 9 T= 90
C. Explain your answer not more than 5 sentences. You will be graded
base on the criteria. Criteria: content 25points and spelling 5points.
• What is validity?
• Why do we need to do validation?
• What are the three main evidence?
• What is reliability?
• What is the relationship between validity and reliability?

REFERENCE

• Assessment in learning 1 author: Rosita L. Navarro Ph.d, Rosita G.


Santos Ph.D, and Brenda B. Corpuz Ph.D

You might also like