Test Development Process Overview
Test Development Process Overview
Initial @02/16/2024
Progress Done
Test Construction
an emerging social phenomenon or pattern of behavior might serve as the stimulus for the
development of a new test
Preliminary Questions
how the test developer defines the construct to be measured and how different it is
from other tests that measure the same construct
what the goal would the test achieve and how different it is from others
what ways will the new test be better than or different from existing ones
Test Development 1
factors such as age range, reading level, and cultural factors that might affect
testtaker response
What special training will be required of test users for administering and interpreting
the test?
norm-referenced
a good item is where high scorers tend to respond correctly and low scorers tend to
respond to the same item incorrectly
criterion-referenced
each item addresses the issue of whether the testtaker has met certain criteria
Test Development 2
may entail exploratory work with at least two groups of testtakers: one group
known to have mastered the knowledge or skills and another group known not to
have mastered them
the items that best discriminate between these two groups would be considered
“good items”
Pilot Work
test items may be pilot studied to evaluate whether they should be included in the final
form of the instrument
test developer typically attempts to determine how best to measure a targeted construct
process may entail literature reviews and experimentation as well as the creation,
revision, and deletion of preliminary test items
a necessity when constructing tests or other measuring instruments for publication and
wide distribution
Test Construction
Scaling
process by which a measuring device is designed and calibrated by which numbers (or
other indices)—scale values—are assigned to different amounts of the trait, attribute, or
characteristic being measured
Types of Scales
scales are instruments used to measure something (usually a trait, state, or ability)
Test Development 3
stanine scale → if all raw scores on the test are to be transformed into scores than
can range from 1 to 9
categorical → stimuli are placed into one of two or more alternative categories
that differ quantitatively with respect to some continuum
Scaling Methods
summative scale → final test score is obtained by summing the ratings across all
items
Likert scale → each item presents the testtaker with five to seven alternate
responses, usually on an agree-disagree or approve-disapprove continuum
they must select one of the stimuli according to some rule (ex: rule that they
agree more with one statement than the other or that they find one stimulus
more appealing than the other)
Test Development 4
for each pair of options, testtakers receive a higher score for selecting the
option deemed more justifiable by the majority of a group of judges
the judges would have been asked to rate the pairs of options before the
distribution of the test
a list of the options selected by the judges would be provided along with
the scoring instructions as an answer key
the test score would reflect the number of times the choices of a testtaker
agreed with those of the judges
Guttman Scale
all respondents who agree with the stronger statements of the attitude will also
agree with milder statements
Thurstone
scaling method used to obtain data that are presumed to be interval in nature
Steps:
Test Development 5
example: judges’ scales may range from 1-9 as if there were an equal
distance between each of the values (hence making it an interval
scale)
a mean and standard deviation of the judges’ ratings are calculated for
each statement
items are selected for inclusion in the final scale based on several criteria
test developer’s degree of confidence that the items have indeed been
sorted into equal intervals
Writing Items
Item Pool → reservoir or well from which items will or will not be drawn for the final
version of a test
comprehensive sampling provides a basis for content validity of the final version of
the test
approximately half of the initial item pool will be eliminated from the final
version, therefore the test developer needs to ensure that the final version
contains items that adequately sample the domain
any new or rewritten items would also be subjected to tryout so as not to jeopardize
the test’s content validity
multiple-choice format
Test Development 6
first draft must contain approximately twice the number of items that the final
version of the test will contain
alternate forms
multiply the number of items required in the pool for one form of the test by
the number of forms planned = total number of items needed for initial item
pool
a test developer may develop items for the item pool by writing a large number of items
from personal experience or academic acquaintance with the subject matter
Item Format
selected-response format
three types:
multiple-choice format
matching item
binary-choice item
Test Development 7
a good item contains a single idea, is not excessively long, and is not
subject to debate
constructed-response format
require testtakers to supply or to create the correct answer, not merely to select
it
three types:
completion item
short-answer item
essay item
item bank
advantages:
Test Development 8
accessibility to a large number of test items conveniently classified by
subject area, item statistics, or other variables
items may be added to, withdrawn from, and even modified in an item
bank
the computer may not permit the testtaker to continue with the test until
the practice items have been responded to in a satisfactory manner and the
testtaker has demonstrated an understanding of the test procedure
test administered may be different for each testtaker, depending on the test
performance on the items presented
only a sample of a total number of items in the item pool is administered to any
one testtaker
item branching
Scoring Items
cumulative model → the higher the score on the test, the higher the testtaker is on the
ability, trait, or other characteristic that the test purports to measure
Test Development 9
ipsative scoring → comparing a testtaker’s score on one scale within a test to another
scale within that same test
Test Tryout
the test should be tried out on people who are similar in critical respects to the people for
whom the test was designed
informal rule of thumb: no fewer than 5 and preferably 10 subjects per item
too few subjects = phantom factors → factors that are just artifacts of the small sample
size
test tryout should be executed under conditions as identical as possible to the conditions
under which the standardized test will be administered
in general, the test developer endeavors to ensure that differences in response to the test’s
items are due in fact to the items, not to extraneous factors
it is also the case than an item that is answered correctly by low scorers on the test
as a while may not be a good item
Item Analysis
Item-Difficulty Index
obtained by calculating the proportion of the total number of testtakers who answered
the item correctly
Test Development 10
p1
example:
p1 = 50/100
p1 = 0.5
calculated by obtaining the item-difficulty indices for all the test’s items
summing the item-difficulty indices for ALL test items and dividing by the
total number of items on the test
is approximately 0.5
optimal average item difficulty is usually the midpoint between 1.00 and the
chance success proportion
therefore, the optimal item difficulty is between .50 and 1.00 (could
be 0.75)
Test Development 11
2
in binary-choice items
equation:
equation
equation
Item-Reliability Index
equal to the product of the item-score standard deviation (s) and the correlation (r)
between the item score and the total test score
items that do not “load on” the factor they were written to tap can be revised or
eliminated
in other words, items that appear to not measure the construct they were
designed to measure can be revised or removed
Test Development 12
if too many items appear to be tapping a particular area, the weakest of such items
can be eliminated
Item-Validity Index
denoted by r1 c
item-score standard deviation multiplied by correlation between item score and criterion
score = item-validity index
calculating this index is important when the test developer’s goal is to maximize the
criterion-related validity of the test
Item-Discrimination Index
denoted by d
indicates how adequately an item separates or discriminates between high scorers and
low scorers on an entire test
in normal distribution, upper and lower areas will set the limit of 27% of
distribution of scores
Test Development 13
in platykurtic, upper and lower areas will set the limit of 33% of distribution of
scores
for most applications, any percentage between 25% to 33% will yield similar
estimates
the higher d, the greater the number of high scorers answering the item correctly
negative d-value → indicates that low-scoring examinees are more likely to answer the
item correctly than high-scoring examinees
after the test is given, the developer will take the 27% of high scorers and low scorers
respectively
high scorers = upper limit (U); low scorers = lower limit (L)
all scores in the upper limit are average as well as those in the lower limit
d = (U-L)/n
when:
more U-group members answer an item correctly than L-group members, the item
is reasonable or good
all U-group members answer an item correctly and all L-group members answer
incorrectly, the item is excellent
equal numbers of U- and L-group members answer the test correctly, the item is not
discriminating between testtakers at all
results in d = 0
all L-group members answer an item correctly and all U-group members answer
incorrectly, the item is bad
Test Development 14
by charting the number of testtakers in the U and L groups who chose each
alternative, the test developer can get an idea of the effectiveness of a distractor by
means of a simple eyeball test
cases:
the item is good because more U group members answered correctly, and each
of the distractors attracted some testtakers
Test Development 15
this is acceptable because there is still more correct U group members,
although distractor e was too effective because most L group members
answered it
this is a poor item because more L group members answered it correctly and
more U group members were attracted by distractors
Test Development 16
Test Development 17
Item-Characteristic Curves
cases:
this is a bad item because people of low ability got it correct and people of high
ability got it wrong
Test Development 18
this is also a bad item because people of moderate ability got it correct while both
people of high and low ability got it wrong
this is a good item because people of low ability got it wrong while people of high
ability got it correctly
this is an excellent item because it indicates that the probability is great that all
testtakers at or above moderate ability will respond correctly to the item and
probability is great that those falling below moderate ability will respond
incorrectly to the item
not desirable in tests measuring testtaker ability across all ability levels
Guessing
Test Development 19
three criteria that any correction for guessing must meet:
correction must recognize that a guess is not typically made on a random basis
but sometimes based on the testtaker’s knowledge of the subject matter and the
ability to rule out one or more of the distractors
some testtakers may be luckier than others in guessing the choices that are
keyed correct
test developer addresses the problem of guessing by including in the test manual:
explicit instructions regarding this guessing for the examiner to convey to the
examinees
Item Fairness
biased test item → item that favors one particular group of examinees in relation to
another when differences in group ability are controlled
the same proportion of persons from each group should pass any given item of
the test, provided that the persons all earned the same total score on the test
when majority of items are biased, it cannot be said that the test measures the same
abilities in the two groups
Speed Tests
the closer an item is to the end of the test, the more difficult it may appear to be
Test Development 20
this is because testtakers simply may not get to items near the end of the test
before the time runs out
because testtakers who know the material better may work faster and are thus
more likely to answer the latter items
may show positive item-total correlations because of the select group of examinees
reaching those items
not recommended solution: restrict the item analysis of items only to the items
completed by the testtaker
more knowledgeable examinees reach the later items, therefore part of the
analysis will be based only on a selected sample (low fairness)
recommended:
once item analysis is completed, norms should be established using the speed
conditions intended for use with the test in actual practice
qualitative methods → techniques of data generation and analysis that rely primarily on
verbal rather than mathematical or statistical procedures
compares individual test items to each other and to the test as a whole
Test Development 21
involve exploration of the issues through verbal means such as interviews and
group discussions conducted with testtakers and other relevant parties
designed to shed light on the testtaker’s thought processes during the administration
of a test
one-to-one basis where examinees are asked to take a test, thinking aloud as they
respond to each item
these yield valuable insights regarding the way individuals perceive, interpret, and
respond to the items
Expert Panels
sensitivity review
study of test items, typically conducted during the test development process, in
which items are examined for fairness to all prospective testtakers and for the
presence of offensive language, stereotypes, or situations
Test Revision
Approaches
items that have many weaknesses will be prime candidates for deletion or
revision
very difficult items have a restricted range because almost all testtakers get
them wrong → this will lack reliability and validity
Test Development 22
balance various strengths and weaknesses across items
test developers may purposefully include some more difficult items on a test
that has good items but is somewhat easy
as revision proceeds, the advantage of writing a large item pool becomes more and more
apparent
poor items can be eliminated in favor of those that were shown in the test tryout to
be good items
after the concerns have been balanced, the revised test will be administered under
standardized conditions to a second appropriate sample of examinees
once the test is deemed finished after item analysis, the norms may be developed from
the data, and the test will be deemed as “standardized” on the second sample
if item analysis indicate that the test is not yet in finished form, the steps of revision,
tryout, and item analysis are repeated until the test is satisfactory and standardization
can occur
an existing test should be kept in its present form as long as it remains “useful” but
should be revised “when significant changes in the domain represented, or new
conditions of test use and interpretation, make the test inappropriate for its intended
use”
stimulus materials look dated and current testtakers cannot relate to them
Test Development 23
meanings
test norms are no longer adequate due to membership changes in the population of
potential testtakers
test norms are no longer adequate due to age-shifts in the abilities measured over
time (age extension of the norms is necessary)
theory on which the test was originally based has been improved significantly, and
these changes should be reflected in the design and content of the test
note: test revision process in this sense is similar to the steps in making a brand new test
—especially when a new or improved theory is involved
any change in performance between the old and revised test cannot automatically be
viewed as a change in examinee performance
it is expected that items selected for the final version of the test will have smaller
item validities when administered to a second sample of testtakers because of the
operation of chance
validity shrinkage
co-validation → test validation process conducted on two or more tests using the same
sample of testtakers
can also be referred to as co-norming when used in conjunction with the creation of
norms or revision of existing norms
it is more economical for the test publisher and it eliminates the impact of sampling
error due to having only one sample be the norm for the tests being co-validated
Test Development 24