0% found this document useful (0 votes)
11 views24 pages

Test Development Process Overview

The document outlines the test development process, including conceptualization, construction, tryout, item analysis, and revision. It discusses key considerations for test construction, such as defining the construct, objectives, target users, and item formats. Additionally, it covers scaling methods, item writing, scoring, and the importance of pilot testing and item analysis to ensure reliability and validity of the test.

Uploaded by

cellyreads04
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
11 views24 pages

Test Development Process Overview

The document outlines the test development process, including conceptualization, construction, tryout, item analysis, and revision. It discusses key considerations for test construction, such as defining the construct, objectives, target users, and item formats. Additionally, it covers scaling methods, item writing, scoring, and the importance of pilot testing and item analysis to ensure reliability and validity of the test.

Uploaded by

cellyreads04
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Test Development

Book Chapters Cohen (8) ; Kaplan (6)

Initial @02/16/2024

Progress Done

Test Development Process

Conceptualization → Construction → Tryout →


Item Analysis → Revision
revision may go back to tryout to repeat the process

Test Construction

an emerging social phenomenon or pattern of behavior might serve as the stimulus for the
development of a new test

Preliminary Questions

What is the test designed to measure?

how the test developer defines the construct to be measured and how different it is
from other tests that measure the same construct

What is the objective of the test?

what the goal would the test achieve and how different it is from others

Is there a need for this test?

what ways will the new test be better than or different from existing ones

Who will use this test?

purpose or setting of the test (clinicians, educators, etc.)

Who will take this test?

Test Development 1
factors such as age range, reading level, and cultural factors that might affect
testtaker response

What content will the test cover?

pertaining to content coverage

How will the test be administered?

either or both individual or group

either or both pen-and-paper or computerized

What is the ideal format of the test?

ex: true-false, essay, multiple-choice

Should more than one form of the test be developed?

What special training will be required of test users for administering and interpreting
the test?

What types of responses will be required of testtakers?

pertaining to adaptations and accommodations for persons with disabilities

Who benefits from an administration of this test?

Is there any potential harm as the result of an administration of this test?

How will meaning be attributed to scores on this test?

either norm-referenced or criterion-referenced

Norm-referenced vs. Criterion-referenced

norm-referenced

a good item is where high scorers tend to respond correctly and low scorers tend to
respond to the same item incorrectly

criterion-referenced

each item addresses the issue of whether the testtaker has met certain criteria

commonly employed in licensing contexts and where mastery of a particular


material must be demonstrated

Test Development 2
may entail exploratory work with at least two groups of testtakers: one group
known to have mastered the knowledge or skills and another group known not to
have mastered them

the items that best discriminate between these two groups would be considered
“good items”

Pilot Work

preliminary research surrounding the creation of a prototype of a test

test items may be pilot studied to evaluate whether they should be included in the final
form of the instrument

test developer typically attempts to determine how best to measure a targeted construct

process may entail literature reviews and experimentation as well as the creation,
revision, and deletion of preliminary test items

a necessity when constructing tests or other measuring instruments for publication and
wide distribution

Test Construction

Scaling

process of setting rules for assigning numbers in measurement

process by which a measuring device is designed and calibrated by which numbers (or
other indices)—scale values—are assigned to different amounts of the trait, attribute, or
characteristic being measured

L.L. Thurstone → introduced the notion of absolute scaling—procedure for obtaining a


measure of item difficulty across samples of testtakers who vary in ability

Types of Scales

scales are instruments used to measure something (usually a trait, state, or ability)

age-based scale → testtaker’s performance as a function of age

grade-based scale → testtaker’s performance as a function of grade

Test Development 3
stanine scale → if all raw scores on the test are to be transformed into scores than
can range from 1 to 9

a test may be either unidimensional or multidimensional

unidimensional → only one dimension is presumed to underlie the ratings

multidimensional → more than one dimension is thought to guide the


testtaker’s responses

a test may be comparative as opposed to categorical

comparative → entails judgments of a stimulus in comparison with every other


stimulus on the scale

categorical → stimuli are placed into one of two or more alternative categories
that differ quantitatively with respect to some continuum

Scaling Methods

rating scale → grouping of words, statements, or symbols on which judgments of


the strength of a particular trait, attitude, or emotion are indicated by the testtaker

can be used to record judgments of oneself, others, experiences, or objects, and


can take many forms (symbols, true-false, numerical value)

results in ordinal-level data

summative scale → final test score is obtained by summing the ratings across all
items

Likert scale → each item presents the testtaker with five to seven alternate
responses, usually on an agree-disagree or approve-disapprove continuum

Likert scales are usually reliable

Likert concluded that assigning weights of 1 through 5 generally worked best

method of paired comparisons

testtakers are presented with pairs of stimuli

they must select one of the stimuli according to some rule (ex: rule that they
agree more with one statement than the other or that they find one stimulus
more appealing than the other)

Test Development 4
for each pair of options, testtakers receive a higher score for selecting the
option deemed more justifiable by the majority of a group of judges

the judges would have been asked to rate the pairs of options before the
distribution of the test

a list of the options selected by the judges would be provided along with
the scoring instructions as an answer key

the test score would reflect the number of times the choices of a testtaker
agreed with those of the judges

advantage → forces testtakers to choose between items

Guttman Scale

items range sequentially from weaker to stronger expressions of attitude,


belief, or feeling being measured

all respondents who agree with the stronger statements of the attitude will also
agree with milder statements

developed through the administration of a number of items to a target group

resulting data are analyzed by scalogram analysis

involves a graphic mapping of a testtaker’s responses

objective is to obtain an arrangement of items wherein endorsement of one


item automatically connotes endorsement of less extreme positions

Method of Equal-Appearing Intervals

Thurstone

scaling method used to obtain data that are presumed to be interval in nature

Steps:

reasonably large number of statements reflecting positive and negative


attitudes are collected

judges evaluate each statement in terms of how strongly it indicates a


variable is present

Test Development 5
example: judges’ scales may range from 1-9 as if there were an equal
distance between each of the values (hence making it an interval
scale)

judges are cautioned to focus their ratings on the statements, not on


their own views on the matter

a mean and standard deviation of the judges’ ratings are calculated for
each statement

items are selected for inclusion in the final scale based on several criteria

degree to which the item contributes to a comprehensive measurement


of the variable in question

test developer’s degree of confidence that the items have indeed been
sorted into equal intervals

note: a low standard deviation is indicative of a good item

administration → the values of items that the respondent selects (based on


the judges’ ratings) are averaged, producing a score on the test

this is an example of direct estimation scaling method wherein there is no need


to transform the testtaker’s responses into some other scale as opposed to
indirect estimation methods

Writing Items

Item Pool → reservoir or well from which items will or will not be drawn for the final
version of a test

comprehensive sampling provides a basis for content validity of the final version of
the test

approximately half of the initial item pool will be eliminated from the final
version, therefore the test developer needs to ensure that the final version
contains items that adequately sample the domain

any new or rewritten items would also be subjected to tryout so as not to jeopardize
the test’s content validity

multiple-choice format

Test Development 6
first draft must contain approximately twice the number of items that the final
version of the test will contain

alternate forms

multiply the number of items required in the pool for one form of the test by
the number of forms planned = total number of items needed for initial item
pool

a test developer may develop items for the item pool by writing a large number of items
from personal experience or academic acquaintance with the subject matter

help may be sought from others, including experts

Item Format

form, plan, arrangement, and layout of individual test items

selected-response format

require testtakers to select a response from a set of alternative responses

three types:

multiple-choice format

has three elements: stem, correct alternative or option, and several


distractors or foils

stem → question or statement

distractors/foils → incorrect options

matching item

testtaker is presented with two columns:

premises on the left

responses on the right

task is to determine which response is best associated with which


premise

binary-choice item

usually takes the form of a sentence that requires the testtaker to


indicate whether the statement is or is not a fact

Test Development 7
a good item contains a single idea, is not excessively long, and is not
subject to debate

the correct response must undoubtedly be one of the two choices

binary-choice items cannot contain distractor alternatives

constructed-response format

require testtakers to supply or to create the correct answer, not merely to select
it

three types:

completion item

requires the examinee to provide a word or phrase that completes a


sentence

characterized by “fill in the blank” types of tests

short-answer item

another form of completion item that is more on identification rather


than sentence completion

short-answer may be a word, a term, a sentence, or a paragraph

essay item

beyond a paragraph or two

a test item that requires the testtaker to respond to a question by


writing a composition, typically one that demonstrates recall of facts,
understanding, analysis, and/or interpretation

useful when the test developer wants the examinee to demonstrate a


depth of knowledge about a single topic

Writing Items for Computer Administration

item bank

relatively large and easily accessible collection of test questions

advantages:

Test Development 8
accessibility to a large number of test items conveniently classified by
subject area, item statistics, or other variables

items may be added to, withdrawn from, and even modified in an item
bank

computerized adaptive testing (CAT)

interactive, computer-administered test-taking process wherein items presented


are based in part on the testtaker’s performance on previous items

test might begin with some sample, practice items

the computer may not permit the testtaker to continue with the test until
the practice items have been responded to in a satisfactory manner and the
testtaker has demonstrated an understanding of the test procedure

test administered may be different for each testtaker, depending on the test
performance on the items presented

only a sample of a total number of items in the item pool is administered to any
one testtaker

based on previous response patterns, items that have a high probability of


being answered in a particular fashion are not presented, thus providing
economy in terms of testing time and total number of items presented

CAT tends to reduce floor effects and ceiling effects

item branching

ability of the computer to tailor the content and order of presentation of


test items on the basis of responses to previous items

Scoring Items

cumulative model → the higher the score on the test, the higher the testtaker is on the
ability, trait, or other characteristic that the test purports to measure

class scoring/category scoring → testtaker responses earn credit toward placement in a


particular class or category with other testtakers whose pattern of responses is
presumably similar in some way

used in some diagnostic systems wherein individuals must exhibit symptoms in


order to be diagnosed

Test Development 9
ipsative scoring → comparing a testtaker’s score on one scale within a test to another
scale within that same test

“John’s openness to experience is higher than his extraversion”

Test Tryout

the test should be tried out on people who are similar in critical respects to the people for
whom the test was designed

informal rule of thumb: no fewer than 5 and preferably 10 subjects per item

more subjects, the better

decreases the role of chance in data analysis

too few subjects = phantom factors → factors that are just artifacts of the small sample
size

test tryout should be executed under conditions as identical as possible to the conditions
under which the standardized test will be administered

in general, the test developer endeavors to ensure that differences in response to the test’s
items are due in fact to the items, not to extraneous factors

a good test item

reliable and valid

helps to discriminate testtakers

answered correctly (or in an expected manner) by high scorers on the test as a


whole

it is also the case than an item that is answered correctly by low scorers on the test
as a while may not be a good item

Item Analysis

Item-Difficulty Index

obtained by calculating the proportion of the total number of testtakers who answered
the item correctly

Test Development 10
p1 

p → denotes item difficulty

1 → (must be subscript) denotes item number

taken as a whole, it is read as “item-difficulty index for item 1”

the larger the item-difficulty index, the easier the item

higher p, the easier the item

example:

50 out of 100 examinees correctly answered item 1

p1 = 50/100

p1 = 0.5

index of difficulty of the average test item

calculated by obtaining the item-difficulty indices for all the test’s items

summing the item-difficulty indices for ALL test items and dividing by the
total number of items on the test

for maximum discrimination among the abilities of testtakers, the optimal


average item difficulty

is approximately 0.5

individual items on the test ranging from about 0.3 to 0.8

optimal average item difficulty is usually the midpoint between 1.00 and the
chance success proportion

probability of answering correctly by random guessing

example: in a binary-choice item, the probability of guessing correctly


based on chance is 1/2 or .50

therefore, the optimal item difficulty is between .50 and 1.00 (could
be 0.75)

generally, the midpoint representing the optimal item difficulty is obtained by


summing the chance success proportion and 1.00 and then dividing the sum by

Test Development 11
2

in binary-choice items

chance success proportion = 1/2 or 0.5

equation:

0.5 + 1.00 = 1.5


1.5 / 2 = .60 ← optimal item difficulty

in five-option multiple-choice items

chance success proportion = 1/5 or 0.2

equation

0.2 + 1.00 = 1.2

1.2 / 2 = .60 ← optimal item difficulty

in seven-option multiple-choice items

chance success proportion = 1/7 or 0.14

equation

0.14 + 1.00 = 1.14

1.14 / 2 = 0.57 ← optimal item difficulty

Item-Reliability Index

provides an indication of the internal consistency of a test

the higher this index, the greater the internal consistency

equal to the product of the item-score standard deviation (s) and the correlation (r)
between the item score and the total test score

statistical tool → factor analysis

items that do not “load on” the factor they were written to tap can be revised or
eliminated

in other words, items that appear to not measure the construct they were
designed to measure can be revised or removed

Test Development 12
if too many items appear to be tapping a particular area, the weakest of such items
can be eliminated

Item-Validity Index

statistic designed to provide an indication of the degree to which a test is measuring


what it purports to measure

higher index = greater criterion-related validity

two statistics must be known first:

item-score standard deviation

p is equal to the item difficulty index of item 1

s is equal to inter-score standard deviation

correlation between the item score and the criterion score

denoted by r1 c

item-score standard deviation multiplied by correlation between item score and criterion
score = item-validity index

calculating this index is important when the test developer’s goal is to maximize the
criterion-related validity of the test

Item-Discrimination Index

denoted by d

indicates how adequately an item separates or discriminates between high scorers and
low scorers on an entire test

the estimate of item discrimination compares performance on a particular item with


performance in the upper and lower regions of a distribution of continuous test scores

in normal distribution, upper and lower areas will set the limit of 27% of
distribution of scores

Test Development 13
in platykurtic, upper and lower areas will set the limit of 33% of distribution of
scores

for most applications, any percentage between 25% to 33% will yield similar
estimates

the higher d, the greater the number of high scorers answering the item correctly

negative d-value → indicates that low-scoring examinees are more likely to answer the
item correctly than high-scoring examinees

after the test is given, the developer will take the 27% of high scorers and low scorers
respectively

high scorers = upper limit (U); low scorers = lower limit (L)

all scores in the upper limit are average as well as those in the lower limit

d = (U-L)/n

n = number of members in each group

when:

more U-group members answer an item correctly than L-group members, the item
is reasonable or good

all U-group members answer an item correctly and all L-group members answer
incorrectly, the item is excellent

it reaches d = +1.00 which is the highest possible value for d

equal numbers of U- and L-group members answer the test correctly, the item is not
discriminating between testtakers at all

results in d = 0

all L-group members answer an item correctly and all U-group members answer
incorrectly, the item is bad

worst possible value is d = -1.00

this means that the item should be revised or eliminated

Analysis of Item Alternatives

Test Development 14
by charting the number of testtakers in the U and L groups who chose each
alternative, the test developer can get an idea of the effectiveness of a distractor by
means of a simple eyeball test

cases:

the item is good because more U group members answered correctly, and each
of the distractors attracted some testtakers

this is acceptable because even though more U group members answered


correctly, a large number was attracted by a particular distractor

can be improved through revision

this is an excellent item because ALL U group members answered correctly,


while some L group members were attracted by distractors

Test Development 15
this is acceptable because there is still more correct U group members,
although distractor e was too effective because most L group members
answered it

this is a poor item because more L group members answered it correctly and
more U group members were attracted by distractors

item can be revised or eliminated

Summary of Item Analysis (Item Discrimination and Item Difficulty)

Test Development 16
Test Development 17
Item-Characteristic Curves

graphic representation of item difficulty and discrimination

ability → horizontal axis; probability of correct response (PCR) → vertical axis

item discrimination = slope

steeper the slope, the greater item discrimination

item difficulty = skewness

easy item → negative skew

difficult item → positive skew

cases:

this is a bad item because people of low ability got it correct and people of high
ability got it wrong

Test Development 18
this is also a bad item because people of moderate ability got it correct while both
people of high and low ability got it wrong

this is a good item because people of low ability got it wrong while people of high
ability got it correctly

a desirable item-characteristic curve is one that shows a linear increase in


scores as ability increases

this is an excellent item because it indicates that the probability is great that all
testtakers at or above moderate ability will respond correctly to the item and
probability is great that those falling below moderate ability will respond
incorrectly to the item

this item may be desirable in selection of applicants based on a cutoff score

not desirable in tests measuring testtaker ability across all ability levels

Other Considerations in Item Analysis

Guessing

Test Development 19
three criteria that any correction for guessing must meet:

correction must recognize that a guess is not typically made on a random basis
but sometimes based on the testtaker’s knowledge of the subject matter and the
ability to rule out one or more of the distractors

correction must deal with omitted items (hindi sinagutan or nalaktawan)

some testtakers may be luckier than others in guessing the choices that are
keyed correct

to date, no solution to the problem of guessing has been deemed entirely


satisfactory

test developer addresses the problem of guessing by including in the test manual:

explicit instructions regarding this guessing for the examiner to convey to the
examinees

specific instructions for scoring and interpreting omitted items

Item Fairness

refers to the degree, if any, a test item is biased

biased test item → item that favors one particular group of examinees in relation to
another when differences in group ability are controlled

item-characteristic curves can be used to identify biased items

when two groups exhibit significantly differ in item-characteristic curves even


though the groups do not differ in total test score

the same proportion of persons from each group should pass any given item of
the test, provided that the persons all earned the same total score on the test

biased items must be revised or eliminated from the test

when majority of items are biased, it cannot be said that the test measures the same
abilities in the two groups

Speed Tests

speed tests yield misleading or uninterpretable results

the closer an item is to the end of the test, the more difficult it may appear to be

Test Development 20
this is because testtakers simply may not get to items near the end of the test
before the time runs out

may show high item discrimination in late-appearing items

because testtakers who know the material better may work faster and are thus
more likely to answer the latter items

may show positive item-total correlations because of the select group of examinees
reaching those items

not recommended solution: restrict the item analysis of items only to the items
completed by the testtaker

reasons for opposition:

item analyses of the later items would be based on a progressively smaller


number of testtakers, become less reliable

more knowledgeable examinees reach the later items, therefore part of the
analysis will be based only on a selected sample (low fairness)

because the more knowledgeable testtakers are more likely to score


correctly, it would make later items seem easier than they are (low
difficulty)

recommended:

administer the test to be item-analyzed with generous time limits to complete


the test

once item analysis is completed, norms should be established using the speed
conditions intended for use with the test in actual practice

Qualitative Item Analysis

qualitative methods → techniques of data generation and analysis that rely primarily on
verbal rather than mathematical or statistical procedures

qualitative item analysis

nonstatistical procedures designed to explore how individual test items work

compares individual test items to each other and to the test as a whole

Test Development 21
involve exploration of the issues through verbal means such as interviews and
group discussions conducted with testtakers and other relevant parties

“Think aloud” test administration

designed to shed light on the testtaker’s thought processes during the administration
of a test

one-to-one basis where examinees are asked to take a test, thinking aloud as they
respond to each item

these yield valuable insights regarding the way individuals perceive, interpret, and
respond to the items

Expert Panels

sensitivity review

study of test items, typically conducted during the test development process, in
which items are examined for fairness to all prospective testtakers and for the
presence of offensive language, stereotypes, or situations

forms of content bias may include status, stereotype, familiarity, offensive


choice of words, and others

experts on a particular culture can inform test developers on optimal ways to


achieve desired measurement ends with specific populations of testtakers

Test Revision

Approaches

characterize each item according to its strengths and weaknesses

some items may be highly reliable but lack criterion validity

other items may be purely unbiased but too easy

items that have many weaknesses will be prime candidates for deletion or
revision

very difficult items have a restricted range because almost all testtakers get
them wrong → this will lack reliability and validity

Test Development 22
balance various strengths and weaknesses across items

test developers may purposefully include some more difficult items on a test
that has good items but is somewhat easy

difficult items are the target for rewriting

change the blueprint of the test according to its purpose

for educational placement or employment → biggest concern is item bias

for testing skills and abilities → item discrimination is the priority

as revision proceeds, the advantage of writing a large item pool becomes more and more
apparent

poor items can be eliminated in favor of those that were shown in the test tryout to
be good items

after the concerns have been balanced, the revised test will be administered under
standardized conditions to a second appropriate sample of examinees

once the test is deemed finished after item analysis, the norms may be developed from
the data, and the test will be deemed as “standardized” on the second sample

if item analysis indicate that the test is not yet in finished form, the steps of revision,
tryout, and item analysis are repeated until the test is satisfactory and standardization
can occur

Test Revision in the Life Cycle of an Existing Test

an existing test should be kept in its present form as long as it remains “useful” but
should be revised “when significant changes in the domain represented, or new
conditions of test use and interpretation, make the test inappropriate for its intended
use”

tests are deemed to be due for revision on the following conditions:

stimulus materials look dated and current testtakers cannot relate to them

verbal content of the test, including administration instructions, contains dated


vocabulary not readily understood by current testtakers

certain words or expressions in the test may be perceived as inappropriate or


offensive to a particular group as popular culture changes and words take on new

Test Development 23
meanings

test norms are no longer adequate due to membership changes in the population of
potential testtakers

test norms are no longer adequate due to age-shifts in the abilities measured over
time (age extension of the norms is necessary)

reliability or validity, as well as effectiveness of test items, can be significantly


improved by a revision

theory on which the test was originally based has been improved significantly, and
these changes should be reflected in the design and content of the test

note: test revision process in this sense is similar to the steps in making a brand new test
—especially when a new or improved theory is involved

any change in performance between the old and revised test cannot automatically be
viewed as a change in examinee performance

Cross-validation and Co-validation

cross-validation → revalidation of the test on a sample of testtakers other than those on


whom test performance was originally found to be a valid predictor of some criterion

it is expected that items selected for the final version of the test will have smaller
item validities when administered to a second sample of testtakers because of the
operation of chance

validity shrinkage

decrease in item validities that inevitably occurs after cross-validation of


findings

co-validation → test validation process conducted on two or more tests using the same
sample of testtakers

can also be referred to as co-norming when used in conjunction with the creation of
norms or revision of existing norms

it is more economical for the test publisher and it eliminates the impact of sampling
error due to having only one sample be the norm for the tests being co-validated

Test Development 24

You might also like