• an umbrella term for all that goes into the process of creating a test
.
.
o TEST CONCEPTUALIZATION – brain storming of ideas about what kind
of test a developer wants to publish
o Questions to ponder on when conceptualizing for new tests:
What is the test designed to measure?
1.
What is the objective?
2.
Is there a need for this kind of test?
3.
Who will use the test?
4.
Who will take the test?
5.
How will the test be administered?
6.
What is the ideal format of the test?
7.
Should more than one form of test be developed?
8.
.
o Questions to ponder on when conceptualizing for new tests:
What special training will be required of test users for
1.
administering or interpreting the test?
What types of responses will be required of test takers?
2.
Who benefits from an administration of this test?
3.
o PILOT WORK/PILOT STUDY/PILOT RESEARCH – preliminary research
surrounding the creation of a prototype of the test
oAttempts to determine how best to measure a targeted construct
oEntail lit reviews and experimentation, creation, revision, and deletion of
preliminary items
.
o TEST CONSTRUCTION – stage in the process that entails writing test items, revisions,
formatting, setting scoring rules
o Scaling – process of setting rules for assigning numbers in measurement
o Unidimensional Scaling – the quality of measuring a single construct, trait or other
attribute
o Example: Self- esteem, which are assumed to have a single dimension going from low
to high
o Multidimensional Scaling – the quality of a scale, test or so forth that is capable of
measuring more than one dimension of a construct
- employ different items or tests to measure each dimension of the construct separately,
and
then combine the scores on each dimension to create an overall measure of the
multidimensional construct
oExample: academic aptitude (mathematical, verbal ability), intelligence
.o Louis Leon Thurstone– was one the first and most productive scaling theorists. He invented
different methods for developing a unidimensional scale:
o The method of paired comparison
o The method of equal- appearing intervals
The method of paired comparison
– produces ordinal data by presenting with
pairs of two stimuli which they are asked to
compare:
The method of equal- appearing intervals
.
– one scaling method used to obtain data that are presumed to be interval
- Used to measure attitudes of people in interval nature
- Should have positive, negative and neutral statements
.
o Rating Scale– grouping of words, statements, or symbols on which judgments of the strength of a
particular trait are indicated by the test taker
o Summative Scale – final score is obtained by
summing the ratings across all the items
• Likert Scale – scale attitudes, usually
reliable
.
o Semantic Differential Scale – a person rates a series of concepts on several seven-point, bipolar
adjectival scales
o Gutmann Scale –It is designed so that if a person agrees with (or endorses) a particular item, they
are expected to agree with all previous items that are ranked lower in difficulty or extremity.
.o WRITING ITEMS
o ITEM POOL – reservoir or well from which the items will or will not be drawn for
the final version of the test
o The test developer may write a large number of items from personal
experience or academic acquaintance with the subject matter or experts
o ITEM FORMAT – form, plan, structure, arrangement, and layout of individual
test items
o The two types are:
o SELECTED RESPONSE FORMAT
o CONSTRUCTED- RESPONSE FORMAT
.
o SELECTED RESPONSE FORMAT
.
o Multiple Choice –Has three elements: stem (question), a correct option, and several incorrect
alternatives (distractors or foils)
o Should’ve one correct answer, has grammatically parallel alternatives, similar length, alternatives
that fit grammatically with the stem, avoid ridiculous distractors, not excessively long, “all of
the above”, “none of the above”
o Probability of getting the correct answer is 25%
o Matching Type –Test taker is presented with two columns: Premises and Responses
o Binary Choice –Usually takes the form of a sentence that requires the test taker to indicate
whether the statement is or is not a fact, e.g. True or False
o Probability of getting the correct answer is 50%
.o CONSTRUCTED RESPONSE FORMAT
o Completion Item–Requires the examinee to provide a
word or phrase that completes a sentence
o Should be worded properly so that the correct answer
is specific
o Short-answer Item –Should be written clearly enough
that the test taker can respond succinctly, with short
answer
o Essay Item– Respond by writing a composition
.o ITEM BANKS – relatively large and easily accessible collection of test questions
o Computerized Adaptive Testing – refers to an interactive, computer administered test-
taking process wherein items presented to the test taker are based in part on the test
taker’s performance on previous items
o The test administered may be different for each test taker, depending on the test
performance on the items presented
o Reduces floor and ceiling effects
oFloor Effects – occurs when there is some lower limit on a survey or questionnaire
and a large percentage of respondents score near this lower limit (test takers have
low scores)
oCeiling Effects – occurs when there is some upper limit on a survey or questionnaire
and a large percentage of respondents score near this upper limit (test takers have
high scores)
.o SCORING – the process of assigning scores
to performances
o Cumulative Scoring– the higher score
one achieved on the test, the higher the
test taker is on the ability that the test
purports to measure
o Ex: Ability Test
o Class Scoring/Category Scoring – test
taker responses earn credit toward
placement in a particular class or category
with other test takers who pattern of
responses is presumably similar in some
way
o Ex: DSM 5
. o Ipsative Scoring – comparing a test taker’s score on one scale within a test to another
scale within that same test
o Ex: EPPS (Edward Personal Preference Schedule
.
o The test should be tried out on people who are similar in critical
respects to the people for whom the test was designed
o An informal rule of thumb should be no fewer that 5 and preferably
as many as 10 for each item (the more, the better)
o Risk of using few subjects = phantom factors emerge
o Should be executed under conditions as identical as possible
.o After the first draft of the test has been administered to a representative group of examinees,
the test developer analyzes test scores and responses to individual items
o ITEM ANALYSIS - Statistical procedure used to analyze items
- refers to statistical methods used for selecting items for inclusion in a psychological test
- Indices to be computed:
* Item Difficulty Index
* Item Discrimination Index
* Item Reliability Index
* Item Validity Index
o Item Difficulty - defined by the number of people who get a particular item correct
o Item Difficulty Index- calculating the proportion of the total number of test takers who
answered the item correctly
.
o Item Difficulty Index- calculating the proportion of
the total number of test takers who answered the
item correctly
o The larger the Item Difficulty Index, the easier
the item
o The smaller the Item Difficulty Index, the more
difficult the item
o Formula = P = R/T
o Where P = the item difficulty index
R = number of correct responses
T = total number of responses
Exercise: Get the item Difficulty Index
Item Endorsement Index- in the context of Item 1 = Correct (45) Total (50)
personality test. Measure the percent of people Item 2 = Correct (10) Total (50)
who said yes, agreed with, or otherwise endorse
the item.
o Item Discrimination Index- indicates how adequately an item separates or discriminates
. the scores between high and low scorers
o Measure of the difference between the proportion of high scorers answering an item
correctly and the proportion of low scorers answering the item correctly
o Formula = D = U-L/ N
oWhere D = the item discrimination index
U = Number of correct responses in the upper 27% of test-takers
L = Number of correct responses in the lower 27% of test-takers
N = Number of students in each group (upper and lower)
Exercise: Get the item Discrimination Index
Item 1 = U (10) L (5) T(11)
Item 2 = U (6) L (6) T(11)
Item 3 = U (4) L (5) T(11)
.
.o ITEM RELIABILITY INDEX - provides an indication of the internal
consistency of a test
-The higher this index, the greater the test’s internal consistency
o ITEM VALIDITY INDEX - designed to provide an indication of the degree
to which a test is measure what it purports to measure
-The higher this index, the greater the test’s criterion-related validity
o Characterize each item according to its strength and weaknesses
.
o As revision proceeds, the advantage of writing a large item pool becomes more apparent because
some items were removed and must be replaced by the items in the item pool
o Administer the revised test under standardized conditions to a second appropriate sample of
examinee
o Cross-Validation – revalidation of a test on a sample of test takers other than those on who test
performance was originally found to be a valid predictor of some criterion
o Often results to validity shrinkage
oValidity Shrinkage – decrease in item validities that inevitably occurs after cross-
validation
o Co-validation – conducted on two or more test using the same sample of test takers
o Co-norming – creation of norms or the revision of existing norms
.