Chapter 7: Utility Multiple Cut Scores: The use of more than one
cutoff point on a test to categorize testtakers
Benefit: A general term referring to a positive
into several different groups.
outcome or advantage.
Fixed Cut Score: A cutoff score established
Economic Benefit: A positive outcome or
based on an absolute standard or judgment of
advantage that can be measured in monetary
minimal competency.
terms.
Method of Predictive Yield: A technique for
Revenue: The total amount of money received
setting cut scores that considers the number of
from sales of goods or services.
positions to be filled and the predicted success
Noneconomic Benefit: A positive outcome or rate of those selected.
advantage that is not easily quantifiable in
Bookmark Method: A method for setting cut
monetary terms (e.g., improved morale, better
scores on educational tests where subject
public image).
matter experts review test items and place a
Specificity (True Negative Rate): The "bookmark" to indicate the minimum level of
probability that a test will correctly identify knowledge required.
individuals who do not have a particular
Item-Mapping Method: A cut score setting
condition.
technique that involves arranging test items
Positive Predictive Value: The probability that based on their difficulty and identifying a
an individual actually has a condition if the point that reflects the desired level of mastery.
test indicates its presence.
Known Groups Method: A technique for
Negative Predictive Value: The probability setting cut scores that involves administering
that an individual truly does not have a the test to groups known to possess or not
condition if the test indicates its absence. possess the characteristic of interest and then
determining a score that best differentiates the
Base Rate (Prevalence): The proportion of groups.
individuals in a population who have a
particular condition. Top-Down Selection: A selection process
where candidates are ranked based on their
Miss (False Negative): An instance where the total scores, and those with the highest scores
test incorrectly indicates the absence of a are selected until all positions are filled.
condition when it is actually present.
Multiple Hurdles: A selection process where
False Alarm (False Positive): An instance applicants must meet a minimum cutoff score
where the test incorrectly indicates the on each predictor before moving to the next
presence of a condition when it is actually stage.
absent.
Multistage or Multiple Hurdle Selection
Hit (True Positive): An instance where the test Process: A selection strategy involving a
correctly indicates the presence of a condition sequence of steps, where candidates must pass
when it is actually present. each "hurdle" (often a test or assessment) to
Correct Rejection (True Negative): An proceed.
instance where the test correctly indicates the Compensatory Model of Selection: A selection
absence of a condition when it is truly absent. approach where high scores on one predictor
Relative or Norm-Referenced Cut Score: A can offset low scores on another.
cutoff score determined by comparing an Taylor-Russell Table: A table used to estimate
individual's score to the scores of a norm the improvement in selection decisions (utility
group. gain) resulting from the use of a test, based on
Cut Score: A specific point on a test score its validity, the selection ratio, and the base
scale used to divide testtakers into different rate of success.
groups for decision-making purposes. Naylor-Shine Table: Similar to the Taylor-
Russell table, it provides an estimate of the
average increase in criterion performance
resulting from the use of a selection test.
Multiple Regression: A statistical technique Cost: The expenditure of resources (e.g., time,
used to predict a criterion variable based on a money, effort) associated with a particular
linear combination of two or more predictor action.
variables.
Index of Reliability: A numerical indicator of
Decision Theory: A framework for making the consistency or stability of a test's scores.
optimal choices under conditions of
Index of Validity: A numerical indicator of the
uncertainty, often involving the consideration
extent to which a test measures what it is
of potential outcomes and their associated
intended to measure.
probabilities and values.
Index of Soundness: A general term referring
Productivity Gain: An increase in the amount
to the overall quality and appropriateness of a
of output or work achieved.
test, encompassing reliability and validity.
Cost-Benefit Index: A ratio comparing the
Index of Utility: A numerical indicator of the
advantages (benefits) of a course of action to
practical value or usefulness of a test in a
its disadvantages (costs).
specific context.
Utility Gain: The increase in overall
Utility: In the context of testing, the overall
effectiveness or value resulting from the use of
usefulness or practical value of a test or
a particular selection procedure or
assessment procedure.
intervention.
Testing: The process of administering and
Return on Investment: A performance measure
interpreting tests to gather information.
used to evaluate the efficiency or profitability
of an investment. Test Utility: The degree to which the use of a
test contributes to the efficiency or
Expectancy Table: A table that displays the
effectiveness of decision-making.
probability of achieving different levels of
performance on a criterion based on test Psychometric Soundness: The overall
scores. adequacy of a test based on its reliability,
validity, and other relevant technical
Expectancy Data: Information, often presented
characteristics.
in a table or graph, that shows the relationship
between predictor scores and the likelihood of Discriminant Analysis: A statistical technique
success on a criterion. used to predict group membership based on a
set of predictor variables.
Sensitivity (Hit Rate): The probability that a
test will correctly identify individuals who do Factor Analysis: A statistical method used to
have a particular condition. identify underlying factors or dimensions that
explain the relationships among a set of
Overall Accuracy: The proportion of all
observed variables.
decisions (both positive and negative) made by
a test that are correct. Utility Analysis: A family of techniques used
to evaluate the costs and benefits of using a
Selection Ratio: The proportion of applicants
particular selection or assessment procedure.
who are hired out of the total number of
applicants. Brodgen-Cronbach-Gleser Formula: A specific
formula used in utility analysis to estimate the
Loss: A disadvantage, cost, or negative
monetary gain resulting from the use of a
consequence.
selection test.
Noneconomic Cost: A cost that is not easily
Adaptive Treatment: Tailoring interventions or
quantifiable in monetary terms (e.g., damage
job requirements to an individual's specific
to reputation, decreased morale).
abilities or needs.
Economic Cost: A cost that can be measured
Benefit Treatment: An approach that focuses
in monetary terms (e.g., expenses for testing
on maximizing the positive outcomes or
materials, personnel time).
advantages for individuals.
Productivity Treatment: An approach aimed at Scalogram Analysis: Used to analyze data
enhancing the output or efficiency of
individuals or groups. from a Guttman scale.
Utility Treatment: An approach that Completion or Short-Answer Item: Fill-in-the-
emphasizes the practical value and
effectiveness of interventions or procedures. blank.
Discriminant Function Analysis: A statistical Binary-Choice Item: Two choices (e.g.,
method similar to discriminant analysis, used
true/false).
to find the linear combination of predictor
variables that best separates groups.
Essay Item: Writing a longer response.
Angoff Method: A widely used method for
setting cut scores on multiple-choice tests, Multiple-Choice Format: Choosing from a list
where subject matter experts estimate the
probability that a minimally competent of answers.
testtaker would answer each item correctly.
Matching Item: Connecting items from two
lists.
CHAPTER 8
Constructed-Response Format: You create the
Rating Scale: A way to measure how much of answer.
something (like an attitude) someone has,
where they choose from a set of options. Selected-Response Format: You choose the
answer.
Likert Scale: A common rating scale (often for
attitudes) with choices like "strongly agree" to Test Conceptualisation: The initial idea for a
"strongly disagree." test.
Unidimensional Scale: Measures one thing. Test Construction: Writing items, setting rules.
Multidimensional Scale: Measures more than Test Tryout: Testing the test on a sample
one thing. group.
Summative Scale: Your total score is found by Test Development: The whole process of
adding up the scores on all the items. creating a test.
Guttman Scale: A scale where agreeing with a Pilot Work: Preliminary research for a test.
"stronger" statement means you'll also agree
Item Analysis: Checking if test questions are
with "weaker" ones.
good.
Ipsative Scoring: Comparing your scores on
Item-Difficulty Index: How many people got
different parts of the same test.
the question right.
Class or Category Scoring: Putting test-takers
Item-Discrimination Index: How well a
into groups based on their scores.
question separates high and low scorers.
Cumulative Model: Higher test score = more
Item-Endorsement Index: How many people
of the trait.
"agreed" with the item.
Item-Validity Index: Does the item measure Item Fairness: Whether a test item is fair.
what it should?
Ceiling Effect: Test is too easy, can't measure
Item-Reliability Index: Does the item high ability.
consistently measure?
Floor Effect: Test is too hard, can't measure
Item-Characteristic Curve: A graph showing low ability.
item difficulty and discrimination.
Validity Shrinkage: Test is less accurate with a
DIF Items/Differential Item Functioning: new group.
When an item works differently for different
Co-Norming: Developing norms for several
groups.
tests at once.
Qualitative Item Analysis: Non-statistical
Cross-Validation: Checking if a test works on
ways to study test items.
a new group.
Sensitivity Review: Checking for bias in test
Method of Paired Comparisons: Comparing
items.
items in pairs to rank them.
Scaling: Setting rules for assigning numbers.
Method of Equal-Appearing Intervals:
Scoring: Assigning numbers to test answers. Creating scales where items seem equally
spaced apart.
Measurement: Assigning numbers to things.
Categorical Scaling: Sorting into categories.
Computerized Adaptive Testing (CAT): A test
that changes questions based on your answers. Comparative Scaling: Comparing items to
each other.
Item Bank: A collection of test questions.
Think Aloud Test Administration: Observing
Item Pool: A collection of test questions.
test-takers' thoughts.
Item Format: The structure of test questions.
Giveaway Item: An easy item to start a test.
Item Branching: Computer changes questions
Anhedonia: Inability to experience pleasure.
based on your answers.
Asexuality: Lack of sexual attraction.
Anchor Protocol: A model for scoring
answers. Pseudobulbar Affect (PBA): Involuntary
laughing or crying.
Scoring Drift: Changes in scoring accuracy
over time. CHAPTER 9
Bias: A test unfairly favors a group.
Biased Test Item: A test question that unfairly
favors a group.