Psychometric Test Development and Analysis
Introduction to Psychometric Theories and Frameworks
Psychometric theories provide the foundational understanding for developing and interpreting
psychological tests. The Measurement and Assessment Framework (APA 2013) outlines key areas within this
field.
Observed Score Approach
This approach centers on the relationship between observed scores and true scores.
Classical Test Theory (CTT): A foundational theory in psychology, CTT posits that an observed score
(X ) is a combination of a true score (T ) and an error component (E ), represented by the equation
X = T + E . This fundamental equation is accepted across various psychometric theories, including
Item Response Theory (IRT) and the Rasch Model.
Generalizability Theory: An extension of CTT, this theory emerged in the 1970s. It seeks to
decompose the error component (E ) in the X = T + E equation into different facets or sources of
error. These sources can include the specific items selected, the raters involved, or characteristics of
the test administrator or examinee (e.g., gender). By unpacking error, Generalizability Theory
implicitly redefines the true score (T ).
Latent Variable Approach
This approach focuses on unobservable constructs or traits that underlie test performance.
structure
Factor-Analytic Theory: The oldest of the latent variable theories, introduced by Charles Spearman
+
underlyin
in the early 1900s. Advances from the 1960s to the 1990s saw factor-analytic theory integrated into
g factor
statistical modeling frameworks, utilizing likelihood theory for estimation and testing. This led to
confirmatory modeling strategies and, more recently, Exploratory Structural Equation Modeling
(ESEM), which blends exploratory and confirmatory approaches.
items +
abilty
Item-Response Theory (IRT): IRT gained popularity in psychology in the 1990s, though it emerged in
the 1960s. It requires large sample sizes for parametric models. IRT focuses on the range of latent
ability (often denoted as θ ) and the characteristics of items at different points along this continuum. A
key advantage of IRT over CTT is its ability to provide estimates of reliability and standard error of
measurement across the entire range of the latent variable, rather than a single estimate for the test
as a whole.
1/8
Rasch Theory: Mathematically, Rasch Theory can be characterized as a specific type of IRT with only
strict IRT -
measures
difficult +
one item parameter: item difficulty. It can be seen as a special case of a three-parameter IRT model
ability
where item discrimination and the lower asymptote (guessing parameter) are fixed. While
mathematically accurate, this description doesn't fully capture the philosophical orientation of Rasch
theory, which often prioritizes the model's fit to the data, sometimes leading to the exclusion of
respondents or items until the model fits.
Mixed-Models: These models represent statistical interconnections among CTT, factor-analytic
theory, and IRT, all within the framework of latent variable modeling.
Approaches in Test Development - rational- theoretical approach
- factor analytic approach
- empirical criterion-keyed approach
Several approaches can be used when developing tests. - projective approach
Rational-Theoretical Approach: This is the most common approach. It involves using existing theory
or an intuitive, common-sense method to develop test items. Expert opinion, whether from a single
researcher, a group of experts, or a theoretical framework, forms the basis for item development and
selection.
Factor-Analytic Approach: In this approach, items are selected based on whether they load onto a
particular factor. Statistical rules guide the development and selection of items.
Empirical Criterion-Keyed Approach: Items are chosen based on their ability to discriminate
between a group of interest and a control group. This method is less frequently used today.
Projective Approach: This approach utilizes ambiguous stimuli (e.g., inkblots, pictures) or tasks
where individuals create their own output (e.g., drawing a person). The underlying idea is that
individuals will project their own concerns, fears, attitudes, and beliefs onto their interpretations or
creations.
Standardization of Tests
Standardization is crucial for ensuring that tests are administered and scored consistently, allowing for
meaningful comparisons between individuals.
When to Standardize a Test
No suitable test exists for a particular purpose.
Existing tests for a specific purpose are inadequate for various reasons.
Basic Premises of Standardization
The independent variable is the individual being tested.
The dependent variable is their behavior.
2/8
Behavior is a function of the person and the situation: Behavior = Person × Situation.
In psychological testing, the goal is to ensure the person factor is the primary determinant of
behavior, while situational factors are controlled. Control of extraneous variables is the essence of
standardization.
What Should Be Standardized?
Test Conditions:
Physical Conditions: Uniformity in the testing environment (e.g., lighting, temperature, noise
levels).
Motivational Conditions: Ensuring consistent encouragement or instructions to motivate test-
takers.
Test Administration Procedure:
Uniformity in instructions and the administration process itself. This involves carefully following
the procedures specified by the test developers to ensure the test is used as intended.
Test administrators should create conditions that maximize the opportunity for optimal
performance.
Involvement of test-takers, parents, and organizations, as appropriate, in the testing process.
Sensitivity to Disabilities: Making accommodations for individuals with disabilities (e.g.,
increasing voice volume, referring to alternative tests).
Desirable Procedures for Group Testing: Attention to time limits, clarity of instructions,
physical conditions, and managing potential guessing.
Scoring: A consistent mechanism and procedure for scoring responses is essential for accurate
measurement. Scoring procedures should be audited regularly to ensure consistency and accuracy.
Interpretation: Similar results should lead to common interpretations. Factors influencing
interpretation include:
Psychometric Factors: Reliability, norms, standard error of measurement, and validity of the
instrument.
Test Taker Factors: Characteristics of the individual, such as group membership (gender, age,
ethnicity, race, socioeconomic status, marital status), and how these might impact test results.
Contextual Factors: The relationship of the test to the instructional program, opportunity to
learn, quality of the educational program, work and home environments, and other factors that
provide context for understanding test results. For example, if a test does not align with
curriculum standards and how they are taught, its results may not be informative.
- test condition
- test administration
- scoring
- interpretation
3/8
Tasks for Ensuring Uniformity
Test Developers: Prepare a comprehensive test manual that includes:
Materials needed (test booklets, answer sheets).
Time limits.
Oral instructions.
Demonstrations or examples.
Procedures for handling examinee queries.
Examiners/Test Users:
Ensure test user qualifications are met (training, licensing).
Conduct advance preparations, including familiarity with the test(s), testing procedures, and
instructions.
Prepare test materials.
Orient proctors for group testing.
Standardization Sample
A random sample of test-takers used to evaluate the performance of others. This sample is considered
representative if it consists of individuals similar to the population for whom the test is intended.
Objectivity in Testing
Objectivity in testing refers to ensuring that scoring and administration are free from bias.
Time-Limit Tasks: All examinees are given the same amount of time to complete a task.
Work-Limit Tasks: All examinees are required to complete the same amount of work.
Issue of Guessing: Strategies to mitigate the impact of guessing are considered.
Stages in Test Development
The process of developing a test typically involves several stages:
1. Test Conceptualization:
Define the Target Population/Clientele.
Specify the Objective of the Test.
Clearly define the variables/constructs to be measured.
Identify Test Constraints and Conditions.
Establish Content Specifications (Topics, Skills, Abilities).
Determine the Scaling Method (Comparative or Non-comparative scaling).
Define the Test Format:
4/8
1. Test Construction: Adhere to guidelines such as:
Deal with only one central thought per item.
Stimulus: Interrogative, Declarative, Blanks, etc.
Mechanism of Response: Structured vs. Free.
Multiple Choice:
More answer options (e.g., 4-5) reduce the chance of guessing.
Many items can aid in student comparison and reduce ambiguity, increasing reliability.
Easier to score.
Measures narrow facets of performance.
Potential drawbacks include increased reading time, transparent clues that may
encourage guessing, difficulty in writing reasonable choices, and the possibility of test-
takers getting correct answers by guessing.
True or False: Ideally, these questions should indicate a student's misunderstanding when
answered incorrectly, though this can be challenging to construct.
Be precise and brief.
Avoid awkward wording or dangling constructs.
Avoid irrelevant information.
- test conceptualization
Present items in positive language. - test construct
- test tryout
Avoid double negatives. - item analysis
Avoid terms like "all" and "none."
1.
Test Tryout: Administering the draft test to a sample group.
1.
Item Analysis: Evaluating the quality and appropriateness of test questions. This measures how well items
can assess the intended ability or trait.
1.
Test Revision: Modifying the test based on item analysis and other feedback.
Item Analysis
Item analysis is a critical process for evaluating the quality of test questions and how well they measure a
specific ability or trait.
5/8
Classical Test Theory (CTT) Analysis
These are the easiest and most widely used forms of analysis.
Known as the "true-score model," it uses the formula: Xte
= Txx ⋅ X + Xmean
Xte : True Score
Txx : Correlation Coefficient
X : Raw Score
Xmean : Mean Score
CTT assumes that a person's test score (X ) is composed of their "true score" (T ) plus some
measurement error (e), i.e., X = T + e.
Item-Response Theory (IRT) / Latent Trait Theory Analysis
Often referred to as "modern psychometrics."
Latent trait models aim to understand the underlying traits that produce test performance.
6/8
Key statistics employed include:
Item Difficulty: The proportion of examinees who answered the item correctly.
Number of correct responses
Formula: Item Difficulty = Total number of examinees
A higher item mean indicates an easier item; a lower item mean indicates a more difficult
item.
Difficulty Level Interpretation:
0.00-0.20: Very Difficult / Unacceptable
0.21-0.40: Difficult / Acceptable
0.41-0.60: Moderate / Highly Acceptable
0.61-0.80: Easy / Acceptable
0.81-1.00: Very Easy / Unacceptable
Item Discrimination: Measures how well an item distinguishes between knowledgeable and
less knowledgeable examinees. It indicates how well the item relates to the underlying trait.
Range: -1.00 to +1.00.
An index closer to +1 indicates more effective discrimination.
An acceptable index is 0.30 and above.
N u−N l
Formula (using upper and lower groups): Discrimination Index = N /2
N u: Number of students from the upper group who answered correctly.
N l: Number of students from the lower group who answered correctly.
N : Total number of examinees.
Discrimination Index Interpretation:
0.40-above: Very Good Item / Highly Acceptable
0.30-0.39: Good Item / Acceptable
0.20-0.29: Reasonably Good Item / For Revision
0.10-0.19: Difficult Item / Unacceptable
Below 0.19: Very Difficult Item / Unacceptable
Item Reliability Index: The higher this index, the greater the test's internal consistency.
Item Validity Index: The higher this index, the greater the test's criterion-related validity.
Distracter Analysis: Incorrect options (distractors) should be equally plausible and preferably
selected by a greater proportion of lower-scoring examinees than top-scoring ones.
Overall Evaluation of Test Items: Considers both difficulty level and discriminative power.
Book-Mark Method for Setting Cut Scores
The Book-Mark method is a technique for establishing cut scores (pass/fail points) on a test.
Process:
1. Train experts on the minimal knowledge, skills, or abilities (KSAs) required to "pass."
2. Present experts with a book of items, one per page, ordered by difficulty (easiest to hardest).
7/8
3. Experts place a bookmark to divide test-takers who possess the minimal KSAs from those who do not.
Bookmarks: These serve as the cut scores, decided upon by the test developers.
Potential Problems:
Training of experts.
Possible floor and ceiling effects (where scores are too low or too high to differentiate
meaningfully).
Determining the optimal length of item booklets.
Other Methods for Setting Cut Scores
Predictive Method: This approach considers:
The number of positions to be filled.
Projections regarding the likelihood of offer acceptance.
The distribution of applicant scores.
Discriminant Analysis: A family of statistical techniques used to understand the relationship
between variables, often applied to differentiate between groups (e.g., successful vs. unsuccessful job
performers) based on test scores.
Psychological Report Writing
The ability to write clear, accurate, and useful psychological reports is a key skill for test users. Reports
should communicate test findings effectively to relevant audiences, considering various factors that can
impact interpretation.
8/8