0% found this document useful (0 votes)
29 views7 pages

Test Development: Concepts and Methods

The document outlines the process of test development, including conceptualization, construction, and revision. It discusses various testing formats, scaling methods, item analysis, and the importance of validity and reliability in test items. Additionally, it covers qualitative analysis techniques and the need for cross-validation and co-validation to ensure the accuracy of tests across different populations.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
29 views7 pages

Test Development: Concepts and Methods

The document outlines the process of test development, including conceptualization, construction, and revision. It discusses various testing formats, scaling methods, item analysis, and the importance of validity and reliability in test items. Additionally, it covers qualitative analysis techniques and the need for cross-validation and co-validation to ensure the accuracy of tests across different populations.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Week 6

Test Development
1-) Test Conceptualization

Begins with the idea: “A test is needed to measure something."

When you make a test, there are a lot of questions but you can guess what they are when
you see on exam. Ones I thought one could confuse were:

-What is the ideal format of the test? ⇒ True-false, essay, multiple choice
-Should more than one form of the test be developed? ⇒ Should there be alternate/parallel
forms of the test?
-How will it be administered? ⇒ Individually/in groups? With computer?
-How will meaning be attributed to scores on this test? ⇒ Will a test score be compared to
others taking the test at the same time? Or to a criterion group?

Pilot work: It’s a personal, not official prototype of a test. Helps the developer see how to
measure the construct effectively. Provides early insights before creating the formal test.
For example: Interviews with people (and those who know them) to measure
extraversion/introversion.

2-) Test Construction

Scaling: Designing and calibrating a test, assigning numbers to levels of a trait or


characteristic.
Types of scales: Nominal, Ordinal, Interval, Ratio (NOIR).

Rating scale: Uses words, statements, or symbols for respondents to indicate strength of a
trait, attitude, or emotion.
For example, Morally Debatable Behaviors Scale (MDBS) – 30 item scale

Paired compression: Test takers are presented with pairs of stimuli and asked to select one.
For example: E.g., Select the behavior you think would be more justified:
1-) Cheating on taxes if one has a chance
2-) Accepting a bribe in the course of one’s duties

Sorting tasks
1-) Comparative scaling: Rank each stimulus against all others (e.g., sort MDBS cards from

1/7
Week 6

most to least justifiable).

2-) Categorical scaling: Place stimuli into predefined categories along a continuum (e.g.,
never, sometimes, always justified).

Guttman Scale: Items increase sequentially in intensity. Agreement with stronger items
implies agreement with milder items.
Example: Do you agree or disagree with each of the following:
A-) All people should have the right to decide whether they wish to end their lives.
B-) People who are terminally ill and in pain should have the option to have a doctor assist
them in ending their lives.
C-) People should have the option to sign away the use of artificial life-support equipment
before they become seriously ill.
D-)People have the right to a comfortable life.

-Developed by administering items to a target group.

"But then how do we turn it into Guttman Scale?"

2/7
Week 6

Item Pool: The reservoir from which items will or will not be drawn for the final version of the
test (Oyunlardaki item havuzu gibi düşünün).

Item Format: The form, plan, structure, arrangement, and layout of individual test items
-Selected-response format: Requires test takers to select a response from a set of
responses
-Multiple-choice format: Stem (Question), a correct option and several incorrect options
(distractors).

-Matching item: Test taker is presented with two columns - premises on the left, responses
on the right – Which response is best associated with best premise.

-Binary-choice item: True/False

-Constructed-response format: Test takers create the answers themselves.


-Completion item: Provide a word or phrase that completes a sentence (fill-in-the-blank)
-Short-answer item
-Essay: Write extended responses

-Computer Administration: Stores items in an item bank and allows individualized testing.
-Computerized Adaptive Testing (CAT): Items adjust based on previous responses.
Advantages of CAT: Reduces floor and ceiling effects.
-Floor: Low performers are not well distinguished.
-Ceiling: High performers are not well distinguished.
Item branching: Adjusts content and order based on responses.

Test Tryout

-Test group must be of similar people


-Sample size must be large enough to reduce the influence of random chance in the results.
-Environment: Environment conditions must be identical to the conditions used for the final,
standardized test.

What is a good item?

A good item must:


-Be valid and reliable (obviously)
-Its main job is to discriminate (tell the difference) between test takers.
-Must be answered correctly by high scorers and incorrectly by low scorers

Item Analysis
1-) Item-Difficulty Index

Goal: To select the best items for a final test


Tool: The Item-Difficulty Index is used to evaluate how easy or hard each item is.

3/7
Week 6

Key Rule (for a 'good' item): If everyone answers an item correctly or everyone answers it
incorrectly, that item is not good because it fails to discriminate between test takers.

Exception: A Giveaway Item is an exception. It's a very easy item placed at the beginning of
the test specifically to:
-Increase motivation.
-Decrease anxiety.

Index: The proportion of the total number of test takers who answered the item correctly
T For example: If 50 out of 100 examinees answered item 2 correctly then the index would
be .5 (p2 = .5)

For the whole test take average of each index.

Optimum difficulty: Average .5, range .3 to .8 For true-false items .75 (taking the role of
guessing and chance into account)

2-) Item-Reliability Index

Indicates how much an item contributes to the test’s internal consistency.


Calculated as: item’s SD × correlation between item score and total test score.
Used with factor analysis/inter-item correlations to identify weak items.
Items that don’t load well on the intended factor are revised or removed.

3-) Item-Validity Index

Internal consistency of a test using standard deviation and the correlation between the item
score and the total test score.

Formula:

s1 ⇒ Standard deviation of item's scores.


p1 ⇒ The Proportion of test takers who answered the item correctly.
(1-p1) ⇒ The Proportion of test takers who answered the item incorrectly.

4-) Item-Discrimination Index

Shows how well an item separates high scorers from low scorers on the overall test.
Good item: High scorers answer correctly; low scorers don’t.
Compares item performance between top 27% and bottom 27% of test scorers.
Symbol: d
Higher d = Better discrimination
Negative d = Serious problem (Red flag yazmış hoca çıkabilir bile)
4/7
Week 6

Item ⇒ Kaçıncı soru


U ⇒ Upper Group (top scores, answered correctly)
L ⇒ Lower Group (bottom scores, answered incorrectly)
U - L ⇒ The Difference in correct answers between the Upper and Lower groups.
n ⇒ Kişi sayısı
d ⇒ The Item-Discrimination Index. This is the difference in proportions of correct answers
between the two groups.

Other Considerations

-Item Fairness: An item is unfair if it advantages one group over another after controlling for
true ability.
ICCs (Item Characteristic Curves) help detect bias by comparing groups (e.g., men vs.
women).
-Speed tests can be misleading because many low scorers simply did not reach the items,
not because they got them wrong. Therefore, conduct item analysis with ample time, but set
norms using the actual intended time limits.

Qualitative Item Analysis:


-Uses verbal, non-statistical methods to evaluate items. Involves asking test-takers
(individually or in groups) to describe their test-taking experience.
-Helps reveal how items are interpreted, whether wording is confusing, and how items
function beyond what statistics can show.

Not: Aşağıdaki uzun gibi gözükebilir 3-5 tanesini okusanız anlarsınız zaten.

5/7
Week 6

Qualitative Item Analysis Techniques:


-Think-aloud: Respondents verbalize thoughts while answering, revealing their cognitive
processes.
-Expert panels: Experts review items and provide qualitative feedback.
Sensitivity review: Items checked for fairness and to avoid offensive language, stereotypes,
or biased situations.

Test Revision

-Remove or rewrite items based on item analysis.


-Assess each item’s strengths and weaknesses.
-Balance items according to the test’s purpose (e.g., high discrimination if identifying top
performers).
-After revision, conduct a test tryout with the new version.
-Update tests that have aged poorly:
6/7
Week 6

Dated stimulus materials or vocabulary


Inappropriate/offensive words
Outdated norms
-Revision can improve reliability and validity.
-Incorporate updated theory.
-Follows the same development stages as a new test.

Cross-validation and Co-validation

Cross-validation: Re-validating a test on a new sample different from the original validation
group.
For example: A new depression inventory was validated on college students. To cross-
validate, researchers test it on a different group, like community adults, to see if it still
predicts depression accurately.

Co-validation: Validating two or more tests simultaneously using the same sample.
For example: A study wants to validate both a new anxiety scale and a new stress scale.
Both tests are given to the same group of participants, and their relationships to relevant
outcomes (like physiological stress markers) are analyzed simultaneously.

7/7

You might also like