WRITING ITEMS AND
DEVELOPING TESTS
CONSTRUCTION OF TEST
Identify a need
• The first step is the identification of a need that a test may be
able to fulfill. A school system may require an intelligence test
that can be administered to children of various ethnic
backgrounds in a group setting; a literature search may indicate
that what is available doesn’t fit the particular situation.
The role of theory
• Every test that is developed is implicitly or explicitly influenced or
guided by the theory or theories held by the test constructor. The
theory may be very explicit and formal.
• . A theory might also yield some very specific guidelines. For example,
a theory of depression might suggest that depression is a disturbance
in four areas of functioning: self-esteem, social support, disturbances
in sleep, and negative affect.
• The theory may also be less explicit and not well formalized. The test
constructor may, for example, view depression as a troublesome state
composed of negative feelings toward oneself, a reduction in such
activities as eating and talking with friends, and an increase in
negative thoughts and suicide ideation.
Practical choices
• The following questions should be answered:
What format will the items have?
Will they be true-false, multiple choice, 7-point rating scales, etc.? Will
there be a time limit or not?
Will the responses be given on a separate answer sheet?
Will the response sheet be machine scored?
Will my instrument be a quick “screening” instrument or will it give
comprehensive coverage for each life stage?
Will I need to incorporate some mechanism to assess honesty of
response?
Will my instrument be designed for group administration?
Pool of items
• The next step is to develop a table of specifications, much like the
blueprint needed to construct a house.
• This table of specifications may reflect not only my own thinking, but
the theoretical notions present in the literature, other tests that are
available on this topic, and the thinking of colleagues and experts.
• Test companies that develop educational tests such as achievement
batteries often go to great lengths in developing such a table of
specifications by consulting experts, either individually or in group
conferences; the construction of these tests often represent major
efforts of many individuals, at a high cost beyond the reach of any one
person.
Tryouts and Refinement
• The initial pool of items will probably be large and rather unrefined.
Items may be near duplications of each other, perhaps not clearly
written or understood.
• The intent of this step is to refine the pool of items to a smaller but
usable pool. To do this, we might ask colleagues (and/or enemies) to
criticize, the items, or we might administer them to a captive class of
psychology majors to review and identify items that may not be clearly
written.
• This step then, consists of a series of procedures, some requiring
logical analysis, others statistical analysis, that are often repeated
several times, until the initial pool of items has been reduced to
manageable size, and all the evidence indicates that our test is
working the way we wish it to.
Reliability and validity
• Once we have refined our pool of items to manageable size, and
have done the preliminary work of the above steps, we need to
establish that our measuring instrument is reliable, that is,
consistent, and measures what we set out to measure, that is,
the test is valid.
Standardization and norms
• Once we have established that our instrument is both reliable
and valid, we need to standardize the instrument and develop
norms. To standardize means that the administration, time limits,
scoring procedures, and so on are all carefully spelled out so that
no matter who administers the test, the procedure is the same.
Further refinements
• Once a test is made available, either commercially or to other
researchers, it often undergoes refinements and revisions. Well-
known tests such as the StanfordBinet have undergone several
revisions, sometimes quite major and sometimes minor. Sometimes
the changes reflect additional scientific knowledge, and sometimes
societal changes, as in our greater awareness of gender bias in
language.
• One type of revision that often occurs is the development of a short
form of the original test. Typically, a different author takes the original
test, administers it to a group of subjects, and shows by various
statistical procedures that the test can be shortened without any
substantial loss in reliability and validity.
WRITING ITEMS
Guidelines of Writing tests items
• Define clearly the test intending to measure
• Generate an item pool
• Avoid exceptionally long items
• Keep the level of reading difficulty appropriate to the intended
test takers
• Avoid double- barreled items
• Consider mixing positively and negatively worded items
Design Phase
❖Should have clearly defined purpose
❖Should have specific and standard
content
❖Should have set of standard
administration procedures
❖Should have standard scoring procedure
Evaluation Phase
❖Reliable
❖Valid
❖Should have good item statistics
ITEM FORMATS
Dichotomous Format
• Offers two alternatives for each item. Usually a point is
given for the selection of one of the alternatives.
• Advantages: simple, easy to administer, can be scored
quickly, requires absolute judgment.
• Disadvantages: if used in ability tests, it encourages
memorization, test of luck and response style; forced
choice
• Example: True of False
Polytomous/Polychotomous
• Offers more than two alternatives. Typically, a point is given for
the selection of one of the alternatives and no point is given for
selecting any other choice.
• Advantages: easy to administer and score, uses distractors or
incorrect choices, requires absolute judgment, reduces guessing
• Disadvantages: if used in ability tests, it can encourage test of
luck and response style, forced choice
Example: Multiple Choice
Likert Format
• Requires the respondent indicate the degree of agreement with a
particular attitudinal question.
• Instead of asking for a yes/no reply, five alternatives are offered:
strongly agree, agree, neutral, disagree and strongly disagree.
• Positively worded items are normally scored while negatively
worded items are reversely scored.
• Example: questionnaires with choices from: strongly agree-
strongly disagree
Category Format
• Similar to likert format but uses a scale which
holds greater number of choices.
• Example: Rating of a certain item from 1 is the
lowest 10 is the highest
Checklist
• The subject receives a long list of descriptive
statements and indicate whether each one is
characteristics of him/herself or others.
• Example: Check all that are applicable to you
Q-Sort
• Combination of checklist and category format.
Subjects are give statements and asked to
sort them into 9 piles
Example:
Place from 1-9 from least to the most favorite
Spiral Omnibus Format
• Usually used in ability test wherein items
are arranged from easy to difficult.
Example:
Solve the following equations:
. 12 + 7
2. 50 – 9
3. 12 x 15
4. 473 / 11
5. (261 + 90) – 46
6. (837 – 41) x (63 + 37)
ITEM ANALYSIS
Item Analysis
• It is a general term for a set of methods used
to evaluate test items in order to come up with
a cluster of valid and reliable test items.
• The basic method involve in assessment of
item difficulty and item discriminability.
Item diffculty index (p)
• Defined by the number of people who get a particular item
correct.
• If there is a higher proportion who get the item correct, the easier
the item is.
• Formula:
p= Np/N
Np- # of test takers who got the item correct
N- total # of test takers
Interpretation of Item-difficulty index (p)
Item-difficulty index (p) Interpretation
0.81 and above Very easy
0.61 – 0.80 Easy
0.41 – 0.60 Optimum
0.21 – 0.40 Difficult
0.20 and below Very difficult
Item Discriminability Index (d)
• Determines whether the people who have done well on a
particular test items have also done well on the whole test.
• The higher the value of d the better the test.
• ***An item can have a negative or positive discriminating power.
• Formula:
d= Up-Lp: # of test takers who get the item correct
U: total number of test taker in the upper group
Sample Situation of 50 Total Test Takers
• For this calculation, we divide the test takers into three groups
according to their scores on the test as a whole:
• • an upper group consisting of the 27% who make the highest
scores (U)
• • a lower group consisting of the 27% who make the lowest
scores (L)
• • a middle group consisting of the remaining 46%. (M)
Interpretation of Item-discriminability index (d)
Item-difficulty index (p) Interpretation
0.40 and above Very good item
0.30 – 0.39 Good item
0.20 – 0.29 Fair Item
0.09 – 0.19 Poor Item
0.08 and below Very poor item
After Calculation or Solving, next Step..
• Arrange test scores from highest to lowest.
• Get 27% of the papers from the highest scores
and another 27% of the papers from the lowest
scores.
• Record separately the number of times the
correct answer was chosen by the test takers in
each group
WHEN SHOULD A TEST ITEM BE
REJECTED? RETAINED?
MODIFIED OR REVISED?
• A test item can be retained if it’s level of difficulty is easy,
optimum, or difficult and discriminating power is fair to
very good.
• It has to be rejected if it is either very easy or very difficult
and its discriminating power is poor, negative, or zero.
• An item can be modified if its difficulty level is optimum
and its discriminating power is negative.
• Once you have completed getting the raw scores of your sample,
create a table showing the equivalent Z score, percentile and
stanine for each score. Include a qualitative interpretation of the
scores.
STANDARDS FOR TEST
ADMINISTRATION SCORING AND
INTERPRETATION
TEST
ADMINISTRATION
• Psychological testing is a dynamic process influenced by many
factors.
• Although examiners strive to ensure that test results accurately
reflect the traits or capacities being assessed, many extraneous
factors can sway the outcome of psychological testing.
• Invalid test results do not originate only from obvious sources like
non-standardized administration, inexperienced examiner, noisy
testing room, scared examinee or careless scoring
• The interpretation of a psychological test is most reliable when
the measurements are obtained under the standardized
conditions outlined in the test manual.
• Non-standardized testing procedures can alter the meaning of
the test results, rendering them invalid and misleading.
• Even though standardized testing procedures are normally
essential, there are instances in which flexibility in procedures is
desirable or even necessary.
Conditions of Testing
• Physical Condition- Ventilation, Lighting, can affect the
test scores.
• Conditions of the Person- the state of the test taker
• Test Condition- testing materials conditions and spacing of
giving the test
• Condition of the Day- the time the test was given may also
influence the test scores
Common Assumption–examination procedures
are so simple and straightforward that a quick
once-through reading of the test manual will
suffice as preparation for testing.
Test administrators should follow carefully
the standardized procedures for
administration and scoring specified by the
test developer and any instructions from the
test user.
When formal procedures have been
established for requesting and receiving
accommodations, test takers should be
informed of these procedures in advance
of testing.
SENSITIVITY TO DISABILITIES
Impairments in hearing, vision, speech or motor control may
seriously distort test results.
If the examiner does not recognize the physical disability
responsible for the poor test performance, a subject may be
branded as intellectually or emotionally impaired when, in fact,
the essential problem is a sensory or motor disability.
An examiner must modify test procedures that can suit the needs
of people with disabilities or he/she can use alternative
instruments developed exclusively for the target population.
Assumption: any adult can accurately
administer group tests as long as he or
she has the manual.
Administration of a group test requires more accurate
and more rigid procedures than individual assessment.
The standard ratio of examiner to examinee is 1:25 or
less
ERRORS IN GROUP TEST
ADMINISTRATION
1. Incorrect timing of tests that require a time limit.
a. Cutting short the designated time limit and allowing too much
time for a test will make the norms completely invalid.
b. Examiners must allot sufficient time for the entire testing
process: set-up, reading instructions aloud and actual test taking.
2. Lack of clarity in the directions to the examinees.
a. Instructions must not be paraphrased.
b. Examiners must read the instructions slowly in a clear, loud
voice that commands the attention of the subjects
ERRORS IN GROUP TEST
ADMINISTRATION
3. Variations in the physical conditions under which tests are given.
a. Examiners must ensure that the testing room is well iluminated and
extreme variations in temperature and humidity are controlled.
b. The quality of the writing surface can be crucially important and it is
magnified by the current tendency to use separate answer sheets.
c. Loud noises, especially if intermittent and unpredictable, will cause
test scores to be invalid.
4. Failure to explain when and if examinees should guess.
a. Examiners should not give supplementary advice on guessing – this
would constitute a serious deviation from standardized procedures
Changes or disruptions to standardized test
administration procedures or scoring should
be documented and reported to the test
user.
Testing environment should furnish
reasonable comfort with minimal
distractions to avoid construct-
irrelevant variance.
Test takers should be provided
appropriate instructions, practice, and
other support necessary to reduce
construct-irrelevant variance.
Reasonable efforts should be made to
ensure the integrity of test scores by
eliminating opportunities for test takers
to attain scores by fraudulent or
deceptivemeans.
Test users have the responsibility of
protecting the security of test materials at
all times.
INFLUENCE OF EXAMINER
Importance of Rapport
• Test publishers urge examiners to establish rapport – a
comfortable, warm atmosphere in which serves to motivate
examinees and elicit motivation.
• A tester who fails to establish rapport may cause a subject to
react with anxiety, passive-aggressive non cooperation, or open
hostility.
• Failure to establish rapport distort test findings – ability is
underestimated and personality is misjudged
Expectancy Effect / Rosenthal Effect
• The tendency for results to be influenced by what the test
administrators expect to find.
• In this phenomenon, the test administrator have communicated
the expected results to the examinees.
• The greater the expectation placed on people, the better they
will
Effect of Reinforcing Responses
• The use of reinforcements can damage the reliability and validity
of test scores.
• Reinforcement and feedback guide the examinee toward a
preferred response.
• Random reinforcement destroys the accuracy of performance
and decreases the motivation to respond
BACKGROUND AND MOTIVATION OF THE
EXAMINEE
• Examinees differ not only in the characteristics which examiners
desire to assess, but also in each other extraneous ways that
might confound the test results.
• Test results may be inaccurate due to the filtering and distorting
effects of certain examinee characteristics such as anxiety,
malingering, coaching or cultural background
What to do for Examinee with Test Anxiety
• Test Anxiety- refers to those phenomenological, physiological
and behavioral responses that accompany concern about
possible failure on a test.
• Make sure your instructions are neutral and non-threatening.
Test-anxious subjects show significant decrements in
performance when they perceive the situation as a test.
• If a test is timed, make sure that the timer is out of the
examinee’s view.
Effects of Coaching on Test Results
• Coaching may include several components: extra practice on
test-like materials, review of fundamental concepts likely to be
covered by the test, and advice about optimal test-taking
strategies.
• Coaching can inflate a subject’s score without correspondingly
improving his or her overall abilities or behaviors in the domain
being tested
TEST SCORING
Two ways to score Psychological Tests
• Hand Scoring- used only if there are small
number of answer sheets to be scored.
• Machine Scoring- used in scoring a large
amount of answer sheets in a small amount of
time.
Frequency Distribution
A single test score means more if a psychologist relates it to other
test scores.
More meaning will arise if a single test score is included in a
distribution of scores which summarizes the scores of a group of
individuals.
Frequency distribution – a technique for systematically displaying
scores on a variable or a measure to reflect how frequently each
value was obtained
Percentile and Percentile Ranks
• Percentile and percentile ranks are basically similar. However,
percentiles indicate the particular score below which a defined
percentage of scores fall.
• Percentile rank of a score is the percentage of scores in its frequency
distribution that are the same or lower than it.
• It answers the question “What percent of the scores fall below a
particular score?”.
• Formula for Simple Frequency Distribution: Pr = B/ N x 100
• Pr = Percentile Rank
• B = the number of scores/cases below the score of interest
• N= the total number of scores
Measure of Central Tendency
• Mean- This refers to the sum of all the given values or items in a
distribution divided by number of values or items summed.
• Median - This refers to the point in a distribution that divides the
group into 2 parts so that 50% fall below and another 50% fall
above that point.
• Steps and Formula for Ungrouped data:
• 1. Arrange the data in increasing/ascending order. 2. Let n
denote the number of pieces of data and locate the median using
the formula: (n + 1) / 2 3. The value obtained from the formula
points to the ordinal position of the median.
Measure of Central Tendency
• Mode-This refers to the value or item in a distribution with the
most number of cases or highest frequency.
This can be:
• Unimodal (e.g. 20,18, 18, 18, 17, 16, 11)
• Bimodal (e.g. 4, 4, 7, 11, 6, 5, 8, 5, 2)
• Inexistent (e.g. 9, 2, 15, 4, 6, 1, 5, 13)
Standard Deviation (sd)
• This is a measure of variability which indicates
how far, on the average, the data values are
from the mean.
• This basically shows the distance of the
scores from one another using the mean as a
reference point.
Raw Scores, Standard Scores and Norms
• Raw Score- measure of performance that is given directly by
scoring according to the procedure of scoring for the test.
• Standard Scores- it indicates the precise location of any score in
a given distribution.
a. Z score- standardized unit of a given score or data that is
much easier to interpret.
b. T score- it is a converted score from Z score. It indicates
the exact location of a score within the distribution with a mean of
50.
c. Stanine score- system that converts scores into a
transformed scale that ranges from 1to 9.
Z score / Standard score
• A statistical procedure which transforms data into
standardized units that are easier to interpret.
• The Z score is simply the number of standard
deviations between the mean and the raw score
Converting Z score to Percentile
• If the Z score is positive, the converted value will be added to
0.50, The sum will multiplied by 100, resulting to the equivalent
percentile.
• If the Z score is negative, the converted value will be subtracted
from 0.50, resulting to the equivalent value. The difference will be
multiplied by 100, resulting to the equivalent percentile.
Example
• Example:
• A score of +1.67 has a converted value of 0.4525
• The converted value will be added to 0.50 0.50 + 0.4525 =
0.9525
• The sum will be multiplied to 100, resulting to the equivalent
percentile:
(0.9525) (100) = 95.25th percentile
Example
• Example:
• A score of -0.33 has a converted value of 0.1293
• The converted value will be subtracted from 0.50 0.50 - 0.1293 =
0.3707
• The difference will be multiplied to 100, resulting to the equivalent
percentile:
(0.3707) (100) = 37.07th percentile
T Scores / McCall’s T
• A system of transforming raw scores created by W.A.
McCallin with the purpose of increasing their interpretive
value.
• This is basically the same as Z scores except that the
mean in McCall’s system is 50 rather than 0 and the
standard deviation is 10 rather than one.
• Formula: T = 10Z + 50
Quartiles and Deciles
• Quartiles – divides the frequency distribution into 4 equal parts:
Q1 , Q2 , Q3 and Q4 .
• Deciles – divides the frequency distribution into 10 equal parts:
D1 to D10.
Stanine System
A system which
converts any set of
scores into a
transformed scale,
which ranges from 1
to 9. It has a mean of
5 and a standard
deviation of
approximately 2.
Norms
Norms – refers to the performances by defined groups on
particular tests. It serves as a reference when interpreting or
evaluating a test score.
The norms for a test are based on the distribution of scores
obtained from a defined sample of individuals.
Norms are obtained by administering the test to a sample of
people and obtaining the distribution of scores for that group.
Norms
• Normative sample – is the group of people whose performance
on particular test is analyzed for reference in evaluating the
performance of individual test takers.
• Norming – refers to process of deriving norms; may be modified
to describe a particular type of norm derivation.
• Standardization – refers to the process of administering a test to
a representative sample of test takers for the purpose of
establishing norms
Sampling to Develop Norms
• Sampling - the process of selecting individuals for a study.
Sampling methods fall into two major categories:
• Probability sampling – the entire population is known, each
individual in the population has a specifiable probability of
selection, and sampling occurs by a random process based on
probabilities.
• Non-probability sampling – the population is not completely
known, individual probabilities cannot be known and the sampling
method is based on common sense or ease but still maintains
representativeness and avoids bias.
Age-Related Norms
• Certain tests have different normative groups for particular age
groups. Most IQ tests are of this sort.
• There are times that you cannot use the norms of a particular
age group on a younger or older age group. Therefore, you have
to use age-related norms.
• The purpose of establishing norms for a test is to determine how
a test taker compares with others.
Norm-referenced tests vs. Criterion-referenced
tests
• Norm-referenced test
A type of test which compares each person with norm.
Results are used to rank people according to performance.
• Criterion-referenced test
A type of test which describes the specific types of skills, tasks or
knowledge that the test taker can demonstrate.
Results are not used to make comparisons among test takers but
to diagnose, document and identify problems that need
remediation.
1) Those responsible for test scoring should
establish scoring protocols. Test scoring that
involves human judgment should include rubrics,
procedures, and criteria for scoring. When scoring
of complex responses is done by computer, the
accuracy of the algorithm and processes should be
documented.
2)Those responsible for test scoring should
establish and document quality control
processes and criteria. Adequate training should
be provided. The quality of scoring should be
monitored and documented.
ISSUES IN SCORING
Clerical scoring errors are the most commonly committed mistakes
in the scoring of psychological tests and they cause test scores to
be wildly inaccurate:
• Using the wrong template to count the scores.
• Not double checking the score count.
• Adding columns of scores incorrectly.
• Consulting the wrong reference table.
• Subjective scoring of test answers.
INTERPRETATION
Level 1- Descriptive
• Data are primarily treated in a sampling
or correlate way
• Minimal concern with intervening
processes
• NO concern on underlying constructs
Level 2
•Constructs descriptive
generalization
•Identification of a Hypothetical
Construct
Level 3- Diagnosis
• There is an attempt to develop a theory
about the individual’s life, school
performance or working image. There is
an exploration of the individual’s
personality and psychological
developmental history and psychosocial
situation.
1) When test score information is released,
those responsible for testing programs
should provide interpretations appropriate
to the audience.
2)When automatically generated interpretations of test
response protocols or test performance are reported, the
sources, rationale, and empirical basis for these
interpretations should be available, and their limitations
should be described.
5)When group-level information is obtained by aggregating the
results o f partial tests taken by individuals, evidence o f validity and
reliability/precision should be reported for the level o f aggregation at
which results are
reported. Scores should not be reported for individuals without
appropriate evidence to support the interpretations for intended
uses.
6)When a material error is found in test scores or other
important information issued by a testing organization
or other institution, this information and a corrected
score report should be distributed as soon as practicable
to all known recipients who might otherwise use the
erroneous scores as a basis for decision
making.
7) Organizations that maintain individually
identifiable test score information should
develop a clear set o f policy guidelines on the
duration of retention o f an individual’s
records and on the availability and use over time of
such data for researchor other purposes.
8)When individual test data are
retained, both the test protocol and
any written report should also be
preserved in some form.
9)Transmission of individually identified test scores to
authorized individuals or institutions should be done in
a manner that protects the confidential nature o f the
scores and pertinent ancillary information.