0% found this document useful (0 votes)
8 views60 pages

Test Construction Principles Explained

Uploaded by

abeidmsanganzila
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPT, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
8 views60 pages

Test Construction Principles Explained

Uploaded by

abeidmsanganzila
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPT, PDF, TXT or read online on Scribd

Module 2: Principles of Test

construction
Expected Learning outcomes:
•To determine the purpose of Testing
•To establish Consequences of Testing or
not Testing
•To explain the Characteristics of a good
Test
•To describe the Classification of tests and
Construction of test items
What

is a test?
A test is a particular assessment that
typically consists of set of questions
administered during a fixed period of time
under reasonably comparable conditions for
all students.
 Because a test is a form of assessment, tests
also answer the question: “How well does the
individual perform either in comparison with
others (NRT) or in comparison with a domain
of performance tasks (CRT)”.
Purposes of Testing
 Testing shapes what students learn and
how they learn.
 It may serve the following functions:
 Determine whether students are learning what
we are expecting them to learn.
 Motivate and help students structure their
academic efforts.
 Help the teacher understand how successfully
he is presenting the material.
 Reinforce learning by showing students what
topics or skills they have not yet mastered and
should concentrate on.
Test, quizzes, exams, portfolio e.t.c all these
need to have special qualities in order to judge
them as either good, poor or bad.
1. Cognitive Complexity
Standard
The test questions will focus on appropriate
intellectual activity ranging from simple
recall of facts to problem solving, critical
thinking, and reasoning.
Note:
CC refers to various levels of learning that
can be tested.
Test items should reflect how teaching was
done
5
2. Language Appropriateness
Standard
The language demands will be clear
and appropriate to the assessment
tasks and to students.
Notes
Test questions should reflect the
language that is used in the classroom.
The vocabulary (uncommon usage;
non-literal usage) and the syntax of the
test (atypical parts of speech; complex
structures) may create language
barriers. 6
3. Content Quality
Standard
The test questions will permit students to
demonstrate their knowledge of challenging
and important subject matter.
Notes
Relate to course content that the test will
cover:
test specifications: what skills will be tested?
number of questions? how many topics will be
covered?
What formats will be used to test?
7
4. Meaningfulness
Standard
The test questions will be worth students’
time and students will recognize and
understand their value.
Notes
It is very easy to write items which require
only rote recall but are nonetheless difficult
because they are taken from obscure
passages, e.g., footnotes.

8
[Link]
Test practicability refers to all those organizational
and physical procedures that affect the efficiency of
a test.
Aspects to consider
Clarity in instruction such as time allowed, where
and how to place the responses to the questions
Availability of necessary materials and tools that is
going to be used
Scoring key/marking scheme should be clear

11/26/23 9
[Link]
Assessment should be economical in time and
materials.
It should not take too much of the teachers’
time and students’ time
Should also not take too much paper.

11/26/23 10
7. Transfer and
Generalizability
Standard
Successful performance on the test will allow
valid generalizations about achievement to be
made.
Notes
Well constructed tests-whether they are
objective or performance oriented-allow
teachers to understand what needs to be
taught next.

11
8. Fairness
Standard
Student performance will be measured in a way
that does not give advantage to factors irrelevant
to school learning; scoring schemes will be
similarly equitable.
Notes
Test questions should reflect the objectives of
the unit;
Expectations should be clearly known by the
students;
Each test item should present a clearly
formulated task;
One item should not aide in answering another;
12
9;Reliability
This is the extent to which test produces the
same results when used repeatedly under the
same conditions.
Standard
Answers to test questions will be
consistently trusted to represent what
students know.
Notes
Design test items such that no guessing is possible
without proper studying
Example;
If you create a quiz to measure students’
ability to solve quadratic equations, you
should be able to assume that if student gets
an item correct, he or she will also get other
similar items correct
Characteristics of Reliability
1. Reliability is consistency of a test scores
2. It is the measure of variable error or measurement error
3. It is the function of a test length
4. It refers, the stability of a test for a certain population
5. It is the temporal stability of a measuring instrument
6. It is the coefficient of stability
7. It is the coefficient of internal consistency
8. It is the self correlation
9. It is reproducibility of the scores
10. Importantly it refers to the accuracy or precision of a
measuring instrument
11. It does not ensure the validity of the test always
NOTE; The value for reliability coefficients ranges
from 0 to 1.0
-Coefficient of 0 means NO reliability and 1.0
Means PERFECT reliability
-BUT reliability never reach 1.0 because of error.
-All test have errors
-If standardized test is above 0.8 =Very good
reliability
-Below 0.5 = would not considered very reliability test
Factors influencing reliability
The reliability of test scores is influenced
mainly by two factors: extrinsic and
intrinsic.
Extrinsic factors are those factors which lie
outside the test itself and tend to make the test
reliable or unreliable. For example, variability
in the range of ability of a group,
environmental conditions, guessing by the
examinee etc.
Intrinsic factors refer to those factors which
lie within the test itself and influence the
reliability of the test. For example,
characteristics of items, total score, length of
Cont…..
Important extrinsic factors affecting the
reliability of a test may be enumerated as
follows:
• Group variability: When the group of
examinees being tested is homogeneous in
ability, the reliability of the test scores is likely
to be lowered. But when the examinees vary
widely in their range of ability, that is, the group
of examinees is heterogeneous one, the
reliability of the test scores is likely to be high.

• Scorer reliability: By scorer reliability is meant


how closely two or more scorers agree in scoring or
rating the same set of responses. If they do not
Cont…
2. Guessing by the examinees:
Guessing in a test is an important source of
unreliability. In two-alternative response
options there is a 50% chance of answering
the items correctly on the basis of the guess.
In multiple–choice items the chances of
getting the answer correct purely by guessing
are reduced.
- Guessing has two important effects upon the
total test scores.
First, it tends to raise the total score and
thereby makes the reliability coefficient
spuriously high.
Second, guessing contributes to the
Cont….
[Link] conditions: As far as
possible, the testing environment should
be uniform. Arrangement should be such
that light, sound and other comforts are
equal and uniform to all the examinees,
otherwise it will tend to lower the
reliability of the test scores.
4. Momentary fluctuations in the
examinee - influence the test score. For
example, A broken pencil, momentary
distraction by the sudden sound of an
aeroplane flying above, anxiety regarding
non-completion of home work, mistake in
Intrinsic factors
Cont…
Example-1: Suppose an intelligence test of
100 items has a reliability coefficient of 0.80.
If the test is increased four times its present
length, that is, 300 more items are added so
that now the test becomes 400 items then
how much would be reliability index of a test.
Cont…
Example-1: Suppose an intelligence test of
100 items has a reliability coefficient of 0.80.
If the test is increased four times its present
length, that is, 300 more items are added so
that now the test becomes 400 items then
how much would be reliability index of a test.
Ґnn=(4)(.80)/1+(4-1)(.80)
=3.2/1+3x.80
=1+2.4 = 3.2/3.4 = .94
Cont…..
Example-2
Suppose the reliability of an intelligence test
is .60. For how much time should the test be
lengthened in order to reach a reliability
coefficient of .90
Cont…..
Example-2
Suppose the reliability of an intelligence test
is .60. For how much time should the test be
lengthened in order to reach a reliability
coefficient of .90
n= Ґnn(1- Ґtt)
Ґtt(1- Ґnn)
Where n=number of time the test is to be
lengthened; Ґnn=level of reliability coefficient
required; Ґtt=reliability of the existing test
n=.90(1-.60)
.60(1-.90)
=.90 x .40 = .36 = 6
Cont..
Range of the total scores: If the obtained
total scores on the test are very close to each
other, that is, if there is lesser variability
among them, the reliability of the test is
lowered.
On the other hand, if the total scores on the
test vary widely, the reliability of the test is
increased.
In statistically it can be said that when the
standard deviation of the total score is high,
the reliability is also high and vice versa.
Cont…
Homogeneity of items:
The concept of homogeneity of items includes
two things-item reliability(inter-item
correlation) and the homogeneity of function
or trait measured from one item to another.
When the items measure different functions
and the inter correlations of items are zero or
near it (that is, when the test is
heterogeneous one), the reliability is zero or
very low.
When all items measure the same function or
trait and when the inter-item correlation is
high, the reliability of the test is also high.
Cont…..
Difficulty value of items:
In general, items having index of difficulty at
o.5 or close to it, yield higher reliability than
items of extreme index of difficulty. In other
words, when items are too easy or too
difficult, the test yields very poor
reliability(because such items do not
contribute to the reliability)-than when items
are of moderate difficulty values.
Cont…
Discrimination value:
When the test is composed of discriminating
items, the item – total test score is likely to be
high and then, the reliability is also likely to
be high. But when items do not discriminate
well between superior and inferior, that is
when items have poor discrimination values,
the item-total correlation is affected and then
it would decrease the reliability of the test.
Cont…
Suggestions for improving reliability:
-the group of examinees should be
heterogeneous, that is, the examinees should
vary widely in their ability or trait being
measured.
-items should be homogeneous
-the items should be of moderate difficulty
value
- items should be of high discriminating index
Cont….
Generally; a test which yields inconsistent
results (poor reliability) is ordinarily not
expected to correlate with some out side
independent criterion. In other words, a
test which has poor reliability is not
expected to yield high validity. Thus,
validity is dependent upon reliability.
This prediction is true for the homogeneous
test only.
If a test is heterogeneous, validity may be
high even without high reliability.
This is because in a heterogeneous test
each part measures an independent
TYPES OF RELIABILITY
(Methods of estimating the consistency of test)
1)Stability (Test retest method)
 Give the same test twice, separated by days, weeks or months
 Reliability ≈correlation between scores at time 1 and 2
 Spearman rank order correlation.
rho≈ρ
ρ=1-6Ƹd2
n(n2-1)

Where d =Square of difference between ranks

n= total number of candidates

Ƹ= sum
Assumptions of the method
1. No. of item in the test should be large,
therefore memory, practice and carry over
will not effect the retest score
2. Innate ability of an individual should
remains constant so the growth the maturity
will not effect the retest scores.
3. The most appropriate and convenient time
gap between the two administrations should
be fortnight, which is considered neither too
short no too long.
Limitations of the method
This method is less accurate than the other
methods.
Memory, practice, carry over effects are
observed while the test is repeated
immediately.
If the interval between tests is long (six
months or more), growth and maturity will
effect the retests scores and tends to lower
down the index of reliability.
There is no agreement among the
psychometricians regarding the time gap
between the test.
The individual’s Physical and mental health,
emotional and motivational conditions do not
2; Equivalent form /alternate form/parallel method

 The teacher construct two tests on the


same content area, testing the same
skills, having the same number of
item.
 Or create two forms of the same test
vary the items slightly.
 Reliability is stated as correlation
between scores of test 1 and test 2
Alternative-forms reliability or
Coefficient of Equivalence cont…
Alternative forms reliability is known by
various names such as the parallel-forms
reliability, equivalent-forms reliability and the
comparable-forms reliability.
It is an improvement over the earlier method
and it is one way of overcoming the problems of
memory, practice, carry over and recall factors.
It is this method which requires that the test be
developed in two forms and it should be
comparable or equivalent.
Cont…..
Two forms of the test are administered on
same sample of subjects on the same day
after a considerable time interval.
Pearson’s method of correlation is used for
calculating the coefficient of correlation
between two sets of scores obtained by
administering the two forms of the test. Such
a coefficient is known as the coefficient of
equivalence.
Assumptions of the method
The number of items in both forms should be
equal.
Two forms of a test should be alike with
reference to: content and type of items, the
range of difficulty and discrimination indexes,
mean and variance of both the forms,time
administration
Limitations of the method
Practice and carry over factors can not be
controlled, second form of the test scores are
generally high.
It is difficult to construct parallel forms of
test and satisfy all the conditions mentioned.
There is no agreement among the
psychometrican about the interval between
the two forms of test
Interval for administration the two forms will
not be more than two weeks.
It is not possible to provide alike situations to
3: Internal consistency /Alpha /Split-half method
This method of reliability is an improvement
over the earlier two methods-coefficient of
stability and coefficient of equivalence, as it
involves both the characteristics of stability and
equivalence.
This method particularly determines the internal
consistency of the test and internal consistency
reliability indicates the homogeneity of the test.
If all the items of the test measure the same
function or trait, the test is said to be a
homogeneous one and its internal consistency
reliability would be high.
3: The Split-half Method cont…
Compare one half of the test to the other half by
using the method such as Kuder-Richardson formula
(KR 20)
Or Obtain scores for the odd-numbered items and
even –numbered items by using Spearman-rank order
correlation for half test and again use Spearman –
brown formula to get full test.
Ґtt = 2 x reliability of half test
1+reliability of half test

i.e. r=2ρ
1+ρ
Split half Cont…..
Specifically in this method, the test is divided into
two equal or nearly equal halves. The common way
of splitting the test is the odd-even method.
However, all odd-numbered items (like 1,3,5,7,9
etc.), constitute one part of the test and all even-
numbered items (like 2,4,6,8,10,12etc.), constitute
another part of the test.
In this way, each examinee receives two scores:
scores on odd-numbered items and scores on even-
numbered items. In this way from single
administration of the single form of the test two sets
of scores are obtained and then Pearson’s method of
correlation or Spearman –Prophecy formula can be
used for calculating the coefficient of correlation
between the two parts of the test.
Assumptions of the method
The test should be divided into two equal or
nearly equal halves.
All the items of the test should measure the
same trait or ability
All the items of the test should be the same
difficulty value
The assumptions of Pearson’s method i.e.
linearity is applied to this method
Limitations of the method
Chance errors may effect scores on the two
halves of the test in the same way, it tends to
make the reliability index too high
A test can be divided into two parts in a number
of ways, so that the reliability coefficient is not a
unique value
This method can not be used in power test and
heterogeneous tests
It is not possible to split the test items in two
equivalent forms, because items of a test
measures the different aspect of the same trait or
ability
4; Scorer reliability /independent judge method

Two examiners independently score


a set of test papers then correlate their
score.
Use spearman rank order correlation
Example; A class of 8 students scored out
of 10 marks. The test had 5 items.
calculate the reliability of a full test.
10; VALIDITY
Refers to the accuracy of an assessment
whether or not it measures what is supposed to
measure
Any measurement device is valid if it test what
it is supposed to test
NOTE;
A test which is valid should necessarily be
reliable, but a reliable test is not necessarily valid
 A good test should be both reliable and valid
Example; If the test is given to measure a learner’s
ability to use the four fundamental operations, the
test should consists of items that ask Students to
add, subtract, multiply and divide
•In broad sense, validity is concerned with
generalizability.
•When a test is a valid one, it means its conclusion
can be generalized in relation to the general
population.
Characteristics of validity of a test
scores
It is one of the most important characteristics of
a measuring instrument.
It is an index of external correlation. The test
scores are correlated with external criterion
scores.
It relates to the purpose or objective of a test
scores.
Validity ensures the reliability of a test. If a test
is valid, it must be reliable.
It is also the function of a test length.
WAYS OF MEASURING VALIDITY RELATED
EVIDENCE (TYPES OF VALIDITY
1;Content/face/logically validity
Used in evaluating achievement test
Or the extent to which the content of the test
matches the instructional objectives
In order to determine the validity use the table of
specification (shows instructional objectives
+topics to be tested)
Content validity cont……
For example, a test designed to measure
knowledge of biology might have good item
validity because all the items indeed deal with
good biological facts but might have poor
sampling validity, that is, all the items may
deal only with vertebrates.
Thus a test with good content validity also
samples the appropriate content area. This
becomes important because we can not
possibly measure each and every aspect of a
certain content area. Therefore the inferences
about performance in the whole content may
not be judged correctly.
Judgment

of content validity
Content validity of a test is examined in two
ways:
(i)by the expert’s judgment, and (ii) by
statistical analysis.
For example, an investigator wants to examine
the content validity of a test on Tanzanian
history. For this purpose, the content or items of
the test will be submitted to a group of subject-
matter experts.
These experts will judge whether or not the
items represent all the important events of
Tanzanian history, whether or not some
additional items should be added for complete
coverage, what should the relative weights of
the items of a particular event be, etc.
Cont..
ii. Statistical analysis: In this techniques,
scores on the two independent test are
correlated and both of which are said to
measure the same thing.
Suppose one wants to know the content
validity of English spelling test. Then the
teacher can correlate the scores on the said
test with another similar English spelling
test. A high correlation coefficient would
provide an index for the content validity.
Cont…
Although a high correlation coefficient can
easily be demonstrated in two sets of scores
obtained from two similar test, it does not
fully guarantee content validity because high
correlation may be due to the fact that both
the tests measure the same incorrect things.
Therefore test developer should specify:
- the area of content explicitly so that all
major portions in equal proportion be
adequately covered by the items.
- content area should be fully defined in clear
words and must include the objects.
- the relevance of contents or items should be
2;criterion

The extent to which scores on the test are in


agreement with the concurrent or predictive.
i.e. Concurrent validity (scores in the present test
correlate with or may predict current performance)
Predictive validity (scores in the present test
help to predict future performance)
Example; Entrance examination
3; Construct validity
The extent to which an assessment corresponds
to other variables as predicted by some rationale or
theory
Or the agreement of the test with a theoretical
construct or trait.
Used to interpret the psychological traits of the
learners such as sociability, honesty, or anxiety

Example; Intelligent quotient (IQ),math's


reasoning, logic etc.
Construct

Validity cont……
Construct validity has also other names such as
factorial validity and trait validity. In construct
validity the meaning of the test is examined in
terms of a construct.

What is construct? A construct is non- observable


trait such as intelligence, anxiety etc. which
explains our behavior.

Anastasi (1968) has defined it as ‘the extent to


which the test may be said to measure a
theoretical construct or trait.

It may be defined as the degree to which the


individual possesses some hypothetical trait or
ability or quality (construct)presumed to be
reflected in the test performance.
STEPS OF CALCULATING CORRELAT
[Link]
Pearson product moment correlation
Coefficient
FACTORS INFLUENCING TEST VALIDITY
[Link] instrument
• Language difficult
• Irrelevance of test items
• Length of the test (many concepts to be tested but
few of them are tested)
• Time limit (too short time for administering)
• Clues and patterns of answers which encourages
guessing e.g. AA BB CCC
2. Administering and scoring a test

3. Pupils responses

[Link] of the group being tested

You might also like