0% found this document useful (0 votes)
91 views34 pages

Test Administration and Scoring Guide

The document provides guidance on test administration, scoring, and interpretation. It discusses: 1) The best time to administer a test is at the beginning of a class period to ensure enough time and focus from students. 2) Elements of good test administration include a favorable testing environment, clear written directions, and minimizing distractions and test anxiety. 3) Scoring involves correcting tests to assign raw scores according to marking schemes. Objective tests are easy to score but may require correction formulas to account for guessing. Essay questions require analytical scoring against an outline of expected answers.

Uploaded by

marubegeoffrey41
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
91 views34 pages

Test Administration and Scoring Guide

The document provides guidance on test administration, scoring, and interpretation. It discusses: 1) The best time to administer a test is at the beginning of a class period to ensure enough time and focus from students. 2) Elements of good test administration include a favorable testing environment, clear written directions, and minimizing distractions and test anxiety. 3) Scoring involves correcting tests to assign raw scores according to marking schemes. Objective tests are easy to score but may require correction formulas to account for guessing. Essay questions require analytical scoring against an outline of expected answers.

Uploaded by

marubegeoffrey41
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

LECTURE 10

TEST ADMINISTRATION, SCORING AND INTERPRETATION

Introduction

The effort through this Module is to show what test is, why test is
important, how tests are constructed and what precautions are taken to
ensure validity of tests. The module will round off by explaining how
tests are scored and interpreted. In order to enjoy the study of the unit,
you should have other units by your side and cross-check aspects
relevant to this unit that were discussed in the previous units.

OBJECTIVES

By the end of this unit, you should be able to: Administer a test ,

a. score and interpret tests in general and continuous assessment in parti


cular;

b. analyse test items;

c. compute some measures of general tendency and variability; and

d. compute Z – score and the Percentile.

Test Administration

Just as constructing a well- written test is crucial, administering it properly is also


important. The correct/or good administration of tests is basic to obtaining
good results.

What is the best time to give tests?


- The best time to give a test is at the beginning of a class period. This will
ensure that there are enough time available for all students to complete the
test. It is also a favorable time because these students have moved above
and relaxed during the interval between classes, placing them in the best
situation to begin concentrating.
- A number of problems can be created when a test is administered toward
the end of the period. Students are to be restless and nervous in anticipation
of taking the test and are not concentrating during the period of time before
the test is administered.

Elements/principles of good test administration

1. Time of administration.
2. Provide for favorable testing environment.
- Give the test in a room away from direct noise.
- Give the test in familiar surroundings.
- Provide for proper light, heat and ventilation.
- Arrange the darks and/or tables and chairs properly with enough space.
- This will avoid/discourage cheating.
- Plan to have all distractions eliminated i.e. noise from the next
classroom, games in the field. (put a sign in the door :testing do not
disturb)

3. Written directions should be clear enough to make the test self-


administering, but in some situations it may be desirable to give the
directions orally as well. With young students a blackboard illustration may
also be useful.
4. Eliminate any psychological influences

- Students will not perform either best if they are tense and anxious
during testing. Some of the things that create excessive test anxiety
are:-
i Threatening students with tests if they do not behave.
ii Warning them to do items to their level best because this test is
important.

iii Telling pupils they must work fast in order to complete the test on
time.

iv Threatening dire consequences if they fail the test.

5. Do not talk unnecessarily before the test.


- Do not waste/or deprive students of their testing time. Before a test,
students are mentally set for the test and win ignore anything that
pertaining to the test for fear in it will hinder their recall of information
needed to answer the question.

- Avoid giving hints to pupils who ask about individual items/questions.


If a question is ambiguous or some words are spelled incorrectly, it
should be clarified for the entire group. If it is not ambiguous refrain
from helping the pupil answer it.

6. Keep interruptions to the minimum during the test. At times a student will
ask to have an ambiguous question clarified. This should be done for the
entire group. Such one time interruptions are necessary but should be
kept to a minimum.

Scoring of Tests

This section introduces to you the pattern of scoring of tests, be they


continuous assessment tests or other forms of tests.

Scoring means the process of correcting assignments, test, projects, term papers
or giving so many points for a performance or a product. The scores are usually
in raw score form.

Scoring is an important part of the testing process.


In scoring, it is assumed that the teacher has stated objectives for teaching and
learning (step 1 in teaching process)

- Has provided the best opportunities for the students to grow toward the
objective and:
- Is now ready to evaluate student progress toward the goals of instruction.

NB. The measuring instrument which is used must be the best one to measure
growth toward the objective. Some examples of measuring instruments and
how they are usually scored follows.

The following guidelines are suggested for scoring of tests:

i. You must remember that multiple choice tests are difficult to design,
difficult toadminister, especially in a large class, but easy to score. In
some cases, they are scored by machines. The reasons for easy
scorability of multiple-choice tests are because they usually have one
correct answer which must be accepted across the board.
ii. Essay or subject types of tests are relatively easy to set and administe
r, especially in alarge class. They are, however, difficult to mark or
assess. The reason is because easy questions require a lot of writing
of sentences and paragraphs. The examiner must read all these.
iii. Whether an objective or subjective tests, all tests must have marking
[Link] schemes are the guide for marking any test. They
consist the points, demands and issues that must be raised before
the candidate can be said to have responded satisfactorily to the
test. Marking schemes should be drawn before testing not after the
test has been taken. All marking schemes should carry mark
allocation. They should also indicate scoring points and how the
scores are totaled up to represent the total score for the question or
the test.
iv. Scoring or marking on impression is dangerous. Some students are v
ery good atimpressing examiners with flowery language without real
academic substance. If you mark on impression, you may be carried
away by the language and not the relevant acts. Again, mood may
change impression; your impression can be changed by joy, sadness,
tiredness, time of the day and son on. That is why you must always
insist on a comprehensive marking scheme.
v. Scoring can be done question-by-
question or all questions at a time. The best way isto score or mark
one question across the board for all students. Sometimes, this may
be feasible and tedious, especially in a large class.
vi. Scores can be interpreted into grades, A, B, C, D, E and F. They may b
e interpreted in terms
of percentages: 10%, 20%, 50% etc. Scores may be presented in acom
parative way in terms of 1st position, 2nd position, and 3rd position to
the last. Scores can be coded in what is called BAND. In band system,
certain criteria are used to determine those who will be in Excellent,
Very Good categories, etc. An example of a band system is the one
given by the International English Testing Services (IETS) and the one
by Teaching English as a Foreign Language (TOEFL)test.

Scoring of Objective Tests

- Usually are points for each correct item unless the item contains more than
one response –e.g. matching tests. Sometimes a guessing correction is
applied the usual formula is:

Score = rights – wrongs.

(n-1)

n=meaning the number of scores in the recognition item. Thus, true- false item
would be rights – wrongs. In a four – choice multiple – choice, it would be rights
–wrongs or the like.

Scoring Correction Formula

As stated in the main paper, objective test is very easy to score. All other
advantages of objective test are well known to all by now. However, the
chances of guessing the correct answer are high. To discourage
guessing, some objective tests give instructions to candidates that they
may be penalized for guessing. In such situation, the correction
formular is applied after scoring. This is given as:

-No. of questions marked right (R)

- No of questions marked wrong (W)

-No. of options per item (N) –

If in an objective test of 50 questions where guessing is prohibited, a


candidate attempted all questions and gets 40 of them correctly, then
the actual score after correction is.

S = 40 = (assuming the options per item is 5.)= 40 = 2.5 = 37.5 = 38 out


of [Link] the corrected scores of two candidates. A and B who both
scored 35 in an objective test of 50, if a attempted 38 questions while B
attempted all the questions. S

= 35 - ¾ = 34 and S

= 35 – 15/4 = 31) Note that under rights only, each of the students gets
35 out of 50.

2. Essay tests.

- The examiner usually decides the number of points an essay question is


worth by the number and quality of ideas which are expected for the most
complete response. It might be one point for each idea or, to give some
opportunity to be more discriminating in quality of response, each idea
might be worth 0, 1, or 2 points. The points are then added to obtain a scale
for the item.
(See suggestions for scoring an essay test)

Suggestions for scoring Essay questions

(1) Prepare an outline of the expected answer in advance.

Preparing a scoring key / or marking scheme provides a common basis for


evaluating the pupils’ answers and increases the likelihood that our standards
for each question will remain feasible.

There are two ways of making

- Impressionistic/rating/indistinct approach.
- A situation where you are required to read through the paper and grade the
paper from the impression you get
- Analytical approach
- Looking at number of points written and give accordingly. It is the best
method
Its advantages is that

- It is very reliable – gives reliable score throughout the scoring.


(2) Decide how to handle factors that are irrelevant to the learning outcomes
being measured.

- Several factors influence our evaluations of answers to essay questions that


are not directly pertinent to the purposes of the measurement. Prominent
among these are legibility of handwriting, spelling, sentence structure, neatness
etc.

(3) Evaluate all the answers to one question before going on to the next one.

- One factor that contributes to unreliable scoring of essay questions is a shifting


of standards from one paper to the next. A paper of which average answers may
appear to be of much quality when it follows a failing paper than when it follows
one with near- perfect answers.

(4) Evaluate the answers without looking at the pupil’s /students name.

The general impression we form about each pupil during our teaching is also a
source of bias in evaluating essay questions.

(5) If especially important decisions are to be based on the results, have more
independent ratings.e.g.

- If it is for scholarships, special training etc.

Interpreting Test Results

After a test has been given and scored, the task that remains is to interpret the
score /results. Methods of interpreting results included:

(1) The simple raw score


- test marks which are obtained by adding up the marks for each question or
part of a question are referred to as raw scores. The word raw is used here
to imply that these marks have not been treated or modified in any way.
e.g. – marking a question out of 20

13,8,16,9,12, 11, 9, 11, 10 etc.

(2) The percentage score

This is determined by collecting what % of the total possible mark


(maximum mark) each candidate’s mark is. In the above example –
maximum Mark = 20

If the highest mark is 16 the % becomes

16 X 100 = 80%

20
(3) The simple rank.

When the pupils’ achievement scores are put in order of merit, a simple
rank results. The higher mark is given a rank of 1,the next 2 etc. If no two
pupils have the same mark, there is no question in assessing ranks from raw
scores or % scores. If however, there are cases where many pupils earn the
same marks, it is more difficult to assign ranks.

- This could be done also by calculating the rank difference correlation.


The simple rank tells a pupil’s relative position in a class.

- In ranks, the highest ranks show excellence and the bottom ranks – poor
performance. But the ranks in between are difficult to interpret for both the
parents and majority of teachers. It becomes more confusing if the school
report issues the pupils ranks in all school subjects. How can parents tell
which was his best subject in terms of class performance?
(4) The percentile rank

- A percentile rank of any mark indicate what % of the total number of


candidate have earned marks below that mark. Suppose a student scored 13
in a twenty- item word memory test, his simple rank is 7th. If his parents can
be told what % of students in his /her class have performed below the mark,
they can visualize more easily their child’s attainments in relation to the rest
of the pupils in his class.

(5) Pass mark

- The teacher and examiner are consistently being called upon to establish a
boundary between success and failure.

Item Analysis

Item analysis helps to decide whether a test is good or poor in two ways:

i. It gives information about the difficulty level of a question.


ii. It indicates how well each question shows the difference (discri
minate) between thebright and dull students. In essence, item
analysis is used for reviewing and refining a test.
These are the process of determining the student’s answers to each item
/question to asses the quality of that item. This item analysis is usually designed
to answer questions such as the following.

(1) Did the item function as intended?


(2) Were the test items of appropriate difficulty
(3) Were the test items source of irrelevant clues and other defects?
(4) Was each of the destructors effective (in multiple choice questions)?
Answers to such questions are of obvious value in selecting or revising items
for future use. The benefits of item analysis are not limited to the
improvement of individual test items, however.

Significance / importance of item analysis

(1)Item analysis data provide a basis for efficient class discussion of test
results.

- Knowing how effectively each item functioned in measuring achievement


makes it possible to confine the discussion to those areas most helpful to
pupils.
- Concepts in those items causing the pupils the greatest difficulty can
receive special emphasis while the others can be omitted.
- Item analysis will also expose technical defects in items.
(2)Item –analysis data provide a basis for remedial work.

- Although discussing the test results in class can clarify and correct many
specific points, item analysis frequently brings to light general areas of
weakness requiring more extended attention. In a mathematics test for
example, item analysis may reveal that pupils are proficient in mathematics
skills but are having difficulty with problems requiring the application of
these skills.
(4) Item analysis data provide a basis for the general improvement of
classroom instruction.
- Item analysis data can assist in evaluating appropriateness of the learning
outcomes and the course content for the particular pupils being taught.
Material that is consistently too simple or too difficult for the pupils might
suggest curriculum revisions or shifts in teaching emphasis.
(5) Item – analysis procedures provide a basis for increased skill in item
construction
- Item analysis reveals ambiguities, clues, ineffective distracters, and other
technical defects that were missed during the test’s preparation. This
information is used directly in revising the test items for future use. As we
analyse student’s responses to items.

Procedure of Item Analysis for Norm-referenced tests

1)Mark the papers.

2)Arrange the papers from high side and take 1/3 or 27%of the papers from
high side and take I/3 or 27% of the papers from the low side.

4)For each question, count the number of students in each group who choose
each alternative/option.

record the count e.g. 1

A B* C D OMITS

Upper score (20) 0 10 0 0 0

Lower (20) 2 4 1 3 0

B* is the Key

1. Compute the difficulty of each item (% of pupils what got the item
right)
2. Compute the discriminating power of each item.( difference
between the number of pupils in the upper and lower groups who
got the item right)
3. Evaluate the effectiveness of destructors in each item
(attractiveness) of the incorrect alternative.
Item Difficulty

By difficulty level we mean the number of candidates that got a


particular item right in any given test. For example, if in a class of 45
students, 30 of the students got a question correctly, then the difficulty
level is 67% or 0.67. The proportion usually ranges from 0 to 1or 0 to
100%. An item with an index of 0 is too difficult hence everybody missed
it while that of 1 is too easy as everybody got it right. Items with index of
0.5 are usually suitable for inclusion in a test. Though the items with
indices of 0 and 1 may not really contribute to an achievement test, they
are good for the teacher in determining how well the students are doing
in that particular area of the content being tested. Hence, such items
could be included. However, the mean difficult level of the whole test
should be 0.5 or 50%.Usually, the formular for their difficulty is

pn x N = 100 where
P = item difficultn = the no of students who got the item correct.
N = the number of students involved in the test. However, in the
classroom setting, it is better to use the upper 13 of the students that
got the item right (U) and the lower 13 of the students that got it right
(L)

Where N is the number of students actually involved in the item analysis


(upper 13 + lower1/3 of the tests).Consider a class of two arms with a
population of 60 each. If 36 candidates of the upper 13 population and
20 of the lower 1/3 got question number 2 correctly, what is the index of
difficulty (difficulty level) of the question?

The difficulty of a test item is indicated by the % of pupils who get the item right.
Hence we can compute item difficulty by means of the following formulae.

Item difficulty =R/T x 100


Where:

R= the number of pupils who get the item right.

T= the total number of pupils who tried the item.

Using the example given, our item difficulty/difficulty index will be;

P = 14 x 100 = 35% = 0.35

40

In computing item difficulty from item analysis data, or calculation based on the
upper and lower groups only. We assume that the response of pupils in the
middle group follow essentially the same pattern.

The range of difficulty index is 0 to 1 or 0 to 100%

The difficulties index is needed in trying out (piloting) an exam to be given out.
If it is 0 the exam is hard.

Computing Item Discriminatory Power

The discrimination index shows how a test item discriminates between


the bright and the dull students. A test with many poor questions will
give a false impression of the learning situation. Usually, a discrimination
index of 0.4 and above are acceptable. Items which discriminate
negatively are bad. This may be because of wrong keys, vagueness or
extreme difficulty.

The discriminatory power of an achievement test item refers to the degree to


which it discriminates between pupils with high and low achievement. Item
Discriminatory power can be obtained by subtracting the number of pupils in
the lower group who get the item right (R L) from the number of pupils in the
upper group who get the item right (R U) and dividing by one half of the total
number of pupils included in the item analysis (½ T).

The formula is given as:

Item Discriminatory Power = RU - RL

½T

D = R U - RL
1/2
T

Using our example;

D = 10 – 4

½ (40)

= 0.6 = 0.3

20

NB. An item with no discriminatory power tone in which an equal number of


pupils in both the upper and lower groups get the item right. These results in
an index of 0.00, as follows:

D = 10 – 10 =0.00

10

With the formula, it is also possible to calculate an index of negative


discriminating power, i.e. one in which pupils in the lower group than the upper
group gets the item right. This is generally wasted effort, however, because we
are not interested in using items that discriminate in the wrong direction. Such
items should be revised so that they discriminate positively, or they should be
discarded.
NB. A low index of discriminatory power does not necessarily indicate a
defective item. Items that discriminate poorly between high and low achievers
should be examined for the possible presence of ambiguity, clues and other
technical defects. If none is found and the items measure an important learning
outcome, they should be retained for future use.

Item Analysis Procedure for Criterion Referenced Tests.

Since criterion referenced tests are designed to describe which learning tasks a
student can and cannot perform, rather than to discriminate among students,
the traditional indices of item difficulty and item discriminating power are of
little value. A set of items in a criterion-referenced mastery test, for example
might be answered correctly by all students (zero discriminating power) and said
be effective items. If the items closely match an important outcome the results
simply tell us that here is an outcome that all students have mastered. This is
valuable information for describing the types of tasks students can perform and
to eliminate such items from the test will distort our description of student
learning.

The difficulty of an item in a criterion-referenced test is determined by the


learning task if it is designed to measure. If the task is easy the item should be
easy. If the task is difficult the item should be difficult. No attempt should be
made to eliminate easy items or alter item difficulty simply to obtain a spread
of test scores. Although an index of item difficulty can be computed for items in
a criterion referenced test, there is seldom a need to do so. If mastery is being
measured and the instruction has been effective, criterion-referenced test items
are typically answered correctly by a large percentage of the students.

NB. A basic concern in evaluating the items in a criterion referenced mastery test
is the extent to which each item is measuring the effects of instruction. If an item
can be answered correctly by all students before and after instruction, the item
obviously is not measuring instructional effects. Similarlarly, if an item is
answered in correctly by all students between, before and after instruction, the
item is not serving its intended function. These are extreme example, of course,
but they highlight the importance of obtaining a measure of instructional effects
at one basis for determining item quality.

To obtain a measure of item effectiveness based on instructional effects, the


teacher must give the same test before instruction and after instruction. Effective
items will be answered correctly by a larger number of students after instructions
than before instruction. An index of sensitivity to instructional effect(s) can be
computed by using the following formula:

RA-RB

S =

Where

RA=the number of students answering the item correctly after instruction.

RB=the number of students answering correctly before instruction.

T=the total number answering the item both times.

E.g. an item that has been answered incorrectly by all students before instruction
and correctly by all students after instruction (T=32), on result would be as
follows:

S=32-0=1.0

32

Maximum sensitivity to instructional effects is indicated by an index of 1.00. the


index of effective items will fall between 0.00 and 1.00, and larger positive values
will indicate items with greater sensitivity to the effects of instruction.

Limitations
General limitations are answered using the sensitivity index:

1. The test must be given twice to complete the index.


2. A low index may be due to either an ineffective item or effective
instruction.
3. The item responses after instruction may be influenced to some extent by
having taken the same test earlier. {This limitation is likely to be most
serious where instruction time is short}.
NB. Despite these limitations, the sensitivity index provides a useful means of
evaluating the effectiveness of items in a criterion referenced mastery test.

Importance of Score Interpretation

There are two important functions of classroom tests;

i. To compare learners achievement with those of other testees who wrote the
same tests under the same condition.
ii. To determine how much and what each pupil knows the learning content over
certain period of time.
These two approaches are referred to as the norm-referenced and criterion-
referenced interpretation of the test scores respectively.

Norm-referenced Interpretation

In norm-referenced test, a testees performance is evaluated relative to the


performance of in the some well defined comparison or norm group.

This implies that the academic performance of pupil is compared with his/her
classmates. It may at times necessary to know how the achievement of two or
more students compares with each other. The selection of the top or poor
achievers is usually done by the ratting of raw or % scores.

Depending on what the function of the test will be, achievers may now be
selected to serve a specific function. Unfortunately the achievement of different
students in different subjects can not be compared unless the raw scores and
measuring instruments have been starndadised. Norm–reference interpretation
tells us very little to execute a specific task or to perform a specific function.
Therefore there is tendency to move away from norm-referenced to criterion
referenced-Interpretation where more emphasis is placed on ability of pupil to
measure up to prescribed prescribed criterion.

Criterion-referenced interpretation of scores

Most educationists prefer criterion referenced tests to norm-referenced test like


wise the criterion referenced examination is preferred to norm- referenced one.
These tests serve a valuable purpose of describing factors of the pupil with
respect to explicit and defined instructional objectives. This should be the
overriding function of the of all academic tests. It helps to know how well to
apply the knowledge and ability he/she has mastered.

The setting of a minimum or performance standard is one of the main


problem in criterion referencing and this score remain fixed for all
testees. Pupils with the scores with minimum starndard is considered
unsuitable for the specific tasks while student with scores above the
minimum scores will be graded as “suitable

Using Test Results

As earlier mentioned, conducting tests is not an end in itself. However,


before tests could be used for those purposes, the teacher needs to
know how well designed the test is in terms of difficulty level and
discrimination power, then he should be able to compare a child’s
performance with those of his peers in the class. Occasionally, he may
like to compare the child’s performance in one subject area with
[Link] do this, he carries out the following activities at various times:

i. Item analysis.
ii. Drawing of frequency Distribution Tables.
iii. Finding measures of central tendency (Mean, Mode, Median)
iv. Finding measures of Variability and Derived Scores.
v. Assigning grades.

Application of Assessment and Evaluation to Educational Decision Making

Assessment and evaluation are terms often used in connection with


achievement testing. What’s the distinction between the two?

Assessment – refers to the act or process of determining the present level


(usually of achievement) of a group or individual.

Evaluation – is a statement of test results that includes a judgment factor (e.g. if


the class is achieving higher than others in the school or Mary is doing better in
arithmetic than in English.

The assessment and evaluation information is important in educational decision


making. At the college and university level, administrators need information
about academic abilities of prospective students in order to make appropriate
admission decisions. They are interested in selecting those students who have
the ability to succeed and eliminating those who are likely to fail. At the primary
and secondary school levels, educators often are more interested in identifying
those students who have special educational needs. In the Kenyan case,
assessment and evaluation are meant for coaching and promotion of the
learners to the next level of education. Also at primary and secondary levels,
tests/ assessment are needed to identify learners who may be in need of special
remedial help. In other countries like U.S.A there is a need to identify the
intellectually gifted learners.

Institutions have long had a need to identify those individuals who are at either
end of the ability continuum. One group of students needing rather early
detection and remedial work work concerns those who experience
developmental delays in their maturation. The earlier these people can be
identified and provided with learning environments designed to stimulate their
intellectual and emotional development, the more likely it is that they will be
able to maximize their talents. Aptitude tests are employed to make these
decisions, while recognizing that some decisional errors inevitably will occur.
Typically, the inaccurate decisions that are made usually are more likely to be
false negative errors.

In essence, tests are used in making decisions about how to best group students
within a school system so that the instructional level chosen will be appropriate
for their developmental level.

While institutions usually are concerned with admissions and selection


decisions, individuals want information about their interests, strengths, and
weaknesses to help them for their educational and vocational futures. They are
interested in knowing what kinds of skills they posses, what they are capable of
learning, and what course of study and type of work they should pursue, i.e.
individuals being classified as less able than they are rather than false positive
errors (i.e. students being classified as more talented than they really are )
consequently, even when errors are made, it is likely that those students who
have been mis-classified will be given special remedial treatment until the errors
are detected. Testing errors can lead to the potential loss of educational
opportunities and possible stigma for the students involved.

An area receiving much attention today involves the identification of


intellectually gifted students. Institutions need to identify those talented
individuals who will benefit most from an enriched educational curriculum. In
order to properly challenge these students, more demanding and complex
learning environments need to be made available. Failure to provide adequate
stimulation often leads to poor motivation and performance on the part of
gifted students. They often become discouraged and fail to channel their talents
in positive directions.

Why would we necessarily see an improvement in scores when a new test group
is compared with the older norm group?

One factor that teachers often ‘teach to the test’ i.e. they note what content is
covered on the various standardized tests in use at their school and then modify
their teaching to cover that material in their classes. Curricular are often
modified by teachers / administrators to better reflect the test content so that
students will have exposure to the material and consequently perform well when
tested. This is in curriculum alignment.

Is teaching to the test and curriculum alignment reasonable methods of


attaining educational goals? Proponents claim that modifying curricula to reflect
test content ensures that our educational programmes are up to date and
responsive to the needs of the society.

Critics – claim that as long as we use test content to guide our teaching, we
jeopardize our ability to effectively evaluate how well we are really teaching.

NB. With all of the external pressure from education stake holders and other
interested parties outside the school system it appears that teaching to the test
will probably continue.

Is testing necessary?

Relationships between evaluation procedures and instructional objectives

Instructional objectives encompass a variety of learning outcomes, and


evaluation includes a variety of procedures. The key to effective evaluation of
student learning is to relate the evaluation procedures as directly as possible to
the intended learning outcomes. This is made easiest to accomplish if the
instructional objectives and learning outcomes have been clearly stated in terms
of student performance. It is then simply a matter of constructing or selecting
evaluation instruments that provide the most direct evidence concerning the
attainment of the stated outcomes. Evaluation is a continuous comprehensive
process which utilizes a variety of procedures and which is incapably related to
the objectives of the educational programme.

Activity
Preparing test items that are directly relevant to the instructional objectives to
be measured requires matching the performance measured by the test items to
the types of performance specified by the intended outcomes.

Explain a clear statement of instructional/educational objectives extremely


important at the basis for planning good evaluation.

Distribution and Measures of Central Tendency

We shall not dwell so much on the drawing of frequency distribution tables and
calculating measures of central tendency. This will be taken care of elsewhere. A
measure of central tendency is designed to give us a single value that is most
characterized or typical of a set of scores. Three such measures are common in
testing, ie. mode, median and mean

Mode: The mode is the most frequent or popular score in the population. This
is usually evident during the drawing of frequency tables. It is not frequently
used as the median and mean in the classroom because it can fall anywhere
along the distribution of scores (top, middle or bottom) and a distribution may
have more than one mode. If the group of scores is large, instead of working
out the mean score on estimate of central tendency called the mode is
sometimes used. The mode is the most commonly obtained score or the mid
point of the score interval having the highest frequency. The mode is a quick
(and rough) indication of central tendency but it is not especially useful in
connection with test scores. The mode should not be used for small samples.

Median: This is the middle score after all the scores have been arranged in order
of magnitude i.e.50% of the score are on either side of it. Median is very good
where there are deviant or extreme scores in a distribution, however, it does not
take the relative size of all the scores into consideration. Also, it cannot be
used for further statistical computations.
With test scores, we are likely to have one very high score (or, at best, a few very
high scores) and many lower scores. The result is that the mean tends to
exaggerate the scores i.e. negatively pulled towards the extreme scores or value,
and the median becomes the preferred measure.

NB. The median is that value above which falls 50% of the scores and below
which falls 50% of the scores; thus it is less likely to be drawn in the direction of
the extreme cases.

The median is also preferred when a distribution is truncated (cut off in some
way so that it can be no ease beyond a certain point) as in the figure below, the
distribution in truncated, perhaps is because of a very difficult test on which zero
was the lowest score assigned. The doted line suggests the distribution we
might have obtained in the scoring had permitted the assignment of the
negative scores.

If there is an odd number of candidates, the median is the middle score obtained
by the candidates eg. 10, 15, 22, 27, 37, 45, 51, 62, 73, and 91

The median is 45.

If there is an even number of candidate, as in our example of history scores, the


median is the average of the two central scores.

10, 20, 31, 50, 60, 70, m 80, 81, 94, 95, 100,

The median = 70+80 =75

2
The Mean: This is the average of all the scores and it is obtained by
adding the scores together and dividing the sum by the number of
scores. M or = X = S u m o f a l l S c o r e s N u m b e r o f
S c o r e s . Though, the mean is influenced by deviant scores, it is very
important in that it takes into cognizance the relative size of each score
in the distribution and it is also useful for other statistical calculations.

The most common measure of position and of central tendenancy in the


arithmetic mean (usually called simply the mean). This statistics is computed by
simply adding up all the scores and dividing by the number of scores.

i.e. ‾X = ∑X= or ‾X =∑fX

N N

Where

‾X=the mean of test X.

∑=summation of marks.

X=raw scores of test X.

N=Number of faces (population).

The value of ‾X it referred to as a mean score when it comes to tests and


examination.

Eg in a history examination the raw scores obtained by students were, 80, 95,
50, 81, 94, 60, 100, 10, 31, 80, 31, 80, 20, 70.

The sum of the raw scores = 771.

The number of students being tested =12.

772

Mean X= =64.25.

12
This method of finding the mean is quite satisfactory for small groups. If
however we require finding the mean of a large number of scores it is quicker
to estimate the mean to the nearest round number and find the sum of the
divisions of each score from the estimated mean.

Looking at the history scores we can estimate the mean to be between 60 and70.
Let us assume the mean is 60 using the formulae of deviations from the assumed
mean which is given as:

Deviation from assumed mean = raw scores -assumed mean,

i.e. deviation = X- M, then we can compute the mean of large groups.

Score Deviation Assumed Mean

X X -

80 20 -

95 35 -

50 - 10

81 27 -

91 34 -

60 0 0

100 40 -

10 - -

31 - 29

80 20 -

20 - 40
70 10 -

∑+ = 180

∑-=-
129

We then find the sum of the deviations, divide by the number in the group
and add this result to the assumed mean.

Sum of deviations: = 180 – 129.

=51.

Therefore: 51/12=4.25

Mean=60+4.25=64.25

NB satistitician note that the mean should be used to describe only the middle
of a set of scores larger than 30.

If we take 64.5 to be the mean then 7 students have scores higher than the
mean. This will not reflect a normal curve.

The mean in a female group in unduly influenced by the extreme scores in this
case by the three lower scores.

Therefore, the measure of central tendency which should be used when the
sample is small, is called the median.

Measures of Variability

Measures of variability tell us how much variability (or dispersion) there is in a


distribution, that is, they tell us how scattered the scores are. They include range
variance and standard deviation.
Range

The range is familiar to all of us, representing the difference between the highest
and the lowest scores. The range is easily found and easily understood, but is
valuable only as a rough indication of variability.

It is the least stable measure of variability, depending entirely on the two most
extreme (and therefore least typical) scores. It is less useful inconnection with
other statistics than other measures of variability are. It is obtained by the
Formula:

higher score – lower score.

I.e. R=Hs-Ls

Standard deviation

It is the most dependable measure of variability for it varies less than other
measures from one sample to the next. It is orderly accepted as the best
measure of variability and is of special value to test users because it is the basis
for:

- Standard scores
- A way of expressing the reliability of a test score.
- A way of indicating the accuracy of values predicted from a correction
coefficient.
- A common statistical test of significance.
NB. This statistics is one which every test user should know thoroughly.

The std deviation is equal to the square root of the mean of the squared
deviation from the distribution mean. The following formula is usually used.

Sx = ∑ (X-X‾ )2
N

Sx = ∑ X2-(X‾ )2

Sx = Standard deviation of test x

X= Raw scores on test x

X‾= Mean of the Test X

N= Number of Persons Whose Scores are Involved

Standard deviation is also used in making interpretation from the normal curve

Variance is calculated using the formula bellow.

S2 x = ∑ X2-(∑X‾ )2 / N

The Value of Measures of Central Tendency

An important use of measure of central tendency is to decide which measure is


representative of the performance of the group of a whole. It is important
because knowledge of a group performance enable the teacher or examiner to
interpret the value of the individual marks earned by members of the class. A
single mark has no meaning on its own except in relation to marks earned by
other members of the group, i.e. in relation to a standard to which in this case,
is a measure of central tendency.

Activity II

The mean score is the same as the average score i.e. Sum of all
scores/the number of scores. This is the most common statistical
instrument used in our classroom If in a class of 9, the scores are 29, 85,
78, 73, 40, 35, 20, 10 and 5. Find the mean.

Measures of Variability

Measure of variability indicates the spread of the scores. The usual


measures of variability are Range, Quartile Deviation and Standard
Deviation. Their computations are as illustrated below.

Range

The range is usually taken as the difference between the highest and the
lowest scores in a set of distribution. It is completely dependent on the
extreme scores and may give a wrong picture of the variability in the
distribution. It Is the simplest measure of
[Link]: 7, 2, 5, 4, 6, 3, 1, 2, 4, 7, 9, 8, 10. Lowest score = 1, Hig
hest = 10. Range =10 -1 = 9

Quartile Deviation

Note that Quartiles are points on the distribution which divide it into
“quartiles”, thus, we have 1st , 2nd and 3rd quartiles. Inter-quartile range
is the difference between Q3 and Q1 i.e. Q3 = Q1. This is often used than
the Range as it cuts off the extreme score. Semi inter-quartile Range
is thus half of inter-quartile range. This is also known as the semi-inter
quartile range. It is half the difference between the upper quartile (Q3)
and the lower quartile (Q1) of the set of scores.

QQ

312

Where Q3 = P75 = point in the distribution below which lie 75% of the
scores. Q1 = P25 = Point in the frequency distribution below which lies
25% of the scores. In cases where there are many deviant scores, the
quartile deviation is the best measure of variability.

Standard Deviation

This is the square root of the mean of the squared deviations. The mean
of the squared deviations is called the variance (S2). The deviation is the
difference between each score and the [Link] (Μ) =
∑ x N 2 x = X - X - deviation of each score from the mean
N = number of scores. The SD is the most reliable of all measures of
variability and lend itself for use in other statistical calculations. Deviation
is the difference between each score (X) and the mean (M). To calculate
the standard deviation: (i) find the mean (m) (ii) find the deviation (x-
m) and square each. (iii) sum up the squares and divide by the
number of the population (N) (iv) find the positive square root.

Activity

Find the mean and standard deviation for the following marks.20, 45. 39,
40, 42, 48, 30, 46 and 41.
DERIVED SCORES

In practice, we report on our students after examinations by adding


together their scores in the various subjects and thereafter calculate the
average or percentage as the case may be. This does not give a fair and
reliable assessment. Instead of using raw scores, it is better to use
derived scores”. A derived score usually expresses every raw score in
terms of other raw score on the test. The commonly used ones in the
class room are the Z-Scores, T-Score and Percentiles. The computation of
each of these will be demonstrated.

STANDARD SCORE OR Z-SCORE

Standard score is the deviation of the raw score from the mean divided
by the standard deviation i.e. Z =

X X SD −

Where Z = Z – score

X = any raw score

X = the mean SD = Standard Deviation Raw scores above the mean


usually have positive Z-scores while those below the mean have negative
Z-scores. Z-scores can be used to compare a child’s performance with
his peers in a test or his performance in one subject with another.

T-Score

This is another derived score often used in conjunction with the Z-


score. It is defined by the equation. T = 50 + 10Z Where z is the standard
score. It is also used in the same way as the Z-score except that the
negative signs are eliminated in T-Scores.
The use of measure of variability

A measure of variability takes the teacher or examiner to the degree of scatter


or divergence of the marks. This information is of practical importance.

[Link] a measure of variability is large, we can conclude that the marks are
widely distributed i.e. there is a marked distance in the performance of the
learners.

Generally, measures of variability are used in association with measures of


central tendency. E.g. suppose the objective of a unit in the university is to
acquire mastery of statistics - suppose after the unit had been taught the
learners was tested.

If the class mean is high while the standard deviation is low, then the objective
of the course unit had been achieved and the learner can move on to another
unit.

If the class is mean is low and standard deviation is low, then we can conclude
that the objective has not been achieved.

The unit should be repeated by the same teacher using another approach, or by
another teacher or by a contribution of both techniques. If the class is average
and the standard deviation is also average then we can conclude that the
objective of the unit has been partly achieved.

Achievement of the objectives of mastery must therefore be seen jointly in terms


of the magnitude of the mean and the std deviation.

Simplify comparisons of sets of number, especially large sets of


number, by calculating the center values using mean, mode and
median. Use the ranges and standard deviations of the sets to examine
the variability of data.
Course Credit System and Grade Points

Perhaps the most precious and valuable records after evaluation are the
marked scripts and the transcripts of a student. At the end of every
examination e.g. semester examination, the marked scripts are submitted
through the head of department or faculty to the Examination
Officer. Occasionally, the Examination Officer can round off the marks
carrying decimal, either up or down depending on whether or not the
decimal number is greater or less than 0.5The marks so received are
thereafter translated/interpreted using the Grade Point (GP),Weighted
Grade Point (WGP), Grade Point Average (GPA) or Cumulative Grade
Point Average (CGPA).

CREDIT UNITS

Courses are often weighed according to their credit units in the course
credit system. Credit units of courses often range from 1 to 4. This is
calculated according to the number of contact hours as
follows:1 credit unit = 15 hours of teaching.2 credit units = 15 x 2 or 30 h
ours3 credit units = 15 x 3 or 45 hours4 credits units = 15 x 4 or 60 hours
Number of hours spent on practicals are usually taken into consideration
in calculating credit loads.

GRADE POINT (GP)

This is a point system which has replaced the A to F Grading System as


shown in the summary table below.

WEIGHTED GRADE POINT (WGP)


This is the product of the Grade Point and the number of Credit Units
carried by the course i.e. WGP = GP x No of Credit Units.

RADE POINT AVERAGE (GPA)

This is obtained by multiplying the Grade Point attained in each course


by the number of Credit Units assigned to that course, and then
summing these up and dividing by the total number of credit units
taken for that semester (total registered for).GPA =

Total Points Scored Total Credit Units registered = Total WGP Total Credit Units
registered

CUMMULATIVE GRADE POINT AVERAGE (CGPA)

This is the up-to-date mean of the Grade Points earned by the student. It
shows the student’s overall performance at any point in the programme.
CGPA = Total Points so far Scored Total Credit Units so far taken or registered A

Common questions

Powered by AI

Item analysis plays a critical role in improving classroom instruction by identifying which test items were effective in measuring student achievement and by pinpointing areas where students experience difficulties. It involves reviewing student responses to test questions to assess each item's quality, determining whether the item functioned as intended, was of appropriate difficulty, lacked irrelevant clues, and had effective distractors. This process not only helps refine individual test items but also informs curriculum adjustments by highlighting topics that need more attention or are too easy or difficult .

Measures of central tendency, such as the mean, provide a benchmark against which individual student scores can be compared. Understanding the average performance of the group allows educators to contextualize each student's scores, revealing whether they are above or below average in relation to their peers. This is crucial for identifying students who may need additional support or enrichment, thus guiding instructional decisions .

Item analysis can significantly enhance future test construction and curriculum design by identifying technical flaws and adjusting item difficulty and discrimination levels. By analyzing which items effectively measure learning outcomes, educators can refine test content, ensuring that future tests are both valid and reliable. Moreover, insights from item analysis can inform curriculum adjustments by highlighting areas where students need more support or challenging content, promoting targeted instructional strategies .

Standard deviation is highly utilized in interpreting test results because it provides a reliable measure of score dispersion around the mean, helping educators assess the reliability of test scores and the consistency of test items. It is preferable because it varies less between samples and offers a standardized method for expressing score variability. Additionally, standard deviation is crucial in deriving standard scores, assessing test reliability, and analyzing prediction accuracy with correlation coefficients .

Z-scores and T-scores play significant roles in assessing student performance by providing a standard way to compare scores across different tests and subject areas. A Z-score indicates how many standard deviations a raw score is from the mean, with scores above the mean having positive values and those below having negative. T-scores, calculated from Z-scores using the formula T = 50 + 10Z, eliminate negative values, making them easier to interpret. Both scores allow for performance comparison within a peer group or across subjects .

The discrimination index is crucial for evaluating test items because it measures how well an item differentiates between high-performing and low-performing students. A high discrimination index (0.4 or above) indicates that an item effectively distinguishes between different levels of student achievement. It is computed by taking the difference in the number of high-achieving and low-achieving students who answered the item correctly, divided by half of the total number of students analyzed. A positive discrimination index suggests that the item performed well, whereas a negative index may indicate flaws such as ambiguity or excessive difficulty .

The range is the simplest measure of variability, representing the difference between the highest and lowest scores. Its main benefit is ease of calculation, making it a quick indicator of score dispersion. However, its limitations include a dependency on extreme scores, which can skew results and fail to provide a comprehensive picture of variability. As it relies solely on the two most extreme scores, it may not accurately reflect the overall distribution of scores .

The difficulty index influences test construction by ensuring that items are neither too easy nor too hard, thus providing a balanced assessment of student performance. An item difficulty index of 0 indicates that the item is too difficult as nobody got it right, whereas an index of 1 suggests that the item is too easy. Ideally, items should have a difficulty index around 0.5 to contribute effectively to an achievement test. However, items with extreme indices can still be useful to gauge students' understanding of specific content areas, guiding teachers in instructional adjustments .

High variability measures suggest a wide distribution of student scores, indicating a significant range in student performance levels. This scenario may signal differing levels of understanding, potentially due to varied instructional methods or diverse learner abilities. Such findings necessitate a review of teaching strategies to address gaps and ensure that all students can meet learning objectives. Educators might need to provide differentiated instruction to cater to the diverse needs reflected in such variability .

Items with low or negative discrimination indices might still be considered for use if they align closely with important learning outcomes and are free from technical issues like ambiguity or clues. These items may not distinguish well between high and low achievers but could provide valuable information on whether students have mastered specific content areas. In such cases, retaining these items helps maintain a comprehensive assessment of student learning, especially in criterion-referenced tests where mastery of content takes precedence over relative performance .

You might also like