EVALUATION
Evaluation is a broad term which involves the systematic way of gathering reliable and relevant
information for the purpose of making decisions.
Focus of Evaluation:
Evaluation may focus on different components of a course:
the achievement of the learners,
the teachers,
quality of the materials,
the appropriateness of the objectives,
the teaching methodology,
the syllabus etc.
Kinds of Evaluation:
Evaluative information can be both
quantitative (e.g. test scores)
qualitative (e.g. comments/opinions) in form (Bachman,1990 Lynch, 1996)
Methods for Collecting Evaluative Information:
It can be collected through different methods, which require the use of different data gathering
instruments, such as
questionnaires,
interviews,
classroom observation,
study of documents,
tests,
ratings.
Basis and Concerns of Evaluation:
In practice, most evaluations are concerned with the evaluation of programs or courses as
a whole.
Evaluation can make use of tests but it is not limited to such forms.
Apart from using tests, evaluation in the classroom may be based on the teacher’s own
subjective assessment (overall impression), and the assessment of the students’ classwork
and their homework.
Assessment
Although in several cases the terms evaluation and assessment can be used interchangeably, we
need to distinguish between these two terms.
What is Assessment?
Assessment involves testing, measuring or judging the progress, the achievement or the language
proficiency of the learners.
Focus of Assessment:
The focus is on the students’ learning and the outcomes of teaching.
Assessment may be one part of an evaluation.
In some cases classroom assessment can be non-judgemental, and does not provide
evidence for evaluating or grading students, it is simply used to assess and provide early
feedback on the students’ learning before tests, midterm and final exams are
administered.
Forms of assessment
There is a distinction between formative and summative assessment.
Formative assessment
It is used to monitor the students’ progress during a course, to check how much they have
learnt of what they should have learnt, and then using this information future teaching
might be modified if necessary.
It can be carried out in the form of informal tests and quizzes
It can be the basis for feedback to the students.
Summative assessment
It is used at the end of a term, a semester, or a year, to assess how much has been
achieved by individuals or groups.
It is usually carried out by using more formal tests.
Measurement
The terms measurement, test and evaluation are often used synonymously, and in practice they
may refer to the same activity. However, apart from their superficial similarities they are distinct
from each other.
Measurement in the social sciences is the process of quantifying the characteristics of
persons according to explicit procedures and rules (Bachman, 1990)
It means that we assign numbers to the different mental characteristics, attributes and
abilities, such as aptitude, intelligence, motivation, fluency in speaking, achievement in
reading comprehension, and this quantification must be done according to well defined
and set rules and procedures.
Tests
Definition:
An educational test is a measurement instrument which is designed to elicit a specific
sample of an individual’s language behaviour. (Bachman, 1990)
Language tests provide the means for focusing on the specific language abilities that we
are interested in.
The elicited specific kinds of language behavior then can be interpreted as the evidence
of the abilities which we are interested in.
Tests are often used for several pedagogical purposes.
They may give diagnostic information to the teacher about where the learners are at the
given moment to help decide what to teach next.
Tests also give information to the learners about what they know, so that they can decide
what they need to learn or review. In this way tests provide students with a sense of
achievement and progress in their learning and this way may motivate students to learn
specific material.
At the same time tests give tasks which themselves can provide useful practice.
Tests assessing performance provide a clear indication of reaching a certain phase in the
course such as the end of a unit, or the end of the course.
Tests may also be used for purely descriptive purposes in doing research.
Sometimes tests are misused in the classroom, and are used as a means to get a noisy
class to keep quiet and concentrate, which is a very bad practice.
In summary, it can be said that not all measurements are tests, not all tests are evaluative, and not
all evaluation involves either measurement or tests.
Criteria of good tests
It is always the tester’s task to provide the best solution to a particular testing problem. But what
is common is that every test or testing system has to fulfil the following requirements:
consistently provide accurate measures of precisely the abilities in which we are
interested (validity and reliability)
have a beneficial effect on teaching (in those cases where the tests are likely to influence
teaching (washback)
be economical in terms of time and money (practicality) (Hughes, 2003)
Validity
Definition:
Validity always refers to the degree to which the gathered empirical evidence supports the
adequacy and appropriateness of the inferences that are made from the scores (Bachman 1990)
It means that the interpretations and uses that we make of test scores are to be valid, in
other words, a test is said to be valid to the extent that is measures what it is supposed to
measure, (Alderson, 1995)
If the test is not valid for the purpose for which it was designed, the scores do not mean
what they are supposed to mean.
Types of Validity:
There are different types of validity, which in reality are different ’methods’ of assessing validity
(Bachman 1990)
There are three main types/aspects of validity that can be distinguished: internal, external and
construct validity.
1. Internal validity relates to the perceived content of the test and its perceived effect.
There are three aspects:
I. face validity,
II. content validity
III. response validity
i. Face validity refers to the surface credibility or public acceptability of the test.
It involves intuitive judgment about the test content made by so called ’lay’ people,
who are involved in the testing process but who are not experts in testing.
Such people include non-expert users, students, their teachers, administrators.
If test takers accept the test as a face valid test, they are more likely to perform to the
best of their ability on that test.
Data on face validity can be collected by interviewing test takers, students or asking
them to complete questionnaires about their feelings about or reactions to the test that
they have taken.
ii. Content validity shows whether the test contains a representative sample of the relevant
language skills.
It can be proved by a systematic analysis of the test content carried out by experts,
who make their judgments by comparing it with e.g. the test specification, which
states what the content should be or with a formal syllabus or curriculum, or by rating
test items and texts following a precise list of criteria .
A further alternative can be interviewing teachers of a range of academic subjects, or
administering a questionnaire survey in which respondents are asked to make
judgments about the texts and tasks.
iii. Response validity can be checked by gathering information on how test takers respond to the
test items.
It can be collected by asking learners and test takers to tell how they responded to the
test item, what their test taking behaviour was, because the reasoning and the
processes they follow when they are solving the items give important indications of
what the test is testing.
2. External validity relates to procedures which compare students’ test scores with measures of
their ability taken from outside the test.
It has two types:
i. concurrent validity
ii. Predictive validity.
i. Concurrent validation involves comparing the test scores of the candidates with some
other measure for the same candidates taken roughly at the same time as the test.
This measure can be expressed numerically with statistical methods by correlating
students’ test results with their test scores gained on other tests, with teachers’
rankings or with the students’ own ratings of their language ability in the form of self-
assessment.
ii. Predictive validation involves comparing the students’ test scores with some other
external measure taken some time after the test has been administered.
This measure also can be expressed numerically with statistical methods by
correlating students’ test results with their scores gained on other tests taken some
time later (e.g. correlate entrance test results with scores of final tests).
iii. Construct validity shows to what extent the test is based upon its underlying theory, that
is, how well test performance can be interpreted as a meaningful measure of some
characteristic or quality.
The term ‘construct’ refers to a psychological construct, a theoretical concept about a
kind of language behaviour that the test makers want to measure.
It can be regarded as a kind of attribute or ability of people which is assumed to be
reflected in test performance.
Construct validity refers to the extent to which performance on tests is consistent with
the predictions that test makers make based on a theory of abilities, or constructs.
(Bachman, 1990)
It can be assessed by correlating different test components (sub tests) with each other
or by complex statistics, a combination of internal and external validation.
Reliability
Definition:
Reliability refers to the consistency with which a test can be scored, that is, consistency from
person to person, time to time or place to place.
It means that tests are to be constructed, administered and scored in such a way that the scores
obtained on a test on a particular occasion are likely to be very similar to those which would
have been obtained if it had been administered with the same students with the same ability, but
at a different time (Hughes, 1991)
Components:
There are two components of test reliability:
The reliability of the scores on the performance of candidates from occasion to occasion,
which can be ensured by the construction and the administration.
The reliability of scoring
Reliability of scoring
Reliability of scoring can be achieved more easily with objectively scored tests (e.g. tests of
reading and listening comprehension), in which scoring does not require the scorer’s personal
judgment of the correctness, because the test items can be marked on the basis of right or wrong.
Importance:
Scorer reliability is especially important in the case of subjectively scored tests (i.e. tests of
writing and speaking skills), because they cannot be assessed on a right or wrong basis,
assessment requires a judgment on the part of the scorers.
Aspects of scorer reliability:
There are two aspects of scorer reliability: intra-rater reliability and inter-rater reliability.
Intra-rater reliability is achieved if the same scorer gives the same set of oral
performances or written texts the same scores on two different occasions. It can be
measured by means of a correlation coefficient
Inter-rater reliability refers to the degree of consistency of scores given by two or more
scorers to the same set of oral performances or written texts.
Reliability Coefficient
The reliability of a test can be quantified in the form of a reliability coefficient.
It can be worked out by comparing two sets of test scores.
These two sets can be obtained by administering the same test to the same group of test
takers twice (test retest method), or by splitting the test into two equivalent halves and
giving separate scores for the two halves, then correlating the scores (split half method)
The more similar are the two sets of scores the more reliable is the test said to be.
(Alderson, 1995)
The relationship of validity and reliability
A test cannot be valid if it cannot provide consistently accurate measurements.
It means that a valid test must be reliable, too.
However, a reliable test may not be valid.
It depends partly of what exactly we want to measure. For example multiple choice tests
can be made highly reliable especially if there are enough items, but performance on a
multiple choice test cannot be regarded a highly valid measure of one’s overall language
ability.
There is always some tension between reliability and validity.
In order to maximize reliability it is often necessary to reduce validity.
An oral test may be a valid measure but performance on it may be difficult to assess
reliably.
In practice there are only degrees of both that testers want to achieve, shortly there is a
trade-off between the two: one is maximised at the expense of the other.
The relationship between teaching and testing
We have discussed two very important criteria of a good test: validity and reliability.
Washback /backwash:
Another requirement is that a test should exert beneficial effect on teaching in cases when
the tests are likely to affect teaching. This effect on teaching and learning is called
washback /backwash.
It might happen that a test is considered to be so important that the preparation for it may
dominate the whole teaching and learning.
If the teaching is poor and inappropriate and the testing is good, that is, the test
administered is a valid test, based on the real communicative needs of the students and
includes tasks very similar to those that they have to perform in real life, testing will have
beneficial washback.
But if the content and the testing techniques are very far from the objectives of the
course, if the teaching is good and appropriate, testing is not, there may be a harmful
washback.
The proper relationship between teaching and testing should be that of a partnership,
testing should support good teaching, and if it is necessary it should exert a corrective
influence on bad teaching.
Within teaching systems, individuals need to be given feedback of their achievement, but
in some cases teachers’ assessments of their students might not be sufficient, especially if
the achievements of groups of learners are to be compared in order to make decisions.
That is why tests are necessary, but it is very important that tests should be of good
quality.
Practicality
Practicality refers to the efficiency in terms of the necessary equipment, the time needed for
setting, administering or marking the test, that is how easy and quick it is to set or score the test,
how much it costs, how simple it is, how much equipment is required to administer it.
Test types
Tests can be distinguished by the purposes for which testing is carried out. There can be several
purposes according to the kind of information that is sought by the tester.
The most common categories are the following:
1. measure overall language proficiency independent of any language course that the
candidates may have attended - proficiency test
2. measure the students’ achievement on a completed course of studies - achievement test
3. measure how much the students have learnt of the recently taught material progress test
4. diagnose students’ strengths and weaknesses in language knowledge and use , to find out
what they know and what they do not know - diagnostic test
5. assist placement of students by identifying the stage or part of a teaching programme
which is the most appropriate to the level of their proficiency – placement test
6. identify general abilities, to find out who is to be good at learning languages – aptitude
test
1. Aptitude tests
With the help of aptitude tests we can predict how good language learners students are
likely to become. These are given before the start of a language course.
These tests are not target language tests they are rather tests of general intelligence and
linguistic ability, and
Focus:
It tries to focus on factors thought to contribute to language learning:
memory,
grammatical ability,
vocabulary level in native language,
interest in the foreign language,
language analysis,
sound discrimination,
language background,
language learning attitudes,
Verbal intelligence, language aptitude, educational level, age, learning style etc.)
2. Placement tests
They measure the students’ general knowledge of the language; test their previous language-
learning experience in order to separate them into different levels of language proficiency so that
they can be arranged in groups or language classes of the appropriate level.
Administration:
Placement tests are administered at the start of a new language course or at the start of a
new phase of a language course.
They tend to be quick, simple and easy to administer and that is why they aim to make
only a rough estimate of language proficiency.
Principle:
Placement tests work on the principle of representative sampling; they select one or
two areas of the students’ language knowledge and take this sample as a representative of
their entire proficiency.
The most commonly used areas are grammar and vocabulary.
Technique:
Most placement tests rely on objective techniques for reasons of reliability and
practicality.
3. Achievement tests
They look back over a longer period to check how much of the language syllabus has been
acquired by the students, whether they have achieved the course objectives.
Purpose:
There can be several purposes why this information is needed:
Achievement tests are used for certification or promotion to a more advanced course, or
as an entry qualification for higher education.
Besides, by giving assessment of their achievement the test result can motivate students
to go further, the information can be used for planning the next phase, it can show how
successful the whole course has been in achieving objectives,
Moreover it can show the possible strengths and weaknesses of the whole language
programme, which can be used to make amendments in the course programme.
Base of Achievement Tests:
Achievement tests are always based on the taught syllabus and the teaching methods used
in the course.
Their content is closely related to the course content, to the materials and books that were
used and the testing techniques reflect the recommended methodology.
Examples of achievement test:
Typical examples of achievement tests are end of term tests, and they may be internal, set by
the teacher or the school,
but quite often they are quite formal and set externally by a ministry of education or a testing
authority. As they are looking back over a long period and have to test a broad range of
language, they are large-scale tests covering most or all the four skills.
Nature and administration
They are summative in nature, and are administered at the end of the course (e.g. the end of
the semester or school year)
4. Progress tests
They are very similar to achievement tests, in as much that they assess how much of what has
been taught has been learnt, but they look back to a shorter period e.g. a teaching unit, a chapter
of a textbook and they intend to measure the progress that the students are making.
Base:
Their content is also based on the course material, but they cover only one and two language
points and assess whether the students have mastered these points adequately.
Nature:
They are formative in nature, and given during the course, at the end of a unit of language
teaching
Range:
They are smaller scale tests (often they are called quizzes) because they concentrate only
one or two aspects of language and
they are informal tests, set by the teacher,
purpose:
the purpose of obtaining information about which areas need further attention in order to plan
future teaching. Theoretically they do not need to be graded, but in practice they are.
5. Diagnostic tests
Purpose:
They are used to identify students’ language problems, weaknesses or deficiencies with the
purpose of obtaining information of which language areas require further teaching in order to
plan future teaching priorities.
Then this information can be used to design a syllabus.
Range
Diagnostic tests are usually largescale tests, they look back over a wide range of language that
students should have learnt over a long period and assess which areas might be problematic.
Administration:
They are usually administered at the start of a new phase of language teaching (e.g. the start of
a new school year)
6. Proficiency tests
These tests are designed to measure the test takers’ ability in a language, their present level of
mastery regardless of any previous training.
Base:
The content of a proficiency test is not based on the content, syllabus or objectives of language
courses; it is rather based on a specification of what test takers have to be able to do in order
to be proficient.
Range:
Basically proficiency tests are large scale, complex tests, covering a wide range of language and
they look forward to language applications in the real world.
Nature:
They are summative, and try to simulate the target language tasks and cover relevant language
skills in an authentic way.
Main function:
The main function is to test if the test takers have the necessary language skills or the degree of
mastery of these skills.
Administration;
Proficiency tests are usually devised and administered by external testing bodies.