0% found this document useful (0 votes)
18 views7 pages

Understanding Construct Validity in Testing

The document discusses the importance of construct validity in psychological testing, particularly in the context of translation ability assessment. It emphasizes the need for clear definitions of constructs, the distinction between norm-referenced and criterion-referenced tests, and the significance of validity, reliability, and authenticity in test design. Ultimately, a valid translation test must accurately measure the defined construct of translation ability.

Uploaded by

shadmahya89
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
18 views7 pages

Understanding Construct Validity in Testing

The document discusses the importance of construct validity in psychological testing, particularly in the context of translation ability assessment. It emphasizes the need for clear definitions of constructs, the distinction between norm-referenced and criterion-referenced tests, and the significance of validity, reliability, and authenticity in test design. Ultimately, a valid translation test must accurately measure the defined construct of translation ability.

Uploaded by

shadmahya89
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Construct Validity

An understanding of the concept of a psychological construct is


prerequisite to understanding construct validity. A psychological construct is an
attribute, proficiency, ability, or skill defined in psychological theories.
CRITERION - RELATED VALIDITY: A STRATEGY FOR NRTS
The concept of criterion-reZated validity (not to be confused with criterion -
referenced tests) is basically a subset of the ideas discussed under construct
validiw. Demonstration of criterion - related validity usually entails designing an
experiment, too, but in this case, one group of students takes two tests: the
new test that testers are developing, and another test that is already a well-
established measure of the construct involved
One source of confusion that arises from reports about criterion - related
validity studies is that it is sometimes called concurrent or predictive validity.
These avo labels are just variations on the same theme. Coltcurrent ualidity is
criterion - related validity but indicates that both measures were administered
at about the same time, as in the TOESLP example. Predictive validity is also
a variant of criterion - related validity, but this time the two sets of numbers
are collected at different times. In fact, for predictive validity, the purpose of
the test should logically be “predictive.”
Using a rubric to assess translation ability
Introduction
Translation has been characterized as both a process and a product (Cao 1996),
more pointedly a very complex process and product. The fact that translation is a
multi-dimensional and complex phenomenon may explain why there have been
few attempts to validly and reliably measure translation competence/ability. This
is evident when comparing the research produced in translation testing with that
produced in testing in related fields.
Valid and reliable procedures for measuring translation or interpreting (or any
other construct for that matter) start by posing essential questions about the pro-
cedures (Cohen 1994: 6) such as: for whom the test is written, what exactly the test
measures, who receives the results of the test, how results are used, etc. Testing for
both translation and interpreting share some similarities, specifically in the appli-
cation of basic principles of measurement. But, because of the differences between
these two the remainder of the discussion will focus solely on translation.
.2 Nature of the test
Among further relevant questions there are those concerning the nature of the
test to be developed. Is the test a norm-referenced or a criterion referenced one?
This distinction is important since it allows for different things. “A norm-ref-
erenced assessment provides a broad indication of a relative standing, while
criterion-referenced assessment produces information that is more descriptive
and addresses absolute decisions with respect to the goal” (Cohen 1994: 25). The
norm-referenced approach allows for an overall estimate of the ability relative to
the other examinees. Norm-referenced tests are normed using a group of exam-
inees (e.g. professional translators with X amount of years of experience, or trans-
lators who have graduated from translation programs 6 months before taking
thetest, etc.).
1.3 Validity
. As validity is multifaceted and multi-componential
(Weir 2005), different types of evidence are needed to support any kind of claims
for the validity of scores on a test. In 1985 the American Psychological Association
defined validity as “the appropriateness, meaningfulness and usefulness of the spe-
cific inferences made from test scores (in Bachman 1990: 243). Therefore, test va-
lidity is not to be considered in isolation, as a property that can only be attributed
to the test (or test design), nor as an all-or-none, but rather it is immediately linked
to the inferences that are made on the basis of test scores (Weir 2005).
As an example, let’s consider construct validity. This category is used in test-
ing to examine the extent to which test users can make statements and infer-
ences about a test taker’s abilities based on the test results. Bachman and Palmer
(1996) suggest doing a logical analysis of a testing instrument’s construct validity
by looking at the clarity and appropriateness of the test construct, the ways that
the test tasks do and do not test that construct, and by examining possible ar-
eas of bias in the test tasks themselves. Construct validity relates to scoring and
test tasks. Scoring interacts with the construct validity of a testing instrument in
two primary ways. Firstly, it is important that the methods of scoring reflect the
range of abilities that are represented in the definition of competency in the test
construct. Similarly, it is important to ask if the scores generated truly reflect the
measure of the competency described in the construct. Both of these questions
essentially are concerned with whether or not test scores truly reflect what the test
developers intended them to reflect.
Ideally, the only thing that should influence test scores is the candidate’s
competence, or lack thereof, as defined by the test’s construct. With the question
on validity comes the question of reliability.
.4 Reliability
Primarily, reliability is used as a technical term to describe the amount of
consistency of test measurement (Bachman 1990; Cohen 1994; Bachman & Palmer
1996) in a given construct. One way of judging reliability is by examining the
consistency of test scores.
To determine reliability, we can use the questions for making a logical evalu-
ation of reliability as set forth in Bachman & Palmer (1996). The factors impact-
ing reliability are: (1) variation in test administration settings, (2) variations in
test rubrics (scoring tool for subjective assessment), (3) variations in test input,
(4) variation in expected response, and, (5) variation in the relationship between
input and response types
1.5 Test authenticity
Another important aspect of testing is test authenticity. Authenticity is the term
that the testing community uses to talk about the degree to which tasks on a test
are similar to, and reflective of a real world situation towards which the test is tar-
geted. It is important that test tasks be as authentic as possible so that a strong re-
lationship can be claimed between performance on the test and the performance in
the target situation.
[Link] authenticity
When developers are creating a test, another important aspect of task authenticity
is the test format, and its impact on the security of the test. When organizations
require test takers to produce a translation, this task is reflective of the type of
task that a professional will perform in the target situation. The main areas in
which a test task may seem to be inauthentic are the response format, the avail-
ability of tools, and the lack of time for editing. Some of these are authenticity
problems which are logistically difficult to solve. The handwritten nature of the
response format is seen as being fairly inauthentic for most contemporary transla-
tion workplaces. However, this is not a simple problem to solve as there are im-
plications for test security and fairness, among others. It is possible, with current
technology improvements, to disable e-mail functions temporarily and prevent
the exam from leaving the examination room via the internet. Additionally exam
proctors can be asked to control the environment and not allow electronic devices
such as flash drives into the testing room to avoid downloading exam originals.
To increase authenticity, it is important that the computerized test format mirror
the tools and applications currently used by professional translators as closely as
possible while maintaining test security.
2. Defining the test construc
A construct consists of a clearly spelled out definition of exactly what a test de-
veloper understands to be involved in a given ability. If we are testing an ability
to translate, it is important that we first clearly and meticulously define exactly
what it is that we are trying to measure. This task not only involves naming the
ability, knowledge, or behavior that is being assessed but also involves breaking
it down into its constituent elements (Fulcher 2003). Thus, in order to measure
a translator’s professional ability in translating from one specific language into
another, we need to first define the exact skills and sub-skills that constitute a
translator’s professional ability. In order to design and develop a test that assesses
the ability to translate at a professional level, we have to define what the transla-
tion ability is. We have to operationalize it. The goal is to consider what type of
knowledge and skills (in the broadest sense) might contribute to an operational
definition of ‘translation ability’ that will inform the design and the development
of a test of translation competency (Fulcher 2003). That is, we must say exactly
what knowledge a translator needs to have and what skills a candidate needs to
have mastered in order to function as a qualified professional translator. These
abilities cannot be vague or generic.
To illustrate this we look at definitions (operationalizations) of translation
competence. One definition of translation competence (Faber 1998) states the fol-
lowing: “The concept of Translation Competence (TC) can be understood in terms
of knowledge necessary to translate well (Hatim & Mason 1990: 32f; and Beeby
1996: 91 in Faber 1998: 9). This definition does not provide us with specific de-
scriptions of the traits that are observable in translation ability, and therefore it
does not help us when naming or operationalizing the construct to develop a test.
To find an example of a definition developed by professional organizations,
we can look at the one published by the American Translators Association (ATA).
The ATA defines translation competence as the sum of three elements: (1) com-
prehension of the source-language text; (2) translation techniques; and (3) writ-
ing in the target language. In a descriptive article, Van Vraken, Diel-Dominique
& Hanlen (2008 [Link]
define criterion for comprehension of the source text as “translated text reflects
a sound understanding of the material presented.” The criterion for translation
techniques is defined as “conveys the full meaning of the original. Common
translation pitfalls are avoided when dictionaries are used effectively. Sentences
are recast appropriately for target language style and flow.” Finally, evaluation of
writing in the target language is based on the criterion of coherence and appropri-
ate grammar such as punctuation, spelling, syntax, usage and style. In this profes-
sional organization, the elements being tested (according to their definition) are
primarily those belonging to the sub-components of grammatical competence
(language mechanics) and textual competence (cohesiveness and style). But while
this definition is broader than that of Beeby, Faber, or Hatim and Mason (in Faber
1998), it still does not account for all the elements present in the translation task
required by their test.
We could argue that translation involves various traits that are observable
and/or visible which include, but are not limited to conveyance of textual mean-
ing, socio-cultural as well as sociolinguistic appropriateness, situational adequacy,
style and cohesion, grammar and mechanics, translation and topical knowledge.
These traits contribute to an operational definition of translation ability, and they
are essential to the development of a test.
A test can only be useful and valid if it measures exactly what it intends to
measure; that is, if it measures the construct it claims to measure. Therefore, for a
translation test to be valid, it must measure the correct construct, i.e. translation
ability. The first crucial task of the test developer is to define the construct clearly.
Once the construct is defined clearly, then and only then can the test developer
begin to create a test that measures that construct.

You might also like