Validity
According to Oriondo & Antonio (1989), validity refers to the extent to which the test serves its
purpose or the efficiency with which it measures what it intends to measure. Evaluation should
utilize appropriate and efficient assessment instruments. Validity is the degree to which
assessment instruments can gather accurate data. The teacher must construct assessment
instruments that can evaluate the contents and behaviors which he wants to assess. The contents
and behaviors, as subjects of assessment are clearly stated in the instructional objectives in the
teacher's lesson plan.
The Meaning of Validity
Validity refers to the degree to which a test actually measures what it tries to measure. The
validity of a test concerns what the test measures and how well it does so. There are two major
forms of validity. These are external validity and internal validity. External validity refers to the
ability to be generalized across persons, settings, and times. Internal validity refers to the ability
of a test to measure what it purports to measure.
For a test to be reliable, it must be valid. A test is valid if it measures what it purports to measure.
For example, if a scale is off by 5 pounds, it read weight every day with excess of 5 pounds. The
scale is reliable because it consistently reports the same weight every day, but it is NOT valid
because it adds 5 pounds to the true weight. It is not a valid measure of the true weight (Linn and
Gronlund, 2000).
Validity is the extent to which a test measures what it purports to measure or as referring to the
appropriateness, correctness, meaningfulness and usefulness of the specific decisions a teacher
makes based on the test results. These two definitions of validity differ in the sense that the first
definition refers to the test itself while the second refers to the decisions made by the teacher
based on the test.
Factors Affecting Validity
Factors that Affect the Validity of a Test
There are some factors that greatly affect the validity of a test.
These include the following:
1. Inappropriateness of the test items. Measuring the understandings, thinking skills, and other
complex types of achievement with test forms that are appropriate only for measuring factual
knowledge will invalidate the results.
2. Directions of the test items. Directions that are not clearly stated as to how the students
respond to the items and record their answers will tend to lessen the validity of the test items.
3. Reading vocabulary and sentence structure. Vocabulary and sentence structures that do not
match the level of the students will result in the test of measuring reading comprehension or
intelligence rather than what it intends to measure.
4. Level of difficulty of the test item. When the test items are too easy and too difficult they
cannot discriminate between the bright and the poor students. Thus, it will lower the validity of
the test.
5. Poorly constructed test items. Test items which unintentionally provide clues to the answer
will tend to measure the students' alertness in detecting clues and the important aspects of
students performance that the test is intended to measure will be affected.
6. Length of the test items. A test should be of sufficient number of items to measure what it is
supposed to measure. If a test is too short to provide a representative sample of the performance
that is to be measured, validity will suffer accordingly.
7. Arrangement of the test items. Test items should be arranged in an increasing difficulty.
Placing difficult items early in the test may cause mental blocks and it may take up too much
time for the students; hence, students are prevented from reaching items they could easily
answer. Therefore, improper arrangement may also affect validity by having a detrimental effect
on students' motivation.
8. Pattern of the answers. A systematic pattern of correct answers will enable students to guess
answers, and this will lower again the validity of the test.
9. Ambiguity. Ambiguous statements in test items contribute to misinterpretations and confusion.
Ambiguity sometimes confuses the bright students more than the poor students, causing the
items to discriminate in a negative direction.
Types of Validity
Types of Validity
1. Face Validity. It ascertains that the measure appears to be assessing the intended construct
under study.
Example: If the measure of art appreciation is created, all of the items should be related to the
different components and types of art. If the questions are regarding historical time periods with
no reference to any artistic movement, teacher may be motivated to give their best
Effort in this measure because they do believe it is a true assessment of art appreciation.
Face validity tells us nothing about what a test actually measures. Face validity refers to how test
takers perceive the attractiveness and appropriateness o test. Why then is it important? If test
takers consider the test to have face validity, they may offer a more conscientious effort to
complete the test. If a test does not have face validity they might hurry through a test and take it
less seriously.
2. Construct Validity. This is used to ensure that the measure is actually measure of what it is
intended to measure (construct), and not other variables. Using panel of experts who are familiar
with the construct is a way in which this type of validity can be assessed. The experts can
examine the items and decide what that specific item is intended to measure. Students can be
involved in this process to obtain their feedback.
Example: A women’s studies program may design a cumulative assessment of learning
throughout the major. The questions are written with complicated wording and phrasing.
This can cause the test inadvertently becoming a test or reading comprehension, rather than a test
of women’s studies. It is important that the measure is actually assessing the intended construct,
rather than extraneous factor.
A test has construct validity if it accurately measures a theoretical, non-observable construct or
trait. The construct validity of a test is worked out over a period of time on the basis of an
accumulation of evidence. There are a number of ways to establish construct validity.
What is a Construct? Constructs are attributes that exist in the theoretical sense. Thus, they do
not exist in either the literal or physical sense. Despite this, we can observe and measure
behaviors that provide evidence of these constructs. For example, consider gravity. We cannot
see gravity, but we can see what we assume to be its results – a falling apple.
Definitions of constructs often vary from person to person, even among persons who are
considered experts in an area of study. For example, take the construct of alcoholism. Consider
the construct introduced by Corbett and Wilson (1991) of self-efficacy. It is defined as a person’s
expectations about his or her own competence and ability to accomplish an activity or task.
From the model Bandura (1977) proposed the following about the construct of self-efficacy,
expectations of personal efficacy determine whether coping behavior will be initiated, hote much
effort will be expanded, and how long it will be sustained in the face of obstacles and aversive
experiences.
Since our ability to measure an abstract concept like self-efficacy depends on our ability to
observe and measure related behavior, how should we go about defining or explaining a
psychological construct?
Three (3) steps referred to as construct explication outlines the process of defining a construct.
1. Identify the behaviors that relate to the construct. The more you can generate the better
able you are to define the construct.
2. Identify other constructs that may be related or unrelated to the construct being explained.
This will help determine the boundaries of the construct.
3. Identify behaviors related to these similar and dissimilar constructs and determine
whether these behaviors are related to the current construct being measured.
Two methods of establishing a test’s construct validity are convergent/divergent validation and
factor analysis.
a. Convergent divergent validation. A test has convergent validity
If it has a high correlation with another test that measures the same construct. By contrast, a
test’s divergent validity is demonstrated through a low correlation with a test that measures a
different construct. Note: this is the only case when a low correlation coefficient (between two
test that measure different traits) provides evidence of high validity.
b. Factor analysis. Factor analysis is a complex statistical procedure which is
conducted for a variety of purposes, one of which is to assess the construct
validity of a test or a number of tests.
Other Methods of Assessing Construct Validity
Item Analysis. There are a variety of techniques for performing an item analysis, which is often
used, for example, to determine which items will be kept for the final version of a test. Item
analysis is used to help “build” reliability and validity are into the test from the start. Item
analysis can be both qualitative and quantitative. The former focuses on issues related to the
content of the test, e.g., content validity. The latter primarily includes measurement of item
difficulty and item discrimination.
Item difficulty. An item’s difficulty level is usually measured in terms of the percentage of
examinees who answer the item correctly. This percentage is referred to as the item difficulty
index, or “p”.
Item discrimination. It refers to the degree to which items
Differentiate among examinees in terms of the characteristic being measured (e.g., between high
and low scorers). This can be measured in many ways. One method is to correlate item responses
with the total test score; items with the highest test correlation with the total score are retained
for the final version of the test. This would be appropriate when a test measures only one
attribute and internal consistency is important.
Another way is a discrimination index (D).
A teacher can asses the test’s internal consistency. That is, if a test has construct validity, scores
on the individual test iterns should correlate highly with the total test score. This is evidence that
the test is measuring a single construct
Developmental Changes are tests measuring certain constructs can be shown to have construct
validity if the scores on the tests show predictable developmental changes over time.
Experimental Intervention, that is if a test has construct valid-ity, scores should change following
an experimental manipu-lation, in the direction predicted by the theory underlying the construct.
4. Criterion-related Validity. This is used to predict future or current performance – it
correlates test results with another criterion of interest.
Example: If a physics program designed to measure to assess cumulative student learning
throughout the major, the new measure could be correlated with a standardized measure of
ability in this disci-pline. The higher the correlation between the established measure and new
measure, the more faith the teachers can have in the new assess-ment tool.
Criterion-related validity is generally used for tests that claim to predict outcomes. Construct
validity is appropriate when your test is measuring an abstract concept like beauty. Construct
validity involves the accumulation of a variety of evidence (reliability and other types of
validity) that shows the test is functioning as it was intended to. For example, your test of beauty
may correlate highly with another test of beauty, or may cause specific behaviors (e.g., dilation
of pupils) as predicted by some theory.
5. Formative Validity. When applied to outcomes assessment, it is used to assess how well a
measure is able to provide information to help improve the program under study.
Example: When designing a rubric for history, one could assess student’s knowledge across the
discipline. If the measure can provide information that students are lacking knowledge in a
certain area, then that assessment tool is providing meaningful information that can be used to
improve the course or program requirements.
6. Sampling Validity. This is similar to content validity. It ensures that the measure covers
broad range of areas within the concept under study. Not everything can be covered, so
items need to be sampled from all of the domains. This may need to be completed using a
panel of experts to ensure that the content area is adequately sampled. In addition, a panel
can help limit biasa test reflecting what an individual personality feels is important or
relevant.
Example: When designing an assessment of learning in the theater department, it would not be
sufficient to only cover issues related to acting. Other areas of theater such as lighting, sound,
functions of stage managers should all be included. The assessment should reflect the content
area in its entirety.
Face Validity. This is done by examining the test to find out if it is a good one. In judging face
validity, there should be at least three knowledgeable persons who are qualified to pass judgment
on the appropriateness, suitability, and mechanics in the construction of the tests. There is no
common numerical method for face validity.
Content Validity. Like the face validity, it is done by examining the test by at least three
knowledgeable persons but this time there is a more serious query to find out if the test really
measures what it seeks to measure. In judging content validity, one should look at both the
subject matter covered in the test and the type of behavior to be measured. Mathematical method
may be applied to the problem of content validity. In this case, item analysis is employed.
Construct validity. Construct validity refers to the degree to which the test can be described
psychologically. Here, it is assumed that there are factors that make up psychological construct.
For example, a psychologist developed a new test which consisted of 15 subtests to measure
three distinct psychological constructs and wanted to validate that test. Say a sample of 150
subjects was drawn from the population and measured on the 15 subtests. To establish the the
construct validity, the 150 by 15 data matrix needs to be analysis statistically using factor
analysis. Factor analyzed has been widely used to access the construct validity of the test, and
this topic is beyond the scope of this book since it falls under applied multivariate analysis, a
higher statistical method which requires statistical softwares. Construct validity. Construct
validity refers to the degree to which the test can be described psychologically. Here, it is
assumed that there are factors that make up psychological construct. For example, a psychologist
developed a new test which consisted of 15 subtests to measure three distinct psychological
constructs and wanted to validate that test. Say a sample of 150 subjects was drawn from the
population and measured on the 15 subtests. To establish the the construct validity, the 150 by 15
data matrix needs to be analysis statistically using factor analysis. Factor analyzed has been
widely used to access the construct validity of the test, and this topic is beyond the scope of this
book since it falls under applied multivariate analysis, a higher statistical method which requires
statistical softwares.
Construct validity. Construct validity refers to the degree to which the test can be described
psychologically. Here, it is assumed that there are factors that make up psychological construct.
For example, a psychologist developed a new test which consisted of 15 subtests to measure
three distinct psychological constructs and wanted to validate that test. Say a sample of 150
subjects was drawn from the population and measured on the 15 subtests. To establish the the
construct validity, the 150 by 15 data matrix needs to be analysis statistically using factor
analysis. Factor analyzed has been widely used to access the construct validity of the test, and
this topic is beyond the scope of this book since it falls under applied multivariate analysis, a
higher statistical method which requires statistical softwares.
Construct validity. Construct validity refers to the degree to which the test can be described
psychologically. Here, it is assumed that there are factors that make up psychological construct.
For example, a psychologist developed a new test which consisted of 15 subtests to measure
three distinct psychological constructs and wanted to validate that test. Say a sample of 150
subjects was drawn from the population and measured on the 15 subtests. To establish the the
construct validity, the 150 by 15 data matrix needs to be analysis statistically using factor
analysis. Factor analyzed has been widely used to access the construct validity of the test, and
this topic is beyond the scope of this book since it falls under applied multivariate analysis, a
higher statistical method which requires statistical softwares.
Construct validity. Construct validity refers to the degree to which the test can be described
psychologically. Here, it is assumed that there are factors that make up psychological construct.
For example, a psychologist developed a new test which consisted of 15 subtests to measure
three distinct psychological constructs and wanted to validate that test. Say a sample of 150
subjects was drawn from the population and measured on the 15 subtests. To establish the the
construct validity, the 150 by 15 data matrix needs to be analysis statistically using factor
analysis. Factor analyzed has been widely used to access the construct validity of the test, and
this topic is beyond the scope of this book since it falls under applied multivariate analysis, a
higher statistical method which requires statistical softwares.
Construct validity. Construct validity refers to the degree to which the test can be described
psychologically. Here, it is assumed that there are factors that make up psychological construct.
For example, a psychologist developed a new test which consisted of 15 subtests to measure
three distinct psychological constructs and wanted to validate that test. Say a sample of 150
subjects was drawn from the population and measured on the 15 subtests. To establish the the
construct validity, the 150 by 15 data matrix needs to be analysis statistically using factor
analysis. Factor analyzed has been widely used to access the construct validity of the test, and
this topic is beyond the scope of this book since it falls under applied multivariate analysis, a
higher statistical method which requires statistical softwares.
Construct validity. Construct validity refers to the degree to which the test can be described
psychologically. Here, it is assumed that there are factors that make up psychological construct.
For example, a psychologist developed a new test which consisted of 15 subtests to measure
three distinct psychological constructs and wanted to validate that test. Say a sample of 150
subjects was drawn from the population and measured on the 15 subtests. To establish the the
construct validity, the 150 by 15 data matrix needs to be analysis statistically using factor
analysis. Factor analyzed has been widely used to access the construct validity of the test, and
this topic is beyond the scope of this book since it falls under applied multivariate analysis, a
higher statistical method which requires statistical softwares.
Construct validity. Construct validity refers to the degree to which the test can be described
psychologically. Here, it is assumed that there are factors that make up psychological construct.
For example, a psychologist developed a new test which consisted of 15 subtests to measure
three distinct psychological constructs and wanted to validate that test. Say a sample of 150
subjects was drawn from the population and measured on the 15 subtests. To establish the the
construct validity, the 150 by 15 data matrix needs to be analysis statistically using factor
analysis. Factor analyzed has been widely used to access the construct validity of the test, and
this topic is beyond the scope of this book since it falls under applied multivariate analysis, a
higher statistical method which requires statistical softwares.
Construct validity. Construct validity refers to the degree to which the test can be described
psychologically. Here, it is assumed that there are factors that make up psychological construct.
For example, a psychologist developed a new test which consisted of 15 subtests to measure
three distinct psychological constructs and wanted to validate that test. Say a sample of 150
subjects was drawn from the population and measured on the 15 subtests. To establish the the
construct validity, the 150 by 15 data matrix needs to be analysis statistically using factor
analysis. Factor analyzed has been widely used to access the construct validity of the test, and
this topic is beyond the scope of this book since it falls under applied multivariate analysis, a
higher statistical method which requires statistical softwares.
Construct validity. Construct validity refers to the degree to which the test can be described
psychologically. Here, it is assumed that there are factors that make up psychological construct.
For example, a psychologist developed a new test which consisted of 15 subtests to measure
three distinct psychological constructs and wanted to validate that test. Say a sample of 150
subjects was drawn from the population and measured on the 15 subtests. To establish the the
construct validity, the 150 by 15 data matrix needs to be analysis statistically using factor
analysis. Factor analyzed has been widely used to access the construct validity of the test, and
this topic is beyond the scope of this book since it falls under applied multivariate analysis, a
higher statistical method which requires statistical softwares.
Construct validity. Construct validity refers to the degree to which the test can be described
psychologically. Here, it is assumed that there are factors that make up psychological construct.
For example, a psychologist developed a new test which consisted of 15 subtests to measure
three distinct psychological constructs and wanted to validate that test. Say a sample of 150
subjects was drawn from the population and measured on the 15 subtests. To establish the the
construct validity, the 150 by 15 data matrix needs to be analysis statistically using factor
analysis. Factor analyzed has been widely used to access the construct validity of the test, and
this topic is beyond the scope of this book since it falls under applied multivariate analysis, a
higher statistical method which requires statistical softwares.
Construct validity. Construct validity refers to the degree to which the test can be described
psychologically. Here, it is assumed that there are factors that make up psychological construct.
For example, a psychologist developed a new test which consisted of 15 subtests to measure
three distinct psychological constructs and wanted to validate that test. Say a sample of 150
subjects was drawn from the population and measured on the 15 subtests. To establish the the
construct validity, the 150 by 15 data matrix needs to be analysis statistically using factor
analysis. Factor analyzed has been widely used to access the construct validity of the test, and
this topic is beyond the scope of this book since it falls under applied multivariate analysis, a
higher statistical method which requires statistical softwares.
Construct validity. Construct validity refers to the degree to which the test can be described
psychologically. Here, it is assumed that there are factors that make up psychological construct.
For example, a psychologist developed a new test which consisted of 15 subtests to measure
three distinct psychological constructs and wanted to validate that test. Say a sample of 150
subjects was drawn from the population and measured on the 15 subtests. To establish the the
construct validity, the 150 by 15 data matrix needs to be analysis statistically using factor
analysis. Factor analyzed has been widely used to access the construct validity of the test, and
this topic is beyond the scope of this book since it falls under applied multivariate analysis, a
higher statistical method which requires statistical softwares.
Construct validity. Construct validity refers to the degree to which the test can be described
psychologically. Here, it is assumed that there are factors that make up psychological construct.
For example, a psychologist developed a new test which consisted of 15 subtests to measure
three distinct psychological constructs and wanted to validate that test. Say a sample of 150
subjects was drawn from the population and measured on the 15 subtests. To establish the the
construct validity, the 150 by 15 data matrix needs to be analysis statistically using factor
analysis. Factor analyzed has been widely used to access the construct validity of the test, and
this topic is beyond the scope of this book since it falls under applied multivariate analysis, a
higher statistical method which requires statistical softwares.
Construct validity. Construct validity refers to the degree to which the test can be described
psychologically. Here, it is assumed that there are factors that make up psychological construct.
For example, a psychologist developed a new test which consisted of 15 subtests to measure
three distinct psychological constructs and wanted to validate that test. Say a sample of 150
subjects was drawn from the population and measured on the 15 subtests. To establish the the
construct validity, the 150 by 15 data matrix needs to be analysis statistically using factor
analysis. Factor analyzed has been widely used to access the construct validity of the test, and
this topic is beyond the scope of this book since it falls under applied multivariate analysis, a
higher statistical method which requires statistical softwares.
Construct validity. Construct validity refers to the degree to which the test can be described
psychologically. Here, it is assumed that there are factors that make up psychological construct.
For example, a psychologist developed a new test which consisted of 15 subtests to measure
three distinct psychological constructs and wanted to validate that test. Say a sample of 150
subjects was drawn from the population and measured on the 15 subtests. To establish the the
construct validity, the 150 by 15 data matrix needs to be analysis statistically using factor
analysis. Factor analyzed has been widely used to access the construct validity of the test, and
this topic is beyond the scope of this book since it falls under applied multivariate analysis, a
higher statistical method which requires statistical softwares.
Construct validity. Construct validity refers to the degree to which the test can be described
psychologically. Here, it is assumed that there are factors that make up psychological construct.
For example, a psychologist developed a new test which consisted of 15 subtests to measure
three distinct psychological constructs and wanted to validate that test. Say a sample of 150
subjects was drawn from the population and measured on the 15 subtests. To establish the the
construct validity, the 150 by 15 data matrix needs to be analysis statistically using factor
analysis. Factor analyzed has been widely used to access the construct validity of the test, and
this topic is beyond the scope of this book since it falls under applied multivariate analysis, a
higher statistical method which requires statistical softwares.
Construct validity. Construct validity refers to the degree to which the test can be described
psychologically. Here, it is assumed that there are factors that make up psychological construct.
For example, a psychologist developed a new test which consisted of 15 subtests to measure
three distinct psychological constructs and wanted to validate that test. Say a sample of 150
subjects was drawn from the population and measured on the 15 subtests. To establish the the
construct validity, the 150 by 15 data matrix needs to be analysis statistically using factor
analysis. Factor analyzed has been widely used to access the construct validity of the test, and
this topic is beyond the scope of this book since it falls under applied multivariate analysis, a
higher statistical method which requires statistical softwares.
Construct validity. Construct validity refers to the degree to which the test can be described
psychologically. Here, it is assumed that there are factors that make up psychological construct.
For example, a psychologist developed a new test which consisted of 15 subtests to measure
three distinct psychological constructs and wanted to validate that test. Say a sample of 150
subjects was drawn from the population and measured on the 15 subtests. To establish the the
construct validity, the 150 by 15 data matrix needs to be analysis statistically using factor
analysis. Factor analyzed has been widely used to access the construct validity of the test, and
this topic is beyond the scope of this book since it falls under applied multivariate analysis, a
higher statistical method which requires statistical softwares.
Concurrent validity, Concurrent validity refers to the degree to which the test correlates
with a criterion, which is set up as an acceptable measure or standard other than the test
itself. The criterion is always available at the time of testing.
Mathematical method may be applied to the problem of concurrent validity! The appropriate
statistical method is the correlation technique. The formula is shown below.