0% found this document useful (0 votes)
24 views6 pages

Understanding Test Validity in Research

Research

Uploaded by

Pankaj Singh
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
24 views6 pages

Understanding Test Validity in Research

Research

Uploaded by

Pankaj Singh
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

VALIDITY

The degree to which the test actually measures what


it claims to measure.
Validity can be defined as the accuracy with which the scale measure what it claims to
measure. Validity and purpose are like two sides of a coin. Any measuring instrument which
fulfils the purpose for which it is developed can be called a valid measuring instrument. It is also
the extent to which the inferences and conclusions made on the basis of scores earned on
measuring are appropriate and meaningful.
According to H. E. Garrett (1965)
“The validity of a test or any measuring instrument depends
upon the fidelity with which it measures what it proposes to
measure.”
According to Freeman (1960)
"An index of validity shows the degree to which a test
measures what it purpose to measure when compared with
accepted criteria"
According to Anastasi (2007)

“The validity of a test concerns with what the test measures and
how well it does so.”
The first essential quality that any valid test should possess is Reliability. Reliability of
any test can be estimated by repetition of measurements but validity of a test can be estimated by
comparing the performance with some standard criterion. The Validation of a test score is the
most important step in the process of standardization of any tool. Therefore, every constructor
has to establish the validity of the tool to ensure its acceptability.
 In psychometrics, validity has a particular application known as test validity: "the degree
to which evidence and theory support the interpretations of test scores.
 Validity refers to how well a test measures what it purport to measure.
 Validity refers to the credibility or believability of the research.
 Validity is a measure of how well a test measures what it claims to measure.

1
2
Aspects/Nature of Validity

 There are two aspects of validity:

Internal validity - the instruments or procedures used in the research measured what they
were supposed to measure. Example: As part of a stress experiment, people were shown
photos of war. Afterwards they were asked how the pictures made them feel. They
responded that the pictures were very upsetting. In this study, the photos have good internal
validity as stress producers.

External validity - the results can be generalized beyond the immediate study. If the data is
collected from students of multiple universities, the results can be generalized to same
sample of the region.
Types of Validity
1. Face Validity

 Ascertains that the measure appears to be assessing the intended construct under study.
 Example: if a test of DEPRESSION appears/looks like measuring symptoms of
depression. Its face validity is considered high.
 Face validity is one of the most basic measures of validity. Essentially, researchers are
simply taking the validity of the test at face value by looking at whether a test appears to
measure the target variable. On a measure of happiness, for example, the test would be
said to have face validity if it appeared to actually measure levels of happiness.
 Obviously, face validity only means that the test looks like it works. It does not mean that
the test has been proven to work. However, if the measure seems to be valid at this point,
researchers may investigate further in order to determine whether the test is valid and
should be used in the future.
 Obviously, while face validity might be a good tool for determining whether a test seems
to measure what it purports to measure, having face validity alone does not mean that a
test is actually valid. Sometimes a test looks like it is measuring one thing, while it is
actually measuring something else entirely. So face validity is not technically valid.
2. Criterion Validity

 Criterion Validity assesses whether a test reflects a certain set of abilities.

3
 A test is said to have criterion-related validity when the test has demonstrated its
effectiveness in predicting criterion or indicators of a construct — for instance, when an
employer hires new employees based on normal hiring procedures like interviews,
education, and experience. This method demonstrates that people who do well on a test
will do well on a job, and people with a low score on a test will do poorly on a job.
 Criterion-Related Validity is used to predict future or current performance - it correlates
test results with another criterion of interest.
 Example: If a physics program designed a measure to assess cumulative student learning
throughout the major. The new measure could be correlated with a standardized measure
of ability in this discipline, such as an ETS field test or the GRE subject test. The higher
the correlation between the established measure and new measure, the more faith
stakeholders can have in the new assessment tool.
Types of criterion validity

 There are two different types of criterion validity


Concurrent validity
 measures the test against a benchmark (standard) test and high correlation indicates that
the test has strong criterion validity.
 It occurs when the criterion measures are obtained at the same time as the test scores.
This indicates the extent to which the test scores accurately estimate an individual’s
current state with regard to the criterion. For example, a new developed test of depression
is administered on the sample along with already developed test of DEPRESSION as a
standard measure.
 Concurrent validity is a measure of how well a particular test correlates with a previously
validated measure. It is commonly used in social science, psychology and education.
 As the name suggests, concurrent validity relies upon tests that took place at the same
time. Ideally, this means testing the subjects at exactly the same moment, but some
approximation is acceptable.
 For example, testing a group of students for intelligence, with an IQ test, and then
performing the new intelligence test a couple of days later would be perfectly acceptable.

4
Predictive Validity

 Predictive validity is a measure of how well a test predicts abilities. It involves testing a
group of subjects for a certain construct and then comparing them with results obtained at
some point in the future.
 This measures the extent to which a future level of a variable can be predicted from a
current measurement. This includes correlation with measurements made with different
instruments.
 For example, a political poll intends to measure future voting intent.

College entry tests should have a high predictive validity with regard to final exam results. If a
student passes entry test in good marks. His/her future performance can be predicted on the
basis of that entry test. It is assumed that good performance leads to good marks in that
course
Content Validity

 Content validity is the estimate of how much a measure represents every single element
of a construct.
 Content validity requires the use of recognized subject matter experts to evaluate whether
test items assess defined content.
 Content validity is most often addressed in academic and vocational testing, where test
items need to reflect the knowledge actually required for a given topic area (e.g., history)
or job skill (e.g., accounting).
 In clinical settings, content validity refers to the correspondence between test items and
the symptom content of a syndrome.
 In psychometrics, content validity (also known as logical validity) refers to the extent to
which a measure represents all facets of a given construct. For example,
a depression scale may lack content validity if it only assesses the affective dimension of
depression but fails to take into account the behavioral dimension.
 For example, an educational test with strong content validity will represent the subjects
actually taught to students, rather than asking unrelated questions.
 Construct Validity is used to ensure that the measure is actually measure what it is
intended to measure (i.e. the construct), and not other variables. Using a panel of

5
“experts” familiar with the construct is a way in which this type of validity can be
assessed. The experts can examine the items and decide what that specific item is
intended to measure. Students can be involved in this process to obtain their feedback.
 Example: A women’s studies program may design a cumulative assessment of learning
throughout the major. The questions are written with complicated wording and
phrasing. This can cause the test inadvertently becoming a test of reading
comprehension, rather than a test of women’s studies. It is important that the measure is
actually assessing the intended construct, rather than an extraneous factor.
Construct Validity

 Construct validity defines how well a test or experiment measures up to its claims. A
test designed to measure depression must only measure that particular construct, not
closely related ideals such as anxiety or stress. There are two types construct validity.
 Convergent validity tests that constructs that are expected to be related are, in fact,
related.
 Measures of constructs that theoretically should be related to each other are, in fact,
observed to be related to each other.
 E.g., Self Esteem is related to self concept, self appraisal and confidence
Discriminant Validity

 Discriminant validity tests that constructs that should have no relationship do, in fact, not
have any relationship. (also referred to as divergent validity).
 Measures of constructs that theoretically should not be related to each other are, in fact,
observed to not be related to each other.
 E.g., Students face less academic problems if they have high resilience.
Validity Coefficient

 The validity coefficient is calculated as a correlation between the two items being
compared, very typically success in the test as compared with success in the job.
 Validity coefficient (r) greater than .4 considered good
 A validity of 0.6 and above is considered high, which suggests that very few tests give
strong indications of job performance.
 Example: Is typing speed a valid measure of secretary performance?

You might also like