MALABAR COLLEGE OF COMMERCE AND SCIENCE,
MANOOR
DEPT. OF PSYCHOLOGY
PSYCHOLOGICAL ASSESSMENT
Psychometric properties of a test: Reliability, validity and
norms
Psychometric properties are most often expressed quantitatively.
Numerical quantities such as a coefficient or an index represent the
property. The awareness of the different psychometric properties of a
test ensures that the information gained using it will provide a firm
foundation for making the right decisions.
A standardized test is administered and scored in a consistent or
standard manner. They are designed to stabilize the questions,
conditions for administering, scoring procedures, and interpretations.
Standardized testing can consist of true-false, multiple-choice,
authentic assessments or essay questions. It’s possible to shape any
form of assessment into standardized tests.
Here are the three psychometric characteristics that must be
considered when creating or standardizing tests:
Reliability is a fundamental concept in research and measurement that
ensures consistency, stability, and dependability in data collection and
analysis. In both qualitative and quantitative studies, reliability
determines whether a test, tool, or method produces the same results
under consistent conditions
Psychometric reliability is the extent to which test scores are accurate.
A reliable test score is precise and consistent during all the instances
of tests taken. An assessment is considered reliable only if it produces
similar results under variable conditions across multiple testing
instances, numerous test editions, or multiple raters grading the
participant’s responses. Reliability is an essential component of a
perfect assessment test.
types of reliability
Test-Retest Reliability
This type assesses the consistency of results when the same test is
administered to the same subjects at different times.
• Example: A psychological questionnaire administered to
participants twice, two weeks apart, yielding similar scores
indicates high test-retest reliability.
• Purpose: Evaluates the stability of a test over time.
Parallel-Forms Reliability
This type examines the equivalence of two different forms of a test
designed to measure the same construct.
• Example: A teacher creates two versions of a math test, and
students’ scores are consistent across both versions,
demonstrating parallel-forms reliability.
• Purpose: Assesses consistency between equivalent test forms.
Split-Half Reliability
This method involves dividing a test into two halves and comparing
the consistency of results between the halves.
• Example: A 20-question exam split into two sets of 10
questions yields similar scores for both halves.
• Purpose: Tests the consistency of items within a single
instrument.
Reliability is a cornerstone of research and measurement, ensuring
consistent and dependable results across studies. By understanding its
types—such as test-retest, inter-rater, and internal consistency—
researchers can select appropriate methods and tools to assess
reliability. While challenges exist, adopting standardized procedures,
refining instruments, and employing statistical methods can enhance
the reliability of any research process. Ultimately, reliable data forms
the foundation for valid conclusions and evidence-based decision-
making.
validity
Validity is the degree to which the test measures what it claims to
measure. As per the definition put forward by The Standards for
Educational and Psychological Testing (2014), validity is the ‘degree
to which evidence and theory support the interpretations of test scores
for proposed uses of tests.’
Even though an assessment might be reliable, it may fail to provide
the correct measure of the test-takers’ traits if it is not valid. Since the
assessor will make decisions about the test takers based on the
assessment, the validity inferred from it is crucial. Four types of
validity can be measured, and all four should be considered to ensure
a test is valid.
The four types of validity are:
Construct validity
Construct validity evaluates whether a measurement tool really
represents the thing we are interested in measuring. It’s central to
establishing the overall validity of a method.
A construct refers to a concept or characteristic that can’t be directly
observed, but can be measured by observing other indicators that are
associated with it.
Constructs can be characteristics of individuals, such as intelligence,
obesity, job satisfaction, or depression; they can also be broader
concepts applied to organizations or social groups, such as gender
equality, corporate social responsibility, or freedom of speech.
Construct validity is about ensuring that the method of measurement
matches the construct you want to measure. If you develop a
questionnaire to diagnose depression, you need to know: does the
questionnaire really measure the construct of depression? Or is it
actually measuring the respondent’s mood, self-esteem, or some other
construct?
Content validity
Content validity assesses whether a test is representative of all aspects
of the construct.
To produce valid results, the content of a test, survey or measurement
method must cover all relevant parts of the subject it aims to measure.
If some aspects are missing from the measurement (or if irrelevant
aspects are included), the validity is threatened and the research is
likely suffering from omitted variable bias.
Face validity
Face validity considers how suitable the content of a test seems to be
on the surface. It’s similar to content validity, but face validity is a
more informal and subjective assessment. As face validity is a
subjective measure, it’s often considered the weakest form of validity.
However, it can be useful in the initial stages of developing a method.
Criterion validity
Criterion validity evaluates how well a test can predict a concrete
outcome, or how well the results of your test approximate the results
of another test. A criterion variable is an established and effective
measurement that is widely considered valid, sometimes referred to as
a “gold standard” measurement. Criterion variables can be very
difficult to find.
To evaluate criterion validity, we can calculate
the correlation between the results of your measurement and the
results of the criterion measurement. If there is a high correlation, this
gives a good indication that your test is measuring what it intends to
measure.
What are norms
Norms refer to a sample of test-takers who represent the intended
population for the assessment.
Norming helps the test designer understand the group they are
assessing and identify what is considered normal within the target
group. For example, a test that is designed to evaluate the coding
skills of an experienced programmer in Java and will be used to hire
coders with five years of experience will have a norming group
comprising Java programmers with five years of experience