DEFINITION OF PSYCHOLOGICAL TEST
• Psychological test was defined by Gregory (2010, page 16) as “a standardised procedure for sampling
behavior and describing it with categories or scores”.
• Psychological test can be described as measurement of sample of behaviour that is standardised and
objective (Anastasi, 1969).
• Kaplan and Saccuzzo (2013) explained psychological test a device or technique used in quantification of
behaviour that helps in not only understanding behaviour but also to predicting it.
• Cohen and Swerdlik (2010, page 2), defined psychological testing “ as the process of measuring psychology-
related variables by means of devices or procedures designed to obtain a sample of behavior”
There are some of the assumptions in this regard
1) A psychological test needs to be valid or should measure what it is supposed to measure.
2) It should be reliable or consistent.
3) It should be objective and it is assumed that the individuals taking the test, understand the test items in a similar manner.
4) It is assumed that the individuals taking the test will answer the test and will be able to accurately express their feeling in
that regard.
5) It is assumed that individuals will answer the items honestly. Though there is always a possibility that individuals are
answering in a certain way due to social desirability.
6) Error variance is assumed to occur due to administrator ( bias, expectations), test taker (anxiety, fatigue) as well as testing
conditions (temperature, distractions
CHARACTERISTICS OF A GOOD PSYCHOLOGICAL TEST
1) Psychological tests are objective in nature: Any good psychological test needs to be objective and not subjective. There
should be no place for any kind of bias. An objective psychological test also denotes that it is valid and reliable.
2) Validity of a psychological test: The next characteristic that a good psychological test should possess is validity. Validity can
be explained as the ability of the test to measure what it is supposed to measure. A weighing machine is a valid tool to
measure weight and it is not valid to measure length.
3) Reliability of a psychological test: A good psychological test is also reliable or consistent. For example, if you measure
length of a table with a ruler on certain day and if you measure the length of the same table with same ruler after six
months, the length obtained in centimetres will remain same, thus indicting that this ruler is reliable
4) A good psychological test will have discriminant feature: The test should be able to denote any difference between one
individual from the other on a given aspect or variable. For example, if two individuals differ in their music aptitude, the the
test should be able to differentiate between the two on this aptitude.
5) A good test will be comprehensive: This denotes that the test measures all the dimensions or aspects of the construct
that it measures
TYPES OF TESTS
A) On basis of administration
• Individual test: Tests that are administered on a single individual. For example, Wechsler Adult Intelligence Scale (WAIS),
Stanford-Binet Intelligence Scale (SB), Bhatia battery. Group test: Such tests can be administered to a group of
individuals at the same time. For example, NEO PI and Minnesota Multiphasic Personality Inventory
• Speed test: A speed test constitutes items that are of same difficulty level, however a certain time period is provided to
complete the test.
• Power test: A power test constitutes items that increase gradually in terms of their difficult level. Though there is no
time limit to complete the test. Verbal test: A paper pencil test can be termed as a verbal test where the items are
mentioned using language. For example: 16 PF and Eysenck’s Personality Inventory.
• Non-verbal test: In this type of test certain figures and symbols are used. For example, Raven’s Progressive Matrices. In
this the language may be used only to provide instructions to the individual taking the test.
• Performance test: In performance test, the individual taking the test has to perform certain tasks. For example:
Alexander’s passalong test and Koh’s block design test.
• Objective tests: In objective tests, the individual will choose from certain correct answers that are decided in advance.
This avoids any subjectivity on behalf of the scorer. The responses could be in terms of true or false or multiple choices
or even a rating scale like Likert scale or Thurston’s scale may be used. For example: NEO PI.
• Projective Tests: These are subjective in nature. Here, the test taker may be asked to respond to certain semi-
structured or unstructured stimuli. The responses are then to be interpreted by the administrator, where subjectivity
may creep in. Examples of projective tests are Rorschach Inkblot test, Somatic Inkblot Series, Sentence Completion
Test, Thematic Apperception Test and Children’s Apperception Test.
B) On basis of purpose
• Intelligence tests: There are various intelligence tests that are used to measure intelligence of individuals. Intelligence
can be described as one’s ability to adjust and cope with the environment. Binet and Simon (1960) defined intelligence
as an individual’s capacity to make adequate judgements, carry out reasoning and ability to comprehend. Wechsler
(1944, page 3) defined intelligence as “the aggregate or global capacity of the individual to act purposefully, to think
rationally, and to deal effectively with his environment”. These tests are often used in educational and clinical set ups.
Examples of intelligence test are Wechsler Intelligence Scale for Children (WISC), Stanford-Binet Intelligence Scale (SB),
Bhatia battery.
• Personality tests: These are used to measure personality of individuals. Larsen and Buss (2018) defined personality as
a collection of psychological traits and mechanisms that are stable and organised and that have an influence an
individual’s interaction and also has an impact on how he/she modifies his/ her physical, social and psychological
environment. It can also be explained as differences amongst individuals with regard to their patterns of thinking,
feeling and the way they behave (American Psychological Association, 2019). Personality tests are used widely in
varied setups including clinical, educational, counselling, industrial and organisational setup and so on. Examples of
personality test are Eysenck’s Personality Inventory, Thematic Apperception Test (TAT), Somatic Inkblot Series (SIS).
• Aptitude tests: There are tests that measure the potential/ abilities possessed by an individual in certain area. These
find their application in schools and even in industrial set up for selection [Link] denote whether a person will
be able to perform effectively if he/ she is given training in that area. For instance, a person with aptitude for dance
or music will do well in the area if given training. Examples of aptitude tests are Differential Aptitude Test, Seashore
Musical Aptitude Test. Interest inventories: These measure interests of individuals. Interest is important as aptitude in
making career decisions and thus these tests are also used often in educational setup. Example of interest inventory
is Vocational Interest Inventory.
• Attitude tests: These tests measure attitude of an individual towards events, other individuals, objects and so on.
Often in attitude tests, Thurston and Likert scales are used. These could measure attitude towards women, health
and so on. Achievement tests: There are also tests that measure achievement of individuals. They mainly test an
individual’s learning in certain academic area. Such tests are often used in educational setup. Academic achievement
test and Mathematics Achievement Test are examples of achievement tests
RELIABILITY
Reliability means consistency with which the instrument yields similar results. Reliability concerns the ability of
different researchers to make the same observations of a given phenomenon if and when the observation is
conducted using the same method(s) and procedure(s).
The reliability of measuring instruments can be improved by two ways.
i) By standardizing the conditions under which the measurement takes place i.e. we must ensure that external
sources of variation such as boredom, fatigue etc., are minimized to the extent possible to improve the stability
aspect.
ii) By carefully designing directions for measurement with no variation from group to group, by using trained and
motivated persons to conduct the research and also by broadening the sample of items used to improve equivalence
aspect
The three basic methods for establishing the reliability of empirical measurements are:
1) Test - Retest Method :One of the easiest ways to estimate the reliability of empirical measurements is by the test -
retest method in which the same test is given to the same people after a period of time. Two weeks to one month is
commonly considered to be a suitable interval for many psychological tests. The reliability is equal to the correlation
between the scores on the same test obtained at two points in time. If one obtains the same results on the two
administrations of the test, then the test – retest reliability coefficient will be 1.00. But, invariably, the correlation of
measurements across time will be less than perfect. This occurs because of the instability of measures taken at multiple
points in time. For example, anxiety, motivation and interest may be lower during the second administration of the test
simply because the individual is already familiar with it.
Advantages
• This method can be used when only one form of test is available.
• Test – retest correlation represent a naturally appealing procedure. Limitations • Researchers are often able to obtain
only a measure of a phenomenon at a single point in time. • Expensive to conduct test and retest and some time
impractical as well.
• Memory effects lead to magnified reliability estimates. If the time interval between two measurements is short, the
respondents will remember their early responses and will appear more consistent than they actually are.
• Require a great deal of participation by the respondents and sincerity, devotion by the research worker. Because,
behaviour changes and personal characteristics may likely to influence the re-test as they are changing from day to day.
2) Alternative Form Method/Equivalent Form/Parallel Form
The alternative form method which is also known as equivalent / parallel form is used extensively in education, extension
and development research to estimate the reliability of all types of measuring instruments. It also requires two testing
situations with the same people like test- retest method. But it differs from test – retest method on one very important
regard i.e., the same test is not administered on the second testing, but an alternate form of the same test is
administered. Thus two equivalent reading tests should contain reading passages and questions of the same difficulty.
But the specific passages and questions should be different i.e., approach is different. It is recommended that the two
forms be administered about two weeks apart, thus allowing for day –to- day fluctuations in the person to occur. The
correlation between two forms will provide an appropriate reliability coefficient.
Advantages
• The use of two parallel tests forms provides a very sound basis for estimating the precision of a psychological or
educational test
• Superior to test- retest method, because it reduces the memory related inflated reliability. Limitations
• Basic limitation is the practical difficulty of constructing alternate forms of two tests that are parallel.
• Requires each person’s time twice.
• To administer a secondary separate test is often likely to represent a somewhat burdensome demand upon available
resources.
3) Split-Half Method Split - half method is also a widely used method of testing reliability of measuring instrument for
its internal consistency. In split-half method, a test is given and divided into halves and are scored separately, then the
score of one half of test are compared to the score of the remaining half to test the reliability. In split-half method,
1st-divide test into halves. The most commonly used way to do this would be to assign odd numbered items to one
half of the test and even numbered items to the other, this is called, Odd-Even reliability.
Advantages
• Both, the test – retest and alternative form methods require two test administrations with the same group of
people. In contrast the split –half method can be conducted on one occasion.
• Split-half reliability is a useful measure when impractical or undesirable to assess reliability with two tests or to have
two test administrations because of limited time or money. Limitations
• Alternate ways of splitting the items results in different reliability estimates even though the same items are
administered to the same individuals at the same time.
VALIDITY
• According to Goode and Hatt, a measuring instrument (scale, test etc) possesses validity when it actually measures
what it claims to measure.
• There are three types of validity: i) Content validity ii) Criterion validity (Predictive validity and Concurrent validity) iii)
Construct validity
Content Validity
The term content validity is used, since the analysis is largely in terms of the content. Content validity is the representative
ness or sampling adequacy of the content. Consider a test that has been designed to measure competence in using the
English language. How can we tell how well the test in fact measures that achievement? First we must reach at some
agreement as to the skills and knowledge that comprise correct and effective use of English, and that have been the
objectives of language instruction. Then we must examine the test to see what skills, knowledge and understanding it calls
for. Finally, we must match the analysis the test content against of course content and instrumental objectives, and see
how well the former represents the latter. If the test represents the objectives, which are the accepted goals for the
course, then the test is valid for use
Criterion Validity
The two types of criterion validities are predictive validity and concurrent validity.
They are much alike and with some exceptions, they can be considered the same, because they differ only in the time
dimension. They are characterized by prediction and by checking the measuring instrument either now or in future against
some outcome. Example: A test that help researcher / teacher to distinguish between students who can study by themselves
after attending the class and those who are in need of extra and special coaching, is said to have concurrent validity. The test
distinguishes individually who differ in their present status. On the other hand, the investigator may wish to predict the
percentage of passes during the final examination for that particular period. The adequacy of the test for distinguishing
individuals who differ in the future may be called as predictive validity.
Predictive Validity Vs. Concurrent Validity
Predictive validity concerns a future criterion which is correlated with the relevant measure. Example: Tests used for selection
purposes in different occupations are, by nature, concerned with predictive validity. Thus a test used to screen applications for
the post of ‘health extension and development workers’ could be validated by correlating their test scores with future
performance in fulfilling the duties associated with health extension work.
Concurrent criterion is assessed by correlating a measure and the criterion at the same point in time. Example: A verbal
report of voting behaviour could be correlated with participation in an election, as revealed by official voting records
Construct Validity
Both content and criterion validities have limited usefulness for assessing the validity of empirical measures of theoretical
concepts employed in extension and development studies. In this context, construct validity must be investigated whenever
no criterion or universe of content is accepted as entirely adequate to define the quality to be measured.
Examination of construct validity involves validation not only of the measuring instrument but of the theory underlying it. If
the predictions are not supported, the investigator may have no clear guide as to whether the shortcoming is in the
measuring instrument or in the theory.
Construct validation involves three distinct steps. a) specify the theoretical relationship between the concepts themselves
b) examine the empirical relationship between the measures of the concepts c) interpret the empirical evidence in terms of
how it clarifies the construct validity of the particular measure