Overview of Psychological Testing
Overview of Psychological Testing
HISTORICAL PERSPECTIVE
The origins of testing are not recent. Evidence suggests that the Chinese had a relatively
sophisticated civil service testing program more than 4000 years ago. Every third year in China,
oral examinations were given to help determine work evaluations and promotion decisions.
By the Han Dynasty (206 B.C.E. to 220 C.E.), the use of test batteries (two or more
tests used in conjunction) was quite common. These early tests related to such diverse topics as
civil law, military affairs, agriculture, revenue, and geography. Tests had become quite well
developed by the Ming Dynasty (1368–1644 C.E.). During this period, a national multistage
testing program involved local and regional testing centers equipped with special testing booths.
Those who did well on the tests at the local level went on to provincial capitals for more
extensive essay examinations. After this second testing, those with the highest test scores went
on to the nation’s capital for a final round. Only those who passed this third set of tests were
eligible for public office.
The Western world most likely learned about testing programs through the Chinese. Reports by
British missionaries and diplomats encouraged the English East India Company in 1832 to
copy the Chinese system as a method of selecting employees for overseas duty. Because testing
programs worked well for the company, the British government adopted a similar system of
testing for its civil service in 1855. After the British endorsement of a civil service testing system,
the French and German governments followed suit. In 1883, the U.S. government established the
American Civil Service Commission, which developed and administered competitive
examinations for certain government jobs. The impetus of the testing movement in the Western
world grew rapidly at that time.
An important step toward understanding individual differences came with the publication of
Charles Darwin’s highly influential book, The Origin of Species, in 1859. According to
Darwin’s theory, higher forms of life evolved partially because of differences among individual
forms of life within a species. Given that individual members of a species differ, some possess
characteristics that are more adaptive or successful in a given environment than are those of
other members. Darwin also believed that those with the best or most adaptive characteristics
survive at the expense of those who are less fit and that the survivors pass their characteristics on
to the next generation. Through this process, he argued, life has evolved to its currently complex
and intelligent levels.
Sir Francis Galton, a relative of Darwin’s, soon began applying Darwin’s theories to the study
of human beings. Given the concepts of survival of the fittest and individual differences, Galton
set out to show that some people possessed characteristics that made them more fit than others, a
theory he articulated in his book Hereditary Genius, published in 1869. Galton (1883)
subsequently began a series of experimental studies to document the validity of his position. He
concentrated on demonstrating that individual differences exist in human sensory and motor
functioning, such as reaction time, visual acuity, and physical strength. In doing so, Galton
initiated a search for knowledge concerning human individual differences, which is now one of
the most important domains of scientific psychology. Galton’s work was extended by the U.S.
psychologist James McKeen Cattell, who coined the term mental test (Cattell, 1890). Cattell’s
doctoral dissertation was based on Galton’s work on individual diff erences in reaction time. As
such, Cattell perpetuated and stimulated the forces that ultimately led to the development of
modern tests.
A second major foundation of testing can be found in experimental psychology and early
attempts to unlock the mysteries of human consciousness through the scientific method. Before
psychology was practiced as a science, mathematical models of the mind were developed, in
particular those of J. E. Herbart. Herbart eventually used these models as the basis for
educational theories that strongly influenced 19th-century educational practices. Following
Herbart, E. H. Weber attempted to demonstrate the existence of a psychological threshold, the
minimum stimulus necessary to activate a sensory system. Then, following Weber, G. T. Fechner
devised the law that the strength of a sensation grows as the logarithm of the stimulus intensity.
Wilhelm Wundt, who set up a laboratory at the University of Leipzig in 1879, is credited with
founding the science of psychology, following in the tradition of Weber and Fechner (Hearst,
1979). Wundt was succeeded by E. [Link], whose student, G. Whipple, recruited L. L.
Thurstone. Whipple provided the basis for immense changes in the field of testing by conducting
a seminar at the Carnegie Institute in 1919 attended by Thurstone, E. Strong, and other early
prominent U.S. psychologists. From this seminar came the Carnegie Interest Inventory and later
the Strong Vocational Interest Blank. Later in this book we discuss in greater detail the work of
these pioneers and the tests they helped to develop. Thus, psychological testing developed from
at least two lines of inquiry: one based on the work of Darwin, Galton, and Cattell on the
measurement of individual differences, and the other (more theoretically relevant and probably
stronger) based on the work of the German psychophysicists Herbart, Weber, Fechner, and
Wundt. Experimental psychology developed from the latter. From this work also came the idea
that testing, like an experiment, requires rigorous experimental control. Such control, as you will
see, comes from administering tests under highly standardized conditions. The efforts of these
researchers, however necessary, did not by themselves lead to the creation of modern
psychological tests. Such tests also arose in response to important needs such as classifying and
identifying the mentally and emotionally handicapped. One of the earliest tests resembling
current procedures, the Seguin Form Board Test (Seguin, 1866/1907), was developed in an effort
to educate and evaluate the mentally disabled. Similarly, Kraepelin (1912) devised a series of
examinations for evaluating emotionally impaired people.
An important breakthrough in the creation of modern tests came at the turn of the 20th century.
The French minister of public instruction appointed a commission to study ways of identifying
intellectually subnormal individuals in order to provide them with appropriate educational
experiences. One member of that commission was Alfred Binet. Working in conjunction with the
French physician T. Simon, Binet developed the first major general intelligence test. Binet’s
early effort launched the first systematic attempt to evaluate individual differences in human
intelligence
World War I- The testing movement grew enormously in the United States because of the
demand for a quick, efficient way of evaluating the emotional and intellectual functioning of
thousands of military recruits in World War I. The war created a demand for large-scale group
testing because relatively few trained personnel could evaluate the huge influx of military
recruits. However, the Binet test was an individual test. Shortly after the United States became
actively involved in World War I, the army requested the assistance of Robert Yerkes, who was
then the president of the American Psychological Association. Yerkes headed a committee of
distinguished psychologists who soon developed two structured group tests of human abilities:
the Army Alpha and the Army Beta. The Army Alpha required reading ability, whereas the
Army Beta measured the intelligence of illiterate adults. World War I fueled the widespread
development of group tests. About this time, the scope of testing also broadened to include tests
of achievement, aptitude, interest, and personality. Because achievement, aptitude, and
intelligence tests overlapped considerably, the distinctions proved to be more illusory than real.
Even so, the 1916 Stanford-Binet Intelligence Scale had appeared at a time of strong demand and
high optimism for the potential of measuring human behavior through tests. World War I and the
creation of group tests had then added momentum to the testing movement. Shortly after the
appearance of the 1916 Stanford-Binet Intelligence Scale and the Army Alpha test, schools,
colleges, and industry began using tests. It appeared to many that this new phenomenon, the
psychological test, held the key to solving the problems emerging from the rapid growth of
population and technology.
The Period of Rapid Changes in the Status of Testing
The 1940s saw not only the emergence of a whole new technology in psychological testing but
also the growth of applied aspects of psychology. The role and significance of tests used in
World War I were reaffirmed in World War II. By this time, the U.S. government had begun to
encourage the continued development of applied psychological technology. As a result,
considerable federal funding provided paid, supervised training for clinically oriented
psychologists. By 1949, formal university training standards had been developed and accepted,
and clinical psychology was born. Other applied branches of psychology—such as industrial,
counseling, educational, and school psychology—soon began to blossom. One of the major
functions of the applied psychologist was providing psychological testing.
The Current Environment
During the 1980s, 1990s, and 2000s several major branches of applied psychology emerged and
flourished: neuropsychology, health psychology, forensic psychology, and child psychology.
Because each of these important areas of psychology makes extensive use of psychological tests,
psychological testing again grew in status and use. Neuropsychologists use tests in hospitals and
other clinical settings to assess brain injury. Health psychologists use tests and surveys in a
variety of medical settings. Forensic psychologists use tests in the legal system to assess mental
state as it relates to an insanity defense, competency to stand trial or to be executed, and
emotional damages. Child psychologists use tests to assess childhood disorders.
Test Construction
Test construction is the process of developing and creating psychological or educational tests.
These tests are designed to measure specific traits, abilities, knowledge, or characteristics in
individuals. Proper test construction is essential to ensure that the test is valid, reliable, and fair.
Steps in test construction-
1) Planning of the Test (Objectives): In this initial step, test constructors define the purpose and
objectives of the test. They determine what knowledge, skills, or abilities the test should assess
and for whom it is intended. Clear objectives help guide the entire test development process.
2) Writing Items of the Test: This step involves crafting the actual test questions or items. Test
constructors must create items that align with the defined objectives. They should strive for
clarity, relevance, and appropriateness for the target audience. These items can be multiple-
choice, essay questions, true/false statements, or any other format that suits the assessment's
goals.
3) Preliminary Administration: Before the full-scale administration, it is advisable to conduct a
preliminary administration or a pilot test. This helps identify any issues with the items, such as
ambiguous wording, difficulty levels, or time constraints. The results from the preliminary
administration can inform item selection and refinement.
4) Reliability: Reliability refers to the consistency and stability of test scores over time. Test
constructors assess the reliability of the test by using statistical techniques, such as test-retest
reliability or internal consistency (Cronbach's alpha). A reliable test consistently measures what
it is intended to measure.
5) Validity: Validity is the extent to which a test measures what it is supposed to measure.
Establishing validity is a crucial step in test construction. Different types of validity, such as
content, criterion-related, and construct validity, need to be considered and assessed. Validity
evidence can come from expert reviews, correlation with other measures, or experimental data.
6) Norms: Once the test is administered to a representative sample of the target population,
norming is conducted. This process involves collecting data on how the test performs with a
diverse group of test-takers. The resulting norms provide a basis for interpreting individual
scores relative to the group.
7) Preparation of Manual and Reproduction: A test manual is a comprehensive document that
includes detailed information about the test's purpose, administration procedures, scoring, and
interpretation. It also describes the test's reliability, validity, and norms. The manual serves as a
guide for test administrators and users. After finalizing the manual, the test is reproduced for
distribution, ensuring it maintains its quality and integrity.
Item Analysis
Item analysis is a method used to evaluate the quality of individual test questions or items. It
helps test developers and instructors understand how well each question performs and whether
they should keep or revise it. Here's a simple explanation of item analysis:
Calculate Difficulty: Item analysis begins by calculating the difficulty of each question.
Difficulty is the proportion of test-takers who answered the item correctly. A question that most
test-takers answer correctly is considered easy, while one that most answer incorrectly is
considered difficult. It helps identify how well the item discriminates between high and low
performers. The index of difficulty is a concept used to assess the level of difficulty associated
with individual test items or questions.
When constructing a test, it's important to balance the indices of difficulty across items. The test
should include items with varying levels of difficulty to effectively assess the full range of
abilities within the target population.
Calculate Discrimination: Discrimination assesses how well a question differentiates between
high- and low-performing students. It's calculated by comparing the performance of the top-
scoring students (e.g., the top 27%) with the bottom-scoring students (e.g., the bottom 27%).
Questions that high-performing students answer correctly more often than low-performing
students are considered good discriminators.
A positive discrimination index indicates that the item is more likely to be answered correctly by
high scorers, suggesting that the item effectively differentiates between the two groups. Such
items are considered good discriminators. A negative discrimination index suggests that the item
is more likely to be answered correctly by low scorers, indicating a poor item discriminator.
Negative item discrimination can be problematic and may need to be reviewed or revised.
Types of tests
Objectivity: A good test should be objective, meaning that its administration and scoring
should be standardized and free from subjective interpretation. Different examiners or
graders should obtain consistent results. Objective tests often have clear and unambiguous
questions with predetermined answer choices.
Practicality: The test should be practical to administer and score within the given
constraints, including time, resources, and personnel. It should not be overly burdensome or
time-consuming for both administrators and test-takers.
Reliability: Reliability refers to the consistency and stability of test scores. A good test
should produce consistent results when administered multiple times or by different
examiners. It should be free from errors that might lead to variations in scores. Common
measures of reliability include test-retest reliability, internal consistency (Cronbach's alpha),
and inter-rater reliability.
Validity: Validity is the extent to which a test measures what it is intended to measure. A
good test should demonstrate evidence of validity, showing that it effectively assesses the
construct, knowledge, or skill it was designed to measure. Different types of validity, such as
content validity, criterion-related validity, and construct validity, should be considered and
supported.
Norms: A good test should have established norms, which provide a basis for interpreting
individual test scores. Norms are statistics that describe the performance of a representative
sample of the population against which a test-taker's performance can be compared. They
help identify how an individual's score ranks relative to the group, enabling meaningful
comparisons.
Reliability
Reliability refers to the consistency of a measure. A test is considered reliable if we get the same
result repeatedly. For example, if a test is designed to measure a trait (such as introversion), then
each time the test is administered to a subject, the results should be approximately the same.
Types of reliability
1) Test-retest reliability: The test-retest reliability method in research involves giving a
group of people the same test more than once over a set period of time. In this assessment,
the research method and sample group stay the same, but when administering the method
to the group changes. If the results of the test are similar each time the test given to the
sample group, that shows the research method is likely reliable and not influenced by
external factors, like the sample group's mood or the day of the week. Example: Give a
group of college students a survey about their satisfaction with their school's parking lots
on Monday and again on Friday, then compare the results to check the test-retest
reliability.
2) Parallel forms reliability: When using parallel forms reliability to assess research, the
researcher may give the same group of people multiple different types of tests to
determine if the results stay the same when using different research methods. The theory
behind this assessment is that consistent results across research methods ensure each
method is looking for the same information from the group and the group is behaving
similarly for each test. This means the methods are likely reliable because, if they weren't,
the participants in the sample group may behave differently and change the results.
Example: In marketing, the researcher may interview customers about a new product,
observe them using the product and give them a survey about how easy the product is to
use and compare these results as a parallel forms reliability test.
3) Inter-rater reliability: With inter-rater reliability testing, researcher may have multiple
people performing assessments on a sample group and comparing their results to avoid
influencing factors, like an assessor's personal bias, mood or human error. If most of the
results from different assessors are similar, it's likely the research method is reliable and
can produce usable research because the assessors gathered the same data from the group.
This is useful for research methods like observations, interviews and surveys where each
assessor may have different criteria but can still end up with similar research results.
Example: Multiple behavioral specialists may observe a group of children playing to
determine their social and emotional development and then compare notes to check for
inter-rater reliability.
4) Internal consistency reliability: split-half reliability is one of the method in which
researcher can perform this test by splitting a research method, like a survey or test, in
half, delivering both halves separately to a sample group and comparing the results to
ensure the method can produce consistent results. If the results are consistent, then the
results of the research method are likely reliable. Example: You may give a company's
cleaning department a questionnaire about which cleaning products work the best, but
you split it in half and give each half to the department separately and calculate the
correlation to test for split-half reliability.
Validity
One of the greatest concerns when creating a psychological test is whether or not it actually
measures what we think it is measuring.
Types of validity
Norms
Norms are sets of score obtained by whom the test is intended. The scores obtained by these
groups provide a basic for interpreting any individual score. Norm refers to the typical
performance level for a certain group of individuals. Any psychological test with just the raw
score is meaningless until it is supplemented by additional data to interpret it further.
The terms Norm-Referenced and Criterion-Referenced refer to score interpretations. Norm-
referenced refers to standardized tests that are designed to compare and rank test takers in
relation to one another. A criterion-referenced test is designed to measure a student's academic
performance against some standard or criteria. This standard or criteria is predetermined before
students begin the test.
Types of norms
1) Percentiles: They refer to the percentage of people in a standardized sample that are below a
certain set of score. They depict an individual’s position with respect to the sample. Here the
counting begins from bottom, so the higher the percentile the better the rank. For example if a
person gets 97 percentile in a competitive exam, it means 97% of the participants have scored
less than him/her
2) Standard Score: It signifies the gap between the individuals score and the mean depicted as
standard deviation of the distribution. It can be derived by linear or nonlinear transformation of
the original raw scores. They are also known as T and Z scores.
3) Age Norms: To obtain this, we take the mean raw score gathered from all in the common age
group inside a standardized sample. Hence, the 15 year norm would be represented and be
applicable by the mean raw score of students aged 15 years.
4) Grade Norms: It is calculated by finding the mean raw score earned by students in a specific
grade.
Attitude
Attitudes are constructs which are not open to direct observation. Therefore, the common way of
measuring attitudes is to study some aspect of behavior and from that makes inferences about the
attitudes that may be responsible for behavior.
Many different ways of measuring attitudes have been developed, though the more common
techniques make use of self-report questionnaires. In general, an attitude scale or questionnaire
will present the individual with a number of statements to which she/he has to respond and from
these responses the investigator will arrive at some conclusions about the attitudes of that
individual.
Likert (1932) proposed a technique, which makes this process simpler by allowing the
participant to make a range of possible responses, usually in the form of a five-point scale,
ranging from strongly agree, agree, undecided, to disagree, and strongly disagree. With this
method, there is no requirement for the judges to categorize each statement, as the categorization
is built into the scale. Assigning the score from 1 to 5 to each of the responses, scores the Likert
scales. Then the totaling of the score is done, to give a final measure of the individual’s attitude.
The summated rating scale, more commonly known as the Likert scale, is based upon the
assumption that each statement/item on the scale has equal attitudinal value, ‘importance’ or
‘weight’ in terms of reflecting an attitude towards the issue in question. This assumption is also
the main limitation of this scale as statements on a scale seldom have equal attitudinal value.
Thurstone (1929) is credited with having first created the attitude-measurement methodology. He
is considered to be the ‘father’ of attitude scaling. He developed one of the earliest methods for
measuring the attitudes towards religion. It is made up of statements about a particular issue, and
each statement has a numerical value, indicating how favorable or unfavorable it is judged to be.
People check each of the statements to which they agree, and a mean score is computed,
indicating their attitude
Unlike the Likert scale, the Thurstone scale calculates a ‘weight’ or ‘attitudinal value’ for each
statement. The weight (equivalent to the median value) for each statement is calculated on the
basis of rating assigned by a group of judges. Each statement with which respondents express
agreement (or to which they respond in the affirmative) is given an attitudinal score equivalent to
the ‘attitudinal value’ of the statement.
Louis Guttman (1947) developed a scale much like Likert and Thurstone, but he placed extreme
stress on the instrument, measuring only a single trait (a property called unidimensionality, a
single dimension underlies the responses to the scale). The items are ordered hierarchically, such
that, a positive response to one means positive responses to each of the items prior to it.
This scale asks respondents to rate their attitude toward an object, concept, or idea on a set of
bipolar adjectives or pairs of opposite adjectives. For example, respondents might rate their
feelings toward a product on a scale from "Good" to "Bad," "Happy" to "Sad," or "Exciting" to
"Boring."
Aptitude test
Aptitude tests are assessments designed to measure an individual's inherent or potential ability in
a specific area or domain. Two common types of aptitude tests are "tests of special abilities" and
"differential aptitude tests."
Tests of Special Abilities: These are aptitude tests that focus on assessing a person's aptitude or
potential in a particular skill or area. Special abilities tests are designed to identify innate talents
or skills that an individual may possess. For example, tests of special abilities might measure a
person's aptitude for music, art, mathematics, or mechanical skills. These tests are often used in
educational and vocational settings to help individuals understand their strengths and interests.
Differential Aptitude Tests: These tests are designed to assess an individual's potential for
success in various areas, often related to academic or career performance. Differential aptitude
tests are used to predict how a person might perform in different domains and can provide
insights into career or educational paths.
Achievement test
Achievement tests are assessments designed to measure what a person has learned or achieved in
a specific subject, skill, or domain. There are two main types of achievement tests: general
achievement tests and special achievement tests.
General achievement tests are designed to assess a person's overall knowledge and skills in a
broad area or subject, such as mathematics, reading, science, or language arts. These tests
provide a comprehensive evaluation of a person's performance in a particular subject area.
Use: General achievement tests are commonly used in educational settings to determine a
student's level of proficiency in a specific subject, to track academic progress, and to make
educational decisions, such as grade placement and curriculum adjustments.
Special achievement tests are more specific and focused assessments designed to measure a
person's knowledge or skills in a particular, well-defined area or domain. These tests are often
tailored to assess proficiency in a specific subtopic within a broader subject.
Use: Special achievement tests are used to evaluate a person's expertise or proficiency in a
narrow subject area. They can be employed in educational settings to assess mastery of specific
content, in professional contexts to certify expertise, or in research to gather data on specific
skills.
Personality test