Stanford-Binet IQ
10 Chronological age
Heritability
Flynn effetc
Testing and Individual
Overview
Differences We all take many standardized tests and receive scores that tell us how we
perform. Given the world in which we have grown up, it is almost
unimaginable that there ever could have been a time during which people’s
Learning Objectives mental abilities were not measured and tested. Francis Galton was a pioneer
In this chapter, you will learn about: in the study of human intelligence and testing. He initiated the use of
→ Standardization and norms surveys for collecting data, and he both developed and applied statistics for
→ Reliability and validity analyzing that data. In this chapter, we will review what makes for a good
test, how to interpret your scores on such tests, and what different kinds of
→ Theories of intelligence
tests exist. Then we will focus on one of the most tested characteristics of
→ Intelligence tests all, intelligence.
→ Nature vs. nurture: intelligence
Standardization and Norms
As a student, you probably take a lot of tests. Although most teachers are
Key Terms experienced with creating tests, psychometricians are psychologists who
Standardized specialize in making standardized tests. When we say that a test is
Reliability standardized, we mean that the test items have been piloted on a similar
Split-half reliability population of people as those who are meant to take the test (the
standardization sample) and that achievement norms have been established.
Test-retest reliability
For standardized tests, like Advanced Placement tests, we want to be
Validity confident that scoring a 5 is indicative of a similar level of mastery on each
Predictive validity exam.
Construct validity The purpose of tests is to distinguish among people. Therefore, test
Aptitude tests questions that virtually everyone answers correctly as well as questions that
Achievement tests almost no one can answer correctly are discarded. Such items do not
Intelligence provide information that differentiates among the people taking the test.
Fluid intelligence
Crystallized intelligence Reliability and Validity
Mental age In order for us to have any faith in the meaning of a test score, we must
believe the test is both reliable and valid. Reliability refers to the Just as several different kinds of reliability exist, several different kinds
repeatability or consistency of the test as a means of measurement. For of validity exist. Face validity refers to a superficial measure of accuracy. A
instance, if you were to take a test three times that purportedly determined test of cake-baking ability has high face validity if you are looking for a
what career you should pursue and if on each occasion you received chef but has low face validity if you are in the market for a doctor. Face
radically different recommendations, you might question the reliability of validity is a type of content validity. Content validity refers to how well a
the test. Similarly, if you scored 115, 92, and 133 on three different measure reflects the entire range of material it is supposed to be testing. If
administrations of the same IQ (intelligence quotient) test, you would have one really wanted to design a test to find a good chef (as opposed to a cake
little reason to believe your intelligence had been accurately measured. baker), a test that requires someone to create an entrée and whip up a salad
The reliability of a test can be measured in several different ways. Split- dressing in addition to baking a cake would have greater content validity.
half reliability involves randomly dividing a test into two different sections Another kind of validity is criterion-related validity. Tests may have two
and then correlating people’s performances on the two halves. The closer kinds of criterion-related validity: concurrent and predictive. Concurrent
the correlation coefficient is to +1, the greater the split-half reliability of the validity measures how much of a characteristic a person has now. For
test. Many tests are available in several equivalent forms. The correlation example, is a person a good chef now? Predictive validity is a measure of
between performance on the different forms of the test is known as future performance. For example, does a person have the qualities that
equivalent-form reliability. Finally, test-retest reliability refers to the would enable him or her to become a good chef?
correlation between a person’s score on one administration of the test with Finally, construct validity is thought to be the most meaningful kind of
the same person’s score on a subsequent administration of the test. validity. If an independent measure already exists that has been established
A test is valid when it measures what it is supposed to measure. Validity to identify those who will make fine chefs and love their work, we can
is often referred to as the accuracy of a test. A personality test is valid if it correlate prospective chefs’ performance on this measure with their
truly measures an individual’s personality, and the career inventory performance on any new measure. The higher the correlation, the more
described above is valid only if it truly measures for which jobs a person is construct validity the new measure has. The limitation, of course, is the
best suited. The latter example should serve to highlight an important point: difficulty in creating any measure that we believe is perfectly valid in the
a test cannot be valid if it is not reliable. If subsequent administrations of the first place.
career inventory yield grossly disparate results for the same person, it
clearly does not accurately reflect a person’s vocational strengths or Table 10.1 Reliability Versus Validity in Archery
interests. However, a test may be reliable without being valid. Even if
someone’s performance on the test repeatedly indicates that he or she Neither An archer always misses the target: sometimes shooting
should be a chef and thus is reliable, if the person hates to cook, the test is reliable too high, sometimes too low, sometimes too far to the
not a valid measure of his or her interest. (See Table 10.1 for a comparison nor valid left, and sometimes too far to the right.
of reliability and validity.) Reliable An archer always misses the target, but the arrows
but not consistently go just over the right side of the top of the
TIP valid target.
Reliability and validity are important terms for you to know. The psychological Both An archer always puts the arrow in or near the bull’s-eye.
meaning ascribed to these two terms may differ somewhat from how they are used by reliable
the general population. Reliability refers to a test’s consistency, and validity refers to and valid
a test’s accuracy.
Types of Tests are provided to the group, and then people are given a certain amount of
time to complete the various sections of the test. Group tests are less
Two common types of tests are aptitude tests and achievement tests. expensive to administer and are thought to be more objective than
Aptitude tests measure ability or potential, while achievement tests individual tests. Individual tests involve greater interaction between the
measure what one has learned or accomplished. For instance, any examiner and the examinee.
intelligence test is supposed to be an aptitude test. These tests are made to
express someone’s potential, not their current level of achievement. Theories of Intelligence
Conversely, most, if not all, the tests you take in school are supposed to be
achievement tests. They are meant to indicate how much you have learned Although intelligence is a commonly used term, it is an extremely difficult
in a given subject. However, making a test that exclusively measures one of concept to define. Typically, intelligence is defined as the ability to gather
these qualities is virtually impossible. Whatever one’s aptitude for a and use information in productive ways. However, we will not present any
particular field or skill, one’s experience affects it. Someone who has had a one correct definition of intelligence because nothing that approaches a
lot of schooling will score better on a test of mathematics aptitude than consensus has been achieved. Rather, we will present brief summaries of
someone who might have an equally great potential to be a mathematician some of the most widely known theories of intelligence. Table 10.2
but who has never had any formal training in math. Similarly, two people compares three of these theories.
who have achieved equally in learning biology will not necessarily score the Many psychologists differentiate between fluid intelligence and
same on an achievement test. If one has far greater test-taking aptitude, she crystallized intelligence. Fluid intelligence refers to our ability to solve
or he will likely outscore the other. abstract problems and pick up new information and skills, while
crystallized intelligence involves using knowledge accumulated over time.
TIP Although fluid intelligence seems to decrease as adults age, research shows
that crystallized intelligence holds steady or may even increase. For
Even though it is essentially impossible to create a pure aptitude or pure achievement instance, a 20-year-old may be able to learn a computer language more
test, tests that purport to measure aptitude seek to measure someone’s ability or
potential, whereas achievement tests seek to measure how much of a body of
quickly than a 60-year-old, whereas the older person may well have the
material someone has learned. advantage on a vocabulary test or an exercise dependent on wisdom.
One fundamental issue of debate is whether intelligence refers to a single
ability, a small group of abilities, or a wide variety of abilities. Charles
Distinguishing between speed and power tests is also possible. Speed Spearman argued that intelligence could be expressed by a single factor. He
tests generally consist of a large number of questions asked in a short used factor analysis, which is a statistical technique that measures the
amount of time. The goal of a speed test is to see how quickly a person can correlations between different items, to conclude that underlying the many
solve problems. Therefore, the amount of time allotted should be different specific abilities, s, that people regard as types of intelligence is a
insufficient to complete the problems. The goal of a power test is to gauge single factor that he named g for general.
the difficulty level of problems an individual can solve. Power tests consist Howard Gardner subscribes to the idea of multiple intelligences. Unlike
of items of increasing difficulty levels. Examinees are given sufficient time many other researchers, however, the kinds of intelligences that this
to work through as many problems as they can since the goal is to determine contemporary researcher has named thus far encompass a large range of
the ceiling difficulty level, not their problem-solving speed. human behavior. Three of Gardner’s multiple intelligences—linguistic,
Finally, some tests are group tests while others are individual tests. Group logical-mathematical, and spatial—fall within the bounds of qualities
tests are administered to many people at a time. Interaction between the traditionally labeled as intelligences. To that list Gardner has added musical,
examiner and the people taking the test is minimal. Generally, instructions bodily-kinesthetic, intrapersonal, interpersonal, and naturalist intelligences.
Musical intelligence, as one might suspect, includes the ability to play an Recently there has been a lot of discussion of EQ, which is also known as
instrument or compose a symphony. A dancer or an athlete would have a lot emotional intelligence. One of the main proponents of EQ is Daniel
of bodily-kinesthetic intelligence as would a hunter. Intrapersonal Goleman. EQ roughly corresponds to Gardner’s notions of interpersonal and
intelligence refers to one’s ability to understand oneself. People who can intrapersonal intelligence. Researchers who argue for the importance of EQ
persevere without becoming discouraged or who can differentiate between point out that the people with the highest IQs are not always the most
situations in which they will be successful and those that may simply successful people. They contend that both EQ and IQ are needed to succeed.
frustrate them have intrapersonal intelligence. Interpersonal intelligence, on
the other hand, corresponds to a person’s ability to get along with and be Table 10.2 Theories of Intelligence
sensitive to others. Successful psychologists, teachers, and salespeople
typically have a lot of interpersonal intelligence. Finally, naturalist Spearman Intelligence can be measured by a single, general ability
intelligence is found in people gifted at recognizing and organizing the (g).
things they encounter in the natural environment. Such people would be Gardner Theory of multiple intelligences—the term “intelligence”
successful in fields such as biology and ecology. should be applied to a wide variety of abilities including
Robert Sternberg is another contemporary researcher who has offered a kinesthetic, musical, interpersonal, intrapersonal,
somewhat nontraditional definition of intelligence. Sternberg’s triarchic naturalistic, verbal, spatial, and mathematical.
theory holds that three types of intelligence exist. Componential, or analytic, Sternberg Triarchic theory of intelligence—people can be
intelligence involves the skills traditionally thought of as reflecting intelligent in different ways; they can evidence analytic,
intelligence. Most of what we are asked to do in school involves this type of practical, and creative intelligence.
intelligence: the ability to compare and contrast, explain, and analyze. The
second type, experiential or creative intelligence, focuses on people’s ability
to use their knowledge and experiences in new and innovative ways. Rather Intelligence Tests
than comparing the different definitions of intelligence that others have Not surprisingly, the ongoing debate over what constitutes intelligence
offered, someone with this type of intelligence might prefer to come up with makes constructing an assessment particularly difficult. Two widely used
his or her own theory of what constitutes intelligence. The third kind of individual tests of intelligence are the Stanford-Binet and the Wechsler.
intelligence Sternberg discusses is contextual or practical intelligence. Alfred Binet was a Frenchman who wanted to design a test that would
People with this type of intelligence are what we consider street-smart; they identify which children needed special attention in schools. His purpose was
can apply what they know to real-world situations. not to rank or track children but, rather, to improve the children’s education
This last aspect of Sternberg’s theory, the idea of practical intelligence, by finding a way to tailor it better to their specific needs. Binet came up
raises another important and unresolved issue in the study of intelligence: with the concept of mental age, an idea that presupposes that intelligence
does intelligence depend on context? The other theories of intelligence increases as one gets older. The average 10-year-old child has a mental age
discussed above essentially posit that intelligence is an ability, something or of 10. When this average child grows to age 12, she or he will seem more
some collection of things that one has or does not have. Sternberg, on the intelligent and will have a mental age of 12. By using this method, Binet
other hand, asserts that what is intelligent behavior depends on the context created a test that would identify children who lagged behind most of their
or situation in which it occurs. If intelligence does, indeed, depend on peers, who were in step with their peer group, and who were ahead of their
context, devising an intelligence test becomes a particularly difficult task. peers. Binet created a standardized test using the method described earlier in
The most common intelligence tests used (described in the next section) are this unit. He administered questions to a standardization sample and
based on the view of intelligence as ability based. constructed a test that would differentiate among children functioning at
different levels.
Lewis Terman, a Stanford professor, used this system to create the
measure we know as IQ and the test known as the Stanford-Binet IQ test.
IQ stands for intelligence quotient. A person’s IQ score on this test is
computed by dividing the person’s mental age by his or her chronological
age and multiplying by 100. Thus, the 10-year-old child described above
has an IQ of 100 because 10/10 × 100 = 100. A child who has a mental age
of 15 at age 10 would have an IQ of 150, 15/10 × 100 = 150. A commonly
asked question about this system is how it deals with adults. While talking
about a mental age of 8 or 11 or 17 makes sense, what does having a mental
age of 25 or 33 or 58 mean? To address this problem, Terman assigned all
adults an arbitrary age of 20.
David Wechsler used a different way to measure intelligence. Although it
does not involve finding a quotient, it is still known as an IQ test. Three
different Wechsler tests actually exist. The Wechsler Adult Intelligence
Scale (WAIS) is used in testing adults, the Wechsler Intelligence Scale for Figure 10.1 The normal distribution on an IQ test
Children (WISC) is given to children between the ages of 6 and 16, and the
Wechsler Preschool and Primary Scale of Intelligence (WPPSI) can be For instance, approximately 68 percent of scores fall within 1 standard
administered to children as young as 4. The Wechsler tests yield IQ scores deviation of the mean, approximately 95 percent fall within 2 standard
based on what is known as deviation IQ. The tests are standardized so that deviations of the mean, and nearly 99 percent of scores fall within 3
the mean is 100, the standard deviation is 15, and the scores form a normal standard deviations of the mean. People’s scores are determined by how
distribution. Remember that in a normal distribution, the percentages of many standard deviations they are away from the mean. Thus, Peter who
scores that fall under each part of the normal curve are predetermined (see scores at the 15.87th percentile falls at 1 standard deviation below the mean
Figure 10.1). and is assigned a score of 85, while Juanita who scores at the 97.72nd
percentile has scored 2 standard deviations above the mean and has scored
130. Of course, most people do not fall exactly 1 or 2 standard deviations
above the mean. However, using such an example would necessitate less
obvious mathematical calculations. For more information on the normal
curve, you might want to refer to Chapter 3, “Statistics.”
Whereas the Stanford-Binet IQ test utilizes a variety of different kinds of
questions to yield a single IQ score, the Wechsler tests result in scores on
several subscales as well as a total IQ score. For instance, the WAIS has 11
subscales. Six of them are combined to produce a verbal IQ score. Five are
used to indicate performance IQ. The kinds of questions used to measure
verbal IQ ask people to define words, solve mathematical word problems,
and explain ways in which different items are similar. The items on the
performance section involve tasks like duplicating a pattern with blocks,
correctly ordering pictures so they tell a story, and identifying missing intelligence? That heritability does not apply to an individual but, rather, to
elements in pictures. Differences between a person’s score on the verbal and a population is important to point out. Whatever the heritability ratio for
performance sections of this exam can be used to identify learning intelligence, it will not tell us how much of any person’s intelligence was
disabilities. determined by nature or by nurture.
Solving this controversy once and for all is essentially impossible because
Bias in Testing we cannot ethically set up the kind of controlled experiment necessary to
Much discussion has centered on whether widely used IQ tests and the SAT provide definitive answers to this question. However, many researchers
are biased against certain groups. Interestingly, researchers seem to agree have studied this issue, and some of their findings are presented below:
that although different races and genders may score differently on these Performance on intelligence tests has been increasing steadily
tests, the tests have the same predictive validity for all groups. In other
throughout the century, a finding known as the Flynn effect. Since the
words, SAT scores are equally good predictors of college grades for
different genders and for different racial groups. Thus, in a sense, the test is gene pool has remained relatively stable, this finding suggests that
clearly not biased. However, other researchers have argued that both the environmental factors such as nutrition, education, and, perhaps,
tests and the college grades are biased in a far more fundamental way. television and video games play a role in intelligence.
Advantages seem to accrue to the white, middle-class, and upper-class Monozygotic (identical) twins, who share 100 percent of their genetic
students. The experiences of other cultural groups seem to work to their material, score much more similarly on intelligence tests than do
detriment both on these tests and in college. Writers of the test may assume dizygotic (fraternal) twins, who have, on average, only 50 percent of
a level of vocabulary and a range of experiences that members of these their genes in common. Nonetheless, some researchers have suggested
groups have not been exposed to. To the extent that the tests are supposed to that monozygotic twins tend to be treated more similarly than dizygotic
identify academic potential, they may then be both flawed and biased. twins, thus confounding the effects of nature with those of nurture.
Research on identical twins separated at birth has found strong
Nature vs. Nurture: Intelligence correlations in intelligence scores. However, researchers advocating
One of the most difficult and controversial issues in psychology involves more of an environmental influence point out that usually the twins are
sorting out the relative effects of nature and nurture. Keep in mind that placed into similar environments, again making it difficult to sift out the
nature refers to the influence of genetics, while nurture stresses the relative effects of nature and nurture. For instance, if each of the twins
importance of the environment and learning. One of the more hotly is placed into a white, middle-class, suburban home, concluding that all
contested aspects of the nature-nurture debate is intelligence. Human their similarities are genetically based does not make sense.
intelligence is clearly affected by both nature and nurture. Research Psychologists agree that racial differences in IQ scores are explained by
suggests that both genetic and environmental factors play a role in molding
differences in environment.
intelligence.
An important term that researchers use in discussing the effects of nature Participation in government programs such as Head Start, meant to
and nurture is heritability. Heritability is a measure of how much of a redress some of the disadvantages faced by impoverished groups, has
trait’s variation is explained by genetic factors. Heritability can range from been shown to correlate with higher scores on intelligence tests.
0 to 1, where 0 indicates that the environment is totally responsible for However, opponents of such programs assert that these gains are
differences in the trait and 1 means that all the variation in the trait can be limited and of short duration. Advocates of such interventions respond
accounted for genetically. Thus, the question can be asked: How heritable is that expecting the gains to outlast the programs is unreasonable.
After putting the issue of cause aside, when comparing groups of people
on any characteristic, keep in mind that differences within groups generally
dwarf differences between groups. In other words, within any one group
there will be more diversity than between any two groups. Practically
speaking, if we find that boys perform better on a certain test than girls do,
more of a difference will exist between the highest-scoring boy and the
lowest-scoring boy than between the average-scoring boy and the average-
scoring girl. Furthermore, knowing that boys generally outperform girls on
this test tells us nothing about the performance of any particular girl
compared with the performance of any particular boy. Therefore, we need to
be careful about how we use information about differences between groups.
Essentially, we should not use it. We should ignore it and evaluate each
person, regardless of group membership, as an individual.
A CAUTIONARY NOTE
It is often said that we live in a testing society. We like to be able to
measure things and assign them a number. Therefore, keeping in mind
the limitations and extraordinary labeling power of these instruments is
particularly important. As we have discussed, the definition of
intelligence (and many other concepts) remains hotly debated and many
factors affect people’s performances on tests. Thus, we need to take care
not to ascribe too great a meaning to a test score. Many schools that used
to measure all their students’ IQs periodically have abandoned that
practice. Schools that used to base admission to programs for exceptional
children solely on these tests now frequently gather information in other
ways as well. When IQ tests are given, the results remain confidential so
as not to create expectations about how people ought to perform.
Although well-designed tests can be extremely useful, we must recognize
their limitations.
Unit 2 Multiple-Choice Questions
1. If you had sight in only one eye, which of the following depth cues