UNIT 4 – PERCEIVING OBJECTS AND
RECOGNIZING PATTERNS
PERCEPTION
INTRODUCTION
Perception is the set of processes by which we recognize, organize, and make sense of the sensations
we receive from environmental stimuli. James Gibson (1966, 1979) provided a useful framework for
studying perception. He introduced the concepts of distal (external) object, informational medium,
proximal stimulation, and perceptual object.
Distal Objects: These are things that exist in the environment and emit signals that are available to be
perceived.
Informational Medium: The means by which data is conveyed from the object to the perceiver: light
rays, sound waves, chemical particles, heat, etc.
Proximal Stimulation: Refers to activation that occurs at the specific location where the information is
received: the eardrums, the retina, the skin, etc.
Perceptual Object: An internal conception, based on the proximal stimulation, of the distal object
PROCESS OF PERCEPTION
Each such object is a distal stimulus. For a living organism to process information about these stimuli,
it must first receive the information through one or more sensory systems—in this example, the visual
system. The reception of information and its registration by a sense organ make up the proximal
stimulus. In our earlier example, light waves reflect from the trees and cars to your eyes, in particular
to a surface at the back of each eye known as the retina. There, an image of the trees and cars, called
the retinal image, is formed. This image is two-dimensional, and its size depends on your distance
from the window and the objects beyond (the closer you are, the larger the image). In addition, the
image is upside down and is reversed with respect to left and right. The meaningful interpretation of
the proximal stimulus is the percept— your interpretation that the stimuli are trees, cars, people, and
so forth. From the upside-down, backward, two-dimensional image, you quickly (almost
instantaneously) “see” a set of objects you recognize. You also “recognize” that, say, the giant oak tree
is closer to you than are the lilac shrubs, which appear to recede in depth away from you. This
information is not part of the proximal stimulus. Somehow, you must interpret the proximal stimulus
to know it. Although researchers studying perception disagree about much, they do agree that percepts
are not the same things as proximal stimuli. Consider a simple demonstration of size constancy.
Extend your arm away from your body, and look at the back of your hand. Now, keeping the back of
your hand facing you, slowly bring it toward you a few inches, then away from you. Does your hand
seem to be changing size as it moves? Probably not, although the size of the hand in the retinal image
is most certainly changing. The point here is that perception involves something different from the
formation of retinal images. Related to perception is a process called pattern recognition. This is the
recognition of a particular object, event, and so on, as belonging to a class of objects, events, and so
on. Your recognition of the object you are looking at as belonging to the class of things called
“shrubs” is an instance of pattern recognition. Because the formation of most percepts involves some
classification and recognition, most, if not all, instances of perception involve pattern recognition.
The precondition for vision is the existence of light. Light is electromagnetic radiation that can be
described in terms of wavelength. Humans can perceive only a small range of the wavelengths that
exist; the visible wavelengths are from 380 to 750 nanometers. Vision begins when light passes
through the cornea, is a clear dome that protects the eye. The light then passes through the pupil, the
opening in the center of the iris. It continues through the crystalline lens and the vitreous humor.
Eventually, the light focuses on the retina where electromagnetic light energy is transduced, i.e.,
converted into neural electrochemical impulses. Vision is most acute in the fovea, which is a small,
thin region of the retina. When looking straight at an object, the eyes rotate so that the image falls
directly onto the fovea.
The retina consists of three layers of neuronal tissue, namely, ganglion cells, interneuron cells, and
photoreceptors.
Ganglion cells: Retinal ganglion cells (RGC) bear the sole responsibility of propagating visual stimuli
to the brain. Their axons, which make up the optic nerve, project from the retina to the brain through
the lamina cribrosa.
Interneuron cells: Three types of interneuron cells are present on the retina. Amacrine cells and
horizontal cells the middle layer of cells. Bipolar cells make dual connections forward and outward to
the ganglion cells, as well as backward and inward to the third layer of retinal cells.
Photoreceptors: They convert light energy into electrochemical energy that is transmitted by neurons
to the brain. There are two kinds of photoreceptors—rods and cones. Each eye contains roughly 120
million rods and 8 million cones. Rods and cones differ not only in shape but also in their
compositions, locations, and responses to light. Within the rods and cones are photopigments,
chemical substances that react to light and transform physical electromagnetic energy into an
electrochemical neural impulse that can be understood by the brain. The rods are long and thin
photoreceptors. They are more highly concentrated in the periphery of the retina than in the foveal
region. The rods are responsible for night vision and are sensitive to light and dark stimuli. The cones
are short and thick photoreceptors and allow for the perception of color. They are more highly
concentrated in the foveal region than in the periphery of the retina.
The axons of the ganglion cells in the eye collectively form the optic nerve for that eye. The optic
nerves of the two eyes join at the base of the brain to form the optic chiasma. At this point, the
ganglion cells from the inward, or nasal, part of the retina cross through the optic chiasma and extend
to the opposite hemisphere of the brain. The ganglion cells from the temporal area of the retina go to
the hemisphere on the same side of the body.
The lens of each eye naturally inverts the image of the world as it projects the image onto the retina.
In this way, the message sent to the brain is upside-down and backward. After being routed via the
optic chiasma, about 90% of the ganglion cells then go to the lateral geniculate nucleus of the
thalamus. From the thalamus, neurons carry information to the primary visual cortex (V1 or striate
cortex) in the occipital lobe of the brain. The visual cortex contains several processing areas. Each
area handles different kinds of visual information relating to intensity and quality, including color,
location, depth, pattern, and form.
There are two visual pathways in the brain. Work on visual perception has identified separate neural
pathways in the cerebral cortex for processing different aspects of the same stimuli. Perception
deficits like ataxia and agnosia also point toward the existence of different pathways. The information
from the primary visual cortex in the occipital lobe is forwarded through two fasciculi (fiber bundles):
One ascends toward the parietal lobe (along the dorsal pathway), and one descends to the temporal
lobe (along the ventral pathway). The dorsal pathway is also called the where pathway and is
responsible for processing location and motion information; the ventral pathway is called the what
pathway because it is mainly responsible for processing the color, shape, and identity of visual
stimuli.
GESTALT PRINCIPLES OF PERCEPTION
INTRODUCTION
The Gestalt approach to form perception that was developed in Germany in the early 20th century is
useful particularly for understanding how we perceive groups of objects or even parts of objects to
form integral wholes. It was founded by Kurt Koffka (1886–1941), Wolfgang Köhler (1887–1968),
and Max Wertheimer (1880–1943) and was based on the notion that the whole differs from the sum of
its individual parts.
The overarching law is the law of Prägnanz. Any given visual array is perceived in a way that most
simply organizes the different elements into a stable and coherent form. Gestalt principles include
figure-ground perception, proximity, similarity, continuity, closure, and symmetry. Each of these
principles supports the overarching law of Prägnanz.
GESTALT PRINCIPLES
TOP - DOWN THEORIES OF PERCEPTION
INTRODUCTION
Top-down processing is defined as the development of pattern recognition through the use of
contextual information. As an impact of this approach, an individual's perception begins with the most
general and gradually moves towards the most specific. The perceptions of the individuals are largely
influenced by the expectations that they have and their prior knowledge. Gregory suggests that using
prior knowledge and experience of a stimulus helps an individual to draw inferences. Richard Gregory
estimated that about 90% of the information is lost between the time it takes to go from the eye to the
brain, which is why the brain must guess what the individual sees based on previous experience. In
other words, individuals construct their perception of reality, and these perceptions are hypotheses or
propositions based on prior experiences and prior knowledge.
TOP - DOWN APPROACH THEORIES
Constructive perception theory: the perceiver builds (constructs) a cognitive understanding
(perception) of a stimulus. The concepts of the perceiver and his or her cognitive processes influence
what he or she sees. The perceiver uses sensory information as the foundation for the structure but
also uses other sources of information to build the perception. This viewpoint also is known as
intelligent perception because it states that higher-order thinking plays an important role in
perception. It also emphasizes the role of learning in perception. Some investigators have pointed out
that not only does the world affect our perception but also the world we experience is actually formed
by our perception. In other words, perception is reciprocal with the world we experience. Perception
both affects and is affected by the world as we experience it. According to this theory, perception
comprises not merely a low-level set of cognitive processes, but actually a quite sophisticated set of
processes that interact with and are guided by human intelligence. Phenomena such as perceptual
constancies demonstrate the constructive perception assumption.
According to constructivists, we usually make the correct attributions regarding our visual sensations.
The reason is that we perform unconscious inference, the process by which we unconsciously
assimilate information from a number of sources to create a perception. In other words, using more
than one source of information, we make judgments that we are not even aware of making. Successful
constructive perception requires intelligence and thought in combining sensory information with
knowledge gained from previous experience.
One reason for favoring the constructive approach is that bottom-up (data driven) theories of
perception do not fully explain context effects. Context effects are the influences of the surrounding
environment on perception. Objects that were appropriate to the established context are recognized
more rapidly than were objects that are inappropriate to the established context. The strength of the
context also plays a role in object recognition. Perhaps even more striking is a context effect known as
the configural-superiority effect, by which objects presented in certain configurations are easier to
recognize than the objects presented in isolation, even if the objects in the configurations are more
complex than those in isolation.
In this view, intelligence and perceptual processes interact in the formation of our beliefs about what it
is that we are encountering in our everyday contacts with the world at large. An extreme top-down
position would drastically underestimate the importance of sensory data. If it were correct, we would
be susceptible to gross inaccuracies of perception. We frequently would form hypotheses and
expectancies that inadequately evaluated the sensory data available.
BOTTOM - UP THEORIES OF PERCEPTION
INTRODUCTION
In the bottom-up processing approach, perception starts at the sensory input, the stimulus. Thus,
perception can be described as data-driven. Bottom-up processing is the process of "sensing," in
which our sensory receptors take in sensory data from the outside environment. Our brains select,
arrange and interpret these sensations through perception. In the bottom-up processing approach, the
perception begins at the sensory input, the stimulus. Therefore, while describing perception, the word
data-driven can be used. The bottom-up approach was developed by psychologist E.J. Gibson, who
set a solid foundation for the perception of human beings. According to Gibson, sensation is
perception, and there is no need for extra interpretation, as there is enough information in our
environment to directly make sense of the world. The bottom-up approach deconstructs perception
rather than looking at it comprehensively, involving how sensory information, visual processes, and
expectations contribute to how the world is perceived.
BOTTOM - UP APPROACH THEORIES
The four main bottom-up theories of form and pattern perception are direct perception, template
theories, feature theories, and recognition-by-components theory.
Direct perception: According to Gibson’s theory of direct perception, the information in our sensory
receptors, including the sensory context, is all we need to perceive anything. As the environment
supplies us with all the information we need for perception, this view is sometimes also called
ecological perception. In other words, we do not need higher cognitive processes or anything else to
mediate between our sensory experiences and our perceptions. Existing beliefs or higher-level
inferential thought processes are not necessary for perception. Gibson believed that, in the real world,
sufficient contextual information usually exists to make perceptual judgments. He claimed that we
need not appeal to higher level intelligent processes to explain perception. Gibson (1979) believed
that we use this contextual information directly. In essence, we are biologically tuned to respond to it.
According to Gibson, we use texture gradients as cues for depth and distance. Those cues aid us to
perceive directly the relative proximity or distance of objects and of parts of objects. Gibson’s model
sometimes is referred to as an ecological model.
Template theories: These theories suggest that we have stored myriad sets of templates. Templates are
highly detailed models for patterns we potentially might recognize. We recognize a pattern by
comparing it with our set of templates. We then choose the exact template that perfectly matches what
we observe. Template matching theories belong to the group of chunk-based theories that suggest that
expertise is attained by acquiring chunks of knowledge in long-term memory that can later be
accessed for fast recognition. Studies with chess players have shown that the temporal lobe is indeed
activated when the players access the stored chunks in their long-term memory. Template-matching
theories fail to explain some aspects of the perception of letters. We can recognize an A as an A
despite variations in the size, orientation, and form in which the letter is written.
Feature matching theories: According to these theories, we attempt to match features of a pattern to
features stored in memory, rather than to match a whole pattern to a template or a prototype. One such
feature-matching model has been called Pandemonium model. In it, metaphorical “demons” with
specific duties receive and analyze the features of a stimulus. In Oliver Selfridge’s Pandemonium
Model, there are four kinds of demons: image demons, feature demons, cognitive demons, and
decision demons. The “image demons” receive a retinal image and pass it on to “feature demons.”
Each feature demon calls out when there are matches between the stimulus and the given feature.
These matches are yelled out at demons at the next level of the hierarchy, the “cognitive (thinking)
demons.” The cognitive demons in turn shout out possible patterns stored in memory that conform to
one or more of the features noticed by the feature demons. A “decision demon” listens to the
pandemonium of the cognitive demons. It decides on what has been seen, based on which cognitive
demon is shouting the most frequently (i.e., which has the most matching features).
Recognition by components theory: According to this theory, our ability to perceive 3-D objects is
due to the help of simple geometric shapes. Irving Biederman (1987) suggested that we achieve this
by manipulating a number of simple 3-D geometric shapes called geons (for geometrical ions). They
include objects such as bricks, cylinders, wedges, cones, and Biederman’s recognition-by-components
(RBC) theory, we quickly recognize objects by observing the edges of them and then decomposing
the objects into geons. The geons also can be recomposed into alternative arrangements. You know
that a small set of letters can be manipulated to compose countless words and sentences. Similarly, a
small number of geons can be used to build up many basic shapes and then myriad basic objects.