1.
Introduction: How We Perceive Depth
We live in a three-dimensional world, yet each retina receives a flat, two-
dimensional image. The brain reconstructs depth and distance from this
limited information by interpreting depth cues—features in the image or
within our eyes that signal how far things are.
Depth perception also supports size and shape constancy, allowing us to
perceive objects as stable even when their retinal image changes with
distance or angle.
Because a single 2-D projection can correspond to many possible 3-D
scenes, the visual system combines multiple, partly redundant cues. These
fall into three broad categories: oculomotor, monocular (static and
dynamic), and binocular.
2. Case Study — “Stereo Sue”
Neuroscientist Sue Barry was born with strabismus (misaligned eyes). To prevent
double vision, her brain suppressed one eye’s input, leaving her stereoblind—she
could judge depth only from monocular and motion cues.
After undergoing visual therapy with prism glasses, she developed stereopsis,
perceiving the world in vivid depth for the first time. She described the steering
wheel as “popping out” and snowflakes as floating in three-dimensional space. Her
case demonstrates that stereoscopic depth arises from coordinated binocular
processing, and that even adults can sometimes retrain the brain to achieve it.
3. Oculomotor Depth Cues (Short-Range, ≤ 2 m)
Oculomotor cues come from the feedback of muscles that control eye position and
lens curvature. They operate only at near distances but provide strong signals of
proximity.
Accommodation: The ciliary muscles adjust the lens shape to focus an
image on the retina; the effort required signals distance.
Convergence: The eyes rotate inward for near targets. Greater inward
rotation indicates a closer object; beyond about two meters, the convergence
angle changes too little to be a significant cue.
4. Monocular Depth Cues
Monocular cues use information from a single retinal image and operate over long
distances. They include static (pictorial) cues—visible even in still pictures—and
dynamic (motion-based) cues that appear when observer or scene moves.
A. Static / Pictorial Cues
Position-Based:
Partial Occlusion (Interposition): When one object overlaps another, the
occluder appears closer. Occlusion boundaries form T-junctions where edges
meet.
Relative Height: Below the horizon, lower = closer; above the horizon,
higher = closer.
Size-Based:
Visual Angle: Farther objects project smaller retinal images.
Familiar Size: Knowing an object’s real size lets us estimate distance.
Relative Size: Among similarly sized objects, smaller images mean greater
distance.
Texture Gradient: Repetitive patterns shrink and crowd together with
distance.
Linear Perspective: Parallel lines converge with depth, as seen in railroad
tracks or hallways.
Lighting-Based:
Atmospheric Perspective: Distant objects appear hazy or bluish due to light
scattering.
Shading: Light and shadow define curvature, interpreted under the
assumption that light comes from above.
Cast Shadows: Shadow placement indicates relative height and distance
between object and surface.
B. Dynamic (Motion-Based) Cues
Movement provides powerful information about spatial layout.
Motion Parallax: As the observer moves, nearby objects sweep faster
across the retina than far ones. Fixating one point makes near objects seem
to move opposite your motion and far ones move with it.
Optic Flow: During forward motion, the visual field radiates outward from
a central focus of expansion (FOE) that marks the travel direction.
Deletion and Accretion: As you or objects move, nearer surfaces gradually
hide (delete) or reveal (accrete) more distant ones, signaling relative depth.
5. Binocular Cues — Stereopsis
Stereopsis is the vivid three-dimensional perception that results from combining
slightly different views from each eye.
Binocular Disparity: The positional difference between left- and right-eye
images.
Corresponding Points: Retinal locations that would coincide if the retinas
were overlaid.
Horopter: Imaginary surface where images fall on corresponding points
(zero disparity).
Crossed Disparity: For objects nearer than the fixation point.
Uncrossed Disparity: For objects farther away.
Magnitude of Disparity: Increases with an object’s distance from the
horopter.
Correspondence Problem: The brain must decide which features in each
eye match; random-dot stereograms show this matching precedes object
recognition.
Neural Basis: Binocular neurons in V1 respond when both eyes are
stimulated together and are tuned to specific disparity ranges.
6. Integrating Depth Cues
The visual system integrates multiple cues through unconscious inference—often
modeled by Bayesian cue combination, weighting each cue by its reliability and
prior knowledge. No single cue dominates, but partial occlusion tends to be most
universal. The more cues available, the more precise the perception of depth.
7. Perceptual Constancy
A. Size Constancy
We perceive object size as constant despite retinal changes. This follows the Size–
Distance Invariance Relation:
[
S=R×D
]
where S = perceived size, R = retinal size, D = perceived distance.
Emmert’s Law extends this principle to afterimages—an afterimage appears
larger on a distant wall because its retinal size stays constant while perceived
distance increases.
B. Shape Constancy
Perceived shape remains stable even when viewing angle changes. The Shape–
Slant Invariance principle links perceived shape to perceived surface orientation.
8. Illusions of Depth, Size, and Shape
Our brain’s reliance on learned depth rules can create striking misperceptions.
Forced Perspective: Misaligned near and far elements seem equally distant,
distorting size.
Ponzo Illusion: Converging lines imply depth; identical objects placed
“farther” up the perspective appear larger. fMRI shows correlated activation
in V1.
Ames Room: A trapezoidal room viewed through a peephole appears
rectangular, making two identical-sized people look very different—an error
in shape constancy.
Moon Illusion: The horizon moon looks larger than the zenith moon
because contextual depth cues make it seem farther away.
Tabletop Illusion: Two identical trapezoid tabletops appear different when
slant cues suggest differing orientations.
9. Applications of Depth Perception
3-D Movies: Add binocular disparity to monocular cues; viewed through
polarized or liquid-crystal shutter glasses.
Autostereoscopic Displays: Use parallax barriers or lenticular lenses for
glasses-free 3-D (limited viewpoints).
Holography: Reconstructs full light fields so the image changes naturally
with observer movement, introducing dynamic cues like motion parallax and
deletion/accretion.
10. Motion Perception (Chapter 7)
A. Overview and Case Study
Motion is central to survival and object recognition. After a stroke damaged her
area MT, patient L.M. lost motion vision (akinetopsia). Liquids appeared frozen
mid-pour, and cars seemed to jump between positions. Her case reveals how
crucial dedicated motion areas are for everyday life.
B. Perceptual Organization from Motion
Common Motion: Elements moving together are grouped as one object.
Apparent Motion: Alternating flashes at separate locations create the
illusion of continuous movement (the basis of film).
Figure–Ground Segregation: Motion discontinuities alone can distinguish
moving figures from static backgrounds, demonstrated in random-dot
kinematograms.
Biological Motion: Point-light walker displays instantly reveal human form
and action. The posterior superior temporal sulcus (STS p) specializes in
such animate motion, which captures attention more than inanimate motion.
C. Eye Movements and Stability
Because the fovea handles high detail, the eyes constantly move.
Saccades: Rapid gaze shifts; saccadic suppression blocks visual input to
prevent blur.
Smooth Pursuit: Keeps a moving target stable on the retina.
Real-Motion Cells: Respond to actual environmental motion, not self-
generated retinal motion.
Corollary Discharge Signal (CDS): A copy of the eye-movement
command from the superior colliculus predicts the sensory consequences of
movement, allowing visual stability.
D. Neural Mechanisms of Motion
Reichardt Detector: Simplest model—two inputs, one delay, one
comparator—to encode speed and direction.
Direction Tuning: Motion-sensitive neurons respond maximally to a
preferred direction.
Motion Aftereffect (MAE): Prolonged motion stimulation fatigues one
directional pool, producing an opposite-motion illusion when viewing a
stationary scene.
Area V1: Detects small-scale local motion.
Area MT (V5): Integrates local signals into global motion, tuned for both
direction and speed; stimulation can bias perceived motion, and lesions
cause akinetopsia.
Aperture Problem: Individual V1 neurons see motion through narrow
“windows.” MT integrates across them to infer true object trajectory.
CHAPTER 6: PERCEIVING DEPTH
Born with strabismus (misaligned eyes), leading to automatic suppression of one
eye’s signal and resulting in her being stereoblind. Judged depth only using
monocular and motion cues. After corrective therapy (prism glasses), she gained
stereopsis, perceiving depth vividly ("steering wheel had ‘popped out’,"
snowflakes perceived in 3-D space)
Introduction (Overview of Depth Perception)
The fundamental goal of depth perception is the accurate perception of a 3-D
world based on two 2-D retinal images.
The vertical and horizontal dimensions of 3-D space are explicitly
represented in the retinal image.
The representation of 3-D space in the 2-D retinal image is inherently many-
to-one, meaning an infinite variety of 3-D scenes can produce one and the
same retinal image. For instance, if an object's size is changed, moving it
correspondingly closer or farther away will result in an unchanged retinal
image.
The ambiguity of the retinal image means depth perception is closely linked
to perceptions of size and shape.
Depth perception is achieved by integrating information from various
reliable (but fallible) cues. The cues fall into three categories: oculomotor,
monocular (static and dynamic), and binocular [1054f].
Oculomotor Depth Cues
Oculomotor depth cues arise from the feedback from the muscles that control the
shape of the lens and the position of the eyes. They provide depth information only
for objects up to about 2 meters away.
Accommodation: The process where the lens adjusts its shape to focus a
sharp image on the retina. Sensation of this adjustment acts as a cue to the
object’s distance.
Convergence: The eyes turn inward toward each other to focus on close
objects [196f, 1063]. Tension felt in the eye muscles cues the decreasing
distance. For objects farther than about 2 meters, the convergence angle
changes too little to be a significant cue.
Monocular Depth Cues
Monocular depth cues are based solely on information within the retinal image and
operate across a much greater range of distances than oculomotor cues.
Static Cues: Position, Size, and Lighting in the Retinal Image
Static cues, also called pictorial cues, are present in motionless 2-D depictions of
3-D scenes.
Position in the Retinal Image
o Partial Occlusion (Interposition): If one object partially hides
another, the occluding object is perceived as being closer [197, 1068,
1069f]. This works because the visual system unconsciously assumes
that the occluded shape is a simple continuation of the visible part,
rejecting the highly unlikely possibility of accidental alignment.
Intersections between edges that signal occlusion are called T-
junctions.
o Relative Height: Objects below the horizon/eye level that are lower
in the retinal image are perceived as closer [199, 1078, 1079f]. Above
the horizon/eye level, higher objects are closer.
Size in the Retinal Image
o The size—distance relation states that a farther object produces a
smaller retinal image. Retinal image size is measured by visual angle
[199, 200f].
o Familiar Size: Knowing the retinal image size of a familiar object at
a known distance allows the observer to gauge its actual distance
[200, 200f, 1086].
o Relative Size: Assuming objects are of approximately equal size, the
relative size of their retinal images dictates their relative distances
(e.g., a person with half the retinal image size is twice as far away)
[200, 201f, 1088, 1089].
o Texture Gradients: On surfaces with regular elements, the retinal
image size of these elements decreases as distance increases,
functioning as a depth cue [201, 201f].
o Linear Perspective: Parallel lines appear to converge as they recede
in depth [201, 202f].
Lighting in the Retinal Image
o Atmospheric Perspective: Distant objects appear less distinct/more
bluish due to light scattering by air particles [202, 202f].
o Shading: Patterns of illumination on curved surfaces cue depth and
surface orientation [202, 203f]. Interpretation relies on the
unconscious assumption that the light source is typically from above
[202, 203f].
o Cast Shadows: Shadows cast by objects on a surface serve as cues for
depth [203, 203f].
Dynamic Cues: Movement in the Retinal Image
Dynamic cues provide depth information based on movement and a constantly
changing viewpoint.
Motion Parallax: As an observer moves across a scene, nearby objects
move more quickly across the retina than distant objects. If the observer
fixates an object, nearer objects appear to move opposite the direction of
motion, while farther objects move in the same direction.
Optic Flow: The relative motions of objects and surfaces as the observer
moves forward/backward. Objects flow outward from the focus of
expansion (the navigation goal), with nearer objects flowing more rapidly.
Deletion and Accretion: The gradual hiding (deletion) and uncovering
(accretion) of an object as it moves behind another acts as a dynamic depth
cue.
Binocular Depth Cue: Disparity in the Retinal Images
Stereopsis (Stereoscopic Depth Perception): The vivid sense of depth
resulting from combining the slightly different retinal images from the two
eyes.
Binocular Disparity: The difference in the relative positions of the retinal
images of objects in the two eyes [208, 208f].
Corresponding and Noncorresponding Points, and the Horopter:
o Corresponding points on the retinas coincide if the retinas are
superimposed (e.g., foveas) [208, 209f, 1128, 1129].
o Noncorresponding points do not coincide [208, 209f, 1129].
o The horopter is the imaginary surface where objects project images
onto corresponding points [208, 209f, 1131].
Crossed Disparity, Uncrossed Disparity, and Zero Disparity:
o Zero disparity: Occurs when an object falls on corresponding points
(e.g., objects on the horopter or the fixated object) [211, 210f, 1141].
o Crossed disparity: Occurs for objects closer than the horopter [209,
210f].
o Uncrossed disparity: Occurs for objects farther away than the
horopter [211, 210f].
o The magnitude of disparity increases as the object's distance from
the horopter increases [211, 210f, 1141, 1142].
Correspondence Problem: The problem of determining which feature in
one eye's image matches which feature in the other eye's image.
o Depth perception in Random Dot Stereograms (RDS) proves that
matching (correspondence) precedes object recognition.
o The brain solves this by assuming that each feature matches one and
only one feature, and that scenes consist of smooth, continuous
surfaces.
Neural Basis of Stereopsis: Binocular cells respond best when receptive
fields in both eyes are stimulated simultaneously [216, 217f]. These cells are
tuned to specific types and magnitudes of disparity.
Integrating Depth Cues
The visual system uses a wide variety of cues because redundancy ensures
accuracy and different cues are useful under different conditions.
No single cue dominates or is necessary in all situations (though partial
occlusion may be the most dominant).
The more cues present, the greater the accuracy and consistency of depth
perception.
Cues are integrated via unconscious inference, often modeled using the
Bayesian approach, where cues are combined into a weighted average
based on their reliability and prior knowledge.
Depth and Perceptual Constancy
Size Constancy and Size—Distance Invariance: Size constancy is
perceiving an object's size as constant despite changing retinal image size
due to distance. This is governed by the principle of size—distance
invariance (perceived size depends on perceived distance, and vice versa).
Emmert’s law describes this principle applied to retinal afterimages [220,
220f, 1195].
Shape Constancy and Shape-Slant Invariance: Shape constancy is
perceiving an object's shape as constant despite changes in retinal image
shape due to its slant (orientation). This is governed by the principle of
shape—slant invariance (perceived shape depends on perceived slant, and
vice versa) [221, 221f, 1197].
Illusions of Depth, Size, and Shape
Forced Perspective: Misalignment of near and far objects makes them seem
equally distant, leading to size misperception [222, 222f, 1204].
Ponzo Illusion: Linear perspective cues lead to the misperception that two
identical objects are at different distances, making the apparently farther
object seem larger (size—distance invariance) [223, 223f, 1206]. The
illusion affects activity in V1 [224, 224f, 1210].
Ames Room: A trapezoidal room designed to look rectangular from one
viewpoint (a peephole); causes perceived size distortion of people standing
inside. It is a failure of shape constancy [224, 224f, 1212, 1215].
Moon Illusion: The horizon moon appears larger than the zenith moon,
likely because depth cues make it seem farther away (size—distance
invariance) [225, 226f, 1218].
Tabletop Illusion: Identical retinal shapes of trapezoids are perceived as
different due to perceived slant, demonstrating the automatic correction for
slant (shape constancy) [226, 226f, 1222, 1223].
APPLICATIONS: 3-D Motion Pictures and Television
3-D motion pictures add binocular disparity cues to the static and dynamic
monocular cues found in conventional movies.
Viewing methods include polarized lenses and liquid-crystal shutter
glasses [228, 228f].
Autostereoscopy (using parallax barriers or lenticular lenses) enables 3-D
viewing without glasses but is limited to specific viewpoints [228, 229f,
1229].
Holography aims to provide images that change view as the observer
moves, allowing for dynamic depth cues (motion parallax,
deletion/accretion) resulting from the observer’s own motion.
CHAPTER 7: PERCEIVING MOTION
Suffered a stroke that damaged regions including area MT. Her most serious
complaint was the loss of her ability to see motion (akinetopsia). Experienced
difficulties such as pouring tea (level appeared "frozen") and crossing streets (cars
suddenly appeared close).
Introduction (Overview of Motion Perception)
The world is characterized by moving objects and continuously moving observers,
which makes motion perception vital for survival and interaction. Impairment to
motion vision (as seen in L.M.) can severely disrupt daily life.
Perceptual Organization from Motion
Motion processing contributes significantly to how the visual system organizes
sensory input.
Perceptual Grouping Based on Real and Apparent Motion:
o Motion helps group successive views of an object into a stable, single
perception.
o Apparent motion is the illusion of motion created by rapidly
alternating presentation of stimuli separated in time and location [234,
234f, 1259, 1261].
o The perceived direction in an apparent motion quartet can
spontaneously flip, and perceived direction is strongly biased by the
relative proximity of the dots [234f, 235, 1263, 1265].
Figure—Ground Organization: Discontinuity in movement alone can
separate a moving object (figure) from a stationary background (ground)
[235, 236f].
o Random Dot Kinematograms (RDKs) use this principle to make the
shape of a rigidly moving region of dots visible against a randomly
moving background [235, 236f, 1266, 1270].
Sensitivity to Biological Motion:
o Point-light walker displays show the motion of people/animals using
lights attached to joints [236, 237f, 1271]. Perception of motion,
identity, and gender is almost immediate, based on the coordinated
timing of the points.
o Perception of biological motion is associated with the posterior
superior temporal sulcus (STSp).
o The visual system exhibits preferential attention capture toward
animate motion (e.g., changes in direction without external cause)
compared to inanimate motion [237, 238f, 1278].
Eye Movements and the Perception of Motion and Stability
Eye movements are frequent because high visual acuity is limited to the fovea.
Saccadic eye movements (saccades) are brief, rapid movements that shift
the gaze. Smooth pursuit eye movements track moving objects.
Saccadic suppression temporarily shuts down retinal input during saccades
to prevent motion blur.
The stability of the visual world, despite retinal motion, requires the visual
system to use extraretinal information about eye position/movement.
Real-motion cells respond only to actual object movement in the
environment, not to retinal motion caused by eye movements [240, 241f,
1297].
The extraretinal information is conveyed by a corollary discharge signal
(CDS), a copy of the eye-movement commands (from the superior
colliculus) sent to the brain. Evidence shows that this signal provides
advance warning of intended movements, enabling anticipatory adjustments
for stability [241, 242f, 243f, 1302, 1303, 1308].
Neural Basis of Motion Perception in Area V1 and Area MT
The visual system represents motion direction and speed.
A Simple Neural Circuit That Responds to Motion:
o To encode motion direction and speed, a neural circuit must monitor
two retinal locations and the timing of their stimulation.
o A critical element is a delay built into the signal transmission from
one of the receptive fields so that the signals arrive simultaneously at
the motion-selective neuron only if the stimulus is moving at the
neuron's specific preferred speed/direction [244, 245f, 1319].
o This simple circuit accounts for apparent motion [246, 246f].
o Neurons exhibit direction tuning, meaning response strength falls off
gradually from the preferred direction [246, 247f].
The Motion Aftereffect (MAE):
o The MAE is explained by an opponent motion circuit [246, 247f].
This circuit compares the activity of two subunits tuned to opposite
directions. Prolonged viewing fatigues the responding subunit,
causing the opponent subunit's resting activity to dominate, signaling
motion in the opposite direction. This process supports the perception
of motion contrast.
Area MT:
o MT is highly specialized for motion, generating a significant burst of
activity in response only to moving stimuli (unlike V1) [249, 249f,
1345].
o MT neurons are tuned to both direction and speed [105, 105f].
o MT activity correlates strongly with perceptual judgments of motion
direction, even at low motion coherence levels [250-252, 251f, 252f,
1357]. Damage to MT severely impairs motion perception.
The Aperture Problem: Perceiving the Motion of Objects:
o The aperture problem arises because V1 neurons have small
receptive fields ("apertures") and can only detect the component of
motion perpendicular to an object's edge, not its true direction [253,
253f, 1363].
o This is solved by MT neurons, which have larger receptive fields and
integrate information from multiple V1 neurons to perceive the
object's actual direction.