unit 3
unit 3
Milestones in Acquisition - Theoretical Perspectives- Semantics and Cognitive Science - Meaning and
Entailment -Reference Sense Cognitive and Computational Models of Semantic Processing - Information
Processing Models of the Mind- Physical symbol systems and language of thought- Applying the
Symbolic Paradigm- Neural networks and distributed information processing- Neural network models of
Cognitive Processes.
Milestones in Acquisition
Milestones in acquisition usually refers to the key stages in how humans acquire knowledge, skills, and
language across development. Cognitive science maps these stages to understand how the mind grows
from infancy to adulthood.
First-language acquisition follows a natural, predictable progression. Children move from pre-
linguistic vocalizations to single words around their first birthday, and transition into multi-word
combinations and complex grammar by age five.
A breakdown of the essential stages in early language acquisition includes:
Pre-Linguistic Stage (0–6 months): Infants begin communicating by crying, cooing, and
making vowel-like sounds. They learn to recognize the voices of primary caregivers and
distinguish basic sounds in their native language.
Babbling Stage (6–12 months): Babies start producing repetitive consonant-vowel
combinations (e.g., "ma-ma" or "da-da") and begin to understand simple requests or gestures.
Holophrastic Stage (12–18 months): Also known as the one-word stage, children begin to utter
their first real words. A single word often conveys an entire thought or request (e.g., saying
"Milk!" when they are thirsty).
Two-Word Stage (18–24 months): Once a child's vocabulary reaches around 50 words, they
start forming basic two-word phrases (e.g., "more juice" or "daddy go").
Telegraphic Stage (24–30 months): Children experience a vocabulary spurt and create longer,
"pidgin-like" multi-word sentences. They drop function words (like 'the', 'a', or 'is') while
preserving core grammatical word order.
Complex/Multi-Word Stage (30+ months to 5 years): Grammar skills rapidly improve as
children begin using suffixes, prefixes, pronouns, and tenses. By age five, they can engage in
fluid conversations, follow multi-step instructions, and tell stories.
six dimensions
These six dimensions of acquisition highlight specific mental processes and cognitive
milestones:
Language acquisition is primarily explained by four major theoretical perspectives: Behaviorist, Nativist,
Cognitivist, and Interactionist. These theories explore the complex interplay between innate biological
faculties and environmental influences in how humans—especially children—develop language skills.
1. Behaviorist Perspective
Championed by B.F. Skinner, the behaviorist theory posits that language is acquired through operant
conditioning, environmental conditioning, imitation, and reinforcement. Children imitate the sounds they
hear. If their speech receives positive reinforcement (e.g., a smile, praise, or getting what they want), the
behavior is repeated and becomes a learned [Link] struggles to explain how children generate
novel, grammatically correct sentences they have never heard before.
2. Nativist Perspective
Developed by Noam Chomsky, this biological approach argues that humans are born with an innate,
hardwired mechanism for acquiring language, known as the Language Acquisition Device (LAD).The
LAD contains Universal Grammar, a foundational blueprint of linguistic rules common to all human
languages. This allows children to effortlessly absorb and produce the complex grammatical structures of
their native language despite receiving limited, unstructured input. While highly influential, it tends to
downplay the critical role of social interaction and environmental input.
3. Cognitive Perspective
Associated with Jean Piaget, this framework argues that language is simply one aspect of a child's overall
cognitive and intellectual development. Children must first develop the cognitive capacity to understand
concepts (e.g., object permanence, categorization, and cause-and-effect) before they can articulate them
through language. There are documented cases of individuals with severe cognitive impairments who still
possess excellent linguistic and grammatical skills, challenging the idea that language is strictly a
byproduct of cognition.
4. Interactionist Perspective
Integrating both biological and social factors, the interactionist perspective (developed by theorists like
Lev Vygotsky and Jerome Bruner) emphasizes that language develops from the innate desire to
communicate with others. Caregivers play a crucial role by supporting children's language attempts
through "scaffolding," simplified speech (often called child-directed speech), and routine (Bruner's
Language Acquisition Support System - LASS). It views language as a socio-historical process driven
by social interaction. Contemporary research typically embraces an interactionist view, recognizing that
while humans possess the biological wiring (nature) for language, it is the social environment (nurture)
that activates and shapes this ability.
The relationship between semantics and cognitive science centers on how the human mind
constructs, stores, and processes meaning. While traditional formal semantics treats meaning
as a set of logical, abstract truth-conditions mapping directly to the objective world, cognitive
science treats semantic structure as conceptual structure. In this view, language does not just
describe an objective reality; it reflects how the human brain categorizes and interacts with the
mind's internal conceptualizations.
Cognitive semantics is a major subfield within cognitive linguistics that bridges psychology,
neuroscience, and philosophy. It operates under several foundational principles:
Cognitive scientists and linguists study semantics through several prominent models:
Applications
Entailment is the foundational logical relationship where the truth of one sentence guarantees the
truth of another. While cognitive semantics focuses on how meaning is shaped by the human mind,
formal semantics relies heavily on entailment to define meaning itself: to know what a sentence means is
to know what else must be true if that sentence is true.
Entailment Logically required truth. Fails. The inference is A: "I bought a laptop."
destroyed. → Entails: I bought
something.
(If you didn't buy a laptop,
you didn't necessarily buy
nothing).
Types of Entailment
1. Lexical (Hyponymic) Entailment: Derived from the inherent meaning of words within a
hierarchy.
o Example: "She owns a retriever" entails "She owns a dog" (because a retriever is a
hyponym of dog).
2. Syntactic (Structural) Entailment: Derived from the grammatical construction of the sentence.
o Example: Active-to-passive transformations. "The chef prepared the meal" entails "The
meal was prepared by the chef.
In formal semantic theories, meaning is defined through truth conditions. This approach argues that
understanding the meaning of a sentence does not require looking into psychological concepts, but rather
understanding exactly what conditions must look like in the real world for that sentence to be true.
It address a central question in cognitive psychology: How does the human mind represent, store, and
process the meaning of words and concepts?
The actual object, entity, or state of affairs in the real (or a possible) world that a word points to.
For example, the actual physical planet Venus is a reference.
Sense (Sinn): The "mode of presentation" or the cognitive way the reference is understood by
the mind. The terms "The Morning Star" and "The Evening Star" have different senses (mental
concepts), even though they share the exact same reference (the planet Venus).
Processing Implications: In cognitive architecture, reference acts as the input from the world,
while sense is the mental output or semantic encoding produced by the brain to understand it.
The distinction between reference and sense describes how humans and machines map
language onto reality by separating the target of our thoughts from the concepts we use to
describe them. Reference is the objective entity, person, or object in the world that a word
points to, acting as the concrete destination of our communication. Sense, by contrast, is
the mental route, description, or mode of presentation chosen to reach that destination.
For example, while the phrases "the author of Hamlet" and "the Bard of Avon" share the
exact same real-world reference—William Shakespeare—they express entirely different
senses that highlight distinct aspects of his [Link] human cognition, this separation
allows us to learn new information through identity statements; hearing that "the morning
star is the evening star" is an astronomical revelation because it connects two different
conceptual routes (senses) to a single planet (the reference). This dynamic presents a
complex hurdle for computational models, which must use advanced algorithms to
determine when different linguistic expressions refer to the same entity in a text, or how a
single word's mathematical vector coordinates should shift to capture its exact contextual
meaning.
Mechanism: Concepts are treated as distinct "nodes" arranged in a logical, hierarchical tree
structure.
Example: The Collins and Quillian model places "Animal" at the top, branching down to "Bird,"
which branches down to "Canary".
Processing: Knowledge is retrieved via spreading activation—when you think of "canary,"
activation spreads to "bird" and "canary can sing".
Feature-Comparison Models
Mechanism: Concepts are stored not as unified nodes, but as lists of structured, binary features.
Attributes: Features are split into defining features (essential to the meaning, like
<has_wings> for a bird) and characteristic features (common but not mandatory, like
<can_fly> ).
Processing: The mind verifies meaning by mathematically calculating the overlap between
feature lists.
Cognitive and computational models of semantic processing explain how the mind and machines
organize knowledge through three primary architectural frameworks. Semantic Network
Theory models memory as a vast web of discrete conceptual nodes connected by relational
links, using a mechanism called spreading activation to explain why humans recognize related
words faster—a phenomenon mirrored computationally in structured databases like WordNet.
Shifting away from rigid nodes, Parallel Distributed Processing (PDP) or Connectionist
Models represent meaning as unique, distributed patterns of activity across interconnected,
neuron-like units. These systems learn implicitly by adjusting mathematical weights based on
error feedback, accurately mimicking human "graceful degradation" where brain damage or
dementia causes a gradual loss of specific details while preserving broad category knowledge.
Finally, Distributional or Vector Space Models operate on the statistical principle that words
appearing in similar contexts share similar meanings, mapping language into a high-dimensional
geometric space where spatial proximity reflects semantic similarity. This statistical framework
underpins modern Large Language Models and word embeddings, capturing how both human
learners and advanced computer algorithms infer the meaning of unfamiliar terms purely from
the company they keep.
Information processing models of the mind view human cognition through the lens of a
computer metaphor, framing the brain as organic hardware and the mind as software that inputs,
encodes, stores, retrieves, and outputs data.
The most enduring linear framework divides human memory into three sequential hardware
stations:
Sensory Memory: A high-capacity but ultra-short buffer that holds raw environmental
stimuli (visual iconic memory lasts $<0.5$ seconds; auditory echoic memory lasts $\
sim3-4$ seconds).
Short-Term/Working Memory: A conscious workspace with a strictly limited capacity
—traditionally cited as $7 \pm 2$ chunks of information—where data is actively
manipulated for up to 30 seconds unless maintained through rehearsal.
Long-Term Memory: An infinite, permanent storage repository where information is
indexed semantically, ready to be retrieved back into working memory when triggered by
internal or external cues.
The Stage Theory of Memory, or Atkinson-Shiffrin Model, frames human cognition as a linear
pipeline where environmental data must successfully transition through three distinct hardware
stations to achieve permanent retention. Processing begins in Sensory Memory, an ultra-short
buffer that captures massive amounts of raw sensory input—such as fleeting visual snapshots or
brief auditory echoes—which instantly decays unless selected by targeted attention. Information
that survives this filter enters Short-Term Memory (STM), a conscious but strictly limited
workspace capable of holding only about \(7 \pm 2\) structural chunks of data for roughly 15 to
30 seconds. Because STM is a severe cognitive bottleneck, information is easily bumped out by
incoming stimuli unless kept alive through active mental repetition, or encoded deeply via
meaningful association into Long-Term Memory (LTM). Once inside LTM, this stored
knowledge enjoys a theoretically infinite capacity and lifetime duration, remaining dormant until
a retrieval cue pulls a copy of the archive back into short-term awareness for behavioral use. This
strict architectural division is strongly supported by the serial position effect—where people best
recall the beginning and end of lists—and neuropsychological cases like Patient H.M., whose
fully intact short-term memory could no longer transfer new experiences into long-term storage.
Expanding on the limitations of a passive short-term buffer, this model treats working memory
as an active, multi-component processor managed by a Central Executive controller. This
system dynamically allocates attention between two slave storage subsystems: the Phonological
Loop (handling speech and sound-based data) and the Visuospatial Sketchpad (handling mental
imagery and spatial tracking).
An Episodic Buffer acts as a temporary integration hub, binding inputs from these subsystems
and long-term memory into coherent, chronological experiences. The Central Executive
Architecture, proposed by Alan Baddeley and Graham Hitch, reframes short-term memory from a passive
storage container into an active, multi-component workspace known as working memory. At the core of
this system sits the Central Executive, a limited-capacity supervisory controller that manages cognitive
processing rather than storing data itself. It functions like a dynamic coordinator: it drives attention, shifts
between tasks, selects processing strategies, and suppresses irrelevant distractions.
To avoid cognitive overload, the Central Executive delegates the maintenance of information to three
specialized slave systems: the Phonological Loop, which acts as an inner voice to rehearse verbal and
auditory data, the Visuospatial Sketchpad, which functions as an inner eye to manipulate visual features
and spatial layouts, and the Episodic Buffer, which temporarily integrates information from these
modalities along with long-term memory into a coherent, chronological stream of experience. This
architecture explains why human multi-tasking succeeds when using different modalities—such as
driving while listening to a podcast—but instantly suffers severe bottlenecks when two tasks compete for
the exact same slave processor or overwhelm the Central Executive's attentional capacity
Moving away from linear, step-by-step assembly lines, connectionist models argue that
information processing happens simultaneously across massive, overlapping neural pathways.
Instead of a single central processor retrieving data from a specific slot in memory, knowledge is
stored distributively across the strengths (weights) of connections between basic processing
units. When a stimulus occurs, the entire network shifts its activation pattern in parallel, allowing
for rapid pattern recognition, context adaptation, and human-like intuition. Connectionist and
Parallel Distributed Processing (PDP) models reject the idea of a centralized computer processor
or a single storage slot for a specific concept, arguing instead that human cognition occurs
simultaneously across a massive, interconnected network of simple processing units. Inspired by
the biological structure of the brain, these models represent information not as static symbols, but
as unique, distributed patterns of activation spreading across thousands of neuron-like units all at
once. Knowledge is stored implicitly within the network in the varying strengths, or weights, of
the connections between these units. When the network receives a stimulus, it processes the data
through parallel computational layers—input, hidden, and output—adjusting its internal
connections via error-feedback algorithms like backpropagation until a stable answer is reached.
This distributed architecture explains why human memory exhibits "graceful degradation,"
where brain injury or neurological decline causes a fuzzy, gradual loss of fine details while
keeping broad conceptual categories intact, rather than deleting an entire memory or word
cleanly out of existence.
Processing Operations
[ENVIRONMENTAL INPUT] ──► [ENCODING] ──► [STORAGE] ──► [RETRIEVAL] ──►
[BEHAVIORAL OUTPUT]
│ │ │
(Transformation (Maintenance (Accessing text/
into code) over time) experience)
Encoding: The initial transformation of sensory input into a structured mental code or
representation that the brain's internal architecture can interpret.
Storage: The maintenance and structural consolidation of encoded information over brief
or protracted intervals of time.
Retrieval: The intentional or automatic activation and extraction of stored data back into
conscious awareness to guide decision-making.
The Physical Symbol System Hypothesis (PSSH) and the Language of Thought (LoT)
Hypothesis are two cornerstone theories in classical cognitive science and symbolic AI.
Together, they argue that both human minds and digital computers process meaning by
manipulating discrete, rule-governed symbolic representations.
Proposed by Allen Newell and Herbert Simon in 1976, this hypothesis presents a foundational
requirement for intelligence.
The Core Premise: A physical symbol system has the necessary and sufficient means for
general intelligent action.
What is a "Symbol"?: Physical patterns (like tokens in a computer's memory or neural
inscriptions in the brain) that can be combined into complex structures.
The Engine: Intelligence is achieved entirely through the mechanical manipulation of these
symbols based on formal rules (algorithms), independent of what physical matter the system is
made of (substrate independence).
Proposed by philosopher Jerry Fodor in 1975, LoT (often called Mentalese) applies symbolic
computation specifically to human cognition and language.
The Core Premise: Thinking does not happen in natural languages (like English or Tamil) or
vague images, but in an internal, innate, language-like mental code.
Key Characteristics:
o Productivity: A finite set of mental symbols can generate an infinite number of novel thoughts.
o Systematicity: The ability to think one thought implies the ability to think structurally related
thoughts (e.g., if you can think "The cat chased the mouse," you can inherently think "The mouse
chased the cat").
o Compositionality: The meaning of a complex mental expression is determined by the meanings
of its component symbols and the rules used to combine them.
Core Pillars of LOTH
Fodor argued that mental representations are structural, language-like tokens. The theory relies
on four defining properties:
Compositionality: Complex mental states are built systematically from simpler constituent
parts. The meaning of the thought [JOHN] [LOVES] [MARY] is derived strictly from its individual
concepts and how they are structurally combined.
Systematicity: The capacity to produce/understand certain thoughts is intrinsically connected to
the capacity to produce/understand structurally related thoughts. If a mind can formulate the
thought "The scientist created the AI," it automatically possesses the architectural capability to
formulate "The AI created the scientist."
Productivity: A finite set of mental tokens and computational rules can generate an infinite
number of unique, novel thoughts, independent of direct environmental stimuli.
Logical Syntax: Thinking is fundamentally a process of computational symbol manipulation.
The mind responds only to the formal, physical shapes of these mental tokens (syntax), yet these
operations perfectly preserve the underlying truth-values of the ideas (semantics).
In the context of semantic processing, these classical architectures handle meaning through strict
formal syntax:
Sense: Represented as a structurally unique symbol or token string (e.g., MORNING_STAR vs.
EVENING_STAR). Each string carries its own distinct computational path and logical
relationships within the system.
Reference: Achieved when a symbol successfully "points to" or designates an external physical
object in the world (VENUS) through an indexical or causal link.
To apply the Symbolic Paradigm (rooted in the Physical Symbol System Hypothesis and the Language of
Thought) to semantic processing, a system must treat meaning as the rule-based manipulation of
discrete, explicit tokens.
In this paradigm, Sense is defined as a specific configuration of symbolic structures, while Reference is a
formal pointer mapping those structures to a model of the world.
The system translates the natural language input into unambiguous mental predicates or logical
assertions. It assigns unique internal identifiers (tokens) to ensure there is no lexical ambiguity.
The system looks up the tokens in its symbolic database. The Sense is the collection of logical relations,
rules, and attributes bound to that specific token.
The system maintains a distinct "World Model" database that catalogs real-world entities. To resolve the
reference, it executes an identity operation ( = ) based on its rule base:
[Execution]
Assert: IdentityAssertion(CONCEPT_MORNING_STAR, ENTITY_VENUS_PLANET)
Result: ReferentOf(CONCEPT_MORNING_STAR) == ReferentOf(CONCEPT_EVENING_STAR)
== RealWorldID_43921 ("Venus")
The Advantages
Perfect Explainability: Every step of semantic processing can be traced back to an explicit logical
rule. There are no "black box" weights.
Absolute Systematicity: If the system understands Like(John, Mary) , it inherently
understands Like(Mary, John) by swapping tokens within the syntactic template.
Deterministic Reference: References are exact. A token either points to a specific database
entity or it doesn't, leaving no room for "hallucinations."
The distributed paradigm completely alters how data is stored, represented, and updated compared to
classical symbolic machines.
Sub-symbolic Micro-features: Individual nodes in a neural layer do not represent whole words
or complete concepts (like "Apple" or "Car"). Instead, nodes represent minute, abstract micro-
features (e.g., has-wheels, is-red, organic, metallic).
Distributed Representations: A full concept exists only as a pattern of activation across a
massive layer of hidden units. The concept "Apple" is represented by the simultaneous activation
of specific nodes, while "Fire Engine" shares some of those overlapping nodes (is-red) but
diverges sharply on others (organic vs. metallic).
Superpositional Memory: Concepts are not filed away in specific, isolated memory addresses.
Multiple concepts are stored superpositioned within the exact same set of connection weights.
Adjusting a weight shifts the relationships of many concepts simultaneously.
Parallel Processing: Unlike a traditional CPU that processes commands sequentially (one step at
a time), neural architectures update all node activation levels in parallel, mirroring biological
brain networks.
When shifted into a neural or distributed environment, the classic semantic dichotomy of Sense and
Reference becomes entirely mathematical:
In connectionist frameworks and their modern scalable descendants (such as Transformer embeddings),
the Sense of a word is represented as a high-dimensional vector.
The meaning of a term is defined entirely by its mathematical relationship to other terms in the
system (distributional semantics).
Similarity of sense is computed geometrically, using metrics like Cosine Similarity to measure
the angle between vectors in high-dimensional space.
Contextual nuance is fluid; the vector for a word dynamically drifts based on surrounding token
activations, allowing the system to handle polysemy (e.g., distinguishing "financial bank" from
"river bank") seamlessly.
Inputting a messy, noisy, or incomplete set of features (e.g., has feathers, sings, yellow, lives in a
cage) acts as an initial nudge to a recurrent or deep network.
This input places the network's hidden layer state into a high-dimensional mathematical
landscape consisting of computational hills and valleys.
Through successive iterations, the network's weights pull the active state downward into a
stable valley called an attractor basin.
Settling into that specific basin represents the network arriving at a definitive categorization or
identifying the correct external referent (e.g., concluding: This is a Canary).
The distributed paradigm solves several fundamental engineering and cognitive bottlenecks that
plagued classical artificial intelligence:
The application of neural networks to model human mental faculties is known as Computational
Cognitive Modeling or Connectionist Cognitive Science. Instead of treating the brain as a digital
computer executing abstract code, these models simulate cognitive processes—such as reading,
memory retrieval, and language acquisition—using networks of simulated neurons.
Developed by McClelland and Rumelhart (1981), the IA model explains how humans perceive written
words.
The Architecture: A multi-layered, localist neural network containing three distinct processing
tiers: Visual Features → Letters → Words.
Cognitive Insight: The model uses top-down feedback loops (excitatory and inhibitory
connections) from the Word layer back to the Letter layer. This mathematically explains the
Word-Superiority Effect—the cognitive phenomenon where humans identify a letter (e.g., 'T')
much faster when it is embedded in a real word ("PROP") than when it appears in a random string
of letters ("PZOR").
Developed by Seidenberg and McClelland (1989), the Triangle Model describes how the brain processes
written text, sound, and meaning.
One of connectionism's historic breakthroughs was modeling how children learn the English past tense
(e.g., changing "walk" to "walked", but "go" to "went").
The Cognitive Phenomenon: Children display a U-shaped learning curve. First, they correctly
use irregulars ("broke"). Next, as they generalize the rules, their performance drops and they
over-regularize ("broked"). Finally, they master both categories.
The Neural Solution: Rumelhart and McClelland (1986) demonstrated that a simple pattern-
associator network trained on verb pairs naturally mimics this exact U-shaped curve. As the
network shifts from memorizing individual examples to adjusting its global weights to capture
statistical regularities, it naturally goes through an over-regularization phase, matching human
developmental data without needing explicit "if/then" rule structures.
Primary Goal To explain the human mind: Mirror To maximize performance: Solve complex
human errors, reaction times, and operational tasks with high efficiency and
developmental stages. accuracy.
Data Scaling Human-constrained: Must learn from Brute-force scaling: Requires trillions of
the limited volume of data a human child tokens harvested across the entire
hears/sees in early life. internet.
Error Model is successful if its mistakes match Model is successful if it minimizes errors
Matching human psychological errors. completely.