Models of Language Emergence Explained
Models of Language Emergence Explained
net/publication/5284369
CITATIONS READS
436 7,382
1 author:
Brian Macwhinney
Carnegie Mellon University
369 PUBLICATIONS 31,161 CITATIONS
SEE PROFILE
All content following this page was uploaded by Brian Macwhinney on 05 November 2014.
2-1-1998
Recommended Citation
MacWhinney, Brian, "Models of the Emergence of Language" (1998). Department of Psychology. Paper 169.
[Link]
This Article is brought to you for free and open access by the College of Humanities and Social Sciences at Research Showcase. It has been accepted
for inclusion in Department of Psychology by an authorized administrator of Research Showcase. For more information, please contact
kbehrman@[Link].
Models of the Emergence of Language
Brian MacWhinney
Department of Psychology, Carnegie Mellon University, Pittsburgh, PA 15213,
macw@[Link]
Appeared in: Annual Review of Psychology 1998 49: 199-227
1
Models of the Emergence of Language
2
Models of the Emergence of Language
they listen to. Using the sucking habituation technique, Mandel et al. (1994) showed that
2-month-olds tend to remember word strings better when they are presented with normal
sentence intonation, than when they are presented as unintegrated lists of words with flat
prosody. It appears that stressed intonation may have a particularly important role in
picking up auditory strings. Jusczyk and Pisoni (1995) have shown that children tend to
pick up and learn stressed syllables above unstressed syllables. However, it also appears
that syllables which directly follow after a stressed syllable are also well encoded (Aslin,
Jusczyk & Pisoni, 1997). As a result, many of the first sound sequences recorded by the
child consist of a stressed peak followed by one or two further weak syllables. This pattern
of sound learning has been discussed as a “trochaic bias”. However, it can also be viewed
as emerging from the combination of a bias to track stressed syllables together with a linear
sequence recorder that fires when a stressed syllable is detected.
3
Models of the Emergence of Language
Lexical principles
Markman (1989) and Golinkoff, Mervis, and Hirsh-Pasek (1994) have proposed that
Quine’s problem can be solved by imagining that the child’s search for word meanings is
guided by lexical principles. For example, children assume that words refer to whole
objects, rather than parts of objects. Thus, a child would assume that the word “rabbit”
refers to the whole rabbit and not just some parts of the rabbit. However, there is reason to
believe that such principles are themselves emergent properties of the cognitive system. For
example, Merriman and Stevenson (1997) have argued that the tendency to avoid learning
two names for the same object emerges naturally from the competition (MacWhinney,
1989) between closely-related lexical items.
Another proposed lexical principle is the tendency to focus on object names and
nominal categories over other parts of speech. Gentner (1982) compared the relative use of
nominal terms, predicative terms, and expressive terms in English, German, Japanese, Kaluli,
and Turkish. She found that, in all five languages, words for objects constituted the largest
group of words learned by the child. Like Gentner, Tomasello (1992) has argued that
nouns are easier to “package” cognitively than verbs. Nouns refer to objects that can be
repeatedly touched and located in space, whereas verbs refer to transitory actions that are
often hard to repeat and whose contour varies markedly for different agents. However,
Gopnik and Choi (1990) and Choi and Bowerman (1991) have reported that the first words
of Korean-speaking children include far more verbs than do those of English-speaking
children. Findings of this type indicate that the nominal bias emerges only in languages
that tend to emphasize nouns.
Even in English, we know that children will often treat a new word as a verb or an
adjective (Hall, Waxman & Hurwitz, 1993), because words like “run”, “want”, “hot” and
“good” are included in some of the child’s first words. Children are also quick to pick up
socially-oriented words such as “hi” and “please”. As Bloom, Tinker, and Margulis
(1993) and Vihman and McCune (1994) have argued, the nominal bias is far from a
predominant force, even in English.
Social support
The idea that early word learning depends heavily on the spatio-temporal contiguity
of a novel object and a new name can be traced back to Aristotle, Plato, and Augustine.
Recently, Baldwin (1991; 1989) has shown that children try to acquire names for the objects
that adults are attending to. Similarly, Akhtar, Carpenter, and Tomasello (1996) and
Tomasello and Akhtar (1995) have emphasized the crucial role of mutual gaze between
mother and child in the support of early word learning. Moreover, Tomasello has argued
that human mothers differ significantly from primate mothers in the ways that they
encourage mutual attention during language. While not rejecting the role of social support
in language learning, Samuelson and Smith (in press) have noted that one can also interpret
the findings of Akhtar, Carpenter, and Tomasello in terms of low-level perceptual and
attentional matches that help focus the child’s attention to novel objects to match up with
new words.
Child-based meanings
Several researchers have emphasized the extent to which the shape of the meanings of
the first words is governed by a “child-based agenda” (Mervis, 1984; Slobin, 1985).
Children seem to be particuarly interested in finding ways of talking about their favorite
toys, friends, and foods (Dromi, 1997). They also like to learn words to discuss social
activities and functions. In fact, Ninio and Snow (1988) have argued that the basic
orientation of the child’s first words and early grammar is not towards some objective,
nominal, cognitive reality, but towards the interpersonal world involving people and social
roles.
4
Models of the Emergence of Language
5
Models of the Emergence of Language
as “Here’s the nice (toy name)” or “Show me your (body part name)”. Having learned
these frames, children can quickly pick up a large quantity of new words in the context of
each frame. In this way, the vocabulary spurt could be dependent upon syntactic
development. In fact, Bates et al, (1988) reported a correlation of between .70 and .84
between lexical size at 20 months and syntactic abilities at 28 months. This level of
correlation is exactly what is predicted by a model that views lexical learning as facilitated
by the appearance of words in the context of well-understood syntactic frames.
In accord with the Piagetian emphasis on cognitive determination of developmental
stages, a third group of authors has attributed the vocabulary spurt to the underlying growth
in those cognitive capacities (Bloom, 1970; Gopnik & Meltzoff, 1987) that allow children to
understand the meanings of new words. For example, one could argue that 14-month-olds
are not yet ready conceptually to acquire the meanings of comparative adjectives,
conjunctions, abstract nouns, speech act verbs, and superordinates. To be sure, very young
children have not yet acquired complex relational concepts, such as the ones required to
support the learning of form like “nonetheless”, “preamble”, or “next Thursday”
(Kenyeres, 1926). However, attempts to relate overall aspects of linguistic development to
fundamental changes in cognitive development have seldom demonstrated strong linkages
(Corrigan, 1978; Corrigan, 1979). Instead, it appears that the links between cognitive and
lexical development are fragmentary and specific to particular lexical fields (Gopnik &
Meltzoff, 1986).
Each of these three accounts is compatible with attempts (Bates & Carnevale, 1993; van
Geert, 1991) to model vocabulary growth as a dynamic system using logistic growth
functions. The nonlinear effects that emerge during the vocabulary spurt can be viewed as
arising from the dynamic coupling of the lexical system with a quickly developing system
of syntactic patterns, phonological advances, or cognitive advances. As these various
patterns develop, they feed into vocabulary growth in a nonlinear and interactive fashion, as
growth in vocabulary leads to further growth in syntactic structures, at least during the
several months of the vocabulary spurt.
6
Models of the Emergence of Language
Dyer (1990; 1991). These self-organizing networks treat word learning as occurring in
maps of connected neurons in small areas of the cortex. Three local maps are involved in
word learning: an auditory map, a concept map, and articulatory maps. Emergent self-
organization on each of these three maps uses the same learning algorithm. Word learning
involves the association of elements between these three maps. What makes this mapping
process self-organizing is the fact that there is no pre-established pattern for these mappings
and no preordained relation between particular nodes and particular feature patterns.
Evidence regarding the importance of syllables in early child language (Bijeljac,
Bertoncini & Mehler, 1993; Jusczyk, Jusczyk, Kennedy, Schomberg & Koenig, 1995)
suggests that the nodes on the auditory map may best be viewed as corresponding to full
syllabic units, rather than separate consonant and vowel phonemes. The recent
demonstration by Saffran et al. (1996) of memory for auditory patterns in four-month-old
infants indicates that children are not only encoding individual syllables, but are also
remembering sequences of syllables. In effect, prelinguistic children are capable of
establishing complete representations of the auditory forms of words. Within the self-
organizing framework, these capabilities can be represented in two alternative ways. One
method uses a slot-and-frame featural notation from MacWhinney, Leinbach, Taraban, and
McDonald (1989). An alternative approach views the encoding as a temporal pattern that
repeatedly accesses a basic syllable map. A lexical learning model developed by Gupta and
MacWhinney (1997) uses serial processes to control word learning. This model couples a
serial order mechanism known as an “avalanche” (Grossberg, 1978) with a lexical feature
map model. The avalanche controls the order of syllables within the word. Each new word
is learned as a new avalanche.
The initial mapping process involves the association of auditory units to conceptual
units. Initially, this learning links concepts to auditory images (Naigles & Gelman, 1995;
Reznick, 1990). For example, the 14-month-old who has not yet produced the first word,
may demonstrate an understanding of the word “dog” by turning to a picture of a dog,
rather than a picture of a cat, when hearing the word “dog”. It is difficult to measure the
exact size of this comprehension vocabulary in the weeks preceding the first productive
word, but it is probably at least 20 words in size.
In the self-organizing framework, the learning of a word is viewed as the emergence of
an association between a pattern on the auditory map and a pattern on the concept map
through Hebbian learning (Hebb, 1949; Kandel & Hawkins, 1992). When the child hears a
given auditory form and sees an object at the same time, the coactivation of the neurons that
respond to the sound and the neurons that respond to the visual form produces an
association across a third pattern of connections which maps auditory forms to conceptual
forms. Initially, the pattern of these interconnections is unknown, because the relation
between sounds and meanings is arbitrary (de Saussure, 1966). This means that the vast
majority of the many potential connections between the auditory and conceptual maps will
never be used, making it a very sparse matrix (Kanerva, 1993). In fact, it is unlikely that all
units in the two maps are fully interconnected (Shrager & Johnson, 1995). In order to
support the initial mapping, some researchers (Schmajuk & DiCarlo, 1992) have suggested
that the hippocampus may provide a means of maintaining the association until additional
cortical connections have been established. As a result, a single exposure to a new word is
enough to lead to one trial learning. However, if this initial association is not supported by
later repeated exposure to the word in relevant social contexts, the child will no longer
remember the word.
7
Models of the Emergence of Language
auditions is a major challenge. The simple control of the articulatory system is still a major
challenge for the two-year-old. Apart from this, the child have must acquire a mapping from
individual auditory features to articulatory gestures, and must also encode the sequence and
prosodic contour of each of the syllables in the word. Just like the learning of auditory
sequences requires the mediation of memory systems, the learning of articulatory sequences
may involve support from rehearsal loops or hippocampal systems.
Models of word learning in adults (Burgess & Hitch, 1992; Grossberg, 1978;
Grossberg, 1987; Houghton, 1990)} have tended to emphasize the role of working memory.
Gupta and MacWhinney (1997) have shown that a model based on the encoding of syllable
strings for output phonology in avalanches does a good job of accounting for a wide variety
of well-researched phenomena in the literature on word learning, immediate serial recall,
interference effects, and rehearsal in both adults and children (Gathercole & Baddeley,
1993).
8
Models of the Emergence of Language
knowing that balls are round and can roll or knowing that tables are flat and that rolling
involves movement.
9
Models of the Emergence of Language
200 hidden
7 units
HIDDEN UNITS
20200
gender/
hidden
10 case units
number units
As noted above, a central feature of such connectionist models is the very large number
of connections among processing units. As shown in Figure 1, each input-level unit is
connected to first-level hidden units; each first-level hidden unit is connected to second-level
hidden units; and each second-level hidden unit is connected to each of the six output units.
None of these hundreds of individual node-to-node connections is illustrated in Figure 1,
since graphing each individual connection would lead to a blurred pattern of connecting
lines. Instead a single line is used to stand in place of a fully interconnected pattern between
levels. Learning is achieved by repetitive cycling through three steps. First, the system is
presented with an input pattern that turns on some, but not all of the input units. In this
case, the pattern is a set of sound features for the noun being used. Second, the activations
of these units send activations through the hidden units and on to the output units. Third,
the state of the output units is compared to the correct target and, if it does not match the
target, the weights in the network are adjusted so that connections that suggested the correct
answer are strengthened and connections that suggested the wrong answer are weakened.
MacWhinney et al. tested this system’s ability to master the German article system by
repeatedly presenting 102 common German nouns to the system. Frequency of presentation
of each noun was proportional to the frequency with which the nouns are used in German.
The job of the network was to choose which article to use with each noun in each particular
context. After it did this, the correct answer was presented, and the simulation adjusted
connection strengths so as to optimize its accuracy in the future. After training was
finished, the network was able to choose the correct article for 98 percent of the nouns in the
original set.
To test its generalization abilities, we presented the network with old nouns in new case
roles. In these tests, the network chose the correct article on 92 percent of trials. This type
of cross-paradigm generalization is clear evidence that the network went far beyond rote
memorization during the training phase. In fact, the network quickly succeeded in learning
the whole of the basic formal paradigm for the marking of German case, number, and
gender on the noun. In addition, the simulation was able to generalize its internalized
knowledge to solve the problem that had so perplexed Mark Twain -- guessing at the gender
of entirely novel nouns. The 48 most frequent nouns in German that had not been included
10
Models of the Emergence of Language
in the original input set were presented in a variety of sentence contexts. On this completely
novel set, the simulation chose the correct article from the six possibilities on 61 percent of
trials, versus 17 percent expected by chance. Thus, the system’s learning mechanism,
together with its representation of the noun's phonological and semantic properties and the
context, produced a good guess about what article would accompany a given noun, even
when the noun was entirely unfamiliar.
The network’s learning paralleled children’s learning in a number of ways. Like real
German-speaking children, the network tended to overuse the articles that accompany
feminine nouns. The reason for this is that the feminine forms of the article have a high
frequency, because they are used both for feminines and for plurals of all genders. The
simulation also showed the same type of overgeneralization patterns that are often
interpreted as reflecting rule use when they occur in children’s language. For example,
although the noun Kleid (which means clothing) is neuter, the simulation used the initial
“kl” sound of the noun to conclude that it was masculine. Because of this, it invariably
chose the article that would accompany the noun if it were masculine. Interestingly, the same
article-noun combinations that are the most difficult for children proved to be the most
difficult for the simulation to learn and to generalize to on the basis of previously learned
examples.
How was the simulation able to produce such generalization and rule-like behavior
without any specific rules? The basic mechanism involved adjusting connection strengths
between input, hidden, and output units to reflect the frequency with which combinations of
features of nouns were associated with each article. Although no single feature can predict
which article would be used, various complex combinations of phonological, semantic, and
contextual cues allow quite accurate prediction of which articles should be chosen. This
ability to extract complex, interacting patterns of cues is a characteristic of the particular
connectionist algorithm, known as back-propagation, that was used in the MacWhinney et
al. simulations. What makes the connectionist account for problems of this type
particularly appealing is the fact that an equally powerful set of production system rules for
German article selection would be quite complex (Mugdan, 1977) and learning of this
complex set of rules would be a challenge in itself.
11
Models of the Emergence of Language
U-shaped learning
A major shortcoming of nearly all connectionist models of inflectional learning has been
their inability to capture the patterns of overgeneralization and recovery from
overgeneralization that have been called “u-shaped” learning. In u-shaped learning, the
child begins by correctly producing an irregularly inflected form such as “went”. Next,
under the pressure of the general pattern, the child produces the overgeneralized form
“goed”. Finally, the child recovers from overgeneralization and returns to saying “went”.
Some writers have mistakenly assumed that this type of u-shaped learning applies across all
verbs to create three major periods in language learning. However, empirical work by
Marcus, Ullman, Pinker, Hollander, Rosen, and Xu (1992) has shown that strong u-shaped
learning patterns occur only for some verbs and only for some children.
The modelling of even these weaker u-shaped patterns has proven difficult for neural
networks. In order to correctly model the child’s learning of inflectional morphology,
models must go through a period of virtually error free learning of irregulars, followed by a
period of learning of regulars accompanied by the first overregularizations (Marcus et al.,
1992). No current model consistently displays all of these features in exactly the right
combination. MacWhinney (1997) has argued that models that rely exclusively on
backpropagation will never be able to display the correct combination of developmental
patterns and that a two-process connectionist approach may be needed (Kawamoto, 1994;
Stone, 1994). The basic process is one that learns new inflectional formations, both regular
and irregular, as items in self-organizing feature maps. The secondary process is a network
that generalizes the information inherent in feature maps to extract secondary productive
generalizations. Unlike Pinker’s dual-route account, this proposed account works on a
uniform underlying connectionist architecture without relying on formal, symbolic linguistic
rules.
12
Models of the Emergence of Language
connectionist symbolic pattern associator which did a better job modeling the Prasada and
Pinker data. However, MacWhinney (1993a) found that the network model of
MacWhinney and Leinbach (1991) worked as well as Ling and Marinov’s symbolic model
in terms of matching up to the Prasada and Pinker generalization data.
go + PAST
went competition go + ed
emergent
episodic lexical
support properties
13
Models of the Emergence of Language
and the child recovers from the overgeneralization. This is done without negative evidence,
solely on the basic of positive support for the form receiving episodic confirmation.
14
Models of the Emergence of Language
category structure. Although not all nouns are objects, the best or most prototypical nouns
all share this feature. As the category of “noun” radiates out (Lakoff, 1987), non-central
members start to share fewer of the core features of the prototype. Maratsos and Chalkley
(1980) point out that words like “justice” and “lightning” are so clearly non-objects that
their membership in the class of nouns cannot be predicted from their semantic status and
can only be inferred from the fact that the language treats them as nouns. Although Bates
and MacWhinney (1982) and Maratsos and Chalkley (1980) staked out strongly
contrasting positions on this issue, each of the approaches granted the possibility that both
cooccurence and semantic factors play a major role in the emergence of the parts of speech.
At this point, language researchers are primarily interested in exploring detailed models
that show exactly how the parts of speech and argument frames can be induced. Elman
(1993) has presented a connectionist model that does just this. The model relies on a
recurrent architecture of the type presented in Figure 3. This model takes the standard
three-layer architecture of pools A, B, and C and adds a fourth input pool D of context units
which has recurrent connections to pool B. Because of the recurrent or bidirectional
connections between B and D, this architecture is know as “recurrent backpropagation”.
A pr edict category
B int er nal st at e
C D
15
Models of the Emergence of Language
that model, part-of-speech information is assumed and the goal of the model is to select the
agent and the patient using a variety of grammatical and pragmatic cues.
The training set for the model consists of dozens of simple English sentences such as
“The big dog chased the girl.” By examining the weight patterns on the hidden units in the
fully trained model, Elman showed that the model was conducting implicit learning of the
parts of speech. For example, after the word “big” in our example sentence, the model
would be expecting to activate a noun. The model was also able to distinguish between
subject and object relative structures, as in “the dog the cat chased ran” and “the dog that
chased the cat ran”.
CONCLUSION
In this chapter, we have seen how neural network models can help us organize our
growing understanding of auditory, articulatory, lexical, inflectional, and syntactic
development. There are many aspects of language development to which these models have
not yet been applied. We do not yet have models that can learn to control sociolinguistic
relations, conversational patterns, narrative structures, intonational contours, and gestural
markings. Even in the areas to which they have been applied, emergentist models are
limited in many ways. The treatment of the more complex aspects of syntax remains
unclear, the modelling of lexical extensions is still quite primitive, and the development of
the auditory and articulatory systems is not yet sufficiently grounded in physiological and
neurological facts. Despite these limitations, we can see that, by treating language learning
as an emergent process, these models have succeeded in providing an exciting new
perspective on questions about language learning that have intrigued scholars for centuries.
16
Models of the Emergence of Language
LITERATURE CITED
Akhtar, N., Carpenter, M., & Tomasello, M. (1996). The role of discourse novelty in early
word learning. Child Development, 62, 635-645.
Allen, R., & Gardner, B. (1969). Teaching sign language to a chimpanzee. Science, 165,
664-672.
Anglin, J. M. (Ed.). (1977). Word, object, and conceptual development. New York: Norton.
Aslin, R., Jusczyk, P., & Pisoni, D. (1997). Speech and auditory processing during infancy:
Constraints on and precursors to language. In D. Kuhn & R. Siegler (Eds.), Handbook
of child psychology. Volume 2, . New York: Wiley.
Atkinson, K., MacWhinney, B., & Stoel, C. (1970). An experiment on the recognition of
babbling. Papers and Reports on Child Language Development, 5, 1-8.
Atkinson, M. (1992). Children's syntax. Oxford: Blackwells.
Baker, C. L. (1979). Syntactic theory and the projection problem. Linguistic Inquiry, 10,
533-581.
Baker, C. L., & McCarthy, J. J. (Eds.). (1981). The logical problem of language acquisition.
Cambridge: MIT Press.
Baldwin, D. A. (1991). Infants' contribution to the achievement of joint reference. Child
Development, 62, 875-890.
Baldwin, D. A., & Markman, E. M. (1989). Establishing word-object relations: A first step.
Child Development, 60, 381-398.
Barrett, M. (1995). Early lexical development. In P. Fletcher & B. MacWhinney (Eds.),
Handbook of Child Language, . Oxford: Basil Blackwell.
Bates, E., Bretherton, I., & Snyder, L. (1988). From first words to grammar: Individual
differences and dissociable mechanisms. Cambridge, MA: Cambridge University Press.
Bates, E., & Carnevale, G. (1993). New directions in research on language development.
Developmental Review, 13, 436-470.
Bates, E., & MacWhinney, B. (1982). Functionalist approaches to grammar. In E. Wanner
& L. Gleitman (Eds.), Language acquisition: The state of the art, (pp. 173-218). New
York: Cambridge University Press.
Bechtel, W., & Abrahamsen, A. (1991). Connectionism and the mind: An introduction to
parallel processing in networks. Cambridge, MA: Basil Blackwell.
Bickerton, D. (1990). Language and species. Chicago: Chicago University Press.
Bijeljac, B., R., Bertoncini, J., & Mehler, J. (1993). How do four-day-old infants categorize
multisyllabic utterances? Developmental Psychology, 29, 711-721.
Bloom, L. (1970). Language development: Form and function in emerging grammars.
Cambridge, MA: MIT Press.
Bloom, L. (1993). The transition from infancy to language: Acquiring the power of
expression. Cambridge: Cambridge University Press.
Bloom, L., Tinker, E., & Margulis, C. (1993). The words children learn: Evidence against a
noun bias in early vocabularies. Cognitive Psychology, 8, 431-450.
Bloom, P. (1994). Overview: Controversies in language acquisition. In P. Bloom (Ed.),
Language acquisition: Core readings, (pp. 5-48). Cambridge, MA: MIT Press.
Bowerman, M. (1982). Reorganizational processes in lexical and syntactic development. In
E. Wanner & L. Gleitman (Eds.), Language acquisition: The state of the art, (pp. 319-
346). New York: Cambridge University Press.
Bowerman, M. (1988). The "no negative evidence" problem. In J. Hawkins (Ed.),
Explaining language universals, (pp. 73-104). London: Blackwell.
17
Models of the Emergence of Language
18
Models of the Emergence of Language
Gentner, D. (1982). Why nouns are learned before verbs: Linguistic relativity versus natural
partitioning. In S. Kuczaj (Ed.), Language development: Language, culture, and
cognition, (pp. 301-334). Hillsdale, NJ: Lawrence Erlbaum.
Gleitman, L. (1990). The structural sources of verb meanings. Language Acquisition, 1, 3-
55.
Gleitman, L. R., Newport, E. L., & Gleitman, H. (1984). The current status of the motherese
hypothesis. Journal of Child Language, 11, 43-79.
Goldberg, A. (1995). Constructions. Chicago: University of Chicago Press.
Golinkoff, R., Hirsh-Pasek, K., Cauley, K., & Gordon, L. (1987). The eyes have it: lexical
and syntactic comprehension in a new paradigm. Journal of Child Language, 14, 23-46.
Golinkoff, R. M., Mervis, C. B., & Hirsh-Pasek, K. (1994). Early object labels: The case
for a developmental lexical principles framework. Journal of Child Language, 21, 125-
155.
Gopnik, A., & Choi, S. (1990). Do linguistic differences lead to cognitive differences? A
crosslinguistic study of semantic and cognitive development. First Language, 10, 199-
215.
Gopnik, A., & Meltzoff, A. (1987). The development of categorization in the second year
and its relation to the other cognitive and linguistic developments. Child Development,
58, 1523-1531.
Gopnik, A., & Meltzoff, A. N. (1986). Relations between semantic and cognitive
development in the one-word stage: The Specificity Hypothesis. Child Development, 57,
1040-1053.
Gopnik, M. (1990). Feature blindness: A case study. Language Acquisition, 1, 139-164.
Grossberg, S. (1978). A theory of human memory: Self-organization and performance of
sensory-motor codes, maps, and plans. Progress in Theoretical Biology, 5, 233-374.
Grossberg, S. (1987). Competitive learning: From interactive activation to adaptive
resonance. Cognitive Science, 11, 23-63.
Gupta, P., & MacWhinney, B. (1992). Integrating category acquisition with inflectional
marking: A model of the German nominal system, Proceedings of the Fourteenth
Annual Conference of the Cognitive Science Society, (pp. 253-258). Hillsdale, NJ:
Lawrence Erlbaum Associates.
Gupta, P., & MacWhinney, B. (1997). Vocabulary acquisition and verbal short-term
memory: Computational and neural bases. Brain and Language, 59, 267-333.
Hall, D., Waxman, S., & Hurwitz, W. (1993). How two- and four-year-old children
interpret adjectives and count nouns. Child Development, 64, 1651-1664.
Harris, C. (1990). Connectionism and cognitive linguistics. Connection Science, 2, 7-33.
Harris, C. L. (1994). Back-propagation representations for the rule-analogy continuum. In
J. Barnden & K. Holyoak (Eds.), Analogical connections, (pp. 282-326). Norwood, NJ:
Ablex.
Harris, M., Barrett, M. D., Jones, D., & Brookers, S. (1988). Linguistic input and early
word meaning. Journal of Child Language, 15, 77-94.
Hebb, D. (1949). The organization of behavior. New York: Wiley.
Houghton, G. (1990). The problem of serial order: A neural network model of sequence
learning and recall. In R. Dale, C. Mellish, & M. Zock (Eds.), Current research in
natural language generation, (pp. 287-319). London: Academic.
Hunt, E. (1962). Concept learning: an information processing approach. New York: Wiley.
Huttenlocher, J. (1974). The origins of language comprehension. In R. Solso (Ed.),
Theories in cognitive psychology: The Loyola symposium, (pp. 331-388). Potomac,
Maryland: Lawrence Erlbaum.
Hyams, N. (1995). Nondiscreteness and variation in child language: Implications for
Principle and Parameter modesl of language development. In Y. Levy (Ed.), Other
children, other languages, (pp. 11-40). Hillsdale, NJ: Lawrence Erlbaum.
Hyams, N., & Wexler, K. (1993). On the grammatical basis of null subjects in child
language. Linguistic Inquiry, 24(3), 421-459.
19
Models of the Emergence of Language
Jaeger, J. J., Lockwood, A. H., Kemmerer, D. L., Van Valin, R. D., & Murphy, B. W.
(1996). A positron emission tomographic study of regular and irregular verb
morphology in English. Language, 72, 451-497.
Jakobson, R. (1968). Child language, aphasia and phonological universals. The Hague:
Mouton.
Jusczyk, P., & Aslin, R. (1995). Infants' detection of the sound patterns of words in fluent
speech. Cognitive Psychology, 29, 1-23.
Jusczyk, P. W., Jusczyk, A. M., Kennedy, L. J., Schomberg, T., & Koenig, N. (1995).
Young infants' retention of information about bisyllabic utterances. Journal of
Experimental Psychology: Human Perception and Performance, 21, 822-836.
Kandel, E. R., & Hawkins, R. D. (1992). The biological basis of learning and individuality.
Scientific American, 266, 40-53.
Kanerva, P. (1993). Sparse distributed memory and related models. In M. Hassoun (Ed.),
Associative neural memories: Theory and implementation, . New York: Oxford
University Press.
Katz, N., Baker, E., & Macnamara, J. (1974). What's in a name? A study of how children
learn common and proper names. Child Development, 45, 469-473.
Kawamoto, A. (1994). One system or two to handle regulars and exceptions: How time-
course of processing can inform this debate. In S. D. Lima, R. L. Corrigan, & G. K.
Iverson (Eds.), The reality of linguistic rules, (pp. 389-416). Amsterdam: John
Benjamins.
Kenyeres, E. (1926). A gyermek elsö szavai es a szófajók föllépése. Budapest:
Kisdednevelés.
Kohonen, T. (1982). Self-organized formation of topologically correct feature maps.
Biological Cybernetics, 43, 59-69.
Köpcke, K.-M. (1994). Funktionale Untersuchungen zur deutschen Nominal- und
Verbalmorphologie. Linguistiche Arbeiten, 319, 81-95.
Köpcke, K. M., & Zubin, D. A. (1983). Die kognitive Organisation der Genuszuweisung zu
den einsilbigen Nomen der deutschen Gegenwartssprache. Zeitschrift für
germanistische Linguistik, 11, 166-182.
Köpcke, K. M., & Zubin, D. A. (1984). Sechs Prinzipien fur die Genuszuweisung im
Deutschen: ein Beitrag zur natürlichen Klassifikation. Linguistische Berichte, 93, 26-50.
Kuhl, P. K. (1991). Human adults and human infants show a "perceptual magnet effect" for
the prototypes of speech categories, monkeys do not. Perception and Psychophysics,
50, 93-107.
Kuhl, P. K., & Miller, J. D. (1975). Speech perception by the chinchilla: Voiced-voiceless
distinction in alveolar plosive consonsants. Science, 190, 69-72.
Kuhl, P. K., & Miller, J. D. (1978). Speech perception by the chinchilla: Identification
functions for synthetic VOT stimuli. Journal of the Acoustical Society of America, 63,
905-917.
Kuhl, P. K., & Padden, D. M. (1982). Enhanced discriminability at the phonetic boundaries
for the voicing feature in macaques. Perception and Psychophysics, 32, 542-550.
Kuhl, P. K., & Padden, D. M. (1983). Enhanced discriminability at the phonetic boundaries
for the voicing feature in macaques. Journal of the Acoustical Society of America, 73,
1003-1010.
Lachter, J., & Bever, T. (1988). The relation between linguistic structure and associative
theories of language learning: A constructive critique of some connectionist learning
models. Cognition, 28, 195-247.
Lakoff, G. (1987). Women, fire, and dangerous things. Chicago: Chicago University Press.
Landau, B., Smith, L., & Jones, S. (1992). Syntactic context and the shape bias in children's
and adults' lexical learning. Journal of Memory and Language, 31, 807-825.
Levitt, A. G., Utman, J., & Aydelott, J. (1993). From babbling towards the sound systems of
English and French: A longitudinal two-case study. Journal of Child Language, 19, 19-
49.
20
Models of the Emergence of Language
Lewis, M. M. (1936). Infant speech: A study of the beginnings of language. New York:
Harcourt, Brace and Co.
Li, P., & MacWhinney, B. (1996). Cryptotype, overgeneralization, and competition: A
connectionist model of the learning of English reversive prefixes. Connection Science, 8,
3-30.
Ling, C., & Marinov, M. (1993). Answering the connectionist challenge. Cognition, 49,
267-290.
MacWhinney, B. (1978). The acquisition of morphophonology. Monographs of the
Society for Research in Child Development, 43, Whole no. 1, pp. 1-123.
MacWhinney, B. (1982). Basic syntactic processes. In S. Kuczaj (Ed.), Language
acquisition: vol 1. Syntax and semantics, (pp. 73-136). Hillsdale, NJ: Lawrence
Erlbaum.
MacWhinney, B. (1984). Where do categories come from? In C. Sophian (Ed.), Child
categorization, (pp. 407-418). Hillsdale, N.J.: Lawrence Erlbaum.
MacWhinney, B. (1988). Competition and teachability. In R. Schiefelbusch & M. Rice
(Eds.), The teachability of language, (pp. 63-104). New York: Cambridge University
Press.
MacWhinney, B. (1989). Competition and lexical categorization. In R. Corrigan, F.
Eckman, & M. Noonan (Eds.), Linguistic categorization, (pp. 195-242). New York:
Benjamins.
MacWhinney, B. (1993a). Connections and symbols: Closing the gap. Cognition, 49, 291-
296.
MacWhinney, B. (1993b). The (il)logical problem of language acquisition, Proceedings of
the Fifteenth Annual Conference of the Cognitive Science Society, (pp. 61-70).
Hillsdale, NJ: Lawrence Erlbaum Associates.
MacWhinney, B. (1997). Lexical connectionism. In P. Broeder & J. Murre (Eds.), Models
of language acquisition: Inductive and deductive approaches, . Cambridge, MA: MIT
Press.
MacWhinney, B., & Bates, E. (Eds.). (1989). The crosslinguistic study of sentence
processing. New York: Cambridge University Press.
MacWhinney, B., & Leinbach, J. (1991). Implementations are not conceptualizations:
Revising the verb learning model. Cognition, 29, 121-157.
MacWhinney, B. J., Leinbach, J., Taraban, R., & McDonald, J. L. (1989). Language
learning: Cues or rules? Journal of Memory and Language, 28, 255-277.
Mandel, D. R., Jusczyk, P. W., & Kemler Nelson, D. G. (1994). Does sentence prosody
help infants to organize and remember speech information? Cognition, 53, 155-180.
Maratsos, M., & Chalkley, M. (1980). The internal language of children's syntax: The
ontogenesis and representation of syntactic categories. In K. Nelson (Ed.), Children's
language: Volume 2, (pp. 127-214). New York: Gardner.
Marcus, G., Ullman, M., Pinker, S., Hollander, M., Rosen, T., & Xu, F. (1992).
Overregularization in language acquisition. Monographs of the Society for Research in
Child Development, 57(4), 1-182.
Markman, E. (1989). Categorization and naming in children: Problems of induction.
Cambrdige, MA: MIT Press.
Marler, P. (1991). Song-learning behavior: the interface with neuroethology. Trends in
Neuroscience, 14, 199-206.
Merriman, W. E., & Stevenson, C. M. (1997). Restricting a familiar name in response to
learning a new one: Evidence for the mutual exclusivity bias in young 2-year-olds. Child
Development, 68, 211-258.
Mervis, C. (1984). Early lexical development: The contributions of mother and child. In C.
Sophian (Ed.), Origins of cognitive skills, (pp. 339-370). Hillsdale, N.J.: Lawrence
Erlbaum.
Mervis, C., & Bertrand, J. (1994). Acquisition of the novel name-nameless category (NC3)
principle. Child Development, 65, 1646-1662.
21
Models of the Emergence of Language
Mervis, C., & Bertrand, J. (1995). Early lexical acquisition and the vocabulary spurt: a
response to Goldfield and Reznick. Journal of Child Language, 22, 461-468.
Miikkulainen, R. (1990). A distributed feature map model of the lexicon, Proceedings of the
12th Annual Conference of the Cognitive Science Society, . Hillsdale, NJ: Lawrence
Erlbaum Associates.
Miikkulainen, R., & Dyer, M. (1991). Natural language processing with modular neural
networks and distributed lexicon. Cognitive Science, 15, 343-399.
Moon, C., Cooper, R. P., & Fifer, W. P. (1993). Two-day infants prefer their native
language. Infant Behavior and Development, 16, 495-500.
Morgan, J., & Travis, L. (1989). Limits on negative information in language input. Journal
of Child Language, 16, 531-552.
Mugdan, J. (1977). Flexionsmorphologie und Psycholinguistik. Tübingen: Gunter Narr.
Naigles, L. G., & Gelman, S. A. (1995). Overextensions in comprehension and production
revisited: Preferential looking in a study of dog, cat, and cow. Journal of Child
Language, 22, 19-46.
Ninio, A., & Snow, C. (1988). Language acquisition through language use: The functional
sources of children's early utterances. In Y. Levy, I. Schlesinger, & M. Braine (Eds.),
Categories and processes in language acquisition, (pp. 11-30). Hillsdale, NJ: Lawrence
Erlbaum.
Piaget, J. (1954). The construction of reality in the child. New York: Basic Books.
Pinker, S. (1984). Language learnability and language development. Cambridge, Mass:
Harvard University Press.
Pinker, S. (1989). Learnability and cognition: the acquisition of argument structure.
Cambridge: MIT Press.
Pinker, S. (1991). Rules of Language. Science, 253, 530-535.
Polka, L., & Werker, J. F. (1994). Developmental changes in perception of non-native
vowel contrasts. Journal of Experimental Psychology: Human Perception and
Performance, 20, 421-435.
Port, R. F., & van Gelder, T. (Eds.). (1995). Mind as motion. Cambridge, MA: MIT Press.
Prasada, S., & Pinker, S. (1993). Generalisation of regular and irregular morphological
patterns. Language and Cognitive Processes, 8, 1-56.
Quine, W. V. O. (1960). Word and object. Cambridge, MA: MIT Press.
Reznick, S. (1990). Visual preference as a test of infant word comprehension. Applied
Psycholinguistics, 11, 145-166.
Rosch, E., & Mervis, C. B. (1975). Family resemblances: Studies in the internal structure of
categories. Cognitive Psychology, 7, 573-605.
Rumelhart, D. E., & McClelland, J. L. (1986). On learning the past tense of English verbs.
In J. L. McClelland & D. E. Rumelhart (Eds.), Parallel distributed processing:
Explorations in the microstructure of cognition, (pp. 216-271). Cambridge: MIT Press.
Rumelhart, D. E., & McClelland, J. L. (1987). Learning the past tenses of English verbs:
Implicit rules or parallel distributed processes? In B. MacWhinney (Ed.), Mechanisms
of Language Acquisition, (pp. 195-248). Hillsdale, N.J.: Lawrence Erlbaum.
Saffran, J., Aslin, R., & Newport, E. (1996). Statistical learning by 8-month-old infants.
Science, 274, 1926-1928.
Samuelson, L. K., & Smith, L. B. (in press). Memory and attention make smart word
learning: An alternative account of Akhtar, Carpenter, and Tomasello. Child
Development, xx, xx.
Savage-Rumbaugh, S., Sevcik, R. A., & Hopkins, W. D. (1988). Symbolic cross-modal
transfer in two species of chimpanzees. Child Development, 59, 617-625.
Schafer, G., & Plunkett, K. (in press). Rapid word learning by 15-month-olds under tightly
controlled conditions. Child Development, xx, xx-xx.
Schmajuk, N., & DiCarlo, J. (1992). Stimulus configuration, classical conditioning, and
hippocampal function. Psychological Review, 99, 268-305.
22
Models of the Emergence of Language
23
Models of the Emergence of Language
Zubin, D. A., & Köpcke, K. M. (1981). Gender: A less than arbitrary grammatical category.
In R. Hendrick, C. Masek, & M. Miller (Eds.), Papers from the Seventeenth Regional
Meeting, (pp. 439-449). Chicago: Chicago Linguistic Society.
Zubin, D. A., & Köpcke, K. M. (1986). Gender and folk taxonomy: The indexical relation
between grammatical and lexical categorization. In C. Craig (Ed.), Noun classes and
categorization, (pp. 139-180). Amsterdam: John Benjamins.
24