Phonetic and Phonological Analysis in Linguistics
Phonetic and Phonological Analysis in Linguistics
Peter Ladefoged
Linguistics Department, UCLA, Los Angeles, CA 90095-1543
Abstract
The phonetic description of a language must be related to the phonology. A
computerized description of a language can have a very faithful phonetic
component, but its phonetic structures are not appropriate for a phonological
description. In current systems of linguistic analysis there are three aspects of
phonology: (1) the representation of the lexical contrasts in a language; (2) the
specification of the constraints on the sounds in lexical items; and (3) the
description of phonological patterns of sounds as evident in the relations between
the underlying lexical items and the observable phonetic output. There is a
conflict between the phonetic component required for the first of these goals and
that required for the other two. Characterizing the sounds of languages can be
done most efficiently by using a large number of features, all defined in
articulatory terms. This will result in having more features than are necessary to
characterize phonological patterns efficiently. In addition, some phonological
patterns depend on auditory characteristics which will require auditorily defined
features. Yet other patterns are observable in a language considered as a social
institution rather than a mental concept.
There are many ways in which one can make a description of the sounds of a language, and
linguists often forget about the most obvious one. Figure 1 is an example of part of a description
of the sounds of English, namely their waveforms.
There is no known language that contrasts the 8 distinct gestures involved in producing the
sibilants shown in table 2. If we are considering phonological distinctions, do we need a
classificatory system that formalizes the possibility of all these articulations? The claim here is
that we do need such a system. Table 3 in part validates this claim by showing some of the
contrasts that have been observed using the data in Tables 1 and 2, with the addition of data from
English
It is difficult to know precisely how many of the gaps in Table 3 are due to our not having
found a language using a given contrast as opposed to being due to over-differentiation of
categories. Consider, for example, dental vs. alveolar sibilants. There is no doubt that several
languages have dental sibilants. In addition to Toda, exemplified above, many languages of
California, such as Karok (Bright 1978), have dental [s1]. But in these languages the nearest other
sibilant is post-alveolar [s¢] rather than alveolar [s]. It may well be that dental and alveolar
sibilants are not sufficiently different to provide reliable linguistic contrasts. But from a
classificatory point of view we are not extending the system by allowing for the possibility of
contrasting dental and alveolar sibilants, as we already need the dental–alveolar contrast for other
sounds, such as nasals in Malayalam (Ladefoged and Maddieson 1996).
Table 3. A matrix showing observed contrasts among sibilant fricatives in
different languages.
Alveolar Laminal Apical Laminal laminal Closed Sub-apical
flat Post- post - domed Post- palatalized post retroflex
alveolar Alveolar alveolar post- alveolar
alveolar
s s¢ s2 S Ç s$ ß
Dental s1 xxx Polish Toda Toda Polish Toda
Alveolar s Chinese,
Ubykh
English Chinese,
Ubykh
Ubykh
alveolar
In the second row in Table 3, the contrast between alveolar [s] and post-alveolar [s] cannot
be demonstrated from the sibilant data at hand. But these appear to be distinct gestures and could
well contrast. The same is even more true for the other gap in that row, between alveolar [s] and
sub-apical alveolar [ß]. In fact we can safely assert that all the gaps in the sub-apical column are
due to the rarity of these sounds, rather than to their being a non-contrastive variant of some
other gesture. The same applies to the gaps in the closed post-alveolar column. These are unusual
sounds, but clearly sounds with distinct gestures.
This leaves us with two suspicious gaps to consider. There is no contrast that we know of
between laminal flat post-alveolar [s¢] and apical post-alveolar [s2], nor is there one between
laminal domed post-alveolar (palatoalveolar) [S] and alveolo-palatal sibilants [Ç]. The first of
these gaps does not cause any extension of the classificatory system as we have to have a feature
characterizing the apical — laminal contrast for other sounds. The feature Distributed is usually
assigned to this task. However, within a hierarchical feature system, we should note that it seems
likely that Distributed should have no value in this context.
The lack of a contrast between laminal domed post-alveolar (palatoalveolar) [S] and alveolo-
palatal [Ç]may be caused by the auditory similarity of these sounds. They may be not sufficiently
distinct to be able to sustain a reliable linguistic contrast. Nevertheless speakers of languages
such as Polish and Chinese are keenly aware of the difference between their [Ç] and the similar
but distinguishable [S] sound that occurs in English. They protest loudly when phoneticians
confuse the two. We will take it that this is a case where we have not yet found a language that
uses this contrast rather than ruling out the possibility of a contrast on the grounds that the two
sounds are too similar. However, this may be a wrong judgment. It must be admitted that these
sounds are considered to be distinct because, simply as a matter of opinion, it seems likely that
they could be used as distinctive categories, and with no real scientific basis for this conclusion.
How can the 8 sibilant gestures that we have established as being distinct be described in
terms of phonological features? They are all CORONAL, and [+ strident]. They are also all [–
voice], [+ spread glottis], [– constricted glottis], [– syllabic], [– sonorant, [+ consonantal], [–
ATR], [– lateral], [– nasal], [+ continuant], and [– round]. Considering the features listed above,
this leaves us with High, Low, Back, Distributed, Anterior, and Tense to account for these 8
distinctions. Low and Back as usually defined for vowels and seem non-definable for any of the
sibilant gestures, as does Tense. This leaves us with Anterior, Distributed and High to distinguish
these 8 sounds.
Table 4. The classification of sibilants in terms of the features Anterior, Distributed and High.
Dental Alveolar Laminal Apical Laminal Laminal Closed Sub-
flat post- post - domed palatalized post apical
alveolar alveolar post- post- alveolar retroflex
alveolar alveolar
s1 s s¢ s2 S Ç s$ ß
Anterior + + – – – – – –
Distributed (+) (–) + – + + + –
High – – – – – + – –
Retroflex – – + – – – – +
Closed – – – – – – + –
Table 4 shows how these three features can be used, together with two more that prove
necessay. The usual definition of Anterior is that [+ anterior] sounds are articulated forward of
the alveolar ridge, i.e. equivalent to what we have been calling dental or alveolar, [– anterior]
sounds are post-alveolar. We have already noted that we do not know of a language that
distinguishes dental and alveolar sibilants, but if there were one it would seem likely that one
would be [+ distributed] and the other [– distributed]. The remaining 6 sounds cannot be
distinguished by just the two features Distributed and High. If Distributed is defined such that [+
distributed] is equivalent to what we have been calling laminal, and [– distributed] equivalent to
apical we can fill in the third row in Table 4 as shown. This leaves us with having to distinguish
[ s¢, S, Ç, s$], all of which are [+ distributed], and [ s2] and [ ß], which are both [– distributed]. The
feature High can be used to distinguish [ Ç] from [ s¢, S, s$] by calling [ Ç] [+ high], but this
feature does not really distinguish any of the others. If we are to distinguish all these sounds in
terms of a Universal Grammar we would have to add two new features as shown in the last two
rows of Table 4. Retroflex will separate [ s¢] from [ S, s$], and Closed will separate [ s$] from
[ S ].
Similar arguments can easily be made showing that more features are needed for classifying
vowels. Thinking just of how many vowel heights there are, Danish has four front vowels, each
of which can be long or short. These vowels cannot be said to differ in terms of the features
Tense or ATR as usually defined. As well as the features High and Low, we will have to add
some other feature, such as Mid, which we can define as having first formant frequencies in the
middle of the range. We can then say that the Danish vowels are high, mid high, mid low, low,
much as the IPA does. This would also enable us to characterize Germanic dialects, such as the
Dutch dialect of Weert (see Table 5) that have 5 vowel heights among front vowels. We can
describe these vowels in terms of a single parameter of vowel height as indicated by the terms in
the first column of the table, and in terms of binary features as shown in the last three columns.
Table 5. Words illustrating the front unrounded vowels of the Dutch dialect of Weert (data from
Heijmans and Gussenhoven,1998).
Long Short High Mid Low
High B §i…t Rit + – –
‘far’ ‘Mary’
Mid-high Re…t hItst + + –
‘reed’ ‘heat’
Mid ‘bl”…cE ‘z”gE – + –
‘leaf’ (dim.) ‘to say’
Mid-low tœ…nt slœt – + +
‘tent’ ‘dishcloth’
Low na…t – – +
‘wet’
It is easy to carry this line of argument further and show that we need still more features for a
truly Universal grammar. For example, we need features for the lexical contrasts formed by the
83 clicks that occur in !Xóõ, and for the phonation types in the Austro-Asiatic languages of
South-East Asia, among many other contrasting sounds.
We should also note that any system based on our current linguistic knowledge must be
incomplete for two reasons. Firstly, as we have admitted, some distinctions permitted within the
system are simply estimates of what distinctions are possible within a language. Some future
language may arise that proves these estimates wrong. We may find, for example, a language
that distinguished not only the tense and modal phonations of Bruu, but also murmured and
creaky phonations. Secondly, we cannot account for what is not systematic. Despite de
Saussure’s claim that ‘la langue est une systéme ou toute sa tient” (Sausure 1968), it is not true
that each language forms a system in which everything holds together. There are always odds
and ends that may, or may not, be considered part of the language. Ladefoged and Everett (1996)
point to a number of cases in which a language has an usual sound that occurs in a dozen or
fewer regular lexical items, such as the alveolar released bilabial trill that occurs in Oro Win in
words such as [ tı° um ]‘a small boy’.
Other properties of feature systems
So far the burden of the argument has been that the set of features necessary for describing
lexical contrasts in a Universal Grammar is large, cumbersome and different from the set of
features needed for describing phonological patterns. Describing lexical contrasts necessarily
requires different features than those required for describing phonological patterns. The reverse
is also true. Describing phonological patterns requires different features from those needed for
describing the lexicon. We can make complete descriptions of the lexical contrasts that occur in
the world’s languages by referring to properties of the ways in which the sounds are made. We
do not have to refer to ways in which they are heard. Insofar as features are simply within the
minds of speakers they can be said to have both articulatory and auditory correlates. But when
we are accounting for patterns that occur in a language we need some features that have
articulatory correlates and others that have auditory correlates.
Many years ago Martinet (1955) described the two principal causes of sound changes and the
resulting patterns of sounds: articulatory ease (which produces, for example, assimilation in
plural forms such as cats and ducks with [s], as opposed to dogs and lions with [z]), and auditory
distinctiveness (which produces plural forms such as horses and fishes, in which the two sibilants
are kept separate by an epenthetic vowel). Even in this single, well known, linguistic
phenomenon, we need to refer to the assimilation of an articulatory feature, Voice, and the
separation of sounds with the feature, Strident, an auditory feature that is plainly distinguished by
its acoustic characteristics rather than by its manner of articulation.
Ladefoged (1971, 1992) has described several other features that are characterized by
auditory properties rather than by articulations. These include Sonorant, Rhotic, and the features
that specify vowels, such as High, Low and Back, which, despite popular belief, do not have well
defined articulatory correlates (which is why the correlates of Mid were given in terms of
formant frequencies in the preceding section)
The implications of there being more than one origin of phonological patterns have not been
apparent in classificatory systems. Neither Jakobsonian distinctive features nor Chomsky and
Halle (1968) phonological features took articulatory ease and auditory distinctiveness into
account. They had different aims. Jakobson, Fant and Halle (1951) were interested in developing
a minimal classificatory system rather than one that helped explain the observed patterns.
Chomsky and Halle were interested in explaining observed sound patterns, but they considered
their feature set to have both articulatory and acoustic properties that speakers know about. They
did not consider sound patterns in terms of two distinct sets of features.
To appreciate the difference between the view presented here and that of Chomsky and Halle
we have to consider the nature of language. What is a grammar trying to specify? From the
viewpont being presented here it is not just something in a speaker’s mind. There are around half
million words in the Oxford English Dictionary, and no individual knows all of them. Yet they
are all good lexical items that a grammar must contain. This cannot be considered as simply a
matter of the competence of an ideal speaker as opposed to the performance of an individual. We
cannot claim that an ideal speaker could know all these words. It would not be true. The brain is
not structured that way. Nevertheless there are around half a million words in present day
English. They are all part of the language considered as a social institution.
There are other properties of the language that are not part of a speaker’s competence, but are
part of the language as a social institution. There are many observable patterns of English sounds
that a grammar should describe as part of the language. For example there is a constraint in
English against having two non-coronal stops, oral or nasal, at the end of a word, as in bomb vs,
bombardier, iamb vs. iambic, paradigm vs. paradigmatic etc. The same provision applies at the
beginning of a word, as in mnemonic vs. amnesia, Gnostic vs. agnostic, and even pterygoid vs.
helicopter. The pattern is there, and if we were ever invaded by pterygoid aliens with six wings
we would no doubt coin a new word and regard them as hexapters. The existence of a pattern,
even one that is productive for literate speakers, is not proof that this is part of a speaker's
knowledge of the language, just as the existence of half a million words is not proof that they
could all be part of an ideal speaker’s competence. But if a pattern is there it should be described
in a grammar that regards a language as a social institution
Our current linguistic analyses may be far from reflecting mental processes for many reasons.
In writing grammars we have been strongly influenced not so much by observations of what goes
on in a person’s mind as by observations of the development of the language. Viewed as a social
institution, a language changes because of military or economic conquests, and because social
groups are always striving for a new way of speaking to mark their cohesiveness. The current
form of a language is a reflection of its past history, and any grammar describing the sound
patterns that occur must take this into account.
A language is an intricate social institution, like the national economy. Languages and
economies are best discussed as self organizing systems subject to the pressures of particular
societies. No one would describe the economy in terms of the competence of an ideal user of
money. Similarly we have to think of languages as institutions that are molded by many factors.
We linguists are a little better off than economists. We have identified some of the more
important factors — the influences of other languages, the pressures exerted by laziness or
efficiency resulting in articulatory changes, and the desire to maintain clarity and produce a
distinct auditory message. To these well known influences on languages we must add a third that
we will call organizational economy. This is the counterpart within a self-organizing system of
what Hockett (1955) calls pattern congruity, the tendency of a language to fill in the gaps so that
if it has a set of voice stops /b, d, g/ it is also likely to have a set of voiceless stops /p, t, k/.
Maddieson (1996) refers to similar tendencies as gestural economy, and Clements (2004)
provides good arguments for considering it as simply feature economy.
Failure to recognize the nature of language may explain why the features of a Universal
Grammar as defined by various authorities are not as useful as might be expected in describing
phonological patterens. In a survey of 561 languages Mielke (2004) found that the best known
feature sets frequently fail to define appropriate natural classes that operate in phonological rules.
It seems that on many occasions languages act in individual ways. The different forces acting on
a language are too complex to be encapsulated by the feature sets of a Universal Grammar.
The three constraints on the evolution of the self-organizing sound systems of languages
must be reflected in any feature system that is trying to account for phonological patterns. Some
of the features must be defined with reference to articulatory properties, and some with reference
to auditory properties such as those described by Ladefoged (1971, 1992). A hierarchical
arrangement of the features will reflect constraints imposed to achieve organizational economy.
But this feature system is redundant with regard to specifying the lexicon. As we have noted, the
lexicon can be specified without using auditory features.
Learning and speaking a language
Arguing within the assumptions of the current paradigm, we have shown that language
patterns arise through a variety of causes that cannot be explicated by reference to a single set of
features. We have also shown that the features necessary to specify all the distinct sounds in the
languages of the world would be a large set that is not likely to be a part of a Universal
Grammar, and would be difficult for a child to manage. But at this point we should step outside
the current paradigm and ask why should we imagine that a child learns a language by reference
to an innate set of features? Learning a language involves tuning the phonetic parameters to the
right values, which can be done without any thought of features. Indeed, agents governed by a
computer program can learn the properties of a given vowel system (de Boer 2000), and systems
for learning consonants in terms of articulatory phonology have been described by Goldstein and
Fowler (2003). Children almost certainly learn whole syllables or larger utterances without
reference to features. Toddlers leaving a room know that something has to be said. Some say
["baI"baI] bye bye, others know that the correct thing to say is the single word ["bIraI"bœ] be right
back. This is unlikely to be represented as shown in part in Table 6. We certainly cannot
demonstrate that a feature specification could be used to control the vocal organs. No one has yet
produced a computer model that will generate speech sounds directly from a feature matrix.
Phonological features are best regarded as artifacts that linguists have devised in order to
describe linguistic systems. What linguists are describing is a social institution, which, like the
economy, or the way we dress, is subject to the effects of wars, fashion, laziness and
communicative efficiency. Phonological features are great for describing the patterns that occur
in a language, but learning a language, and the acts of speaking and listening all involve
adjusting articulatory parameters not phonological features. What speakers and listeners do may
be better described in terms of articulatory phonology and direct perception as suggested by
Goldstein and Fowler (2003), rather than by the features that are needed to describe linguistic
patterns. In summary phonologists can make good descriptions of languages using features.
Phoneticians can make good descriptions using parameters such as those of articulatory
phonology or acoustic phonetics. Both phonologists and phoneticians should also be describing
language as a social institution. Language is not only in the mind. There’s much more of it
outside.
References
Association Phonétique Internationale. 1900. Exposé des Principes. Bourg-la-Reine, France.
Bright, W. 1978. Sibilants and naturalness in aboriginal California. Journal of California
Anthropology. Papers in Linguistics, 1, 39-63.
Chomsky, N. 1964. Current Issues in Linguistic Theory. The Hague: Mouton and Co., 1964.
Chomsky, N., & Halle , M. 1968. The Sound Pattern of English. New York: Harper and Row.
Clements, G. N. 2004. Feature economy in sound systems, Phonology. 20.3. 287-333.
Clements, G. and E. Hume. 1995. The Internal Organization of Speech Sounds. In: J. Goldsmith
(ed.), The Handbook of Phonological Theory. Cambridge, Mass.: Blackwell. 245-306.
deBoer, B. 2000. Self-organization in vowel systems. Journal of Phonetics 28: 441-465.
Goldsmith, J.A. 1996. Handbook of Phonological Theory. Blackwells.
Goldstein, L. & Fowler, C. 2003. Articulatory phonology: a phonology for public language use.
In Phonetics and Phonology in Language Comprehension and Production: Differences
and Similarities, edited by Antje S. Meyer and Niels O. Schiller. Mouton de Gruyter.
Hockett, C. F. 1955. A manual of phonology. Baltimore: Waverly Press.
International Phonetic Association. 1904. Aim and Principles. Bourg-la-Reine France.
International Phonetic Association 1999. Handbook of the International Phonetic Association,
Cambridge: Cambridge University Press
Jakobson, R., Fant, G., & Halle, M. 1951. Preliminaries to speech analysis: The distinctive
features and their correlates. Cambridge, MA: MIT.
Ladefoged, P, 1971. Preliminaries to linguistic phonetics, Chicago: Chicago University Press.
Ladefoged, P. 1992. The many interfaces between phonetics and phonology. In W. U.
Dressler,H. C. Luschitzky,O. E. Pfeiffer, & J. R. Rennison (Eds.), Phonologica 1988 , pp.
165-179. Cambridge: Cambridge University Press.
Ladefoged, P., & Everett, D. 1996. The status of phonetic rarities. Language, 72 (4), 794-800.
Ladefoged, P., & Maddieson, I. 1996. Sounds of the World's Languages. Oxford: Blackwells.
Local, J. & Lodge 2004. Some auditory and acoustic observations on the phonetics of [ATR]
harmony in a speaker of a dialect of Kalenjin, Jornal of the International Phonetic
Association, 34.1, 1-16.
Maddieson, I. 1995 Gestural economy. In Proceedings of the 13th International Congress of
Phonetic Sciences Stockholm (Edited K. Elenius and P. Branderud), Volume 4: 574-577.
Martinet, A. 1955. Economie des changements phonétiques. Berne: Francke.
Mielke, J. 2004. The emergence of distinctive features. Ph.D. Dissertation, Ohio State
University.
Saussure, Ferdinand de 1968. Cours de linguistique générale. Paris: Payot.