0% found this document useful (0 votes)
108 views9 pages

Acoustic Phonetics Instrumentation Overview

The document discusses the history and development of instrumentation used in acoustic phonetics. It describes early instruments like kymographs and palatographs that recorded vocal cord vibrations and places of articulation. The development of the sound spectrograph in the 1940s was pivotal, allowing visualization of the acoustic properties of speech. This enabled more detailed analysis of vowels and consonants. Later, technologies like electropalatography and computer programs improved analysis capabilities and user experience. Overall, the document traces the progression from early mechanical instruments to modern digital tools that have advanced the scientific study of speech sounds.

Uploaded by

Chaitanya Kirti
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
108 views9 pages

Acoustic Phonetics Instrumentation Overview

The document discusses the history and development of instrumentation used in acoustic phonetics. It describes early instruments like kymographs and palatographs that recorded vocal cord vibrations and places of articulation. The development of the sound spectrograph in the 1940s was pivotal, allowing visualization of the acoustic properties of speech. This enabled more detailed analysis of vowels and consonants. Later, technologies like electropalatography and computer programs improved analysis capabilities and user experience. Overall, the document traces the progression from early mechanical instruments to modern digital tools that have advanced the scientific study of speech sounds.

Uploaded by

Chaitanya Kirti
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Linguistics

Instrumentation and Technology in Acoustic Phonetics

Principal Investigator Prof. Pramod Pandey


Centre for Linguistics, SLL&CS, Jawaharlal Nehru University, New
Delhi-110067
Email id: pkspandey@[Link]
011-26704226; 011-26741258, 9810979446

Paper Tiltle Introduction to Phonetics and Phonology

Paper Coordinator Prof. Pramod Pandey

Module ID & Name Lings_P2_M16: Instrumentation and Technology in Acoustic


Phonetics
Content Writer Vaishna Narang
Centre for Linguistics, SLL&CS, Jawaharlal Nehru University,
New Delhi- 110067
Email id: [Link]@[Link]
+91-11-26704664, 91-9810608936

Content Reviewer Hemalatha Nagarajan

MODULE 16
Instrumentation and Technology in Acoustic Phonetics
The present unit includes a brief history of acoustic phonetics and instrumentation for the
study of speech sounds which shows how various methods and instruments were used to
understand human speech in the past especially in the last three decades. One cannot ignore
the ancient classical tradition of the 3rd/4th century BC when scholars wrote very
comprehensive accounts of Greek, Arabic and Sanskrit grammars. But the developments in
modern electro- acoustic phonetics can perhaps be traced back to 18th century “Talking
Machines” Lat mettrie (1747), Kratzenstein (1780), Von Kempelen (1791), “Source Filter
Theory of Speech Production” (Muller, 1894). Later version of Talking Machines appeared as
“Pattern play back” in the 20th century. The source-filter theory was further developed by
phoneticians in the 19th and the 20th centuries, Hermann (1894), Fant (1960), Lieberman &
Blumstein (1988). The Source-Filter Theory developed by Fant (1960), with the writing of
his book entitled “Acoustic Theory of Speech Production” is now commonly known as the
“Linear Source-Filter Theory of Speech Production”. According to this theory (Fant 1960),
speech sounds are produced with three major mechanisms i.e. the sound generating source
(vocal cords), filtering feature of the vocal tract, and radiation characteristics into the external
environment.

Coming back to the instrumental phonetics beginning with tuning forks and resonators,
kymographs and palatographs, as Malmberg (1963) says “Fifty years ago, (which would
mean early 20th Century since Malmberg’s publication is dated 1963) in the field of acoustics,
phoneticians had at their disposal only very modest resources: tuning forks and resonators
mechanical recordings of vibrations, which were analysed according to Fourier’s theorem.
Despite this imperfection of the instruments, an astonishingly exact knowledge of the
structure of vowels had already been arrived at towards the end of the last century (i.e. the
19th century) thanks to the genius of a few great physicists and phoneticians (Helmholtz,
Hermann, Rousselot, Pipping). It is, however, only with modern electro-acoustics that we
have succeeded in going farther than the phoneticians of the last century. Thanks to the
microphone, the cathode oscillograph, filters, and the different acoustic spectrometers, to
“visible speech,” if and to synthetic speech, there is no longer anything, from a technical
point of view, to prevent a complete and integral analysis of all details of the sounds used in
human language.”(p.87)

The so called physiological instruments the kymograph and the palatograph were the earliest
to develop for recording oral, nasal sound output and vocal cord vibrations using kymograph
and places of articulation using an artificial palate in palatography.

Fig. 16.1 – 16.3 - Kymographs


References for images: Fig 16.1- [Link]
[Link]/sites/default/files/styles/full_size/public/chang_heinitz3.jpeg?itok=VYVyKn5
Z Fig 16.2- [Link] Fig 16.3-
[Link]

With the help of the kymograph cylinder, it is possible to register oral and nasal speech
output on blackened, soot covered paper covering the cylinder, and thus obtain the waveform
on the cylinder which rotates at a controlled speed. With the help of a mouth piece, Marey’s
capsule and a rubber membrane it is also possible to capture the oral air stream, the nasal air
flow as well as the vocal cord vibrations. With the rotation of the drum at a controlled speed
it is thus possible to establish the difference between vowels and consonants, stop closure
and stop plosion, and aspiration, voiced and voiceless productions, nasal and non-nasal
productions. The tracings obtained from the kymograph, i.e. the kymograms thus show
tracings of the sound waves as produced through the oral, nasal air stream and vocal cord
vibrations.

Palatography: then and now


The artificial palate, custom made to fit an individual palate was made, smeared with edible
oil and with refined flour sprinkled, which when inserted in the mouth of the speaker, would
carry all the marks of the contact between the articulator and the place of articulation.
Palatograms thus produced with specific segments could indicate the places of articulations
for different speech sounds. See some palatograms in figures 16.4 and 16.5 below.
Fig. 16.4: Palatograms /s/, /ʃ/, /c/ in Japanese and English speakers’ articulations
[Link]
:301844090900486@1448976647726/Figure-1-Palatograms-from-Edwards-1-showing-the-
[Link]

Fig. 16.5 Palatograms of /t/, /ʈ/, /n/, /s/, /š/, /pi/, /pu/, /pɨ/ in a Hindi speaker’s
articulations (in order) (Narang, 1972)

Fig. 16.6 (a): The articulate instrument for electropalatography.


[Link]
[Link]
Fig. 16.6 (b): A diagram demarcating the areas in a palatograph. (Narang, 1972)

Those time consuming experiments on places of articulation using artificial palates soon gave
way to electro-palatography in the 70’s and the 80’s which could give us the articulator
movements and points of contact in time, frame by frame for a much more detailed and
precise descriptions of places of articulations. The electropalatography as a tool in linguistic/
phonetic studies later found applications in clinical studies and practices, supporting speech
therapists with diagnostic as well as therapeutic tools and practices. See Fig16.6 for
electropalatograms of alveolar stops – t, d, n, alveolar, apical and the last one in fig 16.7 is
the EPG plate. (2013)
Fig. 16.7: Electropalatography Plate. [Link]
[Link]

Sound spectrograph

As stated by Gunner Fant in his 2004 overview “Phonetics and Phonology in the last 50
years” of major importance for this field was “the development of the sound spectrograph at
Bell Laboratories (Potter et al.,1947). An early and initiated study of acoustic phonetics was
that of Martin Joos (1948). My early work at Ericsson was reported in Fant (1948, 1949) and
later in Fant (1959). A forerunner of vocal tract theory was that of Chiba and Kajiyama
(1941). My first contact with the field of speech and hearing was through Fletcher (1928)”.

The sound was used extensively for experiments on visible speech. A schematic diagram of
the filters as used in Visible Speech as in Malmberg, (1963: 89) is reproduced below.

Fig. 16.8: Filters of sound spectrograph, a schematic representation from Malmberg


(1963: 89)
Unlike a spectral representation which gives frequency on X-axis and amplitude on Y-axis,
which shows resonant frequencies/ formants as spectral peaks, a spectrogram gives all the
three dimensions, time on X-axis, frequency on y-axis and intensity/ amplitude as depth,
lightness/ darkness of the lines, as the third dimension. Fant (1962) “Sound Spectrograph”
gives the details of the instrument as well as the studies conducted on the machine in the
early 40’s, 50’s. By convention one would use a narrow band filter, say around 45 Hz for
getting the f0 and the overtones/ harmonics and one would choose a wide band spectrogram
with a filter of say 300 or 350 Hz to see the formant patterns (see details of spectrograph in
the study of vowels in module 14, and some details of spectrographs for the study of
consonants in module 15). Once we had a spectrograph to show all the three dimensions of
all the constituent frequencies of speech sounds, there was no looking back as far as the
acoustics of resonant speech sounds was concerned. The instruments and the models were
improving with time and with extensive usage (Cepstral Analysis and LPC i.e. Linear
Predictive Code of speech and formant tracking etc in the 60’s). In the 70’s and the 80’s a
number of such tools and technologies developed as softwares and programmes which were
much more efficient and user friendly such as PRAAT, Wavesurfer, Gold wave, audacity to
mention a few. Details of these tools and instruments are given in the following sections (2.1
to 2. 7)

For a very broad and general classification of speech sounds as consonants one could use
spectrograms which would show vowels with formant patterns as distinct from consonants,
stops and plosives distinct from say fricatives and affricates, with or without voicing
indicated by light striations at the bottom / low frequency range, fricatives identified as high
frequency, high intensity noise patterns etc but a major breakthrough in the identification of
consonants was made only after the scientists at Haskins laboratories used spectrography
along with ‘Pattern Play Back” invented in the fifties.

Fig 16.9: Pattern Play Back. [Link]

Pattern Play Back

In 1950, the advances in vocal tract modelling and speech synthesis (Dunn 1950, Lawrence
1953, Fant 1956 & 1959) and a large number of experiments done at the Haskins
Laboratories (Cooper et al. 1951) using synthesis from hand painted spectrograms provided
the technology for carrying out many types of investigation in speech perception.

The pattern playback used in speech synthesis experts at Haskin laboratory employs a
variable-density tone wheel to modulate the light from a mercury arc, producing a
fundamental of 120 cps and its entire harmonics through the fiftieth at 6000 cps. The
modulated light beams are imaged on the spectrogram, and are so spread across it as to match
its frequency scale. As the hand-painted spectrogram is moved through the light, the white
paint reflects beams whose modulation frequencies correspond to the position of the paints on
the frequency scale of the spectrogram. The reflected beams are led by plastic light guides to
a phototube, the current of which is amplified and converted to sound.

As explained in the earlier modules (14 & 15) the huge range of experiments conducted with
synthetic speech produced by hand painted spectrograms run through pattern play back and
recordings used for a number of perception experiments led to a major breakthrough which
is the three acoustic cues for the identification of consonants viz. formant transitions,
placement of a burst of noise along the formants of the following vowels, and locus
hypothesis. See details of the study of consonants and the study of vowels and consonants in
continuous speech in modules 14 & 15. (Cooper et al, 1952, Delattre et al. 1955, Lieberman,
1956)

2. TOOLS and SOFTWARES: Modern technologies for the study of speech sounds:

As indicated in the previous sections all the tools, techniques and models developed for
speech acoustics received a great boost with the computer aided instrumentation in the 80’s
and the 90’s. In other words when the earlier research on the study of speech sounds came
together with the computer technology the result was as expected, unprecedented pace of
developments in the science of acoustics in general and speech acoustics in particular. A brief
introduction to some of these tools and softwares is given below in sections 2.1.1 to 2.1.6.

2.1 Tools and Software for Acoustics of Speech

The analysis of vowel sounds in terms of their constituent frequencies has been made
convenient through various interfaces and softwares such as PRAAT, Wavesurfer.
Goldwave, Goldsurfer, Computerized Speech Lab (CSL) to name a few. PRAAT is
software for analysis and synthesis of speech which was written by Paul Boersma of the
Department of Phonetics of The University of Amsterdam. The program is constantly being
improved and a new build is published almost every week. Version 4.2 was published in
March 2004 and the last build was 4.2.34. Version 4.3 was introduced in February 2005 and
the current builds (22 February 2005) are 4.3.00 for Solaris, 4.3.01 for Linux and 4.3.02 for
Windows and Macintosh. PRAAT has a variety of features which include the facility to
analyze speech in terms of its waveform, fundamental frequency, constituent formants, and
spectrographic analysis.

2.1.1 Computerized Speech Lab (CSL):

KAY Elemetrix earlier known as KAY Pentax as a research lab have produced a number of
instruments and programmes for application in the field of acoustic phonetics, forensic
sciences, speech pathology and of course in the field of descriptive linguistic studies. CSL,
Multispeech, MDVP and a number of other programmes are dedicated to speech and voice
analysis.

Computerized Speech Lab (CSL) is a system with requisite software as well as hardware
used for speech and signal processing. In CSL, speech samples can be recorded directly and
then further analyzed. For the analysis of the vowel sounds as they occur in continue speech,
the vowel phonemes can be separated from the adjoining consonants sounds by trimming
them. This is possible through the ‘Trim Data Waveform’ option that is available in the CSL.
Through this option we are also able to erase the entire unnecessary noise disturbance that is
captured from the surroundings during the course of the recording, voice patterns, spectrum,
spectrogram etc. waveform analysis can then proceed on this data, using a number of
software options.

2.1.2 GoldWave:

GoldWave is digital audio editing software with various graphic visuals including
spectrogram and spectrum. And Wavesurfer has a simple and logical user interface that
provides functionality in an intuitive way and which can be adapted to different tasks. It can
be used as a stand-alone tool for a wide range o tasks in speech research and education.
Typical applications are speech/sound analysis and sound annotation/transcription.
Wavesurfer can also serve as a platform for more advanced/specialized applications. This is
accomplished either through extending the wavesurfer application with new custom plugins
or by embedding wave-surfer visualization components in other applications.

Fig 16.10: C2: T1 GoldWave Window

Once a sound file is transferred to the machine and given a specific code, the individual
words needed can be selected and converted to individual files using Goldwave. The voice
files can also be converted to the MP3 format. Sometimes due to unavoidable background
noises the word recordings have to be repeated.

Goldwave can be used for noise reduction sound samples recorded. Using PRAAT, each
sound file can, the, be analysed one by one.

2.1.3 Sensimetrics Speech Station (SSS):


Sensimetrics Speech Station (SSS), CSRE (Computer Speech Research Environment),
SPEECHLAB and SPECTO has been discussed in detail by Pickett,1999. Sensimetrics
Speech Station which is Window 95 compatible and it used for standard spectrographic
analysis and also for reading frequency values and time duration for sounds. The Computer
Speech Research Environment developed by the Dept. of Communicative Discorders,
University of Western Ontario and the Signalyse System from Macintosh computers also
help perform Fourier analysis and standard spectrographic study. SPEECHLAB and
SPECTO are similar tools for speech analysis (Pickett, 1999:344).

2.1.4 PRAAT

Praat is a computer program which allows phoneticians (or anyone else) to analyse,
synthesise, and manipulate speech. It is quite an unintuitive program to use, especially for
people who are only familiar with Windows. Praat was created by Paul Boersma and David
Weenink from the Institute of Phonetic Sciences at the University of Amsterdam - initially
for their own use. Over time the user base has grown: at first, other people in the UoA
Phonetic Sciences department, later, people at other institutions. Today, it is the most
powerful and (pretty much) the only program that allows you to do phonetics by computer.

Praat was not designed for Windows - it was ported. For this reason, it does not perform (or
look) like a typical Windows program. There is no drag and drop, control+A does not select
all, and pressing tab will not allow you to navigate around the window.

Praat can be downloaded for free from the Praat homepage. You only download one file:
[Link] - this is the Praat program. Place the file in a folder somewhere on your hard drive,
and create a shortcut to it on your desktop and/or in your start menu. New versions of Praat
are available for download every few weeks. Updating is as simple as overwriting your old
[Link] with the newly downloaded one - which is kind of nice.

2.1.5 Wavesurfer

Wavesurfer is a sound analysis tool that can be used to visualize and manipulate sound files.
It is freeware that can be modified and redistributed as wanted. The documentation provided
and suggests that the program could be modified easily. Specific information can be found
at the Wavesurfer website at [Link]

2.2.6 Audacity:

Audacity is an easy to use but powerful audio recording and editing package. It is also free to
download and use (see Appendix A – Downloading and Installing Audacity and the LAME
MP3 encoder at the end of this document for information on how to do this). Audacity
enables you to record your voice, edit your recording to correct any mistakes you might
make, and to combine sound recordings from various sources such as interviews, music, or
other sound recordings you may have. Audacity also enables you to export your recording as
an MP3 file, and because of this it is ideal for producing podcasts. This guide will serve as an
introduction to the key features of Audacity 1.2.6 for Windows. A full reference guide for the
software can be found at [Link]

 Audacity is a free, easy-to-use audio editor and recorder for Windows, Mac OS X,
GNU/Linux, and other operating systems. Audacity was created by Dominic Mazzoni of
Google, while he was a graduate student at Carnegie-Mellon University. Dominc Mazzoni
is still the main developer and maintainer of Audacity, with help from many others around
the world. Audacity is extremely popular in the podcasting world due to its wide
availability, multiplatform support, and the fact that it is free. Some of Audacity's features
include:

* Importing and exporting WAV, MP3 (via the LAME MP3 Encoder, downloaded
separately), Ogg Vorbis, and other file formats
* Recording and playing sounds
* Editing via Cut, Copy, Paste (with unlimited Undo)
* Multi-track mixing
* Digital effects and effect plug-ins. Additional effects can be written with Nyquist
* Amplitude envelope editing
* Noise removal
* Support for multichannel modes with sampling rates up to 96 kHz with 24 bits per
sample
* The ability to make precise adjustments to the audio's speed, while maintaining pitch, in
order to synchronise it with video, run for the right length of time, etc.
* Unlike many other programs, Audacity has very few audible artifacts (doubling, or
chorus type effects) when lengthening or shortening a file.
* High ease of use
* Large array of plug-ins available In this free online Audacity Course you will learn to:
* Record live audio.
* Convert tapes and records into digital recordings or CDs.
* Edit Ogg Vorbis, MP3, and WAV sound files.
* Cut, copy, splice, and mix sounds together.
* Change the speed or pitch of a recording

2.1.7 Irregular Polygon Area Calculator

The Irregular Polygon Area Calculator can be used to calculate the area of the acoustic space,
for comparison across groups under study. The correlation between the formantscan be
studied to see if these patterns change in cases of different type of samples. Data can then be
tabulated for different subjects groups and correlation between different parameters can be
studied.

You might also like