Understanding Spectrograms in Speech
Understanding Spectrograms in Speech
A spectrogram distinguishes voiced from voiceless fricatives by capturing the dual characteristics of vibrations and turbulence. Voiced fricatives show aspects of regular vocal fold vibrations in addition to the chaotic, random frequencies due to a turbulent airstream. This combination results in visible periodic vocal striations or bands alongside noise patterns, differentiating them from purely voiceless fricatives, which only display randomness resembling static noise without such striations .
Identifying formants in back vowels like [ɑ] can be challenging because F1 and F2 often occur so close together that they appear as a single broad formant band on a spectrogram. This close positioning causes difficulty in distinguishing individual formants and can lead to misidentification unless there is pre-existing knowledge about the expected formant location. These challenges highlight the need for an analytical understanding of spectrogram interpretation aided by contextual linguistic knowledge .
Amplitude, reflecting the intensity of sound frequencies, plays a critical role in the analysis of speech sounds using spectrograms. Visually, amplitude is indicated by the brightness or color intensity of points on the spectrogram, with higher amplitude sounds appearing brighter or more colorful. This allows for the identification of sound energy levels over time, aiding in distinguishing between voiced consonants, vowels, and noise-like sounds such as fricatives based on their relative intensity .
The release burst is a critical feature in identifying plosives on a spectrogram. It appears as a very thin fricative and is crucial for differentiating plosives because it marks the precise moment when the built-up air pressure behind the closure is released. This feature helps listeners understand the identity of the plosive even if the voicing characteristics are masked by noise or equipment constraints . The presence of formant transitions, which look like the distortion of formants away from their stable frequencies, further assists in recognizing plosives when visible .
Vowels on a spectrogram exhibit clearly defined formant bars, which represent the resonant frequencies of the vocal tract. These formants are indicative of the specific vowel sound being produced, with diphthongs showing a change in formant frequencies as the tongue moves . In contrast, fricatives appear as a chaotic mix of random frequencies resembling static noise due to their turbulent airstream. High-frequency fricatives like [s] have a higher average frequency compared to [ʃ], [f], or [θ]. Thus, the presence of structured formants helps in identifying vowels, while random noise patterns assist in identifying fricatives.
Nasals and lateral approximants present identification challenges on a spectrogram because they display acoustic properties that differ markedly from vowels. While vowels exhibit strong, clearly defined formant bands, nasals and laterals appear as faint vowels with low amplitude at higher frequencies . This reduced amplitude is due to the complex resonant properties involving tubes with branches and side-chambers, resulting in both formants and anti-formants. These anti-formants create less pronounced formant bands, making it difficult to distinguish specific nasals or laterals solely from a spectrogram .
Formant transitions on a spectrogram are indicative of the onset or release phases of consonants such as [t] and [k]. At the end of a vowel, formant transitions exhibit changes wherein formant frequencies are distorted away from their steady-state positions as they move towards or away from the frequencies associated with the consonant articulation. For example, towards the end of the vowel [æ], the F2 and F3 formants move towards each other, indicating the onset of a velar consonant, a pattern often referred to as the "velar pinch" . This transition aids in identifying the place of articulation for the following or preceding consonant.
Aspiration and release bursts for voiceless plosives differ significantly in duration and representation on a spectrogram. The release burst is a very brief occurrence that resembles a thin fricative and indicates the release of occluded air pressure . Aspiration, however, lasts longer than the release burst and appears as noise resembling an [h], positioned between the silent phase and vowel onset on the spectrogram. It represents the delay in voicing onset, characterized by an absence of vocal fold vibration despite the vocal tract being in position for the next vowel .
The medial phase of a voiceless plosive is represented on a spectrogram as a region of complete silence, which appears as a white blank space. This phase signifies the moment of occlusion where no sound energy is released before the plosive's release burst occurs . Following this silent gap, the release burst and potential aspiration mark the plosive's continuation, helping to visually segment the sound on the spectrogram.
The distribution of static noise on a spectrogram differs between fricatives like [s] and [f] based on their average frequency characteristics. [s] exhibits a higher average frequency compared to [f], resulting in a more concentrated area of bright noise at higher frequencies on the spectrogram. In contrast, [f] displays noise more prominently at lower frequencies, with a less concentrated high-frequency band . This variation in noise distribution aids in distinguishing specific fricative sounds through their unique frequency signatures.