UNIVERSITY OF SALFORD
School of Acoustics and Electronic Engineering
Level 3 Module AGA03304 Electroacoustic System Design
Intelligibility
Depends on the following:-
1. The articulation of the speaker
2. The hearing acuity of the listener
3. The acoustics (e.g. RT etc.) of the room
4. The level of background noise
We only have control over 3 and 4.
These reduce to
Direct/Reverberation ratio
Direct/Noise ratio
We have 2 current methods of measuring/predicting
intelligibility which take account of these factors
% ALcons
STI or RASTI
Note that the critical distance Dc has a direct influence on
intelligibility
QWA QA
Dc = →
16πWT 16π (1 + n)
where the symbols are as defined before, and the last part of the
equation is valid only if all sources have equal power W.
%ALcons
This refers to the loss of consonants in speech measured with
single phonetically balanced words placed in short sentences.
Consonants are more critical than vowels.
Calculations are based on octave band values averaged at 1kHz
and 2kHz.
(A) Effect of Reverberation Time T60
Reverberation can cause one syllable to mask the next.
Experimentally it was found that:
For r < 3.5Dc (at which Direct/Reverberation ratio = -11dB)
2
200 r 2 T60 (1 + n)
% ALcons = +K
V QM
where
r = distance from the nearest loudspeaker
T60 = RT (averaged over 1 and 2 kHz)
V = Volume of the room
Q = Directivity of the nearest source
M = Acoustic modifier for reverberant power
= 1 is a conservative assumption
(1+n)= total number of EQUAL sources
K = listener factor ≈ 2% for a good listener.
For r > 3.5Dc (reverberant field, always worse than above)
%ALcons = 9 T60 + K
%ALcons < 10 % intelligibility is very good
%ALcons < 15% intelligibility is acceptable
%ALcons > 15% intelligibility will be a problem
Note that this chart does not agree totally with the above
equation. The equation is generally preferred.
(B) Effect of Background Noise
A poor S/N ratio is clearly bad in itself and it combines with
reverberation to increase %Alcons in the reverberant field. For
S/N ratio < 25dB (1&2kHz) %ALcons is degraded.
%ALcons as a function of S/N ratio for different T (=RT).
Receiver at r > 3.5Dc
Other researches suggested that S/N ratio of > 6dB(A) is
adequate and S/N ratio =15dB(A) is the maximum necessary.
Not directly comparable with 25dB S/N limit in 1 & 2kHz.
Conclusions
Intelligibility is going to be OK if:
1. RT (1 & 2kHz) < 1.5s (≈ 15% ALcons)
2. S/N (1 & 2kHz) ≥ 25dB
Speech Transmission Index (STI)
Primarily a measurement process, can also be estimated from
the room impulse response.
Speech consists of a wide range of frequencies which are
modulated by low frequency envelops to produce the individual
sounds. This modulation must be maintained in order for the
speech to remain intelligible.
The modulation can be degraded by reverberation and
background noise. Note that the modulation is defined in terms
of the intensity or p2.
Illustration of the concept of STI.
Illustration of the effect of reverberation and background noise
on STI. The reduction of the fluctuations in the (octave band
specific) envelop of the output signal (A or B in the diagram)
relative to the original signal can be expressed by the
Modulation Transfer Function (MTF or the m(F) in the
diagram). The two conditions considered (reverberation or noise
interference), lead to the characteristic MTFs, which can be
calculated according to the theoretical expressions at the right
hand side under diffuse field conditions. Note that the MTF is
frequency dependent.
To simulate human speech in the measurement, the simulation
uses shaped noise split into 7 octave bands each of which is
fully modulated by 14 different modulation frequencies F. The
modulation is reduced by the room (reverberation + noise) to
m(F) for each F in each octave band. In total 98 (=7x14) values
of m will be obtained. The m(F) can be measured or calculated.
Notes:
1. If a human speaker is being simulated, then the directivity
of the actual source should be the same as for a human
source.
2. If background noise is a problem then source level and
spectrum must be correct. If only RT is important then
source level and spectral shape do not matter.
Analysis Procedure
1. Each value of m is converted into an apparent S/N ratio
using
m
( S / N ) app = 10 log10 dB
1− m
2. Truncate (S/N)app to +/- 15dB. Nothing outside this range
is significant.
3. A simple average is taken of the 14 (S/N)app values in each
octave band to give 7 (S/N)j where j indexes the octave
bands .
4. A weighted average is taken of the 7 (S/N)j values
7
S / N = ∑ w j (S / N ) j
j =1
Where wj = 0.13, 0.14, 0.11, 0.12, 0.19, 0.17, 0.14
5. STI is then calculated as:
( S / N + 15)
STI =
30
Illustration of STI Analysis Procedure. Shaped areas are those
used by RASTI.
Criteria
STI < 0.4 Poor
0.4 < STI < 0.6 Fair
0.6 < STI < 0.8 Good
0.8 < STI < 1.0 Excellent
Rapid Speech Transmission Index (RASTI)
Introduced to speed up procedure.
Operates with 4 modulation frequencies in the 500Hz band
5 modulation frequencies in the 2kHz band
Analysis is the same except that there are no weighting factors.
So S / N is the straight average of the 9 values.
Relationship between STI and %ALcons obtained over a wide
variety of conditions comprising combinations of various S/N
ratios, reverberation times and echo-delay times. The %ALcons
score refers to the mean loss of consonants in phonetically
balanced monosyllabic (CVC) nonsense words embodied in
neutral carrier phases.
In summary:
%ALcons
Measured directly with suitable word list – needs listening
panel.
Calculated using 1kHz and 2kHz only.
Effect of RT – reliable
Effect of S/N ratio – limited to r > 3.5Dc
Relatively easy.
STI (or RASTI)
Measured using appropriate noise source in-situ
Calculated – combines RT and background noise over full
frequency range.
More difficult.