Sound Source Localization for Mobile Robots
Sound Source Localization for Mobile Robots
Abstracl-Mobile robots in real-life settings would benefit from readings when the sound source is in the same axis of the pair
being able to localize sound sources. Sucb a capability can nicely of microphones.
complement vision to help localize a person or an internsting
event in the environment, and also to provide enhanced pro- Robots are not inherently limited to two microphones; we
cessing for other capabilities such as speech mcognition. In this decided to use more microphones to better approach the
paper we present a robust sound source localization method in localization abilities of the human auditory system. This way,
three-dimensional space using an array of 8 microphones. The increased resolution can be obtained in three-dimensional
method is based on a frequency-domain implementation of a space. This also means increased robustness. since multiple
steered beamformer along with a probabilistic post-processor.
Results show that a mobile robot can localize in real time multiple signals greatly helps reduce the effects of noise (instead of
moving sources of different types over a range of 5 meters with trying to isolate the noise source by putting sensors inside the
a response time of 200 ms. robot's head, as with SIG) and discriminate multiple sound
sources. There are already robots available with more than
1. INTRODUCTION
two microphones; the Sony SDR-4X bas seven.
The sense of hearing is quite important in providing in-
An artificial audition system can be used for three things:
formation in a real life environment: it can draw attention
1) localizing sound sources, 2) separating sound sources in
to particular and discriminate events in the world that can be
order to process only signals ihat are relevant to a particular
further analyzed using other senses such as vision, or it allows
event in the environment, and 3) processing sound sources to
to exchange information through language. For those who do
extract useful information from the environment (like speech
not have hearing impairments, it is hard to imagine going a
recognition for instance). This paper focuses on sound source
day without being able to hear, especially given the fact that
localization. In previous work [ 5 ] , we presented a method
we are moving in many different environments (indoor and
based on time delay of arrival (TDOA) estimation. The method
outdoor).
works for far-field and near-field sound sources and was
Signal processing research that address artificial audition is
validated using a Pioneer 2 mobile robotic platform.
often geared toward specific tasks such as speaker tracking
for videwonferencing. However, artificial hearing for mobile In this paper, we present an approach with the same
robots is still in its infancy. The SAIL robot uses one mi- objective, hut is based on a frequency-domain beamformer
crophone to develop online audio-driven behaviors [I]. The that is steered in all possible directions to detect sources.
robot ROBlTA uses two microphones to follow a conversation Instead of measuring TDOAs and then converting to a position,
between two people [2]. SIG, a humanoid robot uses two pairs the search is performed in a single step. This makes the
of microphones; one pair is installed on both sides of the system more robust, especially in the case where an obstacle
head, while the other pair is placed inside the head to record prevents one or more microphones from properly receiving the
internal sounds (such as motor noise) for noise cancellation signals. The results are then enhanced by probability-based
[31, [4]. Like humans, these last two robots use binaural post-processing which prevents false detection of sources.
localization, i.e. the ability to locate the source of sound in This makes the system sensitive enough for simultaneous
three dimensional space. localization of multiple moving sound sources.
It is difficult to localize sounds with only two input sources. The paper is organized as follows. Section U presents a
The human auditory system accounts for the acoustic shadow brief overview of the system and Section III describes our
of the head and the ridges of the outer ear. Without this ability, frequency-domain implementation of a steered beamformer.
only localization in two dimensions is possible without the Section N explains how we enhance the results from the
possibility to distinguish if the sounds come from the front beamformer using a probabilistic post-processor, followed by
or the back. Also, it may be difficult to obtain high-precision experimental results in Section V.
Authorized licensed use limited to: INDIAN INSTITUTE OF TECHNOLOGY ROORKEE. Downloaded on June 16,2025 at 12:45:14 UTC from IEEE Xplore. Restrictions apply.
11. SYSTEM OVERVIEW Assuming that one sound source is present, we can see that
E will be maximal when the delays T, are such that the micro-
The proposed localization system as shown in Figure 1 is
phone signals are in phase (and therefore add constructively).
..
composed of three parts:
A microphone array;
A memoryless localization algorithm based on a steered
There is, however, a problem with that technique in that en-
ergy peaks are very wide [7],which means that the resolution
is poor. Moreover, in the case of multiple sources, it makes it
beamformer ; more likely to have sources responses overlap.
A probability-based post-processor. One way to narrow the peaks is to whiten the microphone
signals prior to computing the energy [PI. Unfortunately, the
coarse-fine search methods as proposed in [7]cannot be used
because the narrow peaks can be missed during the coarse
search. Therefore. a fine search is necessary, which requires
increased computing power. It is however possible to reduce
the amount of computation by calculating the beamformer
energy in the frequency domain. This also has the advantage
Bsamfamo SaurC
'now pmbrblllocr of making the whitening of the signal easier.
We first notice that the beamformer output energy in Equa-
Fig. 1. Overview of the sysiern tion 2 can be expanded as:
1034
Authorized licensed use limited to: INDIAN INSTITUTE OF TECHNOLOGY ROORKEE. Downloaded on June 16,2025 at 12:45:14 UTC from IEEE Xplore. Restrictions apply.
~
1035
Authorized licensed use limited to: INDIAN INSTITUTE OF TECHNOLOGY ROORKEE. Downloaded on June 16,2025 at 12:45:14 UTC from IEEE Xplore. Restrictions apply.
a post-processing step that provides some smoothing in time, We assume that the transitions between Ho and H1 can
while combining the results of the short- and medium-term be modeled as a first order Markov process with transition
estimators, Using the same quantized sphere as in the previous probabilities a,j = P (H;1 H:-'). This leads to:
section, we associate a probability of source presence to each
region of the grid (we omit the grid region index for clarity).
P(H; Ion-I ) = 001[l - P ( q - 1 lo"-l)]
We note H ; the hypothesis of source presence at discrete
time n and H; the hypothesis of no source being present at + a11P (H;-1lo"-l) (14)
that time. Also, the steered beamformer observation for time For this work, we use a01 = 0.00004, all = 0.992 for
n is denoted on, with 0, = ( 0 1 ; 0 2 , . . . :on) the set of all the short-term estimator and aol = 0.0002, all = 0.96 for
observations up to time n. the medium-term estimator. The reason for the differences in
We first introduce an instantaneous probability estimation values is that the medium-term estimator is updated less often.
that uses the results of the steered beamformer of Section In order to avoid computing P (On), P and P ( 0 , )
ID. The idea is that the higher the output energy of the terms that do not depend on Ho or HI, we introduce the un-
beamformer, the more likely that a source is present. We thus normalized probabilities R ( H y IO,) and R ( H t IO,) that
approximate the instantaneous probability of a source being omit these terms. For example, from Equation 12, we have the
present as: unnormalized probability:
where E is the energy at the output of the beamformer, E,,, From there, it is easy to compute P ( H y IO,) as:
is an energy threshold corresponding to the value when no
source is present, and p,,, is the minimal probability we
want to assign for a source that is detected by the steered
beamformer (with p,,, = 0.1). In the case where there is
no source detected by the beamformer at a certain point. we
assign a floor probability pjloor= 0.005 that accounts for the
possibility that the beamformer does not detect anything even
B. Combination of esrimaror probabilities
though a sound source is present.
After using the temporal integration method to derive the
A. Temporal integration short-term and medium-term estimators, the last step consists
At time N,we use Bayes' rule t o express the probability of combining these probabilities to infer a unique probability
of source presence given all observations as: of source presence. Let O s ;Om respectively be all the obser-
vations made by the short- and medium-term estimators up to
a certain time, we can first write using Bayes' rule:
1036
Authorized licensed use limited to: INDIAN INSTITUTE OF TECHNOLOGY ROORKEE. Downloaded on June 16,2025 at 12:45:14 UTC from IEEE Xplore. Restrictions apply.
~
TABLE I
DETECTION
RATE AS A FUNCTION OF DISTANCE FOR DIFFERENT SOUNDS
Our choice of the geometric mean is based on the fact that
the probabilities can have a very wide dynamic range that is Sound source 3m 5m 7")
not suitable for the arithmetic mean. Hands clapping 92% 94% 84%
Speech (%st'') 1W% 90% 42%
V. RESULTS Noise bunt (250 ms) 1W% lW% 100%
The array used for experimentation is composed of eight
microphones arranged on the summits of a rectangular prism. The second task for which the system is evaluated is speaker
The array is mounted on an ActivMedia Pioneer 2 robot. tracking. In this experiment, several people talk to the robot
as shown in Figure 3. However, due to processor and space simultaneously and in two of the three cases presented, the
limitations (the acquisition is performed using an 8-channel speakers are moving while they talk. In Figure 4, we plot
PCI soundcard that cannot be installed on the robot), the signal the regions where the probability of source presence is at
acquisition and processing is performed on a desktop computer least 0.6. Only azimuth is shown, since the sources are all
(Athlon XP 2000+). The algorithm currently requires 30% located in the same elevation range. From the Figure, it can
CPU to work in real-time, but this amount could be reduced be observed that the system has no difficulty tracking up to 4
by lowering the grid resolution or by using approximations in moving speakers. With I speakers, the system becomes unable
computing the source probabilities. It is worth mentioning that to detect all speakers simultaneously, but nonetheless succeeds
the CPU time does not increase with the number of sources. in localizing them all at over a period of time.
For all results presented in this paper, we used real multi- A third test is performed with two stationary speakers, and
channel recordings in a noisy environment with moderate [Link] robot. Figure 5 shows how the robot localizes the
reverberation. The system is tested under different conditions. speakers as it moves. This demonstrates that the system is able
First, we measure the maximum distance at which the system to function despite the noise caused by its motors. The two
is able to detect different sound sources. During the test, the sources that are sometimes detected at 0' and 90' elevation
sound source is produced 50 times with the robot placed are respectively a computer fan located at 1.5 meter and a
in different positions. The source detection rates (number of ceiling ventilation trap.
detectionslnumber of occurrences) are shown in Table I. We A last experiment was conducted in which we verified
note that the system is able to reliably detect sources at that the system still works when the microphone array is not
distances up to 5 meters. Also, while the system is able to completely open. Even when some sides of the array are filled
detect bursts of white noise reliably at great distance, it is and some microphones no longer have a line of sight with the
1037
Authorized licensed use limited to: INDIAN INSTITUTE OF TECHNOLOGY ROORKEE. Downloaded on June 16,2025 at 12:45:14 UTC from IEEE Xplore. Restrictions apply.
source, the system’s reliability is not significantly affected. Special thanks to Dominic Etoumeau, Serge Caron, Nicolas
BCgin, Mathieu Lemay, Pierre Lepage and Nathan S h a h for
VI. CONCLUSION their help in this work.
Using an array of 8 microphones, we have implemented REF ER EN CE S
a system that is able to reliably localize sounds up to five
[ l ] [Link] and 1. Weng, ‘Grounded auditory development by a develop-
meters away, even in the presence of noise. It is also possible mental robot,” in Proceedings INNWIEEE htemarional Joint Conference
to detect and track simultaneous and moving sound sources. on Neuml Nemo&. 2001, pp. 1059-1064.
Moreover, our system is adapted to both short-duration sounds [2] [Link], T. Tojo, S. Kubota, K. Fwkawa, [Link]. K. Hayata,
[Link], and [Link], “Multiperson conversation via multi-modal
like handclaps and longer duration sounds like speech. interface - a robot who communicate with multi-user,” in Pmceedings
In the proposed system, localization is performed in two EUROSPEECH, 1999, pp. 1723-1726.
[3] K. Nakadai. H. G. Okuno. and H. Kit-, “Real-time sound roum
steps. The first step consists of a beamformer that is steered localization and separation for robot audition,” in Proceedings IEEE
in all possible directions, trying t o maximize output power. lnrernorio~lConference on Spoken Longuoge Processing. 2002. pp. 193-
The second step uses Bayesian probability combinations to 196.
[4] H. G. Okuna, K. Nakadai. and H. Kitano, “Social interaction of humanoid
enhance the steered beamformer results, removing most of the robot based on audio-visual tracking,” in Pmceedings of Eighfcenlh
false detections while maintaining a good detection rate. Intemrionni Conference on Industrid ond Engineering Applicnrions of
In its current form, the localization system is very sensitive An$cid Inreiligrnce ond Expen System. 2002, pp. 725-735.
[5] I.-M. Valin, F. Michaud, I. Rouat. and D. Utourneau, “Robust round
and is sometimes able to detect weak sounds like computer source localization using a microphone array on a mobile robot:’ in
fans located within 2-3 meters. While this may in some cases Proceedings Intemaliowl Conference on lnlelligent Robots and Syslem,
be desirable, it may be desirable in the future to design 2003.
[6] I.-M. Win, 1. Rouat, and F. Miehaud, “Microphone array pmt-filter
an algorithm capable of ranking sound sources in terms of for separation of rimullaneous non-stationary sources,’’ in Submilled to
potential interest to the robot. ICASSP 2004.
[71 R. Duraiswami. D. Zotlon. and L. Davis, “Active speech SOWCC localiza-
ACKNOWLEDGMENT tion by a dual caarse-tDfine search:’ in Proceedings IEEE lntemOIiOMl
Conference on Acoustic& Speech, and Signal Pmccssing, 2001.
FranGois Michaud holds the Canada Research Chair (CRC) [SI M. Ornologo-and P. Svaizer, ‘Xcoustic event localization using a
crosspower-specmm phase based technique:’ in Proceedings IEEE Infer-
in Mobile Robotics and Autonomous Intelligent Systems. This notional Conference on Aeoudcs, Speech. ond Signal Processing. 1994.
research is supported financially by the CRC Program, the pp. 11-273-n-276.
Natural Sciences and Engineering Research Council of Canada [9] F. Giraldo, ‘Zagrange-galerkin methods on spherical geodesic grids,”
Journal of Computotionoi Phy$io, vol. 136, pp. 197-213, 1997.
(NSERC) and the Canadian Foundation for Innovation (CFI).
1038
Authorized licensed use limited to: INDIAN INSTITUTE OF TECHNOLOGY ROORKEE. Downloaded on June 16,2025 at 12:45:14 UTC from IEEE Xplore. Restrictions apply.