Valin Localization
Valin Localization
Particle Filtering
Abstract
Mobile robots in real-life settings would benefit from being able to localize and track sound
sources. Such a capability can help localizing a person or an interesting event in the en-
vironment, and also provides enhanced processing for other capabilities such as speech
recognition. To give this capability to a robot, the challenge is not only to localize simulta-
neous sound sources, but to track them over time. In this paper we propose a robust sound
source localization and tracking method using an array of eight microphones. The method
ticle filter-based tracking algorithm. Results show that a mobile robot can localize and track
in real-time multiple moving sources of different types over a range of 7 meters. These new
capabilities allow a mobile robot to interact using more natural means with people in real
life settings.
sound sources in relation to a point in space. The auditory system of living creatures
provides vast amounts of information about the world, such as localization of sound
sources. For us humans, it means to be able to focus our attention on events and
person who is talking to us, etc. Hearing complements well other sensors such as
vision by being omni-directional, capable to work in the dark and not limited by
physical structure (such as walls). For those who do not have hearing impairments,
it is hard to imagine going a day without being able to hear, especially having to
move in a very dynamic and unpredictable world. Marschark [1] even suggests that
although deaf children have similar IQ results compared to other children, they
capabilities.
To perform sound localization, our brain combines timing (more specifically de-
lay or phase) and amplitude information from the sound perceived by two ears [2],
sources using only two inputs is a challenging task. The human auditory system is
very complex and resolves the problem by accounting for the acoustic diffraction
Email address:
{[Link],[Link],[Link]}@[Link]
2
around the head and the ridges of the outer ear. Without this ability, localization
with two microphones is limited to azimuth only, along with the impossibility to
distinguish if the sounds come from the front or the back. Also, obtaining high-
precision readings when the sound source is in the same axis as the pair of micro-
One advantage with robots is that they do not have to inherit the same limitations
as living creatures. Using more than two microphones allows reliable and accurate
localization in both azimuth and elevation. Also, having multiple signals provides
additional redundancy, reducing the uncertainty caused by the noise and non-ideal
sible directions to detect sources. Instead of measuring TDOAs and then converting
to a position, the localization process is performed in a single step. This makes the
system more robust, especially in the case where an obstacle prevents one or more
microphones from properly receiving the signals. The results of the localization
false detection of sources. This makes the system sensitive enough for simultane-
of earlier work [3] and works for both far-field and near-field sound sources. De-
tection reliability, accuracy, and tracking capabilities of the approach are validated
using a mobile robot, with different types of sound sources. We consider both our
3
The paper is organized as follows. Section 2 situates our work in relation to other
research projects in the field. Section 3 presents a brief overview of the system.
Section 5 explains how we enhance the results from the beamformer using a proba-
ing how the system behaves under different conditions. Section 7 concludes the
2 Related work
Signal processing research that addresses artificial audition is often geared toward
specific tasks such as speaker tracking for videoconferencing [4]. An artificial au-
dition system for a mobile robot can be used for three purposes: 1) localizing sound
sources; 2) separating sound sources in order to process only signals that are rel-
Even though artificial audition on mobile robots is a research area still in its in-
fancy, most of the work has been done in relation to localization of sound sources
and mostly using only two microphones. This is the case of the SIG robot that uses
both inter-aural phase difference (IPD) and inter-aural intensity difference (IID) to
locate sounds [5]. The binaural approach has limitations when it comes to evalu-
ating elevation and usually, the front-back ambiguity cannot be resolved without
More recently, approaches using more than two microphones have been developed.
One approach uses a circular array of eight microphones to locate sound sources
4
[7]. In our previous work also using eight microphones [8], we presented a method
for localizing a single sound source where time delay of arrival (TDOA) estimation
was separated from the direction of arrival (DOA) estimation. It was found that
a system combining TDOA and DOA estimation in a single step improves the
sources [3]. Kagami et al. [9] reports a system using 128 microphones for 2D sound
Most of the work so far on localization of source sources does not address the
for tracking a moving source. However the proposed method assumes that a single
source is present. In the past years, particle filtering [11] (a sequential Monte Carlo
method) has been increasingly popular to resolve object tracking problems. Ward et
al. [12,13] and Vermaak [14] use this technique for tracking single sound sources.
Asoh et al. [15] even suggested to use this technique for mixing audio and video
data to track speakers. But again, the technique is limited to a single source due to
the problem of associating the localization observation data to each of the sources
lem. Some attempts are made at defining multi-modal particle filters in [16], and the
But so far, the technique has not been applied to sound source tracking. Our work
demonstrates that it is possible to track multiple sound sources using particle filters
5
3 System Overview
of three parts:
• A microphone array;
robot. Since the system is designed to be installed on any robot, there is no strict
constraint on the position of the microphones: only their positions must be known in
relation to each other (measured with ∼0.5 cm accuracy). The microphone signals
are used by a beamformer (spatial filter) that is steered in all possible directions
in order to maximize the output energy. The initial localization performed by the
steered beamformer is then used as the input of a post-processing stage that uses
particle filtering to simultaneously track all sources and prevent false detections.
The output of the localization system can be used to direct the robot attention to
the source. It can also be used as part of a source separation algorithm to isolate the
Omni-directional
microphones
. Steered Particle
. beamformer filtering
.
Beamformer Source
energy positions
6
4 Localization Using a Steered Beamformer
The basic idea behind the steered beamformer approach to source localization is
to direct a beamformer in all possible directions and look for maximal output.
the microphone signals and Section 4.3 shows how the search is performed. A
M −1
(1)
X
y(n) = xm (n − τm )
m=0
where xm (n) is the signal from the mth microphone and τm is the delay of arrival
for that microphone. The output energy of the beamformer over a frame of length
L−1
[y(n)]2
X
E=
n=0
L−1
[x0 (n − τ0 ) + . . . + xM −1 (n − τM −1 )]2 (2)
X
=
n=0
Assuming that only one sound source is present, we can see that E will be maximal
when the delays τm are such that the microphone signals are in phase, and therefore
add constructively.
7
One problem with this technique is that energy peaks are very wide [20], which
means that the resolution is poor. Moreover, in the case where multiple sources
are present, it is likely for the two or more energy peaks to overlap, making them
impossible to differentiate. One way to narrow the peaks is to whiten the micro-
phone signals prior to computing the energy [21]. Unfortunately, the coarse-fine
search method as proposed in [20] cannot be used in that case because the narrow
peaks can then be missed during the coarse search. Therefore, a full fine search
domain. This also has the advantage of making the whitening of the signal easier.
M −1 L−1
x2m (n − τm )
X X
E=
m=0 n=0
M −1 mX1 −1 L−1
(3)
X X
+2 xm1 (n − τm1 ) xm2 (n − τm2 )
m1 =0 m2 =0 n=0
M −1 mX
1 −1
(4)
X
E =K +2 Rxm1 ,xm2 (τm1 − τm2 )
m1 =0 m2 =0
delays and can thus be ignored when maximizing E. The cross-correlation function
L−1
Xi (k)Xj (k)∗ e2πkτ /L (5)
X
Rij (τ ) ≈
k=0
where Xi (k) is the discrete Fourier transform of xi [n], Xi (k)Xj (k)∗ is the cross-
spectrum of xi [n] and xj [n] and (·)∗ denotes the complex conjugate. The power
8
spectra and cross-power spectra are computed on overlapping windows (50% over-
by averaging the cross-power spectra Xi (k)Xj (k)∗ over a time period of 4 frames
(40 ms). Once the Rij (τ ) are pre-computed, it is possible to compute E using only
follows that the complexity of the search itself is reduced from 1.2 Gflops to only
only 48.4 Mflops, 25 times less than a time domain search with the same resolu-
tion.
L−1
(w) Xi (k)Xj (k)∗ 2πkτ /L
(6)
X
Rij (τ ) ≈ e
k=0 |Xi (k)| |Xj (k)|
tion has one drawback: each frequency bin of the spectrum contributes the same
amount to the final correlation, even if the signal at that frequency is dominated by
noise. This makes the system less robust to noise, while making detection of voice
(which has a narrow bandwidth) more difficult. In order to alleviate the problem,
ξin (k)
ζin (k) = (7)
ξin (k) + 1
9
where ξin (k) is an estimate of the a priori SNR at the ith microphone, at time frame
where αd = 0.1 is the adaptation rate and σi2 (k) is the noise estimate for micro-
phone i. It is easy to estimate σi2 (k) using the Minima-Controlled Recursive Aver-
age (MCRA) technique [23], which adapts the noise estimate during periods of low
energy.
It is also possible to make the system more robust to reverberation by modifying the
weighting function in Equation 8 to use a new noise estimate σ̃i2 (k) that includes a
2
λrev rev n n−1
n,i (k) = γλn−1,i (k) + (1 − γ)δ ζi (k)Xi (k) (10)
where γ represents the reverberation decay for the room, δ is the level of reverber-
precedence effect [24,25] in order to give less weight to frequency bins where a loud
sound was recently present. The resulting enhanced cross-correlation is defined as:
L−1
(e) ζi (k)Xi (k)ζj (k)Xj (k)∗ 2πkτ /L
(11)
X
Rij (τ ) = e
k=0
|Xi (k)| |Xj (k)|
10
4.3 Direction Search on a Spherical Grid
In order to reduce the computation required and to make the system isotropic, we
define a uniform triangular grid for the surface of a sphere. To create the grid, we
start with an initial icosahedral grid [26]. Each triangle in the initial 20-element
The resulting grid is composed of 5120 triangles and 2562 points. The beamformer
energy is then computed for the hexagonal region associated with each of these
points. Each of the 2562 regions covers a radius of about 2.5 ◦ around its center,
Ed ← 0
τ ← lookup(d, ij)
(e)
Ed ← Ed + Rij (τ )
end for
end for
(e)
Once the cross-correlations Rij (τ ) are computed, the search for the best direction
11
on the grid is performed as described by Algorithm 1. The lookup parameter is a
pre-computed table of the time delay of arrival (TDOA) for each microphone pair
and each direction on the sphere. Using the far-field assumption [8], the TDOA in
Fs
τij = (~pi − ~pj ) · ~u (12)
c
direction of the source, c is the speed of sound and Fs is the sampling rate. Equation
12 assumes that the time delay is proportional to the distance between the source
and microphone. This is only true when there is no diffraction involved. While this
hypothesis is only verified for an “open” array (all microphones are in line of sight
the approximation is good enough for our system to work for a “closed” array (in
configuration (N = 2562, M = 8), the accessed data can be made to fit entirely in
τ ← lookup(Dq , ij)
(e)
Rij (τ ) = 0
end for
end for
12
Using Algorithm 1, our system is able to find the loudest source present by maxi-
mizing the energy of a steered beamformer. In order to localize other sources that
may be present, the process is repeated by removing the contribution of the first
how many sources are present, we always look for four sources, as this is the max-
imum number of sources our beamformer is able to locate at once. This situation
leads to a high rate of false detection, even when four or more sources are present.
the size of the grid used. It is however possible, as an optional step, to further refine
the source location estimate. In order to do so, we define a refined grid for the
surrounding of the point where a source was found. To take into account the near-
field effects, the grid is refined in three dimensions: horizontally, vertically and over
distance. Using five points in each direction, we obtain a 125-point local grid with
Fs
τij = (kd~u − ~pj k − kd~u − ~pi k) (13)
c
where d is the distance between the source and the center of the array. Equation
the direction of the source with improved accuracy. Unfortunately, it was observed
that the value of d found in the search is too unreliable to provide a good estimate
of the distance between the source and the array. The incorporation of the distance
13
nonetheless provides improved accuracy for the near field case.
5 Particle-Based Tracking
The steered beamformer detailed in Section 4 provides only instantaneous, noisy in-
formation about sources being possibly present and provides no information about
the behavior of the source in time (tracking). For that reason, it is desirable to use a
probabilistic temporal integration to track the different sound sources based on all
measurements available up to the current time. It has been shown [12,13,15] that
particle filters are an effective way of tracking sound sources. Using this approach,
all hypotheses about the location of each source are represented as a set of particles
composed of six dimensions, three for position and three for its derivative:
(t)
xj,i
(t)
(14)
sj,i =
(t)
ẋj,i
Since the particle position is constrained to lie on a unit sphere and the speed
is tangent to the sphere, there are only four degrees of freedom. The sampling
sources. The probability density function (pdf) for the location of each source is
approximated by a set of particles that are given different weights. The weights are
updated by taking into account observations obtained from the steered beamformer
14
Algorithm 3 Particle-based tracking algorithm. Steps 1 to 7 correspond to Subsec-
beamformer response
(t)
(3) Compute probabilities Pq,j associating beamformer peaks to sources being
tracked
(t)
(4) Compute updated particle weights wj,i
(7) Resample particles for each source if necessary and go back to step 1.
and by computing the assignment between these observations and the sources being
tracked. From there, the estimated location of the source is the weighted mean of
5.1 Prediction
it has been observed to work well in practice and can easily model different source
(t) (t−1)
ẋj,i = aẋj,i + bFx (15)
(t) (t−1) (t)
xj,i = xj,i + ∆T ẋj,i (16)
√
where a = e−α∆T controls the damping term, b = β 1 − a2 controls the excitation
15
• Stationary source (α = 2, β = 0.04);
(t) (t)
A normalization step ensures that xi still lies on the unit sphere ( xj,i = 1) after
source locations yq found by Algorithm 2. We also denote O(t) , the set of all
source q is a true source (not a false detection). The value of Pq can be interpreted
as our confidence in the steered beamformer output. We know that the higher the
beamformer energy, the more likely a potential source is to be true. For q > 0, false
alarms are very frequent and independent of energy. With this in mind, we define
Pq empirically as:
ν 2 /2, q = 0, ν ≤ 1
1 − ν −2 /2, q = 0, ν > 1
Pq = 0.3, q=1 (17)
0.16, q=2
0.03, q=3
16
with ν = E0 /ET , where E0 is the beamformer output energy for the first source
found and ET is a threshold that depends on the number of microphones, the frame
size and the analysis window used (we use ET = 150). Figure 3 shows an example
of Pq values for potential sources found by the steered beamformer with four people
150
100
50
Azimuth [deg]
-50
-100
-150
0 5 10 15 20 25 30 35
Time [s]
tions with Pq > 0.5 shown in red, 0.2 < Pq < 0.5 in blue, Pq < 0.2 in green.
At time t, the probability density of observing Oq(t) for a source located at particle
(t)
position xj,i is given by:
(t)
p Oq(t) xj,i = N yq ; xj,i ; σ 2 (18)
We use σ = 0.05, which corresponds to an RMS error of 3 degrees for the loca-
tion found by the steered beamformer. This error takes into account the resolution
17
5.3 Probabilities for Multiple Sources
(t)
Before we can derive the update rule for the particle weights wj,i , we must first
most one tracked source and that a tracked source can correspond to at most one
potential source.
Potential Tracked
source q source j
False
detection
Not observed
New source
Figure 4. Assignment example where two of the tracked sources are observed,
with one new source and one false detection. The assignment can be described as
observation q to the source j (values -2 is used for false detection and -1 is used for
a new source). Figure 4 illustrates a hypothetical case with four potential sources
detected by the steered beamformer and their assignment to the tracked sources.
Knowing P f O (t) (the probability that f is the correct assignment given obser-
18
vation O (t) ) for all possible f , we can derive Pq,j , the probability that the tracked
(t)
δj,f (q) P f O (t) (19)
X
Pq,j =
f
Pq(t) (H0 ) = δ−2,f (q) P f O (t) (20)
X
f
Pq(t) (H2 ) = δ−1,f (q) P f O (t) (21)
X
where δi,j is the Kronecker delta. Equation 19 is in fact the sum of the probabilities
of all f that assign potential source q to tracked source j and similarly for Equations
20 and 21.
p(O|f )P (f )
P (f |O) = (22)
p(O)
Knowing that there is only one correct assignment ( P (f |O) = 1), we can avoid
P
f
(23)
Y
p ( O| f ) = p (Oq | f (q))
q
We assume that the distribution of the false detections (H0 ) and the new sources
(H2 ) are uniform, while the distribution for tracked sources (H1 ) is the pdf approx-
imated by the particle distribution convolved with the steered beamformer error
19
pdf:
1/4π, f (q) = −2
p ( Oq | f (q)) = 1/4π, f (q) = −1 (24)
i wf (q),i p ( Oq | xj,i ) , f (q) ≥ 0
P
The a priori probability of f being the correct assignment is also assumed to come
(25)
Y
P (f ) = P (f (q))
q
with:
(1 − Pq ) Pf alse , f (q) = −2
P (f (q)) = Pq Pnew f (q) = −1 (26)
(t)
Pq P Obsj O(t−1) f (q) ≥ 0
where Pnew is the a priori probability that a new source appears and Pf alse is the a
(t)
priori probability of false detection. The probability P Obsj O(t−1) that source
(t) (t)
P Obsj O(t−1) = P Ej O(t−1) P Aj O(t−1) (27)
(t)
where Ej is the event that source j actually exists and Aj is the event that it is
active (but not necessarily detected) at time t. By active, we mean that the sig-
nal it emits is non-zero (for example, a speaker who is not making a pause). The
20
probability that the source exists is given by:
(t−1)
(t−1)
Po P Ej O(t−2)
P Ej O(t−1) = Pj + 1− Pj (28)
1 − (1 − Po ) P (Ej |O(t−2) )
where Po is the a priori probability that a source is not observed (i.e., undetected
(t)
by the steered beamformer) even if it exists (with P0 = 0.2 in our case) and Pj =
(t)
Pq,j is the probability that source j is observed (assigned to any of the potential
P
q
sources).
Assuming a first order Markov process, we can write the following about the prob-
(t) (t−1)
with P Aj Aj the probability that an active source remains active (set to
(t) (t−1)
0.95), and P Aj ¬Aj the probability that an inactive source becomes active
again (set to 0.05). Assuming that the active and inactive states are equiprobable,
the activity probability is computed using Bayes’ rule and usual probability manip-
ulations:
(t)
1
P Aj O(t) = h
(t)
ih
(t)
i (30)
1−P Aj |O(t−1) 1−P Aj |O(t)
1+
(t)
(t)
P Aj |O(t−1) P Aj |O(t)
At times t, the new particle weights for source j are defined as:
(t) (t)
wj,i = p xj,i O(t) (31)
21
Assuming that the observations are conditionally independent given the source
(t)
position, and knowing that for a given source j, wj,i = 1, we obtain through
PN
i=1
Bayesian inference:
(t) (t)
(t)
p O(t) xj,i p xj,i
wj,i =
p (O(t) )
(t) (t) (t)
p O (t) xj,i p O(t−1) xj,i p xj,i
=
p (O(t) )
(t)
p xj,i O (t) p xj,i O(t−1) p O (t) p O(t−1)
=
(t)
p (O(t) ) p xj,i
(t) (t−1)
p xj,i O (t) wj,i
= PN
(t)
(t−1)
(32)
(t) w
i=1 p xj,i |O j,i
(t)
Let Ij denote the event that source j is observed at time t and knowing that
(t) (t) (t)
P Ij = Pj = Pq,j , we have:
P
q
In the case where no observation matches the source, all particles have the same
probability, so we obtain:
(t) (t)
PQ
(t)
(t)
(t)
1 q=1 Pq,j p Oq xj,i
p xj,i O (t) = 1 − Pj + Pj PN PQ (t)
(t) (t)
(34)
N i=1 q=1 Pq,j p Oq xj,i
where the denominator on the right side of Equation 34 provides normalization for
(t) (t) (t)
the Ij case, so that
PN (t)
i=1 p xj,i O , Ij = 1.
In a real environment, sources may appear or disappear at any moment. If, at any
time, Pq (H2 ) is higher than a threshold equal to 0.3, we consider that a new source
22
is present. In that case, a set of particles is created for source q. Even when a
In the same way, we set a time limit on sources. If the source has not been observed
(t)
(Pj < Tobs ) for a certain amount of time, we consider that it no longer exists. In
that case, the corresponding particle filter is no longer updated nor considered in
future calculations.
The estimated position of each source is the mean of the pdf and can be obtained
N
(t) (t) (t)
(35)
X
x̄j = wj,i xj,i
i=1
algorithm. This can be achieved by augmenting the state vector by past position
N
(t−T ) (t) (t−T )
(36)
X
x̄j = wj,i xj,i
i=1
Using the same example as in Figure 3 we show in Figure 5 how the particle
filter is able to remove the noise and produce smooth trajectories. The added delay
23
150 150
100 100
50 50
Azimuth [deg]
Azimuth [deg]
0 0
-50 -50
-100 -100
-150 -150
0 5 10 15 20 25 30 35 0 5 10 15 20 25 30 35
Time [s] Time [s]
Figure 5. Tracking of four moving sources, showing azimuth as a function of time. Left: no
5.7 Resampling
P −1
Resampling is performed only when Nef f ≈ N 2
i=1 wj,i < Nmin [27] with
Nmin = 0.7N . That criterion ensures that resampling only occurs when new data is
available for a certain source. Otherwise, this would cause unnecessary reduction
6 Results
ber of analog input channels on commercially available soundcards. Two array con-
figurations are used for the evaluation of the system. The first configuration (C1)
left in Figure 6). The second configuration (C2) is a closed array and uses smaller,
24
on the body of the robot (shown right in Figure 6). For both arrays, all channels are
currently requires 30% of a 1.6 GHz Pentium-M CPU. Due to the low complexity
of the particle filtering algorithm, we are able to use 1000 particles per source with-
out noticeable increase in complexity. This also means that the CPU time does not
time (-60 dB) of 350 ms. The second environment (E2) is a hall (16 m × 17 m, 3.1
m ceiling, connected to other rooms) with 1.0 s reverberation time. For all tasks,
configurations and environments, all parameters have the same value, except for
the reverberation decay γ, which is set to 0.65 in the E1 environment and 0.85 in
the E2 environment.
25
6.1 Characterization
and accuracy. Detection reliability is defined as the capacity to detect and localize
sounds to within 10 degrees, while accuracy is defined as the localization error for
sources that are detected. We use three different types of sound: a hand clap, the
test sentence “Spartacus, come here”, and a burst of white noise lasting 100 ms.
The sounds are played from a speaker placed at different locations around the robot
Detection reliability is tested at distances (measured from the center of the array)
the room). Three indicators are computed: correct localization (within 10 degrees),
reflections (incorrect elevation due to roof of ceiling), and other errors. For all
played. This test includes 1440 sounds at a 22.5◦ interval for 1 m and 3 m and
360 sounds at a 90◦ interval for 5 m and 7 m. Because of the limited size of the
room used for the experiment, the tests for 5 m and 7 m had to use fixed positions
for the robot and the source, leading to less variability in the conditions. This can
explain differences between these results and those obtained for shorted distances,
C1, results show near-perfect reliability even at seven meter distance. For C2, we
noted that the reliability depends on the sound type, so detailed results for different
26
sounds are provided in Table 2, showing that only hand clap sounds cannot be
reliably detected passed one meter. We expect that a human would have achieved a
Like most localization algorithms, our system is unable to detect pure tones. This
behavior is explained by the fact that sinusoids occupy only a very small region
of the spectrum and thus have a very small contribution to the cross-correlations
with the proposed weighting. It must be noted that tones tend to be more difficult
C1 C2 C1 C2 C1 C2
In order to measure the accuracy of the localization system, we use the same setup
as for measuring reliability, with the exception that only distances of 1m and 3m are
tested (1440 sounds at a 22.5◦ interval) due to limited space available in the testing
environment. Neither distance nor sound type has significant impact on accuracy.
The root mean square accuracy results are shown in Table 3 for configurations C1
27
Table 2
Correct localization rate as a function of sound type and distance for C2 configuration
and C2. Both azimuth and elevation are shown separately. According to [28,29],
human sound localization accuracy ranges between two and four degrees in similar
conditions. The localization accuracy of our system is thus equivalent or better than
We measure the tracking capabilities of the system for multiple sound sources.
In all cases, the distance between the robot and the sources is approximately two
meters. The azimuth is shown as a function of time for each source. The elevation is
not shown as it is almost the same for all sources during these tests. The trajectories
28
for the three experiments are shown in Figure 7.
Figure 7. Source trajectories (robot represented as an X). Left: moving sources. Center:
In a first experiments, four people were told to talk continuously (reading a text
with normal pauses between words) to the robot while moving, as shown on the
left of Figure 7. Each person walked 90 degrees towards the left of the robot before
Results are presented in Figure 8 for delayed estimation (500 ms). In both environ-
ments, the source estimated trajectories are consistent with the trajectories of the
four speakers and only one false detection was present (in E1, at t = 15 s) for a
150 150
100 100
50 50
Azimuth [deg]
Azimuth [deg]
0 0
-50 -50
-100 -100
-150 -150
0 5 10 15 20 25 30 35 5 10 15 20 25 30 35 40
Time [s] Time [s]
Figure 8. Four speakers moving around a stationary robot. Left: E1, right: E2. False detec-
29
6.2.2 Moving Robot
Tracking capabilities of our system are also evaluated in the context where the
robot is moving, as shown in the center of Figure 7. In this experiment, two people
are talking continuously to the robot as it is passing between them. The robot
then makes a half-turn to the left. Results are presented in Figure 9 for delayed
estimation (500 ms). Once again, the estimated source trajectories are consistent
with the trajectories of the sources relative to the robot for both environments. Only
one false detection was present (in E1, at t = 38 s) for a short period of time.
150 150
100 100
50 50
Azimuth [deg]
Azimuth [deg]
0 0
-50 -50
-100 -100
-150 -150
5 10 15 20 25 30 35 40 5 10 15 20 25 30
Time [s] Time [s]
Figure 9. Two stationary speakers with the robot moving. Left: E1, right: E2. False detection
shown in black.
In this experiment, two moving speakers are talking continuously to the robot, as
shown on the right of Figure 7. They start from each side of the robot, intersecting
in front of the robot before reaching the other side. Results in Figure 10 show that
the particle filter is able to keep track of each source. This result is possible because
the prediction step imposes some inertia to the sources and despite the fact that the
steered beamformer typically only “sees” one source when the two sources are very
close.
30
100
100
50
50
Azimuth [deg]
Azimuth [deg]
0 0
-50
-50
-100
-100
2 4 6 8 10 12 14 16 18 4 6 8 10 12 14 16
Time [s] Time [s]
Figure 10. Two speakers intersecting in front of the robot. Left: E1, right: E2.
These results evaluate how the number of microphones used affect the system ca-
pabilities. To do so, we use the same recording as in 6.2.1 for C2 in E1 with only a
the system for four to seven microphones (selected arbitrarily as microphones num-
are removed. While using seven microphones makes little difference compared to
the baseline of eight microphones, the system is unable to reliably track more than
two of the sources when only four microphones are used. Although there is no theo-
retical relationship between the number of microphones and the maximum number
of sources that can be tracked, this clearly shows the how redundancy added by
using more microphones can help in the context of sound source localization.
This experiment is performed in real-time and consists of making the robot follow
the person speaking to it. At any time, only the source present for the longest time is
31
4 microphones 5 microphones
150 150
100 100
50 50
Azimuth [deg]
Azimuth [deg]
0 0
-50 -50
-100 -100
-150 -150
0 500 1000 1500 2000 2500 3000 3500 4000 0 500 1000 1500 2000 2500 3000 3500 4000
Time [s] Time [s]
6 microphones 7 microphones
150 150
100 100
50 50
Azimuth [deg]
Azimuth [deg]
0 0
-50 -50
-100 -100
-150 -150
0 500 1000 1500 2000 2500 3000 3500 4000 0 500 1000 1500 2000 2500 3000 3500 4000
Time [s] Time [s]
Figure 11. Tracking of four sources using C2 in the E1 environment, using 4 to 7 micro-
phones.
considered. When the source is detected in front (withing 10 degrees) of the robot,
it is made to go forward. At the same time, regardless of the angle, the robot turns
toward the source in such a way as to keep the source in front. Using this simple
control system, it is possible to control the robot simply by talking to it, even in
This has been tested by controlling the robot going from environment E1 to envi-
ronment E2, having to go through corridors and an elevator, speaking to the robot
with normal intensity at a distance ranging from one meter to three meters. The sys-
delay on the estimator) with the robot reaction time limited mainly by the inertia of
the robot. One problem we encountered during the experiment is that when going
through corridors, the robot would sometimes mistake reflections on the walls for
real sources. Fortunately, the fact that the robot considers only the oldest source
32
present reduces problems from both reflections and noise sources.
7 Conclusion
localize and track simultaneous moving sound sources in the presence of noise and
system is capable of controlling in real-time the motion of a robot, using only the
ment problem is also applicable to other multiple objects tracking problems. Other
A robot using the proposed system has access to a rich, robust and useful set of in-
formation derived from its acoustic environment. This can certainly affect its ability
of making autonomous decisions in real life settings, and show higher intelligent
behavior. Also, because the system is able to localize multiple sound sources, it can
performed. This will allow to identify the localized sound sources so that additional
33
Acknowledgment
Jean-Marc Valin was supported by the National Science and Engineering Research
Council of Canada (NSERC) and the Quebec Fonds de recherche sur la nature et
les technologies (FQRNT). François Michaud holds the Canada Research Chair
also supported financially by the CRC Program and the Canadian Foundation for
Innovation (CFI). Special thanks to Brahim Hadjou for help formalizing the particle
filtering notation and to Dominic Létourneau and Pierre Lepage for help controlling
References
[1] M. Marschark, Raising and Educating a Deaf Child. Oxford University Press, 1998,
[Link] memrtl/course/interpreting/modules/[Link].
[2] W. M. Hartmann, “How we localize sounds,” Physics Today, pp. 24–29, 1999.
moving sound sources for mobile robot using a frequency-domain steered beamformer
Systems, Man, and Cybernetics Part B, vol. 34, no. 3, pp. 1526–1540, 2004.
34
[6] K. Nakadai, T. Lourens, H. G. Okuno, and H. Kitano, “Active audition for humanoid,”
[7] F. Asano, M. Goto, K. Itou, and H. Asoh, “Real-time source localization and separation
[8] J.-M. Valin, F. Michaud, J. Rouat, and D. Létourneau, “Robust sound source
[9] S. Kagami, Y. Tamai, H. Mizoguchi, and T. Kanade, “Microphone array for 2D sound
[10] D. Bechler, M. Schlosser, and K. Kroschel, “System for robust 3D speaker tracking
Conference on Acoustics, Speech, and Signal Processing, vol. II, 2002, pp. 1777–1780.
[14] J. Vermaak and A. Blake, “Nonlinear filtering for speaker tracking in noisy
35
Acoustics, Speech, and Signal Processing, vol. 5, 2001, pp. 3021–3024.
and J. Ogata, “An application of a particle filter to bayesian multiple sound source
tracking with audio and video information fusion,” in Proceedings of 7th International
pp. 1950–1954.
multiple objects,” International Journal of Computer Vision, vol. 39, no. 1, pp. 57–
71, 2000.
[18] C. Hue, J.-P. L. Cadre, and P. Perez, “A particle filter to track multiple objects,” in
[19] J. Vermaak, S. Godsill, and P. Pérez, “Monte carlo filtering for multi-target tracking
and data association,” IEEE Transactions on Aerospace and Electronic Systems, 2005.
(To appear).
[20] R. Duraiswami, D. Zotkin, and L. Davis, “Active speech source localization by a dual
[22] Y. Ephraim and D. Malah, “Speech enhancement using minimum mean-square error
36
[23] I. Cohen and B. Berdugo, “Speech enhancement for non-stationary noise
model of the precedence effect,” Speech Communication, vol. 27, no. 3-4, pp. 223–
233, 1999.
[27] A. Doucet, S. Godsill, and C. Andrieu, “On sequential Monte Carlo sampling methods
for bayesian filtering,” Statistics and Computing, vol. 10, pp. 197–208, 2000.
37