MPEG Surround
Synthesis by Tewabe chekole, ctewabe@[Link]
Abstract
The paper describes a proposal for an extension of the MPEG set of standards to surround audio
(5.1 audio). It describes an audio processing technique able to code multichannel audio into a
compressed stereo bitstream (downmixing) plus an additional low bitrate channel carrying
spatial information for multichannel sound rebuilding (upmixing). The paper surveys the coding
scheme, discusses the backward compatiiblity with conventional stereo equipment and evaluates
the benefits of this solution for several application environments, in particular for Digital Audio
Broadcasting.
Synthesis of the article
The problem approached by the paper [1] can be defined as the coding and transmission of multichannel
surround audio using traditional stereo media formats, to achieve three goals: (1) to be backward compatible
with stereo audio devices, without additional signal processing functions to mix multichannel audio to stereo
audio; (2) to use bandwith adequate for stereo transmission also for transmitting multichannel audio content;
(3) to rebuild the multichannel audio at destination using a limited amount of additional information whose
impact on the bandwith is small. The problem is framed into the MPEG standard and is a proposal for a new
MPEG surround coding scheme.
The proposal is based on the analysis of the human perception and the architecture of surround audio
systems. The human localization of sound in space is based on interaural difference, i.e., the differences in
perception between the signals received by the two ears; such differences are due to three casuses: (1) level
difference, due to the more or less direct arrival of sound to the ears depending on the distance and
position of the sound source; (2) time difference, due to the same cause; (3) interaural coherence,
due to the reflections in a reverberant environment, which solicits both ears with correlated signals
independently from their origin. Figure 1 (corresponding to Figure 1 in the paper), show these concepts.
Stereo sound perception is limited to the front of the listener, since no sound is generated behind him/her.
Surround audio (commonly known as 5.1 audio in domestic appliances) uses five channels to create a
360° spatial distribution. Besides the usual frontal left (L) and right (R) channels, an additional central
channel (C) and two rear channels, called Ls and Rs (s standing for surround) are added, covering the whole
space around the listener; a sistxh channel (the “.1” in the 5.1 scheme) is called LFE (Low Frequency
Effects) and is used for special audio effects at very low frequency.
The surround coding scheme proposed in the paper uses a dowmixing algorithm (not described in the
paper) able to linearly combine N channels into two channels corresponding to the left and right areas of the
sound space. Additional information in form of spatial parameter estimation is coded based on the
correlation of signals, and sent over a low bandwidth additional channel.
At decode time the decompression of the two channels signal rebuilds left and right components of the
original sound. A stereo decoder could ignore this information and stop here, while a multichannel decoder
can use the estimated spatial information, transmitted separately from the mixed audio, to rebuild front, rear
and center signals. Figure 2 (corresponding to Figure 2 in the original paper), shows the block diagrams of
this process. The paper does not explain the details of the downmixing and upmixing algorithms, noting that
dowmixing can be performed by hand (the so called artistic downmixing) for better results. Similarly,
the spatial parameters estimation is only surveyed; it is based on a hierarchical data structure containing four
information types: the channel level differences, the interchannel correlation, the channel prediction
coefficients and residual prediction errors. Further details are left to the referenced paper.
The proposed scheme does not define the audio compression algorithm, which can be selected among the
standards like MP3, AAC, etc..
The paper then discusses the quality of the two channel transmitted audio with respect to stereo and to
sorround playback, noting that the quality of the stereo playback depends on the quality of the downmix; an
interesting cue is that a hand-made downmix should produce the best quality for stereo listening, being able
to balance the different channel to create a correct left/right composition of the original signals.
1
The decoder uses the spatial cues to upmix the stereo signal to a multichannel signal. The upmix occurs
through a filter bank in the frequency domain; details of the algorithm are not provided in the paper but left
to the references.
After a brief discussion about the state of this proposal in the MPEG committee, the paper discusses the
benefits with respect to a set of applications: Digital Audio Broadcasting (DAB), music download services,
Internet streaming, teleconferencing and gaming. The contribution of multichannel audio to each of these
application environments are briefly discussed; a more complete discussion is devoted to Digital Audio
Broadcasting, which in the authors’ opinion is the most promising application field from a commercial
perspective. Comparing the diffusion of the 5.1 audio in home theatre equipment, the extension of digital
broadcast with 5.1 audio support is considered a killer applicaton. The authors cite, among others, the
market of luxury cars, already equipped with surround audio and prone to high cost optional equipments.
Comments
The paper is a survey paper and does not enter into the details of the coding scheme; the technical adequacy
of such approach cannot be evaluated on the base of this paper only. The benefits noted by the authors seem
interesting, mainly from two points of view: backward compatibility and independance from the audio
compression algorothm. Backward compatibility should assure the adoption of this schema without need for
a conversion of the user equipment; independence from the compression algorithm should ensure the widest
applicability of this scheme also in legacy applications.
The paper was published in 2005. Since then, the proposed MPEG Surround has become an ISO/IEC
standard [2], also known as MPEG-D standard; software solutions are distributed by the Fraunhofer Institute
[3], developer of the MP3 compression standard, to which one of the authors of the paper was affiliated.
References
1. S. Quackenbush, J. Herre. MPEG Surround, IEEE Multimedia, vol. 12, no. 4, pp. 18-23, Oct.-Dec. 2005
2. [Link]
3. [Link]
2
Figure [Link] of interaural level differences(ILD), interaural
time differences (ITD), and interaural coherence (IC).
Figure 2. Principles of MPEG surround