0% found this document useful (0 votes)
13 views26 pages

Neural Networks in Signal Processing

The document discusses the integration of neural networks into statistical signal processing, emphasizing their ability to handle nonlinearity, nonstationarity, and non-Gaussianity. It outlines the advantages of neural networks over traditional methods, including adaptability and robustness, and introduces the principle of empirical risk minimization as a key concept. The article also highlights the importance of preserving information in signal processing tasks for optimal performance.

Uploaded by

Velu Siva
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
13 views26 pages

Neural Networks in Signal Processing

The document discusses the integration of neural networks into statistical signal processing, emphasizing their ability to handle nonlinearity, nonstationarity, and non-Gaussianity. It outlines the advantages of neural networks over traditional methods, including adaptability and robustness, and introduces the principle of empirical risk minimization as a key concept. The article also highlights the importance of preserving information in signal processing tasks for optimal performance.

Uploaded by

Velu Siva
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

ral Networks Expan

orizons

Advanced algorithms for


signal processing
simultaneously account
for non1inearity,
nonstationarity,
and non-Gaussianity

SIMON HAYKIN

tatistical signal processing covers an area where phys- Interest in neural networks, or to be more precise, artificial
ics and mathematics meet and interact to solve a wide neural networks, has always been motivated by the fact that
range of problems. Its origins may be traced back to the the human brain functions in a manner entirely different from
1943 classified RCA report by North, republished in [l], the the conventional digital computer. The human brain is a
1946 classic paper [Z]by Van Vleck and Middleton, and the gigantic, and yet highly efficient, information-processing
pioneering work by Wiener [ 3 ] . In particular, the classical machine that encompasses a wide variety of complex signal
methods of statistical signal processing are founded on three processing operations. To appreciate the enormous scale of
basic assumptions: linearity, stationarity, and second-order these operations, we need only look at our visual and auditory
statistics with particular emphasis on Gaussianity. These systems and be amazed at the “seamless” nature of the way
assumptions are invoked for the sake of mathematical tracta- in which different forms of information gathered by our eyes
bility. Yet most, if not all, the physical signals that we have and ears are individually processed and then finally fused
to deal with in real-life applications are generated by dynamic together.
processes that are simultaneously nonlinear, nonstationary, Work on neural networks may be traced back to the
and non-Gaussian. The end result of designing a signal-proc- pioneering paper [4] by McCulloch and Pitts in 1943, which
essing system along traditional lines is a suboptimal solution. was followed by Rosenblatt’s development of the perceptron
One way in which the performance of the system can be [5] and Widrow’s development of the adaline [6] in the late
improved is to consider the use of neural networks in combi- 1950s. After going through a period of dormancy (in an
nation with other suitable techniques (e.g., time-frequency engineering context) in the 1970s, neural networks re-
analysis), depending on the task at hand. emerged in the 1980s with the publication of Hopfield’s

24 IEEE SIGNAL PROCESSING MAGAZINE MARCH 1996


1053-5888/96/$5 0001YY6IEEE
paper on recurrent networks [7] and the two-volume seminal Neural networks have a natural ability to adapt their free
book [8] by Rumelhart and McClelland on parallel distrib- parameters to statistical changes in the environment in which
uted processing (PDP). We may look back on the 1980s not they operate.
only as the decade of re-emergence of neural networks but As a rule of thumb, we may say that the more we make a
also as one of consolidation. nonlinear system adaptive, the more robust the performance
Insofar as this article is concerned, the primary interest is of that system is likely to be when it operates in a nonstation-
in the use of neural networks as an engineering tool for signal ary environment, subject, of course, to the requirement that
processing applications. The aim of the article is three fold: the system remains stable. (We ourselves are a living example
of this rule.) However, for the full benefits of adaptivity to be
Articulate a new philosophy in the approach to statistical realized, there has to be a successful resolution to the stabil-
signal processing using neural networks;, which (either by ity-plasticity dilemma. This means that the principal time
themselves or in combination with other suitable tech- constants of the system should be long enough to ignore
niques) account for the practical realities of nonlinearity, spurious disturbances, and yet short enough to respond to
nonstationarity, and non-Gaussianity meaningful changes in the environment. Ordinary adaptive
filters also have the ability to adjust their parameters auto-
Describe three case studies using real-life data, which matically in accordance with statistical variations of their
clearly demonstrate the superiority of this new approach environment [10,l I]; however, their adaptive signal process-
over the classical approaches to statistical signal process- ing capability is limited by their structural formulation as
ing simple linear combiners.

Discuss mutual information as a criterion for designing Neural networks provide a nonparametric approachfor the
unsupervised neural networks, thus moving away from the nonlinear estimation of data
mean-square error criterion The nonlinear, feedforward multilayer class of neural net-
works (encompassing multilayer perceptrons and radial ba-
Rationale for Using Neural Networks sis-function networks) learns about its environment in a
supervised manner. (The design of multilayer perceptrons
Neural networks have a number of important properties that and radial-basis function networks is discussed in the book
befit their use for signal processing applications. In particu- by Haykin [ 121, and the review papers by Lippmann [ 131and
lar, we mention the following five properties: Hush and Horne [ 141.) Specifically, these neural networks
undergo a training session during which their free parameters
Neural networks are distributed nonlinear devices (i.e., synaptic weights and biases) are adjusted in a systematic
This property is a direct result of the fact that each processing way so as to minimize a cost function. Typically, the cost
unit (i.e., neuron) of a neural network has a built-in activation function is defined on the basis of a mean square-error
function (for example, in the form of a logistic function) that criterion, with the error signal itself being defined as the
is nonlinear. Accordingly, neural networks have the inherent difference between a desired response and the actual output
ability to model underlying nonlinearities contained in the of the network produced in response to a corresponding input
physical mechanism responsible for generating the input signal. The neural network learns from examples by con-
data. structing an input-output mapping for the problem at hand,
which brings to mind the notion of nonparametric statistical
A neural network consists of a massively parallel processor inference; see Table 1. The term “nonparametric”is used here
that has the potential to be fault tolerant in a statistical sense, meaning that no knowledge of the
For example, a multilayer perceptron, representing a popular underlying probability distribution is required.
structure for the implementation of a neural network, consists In the traditional approach to mathematical statistics as
of a large number of neurons arranged in the form of layers, taught in a statistics department, the issues of primary con-
with each neuron in a particular layer connected to a large cern are two-fold:
number of source nodestneurons in the previous layer. This
form of global interconnectivity has the potential to be fault
tolerant, in the sense that the performance is degraded grace- The use of mathematically tractable models, assuming the
fully under adverse operating conditions. If a neuron or its idealized conditions of linearity, wide-sense stadonarity,
synaptic links are damaged, the recall quality of a stored and Gaussianity, for the derivation of parameter estimators.
pattern is impaired, but owing to the highly distributed nature
of the network, the damage has to be extensive before the Derivation of exact properties (e.g.,mean and variance) of
performance is seriously degraded. Nevertheless, to be as- estimators for small sample-sizes; if the exact properties
sured that the neural network is in fact fault tolerant, we may are not mathematically tractable, then one would consider
find it necessary to take proper measures in designing the the asymptotic properties of the estimators as the number
algorithm used to do the training [8]. of samples approaches infinity.

MARCH 1996 IEEE SIGNAL PROCESSING MAGAZINE 25


In contrast, neural network-based methods are attractive of basis functions represented by the outputs of the hidden
for practical applications by virtue of their ability to deal with neurons are optimized simultaneously in an iterative fashion,
nonlinearity, nonstationarity, and non-Gaussianity. More- hence the more robust behavior.
over, they offer robustness with respect to parameter tuning In addition to robustness, there are other considerations to
and sample properties, which is important for a good setting be taken into account, such as prediction accuracy in the
of user-tunable parameters by non-expert users. It is not quite context of dense samples or sparse samples, where “dense-
clear why, in this respect, neural networks appear to behave ness” is measured with respect to the target function com-
better than comparable statistical techniques such as projec- plexity (i.e., smoothness). Insofar as prediction accuracy is
tion pursuit [27,28], splines [29], and multivariate adaptive concerned, it can be said that there is no single method that
regression splines (MARS) [30].Projection pursuit is simi- provides a superior performance under all possible situations
lar and mathematically equivalent to the multilayer percep- [31]. Evidence supporting this claim, using computer simu-
tron in terms of representation. Splines are closely related to lations on various statistical methods including neural net-
radial-basis function (RBF) networks. MARS may be viewed works, is presented in [25].
as a tree of neurons with each leaf of the tree consisting of a
neuron; the neuron may itself be modeled as a piecewise Neural networks, operating in a supervised manner, are
linear polynomial or a cubic polynomial with the knot of the universal approximators
spline treated as a variable. Multilayer feedforward networks (i.e., multilayer percep-
A possible explanation for the superiority of neural net- trons and radial-basis function networks) are universal ap-
works may be found in the differences in the way in which proximators, in the sense that they can approximate any
the respective optimization procedures are pursued [32].In continuous input-output mapping to any desired degree of
statistical methods, particularly those that use a greedy form approximation, given a sufficient number of hidden units
of optimization with the basis functions tuned one at a time, [33-351. This property is also shared by classical methods
it may be difficult or perhaps impossible to recover from any based on the use of smooth functions such as algebraic or
wrong decisions made in some early stages of the optimiza- trigonometric polynomials. What is really important, there-
tion process. In contrast, in a neural network the complete set fore, is the rate of convergence with which the unknown

26 IEEE SIGNAL PROCESSING MAGAZINE MARCH 1996


function is approximated for a prescribed set of basis func- sired response vector, di, corresponding to an input vector,
tions. In classical approximation theory involving bounded Xi, and the actual response, F(xi,w), produced by the neural
norms of the derivatives of orders for some s > 0, the rate of network.
convergence is O(n-2S’(2S+P)), where n is the degree of the Let Wemp denote the parameter vector that minimizes
polynomial, and p is the dimensionality of the input space. Remp(W) over the parameter space. According to the principle
The dependence of the rate of convergence on p in the of empirical risk minimization, the functional Remp(W) con-
exponent is a manifestation of the curse of dimensionality. verges in probability to the minimum possible value of the
The implication here is that in a high-dimensional space (i.e., actual risk functional R(w) as the size, N , of the training set
large p ) one can only approximate very smooth functions is made infinitely large, provided that the empirical risk
(i.e., large s) for a given number of samples (i.e., prescribed functional Remp(w) converges uniformly to the actual risk
n). For the corresponding case of neural networks, Barron functional, R(w). The theory of uniform convergence of
[36] has shown that it is possible to approximate any function Remp(w) to R(w) includes bounds on the rate of convergence,
satisfying a certain condition on its Fourier transform by a which are, in turn, based on the VC-dimension. For a discus-
multilayer perceptron, with the rate of convergence being sion of these issues, see [12,39,40].
O( 1/ & ), where n is the number of sigmoid basis functions
(i.e., hidden neurons). Even though this result has been Information Preservation Rule
(mis)interpreted as if the use of neural networks overcomes
the curse of dimensionality (i.e., the rate of convergence does
Neural networks may not be adequate to tackle all statistical
not depend on the dimensionality p of the input space),
signal processing applications by themselves. Rather, neural
careful examination of the result shows that with increasing
networks may have to be integrated with other related tech-
dimensionality p , the smoothness of the function being ap-
niques in a principled way in order to capture the full infor-
proximated would have to be increased to ensure that the
mation content of the input data and exploit the information
condition on the bounded norm of its Fourier transform is
in an efficient manner. A particular technique that lends itself
satisfied [37].
to this approach is that of time-frequency (scale) analysis
[41-421, by means of which a one-dimensional signal is
Principle of Empirical Risk Minimization
transformed into a time-frequency (scale) image. (For exam-
ple, as illustrative ways in which wavelets, a popular method
In a real-life situation, we have to work with a finite sample
for performing time-scale analysis, can be integrated with
size, irrespective of the statistical estimation procedure used.
neural networks for different applications, see [43-511.) The
With the notable exception of Vapnik’s pioneering work that
useful feature of time-frequency transformation is that it
remains largely unknown to the signal processing commu-
displays the temporal localization of the signal’s spectral
nity, there exists no widely accepted theory for small-size
components in a more discernible fashion than would be the
nonparametric estimation.
case directly from the signal or its spectrum.
Vapnik’s work hinges on an important parameter called
Whatever form of system integration is used, the design
the Vapnik-Chervonenkis dimension, or simply the VC-di-
objective should be in accord with an information-theoretic
mension [38]. In the context of pattern classification, the
rule of thumb that may be stated as follows:
VC-dimension provides a measure of the capacity of the
family of classification functions realizedl by a learning ma-
In designing a receiver, the available information per-
chine.
taining to a signal-processing task (e.g., target detection or
The VC-dimension plays a central role in the principle of
parameter estimation) should be preserved optimally (in a
empirical risk minimization, an inductive principle that does
statistical sense) and used efficiently (in a computational
not require probability density estimation. This makes it
sense), until the receiver is ready for final decision-making.
perfectly suited to the underlying premise of neural networks.
(This rule is based on one of three lessons learned from
The basic idea of the method is to use a set of N identically
information theory; see Viterbi [52].)
and independently distributed (iid) training examples
(xi ,di),(x2,d~),...,(x~,d~)to construct the empirical risk func-
In the sequel, we refer to this rule as the information
tional [39]:
preservation rule. To appreciate the practical significance of
this rule, consider the case of remote sensing. In this kind of
application, we typically find that the sensors (e.g., the an-
tenna in a weather radar system and its electromagnetic
which does not depend on the unknown probability that accessories) represent a highly significant part of the capital
pertains to the generation of the training examples. In this investment involved in building the system. Most impor-
equation, Xi and di denote the input vector and the desired tantly, it is unlikely that the cost of the sensors would go
response vector for the ith training example, respectively, and down. In direct contrast, signal-processing subsystems are
w is the set of free pararneters (weights) selected by the becoming progressively cheaper, thanks to very-large-scale
learning machine (i.e., neural network). The function L(di; integration (VLSI) technology. This is all the more reason for
F(xi,w)) represents the loss or discrepancy between the de- adhering to the information preservation rule, thereby putting

MARCH 1996 IEEE SIGNAL PROCESSING MAGAZINE 27


uncertainty to the reconstructed input-output mapping. To
make the learning process well posed, some prior information
(e.g., smoothness constraints) on the input-output mapping
must be included in the formulation of the learning algorithm.
environment This is achieved by adding aregularizing term (i.e., stabilizer)
to the cost function used to derive the algorithm for training
the neural network [12, 53, 541.
To sum up, we may identify two different approaches to
statistical signal processing, as indicated in Fig. 1. In the
parametric approach depicted in Fig. la, we start with a
Test the receiver performance statistical model of the underlying physical mechanism re-
with real-lofe data
sponsible for generating the input data, and then use the
model to design the receiver. The success of this approach .
depends on how closely the model describes the realities of
the physical mechanism responsible for generating the data.
In direct contrast, in the nonparametric approach depicted in
Fig. lb, the information-processing machine provides not
only a statistical model of the environment in which it oper-
ates, but also the final receiver design. The success of this
latter approach depends on how representative the training
data are of the physical environment, and how adequate the
size of the training data is. Neural networks, viewed in a
statistical sense, belong to the approach described in Fig. lb.

Train the machine with


real-life data Criteria for Acceptance of Neural Networks
(b)
. Two classes of statistical signal processing techniques.
In assessing the engineering attributes of a “good” signal
processor, we come upon two particular attributes:
the sensors to their most cost-effective use. Neural networks
have the potential to preserve information by virtue of their * Optimal preservation of the available information, and
ability to learn a model of their environment through expo- therefore optimality of performance in some statistical
sure to input-output examples that are representative of the sense
environment. For the information preservation rule to be
satisfied, however, the structural complexity of the neural 0 Robustness of performance with respect to small variations
network should closely match the underlying complexity of in environmental conditions
the input data. This raises the issue of network complexity
that has attracted the attention of many researchers, building Given these attributes, neural networks can gain acceptance
on statistical criteria such as Rissanen’s minimum-descrip- as tools for solving statistical signal processing problems, in
tion length (MDL) criterion and cross validation; see [ 18,261 preference to traditional methods, if
for a discussion of this important design issue.
Successful design of a neural network rests not only on the (i) using a neural network makes a significant difference
right selection of a network structure but also the availability in the statistical performance of a system for a real-world
of a reliable training set, (i.e., a training set that is precise and application, or can provide a significant reduction in the cost
relatively noise-free). If we recognize that, inespective of the of implementation without compromising performance
design methodology, the statistical performance of a signal-
processing system must be evaluated with real-life data prior (ii) by virtue of its massively parallel and distributed
to use in an operational setting, we may just as well start with structure, a neural network offers a more graceful degrada-
the collection of a “labeled” dataset (i.e., ground truthed) that tion of performance due to the unavoidable failure of network
will be representative of the particular environment of inter- components than would be possible with other nonparametric
est. Part of the dataset is used to train the neural network, and methods
the remaining part is subsequently used to test it. Unfortu-
nately, the learning process is an ill-posed inverse problem (iii) the tuning of adjustable parameters in a neural net-
for the following reasons [12]. First, the information content work is a more straightforward task (and therefore easily
of the training data may not be sufficient to reconstruct the accomplished by a nonexpert user) than would be the case
input-output mapping uniquely. Second, the unavoidable with other nonparametric methods
presence of noise or imprecision in the training data adds

28 IEEE SIGNAL PROCESSING MAGAZINE MARCH 1996


(iv) through the use of a neural network, by itself or in paper with Li [88] probably introduced the term as a mathe-
combination with some other devices, we are able to solve matical concept, though many others had previously talked
difficult signal processing problems for which there are no about chaotic fluid behaviors. Chaos, representing a new
viable solutions using standard methods paradigm, owes its origin to the pioneering work of Lorenz
on simulated weather data [89]; some of Lorenz’s personal
A practical limitation of neural networks is that when recollections of his first computer model of the atmosphere
working with real-life data, training for an application may appear in [90]. What came out of that computer simulation is
take a long time; the length of training would naturally have now known as the Lorenz attractor. An attractor represents
to be viewed in the context of available computing resources. the equilibrium state of a nonlinear dynamical system, which
The relatively long time needed to train a neural network is may be observed experimentally after the transients have
largely due the computer architecture (serial in nature) in died. The Lorenz attractor is an example of a strange attractor.
current use, which is ill suited to programming neural net- The strangeness comes from two important features: unlike
works. Special-purposeprocessors (e.g., the ANNA chip [55] a smooth curve or surface, a strange attractor is an object with
and CNAPS [56]) are available today that can speed up the a fractal (i.e., non-integer) dimension; unlike an ordinary
training process significantly for specific types of neural attractor, the motion of a strange attractor exhibits sensitive
networks. In addition, the back-propagation algorithm dependence on initial conditions.
(widely recognized as the workhorse for the design of neural The term “strange attractor” was coined in a paper by
networks) lends itself to parallelism. Indeed, many papers Ruelle and Takens [91], in which they claimed that turbulent
have been written on this issue; see [57] and the references flow is not described by a superposition of many modes (as
listed herein. Through the use of parallelism, the training previously proposed) but by strange attractors. The existence
process of a multilayer perceptron required to tackle a large of chaos in fluid turbulence is confirmed in [92].
problem may be facilitated by using a large number of paral- Using an extensive and ground-truthed database collected
lel processors and distributing the synaptic weights of the by means of an instrument-quality X-band radar (called the
network over these processors. IPIX radar) pointing along a fixed direction and dwelling
Another weakness is that it is often difficult to see how onto a patch of the ocean surface, researchers have demon-
knowledge gained by the neural network about its environ- strated the chaotic nature of sea clutter in light of what is
ment is actually represented inside the network. Some dis- known about chaos theory. The clutter-to-noise ratio of the
play/graphical tools such as the Hinton diagram and the bond data collected with this radar was on the order of 30 dB, and
diagram have been developed to remedy this difficulty [ 12, the wordlength of the A/D converter was 8 bits (equivalent
58, 591. to a dynamic range of 48 dB). Important aspects of the
research findings reported in [65,66] may be summarized as
Case Studies follows:

Now consider three case studies (based on real-life data), with 1. The largest Liapunov exponent, X i , is always positive. For
which I and some of my research colleagues have been the particular radar used to do the data collection, X i is
involved for the past six years. These studies, in their own estimated to be about 0.03, which is normalized with respect
ways, testify to the computing power of neural networks in to the pulse-repetition period of the radar. This value is
solving difficult signal processing problems. essentially independent of the following (for a given radar
system):

Case Study I: Chaotic Modeling of Sea Clutter and radar parameter (i.e., amplitude, in-phase component, or
its Cancellation quadrature component
sea state
For nearly half a century, sea clutter (i.e., the radar backscat- radar location
ter from an ocean surface) has been modeled as a stochastic
process, with a variety of probability distributions proposed Moreover, the second Liapunov exponent, X2, is very close
for describing its stochasticity [60-631. However, there is to zero, and for a prescribed embedding dimension, the sum
now strong experimental evidence that shows that sea clutter of all the Liapunov exponents is negative. The implications
is indeed a chaotic process [64-661. of these latter observations are twofold:
A chaotic process is generated by a deterministic mecha-
nism of a relatively low dimension, and yet it generates a Sea clutter is generated by a coupled system of nonlinear
randomlike waveform that exhibits many of the charac- differential equations
teristics that are normally associated with a stochastic proc- The dynamic mechanism responsible for the generation of
ess. Table 2 presents a summary of the important properties sea clutter is a dissipative one
of a chaotic process.
The term “chaos” was coined by J.A. Yorke, an applied 2. The correlation dimension,Dc, is fractal (i.e., non-integer),
mathematician at the University of Maryland [87]. Yorke’s lying in the range of 6 to 9. Moreover, it is also essentially

MARCH 1996 IEEE SIGNAL PROCESSING MAGAZINE 29


30 IEEE SIGNAL PROCESSING MAGAZINE MARCH 1996
Input data
vector, of size
N=zD,
- Neural
Network
Prediction one step
into the future
x(n-2),...,x( n-p) is applied to the input layer of the network,
and its synaptic weights are adjusted to minimize the
prediction error [i.e., the difference between the actual
sample value x(n) and the predicted value, i ( n ) ]in a
mean-square sense. For this to be attained, the size of the
training set has to be large enough, and the training session
would have to be continued until the synaptic weights of
the network reach steady-state values, whereafter they are
fixed.

Unit delay The neural network is next tested for its generalization
x(n-1) performance, as depicted in Fig. 2b. The network is initial-
ized by presenting it a set of samples xk-1), @2), ...,x( 1)
that have not been seen by the network before. The result-
Recursive prediction. ing prediction i (p) is delayed by one time unit and then
fed back to the input. Correspondingly, the samples x@-
independent of radar parameter, sea state, and radar location. 1),..4(2) are each delayed by one time unit, and the oldest
However, unlike the Liapunov spectrum, the correlation di- sample x(1) is dropped to make room for the delayed
mension is essentially independent of the radar system used prediction i ( p ) . The set of p samples so obtained is used
to perform the data collection. to make a new prediction, and the process is repeated until
all the original samples used to do the initialization have
Sea clutter is indeed generated by a chaotic process. But been removed from the recursive prediction process. From
the mechanism by which this chaotic process actually arises that point on, the neural network operates in a completely
in physical terms is unknown. autonomous fashion, producing a time series that is repre-
With this background on the chaotic nature of sea clutter, sentative of the dynamics learned by the neural network as
we may now turn attention to its signal processing implica- a result of the training process.
tions. Specifically, we wish to (1) demonstrate that sea clutter
permits a nonlinear predictive model with a significant hori- For the predictive modelling of sea clutter, we used a
zon of predictability, and (2) describe a novel radar applica- multilayer perceptron trained with the backpropagation algo-
tion exploiting this predictive capability. rithm. The size of the input layer, denoted by p , is chosen in
The nonlinear predictive modelling of sea clutter involves accordance with the formula p 2 rDE, where DE is the
the use of recursive (iterated) prediction [12,93], illustrated embedding dimension, and 7 is the embedding delay (nor-
in Fig. 2. This is a difficult procedure designed to test the malized with respect to the pulse repetition period). In prac-
generalization capability of the model. For the case study tice, it is inadvisable to choose p much larger than the lower
presented here, a neural network is used as the predictive bound, TDE,as the effect of additive noise contaminatingthe
model. There are two separate operations to be considered: input radar data would become more pronounced. Based on
measurements on real-life data reported in [66], the embed-
The neural network is trained to operate as a one-step ding dimension DE for sea clutter is estimated to be 10. Also,
predictor, as depicted in Fig. 2a. A set o f p samples x(n-1), for the particular radar (operating at a pulse repetition fre-

MARCH 1996 IEEE SIGNAL PROCESSING MAGAZINE 31


quency of 1kHz)used in those measurements, the normalized
embedding delay, T , is estimated to be 5. Thus, for the
problem at hand, choosing p 50 is logical.
To illustrate the importance of this lower bound onp, Fig.
3 shows the results obtained using recursive prediction per-
formed on a multilayer perceptron that had already been
trained (using the back-propagation algorithm) on actual sea
clutter data. The size of the training dataset was lo4 clutter
samples. The multilayer perceptron had two hidden layers:

Input layer: 50 source nodes


First hidden layer: 80 neurons
Second hidden layer: 55 neurons
Output layer: 1 neuron

The neurons in both hidden layers used a logistic function for 50 100 150 200 250
Time
their activation functions, whereas the output neuron was
linear. The solid line in Fig. 3 is the actual sea clutter wave- I
4. Sensitivity o f t h e recursive prediction process to a change in
form, and the dashed curve is the result of recursive predic- the neural network design. The network used f o r this experiment
tion. The origin in this figure corresponds to the end of the consists of a n input layer with 45 source nodes, two hidden layers
initialization procedure. We see that for about the first 50 with 80 and 55 neurons, respectively, and one output neuron.
points shown in Fig. 3, the predicted and actual waveforms
of sea clutter match fairly closely and thereafter they diverge.
made previously: the generation of sea clutter is governed by
a coupled system of nonlinear differential equations. In ef-
100 fect, the neural network provides an approximation to such a
1
system.
90
Figure 4 shows the result of a recursive prediction per-
formed by a multilayer perceptron with p = 45, which is
slightly smaller than the lower bound of 50 defined above.
Except for this difference, the model has two hidden layers,
with 80 neurons in the first one and 55 neurons in the second
one, and a single linear output neuron as before. Moreover,
the model is trained with the same data set and of the same
size used to obtain the result shown in Fig. 3, and the recursive
prediction procedure is used to test the model after complet-
ing the training session in exactly the same way as before.
There is a dramatic difference between the results shown in
I Figs. 3 and 4. In particular, when the size of the input layer
Od 50 100 150 200 250 of the multilayer perceptron model is not large enough, the
Time model fails to capture the underlying dynamics of sea clutter.
7. Recursive prediction of sea clutter, using a multilayer percep- Clearly, the choice of a neural network predictor that violates
tron with 50 source nodes, two hidden layers with 80 and 55 neu- the lower bound on the size of p is unacceptable.
rons, respectively, and one output neuron. The solid curve refers
To emphasize the need for a nonlinear predictive model,
to the original sea clutter wavefomz. The dashed c u w e refers to
we show Fig. 5, the recursive prediction results obtained
the recursive predicted waveform, for which thefirst 50points of
the sea clutter set (not shown in the figure) are used as the initial
using (a) an autoregressive (AR) model, and (b) a multilayer
starting point. perceptron model. Both models used 50 delay taps for the
input, in accordance with the lower bound ofp 2 50. Clearly,
the AR model, which is linear, fails completely to capture the
This result confirms that sea clutter produced by a nonco- underlying dynamics of sea clutter.
herent radar (i.e., one that relies on amplitude information Turning next to the radar application of the predictive
alone) is locally predictable. Moreover, the horizon of pre- modelling of sea clutter, Fig. 6 shows the results of another
dictability, namely, 50, is approximately equal to the inverse experiment involving an off-the-shelf commercial noncoher-
of the largest Liaponuv exponent, 0.03. (For an accurate ent marine radar operating in a scanning mode [94]. In this
calculation of the horizon of predictability, see [66].)The fact application, the multilayer perceptron, trained on examples
that a neural network with an input layer of the right size can drawn from sea clutter and then having its synaptic weights
be trained to learn the underlying nonlinear dynamics of sea fixed, acts as a clutter (interference) canceller. In particular.
clutter is further testimony for the important observation through training, the network acquires the function of a

32 IEEE SIGNAL PROCESSING MAGAZINE MARCH 1996


100 clutter model. When it is fed with a received signal that
- Natural
90 --- MLP consists of a target signal plus clutter, as in Fig. 6a, the
presence of the target signal causes a corresponding pertur-
bation in the output of the network. That is, the network
suppresses the clutter component and thereby enhances the
presence of the target signal at its output, as illustrated in Fig.
6b.
Figure 7 demonstrates another interesting property of the
clutter canceller (nonlinear predictive model), using data
collected with the same noncoherent marine radar employed
for Fig. 6. The so-called B-scan (azimuth versus range)
images shown in the two parts of Fig. 7 represent (a) the
output of a conventional constant false-alarm rate (CFAR)
0 50 100 150 200 250 processor, and (b) the output of the clutter canceller. In both
Time cases, the images show the respective processor outputs prior
'. A comparison of the AR model and the multilayerperceptron to the application of a detection threshold. The input radar
(MLP) model for reconstruction of the underlying dynamics of data set contains two closely spaced targets. While the echoes
sea clutter.
from these two targets are blurred together in the conven-

12/
10
6

-4I '
0 50 100 150 200 250 300 350 400 450 E
Samples in the azimuthal direction 5 10 15 20 25 30 35 40 45 50
(4 (a)
1.6

1.4
5
1.2
10
m 1
3
-
-
50.8 15

' 0.6

0.4
25
0.2
30
0
35
-0.2b '
50 100 150 200 250 300 350 400 450 !
Samples in the azimuthal direction IO
(b) 5 10 15 20 25 30 35 40 45 50
(b)
(a)Azimuthal time series containing sea clutter and target. (b)
Prediction error at output of neural network. Target is clearly evi- . B-Scan images: ( a ) Output of conventional CFAR processor.
dent. ( b ) Output of neural network-based clutter canceller.

MARCH 1996 IEEE SIGNAL PROCESSING MAGAZINE 33


tional CFAR processor image, they are clearly separable in
the neural-network processor image.
Time-frequency Pattern
analysis Feature extraction
classification

Case Study II: Modular Learning Strategy for


Signal Detection in a Nonstationary Environment 8. Functional diagram of the receiver.

In the first case study, we showed how aneural network, used terms so that they bounce up and down less, when conipared
as a nonlinear predictive model, can exploit prior knowledge, io the standard form of the WVD. In so doing, the RID
namely, the fact that sea clutter is chaotic. The only informa- provides an “almost” positive distribution, which is what a
tion used in that case study was the information contained in time-frequency energy distribution should be, particularly for
the amplitude of the received signal, which is provided by a those applications that require the analysis of signals. For
noncoherent radar. This second case study pertains to the signal detection, it would be tempting to do the opposite (i.e.,
detection of a weak target signal corrupted by an interfering retain the cross Wigner-Ville distributions and suppress the
signal. Here, one or the other or both of these signals may be auto-terms). The detection strategy would then focus solely
nonstationary, and no prior knowledge about the environ- on the presence or absence of the cross Wigner-Ville distri-
ment is invoked. However, these problems are ameliorated butions. Such a procedure would, however, violate the infor-
through the use of Doppler information in addition to ampli- mation preservation rule by removing useful information
tude information, which requires the use of a coherent radar. contained in the auto-terms; its use is therefore not recom-
Case Studies I and I1 do have one thing in common: in both mended for signal detection.
cases, the radar operates in an ocean environment, with sea
In a clutter-dominated environment, which is the environ-
clutter being the primary source of interference.
ment of interest in Case Study 2, the cross-terms arise only
Now consider a novel modular learning strategy for signal
when a target signal is present. Thus, the presence of such
detection that is motivated by the echo-location (sonar) of a
terms is in fact an asset. We say this because the terms provide
bat, which detects, pursues, and captures its target (e.g., an
another feature that can enhance the visibility of a target in
insect) with a facility and success rate that is the envy of every
the time-frequency image resulting from the application of
radar or sonar engineer [95]. We are not suggesting that the
the WVD. Indeed, the cross-terms are essential to the optimal
modular detection strategy describe$ in this article involves
information-preserving property of the WVD. To appreciate
all the signal processing functions performed in the bat’s
the importance of the WVD for the radar detection problem,
echo-location system. What we are saying is that the principal
we present three sample WVD images of real-life radar
functions that characterize the modular learning strategy are
returns. These represent three different situations pertaining
found in one form or another in the bat’s echo-location
to an ice-infested ocean environment using a coherent radar
system.
[110-1111:
Figure 8 shows a block diagram of the basic detection
strategy consisting of three fundamental functional blocks
that are designed to perform time-frequency analysis, feature 0 Strong radar retum from a large ice target, shown in Fig.
extraction, and pattern classification, in that order. This form 9a.
of front-end processing is commonly used in pattern recog-
nition tasks [106]. 0 Relatively weak radar return from a small ice target, shown
For the time-frequency analysis, we have chosen the in Fig. 9b.
Wigner-Ville distribution (WVD); Table 3 presents a sum-
mary of the important properties of the WVD. Among the
family of bilinear time-frequency distributions, the WVD 0 Sea clutter alone, shown in Fig. 9c.
possesses two distinct advantages over other members of the
family for signal detection [ 1071: The WVD images presented in this figure significantly dif-
ferentiate between these three scenarios. Unfortunately, the
1. It is always a real-valued function. use of the WVD leads to a significant increase in the amount
2. It exhibits the least amount of spread in the time-frequency of redundant information contained in the time-frequency
plane. image of a radar signal. To improve computational effi-
ciency, it is therefore necessary to follow up the WVD with
One criticism that is often made against the WVD is the some form of data compression. (This point is also made in
generation of cross-terms, or more precisely, cross Wigner- [112, 1131, where singular value decomposition is used for
Ville distributions, due to the combined presence of two (or the extraction of features from the WVD image of a signal
more) components in the received signal. Various procedures for the purpose of signal detection or [Link],
have been developed in the literature for dealing with the the scheme described therein is primitive compared to the
cross Wigner-Ville distributions. In [ 108,1091,for example, modular learning strategy embodied in Fig. 10, as it lacks a
an algorithm known as the reduced interference distributions learning capability and does not address the two fundamental
(RID) is described, which is designed to flatten the cross- questions raised later in this case study.)

34 IEEE SIGNAL PROCESSING MAGAZINE MARCH 1996


A signal processing tool that is well suited for data com- features by the PCA is minimized and quantifiable, and so
pression is principal components analysis (PCA) [ 1061. Ba- we are still operating in the realm of the information preser-
sically, the PCA performs an eigendecompositionon a square vation rule. It is conceivablethat for our applicationnonlinear
matrix (in our case, the covariance matrix of the time series devices known as principal curves/surfaces[ 1141in statistics
obtained by scanning the WVD image of the incoming radar can do better than PCA. A similar capability is provided by
signal on a column-by-columnbasis, with each column rep- Kohonen’s self-organizing feature map [115], which is of a
resenting a time slice), orders the eigenvalues in descending low dimension and used to approximatea higher-dimensional
order, and retains the eigenvectorsassociated with the largest scatter-plot of samples. For a comparison of statistical and
[Link] compressedsignal is representedby a linear SOFM approaches, see Mulier and Cherkassky [116].
combination of the eigenvectors retained by the PCA. Thus, The final operation in our modular learning detection
the PCA is instrumental in extracting a finite set of features strategy is that of pattern classification, the purpose of which
for the WVD image that is optimum (among linear tech- is to distinguish between two different time-frequency im-
niques), in that the original WVD image (and therefore the ages on the basis of features extractedby the PCA. One image
original received signal) can be reconstructed from these pertains to the presence of clutter alone. The other image
features in a minimum mean-square error sense. In other pertains to the combined presence of clutter and the target
words, information loss brought on by the extraction of signal of interest. The pattern classification process is typi-

MARCH 1996 IEEE SIGNAL PROCESSING MAGAZINE 35


Target signal is present (hypothesis H,)

Classification

Target
channel

1
channel

4.-
0.00 005
A...
0.10 015 0.20 025
Principalmmponents
analyser 1, matched
to the WVD of
interference
Feature
extraction
j Principaicomponents
analyser 1, matched
to the WVD of signal
plus interference

Wagner-Ville
distribution (WVD)
computed

Input Data (complex time-series)


-10 -5 0 5 10 0. Block diagram of the two-channel receiver using modular
learning strategy.

ibi
generalized Hebbian algorithm (GHA) due to Sanger [ 1171.
Table 4 presents a summary of this algorithm. The PCA
network in the clutter channel is trained by presenting it with
WVD images known to represent clutter only, under varying
400 environmentalconditions. Once the training is completed, the
synaptic weights of that PCA network are fixed. The training
200
procedure of the PCA network for the target channel follows
0 a similar procedure, except for the fact that its training
examples consist of WVD images known to contain target
200
plus clutter, under varying conditions. The outputs of the
400 PCA networks may be viewed as a specific number of domi-
-10 -5 0 5 10 nant projections of the input WVD space on two subspaces,
with one subspace representing clutter alone and the other
subspace representing target plus clutter. Typically, these two
subspaces are unknown and nonlinear; projections of the
WVD space onto them are therefore best learned by way of
000 005 010 015 020 025 real-life examples that are representative of the two scenar-
ios. The end result is that the PCA network in one channel is
. (a) WVDfor a clearly visible growler; (b) WVDfor a barely adaptively matched to clutter alone, and the PCA network in
visible growler; (c) WVDfor sea clutter. the other channel is adaptively matched to target plus clutter,
hence the designations of the two channels in Fig. 10 as
cally nonlinear, making the task that much more challenging clutter (interference) and target channels, respectively.
to implement. Each multilayer perceptron has two hidden layers and an
Figure 10 shows a block diagram of a neural network- output layer with three output nodes. The output nodes are
based implementation of the modular detection strategy de- linearly combined into a single decision-making node. Thus,
scribed herein [110-1111. It consists of two channels, one the decision as to whether a target is present or not is deferred
termed the clutter or interference channel, and the other to the very output of the system, in accordance with the
termed the target channel. Both channels are fed from a information preservation rule. Specifically, if a threshold set
common input representing the WVD image of the received for a prescribed probability of false alarm is exceeded by the
signal. Each channel consists of a PCA network followed by overall output of the receiver, a decision is made that a target
a multilayer perceptron for pattern classification. The PCA is present; otherwise, a decision is made that the received
networks are trained in a self-organized fashion, using the radar signal consists of clutter alone. The synaptic weights of

36 IEEE SIGNAL PROCESSING MAGAZINE MARCH 1996


the two hidden layers and output layers of both perceptrons tions and possibly improve the generalization performance
and those of the linear combiner are trained simultaneously of the multilayer perceptrons.
in a supervised manner by presenting the whole network with There are two basic questions that need to be addressed in
WVD images that are known to represent the following the context of the modular learning strategy for signal detec-
situations that can arise: tion described in Fig. 10. First, what is the rationale for using
two channels? To answer this question, we first note that in
Clutter alone the traditional approach to radar target detection in a clutter-
dominated environment, for example, we may use a “best”
Strong target return plus clutter mismatched filter for clutter discrimination [ 1211. In such an
approach involving a single channel in the receiver, the
Barely visible target return plus clutter requirement for best performance in additive noise is traded
for an improvement in performance in clutter by purposely
The training of these layers is performed using the back- mismatching the filter. We may avoid the need for this
propagation algorithm. As a matter of interest, the first hidden trade-off in performance by using two nonlinear matched
layer of each multilayer perceptron uses the notions of recep- filters as depicted in Fig. 10, with each filter being adaptively
tive fields and weight sharing described in [12, 1201. By matched to the received signal arising under one of the two
“receptive field,” we mean that each neuron in the first hidden
hypotheses that are to be distinguished. In addition, the use
layer is connected only to a finite set of neurons that lie in its
of two different channels as described herein provides two
local neighborhood in the input layer. By “weight sharing,”
independent assessments of the decision that should be taken,
we mean that all the receptive fields of the layer share the
given the received signal. A simple and yet effective method
same set of weights. The use of receptive fields and weight
of integrating the two channel outputs is through the use of
sharing is designed to reduce the number of synaptic connec-
linear combining [122], which is precisely what has been
done in designing the modular learning strategy of Fig. 10.

MARCH 1996 IEEE SIGNAL PROCESSING MAGAZINE 37


This strategy therefore makes it possible to learn how to
Dotmler CFAR
classify the two sets of features that have been learned by the
PCA networks in the target and clutter channels. An altema- 250
tive way to state this is that it exhibits a “learning to learn”
capability.
The second question is: why does each channel have three
output nodes? The nonlinear input-output mapping produced 200
by the modular learning strategy of Fig. 10 depends, among h
0
other things, on the number of output nodes per channel. The a,
use of two output nodes per channel severely limits the !i
capacity of each multilayer perceptron to classify the re- CD
150
ceived signal. On the other hand, the use of three output c
0
nodes per channel makes it possible to provide a finer classi- v)
Q
fication of the received signal by saying c
a,
v)

100
0 the received signal contains a strong target signal

e the received signal contains a weak target signal, or


50
0 the received signal contains clutter alone

This, in turn, has the beneficial effect of reducing the overlap


between the two primary classes of interest: target is not
present (null hypothesis), and target is present (the other n
-
0 10 20 30 40
hypothesis). Consequently, the receiver with three output
nodes per channel has the potential of outperforming the range gate (5m apart)
receiver with two output nodes per channel. Indeed, experi- (ai
mental results presented in [ 110, 1113 bear out the validity of NN based detector
this statement.
Figure 11 presents a comparison of the detection results 250
obtained for the modular receiver of Fig. 10 with those of a
conventional Doppler CFAR (constant false alarm rate) proc-
essor for a false alarm rate set at Here, black denotes the
presence of a target signal, and white denotes clutter. With 200
0
the target constantly being in the range of the radar (i.e., for (U

all time), we should ideally see a continuous black strip


(representing the target) in a light background (representing
c
cc)
Ln
the clutter). In light of this observation, a significant improve- @.I
3 150
ment in the modular learning strategy is found by filling in
the periods of “silence” that are observed in the detection
performance of the conventional Doppler CFAR processor.
This “silence” is caused by the partial obscuration of the
target (a small piece of ice in the experiment described
herein) by an ocean wave in front of it, or the dipping of the
target in a wave trough. The performance displayed in Fig.
11 is quite remarkable, since the modular learning system is
able to perform satisfactorily even in a situation when the
target returns are weak; in other words, a barely visible target
has been made “visible in signal processing terms”. The other
observation made from Fig. 11 is the occasional blanking of
a signal from the target (as seen in the middle of the plot); in
such cases, there is no way any method would be able to 0 10 20 30 40
detect the target since, insofar as the radar is concerned, the range gate (5m apart)
target is simply not there to be seen. In [1111, experimental (b)
results are presented demonstrating that the modular learning 11. Postdetection results f o r two different schemes:(a) COG-
system of Fig. 10 has a robust performance with respect to tional Doppler CFAR processor; (b)modular learning scheme of
wide variations in sea state. Fig. 10.

38 IEEE SIGNAL PROCESSING MAGAZINE MARCH 1996


In summation, it may be difficult, in the traditional ap-
proach to radar receiver design, to make provisions for real-
life situations in a manner similar to that attainable with the
modular learning strategy described in Fig. 10. Unfortu-
nately, however, the highly complex composition of this
structure defies a detailed mathematical analysis. Moreover,
to design it, one would need a sufficiently large training set
that is truly representative of the operational environment. Principal Components Vector Quantization
Once the training process is completed and all the synaptic
weights of the PCA networks and multilayer perceptron
classifiers are computed and thereafter fixed, and the receiver
is ready for normal operation, signal propagation through the
network is very rapid. This is basically due to the fact that
both the PCA networks and the multilayer perceptron classi-
fiers consist of feedforward structures.

Case Study Ill: Mixture of Principal Components


for Image Compression and Segmentation
Mixture of PCs
Our third and final case study pertains to the use of neural 2. A spectrum of representations in two dimensions.
networks for image compression and segmentation. The
study of image compressionmethods has been an active area described. Specifically, it combines desirable attributes of
of research since the inception of digital imaging. Successful both principal components analysis (PCA) [lo61 and vector
image compression schemes must satisfy two conflicting quantization (VQ) [125]. Within a class, an input vector is
requirements: represented by a continuous, linear combination of M basis
vectors of the subspace in a manner analogous to the PCA
During the coding phase of image compression, data are representation. But, because of the partitioning of the data
transformed from their native format, typically an array of into a discrete number of regions or classes, the MPC effects
gray level or trichromatic pixels, into a new format that a nonlinear mapping of the input data in a manner analogous
requires less bandwidth or storage. to VQ. The relations between these three methods of repre-
sentation are illustrated in Fig. 12 for a two-dimensional
The transformationmust preserve the essential information example:
content of the original image, so that the difference be-
tween the original and decoded images is not perceptually 1. The PCA approach forms a complete, continuous rep-
discernible. The significance of this difference must be resentation of input data using, in this example, a linear
clearly evaluated within the context of the end use of the combination of two basis vectors, as indicated in Fig. 12(a).
image. For example, medical images must not lose their
diagnostic value after compression.
2. With VQ, the input data are represented in a purely
discrete manner by partitioning the input space, in this exam-
A major problem with many image processing applications,
ple, into 10 distinct regions and representing each region by
is their implicit assumptionof stationarity. The fallacy of this,
a Voronoi center, as indicated in Fig. 12(b).
assumption is the reason why many conventional image
processingtechniquesperform poorly in the vicinity of edges
Here, the image statistics tend to be radically different from 3. The MPC lies between these two extremes of data
the global statistics of the image. Conventional image com- representation, as indicated in Fig. 12(c):
pression methods, such as the Karhunen-Lobve transform
(KLT) [106], are designed according to a globally optimal1 In a manner similar to the VQ, the input space is partitioned,
mean-square error criterion. However, the aforementioned in this example, into 4 distinct regions.
nonstationarityof edge regions makes this criterion far from
ideal. Therefore we may say that if an image compression Within each region, the input data are represented, in this
method can be made to adapt to local nonstationaritiesin the example, by a single basis vector. Thus, like PCA, the data
image, then its performance would be superior to that of the: are given a continuous representation.
KLT.
To account for variationsin the local statistics of an image, For higher-dimensional input spaces, the number of basis
a transformationmust have the capability to adapt locally. In vectors used in MPC may be two or more, in which case we
[ 123-1241, a new family of adaptive transform coding meth- find that planes, hyperplanes, or higher-dimensional sub-
ods, called a mixture of principal components (MPC), is spaces are formed within the input space.

MARCH 1996 IEEE SIGNAL PROCESSING MAGAZINE 39


organized, combining both Hebbian learning and competitive
learning; it produces an adaptive linear transformation that
r”‘ tries to minimize the mean-squared error between the input
data and the decoded data, beyond that attainable with the
KLT. As such, the learning algorithm is well suited to the task
of image compression. A summary of OIAL is presented in
Table 5.
Figure 14a shows the magnetic resonance image (MRI)

1w, ikM
+x
!
] .Li
..................................
rk7
.
;
used for training. The image in Fig. 14b is the adjacent section
from the same study (patient), which was used for testing.
U
Each image consists of 256 x 256 pixels, with the dynamic
IQM 1 : ......................................... !
range of 8 bits or 256 gray levels. The training image was
divided into blocks of 8 x 8 pixels for an input dimension of
I
N = 64. The blocks were overlapped at two pixel intervals for
’[Link] structure of OIAL scheme. a total number of 15,625 training samples. During training,
the samples were presented in random order. For comparison,
Figure 13 shows a network structure for implementing one the KLT was calculated based on the same training data.
The test image was divided into 8 x 8 non-overlapping
particular form of the MPC. The system is modular, consist-
blocks. These blocks were transformed by the previously
ing of a number of modules corresponding to different classes
computed system into a set of coefficients, quantized, and
of input data. Each module consists of a linear transformation, then transformed back into image blocks. The coefficients
whose basis vectors are computed using an initial training were quantized in a similar manner to that of the JPEG
period. The appropriate class for a given input vector is standard. The first coefficient was coded via first-order
determined by the subspace classifier. The system utilizes a DPCM using a uniform quantizer. The remaining coefficients
learning algorithm referred to as the optimallv integrated were coded via PCM using a uniform quantizer. For a given
adaptive learning (OIAL) algorithm. The algorithm is self- coding rate, the same quantization interval was used for all

40 IEEE SIGNAL PROCESSING MAGAZINE MARCH 1996


of these two images, it is clear that the OIAL image preserves
more features than the KLT image. In the upper forehead
region near the skull, the dark line of the outer table of the
skull between the white line of the skin and the white line of
the diploe (i.e., hard porous tissue between the walls of the
cranial bones) is visible in the OIAL, but completely ob-
scured in the U T . The same is true of the detail in the top
portion of the orbit. Not only does the KLT lose information,
it also introduces texture variations that are not present in the
original image nor in the OIAL image. This texture interferes
with the visibility of the sulci in the outer portion of the brain.

the coefficients. The quantized data were then Huffman-


coded with a codebook optimized for Laplacian distributions.
The number of bits assigned for the class information was
simply log& bits per block. Different bit rates were used in
the quantization step. For the KLT, an identical coding
scheme was used except, of course, that no additionalbits per
block were required to code the class assignment.
For the new approach, the overall mean-squared error was
reduced and the perceptual image quality was improvedl,
when compared to the KLT. More image details were pre-
served and fewer artifacts were [Link] particular, Fig.
15a shows the details of the new coding scheme with 1218
classes, 4 coefficientsper block, at 0.25 bpp, while Figs. 151b (b)
shows the corresponding details of KLT coding at the same 5. ( a ) Reconstruction of test image using OIAL; (b) reconstruc-
bit rate of 0.25 bpp. When examining the detailed structure tion of test image using KLT.

MARCH 1996 IEEE SIGNAL PRIOCESSING MAGAZINE 41


of the OWL algorithm [123], two useful properties are
achieved:

0 The OIAL algorithm acquires the ability to perform image


segmentation.

* The problem of choosing the initial set of transformation


matrices is removed from the algorithm.

The result of this integration is a topological ordering of


classes, during training with like classes being close together

16. Original Lena image.

The image compression results just presented are a good


indication that the OIAL generalizes within the pertinent
class of image. While the “within class condition” may seem
restrictive at first, in practice this need not be so. Moreover,
while we do not claim that there exists a single network
configuration that would perform as well as a general-pur-
pose image compression scheme across a wide variety of
images, it is interesting, nevertheless, to see how well a
system trained on one class of images generalizes outside that
image class. Figure 16 shows the Lena image that is obvi-
ously quite different from the image used for training, as
shown in Fig. 14a. Figure 17a shows theresulting compressed
image using the same network (4 coefficients and 128
classes) and bit rate (0.25 bpp) as that used for the image
shown in Fig. 14b. The mean-squared error from this image
was 54.9, referred to the original image. For comparison, the
image was compressed using the KLT of itself and quantized
to the same number of bits (4 coefficients with 0.25 bpp). The
resulting image is shown in Fig. 17b, and has a mean-squared
error of 71.0, also referred to the original image. These two
images clearly show that the OIAL system trained on a
magnetic resonance image of a head performs better than the
KLT optimized for the specific image being coded.
For many applications, it may be advantageous to have
similarity between adjacent classes. The self-organizing fea-
ture map (SOFM) introduced by Kohonen [115] makes for
such a provision in a simple and yet effective fashion. (The
use of the SOFM algorithm as the basis of subspace classifiers
is also discussed by Kohonen [ 1261; the resulting structure is
referred to as the “adaptive subspace self-organizing algo-
rithm”, which is quite different from the integration of the
OIAL and SOFM algorithms.)
During training, each training vector is used not only to
update the winning class, but also classes that are adjacent to 17. (a) Reconstruction of Lena image using OIAL; (b)reconstruc-
it. By integrating the SOFM algorithm into the composition tion of Lenna image using KLT.
I
42 IEEE SIGNAL PROCESSING MAGAZINE MARCH 1996
each image were shown simultaneously to each radiologist,
who was asked to rate each one on a five-point scale for image
quality and visibility of pathology. Only for the 40: 1 versions
were there any unacceptable ratings of the McMEC com-
pressed images; even then, these images received a substan-
tial number of top ratings. Many times, the radiologists
commented on how little difference there was, if any, be-
tween the images. Occasionally, a radiologist would pick the
40:l compressed image as the best. When compared to the
KLT, the McMEC versions ranked better than or as good as
the KLT versions 17 times out of 18. In addition, four out of
nine times the 40:1 McMEC version was ranked as good or
better than the 30: 1 KLT version, which is quite remarkable.
Currently, the performance of the MPC approach of image
compression is being explored for synthetic aperture radar
(SAR) images. In preliminary results, the OIAL approach has
reduced the reconstruction error by 3 dB over the KLT for
the same compression ratio. This margin of improvement
appears to be valid up to compression ratios of 50:l. One
explanation for this marked reduction in distortion is the fact
18. OIAL segmentation map of test image with 32 classes, two C O . that SAR image formation process is highly nonlinear and
efficients per class. Color indicates class membership; intensity is includes a high degree of “speckle” noise. As a result, the
weighted by the magnitude of the second coeficient for each nonlinear nature of the MPC approach matches the signal
block. characteristics better than the linear KLT.

in a manner analogous to the ordering of directionally sensj- Information-Theoretic Models for


tive columns in the visual cortex [127]. This is illustrated in Unsupervised Learning
Fig. 18, where each basis block acts as a feature detector. The
features corresponding to the basis vectors are either lines or In the previous section we discussed three different signal
edges of a specific orientation. When comparing adjacent processing applications of neural networks that require the
classes, the angles of the features are similar. Moreover, the use of unsupervised learning, exemplified by the nonlinear
angles change in a somewhat regular manner as the class predictive model in Case Study I, and linear PCA networks
number progresses. One other important property resulting in Case Studies I1 and [Link] this section, we discuss another
from the combined use of OIAL and SOFM algorithms is that powerful approach to unsupervised learning, which is rooted
the segmentation is independent of variations in illumination, in information theory. The approach builds on the so-called
as it is natural in the human visual system [ [Link] property principle of maximum information preservation, also re-
may prove to be of significant practical value in image ferred to as Infomax for short, which is due to Linsker [ 12,
analysis (e.g., computer aided tomography preprocessing, [Link] may be stated as follows:
and objecthackground discrimination).
The OIAL algorithm summarized in Table 5 is one way “The transformation of a vector x in the input layer of a
of implementing the MPC method. In [124], another algo- neural network to a vector y in the output layer of the network
rithm, called the multi-class maximum entropy coder should be so chosen that activities of the neurons in the output
(McMEC), is described for implementing the MPC method. layer jointly maximize information about the activities in the
The McMEC algorithm uses only one basis vector per mod- input layer. The parameter to be maximized is the mutual
ule, while the OIAL algorithm has M basis vectors. As a information between the input vector x and the output vector
consequence, the McMEC algorithm is required to use a y in the presence of processing noise.”
much larger number of modules than the OIAL algorithm.
One of the most demanding application areas of image This principle may be viewed as the neural network coun-
compression is compressing medical images, where the ini- terpart to the concept of channel capacity, which defines the
plications of any sort of distortion are grave indeed. Dony, let Shannon limit on the rate of information transmission
al., [129] have investigated the application of the McMEC through a communication channel.
algorithm to the compression of clinical chest radiographs Becker and Hinton [12, 134-1371have extended the idea
acquired digitally using the Fuji computed radiography (CR) of maximizing mutual information to unsupervised process-
system. Comparative evaluations with the KLT were also ing of the image of a natural scene. Specifically, for a given
included in the study. Four degrees of compression were image, the mutual information between the outputs of two
used: 10:1,20: 1,30:1, and 40: 1. Seven radiologists evaluated neural network modules is maximized, with adjacent and
the images. The original and four compressed versions of nonoverlapping patches of the image providing the inputs.

MARCH 1996 IEEE SIGNAL PROCESSING MAGAZINE 43


This latter principle is referred to in the literature as Z m a . A suited for image processing with emphasis on the discovery
polarimetric radar application involving navigation along a of properties of a noisy sensory input that exhibit some form
confined waterway, which builds on a variant of Zmax, is of coherence across space and time.
described in [137-1381. Bell and Sejnowski [139, 1401 have built on these infor-
Both of these unsupervised learning procedures, Infomax mation-theoretic models for unsupervised learning by devel-
and Zmax, rely on the use of noisy models for their operation, oping their own algorithms to tackle the difficult signal
which makes their application all the more realistic. Infomax processing problems of blind signal separation and blind
is well suited for the development of unsupervised learning deconvolution. The paper by Herault and Jutten [141] is the
models and feature maps. Zmax, on the other hand, is well first neural net paper on blind signal separation using Heb-
I 44 IEEE SIGNAL PROCESSING MAGAZINE MARCH 1996
I
bian learning. For a list of references on blind signal separa- presents a summary of the underlying principle in the Bell-
tion, see Comon [142]. Note, however, toward the end of‘ Sejnowski procedure for their blind signal separation and
1995, close to 50 papers have been published on this subjecl blind decanvolution algorithms. In [139-1401, experimental
since Comon’s paper. Also, for a review of traditional signal results are presented that demonstrate the capability of this
processing methods used in blind deconvolution, see [ 143- new approach for solving blind signal separation and blind
1441. It appears that the first application of Infomax to the: deconvolution problems. Although these demonstrationsare
blind deconvolution problem in the context of blind equali-. indeed impressive, the unsupervised learning algorithms de-
zation of a communication channel was described in [145]. veloped by Bell and Sejnowski in their present forms are of
A typical blind signal separation may be represented by a limited use in certain fundamental respects, as summarized
set of sources corresponding to a number of people engaged here:
in a conversationwith music in the background. The signals,
si(& s2(t),...,s d t ) produced by these different sources are The neural network models considered are of the single-
mixed together by an N-by-N matrix, A. The sources of these layer type, with the result that the optimal mappings dis-
signals and the mixing matrix A are all unknown. All that is covered by the algorithms are constrained to be linear; the
available for processing is a corresponding set of N received use of multilayer models may lead to the development of
signals x i ( t ) , x2(t), ...&fit), which are linear superpositionsor more powerful input-output mappings.
the original signals si@), s2(t),...,sfit). The problem is to
reconstruct these original signals by finding a separating For the blind signal separation problem, for N inputs it is
matrix, W, that is a permutation and rescaling of the inverse assumedthat there are an equal number of outputs available
of the unknown matrix, A. The problem described herein i!; for processing. There is no corresponding theory for the
sometimes referred to as the “cocktail-party problem.” more general case when the number of inputs is not equal
A similar and equally difficult problem is blind deconvo- to the number of outputs.
lution, where a source signal s(t) is operated on by a linear
filter of impulse response, h(t),to produce a received signal, In a realistic environment pertaining to the blind signal
x(t). The original source signal, s(t), and the impulse re- separation problem, there are unavoidable propagationde-
sponse,h(t),are both unknown. All that is given is a statistical lays associated with the individual signal paths before they
model of the source responsiblefor generatingthe signal, s(t). are mixed [Link], these propagationdelays are
Given the received signal, x(t), the problem is to reconstruct unknown. Some adaptive mechanism would therefore
the original signal, s(t), with little or no distortion. Areas of have to be incorporated into the blind signal separation
application of blind deconvolution include the following algorithm to take account of this practical issue.
[ 144,1451:
In the blind deconvolution [Link], it is assumed that the
1. Cancellation of reverberation due to the barrel effect original signal consists of statistically independent sym-
encountered in hands-free telephone operation. bols. Although this assumption can be justified in certain
situations (e.g., channel equalization), it would be useful
2. Seismic deconvolution, where the problem is compli- to relax the assumption of statistical independence.
cated by a lack of precise knowledge of the short-duratioin
pulse used for excitation. These four important research issues, the first three of which
are highlighted in [139], certainly deserve more attention.
3. Image restoration, where difficulties arise due to un- Perhaps the most important point to emphasizehere is that
known blurring effects caused by photographic and/or elec- the use of information theoretic models for unsupervised
tronic imperfections. learning, as in the work of Bell and Sejnowski, is a move
away from the mean-squareerror criterion that has permeated
so much of the traditional approach to the design of neural
4. Blind equalization of a communication channel where
networks.
it is not feasible to send a training sequence of long enough
duration.
Concluding Remarks
Going back to the important contribution by Bell and
Sejnowski [139, 1401, the essence of their approach may be Neural networks, often referred to as an emerging technol-
summarized as follows. In a neural network whose individual ogy, have grown very rapidly on many fronts during the past
neurons are characterizedby sigmoidal activation functions, 10 years. Their theory and design principles have benefited
maximization of the information transfer across the network enormously from contributions made by workers in many
tends to reduce the redundancy between the neurons in the diverse fields. As such, they represent a significant addition
output layer of the network. It is the latter property that to the “kit of tools” available to system designers. Their
enables the network to perform signal separation or decon- ability to learn in a supervised or unsupervised manner,
volution in an unsupervised manner. It is assumed that thie depending on the way in which they are applied, makes them
original signal consists of independent symbols. Table 6 well suited for solving difficult signal processing tasks. In

MARCH 1996 IEEE SIGNAL PROCESSING MAGAZINE 45


particular, they can naturally account for the commonly References
encountered properties of real-life data, including nonlinear-
ity, nonstationatity, and non-Gaussianity. 1. North, D.O., “An analysis of the factors which determine signal-noise
discrimination in pulsed carrier systems,” Proc. IEEE, vol. 51, pp. 1016-
Perhaps the biggest virtue of neural networks is that they 1027, 1963.
learn about their environment by way of examples, and 2. Van Vleck, J.H., and D. Middleton, “A theoretical comparison of visual,
thereby construct an input-output mapping that brings to aural, and meter reception of pulsed signals in the presence of noise,” J. Appl.
mind the notion of nonparametric statistical inference. In so Phys., Vol. 17, pp. 940-971, 1946
doing, the solutions that they compute may not be guaranteed 3. Wiener, N, Extrapolation, Intepolation, andSmoothing of Stationary Time
to be optimum, but they are usually found to be good engi- Series with Engineering Applications, Wiley, 1949. (This book was origi-
nally issued as a classified National Defense Research Council Report,
neering solutions. Most importantly, from a signal processing February, 1942.)
perspective, they have the potential (by themselves or in
4. McCullvch, W.S., and W. Pitts, “A logical calculus of the ideas immanent
combination with other technologies) to outperform their in nervous activity,” Bulletin Math. Biophys., vol. 5, pp. 115-133, 1943.
traditional counterparts, as demonstrated by the three case
5. Rosenblatt, F., ‘The perceptron: A probabilistic model for information
studies presented in this article. storage and organization in the brain,” Psych. Rev., vol. 65, pp. 386-408,
1958.

Acknowledgments 6. Widrow, B., and M.E. Hoff, Jr., “Adaptive switching circuits,” IRE
WESCON Convention Record, pp. 96-104,1960,
7. Hopfield, J.J., “Neural networks and physical systems with emergent
I am truly grateful to Dr. Vladimir Cherkassky, University of collective computational abilities,” Proc. Nat. Acad. Sci. USA, vol. 79, pp.
Minnesota, Minneapolis, and Dr. William J. Williams, Uni- 2554-2558,1982.
versity of Michigan, Ann Arbor, for their many critical inputs 8. Rumelhart, D.E., and J.L. McClelland, editors, Parallel Distributed Proc-
and useful suggestions that have impacted the finalizing of essing: Explorations in the Micvostructure of Cognition, vol. 1, Cambridge,
this article in a significant way. I also wish to thankDr. Henry MA: MIT Press. 1986.
D.I. Aberbanel, University of Califomia, San Diego, Dr. 9. Chiu, C.-T., K. Mehrota, C. K. Mohan, and S. Ranka, “Robustness of
feedforward neural networks.” In P. K. Simpson, editor, Neural Networks:
Anthony Bell, Salk Institute, California, Dr. Fay Boudreaux- Theory, Technology, andAppZications, pp. 348-353,1996, E E E Press, New
Bartels, University of Rhode Island, Dr. Bemie Mulgrew, York.
University of Edinburgh, Scotland, and Dr. Kari Torkkola, 10. Haykin, S . , Adaptive Filter Theory, Third Edition, Prentice-Hall, 1996.
Motorola, Utah, for their helpful inputs on selected parts of 11. Wldrow, B., and S. Steams, Adaptive Signal Processing, Prentice-Hall,
the article. I am indebted to my own colleagues, Dr. Sue 1985.
Becker, Department of Psychology, Dr. Nanda Kambhatla, 12. Haykin, S., Neural Networks: A Comprehensive Foundation, New York
Communications Research Laboratory, Dr. James P. Reilly, Macmillan College Publishing Company, 1994.
Department of Electrical and Computer Engineering, and Dr. 13. Lippmann, R.P., “An introduction to computing with neural nets,” IEEE
Patrick Yip, Department of Mathematics and Statistics, my ASSP Magazine, vol. 4, pp. 4-22,1987.
former research colleagues Dr. Tarun Bhattacharya, 14. Hush, D.R., and B.G. Home, “Progress in supervised neural networks:
Raytheon Canada, Dr. Robert D. Dony, Sir Wilfrid Laurier What’s new since Lippmann?,” IEEE Signal Processing Magazine, vol. 10,
University, and Dr. Graeme Jones, Raytheon Canada, and my pp. 8-39, 1993.
current graduate students Hugh Pasika and Paul Yee for 15. Rao, C.R., Linear Statistical Inference and its Applications, Second
reading the entire manuscript and offering numerous sugges- Edition, Wiley, New York. 1973.
tions for improving it. I wish to acknowledge the many [Link], A.R., and R.L. Barron, “Statistical learning networks: Aunifying
view,” In Proceedings of the 198%Symposium on the Inte~ace:Statistics
helpful inputs on chaos that I received on the Intemet from and Computing Science (Ed., E.J. Weaman), pp. 192-203,American Statis-
Dr. Martin Casdagli, University of Michigan, Ann Arbor, Dr. tical Association, Washington, DC, 1988.
Daniel Kaplan, McGill University, Quebec, Dr. Mathew 17. White, H., ‘‘Learningin artificial neural networks: A statistical perspec-
Kennel, Oak Ridge National Laboratory, Tennessee, Dr. Lou tive,” Neural Computation, vol. 1, pp. 425-464, 1989.
Pecora, Naval Research Laboratory, Washington, DC, Dr. 18. Ripley, B.D., “Neural networks and related methods for classification,”
Florin Takens, University of Gronigen, The Netherlands, Dr. J. Royal Statistical Society Series B, vol. 56, pp. 409-456, 1995.
James Theiler, Los Alamos National Laboratory, New Mex- 19. Ripley, B.D., “Flexible nonlinear methods for classification,” World
ico, Dr. Howell Tong, University of Kent, England, and Dr. Congress on Neural Networks, vol. 2, pp. 927-934, Washington, DC., 1995.
James Yorke, University of Maryland. Last, but by no means 20. Murtagh, F., “Neural networks and related massively parallel methods
least, 1wish to thank Dr. Jose Principe, University of Florida, for statistics: A short review,” Intemational Statistical Review, vol. 62, pp.
275-288, 1994.
Gainsville. Dr. Andrew Webb, Defence Research Agency,
Great Malvern, England, three anonymous reviewers and Dr. 21. Cheng, B., and D.M. Titterington, “Neural Networks: A review from a
statistical perspective,” Statistical Science, vol. 9, pp. 2-54, 1994.
Don Hush, Associate Editor of the SP Magazine for their
careful reviews of an earlier version of the article. 22. Cherkassky, V., J.H. Friedman, andH. Wechsler, editors, From Statistics
to Neural Networks: Theory and Pattern Recognition Applications, Sprin-
ger-Verlag, 1994.
Simon Haykin is Professor of Electrical and Computer 23. Helstrom, C.W., Statistical Theory ofsignal Detection, Second Edition,
Engineering, McMaster University, Hamilton, Ontario, Can- Pergamon Press, Oxford. 1968.
ada. 24. Poor, H.V., An Introduction into Signal Detection and Estimation,
Springer-Verlag, New York. 1988.

I 46 IEEE SIGNAL PROCESSING MAGAZINE MARCH 1996


25. Cherkassky, V., D. Gerhing, and F. Mulier, “Pragmatic comparison of 48. Lo, S.-[Link] al., “Wavelet-basedconvolutionneural network for pattem
statistical and neural network methods for function estimation,” World recognition,” World Congress on Neural Networks, vol. 2, pp. 838-843,
Congress on Neural Networks, vol. 2, pp. 917-926, Washington, DC., 1995. Washington, DC, 1995.
26. Kearns, M., “A bound on the error of cross validation using the 49. Szu, H., B.A. Telfer, J. Anandkumar, and M. Zaghlul, “Remote ECG
approximation and estimation rates, with consequences for the training-test diagnosis using wavelet transform and artificial neural networks,” World
split.” In M. C. Mozer, editor, Advances in Neural Information Processing Congress on Neural Networks, vol. 2, pp. 844-848, Washington, DC, 1995.
Systems, MIT, Press, 1996.
50. Qian, W., L. Li, B. Zheng, and L.P. Clarke, “Digital mamography:
27. Friedman,J.H., and W. Stuetzle,“Projectionpursuitregression,”[Link]. Wavelet-based mixed feature ANN for automatic detection of microcalcifi-
Statist. Soc., vol. 76, pp. 817-823, 1981. cations,” World Congress on Neural Networks, vol. 2, pp.849-853, Wash-
ington, DC, 1995.
28. Friedman, J.H., “Exploratory projection pursuit,” J. Amer. Statist. Soc.,
vol. 82, pp. 249-266, 1987. 51. Rao, S. S., and R. S. Pappu, “Hierarchicalwavelet neural networks.” In
P. K. Simpson, editor, Neural Networks: Theory, Technology, and Applica-
29. Wahba, G., “SplineModels for ObservationalData,” SIAM, Philadelphia, tions, pp. 83-90, IEEE Press, 1996.
1990.
52. Viterbi, A.J., “Wireless digital communication: A view based on three
30. Friedman, J.H., “Multivariate adaptive regression splines (with discus- lessons learned,”IEEE CommunicationsMagazine, vol. 29, No. 9, pp. 33-36,
sion),” Annals of Statistics, vol. 19, pp. 1-141, 1991. 1991.
[Link], J.H., “An overview of predictive learning and function ap- 53. Poggio T., and F. Girosi, “Networks for approximation and learning,”
proximation”. In V. Cherkassky, J.H. Friedman, and H. Wechsler (editors), Proc. IEEE, vol. 78, pp. 1481-1497, 1980.
From Statistics to Neural Networks, Springer-Verlag, 1994..
54. Bishop, C.M., “Curvature-driven smoothing in backpropagation neural
32. Cherkassky, V., Private communication, 1995. networks,” CLM-PILO, AEA Technology, Cullham Laboratory, Abingdon,
VIC.
33. Cybenko, G., “Approximation by superpositions of a sigmoidal func-
tion,” Math. Control, Signals, and Systems, vol. 2, pp. 303-314, 1989. 55. Sackinger, C., et al., “Application of the ANNA neural network chip to
high-speed character recognition,”IEEE Trans. Neural Networks, vol. 3, pp.
34. Homik, K., M. Stinchcombe, and H. White, “Multilayer feedforward 498-505,1992,
networks are universal approximators,”Neural Networks, vol. 2, pp. 359-
366,1989. 56. Hammerst”, D., “A VLSI architecture for high-performance, low-
cost, on-chip learning,” IJCNN, vol. 2, pp. 537-544, San Diego, CA., 1990.
35. Park, J., and I.W. Sandberg,“Universal approximationusing radial-basis
function networks,” Neural Computation, vol. 3, pp. 246-257, 1991. 57. Soderstrom, T., and B. Svensson, “Using and designing massively
parallel computersfor artificialneural networks,” J. Parallel and Distributed
36. Barron, A.R., “Universal approximation bounds for superposition of Computing, vol. 14, pp.260-285, 1992.
sigmoid functions,” IEEE Trans. Information Theory, vol. 39, pp. 930-945,
1993. 58. Hinton, G.E., and T.J. Sejnowski,“Learning and relearningin Boltzmann
machines”. In Parallel Distributed Processing edited by D.E. Rumelhart and
37. Kurkova, V., P.C. Kainen, and V. Kreinovich, “Dimension-independent J.L. McClelland, MIT Press, 1986.
rates of approximation by neural networks and variation with respect to
half-spaces”. In World Congress on Neural Networks, vol. I, pp. 54-57, 59. Wejchert, J., and G. Tesauro, “Visualizingprocesses in neural networks,”
Washington, DC, 1995. IBM J. Res. and Dev., vol. 35, pp. 244-253, 1991.

38. Vapnik, V.N., and A.Y. Chervonenkis, “On the uniform convergence of 60. Goldstein, H., “Sea echo in propagation of short radio waves.” In D.E.
relative frequencies of events to their probabilities,” Theoretical Probability Kerr, editor, MIT Radiation Laboratory Series, vol. 13, Section 6.6,
and its Applications, vol. 17, pp. 264280, 1971. McGraw-Hill, New York, 1951.
61. Jakeman,E., andP.N. Pusey, “A model for non-Rayleigh sea echo,” IEEE
39. Vapnik, V.N., The Nature of Statistical Learning Theory, Springer-Vler-
lag, 1995. Transactions on Antennas and Propagation, vol. AP-24, pp.806-814,1976.
62. Skolnik, M. editor. Radar Handbook, Second Edition, McGraw-Hill,
40. Natrajan, B.K., Machine Leanzing: A Theoretical Approach, Morgan
New York, 19...
Kaufmann, 1991.
63. Ward, K.D., C.J. Baker, and S. Watts, “Maritime surveillanceradar, Part
41. Hlawatsch, F., and G.F. Boudreaux-Bartels,“Linear and quadratic tinie- I Radar scattering from the ocean surface,” IEE Proceedings, vol. 137, Pt.
frequency signal representations,”IEEE Signal Processing Magazine, vol. F,pp. 51-62, 1990.
9, pp. 21-61, 1992.
64. Leung, H., and S. Haykin, “Is there a radar clutter attractor?” Applied
42. Rioul, O., and M. Vetterli, “Wavelets and signal processing,” [Link] Physics Letters, vol. 56, pp. 592-595, 1990.
Signal Processing Magazine, vol. 8. No. 10, 1991.
65. Li, X., and S. Haykin, “Chaotic characterizationof sea clutter,” l’Onide
43. Brotherton,T., T. Pollard, andD. Jones, “Applicationsof time-frequency Electrique, Special Issue on Radar, SEE, France, pp. 60-65, March 1994.
and time-scale representations to fault detection and classification,” Pro-
ceedings of the IEEE-SP International Symposium on Time-frequency and 66. Haykin; S., and X. Li. “Detection of signals in chaos,” Proceedings of
Time-scale Analysis, pp. 95-97, Victoria, BC. 1992. the IEEE, vol. 83, pp. 95-122. 1995.
44. Pati, Y.C., and P.S. Krishnaprasad, “Analysis and synthesis of feedfor- 67. Schuster,H.G., Deterministic Chaos: An Introduction, VCH, Weinheim,
ward neural networks using discrete affine wavelet transformations,”IEEE Germany, 1980.
Transactions on Neural Networks, vol. 4, pp. 73-85, 1993.
68. Ruelle, D., Chaotic Evolution and Strange Attractors, Cambridge Uni-
45. Manjunath, B.S., and R. Chelleppa, “A unified approach to boundary versity Press, 1989.
perception: Edges, textures, and illusory contours,” IEEE Transactions on 69. Ott, E., Chaos in Dynamical Systems, CambridgeUniversity Press, 1993.
Neural Networks, vol. 4, pp. 96-108, 1993.
70. Newhouse, S., “Understanding chaotic dynamics”. In J. Chandra, editor,
46. Kreinovich, V., 0. Sirisaengtaksin, and S. Cabrera, “Wavelet neural Chaos in Nonlinear Dynamical Systems, SIAM, 1984.
networks are asymptotically optimal approximators for functions of one
variable,” Proceedings of IEEE International Conference on Neural Net- 71. Wolf, A., J.B. Swift, H.L. Swinney, and J.A. Vastano, “Determining
works, vol. 1, pp. 299-304, Orlando, Florida, 1994. Liapunov exponents from a time series,” Physica 16D, pp.285-317,1985.
47. Szu, H., J. Garcia, and J. DeWitte, “Telemedicine:Recognition from 2D 72. Brown, R., P. Bryant, and H. Aberbanal, “Computing the Liapunov
views and 1D projections,” World Congress on Neural Networks, vol. 2, pp. exponents of a dynamical system from observed time series,” Physical
828-837, Washington, DC, 1995. Review A, vo1.43, pp.2787-2806, 1991.

MARCH 1996 IEEE SIGNAL PROCESSING MAGAZINE 47


73. Schouten, J.C., F. Takens, and C.M. van den Bleek, “Maximum-likeli- 99. Cohen, L., Timeyrequency Analysis, Prentice-Hall, Englewood Cliffs,
I

hood estimation of the entropy of an attractor,” Physical Review E, ~01.49, NJ, 1995.
pp.126-129, 1994.
100. Boashash, B., “Time-frequency signal analysis”. In S. Haykin, editor,
74. Farmer, J.D., E. Ott, and J. A. Yorke, “The dimension of chaotic Advances in Spectrum Analysis and Array Processing, vol. I, pp-418-517,
attractors,” Physica D, vol7, pp. 153-180, 1983. Prentice-Hall, Englewood Cliffs, NJ. 1991.
75. Grassberger, P., and I. Procaccia, “On the characterization of strange 101. Wigner, E.P., “On the quantum correction for thermodynamic equilib-
attractors,” Phys. Rev. Letters, vol. 50, p.346, 1983. rium,” Phys. Rev., vol. 40, pp. 749-759, 1932.
76. Pineda, F.J., and J.C. Sommerer, “A fast algorilhm for estimating the [Link], J., ‘Theorie et applications de la notion de signal analytique,”
generalizeddimension and choosing time delays.”In Time Series Prediction: Cables et Transmissions, vol.2A, pp.61-74, 1948.
Forecasting the Future and Understanding the Past, edited by A.S. Weigend
and N.A. Gershenfeld, pp.367-385, AddisonWesley, 1994. [Link], L., “Generalized phase-space distribution functions,” J. Math.
Phys., vol. 7, pp.781786. 1966.
77. Schouten, J.C., F. Takens, and C.M. van den Bleek, “Estimation of the
dimension of a noisy attractor,” Physical Review E, vo1.50, pp. 1851-1861, 104. Hlawalsch, F., “Regularity and unitarity of bilinear time-frequency
1994. signal representations,”IEEE Transactions on Information Theory, vol. 38,
pp. 82-94, 1992.
78. Kaplan, J.L., and J.A. Yorke, in Functional Differential Equations and
Approximation of Fixed Points, edited by H.-0. Peitgen and H.O. Walther, 105. Hlawatsch, F., “Bilinear time-frequency representationsof signals: The
Springer, Verlag, 1979. shift-scale invariant class,” IEEE Transactions on Signal Processing, vol.
62, pp. 357-366, 1994.
79. Packard, N., J.P. Crutchfield, J.D. Farmer, and R.S. Shaw, ‘‘Geogrmtq
from a time series,’’ Phys. Rev. Letters, vol. 45, p.712, 1980. 106. Fukunaga, J., Statistical Pattem Recognition, Second Edition, New
York: Academic Press, 1990.
80. Takens, F., “On the numerical determination of the dimension of an
attractor”. In D. Rand and L . 3 . Yong, editors, Dynamical Systems and 107. Flandnn, P., “A time-frequency formulation of optimum detection,”
Turbulence, Warwick 1980. Lecture Notes in Mathematics, vol. 898, pp. IEEE Trans. Signal Process., vo1.36, pp.1377-1384, 1988.
366-381, Springer-Verlag, 1981.
108. Williams, W.J., H.P. Zaveri, and J.C. Sackallares, “Time-frequency
81. Sauer, T., J.A. Yorke, and M. Casdagli, “Embodology,” Joumal of analysis of electrophysiology signals in epilepsy,” IEEE Engineering in
Statistical Physics, vol. 65, pp. 579-616, 1991. Medicine and Biology, pp.133-143, MarcWApril 1995.
82. Whitney, H., “Differentialmanifolds,” Ann. Math., vol. 37, pp. 645-680, 109. Zaveri, H.P., W.J. Williams, L.D. Iasemedis, and J.C. Sackallares,
1936. “Time-frequency representation of electrocorticograms in temporal like
83. Fraser, A.M., “Information and entropy in strange attractors,” IEEE epilepsy,” IEEE Trans. Biomedical Engr., vo1.39, pp. 502-509, 1992.
Trans. Information Theory, vo1.35, pp.245-262, 1989. 110. Baykin, S., andT. Bhattacharya,“Wigner-Villedistribution:An impor-
84. Aberbanel, H., and M.B. Kennel, “Local false nearest neighbors and tant functional block for radar target detection in clutter,” in ASILUMAR
dynamical dimensions from observed chaotic data,” Physical Review E, Conference on Signals, Systems, and Computers, Pacific Grove, CA., 1994.
~01.47,pp.3057-3068, 1993. 111. Haykin, S., and Bhattacharya, T.K., “Modular learning strategy for
85. Abarbanel, H.D.I., R. Brown, J.J. Sidorowich, andL.S. Tsimring, “The signal detection in a nonstationary environment,” IEEE Transactions on
analysis of observed chaotic data in physical systems,” Reviews of Modern Signal Processing, under review.
Physics, vol. 65, pp. 1331-1392, 1993. 112. White, L.B., and Boashash, B., “Time-frequency coherence - a
86. Tong, H., “A personal overview of nonlinear time series analysis from a theoreticalbasis for cross spectral analysis of nonstationary signals,”in Proc.
chaos perspective,” 15th Nordic Conference on Mathematical Statistics, IASTED Int. Symp. Signal Processing and its Applications, pp. 18-23, Bris-
Lund, Sweden, August 1994. bane, Australia, 1987.
87. Ruelle, D., Chance and Chaos, Princeton University Press, p.67, 1991. 113. Abeyskera, S.S., and [Link], “Methods of signal classification
using the images produced by the Wigner-Ville distribution,” Pattem Rec-
88. Li, T., and J.A. Yorke, “Periodthree implies chaos,”[Link]. Monthly, ognition Letters, vol. 12, pp. 717-729, 1991.
vol. 82, p.985, 1975.
[Link], T., and W. Stuetzle, “Principalcurves,”J. American Stat. Assoc.,
89. Lorenz, E.N., “Deterministic nonperiodic flow,” J. Amos. Sci., vol. 20, vol. 84, pp. 502516. 1989.
pp. 130-141, 1963.
115. Kohonen, T., “The self-organizingmap,” Proceedings of the IEEE, vol.
90. Lorenz, E.N., “On the prevalence of periodicity in simple systems.” In
78, pp. 1464-1480, 1990.
Global Analysis, edited by Mgrmele and J. Marsden, pp.53-75, Springer-
Verlag, 1979. 116. Mulier, F., and V. Cherkassky, “Self-organizationas an iterative kernel
smoothing process,” Neural Computation, vol. 7, pp. 1141-1153, 1995.
9 1. Ruelle, D., and F. Takens, “On the nature of turbulence,” Commun. Math.
Phys., ~01.20,pp.167-192, 1971. 117. Sanger, T.D., “Optimal unsupervised learning in a single-layerfeedfor-
92. Aberbanel, H.D.I., R. Katz, J. Cembrole, T. Galeb, and T. Frison, ward neural network,” Neural Networks, vol. 1, pp. 459-473, 1989.
“Nonlinearanalysis of high Reynold number flows over a buoyant asymmet- 118. Hebb, D.O., The Organization ofBehavior, Wiley, New York, 1949.
ric body,” Phys. Rev. E, vo1.49, pp. 40034018, 1994.
119. Oja, E., “A simplifiedneuronmodel as aprincipal component analyzer,”
93. Casdagli, M., “Nonlinear prediction of chaotic time series,” Physica, Journal ofMathematica1 Biology, vol. 15, pp. 267-273, 1982.
vo1.35D, pp.335-356, 1989.
120. LeCun, Y. et al., “Handwritten digit recognition with a back-propaga-
94. Haykin, S., A. Ukrainec, B. Cnrrie, X. Li and M. Audette, “A neural tion network”. In D.S. Touretsky, editor, Advances in Neural Information
network-based noncoherent radar processor for a chaotic ocean environ- Processing Systems, vol. 2, pp. 396-404,Morgan Kaufmann, SanMeteo, CA.
ment,” ANNIE’95, St. Louis.
121. Stutt, C.A., and L.J. Spafford, “A ‘best’ mismatched filter response for
95. Suga, N., “Bisonar and neural computationin bats,” Scient$% American, radar clutter discrimination,”IEEE Transactions on Information Theory, vol.
vol. 262(6), pp. 6068-, 1990. IT-14, pp. 280-287, 1968.
96. Gabor, D., “Theory of Communication,” J. IEE, vo1.93, pp.429-457,
1946. 122. Perrone, M.P., and L.N. Cooper, “Learning from what’s been learned
Supervised learning in multi-neural network systems,” World Congress on
97. Haykin, S., Communication Systems, Third Edition, Vlley, 1994. Neural Networks, vol. 3, pp. 354-357, Portland. OR. 1993.
98. Cohen, L., “Time-frequency distributions - A review,” Proceedings of 123. Dony, R., and S. Haykin, “Optimally Adaptive Transform Coding,”
the IEEE, vol. 77, pp. 941-981, 1989. IEEE Tram. Image Processing, October 1995.

48 IEEE SIGNAL PROCESSING MAGAZINE MARCH 1996


124. Dony R., and S. Haykin, “Multi-classmaximum entropy coder,” IEEIl 136. Becker, S., and G.E. Hinton, “Learning Mixture Models of Spatial
International Con& on Systems, Man, and Cybernetics, Vancouver, BC, Coherence,”Neural Computation, vol. 5 , pp. 267-277, 1993.
Canada, Oct. 1995.
[Link], A., and S. Haykin, “A Modular Neural Network for Enhance-
125. Gray, R.M., “Vector quantization,” IEEE ASSP Magazine, vol. 1, pp. ment of Cross-polar Radar Targets,” Neural Networks, vol9, 1996.
4-29, 1984.
138. Ukrainec, A., and S. Haykin, “Mutual Information-based Learning and
126. Kohonen, T., Self-organizing Maps, Springer-Verlag, 1995. its Application to Radar,” in Fuzzy Logic and Neural Networks Handbook,
to be published by McGraw-Hill Inc., C.H. Chen, Editor, 1996.
[Link], D.H., and T.N. Wiesel, “Brain mechanisms of vision,” Scientific
American, pp. 130146, September 1979. 139. Bell, A.J., andT.J. Sejnowski,“An information-maximization approach
to blind separation and blind deconvolution,” Neural Computation, vol. 7,
128. Comsweet, T.N., “Visual Perception,” Academic Press, New York,
1995.
1970.
140. Bell, A.J., and T.J. Sejnowski, “Blind separation and blind deconvolu-
129. Dony, R.D., S. Haykin, C. Coblenz, and C. Nahmias, “Compressionof
tion: An information-theoreticapproach,” Proc. ICASSP, vol. 5, pp. 3415-
digital chest radiographs using a mixture of principal components neural
3418, Detroit, Michigan, 1995.
networks,” Radiological Society of North America, Chicago, IL., November
26-December 1, 1995. 141. Herault, J., and C. Jutten, “Space or time adaptive signal processing by
neural network models,” In Neural Networks for Computing edited by J.S.
130. Linsker, R., “Self organization in aperceptual network,” Computer, voll.
Denker, AIP Conference Proceedings ISI, American Institute for Physics,
21, pp. 105-117, 1988.
1986.
131. Linsker, R., “An application of the principle of maximum information
preservation to linear systems”. In D.S. Touretzky, editor, Advances in 142. Comon, P., “Independentcomponent analysis, a new concept?,”Signal
Neural Information Processing Systems, vol. 1, pp. 186-194,Morgan Kauf- Processing, vol. 26, pp. 287-314, 1994.
mann, San Mateo, CA, 1989. 143. Haykin, S. editor, Blind Deconvolution, Englewood Cliffs, NJ: Pren-
132. Linsker, R., “How to generate ordered maps by maximizing the mutual tice-Hall, 1994.
information between input and output signals,” Neural Computation, vol. I,, 144. Duhammel, P., “Tutorial: Blind Equalization,” The 1995 IEEE Inter-
pp. 402-411, 1989. national Conference on Acoustics, Speech, and Signal Processing, Detroit,
133. Linsker, R., “Self organization in a perceptual system: How network Michigan, 1995.
models and information theory may shed light on neural organization”. In 145. Haykin, S., “Blind equalization formulated as a self-organized learning
S.J. Hanson and C.R. Olson, editors, Connectionist Modeling and Brain process,” 26th Annual Asilomar Conference on Signals, Systems, and Com-
Function: The Developing Interface, pp. 351-392, MIT Press, Cambridge, puters, vol. 1, pp. 346-350, Pacific Grove, California, 1992.
MA, 1990.
146. Cover, T.M., and J.A. Thomas, Elements oflnformation Theory, Wiley,
134. Becker, S., “An Information-theoreticUnsupervised Learning Algo- New York, 1991.
rithm for Neural Networks,” Ph.D. Thesis, University of Toronto, Ontario,
Canada. 147. Comon, P., C. Jutten and J. Herault, “Blind separation of sources, part
I1 problem statement,” Signal Processing, vol. 24, pp. 11-21, 1992.
135. Becker, S., and G.E. Hinton, “A self-organizing neural network that
discovers surfaces in random-dot stereograms,” Nature (London), vol. 355, [Link] H.B., Unsupervised learning.”Neural Computation, vol. 1, pp.
pp. 161-163, 1992. 295-311, 1989.

MARCH 1996 IEEE SIGNAL PROCESSING MAGAZINE 49

Common questions

Powered by AI

Neural networks exhibit several properties that are advantageous for statistical signal processing applications. Firstly, they are distributed nonlinear devices, meaning each neuron in the network uses a nonlinear activation function, enabling the network to model the nonlinearities inherent in signal data . Secondly, they offer fault tolerance due to their massively parallel structure, which allows for graceful degradation of performance when components fail . Thirdly, they have adaptive capabilities, allowing them to adjust their parameters in response to statistical changes in the environment . Finally, they work well in nonstationary environments if the stability-plasticity dilemma is addressed .

The massively parallel and distributed structure of neural networks contributes to their fault tolerance by spreading information across many neurons. Consequently, the failure of some neurons or their connections has a limited impact on overall performance, allowing the network to degrade gracefully rather than catastrophically . This distributed approach means the network is not reliant on a single component, which contrasts with other methods where failures can have significant, uncompensated impacts. Thus, neural networks maintain operability even under adverse conditions, providing robustness and resilience in their applications .

Back-propagation plays a crucial role in training multilayer perceptrons by providing a systematic method for updating the network's weights to minimize the error between actual and predicted outcomes. In signal processing tasks, back-propagation allows the network to learn complex patterns and dynamics from the data, adjusting weights based on the error gradient. This algorithm is particularly effective in enabling the perception of non-linear and non-stationary data changes, thus equipping the networks to handle intricate signal processing tasks that traditional methods might struggle with .

Mutual information is preferred over the mean-square error for designing unsupervised neural networks because it provides a more comprehensive measure of dependency between input and output variables. This criterion considers the entire distribution of the data rather than just fitting the data points, which allows for capturing the underlying data structure more effectively. It moves beyond merely minimizing errors to better capturing complex patterns and relationships, especially in nonlinear and nonstationary environments commonly encountered in statistical signal processing .

Neural networks can be effectively integrated with other techniques, such as wavelet transforms, to enhance their function in solving complex signal processing problems. This combination allows neural networks to handle both time and frequency dimensions of data, improving their capacity to detect and classify signals with high precision. Additionally, employing hybrid systems where neural networks are used alongside traditional statistical methods can optimize performance and exploit the strengths of both approaches. Furthermore, the use of specialized hardware accelerators can tackle the challenge of long training times, enabling faster convergence on complex datasets without compromising accuracy .

Recursive prediction processes in neural networks illustrate their dynamic modeling capabilities by enabling the network to generate future data points based on learned patterns. In applications like the modeling of sea clutter, the network is trained on initial data and continues to make predictions by feeding its own output back into the input, effectively simulating the signal over time autonomously. This demonstrates the network's ability to comprehend and model the underlying dynamics of the data, maintaining predictive accuracy for multiple steps before divergence occurs due to accumulated error or intrinsic variability in the signal .

Neural networks can reduce implementation costs primarily through their ability to solve complex problems more efficiently than traditional methods, which might require more extensive and costly input data processing. Their adaptability to different environments without needing extensive reprogramming for each new application lowers ongoing operational costs. Moreover, the parallel nature of their structure allows concurrent processing, further reducing the time and resources needed for computation, which translates into cost savings. Ultimately, for many applications, these reductions occur without a performance trade-off, as neural networks can adaptively refine their processing parameters rather than relying on rigid, conventional algorithms .

One practical limitation of neural networks is the long training time required, which arises from the necessity of extensive data processing and the current limitations of computing architecture. This can be partially mitigated by using special-purpose processors like the ANNA chip and employing parallel processing techniques . Another limitation is the difficulty in understanding how neural networks represent learned knowledge, which can make interpretation and debugging challenging. Visualization tools such as Hinton and bond diagrams can assist in making these representations more comprehensible .

Neural networks outperform traditional signal processing methods in several ways. They provide significant improvements in statistical performance in complex real-world applications due to their ability to model non-linear and non-stationary data . Additionally, their massively parallel structure allows for more robust handling of component failures compared to traditional methods . Neural networks can also tackle problems unsolvable by standard methods by either using neural networks alone or in combination with other techniques . However, challenges such as long training times and difficulty in understanding the learned representations in neural networks can limit their practicality .

Training neural networks for real-world applications presents several challenges. One significant issue is the long training time required, often constrained by the serial nature of current computer architectures, which are not well-suited to programming neural networks . This process can be accelerated with special-purpose processors like the ANNA chip and CNAPS, which enhance the training speed for specific types of networks . Another challenge is understanding how neural networks represent the knowledge they gain during training, though tools like the Hinton and bond diagrams can help visualize the processes .

You might also like