Neural Networks in Signal Processing
Neural Networks in Signal Processing
orizons
SIMON HAYKIN
tatistical signal processing covers an area where phys- Interest in neural networks, or to be more precise, artificial
ics and mathematics meet and interact to solve a wide neural networks, has always been motivated by the fact that
range of problems. Its origins may be traced back to the the human brain functions in a manner entirely different from
1943 classified RCA report by North, republished in [l], the the conventional digital computer. The human brain is a
1946 classic paper [Z]by Van Vleck and Middleton, and the gigantic, and yet highly efficient, information-processing
pioneering work by Wiener [ 3 ] . In particular, the classical machine that encompasses a wide variety of complex signal
methods of statistical signal processing are founded on three processing operations. To appreciate the enormous scale of
basic assumptions: linearity, stationarity, and second-order these operations, we need only look at our visual and auditory
statistics with particular emphasis on Gaussianity. These systems and be amazed at the “seamless” nature of the way
assumptions are invoked for the sake of mathematical tracta- in which different forms of information gathered by our eyes
bility. Yet most, if not all, the physical signals that we have and ears are individually processed and then finally fused
to deal with in real-life applications are generated by dynamic together.
processes that are simultaneously nonlinear, nonstationary, Work on neural networks may be traced back to the
and non-Gaussian. The end result of designing a signal-proc- pioneering paper [4] by McCulloch and Pitts in 1943, which
essing system along traditional lines is a suboptimal solution. was followed by Rosenblatt’s development of the perceptron
One way in which the performance of the system can be [5] and Widrow’s development of the adaline [6] in the late
improved is to consider the use of neural networks in combi- 1950s. After going through a period of dormancy (in an
nation with other suitable techniques (e.g., time-frequency engineering context) in the 1970s, neural networks re-
analysis), depending on the task at hand. emerged in the 1980s with the publication of Hopfield’s
Discuss mutual information as a criterion for designing Neural networks provide a nonparametric approachfor the
unsupervised neural networks, thus moving away from the nonlinear estimation of data
mean-square error criterion The nonlinear, feedforward multilayer class of neural net-
works (encompassing multilayer perceptrons and radial ba-
Rationale for Using Neural Networks sis-function networks) learns about its environment in a
supervised manner. (The design of multilayer perceptrons
Neural networks have a number of important properties that and radial-basis function networks is discussed in the book
befit their use for signal processing applications. In particu- by Haykin [ 121, and the review papers by Lippmann [ 131and
lar, we mention the following five properties: Hush and Horne [ 141.) Specifically, these neural networks
undergo a training session during which their free parameters
Neural networks are distributed nonlinear devices (i.e., synaptic weights and biases) are adjusted in a systematic
This property is a direct result of the fact that each processing way so as to minimize a cost function. Typically, the cost
unit (i.e., neuron) of a neural network has a built-in activation function is defined on the basis of a mean square-error
function (for example, in the form of a logistic function) that criterion, with the error signal itself being defined as the
is nonlinear. Accordingly, neural networks have the inherent difference between a desired response and the actual output
ability to model underlying nonlinearities contained in the of the network produced in response to a corresponding input
physical mechanism responsible for generating the input signal. The neural network learns from examples by con-
data. structing an input-output mapping for the problem at hand,
which brings to mind the notion of nonparametric statistical
A neural network consists of a massively parallel processor inference; see Table 1. The term “nonparametric”is used here
that has the potential to be fault tolerant in a statistical sense, meaning that no knowledge of the
For example, a multilayer perceptron, representing a popular underlying probability distribution is required.
structure for the implementation of a neural network, consists In the traditional approach to mathematical statistics as
of a large number of neurons arranged in the form of layers, taught in a statistics department, the issues of primary con-
with each neuron in a particular layer connected to a large cern are two-fold:
number of source nodestneurons in the previous layer. This
form of global interconnectivity has the potential to be fault
tolerant, in the sense that the performance is degraded grace- The use of mathematically tractable models, assuming the
fully under adverse operating conditions. If a neuron or its idealized conditions of linearity, wide-sense stadonarity,
synaptic links are damaged, the recall quality of a stored and Gaussianity, for the derivation of parameter estimators.
pattern is impaired, but owing to the highly distributed nature
of the network, the damage has to be extensive before the Derivation of exact properties (e.g.,mean and variance) of
performance is seriously degraded. Nevertheless, to be as- estimators for small sample-sizes; if the exact properties
sured that the neural network is in fact fault tolerant, we may are not mathematically tractable, then one would consider
find it necessary to take proper measures in designing the the asymptotic properties of the estimators as the number
algorithm used to do the training [8]. of samples approaches infinity.
Now consider three case studies (based on real-life data), with 1. The largest Liapunov exponent, X i , is always positive. For
which I and some of my research colleagues have been the particular radar used to do the data collection, X i is
involved for the past six years. These studies, in their own estimated to be about 0.03, which is normalized with respect
ways, testify to the computing power of neural networks in to the pulse-repetition period of the radar. This value is
solving difficult signal processing problems. essentially independent of the following (for a given radar
system):
Case Study I: Chaotic Modeling of Sea Clutter and radar parameter (i.e., amplitude, in-phase component, or
its Cancellation quadrature component
sea state
For nearly half a century, sea clutter (i.e., the radar backscat- radar location
ter from an ocean surface) has been modeled as a stochastic
process, with a variety of probability distributions proposed Moreover, the second Liapunov exponent, X2, is very close
for describing its stochasticity [60-631. However, there is to zero, and for a prescribed embedding dimension, the sum
now strong experimental evidence that shows that sea clutter of all the Liapunov exponents is negative. The implications
is indeed a chaotic process [64-661. of these latter observations are twofold:
A chaotic process is generated by a deterministic mecha-
nism of a relatively low dimension, and yet it generates a Sea clutter is generated by a coupled system of nonlinear
randomlike waveform that exhibits many of the charac- differential equations
teristics that are normally associated with a stochastic proc- The dynamic mechanism responsible for the generation of
ess. Table 2 presents a summary of the important properties sea clutter is a dissipative one
of a chaotic process.
The term “chaos” was coined by J.A. Yorke, an applied 2. The correlation dimension,Dc, is fractal (i.e., non-integer),
mathematician at the University of Maryland [87]. Yorke’s lying in the range of 6 to 9. Moreover, it is also essentially
Unit delay The neural network is next tested for its generalization
x(n-1) performance, as depicted in Fig. 2b. The network is initial-
ized by presenting it a set of samples xk-1), @2), ...,x( 1)
that have not been seen by the network before. The result-
Recursive prediction. ing prediction i (p) is delayed by one time unit and then
fed back to the input. Correspondingly, the samples x@-
independent of radar parameter, sea state, and radar location. 1),..4(2) are each delayed by one time unit, and the oldest
However, unlike the Liapunov spectrum, the correlation di- sample x(1) is dropped to make room for the delayed
mension is essentially independent of the radar system used prediction i ( p ) . The set of p samples so obtained is used
to perform the data collection. to make a new prediction, and the process is repeated until
all the original samples used to do the initialization have
Sea clutter is indeed generated by a chaotic process. But been removed from the recursive prediction process. From
the mechanism by which this chaotic process actually arises that point on, the neural network operates in a completely
in physical terms is unknown. autonomous fashion, producing a time series that is repre-
With this background on the chaotic nature of sea clutter, sentative of the dynamics learned by the neural network as
we may now turn attention to its signal processing implica- a result of the training process.
tions. Specifically, we wish to (1) demonstrate that sea clutter
permits a nonlinear predictive model with a significant hori- For the predictive modelling of sea clutter, we used a
zon of predictability, and (2) describe a novel radar applica- multilayer perceptron trained with the backpropagation algo-
tion exploiting this predictive capability. rithm. The size of the input layer, denoted by p , is chosen in
The nonlinear predictive modelling of sea clutter involves accordance with the formula p 2 rDE, where DE is the
the use of recursive (iterated) prediction [12,93], illustrated embedding dimension, and 7 is the embedding delay (nor-
in Fig. 2. This is a difficult procedure designed to test the malized with respect to the pulse repetition period). In prac-
generalization capability of the model. For the case study tice, it is inadvisable to choose p much larger than the lower
presented here, a neural network is used as the predictive bound, TDE,as the effect of additive noise contaminatingthe
model. There are two separate operations to be considered: input radar data would become more pronounced. Based on
measurements on real-life data reported in [66], the embed-
The neural network is trained to operate as a one-step ding dimension DE for sea clutter is estimated to be 10. Also,
predictor, as depicted in Fig. 2a. A set o f p samples x(n-1), for the particular radar (operating at a pulse repetition fre-
The neurons in both hidden layers used a logistic function for 50 100 150 200 250
Time
their activation functions, whereas the output neuron was
linear. The solid line in Fig. 3 is the actual sea clutter wave- I
4. Sensitivity o f t h e recursive prediction process to a change in
form, and the dashed curve is the result of recursive predic- the neural network design. The network used f o r this experiment
tion. The origin in this figure corresponds to the end of the consists of a n input layer with 45 source nodes, two hidden layers
initialization procedure. We see that for about the first 50 with 80 and 55 neurons, respectively, and one output neuron.
points shown in Fig. 3, the predicted and actual waveforms
of sea clutter match fairly closely and thereafter they diverge.
made previously: the generation of sea clutter is governed by
a coupled system of nonlinear differential equations. In ef-
100 fect, the neural network provides an approximation to such a
1
system.
90
Figure 4 shows the result of a recursive prediction per-
formed by a multilayer perceptron with p = 45, which is
slightly smaller than the lower bound of 50 defined above.
Except for this difference, the model has two hidden layers,
with 80 neurons in the first one and 55 neurons in the second
one, and a single linear output neuron as before. Moreover,
the model is trained with the same data set and of the same
size used to obtain the result shown in Fig. 3, and the recursive
prediction procedure is used to test the model after complet-
ing the training session in exactly the same way as before.
There is a dramatic difference between the results shown in
I Figs. 3 and 4. In particular, when the size of the input layer
Od 50 100 150 200 250 of the multilayer perceptron model is not large enough, the
Time model fails to capture the underlying dynamics of sea clutter.
7. Recursive prediction of sea clutter, using a multilayer percep- Clearly, the choice of a neural network predictor that violates
tron with 50 source nodes, two hidden layers with 80 and 55 neu- the lower bound on the size of p is unacceptable.
rons, respectively, and one output neuron. The solid curve refers
To emphasize the need for a nonlinear predictive model,
to the original sea clutter wavefomz. The dashed c u w e refers to
we show Fig. 5, the recursive prediction results obtained
the recursive predicted waveform, for which thefirst 50points of
the sea clutter set (not shown in the figure) are used as the initial
using (a) an autoregressive (AR) model, and (b) a multilayer
starting point. perceptron model. Both models used 50 delay taps for the
input, in accordance with the lower bound ofp 2 50. Clearly,
the AR model, which is linear, fails completely to capture the
This result confirms that sea clutter produced by a nonco- underlying dynamics of sea clutter.
herent radar (i.e., one that relies on amplitude information Turning next to the radar application of the predictive
alone) is locally predictable. Moreover, the horizon of pre- modelling of sea clutter, Fig. 6 shows the results of another
dictability, namely, 50, is approximately equal to the inverse experiment involving an off-the-shelf commercial noncoher-
of the largest Liaponuv exponent, 0.03. (For an accurate ent marine radar operating in a scanning mode [94]. In this
calculation of the horizon of predictability, see [66].)The fact application, the multilayer perceptron, trained on examples
that a neural network with an input layer of the right size can drawn from sea clutter and then having its synaptic weights
be trained to learn the underlying nonlinear dynamics of sea fixed, acts as a clutter (interference) canceller. In particular.
clutter is further testimony for the important observation through training, the network acquires the function of a
12/
10
6
-4I '
0 50 100 150 200 250 300 350 400 450 E
Samples in the azimuthal direction 5 10 15 20 25 30 35 40 45 50
(4 (a)
1.6
1.4
5
1.2
10
m 1
3
-
-
50.8 15
' 0.6
0.4
25
0.2
30
0
35
-0.2b '
50 100 150 200 250 300 350 400 450 !
Samples in the azimuthal direction IO
(b) 5 10 15 20 25 30 35 40 45 50
(b)
(a)Azimuthal time series containing sea clutter and target. (b)
Prediction error at output of neural network. Target is clearly evi- . B-Scan images: ( a ) Output of conventional CFAR processor.
dent. ( b ) Output of neural network-based clutter canceller.
In the first case study, we showed how aneural network, used terms so that they bounce up and down less, when conipared
as a nonlinear predictive model, can exploit prior knowledge, io the standard form of the WVD. In so doing, the RID
namely, the fact that sea clutter is chaotic. The only informa- provides an “almost” positive distribution, which is what a
tion used in that case study was the information contained in time-frequency energy distribution should be, particularly for
the amplitude of the received signal, which is provided by a those applications that require the analysis of signals. For
noncoherent radar. This second case study pertains to the signal detection, it would be tempting to do the opposite (i.e.,
detection of a weak target signal corrupted by an interfering retain the cross Wigner-Ville distributions and suppress the
signal. Here, one or the other or both of these signals may be auto-terms). The detection strategy would then focus solely
nonstationary, and no prior knowledge about the environ- on the presence or absence of the cross Wigner-Ville distri-
ment is invoked. However, these problems are ameliorated butions. Such a procedure would, however, violate the infor-
through the use of Doppler information in addition to ampli- mation preservation rule by removing useful information
tude information, which requires the use of a coherent radar. contained in the auto-terms; its use is therefore not recom-
Case Studies I and I1 do have one thing in common: in both mended for signal detection.
cases, the radar operates in an ocean environment, with sea
In a clutter-dominated environment, which is the environ-
clutter being the primary source of interference.
ment of interest in Case Study 2, the cross-terms arise only
Now consider a novel modular learning strategy for signal
when a target signal is present. Thus, the presence of such
detection that is motivated by the echo-location (sonar) of a
terms is in fact an asset. We say this because the terms provide
bat, which detects, pursues, and captures its target (e.g., an
another feature that can enhance the visibility of a target in
insect) with a facility and success rate that is the envy of every
the time-frequency image resulting from the application of
radar or sonar engineer [95]. We are not suggesting that the
the WVD. Indeed, the cross-terms are essential to the optimal
modular detection strategy describe$ in this article involves
information-preserving property of the WVD. To appreciate
all the signal processing functions performed in the bat’s
the importance of the WVD for the radar detection problem,
echo-location system. What we are saying is that the principal
we present three sample WVD images of real-life radar
functions that characterize the modular learning strategy are
returns. These represent three different situations pertaining
found in one form or another in the bat’s echo-location
to an ice-infested ocean environment using a coherent radar
system.
[110-1111:
Figure 8 shows a block diagram of the basic detection
strategy consisting of three fundamental functional blocks
that are designed to perform time-frequency analysis, feature 0 Strong radar retum from a large ice target, shown in Fig.
extraction, and pattern classification, in that order. This form 9a.
of front-end processing is commonly used in pattern recog-
nition tasks [106]. 0 Relatively weak radar return from a small ice target, shown
For the time-frequency analysis, we have chosen the in Fig. 9b.
Wigner-Ville distribution (WVD); Table 3 presents a sum-
mary of the important properties of the WVD. Among the
family of bilinear time-frequency distributions, the WVD 0 Sea clutter alone, shown in Fig. 9c.
possesses two distinct advantages over other members of the
family for signal detection [ 1071: The WVD images presented in this figure significantly dif-
ferentiate between these three scenarios. Unfortunately, the
1. It is always a real-valued function. use of the WVD leads to a significant increase in the amount
2. It exhibits the least amount of spread in the time-frequency of redundant information contained in the time-frequency
plane. image of a radar signal. To improve computational effi-
ciency, it is therefore necessary to follow up the WVD with
One criticism that is often made against the WVD is the some form of data compression. (This point is also made in
generation of cross-terms, or more precisely, cross Wigner- [112, 1131, where singular value decomposition is used for
Ville distributions, due to the combined presence of two (or the extraction of features from the WVD image of a signal
more) components in the received signal. Various procedures for the purpose of signal detection or [Link],
have been developed in the literature for dealing with the the scheme described therein is primitive compared to the
cross Wigner-Ville distributions. In [ 108,1091,for example, modular learning strategy embodied in Fig. 10, as it lacks a
an algorithm known as the reduced interference distributions learning capability and does not address the two fundamental
(RID) is described, which is designed to flatten the cross- questions raised later in this case study.)
Classification
Target
channel
1
channel
4.-
0.00 005
A...
0.10 015 0.20 025
Principalmmponents
analyser 1, matched
to the WVD of
interference
Feature
extraction
j Principaicomponents
analyser 1, matched
to the WVD of signal
plus interference
Wagner-Ville
distribution (WVD)
computed
ibi
generalized Hebbian algorithm (GHA) due to Sanger [ 1171.
Table 4 presents a summary of this algorithm. The PCA
network in the clutter channel is trained by presenting it with
WVD images known to represent clutter only, under varying
400 environmentalconditions. Once the training is completed, the
synaptic weights of that PCA network are fixed. The training
200
procedure of the PCA network for the target channel follows
0 a similar procedure, except for the fact that its training
examples consist of WVD images known to contain target
200
plus clutter, under varying conditions. The outputs of the
400 PCA networks may be viewed as a specific number of domi-
-10 -5 0 5 10 nant projections of the input WVD space on two subspaces,
with one subspace representing clutter alone and the other
subspace representing target plus clutter. Typically, these two
subspaces are unknown and nonlinear; projections of the
WVD space onto them are therefore best learned by way of
000 005 010 015 020 025 real-life examples that are representative of the two scenar-
ios. The end result is that the PCA network in one channel is
. (a) WVDfor a clearly visible growler; (b) WVDfor a barely adaptively matched to clutter alone, and the PCA network in
visible growler; (c) WVDfor sea clutter. the other channel is adaptively matched to target plus clutter,
hence the designations of the two channels in Fig. 10 as
cally nonlinear, making the task that much more challenging clutter (interference) and target channels, respectively.
to implement. Each multilayer perceptron has two hidden layers and an
Figure 10 shows a block diagram of a neural network- output layer with three output nodes. The output nodes are
based implementation of the modular detection strategy de- linearly combined into a single decision-making node. Thus,
scribed herein [110-1111. It consists of two channels, one the decision as to whether a target is present or not is deferred
termed the clutter or interference channel, and the other to the very output of the system, in accordance with the
termed the target channel. Both channels are fed from a information preservation rule. Specifically, if a threshold set
common input representing the WVD image of the received for a prescribed probability of false alarm is exceeded by the
signal. Each channel consists of a PCA network followed by overall output of the receiver, a decision is made that a target
a multilayer perceptron for pattern classification. The PCA is present; otherwise, a decision is made that the received
networks are trained in a self-organized fashion, using the radar signal consists of clutter alone. The synaptic weights of
100
0 the received signal contains a strong target signal
1w, ikM
+x
!
] .Li
..................................
rk7
.
;
used for training. The image in Fig. 14b is the adjacent section
from the same study (patient), which was used for testing.
U
Each image consists of 256 x 256 pixels, with the dynamic
IQM 1 : ......................................... !
range of 8 bits or 256 gray levels. The training image was
divided into blocks of 8 x 8 pixels for an input dimension of
I
N = 64. The blocks were overlapped at two pixel intervals for
’[Link] structure of OIAL scheme. a total number of 15,625 training samples. During training,
the samples were presented in random order. For comparison,
Figure 13 shows a network structure for implementing one the KLT was calculated based on the same training data.
The test image was divided into 8 x 8 non-overlapping
particular form of the MPC. The system is modular, consist-
blocks. These blocks were transformed by the previously
ing of a number of modules corresponding to different classes
computed system into a set of coefficients, quantized, and
of input data. Each module consists of a linear transformation, then transformed back into image blocks. The coefficients
whose basis vectors are computed using an initial training were quantized in a similar manner to that of the JPEG
period. The appropriate class for a given input vector is standard. The first coefficient was coded via first-order
determined by the subspace classifier. The system utilizes a DPCM using a uniform quantizer. The remaining coefficients
learning algorithm referred to as the optimallv integrated were coded via PCM using a uniform quantizer. For a given
adaptive learning (OIAL) algorithm. The algorithm is self- coding rate, the same quantization interval was used for all
Acknowledgments 6. Widrow, B., and M.E. Hoff, Jr., “Adaptive switching circuits,” IRE
WESCON Convention Record, pp. 96-104,1960,
7. Hopfield, J.J., “Neural networks and physical systems with emergent
I am truly grateful to Dr. Vladimir Cherkassky, University of collective computational abilities,” Proc. Nat. Acad. Sci. USA, vol. 79, pp.
Minnesota, Minneapolis, and Dr. William J. Williams, Uni- 2554-2558,1982.
versity of Michigan, Ann Arbor, for their many critical inputs 8. Rumelhart, D.E., and J.L. McClelland, editors, Parallel Distributed Proc-
and useful suggestions that have impacted the finalizing of essing: Explorations in the Micvostructure of Cognition, vol. 1, Cambridge,
this article in a significant way. I also wish to thankDr. Henry MA: MIT Press. 1986.
D.I. Aberbanel, University of Califomia, San Diego, Dr. 9. Chiu, C.-T., K. Mehrota, C. K. Mohan, and S. Ranka, “Robustness of
feedforward neural networks.” In P. K. Simpson, editor, Neural Networks:
Anthony Bell, Salk Institute, California, Dr. Fay Boudreaux- Theory, Technology, andAppZications, pp. 348-353,1996, E E E Press, New
Bartels, University of Rhode Island, Dr. Bemie Mulgrew, York.
University of Edinburgh, Scotland, and Dr. Kari Torkkola, 10. Haykin, S . , Adaptive Filter Theory, Third Edition, Prentice-Hall, 1996.
Motorola, Utah, for their helpful inputs on selected parts of 11. Wldrow, B., and S. Steams, Adaptive Signal Processing, Prentice-Hall,
the article. I am indebted to my own colleagues, Dr. Sue 1985.
Becker, Department of Psychology, Dr. Nanda Kambhatla, 12. Haykin, S., Neural Networks: A Comprehensive Foundation, New York
Communications Research Laboratory, Dr. James P. Reilly, Macmillan College Publishing Company, 1994.
Department of Electrical and Computer Engineering, and Dr. 13. Lippmann, R.P., “An introduction to computing with neural nets,” IEEE
Patrick Yip, Department of Mathematics and Statistics, my ASSP Magazine, vol. 4, pp. 4-22,1987.
former research colleagues Dr. Tarun Bhattacharya, 14. Hush, D.R., and B.G. Home, “Progress in supervised neural networks:
Raytheon Canada, Dr. Robert D. Dony, Sir Wilfrid Laurier What’s new since Lippmann?,” IEEE Signal Processing Magazine, vol. 10,
University, and Dr. Graeme Jones, Raytheon Canada, and my pp. 8-39, 1993.
current graduate students Hugh Pasika and Paul Yee for 15. Rao, C.R., Linear Statistical Inference and its Applications, Second
reading the entire manuscript and offering numerous sugges- Edition, Wiley, New York. 1973.
tions for improving it. I wish to acknowledge the many [Link], A.R., and R.L. Barron, “Statistical learning networks: Aunifying
view,” In Proceedings of the 198%Symposium on the Inte~ace:Statistics
helpful inputs on chaos that I received on the Intemet from and Computing Science (Ed., E.J. Weaman), pp. 192-203,American Statis-
Dr. Martin Casdagli, University of Michigan, Ann Arbor, Dr. tical Association, Washington, DC, 1988.
Daniel Kaplan, McGill University, Quebec, Dr. Mathew 17. White, H., ‘‘Learningin artificial neural networks: A statistical perspec-
Kennel, Oak Ridge National Laboratory, Tennessee, Dr. Lou tive,” Neural Computation, vol. 1, pp. 425-464, 1989.
Pecora, Naval Research Laboratory, Washington, DC, Dr. 18. Ripley, B.D., “Neural networks and related methods for classification,”
Florin Takens, University of Gronigen, The Netherlands, Dr. J. Royal Statistical Society Series B, vol. 56, pp. 409-456, 1995.
James Theiler, Los Alamos National Laboratory, New Mex- 19. Ripley, B.D., “Flexible nonlinear methods for classification,” World
ico, Dr. Howell Tong, University of Kent, England, and Dr. Congress on Neural Networks, vol. 2, pp. 927-934, Washington, DC., 1995.
James Yorke, University of Maryland. Last, but by no means 20. Murtagh, F., “Neural networks and related massively parallel methods
least, 1wish to thank Dr. Jose Principe, University of Florida, for statistics: A short review,” Intemational Statistical Review, vol. 62, pp.
275-288, 1994.
Gainsville. Dr. Andrew Webb, Defence Research Agency,
Great Malvern, England, three anonymous reviewers and Dr. 21. Cheng, B., and D.M. Titterington, “Neural Networks: A review from a
statistical perspective,” Statistical Science, vol. 9, pp. 2-54, 1994.
Don Hush, Associate Editor of the SP Magazine for their
careful reviews of an earlier version of the article. 22. Cherkassky, V., J.H. Friedman, andH. Wechsler, editors, From Statistics
to Neural Networks: Theory and Pattern Recognition Applications, Sprin-
ger-Verlag, 1994.
Simon Haykin is Professor of Electrical and Computer 23. Helstrom, C.W., Statistical Theory ofsignal Detection, Second Edition,
Engineering, McMaster University, Hamilton, Ontario, Can- Pergamon Press, Oxford. 1968.
ada. 24. Poor, H.V., An Introduction into Signal Detection and Estimation,
Springer-Verlag, New York. 1988.
38. Vapnik, V.N., and A.Y. Chervonenkis, “On the uniform convergence of 60. Goldstein, H., “Sea echo in propagation of short radio waves.” In D.E.
relative frequencies of events to their probabilities,” Theoretical Probability Kerr, editor, MIT Radiation Laboratory Series, vol. 13, Section 6.6,
and its Applications, vol. 17, pp. 264280, 1971. McGraw-Hill, New York, 1951.
61. Jakeman,E., andP.N. Pusey, “A model for non-Rayleigh sea echo,” IEEE
39. Vapnik, V.N., The Nature of Statistical Learning Theory, Springer-Vler-
lag, 1995. Transactions on Antennas and Propagation, vol. AP-24, pp.806-814,1976.
62. Skolnik, M. editor. Radar Handbook, Second Edition, McGraw-Hill,
40. Natrajan, B.K., Machine Leanzing: A Theoretical Approach, Morgan
New York, 19...
Kaufmann, 1991.
63. Ward, K.D., C.J. Baker, and S. Watts, “Maritime surveillanceradar, Part
41. Hlawatsch, F., and G.F. Boudreaux-Bartels,“Linear and quadratic tinie- I Radar scattering from the ocean surface,” IEE Proceedings, vol. 137, Pt.
frequency signal representations,”IEEE Signal Processing Magazine, vol. F,pp. 51-62, 1990.
9, pp. 21-61, 1992.
64. Leung, H., and S. Haykin, “Is there a radar clutter attractor?” Applied
42. Rioul, O., and M. Vetterli, “Wavelets and signal processing,” [Link] Physics Letters, vol. 56, pp. 592-595, 1990.
Signal Processing Magazine, vol. 8. No. 10, 1991.
65. Li, X., and S. Haykin, “Chaotic characterizationof sea clutter,” l’Onide
43. Brotherton,T., T. Pollard, andD. Jones, “Applicationsof time-frequency Electrique, Special Issue on Radar, SEE, France, pp. 60-65, March 1994.
and time-scale representations to fault detection and classification,” Pro-
ceedings of the IEEE-SP International Symposium on Time-frequency and 66. Haykin; S., and X. Li. “Detection of signals in chaos,” Proceedings of
Time-scale Analysis, pp. 95-97, Victoria, BC. 1992. the IEEE, vol. 83, pp. 95-122. 1995.
44. Pati, Y.C., and P.S. Krishnaprasad, “Analysis and synthesis of feedfor- 67. Schuster,H.G., Deterministic Chaos: An Introduction, VCH, Weinheim,
ward neural networks using discrete affine wavelet transformations,”IEEE Germany, 1980.
Transactions on Neural Networks, vol. 4, pp. 73-85, 1993.
68. Ruelle, D., Chaotic Evolution and Strange Attractors, Cambridge Uni-
45. Manjunath, B.S., and R. Chelleppa, “A unified approach to boundary versity Press, 1989.
perception: Edges, textures, and illusory contours,” IEEE Transactions on 69. Ott, E., Chaos in Dynamical Systems, CambridgeUniversity Press, 1993.
Neural Networks, vol. 4, pp. 96-108, 1993.
70. Newhouse, S., “Understanding chaotic dynamics”. In J. Chandra, editor,
46. Kreinovich, V., 0. Sirisaengtaksin, and S. Cabrera, “Wavelet neural Chaos in Nonlinear Dynamical Systems, SIAM, 1984.
networks are asymptotically optimal approximators for functions of one
variable,” Proceedings of IEEE International Conference on Neural Net- 71. Wolf, A., J.B. Swift, H.L. Swinney, and J.A. Vastano, “Determining
works, vol. 1, pp. 299-304, Orlando, Florida, 1994. Liapunov exponents from a time series,” Physica 16D, pp.285-317,1985.
47. Szu, H., J. Garcia, and J. DeWitte, “Telemedicine:Recognition from 2D 72. Brown, R., P. Bryant, and H. Aberbanal, “Computing the Liapunov
views and 1D projections,” World Congress on Neural Networks, vol. 2, pp. exponents of a dynamical system from observed time series,” Physical
828-837, Washington, DC, 1995. Review A, vo1.43, pp.2787-2806, 1991.
hood estimation of the entropy of an attractor,” Physical Review E, ~01.49, NJ, 1995.
pp.126-129, 1994.
100. Boashash, B., “Time-frequency signal analysis”. In S. Haykin, editor,
74. Farmer, J.D., E. Ott, and J. A. Yorke, “The dimension of chaotic Advances in Spectrum Analysis and Array Processing, vol. I, pp-418-517,
attractors,” Physica D, vol7, pp. 153-180, 1983. Prentice-Hall, Englewood Cliffs, NJ. 1991.
75. Grassberger, P., and I. Procaccia, “On the characterization of strange 101. Wigner, E.P., “On the quantum correction for thermodynamic equilib-
attractors,” Phys. Rev. Letters, vol. 50, p.346, 1983. rium,” Phys. Rev., vol. 40, pp. 749-759, 1932.
76. Pineda, F.J., and J.C. Sommerer, “A fast algorilhm for estimating the [Link], J., ‘Theorie et applications de la notion de signal analytique,”
generalizeddimension and choosing time delays.”In Time Series Prediction: Cables et Transmissions, vol.2A, pp.61-74, 1948.
Forecasting the Future and Understanding the Past, edited by A.S. Weigend
and N.A. Gershenfeld, pp.367-385, AddisonWesley, 1994. [Link], L., “Generalized phase-space distribution functions,” J. Math.
Phys., vol. 7, pp.781786. 1966.
77. Schouten, J.C., F. Takens, and C.M. van den Bleek, “Estimation of the
dimension of a noisy attractor,” Physical Review E, vo1.50, pp. 1851-1861, 104. Hlawalsch, F., “Regularity and unitarity of bilinear time-frequency
1994. signal representations,”IEEE Transactions on Information Theory, vol. 38,
pp. 82-94, 1992.
78. Kaplan, J.L., and J.A. Yorke, in Functional Differential Equations and
Approximation of Fixed Points, edited by H.-0. Peitgen and H.O. Walther, 105. Hlawatsch, F., “Bilinear time-frequency representationsof signals: The
Springer, Verlag, 1979. shift-scale invariant class,” IEEE Transactions on Signal Processing, vol.
62, pp. 357-366, 1994.
79. Packard, N., J.P. Crutchfield, J.D. Farmer, and R.S. Shaw, ‘‘Geogrmtq
from a time series,’’ Phys. Rev. Letters, vol. 45, p.712, 1980. 106. Fukunaga, J., Statistical Pattem Recognition, Second Edition, New
York: Academic Press, 1990.
80. Takens, F., “On the numerical determination of the dimension of an
attractor”. In D. Rand and L . 3 . Yong, editors, Dynamical Systems and 107. Flandnn, P., “A time-frequency formulation of optimum detection,”
Turbulence, Warwick 1980. Lecture Notes in Mathematics, vol. 898, pp. IEEE Trans. Signal Process., vo1.36, pp.1377-1384, 1988.
366-381, Springer-Verlag, 1981.
108. Williams, W.J., H.P. Zaveri, and J.C. Sackallares, “Time-frequency
81. Sauer, T., J.A. Yorke, and M. Casdagli, “Embodology,” Joumal of analysis of electrophysiology signals in epilepsy,” IEEE Engineering in
Statistical Physics, vol. 65, pp. 579-616, 1991. Medicine and Biology, pp.133-143, MarcWApril 1995.
82. Whitney, H., “Differentialmanifolds,” Ann. Math., vol. 37, pp. 645-680, 109. Zaveri, H.P., W.J. Williams, L.D. Iasemedis, and J.C. Sackallares,
1936. “Time-frequency representation of electrocorticograms in temporal like
83. Fraser, A.M., “Information and entropy in strange attractors,” IEEE epilepsy,” IEEE Trans. Biomedical Engr., vo1.39, pp. 502-509, 1992.
Trans. Information Theory, vo1.35, pp.245-262, 1989. 110. Baykin, S., andT. Bhattacharya,“Wigner-Villedistribution:An impor-
84. Aberbanel, H., and M.B. Kennel, “Local false nearest neighbors and tant functional block for radar target detection in clutter,” in ASILUMAR
dynamical dimensions from observed chaotic data,” Physical Review E, Conference on Signals, Systems, and Computers, Pacific Grove, CA., 1994.
~01.47,pp.3057-3068, 1993. 111. Haykin, S., and Bhattacharya, T.K., “Modular learning strategy for
85. Abarbanel, H.D.I., R. Brown, J.J. Sidorowich, andL.S. Tsimring, “The signal detection in a nonstationary environment,” IEEE Transactions on
analysis of observed chaotic data in physical systems,” Reviews of Modern Signal Processing, under review.
Physics, vol. 65, pp. 1331-1392, 1993. 112. White, L.B., and Boashash, B., “Time-frequency coherence - a
86. Tong, H., “A personal overview of nonlinear time series analysis from a theoreticalbasis for cross spectral analysis of nonstationary signals,”in Proc.
chaos perspective,” 15th Nordic Conference on Mathematical Statistics, IASTED Int. Symp. Signal Processing and its Applications, pp. 18-23, Bris-
Lund, Sweden, August 1994. bane, Australia, 1987.
87. Ruelle, D., Chance and Chaos, Princeton University Press, p.67, 1991. 113. Abeyskera, S.S., and [Link], “Methods of signal classification
using the images produced by the Wigner-Ville distribution,” Pattem Rec-
88. Li, T., and J.A. Yorke, “Periodthree implies chaos,”[Link]. Monthly, ognition Letters, vol. 12, pp. 717-729, 1991.
vol. 82, p.985, 1975.
[Link], T., and W. Stuetzle, “Principalcurves,”J. American Stat. Assoc.,
89. Lorenz, E.N., “Deterministic nonperiodic flow,” J. Amos. Sci., vol. 20, vol. 84, pp. 502516. 1989.
pp. 130-141, 1963.
115. Kohonen, T., “The self-organizingmap,” Proceedings of the IEEE, vol.
90. Lorenz, E.N., “On the prevalence of periodicity in simple systems.” In
78, pp. 1464-1480, 1990.
Global Analysis, edited by Mgrmele and J. Marsden, pp.53-75, Springer-
Verlag, 1979. 116. Mulier, F., and V. Cherkassky, “Self-organizationas an iterative kernel
smoothing process,” Neural Computation, vol. 7, pp. 1141-1153, 1995.
9 1. Ruelle, D., and F. Takens, “On the nature of turbulence,” Commun. Math.
Phys., ~01.20,pp.167-192, 1971. 117. Sanger, T.D., “Optimal unsupervised learning in a single-layerfeedfor-
92. Aberbanel, H.D.I., R. Katz, J. Cembrole, T. Galeb, and T. Frison, ward neural network,” Neural Networks, vol. 1, pp. 459-473, 1989.
“Nonlinearanalysis of high Reynold number flows over a buoyant asymmet- 118. Hebb, D.O., The Organization ofBehavior, Wiley, New York, 1949.
ric body,” Phys. Rev. E, vo1.49, pp. 40034018, 1994.
119. Oja, E., “A simplifiedneuronmodel as aprincipal component analyzer,”
93. Casdagli, M., “Nonlinear prediction of chaotic time series,” Physica, Journal ofMathematica1 Biology, vol. 15, pp. 267-273, 1982.
vo1.35D, pp.335-356, 1989.
120. LeCun, Y. et al., “Handwritten digit recognition with a back-propaga-
94. Haykin, S., A. Ukrainec, B. Cnrrie, X. Li and M. Audette, “A neural tion network”. In D.S. Touretsky, editor, Advances in Neural Information
network-based noncoherent radar processor for a chaotic ocean environ- Processing Systems, vol. 2, pp. 396-404,Morgan Kaufmann, SanMeteo, CA.
ment,” ANNIE’95, St. Louis.
121. Stutt, C.A., and L.J. Spafford, “A ‘best’ mismatched filter response for
95. Suga, N., “Bisonar and neural computationin bats,” Scient$% American, radar clutter discrimination,”IEEE Transactions on Information Theory, vol.
vol. 262(6), pp. 6068-, 1990. IT-14, pp. 280-287, 1968.
96. Gabor, D., “Theory of Communication,” J. IEE, vo1.93, pp.429-457,
1946. 122. Perrone, M.P., and L.N. Cooper, “Learning from what’s been learned
Supervised learning in multi-neural network systems,” World Congress on
97. Haykin, S., Communication Systems, Third Edition, Vlley, 1994. Neural Networks, vol. 3, pp. 354-357, Portland. OR. 1993.
98. Cohen, L., “Time-frequency distributions - A review,” Proceedings of 123. Dony, R., and S. Haykin, “Optimally Adaptive Transform Coding,”
the IEEE, vol. 77, pp. 941-981, 1989. IEEE Tram. Image Processing, October 1995.
Neural networks exhibit several properties that are advantageous for statistical signal processing applications. Firstly, they are distributed nonlinear devices, meaning each neuron in the network uses a nonlinear activation function, enabling the network to model the nonlinearities inherent in signal data . Secondly, they offer fault tolerance due to their massively parallel structure, which allows for graceful degradation of performance when components fail . Thirdly, they have adaptive capabilities, allowing them to adjust their parameters in response to statistical changes in the environment . Finally, they work well in nonstationary environments if the stability-plasticity dilemma is addressed .
The massively parallel and distributed structure of neural networks contributes to their fault tolerance by spreading information across many neurons. Consequently, the failure of some neurons or their connections has a limited impact on overall performance, allowing the network to degrade gracefully rather than catastrophically . This distributed approach means the network is not reliant on a single component, which contrasts with other methods where failures can have significant, uncompensated impacts. Thus, neural networks maintain operability even under adverse conditions, providing robustness and resilience in their applications .
Back-propagation plays a crucial role in training multilayer perceptrons by providing a systematic method for updating the network's weights to minimize the error between actual and predicted outcomes. In signal processing tasks, back-propagation allows the network to learn complex patterns and dynamics from the data, adjusting weights based on the error gradient. This algorithm is particularly effective in enabling the perception of non-linear and non-stationary data changes, thus equipping the networks to handle intricate signal processing tasks that traditional methods might struggle with .
Mutual information is preferred over the mean-square error for designing unsupervised neural networks because it provides a more comprehensive measure of dependency between input and output variables. This criterion considers the entire distribution of the data rather than just fitting the data points, which allows for capturing the underlying data structure more effectively. It moves beyond merely minimizing errors to better capturing complex patterns and relationships, especially in nonlinear and nonstationary environments commonly encountered in statistical signal processing .
Neural networks can be effectively integrated with other techniques, such as wavelet transforms, to enhance their function in solving complex signal processing problems. This combination allows neural networks to handle both time and frequency dimensions of data, improving their capacity to detect and classify signals with high precision. Additionally, employing hybrid systems where neural networks are used alongside traditional statistical methods can optimize performance and exploit the strengths of both approaches. Furthermore, the use of specialized hardware accelerators can tackle the challenge of long training times, enabling faster convergence on complex datasets without compromising accuracy .
Recursive prediction processes in neural networks illustrate their dynamic modeling capabilities by enabling the network to generate future data points based on learned patterns. In applications like the modeling of sea clutter, the network is trained on initial data and continues to make predictions by feeding its own output back into the input, effectively simulating the signal over time autonomously. This demonstrates the network's ability to comprehend and model the underlying dynamics of the data, maintaining predictive accuracy for multiple steps before divergence occurs due to accumulated error or intrinsic variability in the signal .
Neural networks can reduce implementation costs primarily through their ability to solve complex problems more efficiently than traditional methods, which might require more extensive and costly input data processing. Their adaptability to different environments without needing extensive reprogramming for each new application lowers ongoing operational costs. Moreover, the parallel nature of their structure allows concurrent processing, further reducing the time and resources needed for computation, which translates into cost savings. Ultimately, for many applications, these reductions occur without a performance trade-off, as neural networks can adaptively refine their processing parameters rather than relying on rigid, conventional algorithms .
One practical limitation of neural networks is the long training time required, which arises from the necessity of extensive data processing and the current limitations of computing architecture. This can be partially mitigated by using special-purpose processors like the ANNA chip and employing parallel processing techniques . Another limitation is the difficulty in understanding how neural networks represent learned knowledge, which can make interpretation and debugging challenging. Visualization tools such as Hinton and bond diagrams can assist in making these representations more comprehensible .
Neural networks outperform traditional signal processing methods in several ways. They provide significant improvements in statistical performance in complex real-world applications due to their ability to model non-linear and non-stationary data . Additionally, their massively parallel structure allows for more robust handling of component failures compared to traditional methods . Neural networks can also tackle problems unsolvable by standard methods by either using neural networks alone or in combination with other techniques . However, challenges such as long training times and difficulty in understanding the learned representations in neural networks can limit their practicality .
Training neural networks for real-world applications presents several challenges. One significant issue is the long training time required, often constrained by the serial nature of current computer architectures, which are not well-suited to programming neural networks . This process can be accelerated with special-purpose processors like the ANNA chip and CNAPS, which enhance the training speed for specific types of networks . Another challenge is understanding how neural networks represent the knowledge they gain during training, though tools like the Hinton and bond diagrams can help visualize the processes .