Introduction To Adaptive Arrays
Introduction To Adaptive Arrays
Introduction to Adaptive Arrays, 2nd Edition is organized as a tutorial, taking the readers by the hand
and leading them through the maze of jargon that often surrounds this highly technical subject. It is
easy to read and easy to follow as fundamental concepts are introduced with examples before more
current developments and techniques are introduced.
Problems at the end of each chapter serve both instructors and professional readers by illustrating and
extending the material presented in the text. Both students and practicing engineers will easily gain
familiarity with the modern contribution that adaptive arrays have to offer practical signal reception
systems.
AUDIENCE
Radar • Sonar • Communications • Seismology • Radio Astronomy
Randy Haupt is a Fellow of the IEEE and Applied Computational Electromagnetics Society (ACES) and
is a Senior Scientist at the Penn State Applied Research Lab. From 1999-2003 he was Professor and
Department Head of ECE at Utah State University. He also was a Professor of EE at the Air Force Academy
and the University of Nevada-Reno. He is a retired Lt. Col. from the US Air Force. He has many journal
articles, conference publications, and book chapters to his credit and is the author of Antenna Arrays:
A Computational Approach by Wiley (2010).
Thomas Miller currently works for Raytheon Corporation where he has been cited for leadership
in advanced early warning surveillance radar and adaptive signal processing. He is the author and
coauthor of multiple journal articles as well as contributor to several books.
Monzingo-7200014 monz7200014˙fm ISBN : XXXXXXXXXX November 24, 2010 16:30 i
2nd Edition
Robert A. Monzingo
Randy L. Haupt
Thomas W. Miller
Raleigh, NC
[Link]
Monzingo-7200014 monz7200014˙fm ISBN : XXXXXXXXXX November 24, 2010 16:30 iv
No part of this publication may be reproduced, stored in a retrieval system or transmitted in any form or by any
means, electronic, mechanical, photocopying, recording, scanning or otherwise, except as permitted under Sections
107 or 108 of the 1976 United Stated Copyright Act, without either the prior written permission of the Publisher,
or authorization through payment of the appropriate per-copy fee to the Copyright Clearance Center, 222 Rosewood
Drive, Danvers, MA 01923, (978) 750-8400, fax (978) 646-8600, or on the web at [Link]. Requests to the
Publisher for permission should be addressed to the Publisher, SciTech Publishing, Inc., 911 Paverstone Drive, Suite
B, Raleigh, NC 27615, (919) 847-2434, fax (919) 847-2568, or email editor@[Link].
The publisher and the author make no representations or warranties with respect to the accuracy or completeness of
the contents of this work and specifically disclaim all warranties, including without limitation warranties of fitness
for a particular purpose.
ISBN: 978-1-891121-57-9
Dedication
Acknowledgements are generally used to thank those people who, by virtue
of their close association with and support of the extended effort required
to produce a book. Occasionally, however, it becomes apparent that rip-
ples across a great span of distance and time should receive their due. I
only recently learned that I am a direct ancestor of a Bantu warrior named
Edward Mozingo who, some 340 years ago, filed a successful lawsuit in the
Richmond County Courthouse thereby gaining his freedom after 28 years of
indentured servitude to become that rarest of creatures in colonial Virginia:
a free black man. While that accomplishment did not directly result in the
production of this volume, the determination and effort that it involved rep-
resent a major human achievement that, it seems to me, deserves to be
recognized, however belatedly.
With deep appreciation to Richard Haupt, a great brother.
Monzingo-7200014 monz7200014˙fm ISBN : XXXXXXXXXX November 24, 2010 16:30 vi
Monzingo-7200014 monz7200014˙fm ISBN : XXXXXXXXXX November 24, 2010 16:30 vii
Brief Contents
Preface xiv
vii
Monzingo-7200014 monz7200014˙fm ISBN : XXXXXXXXXX November 24, 2010 16:30 viii
Monzingo-7200014 monz7200014˙fm ISBN : XXXXXXXXXX November 24, 2010 16:30 ix
Contents
Preface xiv
1 Introduction 3
1.1 Motivation For Using Adaptive Arrays 4
1.2 Historical Perspective 5
1.3 Principal System Elements 6
1.4 Adaptive Array Problem Statement 7
1.5 Existing Technology 9
1.6 Organization of the Book 21
1.7 Summary and Conclusions 22
1.8 Problems 23
1.9 References 24
ix
Monzingo-7200014 monz7200014˙fm ISBN : XXXXXXXXXX November 24, 2010 16:30 x
x Contents
Contents xi
PA R T I I I Advanced Topics
xii Contents
Contents xiii
Preface
This book is intended to serve as an introduction to the subject of adaptive array sensor
systems whose principal purpose is to enhance the detection and reception of certain
desired signals. Array sensor systems have well-known advantages for providing flexible,
rapidly configurable, beamforming and null-steering patterns. The advantages of array
sensor systems are becoming more important, and this technology has found applications
in the fields of communications, radar, sonar, radio astronomy, seismology and ultrasonics.
The growing importance of adaptive array systems is directly related to the widespread
availability of compact, inexpensive digital computers that make it possible to exploit
certain well-known theoretical results from signal processing and control theory to provide
the critical self-adjusting capability that forms the heart of the adaptive structure.
There are a host of textbooks that treat adaptive array systems, but few of them take
the trouble to present an integrated treatment that provides the reader with the perspective
to organize the available literature into easily understood parts. With the field of adaptive
array sensor systems now a maturing technology, and with the applications of these systems
growing more and more numerous, the need to understand the underlying principles of
such systems is a paramount concern of this book. It is of course necessary to appreciate
the limitations imposed by the hardware adopted to implement a design, but it is more
informative to see how a choice of hardware “fits” within the theoretical framework of
the overall system. Most of the contents are derived from readily available sources in the
literature, although a certain amount of original material has been included.
This book is intended for use both as a textbook at the graduate level and as a reference
work for engineers, scientists, and systems analysts. The material presented will be most
readily understood by readers having an adequate background in antenna array theory,
signal processing (communication theory and estimation theory), optimization techniques,
control theory, and probability and statistics. It is not necessary, however, for the reader
to have such a complete background since the text presents a step-by-step discussion of
the basic theory and important techniques required in the above topics, and appropriate
references are given for readers interested in pursuing these topics further. Fundamental
concepts are introduced and illustrated with examples before more current developments
are introduced. Problems at the end of each chapter have been chosen to illustrate and
extend the material presented in the text. These extensions introduce the reader to actual
adaptive array engineering problems and provide motivation for further reading of the
background reference material. In this manner both students and practicing engineers
may easily gain familiarity with the modern contributions that adaptive arrays have to
offer practical signal reception systems.
The book is organized into three parts. Part One (Chapters 1 to 3) introduces the
advantages that obtain with the use of array sensor systems, define the principal system
components, and develop the optimum steady-state performance limits that any array
system can theoretically achieve. This edition also includes two new topics that have
practical interest: the subject of a performance index to grade the effectiveness of the
overall adaptive system, and the important theme of polarization sensitive arrays. Part Two
xiv
Monzingo-7200014 monz7200014˙fm ISBN : XXXXXXXXXX November 24, 2010 16:30 xv
Preface xv
(Chapters 4 through 9) provides the designer with a survey of adaptive algorithms and
a performance summary for each algorithm type. Some important modern developments
in matrix inversion computation and random search algorithms are treated. With this
information available, the designer may then quickly identify those approaches most likely
to lead to a successful design for the signal environment and system constrains that are
of concern. Part Three (Chapters 10, 11, and 12) considers the problem of compensation
for adaptive array system errors that inevitably occur in any practical system, explores the
important topic of direction of arrival (DOA) estimation, and introduces current trends in
adaptive array research. It is hoped that this edition succeeds in presenting this exciting
field using mathematical tools that make the subject interesting, accessible, and appealing
to a wide audience.
The authors would like to thank Northrop Grumman (Dennis Lowes and Dennis
Fortner), the National Electronics Museum (Ralph Strong and Michael Simons), Material
Systems Inc. (Rick Foster), and Remcom Inc. (Jamie Knapil Infantolino) for providing
some excellent pictures.
Monzingo-7200014 monz7200014˙fm ISBN : XXXXXXXXXX November 24, 2010 16:30 xvi
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:5 1
PART I
Adaptive Array
Fundamental Principles:
System Uses, System
Elements, Basic Concepts,
and Optimum Array
Processing
CHAPTER 1 Introduction
CHAPTER 2 Adaptive Array Concept
CHAPTER 3 Optimum Array Processing:
Steady-State Performance Limits
and the Wiener Solution
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:5 2
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:5 3
CHAPTER
Introduction
1
' $
Chapter Outline
1.1 Motivation for Using Adaptive Arrays . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4
1.2 Historical Perspective . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5
1.3 Principal System Elements . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
1.4 Adaptive Array Problem Statement . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
1.5 Existing Technology . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9
1.6 Organization of the Book . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21
1.7 Summary and Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22
1.8 Problems. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23
1.9 References . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24
& %
An array of sensor elements has long been an attractive solution for severe reception
problems that commonly involve signal detection and estimation. The basic reason for
this attractiveness is that an array offers a means of overcoming the directivity and sen-
sitivity limitations of a single sensor, offering higher gain and narrower beamwidth than
that experienced with a single element. In addition, an array has the ability to control
its response based on changing conditions of the signal environment, such as direction
of arrival, polarization, power level, and frequency. The advent of highly compact, inex-
pensive digital computers has made it possible to exploit well-known results from signal
processing and control theory to provide optimization algorithms that automatically ad-
just the response of an adaptive array and has given rise to a new domain called “smart
arrays.” This self-adjusting capability renders the operation of such systems more flexible
and reliable and (more importantly) offers improved reception performance that would
be difficult to achieve in any other way. This revised edition acquaints the reader with
the historical background of the field and presents important new developments that have
occurred over the last quarter century that have improved the utility and applicability of
this exciting field.
3
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:5 4
4 CHAPTER 1 Introduction
distorted by near-field effects. The adaptive capability overcomes any distortions that occur
in the near field (i.e., at distances from the radiating antenna closer than λ/2π where λ is
the wavelength) and merely responds to the signal environment that results from any such
distortion. Likewise, in the far field (at distances from the radiating antenna greater than
2λ) the adaptive antenna is oblivious to the absence of any distortion.
An adaptive array improves the SNR by preserving the main beam that points at the
desired signal at the same time that it places nulls in the pattern to suppress interference
signals. Very strong interference suppression is possible by forming pattern nulls over a
narrow bandwidth. This exceptional interference suppression capability is a principal ad-
vantage of adaptive arrays compared to waveform processing techniques, which generally
require a large spectrum-spreading factor to obtain comparable levels of interference sup-
pression. Sensor arrays possessing this key automatic response capability are sometimes
referred to as “smart” arrays, since they respond to far more of the signal information
available at the sensor outputs than do more conventional array systems.
The capabilities provided by the adaptive array techniques to be discussed in this
book offer practical solutions to the previously mentioned realistic interference problems
by virtue of their ability to sort out and distinguish the various signals in the spatial do-
main, in the frequency domain, and in polarization. At the present time, adaptive nulling
is considered to be the principal benefit of the adaptive techniques employed by adap-
tive array systems, and automatic cancellation of sidelobe jamming provides a valuable
electronic counter–countermeasure (ECCM) capability for radar systems. Adaptive arrays
are designed to incorporate more traditional capabilities such as self-focusing on receive
and retrodirective transmit. In addition to automatic interference nulling and beam steer-
ing, adaptive imaging arrays may also be designed to obtain microwave images having
high angular resolution. It is useful to call self-phasing or retrodirective arrays adaptive
transmitting arrays to distinguish the principal function of such systems from an adaptive
receiving array, the latter being the focus of this book.
6 CHAPTER 1 Introduction
self-optimizing control work established the least mean square (LMS) error algorithm that
was based on the method of steepest descent. The Applebaum and the Widrow algorithms
are very similar, and both converge toward the optimum Wiener solution.
The use of sensor arrays for sonar and radar signal reception had long been common
practice by the time the early adaptive algorithm work of Applebaum and Widrow was
completed [10,11]. Early work in array processing concentrated on synthesizing a “desir-
able” pattern. Later, attention shifted to the problem of obtaining an improvement in the
SNR [12–14]. Seismic array development commenced about the same period, so papers
describing applications of seismic arrays to detect remote seismic events appeared during
the late 1960s [15–17].
The major area of current interest in adaptive arrays is their application to problems
arising in radar and communications systems, where the designer almost invariably faces
the problem of interference suppression [18]. A second example of the use of adaptive
arrays is that of direction finding in severe interference environments [19,20]. Another
area in which adaptive arrays are proving useful is for systems that require adaptive
beamforming and scanning in situations where the array sensor elements must be organized
without accurate knowledge of element location [21]. Furthermore, large, unstructured
antenna array systems may employ adaptive array techniques for high angular resolution
imaging [22,23]. Adaptive antennas are a subset of smart antennas and include topics
such as multiple input, multiple output (MIMO) [24], element failure compensation [25],
reconfigurable antennas [26], and beam switching [27].
Sensor
array
Adaptive
algorithm
w1 w2 w3 … wN
Beamforming
network Σ
Array output
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:5 7
The array consists of N sensors designed to receive (and transmit) signals in the
propagation medium. The sensors are arranged to give adequate coverage (pattern gain)
over a desired spatial region. The selection of the sensor elements and their physical
arrangement place fundamental limitations on the ultimate capability of the adaptive array
system. The output of each of the N elements goes to the beamforming network, where
the output of each sensor element is first multiplied by a complex weight (having both
amplitude and phase) and then summed with all other weighted sensor element outputs to
form the overall adaptive array output signal. The weight values within the beamforming
network (in conjunction with the sensor elements and their physical arrangement) then
determine the overall array pattern. It is the ability to shape this overall array pattern that
in turn determines how well the specified system requirements can be met for a given
signal environment.
The exact structure of the adaptive algorithm depends on the degree of detailed in-
formation about the operational signal environment that is available to the array. As the
amount of a priori knowledge (e.g., desired signal location, jammer power levels) concern-
ing the signal environment decreases, the adaptive algorithm selected becomes critical to
a successful design. Since the precise nature and direction of all signals present as well
as the characteristics of the sensor elements are not known in practice, the adaptive algo-
rithm must automatically respond to whatever signal environment (within broad limits)
confronts it. If any signal environment limits are known or can reasonably be construed,
such bounds are helpful in determining the adaptive processor algorithm used.
8 CHAPTER 1 Introduction
obtained by maximizing the output SNR, so there is an underlying unity to problems that
initially appear to be quite different.
An adaptive array design includes the sensor array configuration, beamforming net-
work implementation, signal processor, and adaptive algorithm that enables the system to
meet several different requirements on its resulting performance in as simple and inexpen-
sive a manner as possible. The system performance requirements are conveniently divided
into two types: transient response and steady-state response. Transient response refers to
the time required for the adaptive array to successfully adjust from the time it is turned on
until reaching steady-state conditions or successfully adjusting to a change in the signal
environment. Steady-state response refers to the long-term response after the weights are
done changing. Steady-state measures include the shape of the overall array pattern and
the output signal-to-interference plus noise ratio. Several popular performance measures
are considered in detail in Chapter 3. The response speed of an adaptive array depends on
the type of algorithm selected and the nature of the operational signal environment. The
steady-state array response, however, can easily be formulated in terms of the complex
weight settings, the signal environment, and the sensor array structure.
A fundamental trade-off exists between the rapidity of change in a nonstationary
noise field and the steady-state performance of an adaptive system: generally speaking,
the slower the variations in the noise environment, the better the steady-state performance
of the adaptive array. Any adaptive array design needs to optimize the trade-off between
the speed of adaptation and the accuracy of adaptation.
System requirements place limits on the transient response speed. In an aircraft com-
munication system, for example, the signal modulation rate limits the fastest response
speed (since if the response is too fast, the adaptive weights interact with the desired
signal modulation). Responding fast enough to compensate for aircraft motion limits the
slowest speed.
The weights in an adaptive array may be controlled by any one of a variety of different
algorithms. The “best” algorithm for a given application is chosen on the basis of a host
of factors including the signal structures, the a priori information available to the adaptive
processor, the performance characteristics to be optimized, the required speed of response
of the processor, the allowable circuit complexity, any device or other technological limi-
tations, and cost-effectiveness.
Referring to Figure 1-1, the received signal impinges on the sensor array and arrives
at each sensor at different times as determined by the direction of arrival of the signal
and the spacing of the sensor elements. The actual received signal for many applications
consists of a modulated carrier whose information-carrying component consists only of
the complex envelope. If s(t) denotes the modulated carrier signal, then s̃(t) is commonly
used to denote the complex envelope of s(t) (as explained in Appendix B) and is the only
quantity that conveys information. Rather than adopt complex envelope notation, however,
it is simpler to assume that all signals are represented by their complex envelopes so the
common carrier reference never appears explicitly. It is therefore seen that each of the N
channel signals xk (t) represents the complex envelope of the output of the element of a
sensor array that is composed of a signal component and a noise component, that is,
In a linear sensor array having equally spaced elements and assuming ideal propagation
conditions, the sk (t) are determined by the direction of the desired signal. For example, if
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:5 9
the desired signal direction is located at an angle θ from mechanical boresight, then (for
a narrowband signal)
2πkd
sk (t) = s(t) exp j sin θ (1.2)
λ
where d is the element spacing, λ is the wavelength of the incident planar wavefront, and
it is presumed that each of the sensor elements is identical.
For the beamforming network of Figure 1-1, the adaptive array output signal is written
as
N
y(t) = w k xk (t) (1.3)
k=1
y(t) = wT x = xT w (1.4)
where the superscript T denotes transpose, and the vectors w and x are given by
wT = [w 1 w 2 . . . w N ] (1.5)
xT = [x1 x2 . . . x N ] (1.6)
Throughout this book the boldface lowercase symbol (e.g., a) denotes a vector, and a
boldface uppercase symbol (e.g., A) denotes a matrix.
The adaptive processor must select the complex weights, w k , to optimize a stipulated
performance criterion. The performance criterion that governs the operation of the adap-
tive processor is chosen to reflect the steady-state performance characteristics that are of
concern. The most popular performance measures that have been employed include the
mean square error [9,28–31]; SNR ratio [6,14,32–34]; output noise power [35]; maximum
array gain [36,37]; minimum signal distortion [38,39]; and variations of these criteria
that introduce various constraints into the performance index [16,40–43]. In Chapter 3,
selected performance measures are formulated in terms of the signal characterizations
of (1.1)–(1.4). Solutions are found that determine the optimum choice for the complex
weight vector and the corresponding optimum value of the performance measure. The
operational signal environment plays a crucial role in determining the effectiveness of
the adaptive array to operate. Since the array configuration has pronounced effects on the
resulting system performance, it is useful to consider sensor spacing effects before pro-
ceeding with an analysis using an implicit description of such effects. The consideration
of array configuration is undertaken in Chapter 2.
10 CHAPTER 1 Introduction
FIGURE 1-2 t
Pulse modulated
carrier signal.
Carrier signal
where is the velocity of propagation of the transmitted signal. For underwater applications
the velocity of sound in water varies widely with temperature, although a nominal value
of 1,500 m/sec can be used for rough calculations. The velocity of electromagnetic wave
propagation in the atmosphere can be taken approximately to be the speed of light or
3 ×108 m/sec.
If the range discrimination capability between targets is to be rd , then the maximum
pulse length tmax (in the absence of pulse compression) is given by
2rd
tmax = (1.8)
It will be noted that rd also corresponds to the “blind range”— that is, the range within
which target detection is not possible. Since the signal bandwidth ∼ = 1/pulse length, the
range discrimination capability determines the necessary bandwidth of the transducers
and their associated electrical channels.
The transmitted pulses form a pulse train in which each pulse modulates a carrier fre-
quency as shown in Figure 1-2. The carrier frequency f 0 in turn determines the wavelength
of the propagated wavefront since
λ0 = (1.9)
f0
where λ0 is the wavelength. For sonar systems, frequencies in the range 100–100,000
Hz are commonly employed [44], whereas for radar systems the range can extend from
a few megahertz up into the optical and ultraviolet regions, although most equipment is
designed for microwave bands between 1 and 40 GHz. The wavelength of the propagated
wavefront is important because the array element spacing (in units of λ) is an important
parameter in determining the array pattern.
control for both missiles and guns against both airborne and ground (or sea) targets,
and additionally to provide navigational aid and perform reconnaissance. Current civil
aeronautical needs include air traffic control, collision avoidance, instrument approach
systems, weather sensing, and navigational aids. Additional applications in the fields of
law enforcement, transportation, and Earth resources are just beginning to grow to sizable
proportions [48].
Figure 1-3 is a block diagram of a typical radar system. These major blocks and their
corresponding functions are described in Table 1-1 [46]. The antenna, receiver, and signal
processing blocks are of primary interest for our purposes, and these are now each briefly
discussed in turn.
Block Function
Transmitter Generates high power RF waveform
Antenna Determines direction and shape of transmit-and-receive beam
Receiver Provides frequency conversion and low-noise amplification
Signal processing Provides target detections, target and clutter tracking,
and target trajectory estimates
Display Converts processed signals into meaningful tactical information
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:5 12
12 CHAPTER 1 Introduction
during World War II) operated in the very high frequency (VHF) and ultra high frequency
(UHF) bands. Sometimes, a parabolic dish was used, such as the FuG 65 Wurzburg Riese
radar antenna in Figure 1-4. Its 3 m parabolic dish operated at 560 MHz and was used
to guide German intercept fighters during WWII [53]. Arrays, such as the SCR-270 in
Figure 1-5, were used by the United States in WWII for air defense [54]. It has four rows
of eight dipoles that operate at 110 MHz. After WWII, radar and communications systems
began operating at higher frequencies. Different types of antennas were tried for various
systems. A microwave lens was used in the Nike AJAX MPA-4 radar shown in Figure 1-6
[55]. In time, phased array antennas became small enough to place in the nose of fighter
airplanes. In the 1980s, the AN/APG-68 (Figure 1-7) was used in the F-16 fighter [56]. The
array is a planar waveguide with slots for elements. Active electronically scanned arrays
(AESA) provide fast wide angle scanning in azimuth and elevation and include advanced
transmit/receive modules [57]. An example is the AN/APG-77 array for the F-22 fighter
shown in Figure 1-8.
FIGURE 1-5
SCR-270 antenna
array (Courtesy of
the National
Electronics
Museum).
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:5 13
FIGURE 1-6
Waveguide lens
antenna for the
Nike AJAX MPA-4
radar (Courtesy of
the National
Electronics
Museum).
The two most common forms of antenna arrays for radar applications are the linear
array and the planar array. A linear array consists of antenna elements arranged in a straight
line. A planar array, on the other hand, is a two-dimensional configuration in which the
antenna elements are arranged to lie in a plane. Conformal arrays lie on a nonplanar
surface. The linear array generates a fan beam that has a broad beamwidth in one plane
and a narrow beamwidth in the orthogonal plane. The planar array is most frequently used
in radar applications where a pencil beam is needed. A fan-shaped beam is easily produced
by a rectangular-shaped aperture. A pencil beam may easily be generated by a square- or
circular-shaped aperture. With proper weighting, an array can be made to simultaneously
generate multiple search or tracking beams with the same aperture.
Array beam scanning requires a linear phase shift across the elements in the array.
The phase shift is accomplished by either software in a digital beamformer or by hardware
phase shifters. A phase shifter is often incorporated in a transmit/receive module. Com-
mon phase-shifter technology includes ferroelectrics, monolithic microwave integrated
FIGURE 1-7
AN/APG-68 array
(Courtesy of
Northrop Grumman
and available at the
National Electronics
Museum).
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:5 14
14 CHAPTER 1 Introduction
FIGURE 1-8
AN/APG-77 array
(Courtesy of
Northrop Grumman
and available at the
National Electronics
Museum).
[Link] Receivers
A receiver design based on a matched filter or a cross-correlator maximizes the SNR in
the linear portion of the receiver. Different types of receivers that have been employed
in radar applications include the superheterodyne, superregenerative, crystal video, and
tuned radio frequency (TRF) [58]. The most popular and widely applied receiver type is
the superheterodyne, which is useful in applications where simplicity and compactness
are especially important. A received signal enters the system through the antenna, then
passes through the circulator and is amplified by a low-noise RF amplifier. Following RF
amplification, a mixer stage is entered to translate the RF to a lower intermediate frequency
(IF) where the necessary gain is easier to obtain and filtering is easier to synthesize. The
gain and filtering are then accomplished in an IF amplifier section.
(2) extraction of information from the received waveform to obtain target trajectory data
such as position and velocity. Detecting a signal imbedded in a noise field is treated by
means of statistical decision theory. Similarly, the problem of the extraction of information
from radar signals can be regarded as a problem concerning the statistical estimation of
parameters.
FIGURE 1-9
Control Sonar receiver block
diagram.
16 CHAPTER 1 Introduction
a
c
b
g
d
f
FIGURE 1-10 Various acoustic transducers. a: Seabed mapping—20 kHz multibeam receive
array module; 8 shaded elements per module; 10 modules per array. b: Subbottom profiling
(parametric)—200 kHz primary frequency; 25 kHz secondary frequency. c: Port and harbor
Security—curved 100 kHz transmit/receive array. d. Obstacle avoidance—10 × 10 planar re-
ceive array with curved transmitter. e. ACOMMS—Broadband piezocomposite transducers for
wideband communication signals. f. AUV FLS—high-frequency, forward-looking sonar array.
g. Mine hunting—10 × 10 transmit/receive broadband array; available with center frequen-
cies between 20 kHz to 1MHz h. Side scan—multibeam transmit/receive array (Courtesy of
Materials Systems Inc.).
A transmitting power ranging from a few acoustic watts up to several thousand acoustic
watts at ocean depths up to 20,000 ft can be achieved [61]. Figure 1-10 shows a sampling
of different acoustic transducers manufactured by Materials Systems Inc. for various
applications.
The basic physical mechanisms most widely used in transducer technology include
the following [62]:
1. Moving coil. This is long familiar from use as a loudspeaker in music reproduction
systems and used extensively in water for applications requiring very low frequencies.
2. Magnetorestrictive. Magnetic materials vibrate in response to a changing magnetic
field. Magnetorestrictive materials are rugged and easily handled, and magnetorestric-
tive transducers were highly developed and widely used during World War II.
3. Piezoelectric. The crystalline structure of certain materials results in mechanical vi-
bration when subjected to an alternating current or an oscillating electric field. The
relationship between mechanical strain and electric field is linear. Certain ceramic ma-
terials also exhibit a similar effect and have outstanding electromechanical properties.
Consequently, over the last decade the great majority of underwater sound transducers
have been piezoceramic devices that can operate over a wide frequency band and have
both high sensitivity and high efficiency.
4. Electrostrictive. Similar to piezoelectric but has a nonlinear relationship between me-
chanical strain and electric field.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:5 17
5. Electrostatic. These capacitive transducers use the change in force between two charged
parallel plates due to mechanical movement. These have found use with MEMS but
not in underwater acoustics.
6. Variable reluctance and hydroacoustic transducers have also been used for certain
experimental and sonar development work, but these devices have not challenged the
dominance of piezoceramic transducers for underwater sound applications [59].
FIGURE 1-11
Picture of a
100-element receive
acoustic array
manufactured in
four layers:
matching layer,
piezocomposite,
flex circuit, and
absorbing back
(Courtesy of
Materials Systems
Inc.).
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:5 18
18 CHAPTER 1 Introduction
FIGURE 1-12 Left—Port and harbor surveillance piezocomposite array; 100 kHz
transmit/receive array. Right—Forward-looking piezocomposite sonar for AUV; piezocompos-
ite facilitates broad bandwidth, high element count arrays, and curved geometries (Courtesy
of Materials Systems Inc.).
Surface ships often use a bubble-shaped bow dome in which a cylindrical array is
placed like that shown in Figure 1-15. This array uses longitudinal vibrator-type elements
composed of a radiating front end of light weight and a heavy back mass, with a spring
having active ceramic rings or disks in the middle [63]. The axial symmetry of a cylindrical
array renders beam steering fairly simple with the azimuth direction in which the beam
is formed, because the symmetry allows identical electronic equipment for the phasing
and time delays required to form the beam. Planar arrays do not have this advantage,
FIGURE 1-15
Cylindrical sonar
array used in
bow-mounted
dome.
since each new direction in space (whether in azimuth or in elevation) requires a new
combination of electronic equipment to achieve the desired pointing.
A spherical array is the ideal shape for the broadest array coverage in all directions.
Spherical arrays like that shown in the diagram of Figure 1-16 have been built with a
diameter of 15 ft and more than 1,000 transducer elements. This spherical arrangement
can be integrated into the bow of a submarine by means of an acoustically transparent
dome that provides minimum beam distortion. For instance, the bow dome of a Virginia
class submarine is a 25 ton hydrodynamically shaped composite structure that houses
a sonar transducer sphere. The bow dome is 21 feet tall and has a maximum diameter
of 26 feet. A two-inch thick, single-piece rubber boot is bonded to the dome to enhance
acoustic performance. Minimal sound energy absorption and reflection properties inherent
in the rubber material minimally reflect and absorb acoustic signals. Figure 1-17 shows
a spherical microphone array that was constructed by placing a rigid spherical array at
the center of a larger open spherical array [66]. Both arrays have 32 omnidirectional
microphones and a relatively constant directivity from about 900 Hz to 16 kHz.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:5 20
20 CHAPTER 1 Introduction
FIGURE 1-16
Spherical array
having 15 ft diameter
and more than 1,000 Transducer
transducer elements
elements.
FIGURE 1-17 A
dual, concentric
SMA. (A. Parthy,
C. Jin, and A. van
Schaik, “Acoustic
holography with a
concentric rigid and
open spherical
microphone array,”
IEEE International
Conference on
Acoustics, Speech
and Signal
Processing, 2009,
pp. 2173–2176.)
[Link] Beamformer
Beamforming ordinarily involves forming multiple beams from multielement arrays
through the use of appropriate delay and weighting matrices. Such beams may be di-
rectionally fixed or steerable. After that, sonar systems of the 1950s and 1960s consisted
largely of independent sonar sets for each transducer array. More recently, the sophisti-
cated use of multiple sensors and advances in computer technology have led to integrated
sonar systems that allow the interaction of data from different sensor arrays [67,68]. Such
integrated sonar systems have software delays and weighting matrices, thereby general-
izing the structure of digital time domain beamformers. Consequently several units of a
single (programmable) beamformer design may be used for all the arrays in an integrated
system. Furthermore, programmable beamformer matrices make it possible to adapt the
receive pattern to the changing structure of the masking noise background.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:5 21
1.6.2 Part 2
The heart of the adaptive capability within an adaptive array system is the adaptive algo-
rithm that adjusts the array pattern in response to the signal information found at the sensor
element outputs. Part 2, including Chapters 4 through 8, introduces different classes of
adaptation algorithms. In some cases adaptation algorithms are selected according to the
kind of signal information available to the receiver:
1. The desired signal is known.
2. The desired signal is unknown, but its direction of arrival is known.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:5 22
22 CHAPTER 1 Introduction
1.6.3 Part 3
The adaptive array operating conditions considered so far were nonideal only in that in-
terference signals were present with which the array had to contend. In actual practice,
however, the effects of several other nonideal operating conditions often result in unaccept-
able degradation of array performance unless compensation of such effects is undertaken.
Such nonideal operating conditions include processing of broadband signals, multipath
effects, channel mismatching, and array propagation delay effects. Compensation for these
factors by means of tapped delay-line processing is considered, and the question of how
to design a tapped delay line to achieve a desired degree of compensation is addressed.
Finally, current trends in adaptive array research that provide an indication of the direction
that future developments are likely to take are discussed.
1.8 Problems 23
(1) small changes in interference signal bearing, (2) small errors in the adaptive weight val-
ues, and (3) statistical fluctuations of measured correlations due to finite integration time.
A lightweight four-element adaptive array using hybrid microwave integrated circuitry
and weighing only 1 pound, intended for communication applications, was built and tested
[71]. This unit employed a null-steering algorithm appropriate for a coherent sidelobe
canceller and succeeded in forming broadband nulls over a 60–100 MHz bandwidth having
a cancellation depth of 25–30 dB under weak desired signal and strong interference signal
conditions. To attain this degree of interference signal cancellation, it was essential that
the element channel circuitry be very well matched over a 20% bandwidth.
Another experimental four-element adaptive array system for eliminating interference
in a communication system was also tested [48]. Pattern nulls of 10–20 db for suppressing
interference signals over a 200–400 MHz band were easily achieved so long as the desired
signal and interference signal had sufficient spatial separation (greater than the resolution
capability of the antenna array), assuming the array has no way to distinguish between
signals on the basis of polarization. Exploiting polarization differences between desired
and interference signals by allowing full polarization flexibility in the array, an interference
signal located at the same angle as the desired signal can be suppressed without degrading
the reception of the desired signal. Yet another system employing digital control was
developed for UHF communications channels and found capable of suppressing jammers
by 20–32 dB [72].
In summary, interference suppression levels of 10–20 dB are consistently achieved
in practice. It is more difficult but nevertheless practicable to achieve suppression levels
of 20–35 dB and usually very difficult to form cancellation nulls greater than 35 dB in a
practical operating system.
The rapid development of digital technology is presently having the greatest impact
on signal reception systems. The full adaptation of digital techniques into the processing
and interpretation of received signals is making possible the realization of practical sig-
nal reception systems whose performance approaches that predicted by theoretical limits.
Digital processors and their associated memories have made possible the rapid digestion,
correlation, and classification of data from larger search volumes, and new concepts in the
spatial manipulation of signals have been developed. Adaptive array techniques started
out with limited numbers of elements in the arrays, and the gradual increase in the num-
bers of elements and in the sophistication of the signal processing will likely result in an
encounter with techniques employed in optical and acoustical holography [69,73]. Holog-
raphy techniques are approaching such an encounter from the other direction, since they
start out with a nearly continuous set of spatial samples (as in optical holography) and
move down to a finite number of samples (in the case of acoustic holography).
1.8 PROBLEMS
1. Radar Pulse Waveform Design Suppose it is desired to design a radar pulse waveform that
would permit two Ping-Pong balls to be distinguished when placed only 6.3 cm apart in range
up to a maximum range from the radar antenna of 10 m.
(a) What is the maximum PRF of the resulting pulse train?
(b) What bandwidth is required for the radar receiver channel?
(c) If it is desired to maintain an array element spacing of d = 2 cm where d = λ0 /2, what
pulse carrier frequency should the system be designed for?
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:5 24
24 CHAPTER 1 Introduction
2. Sonar Pulse Carrier Frequency Selection In the design of an actual sonar system many
factors must be considered—all the sonar parameters (e.g., source level, target strength) and the
environment parameters. The effect of environmental parameters depends largely on frequency.
Suppose in a highly oversimplified example that only the factors of transmission loss (due to
attenuation) and ambient noise are of concern. Let the attenuation coefficient α be given by
1
log10 (α) = [−21 + 5 log10 ( f )]
4
Furthermore, let the ambient noise spectrum level N0 be given by
1
10 log10 (N0 ) = [20 − 50 log10 ( f )]
3
If the cost to system performance is given by J = C1 α + C2 N0 where C1 and C2 denote the
relative costs of attenuation and noise to the system, what value of pulse carrier frequency f
should be selected to optimize the system performance?
1.9 REFERENCES
[1] L. C. Van Atta, “Electromagnetic Reflection,” U.S. Patent 2908002, October 6, 1959.
[2] IEEE Trans. Antennas Propag. (Special Issue on Active and Adaptive Antennas), Vol. AP-12,
March 1964.
[3] D. L. Margerum, “Self-Phased Arrays,” in Microwave Scanning Antennas, Vol. 3, Array
Systems, edited by R. C. Hansen, Academic Press, New York, 1966, Ch. 5.
[4] P. W. Howells, “Intermediate Frequency Sidelobe Canceller,” U.S. Patent 3202990, August
24, 1965.
[5] P. W. Howells, “Explorations in Fixed and Adaptive Resolution at GE and SURC,” IEEE
Trans. Antennas Propag., Special Issue on Adaptive Antennas, Vol. AP-24, No. 5, pp. 575–
584, September 1976.
[6] S. P. Applebaum, “Adaptive Arrays,” Syracuse University Research Corporation, Rep. SPL
TR66-1, August 1966.
[7] B. Widrow, “Adaptive Filters I: Fundamentals,” Stanford University Electronics Laboratories,
System Theory Laboratory, Center for Systems Research, Rep. SU-SEL-66-12, Tech. Rep.
6764-6, December 1966.
[8] B. Widrow, “Adaptive Filters,” in Aspects of Network and System Theory, edited by R. E.
Kalman and N. DeClaris, Holt, Rinehard and Winston, New York, 1971, Ch. 5.
[9] B. Widrow, P. E. Mantey, L. J. Griffiths, and B. B. Goode, “Adaptive Antenna Systems,” Proc.
IEEE, Vol. 55, No. 12, p. 2143–2159, December 1967.
[10] L. Spitzer, Jr., “Basic Problems of Underwater Acoustic Research,” National Research Council,
Committee on Undersea Warfare, Report of Panel on Underwater Acoustics, September 1,
1948, NRC CUW 0027.
[11] R. H. Bolt and T. F. Burke, “A Survey Report on Basic Problems of Acoustics Research,”
Panel on Undersea Warfare, National Research Council, 1950.
[12] H. Mermoz, “Filtrage adapté et utilisation optimale d’une antenne,” Proceedings NATO Ad-
vanced Study Institute, September 14–26, 1964. “Signal Processing with Emphasis on Un-
derwater Acoustics” (Centre d’Etude des Phénomenès Aléatoires de Grenoble, Grenoble,
France), 1964, pp. 163–294.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:5 25
1.9 References 25
26 CHAPTER 1 Introduction
1.9 References 27
[56] M. I. Skolnik, Introduction to Radar Systems, McGraw-Hill, New York, 1962, Ch. 2.
[57] C. H. Sherman, “Underwater Sound - A Review: I. Underwater Sound Transducers,” IEEE
Trans. Sonics and Ultrason., Vol. SU-22, No. 5, September 1975, pp. 281–290.
[58] A. A. Winder, “Underwater Sound - A Review: II. SKonar System Technology,” IEEE Trans.
Sonics and Ultrason., Vol. SU-22, No. 5, September 1975, pp. 291–332.
[59] E. A. Massa, “Deep Water trnasducers for Ocean Engineering Applications,” Proceedings
IEEE, 1966 Ocean Electronics Symposium, August 29–31, Honolulu, HI, pp. 31–39.
[60] C. H. Sherman and J. L. Butler, Transducers and Arrays for Underwater Sound, New York,
Springer, 2007.
[61] T. F. Hueter, “Twenty Years in Underwater Acoustics: Generation and Reception,” J. Acoust.
Soc. Am., Vol. 51, No 3 (Part 2), March 1972, pp. 1025–1040.
[62] J. S. Hickman, “Trends in Modern Sonar Transducer Design,” Proceedings of the Twenty-
Second National Electronics Conference, October 1966, Chicago, Ill.
[63] E. L. Carson, G. E. Martin, G. W. Benthien, and J. S. Hickman, “Control of Element
Velocity Distributions in Sonar Projection Arrays,” Proceedings of the Seventh Navy Science
Symposium, May 1963, Pensacola, FL.
[64] A. Parthy, C. Jin, and A. van Schaik, “Acoustic Holography with a Concentric Rigid and
Open Spherical Microphone Array,” IEEE International Conference on Acoustics, Speech
and Signal Processing, 2009, pp. 2173–2176.
[65] J. F. Bartram, R. R. Ramseyer, and J. M Heines, “Fifth Generation Digital Sonar Signal
Processing,” IEEE EASCON 1976 Record, IEEE Electronics and Aerospace Systems
Convention, September 26–29, Washington, D.C., pp. 91A–91G.
[66] T. F. Hueter, J. C. Preble, and G. D. Marshall, “Distributed Architecture for Sonar Systems
and Their Impact on Operability,” IEEE EASCON 1976 Record, IEEE Electronics and
Aerospace Convention, September 16–29, Washington, D.C., pp. 96A–96F.
[67] V. C. Anderson, “The First Twenty Years of Acoustic Signal Processing,” J. Acoust. Soc.
Am., Vol. 51, No. 3 (Part 2), March 1972, pp. 1062–1063.
[68] M. Federici, “On the Improvement of Detection and Precision Capabilities of Sonar Systems,”
Proceedings of the Symposium on Sonar Systems, July 9–12, 1962, University of Birmingham,
pp. 535–540, Reprinted in J. Br. I.R.E., Vol. 25, No.6, June 1963.
[69] G. G. Rassweiler, M. R. Williams, L. M. Payne, and G. P. Martin, “A Miniaturized Lightweight
Wideband Null Steerer,” IEEE Trans. Ant. & Prop., Vol. AP-24, September 1976, pp. 749–754.
[70] W. R. Jones, W. K. Masenten, and T. W. Miller, “UHF Adaptive Array Processor Development
for Naval Communications,” Digest AP-S Int. Symp., IEEE Ant. & Prop Soc., June 2–22,
1977, Stanford, CA., IEEE Cat. No. 77CH1217-9 AP, pp. 394–397.
[71] G. Wade, “Acoustic Imaging with Holography and Lenses,” IEEE Trans. Sonics & Ultrason.,
Vol. SU-22, No. 6, November 1975, pp 385–394.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:5 28
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 29
CHAPTER
To understand why an array of sensor elements has the potential to improve the reception
of a desired signal in an environment having several sources of interference, it is neces-
sary to understand the nature of the signals as well as the properties of an array of sensor
elements. Furthermore, the types of elements and their arrangement impact the adaptive
array performance. To gain this understanding the desired signal characteristics, inter-
ference characteristics, and signal propagation effects are first discussed. The properties
of sensor arrays are then introduced, and the possibility of adjusting the array response
to enhance the desired signal reception is demonstrated. Trade-offs for linear and planar
arrays are presented to aid the designer in finding an economical array configuration.
In arriving at an adaptive array design, it is necessary to consider the system constraints
imposed by the nature of the array, the associated system elements with which the designer
has to work, and the system requirements the design is expected to satisfy. Adaptive array
requirements may be classified as either (1) steady-state or (2) transient depending on
whether it is assumed the array weights have reached their steady-state values (assuming
a stationary signal environment) or are being adjusted in response to a change in the signal
environment. If the system requirements are to be realistic, they must not exceed the
predicted theoretical performance limits for the adaptive array system being considered.
Formulation of the constraints imposed by the sensor array is addressed in this chapter.
Steady-state performance limits are considered in Chapter 3. The formulation of transient
performance limits, which is considerably more involved, is addressed in Part 2. For the
performance limits of adaptive array systems to be analyzed, it is necessary to develop a
generic analytic model for the system. The development of such an analytic model will
be concerned with the signal characteristics and the subsequent processing necessary to
obtain the desired system response.
29
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 30
signal model possessing some of the signal characteristics of the unknown signal can be
adopted in certain circumstances. Communications signals pass through environments that
randomly add scattering and noise to the desired signal. Thermal sensor noise, ambient
noise, and interference signal sources are also random in nature. These noises typically
arise from the combined effect of many small independent sources, and application of the
central limit theorem of statistics [6] permits the designer to model the resulting noise
signal as a Gaussian (and usually stationary) random process. Quite frequently, the phys-
ical phenomena responsible for the randomness in the signals of concern are such that it
is plausible to assume a Gaussian random process. The statistical properties of Gaussian
signals are particularly convenient because the first- and second-moment characteristics
of the process provide a complete characterization of the random signal.
The statistical properties of the signal are not always known, so a selected deterministic
signal is used instead. This deterministic signal does not have to be a perfect replica of the
desired signal. It needs only to be reasonably well correlated with the desired signal and
uncorrelated with the interference signals.
Sometimes the desired signal is known, as in the case of a coherent radar with a target
of known range and character. In other cases, the desired signal is known except for the
presence of uncertain parameters such as phase and signal energy. The case of a signal
known except for phase occurs with an ordinary pulse radar having no integration and
with a target of known range and character. Likewise, the case of a signal known except
for phase and signal energy occurs with a pulse radar operating without integration and
with a target of unknown range and known character. The frequency and bandwidth of
communication signals are typically known, and such signals may be given a signature by
introducing a pilot signal.
For a receive array composed of N sensors, the received waveforms correspond to N
outputs, x1 (t), x2 (t), . . . , x N (t), which are placed in the received signal vector x(t) where
⎡ ⎤
x1 (t)
⎢ x (t) ⎥
⎢ 2 ⎥
x(t) = ⎢ . ⎥ for 0 ≤ t ≤ T (2.1)
⎣ .. ⎦
x N (t)
over the observation time interval. The received signal vector is the sum of the desired
signal vector, s(t), and the noise component, n(t).
techniques assume slowly varying Gaussian ambient noise fields. Adaptive processors
designed for Gaussian noise are distinguished by the pleasant fact that they depend only
on second-order noise moments. Consequently, when non-Gaussian noise fields must be
dealt with, the most convenient approach is to design a Gaussian-equivalent suboptimum
adaptive system based on the second-order moments of the non-Gaussian noise field.
In general, an adaptive system works best when the variations that occur in the noise
environment are slow.
where the ith component of m(t) is m i (t) and represents the propagation effects from the
source to the ith sensor as well as the response of the ith sensor. For the ideal case of
nondispersive propagation and distortion-free sensors, then m i (t) is a simple time delay
δ(t − τi ), and the desired signal component at each sensor element is identical except for
a time delay so that (2.4) can be written as
⎡ ⎤
s(t − τ1 )
⎢ s(t − τ2 ) ⎥
⎢ ⎥
s(t) = ⎢ .. ⎥ (2.5)
⎣ . ⎦
s(t − τ N )
When the array is far from the source, then the signal is represented by a plane wave from
the direction α as shown in Figure 2-1 (where α is taken to be a unit vector). In this case,
FIGURE 2-1 z
Three-dimensional
array with plane
Wa
wave signal
ve
propagation.
fro
nt
Sensor s(t)
locations
r1
r2
a
x
r4
r3
y
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 33
q x(t)
d sin q
Sensor 2 Sensor 1
Array axis
signal x(t) impinge on the two sensor elements in a plane containing the two elements and
the signal source from a direction θ with respect to the array normal. Figure 2-2 shows
that element 2 receives the signal τ after element 1.
d sin θ
τ= (2.8)
Let the array output signal y(t) be given by the sum of the two sensor element signals so
that
If x(t) is a narrowband signal having center frequency f 0 , then the time delay τ
corresponds to a phase shift of 2π(d/λ0 ) sin θ radians, where λ0 is the wavelength corre-
sponding to the center frequency,
λ0 = (2.10)
f0
The overall array response is the sum of the signal contributions from the two array
elements. That is,
2
y(t) = x(t)e j (i−1)ψ (2.11)
i=1
where
The directional pattern or array factor of the array (sensitivity to signals vs. angle at
a specified frequency) may be found by considering only the term
2
AF(θ ) = e j (i−1)ψ (2.13)
i=1
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 35
The normalized array factor in decibels for the two elements is then given by
|AF(θ )2 |
AF(θ)(decibels) = 10 log10 (2.14)
22
so that the peak value is unity or 0 dB. A plot of the normalized AF(θ ) for this two-
element example is given in Figure 2-3 for d/λ0 = 0.5, 1.0, and 1.5. From Figure 2-3a
it is seen that for d/λ0 = 0.5 there is one principal lobe (or main beam) having a 3 dB
beamwidth of 60◦ and nulls at θ = ±90◦ off broadside. The nulls at θ = ±90◦ occur
because at that direction of arrival the signal wavefront must travel exactly λ0 /2 between
0 0
−5 −5
−10 −10
AF (dB)
AF (dB)
−15 −15
−20 −20
−25 −25
−30 −30
−90 −45 0 45 90 −90 −45 0 45 90
q (degrees) q (degrees)
(a) (b)
−5
−10
AF (dB)
−15
−20
−25
−30
−90 −45 0 45 90
q (degrees)
(c)
FIGURE 2-3 Array beam patterns for two-element example. (a) d/λ0 = 0.5. (b) d/λ0 = 1.
(c) d/λ0 = 1.5.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 36
the two sensors, which corresponds to a phase shift of 180◦ between the signals appearing
at the two sensors and therefore yields exact cancellation of the resulting phasor sum.
If the element spacing is less than 0.5λ0 , then exact cancellation at θ = ±90◦ does not
result, and in the limit as the element spacing approaches zero (ignoring mutual coupling
effects) the directional pattern becomes an isotropic point source pattern. There is very
little difference in the directional pattern between a single element and two closely spaced
elements (less than λ/4 apart); consequently, arrays employing many elements very closely
spaced are considered “inefficient” if it is desired to use as few array elements as possible
for a specified sidelobe level and beamwidth. If the element spacing increases to greater
than 0.5λ0 , the two pattern nulls migrate in from θ = ±90◦ , occurring at θ = ±30◦ when
d = λ0 , as illustrated in Figure 2-3b. The nulls at θ = ±30◦ in Figure 2-3b occur because
at that angle of arrival the phase path difference between the two sensors is once again
180◦ , and exact cancellation results from the phasor sum. Two sidelobes at θ = ±90◦
having an amplitude equal to the principal lobe at θ = 0◦ . They appear because the phase
path difference between the two sensors is then 360◦ , two phasors exactly align, and
the array response is the same as for broadside angle of arrival. As the element spacing
increases to 1.5λ0 , the main lobe beamwidth decreases still further, thereby improving
resolution, the two pattern nulls migrate further in, and two new nulls appear at ±90◦ ,
as illustrated in Figure 2-3c. Further increasing the interelement spacing results in the
appearance of even more pattern nulls and sidelobes and a further decrease in the main lobe
beamwidth.
When N > 2, the array factor for N isotropic point sources becomes
N
2π
[xn sin θ cos φ+yn sin θ sin φ+z n cos θ ]
AF(θ, φ) = wne j λ (2.15)
n=1
where
Three common examples of multiple element arrays used for adaptive nulling are [11]
N
2π
xn sin θ cos φ
Linear array along the x-axis: AF(θ, φ) = wne j λ (2.16)
n=1
N
2π
[xn sin θ cos φ+yn sin θ sin φ]
Planar array in the x-y plane: AF(θ, φ) = wne j λ (2.17)
n=1
N
2π
[xn sin θ cos φ+yn sin θ sin φ+z n cos θ ]
3-dimensional array: AF(θ, φ) = wne j λ (2.18)
n=1
The sensors take samples of the signals incident on the array. As long as the sensors
take two samples in one period or wavelength, then the received signal is adequately
reproduced.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 37
The directional pattern in a plane containing the array may therefore be found by consid-
ering the array factor
N
AF(θ ) = w i e j (i−1)ψ (2.20)
i=1
When w i = 1.0, the array is called a “uniform array.” The uniform array has an array
factor given by
Nψ
sin 2
AF(θ ) = (2.21)
ψ
sin 2
If mutual coupling is ignored or if the element patterns are averaged, the individual element
patterns are all the same, so the array pattern is the product of an element pattern times
the array factor.
N
AF(θ) = F( f 0 , θ ) w i e j (i−1)ψ (2.25)
i=1
0 0
−5 −5
−10 −10
AF (dB)
AF (dB)
−15 −15
−20 −20
−25 −25
−30 −30
−90 −45 0 45 90 −90 −45 0 45 90
q (degrees) q (degrees)
(a) (b)
FIGURE 2-4 inear array beam patterns for d/λ0 = 0.5. (a) Three-element array.
(b) Four-element array.
3. Multiplying the array pattern resulting from step 2 by the beam pattern of the individual
elements of the original array
Maintaining the interelement spacing at d/λ0 = 0.5 and increasing the number of
point sources, the normalized array directional (or beam) pattern may be found from
(2.21), and the results are shown in Figure 2-4 for three and four elements. It is seen that
as the number of elements increases the main lobe beamwidth decreases, and the number
of sidelobes and pattern nulls increases.
To illustrate how element spacing affects the directional pattern for a seven-element
linear array, Figures 2-5a through 2.5d show the directional pattern in the azimuth plane
for values of d/λ0 ranging from 0.1 to 1.0. So long as d/λ0 is less than 17 , the beam pattern
has no exact nulls as the −8.5 dB null occurring at θ = ±90◦ for d/λ0 = 0.1 illustrates.
As the interelement spacing increases beyond d/λ0 = 17 , pattern nulls and sidelobes (and
grating lobes) begin to appear, with more lobes and nulls appearing as d/λ0 increases and
producing an interferometer pattern. When d/λ0 = 1 the endfire sidelobes at θ = ±90◦
have a gain equal to the main lobe since the seven signal phasors now align exactly and
add coherently.
Suppose for the linear array of Figure 2-6 that a phase shift (or an equivalent time
delay) of δ is inserted in the second element of the array, a phase shift of 2δ in the third
element, and a phase shift of (n − 1)δ in each succeeding nth element. The insertion of
this sequence of phase shifts has the effect of shifting the principal lobe (or main beam)
by
−1 1 λ0
θs = −sin δ (2.26)
2π d
so the overall directional pattern has in effect been “steered” by insertion of the phase-
shift sequence. Figure 2-7 is the array factor in Figure 2-5c steered to θs = −30◦ using
δ = −90◦ .
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 39
0
0
−5
−5
−10
−10
AF (dB)
AF (dB)
−15
−15
−20 −20
−25 −25
−30 −30
−90 −45 0 45 90 −90 −45 0 45 90
q (degrees) q (degrees)
(a) (b)
0 0
−5 −5
−10 −10
AF (dB)
AF (dB)
−15 −15
−20 −20
−25 −25
−30 −30
−90 −45 0 45 90 −90 −45 0 45 90
q (degrees) q (degrees)
(c) (d)
FIGURE 2-5 Seven-element linear array factors. (a) d/λ0 = 0.1. (b) d/λ0 = 0.2. (c) d/λ0 =
0.5. (d) d/λ0 = 1.0.
The gain of an array factor determines how much the signal entering the main beam
is magnified by the array. For a radar/sonar system the received power is given by [12]
Pt G 2 λ2 σ
Pr = (2.27)
(4π )3 R 4
where
Pt = power transmitted
G = array gain
σ = target cross section
R = distance from array to target
and for a communications system is given by Friis transmission formula [13]
Pt G t G r λ2
Pr = (2.28)
(4π R)2
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 40
FIGURE 2-6
Seven-element
q
linear array steered
with phase-shift
elements.
d = l 0 /2 d
Sensor
elements
6d 5d 4d 3d 2d d
Array
output
y(t )
FIGURE 2-7 0
Seven-element
linear array factor
steered to θs = 30◦ −5
with δ = −90◦ .
−10
AF (dB)
−15
−20
−25
−30
−90 −45 0 45 90
q (degrees)
The array gain is the ratio of the power radiated in one direction to the power delivered
to the array. Directivity is similar to gain but does not include losses in the array. As a
result, directivity is greater than or equal to gain. Realized gain includes the mismatch
between the array and the transmission line feeding the array. If gain is not written as a
function of angle, then G is the maximum gain of the array. Gain is usually expressed
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 41
in dB as
G dB = 10 log10 G = 10 log G (2.29)
The directivity of an array is found by solving
4π |AFmax |2
D= (2.30)
π
2π
|AF(θ, φ)|2 sin θ dθ dφ
0 0
If the elements are spaced 0.5λ apart in a linear array, then the directivity formula simplifies
to
2
N
w n
D = n=1 (2.31)
N
|w n | 2
n=1
The z-transform converts the linear array factor into a polynomial using the substitu-
tion
z = e jψ (2.32)
Substituting z into (2.16) yields a polynomial in z
N
AF = w n z (n−1) = w 1 + w 2 z + · · · + w N z N −1 (2.33)
n=1
p 0
−p /2
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 42
z7 + z6 + z5 + z4 + z3 + z2 + z + 1
= z + e jπ/4 z + e− jπ/4 z + e jπ/2 z + e− jπ/2 z + e j3π/4 z + e j3π/4 z + e jπ
(2.35)
FIGURE 2-9 z
Rectangular-shaped
planar array.
P(r, q, f)
r
Projection of
P(r, q, f) onto the
dx x–y plane
dy
Sensor x
elements
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 43
where
dx dy
ψx = 2π sin θ cos φ and ψ y = 2π sin θ sin φ (2.38)
λ0 λ0
The output signal now depends on both the projected azimuth angle, φ, and the elevation
angle, θ. The total sum of signal contributions from all array elements is given by
Nx Ny
y(t) = w i,k x(t)e j (i−1)ψx e j (k−1)ψ y (2.39)
i=1 k=1
If the weights are separable, w i,k = w i × w k , then the planar array factor is
where
⎫
Nx ⎪
⎪
⎪
⎪
AFx (θ, φ) = w i e j (i−1)ψx ⎬
i=1 (2.41)
and ⎪
⎪
Ny ⎪
⎪
⎭
AFy (θ, φ) = w k e j (k−1)ψ y
k=1
If w n = 1.0 for all n, then the array factor for rectangular spacing is
N x ψx N y ψy
sin 2
sin 2
AF = ψy
(2.43)
ψx
N x sin 2
N y sin 2
where
2π
ψx = dx sin θ cos φ
λ
2π
ψy = d y sin θ sin φ
λ
Usually, the beamwidth of a planar array is defined for two orthogonal planes. For
example, the beamwidth is usually defined in θ for φ = 0◦ and φ = 90◦ . As with a linear
array, nulls in the array factor of a planar array are found by setting the array factor equal
to zero. Unlike linear arrays, the nulls are not single points.
The directivity of a planar array can be found by numerically integrating (2.24). A
reasonable estimate for the 3 dB beamwidth in orthogonal planes is given by
32400
D= ◦ ◦ (2.44)
θ3dBφ=0 ◦ θ3dBφ=90◦
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 44
where
◦ ◦
θ3dBφ=0 ◦ = 3 dB beamwidth in degrees at φ= 0
◦ ◦
θ3dBφ=90 ◦ = 3 dB beamwidth in degrees at φ= 90
The array factor for a planar array with rectangular element spacing can be written as
Ny Nx dy
(n−1) dλx (u−u s )+(m−1) (v−vs )
AF = w mn e j2π λ (w) (2.45)
m=1 n=1
where
u = sin θ cos φ
v = sin θ sin φ
(2.46)
us = sin θs cos φs
vs = sin θs sin φs
N
2π
(xn sin θ cos φ+yn sin θ sin φ+z n cos θ )
AF = wne j λ (2.47)
n=1
The array curvature causes phase errors that distort the array factor unless phase weights at
the elements compensate. As an example, consider a 12-element linear array bent around
a cylinder of radius r = 3.6λ as shown in Figure 2-10. If no phase compensation is
applied, then the array factor looks like the dashed line in Figure 2-11. Adding a phase
delay of yn 2π/λ results in the array factor represented by the solid line in Figure 2-11.
The compensation restores the main beam and lowers the sidelobe level.
FIGURE 2-10 y
Twelve-element
linear array bent
around a cylinder of
radius r = 3.6λ.
(xn, yn)
f
x
r
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 45
10 FIGURE 2-11
Array factor for the
12-element
0 Uncompensated conformal array (with
and without phase
compensation).
− 10
AF (dB)
−20
−30 Compensated
−40
0 45 90 135 180
f (degrees)
d = l 0 /2
w1 + jw2 w1 + jw2
Array output
consider the two element array of isotropic point sources in Figure 2-12 in which a desired
signal arrives from the normal direction θ = 0◦ , and the interference signal arrives from
the angle θ = 30◦ . For simplicity, both the interference signal and the desired pilot signal
are assumed to be at the same frequency f 0 . Furthermore, assume that at the point exactly
midway between the array elements the desired signal and the interference are in phase
(this assumption is not required but simplifies the development). The output signal from
each element is input to a variable complex weight, and the complex weight outputs are
then summed to form the array output.
Now consider how the complex weights can be adjusted to enhance the reception of
p(t) while rejecting I (t). The array output due to the desired signal is
For the output signal of (2.48) to be equal to p(t) = Pe jω0 t , it is necessary that
w1 + w3 = 1
(2.49)
w2 + w4 = 0
The incident interfering noise signal exhibits a phase lead with respect to the array midpoint
when impinging on the element with complex weight w 3 + jw 4 of value 2π( 14 ) sin(π/6) =
π/4 and a phase lag when striking the other element of value −π/4. Consequently, the
array output due to the incident noise is given by
Now
1
e j (ω0 t−π/4) = √ [e jω0 t (1 − j)]
2
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 47
and
1
e j (ω0 t+π/4) = √ [e jω0 t (1 + j)]
2
so for the array noise response to be zero it is necessary that
w1 + w2 + w3 − w4 = 0
(2.51)
−w 1 + w 2 + w 3 + w 4 = 0
1 1 1 1
w1 = , w2 = − , w3 = , w4 = (2.52)
2 2 2 2
With the previous weights, the array will accept the desired signal while simultaneously
rejecting the interference.
While the complex weight selection yields an array pattern that achieves the desired
system objectives, it is not a very practical way of approaching adaptive nulling. The
method used in the previous example exploits the facts that there is only one directional
interference source, that the signal and interference sources are monochromatic, and that
a priori information concerning the frequency and the direction of arrival of each signal
source is available. A practical processor must work without detailed information about
the location, number, and nature of signal sources. Nevertheless, this example has demon-
strated that an adaptive algorithm achieves certain performance objectives by adjusting the
complex weights. The development of a practical adaptive array algorithm is undertaken
in Part 2.
FIGURE 2-13 p /2
Unit circle for the
eight-element 30 dB
Chebyshev taper.
p 0
−p/2
where
1
ζ = cosh−1 10sll/20 (2.54)
π
The factored polynomial is found using ψn . Once the ψn are known, the polynomial of
degree N − 1 in factored form easily follows. The polynomial coefficients are the array
amplitude weights.
As an example, designing an eight-element 30 dB Chebyshev array starts with calcu-
lating ζ = 1.1807. Substituting into (2.53) results in angular locations (in radians) on the
unit circle shown in Figure 2-13 and given by
Finally, the normalized amplitude weights are shown in Figure 2-14 with the corresponding
array factor in Figure 2-15.
The Chebyshev taper is not practical for large arrays, because the amplitude weights
are difficult to implement in a practical array. The Taylor taper [16] is widely used in the
design of low sidelobe arrays. The first n̄ − 1 sidelobes on either side of the main beam
are sll, whereas the remaining sidelobes decrease as 1/ sin θ . The Taylor taper moves the
first n̄ − 1 nulls on either side of the main beam to
⎧ "
⎨sin−1 ±λn̄ ζ 2 +(n−0.5)22 n < n̄
Nd ζ 2 +(n̄−0.5)
θn = (2.55)
⎩
sin−1 ±λnNd
n ≥ n̄
As with the Chebyshev taper, ζ is first found from (2.54). Next, the array factor nulls
are found using (2.55) and then substituting into ψn . Finally the factored array polynomial
is multiplied together to get the polynomial coefficients, which are also the Taylor weights.
A Taylor amplitude taper is also available for circular apertures as well [17].
The design of a 16-element 25 dB n̄ = 5 Taylor taper starts by calculating ζ = 1.1366.
Substituting into (2.55) results in angular locations (in radians) on the unit circle shown
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 49
1 FIGURE 2-14
Amplitude weights
for the eight-element
30 dB Chebyshev
0.8 taper.
0.6
Amplitude
0.4
0.2
0
1 2 3 4 5 6 7 8
Element
0 FIGURE 2-15
Array factor for the
eight-element 30 dB
−10 Chebyshev taper.
−20
AF (dB)
−30
−40
−90 −45 0 45 90
q (degrees)
Finally, the normalized amplitude weights are shown in Figure 2-17 with the corresponding
array factor in Figure 2-18.
FIGURE 2-16 p /2
Unit circle for 20 dB
Taylor n̄ = 5 taper.
p 0
−p /2
FIGURE 2-17 1
Amplitude weights
for Taylor 25 dB
n̄ = 5 taper.
0.8
0.6
Amplitude
0.4
0.2
0
0 5 10 15
Element
large arrays is impractical due to the complexity of the feed network and the inconsistency
in the mutual coupling environment of the elements. As a result, density tapering in a large
array is accomplished by “thinning” or removing active elements from the element lattice
in the array aperture.
If the desired amplitude taper function is normalized, then it looks like a probability
density function. A uniform random number generator assigns a random number to each
element. If that random number exceeds the desired normalized amplitude taper for that
element, then the element is passive in the array; otherwise, it is active. An active element
is connected to the feed network, whereas an inactive element is not. The advantages of
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 51
0 FIGURE 2-18
Array factor for 20
dB Taylor n̄ = 5
−10 taper.
−20
AF (dB)
−30
−40
−90 −45 0 45 90
q (degrees)
where sll2 is the power level of the average sidelobe level, and Nactive is the number of
active elements out of N elements in the array. An expression for the peak sidelobe level
of linear and planar arrays having half-wavelength spacing is found by assuming all the
sidelobes are within three standard deviations of the rms sidelobe level and has the form
[25]
N /2
P all sidelobes < sll2p 1 − e−sll p /sll
2 2
(2.57)
Statistical thinning was used until the 1990s when computer technology enabled
engineers to numerically optimize the thinning configuration of an array to find the desired
antenna pattern. Current approaches to array thinning include genetic algorithms [26] and
particle swarm optimization [27]. In addition, array thinning with realistic elements, like
dipoles [28], spirals [29], and thinned planar arrays [30], are used in place of point sources.
As an example, Figure 2-19 shows the array factor for a 100-element thinned linear
array with half-wavelength spacing. The thinning was performed using a 25 dB Taylor
n̄ = 5 amplitude taper as the probability density function. The array factor shown has
a peak sidelobe level of −18.5 dB below the peak of the main beam and has 77% of
the elements turned on. This array factor had the lowest relative sidelobe level of 20
independent runs. The genetic algorithm was used to find the thinning configuration that
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 52
FIGURE 2-19 0
Statistically thinned
100-element array.
−5
−10
AF (dB)
−15
−20
−25
−30
−90 −45 0 45 90
q (degrees)
FIGURE 2-20 0
Genetic algorithm
thinned 100-element
−5
array.
−10
AF (dB)
−15
−20
−25
−30
−90 −45 0 45 90
q (degrees)
yields the lowest relative sidelobe level. Figure 2-20 shows the optimized array factor with
a −23.2 dB peak relative sidelobe level when 70% of the elements are turned on.
0
Nulled Quiescent
−10
p /2
−20
AF (dB)
−30
p 0
−40
−90 −45 0 45 90
−p/2 q (degrees)
(a) (b)
0
Nulled
Quiescent
−10
p/2
−20
AF (dB)
−30
p 0
−40
−90 −45 0 45 90
−p /2 q (degrees)
(c) (d)
FIGURE 2-21 Null synthesis with the unit circle. a: Moving one zero from −3π/4 to
π sin(−38◦ ). b: Array factor nulls moves from −48.6◦ to −38◦ . c: Moving a second zero from
3π/4 to π sin(38◦ ). d: Second-array factor null moves from 48.6◦ to 38◦ .
A null can be placed at −38◦ using real weights by also moving the null in the array factor
at 48.6◦ to 38◦ to form complex conjugate pairs:
◦ ◦
z + e jπ/4 z + e− jπ/4 z + e jπ/2 z + e− jπ/2 z + e− jπ sin(38 ) z + e jπ sin(38 ) z + e jπ
= z7 + 0.30z 6 + 1.29z 5 + 0.59z 4 + 0.59z 3 + 1.29z 2 + 0.30z + 1 (2.59)
These examples of using null synthesis with the unit circle illustrate several points
about adaptive nulling:
1. There are an infinite number of ways to place a null at −38◦ . After one of the zeros
is placed at the desired angle, then the other zeros can be redistributed as well. Thus,
placing a null at a given location does not require a unique set of weights unless further
restrictions are applied.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 54
2. Some weights produce greater distortion to the array factor than other weights. The
weight selection impacts the gain of the array factor as well as the sidelobe levels and
null positions.
3. An N element array with complex weights can uniquely place N − 1 nulls. If the
weights are real, then only N /2 − 1 nulls can be uniquely placed.
4. Adaptive algorithms maneuver the nulls on the unit circle indirectly by changing the
element weights. The root music algorithm presented in Chapter 10 makes use of the
zeros on the unit circle.
The next section shows an approach to null synthesis that results in small changes in the
array weights and minimal distortion to the quiescent array factor.
w n = an (1 + n ) (2.59)
where an is the amplitude taper, and n is the complex weight perturbation that causes
the nulls. Substituting (2.59) into the equation for a linear array factor results in a nulled
array factor that is the sum of the quiescent array factor and a cancellation array factor
[31].
N N N
wne jkxn u
= an e jkxn u
+ n an e jkxn u (2.60)
n=1 n=1 n=1
# $% & # $% &
quiescent array factor cancellation array factor
AWT = BT (2.61)
where
⎡ ⎤
a1 e jkx1 u 1 a2 e jkx2 u 1 · · · a N e jkx N u 1
⎢ a1 e jkx1 u 2 a2 e jkx2 u 2 · · · a N e jkx N u 2 ⎥
⎢ ⎥
A=⎢ .. .. .. .. ⎥
⎣ . . . . ⎦
a1 e jkx1 u M
a2 e jkx2 u M
· · · aN e jkx N u M
W = [1 2 · · · N ]
N
N N
B=− an e jkxn u 1
an e jkxn u 2
··· an e jkxn u M
n=1 n=1 n=1
Since (2.61) has more columns than rows, finding the weights requires a least squares
solution in the form
Ideally, the quiescent pattern should be perturbed as little as possible when placing
the nulls. The weights that produce the nulls are [32]
M
w n = an − γm cn e− jnkdu m (2.63)
m=1
where
γm = sidelobe level of the quiescent pattern at u m
cn = amplitude taper of the cancellation beam
When cn = an , the cancellation pattern looks the same as the quiescent pattern. When
cn = 1.0, the cancellation pattern is a uniform array factor and is the constrained least
mean square approximation to the quiescent pattern over one period of the pattern. The
nulled array factor can now be written as
N N M N
wne jkxn u
= an e jkxn u
− γm an e jkxn (u−u m ) (2.64)
n=1 n=1 m=1 n=1
The sidelobes of the cancellation pattern in (2.64) produce small perturbations to the
nulled array factor.
As an example, consider an eight-element 30 dB Chebyshev array with interference
entering the sidelobe at θ = −40◦ . When cn = an , the cancellation beam is a Chebyshev
array factor as shown in Figure 2-22(a), whereas when cn = 1.0, the cancellation beam is
a uniform array factor as shown in Figure 2-22(b).
This procedure also extends to phase-only nulling by taking the phase of (2.63). When
the phase shifts are small, then e jφn ≈ 1 + jφn , and the array factor can be written as [33]
N N N N N
w n e jkxn u = an e jδn e jkxn u ≈ an (1 + jδn )e jkxn u = an e jkxn u + j an δn e jkxn u
n=1 n=1 n=1 n=1 n=1
(2.65)
Adapted
0 Cancellation Adapted
0 Cancellation
Quiscent
Quiscent
−20
Gain (dB)
−20
Gain (dB)
−40
−40
−60
−60
−90 −45 0 45 90
−90 −45 0 45 90
q (degrees) q (degrees)
(a) (b)
FIGURE 2-22 Array factors for an eight-element array with a 30 dB Chebyshev taper when
a null is synthesized at θ = −40◦ . (a) cn = an . (b) cn = 1.0.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 25, 2010 21:5 56
FIGURE 2-23
Array factors for an Adapted
eight-element array 0 Cancellation
with a 30 dB Quiscent
Chebyshev taper
when a phase-only
null is synthesized at −20
Gain (dB)
θ = −40◦ .
−40
−60
−90 −45 0 45 90
q (degrees)
N
N
an δn [cos (kxn u m ) + j sin (kxn u m )] e jkxn u = j an [cos (kxn u m ) + j sin (kxn u m )]
n=1 n=1
(2.66)
Equating the real and imaginary parts of (2.66) leads to
N
N
an δn sin (kxn u m ) e jkxn u = an cos(kxn u m ) (2.67)
n=1 n=1
N
N
an δn cos (kxn u m ) e jkxn u
=− an sin(kxn u m ) = 0 (2.68)
n=1 n=1
p /2 p/2
p 0 p 0
−p /2 −p /2
(a) (b)
FIGURE 2-24 Unit circle representations of the array factors in (a) Figure 2-22 (b) Figure 2-23.
FIGURE 2-25
Diagram of an array
with low sidelobe
w1 w2 … wN sum and difference
channels.
a1 b1 a2 b2 aN bN
Σ − Σ +
and difference patterns. The matrixes in (2.61) are now written as [35]
⎡ ⎤
a1 e jkx1 u 1 · · · a N e jkx1 u M
⎢ .. .. .. ⎥
⎢ . . . ⎥
⎢ ⎥
⎢a1 e jkx N u 1 · · · aN e jkx N u M ⎥
A=⎢ ⎢ b1 e jkx1 u 1 · · ·
⎥
⎢ b N e jkx1 u M ⎥
⎥
⎢ .. .. .. ⎥
⎣ . . . ⎦
b1 e jkx N u 1
··· bN e jkx N u M
' (
W = 1 2 · · · N
N
N
N
N
B = −j an e jkxn u 1 ··· an e jkxn u M
bn e jkxn u 1
··· bn e jkxn u M
n=1 n=1 n=1 n=1
10 10
−10 −10
Gain (dB)
Gain (dB)
−30 −30
−50 −50
FIGURE 2-26 Nulls are simultaneously placed at θ = −40◦ in the sum and difference array
factors for a 16-element monopulse array. (a) 25 dB n̄ = 5 Taylor sum pattern. (b) 25 dB n̄ = 5
Bayliss difference pattern. The dotted lines are the quiescent patterns.
To exactly cancel the interference signal at the array output at a particular frequency,
f 0 (termed the center frequency of the jamming signal), it is necessary that
w 2 = −w 1 e jω0 τ (2.71)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 59
Pj , f0 FIGURE 2-27
Two-element
s(t) adaptive array with
PN jammer located at θ .
w1
q
d Σ y (t)
w2
s(t + τ)
If we select the two adaptive weights to satisfy (2.69), it follows that the interference signal
component of the array output at any frequency is given by
and consequently the interference component of output power at any frequency may be
expressed as (assuming momentarily for simplicity that |S(ω)|2 = 1)
Let |S(ω)|2 = PJ and denote the constant spectral density of the internal noise power in
each channel by PN , it follows that the total output noise power spectral density, Po (ω),
can be written as
If we recognize that the output thermal noise power spectral density is just Pn = 2|w 1 |2 PN
and normalize the previous expression to Pn , it follows that the ratio of the total output
interference plus noise power spectral density, P0 , to the output noise power spectral
density, Pn , is then given by
P0 {1 − cos[τ (ω − ω0 )]}PJ
(ω) = 1 + (2.75)
Pn PN
Where
PJ /PN = jammer power spectral density to internal noise power spectral density per
channel
τ = (d/v) sin θ
d = sensor element spacing
θ = angular location of jammer from array broadside
ω0 = center frequency of jamming signal
At the center frequency, a null is placed exactly in the direction of the jammer so
that P0 /Pn = 1. If, however, the jamming signal has a nonzero bandwidth, then the other
frequencies present in the jamming signal will not be completely attenuated. Figure 2-28
is a plot of (2.75) and shows the resulting output interference power density in decibels as
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 60
FIGURE 2-28 12
Two-element
cancellation 10 q = 90°
performance: P0 /Pn d = l 0 /2
versus frequency for f0 = 300 MHz
8
PJ /PN = 30 and
(dB)
40 dB.
6
PJ
= 40 dB
Pn
P0
PN
4
30 dB
2
0
295 300 305
a function of frequency for a 10 MHz bandwidth jamming signal located at θ = 90◦ and
an element spacing of λ0 /2 for two values of PJ /PN . The results shown in Figure 2-28
indicate that for PJ /PN = 40 dB about 12 dB of uncanceled output interference power
remains at the edges of the 10 MHz band (with f 0 = 300 MHz). Consequently, the output
residue power at that point is about 40 − 12 = 28 dB below the interfering signal that
would otherwise be present if no attempt at cancellation were made.
It can also be noted from (2.75) and Figure 2-28 that the null bandwidth decreases as
the element spacing increases and the angle of the interference signal from array broadside
increases. Specifically, the null bandwidth is inversely proportional to the element spacing
and inversely proportional to the sine of the angle of the interference signal location.
Interference signal cancellation depends principally on three array characteristics: (1)
element spacing; (2) interference signal bandwidth; and (3) frequency-dependent interele-
ment channel mismatch across the cancellation bandwidth. The effect of sensor element
spacing on the overall array sensitivity pattern has already been discussed. Yet another
effect of element spacing is the propagation delay across the array aperture.
Assume that a single jammer having power PJ is located at θ and that the jamming
signal has a flat power spectral density with (double-sided) bandwidth B Hz. It may be
shown that the cancellation ratio P0 /PJ , where P0 is the (canceled) output residue power
in this case, is given by
P0 sin2 (π Bτ )
=1− (2.76)
Pj (π Bτ )2
where τ is given by (2.8). Note that (2.76) is quite different from (2.75), a result that
reflects the fact that the weight value that minimizes the output jammer power is different
from the weight value that yields a null at the center frequency. Equation (2.76) is plotted
in Figure 2-29, which shows how the interference signal cancellation decreases as the
signal bandwidth–propagation delay product, Bτ , increases. It is seen from (2.76) that the
cancellation ratio P0 /PJ is inversely proportional to the element spacing and the signal
bandwidth.
An elementary interchannel amplitude and phase mismatch model for a two-element
array is shown in Figure 2-30. Ignoring the effect of any propagation delay, the output
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 61
−30
−40
P0
PJ
−50
−60
0.001 0.01 0.1
Bt
−40 nly
so
P0
PJ
r
erro
a se
−50 Ph
−60
0.01 0.1 1.0 10
Error: amplitude (dB), phase (degrees)
y(t) = 1 − ae jφ (2.77)
Figure 2-30 shows a plot of the cancellation ratio P0 /PJ based on (2.78) for amplitude
errors only (φ = 0) and phase errors only (a = 1). In Part 3, where adaptive array
compensation is discussed, more realistic interchannel mismatch models are introduced,
and means for compensating such interchannel mismatch effects are studied. The results
obtained from Figure 2-30 indicates that the two-sensor channels must be matched to
within about 0.5 dB in amplitude and to within approximately 2.8◦ in phase to obtain 25
dB of interference signal cancellation.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 62
FIGURE 2-31
Realization of a w1
complex weight by
Sensor
means of a Σ
quadrature hybrid element
circuit. 90° w2
−20
AF (dB)
1.05 f0
0.95 f0
−30
−40
−50
40 45 50 55 60 65
f (degrees)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 63
demonstrates that the desired null is narrowband, because it moves as the frequency
changes. This fact leads to the conclusion that different complex weights are required
at different frequencies if an array null is to be maintained in the same direction for
all frequencies of interest. A simple and effective way of obtaining different amplitude
and phase weightings at a number of frequencies over the band of interest is to replace
the quadrature hybrid circuitry of Figure 2-31 by a transversal filter having the transfer
function h(ω). Such a transversal filter consists of a tapped-delay line having L complex
weights as shown in Figure 2-33 [33,34]. A tapped-delay line has a transfer function
that is periodic, repeating over consecutive filter bandwidths B f , as shown in Appendix
A. If the tap spacing is sufficiently close and the number of taps is large, this network
approximates an ideal filter that allows exact control of gain and phase at each frequency
within the band of interest. An upper limit on the tap spacing is given by the desired
array cancellation bandwidth Ba , since B f ≥ Ba and B f = 1/ for uniform tap spacing.
The transversal filter not only is useful for providing the desired adjustment of gain and
phase over the frequency band of interest for wideband signals but is also well suited for
providing array compensation for the effects of multipath, finite array propagation delay,
and interchannel mismatch effects; these additional uses of transversal filters are explored
further in Part 3. Transversal filters are also enhance the ability of an airborne moving
target indication (AMTI) radar system to reject clutter, to compensate for platform motion,
and to compensate for near-field scattering effects and element excitation errors [35].
To model a complete multichannel processor (in a manner analogous to that of Section
1.4 for the narrowband processing case), consider the tapped-delay line multichannel
processor depicted in Figure 2-34. The multichannel wideband signal processor consists
of N sensor element channels in which each channel contains one tapped-delay line
like that of Figure 2-33 consisting of L tap points, (L − 1) time delays of seconds
each, and L complex weights. On comparing Figure 2-34 with Figure 1-1, it is seen that
x1 (t), x2 (t), · · · x N (t) in the former correspond exactly with x1 (t), x2 (t) . . . x N (t) in the
latter, which were defined in (1.6) to form the elements of the vector x(t). In like manner,
therefore, define a complex vector x 1 (t) such that
The signals appearing at the second tap point in all channels are merely a time-delayed
version of the signals appearing at the first tap point, so define a second complex signal
vector x 2 (t) where
Sensor
Δ Δ Δ
no. k
xk [t − (l − 1)Δ] xk [t − (L − 2)Δ] xk [t − (L − 1)Δ]
xk(t )
Processor
wk1 wkl wk (L−1) wkL output
Σ signal
y (t)
Sensor
Δ Δ Δ
no. N
xN [t − (l − 1)Δ] xN [t − (L − 2)Δ] xN [t − (L − 1)Δ]
xN(t )
wN 1 wN l wN (L−1) wN L
FIGURE 2-34 Tapped-delay line multichannel processor for wideband signal processing.
Continuing in the previously given manner for all L tap points, a complete signal vector
for the entire multichannel processor can be defined as
⎡ ⎤ ⎡ ⎤
x1 (t) x (t)
⎢ x2 (t) ⎥ ⎢
x (t − ) ⎥
⎢ ⎥ ⎢ ⎥
x(t) = ⎢ . ⎥ = ⎢ .. ⎥ (2.81)
⎣ .. ⎦ ⎣ . ⎦
xL (t) x [t − (L − 1)]
It is seen that the signal vector x(t) contains L component vectors each having dimen-
sion N .
Define the weight vector
(w1 )T = [w 11 w 21 · · · w N 1 ] (2.82)
The weight vector for the entire multichannel processor is then given by
⎡
⎤
w1
⎢w ⎥
⎢ 2⎥
w=⎢ . ⎥ (2.83)
⎣ .. ⎦
wL
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 65
L L
y(t) = (wl )T xl (t) = (wl )T x [t − (l − 1)]
l=1 l=1 (2.84)
= w x(t)
T
which is the same form as (1.4). Yet another reason for expressing the signal vector in the
form of (2.81) is that this construction leads to a Toeplitz form for the correlation matrix
of the input signals, as shown in Chapter 3.
The array processors of Figures 2-33 and 2-34 are examples of classical time-domain
processors that use complex weights and delay lines. Fast Fourier transform (FFT) tech-
niques replace a conventional time-domain processor by an equivalent frequency domain
processor using the frequency-domain equivalent of time delay [39–41]. When a tapped
delay-line array and an FFT array use the same time delay between samples and the same
number of samples, their performance is identical. FFT processing has two advantages
compared with its time-domain equivalent: (1) the hardware problem can be alleviated;
and (2) the computational burden can be reduced. The use of phase shifters or delay lines to
form a directional beam in the time domain becomes cumbersome from a hardware stand-
point as the number of delay elements increases, whereas using the frequency-domain
equivalent permits beamforming to be accomplished with a digital computer, thereby sim-
plifying the hardware. Likewise, taking FFTs tends to reduce the correlation between sam-
ples in different frequency subbands. When samples in different subbands are completely
decorrelated, the signal covariance matrix has a block diagonal form, and the optimal
weights can be computed separately in each subband, thus reducing the computational
burden.
As more taps are added to each delay line, the bandwidth performance of the array
improves. The bandwidth cutoff, Bc , of an array is defined as the maximum bandwidth
such that the array SINR is within 1dB of its continuous wave (narrowband) value. It is
better to divide the taps equally among the array elements rather than to employ an unequal
distribution, although higher bandwidth performance can usually be realized by placing
an extra tap behind the middle element in an array. A piecewise linear curve is plotted in
Figure 2-35, giving the optimal number of taps per element (number of samples for FFT
processing) versus signal fractional bandwidth, B, for 3 to 10 elements in a linear array
(B = BW/carrier frequency). Since it is easiest to use 2n samples in an FFT processor,
where n is some integer, one would use the smallest value of n such that 2n is at least
as large as the value shown in Figure 2-35. The time delay between taps, To , can be
described in terms of the parameter r , where r is the number of quarter wavelength delays
found in To . For optimal SINR array performance, the value of r must be in the range
0 < r < 1/B. If r < 1/B, the performance is optimal but results in large array weights.
Therefore, the choice r = 1/B is dictated by good engineering practice. The notion of
the frequency-domain equivalent of time delay is introduced in the Problems section, and
more detailed presentations of the techniques of digital beamforming may be found in
[42–45].
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 66
FIGURE 2-35 6
Number of taps
versus fractional
bandwidth, B. From 5
Vook and Compton,
IEEE Trans. Aerosp. 10
and Electron. Sys., 4 Number of elements
Vol. 38, No. 3, July Number of taps indicated along
each graph) 6
1992.
7, 8, 9, 10
3
8, 9, 10 4, 5, 6, 7 3, 4, 5, 6
2
3 elements
3, 4, 5, 6, 7
1
0
0 0.05 0.10 0.15 0.20
Fractional bandwidth, B
45 55
35
40 50
sky elevation
45
50 40
38 45
45 50 55 40
40 60
35
36
35
35
38
sky azimuth
50
40
30
20
10
0
20 30 40 50 60 70 80 90 100
J/S (dB)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 68
The coverage improvement factor (CIF) of the array at each point in the field of view
can be computed as
n
G P +N0q
i=1 iq i
G sq Ps
CIF = n
before nulling
(2.86)
G i Pi +N0
i=1
G s Ps
after nulling
where the numerator gains G iq and G sq denote the array gain toward the ith jammer and
the desired signal, respectively, for the quiescent pattern case, and N0q denotes the thermal
noise power before nulling.
The coverage improvement factor statistic (CIFS) can then be obtained by plotting
the percentage of the field of view where the nulling system provides X dB of protection
compared with the quiescent system in the same manner as the J/S coverage statistic plot
was obtained and yields a convenient way of assessing how effective the adaptive array is
in yielding a performance advantage over a nonadaptive array.
2.8 Problems 69
The best possible steady-state performance that can be achieved by an adaptive array
system can be computed theoretically, without explicit consideration of the array factors
affecting such performance. The theoretical performance limits for steady-state array
operation are addressed in Chapter 3.
2.8 PROBLEMS
1. Circular Array. For the circular array of Figure 2-38 having N equally spaced elements, select
as a phase reference the center of the circle. For a signal whose direction of arrival is θ with
respect to the selected angle reference, define the angle φk for which φk = θ − ψk .
(a) Show that ψk = (2π/N )(k − 1) for k = 1, 2, . . . , N .
(b) Show that the path length phase difference for any element with respect to the center of
the circle is given by u k = π(R/λ) cos φk .
(c) Show that an interelement spacing of λ/2 is maintained by choosing the radius of the array
circle to be
λ
R=
4 sin(π/N )
2. Linear Array with Spherical Signal Wavefront. Consider the linear array geometry of
Figure 2-39 in which v is the vector from the array center to the signal source, θ is the angle
between v and the array normal, and is the propagation velocity.
FIGURE 2-38
fk Azimuth plane of
circular array having
N equally spaced
q elements.
Yk
0°
R
y FIGURE 2-39
Linear array
Signal source geometry.
v
z
l
q
x
L L
−
2 2
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 70
(a) Show that for a spherical signal wavefront the time delay experienced by an element
located at position x in the array with respect to the array center is given by
) * +
||v|| x2 x
τ (x) = 1− 1+ −2 sin θ
||v||2 ||v||
(τ is positive for time advance and negative for time delay) where ||v|| denotes the length
of the vector v.
(b) Show that for a spherical signal wavefront the attenuation factor experienced by an element
located at position x in the array with respect to the array center is given by
||v|| 1
ρ(x) = = ,
||l|| x2
1+ ||v||2
−2 x
||v||
sin θ
where ||l|| denotes the distance between the signal source and the array element of concern.
3. Uniformly Spaced Linear Array [11]. Show that for a uniformly spaced linear array having N
elements the normalized array factor is given by
1 sin(N u/2)
AF(u) =
N sin(u/2)
where u = 2π(d/λ) sin θ.
4. Representation of Nonuniformly Spaced Arrays [20]
(a) For nonuniformly spaced elements in a linear array, a convenient “base separation” d can
be selected and the element position then represented by
i
di = + εi d
2
where εi is the fractional change from the uniform spacing represented by d. The normal-
ized field pattern for a nonuniformly spaced array is then given by
N N
1 1 i
AF = cos u i = cos + εi u
N N 2
i=1 i=1
where Au is the pattern of a uniform array having element separation equal to the base
separation and is given by the result obtained in Problem 3.
(b) Assuming all εi u are small, show that the result obtained in part (a) reduces to
u u
AF = AF u − εi sin i
N 2
i
2.8 Problems 71
which may be regarded as a Fourier series representation of the quantity on the right-hand
side of the previous equation. Consequently, the εi are given by the formula for Fourier
coefficients:
0
2N 1 u
εi = (AF u − AF) sin i du
π π u 2
Let
L
AF u − AF 1
= ak δ(u − u k )
u u
k=1
(d) For uniform arrays the pattern sidelobes have maxima at positions approximately given
by
π
u k = (2k + 1)
N
Consequently, u k gives the positions of the necessary impulse functions in the represen-
tation of (AF u − AF)/u given in part (c). Furthermore, since the sidelobes adjacent to the
main lobe drop-off as 1/u, (AF u − AF)/u is now given by
AF u − AF 1
L - π .
= AF 2 (−1)k δ u − (2k + 1)
u u N
k=1
5. Nonuniformly Spaced Arrays and the Equivalent Uniformly Spaced Array [19]. One method
of gaining insight into the gross behavior of a nonuniformly spaced array and making it
amenable to a linear analysis is provided by the correspondence between a nonuniformly
spaced array and its equivalent uniformly spaced array (EUA). The EUA provides a best mean
square representation for the original nonuniformly spaced array.
(a) Define the uniformly spaced array factor A(θ) as the normalized phasor sum of the re-
sponses of each of the sensor elements to a narrowband signal in the uniform array of
Figure 2-40, where the normalization is taken with respect to the response of the center
FIGURE 2-40
Uniformly spaced
q array having an odd
number of sensor
elements.
3′ 2′ 1′ 1 2 3
d
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 72
FIGURE 2-41
Symmetrical
q
nonuniformly
spaced array having
an odd number of
sensor elements.
3′ 2′ 1′ 1 2 3
d1
d2
d3
where l = d/λ. Furthermore, show that when d/λ = 2 the array factor has its second
principal maxima occurring at θ = π/6.
(b) Show that the nonuniformly spaced array factor for the array of Figure 2-41 is given by
(N −1)/2
where li = di /λ.
(c) The result obtained in part (b) may be regarded as AF(θ ) = 1 + 2i cos ωi where
2
ωi = li (π R sin θ ), R = arbitrary scale factor
R # $% &
ω1
Regarding ω1 = π g sin θ as the lowest nonzero superficial array frequency in the EUA,
the scale factor R determines the element spacing in the EUA. For example, R = 1
corresponds to half-wavelength spacing, and R = 12 corresponds to quarter-wavelength
spacing. Each term in AF(θ) may be expanded (in a Fourier series representation for each
higher harmonic) into an infinite number of uniformly spaced equivalent elements. Since
the amplitude of the equivalent elements varies considerably across the array, only a few
terms of the expansion need be considered for a reasonable equivalent representation in
terms of EUAs. The mth term in AF(θ) is given by
2
cos ωm = cos lm ω1 = cos μm ω1
R
2.8 Problems 73
where
πR
2
amv = cos μm ω1 cos vω1 dω1
π 0
where
AF 0 = a10 + a20 + a30 + · · ·
AF 1 = a11 + a21 + a31 + · · ·
Note that the representation given by the previous result is not unique. Many choices are
possible as a function of the scale factor R. From the least mean square error property of
the Fourier expansion, each representation will be the best possible for a given value of ω1 .
6. Array Propagation Delay Effects. Array propagation delay effects can be investigated by
considering the simple two-element array model of Figure 2-42. It is desired to adjust the value
of w 1 so the output signal power P0 is minimized for a directional interference signal source.
(a) Since x2 (ω) = e− jωτ x1 (ω) in the frequency domain, the task of w 1 may be regarded as
one of providing the best estimate (in the mean square error sense) of x1 (ω)e− jωτ , and
the error in this estimate is given by ε(ω) = x1 (ω)(e− jωτ − w 1 ). For the weight w 1 to be
optimal, the error in the estimate must be orthogonal to the signal x1 (ω) so it is necessary
that
E{|x1 (ω)|2 [e− jωτ − w 1 ]} = 0
where the expectation E{·} is taken over frequency. If |x1 (ω)|2 is a constant indepen-
dent of frequency, it follows from the previously given orthogonality condition that w 1
q
l PJ
gna
nce si x1(t)
rfere w1
Inte
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 74
must satisfy
E{e− jωτ − w 1 } = 0
If the signal has a rectangular spectrum over the bandwidth −π B ≤ ω ≤ π B, show that
the previous equation results in
sin(π Bτ )
w1 =
π Bτ
(b) The output power may be expressed as
πB
P0 = φ R R (ω, θ )dω
−π B
where φ R R (ω, θ ) is the output power spectral density. The output power spectral density
is in turn given by
where φss (ω) is the power spectral density of the directional interference signal. Assuming
that φss (ω) is unity over the signal bandwidth (so that PJ = 2π B), then from part (a) it
follows that
− jωτ sin(π Bτ ) 2
πB
P0 = − dw
e π Bτ
−π B
λ
σ x = σ y = σz =
4π θs
where λ is the radiated wavelength, and θs is the maximum scan angle from the initial
pointing direction of the array (i.e., θs is half the field of view).
2. The rms tolerance in the estimate of initial pointing direction is
θ
σθ ∼
=
2θs
2.8 Problems 75
θs D 2
Rm ≈
2λ
where D is the linear dimension of the array.
4. The rms tolerance in the range estimate of the target is
R0 θ
σ R0 ∼
=
2θs
where R0 < Rm is the target range.
(a) Assuming λ = σx = σ y = σz = 10−1 ft, what field of view can successfully be scanned
with open-loop techniques?
(b) What is the minimum range at which far-field scanning may be used?
(c) For the field of view found in part (a), determine the rms tolerance in the initial pointing
error.
8. Frequency-Domain Equivalent of Time Delay [45]. The conventional time-domain proces-
sor for beamforming involving sum and delay techniques can be replaced by an equivalent
frequency-domain processor using the FFT.
The wavefront of a plane wave arriving at a line of sensors having a spacing d between
elements will be delayed in time by an amount τ1 between adjacent sensors given by
d
τ1 = sin θ1
where θ1 is the angle between the direction of arrival and the array normal. Consequently, for
a separation of n elements
τn1 = nτ1
Note that φn represents the phase shift equivalent to the time delay τn .
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 76
(b) The plane wave arriving at the array with arrival angle θ1 can be represented by
- z .
x(t, z) = cos ω t − = cos[ωt − kz]
where k is the wavenumber of the plane wave, and z is the axis of propagation at an angle
θ1 with respect to the array normal. Show that the phase shift φn associated with a time
delay τn is given by
φn = ωτn
where
d
τn = n sin θ1
(c) It is convenient to index θ and ω to account for different combinations of the angle θl and
frequency ω so that
θ = θl = lθ, l = 0, 1, . . . , N − 1
ω = ωm = mω, m = 0, 1, . . . , M − 1
t = ti = it, i = 0, 1, . . . , M − 1
X n (mω) = X n (ωm ) = X mn
Likewise, the sampled time waveform from the nth array element can be represented by
Show that φnl = nkd sin θl . Consequently, show that the previously summed expressions
can be rewritten in terms of both frequency and angle as
N −1
X m (1θ ) = X n (mω)e− jφnl
n=0
nl
φnl =
N
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 77
2.8 Problems 77
xn (it) = xni
The array output has therefore been obtained in the form of spectral sample versus beam
number, which results after a two-dimensional DFT transformation of array input. The
required DFTs can then be implemented using the FFT algorithm.
9. Linear Array Directivity. Calculate the directivity of a 10-element uniform array with d = 0.5λ,
and compare it with a 10-element uniform array with d = 1.0λ. Why are they different?
10. Moving Nulls on the Unit Circle. Plot the unit circle representations of a four-element uniform
array. Now, move the null locations from ψ = ±90◦ to ψ = ±120◦ .
11. Chebyshev Array. Plot the unit circle representation, the amplitude weights, and the array
factor for a six-element 20 dB Chebyshev array.
12. Taylor Array. Plot the unit circle representation, the amplitude weights, and the array factor
for a 20-element 20 dB n̄ = 5 Taylor array.
13. Taylor Array. Start with a Taylor n̄ = 4 sll = 25dB taper, and place a null at u = 0.25 when
d = 0.5λ. Do not allow complex weights.
14. Thinned Array. A 5,000-element linear array has a taper efficiency of 50% and elements spaced
a half wavelength apart. What are the average and peak sidelobe levels of this array?
15. Thinned Array with Taylor Taper. Start with a 100-element array with element spacing d = 0.5,
and thin to a 20 dB n̄ = 5 Taylor taper.
16. Cancellation Beams. A 40-element array of isotropic point sources spaced λ/2 apart has a
30 dB n̄ = 7 Taylor taper. Plot the array factors and cancellation beam for a null at 61o .
17. Cancellation Beams. Repeat the previous example using both uniform and weighted cancel-
lation beams.
18. Cancellation Beams. A 40-element array of isotropic point sources spaced λ/2 apart has a
40 dB n̄ = 7 Taylor taper. Plot the array factors and cancellation beams for nulls at 13◦ and
61◦ . Use phase-only nulling.
19. Simultaneous Nulling in the Sum and Difference Patterns. A 40-element array of isotropic
point sources spaced λ/2 apart has a 30 dB n̄ = 7 Taylor taper and a 30 dB n̄ = 7 Bayliss
taper. Plot the array factors and cancellation beams for nulls at 13◦ and 61◦ . Use simultaneous
phase-only nulling in the sum and difference patterns.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 78
20. Triangular Spacing in a Planar Array [11]. Show that the uniform array factor for triangular
spacing is
sin (N xe ψx ) sin N ye ψ y sin (N xo ψx ) sin N yo ψ y
AF = + e− j (ψx +ψ y )
N xe sin (ψx ) N ye sin ψ y N xo sin (ψx ) N yo sin ψ y
2.9 REFERENCES
[1] P. W. Howells, “Explorations in Fixed and Adaptive Resolution at GE and SURC,” IEEE
Trans. Antennas Propag., Vol. AP-24, No. 5, September 1976, pp. 575–584.
[2] R. L. Deavenport, “The Influence of the Ocean Environment on Sonar System Design,”
IEEE 1975 EASCON Record, Electronics and Aerospace Systems Convention, September
29–October 1, pp. 66.A–66.E.
[3] A. A. Winder, “Underwater Sound—A Review: II. Sonar System Technology,” IEEE Trans.
Sonics and Ultrason., Vol. SU-22, No. 5, September 1975, pp. 291–332.
[4] L. W. Nolte and W. S. Hodgkiss, “Directivity or Adaptivity?” IEEE 1975 EASCON Record,
Electronics and Aerospace Systems Convention, September 29–October 1, pp. 35.A–35.H.
[5] G. M. Wenz, “Review of Underwater Acoustics Research: Noise,” J. Acoust. Soc. Am., Vol.
51, No. 3 (Part 2), March 1972, pp. 1010–1024.
[6] B. W. Lindgren, Statistical Theory, New York, MacMillan, 1962, Ch. 4.
[7] F. Bryn, “Optimum Structures of Sonar Systems Employing Spatially Distributed Receiv-
ing Elements,” in NATO Advanced Study Institute on Signal Processing with Emphasis on
Underwater Acoustics, Vol. 2, Paper No. 30, Enschede, The Netherlands, August 1968.
[8] H. Cox, “Optimum Arrays and the Schwartz Inequality,” J. Acoust. Soc. Am., Vol. 45, No. 1,
January 1969, pp. 228–232.
[9] R. S. Elliott, “The Theory of Antenna Arrays,” in Microwave Scanning Antennas, Vol. 2, Array
Theory and Practice, edited by R. C. Hansen, New York, Academic Press, 1966, Ch. 1.
[10] I. J. Gupta and A. A. Ksienski, “Effect of Mutual Coupling on the Performance of Adaptive
Arrays,” IEEE AP-S Trans., Vol. 31, No. 5, September 1983, pp. 785–791.
[11] R. L. Haupt, Antenna Arrays: A Computational Approach, New York, Wiley, 2010.
[12] M.I. Skolnik, Introduction to Radar Systems, New York, McGraw-Hill, 2002.
[13] H. T. Friis, “A Note on a Simple Transmission Formula,” IRE Proc., May 1946, pp. 254–256.
[14] J. S. Stone, United States Patent No. 1,643,323 and 1,715,433.
[15] C. L. Dolph, “A Current Distribution for Broadside Arrays which Optimizes the Relationship
Between Beam Width and Side-Lobe Level,” Proceedings of the IRE, Vol. 34, No. 6, June
1946, pp. 335–348.
[16] T. T. Taylor, “Design of Line Source Antennas for Narrow Beamwidth and Low Side Lobes,”
IRE AP Trans., Vol. 3, 1955, pp. 16–28.
[17] T. Taylor, “Design of Circular Apertures for Narrow Beamwidth and Low Sidelobes,” IEEE
AP-S Trans., Vol. 8, No. 1, January 1960, pp. 17–22.
[18] F. Hodjat and S. A. Hovanessian, “Nonuniformly Spaced Linear and Planar Array Antennas
for Sidelobe Reduction,” IEEE Trans. Antennas Propag., Vol. AP-26, No. 2, March 1978,
pp. 198–204.
[19] S. S. Sandler, “Some Equivalences Between Equally and Unequally Spaced Arrays,” IRE
Trans. Antennas and Progag., Vol. AP-8, No. 5, September 1960, pp. 496–500.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 79
2.9 References 79
[20] R. F. Harrington, “Sidelobe Reduction by Nonuniform Element Spacing,” IRE Trans. Antennas
Propag., Vol. AP-9, No. 2, March 1961, pp. 187–192.
[21] D. D. King, R. F. Packard, and R. K. Thomas, “Unequally Spaced, Broadband Antenna
Arrays,” IRE Trans. Antennas Propag., Vol. AP-8, No. 4, July 1960, pp. 380–385.
[22] M. I. Skolnik, “Nonuniform Arrays,” in Antenna Theory, Part I, edited by R. E. Collin and F.
J. Zucker, New York, McGraw-Hill, 1969, Ch. 6.
[23] Y. T. Lo, “A Mathematical Theory of Antenna Arrays with Randomly Spaced Elements,”
IEEE Trans. Antennas Propag., Vol. AP-12, May 1964, pp. 257–268.
[24] Y. T. Lo and R. J. Simcoe, “An Experiment on Antenna Arrays with Randomly Spaced
Elements,” IEEE Trans. Antennas Propag., Vol. AP-15, March 1967, pp. 231–235.
[25] E. Brookner, “Antenna Array Fundamentals—Part 1,” in Practical Phased Array Antenna
Systems, Norwood, MA, Artech House, 1991, pp. 2-1–2-37.
[26] R.L. Haupt, “Thinned Arrays Using Genetic Algorithms,” IEEE AP-S Transactions, Vol. 42,
No. 7, July 1994, pp. 993–999.
[27] R. L. Haupt and D. Werner, Genetic Algorithms in Electromagnetics, New York, Wiley, 2007.
[28] R. L. Haupt, “Interleaved Thinned Linear Arrays,” IEEE AP-S Trans., Vol. 53, September
2005, pp. 2858–2864.
[29] R. Guinvarc’h and R. L. Haupt, “Dual Polarization Interleaved Spiral Antenna Phased Array
with an Octave Bandwidth” IEEE Trans. Antennas Propagat., Vol. 58, No. 2, 2010, pp. 397–
403.
[30] R. L. Haupt, “Optimized Element Spacing for Low Sidelobe Concentric Ring Arrays,” IEEE
AP-S Trans., Vol. 56, No. 1, January 2008, pp. 266–268.
[31] H. Steyskal, “Synthesis of Antenna Patterns with Prescribed Nulls,” IEEE Transactions on
Antennas and Propagation, Vol. 30, no. 2, 1982, pp. 273–279.
[32] H. Steyskal, R. Shore, and R. Haupt, “Methods for Null Control and Their Effects on the
Radiation Pattern,” IEEE Transactions on Antennas and Propagation, Vol. 34, No. 3, 1986,
pp. 404–409.
[33] H. Steyskal, “Simple Method for Pattern Nulling by Phase Perturbation,” IEEE Transactions
on Antennas and Propagation, Vol. 31, No. 1, 1983, pp. 163–166.
[34] R. Shore, “Nulling a Symmetric Pattern Location with Phase-Only Weight Control,” IEEE
Transactions on Antennas and Propagation, Vol. 32, No. 5, 1984, pp. 530–533.
[35] R. Haupt, “Simultaneous Nulling in the Sum and Difference Patterns of a Monopulse Antenna,”
IEEE Transactions on Antennas and Propagation, Vol. 32, No. 5, 1984, pp. 486–493.
[36] B. Widrow and M. E. Hoff, Jr., “Adaptive Switching Circuits,” IRE 1960 WESCON Convention
Record, Part 4, pp. 96–104.
[37] J. H. Chang and F. B. Tuteur, “A New Class of Adaptive Array Processors,” J. Acoust. Soc.
Am., Vol. 49, No. 3, March 1971, pp. 639–649.
[38] L. E. Brennan, J. D. Mallet, and I. S. Reed, “Adaptive Arrays in Airborne MTI Radar,” IEEE
Trans. Antennas Propag., Special Issue Adaptive Antennas, Vol. AP-24, No. 5, September
1976, pp. 607–615.
[39] R. T. Compton, Jr., “The Bandwidth Performance of a Two-Element Adaptive Array with
Tapped Delay-Line Processing,” IEEE Trans. Ant. & Prop., Vol. AP-36, No.1, Jan. 1988,
pp. 5–14.
[40] R. T. Compton, Jr., “The Relationship Between Tapped Delay-Line and FFT Processing in
Adaptive Arrays,” IEEE Trans. Ant. & Prop., Vol. AP-36, No. 1, January 1988, pp. 15–26.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:9 80
[41] F. W. Vook and R. T. Compton, Jr., “Bandwidth Performance of Linear Adaptive Arrays with
Tapped Delay-Line Processing,” IEEE Trans. Aerosp. & Electron. Sys., Vol. AES-28, No. 3,
July 1992, pp. 901–908.
[42] P. Rudnick, “Digital Beamforming in the Frequency Domain,” J. Acoust. Soc. Am., Vol. 46,
No. 5, 1969.
[43] D. T. Deihl, “Digital Bandpass Beamforming with the Discrete Fourier Transform,” Naval
Research Laboratory Report 7359, AD-737-191, 1972.
[44] I. M. Weiss, “GPS Adaptive Antenna/Electronics—Proposed System Characterization Per-
formance Measures,” Aerospace Corp. Technical Report No. TOR-95 (544)-1, 20 December
1994.
[45] L. Armijo, K. W. Daniel, and W. M. Labuda, “Applications of the FFT to Antenna Array
Beamforming,” EASCON 1974 Record, IEEE Electronics and Aerospace Systems Convention,
October 7–9, Washington, DC, pp. 381–383.
[46] A. M. Vural, “A Comparative Performance Study of Adaptive Array Processors,” Proceedings
of the 1977 International Conference on Acoustics, Speech, and Signal Processing, May,
pp. 695–700.
[47] S. H. Taheri and B. D. Steinberg, “Tolerances in Self-Cohering Antenna Arrays of Arbitrary
Geometry,” IEEE Trans. Antennas Propag., Vol. AP-24, No. 5, September 1976, pp. 733–739.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 81
CHAPTER
Optimum array processing is an optimum multichannel filtering problem [1–12]. The ob-
jective of array processing is to enhance the reception (or detection) of a desired signal that
may be either random or deterministic in a signal environment containing numerous inter-
ference signals. The desired signal may also contain one or several uncertain parameters
(e.g., spatial location, signal energy, phase) that it may be advantageous to estimate.
Optimum array processing techniques are broadly classified as (1) processing appro-
priate for ideal propagation conditions and (2) processing appropriate for perturbed prop-
agation conditions. Ideal propagation implies an ideal nonrandom, nondispersive medium
where the desired signal is a plane (or spherical) wave and the receiving sensors are distor-
tionless. In this case the optimum processor is said to be matched to a plane wave signal.
Any performance degradation resulting from deviation of the actual operating conditions
from the assumed ideal conditions is minimized by the use of complementary methods,
such as the introduction of constraints. When operating under the aforementioned ideal
conditions, vector weighting of the input data succeeds in matching the desired signal.
When perturbations in either the propagating medium or the receiving mechanism
occur, the plane wave signal assumption no longer holds, and vector weighting the input
data will not match the desired signal. Matrix weighting of the input data is necessary [13]
for a signal of arbitrary characteristics. The principal advantage of such an element space-
matched array processor operates on noncoherent wavefront signals where matching can
be performed in only a statistical sense.
Steady-state performance limits establish a performance measure for any selected
design. Quite naturally, optimum array processing has mostly assumed ideal propagation
81
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 82
conditions, and various approaches to this problem have been proposed for both narrow-
band and broadband signal applications. By far, the most popular and widely reported
approaches involve the adoption of an array performance measure that is optimized by ap-
propriate selection of an optimum weight vector. Such performance measure optimization
approaches have been widely used for radar, communication, and sonar systems and there-
fore are discussed at some length. Approaches to array processing that use the maximum
entropy method (which is particularly applicable to the passive sonar problem) [14,15]
and eigenvalue resolution techniques [16,17] have been reported.
To determine the optimal array weighting and its associated performance limits, some
mathematical preliminaries are first discussed, and the signal descriptions for conventional
and signal-aligned arrays are introduced. It is well known that when all elements in an
array are uniformly weighted, then the maximum signal-to-noise ratio (SNR) is obtained
if the noise contributions from the various element channels have equal power and are
uncorrelated [18]. When there is any directional interference, however, the noise from the
various element channels is correlated. Consequently, selecting an optimum set of weights
involves attempting to cancel correlated noise components. Signal environment descrip-
tions in terms of correlation matrices therefore play a fundamental role in determining the
optimum solution for the complex weight vector.
The problem of formulating some popular array performance measures in terms of
complex envelope signal characterizations is discussed. It is a remarkable fact that the
different performance measures considered here all converge (to within a constant scalar
factor) toward the same steady-state solution: the optimum Wiener solution. Passive de-
tection systems face the problem of designing an array processor for optimum detection
performance. Classical statistical detection theory yields an array processor based on
a likelihood ratio test that leads to a canonical structure for the array processor. This
canonical structure contains a weighting network that is closely related to the steady-state
weighting solutions found for the selected array performance measures. It therefore turns
out that what at first appear to be quite different optimization problems actually have a
unified basis.
where Re{ } denotes “real part of.” This result is further discussed in Appendix B. Applying
this notation to the array output yields
representation is the analytic signal ψ(t) for which y(t) = Re{ψ(t)} and
ψ(t) = y(t) + j y̌(t) (3.3)
where y̌(t) denotes the Hilbert transform of y(t). It follows that for a complex weight
representation w = w 1 + jw 2 having an input signal with analytic representation x1 + j x2 ,
the resulting actual output signal is given by
where x2 (t) = x̌1 (t). Note that (complex envelope) e jω0 t = ψ(t). As a consequence of
the previously given meaning of complex signal and weight representations, two alternate
approaches to the optimization problems for the optimal weight solutions may be taken
as follows:
1. Reformulate the problem involving complex quantities in terms of completely real
quantities so that familiar mathematical concepts and operations are conveniently car-
ried out.
2. Revise the definitions of certain concepts and operations (e.g., covariance matrices,
gradients) so that all complex quantities are handled directly in an appropriate manner.
Numerous examples of both approaches are found in the adaptive array literature. It is
therefore appropriate to consider how to use both approaches since there are no compelling
advantages favoring one approach over the other.
It follows that the n-component complex vector z is completely represented by the 2n-
component real vector x. By representing all complex quantities in terms of corresponding
real vectors, the adaptive array problem is solved using familiar definitions of mathematical
concepts and procedures. If we carry out the foregoing procedure for the single complex
weight w 1 + jw 2 and the complex signal x1 + j x2 , there results
x1
w x = [w 1 w 2 ]
T
= w 1 x1 + w 2 x2 (3.6)
x2
which is in agreement with the result expressed by (3.4). If the foregoing approach
is adopted, then certain off-diagonal elements of the corresponding correlation matrix
(defined in the following section) will be zero. This fact does not affect the results
obtained in this chapter; the point is discussed more fully in connection with an example
presented in Chapter 11.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 84
where E{ } denotes the expected value, and τ is a running time-delay variable. Likewise,
the autocorrelation matrix for the vector x(t) is defined by
Rx x (τ ) = E{x(t)xT (t − τ )} (3.8)
If a signal vector x(t) consists of uncorrelated desired signals and noise components
so that
then
The correlation matrices of principal concern in the adaptive array analysis are those
for which the time-delay variable τ is zero. Rather than write the correlation matrix
argument explicitly as Rx y (0) and Rx x (0), we define
Rx x = Rx x (0) (3.11)
and
Rx y = Rx y (0) (3.12)
It follows that for an N-dimensional vector x(t), the autocorrelation matrix is simply written
as Rx x where
⎡ ⎤
x1 (t)x1 (t) x1 (t)x2 (t) . . . x1 (t)x N (t)
⎢ ⎥
⎢ x 2 (t)x 1 (t) x 2 (t)x 2 (t) . . . ⎥
Rx x = E{xx } = ⎢
T ⎢ ⎥ (3.13)
.. ⎥
⎣ . ⎦
x N (t)x1 (t) ... x N (t)x N (t)
where x (t) is the (real) signal vector defined in (2.80) for a tapped delay-line multichannel
processor. Likewise, for the NL-dimensional vector of all signals observed at the tap points,
the autocorrelation matrix is given by
Rx x (τ ) = E{x(t)xT (t − τ )} (3.16)
The (NL × NL)-dimensional matrix Rx x (τ ) given by (3.18) has the form of a Toeplitz
matrix—a matrix having equal valued matrix elements along any diagonal [20]. The
desirability of the Toeplitz form lies in the fact that the entire matrix is constructed from
the first row of submatrices; that is, Rx x (τ ), Rx x (τ + ), . . . , Rx x [τ + (L − 1)]. Con-
sequently, only an (N × NL)-dimensional matrix need be stored to have all the information
contained in Rx x (τ ).
Covariance matrices are closely related to correlation matrices since the covariance
matrix between the vectors x(t) and y(t) is defined by
cov[x(t), y(t)] = E{(x(t) − x)(y(t) − y)T } (3.19)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 86
where
Thus, for zero-mean processes and with τ = 0, correlation matrices and covariance ma-
trices are identical, and the adaptive array literature frequently uses the two terms inter-
changeably.
Frequency-domain signal descriptions as well as time-domain signal descriptions
are valuable in considering array processors for broadband applications. The frequency-
domain equivalent of time-domain descriptions may be found by taking the Fourier trans-
form of the time domain quantities, (ω) = { f (t)}. The Fourier transform of a signal
correlation matrix yields the signal cross-spectral density matrix
Cross-spectral density matrices therefore present the signal information contained in cor-
relation matrices in the frequency domain.
The gradient of a scalar function, ∇ y , consists of the partial derivatives of f (·) along each
component direction of y. In the case of real variables, the gradient operator is a vector
operator given by
T
∂ ∂
∇y = ... (3.23)
∂ y1 ∂ yn
so that
∂f ∂f ∂f
∇ y f (y) = e1 + e2 + · · · + en (3.24)
∂ y1 ∂ y2 ∂ yn
where e1 , e2 , . . . , en is the set of unit basis vectors for the vector y. For a complex vector
y each element yk has a real and an imaginary component:
yk = xk + j z k (3.25)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 87
Therefore, each partial derivative in (3.24) now has a real and an imaginary component so
that [21]
∂f ∂f ( j)∂ f
= + (3.26)
∂ yk ∂ xk ∂z k
In the optimization problems encountered in this chapter, it is frequently desired to
obtain the gradient with respect to the vector x of the scalar quantity xT a = aT x and of
the quadratic form xT Ax (which is also a scalar), where A is a symmetric matrix. Note
that xT a is an inner product
(x, a) = xT a = aT x (3.27)
where
A = vvT (3.29)
has all the properties of an inner product, so formulas for the differentiation of the trace
of various matrix products are of interest in obtaining solutions to optimization problems.
A partial list of convenient differentiation formulas is given in Appendix C. From these
formulas it follows for real variables that
and
and
The definition of (3.36) yields a matrix that is the complex conjugate (or the transpose) of
the definition given by (3.35). So long as one adheres consistently to either one definition
or the other, the results obtained (in terms of the selected performance measure) will turn
out to be the same; therefore, it is immaterial which definition is used. With either of the
aforementioned definitions, it immediately follows that
FIGURE 3-1 x1
Conventional w1
narrowband array.
x2
w2 Σ y(t) = wTx (t)
xN
wN
FIGURE 3-2 x1 z1
Signal aligned f1 w1
Time
narrowband array. delay or
phase
shift
elements x2 z2
f2 w2 Σ y(t) = wTz(t)
zi (t) = e jf i xi (t)
xN zN
fN wN
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 89
and this information is used to obtain time-coincident desired signals in each channel.
One advantage of the signal aligned array structure is that a set of weights is found that
is independent of the desired signal time structure, which provides a distortionless output
for any waveform incident on the array from the (assumed) known desired signal direction
[22]. Such a processor is useful for extracting impulse “bursts,” which are present only
during relatively short time intervals.
The outputs of the arrays illustrated in Figures 3-1 and 3-2 are expressed, respectively,
as
where x(t) = s(t) + n(t) is the vector of received signals that are complex valued func-
tions. The noise vector n(t) may be assumed to be stationary and ergodic and to have
both directional and thermal noise components that are independent of the signal. The
signal vector s(t) induced at the sensor elements from a single directional signal source is
assumed to be
√
s(t) = S e jω0 t (3.39)
where ω0 is the (radian) carrier frequency, and S represents the signal power. Assuming
identical antenna elements, the resulting signal component in each array element is just a
phase-shifted version (due to propagation along the array) of the signal appearing at the
first array element encountered by the directional signal source. It therefore follows that
the signal vector s(t) is written as
√ jω0 t √ jω0 t+θ1 √
sT (t) = Se , Se , . . . , S e jω0 t+θ N −1
= s(t)vT (3.40)
Consequently, the received signal vector for the conventional array of Figure 3-1 is written
as
The received signal vector (after the time-delay or phase-shift elements) for the signal
aligned array of Figure 3-2 is written as
where now v of (3.42) has been replaced by 1 = (1, 1, . . . , 1)T since the desired signal
terms in each channel are time aligned and therefore identical. The components of the
noise vector then become
In developing the optimal solutions for selected performance measures, four corre-
lation matrices will be required. These correlation matrices are defined as follows for
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 90
and
Rx x = E{x ∗ (t)xT (t)} = Rss + Rnn (3.48)
and further extended [23–25]. Suppose the desired directional signal s(t) is known and
represented by a reference signal d(t). This assumption is never strictly met in practice
because a communication signal cannot possibly be known a priori if it is to convey in-
formation; hence, the desired signal must be unknown in some respect. Nevertheless, it
turns out that in practice enough is usually known about the desired signal that a suitable
reference signal d(t) is obtained to approximate s(t) in some sense by appropriately pro-
cessing the array output signal. For example, when s(t) is an amplitude modulated signal,
it is possible to use the carrier component of s(t) for d(t) and still obtain suitable operation.
Consequently, the desired or “reference” signal concept is a valuable tool, and one can
proceed with the analysis as though the adaptive processor had a complete desired signal
characterization.
The difference between the desired array response and the actual array output signal
defines an error signal as shown in Figure 3-3:
where
⎡ ⎤
x1 (t)d(t)
⎢ x 2 (t)d(t) ⎥
⎢ ⎥
rxd =⎢ .. ⎥ (3.52)
⎣ . ⎦
x N (t)d(t)
x1 (t)
FIGURE 3-3
w1 Basic adaptive array
structure with known
desired signal.
xi (t)
wi Σ Output y(t)
xN (t)
wN
−
Error signal e (t)
+
d (t)
Reference signal
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 92
Since (3.53) is a quadratic function of w, its extremum is a minimum. Therefore, the value
of w that minimizes of E{ε 2 (t)} is found by setting the gradient of (3.53) with respect to
the weight vector equal to zero, that is,
∇w (ε 2 ) = 0 (3.54)
Since
it follows that the optimum choice for the weight vector must satisfy
Equation (3.56) is the Wiener-Hopf equation in matrix form, and its solution, wopt , is
consequently referred to as the optimum Wiener solution.
If we use d(t) = s(t) and (3.39) and (3.42), it then follows that
so that
wMSE = SR−1
xx v (3.58)
εmin
2
= S − rTxd R−1
x x rxd (3.59)
where the input signal vector may be regarded as composed of a signal component s(t)
and a noise component n(t) so that
The signal and noise components of the array output signal may therefore be written as
and
where
⎡ ⎤ ⎡ ⎤
s1 (t) n 1 (t)
⎢ s2 (t) ⎥ ⎢ n 2 (t) ⎥
⎢ ⎥ ⎢ ⎥
s(t) = ⎢ . ⎥ and n(t) = ⎢ . ⎥ (3.67)
⎣ . ⎦. ⎣ . ⎦.
s N (t) n N (t)
Consequently, the output signal power may be written as
Equation (3.70) may be recognized as a standard quadratic form and is bounded by the
minimum and maximum eigenvalues of the symmetric matrix R−1/2 nn Rss Rnn
−1/2
(or, more
−1
conveniently, Rnn Rss ) [29]. The optimization of (3.70) by appropriately selecting the
weight vector w consequently results in an eigenvalue problem where the ratio (s/n) must
satisfy the relationship [30]
s
Rss w = Rnn w (3.73)
n
in which (s/n) now represents an eigenvalue of the symmetric matrix noted already. The
maximum eigenvalue satisfying (3.73) is denoted by (s/n)opt . Corresponding to (s/n)opt ,
a unique eigenvector wopt represents the optimum element weights. Therefore
s
Rss wopt = Rnn wopt (3.74)
n opt
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 94
Substituting Rss = [ssT ] and noting that sT wopt is a scalar quantity occurring on both
sides of (3.75) that may be canceled, we get the result
T
wopt s
s= T
· Rnn wopt (3.76)
wopt Rnn wopt
T T
The ratio wopt s/wopt Rnn wopt may be seen as just a complex (scalar) number, denoted
here by c. It therefore follows that
1
wopt = R−1
nn s (3.77)
c
√
Since from (3.39) the envelope of s is just S v, it follows that
wSNR = αR−1
nn v (3.78)
where
√
S
α=
c
The maximum possible value that (s/n)opt is derived by converting the original system
into orthonormal system variables. Since Rnn is a positive definite Hermitian matrix,
it is diagonalized by a nonsingular coordinate transformation as shown in Figure 3-4.
Such a transformation is selected so that all element channels have equal noise power
components that are uncorrelated. Denote the transformation matrix that accomplishes
this diagonalization by A so that
s = As (3.79)
and
n = An (3.80)
FIGURE 3-4
s1, n1 s ′1, n ′1
Functional w ′1
representation of
orthonormal
transformation and
adaptive weight s2, n2 Matrix s ′2, n ′2 Array
combiner equivalent transformation w ′2 Σ Output
A
to system of
Figure 3-1.
sN, nN s ′N, n ′N
w ′N
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 95
For the array output of the system in Figure 3-4 to be equivalent to the output of the
system in Figure 3-1, it is necessary that
w = AT w (3.83)
Since the transformation matrix A decorrelates the various noise components and equalizes
their powers, the covariance matrix of the noise process n (t) is just the identity matrix,
that is,
The output noise power of the original system of Figure 3-1 is given by
For the output noise power of the two systems to be equivalent, it is necessary that
ARnn AT = I N (3.89)
or
Equation (3.90) simply expresses the fact that the transformation A diagonalizes and
normalizes the matrix Rnn .
The signal output of the orthonormal array system is given by (3.81). Applying the
Cauchy-Schwartz inequality (see Appendix D and [4]) to this expression immediately
yields an upper bound on the array output signal power as
where
From (3.86) and (3.91) it follows that the maximum possible value of the SNR is given by
Substituting (3.79) into (3.83) and using (3.93) and (3.90), we find it then follows that
SNRopt = sT R−1
nn s (3.94)
wSNR = αR−1
nn v
∗
(3.96)
(3.90) is now
SNRopt = sT R−1
nn s
∗
(3.98)
Designing the adaptive processor so that the weights satisfy Rnn w = αv∗ means that
the output SNR is the governing performance criterion, even for the quiescent environment
(when no jamming signal and no desired signal are present). It is usually desirable, however,
to compromise the output SNR to exercise some control over the array beam pattern
sidelobe levels. An alternative performance measure that yields more flexibility in beam
shaping is introduced in the manner described in the following material.
Suppose in the normal quiescent signal environment that the most desirable array
weight vector selection is given by wq (where now wq represents the designer’s most
desirable compromise among, for example, gain or sidelobe levels). For this quiescent
environment the signal covariance matrix is Rnnq . Define the column vector t by
On comparing (3.99) with (3.96), we see that the ratio being maximized is no longer
(3.95) but is the modified ratio given by
|w† t|2
(3.100)
w† Rnn w
This ratio is a more general criterion than the SNR and is often used for practical appli-
cations. The vector t is referred to as a generalized signal vector, and the ratio (3.100) is
called the generalized signal-to-noise ratio (GSNR). For tutorial purposes, the ordinary
SNR is usually employed, although the GSNR is often used in practice.
The adaptive array of Figure 3-1 is a generalization of a coherent sidelobe canceller
(CSLC). As an illustration of the application of the foregoing SNR performance measure
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 97
1 2 N
y1 (t) y2(t) yN (t)
w0 w1 w2 wn
x(t)
Σ
+
−
concepts, it is useful to show how sidelobe cancellation may be regarded as a special case
of an adaptive array.
The functional diagram of a standard sidelobe cancellation system is shown in
Figure 3-5. A sidelobe canceller consists of a main antenna with high gain designated
as channel o and N auxiliary antenna elements with their associated channels. The auxil-
iary antennas have a gain approximately equal to the average sidelobe level of the main
antenna gain pattern [18,31]. If the gain of the auxiliary antenna is greater than the highest
sidelobe, then the weights in the auxiliary channel are less than one. A properly designed
auxiliary antenna provides replicas of the jamming signals appearing in the sidelobes
of the main antenna pattern. These replica jamming signals provide coherent cancel-
lation in the main channel output signal, thereby providing an interference-free array
output response from the sidelobes. Jamming signals entering the main beam cannot be
canceled.
In applying the GSNR performance measure, it is first necessary to select a t column
vector. Since any desired signal component contributed by the auxiliary channels to the
total array desired signal output is negligible compared with the main channel contribution,
and the main antenna has a carefully designed pattern, a reasonable choice for t is the N + 1
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 98
component vector
⎡ ⎤
1
⎢0⎥
⎢ ⎥
⎢ ⎥
t = ⎢0⎥ (3.101)
⎢.⎥
⎣ .. ⎦
0
which preserves the main channel response signal. From (3.99) the optimum weight vector
must satisfy
Rnn w = αt (3.102)
where Rnn is the (N + 1) × (N + 1) covariance matrix of all channel input signals (in the
absence of the desired signal), and w is the N + 1 column vector of all channel weights.
Let the N × N covariance matrix of the auxiliary channel input signals (again in the
absence of the desired signal) be denoted by Rnn , and let w be the N-component column
vector of the auxiliary channel weights. Equation (3.102) may now be partitioned to yield
†
P0 w0 α
= (3.103)
Rnn w 0
Since ŵ 0 is fixed and (3.106) optimizes the GSNR ratio of ŵ 0 to the output noise power, it
follows that wopt must be minimizing the output noise power resulting from the sidelobe
response.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 99
where
for the conventional array of Figure 3-1, and an estimate of s(t) is desired. Define the
likelihood function of the input signal vector as
where P{z/y} is the probability density function for z given the event y. Thus, the likelihood
function defined by (3.109) is the negative natural logarithm of the probability density
function for the input signal vector x(t) given that x(t) contains both the desired signal
and interfering noise.
Now assume that the noise vector n(t) is a stationary, zero-mean Gaussian random
vector with a covariance matrix Rnn . Furthermore, assume that x(t) is also a stationary
Gaussian random vector having the mean s(t)v, where s(t) is a deterministic but unknown
quantity. With these assumptions, the likelihood function is written as
ŝ(t)vT R−1 T −1
nn v = v Rnn x (3.112)
vT R−1
ŝ(t) = nn
x(t) (3.113)
vT R−1
nn v
ŝ(t)v† R−1 † −1
nn v = v Rnn x(t) (3.116)
R−1
nn v
wML = (3.117)
v R−1
†
nn v
N
N
y(t) = wT z(t) = s(t) wi + w i n i (3.118)
i=1 i=1
where the n i represent the noise components after the signal aligning phase shifts. When
we constrain the sum of the array weights to be unity, then the output signal becomes
The relationship between n(t), the noise vector appearing before the signal aligning phase
shifts, and n (t) is given by
The variance of the array output remains unaffected by such a unitary transformation so
that
wT 1 = 1 (3.125)
where
so that
wMV = R−1
nn 1λ (3.129)
R−1
nn 1
wMV = (3.132)
1 R−1
T
nn 1
If complex quantities are introduced, all the foregoing expressions remain unchanged,
except the definition of the covariance matrix Rnn must be appropriate for complex vectors.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 102
wMSE = R−1 ∗ ∗ T −1
x x Sv = [Sv v + Rnn ] Sv
∗
(3.134)
From (3.135), the minimum MSE weights are the product of a matrix filter R−1 ∗
nn v (which
is also common to the other weight vector solutions) and a scalar factor. Since the MV
solution is applied only to a signal aligned array, for the other solutions to pertain to the
signal aligned array it is necessary to replace the v vector only wherever it occurs by the
1 vector.
Now consider the noise power N0 and the signal power S0 that appear at the output
of the linear matrix filter corresponding to w = R−1 ∗
nn v as follows:
N0 = w† Rnn w = vT R−1
nn v
∗
(3.136)
and
The optimal weight vector solution given in Sections 3.3.1 and 3.3.3 for the MSE and the
ML ratio performance measures are now written in terms of the previous quantities as
1 S0
wMSE = · · R−1
nn v
∗
(3.138)
N0 N0 + S0
and
1
wML = · R−1
nn v
∗
(3.139)
N0
Likewise, for the signal aligned array (where v = 1) the ML weights reduce to the unbiased,
MV weights, that is,
R−1
nn 1
wML |v = 1 = = wMV (3.140)
1T R−1
nn 1
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 103
The previous expressions show that the minimum MSE processor can be factored into a
linear matrix filter followed by a scalar processor that contains the estimates corresponding
to the other performance measures, as shown in Figure 3-6. The different optimum weight
vector previously derived be expressed by
w = βR−1
nn v
∗
(3.141)
where β is an appropriate scalar gain; hence, they all yield the same SNR, which can then
be expressed as
∗ T −1 ∗
s w† Rss w β 2 SvT R−1nn v v Rnn v −1 ∗
= † = = SvT Rnn v (3.142)
n w Rnn w β 2 vT R−1
nn v ∗
For the case of a wideband processor, it can similarly be shown that the solutions to
various estimation and detection problems are factored into a linear matrix filter followed
by appropriate scalar processing. This development is undertaken in the next section.
The fact that the optimum weight vector solutions for an adaptive array using the
different performance criteria indicated in the preceding section are all given (to within
a constant factor) by the Wiener solution underscores the fundamental importance of the
Wiener-Hopf equation in establishing theoretical adaptive array steady-state performance
limits. These theoretical performance limits provide the designer with a standard for
determining how much any improvement in array implementation can result in enhanced
array steady-state performance, and they are a valuable tool for judging the merit of
alternate designs.
Let the observation vector x consist of elements representing the outputs from each
of the array sensors xi (t), i = 1, 2, . . . , N . The likelihood ratio is then given by the ratio
of conditional probability density functions [39]
p[x/signal present]
(x) = (3.143)
p[x/signal absent]
If (x) exceeds a certain threshold η then the signal is assumed present, whereas if (x) is
less than this threshold the signal is assumed absent. The ratio (3.143) therefore represents
the likelihood that the sample x was observed, given that the signal is present relative to the
likelihood that it was observed given that the signal is absent. Such an approach certainly
comes far closer to determining the “best” system for a given class of decisions, since
it assumes at the outset that the system makes such decisions and obtains the processor
design accordingly.
It is worthwhile to mention briefly some extensions of the likelihood ratio test repre-
sented by (3.143). In many practical cases, one or several signal parameters (e.g., spatial
location, phase, or signal energy) are uncertain. When uncertain signal parameters (denoted
by θ) are present, one intuitively appealing approach is to explicitly estimate θ (denoted
by θ̂) and use this estimate to form a classical generalized likelihood ratio (GLR) [40]
optimum detection processor may also be obtained very simply by working instead with
the “detection index” d at a specified observation time T where
z = ζ + jγ (3.148)
is required to have the following two properties, which remain invariant under any linear
transformation:
1. The real part ζ and the imaginary part γ are both real random vectors having the same
covariance matrix.
2. All components of ζ and γ satisfy
The development that follows involves some useful matrix properties and generalizations
of the Schwartz inequality that are summarized for convenience in Appendix D. Gaussian
random vectors are important in the subsequent analysis, so some useful properties of both
real and complex Gaussian random vectors are given in Appendix E.
E{s} = u (3.150)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 106
where u and Rss are both known. The best estimate of s, given x, for a quadratic cost
criterion, is just the conditional mean E{s/x}.
The associated covariance matrix of ŝ is obtained from (E.15) or (E.43) from Appendix E
so that
−1
cov(ŝ) = Rss − Rss M† MRss M† + Rnn MRss (3.155)
Applying the matrix identities (D.10) and (D.11) of Appendix D to (3.154) and (3.155)
yields
−1 † −1
ŝ = u + R−1 † −1
ss + M Rnn M M Rnn (x − Mu) (3.156)
or
−1
ŝ = I + Rss M† R−1
nn M u + Rss M† R−1
nn x (3.157)
and
−1
cov(ŝ) = R−1 †
ss + M Rnn M (3.158)
or
−1
cov(ŝ) = I + Rss M† R−1
nn M Rss (3.159)
Letting u = 0 yields the result for the interesting and practical zero-mean case, then (3.156)
yields
−1 † −1
ŝ = R−1 † −1
ss + M Rnn M M Rnn x (3.160)
J = (x − Ms)† R−1 † −1
nn (x − Ms) + (s − u) Rss (s − u) (3.163)
ŝ = a + Bx (3.164)
Consequently,
By setting the gradient of (3.167) with respect to a equal to zero, the value of a that
minimizes (3.165) is given by
a = [I − BM]u (3.168)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 108
After we substitute (3.168) into (3.167) and complete the square in the manner of (D.9)
of Appendix D, it follows that
−1
E{(s − ŝ)(s − ŝ)† } = Rss − Rss M† MRss M† + Rnn MRss
−1
+ B − Rss M† MRss M† + Rnn MRss M† + Rnn
· {B† − [MRss M† + Rnn ]−1 MRss } (3.169)
The value of B that minimizes (3.165) may easily be found from (3.169)
−1
B = Rss M† MRss M† + Rnn (3.170)
By substituting the results from (3.168) and (3.170) into (3.164), the expression for ŝ
corresponding to (3.154) is once again obtained. Likewise, by substituting (3.170) into
(3.169) the same expression for the error covariance matrix as appeared in (3.155) also
results. In the Gaussian case the estimate ŝ given by (3.154) is the conditional mean,
whereas in the non-Gaussian case the same estimate is the “best” linear estimate in the
sense that it minimizes the MSE.
where m(t) is a linear transformation that represents propagation effects and any signal
distortion occurring in the sensor. In the case of ideal (nondispersive) propagation and
distortion-free electronics, the elements of m(t) are time delays, δ(t − Ti ), whereas the
scalar function s(t) is the desired signal.
When the signal and noise processes are stationary and the observation interval is long
(t → ∞), then the convolution in (3.171) is circumvented by working in the frequency
domain using Fourier transform techniques. Taking the Fourier transform of (3.171) yields
where the denote Fourier transformed variables. Note that (3.172)
is the same form as (3.147); however, it is a frequency-domain equation, whereas (3.147) is
a time-domain equation. The fact that (3.172) is a frequency-domain equation implies that
cross-spectral density matrices (which are the Fourier transforms of covariance matrices)
now fill the role that covariance matrices formerly played for (3.147).
Now apply the frequency-domain equivalent of (3.160) to obtain the solution
−1 −1 †
ˆ
(ω) = φss (ω) + † (ω)−1
nn (ω)(ω) (ω)−1
nn (ω)(ω) (3.173)
where
φss (ω)
| (ω)|2 = (3.175)
1 + φss (ω)† (ω)−1nn (ω)(ω)
and φss (ω) is simply the power of s(t) appearing at the frequency ω. In general, the
frequency response given by (3.174) is not realizable, and it is necessary to introduce
a time delay to obtain a good approximation to the corresponding time waveform ŝ(t).
Taking the inverse Fourier transform of (3.174) to obtain ŝ(t) then yields
" ∞
1
ŝ(t) = | (ω)|2 † (ω)−1
nn (ω)(ω)e
jωt
dω (3.176)
2π −∞
ˆ
(ω) = | (ω)|2 (ω)(ω) (3.177)
where
(ω) = † (ω)−1
nn (ω) (3.178)
is a 1 × n row vector of individual filters. The single filter j (ω) operates on the received
signal component j (ω) so that it performs the operation of spatial prewhitening and then
matching to the known propagation and distortion effects imprinted on the structure of the
signal. The processor obtained for this estimation problem therefore employs the principle:
first prewhiten, then match. A block diagram of the processor corresponding to (3.174) is
shown in Figure 3-7.
x1 (t ) FIGURE 3-7
η1* (w) Optimum array
processor for
x2(t ) estimation of a
η*2 (w) random signal.
Σ la (w) l2 s^ (t )
xN (t )
ηN* (w)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 110
It is seen from (3.180) that ŝ is obtained from a linear transformation on x, so the distribution
of ŝ is also Gaussian. From (3.180) it immediately follows that since E{x} = Ms, then
E{ŝ} = s (3.181)
and
−1
cov(ŝ) = M† R−1
nn M (3.182)
J = (x − Ms)† R−1
nn (x − Ms) (3.183)
ŝ = a + Bx (3.184)
E{ŝ} = s (3.185)
ŝ = a + BMs + Bn (3.187)
a=0 (3.188)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 111
and
BM = I (3.189)
and (3.186) is minimized by choosing B to minimize the quantity trace (BRnn B† ) subject
to the constraint (3.189). This constrained minimization problem is solved by introducing
a matrix Lagrange multiplier λ and minimizing the quantity
J = trace BRnn B† + [I − BM]λ + λ† [I − M† B† ] (3.191)
On completing the square of (3.191) using the formula (D.9) in Appendix D, the previous
expression is rewritten as
J = trace λ + λ† − λ† M† R−1 Mλ
nn †
+ B − λ M Rnn Rnn B − R−1
† † −1
nn Mλ (3.192)
B = λ† M† R−1
nn (3.193)
λ† M† R−1
nn M = I (3.194)
or
−1
λ† = M† R−1
nn M (3.195)
Hence
−1 † −1
B = M† R−1
nn M M Rnn (3.196)
Consequently, the estimate (3.184) once again reduces to (3.180), and the mean and
covariance of the estimate are given by (3.181) and (3.182), respectively.
where | (ω)|2 and (ω) are given by (3.175) and (3.178), respectively. Therefore, the
only difference between the optimum array processor for random signal estimation and
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 112
for unknown, nonrandom signal estimation is the presence of an additional scalar transfer
function. The characteristic prewhitening and matching operation represented by (ω) =
† (ω)−1
nn (ω) is still required.
where α = 2 for complex x and α = 1 for real x. Clearly (3.199) is written in terms of a
single exponential function so that
!
1 † † −1 † † −1 † −1
(x) = exp − α s M Rnn Ms − s M Rnn x − x Rnn Ms (3.200)
2
Since the term s† M† R−1nn Ms does not depend on any observation of x, a sufficient test
statistic for making a decision is the variable
1 † † −1 † † −1
y= s M Rnn x + x† R−1
nn Ms = Re s M Rnn x (3.201)
2
The factor α is assumed to be incorporated into the threshold level setting. The distribution
of the sufficient statistic y is Gaussian, since it results from a linear operation on x, which
in turn is Gaussian both when the signal is present and when it is absent.
When the signal is absent, then
E{y} = 0 (3.202)
s† M† R−1
nn Ms
var(y) = (3.203)
α
Likewise, when the signal is present, then
E{y} = s† M† R−1
nn Ms (3.204)
where the vector k is selected so the output SNR is maximized. The ratio given by
change in mean-squared output due to signal presence
r0 =
mean-squared output for noise alone
or equivalently
Rnn = T† T (3.208)
α[Re{k† Ms}]2
r0 = ≤ αs† M† R−1
nn Ms (3.210)
k† Rnn k
where equality results when
k† = s† M† R−1
nn (3.211)
On substituting k† of (3.211) into (3.205), it is found that the test statistic is once again
given by (3.201).
A different approach to the problem of detecting a known signal embedded in non-
Gaussian noise is to maximize the detection index given by (3.146). For y given by (3.205),
it is easily shown that
√
√ Re{k† Ms} αk† Mss† M† k
d= α √ † ≤ √ (3.212)
k Rnn k k† Rnn k
Now applying (D.17) of Appendix D to the upper bound in (3.212) in the same manner as
(D.16) was applied to the upper bound in (3.209), the detection index is shown to satisfy
√ Re{k† Ms} # † † −1
d= α √ † ≤ αs M Rnn Ms (3.213)
k Rnn k
Equality results in (3.213) when (3.211) is satisfied, so (3.201) once again results for the
test statistic.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 114
FIGURE 3-8 x1 (t )
Linear processor 1* (w)
structure for known
signal detection
problem.
x2(t )
2* (w)
Σ y (T )
T
xN (t )
N* (w)
and
" ∞
1
var[y(T )] = † (ω)nn (ω)(ω)dω (3.215)
2π −∞
The Schwartz inequality in the form of (D.14) of Appendix D may be applied to (3.217)
by letting † (ω)† (ω) play the role of f † and [† (ω)]−1 (ω)(ω)e jωT play the role of g.
It then follows that
√ " ∞ 1/2
d 2π ≤ ∗ (ω)† (ω)−1 nn (ω)(ω)(ω)dω (3.218)
−∞
Once again, the usual spatial prewhitening and matching operation represented by
† (ω)−1
nn (ω) appears in the optimum processor.
y = x† KK† x (3.228)
where K maximizes the detection index defined by (3.146). Note that since y is quadratic
in x and the variance of y in the denominator of (3.146) involves E{y 2 /signal absent}, then
fourth-order moments of the distribution for x are involved. By assuming the noise field
is Gaussian, the required fourth-order moments are expressed in terms of the covariance
matrix by applying (E.51) of Appendix E.
The numerator of (3.146) when y is given by (3.228) is written as
K† = A† M† R−1
nn (3.233)
where
y = x† R−1 † −1
nn MRss M Rnn x (3.235)
which is identical to the test statistic (3.227) obtained from the likelihood ratio in the small
signal case.
1 † †
(ω) = | (ω)|2 † (ω)−1
nn (ω)(ω) (ω)−1
nn (ω)(ω)
2
1$ $2
= $ (ω)† (ω)−1nn (ω)(ω)
$ (3.236)
2
where | (ω)|2 is given by (3.175). The structure of the optimum processor corresponding
to (3.176) is shown in Figure 3-9, where Parseval’s theorem was invoked to write (3.236)
in the time domain
" T
1
y(T ) = |z(t)|2 dt (3.237)
2T 0
If the small signal assumption is applicable, then using (3.227) instead of (3.223)
leads to
1 $ $2
(ω) = φss (ω)$† (ω)−1
nn (ω)(ω)
$ (3.238)
2
The result expressed by (3.228) is different from (3.236) only in that the φss (ω) has
replaced the scalar factor | (ω)|2 in (3.236).
x1 (t ) FIGURE 3-9
η1* (w) Optimum array
processor for
x2(t ) detection of a
η*2 (w) T random signal.
∫0 (
z (t) 1
Σ a (w) ( )2 T
) dt
Squarer
Averaging
xN (t ) circuit
ηN* (w)
−1 (w)
η(w) = m† (w) F nn
fss (w)
⎜a (w) ⎜2 =
−1 (w) m† (w)
1 + fss (w) m† (w) F nn
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 118
where ŝ is selected to maximize the conditional density function p(x/s, signal present) and
is the maximum likelihood estimate of s.
Substituting the likelihood estimate of ŝ in (3.180) into (3.240) yields the generalized
likelihood ratio
!
α † −1 † −1 −1 † −1
G (x) = exp x Rnn M M Rnn M M Rnn x (3.241)
2
The sufficient test statistic for (3.241) is obviously
1 † −1 † −1 −1 † −1
y= x Rnn M M Rnn M M Rnn x (3.242)
2
or
1 † † −1
y= ŝ M Rnn x (3.243)
2 1
where
−1 † −1
ŝ1 = M† R−1
nn M M Rnn x (3.244)
Setting R−1
ss = 0 in (3.223) for the case of a random signal vector reduces (3.223) to
(3.242).
† (ω)
˜ ss (ω)(ω)
G(ω) = (3.246)
† (ω)
˜ nn (ω)(ω)
and
where σn2 (ω) is the noise power spectral density averaged over the N sensors so that
1
σn2 (ω) = trace[nn (ω)] (3.249)
N
and
1
σs2 (ω) = trace[ss (ω)] (3.250)
N
The array gain corresponding to (3.246) is the ratio of the output signal-to-noise spectral
ratio to the input signal-to-noise spectral ratio.
Whenever the signal vector s(t) is related to a scalar signal s(t) by a known transfor-
mation m(t) such that
" t
s(t) = m(t − τ )s(τ )dτ (3.251)
−∞
˜ ss (ω) = (ω)
˜ ˜ † (ω) (3.252)
where (ω)
˜ denotes the normalized Fourier transform of m(t) so that
˜ † (ω)(ω)
˜ =N (3.253)
When
˜ ss (ω) is given by (3.252), then the array gain becomes
|† (ω)(ω)|
˜ 2
G(ω) = (3.254)
† (ω) ˜ nn (ω)(ω)
between and
˜ as described in Appendix F such that
|† (ω)(ω)|
˜ 2
cos2 (γ ) = (3.255)
(† (ω)(ω))(˜ † (ω)(ω))
˜
In a conventional beamformer the vector is chosen to be proportional to ,
˜ thus making
γ equal to zero and “matching to the signal direction.” This operation also maximizes the
array gain against spatially white noise as shown subsequently.
1
Substituting (3.252) into (3.246) and using = ˜ nn
2
yields
$ † $2
˜ − 2 (ω)(ω)
1
$ (ω) ˜ $
G(ω) = nn
(3.256)
† (ω)(ω)
Applying the Schwartz inequality (D.14) to (3.256) then gives
˜ −1
˜ † (ω)
G(ω) ≤ nn (ω)(ω)
˜ (3.257)
x1 (t ) *(w) e−jwT I
η1* (w)
x2(t ) T
∫0 (
η*2 (w) 1
Σ a (w) ( )2 T
) dt II
Averaging
xN (t ) filter
ηN* (w)
a*(w) III
Scalar
−1 (w) Wiener IV
η† (w) = m† (w) F nn
filter
fss (w) b (w) V
⎜a (w) ⎜2 =
−1 (w) m (w)
1 + fss (w) m† (w) F nn
1
b (w) = I – Detection, known signal
−1 (w) m (w)
m† (w) Fnn II – Detection, random signal, and unknown nonrandom signal
III – Estimation (maximum likelihood), random signal, and unknown
nonrandom signal
IV – Estimation (MMSE), random signal, and unknown nonrandom signal
V – Maximum array gain and minimum output variance
processor depends on the inverse of the noise cross-spectral density matrix, but in practice
the only measurable quantity is the cross-spectral matrix of the sensor outputs (which in
general contains desired signal-plus-noise components), the use of the signal-plus-noise
spectral matrix may result in performance degradation unless provision is made to obtain
a signal-free estimate of the noise cross-spectral matrix or the use of the signal-plus-noise
spectral matrix is specifically provided for. The consequences involved for providing an
inexact prewhitening and matching operation are discussed in reference [45]. It may be
further noted that the minimum mean square error (MMSE) signal estimate is different
from the maximum likelihood (distortionless) estimate (or any other estimate) only by a
scalar Wiener filter. A Wiener processor is therefore regarded as forming the undistorted
signal estimate for observational purposes before introducing the signal distortion result-
ing from the scalar Wiener filter. The nature of the scalar Wiener filter is further considered
in the Problems section.
Perturbations in the propagation process destroys the dyad nature of ss (ω), so the pro-
cessor structure must be more general than that of Figure 3-8. To determine the nature of
a more general processor structure, the optimum processor for a noncoherent narrowband
signal will be found.
Consider the problem of detecting a signal that is narrowband with unknown amplitude
and phase. Such conditions frequently occur when the received signal undergoes unknown
amplitude and phase changes during propagation. When the signal is present, the received
waveform is expressed as
where
a = unknown amplitude
θ = unknown phase
r (t) = known amplitude modulation
φ(t) = known phase modulation
ω0 = known carrier frequency
Using real notation we can rewrite (3.259) as
where
s1 = a cos θ
s2 = −a sin θ
m 1 (t) = r (t) cos[ω0 t + φ(t)] = f (t) cos(ω0 t) − g(t) sin(ω0 t) (3.261)
m 2 (t) = r (t) sin[ω0 t + φ(t)] = f (t) sin(ω0 t) + g(t) cos(ω0 t) (3.262)
and
where
#
r (t) = f 2 (t) + g 2 (t)
!
−1 g(t)
φ(t) = tan
f (t)
Equation (3.259) may now be rewritten as
x = ms + n (3.265)
where
s1
m = [m 1 m 2 ], s= , and var(n) = σn2
s2
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 123
The results of Section 3.4.5 for the detection of an unknown, nonrandom signal may now
be applied, for which the sufficient test statistic is given by (3.242). For real variables
(3.242) becomes
1 T −1 T −1 −1 T −1
y= x Rnn M M Rnn M M Rnn x (3.266)
2
Note that m 1 (t) and m 2 (t) are orthogonal functions having equal energy so that |m 1 |2 =
|m 2 |2 , and m 1 (t) · m 2 (t) = 0. Consequently, when the functions are sampled in time to
form xT = [x(t1 ), x(t2 ), . . . ], miT = [m i (t1 ), m i (t2 ), . . . ], then (3.266) for the test statistic
is rewritten as
1 m1 −1
y = x2 m1T m2T |m1 |2 σn2
2 m2
1 2 −1 2 2
= σn |m1 |2 m1T x + m2T x (3.267)
2
−1
Since the scalar factor 12 σn2 |m1 |2 in (3.267) is incorporated into the threshold setting
for the likelihood ratio test, a suitable test statistic is
2 2
z = m1T x + m2T x (3.268)
For a continuous time observation, the test statistic given by (3.268) is rewritten as
" T 2 " T 2
z= m 1 (t)x(t)dt + m 2 (t)x(t)dt (3.269)
0 0
The test statistic represented by (3.269) is conveniently expressed in terms of the “sine”
and “cosine” components, f (t) and g(t). Using (3.261) and (3.262), we then can rewrite
(3.269) as shown.
" T " T 2
z= f (t) cos(ω0 t)x(t)dt − g(t) sin(ω0 t)x(t)dt
0 0
" T " T 2
+ f (t) sin(ω0 t)x(t)dt + g(t) cos(ω0 t)x(t)dt (3.270)
0 0
The previous test statistic suggests the processor shown in Figure 3-11.
The optimum detector structure for a noncoherent signal shown in Figure 3-11 leads
to the more general array processor structure illustrated in Figure 3-12. This more general
processor structure is appropriate when propagating medium or receiving mechanism
perturbations cause the plane wave desired signal assumption to hold no longer or when it
is desired to match the processor to a signal of arbitrary covariance structure. A matched
array processor like that of Figure 3-12 using matrix weighting is referred to as an element
space-matched array processor [48].
Regarding the quantity Ms in (3.147) as the signal having arbitrary characteristics,
then it is appropriate to define an array gain in terms of the detection index at the output of
a general quadratic processor [17]. For Gaussian noise, the results summarized in (3.231)
may be used to give
trace † (ω)ss (ω)(ω)
G= 1/2 (3.271)
trace († (ω)nn (ω)(ω))2
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 124
Quadratic
processor
f(T− t )
sin w 0t
Matched filter
FIGURE 3-12 x1 (t )
Structure for general ( )2
array processor. x2(t )
† (w) Σ y (t)
xN (t ) ( )2
Matrix Quadratic
filter processor
It may be seen that (3.271) reduces to (3.246) when is a column vector. Under perturbed
propagation conditions, ss is given by [13]
where the matrix has N rows and r columns, and r denotes the rank of the matrix ss .
For a plane wave signal, the cross-spectral density matrix ss has rank one, and the dyad
structure of (3.258) holds. The array gain given by (3.271) may be maximized in the same
manner as (3.246), now using (D.16) instead of (D.14), to yield [13]
' ( 2 )*1/2
G ≤ trace ss (ω)−1 nn (ω) (3.273)
where equality obtains when the matrix is chosen to be a scalar multiple of † −1
nn . It
follows that the maximum gain cannot be achieved by any having less than r columns.
z FIGURE 3-13
Crossed Dipole
Array Geometry.
L From Compton,
z=+
2 IEEE Trans. Ant. &
q Prop., Sept. 1981.
z =– L
2
x
Eq FIGURE 3-14
Polarization Ellipse.
From Compton,
IEEE Trans. Ant. &
a Prop., Sept. 1981
b
Ef
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 126
E = Eφ φ + Eθ θ (3.276)
The electric field components for a given state of polarization are related to the polarization
ellipse and are given by (aside from a common phase factor)
Eφ = A cos γ (3.277a)
Eθ = A sin γ ejη (3.277b)
The previous equations relating the four variables α, β, γ , and η have a geometrical
relationship shown on the Poincare sphere in Figure 3-15. For a given point on the sphere,
M, the quantities 2γ , 2β, and 2α form the sides of a right spherical triangle. 2γ is the
side of the triangle between the point M and the point labeled H (H is the point on
the sphere representing horizontal polarization. The point V correspondingly represents
vertical polarization and lies 180◦ removed from H in the equatorial plane of the sphere).
The side 2β lies along the equator and extends to the point where it meets the perpendicular
projection of M onto the equator. The angle η then lies between the sides 2γ and 2β. The
special case α = 0 corresponds to linear polarization in which case the point M lies on the
equator: if in addition β = 0 then M lies at the point H, and only Eφ is nonzero so we have
FIGURE 3-15
Poincaré Sphere.
From Compton, M
IEEE Trans. Ant. &
V
Prop., Sept. 1981
2g
2a
h
H 2b
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 127
horizontal polarization; if, however, β = π/2 , then M lies at the point V, and we have vertical
polarization. The poles of the sphere correspond to circular polarization (α = ±45◦ ), with
clockwise circular polarization (α = +45◦ ) at the upper pole. We may conclude that five
parameters characterize the polarization of an electric field: (θ, φ, α, β, and A).
The electric field in Figure 3-16 has an Eφ component with an x-component of
−Eφ sin φ and a y-component of Eφ cos φ. The component of Eθ lying in the x-y plane
is Eθ cos θ, while the z-component is −Eθ sin θ . It immediately follows that the x and y
components of Eθ are given by Eθ cos θ cos φ and Eθ cos θ sin φ. The conversion of the
electric field from spherical to rectangular coordinates is given by
E = (Eθ cos θ cos φ − Eφ sin φ)x̂ + (Eθ cos θ sin φ + Eφ cos φ)ŷ
− (Eθ sin θ)ẑ (3.279)
A crossed dipole antenna has two or three orthogonal dipoles that can excite two or
three orthogonal linear polarizations. Adaptive crossed dipoles control the currents fed to
the dipoles to change the polarization and radiation pattern. The array geometry is shown in
Figure 3-16, where the top dipole consists of elements x1 (vertical) and x2 (horizontal along
the x-axis) and the bottom dipole consists of elements x3 (vertical) and x4 (horizontal along
the x-axis). We may then write the response of the four linear elements to the incoming
electric field of Figure 3-16 (now including time and space phase factors) in vector form as
X = Aej(ωt+ψ) U (3.281)
yr FIGURE 3-16
Diagram of an
adaptive crossed
fr dipole
xr
communications
system.
zt
Eq r
qt
qr
zr
Ef
yt
ft
xt
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 128
where
⎡ ⎤
(− sin γ sin θ e jη )e j p
⎢ (sin γ cos θ cos φ e jη − cos γ sin φ)e j p ⎥
U=⎢
⎣
⎥
⎦ (3.282)
(− sin γ sin θ e jη )e− j p
(sin γ cos θ cos φ e jη − cos γ sin φ)e− j p
and ω is the carrier signal frequency, ψ is the carrier phase of the signal at the coordinate
origin when t = 0, and p is the phase shift of the signals at the two dipoles with respect to
the origin as a result of spatial delay where
πL
p= cos θ (3.283)
λ
Now suppose there is a desired signal specified by (Ad , θd , φd , αd , βd ) and an interfer-
ence signal by (Ai , θi , φi , αi , βi ). If thermal noise is present on each signal then the total
signal vector is given by
X = Xd + Xi + Xn (3.284a)
= Ad e j ( t+ϕd ) Ud + Ai e j ( t+ϕi ) Ui + Xn (3.284b)
where Ud , Ui are given by (3.282). The corresponding signal covariance matrix is then
given by
xx = d + i + n (3.285a)
where
d = E Xd XdH = A2d Ud UH d (3.285b)
i = E Xi XH
i = A 2
i U U
i i
H
(3.285c)
and
n = σ 2 I (3.285d)
A small difference in polarization between two received signals (measured in terms of
the angular separation between two points on the Poincare sphere) is enough for an adaptive
polarized array to provide substantial protection against the interference signal. Compton
[34] shows( that) the signal-to-noise plus interference ratio (SINR) is nearly proportional
to cos2 Md2Mi , where Md Mi represents the (radian) distance on the sphere between the
two polarizations.
The currents fed to the crossed dipoles in an adaptive communications system can be
adjusted to optimize the received signal [54]. Figure 3-16 shows the transmit antenna is
at an angle of (θr , ϕr ) from the receive antenna, and the receive antenna is at an angle of
(θt , ϕt ). Maximum power transfer occurs when θt = 0◦ and θr = 0◦ or when the main
beams point at each other. Controlling the complex signal weighting at the dipoles of the
transmit and receive antennas modifies the directivity and polarization of both antennas.
If the crossed dipole has three orthogonal dipoles, then the transmitted electric field is
written as
ωμe− jkr
E = −j Ix L x cos θ cos ϕ + I y L y cos θ sin ϕ − Iz L z sin θ θ̂
4πr
+ − Ix L x sin ϕ + I y L y cos ϕ φ̂
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 129
where
r = distance from the origin to the field point at (x,y,z),
L x,y,z = dipole length in the x, y, and z directions
ω = radial frequency
k = wave number
μ = permeability
Ix,y,z = constant current in the x, y, or z direction
The directivity and polarization loss factors are given by
where 0 ≤ PLF ≤ 1 with PLF = 1 a perfect match. The t and r subscripts represent
transmit and receive, respectively.
When the transmitting antenna is a pair of orthogonal crossed dipoles in the x-y plane
that emits a circularly polarized field in the z-direction, then increasing θt causes the
transmitted electric field to transition from circular polarization through elliptical until
linear polarization results at θt = 90◦ . Assume the transmit antenna is a ground station
that tracks a satellite and has currents given by Ix = 1, I y = j, and Iz = 0. The transmit
antenna delivers a circularly polarized signal at maximum directivity to the moving receive
antenna. If the receive antenna has two crossed dipoles, then the maximum receive power
transfer occurs when the receive antenna is directly overhead of the transmit antenna.
If the receive antenna remains circularly polarized as it moves, then the power received
decreases, because the two antennas are no longer polarization matched. The loss in power
transfer comes from a reduction in the directivity and the PLF. If the currents at each dipole
are optimally weighted using an adaptive algorithm, then the power loss is compensated.
The curve in Figure 3-17 results from finding the optimum weights for a range of angles.
The link improvement over the nonadaptive antenna is as much as 3dB at θr = 90◦ .
4 FIGURE 3-17
Power loss as a
3 function of receive
Adapted dipoles angle when two
2 dipoles are adaptive.
Link loss (dB)
0
Unadapted dipoles
−1
−2
−3
0 20 40 60 80
q r (degrees)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 130
FIGURE 3-18 4
Power loss as a
function of receive 3
Adapted dipoles
angle when three
dipoles are adaptive. 2
0
Unadapted dipoles
−1
−2
−3
0 20 40 60 80
q r (degrees)
A third orthogonal dipole adds another degree of freedom to the receive antenna as
shown in Figure 3-18. Since a tri-dipole antenna is able to produce any θ or ϕ polarization,
the receive antenna adapts to maximize the received signal in any direction, and there is
no change in the link budget as a function of θr . The three orthogonal dipoles compensate
for changes in the receive antenna directivity and polarization as it moves and produces
up to 6 dB improvement in the link budget at θr = 90◦ .
processor are extremely important. Since the spatial properties of the steady-state solu-
tions resulting from various different performance measures are either identical or very
similar, the question of which array performance measure to select for a given application
is usually not very significant; rather, it is the temporal properties of the adaptive control
algorithm to be used to reach the steady-state solutions that are of principal concern to the
designer. The characteristics of adaptive control algorithms to be used for controlling the
weights in the array pattern-forming network are therefore highly important, and it is to
these concerns that Part 2 of this book is addressed.
Finally, the concept of a polarization sensitive array using polarization sensitive el-
ements was introduced. Such an array possesses more degrees of freedom than a con-
ventional array and permits signals arriving from the same direction to be distinguished
from one another. The notion of polarization “distance” is introduced using the Poincare
sphere. The polarization and gain of crossed dipoles can be adapted to maximize a dynam-
ically changing communications link. Adaptively adjusting the currents on the dipoles in
a crossed dipole system can significantly improve the link budget.
3.8 PROBLEMS
1. Use of the Maximum SNR Performance Measure in a Single Jammer Encounter [18]
The behavior of a linear adaptive array controlled by the maximum SNR performance measure
when the noise environment consists of a single jammer added to the quiescent environment
is of some practical interest. Consider a linear, uniformly spaced array whose weights are
determined according to (3.96). Assume that the quiescent noise environment is characterized
by the covariance matrix
⎡ ⎤
pq
⎢ 0 ⎥
⎢ ⎥
⎢ pq ⎥
Rnnq =⎢
⎢
⎥ = pq Ik
⎥
⎢ .. ⎥
⎣ 0 . ⎦
pq
K
G q (β) = ak e j (k−1)(β−βs )
k=1
where
2π d
β= sin θ
λ
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 132
In view of the expressions for wq and G q (β) previously given, it follows that
⎡ ⎤
1
⎢1⎥
UHwq = G q (β J )⎢ ⎥
⎣ .. ⎦
.
1
RJJ = p J H∗ UH
(c) Using the result from part (b), show that the optimum weight vector is given by
pJ
w = wq − H∗ UHwq
pq + Kp J
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 133
(d) Let the optimum quiescent weight vector for a desired signal located in the direction θs
from mechanical boresight be expressed as
⎡ ⎤
a1
⎢ a2 e− jβs ⎥
⎢ a e− j2βs ⎥
wq = ⎢
⎢.
3 ⎥
⎥
⎣ .. ⎦
− j (K −1)βs
aK e
where
2πd
βs = sin θs
λ
and the ak represent weight amplitudes. If the various ak are all equal, the resulting array
beam pattern is of the form (sinKx)/(sinx).
It then follows from the definition of H in part (a) that
⎡ ⎤
a1
⎢ a2 e j (β J −βs ) ⎥
Hwq = ⎢
⎣ ..
⎥
⎦
.
j (K −1)(β J −βs )
aK e
From the foregoing expressions and the results of part (c), show that the pattern obtained
with the optimum weight vector may be expressed as
PJ
G(β) = bT w = bT wq − G q (β J )b∗J
Pq + KP J
where b J is just b with the variable β J replacing β.
(e) Recalling from part (d) that bT wq = G q (β), show that bT b∗J = C(β − β J ) where
!
(K − 1)x sin K x/2
C(x) = exp j
2 sin x/2
(f) Using the fact that bT wq = G q (β) and the results of parts (d) and (e), show that
PJ
G(β) = G q (β) − G q (β J )C(β − β J )
Pq + KP J
This result expresses the fact that the array beam pattern of an adaptively controlled linear
array in the presence of one jammer consists of two parts. The first part is the quiescent
pattern G q (β), and the second part (which is subtracted from the first part) is a (sin Kx)/(sin
x)-shaped cancellation beam centered on the jammer.
(g) From the results of part (e) it may be seen that C(x)|x=0 = K . Using this fact show that
the gain of the array in the direction of the jammer is given by
Pq
G(β J ) = G q (β J )
Pq + KP J
With the array weights fixed at wq , the array pattern gain in the direction of the jammer
would be G q (β J ). The foregoing result therefore shows that the adaptive control reduces
the gain in the direction of the jammer by the factor
Pq 1
=
Pq + KP J 1 + K (PJ /Pq )
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 134
(h) The proper measure of performance improvement against the jammer realized by the
introduction of adaptively controlled weights is the cancellation ratio defined by
γ
=
1 − γ · (J/N )q
where
noise power Pn . Under these conditions, each diagonal entry of R yy is equal to Pn + PIa ,
and hence trace (R yy ) = N (Pn + PIa ). Since the largest eigenvalue of R yy is less than the
trace of R yy , it follows that Padd ≤ N (Pn + PIa )||||2 . Assume that the weight errors are
due to quantization errors where the quanta size of the in-phase and quadrature channel
is q. Under worst-case conditions, each complex weight component quantization error is
identical and equal to (q/2)(i ± j). Hence, show that
N 2q 2
Padd ≤ (Pn + PIa )
2
(d) The interference-to-noise ratio (PI /Pn )main for the main element is related to the ratio
(PIa /Pn )aux for each auxiliary element by
PI PIa
= |α|2
Pn main
Pn aux
where α is the average voltage gain of the main antenna in the sidelobe region. Assume
that α is given by
α = q · 2 B−1
where B represents the number of bits available to implement the quantizer. Show that by
assuming (PI /Pn )main 1 and using the results of part (c), then
Padd N2
= R0
Prmin 22B−1
where
Prmin
R0 =
Pn
(e) The principle use of a worst-case analysis is to determine the minimum number of bits
required to avoid any degradation of the SIR performance. When the number of available
quantization bits is significantly less than the minimum number predicted by worst-case
analysis, a better prediction of SIR performance is obtained using an MSE analysis. The
only change required for an MSE analysis is to treat ||||2 in an appropriate manner.
Assuming the weight errors of the in-phase and quadrature channels are independent and
uniformly distributed, show that
1
||||2average = ||||2max
6
and develop corresponding expressions for Padd and Padd /Prmin .
3. Wiener Linear MMSE Filtering for Broadband Signals [56]
Consider the scalar signal x(t)
where the desired signal s(x) and the noise n(t) are uncorrelated. The signal x(t) is to be passed
through a linear filter h(t) so that the output y(t) will best approximate s(t) in the MMSE sense.
The error in the estimate y(t) is given by
show that
" ∞
E{n 2 (t)} = H (ω)H ∗ (ω)φnn (ω)dω
−∞
where φnn (ω) = {rnn (τ )}, the spectral density function of n(t). Likewise, show that
" ∞
E{[s (t) − s(t)] } = 2
[1 − H (ω)][1 − H ∗ (ω)] · φss (ω)dω
−∞
so that
" ∞
E{e2 (t)} = H (ω)H ∗ (ω)φnn (ω) + [1 − H (ω)][1 − H ∗ (ω)]φss (ω) dω
−∞
(d) It is desired to minimize E{e2 (t)} by appropriately choosing H (ω). Since the integrand
of the expression for E{e2 (t)} appearing in part (c) is positive for all choices of H (ω), it
is necessary only to minimize the integrand by choosing H (ω). Show that by setting the
gradient of the integrand of E{e2 (t)} with respect to H (ω) equal to zero, there results
φss (ω)
Hopt (ω) =
φss (ω) + φnn (ω)
which is the optimum scalar Wiener filter. Therefore, to obtain the scalar Wiener filter, it
is necessary to know only the signal spectral density and the noise spectral density at the
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 137
point in question. For scalar processes, the foregoing result may be used to obtain the MMSE
signal estimate by introducing the appropriate scalar Wiener filter at the point where a MMSE
signal estimate is desired. It is easy to show that the corresponding result for vector processes
is given by
and when x(t) = s(t) + n(t) where s(t) = vd(t) then opt (ω) is expressed as
(a) Complete the square of the above expression for ???z to obtain the equivalent result
z = k† + λ† H† −1 −1 † †−1 † †
x x x x x x Hλ + k − λ Hx x Hλ − λ g − g λ
Since k appears only in the previously given first quadratic term, the minimizing value of
k is obviously that for which the quadratic term is zero or
kopt = −−1
x x Hλ
(b) Use the constraint equation H† k = g to eliminate λ from the result for kopt obtained in
part (a) and thereby show that
† −1
−1
kopt = −1
x x H H x x H g
(c) For the general array processor of Figure 3-12, the output power is given by
z = trace(† x x ), and the problem of minimizing z subject to multiple linear constraints
of the form H† = L is handled by using a matrix Lagrange multiplier, , and considering
z = trace † x x + † [H† − L] + [† H − L† ]
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 138
Complete the square of the previous expression to show that the optimum solution is
† −1
−1
opt = −1
x x H H x x H L
† −1
for which opt x x opt = L† H† −1
xx H L
5. Introduction of Linear Constraints to the MSE Performance Measure [15]
The theoretically optimum array processor structure for maximizing (or minimizing) some
performance measure may be too complex or costly to fully implement. This fact leads to
the consideration of suboptimal array processors for which the processor structure is properly
constrained within the context of the signal and interference environment.
The K-component weight vector w is said to be linearly constrained if
f = c† w
The number of linear, orthonormal constraints on w must be less than K if any remaining
degrees of freedom are to be available for adaptation.
(a) The MSE is expressed as
E{|y − y A |2 }
where y A = w† x.
Consequently
Append the constraint equation to the MSE by means of a complex Lagrange multiplier
to form
Take the gradient of the foregoing expression with respect to w and set the result equal to
zero to obtain
wopt = R−1 ∗
x x rx y + λ c
(b) Apply the constraint f = c† w to the result in part (a) thereby obtaining a solution for λ
f ∗ − r†x y R−1
xx c
λ=
c† R−1
xx c
This solution may be substituted into the result of part (a) to obtain the resulting solution
for the constrained suboptimal array processor.
f = C† w
where the matrix C has the vector ci as its ith column, and the set {ci /i = 1, 2, . . . , M} must
be a set of orthonormal constraint vectors.
Append the multiple linear constraint to the expected output power of the array output
signal to determine the optimum constrained solution for w.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 139
g = w† Qw
Appending the constraint equation to the expected output power with a complex Lagrange
multiplier yields
J = w† Rx x w + λ[g − w† Qw]
Take the gradient of the previous expression with respect to w and set the result equal to
zero (to obtain the extremum value of w) to obtain
R−1 −1
x x Qw = λ w
(b) Note that the previous result for w is satisfied when w is an eigenvector of R−1 x x Q and
λ is the corresponding eigenvalue. Therefore, maximizing
(minimizing) J corresponds to
selecting the largest (smallest) eigenvalue of R−1
xx Q . It follows that a quadratic constraint
is regarded simply as a means of scaling the weight vector w in the array processor.
(c) Append both a multiple linear constraint and a quadratic constraint to the expected output
power of the array to determine the optimum constrained solution for w. Note that once
again the quadratic constraint merely results in scaling the weight vector w.
8. Introduction of Single-Point, Multiple-Point, and Derivative Constraints to a Minimum
Power Output Criterion for a Signal-Aligned Array [49]
(a) Consider the problem of minimizing the array output power given by
P0 = w† Rx x w
subject to the constraint C† w = f. Show that the solution to this problem is given by
† −1
−1
wopt = R−1
x x C C Rx x C f
(b) A single-point constraint corresponds to the case when the following conditions hold:
w† 1 = N
N R−1
xx 1
wopt =
1T R−1
xx 1
Under the single-point constraint, the weight vector minimizes the output power in all
directions except in the look (presteered) direction.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 140
C = P = [e1 , e, e2 ]
where e1 and e2 are direction vectors referenced to the beam axis on either side, and
⎡ †
⎤
e1 e
f = P† e = ⎣ N ⎦
†
e2 e
where C† w = f. Show that the corresponding optimal constrained weight vector is given by
† −1
−1
wopt = R−1
x x P P Rx x P P† e
(d) Derivative constraints are used to maintain a broader region of the main lobe by specifying
both the response on the beam axis and the derivatives of the response on the beam axis.
The constraint matrix now has dimension N × k and is given by
C = D = e0 , e0 , e0 , . . .
(a) Using the expression for p(x/α), show that the likelihood functional is expressed as
∂M −1 ∂M
y(α) = x† M−1 M x − trace M−1
∂α ∂α
where T is a weighting matrix that incorporates all geometrical properties of the array. In
particular, for linear arrays the T matrix has elements given by the following:
sin θ
For bearing estimation: ti j = (z i − z j )
sin2 θ 2
For range estimation: ti j = − z − z 2j
2r 2 i
where θ = signal bearing with respect to the array normal
z n = position of nth sensor along the array axis (the z-axis)
= velocity of signal propagation
r = true target range
Show that for an array having length L, and with K 1 equally spaced sensors then
62
For bearing estimation: [trace(TT† )]−1 =
K 2 L 2 sin2 θ
452 r 4
For range estimation: [trace(TT† )]−1 =
2L 4 K 2 sin4 θ
The foregoing results show that the range estimate accuracy is critically dependent on the
true range, whereas the bearing estimate is not (except for the range dependence of the
SNR). The range estimate is also more critically dependent on the array aperture L than
the bearing estimate.
10. Suboptimal Bayes Estimate of Target Angular Location [9]
The signal processing task of extracting information concerning parameters such as target an-
gular location can also be carried out by means of Bayes estimation. A Bayes estimator is just
the expectation of the variable being estimated conditioned on the observed data, that is,
" ∞
û = E{u k /x} = u k p(u k /x)du
−∞
where u k = sin θ denotes the angular location of the kth target, x denotes the observed data
vector, and the a posteriori probability density function p(u/x) is rewritten by application of
Bayes’s rule as
p(x/u) p(u)
p(u/x) =
p(x)
The optimum estimators that result using the foregoing approach are quite complex, re-
quiring the evaluation of multiple integrals that may well be too lengthy for many portable
applications. Consequently, the development of a simple suboptimum estimator that approxi-
mates the optimum Bayes estimator is of practical importance.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 142
(a) The lowest-order nonlinear approximation to the optimum Bayes location estimator is
given by [9]
û = xT B
û = [û 1 , û 2 , . . . , û N ]
Assuming a (2K + 1) element array where each element output is amplified and detected
with synchronous quadrature detectors, then the output of the mth pair of quadrature de-
tectors is
and
x = [x ym x zn ]
in which each index pair m, n occupies a separate row. Define a 2K (2K + 1) × 1 column
vector of coefficients
(k)
bk = bmn , m = n
where, again, each index pair m, n occupies a separate row. The full matrix of coefficients
B is then given by the 2K (2K + 1) × N matrix
B = [b1 , b2 , . . . , b N ]
Show that, by choosing the coefficient matrix B so that the resulting estimator is orthogonal
to the estimation error, that is, E[ûT (u − û)] = 0, then B is given by
With B selected as indicated, the MSE in the location estimator for the kth target is then
given by
k = E[u k (u k − û k )] = E u 2k − [E{xu k }]T · [E{xxT }]−1 · [E{xu k }]
(b) Consider a one-target location estimation problem using a two-element array. From the
results of part (a), it follows that
û = b12 x y1 x z2 + b21 x y2 x y1
s ym = α cos mπ u + β sin mπ u
and
The aforementioned signal and noise models correspond to a Rayleigh fading environ-
ment and additive noise due to scintillating clutter. Show that the optimum coefficients for
determining û are given by
E 2 {u sin π u}
MSE = E{u 2 } −
2E{sin π u} + 1/γ + 2/2γ 2
2
R−1 −1
x x Rss Rx x
Kopt = 2
tr R−1
x x Rss
(b) Show that when Rss has rank one and is consequently given by the dyad Rss = vv† , then
where
h = scalar · v† R−1
nn
This result demonstrates that under plane wave signal assumptions the element space-
matched array processor degenerates into a plane wave matched processor with a quadratic
detector.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 144
FIGURE 3-19
Points Md and Mi on
the Poincare Sphere
From Compton,
Md
IEEE Trans. On Ant.
and Prop, 2yd
September, 1981. nd–ni Mi
2yi
H
12. Signal-to-Interference–plus-Noise Ratio (SINR) for a Two Element Polarized Array [49]
Let Md and Mi be points on the Poincare sphere representing the polarizations of the desired
and interference signals, respectively as shown in Figure 3-20. From the figure it is seen that
2γd , 2γi , and the arc Md Mi form the sides of a spherical triangle. The angle ηd − ηi is the angle
opposite side Md Mi . Using a well-known spherical trigonometric identity, the sides on the
spherical triangle are related by cos 2γd cos 2γi + sin 2γd sin 2γi cos(ηd − ηi ) = cos(Md Mi ).
(a) For the case where both the desired and interference signals arrive from broadside, and
θd = φd = θi = φi = 90◦ , show that
Md Mi
|Ud Ui |2 = 2[1 + cos(Md Mi )] = 4 cos2
2
Pd A2 $ $2 A2 $ $2
(b) Using the fact that SINR = , where Pd = d $UTd w$ , Pi = i $UiT w$ , and
Pi + Pn 2 2
σ2 2
Pn = |w| where w denotes the weight vector, use the matrix inversion lemma (m.i.l.)
2
to show that the SINR is written as
$ H $2
$Ud Ui $
SINR = ξd UHd Ud −
ξi−1 + UiH Ui
A2d
where ξd = is the desired signal-to-noise ratio and similarly for ξi .
σ2
(c) Using the result of part (a) in the SINR of part (b) show that
Md Mi
4 cos2
SINR = ξd 2 − 2
ξi−1 + 2
where the U denote the steering vectors for the desired and interference signals, respectively
It is first desired to find −1 .
(a) Use the matrix inversion lemma,
τ −1 + β −1 = ZB−1 ZH
Applying the m.i.l. to the quantity σ 2 I + Ai2 Ui UiH , where B = σ 2 I, Z = Ui , and β = −Ai2 ,
show that
−1
−1 −1 −1 1 1
τ = (ZB Z − β ) H
= Ui UiH + 2
σ 2 Ai
and hence
2
H −1 1 Ui UiH
σ I + Ai Ui Ui
2
= 2 I − −1
σ ξi + Ui UiH
τ Ud UHd τ Ui UHd
+ 4 −1 U d UH
+ Ui UHd
σ ξi + UiH Ui i
σ 4 ξi−1 + UiH Ui
where
1 1 1 Ui UHd Ud UiH
τ −1 = 2 + 2 Ud UHd − 2
Ad σ σ ξi−1 + UiH Ui
(c) Use the fact that w = −1 S where S = E{X∗ R(t)} = Ar Ad U∗d along with the result of
part (b) to show that
Ar A d τ γ Ui UHd
w= 1− U∗d − −1
Ui∗ .
σ2 σ 2 A2d ξi + UiH Ui
where Ar denotes reference signal amplitude, and γ is given by
Ui UHd Ud UiH
γ = A2d Ud UHd −
ξi−1 + UiH Ui
A2d $$ T $$2
(f) Use the result of part (e) along with Pd = Ud w to show that
2
2
A2 γ
Pd = r
2 σ2 + γ
(g) In a manner similar to part (d) show that
Ar A d τ γ ξi−1
UiT w = 1− 2 2 Ui UHd −1
σ2 σ Ad ξi + UiH Ui
Ai2 $$ T $$2
(h) Use the result of part (g) along with Pi = Ui w to show that
2
2 2
Ar2 A2d Ai2 1 ξi−1
Pi = UiH Ud UHd Ui
2 σ +γ
2
ξi−1 + UiH Ui
(i) Finally, from part (c) show that
2
1 γ H H ξi−1
|w| =
2
Ar2 A2d − U Ui U U d 2
σ2 + γ A2d d i
ξ −1 + UH Ui i i
σ 2
(j) From the fact that Pn = |w|2 , use the result of part (i) to show that
2
2 / 2 0
A2 A2 A2 1 γ ξi−1 H H ξi−1
Pn = r d i − Ud Ui Ui Ud −1
2 σ2 + γ A2d ξi + UiH Ui
(k) Combining the results of part (h) with part (j), show that
2
Ar2 Ai2 1
Pi + Pn = γ ξi−1
2 σ2 + γ
(l) Using the results of part (f) for Pd , show that the SINR reduces to the form
Pd γ
SINR = = 2.
Pi + Pn σ
Using the definition of γ given in part (c), it follows that an alternative expression for the
SINR is given by
$ H $2
$Ud Ui $
SINR = ξd UHd Ud −
ξi−1 + UiH Ui
3.9 REFERENCES
[1] P. L. Stocklin, “Space-Time Sampling and Likelihood Ratio Processing in Acoustic Pressure
Fields,” J. Br. IRE, July 1963, pp. 79–90.
[2] F. Bryn, “Optimum Signal Processing of Three-Dimensional Arrays Operating on Gaussian
Signals and Noise,” J. Acoust. Soc. Am., Vol. 34, No. 3, March 1962, pp. 289–297.
[3] D. J. Edelblute, J. M. Fisk, and G. L. Kinneson, “Criteria for Optimum-Signal-Detection
Theory for Arrays,” J. Acoust. Soc. Am., Vol. 41, January 1967, pp. 199–206.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 147
[4] H. Cox, “Optimum Arrays and the Schwartz Inequality,” J. Acoust. Soc. Am., Vol. 45, No. 1,
January 1969, pp. 228–232.
[5] B. Widrow, P. E. Mantey, L. J. Griffiths, and B. B. Goode, “Adaptive Antenna Systems,” Proc.
IEEE, Vol. 55, December 1967, pp. 2143–2159.
[6] L. J. Griffiths, “A Simple Algorithm for Real-Time Processing in Antenna Arrays,” Proc.
IEEE, Vol. 57, October 1969, pp. 1696–1707.
[7] N. Owsley, “Source Location with an Adaptive Antenna Array,” Naval Underwater Systems
Center, Rept. NL-3015, January 1971.
[8] A. H. Nuttall and D. W. Hyde, “A Unified Approach to Optimum and Suboptimum Processing
for Arrays,” U.S. Navy Underwater Sound Laboratory Report 992, April 1969, pp. 64–68.
[9] G. W. Young and J. E. Howard, “Applications of Spacetime Decision and Estimation Theory
to Antenna Processing System Design,” Proc. IEEE, Vol. 58, May 1970, pp. 771–778.
[10] H. Cox, “Interrelated Problems in Estimation and Detection I and II,” Vol. 2, NATO Advanced
Study Institute on Signal Processing with Emphasis on Underwater Acoustics, Enchede, The
Netherlands, August 12–23, 1968, pp. 23-1–23-64.
[11] N. T. Gaarder, “The Design of Point Detector Arrays,” IEEE Trans. Inf. Theory, part 1, Vol. IT-
13, January 1967, pp. 42–50; part 2, Vol. IT-12, April 1966, pp. 112–120.
[12] J. Capon, “Applications of Detection and Estimation Theory to Large Array Seismology,”
Proc. IEEE, Vol. 58, May 1970, pp. 760–770.
[13] H. Cox, “Sensitivity Considerations in Adaptive Beamforming,” Proceedings of NATO
Advanced Study Institute on Signal Processing, Loughborough, England, August 1972,
pp. 621–644.
[14] A. M. Vural, “An Overview of Adaptive Array Processing for Sonar Applications,” IEEE 1975
EASCON Record, Electronics and Aerospace Systems Convention, September 29–October 1,
pp. 34.A–34 M.
[15] J. P. Burg, “Maximum Entropy Spectral Analysis,” Proceedings of NATO Advanced Study
Institute on Signal Processing, August 1968.
[16] N. L. Owsley, “A Recent Trend in Adaptive Spatial Processing for Sensor Arrays: Constrained
Adaptation,” Proceedings of NATO Advanced Study Institute on Signal Processing, Lough-
borough, England, August 1972, pp. 591–603.
[17] R. R. Kneiper, et al., “An Eigenvector Interpretation of an Array’s Bearing Response Pattern,”
Naval Underwater Systems Center Report No. 1098, May 1970.
[18] S. P. Applebaum, “Adaptive Arrays,” IEEE Trans. Antennas Propag., Vol. AP-24, No. 5,
September 1976, pp. 585–598.
[19] A. Papoulis, Probability, Random Variables, and Stochastic Processes, New York, McGraw-
Hill, 1965, Ch. 8.
[20] R. A. Wiggins and E. A. Robinson, “Recursive Solution to the Multichannel Filtering Prob-
lem,” J. Geophys. Res., Vol. 70, No. 8, April 1965, pp. 1885–1891.
[21] R. V. Churchill, Introduction to Complex Variables and Applications, New York, McGraw-
Hill, 1948, Ch. 8.
[22] M. J. Levin, “Maximum-Likelihood Array Processing,” Lincoln Laboratories, Massachusetts
Institute of Technology, Lexington, MA, Semiannual Technical Summary Report on Seismic
Discrimination, December 31, 1964.
[23] J. Chang and F. Tuteur, Symposium on Common Practices in Communications, Brooklyn
Polytechnic Institute, 1969.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 148
[24] P. E. Mantey and L. J. Griffiths, “Iterative Least-Squares Algorithms for Signal Extraction,”
Proceedings of the 2nd Hawaii Conference on System Sciences, 1969, pp. 767–770.
[25] R. L. Riegler and R. T. Compton, Jr., “An Adaptive Array for Interference Rejection,” Proc.
IEEE, Vol. 61, No. 6, June 1973, pp. 748–758.
[26] S. P. Applebaum, Syracuse University Research Corporation Report SPL TR 66-1, Syracuse,
NY, August 1966.
[27] S. W. W. Shor, “Adaptive Technique to Discriminate Against Coherent Noise in a Narrow-Band
System,” J. Acous. Soc. Am., Vol. 34, No. 1, pp. 74–78.
[28] R. T. Adams, “An Adaptive Antenna System for Maximizing Signal-to-Noise Ratio,”
WESCON Conference Proceedings, Session 24, 1966, pp. 1–4.
[29] R. Bellman, Introduction to Matrix Analysis, New York, McGraw-Hill, 1960.
[30] R. F. Harrington, Field Computation by Moment Methods, New York, Macmillan, 1968, Ch. 10.
[31] B. D. Steinberg, Principles of Aperture and Array System Design, New York, Wiley, 1976,
Ch. 12.
[32] L. J. Griffiths, “Signal Extraction Using Real-Time Adaptation of a Linear Multichannel
Filter,” SEL-68-017, Technical Report No. 6788-1, System Theory Laboratory, Stanford
University, February 1968.
[33] R. T. Lacoss, “Adaptive Combining of Wideband Array Data for Optimal Reception,” IEEE
Trans. Geosci. Electron., Vol. GE-6, No. 2, May 1968, pp. 78–86.
[34] C. A. Baird, Jr., and J. T. Rickard, “Recursive Estimation in Array Processing,” Proceedings
of the Fifth Asilomar Conference on Circuits and Systems, 1971, pp. 509–513.
[35] C. A. Baird, Jr., and C. L. Zahm, “Performance Criteria for Narrowband Array Process-
ing,” IEEE Conference on Decision and Control, December 15–17, 1971, Miami Beach, FL,
pp. 564–565.
[36] W. W. Peterson, T. G. Birdsall, and W. C. Fox, “The Theory of Signal Detectability,” IRE
Trans., PGIT-4, 1954, pp. 171–211.
[37] D. Middleton and D. Van Meter, “Modern Statistical Approaches to Reception in Communi-
cation Theory,” IRE Trans., PGIT-4, 1954, pp. 119–145.
[38] T. G. Birdsall, “The Theory of Signal Detectability: ROC Curves and Their Character,”
Technical Report No. 177, Cooley Electronics Laboratory, University of Michigan, Ann Arbor,
MI, 1973.
[39] W. J. Bangs and P. M. Schultheiss, “Space-Time Processing for Optimal Parameter Estima-
tion,” Proceedings of NATO Advanced Study Institute on Signal Processing, Loughborough,
England, August 1972, pp. 577–589.
[40] L. W. Nolte, “Adaptive Processing: Time-Varying Parameters,” Proceedings of NATO
Advanced Study Institute on Signal Processing, Loughborough, England, August 1972,
pp. 647–655.
[41] D. O. North, “Analysis of the Factors which Determine Signal/Noise Discrimination in
Radar,” Report PTR-6C, RCA Laboratories, 1943.
[42] L. A. Zadeh and J. R. Ragazzini, “Optimum Filters for the Detection of Signals in Noise,”
IRE Proc., Vol. 40, 1952, pp. 1223–1231.
[43] G. L. Turin, “An Introduction to Matched Filters,” IRE Trans., Vol. IT-6, June 1960,
pp. 311–330.
[44] H. L. Van Trees, Detection, Estimation, and Modulation Theory, Part I, New York, Wiley,
1968, Ch. 2.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 149
[45] H. Cox, “Resolving Power and Sensitivity to Mismatch of Optimum Array Processors,”
J. Acoust. Soc. Am., Vol., 54, No. 3, September 1973, pp. 771–785.
[46] H. L. Van Trees, “Optimum Processing for Passive Sonar Arrays,” IEEE 1966 Ocean Elec-
tronics Symposium, August, Hawaii, pp. 41–65, 1966.
[47] L. W. Brooks and I. S. Reed, “Equivalence of the Likelihood Ratio Processor, the Maximum
Signal-to-Noise Ratio Filter, and the Wiener Filter,” IEEE Trans. Aerosp. Electron. Syst.,
Vol. AES-8, No. 5, September 1972, pp. 690–691.
[48] A. M. Vural, “Effects of Perturbations on the Performance of Optimum/Adaptive Arrays,”
IEEE Trans. Aerosp. Electron. Syst., Vol. AES-15, No. 1, January 1979, pp. 76–87.
[49] R. T. Compton, “On the Performance of a Polarization Sensitive Adaptive Array,” IEEE Trans.
Ant. & Prop., Vol. AP-29, September 1981, pp. 718–725.
[50] A. Singer, “Space vs. Polarization Diversity,” Wireless Review, February 15, 1998,
pp. 164–168.
[51] M. R. Andrews, P. P. Mitra, and R. deCarvalho, “Tripling the Capacity of Wireless Communi-
cations Using Electromagnetic Polarization,” Nature, Vol. 409, January 18, 2001, pp. 316–318.
[52] R. E. Marshall and C. W. Bostian, “An Adaptive Polarization Correction Scheme Using
Circular Polarization,” IEEE International Antennas and Propagation Society Symposium,
Atlanta, GA, June 1974, pp. 395–397.
[53] R. T. Compton, “The Tripole Antenna: An Adaptive Array with Full Polarization Flexibility,”
IEEE Trans. Ant. & Prop., Vol. AP-29, November 1981, pp. 944–952.
[54] R. L. Haupt, “Adaptive Crossed Dipole Antennas Using a Genetic Algorithm,” IEEE Trans.
Ant. & Prop., Vol. AP-52, August 2004, pp. 1976–1982.
[55] R. Nitzberg, “Effect of Errors in Adaptive Weights,” IEEE Trans. Aerosp. Electron. Syst.,
Vol. AES-12, No. 3, May 1976, pp. 369–373.
[56] J. P. Burg, “Three-Dimensional Filtering with an Array of Seismometers,” Geophysics, Vol. 29,
No. 5, October 1964, pp. 693–713.
[57] A. Ishide and R.T. Compton, Jr., “On grating Nulls in Adaptive Arrays,” IEEE Trans. on Ant.
& Prop., Vol. AP-28, No. 4, July 1980, pp. 467–478.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:36 150
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 151
P A R T II
Adaptive Algorithms
CHAPTER
Gradient-Based Algorithms
4
' $
Chapter Outline
4.1 Introductory Concepts . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 154
4.2 The LMS Algorithm . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 158
4.3 The Howells–Applebaum Adaptive Processor . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 171
4.4 Introduction of Main Beam Constraints . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 191
4.5 Constraint for the Case of Known Desired Signal Power Level . . . . . . . . . . . . . . . . . . . 199
4.6 The DSD Algorithm . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 201
4.7 The Accelerated Gradient Approach (AG) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 209
4.8 Gradient Algorithm with Constraints . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 213
4.9 Simulation Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 224
4.10 Phase-Only Adaptive Nulling Using Steepest Descent. . . . . . . . . . . . . . . . . . . . . . . . . . . 227
4.11 Summary and Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 228
4.12 Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 230
4.13 References . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 235
& %
Gradient algorithms are popular, because they are simple, easy to understand, and solve
a large class of problems. The performance, (w), and adaptive weights determine the
nature of the performance surface. When (w) is a quadratic function of the weight
settings, then it is a bowl-shaped surface with a minimum at the “bottom of the bowl.” In
this case, local optimization methods, such as gradient methods, can find the bottom. In
the event that the performance surface is irregular, having several relative optima or saddle
points, then the transient response of the gradient-based minimum-seeking algorithms get
stuck in a local minimum. The gradient-based algorithms considered in this chapter are
as follows:
1. Least mean square (LMS)
2. Howells–Applebaum loop
3. Differential steepest descent (DSD)
4. Accelerated gradient (AG)
5. Steepest descent for power minimization
153
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 154
Variations of these algorithms come from introducing constraints into the adjustment rule,
and one section develops the procedure for deriving such variations. Finally, changes in
the modes of adaptation are discussed, illustrating how two-mode adaptation enhances the
convergence.
wopt = R−1
x x rxd (4.6)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 155
w2 FIGURE 4-1
Steepest descent
with very small step
size (overdamped
case).
Initial guess
w1
w2 FIGURE 4-2
Steepest descent
with large step size
(underdamped
case).
Initial guess
w1
2
ξmin = d (t) − wopt
T
rxd (4.7)
The method of steepest descent begins with an initial guess of the weight vector
components. Having selected a starting point, we then calculate the gradient vector and
perturb the weight vector in the opposite direction (i.e., in the direction of the steepest
downward slope). Contour plots of a quadratic performance surface (corresponding to a
two-weight adjustment problem) are shown in Figures 4-1 and 4-2. In these figures the
MSE is measured along a coordinate normal to the plane of the paper. The ellipses in these
figures are contours of constant MSE. The gradient is orthogonal to these constant value
contours (pointing in the steepest direction) at every point on the performance surface. If
the steepest descent uses small steps, it is “overdamped,” and the path taken to the bottom
appears continuous as shown in Figure 4-1. If the steepest descent uses large steps, it is
“underdamped,” and each step is normal to the error contour as shown in Figure 4-2.
The discrete form of the method of steepest descent is [1]
w(k + 1) = w(k) − s ∇ e2 (k) (4.8)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 156
where
w(k) = old weight vector guess at time kT
+ 1) = new weight vector guess at time (k + 1)T
w(k
∇ e2 (k) = gradient vector of the MSE determining the direction in which to move
from w(k)
s = step size
Substituting the gradient of (4.5) into (4.8) then yields
2Rxx
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 157
To diagonalize the flow graph of Figure 4-3, consider the expression for the MSE
given by (4.4). Using wopt and ξmin in (4.6) and (4.7), the MSE becomes
T
E e2 (k) = ξ(k) = ξmin + w(k) − wopt Rx x w(k) − wopt (4.10)
Since the matrix Rx x is real, symmetric, and positive definite (for real variables), it is
diagonalized by means of a unitary transformation matrix Q so that
Rx x = Q−1 Q (4.11)
where is the diagonal matrix of eigenvalues, and Q is the modal square matrix of
eigenvectors. If Q is constructed from normalized eigenvectors, then it is orthonormal so
that Q−1 = QT , and the MSE becomes
T
ξ(k) = ξmin + w(k) − wopt QT Q w(k) − wopt (4.12)
Now define
Qw(k) = w (k) (4.13)
Qwopt = wopt (4.14)
2L
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 158
2l p
where λ p is the pth eigenvalue of Rx x . The impulse response of (4.16) is found by letting
r p (z) = 1 and taking the inverse Z -transform of the resulting output −1 {w p (z)}. It follows
that the impulse response is of the form
where
1
α p = − ln(1 − 2s λ p ) (4.17)
T
and T = one iteration period. The time response of (4.17) is a stable system when
Since Rx x is positive definite, λ p > 0 for all p. Consequently, the stability of the mul-
tidimensional flow graph of Figure 4-4 is guaranteed if and only if λ p = λmax in (4.19)
and
The stability of the steepest descent adaptation process is therefore guaranteed so long as
1
> s > 0 (4.21)
λmax
sonar systems), it is pointless to try to generate a fictitious desired signal. Thus, the LMS
algorithm described here is usually employed to improve communications system per-
formance. The LMS algorithm is exactly like the method of steepest descent except that
now changes in the weight vector are made in the direction given by an estimated gradient
vector instead of the actual gradient vector. In other words, changes in the weight vector
are expressed as
where
w(k) = weight vector before adaptation step
w(k + 1) = weight vector after adaptation step
s = step size that controls rate of convergence and stability
∇[ξ(k)]
ˆ = estimated gradient vector of ξ with respect to w
The adaptation process described by (4.22) attempts to find a solution as close as
possible to the Wiener solution given by (4.6). It is tempting to try to solve (4.6) directly,
but such an approach has several drawbacks:
1. Computing and inverting an N × N matrix when the number of weights N is large
becomes more challenging as input data rates increase.
2. This method may require up to [N (N + 3)]/2 autocorrelation and cross-correlation
measurements to find the elements of Rx x and rxd . In many practical situations, such
measurements must be repeated whenever the input signal statistics change.
3. Implementing a direct solution requires setting weight values with high accuracy in
open loop fashion, whereas a feedback approach provides self-correction of inaccurate
settings, thereby giving tolerance to hardware errors.
To obtain the estimated gradient of the MSE performance measure, take the gradient
of a single time sample of the squared error as follows:
ˆ k = ∇[ξ(k)] = 2e(k)∇[e(k)]
∇ (4.23)
Since
it follows that
so that
ˆ k = −2e(k)x(k)
∇ (4.26)
It is easy to show that the gradient estimate given by (4.26) is unbiased by considering the
expected value of the estimate and comparing it with the gradient of the actual MSE. The
expected value of the estimate is given by
ˆ k } = −2E{x(k)[d(k) − xT (k)w(k)]}
E{∇ (4.27)
= −2[rxd (k) − Rx x (k)w(k)] (4.28)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 160
so the expected value of the estimated gradient equals the true value of the gradient of
the MSE.
Substituting the estimated gradient of (4.26) into the weight adjustment rule of (4.22)
then yields the weight control rule
The LMS algorithm given by (4.32) can be rewritten for complex quantities as
w(k + 1) − w(k)
= 2ks e(k)x∗ (k) (4.33)
t
where t is the elapsed time between successive iterations, and s = ks t. In the limit
as t → 0, (4.33) yields an equivalent differential equation representation of the LMS
algorithm that is appropriate for use in continuous systems as
dw(t)
= 2ks e(t)x∗ (t) (4.34)
dt
Equation (4.34) can also be written as
t
w(t) = 2ks e(τ )x∗ (τ )dτ + w(0) (4.35)
0
FIGURE 4-6
Complex conjugate
Analog realization
of the LMS weight
adjustment
algorithm.
wi (t) xi (t)
ƒ ith Element
Integrator
Other
2ks Σ element
outputs
e(t)
_
+
d(t)
Array
output
FIGURE 4-7
Complex conjugate
Digital realization
of the LMS weight
adjustment
xi (k) algorithm.
w i (k + 1) Unit time w i (k)
Σ delay (Δt)
ith Element
Δt
Other
2ks Δt Σ element
outputs
e(k)
_
+
d(k)
Array
output
Now let
Starting with an initial guess w(0), the (k + 1)th iteration of (4.40) yields
When the magnitude of all the terms in the diagonal matrix [I − 2ks t] are less than
one, then
Therefore, the first term of (4.42) vanishes after a sufficient number of iterations, and the
summation factor in the second term of (4.42) becomes
k
1
lim [I − 2ks t]i = −1 (4.44)
k→∞
i=0
2ks t
1
lim E{w(k + 1)} = 2ks tQ−1 −1 Qrxd
k→∞ 2ks t
= R−1
x x rxd (4.45)
This result shows that the expected value of the weight vector in the LMS algorithm does
converge to the Wiener solution after a sufficient number of iterations.
Since all the eigenvalues in are positive, it follows that all the terms in the afore-
mentioned diagonal matrix, I − 2ks t, have a magnitude less than one provided that
where
N
trace[Rx x ] = E{x† (k)x(k)} = E{|xi |2 } = PIN (4.48)
i=1
FIGURE 4-8 w2
Steepest descent
transient response
with widely diverse
eigenvalues.
w1
Initial guess
Since the square of an exponential function is an exponential having half the time
constant of the original exponential function, it follows that when all the time constants
are equal the MSE learning curve is an exponential having the time constant
τ 1
τMSE = = (4.52)
2 4(ks t)λ
In general, of course, the eigenvalues of Rx x are unequal so that
τp 1
τ pMSE = = (4.53)
2 4(ks t)λ p
where τ pMSE is the time constant for the MSE learning curve, τ p is the time constant in the
weights, and λ p is the eigenvalue of the pth normal mode. The adaptive process uses one
signal data sample/iteration, so the time constant expressed in terms of the number of data
samples is
Plots of actual experimental learning curves look like noisy exponentials—an effect
due to the inherent noise that is present in the adaptation process. A slower adaptation
rate (i.e., the smaller the magnitude of ks ) has a smaller noise amplitude that corrupts the
learning curve.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 165
where ξ = E{e2 }. The LMS algorithm misadjustment can be evaluated for a specified
value of ks t by considering the noise associated with the gradient-estimation process.
Assume that the adaptive process converged to a steady state in the neighborhood of
the MSE surface minimum point. The gradient-estimation noise of the adaptive algorithm
at the minimum point (where the true gradient is zero) is just the gradient estimate itself.
Therefore, the gradient noise vector g is given by
g(k) = ∇(k)
ˆ = −2e(k)x(k) (4.56)
When the weight vector is optimized (w(k) = wopt ), then the error e(k) is uncorrelated
with the input vector x(k). If e(k) and x(k) are Gaussian processes, then not only are
they uncorrelated at the minimum point of the MSE surface, but they are also statistically
independent. With these conditions (4.57) becomes
Adaptation based on noisy gradient estimates results in noise in the weight vector.
Recall that the noise-free method of steepest descent is described by the iterative relation
where s is the constant that controls stability and rate of convergence, and ∇(k) is the
gradient at the point on the performance surface corresponding to w = w(k). Following
Widrow and McCool [13], subtract wopt from both sides of (4.60), and define v(k) =
w(k) − wopt to obtain
which represents a first-order vector difference equation with a stochastic driving function—
s g(k). Multiplying (4.64) by Q produces
After initial transients have died out and the steady state is reached, v (k) responds to
the stationary driving function −s g (k) in the manner of a stationary random process.
The absence of any cross-coupling in the primed normal coordinate system means that the
components of both g (k) and v (k) are mutually uncorrelated, and the covariance matrix
of g (k) is therefore diagonal. To find the covariance matrix of v (k) consider
Taking expected values of both sides of (4.66) (and noting that v (k) and g (k) are un-
correlated since v (k) is affected only by gradient noise from previous iterations), we
find
cov [v (k)] = (I − 2s )cov [v (k)](I − 2s ) + 2s cov [g (k)]
−1
= 2s 4s − 42s 2 cov [g (k)] (4.67)
In practical applications, the LMS algorithm uses a small value for s , so that
s I (4.68)
With (4.68) satisfied, the squared terms involving s in (4.67) may be neglected, so
s −1
cov [v (k)] = cov [g (k)] (4.69)
4
Using (4.59), we find
s −1
cov [v (k)] = (4ξmin ) = s ξmin I (4.70)
4
Therefore, the covariance of the steady-state noise in the weight vector (near the minimum
point of the MSE surface) is
Without noise in the weight vector, the actual MSE experienced would be ξmin . The
presence of noise in the weight vector causes the steady-state weight vector solution
to randomly meander about the minimum point. This random meandering results in an
“excess” MSE— that is, an MSE that is greater than ξmin . Since
2
ξ(k) = d (k) − 2rTxd w(k) + wT (k)Rx x w(k) (4.72)
where
2
ξmin = d (k) − wopt
T
rxd (4.73)
wopt = R−1
x x rxd (4.74)
N
2
E{vT (k)v (k)} = λp E vp (k) (4.77)
p=1
Using (4.70) to recognize that E{[vp (k)]2 } is just s ξmin for each p, we see it then
follows that
N
E{vT (k)v (k)} = s ξmin λp
p=1
= s ξmin tr (Rx x ) (4.78)
Furthermore
N
N
1 N 1
s tr (Rx x ) = s λp = = (4.81)
p=1 p=1
4τ pMSE 4 τ pMSE av
where
1
N
1 1
= (4.82)
τ pMSE av N p=1 τ pMSE
the network responsible for generating the reference signal (when the reference signal is
derived from the array output) is discussed in [15].
LMS algorithm adaptation with an injected pilot signal causes the array to form a
beam in the pilot-signal direction. This array beam has a flat spectra response and linear
phase shift characteristic within the passband defined by the spectral characteristic of the
pilot signal. Furthermore, directional noise incident on the array manifests as correlated
noise components that the array will respond by producing beam pattern nulls in the noise
direction within the array passband.
Since injection of the pilot signal could “block” the receiver (by rendering it insensitive
to the actual signal of interest), mode-dependent adaptation schemes have been devised
to overcome this difficulty. Two such adaptation algorithms are discussed in the following
section.
dN
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 170
in the desired look direction to have a specific amplitude and phase shift at that frequency.
On the other hand, if the pilot signal is chosen to be the sum of several sinusoids having
different frequencies, then the adaptation process forces the array gain and phase in the
desired look direction to have specific values at each one of the pilot-signal frequencies.
Finally, if several pilot signals corresponding to different look directions are added to-
gether, then the array gain is simultaneously constrained at the various frequencies and
angles corresponding to the different pilot signals selected. In summary, the two-mode
adaptation process minimizes the total power of all signals received that are uncorrelated
with the pilot signals while constraining the gain and phase of the array beam to values
corresponding to the frequencies and angles dictated by the pilot-signal components.
Figure 4-10 illustrates a practical one-mode method for simultaneously eliminating all
noises uncorrelated with the pilot signal and forming a desired array beam. The circuitry
of Figure 4-10 circumvents the difficulty of being unable to receive the actual signal, while
the processor is connected to the pilot-signal generator by introducing an auxiliary adaptive
processor. For the auxiliary adaptive processor, the desired response is the pilot signal,
and both the pilot signal and the actual received signals enter the processor. A second
processor performs no adaptation (its weights are slaved to the weights of the adaptive
processor) and generates the actual array output signal. The slaved processor inputs do
not contain the pilot signal and can therefore receive the transmitted signal at all times.
In the one-mode adaptation method, the pilot signal is on continuously so the adap-
tive processor that minimizes the MSE forces the adaptive processor output to closely
reproduce the pilot signal while rejecting all signals uncorrelated with the pilot signal.
Second Array
(slaved) output
processor
dN
+
xN +
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 171
The adaptive processor therefore preserves the desired array directivity in the look direc-
tion (over the pilot-signal passband) while placing nulls in the directions of noise sources
(over the noise frequency bands).
+ + + + + +
_ _ _ _ _ _
b*1 b*2 b*3 b*4 b*5 b*6
G G G G G G
6
Array output = Σ wi xi
i=1
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 172
wqT = [w q1 , w q2 , . . . , w q N ] (4.87)
and
where
bk∗ = ck w qk (4.95)
where
dz k
N
τ0 + zk = γ xk∗ w i xi (4.98)
dt i=1
γ = k2G (4.99)
The constant γ represents a conversion-factor gain constant that is assumed to be the same
for all the loops. It is convenient to use (4.96) to convert from z k to w k , so that (4.98) now
becomes
dw k N
∗ ∗
τ0 + w k = bk − γ xk w i xi (4.100)
dt i=1
Using matrix notation, we may write the complete set of N differential equations corre-
sponding to (4.100) as
dw
τ0 + w = b∗ − γ x∗ wT x (4.101)
dt
N
Since (wT x) = (xT w) = i=1 w i xi , the bracketed term in (4.101) can be rewritten as
[x∗ wT x] = [x∗ xT ]w (4.102)
The expected (averaged) value of x∗ xT yields the input signal correlation matrix
Rx x = E{x∗ xT } (4.103)
The averaged values of the correlation components forming the elements of Rx x are
given by
⎧
⎪
⎪ I
⎪
⎨ |J i |2 exp[ jψi (l − k)] l = k (4.104)
∗
xk xl = i=1 I
⎪
⎪
⎩|x k | = |n k | + |J i |2 l = k
2 2
⎪ (4.105)
i=1
Since the correlation matrix in the absence of the desired signal is the sum of the quies-
cent receiver noise matrix Rnnq and the individual interference source matrixes Rnni , it
follows that
I
Rnn = Rnnq + Rnni (4.106)
i=1
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 174
and
⎡ ⎤
1 e jψi e j2ψi · · ·
⎢ e− jψi 1 e jψi · · · ⎥
⎢ ⎥
⎢ − j2ψi ⎥
Rnni = |J i |2 ⎢e e− jψi 1 · · · ⎥ (4.108)
⎢ ⎥
⎢ .. ⎥
⎣ . ⎦
1
Substituting Rnn of (4.106) into (4.101) and rearranging terms, the final expression for the
adaptive weight matrix differential equation becomes
dw
τ0 + [I + γ Rnn ]w = b∗ (4.109)
dt
where I is the identity matrix.
In general, Rnn is not diagonal, so multiplying Rnn by a nonsingular orthonormal
model matrix, Q, results in a simple transformation of coordinates that diagonalizes Rnn .
The resulting diagonalized matrix has diagonal elements that are the eigenvalues of the
matrix Rnn . The eigenvalues of Rnn are given by the solutions of the equation
|Rnn − λi I| = 0, i = 1, 2, . . . , N (4.110)
Rnn ei = λi ei (4.111)
These eigenvectors (which are normalized to unit length and are orthogonal to one another)
make up the rows of the transformation matrix Q, that is,
⎡ ⎤
e11 e12 e13 · · · ⎡ ⎤
⎢ e21 e22 e23 · · ·⎥ ei1
⎢ ⎥ ⎢ ei2 ⎥
⎢ ⎥ ⎢ ⎥
Q = ⎢ e31 e32 e33 · · ·⎥ , where ei = ⎢ . ⎥ (4.112)
⎢ . ⎥ ⎣ .. ⎦
⎣ .. ⎦
ei N
eN 1 eN 2 eN 3 · · ·
Once Rnn is diagonalized by the Q-matrix transformation, there results
⎡ ⎤
λ1 0 0
⎢0 λ 0⎥
⎢ 2 ⎥
⎢ ⎥
∗
[Q Rnn Q ] = ⎢
T ⎢ 0 0 λ 0 ⎥ (4.113)
3 ⎥
⎢ .. .. ⎥
⎣ . . ⎦
··· ··· · λN
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 175
b = Qb (4.119)
where the kth component of b is determined by the kth eigenvector appearing in Q.
The Q-coordinate transformation operating on both x and b∗ suggests an equivalent
circuit representation for the system that is illustrated in Figure 4-12b, where an equivalent
“orthonormal adaptive array” system is shown alongside a simplified representation of the
real system in Figure 4-12a. There are a set of weights forming the weight vector w in the
orthonormal system, and the adaptive weight matrix equation for the equivalent system is
dw
τ0 + I + γ Rnn w = b∗ (4.120)
dt
where
Rnn = E{x∗ xT } = (4.121)
This diagonalization results in an orthonormal system, a set of independent linear dif-
ferential equations, each of which has a solution when the eigenvalues are known. Each
of the orthonormal servo loops in the equivalent system responds independently of the
other loops, because the xk input signals are orthogonalized and are therefore completely
uncorrelated with one another. The weight equation for the kth orthonormal servo loop
can therefore be written as
dw k
τ0 + (1 + γ λk )w k = bk∗ (4.122)
dt
Note that the equivalent servo gain factor can be defined from (4.122) as
μk = γ λk (4.123)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 176
FIGURE 4-12
Equivalent circuit x1 x2 x3 x4 x5 x6
representations for a
six-element adaptive
Σ
array system.
a: Real adaptive
array system.
b: Equivalent w1 w2 w3 w4 w5 w6
orthonormal
adaptive array Amp Amp Amp Amp Amp Amp
system. From int int int int int int
Gabriel, Proc. IEEE,
b*1 b*2 b*3 b*4 b*5 b*6
February 1976.
Output
(a)
x1 x2 x3 x4 x5 x6
Q − Transformation network
x′5
x′1 x′2 x′3 x′4 x′5 x′6
so the equivalent servo gain factors for the various orthonormal loops are now determined
by the eigenvalues of the input signal covariance matrix. The positive, real eigenvalues
λk correspond to the square of a signal voltage amplitude, and any given eigenvalue is
proportional to the power appearing at the orthonormal network output port.
For the input beam steering vector b∗ , the output desired signal power is given by
where the signal vector x is assumed to be composed only of quiescent receiver channel
noise plus the directional noise signal components due to external sources of interference.
The signal-to-noise performance measure is therefore just a ratio of the aforementioned
two quadratic forms
s |wT b|2 w† [b∗ bT ]w
= = (4.126)
n |wT x|2 w† Rnn w
The optimum weight vector (see Chapter 3) that yields the maximum SNR for (4.126) is
1
wopt = R−1 b∗ (4.127)
(constant) nn
On comparing (4.127) with (4.45), it is seen that both the LMS and maximum SNR
algorithms yield precisely the same weight vector solution (to within a multiplicative con-
stant when the desired signal is absent) provided that rxd = b∗ , since these two vectors
play exactly the same role in determining the optimum weight vector solution. Conse-
quently, adopting a specific vector rxd for the LMS algorithm is equivalent to selecting b∗
for the maximum SNR algorithm, which represents direction of arrival information—this
provides the relation between a reference signal and a beam steering signal for the LMS
and maximum SNR algorithms to yield equivalent solutions.
From the foregoing discussion, it follows that the optimum orthonormal weight is
1
w k opt = bk∗ (4.128)
μk
Substitute (4.123) and (4.128) into (4.122) results in
dw k
τ0 + (1 + μk )w k = μk w k opt (4.129)
dt
For a step-function change in the input signal the solution may be written as follows:
w k (t) = w k (0) − w k (∞) e−αk t + w k (∞) (4.130)
where
μk
w k (∞) = w k opt (4.131)
1 + μk
1 + μk
αk = (4.132)
τ0
In the foregoing equations w k (∞) represents the steady-state weight, w k (0) is the initial
weight value, and αk is the transient decay factor. The adaptive weight transient responses
can now be determined by the eigenvalues. The kth orthonormal servo loop may be
represented by the simple type-0 position servo illustrated in Figure 4-13.
To relate the orthonormal system weights w k to the actual weights w k note that the
two systems shown in Figure 4-12 must be exactly equivalent so that
FIGURE 4-13
Type-O servo model
for kth orthonormal
adaptive control _
+ R
loop. From Gabriel, ′
wkopt mk w′k
Proc. IEEE, February
1976.
mk = glk
C
Consequently
w = QT w (4.134)
From (4.134) it follows that the solution for the kth actual weight can be written as
w k = e1k w 1 + e2k w 2 + · · · + e Nk w N (4.135)
where
λ0 = |n 0 |2 (4.137)
so the smallest eigenvalue is simply equal to the receiver channel noise power. This smallest
eigenvalue then defines the minimum servo gain factor μmin as
μmin = γ λ0 (4.138)
Since the quiescent steady-state weight w(∞) must by definition be equal to wq , (4.131),
(4.132), and (4.95) can be applied to yield
1 ck
w qk = b∗ = w qk
1 + μmin k 1 + μmin
or
ck = (1 + μmin ) (4.139)
From (4.130)–(4.131) and (4.123), it follows that the effective time constant with
which the kth component of w converges to its optimum value is τ0 /(1 + γ λk ). In effect,
λmin determines how rapidly the adaptive array follows changes in the noise environment.
Equation (4.135) shows that each actual weight can be expressed as a weighted sum of
exponentials, and the component that converges most slowly is the λmin component.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 179
When the adaptive array in Figure 4-11 operates with a distributed external noise field,
the loop convergence is very slow for some angular noise distributions [23]. Furthermore,
if γ is increased or τ0 is decreased to speed the weight convergence, the loop becomes
“noisy.” Slow weight convergence occurs whenever trace(Rnn )/λmin is large, and in these
cases there is no choice of γ and τ0 that yields rapid convergence without excessive
loop noise. These facts suggest that the effects of noise on the solutions represented by
(4.128)–(4.132) are important.
Griffiths provides a discrete form of the Howells–Applebaum weight update formula
given by [24]
w(k + 1) = w(k) + γ μb∗ − x∗ (k)x† (k)w(k) (4.140)
where γ and μ are constants. The weights converge if γ is less than one over the largest
eigenvalue. Compton shows that this is equivalent to [25]
1
0<γ < (4.141)
PIN
where PIN is the total received power in (4.48). If γ is close to 1/PIN then convergence is
fast, but weight jitter is large. The weight jitter causes SNR fluctuations of several dB at
steady state. If γ is small, then weight jitter is small, but the convergence is slow. A gain
constant of [26]
1
γ = (4.142)
2.5PIN
was found to provide a reasonably stable steady-state weights and rapid conversion.
w=w+ξ (4.143)
Rnn = Rnn + (4.144)
where now w and Rnn denote average values. The adaptive weights must satisfy
dw
τ0 + (I + γ Rnn )w = b∗ (4.145)
dt
Substitute the values w and Rnn into (4.145) and subtract the result from the equation
resulting with (4.143) and (4.144) substituted into (4.145) to give
dξ
τ0 + I + γ Rnn ξ = −γ w (4.146)
dt
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 180
Premultiplying (4.146) by the transformation matrix Q∗ and using the fact that Q∗ QT = I,
then
dζ 1
+ (I + γ )ζ = −βQ∗ w = u (4.147)
dt τ0
where
ζ = Q∗ ξ (4.148)
γ
β= (4.149)
τ0
Equation (4.147) represents a system of N independent linear differential equations of
which the nth component can be written as
dζn + σn ζn dt = u n dt (4.150)
where
1 + γ λn
σn = (4.151)
τ0
u n = u n (τ, w, ζ ) = (−βQ∗ w)n (4.152)
Multiplying (4.150) by the factor eσn t and integrating each term from t0 to t then yields
t
ζn (t) = ζn (t0 ) exp[−σn (t − t0 )] + t0 e−σn (r −τ )
(4.153)
· u n (τ, w, ζ )dτ
If only the steady-state case is considered, then the weights are near their mean steady-
state values. The steady-state solution for variations in the element weights can be obtained
from (4.153) by setting t0 = −∞ and ignoring any effect of the initial value ζn (t0 ) to give
∞
ζn (t) = e−σn τ u n (t − τ )dτ (4.154)
0
One important measure of the noise present in the adaptive loops is the variance of
the weight vector denoted by var(w):
N
var(w) = E |wn − wn | = E{ξ † ξ }
2
(4.155)
n=1
where N is the dimension of the weight vector (or the number of degrees of freedom in
the adaptive array system). Now since ζ = Q∗ ξ , (4.155) becomes
var(w) = E{ζ † ζ } (4.156)
The elements of the covariance matrix of ζ (t) in (4.156) are obtained from (4.154) and
the definition of u n
∞ ∞
∗
E{ζ j ζk } = β 2
dτ1 E [Q∗ (t − τ1 )w]∗j
0 0
· exp(−σ j τ1 − σk τ2 )[Q ∗ (t − τ2 )w]k } dτ2
∞ ∞
+β 2
dτ1 E [Q∗ (t − τ1 )ξ (t − τ1 )]∗j (4.157)
0 0
· exp(−σ j τ1 − σk τ2 )[Q∗ (t − τ2 )ξ (t − τ2 )]k } dτ2
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 181
where the cross-product terms do not appear since E{ξ (t)} = 0 and (t) and ξ (t) are
independent noise processes. A useful lower bound for (5.157) is given by [23].
β 2 1 ∗
N
var(w) ≥ E |(Q w)n |2 (4.158)
2 n=1 σn
where represents the time interval between successive independent samples of the input
signal vector. For a pulse radar, is approximately the same as the pulse width. For a
communications system, is approximately 1/B, where B is the signal bandwidth.
The bound in (4.158) is useful in selecting parameter values for the Howells–Applebaum
servo loops. If this bound is not small, then the noise fluctuations at the output of the adap-
tive loops are correspondingly large. For cases of practical interest [when var(w) is small
compared with w† w], the right-hand side of (4.158) is an accurate estimate of var(w).
Equation (4.158) simplifies (after considerable effort) to yield the expression
β β
N
1
var(w) ≥ − w† Rnn w (4.159)
2 2γ n=1 λn + 1/γ
γ
N
γ
Kn ≥ λn = trace(Rnn ) (4.164)
2τ0 n=1 4Bτ0
where = 1/2B (i.e., B is the bandwidth of the input signal process), so that K n is a
direct measure of algorithm misadjustment due to noise in the weight vector. Recalling the
solution to (4.129), we see that the effective time constant of the normal weight component
w k having the slowest convergence rate is
τ0 ∼ τ0
τeff = = (4.165)
1 + γ λmin γ λmin
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 182
where γ λmin ≥ 1 to avoid a steady-state bias error in the solution. On combining (4.164)
and (4.165) there results
τeff 1 N
trace(Rnn )
≥ λn = (4.166)
2K n λmin n=1 2K n λmin
Equation (4.166) shows that, when the smallest eigenvalue λmin is small compared with
trace(Rnn ), many independent samples of the input signal are required before the adaptive
array settles to a near-optimum set of weights without excessive loop noise; no set of loop
parameters yields both low loop noise and rapid convergence in this case. Berni [27] gives
an analysis of steady-state weight jitter in Howells–Applebaum control loops when there
is no statistical independence between the input signal and weight processes. Steady-state
weight jitter is closely related to the statistical dependence between the weight and signal
processes.
where s and its components si for a linear N -element array are defined by
sT = [s2 , s2 , . . . , s N ] (4.168)
si = e jψ(2i−N −1)/2 (4.169)
q
q0 Null
qJ
Jammer
azimuth
The overall array beam pattern is most easily derived by considering the output of the
orthonormal system represented in Figure 4-12b for the input signal vector s, defined in
(4.169). Since the output for the real orthonormal systems are identical, it follows that
N
N
AF(θ, t) = w i si = w i si = w T s (4.173)
i=1 i=1
where
s = Qs (4.174)
Now the ith component of s is given by
N
si = eiT s = eik sk (4.175)
k=1
but this summation defines the ith eigenvector beam [as can be seen from (4.167)], so that
si = eiT s = gi (θ ) (4.176)
Consequently, the overall array factor can be expressed as
N
AF(θ, t) = w i gi (θ ) (4.177)
i=1
which shows that the output array factor is the summation of the N eigenvector beams
weighted by the orthonormal system adaptive weights.
Since the kth component of the quiescent orthonormal weight vector is given by
w q k = e†k wq (4.178)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 184
the steady-state solution for the kth component of the orthonormal weight vector given by
(4.131) can be rewritten using (4.132), (4.95), and (4.139) to yield
1 + μmin
w k (∞) = w q k (4.179)
1 + μk
Assume as before that quiescent signal conditions up to time t = 0 consist only of
receiver noise and that the external interference sources are switched on at t = 0; then
w k (0) = w q k (4.180)
and the solution for w k expressed by (4.131) is rewritten in the more convenient form
!
−αk t μk − μmin
w k = w qk − (1 − e ) w q k (4.181)
1 + μk
It is immediately apparent that at time t = 0 (4.177) results in
N
AF(θ, 0) = w q i gi (θ ) = wqT s = wqT Qs (4.182)
i=1
The foregoing result emphasizes that the adaptive array factor consists of two parts:
1. The quiescent beam pattern AFq (θ )
2. The summation of weighted orthogonal eigenvector beams that is subtracted from
AFq (θ)
Note also from (4.184) that the weighting associated with any eigenvector beams corre-
sponding to eigenvalues equal to λ0 (the quiescent eigenvalue) is zero since the numerator
(μi − μmin ) is zero for such eigenvalues. Consequently, any eigenvector beams associated
with λ0 is disregarded, leaving only unique eigenvector beams to influence the resulting
pattern. The transient response time of (4.184) is determined by the value of αi , which in
turn is proportional to the eigenvalue. Therefore, a large eigenvalue yields a fast transient
response for its associated eigenvector beam, whereas a small eigenvalue results in a slow
transient response.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 185
−30
−35
−40
−10° 0° 10° 20° 30° 40° 50°
Azimuth angle
Jammer
positions
E{x1∗ x2 } = E{|J1 |2 }g1 (θ1 )g2 (θ1 ) + E{|J2 |2 }g1 (θ2 )g2 (θ2 ) (4.187)
This cross-correlation product can be zero if the product [g1 (θ )g2 (θ )] is positive when
θ = θ1 and negative when θ = θ2 , thereby resulting in decorrelation between the two
eigenvector beam signals. Figure 4-16 shows the overall quiescent beam pattern and the
resulting steady-state adapted pattern for this two-source example. Figure 4-17 illustrates
the transient response (in terms of increase in output noise power) of the adaptive array for
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 186
FIGURE 4-16 0
Steady-state
adapted array −5
pattern and
−25
−30
−90° −60° −30° 0° 30° 60° 90°
Azimuth angle
Jammer
positions
FIGURE 4-17
Increase in output noise power (dB)
10 Time constant a 2
0
0 2 4 6 8 10 12 14 16 18 20 22
Time (msec)
this two-interference source example, where it is seen that the response has two distinct
slopes associated with the two distinct (and widely different) eigenvalues.
The quiescent steered-beam pattern AFq (θ ) and its associated quiescent weight vector
wq are given by (4.87)–(4.90). The eight-element linear array has an element spacing
λ/2, μ = π/2 sin θ, and ak = 1. The quiescent weights and array factor are given by
w qk = e− jψ0 (2k−9)/2 (4.189)
sin [8(ψ − ψ0 )/2]
AFq (θ ) = (4.190)
sin [(ψ − ψ0 )/2]
The coefficients of the input beam steering vector b∗ are found from (4.140) and (4.88)
ck = (1 + μmin ) = 2 (4.191)
bk∗ = ck w qk = 2e− jψ0 (2k−9)/2 (4.192)
The maximum power condition for each of the orthonormal loops of Figure 4-12b is
λmax π Bc τ0
μmax = μmin = −1 (4.193)
λ0 10
where λmax represents the maximum eigenvalue. The channel bandwidth Bc and filter time
constant τ0 are the same for all element channel servo loops. Solving for τ0 from (4.193)
yields
!
10 λmax 10 R
τ0 = 1 + μmin = 1 + μmin + μmin Pr gm (θr )
2
π Bc λ0 π Bc r =1
(4.194)
The maximum power (maximum eigenvalue) is much larger than the jammer-to-receiver-
noise power ratios, because the various Pr are multiplied by the power gain of the eigen-
vector beams.
where
!
μi − μ0
Ai (t) = (1 − e−αi t ) (4.197)
1 + μi
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 188
From (4.197) it is seen that Ai (t) is zero for t = 0 and for μi = μ0 (for nonunique
eigenvalues). Therefore, for quiescent conditions at t = 0, it follows that
N
N
|y0n (0)|2 = |n 0 |2 |w q i |2 = |n 0 |2 |w qk |2 (4.198)
i=1 k=1
since the output noise power must be the same for either the real system or the equivalent
orthonormal system. Consequently, (4.196) is rewritten as
N
N
|y0n (t)|2 = |n 0 |2 |w qk |2 − [2 − Ai (t)]Ai (t)|w q i |2 (4.199)
k=1 i=1
Equation (4.199) is a particularly convenient form because the w q i associated with nonunique
eigenvalues need not be evaluated since A(t) = 0 for such eigenvalues.
The output noise power contributed by R external interference sources is given by the
sum of their output power pattern levels:
R
|y0 j (t)|2 = |n 0 |2 Pr AF 2 (θr , t) (4.200)
r =1
where Pr is the r th source power ratio, θr is its angular location, and AF(θr , t) is given
by (4.184).
The total output noise power is the sum of (4.199) and (4.200), and the increase in the
output noise power (with interference sources turned on) is this sum over the quiescent
noise (4.198).
⎧ ⎫
⎪
⎪ R
N ⎪
⎪
⎨ Pr AF (θr , t) − [2 − Ai (t)]Ai (t)|w qi | ⎪
2 2 ⎪
⎬
|y0 (t)|2 r =1 i=1
= 1 + (4.201)
|y0n (0)|2 ⎪
⎪ N ⎪
⎪
⎪
⎩ |w qk | 2 ⎪
⎭
k=1
The output noise power increase in (4.201) indicates the system transient behavior. An
increase in output noise power indicates the general magnitude of the adapted (steady-
state) weights.
The degradation in the SNR, Dsn , enables one to normalize the effect of adapted-
weight magnitude level. This degradation is the quiescent SNR divided by the adapted
SNR.
where the ratio in the second factor is just (4.201), the increase in output noise power.
Frequency
f0
matrix becomes
R
Rnn = I + Pr Mr (4.203)
r =1
where Mr now represents the covariance matrix due to the r th interference source.
Wideband interference sources are represented by dividing the jammer power spec-
trum into a series of discrete spectral lines. A uniform amplitude spectrum of uncorrelated
lines spaced apart by a constant frequency increment ε is once again assumed as illus-
trated in Figure 4-18. If Pr is the power ratio of the entire jammer power spectrum, then
the power ratio of a single spectral line (assuming a total of L r spectral lines) is
Pr
Prl = (4.204)
Lr
Furthermore, if Br (Br < element channel receiver bandwidth, Bc ) denotes the percent
bandwidth of the jamming spectrum, then the frequency offset of the lth spectral line is
!
fl Br 1 l −1
= − + (4.205)
f0 100 2 Lr − 1
The covariance matrix with R broadband interference sources is written as
R
Lr
Rnn = I + Prl Mrl (4.206)
r =1 l=1
The mnth component (mth row and nth column) of the matrix Mrl is in turn given by
beams required to place nulls at the jammer locations. If the adaptive weight adjustments
are large, there may be appreciable main beam distortion in the overall adapted pattern.
The Howells–Applebaum adaptive loop has one adaptive weight in each element
channel of the array; this configuration works interference sources with a bandwidth of
up to about 20%. Gabriel [22] gives two examples as follows: a 2% bandwidth source
in the sidelobe region for which two degrees of freedom (two pattern nulls) are required
to provide proper cancellation; and a 15% bandwidth source in the sidelobe region for
which three degrees of freedom are required. Broadband interference sources require a
transversal equalizer in each element channel (instead of a single adaptive weight) for
proper compensation, with a Howells–Applebaum adaptive loop then required for every
tap appearing in the tapped delay line.
The adapted pattern for main beam nulling exhibits severe distortions. For interference
sources located in the main beam, the increase in output noise power is an unsatisfactory
indication of array performance, because there is a net SNR degradation due to the resulting
main beam distortion in the adapted pattern. Main beam constraints for such cases can be
introduced.
N
vk =k 2
u ∗k w i xi (4.209)
i=1
On comparing (4.209) with (4.97), it is seen that u ∗k has simply replaced xk∗ , so the
resulting adaptive weight matrix differential equation now becomes
dw
τ0 + [I + γ M]w = b∗ (4.210)
dt
which is analogous to (4.109) with M replacing Rnn and γ = k 2 G where M is the
modified noise covariance with envelope limiting having elements given by
% ∗ &
x m xl
Mml = E (4.211)
|xm |
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 191
FIGURE 4-19
x1 x6 Hard limiter
x2 x5 modification of linear
x3 x4 six-element adaptive
array system. From
Gabriel, Proc. IEEE,
Σ February 1976.
w1 w2 w3 w4 w5 w6
+ + + + + +
_ _ _ _ _ _
b*1 b*2 b*3 b*4 b*5 b*6
G G G G G G
6
Output = Σ wi xi
i=1
Assuming the quadrature components of each signal xk are zero-mean Gaussian random
variables having variance σ 2 , we can then compute the elements of the covariance matrix
M directly from the elements of Rnn by using the relation [28]
'
π 1
Mml = (Rnn )ml (4.212)
8σ
It follows
√ that the elements of M differ from the elements of Rnn by a common factor
(1/σ ) (π/8). Consequently, the effective time constants that determine the rate of con-
vergence and control loop noise are changed by this same factor, thereby reducing the
dependence of array performance on the strength of the external noise field.
It is worthwhile noting that limiting does not change the relative values of the sig-
nal covariance matrix elements or the relative eigenvalue magnitudes presuming identical
channels. Thus, for widely different eigenvalues, limiting does reduce the eigenvalue
spread to provide rapid transient response and low control loop noise. Nevertheless, limit-
ing always reduces the dynamic range of signals in the control loops, thereby simplifying
the loop implementation.
desired mainlobe signals while realizing good cancellation of interference in the sidelobes.
The constraint methods discussed here follow the development that is given by Applebaum
and Chapman [29].
Techniques for applying main beam constraints to limit severe array pattern degrada-
tion include the following:
1. Time domain: The array adapts when the desired signal is not present in the main beam.
These weights are kept until the next adaptation or sampling period. This approach does
not protect against main beam distortion resulting from main beam jamming and is
also vulnerable to blinking jammers.
2. Frequency domain: When the interference sources have much wider bandwidths than
the desired signal, the adaptive processor is constrained to adapt to signals only out-
side the desired signal bandwidth. This approach somewhat degrades the cancellation
capability and distorts the array factor.
3. Angle domain: Three angle domain techniques provide main beam constraints in the
steady state: (a) pilot signals; (b) preadaptation spatial filters; and (c) control loop
spatial filters. These techniques are also helpful for constraining the array response
to short duration signals, since they slow down the transient response to main beam
signals. The angle domain techniques provide the capability of introducing main beam
constraints into the adaptive processor response.
FIGURE 4-20
Beam
Multiple sidelobe
steering
canceller (MSLC)
adaptive array f1 f2 f3 f4
configuration with
beam steering pilot x1 x2 x3 x4
signals and main Pilot
beam control. Σ ms 1 Σ ms2 Σ ms3 Σ ms4
signals
w1 w2 w3 w4
Low
∫ pass ∫ ∫ ∫
filter
Reference − − − −
pilot
+ + + +
signal
ms0
u*1 u*2 u*3 u*4
+
_ Residue, ε
Output
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 193
adaptive loop. The “pilot signals” shape the array beam and maintain the main beam gain
(avoiding SNR degradation). The pilot signals are continuous wave (CW) tones injected
into each element channel at a frequency that is easily filtered out of the signal bandwidth.
It is not necessary to use the beam steering phase shifters shown in Figure 4-20, since if
they are not present the pilot signals may be injected with the proper phase relationship
corresponding to the desired main beam direction instead of in phase with each other
as shown. The amplitudes and phases of the injected pilot signals s1 , . . . , s4 may be
represented by the vector μs, where s has unit length, and μ is a scalar amplitude factor.
The reference channel (or main beam) signal is represented by the injected pilot signal s0 .
For the adaptive control loops shown in Figure 4-20, it follows that the vector differ-
ential equation for the weight vector is written as
dw
= u∗ (t)ε(t) − w(t) (4.213)
dt
Since ε = μs0 − xT w, it follows (4.213) and the results of Section 4.3.5 that
dw
= gμrxs0 − [I + gRx x ]w (4.214)
dt
where g is a gain factor representing the correlation mixer gain and the effect of the limiter.
The steady-state solution of (4.214) is given by
x = n + μs (4.216)
where n is the noise signal vector, and μs is the injected pilot signal vector. Consequently,
Rx x = E{x∗ xT } = Rnn + μs∗ sT (4.217)
∗ ∗
rxs0 = E{x s0 } = μs s0 (4.218)
If s has equal amplitude components, then the main beam response from (4.221) is
sT wss ∼
= s0 (4.222)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 194
which is a constant, independent of Rnn (and hence independent of any received wave-
forms).
The array configuration of Figure 4-20 uses one set of pilot signals for a single main
beam constraint. Multiple constraints require multiple sets of pilot signals, with each set at
a different frequency. Pilot signals are inserted close to the input of each element channel
to compensate for any amplitude and phase errors. Strong pilot signals require channel
elements with a large dynamic range, so they must be filtered to avoid interfering with the
desired signal.
FIGURE 4-21
General structure of
preadaption spatial
filtering. f1 f2 f3 f4 Beam steering
s* A As = 0
Residue signal
e0
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 195
The composite weight vector for the entire system can therefore be written as
w = s∗ − AT y (4.227)
Consequently
sT w = sT (s∗ − AT y) = s2 − sT AT y (4.228)
Since A was selected so that As = 0, it follows that
sT w = s2 = 1 (4.229)
Denote the covariance matrix associated with u by
Ruu = E{u∗ uT } = E{A∗ x∗ xT AT } = A∗ Rx x AT (4.230)
where g is a gain factor, and the right side of (4.232) represents the cross-correlation vector
of em with each component of u. Using (4.223), (4.225), and (4.230) in (4.232), we find
that
I + gA∗ Rx x AT y = gA∗ Rx x s∗ (4.233)
Premultiply (4.233) by AT and use (4.227); it then follows that the composite weight
applied to the input signal vector satisfies the steady-state relation
I + gAT A∗ Rx x wss = s∗ (4.234)
when x does not contain a desired signal component, then Rx x may be replaced by Rnn .
Now allow g to become very large so that (4.233) yields
A∗ Rx x AT y = A∗ Rx x s∗ (4.235)
A∗ Rx x (s∗ − AT y) = A∗ Rx x wss = 0 (4.236)
Since As = 0 and the rank of the transformation matrix A is N − 1, (4.236) implies that
Rx x w is proportional to s∗ so that
as the solution that the composite weight vector approaches when g becomes very large.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 196
Preadaption spatial filtering avoids dynamic range problems, so it does require the
implementation of multiple beams. The accuracy of the beam steering phase shifters limits
the effectiveness of the constraints, but this limit is true of all three methods considered
here. Two realizations of preadaption spatial filtering represented by Figure 4-21 include
[29]: (1) the use of a Butler matrix to obtain orthogonal beams, one of which is regarded
as the “main” beam; and (2) the use of an A matrix transformation obtained by fixed
element-to-element subtraction.
FIGURE 4-22
Adaptive processor
with control loop x1 x2 xk xk + 1 xN
spatial filtering.
wk
+
b*k
_
zk
Spatial
∧
matrix P
filter
vk
u*k
wTx
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 197
Substituting w = b∗ − z into (4.245) means the steady-state weight values must satisfy
When the beam steering vector is uniformly weighted, the projection performed by
the spatial filter to remove signal components in the direction of b is
P̂ = I − b∗ bT (4.244)
Rx x (I + gRx x )−1 b∗
[Q − gb∗ bT ]−1 b∗ = (4.248)
1 − gbT Rx x (I + gRx x )−1 b∗
The denominator of (4.243) may be simplified as
where m n = δmn , δmn = the Kronecker delta, and the m are the constraint vectors.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 198
The constraint that maintains the array response at the peak of the beam is the “zero-
order” constraint. The weight vector solution obtained with a zero-order constraint differs
from the unconstrained solution (P̂ = I) by a multiplicative scale factor. Multiple con-
straints are typically used to increase the beam constraint zone by controlling the first few
derivatives of the pattern function in the direction of interest. A constraint that controls
the mth derivative is referred to as an “mth-order” constraint.
To synthesize a m constraint vector corresponding to the mth derivative of the pattern
function, note that the pattern function of a linear array can be written as
N
AF(θ ) = w k e jkθ (4.252)
k=1
N
AF m (θ ) = ( jk)m w k e jkθ (4.253)
k=1
0i = d0 (4.254)
1i = e0 + e1 i (4.255)
2i = f 0 + f 1 i + f 2 i 2 (4.256)
The constants defining the m elements are made unit length and mutually orthogonal.
Consider how to establish a beam having nonuniform weighting as well as zero-,
first-, and second-order constraints on the beam shape at the center of the main beam. First
expand wq in terms of the constraint vectors m (for m = 0, 1, and 2) and a remainder
vector r as
wq = a0 0 + a1 1 + a2 2 + ar r (4.257)
where
The subspace spanned by the constraint vectors in the N -dimension space of the adaptive
processor is preserved by the foregoing construction. The spatial matrix filter constructed
according to (4.263) results in a signal vector z containing no components in the direction
of wq or its first and second derivatives. The vector wq is now added back in (at the point
in Figure 4-21 where b∗ is inserted) to form the final weight vector w.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 199
4.5 Constraint for the Case of Known Desired Signal Power Level 199
xN(t)
Aux. N
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 200
If a desired signal is present then the output signal-to-interference plus noise ratio (SINR) as
1
SNR = (4.264)
(Pe /|s0 − w† s|2 ) − 1
where Pe represents the total output power, s is the desired signal direction vector, and s0 is
the main channel desired signal component. It will be convenient to define the parameter
|s0 − w† s|2
SN = (4.265)
Pe
so that
1
SNR = (4.266)
(1/S N ) − 1
It can be shown that
2
N (Qrx x0 )i∗ (Qs)i
s0 −
λi +a
SN = i=1
(4.267)
N |(Qrx x0 )i |2
a2 λi (λi +a)2
+ Pe0
i=1
Rxx = J v J v J + Ps vs vs (4.268)
( (
rxx0 = J0 J v J e jφ J + Ps0 Ps vs e jφs (4.269)
where
J0 = main channel jammer power
J = auxiliary channel jammer power (assumed equal in all auxiliary channels)
v J = jammer direction delay vector
Ps0 , Ps , and vs are similarly defined for the desired signal. φ J and φs represent the relative
phase between the main and auxiliary channel signals for the jamming and desired signals,
respectively.
If the desired signal and the interference signal angles of arrival are such that vs and
v J are orthogonal (which simplifies the discussion for tutorial purposes), then (4.270)
reduces to
2
σ2 + a
Ps0
σ 2 + a + NPs
SN = !
Ps0 Ps J0 J
a2 N + + Pe0
(NPs + σ 2 + a)2 (NPs + σ 2 ) (NJ + σ 2 + a)2 (NJ + σ 2 )
(4.270)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 201
where σ 2 = auxiliary channel thermal noise power. For a = 0 (which corresponds to the
conventional LMS null-steering algorithm) SN becomes
2
σ2
Ps0
NPs + σ 2
SN = ; a=0 (4.271)
Pe0
This result shows that S N decreases as the input desired signal power in the auxiliary
channels NPs increases above the thermal noise level σ 2 . When NPs σ 2 in (4.271), SN
is inversely proportional to the input desired signal power, which is the power inversion
characteristic of the minimum MSE performance criterion. When a NPs + σ 2 and
a NJ + σ , (4.270) becomes
2
Ps0
SN = (4.272)
NPs NJ
Ps + J0 + Pe0
NPs + σ 2 0 NJ + σ 2
Suppression of the main channel signal is prevented by selecting a to be sufficiently
large. However, a is too large in this example, because jammer suppression has also been
prevented, as indicated by the presence of the term NJJ 0 /(NJ + σ 2 ) in the denominator
of (4.272).
Next, assume that the main channel jammer power J0 is nominally equal to the aux-
iliary channel jammer power, and choose a = NPs . Then S N becomes
0.25Ps0
SN = (4.273)
Ps0 (NPs − σ 2 )2 (NPs − σ 2 )2
+ + Pe0
4NPs (NPs + σ 2 ) N [1 + (Ps /J )]2 (NJ + σ 2 )
For J Ps , Ps σ 2,
0.25Ps0
SN ≈
0.25Ps0 + Pe0
An approximation for the output signal-to-interference plus noise ratio in (4.269) is
1 Ps0
SNR ∼ = (4.274)
4 Pe0
Thus, the output signal-to-interference plus noise ratio is now proportional to the main
channel signal power divided by the output residue power Pe0 (recall that Pe0 is the
minimum output residue power obtained when a = 0). Equation (4.274) shows that when
J Ps and Ps σ 2 , the output signal-to-interference plus noise ratio can be significantly
improved by selecting the weight feedback gain as
a ≈ NPs (4.275)
This value of a (when J Ps and Ps σ ) then prevents suppression of the relatively
2
weak desired signal while strongly suppressing higher power level jamming signals.
FIGURE 4-24
One-dimensional
x
gradient estimation
by way of direct
measurement.
d d
w
w(k)
obtains gradient vector estimates by direct measurement and is straightforward and easy
to implement [13].
The parabolic performance surface representing the MSE function of a single variable
w is defined by
ξ [w(k)] = ξ(k) = ξmin + αw 2 (k) (4.276)
Figure 4-24 represents the parabolic performance surface as a function of a single com-
ponent of the weight vector w. The first and second derivatives of the MSE are
!
dξ(k)
= 2αw(k) (4.277)
dw w=w(k)
d 2 ξ(k)
= 2α (4.278)
dw 2
w=w(k)
The procedure for estimating the first derivative illustrated in Figure 4-24 requires
that the weight adjustment be altered to two distinct settings while the gradient estimate
is obtained. If K data samples are taken to estimate the MSE at the two weight settings
w(k) + δ and w(k) − δ, then the average MSE experienced (over both settings) is greater
than the MSE at w(k) by an amount γ . Consequently, a performance penalty is incurred
that results from the weight alteration used to obtain the derivative estimate.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 203
γ αδ 2
P= = (4.282)
ξmin ξmin
The perturbation is the average increase in the MSE normalized with respect to the mini-
mum achievable MSE.
A two-dimensional gradient is needed for the input signal correlation matrix
!
r11 r12
Rx x = (4.283)
r21 r22
Measuring the partial derivative of the previous performance surface along the coordinate
w 1 yields a perturbation
r11 δ 2
P= (4.285)
ξmin
Likewise, the perturbation for the measured partial derivative along the coordinate w 2 is
r22 δ 2
P= (4.286)
ξmin
If we allot equal time for the measurement of both partial derivatives (a total of 2K data
samples are used for both measurements), the average perturbation experienced during
the complete measurement process is given by
δ 2 r11 + r22
Pav = · (4.287)
ξmin 2
For N dimensions, define a general perturbation as the average of the perturbations
experienced for each of the individual gradient component measurements so that
δ 2 tr(Rx x )
P= · (4.288)
ξmin N
where “tr” denotes trace, which is defined as the sum of the diagonal elements of the
indicated matrix. When we convert the Rx x matrix to normal coordinates, the trace of Rx x
is the sum of its eigenvalues. Since the sum of the eigenvalues divided by N is the average
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 204
1 2
K
ξ̂ = e (k) (4.290)
K k=1
It is well known that the variance of a sample average estimate of the mean square obtained
from K independent samples is given by the difference between the mean fourth and the
square of the mean square all divided by K. Consequently the variance of ξ̂ may be
expressed as [32]
E{e4 (k)} − [E{e2 (k)}]2
var[ξ̂ ] = (4.291)
K
If the random variable e(k) is normally distributed with zero mean and variance σ 2 ,
then its mean fourth is 3σ 4 , and the square of its mean square is σ 4 . Consequently, the
variance in the estimate of ξ is given by
1 2σ 4 2ξ 2
var[ξ̂ ] = (3σ 4 − σ 4 ) = = (4.292)
K K K
From (4.292) we find that the variance of ξ̂ is proportional to the square of ξ and inversely
proportional to the number of data samples. In general, the variance can be expressed as
ξ2
var[ξ̂ ] = η (4.293)
K
where η has the value of 2 for an unbiased Gaussian density function. In the event that
the probability density function for ξ̂ is not Gaussian, then the value of η is generally less
than but close to 2. It is therefore convenient to assume that the final result expressed in
(4.292) holds for the analysis that follows.
The derivatives required by the DSD algorithm are measured in accordance with
(4.279). The measured derivative involves taking finite differences of two MSE estimates,
so the error in the measured derivative involves the sum of two independent components
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 205
[since the error samples e(k) are assumed to be independent]. The variance of each compo-
nent to the derivative error is given by (4.292). Assume that we are attempting to measure
the derivative at a point on the performance surface where the weight vector is near the
minimum point of the MSE surface and that the perturbation P is small, then the two
components of measured derivative error will have essentially the same variances. The
total variance of the measured derivative error will then be the sum of the variances of the
two components. From (4.279) and (4.292) it follows that the variance in the estimate of
the derivative is given by
!
dξ 1 2ξ 2 [w(k) + δ] 2ξ 2 [w(k) − δ]
var = 2 +
dw w=w(k) 4δ K K
∼ ξ2
= min2 (4.294)
Kδ
When an entire gradient vector is measured, then the errors in each component are
independent. It is convenient to define a gradient noise vector g(k) in terms of the true
gradient ∇(k) and the estimated gradient ∇(k):
ˆ
∇(k)
ˆ = ∇(k) + g(k) (4.295)
where g(k) is the gradient noise vector. Under the previously assumed conditions, the
covariance of the gradient noise vector can be expressed as
ξmin
2
cov [g(k)] = I (4.296)
K δ2
Transforming the gradient noise vector into normal coordinates, we have
g (k) = Qg(k) (4.297)
We see from (4.296) that the covariance matrix of g(k) is a scalar multiplying the identity
matrix, so projecting into normal coordinates through the orthonormal transformation Q
yields the same covariance for g (k):
ξ2
cov [g (k)] = E Qg(k)gT (k)Q−1 = min2 I (4.298)
Kδ
This result merely emphasizes that near the minimum point of the performance surface
the covariance of the gradient noise is essentially a constant and does not depend on w(k).
The fact that the gradient estimates are noisy means that weight adaptation based on
these gradient estimates will also be noisy, and it is consequently of interest to determine
the corresponding noise in the weight vector. Using estimated gradients, the method of
steepest descent yields the vector difference equation
which is a first-order difference equation having a stochastic driving function –s g(k).
Projecting the previous difference equation into normal coordinates by premultiplying by
Q then yields
After initial adaptive transients have died out and the steady state is reached, the
weight vector v (k) behaves like a stationary random process in response to the stochastic
driving function –s g (k). In the normal coordinate system there is no cross-coupling
between terms, and the components of g (k) are uncorrelated; thus, the components of
v (k) are also mutually uncorrelated, and the covariance matrix of g (k) is diagonal. The
covariance matrix of v (k) describes how noisy the weight vector will be in response to the
stochastic driving function, and we now proceed to find this matrix. Since cov[v (k)] =
E{v (k)vT (k)}, it is of interest to determine the quantity v (k + 1)vT (k + 1) by way of
(4.302) as follows:
Taking expected values of both sides of (4.303) and noting that v (k) and g (k) are uncor-
related since v (k) is affected only by gradient noise from previous adaptive cycles, we
obtain for the steady state
In practice, the step size in the method of steepest descent is selected so that
s I (4.306)
Without any noise in the weight vector, the method of steepest descent converges to a
steady-state solution at the minimum point of the MSE performance surface (the bottom
of the bowl). The MSE would then be ξmin . The noise present in the weight vector causes
the steady-state solution to randomly wander about the minimum point. The result of this
wandering is a steady-state MSE that is greater than ξmin and hence is said to have an
“excess” MSE. We will now consider how severe this excess MSE is for the noise that is
in the weight vector.
We have already seen in Section 4.1.3 that the MSE can be expressed as
N
2
E vT (k)v (k) = λ p E v p (k) (4.311)
p=1
2 s ξmin
2
1
E v p (k) = (4.312)
4kδ 2 λp
N s ξmin
2
E{vT (k)v (k)} = (4.313)
4K δ 2
Recalling that the misadjustment M is defined as the average excess MSE divided by
the minimum MSE there results for the DSD algorithm
N s ξmin
M= (4.314)
4K δ 2
The foregoing result is more usefully expressed in terms of time constants of the learning
process and the perturbation of the gradient estimation process as developed next.
Each measurement to determine a gradient component uses 2K samples of data. Each
adaptive weight iteration involves N gradient component measurements and therefore
requires a total of 2KN data samples. From Section 4.2.3 it may be recalled that the MSE
learning curve has a pth mode time constant given by
1 τp
τ pMSE = = (4.315)
4s λ p 2
in time units of the number of iterations. It is useful to define a new time constant T pMSE
whose basic time unit is the data sample and whose value is expressed in terms of the
number of data samples. It follows that for the DSD algorithm
T pMSE = 2KNτ pMSE (4.316)
The time constant T pMSE relates to real time units (seconds) once the sampling rate is
known.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 208
By using the perturbation formula (4.282) to substitute for ξmin in (4.314), the misad-
justment for the DSD algorithm is rewritten as
N s λav
M= (4.317)
4KP
The time constant defined by (4.316) is rewritten using (4.315) as
NK
T pMSE = (4.318)
2s λ p
from which one can conclude that
NK 1
λp = (4.319)
2s T pMSE
so that
NK 1
λav = (4.320)
2s TMSE av
Equation (4.321) shows that for the DSD algorithm, misadjustment is proportional
to the square of the number of weights and inversely proportional to the perturbation. In
addition, the misadjustment is also inversely proportional to the speed of adaptation (fast
adaptation results in high misadjustment). Since the DSD algorithm is based on steepest
descent, it suffers from the disparate eigenvalue problem discussed in Section 4.2.2.
It is appropriate here to compare the misadjustment for the DSD algorithm given by
(4.321) with the misadjustment for the LMS algorithm given by (4.83). With a specified
level of misadjustment for the LMS algorithm, the adaptive time constants increase linearly
with the number of weights rather than with the square of the number of weights as
is the case with the DSD algorithm. Furthermore, with the LMS algorithm there is no
perturbation. As a result, in typical circumstances much faster adaptation is possible with
the LMS algorithm than with the DSD algorithm.
M is defined as a normalized performance penalty that results from noise in the
weight vector. In an actual adaptive system employing the DSD algorithm, the weight
vector is not only stochastically perturbed due to the presence of noise but in addition
is deterministically perturbed so the gradient can be measured. As a consequence of
the deterministic perturbation, another performance penalty accrues as measured by the
perturbation P, which is also a normalized ratio of excess MSE. The total excess MSE is
therefore the sum of the “stochastic” and “deterministic” perturbation components. The
total misadjustment can be expressed as
Mtot = M + P (4.322)
Since P is a design parameter given by (4.282), it can be selected by choosing the deter-
ministic perturbation size δ. It is desirable to minimize the total misadjustment Mtot by
appropriately selecting P. The result of such optimization is to make the two right-hand
terms of (4.323) equal so that
1
Popt = Mtot (4.324)
2
The minimum total misadjustment then becomes
1/2
N2 1 N2 1
(Mtot )min = = (4.325)
4Popt TMSE av 2 TMSE av
Unlike the LMS algorithm, the DSD algorithm is sensitive to any correlation that
exists between successive samples of the error signal e(k), since such correlation has the
effect of making the effective statistical sample size less than the actual number of error
samples in computing the estimated gradient vector. Because of such reduced effective
sample size, the actual misadjustment experienced is greater than that predicted by (4.325),
which was derived using the assumption of statistical independence between successive
error samples.
FIGURE 4-25
Two-dimensional
diagram showing Negative gradient
A direction for point B
directional
relationships for the
Powell descent
method. D C
B Desired direction
for point A
Negative gradient
direction for point A
joining the point A with the point D in Figure 4-25 passes through the point C where the
derivative of the performance measure (w) with respect to distance along the line AD
is zero.
Given an initial estimate w0 at point A, first find the gradient direction that is normal
to the tangent of the constant performance measure contour. Proceed along the line defined
by the negative gradient direction to the point B where the derivative of (w) with respect
to distance along the line is zero. The point B may in fact be any arbitrary point on the
line that is a finite distance from A; however, by choosing it in the manner described the
convergence of the method is assured.
Having found point B, the negative gradient direction that is parallel to the original tan-
gent at (w0 ). Traveling in this new normal direction, we find a point where the derivative
of (w) with respect to distance along the line is zero (point D in Figure 4-25). The line
passing through the points A and D also passes through the point C. The desired point C
is the point where the derivative of (w) with respect to distance along the line AD is zero.
The generalization of the foregoing procedure to an N -dimensional space can be
obtained by recognizing that the directional relationships (which depend on the equal
angle property) given in Figure 4-25 are valid only in a two-dimensional plane. The
first step (moving from point A to point B) is accomplished by moving in the negative
gradient direction in the N -dimensional space. Having found point B, we can construct
(N − 1) planes between the original negative gradient direction and (N − 1) additional
mutually orthogonal vectors, thereby defining points C, D, E, . . . , until (N −1) additional
points have been defined. The last three points in the N -dimensional space obtained in
the foregoing manner may now be treated in the same fashion as points A, B, and D of
Figure 4-25 by drawing a connecting line between the last point obtained and the point
defined two steps earlier. Traveling along the connecting line one may then define a new
point C of Figure 4-25. This new point may then be considered as point D in Figure 4-25,
and a new connecting line may be drawn between the new point and the point obtained
three steps earlier.
The steps corresponding to one complete Powell descent cycle for five dimensions
are illustrated in Figure 4-26. The first step from A to B merely involves traveling in the
negative gradient direction v1 with a step size α1 chosen to satisfy the condition
d
{[w(0) + α1 v1 ]} = 0 (4.326)
dα1
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 211
A FIGURE 4-26
Illustration of Powell
1 descent method
steps required in five
J dimensions for one
B
complete cycle.
9
2
I
C
8
3
n s)
N Steps
io
ct
in ps
(orthogonal H
ire
D
g l te
ed
gradient
in 1 S
directions)
−
7
N
4
ct
ne
G
on
E
6 (c
5
so that
d
[(w + αv)] = 0 (4.331)
dα
For complex weights the MSE performance measure is given by
v† rxd + v† Rx x w
α=− (4.334)
v† Rx x v
Since rxd and Rx x are unknown, some estimate of the numerator is employed to obtain an
appropriate step size estimate. Noting that rxd + Rx x w is one-half the gradient of ξ(w),
it follows that the numerator of (4.334) are approximated by v† Av{e(k)x(k)}. Note that
the quantity v† x is regarded as the output of a processor whose weights correspond to v
and that Av{(v† x)(x† v)} is an approximation of the quantity v† Rx x v, where the average
Av{ } is taken over K data samples. The simultaneous generation of the estimates ∇ ˆ w and
† †
Av{v xx v} requires parallel processors: one processor with weight values equal to w(k)
and another processor with weight values equal to v(k). Having described the procedure
for determining the appropriate step size along a direction v, we may now consider the
steps required to implement an entire Powell descent cycle.
The steps required to generate one complete Powell descent cycle are as follows.
Step 1 Starting with the initial weight setting w(0), estimate the negative gradient
direction v(0) using K data samples then travel in this direction with the appropriate step
size to obtain w(1). The step size determination requires an additional K data samples to
obtain by way of (4.334).
Steps 2 →N Estimate the negative gradient direction at w(k) using K data samples.
If the gradient estimates and the preceding step size were error free, the current gradient
is automatically orthogonal to the previous gradient directions. Since the gradient esti-
mate is not error free, determine the new direction of travel v(k) by requiring it to be
orthogonal to all previous directions v(0), v(1), . . . , v(k − 1) by employing the Schmidt
orthogonalization process so that
k−1
[v† (i)∇(k)]
ˆ
v(k) = ∇(k)
ˆ − · v(i) (4.335)
i=0
[v† (i)v(i)]
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 213
Travel in the direction −v(k) using the appropriate step size (which requires an additional
K data samples to obtain) to arrive at w(k + 1).
Steps N + 1 → 2N − 1 Determine the new direction of travel at w(k) by forming
Travel in the direction −v(k) from w[2(N − 1) − k] using the appropriate step size to
arrive at w(k + 1). These steps require only K data samples since now the direction of
travel does not require that a gradient estimate be obtained.
Delay settings
Σ Σ Σ
x2 xN +2 x(J − 1)N + 2
2d t t t
2
Σ Σ Σ Σ Array output
xN x2N xNJ
t t t
N
Nd
Σ Σ Σ
to the line of sensors (as the previous discussion has assumed), then the time delays in
the spatial correction filter are adjusted so the signal components of each channel at the
output of the preprocessor are in phase.
The adaptive signal processor of Figure 4-27 has N sensors and J taps per sensor
for a total of NJ adjustable weights. Using J constraints to determine the look direction
frequency response leaves NJ − J degrees of freedom to minimize the total array output
power. Since the J constraints fix the look direction frequency response, minimizing the
total output power is equivalent to minimizing the nonlook direction noise power (provided
the signal voltages at the taps are uncorrelated with the corresponding noise voltages at
these taps). If signal-correlated noise in the array is present, then part or all of the signal
component of the array output may be cancelled. Although signal-correlated noise may
not occur frequently, sources of such noise include multiple signal-propagation paths, and
coherent radar or sonar “clutter.”
It is desirable for proper noise cancellation that the noise voltages appearing at the
adaptive processor taps be correlated among themselves (although uncorrelated with the
signal voltages). Such noise sources may be generated by lightning, “jammers,” noise
from nearby vehicles, spatially localized incoherent clutter, and self-noise from the struc-
ture carrying the array. Noise voltages that are uncorrelated between taps (e.g., amplifier
thermal noise) are partially rejected by the adaptive array either as the result of incoherent
noise voltage addition at the array output or by reducing the weighting applied to any taps
that may have a disproportionately large uncorrelated noise power.
sample is defined by
xT (k) = [x1 (k), x2 (k), . . . , xNJ (k)] (4.337)
At any tap the voltages that appear may be regarded as the sums of voltages due to look
direction signals s and nonlook direction noises n, so that
We assume that the signals and noises are zero-mean random processes with unknown
second-order statistics. The covariance matrices of x, s, and n are given by
Since the vector of look direction signals is assumed uncorrelated with the vector of
nonlook direction noises
Assume that the noise environment is such that Rx x and Rnn are positive definite and
symmetric.
The adaptive array output (which forms the signal estimate) at the kth sample is
given by
Suppose that the weights in the jth vertical column of taps sums to a selected number
f j . This constraint may be expressed by the relation
Tj w = f j , j = 1, 2, . . . , J (4.348)
Now consider the requirement of constraining the entire weight vector to satisfy all J
equations given by (4.348). Define a J × NJ constraint matrix C having j as elements.
C = [1 · · · j · · · J ] (4.350)
Furthermore, define f as the J -dimensional vector of summed weight values for each of
the j vertical columns that yield the desired frequency response characteristic in the look
direction as
⎡ ⎤
f1
⎢ f2 ⎥
⎢ ⎥
f=⎢ . ⎥ (4.351)
⎣ .. ⎦
fJ
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 217
It immediately follows by inspection that the full set of constraints (4.348) can be written
in matrix form as
CT w = f (4.352)
Now that the look direction frequency response is fixed by the constraint equation
(4.352), minimizing the nonlook direction noise power is equivalent to minimizing the
total output power given by (4.347). The constrained optimization problem reduces to
Minimize wT Rx x w (4.353)
w
subject to C w = f
T
(4.354)
Lagrange multipliers are used to find wopt that satisfy (4.353) and (4.354) [41]. Ad-
joining the constraint equation (4.354) to the cost function (4.353) by a J -dimensional
vector λ, whose elements are undetermined Lagrange multipliers (and including a factor
of 12 to simplify the arithmetic), then yields
1 T
Minimize (w) = w Rxx w + λT [CT w − f] (4.355)
w 2
The gradient of (4.355) with respect to w is given by
∇w (w) = Rx x w + Cλ (4.356)
A necessary condition for (4.355) to be minimized is that the gradient be equal to zero so
that
Rx x w + Cλ = 0 (4.357)
wopt = −R−1
x x Cλ (4.358)
where the vector λ remains to be determined. The vector of Lagrange multipliers may now
be evaluated from the constraint equation
CT wopt = f = CT − R−1 x x Cλ (4.359)
If we substitute wopt into (4.346), it follows that the constrained least squares estimate of
the look direction signal provided by the array is
If the vector of summed weight values f is selected so the frequency response char-
acteristic in the look direction is all-pass and linear phase (distortionless), then the output
of the constrained LMS signal processor is the maximum likelihood (ML) estimate of
a stationary process in Gaussian noise (provided the angle of arrival is known) [42]. A
variety of other optimal processors can also be obtained by a suitable choice of the vector
f [43]. It is also worth noting that the solution (4.361) is sensitive to deviations of the
actual signal direction from that specified by C and to various random errors in the array
parameters [44].
where the quantity C[CT C]−1 represents the pseudo-inverse of the singular matrix CT [45].
For a gradient type algorithm, after the kth iteration the next weight vector is given by
where s is the step size constant, and denotes the performance measure. Requiring
w(k + 1) to satisfy (4.352) then yields
In the actual system the input correlation matrix is not known, and it is necessary to adopt
some estimate of this matrix to insert in place of Rx x in the iterative weight adjustment
equation. An approximation for Rx x at the kth iteration is merely the outer product of the
tap voltage vector with itself: x(k)xT (k). Substituting this estimate of Rx x into (4.370)
and recognizing that y(k) = xT (k)w(k) then yields the constrained LMS algorithm
w(0) =
(4.371)
w(k + 1) = P[w(k) − s y(k)x(k)] +
If it is merely desired to ensure that the complex response of the adaptive array system
to a normalized signal input from the look direction is unity, then the spatial correction
filter is dispensed with and the compensation for phase misalignment incorporated directly
into the variable weight selection as suggested by Takao et al. [46]. Denote the complex
response (amplitude and phase) of the array system by Y (θ ), where θ is the angle measured
from the normal direction to the array face. The appropriate conditions to impose on the
adaptive weights are found by requiring that e{Y (θ )} = 1 and Im{Y (θ )} = 0 when
θ = θc , the look direction.
and recognize that y(k) = xT (k)w(k), then taking the expected value of both sides of
(4.372) yields
Substitute (4.373) into (4.374) and use = (I − P)wopt and PRx x wopt = 0 [which may be
verified by direct substitution of (4.361) and (4.369)], then the difference vector satisfies
Note from (4.369) that P is idempotent (i.e., P2 = P), then premultiplying (4.375) by
P reveals that Pv(k + 1) = v(k + 1) for all k, so (4.375) can be rewritten as
From (4.376) it follows that the matrix PRx x P determines both the rate of convergence of
the mean weight vector to the optimum solution and the steady-state variance of the weight
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 220
vector about the optimum. The matrix PRx x P has J zero eigenvalues (corresponding to
the column vectors of the constraint matrix C) and NJ − J nonzero eigenvalues σi , i =
1, 2, . . . , NJ − J [48]. The values of the NJ − J nonzero eigenvalues are bounded by the
relation
where λmin and λmax denote the smallest and largest eigenvalues of Rx x , respectively, and
σmin and σmax denote the smallest and largest nonzero eigenvalues of PRx x P, respectively.
The initial difference vector v(0) = − wopt can be expressed as a linear combination
of the eigenvectors of PRx x P corresponding to the nonzero eigenvalues [47]. Consequently,
if v(0) is equals an eigenvector ei of PRx x P corresponding to the nonzero eigenvalue
σi , then
From (4.378) it follows that along any eigenvector of PRx x P the mean weight vector
converges to the optimum weight vector geometrically with the geometric ratio (1−s σi ).
Consequently, the time required for the difference vector length to decay to 1/e of its initial
value is given by the time constant
t
τi =
ln(1 − s σi )
∼ t
= if s σi 1 (4.379)
s σi
where t denotes the time interval corresponding to one iteration.
If the step size constant s is selected so that
1
0 < s < (4.380)
σmax
then the length (given by the norm) of any difference vector is bounded by
It immediately follows that if the initial difference vector length is finite, then the mean
weight vector converges to the optimum so that
where the convergence occurs with the time constants given by (4.379).
The LMS algorithm is designed to cope with nonstationary noise environments by
continually adapting the weights in the signal processor. In stationary environments, how-
ever, this adaptation results in the weight vector exhibiting an undesirable variance about
the optimum solution thereby producing an additional (above the optimum) component
of noise to appear at the adaptive array output.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 221
The additional noise caused by adaptively adjusting the weights can be compared with
(4.383) to determine the penalty incurred by the adaptive algorithm. A direct measure of
this penalty is the “misadjustment” M defined by (4.55). For a step size constant satisfying
1
0 < s < (4.384)
σmax + 1
2
tr(PRx x P)
The steady-state misadjustment has been shown to be bounded by [48]
s tr(PRx x P) s tr(PRx x P)
· ≤M≤ ·
2 1 − (s /2)[tr(PRx x P) + 2σmin ] 2 1 − (s /2)[tr(PRx x P) + 2σmax ]
(4.385)
If s is chosen to satisfy
2
0 < s < (4.386)
3tr(Rx x )
then it will automatically also satisfy (4.384). It is also worth noting that the upper bound in
(4.383) can be easily calculated directly from observations since tr(Rx x ) = E{xT (k)x(k)},
the sum of the powers of the tap voltages.
FIGURE 4-28
Representation of
the constraint plane,
the constraint
subspace plane, and
initial weight vector
in the w-space. Initial weight vector
= C(CTC)−1 f
Constraint plane
Λ = {w : CTw = f}
By setting the constraint weight vector f equal to zero, the homogeneous form of the
constraint equation
CT w = 0 (4.388)
defines a second plane [that is also (NJ−J )-dimensional] that passes through the coordinate
space origin. This constraint subspace is depicted in Figure 4-28.
The constrained LMS algorithm (4.371) premultiplies a certain vector in the W-space
by the matrix P, a projection operator. Premultiplication of any weight vector bythe matrix
P results in the elimination of any vector components perpendicular to the plane , thereby
projecting the original weight vector onto the constraint subspace plane as illustrated in
Figure 4-29.
The only factor in (4.371) remaining to be discussed is the vector y(k)x(k), which
is an estimate of the unconstrained gradient of the performance measure. Recall from
(4.355) that the unconstrained performance measure is 12 wT Rx x w and from (4.356) that
the unconstrained gradient is given by Rx x w. Since the covariance matrix Rx x is unknown
a priori, the estimate provided by y(k)x(k) is used in the algorithm.
The constrained optimization problem posed by (4.353) and (4.354) is illustrated
diagramatically in the w-space as shown in Figure 4-30. The algorithm must succeed
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 223
FIGURE 4-30
Diagrammatic
representation of
wopt the constrained
optimization
w(k) problem showing
Contours of constant = C(CTC)−1 f contours of constant
output power output power, the
wTRxxw constraint plane ,
the initial weight
vector , and the
optimum
constrained weight
vector wopt that
minimizes the output
power.
Λ = {w : CTw = f}
in moving from the initial weight vector to the optimum weight vector wopt along the
constraint plane . The operation of the constrained LMS algorithm (4.371) in solving
the previously given constrained optimization problem is considered.
In Figure 4-31 the current value of the weight vector, w(k), is to be modified by
taking the unconstrained negative gradient estimate −y(k)x(k), scaling it by s , and
adding the result to w(k). In general, the resulting vector lies somewhere off the constraint
plane. Premultiplying the vector [w(k) − s y(k)x(k)] by the matrix P, the projection
onto the constraint subspace plane is obtained. Finally, adding to constraint subspace
plane projection produces a new weight vector that lies on the constraint plane. This new
FIGURE 4-31
Operation of the
constrained LMS
w(k) − Δsy(k) x(k)
w(k+1) algorithm:
w(k + 1) = P[w(k) −
s y(k)x(k)]+ .
w(k)
P[w(k) − Δsy(k) x(k)]
S Λ
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 224
weight vector w(k + 1) satisfies the constraint to within the numerical accuracy of the
computations. This error-correcting feature of the constrained LMS algorithm prevents
any computational errors from accumulating.
The convergence properties of the constrained LMS algorithm are closely related to
those for the unconstrained LMS algorithm and have been previously discussed. Likewise,
the same procedures that increased convergence speed for the LMS algorithm also work
for the constrained LMS algorithm.
1 s 3
Ji †
Rx x = (uu† ) + vi vi + I (4.389)
n n i=1
n
where n denotes the thermal noise power (taken to be unity), s/n denotes the signal-to-
thermal noise ratio, and Ji /n denotes the jammer-to-thermal noise ratios for each of the
three jammers (i = 1, 2, 3). The elements of the signal steering vector u and the jammer
steering vectors vi are easily defined from the array geometry and the signal arrival angles.
The desired signal is a biphase modulated signal having a phase angle of either 0◦ or 180◦
with equal probability at each sample.
Two signal conditions were simulated corresponding to two values of eigenvalue
spread in the received signal covariance matrix. The first condition represents a respectable
FIGURE 4-32 y
Four-element
Y-array geometry
with signal and (0.394l, 0.682l)
jammer locations for
selected example.
J2 /n
J3 /n J1 /n
60° 60°
(−0.787l, 0) x
s/n = 10
l
87
0.7
(0.394l, −0.682l)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 225
eigenvalue spread of λmax /λmin = 2440, whereas the second condition represents a more
extreme eigenvalue spread of λmax /λmin = 16, 700. Choosing the jammer-to-thermal noise
ratios to be J1 /n = 500, J2 /n = 40, and J3 /n = 200 together with s/n = 10 yields the
corresponding eigenvalues λ1 = 2.44 × 103 , λ2 = 4.94 × 102 , λ3 = 25.62, and λ4 = 1.0
for which the optimum output SNR is SNRopt = 15.0 (11.7 dB). Likewise, choosing
the jammer-to-thermal noise ratios to be J1 /n = 4000, J2 /n = 40, and J3 /n = 400
along with s/n = 10 yields the eigenvalues λ1 = 1.67 × 104 , λ2 = 103 , λ3 = 29, and
λ4 = 1.0 for which the optimum output SNR is also SNRopt = 15.0. In all cases the
initial weight vector setting was taken to be wT (0) = [0.1, 0, 0, 0]. Figures 4-33 and 4-34
show the convergence results for the LMS and PAG algorithms, respectively, plotted as
output SNR in decibels versus number of iterations for an eigenvalue spread of 2,440
(here output SNR means output signal-to-jammer plus thermal noise ratio). The expected
value of the gradient and v† Rx x v required by the PAG algorithm was taken over K = 9
data samples, and one iteration of the PAG algorithm occurred every nine data samples,
0 α L = 0.1.
−5.00
−10.00
−15.00
−20.00
−25.00
1 2 3 4 5 6 7 8 1E1 2 3 4 5 6 7 8 1E2 2 3 4 5 6 7 8 1E3 2 3 4 5
Number of iterations
0 K = 9.
−5.00
−10.00
−15.00
−20.00
−25.00
1 2 3 4 5 6 7 8 1E1 2 3 4 5 6 7 8 1E2 2 3 4 5 6 7 8 1E3 2 3 4 5
Number of iterations
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 226
even though a weight update does not occur on some iterations. The loop gain of the
LMS loop was selected in accordance with (4.49), which requires that s tr(Rx x ) < 1 for
stability. Letting s tr(Rx x ) = α L and choosing α L = 0.1 therefore ensures stability while
giving reasonably fast convergence with an acceptable degree of misadjustment error. As
a consequence of the manner in which an iteration was defined for the PAG algorithm,
the time scale for Figure 4-34 is nine times greater than the time scale for Figure 4-33.
In Figure 4-34 the PAG algorithm is within 3 dB of the optimum after approximately
80 iterations (720 data samples), whereas in Figure 4-33 the LMS algorithm requires
approximately 1500 data samples to reach the same point. Furthermore, it may be seen
that the steady-state misadjustment for the two algorithms in these examples is very
comparable so the PAG algorithm converges twice as fast as the LMS algorithm for a
given level of misadjustment in this example.
Figures 4-35 and 4-36 show the convergence of the LMS and PAG algorithms for the
same algorithm parameters as in Figures 4-33 and 4-34 but with the eigenvalue spread =
16,700. In Figure 4-36 the PAG algorithm is within 3 dB of the optimum after approx-
and α L = 0.1. 0
−5.00
−10.00
−15.00
−20.00
−25.00
1 2 3 4 5 6 7 8 1E1 2 3 4 5 6 7 8 1E2 2 3 4 5 6 7 8 1E3 2 3 4 5
Number of iterations
and K = 9. 0
−5.00
−10.00
−15.00
−20.00
−25.00
1 2 3 4 5 6 7 8 1E1 2 3 4 5 6 7 8 1E2 2 3 4 5 6 7 8 1E3 2 3 4 5
Number of iterations
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 227
imately 200 iterations (1,800 data samples), whereas the LMS algorithm in Figure 4-35
does not reach the same point even after 4,500 data samples. The degree of convergence
speed improvement that is attainable therefore increases as the degree of eigenvalue spread
increases.
A word of caution is needed concerning the expected convergence when using the PAG
algorithm. The simulation results given here were compiled for an array having only four
elements; as the number of array elements increases, the number of consecutive steps in or-
thogonal gradient directions also increases, thereby yielding significant direction errors in
the later steps (since estimation errors accumulate over the consecutive step directions). Ac-
cordingly, for a given level of misadjustment the learning curve time constant does not in-
crease linearly with N (as with LMS adaptation), but rather increases more rapidly. In fact,
when N > 10, the PAG algorithm actually converges more slowly than the LMS algorithm.
P(κ) − P(κ − 1)
δn (κ + 1) = δn (κ) + μ (4.390)
(κ)
where
P(κ) = array output power at time step κ
δn (κ) = phase shift at element n
(κ) = small phase increment
2
μ= )
N
[P(κ) − P(−1)]2
n=1
This algorithm worked well for phase-only simultaneous nulling of the sum and difference
patterns of an 80-element linear array of H plane sectoral horns [50]. A diagram of the
array appears in Figure 2-25 of Chapter 2. The sum channel has a 30 dB low sidelobe
Taylor taper and the difference channel has a 30 dB low sidelobe Bayliss taper. These
channels share eight-bit beam steering phase shifters. Experiments used a CW signal
incident on a sidelobe but no signal incident on the main beam. The cost function takes
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 228
FIGURE 4-37 0
Adapted sum
pattern for −10
simultaneous
−40
−50
Adapted
−60
0 10 20 23 30 40
q (degrees)
FIGURE 4-38 0
Adapted difference
pattern for −10
simultaneous
Array Pattern (dB)
phase-only nulling in
−20
the sum and
difference channels.
−30 Quiescent
−40
−50
Adapted
−60
0 10 20 23 30 40
q (degrees)
into consideration both the sum and difference channel output powers; otherwise, a null
will not appear in both patterns. Minimizing the output power of both channels when
an interfering signal appears at θ = 23◦ results in the patterns shown in Figure 4-37 and
Figure 4-38. The desired nulls are place with relatively small deviations from the quiescent
patterns.
4.12 PROBLEMS
1. Misadjustment-Speed of Adaptation Trade-off for the LMS and DSD Algorithms [13] For the
LMS algorithm the total misadjustment in the steady state is given by (4.83), whereas the total
(minimum) misadjustment for the DSD algorithm is given by (4.317).
(a) Assuming all eigenvalues are equal so that (T pMSE )av = TMSE and that M = 10% for the
LMS algorithm, plot TMSE versus N for N = 2, 4, 8, . . . , 512.
(b) Assuming all eigenvalues are equal so that (T pMSE )av = TMSE and that (Mtot )min = 10% for
the DSD algorithm, plot TMSE versus N for N = 2, 4, 8, . . . , 512 and compare this result
with the previous diagram obtained in part (a).
2. Reference Signal Generation for LMS Adaptation Using Polarization as a Desired Signal
Discriminant [53] LMS adaptation requires a reference signal to be generated having properties
sufficiently correlated either to the desired signal or the undesired signal to permit the adaptive
system to preserve the desired signal in its output. Usually, the desired signal waveform properties
(e.g., frequency, duration, type of modulation, signal format) are used to generate the reference
signal, but if the signal and the interference can be distinguished by polarization, then polarization
may be employed as a useful discriminant for reference signal generation.
Let s denote a linearly polarized desired signal having the known polarization angle θ , and
let n denote a linearly polarized interference signal having the polarization angle α (where it is
known only that α = θ). Assume that the desired signal and interference impinge on two linearly
polarized antennas (A and B) as shown in Figure 4-39 where the antennas differ in orientation
by the angle β. The two signals va and vb may then be expressed as
va = s cos θ + n cos α
vb = s cos(β − θ ) + n cos(β − α)
(a) Show that by introducing the weight w 1 as illustrated in Figure 4-39, then the signal vb =
vb − w 1 va can be made to be signal free (have zero desired signal content) by setting
cos(β − θ )
w1 =
cos θ
so that
sin β
vb = n sin(α − θ ) = n f (α, β, θ )
cos θ
(b) From the results of part (a), show that
FIGURE 4-39 A va
Adaptive array
configuration for Linealy +
w1 Fixed v0
interference rejection polarized weight Output
on the basis of antennas _ __
vb + v′b
polarization using w2
B
LMS adaptation.
Adaptive
control
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 231
Since the output signal v0 contains both desired signal and interference components, corre-
lating it with the signal-free voltage vb yields a measure of the interference remaining in the
output signal, and the adaptive weight w 2 can then be adjusted to reduce the interference
content in the output.
(c) The error in the output signal v0 is the interference signal component that is still present
after w 2 bb is subtracted from va . Assume that the interference and the desired signal are
uncorrelated; then
E v02 = E{s 2 } cos2 θ + E{n 2 }[cos α − w 2 f (α, β, θ )]2
If the rate of change of w2 isproportional to ∂ E v02 /∂w 2 , show that the final value of the
weight occurs when ∂ E v02 /∂w 2 = 0 so that
cos α
w2 =
f (α, β, θ )
(d) With w 2 set to the final value determined in part (c), show that the steady-state system output
is given by
v0 = s cos θ
thereby showing that the system output is free of interference under steady-state conditions.
The previous result assumes that (1) knowledge of θ and the setting of w 1 are error free;
(2) the circuitry is noiseless; And (3) the number of input signals equals the number of
antennas available. These ideal conditions are not met in practice, and [53] analyzes the
system behavior under nonideal operating conditions.
3. Relative Sensitivity of the Constrained Look-Direction Response Processor to Perturbed
Wavefronts [44] The solution to the problem of minimizing the expected output power of
†
an array η = E{w† xx† w} subject to x0 w = f (or equivalently, η = f 2 ) is given by (4.366).
†
Since the look direction response is constrained by x0 w = f where x0 denotes a plane wave
signal arriving from the angle θ0 , the rationale behind this constraint is to regard the processor
as a filter that will pass plane waves from the angle θ0 but attenuate plane waves from all other
directions.
Let a perturbed plane wave be represented by x, having components
where αk represents amplitude deviations, and ξk represents phase deviations from the nominal
plane wave signal x0 . Assume that αk , ξk are all uncorrelated zero-mean Gaussian random
variables with variances σα2 , σξ2 at each sensor of the array.
(a) Using η = w† E{xx† }w and the fact that
E{xi x ∗j } = xi0 x ∗j0 exp −σξ2 for i = j
and
E xi x ∗j = |xi0 |2 1 + σα2 for i = j
show that
†
η = exp −σξ2 w† x0 x0 w + 1 − exp −σξ2 + σα2 w† w
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 232
or η ∼
= f 2 + σξ2 + σα2 w† w for small values of σξ2 , σα2 assuming that |xi0 |2 = 1 (which is
the case for a planar wave).
(b) The result in part (a) can be rewritten as
η = f2 1+ σξ2 + σα2
where
w† w
=
f2
Consequently, the ratio can be regarded as the relative sensitivity of the processor to the
perturbations whose variances are σξ2 , σα2 . Using the weights given by (4.169), show that
†
x0 R−2
x x x0
= † 2
x0 Rx−1
x x0
The previous relative sensitivity can become large if the eigenvalues of Rx x have a large
spread, but if the eigenvalues of Rx x have a small spread then cannot become large.
4. MSLC Relationships [29] Show that (4.219) results from (4.215) by the following:
(a) Substitute (4.217) and (4.218) into (4.219).
(b) Let K = I + gRnn .
(c) Apply the matrix inversion lemma [(D.10) of Appendix D] to the resulting expression.
5. MSLC Relationships [29] Show that (4.234) follows from the steady-state relationship given
by (4.232).
6. Control Loop Spatial Filter Relationships [29] Apply the matrix inversion identity
Q−1 e
[Q + efT ]−1 e =
1 + fT Q−1 e
where Q is a nonsingular N × N matrix and e and f are N × 1 vectors to (4.247), and show that
(4.248) results.
7. Control Loop Spatial Filter Relationships [29] By substituting the relationships expressed by
(4.248) and (4.249) into (4.247), show that the steady-state weight vector relationship given by
(4.250) results.
8. Control Loop Spatial Filter Relationships [29] To show that (4.267) can be developed from
(4.265), define the ratio
w† s s† w
SN =
w† Rx x w
where
* −w + !
s
w = , s =
1 s0
and
!
Rx x rx x0
Rx x =
r†x x0 P0
(b) Show that Pe = w† Rx x w = w† Rx x w − w† rx x0 − r†x x0 w + P0 .
(c) Since w = [Rx x + aI]−1 rx x0 from (4.263) show that
w = wopt + w
(d) Substitute w = wopt + w into Pe from part (b) and show that
Pe = Pe0 + w† Rx x w
where
† †
Pe0 = P0 − wopt rx x0 − r†x x0 wopt + wopt Rx x wopt
= P0 − r†x x0 R−1
x x rx x0
†
= P0 − wopt Rx x wopt
because
= w† rx x0
N
|(Qrx x0 )i |2
†
w Rx x w = a 2
λi (λi + a)2
i=1
by using w = −aR−1
x x w.
QR−1
xx Q
−1
= and QQ−1 = I
9. Performance Degradation Due to Errors in the Assumed Direction of Signal Incidence [54]
The received signal vector can be represented by
m
x(t) = s(t) + gi (t) + n(t)
i=2
where
and
and τik represents the delay of the ith directional signal at the kth sensor relative to the geometric
center of the array; ωc is the carrier signal frequency.
The optimum weight vector should satisfy
wopt = R−1
x x rxd
where Rx x is the received signal covariance matrix, and rxd is the cross-correlation vector
between the desired signal s and the received signal vector x. Direction of arrival information is
contained in rxd , and if the direction of incidence is assumed known, then rxd can be specified
and only R−1
x x need be determined to find wopt . If the assumed direction of incidence is in error,
however, then w = R−1 x x r̃xd where r̃xd represents the cross-correlation vector computed using
the errored signal steering vector ṽ1 .
(a) For the foregoing signal model, the optimum weight vector can be written as wopt =
†
[Sv1 v1 + Rnn ]−1 · (Sv1 ), where Rnn denotes the noise covariance matrix, and S denotes
the desired signal power per sensor. If v1 is in error, then r̃xd = (S ṽ1 ). Show that the
resulting weight vector computed using r̃xd is given by
S † †
w= †
1 + Sv1 R−1 −1 −1 −1
nn v1 Rnn ṽ1 − Sv1 Rnn ṽ1 Rnn v1
1+ Sv1 R−1
nn v1
(b) Using the result obtained in part (a), show that the output signal-to-noise power ratio (when
only the desired signal and thermal noise are present) from the array is given by
†
S w† E{ss† }w Sw† v1 v1 w
= =
N out
w† Rnn w w† Rnn w
†
S|v1 R−1
nn ṽ1 |
2
= † † † † † † †
nn ṽ1 − 2S|v1 Rnn ṽ1 | + v1 Rnn v1 [S {(v1 Rnn v1 ) × (ṽ1 Rnn ṽ1 ) − |v1 Rnn ṽ1 | } + 2S ṽ1 Rnn ṽ1 ]
ṽ1 R−1 −1 2 −1 2 −1 ∗ −1 −1 2 −1
(c) Use the fact that Rnn = σ 2 I and the result of part (b) to show that
† 2
N v1 ṽ1
S
S σ2 N2
= † 2 † 2
N out NS
2 v1 ṽ1 NS v1 ṽ1
1+ 2 1− +2 1−
σ N 2 σ2 N2
where d represents the separation between sensors, and θ̃ represents the angular uncertainty
from boresight.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 235
4.13 REFERENCES
[1] B. Widrow, “Adaptive Filters,” in Aspects of Network and System Theory, edited by R. E.
Kalman and N. DeClaris, New York, Holt, Rinehart and Winston, 1970, pp. 563–587.
[2] F. R. Ragazzini and G. F. Franklin, Sampled-Data Control Systems, New York, McGraw-Hill,
1958.
[3] E. I. Jury, Sampled-Data Control Systems, New York, Wiley, 1958.
[4] J. T. Tou, Digital and Sampled-Data Control Systems, New York, McGraw-Hill, 1959.
[5] B. C. Kuo, Discrete-Data Control Systems, Englewood Cliffs, NJ, Prentice-Hall, 1970.
[6] B. Widrow, “Adaptive Filters I: Fundamentals,” Stanford Electronics Laboratories, Stanford,
CA, Rept. SEL-66-126 (Tech. Rept. 6764-6), December 1966.
[7] J. S. Koford and G. F. Groner, “The Use of an Adaptive Threshold Element to Design a Linear
Optimal Pattern Classifier,” IEEE Trans. Inf. Theory, Vol. IT-12, January 1966, pp. 42– 50.
[8] K. Steinbuch and B. Widrow, “A Critical Comparison of Two Kinds of Adaptive Classifi-
cation Networks,” IEEE Trans. Electron. Comput. (Short Notes), Vol. EC-14, October 1965,
pp. 737–740.
[9] F. W. Smith, “Design of Quasi-Optimal Minimum-Time Controllers,” IEEE Trans. Autom.
Control, Vol. AC-11, January 1966, pp. 71–77.
[10] B. Widrow, P. E. Mantey, L. J. Griffiths, and B. B. Goode, “Adaptive Antenna Systems,” Proc.
IEEE, Vol. 55, No. 12, December 1967, pp. 2143–2159.
[11] L. J. Griffiths, “Signal Extraction Using Real-Time Adaptation of a Linear Multichannel
Filter,” Ph.D. Disseration, Stanford University, December 1967.
[12] K. D. Senne and L. L. Horowitz, “New Results on Convergence of the Discrete-Time LMS
Algorithm Applied to Narrowband Adaptive Arrays,” Proceedings of the 17th IEEE Confer-
ence on Decision and Control, San Diego, CA, January 10–12, 1979, pp. 1166–1167.
[13] B. Widrow and J. M. McCool, “A Comparison of Adaptive Algorithms Based on the Methods
of Steepest Descent and Random Search,” IEEE Trans. Antennas Propag., Vol. AP-24, No. 5,
September 1976, pp. 615–638.
[14] R. T. Compton, Jr., “An Adaptive Array in a Spread-Spectrum Communication System,” Proc.
IEEE, Vol. 66, No. 3, March 1978, pp. 289–298.
[15] D. M. DiCarlo and R. T. Compton, Jr., “Reference Loop Phase Shift in Adaptive Arrays,”
IEEE Trans. Aerosp. Electron. Syst., Vol. AES-14, No. 4, July 1978, pp. 599–607.
[16] P. W. Howells, “Intermediate Frequency Side-Lobe Canceller,” U.S. Patent 3202990, August
24, 1965.
[17] S. P. Applebaum, “Adaptive Arrays,” Syracuse University Research Corp., Report SPL TR
66-1, August 1966.
[18] L. E. Brennan and I. S. Reed, “Theory of Adaptive Radar,” IEEE Trans. Aerosp. Electron.
Syst., Vol. AES-9, No. 2, March 1973, pp. 237–252.
[19] L. E. Brennan and I. S. Reed, “Adaptive Space-Time Processing in Airborne Radars,” Tech-
nology Service Corporation, Report TSC-PD-061-2, Santa Monica, CA, February 24, 1971.
[20] A. L. McGuffin, “Adaptive Antenna Compatibility with Radar Signal Processing,” in
Proceedings of the Array Antenna Conference, February 1972, Naval Electronics Labora-
tory, San Diego, CA.
[21] L. E. Brennan, E. L. Pugh, and I. S. Reed, “Control Loop Noise in Adaptive Array Antennas,”
IEEE Trans. Aerosp. Electron. Syst., Vol. AES-7, No. 2, March 1971, pp. 254–262.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 236
[22] W. F. Gabriel, “Adaptive Arrays—An Introduction,” Proc. IEEE, Vol. 64, No. 2, February
1976, pp. 239–272.
[23] L. E. Brennan, E. L. Pugh, and I. S. Reed, “Control Loop Noise in Adaptive Array Antennas,”
IEEE Trans. Aerosp. Electron. Syst., Vol. AES-7, No. 2, March 1971, pp. 254–262.
[24] L. J. Griffiths, “A Simple Adaptive Algorithm for Real-Time Processing in Antenna Arrays,”
Proc. of IEEE, Vol. 57, No. 10, October 1969, pp. 1696–1704.
[25] R. T. Compton, Jr., Adaptive Antennnas Concepts and Performance, Englewood Cliffs, NJ,
Prentice-Hall, Inc., 1988.
[26] M. W. Ganz, “Rapid Convergence by Cascading Applebaum Adaptive Arrays,” IEEE Trans.
AES, Vol. 30, No. 2, April 1994, pp. 298–306.
[27] A. J. Berni, “Weight Jitter Phenomena in Adaptive Control Loops,” IEEE Trans. Aerosp.
Electron. Syst., Vol. AES-14, No. 4, July 1977, pp. 355–361.
[28] L. E. Brennan and I. S. Reed, “Effect of Envelope Limiting in Adaptive Array Control Loops,”
IEEE Trans. Aerosp. Electron. Syst., Vol. AES-7, No. 4, July 1971, pp. 698–700.
[29] S. P. Applebaum and D. J. Chapman, “Adaptive Arrays with Main Beam Constraints,” IEEE
Trans. Antennas Propag., Vol. AP-24, No. 5, September 1976, pp. 650– 662.
[30] R. T. Compton, Jr., “Adaptive Arrays: On Power Equalization with Proportional Control,”
Ohio State University, Columbus, Quarterly Report 3234-1, Contract N0019-71-C-0219,
December 1971.
[31] C. L. Zahm, “Application of Adaptive Arrays to Suppress Strong Jammers in the Presence
of Weak Signals,” IEEE Trans. Aerosp. Electron. Syst., Vol. AES-9, No. 2, March 1973,
pp. 260–271.
[32] B. W. Lindgren, Statistical Theory, New York, Macmillan, 1960, Ch. 5.
[33] M. R. Hestenes and E. Stiefel, “Method of Conjugate Gradients for Solving Linear Systems,”
J. Res. Natl. Bur. Stand., Vol. 29, 1952, p. 409.
[34] J. D. Powell, “An Iterative Method for Finding Stationary Values of a Function of Several
Variables,” Comput. J. (Br.), Vol. 5, No. 2, July 1962, pp. 147–151.
[35] R. Fletcher and M. J. D. Powell, “A Rapidly Convergent Descent Method for Minimization,”
Comput. J. (Br.), Vol. 6, No. 2, July 1963, pp. 163–168.
[36] R. Fletcher and C. M. Reeves, “Functional Minimization by Conjugate Gradients,” Comput.
J. (Br.), Vol. 7, No. 2, July 1964, pp. 149–154.
[37] D. G. Luenberger, Optimization by Vector Space Methods, New York, Wiley, 1969, Ch. 10.
[38] L. Hasdorff, Gradient Optimization and Nonlinear Control, New York, Wiley, 1976.
[39] L. S. Lasdon, S. K. Mitter, and A. D. Waren, “The Method of Conjugate Gradients for
Optimal Control Problems,” IEEE Trans. Autom. Control, Vol. AC-12, No. 2, April 1967,
pp. 132–138.
[40] O. L. Frost, III, “An Algorithm for Linearly Constrained Adaptive Array Processing,” Proc.
IEEE, Vol. 60, No. 8, August 1972, pp. 926–935.
[41] A. E. Bryson, Jr. and Y. C. Ho, Applied Optimal Control, Waltham, MA, Blaisdell, 1969,
Ch. 1.
[42] E. J. Kelly and M. J. Levin, “Signal Parameter Estimation for Seismometer Arrays,”
Massachusetts Institute of Technology, Lincoln Laboratories Technical Report 339, January
1964.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 237
[43] A. H. Nuttall and D. W. Hyde, “A Unified Approach to Optimum and Suboptimum Processing
for Arrays,” U.S. Navy Underwater Sound Laboratory, New London, CT, USL Report 992,
April 1969.
[44] R. N. McDonald, “Degrading Performance of Nonlinear Array Processors in the Presence of
Data Modeling Errors,” J. Acoust. Soc. Am., Vol. 51, No. 4, April 1972, pp. 1186–1193.
[45] D. Q. Mayne, “On the Calculation of Pseudoinverses,” IEEE Trans. Autom. Control, Vol.
AC-14, No. 2, April 1969, pp. 204–205.
[46] K. Takao, M. Fujita, and T. Nishi, “An Adaptive Antenna Array Under Directional Constraint,”
IEEE Trans. Antennas Propag., Vol. AP-24, No. 5, September 1976, pp. 662–669.
[47] O. L. Frost, III, “Adaptive Least Squares Optimization Subject to Linear Equality Con-
straints,” Stanford Electronics Laboratories, Stanford, CA, DOC. SEL-70-053, Technical
Report TR6796-2, August 1970.
[48] J. L. Moschner, “Adaptive Filtering with Clipped Input Data,” Stanford Electronics Labora-
tories, Stanford, CA, Dec. SEL-70-053, Technical Report TR6796-1, June 1970.
[49] C. Baird, and G. Rassweiler, “Adaptive Sidelobe Nulling Using Digitally Controlled Phase-
Shifters,” IEEE Transactions on Antennas and Propagation, Vol. 24, No. 5, 1976, pp. 638–649.
[50] H. Steyskal, “Simple Method for Pattern Nulling by Phase Perturbation,” IEEE Transactions
on Antennas and Propagation, Vol. 31, No. 1, 1983, pp. 163–166.
[51] R. L. Haupt, “Adaptive Nulling in Monopulse Antennas,” IEEE Transactions on Antennas
and Propagation, Vol. 36, No. 2, 1988, pp. 202–208.
[52] D. S. De Lorenzo, J. Gautier, J. Rife, P. Enge, and D. Akos, “Adaptive Array Processing for
GPS Interference Rejection,” Proc. ION GNSS 2005, pp. 618–627.
[53] H. S. Lu, “Polarization Separation by an Adaptive Filter,” IEEE Trans. Aerosp. Electron. Sys.,
Vol. AES-9, No. 6, November 1973, pp. 954–956.
[54] C. L. Zahm, “Effects of Errors in the Direction of Incidence on the Performance of an Adaptive
Array,” Proc. IEEE, Vol. 60, No. 8, August 1972, pp. 1008–1009.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:47 238
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:51 239
CHAPTER
The usefulness of an adaptive array often depends on its convergence rate. For example,
when adaptive radars simultaneously reject jamming and clutter while providing auto-
matic platform motion compensation, then rapid convergence to steady-state solutions is
essential. Convergence of adaptive sensor arrays using the popular maximum signal-to-
noise ratio (SNR) or least mean squares (LMS) algorithms depend on the eigenvalues
of the noise covariance matrix. When the covariance matrix eigenvalues differ by orders
of magnitude, then convergence is exceedingly long and highly example dependent. One
way to speed convergence and circumvent the convergence rate dependence on eigenvalue
distribution is to directly compute the adaptive weights using the sample covariance matrix
of the signal environment [1–3].
Rx x = E{xx† } (5.1)
When the desired signal is absent, then only noise and interference are present and
Rx x = Rnn (5.2)
239
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:51 240
When the desired signal is present, then from Chapter 3 the optimal weight vector solution
is given by
wopt = R−1
x x rxd (5.3)
where rxd is the cross-correlation between the random vector x(t) and the reference signal
d(t). When the desired signal is absent, then the optimal weight vector solution is given by
wopt = R−1 −1 ∗
nn rxd = Rnn b (5.4)
where b∗ is the vector of beam steering signals matched to the target Doppler frequency
and angle of incidence. Note that specifying rxd is equivalent to specifying b∗ .
If the signal, clutter, and interference situation are known a priori, then the covariance
matrix is evaluated and the optimal solution for the adaptive weights is computed directly
using either (5.3) or (5.4). In practice the signal, clutter, and interference situation are
not known a priori, and furthermore the interference environment frequently changes
due to the presence of moving near-field scatterers, antenna motion, interference, and
jamming. Consequently, the adaptive processor continually updates the weight vector to
respond to the changing environment. In the absence of detailed a priori information, the
weight vector is updated using estimates of Rx x or Rnn , and rxd from a finite observation
interval and substituting into (5.3) or (5.4). This method for implementing the adaptive
processor is referred to as the DMI or sample matrix inversion (SMI) technique. The
estimates of Rx x , Rnn , and rxd are based on the maximum likelihood (ML) principle,
which yields unbiased estimates having minimum variance [4]. Although an algorithm
based on DMI theoretically converges faster than the LMS or maximum SNR algorithms,
the covariance matrix could be ill conditioned, so the degree of eigenvalue spread also
affects the practicality of this approach.
It is worth noting that when the covariance matrix to be inverted has the form of a
Toeplitz matrix (a situation that arises when using tapped-delay line channel processing),
then the matrix inversion algorithm of W. F. Trench [5] can be exploited to facilitate the
computation. The convergence results discussed in this chapter assume that all compu-
tations are done with sufficient accuracy to overcome the effects of any ill-conditioning
and therefore represent an upper limit on how well any DMI approach can be expected to
perform.
ŵ1 = R̂−1
x x rxd (5.5)
and assuming x(t) contains the desired signal, where R̂x x is the sample covariance estimate
of Rx x , or by using
ŵ2 = R̂−1 −1 ∗
nn rxd = R̂nn b (5.6)
if we assume x(t) does not contain the desired signal where R̂nn is the sample covariance
estimate of Rnn .
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:51 241
where s denotes the desired signal vector component of x [it will be recalled from Chapter 3
that s(t) = s(t)v]. The SNR (s/n)2 has meaning only during time intervals when a desired
signal is present; the weight adjustment in this case takes place when the desired signal
is absent. The “rate of convergence” of the two algorithms (5.5) and (5.6) depends on the
output SNR normalized to the optimum output SNR, SN o , compared with the number of
independent signal samples K used to obtain the required sample covariance matrices.
Assuming that all signals present at the array input are modeled as sample functions
from zero-mean Gaussian processes, then an ML estimate of Rx x (or Rnn when the desired
signal is not present) is formed using the sample covariance matrix given by
1
K
R̂xx = x( j)x† ( j) (5.9)
K j=1
where x( j) denotes the jth time sample of the signal vector x(t). Note that the assumption
of independent zero-mean samples implies E[x(i)x† ( j)] = 0 for i = j.
Since each element of the matrix R̂x x is a random variable, the output SNR is also a
random variable. It is instructive to compare the actual SNR obtained using ŵ1 and ŵ2 of
(5.5) and (5.6) with the optimum SNR obtained using (5.3) and (5.4) (S No = s† R−1 nn s), by
forming the normalized SNR as follows:
(s/n)1
ρ1 = (5.10)
SN o
(s/n)2
ρ2 = (5.11)
SN o
It can be shown [2] that the probability distribution of ρ2 is described by the incomplete
beta distribution given by
y
K!
Pr(ρ2 ≤ y) = (1 − u) N −2 u K +1−N du (5.12)
(N − 2)!(K + 1 − N )!
0
where
K = total number of independent time samples used in obtaining R̂nn
N = number of adaptive degrees of freedom
The probability distribution function of (5.12) contains important information concerning
the convergence of the DMI algorithm that is easily seen by considering the mean and the
variance of ρ2 . From (5.12) it follows that the average value of ρ2 is given by
K +2− N
E{ρ2 } = ρ 2 = (5.13)
K +1
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:51 242
(K + 2 − N )(N − 1)
var(ρ2 ) = (5.14)
(K + 1)2 (K + 2)
For fixed K and N, (5.14) suggests that var(ρ2 ) is independent of the amount of noise
the system must contend with and the eigenvalue spread of the noise covariance matrix.
Recalling that ρ is a normalized SNR, however, we see that both the actual SNR (s/n)
and the optimum SNR SN o are affected in the same way by any noise power increase, so
the normalized ratio remains the same, and the variance of the normalized ratio likewise
remains unchanged. Eigenvalue spread has no effect on (5.14), since this expression as-
sumes that the sample matrix inversion is computed exactly. As a result, (5.14) contains
only the sample covariance matrix estimation errors. The effect of eigenvalue spread on
the matrix inversion computation is addressed in a later section.
A plot of (5.13) in Figure 5-1 (we assume that N is significantly larger than 2) shows
that so long as K ≥ 2N , the loss in ρ 2 due to nonoptimum weights is less than 3 dB. This
result leads to the convenient rule of thumb that the number of time samples required to
obtain a useful sample covariance matrix (when the desired signal is absent) is twice the
number of adaptive degrees of freedom.
Next, consider the convergence behavior when the signal is present while estimating
w with (5.5). Rather than attempt to derive the probability distribution function of ρ1
directly, it is more convenient to exploit the results obtained for ρ2 by defining the random
variable
r†xd R̂−1 † −1
x x ss R̂x x rxd
ρ1 = (5.15)
r†xd R̂−1 −1 † −1
x x Rx x R̂x x rxd s Rx x s
which has the same probability distribution function as that of ρ2 . Knowing the statistical
properties of ρ1 and the relationship between ρ1 and ρ1 then enables the desired information
about ρ1 to be easily obtained. It can be shown that the relationship between ρ1 and ρ1 is
r 2 0.5
0
N 2N 3N 4N 5N 6N 7N 8N
K
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:51 243
given by [3]
ρ1
ρ1 = (5.16)
S No (1 − ρ1 ) + 1
Since ρ1 has the same probability distribution function as ρ2 , it immediately follows that
and
The inequality expressed by (5.17) implies that, on the average, the output SNR
achieved by using ŵ1 = R̂−1 −1
x x rxd is less than the output SNR achieved using ŵ2 = R̂nn rxd
(except in the limit as K → ∞, in which case both estimates are equally accurate). This
behavior derives from the fact that the presence of the desired signal increases the time
required (or number of samples required) to obtain accurate estimates of Rx x from the
sample covariance matrix compared with the time required to obtain accurate estimates
of Rnn when the desired signal is absent. The limit expressed by (5.18) indicates that for
S N o < 1, the difference in SNR performance obtained using ŵ1 or ŵ2 is negligible.
The mean of ρ1 in (5.16) can be expressed as the following infinite series [3]:
∞
a b b + 1 i + b − 1
E{ρ1 } = 1+ (−S No )i ··· ·
a+b i=1
a+b+1 a+b+2 a+b+i
(5.19)
= 20
0.4
−1 r = 10
ŵ1 = R̂xx xd
0.2 N=4
0
0.1 1.0 10 100
SNo
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:51 244
x(t) enables a desired signal-free noise vector to be formed that can be used to generate
the sample noise covariance matrix R̂nn . An improper procedure that is occasionally
suggested for eliminating the desired signal component is to form the sample covariance
matrix D̂, where
D̂ = R̂x x − ss† (5.20)
so that
E{D̂} = Rnn (5.21)
The procedure suggested by (5.20) is unsatisfactory for obtaining the fast convergence
associated with ŵ2 because even though E {D̂} = Rnn , the weight vector estimate obtained
using ŵ3 = D̂−1 rxd results only in an estimate that is a scalar multiple of ŵ1 = R̂−1
x x rxd
and therefore has the associated convergence properties of ŵ1 . This fact may easily be
seen by forming
ŵ4 = R̂−1
x x r̂xd (5.23)
1
K
r̂xd = x( j)d ∗ ( j) (5.24)
K j=1
The transient behavior of the DMI algorithm represented by ŵ4 of (5.23) is different from
that found for ŵ1 and ŵ2 of (5.5) and (5.6), respectively.
As an example, consider a four-element uniform array with λ/2 spacing. The desired
signal is incident at 0◦ and one interference signal is incident at 45◦ with σnoise = 0.01.
The amplitude and phase of the weights as a function of sample are shown in Figure 5-3
and Figure 5-4. The adapted pattern after 500 samples is shown in Figure 5-5. The weights
do not vary much after 50 samples. A null appears where a sidelobe peak used to be.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 25, 2010 21:9 245
1 FIGURE 5-3
Magnitude of the
DMI weights versus
0.8 sample.
0.6
|wn|
0.4
0.2
0
0 100 200 300 400 500
Sample
50
−50
−100
−150
−200
0 100 200 300 400 500
Sample
−5
−15
Quiescent
−25
−90 −45 0 45 90
q (degrees)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:51 246
FIGURE 5-6
Adaptive array x1 (t) x2(t) xN(t)
configuration for
considering transient
w1 w2 wN
behavior of ŵ4 .
Σ
Reference
signal + −
w† x(t)
x0 (t)
Error
signal
e (t)
The transient response characteristics of ŵ4 are determined by considering the least
squares estimate of wopt based on K independent samples of the input vector x. It is
convenient in posing this problem to assume the adaptive array configuration shown in
Figure 5-6. This configuration represents two important adaptive array structures in the
following manner. If x0 (t) = d(t), the reference signal representation of the desired signal,
then the minimum mean square error (MMSE) and the maximum output SNR are both
given by the Wiener solution wopt . Likewise, if x0 (t) represents the output of a reference
antenna (usually a high-gain antenna pointed in the direction of the desired signal source),
then the configuration represents a coherent sidelobe canceller (CSLC) system, for which
obtaining the MMSE weight vector solution minimizes the output error (or residue) power
and hence minimizes the interference power component of the array output.
The transient response of ŵ4 given by (5.23) is characterized in terms of the output SNR
versus the number of data samples K used to form the estimates R̂x x and r̂xd . An alternate
way of characterizing performance that is appropriate for CSLC applications considers
the output residue power versus the number of data samples. System performance in
terms of the output residue power is considered. The output residue power (or MSE) is
given by
1
k
ξ̂ (ŵ) = | e(i) |2 = σ̂02 − ŵ† r̂xd − r̂†xd ŵ + ŵ† R̂x x ŵ (5.26)
K i=1
1
K
σ̂02 = x0 (i)x0∗ (i) (5.27)
K i=1
The transient behavior of the system is characterized by evaluating the statistical properties
of ξ(ŵ) for a given number of data samples K. Assume that x(i) and x0 (i) are sample
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:51 247
where rx x0 plays the role of rxd so that R̂x x is given by the estimates R̂x x , r̂xd , and σ̂02 .
The sample covariance matrix R̂x x has the following important properties:
1. The elements of R̂x x are jointly distributed according to the complex Wishart probability
density function [6]
|A| K −N −1 exp[−tr(R −1
x x A)]
p(A) = (5.30)
π 1/2(N +1)N (K )(K − 1) · · · (K − N )|R x x | K
where
A = K R̂x x
(k) = (k − 1)!
2. R̂xx is the ML estimate of R x x [6]. Therefore R̂x x , r̂xd , and σ̂02 are the ML estimates
of Rx x , rxd , and σ02 , respectively.
By using a series of transformations on the partitioned matrix R̂x x , the following
important results are obtained [7]:
1. The mean and variance of the sample MSE ξ̂ , realized using ŵ4 of (5.23) is given by
N
E{ξ̂ } = 1 − ξmin (5.31)
K
1 N
var{ξ̂ } = 1− ξmin
2
(5.32)
K K
where ξmin = σ02 − r†xd R−1
x x rxd .
2. The difference between the output residue power ξ(ŵ4 ) and the minimum output residue
power ξmin is a direct measure of the quality of the array performance relative to
the optimum. The normalized performance quality parameter (or “misadjustment”)
M = r 2 is defined as
ξ(ŵ4 ) − ξmin
r2 = (5.33)
ξmin
The parameter r is a random variable having the density function
K! r 2N −1
p(r ) = 2 · , 0<r <∞ (5.34)
(K − N )!(N − 1)! (1 + r 2 ) K +1
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:51 248
The probability density function of ρ3 is difficult to evaluate in closed form, but the
mean and variance of ρ3 can be determined numerically using the relations given as
follows:
1
ρ3 = 2
(5.39)
C + sin φ1
(1 + S No ) 2 − S No
C+sin φ1 cos2 φ2
√
where C = (1/r ) (s/n)3 and where the joint density function of r, φ1 , and φ2 is
given by
2 K! r 2N −1
p(r, φ1 , φ2 ) = · (sin φ1 )2N −2 (sin φ2 )2N −3
π (K − N )!(N − 2)! (1 + r 2 ) K +1
(5.40)
for 0 < r < ∞, 0 ≤ φ1 < π , 0 ≤ φ2 < π . Note that r , φ1 , and φ2 are statistically
independent and that r has the same density as (5.34). With the foregoing expressions
the numerical evaluation of E{ρ3 } can be obtained from
E{ρ3 } = P(r ) P(φ1 ) ρ3 P(φ2 )dφ2 dφ1 dr (5.41)
R 1 2
E {ρ32 } can likewise be obtained from (5.41) with ρ32 replacing ρ3 . The variance of ρ3
is then given by var{ρ3 } = E{ρ32 } − E 2 {ρ3 }.
4. The normalized MSE performance measure defined by
2K ξ̂
ξ̂ N = (5.42)
ξmin
The results just summarized place limits on the transient performance of the DMI
algorithm. Let us first consider the results expressed by (5.36) and (5.37). From (5.36)
it is seen that the output residue power is within 3 dB of the optimum value after only
2N distinct time samples, or within 1 dB after 5N samples, thereby indicating rapid con-
vergence independent of the signal environment or the array configuration. We see this
rapid convergence property, however, applies directly to the interference suppression of
a sidelobe canceller (SLC) system assuming no desired signal is present when form-
ing R̂x x .
In communications systems the desired signal is usually present, and the SNR perfor-
mance measure is the primary quantity of interest rather than the MSE, ξ . Furthermore, in
radar systems the SLC is often followed by a signal processor that rejects clutter returns,
so only that portion of the output residue power due to radiofrequency (RF) interference
(rather than clutter) must be suppressed by the SLC system. Let us now show that the
presence of either the desired signal or clutter returns in the main beam of an SLC system
acts as a disturbance that tends to slow the rate of convergence of a DMI algorithm.
Consider the radar SLC configuration shown in Figure 5-7. The system consists of an
SLC designed to cancel only interference followed by a signal processor to remove clutter.
To simplify the discussion, assume the clutter power ξc2 received in the main antenna is
much larger than clutter power entering the low-gain auxiliary antennas so that clutter in
the auxiliary channels is neglected and ξmin = σc2 + ξ N0 where ξ N0 represents the minimum
output interference plus thermal noise power. Furthermore, assume that the clutter returns
are represented as a sample function from a stationary, zero-mean Gaussian process. It
then follows from (5.25), (5.33), and (5.36) that
E[ξ(ŵ) − ξmin ] N σ2
= 1+ c (5.43)
ξ N0 K−N ξ N0
From (5.43) it is evident that the presence of main beam clutter prolongs conver-
gence of the DMI algorithm by an amount that is approximately proportional to the main
antenna clutter power divided by the minimum output interference power. This slower
convergence is due to the presence of clutter-jammer cross-terms in r̂xd , which results in
noisier estimates of the weights. For rapid convergence, adaptation should be confined
to time intervals that are relatively clutter free, or a means for minimizing clutter terms
present in the estimate of r̂xd must be found.
The foregoing result is also applicable to communications systems where the main
antenna is pointed toward an active desired signal source. Let σs2 represent the desired
signal power received in the main channel, and assume the desired signal entering the
auxiliary channels can be neglected. Then, the SLC output signal to noise ratio SNR can
N
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:51 250
σs2
SNR = (5.44)
σs2r 2 + ξ N0 (1 + r 2 )
SNR 1
ρ4 = = (5.45)
SN o 1 + [1 + (SN o )]r 2
Note that 0 ≤ ρ4 ≤ 1 so the optimum value of ρ4 is unity. When the output signal to
interference plus noise ratio is small so that S No 1, then the probability distribution of
ρ4 is approximated by the following beta probability density function:
K!
P(ρ4 ) = (1 − ρ4 ) N −1 ρ4K −N (5.46)
(N − 1)!(K − N )!
On comparing (5.46) with the probability density function contained in (5.12), it imme-
diately follows that the two density functions are identical provided that N in (5.12) is
replaced by N + 1. It then follows from (5.13) for the CSLC system of Figure 5-6 that
K − N +1
E{ρ4 } = ; SN o 1 (5.47)
K +1
Hence, for small S N o only K = 2N − 1 independent samples are needed to converge
within 3 dB of S N o . For large SNR ( K /N ), the expected SNR (unnormalized) is
approximated by
K K
(SNR) ∼
= − 1; SN o ; K >N (5.48)
N N
The presence of a strong desired signal in the main channel therefore slows convergence
to the optimum SNR but does not affect the average output SNR after K samples under
the conditions of (5.48).
Finally, consider the results given in (5.39) and (5.40) for the case x0 (t) = d(t)
(continuous reference signal present). For large S N o , the distribution of ρ3 in (5.39) is
approximated by the density function of (5.46) so that
K − N +1
E{ρ3 } ∼
= ; SN o 1 (5.49)
K +1
However, the rate of convergence decreases as S N o decreases below zero decibels. This
effect is illustrated by the plot of E{ρ3 } versus K in Figure 5-8.
The behavior described for a reference signal configuration is just the converse of that
obtained for the SLC configuration with the desired signal present. The presence of the
desired signal in the main channel of the SLC introduced a disturbance that reduced the
accuracy of the weight estimate and slowed the convergence. For the reference signal
configuration, however, the estimates r̂xd and R̂x x are highly correlated under strong
desired signal conditions, and the errors in each estimate tend to compensate each other,
thereby yielding an improved weight estimate and faster convergence.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:51 251
0.5 of SN o.
0.4
= −20 dB
0.3
0.2
0.1
30
Input IRN = 0 dB,
No Diagonal 3 dB
Loading. From Ganz
20 SINR = 15.930 et al., IEEE Trans.
Ant. & Prop., March
10 1990.
SIR-theory SINR-asymptotic value
0 SINR-theory SIR-100 run simulation
SIR-asymptotic value SINR-100 run simulation
−10
5 10 100 1000 10000
Sample size = K
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:51 252
FIGURE 5-10 50
SINR and SIR vs. K,
the number of 40 SIR = 36.364
samples. Input SNR
−10
5 10 100 1000 10000
Sample size = K
FIGURE 5-11 50
SINR and SIR vs. K,
SIR = 45.867
the number of 40 SIR-theory
samples. Input SNR SINR-theory
SIR and SINR in dB
= 0 dB, Negative
Diagonal 3 dB
20
Loading. From Ganz SINR = 15.946
et al., IEEE Trans.
Ant. & Prop., March 10
1990. SINR-asymptotic value
0 SIR-100 run simulation
SINR-100 run simulation
−10
5 10 100 1000 10000
Sample size = K
the SIR and SINR versus number of samples for the case when there is no diagonal loading.
Figure 5-10 shows the same results when 3 dB of positive loading is added to the covariance
matrix (i.e., σ 2 is added to each diagonal element). The SIR asymptote is approximately
5.8 dB lower than it was without loading, but the SINR asymptote is essentially unchanged.
The positive loading has decreased the number of samples required for convergence of the
SINR curves to their asymptotic value. Figure 5-11 shows the results of the SIR and SINR
versus number of samples for the case when 0.5σ 2 is subtracted (negative loading) from
each diagonal element of the covariance matrix. The SINR asymptote remains essentially
unchanged, but the SIR asymptote is 5.9 dB higher than it was without loading. Negative
loading achieves a better SIR but a slower convergence to the asymptotic value compared
with the unloaded case. Notice, however, that with negative loading both the output SIR
and SINR curves are quite erratic when a small number of samples are taken; this behavior
is due to the variance of the output powers is large for small sample sizes.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:51 253
5.3.1 Triangularization of R
Suppose there is a transformation matrix, Q, such that QR = U where U is an upper
triangular matrix. In this case it is possible to solve the system Uw = Qv = y for w
directly without finding a matrix inverse by means of a simple back-substitution algorithm
[11]. The following back-substitution algorithm computes the elements of w directly.
For j = N, N − 1, . . . , 1, compute
⎛ ⎞
N
w( j) = ⎝y( j) − u( j, k)w(k)⎠ u( j, j) (5.50)
k= j+1
The algorithm begins with w(N) = y(N)/U(N,N) and progresses backward until
computing element w(1). Suppose instead that QR = L where L is a lower triangular
matrix; then a simple forward-substitution algorithm (which is an analog of the forward-
substitution algorithm) yields the desired solution for the vector w.
Sometimes there is a need to explicitly have the inverse Z = U−1 . A simple algorithm
for obtaining the elements of Z (which is also upper triangular) is given by
z( j, j) = 1/u( j, j) (5.52)
j−1
z(k, j) = − z(k, m)u(m, j) u( j, j), k = 1, . . . , j − 1 (5.53)
m=k
Givens rotations annihilate one element at a time using the planar rotation transfor-
mation
∗
c s
G= (5.54)
−s ∗ c
v
With the complex vector v = 1 , we find that
v2
∗
c v1 + sv2
Gv = (5.55)
−s ∗ v1 + cv2
and require that s ∗ v1 = cv2 with the unitary condition |c|2 + |s|2 = 1. It follows that
v1
c= (5.56)
|v1 |2 + |v2 |2
v∗2
s= (5.57)
|v1 |2 + |v2 |2
v1
Gv = v with v1 = (5.58)
0
where v1 = |v1 |2 + |v2 |2 (5.59)
a11 a∗
where the 3,1 product element becomes zero since c1 a31 = s1∗ a11 , c1 = , s1 = 31
r1 r1
and r1 = |a11 | + |a31 | . The next step is to annihilate element a21 , which is done by
2 2
0 0 1 0 a32 a33
(5.65)
The operation for computing the triangularization of a matrix, R, with elements r(i,j)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:51 256
s
Tu R = R̃ (5.70)
0
N
γ =β · u(i)R(i, j) (5.75)
i=1
R̃(i, j − 1) = R(i, j) + γ u(i), i = 1, . . . , N (5.76)
where R̃(i,j) may be replaced by R(i,j). This completes the reduction of an N × N matrix
to a triangular form.
The Cholesky factorization of R with positive diagonal elements is given by the following
algorithm [11,13]:
For j = 1, 2, . . . , N − 1
L( j, j) = R( j, j)1/2 (5.77)
For k = j + 1, . . . , N
L(k, j) = R(k, j)/L(j, j) (5.78)
For i = k, . . . , N
R(i, k) = R(i, j) − L(i, j)L(k, j) (5.79)
end for
L(N, N) = P(N , N )1/2 (5.80)
An analogous algorithm starting with U(N, N) obviously applies for the upper triangular
Cholesky factorization.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:51 257
for j = N, N − 1, . . . , 2
d(j) = R(j, j) (5.81)
U(j, j) = 1
for k = 1, 2, . . . , j − 1
U(k, j) = R(k, j)/d(j) (5.82)
for i = 1, 2, . . . , k
R(i, k) = R(i, k) − U(i, j)U(k, j)d(j) (5.83)
end for
U(1, 1) = 1 and d(1) = R(1, 1) (5.84)
The U-D factorization given by (5.81)–(5.84) as well as the L-D factorization is different
from the spectral factorization involving eigenvalues and normalized eigenvectors. Spec-
tral factorization renders finding the desired inverses particularly simple, thereby making
the additional computation required by the accompanying eigenvalue and eigenvector
computations worthwhile. Note that the U-D and L-D factorizations are also different
from each other.
The matrix R−1 is formed by inspection from the spectral factorization as follows.
Since R−1 = [MMT ]−1 = (MT )−1 −1 M−1 where the elements of −1 are merely
1/λ(i) for i = 1, . . . , N , MH = M−1 and (MT )−1 = M since M is unitary. The extra
computation involved in finding the eigenvalues and associated eigenvectors is rewarded
by the ease in finding the desired inverses.
The inverses required by the U-D factorization are U−1 , D−1 , where the elements of
−1
D are merely the inverse of the elements of D, and the inverse of U is given by equations
(5.51)–(5.53).
as a function of K (the number of independent samples used in obtaining R̂x x and r̂xd ) for
selected signal conditions. Assuming an input interference-to-signal ratio of 14 dB and an
input signal-to-thermal noise ratio of 0 dB, ρ was obtained for ŵ1 and ŵ4 by averaging
50 independent responses for a four-element linear array for which d/λ = 12 . A single
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:51 258
0.01
1 10 100 400
Number of samples (K )
directional interference source was assumed that was separated from the desired signal
location by 75◦ . The resulting transient response is given in Figure 5-12.
The transient response obtained using ŵ1 and ŵ4 can be compared with the transient
response that would result using the LMS algorithm having an iteration period equal to the
intersample collection period of the DMI algorithm by considering the behavior of ρ(K ).
For the DMI approach, K represents the number of independent signal samples used to
compute the sample covariance matrix, whereas for the LMS algorithm K represents the
number of iterations completed. This comparison of the DMI and LMS algorithms is not
satisfactory for the following reasons:
1. The transient response of the LMS algorithm depends on the selection of the step size,
which can be made arbitrarily small (thereby resulting in an arbitrarily long transient
response time constant).
2. The transient response of the LMS algorithm depends on the starting point at which the
initial weight vector guess is set. A good initial guess may result in excellent transient
response.
3. Whereas the variance of ρ(K) decreases as K increases with the DMI approach, the
variance of ρ(K) remains constant (in the steady state) as K increases for the LMS
algorithm since the steady-state variance is determined by the step size selection.
Nevertheless, a comparison of ρ(K ) behavior does yield an indication of the output SNR
response speed and is therefore of some value.
The LMS algorithm convergence condition (4.49) is violated if the step size s exceeds
1/PIN , where PIN = total array input power from all sources. Consequently, by selecting
s = 0.4/PIN , the step size is only 4 dB below the convergence condition limit, the LMS
control loop is quite noisy, and the resulting speed of the transient response is close to the
theoretical maximum. Since the maximum time constant for the LMS algorithm transient
response is associated with the minimum covariance matrix eigenvalue
1
τmax ∼
= (5.85)
ks tλmin
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:51 259
A convenient starting point for the weight vector with the LMS algorithm is w0T =
[1, 0, 0, 0], since this corresponds to an omnidirectional array pattern. For the foregoing
conditions, the behavior of ρ(K ) resulting from the use of ŵ3 , ŵ4 and the two versions of
the LMS algorithm given by
w(k + 1) = w(k) + ks t [rxd − x(k)x† (k)w(k)], rxd given (5.88)
and
w(k + 1) = w(k) + ks t [x(k) d ∗ (k) − x(k)x† (k)w(k)], d(k) given (5.89)
was determined by simulation. The results are illustrated in Figure 5-12 where for the
specified conditions S N o = 3.8.
The results of Figure 5-12 indicate that the response resulting from the use of the DMI
derived weights ŵ1 and ŵ4 is superior to that obtained from the LMS derived weights.
Whereas the initial response of the LMS derived weights indicated improved output SNR
with increasing K, this trend reverses when the LMS algorithm begins to respond along
the desired signal eigenvector, since any decrease in the desired signal response with-
out a corresponding decrease in the thermal noise response causes the output signal to
thermal noise ratio to decrease. Once the array begins to respond along the thermal noise
eigenvectors, then ρ again begins to increase.
The undesirable transient behavior of the two LMS algorithms in Figure 5-12 can
be avoided by selecting an initial starting weight vector different from that chosen for
the foregoing comparison. For example, by selecting w(0) = rxd the initial LMS algo-
rithm response is greatly improved since this initial starting condition biases the array
pattern toward the desired signal direction thereby providing an initially high output SNR.
Furthermore, by selecting w(0) = αrxd where α is a scalar constant the initial starting
condition can also result in a monotonically increasing ρ(K ) as K increases for the two
1 It is assumed that each weighted channel contains thermal noise (with noise power σ 2 ) that is uncorrelated
between channels.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:51 260
LMS algorithms since by appropriately weighting rxd , the magnitude of the initial array
response to the desired signal can be made small enough so that as adaptation proceeds,
the resulting ρ(K ) always increases.
Improvement of the transient response through judicious selection of the initial starting
condition can also be introduced into the DMI derived weight vectors as well. For example,
by selecting
⎡ ⎛ ⎞⎤−1
1 K
ŵ1 = ⎣ ⎝ x( j)x†( j) + αI⎠⎦ rxd (5.90)
K j=1
then even before forming an estimate of Rx x the weight vector is biased toward the
desired signal direction, and the transient responses corresponding to ŵ1 and ŵ4 in Fig-
ure 5-12 is greatly improved. Nevertheless, even the improved transient response for an
LMS algorithm obtained by appropriately biasing the initial weight vector is slower than
the convergence speed of DMI.
It is also interesting to obtain the transient response of ρ(K ) with two interference sig-
nals present. Assume one interference-to-signal ratio of 30 dB and a second interference-
to-signal ratio of 10 dB, where the stronger interference signal is located 30◦ away from the
desired signal, the weaker interference signal is located 60◦ away from the desired signal,
with all other conditions the same as for Figure 5-12. The resulting transient response for
the two DMI derived weight vectors and for the two LMS algorithms [with d(t) given and
with rxd given] with initial starting weight vector = [1, 0, 0, 0] is illustrated in Figure 5-13.
The presence of two directional interference sources with widely different power levels
results in a wide eigenvalue spread and a consequent slow convergence rate for the LMS
algorithms. Since PIN /λ1 is now 40 times larger than for the conditions of Figure 5-12, the
time constant τ1 is now 40 times greater than before. The DMI derived weight transient
response, however, is virtually unaffected by this eigenvalue spread.
The principal convergence results for DMI algorithms under various array configura-
tions and signal conditions can be conveniently summarized as shown in Table 5-1. The
derivation of these results may be found in [2,3].
0.01
1 10 100 400
Number of samples (K )
Monzingo-7200014
book
TABLE 5-1 Comparison of DMI Algorithm Convergence Rates for Selected Array Configurations and Signal Conditions
2N , SN o 1
Sidelobe canceller R̂−1
x x r̂x x0
σc2
Clutter returns in main beam only Minimum output 2N 1 + , σc2 = main channel clutter power
2ξ N0
(interference + noise)
power, ξ N0
5.4
2SN o
R̂−1
x x rxd Desired signal direction of arrival Maximum SNR, SN o 2N , SN o 1
known and desired signal present S No
2N 1 + , SN o 1
261
2
Transient Response Comparisons
261
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:51 262
FIGURE 5-14 12
Output SNR 11 4 Bits
degradation versus
10 6 Bits
eigenvalue spread
matrix. When the desired signal is absent from the signal environment, then use of the beam
steering vector rxd yields the most desirable transient response characteristics. When the
desired signal is strong, however, formation of the estimate r̂ xd yields transient response
characteristics that are superior to the use of apriori information represented by rxd alone.
DMI algorithms are insensitive to eigenvalue spread until a certain critical level is
exceeded that depends on the number of bits available in the computer to perform the matrix
inversion computation. The practicality of the DMI approach is restricted by the number
of degrees of freedom of the adaptive processor. When the feasibility of the DMI approach
is not precluded, however, the additional complexity introduced by directly obtaining the
sample covariance matrix is rewarded by the rapid convergence and insensitivity to the
previously noted eigenvalue spread.
The difficult problem of the direct computation of a matrix inverse can be circum-
vented by recourse to factorization methods that boast superior accuracy and numerical
stability properties. The three methods that have been presented here (triangularization
of the covariance matrix and solution of the triangular system by forward and backward
substitution, Cholesky factorization, and spectral factorization [U-D or L-D]) offer attrac-
tive alternatives to direct matrix inversion computations. Furthermore, algorithms based
on these methods are extremely fast.
5.7 PROBLEMS
1. Development of a Recursion Formula. The DMI algorithm requires a matrix inversion
each time a new weight is to be calculated (as, e.g., when the signal environment
changes). By applying the matrix inversion lemma
[A + u† Ru]−1 = A−1 − A−1 u† [R−1 + uA−1 u† ]−1 uA−1
show that the weights can be updated at each sample using the recursive formulas
P(k − 1)x(k) ε∗ (k)
W (k) = W (k − 1) +
1 + x† (k)P(k − 1)x(k)
where
ε∗ (k) = −x† (k)P(k − 1) b∗
and
P(k − 1)x(k)x† (k)P(k − 1)
P(k) = P(k − 1) −
1 + x† (k)P(k − 1)x(k)
Hint: Apply the lemma to the inversion of
[R̂x x (k − 1) + x(k)x† (k)]
2. Similarity between Array Element Outputs and Tapped-Delay Line Outputs.
Consider an adaptive filter consisting of a tapped delay line with a complex weight at
each tap.
(a) Show that the DMI algorithm for minimizing the MSE between the filter output
and the reference signal is of the same form as for an adaptive array if the tap
outputs are taken as analogous to the array inputs.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:51 264
show that
1
ρ=
N
| qi |2
1 + (1 + SN o ) √
i=2
| SN o + q1 |2
where
⎡√ ⎤
SN o
⎢ 0 ⎥
√ ⎢ ⎥
⎢ 0 ⎥
q= SN o Ĝ − ⎢ ⎥
⎢ .. ⎥
⎣ . ⎦
0
!
SN o
Ĝ = P−1 R1/2
x x ŵ
1 + SN o
where
0 ≤ r < ∞; i + 1, 2, · · · , 2N − 1
0 ≤ φi < π
derive the density p(r , φ1 , φ2 ) of (5.40) and the expression for ρ(r , φ1 , φ2 ).
Hint: The Jacobian of the transformation is
and
π [(n + 1)/2] √
(sin φ)n dφ = π
0 [(n + 2)/2]
6. Statistical Relations [7]. Derive the result in (5.44) using (5.33) and the fact that for a
CSLC ξ(ŵ) = σs2 + σ N2 and ξmin = ξ N0 + σs2 , and σ N2 is the output noise plus jammer
power.
7. Statistical Relations [7]. Define the following transformation of R x x in (5.29):
" #
x = K σ̂02 − r̂†x x0 R̂−1
x x r̂x x0
Y = K R̂x x
ŵ = R̂−1
x x r̂xd
where
|Y |k−N +1 exp{−tr[I + (1/ξ0 ) w w† Rx x ] R−1 x x Y}
p(ŵ, Y) = (N −1)
π π
N 1/2N (k) · · · (k − N + 1)|Rx x | |ξ0 | N
k
AE − BF = I
AF + BE = 0
for the unknown matrices E and F. Premultiply the two previous equations by B
and A, respectively, and by subtracting show that
If A and B commute (so that AB = BA), show that M −1 = [A2 + B 2 ]−1 [A − jB].
The foregoing result involves the inverse of a real n × n matrix to obtain M −1
but is restricted by the requirement that A and B commute.
(b) Let C = A + B and D = A − B. Show that the original equation pair in part (a)
reduce to
CE + DF = I
−DE + CF = −I
From the results expressed in the equation pair immediately preceding, show that
either
or
provided the indicated inverses exist. The foregoing equation pair represents
alternate ways of obtaining M −1 by inverting real n × n matrices without the
restriction that A and B commute.
(c) An isomorphism exists between the field of complex numbers and a special set
of 2 × 2 matrixes, that is,
a b
a + jb ∼
−b a
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:51 267
Show that H is the inverse of G only if the original equation pair in part (a) is
satisfied. Therefore, one way of obtaining M −1 is to compute G−1 and identify
the n × n submatrices E and F appearing in G −1 . Then M −1 = E + jF. This
approach does not involve the restrictions that beset the approaches of (a) and
(b) but suffers from the fact that it requires the inversion of a 2n × 2n matrix
that drastically increases computer storage requirements, and therefore should be
used only as a last resort.
10. Development of the Cholesky Decomposition Algorithm [11]
Consider the quadratic form xT Px for the case n = 3, where P is positive definite and
symmetric.
(a) Express the quadratic form as a difference of squares plus a remainder where the
two squares involve only P(1, 1), P(1, 2), and P(1, 3) and the remainder involves
P(2, 2), P(2, 3) and P(3, 3).
(b) Now set L(1, 1) = P(1, 1)1/2 , L(2, 1) = P(2, 1)/L(1, 1), and L(3, 1) = P(3, 1)/L(1, 1).
$
3
Furthermore let y1 = L(i, 1)xi . With this notation rewrite xT Px as
i=1
P(2, 2)x22 + 2P(2, 3)x2 x3 + P(3, 3)x23 = [P(2, 2)1/2 x2 + (P(2, 3)/P(2, 2)1/2 )x3 ]2
−[(P(2, 3)/P(2, 2)1/2 )x3 ]2 + P(3, 3)x23
(d) Consistent with part (b), set L(2, 2) = P(2, 2)1/2 , L(3, 2) = P(3, 2)/L(2, 2),
and y2 = L(2, 2)x2 + L(3, 2)x3 .
Now by setting L(3, 3) = [P(3, 3) − L(3, 2)L(3, 2)]1/2 and y3 = L(3, 3)x3 , we
now have xT Px = yT y = xT LLT x.
11. Development of the Inversion of a Triangular Matrix [11]
The following identity for triangular matrices is easily verified:
−1
Rj y R −1 −1 −1
j −R j yσ j+1
= −1 = R −1
0 σ j+1 0 σ j+1 j+1
This relation enables one to compute recursively the inverse of an ever larger
matrix, that is, if R−1
j = Uj , where Rj is the upper left j × j partition of R, then show
that
U j −U j (R(1, j + 1), . . . , R( j, j + 1))T σ j+1
Uj+1 =
0 σ j+1
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:51 268
where σj+1 = 1/R(j + 1, j + 1). This result is readily transformed into the desired
algorithmic form given by (5.57), (5.58), and (5.59).
12. Structure of the Householder Transformation [16]
The Householder transformation can be written as
2uu H
H=I− = I − 2u[uH u]−1 uH = I − 2Pu
u2
where Pu = u[uH u]−1 uH
(a) Using the fact that H is Hermitian unitary (H−1 = HH = H), show that
(b) Making the change of variable u = sin(θ), show that Rn,m of part (a) becomes
1
2π
Rn,m = exp j d(n − m)u du
λ
−1
14. Computer Simulation Problem An eight-element uniform array with λ/2 spacing
has the desired signal incident at 0◦ and two interference signals incident at −21◦ and
61◦ . Use the DMI algorithm to place nulls in the antenna pattern. Assume σnoise =
0.01. In MATLAB, represent the signal by cos (2π (1 : K ) /K ) exp ( j rand) and the
interference by sign(randn(1, K )).
15. Development of the Inverse of a Toeplitz Matrix [5]
(a) A Toeplitz matrix has the form
⎡ ⎤
τ0 τ−1 τ−2 τ−3
⎢ τ1 τ0 τ−1 τ−2 ⎥
⎢ ⎥
⎣ τ2 τ1 τ0 τ−1 ⎦
τ3 τ2 τ1 τ0
to show that
λ B + ê ĝ t ê
λn Bn+1 = n n t
ĝ 1
where Eg = ĝ and Ee = ê. This results shows that given an element of Bn+1, all
the remaining elements along the same diagonal are given if we know λn , gn , and
en (the elements of the 1st row and column of Bn+1 .
(e) Since Bi+1 is persymmetric, we can write
5.8 REFERENCES
[1] I. S. Reed, J. D. Mallett, and L. E. Brennan, “Sample Matrix Inversion Technique,” Proceedings
of the 1974 Adaptive Antenna Systems Workshop, March 11–13, Vol. 1, NRL Report 7803,
Naval Research Laboratory, Washington, DC, pp. 219–222.
[2] I. S. Reed, J. D. Mallett, and L. E. Brennan, “Rapid Convergence Rate in Adaptive Arrays,”
FIEEE Trans. Aerosp. Electron. Sys., Vol. AES-10, No. 6, November 1974, pp. 853–863.
[3] T. W. Miller, “The Transient Response of Adaptive Arrays in TDMA Systems,” Ph.D. Disser-
tation, Department of Electrical Engineering, The Ohio State University, 1976.
[4] H. L. Van Trees, Detection, Estimation and Modulation Theory, Part 1, New York, Wiley,
1968, Ch. 1.
[5] S. Zohar, “Toeplitz matrix inversion: the algorithm of W. F. Trench,” J. Assoc. Comput. Mach,
Vol 16, Oct 1969, pp. 592–601.
[6] N. R. Goodman, “Statistical Analysis Based on a Certain Multivariate Gaussian Distribution,”
Ann. Math. Stat., Vol. 34, March 1963, pp. 152–177.
[7] A. B. Baggeroer, “Confidence Intervals for Regression (MEM) Spectral Estimates,” IEEE
Trans. Info. Theory, Vol. IT-22, No. 5, September 1976, pp. 534–545.
[8] B. D. Carlson, “Covariance Matrix Estimation Errors and Diagonal Loading in Adaptive
Arrays,” IEEE Trans. on Aerosp. & Electron. Sys., Vol. AES-24, No. 4, July 1988, pp. 397–
401.
[9] M. W Ganz, R. L. Moses, and S. L. Wilson, “Convergence of the SMI and the Diagonally
Loaded SMI Algorithms with Weak Interference,” IEEE Trans. Ant. & Prop., Vol. AP-38,
No. 3, March 1990, pp. 394–399.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:51 271
[10] I. J. Gupta, “SMI Adaptive Antenna Arrays for Weak Interfering Signals,” IEEE Trans. Ant.&
Prop., Vol. AP-34, No. 10, October 1986, pp. 1237–1242.
[11] G. J. Bierman, Factorization Methods for Discrete Sequential Estimation, New York, Aca-
demic Press, 1977.
[12] P. E. Gill, G. H. Golub, W. Murray, and M. A. Saunders, “Methods for Modifying Matrix
Factorizations,” Mathematics of Computation, Vol. 28, No. 126, April 1974, pp. 505–535.
[13] J. E. Gentle, “Cholesky Factorization,” Section 3.2.2 in Numerical Linear Algebra for Appli-
cations in Statistics, Berlin: Springer-Verlag, 1998, pp. 93–95.
[14] H. L. Van Trees, Optimum Array Processing, Part IV of Detection, Estimation and Modulation
Theory, New York, Wiley-Interscience, 2002.
[15] W. W. Smith, Jr. and S. Erdman, “A Note on Inversion of Complex Matrices,” IEEE Trans.
Automatic Control, Vol. AC-19, No. 1, February 1974, p. 64.
[16] M. E. El-Hawary, “Further Comments on ‘A Note on Inversion of Complex Matrices,’ ” IEEE
Trans. Automatic Control, Vol. AC-20, No. 2, April 1975, pp. 279–280.
[17] B. D. Carlson, “Equivalence of Adaptive Array Diagonal Loading and Omnidirectional Jam-
ming,” IEEE Trans. Ant. & Prop., Vol. Ap-43, No. 5, May 1995, pp. 540–541.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:51 272
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:54 273
CHAPTER
The least mean squares (LMS) and maximum signal-to-noise ratio (SNR) algorithms
avoid the computational problems associated with the direct calculation of a set of adap-
tive weights. Chapter 8 shows that random search algorithms also circumvent computa-
tional problems. These algorithms have the advantage that the required calculations are
usually much simpler than the corresponding direct calculation, are less susceptible to
hardware inaccuracy, and are continually updated to compensate for a time-varying signal
environment.
Recursively inverting the matrix circumvents many computational problems [1–4].
The recursive algorithms exhibit a steady-state sensitivity to eigenvalue spread in the signal
covariance matrix as found for direct matrix inversion (DMI) algorithms. Furthermore,
since the principal difference between the recursive methods and the DMI algorithms lies
in the manner in which the matrix inversion is computed, their rates of convergence are
comparable. The recursive algorithms are based on least square estimation techniques and
are closely related to Kalman filtering methods [5]. For stationary environments, these
recursive procedures compute the best possible selection of weights (based on a least
squares fit to the data received) at each sampling instant, whereas in contrast the LMS,
maximum SNR, and random search methods are only asymptotically optimal.
FIGURE 6-1
Conventional x1 x2 xN
N-element adaptive
array processor.
w1 w2 wN
y = wTx
Assume that the received signals xi (t) contain a directional desired signal component
si (t) and a purely random component n i (t), due to both directional and thermal noise
so that xi (t) = si (t) + n i (t). Collecting the received signals xi (t) and the multiplicative
weights w i (t) as components in the N -dimensional vectors x(t) and w(t), we write the
adaptive processor output signal y(t) as
The narrowband processor model of Figure 6-1 has been chosen instead of the more general
tapped delay line wideband processor in each element channel because the mathematical
manipulations are simplified. The weighted least squares error processor extends to the
more general tapped delay line form.
Consider the weighted least squares performance measure based on k data samples
following Baird [6]
1
k
1
(w) = αi [wT x(i) − d(i)]2 = [X(k)w − d(k)]T A−1
k [X(k)w − d(k)] (6.2)
2 i=1 2
where the elements of X(k) are received signal vector samples, and the elements of d(k)
are desired (or reference) signal samples as follows:
⎡ ⎤
xT (1)
⎢ xT (2) ⎥
⎢ ⎥
X(k) = ⎢ .. ⎥ (6.3)
⎣ . ⎦
xT (k)
⎡ ⎤
d(1)
⎢ d(2) ⎥
⎢ ⎥
d(k) = ⎢ .. ⎥ (6.4)
⎣ . ⎦
d(k)
Both (6.4) and (6.2) presume that the desired array output signal, d(t), is known and
sampled at d(1), d(2), . . . , d(k). In the performance measure of (6.2), Ak is a diagonal
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:54 275
weighting matrix that deemphasizes old data points and is of the form
⎡ k−1 ⎤
α 0 ... . 0
⎢ 0 α k−2
. . . . .⎥
⎢ ⎥
⎢ .. ⎥
Ak = ⎢ .⎥ (6.5)
⎢ ⎥
⎣ 0 ... α 0⎦
0 0 1
where 0 < α ≤ 1, so that older data have increasingly less importance. If the signal envi-
ronment is stationary so that all data samples are equally important, then Ak = I, the iden-
tity matrix. The performance measure given by (6.2) is minimized by selecting the weight
vector to yield the “best” (weighted least squares) estimate of the desired signal vector d(k).
To minimize the weighted least squares performance measure (6.2), set the derivative
of (w) with respect to w equal to zero, thereby yielding the optimum weight setting as
−1
wls (k) = XT (k)A−1
k X(k) XT (k)A−1
k d(k) (6.6)
When an additional data sample is taken, the foregoing weight vector solution is updated
in the most efficient manner. The updated signals X(k + 1) and d(k + 1) as well as the
updated matrix Ak+1 can each be partitioned as follows:
X(k)
X(k + 1) = (6.7)
xT (k
+ 1)
d(k)
d(k + 1) = (6.8)
d(k + 1)
and
⎡ ⎤
..
⎢ . 0
⎢ ·. . ⎥ ⎥
⎢ αAk .. .. ⎥
Ak+1 ⎢
=⎢ · .. ⎥ (6.9)
⎢ . 0⎥ ⎥
⎣ · · · · · ··· · · · ⎦
..
0···0 . 1
With this partitioning, the updated weight vector can be written as
−1
wls (k + 1) = XT (k + 1)A−1
k+1 X(k + 1) XT (k + 1)A−1
k+1 d(k + 1)
−1
= XT (k + 1)A−1
k+1 X(k + 1)
· αXT (k)A−1
k d(k) + x(k + 1) d(k + 1) (6.10)
From (6.10) it is seen that the updated weight vector solution requires the inversion of
the matrix [XT (k + 1)A−1k+1 X(k + 1)], which can also be expanded by the partitioning
previously given as
⎧ ⎡ ⎤ ⎫−1
⎪ .. ⎪
⎪
⎪ . 0 ⎪
⎪
⎪
⎪ ⎢ ·. . ⎥ ⎪
⎪
⎪
⎨ ⎢ −1 . . ⎥ ⎪
⎬
−1 ⎢ αA . . ⎥ X(k)
−1
X (k + 1)Ak+1 X(k + 1)
T
= [X (k) x(k + 1)] ⎢
T ⎢ k
·.. ⎥
⎪
⎪
⎪ ⎢ . 0⎥ ⎥ xT (k + 1) ⎪
⎪
⎪
⎪
⎪ ⎣ ·
· · · · · ·.· · · · ⎦ ⎪
⎪
⎪
⎩ . ⎪
⎭
0···0 . 1
−1
= α XT (k)A−1 k X (k) + x(k + 1)x (k + 1)
T T
(6.11)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:54 276
Now define
P−1 (k) = XT (k)A−1
k X(k) (6.12)
Likewise, define
P−1 (k + 1) = XT (k + 1)A−1
k+1 X(k + 1)
= α XT (k)A−1
k X(k) + x(k + 1)x (k + 1)
T
1
= α P−1 (k) + x(k + 1)xT (k + 1) (6.13)
α
Inverting both sides of (6.13) and applying the matrix inversion lemma [(D.10) of
Appendix D] then yields
1 P(k)x(k + 1)xT (k + 1)P(k)
P(k + 1) = P(k) − (6.14)
α α + xT (k + 1)P(k)x(k + 1)
By our use of (6.14) in (6.10) and recognition that wls (k) = P(k)X(k)A−1
k d(k), the updated
solution for the weight vector now becomes
P(k)x(k + 1)
wls (k + 1) = wls (k) +
α+ xT (k+ 1)P(k)x(k + 1)
· d(k + 1) − wlsT (k)x(k + 1) (6.15)
Equations (6.14) and (6.15) are the iterative relations of the recursive least squares algo-
rithm. Equations (6.14) and (6.15) are started by adopting an initial guess for the weight
vector w(0) and the initial Hermitian matrix P(0). It is common practice to select as an
initial weight vector w(0) = [1 0◦ , 0, 0, . . . , 0], thereby obtaining an omnidirectional
array pattern (provide the sensor elements each have omnidirectional patterns) and to
select P(0) as the identity matrix.
Equations (6.14) and (6.15) yield the updated weight vector in a computationally
efficient manner that avoids calculating the matrix inverses present in (6.6) and (6.10). It
is instructive to consider (6.6) in more detail for the additional insight to be gained into the
mechanics of the processor. Since the trace of Ak , tr [Ak ], is a scalar, (6.6) can be rewritten
as
−1
XT (k)A−1
k X(k) XT (k)A−1k d(k)
wls (k) = (6.16)
tr [Ak ] tr [Ak ]
The first bracketed term on the right-hand side of (6.16) is an estimate of the autocorrelation
matrix R̂xx based on k data samples, that is,
k
R̂xx i, j
(k) = α k−n xi (k)x j (k) (6.17)
n=1
R̂−1
xx (k) = tr [Ak ]P(k) = (1 + α + α + · · · + α
2 k−1
)P(k)
1−α k
= P(k) (6.18)
1−α
The form of the algorithm given by (6.14) and (6.15) requires that the desired signal be
known at each sample point, which is an unrealistic assumption. An estimate of the desired
signal is used in practice, Consequently, replacing d(k + 1) by d(k
ˆ + 1) in (6.15) results
in a practical weighted least square error recursive algorithm.
Other useful forms of (6.15) arise from replacing certain instantaneous quantities by
their known average values [3,7]. To obtain these equivalent forms, rewrite (6.15) as
P(k)
w(k + 1) = w(k) +
α + x T (k + 1)P(k)x(k + 1)
· x(k + 1)d(k + 1) − x(k + 1)y(k + 1) (6.19)
where y(k + 1) is the array output. The product x(k + 1) · d(k + 1) is replaced by its
average value, which is an estimate of the cross-correlation vector r̂xd . Since the estimate
r̂xd does not follow instantaneous fluctuations of x(t) and d(t), it may be expected that
the convergence time would be greater using r̂xd than when using x(k) d(k) as shown in
Chapter 5.
In the event that only the direction of arrival of the desired signal is known and the
desired signal is absent, then the cross-correlation vector rxd (which conveys direction of
arrival information) is known, and the algorithm for updating the weight vector becomes
P(k)
w(k + 1) = w(k) + rxd − x(k + 1)y(k + 1) (6.20)
α+ xT (k + 1)P(k)x(k + 1)
Equation (6.20) may of course also be used when the desired signal is present, but the
rate of convergence is then slower than for (6.19) with the same desired signal present
conditions.
For stationary signal environments α = 1, but this choice is not practical. As long as
0 < α < 1, (6.14) and (6.15) lead to stable numerical procedures. When α = 1, however,
after many iterations the components of P(k) become so small that round-off errors have a
significant impact. To avoid this numerical sensitivity problem, both sides of (6.14) are mul-
tiplied by the factor (k + 1) to yield numerically stable equations, and (6.18) then becomes
R̂−1
xx (k) = kP(k) (6.21)
where 0 ≤ α ≤ 1 and P−1 (k) = XT (k)A−1 k X(k). A closely related alternative data
weighting scheme uses the sample covariance matrix, R̂xx , for summarizing the effect of
old data, so that
R̂xx (k + 1) = α R̂xx (k) + x∗ (k + 1)xT (k + 1) (6.23)
where 0 ≤ α ≤ 1. Inverting both sides of (6.24) yields
−1
1 1 ∗
R̂−1 (k + 1) = R̂ xx (k) + x (k + 1)x T
(k + 1) (6.24)
xx
α α
Applying the matrix inversion lemma to (6.24) then results in
−1 1 −1 R̂−1 ∗
xx (k)x (k + 1)x (k + 1)
T
−1
R̂xx (k + 1) = R̂xx (k) − R̂xx (k) (6.25)
α α + xT (k + 1)R̂−1 ∗
xx (k)x (k + 1)
This is known as the recursive least squares (RLS) algorithm, because it recursively updates
the correlation matrix such that more recent time samples receive a higher weighting than
past samples [19].
As an example, consider a four-element uniform linear array with λ/2 spacing. The
desired signal is incident at 0◦ , and one interference signal is incident at 45◦ with σnoise =
0.01. After K = 25 iterations and α = 0.9, the antenna pattern appears in Figure 6-2
with a directivity of 5.6 dB, which is less than the quiescent pattern directivity of 6 dB.
Figure 6-3 and Figure 6-4 are the RLS weights as a function of time. They converge in
about 20 iterations.
The data weighting represented by (6.26) implies that past data [represented by R̂xx (k)]
is never more important than current data [represented by x∗ (k + 1)xT (k + 1)]. An
−5
−15
−25
−90 −45 0 45 90
q (degrees)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:54 279
2
|wn|
1.5
0.5
0
0 5 10 15 20 25 30
Iteration
0
∠wn(degrees)
−50
−100
−150
−200
0 5 10 15 20 25 30
Iteration
alternative data weighting scheme that permits past data to be regarded either as less
important or more important than current data is to use
The data weighting scheme represented by (6.27) has been successfully employed [using
β = 1/(k +1) so that each sample is then equally weighted] to reject clutter, to compensate
for platform motion, and to compensate for near-field scattering effects in an airborne
moving target indication (AMTI) radar system [8]. Inverting both sides of (6.27) and
applying the matrix inversion lemma results in [9]
1 β
R̂−1
xx (k + 1) = R̂−1 (k) −
(1 − β) xx (1 − β)
R̂−1 ∗ −1
xx (k)x (k + 1) x (k + 1)R̂xx (k)
T
· (6.28)
(1 − β) + β xT (k + 1)R̂−1 ∗
xx (k)x (k + 1)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:54 280
To obtain the updated weight vector R̂−1 xx (k + 1) must be postmultiplied by the beam
steering vector b∗ . The beam steering vector for AMTI radar systems is matched to a
moving target by including the expected relative Doppler and polarization phase factors,
thereby minimizing the effects of main beam clutter due to (for example) stationary targets
[8,10]. Carrying out the postmultiplication of (6.28) by b∗ then yields
1 R̂−1 ∗
xx (k)x (k + 1)
ŵ(k + 1) = ŵ(k) − β x (k + 1)ŵ(k)
T
(1 − β) (1 − β) + β xT (k + 1)R̂−1 ∗
xx (k)x (k + 1)]
(6.29)
where x(n) is an N × 1 data vector describing the outputs of each array element at time
nT and † denotes complex conjugate transpose. A linear constraint for (6.30) is given by
C† w(n) = f (6.31)
where x(i) is a data vector of length N at time iT, then the weighted output power is
given by
! !
ξ = !1/2 X(n)w(n)! (6.34)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:54 281
where denotes the Euclidian norm, and X(n) is the data matrix of (6.33). The term
“Q-R decomposition” is usually employed to describe the problem of obtaining an upper
triangular transformation of an arbitrary matrix, A, by applying an orthogonal matrix,
Q, obtained by a series of Givens rotations as described in Section [Link]. When this
procedure is applied to the n × N matrix 1/2 X(n), the result can be described by
R(n)
Q(n)1/2 X(n) = (6.35)
0
where R(n) is an N × N upper triangular matrix, and 0 is an (n − N ) × N null matrix.
In keeping with notation introduced earlier, we retain U(n) to denote the upper triangular
matrix in (6.35), even though R(n)—called the Cholesky factor—is what the terminology
“Q-R decomposition” refers to. The weight vector that minimizes ξ of (6.34) is then
given by
−1
w(n) = [U† (n)U(n)]−1 C C† [U† (n)U(n)]−1 C f (6.36)
The similarity of equation (6.36) to equation (4.169) is duly noted. The Q-R decomposition
of the input data matrix using Givens rotations enables the weight vector to be obtained
by using back substitution and can be implemented in parallel and systolic structures.
Back substitution is a costly operation to perform in an algorithm and impedes a parallel
implementation, so the inverse Q-R decomposition, which uses the inverse Cholesky factor,
U−1 (n), is used instead.
To develop a recursive implementation of (6.36), a recursive update of the upper
triangular matrix U(n) is implemented [11]
√
U(n) μ U(n − 1)
= T(n) (6.37)
0T x H (n)
where T(n) is an (N + 1) × (N + 1) orthogonal matrix that annihilates the row vector
√
x† (n) by rotating it into μ U(n − 1). By premultiplying both sides of (6.37) by their
respective Hermitian forms and recognizing that T† (n)T(n) = I, it follows that
where
U−1 (n − 1)z(n)
g(n) = √ (6.43)
μ t (n)
Equation (6.44) may easily be verified by forming the product of each side of (6.44) with its
respective complex conjugate transpose and verifying that (6.43) results. There is a major
difference, however, between the orthogonal matrix P(n) in (6.44) and the orthogonal
matrix T(n) in (6.37). The derivation of the orthogonal matrix P(n) is given in [13], and
the result is
z(n) 0
P(n) = (6.45)
1 t (n)
In other words, P(n) is a rotation matrix that successively annihilates the elements of the
vector z(n), starting from the top, by rotating them into the last element at the bottom.
To relate these results to equation (6.36), it will be convenient to define a new N × N
matrix S(n), given by
Since it can easily be shown that S−1 (n) = X† (n)(n)X(n), S−1 (n) is referred to as a
“correlation matrix” of the exponential weighted sensor outputs averaged over n samples.
It is convenient to define (n) = S(n)C and (n) = C† (n) = C† S(n)C. Then we may
rewrite (6.36) as
Substituting S(n) of (6.46) into (6.39) and using the matrix inversion lemma, it follows
that
1 1
S(n) = S(n − 1) − k(n)x† (n)S(n − 1) (6.48)
μ μ
where
1
μ
S(n − 1)x(n)
k(n) = = S(n)x(n). (6.49)
1+ x (n)S(n − 1)x(n)
1 H
μ
Using the previously given definition for (n) and right multiplying both sides of (6.48)
by C yields
1 1
(n) = (n − 1) − k(n)x† (n)(n − 1) (6.50)
μ μ
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:54 283
Likewise, premultiplying both sides of (n) = S(n)C by C† and using the previously
given definition for (n) gives the result
1 1
(n) = (n − 1) − k(n)x† (n)(n − 1) (6.51)
μ μ
Postmultiplying both sides of (6.43) by C then yields
where
Applying the matrix inversion lemma to (6.55) then yields the recursive relation
√
−1 (n) = μ[I + μ q(n)α(n)]−1 (n − 1) (6.56)
where
√
μ−1 (n − 1)α † (n)
q(n) = = μ−1/2 [−1 (n)α † (n)] (6.57)
1 − μα(n)−1 (n − 1)α † (n)
Finally, to apply the previous recursive relationships to the weight vector of (6.47), we
have
√
w(n) = w(n − 1) − μ[g(n) − μ(n)q(n)]α(n)−1 (n − 1)f (6.58)
The recursive relationship of (6.58) can be further simplified by applying the definitions
of (6.40) and (6.42) to obtain
where
√
μ
ρ(n) = k(n) − (n)q(n) (6.60)
t (n)
ξ(n, n − 1) = x† (n)w(n − 1) (6.61)
U−† (n − 1)x(n)
z(n) = √
μ
– Evaluate the rotations that define P(n), which annihilate z(n) and compute the scalar t (n) from
z(n) 0
P(n) =
1 t (n)
– Update the lower triangular matrix U−† (n), and compute the vector g(n) and α(n) = g† (n) C from
where (k/k) denotes a filtered quantity at sample time k based on measurements through
(and including) k, (k/k − 1) denotes a predicted quantity at sample time k based on
measurements through k − 1, and K(k) is the Kalman gain matrix. For complex quantities
(6.65) is rewritten as
Now
ŵopt (k/k − 1) = (k, k − 1)ŵopt (k − 1/k − 1) (6.67)
and the quantity in brackets of (6.66) is the difference between the optimal array output (or
desired reference signal) and the actual array output. The Kalman-type processor based on
the foregoing equations is shown in Figure 6-5 where the Kalman gain vector is given by
P(k/k − 1)x(k)
K(k) = T (6.68)
[x (k)P(k/k − 1)x(k) + σ 2 (k)]
286
Monzingo-7200014
book
v(k)
ISBN : XXXXXXXXXX
Unit wopt (k) d(k) + e (k) ŵopt (k/k) Unit ŵopt (k −1/k − 1)
Desired signal
Σ time xT (k) Σ K(k) Σ time
approximation
delay − delay
R̂−1
xx (k) = kP(k/k) (6.75)
To begin the recursive equations (6.65)–(6.73), the initial values for ŵ(0/0) and
P(0/0) must be specified. It is desirable to select ŵopt (0/0) = E{wopt } and P(0/0) =
E{w(0)wT (0)} [5] where w(0) = ŵ(0/0) − wopt . From Chapter 3, wopt is given by
the Wiener solution.
wopt = R−1
xx (k)rxd (k) (6.76)
In general, the signal statistics represented by the solution (6.76) are unknown, so a different
procedure (discussed subsequently) is employed to initialize the recursive equations. In
the event that a priori environment data are available, then such information is used to
form a refined initial estimate of wopt using (6.76) as well as to construct a dynamical
system model by way of (6.62).
In situations where a priori information concerning the signal environment is not
available, then ŵopt (0) is generally chosen to yield an omnidirectional array beam pattern,
and P(0/0) can merely be set equal to the identity matrix. Furthermore, some means of
selecting a value for the noise statistic σ 2 (k) must be given.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:54 288
As can be seen from (6.65), the bracketed quantity represents the difference between
the actual array output (using the estimated weights) and a noise corrupted version of the
optimal array output derived from an identical array employing the optimal weight vector
on the received signal set. This optimal output signal d(t) is interpreted as a reference
signal approximation for the actual desired signal since the optimal array weights are
designed to provide an MMSE estimate of the desired signal. Such an interpretation of the
optimal array output signal in turn suggests a procedure for selecting a value for the noise
statistic σ 2 (k). Since d(t) is an approximation for the actual desired signal, s(t), one may
write [5]
where η(k) indicates the error in this approximation for the desired signal s(k). Conse-
quently,
If s(k), η(k), and x(k) are all zero mean processes, then E{v(k)} = 0. Furthermore, if the
noise sequence η(k) is not correlated with s(k) and the noise components of x(k), then
The elements of Q represent the degree of uncertainty associated with adopting the sta-
tionary environment assumption represented by using the identity state transition matrix
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:54 289
The practical effect of the previous modification is to prevent the Kalman gains in
K(k) from decaying to values that are too small, so when variations in the environment
occur sufficient importance is attached to the most recent measurements. The estimate
ŵopt then “follows” variations in the actual value of wopt , although the resulting optimal
weight vector estimates are more “noisy” than when the Q matrix was absent.
or
so that
P(k/k)x(k)
K(k) = (6.87)
σ 2 (k)
When we substitute the result expressed by (6.87) into (6.72) there results
P(k/k)x(k) T
P(k/k) = P(k/k − 1) − x (k)P(k/k − 1) (6.88)
σ 2 (k)
or
[P(k/k)x(k)xT (k)P(k/k − 1)]
P(k/k) = P(k/k − 1) − (6.89)
σ 2 (k)
Premultiplying both sides of (6.89) by P−1 (k/k) and postmultiplying both sides of (6.90)
by P−1 (k/k − 1), we see it follows that [also see (6.74)]
x(k)xT (k)
P−1 (k/k) = P−1 (k/k − 1) + (6.90)
σ 2 (k)
Equation (6.90) can be rewritten as
1
P−1 (k/k) = [σ 2 (k)P−1 (k/k − 1) + x(k)xT (k)] (6.91)
σ 2 (k)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:54 290
so that
The recursive relationship expressed by (6.92) can be repeatedly applied beginning with
P−1 (0/−1) to obtain
−1
k
−1
P(k/k) = σ (k) σ (k)P (0/ − 1) +
2 2
x(i)x (i)
T
(6.93)
i=1
For cases where the desired signal approximation is quite good σ̂ 2 (k) ≈ MMSE, and
"k
the diagonal matrix σ̂ 2 (k)P−1 (0/−1) can be neglected in comparison with x(i)xT (i)
i=1
so that
−1
k
P(k/k) ∼
= σ̂ (k) 2
x(i)x (i)
T
(6.94)
i=1
1
k
x(i)xT (i) → Rxx (k) as k → ∞ (6.96)
k i=1
The MSE at the kth sampling instant, ξ 2 (k), can be written as [16]
trace[P(k/k)Rxx (k)] ∼
= σ 2 (k)N k −1 (6.98)
where N is the dimension of x(k) so the MSE at the kth sampling instant becomes
ξ 2 (k) ∼
= MMSE[1 + N k −1 ] (6.99)
The result in (6.99) means that convergence for this Kalman-type algorithm is theoretically
obtained within 2N iterations, which is similar to the convergence results for the DMI
algorithm.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:54 291
⎡ ⎤
z 1 (t − τ1 )
⎢ z 2 (t − τ2 ) ⎥
⎢ ⎥
x(t) = ⎢ .. ⎥ = d(t)1 + n(t) (6.100)
⎣ . ⎦
z N (t − τN )
where 1 = [1, 1, . . . , 1]T , d(t) is the desired reference signal, and n(t) is the vector of
the interference terms after the time delays. Collecting the steered received signal vector
x(t) and its delayed components along the tapped delay line into a single (M + 1)N × 1
FIGURE 6-6
z1 z1 zN Broadband signal
z(t)
aligned array
t1 SCF t2 SCF SCF tN processor.
x1 (t) x2(t) xN(t)
x (t)
Δ Δ Δ
x (t −Δ)
Δ Δ Δ
Δ Δ Δ
x(t −MΔ)
w01 w11 wM
1 w02 w12 wM
2 wN0 wN1 wNM
y (t) = wT (t) ~
x(t)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:54 292
N
y(t) = wT x(t) = d(t) w i0 (6.103)
i=1
N
N
N
= d(t) w i0 + d(t − ) w i1 + L · · · + d(t − M) w iM + wT n(t)
i=1 i=1 i=1
T
N
since wl 1 = w il . If the adaptive weights are constrained according to
i=1
N
1, l = 0
w il = (6.104)
0, l = 1, 2, · · · , k
i=1
if we assume the noise vector has zero mean. The variance of the output signal is then
given by
where 0 is an N × 1 vector with all zero components. If it is desired to minimize the output
noise variance, then the noise variance performance measure can be defined by
It is now desired to minimize (6.109) subject to the constraint (6.104), which can be
rewritten as
⎡ ⎤
1
⎢ 0 ⎥
⎢ ⎥
I1T w = ⎢ .. ⎥ = c (6.110)
⎣ . ⎦
0
The weight vector w that minimizes (6.109) subject to (6.110) can be chosen by using a
vector Lagrange multiplier to form the modified performance measure
1 T
mvm = w Rnn w + λ c − I1T w (6.111)
2
Setting the derivative of mvm with respect to w equal to zero to obtain wmv , requiring
wmv to satisfy (6.110) to evaluate λ, and substituting the resulting value of λ into wmv
gives the minimum variance weight vector solution
−1
wmv = R−1 T −1
nn I1 I1 Rnn I1 c (6.112)
For a signal-aligned array like that of Figure 6-6, it can also be established (as was done
in Chapter 3 for a narrowband processor) that the minimum variance estimator resulting
from the use of (6.112) is identical to the maximum likelihood estimator [18].
The weight vector computation defined by (6.112) requires the measurement and
inversion of the noise autocorrelation matrix Rnn . The noise autocorrelation matrix can
be obtained from the data that also include desired signal terms, and use of a recursive
algorithm will circumvent the necessity of directly inverting Rnn .
The difficulty of measuring the noise autocorrelation matrix when desired signal terms
are present can be avoided by reformulating the optimization problem posed by (6.109)
and (6.110). The minimization of (6.109) subject to the constraint (6.110) is completely
equivalent to the following problem:
I1 w = c (6.114)
The solution of (6.113), (6.114) may be found by once again using Lagrange multipliers
with the result that
wopt = R−1 T −1 −1
xx I1 [I1 Rxx I1 ] c (6.115)
To show the equivalence between the problem (6.113) and (6.114) and the original problem
(6.109) and (6.110), expand the matrix Rxx as follows:
where
⎧⎡ ⎤ ⎫
⎪
⎪ d(t)1 ⎪
⎪
⎪
⎨⎢ d(t − )1 ⎥ ⎪
⎬
⎢ ⎥
Rdd =E ⎢ .. ⎥ [d(t)1d(t − )1, . . . , d(t − M)1] (6.117)
⎪
⎪⎣ . ⎦ ⎪
⎪
⎪
⎩ ⎪
⎭
d(t − M)1
Now the minimization of (6.119) subject to (6.114) must give exactly the same solution
as the optimization problem of (6.109) and (6.110), since E{d 2 (t)} is not a function of w,
and the two problems are therefore completely equivalent.
The received signal correlation matrix at the kth sample time Rxx (k) can be measured
using the exponentially deweighted finite time average
1
k
R̂xx (k) = # $ α k−n x(n)xT (n) (6.120)
"
k
α k−n n=1
n=1
and
⎡ ⎤
α k−1 0 ··· 0
⎢ .. ⎥
⎢ 0 α k−2 .⎥
A(k) = ⎢
⎢ ..
⎥
⎥ (6.122)
⎣ . α 0⎦
0 ··· 0 1
The inverse of I1T P(k + 1)I1 required in (6.124) can be efficiently computed by application
of the matrix inversion lemma to yield
P(k)
wopt (k + 1) = I − − P(k + 1)
α + x (k + 1)P(k)x(k + 1)
T
· x(k + 1)xT (k + 1) I−11 c (6.125)
The minimum variance recursive processor for the narrowband case takes a particu-
larly simple form. It was found in Chapter 3 for this case that
R−1
nn 1
wmv = (6.126)
1 R−1
T
nn 1
The use of (6.126) presents an additional difficulty since the received signal vector x(t)
generally contains signal as well as noise components. This difficulty can be circumvented,
though.
The input signal covariance matrix is given by
Rxx = E{x∗ (t)xT (t)} = E{d 2 (t)}11T + Rnn (6.127)
where E{d 2 (t)} = β, a scalar quantity. Inverting both sides of (6.127) and applying the
matrix inversion lemma yields
βR−1 T −1
nn 11 Rnn
R−1
xx = [β11 + Rnn ]
T −1
= R−1
nn − (6.128)
1 + β1T R−1
nn 1
Exploit (6.21) for the case when α = 1; it then follows from (6.129) and (6.126) that
(k + 1)P(k + 1)1 P(k + 1)1
wmv (k + 1) = = T (6.130)
(k + 1)1T P(k + 1)1 1 P(k + 1)1
where P(k + 1) is given by (6.14). Note that when there is no desired signal, wmv of
(6.130) converges faster than when the desired signal is present, a result already found in
Chapter 5.
J3 /n J1 /n
60° 60°
(−0.787l, 0) x
s/n = 10
l
87
0.7
(0.394l, −0.682l)
an array geometry and signal environment as in Figure 6-7 (which duplicates Figure 4-30
and is repeated here for convenience). This figure depicts a four-element Y array having
d = 0.787λ element spacing with one desired signal at 0◦ and three distinct narrowband
Gaussian jamming signals located at 15◦ , 90◦ , and 165◦ .
The desired signal in each case was taken to be a biphase modulated signal having
a phase angle of either 0◦ or 180◦ with equal probability at each sample. Two signal
environments are considered corresponding to eigenvalue spreads of 16,700 and 2,440.
Figures 6-8 to 6-11 give convergence results for an eigenvalue spread of 16,700, where the
jammer-to-thermal noise ratios are J1 /n = 4000, J2 /n = 400, and J3 /n = 40, for which
the corresponding noise covariance matrix eigenvalues are given by λ1 = 1.67 × 104 ,
λ2 = 1 × 103 , λ3 = 29.0, and λ4 = 1.0. The input signal-to-thermal noise ratio is
s/n = 10 for Figures 6-8 and 6-9, s/n = 0.1 for Figure 6-10, and s/n = 0.025 for
Figure 6-11. The performance of the algorithm in each case is recorded in terms of the
output SNR versus number of iterations, where one weight iteration occurs with each new
independent data sample.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:54 297
processor with
0 eigenvalue spread of
16,700. Input
−5.00 s/n = 10 for which
output SNRopt = 15
−10.00 with algorithm
parameters α = 1,
−15.00 wT (0) = [1, 0, 0, 0],
and P(0) = I.
−20.00
1 2 3 4 5 6 7 8 1E1 2 3 4 5 6 7 8 1E2 2 3 4
Number of iterations
−5.00
−10.00 Optimum
Output SNR (dB)
−15.00
−20.00
−25.00
−30.00
1 2 3 4 5 6 7 8 1E1 2 3 4 5 6 7 8 1E2 2 3 4 5 6 7 8 1E3
Number of iterations
FIGURE 6-10 Output SNR versus number of iterations for Kalman processor with
eigenvalue spread of 16,700. Input s/n = 0.1 for which output SNRopt = 0.15 (−8.24 dB)
with algorithm parameters σ 2 = MMSE, w(0) = 0, and P(0) = I.
eigenvalue spread
of 16,700. Input
−25.00 s/n = 0.025 for
which output
−30.00 SNRopt =
0.038 (−14.2 dB)
with algorithm
−35.00
parameters
σ 2 = MMSE,
−40.00 w(0) = 0, and
1 2 3 4 5 6 7 8 1E1 2 3 4 5 6 7 8 1E2 2 3 4 5 6 7 8 1E3
P(0) = I.
Number of iterations
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:54 298
eigenvalue spread of 0
2440. Input s/n = 10
for which output −5.00
SNRopt =
−10.00
15 (11.76 dB) with
algorithm −15.00
parameters
σ 2 = MMSE, −20.00
w(0) = 0, and
P(0) = I. −25.00
1 2 3 4 5 6 7 8 1E1 2 3 4 5 6 7 8 1E2 2 3 4 5 6 7 8 1E3
Number of iterations
Figures 6-12 and 6-13 give convergence results for an eigenvalue spread of 2,440,
where the jammer-to-thermal noise ratios are J1 /n = 500, J2 /n = 200, and J3 /n = 40,
for which the corresponding noise covariance matrix eigenvalues are given by λ1 =
2.44 × 103 , λ2 = 4.94 × 102 , λ3 = 25.62, and λ4 = 1.0. The input signal-to-thermal noise
ratio is s/n = 0.1 for Figure 6-12 and s/n = 10.0 for Figure 6-13.
The simulation results shown in Figures 6-8 to 6-13 illustrate the following important
properties of the recursive algorithms:
1. Recursive algorithms exhibit fast convergence comparable to that of DMI algorithms,
especially when the output SNRopt is large (approximately five or six iterations when
SNRopt = 15.0 for the examples shown). On comparing Figure 6-8 (for an eigenvalue
spread of 1.67 × 104 ) with Figure 6-13 (for an eigenvalue spread of 2.44 × 103 ), it is
seen that the algorithm convergence speed is insensitive to eigenvalue spread, which
also reflects the similar property exhibited by DMI algorithms.
2. Algorithm convergence is relatively insensitive to the value selected for the parameter
σ 2 (k) in (6.69). On comparing the results obtained in Figure 6-8 where σ 2 = MMSE
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:54 299
with the results obtained in Figure 6-9 where σ 2 = 1, it is seen that virtually the same
number of iterations are required to arrive within 3 dB of SNRopt even with different
initial weight vectors.
3. The convergence speed of the recursive algorithm for the examples simulated here is
slower for small values of SNRopt and faster for large values of SNRopt , as seen in
Figures 6-8, 6-10, and 6-11. This behavior is also exhibited by the DMI algorithm
(and to some degree by all algorithms that do not assume direction-of-arrival informa-
tion). When no direction-of-arrival information is assumed, such information can be
“learned” from a strong desired signal component.
6.7 PROBLEMS
1. The Minimum Variance Weight Vector Solution
(a) Show that setting the derivative of mvm of (6.111) equal to zero yields
wmv = R−1
nn I1 λ
(b) Show that requiring wmv obtained in part (a) to satisfy (6.110) results in
−1
λ = I1T R−1
nn I1 c
(c) Show that substituting the result obtained in part (b) into wmv of part (a) results in (6.112).
2. Equivalence of the Maximum Likelihood and Minimum Variance Estimates [15]. In some
signal reception applications, the desired signal waveform is completely unknown and cannot
be treated as a known waveform or even as a known function of some unknown parameters.
Hence, no a priori assumptions regarding the signal waveform are made, and the waveform is
regarded as an unknown time function that is to be estimated. One way of obtaining an undistorted
estimate of an unknown time function with a signal-aligned array like that of Figure 6-6 is to
employ a maximum likelihood estimator that assumes that the noise components of the received
signal have a multidimensional Gaussian distribution. The likelihood function of the received
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:54 300
where n, m denote distinct sample times, and ρ is the noise covariance matrix that is a matrix
of N × N submatrices corresponding to the various tap points along the tapped delay line.
Differentiate the logarithm of the likelihood function with respect to sn and equate the result to
zero to obtain ŝm , the maximum likelihood estimator for sm . Show that this result corresponds
to the signal estimate obtained from the processor defined by the result in Problem 1.
3. Derivation of Optimum Weights via Lagrange Multipliers [18]. Show that the solution to the
optimization problem posed by (6.113) and (6.114) is given by (6.115).
4. Development using the M.I.L [18]. Show that (6.124) leads to (6.125) by means of the following
steps:
(a) Pre- and postmultiply (6.14) by I1T and I1 , respectively, to obtain
(b) Apply the matrix identity (D.4) of Appendix D to the result obtained in part (a) to show that
−1 −1
I1T P(k + 1)I1 = α I1T P(k)I1 + I−1 −T
1 x(k + 1)x (k + 1)I1
T
1 P(k)x(k + 1)xT (k + 1)
P(k + 1)P−1 (k) = I−
α α + xT (k + 1)P(k)x(k + 1)
R−1
xx 1
1 R−1
T
xx 1
E ψ
P(n) =
ηH h
EU−H (n − 1) η H U−H (n − 1)
√ = U−H (n) and √ = g H (n)
μ μ
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:54 301
(b) Recognizing that P(n) is orthogonal (P(n)P H (n) = I), show that substituting P(n) of part
(a) into the orthogonal condition yields the following relationships:
EE H + H
=I
Eη + h=0
h2 + η H η = 1
(c) Show that substituting (6.75) into the corresponding result of part (a) and defining
U−H (n − 1)x(n)
a(n) = √
μ
results in η = a(n)
t (n)
(d) Show that substituting η from part (c) into the third relationship of part (b) gives
1
h=
t (n)
(e) The unknown vector ψ may be eliminated by using h of part (d) and η of part (c) in the
second relationship of part (b) to give
ψ = −Ea(n)
(f) Form the product of P(n) with the augmented vector [z H (n)1] H and use the partitioning of
part (a) for P(n) to produce
z(n) E z(n)
P(n) =
1 η h 1
(g) Finally, show that substituting the results from parts (c) and (d) into the result of part (f)
yields the desired result
z(n) 0
P(n) =
1 t (n)
7. RLS algorithm. An eight-element uniform array with λ/2 spacing has the desired signal incident
at 0◦ and two interference signals incident at −21◦ and 61◦ . Use the RLS algorithm to place
nulls in the antenna pattern. Assume that σn = 0.01.
8. RLS algorithm. Plot the received signal as a function of iteration for the RLS and LMS algorithms
when a 0 dB desired signal is incident on the eight-element array at 0◦ while 12 dB interference
signals are incident at −21◦ and 61◦ .
6.8 REFERENCES
[1] C. A. Baird, “Recursive Algorithms for Adaptive Arrays,” Final Report, Contract No. F30602-
72-C-0499, Rome Air Development Center, September 1973.
[2] P. E. Mantey and L. J. Griffiths, “Iterative Least-Squares Algorithms for Signal Extraction,”
Second Hawaii International Conference on System Sciences, January 1969, pp. 767–770.
[3] C. A. Baird, “Recursive Processing for Adaptive Arrays,” Proceedings of the Adaptive Antenna
Systems Workshop, March 11–13, 1974, Vol. I, NRL Report 7803, Naval Research Laboratory,
Washington, DC, pp. 163–182.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 18:54 302
CHAPTER
Cascade Preprocessors
7
' $
Chapter Outline
7.1 Nolen Network Preprocessor . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 304
7.2 Interference Cancellation with a Nolen Network Preprocessor . . . . . . . . . . . . . . . . . . . 311
7.3 Gram–Schmidt Orthogonalization Preprocessor . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 315
7.4 Simulation Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 324
7.5 Summary and Conclusions. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 328
7.6 Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 328
7.7 References . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 332
& %
The least mean squares (LMS) and maximum signal-to-noise ratio (SNR) algorithms
converge slowly whenever there is a wide spread in the eigenvalues of the input signal
correlation matrix. A wide eigenvalue spread occurs if the signal environment includes a
very strong source of interference together with other weaker but nevertheless potent in-
terference sources. This condition also happens when two or more very strong interference
sources arrive at the array from closely spaced but not identical directions.
It was shown in Chapter 4 that by appropriately selecting the step size and moving in
suitably chosen directions an accelerated gradient procedure offers marked improvement
in the convergence rate over that obtained with an algorithm that moves in directions deter-
mined by the gradient alone. Another approach for obtaining rapid convergence rescales
the space in which the minimization is taking place by appropriately transforming the input
signal coordinates so that the constant cost contours of the performance surface (repre-
sented by ellipses in Chapter 4) are approximately circular and no eigenvalue spread is
present in the rescaled space. If such a rescaling is done, then in principle it would be possi-
ble to correct all the error components in a single step by choosing an appropriate step size.
With a method called scaled conjugate gradient descent (SCGD) [1], this philosophy
is followed with a procedure that employs a CGD cycle of N iterations and uses the
information gained from this cycle to construct a scaling matrix that yields very rapid
convergence on the next CGD cycle.
The philosophy behind the development of cascade preprocessors is similar to that of
the SCGD method. The cascade preprocessor introduced by White [2,3] overcomes the
problem of (sometimes) slow convergence by reducing the eigenvalue spread of the input
signal correlation matrix by introducing an appropriate transformation (represented by
the preprocessing network). Used in this manner, the cascade network resolves the input
303
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:1 304
signals into their eigenvector components. Equalizing the resolved signals with automatic
gain control (AGC) amplifiers reduces the eigenvalue spread, thereby simplifying the task
of any gradient algorithm.
By modifying the performance measure governing the control of the adaptive ele-
ments in a cascade preprocessor, a cascade network performs the complete task of array
pattern null steering without any need for a gradient type processor [4]. Using a cascade
preprocessor in this manner reduces the complexity and cost of the overall processor and
represents an attractive alternative to conventional gradient approaches.
Finally, we introduce the use of a cascade preprocessor developed by Brennan et al.
[5–7] to achieve adaptive null steering based on the Gram–Schmidt orthogonalization pro-
cedure [8–10]. A Gram–Schmidt cascade preprocessor is simpler than the other cascade
networks discussed and possesses very fast convergence properties that make it a most ap-
pealing practical alternative [11]. The discussion of eigenvector component preprocessing
networks gives perspective to the development and use of cascade preprocessors.
FIGURE 7-1
Five-element
x1 x2 x3 x4 x5
adaptive array
with eigenvector
5 × 5 Eigenvector component
transformation matrix transformation
and eigenvalue
e1 e2 e3 e4 e5 Eigenvector equalization
components network.
Equalization
AGC AGC AGC AGC AGC
network
d1 d2 d3 d4 d5
Adaptive
weights
Array output
FIGURE 7-2
Single-stage
x1 = v11 x2 = v12 x3 = v13 x4 = v14 x5 = v15
lossless, passive,
reflectionless Nolen
Phase Phase Phase Phase transformation
shifter shifter shifter shifter network for N = 5.
Now suppose we want to maximize the output power resulting from the single-level
Nolen network. The output power is expressed as
N
N
P1 = E{|v12 |2 } = E{v12 v12∗ } = an E{vn1 vl1∗ }al∗
n=1 l=1
(7.2)
N
N
= an m 1nl al∗
n=1 l=1
where
is an element
of the correlation matrix of the input signals. To introduce the unitary
constraint n |an |2 = 1 while maximizing P1 , employ the method of Lagrange multipliers
by maximizing the quantity
N
Q = P1 + λ 1 − |an |2 (7.4)
n=1
Setting both ∂ Q/∂u L and ∂ Q/∂v L = 0, the solution for a L must satisfy the following
conditions:
⎫
λ(a L + a L∗ ) = an m 1n L + m 1Ll al∗ ⎪
⎬
n l
(7.7)
λ(a L − a L∗ ) = an m 1n L − m 1Ll al∗ ⎪
⎭
n l
which is a classical eigenvector equation. Nontrivial solutions of (7.8) exist only when λ
has values corresponding to the eigenvalues of the matrix M1 (of which m 1n L is the n Lth
element).
Substituting (7.8) into (7.2) yields
N
P1 = λ an an∗ (7.9)
n=1
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:1 307
In view of the unitary constraint |an |2 = 1, (7.9) becomes
P1 = λ (7.10)
Consequently, the largest value of P1 (which is the maximum power available at the
right-hand output port) is just the largest eigenvalue of the matrix M1 .
When the phase shifters and directional couplers are adjusted to maximize P1 , then
the signal v12 in Figure 7-2 is orthogonal to all the other signals vk2 . That is,
E{v12 vk2∗ } = 0 for k = 2, 3, . . . , N (7.11)
This orthogonality condition reflects the fact that no other unitary combination of v12 and
any of the other vk2 signals delivers more power than is obtained with v12 alone—such a
result would contradict the fact that P1 has been maximized.
The result of (7.11) shows that all the off-diagonal elements in the first row and the
first column of the covariance matrix of the signals vk2 are set equal to zero. Consequently,
the transformation introduced by the single-stage network of Figure 7-2 is step one in
diagonalizing the covariance matrix of the input signals. To complete the diagonalization,
additional networks are introduced as described in the next section.
FIGURE 7-3
Nolen cascade x1 x2 x3 x4 x5
network for
five-element array.
f12 f13 f14 f15 Phase
shifter
Level 1
Directional
coupler v21
y12 y13 y14 y15
v32
y23 y24 y25
f34 f35
Level 3
v43
y34 y35
v44 v45
f45
Level 4
v54
y45
Level 5
e5 e4 e3 e2 e1
power strictly equal to the total input power, but also the eigenvalues of the covariance
matrix are unchanged. If the element G J represents the overall transfer matrix of the first J
stages, then G J is a product of factors in which each factor represents a single stage, that is,
G J = F J · F J −1 · · · F2 · F1 (7.12)
where
⎡ ⎤
ε1
⎢ ε2 ⎥
⎢ ⎥ ⎡ ⎤
⎢ .. ⎥ x1
⎢ . ⎥
⎢ ⎥ ⎢ x2 ⎥
⎢ ε ⎥ ⎢ ⎥
⎢ J ⎥ = GJ ⎢ .. ⎥ (7.13)
⎢ J +1 ⎥ ⎣ . ⎦
⎢v J +1 ⎥
⎢ . ⎥ xN
⎢ . ⎥
⎣ . ⎦
v NJ +1
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:1 309
Furthermore, the second and lower stages of the network have no effect on the signal ε1 .
It follows that the first row of G J is the same as the first row of
G1 = F 1 (7.14)
Likewise, since the third and lower stages of the network have no effect on F2 , the second
row of G J is identical to the second row of
G2 = F2 · F1 (7.15)
Continuing in the foregoing manner, the entire G N matrix is implemented in the Nolen
form using a step-by-step process. Since the first row of G J is identical to the first row
of G1 , the second row of G J is identical to the second row of G2 , . . ., and the (J − 1)st
row of G J is identical to the (J − 1)st row of G J −1 , it follows that the first J − 1 rows
and J − 1 columns of F J are the same as those of an identity matrix. Furthermore, since
there are no phase shifters between the leftmost input port at any stage and the right-hand
output port, the J th diagonal element of F J is real. These constraints in addition to the
unitary constraint on the transfer matrix define the bounds of the element values (phase
shift and directional coupling) of the J th row.
If the output power maximization at each stage of the cascade transformation net-
work is only approximate, the off-diagonal elements of the covariance matrix are not
completely nulled. However, with only a rough approximation, the off-diagonal elements
are at least reduced in amplitude, and, although the equalization network will no longer
exactly equalize the eigenvalues, it reduces the eigenvalue spread.
of the cascade is reached. The eigenvector beams produced by the piecemeal adjustment
procedure are shown by White [3] to be surprisingly close to those obtained by the com-
plete recursive adjustment procedure. If the adjustment of each phase shifter directional
coupler combination is one iteration, then a total of N (N − 1)/2 iterations completes the
piecemeal adjustment procedure.
⎡ ⎤ ⎡ ⎤
x1 ε1
⎢ x2 ⎥ ⎢ ε2 ⎥
⎢ ⎥ ⎢ ⎥
x = ⎢ . ⎥, ε=⎢ . ⎥ (7.27)
⎣ .. ⎦ ⎣ .. ⎦
xN εN
FIGURE 7-4
x1 x2 xN Nolen beamforming
network.
e1
e2
Passive lossless matched
Nolen network
en
eN
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:1 312
The vectors xs and x I represent the signal and interference components, respectively, of
the input signal envelopes. At the nth output port denote the desired signal power by n
and the interference power by n so that
n = [Rss ]nn (7.36)
n = [RII ]nn (7.37)
That is, the output power of interest is the nth diagonal element of the corresponding
covariance matrix.
The goal is to maximize the signal power n and to minimize the interference power
n at the output port n. Since these two objectives conflict, a trade-off between them
is necessary. One approach to this trade-off selects the unitary transfer matrix, G, that
maximizes n = n /n . A more convenient way of attacking the problem is to adopt as
the performance measure
n = n − n (7.38)
where is a fixed scalar constant that reflects the relative importance on minimizing
interference compared with maximizing the signal. If = 0, then the desired signal is
maximized, while if → ∞, then only the interference is minimized.
There is a value = opt for which maximizing n produces exactly the same result
as maximizing the ratio n . The value opt is a function of the environment and is not
ordinarily known in advance. By setting = min where min is the minimum acceptable
signal-to-interference ratio that provides acceptable performance, then maximizing n
ensures that n is maximized under the conditions when it is most needed. If the signal
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:1 313
environment improves, then n also improves, although not quite as much as if n were
maximized directly.
It is useful to define the matrix
Z = G∗ Rss GT − G∗ RII GT (7.39)
The performance measure n is then the nth diagonal element of Z. Furthermore
Z = G∗ MGT (7.40)
where
M = Rss − RII (7.41)
A consequence of (7.41) is that the matrix M (and hence Z) have both positive and negative
eigenvalues. Selecting the transfer matrix, G, that maximizes the nth diagonal element of
Z subject to (7.30) maximizes the element n . Using the method of Lagrangian multipliers
to maximize the element Z nn with the unitary constraint on G and setting the resulting
gradient equal to zero yield the relation
Mn L G Lm = λG mn (7.42)
L
Equation (7.42) is precisely the form of an eigenvector equation. Consequently, the nth
column of GT is the eigenvector of the matrix M corresponding to the largest eigenvalue.
When the nth column of GT is so constructed, then the element Z nn equals this largest
eigenvalue, and all other elements of the nth row and the nth column of Z vanish.
It is desired to obtain a unitary transfer matrix G for which the elements of the nth
row are the elements of the eigenvector of the matrix M corresponding to the maximum
eigenvalue (resulting in maximizing n ). If the first stage of the cascade Nolen network
is adjusted to minimize 1 , if there is no conflict with maximizing n , and if the first
diagonal element of Z is set equal to the most negative eigenvalue of M, then the off-
diagonal elements of the first row and first column of Z will disappear, and the first row of
G will correspond to the appropriate eigenvector. Likewise, proceed to adjust the second
stage to minimize 2 . This second adjustment results in the diagonalization of the second
row and the second column of Z, and the second row of G corresponds to the second
eigenvector. Continue in this manner adjusting in turn to minimize the corresponding
until reaching stage n. At this point (the nth stage) it is desired to maximize n so the
adjustment criterion must be reversed.
The physical significance of the foregoing adjustment procedure is that the upper
stages of the Nolen network are adjusted to maximize the interference and minimize the
desired signal observed at the output of each stage. This process is the same as maximizing
the desired signal and minimizing the interference that proceeds downward to the lower
stages. When the nth stage is reached where the useful output is desired, then the desired
signal should be maximized and the interference minimized so the adjustment criterion is
reversed.
where A is the complex envelope of one waveform having desired signal and interference
components As and A I . Likewise, B represents the complex envelope of a second wave-
form having desired signal and interference components Bs and B I . If we assume that
the desired signal is modulated with pseudo-random phase reversals and that a synchro-
nized key generator is available for demodulation at the receiver, Figure 7-5 shows the
block diagram of a generalized correlator that forms an estimate of A∗ ⊗ B. After passing
the received signal through a synchronized demodulator to obtain the original unspread
signal, narrow passband filters extract the desired signal while band reject filters extract
the interference. The narrow passband outputs are applied to one correlator (the “signal
correlator”) that then forms estimates of E{A∗s Bs }. The band reject outputs likewise are
applied to a second correlator (the “interference correlator”) that then forms estimates of
E{A∗I B I }. A weighted combination of the outputs (signal component weighted by K s and
interference component weighted by K I ) then forms an estimate of the complex quantity
A∗ ⊗ B. The ratio of K s to K I determines the effective value of .
If the correlation product operator ⊗ is taken to include operations on vector quantities,
then we may write
M = x∗ ⊗ xT (7.44)
Z = h∗ ⊗ h T (7.45)
Synchronized
key generator
I +
Interference
Kl lm (A∗ ⊗ B)
correlator Q −
Balanced Kl
B
modulator
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:1 315
The generalized correlator network of Figure 7-5 provides a basis for estimating the
elements of the matrix Z and determining whether the diagonalization of this matrix is
complete.
The value of the arctangent function lying between −π and zero radians minimizes l ,
whereas the value lying between zero and π radians maximizes l . The complete piece-
meal adjustment of a full Nolen cascade network for an N -element array using (7.46)
and (7.47) requires N (N − 1)/2 iterations, making the practical use of a Nolen eigenvec-
tor component cascade processor less attractive when compared with the Gram–Schmidt
cascade described in the next section.
FIGURE 7-6
Gram–Schmidt
x1 x2 x3 x4 x5
orthogonalization
network with
Gram − Schmidt orthogonalization Howells–Applebaum
preprocessor adaptive processor
for accelerated
y1 y2 y3 y4 y5 convergence.
Array output
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:1 316
FIGURE 7-7
Gram–Schmidt x x1 = v11 x2 = v12 x3 = v13 x4 = v14 x5 = v15
transformation to
obtain independent
variables. Level 1
Level 2
Level 3
Level 4
FIGURE 7-8
x1 x2 x3 x4 x5 Gram–Schmidt
orthogonalization for
+ + + +
− − − − five-element array
y1
realized with
u11 u12 u13 u14 Howells–Applebaum
Filter Filter Filter Filter
adaptive loops.
* * * *
y2
+ + +
− − −
* * *
y3
+ +
− −
u33 u34
Filter Filter
* *
y4
+
−
u44
Filter
y5
shown in Figure 7-8. The transformation occurring at each node in Figure 7-7 (using the
weight indexing of Figure 7-8) are expressed as
vnk+1 = vnk − u k(n−1) vkk , k+1≤n ≤ N (7.49)
where N = number of elements in the array. In the steady state, the adaptive weights have
values given by
k∗
v vk
u k(n−1) = k n (7.50)
vkk∗ vkk
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:1 318
where the overbars denote expected values. The transformation represented by (7.48) and
(7.49) is a close analog of the familiar Gram–Schmidt orthogonalization equations. For
an N -element array, N (N − 1)/2 adaptive weights produce the Gram–Schmidt orthogo-
nalization. Since an adaptive Howells–Applebaum maximum SNR processor requires N
adaptive weights, the configuration of Figure 7-6 requires a total of N (N + 1)/2 adaptive
weights.
The transformation of the input signal vector x into a set of independent output
signals y is not unique. Unlike the eigenvector transformation, which yields output signals
in a normal coordinate system in which the maximum power signal is the first output
component, the Gram–Schmidt transformation adopts any component of x as the first
output component and any of the remaining components of x as the signal component vnk
to be transformed by way of (7.49).
The fact that the transformation network of Figure 7-7 yields a set of uncorrelated
output signals suggests that this network functions in the manner of a CSLC system whose
output signal (in the steady state) is uncorrelated with each of the auxiliary channel input
signals. Recall from the discussion of the SNR performance measure in Chapter 3 that
an N − 1 element CSLC is equivalent to an N -element adaptive array with a generalized
signal vector given by
⎡ ⎤
1
⎢0⎥
⎢ ⎥
t = ⎢ .. ⎥ (7.51)
⎣.⎦
0
It is shown in what follows that by selecting x5 of Figure 7-7 as the main beam channel
signal b (so that tT = [0, 0, . . . , 0, 1]) and z = y5 as the output signal, then the cascade
preprocessor yields an output that converges to z = b − wT x. Here x is the auxiliary
channel signal vector, and w is the column vector of auxiliary channel weights for the
equivalent CSLC system of Figure 7-9 that minimizes the output noise power. The CSLC
system of Figure 7-9 and the cascade preprocessor of Figure 7-7 are equivalent in the
+
−
Array
output z
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:1 319
sense that the cascade network output is the same as the CSLC system output once the
optimum solutions are reached. Recall from Chapter 3 that the optimum weight vector for
the CSLC system is given by
wopt = R−1 ∗
x x (x b) (7.52)
If the main beam channel signal is replaced by a locally generated pilot signal p(t), then
wopt = R−1 ∗
x x (x p) (7.53)
FIGURE 7-10
Conventional
x1 x2 x3 = b three-element CSLC
system having two
Howells–Applebaum
SLC control loops.
SLC SLC
w1 w2
+
−
Σ
z z
z
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:1 320
FIGURE 7-11
Cascade x1 x2 x3 = b
Gram–Schmidt
three-element CSLC
system having three
Howells–Applebaum + +
SLC control loops. SLC − SLC −
x1 = y1 u11 u12
v22 = y2 v23
+
SLC −
u22
v33 = y3 = z
so that
For the cascade system of Figure 7-11, it follows from (7.49) and (7.50) that
y2 = x2 − u 11 x1 (7.58)
v32 = b − u 12 x1 (7.59)
z = v32 − u 22 y2 (7.60)
(x1∗ x2 )
u 11 = (7.61)
(x1∗ x1 )
(x ∗ b)
u 12 = ∗1 (7.62)
(x1 x1 )
(y2∗ v32 )
u 22 = (7.63)
(y2∗ y2 )
Comparing (7.64) with Figure 7-10, we see that w 2 and w 1 fill the roles taken by u 22 and
u 12 − u 11 u 22 , respectively, in the cascade system.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:1 321
The jth row of the transformation matrix G contains the desired weight equivalences for
a j-element CSLC system where the auxiliary channel weights are given by
⎫
w j−1 = −g j ( j−1) ⎪
⎪
w j−2 = −g j ( j−2) ⎪⎬
.. ⎪
(7.71)
. ⎪
⎪
⎭
w 1 = −g j1
Having found the steady-state weight element equivalence relationships, it is now
appropriate to consider the transient response of the networks of Figures 7-10 and 7-11.
The transient responses of the two CSLC systems under consideration are investigated
by examining the system response to discrete signal samples of the main beam and the
auxiliary channel outputs. The following analysis assumes that the adaptive weights reach
their steady-state expected values on each iteration and therefore ignores errors that would
be present due to loop noise. Let xkn denote the nth signal sample for the kth element
channel. The main beam channel samples are denoted by bn . For a system having N
auxiliary channels and one main beam channel, after N independent samples have been
collected, a set of auxiliary channel weights are computed using
N
w k xkn = bn for n = 1, 2, . . . , N (7.72)
k=1
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:1 322
The foregoing system of equations yields a unique solution since the matrix defined
by the signal samples xkn is nonsingular provided that either thermal receiver noise
or N directional interference sources are present. Equation (7.72) is closely related to
∗
(7.52), since multiplying both sides of (7.72) by xmn and summing over the index n
yields
N
N
N
∗ ∗
w k xmn xkn = xmn bn (7.73)
n=1 k=1 n=1
Comparing (7.74) with (7.52) reveals that these two equations yield similar solutions for
the weight vector w in a stationary signal environment.
The iterative weight correction procedure for a single Howells–Applebaum SLC loop
is modeled as shown in Figure 7-12. On receipt of the ith signal sample, the resulting
change in the weight u k(n−1) is computed by applying (7.50) and assuming that the signal
samples are approximately equal to their expected values to yield
vkk∗ [vnk (i) − u k(n−1) (i − 1) · vkk (i)]
u k(n−1) (i) = (7.75)
(vkk∗ vkk )
The element weight value is then updated in accordance with
z−1
uk(n−1)(i)
Δuk(n−1)(i)
With all weights in the cascade Gram–Schmidt network initially set to zero, it follows that
on receipt of the first signal sample v22 = x21 , and v32 = v33 = b1 so the change in weight
settings after the first iteration results in
∗ 1 ∗
x11 v2 x11 x21
u 11 (1) = = (7.77)
|x11 | 2 |x11 |2
∗ 1
x v x ∗ b1
u 12 (1) = 11 32 = 11 2 (7.78)
|x11 | |x11 |
and
∗
v22∗ v32 x21 b1
u 22 (1) = = (7.79)
|v2 |
2 2
|x21 |2
On receipt of the second signal sample v22 = x22 − u 11 (1)x12 , v32 = b2 − u 12 (1)x12 ,
and v33 = v32 − u 22 (1)v22 so that
∗
x12 [x22 − u 11 (1)x12 ]
u 11 (2) = (7.80)
|x12 |2
x ∗ [b2 − u 12 (1)x12 ]
u 12 (2) = 12 (7.81)
|x12 |2
(x22 − u 11 (1)x12 ) ∗ [b2 − u 12 (1)x12 − u 22 (1)(x22 − u 11 (1)x12 )]
u 22 (2) = (7.82)
|x22 − u 11 (1)x12 |2
Substitute (7.77) to (7.79) into (7.80) to (7.82), it then follows that
∗
x12 x22
u 11 (2) = u 11 (1) + u 11 (2) = (7.83)
|x12 |2
x ∗ b2
u 12 (2) = u 12 (1) + u 12 (2) = 12 2 (7.84)
|x12 |
x11 b2 − x12 b1
u 22 (2) = u 22 (1) + u 22 (2) = (7.85)
x11 x22 − x12 x21
By use of the weight equivalence relationships of (7.67) it follows that w 2 (2) = u 22 (2)
and
x22 b1 − x21 b2
w 1 (2) = u 12 (2) − u 11 (2)u 22 (2) = (7.86)
x11 x22 − x12 x21
Note, however, that w 1 (2) and w 2 (2) of (7.85) and (7.86) are the weights that satisfy (7.72)
when N = 2. Therefore, for a two-auxiliary channel CSLC, the cascade Gram–Schmidt
network converges to a set of near-optimum weights in only two iterations. The foregoing
analysis can also be carried out for the case of an N -auxiliary channel CSLC and leads to
the conclusion that convergence to a set of near-optimum weights in the noise-free case
occurs after N iterations. Averaging the input signals mitigates the effects of noise-induced
errors.
Note that using (7.75) to update the weight settings results in values for u 11 and u 12 (the
weights in the first level) that depend only on the current signal sample. The update setting
for u 22 , however, depends on the last two signal samples. Therefore, an N -level cascade
Gram–Schmidt network always updates the weight settings based on the N preceding
signal samples.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:1 324
cascade
preprocessor with
0.00
eigenvalue spread =
−5.00 16,700. Algorithm
parameters are K =
−10.00 3, α = α L = 0.1,
and w(0) = 0 with
−15.00 s/n = 10 for which
SNRopt = 15.
−20.00
1 2 3 4 5 6 78 1E1 2 3 4 5 6 78 1E2 2 3 4 5 6 78 1E3
Number of iterations
cascade
preprocessor with
0.00
eigenvalue spread =
−5.00 16,700. Algorithm
parameters are
−10.00 K = 9, α = 0.3, α L =
0.25α, and w(0) = 0
−15.00 with s/n = 10 for
which SNRopt = 15.
−20.00
1 2 3 4 5 6 78 1E1 2 3 4 5 6 78 1E2 2 3 4 5 6 78 1E3
Number of iterations
cascade
preprocessor with
0.00
eigenvalue spread =
16,700. Algorithm −5.00
parameters are
K = 9, α = α L = −10.00
0.65, and w(0) = 0
with s/n = 10 for −15.00
which SNRopt = 15.
−20.00
1 2 3 4 5 6 78 1E1 2 3 4 5 6 78 1E2 2 3 4 5 6 78 1E3
Number of iterations
cascade
0.00 preprocessor with
−5.00 eigenvalue spread =
2,440. Algorithm
−10.00 parameters are
−15.00
K = 3, α = α L =
0.1, and w(0) = 0
−20.00 with s/n = 10 for
which SNRopt = 15.
−25.00
1 2 3 4 5678 1E1 2 3 4 5678 1E2 2 3 4 5678 1E3 2 3 4 5
Number of iterations
3. Only a slight increase in the loop gain α causes the weights to become excessively
noisy (compare Figure 7-16 with Figure 7-13 and Figure 7-17).
4. The appropriate value to which the product K · α should be set for an acceptable level of
loop noise depends on the value of SNRopt (compare Figures 7-19 and 7-20). Smaller
values of SNRopt require smaller values of the product K · α to maintain the same
loop noise level. It may also be seen that the adaptation time required with these
parameter values for the Gram–Schmidt orthogonalization preprocessor is greater than
that required for either a recursive or direct matrix inversion (DMI) algorithm. The
question of relative adaptation times for the various algorithms is pursued further in
Chapter 10.
5. The degree of transient response improvement that are obtained with a cascade pre-
processor is shown in Figure 7-18 where the Gram–Schmidt cascade preprocessor
response is compared with the LMS processor response. The LMS curve in this fig-
ure was obtained simply by removing the cascade preprocessing stage in Figure 7-6,
thereby leaving the final maximum SNR stage consisting of four LMS adaptive loops.
The greater the eigenvalue spread of the Rxx matrix, then the greater is the degree of
improvement that are realized with a cascade preprocessor compared with the LMS
processor.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:1 328
7.6 PROBLEMS
Problems 1 through 12 all concern Cascade SLC Control Loops
1. Consider the cascade control loop configuration of Figure 7-21.
(a) Show that the weight equivalences between the configuration of Figure 7-21 and the
standard CSLC configuration are given by
w1 = u1 + u3 w 2 = −u 2 u 3
(b) Show that the steady-state weight values of Figure 7-21 are given by
x1∗ b x2∗ x1
u1 = u 2 =
x1∗ x1 x ∗ x2
∗ 2 ∗
(x1 x2 ) (x1 x1 )(x2 b) − (x2∗ x1 )(x1∗ b)
∗
u3 = ∗
(x1 x1 ) (x2∗ x2 )(x1∗ x1 ) − (x1∗ x2 )(x2∗ x1 )
FIGURE 7-21
Cascade x2 x1 b
arrangement of three
SLC +
Howells–Applebaum −
SLC control weight
u1
loops yielding SLC +
−
convergence to weight
incorrect solution. u2
SLC +
−
weight
u3
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:1 329
Note that these steady-state weights do not correspond to the correct solution for the
standard CSLC configuration.
Note: Problems 2 through 12 concern derivations that may be found in [7].
2. Consider the three-element array cascade configuration of Figure 7-11.
(a) Show that the steady-state weight values for this configuration are given by (7.61), (7.62),
and (7.65).
(b) Using (7.64) for the array output, show that
(c) The average noise output power is defined by N0 = |z|2 = zz ∗ . For notational simplicity
let u 1 = u 11 , u 2 = u 12 , u 3 = u 22 , and show that
3. Using the weight notation of Problem 2(c) for the configuration of Figure 7-11, let
u n = u n + δn
where δn represents the fluctuation component of the weight. The total output noise power is
then given by
N0TOT = N0 + Nu
= |x3 − (u 3 + δ3 )x2 − [(u 2 + δ2 ) − (u 1 + δ1 )(u 2 + δ3 )]x1 |2
where Nu represents the excess weight noise due to the fluctuation components. Note that
δ n = 0 and neglect third- and fourth-order terms in δn to show that
where G 1 denotes amplifier gain, and τ1 denotes time constant of the integrating filter in the
Howells–Applebaum control loop for u 1 . The foregoing equation is rewritten as
1 1 G1
u̇ 1 + x1∗ x1 + u 1 = x1∗ x2 where α1 =
α1 G1 τ1
so that
1 1
u̇ 1 + x1∗ x1 + u 1 = x1∗ x2
α1 G1
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:1 330
Subtracting the mean value differential equation from the instantaneous value differential
equation and recalling that u 1 = u 1 + δ1 then yields
1 1
δ̇1 + + x1∗ x1 δ1 + (x1∗ x1 − x1∗ x1 )u 1 = (x1∗ x2 − x1∗ x2 )
α1 G1
provided that the second-order term δ1 (x1∗ x1 − x1∗ x1 ) is ignored. In the steady state where
x1∗ x2
u1 =
x1∗ x1
show that
1 1 (x1∗ x2 ) ∗ 1
δ̇1 + x1∗ x1 + δ1 = x1∗ x2 − x1 x1 = f 1 (t)
α1 G1 (x1∗ x1 ) α1
so that
1
δ̇1 (t) = f 1 (t) − α1 x1∗ x1 + δ1 (t)
G1
Substitute δ̇1 (t) into the differential equation for δ̇1 (t) obtained in Problem 4, and show that this
reduces the resulting expression to an identity, thereby proving that δ1 (t) is indeed a solution
to the original differential equation.
6. The second moment of the fluctuation δ1 is given by δ1∗ δ1 .
(a) Use the expression for δ1 (t) developed in Problem 5 to show that
t t
1 1
δ1∗ δ1 ∗
= exp −2α1 x1 x1 + ∗ ∗
f 1 (τ ) f 1 (u) exp α1 x1 x1 + (τ + u) dτ du
G1 0 0 G1
has a correlation interval denoted by ε, such that values of f 1 (t) separated by more than ε
are independent. Let
t
f 1∗ (τ ) f 1 (u) du = f 1∗ f 1 ε for t > τ + ε.
0
ε f 1∗ f 1
δ1∗ δ1 = for t > τ + ε
2α1 [(x1∗ x1 ) + 1/G 1 ]
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:1 331
8. Substitute f 1∗ f 1 obtained in Problem 7 into δ1∗ δ1 obtained in Problem 6 (b) to show that
Note that to obtain δ2∗ δ2 , merely replace x2 with x3 , α1 by α2 , and G 1 by G 2 in the previous
expression.
9. From Figure 7-11 it may be seen that the inputs to u 3 (u 22 of the figure) are y1 and y2 . Conse-
quently, replace x1 by y2 , x2 by y1 , and G 1 by G 3 in the expression obtained in Problem 8. This
replacement requires the second moments y1∗ y1 and y2∗ y2 and y2∗ y1 to be computed. Show that
(x1∗ x3 )(x3∗ x1 )
y1∗ y1 = |x3 − u 3 x1 |2 = (x3∗ x3 ) −
(x1∗ x1 )
and
2
(x1∗ x2 )
y2∗ y2 = |x2 − u 1 x1 | = x2 − ∗ 2
x1
(x1 x1 )
(x1∗ x2 )(x2∗ x1 )
= (x2∗ x2 ) −
(x1∗ x1 )
Furthermore
∗
(x1∗ x2 ) (x ∗ x3 )
y2∗ y1 = x2 − ∗
x1 x3 − 1∗ x1
(x1 x1 ) (x1 x1 )
∗ ∗
(x x1 )(x x3 )
= (x2∗ x3 ) − 2 ∗ 1
(x1 x1 )
10. The computation of Nu found in Problem 3 requires the second moment δ2∗ δ1 . Corresponding
to the expression for δ1 (t) in Problem 5 we may write
t
1
δ2 (t) = f 2 (τ ) exp − α2 x1∗ x1 + (t − τ )dτ
0 G2
where
(x1∗ x3 ) ∗
f 2 (t) = α2 x1∗ x3 − x1 x1
(x1∗ x1 )
With the previous results, show that
t t
δ2∗ δ1 = f 2∗ (τ ) f 1 (u)
0
0
∗ 1 ∗ 1
· exp −α2 x1 x1 + (t − τ ) − α1 x1 x1 + (t − u)dτ du
G2 G1
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:1 332
show that
ε f 2∗ f 1
δ2∗ δ1 =
[α2 (x1∗ x1 ) + 1/G 2 ] + α1 [(x1∗ x1 ) + 1/G 1 ]
11. Evaluate f 2∗ f 1 by retaining only terms of the form xm∗ xn in the expansion
(x ∗ x1 ) (x ∗ x2 )
f 2∗ f 1 = α2 α1 x1 x3∗ − 3∗ x1∗ x1 x1∗ x2 − 1∗ x1∗ x1
(x1 x1 ) (x1 x1 )
then the previous expression for Nu shows that the control loop noise contribution due to
fluctuations in u 3 is just equal to the sum of the contributions from fluctuations in u 1 and u 2 .
7.7 REFERENCES
[1] L. Hasdorff, Gradient Optimization and Nonlinear Control, New York, Wiley, 1976, Ch. 3.
[2] W. D. White, “Accelerated Convergence Techniques,” Proceedings of the 1974 Adaptive
Antenna Systems Workshop, March 11–13, Vol. 1, Naval Research Laboratory, Washington,
DC, pp. 171–215.
[3] W. D. White, “Cascade Preprocessors for Adaptive Antennas,” IEEE Trans. Antennas Propag.,
Vol. AP-24, No. 5, September 1976, pp. 670–684.
[4] W. D. White, “Adaptive Cascade Networks for Deep Nulling,” IEEE Trans. Antennas Propag.,
Vol. AP-26, No. 3, May 1978, pp. 396–402.
[5] L. E. Brennan, J. D. Mallett, and I. S. Reed, “Convergence Rate in Adaptive Arrays,” Tech-
nology Service Corporation Report No. TSC-PD-A177-2, July 15, 1977.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:1 333
[6] R. C. Davis, “Convergence Rate in Adaptive Arrays,” Technology Service Corporation Report
No. TSC-PD-A177-3, October 18, 1977.
[7] L. E. Brennan and I. S. Reed, “Convergence Rate in Adaptive Arrays,” Technology Service
Corporation Report No. TSC-PD-177-4, January 13, 1978.
[8] C. C. Ko, “On the Performance of Adaptive Gram–Schmidt Algorithm for Interference Can-
celling Arrays,” IEEE Transactions on Antennas and Propagation, Vol. 39, No. 4, 1991,
pp. 505–511.
[9] H. Liu, A. Ghafoor, and P. H. Stockmann, “Application of Gram–Schmidt Algorithm to Fully
Adaptive Arrays,” IEEE Transactions on Aerospace and Electronic Systems, Vol. 28, No. 2,
1992, pp. 324–334.
[10] R. W. Jenkins and K. W. Moreland, “A Comparison of the Eigenvector Weighting and Gram–
Schmidt Adaptive Antenna Techniques,” IEEE Transactions on Aerospace and Electronic
Systems, Vol. 29, No. 2, 1993, pp. 568–575.
[11] C. C. Ko, “Simplified Gram–Schmidt Preprocessor for Broadband Tapped Delay-Line
Adaptive Array,” IEE Proceedings G Circuits, Devices and Systems, Vol. 136, No. 3, 1989,
pp. 141–149.
[12] W. F. Gabriel, “Adaptive Array Constraint Optimization,” Program and Digest 1972 G-AP
International Symposium, December 11–14, 1972, pp. 4–7.
[13] J. C. Nolen, “Synthesis of Multiple Beam Networks for Arbitrary Illuminations,” Bendix
Corporation, Radio Division, Baltimore, MD, April 21, 1960.
[14] N.J.G. Fonseca, “Printed S-Band 4 4 Nolen Matrix for Multiple Beam Antenna Applications,”
IEEE Transactions on Antennas and Propagation, Vol. 57, No. 6, June 2009, pp. 1673–1678.
[15] P. R. Halmos, Finite Dimensional Vector Spaces, Princeton, NJ, Princeton University Press,
1948, p. 98.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:1 334
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:10 335
CHAPTER
Gradient-based algorithms assume the performance measures are either quadratic or uni-
modal. For some classes of problems [1–3], the mathematical relation of the variable
parameters to the performance measure is either unknown or is too complex to be useful.
In yet other problems, constraints are placed on the variable parameters of the adap-
tive controller with the result that the performance surface is no longer unimodal. When
the performance surface of interest is multimodal and contains saddlepoints, then any
gradient-based algorithms finds only a local minimum. A random search algorithm has
the ability to jump out of one valley with a local minimum into another valley with a
potentially lower local minimum. Random algorithms have global search capabilities that
work for any computable performance measure [4–14]. Random search algorithms tend to
have slow convergence, especially in unimodal applications. They do, however, have the
advantages of being simple to implement in logical form, of requiring little computation,
of being insensitive to discontinuities, and of exhibiting a high degree of efficiency where
little is known about the performance surface.
Systematic searches exhaustively survey the parameter space within specified bounds,
making them capable of finding the global extremum of a multimodal performance mea-
sure. As a practical matter, however, this type of search is very time-consuming and incurs
a high search loss, since most of the search period occurs in regions of poor performance.
Random searches are classified as either guided or unguided, depending on whether
information is retained whenever the outcome of a trial step is learned. Furthermore, both
the guided and unguided varieties of random search are given accelerated convergence
by increasing the adopted step size in a successful search direction. Four representative
examples of random search algorithms used for adaptive array applications are considered
335
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:10 336
in this chapter: linear random search (LRS), accelerated random search (ARS), guided
accelerated random search (GARS), and genetic algorithm (GA).
where [·] denotes the selected array performance measure, and μs is a step size constant.
The random vector wk has components generated from a normal probability density
function with zero mean and variance σ 2 . The constants μs and σ 2 are selected to ensure
a fast, stable algorithm convergence. The LRS algorithm is “linear” because the weight
change is proportional to the change in the performance measure.
The true change in the performance measure resulting from adding wk to wk is
()k = [wk + wk ] − [wk ] (8.2)
When the performance measure value is estimated, then the corresponding estimated
change in the performance measure is given by
ˆ = ˆ ˆ
()k [wk + wk ] − [wk] (8.3)
The alternate (but equivalent) form of (8.1) represented by (8.15) is more useful for analysis
even though the algorithm is implemented in the form suggested by (8.1). Equation (8.15)
emphasizes that the adaptive weight vector is regarded as the solution of a first-order linear
vector difference equation having a randomly varying coefficient I − 2μs wk wkT Rxx
and a random driving function μs γk wk .
Premultiplying both sides of (8.15) by the transformation matrix Q of Section 4.1.3
converts the foregoing linear vector difference equation into normal coordinates
vk+1 = I − 2μs wk wT
k vk + μs γk wk
(8.16)
Although (8.16) is somewhat simpler than (8.15), the matrix coefficient of vk still contains
cross-coupling and randomness, thereby rendering (8.16) a difficult equation to solve.
Stability conditions for the LRS algorithm are obtained without an explicit solution to
(8.16) by considering the behavior of the adaptive weight vector mean.
Taking the expected value of both sides of (8.16) and recognizing that wk is a random
vector that is uncorrelated with γk and vk , we find that
E{vk+1 } = E I − 2μs wk wT
k vk + μs E{γk wk }
= I − 2μs E wk wT
k E{vk } + 0
= (I − 2μs σ 2 )E{vk } (8.17)
For the initial conditions v0 , (8.18) gives the expected value of the weight vector’s transient
response. If (8.18) is stable, then the mean of vk must converge. The stability condition
for (8.18) is
1
> μs σ 2 > 0 (8.19)
λmax
If we choose μs σ 2 to satisfy (8.19), it then follows that
Since the foregoing transient behavior is analogous to that of the method of steepest
descent discussed in Section 4.1.2, it is argued by analogy that the time constant of the
pth mode of the expected value of the weight vector is given by
1
τp = (8.21)
2μs σ 2 λ p
Furthermore, the time constant of the pth mode of the MSE learning curve is one-half the
aforementioned value so that
1
τ pmse = (8.22)
4μs σ 2 λ p
Satisfying the stability condition (8.19) implies only that the mean of the adaptive
weight vector will converge according to (8.20); variations in the weight vector about
the mean value may be quite severe, however. It is therefore of interest to obtain an
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:10 339
indication of the severity of variations in the weight vector when using the LRS algorithm
by deriving an expression for the covariance of the weight vector. In obtaining such an
expression, it will simply be assumed that the weight vector covariance is bounded and
that the weight vector behaves as a stationary stochastic process after initial transients
have died out. Assuming that a bounded steady-state covariance matrix exists, we may
calculate an expression for such a covariance by multiplying both sides of (8.16) by their
respective transposes to obtain
vk+1 vT T T
k+1 = I − 2μs wk wk vk vk I − 2μs wk wk
T
Now we take expected values of both sides of (8.23) recalling that γk and wk are zero-
mean uncorrelated stationary processes so that
E vk+1 vT
k+1 = E I − 2μs wk wT T
k vk vk I − 2μs wk wk
T
+μ2s E γk2 E wk wT k +0 (8.24)
Since var[γk ] ∼
= (4/K )ξmin
2
and cov[wk ] = σ 2 I, it follows that (8.24) is expressed
as
E vk+1 vT
k+1 = E I − 2μs wk wT
k
4 2 2
·vk vT
k I − 2μs wk wk
T
+ μ2s ξmin σ I (8.25)
K
In the steady state vk is also a zero-mean stationary random process that is uncorrelated
with wk so (8.25) is written as
T
E vk+1 vT
k+1 = E I − 2μs wk wT
k E vk vk I − 2μs wk wT
k (8.26)
4
+μ2s ξmin 2
σ 2I
K
Consequently, the steady-state covariance of the adaptive weight vector is
4 2 2
cov[vk ] = E I − 2μs wk wT
k cov[vk ] I − 2μs wk wk
T
+ μ2s ξmin σ I
N
= cov[vk ] − 2μs E wk wTk cov[vk ]
− 2μs cov [vk ]E wk wT k
4 2 2
+ 4μs E wk wk cov[vk ]wk wT
2 T
k + μ2s ξmin σ I
N
= cov[vk ] − 2μs σ cov[vk ] − 2μs σ cov[vk ]
2 2
4 2 2
+ 4μ2s E wk wT
k cov[vk ]wk wk
T
+ μ2s ξmin σ I (8.27)
K
Equation (8.27) is not easily solved for the covariance of vk , because the matrices
appearing in the equation cannot be factored. It is likely (although not proven) that the
steady-state covariance matrix of vk is diagonal. The results obtained with such a simplify-
ing assumption do, however, indicate that there is some merit in the plausibility argument.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:10 340
The random driving function appearing in (8.16) consists of components that are
uncorrelated with each other and uncorrelated over time. Furthermore, the random coeffi-
cient I − 2μs wk wTk is diagonal on the average (though generally not for every value
of k) and uncorrelated both with vk and with itself over time. Thus, it is plausible that the
covariance of vk is a diagonal matrix.
Assuming that cov[vk ] is in fact diagonal, then (8.27) can immediately be rewritten
by merely rearranging terms as
4 2 2
4μs σ 2 cov[vk ] − 4μ2s E wk wT T
k cov[vk ]wk wk = μ2s ξmin σ I (8.28)
K
Of greatest interest is the case for which adaptation is slow, in which
μs σ 2 I (8.29)
Furthermore, it may be noted that
T ∼
μ2s E wk wT
k cov[vk ]wk wk = (μs σ 2 )2 cov[vk ] (8.30)
and from (8.29) it follows that
(μs σ 2 )2 cov[vk ] μs σ 2 cov[vk ] (8.31)
With the result of (8.31) it follows that the term −4μ2s E{ } appearing in (8.28) is neglected.
Consequently, (8.28) is rewritten as
μs 2 −1
cov[vk ] = ξ (8.32)
K min
The steady-state covariance matrix of vk given by (8.32) is based on a plausible assumption,
but experience indicates that the predicted misadjustment obtained using this quantity
generally yields accurate results [15].
The misadjustment experienced using the LRS algorithm is obtained by considering
the average excess MSE due to noise in the weight vector, which is given by
N
E vT
k vk = λ p E{(v pk )2 } (8.33)
p=1
where N is the number of eigenvalues of . If we use (8.32), it follows that for the LRS
algorithm
N
μs 2 1 N μs 2
E vT
k vk = λp ξ = ξ (8.34)
p=1
K min λ p K min
Since the misadjustment M is defined to be the average excess MSE divided by the
minimum MSE
T
E vk vk
M= (8.35)
ξmin
It follows that for the LRS algorithm
N μs
M= ξmin (8.36)
K
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:10 341
The result given by (8.36) can be expressed in terms of the perturbation of the LRS process
as
N μs σ 2 tr(Rxx ) N 2 μs σ 2 λav
M= = (8.37)
2KP 2KP
Now recall that the time constant of the pth mode of the learning curve for the LRS
algorithm (in terms of the number of iterations required) is given by (8.21). Since one
iteration of the weight vector requires two estimates of ξ̂ , 2K samples of data are used per
iteration, and the learning curve time constant expressed in terms of the number of data
samples is
K
T pmse = 2K τ pmse = (8.38)
2μs σ 2 λ p
From (8.38) it follows immediately that
K 1
λp = (8.39)
2μs σ 2 T pmse
and
K 1
λav = (8.40)
2μs σ 2 T pmse av
If the deterministic component of the total misadjustment is optimally chosen, then both
M and P are equal and P is one-half the total misadjustment so that
1/2
N2 1 1
(Mtot )min = =N (8.43)
2Popt T pmse av T pmse av
It is informative to compare this result with the corresponding result (4.83) for the LMS
algorithm.
where μs (k) is the step size initially set at μs (0) = μ0 , and w(k) is a random vector
whose components are given by
where θi is a uniformly distributed random angle on the interval {0, 2π} so that
|w i (k)| = 1 and w(k) controls the direction of the weight vector change while μs
controls the magnitude.
Initially the weight vector, w(0), and the corresponding performance measure, [w(0)],
ˆ
(or an estimate thereof [w(0)]) is evaluated. The weight vector changes in accordance
with (8.44) using μs (0) = μ0 . The performance index [w(1)] is evaluated and com-
pared with [w(0)]. If this comparison indicates improved performance, then the weight
direction change vector w is retained, and the step size μs is doubled (resulting in
“accelerated” convergence). If, however, the resulting performance is not improved, then
the previous value of w is retained as the starting point, a new value of w is selected,
and μs is reset to μ0 . As a consequence of always returning to the previous value of w as
the starting point for a new weight perturbation in the event the performance measure is
not improved, the ARS approach is inherently stable, and stability considerations do not
play a role in step size selection. A block diagram of this simplified version of accelerated
random search is given in Figure 8.1.
Consider a single component of the complex weight vector for which vi = w i − w opti .
If vi (k) lies within μ0 /2 of w opt , then any further perturbation in that component of the
weight vector of step size μ0 will result in vi (k + 1) ≥ vi (k) as shown in Figure 8.2.
Consequently, if all components of the weight vector lie within or on the best performance
surface contour contained within the circle of radius μ0 /2 about w opt , then no further
improvement in the performance measure can possibly occur using step size μ0 . The
condition where all weight vector components lie within this best performance surface
contour therefore represents a lower limit on the possible improvement that is achieved
with the ARS procedure, and this is the ultimate condition to which the weight vector is
driven in the steady state.
FIGURE 8-1
Block diagram for
ARS algorithm. Calculate
wi(k+1) = wi(k) + ms(k +1) Δwi(k+1) Array weighting
i = 1, 2, …, N
Array output
Reset w = w(k)
≥0 b [w(k+1)] − <0 Δw(k+1) = Δw(k)
choose qi, i = 1, 2, …, N
b [w(k)] ms(k+1) = 2ms(k)
qi uniform on {0, 2p}
Random Deterministic
phase phase
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:10 343
0
m
3
Re(vi)
wopt m
0
2
m0
To simplify the development, assume it is equally likely that the weight vector com-
ponent w i (k) lies anywhere within or on the circle for which vi = μ0 /2, it then follows
that the steady-state expected value of vi is vss (k) = 0, and the average excess MSE in
this steady-state condition is (adopting = ξ and noting that E{|vi |2 } = μ20 /8)
μ20
E{vT (k)v (k)} = ξmin tr(Rxx ) (8.46)
8
The average misadjustment for this steady-state condition is therefore
μ20
Mav = tr(Rxx ) (8.47)
8
On each successive iteration the weight vector components are perturbed by μ0 from
their steady-state values. Assume that the perturbation is taken from vss = ρ as shown in
Figure 8.3, then E{v p } = μ0 , E{|v p |2 } = 9μ20 /8, and the average total misadjustment for
the random search perturbation is therefore
9μ20
Mtot = tr(Rxx ) (8.48)
8
From the foregoing discussion, it follows that in the steady state, as long as a correct
decision is made concerning ξ [w(k + 1)] − ξ [w(k)], the average total misadjustment is
given by (8.48).
ˆ
In practice, the ARS algorithm examines the statistic [w(k ˆ
+ 1)] − [w(k)] instead
of [w(k + 1)] − [w(k)], and the measured statistic contains noise that may yield a
misleading indication of the performance measure difference. The performance measure
difference due to the weight vector perturbation must be significantly larger than the
standard deviation of the error in the estimated change in the performance measure: this
is done by selecting = [w(k + 1)] − [w(k)] > σγ where σγ2 is given by (8.7).
Selecting K and μs so that > σγ results in the average steady-state misadjustment
approximated by (8.48). Furthermore, the performance measure difference due to the
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:10 344
FIGURE 8-3
vp = √ r 2 + m 20 − 2rm0 cos q vmax = m0 + r
Weight vector
component
perturbation
m0
resulting from step
size μ0 starting from
vss = μ0 /3.
m0
r
wopt 2
vmin = m0 − r
weight vector perturbation should be less than E{[w(k)]}. Therefore, for selected to
be ξ , in the steady state the constants K and μs should be selected to satisfy
9 2 2ξmin
ξmin > μ tr(Rxx ) > √ (8.49)
8 0 K
Even with K and μs selected to satisfy > σγ , it is still possible for noise present
in the measurements of system performance to produce deceptively good results for any
one experiment. Such spurious results will on the average be corrected on successive trials.
σ 2 = K 1 + K 2 ∗ (8.50)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:10 345
FIGURE 8-4
Block diagram for
GARS algorithm.
Calculate
w(k+1) = w(k) + Δw(k) Array weighting
Array output
Start
Generate w0
≥0 b [w(k+1)] − <0
b [w(k)]
where K 1 and K 2 are design constants for the GARS algorithm selected so the step size
is small enough when the optimum performance is realized yet big enough to gain useful
performance surface information when the trial weight vector is far from optimum.
The next trial adaptive weight vector is then computed using
w(k + 1) = w(k) + w(k) (8.51)
and the corresponding performance measure, [w(k + 1)], is evaluated. If no improve-
ment in the performance measure is realized, the algorithm remains in the random phase
for the next trial weight vector, returning to the previous value of w as the starting point
for the next weight perturbation. Once a direction in which to move for improved perfor-
mance is determined, the deterministic phase of the algorithm is entered, and convergence
is accelerated by continuing to travel in the direction of improved performance with twice
the previous step size. The weight vector step w is continually doubled as long as per-
formance measure improvements are realized. Once the performance measure begins to
degrade, the search is returned to the random phase where the adaptive weight vector
perturbations w are considerably smaller than before due to the smaller value of σ used
in generating new search directions.
From the foregoing description of the simplified version of GARS, it is seen that the
principal difference between GARS and ARS lies in how the random phase of the search
is conducted. Not only is the search direction random (as it was before), but the step size
is also random and governed by the parameter σ whose assigned value depends on the
minimum value that the selected performance measure has attained. As a result, the search
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:10 346
step reduces as performance measure improvements are realized. The observations made
in the previous section for the ARS algorithm with minimum step size μ0 now apply in a
statistical sense to the GARS algorithm. When is selected to be ξ the condition expressed
by (8.49) should also be satisfied where now the expected change in performance measure
due to weight perturbations is given by
E{ξ } = σ 2 tr(Rxx ) (8.52)
Noise in the measurements of the performance measure can produce deceptively
good results for any one experiment, thereby creating the risk of spurious measurements
locking the search to a false solution. Local minima are avoided in the random phase of the
algorithm by periodically reexamining the performance measure of the best weight vector
found so far. Furthermore, the algorithm handles nonstationary operating conditions by
periodically using large step sizes and conducting the exploration perturbations uniformly
throughout the parameter space.
The general GARS algorithm incorporates features that are not simulated here. The
most important such feature is the provision for a long-term memory due to employing a
nonuniform multivariate probability distribution function [20] to generate search directions
and thereby to guide the search to increase the probability that future trials will yield better
performance scores than past trials. This multivariate PDF is shaped according to the
results of a series of initial trials conducted during the opening stage of the search when
no preference is given to any search direction. During the middle stage of the search,
the multivariate PDF formed during the opening stage guides the search by generating
new search directions. In the final search stage, the dimensionality of the parameter space
search is reduced by converting from a simultaneous search involving all the parameters to
a nearly sequential search involving only a small fraction of the parameters at any step. This
selected fraction of the parameters to search is chosen randomly for each new iteration.
chromNpop costNpop
2. Natural selection
chrom1 cost1
⇒
chromNmate costNmate
3. Mating
chroma
chromb ⇒ chromc
chrom1
chromc chromNmate +1
⇒ ⇒ binary mask ⇒
chromd chromNmate +2
chromNmate
chrome ⇒ chrome
chromf
4. Mutation
chrom1 100011 10001 100011 10001
111000 01010 101000 01010
001110 01100 001110 01100
chromNmate
⇒ 100111 00111 ⇒ 100111 01111
chromNmate +1
No
Converged?
Yes
1. Form population and evaluate cost. A GA starts with a random population matrix
with N pop rows . Each row is a chromosome and contains the adaptive weights for all the
array elements. Since the adapted weights are normally digital, the population matrix is
binary. If the adaptive weights have Nb bits and there are Na adaptive elements, then each
chromosome contains Nb × Na bits. Thus, the population matrix is given by
⎡ ⎤
⎢ b ⎥ ⎡ chrom ⎤
⎢ 1,1,1 b1,1,2 · · · b1,1,Nbits · · · b1,Na ,1 b1,Na ,2 · · · b1,Na ,Nbits ⎥ 1
⎢ ⎥ ⎢
⎢ b2,1,1 b2,1,2 b2,Na ,1 b2,Na ,2 ⎥ ⎢ chrom2 ⎥ ⎥
P=⎢
⎢ .. .. .. .. .. .. ⎥=⎢
⎥ ⎣ .. ⎥
⎢ . . . . . . ⎥ . ⎦
⎢ ⎥
⎣ b Npop ,1,1 b Npop ,1,2 · · · b Npop ,Nbits · · · b Npop ,Na ,1 b Npop ,Na ,2 · · · b Npop ,Na ,Nbits ⎦ chrom Npop
Element 1 Element Na
(8.53)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:10 348
A cost function generates the cost or output from the chromosomes in the population
matrix and places them in a corresponding cost vector. Each row is sent to the cost function
for evaluation, so that cost, cn , of chromn is given by
cn = f (chromn ) (8.54)
Then, all the costs are placed in a vector.
C = [ c1 c2 · · · c N pop ]T (8.55)
In the case of an adaptive array, the cost is the array performance measure.
2. Natural selection. Once all of the costs are determined, then natural selection occurs
each generation (iteration) and discards unacceptable chromosomes that have a high cost.
Of the N pop chromosomes in a generation, the Nmate chromosomes with the lowest cost
survive and form the mating pool, whereas the bottom Npop − Nmate are discarded to make
room for the new offspring.
3. Mating. Mating is the process of selecting two chromosomes from the mating pool
of Nmate chromosomes to produce two new offspring. Chromosomes from the mating
population are selected and paired to create Npop −Nmate offspring that replace the discarded
chromosomes. Tournament selection is a popular way to select chromosomes for mating.
It randomly selects a subset of chromosomes from the mating pool, and the chromosome
with the lowest cost in this subset becomes a parent. The tournament repeats Npop − Nmate
times.
Mating combines the characteristics of two chromosomes to form two new chro-
mosomes that replace two chromosomes discarded in the natural selection step. Uniform
crossover is a general procedure that selects variables from each parent chromosome based
on a mask and then places them in a new offspring chromosome. First, a random binary
mask is created. A 1/0 in the mask column means the offspring receives the variable value
from chromm/n . If it has a 0/1, then the offspring receives the variable value in chromn/m .
where
vm,n = particle velocity
pm,n = particle variables
r1 , r2 = independent uniform random numbers
1 = cognitive parameter
2 = social parameter
local best
pm,n = best local solution
global best
pm,n = best global solution
PSO updates the velocity vector for each particle and then adds that velocity to the particle
position. Velocity updates depend on the estimate of the global solution found thus far
and the best local solution in the present population. If the best local solution has a cost
less than the cost of the estimate of the global solution, then the best local solution replaces
the global solution estimate.
As an example, consider a 40-element array along the x-axis with elements spaced
d = λ/2 apart and a 30 dB Chebyshev amplitude taper (an ) [24]. Elements have a sin φ
element pattern and six-bit phase shifters. Assume there are two interference sources at
φ = 43.9◦ and 51.7◦ that are 60 dB stronger than the desired signal power in the main
beam. The cost function assumes the phase shifts are antisymmetric about the center of
the array.
2
20
2π
cost = 20 log10 si sin φi an cos (n − 1) du i + δn (8.59)
λ
i=1 n=1
where si is the signal strength of the two interference signals, u i = cos φi , and δn are the
adapted quantized phases.
Only two least significant bits are needed to perform the nulling (11.25◦ and 5.625◦ ).
The adapted phase settings are shown in Figure 8.6. Figure 8.7 shows the adapted pattern
FIGURE 8-6
15 Adapted phase
weights that place
nulls in the far-field
10
pattern at u = 0.62
and u = 0.72.
5 From R. L. Haupt,
“Phase-only
Phase
5 10 15 20 25 30 35 40
Element
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:10 350
FIGURE 8-7 0
Adapted (solid line)
and quiescent
(dotted line) array
−10
patterns for two
60 dB interference
Relative pattern in dB
sources at u = 0.62
and u = 0.72. From −20
R. L. Haupt,
“Phase-only
adaptive nulling with −30
a genetic algorithm,”
IEEE Transactions on
Antennas and
Propagation, Vol. 45, −40
No. 6, 1997,
pp. 1009–1015.
−50
−1 −0.5 0 0.5 1
u
−50
adaptive nulling with
a genetic algorithm,” −55
IEEE Transactions on
Antennas and −60
Propagation, Vol. 45,
No. 6, 1997, pp. −65
1009–1015.
−70
−75
0 5 10 15 20 25
Iteration
FIGURE 8-9
30
Output power of the
best chromosome
25 (solid line) and
average
chromosome of the
Output power in dB
20
population (dashed
line). From R. L.
15 Haupt, “Phase-only
adaptive nulling with
a genetic algorithm,”
10
IEEE Transactions on
Antennas and
5 Propagation, Vol. 45,
No. 6, 1997,
0
pp. 1009–1015.
0 5 10 15 20 25
Iteration
10
−5
SNR in dB
−10
−15
−20
−25
−30
−35
0 5 10 15 20 25
Iteration
FIGURE 8-10 Signal-to-noise ratio (SNR) of the array as a function of the number of iter-
ations of the genetic algorithm. The solid line is the SNR of the best chromosome, and the
dashed line is the SNR of the average chromosome of the population. From R. L. Haupt,
“Phase-only adaptive nulling with a genetic algorithm,” IEEE Transactions on Antennas and
Propagation, Vol. 45, No. 6, 1997, pp. 1009–1015.
to form a fixed elevation main beam pointing 3◦ above horizontal. Eight consecutive
elements are active at a time and with the elements spaced 0.42λ apart at 5 GHz. Each
element has an eight-bit phase shifter (least significant bit equal to 0.0078125π radians)
and eight-bit attenuator (least significant bit equal to .3125 dB). The antenna has a 25 dB
n = 3 Taylor amplitude taper.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:10 352
FIGURE 8-12
Cross section of q
experimental
cylindrical array.
Elements
Active
elements
A 5 GHz continuous source served as the interference. Only the four least significant
bits of the phase shifters and attenuators were used to perform the adaptive nulling, so
minimal distortion occurs to the main beam. The genetic algorithm had a population size
of 16 chromosomes, and only one bit in the population was mutated every generation (mu-
tation rate of 0.1%). The algorithm placed a deep null in less than 30 power measurements
as shown in Figure 8.13 when the interference was at 45◦ . Figure 8.14 is the convergence
plot for placing the null in the antenna pattern in Figure 8.13.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:10 353
0 FIGURE 8-13
Measured adapted
and quiescent
−10 far-field patterns for
an interference
source at 45◦ .
Far field pattern (dB)
−20
−30
−40
Quiescent
−50
Adapted
−60
−60 −30 0 30 60
q (degrees)
FIGURE 8-14 GA
convergence for
−30 placing a null at 45◦ .
Sidelobe level (dB)
−40
−50
−60
0 1 2 3 4 5
Generation
−10.00
−15.00
−20.00
−25.00
1 2 3 4 5678 1E1 2 3 4 5678 1E2 2 3 4 5678 1E3 2 3 4 5
Number of iterations
algorithm with
eigenvalue 0
spread = 153.1. −5.00
−10.00
−15.00
−20.00
−25.00
1 2 3 4 5 678 1E1 2 3 4 5 678 1E2 2 3 4 5 678 1E3 2
Number of iterations
algorithm with
eigenvalue 0
spread = 153.1. −5.00
−10.00
−15.00
−20.00
−25.00
1 2 3 4 5 678 1E1 2 3 4 5 678 1E2 2 3 4 5 678 1E3 2
Number of iterations
give typical convergence results for the case where λmax /λmin = 153.1. Likewise, when
we select jammer-to-thermal noise ratios of J1 /n = 500, J2 /n = 40, and J3 /n = 200
and a signal-to-thermal noise ratio of s/n = 10, the corresponding eigenvalues are then
λ1 = 2440, λ2 = 494, λ3 = 25.6, and λ4 = 1 for which SNRopt = 15.08 (11.8 dB).
Figures 8.19–8.24 then give typical convergence results for the case where λmax /λmin =
2, 440.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:10 355
−10.00
−15.00
−20.00
−25.00
1 2 3 4 5 678 1E1 2 3 4 5 678 1E2 2 3 4 5 678 1E3 2
Number of iterations
algorithm with
0 eigenvalue
−5.00 spread = 2,440.
−10.00
−15.00
−20.00
−25.00
1 2 3 4 5678 1E1 2 3 4 5678 1E2 2 3 4 5678 1E3 2 3 4 5
Number of iterations
algorithm with
0 eigenvalue
−5.00 spread = 2,440.
−10.00
−15.00
−20.00
−25.00
1 2 3 4 5 678 1E1 2 3 4 5 678 1E2 2 3 4 5 678 1E3 2
Number of iterations
The ARS weight adjustment scheme in Figure 8.1 needs modification, because the
farther w(k + 1) is from wopt the greater is the variance in the estimate ξ̂ [w(k + 1)].
Consequently, if the step size μ0 is selected to obtain an acceptable steady-state error in
the neighborhood of wopt , it may well be that the changes in ξ [w(k + 1)] occurring as a
consequence of the perturbation w(k + 1) are overwhelmed by the random fluctuations
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:10 356
−10.00
−15.00
−20.00
−25.00
1 2 3 4 5 678 1E1 2 3 4 5 678 1E2 2 3 4 5 678 1E3 2
Number of iterations
FIGURE 8-22 15
Output SNR versus
number of iterations Optimum
for the GA with 10
eigenvalue
spread = 2,440. 5
SNR (dB)
−5
−10
−15
100 101 102 103
Iteration
FIGURE 8-23 90
Adapted pattern 120 60
after 1,000 iterations
of the GA. 10 dB
150 30
0 dB
−10 dB
180 0
210 330
240 300
270
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:10 357
−10.00
−15.00
−20.00
−25.00
1 2 3 4 5678 1E1 2 3 4 5678 1E2 2 3 4 5678 1E3 2 3 4 5
Number of iterations
experienced in ξ̂ [w(k + 1)] when w(k + 1) is far from wopt . When this situation occurs,
the adjustment algorithm yields a succession of weights that slowly meander aimlessly
with step size μ0 . As a result, the step size μ0 should reflect the changes in the variance
of ξ̂ [w(k + 1)] that occur when ! w(k + 1) is far from wopt . This correction involves
incorporating a step size μs = K 1 + K 2 ∗ into the ARS algorithm in accordance with
the philosophy expressed! by the GARS algorithm in Figure 8.4. Of course, it would be
preferable to use μs = K 1 + K 2 (∗ − min ), but in general min is unknown.
The LRS, ARS, and GARS algorithms were all simulated using K = 90 to obtain
the estimate of MSE, which was the performance measure used in all cases. To satisfy the
condition imposed by (8.49), the GARS algorithm was simulated using
σ 2 tr(Rxx ) = 1 + 2 ξ ∗ (8.60)
where 1 = 1
160
and 2 = 0.1. Likewise, the ARS algorithm was simulated using
with 1 and 2 assigned the same values as for the GARS algorithm. The LRS algorithm was
simulated using the constants μs = 1.6 and σ 2 tr(Rxx ) = 0.05, thereby yielding a greater
misadjustment error than either the ARS or GARS algorithms. The LMS algorithm was
also simulated for purposes of comparison with step size corresponding to μs tr(Rxx ) = 0.1
and using an estimated gradient derived from the average value of three samples of e(k)x(k)
so K = 3 instead of the more common K = 1. In all cases the initial weight vector was
taken to be wT (0) = [0.1, 0, 0, 0].
The results of Figures 8.15–8.17 show that both the ARS and GARS algorithms are
within 3 dB of the optimum SNR after about 800 iterations, whereas the LRS algorithm
does not reach this point after 4,000 iterations, even though the misadjustment is more
severe than for the ARS and GARS algorithms. This result indicates that the misadjustment
versus speed of adaptation trade-off favors the ARS and GARS algorithms more than the
LRS algorithm. The LMS algorithm by contrast is within 3 dB of the optimum output
SNR after only 150 iterations with only a small degree of misadjustment. The extreme
disparity in speed of convergence between the LMS algorithm and the three random search
algorithms is actually more pronounced than the comparison of number of iterations
indicates because each iteration in the random search algorithms represents 90 samples,
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:10 358
whereas each iteration in the LMS algorithm represents only three samples. Consequently,
the time scale on Figures 8.15–8.17 is 30 times greater than the corresponding time scale
on Figure 8.18.
The results given in Figures 8.18–8.24 for the case where eigenvalue spread = 2,440
confirm the previous results obtained with an eigenvalue spread of only 153.1. These
results also show that both the random search algorithms and the LMS algorithm are
sensitive to eigenvalue spread in the Rxx matrix.
Figure 8.22 is a plot of the signal to noise ratio versus iteration when a GA controls the
complex element weights of the Y -array. The GA has a population size of 8 and a mutation
rate of 15%. The best result for each iteration appears in the plot. The number of power
measurements per iteration is less than 7. The GA optimized the SNR without knowledge
of the signal and jammer directions and powers. Figure 8.23 shows the adapted pattern
after 1,000 iterations of the GA. Phase-only adaptive nulling is not a good alternative in
this case, because there are not enough degrees of freedom to null all the jammers. It is
remarkable that the convergence results of the GA in Figure 8-22 are very close to the
results obtained for the LMS algorithm shown in Figure 8-24. The GA algorithm is the
only random search algorithm that is actually competitive with the LMS algorithm in this
extreme eigenvalue spread condition.
8.7 PROBLEMS
1. Misadjustment versus Speed of Adaptation Trade-off for the LRS Algorithm [15]
Assuming all eigenvalues are equal so that (T pmse )av = Tmse , plot Tmse versus N for
the LRS algorithm assuming (Mtot )min = 10% in (8.43), and compare this result with
the corresponding plots obtained for the LMS and DSD algorithms in Problem 1 of
Chapter 4.
2. Search Loss for a Simple Random Search with Reversing Step [27]
Consider the simple random search algorithm described by
where n is the number of degrees of freedom and p(φ) is the probability density
function of the angle φ for a uniform distribution of directions of the random step
in the n-dimensional space. Show that
sinn−2 φ (n − 1)
p(φ) = # π/2 = n−1 2 sin
n−2
φ
2 0
n−2
sin φdφ 2n−2 2
where (·) is the gamma function.
Hint: Note that the area of a ring-shaped zone on the surface of an n-dimensional
sphere corresponding to the angle dφ is An−2 × sinn−2 φdφ. Consequently, the area
FIGURE 8-25
Parameter space
section showing
displacement vector
X and direction φ.
f
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:10 360
of the surface of the hypersphere included in the hypercone with vertical angle 2φ
is given by
" φ
S(φ) = An−2 sinn−2 φdφ
0
The probability that a random vector lies in this cone for a uniform probability
choice of the random direction is equal to the ratio of the “areas” S(φ) and S(π).
The desired probability density is then the derivative of this ratio with respect to the
angle φ.
(b) Show that U (n) defined in part (a) is given by
# π/2
cos φ sinn−2 φdφ (n − 1)
U (n) = 0
# π/2 = 2
0 sin n−2
φdφ 2n−3 (n − 1) n−1
2
Since the probability of successful and unsuccessful steps is the same for the pa-
rameter space of Figure 8.25, then on the average for one successful step there is
one unsuccessful step and a corresponding reverse step—three steps in all. There-
fore, the mean displacement for one successful step is reduced by two-thirds and is
only 13 U (n).
Defining search loss to be the number of steps required by the search such
that the vector sum of these steps has the same length as one operating step in the
successful direction, it follows that the mean search loss for the aforementioned
algorithm is just 3/U (n).
3. Relative Search Efficiency Using the Search Loss Function [27]
Consider a fixed step size gradient search defined by
(i)
x(i + 1) = x(i) − a(i)μs
|(i)|
where
μs = step size
∂F ∂F
(i) = ,...,
∂ x1 ∂ xn x(i)
1 if F[x(i + 1)] < F[x(i)]
a(i) =
0 otherwise
and where F[·] denotes a known performance measure. Likewise, consider a fixed
step size random search defined by
Define the search loss function to be the performance measure F[x] divided by the
ratio of performance improvement per performance evaluation, that is,
F(x)
SL(x) =
F(x)−F(x+x)
N
ρ 2 (n + 1)
S L(x) =
2ρμs − μ2s
(c) For a given base point x(i), the successor trial state x(i + 1) for the random search
defines an angle φ with respect to a line connecting x(i) with the extremum point
of the performance measure for which the probability density function p(φ) was
obtained in Problem 2. Show that the expected value of performance measure im-
provement using the previously defined random search is
" φ0
E{−F} = E{ρ } = 2
ρ 2 p(φ)dφ
0
where
−1 μs
φ0 = cos and ρ 2 = 2ρμs cos φ − μ2s
2ρ
The search loss function of parts (b) and (d) is compared for specific values of
ρ, μs , and n to determine whether the gradient search or the random search is more
efficient.
4. Search Loss Function Improvement Using Step Reversal [28]
The relative efficiency of the fixed step size random algorithm introduced in Problem
3 is significantly improved merely by adding a “reversal” feature to the random search.
The fixed step size random search algorithm with reversal is described by
where
1 if F[x(i)] ≤ F[x(i − 1)]
c(i) =
0 otherwise
8.8 REFERENCES
[1] M. K. Leavitt, “A Phase Adaptation Algorithm,” IEEE Trans. Antennas Propag., Vol. AP-24,
No. 5, September 1976, pp. 754–756.
[2] P. A. Thompson, “Adaptation by Direct Phase-Shift Adjustment in Narrow-Band Adaptive
Antenna Systems,” IEEE Trans. Antennas Propag., Vol. AP-24, No. 5, September 1976,
pp. 756–760.
[3] G. J. McMurty, “Search Strategies in Optimization,” Proceedings of the 1972 International
Conference on Cybernetics and Society, October, Washington, DC, pp. 436–439.
[4] C. Karnopp, “Random Search Techniques for Optimization Problems,” Automatica, Vol. 1,
August 1963, pp. 111–121.
[5] G. J. McMurty and K. S. Fu, “A Variable Structure Automation Used as a Multimodal Search-
ing Technique,” IEEE Trans. Autom. Control, Vol. AC-11, July 1966, pp. 379–387.
[6] R. L. Barton, “Self-Organizing Control: The Elementary SOC—Part I,” Control Eng., February
1968.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:10 363
[7] R. L. Barton, “Self-Organizing Control: The General Purpose SOC— Part II,” Control Eng.,
March 1968.
[8] M. A. Schumer and K. Steiglitz, “Adaptive Step Size Random Search,” IEEE Trans. Autom.
Control, Vol. AC-13, June 1968, pp. 270–276.
[9] R. A. Jarvis, “Adaptive Global Search in a Time-Variant Environment Using a Probabilistic
Automaton with Pattern Recognition Supervision,” IEEE Trans. Syst. Sci. Cybern., Vol. SSC-6,
July 1970, pp. 209–217.
[10] A. N. Mucciardi, “Self-Organizing Probability State Variable Parameter Search Algorithms
for the Systems that Must Avoid High-Penalty Operating Regions,” IEEE Trans. Syst., Man,
Cybern., Vol. SMC-4, July 1974, pp. 350–362.
[11] R. A. Jarvis, “Adaptive Global Search by the Process of Competitive Evolution,” IEEE Trans.
Syst., Man, Cybern., Vol. SMC-5, May 1975, pp. 297–311.
[12] R. L. Haupt, “Beamforming with Genetic Algorithms,” Adaptive Beamforming, S. Chandran,
ed., pp. 78–93, Berlin, Springer-Verlag, 2004.
[13] R. L. Haupt, “Genetic Algorithms for Antennas,” Modern Antenna Handbook, C. Balanis,
New York, Wiley, 2008.
[14] R. L. Haupt, Antenna Arrays: A Computational Approach, New York, Wiley, 2010.
[15] B. Widrow and J. M. McCool, “A Comparison of Adaptive Algorithms based on the Methods
of Steepest Descent and Random Search,” IEEE Trans. Antennas Propag., Vol. AP-24, No. 5,
September 1976, pp. 615–637.
[16] C. A. Baird and G. G. Rassweiler, “Search Algorithms for Sonobuoy Communication,” Pro-
ceedings of the Adaptive Antenna Systems Workshop, March 11–13, 1974, NRL Report 7803,
Vol. 1, September 27, 1974, pp. 285–303.
[17] R. L. Barron, “Inference of Vehicle and Atmosphere Parameters from Free-Flight Motions,”
AIAA J. Spacecr. Rockets, Vol. 6, No. 6, June 1969, pp. 641–648.
[18] R. L. Barron, “Guided Accelerated Random Search as Applied to Adaptive Array AMTI
Radar,” Proceedings of the Adaptive Antenna Systems Workshop, March 11–13, 1974, Vol. 1,
NRL Report 7803, September 27, 1974, pp. 101–112.
[19] A. E. Zeger and L. R. Burgess, “Adaptive Array AMTI Radar,” Proceedings of the Adaptive
Antenna Systems Workshop, March 11–13, 1974, NRL Report 7803, Vol. 1, September 27,
1974, pp. 81–100.
[20] A. N. Mucciardi, “A New Class of Search Algorithms for Adaptive Computation,” Proc.
1973 IEEE Conference on Decision and Control, December, San Diego, Paper No. WA5-3,
pp. 94–100.
[21] R. L. Haupt and S. E. Haupt, Practical Genetic Algorithms, 2d ed., New York, John Wiley &
Sons, 2004.
[22] R. L. Haupt and D. Werner, Genetic Algorithms in Electromagnetics, New York, Wiley, 2007.
[23] J. Kennedy and R. C. Eberhart, Swarm Intelligence, San Francisco, Morgan Kaufmann
Publishers, 2001.
[24] R. L. Haupt, “Phase-Only Adaptive Nulling with a Genetic Algorithm,” IEEE Transactions
on Antennas and Propagation, Vol. 45, No. 6, 1997, pp. 1009–1015.
[25] R. L. Haupt and H. L. Southall, “Experimental Adaptive Nulling with a Genetic Algorithm,”
Microwave Journal, Vol. 42, No. 1, January 1999, pp. 78–89.
[26] R. Hooke and T. A. Jeeves, “Direct Search Solution of Numerical and Statistical Problems,”
J. of the Assoc. for Comput. Mach., Vol. 8, April 1961, pp. 212–229.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:10 364
[27] L. A. Rastrigin, “The Convergence of the Random Search Method in the External Con-
trol of a Many-Parameter System,” Autom. Remote Control, Vol. 24, No. 11, April 1964,
pp. 1337–1342.
[28] J. P. Lawrence, III and F. P. Emad, “An Analytic Comparison of Random Searching and
Gradient Searching for the Extremum of a Known Objective Function,” IEEE Trans. Autom.
Control, Vol. AC-18, No. 6, December 1973, pp. 669–671.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:44 365
CHAPTER
Adaptive Algorithm
Performance Summary 9
−5.00
−10.00
−15.00
−20.00
−25.00
1 2 3 4 5678 1E1 2 3 4 56 8 1E2 2 3 4 56 8 1E3 2 3 4 5
Number of iterations
per iteration. 0
−5.00
−10.00
−15.00
−20.00
−25.00
1 2 3 4 5 6 7 8 1E1 2 3 4 5 6 7 8 1E2 2 3 4 56
Number of iterations
−5.00
−10.00
−15.00
−20.00
−25.00
1 2 3 4 5 6 7 8 1E1 2 3 4 5 6 7 8 1E2 2 3 4 56
Number of iterations
recursive algorithms decreases as the number of iterations increases and cannot be mod-
ified by altering algorithm parameters. These results indicate that the DMI and recursive
algorithms offer by far the best misadjustment versus speed of convergence trade-off, fol-
lowed (in order) by the GSCP algorithm, the PAG algorithm, and the LMS algorithm for
this moderate eigenvalue spread condition of λmax /λmin = 2,440.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:44 367
0 and P(0) = I.
−5.00
−10.00
−15.00
−20.00
−25.00
1 2 3 4 5 6 7 8 1E1 2 3 4 5 6 7 8 1E2 2 3 4 56
Number of iterations
0 K = 3 samples per
−5.00 iteration.
−10.00
−15.00
−20.00
−25.00
1 2 3 4 5 6 7 8 1E1 2 3 4 5 6 7 8 1E2 2 3 4 5 6 7 1E3
Number of iterations
Table 9.1 summarizes the principal operational characteristics associated with the
adaptive algorithms considered throughout Part 2. The algorithms that achieve the simplest
possible instrumentation by requiring only direct measurement of the selected performance
measure pay a severe penalty in terms of the increased convergence time required to reach
the steady-state solution for a given degree of misadjustment. Accepting the instrumenta-
tion necessary to incorporate one correlator for each controlled array element enables the
misadjustment versus convergence speed trade-off for the LMS and Howells–Applebaum
interference suppression loops to be achieved. Further improvement in the misadjustment
versus convergence speed trade-off is obtained where sensitivity to eigenvalue spread is
a concern by paying the price of additional instrumentation or additional computational
power. As the shift to digital processing continues, algorithms requiring more sophisti-
cated computation become not only practicable but also preferable in many cases. Not
only can high performance be achieved that was impractical before, but also the low cost
of the increased computational power may in some cases render a sophisticated algorithm
more economical.
The genetic algorithm (GA) is a “global” minimum seeker, so it is less likely to
get stuck in a local minimum like the steepest descent algorithm. It minimizes the total
output power, if the number of adaptive elements or the adaptive weight range are limited.
TABLE 9-1 Operational Characteristics Summary of Selected Adaptive Algorithms
368
LMS MSNR PAG DSD DMI R GSCP RS GA
Monzingo-7200014
Algorithm Steepest descent Similar Conjugate Perturbation Direct estimate Data weighting Orthogonalize Trial and error Minimize total
philosophy to LMS gradient descent technique of covariance; similar to input signals power or
open loop Kalman filter; maximize SNR
book
closed loop
Transient Misadjustment Nearly More favorable Unfavorable Achieves the Same as DMI More favorable Unfavorable Limited by the
response vs. convergence the misadjustment misadjustment fastest misadjustment misadjustment hardware
characteristic speed trade-off same as vs. convergence vs. convergence convergence vs. convergence vs. convergence
is acceptable for LMS speed trade-off speed trade-off with most speed trade-off speed trade-off
numerous than LMS compared to favorable than PAG. compared to
applications LMS misadjustment Convergence LMS;
vs. convergence speed accelerated
speed trade-off approaches DMI steps improve
ISBN : XXXXXXXXXX
speed
Algorithm Easy to Same as Convergence Easy to Very fast Same as DMI; Convergence Can be applied Minimal
strengths implement, LMS speed less implement; convergence different data speed enjoys to any directly hardware, fast,
requiring N sensitive to requires only speed weighting reduced measurable does not get
correlators and eigenvalue instrumentation independent of schemes are sensitivity to performance stuck in local
integrators; spread than to directly eigenvalue easily eigenvalue index; easy to minimum,
tolerant of LMS; fast measure the spread incorporated spread implement, with independent of
hardware errors convergence for performance compared with meager eigenvalue
November 24, 2010
Algorithm Convergence Same as Relatively Rate of Requires Requires Requires a Convergence Some
weaknesses speed sensitive LMS difficult to convergence N (N + 1)/2 N (N + 1)/2 respectable speed sensitive measurements
to eigenvalue implement and sensitive to correlators to correlators and amount of to eigenvalue are bad each
368
spread requires parallel eigenvalue implement; heavy hardware— spread and the iteration
processors; spread, with matrix inversion computational N (N + 1)/2 slowest of all
convergence speed requires load adaptive loops algorithms
speed sensitive comparable to adequate to implement considered
to number of that of RS with precision and
degrees of accelerated step N 3 /2 + N 2
freedom complex
multiplies
Chapter 4 4 4 4 5 6 7 8 8
reference
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:44 369
Since it works with standard phased array hardware, it can be implemented on existing
arrays and is much cheaper than requiring a receiver at each element. The convergence
is relatively independent of the interference and signal power levels. Each iteration, the
GA must evaluate the entire population, which can be small. As a result, there are bad
measurements in addition to the improved power and SNR measurements. The transient
response is limited by the hardware switching speed and settling time. It may be necessary
to average a few power measurements to get the desired results.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:44 370
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 371
P A R T III
Advanced Topics
CHAPTER
Compensation of
Adaptive Arrays 10
' $
Chapter Outline
10.1 Array Errors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 374
10.2 Array Calibration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 377
10.3 Broadband Signal Processing Considerations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 380
10.4 Compensation for Mutual Coupling . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 396
10.5 Multipath Compensation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 398
10.6 Analysis of Interchannel Mismatch Effects . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 406
10.7 Summary and Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 415
10.8 Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 416
10.9 References . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 419
& %
Narrowband adaptive arrays need only one complex adaptive weight in each element chan-
nel. Broadband adaptive arrays, however, require tapped delay lines (transversal filters)
in each element channel to make frequency-dependent amplitude and phase adjustments.
The analysis presented so far assumes that each element channel has identical electron-
ics and no reflected signals. Unfortunately, the electrical characteristics of each channel
are slightly different and lead to “channel mismatching” in which significant differences
in frequency-response characteristics from channel to channel may severely degrade an
array’s performance without some form of compensation. This chapter starts with an anal-
ysis of array errors and then addresses array calibration and frequency-dependent mis-
match compensation using tapped delay line processing, which is important for practical
broadband adaptive array designs.
The number of taps used in a tapped delay line processor depends on whether the
tapped delay line compensates for broadband channel mismatch effects or for the effects
of multipath and finite array propagation delay. Minimizing the number of taps required
for a specified set of conditions is an important practical design consideration, since each
additional tap (and associated weighs) increases the cost and complexity of the adaptive
array system.
373
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 374
N
an + δna e j ( pn +δn ) e jk (sn +δn )u
p s
AF err = (10.1)
n=1
Element failures result when an element no longer transmits or receives. The probability
that an element has failed, 1 − Pe , is the same as a root mean square (rms) amplitude error,
δna 2 . Position errors are not usually a problem, so a reasonable formula to calculate the
rms sidelobe level of the array factor for amplitude and phase errors with element failures
is [1]
p2
(1 − Pe ) + δna 2 + Pe δn
sllrms = (10.2)
p2
Pe 1 − δn ηt N
Figure 10-1 is an example of a typical corporate-fed array. A random error that occurs
at one element is statistically uncorrelated with a random error that occurs in another
element in the array as long as that error occurs after the last T junction and before an
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 375
1 2 3 4 5 6 7 8 FIGURE 10-1
Corporate-fed array
with random errors.
A
B
C
10 FIGURE 10-2
Array factor
5 with random,
uncorrelated errors
0 Errors superimposed on
the error-free array
factor.
Directivity (dB)
−5
Error free
−10
−15
−20
−25
−30
−50 0 50
q (degrees)
element. If a random error occurs prior to A, for instance, then the random error becomes
correlated between the elements that share the error. For instance, a random error between
A and B results in a random correlated error shared by elements 1 and 2. Likewise, a
random error between B and C results in a random correlated error shared by elements 1,
2, 3, and 4.
As an example, consider an eight-element, 20 dB Chebyshev array that has elements
spaced λ/2 apart. If the random errors are represented by δna = 0.15 and δnp = 0.15, then
an example of the array factor with errors is shown in Figure 10-2. Note that the random
errors lower the main beam directivity, induce a slight beam-pointing error, increase the
sidelobe levels, and fill in some of the nulls.
a = 2−Nba (10.3)
p = 2π × 2−Nbp (10.4)
If the difference between the desired and quantized amplitude weights is a uniformly dis-
tributed random number with the bounds being√ the maximum amplitude error of ±a /2,
then the rms amplitude error is δna = a / 12. The quantization error is random only
when no two adjacent elements receive the same quantized phase shift. The difference
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 376
between the desired and quantized phase shifts is treated as uniform random variables
between ±√ p /2. As with the amplitude error, the random phase error formula in this case
is δnp = p / 12. Substituting this error into (b) yields the rms sidelobe level.
The phase quantization errors become correlated when the beam steering phase shift
is small enough that groups of adjacent elements have their beam steering phase quantized
to the same level. This means that N /N Q subarrays of N Q elements receive the same
phase shift. The grating lobes due to these subarrays occur at [2]
mλ m (N − 1) 2 Nbp
sin θm = sin θs ± = sin θs 1 ± sin θs 1 ± m2 Nbp (10.5)
N Q de N
The approximation in (10.5) assumes that the array has many elements. For large scan
angles, quantization lobes do not form, because the element-to-element phase difference
appears random. The relative peaks of the quantization lobes are given by [1]
√
1 1 − sin θ 2
AFNQ L = Np (10.6)
2 1 − sin θs2
Figure 10-3 shows an array factor with a 20 dB n = 3 Taylor amplitude taper for a
20-element, d = 0.5λ array with its beam steered to θ = 3◦ when the phase shifters
have three bits. Four quantization lobes appear. The quantization lobes decrease when
higher-precision phase shifters are used and when the beam is steered to higher angles.
Significant distortion also results from mutual coupling, variation in group delay
between filters, differences in amplifier gain, tolerance in attenuator accuracy, and aperture
jitter in a digital beamforming array. Aperture jitter is the timing error between samples
in an analog-to-digital (A/D) converter. Without calibration, beamforming or estimation
of the direction of arrival (DOA) of the signal is difficult, as the internal distortion is
uncorrelated with the signal. As a result, the uncorrelated distortion changes the weights
at each element and therefore distorts the array pattern.
FIGURE 10-3 0
Array factor steered
to 3 degrees with
three-bit phase
shifters compared −10
with phase shifters
Array factor (dB)
with infinite
3 bit phase shifter
precision.
−20
−30
−90 −45 0 45 90
q (degrees)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 377
FIGURE 10-4 The uncalibrated array output is less than the calibrated array output, because
errors in the uncalibrated array do not allow the signal vectors from the elements to align.
Target
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 378
FIGURE 10-6
Layout of the smart
antenna test bed.
Making power measurements for every phase setting at every element in an array
is extremely time-consuming. Calibration techniques that measure both amplitude and
phase of the calibrated signal tend to be much faster. Accurately measuring the signal
phase is reasonable in an anechoic chamber but difficult in the operational environment.
Measurements at four orthogonal phase settings yield sufficient information to obtain
a maximum likelihood estimate of the calibration phase [4]. The element phase error is
calculated from power measurements at the four phase states, and the procedure is repeated
for each element in the array. Additional measurements improve signal-to-noise ratio, and
the procedure can be repeated to achieve desired accuracy within resolution of the phase
shifters, since the algorithm is intrinsically convergent.
Another approach uses amplitude-only measurements from multiple elements to find
the complex field at an element [5]. The first step measures the power output from the
array when the phases of multiple elements are successively shifted with the different
phase intervals. Next, the measured power variation is expanded into a Fourier series to
derive the complex electric field of the corresponding elements. The measurement time
reduction comes at the expense of increased measurement error.
Transmit/receive module calibration is an iterative process that starts with adjusting
the attenuators for uniform gain at the elements [6]. The phase shifters are then adjusted to
compensate for the insertion phase differences at each element. Ideally, when calibrating
the array, the phase shifter’s gain remains constant as the phase settings are varied, but the
attenuator’s insertion phase can vary as a function of the phase setting. This calibration
should be done across the bandwidth, range of operating temperatures, and phase settings.
If the phase shifter’s gain varies as a function of setting, then the attenuators need to be
compensated as well. After iterating over this process, all the calibration settings are saved
and applied at the appropriate times.
Figure 10-6 shows an eight-element uniform circular array (UCA) in which a cen-
ter element radiates a calibration signal to the other elements in the array [7]. Since the
calibration source is in the center of the array, the signal path from the calibration source
to each element is identical. As previously noted, random errors are highly dependent
on temperature [8]. An experimental model of the UCA in Figure 10-6 was placed in-
side a temperature-controlled room and calibrated at 20◦ C. The measured amplitude and
phase errors at three temperatures are shown in Figure 10-7 and Figure 10-8, respectively.
Increasing the temperature of the room to 25◦ C then to 30◦ C without recalibration in-
creases the errors shown in Figure 10-7 and Figure 10-8. This experiment demonstrates
the need of dynamic calibration in a smart antenna array.
FIGURE 10-7
Amplitude error for
0 20° C the UCA antenna
as the system
temperature
changes from 20◦ C
Phase error (degrees)
with calibration to
−20 25° C 25◦ C without
recalibration and to
30° C 30◦ C without
recalibration.
−40
2 3 4 5 6 7 8
Signal path
FIGURE 10-8
Phase error for the
UCA antenna as the
system temperature
1.0 changes from 20◦ C
20° C
Amplitude error (Vrms)
with calibration to
25° C 25◦ C without
recalibration and
0.8 to 30◦ C without
30° C recalibration.
0.6
2 3 4 5 6 7 8
Signal path
beamforming arrays is injecting a calibration signal into the signal path of each element
in the array behind each element as shown in Figure 10-9 [9]. This technique provides
a high-quality calibration signal for the circuitry behind the element. Unfortunately, it
does not calibrate for the element patterns that have significant variations due to mutual
coupling, edge effects, and multipath.
Array elements
beamformer.
Computer
Receiver A/D
Receiver A/D
FIGURE 10-10 Alignment results (measured phase deviation from desired value).
a: Unaligned. b: After single alignment with uncorrected measurements. c: After alignment
with fully corrected measurements. From W. T. Patton and L. H. Yorinks, “Near-field alignment
of phased-array antennas,” IEEE Transactions on Antennas and Propagation, Vol. 47, No. 3,
March 1999, pp. 584–591.
element. The calibration algorithm iterates between the measured phase and the array
weights until the phase at all the elements is the same. Figure 10-10 shows the progression
of the phase correction algorithm from left to right. The picture on the left is uncalibrated,
the center picture is after one iteration, and the picture on the right is after calibration
is completed. This techniques is exceptionally good at correcting static errors prior to
deploying an antenna is not practical for dynamic errors.
qs
sin
d
qi
in
ds
H2 (w) H1(w)
Array
output
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 382
To determine whether it is possible to select H1 (ω) and H2 (ω) to satisfy (10.9) and (10.10),
solve (10.9) and (10.10) for H1 (ω) and H2 (ω). Setting H1 (ω) = |H1 (ω)| exp[ jα1 (ω)] and
H2 (ω) = |H2 (ω)| exp[ jα2 (ω)] results in
πω
|H1 (ω)| exp[ jα1 (ω)] + |H2 (ω)| exp j α2 (ω) − sin θs = exp(− jωT1 ) (10.11)
ω0
πω
|H1 (ω)| exp[ jα1 (ω)] + |H2 (ω)| exp j α2 (ω) − sin θi =0 (10.12)
ω0
To satisfy (10.9) and (10.10), it follows from (10.11) and (10.12) (as shown by the devel-
opment outlined in the Problems section) that
1
H1 (ω) = H2 (ω) = (10.13)
2 1 − cos πω ω0
(sin θi − sin θs )
π ω π
α2 (ω) = [sin θs + sin θi ] ∓ n − ωT1 (10.14)
2 ω0 2
π ω π
α1 (ω) = [sin θs − sin θi ] ± n − ωT1 (10.15)
2 ω0 2
where n is any odd integer. This result means that the amplitude of the ideal transfer
functions are equal and frequency dependent. Equations (10.14) and (10.15) furthermore
show that the phase of each filter is a linear function of frequency with the slope dependent
on the spatial arrival angles of the signals as well as on the time delay T1 of the desired
signal.
Plots of the amplitude function in (10.13) are shown in Figure 10-12 for two choices
of arrival angles (θs = 0◦ and θs = 80◦ ), where it is seen that the amplitude of the
distortionless transfer function is nearly flat over a 40% bandwidth when the desired signal
is at broadside (θs = 0◦ ) and the interference signal is 90◦ from broadside (θi = 90◦ ).
Examination of (10.13) shows that whenever (sin θ I −sin θs ) is in the neighborhood of ±1,
Bandwidth
Report ESL 3832-3, 50
40%
1975 [12]. 40
30
20
10 θs = 0°
θi = 90°
0
0 0.5 1 1.5 2
w/w 0
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 383
then the resulting amplitude function will be nearly flat over the 40% bandwidth region.
If, however, both the desired and interference signals are far from broadside (as when
θd = 80◦ and θi = 90◦ ), then the amplitude function is no longer flat.
The degree of “flatness” of the distortionless filter amplitude function is interpreted in
terms of the signal geometry with respect to the array sensitivity pattern. In general, when
the phases of H1 (ω) and H2 (ω) are adjusted to yield the maximum undistorted response to
the desired signal, the corresponding array sensitivity pattern will have certain nulls. The
distortionless filter amplitude function is then the most flat when the interference signal
falls into one of these pattern nulls.
Equation (10.13) furthermore shows that singularities occur in the distortionless chan-
nel transfer functions whenever (ω/ω0 )π(sin θi − sin θs ) = n2π where n = 0, 1, 2, . . ..
The case when n = 0 occurs when the desired and interference signals arrive from exactly
the same direction, so it is hardly surprising that the array would experience difficulty
trying to receive one signal while nulling the other in this case. The other cases when
n = 1, 2, . . ., occur when the signals arrive from different directions, but the phase shifts
between elements differ by a multiple of 2π radians at some frequency ω in the signal
band.
The phase functions α1 (ω) and α2 (ω) of (10.14) and (10.15) are linear functions of
frequency. When T1 = 0, the phase slope of H1 (ω) is proportional to sin θs − sin θi ,
whereas that of H2 (ω) is proportional to sin θi + sin θs . Consequently, when the desired
signal is broadside, α1 (ω) = −α2 (ω). Furthermore, the phase difference between α1 (ω)
and α2 (ω) is also a linear function of frequency, a result that would be expected since this
allows the interelement phase shift (which is also a linear function of frequency) to be
canceled.
wopt = R−1
x x rxd (10.16)
If the signal appearing at the output of each sensor element consists of a desired signal, an
interference signal, and a thermal noise component (where each component is statistically
independent of the others and has zero mean), then the elements of Rx x can readily be
evaluated in terms of these component signals.
Consider the tapped delay line employing real (instead of complex) weights shown
in Figure 10-13. Since each signal xi (t) is just a time-delayed version of x1 (t), it follows
that
⎫
x2 (t) = x1 (t − ) ⎪
⎪
⎪
⎬
x2 (t) = x1 (t − 2)
.. ⎪
(10.17)
. ⎪
⎪
⎭
x L (t) = x1 [t − (L − 1)]
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 384
w1 w2 w3 w4 wL
Channel
output
FIGURE 10-14
Quadrature hybrid
processing for a
two-element array.
Quadrature Quadrature
hybrid hybrid
w1 w2 w3 w4
Array output
can then be found by making use of certain Hilbert transform relations as follows [14,15]:
so that
E{x̌(t)x(t)} = 0 (10.27)
E{x(t) y̌(s)} = Ě{x(t)y(s)} (10.28)
When two different sensor element channels are involved [as with x1 (t) and x3 (t), for
example], then
E{x1 (t) x3 (t)} = rdd (τd13 ) + rII (τ I13 ) (10.32)
where τd13 and τ I13 represent the spatial time delays between the sensor elements of Fig-
ure 10-14 for the desired and interference signals, respectively. Similarly
Once Rxx and rxd have been evaluated for a given signal environment, the optimal LMS
weights can be computed from (10.16), and the steady-state response of the entire array
can then be evaluated.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 387
The tapped delay line in the element channel of Figure 10-13 has a channel transfer
function given by
Likewise, the quadrature hybrid processor of Figure 10-14 has a channel transfer function
H1 (ω) = w 1 − jw 2 (10.41)
The array transfer function for the desired signal and the interference accounts for the
effects of spatial delays between array elements. A two-element array transfer function
for the desired signal is
The spatial time delays associated with the desired and interference signals are represented
by τd and τ I , respectively, between element 1 [with channel transfer function H1 (ω)] and
element 2 [with channel transfer function H2 (ω)]. With two sensor elements spaced apart
by a distance d as in Figure 10-11, the two spatial time delays are given by
d
τd = sin θs (10.44)
d
τ I = sin θ I (10.45)
The output signal-to-total-noise ratio is defined as
Pd
SNR = (10.46)
PI + Pn
where Pd , PI , and Pn represent the output desired signal power, interference signal power,
and thermal noise power, respectively. The array output power for each of the foregoing
three signals may now be evaluated. Let φdd (ω) and φII (ω) represent the power spectral
densities of the desired signal and the interference signal, respectively; then the desired
signal output power is given by
∞
Pd = φdd (ω)|Hd (ω)|2 dω (10.47)
−∞
where Hd (ω) is the overall transfer function seen by the desired signal, and the interference
signal output power is
∞
PI = φII (ω)|H1 (ω)|2 dω (10.48)
−∞
where HI (ω) is the overall transfer function seen by the interference signal. The thermal
noise present in each element output is statistically independent from one element to the
next. Let φnn (ω) denote the thermal noise power spectral density; then the noise power
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 388
Consequently, the total thermal noise output power from a two-element array is
∞
Pn = φnn (ω)[|H1 (ω)|2 + |H2 (ω)|2 ] dω (10.51)
−∞
The foregoing expressions may now be used in (10.46) to obtain the output signal-to-total-
noise ratio.
where φ(t) denotes a phase angle that is either zero or π over each bit interval, and θ is
an arbitrary constant phase angle (within the range [0, 2π ]) for the duration of any signal
pulse. The nth bit interval is defined over T0 + (n − 1)T ≤ t ≤ T0 + nT , where n is any
integer, T is the bit duration, and T0 is a constant that determines where the bit transitions
occur, as shown in Figure 10-16.
Assume that φ(t) is statistically independent over different bit intervals and is zero
or π with equal probability and that T0 is uniformly distributed over one bit interval;
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 389
x5 x1
w3 w1
−90° 0° −90° 0°
QH QH
x4 x3 x2 x1 λ0
Δ Δ
4
w1 w2 w3 w4 w4 w2
x4 x2
Σ Σ
(a) (b)
x6 x1
w6 w1
λ0
Δ Δ
8
x7 x2
w7 w2
x4 x1
w4 w1 λ0
Δ Δ
8
x8 x3
λ0 w8 w3
Δ Δ
4
x5 x2 λ0
w5 w2 Δ Δ
8
λ0 x9 x4
Δ Δ w9 w4
4
λ0
x6 x3 Δ Δ
w6 w3 8
x10 x5
w10 w5
Σ Σ
(c) (d)
FIGURE 10-15 Four adaptive array processors for broadband signal processing
comparison. a: Quadrature hybrid. b: Two-tap delay line. c: Three-tap delay line. d: Five-tap
delay line. From Rodgers and Compton, IEEE Trans. Aerosp. Electron. Syst., January 1979 [13].
then, sd (t) is a stationary random process with power spectral density given by
A2 T sin(T /2)(ω − ω0 ) 2
φdd (ω) = (10.53)
2 (T /2) (ω − ω0 )
This power spectral density is shown in Figure 10-17.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 390
t
T0
ω
w0 − w1 w0 w0 + w1
The reference signal equals the desired signal component of x1 (t) and is time aligned
with the desired component of x2 (t). The desired signal “bandwidth” will be taken to be
the frequency range defined by the first nulls of the spectrum given by (10.53). With this
definition, the fractional bandwidth then becomes
2ω1
desired signal bandwidth = (10.54)
ω0
where ω1 is the frequency separation between the center frequency ω0 and the first null
2π
ω1 = (10.55)
T
Assume that the interference signal is a Gaussian random process with a flat, bandlim-
ited power spectral density over the range ω0 − ω1 < ω < ω0 + ω1 ; then the interference
signal spectrum appears in Figure 10-18. Finally, the thermal noise signals present at
each element are statistically independent between elements, having a flat, bandlimited,
Gaussian spectral density over the range ω0 − ω1 < ω < ω0 + ω1 (identical with the
interference spectrum of Figure 10-18).
With the foregoing definitions of signal spectra, the integrals of (10.48) and (10.51)
yielding interference and thermal noise power are taken only over the frequency range ω0 −
ω1 < ω < ω0 + ω1 . The desired signal power also is considered only over the frequency
range ω0 − ω1 < ω < ω0 + ω1 to obtain a consistent definition of signal-to-noise ratio
(SNR). Therefore, the integral of (10.47) is carried out only over ω0 − ω1 < ω < ω0 + ω1 .
ω
w0 − w1 w0 w0 + w1
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 391
To compare the four adaptive array processors of Figure 10-15, the output SNR
performance is evaluated for the aforementioned signal conditions. Assume the element
thermal noise power pn is 10 dB below the element desired signal power ps so that ps / pn =
10 dB. Furthermore, suppose that the element interference signal power pi is 20 dB stronger
than the element desired signal power so that ps / pi = −20 dB. Now assume that the
desired signal is incident on the array from broadside. The output SNR given by (10.46) can
be evaluated from (10.47), (10.48), and (10.49) by assuming the processor weights satisfy
(10.16) for each of the four processor configurations. The resulting output signal-to-total
noise ratio that results using each processor is plotted in Figures 10-19–10-22 as a function
of the interference angle of arrival for 4, 10, 20, and 40% bandwidth signals, respectively.
In all cases, regardless of the signal bandwidth, when the interference approaches
broadside (near the desired signal) the SNR degrades rapidly, and the performance of
15 FIGURE 10-19
Pi Output signal-to-
= 20 dB
Ps 3 & 5 Taps interference plus
Pn noise ratio
= −10 dB
10 Ps interference angle
for four adaptive
processors with 4%
3 Taps bandwidth signal.
(dB)
−5
−10
15 FIGURE 10-20
Pi Output signal-to-
= 20 dB
Ps 3 & 5 Taps interference plus
Pn noise ratio versus
= −10 dB
10 Ps interference angle
for four adaptive
(dB)
processors with
10% bandwidth
PI + PN
Compton, IEEE
Quadrature hybrid Trans. Aerosp.
0 Electron. Syst.,
20° 40° 60° 80°
January 1979 [13].
Interference angle
−5
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 392
FIGURE 10-21 15
Output signal-to- Pi
= 20 dB 3 & 5 Taps
interference plus Ps
noise ratio versus Pn
= −10 dB
interference angle 10 Ps
3 Taps
for four adaptive
5 Taps
processors with
2 Taps
20% bandwidth
(dB)
Quadrature hybrid
signal. From 5
PI + PN
Rodgers and
Compton, IEEE Pd
Trans. Aerosp. Output
Electron. Syst., 0
20° 40° 60° 80°
January 1979 [13].
Interference angle
−5
−10
FIGURE 10-22 15
Output signal-to- Pi
= 20 dB
interference plus Ps
noise ratio versus Pn
= −10 dB
interference angle 10 Ps
3 Taps
for four adaptive
5 Taps
processors with
2 Taps
40% bandwidth
(dB)
Quadrature hybrid
signal. From 5
PI + P N
Rodgers and
Pd
Compton, IEEE
Trans. Aerosp.
Output
Electron. Syst., 0
20° 40° 60° 80°
January 1979 [13].
Interference angle
−5
−10
all four processors becomes identical. This SNR degradation is expected since, when the
interference approaches the desired signal, the desired signal falls into the null provided
to cancel the interference, and the output SNR consequently falls. Furthermore, as the
interference approaches broadside, the interelement phase shift for this signal approaches
zero. Consequently, the need to provide a frequency-dependent phase shift behind each
array element to deal with the interference signal is less, and the performance of all four
processors becomes identical.
When the interference signal is widely separated from the desired signal, then the
output SNR is different for the four processors being considered, and this difference
becomes more pronounced as the bandwidth increases. For 20 and 40% bandwidth signals,
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 393
for example, neither the quadrature hybrid processor nor the two-tap delay line proces-
sor provides good performance as the interference signal approaches endfire. The per-
formance of both the three- and five-tap delay line processors remains quite good in the
endfire region, however. If 20% or more bandwidth signals are accommodated, then tapped
delay line processing becomes a necessity. Figure 10-22 shows that there is no significant
performance advantage provided by the five-tap processor compared with the three-tap
processor, so a three-tap processor is adequate for up to 40% bandwidth signals in the case
of a two-element array.
Figures 10-21 and 10-22 show that the output SNR performance of the two-tap
delay line processor peaks when the interference signal is 30◦ off broadside, because
the interelement delay time is λ/4 (since the elements are spaced apart by λ/2). Conse-
quently, the single-delay element value of λ/4 provides just the right amount of time delay
to compensate exactly for the interelement time delay and to produce an improvement in
the output SNR.
The three-tap and five-tap delay line processors both produce a maximum SNR of
about 12.5 dB at wide interference angles of 70◦ or greater. For ideal channel processing,
the interference signal is eliminated, the desired signal in each channel is added coherently
to produce Pd = 4 ps , and the thermal noise is added noncoherently to yield PN = 2 pn .
Thus, the best possible theoretical output SNR for a two-element array with thermal noise
10 dB below the desired signal and no interference is 13 dB. Therefore, the three-tap and
five-tap delay line processors are successfully rejecting nearly all the interference signal
power at wide off-boresight angles.
−90
−105
−120
−135
0.98 0.984 0.988 0.992 0.996 1.001 1.004 1.008 1.012 1.016 1.02
Frequency (w/w 0)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 394
Amplitude (dB)
Report ESL 3832-3,
1975 [12]. −60
−75
−90
−105
−120
−135
0.98 0.984 0.988 0.992 0.996 1.001 1.004 1.008 1.012 1.016 1.02
Frequency (w/w 0)
Figure 10-15 with a 4% signal bandwidth and various interference signal angles. The
results shown in these figures indicate that for all four processors and for all interference
angles the desired signal response is quite flat over the signal bandwidth. As the interfer-
ence approaches the desired signal angle at broadside, however, the (constant) response
level of the array to the desired signal drops because of the desired signal partially falling
within the array pattern interference null.
The results in Figure 10-23 for quadrature hybrid processing show that the array
response to the interference signal has a deep notch at the center frequency when the
interference signal is well separated (θi > 20◦ ) from the desired signal. As the interference
signal approaches the desired signal (θi < 20◦ ), the notch migrates away from the center
frequency, because the processor weights must compromise between rejection of the
interference signal and enhancement of the desired signal when the two signals are close.
Migration of the notch improves the desired signal response (since the desired signal power
spectral density peaks at the center frequency) while affecting interference rejection only
slightly (since the interference signal power spectral density is constant over the signal
band).
The array response for the two-tap processor is shown in Figure 10-24. The response
to both the desired and interference signals is very similar to that obtained for quadrature
hybrid processing. The most notable change is the slightly different shape of the transfer
function notch presented to the interference signal by the two-tap delay line processor
compared with the quadrature hybrid processor.
Figure 10-25 shows the three-tap processor array response. The interference signal
response is considerably reduced, with a minimum rejection of the interference signal
of about 45 dB. When the interference signal is close to the desired signal, the array
response has a single mild dip. As the separation angle between the interference signal
and the desired signal increases, the single dip becomes more pronounced and finally
develops into a double dip at very wide angles. It is difficult to attribute much significance
to the double-dip behavior since it occurs at such a low response level (of more than
75 dB attenuation). The five-tap processor response of Figure 10-26 is very similar to the
three-tap processor response except slightly more interference signal rejection is achieved.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 395
−75 90°
−90
−105
−120
−135
0.98 0.984 0.988 0.992 0.996 1.001 1.004 1.008 1.012 1.016 1.02
Frequency (w/w 0)
20°
−60 1975 [12].
50°
−75
90°
−90
−105
−120
−135
0.98 0.984 0.988 0.992 0.996 1.001 1.004 1.008 1.012 1.016 1.02
Frequency (w/w 0)
As the signal bandwidth increases, the processor response curves remain essentially
the same as in Figures 10-23–10-26 except the following:
1. As the interference signal bandwidth increases, it becomes more difficult to reject the
interference signal over the entire bandwidth, so the minimum rejection level increases.
2. The desired signal response decreases because the array feedback reduces all weights
to compensate for the presence of a greater interference signal component at the array
output, thereby resulting in greater desired signal attenuation.
The net result is that as the signal bandwidth increases, the output SNR performance
degrades, as confirmed by the results of Figures 10-19–10-22.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 396
where u = sin θ, θ is the angle of arrival, and the matrix C describes the effects of mutual
coupling and is independent of the signal scan angle. If the array is composed of multimode
elements, then the matrix C would be scan angle dependent.
It follows that the unperturbed signal vector, vd can be recovered from the perturbed
signal vector by introducing compensation for the mutual coupling
vd = C−1 v (10.57)
Introducing the compensation network C−1 as shown in Figure 10-27 then allows all
subsequent beamforming operations to be performed with ideal (unperturbed) element
signals, as are customarily assumed in pattern synthesis.
This mutual coupling compensation is applied to an eight-element linear array having
element spacing d = 0.517 λ consisting of identical elements. Figure 10-28(a) shows the
effects of mutual coupling by displaying the difference in element pattern shape between
a central and an edge element in the array.
Figure 10-28 displays a synthesized 30 dB Chebyshev pattern both without (a) and with
(b) mutual coupling compensation. It is apparent from this result that the compensation
network gives about a 10 dB improvement in the sidelobe level.
WN
N υN υNd
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 397
0
Measured
Theory
−10
−20
Power, dB
−30
−40
0
Measured
Theory
−10
−20
Power, dB
−30
−40
FIGURE 10-28 30 dB Chebyshev pattern (a) without and (b) with Coupling Compensation
with a Scan Angle of 0◦ . From Steyskal & Herd, IEEE Trans. Ant. & Prop. Dec. 1995.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 398
w2
w1
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 399
Let x0 (t), x1 (t), and e(t) represent the complex envelope signals of the main channel
input signal, the auxiliary channel input signal, and the output residue signal, respectively.
Define the complex signal vector
xT = [x1 (t), x2 (t), . . . , x L (t)] (10.58)
where
x2 (t) = x1 (t − )
..
.
x L (t) = x1 [t − (L − 1)]
it follows that
Minimize (10.66) by appropriately selecting the complex weight vector w. Assume the
matrix Rxx (0) is nonsingular: the value of w for which this minimum occurs is given by
wopt = R−1
xx (0)rx x0 (−D) (10.67)
where ω0 is the center frequency of the interference signal. It then follows that
where τ12 and τ22 represent the propagation delay between the main channel element
and the auxiliary channel element for the wavefronts of s(t, θ1 ) and sm (t, ρm , dm , θ2 ),
respectively.
Assuming the signals s(t, θ1 ) and sm (t, ρm , Dm , θ2 ) possess flat spectral density
functions over the bandwidth B, as shown in Figure 10-30a, then the corresponding
auto- and cross-correlation functions of x0 (t) and x1 (t) can be evaluated by recognizing
that
where −1 {·} is the “inverse Fourier transform,” and xx (ω) denotes the cross-spectral
density matrix of x(t).
From (10.74), (10.76), and (10.77) it immediately follows that
sin π BDm
r x0 x0 (0) = 1 + |ρm |2 + ρm e− jω0 Dm + ρm∗ e jω0 Dm (10.78)
π BDm
sin π B [ψ + sgn1 · (i − 1) + sgn2 · D]
Likewise, defining f [ψ, sgn1, sgn2] =
π B [ψ + sgn1 · (i − 1) + sgn2 · D]
sin π B [ψ + sgn · (i − k)]
and g[ψ, sgn] = , then
π B[ψ + sgn · (i − k)]
r xi x0 (−D) = f [τ12 , +, −] exp{− jω0 [τ12 + (i − 1)]}
+ f [Dm + τ22 , +, −]ρm exp{− jω0 [τ22 + (i − 1) + Dm ]} (10.79)
∗
+ f [Dm − τ12 , −, +]ρm exp{− jω0 [τ12 + (i − 1) − Dm ]}
+ f [τ22 , +, −]|ρm |2 exp{− jω0 [τ22 + (i − 1)]}
t
0
1 2
B B
(b)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 402
The quantities τ12 and τ22 are related to the CSLC array geometry by
⎫
d
τ12 = sin θ1 ⎪
⎬
(10.83)
d ⎪
τ22 = sin θ2 ⎭
where
d = interelement array spacing
= wavefront propagation speed
θ1 = angle of incidence of direct ray
θ2 = angle of incidence of multipath ray
Referring to (10.76), (10.79), and (10.80), we see that the parameters ω0 , τ12 , τ22 ,
Dm , and enter the evaluation of the output residue power in the form of the products
ω0 τ12 , ω0 τ22 , ω0 Dm , and ω0 . These products represent the phase shift experienced at the
center frequency ω0 as a consequence of the four corresponding time delays. Likewise, the
parameters B, D, Dm , τ12 , τ22 , and enter the evaluation of the output residue power in
the form of the products BD, BDm , Bτ12 , Bτ22 , and B; these time–bandwidth products
are phase shifts experienced by the highest frequency component of the complex envelope
interference signal as a consequence of the five corresponding time delays. Both the intertap
delay and the multipath delay Dm are important parameters that affect the CSLC system
performance through their corresponding time–bandwidth products; thus, the results are
given here with the time–bandwidth products taken as the fundamental quantity of interest.
Since for this example θ1 = −θ2 , the product ω0 τ12 is specified as
π ⎫
ω0 τ12 = ⎪
⎪
4 ⎬
then the product (10.85)
π⎪⎪
ω0 τ22 = − ⎭
4
Furthermore, let the products ω0 Dm and ω0 be given by
ω0 Dm = 0 ± 2kπ, k any integer
(10.86)
ω0 = 0 ± 2lπ, l any integer
For the element spacing d = 2.25λ0 and θ1 = 30◦ , then specify
1
Bτ12 = −Bτ22 = , P = 72 (10.87)
P
Finally, specifying the multipath delay time to correspond to 46 meters yields
Since
N −1
D= (10.89)
2
Only N and B need to be specified to evaluate the output residue power by way of
(10.68).
To evaluate the output residue power by way of (10.68) resulting from the array
geometry and multipath conditions specified by (10.84)–(10.89) requires that the cross-
correlation vector rx x0 (−D), the N × N autocorrelation matrix Rxx (0), and the autocorre-
lation function r x0 x0 (0) be evaluated by way of (10.78)–(10.80). A computer program to
evaluate (10.68) for the multipath conditions specified was written in complex, double-
precision arithmetic.
Figure 10-31 shows a plot of the output residue power where the resulting minimum
possible value of canceled power output in dB is plotted as a function of B for various
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 404
FIGURE 10-31 0
Decibel cancellation N = number of taps
N=1
versus B for
multipath. −10
N=3
Cancellation (dB)
−20
−30 N=5
Evaluated
with
1
Bτ =
−40 78
N=7 BDm = 0.45
Rm = 0.5
−50
0.1 0.2 0.3 0.4 0.6 0.8 1
BΔ
specified values of N. It will be noted in Figure 10-31 that for N = 1 the cancellation
performance is independent of B since no intertap delays are present with only a single
tap. As explained in Appendix B, the transfer function of the tapped delay line transversal
filter has a periodic structure with (radian) frequency period 2π B f , which is centered at
the frequency f 0 . It should be noted that the transversal filter frequency bandwidth B f
is not necessarily the same as the signal-frequency bandwidth B. The transfer function
of a transversal filter within the primary frequency band (| f − f 0 | < B f /2) may be
expressed as
N
F( f ) = [Ak e jφk ] exp[− j2π(k − 1)δ f ] (10.90)
k=1
Bf ≥ B (10.92)
Figure 10-31 shows that, as B decreases from 1, for values of N > 1 the cancellation
performance rapidly improves (the minimum canceled residue power decreases) until
B = BDm (0.45 for this example), after which very little significant improvement
occurs. As B becomes very much smaller than BDm (approaching zero), the cancellation
performance degrades since the intertap delay is effectively removed. The simulation could
not compute this result since as B approaches zero the matrix Rxx (0) becomes singular
and matrix inversion becomes impossible. Cancellation performance of −30 dB is virtually
assured if the transversal filter has at least five taps and is selected so that = Dm .
Suppose for example that the transversal filter is designed with B = 0.45. Using the
same set of selected constants as for the previous example, we find it useful to consider what
results would be obtained when the actual multipath delay is different from the anticipated
value corresponding to BDm = 0.45. From the results already obtained in Figure 10-31, it
may be anticipated that, if BDm > B, then the cancellation performance would degrade.
If, however, B Dm B, then the cancellation performance would improve since in the
limit as Dm → 0 the system performance with no multipath present would result.
5
N=
−60
−70
0.1 0.2 0.3 0.4 0.6 0.8 1
BΔ
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 406
d w1 Adaptive
electronics
q
S(w)
Auxiliary channel
T1 (w, q )
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 407
w1 w1 wN+ 1 w 2N +1
Auxiliary
channel e−jwΔ e−jw Δ e−jw Δ
N
F(ω) =
k = −N
Σ
wN+ 1 + k e−jwkΔ
function, A(ω). Assume for analysis purposes that all channel distortion is confined to the
main channel and that T1 (ω, θ) = 1. The transversal filter transfer function, F(ω), can be
expressed as
N
F(ω) = w N +1+k e− jωk (10.96)
k=−N
E{[A0 (ω) exp[− jφ0 (ω)] − F(ω)] exp( jωk )} = 0 for k = −N , . . . , 0, . . . , N
(10.99)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 408
(10.100)
Note that
sin [π B(l − k)]
E{exp[− jω(l − k)]} = (10.101)
π B(l − k)
it follows that
" N #
N
sin[π B(l − k)]
E W N +1+l exp(− jωl) exp( jωl) = W N +1+l
l=−N l=−N
π B(l − k)
(10.102)
so that (10.100) can be rewritten in matrix form as
v = Cw (10.103)
where
vk = E{A0 (ω) exp[ j (ωk − φ0 (ω))]} (10.104)
sin[π B(l − k)]
Ck,l = (10.105)
π B(l − k)
Consequently, the complex weight vector must satisfy the relation
w = C−1 v (10.106)
Using (10.106) to solve for the optimum complex weight vector, we can find the output
residue signal power by using
πB
1
Ree (0) = |A(ω) − F(ω)|2 φJJ (ω)dω (10.107)
2π B −π B
where φJJ (ω) is the constant interference signal power spectral density. Assume the inter-
ference power spectral density is unity across the bandwidth of concern; then the output
residue power due only to main channel amplitude variations is given by
πB
1
Ree A = |A0 (ω) − F(ω)|2 dω (10.108)
2π B −π B
Since A(ω) − F(ω) is orthogonal to F(ω), it follows that [15]
and hence
πB
1
Ree A = [A20 (ω) − |F(ω)|2 ] dω (10.110)
2π B −π B
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 409
It likewise follows from (10.107) that the output residue power contributed by main
channel phase variations is given by
πB
1
Ree p = |e− jφ0 (ω) − F(ω)|2 φJJ (ω) dω (10.111)
2π B −π B
where φ0 (ω) represents the main channel phase variation. Once again assuming that
the input signal spectral density is unity across the signal bandwidth and noting that
[e− jφ0 (ω) − F(ω)] must be orthogonal to F(ω), it immediately follows that
πB
1
Ree p = [1 − |F(ω)|2 ] dω
2π B −π B
N N
sin[π B(k − j)]
=1− w k w ∗j (10.112)
j=−N k=−N
π B(k − j)
where the complex weights used to obtain F(ω) must again satisfy (10.102)–(10.106),
which now involve both a magnitude and a phase component and it is assumed that φJJ (ω)
is a constant.
ω
−π B 0 πB
Array bandwidth
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 410
where
2n + 1
T0 = for n = 0, 1, 2, . . .
2B
and the integer n corresponds to (2n + 1)/2 cycles of amplitude mismatching across the
bandwidth B. Letting the phase error φ0 (ω) = 0, it follows from (10.104) that
πB
1
vk = [1 + R cos ωT0 ]e jωk dω (10.115)
2π B −π B
or
sin(π Bk) R sin(π B[T0 + k]) sin(π B[T0 − k])
vk = + +
π Bk 2 π B [T0 + k] B[T0 − k]
for k = −N , . . . , 0, . . . , N (10.116)
Evaluation of (10.116) permits the complex weight vector to be found, which in turn may
be used to determine the residue power by way of (10.110).
Now
where
⎡ ⎤
e jωN
⎢ e jω(N −1) ⎥
⎢ ⎥
β = ⎢. ⎥ (10.118)
⎣ .. ⎦
e− jωN
Carrying out the vector multiplications indicated by (10.117) then yields
2N +1 2N
+1
|F(ω)|2 = w i w k∗ e jω(k−i) (10.119)
i=1 k=1
−40
−50
−60
1
1 1 Cycles
Cycle 2
2
−70
0 5 10 15
N = number of taps
B = 0.5.
−40
−50
1
2 Cycles
2
−60
1
1 1 Cycles
Cycle 2
2
−70
0 5 10 15
N = number of taps
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 412
Cancellation (dB)
B = 0.75.
−40
−50
1
2 Cycles
2
−60
1 1
Cycle 1 Cycles
2 2
−70
0 5 10 15
N = number of taps
B = 1.0. 1
3 Cycles
2
−40
1
2 Cycles
2
1
−50 1 Cycles
2
−60
1
Cycle
2
−70
0 5 10 15
N = number of taps
the amplitude mismatch model. The sufficient number of taps for the selected amplitude
mismatch model was found empirically to be given by
Nr − 1
Nsufficient ≈ [7 − 4(B)] + 1 (10.123)
2
where Nr is the number of half-cycles of ripple appearing in the mismatch model.
If there are a sufficient number of taps in the transversal filter, the cancellation perfor-
mance improves when more taps are added depending on how well the resulting transfer
function of the transversal filter matches the gain and phase variations of the channel
mismatch model. Since the transversal filter transfer function resolution depends in part
on the product B, a judicious selection of this parameter ensures that providing addi-
tional taps provides a better match (and hence a significant improvement in cancellation
performance), whereas a poor choice results in very poor transfer function matching even
with the addition of more taps.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 413
0 FIGURE 10-40
−10 For N = 3 and Rm = 0.9 Decibel cancellation
−20 versus B for
one-half-cycle
amplitude mismatch
−50 model.
Cancellation (dB)
−100
−150
0 0.5 1 1.5
BΔ
Taking the inverse Fourier transform of (10.114) −1 {A(ω)} yields a time function
corresponding to an autocorrelation function f (t) that can be expressed as
f (t) = s(t) + Ks(t ± T0 ) (10.124)
The results of Section 10.5.3 and equation (10.124) imply that = T0 (or equivalently,
B = number cycles of ripple mismatch) if the product B is to “match” the amplitude
mismatch model. This result is illustrated in Figure 10-40 where decibel cancellation is
plotted versus B for a one-half-cycle ripple mismatch model. A pronounced minimum
occurs at B = 12 for N = 3 and Rm = 0.9.
When the number of cycles of mismatch ripple exceeds unity, the foregoing rule
of thumb leads to the spurious conclusion that B should exceed unity. Suppose, for
example, there were two cycles of mismatch ripple for which it was desired to compensate.
By setting B = 2 (corresponding to B f = 12 B), two complete cycles for the transversal
filter transfer function are found to occur across the cancellation bandwidth. By matching
only one cycle of the channel mismatch, quite good matching of the entire mismatch
characteristic occurs but at the price of sacrificing the ability to independently adjust the
complex weights across the entire cancellation bandwidth, thereby reducing the ability
to appropriately process broadband signals. Consequently, if the number of cycles of
mismatch ripple exceeds unity, it is usually best to set B = 1 and to accept whatever
improvement in cancellation performance can be obtained with that value, or increase the
number of taps.
Since
πB
1
vk = exp( j{A cos ωT0 + ωk}) dω for k = −N , . . . , 0, . . . , N
2π B −π B
(10.126)
it can easily be shown by defining
k=1
2
1
B = 0.2. 2
2
Cycles
−40
−50
−60
1
Cycle
2
−70
0 1 3 5 7
N = number of taps
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 415
−50
−60
1
Cycle
2
−70
0 1 3 5 7
N = number of taps
−30
B = 1.0.
1
2 Cycles
−40 2
−50
1
Cycle
2
−60
0 1 3 5 7
N = number of taps
10.8 PROBLEMS
Distortionless Transfer Functions
1. From (10.11) and (10.12) it immediately follows that |H1 (ω) = |H2 (ω)|, thereby yielding the
pair of equations
f 1 {|H1 |, α1 , α2 , θs } = exp(− jωT1 )
and
f 2 {|H1 |, α1 , α2 , θi } = 0
(a) Show from the previous pair of equations that α1 (ω) and α2 (ω) must satisfy
πω
α2 (ω) − α1 (ω) = sin θi ± nπ
ω0
where n is any odd integer.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 417
(b) Since the magnitude of exp(− jωT1 ) must be unity, show using f 1 { } = exp(− jωT1 ) that
(10.13) results.
(c) Show that the angle condition associated with f 1 { } = exp(− jωT1 ) yields (10.14).
(d) Show that substituting (10.14) into the results from part (a) yields (10.15).
2. For a three-element linear array, the overall transfer function encountered by the desired signal
in passing through the array is
ωd ω2d
Hd (ω) = H1 (ω) + H2 (ω) exp − j sin θs + H3 (ω)e − j sin θs
c c
What does imposing the requirements (10.9) and (10.10) now imply for the three-channel
transfer functions?
Hilbert Transform Relations
3. Prove the Hilbert transform relations given by (10.25)–(10.28).
4. Using (10.61), (10.62), and the results of (10.63)–(10.65), show that Ree is given by (10.66).
5. Derive the correlation functions given by (10.78)–(10.80) for the signal environment assump-
tions (10.75) and (10.76)
6. Show that as the time–bandwidth product B approaches zero, then the matrix Rxx (0) [whose
elements are given by (10.80)] becomes singular so that matrix inversion cannot be accom-
plished.
Compensation for Channel Phase Errors
7. For the phase error φ(ω) given by (10.125), show that vk given by (10.127) follows from the
application of (10.126).
8. Let φ(ω) correspond to the phase error model be given by
& '
A 1 − cos 2ω for |ω| ≤ π B
φ(ω) = B
0 otherwise
πB
2 2
vk = cos A 1 − cos ω + j sin A 1 − cos ω exp( jωkdω )
−π B B B
∞
cos(A cos ωT0 ) = J0 (A) + 2 (−1)k · J2k (A) cos[(2k)ωT0 ]
k=1
∞
sin(A cos ωT0 ) = 2 (−1)k J2k+1 (A) · cos[(2k + 1)ωT0 ]
k=0
where Jn (·) denotes a Bessel function of the nth order and define
to show that
Letting u = ω/π B, applying Euler’s formula, and ignoring all odd components of the resulting
expression, show that
1
27A 2
vi = exp j u (1 − u) cos π[u(i − (N + 1))B]du
0 4
where A = 4b(π B/3)3 for i = 1, 2, . . . , 2N +1. The foregoing equation for vi can be evaluated
numerically to determine the output residue power contribution due to the previous phase error
model.
Computer Simulation Problems
10. A 30-element linear array (d = 0.5λ) has a 20 dB, n = 2 Taylor taper applied at the elements.
Plot the array factor when δna = 0.1 and δna = 0.1.
11. A 30-element linear array (d = 0.5λ) has a 30 dB, n = 7 low sidelobe taper. Plot the array
factors for a single element failure at (1) the edge and (2) the center of the array.
12. Find the location and heights of the quantization lobes for a 20-element array with d = 0.5λ
and the beam steered to θ = 3◦ when the phase shifters have three, four, and five bits.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 419
10.9 REFERENCES
[1] E. Brookner, Antenna Array Fundamentals—Part 2, Practical Phased Array Antenna Systems,
E. Brookner, ed., Norwood, MA, Artech House, 1991
[2] R.J. Mailloux, “Array Grating Lobes Due to Periodic Phase, Amplitude, and Time Delay
Quantization,” IEEE AP-S Trans., Vol. 32, December 1984, pp. 1364–1368.
[3] A. J. Boonstra, and A. J. van der Veen, “Gain Calibration Methods for Radio Telescope
Arrays,” IEEE Signal Processing Trans., Vol. 51, No. 1, 2003, pp. 25–38.
[4] R. Sorace, “Phased Array Calibration,” IEEE AP-S Trans., Vol. 49, No. 4, 2001, pp. 517–525.
[5] T. Takahashi, Y. Konishi, S. Makino, et al., “Fast Measurement Technique for Phased Array
Calibration,” IEEE AP-S Trans., Vol. 56, No. 7, 2008, pp. 1888–1899.
[6] M. Borkowski, “Solid-State Transmitters,” Radar Handbook, M. I. Skolnick, ed.,
pp. 11.1–11.36, New York, McGraw Hill, 2008.
[7] N. Tyler, B. Allen, and H. Aghvami, “Adaptive Antennas: The Calibration Problem,” IEEE
Communications Magazine, Vol. 42, No. 12, 2004, pp. 114–122.
[8] N. Tyler, B. Allen, and A. H. Aghvami, “Calibration of Smart Antenna Systems: Measurements
and Results,” IET Microwaves, Antennas & Propagation, Vol. 1, No. 3, 2007, pp. 629–638.
[9] R. L. Haupt, Antenna Arrays: A Computational Approach, New York, Wiley, 2010.
[10] W. T. Patton and L. H. Yorinks, “Near-Field Alignment of Phased-Array Antennas,” IEEE
AP-S Trans., Vol. 47, No. 3, 1999, pp. 584–591.
[11] W. E. Rodgers and R. T. Compton, Jr., “Tapped Delay-Line Processing in Adaptive Arrays,”
Report 3576-3, April 1974, prepared by The Ohio State University Electro Science Labora-
tory, Department of Electrical Engineering under Contract N00019-73-C-0195 for Naval Air
Systems Command.
[12] W. E. Rodgers and R. T. Compton, Jr., “Adaptive Array Bandwidth with Tapped Delay-
Line Processing,” Report 3832-3, May 1975, prepared by The Ohio State University Electro
Science Laboratory, Department of Electrical Engineering under Contract N00019-74-C-0141
for Naval Air Systems Command.
[13] W. E. Rodgers and R. T. Compton, Jr., “Adaptive Array Bandwidth with Tapped Delay-
Line Processing,” IEEE Trans. Aerosp. Electron. Syst., Vol. AES-15, No. 1, January 1979,
pp. 21–28.
[14] T. G. Kincaid, “The Complex Representation of Signals,” General Electric Report No.
R67EMH5, October 1966, HMED Publications, Box 1122 (Le Moyne Ave.), Syracuse, NY,
13201.
[15] A. Papoulis, Probability, Random Variables, and Stochastic Processes, New York,
McGraw-Hill, 1965, Ch. 7.
[16] H. Steyskal & J. Herd, “Mutual Coupling Compensation in Small Array Antennas,” IEEE
Trans. Ant. & Prop., Vol. AP-38, No. 12, December 1995, pp. 603–606.
[17] R. A. Monzingo, “Transversal Filter Implementation of Wideband Weight Compensation
for CSLC Applications,” unpublished Hughes Aircraft Interdepartmental Correspondence
Ref. No. 78-1450.10/07, March 1978.
[18] A. M. Vural, “Effects of Perturbations on the Performance of Optimum/Adaptive Arrays,”
IEEE Trans. Aerosp. Electron. Syst., Vol. AES-15, No. 1, January 1979, pp. 76–87.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:46 420
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:52 421
CHAPTER
An array’s ability to resolve signals depends on the beamwidth of the array, so high-
resolution algorithms have been developed, in order for small arrays can resolve closely
spaced signals by using narrow nulls in place of the wide main beam. A linear array with
elements separated by half-wavelength spacing has N − 1 nulls in which to locate up to
N − 1 signals.
In this chapter, we apply maximum likelihood (ML) estimation methods to estimate the
direction of arrival (DOA), or angle of arrival (AOA), of one or more signal sources, using
data received by the elements of an N-element antenna array. The Cramer–Rao (CR) lower
bound on angle estimation error is derived under several different signal assumptions. The
CR bound helps determine system performance versus signal-to-noise ratio (SNR) and
array size.
The advantage of optimal array estimation methods (processing the antenna element
signals in an optimum fashion) lies in its application to multiple signal environments
and in conditions where interfering signals are present. If only one signal is present in
a white noise background, conventional monopulse processing achieves the same AOA
estimation accuracy. High angular resolution of desired signals is achieved beyond the one
421
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:52 422
11.1 PERIODOGRAM
The simplest approach to finding the direction of a signal is to scan the main beam of
the array by adjusting the steering vector until the signal is detected. The relative output
power of a linear array lying along the x-axis is given by
and RT is the signal plus noise correlation matrix where the subscript “T ” denotes where
the uniform array steering vector is given by
T
A(θ ) = e jkx1 cos θ · · · e jkx N cos θ , θmin ≤ θ ≤ θmax (11.2)
A periodogram is a plot of the output power versus angle, where a window function
that is independent of the data being analyzed must be adopted. By weighting all angles
equally, a rectangular window function is in effect being adopted. Peaks in the periodogram
correspond to signal locations. Large arrays have a narrower beamwidth than smaller arrays
and can resolve closely spaced signals better. Figure 11-1 shows the periodigram for a
12-element array with λ/2 spacing and three sources incident at θ = −50◦ , 10◦ , and 20◦ .
The source at θ = −50◦ is easy to distinguish, but the sources at θ = 10◦ , and 20◦ appear
to be a single source, because the beamwidth is too wide. This example demonstrates the
need for super-resolution techniques.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:52 423
−10
−20
−90 −45 0 45 90
q (degrees)
R−1
xx A
w= (11.3)
A R−1
†
xx A
Figure 11-2 shows the periodigram for a 12-element array with λ/2 spacing and three
sources incident at θ = −50◦ , 10◦ , and 20◦ . Capon’s method distinguishes between two
closely spaced sources much better than the periodogram.
|P (dB)|
−10
−20
−90 −45 0 45 90
q (degrees)
−10
−20
−90 −45 0 45 90
q (degrees)
The signal plus noise correlation matrix in the denominator of (11.3) is replaced by Vλ V†λ
where the columns of Vλ are the eigenvectors of the noise subspace. In the numerator,
A† (θ) replaces R−1
x x and corresponds to the N − Ns smallest eigenvalues of the correlation
matrix. Figure 11-3 shows the MUSIC spectrum for a 12-element array with λ/2 spacing
and three sources incident at θ = −50◦ , 10◦ , and 20◦ . The MUSIC spectrum is similar
to the Capon spectrum, except the floor between the peaks is much lower for the MUSIC
spectrum.
The root-MUSIC algorithm is a more robust alternative that accurately locates the
direction of arrival by finding the roots of the array polynomial that corresponds to the
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:52 425
†
M−1
†
A (θ)Vλ Vλ A(θ ) = cn z n (11.6)
n=M+1
where
z = e j λ nd sin θ
2π
cn = Vλ V†λ
r −c=n
The cn results from summing the diagonal n = r − c of Vλ Vλ† in (11.6) with r and c
indicating the row and column, respectively, of the matrix. Polynomial roots, z m , on the
unit circle (Chapter 2) are the poles of the MUSIC spectrum. The phase of the polynomial
roots in (11.6) are given by
−1 arg(z m )
θm = sin (11.7)
kd
Roots on the unit circle correspond to the signals. Roots off the unit circle are spurious.
The 2N − 1 diagonals of Vλ Vλ† form a polynomial with 2N − 2 roots. Table 11-1 contains
the roots of the polynomial for a 12-element array with λ/2 spacing and three sources
incident at θ = −50◦ , 10◦ , and 20◦ . Roots on the unit circle correspond to signals and have
a “yes” in column 3. Spurious roots are off the unit circle and have a “no” in column 3.
All roots appear in the unit circle plot in Figure 11-4. Note that each root on the unit circle
is actually a double root (see Table 11-1), so it appears that there are only 19 roots in
Figure 11-4 when there are actually 22 roots.
A unitary (real-valued) root-MUSIC algorithm reduces the computational complexity
of the root-MUSIC algorithm by exploiting the eigen decomposition of a real-valued
correlation matrix. Unitary root MUSIC improves threshold and asymptotic performances
relative to conventional root MUSIC.
Real
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:52 426
TABLE 11-1 The roots found using root MUSIC. The ones
close to the unit circle represent the correct signal directions.
under the constraint that φxx ( f ) satisfies a set of N linear measurement equations
W
φxx ( f )G n ( f )d f = gn , n = 1, . . . , N (11.9)
−W
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:52 427
−10
−20
−90 −45 0 45 90
q (degrees)
where the time series is sampled with the uniform period t so the Nyquist fold-over
frequency is W = 1/2t, and the power spectrum of the time series is bandlimited to
±W . The functions G n ( f ) in the measurement equations are known test functions, and
the gn are the observed values resulting from the measurements.
MEM is based on a rational function model of the spectrum that has only poles and
not zeros [18]. The MEM spectrum is given by [19]
1
P(θ) = (11.10)
A† (θ )R−1 [:, n]R†−1 [:, n]A(θ )
xx xx
where n corresponds to the nth column of the inverse correlation matrix. Results depend
on which n is chosen. Figure 11-5 shows the MEM spectrum for a 12-element array with
λ/2 spacing and three sources incident at θ = −50◦ , 10◦ , and 20◦ . Very sharp peaks in
the spectrum occur in the signal directions.
Two cases can now be considered: the first where the autocorrelation function is
partially known; and the second where the autocorrelation function is unknown.
ε N = aTN x (11.15)
where xT = [x N , x N −1 , . . . , x0 ], and aTN = [1, a(N , 1), a(N , 2), . . . , a(N , N )], where
the coefficient a(N , N ) = C N is called the reflection coefficient of order N . The error
ε N is regarded as the output of an Nth order prediction error filter whose coefficients are
given by the vector a N and whose power output is
PN = E ε2N (11.16)
Furthermore, the error ε N must be uncorrelated with all past estimation errors so that
Equations (11.16) and (11.8) are the conditions required for a random process to have
a white power spectrum of total power PN (or a power density level of PN /2W where
W = 1/2t). The prediction error filter is regarded as a whitening filter that operates
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:52 429
FIGURE 11-6
Input time series Whitening Output time series Whitening filter
[Spectrum fxx ( f )] filter (white spectrum) representation of
prediction error filter.
on the input data {x0 , x1 , . . . , x N −1 } to produce output data having a white power density
spectrum of level PN /2W as indicated in Figure 11-6. It immediately follows that the
estimate φ̂xx ( f ) of the input power spectrum φxx ( f ) is given by
PN /2W
φ̂xx ( f ) = 2 (11.19)
N
1 + a(N , n) exp(− j2πfnt)
n=1
where the denominator of (11.19) is recognized as the power response of the predic-
tion error filter. Equation (11.19) yields the MEM estimate of φxx ( f ) provided that the
coefficients a(N , n), n = 1, . . . , N , and the power PN can be determined.
A relationship among the coefficients a(N , n), n = 1, . . . , N , the power PN , and
the autocorrelation function values r (−N ), r (−N + 1), . . . , r (0), . . . , r (N − 1), r (N ) is
provided by the well-known prediction error filter matrix equation [14]
⎡ ⎤⎡ ⎤ ⎡ ⎤
r (0) r (−1) ... r (−N ) 1 PN
⎢ r (1) r (0) . . . r (−N + 1) ⎥ ⎢ ⎥ ⎢ ⎥
⎢ ⎥ ⎢ a(N , 1) ⎥ ⎢ 0 ⎥
⎢ .. .. .. ⎥⎢ .. ⎥ = ⎢ .. ⎥ (11.20)
⎣ . . . ⎦⎣ . ⎦ ⎣ . ⎦
r (N ) r (N − 1) . . . r (0) a(N , N ) 0
Equation (11.20) can be derived [as done in [16] by maximizing the entropy of (11.8)]
subject to the constraint equations (11.12). If we know the autocorrelation values {r (−N ),
r (−N + 1), . . . , r (−1), r (0), r (1), . . . , r (N − 1), r (N )}, the coefficients a(N , n) and the
power PN may then be found using (11.20). Equation (11.20) can be written in matrix
form as
⎡ ⎤ ⎡ ⎤
PN PN
⎢ 0 ⎥ ⎢ 0 ⎥
⎢ ⎥ ⎢ ⎥
Rnn a N = ⎢ . ⎥ or a N = R−1 nn ⎢ .. ⎥ (11.21)
.
⎣ . ⎦ ⎣ . ⎦
0 0
Let
⎡ ⎤
z 11 z 12 ...
⎢ z 21 z 22 ⎥
⎢ ⎥
R−1 =⎢ . ⎥ (11.22)
nn
⎣ .. ⎦
zN1
It then follows that
z 21 z 31 zN1
aTN = 1, , ,..., (11.23)
z 11 z 11 z 11
Ulrych and Bishop [6] also give a convenient recursive procedure for determining the
coefficients in a N .
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:52 430
Having determined the prediction error filter coefficients and the corresponding MEM
spectral estimate, we must now consider how the autocorrelation function can be extended
beyond r (N ) to r (N + 1) where the autocorrelation values r (0), r (1), . . . , r (N ) are all
known. Suppose for example that r (0) and r (1) are known, and it is desired to extrapolate
the autocorrelation function to the unknown value r (2). The prediction error filter matrix
equation for the known values r (0) and r (1) is given by
r (0) r (−1) 1 P1
= (11.24)
r (1) r (0) a(1, 1) 0
where r (−1) = r ∗ (1). To determine the estimate r̂ (2), append one more equation to the
previous matrix equation by incorporating r̂ (2) into the autocorrelation matrix to yield [6]
⎡ ⎤⎡ ⎤⎡ ⎤
r (0) r (−1) r̂ (−2) 1 P1
⎣ r (1) r (0) r (−1) ⎦ ⎣ a(1, 1) ⎦ ⎣ 0 ⎦ (11.25)
r̂ (2) r (1) r (0) 0 0
Solving (11.25) for r̂ (2) then yields r̂ (2) + a(1, 1)r (1) = 0 or
Continuing the extrapolation procedure to still more unknown values of the autocorrelation
function simply involves the incorporation of these additional r (n) into the autocorrelation
matrix along with additional zeros appended to the two vectors to give the appropriate
equation set that yields the desired solution. Since the prediction error filter coefficients
remain unchanged by this extrapolation procedure, it follows that the spectral estimate
given by (11.19) remains unchanged by the extrapolation as well.
With finite data sets, however, (11.27) implicitly assumes that any data outside the finite
data interval are zero. The application of Fourier transform techniques to a finite data
interval likewise assumes that any data that may exist outside the data interval are peri-
odic with the known data. These unwarranted assumptions about unknown data represent
“end effect” problems that may be avoided using the MEM approach, which makes no
assumptions about any unmeasured data.
The MEM approach to the problem of spectral estimation when the autocorrelation
function is unknown estimates the coefficients of a prediction error filter that never runs
off the end of a finite data set, thereby making no assumptions about data outside the data
interval. The prediction error filter coefficients are used to estimate the maximum entropy
spectrum. This approach exploits the autocorrelation reflection–coefficient theorem, which
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:52 431
FIGURE 11-7
a (1, 1) 1 Forward filter
Two-point prediction
error filter operating
x1 x2 x3 x4 xN−3 xN−2 xN−1 xN forward and
backward over an
Backward filter 1 a (1, 1) N-point data set.
1
N
r̂ (0) = |xi |2 (11.28)
N i=1
Now consider how a two-point prediction error filter coefficient can be estimated from an
N-point long data sample. The problem is to determine the two-point filter (having a first
coefficient of unity) that has the minimum average power output where the filter is not
run off the ends of the data sample. For a two-point filter running forward over the data
sample as shown in Figure 11-7, the average power output is given by
N −1
f 1
P1 = |xi+1 + a(1, 1)xi |2 (11.29)
N − 1 i=1
Since a prediction filter operates equally well running backward over a data set as well as
forward, the average power output for a backward running two-point filter is given by
N −1
1
P1b = |xi + a ∗ (1, 1)xi+1 |2 (11.30)
N − 1 i=1
Since there is no reason to prefer a forward-running filter over a backward-running filter and
since (11.29) and (11.30) represent different estimates of the same quantity, averaging the
two estimates should result in a better estimator than either one alone [and also guarantees
that the estimate of the reflection coefficient a(1, 1) is bounded by unity—a fact whose
significance will be seen shortly] so that
1 f
P1 = P1 + P1b
2 N −1
N −1
1 ∗
= |xi+1 + a(1, 1)xi | +
2
|xi + a (1, 1)xi+1 | 2
(11.31)
2(N − 1) i=1 i=1
Now, minimize P1 by selecting the coefficient a(1, 1). Setting the derivative of P1
with respect to a(1, 1) equal to zero shows that the minimizing value of a(1, 1) is given
by [18]
N
−1
−2 xi∗ xi+1
i=1
a(1, 1) = N
−1
(11.32)
(|xi |2 + |xi+1 |2 )
i=1
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:52 432
Having r̂ (0) and a(1, 1), we can find the remaining unknown parts of the prediction error
filter matrix equation using (11.20), that is,
r̂ (0) r̂ (−1) 1 P1
= (11.33)
r̂ (1) r̂ (0) a(1, 1) 0
so that
r̂ (1) = −a(1, 1)r̂ (0) (11.34)
and the output power of the two-point filter is given by
P1 = r̂ (0)[1 − |a(1, 1)|2 ] (11.35)
The result expressed by (11.35) implies that |a(1, 1)| ≤ 1, which is also the necessary and
sufficient condition that the filter defined by {1, a(1, 1)} be a prediction error filter.
Next, consider how to obtain the coefficients for a three-point prediction error filter
from the two-point filter just found. The prediction error filter matrix equation takes the
form
⎡ ⎤⎡ ⎤ ⎡ ⎤
r̂ (0) r̂ (−1) r̂ (−2) 1 P2
⎣ r̂ (1) r̂ (0) r̂ (−1) ⎦ ⎣ a(2, 1) ⎦ = ⎣ 0 ⎦ (11.36)
r̂ (2) r̂ (1) r̂ (0) a(2, 2) 0
From the middle row it follows that
Consequently, the coefficient vector for the three-point filter takes the form
a2T = [1, a(1, 1) + a(2, 2)a ∗ (1, 1), a(2, 2)] (11.40)
Since a(2, 1) is given by (11.38) and a(1, 1) is already known, it follows that the minimiza-
tion of P2 is carried out by varying only a(2, 2) = C2 . As was the case with the two-point
filter, the magnitude of the three-point filter coefficient a(2, 2) must not exceed unity. On
minimizing of P2 , it turns out that |a(2, 2)| ≤ 1, which is the necessary and sufficient
condition that the filter defined by {1, a(1, 1) + a(2, 2)a ∗ (1, 1), a(2, 2)} be a prediction
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:52 433
where A(N , i) now denotes the matrix of N-long forward prediction filter coefficients.
The error associated with x̂ N is then given by
N
ε N = x N − x̂ N = x N + A† (N , i)x N −i = A†N x f (11.46)
i=1
where
A prediction error filter running backward over a data set in general will not have the
same set of matrix prediction filter coefficients as the forward-running prediction filter, so
the error vector associated with a backward-running prediction filter is denoted by
N
b N = x0 − x̂0 = x0 + B† (N , i)xi = B†N xb (11.47)
i=1
The fact that A(N , i) = B† (N , i) reflects the fact that the multichannel backward predic-
tion error filter is not just the complex conjugate time reverse of the multichannel forward
prediction error filter (as it was in the scalar case).
The matrix generalization of (11.20) is given by
⎡ ⎤⎡ ⎤
R(0) R(−1) ... R(−N ) I
⎢ R(1) . . . R(−N + 1) ⎥ ⎢ ⎥
⎢ R(0) ⎥ ⎢ A(N , 1) ⎥
R f AN = ⎢ . ⎥ ⎢ . ⎥
⎣ .. ⎦⎣ .. ⎦
R(N ) R(N − 1) . . . R(0) A(N , N )
⎡ f ⎤
PN
⎢ 0 ⎥
⎢ ⎥
=⎢ . ⎥ (11.48)
⎣ .. ⎦
0
where the p × p block submatrices R(k) are defined by
so that
P N = E{ε N ε †N } = A†N R f A N
f
(11.51)
The backward power matrix PbN for the prediction error filter satisfying (11.52) is then
The matrix coefficients A(N , N ) and B(N , N ) are referred to as the forward and backward
reflection coefficients, respectively, as follows:
f
C N = A(N , N ) and CbN = B(N , N ) (11.54)
The maximum entropy power spectral density matrix can be computed either in terms
of the forward filter coefficients using
†
1 1
xx ( f ) = t A−1 P N A−1
f
(11.55)
z z
where
A(z) = I + A(N , 1)z + · · · + A(N , N )z N (11.56)
and z = e− j2π f t , or in terms of the backward filter coefficients using
where
vm = PbN −1 C N P−1
f
N −1 ε m + bm
N N
(11.63)
Equations (11.61) and (11.63) show that the forward and backward prediction error
f
filter residual outputs depend only on the forward reflection coefficient C N . The coefficient
f
C N are chosen to minimize a weighted sum of squares of the forward and backward residual
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:52 436
f
outputs of the filter of length N; that is, minimize SS C N where
f 1 M
f −1 −1
SS C N = K m u†m P N −1 um + v†m PbN −1 vm (11.64)
2 m=1
and K m is a positive scalar weight = 1/M. The equation that yields the optimum value of
f
C N for (11.64) is then
f f f −1
BC N + PbN −1 C N P N −1 E = −2G (11.65)
where
M
†
B= K m bmN bmN (11.66)
m=1
M
†
E= K m ε mN εmN (11.67)
m=1
M
†
G= K m bmN εmN (11.68)
m=1
f
After obtaining C N from (11.65) (which is a matrix equation of the form AX + XB = C)
then CbN can be obtained from (11.59), and the desired spectral matrix estimate can be
computed from (11.55) or (11.57).
FIGURE 11-8 10
Comparison of the MUSIC
DOA algorithms
Periodogram
when the three
signals incident on Capon
0
the 12-element MEM
uniform array at θ =
−50◦ , 10◦ , and 20◦
|P (dB)|
−20
−90 −45 0 45 90
q (degrees)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:52 437
10 FIGURE 11-9
MUSIC
Comparing four DOA
techniques when the
Periodogram
standard deviation
Capon of the noise is 2.
0
MEM
|P (dB)|
−10
−20
−90 −45 0 45 90
q (degrees)
strengths of 16, 1, and 4, respectively. The lowest amplitude signal is almost lost in the
periodogram and is off by several degrees from the true signal position.
Increasing the noise variance causes the noise floor of the periodogram and Capon
spectrum to rise as shown in Figure 11-9. The Capon spectrum barely distinguishes be-
tween the two closely spaced signals, and the peaks no longer accurately reflect the signal
strengths. The noise floor of the MUSIC spectrum goes up, but it has stronger peaks than
the MEM spectrum.
xN (t)
p(q )
and then using the resulting estimate in the likelihood ratio test statistic as if it were known
exactly. The Bayes approach to the parameter estimation problem incorporates any a priori
knowledge concerning the signal parameters in the form of a probability density function
xN (t)
qˆ
Parameter
estimation
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:52 439
p(θ /signal present). To obtain a signal parameter estimate for an estimate and plug array
processor, note that
(x) p(x/θ, signal present) p(θ /signal present)dθ
= (11.71)
(x/θ̂) p(x/θ̂ , signal present)
The optimal Bayes processor explicitly incorporates a priori knowledge about the unknown
signal parameters θ into the likelihood ratio (x) through the averaging process expressed
by the numerator of (11.71). To obtain an estimate θ̂ to use in a suboptimal estimate and
plug structure, require
(x)
=1 (11.72)
(x/θ̂ )
Having evaluated (x) using the averaging process, we find θ̂ as the solution to (11.72),
and it is referred to as a “pseudo-estimate” θ̂ PSE [32]. The performance of the Bayes optimal
processor for the case of a signal known except for direction (SKED) was investigated
by Gallop and Nolte [33]. A comparison between the two estimate and plug structures
obtained with a MLE and a Bayes pseudo-estimate is given in [34] for the case of target
location unknown. The results indicated that the ML detector performs the same as the
Bayes pseudo-estimate detector when the a priori knowledge about the uncertain parameter
is uniformly distributed. When the a priori knowledge available is more precise, however,
the performance of the Bayes pseudo-estimate detector improves whereas that of the
ML detector does not, and this difference between the two processors becomes more
pronounced as the array size becomes larger.
When unknown signal parameters are present, application of the averaging process
described in the previous section to p(θ ) yields
L
P(x , x , . . . , x ) =
1 2 L
p(xi /xi−1 , . . . , x1 , θ ) p(θ )dθ (11.74)
i=1
where p(θ/xi−1 , . . . , x1 ) represents the updated form of the a priori information contained
in p(θ ).
A sequential processor may now be implemented using (11.79) and (11.80) to form
the marginal density functions required in the likelihood ratio as follows:
p(x1 , . . . , x L /signal present)
(x1 , . . . , x L ) = (11.81)
p(x1 , . . . , x L /signal absent)
A block diagram of the resulting sequential array processor based on (11.79)–(11.81) is
given in Figure 11-12. The sequential Bayesian updating of p(θ) represented by (11.79)
results in an optimal processor having an adaptive capability.
Performance results using an optimal sequential array processor were reported in [29]
for a detection problem involving a signal known exactly imbedded in Gaussian noise
where the noise has an additive component arising from a noise source with unknown
direction. The adaptive processor in this problem must both succeed in detecting the
presence or absence of the desired signal and in “learning” the actual direction of the
x1i
L
noisy signal source. The results obtained indicated that even though the directional noise to
thermal noise ratio was relatively low, the optimal sequential processor could nevertheless
determine the directional noise source’s location.
where ai (θ j ) is a complex scalar representing the sensor response to the jth emitter signal
(there are a total of “d” signals). The jth emitter signal is denoted by s j (t), and the additive
noise is n i (t). In matrix notation, (11.82) can be written as
The emitter covariance matrix is give by S = E[s(t)s H (t)], where “E” denotes the expected
value. Likewise, the output signal covariance matrix is given by
measurements X N in a least squares sense. The basic subspace fitting problem is defined
by
min
Â, T̂ = arg M − AT2F (11.87)
A, T
min
where A2F = trace(A H A) is the Frobenius norm, and θ̂ = arg V(θ) is the mini-
θ
mizing argument of V(θ ) The k × q matrix M in (11.87) represents the data, whereas T
is any p × q matrix. For a fixed A, the minimum with respect to T is a measure of how
well the range spaces of A and M match. The subspace fitting estimate selects A so these
subspaces are as close as possible. The estimate of θ is then obtained from the parameters
of Â. It is of some practical interest to note that the subspace fitting problem is separable in
A and T. By substituting the pseudo-inverse solution T̂ = A p M back into (11.87) where
A p = (A H A)−1 A H , one obtains the following equivalent problem
max
 = arg tr{P A MM H } (11.88)
A
where P A = AA p is a projection matrix that projects onto the column space of A. The
subspace fitting problem then resolves into a familiar parameter optimization problem
described by
max
θ̂ = arg tr{P A (θ )R̂xx } (11.89)
θ
where R̂xx is the sample covariance matrix. Notice from (11.88) that the same result could
be obtained by simply taking R̂xx = MM H . This approach is the deterministic maximum
likelihood method for obtaining direction-of-arrival estimates.
The ESPRIT algorithm assumes that the array is composed of two identical subarrays,
each having k/2 elements (so the total number of array elements is even). The subarrays
are displaced from each other by a known displacement vector so the propagation between
the subarrays can be described by the diagonal matrix = diag[e jωτ1 e jωτ2 . . . e jωτe ],
where τi is the time delay in the propagation of the ith emitter signal between the two
subarrays, and ω is the center frequency of the emitters. The time delay is then related to
the angle of arrival by τi = || sin θi /c, where c is the speed of propagation, and is the
displacement vector between the two subarrays. The output of the array is then modeled as
n (t)
x(t) = s(t) + 1 (11.90)
n 2 (t)
where the k/2 × d matrix contains the common array manifold vectors of the two
subarrays, and is the diagonal matrix discussed already. Since the matrices defined by
[E1T E2T ]T (eigenvectors for the two subarrays) and [ T T T ] have the same range space,
there is a full rank d × d matrix T such that
E1
= T (11.91)
E2
Eliminating in (12.91) then yields
E2 = E1 T−1 T = E1 (11.92)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:52 443
where L = diag[l1 , l2 , . . . , l2d ]. Since = T−1 T, the elements of are estimated by
the eigenvalues of ˆ TLS . The principal angles of these eigenvalues give estimates of the
time delays τi , which in turn give the DOA estimates. Solving for and equating its
eigenvalues to results in the signal angle estimates of
arg(λm)
θm = sin−1 (11.94)
kd
where λ
m = eigenvalues of .
⎡ ⎤
−0.0816 +j0.0000 −0.3214 −j0.0000 0.3239 −j0.0000
⎢ −0.0214 +j0.1413 0.1728 −j0.1651 0.2693 −j0.3461 ⎥
⎢ ⎥
⎢ 0.0775 +j0.2593 −0.0428 +j0.3203 −0.0252 −j0.2482 ⎥
⎢ ⎥
⎢ 0.3061 +j0.0678 −0.1822 −j0.1796 −0.2547 −j0.1985 ⎥
⎢ ⎥
⎢ 0.3649 −j0.1551 0.2694 +j0.1027 −0.0855 +j0.0202 ⎥
⎢ ⎥
⎢ ⎥
⎢ 0.0744 −j0.3899 −0.2220 +j0.1963 −0.0774 +j0.0945 ⎥
Vs = ⎢ ⎥
⎢ −0.1733 −j0.3619 0.1385 −j0.2273 −0.0089 −j0.1292 ⎥
⎢ ⎥
⎢ −0.3708 −j0.0374 0.1865 +j0.2484 −0.0916 −j0.0330 ⎥
⎢ ⎥
⎢ −0.2907 +j0.1344 −0.2014 −j0.1534 −0.3170 +j0.0412 ⎥
⎢ ⎥
⎢ −0.0037 +j0.2384 0.2845 −j0.1651 −0.0919 +j0.2447 ⎥
⎢ ⎥
⎣ 0.0553 +j0.1390 −0.1668 +j0.1730 0.0681 +j0.4281 ⎦
0.0542 −j0.0455 −0.1122 −j0.2981 0.3375 +j0.1231
Next, compute to get
⎡ ⎤
0.6322 −j0.6792 −0.1035 +j0.2722 0.2247 −j0.0474
= ⎣ −0.2434 −j0.0108 −0.6666 +j0.5904 −0.3027 +j0.3979 ⎦
−0.1525 +j0.2374 −0.2235 +j0.1595 0.6221 −j0.6479
The eigenvalues of ψ are found and substituted into (11.94) to find an estimate of the
angle of arrival.
θm = −50.15◦ 9.99◦ 20.02◦
signals to be Gaussian zero mean (“stochastic” signal model); and (2) the second model
assumes that signals are “deterministic” but unknown. We start with the stochastic model.
xk = sk + nk (11.95)
For the stochastic signal model, the desired signal sk is also assumed to be a sample
function from a zero-mean N -variate complex Gaussian process, with covariance
Rss = E sk s†k (11.97)
Under these assumptions, the probability density for the data sample xk is given by [7]
p(xk ) = (π )−N |R|−1 exp −x†k R −1 xk (11.98)
It is assumed that the desired signal and noise are uncorrelated, so that
Our objective is to estimate the angle of arrival of the signal sk in the presence of
internal and external interference denoted by nk , based on K independent data samples
(“snapshots”) of xk , k = 1, . . . , K . The ML estimation procedure is based on determining
the value of the unknown parameters (parameters to be estimated) that maximizes the
conditional joint density function of the K independent data samples, which from (11.97)
has the form
K
†
−N K −K −1
p(x1 , x2 , . . . x K |θ, S ) = (π ) |R| exp − xk R xk (11.101)
k=1
The density in (11.100) is conditional on the values of the unknown signal parameters to be
estimated, namely, the AOA (θ ) and the signal intensity (S). To maximize (11.100) with
respect to θ and S, Equation (11.100) must be reformulated to show explicit dependence
on these variables. For narrowband signals, where the element-to-element time delay
experienced by the desired signal can be represented as phase shifts of the signal, this
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:52 445
equation reduces to
K
1 S K
† −1 † −1
P(x1 , x2 , . . . , x K |θ, S ) = C exp x R d(θ )d (θ )Rnn xk
1 + SB 1 + S B k=1 k nn
(11.102)
where
K
†
−N K −K −1
C = (π) |Rnn | exp − xk Rnn xk (11.103)
k=1
Rss = Sd(θ )d † (θ ) (11.104)
−1
B = d† (θ )Rnn d(θ ) (11.105)
Here the scalar C is a constant, the scalar B depends on θ , and the N × N covariance
matrix Rss depends on both S and θ. Equation (11.102) was derived under the narrowband
signal assumption, which permits the signal covariance matrix to be written as in (11.104).
The N -component vector d(θ) is the (unknown) vector of phase delays corresponding to
the signal angle of arrival. For the case of a linear array of identical isotropic antenna
elements (received signal power is the same in each element), S represents the (unknown)
received signal power at each element so that d(θ ) can be written as
where bkm is the (complex) amplitude of the kth sample of the mth signal, and dm = dm (θm )
is the direction delay vector for the mth signal. The general form of d for the mth narrowband
signal source is given by [38]
dm (θm ) = g1 (θm )e− jφ1 (θm ) , g2 (θm )e− jφ2 (θm ) . . . g N (θm )e− jφ N (θm ) (11.110)
where θm is the angular location (angle of arrival) of the mth source, gn is the complex
gain of the nth antenna element (generally a function of θ ), and φn (θm ) is the phase delay
of the m th mth source at the nth antenna element relative to a suitable reference point (e.g.,
the array phase center).
It is convenient to rewrite (11.107) in vector/matrix notation as
The M-dimensional complex vector bk denotes the signal amplitude and phase in complex
notation for each of the M signals for the kth snapshot. Under the stochastic signal model,
bk is assumed to be a zero-mean Guassian random vector with covariance
The direction vectors for each of the M signals, dm (θm ), m = 1, . . . M, form the columns
of the N × M matrix D(θ ).
Maximizing (11.111) with respect to θ and Rbb is equivalent to maximizing ln( p),
denoted the log likelihood function
K
ln L (θ 1 , θ 2 , . . . , θ M , Rbb ) = C − K log |R| − x†k R −1 xk (11.114)
k=1
The likelihood function in Equation (1.18) can be simplified by dropping the terms that
do not depend on either θ or Rbb and by rearranging terms
1
K
R̂ = x x† (11.117)
K k=1 k k
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:52 447
where
As noted in [38], G(θ) is the orthogonal projection matrix. For the purpose of simplifica-
tion, the derived result in (11.116) made the assumption that the noise covariance matrix
Rnn is known and is given by Rnn = σn2 I N ; that is, the only noise terms are due to thermal
noise. The result in (11.116) is generalized to handle an arbitrary positive definite matrix
Rnn by transforming xk into a new coordinate system so that the noise covariance matrix
in the new coordinate system is σ 2 I N , that is,
−1/2
yk = Rnn xk (11.120)
Once θ̂ M L is determined, Jaffer [38] wrote the maximum likelihood estimate of Rbb as
−1 † −1
R̂bb θ̂ M L = D † θ̂ M L D θ̂ M L D θ̂ M L R̂ − σn2 I N D θ̂ M L D † θ̂ M L D θ̂ M L
(11.121)
In summary to this point, we have shown that the maximum likelihood estimate of the
angles of arrival of M signal sources is determined by finding those values of θ1 , θ2 , . . . , θ M
that maximize J (θ ) in Equation (11.116). This requires an M-dimensional search over the
angular regions of interest. To evaluate J (θ ), it is assumed that the following parameters
are known a priori: (1) antenna element gain versus θ; (2) antenna element location; and
(3) the noise covariance matrix Rnn = σn2 I N . The sample covariance matrix is computed
from the data samples (11.115). Note the result in (11.116) assumes that the signals are
narrowband and propagate as plane waves. Also note that equation (11.116) assumes that
Rnn is known a priori: in most practical situations, the external interference is not known
beforehand and must be estimated from the data samples. In radar, it is often possible
to estimate Rnn by averaging the sample covariance matrix using adjacent range cells
that do not contain the target (desired signal). Such an estimate is more difficult in a
communications system in which the signal is present in all the data samples. In this case,
it may be necessary to use an a priori estimate of Rnn . For example, if it is assumed that
Rnn is made up of only internal noise (Rnn = σn2 I N ), then all external sources must be
estimated using Equation (11.116).
The M-dimensional search required to find the values of θm that maximize J (θ ) can
become computationally intensive for large M. In practice, M is limited to two or three
sources by limiting the angular search region, typically one or two beamwidths in extent.
Any sources outside the search region (e.g., in the sidelobes) are treated as interference and
therefore must be included in Rnn , which as noted already must be known or estimated.
Methods for obtaining an accurate estimate of Rnn depend on the particular situation
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:52 448
(e.g., radar vs. communications vs. pulsed signals, amplitude of the desired signals),
which is beyond the scope of the current discussion.
In this case, the likelihood function p is maximized by minimizing the exponent, which
is a real scalar given by
1
K
H= Hk (11.123)
K k=1
where
Hk = [bk − T −1 V xk ]† T [bk − T −1 V xk ] + x†k Rnn
−1
xk + x†k V † T −1 V xk (11.124)
−1
T = D † (θ)Rnn D(θ ) (11.125)
† −1
V = D (θ)Rnn
Substituting bk into Equation (11.123) and deleting terms that do not depend on θ, one
obtains an expression for H that depends only on θ
1 † −1
K
−1
Jdet (θ ) = x R D(θ )(D † (θ )Rnn D(θ ))−1 D † (θ )Rnn
−1
xk (11.127)
K k=1 k nn
1 †
K
Jdet (θ) = y U (θ )(U † (θ)U (θ))−1 U † (θ )yk (11.130)
K k=1 k
This has the same form as the first term in the expression for Jsto (θ ), and indicates how
the first term in (11.116) can be transformed to generalize it to any positive definite Rnn .
From Equation (11.100), the log likelihood function for the stochastic signal model is
given by
K
ln L(υ) = −NK ln(π ) − K ln |R| − x†k Rnn
−1
xk (11.134)
k=1
Now to simplify the development (a more general case will be derived later), assume
a single narrowband source with N × 1 direction vector d(μ), as in Equation (11.95).
Then the likelihood function reduces to Equation (11.101) and the log likelihood function
reduces to
S †
−1 −1
ln L = ln C − K ln(1 + SB) + xk Rnn d(μ)d† (μ)Rnn xk (11.135)
(1 + SB)
where C is given by (11.103), B is given by (11.105), and S is the received signal power
as defined in (11.102). Taking the partial derivatives in Equation (11.102) with respect to
μ and S, one obtains for the 2 × 2 Fisher information matrix
∂ 2 ln L KS2 2
11 = −E = F + 2(1 + SB)B 2 D
∂μ 2 (1 + SB) 2
∂ ln L
2
KSB
21 = −E = F = 12
∂ S∂μ (1 + SB)2
∂ 2 ln L KB2
22 = −E =
∂S 2 (1 + SB)2
−1 1 + SB
σμ2 ≥ 11 = (11.136)
2KS2 B 2 Q
−1 −1 S −1
12 = 11 F = 21 (11.137)
B
2
−1 −1 S
σ S2 ≥ 22 = 11 2
F 2 + 2(1 + SB)B 2 Q (11.138)
B
where Q and F are real scalar quantities given by
! 2 "
1 † † −1 † † −1
Q = 2 Bd (μ)A Rnn Ad(μ) − d (μ)A Rnn d(μ)
B
−1 −1
F = d† (μ)Rnn Ad(μ) + d† (μ)A† Rnn d(μ)
and A is defined as the M × M diagonal matrix with diagonal elements given by Ann =
βn = 2π αn /λ; n = 1, . . . N .
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:52 451
11.10 Fisher Information Matrix and CR Bound for General Cases 451
Equation (11.136) gives the CR lower bound on the variance of any unbiased estimate
of the angle of arrival θ, and Equation (11.137) gives the lower bound on the variance of
any unbiased estimate of the signal power S.
For an N-element line array of identical and equally spaced antenna elements, each
with unity gain, and assuming Rnn = σn2 I N (so S/σn2 = the signal-to-noise ratio of the
signal received by each element), the lower bound on the variance of the AOA estimate
about its true value is given by
2
−1 6 1 1 + NS/σn2 1 λ
σμ ≥ E(μ̂ − μtr ue ) = 11 =
2 2
2 (11.139)
(2π ) K NS/σn
2 2 (N − 1) α
2
where α is the separation between antenna elements. Note that NS/σn2 is the array SNR.
Equation (11.101) applies to the deterministic signal case, where the signal is unknown
but nonrandom. Ballance and Jaffer [39] showed that for the stochastic signal model
2
−1 6 1 1 1 λ
σθ ≥ E(μ̂ − μtr ue ) = 11 =
2 2 (11.140)
(2π ) K NS/σn (N − 1) α
2 2 2
For moderate to large SNR (greater than approx 10 dB), the two bounds are nearly the
same.
K
−1
ln L = −NK ln(π) − K ln |Rnn | − (xk − sk (υ))† Rnn (xk − sk (υ)) (11.142)
k=1
Note the difference between the previous deterministic case and the stochastic case in
Equation (11.114). The deterministic case is easier to deal with because only the exponent
depends on the unknown parameter vector υ. Taking the partial derivatives defined in
Equation (11.133) and following the derivation in [39], the general M × M symmetric
FIM is given by
K
=2 Re Hk† Rnn
−1
Hk (11.143)
k=1
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:52 452
The AOA is determined from one of terms of −1 . Consider the case of a single emitter,
−1
with Rnn = σn2 I N . Let the kth snapshot xk = d(θ )bk where bk is a 1 × 1 complex vector,
and d(θ) is the N × 1 vector defined as
d(θ) = g1 (θ )e− jϕ1 (θ ) , g2 (θ )e− jϕ2 (θ ) . . . g N (θ )e− jϕ N (θ ) (11.145)
where ϕn (θ) = 2π τ (θ ). τn (θ ) is the time delay experienced by the desired signal relative
λ n
to fixed reference delay. The components of the FIM are determined by letting υ1 = θ,
υ2 = b1 , υ3 = |b1 |, υ4 = b2 , υ5 = |b2 | , . . . υ2K +1 = |b K |, and then taking the partial
−1
derivatives in (11.144). Following [39], 11 is found to be given by
⎡ ⎤
† 2
1 # #
⎢# #2 d d ⎥
−1
σμ̂2 ≥ 11 = ⎣ d − 2 ⎦
(11.146)
2K (SNR1 ) d
1 |bk |2
N
where SNR1 = is the average array signal-to-noise ratio that would be received
K k=1 σn2
by antenna elements with unity gain.
11.12 PROBLEMS
1. Computer Simulation of Direction of Arrival Estimation Algorithms.
a. Use a periodogram to demonstrate the effect of separation angle between two sources using
an 8 element uniform array with λ/2 spacing when θ1 = −30◦ , 10◦ , 20◦ and θ2 = 30◦ .
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:52 453
b. An 8 element uniform array with λ/2 spacing has 3 signals incident upon it: s1 (−60◦ ) = 1,
s2 (0◦ ) = 2, and s3 (10◦ ) = 4. Find the Capon spectrum.
c. An 8 element uniform array with λ/2 spacing has three signals incident upon it: s1 (−60◦ ) =
1, s2 (0◦ ) = 2, and s3 (10◦ ) = 4. Find the MEM spectrum.
d. An 8 element uniform array with λ/2 spacing has three signals incident upon it: s1 (−60◦ ) =
1, s2 (0◦ ) = 2, and s3 (10◦ ) = 4. Find the MUSIC spectrum.
e. An 8 element uniform array with λ/2 spacing has three signals incident upon it: s1 (−60◦ ) =
1, s2 (0◦ ) = 2, and s3 (10◦ ) = 4. Find the location of the signals using the root MUSIC
algorithm.
2. Prediction Error Filter Equations The prediction error filter matrix equation for a scalar
random process may be developed by assuming that two sampled values of a random process
x0 and x1 are known, and it is desired to obtain an estimate of the next sampled value x̂2 using
a second-order prediction error filter
where r (n) = xi x j , |i − j| = n.
b. If P2 = ε 2 = (x2 − x̂2 )(x2 − x̂2 ), use the fact that the error in the estimate x̂2 is orthogonal
to the estimate itself (i.e., (x2 − x̂2 )x̂2 = 0) and the previously given expression for x̂2 to
show that
The results of part (a) combined with the result from part (b) then yield the prediction error
filter matrix equation for this case.
3. The Relationship between MEM Spectral Estimates and ML Spectral Estimates [19]
Assume the correlation function of a random process x(t) is known at uniformly spaced,
distinct sample times. Then the ML spectrum (for an equally spaced line array of N sensors) is
given by
N x
MLM(k) =
v† (k)R−1
xx v(k)
where
k = wavenumber (reciprocal of wavelength)
x = spacing between adjacent sensors
v(k) = beam steering column vector where vn (k) = e− j2π nkx , n = 0, 1, . . . , N − 1
Rxx = N × N correlation matrix of x(t)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:52 454
where 1, c(2, M), . . . , c(M, M) are the weights of the M-long prediction error filter whose
output power is P(M). Note that
⎡ ⎤
P(N ) −−− −−− −−− −−−
⎢ 0 P(N − 1) −−− −−− −−− ⎥
⎢ 0 −−−
⎥
−−− ⎥
Rxx L = ⎢
⎢ .
0
⎥
⎣ .. .. .. ⎦
. .
0 0 P(1)
The maximum entropy spectrum estimate corresponding to the M-long prediction error
filter is given by
P(M)x
MEM(k, M) = 2
M
c(i, M) exp( j2π k(i − 1)x)
i=1
P ≡ L† Rxx L
Show that P is an N × N diagonal matrix whose diagonal elements are given by P(N ),
P(N − 1), . . . , P(1).
−1 †
b. Using the fact that R−1
xx = LP L , show that
N
x
† † † −1 †
v R−1
xx v = (L v) P (L v) =
MEM(k, n)
n=1
1
N
1 1
=
MLM(k) N MEM(k, n)
n=1
Therefore, the reciprocal of the ML spectrum is equal to the average of the reciprocals of
the maximum entropy spectra obtained from the one-point up to the N-point prediction
error filter. The lower resolution of the ML method therefore results from the “parallel
resistor network averaging” of the lowest- to the highest-resolution maximum entropy
spectra.
4. Equivalence of MEM Spectral Analysis to Least-Squares Fitting of a Discrete-Time All-Pole
Model to the Available Data [15] Assume that the first (N + 1) points {r (0), r (1), . . . , r (N )}
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:52 455
of the autocorrelation function of a stationary Gaussian process are known exactly, and it is
desired to estimate r (N + 1). Consider the Toeplitz covariance matrix
⎡ ⎤
r (0) r (1) · · · r (N ) r (N + 1)
⎢ r (1) r (0) · · · r (N − 1) r (N ) ⎥
R N +1 =⎣ ⎦
··· ··· ···
r (N + 1) r (N ) r (1) r (0)
The basic autocorrelation function theorem states that R N +1 must be semipositive definite
if the quantities r (0), r (1), . . . , r (N + 1) are to correspond to an autocorrelation function.
Consequently det[R N +1 ] must be nonnegative.
MEM spectral analysis seeks to select that value of r (N + 1) that maximizes det[R N +1 ].
The entropy of the (N + 2) dimensional probability density function with covariance matrix
R N +1 is given by
and the choice for r (N +1) maximizes this quantity. To obtain r (N +2), the value of r (N +1) just
found is substituted into R N +2 to find det[R N +2 ], and the corresponding entropy is maximized
with respect to r (N + 2). Likewise, substituting the values of r (N + 1) and r (N + 2) found
already into det[R N +3 ] and maximizing yields r (N + 3). The estimates for additional values
r (N + 4), r (N + 5), · · · may then be evaluated by following the same procedure.
a. Show that maximizing det[R N +1 ] with respect to r (N + 1) is equivalent to the relation
⎡ ⎤
r (1) r (0) · · · r (N − 1)
⎢ r (2) r (1) · · · r (N − 2) ⎥
det ⎢
⎣ .. .. .. ⎥=0
⎦
. . .
r (N + 1) r (N ) ··· r (1)
or
yT a = e(n)
where
aT = [1, a1 , a2 , . . . , a N ]
yT = [y(n), y(n − 1) · · · y(n − N )]
N = order of all-pole model
n = number of data samples, n > N
and where e(n) is a zero-mean random variable with E{e(i)e( j)} = 0 for i = j. Assuming
that E{e(n)y(n − k)} = 0 for k > 0, show that multiplying both sides of the previous
equation for e(n) by y(n − k) and taking expectations yields
where r (k) = E{y(n)y(n − k)}.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:52 456
c. Using the results of part (b) and the fact that r (τ ) = r (−τ ), it follows that
If the first N + 1 exact values {r (0), r (1), . . . , r (N )} of any autocorrelation function are
available, then substituting these values into the first N of the simultaneous linear equations
corresponding to R N +1 a = 0 yields a unique solution for the coefficients {a1 , a2 , . . . , a N }.
Consequently, the value for r (N + 1) for a discrete-time all-pole model having the coeffi-
cients {a1 , a2 , . . . , a N } is uniquely determined by
⎡ ⎤
r (1) r (0) · · · r (N − 1)
⎢ r (2) r (1) · · · r (N − 2) ⎥
det ⎢
⎣ .. .. ⎥=0
⎦
. .
r (N + 1) r (N ) · · · r (1)
This result is identical to the relation obtained in part (a); hence, the same solution would
have been obtained from maximum entropy spectral analysis.
5. Angle of Arrival Estimation [37] The MEM technique has superior capability for resolving
closely spaced spectral peaks that may be exploited for estimating the angular distribution of
received signal power.
a. Using a time–space dualism, reformulate equation (11.19) to give a spatial spectrum φ̂xx (μ),
where μ = cos θ , and θ is the angle from array endfire for an N + 1 element linear array.
Assume narrowband signals with spacing d between elements.
b. Reformulate equation (11.20) for the spatial estimation problem of part (a). What correspon-
dence exists between the prediction error filter coefficients and the weights of a coherent
sidelobe canceller with N auxiliary antennas?
6. The Marple Algorithm [41] A new autoregressive (AR) spectral analysis algorithm has been
proposed that yields spectral estimates with no apparent line splitting (the occurrence of two or
more closely spaced peaks in the AR spectral estimate where only one peak should be present)
and reduced spectral peak frequency estimation biases. It exploits forward and backward linear
prediction and therefore is closely related to the Burg algorithm. With the Burg algorithm the
forward linear prediction error, f M,k , in the single-channel case is given by
M
M
f M,k = xk+M + a M,i xk+M−i = a M,i xk+M−i for 1 ≤ k ≤ N − M
i=1 i=0
where M is the order of the all-pole AR model, xk is the kth sample output of the AR model,
and a M,m is the AR parameter m of the Mth order process. Note that an,0 is defined as unity.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:52 457
M
∗
b M,k = a M,i xk+i also for 1 ≤ k ≤ N − M
i=0
Since stationarity is assumed, the backward AR coefficients are the conjugates of the forward
AR coefficients. To obtain estimates of the AR parameters, Burg minimized the sum of the
backward and forward prediction error energies
N −M
N −M
eM = f M,k 2 + b M,k 2
k=1 k=1
Substituting f M,k and b M,k into e M and setting the derivatives of e M with respect to the param-
eters a M,1 through a M,M to zero, one obtains
M
2 a M, j r M (i, j) = 0 for i = 1, . . . , M(a M,0 = 1 by definition)
j=0
N −M ∗ ∗
where r M (i, j) = k=1 (x k+M− j x k+M−i + x k+ j x k+i ) for 0 ≤ i, j ≤ M. The minimum
prediction energy is then given by
M
eM = a M, j r M (0, j).
j=0
a. Show that the previous three expressions can be written in matrix form as
RM AM = EM
⎡ ⎤ ⎡ ⎤
1 eM
⎢ a M,1 ⎥ ⎢ 0 ⎥
where A M =⎢ ⎥ ⎢ ⎥
⎣ .. ⎦ , E M = ⎣ .. ⎦ ,
. .
a M,M 0
⎡ ⎤
r M (0, 0) · · · r M (0, M)
⎢ .. .. ⎥
and R M =⎣ . . ⎦
r M (M, 0) · · · r M (M, M)
b. The matrix expression found in part (a) has a structure that can be exploited to produce an
algorithm requiring a number of operations ∝ M 2 rather than M 3 . R M has both Hermitian
∗ ∗
symmetry [r M (i, j) = r M ( j, i)] and Hermitian persymmetry [r M (i, j) = r M (M −i, M − j);
it does not have Toeplitz symmetry [r M (i, j) = r M (i − j)], as the covariance matrix does.
However, R M is composed of Toeplitz matrices. Show that
R M = (T M ) H T M + (TνM ) H TνM
⎡ ⎤
x M+1 x M · · · x1
⎢ x M+2 x M+1 · · · x2 ⎥
where T M = ⎢
⎣ .. .. ⎥⎦
. .
x N x N −1 · · · x N −M
⎡ ⎤
x1∗ · · · x M+1
∗
⎢ .. .. ⎥
and TνM =⎣ . . ⎦ (conjugate and reversed matrix)
x N∗ −M · · · x N∗
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:52 458
c. To exploit this structure, introduce two new prediction error energy terms
−M−1
N
−M−1
N
eM = f M,k+1 2 + b M,k 2 and eM = f M,k 2 + b M,k+1 2
k=1 k=1
d. Show that the following relationships exist among the correlation matrices R M , RM , and
RM :
∗
⎡ ⎤ ⎡ ⎤
x M+1 x N −M
⎢ .. ⎥ ⎢ . ⎥ ∗
RM = R M − ⎣ . ⎦ x M+1 , · · ·, x1 − ⎣ .. ⎦ x N −M , · · · , x N∗
x1∗ xN
⎡ ⎤ ⎡ ⎤
x N∗ x1
⎢ .. ⎥ ⎢ ⎥
RM = R M − ⎣ . ⎦ x N , · · · , x N −M − ⎣ ... ⎦ x1∗ , · · · , x M+1
∗
x N∗ −M x M+1
⎡ ⎤
RM |
r M+1 (0, M + 1)
⎢ .. ⎥
R M+1 = ⎣ −− −− −− −− −− | . ⎦
r M+1 (M + 1, 0) · · · r M+1 (M + 1, M + 1)
⎡ ⎤
r M+1 (0, 0) · · · r M+1 (0, M + 1)
⎢ .. ⎥
R M+1 =⎣ . | −− −− −− −− −− ⎦
r M+1 (M + 1, 0) | RM
At this point, we are approaching the development given by Burg but now using time-shifted
AR parameters rather than the AR parameters employed before. The complete derivation
is rather lengthy and will not be pursued here. A complete block diagram of the Maple
algorithm is given in [41] along with a similar diagram for the Burg algorithm for ease in
comparison.
7. Computer Simulation Problem An eight-element uniform array with λ/2 spacing has three
signals incident upon it: s1 (−60◦ ) = 1, s1 (−0◦ ) = 2, and s3 (−10◦ ) = 4. Estimate the incident
angles using ESPRIT.
11.13 REFERENCES
[1] J. Capon, “High-Resolution Frequency-Wavenumber Spectrum Analysis,” Proceedings of the
IEEE, Vol. 57, No. 8, 1969, pp. 1408–1418.
[2] R. Schmidt, “Multiple Emitter Location and Signal Parameter Estimation,” IEEE Transactions
on Antennas and Propagation, Vol. 34, No. 3, 1986, pp. 276–280.
[3] R. Schmidt and R. Franks, “Multiple Source DF Signal Processing: An Experimental System,”
IEEE Transactions on Antennas and Propagation, Vol. 34, No. 3, 1986, pp. 281–290.
[4] A. Barabell, “Improving the Resolution Performance of Eigenstructure-Based Direction-
Finding Algorithms.” pp. 336–339.
[5] R. T. Lacoss, “Data Adaptive Spectral Analysis Methods,” Geophysics, Vol. 36, August 1971,
pp. 661–675.
[6] T. J. Ulrych and T. N. Bishop, “Maximum Entropy Spectral Analysis and Autoregressive
Decompositions,” Rev. Geophys. Space Phys., Vol. 13, 1975, pp. 183–200.
[7] J. P. Burg, “Maximum Entropy Spectral Analysis,” Paper presented at the 37th Annual Meeting
of the Society of Exploration Geophysicists, October 31, 1967, Oklahoma City, OK.
[8] J. P. Burg, “A New Analysis Technique for Time Series Data,” NATO Advanced Study Institute
on Signal Processing with Emphasis on Underwater Acoustics, Vol. 1, Paper No. 15, Enschede,
The Netherlands, 1968.
[9] T. J. Ulrych and R. W. Clayton, “Time Series Modelling and Maximum Entropy,” Phys. Earth
Planet. Inter., Vol. 12, 1976, pp. 188–200.
[10] S. B. Kesler and S. Haykin, “The Maximum Entropy Method Applied to the Spectral Analysis
of Radar Clutter,” IEEE Trans. Inf. Theory, Vol. IT-24, No. 2, March 1978, pp. 269–272.
[11] R. G. Taylor, T. S. Durrani, and C. Goutis, “Block Processing in Pulse Doppler Radar,” Radar-
77, Proceedings of the 1977 IEE International Radar Conference, October 25–28, London,
pp. 373–378.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:52 460
[12] G. Prado and P. Moroney, “Linear Predictive Spectral Analysis for Sonar Applications,” C. S.
Draper Report R-1109, C. S. Draper Laboratory, Inc., 555 Technology Square, Cambridge,
MA, September 1977.
[13] R. N. McDonough, “Maximum-Entropy Spatial Processing of Array Data,” Geophysics,
Vol. 39, December 1974, pp. 843–851.
[14] D. P. Skinner, S. M. Hedlicka, and A. D. Matthews, “Maximum Entropy Array Processing,”
J. Acoust. Soc. Am., Vol. 66, No. 2, August 1979, pp. 488–493.
[15] A. Van Den Bos, “Alternative Interpretation of Maximum Entropy Spectral Analysis,” IEEE
Trans. Inf. Theory, Vol. IT-17, No. 4, July 1971, pp. 493–494.
[16] J. P. Burg, “Maximum Entropy Spectral Analysis,” Ph.D. Dissertation, Stanford University,
Department of Geophysics, May 1975.
[17] J. H. Sawyers, “The Maximum Entropy Method Applied to Radar Adaptive Doppler Filtering,”
Proceedings of the 1979 RADC Spectrum Estimation Workshop, October 3–5, Griffiss AFB,
Rome, NY.
[18] S. Haykin and S. Kesler, “The Complex Form of the Maximum Entropy Method for Spectral
Estimation,” Proc. IEEE, Vol. 64, May 1976, pp. 822–823.
[19] J. P. Burg, “The Relationship Between Maximum Entropy Spectra and Maximum Likelihood
Spectra,” Geophysics, Vol. 37, No. 2, April 1972, pp. 375–376.
[20] R. H. Jones, “Multivariate Maximum Entropy Spectral Analysis,” Paper presented at the
Applied Time Series Analysis Symposium, May 1976, Tulsa, OK.
[21] R. H. Jones, “Multivariate Autoregression Estimation Using Residuals,” Paper presented at
Applied Time Series Analysis Symposium, May 1976, Tulsa, OK; also published in Applied
Time-Series Analysis, edited by D. Findley, Academic Press, New York, 1977.
[22] A. H. Nutall, “Fortran Program for Multivariate Linear Predictive Spectra Analysis Employing
Forward and Backward Averaging,” Naval Underwater System Center, NUSC Tech. Doc.
5419, New London, CT, May 9, 1976.
[23] A. H. Nutall, “Multivariate Linear Predictive Spectral Analysis Employing Weighted Forward
and Backward Averaging: A Generalization of Burg’s Algorithm,” Naval Underwater System
Center, NUSC Tech. Doc. 5501, New London, CT, October 13, 1976.
[24] A. H. Nutall, “Positive Definite Spectral Estimate and Stable Correlation Recursion for Mul-
tivariate Linear Predictive Spectral Analysis,” Naval Underwater System Center, NUSC Tech.
Doc. 5729, New London, CT, November 14, 1976.
[25] M. Morf, A. Vieira, D. T. Lee, and T. Kailath, “Recursive Multichannel Maximum
Entropy Spectral Estimation,” IEEE Trans. Geosci. Electron., Vol. GE-16, No. 2, April 1978,
pp. 85–94.
[26] R. A. Wiggins and E. A. Robinson, “Recursive Solution to the Multichannel Filtering
Problem,” J. Geophys. Res., Vol. 70, 1965, pp. 1885–1891.
[27] O. N. Strand, “Multichannel Complex Maximum Entropy Spectral Analysis,” 1977 IEEE
International Conference on Acoustics, Speech, and Signal Processing, pp. 736–741.
[28] O. N. Strand, “Multichannel Complex Maximum Entropy (Autoregressive) Spectral Analysis,”
IEEE Trans. Autom. Control, Vol. AC-22, No. 4, August 1977, pp. 634–640.
[29] L. W. Nolte and W. S. Hodgkiss, “Directivity or Adaptivity?” EASCON 1975 Record,
IEEE Electronics and Aerospace Systems Convention, September 1975, Washington, DC,
pp. 35.A–35.H.
[30] H. Cox, “Sensitivity Considerations in Adaptive Beamforming,” Proceedings of the NATO
Advanced Study Institute on Signal Processing with Emphasis on Underwater Acoustics,
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:52 461
CHAPTER
Recent Developments in
Adaptive Arrays 12
' $
Chapter Outline
12.1 Beam Switching . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 463
12.2 Space-Time Adaptive Processing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 465
12.3 MIMO . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 473
12.4 Reconfigurable Antennas and Arrays. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 479
12.5 Performance Characteristics of Large Sonar Arrays . . . . . . . . . . . . . . . . . . . . . . . . . . . . 484
12.6 Adaptive Processing for Monopulse Tracking Antennas . . . . . . . . . . . . . . . . . . . . . . . . 486
12.7 Partially Adaptive Arrays . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 488
12.8 Summary and Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 503
12.9 Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 503
12.10 References . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 504
& %
This chapter presents several innovations that have taken place since the first edition of this
book. Wireless communication applications often resort to very simple beam switching,
in which multiple beams simultaneously exist and the one with the best signal reception
is selected. Moving radars or sonars must deal with clutter as well as interfering signals.
Space-time adaptive processing (STAP) combines a spatial adaptive array with a temporal
adaptive array to improve clutter cancellation and null placement. Another relatively recent
development is multiple input, multiple output (MIMO) antenna array systems where
an adaptive array is used for both transmit and receive to increase channel capacity.
Reconfigurable antennas change their physical layout using switches to adapt for example
the pattern, frequency response, and polarization response to match the desired signal.
Partial adaptivity is of interest when only a portion of the total number of elements is
controlled, thereby reducing the number of processors required to achieve an acceptable
level of adaptive array performance.
FIGURE 12-1 1
Five orthogonal
beams from a
10-element uniform
0.8
array with half-
wavelength spacing.
0.6
AF
0.4
0.2
0
−50 0 50
q (degrees)
by dashed arrows in Figure 12-1). An algorithm continuously evaluates each beam and
selects the one that maintains the highest signal quality. The system scans each beam output
and selects the beam with the largest output power as well as suppresses interference.
Hardware configurations, such as the Rotman lens [1,2] (Figure 12-2) and Butler
matrix [3] (Figure 12-3), have physical beam ports that correspond to a beam pointing
in a direction determined by the passive feed network, element spacing, and number of
elements. Each beam port in the Rotman lens receives a signal from all the elements. The
different path lengths from the elements to the beam ports account for the phase shift
that steers the beams. The Rotman lens in Figure 12-2 has a 16-element array with 11
beam ports. Each beam port has a pattern of a 16-element array steered to predetermined
directions based on the geometry. The Butler matrix is a hardware version of a fast Fourier
transform (FFT) [4]. Each port receives the signal from one of the beams. A switch
selects the desired beam formed by these hardware beamformers. If an array has a digital
beamformer, then the beams are formed and selected in software. Usually, the beams cover
a desired azimuth range.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:55 465
FIGURE 12-3
Butler matrix with
four beams.
1 2 3 4 Elements
Quadrature
hybrid couplers
1R 2L 2R 1L
2va cos φ
fD = (12.1)
λ
where va is the velocity of the aircraft, and φ is the angle measured from the velocity
vector. Figure 12-4 is a plot of the received signal power as a function of azimuth angle
and Doppler frequency. The peak of the Doppler clutter occurs normal to the direction of
velocity and at zero Doppler frequency. The interference occurs at a single angle but over
FIGURE 12-4 A
plot of the target,
ce
clutter, and
feren
interference returns
Clutter as a function of
Inter
fD
Target
Azimuth angle
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:55 466
FIGURE 12-5
Linear array types. a:
Spatial. b: Temporal. w1
d
T
w2
Σ T
w1 w2 wN
Σ wM T
(a) (b)
all frequencies. The motion of a radar platform spreads the clutter in Doppler frequency.
The Doppler frequency from clutter at a specific point on the ground depends on the angle
of the clutter position relative to the heading of the platform. Interference from a discrete
source appears at one angle but is spread over all Doppler frequencies.
A spatial adaptive array weights and combines the signals received by the array
elements at the same instant in time but at different spatial locations separated by distance
d (Figure 12-5a). A temporal adaptive array combines signals received at the same spatial
location but sampled at different instances in time separated by time T (Figure 12-5b).
A displaced phase center antenna (DPCA) cancels clutter induced by platform motion
in an MTI radar [6]. The idea is to make the antenna appear stationary over the transmitted
pulse train by electronically shifting the phase center of the receive aperture backward to
compensate for the forward motion of the moving platform. The DPCA was first envisioned
for a rotating monopulse radar [7]. By adding and subtracting the output from the azimuth
difference channel to the output of the azimuth sum channel, a fore and aft beam are
formed. If the output from the aft beam is subtracted from the output from the fore beam
at the same pointing angle, then the clutter return would be canceled.
A better implementation is based on synthetic aperture radar (SAR). When the velocity
vector of the platform is parallel to the linear array axis, then the pulse repetition frequency
(PRF) is adjusted to the platform velocity so that the first, second, and subsequent elements
at the current pulse appear to move to the respective positions of the second, third, and
subsequent elements at the previous pulse [8]. Figure 12-6 shows a four-element array
split into two three-element arrays: fore and aft. The full aperture transmits a pulse train
with Nt pulses at t = 0. Both the fore and aft arrays receive one pulse, then the array
moves in space a distance d, and the fore and aft arrays receive another pulse. Once the
array moves forward by (M − 1)d, then the full aperture transmits another pulse train of
Nt pulses. Figure 12-7 shows the phase centers of the full, fore, and aft apertures. They
Time
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:55 467
Aft Fore
Beamformer Beamformer
T
+ −
Σ
Doppler
Filter Bank
are equally spaced by d/2. Delaying the fore aperture output by T effectively moves the
fore aperture phase center to correspond to the aft aperture phase center. Thus, the array
moves a distance d/2 in one pulse repetition interval (PRI), so [8]
d
va = (12.2)
2PRI
The voltage output from the aft and fore arrays are given by
π π
AF A (φ) = e− j f d cos φ
+ e− j c f d cos φ f d cos φ
3π
Aft beam: c + ej c
π π
AF F (φ) = e− j f d cos φ f d cos φ f d cos φ
3π
Fore beam: c + ej c + ej c (12.3)
Assuming that the clutter does not change from pulse to pulse, when the fore beam is
delayed by T, then it is identical to the aft beam as shown by
d
− j2π 2va λcos φ
AF F (φ)e− j2π fd T = AF F (φ)e 2va
= AF F (φ)e− j f d cos φ
2π
c (12.4)
−j 3π
f d cos φ − j πc f d cos φ j πc f d cos φ
=e c +e +e
= AF A (φ)
Thus, subtracting the fore aperture output delayed by T from the aft aperture output should
cancel the clutter, which does not change much from pulse to pulse.
DCPA works perfectly when there are no errors. Some practical problems with DCPA
are as follows [9]:
1. The platform velocity does not perfectly match the PRI of the radar.
2. Clutter can change from position to position.
3. Error tolerances in the antenna components cause the signal path at each element to be
different.
4. Antenna calibration is necessary to compensate for thermal noise and aging of
components.
5. Unwanted platform motion causes deviations from the desired velocity vector.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:55 468
An adaptive DCPA algorithm (ADPCA) [10] is not an optimal processor, so it can have
significant signal-to-interference plus noise ratio (SINR) loss. DCPA was implemented
on an advanced development model called Pave Mover and then was followed by the
development of the Joint Surveillance and Target Attack Radar System (Joint STARS)
[11]. Joint STARS uses a 24 ft. long, 2 ft. high phased array antenna (Figure 12-8) mounted
on the forward underfuselage of an Air Force E-8A (Figure 12-9).
STAP enables radars and sonars to detect targets obscured by clutter and jamming. It
was first envisioned by Brennan and Reed [12,13] as an adaptive array for MTI radar. STAP
is an improvement over DCPA, because it integrates the spatial adaptive array with the
temporal adaptive array. It improves clutter cancellation performance and integrates spatial
processing (sidelobe control and null placement) with clutter cancellation. STAP also has
FIGURE 12-8
Photograph of the
phased array used
for Pave Mover and
Joint STARS
(Courtesy of the
National Electronics
Museum).
FIGURE 12-9
Picture of array
mounted beneath
the USAF E-8A
(Courtesy of the
USAF Strategic Air
Command).
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:55 469
d FIGURE 12-10
Diagram of a STAP
antenna.
T T T
T T T
T T T
FIGURE 12-11
Jammer
Jammer
STAP reduces
clutter and jammer
power as seen by
the before (left) and
after (right) Doppler
Doppler
r r
angle–Doppler plots. te te
ut ut
Cl Cl
Target
STAP
Angle Angle
The initial nonadaptive filtering can be either a transformation into the frequency
domain (e.g., by performing an FFT, over pulses in each channel) or a transformation into
beam space (e.g., by performing nonadaptive beamforming in each pulse). We can perform
both spatial and temporal transformations, if desired, or we can eliminate nonadaptive
filtering altogether.
The nonadaptive filtering determines the domain (frequency or time, element or beam)
in which adaptive weight computation occurs. Figure 12-12 is a diagram representing the
data domain for a single range gate after a different type of nonadaptive transform [14,15].
For example, the upper right quadrant represents PRI data that have been transformed
into Doppler space. Thus, each sample is a radar return for a specific Doppler frequency
and receiver element. The lower left quadrant represents element data that have been
transformed into beam space. Each sample is a radar return for a specific PIR and look
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:55 472
Element space
involve either
Element space Element space
one-dimensional
pre-Doppler pre-Doppler
space and time
transforms or
two-dimensional
space–time Spatial 2D Space- Spatial
transform. FFT Time FFT FFT
Beam space
Temporal FFT
direction. For example, a STAP kernel that is adaptive in the frequency domain falls
into the element-space post-Doppler quadrant of the taxonomy. A Doppler filter-bank
FFT transforms the signals from each element. Low-sidelobe Doppler filtering localizes
competing clutter in angle to reduce the extent of clutter to be adaptively canceled. The
adaptation occurs over all elements and a number of Doppler bins. The number of Doppler
bins is a parameter of the element-space post-Doppler algorithm. Factored post-Doppler
algorithms perform spatial adaptation in a single bin and are not adaptive in time.
A partially adaptive STAP algorithm breaks down the fully adaptive STAP prob-
lem into a number of independent, smaller, and more computationally tractable adaptive
problems while achieving near-optimum performance [8,14]. Figure 12-13 shows that a
FIGURE 12-13
STAP enhances the
target while N Original
Element
te
ga
n ge
Doppler Ra
Subarrayed
data cube
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:55 473
partially adaptive algorithm starts by nonadaptive filtering of the input signal data to reduce
dimensionality. Once the input data are transformed and bins and beams (or channels and
pulse repetition intervals) are selected to span the target and interference subspaces, mul-
tiple separate adaptive sample matrix inversion problems are solved, one for each Doppler
frequency bin or pulse repetition interval, across either antenna elements or beams, de-
pending on the domain of the adaptation.
Some current STAP research focuses on the following [16]:
1. Bistatic configurations where the transmit and receive platforms move separately to
keep the receiving platform covert
2. Conformal arrays
3. Nonstationary received signals that arise from bistatic STAP, conformal arrays, terrain
with different reflection coefficients, and clutter motion like vegetation blowing in the
wind.
4. Knowledge-aided STAP, which attempt to remove as much of the heterogeneity from
the snap shots prior to using conventional estimation methods. This is done by using a
priori knowledge, typically stored in databases.
5. STAP applications in sonar, telecommunications, and detection of plastic landmines.
12.3 MIMO
A communications system falls into one of four categories shown in Figure 12-14 based
on the number of antennas used for transmit and receive:
1. A single input, single output (SISO) system has one antenna for transmit and one
antenna for receive. SISO is the most common communications system category and
offers no spatial diversity.
2. A single input, multiple output (SIMO) system has one antenna for transmit and multi-
ple antennas for receive. This configuration is common when an antenna array is used
on receive.
3. A multiple input, single output (MISO) system has multiple antennas for transmit and
one antenna for receive.
4. A MIMO system has multiple antennas for transmit and multiple antennas for receive.
This configuration offers the greatest spatial diversity and control but also has the
greatest hardware cost.
Adaptive antennas have traditionally been SIMO. Some work has been done in MISO
systems, but today considerable attention has fallen on MIMO systems because they have
Receive Antenna
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:55 474
the greatest promise to deliver the highest capacity. MIMO has also found use in acoustic
arrays [17].
In a SISO system with two isotropic point sources separated by R and no multipath,
if the transmitted signal is s, then the received signal, r, is given by
se− j λ
2π
R
r= (12.13)
R
If the signal also takes one bounce from an object with a reflection coefficient, , then the
received signal includes a second path of distance R1 .
se− j λ se− j λ
2π 2π
R R1
r= + (12.14)
R R1
As the number of scattering objects increase, the number of paths increases until the
multipath formulation is given by
se− j λ m se− j
2π M 2π
R λ Rm
r= + (12.15)
R m=1
Rm
Multiple paths turn the channel transfer function, h, into a random Gaussian variable, so
(12.15) becomes
r = hs (12.16)
When the transmit and receive antennas are arrays, then each of the Nt transmit
elements send a signal to each of the Nr receive elements. Figure 12-15 shows a diagram
of a MIMO system [18]. To take interactions between the elements in the transmit and
receive arrays into account, (12.16) is written into matrix form
r = Hs + N (12.17)
where N is a noise vector, and H is the channel matrix given by (Figure 12-16)
⎡ ⎤
h 11 h 12 · · · h 1N
⎢ .. ⎥
⎢ h 21 h 22 . ⎥
H=⎢ ⎢ ..
⎥
⎥ (12.18)
⎣ . ⎦
h M1 · · · hMN
The SVD of H is given by
The columns of USVD are the receive weight vectors corresponding to the singular values,
and columns of VSVD are the transmit weight vectors corresponding to the singular values
[19]. Since the channel is reciprocal, exactly the same weights may be used for transmission
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:55 475
Pulse shaping
RF up-conversion
filtering and matching
I Nt Transmit array
Channel
I Nr Receive array
RF down-conversion
filtering and matching
Matched filtering
and sampling
Space/time decoder
Symbols
hM1
hM2
Nt Nr
hMN
as for reception. Beamforming using the singular vectors as array weights produces eigen-
patterns that create independent (spatially orthogonal) parallel communication channels
in the multipath environment.
Alternatively, (12.17) can be written in terms of power such that the power received
by the array is given by
P = r† r = s† HH† s (12.20)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:55 476
where Vλ are the eigenvectors, and λm are the eigenvalues. The off-diagonal elements
of R H represent the correlation between the transmitted signal streams, with increased
correlation resulting in decreased capacity. The eigenvalue represents the received signal
power level in the the eigenchannel. Once an accurate estimate of H is established, then
the transmitted data, s, is calculated from the received data by inverting the channel matrix.
s = H−1 r (12.22)
In free space with no multipath, the elements of H are the free-space Greens function
⎡ − jk R11 ⎤
e e− jk Rmn e− jk Rmn
⎢ R ···
⎢ 11 Rmn Rmn ⎥ ⎥
⎢ − jk R21 ⎥
⎢e e− jk Rmn ⎥
⎢ ⎥
⎢ R21 Rmn ⎥
H=⎢ ⎢ ⎥ (12.23)
. .. . ⎥
⎢ .. . .
. ⎥
⎢ ⎥
⎢ ⎥
⎢ ⎥
⎣e − jk R mn
e − jk R mn ⎦
···
RM N Rmn
The number of data streams supported must be less than or equal to the rank of H. The
rank of a matrix is the number of nonzero singular values. H is ill conditioned as presented
in (12.23). Increasing the separation between array elements or adding random components
to the matrix elements decreases the matrix condition number and improves the accuracy
of inverting H. Multipath adds random components, so it significantly improves the ability
of inverting H. Low-rank MIMO channels are associated with minimal multipath or large
separation distance between the transmit and receive antennas [20]. The low-rank MIMO
channel is equivalent to a SISO channel with the same total power. High-rank MIMO
channels are associated with a high multipath environment, and the separation distance
between transmit and receive antennas is small. As we pack more antennas into our array,
the capacity per antenna drops due to higher correlation between adjacent elements. MIMO
systems perform best with a full rank channel matrix, which means a low correlation
between signals on the different antennas.
The instantaneous eigenvalues have limits defined by [19]
2 2
Nt − Nr < λn < Nt + Nr (12.24)
When Nt is much larger than Nr , then all the eigenvalues cluster around Nt . Each eigen-
value is nonfading due to the high-order diversity. Thus, the uncorrelated asymmetric
channel with many antennas has a very large theoretical capacity of Nr equal, constant
channels with gains of Nt .
Each element of a MIMO transmit array uses the same frequency spectrum but may
transmit symbols from different modulation schemes and carry independent data streams.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:55 477
The propagation channel is assumed to be Rayleigh and unknown to the transmitter. Each
burst contains a training sequence that allows the receiver to get an accurate estimate of
the propagation conditions. The channel may change from one burst to the next due to
motion of the transmitter or receiver.
When the channel is unknown at the transmitter, the power is uniformly distributed
over the antennas
Pn = P/Nt (12.25)
When the channel is known, then water filling determines the power allocation to the
channels [21]. In water filling, more power is allocated to better subchannels with higher
SNR to maximize the sum of data rates in all subchannels. Since the capacity is a logarith-
mic function of power, the data rate is usually insensitive to the exact power allocation,
except when the SNR is low. Assuming all noise powers to be the same, water filling is the
solution to the maximum capacity, where each channel is filled up to a common level
1
+ Pn = where Pn = P (12.26)
λn
The capacity difference between the known and unknown channels is small for large P.
Thus, the highest-gain channel receives the largest share of the power. In the limit where
the SNR is small ( p < 1/λ2 − 1/λ1 ), only one eigenvalue, the largest, is left.
The capacity of a MIMO system depends on the propagation environment, array
geometry, and antenna patterns. If the sources are uncorrelated and equal power and the
channel is random, then the ergodic (mean) capacity is given by [18]
P †
C = E log2 det I N + HH bits/s/Hz (12.27)
Nt σn2
MIMO capacity increases linearly with the number of elements when the number of
transmit and receive antennas are the same. Winters suggests that Nt should be of the
order 2Nr [22]. When Nt and Nr are large and Nt > Nr , the capacity is [19]
As long as the ratio of Nt /Nr is constant, then the capacity is a linear function of Nr .
The vertically layered blocking structure (V-BLAST) algorithm was developed at Bell
Laboratories for spatial multiplexing [20]. It splits a single high data rate data stream into
Nt lower rate substreams in which the bits are mapped to symbols. These symbols are
transmitted from Nt antennas. Assuming the channel response is constant over the system
bandwidth (flat fading), the total channel bandwidth is a fraction of the original data
stream bandwidth. The receive array has an adaptive algorithm where each substream is a
desired signal while the rest are nulled. As a result, the receive array forms Nt beams while
simultaneously creating nulls. The received substreams are then multiplexed to recover
the original transmitted signal. V-BLAST transmits spatial uncoded data streams without
equalizing the signal at the receiver. VB cannot separate the streams and suffers from
multistream interferences (MSI). Thus, the transmission is unsteady, and forward error
coding is not always able to resolve this issue.
Space-time codes deliver orthogonal and independent data streams. Orthogonal fre-
quency division multiplexing (OFDM) is commonly used in MIMO systems [21]. It splits
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:55 478
the high-speed serial data to be transmitted into many much lower-speed serial data signals
that are sent over multiple channels. The bit or symbol periods are longer, so multipath
time delays are not as deleterious. Increasing the subcarriers and bandwidth makes the
signal immune from multipath. When the subcarrier spacing equals the reciprocal of the
symbol period of the data signals, they are orthogonal. The resulting sinc frequency spectra
have their first nulls at the subcarrier frequencies on the adjacent channels.
The channel matrix is found through measurement or computer model [18]. A switched
array design employs a single transmitter and single receiver to measure H by using high-
speed switches to sequentially connect all array combinations of array elements. Switching
times range from 2 to 100 ms, so the measurement of all antenna pairs is possible before
the channel appreciably changes for most environments. Virtual array instruments displace
or rotate a single antenna element. This method eliminates mutual coupling effects, but a
complete channel matrix measurement takes several seconds or minutes, so the channel
must remain stationary over that time. As a result, virtual arrays work best for fixed indoor
measurements when activity is low.
Channel models compute H based on a statistics. It is common practice to assume
that the transfer function between one transmit and one receive antenna has Rayleigh
magnitude and uniform phase distributions for a non-line-of-sight (NLOS) propagation
environment. Ray tracing programs can be used to calculate H. Figure 12-17 shows an
(a) (b)
R R
T
T
(c) (d)
R R
T T
example of ray tracing from a transmit antenna to a receive antenna in a city when the
transmit antenna and receive antenna are not line of sight [22]. Figure 12-17a is an example
when only two bounces occur between the transmit and receive antenna. The signal paths
in Figure 12-17b take two very different routes. Figure 12-17c and Figure 12-17d are
examples of paths that take many bounces. Signal paths are highly dependent on the
environment and the positions of the transmitters and receivers. Small changes in position
as shown in Figure 12-17 create large changes in H.
U-slot L-stub
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:55 480
MEMS switch
Feed
FIGURE 12-20 A B
Patch with edges Patch
made from variable
conductivity
material.
Feed
from one linear polarization to the orthogonal linear polarization or to circular polariza-
tion (Figure 12-19) [24].
A patch resonates at several frequencies if conducting extensions to the edges are
added or removed [26]. The patch antenna in Figure 12-20 lies on a 3.17 mm thick
polycarbonate substrate (εr = 2.62), and its edges at A and B made from a material with
variable conductivity, such as silicon. The size of the copper patch as well as the width of
A and B are optimized along with the feed location to produce a 50 input impedance at
1.715, 1.763, 1.875, and 1.930 GHz, depending on whether an edge is conductive. Graphs
of the reflection coefficient, s11 , from 1.6 to 2.0 GHz for the four cases are shown in
Figure 12-21. Each combination has a distinct resonant frequency.
The five-element array in Figure 12-22 has rectangular patches made from a perfect
conductor that is 58.7 × 39.4 mm [25]. The 88.7 × 69.4 × 3 mm optically transparent
fused quartz substrate (εr = 3.78) is backed by a perfectly conducting ground plane.
The patch has a thin strip of silicon 58.7 × 2 mm with εr = 11.7 that separates the
narrow right conducting edge of the patch (58.7 × 4.2 mm) from the main patch. A
light-emitting diode (LED) beneath the ground plane illuminates the silicon through small
holes in the ground plane or by making the ground plane from a transparent conductor.
The silicon conductivity is proportional to the light intensity. A graph of the amplitude
of the return loss is shown in Figure 12-23 when the silicon is 0, 2, 5, 10, 20, 50, 100,
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:55 481
0 0
−5 −5
Experimental
Computed
s11 (dB)
s11 (dB)
−10 −10
Computed
−15 −15
Experimental
−20 −20
1.6 1.65 1.7 1.75 1.8 1.85 1.9 1.95 2 1.6 1.65 1.7 1.75 1.8 1.85 1.9 1.95 2
Frequency (GHz) Frequency (GHz)
(a) (b)
0 0
Computed Computed
−5 −5
Experimental Experimental
s11 (dB)
s11 (dB)
−10 −10
−15 −15
−20 −20
1.6 1.65 1.7 1.75 1.8 1.85 1.9 1.95 2 1.6 1.65 1.7 1.75 1.8 1.85 1.9 1.95 2
Frequency (GHz) Frequency (GHz)
(c) (d)
FIGURE 12-21 Reflection coefficient (s11 ) versus frequency for the following: a: No edge.
b: Back edge. c: Front edge. d: Both edges are conductive.
Patch
200, and 1,000 S/m. There is a distinct resonance at 2 GHz when the silicon has no
conductivity. Increasing the conductivity gradually changes the resonance to 1.78 GHz.
As the silicon conductivity increases, the resonant frequency of the patch changes, so the
photoconductive silicon acts as an attenuator. The element spacing in the array is 0.5λ.
Carefully tapering the illumination the LEDs creates so that the silicon conductivity at the
five elements is [16 5 0 5 16] S/m results in the antenna pattern in Figure 12-24. It has a
gain of 10.4 dB with a peak relative sidelobe level 23.6 dB.
An experimental adaptive array was constructed with photoconductive attenuators
integrated with broadband monopole antennas (Figure 12-25) [26]. A close-up of one
of the array elements appears in Figure 12-26. The attenuator consists of a meandering
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:55 482
FIGURE 12-23 1
Plots of the
magnitude of s11 for
silicon conductivities 0.8
of 0, 1, 2, 5, 10, 20,
30, 50, 75, 100, 200,
500, and 1,000 S/m.
0.6
|s11| 10 S/m
0.4
0 S/m
0
1.75 1.8 1.85 1.9 1.95 2 2.05 2.1
Frequency (GHz)
FIGURE 12-24 15
The quiescent
pattern is the Quiescent
10 G = 10.14 dB
dashed line and has
all the conductivities
set to 0. The 5
adapted pattern is
s11 = 24 dB
the solid line and Adapted
Gain (dB)
0
has the silicon
conductivities set to
−5
[16 8 0 8 16] S/m.
−10
−15
−20
−80 −60 −40 −20 0 20 40 60 80
q (degrees)
center conductor in coplanar waveguide mounted on a silicon substrate [27]. This attenu-
ator is flip mounted over a hole in the coplanar waveguide feed of the broadband element
(Figure 12-26). An infrared (IR) LED illuminates the meandering line in the center rect-
angular section to increase in conductivity of the silicon and to create a short between the
center and outer conductors of the coplanar waveguide. Connecting this attenuator to an
array element attenuates the signal received by the element. The element bandwidth was
measured from 2.1 to 2.5 GHz (S11 < −10 dB). Figure 12-27 shows the array patterns the
uniform array (all LEDs off) and for two Chebyshev amplitude tapers found by relating
the LED current to the signal amplitude at the elements.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:55 483
FIGURE 12-25
Picture of
experimental array
6 7 8 with numbered
3 4 5
1 2 elements.
FIGURE 12-26
The attenuator on
the left is flipped and
soldered to the
antenna element on
the right.
0 FIGURE 12-27
Measured low
sidelobe antenna
patterns for the
eight-element array
−10 Uniform
when the diode
Arrray pattern (dB)
−30
25 dB Chebyshev
0 15 30 45 60 75 90
q (degrees)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:55 484
There are limits determining the number of time samples (snapshots), K, which are avail-
able in forming the estimate R̂x x : (1) the time duration limit; and (2) the bandwidth limit,
over which frequency averaging can be done. At broadside the main lobe of a sonar
resolution cell has a cross-range extent given by
λr
x≈ (12.30)
L
where r is the range to the source, L is the array aperture length, and λ is the wavelength
of the transmitted sonar signal. For a source traveling with speed v, the associated bearing
rate is
v
φ̇ = (12.31)
r
Hence, the time spent within a single resolution cell is
λr
x λ
T= = L
= (12.32)
v v L φ̇
The transit time, Ttransit , of an acoustic wave across the face of the array at endfire is
given by
L
Ttransit = (12.33)
v
where v is the acoustic wave velocity in water (nominally 1,500 m/sec). The estimate of
the phase in the cross-spectra will be smeared if one averages over more than one-eighth
the transmit time bandwidth, so
1 v
B< = (12.34)
8Ttransit 8L
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:55 485
The time–bandwidth product given by the product of (12.33) and (12.34) yields the
approximate number of time samples (snapshots) available to form R̂x x , or
2 2
λ 1 1 v f λ
K< T∗B= × = = (12.35)
L φ̇ 8Ttransit 8 f φ̇ L 8φ̇ L
where λ = vf .
Equation (12.35) expresses the inverse square law dependence on array lengths and
thereby forces L to be quite small (e.g., a 200 Hz source moving at 20 knots and a 10 km
distance transiting a 100 wavelength array limits the number of time samples to only 3).
This time–bandwidth product limit dependence on propagation speed illustrates a big
difference between sonar and radar.
The transit of a signal source across a resolution cell introduces eigenvalue spread into
the signal source spectrum. To see how this occurs, introduce the parameter μ, representing
the fraction of motion relative to a beamwidth:
L (cos(φ))
μ= (12.36)
λ
where denotes the beamwidth extent in radians, and φ is the angle off-broadside. When
the source is within a beamwidth, there is a single large eigenvalue associated with R̂x x ;
however, as soon as the motion occupies approximately one resolution cell, the second
eigenvalue becomes comparable. Splitting the source into two sources separated at one-half
the distance traveled approximates this behavior. Solving this two-component eigenvalue
problem leads to
1 πμ (π μ)2
λ̂2 = 1 − sin c2 ≈ (12.37)
4 2 48
The resulting eigenvalue distribution for a 10λ array with L = 500 is illustrated in Fig-
ure 12-28. A linear approximation for λ2 , denoted by “A”, is about 1 dB below the actual
value. A linear fit, denoted by “B,” is also indicated.
SNR = 300 dB, Monte Carlo runs N = 20, L = 500 (10l) FIGURE 12-28
5 Eigenvalues as a
A
Function of Motion
0 for a Moving Source
l1 B
A = −6.87 + 20 Log10 (m) Transiting a Sonar
−5
l2
10 Log10 (lK) - 10 Log10 (N * SNR)
−30
−35
−40
−45
−20 −15 −10 −5 0 10 15
10 Log10 (m), m is motion in beam widths
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:55 486
Most array processing algorithms assume that plane waves impinge on the array. For
large apertures, however, wavefront curvature becomes significant, and plane wave models
are no longer appropriate. The signal source range below which wavefront curvature must
be taken into account is given by the Fresnel range,
2
L2 λ L
RFresnel = = (12.38)
2λ 2 λ
Near-field effects become important even for modest frequencies (i.e., f > 100 Hz), and
range-dependent processing must be introduced. The general solution to this problem,
which is beyond the scope of this text, is known as matched-field processing, where the
full-field propagation model occurs in the array processing [29].
One common way of dealing with the problem of a scarce number of samples in form-
ing the sample covariance matrix is to introduce diagonal loading (discussed in Chapter 5).
The introduction of diagonal loading improves the poor conditioning of the covariance
matrix (although it hinders the detection of weak signals and introduces biases into the
resulting direction-of-arrival estimates).
w = −1 s (12.39)
where is the interference signal covariance matrix, and s is the steering vector for the
target of concern. The next step forms the difference beam (θ0 , f 0 ), where θ0 is the
target azimuth angle, and f 0 is the target Doppler frequency. The basic idea is to form the
difference beam such that the received interference in minimized while satisfying three
distinct constraints as follows:
1. (θ0 , f 0 ) = 0 (the difference beam response is zero at the target location)
(θ0 + θ, f 0 )
2. = ks θ (maintain a constant slope where ks is a slope constant, and
(θ0 + θ, f 0 )
θ is the angle excursion from θ0 where the constraint holds)
(θ0 + θ, f 0 )
3. = −ks θ (negative angle excursion)
(θ0 + θ, f 0 )
The steering vector can is the NM × 1 vector defined as
reflecting an N -element array having a tapped delay line at each element and each tapped
delay line has M time taps with a T sec time delay between each tap. The elements of s are
snm = e j2 p[(n−1) λ d(sin θ0 )−(m−1) f0 T ]
2π
(12.41)
It is convenient to define the NM × 1 vector g(θ0 , f 0 ) as the previous vector s. With
these vectors, the constraints can be written in matrix notation as
HT w = ρ (12.42)
where
⎡ ⎤
gT (θ0 + θ, f 0 )
⎢ ⎥
HT = ⎣ gT (θ0 , f 0 ) ⎦ (12.43)
gT (θ0 − θ, f 0 )
⎡ T ⎤
w g(θ0 + θ, f 0 )
ρ = ks ⎣ 0 ⎦ θ (12.44)
−w g(θ0 − θ, f 0 )
T
The weight vector w that minimizes the difference beam interference power, w H w ,
subject to the constraint (12.42) is
To illustrate these results, consider a 13-element linear array with 14 taps in each
tapped delay line. Assume that the clutter-to-noise ratio is 65 dB per element. The sum
beam weight vector given by (12.1) produces a beam that has an interference plus noise
power after adaptation that is close to the noise floor for all target speeds V such that
0.05 < V/Vb < 0.95, where Vb is the radar blind speed. Likewise, the weight vector given
by (12.47) also produces a difference beam with an adapted, interference plus noise power
close to the noise floor for all target speeds. Figure 12-29 shows the adapted monopulse
pattern, / , for two different target speeds where θ = 0.644×(3 dB -beam angle).
0.3 Desired
0.2
0.1 3 dB angle
for Σ beam
0
0 0.5 1
Normalized Azimuth Angle
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:55 488
FIGURE 12-30
Interference
x1 x2 x3 x4 xN y1
cancelling adaptive
array block diagram. y2
Adaptive
processor
yM
Σ
Main
array m(t)
output +
− Adaptive element array output a(t)
Total
array
output e0 (t)
P0 = Pm − rym
T r*
ym
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:55 489
Each of these approaches will now be discussed to determine the characteristics that typify
these partially adaptive array concepts.
As shown in Chapter 2, null synthesis can be applied to analyzing partial adaptive
nulling. A first step in picking the appropriate elements for a partially adaptive array is to
look at null synthesis using a subset of the elements. The synthesized array factor when a
subset of the elements has variable weights is given by
N Na
j (kxe(n) cos φ+δe(n) )
AF = an e jkxn cos φ − e(n) an e (12.46)
n=1 n=1
where
0 ≤ an ≤ 1.0
0 ≤ δn < 2π
Na = number of adaptive elements
e = vector containing indexes of the adaptive elements
Equation (12.46) can be placed in matrix form
A = b (12.47)
where
⎡ ⎤
ae(1) e jk(e(1)−1)d cos φ1 · · · ae(Na ) e jk(e(Na )−1)d cos φ1
⎢ .. .. .. ⎥
A = ⎣ . . . ⎦
jk(e(1)−1)d cos φ M jk(e(Na )−1)d cos φ M
ae(1) e · · · ae(Na ) e
T
w = e(1) · · · e(Na )
N N
T
b= w n e jk(n−1)d cos φ1 ··· w n e jk(n−1)d cos φ M
n=1 n=1
The Na adaptive weights in an N element partial adaptive array are written as
an (1 − n )e jδn if element n is adaptive
wn = (12.48)
an if element n is not adaptive
FIGURE 12-31
Adapted and 10 Adapted
cancellation patterns Cancellation
superimposed on Quiscent
the quiescent 0
pattern when a null
is synthesized at
Gain (dB)
θ = −21◦ with all
−10
elements adaptive.
−20
−30
−90 −45 0 45 90
f (degrees)
10 Adapted 10 Adapted
Cancellation Cancellation
Quiscent Quiscent
0 0
Gain (dB)
Gain (dB)
−10 −10
−20 −20
−30 −30
10 Adapted 10 Adapted
Cancellation Cancellation
Quiscent Quiscent
0 0
Gain (dB)
Gain (dB)
−10 −10
−20 −20
−30 −30
FIGURE 12-32 Adapted and cancellation patterns superimposed on the quiescent pattern
when a null is synthesized at θ = −21◦ with four of the eight elements adaptive.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:55 491
four-element uniform array factor with its main beam steered to θ = −21◦ . When two
elements on each end of the array are adaptive, then the cancellation pattern has many
lobes of approximately the same height (Figure 12-32b). Making every other element in the
array adaptive induces grating lobes in the cancellation pattern as shown in Figure 12-32c.
Random spacing of the adaptive elements produces the high sidelobe but narrow main beam
cancellation pattern. Figure 12-32d shows the synthesized array factor and cancellation
beam when the random elements are 2, 4, 5, and 8.
The cancellation pattern for any experimental or computed adapted array pattern is
found by subtracting the quiescent electric field, E Quiescent , from the adapted electric field,
EAdapted .
Adapted Adapted
Cancellation Cancellation
10 Quiscent 10 Quiscent
Gain (dB)
Gain (dB)
0 0
−10 −10
−20 −20
−30 −30
−90 −45 0 45 90 −90 −45 0 45 90
q (degrees) q (degrees)
(a) Elements 11, 12, 13, and 14 (b) Elements 1, 2, 7, and 8 adaptive
adaptive complex weights. complex weights.
Adapted Adapted
Cancellation Cancellation
10 Quiscent 10 Quiscent
Gain (dB)
Gain (dB)
0 0
−10 −10
−20 −20
−30 −30
−90 −45 0 45 90 −90 −45 0 45 90
q (degrees) q (degrees)
(c) Elements 11, 12, 13, and 14 (d) Elements 1, 2, 7, and 8
adaptive phase weights. adaptive phase weights.
FIGURE 12-33 Adapted and cancellation patterns for a 24-element array superimposed
on the uniform quiescent pattern when a null is placed at θ = −22◦ .
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:55 492
elements in the center of the array. The nulling is accomplished with the side of the
main beam of the cancellation beam rather than the peak. As a result, the loss in main
beam gain and sidelobe distortion is significant but better than in the eight-element case.
When the tow edge elements on either end of the array are adaptive, then the problems
are similar to the previous case, but the distortion is different, because the cancellation
pattern has many narrow lobes of nearly the same height rather than a uniform cancellation
pattern (Figure 12-33b). Phase only adaptive nulling does not fare any better as shown in
Figure 12-33c and Figure 12-33d.
Partial adaptive nulling experiments were done using the array in Figure 12-25. The
output from the receiver goes to a computer with a genetic algorithm controller. The
genetic algorithm varies the current fed to the LEDs to control signal attenuation at the
adaptive elements. At least 15 dB of attenuation is available at each element. The largest
possible decrease in gain occurs when the LEDs of the four adaptive elements are fed with
250 mA of current. The array becomes a four-element uniform array, so the main beam
should decrease by 6 dB. Figure 12-34 shows a 5.2 dB decrease in the measured main
beam of the far-field pattern. Thus, the adaptive array reduces the desired signal entering
the main beam but cannot place a null in the main beam.
Figure 12-35a is the adapted pattern when one signal is incident on the array at −19◦
and elements 1, 2, 7, and 8 are adaptive. The adaptation lowered the main beam by 3.9 dB
and the sidelobe level by 15.7 dB at −19◦ . Putting a −10 dBm desired signal at 0◦ and a
15 dBm signal incident at −19◦ produces the adapted pattern in Figure 12-35b. Its main
beam goes down by −3.7 dB, and the sidelobe level at −19◦ goes down by 17.1 dB. These
results show that adaptation with the desired signal present was approximately the same
as when it was absent. Nulls can be placed in any sidelobe of the array pattern.
The adapted pattern when two 15 dBm signals are incident at −35◦ and −19◦ is
shown in Figure 12-36. The main beam is reduced by 3.6 dB, while the sidelobe at
−35◦ goes from −14 dB to −23 dB, and the sidelobe at −19◦ goes from −11.8 dB
to −24.3 dB. Figure 12-37 shows the convergence for six independent random runs.
Significant improvement occurs after five iterations.
FIGURE 12-34
Antenna pattern
when the four 0
Quiescent
adaptive elements
are turned off.
−10
Far field (dB)
−20
Adapted
−30
−90 −45 0 45 90
q (degrees)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:55 493
0 Quiescent 0
Quiescent
−10 −10
−20 −20
FIGURE 12-35 Adapted pattern (b) with and (a) without a signal present in the main beam.
FIGURE 12-36
Adapted pattern
0 when signals are
Quiescent
incident on the array
at −35◦ and −19◦ .
−10
Far field (dB)
−20
Adapted
−30
−90 −45 0 45 90
q (degrees)
.25
0 5 10 15 20 25 30
Generation
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:55 494
The adapted pattern for a signal is incident at 35◦ when elements 1 and 8 are adaptive
is shown in Figure 12-38a. The main beam is reduced by 0.9 dB, whereas the sidelobe at
35◦ is reduced by 22 dB. A −10 dBm signal incident at 0◦ and a 15 dBm signal incident at
−18◦ with elements 1 and 7 adaptive result in the adapted pattern in Figure 12-38b. The
main beam reduction is 2.7 dB, and the sidelobe level at −18◦ is reduced by 13.1 dB. The
main beam gain reduction is less when only two elements are adaptive compared with
when four elements are adaptive.
y = QT x (12.50)
R yy = Q† Rx x Q (12.51)
0 Quiescent 0
Quiescent
Far field (dB)
−10 −10
−20 −20
FIGURE 12-38 Adapted pattern when a signal is incident on the array at 35◦ and two
elements are adaptive.
y1 y2 yM Subarray signals
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:55 495
w yopt = αR−1 ∗
yy v y (12.52)
where the beam steering vector for the subarray v y is related to the beam steering vector
for the total array vx by
v y = QT vx (12.53)
Consequently, the resulting array beam pattern can be computed by using the implied
relationship
wx = Qw yopt (12.54)
is used for each subarray so the resulting structure can be regarded as a multibeam pro-
cessor, avoids the grating lobe problem. Vural [37] likewise showed that a beam-based
processor realizes superior performance in partially adaptive operations under diverse in-
terference conditions. To introduce constraints into beam-space adaptive algorithms, the
only difference compared with element-space data algorithms is in the mechanization of
the constraint requirements [34].
Subarray groups for planar array designs are chosen by combining row (or column)
subarrays. This choice is adaptive in only one principal plane of the array beam pattern
and may therefore be inadequate. A more realistic alternative to the row–column subarray
approach is a configuration called the row–column precision array (RCPA) [38], in which
each element signal is split into two paths: a row path and a column path. All the element
signals from a given row or column are summed, and all the row outputs and column
outputs are then adaptively combined. The number of degrees of freedom in the resulting
adaptive processor equals the number of rows plus the number of columns in the actual
array.
When ideal operating conditions are assumed with perfect element channel matching,
simulation studies have shown that the subarray configurations previously discussed yield
array performance that is nearly the same as that of fully adaptive arrays [33]. When the
array elements have independent random errors, however, the resulting random sidelobe
structure severely deteriorates the quality of the array response. This performance dete-
rioration results from the need for precision in the subarray beamforming transformation
since this precision is severely affected by the random element errors.
W = R∗−1
yy r ym (12.58)
P0 = Pm − rTym R−1 ∗
yy r ym (12.59)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:55 497
where Pm is the main array output power. The straightforward evaluation of (12.59) does
not yield much insight into the relationship between adaptive performance and the array
configuration and jamming environment. A concise mathematical expression of (12.59)
is accordingly more desirable to elucidate the problem.
Consider two narrowband jammers of frequencies ω1 and ω2 . The composite signal
vector y is written as
where J1 and J2 represent the reference amplitude and phase, respectively, of the jammers,
and n is the additive noise vector with independent components of equal power Pn . The
spatial vectors v1 , v2 have components given by
2π
Vk, i = exp j (ri · uk ) , i ∈ A, k = 1, 2 (12.61)
λ
where A denotes the subset of M adaptive elements, ri is the ith element location vector,
and uk is a unit vector pointed in the direction of arrival of the kth jammer. For a linear
array having elements aligned along the x-axis, (12.61) reduces to
2π
Vk, i = exp j xi sin θk (12.62)
λ
where θk is the angle of arrival of the kth jammer measured from the array boresight.
The main array output signal can likewise be written as
N
m(t) = n i (t) + J1 e jω1 t h 1 + J2 e jω2 t h 2 (12.63)
i=1
where
N
hk = vk, i , k = 1, 2 (12.64)
i=1
is the main array factor, which is computed by forming the sum of all spatial vector
components for all the array elements i = 1, . . . , N . It follows that the array output signal
is written as
where P1 and P2 denote the jammer power levels, and 1 is an M × 1 vector of ones.
Since (12.58) requires the inverse of R yy , this inverse is explicitly obtained by the
twofold application of the matrix inversion identity (D.10) of Appendix D to (12.66).
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:55 498
Fk = vk, i (12.69)
i∈ A
denotes the complementary array factor that is computed by summing the spatial vector
components for all the elements that are not adaptively controlled, the factor
1 ∗
ρ= v1, i v2, i (12.70)
M i⊂A
is the complex correlation coefficient of the adaptive element spatial vectors, and
−1
Pn Pn Pn2
γ = M(1 − |ρ| ) + + 2
+ (12.71)
P1 P2 M P1 P2
For a fully adaptive array where N = M, then (12.58) reduces to wopt = 1 in the
absence of any constraints on the main lobe. If it is assumed that both jammers are much
stronger than the noise so that P1 , P2 Pn , then (12.58) becomes
F1 − ρ ∗ F2 ∗ F2 − ρ F1 ∗
Wopt | P1 ,P2 →∞ → 1 + v + v (12.72)
M(1 − |ρ|2 ) 1 M(1 − |ρ|2 ) 2
The previous expression shows that the adaptive weights are ill conditioned whenever
|ρ| ≈ 1. This condition physically corresponds to the situation that occurs when the
adaptive array pattern cannot resolve the two jammers. Such a failure to resolve occurs
either because the two jammers are too close together or because they simultaneously
appear on distinct grating lobes.
Recognizing that the output residue power is given by
we see that it follows that the normalized output residue power is expressed as
P0 ∗ Pn |F1 |2 |F2 |2
= N − M + γ |F1 | + |F2 | − 2Re(ρ F1 F2 ) +
2 2
+ (12.74)
Pn M P2 P1
When N = M (a fully adaptive array) then, in the absence of main lobe constraints,
P0 = 0. Equation (12.64) immediately yields an upper bound for the maximum residue
power as
In the case of strong jammers for |ρ| < 1, then (12.64) reduces to
P0 |F1 |2 + |F2 |2 − 2Re(ρ F1 F2∗ )
→N−M+ , |ρ| < 1 (12.76)
Pn P1 ,P2 →∞ M(1 − |ρ|2 )
This expression emphasizes the fact that the residue power takes on a large value whenever
|ρ| = 1. Furthermore, (12.66) surprisingly does not depend on the jammer power levels P1
and P2 but instead depends only on the geometry of the array and the jammer angles. The
maximum residue power upper bound (12.65) does depend on P1 and P2 , however, since
P1 P2
P0max P1 ,P2 →∞ → (|F1 | + |F2 |)2 (12.77)
P1 + P2
The foregoing expressions for output residue power demonstrate the central role the
correlation coefficient plays between the adaptive element spatial vectors in characterizing
the performance of a partially adaptive array. In simple cases, this correlation coefficient
ρ can be related to the array geometry and jammer angles of arrival thereby demonstrat-
ing that it makes a significant difference which array elements are chosen for adaptive
control in a partially adaptive array. In most cases, the nature of the relationship between
ρ and the array geometry is so mathematically obscured that only computer solutions
yield meaningful solutions. For the two-jammer case, it appears that the most favorable
adaptive element locations for a linear array are edge-clustered positions at both ends of
the array [32].
The previous results described for a partially adaptive array and a two-jammer envi-
ronment did not consider the effect of errors in the adaptive weights, even though they
are very important [31]. The variation in the performance improvement achieved with a
partially adaptive array with errors in the nominal pattern weights is extremely sensitive
to the choice of adaptive element location within the entire array. Consequently, for a
practical design the best choice of adaptive element location depends principally on the
error level experienced by the nominal array weights and only secondarily on the optimum
theoretical performance that can be achieved. It has been found that spacing the elements
in a partially adaptive array using elemental-level adaptivity so the adaptively controlled
elements are clustered toward the center is desirable since then the adapted pattern in-
terference cancellation tends to be independent of errors in the nominal adaptive weight
values [31].
Introducing the desired signal steering vector, vd , which satisfies w H vd = 1, the minimum
power distortionless response (MPDR) beamformer is then given by
R−1
x x vd E S+I −1 H
S+I E S+I vd
wMPDR = = (12.82)
vdH R−1x x vd vdH E S+I −1 H
S+I E S+I vd
since vd is in E S+I and hence is orthogonal to E N . A corresponding result for the model
of (12.78) yields
Er r−1 ErH vd
WMPDR = (12.83)
vdH Er r−1 ErH vd
Several variations on the theme of eigendecomposition have been suggested [40–42], but
the previously given development expresses the core ideas about the various schemes that
are developed.
An interesting application of eigendecomposition to the interference nulling problems
facing radio telescopes is discussed in [43]. In radio astronomy applications, detection of
the desired signal cannot take advantage of traditional demodulation–detection algorithms,
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:55 501
and the signal of interest is so minute that it is many orders of magnitude weaker than the
noise power spectral density, leading to minutes or even hours of integration to achieve
a positive detection. Radio frequency interference (RFI) is a growing problem for ra-
dio astronomy; hence, there is strong motivation to exploit adaptive nulling techniques.
Common sources of RFI in radio astronomy include nongeosynchronous satellites and
land-mobile radio. To deal effectively with such dynamic signals, it is necessary to use a
weight update period that is on the order of 10 msec. On the other hand, a self-calibration
process updated about once per minute is highly important for radio astronomy, since it
removes environmental and instrumental errors. Self-calibration methods require nearly
stationary adapted array patterns between self-calibration updates; hence, most standard-
weight update algorithms, which are subject to “weight jitter,” are inappropriate for radio
astronomy. To circumvent the weight jitter problem, the eigendecomposition of (12.80) is
introduced, and the further decomposition of (12.80) is used. Since the desired signal in
radio astronomy is miniscule, the signal term in (12.80) can safely be ignored, leaving
Rx x = E I I E IH + E N N E NH (12.84)
Equation (12.84) is another way of saying that only interference and noise are in the
observations, since the signals are not visible within the timeframe of a weight update.
The term E I I E IH completely describes the interference, and the column span of E I
is referred to as the interference subspace. Likewise, column span of E N is the noise
subspace. Rather than form an estimate of Rx x as is done with most pattern nulling
algorithms, an attractive approach is to use subspace tracking, in which an estimate of the
interference subspace is formed directly using rank-ordered estimates of the eigenvectors
and associated eigenvalues. This approach has been found to be able to form the desired
interference nulls without experiencing any associated weight (or pattern) jitter.
The estimation of E I is accomplished by recourse to the projection approximation
subspace tracking with deflation (PASTd) method [44] in combination with the Gram–
Schmidt orthonormalization procedure to ensure that the resulting eigenvectors are or-
thonormal. It is beyond the scope of this chapter to present the details of this approach;
suffice it to say that accurate interference subspace identification has been achieved for an
interference-to-noise ratio of unity.
Jammer
No. 1 Jammer
No. 2
Target
Jammer
No. J
Main
array
……
1 2 N
J + 1 output for
adaptive
algorithm input
Σ Σ Σ Σ
0
Adaptive-
1
adaptive
2 SMI
array
…
…
output
J
the signals appearing at the (J + 1) ports. The implementation of this approach is shown
in Figure 12-40. The number of elements in the equivalent small array is only (J + 1), so
the equivalent signal covariance matrix is only (J + 1) × (J + 1).
The reason the adaptive–adaptive approach does not degrade the antenna sidelobes is
that the equivalent small array subtracts one auxiliary beam pointing at the jammer from
the main signal channel beam. The gain of the auxiliary beam in the direction of the jammer
equals the gain of the main channel beam sidelobe in the direction of the jammer. As a
result, the subtraction produces a null at the jammer location in the main channel sidelobe.
Further variations of this basic scheme are discussed in the previously noted reference.
resulting in a series of fixed beams using fixed weight vectors. Each of the resulting fixed
beams may then be treated as a single element by weighting it with a single adaptive
weight. Whereas the original array contains N elements, following the division into K
fixed beams, there will be only K adaptive weights where K < N . It is not surprising
that relatively good interference cancellation performance was observed to occur with
subarrays comprised of elements clustered about the edge of the original array.
12.9 PROBLEMS
1. Multiple Beams. Write a program in MATLAB to generate Figure 12-1.
2. MIMO. A MIMO system has a three-element array of isotropic point sources spaced d apart on
transmit and receive. The system operates at 2.4 GHz, and the arrays are 100 m apart and face
each other. Show how the condition number of H changes as the element spacing increases.
3. Partial Adaptive Nulling. Plot the cancellation patterns associated with placing a null at θ =
16.25◦ in the array factor of a 32-element uniform array with λ/2 element spacing. Four different
configurations of eight adaptive element are considered:
a. 1,2,3,4,29,30,31,32
b. 13,14,15,16,17,18,19,20
c. 1,5,9,13,17,21,25,29
d. 2,8,13,16,18,23,24,30
4. Element Level Adaptivity Recognizing that
A−1 b† bA−1
(A + bb† )−1 = A−1 −
(1 + b† A−1 b)
†
let A = Pn I + P1 v1 v1 in (12.66), and apply the previously given matrix identity once. Finally,
apply the previous matrix identity to the resulting expression and show that (12.68) results from
these operations on substitution of (12.66) into (12.58).
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:55 504
12.10 REFERENCES
[1] W. Rotman and R. Turner, “Wide-Angle Microwave Lens for Line Source Applications,” IEEE
Transactions on Antennas and Propagation, Vol. 11, No. 6, 1963, pp. 623–632.
[2] “Rotman Lens Designer,” Remcom, Inc., 2009.
[3] J. Butler and R. Lowe, “Beam-Forming Matrix Simplifies Design of Electronically Scanned
Antennas,” Electronic Design, Vol. 9, April 1961, pp. 170–173.
[4] J. P. Shelton, “Fast Fourier Transforms and Butler Matrices,” Proceedings of the IEEE, Vol. 56,
No. 3, 1968, pp. 350–350.
[5] G. W. Stimson, Introduction to Airborne Radar, Raleigh, NC, SciTech Pub., 1998.
[6] C. E. Muehe and M. Labitt, “Displaced-Phase-Center Antenna Technique,” Lincoln Lab Jour-
nal, Vol. 12, No. 2, 2000, pp. 281–296.
[7] J. K. Day and F. M. Staudaher, “Airborne MTI,” Radar Handbook, M. I. Skolnik, ed., New York,
McGraw Hill, 2008, pp. 3.1–3.34.
[8] R. Klemm and Institution of Engineering and Technology, Principles of Space-Time Adaptive
Processing, 3rd ed., London, Institution of Electrical Engineers, 2006.
[9] Y. Dong, “A New Adaptive Displaced Phase Centre Antenna Processor,” 2008 International
Conference on Radar, September 2–5, 2008, pp. 343–348.
[10] R. S. Blum, W. L. Melvin, and M. C. Wicks, “An Analysis of Adaptive DPCA,” Proceedings
of the 1996 IEEE National Radar Conference, 1996, pp. 303–308.
[11] H. Shnitkin, “Joint STARS Phased Array Radar Antenna,” IEEE Aerospace and Electronic
Systems Magazine, Vol. 9, No. 10, October 1994, pp. 34–40.
[12] L. E. Brennan, and L. S. Reed, “Theory of Adaptive Radar,” IEEE Transactions on Aerospace
and Electronic Systems, Vol. AES-9, No. 2, 1973, pp. 237–252.
[13] L. Brennan, J. Mallett, and I. Reed, “Adaptive Arrays in Airborne MTI Radar,” IEEE Trans-
actions on Antennas and Propagation, Vol. 24, No. 5, 1976, pp. 607–615, 1976.
[14] K. Gerlach and M. J. Steiner, “Fast Converging Adaptive Detection of Doppler-Shifted,
Range-Distributed Targets,” IEEE Transactions on Signal Processing, Vol. 48, No. 9, 2000,
pp. 2686–2690.
[15] J.O. McMahon, “Space-Time Adaptive Processing on the Mesh Synchronous Processor,”
Lincoln Lab Journal, Vol. 9, No. 2, 1996, pp. 131–152.
[16] W. Dziunikowski, “Multiple-Input Multiple-Output (MIMO) Antenna Systems,” Advances
in Direction of Arrival Estimation, S. Chandran (Ed.), Boston, MA, Artechhouse, 2006,
pp. 3–19.
[17] R. Klemm, “STAP Architectures and Limiting Effects,” in ROT Lecture Series 228: Military
Application of Space-Time Adaptive Processing, April 2003, pp. 5-1–5-27.
[18] V. F. Mecca and J. L. Krolik, “MIMO STAP Clutter Mitigation Performance Demonstration
Using Acoustic Arrays,” in 42nd Asilomar Conference on Signals, Systems and Computers,
2008, pp. 634–638.
[19] M. Rangaswamy, “An Overview of Space-Time Adaptive Processing for Radar,” Proceedings
of the International Radar Conference, 2003, pp. 45–50.
[20] J. B. Andersen, “Array Gain and Capacity for Known Random Channels with Multiple Element
Arrays at Both Ends,” IEEE Journal on Selected Areas in Communications, Vol. 18, No. 11,
November 2000, pp. 2172–2178.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 19:55 505
[39] B. D. Van Veen, “Eigenstructure Based Partially Adaptive Array Design,” IEEE Trans. Ant.
& Prop., Vol. AP-36, No. 3, March 1988, pp. 357–362.
[40] C. Lee and J. Lee, “Eigenspace-Based Adaptive Array Beamforming with Robust Capabili-
ties,” IEEE Trans. Ant. & Prop., Vol. AP-45, No. 12, December 1997, pp. 1711–1716.
[41] N. L. Owsley, “Sonar Array Processing,” Array Signal Processing, S. Haykin, ed., Englewood
Cliffs, NJ, Prentice-Hall, 1985.
[42] I. P. Kirstens and D. W. Tufts, “On the Probability Density of Signal-to-Noise Ratio in an
Improved Adaptive Detector,” Proc. ICASSP, Tampa, FL, Vol. 1, 1985, pp. 572–575.
[43] S. W. Ellingson and G.A. Hampson, “A Subspace-Tracking Approach to Interference Nulling
for Phased Array-Based Radio Telescopes,” IEEE Trans. Ant. & Prop., Vol. AP-50, No. 1,
January 2002, pp. 25–30.
[44] B. Yang, “Projection Approximation Subspace Tracking,” IEEE Trans. Signal Processing,
Vol. 43, No. 1, January 1995, pp. 95–107.
[45] E. Brookner and J. Howell, “Adaptive-Adaptive Array Processing,” Proc. IEEE, Vol. 74, No.
4, April 1986, pp. 602–604.
[46] H. Yang and M. A. Ingram, “Subarray Design for Adaptive Interference Cancellation,” IEEE
Trans. Ant. & Prop., Vol. AP-50, No. 10, October 2002. pp. 1453–1459.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 20:1 507
Appendix A: Frequency
Response Characteristics
of Tapped-Delay Lines
The frequency response characteristics of the tapped-delay line filter shown in Figure A-1
can be developed by first considering the impulse response h(t) of the network. For the
input signal x(t) = δ(t) it follows that
2n+1
h(t) = w i δ[t − (i − 1)] (A.1)
i=1
2n+1
{h(t)} = H (s) = w i e−s(i−1) (A.2)
i=1
Equation (A.1) represents a sequence of weighted impulse signals that are summed to
form the output of the tapped-delay line. The adequacy of the tapped-delay line structure
to represent frequency dependent amplitude and phase variations by way of (A.2) depends
on signal bandwidth considerations.
Signal bandwidth considerations are most easily introduced by discussing the con-
tinuous input signal depicted in Figure A-2. With a continuous input signal, the signals
appearing at the tapped-delay line taps (after time t = t0 + 2n has elapsed where t0
denotes an arbitrary starting time) are given by the sequence of samples x(t0 + (i − 1)),
i = 1, 2, . . . , 2n + 1. The sample sequence x(t0 + (i − 1)), i = 1, . . . , 2n + 1, uniquely
characterizes the corresponding continuous waveform from which it was generated pro-
vided that the signal x(t) is band-limited with its highest frequency component f max less
than or equal to one-half the sample frequency corresponding to the time delay, that is,
1
f max ≤ (A.3)
2
Equation (A.3) expresses the condition that must be satisfied in order for a continuous
signal to be uniquely reconstructed from a sequence of discrete samples spaced seconds
apart, and it is formally recognized as the “sampling theorem” [1]. Since the total (two-
sided) bandwidth of a band-limited signal x(t) is BW = 2 f max , it follows that a tapped-
delay line can uniquely characterize any continuous signal having BW ≤ 1/ (Hz), so
1/ can be regarded as the “signal bandwidth” of the tapped-delay line.
Since the impulse response of the transversal filter consists of a sequence of weighted
impulse functions, it is convenient to adopt the z-transform description for the filter transfer
507
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 20:1 508
FIGURE A-1
Transversal filter Input, x(t)
∇ ∇ ∇
Σ Output, y(t)
having 2n + 1
complex weights. w2n + 1
w1 w2 w3
∇ t
0 t0 t0 + 2n
∇
so that
2n+1
{h(t)} = H (z) = w i z −(i−1) (A.5)
i=1
The frequency response of the transversal filter can then be obtained by setting s = jω
and considering how H ( jω) behaves as ω varies. Letting s = jω corresponds to setting
z = e jω , which is a multiple-valued function of ω since it is impossible to distinguish
ω = +π from ω = −π, and consequently
2π
H ( jω) = H jω ± k (A.6)
Equation (A.6) expresses the fact that the tapped-delay line transfer function is a
periodic function of frequency having a period equal to the signal bandwidth capability
of the filter. The periodic structure of H (s) is easily seen in the complex s-plane, which
is divided into an infinite number of periodic strips as shown in Figure A-3 [2]. The strip
located between ω = −π/ and ω = π/ is called the “primary strip,” while all other
strips occurring at higher frequencies are designated as “complementary strips.” Whatever
behavior of H ( jω) obtains in the primary strip, this behavior is merely repeated in each
succeeding complementary strip. It is seen from (A.5) that for 2n + 1 taps in the tapped-
delay line, there will be up to 2n roots of the resulting polynomial in z −1 that describes
H (z). It follows that there will be up to 2n zeros in the transfer function corresponding to
the 2n delay elements in the tapped-delay line.
We have seen how the frequency response H ( jω) is periodic with period determined
by the signal bandwidth 1/ and that the number of zeros that can occur across the signal
bandwidth is equal to the number of delay elements in the tapped-delay line. It remains
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 20:1 509
jw FIGURE A-3
Complex s – plane Periodic structure of
H (s) as seen in the
5p
complex s-plane.
Δ
s1 + j 4p 0
Δ Complementary strip
3p
Δ
s1 + j 2p 0
Δ Complementary strip
p
s1 0 Δ
s Primary strip
p
−
Δ
s1 − j 2p 0
Δ Complementary strip
3p
−
Δ
s1 − j 4p 0
Δ Complementary strip
5p
−
Δ
to show that the resolution associated with each of the zeros of H ( jω) is approximately
1/N , where N = number of taps in the tapped-delay line. Consider the impulse response
of a transversal filter having N taps and all weights set equal to unity so that
−1
1 − e− jωN
N
H ( jω) = e− jωi = (A.7)
i=0
1 − e− jω
Factoring e− jω(/2)N from the numerator and e− jω(/2) from the denominator of (A.7)
then yields
sin [ω(/2)N ]
H ( jω) = exp jω (1 − N ) (A.8)
2 sin [ω(/2)]
sin [ω(/2)N ]
[ω(/2)N ]
= N exp jω (1 − N ) (A.9)
2 sin[ω(/2)]
[ω(/2)]
The denominator of (A.8) has its first zero occurring at f = 1/, which is outside the range
of periodicity of H ( jω) for which 12 BW = 1/2. The first zero of the numerator of (A.8),
however, occurs at f = 1/N so the total frequency range of the principal lobe is just
2/N ; it follows that the 3 dB frequency range of the principal lobe is very nearly 1/N
so the resolution (in frequency) of any root of H ( jω) may be said to be approximately
the inverse of the total delay in the tapped-delay line. In the event that unequal weighting
is employed in the tapped-delay line, the width of the principal lobe merely broadens, so
the above result gives the best frequency resolution that can be achieved.
The discussion so far has assumed that the frequency range of interest is centered
about f = 0. In most practical systems, the actual signals of interest are transmitted with
a carrier frequency component f 0 as shown in Figure A-4. By mixing the actual transmitted
signal with a reference oscillator having the carrier frequency, the transmitted signal can
be reduced to baseband by removal of the carrier frequency component, thereby centering
the spectrum of the information carrying signal component about f = 0. By writing all
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 20:1 510
FIGURE A-4 BW
Signal bandwidth Spectrum of baseband
BW centered about reduced signal
carrier frequency f0 .
f
0 BW f0 BW
f0 − f0 +
2 2
signals as though they had been reduced to baseband, no loss of generality results since this
merely assumes that the signal spectrum is centered about f = 0; the baseband reduced
information-carrying component of any transmitted signal is referred to as the “complex
envelope,” and complex envelope notation is discussed in Appendix B.
A.1 REFERENCES
[1] R. M. Oliver, J. R. Pierce, and C. E. Shannon, “The Philosophy of Pulse Code Modulation,”
Proc. IRE, Vol. 36, No. 11, November 1948, pp. 1324–1331.
[2] B. C. Kuo, Discrete-Data Control System, Prentice-Hall, Englewood Cliffs, NJ, 1970, Ch. 2.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 20:2 511
where the subscript i denotes the signal from the ith source. There is an analytic signal
associated with each real signal that is defined by
z i (t) = [vi (t) + j v̌i (t)] (B.2)
The baseband-reduced complex envelope of vi (t), denoted by ṽi (t), is then defined in
terms of the analytic signal by
Consequently
gives the relationship for real (in-phase) signal components in terms of their corresponding
complex envelopes. Likewise
v̌i (t) = Im{ṽi (t) exp( jωc t)} (B.7)
gives the relationship for imaginary (quadrature) signal components in terms of their
corresponding complex envelopes. Let the output of the kth array element in response to
several signal sources be denoted by xk (t) so that
xk (t) = vi (t − τik ) + n k (t) (B.8)
i
delayed real thermal
signal noise
component component
of ith signal
where τik represents the differential delay of the ith signal to the kth array element relative
to the array phase center, and ṽi (t − τik ) denotes the complex envelope ṽi (t), time delayed
by τik .
The power spectral density relationships that exist among v(t), z(t), and ṽ(t) are
illustrated in Figure B-1. These relationships are readily derived from the power spectral
density relationships between input and output signals for linear filters [3].
It is also instructive to consider two signals v1 (t) and v2 (t) and determine the cross
product relationships that exist between them for the implications this has for covariance
matrix entries. Using the notation of (B.1), we see that it follows
α1 (t)α2 (t)
v1 (t)v2 (t) = {cos[φ1 (t) − φ2 (t) + 1 − 2 ]
2
+ cos[2ωc t + φ1 (t) + φ2 (t) + 1 + 2 ]} (B.10)
α1 (t)α2 (t)
v1 (t)v̌2 (t) = {sin[φ1 (t) − φ2 (t) + 1 − 2 ]
2
+ sin[2ωc t + φ1 (t) + φ2 (t) + 1 + 2 ]} (B.11)
α1 (t) α2 (t)
v̌1 (t)v̌2 (t) = {cos[φ1 (t) − φ2 (t) + 1 − 2 ]
2
− cos[2ωc t + φ1 (t) + φ2 (t) + 1 + 2 ]} (B.12)
∗
ṽ1 (t)ṽ2 (t) = v1 (t)v2 (t) + v̌1 (t)v̌2 (t) + j[v̌1 (t)v2 (t) − v1 (t)v̌2 (t)]
= α1 (t)α2 (t){cos[φ1 (t) − φ2 (t) + 1 − 2 ]
− j sin[φ1 (t) − φ2 (t) + 1 − 2 ]} (B.13)
If the products of the real valued waveforms are now low-pass filtered to remove the double
frequency terms, then
FIGURE B-1
Spectrum
relationships among
the real signal
representation,
the analytic signal
representation,
and the complex
envelope
representation of
the signal v(t).
(a) Spectrum of the
real signal v(t).
(b) Spectrum of z(t),
the analytic signal
associated with v(t).
(c) Spectrum of ṽ(t),
the complex
envelope
representation
of v(t).
B.1 REFERENCES
[1] K. L. Reinhard, “Adaptive Antenna Arrays for Coded Communication Systems,” M. S. Thesis,
The Ohio State University, 1973.
[2] T. G. Kincaid, “The Complex Representation of Signals,” General Electric Company Report
No. R67EMH5, October 1966, and HMED Technical Publications, Box 1122, Le Moyne Ave.,
Syracuse, NY, 13201.
[3] A. Papoulis, Probability, Random Variables, and Stochastic Processes, McGraw-Hill,
New York, 1965, Ch. 10.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 20:2 514
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 20:3 515
Appendix C: Convenient
Formulas for Gradient
Operations
The scalar function trace [ABT ] has all the properties of an inner product, so the formulas
for differentiation of the trace of various matrix products that appear in reference [1] are of
interest for computing the gradients of certain inner products that appear in adaptive array
optimization problems. For many problems, the extension of vector and matrix derivative
operations to matrix functions of vectors and matrices is important; this generalization is
described in reference [2].
∂
tr [X] = I (C.1)
∂X
∂
tr [AX] = AT (C.2)
∂X
∂
tr [AXT ] = A (C.3)
∂X
∂
tr [AXB] = AT BT (C.4)
∂X
∂
tr [AXT B] = BA (C.5)
∂X
∂
tr [AX] = A (C.6)
∂XT
∂
tr [AXT ] = AT (C.7)
∂XT
∂
tr [AXB] = BA (C.8)
∂XT
∂
tr [AXT B] = AT BT (C.9)
∂XT
∂
tr [Xn ] = n(Xn−1 )T (C.10)
∂X
∂
tr [AXBX] = AT XT BT + BT XT AT (C.11)
∂X
∂
tr [AXBXT ] = AT XBT + AXB (C.12)
∂X
∂
tr [XT AX] = AX + AT X = 2AX if A is symmetric (C.13)
∂X
515
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 20:3 516
C.1 REFERENCES
[1] M. Athans and F. C. Schweppe, “Gradient Matrices and Matrix Calculations,” M.I.T. Lincoln
Laboratory, Lexington, MA, Technical Note 1965-53, November 1965.
[2] W. J. Vetter, “Derivative Operations on Matrices,” IEEE Trans. Autom. Control, Vol. AC-15,
No. 2, April 1970, pp. 241–244.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 20:3 517
This appendix summarizes some matrix properties that are especially useful for solving
optimization problems arising in adaptive array processing. Derivations for these relations
as well as a treatment of the basic properties of matrixes may be found in the references
[1]–[3] for this appendix.
The trace of a square matrix is defined to be the sum of its diagonal elements so that
trace(A) = aii (D.1)
i
The following matrix inversion identities are very useful in many of the derivations.
First,
The Schwartz inequality for two scalar functions f (t) and g(t) is as follows:
2
f ∗ (t)g(t)dt ≤ | f (t)|2
dt | g(t)|2 dt (D.13)
Equality in (D.14) obtains if and only if f(t) is a scalar multiple of g(t). Note that (D.14)
is analogous to
Equality in (D.16) obtains if and only if F(t) is a scalar multiple of G(t). Note that (D.16)
is analogous to
D.1 REFERENCES
[1] R. Bellman, Introduction to Matrix Analysis, McGraw-Hill, New York, 1960.
[2] E. Bodewig, Matrix Calculus, North-Holland Publishing Co., Amsterdam, 1959.
[3] K. Hoffman and R. Kunze, Linear Algebra, Prentice-Hall, Englewood Cliffs, NJ, 1961.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 20:4 519
Appendix E: Multivariate
Gaussian Distributions
Useful properties that apply to real Gaussian random vectors are briefly discussed and
extensions to complex Gaussian random vectors are presented in this appendix. Detailed
discussions of these properties will be found in the references [1]–[4] for this appendix.
y = Ax + b (E.5)
where A and b are nonrandom. Then the vector y is also Gaussian and
E{y} = Au + b (E.6)
cov{y} = AVAT (E.7)
E{w} = 0 (E.8)
519
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 20:4 520
and
E{wwT } = I (E.9)
Since any positive semidefinite matrix V can be factored into the form
V = ZZT (E.10)
it follows that a Gaussian vector x having mean u and covariance matrix V can be obtained
from a normalized Gaussian vector w by means of the following linear transformation:
x = Zw + u (E.11)
Let
x1
x2
be a random vector x comprised of two real jointly Gaussian random vectors with mean
x1 u
E = 1 (E.12)
x2 u2
and
x1 V11 V12
cov = (E.13)
x2 V21 V22
Then the conditional distribution for x1 given x2 is also Gaussian with
and
If x1 , x2 , x3 , x4 are all jointly Gaussian scalar random variables having zero mean,
then
The result expressed by (E.16) can be generalized to jointly Gaussian random vectors (all
having zero means). Let the covariance matrix E{wxT } be denoted by
E {wxT } = Vw x (E.17)
and denoted similarly for the other covariances. The desired generalization can then be
expressed as
Quadratic forms are important in detection problems. Let x be a real Gaussian random
vector with zero mean and covariance matrix V. The quadratic form defined by
y = xT KKT x (E.19)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 20:4 521
E{y} + E{xT KKT x} = trace{E [KT xxT K]} = trace(KT VK) = trace(KKT V) (E.24)
and
V
cov(x) = cov(y) = (E.28)
2
and
wi j
cov(xi y j ) = −cov(yi x j ) = − (E.29)
2
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 20:4 522
As a consequence of (E.28) and (E.29), the joint covariance matrix has the special form
x 1 V −W
E [xT yT ] = (E.30)
y 2 W V
C = V + jW (E.32)
z = x + jy (E.33)
If z 1 , z 2 , z 3 , z 4 are all complex jointly Gaussian scalar random variables having zero mean,
then
E {z 1 z 2∗ z 3∗ z 4 } = E {z 1 z 2∗ } E {z 3∗ z 4 } + E {z 1 z 3∗ } E {z 2∗ z 4 } (E.37)
The probability density function for a complex Gaussian random vector z having mean u
and covariance matrix C can be written as
Any linear transformation of a complex Gaussian random vector also yields a complex
Gaussian random vector. For example, if
s = Ax + b (E.39)
and covariance
If z1 and z2 are two jointly Gaussian complex random vectors with mean values u1
and u2 and joint covariance matrix
z1 C11 C12
cov = (E.42)
z2 C21 C22
then the conditional distribution for z1 given z2 is also complex Gaussian with
and
Hermitian forms are the complex analog of quadratic forms for real vectors and play
an important role in many optimization problems. Suppose that z is a complex Gaussian
random vector with zero mean and covariance matrix C. Then the Hermitian form
y = z† KK† z (E.45)
E{y 2 } = E{z† KK† zz† KK† z} = trace[(KK† C)2 ] + [trace(KK† C)]2 (E.48)
and
2
var(y) = trace{(KK† V)2 } (E.52)
α
One result that applies to certain transformations of non-Gaussian complex random
variables is also sometimes useful. Let z = x + jy be a complex random vector that is
not necessarily Gaussian but has the special covariance properties (E.28), (E.29). Now
consider the real scalar,
r = Re{k† z} (E.53)
k† = aT − jbT (E.54)
r = aT x + bT y (E.55)
and
E.3 REFERENCES
[1] T. W. Anderson, An Introduction to Multivariate Statistical Analysis, Wiley, New York, 1959.
[2] D. Middleton, An Introduction to Statistical Communication Theory, McGraw-Hill, New York,
1960.
[3] N. R. Goodman, “Statistical Analysis Based on a Certain Multivariate Complex Distribution
(An Introduction),” Ann. Math. Stat., Vol. 34, March 1963, pp. 152–157.
[4] I. S. Reed, “On a Moment Theorem for Complex Gaussian Processes,” IEEE Trans. Info.
Theory, Vol. IT-8, April 1962.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 20:4 525
Appendix F: Geometric
Aspects of Complex Vector
Relationships
For purposes of interpreting array gain expressions, which initially appear to be quite
complicated, it is very convenient to introduce certain geometric aspects of various com-
plex vector relationships and thereby obtain quite simple explanations of the results. The
discussion in this appendix is based on the development given by Cox [1].
The most useful concept to be explored is that of a generalized angle between two
vectors. Let the matrix C be positive definite Hermitian, and let a and b represent two
N -component complex vectors. A generalized inner product between a and b may be
defined in the complex vector space H (C) as
inner product (a, b) = a† Cb (F.1)
In this complex vector space the cosine-squared function of the generalized angle can be
defined as
|a† Cb|2
cos2 (a, b; C) = (F.2)
{(a† Ca) (b† Cb)}
In the complex vector space H (C), the length of the vector b is (b† Cb)1/2 . The vectors
a and b are orthogonal in H (C) when cos2 (a, b; C) = 0. Likewise the vectors a and
b are in exact alignment (in the sense that one is a scalar multiple of the other) when
cos2 (a, b; C) = 1. Furthermore the sine-squared function can be naturally defined using
the trigonometric identity
Cases that are of special interest for adaptive array applications are C = I, the identity
matrix, and C = Q−1 (where the matrix Q represents the normalized noise cross-spectral
matrix). When C = I it is convenient to simply write cos2 (a, b) in place of cos2 (a, b; I ).
To obtain a clearer understanding of the effect of the matrix Q−1 on the generalized
angle between a and b, it is helpful to further compare cos2 (a, b) and cos2 (a, b; Q−1 ).
To obtain such a comparison involves a consideration of the eigenvalues and correspond-
ing orthonormal eigenvectors of Q denoted by {λ1 , . . . , λ N } and {e1 , . . . , e N }, respec-
tively. Since Q is normalized (and its trace is consequently equal to N ), i λi = N . The
525
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 20:4 526
two vectors a and b can be represented in terms of their projections on the eigenvectors
of Q as:
N
a= h i ei where h i = ei† a (F.5)
i=1
and
N
b= gi ei where gi = ei† b (F.6)
i=1
Substituting (F.5) and (F.6) into (F.2) and using the orthonormal properties of the ei vectors
then yields
2
N
i=1 h i∗ gi
cos2 (a, b) = (F.7)
N
i=1 |h i | 2 N
i=1 |gi | 2
2
N
i=1 h i∗ gi /λi
−1
cos (a, b; Q ) =
2 (F.8)
N
i=1 |h i | 2 /λ
i
N
i=1 |g i | 2 /λ
i
It will be noted from equations (F.7) and (F.8) that the only effect of the matrix Q−1
is to scale each of the factors h i and gi by the quantity (1/λi )1/2 . This scaling weights
more heavily those components of a and b corresponding to small eigenvalues of Q and
weights more lightly those components corresponding to large eigenvalues of Q. Since
the small eigenvalues of Q correspond to components having less noise, it is natural that
these components would be emphasized in an optimization problem.
F.1 REFERENCE
[1] H. Cox, “Resolving Power and Sensitivity to Mismatch of Optimum Array Processors,”
J. Acoust. Soc. Am., Vol. 54, No. 3, September 1973, pp. 771–785.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 20:6 527
Appendix G: Eigenvalues
and Eigenvectors
det[λI − A] = 0 (G.2)
The characteristic polynomial of A is f (λ) = det[λI − A], and the roots of f (λ) are
the eigenvalues of A:
N
N
f (λ) = (λi − λ) = f k λk (G.3)
i=1 k=0
There is an eigenvector associated with each λ. For distinct eigenvalues, the eigenvectors
are orthonormal vectors. For eigenvalues of multiplicity M, we can construct orthonormal
eigenvectors. The eigenvalues of a Hermitian matrix are all real. Furthermore,
The eigenvectors associated with the larger eigenvalues are called principal
eigenvectors.
There are several important properties of eigenvalues and eigenvectors as follows:
1. the highest coefficient in (G.3) is
f N = (−1) N (G.5)
(I + A)−1 = I − A + A2 − A3 + · · · (G.8)
A = V A V HA (G.9)
and
A−1 = V A −1 V A if A is nonsingular (G.10)
In addition,
N
N
1
A= λi vi viH and A−1 = vi viH (G.11)
i=1 i−1
λi
A = Vs s Vs H + V N N V N H (G.13)
where s = diag[λ1 + σ w 2 , . . . , λ D + σ w 2 ]
N = σ w 2 I(N −D)
Vs = [v1 , v2 , . . . , v D ]
V N = [v D+1 , v D+2 , . . . , v N ]
and viH v j = 0, i = D + 1, . . . , N , j = 1, . . . , D.
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 20:6 529
Chapter 1
1. a. 15 MHz
b. 230 MHz
c. 750 MHz
1/2.4166
2. f = CC12 0.909 × 10−6
Chapter 2
9. They are the same: D = 10.
11.
p/2
1 0
array factor in dB
−20
0.8
p 0 −40
0.6
−60
0.4 −80
0 2 4 6 −1 0 1
Element −p/2 u
15.
0
−5
−10
−15
Thinned
−20
AF (dB)
Desired
−25
−30
−35
−40
−45
−1 −0.5 0 0.5 1
u
529
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 20:6 530
17.
0
−10
−20 Adapted
−30 Cancellation Quiescent
AF (dB)
−40
−50
−60
−70
−80
−90
−90 −45 0 45 90
q (degrees)
19.
0 0
−10 −10
Adapted
−20 −20 Adapted
−30 Quiescent −30 Quiescent
Cancellation
AF (dB)
AF (dB)
−40 −40
−50 −50
−60 −60
−70 −70
−80 −80 cancellation
−90 −90
−50 0 50 −50 0 50
q (degrees) q (degrees)
Chapter 3
7.
10
1
5
−21° 61°
Directivity (dB)
0 0.8
−5 0.6
|wn|
−10
0.4
−15
−20 0.2
−25 0
−50 0 50 0 5 10 15 20
q (degrees) Iteration
80
60
40
∠wn (degrees)
20
0
−20
−40
−60
−80
0 5 10 15 20
Iteration
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 20:6 531
Chapter 4
2. a. We know that
vb = vb − w 1 va
where
va = S cos θ + n cos α
and
vb = S cos(β − θ ) + n cos(β − α)
In other words, vb = S[cos(β − α) − w 1 cos θ ] + n[cos(β − α) − w 1 n cos α]
Setting the bracketed quantity associated with S equal to zero, then yields
w 1 cos θ = cos(β − θ )
Chapter 5
1. Applying the Matrix Inversion Lemma, we have
Noting that the quantity inside the square brackets on the right hand side is just a scalar,
it can immediately be placed in the denominator.
Now post-multiply everything by b∗ and noting that Rx−1 ∗
x (k − 1)b = w(k − 1) and
x ↑ (k)Rx−1 ∗ ∗
x (k − 1)b = −ε (k) yields the desired result by merely placing P(k − 1) =
−1
Rx x (k − 1).
Chapter 6
1 t
1. a. Perf Index = w Rnn w + λ[c – It1 w]
2
Taking the derivative with respect to w then yields
Rnn w + λ[−I1 ] = 0
Chapter 7
1. a. For the standard CSLC network (Figure 7-10), it follows that the output is given by
Ouput = b − w 1 x1 − w 2 x2
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 20:6 532
Output = b − u 1 x1 − (x1 − u 2 x2 )u 3
Carrying out the indicated multiplication and collecting like terms then yields
Comparing the outputs for Figure 7-10 and Figure 7-21 then yields the desired
result.
2. a. From Figure 7-11 we find that
y2 = x2 − u 11 x1
v32 = b − u 12 x1
z = v32 − u 22 y2
Use the relationship (7.50) where now v11 = x1 , v21 = x2 , v31 = b, v22 = y2 to
directly obtain (7.61), (7.62) and (7.65).
Chapter 8
3. a. Consider F(x + x)2 where (x + x)2 = (x − sx/ρ)2 = x 2 − 2sx 2 /ρ + s 2 x 2 /ρ 2
but since F(x 2 ) = ρ 2 , we are left with ρ 2 − 2sρ 2 /ρ + s 2 = ρ 2 − 2ρs + s 2
F(x)
b. SL(x) ≡
F(x)−F(x+x)
N
If each search over n + 1 function evaluations yields a function improvement of
(2ρs − s 2 )
2ρs − s 2 , then the function improvement per function evaluation is just
(n + 1)
ρ2
so that SL(x) per function evaluation becomes
2ρs − s 2
(n + 1)
(n + 1)ρ 2
Hence SL(x) =
(2ρs − s 2 )
c. We are interested in showing that ρ02 − ρ12 = 2sρ cos(φ) − s 2
5.
0 dB 30 dB
30 40
−10 0 Jammer 1
−20 −10
0 50 100 150 0 10 20 30 40 50
f (degrees) Generation
15
Signal to jammer ratio (dB)
10
5
0
5
−10
−15
0 10 20 30 40 50
Generation
Chapter 10
1. a. Using |H1 (ω)|{exp j (α1 (ω) + exp j[α2 (ω) − πω ω0
sin θi ]} = 0 implies that the two
angles above must be 180 degrees out of phase with each other. Therefore we may
write
πω
α2 (ω) − sin θi = α1 (ω) ± nπ, n odd
ω0
or
πω
α2 (ω) − α1 (ω) = sin θi ± nπ, n odd
ω0
b. Using |H1 (ω)|{exp jα1 (ω) + exp j[α2 (ω) − πω ω0
sin θs ]} = 1
Since |H1 (ω)|2 = H1 (ω)H∗1 (ω) = 1, it immediately follows that | H1 (ω)|2 AA∗ =
1 where A is just the bracketed expression above. Carrying out the multiplication
πω
of the complex conjugates and recognizing that α2 (ω) − α1 (ω) = sin θi ± nπ
ω0
− jx
e +e
jx
and = cos(x) then yields
2
1
|H1 (ω)|2 =
πω
2 − 2 cos (sin θi − sin θ S )
ω0
from which the desired result follows. Parts (c) and (d) may be shown by using
the considerations given above.
3. A signal and its Hilbert Transform are orthogonal to each other. i.e., E{x(t)x̌(t)} = 0.
This is most easily approached from the point of view of Fourier Transforms.
∞ ∞
∗
x(t)y (t)dt = X ( f )Y ∗ ( f )d f
−∞ −∞
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 20:6 534
Hence,
∞ ∞ ∞
∗ ∗
x(t)x̌ (t)dt = X ( f ) X̌ ( f )dt = X ( f )[ jsgn( f )X ∗ ( f )]dt
−∞ −∞ −∞
∞
= j sgn( f )|X ( f )|2 d f
−∞
but sgn(f) is an odd function of f , while |X ( f )|2 is an even function, so the integrand
is an odd function and the integral is just zero!
Once again we have
∞
E{x̌(t) y̌(s)} = X ( f ) j sgn( f ) Y ∗ ( f )[− j sgn( f )]d f
=∞
∞
= X ( f )Y ∗ ( f )[+sgn2 ( f )]d f
−∞
0 No failures
−10
−20
0 0.2 0.4 0.6 0.8 1
u
Chapter 11
1. a. b.
0 0
q1 = 10°
q1 = −30° −10
q1 = 20°
P (dB)
P (dB)
−10
−20
−20 −30
−90 −45 0 45 90 −90 −60 −30 0 30 60 90
q (degrees) q (degrees)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 20:6 535
c. d.
0 0
−10 −10
P (dB)
P (dB)
−20 −20
−30 −30
−90 −60 −30 0 30 60 90 −90 −60 −30 0 30 60 90
q (degrees) q (degrees)
e.
p/2
p 0
−p/2
Recognizing that x̄0 x̄2 = r (2), x̄0 x̄1 = r (1), and x̄02 = x̄12 = r (0), the above pair of
equations can immediately be put in the desired matrix form.
Chapter 12
3. a. b.
20 x 20 x
0 106.25 0 106.25
Gain (dB)
Gain (dB)
−20 −20
−40 −40
−60 −60
0 45 90 135 180 0 45 90 135 180
f (degrees) f (degrees)
(a) (b)
Monzingo-7200014 book ISBN : XXXXXXXXXX November 24, 2010 20:6 536
c. d.
20 x 20 x
0 106.25 0 106.25
Gain (dB)
Gain (dB)
−20 −20
−40 −40
−60 −60
0 45 90 135 180 0 45 90 135 180
f (degrees) f (degrees)
(c) (d)
Monzingo-7200014 monz7200014˙Ind ISBN : XXXXXXXXXX November 25, 2010 21:32 537
Index
537
Monzingo-7200014 monz7200014˙Ind ISBN : XXXXXXXXXX November 25, 2010 21:32 538
538 Index
Index 539
Joint Surveillance and Target Attack Misadjustment, 165, 167–168, 204–209, Partially known autocorrelation function,
Radar System (Joint STARS), 468 221, 340, 343, 357, 359, 365–367 427–430
J/S sky contour plot, 66–67 Monolithic microwave integrated circuit Particle swarm optimization (PSO), 51,
J Subbeam approach, 501–502 (MMIC), 13–14 348–349
Monopulse array, 56, 58, 486 Pattern forming network, 4–8, 21, 23
K Monopulse tracking antenna, 486–487 multiplication, principle of, 37–38
Kalman filter methods, 284–290 Morgan, D. R., 496 Pave Mover, 468
Kmetzo, J. L., 55 Moving target indicator (MTI), 465–466, Pencil beam, 13
468 Performance comparison of four
L Multichannel processor, 63, 64 processors, 388–393
Laplace transform, 507, 508 Multipath, 4, 30 Performance limits, 21, 29, 81–82, 103
Learning curve, 163, 338, 341 compensation of, 63, 397–405 Performance measure, array gain, 9, 21,
Least mean square (LMS) algorithm, 6, correlation matrix of, 400 66–68, 81–82, 130–134, 137–138
158–169 Multiple input, multiple output (MIMO), Bayes likelihood ratio, 104
with constraints, 213–224 6, 473–479, 503 cancellation ratio, 134
convergence condition for, 162–163 Multiple sidelobe canceller (MSLC), see detection index, 105, 113, 114,
Likelihood function, 99 Coherent sidelobe canceller 116, 123
ratio, 112, 115, 117 Multiple signal classification (MUSIC), generalized likelihood ratio,
Linear random search (LRS), 336–341 423–426 104, 118
Look direction, 170–171, 213, 215, Multivariate Gaussian distributions, generalized signal-to-noise ratio,
231–232 519–524 96, 171
Loop noise, 179–182, 321 MUSIC algorithm, 423–426 likelihood ratio, 103
Lo, Y. T., 51 Mutual coupling, 33, 396–397 maximum likelihood, 90, 99
mean squared error, 9, 90, 91, 121,
M N 154, 159
Main beam constraints, 191–198 Narrowband processing, 62–66, 90–103, minimum noise variance, 90,
Mantey, P. E., 276 143, 295, 445–446 100–101, 293
Marple Algorithm, 456–459 Narrowband signal processing, see Signal minimum signal distortion, 22, 120
Matched filtering, 104 processing, narrowband output power, 138, 139, 143, 185, 187,
Matrix, autocorrelation, 84, 85, 88 Natural modes, 156, 157 188, 200, 214
Butler, 196, 464, 465 Near field, 5, 379–380, 486 signal-to-interference ratio, 251,
correlation, 84–86, 87, 88, 219 Near-field scattering, 4, 63, 279, 400 312, 423
covariance, see Covariance matrix Nike AJAX MPA-4, 13 signal-to-noise ratio, 33, 82, 90, 113,
cross-spectral density, 86, 108 Noise subspace, 424, 441, 501 144, 303
error covariance, 287, 295 Nolen, J. C., 304 Performance surface, 154–158, 165,
Hermitian, 88, 94, 139, 276, 427, 433 Nolen network, 304–315 202–205
quadratic form, 135 Nolte, L. W., 439 quadratic, 154–157, 163
relations, 300–301 Non-Gaussian random signal, 107–108 Periodogram, 422–423, 436, 437
Toeplitz, 65, 85, 240, 269, 427, 457 Nonstationary environment, 4, 154, 190, Perturbation, 122, 203–204, 205,
trace of, 87, 203 220, 288 208–209, 337, 341–346
Maximum entropy method (MEM), 82, Nulling tree, 50 Phase conjugacy, 190
426–436 Null placement, 45, 463 Phase lock loop, 5
extension to multichannel case, Null synthesis, 52–58, 489 Phase mismatch, 60, 409, 413–415
433–436 Phase-only adaptive nulling,
relation to maximum likelihood O 227–228, 358
estimates, 453–454 One-shot array processor, 437, 438 Phase shifters, 307, 309–310, 311, 349,
Maximum likelihood estimates, 109, 118, Open-loop scanning, 74 351, 352, 375–376, 377, 378
120, 140–141, 218, 299–300, Orthogonal frequency division Phase shift sequence, 38
443–449, 453–454 multiplexing (OFDM), 477–478 shifters, 13
McCool, J. M., 165, 336 Orthonormal system, 94–95, 175, Piecemeal adjustment procedure,
Mean squared error, see Performance 177, 183 310–311, 315
measure, mean square error Output noise power, 93, 95, 177, 181, Pilot signals 5, 168–171, 192–194
Micro-electro-mechanical systems 187–188 Planar array, 13, 17, 18, 29, 42–44, 51, 66,
(MEMS), 14, 479–480 78, 496
Minimum noise variance (MV), 90, P Poincare sphere, 126, 131, 144
100–101, 102, 103 Partially adaptive array, 488–503 Polarization, 3, 5, 23, 124–130, 479–480
Minimum variance processor, 291–295 subarray beamforming, 494–496 as a signal discriminant, 230–231
Monzingo-7200014 monz7200014˙Ind ISBN : XXXXXXXXXX November 25, 2010 21:32 540
540 Index
Polarization loss factor, 129 updated covariance matrix inverse, non-Gaussian random, 32
Polarization sensitive arrays, 124–130 277–284 non-random unknown, 31
Polar projection diagram, 46 weighted least squares, 273–277 Signal parameters, 7, 30, 31, 74–75, 104
Powell descent method, 209–211 Recursive equations for prediction error Signal processing, broadband, 62–66,
algorithm, 212–213, 229 filter, 429–430 103–121, 380–395
Powell, J. D., 209–210 Recursive least squares (RLS) algorithm, narrowband, 62–66
Power centroid, 240 276, 278, 301 Signal spectral matching, 129
ratio, 232, 242 Recursive matrix inversion, 263, 267 Signal-to-interference ratio, see
Preadaption spatial filter, 194–196 Reed, I. S., 468 Performance measure,
Prediction error filter, 428–435, 454, 456 Reference signal, 31, 91, 145, 168–169, signal-to-interference ratio
matrix equation of, 453, 509 230–231, 246, 250, 288, 386 Signal-to-noise ratio, see Performance
whitening filter representation of, Reflection coefficients, 399, 402, 428, measure, signal-to-noise ratio
428–429 430–431 Singular values, 474, 476
Prewhitening and matching operation, Relative search efficiency, 360–361 Sonar, 15–19, 484–486
109, 112, 115, 118, 120, 121, 131 Resolution, 22, 23, 33, 45, 422 Sonar arrays, 17–20, 484–486
Propagation delay, 60–61, 63, Retrodirective array, 5 Sonar technology, 15–21
73–74, 405 beam, 183 operating frequencies, 11, 15
compensation of, 405 transmit, 5 Sonar transducers, 15–17
effects, 22 Reverberation, 30, 469 Space-time adaptive processing (STAP),
ideal model of, 32–33, 81, 108 RF interference (RFI), 4, 501 465–473
perturbed, 81, 121–124 Rodgers, W. E., 391–395 Spatial filter, control loop, 196–198,
vector, 86–87 root-MUSIC algorithm, 424–425 232
Pugh, E. L., 256 Rotman lens, 464 correction, 192, 229
Pulse repetition frequency (PRF), 466 preadaptation, 194–196
Pulse repetition interval (PRI), 467, 469, S Spectral factorization, 257
471–472 Sample covariance matrix, 240–244 Speed of light, 10
Sample cross-correlation vector, of sound, 10
Q 244–251 Spherical array, 19
Quadratic form, 87, 93, 177, 267 Sample matrix inversion (SMI) see Direct Spread spectrum, 30, 314
Quadratic performance surface, 154, matrix inversion (DMI) algorithm Spread spectrum modulation, 4
155, 157 Scaled conjugate gradient descent systems, 30
Quadrature hybrid weighting, 62–63 (SCGD), 303 Static errors, 374, 377, 380
processing, 62, 383–388, 394 Schwartz inequality, 105, 114, 116, 120, Steady-state response, 8, 9, 21, 22, 29
Quantization errors, 375–376 517–518 requirements, 29
Quiescent noise covariance matrix, 132, SCR-270, 12 Steepest descent, feedback model of,
178, 188–189 Search loss function, 360, 361 156–158
Seismology, 11 method of, 154–156
R Self-phased array, 5 stability condition for, 158
Radar technology, 10–15 Sensitivity of constrained processor to Steepest descent for power minimization,
operating frequencies, 11 perturbations, 60, 209, 229, 227–228
Radiating elements, 11–14 231–232 Step reversal, 361–362
Radio astronomy, 500–501 of convergence rate to eigenvalue Strand, O. N., 433, 452
Random errors, 374–375 spread, 163, 209, 227, 229, 328 Subarray beamforming, 494–496
Random search algorithm, accelerated of steady-state accuracy to eigenvalue beamspace, 495
(ARS), 341–344 spread, 273 simple, 495
guided accelerated (GARS), 344–346 Sensitivity to eigenvalue spread, Subarray Design, 502–503
linear (LRS), 336–341 229, 262 Subspace Fitting, 441–443
Range discrimination, 10 Sidelobes, 4, 36, 47, 48, 51, 55, 71, 97 Sufficient statistic, 104, 112, 115, 116,
Rassweiler, G. G., 341 Signal aligned array, see Array, signal 117, 118, 123
Receiver operating characteristic aligned Super-resolution, 422
(ROC), 104 Signal bandwidth-propation delay Synthetic aperture radar (SAR), 466
Receivers, 14 product, 60, 61
Reconfigurable antennas, 479–483 Signal distortion, 9, 22, 30, 108, T
Rectilinear coordinate diagram, 47 120, 393 Takao, K., 219
Recursive algorithm, Kalman type, Signal injection, 378, 379 Tapped delay line, 63–65
284–290 Signal model, Gaussian random, 31 processing, 22, 373, 383–388
minimum variance, 291–295 known, 31 frequency response of, 405–407
Monzingo-7200014 monz7200014˙Ind ISBN : XXXXXXXXXX November 25, 2010 21:32 541
Index 541
Introduction to Adaptive Arrays, 2nd Edition is organized as a tutorial, taking the readers by the hand
and leading them through the maze of jargon that often surrounds this highly technical subject. It is
easy to read and easy to follow as fundamental concepts are introduced with examples before more
current developments and techniques are introduced.
Problems at the end of each chapter serve both instructors and professional readers by illustrating and
extending the material presented in the text. Both students and practicing engineers will easily gain
familiarity with the modern contribution that adaptive arrays have to offer practical signal reception
systems.
AUDIENCE
Radar • Sonar • Communications • Seismology • Radio Astronomy
Randy Haupt is a Fellow of the IEEE and Applied Computational Electromagnetics Society (ACES) and
is a Senior Scientist at the Penn State Applied Research Lab. From 1999-2003 he was Professor and
Department Head of ECE at Utah State University. He also was a Professor of EE at the Air Force Academy
and the University of Nevada-Reno. He is a retired Lt. Col. from the US Air Force. He has many journal
articles, conference publications, and book chapters to his credit and is the author of Antenna Arrays:
A Computational Approach by Wiley (2010).
Thomas Miller currently works for Raytheon Corporation where he has been cited for leadership
in advanced early warning surveillance radar and adaptive signal processing. He is the author and
coauthor of multiple journal articles as well as contributor to several books.