0% found this document useful (0 votes)
5 views248 pages

Notes

Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views248 pages

Notes

Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

LECTURE NOTES

on
LINEAR ALGEBRA, LINEAR DIFFERENTIAL EQUATIONS,
and
LINEAR SYSTEMS

A. S. Morse

August 29, 2020


Contents

I Linear Algebra and Linear Differential Equations 1

1 Fields and Polynomial Rings 3

2 Matrix Algebra 7

2.1 Matrices: Notation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7

2.2 Types of Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8

2.3 Determinants: Review . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9

2.4 Matrix Rank . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9

2.5 Basic Matrix Operations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10

2.5.1 Transposition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10

2.5.2 Scalar Multiplication . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11

2.5.3 Addition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11

2.5.4 Multiplication . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12

2.5.5 Linear Algebraic Equations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14

2.5.6 Matrix Inversion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15

2.6 Linear Recursion Equations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17

2.7 Invariants under Premultiplication . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19

2.7.1 Elementary Row Operations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20

2.7.2 Echelon Form . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21

2.7.3 Gauss Elimination . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22

2.7.4 Elementary Column Operations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27

2.7.5 Equivalent Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28

3 Linear Algebra 31

i
3.1 Linear Vector Spaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31

3.1.1 Subspaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 32

3.1.2 Subspace Operations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 33

3.1.3 Distributative Rule . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 34

3.1.4 Independent Subspaces and Direct Sum . . . . . . . . . . . . . . . . . . . . . . . . . . 36

3.1.5 Linear Combination and Span . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 36

3.1.6 Linear Independence . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 36

3.1.7 Basis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 37

3.1.8 Dimension . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 40

3.2 Functions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 40

3.2.1 Linear Transformations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 41

3.2.2 Operations With Linear Transformations . . . . . . . . . . . . . . . . . . . . . . . . . 42

3.2.3 Representations of Linear Transformations . . . . . . . . . . . . . . . . . . . . . . . . 42

3.2.4 Coordinate Transformations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 43

3.2.5 Defining a Linear Transformation on a Basis . . . . . . . . . . . . . . . . . . . . . . . 44

3.2.6 Linear Equations: Existence, Image, Epimorphism . . . . . . . . . . . . . . . . . . . . 45

3.2.7 Linear Equations: Uniqueness, Kernel, Monomorphism . . . . . . . . . . . . . . . . . . 47

3.2.8 Isomorphisms and Isomorphic Vector Spaces . . . . . . . . . . . . . . . . . . . . . . . 48

3.2.9 Endomorphisms and Similar Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . 48

4 Basic Concepts from Analysis 51

4.1 Normed Vector Spaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 51

4.1.1 Open Sets . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 53

4.2 Continuous and Differentiable Functions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 53

4.3 Convergence . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 55

5 Ordinary Differential Equations - First concepts 57

5.1 Types of Equation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 57

5.2 Modeling . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 59

5.3 State Space Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 62

5.3.1 Conversion to State-Space Form . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 63

5.4 Initial Value Problem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 65


5.5 Picard Iterations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 68

5.5.1 Three Relationships . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 69

5.5.2 Convergence . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 70

5.5.3 Uniqueness . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 72

5.5.4 Summary of Findings . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 74

5.6 The Concept of State . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 77

5.6.1 Trajectories, Equilibrium Points and Periodic Orbits . . . . . . . . . . . . . . . . . . . 78

5.6.2 Time-Invariant Dynamical Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 80

5.7 Simulation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 81

5.7.1 Numerical Techniques . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 81

5.7.2 Analog Simulation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 83

5.7.3 Programming Diagrams . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 84

6 Linear Differential Equations 87

6.1 Linearization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 87

6.2 Linear Homogenous Differential Equations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 90

6.2.1 State Transition Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 91

6.2.2 The Matrix Exponential . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 92

6.2.3 State Transition Matrix Properties . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 93

6.3 Forced Linear Differential Equations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 96

6.3.1 Variation of Constants Formula . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 96

6.4 Periodic Homegeneous Differential Equations . . . . . . . . . . . . . . . . . . . . . . . . . . . 97

6.4.1 Coordinate Transformations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 98

6.4.2 Lyapunov Transformations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 99

6.4.3 Floquet’s Theorem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 100

7 Matrix Similarity 103

7.1 Motivation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 103

7.1.1 The Matrix Exponential of a Diagonal Matrix . . . . . . . . . . . . . . . . . . . . . . 104

7.1.2 The State Transition Matrix of a Diagonalizable Matrix . . . . . . . . . . . . . . . . . 104

7.2 Eigenvalues . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 105

7.2.1 Characteristic Polynomial . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 105


7.2.2 Computation of Eigenvectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 106

7.2.3 Semisimple Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 107

7.2.4 Eigenvectors and Differential Equations . . . . . . . . . . . . . . . . . . . . . . . . . . 108

7.3 A - Invariant Subspaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 109

7.3.1 Direct Sum Decomposition of IKn . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 111

7.4 Similarity Invariants . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 112

7.5 The Cayley-Hamilton Theorem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 112

7.6 Minimal Polynomial . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 113

7.6.1 The Miminal Polynomial of a Vector . . . . . . . . . . . . . . . . . . . . . . . . . . . . 114

7.6.2 The Minimal Polynomial of a Finite Set of Vectors . . . . . . . . . . . . . . . . . . . . 114

7.6.3 The Minimal Polynomial of a Subspace . . . . . . . . . . . . . . . . . . . . . . . . . . 114

7.6.4 The Minimal Polynomial of A . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 115

7.6.5 A Similarity Invariant . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 115

7.7 A Useful formula . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 116

7.7.1 Application . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 117

7.8 Cyclic Subspaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 117

7.8.1 Companion Forms . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 118

7.8.2 Rational Canonical Form . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 120

7.9 Coprime Decompositions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 121

7.9.1 The Jordon Normal Form . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 124

7.10 The Structure of the Matrix Exponential . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 126

7.10.1 Real Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 128

7.10.2 Asymptotic Behavior . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 129

7.11 Linear Recursion Equations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 129

8 Inner Product Spaces 131

8.1 Definition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 131

8.1.1 Triangle Inequality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 132

8.2 Orthogonality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 133

8.2.1 Orthogonal Vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 133

8.2.2 Orthonormal Sets . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 133


8.3 Gram’s Criterion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 134

8.4 Cauchy-Schwartz Inequality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 135

8.5 Orthogonalization of a Set of Vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 136

8.6 Orthogonal Projections . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 137

8.7 Bessel’s Inequality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 138

8.8 Least Squares . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 139

9 Normal Matrices 143

9.1 Algebraic Groups . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 143

9.1.1 Orthogonal Group . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 144

9.1.2 Unitary Group . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 145

9.2 Adjoint Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 145

9.2.1 Orthogonal Complement . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 145

9.2.2 Properties of Adjoint Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 146

9.3 Normal Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 146

9.3.1 Diagonalization of a Normal Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . 147

9.4 Hermitian, Unitary, Orthogonal and Symmetric Matrices . . . . . . . . . . . . . . . . . . . . . 149

9.4.1 Hermitian Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 149

9.4.2 Unitary Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 150

9.4.3 Orthogonal Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 150

9.4.4 Symmetric Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 150

9.5 Real Quadratic Forms . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 152

9.5.1 Symmetric Quadratic Forms . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 153

9.5.2 Change of Variables: Congruent Matrices . . . . . . . . . . . . . . . . . . . . . . . . . 153

9.5.3 Reduction to Principal Axes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 153

9.5.4 Sum of Squares . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 154

9.6 Positive Definite Quadratic Forms . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 156

9.6.1 Conditions for Positive Definiteness . . . . . . . . . . . . . . . . . . . . . . . . . . . . 156

9.6.2 Matrix Square Roots . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 157

9.7 Simultaneous Diagonalization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 158

9.7.1 Constrained Optimization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 159


9.7.2 A Special Case . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 159

II Linear Systems 163

10 Introductory Concepts 165

10.1 Linear Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 165

10.2 Continuous-Time Linear Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 166

10.2.1 Time-Invaraint, Continuous-Time, Linear Systems . . . . . . . . . . . . . . . . . . . . 166

10.3 Discrete-Time Linear Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 168

10.3.1 Time-Invariant Discrete Time Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . 168

10.3.2 Sampled Data Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 169

10.4 The Concept of a Realization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 170

10.4.1 Existence of Continuous-Time Linear Realizations . . . . . . . . . . . . . . . . . . . . 171

11 Controllability 173

11.1 Reachable States . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 173

11.2 Controllability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 175

11.2.1 Controllability Reduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 176

11.3 Time-Invariant Continuus-Time Linear Systems . . . . . . . . . . . . . . . . . . . . . . . . . . 178

11.3.1 Properties of the Controllable Space of (A, B) . . . . . . . . . . . . . . . . . . . . . . . 181

11.3.2 Control Reduction for Time-Invariant Systems . . . . . . . . . . . . . . . . . . . . . . 182

11.3.3 Controllable Decomposition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 184

12 Observability 187

12.1 Unobservable States . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 187

12.2 Observability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 189

12.2.1 Observability Reduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 190

12.3 Minimal Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 192

12.4 Time-Invariant Continuous-Time Linear Systems . . . . . . . . . . . . . . . . . . . . . . . . . 193

12.4.1 Properties of the Unobservable Space of (C, A) . . . . . . . . . . . . . . . . . . . . . . 196

12.4.2 Observability Reduction for Time-Invariant Systems . . . . . . . . . . . . . . . . . . . 198

12.4.3 Observable Decomposition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 199


13 Transfer Matrix Realizations 201

13.1 Realizing a Transfer Function . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 201

13.1.1 Realization Using the Control Canonical Form . . . . . . . . . . . . . . . . . . . . . . 202

13.1.2 Realization Using Partial Fractions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 202

13.2 Realizing a p × 1 Transfer Matrix Using the Control Canonical Form . . . . . . . . . . . . . . 204

13.3 Realizing a p × m Transfer Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 205

13.4 Minimal Realizations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 205

13.5 Isomorphic Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 206

14 Stability 209

14.1 Uniform Stability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 209

14.2 Exponential Stability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 210

14.3 Time-Invariant Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 212

14.3.1 The Structure of the Matrix Exponential . . . . . . . . . . . . . . . . . . . . . . . . . 212

14.3.2 Asymptotic Behavior . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 214

14.3.3 Discrete Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 215

14.4 Lyapunov Stability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 216

14.4.1 Linear Time-Varing Differential Equations . . . . . . . . . . . . . . . . . . . . . . . . . 217

14.4.2 Time-Invariant Linear Differential Equations . . . . . . . . . . . . . . . . . . . . . . . 220

14.5 Perturbed Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 222

14.5.1 Nondestabilizing Signals . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 222

15 Feedback Control 225

15.1 State-Feedback . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 226

15.2 Spectrum Assignment . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 227

15.3 Observers . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 229

15.3.1 Definition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 229

15.3.2 Full-State Observers . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 231

15.3.3 Forced Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 232

15.4 Observer-Based Controllers . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 232

Index 236
Part I

Linear Algebra and Linear Differential


Equations

1
Chapter 1

Fields and Polynomial Rings

An algebraic system {e.g., field, ring, vector space} is a set together with certain operations under which
the set is closed. Each algebraic system is defined by a list of postulates. We denote by IR the algebraic
system consisting of the set of all real numbers together with the operations of ordinary addition {+} and
multiplication {·}. Let us note that addition and multiplication in IR are each associative

r1 + (r2 + r3 ) = (r1 + r2 ) + r3

r1 · (r2 · r3 ) = (r1 · r2 ) · r3
and commutative
r1 + r2 = r2 + r1
r1 · r2 = r2 · r1
operations and that multiplication distributes over addition:

r · (r1 + r2 ) = r · r1 + r · r2 .

Furthermore, IR contains distinct additive and multiplicative units, namely 0 and 1 respectively, which satisfy
0 + r = r, 1 · r = r, ∀r ∈ IR. Finally, for each nonzero r ∈ IR, there exist additive and multiplicative inverses,
namely −r and r−1 respectively , such that r + (−r) = 0 and r · r−1 = 1. These properties are enough to
guarantee that IR is an algebraic field.

The set of all complex numbers together with ordinary complex addition and complex multiplication is
also an algebraic field since it has all of the properties outlined above. The complex field is denoted by C.
l
Although much of what follows holds for arbitrary fields, we shall limit the scope of these notes to just IR
and C.l The symbol IK is sometimes used to denote either IR or Cl when discussing a concept valid for both
fields.

Example 1 The set consisting of just the integers 0 and 1 together with addition {+} and multiplication
{·} operations defined by

∆ ∆
0+1 = 1 0·1 = 0
∆ ∆
1+0 = 1 1·0 = 0
∆ and ∆
0+0 = 0 0·0 = 0
∆ ∆
1+1 = 0 1·1 = 1

3
is a field. This field, called GF(2) {Galois Field 2} or the field of integers modulo 2, is used in the coding
of messages in communication systems. Note that all of these definitions are the same as the standard ones
from integer arithmetic except for the definition of 1 + 1. To ensure closure, this sum must be defined to be

either 0 or 1. If 1 + 1 = 1 rather than as above, the resulting algebraic system would not be a field because
there would be no additive inverse for 1; i.e., there would be no number in the set {0, 1} which, when added
∆ ∆
to 1 would give the additive unit 0. In fact if one were to define 1 + 1 = 1 rather than 1 + 1 = 0, what
one would have would be the “Boolean Algebra” upon which digital logic is based; in this case the symbols
1, 0, + and · would denote true, false, or and and respectively.

Consider next the set of all polynomials ρ(s) with indeterminate (i.e., ‘variable’) s and coefficients in IR;
i.e.,
ρ(s) = rn+1 sn + rn sn−1 + . . . + r2 s + r1 ,
where ri ∈ IR, n = deg ρ(s) and rn+1 6= 0. Include in this set all polynomials of degree zero and the
zero polynomial. Let IR[s] denote the algebraic system consisting of all such polynomials together with
ordinary polynomial addition and multiplication. Note that IR[s] possesses all of the algebraic properties
which defined the field IR except one: there are nonzero elements in IR[s] which don’t have multiplicative
inverses {in IR[s]}. In particular, if ρ(s) ∈ IR[s] is any polynomial of positive degree, there is no polynomial
γ(s) ∈ IR[s] such that ρ(s)γ(s) = 1. It is true that ρ(s)(ρ(s)−1 ) = 1, but ρ(s)−1 is not a polynomial since
deg ρ(s) > 0. Algebraic systems such as IR[s] which have all field properties except the one just described
are called rings. The set of all polynomials with coefficients in Cl together with polynomial addition and
multiplication is the ring C[s].
l We often denote either IR[s] or C[s]
l by IK[s].

The following elementary property of IK[s] is important.

Proposition 1 For each pair of polynomials α, β ∈ IK[s] with β 6= 0, there exist unique polynomials γ, δ ∈
IK[s] such that
α = βγ + δ and deg δ < deg β

{The degree of the zero polynomial is defined as −∞.} The quotient γ and remainder δ of the pair (α, β)
are sometimes denoted by q(α, β) and r(α, β) respectively. These polynomials can be computed by ordinary
polynomial division. In case r(α, β) = 0, i.e., α = βγ, we say β divides α; if β divides 1, β is invertible.
Note that the set of all invertible elements in IK[s] together with 0 is precisely the set IK which, in turn, is
a subset of IK[s].


Example 2 If a = s3 + 3s2 + 3s + 1 and b = s2 + 2s + 3, then q = s + 1 and r = −2s − 2. To compute these
polynomials in Matlab one inputs A = [1,3,3,1] and B = [1,2,3] and then uses the command [Q, R] =
deconv(A, B) to get Q=[1,1] and R=[-2, -2]. Try it with your favorite polynomials.

A polynomial α ∈ IK[s] is prime or irreducible if 1) α 6∈ IK and 2) for each β ∈ IK[s] such that β divides
α, either α divides β or β ∈ IK. Irreducibles in IR[s] turn out to be all polynomials of degree one and those
of degree two with roots not in IR; irreducibles in C[s]
l are all polynomials of degree one. An nth degree
polynomial α(s) = kn+1 sn + kn sn−1 + . . . k1 is monic if kn+1 = 1.

A subset I of IK[s] is called an ideal if for each pair of polynomials α, β ∈ I, there follows αρ + βδ ∈
I, ∀ρ, δ ∈ IK[s]. Thus an ideal is a subset of IK[s] closed under ring addition and multiplication by elements
from IK[s]. The set consisting of just the zero polynomial is the zero ideal. Suppose I is a nonzero ideal and
let β ∈ I be any nonzero polynomial of least {finite} degree. Let α ∈ I be arbitrary and use Proposition 1 to
write α = βq(α, β) + r(α, β). Clearly βq(α, , β) ∈ I. Since r(α, β) = α − βq(α, β), it follows that r(α, β) ∈ I.
But by Proposition 1, deg r(α, β) < deg β. Since β has least degree, it must be that r(α, β) = 0; this means
that β divides α. Clearly β is uniquely determined up to an invertible factor in IK[s]. So we may as well
make things unique by taking β to be monic. We summarize:

Proposition 2 Let I be a nonzero ideal of IK[s]. There exists a unique, monic polynomial β ∈ I such that
for each α ∈ I,
α = γβ
for some γ ∈ IK[s].

Let {α1 , α2 , . . . , αm } be a fixed set of polynomials in IK[s]. The greatest common divisor {gcd} of
{α1 , α2 , . . . , αm } is the unique monic polynomial of greatest degree which divides each αi . The least common
multiple {lcm} of {α1 , α2 , . . . , αm } is the unique monic polynomial of least degree which each αi divides. If
α and β are monic polynomials, then

gcd{α, β}lcm{α, β} = αβ (1.1)

Thus the lcm of two polynomials can easily be computed once the corresponding gcd is known. There are
many procedures for computing gcd{α, β}, the Euclidean Algorithm [1] being the classical method. The
Matlab command for computing the gcd G of two polynomials A and B is G = gcd(A, B). Try computing
gcd{s2 + (2 + ǫ)s + 2 + ǫ, s2 + 5s + 6} for ǫ = 0 and also for very small nonzero values of ǫ.

With α, β as above, consider the set I(α, β) ≡ {απ + βµ : π, µ ∈ IK[s]}; clearly I(α, β) is an ideal. Thus
by Proposition 2, there exists a monic polynomial δ ∈ I(α, β) which divides both α and β; thus if γ is the
greatest common divisor of α and β then δ divides γ. In addition, since δ ∈ I(α, β), we can write δ = ασ + βρ
for some σ, ρ ∈ IK[s]; from this and the fact that γ divides both α and β we conclude that γ divides δ. It is
apparent that if δ is monic, then δ = γ and from this there follows a useful fact.

Proposition 3 For each pair of polynomials α, β ∈ IK[s], there exist polynomials σ, ρ ∈ IK[s] such that

gcd{α, β} = ασ + βρ (1.2)

Note that if α and β are coprime {i.e., if gcd{α, β} = 1}, then σ and ρ can be chosen so that ασ + βρ = 1.
For any two polynomials α and β, represented in Matlab format by matrices A and B, the Matlab command
for computing gcd{α, β} together with polynomials σ and ρ satisfying (1.2), is [G, S, R] = gcd(A, B).
Here G, S and R are Matlab representations of gcd{α, β}, σ, and ρ respectively.
Chapter 2

Matrix Algebra

In the sequel are discussed the basic concepts of matrix algebra.

2.1 Matrices: Notation

A p × q matrix M , over a field IK, is an array of numbers of the form


k k12 · · · k1q 
11
 k21 k22 · · · k2q 
M =
 .. .. .. .. 

. . . .
kp1 kp2 · · · kpq

The elements or entries kij in M , are elements of IK; p × q is the size of M ; i.e., M has p rows and q columns
{e.g., the 3rd row of M is [ k31 , k32 , . . . , k3q ]}. We often use the abbreviated notations

M = [ kij ] or M = [ kij ]p×q

M is a square matrix if p=q; i.e., if M has the same number of rows as columns. Otherwise M is rectangular.

It is possible to partition a matrix into ‘submatrices’, by drawing lines between some of its rows and
columns. For examples, the lines within the matrix
 
1 2 7 7 5
 1 2 1 7 1 
A=
 4

3 2 8 1 
6 4 3 5 1

partitions A into the six submatrices, namely

A11 = [1 2] A12 = [7 7 5]

A21 = [1 2] A22 = [1 7 1]
   
4 3 2 8 1
A31 = A23 =
6 4 3 5 1

7
If we wish, we can then write A as  
A11 A12
A =  A21 A22 
A31 A32
There are of course many different ways to partition a matrix.

2.2 Types of Matrices

There are a number of special types of matrices. Here are a few of the most basic.

Zero matrices: Matrices of any size with zeros in every location. Zero matrices are usually denoted by 0.
In Matlab, the command zeros(n, m) produces the n × m zero matrix.

n-vectors: Matrices with n rows and one column. In these notes such matrices are usually denoted by lower
case Roman letters x, y, z, . . .; e.g.,
 
1
1
x= 
5
0

Unit n - vectors: n - vectors with exactly one nonzero entry which in turn is a 1.

Zero vectors: Zero matrices with one column.

Square matrices: Matrices with the same number of rows as columns; e.g.,
 
1 2
A=
3 4

Diagonal matrices: Square matrices with zeros everywhere except along the main diagonal; e.g.,
 
d1 0 ··· 0
 0 d2 ··· 0 
D=
 0 .. 
0 . 0 
0 0 · · · dp p×p


If d = [ d1 d2 · · · dp ], the Matlab statement D = diag(d) inputs the diagonal matrix D above.

Symmetric matrices: Any square matrix A = [ aij ]n×n which is symmetric with respect to its main
diagonal; i.e., aij = aji , ∀, i, j ∈ {1, 2, . . . , n}

Upper triangular matrices: Square matrices with zeros everywhere below the main diagonal; e.g.,

u u12 · · · u1p 
11
 0 u22 · · · u2p 
U =
 .. .. 

0 0 . .
0 0 · · · upp p×p
A lower triangular matrix is similarly defined.
Identity matrices: Diagonal matrices with 1′ s along the main diagonal. Identity matrices are usually
denoted by I; e.g.,  
1 0 0
∆ 
I3×3 = 0 1 0 
0 0 1
The Matlab statement I = eye(n) defines I to be the n × n identity. A list of Matlab commands
defining other elementary matrix can be seen by typing help elmat.

2.3 Determinants: Review

Let A = [ aij ]n×n be a fixed matrix. The cofactor of the ijth element aij is a number āij defined by the
formula

āij = (−1)(i+j) det Āij
where Āij is the (n − 1) × (n − 1) matrix obtained by deleting row i and column j from A. The Laplace
Expansion {along row i} provides a means for computing det A:
n
X
det A = aij āij
j=1

This is useful for hand computation and theoretical development, but not for machine computation, especially
when n is large. The above formula holds for any fixed i.

If A = [ a ]1×1 , then det A = a. From this and the Laplace expansion, one can derive the following general
formula for 2 × 2 matrices:  
a b
det = ad − bc
c d

Example 3 If  
1 4 9
A = 2 5 8
3 6 9
then  
2 8
4̄ = (−1)(1+2) det =6
3 9
and det A = −6. The Matlab commands for computing āij , [ āij ] and det A are cofactor(A, i, j), cofac-
tor(A) and det(A) respectively.

It is worth noting that the determinant of an upper {or lower} triangular matrix is the product of its diagonal
elements. The reader should try to prove this.

2.4 Matrix Rank

Any matrix obtained by deleting any number of rows and columns from a matrix M is called a subarray of
M . For example,  
m21 m23
 m31 m33 
m51 m53
is the subarrray of M = [ mij ]3×5 obtained by deleting rows 1 and 4 as well as column 2.

By a minor of a matrix M is meant the determinant of a square subarray of M ; if the subarray has size
k × k, then the minor is of order k. The largest among the orders of the non-zero minors of M is the rank
of M . Clearly
rank Mp×q ≤ min{p, q} (2.1)

In other words, the rank of a matrix doesn’t exceed its number of rows or number of columns, whichever is
smaller. M has full rank if (2.1) holds with equality.

Example 4 The determinants of the three 2 × 2 subarrays

     
1 2 1 3 2 3
2 4 2 6 4 6

of the matrix  
1 2 3
M=
2 4 6

are all zero. Since [ 4 ] is a 1 × 1 subarray with non-zero determinant, it must be true that rank M = 1. The
Matlab command for computing rank M is {you guessed it !} rank(M).

2.5 Basic Matrix Operations

Here we review four basic matrix operations, namely transposition, scalar multiplication, addition, and
multiplication.

2.5.1 Transposition

The transpose of a matrix


M = [ mij ]p×q

is denoted by M ′ and is defined by the formula


M ′ = [ mji ]q×p

M is obtained by simply interchanging the rows and columns of M . For example, if


 
1 3  
1 2 4
M = 2 7 then ′
M =
3 7 5
4 5

Note that a matrix is symmetric just in case it equals its transpose.

Facts: The determinant of a square matrix is invariant under transposition. From this and the definition
of rank, it is not hard to see why the rank of a matrix {square or not} is invariant under transposition.

The Matlab statement B = A´ defines B as the transpose of A.


2.5.2 Scalar Multiplication

The scalar multiplication of a matrix


M = [ mij ]p×q

by a scalar k ∈ IK is denoted by either kM or M k and is defined by


kM = [ kmij ]p×q

For example, with M as above


 
1k 3k
kM =  2k 7k 
4k 5k
Matlab treats scalars and 1 × 1 matrices as one and the same. In Matlab, the product kM is written as
k*M.

2.5.3 Addition

The sum of two matrices is defined only if both matrices are of the same size. If A = [ aij ]m×n and
B = [ bij ]m×n , then their sum is denoted by A + B and is defined by


A + B = [ aij + bij ]m×n

For example

     
1 1 3 1 7 9 2 9 12
+ =
1 1 1 0 4 4 1 5 5
The Matlab command for a matrix sum is just what you’d expect, namely A + B.

Observe that matrix addition is

associative: A + (B + C) = (A + B) + C

commutative: A + B = B + A

and for each matrix A, there is an ‘additive inverse,’ namely −A, such that A + (−A) = 0.

The idea of matrix addition extends in a natural way to the sum of two partitioned matrices, provided
both matrices are partitioned in the same way. For example, if
   
Bm×n Cm×(q−n) B̄m×n C̄m×(q−n)
Ap×q =   and Āp×q =  
D(p−m)×n E(p−m)×(q−n) D̄(p−m)×n Ē(p−m)×(q−n)

then  
B + B̄ C + C̄
A + Ā =  
D + B̄ E + Ē
2.5.4 Multiplication

Just like matrix addition, matrix multiplication is defined only for pairs of matrices whose sizes are consistent.
In particular, if Ap×m and Bn×q are matrices, their product AB is defined only if the number of columns of
A equals the number of rows of B. That is if n = m. If A = [ aij ]p×m and B = [ bjk ]m×q , there product AB

is a p × q matrix C = AB whose elements are defined by the formulas
m
X
cik = aij bjk i ∈ {1, 2, . . . , p}, k ∈ {1, 2, . . . , q} (2.2)
j=1

When thinking about the product of Ap×n and Bn×q , it is useful to keep in mind 1) that the ‘inner’ integer
n must be the same for multiplication to be defined and 2) that the size of the product Cp×q is determined
by the outer integers p and q.

The Matlab command for computing the product AB is simply A*B. Although Matlab uses (2.2) as
it stands to carry out this computation, it turns out that this formula is not the most useful way to think
about what matrix multiplication means. In the sequel we will elaborate on this point.

Multiplication of a matrix times a vector

Suppose A is an m × n matrix and x is an n - vector. We can write A and x in partitioned form as


 
x1
 x2 
A = [ a1 a2 · · · am ] and x= 
 ... 
xn

where ai is the ith column of A and xi is the ith element of x. The product of A times x turns out to be an
m-vector of the form
Ax = a1 x1 + a2 x2 + · · · + an xn

In other words, the product Ax can be thought of as the sum of n, m - vectors, the ith m - vector being the
product of the ith element of x times the ith column of A. For example, if

   
1
1 0 1 2
2
A = 2 1 1 2 and x= 
4
0 0 1 2
3

then
         
1 0 1 2 11
Ax = 1  2  + 2  1  + 4  1  + 3  2  =  14 
0 0 1 2 10

The reader may wish to verify that this is consistent with the definition of matrix multiplication given by
formula (2.2).

Note that if x were the ith unit n- vector {i.e., that n - vector whose ith row element is 1 and whose
remaining n − 1 elements are all 0}, then multiplication of A times x would have the effect of ‘extracting’
from A its ith column. In other words, Ax = ai .
Multiplication of a matrix times a matrix

Suppose A is an m × n matrix and that B is a n × q matrix. Since the inner integer n is the same for both
matrices, the product AB is a well-defined m × q matrix whose ikth entry is given by (2.2). An especially
useful way to think about what this product is, is as follows:

1. First think of B as an ordered array of q columns, each column being an n - vector; i.e.,

B = [ b1 b2 · · · bq ]
where bi is the ith column of B.
2. For each bi compute the product Abi - remember bi is an n - vector and we’ve already explained how
to do multiplication in this case.
3. The ith column of the product AB is then just Abi . In other words

AB = [ Ab1 Ab2 · · · Abq ] (2.3)

For example, if  
1 1  
1 1 2
A = 2 3 and B=
1 0 0
4 0
then         
1 1   1 1   1 1   2 1 2
1 2 3 1 2  
AB =   2 3  2 3 = 5 2 4
1 0 0
4 0 4 0 4 0 4 4 8

Multiplication of a matrix times a diagonal matrix

Note that if B were a square diagonal matrix with ith diagonal element bii , then in view of (2.3), the ith
column of the product AB would be the ith column of A times bii . Thus if B were the n × n identity In×n ,
then AIn×n = A. Similar reasoning can be used to prove that if A were a square diagonal matrix with ith
diagonal element aii , then the ith row of the product AB would be the ith row of B times aii . Also if A
were the m × m identity Im×m , then Im×m B = B. Therefore, for any m × n matrix M ,

Im×m M = M = M In×n

Properties of matrix multiplication

• Multiplication is an associative operation:

A(BC) = (AB)C

• Multiplication is not a commutative operation. In other words, except for special cases AB 6= BA,
even if A and B are both square. For example
         
1 0 0 0 0 0 0 0 0 0 1 0
= 6= =
0 0 1 0 0 0 1 0 1 0 0 0

Two square matrices of the same size are said to commute if AB = BA.
• Multiplication distributes over addition:

A(B + C) = AB + AC

• From the equation AD = BD, one cannot in general conclude that A = B. For example, if
     
1 −1 −2 2 1
A= , B= , and D=
0 6 9 −3 1
then AD = BD even though A 6= B.
• In general (AB)′ = B ′ A′ . That is, the transpose of the product of two matrices equals the reverse
product of the transposes of the matrices.
• The following formula is useful in many applications:

rank (AB) ≤ min{rank A, rank B}


In other words, the rank of the product of two matrices A and B does not exceed the rank of either A
or B. A proof of this claim can be found in [1] on pages 10-12.
• Another important fact: If A and B are square matrices of the same size, then the determinant of the
product equals the product of the constituent determinants. In other words,

det(AB) = det(A) det(B)

For a proof of this see [1], pages 10-11.

2.5.5 Linear Algebraic Equations

For sure the most common place one encounters matrices is in dealing with linear algebraic equations.
Consider the system of n simultaneous linear algebraic equations in m unknowns:

yi = ai1 u1 + ai2 u2 + · · · + aim um , i ∈ {1, 2, . . . , n} (2.4)



By introducing the matrix A = [ aij ]m×n and the vectors
   
y1 u1
∆  y2  ∆  u2 
y= 
 ..  and u= 
 ... 
.
yn un

it is possible to rewrite (2.4) compactly as


y = Au (2.5)
This is why multiplication of a matrix times a vector is defined as it is: the definition is crafted as it is to
make the conversion from (2.4) to (2.5) come out right.

Suppose that B is another m × n matrix, and that for each n - vector u satisfying (2.5)

z = Bu (2.6)

Then the definition of matrix addition implies that

y + z = (A + B)u
Also if
w = Dy (2.7)
is a matrix equation relating y to w, then (2.5), (2.7) and the definition of matrix multiplication imply that

w = (DA)u

In other words, matrix addition and multiplication are defined so as to be consistent with the manipulations
of systems of simultaneous linear algebraic equations.

At this point some readers may be thinking that matrix algebra is little more than a convenient shorthand
notation for dealing with simultaneous algebraic equations. This would be something like characterizing
natural language as nothing more that a convenient way for humans to communicate. Natural language
provides us not only with a convenient way to communicate, but also with a high-level conceptual framework
for problem solving and working out new ideas in our own minds. The same is true of matrix algebra.

2.5.6 Matrix Inversion

The notions of matrix addition and matrix multiplication are natural generalizations of the notions of
addition and multiplication of real numbers. In the sequel we discuss various generalizations of the arithmetic
operation of inverting {forming the reciprocal of} a nonzero number.

Nonsingular Matrices

In extending matrix inversion to square matrices, the natural generalization of a nonzero number is not
a nonzero matrix but rather a ‘nonsingular’ one. A square matrix A is nonsingular if its determinant is
nonzero. If det A = 0 then A is singular. Note that the rank of an n × n matrix is n if and only if the matrix
is nonsingular.

Remark 1 The logical statement “X is true if and only if Y is true” means that X and Y are equivalent
assertions. The statement also means that “X implies Y and Y implies X.” The symbol iff abbreviates if
and only if. Sometimes we say that “for X to be true it is necessary and sufficient that Y be true.” This is
another way of saying that “X is true if and only if Y is true.”

Matrix Inverses

The concept of a matrix inverse is not limited to just square matrices.

Right and left inverses

An n × m matrix A is said to have a right inverse AR if AAR = In×n where In×n is the n × n identity
matrix. Similarly A has a left inverse AL if AL A = Im×m where Im×m is the m × m identity.

Suppose A has a right inverse AR . Then

n = rank In×n = rank (AAR ) ≤ min{rank A, rank AR } ≤ rank A ≤ min{m, n} ≤ n

Since n appears at both ends of this chain of inequalities, the only way this can be true is if all of the above
inequalities hold with equality. Therefore a necessary condition for A to have a right inverse is that the rank
of A equal the number of rows of A. Later it will be shown that this rank condition is also sufficient for A
to have a right inverse. Similar reasoning can be used to show that a matrix B has a left inverse if and only
if its rank equals the number of its columns.

Note that the transpose of any right inverse of a matrix M is a left inverse of M ′ and visa versa. Because
of this the problems of computing right and left inverses are essentially the same. The left inverse of a
matrix An×m whose rank equals m can be computed in Matlab using the command pinv(A). Try this on
your favorite 3 × 2 full rank matrix.

The inverse of a square matrix

The preceding imply that a given matrix A will have both a left inverse AL and a right inverse AR just in
case the number of rows of A = rank A = the number of columns of A; in other words, just in case A is a
square and nonsingular matrix. Suppose
AAR = In×n
AL A = Im×m
Clearly
AR = Im×m AR = (AL A)AR = AL (AAR ) = AL In×n = AL
so AL = AR . When this is the case we write A−1 for AR and call A−1 the {unique} inverse of A. It follows
from the preceding discussion that A−1 exists if and only if A is a square nonsingular matrix. Note in
addition that A−1 A = AA−1 = In×n .

Inversion formulas

In general left and right inverses are not unique and because of this there are no specific formulas for calcu-
lating them. Nevertheless left and right inverses can be computed when the exist, using Gauss Elimination
– an important algorithm which we will discuss a little later. The Matlab commands for the left inverse of
a n × m matrix A of rank m is is pinv(A). If A = [ ai j ]n×n is a nonsingular matrix, then then there is an
explicit formula for its inverse. In particular
 
1 ′
A−1 = [ āij ]n×n
det A
where āij is the cofactor of aij . {See pages 13-17 of [1] for a proof of this.} While this formula is convenient
for hand calculation and for theoretical development, it is a very poor way to compute an inverse when n
is larger than (say) 3 or 4. The Matlab command for computing the inverse of a square matrix A is either
inv(A) or Aˆ(-1) .

Remark 2 There is an important point here. There are mathematical objects {e.g., the right inverse of a
matrix A} which may or may not exist and which may or may not be unique even when they exist. Moreover
there may be no explicit formula for calculating such an object even though there may well be a computer
algorithm for doing just that. As we shall see there are many mathematical objects including solutions to
some types of linear algebraic equations which share these properties.

The following formulas are especially useful:

• If A and B are both n × n nonsingular matrices then


(AB)−1 = B −1 A−1
• If A is nonsingular, then
(A−1 )′ = (A′ )−1

2.6 Linear Recursion Equations

Later in this course we will be studying various properties of ‘ordinary differential equations.’ Ordinary
differential equations are of major importance to science and engineering because they provide good models
of many different types of natural phenomena. ‘Recursion equations’ are first cousins to ordinary differential
equations. In fact the way a computer calculates a solution to a differential equation is by first converting it
to a recursion equation and then calculating a solution to the latter. Recursion equations are useful for many
other purposes as well. For example, Newton’s method for computing the nth root of a positive number µ
using only addition, multiplication and division is described by the recursion equation
 
1 µ
y(i + 1) = 1 − y(i) +
n ny(i)(n−1)

This is an example of a 1 - dimensional, nonlinear recursion equation. Nonlinear, because for each fixed i,
the right hand side of the equation is a nonlinear function of y(i), and 1-dimensional because each y(i) is a
scalar. For a given initial value for y(1), the equation is ‘solved’ by first setting i = 1 and evaluating y(2),
then setting i = 2 and evaluating y(3) and so on.

Let A be a given n × n matrix. Then


x(i + 1) = Ax(i) (2.8)
is a linear, n - dimensional recursion equation. Note that x(2) = Ax(1), x(3) = Ax(2) = AAx(1), x(4) =
Ax(3) = AAAx(1) and so on. Let us define
∆ ∆ ∆ ∆
A0 = I, A1 = A, A2 = AA, A3 = AAA

and so on. Then it is easy to see that

x(i) = A(i−j) x(j), ∀i≥j≥0

Moreover Ap Aq = A(p+q) for all integers p ≥ 0 and q ≥ 0.

The notation generalizes in a natural way. For example suppose



α(s) = 5s3 + 6s2 + 7s + 1

The notation α(A) is defined to mean



α(A) = 5A3 + 6A2 + 7A + In×n

The n × n matrix α(A) is called a matrix polynomial. This formalism enables us to say (for example) that
if α(s) and β(s) are given polynomials, then

α(A)β(A) = β(A)α(A)

The matlab command for computing Am is Aˆm.

Example 5 Here is a very simple example of where a linear recursion equation arises. Recall that a fixed-rate
mortgage is an agreement to pay back to a lender over a specified period {usually N years}, a certain amount of
borrowed money called the principal P together with interest computed at a fixed annual rate R%. Payments are

usually made monthly, so for our example the number of payment periods is n = 12N . In accordance with standard
practice, the interest rate per month is then
∆ 1 R
r= ×
12 100
Let y(i) denote the remaining unpaid principal after payment i−1 has been made. We thus have at once the boundary
conditions

y(1) = P (2.9)
y(n + 1) = 0 (2.10)

The ith monthly payment π(i), i ∈ {1, 2, . . . , n} to the lender is made up of two components, the ith interest payment
ry(i), and the agreed upon ith principal payment x(i). Thus

π(i) = x(i) + ry(i), i ∈ {1, 2, . . . , n} (2.11)

Clearly
y(i + 1) = y(i) − x(i), i ∈ {1, 2, . . . , n} (2.12)
For historical reasons {lenders didn’t have computers when mortgages were invented}, the x(i) are usually chosen so
that all n monthly payments are the same. In other words so that π(i + 1)) = π(i), ∀ i ∈ {1, 2, . . . , n − 1}. We’re
going to consider a more progressive policy in which each monthly payment is µ times the previous monthly payment,
µ being to agreed upon positive scale factor. That is

π(i + 1) = µπ(i), i ∈ {1, 2, . . . , n − 1} (2.13)



Setting µ = 1 gives the standard mortgage policy. Clearly there are constraints on what other values µ can take, but
lets not worry about them now. Using (2.11)-(2.13) we get

x(i + 1) = r(µ − 1)y(i) + (µ + r)x(i), i ∈ {1, 2, . . . , n − 1} (2.14)

It is now possible to write (2.12) and (2.14) as a 2 - dimensional linear recursion equation of the form
    
y(i + 1) 1 −1 y(i)
 =   , i ∈ {1, 2, . . . , n − 1} (2.15)
x(i + 1) r(µ − 1) µ + r x(i)

Recursing this equation n − 1 times and making use of (2.9) we get


   (n−1)  
y(n) 1 −1 P
 =   
x(n) r(µ − 1) µ+r x(1)

We also know from (2.12) and (2.10) that  


y(n)
[1 −1 ]  =0
x(n)
Combining these two equations we get the linear equation
 (n−1)  
1 −1 P
[ 1 −1 ]    =0
r(µ − 1) µ + r x(1)

which can then be solved for x(1). Once this is accomplished, the computation of y(i) and x(i) can be carried out
by recursively solving (2.15). The end result is a formula for π(i) of the form
 (i−1)  
1 −1 P
π(i) = [ r 1]   , i ∈ {1, 2, . . . , n}
r(µ − 1) µ+r x(1)

We’ve used (2.11) here. The indicated computations can easily be carried out using Matlab. Here are plots of π(·)
∆ ∆
verses payment period for µ = 1 and µ = 1.0008 assuming P = $1000., R = 7% and N = 30 years.
8.5

monthly payment
7.5
1.0008

6.5 1.0000

6
0 50 100 150 200 250 300 350
months

∆ ∆
It seems that µ = 1.0008 has little to offer. How about trying µ = .99?

2.7 Invariants under Premultiplication

To motivate what comes next, let us consider the following problems:

P1: Given an n × m matrix A, develop an algorithm for computing the rank of A.

P2: Given an n × m matrix A and an n -vector b develop an algorithm for

(a) deciding when the equation Ax = b has a solution x, and


(b) finding a solution if one exists.

Key to the solutions of both of these problems are the following facts.

F1: For any nonsingular n × n matrix T ,


rank T A = rank A

F2: For any nonsingular n × n matrix T , the set of all solutions to the equation

T Ax = T b (2.16)

is the same as the set of all solutions to the equation

Ax = b (2.17)

The validity of Fact 1 follows from the chain of inequalities

rank A = rank (T −1 T A) ≤ min{rank T −1 , rank (T A)}


≤ rank (T A) ≤ min{rank T, rank A} ≤ rank A
As for Fact 2, note first that if y is any solution to (2.17), then Ay = b; premultiplication of both sides of
this equation by T shows that y must also satisfy (2.16). Conversely, if z is any solution to (2.16), then
T Az = T b; premultiplication of both sides of this equation by T −1 shows that z must also satisfy (2.17).
Thus Fact 2 is true.

The reason Fact 1 is useful can be explained as follows. As we shall see, for any matrix Mn×p it is
possible to compute a finite sequence of n × n nonsingular matrices E1 , E2 , . . . , Eq which transform M
into a matrix

M ∗ = Eq · · · E2 E1 M

of very special form. So special if fact that the rank of M ∗ can be read off by inspection. Notice that
Eq · · · E2 E1 is a nonsingular matrix, since each Ei is. Thus if we had chosen M to be A and defined T to be

T = Eq · · · E2 E1 , then this process would have transformed A into a new matrix T A whose rank could have
been read off by inspection. Therefore because of Fact 1, the overall approach would have provided a means
for computing the rank of A.

The utility of Fact 2 admits a similar explanation. In this case one would define M = [ A b ]. Trans-

forming M into M ∗ = Eq · · · E2 E1 M would in this case yield the partitioned matrix M ∗ = [ T A T b ] where

T = Eq · · · E2 E1 as before. Because of M ∗ ’s special form {which we have yet to describe}, T A and T b also
have special forms which make it possible to decide by inspection if there is a solution to T Ax = T b and, if
there is, to read off such a solution in a direct manner. Because of Fact 2, this overall process provides an
especially useful algorithm for solving Problem 2.

2.7.1 Elementary Row Operations

There are three very simple kinds of operations which can be performed on the rows of an arbitrary n × m
matrix A. The operations are as follows.

1. Interchange two rows of A

2. Multiply a row of A by a nonzero number.

3. Add to a row of A, a different row of A multiplied by a number.

Corresponding to each elementary row operation, there is a nonsingular matrix E, which when multiplied
times A on the left, produces in a new matrix EA which is the same as that matrix which would have
resulted had the same elementary row operation been performed directly on A. Moreover E is that matrix
which results when the elementary row operation under consideration is applied to the n × n identity matrix
In×n . To illustrate, consider the 3 × 2 matrix
 
a11 a12
A =  a21 a22 
a31 a32

1. The elementary row operation of interchanging rows 1 and 3 yields


 
a31 a32
B =  a21 a22 
a11 a12
The same operation on I3×3 produces the elementary matrix
 
0 0 1
E{r1 ↔r3 } =  0 1 0 
1 0 0

Moreover E{r1 ↔r3 } A = B. Note that E{r1 ↔r3 } is nonsingular.


2. The elementary row operation of multiplying row 2 of A by some number µ 6= 0 yields
 
a11 a12
C =  µa21 µa22 
a31 a32

The same operation on I3×3 produces the elementary matrix


 
1 0 0
E{r2 →µr2 } =  0 µ 0 
0 0 1

and E{r2 →µr2 } A = C. Note that E{r2 →µr2 } is nonsingular because µ 6= 0.


3. The elementary row operation of adding to row 2 of A, row 3 multiplied by some number µ, yields the
matrix  
a11 a12
D =  µa31 + a21 µa32 + a22 
a31 a32
The same operation on I3×3 produces the elementary matrix
 
1 0 0
E{r2 →(r2 +µr3 )} =  0 1 µ 
0 0 1

and E{r2 →(r2 +µr3 )} A = D. Note again that E{r2 →(r2 +µr3 )} is nonsingular.

2.7.2 Echelon Form

By performing a sequence of elementary row operations on a given n × n matrix A, it is possible to transform


A into a new, specially structured matrix A∗ , called an echelon form of A. The structure of A∗ is as follows.

1. Each of the first k ≥ 0 rows of A∗ are nonzero and each of the remaining n − k rows has only zeros in
it.
2. If k > 0, the first nonzero element in each row i ∈ {1, 2, . . . , k} is a 1.
3. If ji is the index of the column within which the ith such 1 appears, then

j1 < j2 < · · · jk

In other words, the nonzero rows of an echelon matrix all begin with 1’s and to the left and below each such
1 are only 0’s. For example
 
0 1 a13 a14 a15 a16
0 0 0 1 a25 a26 
A∗ =  
0 0 0 0 1 a36
0 0 0 0 0 0

is an echelon matrix with k = 3 nonzero rows.

Sometimes it is useful to perform additional elementary row operations on a matrix to bring it to a


reduced echelon form. A reduced echelon form differs from an echelon form in that all of the entries above
the 1’s are also zero. For example, the echelon form A∗ just described can be transformed to a reduced
echelon form which looks like  
0 1 a13 0 0 a16
 0 0 0 1 0 a26 
B∗ =  
0 0 0 0 1 a36
0 0 0 0 0 0

One nice feature of an echelon form {or a reduced echelon form} is that its rank can be determined by
inspection. For example, consider the echelon form A∗ above. Because the last row of this matrix is zero,
any 4 × 4 subarray will have only zeros in its bottom row; thus all minors of order 4 are zero so the rank of
A∗ must be less than 4. On the other hand, the 3 × 3 subarray
 
1 a14 a15
 0 1 a25 
0 0 1

has determinant 1, so A∗ has a third order nonzero minor. Therefore the rank of A∗ is three, which just
so happens to be the number of nonzero rows of A∗ . This is true in general for both echelon and reduced
echelon forms. The rank of any echelon or reduced echelon form A∗ is equal to the number of nonzero rows
of A∗ .

2.7.3 Gauss Elimination

The algorithm which transforms a given matrix A into echelon form be applying a sequence of elementary row
operations, is called Gauss Elimination. Gauss Elimination is in essence nothing more than the formalization
of the method commonly used to solve simultaneous linear algebraic equations by hand. It is without question
the most basic of all algorithms in real numerical analysis.

The easiest way to understand Gauss Elimination is by watching someone work through an example.
The algorithm can be described in words as follows.

1. Start with a given {nonzero} n × m matrix A = [ aij ].

2. Let j denote the index of the first nonzero column of A.

3. Let i be an integer for which aij 6= 0.

4. Elementary Row Interchange: Interchange rows 1 and i obtaining a matrix B whose 1j entry
b1j = aij 6= 0.

5. Pivot on Element 1j:


1
(a) Elementary Row Multiplication: Multiply row 1 of B by b1j thereby obtaining a matrix C
whose first j − 1 columns are zero and c1j = 1.
(b) Elementary Row Additions: For the first nonzero entry ckj below c1j {i.e., k > 1}, add −ckj
times row 1 of C to row k. Repeat this process successively on each resulting matrix thereby
obtaining in at most n − 1 steps, a matrix D whose first j − 1 columns are zero and whose jth
column has a 1 in row 1 and zeros below the 1.
6. Elementary Row Interchange: Let k be the first column of D beyond the jth which has a nonzero
entry not in its first row. {If k does not exist, stop because D is in echelon form}. Let h be any integer
greater than 1 for which dhk 6= 0. Interchange rows 2 and h obtaining a matrix F whose 2k entry
f2k = dhk 6= 0.
7. Pivot on Element 2k:
1
(a) Elementary Row Multiplication: Multiply row 2 of F by f2k thereby obtaining a matrix G
whose first j − 1 columns are zero, whose jth column has a 1 in the first row and zeros below,
and whose second row has as its first nonzero element a 1 in column k.
(b) Elementary Row Additions: Successively eliminate the nonzero entries below g2k , as in step
5b above1.
8. Continue these steps until all columns have been processed. The result is an echelon matrix.

Here is an example.

     
0 0 0 3 1 0 2 4 2 2 0 1 2 1 1
∆ 0 2 4 2 2 r1 ↔r2 0 0 0 3 1 r1 → 12 r1 0 0 0 3 1
A=  −−−−−−−−→   −−−−−−−−→  
0 1 2 1 0 0 1 2 1 0 0 1 2 1 0
0 3 0 1 0 0 3 0 1 0 0 3 0 1 0

     
0 1 2 1 1 0 1 2 1 1 0 1 2 1 1
r3 →(r3 −r1 )  0 0 0 3 1  r4 →(r4 −3r1 )  0 0 0 3 1 r2 ↔r4 0 0 −6 −2 −3 
−−−−−−−−→   −−−−−−−−→   −−−−−−−−→  
0 0 0 0 −1 0 0 0 0 −1 0 0 0 0 −1
0 3 0 1 0 0 0 −6 −2 −3 0 0 0 3 1

     
0 1 2 1 1 0 1 2 1 1 0 1 2 1 1
r2 →− 61 r2 1 1 1 1 r3 → 31 r3 1 1
0 0 1 3  2r3 ↔r4 0 0 1 3  2 0 0 1 3 2 
−−−−−−−−→   −−−−−−−−→   −−−−−−−−→  1 
0 0 0 0 −1 0 0 0 3 1 0 0 0 1 3
0 0 0 3 1 0 0 0 0 −1 0 0 0 0 −1

 
0 1 2 1 1
1 1
r4 →−r4 0 0 1 3 2  ∗
−−−−−−−−→  1 =A
0 0 0 1 3
0 0 0 0 1
As noted in footnote 1, a reduced echelon form B ∗ of A can be obtained by performing additional elementary
row additions as part of each pivot. Alternatively B ∗ can be obtained by first computing A∗ and then
performing additional elementary row additions A∗ to eliminate the nonzero entries in above the 1’s. The
following illustrates the latter approach.

     
0 1 2 1 1 0 1 0 − 13 0 0 1 0 0 1
9
1 1
 0 0 1  r1 →(r1 −2r2 )  0 0 1 31 1
 r1 →(r1 + 13 r3 )
0 0 1 1 1

A∗ =  3 2
1  −−−−−−−−→  2
1  −−−−−−−−→  3 2
1 
0 0 0 1 3 0 0 0 1 3 0 0 0 1 3
0 0 0 0 1 0 0 0 0 1 0 0 0 0 1
1 If a reduced echelon form is sought, at this point one would also eliminate the nonzero entries above the 1 in position 2k

by additional elementary row additions.


 1     
0 1 0 0 9 0 1 0 0 0 0 1 0 0 0
r2 →(r2 − 13 r3 )  0 7 7
0 1 0 18 r1 →(r1 − 19 r4 )  0 0 1 0  7
r2 →(r2 − 18
18  −−−−−−−−→ 
r4 )  0 0 1 0 0
−−−−−−−−→  1 −−−−−−−−→  1 1 
0 0 0 1 3 0 0 0 1 3 0 0 0 1 3
0 0 0 0 1 0 0 0 0 1 0 0 0 0 1

 
0 1 0 0 0
r3 →(r3 − 13 r4 )
0 0 1 0 0 ∗
−−−−−−−−→  =B
0 0 0 1 0
0 0 0 0 1
The entries in 1 reduced echelon form are not always just 1’s and 0’s even though this turned out to be the
case here. The Matlab command for computing a reduced echelon form of a given matrix A is rref(A).
Matlab uses essentially the method just described with the elimination of nonzero elements above 1’s being
carried out as part of each pivot as noted mentioned in footnote 1. A nice Matlab demo illustrating this can
be seen by typing the command rrefmovie.

Computation of Matrix Rank

Since an echelon form or reduced echelon form A∗ of A is obtained by applying a sequence of elementary
row operations to A, it follows that there is corresponding sequence of elementary matrices E1 , E2 , . . . , Ei ,
one for each elementary row operation, such that

A∗ = Ei Ei−1 · · · E2 E1 A

Since each Ej is nonsingular, the product T = Ei Ei−1 · · · E2 E1 must be nonsingular as well. Therefore

A∗ = T A

This proves that any matrix can be transformed into an echelon form or a reduced echelon form by multiplying
it on the left by a suitably defined nonsingular matrix.

As we’ve already noted, the rank of a matrix doesn’t change when multiplied on the left by a nonsingular
matrix. Therefore it must be true that rank A = rank A∗ . But we’ve already noted that the rank of an
echelon form or a reduced echelon form is equal to its number of nonzero rows. We are therefore lead to the
following conclusion. The rank of any matrix A is the same as the same as the number of nonzero rows in
any echelon form or reduced echelon form of A. Together with Gauss Elimination, this provides a numerical
method for obtaining the rank of a matrix. Note that this method has the desirable feature of not requiring
the computation of determinants.

Matrix Inversion

Suppose that An×n is a square, nonsingular matrix. Then A must have rank n. This means that any reduced
echelon form B ∗ of A must have n nonzero rows. This implies that B ∗ must be the n × n identity In×n ,
since reduced echelon forms are required to have 0s above and below each 1. This means that A can be
transformed into the n×n identity matrix by a applying to it, a sequence of elementary row operations. Thus
if E1 , E2 , . . . , Eq is the set of elementary matrices obtained by applying the same sequence of elementary
row operations to identity matrices, then

Eq Eq−1 · · · E2 E1 A = I
Clearly
A−1 = Eq Eq−1 · · · E2 E1
so
[I A−1 ] = Eq Eq−1 · · · E2 E1 [ A I ]
In other words, it is possible to compute the inverse of A by applying a sequence of elementary operations
to the partitioned matrix [ A I ] which transform it into its reduced echelon form. This form automatically
turns out to be [ I A−1 ]. More symbolically
opq , opq−1 , ..., op2 , op1
[ A I ] −−−−−−−−−−−−−−−−−−−−−−−−−→ [ I A−1 ]

where opi denotes the ith elementary row operation applied to [ A I ] in the process of transforming it
to reduced echelon form. Computing A−1 in this way is called the Gauss Jordan method. It is preferred
method for computing the inverse of an arbitrary nonsingular matrix. To illustrate, suppose
 
1 3
A=
2 4

Then the Gauss Jordan computation would proceed as follows.

     
1 3 1 0 r2 →(r2 −2r1 ) 1 3 1 0 r2 →− 12 r2 1 3 1 0
−−−−−−−−−−−−→ −−−−−−−−−−−−→
2 4 0 1 0 −2 −2 1 0 1 1 − 21

 
1
r1 →(r1 −3r2 ) 0 −2 32
−−−−−−−−−−−−→
0 1 1 − 21
Thus  
3
−2 2
A−1 =  
1 − 21
There are two Matlab commands for computing the inverse of a nonsingular matrix A, namely Aˆ(-1) and
inv(A).

Solutions to Linear Algebraic Equations

As we explained at the beginning of Section 2.7, the set of solutions x to a linear algebraic equation of the
form
An×m xm×1 = bn×1 (2.18)
is the same as the set of solutions to
T Ax = T b
provided T is a n × n nonsingular matrix. Suppose op1 , op2 , · · · , opq is a sequence of elementary row
operations transforming the partitioned matrix [ A b ] into a reduced echelon form M ∗ written in partitioned
form [ Ā b̄ ]. In other words
opq , opq−1 , ..., op2 , op1
[ A b ] −−−−−−−−−−−−−−−−−−−−−−−−−→ [ Ā b̄ ]

Then the set of solutions to (2.18) must be the same as the set of solutions to

Āx = b̄ (2.19)

because Ā = T A and b̄ = T b where T = Eq Eq−1 · · · E2 E1 and Ei is the elementary row matrix obtained by
applying operation opi to the n × n identity matrix. The key point here is that because [ Ā b̄ ] is a reduced
echelon form, the problem of finding solutions to (2.19) {and consequently (2.18)} is especially easy. The
following example serves to illustrate this point.

Suppose that A is 4 × 5, that b is 4 × 1 and consequently that x is a 5 - vector of unknowns; i.e.


 
x1
 x2 
 
x =  x3 
 
x4
x5
Suppose that C ∗ is a reduced echelon form of [ A b ] which looks like

 
0 1 c13 0 0 c16
∗ 0 0 0 1 0 c26 
C =  (2.20)
0 0 0 0 1 c36
0 0 0 0 0 0
Then the set of solutions to (2.18) is the same as the set of solutions to

 
  x1  
0 1 c13 0 0 c16
 x2 
0 0 0 1 0     c26 
 x  =  
0 0 0 0 1  3 c36
x4
0 0 0 0 0 0
x5
By carrying out the indicated multiplications, one obtains the simultaneous equations

x2 + c13 x3 = c16
x4 = c26
x5 = c36

From these equations it is clear that the set of all solutions to (2.18) consists of all vectors of the form
 
µ1
 c16 − c13 µ3 
 
x= µ3 
 
c26
c36
where µ1 and µ3 are real numbers. Because µ1 and µ3 are arbitrary elements of IR, the system of equations
under consideration does not have a unique solution; in fact there are infinitely many solutions, since there
are infinitely many possible values for µ1 and µ3 .

To illustrate the possibility of nonexistence, suppose that in place of


 
c16
 c26 
 
c36
0
the last column of the reduced echelon form of [ A b ] in (2.20) had turned out to be
 
0
0
 
0
1
Then the resulting matrix equation
 
  x1  
0 1 c13 0 0 0
 x2 
0 0 0 1 0  0
 x  =  
0 0 0 0 1  3 0
x4
0 0 0 0 0 1
x5

would not have had a solution at all because any such solution would have had to satisfy
 
x1
 x2 
 
[0 0 0 0 0 ]  x3  = 1
 
x4
x5

which is impossible.

To summarize, existence, uniqueness, and explicit solutions to Ax = b can be obtained directly by


inspection of a reduced echelon form of [ A b ]. Any such form can be obtained in a straight forward
manner via elementary row operations.

Solutions to Linear Matrix Equations

Now suppose that An×m and Bn×q are given matrices and consider the matrix equation

AX = B (2.21)

Let bi and xi denote the ith columns of B and X respectively. Thus each bi is an n - vector, each xi is an
m - vector, and
B = [ b1 b2 · · · bq ] and X = [ x1 x2 · · · xq ]
The point here is that (2.21) is equivalent to the q ‘decoupled’ equations

Axi = bi , i ∈ {1, 2, . . . , q} (2.22)

Because they are decoupled, each such equation can thus be solved independently of the others using the
method discussed in the last section. Defining these q vector solutions as the columns of a matrix X̄ thus
provides a solution X = X̄ to the original matrix equation (2.21). Note that the q independent equations
(2.22) can also be solved simultaneously by computing a reduced echelon form of [ A B ]. The Matlab
command for computing a matrix X which solves AX = B is X = A\B.

2.7.4 Elementary Column Operations

Just as there are three elementary row operations which can be performed on the rows of an arbitrary n × m
matrix A there are three analogous operations which can be performed on the columns of A. The operations
are

1. Interchange two columns of A.

2. Multiply a column of A by a nonzero number.

3. Add to a column of A, a different column of A multiplied by a number.


Corresponding to each elementary column operation, there is a nonsingular matrix E, which when multiplied
times A on the right, produces in a new matrix AE which is the same as that matrix which would have
resulted had the elementary column operation under consideration been performed directly on A. Moreover
E is that matrix which results when the elementary column operation under consideration is applied to the
m × m identity matrix Im×m .

We shall not illustrate all this. Suffice it to say that an elementary column operation on A {e.g.,
interchange columns i and j of A} amounts to performing the corresponding elementary row operation
on A′ {e.g., interchange rows i and j of A′ } and then transposing the resulting matrix.

2.7.5 Equivalent Matrices

Let An×m be a given matrix. As we’ve already seen, by application of a sequence of elementary row operations
it is possible to transform A into a reduced echelon form A∗ . Such an A∗ might look like

 
0 1 a13 0 0 a16
∗ 0 0 0 1 0 a26 
A = 
0 0 0 0 1 a36
0 0 0 0 0 0
We’ve also seen that there must be a nonsingular matrix T such that

A∗ = T A (2.23)

Now suppose that by means of elementary row operations, the transpose of A∗ e.g.
 
0 0 0 0
 1 0 0 0
 
a 0 0 0
(A∗ )′ =  13 
 0 1 0 0
 
0 0 1 0
a16 a26 a36 0

is transformed into a reduced echelon form B ∗ . For the example at hand, B ∗ would have to be
 
1 0 0 0
0 1 0 0
 
0 0 1 0
B∗ =   (2.24)
0 0 0 0
 
0 0 0 0
0 0 0 0

The very special structure of B ∗ is no accident. Rather it is a direct consequence of the fact that B ∗ is
a reduced echelon form of the transpose of a reduced echelon form. Note that (B ∗ )′ could also have been
obtained by applying elementary column operations to A∗ .

We know that there must be a nonsingular matrix R′ such that

B ∗ = R′ (A∗ )′

Combining this with (2.24) and (2.23) and we see that


 
I 03×3
T AR = 3×3
01×3 01×3
The matrix on the right is an example of a “left-right equivalence canonical form.”

For given nonnegative integers n, m and r with r ≤ min{n, m}, a left-right equivalence canonical form
Mn×m = [ mij ] is a n × m matrix whose entries are all zero except for the r entries m11 , m22 , . . . , mrr which
are all 1’s. Note that r is the rank of M . What we’ve just shown by example is that for any matrix A there
are nonsingular matrices T and R which transform A to a left-right canonical form T AR.

Equivalence Relations

There are aspects of the preceding which can be usefully interpreted in a broader context. This requires the
concept of an equivalence relation. An equivalence is a special type of relation.

Let S be a given set. Roughly speaking, a “relation” between objects in S is an association between
certain pairs of objects in S determined by a specific set of rules. For example, if S is the set of graduate
students taking this course this fall, a set of rules defining a relation called “sort of similar” might be as
follows: student sa is sort of similar to student sb if student sa is registered in the same department as
student sb and has been at Yale at least as long as student sb . Notice that the relation “sort of similar”
enables one to form ordered pairs of sort of similar students (sa , sb ). The set of all such ordered pairs is
sometimes considered to be the definition of the relation.

There is a special type of relation, called an ‘equivalence relation’ which is especially useful. We’ll first
define it formally, and then give some examples.

Let S be a set. Let S × S denote the set of all ordered pairs (s1 , s2 ) where s1 and s2 are both elements
of S. {S × S is sometimes called the Cartesian Product of S with itself.} Let E be a given subset of S × S.
Let’s use the notation s1 Es2 to mean (s1 , s2 ) is an element of E. The subset E is called an equivalence
relation on S if the following three requirements are satisfied.

1. Reflexivity: sEs, ∀s ∈ S. In other words each s in S must be equivalent to itself.


2. Symmetry: tEs, ∀s, t ∈ S such that sEt. In other words for each s and t in S such that s is equivalent
to t, t must be equivalent to s.
3. Transitivity: sEu ∀s, t, u ∈ S such that sEt and tEu. In other words for each s, t, and u in S such
that s is equivalent to t and t is equivalent to u, it must follow that s is equivalent to u.

Equality of real numbers is a simple example of an equivalence relation on IR. On the other hand, ‘less
than or equal’ {i.e., ≤} is not an equivalence relation on IR because requirement 2, namely symmetry, is not
satisfied. The relation ‘sort of similar’ defined above is also not an equivalence relation because it fails to be
symmetric.

An equivalence relation E on a set S provides a simple means of partitioning S into disjoint subsets of
like {i.e., equivalent} objects. These subsets are called ‘equivalence classes’ and are defined just as you’d

expect. The equivalence class of s ∈ S under E is the subset [s]E = {t : tEs, t ∈ S}. In other words, [s]E
is simply the set of all elements in S which are equivalent to s. The set of all possible equivalence classes of
S under E is written as S/E and is called the quotient set induced by E. Thus S/E = {[s]E : s ∈ S}. The
rules defining E insure that distinct equivalence classes don’t overlap:
[s]E ∩ [t]E = the empty set if sE
/t
The rules also insure that the union of all such equivalence classes is S:
[
S= [s]E
s∈S

Example 6 Suppose S = {1, 2, 3, 4, 5} and suppose that E is defined so that sEt whenever s − t is an
integer multiple of 3. Then E = {(1, 4), (4, 1), (2, 5), (5, 2), (1, 1), (2, 2), (3, 3), (4, 4), (5, 5)}. Since the
above requirements are satisfied, E is an equivalence relation on S. In addition [1]E = {1, 4}, [2]E = {2, 5},
[3]E = {3}, [4]E = [1]E , [5]E = [2]E and S/E = {[1]E , [2]E , [3]E }.

Left-Right Equivalence

Later in the course we will study two special equivalence relations on sets of matrices, namely “similarity”
and “congruence.” In the sequel we will consider a third which serves to illustrate the ideas we’ve just
discussed.

Let Mn×m denote the set of all n × m matrices over the field IK. Let us agree to call two matrices M1
and M2 in Mn×m left-right equivalent 2 if there exist nonsingular matrices P and Q such that

P M1 Q = M2

We sometimes write M1 7−→ QM1 P to emphasize that such a Q and P are transforming M1 into M2 .
We often call this a left-right equivalence transformation. The reader is encouraged to verify that left-right
equivalence is an equivalence relation on Mn×m .

Earlier in this section we demonstrated by example that for any matrix A in Mn×m there are nonsingular

matrices T and R which transform A into a left-right equivalence canonical form E ∗ = T AR. Therefore A
and E ∗ are left-right equivalent. We also pointed out that the structure of E ∗ is uniquely determined by three
integers n, m, and r, the latter being the rank of E ∗ . Since matrix rank is invariant under premultiplication
and postmultiplication by nonsingular matrices {i.e, under left-right equivalence transformations}, r must
also be the rank of A. In other words, E ∗ itself is uniquely determined by properties of A {namely A’s size
and rank}, which in turn are are invariant under left-right equivalence transformations.

In the light of the preceding, it is possible to determine what’s required of two matrices A and B in Mn×m
in order for them to be left-right equivalent. On the one hand, if they are in fact left-right equivalent, then
they must both have the same rank because rank is invariant under left-right equivalence transformations.
On the other hand, if A and B both have the same rank, they turn out to be left-right equivalent. Here’s a
proof:

• Suppose A and B have the same rank.


• Then A and B must both have the same left-right equivalence canonical form E ∗ .
• By definition, A and B are thus both left-right equivalent to E ∗ .
• By symmetry, E ∗ is left-right equivalent to B
• Since A is left-right equivalent to E ∗ and E ∗ is left-right equivalent to B it must be that A is left-right
equivalent to B because of transitivity.

Therefore two matrices of the same size are left-right equivalent if and only if they both have the same rank.

Let us note that the concept of left-right equivalence enables us to decompose Mn×m into 1 + min{n, m}
equivalence classes Ci , i ∈ {0, 1, . . . , min{n, m}} where M ∈ Ci just in case rank M = i. The rank of a matrix
M serves to uniquely label M ’s left-right equivalence class in the quotient set Mn×m /left-right equivalence.
Said differently, the rank of a matrix uniquely determines the matrix “up to a left-right equivalence trans-
formation.”
2 In many texts, ‘left-right’ equivalent matrices are called simply ‘equivalent’ matrices. Later on, when there is less chance

of confusion, we’ll adopt the shorter name.


Chapter 3

Linear Algebra

It is sometimes useful to think of a matrix as a “representation” of a “linear function” which “assigns” to


each vector in a “vector space,” a vector in another vector space. By doing this it is possible to single out
those properties of a matrix {e.g., rank} which don’t depend on the “coordinate system” in which the linear
transformation is represented. In the sequel we expand on these ideas by first introducing the concept of a
vector space and discussing some of its properties. Then we make precise the concept of a linear function
and explain how it relates to a matrix.

3.1 Linear Vector Spaces

Let IK denote an algebraic field {e.g., IR or C}


l and let Y be a set. Let ◦ and +̇ be symbols such that k ◦ y
and y1 +̇y2 denote well-defined elements of Y for each k ∈ IK and each y, y1 , y2 ∈ Y . The algebraic system
{Y, IK, +̇, ◦} is called a IK - vector space with vectors y ∈ Y and scalars k ∈ IK provided the following
axioms hold:

1. y1 +̇(y2 +̇y3 ) = (y1 +̇y2 )+̇y3 , ∀ y1 , y2 , y3 ∈ Y .

2. y1 +̇y2 = y2 +̇y1 ∀ y1 , y2 ∈ Y .

3. There exists a vector 0̄ ∈ Y {the additive unit} call the zero vector such that
0̄+̇y = y, ∀y ∈ Y .

4. For each y ∈ Y there is a vector −y ∈ Y {the additive inverse} such that y +̇(−y) = 0̄.

5. k ◦ (y1 +̇y2 ) = k ◦ y1 +̇k ◦ y2 , ∀k ∈ IK, y1 , y2 ∈ Y .

6. (k1 + k2 ) ◦ y = k1 ◦ y +̇k2 ◦ y, ∀ k1 , k2 ∈ IK, y ∈ Y .

7. (k1 k2 ) ◦ y = k1 ◦ (k2 ◦ y), ∀ k1 , k2 ∈ IK, y ∈ Y .

8. 1 ◦ y = y, ∀ y ∈ Y .

To summarize, a IK - vector space is a field of scalars IK and a set of vectors Y together with two ‘operations’,
namely scalar multiplication {◦} and vector addition {+̇} which satisfy the familar rules of vector algebra
just listed. To simplify notation we shall henceforth write ky for k ◦ y, y1 + y2 for y1 +̇y2 , and 0 for 0̄. These
notational simplifications are commonly used and should cause no confusion.

31
Throughout these notes capital calligraphic letters X , Y, Z . . . usually denote IK - vector spaces. Vectors
are typically denoted by lower case Roman letters x, y, z, . . .. Vector spaces containing just the zero vector
are denoted by 0. Note that there is no such thing as an empty vector space because every vector space
must contain at least one vector, namely the zero vector!

The set of all n × 1 real-valued matrices together with ordinary matrix addition and {real} scalar multi-
plication provides a concrete example of an IR - vector space; this space is denoted by IRn . Similarly, Cl n is
the Cl - vector space consisting of all n × 1 complex-valued matrices with addition and scalar multiplication
defined in the obvious way. While these two vector spaces arise quite often in physically-motivated problems,
they are by no means the only vector spaces of practical importance.

Example 7 The set of all infinite sequences {k1 , k2 , . . .} of numbers ki ∈ IK together with component-wise
addition defined by

{a1 , a2 , . . .} + {b1 , b2 , . . .} = {a1 + b1 , a2 + b2 , . . .}
and component-wise scalar multiplication defined by

k{k1 , k2 , . . .} = {kk1 , kk2 , . . .}

is a IK-vectos space. In this case the “vectors” are infinite sequences.


Example 8 The set of all real-valued functions f (t) defined on a closed interval of the real line [a, b] = {t :
a ≤ t ≤ b}, together with ordinary pointwise addition and multiplication by real numbers, is a real vector
space. In this case the “vectors” are functions defined on [a, b].

Example 9 The set of all polynomials in one variable s, with complex coefficients, and degree less than or
equal to n, together with ordinary polynomial addition and “scalar” multiplication by complex numbers is
a C-vector
l space. In this case the “vectors” are polynomials.

3.1.1 Subspaces

Recall that a subset V of a set X is itself a set whose elements are also elements of X. Let X be a given
vector space over IK. A nonempty subset V whose elements are vectors in X , is a subspace of X if V is a
vector space under the operations of X . In other words for V to be a subspace of X , each element of V must
be a vector in X and V must be closed under the addition and scalar multiplication operations of X :

k1 v1 + k2 v2 ∈ V, ∀ ki ∈ IK, v2 ∈ V

Let us note that the zero subspace and the whole space X are subspaces of X . Any other subspace of X
{i.e., any subspace other than the whole space X or the zero subspace 0} is said to be a proper subspace of
X . To indicate that V is a subspace of X one often uses the notation V ⊂ X and says that V is contained in
X.

Example 10 The set of all vectors of the form


 
r1
 r2  , r1 , r2 ∈ IR
0

is a subspace of IR3 .
Example 11 The set of all vectors of the form

 
r1
 r2  , r1 , r2 ∈ IR
2
is not a subspace of IR3 because this set is not closed under addition; i.e., although
   
1 1
1 and 0
2 2

are vectors in the set, their sum  


2
1
4
is not since 4 6= 2.

3.1.2 Subspace Operations

Let U and V be subspaces of a given vector space X . The sum of U and V, written U + V, is defined by

U + V = {u + v : u ∈ U, v ∈ V}

In other words, the sum of U and V is the set of all possible vectors of the form u + v where u and v are
vectors in U and V respectively. For example, if U is the subspace of all vectors of the form
 
u1
 0 
  , u1 , u3 ∈ IR
u3
0
and V is the subspace of all vectors of the form
 
v1
0
 , v1 , v4 ∈ IR
0
v4
then U + V is the set of all vectors of the form
 
r1
0
 , r1 , r3 , r4 ∈ IR
r3
r4

Note that “sum” and “union” are not the same thing. The union of the elements of U and V is the subset

U ∪ V = {w : w ∈ U or w ∈ V}

Clearly
(U ∪ V) ⊂ (U + V)
but the reverse inclusion may not hold.
Both U + V and U ∪ V are subsets of X . However only U + V and not U ∪ V is a subspace of X . To prove
that U + V is a subspace it is enough to show that U + V is closed under addition and scalar multiplication.
To prove closure under addition, suppose that x1 and x2 are vectors in U + V. Then, because of the definition
of U + V there must be vectors u1 , u2 ∈ U and v1 , v2 ∈ V such that x1 = u1 + v1 and x2 = u2 + v2 . Clearly
x1 + x2 = (u1 + u2 ) + (v1 + v2 ). Since (u1 + u2 ) ∈ U and (v1 + v2 ) ∈ V, it must be that (x1 + x2 ) ∈ (U + V).
Thus U + V is closed under addition. To establish closure under scalar multiplication, let us note that for
any k ∈ IK, kx1 = ku1 + kv1 ; but ku1 ∈ U and kv1 ∈ V so kx1 ∈ (U + V). Hence U + V is closed under scalar
multiplication. Therefore U + V is a subspace.

Let X , U and V be as above. The intersection of U and V, written U ∩ V, is the set of all vectors which
are in both U and V. In other words

U ∩ V = {x : x ∈ U and x ∈ V}

It turns out that U ∩ V, like U + V, is also a subspace of X . The reader should try to prove that this is so.

To illustrate intersection, suppose that U and V are subspaces of IR4 consisting of all vectors of the forms
   
u1 v1
 0  0
  and  
u3 0
0 v4
respectively. Then U ∩ V is all vectors of the form
 
r1
0
 
0
0

The definitions of sum and intersection of subspaces extend to finite families of subspaces in a natural
∆ ∆
way. For example, if S1 , S2 and S3 are subspaces of X then S1 + S2 + S3 = (S1 + S2 ) + S3 and S1 ∩ S2 ∩ S3 =
(S1 ∩ S2 ) ∩ S3 .

3.1.3 Distributative Rule

Suppose that U, V and W are subspaces of a vector space X . We claim that

(W ∩ U) + (W ∩ V) ⊂ W ∩ (U + V) (3.1)

To prove that this is so, let x be any vector in (W ∩ U) + (W ∩ V). Then x ∈ W and x = u + v for some u ∈ U
and some v ∈ V. Since u + v ∈ (U + V) it must be true that x ∈ (U + V). Since x is also in W, it follows
that x ∈ W ∩ (U + V). As x was selected arbitrarily in (W ∩ U) + (W ∩ V), the containment x ∈ W ∩ (U + V)
must hold for all x ∈ (W ∩ U) + (W ∩ V). In other words, (3.1) is true.

Example 12 Suppose U, V and W are subspaces of IR2 consisting of all vectors of the forms
     
r1 0 r3
, , and
0 r2 r3
respectively. Then
U ∩ W = 0, V ∩ W = 0, and U + V = IR2
Thus
(U ∩ W) + (V ∩ W) = 0 and W ∩ (U + V) = W
Clearly
W ∩ (U + V) 6= (U ∩ W) + (V ∩ W)
Geometrically, things look like this.

V W

U
@
I 1
@
@
@
@
@
@
This is the only point
common to U and W,
common to V and W,
and thus the only point
in (U ∩ W) + (V ∩ W).

Here all points in the plane of the paper represent points in IR2 and the point of intersection of the three
lines is the zero vector. V consists of all points on the extension of the vertical line shown, from −∞ to ∞.
W and U admit similar interpretations.

It can be shown by example that the reverse inclusion in (3.1) does not necessarily hold so the containment
symbol in (3.1) cannot in general be replaced with an equal sign. An exception to this occurs if it happens
to be true that
U ⊂W (3.2)
Our aim is to prove that this is so; i.e., that (3.2) implies that

(W ∩ U) + (W ∩ V) = W ∩ (U + V) (3.3)

The implication
U ⊂ W ⇒ (W ∩ U) + (W ∩ V) = W ∩ (U + V)
is called the modular distributative rule for subspaces. In view of (3.1), to prove the rule’s validity all we
need to show is that (3.2) implies

(W ∩ U) + (W ∩ V) ⊃ W ∩ (U + V) (3.4)

This can be done as follows.

Let x be any vector in W ∩ (U + V). Then x ∈ W and x = u + v for some vectors u ∈ U and v ∈ V.
Since u ∈ U, we have by (3.2) that u ∈ W so u ∈ U ∩ W. Since v = x − u and both x and u are in
W, it must be true that v ∈ W; but v ∈ V so v ∈ V ∩ W. Therefore u + v ∈ (W ∩ U) + (W ∩ V) which
implies that x ∈ (W ∩ U) + (W ∩ V). Since x can be any vector in W ∩ (U + V), all such vectors must be in
(W ∩ U) + (W ∩ V). In other words, (3.4) must hold.
3.1.4 Independent Subspaces and Direct Sum

Two subspaces {U and V} of a vector space X are independent if their intersection is the zero subspace in
X ; i.e., if U ∩ V = 0. Note that every subspace of X contains the zero vector, so the intersection of subspaces
is never the empty set. The subspaces U and W in Example 12 are independent as are the subspaces V and
W.

A finite family of subspaces {Si : i ∈ {1, 2, . . . , n}} is independent if


 
\ X n

Si  Sj  = 0, i ∈ {1, 2, . . . , n}
j=1
j6=i


If {Si : i ∈ {1, 2, . . . , n}} is an independent family and S is the sum S = S1 + S2 + · · · + Sn , then S is called
the direct sum of the Si . The symbol ⊕ is usually used to denote direct sum. Thus if {Si : i ∈ {1, 2, . . . , n}}
is an independent family, then
S = S1 ⊕ S2 ⊕ · · · ⊕ Sn
One consequence of S being a direct sum is that each vector s ∈ S can be written as s = s1 + s2 + · · · + sn
where each vector si ∈ Si is unique.

3.1.5 Linear Combination and Span

Let {xi : i ∈ {1, 2, . . . , n}} be a set of vectors in a vector space X . A vector x ∈ X is said to be a linear
combination of the vectors x1 , x2 , . . . , xn if there exist scalars ki ∈ IK such that x = k1 x1 +k2 x2 +· · ·+kn xn .
By the span of the set of vectors {xi : i ∈ {1, 2, . . . , n}} is meant the set of all possible linear combinations
of x1 , x2 , . . . , xn . It is easy to verify that the span of x1 , x2 , . . . , xn is a subspace of X . We say that
x1 , x2 , . . . , xn spans a given subspace U ⊂ X just in case U is the span of x1 , x2 , . . . , xn . X is said to
be a finite dimensional vector space if it can be spanned by a finite set of vectors. If no such set exists,
X is called an infinite dimensional space. It will soon be clear, if it is not already, that subspaces of finite
dimensional spaces are finite dimensional spaces. Unless otherwise stated, we shall deal only with finite
dimensional subspaces in these notes.

Example 13 Since IRn is spanned by the unit n - vectors1 e1 , e2 , . . . , en , IRn is a finite dimensional vector
space. The same vectors e1 , e2 , . . . , en also span Cl n so Cl n is also finite dimensional. The vector space
of polynomials in one variable s, with coefficients in Cl and degrees not exceeding n {see Example 9} is
finite dimensional. The vector space of all infinite sequences {k1 , k2 , . . .} defined in Example 8 is infinite
dimensional as is the vector space of all real-valued functions on [a, b] defined in Example 7.

3.1.6 Linear Independence

A set of vectors {xi : i ∈ {1, 2, . . . , n}} in X is linearly dependent if there exist scalars ki ∈ IK, not all zero,
such that k1 x1 + k2 x2 + · · · + kn xn = 0. If the set is not linearly dependent it is linearly independent. For
example, the set of vectors      
1 3 −1
x1 =  1  , x2 =  −2  , x3 =  4 
2 0 4
is linearly dependent since 2x1 − x2 − x3 = 0. On the other hand, the vectors x1 and x2 are linearly
independent since the equation k1 x1 + k2 x2 = 0 implies that k1 and k2 must both equal zero.
1e is a n × 1 matrix with a 1 in row i and 0’s elsewhere.
i
3.1.7 Basis

A linearly independent set {xi : i ∈ {1, 2, . . . , n}}, ordered by i, is a basis for X if the set spans X . The set
of unit vectors {e1 , e2 , . . . , en } is a basis for both IRn and Cl n . The ordered set consisting of the vectors x1
and x2 just defined, is a basis for their span which, in turn, is a subspace of IR3 .

Suppose that {z1 , z2 , . . . , zm } is a finite set of vectors which spans X . If this set is independent, it is
a basis for X . If not, it is possible to reduce {z1 , z2 , . . . , zm } to a smaller set {x1 , x2 , . . . , xn } which is a
basis. Here’s how:

Suppose that {z1 , z2 , . . . , zm } is a finite set of linearly dependent vectors which spans X . Then there
must exist scalars ki , not all zero, such that

k1 z1 + k2 z2 + · · · km zm = 0

Without loss of generality, assume km 6= 0. Then zm can be written as a linear combination of the remaining
vectors z1 , z2 , . . . , zm−1 . That is
−k1 −k2 −km−1
zm = z1 + z2 + · · · + zm−1
km km km
From this it should be clear that the reduced set of vectors {z1 , z2 , . . . , zm−1 } also spans X . If this set is
linearly independent, then it is a basis. If not, the process can be continued for at most m − 1 additional
steps until a linearly independent subset {xi : i ∈ {1, 2, . . . , n}} ⊂ {z1 , z2 , . . . , zm } is finally obtained. The
result is a basis for X . We have therefore shown that every finite dimensional vector space has a basis. Note:
the zero subspace has no basis, but it doesn’t need one since it contains exactly one vector, namely 0.

Representations

Let {yi : i ∈ {1, 2, . . . , n}} be a basis for a vector space Y and let y be any given vector in Y. Because
{yi : i ∈ {1, 2, . . . , n}} spans Y there must be scalars ki ∈ IK such that

y = k1 y1 + k2 y2 + · · · + kn yn (3.5)

Moreover the ki must be unique because {yi : i ∈ {1, 2, . . . , n}} is a linearly independent set. In other words,
for each basis {yi : i ∈ {1, 2, . . . , n}} and each vector y there is a unique, ordered set of numbers {ki : i ∈
{1, 2, . . . , n}} for which (3.5) holds. The ki are called the coordinates of y in the basis {yi : i ∈ {1, 2, . . . , n}}.
The column vector  
k1
 k2 
 . 
 .. 
kn
is called a representation of y in the basis {yi : i ∈ {1, 2, . . . , n}}. For example, the vector space of polynomials
of degree not exceeding n defined in Example 9, has as a basis the set of polynomials {1, s, s2 , . . . , sn }. In
this basis, any polynomial cp s(p−1) + cp−1 s(p−1) + · · · c2 s + c1 in the space can be represented by a (n + 1) × 1
matrix of the form  
c1
 c2 
 . 
 . 
 . 
 
 cp 
 
0
 . 
 .. 
0
Representations provide an important link between abstract vectors {which might be polynomials, se-
quences, etc.} and matrices which makes it possible to determine various properties of the original vectors
by studying analogous properties of their matrix representations. In the sequel we shall illustrate this by
developing a procedure for determining when a given set of vectors is linearly independent by evaluating the
rank of a matrix whose columns are representations of the vectors under consideration.

Linear Independence and Rank

Let X be any finite dimensional vector space and let {xj : j ∈ {1, 2, . . . , n}} be a basis. Let {yi : i ∈
{1, 2, . . . , m}} be any set of vectors in X . Let
 
a1i
∆  a2i 
ai =  
 ... 
ani

be the representation of yi in the basis {xj : j ∈ {1, 2, . . . , n}}. That is


n
X
yi = aji xj , i ∈ {1, 2, . . . , m} (3.6)
j=1

Next let {bi : i ∈ {1, 2, . . . , m}} be any set of numbers such that
m
X
bi yi = 0 (3.7)
i=1

Substituting (3.6) into (3.7) we obtain the equation


m
X n
X
bi aji xj = 0 (3.8)
i=1 j=1

Since the xj are independent, their coefficients in (3.8) must each be zero. In other words,
m
X
aji bi = 0, j ∈ {1, 2, . . . , n} (3.9)
i=1

∆ ∆
If we define A = [ aji ]n×m and b = [ bi ]m×1 then (3.9) can be written as

Ab = 0 (3.10)

What we’ve just proved is that (3.7) implies (3.10). The reader should verify that the reverse implication
is also true. In other words (3.7) and (3.10) are equivalent assertions. From this and the definition of linear
independence it follows that {yi : i ∈ {1, 2, . . . , m}} will be a linearly independent set just in case the only
vector b for which (3.10) holds is the zero vector. We now develop a condition for testing when the latter is
so.

Let P and Q be nonsingular matrices which transform A into its left-right equivalence canonical form B ∗ ;
i.e. B ∗ = P AQ. Note that if b is a solution to (3.10), then Q−1 b is a solution d to B ∗ d = 0. Conversely, if d
is a solution to B ∗ d = 0, then Q−1 d is a solution to (3.10). From this and the fact that Q is nonsingular it
follows that (3.10) will have b = 0 as its only solution just in case B ∗ d = 0 has d = 0 as its only solution. But
because of B ∗ ’s special structure, the equation B ∗ d = 0 will have d = 0 as its only solution if and only if the
number of 1’s in B ∗ {i.e., the rank of B ∗ } equals m, the number of columns of B ∗ . But rank B ∗ = rank A.
From this it follows that (3.10) will have b = 0 as its only solution just in case rank A = m. Thus
{yi : i ∈ {1, 2, . . . , m}} will be a linearly independent set if and only if rank A = m. We’ve proved the
following.

Proposition 4 Let X be a finite dimensional space with basis {xj : j ∈ {1, 2, . . . , n}}, and let {yi : i ∈
{1, 2, . . . , m}} be any set of vectors in X . Then {yi : i ∈ {1, 2, . . . , m}} is a linearly independent set if and
only if
rank [ a1 a2 · · · am ]n×m = m
where for i ∈ {1, 2, . . . , m}, ai is the representation of yi in the basis {xj : j ∈ {1, 2, . . . , n}}.

Note that (3.10) can also be written as

a1 b 1 + a2 b 2 + · · · am b m = 0

Clearly the columns of A will be linearly independent just in case b = 0; i.e., just in case rank A = m. In
the sequel, this line of reasoning is carried further.

Suppose that Gn×m is an arbitrary matrix with columns g1 , g2 , . . . , gm . Let G denote the span of
{gi : i ∈ {1, 2, . . . , m}}. As before, we can remove dependent columns from {gi : i ∈ {1, 2, . . . , m}} without
changing span, until we end up with a linearly independent subset which is a basis for G. Without loss of
generality, assume that this subset consists of the first q columns of G. Therefore gq+1 can be written as

gq+1 = k1 g1 + k2 g2 + · · · + kq gq

It follows that if −k1 times g1 plus −k2 times g2 plus . . . plus −kq times gq is added to column q + 1 of G,
the result will be a zero column. By continuing this process, all columns of G, from q + 1 to m, can be made
equal to zero. The end result is a matrix Ḡn×m of the form
 
Ḡ = Gbn×q 0n×(m−q)

where
b = [ g1
G g2 · · · gq ]
Moreover since the steps involved in going from G to Ḡ are elementary column operations, there must be a
nonsingular matrix T such that
Ḡ = GT (3.11)

Now since the columns of G b are independent, the number of columns of G,b namely q, must equal the rank
b b
of G. Thus G and therefore Ḡ must contain a nonzero minor of order q. It folows that rank Ḡ ≥ q. But
any square subarray of Ḡ larger than q × q must contain a zero row or column and hence must be singular.
This means that q is the largest integer for which there exists a nonzero minor of order q in Ḡ. Therefore
rank Ḡ = q. But from (3.11) and the fact that T is nonsingular, it follows that rank G = q. We’ve proved
the following very useful fact.

Proposition 5 The rank of a matrix is equal to the number of linearly independent columns of the matrix.

Since the rank of a matrix equals the rank of its transpose, we also have the following:

Proposition 6 The rank of a matrix is equal to the number of linearly independent rows of the matrix.
Note that the two propositions imply that the number of linearly independent columns of a matrix equals
the number of linearly independent rows.

It is sometimes possible to tell by inspection just how many linearly independent columns or rows a
matrix has. For example, the matrix
 
2 1 3
1 1 2
 
3 1 4
0 1 1
has two independent columns since the first two columns are clearly independent and the third is the sum
of the first and second. Therefore the rank of the matrix is 2.

3.1.8 Dimension

The dimension of a vector space X , written dim(X ), is the least number of vectors required to span X . The
dimension of any zero space is defined to be zero.

Suppose {xj : j ∈ {1, 2, . . . , n}} and {yi : i ∈ {1, 2, . . . , m}} are two bases for X . We can express the
vectors in one basis as linear combinations of vectors in the other basis. For example
n
X
yi = aji xj i ∈ {1, 2, . . . , m}
j=1

Since {yi : i ∈ {1, 2, . . . , m}} is an independent set, it follows from Proposition 4 that rank [ aji ]n×m = m.
Since the rank of a matrix cannot be larger than the number of its rows, it must be true that m ≤ n. Now
the reverse inequality, namely n ≤ m, can also be established by interchanging the roles of the yi and xj
and using similar reasoning. This means that n = m. We’ve thus proved the following.

Proposition 7 All bases for a vector space X contain the same number of vectors.

Observe that the dimension of X cannot exceed the number of vectors in a basis for X because bases
are spanning sets. Conversely, the number of vectors in a basis cannot exceed the dimension of X ; for if
this were false, there would exist a spanning set for X containing less vectors than are in a basis. Since the
spanning set could be reduced to a basis by eliminating dependent vectors, this would imply that X has two
bases containing different numbers of vectors – a clear contradiction of Proposition 7. We have thus proved
the following.

Proposition 8 Each basis for a vector space X contains exactly dim(X ) vectors.

Note that any set of dim(X ) vectors which spans X must be a basis for X .

3.2 Functions

Having completed our discussion of the basic properties of finite dimensional linear vector spaces, we turn to
the concept of a linear transformation. What is a linear transformation? A linear transformation is a linear
function. What is a function? To pin down with reasonable precision just what a function is we need first
two sets R and S. Then a function from domain R to codomain S, written f : R → S, is a rule which assigns
to each element r in its domain R a corresponding element s in its codomain S; s is called the value of f
at r and is often denoted by f (r). We sometimes write r 7−→ f (r) to indicate the action of f on a typical
element r in its domain. As we shall soon see, it is essential in talking about a function that we specify both
its domain and its codomain.

Example 14 The rule which assigns to each real number its cube, determines a function from the real
numbers to the real numbers.

Example 15 The best selling fiction book list published weekly by the New York Times can be viewed as
a function from the set of integers {1, 2, . . . , 10} to the set of all fiction books.

Example 16 An algorithm which computes the length of a shortest path between two vertices on a weighted,
directed graph can be thought of as a function from the set of all such graphs to the reals.

Example 17 A balance scale determines a function which assigns to each object of mass, a real number,
namely the mass of the object.

It is important to recognize that two functions are equal just in case they both have the same domain, the
same codomain and the same value at each point in their common domain. Note that the function which
assigns to each positive real number, its cube, is not quite the same as the function discussed in Example 14,
since the two functions have different domains.

The composition of two functions g : S → T and f : R → S, written gf , is the function gf : R → T


defined by r 7−→ g(f (r)). Thus the composition of g with f is a “function of a function.” Note that the
definition of gf makes sense only if domain g = codomain f .

3.2.1 Linear Transformations

In order to define a linear transformation we first need two linear vector spaces, X and Y, defined over the
same field, IK. A function f : X → Y is then said to be a linear transformation if the superposition rule

f (k1 x1 + k2 x2 ) = k1 f (x1 ) + k2 f (x2 )

holds for all x1 , x2 ∈ X and all k1 , k2 ∈ IK. The reason for requiring X and Y to be linear vector
spaces is clear: in order to define f as a linear transformation, we need well-defined notions of addition and
scalar multiplications in both f ’s domain and codomain. Throughout these notes, linear transformations
are usually denoted by capital Roman letters A, B, C, . . .. When we wish to distinguish sharply between
a linear transformation A and a matrix A, we shall usually express the latter in boldface; e.g., A. We often
write Ax rather than A(x) to denote the value of A at x.

Example 18 Let P n+1 denote the (n+1) - dimensional linear vector space of all real-coefficient polynomials,
of the form α(t) = an+1 tn + an tn−1 + · · · + a2 t + a1 together standard polynomial addition and multiplication
d
by scalars in IR. The function which assigns to each α(t) ∈ P n+1 , its derivative dt α(t), determines a linear
n+1 n
transformation from P to P .

Example 19 Let Y denote the vector space of all directed line segments in the real plane, drawn from a
fixed origin. The function which assigns to each line segment, the same line segment rotated 60 degrees
clockwise, is a linear transformation from Y to Y.
line segment

o
60
rotated line segment

Example 20 Let Z denote the set of all directed line segments in the real plane, drawn from all points in
the plane. The function which assigns to each line segment in the vector space Y defined in Example 19, the
same line segment translated two units to the right, is a function from Y to Z. However this function is not
quite a linear transformation because superposition does not hold.

translation -
 
 
 
 
 
 
 
 
 
 2 -

Example 21 Let An×m be a real matrix. The function which assigns to each vector x in IRm , the vector
Ax in IRn is a linear transformation from IRm to IRn .

3.2.2 Operations With Linear Transformations

There are basically two operations one can perform with linear transformations, namely composition and
addition. The composition of the linear transformations M : X → Y with the linear transformation N :
Y → W is the function N M : X → W defined so that N M (x) = N (M x) ∀x ∈ X . Thus the composition of
functions is defined in the same manner, whether the functions are linear or not. It turns out however that
the composition of two linear functions is a linear function. The reader should try to verify this.

The sum of two linear transformations M : X → Y and L : X → Y, written M + L, is the function


(M + L) : X → Y defined so that (M + L)(x) = M (x) + L(x), ∀x ∈ X . The reader may with to verify that
the sum of two linear transformations is a linear transformation.

3.2.3 Representations of Linear Transformations

As we’ve already noted, is is sometimes useful to represent abstract vectors using column matrices. Doing
this makes it possible to study properties of the original abstract vectors {e.g., linear independence} in terms
of their representations. The same is true of linear transformations. The idea is that a linear transformation
can be “represented” or “characterized” by a matrix, once bases for the domain and codomain of the
transformation have been settled upon. Matrix representations of linear transformations are defined so that
the linear transformation operations of addition and composition translate into the matrix operations of
addition and multiplication respectively. We shall expand on these ideas in the sequel.

Let L : X → Y be a given linear transformation and let {xj : j ∈ {1, 2, . . . , n}} and {yi : i ∈ {1, 2, . . . , m}}
be bases for X and Y respectively. Since L(xj ) is a vector in Y and since {yi : i ∈ {1, 2, . . . , m}} is a basis
for Y, there are unique numbers kij such that

m
X
L(xj ) = kij yi , j ∈ {1, 2, . . . , n} (3.12)
i=1
The array
 
k11 k12 ··· k1n
∆  k21 k22 ··· ··· 
L = [ kij ]m×n =
 ... .. .. .. 
. . . 
km1 ··· · · · kmn
is called the matrix representation of L in the chosen bases. Note that the size of L, namely m × n, is
determined by dim(Y) and dim(X ).

Example 22 Let P 3 denote the vector space of all polynomials α(t) = a3 t2 + a2 t + a1 as in Example
d
18. Let P 3 → P 2 denote the linear transformation which assigns to α(t), its derivative dt α(t). That is
∆ d
L(α(t)) = dt α(t). Let {1, t, t2 } and {1, 1 + t} be bases for P 3 and P 2 respectively. Then

L(1) = 0 = 0(1) + 0(1 + t)


L(t) = 1 = 1(1) + 0(1 + t)
L(t2 ) = 2t = −2(1) + 2(1 + t)
Hence the matrix of L in the chosen bases is
 
0 1 −2
L=
0 0 2

Let X , Y, and Z be vector spaces with bases {xi : i ∈ {1, 2, . . . , n}}, {yi : i ∈ {1, 2, . . . , m}} and {zi : i ∈
{1, 2, . . . , q}} respectively. Suppose that L, M and N are matrix representations of linear transformations L :
X → Y, M : X → Y and N : Y → Z respectively, in these bases. It can be shown in a straightforward manner
that in these same bases, the matrix representations of the sum L+M and the composed transformation N M
are L + M and NM respectively. Because of this there is a close correspondence between matrix addition
and multiplication one the one hand and linear transformation addition and composition on the other.

3.2.4 Coordinate Transformations

Suppose as above, that L = [ kij ] is the matrix representation of L : X → Y in the bases {xj : j ∈
{1, 2, . . . , n}} and {yi : i ∈ {1, 2, . . . , m}}. Suppose that {x̄ : i ∈ {1, 2, . . . , n}} and {ȳi : i ∈ {1, 2, . . . , m}}
is another pair of bases for X and Y respectively, and that L̄ = [ k̄ij ] is the representation of L in these
bases. Our aim is to explain how these two representations of L are related. Since {ȳi : i ∈ {1, 2, . . . , m}}
and {xj : j ∈ {1, 2, . . . , n}} are bases for Y and X respectively, there are numbers pir and qsj such that
m
X
yr = pir ȳi , r ∈ {1, 2, . . . , m} (3.13)
i=1
and n
X
x̄j = qsj xs , j ∈ {1, 2, . . . , n} (3.14)
s=1

Therefore for j ∈ {1, 2, . . . , n}


n
!
X
L(x̄j ) = L qsj xs {by (3.14)}
s=1
n
X
= qsj L(xs ) {by linearity of L}
s=1
n m
!
X X
= qsj krs yr { by (3.12)}
s=1 r=1
n m m
!!
X X X
= qsj krs pir ȳi {by (3.13)}
s=1 r=1 i=1

Thus for i ∈ {1, 2, . . . , m} and j ∈ {1, 2, . . . , n}


n m
!
X X
k̄ij = qsj krs pir
s=1 r=1

Using this formula for the ijth entry of L̄, it is straightforward {but tedious} to verify that

L̄ = PLQ (3.15)
∆ ∆
where P = [ pir ]m×m and Q = [ qsj ]n×n . Moreover P and Q are nonsingular matrices because of Proposition
4. The reader should try to verify these claims.

What (3.15) shows is that two different representations of L, resulting from two different choices of bases
for X and Y, are related by a left-right equivalence transformation. Matrices P and Q correspond to changes
of bases for Y and X respectively. From this discussion we can conclude that the matrix representation of a
linear transformation is uniquely determined only up to a left-right equivalence transformation.

3.2.5 Defining a Linear Transformation on a Basis

One of the useful consequences of linearity is that it is possible to define a linear transformation by merely
specifying its action on the elements of a basis for its domain. To illustrate this, suppose X is a IK - vector
space with basis {xi : i ∈ {1, 2, . . . , n}}. Let Y be any IK - vector space and let {yi : i ∈ {1, 2, . . . , n}} be any
given set of n vectors in Y. The specification xi 7−→ yi , i ∈ {1, 2, . . . , n}, then completely defines a linear
transformation L : X → Y. For if x is any vector in X { i.e.,
n
X
x= ki xi
i=1

where the ki are the coordinates of x in the basis {xi : i ∈ {1, 2, . . . , n}}} then
n
X n
X
L(x) = ki L(xi ) = ki yi
i=1 i=1


because of linearity and the definitions L(xi ) = yi , i ∈ {1, 2, . . . , n}. In other words, because of linearity,
defining L on a basis for X automatically defines L on all of X .
3.2.6 Linear Equations: Existence, Image, Epimorphism

Let L : X → Y and ȳ ∈ Y be given. Our aim is to develop conditions in terms of L and ȳ for the existence
of a vector x ∈ X which solves the linear equation

L(x) = ȳ (3.16)

As a first step, we introduce a special set of vectors called the image of L:



image L = {L(x) : x ∈ X }

In other words, the image of L consists of all vectors y such that y = L(x) for some x ∈ X .

Let us note that image L is a subspace of Y. For if y1 , y2 ∈ image L, then there must be vectors
x1 , x2 ∈ X such that yi = L(xi ), i ∈ {1, 2, . . . , 2}. From this it follows that for any k1 , k2 ∈ IK,

k1 y1 + k2 y2 = k1 L(x1 ) + k2 L(x2 ) = L(k1 x1 + k2 x2 )

so k1 y1 + k2 y2 ∈ image L. Thus image L is closed under addition and scalar multiplication which proves
that image L is a subspace of Y.

Returning to the problem under consideration, let us note that (3.16) will have a solution x = x̄ if and
only if
ȳ ∈ image L (3.17)
The claim is almost self-evident: it follows at once from the definition of image L. For theoretical de-
velopment, (3.17) proves to be a very useful condition for existence of solutions to (3.16). For numerical
computation, something a little more concrete is needed. We’ll discuss this further a little later.

What we want to do next is to develop conditions for the existence of a solution X : Z → X to the more
general linear equation

LX = M (3.18)
where L : X → Y and M : Z → Y are given linear transformations. First suppose that (3.18) holds. Then
for each z ∈ Z
L(x) = M (z) (3.19)
where x = X(z). Thus for each vector z ∈ Z there is a vector x ∈ X such that (3.19) holds. This can be
true only if
image M ⊂ image L (3.20)
so (3.20) is a necessary condition for (3.18) to have a solution X.

Now suppose that (3.20) holds and let {zi : i ∈ {1, 2, . . . , q}} be a basis for Z. Then M (zi ) ∈ image L, i ∈
{1, 2, . . . , q}. Therefore because of the definition of image, there are vectors x1 , x2 , . . . , xq in X such that

M (zi ) = L(xi ), i ∈ {1, 2, . . . , q}

Since {zi : i ∈ {1, 2, . . . , q}} is a basis for Z, we can define X by simply specifying its action on this basis.
In particular, define X so that X(zi ) = xi , i ∈ {1, 2, . . . , q}. Thus

M (zi ) = LX(zi ), i ∈ {1, 2, . . . , q}

Since {zi : i ∈ {1, 2, . . . , q}} is a basis for Z it must therefore be true that (3.18) holds. We have proved the
following.
Proposition 9 Let L : X → Y and M : Z → Y be given linear transformations. There exists a linear
transformation X : Z → X such that
LX = M (3.21)
if and only if
image M ⊂ image L (3.22)

Note that (3.22) will necessarily hold {and (3.21) will consequently always have a solution} for any
M : Z → Y, provided image L = Y. Linear transformations with this property are important enough to be
given a special name. L : X → Y is said to be an epimorphism2 if image L = Y. Thus a sufficient condition
for (3.21) to have a solution is that L be an epimorphism. The reader may wish to verify that L will be an
epimorphism just in case any matrix representation of L has all rows linearly independent.

Condition (3.22) is quite convenient for theoretical development. For computational purposes, one would
apply the matrix version of this condition to matrix representations of L and M .

As a first step toward the development of a matrix condition for existence, let us agree to call the subspace
spanned by the columns of a given matrix A, the image3 of A. Now in terms of matrix representations (3.18)
is equivalent to
LX = M (3.23)
It is straightforward, but a little tedious, to show that (3.22) is equivalent to

image M ⊂ image L (3.24)

This containment can also be interpreted without regard to (3.18), as a necessary and sufficient condition
for the matrix equation (3.23) to have a solution X.

To develop the matrix version of (3.24), we first write (3.24) in the equivalent form

image M + image L = image L (3.25)


Using the identity
image M + image L = image [ M L] (3.26)
(3.25) can be rewritten as

image [ M L ] = image L (3.27)


{The reader should verify (3.26) and the equivalence of (3.25) and (3.24).}

We claim that (3.27) is equivalent to

dim (image [ M L ]) = dim (image L) (3.28)

That (3.27) implies (3.28) is clear. The reverse implication is a consequence of the general subspace impli-
cation dim (U) = dim (V) ⇒ U = V which is valid in the special case when U ⊂ V.

To proceed we need to make use of the following fact.

Proposition 10 If L : X → Y is a linear transformation, and L is any matrix representation of L, then

dim (image L) = dim (image L) = rank L


2 Other names sometimes used are surjective linear transformation or onto linear transformation.
3 Sometimes this subspace is called the column span of A.
A proof of this proposition can be derived using Propositions 4 and 5.

In view of the equivalence of (3.28) and (3.24) we can state the following.

Proposition 11 Let Mn×m and Ln×p be given matrices over IK. A necessary and sufficient condition for
the equation
LX = M
to have a solution X over IK is that

rank [ M L ]n×(m+p) = rank L

The proposition is true for any field IK including IR, C,


l rational numbers, GF(2), etc.

3.2.7 Linear Equations: Uniqueness, Kernel, Monomorphism

Again let L : X → Y and ȳ ∈ Y be given. Suppose that x̄ is a solution to

L(x) = ȳ (3.29)

Our aim is to describe all possible solutions to (3.29). For this we define the following set of vectors called
the kernel of L.

kernel L = {x : L(x) = 0, x ∈ X }
In other words, the kernel of L is the set of all vectors x in X satisfying L(x) = 0. The reader should verify
that kernel L is a subspace of X .

Returning to our problem, suppose that x̃ is any solution to (3.29); i.e., L(x̃) = ȳ. Since L(x̄) = ȳ it
must be that L(x̃) − L(x̄) = 0 and thus becasue L is linear, that L(x̃ − x̄) = 0. Therefore

x̃ − x̄ ∈ kernel L (3.30)

Conversely if (3.30) holds and x̄ is a solution to (3.29) then L(x̃ − x̄) = 0 and L(x̄) = ȳ so x̃ is a solution to
(3.29) as well. In other words the set of all solutions to (3.29) consists of all vectors of the form x̄ + z where
x̄ is a solution and z is a vector in kernel L.

From these observations it is clear that (3.29) will have at most one solution just in case kernel L = 0.
Linear transformations with this property are important enough to be given a special name. L : X → Y
is said to be a monomorphism4 if kernel L = 0. Thus a necessary and sufficient condition for (3.29) to
have at most one solution is that L be a monomorphism. The reader may wish to verify that L will be a
monomorphism just in case any matrix representation of L has all columns linearly independent.

In the sequel we will use the identity

dim (image L) + dim (kernel L) = dim (X ) (3.31)

The formula is not hard to verify. For a derivation see page 64 of [1].

Suppose that L is a m × n matrix representation of L and suppose that L’s rank is r. Then dim X = n,
dim Y = m, and dim (image L) = r. Then dim (kernel L) = n − r, because of (3.31). It follows that the
number of linearly independent solutions to
Lx = ȳ
is n − r + 1 provided at least one solution exists in the first place.
4 Other names sometimes used are injective linear transformation or one-to-one linear transformation.
3.2.8 Isomorphisms and Isomorphic Vector Spaces

A linear transformation L : X → Y is an isomorphism if it is both a monomorphism and an epimorphism.


Note that any matrix representation of an isomorphism must have both linearly independent rows and
linearly independent columns and hence must be a square, nonsingular matrix.

Two IK-vector spaces U and V are isomorphic if there exists an isomorphism T : U → V mapping one
space into the other. Suppose that T is such an isomorphism. Then kernel T = 0 and image T = V.
It follows from (3.31) that dim V = dim U. In other words, if two finite dimensional vector spaces are
isomorphic then they must both have the same dimension. Conversely, if U and V are any two vector
spaces with the same finite dimension - say n - then they must be isomorphic. To prove that this is so, let
{ui : i ∈ {1, 2, . . . , n}} and {vi : i ∈ {1, 2, . . . , n}} be bases for U and V respectively and define T : U → V so
that T( ui ) = vi , i ∈ {1, 2, . . . , n}. We leave it to the reader to verify (1) that T is a monomorphism, and (2)
that T is a epimorphism {and thus an isomorphism} and therefore that U and V are isomorphic as claimed.

The fact that all n-dimensional vector spaces over IK are isomorphic has an important consequence.
What’s implied is {say for IK = IR} that all real n-dimensional spaces differ from IRn by at most an
isomorphism. It turns out that for many practical problems, this difference is not too important. For
example, if a problem begins with an abstract vector space such as the vector space of all real polynomials
of degree less than n, then the problem can be transformed via isomorphism into an equivalent problem in
IRn where analysis can be carried out in more concrete terms.

3.2.9 Endomorphisms and Similar Matrices

A linear transformation A : X → X which maps a vector space into itself is called an endomorphism of X .
In representing A as a matrix, one always chooses the same basis for A’s domain X and A’s codomain X .

Suppose that {xi : i ∈ {1, 2, . . . , n}} and {x̄i : i ∈ {1, 2, . . . , n}} are two different bases for X and that
A and Ā are the matrix representations of A is these respective bases. Our aim is to explain how these two
representations of A are related. In the light of the discussion in section 3.2.4, it is clear that

Ā = PAQ

where P = [ pir ]n×n , Q = [ qsj ]n×n , and the pir and qsj are numbers such that
P 
xr = P ni=1 pir x̄i
n , r, j ∈ {1, 2, . . . , n}
x̄j = s=1 qsj xs

Thus
n
X n
X
x̄j = qsj pis x̄i , j ∈ {1, 2, . . . , n}
s=1 i=1
so
n
X
x̄j = bij x̄i j ∈ {1, 2, . . . , n}
s=1

where
n
X
bij = pis qsj , i, j ∈ {1, 2, . . . , n}
s=1

These last two equations respectively imply that [ bij ] is a matrix representation of the identity map x −→ x
and also the matrix product PQ; i.e.,
PQ = In×n
Therefore Q = P−1 so
Ā = PAP−1 (3.32)
This is an example of “matrix similarity.” In particular, two n × n IK-matrices A and B are said to be
similar if there exists a nonsingular IK-matrix T such that

B = TAT−1

and the function A 7−→ TAT−1 is called a similarity transformation. It is easy to verify that matrix
similarity is an equivalence relation on the class of all n × n matrices over IK.

What (3.32) shows is that two different matrix representations of an endomorphism A, resulting from two
different choices of a basis for X , are related by a similarity transformation. From this we can conclude that
the matrix representation of an endomorphism is uniquely determined only up to a similarity transformation.
In the sequel we will study in detail those properties of a matrix A which remain unchanged or invariant under
similarity transformations. Such coordinate-independent properties typically turn out to be the physically
significant attributes of the mathematical model under consideration.

This finishes our discussion of basis linear algebra. There are two more topics in the area which we intend
to discuss. The first deals with similarity transformations and the second with congruence transformations.
Both topics will be treated later in the notes. Similarity transformations are what result when one makes
changes of variables in ordinary linear equations.
Chapter 4

Basic Concepts from Analysis

In the next chapter we will begin to study ordinary differential equations. To do this it is necessary to have
in hand a number of basic concepts from calculus and real analysis. The aim of this short chapter is to
briefly review the specific concepts which are need.

4.1 Normed Vector Spaces

In order to be able to talk about solutions to differential equations, we need to extend various ideas from
single-variable real analysis or calculus to vectors of n-variables. The device which makes these extensions
possible is the concept of a “norm.”

Let V be any vector space over IK {i.e., IR or C};


l here V need not even be finite dimensional. A
nonnegative-valued function mapping domain V into codomain IR, written || · ||, is a norm on V if the
following rules all hold:

i. ||v1 + v2 || ≤ ||v1 || + ||v2 ||, ∀vi ∈ V

ii. ||µv|| = |µ|||v||, ∀v ∈ V, µ ∈ IK

iii. ||v|| = 0 ⇔ v = 0

A vector spaces equipped with a norm is called a normed vector space. A norm can be thought of as a
generalization to V of the familiar idea of absolute value as defined on IK.

Rule i. is the triangle inequality. The requirement that a norm be nonnegative - valued at each point in
its domain means that a norm is a positive semi-definite function. Rule iii. mean that ||v|| can be zero if
and only if v = 0; positive-semidefinite functions with this property are often call positive-definite functions.

For the present we will need norms defined only on the vector spaces IRn for various values of n. For each
such vector space there are a great many ways to define a norm. For fixed n, one class of norms is those of
the form
n
! p1

X
p
||x||p = |xi |
i=1

51
where p is a positive integer and xi is the ith entry of x ∈ IRn . The most common of these are the one-norm
n
X
||x||1 = |xi |
i=1

and the two-norm v


u n
uX √
||x||2 = t x2i = x′ x
i=1

which is usually called the Euclidean Norm. Also common is the infinity norm

||x||∞ = max |xi |
i

One way to extend these definitions to the linear space IRn×m for m > 1 is to define the p - norm of each
matrix A ∈ IRn×m to be the same as the p - norm of the nm - column vector obtained by stacking the
columns of A on top of each other. While this would be a perfectly good definition, it would fail to have the
“submultiplicative property” discussed below, for values of p > 1. A better way to define the p - norm of A
is suggested by the observation that the definition of the p - norm of x ∈ IRn implies that
||xy||p
||x||p = sup
y∈IR ||y||p

where sup is an abbreviation for “supremum1 .” For m > 1, we define the p - norm of A ∈ IRn×m to be
∆ ||Ay||p
||A||p = sup
y∈IRm ||y||p
It can be shown that
n
X m
X
||A||1 = max |aij | and ||A||∞ = max |aij |
j i
i=1 j=1

where A = [ aij ]. Latter in these notes it will be shown that ||A||2 is the square root of largest “eigenvalue”
of the symmetric matrix A′ A; this number, in turn, is sometimes called the largest “singular value” of A.

As we’ve already noted, the reason for defining p -norms is this way is because each has the sub-
multiplicative property; that is for all p > 1, including p = ∞,
||AB||p ≤ ||A||p ||B||p , ∀A ∈ IRn×m , ∀B ∈ IRm×r (4.1)
Observe that for r = 1, (4.1) is an immediate consequence of the defintion of a p - norm. It should be
noted that the sub-multiplicative property does necessarily not hold for other norms defined on IRn×m nor
for norms defined on infinite dimensional spaces.

It can be verified that for any fixed positive integers p, q, m, n, including possible p = ∞, there are
constants cpq and cqp depending on n and m, such that
||X||p ≤ cpq ||X||q and ||X||q ≤ cqp ||X||p (4.2)
for all matrices X ∈ IRn×m . This means, for example, that if f : S → IRn is bounded with respect to || · ||p ,
then it is bounded with respect to || · ||q for every value of q. To this extent all such p-norms on IRn×m are
equivalent2 . For many purposes, one such p-norm on IRn×m proves to be just as useful as another. In the
sequel we will often drop the subscript p from || · ||p when what’s being discussed is valid for any p.
1 A function f : S → IRn defined on some set S is bounded {with respect to the p - norm || · || } if there is a finite number
p
C such that ||f (x)||p < C for all x ∈ S. If f is bounded, the smallest value of C with the aforementioned property is denoted

by supS ||f ||p and is called the supremum of f over S with respect to || · ||p . If f is not bounded, supS ||f ||p = ∞. Why is
it necessary to introduce the term “supremum” when we already have the term “max?” In other words, what’s the difference
between sup and max?
2 This is not necessarily so for norms defined on infinite dimensional vector spaces.
4.1.1 Open Sets


Recall that an open interval of the real line, written (t1 , t2 ) is the set of point (t1 , t2 ) = {t : t1 < t < t2 }.
Here t1 and t2 can be either finite numbers or −∞ and/or +∞ respectively. Similary, for finite t1 and t2 we
can define a closed interval [t1 , t2 ] to be the set {t : t1 ≤ t ≤ t2 }. The difference between [t1 , t2 ] and (t1 , t2 )
is thus that the former contains its end points t1 and t2 whereas the latter does not. Right half-open intevals
and left-half open intervals, written [t1 , t2 ) and (t1 , t2 ] respectively, are defined in the obvious ways.

Open and closed intervals are examples of open and closed sets in IR. Other examples of open sets in IR
are finite unions of open intervals. Similarly the union of a finite number of closed intervals is an example
of a closed set in IR. Using norms it is possible to extend the ideas of open and closed sets to vector spaces
of any dimension. Let V be any vector space {finite dimensional or not} and let || · || be a norm defined on
V. A neighborhood or open ball of radius r about a vector v̄ ∈ V is the set of points

{v : ||v − v̄|| < r}

An open set S ⊂ V is simply any subset of V whose elements all have neighborhoods which lie completely
inside of S. Thus the interior of a rectangle drawn on this page is an open set R in any real p-normed
two-dimensional vector within which the page resides.


v q


Note that the set R̄ consisting of all points within the rectangle and on its boundary is not open because it
is impossible to construct an open neighborhood of any point on the boundary which lines totally within R̄.


v q


The rectangle plus its boundary is an example of a “closed-set,” which in turn is a generalization of the idea
of a closed interval. For any vector space V, a subset W is a closed set if its complement in V is an open set.

4.2 Continuous and Differentiable Functions

Let f (·) be a real-valued, scalar function whose domain Ω is a connected subset3 of IRn . We assume that
the reader knows what it means for f to be continuous or differentiable at a point x̄ ∈ Ω. Recall that f
is a continuous function on Ω, if it is continuous at each point in its domain. We say that f continuously
differentiable if the row vector of partial derivatives

∂f ∆ ∂f ∂f ∂f
= [ ∂x1 ∂x2 ... ∂xn ]
∂x
3 By connected we mean that it is possible to connect any two points in Ω with a curve which does not leave Ω.
exists at all points in Ω and each partial derivative is a continuous function on Ω. For example, for even

values of p, f (x) = ||x||p is a continuously differentiable function on IRn . All p norms are continuous.

The notions of continuity, differentiability and continuous differntiability extend to vector-valued func-
tions of n variables in a natural way. For out purposes it suffices to say that a function g : Rn → IRm is
continuous {resp. differentiable, continuously differentiable} if each component of g is a continuous {resp.
differentiable, continuously differentiable} scalar-valued function. By the way, note that the derivative of an
m-vector  
g1
∆  g2 
g= 
 .. 
.
gm
whose elements gi are scalar-valued functions of an n-vector

x1
∆  x2 
x= 
 ...  ,
xn

is an m × n matrix of the form  


∂g1 ∂g1 ∂g1
∂x1 ∂x2 ··· ∂xn
 
∂g ∆  .. .. .. .. 
= . 
∂x  .
 . . 

∂gm ∂gm ∂gm
∂x1 ∂x2 ··· ∂xn m×n

In the event that x is a differentiable function on IR, then

dx(t)
dt
is a n-vector whose ith component is
dxi (t)
dt
Note that by the chain rule
dg(x(t)) dx(t)
= gx
dt dt
where gx is the m × n matrix
∆ ∂g(µ)
gx =
∂µ ∆
µ=x(t)

In the event that g’s domain Ω is a interval {i.e., a connected subset of IR} it is possible to generalize
slightly the concept of continuity. We say that g is a piecewise-continuous function if, on each subinterval
of Ω of finite length, (1) g has at most a finite number of points of discontinuty and (2) at each such point
g has unique, finite limits which approached from above or below. A standard square wave on the real time
interval 0 ≤ t < ∞ is a good example of the graph of a piecewise-continuous function.

Suppose g’s domain Ω is a an interval with finite-valued endpoints t1 and t2 . If g is at least piecewise
continuous, the integral of each of its components over this interval is well-defined. We use the notation
Z t2
g(µ)dµ
t1
to denote the m-vector whose ith component is
Z t2
gi (µ)dµ
t1

Using the definition of the Riemann Integral of a continuous function, as the limit of a Riemann Sum, it is
not difficult to prove that for t2 ≥ t1 ,
Z t2 Z t2
g(µ)dµ ≤ kg(µ)k dµ (4.3)
t1 t1

In other words, the p - norm of the integral of a piecewise continuous function is less than or equal to the
absolute value integral of the p - norm of the function. We’ll make use of this inequality a little later.

4.3 Convergence

A given sequence of vectors x1 , x2 , . . . , xi , . . . in IRn is said to converge to a limit x̄ ∈ IRn if the sequence
of numbers
||xi − x̄||
converges to zero. In other words x1 , x2 , . . . , xi , . . . converges to a limit x̄ if for any given positive tolerance
{i.e., number} ǫ > 0, there is an integer N which is sufficiently large so that all for all i ≥ N the normed
difference ||xi − x̄|| is within the given tolerance; i.e.,

||xi − x̄|| < ǫ, ∀i ≥ N

When this so we call x1 , x2 , . . . , xi , . . . a convergent sequence and we write

lim xi = x̄ or xi → x̄
i→∞

The concept of a norm thus links the idea of a convergent sequence of vectors to the familar idea of convergence
of a sequence of numbers.

Using the concept of a convergent sequence it is possible to redefine a closed set to be any set S within
which all sequences converge to limits which are also in S. This definition proves to be entirely equivalent
to the definition of a closed set given earlier.

As a practical matter it is usually of greater interest to determine if a given sequence converges to a limit,
than it is to determine what the sequence’s limit is, if there is a limit at all. An important idea which enables
one to address the convergence question without having to compute a limit, is the concept of a “Cauchy
Sequence.” A given sequence of vectors x1 , x2 , . . . , xi , . . . in IRn is said to be a Cauchy Sequence if for any
given tolerance ǫ > 0 there is an integer N which is sufficiently large so that for all i ≥ N and all j ≥ N the
normed difference ||xi − xj || is within the given tolerance; i.e.,

||xi − xj || < ǫ, ∀i, j ≥ N

Let us note that any convergent sequence must be a Cauchy Sequence; for if ǫ > 0 is any tolerance and N
is such that
ǫ
||xk − x̄|| < , ∀k ≥ N
2
then for all i, j ≥ N

||xi − xj || = ||(xi − x̄) + (x̄ − xj )|| ≤ ||xi − x̄|| + ||x̄ − xj || ≤ ǫ


The notions of limit, convergence and Cauchy Sequence are not limited just to IRn or even to normed
spaces which are finite dimensional. Moreover convergent sequences are always Cauchy Sequences, whether
the vector space in quesion is finite dimensional or not. On the other hand, Cauchy Sequences are not
necessarily convergent sequences. Normed linear spaces which have the property that every Cauchy Sequence
is convergent are said to be complete vector spaces. Such spaces are often called Banach Spaces after the
mathematician who first studied them. Determining what kinds of normed spaces are Banach Spaces is a
topic within the area of functional analysis. It can be shown that IRn equipped with any p - norm is a
Banach Space. It can also be shown that for any closed finite-length interval [t1 , t2 ], the linear vector space
of continuous functions f : [t1 , t2 ] → IRn equipped with the norm

||f ||cont = max ||f (t)|| (4.4)
t∈[t1 , t2 ]

is a Banach Space. Another way of saying this is as follows.

Theorem 1 (Uniform Convergence) Let IRn be equipped with any p-norm || · ||. Let [t1 , t2 ] be a closed,
bounded interval and let f1 , f2 , . . . , fi , . . . be a sequence of continuous functions fi : [t1 , t2 ] → IRn . Suppose
that for any tolerance ǫ > 0, there is a number N sufficiently large so that for every i, j ≥ N

max ||fi (t) − fj (t)|| < ǫ (4.5)


t∈[t1 , t2 ]

Then there is a continuous function f¯ : [t1 , t2 ] → IRn such that

max ||fi (t) − f¯(t)|| → 0 as i → ∞ (4.6)


t∈[t1 , t2 ]

Condition (4.5) says in effect that f1 , f2 , . . . , fi , . . . is a Cauchy Sequence in the space of continuous
functions f : [t1 , t2 ] → IRn with norm (4.4). Condition (4.6) says that for any tolerance ǭ there must be a
number N̄ , not depending on t such that

||fi (t) − f¯(t)|| < ǭ, ∀t ∈ [t1 , t2 ], i ≥ N̄

The modifier uniform is used to emphasize the fact that N̄ need not depend on t. The proof of this theorem
can be found in most basic real analysis texts.
Chapter 5

Ordinary Differential Equations -


First concepts

There is probably no type of equation of greater importance in the sciences and engineering then a differ-
ential equation. Most of the physical laws we know {e.g., Newton’s Laws, Maxwell’s Equations, . . . } are
characterized in terms of differential equations. The aim of this chapter is to discuss “ordinary” differential
equations.

5.1 Types of Equation

There are a number of different types of equations which one might encounter in the modeling and analysis
of physical processes:

Algebraic Equations:

An algebraic equation is typically of the form

f (x, y, z, · · · , ) = 0

where f (·) is an “algebraic function” and x, y, z, · · · are variables. Such an f might look like

√ y2
f (x, y, z) = c1 x2 + c2 xy + c3 zxy + c4
x3 + 1
where the ci are given numbers or “coefficients.” Systems of algebraic equations consist of a family of
simultaneous algebraic equations such as the proceding, with each equation having the same variables.

Partial Differential Equation:

A partial differential equation is typically of the form


 
∂y ∂y ∂ 2 y ∂z ∂z ∂ 2 z
f y, , , , · · · , y, , , , · · · , t, x = 0
∂t ∂x ∂t∂x ∂t ∂x ∂t2

57
where x and t are independent variables and y and z are functions of x and t or dependent variables.
Dependent variables are variables which go their own ways or vary independently without regard to the
equations involved. Typical examples are time and space variables. independent variables are variables whose
evolutions with respect to the equation’s independent variables are goverened by the particular equation
under consideration. Partial differential equations can have more than two independent [Link] can
also have any number of dependent variables. It is also possible to have simultaneous partial differential
equations, each equation with the same dependent and independent variables.

Ordinary Differential Equation:

Partial differential equations with just one independent variable are sufficiently special to be given their own
name: they are called ordinary differential equations. In this course we will deal exclusively with ordinary
differential equations. Partial differential equations will be covered in the sequel to this course given in the
spring.

One fairly general type of ordinary differential equation, with y a dependent scalar-valued variable and
t as the independent variable, is of the form
 n 
d y d(n−1) y d2 y dy
h , , . . . , , , y, t =0
dtn dt(n−1) dt2 dt

where h(·) is some given function of its arguments. The order of this equation is the largest value of i for
which
di y
dti
appears explicitly. An example of such an equation, of order six, is
 6

2d y d3 y
sin y + log(t) + cos(t) = 0
dt6 dt3

There are two ways in which the preceding generalizes. First, the number of dependent variables might
be greater than one and second, the number of equations might be greater than one. For example, one might
have the simultaneous differential equations

dy 2 d4 z
2
+ y 4 e−t = 0
dt dt

d3 z
+y = 0
dt3

Ordinary Recursion Equation:

An ordinary recursion equation or difference equation is very much like an ordinary differential equation
except that the dependent variable t take on only discrete values {i.,e., t = 1, 2, 3, . . .} and the derivative
d
dt (·) is replaced with shift δ{·}. Such an equation would be of the form
 
h δ n {y}, δ (n−1) {y}, . . . , δ 2 {y}, , δ{y} y, δ −1 {y}, δ −2 {y}, . . . , t = 0

where

δ i {y} = y(t + i)
An example of a recursion equation is

sin y 2 δ 6 {y} + log(t)δ 3 {y} + δ −2 {y} + cos(t) = 0

5.2 Modeling

The problem of deciding which differential equation or recursion equation is appropriate to describe the
physical behaviour in a particular application is usually called modeling. Modeling problems range from
very easy to extremely difficult. Although this is not a course on modeling, we shall nevertheless give several
examples of the modeling process.

In the sequel we shall adopt Newton’s dot notation

∆ dx ∆ d2 x
ẋ = , ẍ = ,...
dt dt2
For recursion equations where t take on only discrete values, we shall adopt the diamond notation
⋄∆ ⋄⋄ ∆
x= δ{x} = x(t + 1), x = δ 2 {x} = x(t + 2), . . .

In neither case need the independent variable t necessarily represent time.

Example 23 Newton’s Gravitational Law applied to two masses yields the nonlinear, second order differ-
ential equation
K
mẍ = − 2
x
where K is a constant.

x
mass m

mass M

Example 24 If m were a given decreasing function of time, then in place of mass times acceleration, one
would use the time rate of change of momentum. Thus for this case the correct equation would be

d K
{mẋ} = − 2
dt x
or
K
mẍ + ṁx +
x2
Here x is a dependent variable and m and ṁ are time-varying coefficients.
Example 25 Consider the following mechanical system consisting of the series connection of a stiff {nonlinear}
spring with force/displacement function

f (µ) = K1 µ3 ,
a mass M1 which slides on a surface with friction coefficient B, a linear spring with force/displacement
constant K2 , and a mass M2 which slides on a frictionless surface.

f (µ)
y1 y2

f (·) K2

M1 M2 µ
 µ -
HH
Y
HH
HH
B

To write the equations which model the system’s motion, let y1 and y2 denote the displacements of masses
M1 and M2 from their respective equilibrium positions. Application of Newton’s Law provides

M1 ÿ1 = −B ẏ1 − K1 y13 − K2 (y1 − y2 ) (5.1)


M2 ÿ2 = −K2 (y2 − y1 ) (5.2)

It is possible to combine these two equations into a single equation with one dependent variable. For example,
by twice differentiating (5.1) one gets

d4 y1 d3 y1 
M1 4
= −B 3 − K1 6y1 ẏ12 + 3y12 ẏ1 − K2 (ÿ1 − ÿ2 ) (5.3)
dt dt
Elimination of the term (y1 − y2 ) from (5.1) and (5.2) and then solving for ÿ2 provides
M1 B1 K1 3
ÿ2 = − ÿ1 − ẏ1 − y
M2 M2 M2 1
Substitution into (5.3) results in the 4th - order nonlinear differential equation
 
d4 y1 d3 y1  2 2 M1 B1 K1 3
M1 4 = −B 3 − K1 6y1 ẏ1 + 3y1 ẏ1 − K2 ÿ1 − K1 ÿ1 + ẏ1 + y (5.4)
dt dt M2 M2 M2 1

Example 26 In 1202 Leonardo of Pisa, known as Fibonacci, discovered the sequence of numbers

1, 1, 2, 3, 5, 8, 13, . . .

while studying the population growth patterns of rabbits. This list of numbers y(1), y(2), . . ., now known
as the Fibonacci Sequence, can be generated by the linear recursion equation

y(t + 2) = y(t + 1) + y(t), t ∈ {1, 2, . . .}


∆ ∆
with the initial conditions y(1) = 1 and y(2) = 1. The sequence has a number of curious properties. For
example the limit of the ratio
y(t + 1)
y(t)
as t tends to infinity turns out to be the “golden mean.1 ”” Another curious fact is that

y(12j) = (12)j , j ∈ {1, 2, . . .}

Properties such as these are exploited in a variety of ways in the development of practical computational
algorithms for many purposes. Later in the course, we will know enough about recursion equations to
verify that these properties are indeed correct. By the way, do you think the sequence of prime numbers
1, 2, 3, 5, 7, 11, 13 . . . can be generated be a {linear} recursion equation?

Example 27 Suppose for a given a real-valued scaler function f of a real variable x, it is desired to solve
the equation
f (x) = 0
One way to try to do this is to generate a sequence of successive approximations to a solution x = x̄ using
Newton’s Method. Suppose xi is the ith approximation of x̄. The heurestic idea upon which the generation
of the next approximation xi+1 is based is as follows. If x̄ were a correct solution, then one would have

f (x̄) = 0 or f (xi + δi ) = 0

where

δi = x̄ − xi
Formally expanding the preceding in a Taylor series about xi one would get

f (xi ) + fx (xi )δi + {higher order terms in δi } = 0

where
∆ df (x)
fx (x) =
dx
Therefore to first approximation

f (xi ) f (xi )
δi ≈ − or x̄ ≈ xi −
fx (xi ) fx (xi )

Thus a reasonable way to generate xi+1 would be via the equation

f (xi )
xi+1 = xi −
fx (xi )

Of course without saying more about f , there is no guarantee that this equation will generate a sequence
of values x1 , x2 , x3 , . . . which converges to a value x̄ for which f (x̄) = 0. One definition for f for which
convergence from any initial chosen value x1 can be proved is

f (x) = x2 − d

where d is a positive number. Indeed this particular choice for f leads to the highly efficient algorithm

xi d
xi+1 = +
2 2xi

for recursively computing the square root of d.


1 Geometrically the golden mean is the ratio of length to width of any rectangle which has the property that the ratio of

its length to its width equals the ratio of its length plus its width to its length. Rectangles of this type were thought by the
ancient Greeks to have the most pleasing proportions possible and consequently were often used in their architecture.
5.3 State Space Systems

Given a differential equation {or a recursion equation} there arise several key issues:

1. How might one go about solving the equation?

(a) What is meant by a solution?


(b) When does such a solution exist and when is it unique?
(c) How might a solution be found using analytic methods?
(d) How might such a solution be found using a computer?
(e) How might a solution be approximated when exact solvability is hopelessly difficult?

2. How might one go about determining various properties of whole families of solutions without having
to actually construct the family?

(a) What are the “equilibrium points” of a differential equation?


(b) How might one predict periodic behavior or “limit cycles”?
(c) How might one predict “unstable” or “chaotic” behavior?

In order to address these issues in a systematic manner, it is especially useful to focus on differential
equation represented in “state space” form. By a state space system or dynamical system of dimension n is
meant a system of coupled first-order differential equations of the form

ẋ1 = f1 (x1 , x2 , . . . , xn , t)
ẋ2 = f2 (x1 , x2 , . . . , xn , t)
..
.
ẋn = fn (x1 , x2 , . . . , xn , t)

Here the xi are scalar-valued dependent variables and each fi is a given scalar-valued function of the xi
and the independent variable t. We’ll ofter refer to t as time, but it could just as easily represent another
quantity such as distance. The xi are often called state variables. An analogous discrete dynamical system
would be of the form

x1 = f1 (x1 , x2 , . . . , xn , t)

x2 = f2 (x1 , x2 , . . . , xn , t)
..
.

xn = fn (x1 , x2 , . . . , xn , t)

It is possible to denote the preceding systems a little more concisely using the notations

ẋ = f (x, t) and x= f (x, t)

respectively where
   f (x , x , . . . , x , t) 
x1 1 1 2 n

∆  x2  ∆  f 2 (x1 , x2 , . . . , xn , t) 
x= 
 ...  and f (x, t) = 
 .. 

.
xn fn (x1 , x2 , . . . , xn , t)
5.3.1 Conversion to State-Space Form

Many ordinary differential equations and recursion equations can be converted to equivalent state-space
forms. The conversion process is often straight-forward. Consider for example, the nth order scalar differ-
ential equation  n 
d y d(n−1) y d2 y dy
h , , . . . , , , y, t =0
dtn dt(n−1) dt2 dt
Suppose that it is possible to express
dn y
dtn
as a function of lower derivatives and t:
 
dn y d(n−1) y d2 y dy
=g , . . . , , , y, t
dtn dt(n−1) dt2 dt

To carry out the conversion, one could define


n−1
∆ ∆ dy ∆ d y
x1 = y, x2 = , . . . . . . , xn = n−1
dt dt
This would imply that
dy
ẋ1 = dt = x2

d2 y
ẋ2 = dt2 = x3

.. ..
. .
dn−1 y
ẋn−1 = dy n−1 = xn

dn y
ẋn = dtn = g(xn , xn−1 , . . . , x1 , t)
Thus in vector form, these equations would look like

ẋ = f (x, t)

and
y = Cx
where  
x2
 x3 
∆  ..  ∆
f (x, t) = 
 .

 C = [1 0 · · · 0 ]1×n
 xn 
g(xn , xn−1 , . . . , x1 , t)
This particular representation is sometime called phase variable form. There are lots of other ways to define
the xi .

Example 28 If n = 2 and

h(ÿ, ẏ, y, t) = q(ẏ, y, t)ÿ + p(ẏ, y, t)
where p and q are given functions, then g would be

∆ p(ẏ, y, t)
g(ẏ, y, t) = −
q(ẏ, y, t)
In phase variable form x would be a two-vector and f would be
 
x2

f (x, t) =  
− p(x 2 ,x1 ,t)
q(x2 ,x1 ,t)

Example 29 The original set of equations describing the motion of the masses in the mechanical system
discussed in Example 25 were
M1 ÿ1 = −B ẏ1 − K1 y13 − K2 (y1 − y2 ) (5.5)
M2 ÿ2 = −K2 (y2 − y1 ) (5.6)
Before we explained how to combine these two equations into the single 4th - order nonlinear differential
equation
 
d4 y1 d3 y1  M1 B1 K1 3
M1 4 = −B 3 − K1 6y1 ẏ12 + 3y12 ẏ1 − K2 ÿ1 − K1 ÿ1 + ẏ1 + y1
dt dt M2 M2 M2
It is also possible to convert (5.5) and (5.6) into an equivalent state space system. For example, this can be
done by simply defining
∆ ∆ ∆ ∆
x1 = y1 , x2 = ẏ1 , x3 = y2 x4 = ẏ2
These definitions would result in the state space system

      
ẋ1 0 1 0 0 x1 0
      
   K2 B K2     K1 
 ẋ2   − M1 −M M1 0   x2   M 
   1    1  3
  =    −   x1
      
 ẋ3   0 0 0 1   x3   0 
      
K2 K2
ẋ4 M2 0 −M 2
0 x4 0
 
x1
    


y1 1 0 0 0  x2 
  =   


 
y2 0 0 1 0  x3 
 
x4

Example 30 As a final example, consider the following lumped, linear electrical network driven by a voltage
v.

R1
R2
v L

C y2
y1
To model this network’s behavior, let y1 and y2 denote the current through the inductor and voltage across
the capacitor respectively. Then y1 + C ẏ2 must be the current through resistor R1 . Kirkoff’s voltage law
dictates that the v must equal the sum of the voltage drops across resistor R1 and the capacitor. In other
words,
v = R1 (y1 + C ẏ2 ) + Lẏ1 (5.7)
In addition, the voltage drop across the inductor must equal the sum of the voltage drops across resistor R2
and the capacitor; i.e.,
Lẏ1 = R2 C ẏ2 + y2 (5.8)
It is not very difficult to rewrite these equations in the equivalent state space form
      
ẏ1 − (RR 1 R2
1 +R2 )L
R1
(R1 +R2 )L y1 R2
(R1 +R2 )L
 =   +  v
R1 1 1
ẏ2 − (R1 +R2 )C − (R1 +R2 )C y2 (R1 +R2 )C

There are systematic methods for constructing state space models of this type for extremely complicated
{vlsi} networks with many thousands of circuit elements. Without such models and systematic methods
for analyzing and synthesizing them, the design and implementation very large scale integrated electrical
networks {i.e., computer chips!} would be a hopelessly difficult task.

5.4 Initial Value Problem

What we want to do now is to focus attention on state space differential equations of the form

ẋ = f (x, t) (5.9)

where f : IRn × [0, ∞) → IRn is at least a continuous function of x and a piecewise continuous function of t.
To grasp the main ideas of the theory of ordinary differential equations, it is necessary to understand what
it means for these equations to have a “solution,” when such a solution exists, and when it is unique2 .

By a solution to (14.5) on a finite interval I ⊂ [0, ∞) of positive length, is meant a continuous function
φ : I → IRn which is differentiable at all but at most a finite number of points in I and satisfies

φ̇ = f (φ, t)

where ever its derivative φ̇ exists. φ is a solution to (14.5) on an interval I of infinite length if its restriction3
to each finite subinterval of I is a solution to (14.5).

Example 31 The continuous function



φ(t) = |t − 2|, t ∈ [0, ∞)

is a solution on [0, ∞) to the differential equation

ẋ = f (x, t)
2 It is possible to make precise what is meant by a solution to (14.5) in the case when f is not a continuous function of x

[2]. Equations of this type arise, for example, if one is modeling static friction in a mechanical system or if one is interested in
describing the behavior of an electromechanical system employing relays or other types of switches. An example a differential

equation with a discontinuous right hand side is ẋ = −sign x where sign (x) = 1 if x ≥ 0 and sign (x) = −1 otherwise. The
reader may with to consider what a “solution” might look like for this system assuming x(0) = 0.
3 The restriction of a function h : S → T to a subset U ⊂ S is the function k : U → T for which k(u) = h(u) ∀u ∈ U . Note

that h and k are different functions even though both use the same rule to map points from U to T
where f : IR × [0, ∞) → IR is the function
(
∆ 1 if t ≤ 2
f (x, t) =
−1 if t > 2

for all x ∈ IR. Note that φ̇ is piecewise continuous but not continuous.

For our purposes, the initial value problem is as follows: Given an initial time t0 ∈ [0, ∞) and an initial
state vector x0 ∈ IRn find, if possible, an interval I ⊂ [0, ∞) of positive length which contains t0 , and a
solution φ to (14.5) on I which passes through the “state” x0 at t = t0 . Lets look at some examples.

Example 32 Certainly the most basic of all initial value problems is to find a solution to the first-order,
linear differential equation
ẋ = a(t)x (5.10)
for a given initial time t0 ≥ 0 and state x0 ∈ IR. Here a : [0, ∞) → IR is a given piecewise continuous
function. We claim that Rt
∆ a(µ)dµ
φ(t) = x0 e t0
solves this initial value problem on the whole interval [0, ∞). This is because φ is well-defined and continuous
on the interval [0, ∞), because
Rt
a(µ)dµ
φ̇(t) = a(t)x0 e t0
= a(t)φ(t), t ∈ [0, ∞),
and because
φ(t0 ) = x0
Note that if a(t) were a constant, say a(t) = λ, then φ would be the exponential time function
φ(t) = x0 e(t−t0 )λ

The initial value problem turns out to always have a solution φ on some interval I provided f is a
continuous function of x and a piecewise-continuous function of t. We will not prove this, in part because
the kind of f ’s typically encountered in science and engineering are more than continuous in x, and by
exploiting this one can prove things more easily. Moreover, it turns out that continuity with respect to x is
not a strong enough condition to ensure uniqueness of solutions.

Example 33 The differential equation √


3
ẋ = x
∆ ∆
together with the initial data t0 = 0 and x0 = 0 demonstrates nonuniqueness. In particular,
  32
2
φ(t) = t and φ(t) = 0
3
both satisfy the differential equation and the initial condition φ(0) = 0.

To avoid unpleasent uniqueness questions, we shall impose a slightly stronger smoothness condition on f
then mere continuity with respect to x. The specific properties we shall require f to have are stated below.
Fortunately, the f ’s encountered in the study of most problems of interst to scientists and engineers have
these properties.
Properties:

1. For each closed, bounded subset4 S ⊂ IRn and each bounded interval I ⊂ [0, ∞) there exists a constant
λ such that
||f (x, t) − f (y, t)|| ≤ λ||x − y||
for all x, y ∈ S and all t ∈ I.
2. For each fixed x ∈ IRn , the function f (x, ·) : [0, ∞) → IRn , t 7−→ f (x, t) has at most a finite number
of points of discontinuity on each bounded subinterval of [0, ∞), and at each such point, f (x, ·) has
unique limits when approached from above or from below.

A function f with the preceding properties is said to be locally Lipschitz Continuous on IRn and piecewise-
continuous on [0, ∞); and λ is often called a Lipschitz Constant. If Property 1 holds for S equal to all of
IRn , then f is globally Lipschitz continuous on IRn . Of course in this case λ may still depend on I.

Example 34 The function f (x) = x2 is locally Lipschitz on IR because for any given closed, bounded subset
S ⊂ IR
||x2 − y 2 || ≤ λ||x − y||, ∀x, y ∈ S
where

λ = max ||x + y|| < ∞
x,y∈S

Here we’ve used the relations

||x2 − y 2 || = ||(x + y)(x − y)|| ≤ ||x + y||||x − y||

and the fact that the maximum value {i.e., L} of a continuous function {i.e., ||x + y||} over a closed, bounded
set {i.e., S × S} is finite. Of course λ depends on what S is. Note that f is not globally Lipschitz on IR
because there is no finite number λ such that

||x2 − y 2 || ≤ λ||x − y||, ∀x, y ∈ IR

Let us note that any function of the form

f (x) = An×n x + bn×1

is globally Lipschitz on IRn . This is because

||(Ax + b) − (Ay + b)|| = ||A(x − y)|| ≤ ||A||||x − y||, ∀x, y ∈ IRn

It can be shown that any continuously differentiable function is locally Lipschitz and that any locally
Lipschitz function is continuous. On the other hand, the reverse implications do not hold in general. For ex-

t) = |ax+b| is locally Lipschitz on IR but not continuously differentiable on IR.
ample, the scalar function f (x,√
Similarly the function f (x) = x is continuous but neither locally Lipschitz nor continuously differentiable
on IR. The reader should verify these claims.

In the sequel it will be shown that initial value problem always has a unique solution φ on some interval
I provided f is a locally Lipschitz function of x and a piecewise-continuous function of t. The possibility
that I might have to be something less than the whole real interval [0, ∞) is illustrated by the following
example.
4 A subset S ⊂ IRn is bounded if there is a finite number C such that ||x|| < C for all x ∈ S. If S is bounded, the smallest

value of C with the aforementioned property is called the supremum or sup of S.


Example 35 Consider the first order differential equation

ẋ = x2

As noted before, f (x) = x2 is a locally Lipschitz function on IR. Because of this, for each set of initial time
t0 ≥ 0 and each initial condition x0 , the differential equation must have exactly one continuous solution φ
defined on some interval I which passes through x0 at t0 . Fix t0 ∈ [0, ∞) and x0 > 0 and define the function
θ : [0, ∞) → IR so that

 x0

 1−(t−t0 )x0 if t 6= t̄

θ(t) =


0 if t = t̄
where
1 ∆
t̄ = t0
x0
The reader may wish to verify that θ satisfies the initial condition and also the differential equation except
at the time t = t̄. Note however that since θ is not continuous at t = t̄ it cannot be a solution to the initial
value problem. On the other hand, if we defnie φ to be the restriction of θ to [0, t̄), i.e. φ : [0, t̄) → IR and

φ(t) = θ(t), t ∈ [0, t̄)

then φ must be a solution since it is continuous, and satisfies both the initial condition and the differential
equation. A little thought reveals that [0, t̄) is actually the largest possible interval on which a solution to
the problem exists. Note: It is quite incorrect to say that a solution exists at all points in [0, ∞) except at
t = t̄. The behavior of θ for t > t̄ is irrelevant to the initial value problem.

5.5 Picard Iterations

Consider again the differential equation.


ẋ = f (x, t) (5.11)
Suppose that f : IRn × [0, ∞) → IRn is locally Lipschitz in x and piecewise continuous in t. Suppose that for
given t0 ≥ 0 and x0 , the function φ : [t0 , t1 ] → IRn solves the initial value problem on some closed interval
of positive length [t0 , t1 ]5 . Since φ is required to be continuous, f (φ(t), t) must be piecewise continuous on
[t0 , t1 ]. From this and the requirement that φ must satisfy (5.11) wherever φ̇ exists, it follows that φ must
satisfy the integral equation Z t
φ(t) = x0 + f (φ(µ), µ)dµ, t ∈ [t0 , t1 ] (5.12)
t0

Conversely any function satisfying φ satisfying (5.12) must be continuous and must satisfy φ(t0 ) = x0 .
In addition (5.11) must hold wherever φ̇ exists because of the fundamental theorem of calculus. Therefore
finding a continuous solution φ to (5.11) on [t0 , t1 ] which satisfies the initial condition φ(t0 ) = x0 is equivalent
to finding a solution to (5.12). One way to try to accomplish the latter is to generate a sequence of functions
x1 (t), x2 (t), . . . , xi , . . . using the recursion formula
Z t

xi+1 (t) = x0 + f (xi (µ), µ)dµ, t ∈ [t0 , t1 ], i ∈ {0, 1, 2, . . .} (5.13)
t0
5 All that follows can also be done for φ defined on a closed interval [ta , tb ] ⊂ [0, ∞) of positive length whose interior
contains t0 or whose right end point tb = t0 . The developments for these generalizations are essentially the same.
starting with the initial function

x0 (t) = x0 , t ∈ [t0 , t1 ] (5.14)
For if such a sequence were to converge to a continuous limit x̄(t) then as we shall soon see, x̄(t) would have
to satisfy Z t
x̄(t) = x0 + f (x̄(µ), µ)dµ, t ∈ [t0 , t1 ] (5.15)
t0
In other words if we can insure that the above sequence of functions has a continuous limit x̄, then we can
define φ to be x̄. This procedure is called Picard Iteration.

5.5.1 Three Relationships

In view of Example 35, it should not be supprising that for Picard Iteration to work, a constraint must be
placed on the value of t1 . Basically what we need to do is to pick this number close enough to t0 so as to
insure that the value of each xi (t), at each time t ∈ [t0 , t1 ], remains within some set X on which f satisfies
the Lipschitz condition

||f (x, t) − f (y, t)|| ≤ λ||x − y||, ∀x, y ∈ X, and ∀t ∈ [t0 , t1 ] (5.16)

for some nonnegative Lipschitz Constant λ. In other words t1 and X ⊂ IRn must be chosen so that t0 < t1 ,
x0 ∈ X, and (5.16) and
xi (t) ∈ X, t ∈ [t0 , t1 ], i ∈ {0, 1, 2, . . .} (5.17)
both hold. To do this, two distinct cases need to be considered:

1. f is locally Lipschitz in x: For this case pick any finite positive number b, any finite time t̄ > t0 , and
define

X = {x : ||x − x0 || ≤ b}
Next pick F to be any finite positive number such that
F ≥ sup ||f (x, t)||
x∈X
t∈[t0 , t̄ ]

Note that such a finite F must exist because X and [t0 , t̄ ] are both closed, bounded sets and because
f is continuous in x and piecewise continuous in t. Now define
∆ b
t1 = min{t̄, t0 + }
F
The definitions imply that t0 < t1 ≤ t̄, that
||f (x, t)|| ≤ F, x ∈ X, t ∈ [t0 , t1 ] (5.18)
and that
(t − t0 )F ≤ b, t ∈ [t0 , t1 ] (5.19)
We claim that with these choices, the Picard Iterates xi generated by (5.13) and (5.14) must satisfy
(5.17). Now this is clearly true for i = 0 because of the definitions of x0 (t) and X. Suppose therefore
that for some j ≥ 0, (5.17) hold for all i ∈ {0, 1, . . . , j}. Thus we can use integral inequality (4.3)
together with (5.13), (5.18) and (5.19) to deduce, for t ∈ [t0 , t1 ], that
Z t Z t Z t
||xj+1 (t) − x0 || = f (xj (µ), µ)dµ ≤ ||f (xj (µ), µ)||dµ ≤ F dµ = (t − t0 )F ≤ b
t0 t0 t0

In other words, xj+1 ∈ X for t ∈ [t0 , t1 ]. By induction, (5.17) must therefore hold for all i ≥ 0.
2. f is globally Lipschitz in x: In this case pick t1 to be any finite number greater than t0 and define

X = IRn . The xi then satisfy (5.17). Moreover, since f is globally Lipschitz in x there must be a finite
nonnegative Lipschitz Constant λ {possibly depending on t1 } such that (5.16) holds.

To proceed, we need one additional relationship, namely

||x1 (t) − x0 (t)|| ≤ b, t ∈ [t0 , t1 ] (5.20)

This holds for case 1, when f is only locally Lipschitz in x, because of (5.17) and the definitions of x0 (t) and

X. We can also make (5.20) hold for case 2 by simply defining b = (t1 − t0 )M , where

M = sup ||f (x0 , t)||
t∈[t0 ,t1 ]

This is because for case 2


Z t Z t Z t
||x1 (t) − x0 || = f (x0 (µ), µ)dµ ≤ ||f (x0 , µ)||dµ ≤ M dµ = (t − t0 )M ≤ (t1 − t0 )M
t0 t0 t0

for all t ∈ [t0 , t1 ]. Note that M must be finite because f (x0 , t) is piecewise continuous and [t0 , t1 ] is a
closed, bounded interval.

5.5.2 Convergence

To recap, for either case we’ve been able to define t0 , t1 , b and X so that (5.16), (5.17), and (5.20) all hold.
What we aim to do next is to use these three relations to prove for either case 1 or case 2, that the Picard
Iterates x1 (t), x2 (t), . . . , xi (t) form a uniformly convergent sequence. As a first step toward this end we
shall establish for each i ≥ 0, that
((t − t0 )λ)i
||xi+1 (t) − xi (t)|| ≤ b , t ∈ [t0 , t1 ] (5.21)
i!
where i! denotes i factorial6 .

In view of (5.20) we already know that (5.21) holds for i = 0. Suppose therefore that for some j ≥ 1,
(5.21) holds for all i ∈ {0, 1, . . . , j − 1}. Because of (5.13) and (5.14) we can write
Z t
||xj+1 (t) − xj (t)|| = (f (xj (µ), µ) − f (xj−1 (µ), µ))dµ
t0
Z t
≤ ||f (xj (µ), µ) − f (xj−1 (µ), µ)||dµ
t0

Hence by (5.16) {which holds because of (5.17)}


Z t
||xj+1 (t) − xj (t)|| ≤ λ||xj (µ) − xj−1 (µ)||dµ
t0

But since (5.21) holds at i = j,


Z t
((µ − t0 )λ)j ((t − t0 )λ)(j+1)
||xj+1 (t) − xj (t)|| ≤ λb dµ = b
t0 j! (j + 1)!
6 Recall ∆
that 0! = 1.
Thus by induction, (5.21) holds for all i ≥ 0.

Next we want to do is to prove for any integers i ≥ 0 and j ≥ 1 that

((t − t0 )λ)i (t−t0 )λ


||xi+j (t) − xi (t)|| ≤ b e (5.22)
i!
To do this we write
j−1
X
xi+j (t) − xi (t) = (xi+k+1 (t) − xi+k (t))
k=0

Using this, the triangle inequality and (5.21) we get

j−1
X
||xi+j (t) − xi (t)|| = (xi+k+1 (t) − xi+k (t))
k=0
j−1
X
≤ ||xi+k+1 (t) − xi+k (t)||
k=0
j−1
X ((t − t0 )λ)(i+k)
≤ b
(i + k)!
k=0

But
! 1 1

(i + k)! i! k!
so
j−1
((t − t0 )λ)i X ((t − t0 )λ)k
||xi+j (t) − xi (t)|| ≤ b
i! k!
k=0

((t − t0 )λ)i X ((t − t0 )λ)k
≤ b
i! k!
k=0
((t − t0 )λ)i (t−t0 )λ
= b e
i!
Thus (5.22) is true.

We are now ready to establish the convergence of the xi . First, let us note from (5.22) that for i ≥ 0 and
j≥0
((t1 − t0 )λ)i (t1 −t0 )λ
max ||xi+j (t) − xi (t)|| ≤ b e (5.23)
t∈[t0 , t1 ] i!
Let ǫ > 0 be any positive tolerance. Pick an integer N so large that

((t1 − t0 )λ)i (t1 −t0 )λ


b e ≤ǫ
i!
for i ≥ N . Then (5.23) enables us to write that

max ||xk (t) − xi (t)|| ≤ ǫ, ∀k ≥ N, and ∀i ≥ N


t∈[t0 , t1 ]

Since this is condition (4.5) of the Uniform Convergence Theorem, it must be that x0 (t), x1 (t), . . . , is indeed
a uniformly convergent sequence with a continuous limit; we henceforth denote this limit by x̄(t).
Our next objective is to show that x̄ satisfies (5.15) which, for ease of reference we rewrite here:
Z t
x̄(t) = x0 + f (x̄(µ), µ)dµ, t ∈ [t0 , t1 ] (5.24)
t0

For this let us first note that for i ≥ 0 and t ∈ [t0 , t1 ]


Z t Z t Z t
f (xi (µ), µ)dµ − f (x̄(µ), µ)dµ ≤ ||f (xi (µ), µ) − f (x̄(µ), µ)||dµ
t0 t0 t0
Z t
≤ λ ||xi (µ) − x̄(µ)||dµ
t0
Z t 
≤ λ max ||xi (τ ) − x̄(τ )|| dµ
τ ∈[t0 , t1 ] t0
≤ (t1 − t0 )λ max ||xi (t) − x̄(t)||
t∈[t0 , t1 ]

Since
lim max ||xi (t) − x̄(t)|| = 0
i→∞ t∈[t0 , t1 ]

it must therefore be true that


Z t Z t
lim f (xi (µ), µ)dµ − f (x̄(µ), µ)dµ = 0, ∀t ∈ [t0 , t1 ]
i→∞ t0 t0

Therefore for t ∈ [t0 , t1 ]


Z t Z t
f (xi (µ), µ)dµ → f (x̄(µ), µ)dµ and xi+1 (t) → x̄(t)
t0 t0

as i → ∞. Since the xi satisfy


Z t

xi+1 (t) = x0 + f (xi (µ), µ)dµ, t ∈ [t0 , t1 ], i ∈ {0, 1, 2, . . .},
t0

in the limit (5.24) must hold.

5.5.3 Uniqueness

In this section it will be shown that there is no continuous function defined on on [t0 , t1 ], other than x̄,
which satisfies (5.24). To prove that this is so, we will make use of a special technical result which also finds
application in a number of other areas including most especially the stability analysis of ordinary differential
equations.

Lemma 1 (Bellman-Gronwall) If some numbers tb ≥ ta , some constant c ≥ 0 and some nonnegative,


piecewise-continuous function α : [ta , tb ] → IR, w : [ta , tb ] → IR is a continuous function satisfying
Z t
w(t) ≤ c + α(µ)w(µ)dµ, t ∈ [ta , tb ] (5.25)
ta

then Rt
α(µ)dµ
w(t) ≤ ce ta , t ∈ [ta , tb ] (5.26)
The Bellman-Gronwall Lemma becomes quite plausable as soon as one recognizes that the solution to
the scalar differential equation,

ẇ = αw
w(ta ) = c

or equivalant integral equation Z t


w(t) = c + α(µ)w(µ)du
ta

is Rt
α(µ)dµ
w(t) = ce ta

The lemma remains true if the right and/or left end point is removed from [ta , tb ].

Proof: Set
Rt
v(t) = w(t)e− ta α(µ)dµ
 Z t  R
t
u(t) = c+ α(µ)w(µ)dµ e− ta α(µ)dµ
ta

and note from (5.25) that


v(t) ≤ u(t), t ∈ [ta , tb ] (5.27)
Differentiating the expression for u and then replacing w with
Rt
α(µ)dµ
v(t)e ta

one gets
Rt
u̇ = −αu + αwe− ta
α(µ)dµ

= α(v − u)

Hence by (5.27) u̇ ≤ 0 so u(t) ≤ u(t0 ) = c, t ∈ [ta , tb ]. Therefore by (5.27),

v(t) ≤ c, t ∈ [ta , tb ]
Rt Rt
α(µ)dµ α(µ)dµ
Multiplying through by e ta and then replacing v(t)e ta by w(t) yields (5.26) which is the desired
result.

We now turn to the uniqueness question. For this, fix y0 ∈ IRn and suppose that y : [t0 , t1 ] → IRn is any
continuous function satisfying the initial condition

y(t0 ) = y0 (5.28)

and the equation


ẏ(t) = f (y(t), t) (5.29)
wherever ẏ exists. As we’ve already seen this is equivalent to y satisfying
Z t
y(t) = y0 + f (x̄(µ), µ)dµ, t ∈ [t0 , t1 ] (5.30)
t0

To establish x̄′ s uniqueness it is enough to show that if y0 = x0 then y(t) = x(t), t ∈ [t0 , t1 ].

To begin, let us first note that since x̄ and y are continuous functions and [t0 , t1 ] is a closed bounded
interval, both x̄ and y must be bounded on [t0 , t1 ]. Thus there must be a closed, bounded subset X̄ ⊂ IRn
such that x̄(t) ∈ X̄ and y(t) ∈ X̄ for all t ∈ [t0 , t1 ]. Since f is at least locally Lipschitz in its first argument,
this means that there must be a finite constant λ̄ ≥ 0 such that

||f (y(t), t) − f (x̄(t), t)|| ≤ λ̄||y(t) − x̄(t)||, t ∈ [t0 , t1 ]

From this, (5.30), and (5.24) it thus follows that for t ∈ [t0 . t1 ]
Z t Z t
||y(t) − x̄(t)|| = y0 + f (y(µ), µ)dµ − x0 − f (x̄(µ), µ)dµ
t0 t0
Z t
≤ ||y0 − x0 || + ||f (y(µ), µ) − f (x̄(µ), µ)||dµ
t0
Z t
≤ ||y0 − x0 || + λ̄ ||y(µ) − x̄(µ)||dµ
t0

∆ ∆ ∆ ∆ ∆
Therefore by defining c = ||y0 − x0 ||, ta = t0 , tb = t1 , w = ||y(t) − x̄(t)||, α = λ̄ and applying the
Bellman-Gronwall Lemma, we get

||y(t) − x̄(t)|| ≤ ||y0 − x0 ||e(t−t0 )λ̄ , t ∈ [t0 , t1 ] (5.31)

Therefore, if y0 = x0 it must be true that ||y(t) − x̄(t)|| = 0, t ∈ [t0 , t1 ] and thus that y(t) = x̄(t) for
t ∈ [t0 , t1 ].

5.5.4 Summary of Findings

The last few sections covered a lot of material. Here is a brief summary:

1. The problem of interest has been to find a solution to the initial value problem: Given is a differential
equation
ẋ = f (x, t) (5.32)

where f : IRn × [0, ∞) is a function with the following properties:

Locally Lipschitz Continuity in x: For each closed, bounded subset S ⊂ IRn and each bounded
interval I ⊂ [0, ∞) there exists a constant λ such that

||f (x, t) − f (y, t)|| ≤ λ||x − y|| (5.33)

for all x, y ∈ S and all t ∈ I.


Piecewise Continuity in t: For each fixed x ∈ IRn , the function f (x, ·) : [0, ∞) → IRn , t 7−→ f (x, t)
has at most a finite number of points of discontinuity on each bounded subinterval of [0, ∞), and
at each such point, f (x, ·) has unique, finite limits when approached from above or from below.

Given also is an initial time t0 and an initial state x0 ∈ IRn . The problem then is to find an interval
I ⊂ [0, ∞) of positive length which contains t0 , and a continuous function φ : I → IRn which passes
through x0 at t = t0 and satisfies the equation

φ̇(t) = f (φ(t), t) (5.34)

at all points in I where φ̇ exists.


2. We explained how to construct such a φ for the case when I is a closed interval of positive but finite
length with t0 as its left end point. In particular, we showed how to pick a finite number t1 > t0 in
such a way that the sequence of Picard Iterates
Z t

xi+1 (t) = x0 + f (xi (µ), µ)dµ, t ∈ [t0 , t1 ], i ∈ {0, 1, 2, . . .} (5.35)
t0


initialized with the function x0 (t) = x0 , t ∈ [t0 , t1 ], converges uniformly to such a φ. If t0 is strictly
greater than 0, then using essentially the same ideas, it is possible to pick a second number t2 ∈ [0, t0 ),

so that Picard Iterates defined by x0 (t) == x0 , t ∈ [t2 , t0 ] and
Z t

xi+1 (t) = x0 + f (xi (µ), µ)dµ, t ∈ [t2 , t0 ], i ∈ {0, 1, 2, . . .} (5.36)
t0

also converge uniformly.

3. For the case when f is globally Lipschitz in x {i.e., when for each finite length interval I ⊂ [0, ∞)
there is a constant λ for which (5.33) holds for all x, y ∈ IRn and all t ∈ I} all that is required of t1
for the preceding to be true is that t1 be finite and greater than t0 . It is also true that if t0 > 0, then
the only constraint on the time t2 mentioned above, is that it be in the interval [0, t0 )

4. We then used the Bellman-Gronwall Lemma to prove that if y0 ∈ IRn is arbitrary and y : [t0 , t1 ] → IRn
is any continuous function satisfying the initial condition y(t0 ) = y0 , and the differential equation

ẏ(t) = f (y(t), t)

then
||y(t) − x̄(t)|| ≤ ||y0 − x0 ||e(t−t0 )λ , t ∈ [t0 , t1 ]

5. From the preceding we concluded that φ is the only continuous function defined on [t0 , t1 ] which
satisfies the initial condition φ(t0 ) = x0 and the differential equation (5.34).

From the preceding it is possible to draw a number of useful conclusions:

I. More Uniqueness:

• Suppose φ : [t0 , t1 ] → IRn is the solution to the initial value problem which passes through x0 at
t = t0 .
• Suppose φ̄ : [t̄0 , t̄1 ] → IRn is the solution to the initial value problem which passes through x̄0 at
t = t̄0 .

If there is a time t∗ in the interval [max{t0 , t̄0 }, min{t1 , t̄1 }] at which φ(t∗ ) = φ̄(t∗ ) then, as a
consequence of uniqueness,

φ(t) = φ̄(t), t ∈ [max{t0 , t̄0 }, min{t1 , t̄1 }]

Moreover the the function φ∗ : [min{t0 , t̄0 }, max{t1 , t̄1 }] → IRn


(
∗ ∆ φ(t) for t ∈ [t0 , t1 ]
φ (t) =
φ̄(t) for t ∈ [t̄0 , t̄1 ]

is unambiguously defined, passes through x0 at t = t0 , and satisfies the differential equation (5.32) on
[min{t0 , t̄0 }, max{t1 , t̄1 }]. This too is an immediate consequence of uniqueness.
φ

φ
t0 t0 ∗ t1 t1
φ

II. Maximal Inverval of Existence For fixed t0 and x0 we know that the set

S1 = {t1 : a solution with value x0 at t = t0 exists on [t0 , t1 ] and t0 < t1 < ∞}

is nonempty. Similarly if t0 > 0, the set



S2 = {t2 : a solution with value x0 at t = t0 exists on [t2 , t0 ] and 0 ≤ t2 < t0 }

is also nonempty. Let us define7



t̃1 = sup S1
and 

0 if t0 = 0

t̃2 =


inf S2 otherwise

Let us also define 



 [0, t̃1 ) if t0 = 0

I(t0 ,x0 ) =


(t̃2 , t̃1 ) otherwise
Note whether t̃1 is finite or not there must be a continuous function φ̃ : I(t0 ,x0 ) → IRn which satisfies
(5.32) and passes through the initial state x0 at t = t0 . φ̃ is called a maximal solution because
I(t0 ,x0 ) is the largest possible interval for which there exists a solution φ which passes through x0 at
time t0 . I(t0 ,x0 ) is often called the maximal interval of existence for the initial value problem under
consideration. The uniqueness results discussed earlier, also apply to maximal solutions.
II. Global Existence: There are two important instances when I(t0 ,x0 ) = [0, ∞)
1. If f is Globally Lipschitz: In this case we know that a solution must exist on every closed in-
terval [t2 , t1 ] ⊂ [0, ∞) of finite length which contains t0 . Thus in this case t2 can be taken as
t2 = 0 and the supremum of such values of t1 must be ∞.
2. If the Maximal Solution is Bounded: Suppose that we happen to know that the maximal so-
lution φ̃ is bounded. Then because of continuity, the limits

x̃i = lim φ̃(t) i ∈ {1, 2, . . . , 2}
t→t̃i

must exist. Now if t̃1 were finite there would have to be some interval [t̃1 , t̃] of positive length
on which there were a continuous solution φb to (5.32) which passed through x̃ at t = t̃1 . This
in turn would imply that the concatanated function θ : {It0 ,x0 ) ∪ [t̃1 , t̃)} → IRn defined so that
b
θ(t) = φ̃(t), t ∈ I(t0 ,x0 ) , and θ(t) = φ(t), t ∈ [t̃1 , t̃ ), would be a continuous solution to (5.32)
passing through x0 at time t0 . But this cannot be so because t̃ > t̃1 and I(t0 ,x0 ) is the maximal
interval of existence. In other words, t̃1 cannot be finite. Similar reasoning can be used to prove
that t̃2 must be 0.
7 The inf or infimum of a subset S ⊂ IRn is the smallest number µ such that ||x|| ≤ µ for all x ∈ S.
The following theorems summarize what we’ve uncovered so far.

Theorem 2 (Local Existence) Suppose that f : IRn ×[0, ∞) → IRn is a function which is locally Lipschitz
in x and piecewise continuous in t as defined previously. For each x0 ∈ IRn and each t0 ∈ [0, ∞) there exists
a largest interval I(t0 ,x0 ) ⊂ [0, ∞) and exactly one continuous solution φ to

ẋ = f (x, t) (5.37)

on I(t0 ,x0 ) which passes through x0 at t = t0 . Moreover, the restriction of φ to any given subinterval I ⊂
I(t0 ,x0 ) containing t0 , is the only continuous function which passes through x0 as time t0 and is a solution
to (5.37) on I.

Theorem 3 ( Global Existence ) Suppose that under the conditions of Theorem 2, f is is globally Lips-
chitz in x. Then for each x0 ∈ IRn and each t0 ∈ [0, ∞), the maximal interval of existence is [0, ∞).

Linear differential equations of the form

ẋ = A(t)x + b(t)

are equations for which f = A(t)x + b(t). Since such f ’s are globally Lipschitz x, solutions to such linear
differential equations always exist on [0, ∞), assuming of course that A(t) and b(t) are piecewise continuous
matrices. As we have already seen, the other instance when I(t0 ,x0 ) = [0, ∞) for a given maximal solution
φ to the initial value problem, is when the maximal solution in question is known to be bounded wherever
it exists.

Theorem 4 ( Global Boundedness ) If, under the conditions of Theorem 2, the maximal solution to the
initial value problem which passes through x0 at t = t0 is bounded, then the maximal interval of existence is
[0, ∞).

5.6 The Concept of State

Consider again the differential equation


ẋ = f (x, t) (5.38)
where f : IRn × [0, ∞) → IRn . For the remainder of these notes, unless otherwise stated we will assume that
f (x, t) is at least locally Lipschitz Continuous in x and piecewise continuous in t.

It is sometimes useful to think of (5.38) as a dynamical system Σ whose possible “states” are the vectors
in IRn . Σ is said to be in state z at time t = τ if z is the value at time τ of the solution to (5.38) under
consideration. Σ’s state space IRn is thus the space in which solutions to (5.38) take values.

Example 36 There are dynamical systems whose state spaces have more structure than IRn . For example,
the two scalar differential equations

ẋ1 = x2 g(x1 , x2 )
ẋ2 = −x1 g(x1 , x2 )

with g any scalar-valued function, can be viewed as a dynamical system whose state space is the unit circle
C centered at 0 in IR2 . This is because the differential equations imply that
d 2
{x + x22 } = 0
dt 1
and thus that x21 + x22 = c2 , for some constant c. This means that Σ’s state vector x evolves on the unit
circle - assuming of course x started out there in the first place.

Note that the concept of state as defined here is consistent with what’s usually meant by the term in a
nontechnical context. In particular, knowledge of Σ’s state at time τ is sufficient to uniquely determine the
future behaviour of x for all t ≥ τ . This of course is because of uniqueness - i.e., there can be exactly one
maximal solution to (5.38) passing through state z at time τ.

For each z ∈ IRn and each µ ≥ 0, let ζ(·, µ, z) : I(µ,z) → IRn denote the maximal solution to (5.38)
which passes through state z at time µ. Then because of uniqueness it clearly must be true that for each
t0 ∈ [0, ∞) and each x0 ∈ IRn

ζ(t, t0 , x0 ) = ζ(t, τ, ζ(τ, t0 , x0 )), ∀ t, τ ∈ I(t0 ,x0 ) (5.39)

This is sometimes called the composition rule for dynamical systems. It says that the state of Σ at time t,
assuming x0 was Σ’s state at time t0 , is the same as Σ’s state at time t assuming Σ was in state ζ(τ, t0 , x0 )
at time τ . Thus the value of Σ’s state at time τ is sufficient to determine the value of Σ’s state at any future
time t ∈ I(t0 ,x0 ) .

5.6.1 Trajectories, Equilibrium Points and Periodic Orbits

The set of states which ζ(t, τ, z) passes through at t ranges over the interval I(µ,z) from τ onward to its
right end point is often called the trajectory of Σ which leaves state z at time τ . An important consequence
of uniqueness is that for fixed τ , there is exactly one trajectory leaving each state z ∈ IRn . Moreover,
trajectories starting at the same time but from different states cannot cross in finite time.

The simplist example of a trajectory is an “equilibrium point.” A vector z in IR at which

f (z, t) = 0, ∀t ∈ [0, ∞)

is called an equilibrium point of Σ. A given f can have no, one, a finite number or even a continuum of
equilibrium points. If z is an equilibrium point, then clearly for each t0 ≥ 0, the constant function

ζ(t, t0 , z) = z, ∀t ∈ [0, ∞)

is a maximal solution to (5.38). Thus each equilibrium point of (5.38) is a trajectory. Since trajectories
starting at the same time but in different states cannot cross in finite time, no equilibrium point can be
reached from any other state in finite time.

Example 37 The differential equation


ẋ = x2 + 1
has no equilibrium points, whereas the differential equation

ẋ = x2 + xy
ẏ = y 2 + xy

has infinitely many, namely all points for which x + y = 0. The differential equation

ẋ = x2 + 3x + 2
has two equilibrium points, namely −1 and −2. For A a n × n nonsingular matrix the equation
ẋ = Ax
has only x = 0 as an equilibium point.

An equilibrium point z of (5.38) is an attractor if trajectories starting “close” enough to z tend to z as


t → ∞. In other words, z is an attractor if there is a number ǫ > 0 such that for all t0 ≥ 0 and all x0 ∈ IRn
satisfying ||x0 − z|| ≤ ǫ,
lim ζ(t, t0 , x0 ) = z
t→∞
To determine if an equilibrium point z is attractive, one looks at the behavior of trajectories in the vicinity
of z. There are several ways to do this. One way, which is applicable in the case when f is not dependent
on t, involves examinimg the properties of
∂f (x)
∂x x=z
We will return to this topic later.

Example 38 The equilibrium point x = y = 0 of the dynamical system


ẋ = −x
ẏ = −y
is {globally} attractive because all solutions tend to zero at t → ∞. On the other hand, for any nonnegative
number λ, the equilibrium point x = y = 0 of the dynamical system
ẋ = λx
ẏ = −y
is not attractive since no trajectory starting at y 6= 0 ever tends to the zero state.

A maximal solution ζ(t, t0 , x0 ) to (5.38) is said to be periodic if it is either a constant or if there exists
a finite number T > 0 such that ζ(t + T, t0 , x0 ) = ζ(t, t0 , x0 ) for all t ∈ I(t0 ,x0 ) . The smallest value of T
for which this is true is the period of ζ. Note that if ζ(t, t0 , x0 ) is periodic, it must be bounded. Thus in
this case, because of the Global Boundednes Theorem, the right end point of I(t0 ,x0 ) must be infinity; i.e.,
periodic solutions go on forever.

Note that the trajectory of a periodic solution must be a closed curve or orbit in IRn . The reverse is also
true if f does not depend on t. In other words, if f is independent of t, then maximal solutions to (5.38),
whose trajectories are closed orbits, are periodic solutions. The reader should try to prove this by exploiting
both the time invariance of f and the uniqueness of solutions to (5.38).

Example 39 The differential equation


ẋ = 3y
ẏ = −5x
has one equilibrium point, namely x = y = 0, but it is not attractive. On the other hand all maximal
solutions are periodic. Each generates a closed orbit which is an elipse in IR2 .
5.6.2 Time-Invariant Dynamical Systems

A number of things become easier to deal with in the event that the function f in (5.38) is independent of
t. Here we briefly discuss some of them.

Consider the time-independent or autonomous differential equation

ẋ = f (x) (5.40)
∂f
where f : IRn → IRn is a function whose first partial dirivatives ∂x i
, i ∈ {1, 2, . . . , n} are continuous functions
n
on IR . As before, let ζ(t, τ, z) denote the maximal solution to (5.40) which passes through state z at t = τ .
The first consequence of time invariance is that for each t0 ≥ 0 and each x0 ∈ IRn

ζ(t, 0, z) = ζ(t + t0 , t0 , z), t ∈ I(0,z) (5.41)

In other words, the value at each time t of maximal solution starting in state z at time 0 is the same as the
value at time t + t0 of the maximal solution which passes through z at time t0 . {The reader should verify
that the time-independence of f and the uniqueness of solutions to (5.40) imply that this is so.} Because of
this property, there is no reason to not take t0 = 0. And so for the time-invaraint case we always will.

A set V ⊂ IRn is called is said to be invariant with respect to (5.40) if every maximal solution to (5.40)
which starts at a state in V , remains in V forever. In other words, for V to be an invariant set, all solutions
starting at states in V must exist for all time and moreover, each state along such a solution must lie on a
trajectory which is a subset of V .

Example 40 The set of points in IR2 for which y = 0 is an invariant subset V for the differential equation

ẋ = −x
ẏ = −y

In this case V is actually a subspace.

Equilibrium points, closed orbits, and trajectories of solutions which exist for all time are also examples
of invariant sets. A more interesting class of invariant sets can be characterized as follows.

Call a function F : IRn → IR a first integral of (5.40) if

∂F
f (x) = 0, ∀x ∈ IRn
∂x
If F is a first integral and φ is a solution to (5.40), then

d
F (φ(t)) = 0, t ≥ 0
dt
What this means is that if F is a first integral, then F (x) is constant along trajectories in IRn .

If F is a first integral, then for each constant c, the set



Sc = {x : F (x) = c, x ∈ IRn }

is called an integral manifold. Note that any such integral manifold is an invariant set provided the maximal
solutions which generate the trajectories in the manifold exist for all time.

Example 41 The function F (x, y) = x2 + y 2 is a first integral for

ẋ = y
ẏ = −x

The integral manifold determined by


F (x, y) = 2
is a circle of radius 2 in IR2 .

5.7 Simulation

As we have already seen, it is possible to use Picard Iteration to compute an approximate solution to an
initial value problem on a prespecified interval of existence [t0 , t1 ]. For this, one needs a good numerical
procedure {e.g., Simpson’ Rule} for calculating the integral, as a function of t ∈ [t0 , t1 ], of a function defined
on [t0 , t1 ]. Such procedures are readily available and can be found in most numerical software packages; in
Matlab one could use QUAD or one of its variants.

It turns out that Picard Iteration is usually not the best way to go about computing a solution to an
initial value problem8 . One important reason for this is that Picard Iteration requires one to specify t1 prior
to the initiation of the computation. For many problems of interest, the length of time over which one seeks
a solution in open-ended and is only decided upon by studying the behavior of the solution as it evolves
in time. An alternative approach which avoids this limitation is to compute an approximate solution by
iterating with respect to the independent variable time. The basic idea is quite simple and can be illustrated
with x a scalar.

5.7.1 Numerical Techniques

Suppose f : IR × [0, ∞) → IR is a given scalar valued function and it is desired to compute a solution to

ẋ = f (x, t) (5.42)

with initial state x0 at t = 0. Let us pick a small positive number ∆ called a step size. Suppose that φ(t) is
the exact solution to (5.42) with initial value φ(0) = x0 . What we want to do is to develop a procdure for
computing numbers y(1), y(2) y(3), . . . such that

y(i) ≈ φ(∆i), i ≥ 0

To start things off correctly, let us define



y(0) = x0

We approach the problem of computing the y(i), i ≥ 1 in much the same way as one approaches a proof by
induction. We presume that for some j ≥ 0, we’ve already computed an approximation

y(j) ≈ φ(∆j) (5.43)


8 On the other hand, for “two-point” boundary value problems where some of the components of x are given at t while
0
others are given at t1 , Picard Iteration or something similar may be be preferable.
and we seek a procedure for computing y(j + 1). There are a great many different ways to go about doing
this. Perhaps the most straight forward is “Euler’s Method” which is a simple procedure based on a one
term Taylor series expansion of φ(∆(j + 1)) about the time t = ∆j. In particular φ(∆(j + 1)) admits the
expansion
φ(∆(j + 1)) = φ(∆j) + ∆φ̇(∆j) + Oj (∆)
where Oj (∆) is a sum of higher power terms in ∆ with the property that
O(∆)j
lim =0
∆→0 ∆
Since φ satisfies
φ̇(t) = f (φ(t), t) (5.44)
we can also write
φ(∆(j + 1)) = φ(∆j) + ∆f (φ(∆j), ∆j) + Oj (∆)
This and (5.43) suggest that a reasonable way to try to approximate φ(∆(j + 1)) with y(j + 1) might be to
use the formula
y(j + 1) = y(j) + ∆f (y(j), ∆j)
This equation can be recursively solved thereby a generating a sequence y(0), y(1), y(2), . . . which in some
sense approximates φ(0), φ(1), φ(2), . . .. This particular approximation is called Euler’s method. As a
practical matter it is too rudimentary to work very well, except in very special situations.

There are very much better techniques based on more sophisticated approximations. Note that Euler’s
Method really stems from the approximation
φ(∆ + t) − φ(t)
≈ φ̇(t) (5.45)

Another equally valid approximation of this type is
φ(t) − φ(t + ∆)
≈ φ̇(t + ∆)
−∆

Subtracting this from (5.45) and multiplying through by 2 yields

φ(t + ∆) − φ(t) ≈ {φ̇(t) + φ̇(t + ∆)}
2
Replacing φ̇(t) and φ̇(t + ∆) by f (φ(t), t) and f (φ(t + ∆), t + ∆) thus provides

φ(t + ∆) − φ(t) ≈ {f (φ(t), t) + f (φ(t + ∆), t + ∆)} (5.46)
2
Using the Euler approximation

φ(t + ∆) ≈ φ(t) + ∆φ̇(t) = φ(t) + ∆f (φ(t), t)

one can write


f [φ(t + ∆), t + ∆] ≈ f [φ(t) + ∆f (φ(t), t), t + ∆]
Using this approximation in (5.46) we obtain

φ(t + ∆) − φ(t) ≈ {f (φ(t), t) + f [φ(t) + ∆f (φ(t), t), t + ∆]}
2
This is the basis for the recursion algorithm

y(i + 1) = y(i) + {f (y(i), ∆i) + f [y(i) + ∆f (y(i), ∆i), t + ∆]}
2
which is one form of the Runge-Kutta method. A more elaborate version of the Runge-Kutta method is
described by the equation


y(i + 1) = y(i) + {z1 (i) + 2z2 (i) + 2z3 (i) + z4 (i)}
6
where
z1 (i) = f [y(i), ∆i] z3 (i) = f [y(i) + ∆z2 (i), ∆(i + 12 )]

z2 (i) = f [y(i) + ∆z1 (i), ∆(i + 21 )] z4 (i) = f [y(i) + 2∆z3 (i), ∆(i + 1)]
This particular version of the Runge-Kutta Method is widely used in many software packages including
Matlab.

Much is known about recursive techniques such as the ones we’ve been discussing. Often a variable step
size ∆(i) is used to control the accumulation of errors. A course on numerical anaylsis would typically deal
with this topic in depth. We will not pursue the subject further here.

5.7.2 Analog Simulation

An alternative way to deal with initial value problems is via “simulation” on an “analog computer.” Analog
computers were widely for simulation years ago {especially in real-time applications} prior to the development
modern digital computers. Although still in use, analog computers are at present much less popular than
digital computers for two reasons:

1. Analog computers are limited by the accuracy of their components {e.g., resisters and capacitors} and
for this reason cannot usually match the precision obtainable with digital computers.

2. Because of the limited dynamic ranges of their components, electronic analog computers a much more
difficult to program and to keep from saturating than their digital conterparts.

Dispite these shortcomings analog computation is worth knowing about for several reasons:

1. Analog computers are inherently parallel machines. Partly for this reason, analog computers are often
faster than their digital counterparts, all other things being equal.

2. The detailed study of analog vlsi chips is currectly begin pursued and there is some reason to believe
that analog techniques will re-emerge as the preferred approach for dealing with certain types of
problems, especially those for which speed is of much greater concern than accuracy.

3. Perhaps the most important reason for understanding the analog approach is because it provides an
intuitively appealing conceptual framework for visualizing how a computing machine might go about
solving a complicated system of ordinary differential equations. As we shall see, this framework is
especially useful even if the the machine is digital.

An analog computer is an electronic device composed primariy of precision resistors, precision capacitors,
high gain linear {i.e., operational} amplifiers, and diodes. Solutions to differential equations are represented
on an analog computer as time-dependent voltage signals. Interconnections of an analog computer’s basic
components yields higher level analog building blocks called summers, inverters, integrators, multipliers,
gains and nonlinear function generators. Each of these building blocks performs the function corresponding
to its name. For example, an integrator acts on its input signal x(t) by generating an output signal y(t) which
is the real-time integral of x(t). Analog computers are capable of performing simultaneously many different
time integrations, multiplications, summations, etc. In other words, an analog computer is an inherently
parallel machine.

An analog computer is programmed by physically interconnecting the inputs and outputs of its higher level
building blocks thereby creating an analog program. A simulation is initiated by turning on the integrators
and letting the computer run. The time evolutions of signals are monitored using standard analog displays
including chart recorders and scopes.

In the sequel we will explain how an analog computer is programmed. We will do this using the building
block symbols found in Matlab’s Simulink Package since these are the symbols normally used in an analog
computer programming diagram and since Simulink is programmed as if it were the front end of an analog
computer. Of course Simulink is not truely the front end of an analog computer. Rather it is an especially
useful graphical interface for programming a digital computer to solve ordinary differential equations. Once
Simulink is programmed in an analog fashion, what the computer actually does is to translate the program
into digital c code and and then execute it serially using standard numerical algorithms such as the Runge-
Kutta method. From the users perspective, the computer appears to be doing an analog simulation even
though this is not really what’s happening.

5.7.3 Programming Diagrams

To access Simulink within Matlab, one enters the command simulink in the Command Window. In response
there appears a mouse-driven window called Simulink which looks like this.

Sources Sinks Discrete Linear Nonlinear Connections Extras

SIMULINK Block Library (Version 1.3c)

As a first step one would typically open a new window by selecting new under the file menu. Double clicking
on any of the above windows,{e.g., Sources} opens it. Any item can then be dragged into the window new
opened previously. Items in columns 1 through 4 below come from windows Sources, Sinks, Linear and
Nonlinear respectively. The one exception to this is the Unit Delay in the lower right hand corner which
comes from the Discrete window.

+ *
+
Sum Product
Signal Scope
Generator
sin(u[1])
1/s Fcn
Sine Wave Integrator
Graph

1 1/z
Unit Delay
Step Input Gain
XY Graph
These are some of the most important building blocks in Simulink for solving ordinary differential equations.
Each building block does just what you’d expect:

Signal Generator: Generates a time signal whose specifics are set by opening the Signal Generator window.

Sine Wave: Generates a sine-wave time signal whose amplitude and frequency are set by opening the
Sine-Wave window.

Step Input: Generates a step time signal whose height and time of occurence are set by opening the Step
Input window.

Scope: Plots its input signal(s) verses time.

Graph: Plots its input signals(s) on one graph verses time.


XY Graph: Plots one input signal on a graph verses the other - good for phase plane trajectory plots.

Sum: Generates an output time signal which is the sum of its input signals, each input being weighted by
a specified +1 or −1; handles more than two inputs.

Integrator: Generates an output signal which is the time integral of its input plus a specified initial con-
dition constant.

Gain: Generates an output signal which is its input signal times a specified gain constant.

Product: Generates an output signal which is the product of its input signals.
Fcn: Generates an output signal which is any specified function of its input signal.

Unit Delay: Generates an output signal which is the input signal delayed by a prespecified amount of time;
used in place of an integrator for a discrete time simulation.

Arrows connecting building blocks can be drawn by holding down the mouse button and dragging within
the new window

Example 42 The differential equation


4.5
ÿ + .5ẏ + y − = u̇ + u
y2
can be represented by the state space equations

ẋ1 = −.5x1 + x2 + u
4.5
ẋ2 = −x1 + 2 + u
x1

where
y = x1
Note, by the way, that u̇ does not appear in the state space model, even though it appears in the original
differential equation. This would be especially important if u were a discontinuous input like a step. Figure
42 shows a Simulink simulation diagram for this system with x1 and x2 being the signals coming out of
integrators 1 and 2 respectively. Note that −.5x1 + x2 + u is the the sum of the signals going into integrator
1 and −x1 + 4.5x21
+ u is the corresponding sum going into integrator 2. As a rule, simulation diagrams for
state space models can always be set up with the outputs of the integrators being state variables. Thus an
n-dimensional state space model requires n-integrators for simluation.
+ +
- 1/s + 1/s
Step Input + -
int2 int1
sum2 sum1 Graph
.5

Scope gain2

4.5 u^(-2)

gain1 Fcn

Below is a graph of y for a simulation in which x1 (0) = .5, x2 (0) = 0. and u is five times a unit step applied
at t = 15.
10

0
0 5 10 15 20 25 30 35
Time (second)
Chapter 6

Linear Differential Equations

By an n - dimensional linear state differential equation is meant an equation of the form

ẋ = A(t)x + b(t) (6.1)

where A(t) is a n × n matrix and x and b (t) are n - vectors. If A(t) is constant, (6.1) is constant coefficient;
if A(t) is constant and t is time, we sometimes say that (6.1) is time invariant or stationary or autonomous.
The vector b(t) is often called a forcing function; in the event that it is the zero function, (6.1) is called
unforced or homogeneous.

We will assume throughout that A(t) and b(t) are at least piecewise continuous functions. Then f (x, t) =
A(t)x + b(t) is globally Lipschitz in x and piecewise continuous is t. In view of our earlier findings, this
means that for each initial time t0 ∈ [0, ∞) and initial state x0 ∈ IRn , there is exactly one solution φ to
(6.1) on [0, ∞) which passes through state x0 at t = t0 . Sometimes we will simply write x instead of φ.

The aim of this chapter is to study the basic properties of such equations. We begin by explaining one
of the most common ways in which linear differential equations arise.

6.1 Linearization

More often than not when modeling a physical process, one encounters ordinary differntial equations which
are not linear and which have “inputs”. Suppose that

ẋ = f (x, t, u) (6.2)

is such an equation where f : IRn × [0, t) × IRm → IRn is a continuously differentiable function of x and
a piecewise continuous function of t and a continuously differentiable function of u. Suppose in addition
that for some given initial time t0 and initial state x0 and piecewise continuous function v : [0, ∞) → IRn )

called an input, that φ is the solution to (6.2) {assuming u = v}, which passes through x0 at t = t0 . In
many practical situations it is important to have a good description of how a this solution might change in
response to a small change in x0 or a small change in the input function u. If the presumed changes are
small enough, then their effect can be studied by using a linear model derived from (6.2). The process of
constructing this model is called linearization and is as follows.

Let us suppose that y0 ∈ IRn is any vector with a “small” norm and that w : [0, ∞) → IRm is any
piecewise continuous function whose norm ||w(t)|| is also “small” for each fixed t ∈ [0, ∞). Let θ denote a

87

solution to (6.2) {assuming u = v + w} which passes through x0 + y0 at t = t0 . What we want to do is to
derive a linear differential equation whose solution approximates the difference θ − φ. Now what we know is
that

φ̇(t) = f (φ(t), t, v(t)) (6.3)


φ(t0 ) = x0 (6.4)
θ̇(t) = f (θ(t), t, v(t) + w(t)) (6.5)
θ(t0 ) = x0 + y0 (6.6)

If for each fixed t we now expand the right side of (6.5) in a Taylor Series about the point (φ(t), v(t)) we get

f (θ(t), t, v(t) + w(t)) = f (φ(t), t, v(t)) + A(t){θ(t) − φ(t)} + B(t)w(t)


+ higher order terms in {θ(t) − φ(t)} and w(t)

where

∆ ∂f (x, t, u) ∆ ∂f (x, t, u)
A(t) = and B(t) =
∂x (x,t,u)=(φ(t),t,v(t)) ∂u (x,t,u)=(φ(t),t,v(t))

This and (6.3) imply that

d
{θ(t) − φ(t)} = A(t){θ(t) − φ(t)} + B(t)w(t)
dt
+ higher order terms in {θ(t) − φ(t)} and w(t)

This suggests that by solving the linear differential equation

ẏ = A(t)y + B(t)w(t) (6.7)

with initial state y0 at time t = t0 , one should get a solution ψ(t) which to first approximation satisfies

ψ ≈θ−φ

The point here is that (6.7) is a linear differential equation and consequently an equation amenable to
anaylsis. We call (6.7) the linearization of (6.3) about the nominal solution φ determined by nominal initial
condition x0 and nominal input v(t).

One of the most common applications of the linearization process arises in situations when f (x, t, u) does
not depend on t {i.e., f (x, t, u) = f¯(x, u)} and when the nominal input v(t) and nominal solution φ(t) are
both constant. Note that for v(t) a constant v̄, φ(t) will be the constant solution φ(t) = x0 provided

f¯(x0 , v̄) = 0

Note also that under these conditions both A(t) and B(t) will be constant matrices.

Example 43 Consider a satellite of mass m in orbit about the earth as shown in the following diagram.
'$

@
&%@
9 @ r
θ @
R wm
@
@
@
R
@
u2 u1

Assume that the satellite can be thrusted in the radial direction with thrust u1 and in the tangential direction
with thrust u2 . In polar coordinates the satellite’s equations of motion are

k
r̈ = θ̇2 r − + u1
r2
ṙθ̇ u2
θ̈ = −2 +
r r

where θ̇2 r and rk2 model the effects of centrifugal force and gravity respectively. To develop an equivalent
state space model, define
∆ ∆ ∆ ∆
x1 = r, x2 = ṙ, x3 = θ, x4 = θ̇

Then

ẋ1 = x2
k
ẋ2 = x24 x1 − + u1
x21
ẋ3 = x4
x2 x4 u2
ẋ4 = −2 +
x1 x1
and

r = x1
θ = x3

Let us note that for u1 = u2 = 0, these equations admit solutions of the form

x1 = r0 , x2 = 0, x3 = ω0 t, x4 = ω0

where r0 and ω0 are positive constants satisfying ω02 r03 = k. Any such solution describes a circular orbit of
radius r0 with angular velocity ω0 . Our aim is to linearize the above equations about such an orbit. In other
words we want to develop a system of linear differential equations for deviation variables z1 , z2 , z3 , z4 such
that any solution x1 , x2 , x3 , x4 to the above equations satisfies

x1 ≈ r0 + z1 , x2 ≈ z2 , x3 ≈ ω0 t + z3 , x4 ≈ ω0 + z4
assuming u1 and u2 are small thrust signals and the initial values of x1 , x2 , x3 , and x4 are close to r0 , 0, ω0 t
and ω0 respectively. Carrying out the linearization one gets the equations

ż1 = z2
 
2 k
ż2 = ω0 + 3 3 z1 + 2ω0 r0 z4 + u1
r0
ż3 = z4
ω0 1
ż4 = −2 z2 + u2
r0 r0

which can be written more concisely as


ż = Az + Bu

where
     
z1 0  1 0 0 0 0  
∆  2ω0 r0 
2 k
∆ z  ω + 3 0 0 ∆ 1 0  ∆ u1
z =  2 , A= 0 ,
3
r0
 B= , u=
z3 0 0 0 1  0 0 u2
1
z4 0 −2 ωr00 0 0 0 r0

Note that for this example, the linearized equations are time-invariant even though the solution about which
the linearization is made is not.

6.2 Linear Homogenous Differential Equations

In this section we shall investigate the fundamential properties of the linear, homogeneous, n-dimensional
differential equation
ẋ = A(t)x (6.8)

where A(t) is a n × n matrix of piecewise continuous time functions defined on the interval [0, ∞). Let us
note right away that for each finite interval [t0 , t1 ] ⊂ [0, ∞) we can write a Lipschitz inequality of the form

||A(t)x − A(t)y|| ≤ λ||x − y||, x, y ∈ IRn

where

λ = sup ||A(t)|| < ∞
t∈[t0 ,t1 ]

Thus solutions to (6.8) exist globally and are unique.

Theorem 5 Suppose that A : [0, ∞) → IRn×n is piecewise continuous. For each x0 ∈ IRn and each
t0 ∈ [0, ∞) there exists exactly on continuous function φ which satisfies (6.8) on [0, ∞) and passes through
the state x0 at t = t0 .
6.2.1 State Transition Matrix

In the sequel we shall write ζ(t, τ, z) for the unique solution to (6.8) which passes through state z at t = τ .
Thus ζ : [0, ∞) × [0, ∞) × IRn → IRn is a continuous function uniquely defined by the equations1

ζ̇(t, τ, z) = A(t)ζ(t, τ, z) 
∀t, τ ∈ [0, ∞), ∀z ∈ IRn (6.9)

ζ(τ, τ, z) = z

Let ei , i ∈ {1, 2, . . . , n} denote the ith unit vector in IRn . The state transition matrix2 of A(t) is the n × n
matrix

Φ(t, τ ) = [ ζ(t, τ, e1 ) ζ(t, τ, e2 ) . . . ζ(t, τ, en ) ] (6.10)
The definition implies that

Φ(τ, τ ) = [ ζ(τ, τ, e1 ) ζ(τ, τ, e2 ) . . . ζ(τ, τ, en ) ] = [ e1 e2 . . . en ] = In×n

and also that

Φ̇(t, τ ) = [ ζ̇(t, τ, e1 ) ζ̇(t, τ, e2 ) . . . ζ̇(t, τ, en ) ]

= [ A(t)ζ(t, τ, e1 ) A(t)ζ(t, τ, e2 ) . . . A(t)ζ(t, τ, en ) ]

= A(t) [ ζ(t, τ, e1 ) ζ(t, τ, e2 ) . . . ζ(t, τ, en ) ]

= A(t)Φ(t, τ )

In other words, the definition of Φ(t, τ ) in (6.10) implies that



Φ̇(t, τ ) = A(t)Φ(t, τ ) 
∀t, τ ∈ [0, ∞) (6.11)

Φ(τ, τ ) = In×n

But because of uniqueness of solutions to linear differential equations, there can be only one matrix possessing
these two properties. Thus (6.11) also implies (6.10). Because of this, (6.11) an be taken as an alterative
definition for the state transition matrix of A(t). It is entirely equivalent to the original definition of Φ(t, τ )
given by (6.10).

The importance of the state transition matrix stems from the fact that

ζ(t, τ, z) = Φ(t, τ )z, ∀t, τ ∈ [0, ∞), ∀z ∈ IRn (6.12)


To verify that this is so, fix τ and z and set φ(t) = Φ(t, τ )z; then because of (6.11), φ must satisfy (6.8) and
φ(τ ) = z. Since these equations are satisifed by just one continuous function, namely ζ(t, τ, z), it must be
that (6.12) is true.

Let us note that if the dynamical system (6.8) is in state x1 at time t1 and state x2 at time t2 then the
“transition” from x1 to x2 is described by the equation

x2 = Φ(t2 , t1 )x1
1 Here as always, “·” denotes differentiation with respect to t.
2 In most mathematics texts, the state transition matrix of A(t) is called the fundamental matrix of A(t). The name “state
transition matrix” is commonly used in many fields. We’ve adopted it here because state transition is a more descriptive
modifier than fundamental.
Hence the name “state transition matrix.”

Since Φ(t, τ ) is a continuous function of t, (6.11) is equivalent to the integral equation


Z t
Φ(t, τ ) = I + A(µ)Φ(µ, τ )dµ ∀t, τ ∈ [0, ∞) (6.13)
τ

This equation can be solved by Picard Iteration. In particular the sequence of matrices Θ0 (t, τ ), Θ1 (t, τ ),
Θ2 (t, τ ), . . ., defined for all t, τ ∈ [0, ∞) by
Z t

Θ0 (t, τ ) = I and Θi+1 (t, τ ) = I + A(µ)Θi (µ, τ )dµ, i ∈ {0, 1, . . .}, (6.14)
τ

converges uniformly on every finite interval [0, t1 ] and in the limit

Φ(t, τ ) = lim Θi (t, τ ) ∀t, τ ∈ [0, ∞) (6.15)


i→∞

6.2.2 The Matrix Exponential

It is especially useful at this point to consider briefly what the Θi (t, τ ) look like in the special case when
A(t) is a constant matrix B. If this is so, then because of (6.14)
Z t
Θ1 (t, τ ) = I + Bdµ = I + (t − τ )B
τ

By similar reasoning
Z t
1
Θ2 (t, τ ) = I + (I + µB)dµ = I + (t − τ )B + {(t − τ )B}2
τ 2
Continuing this process, one eventually arrives at the formula

Xi
1
Θi (t, τ ) = {(t − τ )B}j
j=0
j!

In view of (6.15), the series


Xi
1
{(t − τ )B}j
j=0
j!

converges to Φ(t, τ ) as i → ∞ for every finite value of t and τ in [0, ∞) In other words

Φ(t, τ ) = e(t−τ )B , t, τ ∈ [0, ∞) (6.16)

where for any n × n matrix M , the symbol eM stands for the the n × n matrix exponential

X∞
∆ 1 j
eM = M (6.17)
j=0
j!

Note that this series must converge because, as we already know, the series
X∞
1
{(t − τ )B}j
j=0
j!
converges to Φ(t, τ ) for any constant matrix B, and any finite values of t and τ in [0, ∞) - to complete the
proof just take B = M and t − τ = 1.

From the fact that Φ(t, τ ) satisfies (6.11) we arrive at the formulas

d
 tB 
dt e = BetB 
∀t ∈ [0, ∞) (6.18)

etB t=0
= In×n

In these ways, the matrix exponential behaves much like a scalar exponential.

A little latter we will explain how to explicitly calculate e(t−τ )B . For the present, we conclude this
excursion into the study of the case when A(t) is a constant, with a few words of caution: The expression
for Φ(t, τ ) in (6.16) holds only if A(t) is a constant matrix B. If A(t) is nonconstant, it is not even true
that Rt
Φ(t, τ ) = e τ A(µ)dµ
R  R 
t t
except in the very special case when A(t) τ A(µ)dµ = τ A(µ)dµ A(t), ∀t, τ ≥ 0.

6.2.3 State Transition Matrix Properties

We now return to the general case when A(t) is any n × n piecewise continuous matrix. Our aim in this
section is to derive several basis properties of A’s state transition matrix Φ(t, τ ).

Exponential Growth for Bounded A(t)

To begin let us note that (6.13) enables us to write for t ≥ τ ≥ 0


Z t Z t
||Φ(t, τ )|| = I + A(µ)Φ(µ, τ )dµ ≤ ||I|| + ||A(µ)Φ(µ, τ )||dµ
τ τ

Thus Z t
||Φ(t, τ )|| ≤ ||I|| + ||A(µ)||||Φ(µ, τ )||dµ
τ

By the Bellman-Gronwall Lemma, Rt


||A(µ)||dµ
||Φ(t, τ )|| ≤ ||I||e τ (6.19)
Suppose A(t) is bounded in norm on [0, ∞); i.e. suppose

λ= sup ||A(t)|| < ∞
t∈[0, ∞)

Then (6.19) simplifies to


||Φ(t, τ )|| ≤ ||I||e(t−τ )λ
Similar reason leads to the inequality

||ζ(t, τ, z)|| ≤ ||z||e(t−τ )λ

where ζ(t, τ, z) is the solution to (6.8) passing through state z at t = τ . These inequalities show that if A(t)
is bounded on [0, ∞), then the norms of Φ(t, τ ) and all solutions to (6.8) can grow at most exponentially
fast as t → ∞.
Composition Rule and Consequences

As noted in the last chapter, the composition rule for a dynamical system ẋ = f (x, t) is
ζ(t, τ, z) = ζ(t, µ, ζ(µ, τ, z)), ∀z ∈ IRn , ∀t, τ, µ ∈ [0, ∞)
where, as usual, ζ(t, τ, z) denotes the the solution to ẋ = f (x, t) which passes through z at t = τ . For the
linear differential equation ẋ = A(t)x, we know because of (6.12) that
ζ(t, τ, z) = Φ(t, τ )z, ζ(t, µ, ζ(µ, τ, z)) = Φ(t, µ)ζ(µ, τ, z), and ζ(µ, τ, z) = Φ(µ, τ )z
It thus follows that
Φ(t, τ )z = Φ(t, µ)Φ(µ, τ )z, ∀z ∈ IRn , ∀t, τ, µ ∈ [0, ∞) (6.20)
From this it follows that
Φ(t, τ ) = Φ(t, µ)Φ(µ, τ ), ∀t, τ, µ ∈ [0, ∞) (6.21)
This is the composition rule for state transition matrices3 .

Since the composition rule holds for all values of t, τ, µ ∈ [0, ∞), there is nothing to prevent us from
setting τ = t. This results in I = Φ(t, µ)Φ(µ, t) and thus the remarkable fact that

Φ−1 (t, µ) = Φ(µ, t), ∀ t, µ ∈ [0, ∞) (6.22)

In other words, the inverse of Φ(t, µ) is obtainable by simply interchanging its arguments.

Note that we have incidently proved that Φ(t, µ) is nonsingular for all finite values of t and µ. A little
thought reveals that this is in essence a restatement of the fact that trajectories starting at the same time
but in different states, cannot cross in finite time.

Properties of Matrix Exponentials

Let us again briefly return to the case when A(t) is a constant matrix B. Recall that when this is so Φ(t, τ )
is the matrix exponential
Φ(t, τ ) = e(t−τ )B
In this case (6.21) and (6.22) simplify to
e(t−τ )B = e(t−µ)B e(µ−τ )B , ∀t, τ, µ ∈ [0, ∞)
and  −1
e(t−µ)B = e−(t−µ)B , ∀ t, µ ∈ [0, ∞)
respectively. The latter formula proves that for any n × n matrix M
−1
eM = e−M (6.23)

whereas the former showns that


e(t1 +t2 )M = et1 M et2 M , ∀ t1 , t2 ∈ IR (6.24)
To this extent, matrix exponentials are just like scalar exponentials. Note however that if F and G are two
n × n matrices, then in general
e(F +G) 6= eF eG
except in the special case when F and G commute.
3 In arriving at (6.21) from (6.20) we’ve used the algebraic fact that if C and D are two m × n matrices such that Cy =

Dy, ∀y ∈ IRn then C = D. This in turn can be proved by noting that the hypothesis Cy = Dy, ∀y ∈ IRn implies the matrix
equation [ Cy1 Cy2 · · · Cyn ] = [ Dy1 Dy2 · · · Dyn ] for all choices of the yi ∈ IRn . Picking yi to be the ith unit vector
in IRn then yields the desired result.
A Formula for det Φ(t, τ )

In the sequel we shall derive an interesting formula for the determinant of Φ(t, τ ) in terms of the integral of
the trace4 of A(t). To do this we need first to develop a formula for the derivative with respect to t of the
determinant of Φ(t, τ ). In the sequel φij denotes the ijth entry of Φ.

To begin we note that since det Φ is a function of its n2 entries φij ,

XX n n  
d ∂{det Φ}
{det Φ(t, τ )} = φ̇ij (6.25)
dt i=1 j=1
∂φij

Next recall that det Φ can be written as


n
X
det Φ = φ̄ij φ
i=1

where φ̄ is the cofactor of φij . Since φ̄ij does not depend on φij it must be true that

∂ det Φ
= φ̄ij
∂φij

From this and (6.25) it follows that

XX n n
d
{det Φ(t, τ )} = φ̄ij φ̇ij (6.26)
dt i=1 j=1


Let C denote the cofactor matrix C = [ φ̄ij ]n×n . A simple calculation shows that
n
X
φ̄ij φ̇ij
i=1

is jth diagonal element of the product C ′ Φ̇. Therefore


n X
X n
trace C ′ Φ̇ = φ̄ij φ̇ij
j=1 i=1

From this and (6.26) it follows that


d
{det Φ} = trace C ′ Φ̇
dt
Using the relation Φ̇ = AΦ we therefore have
d
{det Φ} = trace {C ′ AΦ}
dt
Thus5
d
{det Φ} = trace {AΦC ′ } (6.27)
dt
To proceed, recall that the inversion formula for Φ is
1
Φ−1 = C′
det Φ
∆ P
4 The trace of a square matrix is the sum of its diagonal elements; i.e. trace [ qij ]n×n = n i=1 qii .
5 Here we are using the identity trace {M N } = trace {N M } which holds for any pair of n × n matrices M and N .
Hence ΦC ′ = (det Φ) I. Substitution into (6.27) thus yields the one - dimensional linear differential equation

d
{det Φ(t, τ )} = trace {A(t)}{det Φ(t, τ )}
dt
Therefore Rt
trace A(µ)dµ
det Φ(t, τ ) = {det Φ(τ, τ )} e τ

Since Φ(τ, τ ) = I, we arrive at the expression


Rt
trace A(µ)dµ
det Φ(t, τ ) = e τ

which is sometimes called the Abel, Jacobi, Liouville Formula.

6.3 Forced Linear Differential Equations

We now turn to the forced linear differential equation

ẋ + A(t)x + b(t) (6.28)

Our main objective is to develop a formula which expresses this equation’s solution in terms of b and the
value of x at some specified time t0 . In doing this, we will make use of the state transition matrix of A(t)
defined in the last section.

6.3.1 Variation of Constants Formula

Let us begin by supposing that x satisfies (6.28) and that Φ(t, τ ) is the state transition matrix of A(t). Then
as noted previously 
Φ̇(t, τ ) = A(t)Φ(t, τ ) 
∀t, τ ∈ [0, ∞) (6.29)

Φ(τ, τ ) = In×n
What we want to do next is to make a change of variable which replaces x by a new state vector y whose
differential equation is easier to solve than (6.28). With the benefit of considerable hindsight we shall define
y to be


y(t) = Φ(t0 , t)x(t) (6.30)
where t0 is a fixed number in [0, ∞) The definition implies that

y(t0 ) = x(t0 ) (6.31)

because of (6.29) and that


x(t) = Φ(t, t0 )y(t) (6.32)
because of (6.22). Differentiating both sides of(6.32) we get

ẋ = Φ̇(t, t0 )y + Φ(t, t0 )ẏ

From (6.29) it thus follows that


ẋ = A(t)Φ(t, t0 )y + Φ(t, t0 )ẏ (6.33)
Meanwhile from (6.28) and (6.32) we have that

ẋ = A(t)Φ(t, t0 )y + b(t)

Substituting this into (6.33) to eliminate ẋ we get

A(t)Φ(t, t0 )y + b(t) = A(t)Φ(t, t0 )y + Φ(t, t0 )ẏ

After cancellation and rearrangment of terms this simplifies to

Φ(t, t0 )ẏ = b(t)

In view of (6.22) we can therefore write


ẏ = Φ(t0 , t)b(t)
Therefore Z t
y(t) = y(t0 ) + Φ(t0 , µ)b(µ)dµ, t ∈ [0, ∞)
t0

Multiplying bothe sides of this equation by Φ(t, t0 ) gives


Z t
Φ(t, t0 )y(t) = Φ(t, t0 )y(t0 ) + Φ(t, t0 )Φ(t0 , µ)b(µ)dµ, t ∈ [0, ∞)
t0

Using (6.32), (6.31) and the composition rule Φ(t, t0 )Φ(t0 , µ) = Φ(t, µ) we arrive at the equation
Z t
x(t) = Φ(t, t0 )x(t0 ) + Φ(t, µ)b(µ)dµ, t ∈ [0, ∞) (6.34)
t0

which is known as the variation of constants formula.

In the special case when A(t) is a constant matrix B, Φ(t, τ ) is the matrix exponential

Φ(t, τ ) = e(t−τ )B

Thus in the case, the variation of constants formula becomes

Z t
x(t) = e(t−t0 )B x(t0 ) + e(t−µ)B b(µ)dµ, t ∈ [0, ∞) (6.35)
t0

It is interesting to note that the integral in this expression is of the form


Z
W (t − µ)h(µ)dµ

Such integral forms arise in many contexts. They are called convolution integrals.

6.4 Periodic Homegeneous Differential Equations

In the light of the discussions in the preceding sections it is clear that the state transition matrix of A(t)
plays a central role in characterizing solutions to the differential equation ẋ = A(t)x + b(t). There remains
the problem of determining just what Φ(t, τ ) is for a given A(t). Unfortunately for general A’s it is more
or less impossible to derive explicit analytical expressions for Φ(t, τ ) in terms of (say) elementary functions.
Even differential equations as simple looking as the famous Mathieu Equation {Here ǫ is a positive number.}

ÿ + (1 + 2ǫ cos 2t)y = 0 (6.36)

defy closed form solutions. In fact, constant A’s are really the only general types of matrices for which
closed-form expressions for state transition matrices can actually be worked out.

Between general A’s and the class of constant A’s lies the intermediate class of A’s which are periodic; i.
e., A’s satisfying
A(t + T ) = A(t), t ≥ 0 (6.37)
where T is a fixed nonnegative number. While it is still not possible to give a complete closed-form description
for the state transition matrix of such an A, it is possible to go further than in the general case when A is
an arbitrary nonconstant matrix. The aim of this section is to explain just how much further.

6.4.1 Coordinate Transformations

Consider the differential equation


ẋ = A(t)x (6.38)
where A(t) is a continuous n × n matrix. For the moment, suppose A is otherwise arbitrary; i.e. not
necessarily periodic. Let Q(t) be a n × n matrix which is nonsingular and continuously differentiable on
[0, ∞). Suppose we define

z = Q(t)x (6.39)
Our aim is to derive a linear homeogeneous differential equation for z. For this, let us note that

ż = Q̇(t)x + Q(t)ẋ = Q̇(t)x + Q(t)A(t)x

In view of (6.39), it must be true that

ż = (Q̇(t)Q−1 (t) + Q(t)A(t)Q−1 (t))z (6.40)

Thus the change of variables


x 7−→ Q(t)x
induces a corresponding change in A(t), namely

A(t) 7−→ Q̇(t)Q−1 (t) + Q(t)A(t)Q−1 (t)

{Note, by the way, that if Q were constant, Q̇ would be zero and the preceding would be a similarity
transformation.}

Now if we’re lucky, Q̇(t)Q−1 (t) + Q(t)A(t)Q−1 (t) might turn to be a constant matrix B in which case
(6.40) would be the constant coefficient differential equation

ż = Bz (6.41)

Actually, in this instance we can control our own luck and even more: we can always find a matrix Q(t)
which is nonsingular on [0, ∞) such that

Q̇(t)Q−1 (t) + Q(t)A(t)Q−1 (t) = B (6.42)


In fact, we can find such a Q(t) for any given n × n matrix B. All we need to do is to solve the linear
{matrix} differential equation6
Q̇ + QA(t) = BQ (6.43)
with the initial value
Q(0) = I (6.44)
For if (6.43) holds and Q(t) is nonsingular on [0, ∞), then (6.42) clearly must hold as well. Therefore to
justify the claim, all we need do is show that the solution to (6.43) and (6.44) is nonsingular on [0, ∞).
What we will actually show, is something a little stronger, namely that the solution to (6.43) and (6.44) is

Q(t) = etB Φ−1 (t, 0) (6.45)

where Φ(t, τ ) is the state transition matrix of A(t).

To justify (6.45), define



P (t) = etB Φ−1 (t, 0)
Then
P (0) = I (6.46)
and
P (t)Φ(t, 0) = etB
Thus
Ṗ Φ + P Φ̇ = BetB
Since Φ̇ = AΦ, it must be true that
Ṗ Φ + P AΦ = BetB
But Φ(t, 0) is nonsingular for all finite t so

Ṗ + P A = BetB Φ(t, 0)−1 = BP

Since this is the same differential equation as (6.43) and since Q(0) and P (0) are both equal the identity
matrix, by uniqueness of solutions to linear differential equations, Q and P must be equal for all t. Thus
(6.45) is true and so Q(t) is nonsingular on [0, ∞).

6.4.2 Lyapunov Transformations

The preceding shows that we can transform any time varying homegeneous linear differential equation into
an essentially aribitrary time invariant one by suitable change of variables. In view of the arbitraryness
of B, such a transformation may of course not be very illuminating. One class of transformations which
leads to useful results, consists of those Q’s which preserve “stability.” Such Q’s are called “Lyapunov”
transformations. More precisely, Q is a Lyapunov Transformation if

1. Q is continuously differentiable on [0, , ∞)

2. Q and Q̇ are bounded on [0, ∞)

3. there exists a positive number µ such that | det Q(t)| ≥ µ for t ∈ [0. ∞)
6 By defining a state vector consisting of the columns of Q stacked one atop the next, one can rewrite this equation as

a standard linear vector differential equation; existence and uniqueness theorems thus extend effortlessly to linear matrix
differential equations such as this.
The last condition stipulates that the deteminant of Q must be bounded away from zero as t → ∞. {Note,

for example, that in the scalar case Q = e− t fails to have this property.} Together conditions 1 and 3 insure
that Q−1 (t) exists and is bounded on [0, ∞).

Let Q be a Lyapunov Transformation and suppose that x and z satisfy (6.38) and (6.40) respectively. It
can then be shown that if all solutions to (6.38) are bounded on [0, ∞) then so are all solutions to (6.40) -
and conversely if all solutions to (6.40) are bounded on [0, ∞) then so are all solutions tp (6.38). It can also
be shown that all solutions to (6.38) must tend to zero as t → ∞ if and only if all solutions to (6.40) must
tend to zero as t → ∞. In other words, if Q is a Lyapunov Transformation, then the stability properties of
the differential equations in(6.38) and (6.40) are equivalent.

The definition of a Lyapunov Transformation Q does not require Q to satisfy (6.42). If A is such that it is
in fact possible to find a Lyapunov Transformation Q for which (6.42) holds with B a constant matrix, then
is (6.38) is said to be reducible to a constant linear differential equation. A reducible differential equations
is thus a differential equation whose stability properties can be decided {in theory at least} by examining
the stability properties of a related constant coefficient differential {i.e., (6.40)}. In the next chapter we will
explain how to determine the stability properties of the later.

Of course if one were to follow this program, one would still have to have methods for determining which
linear differential equations are reducible. In general this is a difficult matter. There is however, one type
of linear differential equation which is necessarily reducible. In particular we will now show that any linear
differential equation with a periodic A is reducible.

6.4.3 Floquet’s Theorem

We now invoke the assumption that A(t) is a continuous matrix for which
A(t + T ) = A(t), t ∈ [0, ∞) (6.47)
for some number T ≥ 0. What we want is for B to be defined so that
eBT = Φ(T, 0) (6.48)
Since explaining just how to do this would take us beyond the scope of this course, we will limit are discussion
of this point to the following remark.

Remark 3 The problem of finding solution a solution Xn×n to the transendental matrix equation
eX = Rn×n (6.49)
is discussed on pages 239 - 241 of [1]. It turns out that all that’s required for such a solution to exist is
that R be nonsingular. However (6.49) may not have a real-valued solution7 even if R is real. Consider for
example the scalar case in which R = −1. In general solutions to (6.49) are not unique. On the other hand,
for given R all solutions have certain properties in common and these are precisely the properties needed to
characterize the the stability of ż = Bz.

What we want to do now is to show that the matrix Q defined by (6.43) and (6.44), or equivalently by
(6.45), is periodic. For ease of reference we rewrite (6.45) here:
Q(t) = etB Φ−1 (t, 0) (6.50)
7 The ∆ 1
definition of eM for complex M is exactly the same as for real M ; i.e., eM = j
P∞
j=0 j! M . Convergence of this series
can be proved in essentially the same way as for the real case.
To begin, let us note that
Q(t + T ) = e(t+T )B Φ−1 (t + T, 0)
By the composition rule

Φ(t + T, 0) = Φ(t + T, T )Φ(T, 0) and e(t+T )B = etB eT B

Thus
Q(t + T ) = etB eT B Φ−1 (T, 0)Φ−1 (t + T, T )
so by (6.48)
Q(t + T ) = etB Φ−1 (t + T, T ) (6.51)
Meanwhile,
Φ̇(t, 0) = A(t)Φ(t, 0)
and
Φ̇(t + T, T ) = A(t + T )Φ(t + T, T )
In view of (6.47)
Φ̇(t + T, T ) = A(t)Φ(t + T, T )
Thus Φ(t, 0) and Φ(t + T, T ) both satisfy differential equation. Since Φ(t, 0) and Φ(t + T, T ) also both equal
the identity matrix at t = 0, by uniqueness of solutions to linear differential equations it must be that
Φ(t, 0) = Φ(t + T, T ) for all t ≥ 0. From this, (6.51) and (6.50) it follows that

Q(t + T ) = etB Φ−1 (t, 0) = Q(t), t ∈ [0, ∞)

Therefore Q(t) is periodic.

The definition of Q in (6.50) and the assumed continuity of A imply that Q is continuously differentiable.
This and the fact that Q is periodic, further imply that Q and Q̇ are bounded on [0, ∞). Moreover, since
etB and Φ(t, 0) are nonsingular for finite t, and since Q is periodic, it must be that the determinant of Q
is bounded away from zero for all t. In other words, Q is a Lyapunov Transformation and ẋ = A(t)x is a
reducible differential equation.

It is possible to carry things a bit further. Note from (6.50) that

Φ(t, 0) = Q−1 (t)etB

Moreover since
Φ(0, τ ) = Φ−1 (τ, 0)
it must be true that
Φ(0, τ ) = e−τ B Q(τ )
Thus
Φ(t, τ ) = Φ(t, 0)Φ(0, τ ) = Q−1 (t)etB e−τ B Q(τ )
We have proved the following theorem.

Theorem 6 (Floquet) It is possible to write the state transition matrix Φ(t, τ ) of any continuous matrix
A(t) which satisfies A(t + T ) = A(t), t ≥ 0 as

Φ(t, τ ) = Q−1 (t)e(t−τ )B Q(τ ), t, τ ∈ [0, ∞)

where B is a constant matrix and Q(t) is a nonsingular, continuously differentiable matrix satisfying Q(t +
T ) = Q(t), t ≥ 0.
Chapter 7

Matrix Similarity

Two n × n matrices A and B with elements in a field IK are similar if there exists a nonsingular n × n matrix
T with elements in IK such that
T AT −1 = B (7.1)
Matrix similarity is an equivalence relation on the set of all n × n matrices with elements in IK. The aim of
this chapter is to explain the significance and implications of matrix similarity.

7.1 Motivation

Consider the differential equation


ẋ = Ax (7.2)
where A is a constant n × n matrix with elements in IK. Suppose that T and B are any IK-matrices such
that (7.1) holds. If we define the new state vector

z = Tx (7.3)

then from (7.2) we see that


ż = T ẋ = T Ax = T AT −1z
Therefore because of (7.1)
ż = Bz (7.4)
Thus the similarity transformation
A 7−→ T AT −1 = B
describes what happens to A when one makes the change of state variables

x 7−→ T x = z (7.5)

We claim that
etB = T etA T −1 (7.6)
and thus that
etA 7−→ T etAT −1 = etB

103
describes what happens to etA under the change of state variables given by (7.5). To understand why (7.6)
is so, first observe that the matrix

Θ(t) = T etA T −1
satisfies   
Θ̇ = T AetA T −1 = T AT −1 T etA T −1 = T AT −1 Θ
and thus
Θ̇ = BΘ
Since etB satisfies the same differential equation and since Θ(t) and etB both equal the identity matrix at
t = 0, by uniqueness it must be true that etB = Θ(t), t ≥ 0. In other words, (7.6) is true.

Equation (7.6) suggests a procedure for calculating etA : Find matrices T and B such that (7.1) holds
and in addition, such that etB can be evaluated by inspection. Let us digress briefly to discuss one class of
matrices whose matrix exponentials are very easy to compute.

7.1.1 The Matrix Exponential of a Diagonal Matrix

Let us consider an arbitrary diagonal matrix


 
d1 0 ··· 0
∆  0 d2 ··· 0 
D= ... .. ..  ,
0 . . 
0 0 · · · dn

whose diagonal entries are elements in IK. For such a matrix, the definition of a matrix exponential

X∞
∆ 1 i
eD = D
i=0
i!

readily reduces to
 P∞ di1

0 ··· 0  d1 
i=0 i! e 0 ··· 0
 P∞ di2 

D ∆  0 ··· 0   0 ed2 ··· 0 
e = i=0 i! = . .. .. .. 
.. .. .. ..   .. . . . 
 . . . 
P∞. din 0 0 · · · edn
0 0 ··· i=0 i!

In other words the matrix exponential eD of a diagonal matrix D is the diagonal matrix
 
ed1 0 ··· 0
 0 ed2 ··· 0 
eD =
 ... .. .. .. 
. . . 
0 0 · · · edn

7.1.2 The State Transition Matrix of a Diagonalizable Matrix

Suppose that we can somehow find a nonsingular n × n matrix T such that

T AT −1 = Λ (7.7)
where Λ is a diagonal matrix; i.e.,
 
λ1 0 ··· 0
 0 λ2 ··· 0 
Λ=
 ... .. ..  (7.8)
0 . . 
0 0 · · · λn
Since tΛ is diagonal, it must therefore be true that
 
etλ1 0 ··· 0
 0 etλ2 ··· 0 
etΛ =
 ... .. ..  (7.9)
0 . . 
0 0 · · · etλn


Moreover by exploiting (7.6) with B = Λ we can write

etA = T −1 etΛ T, t ≥ 0 (7.10)

This suggests a procedure for calculating etA :

1. Find, if possible, a nonsingular matrix T and a diagonal matrix Λ as in (7.8) for which (7.7) holds.
2. Construct the matrix exponential etΛ as in (7.9).
3. Use (7.10) to obtain etA .

For this to be a constructive method, one of course must know if step 1 can in fact be completed and if so,
how one might go about doing it. These are issues in linear algebra. We now begin a discussion aimed at
resolving both of them.

7.2 Eigenvalues

Let A be a given n × n matrix over IK. Whether A is real or complex, a complex number λ ∈ Cl is said to
be an eigenvalue of A is there exists a nonzero vector x ∈ Cl n such that

Ax = λx (7.11)

Any vector x for which (7.11) holds is called an eigenvector of A for eigenvalue λ. Note that the set of all x
for which (7.11) holds is closed under addition and scalar multiplication and is thus a subspace of Cn ; this
subspace is uniquely determined by A and λ and is called the eigenspace of λ. It is easy to see that the
kernel of λI − A is the eigenspace of λ. Clearly any nonzero vector in the eigenspace of λ is an eigenvector
for λ.

7.2.1 Characteristic Polynomial

Suppose we wish to calculate a eigenvalue of A. In view of (7.11), what we need to do is to find a number λ
{the eigenvalue} and a nonzero vector x such that

(λI − A)x = 0 (7.12)


where I is the n × n identity matrix. Now for arbitrary but fixed λ, (7.12) will have a nonzero solution x
if and only if the kernel of λI − A is nonzero. For λI − A to have a nonzero kernel is equivalent to the
requirement that λI − A have linearly dependent columns. Since λI − A is a square matrix, its columns will
be linearly dependent just in case rank (λI − A) < n or equivalently λI − A is singular. In other words λ
will be an eigenvalue of A if and only if
det(λI − A) = 0 (7.13)

With the preceding in mind let us define



α(s) = det(sI − A) (7.14)

where s is a variable. Note that α(s) must be a polynomial in IK[s], the ring of polynomials in the variable s
with coefficients in IK. This is because sI − A is a matrix whose elements are polynomials in IK[s], because
the determinant of any matrix is the sum of products of its elements, and because IK[s] is closed under
polynomial addition and multiplication. We call α(s) the characteristic polynomial of A.

Using basic properties of determinants it is not difficult to prove that α(s) is a polynomial of degree n
whose leading coefficient is one. Polynomials with the latter property are sometimes referred to as monic.

From the definition of α(s) and (7.13) it is clear that for λ to be an eigenvalue of A, it is both necessary
and sufficient that λ be a root of α(s); i.e., λ must be a solution to the characteristic equation

α(λ) = 0 (7.15)

This of course is just another way to write (7.13). From these observations we can draw the following
conclusions:

1. The spectrum of A {i.e., the set of eigenvalues of A} consists of precisely n complex numbers, namely
the n roots of the characteristic polynomial of A.
2. If A is a real matrix, then its characteristic polynomial α(s) is in IR[s] and thus has real coefficients.
Because of this, all of α(s)’s nonreal roots must occur in complex conjugate pairs. Thus all nonreal
eigenvalues of a real matrix A must occur in conjugate pairs. To emphasize this, we sometimes say
that A’s spectrum is symmetric.

Suppose λ is an eigenvalue of A. Then s − λ must be a factor of α(s). The largest value of d for which
(s − λ)d is a factor of α(s) is called the algebraic multiplicity of λ. Meanwhile the dimension of the eigenspace
of λ is called the geometric multiplicity of λ. The geometric multiplicity of λ can never be larger that the
algebraic multiplicity of λ

7.2.2 Computation of Eigenvectors

Suppose λ is an eigenvalue of A. To calculate an eigenvector x for this eigenvalue, what we need to do is to


find a nonzero solution to the linear equation

(λI − A)x = 0 (7.16)

Note that such a solution must exist because (λI − A) has linearly dependent columns. To calculate such an
x for small n, one might attempt to solve (7.16) by hand. For large n, one would turn something like Gauss
Elimination, perhaps transforming (λI − A) into reduced echelon form and then solving for x by inspection.
Note that x is never uniquely determined, so any attempt to come up with an explicit formula for x is out
of the question.
7.2.3 Semisimple Matrices

A matrix An×n over IK is said to be semisimple if it has n linearly independent eigenvectors. Suppose A is
semisimple with a linearly independent eigenvector set {q1 , q2 , . . . , qn }. Then {q1 , q2 , . . . , qn } is a basis
for Cl n . Therefore the matrix

Q = [ q1 q2 · · · qn ]
is nonsingular. Moreover, since
Aqi = λi qi , i ∈ {1, 2, . . . , n}
it must be true that
AQ = [ Aq1 Aq2 · · · Aqn ] = [ λ1 q1 λ2 q2 · · · λn qn ]
and thus that
AQ = QΛ
where  
λ1 0 ··· 0
 0 λ2 ··· 0 
Λ=
 ... .. .. .. 
. . . 
0 0 · · · λn
It follows that A is similar to Λ; i.e.,
Q−1 AQ = Λ
We leave it to the reader to verify that the converse is also true. That is, if B is a matrix for which there exists
a matrix P such that P −1 BP is diagonal, then the columns of P are eigenvectors of B and consequently B
is a semisimple matrix. We summarize:

Theorem 7 A matrix An×n is semisimple if and only if there exists a nonsingular matrix Q such that
Q−1 AQ = Λ where Λ is a diagonal matrix. If Q and Λ are such matrices, then the n diagonal elements of Λ
are A’s eigenvalues and the n columns of Q are eigenvectors of A which together form a linearly independent
set.

Example 44 Suppose  
0 −2 0
A =  1 −3 0
0 0 0
Then  
λ 2 0
α(s) = det(sI − A) = det  −1 s + 3 0  = s(s2 + 3s + 2) = s(s + 1)(s + 2)
0 0 s
∆ ∆
Hence A’s three eigenvalues are 0, −1 and −2. To compute an eigenvectors q1 , q2 , q3 for λ1 = 0, λ2 = −3,

λ3 = −2 we need to solve the equations
     
0 2 0 −2 2 0 −1 2 0
 −1 3 0  q1 = 0,  −1 1 0  q2 = 0,  −1 2 0  q3 = 0
0 0 0 0 0 −2 0 0 −1

One set of solutions is      


0 1 2
q1 =  0  , q2 =  1  , q3 =  1 
1 0 0

Thus if we define Q = [ q1 q2 q3 ], then

 
0 0 0
Q−1 AQ =  0 −2 0 
0 0 −1

It is a sad but true fact that there is no simple way to decide if a given matrix is semisimple. On the
other hand it turns out that “almost every” matrix is semisimple. Confused? Stay tuned . . .

7.2.4 Eigenvectors and Differential Equations

Let An×n be a given constant matrix and consider the differential equation

ẋ = Ax (7.17)

with initial state x(0) = d. We want to describe the form of the solution to (7.17) for various types of d’s.

1. Suppose first that d is an eigenvector of A for eigenvalue λ. Then


x(t) = etλ d (7.18)

must be the solution to (7.17) because with x so defined x(0) = d and

ẋ(t) = λetλ d = etλ Ad = Ax(t)

The form of x(t) in (7.18) clearly shows that if the initial state of (7.17) lies in span {d} then the value
of the corresponding solution x(t) at each instant of time, is in span {d}, since etλt d is a scalar multiple
of d. In other words, any trajectory of (7.17) starting in span {d} never leaves span {d}.

Example 45 If
   
0 −2 0 1
A = 1 −3 0  and d = 1
0 0 0 0

then Ad = −2d. Thus


   −t2 
x1 (t) e
x(t) =  x2 (t)  =  e−t2 
x3 (t) 0
x3

1
x1

1 eigenspace

trajectory
x2

2. Now suppose that d is a vector in some subspace E which in turn is spanned by some set of eigenvectors
{q1 , q2 , . . . , qm }. In other words, suppose that

d ∈ E = E1 + E2 + · · · + Em

where for each i, Ei is spanned by qi . This means there must be numbers µi ∈ Cl such that

d = µ1 q1 + µ2 q2 + · · · + µm qm

as well as eigenvalues λi , i ∈ {1, 2, . . . , m} such that Aqi = λi qi , i ∈ {1, 2, . . . , m}. Using the same
reasoning as before it is easy to establish that under these conditions, the solution to (7.17) is

x(t) = µ1 etλ1 q1 + µ2 etλ2 q2 + · · · + µm etλm qm

Since for each fixed value of t, x(t) is a linear combination of the qi , it must be true that for each t ≥ 0,
x(t) ∈ E. We can therefore conclude that any trajectory of (7.17) starting in E remains in E forever.

7.3 A - Invariant Subspaces

The subspace E just discussed is an example of an “A - invariant subspace.” The aim of this section is to
explain what an A - invariant subspace is and to discuss some of the implications of A - invariance.

Let U be a subspace of Cl n {or perhaps IRn if A is real} with the property that Au ∈ U for every u ∈ U.
We call U an A-invariant subspace. To denote the A-invariance of U we often write AU ⊂ U. The notation

is appropriate if we take AU to mean the subspace AU = {Au : u ∈ U}. Examples of A-invariant subspaces
include the zero space, Cl n , eigenspaces and sums of eigenspaces. Later we will encounter other types of
A-invariant subspaces.

Now suppose that the initial state d of (7.17) is a vector in an A-invariant subspace U. Since ẋ(t) = Ax(t),
it follows that ẋ(0) = Ad ∈ U. Thus the solution to (7.17) which starts at d, is not only starting at a point
in U, but also has a derivative which is initially in U. From this it is plausible {and can easily be shown}
that the entire trajectory starting at d, remains in U forever. In other words, trajectories of (7.17) starting
in an A-invariant subspace never leave the subspace.

The fact that U is an A-invariant subspace has implications concerning A’s algebraic structure under
similarity transformations. Our aim is to explain what these implications are. For this let V be any com-
pletion1 from U to IKn ; i.e., any subspace such that U ⊕ V = IKn . Pick bases {u1 , u2 , . . . , um } and
{v1 , v2 , . . . , vn−m } for U and V respectively. Then {u1 , u2 , . . . , um , v1 , v2 , . . . , vn−m } is a basis for IKn .
It follows that the basis matrices
∆ ∆
U = [ u1 u2 · · · um ]n×m and U = [ v1 v2 · · · vn−m ]n×(n−m)

have linearly independent columns and that the matrix



T = [U V ]n×n

is nonsingular.

Note that since AU ⊂ U, the span of the columns of AU {namely AU} must be a subspace of the span
of the columns of U {namely U}. Thus the necessary and sufficient condition is satisfied for there to exist a
solution Bm×m to the linear matrix equation

AU = U B (7.19)

Moreover, because U has linearly independent columns, kernel U = 0 so B is even unique. We will call B
a restriction of A to U. On your homework, you are asked to prove that all possible restrictions of A to U,
resulting from all possible choices of bases for U, are similar.

In view of (7.19) we can write  


B
AU = [ U V]
0
or  
B
AU = T (7.20)
0
where the 0 in these partitioned matrices denotes the (n − m) × m zero matrix. What’s more, if we defined
the partitioned matrix  
C ∆ −1
= T AV,
D
where C and D are m × (n − m) and (n − m) × (n − m) matrices respectively, then
 
C
AV = T (7.21)
D

Combining (7.20) and (7.21) we thus obtain


 
B C
A[U V ]=T
0 D
or  
B C
AT = T
0 D
1 Such a subspace can be constructed as follows. First pick bases {u , u , . . . , u } and {x , x , . . . , x } for U and IKn
1 2 m 1 2 n
respectively. Second, starting from the left, successively eliminate from the set {u1 , u2 , . . . , um , x1 , x2 , . . . , xn } vectors which
are linearly dependent on their predecessors. What results is a basis for IKn of the form {u1 , u2 , . . . , um , v1 , v2 , . . . , vn−m }
where each vi ∈ {x1 , x2 , . . . , xn }. Defining V to be the span of {v1 , v2 , . . . , vn−m } yields a subspace with the required
properties.
Therefore  
B C
T −1 AT =
0 D
The matrix on the right with square diagonal blocks B and D is called a block triangular matrix. Thus we
can conclude the following If U is an A-invariant subspace, then A is similar to a block triangular matrix
with one diagonal block being a restriction of A to U.

Why is this interesting? Well for one thing, the characteristic polynomial of a block triangular matrix is
equal to the product of the characteristic polynomials of its diagonal blocks; i.e.,
  
B C
det sI − = (det(sI − B)) (det(sI − D))
0 D

Thus the spectrum of a block triangular matrix is the union of the spectra of its diagonal blocks. More on
this a little later.

7.3.1 Direct Sum Decomposition of IKn

Suppose that the completion V constructed above is also A - invariant. In other words, let us consider the
situation in which
U ⊕ V = IKn , AU ⊂ U, AV ⊂ V (7.22)

Then it is reasonably clear and can easily be verified using the same reasoning as above that with T, B, C, D
as just defined, C will turn out to be the zero matrix. In other word, under the conditions of (7.22),
 
B 0
T −1 AT =
0 D

Thus A is similar to a block diagonal matrix whose diagonal blocks B and D are restrictions of A to U and
V respectively.

These observations generalize. Suppose, for example, that X1 , X2 , . . . , Xm are subspaces of IKn satisfying

X1 ⊕ X2 ⊕ · · · ⊕ Xm = IKn and AXi ⊂ Xi , i ∈ {1, 2, . . . , m}

Then by forming the nonsingular matrix

∆  
T = x(1)
1 x2
(1) (1)
· · · xn1
(2)
x1 x2
(2)
· · · xn2
(2)
· · · x1
(m) (m)
x2
(m)
· · · xnm

(i) (i) (i)


where {x1 . x2 , · · · , xni } is a basis for Xi , one can transform A into a block diagonal matrix as
 
A1 0 ··· 0
 0 A2 ··· 0 
T −1 AT = 
 ... .. .. .. 
. . . 
0 0 · · · Am

where Ai is a ni × ni restriction of A to Xi . Note that the diagonalization of a matrix into scalar diagonal
blocks is a special case of this in which the Xi are independent eigenspaces and the Ai are 1 × 1 matrices,
namely the eigenvalues of A. We summerize: If IKn can be decomposed into a direct sum of m A-invariant
subspaces Xi , then A can be transformed by means of a similarity transformation into a block diagonal matrix
whose diagonal blocks are restrictions of A to the Xi .
7.4 Similarity Invariants

There are certain properties of A which are shared by all matrices similar to A. For example, if A is n × n,
then so is every matrix similar to A. We say that the size of A is invaraint under similarity transformations.
the size of A is sometimes called a similarity invariant.

Let us note that if T is any nonsingular matrix, then

det(sI − A) = det(T T −1 (sI − A))


= det(T ) det(T −1 ) det(sI − A)
= det(T (sI − A)T −1 )
= det(sI − T AT −1)

Thus A and T AT −1 have one and the same characteristic polynomial. In other words, the characteristic
polynomial of any square matrix is invariant under similarity transformations. From this it follows that the
spectrum of any square matrix is invariant under similarity transformations.

Note for example that if A is similar to a block triangular matrix with diagonal blocks A1 , A2 , . . . , Am ,
then the characteristic polynomial {spectrum} of A must equal the product {union} of the characteristic
polynomials {spectra} of the Ai . It is largely for these reasons that the direct sum decompositions discussed
in the last section are of interest.

7.5 The Cayley-Hamilton Theorem

The following theorem is one of the most important results in linear algebra.

Theorem 8 (Cayley-Hamilton) Let A be a given n × n matrix with characteristic polynomial α(s). Then

α(A) = 0n×n

The theorem says that if α(s) is the characteristic polynomial of A, then the matrix polynomial α(A) equals
the zero matrix. Although somewhat confusing, the assertion “A satisfies its own characteristic equation” is
often taken as a statement of the Cayley-Hamilton Theorem.

Example 46 As noted earlier, the characteristic polynomial of the matrix


 
0 −2 0
A =  1 −3 0 
0 0 0

is
α(s) = s3 + 3s2 + 2s
According to the Cayley-Hamiliton
A3 + 3A2 + 2A + 0I = 0
To verify that this is so we first compute
   
−2 6 0 6 −14 0
A2 =  −3 7 0  and A3 =  7 −15 0 
0 0 0 0 0 0
Then          
6 −14 0 −2 6 0 0 −2 0 1 0 0 0 0 0
 7 −15 0  + 3  −3 7 0  + 2  1 −3 0  + 0  0 1 0 = 0 0 0
0 0 0 0 0 0 0 0 0 0 0 1 0 0 0

7.6 Minimal Polynomial

So far we have shown that any semisimple matrix An×n can be diagonalized by means of a similarity
transformation. Unfortunately, not every matrix is semisimple as the following example shows.

Example 47 The eigenvalues of the matrix


 
0 1
A=
0 0
are both 0. Thus for  
x1
x=
x2
to be an eigenvector, it must be a nonzero solution to the equations
  
0 1 x1
=0
0 0 x2
But all such solutions must clearly be of the form
 
x1
x=
0
It is therefore impossible to find two eigenvectors which are not scalar multiples of each other. In other
words A does not have two linearly independent eigenvectors so it is not semisimple.

Some readers may think that the A matrix used in this example is pathelogical, and that in practice
one typically does not encounter matrices of this type. While it is true that the matrix has some very
special features {e.g., its eigenvalues are both the same}, this matrix is nevertheless commonly encountered
in practical modeling problems. For example, the cascade interconnection of two Simulink integrators shown
in the following diagram, admits a state space model of the form
      
ẏ1 0 1 y1 0
= + u
ẏ2 0 0 y2 1

u y2 y1
1/s 1/s
Integrator1 Integrator
Signal
Generator

In order to understand more clearly just why some matrices are semisimple while others are not, we need
some new concepts. In the sequel we shall explain what they are.
7.6.1 The Miminal Polynomial of a Vector

Let An×n be a given IK matrix and let x be any nonzero vector in IKn . Let m be the least integer such
that the set of vectors {x, Ax, . . . , Am−1 x} is a linearly independent set. Clearly 1 ≤ m ≤ n. Since
{x, Ax, . . . , Am−1 x} is a linearly independent set and {x, Ax, . . . , Am x} is not, Am x must be a linear
combination of the vectors in {x, Ax, . . . , Am−1 x}. In other words, there must be exactly one set of
numbers ai such that
Am x = −a1 x − a2 Ax − · · · − am Am−1 x
Thus
(Am + am Am−1 + · · · + a2 A + a1 I)x = 0 (7.23)
Therefore if we define the polynomial

αx (s) = sm + am sm−1 + · · · a2 s + a1

then (7.23) can be rewritten as


αx (A)x = 0 (7.24)
We call αx (s) the minimum polynomial of x {relative to A}. Note that αx (s) is uniquely determined by A
and x. {For consistency we define the minimal polynomial of the zero vector to by the polynomial 1}.

Let us note that if Ix is the {nonempty} set of all polynomials β(s) ∈ IK[s] such that β(A)x = 0, then
Ix is closed under polynomial addition and multiplication by polynomials in IK[s]; i.e.,

β(A)x = 0, δ(A)x = 0 ⇒ (β(A) + δ(A))x = 0

β(A)x = 0, π(s) ∈ IK[s] ⇒ π(A)β(A)x = 0

Thus Ix is an ideal in IK[s]. It follows from Proposition 2 that Ix contains a unique monic polynomial γ(s)
of least degree which divides {without remainder} every polynomial in Ix . Since αx (s) is monic, is in Ix ,
and by construction is of least degree it must be true that γ(s) = αx (s). In other words, if β(s) is any
polynomial such that β(A)x = 0 then β(s) must be divisable by αx (s).

7.6.2 The Minimal Polynomial of a Finite Set of Vectors


Now suppose that X = {x1 , x2 , . . . , xm } is a finite set of vectors in IKn . If we define IX to be the set of
all polynomials β(s) such that
β(A)xi = 0, i ∈ {1, 2, . . . , m}
then IX is also an ideal in IK[s]. {The reader may wish to verify this using the same reasoning as above.}
Again from Proposition 2 it follows that there is a unique monic polynomial αX (s) of least degree which
divides every polynomial in IX . We call αX (s) the minimal polynomial of X relative to A. It is not difficult
to prove that
αX (s) = lcm{αx1 (s), αx2 (s), . . . , αxm (s)}
where αxi (s) is the minimal polynomial of xi .

7.6.3 The Minimal Polynomial of a Subspace

Now suppose that U is a subspace of IKn and let IU be the set of all polynomials β(s) ∈ IK[s] such that

β(A)u = 0, ∀u ∈ U
IU is also an ideal and its unique, monic generating polynomial, αU (s), is called the minimal polynomial of
U relative to A. Suppose U is not the zero subspace and let {u1 , u2 , . . . , um } be a basis for U. It is easy
to verify that β(s) ∈ IU if and only if

β(A)ui = 0, i ∈ {1, 2, . . . , m}

From this it follows that the minimal polynomial of U is the same as the minimal polynomial of any basis
for U. Hence
αU = lcm{αu1 (s), αu2 (s), . . . , αum (s)}

7.6.4 The Minimal Polynomial of A

In the spacial case when U = IKn , αU is written without a subscript as α(s) and is called the minimal
polynomial of A. Thus the minimal polynomial of A is the unique monic polynomial of least degree such
that
α(A)x = 0, ∀x ∈ IKn

This is equivalent to saying that α(s) is the unique monic polynomial of least degree such that

α(A) = 0

Let us recall that if β(s) is the characteriatic polynomial of A then β(A) = 0 because of the Cayley-
Hamilton Theorem. Since α(s) devides every polynomial δ(s) such that δ(A) = 0 it must be true that the
minimal polynomial of A divides the characteristic polynomial of [Link] useful consequence of this fact is that
if A had n distinct eigenvalues, the A’s minimal and characteristic polynomials must be one and the same.

It can be shown without difficulty that if Y1 and Y2 are any nonempty subsets or subspaces of IK and
if Y1 ⊂ Y2 , then αY1 (s) divides αY2 (s). Thus the minimal polynomial of any subspace of IKn divides the
minimal polynomial of A.

Example 48 If
   
2 0 0 1
A = 0 2 1 and x = 0
0 0 2 0
then
minimum polynomial of x = s−2 (A − 2I)x = 0

minimum polynomial of A = (s − 2)2 (A − 2I)2 = 0

characteristic polynomial of A = (s − 2)3 (A − 2I)3 = 0

7.6.5 A Similarity Invariant

Let A and B be similar n × n matrices over IK. Then there must be a nonsingular matrix T such that

B = T AT −1
Note that

T A0 T −1 = I = B0
T AT −1 = B
T A2 T −1 = T AAT −1 = T AT −1 T AT −1 = B 2
..
.
i −1
TA T = Bi
..
.

More generally, for any polynomial β(s) ∈ IK[s],

T β(A)T −1 = β(B)

Thus if β is the minimal polynomial of B, then T β(A)T −1 = 0. Since T is nonsingular, β(A) = 0. Therefore
β(s) must be divisable by the minimal polynomial of A. By repeating the same argument with B and A
interchanged, we can establish that the minimal polynomial of A must be divisable by β(s). Since both
polynomials divide each other and are monic, they must be one and the same. We have proved that the
minimal polynomial of A equals the minimal polynomial of any matrix similar to A. In other words, the
minimum polynomial of A is invariant under similarity transformations.

7.7 A Useful formula

The rational matrix (sI − A)−1 , called the resolvent of A, plays a key role in the determination of etA by
means of Laplace Transform Techniques. While (sI − A)−1 can be expanded in the formal power series

X
(sI − A)−1 = Ai−1 s−i
i=1

which converges for all values of |s| sufficiently large, it is sometimes useful to express this matrix in more
explicit terms. To do this, let

α(s) = sm + an sm−1 + · · · + a2 s + a1
be any monic polynomial such that α(A) = 0. Thus for example, α(s) might be either the minimal or
characteristic polynomial of A. Next, successively define polynomials αm (s), αm−1 (s), . . . , α1 (s), α0 (s) so
that

αm (s) = 1

αi−1 (s) = sαi (s) + ai , i ∈ {m, , m − 1, . . . , 1}

{Note that α0 (s) = α(s).} Starting with the easily verifiable identity
m
X m
X
(sI − A) πi (s)Ai−1 = sπ1 (s) − πm (s)Am + {sπi (s) − πi−1 (s)} Ai−1 ,
i=1 i=2

which holds for any set of m polynomials π1 (s), π2 (s), . . . , πm (s), one can show that

m
1 X
(sI − A)−1 = αi (s)Ai−1 (7.25)
α(s) i=1
7.7.1 Application

Recall that the matrix exponential etA uniquely solves the matrix differential equation/initial condition
d tA
e = AetA , etA t=0
=I
dt
Thus if X(s) is the Laplace Transform2 of etA , then

sX(s) − I = AX(s)

Therefore
(sI − A)X(s) = I
so
X(s) = (sI − A)−1
Hence using (7.25) we can write
m
1 X
X(s) = αi (s)Ai−1
α(s) i=1
Therefore
m
X
etA = gi (t)Ai−1
i=1
i (s)
where gi (t) is the inverse Laplace Transform of the rational function αα(s) . Of course to complete the
calculation of etA one would have to actually carry out the evaluations of the gi (t). This, in turn, is not
really any easier that the eigenvalue approach we’re pursuing here.

7.8 Cyclic Subspaces

Let An×n be a given matrix with elements in IK. A nonzero subspace U ⊂ IKn is A-cyclic if there exists a
vector g ∈ U, called a generator, such that

U = span {g, Ag, . . . , An−1 g}

Note that each vector x ∈ IKn generates a unique cyclic subspace, namely

span {x, Ax, . . . , An−1 x}

Cyclic subspaces are A-invariant subspaces. To understand why this is so, suppose u ∈ U where U =
span {g, Ag, . . . , An−1 g}. Then there must be numbers µi ∈ IK such that

u = µ1 g + µ2 Ag + · · · µn An−1 g

Therefore Au = µ1 Ag + µ2 A2 g + · · · µn An g. But by the Cayley-Hamilton Theorem, there must be numbers


ai ∈ IK such that
An + an An−1 + an−1 An−2 + · · · + a2 A + a1 I = 0
Thus
Au = −a1 g + (µ1 − a2 )Ag + (µ2 − a3 )A2 g + · · · + (µn−1 − an )An−1 g
so Au ∈ U. This establishes the the A-invariance of U.
2 Review this material when you study Laplace Transforms.
It is not difficult to prove that if U is cyclic with generator g, and if m is the largest integer such that
{g, Ag, . . . , Am−1 g} is an independent set, then {g, Ag, . . . , Am−1 g} is a basis for U. Using this, it can
then be shown that the minimal polynomial of U is the same as the minimal polynomial of g.

If the whole space IKn is cyclic, A is said to be a cyclic matrix. Thus A is cyclic just in case there is a
vector x ∈ IKn such that {x, Ax, . . . , An−1 x} is a basis for IKn . A little thought reveals that cyclic matrices
are precisely those matrices whose minimal and characteristic polynomials are one and the same. Previously
we noted that matrices with all eigenvalues distinct have the same minimal and characteristic polynomials.
It followes that matrices with distinct eigenvalues are cyclic.

7.8.1 Companion Forms

Let An×n be a cyclic matrix with characteristic polynomial

α(s) = sn + an sn−1 + · · · + a2 s + a1

Let g be any vector which generates IKn . Then {g, Ag, . . . , An−1 g} is a basis for IKn and thus

Q = [g Ag · · · An−1 g ]

is a nonsingular matrix. Define the row vector



c = [0 0 ··· 0 1 ]1×n Q−1

Then
cQ = [ 0 0 ··· 0 1 ]1×n
so
cg = 0, cAg = 0, . . . . . . . . . cAn−2 g = 0 (7.26)
and
cAn−1 g = 1 (7.27)
Define the matrix  
c
∆  cA 
T = 
 ...  (7.28)
cAn−1 n×n

and note that  


0
0
.
Tg =  .
.
0
1 n×1

We claim that T is nonsingular. Suppose it were not; then it would be possible to find numbers bi such
that
n−1
X
cAn−1 = bi cAi−1
i=1

This would imply that


n−1
X
cAn−1 g = bi cAi−1 g = 0
i=1
because of (7.26). But this contradicts (7.27), so T cannot be singular.

In view of the definition of T in (7.28)




cA
 cA2 
TA =  
 ...  (7.29)
cAn
and
cAi = e′i+1 T, i ∈ {1, 2, . . . , n − 1} (7.30)
where ei+1 is the ith unit vector in IKn . Meanwhile, by the Cayley-Hamilton Theorem

cAn = −a1 c − a2 cA − · · · − an cAn−1

so
cAn = [ −a1 −a2 · · · −an ] T (7.31)
because of the definition of T in (7.28). From (7.29) - (7.31) we thus obtain the expression

 e′2 T 
 e′3 T 
 .. 
TA = 
 .


 
e′n T
[ −a1 −a2 · · · −an ] T
Factoring out T on the right we get
T A = AC T (7.32)
where  
0 1 0 ··· 0
 0 0 1 ··· 0 
 . .. .. .. .. 
AC = 
 .
. . . . . 

 0 0 0 ··· 1 
−a1 −a2 −a3 · · · −an n×n

Therefore

AC = T AT −1 (7.33)
AC is called a companion form. Note that the bottom row of AC consists of the negative of the row
vector of coefficients of the characteristic polynomial of A. In other words, AC is uniquely determined by
A’s characteristic polynomial and consequently by A itself. It thus makes sense to call AC the companian
{canonical} form of A. For the equivalence relation similarity defined on the class of n × n cyclic matrices,
companion forms are thus canonical forms and the characteristic polynomial is a complete invariant. The
following theorem says essentially the same thing.

Theorem 9 Each n × n cyclic matrix A with elements in IK is similar to exactly one companion form with
elements in IK and two n × n cyclic matrices are similar if and only if they have the same characteristic
polynomial.

Remark 4 It is worth noting that any monic polynomial β(s) of any degree, uniquely determines a com-
panion matrix BC whose bottom row is the negative of the row vector of β(s)’s coefficients. This provides
an especially easy way to construct a matrix BC whose characteristic polynomial is a given polynomial β(s).
Note in particular, that the roots of β(s) need not be calculated in order to carry out this construction.
7.8.2 Rational Canonical Form

Theorem 9 states that any cyclic matrix A is similar to exactly one companion form AC and moreover that
AC is uniquely determined by the characteristic polynomial of A. Unfortunately not every matrix is cyclic.

Example 49 For the matrix  


3 0
A=
0 3
and any vector x ∈ IK2 , Ax = 3x which is a scalar multiple of x. Thus there is no vector in IK2 which generates
IK2 and so A cannot be cyclic. Another reason A fails to be cyclic is because A’s minimal polynomial s − 3
is not equal to its characteristic polynomial (s − 3)2 .

To treat in detail the generalization of the preceding ideas to matrices which are not cyclic, is beyond the
scope of this course. What we shall do instead is to state without proof the main technical result applicable
to the noncyclic case. The reader interested in more detail should consult a good text on matrix algebra.

Theorem 10 (Rational Canonical Decomposition) For fixed n > 0, let An×n be a given matrix with
elements in IK. There exists a positive integer m and m nonzero cyclic subspaces Xi , i ∈ {1, 2, . . . , m}, with
minimal polynomials αi (s), i ∈ {1, 2, . . . , m} respectively, such that α1 (s) is the minimum polynomial of A,
αi (s) divides αi−1 (s) for i ∈ {2, 3, . . . , m}, and
IKn = X1 ⊕ X2 ⊕ · · · ⊕ Xm
Moreover if X¯1 , X¯2 , . . . , X¯m̄ is any set of subspaces with the preceding properties, then m̄ = m and the
minimal polynomial of X¯i equals αi (s), i ∈ {1, 2, . . . , m}.

A constructive procedure for finding the component subspaces Xi and the αi (s) which satisfy the conditions
of the preceding theorem can be found in [1], Chapter 7, Section 4.

The first implication of the Rational Canonical Decomposition Theorem is that for any given n×n matrix
A, one can calculate a unique set of polynomials {α1 (s), α2 (s), . . . , αm (s)}, which is ordered so that each
αi (s) divides its predecessor. The αi (s) are called the invariant factors of A.

The second implication of the theorem is that A is similar to a block diagonal matrix

 
A1 0 ··· 0
 0 A2 ··· 0 
A∗ = 
 ... .. .. .. 
. . . 
0 0 · · · Am n×n
whose diagonal blocks are companian matrices, αi (s) being the characteristic polynomial of Ai . A∗ is called
the rational canonical form of A∗ .

The block diagonal structure of A∗ is a direct consequence of the fact that IKn is the direct sum of the
A-invariant subspaces Xi , i ∈ {1, 2, . . . , m}. Each Ai is a cyclic matrix with characteristic polynomial αi (s)
because Xi is a cyclic subspaces with minimal polynomial αi (s). Each Ai is actually a companion form
because of Theorem 9.

It can be shown that the invariant factors of A are similarity invariants. Since each n × n matrix A is
similar to exactly one rational canonical form and since each canonical form is uniquely determined by A’s
invariant factors, we have a criterion for deciding when two matrices are similar.
Theorem 11 Two n× n matrices with elements in IK are similar if and only if they have the same invariant
factors.

The block diagonal structure of A∗ implies that the characteristic polynomal of A∗ {and hence A} is
equal to the product of the invariant factors of A. That is

α(s) = α1 (s)α2 (s) · · · αm (s)

where α(s) is A’s characteristic polynomial. Since α1 (s) is the minimum polynomial of A, it follows that the
minimum polynomial of A divides the characteristic polynomial of A, which is something we also concluded
earlier by appealing to the Calyey-Hamilton Theorem. Since α1 (A) = 0, it must be that α(A) = 0 which is a
restatment of the Calyey-Hamilton Theorem. In other words, the Calyey-Hamilton Theorem is a consequence
of the Rational Canonical Decomposition Theorem. Clear?

Example 50 The matrix


 
0 1 0 0 0
 0 0 1 0 0 
 
A= 0 −2 −3 0 0 
 
0 0 0 0 1
0 0 0 0 −2
is in rational canonical form and its invariant factors are α1 (s) = s3 + 3s2 + 2s = s(s + 1)(s + 2) and
α2 (s) = s2 + 2s = s(s + 2). A’s characteristic polynomial and minimal polynomial are s2 (s + 1)(s + 2)2 and
s(s + 1)(s + 2) respectively, and its spectrum is {0, 0, −1, −2, −2}. If B is any matrix similar to A then B
must have the same invariant factors; conversely, if B’s invariant factors are α1 (s) and α2 (s), then B must
be similar to A.

7.9 Coprime Decompositions

While the rational canonical form is useful in certain applications, there is another canonical form which
is even more useful - especially in the characterization of etA . To develop this canonical form, we need a
several new ideas which we now explain.

Again let A be a given n × n matrix with elements in IK. Let α(s) be either the minimal or characteristic
polynomial of A. Suppose that α(s) can be written as

α(s) = α1 (s)α2 (s) (7.34)

where α1 (s) and α2 (s) are monic polynomials which are at the same time coprime; i.e.,

gcd{α1 (s), α2 (s)} = 1 (7.35)

Suppose we define

Xi = kernel αi (A), i ∈ {1, 2} (7.36)
n
Thus Xi is the set of all x ∈ IK such that αi (A)x = 0.

Claim 1: AXi ⊂ Xi , i ∈ {1, 2}

To show that this is so, fix i ∈ {1, 2} and x ∈ Xi . Then αi (A)x = 0 so Aαi (A)x = 0. Since αi (A) and A
commute, we can write αi (A)Ax = 0 or equivalently Ax ∈ kernel αi (A). Therefore Ax ∈ Xi . Since this must
be so for i ∈ {1, 2} and for every x ∈ Xi , claim 1 must be true.
{Note that this proof has not made use of any special properties of the αi (s). We have therefore incidently
proved that for any polynomial β(s), kernel β(A) is an A-invariant subspace.}

Claim 2: X1 + X2 = IKn

To establish this claim we need to make use of (7.35). In particular, the coprimeness assumption insures the
existence of polynomials β1 (s) and β2 (s) such that

β1 (s)α1 (s) + β2 (s)α2 (s) = 1 (7.37)

This is a consequence of Proposition 3 which appears at the very beginning of these notes. Observe that
(7.37) implies the matrix identity

β1 (A)α1 (A) + β2 (A)α2 (A) = In×n (7.38)

{If this surprises you, check it out for (say) α1 (s) = s2 +s+1, α2 (s) = s2 +2, β1 (s) = − 31 (s+1), β2 (s) = s+2,
and any A of any size you like.}

To proceed, let x be any vector in IKn . Then from (7.38)

x = β1 (A)α1 (A)x + β2 (A)α2 (A)x (7.39)

Thus if we define
x1 = β2 (A)α2 (A)x and x2 = β1 (A)α1 (A)x (7.40)
Then
x = x1 + x2 (7.41)

To prove claim 2 it is enough to show that xi ∈ Xi , i ∈ {1, 2}. Since α1 (s)α2 (s) is the minimal or
characteristic polynomial of A
α1 (A)α2 (A) = 0 (7.42)
From this and definitions of the xi in (7.40) it follows that

α1 (A)x1 = α1 (A)β2 (A)α2 (A)x = β2 (A)α1 (A)α2 (A)x = 0


α2 (A)x2 = α2 (A)β1 (A)α1 (A)x = β1 (A)α1 (A)α2 (A)x = 0

Therefore xi ∈ kernel αi (A) = Xi , i ∈ {1, 2}.

Claim 3: X1 ⊕ X2 = IKn

In view of claim 2, to establish claim 3 we need only show that X1 ∩ X2 = 0. For this, pick x ∈ X1 ∩ X2
and let x1 and x2 be defined as in (7.40). Since x ∈ X1 ∩ X2 , it must be true that x ∈ X2 so α2 (A)x = 0.
From this and (7.40) it follows that x1 = 0. By similar reasoning x2 = 0. Therefore by (7.41), x = 0. Thus
x ∈ X1 ∩ X2 implies x = 0 so X1 ∩ X2 = 0. Therefore claim 3 is true.

Putting claims 1 and 3 together we see that if α1 (s)α2 (s) is a coprime factorization of either the minimal

or characteristic polynomial of A, and if Xi = kernel αi (A), i ∈ {1, 2}, then X1 and X2 are independent A-
invaraint subspaces whose direct sum equals IKn . In view of the discussion in section 7.3.1, we can therefore
conclude that A must be similar to a block diagonal matrix of the form
 
A1 0
0 A2

where Ai is a restriction of A to Xi . The T matrix in the similarity transformation A 7−→ T −1 AT which


achieves this is of the form T = [ X1 X2 ] where for each i, Xi is a basis matrix for Xi . In addition, it can
be shown that if α1 (s)α2 (s) is the characteristic {minimal} polynomial of A, then for each i, αi (s) is the
characteristic {minimal} polynomial of Ai . These findings readily generalize to factorizations of the minimal
or characteristic of Aconsisting of any number of pairwise coprime factors:

Theorem 12 (Coprime Decomposition) Let A be any matrix in IKn×n and let α(s) denote either its
characteristic or minimal polynomial. If α1 (s), α2 (s), . . . , αm (s) are monic, polynomials satisfying

gcd{αi (s), αj (s)} = 1, ∀ i 6= j

and
α(s) = α1 (s)α2 (s) · · · αm (s),

and if Xi is a basis matrix for kernel αi (A), then the matrix


T = [ X1 X2 ··· Xm ]

is nonsingular and
 
A1 0 ··· 0
 0 A2 ··· 0 
T −1 AT = 
 ... .. .. .. 
. . . 
0 0 ··· Am

Moreover, if α(s) is the characteristic {minimal} polynomial of A, then for i ∈ {1, 2, . . . , m}, αi (s) is the
characteristic {minimal} polynomial of Ai .

Suppose A has n distinct eigenvalues λ1 , λ2 , . . . , λn . Then A’s characteristic polynomial can be written

as a product of the n pairwise coprime polynomials αi (s) = s − λi , i ∈ {1, 2, . . . , n}. In this case, the
kernel αi (A) are eigenspaces and each Ai = λi . We can thus state the following.

Corollary 1 If An×n has n distinct eigenvalues, then A is semisimple and thus diagonalizable by a similarity
transformation.

Remark 5 At this point we know that any matrix An×n with n distinct eigenvalues is both semisimple and
cyclic. Unfortunately the converse implications are not true; i.e., there are semisimple matrices and cyclic
matrices which do not have distinct eigenvalues. For example, the matrix
 
3 1
A=
0 3

is cyclic but noes not have distinct eigenvalues. And the matrix
 
3 0
A=
0 3

is semisimple but does not have distinct eighevalues. Here’s a tough question: Is there a matrix which does
not have distinct eigenvalues which is both semisimple and cyclic? More generally, is the set of n×n matrices
which have distinct eigenvalues the same as the set of n × n matrices which are both semisimple and cyclic?
7.9.1 The Jordon Normal Form

By making use of the Rational Canonical and Coprime Decomposition Theorems as they apply to any given
matrix An×n , it is possible to construct an n × n matrix T which transforms A via simalarity into a block
diagonal matrix whose diagonal blocks have the smallest sizes possible. The type of matrix we’re talking
about is call a “Jordon Normal Form.” The aim of this section is to explain what the Jordon Normal Form is
and to very briefly outline the ideas upon which the construction of the corresponding transforming matrix
T is based. In what follows, A is any given n × n matrix with elements in IK and α(s) is its characteristic
polynomial.

As a first step, let us use the fact that any polynomial with coefficients in either IR or Cl can be written
as a product of powers of prime polynomials in C[s].l For our purposes this means that α(s) can be written
as
α(s) = (s − λ1 )d1 (s − λ2 )d2 · · · (s − λq )dq
where the λi are distinct roots of α(s) over C,l {i.e., λi 6= λj if i 6= j} and di is λi ’s algebraic multiplicity.
Meanwhile the geometric multiplicity of λi is the dimension of the eigenspace of λi . Recall from elementary
algebra that such a factorization is always possible and is unique up to a reordering of the distinct roots of
α(s). The λi are, of course, the distinct eigenvalues of A and d1 + d2 + · · · dq = n.

As a second step, let us use the Coprime Decomposition Theorem to find a nonsingular matrix R such
that  
A1 0 · · · 0
 0 A2 · · · 0 
R−1 AR = 
 ... .. .. ..  (7.43)
. . . 
0 0 · · · Aq
where for each i ∈ {1, 2, . . . , q}, Ai is a square matrix with characteristic polynomial (s − λi )di .

As a third step, for each i ∈ {1, 2, . . . , q} let us use the Rational Canonical Decomposition Theorem to
find a di × di matrix Pi such that
 
Bi1 0 ··· 0
 0 Bi2 ··· 0 
Pi−1 Ai Pi = 
 ... .. .. ..  (7.44)
. . . 
0 0 · · · Bimi

where the matrix on the right is the rational canonical form of Ai . {We’ve not really explained exactly how
to do this, but this is not so important at this point.} Recall that each diagonal block in a rational canonical
form is a cyclic matrix whose characteristic polynomial divides that of its predecessor. Since (s − λi )di is
the characteristic polynomial of Ai , this means that the characteristic polynomial of the Bij must be of the
form
βij (s) = (s − λi )nij
where the nij are positive integers satisfying

ni1 ≥ ni2 ≥ · · · ≥ nimi

and
di = ni1 + ni2 + · · · + nimi

The forth step is to construct matrices Sij for which


−1
Sij Bij Sij = Jij (7.45)
where  
λi 1 0 ··· 0 0
0 λi 1 ··· 0 0
 .. 
 . 0 
0 0 λi 0
Jij =  . .. .. .. . ..  (7.46)
 . . .. 
 . . . . 
0 0 0 · · · λi 1
0 0 0 ··· 0 λi nij ×nij

The matrix Jij is called the Jordon Block of (s − λi )nij . The existence of Sij is a consequence of three
things:

1. Bij and Jij both have the same characteristic polynomial, namely (s − λi )nij .
2. Bij and Jij are both cyclic. {Cyclicity of Jij can be estabilshed by showing that with x the nij th unit
n −1 
vector in IKnij , the matrix x Jij x · · · Jijij x is nonsingular; this in turn can be verified by
checking that this matrix is upper triangular with one’s on the main diagonal.}
3. All cyclic matrices with the same characteristic polynomial are similar {c.f., Theorem 9}.

Now suppose we define


 
P1 S1 0 ··· 0
∆  0 P2 S2 ··· 0 
T = R
 ... .. .. .. 
. . . 
0 0 · · · Pq Sq
where  
Si1 0 ··· 0
 0 Si2 ··· 0 
Si = 
 ... .. .. .. 
. . . 
0 0 · · · Simi
Then using (7.43)-(7.45) we arrive at the similarity transformation

T −1 AT = J
where J is the block diagonal matrix

J 
11
 J12 
 .. 
 . 
 
 J1m1 
 
 J21 
 
 J22 
 
 .. 
∆  . 
J = 
 J2m2 
 
 .. 
 . 
 
 Jq1 
 
 Jq2 
 
 .. 
 . 
Jqmq
J is called the Jordon Normal Form of A. J’s diagonal blocks are each of the form (7.46) and there is one
such block for each polynomial (s − λi )nij , i ∈ {1, 2, . . . , q}, j ∈ {1, 2, . . . , mi }. The (s − λi )nij are called
the elementary divisors of A. They are uniquely determined by A and in turn uniquely determine A’s Jordon
Normal Form up to the ordering of its diagonal blocks. Given all this, it is not very difficult understand why
two n × n matrices are similar if and only if they both have the same set of elementary divisors.

Note that in the special case when all of A’s elementary divisors are of degree one, the Jordon Normal Form
of A is nothing more than a diagonal matrix whose diagonal entries are A’s eigenvalues. Thus semisimple
matrices are precisely those matrices whose elementary divisors all have degree one. Unfortunately the
problem of explicitly computing the elementary divisors of a matrix is a bit of an exercise. There is no
simple way to conclude that a given square matrix is semisimple.

As with invariant factors and rational canonical forms, our purpose has here has been only to make the
reader aware of the existence of elementary divisors and the Jordon Normal Form. Those individual wishing
to learn more about this topic should consult a good matrix algebra text or reference such as [1], Chapter 7.

Example 51 The matrix


 
−2 1 0 0 0 0 0 0
 0 −2 1 0 0 0 0 0
 
 0 0 −2 0 0 0 0 0
 
 0 0 0 3 0 0 0 0
A= 
 0 0 0 0 −2 1 0 0
 
 0 0 0 0 0 −2 0 0
 
0 0 0 0 0 0 −3 1
0 0 0 0 0 0 0 −3

is in Jordon Normal Form and its elementary divisors are (s + 2)3 , (s − 3), (s + 2)2 , and (s + 3)2 .

7.10 The Structure of the Matrix Exponential

For sure the main application of the Jordon Normal Form of a matrix A is in characterizing A’s state
transition matrix. The aim of this section is to explain this characterization, and to discuss some of its
implications.

Let us begin by assuming that A is a given n × n matrix, that J is its Jordon Normal Form and that T
is a nonsingular matrix such that
A = T JT −1

Then as we’ve already explained at the beginning of the chapter

etA = T etJ T −1

Now since J is a block diagonal matrix of the form


 
J1 0 ··· 0
 0 J2 ··· 0 
J =
 ... .. .. .. 
. . . 
0 0 · · · Jm
we can write
 
etJ1 0 ··· 0
 0 etJ2 ··· 0 
etJ =
 ... .. .. .. 
. . . 
0 0 · · · etJm

Thus to characterize the kinds of time functions which can appear in etA , what we really need to do it to
characterize the kinds of time functions which appear in matrix exponentials of the form etJi where Ji is
a typical Jordon Block. Whatever these time functions are, they are all going to appear in etA in linear
combination with similar time functions from J’s other Jordon Blocks. This is because for any matrix B,
each entry of the matrix T BT −1 is a linear combination of the entries of the matrix B.

Now it is quite clear that if


Ji = [ λi ]1×1

then etJi is simply the scalar exponential


etJi = etλi

Consider next the case when


 
λi 1
Ji =
0 λi 2×2

Then
 
etλi tetλi
etJi =
0 etλi

This can easily be verified by checking to see that with etJi so defined,

d tJi
e = Ji etJi , and etJi t=0
=I
dt

Similarly, for
 
λi 1 0
Ji =  0 λi 1
0 0 λi

one finds that


 t2 tλi

etλi tetλi 2e
etJi = 0 etλi tλi
te 
0 0 etλi

More generally, if

 
λi 1 0 ··· 0 0
0 λi 1 ··· 0 0
 .. 
 . 0 
0 0 λi 0
Ji =  . .. .. .. . .. 
 . . .. 
 . . . . 
0 0 0 · · · λi 1
0 0 0 ··· 0 λi mi ×mi
then  
t2 tλi t(mi −2) tλi t(mi −1) tλi
etλi tetλi 2e ··· (mi −2)! e (mi −1)! e
 
 t(mi −3) tλi t(mi −2) tλi

 0 e tλi
te tλi
··· (mi −3)! e (mi −2)! e

 
 
 t(mi −4) tλi (mi −3) 
 0 0 etλi ··· t tλi 
 (mi −4)! e (mi −3)! e 
etJi =



 .. .. .. .. .. .. 
 . . . . . . 
 
 
 
 0 0 0 ··· etλi tetλi 
 
0 0 0 ··· 0 etλi mi ×mi
mi
Thus we see that the Jordon Block of the elementary divisor (s−λi ) contains within it, linear combinations
{actually just scalar multiples} of the time functions

etλi , tetλi , . . . , t(mi −1) etλi

We’ve arrived at the following characterization of the state transition matrix of A.

Theorem 13 Let A be an n × n matrix with elementary divisors

(s − λ1 )m1 , (s − λ2 )m2 , . . . , (s − λq )mq

Then each of the entries of etA is a linear combination of the time functions

etλi , tetλi , . . . , t(mi −1) etλi , i ∈ {1, 2, . . . , q}

Thus the qualitative behavior of etA and consequently all solutions to ẋ = Ax, is completely characterized
by A’s eigenvalues and the sizes of their Jordan blocks.

7.10.1 Real Matrices

Suppose A is real, that λ is one of its eigenvalues and that m is the size of its Jordan block. Then as we’ve
just noted, linear combinations of the time functions

eλt , teλt , . . . , tm−1 eλt

must appear in etA . Suppose λ is not real; i.e.,

λ = a + jω

where a and ω are real numbers and ω 6= 0. Then

λ∗ = a − jb

must also be an eigenvalue of A with multiplicity m. Because A is real, it can be shown that the linear
combinations of the time functions
∗ ∗ ∗
eλt , teλt , . . . , tm−1 eλt , eλ t , teλ t , . . . , tm−1 eλ t

which appear in etA can also be written as linear combinations of time functions of the forms

eat cos ωt, eat sin ωt, teat cos ωt, teat sin ωt, . . . , tm−1 eat cos ωt, tm−1 eat sin ωt
7.10.2 Asymptotic Behavior

Let A be either real or complex valued. On the basis of Theorem 13 we can draw the following conclusions.

1. If all of A’s eigenvalues have negative real parts, then ||etA || → 0 exponentially fast, as t → ∞3 . Such
A are said to be asymptotic stability matrices.
2. If at least one eigenvalue of A
(a) has a positive real part a, then
||etA ||
must grow without bound, as fast as eat , as t → ∞.
(b) has a zero real part and a Jordan block of size m > 1, then
||etA ||
must grow without bound as fast tm−1 , as t → ∞.
In either case, such A are said to be unstable matrices.
3. If all of A’s eigenvalues have nonpositive real parts, and those with zero real parts all have Jordan
blocks of size 1, then
||etA ||
will remain finite as t → ∞ but will not approach 0. Such A are said to be stable matrices.

The key observation here is summarized as follows:

Theorem 14 A n × n matrix A is asymptotically stable; i.e.,


lim etA = 0
t→∞

if and only if all of A’s eigenvalues have negative real parts.

7.11 Linear Recursion Equations

Just about everything we’ve discussed in the preceding section for the linear differential equation
ẋ = Ax
extends an a natural way to a linear recursion equation of the form
x(t + 1) = Ax(t), t = 0, 1, 2, . . .
For example, for the recursion equation we can write
x(t) = At−τ x(τ ), ∀t, τ ≥ 0
For this reason At−τ is called the {discrete-time} state transition matrix of A. Note that if J is A’s Jordan
normal form and A = T JT −1, then
At = T J t T −1 ,
so the asymptotic behavior of At is characterized by A’s eigenvalues, just as in the continuous time case
discussed above. We leave it to the reader to verify the following:
3 Note that for any finite integer i ≥ 0, and any positive number µ, ti e−µt → 0 as t → ∞.
1. If all of A’s eigenvalues have magnitude less than 1, then ||At || → 0 exponentially fast, as t → ∞4 .
Such A are said to be discrete-time asymptotic stability matrices.
2. If at least one eigenvalue of A
(a) has magnitude a greater than 1, then
||At ||
must grow without bound, as fast as |a|t , as t → ∞.
(b) has magnitude 1 and a Jordan block of size m > 1, then

||At ||

must grow without bound as fast tm−1 , as t → ∞.


In either case, such A are said to be discrete-time unstable matrices.
3. If all of A’s eigenvalues have magnitudes no greater than 1, and those with magnitude 1 each have
Jordan blocks of size 1, then
||At ||
will remain finite as t → ∞ but will not approach 0. Such A are said to be discrete-time stable matrices.

The discrete-time version of Theorem 14 is as follows:

Theorem 15 An n × n matrix A is discrete-time asymptotically stable; i.e.,

lim At = 0
t→∞

if and only if all of A’s eigenvalues have magnitudes less than 1.

NOTE: Throughout this section t is a discrete variable taking values in the set of natural numbers
{0, 1, 2, 3, . . .}!

Later in these notes we will develop alternative methods for testing for discrete and continuous-time
stability of a matrix A, which do not require one to compute the eigenvalues of A.

4 Note that for any finite integer i ≥ 0, and any real number µ with magnitude less than 1, ti µt → 0 as t → ∞.
Chapter 8

Inner Product Spaces

In Chapter 4 we briefly discussed the concepts of a normed vector space. We saw that by exploiting the idea
of a norm, one could extend from the set of real numbers to normed vector spaces, a number of concepts
from real analysis or calculus such as continuity, convergence, openness, boundness, etc. The concept of a
normed space is a special case of a more general concept known as a “metric space1 .” On the other hand,
normed spaces are generalizations of “inner product spaces.” The aim of this chapter is to briefly discuss
some of the basic properties of inner product spaces.

8.1 Definition

Let V be a {possibly infinite dimensional} vector space over IK. A function (·) mapping domain V × V into
codomain IK is said to be an inner product on V if for all x, y, z ∈ X and all µ ∈ IK

1. (x, y) = (y, x)∗


2. (µx, y) = µ∗ (x, y)
3. (x + y, z) = (x, z) + (y, z)

∆ √ ∆
Here ∗ denotes complex conjugation; i.e., for a, b ∈ IR and j = −1, (a + jb)∗ = a − jb.

Let us note that requirements 1 to 3 have the following consequences:

2′ . (x, µy) = µ(x, y)


3′ . (x, y + z) = (x, y) + (x, z)

What’s more, whether IK is IR or C,


l it follows from 1. that for every vector x, the inner product (x, x) is a
real number. For our purposes, we shall impose the additional requirement that (x, x) be positive definite;
i.e.,
1 A metric space is a real or complex vector space X on which is defined for every pair of vectors x and y, a real, scalar-valued

“distance function” or “metric” d(x, y) which is a measure of the “distance” between x and y. For such a function to be a
metric, it (i) must be nonnegative-valued, (ii) must equal 0 just in case x = y, (iii) must satisfy d(x, y) = d(y, x) for all x, y ∈ X
and (iv) must satisfy the triangle inequality d(x, y) ≤ d(x, z) + d(z, y) for every x, y, z ∈ X . If IK = C,
l then X is called a unitary
space with Hermitian metric. If IK = IR, then X is called a Euclidean space with Euclidean metric. The metric on a normed

space is simply d(x, y) = ||x − y||.

131
4. (x, x) > 0 for all x 6= 0

5. (x, x) = 0 if x = 0

In view of properties 4 and 5 it makes sense to define the norm of x by


∆ p
||x|| = (x, x)

It can be shown that properties 1 - 5 make V into a normed vector space.

Example 52 Suppose V = Cl n and (x, y) is defined so that

(x, y) = (x∗ )′ y

where as usual, prime denotes transpose. It is easy to verify that properties 1 - 5 hold and that V is a unitary
space. Note that if
 
x1
 x2 
x= 
 ... 
xn
then
n
X n
X
||x||2 = (x∗ )′ x = x∗i xi = |xi |2
i=1 i=1

where |xi | denotes the magnitude of xi .

Example 53 Suppose V = IRn and (x, y) is defined so that

(x, y) = x′ y

It is quite easy to see that properties 1 - 5 hold. Thus in this case V is a Euclidean space2 .

8.1.1 Triangle Inequality

Note that in IR2 , ||x − y|| can be thought of as the distance between the directed line segment x and the
directed line segment y as shown in Figure 8.1.

|| x
-y
||
y x-y

Figure 8.1:
2 The reader should not confuse the inner product of x and y, namely the scalar x′ y, with the outer product of x and y, namely

the n × n matrix xy ′ . Outer products play an important role in quantum mechanics. Note for example that xy ′ zw ′ = (y ′ z)xw ′ .
Manipulations such as this arise in connection with Dirac notation.
If x, y and z are arbitrary vectors in IR2 , then we can make use of well known properties of triangles to
conclude that the triangle inequality

||x − y|| + ||y − z|| ≥ ||x − z||

is valid for any three vectors in any inner product space of any dimension {cf, Figure 8.2}. The inequality
can be derived directly from the defining properties 1 – 4 above.

y x-y y-z

x-z
x

Figure 8.2:

8.2 Orthogonality

By far the most important feature of an inner product space which distinguishes it from a more general
normed space, is the concept of orthogonality. In the sequel we define orthogonality and discuss some of its
consequences.

8.2.1 Orthogonal Vectors

Two vectors x and y in an inner product space are said to be orthogonal if their inner product equals zero.
In other words
x and y orthogonal ⇐⇒ (x, y) = 0
Note also that the identity
||x + y||2 = ||x||2 + ||y||2 + (x, y) + (y, x)
implies that
||x + y||2 = ||x||2 + ||y||2
whenever x and y are orthogonal. This equation, is valid for orthogonal vectors in any inner product space
- it can be viewed as a generalization of Pythagoras’s Theorem for right triangles.

8.2.2 Orthonormal Sets



A set vectors Y = {y1 , y2 , . . . , yn } whose elements all satisfy ||yi || = 1, is said to be normalized. A set of

vectors X = {x1 , x2 , . . . , xm } whose elements are pairwise orthogonal {i.e., (xi , xj ) = 0, ∀i 6= j} is called
an orthogonal set. It is possible to normalize and such orthogonal set, without changing its span, by simply
replacing each xi by the normalized vector
∆ 1
x̄i = xi
||xi ||
Note that each such x̄i satisfies
s  s
p 1 1 1
||x̄i || = (x̄i , x̄i ) = xi , xi = (xi , xi ) = 1
||xi || ||xi || ||xi ||2


so X̄ = {x̄1 , x̄2 , . . . , x̄m } is both an orthogonal and a normalized set. Such sets are called orthonormal. X
and X̄ both span the same subspace because each x̄i is just a nonzero scalar multiply of xi .

8.3 Gram’s Criterion

Let {x1 , x2 , . . . xm } be a finite set of vectors in a possibly infinite dimensional inner product space X . Our
aim is to develop a simple test for deciding whether or not {x1 , x2 , . . .} is a linearly independent set.

Suppose that {x1 , x2 , . . .} is linearly dependent. Then there must be numbers ai , not all zero, such that

a1 x1 + a2 x2 + · · · + am xm = 0 (8.1)

Clearly
(xi , {a1 x1 + a2 x2 + · · · + am xm }) = 0, i ∈ {1, 2, . . . , m}
Therefore
a1 (xi , x1 ) + a2 (xi , x2 ) + · · · + am (xi , xm ) = 0 i ∈ {1, 2, . . . , m} (8.2)
In matrix form, this is equivalent to
Ga = 0 (8.3)
where

G = [ (xi , xj ) ]m×m (8.4)
and  
a1
∆  a2 
a= 
 ... 
am
But a 6= 0 by hypothesis, so G must have linearly dependent columns. Since G is a square matrix, G must
therefore be singular. In other words, linear dependence of the set {x1 , x2 , . . . , xm } implies that G must be
singular.

Conversely, suppose that the matrix G defined by (8.4) is singular. Then (8.3) must have a nonzero
solution a so there must be numbers ai , not all zero, for which (8.2) holds. Multiplying the ith equation in
(8.2) by a∗i we obtain
a∗i a1 (xi , x1 ) + a∗i a2 (xi , x2 ) + · · · + a∗i am (xi , xm ) = 0
Therefore
(ai xi , {a1 x1 + a2 x2 + · · · + am xm }) = 0, i ∈ {1, 2, . . . , m}
Summing these m equations thus yields

||a1 x1 + a2 x2 + · · · am xm ||2 = 0
or equivalently
a1 x1 + a2 x2 + · · · am xm = 0
Since the ai are not all zero, {x1 , x2 , . . . , xm } must be a dependent set. In other words, if G is singular,
{x1 , x2 , . . . , xm } must be a dependent set.

The matrix G defined by (8.4) is called the Gramian of {x1 , x2 , . . . , xm }. We’ve proved the following.

Theorem 16 A set of vectors {x1 , x2 , . . . , xm } in an inner product space is linearly independent if and only
if its Gramian is a nonsingular matrix.


Example 54 Note that with the inner product (x, y) = x′ y, the Gramian of the columns of the real matrix
Am×n is A′ A. Thus the columns of A are linearly independent if and only if the n × n matrix A′ A is
nonsingular.

Example 55 Let f1 (t), f2 (t), . . . , fn (t) be continuous, real scalar-valued functions of t defined on the interval
[0, 1]. The fi can be viewed as vectors in the infinite dimensional vector space of continuous, real scalar
functions defined on [0, 1]. We can make this vector space into a Euclidean space by introducting the inner
product 3 Z 1

(f, g) = f (t)g(t)dt
0

The Gramian of {f1 , f2 , . . . , fn } is therefore the {constant} matrix


R1 R1 R1 
0 f1 (t)f1 (t)dt 0 f1 (t)f2 (t)dt · · · 0 f1 (t)fn (t)dt
 .. .. .. 
G= . . · · · . 
R1 R1 R1
f (t)f1 (t)dt 0 fn (t)f2 (t)dt · · · 0 fn (t)fn (t)dt
0 n

The set {f1 , f2 , . . . , fn } is thus linearly independent if and only if det G 6= 0.

8.4 Cauchy-Schwartz Inequality

Suppose that x and y are vectors in an inner product space. Observe that the inequality

||||y||2 x − (y, x)y||2 ≥ 0

expands out to
||y||4 ||x||2 − ||y||2 {(x, (y, x)y) + ((y, x)y, x)} + |(y, x)|2 ||y||2 ≥ 0 (8.5)
But since (x, (y, x)y) = |(y, x)|2 = ((y, x)y, x), (8.5) can be rewritten as

||y||2 (||y||2 ||x||2 − |(y, x)|2 ) ≥ 0 (8.6)

Whether ||y|| = 0 or not, (8.6) implies that

||y||2 ||x||2 ≥ |(y, x)|2

which when rewritten as


||y||||x|| ≥ |(y, x)|
is called the Cauchy-Schwartz Inequality.
3 As ∆ ∆ p
always with an inner product space, d(f, g) = ||f − g|| and ||f || = (f, f ).
8.5 Orthogonalization of a Set of Vectors

Let {x1 , x2 , . . . , xm } be a set of vectors in an inner product space - and write S = span {x1 , x2 , . . . , xm }.
Our aim is to use to construct from the xi , another set {y1 , y2 , . . . , ym } which also spans S and which, in
addition, is orthogonal. The procedure for doing this is as follows.

1. Set
y1 = x1
2. Set
(y1 , x2 )
y2 = x2 − y1
(y1 , y1 )
·
·
·
i. Set
(y1 , xi ) (y2 , xi ) (yi−1 , xi )
yi = xi − y1 − y2 − · · · − yi−1
(y1 , y1 ) (y2 , y2 ) (yi−1 , yi−1 )
Note that
(y1 , x2 )
(y1 , y2 ) = (y1 , x2 ) − (y1 , y1 ) = 0
(y1 , y1 )
so y1 and y2 are orthogonal. Similarly, for j < i

(y1 , xi ) (y2 , xi ) (yi−1 , xi )


(yj , yi ) = (yj , y1 ) − (yj , y2 ) − · · · − (yj , yi−1)
(y1 , y1 ) (y2 , y2 ) (yi−1 , yi−1

Since yj is orthogonal to all of the yk for k < i and k 6= j, it follows that

(yj , xi )
(yj , yi ) = (yj , xi ) − (yj , yj ) = 0
(yj , yj )

Therefore the set {y1 , y2 , . . . , ym } is orthogonal.

Let us note finally that

span {y1 } = span {x1 }


span {y1 , y2 } = span {x1 , x2 }
·
·
·
span {y1 , y2 , . . . , yi } = span {x1 , x2 , . . . , xi }

Thus
span {y1 , y2 , . . . , ym } = span {x1 , x2 , . . . , xm }

Example 56 Let X be the vector space of all piecewise-continuous, real-valued scalar functions defined on
the closed interval [−1, 1]. The inner product
Z 1

(f, g) = f (t)g(t)dt, ∀f, g ∈ X
−1
makes X a Euclidean space. Using Gram’s test, one can then verify that {1, t, t2 , . . . , tm } is a linearly
independent set of vectors in this space. This set can be orthogonalized as follows:

Set

y1 = 1
!
R1
∆ τ dτ
y2 = t − R−1
1 y1 = t
−1
τ dτ
R1 2 ! R1 !
∆ τ dτ τ 3 dτ
y3 = t2 − R−1
1 y1 − −1
R1 y2
−1 τ dτ −1 τ 2 dτ

etc. It is easy to verify that the yi coincide up t scale factor with the well known Legendre polynomials

1 d(i−1) 2
{t − 1}(i−1) , i ∈ {1, 2, . . .}
2(i−1) {(i − 1)!} dt(i−1)

By changing the definition of the inner product on X , we can develop other sets of orthogonal functions.
For example, by orthogonalizing the set {1, t, . . .} using the inner product
Z 1
∆ 1
(f, g) = √ f (t)g(t)dt, f, g ∈ X
−1 1 − t2

one obtains the Chebyshev polynomials


1
cos(i arccos{t}), i ∈ {1, 2, . . .}
2(i−1)

8.6 Orthogonal Projections

Let {x1 , x2 , . . . , xm } be a finite set of orthogonal vectors in a {possibly infinte dimensional} inner product

space X . Let S = span {x1 , x2 , . . . xm }. Note that S is finite dimensional because it is spanned be a finite
set of of vectors. Let x be any given vector in X . We claim that there are unique vectors s and y such that

1. x = s + y

2. s ∈ S

3. y is orthogonal to every vector in S; i.e., (y, xi ) = 0, i ∈ {1, 2, . . . , m}.

We call s the orthogonal projection of x on S.

To construct s, we first form the projection of x on the span of each xi :

(x, xi )
xi
(xi , xi )

The definition of s is then simply


Xm
∆ (x, xi )
s= xi
i=1
(xi , xi )

Note that property 2 above must automatically be true because s is a linear combination of the xi .
Set

y =x−s
thereby establishing property 1. Note that for each j ∈ {1, 2, . . . , m},
m
!
X (x, xi )
(y, xj ) = (x, xj ) − (s, xj ) = (x, xj ) − xi , xj
i=1
(xi , xi )

But {x1 , x2 , . . . xm } is an orthogonal set so (xi , xj ) = 0 for i 6= j. Therefore

(x, xj )
(y, xj ) = (x, xj ) − (xj , xj ) = 0
(xj , xj )

Thus y is orthogonal to each vector in a basis for S. This implies that y is orthogonal to each vector in S.
In other words, property 3 above is true.


Example 57 In IR3 suppose that S = span {x1 , x2 } where
   
1 0
x1 =  1  and x2 =  0 
0 1

Then x1 and x2 are orthogonal since x′1 x2 = 0. If


 
1

x = 0
1

then 1
2
     
x′ x1 x′ x2 1 1  
s= x1 + x2 = x1 + x2 =  12 
x′1 x1 x′2 x2 2 1  
1
and  1 
2
 
 
y = x − s =  − 21 
 
0

8.7 Bessel’s Inequality

Now suppose {x1 , x2 , . . . , xm } is an orthonormal set in some inner product space X - and let x be any vector

in X . Then the orthogonal projection of x on the subspace S = span {x1 , x2 , . . . , xm } is
m
X
s= ai xi
i=1

where

ai = (x, xi )
3

x
1 2 x

1
1 1
2

1
2

1 x
1
2

We can write
||x||2 = ||s + x − s||2 = ||s||2 + ||x − s||2 + (s, x − s) + (x − s, s)

But s and x − s are orthogonal so


||x||2 = ||s||2 + ||x − s||2

Therefore
||x||2 ≥ ||s||2

Moreover, since {x1 , x2 , . . . , xm } is an orthonormal set


m
X
||s||2 = |ai |2
i=1

Thus
m
X
||x||2 ≥ |ai |2 (8.7)
i=1

which is called Bessel’s inequality. Note that (8.7) holds with an equal sign whenever x ∈ span {x1 , x2 ,
. . . , xm }.

8.8 Least Squares

Suppose that X is an infinite dimensional inner product space and that x1 , x2 , . . . is an infinite sequence of

orthonormal vectors. Let x be any fixed vector in X with finite norm ||x||. Define ai = (x, xi ), i ≥ 1. Since
(8.7) holds for every m ≥ 1, the sequence of numbers
m
X
|ai |2 , m ∈ {1, 2, . . .}
i=1
is bounded from above by ||x||. From this it follows that the infinite series

X
|ai |2
i=1

converges, and moreover that



X
|ai |2 ≤ ||x||2
i=1

It turns out that for any fixed m > 0, the minimum of the norm square
m
X
||x − bi xi ||2
i=1


is attained when bi = ai , i ∈ {1, 2, . . . , m}. {Verify this !} Thus the projection of x on Sm = span {x1 , x2 ,
. . . , xm }, namely

sm = a1 x1 + a2 x2 + · · · + am xm ,
can be thought of as that vector in Sm which “best” approximates x in the “least squares” sense. The
corresponding least square error is
m
X
e2m = ||x − ai xi ||2
i=1
( m
) ( m
)!
X X
= x− ai xi , x− ai xi
i=1 i=1
m m
! m
! m
!
X X X X
= (x, x) + ai xi , ai xi − x, ai xi − ai xi , x
i=1 i=1 i=1 i=1
m
X Xm
2
= ||x|| + |ai |2 − 2 |ai |2
i=1 i=1

or
m
X
e2m = ||x||2 − |ai |2
i=1

In the event that


lim e2m = 0,
i→∞

the infinite series



X
ai xi
i=1

is said to converge in the mean to the vector x; if this is so,



X
||x||2 = |ai |2
i=1

If for each x ∈ X , the corresponding series



X ∆
ai xi {ai = (x, xi )}
i=1
converges, then the sequence of vectors x1 , x2 , . . . is called complete and X is a complete vector space.
Complete inner product spaces such as this are usually called Hilbert spaces 4 . Thus in a Hilbert space, the
sequence of least squares approximations s1 , s2 , . . . of any given vector x, converges to x at m → ∞.

Example 58 Let X denote the Cl vector space of all complex-valued, scalar functions f that are defined
and piecewise continuous on the closed interval [0, 2π]. Include in the definition of X , the inner product
Z 2π

(f, g) = f (t)g ∗ (t)dt
0

where ∗ denotes complex conjugate. Consider the infinite sequence of functions


1
√ ejkt , k = 0, ±1, ±2, . . .

∆ √
where j = −1. Note that these functions form an orthonormal set since
Z 2π (
1 0 for i 6= k
(ejkt , ejit ) = ej(i−k)t dt =
2π 0 1 for i = k

In can be shown that this set of functions is also complete. Thus if f is any given function in X , then it is
possible to express f as an ‘infinite linear combination’ of the preceding exponentials. In other words,
k=∞
X
f (t) = ak ejkt
k=−∞

in the sense that


k=m
X
lim ||f = ak ejkt || = 0
m→∞
k=−m

where Z 2π
∆ 1
ak = (f, e jkt
)= √ f (t)e−jkt dt, k = 0, ±1, ±2, . . .
2π 0

The preceding expression for f is the Fourier series expansion of f and the ak are called the Fourier
coefficients.

In this chapter we have briefly addressed a number of basic topics from an area of mathematics known as
“functional analysis.” All that we’ve discussed applies to vector spaces of both finite and infinite dimensions.
In the sequel we will focus exclusively on the finite dimensional case.

4 Recall that a complete normed spaced is called a Banach space.


Chapter 9

Normal Matrices

For the remainder of these notes we shall assume that X is either the unitary space Cl n with inner product
(x∗ )′ y or the Euclidean space IRn with inner product x′ y. Since in either case X has a basis, we can
orthogonalize and normalize such a basis to obtain an orthonormal basis for X .

Suppose that {x1 , x2 , . . . , yn } and {y1 , y2 , . . . , yn }are two orthonormal bases for X . Then there must be
a nonsingular matrix T such that
T xi = yi , i ∈ {1, 2, . . . , n} (9.1)
These n equations are equivalent to the single matrix equation
TX = Y (9.2)
∆ ∆
where X = [ x1 x2 · · · xn ] and Y = [ y1 y2 · · · yn ]. Thus
T = Y X −1 (9.3)
In the sequel it will be shown that T, X, X −1 and Y are all examples of “orthogonal matrices” if X = IRn
or “unitary matrices ” if X = Cl n . Orthogonal and unitary matrices as well as “Hermitian” and symmetric
matrices are examples of “normal matrices.” The aim of this chapter is define these matrices different matrix
types and to discuss their properties.

9.1 Algebraic Groups



A group G = {G, ◦} is an algebraic system consisting of a set of elements G together with a single operation
◦ between pairs of elements of G under which the set is closed, and in addition, certain axioms hold. In
particular, for all g1 , g2 ∈ G, ◦ must be defined so that g1 ◦ g2 ∈ G and also so that

i. for all gi ∈ G, g1 ◦ (g2 ◦ g3 ) = (g1 ◦ g2 ) ◦ g3 {i.e., ◦ is associative}.


ii. there is an element e ∈ G, called an identity or unit such that e ◦ g = g ◦ e = g for all g ∈ G.
iii. for each g ∈ G, there is an inverse in G, written g −1 such that g ◦ g −1 = e.

If, in addition, the group operation ◦ is commutative {i.e., g1 ◦ g2 = g2 ◦ g1 , ∀gi ∈ G} then G is called either
a commutative or Abelian group. {What is purple and commutes?}1 It is common practice to write g ∈ G
1 An Abelian grape.

143
whenever one means g ∈ G. There is a quite extensive theory of groups. Algebraic groups arise in many
applications including solid state and atomic physics, the study of nonlinear dynamics, and communications.
There are more advanced course devoted solely various topics in group theory.

Example 59

1. The set of all n × n nonsingular, real-valued matrices together with matrix multiplication is a
{noncommutative} group called the general linear group. For this group, In×n is the identity.

2. The set of all n × n matrices with exactly one 1 in each row and in each column and 0’s elsewhere,
together with matrix multiplication is the n-dimensional permutation group. In×n is also the identity
for this group.

3. The set of all n × n nonsingular, diagonal, real-valued matrices together with matrix multiplication is
an Abelian group. Again for this group, In×n is the identity.

4. The set of all n × m , real-valued matrices together with matrix addition is an Abelian group For
this group, 0n×m is the identity. The inverse of a matrix M in this group {with respect to the group
operation of addition} is simply −M .

9.1.1 Orthogonal Group

A real-valued, nonsingular matrix is said to be orthogonal if its inverse is equal to its transpose. In other
words
A is orthogonal ⇐⇒ A′ = A−1
Let On denote the set of all n × n orthogonal matrices. We claim that On is a multiplicative group.

i. Matrix multiplication is associative. Moreover On is closed under multiplication: in particular, if A


and B are in On , then
(AB)′ = B ′ A′ = B −1 A−1 = (AB)−1

ii. In×n is in On and serves as the group’s identity.

iii. If A ∈ On and B = A−1 , then


B ′ = (A−1 )′ = (A′ )′ = A = B −1
so B ′ = B −1 . Thus B is orthogonal and therefore in On . Thus A−1 serves as the group inverse of A.

Thus On is a group.

Let us note that a given real matrix An×n is orthogonal if and only if

A′ A = In×n (9.4)

Writing A as A = [ a1 a2 · · · an ], we therefore see that


 ′ 
a1 a1 a′1 a2 · · · a′1 an
 .. 
A′ A =  ... ..
.
..
. . 
a′n a1 a′n a2 · · · a′n an
is the Gramian of the columns of A. I clearly follows that (9.4) holds just in case {a1 , a2 , . . . , an } is an
orthonormal set. Therefore a real-valued square matrix is orthogonal if and only if its columns form an
orthonormal set.

Returning to (9.2) above, we see that if X = IRn , the matrices X and Y are orthogonal. Since On is a
group, it follows from (9.3) that T is orthogonal as well.

9.1.2 Unitary Group

A complex-valued, nonsingular matrix A is called a unitary matrix if the inverse of A equals the transpose
of A∗ , A∗ being the matrix whose elements are the complex conjugates of the corresponding elements of A;
i.e., A∗ = [ a∗ij ] where A = [ aij ]. A∗ is called the complex conjugate of A. Thus for A to be unitary means
that
A−1 = (A∗ )′
It is easy to verify that the set of all n × n unitary matrices together with ordinary matrix multiplication is
a group. This group is called the unitary group.
∆ ∆
Now suppose that X = Cl n equiped with the inner product (x, y) = (x∗ )′ y. Using a Gramian, the reader
should verify that a square matrix A is unitary if and only if its columns form an orthonormal set with
respect to this inner product. In (9.2) above, X and Y are thus unitary matrices as is T in (9.3) because
the set of n × n unitary matrices from a multiplicative group.

9.2 Adjoint Matrix


∆ ∆
Again let X denote either IRn or Cl n and define (x, y) == x′ y if X = IRn or (x, y) = (x∗ )′ y if X = Cl n . In
either case, the adjoint of A, written Aa , is that unique matrix such that
(Ax, y) = (x, Aa y) (9.5)
It is easy to verify that (
a A′ if X = IRn
A = (9.6)
(A∗ )′ if X = Cl n
Observe from (9.6) that a real-valued {resp. complex-valued} matrix is orthogonal {resp. unitary} if and
only if it equal to the inverse of its adjoint.

Note from (9.5) that orthogonal and unitary matrices are isometric in the sense that they preserve metric
or length. In other words Aa = A−1 implies that
p p p
||Ax|| = (Ax, Ax) = (x, Aa Ax) = (x, x) = ||x||
and conversely. Said differently, for all x the length of Ax equals the length of x just in case A is an orthogonal
{or unitary} matrix.

9.2.1 Orthogonal Complement

With X as above, let S be a subspace of X . The orthogonal complement of S, written X ⊥ , is the set of all
vectors in X which are orthogonal to all vectors in S. Thus

S ⊥ = {x : (x, s) = 0, ∀s ∈ S}
It is easy to verify that S ⊥ is a subspace of X . Since all vectors in S ⊥ are orthogonal to all vectors in S, it
must be that S ⊥ and S are independent subspaces. What’s more, if x is any given vector in X , {see §8.6}
we know how to construct vectors s ∈ S and y ∈ S ⊥ such that x = s + y. This implies that X ⊂ S + S ⊥
and therefore that X = S + S ⊥ . But since S and S ⊥ are independent subspaces, this sum must be direct:

X = S ⊕ S⊥ (9.7)

9.2.2 Properties of Adjoint Matrices

∆ ∆
Again let X denote either IRn or Cl n and define (x, y) = x′ y if X = IRn or (x, y) = (x∗ )′ y if X = Cl n . Let A
be a fixed matrix in IRn×n if X = IRn or in Cl n×n if X = Cl n . Suppose that S is an A-invariant subspace
and let S ⊥ denote the orthogonal complement of S. What we want to prove is that S ⊥ is Aa -invariant. In
other words, we want to show that
AS ⊂ S =⇒ Aa S ⊥ ⊂ S ⊥ (9.8)
Following the standard game plan for establishing such an implication, let y be any vector in S ⊥ . Then y is
orthogonal to each vector in S; i.e.,
(y, s) = 0, ∀s ∈ S
But since S is A-invariant, As ∈ S for all s ∈ S. Therefore

(y, As) = 0, ∀s ∈ S

But (y, As) = (Aa y, s) so


(Aa y, s) = 0 ∀s ∈ S
Thus Aa y is orthogonal to every vector s ∈ S which means that Aa y ∈ S ⊥ . We have shown that for each
vector y ∈ S ⊥ , Aa y ∈ S ⊥ . Therefore (9.8) is true. Using similar reasoning, it can be easily shown that the
reverse implication in (9.8) is also true.

What we want to show next, is that if A and Aa have a common eigenvector, then the corresponding
eigenvalues must be complex conjugates. In other words, we want to prove that
 
Ax = λx
and x 6= 0 =⇒ λ∗ = µ (9.9)
Aa x = µx

To prove that this is so, let us note that Ax = λx and Aa x = µx imply

(x, Ax) = (x, λx) = λ∗ ||x||2 and (Aa x, x) = (µx, x) = µ||x||2

respectively. But (Aa x, x) = (x, Ax) and ||x||2 6= 0 so λ∗ = µ as claimed.

9.3 Normal Matrices

So far in our discussion of metric spaces we’ve encountered two special types of matrices, namely orthogonal
and unitary. These matrices are examples of a more general class of matrices, called ‘normal matrices’ which
have especially important properties worth understanding. The aim of this section is to explain what these
properies are and why normal matrices possess them. We begin with the definition of a normal matrix.

A square matrix A defined over either IR or Cl is said to be normal if it commutes with its adjoint. Thus
A is normal just in case
AAa = Aa A
Since both orthogonal and unitary matrices equal the inverses of their respective adjoints, it follows that both
are normal matrices. There are two other types of normal matrices of particular importance in engineering
and the sciences:

1. Symmetric matrices: A matrix A with elements in IR is symmetric if it equals its adjoint; i.e.,

A ∈ IRn×n is symmetric ⇐⇒ A = Aa

Note that for such matrices, Aa = A′ , so A is symmetric just in case it equals its transpose.
2. Hermitian matrices: A matrix A with elements in Cl is Hermetian if it equals its adjoint; i.e.,

A ∈ Cl n×n is Hermitian ⇐⇒ A = Aa

Note that for such matrices, Aa = (A∗ )′ , so A is Hermitian just in case it equals its conjugate transpose.

9.3.1 Diagonalization of a Normal Matrix

All normal matrices {e.g., unitary, orthogonal, symmetric, Hermitian} have a special property. In particular,
each can be diagonalized by means of a similarity transformation determined by a unitary matrix. We now
explain why this is so.

We begin by proving the following:

Claim: Any two n × n matrices which commute have a common eigenvector.

To prove this, let An×n and Bn×n be any two square matrices satisfying

AB = BA (9.10)

Suppose that x is an eigenvector of A; i.e., x 6= 0 and

Ax = λ0 x (9.11)

for some {possibly complex} number λ0 . It follows from (9.10) and (9.11) that

AB k x = λ0 B k x, k = 0, 1, 2, . . . (9.12)

so for each k ≥ 0, B k x is an eigenvector of A {provided B k 6= 0}. Indeed, any nonzero vectors in the B-cyclic
subspace
span {x, Bx, . . . , B n−1 x}
is an eigenvector of A. Let α(s) be the minimal polynomial of x relative to B. Therefore

α(B)x = 0 (9.13)

Let λ1 be a root of α(s). Then α(s) can be written as

α(s) = (s − λ1 )β(s) (9.14)

for some polynomial β(s). What’s more, β(A)x 6= 0 because α(s) is the minimal polynomial of x and

deg β(s) < deg α(s). Set y = β(A)x. Thus y is nonzero; moreover y is in the span of the set of vectors
{x, Bx, . . . , B n−1 x}. Therefore y is an eigenvector of A. Furthermore, from (9.13) and (9.14) and the
definition of y it follows that

(B − λ1 I)y = (B − λ1 I)β(B)x = α(B)x = 0


so y is an eigenvector of B. Thus the preceding claim is true.

Now suppose that An×n is a normal matrix. This means that A must commute with Aa . But we’ve just
proved that any two matrices which commute must have a common eigenvector. Therefore there must be a
nonzero vector x1 and {possibly complex} numbers λ1 and µ1 such that
Ax1 = λ1 x1 and Aa x1 = µ1 x1 (9.15)
Because of this, the subspace S1 spanned by x1 is both A-invariant and Aa -invariant. But the Aa -invarance
of S1 together with (9.8) imply that S1⊥ must be A-invariant as well. In summary
AS1 ⊂ S1 , AS1⊥ ⊂ S1⊥ , and S1 ⊕ S1⊥ = X
The direct sum was noted before in (9.7).

Let u1 be a normalized version of x1 and let {u2 , u3 , . . . , un } be an orthonormal basis for S1⊥ . Since
S1 ⊕ S1⊥ = X , {u1 , u2 , u3 , . . . , un } must be an orthonormal basis for X . Therefore

U1 = [ u 1 u2 · · · un ]
must be a unitary matrix. Moreover since AS1 ⊂ S1 , and AS1⊥ ⊂ S1⊥ , it must be true that
 
λ 0
AU1 = U1 1 (9.16)
0 A1
where λ1 is the eigenvalue of A associated with eigenvector x1 and A1 is the unique solution to the linear
equation
A [ u2 u3 · · · un ] = [ u2 u3 · · · un ] A1 (9.17)
Clearly (9.16) can be written as  
λ1 0
U1−1 AU1 =
0 A1
We claim that A1 is a normal matrix. To prove this, let us first note that
 a   a
λ1 0 λ1 0
= = {U1−1 AU1 }a = U1a Aa (U1−1 )a = U1−1 Aa U1
0 Aa1 0 A1
Thus
  
λ1 0 λa1 0
= (U1−1 AU1 )(U1−1 Aa U1 )
0 A1 0 Aa1
= U1−1 AAa U1
= U1−1 Aa AU1
= (U1−1 Aa U1 )(U1−1 AU1 )
 a  
λ1 0 λ1 0
=
0 Aa1 0 A1
Clearly A1 Aa1 = Aaa A1 so A1 is normal. To summarize, we’ve just shown that for any normal matrix An×n ,
there must be a n × n unitary matrix U1 , an eigenvalue λ1 of A and a (n − 1) × (n − 1) normal matrix A1
such that  
−1 λ1 0
U1 AU1 =
0 A1

Because A1 is normal, it is possible to repeat this process. In particular there must be a (n − 1) × (n − 1)


unitary matrix V2 , an eigenvalue λ2 of A1 and a (n − 2) × (n − 2) normal matrix A2 such that
 
−1 λ2 0
V2 A1 V2 =
0 A2
It follows that if we define  
1
∆ 0
W2 =
0 V2 n×n
then W2 must be unitary and  
  λ1 0 0
λ1 0
W2−1 W2 =  0 λ2 0 
0 A1
0 0 A2

Thus with the unitary matrix U2 = U1 W2 we can write
 
λ1 0 0
U2−1 AU2 =  0 λ2 0 
0 0 A2
Clearly this process can be continued for a total of n steps. Thus we have proved that for any normal matrix
A there must be a unitary matrix U which diagonalizes U −1 AU .

The converse is also true. In particular, suppose that A is now any n × n matrix and that U is a unitary
matrix which diagonalizes U −1 AU . The because U −1 AU is diagonal
(U −1 AU )a (U −1 AU ) = (U −1 AU )(U −1 AU )a (9.18)
But
U −1 Aa AU = U −1 Aa U U −1 AU = (U −1 AU )a (U −1 AU )
From this and (9.18) it follows that
U −1 Aa AU = (U −1 AU )(U −1 AU )a = U −1 AU U −1 Aa U = U −1 AAa U
This implies that Aa A = AAa so A is normal. We have proved the following important theorem.

Theorem 17 A matrix A with elements in either IR or Cl is normal if and only if there exists a unitary
matrix U which diagonalizes
U −1 AU

9.4 Hermitian, Unitary, Orthogonal and Symmetric Matrices

The aim of this section is to discuss the implications of Theorem 17 for Hermitian, unitary, orthogonal and
symmetric matrices. We begin with Hermitian matrices.

9.4.1 Hermitian Matrices

Suppose that A is a Hermitian matrix. Then A is normal so by Theorem (17) there must be a unitary matrix
U such that U −1 AU = Λ where  
λ1 0 0 · · · 0
 0 λ2 0 · · · 0 
Λ= ... .. .. . . . ,
. . . .. 
0 0 0 · · · λn
the λi being the eigenvalues of A. But
Λ∗ = (Λ∗ )′ = Λa = (U −1 AU )a = (U −1 Aa U )a = U a A(U −1 )a = U −1 AU = Λ
This can be possible only if Λ is a real matrix. We’ve proved the “only if” part of the following corollary.
Corollary 2 A matrix A is Hermitian if and only if there exista a unitary matrix U which transforms A
into a real diagonal matrix
U −1 AU

The preceding implies, among other things, that whether A is real or not, if it is Hermitian its spectrum
{i.e., set of eigenvalues} is real. The reader should try to prove the “if” part of Corollary 2.

9.4.2 Unitary Matrices

Suppose now that A is a unitary matrix. Then A is normal so by Theorem (17) there must be a unitary
matrix U such that U −1 AU = Λ where
 
λ1 0 0 ··· 0
 0 λ2 0 ··· 0 
Λ=  ... .. .. .. ..  ,
. . . . 
0 0 0 · · · λn
the λi being the eigenvalues of A. But

(Λ)a Λ = (U −1 AU )a (U −1 AU ) = (U a Aa (U −1 )a )(U −1 AU ) = U a Aa AU = U a U = I

Thus the the ith diagonal element of (Λ)a Λ, namely |λi |2 must equal one. We’ve proved the “only if” part
of the following corollary.

Corollary 3 A matrix A is unitary if and only if there exists a unitary matrix U which transforms A into
a diaginal matrix
U −1 AU
whose diagonal elements all have magnitude one.

The preceding implies that whether A is real or not, if it is unitary its eigenvalues all lie on the unit circle
in the complex plane. The reader should try to prove the “if” part of Corollary 3.

9.4.3 Orthogonal Matrices

Orthogonal matrices are simply real unitary matrices. Thus we have at once

Corollary 4 A real matrix A is orthogonal if only if there exists a unitary matrix U which transforms A
into a diaginal matrix
U −1 AU
whose diagonal elements all have magnitude one.

9.4.4 Symmetric Matrices

Now suppose that A is real and symmetric. Then A is clearly Hermitian so Corollary (2) is applicable and A
can be diagonalized buy a unitary matrix U . There is of course no reason to believe that U must be real {and
therefore orthogonal}; in fact it need not be. Nevertheless A can be diagonalized by a real unitary matrix,
but we need to prove that this is so. Let us begin by noting that A muist have real eigenvalues because of
the emphasized statement just below Corollary 2. Thus if λ1 is one of A’s eigenvalues, then λ1 I − A must
be real. This means that we can find a real eigenvector x1 {any vector in the kernel of λ1 I − A will do} such
that
Ax1 = λ1 x1
Thus if S1 is the span of x1 , then AS1 ⊂ S1 . Moreover Aa S1⊥ ⊂ S1⊥ because of (9.7). But A′ = Aa = A, so
AS1⊥ ⊂ S1⊥ . In summary, we have the situation that

AS1 ⊂ S1 , AS1⊥ ⊂ S1⊥ , and S1 ⊕ S1⊥ = IRn

except in this case scrX = IRn .

Let y1 be a normalized version of x1 and let {y2 , y3 , . . . , yn } be an orthonormal basis for S1⊥ . Since
S1 ⊕ S1⊥ = IRn , {y1 , y2 , y3 , . . . , yn } must be an orthonormal basis for IRn . Therefore

T 1 = [ y1 y2 · · · yn ]

must be an orthogonal matrix. Moreover since AS1 ⊂ S1 , and AS1⊥ ⊂ S1⊥ , it must be true that
 
λ1 0
AT1 = T1 (9.19)
0 A1

where A1 is the unique solution to the linear equation

A [ y2 y3 · · · yn ] = [ y2 y3 · · · yn ] A1 (9.20)

Clearly (9.19) can be written as  


λ1 0
T1−1 AT1 =
0 A1
We leave it to the reader to prove that A1 must be symmetric. To summarize, for any real symmetric matrix
An×n , there must be a n × n orthogonal matrix T1 , a real eigenvalue λ1 of A and a (n − 1) × (n − 1) real
symmetric matrix A1 such that  
−1 λ1 0
T1 AT1 =
0 A1

By continuing this process, much like before when we discussed Hermitian matrices, we can eventually
transform A into a diagonal matrix by means of a sequence of orthogonal similarity transformations. Since
the product of orthogonal matrices is orthogonal, we see that any symmetric matrix is orthogonally similar
to a real diagonal matrix. The reader should try to prove the converse. We summarize.

Corollary 5 A real matrix A is symmetric if and only if there exists an orthogonal matrix T such that
T −1 AT is a real diagonal matrix.

The diagonal entries of T −1 AT are of course the eigenvalues of A.

Example 60 Let  
∆ 1 2
A=
2 1
Then A’s characteristic polynomial is
 
s − 1 −2
det [ sI − A ] = det = (s − 1)2 − 4 = (s − 3)(s + 1)
−2 s − 1
Therefore the eigenvalues of A are 3 and −1. To compute an eigenvector for 3, we need to find a nonzero
vector in the kernel of the matrix  
2 −2
3I − A =
−2 2
The vector  
1
x1 =
1
will do. Normalizing we get " #
√1
2
y1 =
√1
2
By a similar process we obtain for the eigenvalue −1 the normalized eigenvector
" 1 #
− √2
y2 = 1 √
2

Note that y1 and y2 are automatically orthogonal as predicted by the theory. Since y1 and y2 are both
normalized, the matrix " 1 #

2
− √12
T = [ y1 y2 ] = 1 1 √ √
2 2
is orthogonal. Finally we see that  
3 0
T −1 AT =
0 −1

9.5 Real Quadratic Forms

By a real quadratic form is meant an algebraic expression of the type


n X
X n
aij xi xj (9.21)
i=1 j=1

where the xi are real scalar variables and the aij are fixed real numbers. A quadratic form can be thought
of as a special type of real bilinear form
Xn X m
bij xi yj (9.22)
i=1 j=1

in which m = n, bij = aij , and yi = xi . The term bilinear stems from the fact that (9.22) is linear in the xi
{or the yj } if the yj {or the xi } are held fixed. The reason (9.21) is called a quadratic form is because it is a
sum of pairwise products of the xi . Quadratic forms find many applications in engineering and the sciences.

It is especially useful to exploit matrix notation when dealing with quadratic forms. By introducing the
vectors    
x1 y1
∆  x2  ∆  y2 
x= 
 ...  , y=  .. 

.
xn ym
and the matrices
∆ ∆
A = [ aij ]n×n , B = [ bij ]n×m ,
(9.22) and (9.21) can be written as
x′ By and x′ Ax
respectively.
9.5.1 Symmetric Quadratic Forms

A quadratic form x′ Ax is said to be symmetric if A is a symmetric matrix. It turns out that any quadratic
form x′ Bx can be rewritten as a symmetric quadratic form. To understand how to do this, let us first note
that
x′ Bx = (x′ Bx)′
because x′ Bx is a scalar. Since (x′ Bx)′ = x′ B ′ x it follows that

x′ Bx = x′ B ′ x

Therefore  
1 1
x Bx = (x′ Bx + x′ B ′ x) = x′
′ ′
(B + B ) x
2 2
Hence if we define
∆ 1
A= (B + B ′ )
2
then A is symmetric and
x′ Bx = x′ Ax, ∀x
Thus, without loss of generality, we may as well assume that every quadratic form is symmetric.

9.5.2 Change of Variables: Congruent Matrices

Suppose x′ Ax is a real, symmetric quadratic form. Let T be a nonsingular matrix and define the new vector
of variables

y = T −1 x
Then x = T y so
x′ Ax = y ′ T ′ AT y
Note that T ′ AT is also a symmetric matrix. We see that corresponding to the change of variables x 7−→ y,
there is the coefficient matrix transformation A 7−→ T ′ AT .

More precisely, two real n × n symmetric matrices A and B are said to be congruent if there exists a
real nonsingular matrix T such that B = T ′ AT - and the mapping A 7−→ T ′ AT is called a congruence
transformation. It is a simple matter to verify that matrix congruence is an equivalence relation on the set
of all real-valued n × n matrices. Thus, for example, if A is congruent to B and B is congruent to C then A
is congruent to C.

9.5.3 Reduction to Principal Axes

There are various ways to simplify the structure of a quadratic form by changing variables. Suppose for
example that we are given the real symmetric quadratic form
n X
X n
x′ Ax = aij xi xj
i=1 j=1

and we wish to define a new set of variables y1 , y2 , . . . , yn by a formula of the form

y = T −1 x (9.23)

in such a way that


i. y has the same length as x and

ii. the new expression for x′ Ax in terms of yi variables involves no cross terms; i.e.
n
X

x Ax = bi yi2 (9.24)
i=1

To satisfy requirement i. it is clearly necessary and sufficient that T be an orthogonal matrix. To satisfy
requirement ii., T ′ AT must be a diagonal matrix whose ith diagonal entry is bi . In other words, T must be an
orthogonal matrix which diagonalizes T ′ AT . Now we’ve already explained how to construct an orthogonal
matrix T which diagonalizes T −1 AT - and with T so defined, the diagonal elements of T −1 AT are necessarily
the eigenvalues of A. Moreover, since T is orthogonal T −1 = T ′ so T diagonalizes T ′ AT as required. In
other words, with T so defined, (9.24) holds and ||y|| = ||x|| for all pairs {x, y} satisfying (9.23). The yi are
sometimes called the principal axes of the quadratic form.

9.5.4 Sum of Squares

Again consider the symmetric quadratic form


n X
X n
x′ Ax = aij xi xj
i=1 j=1

What we aim to do now is to define a new set of variables z1 , z2 , . . . , zn by a formula of the form

z = T −1 x (9.25)

in such a way that the new expression for x′ Ax in terms of zi variables involves no cross terms and in
addition, the coefficients of the zi2 are numbers in the finite set {1, −1, 0}. In other words, we will define the
zi so that
X n
x′ Ax = ci zi2 (9.26)
i=1

where each ci is either 1, −1 or 0. To accomplish this T must be nonsingular, but not necessarily orthogonal,
and T ′ AT must be a diagonal matrix with diagonal elements in the set {1, −1, 0}. The construction of T is
as follows.

First pick an orthogonal matrix M which diagonalizes M −1 AM . Let λ1 , λ2 , . . . , λn denote the diagonal
elements of M −1 AM . The λi are of course the eigenvalues of A. Let n1 denote the number of eigenvalues
are positive and n2 the number which are negative. For simplicity, suppose that M has been defined so that
the first n1 λi are positive, the next n2 λi negative, and the last n − n1 − n2 λi are all zero. Then
 
Λ1 0 0
M AM =  0

Λ2 0
0 0 0

where
  λ 0 ··· 0 
λ1 0 ··· 0 n1 +1
 0 λ2 ··· 0   0 λn1 +2 ··· 0 
Λ1 = 
 ... .. .. ..  and Λ2 =  .. .. .. .. 
. . .  
. . . .

0 0 · · · λn1 0 0 · · · λn1 +n2
It follows that if we define
√  p 
λ1 √0 ··· 0 −λn1 +1 0 ··· 0
p
 0 λ2 ··· 0
  0 −λn1 +2 ··· 0 
∆ ∆  
N1 = 
 ... .. ..  and N2 =
..
  .. .. .. .. ,
. . .  . . . p . 
p
0 0 ··· λn1 0 0 ··· −λn1 +n2

then
Λ1 = N1′ I1 N1 and Λ2 = −N2′ I2 N2
where I1 and I2 are the n1 × n1 and n2 × n2 identity matrices respectively. Thus
 
I1 0 0
M ′ AM = N ′  0 −I2 0N
0 0 0

where N is the n × n nonsingular matrix


 
N1 0 0

N = 0 N2 0
0 0 I3

and I3 is the (n − n1 − n2 ) × (n − n1 − n2 ) identity matrix.

Hence if we now define



T = M N −1
then we can write
T ′ AT = AC (9.27)
where  
I1 0 0

AC =  0 −I2 0 (9.28)
0 0 0
From this it clearly follows that if z is defined as in (9.25), then (9.26) will hold with c1 = c2 = · · · = cn1 = 1,
cn1 +1 = cn1 +2 = · · · cn1 +n2 = −1, and cn1 +n2 +1 = cn1 +n2 +2 = · · · cn = 0.

Let us note that the rank of A, written r, must equal the rank of T ′ AT because T is nonsingular. But
the rank of T ′ AT must equal n1 + n2 because of (9.27) and the structure of AC in (9.28). In other words

r = n1 + n2

Let us define the signature of A, written s, to be the difference between the number positive eigenvalues of
A and the number of negative eigenvalues of A. Thus

s = n1 − n2

It can be shown that s, like r, is invariant under congruent transformations. What’s more, the structure of
AC in (9.28) is uniquely determined by the congruence invariants r and s because

r+s r−s
n1 = and n2 =
2 2
AC is accordingly called the congruence canonical form of A.
9.6 Positive Definite Quadratic Forms

A real {symmetric} quadratic form x‘Ax is said to be positive semidefinite if

x′ Ax ≥ 0, ∀x (9.29)

If, in addition, the only vector for which x′ Ax = 0 is x = 0, i.e.,

x′ Ax = 0 =⇒ x = 0, (9.30)

then the quadratic form is positive definite. A is a positive definite matrix {resp. positive semidefinite
matrix} if x′ Ax is a positive definite {resp. positive semidefinite} quadratic form. A is a negative definite
matrix {resp. negative semidefinite matrix} if −A is a positive definite {resp. positive semidefinite} matrix.
Symmetric matrices which are neither positive semidefinite nor negative semidefinite are called indefinite
matrices. We sometimes write A > 0 or A ≥ 0 to indicate that A is positive definite or positive semidefinite
respectively.

We claim that if A is positive definite, then so is any matrix B which is congruent to A. For if A and
B are any such matrices, then (9.29) and (9.30) must hold and there must be a nonsingular matrix T such
that
B = T ′ AT
Now if B were not positive definite, there would have to be a nonzero vector y such that

y ′ By ≤ 0

But if such a vector were to exist, then with x = T y one would have

x′ Ax ≤ 0 and x 6= 0

which contradicts (9.29) and (9.30). Hence the claim is true.

Similar reasoning can be used to prove that all matrices congruent to a positive semidefinite matrix are
positive semidefinite. In other words, positive definiteness and positive semidefiniteness are properties of
symmetric matrices which are invariant under congruence transformations.

9.6.1 Conditions for Positive Definiteness

Let A be a positive definite matrix and let T be an orthogonal matrix which diagonalizes T −1 AT . Then
 
λ1 0 · · · 0
 0 λ2 · · · 0 
T ′ AT = 
 ... .. . . .. 
. . . 
0 0 · · · λn

where λ1 , λ2 , . . . λn are the eigenvalues of A. We claim that T ′ AT {and consequently A} will be positive
definite if and only if all of the λi are positive numbers. For if λk ≤ 0, then with the kth unit vector
ek we would have e′k T ′ AT ek = λk ≤ 0. Since ek 6= 0, this would mean that T ′ AT could not be positive
definite. This proves that if A is positive definite, then its eigenvalues must be positive numbers. The reverse
implication can be established in a similar manner. Thus we can conclude that a symmetric matrix is positive
definite if and only if all of its eigenvalues are positive numbers2 . It can be shown in much the same manner
that a symmetric matrix is positive semidefinite if and only if all of its eigenvalues are nonnegative.
2 What must be true of the signature and rank of a symmetric matrix in order for it to be positive definite?
Were one to rely on the preceding to test for positive definiteness of A, one would have to compute the
eigenvalues of A. There is another way to test for positive definiteness which avoids this. We caution the
reader to remember that the test we are about to describe applies only to symmetric matrices.

Suppose
 
a11 a12 a13 ··· a1n
 a12 a22 a23 ··· a2n 
 
A= a a23 a33 ··· a3n 
 13 .. 
 ... ..
.
..
.
..
. . 
a1n ··· ··· ··· ann n×n

For i ∈ {1, 2, . . . , n}, define the ith principal subarray of A to be


 
a11 a12 a13 · · · a1i
 a12 a22 a23 · · · a2i 
∆ a a23 a33

· · · a3i 
Di = 
 13
 ... .. .. .. .. 
. . . . 
a1i ··· ··· · · · aii i×i

The numbers
det D1 , det D2 , . . . , det Dn
are called the n principal leading minors of A. A proof of the following theorem can be found in most good
texts on matrix algebra.

Theorem 18 A symmetric matrix An×n is positive definite {resp. positive semidefinite} if and only if its
n leading principal minors are positive {resp. nonnegative}.

Example 61 The matrix  


1 2
2 1
has leading minors 1 and −3 and hence is neither positive definite nor positive semidefinite. Note that the
matrix has a negative eigenvalue, namely −1.

Example 62 The matrix  


1 2 3
A = 2 4 6
3 6 9
has leading minors  
1 2
det 1 = 1, det = 0, det A = 0
2 4
and therefore is positive semidefinite but not positive definite.

9.6.2 Matrix Square Roots

Let A be a positive semidefinite matrix. A real {square} matrix B is called a square root of A if

B2 = A

Such a matrix B is often denoted by A.

Out aim is√to show that each positive semidefinite matrix A has a positive semidefinite square root A.
To construct A we first make use of the fact that there is an orthogonal matrix T which diagonalizes
T −1 AT . Thus
A = T ΛT ′
where Λ is a diagonal matrix whose diagonal elements,
√ λ1 , λ2 , . . . , λn are A’s eigenvalues. Since A is positive
semidefinite,
√ √ all
√ of the λi are nonnegative.
√ √ Let Λ be that diagonal matrix whose diagonal elements are
λ1 , λ2 , . . . , λn . Clearly Λ Λ = Λ. Hence if we now define
√ ∆ √ ′
A = T ΛT

then √ √ √ √ √ √ √
( A)2 = T ΛT ′ T ΛT ′ = T ΛT −1 T ΛT ′ = T Λ ΛT ′ = T ΛT ′ = A
which is as required.

QUESTION: If M is any real, rectangular, n × m matrix, then why is M ′ M

1. symmetric?

2. positive semidefinite?

3. positive definite if and only if rank M = m?

9.7 Simultaneous Diagonalization

Let A and B be real, n × n symmetric matrices with A positive definite. In the sequel we will explain how
to construct a nonsingular {but not necessarily orthogonal} matrix T such that

T ′ AT = I (9.31)
T ′ BT = Λ (9.32)

where I is the n × n identity matrix and Λ is a diagonal matrix. Caution: The diagonal elements of Λ will
not necessarily be the eigenvalues of B!

To begin, let us note that A’s eigenvalues must all be positive because A is positive definite. This means
that A’s signature must be n and thus A’s congruence canonical form must be the n × n identity matrix.

Let P be any nonsingular matrix, constructed as in §9.5.4, which takes A into its congruence canonical
form; i.e.,
P ′ AP = In×n (9.33)
Since P ′ BP is symmetric, it is possible to construct an orthogonal matrix Q which diagonalizes

Q−1 (P ′ BP )Q. Set Λ = Q−1 (P ′ BP )Q. Since Q is orthogonal.

Q′ P ′ BP Q = Λ

Moreover because of (9.33),


Q′ P AP Q = Q′ Q = I
Thus if we define

T = PQ
then (9.31) and (9.32) both hold.
9.7.1 Constrained Optimization

Now suppose that A and B are n × n matrices and that A is positive definite. In the sequel we shall consider
the problem of finding a vector x, within the set of vectors which satisfy
x′ Ax = 1, (9.34)
which maximizes the value of the quadratic form

q(x) = x′ Bx (9.35)
Towards this end, let us first note that if T is any n × n nonsingular matrix, and y is the vector of variables
defined by

y = T −1 x, (9.36)
then maximization of (9.35) with respect to x satisfying constraint (9.34), is equivalent to maximization of

q̄(y) = y ′ T ′ BT y (9.37)
with respect to y satisfying the constraint
y ′ T ′ AT y = 1 (9.38)
In other words, if x0 is a vector satisfying constraint (9.34), which at the same time maxamizes the quadratic

form in (9.35), then y0 = T −1 x0 maximizes the quadratic form in (9.37) subject to (9.38) and conversely.
Moreover, for any such x0 and y0 , q(x0 ) = q̄(y0 ). We leave of to the reader to verify that these statements
are correct.

The preceding suggests that we might significantly simplify the maximization problem by choosing T so
that
T ′ AT = I (9.39)

T BT = Λ (9.40)
where Λ is a diagonal matrix with entries λ1 , λ2 , . . . , λn . For with T so chosen, (9.37) and (9.38) become
n
X
q̄(y) = λi yi2 (9.41)
i=1

and n
X
yi2 = 1 (9.42)
i=1
respectively, where yi is the ith component of y. Suppose λk is the largest number in the set {λ1 , λ2 , . . . , λn }.
It is then reasonably clear {and can be formally verified using, for example, Lagrange multipliers} that the

kth unit vector ek is a value of y which maximizes (9.41) subject to (9.42). Thus with x0 = T ek ,

x′0 Bx0 = ′max x′ Bx = λk


x Ax=1

9.7.2 A Special Case

Now suppose that in the above problem, A is the n × n identity matrix. Then T must be an orthogonal
matrix because of (9.39). Therefore the λi must be the eigenvalues of B because of (9.40). It follows that in
this case, the maximum value of x′ Bx subject to x′ x = 1 is the largest eigenvalue {say λmax } of B. Moreover
if z is a normalized eigenvector corresponding to λmax , then z ′ z = 1 and z ′ Bz = λmax z ′ z = λmax . Therefore
z is a value of x which solves the maximization problem.
Theorem 19 Let B be an n × n symmetric matrix and let λmax denote the largest eigenvalue of B. Then

λmax is the maximum value of the quadratic form q(x) = x′ Bx over the set of all vectors x satisfying
x′ x = 1. Moreover, q(x) attains this maximum value at x = x0 , where x0 is any normalized eigenvector of
B corresponding to the eigenvalue λmax .

Using similar reasoning as above, one can prove just as easily the following theorem.

Theorem 20 Let B be an n × n symmetric matrix and let λmin denote the smallest eigenvalue of B. Then

λmin is the minimum value of the quadratic form q(x) = x′ Bx over the set of all vectors x satisfying
x′ x = 1. Moreover, q(x) attains this minimum value at x = x0 , where x0 is any normalized eigenvector of
B corresponding to the eigenvalue λmin .

The following are direct consequences of the preceding theorems.

x′ Bx
max = λmax
x x′ x
x′ Bx
min = λmin
x x′ x
These expressions imply that

λmin ||x||2 ≤ x′ Bx ≤ λmax ||x||2 , ∀x ∈ IRn

These inequalities are especially useful in the important special case when B is positive definite and conse-
quently λmax and λmin are positive numbers.

Example 63 Suppose  
∆2 1
B=
1 2
in which case
det(sI − B) = (s − 1)(s − 3)
so B has eigenvalues of 1 and 3. Therefore

x′ Bx x′ Bx
max =3 and min =1
x x′ x x x′ x
Since    
√1 − √12
∆ 2 ∆
z=  and w= 
√1 √1
2 2
are normalized eigenvectors of B for the eigenvalues 3 and 1 respectively,

z ′ Bz = 3, ||z|| = 1, w′ Bw = 1, and ||w|| = 1.

Note that for any fixed positive number c, the set of x satisfying

x′ Bx = c

is an ellipse in IR2 . The set of x satisfying


||x|| = 1
is of course the unit circle in IR2 . With  
x1

x=
x2
these equations can be written as
2x21 + 2x1 x2 + 2x22 = c
and
x21 + x22 = 1
respectively. In the plane, the situation is as follows.

x2

w z
1
2

x1
-1 1 1
2 2
c=
1
c=
3

in
c
re
as
in
g
c
Part II

Linear Systems

163
Chapter 10

Introductory Concepts

The term “linear system” arises in many fields and can mean anything from a linear algebraic equation to
a linear map between suitably defined Hilbert spaces. The aim of this chapter is to define what is meant by
a linear dynamical system.

10.1 Linear Systems

We shall be primarily concerned with physical systems which can be represented by linear, finite dimensional
dynamical models. By a real n-dimensional continuous-time linear dyanmical system with input u, state x
and output y is meant a system of equations of the form

ẋ = A(t)x + B(t)u (10.1)


y = C(t)x + D(t)u (10.2)

where t takes values in the set of non-negative reals, u, x, and y are vector-valued functions of t, and overdot
· means differentiation with respect to time; i.e.,

∆ d
ẋ(t) = x(t)
dt
Similarly, by a real n-dimensional discrete-time linear dyanmical system with input u, state x and output y
is meant a system of equations of the form

x = A(t)x + B(t)u (10.3)
y = C(t)x + D(t)u (10.4)

where t takes values in the set of natural numbers N = {0, 1, 2, 3, · · ·}, u, x, and y are vector-valued functions
of t, and overdiamond ⋄ means time advance; i.e.,
⋄ ∆
x (t) = x(t + 1)

In either the continuous or discrete time case, u, x, and y take values in IRm , IRn and IRp respectively and
An×n (t), Bn×m (t), Cp×n (t), and Dp×m (t) are real matrix-valued functions. In the event that A(t), B(t), C(t),
and D(t) are constant matrices, not depending on t, the corresponding system is said to be time-invariant
or stationary. Otherwise the system is time-varying.

165
In the sequel we will often use the coefficient matrix quadruple {A(t), B(t), C(t), D(t)} to denote the
continuous-time system (10.1), (10.2) or the discrete-time system (10.3), (10.4), depending on context.
The coefficient matrix triple {A(t), B(t), C(t)} denotes either system (10.1), (10.2) or (10.3), (10.4), again
depending on context, for the case when D(t) = 0.

10.2 Continuous-Time Linear Systems

When dealing a continuous-time linear system {A(t), B(t), C(t), D(t)} we shall invariably assume that its
input u and coefficient matrices A(t), B(t), C(t), D(t) are at least piecewise continuous functions of t. Thus
the variation of constants formula applies and the solution to the state equation (10.1) can be written at
Z t
x(t) = Φ(t, t0 )x0 + Φ(t, τ )B(τ )u(τ )dτ, t ≥ t0 ≥ 0 (10.5)
t0

where Φ(t, τ ) is the state transition matrix of A(t) and x0 is the state of (10.1) at t = t0 . Using this and the
readout equation (10.2) we thus obtain the input-output formula
Z t
y(t) = C(t)Φ(t, t0 )x0 + C(t)Φ(t, τ )B(τ )u(τ )dτ + D(t)u(t), t ≥ t0 ≥ 0 (10.6)
t0

which completely characterizes the input-output behavior of the system. Let us note that if x0 = 0, the
preceding defines a bona fide linear function
Z t
u 7−→ C(t)Φ(t, τ )B(τ )u(τ )dτ + D(t)u(t)
t0

from the function space of piecewise continuous inputs [t0 , ∞) → IRm to function space of piecewise contin-
uous outputs [t0 , ∞) → IRp .

10.2.1 Time-Invaraint, Continuous-Time, Linear Systems

In the time-invariant case, coefficent matrices A, B, C, and D are constant. Thus in this case, the state
transition matrix of A is of the for
Φ(t, τ ) = e(t−τ )A
where etA is the matrix exponential
X∞
∆ 1
etA = (tA)i (10.7)
i=0
i!

Hence in this case, (10.5) and (10.6) become


Z t
x(t) = e(t−t0 )A x0 + e(t−τ )A Bu(τ )dτ, t ≥ t0 ≥ 0 (10.8)
t0

and
Z t
(t−t0 )A
y(t) = Ce x0 + Ce(t−τ )A Bu(τ )dτ + Du(t), t ≥ t0 ≥ 0 (10.9)
t0

respectively.
For time invariant systems, one is often interested in the situation when t0 = 0 and x0 = 0. Under these
conditions, (10.9) simplifies to the convolution integral
Z t
y(t) = H(t − τ )u(τ )dτ, t ≥ 0 (10.10)
0

where

H(t) = CetA B + δ(t)D (10.11)
and δ is the unit impulse at t = 0. Note that in the special case when m = 1 and u is a unit impulse applied
at time t = 0, {A, B, C, D}’s output response is H(t). It is thus natural to call H(t) the impulse response
matrix of {A, B, C, D}.

For time-invariant linear systems there is much to be gained by working in the frequency domain with
Laplace Transforms. Let us write U (s), X(s), and Y (s) for the Laplace transforms of u, x, and y respectively.
Recall that the Laplace transform is a linear map and that the Laplace transform of ẋ is sX(s) − x(0). From
this and the fact that ẋ = Ax + Bu and y = Cx + Du it follows that
sX(s) − x(0) = AX(s) + BU (s) and Y (s) = CX(s) + DU (s)
Hence by solving for X(s) we get
X(s) = (sI − A)−1 x0 + (sI − A)−1 BU (s)
Substituting X(s) into the expression for Y (s) we thus obtain
Y (s) = C(sI − A)−1 x0 + T (s)U (s)
where

T (s) = C(sI − A)−1 B + D (10.12)
We call T (s) the transfer matrix of Σ or of {A, B, C, D}. We see that T (s) characterizes completely the
relationship between U (s) and Y (s) for the case when x0 = 0. It is perhaps not supprising then that T (s)
is the Laplace transform of the impulse response matrix H(t).

The formula in (10.12) clearly shows that the transfer functions in T (s) are rational functions of s. It is
not difficult to see that they are proper rational functions1 . Note that T (s) will be strictly proper just in
case just in case D = 0.

T (s) and H(t) are Laplace transform pairs and should be regarded as two mathematically equivalent
input-output descriptions of Σ. There is a third input-output description which we will find useful later.
What we are referring to is the sequence of matrices

∆ ∆ d(i−1)
M0 = D, Mi = (H(t) − Dδ(t)) {t=0} , i≥1
dt
which is called Σ’s { or H(s)’s or T (s)’s} Markov sequence. It is easy to verify that the Mi relate to the
coefficient matrix of Σ by the formulas
M0 = D, Mi = CAi−1 B, i≥1
The Mi are also the coefficient matrices of a formal power series expansion of T (s); i.e.,
T (s) ≈ M0 + M1 s−1 + M2 s−2 + · · ·
It is possible to prove that two linear systems have the same transfer matrix if and only if they have the
same Markov sequence.
1 A rational function in one variable is proper if the degree of its numerator is no greater than the degree of its denominator;

a proper rational function is strictly proper if the degree of its numerator is strictly less than the degree of its denominator. A
matrix of rational functions is proper if each of its entries is proper and strictly proper if each of its entries is strictly proper.
10.3 Discrete-Time Linear Systems

The state equation (10.3) of a discrete-time linear system also admits a solution which is similar in form to
(10.5). For the discrete case, the equation is of the form
t−1
X
x(t) = Φ(t, t0 )x(t0 ) + Φ(t − 1, τ )B(τ )u(τ ) (10.13)
τ =t0

where Φ : N × N → IRn×n is the discrete-time state transition matrix


(
A(t)A(t − 1) · · · A(τ + 1) if t > τ
Φ(t, τ ) = (10.14)
I if t = τ

Hence from (10.4)


t−1
X
y(t) = C(t)Φ(t, t0 )x(t0 ) + C(t)Φ(t − 1, τ )B(τ )u(τ ) + D(t)u(t) (10.15)
τ =t0

10.3.1 Time-Invariant Discrete Time Systems

In the time invariant case, A(t) is constant so from (10.14) we get

Φ(t, τ ) = A(t−τ ) , t ≥ τ (10.16)

Hence in the time-invariant case (10.13) and (10.15) simplify to


t−1
X
x(t) = A(t−t0 ) x(t0 ) + A(t−1−τ ) Bu(τ ) (10.17)
τ =t0

and
t−1
X
y(t) = CA(t−t0 ) x(t0 ) + CA(t−1−τ ) Bu(τ ) + Du(t) (10.18)
τ =t0

respectively. Just as in the continuous case, for time invariant discrete time systems, one is often interested
in the situation when t0 = 0 and x0 = 0. Under these contitions, (10.18) simplifies to the convolution sum
t
X
y(t) = H(t − τ )u(τ ), t≥0 (10.19)
0

where (
∆ D if t = 0
H(t) = (10.20)
CA(t−1) B if t > 0
Note that in the special case when m = 1 and u is a unit pulse applied at time t = 0, {A, B, C, D}’s output
response is H(t). It is thus natural to call H(t) the pulse response matrix of {A, B, C, D}.

For time-invariant discrete-time linear systems there is much to be gained by working in the frequency
domain with z-transforms. Let us write U (z), X(z), and Y (z) for the z-transforms of u, x, and y respectively.

Recall that the z-transform is a linear map and that the z-transform of x is zX(z) − x(0). From this and

the fact that x= Ax + Bu and y = Cx + Du it follows that

zX(z) − x(0) = AX(z) + BU (z) and Y (z) = CX(z) + DU (z)


Hence by solving for X(z) we get
X(z) = (zI − A)−1 x0 + (zI − A)−1 BU (z)
Eliminating X(z) from the expression for Y (z) we thus obtain
Y (z) = C(zI − A)−1 x0 + T (z)U (z)
where

T (z) = C(zI − A)−1 B + D (10.21)
We call T (z) the discrete-time transfer matrix of Σ or of {A, B, C, D}. We see that T (z) characterizes
completely the relationship between U (z) and Y (z) for the case when x0 = 0. It is perhaps not supprising
then that T (z) is the z- transform of the pulse response matrix H(t). Note that the formula for the transfer
matrix of a discrete-time linear system has exactly algebraic form as the transfer matrix for a continuous-time
linear system.

T (z) and H(t) are z-transform pairs and should be regarded as two mathematically equivalent input-
output descriptions of Σ. There is a third input-output description which we will find useful later. What we
are referring to is the sequence of matrices

Mi = H(i), i≥0
which is called Σ’s { or H(t)’s or T (s)’s} discrete-time Markov sequence. It is easy to verify that the Mi
relate to the coefficient matrix of Σ by the formulas
M0 = D, Mi = CAi−1 B, i≥1
The Mi are also the coefficient matrices of a formal power series expansion of T (z); i.e.,
T (z) ≈ M0 + M1 z −1 + M2 z −2 + · · ·
It is possible to prove that two discrete-time linear systems have the same transfer matrix if and only if they
have the same Markov sequence.

10.3.2 Sampled Data Systems

Perhaps the most important example of a discrete time system is that which results when the output of a
time-invariant, continuous time system
ẋ = Ax + Bu y = Cx + Du
is sampled. Let T be a fixed positive number called a sampling time and let UT be the subspace of all
piecewise constant input signals of the form
u(t) = ui , iT ≤ t ≤ (i + 1)T, i = 0, 1, 2, . . .
m
where ui ∈ IR . If u is any such input and if x(t) is the resulting state-response, then for each i ≥ 0,
Z (i+1)T
x((i + 1)T ) = eAT x(iT ) + eA((i+1)T −τ ) Bui dτ
iT

From this it follows that if we define s = τ − iT , xi = x(iT ) and yi as the limit of y(t) at t approaches iT
from above, then for i ≥ 0
Z T !
AT
 A(T −s)
xi+1 = e xi + e Bds ui
0

yi = Cxi + Dui
which is a time-invariant, discrete time, linear system.
10.4 The Concept of a Realization

The continuous and discrete state systems discussed thus far should be regarded as candidate systems for
modeling the relationships between physically significant variables, namely inputs and outputs.

u Σ y
Figure 10.1: Linear System Σ

As we have seen, each such system uniquely determines a linear mapping of inputs to outputs. For example,
in the case of a continuous-time linear system {A(t), B(t), C(t), D(t)} initialized at the zero state at t = t0 ,
one has the relation Z t
y(t) = W (t, τ )u(τ )dτ + D(t)u(t)
t0

where

W (t, τ ) = C(t)Φ(t, τ )B(τ )
Φ(t, τ ) being the state transition matrix of A(t). We call W (t, τ ) the weighting pattern of {A(t), B(t), C(t)}.
Similarly in the case of a discrete-time linear system {A(t), B(t), C(t), D(t)} initialized at the zero state at
t = t0 , one has the relation
Xt−1
y(t) = W (t, τ )u(τ ) + D(t)u(t)
τ =t0

where

W (t, τ ) = C(t)Φ(t − 1, τ )B(τ )
Φ(t, τ ) being the discrete-time state transition matrix of A(t). In this case, W (t, τ ) the discrete-time weighting
pattern of {A(t), B(t), C(t)}.

Consider the continuous time case. Suppose that we are given an arbitrary of p × n matrix Q(t, τ )
which depend at least piecewise continously on its arguments. By a realization of Q(t, τ ) is meant any
continuous-time linear system {A(t), B(t), C(t)} whose weighting pattern is equal to Q(t, τ ).

Exactly the same ideas apply in the discrete-time case: Suppose that we are given an arbitrary of p × n
matrix Q : N × N → IRp×m By a discrete-time realization of Q(t, τ ) is meant any discrete-time linear system
{A(t), B(t), C(t)} whose weighting pattern is equal to Q(t, τ ).

Consideration of either the continuous-time or discrete-time case leads naturally to a number of questions:

1. Existence: Under what conditions does Q(t, τ ) admit a finite-dimensional realization?

2. Uniqueness: To what extent are realizations of Q(t, τ ) unique?

3. Time-Invariance: What must be true of Q(t, τ ) in order for it to admit a time-invariant linear realiza-
tion?

4. Minimal-Realization: What must be true of a realization of Q(t, τ ) in order for its {state-space}
dimension to be minimal over the set of all realizations of Q(t, τ )?

Linear realization theory is concerned with development of answers to all of these questions. Answers
for the continuous-time case and for the discrete-time case are of course going to be different. Issues
in continuous-time case are typically somewhat more challenging technically than the discrete-time case,
because for continous-time systems one has to deal with differential equations. On the other hand, the
continuous-time theory is somewhat more concise than the discrete-time theory; this partitally counterbal-
ances the relative technical simplicity the discrete-time case has over the continuous-time case. In any event,
it almost always turns out that once one understands an issue in continous-time, extension to the discrete-
time case is a simple matter. For this reason, in the sequel we will usually limit developments primarily to the
continuous-time case. The reader is encouraged to work out the discrete-time analog of each continuous-time
result which follows.

Transfer Matrix Realizations

In the time-invariant case, the weighting pattern of a linear system {A, B, C} is of the form

W (t, τ ) = Ce(t−τ )A B

Note that
W (t, τ ) = H(t − τ )
where H(t) is the impulse response of {A, B, C}. Thus realizing a given W amounts to finding matrices
A, B, C for which
CetA B = H(t)

Prompted by this, we also use the word realization to describe any constant linear system Σ = {A, B, C, D}
for which either
CetA B + Dδ(t) = H(t) (10.22)
or
C(sI − A)−1 B + D = T (s) (10.23)
where H(t) and T (s) are given matrices and δ(t) is the unit impulse at t = 0. In the former case we say that
Σ is an impulse matrix realization, whereas in the latter Σ is a transfer matrix realization. Of course, any
transfer matrix realization is also a impulse response realization and visa versa.

Recall that the formula for the transfer matrix of a discrete-time system is identical in form to that for a
continuous time system. Because of this, the transfer matrix realization questions for discrete and continuous
time systems are identical. We will address these questions later in the notes.

10.4.1 Existence of Continuous-Time Linear Realizations

The following theorem settles the main existence question for continous-time linear systems.

Theorem 21 A p × m piecewise-continuous matrix Q(t, τ ) is the weighting pattern of a linear system, if


and only if there exists a finite positive integer n and piecewise-continuous matrices M (t) and N (t) such that

Q(t, τ ) = M (t)N (τ ) (10.24)

Moreover, if such Mp×n (t) and Nn×m (t) exist, Q(t, τ ) is realized by the system {A(t), B(t), C(t)}, where
∆ ∆ ∆
A(t) = 0n×n , B(t) = N (t) and C(t) = M (t).


Proof: The state- transition matrix of A(t) = 0 is Φ(t, τ ) = I, ∀t, τ ≥ 0. Hence the weighting pattern of
{A(t), B(t), C(t)} is W (t, τ ) = M (t)In×n N (τ ) = Q(t, τ ).
Although Theorem 21 answers the existence question, it leaves many closely related issues unresolved.
For example, the realization used in the proof of the theorem will typically be a time-varying system, even in
the case when Q(t, τ ) admits a time-invariant realization. Moreover, the dimension of the proof’s realization
may well be much larger than that of some other, more carfully contructed realization of Q(t, τ ). To deal
with these issues, we will need to delve into the underlying properties of continuous-time linear systems.
Among these is are the properties of controllability and observability, which after stability, are probably the
most important concept in linear system theory. Controllability is discussed in the next chapter.
Chapter 11

Controllability

The aim of this chapter is to define and charactize the concept of a controllable continuous-time linear
system. We begin with the closely related idea of reachability.

11.1 Reachable States



Let z(t, τ, x; v) denote the state of Σ = {A(t), B(t), C(t), D(t)} at time t assuming Σ was in the state x at
time τ and u(t) = v(t) for t ∈ [τ, t]. Thus by the variation of constants formula,
Z t
z(t, τ, x; v) = Φ(t, τ )x + Φ(t, s)B(s)v(s)ds (11.1)
τ

where Φ(t, τ ) is the state transition matrix of A(t).

Let t0 and T be fixed times with T > t0 . Let x0 ∈ IRn be given. Let us agree to say that a state x ∈ IRn
of Σ is reachable on [t0 , T ] from x0 if there exists an input v defined on [t0 , T ] such that

x = z(T, t0 , x0 ; v) (11.2)

Thus x is reachable from x0 on [t0 , T ] just in case there is an input v which transfers Σ from state x0 at time
t0 to state x at time T .

Let R[t0 , T ] denote the set of all states of Σ which are reachable from the zero state on [t0 , T ]. Note that
R[t0 , T ] is nonempty because the zero state of Σ is reachable from the zero state on [t0 , T ] with the input
v(t) = 0, t ∈ [t0 , T ]. Now suppose that x1 and x2 are any two states in R[t0 , T ]. This means that there
must be inputs v1 and v2 such that

xi = z(T, t0 , 0; vi ), i ∈ {1, 2}

Then by linearity

r1 x1 + r2 x2 = r1 z(T, t0 , 0; v1 ) + r2 z(T, t0 , 0; v2 ) = z(T, t0 , 0; r1 v1 + r2 v2 ), ∀r1 , r2 ∈ IR

Thus for any numbers r1 and r2 , the linear combination r1 x1 +r2 x2 is reachable from the zero state on [t0 , T ].
Evidently R[t0 , T ] is not only nonempty, but also closed under vector addition and scalar multiplication.
This proves that R[t0 , T ] is a linear subspace of IRn . We call R[t0 , T ] the reachable space of Σ {or of the pair
(A(t), B(t))} on [t0 , T ]. It is possible to characterize R[t0 , T ] as follows.

173
Theorem 22
R[t0 , T ] = image GC (t0 , T ) (11.3)
where Z T
GC (t0 , T ) = Φ(T, τ )B(τ )B ′ (τ )Φ′ (T, τ )dτ (11.4)
t0

The matrix GC (t0 , T ) defined by (11.4) is called the controllability Gramian of (A(t), B(t)) or of Σ on [t0 , T ].
Note that GC (t0 , T ) is a symmetric, positive semi-definite matrix; in addition if GC (t0 , T ) is positive definite,
then so is GC (t0 , T̄ ) for any T̄ ≥ T ≥ 0. From this and Theorem 22 it follows that if every state of Σ is
reachable from the zero state on [t0 , T ], then every state of Σ is also reachable from the zero state on [t0 , T̄ ]
provided T̄ ≥ T .

Proof of Theorem 22: First suppose that x ∈ R[t0 , T ]. Then x is reachable from the zero state on [t0 , T ]
so there must be an input v such that
Z T
x= Φ(T, τ )B(τ )v(τ )dτ (11.5)
t0

Write GC for GC (t0 , T ). In general, (image GC )⊥ = kernel G′C , but since GC is symmetric, this identity
becomes (image GC )⊥ = kernel GC . Hence IRn admits an orthogonal decomposition of the form

IRn = image GC ⊕ kernel GC

Therefore there must be vectors g ∈ image GC and ḡ ∈ kernel GC such that

x = g + ḡ (11.6)

Since ḡ and g are orthogonal


ḡ ′ x = ||ḡ||2
∆ √
where, for all y ∈ IRn , ||y|| denotes the Euclidean norm ||y|| = y ′ y. From this and (11.5) it follows that
Z T
||ḡ||2 = ḡ ′ Φ(T, τ )B(τ )v(τ )dτ (11.7)
t0

Using the definition of GC in (11.4) ,


Z T Z T
ḡ ′ GC ḡ = ḡ ′ Φ(T, τ )B(τ )B ′ (τ )Φ(T, τ )ḡdτ = ||B ′ (τ )Φ′ (T, τ )ḡ||2 dτ (11.8)
t0 t0

But
ḡ ′ GC ḡ = 0
because ḡ is in the kernel of GC . From this and (11.8) it follows that
Z T
||B ′ (τ )Φ′ (T, τ )ḡ||2 dτ = 0
t0

This implies that


B ′ (τ )Φ′ (T, τ )ḡ = 0 ∀τ ∈ [t0 , T ]
because B ′ (τ )Φ′ (T, τ )ḡ is continuous on [t0 , T ]. This shows that the integrand in (11.7) vanishes identically
on [t0 , T ] so ḡ = 0. In view of (11.6), it must be true that x = g and therefore that x ∈ image GC . Since x
was chosen arbitrarily, we have proved that

R[t0 , T ] ⊂ image GC (t0 , T ) (11.9)


To prove the reverse inclusion, now fix x ∈ image GC . Then there must be a vector q ∈ IRn such that

x = GC q

Consider the input


v(t) = B ′ (t)Φ′ (T, t)q, t ∈ [t0 , T ]
With v so defined
Z T Z T
z(T, t0 , 0; v) = Φ(T, τ )B(τ )v(τ )dτ = Φ(T, τ )B(τ )B ′ (τ )Φ′ (T, τ )qdτ = GC (t0 , T )q = x
t0 t0

Thus x is reachable from the zero state on [t0 , T ]. Since x is arbitrary, this must be true for every x in the
image of GC . Therefore
image GC (t0 , T ) ⊂ R[t0 , T ]
From this and (11.9), it now follows that (11.3) is true.

11.2 Controllability

A linear system Σ = {A(t), B(t), C(t), D(t)} {or matrix pair (A(t), B(t)) } is said to be controllable on [t0 , T ]
if for each pair of states x1 , x2 ∈ IRn there is a control input v which transfers Σ from state x1 at t = t0 to
state x2 at time t = t2 . In other words, Σ is controllable on [t0 , T ] just in case for each pair x1 , x2 ∈ IRn ,
there is a v such that
x2 = z(T, t0 , x1 , v) (11.10)
where, as before,
Z t

z(t, τ, x; v) = Φ(t, τ )x + Φ(t, s)B(s)v(s)ds (11.11)
τ

Thus a controllable linear system on a given time interval, is a linear system within which it is possible to
transfer from any state to any other state on the given time intereval.

Let us note that if Σ is controllable on [t0 , T ], then every possible state of Σ must be reachable from the
zero state on [t0 , T ]. In other words, if Σ is controllable on [t0 , T ], then it must be true that R[t0 , T ] = IRn .
The converse is also true:

Theorem 23 A linear system Σ with coefficient matrix quadruple {A(t), B(t), C(t), D(t)} is controllable on
[t0 , T ] if and only if
R[t0 , T ] = IRn (11.12)

In view of Theorem22, we see that controllability of Σ on [t0 , T ] is equivalent to image GC (t0 , T ) =


IRn . Since the latter, in turn, is equivalent to det GC (t0 , T ) 6= 0, we are led to the following alternative
characterization of controllability.

Corollary 6 A linear system with coefficient matrix quadruple {A(t), B(t), C(t), D(t)} is controllable on
[t0 , T ] if and only if its controllablity Gramian GC (t0 , T ) is nonsingular.

Proof of Theorem 23: As we’ve already explained, just before the statement of Theorem 23, controllability
of Σ on [t0 , T ] implies that R[t0 , T ] = IRn . To prove the converse, suppose that (11.12) holds, and that x1
and x2 are any two given states of Σ. In view of Theorem 22, image GC [t0 , T ] = IRn which implies that
GC [t0 , T ] is nonsingular and thus invertible. We claim that the control
v(t) = B ′ (t)Φ′ (T, t)G−1
C (t0 , T ){x2 − Φ(T, t0 )x1 }, t ∈ [t0 , T ]

accomplished the desired state transfer. In particular, with this control


Z T
z(T, t0 , x1 ; v) = Φ(T, t0 )x1 + Φ(T, τ )B(τ )v(τ )dτ
t0
Z T
= Φ(T, t0 )x1 + Φ(T, τ )B(τ )B ′ (τ )Φ′ (T, τ )G−1
C (t0 , T ){x2 − Φ(T, t0 )x1 }dτ
t0
Z T !
= Φ(T, t0 )x1 + Φ(T, τ )B(τ )B (τ )Φ (T, τ )dτ G−1
′ ′
C (t0 , T ){x2 − Φ(T, t0 )x1 }
t0

= Φ(T, t0 )x1 + GC (t0 , T )G−1


C (t0 , T ){x2 − Φ(T, t0 )x1 }
= Φ(T, t0 )x1 + x2 − Φ(T, t0 )x1
= x2
Hence (11.10) holds which proves that v has the desired property.

11.2.1 Controllability Reduction


Let Σ = {A(t), B(t), C(t)} be a given n-dimensional continuous time linear system and let
W (t, τ ) = C(t)Φ(t, τ )B(τ )
be its weighting pattern where Φ(t, τ ) is the state transition matrix of A(t). Let [t0 , T ] be any interval of
positive length on which Σ is defined. The aim of this section is to prove that every weighting pattern has
a controllable realization. We will do this constructively by explaining how to use Σ to compute a linear

system Σ̄ = {Ā(t), B̄(t), C̄(t)} which is controllable on [t0 , T ] and which has the same weighting pattern as
Σ. Since Σ̄’s dimension will turn out to be less than Σ’s, if Σ is not controllable, this construction is referred
to as controllability reduction. The fact that W (t, τ ) can always be realized by a controllable system is one
of the cornerstones of the theory upon which the existence of minimal dimensional realizations of W (t, τ ) is
based.

To begin, let us suppose right away that Σ is not controllable. For if Σ were controllable, we could simply
define Σ̄ to be Σ and we would have at once a controllable realization of W (t, τ ). To proceed, let GC denote
Σ’s controllability Gramian on [t0 , T ]. In other words
Z T

GC = Φ(T, τ )B(τ B ′ (τ )Φ′ (T, τ )dτ
t0

As noted earlier, the symmetry of GC implies that IRn admits an orthogonal decomposition of the form
IRn = image GC ⊕ kernel GC
Let P denote the orthogonal projection matrix on image GC . Hence P GC = GC . Because GC is positive-

semidefinite it is possible1 to construct a full rank matrix Rn×n̄ , with n̄ = rank GC , such that
GC = RR′
 
1 This In̄×n̄ 0
can easily be deduce by first noting that GC is congruent to a block diagonal matrix of the form . In
0 0
 
I 0
other words, for some nonsingular matrix Q, GC = Q n̄×n̄ Q′ . Defining R to be the submatrix consisting of the first n̄
0 0
columns of Q produces the desired factorization.
Our construction relies on the fact that P can be written as
P = R(R′ R)−1 R′
{The reader should verify that this basic algebraic property is so.} Now consider the new system Σ̄ defined
for t ∈ [t0 , T ] by
Ā(t) = 0 B̄(t) = SΦ(T, t)B(t) C̄(t) = C(t)Φ(t, T )R
where

S = (R′ R)−1 R′
It will now be shown that Σ̄ is controllable and that its weighting pattern is W (t, τ ). We establish control-
lability first. Toward this end, note that Ā’s state transition matrix is Φ̄(t, τ ) = In̄×n̄ . Using this fact we
can write Σ̄’s controllability Gramian
Z T
ḠC = Φ̄(T, t)B̄(t)B̄ ′ (t)Φ̄′ (T, t)dt
t0
as Z T
ḠC = B̄(t)B̄ ′ (t)dt
t0
Hence, using the definitions of B̄(t) and C̄(t) we can write
Z T
ḠC = SΦ(T, t)B(t)B ′ (t)Φ(T, t)′ S ′ dt
t0
Z !
T
= S Φ(T, t)B(t)B (t)Φ(T, t) dt S ′
′ ′
t0

= SGC S ′
= (R′ R)−1 R′ RR′ R(R′ R)−1
= In̄×n̄
This prove that ḠC (t0 , T ) is nonsingular. Therefore by Corollary 6, Σ̄ is controllable on [t0 , T ] as claimed.

To prove that Σ̄’s weighting pattern is the same as Σ’s, we first need to show that
P Φ(T, t)B(t) = Φ(T, t)B(t), t ∈ [t0 , T ] (11.13)
Now this will be true provided
Z T
||(B ′ (t)Φ(T, t)′ P ′ − B ′ (t)Φ′ (T, t))z||2 dt = 0, ∀z ∈ IRn (11.14)
t0

Fixing z arbitrarily and expanding the integral on the left above we get
Z T Z T
′ ′ ′ ′ 2
||(B (t)Φ(T, t)P − B (t)Φ (T, t))z|| dt = z ′ P Φ(T, t)B(t)B ′ (t)Φ′ (T, t)P ′ zdt
t0 t0
Z T
− z ′ Φ(T, t)B(t)B ′ (t)Φ′ (T, t)P ′ zdt
t0
Z T
− z ′ P Φ(T, t)B(t)B ′ (t)Φ′ (T, t)zdt
t0
Z T
+ z ′ Φ(T, t)B(t)B ′ (t)Φ′ (T, t)zdt
t0
= z ′ P GC P ′ z − z ′ GC P ′ z − z ′ P GC z + z ′ GC z
= z ′ GC z − z ′ GC z − z ′ GC z + z ′ GC z
= 0
Hence (11.14) is true. Consequently (11.13) is true as well.

It remains to be shown that Σ̄’s weighting pattern is the same as Σ’s. This is now easily accomplished
by noting that Σ̄’s weighting pattern W̄ satisfies
W̄ (t, τ ) = C̄(t)Φ̄(t, τ )B̄(τ )
= C̄(t)B̄(τ )
= C(t)Φ(t, T )RSΦ(T, τ )B(τ )
= C(t)Φ(t, T )P Φ(T, τ )B(τ )
= C(t)Φ(t, T )Φ(T, τ )B(τ )
= C(t)Φ(t, τ )B(τ )
= W (t, τ )
Therefore Σ̄ and Σ have the same weighting pattern as claimed.

We summarize.

Theorem 24 Every continuous-time weighting pattern admits a controllable realization.

11.3 Time-Invariant Continuus-Time Linear Systems

While the constructions and results regarding controllability discussed so far, apply to both time-varying
and time-invariant linear systems, the consequential simplications of the time-invariant case are significant
enough to justfy a separate treatment.

Let Σ = {A, B, C, D} be a time-invariant system. Then Σ’s controllabilty Gramian on [t0 , T ] is

Z T
∆ ′
GC (t0 , T ) = e(T −τ )A BB ′ e(T −τ )A dτ
t0

By introducting the change of variables s = τ − t0 we see that
Z T −t0

GC (t0 , T ) = e(T −t0 −s)A BB ′ e(T −t0 −s)A ds
0

Since the right hand side of the preceding is also the controllability Gramian of Σ on [0, T − t0 ], we can
conclude that in the time-invariant case the controllability Gramian of Σ on [t0 , T ] depends only on the
difference T − t0 , and not independently on T and t0 . For this reason, in the time-invariant case t0 is usually
taken to be zero. We will adopt this practice in the sequel.

As we’ve already seen, the image of Σ’s controllability Gramian on a given interval [t0 , T ] of positive
length is the subspace R[t0 , T ] of states of Σ reachable from the zero state on [t0 , T ]. Evidently this subspace
too, only depends on the time difference T − t0 , and not separately on t0 and T . In fact, R[t0 , T ] doesn’t
even depend on T − t0 , provided this difference is a positive number. Our aim is to explain why this is so
and at the same time to provide an explicit algebraic characterization of R[t0 , T ] expressed directly in terms
of A and B. Towards this end, let us define the subspace

< A|B >= B + AB + · · · A(n−1) B
where n is the dimension of Σ and for i ≥ 0 Ai B is the image or column span of Ai B. We call < A|B > the
controllable space of Σ or of the pair (A, B). Let us also agree to call the n × nm matrix
[B AB · · · A(n−1) B ]
the controllability matrix of Σ or of the pair (A, B). We see at once that the controllable space of (A, B)
is the image or column span of the controllability matrix of (A, B). The following theorem establishes the
significance these definitions.

Theorem 25 For all T > t0 ≥ 0


R[t0 , T ] =< A|B > (11.15)

The theorem implies that the set of states of a time-invariant which are reachable from the zero state on
any interval of positive length, does not depend on either the length of the interval or on when the interval
begins. The proof of Theorem 25 depends on the following lemma.

Lemma 2 For each fixed time t1 > 0


Z t1

< A|B >= image eAt BB ′ eA t dt (11.16)
0

Proof of Theorem 25: By definition,


Z T

GC (t0 , T ) = e(T −τ ) BB ′ eA (T −τ ) dτ
t0

By making the change of variable t = T − τ we see that


Z (T −t0 )

GC (t0 , T ) = eAt BB ′ eA t dt
0

From this and Theorem 22 we see that


Z (T −t0 )

R[t0 , T ] = image eAt BB ′ eA t dt
0


Application of Lemma 2 with t1 = T − t0 then yields (11.15) which is the desired result.

Proof of Lemma 2: Fix t1 > 0 and define


Z t1

M= eAt BB ′ eA t dt (11.17)
0

As pointed out earlier in the math notes, et1 A can be expressed as


n
X
etA = γi (t)A(i−1)
i=1

for suitable functions γi (t). It is therefore possible to write M as


Z t1 n
X ′
M= γi (t)A(i−1) BB ′ eA t dt
0 i=1

Therefore
n
X
M= A(i−1) BNi
i=1
where Z t1

Ni = γi (t)B ′ eA t dt, i ∈ {1, 2, . . . , n}
0

It follows that
n
X n
X n
X
image M = image A(i−1) BNi ⊂ image A(i−1) BNi ⊂ image A(i−1) B =< A|B >
i=1 i=1 i=1

For the reverse inclusion, let x be any vector in kernel M . Then M x = 0, x′ M x = 0,


Z t1 

x′ eAt BB ′ eA t dt x = 0
0

and so Z t1

||B ′ eA t x||2 dt = 0
0

It follows that

B ′ eA t x = 0, ∀t ∈ [0, t1 ]
Repeated differentiation yields

B ′ (A′ )(i−1) eA t x = 0, ∀t ∈ [0, t1 ], i ∈ {1, 2, . . . , n}

Evaluation at t = 0 thus provides

B ′ (A′ )(i−1) x = 0, i ∈ {1, 2, . . . , n}

Clearly
R′ x = 0
where R is the controllability matrix

R = [B AB · · · A(n−1) ]

Therefore x ∈ kernel R′ and since x was chosen arbitrarily in kernel M , it must be true that kernel M ⊂
kernel R′ . Therefore, taking orthogonal complements we have

image R ⊂ image M ′

But M is symmetric and image R =< A|B > so

< A|B >⊂ image M

which is the required reverse containment.

In the light of Theorems 23 and 25 we can now state

Theorem 26 An n- dimensional time-invariant linear system {A, B, C, D} is controllable if and only if

< A|B >= IRn

From this and the fact that the controllable space of (A, B) is the image of the controllability matrix of
(A, B) we can also state the following.
Corollary 7 An n- dimensional time invariant linear system {A, B, C, D} is controllable if and only if

rank [ B AB A2 B · · · A(n−1) B ] = n

We’d like to illustrate next that controllability is a “generic” or typical property of a time-invariant
linear system. To make this a little more precise, suppose we write M for the linear vector space of all
real matrix pairs (An×n , Bn×m ), with addition and scalar multiplication of such pairs defined in the most
natural way. We’d like to show that “almost every” pair in M is controllable. To do this we first provide a
characterization of those pairs which are not controllable. To do this, let us first note that each nth order
minor of the controllablity matrix of (A, B) is a polynomial function of the n2 + nm entries of A and B. Let
µ denote the sum of the squares of all such minors. Then µ is also a polynomial function of the n2 + nm
entries of A and B. Moreover µ will vanish for some pair (A, B), just in case all of the nth order minors
of (A, B)’s controllabilty matrix vanish. Now for (A, B) not to be controllable, all nth order minors must
vanish. This follows from Corollary 7 and from the definition of rank. Thus the set of pairs (A, B) which are
not controllable is the same as the set of (A, B) pairs for which µ vanishes. This set is thus a strictly proper
algebraic subset of M. Its complement – let’s call it S – is the set of all controllable pairs in M. Because µ
is a continuous function, S must be an open subset in M. Moreover, from the fact that µ is a polynomial it
can also be shown that S is dense2 in M. In this sense controllable pairs in M are “almost all” pairs in M.

11.3.1 Properties of the Controllable Space of (A, B)

The controllable space of (A, B) turns out to have several important properties which we will make use of in
the next section. The first, and perhaps most important of these is that < A|B > is an A-invariant subspace.

Lemma 3 For each matrix pair (An×n , Bn×m ),

A < A|B >⊂< A|B > (11.18)

Proof: Fix x ∈< A|B >. Thus there must be vectors bi ∈ B, i ∈ {1, 2, . . . , n} such that

x = b1 + Ab2 + · · · + A(n−2) bn−1 + A(n−1) bn

Then
Ax = Ab1 + A2 b2 + · · · + A(n−1) bn−1 + An bn
Now by the Cayley-Hamilton theorem

An = −a1 I − a2 A − · · · an A(n−1)

where sn + an s(n−1) + · · · + a2 s + a1 is the characteristic polynomial of A. Therefore

Ax = −a1 bn + A(b1 − a2 bn ) + A2 (b2 − a3 bn ) + · · · + A(n−1) (bn−1 − an bn )

Thus Ax ∈< A|B >. Since this is true for all x ∈< A|B >, (11.18) must be true.

By the controllability index of (An×n , Bn×m ) is meant the smallest positive integer nc for which

< A|B >= B + AB + . . . A(nc −1) B

Clearly nc ≤ n, with the inequality typically being strict if m > 1. The following is a consequence of the
preceding lemma and the Cayley-Hamilton theorem.
2 This means that each matrix pair in M can be approximated with arbitrarily small error by a controllable pair in M.
Lemma 4 For each matrix pair (An×n , Bn×m ),

< A|B >= B + AB + . . . A(k−1) B, ∀k ≥ nc (11.19)

where nc is the controllablility index of (A, B).

Proof: In view of the definition of nc it is enough to prove that

< A|B >= B + AB + . . . A(k−1) B, (11.20)

holds for all k > n. To prove that this is so we will use induction. We know, first of all that (11.20) is true
if k = n. Suppose therefore that i ≥ n is a fixed integrer such that (11.20) holds for all n ≤ k ≤ i. Our goal
is use this to prove that that (11.20) also holds for k = i + 1. For this, we first note that

B + AB + . . . Ai B = B + A(B + AB + . . . A(i−1) B)

But by the inductive hypothesis,


B + AB + . . . A(i−1) B =< A|B >
so
B + AB + . . . Ai B = B + A < A|B >
From this and the A-invarance of < A|B > , there follows

B + AB + . . . Ai B ⊂ B+ < A|B >

But since B ⊂< A|B >, it must be true that B+ < A|B >=< A|B >. Therefore

B + AB + . . . Ai B ⊂< A|B >

Since the reverse containment is clearly true, it follows that (11.20) must hold for k = i + 1. Hence, by
induction (11.19) is true.

There is an interesting way to characterize a controllable space, which differs from what we’ve discussed
so far. Let (A, B) be given and fixed. Suppose we define C to be the class of all subspaces of IRn which are
A-invariant and which contain B. That is

C = {S : AS ⊂ S, B ⊂ S}

Note that IRn ∈ C so C is nonempty. Moreover it is easy to show that if S ∈ C and T ∈ C, then S ∩ T ∈ C.
In other words, C is closed under subspace intersection. Because of this and the finite dimensionality of IRn ,
it is possible to prove that C contains a unique smallest element which is a subspace of every other subspace
in C. In fact, < A|B > turns out to be this unique smallest subspace. In other words, < A|B > contains
B, is A-invariant, and is contained in every other subspace which contains B and is A-invariant. The reader
may wish to prove that this is so.

11.3.2 Control Reduction for Time-Invariant Systems

Although the control reduction technique described earlier also works in the time-invariant case, it is possible
to accomplish the same thing in a more algebraic manner without having to use the controllability gramian.
What we shall do is to explain how to construct from a given time-invariant linear system, a new time-
invariant linear system which is controllable and which has the same transfer matrix as the system we
started with. To carry out this program we shall make use of the following fact, which is of interest in its
own right.
Lemma 5 Let {An×n , Bn×m , Cp×n } and {Ān̄×n̄ , B̄n̄×m , C̄p×n̄ } be two constant matrix triples. If there exists
a matrix Tn×n̄ such that
AT = T Ā B = T B̄ CT = C̄ (11.21)
then
C(sI − A)−1 B = C̄(sI − Ā)−1 B̄

Proof: Since AT = T Ā, it must be true that (sI − A)T = T (sI − Ā) and thus that

T (sI − Ā)−1 = (sI − A)−1 T

Therefore
C(sI − A)−1 B = C(sI − A)−1 T B̄ = CT (sI − Ā)−1 B̄ = C̄(sI − Ā)−1 B̄
which completes the proof.

The lemma implies that any two time-invariant linear systems whose coefficient matrices are related as
in (11.21), have the same transfer matrix. It is important to note that such systems need not have the same
dimension, and that T need not be nonsingular or even square. Suppose that that (11.21) holds and that
Σ̄’s is a system of the form
x̄˙ = Āx̄ + B̄u y = C̄ x̄
If we then define

x = T x̄
then it is possible to write
ẋ = Ax + Bu y = Cx
which is a new linear system Σ modeling the same input-output relationship between u and y. In the event
that Σ and Σ̄ do have the same dimension {i.e., n = n̄} and T is nonsingular, (11.21) yields

T −1 AT = Ā T B = B̄ CT −1 = C̄

Under these conditions it is natrual to say that Σ and Σ̄ are similar linear systems.

We now turn to the problem of constructing from a given time-invariant linear system Σ = {A, B, C}, a
new time-invariant linear system Σ̄ which is controllable and which has the same transfer matrix as Σ. As

a first step, define Σ̄’s dimension n̄ = dim < A|B >. Next pick a basis matrix Rn×n̄ for < A|B > . Because
< A|B > is A-invariant, the equation
AR = RĀ (11.22)
has a unique solution Ān̄×n̄ . Moreover, because B ⊂< A|B >, the equation

B = RB̄ (11.23)

also has a unique solution B̄. Define



C̄ = CR (11.24)

It follows from (11.22)-(11.24) and Lemma 5 that Σ̄ = {Ā, B̄, C̄} has the same transfer matrix as Σ. Now
from (11.22), Ai R = RĀi , i ≥ 1. This and (11.23) imply that

Ai B = RĀi B̄, i ≥ 0

Therefore
[B AB · · · A(n−1) ]n×m = Rn×n̄ [ B̄ ĀB̄ . . . Ā(n−1) B̄ ]n̄×m
This implies that
rank [ B AB · · · A(n−1) ]n×m ≤ rank [ B̄ ĀB̄ . . . Ā(n−1) B̄ ]n̄×m
But
rank [ B AB · · · A(n−1) ]n×m = n̄
and
rank [ B̄ ĀB̄ . . . Ā(n−1) B̄ ]n̄×m ≤ n̄
so
rank [ B̄ ĀB̄ . . . Ā(n−1) B̄ ]n̄×m = n̄
This is equivalent to
B̄ + ĀB̄ + · · · Ā(n−1) B̄ = IRn̄
But from Lemma 4,
B̄ + ĀB̄ + · · · Ā(n−1) B̄ =< Ā|B̄ >
Therefore
< Ā|B̄ >= IRn̄
so (Ā, B̄) is a controllable pair. Hence Σ̄ is a controllable system with the same transfer matrix as Σ.

11.3.3 Controllable Decomposition

The preceding discussion was concerned solely with the problem of constructing from a given linear system,
a new controllable linear system with the same transfer matrix as the original. Our aim now is to represent
the original system
ẋ = Ax + Bu y = Cx (11.25)
in a new coordinate system within which {Ā, B̄, C̄} appears naturally as a subsystem. Toward this end, let

S be any n × (n − n̄) matrix for which T = [ R S ] is nonsingular. Define
 
Ãn̄×(n−n̄) ∆ −1 b=∆
b(n−n̄)×(n−n̄) = T AS
A
C CS

Recall that
AR = RĀ, B = RB̄ CR = C̄
From these expressions it follows that
   
−1 Ā Ã −1 B̄  
T AT = T B= CT = C̄ b
C
0 A b 0
Hence if we define z = T −1 x, and partition z as
 


z=
zb
where z̄ is a n̄-vector, then we can write
z̄˙ = Āz̄ + Ãb
z + B̄u (11.26)
zḃ = b
Abz (11.27)
y = b zb
C̄ z̄ + C (11.28)
Note that u has no influence on zb and moreover that if zb(0) = 0, then the input-output relationship is
determined solely by the controllable subsystem
z̄˙ = Āz̄ + B̄u y = C̄ z̄
b is is uniquely determined by (A, B) and is sometimes called the uncontrollable spectrum
The spectrum of A
of (A, B).
Single-Input Systems


Let (An×n , bn×1 ) be a controllable matrix pair. This means that the controllablilty matrix Q = [ b Ab · · · A(n−1) b ]
has rank n and thus that {b, Ab, . . . A(n−1) b} is a basis for IRn . In other words,

Lemma 6 An n-dimensional, single input pair (A, b) is controllable if and only if A is cyclic and b is a
generator for IRn .

By means of similarity, we can transform such a pair (A, b) into an especially useful form. Towards this
end define the row vector

h = [ 0 0 · · · 0 1 ]1×n Q−1
Then
hQ = [ 0 0 ··· 0 1 ]1×n
so
hb = 0, hAb = 0, . . . . . . . . . hAn−2 b = 0 (11.29)
and
hAn−1 b = 1 (11.30)
Define the matrix  
h
∆  hA 
T = 
 ...  (11.31)
hAn−1 n×n

and note that  


0
0
.
Tb =  .
. (11.32)
0
1 n×1

We claim that T is nonsingular. Suppose it were not; then it would be possible to find numbers gi such
that
n−1
X
hAn−1 = gi hAi−1
i=1

This would imply that


n−1
X
hAn−1 b = gi hAi−1 b = 0
i=1

because of (11.29). But this contradicts (11.30), so T cannot be singular.

In view of the definition of T in (11.31)



hA
 hA2 
TA =  
 ...  (11.33)
hAn

and
hAi = e′i+1 T, i ∈ {1, 2, . . . , n − 1} (11.34)
where ei+1 is the ith unit vector in IKn . Meanwhile, by the Cayley-Hamilton Theorem

hAn = −a1 h − a2 hA − · · · − an hAn−1

where
α(s) = sn + an sn−1 + · · · + a2 s + a1
is the characteristic polynomina of A. Hence

hAn = [ −a1 −a2 · · · −an ] T (11.35)

because of the definition of T in (11.31). From (11.33) - (11.35) we thus obtain the expression

 e′2 T 
 e′3 T 
 .. 
TA = 
 .


 
e′n T
[ −a1 −a2 · · · −an ] T
Factoring out T on the right we get
T A = AC T (11.36)
where  
0 1 0 ··· 0
 0 0 1 ··· 0 
 . .. .. .. .. 
AC = 
 .
. . . . . 

 0 0 0 ··· 1 
−a1 −a2 −a3 · · · −an n×n

From this and (11.32) we conclude that



AC = T AT −1 bC = T b (11.37)

where  
0
 .. 

bC =  .
0
1 n×1
(AC , bC ) is uniquely determined by (A, b) and is called the control canonical form of the single-input con-
trollable pair (A, b). Note that (AC , bC ) is also uniquely determined by A’s characteristic polynomial α(s).
Control canonical forms find applications in several different places. One is in realizing transfer functions
Another is in defining a state-feedback to assign to a “closed-loop” linear system a prescribed spectrum. We
will discuss theses applications in later chapters.
Chapter 12

Observability

In this chapter we explore the extent to which the state of a continuous time linear system can be recovered
from available measured signals, namely u and y. In doing the we will come up with a concept, namely
observability, which plays a role dual to controllability in the realization and representation of a linear
system. We begin with a discussion of unobservable states.

12.1 Unobservable States

Let θ(t, t0 , x0 ; v) denote the value of the output of the n-dimensional linear system

Σ: y = C(t)x + D(t)u ẋ = A(t)x + B(t)u (12.1)

at time t, assuming the system was in state x0 at time t0 and that v(t) was the input applied on the time
interval [t0 , t]. Then, by the variation of constants formula,
Z t
θ(t, t0 , x0 ; v) = C(t)Φ(t, t0 )x0 + W (t, τ )v(τ )dτ + D(t)v(t) (12.2)
t0

where Φ(t, τ )is the state transition matrix of A(t) and W (t, τ ) is the weighting pattern

W (t, τ ) = C(t)Φ(t, τ )B(τ )

Fix T > t0 ≥ 0. Two states x1 and x2 of Σ are said to be indistinguishable on [t0 , T ], if for all inputs v(t)
which are defined and piecewise-continuous in [t0 , T ],

θ(t, t0 , x1 ; v) = θ(t, t0 , x2 ; v), ∀t ∈ [t0 , T ] (12.3)

Thus two states are indistinguishable just in case both determine the same mapping of Σ’s inputs into its
outputs on [t0 , T ]. This means that if x1 were x2 are indistinguishable, it would impossible to decide from
input-output measurements on [t0 , T ] whether Σ was in state x1 or x2 at time t0 . It is easy to verify that
indistinguishability is an equivalence relation on Σ’s state space IRn . The most one could expect to learn
from input-output measurements on [t0 , T ] about the state x0 of Σ at time t0 , is thus the indistinguishablilty
equivalence class within which x0 resides.

In the light of (12.2), it is clear that x1 and x2 are indistinguishable if and only if

C(t)Φ(t, t0 )(x1 − x2 ) = 0 ∀t ∈ [t0 , T ] (12.4)

187
This suggests that the set of states indistinguishable from the zero state ought to be easy to characterize.
Accordingly, write

N [t0 , T ] = {x : C(t)Φ(t, t0 )x = 0, ∀t ∈ [t0 , T ]}
and note that N [t0 , T ] is a subspace of IRn . N [t0 , T ] is thus the set of states of Σ which produce the zero
output on [t0 , T ] when Σ is initialized at any such state at time t0 and which the zero input is applied on
[t0 , T ]. We call N [t0 , T ] the unobservable space of Σ {or of (C(t), A(t))} on [t0 , T ]. Note that two states x1
and x2 of Σ are indistinguishable on [t0 , T ] just in case their difference x1 − x2 is an unobservale state in
N [t0 , T ].

Mathematically N [t0 , T ] possesses a set of properties “dual” to those of the reachable space of Σ discussed
earlier.

Theorem 27
N [t0 , T ] = kernel GO (t0 , T ) (12.5)
where Z T

GO (t0 , T ) = Φ′ (τ, t0 )C ′ (τ )C(τ )Φ((τ, t0 )dτ (12.6)
t0

The matrix GO (t0 , T ) defined by (12.6) is called the observability Gramian of (C(t), A(t)) or of Σ on [t0 , T ].
Note that GO (t0 , T ) is a symmetric, positive semi-definite matrix; in addition it is and a non-decreasing
function of T is the sense that
GO (t0 , T2 ) ≥ GO (t0 , T1 ) if T2 ≥ T1
From this and Theorem 27 it follows that Σ’s unobservable space is a non-increasing function of T in the
sense that
N [t0 , T2 ] ⊂ N [t0 , T1 ] if T2 ≥ T1
This means that any state unobservable on [t0 , T2 ] is also unobservable on [t0 , T1 ] provided T2 ≥ T1 .

Proof of Theorem 27: First suppose that x ∈ N [t0 , T ]. Then


C(t)Φ(t, t0 )x = 0, ∀t ∈ [t0 , T ]
Thus
Φ′ (t, t0 )C ′ (t)C(t)Φ(t, t0 )x = 0, ∀t ∈ [t0 , T ]
Therefore Z T
Φ′ (τ, t0 )C ′ (τ )C(τ )Φ(τ, t0 )xdτ = 0
t0
From this and the definition of GO (t0 , T ) in (12.6) it follows that GO (t0 , T )x = 0. Therefore x ∈ kernel GO (t0 , T ).
Since x is chosen arbitrarily, we have proved that
N [t0 , T ] ⊂ kernel GO (t0 , T ) (12.7)

To prove the reverse inclusion, now fix x ∈ kernel GO (t0 , t). Then
Z T !
′ ′
Φ (τ, t0 )C (τ )C(τ )Φ(τ, t0 )dτ x = 0
t0

Therefore Z T
||C(τ )Φ(τ, t0 )x||2 dτ = 0
t0
so
C(t)Φ(t, t0 )x = 0, ∀t ∈ [t0 , T ]
It follows that x ∈ N [t0 , T ]. Thus kernel GO (t0 , T ) ⊂ N [t0 , T ]. Therefore (12.5) is true.
12.2 Observability

A linear system Σ = {A(t), B(t), C(t), D(t)} {or matrix pair (C(t), A(t)) } is said to be observable on [t0 , T ]
if for each initial state x(t0 ) = x0 and each input v, defined on [t0 , T ], it is possible to uniquely determine x0
from from knowledge of v on [t0 , T ] and Σ’s response y on [t0 , T ]. In other words, Σ is observable on [t0 , T ]
just in case for each pair x1 , x2 ∈ IRn and each input v,

x1 = x2 whenever θ(t, t0 , x1 ; v) = θ(t, t0 , x2 ; v), ∀t ∈ [t0 , T ]


where, as before, Z t
θ(t, t0 , x0 ; v) = C(t)Φ(t, t0 )x0 + W (t, τ )v(τ )dτ + D(t)v(t) (12.8)
t0
Thus an observable linear system on a given time interval, is a linear system for which it is possible to
uniquely determine the system’s state at the beginning of the interval from measurements of the system’s
input and output on the interval.

Let us note that if Σ is observable on [t0 , T ], then N [t0 , T ] must be the zero subspace – for if N [t0 , T ]
contained two distinct vectors, and if Σ were initialized at one of them at t = t0 , then it would be impossible
to uniquely recover this initial state from knowledge of u and y, because distinct states in N [t0 , T ] are
indistinguisble on [t0 , T ]. In other words, if Σ is observable on [t0 , T ], then it must be true that N [t0 , T ] = 0.
The converse is also true:

Theorem 28 A linear system Σ with coefficient matrix quadruple {A(t), B(t), C(t), D(t)} is observable on
[t0 , T ] if and only if
N [t0 , T ] = 0 (12.9)

In view of Theorem 27, we see that observablilty of Σ on [t0 , T ] is equivalent to kernel GO (t0 , T ) = 0. Since
the latter, in turn, is equivalent to det GO (t0 , T ) 6= 0, we are led to the following alternative characterization
of observability.

Corollary 8 A linear system with coefficient matrix quadruple {A(t), B(t), C(t), D(t)} is observable on
[t0 , T ] if and only if its observability Gramian GO (t0 , T ) is nonsingular.

Proof of Theorem 28: As we’ve already explained, just before the statement of Theorem 28, observability
of Σ on [t0 , T ] implies that N [t0 , T ] = 0. To prove the converse, suppose that N [t0 , T ] = 0, that x0 is any
given initial state of Σ at t = t0 , and that v is any input which is defined and piece-wise continuous on
[t0 , T ]. Then by the variation of constants formula
Z t
y(t) = C(t)Φ(t, t0 )x0 + W (t, τ )v(τ )dτ + D(t)v(t), t ∈ [t0 , T ]
t0

Therefore
z(t) = C(t)Φ(t, t0 )x0 , t ∈ [t0 , T ] (12.10)
where Z t

z(t) = y(t) − W (t, τ )v(τ )dτ + D(t)v(t), t ∈ [t0 , T ]
t0

Since y and v are known signals, so is z. To complete the proof, it is enough to show that we can solve (12.10)
uniquely for x0 . Towards this end, we multiply both sides of (12.10) by Φ′ (t, t0 )C ′ (t) and then integrate.
What results is Z Z
T T
Φ′ (τ, t0 )C ′ (τ )z(τ )dτ = Φ′ (τ, t0 )C ′ (τ )C(τ )Φ(τ, t0 )x0 dτ
t0 t0
In view of the definition of the observability Gramian in (12.6), this equation becomes
Z T
Φ′ (τ, t0 )C ′ (τ )z(τ )dτ = GO (t0 , T )x0 (12.11)
t0

But N [t0 , T ] = 0 by assumption and kernel GO [t0 , T ] = N [t0 , T ] because of Theorem 27. Hence GO [t0 , T ] is
nonsingular and thus invertible. It follows from (12.11) that x0 is uniquely recovered by the formula
Z T
x0 = G−1
O (t0 , T ) Φ′ (τ, t0 )C ′ (τ )z(τ )dτ
t0

Thus Σ is observable.

12.2.1 Observability Reduction


Let Σ = {A(t), B(t), C(t)} be a given n-dimensional continuous time linear system and let

W (t, τ ) = C(t)Φ(t, τ )B(τ )

be its weighting pattern where Φ(t, τ ) is the state transition matrix of A(t). Let [t0 , T ] be any interval of
positive length on which Σ is defined. The aim of this section is to prove that every weighting pattern has
an observable realization. We will do this constructively by explaining how to use Σ to compute a linear

system Σ̄ = {Ā(t), B̄(t), C̄(t)} which is observable on [t0 , T ] and which has the same weighting pattern as Σ.
Since Σ̄’s dimension will turn out to be less than Σ’s, if Σ is not observable, this construction is referred to
as an observability reduction.

To begin, let us suppose right away that Σ is not observable. For if Σ were observable, we could simply
define Σ̄ to be Σ and we would have at once an observable realization of W (t, τ ). To proceed, let GO denote
Σ’s observability Gramian on [t0 , T ]. In other words
Z T

GO = Φ′ (τ, t0 )C ′ (τ )C(τ )Φ(τ, t0 )dτ
t0

As noted earlier when we discussed controllability, the symmetry of GO implies that IRn admits an orthogonal
decomposition of the form
IRn = image GO ⊕ kernel GO
Let P denote the orthogonal projection matrix on image GO . Hence P GO = GO and therefore G0 P ′ = GO .

Because GO is positive-semidefinite it is possible to construct a full rank matrix Rn×n̄ , with n̄ = rank GO ,
such that
GO = RR′
Our construction relies on the fact that P can be written as

P = R(R′ R)−1 R′

Now consider the new system Σ̄ defined for t ∈ [t0 , T ] by

Ā(t) = 0 B̄(t) = R′ Φ(t0 , t)B(t) C̄(t) = C(t)Φ(t, t0 )S

where

S = R(R′ R)−1
It will now be shown that Σ̄ is observable and that its weighting pattern is W (t, τ ). We establish observability
first. Toward this end, note that Ā’s state transition matrix is Φ̄(t, τ ) = In̄×n̄ . Using this fact we can write
Σ̄’s observability Gramian
Z T
ḠO = Φ̄′ (t, t0 )C̄ ′ (t)C̄(t)Φ̄(t, t0 )dt
t0
as Z T
ḠO = C̄ ′ (t)C̄(t)dt
t0

Hence, using the definition of C̄(t)


Z T
ḠO = S ′ Φ′ (t, t0 )C ′ (t)C(t)Φ(t, t0 )Sdt
t0
Z !
T
′ ′ ′ ′
= S Φ (t, t0 )C (t)C(t)Φ (t, t0 )dt S
t0

= S ′ GO S
= (R′ R)−1 R′ RR′ R(R′ R)−1
= In̄×n̄

This prove that ḠO (t0 , T ) is nonsingular. Therefore by Corollary 8, Σ̄ is observable on [t0 , T ] as claimed.

To prove that Σ̄’s weighting pattern is the same as Σ’s, we first need to show that

C(t)Φ(t, t0 )P = C(t)Φ(t, t0 ), t ∈ [t0 , T ] (12.12)

Now this will be true provided


Z T
||(C(t)Φ(t, t0 )′ P − C(t)Φ(t, t0 ))z||2 dt = 0, ∀z ∈ IRn (12.13)
t0

Fixing z arbitrarily and expanding the integral on the left above we get
Z T Z T
||(C(t)Φ(t, t0 )P − C(t)Φ(t, t0 ))z||2 dt = z ′ P ′ Φ′ (t, t0 )C ′ (t)C(t)Φ(t, t0 )P zdt
t0 t0
Z T
− z ′ Φ′ (t, t0 )C ′ (t)C(t)Φ′ (t, t0 )P zdt
t0
Z T
− z ′ P ′ Φ′ (t, t0 )C ′ (t)C(t)Φ′ (t, t0 )zdt
t0
Z T
+ z ′ Φ′ (t, t0 )C ′ (t)C(t)Φ′ (t, t0 )zdt
t0
= z ′ P ′ GO P z − z ′ GO P z − z ′ P ′ GO z + z ′ GO z
= z ′ GO z − z ′ GO z − z ′ GO z + z ′ GO z
= 0

Hence (12.13) is true. Consequently (12.12) is true as well.

It remains to be shown that Σ̄’s weighting pattern is the same as Σ’s. This is now easily accomplished
by noting that Σ̄’s weighting pattern W̄ satisfies

W̄ (t, τ ) = C̄(t)Φ̄(t, τ )B̄(τ )


= C̄(t)B̄(τ )
= C(t)Φ(t, t0 )SR′ Φ(t0 , τ )B(τ )
= C(t)Φ(t, t0 )P Φ(t0 , τ )B(τ )
= C(t)Φ(t, t0 )Φ(t0 , τ )B(τ )
= C(t)Φ(t, τ )B(τ )
= W (t, τ )
Therefore Σ̄ and Σ have the same weighting pattern as claimed. We summarize.

Theorem 29 Every continuous-time weighting pattern admits an observable realization.

Suppose now that Σ is controllable. Our aim is to show that controllability is not lost under an observ-
ability reduction. In other words, we want to verify that Σ̄ is controllable. To accomplish this we look at
Σ̄’s controllability Gramian
Z T
ḠC = Φ̄(T, t)B̄(t)B̄ ′ (t)Φ̄′ (T, t)dt
t0
Since Φ̄(t, τ ) = I,
Z T
ḠC = B̄(t)B̄ ′ (t)dt
t0
Hence, using the definition of B̄(t)
Z T
ḠC = R′ Φ(t0 , t)B(t)B ′ (t)Φ′ (t0 , t)Rdt
t0
Z T
= R′ Φ(t0 , T )Φ(T, t)B(t)B ′ (t)Φ′ (T, t)Φ′ (t0 , T )Rdt
t0
Z !
T
= R′ Φ(t0 , T ) Φ(t0 , t)B(t)B ′ (t)Φ′ (t0 , t)dt φ′ (t0 , T )R
t0

= R′ Φ(t0 , T )GC (t0 , T )Φ′ (t0 , T )R


But Rn×n̄ has full rank and Φ(t0 , T ) is nonsingular so Φ′ (t0 , T )R has full rank as well. Moreover, controlla-
bility of Σ implies that GC (t0 , T ) is positive definite. Hence R′ Φ(t0 , T )GC (t0 , T )Φ′ (t0 , T )R is also positive
definite and thus nonsingular. This proves that ḠC (t0 , T ) is nonsingular and thus that Σ̄ is controllable. Thus
we see that reducing a controllable system to an observable system results in a system which is controllable
as well.

It is now clear that we have a procedure for realizing any weighting pattern as a controllable and observ-
able system. First realize the weighting pattern. Second reduce the realization to one which is controllable.
Third reduce the controllable realization to one which is observable. It can be shown that the process also
works if steps two and three are interchanged. We summarize.

Theorem 30 Each weighting pattern admits a controllable and observable realization.

In the next section we prove that this procedure also provides us with a “minimal realization” of any weighting
pattern.

12.3 Minimal Systems



Let us agree to call a system Σ = {A(t), B(t), C(t)} minimal if the dimension of Σ’s state space is at least as
small as the dimension of the state space of any linear system with the same weighting pattern as Σ. Thus
Σ is minimal if its dimension is the least among the dimensions of all linear systems with the same weighting
pattern. The following theorem is the centerpiece of linear realization theory.

Theorem 31 A linear system is minimal if and only if it is controllable and observable.

Proof: If Σ is minimal then it must be both controllable and observable. For it it were not controllable
{observable}, a lower dimensional system with the same weighting pattern could be constructed via a con-
trollability {observability} reduction.

To prove the converse, suppose that Σ is controllable and observable and that Σ̄ = {Ā(t), B̄(t), C̄(t)} any
linear system with the same weighting pattern as Σ. Let n and n̄ denote the state space dimensions of Σ
and Σ̄ respectively. It is enough to prove that n ≤ n̄.

Since Σ and Σ̄ have the same weighting pattern,

C(t)Φ(t, τ )B(τ ) = C̄(t)Φ̄(t, τ )B̄(τ )

Using the composition rule for state transition matrices we can write

C(t)Φ(t, t0 )Φ(t0 , T )Φ(T, τ )B(τ ) = C̄(t)Φ̄(t, t0 )Φ̄(t0 , T )Φ̄(T, τ )B̄(τ )

Multiplying both sides of this equation on the left by Φ′ (t, t0 )C ′ (t) and on the right by B ′ (τ )Φ′ (T, τ ) one
gets

Φ′ (t, t0 )C ′ (t)C(t)Φ(t, t0 )Φ(t0 , T )Φ(T, τ )B(τ )B ′ (τ )Φ′ (T, τ ) =


Φ′ (t, t0 )C ′ (t)C̄(t)Φ̄(t, t0 )Φ̄(t0 , T )Φ̄(T, τ )B̄(τ )B ′ (τ )Φ′ (T, τ )

Integrating both sides with respect to both t and τ there folows

GO (t0 , T )Φ(t0 , T )GC (t0 , T ) = M Φ̄(t0 , T )N (12.14)

where
Z T Z T
∆ ∆
GO (t0 , T ) = Φ′ (t, t0 )C ′ (t)C(t)Φ(t, t0 )dt GC (t0 , T ) = Φ(T, τ )B(τ )B ′ (τ )Φ′ (T, τ )dτ
t0 t0

Z T Z T
∆ ∆
Mn×n̄ = Φ′ (t, t0 )C ′ (t)C̄(t)Φ̄(t, t0 )dt Nn̄×n = Φ̄(T, τ )B̄(τ )B ′ (τ )Φ′ (T, τ )dτ
t0 t0

Since Σ is controllable and observable, both GC (t0 , T ) and GO (t0 , T ) are nonsingular n × n matrices. Since
Φ(t0 , T ) also is nonsingular, the matrix on the left in (12.14) is nonsingular. It follows that M Φ̄(t0 , T )N is
nonsingular or equivalently that
rank M Φ̄(t0 , T )N = n
But Φ̄(t0 , T ) is an n̄ × n̄ matrix so n ≤ n̄.

12.4 Time-Invariant Continuous-Time Linear Systems

While the constructions regarding observability discussed so far, apply to both time-varying and time-
invariant linear systems, the consequential simplications of the time-invariant case are significant enough to
justfy a separate treatment.

Let Σ = {A, B, C, D} be a time-invariant system. Then Σ’s observability Gramian on [t0 , T ] is
Z T
∆ ′
GO (t0 , T ) = e(τ −t0 )A C ′ Ce(τ −t0 )A dτ
t0


By introducting the change of variables s = τ − t0 we see that
Z T −t0

GC (t0 , T ) = esA C ′ CesA ds
0

Since the right hand side of the preceding is also the observability Gramian of Σ on [0, T −t0 ], we can conclude
that in the time-invariant case the observability Gramian of Σ on [t0 , T ] depends only on the difference T −t0 ,
and not independently on T and t0 . For this reason, in the time-invariant case t0 is usually taken to be zero.
We will adopt this practice in the sequel.

As we’ve already seen, the kernel of Σ’s observability Gramian on a given interval [t0 , T ] of positive length
is the subspace of unobservable states N [t0 , T ] of Σ on [t0 , T ]. Evidently this subspace too, only depends on
the time difference T − t0 , and not separately on t0 and T . In fact, N [t0 , T ] doesn’t even depend on T − t0 ,
provided this difference is a positive number. Our aim is to explain why this is so and at the same time to
provide an explicit algebraic characterization of N [t0 , T ] expressed directly in terms of C and A. Towards
this end, let us define the subspace
\n

[C|A] = kernel CA(i−1)
i=1

where n is the dimension of Σ.

Let us also agree to call the pn × n matrix


 
C
 CA 
 .. 
 . 
CA(n−1)

the observability matrix of Σ or of the pair (C, A). We see at once that [C|A] is the kernel of the observability
matrix of (C, A). The following theorem establishes the significance these definitions.

Theorem 32 For all T > t0 ≥ 0


N [t0 , T ] = [C|A] (12.15)

The theorem implies that the set of states of a time-invariant which are unobservable on any interval of
positive length, does not depend on either the length of the interval or on when the interval begins. In view
of Theorem 32 we will also call [C|A] the unobservable space of Σ or of the pair (C, A).

The proof of Theorem 32 depends on the following lemma.

Lemma 7 For each fixed time t1 > 0


Z t1

[C|A] = kernel eA t C ′ CeAt dt (12.16)
0
Proof of Theorem 32: By definition,
Z T

GO (t0 , T ) = e(τ −t0 )A C ′ Ce(τ −t0 )A dτ
t0

By making the change of variable t = τ − t0 we see that


Z (T −t0 )

GO (t0 , T ) = eA t C ′ CeAt dt
0

From this and Theorem 27 we see that


Z (T −t0 )
′ ′
N [t0 , T ] = kernel eA t C ′ CeA dt
0


Application of Lemma 7 with t1 = T − t0 then yields (12.15) which is the desired result.

Proof of Lemma 7: Fix t1 > 0 and define


Z t1

M= eA t C ′ C ′ eAt dt (12.17)
0

As pointed out earlier in the math notes, et1 A can be expressed as


n
X
etA = γi (t)A(i−1)
i=1

for suitable functions γi (t). It is therefore possible to write M as


Z t1 n
X ′
M= γi (t)eA t C ′ CA(i−1) dt
0 i=1

Therefore
n
X
M= Ri CA(i−1)
i=1

where Z t1

Ri = γi (t)eA t C ′ dt, i ∈ {1, 2, . . . , n}
0
It follows that
n
X n
\ n
\
kernel M = kernel Ri CA(i−1) ⊃ kernel Ri CA(i−1) ⊃ kernel CA(i−1) = [C|A]
i=1 i=1 i=1

Thus kernel M ⊃ [C|A].

For the reverse inclusion, let x be any vector in kernel M . Then M x = 0, x′ M x = 0,


Z t1 
′ A′ t ′ At
x e C Ce dt x = 0
0

and so Z t1
||CeAt x||2 dt = 0
0
It follows that
CeAt x = 0, ∀t ∈ [0, t1 ]
Repeated differentiation yields

C(A)(i−1) eAt x = 0, ∀t ∈ [0, t1 ], i ∈ {1, 2, . . . , n}

Evaluation at t = 0 thus provides

C(A)(i−1) x = 0, i ∈ {1, 2, . . . , n}

Therefore x ∈ [C|A] and since x was chosen arbitrarily in kernel M , it must be true that kernel M ⊂ [C|A].
This is the required reverse containment.

In the light of Corollary 8 we can now state

Theorem 33 An n- dimensional time invariant linear system {A, B, C, D} is observable if and only if

[C|A] = 0

From this and the fact that the unobservable space of (C, A) is the kernel of the observability matrix of
(C, A) we can also state the following.

Corollary 9 An n- dimensional time invariant linear system {A, B, C, D} is observable if and only if
 
C
 CA 
 
rank  CA2 =n
 .. 
 . 
CA(n−1)

It has by now no doubt occured to the reader that there are certain algebraic similarities between the
ideas we’ve been discussing in this chapter and those discussed earlier in the chapter on controllability. For
example, the transpose of the observability matrix of (C, A) is the controllability matrix of the transposed
pair (A′ , C ′ ). The unobservable space of (C, A) is the orthogonal complement of the controllable space of
(A′ , C ′ ). The controllability Gramian of (A, B) and the observability Gramian of (C, A) are similar in form.
While these similarities are especially useful in gaining insights into proofs and constructions, they actually
have no system theoretic significance. We sometimes say that controllability and observability are dual
concepts. It is possible to be more precise about this by appealing to an idea from linear algebra called
“duality,” but the abstraction one must envoke hardly seems worth the effort1 .

12.4.1 Properties of the Unobservable Space of (C, A)

Like the controllable space of (A, B), the unobservable sapce of (C, A) has several important properties which
we will make use of in the next section. The first, and perhaps most important of these is that [C|A] is an
A-invariant subspace.

Lemma 8 For each matrix pair (Cp×n , An×n ),

A[C|A] ⊂ [C|A] (12.18)


1 The concept of duality being eluded to here, should not be confused with the very different and very useful concept of

duality which arises in convex analysis and linear programming.


Proof: Fix x ∈ [C|A]. Thus
CA(i−1) x = 0, i ∈ {1, 2, . . . n} (12.19)
Thus
CA(i−1) (Ax) = 0, i ∈ {1, 2, . . . n − 1} (12.20)
By the Cayley-Hamilton theorem An = −an A(n−1) − · · · − a2 A − a1 I, where sn + an s(n−1) + · · · a2 s + a1 is
the characteristic polynomial of A. Thus

CA(n−1) (Ax) = −an CA(n−1) x − · · · − a2 CAx − a1 Cx

so from (12.19), CA(n−1) (Ax) = 0. From this and (12.20) it follows that

CA(i−1) (Ax) = 0, i ∈ {1, 2, . . . n}

and thus that Ax ∈ [C|A]. Since this is true for all x ∈ [C|A], (12.18) must be true.

By the observability index of (Cp×n , An×n ) is meant the smallest positive integer no for which
no
\
[C|A] = kernel CA(i−1)
i=1

Clearly no ≤ n, with the inequality typically being strict if p > 1. The following is a consequence of the
preceding lemma and the Cayley-Hamilton theorem.

Lemma 9 For each matrix pair (Cp×n , An×n ),


   
C C
 CA   CA 
rank 
 ..  = rank 
  .. 
 ∀k ≥ no (12.21)
. .
CA(k−1) CA(n−1)

where no is the observability index of (C, A).

This lemma can be proved by exploiting duality and appealing to Lemma 4. The first step in this direction
would be to replace (12.21) with the equivalent expression

< A′ |C ′ >= C ′ + A′ C ′ + · · · + (A′ )(k−1) C ′


where C ′ = image C ′ . The reader may wish to develop a proof along these lines.

There is an interesting way to characterize an unobservable space, which differs from what we’ve discussed
so far. Let (C, A) be given and fixed. Suppose we define C to be the class of all subspaces of IRn which are
A-invariant and which are contained in kernel C. That is

C = {S : AS ⊂ S, S ⊂ kernel C}

Note that the zero subspace is in ∈ C so C is nonempty. Moreover it is easy to show that if S ∈ C and
T ∈ C, then S + T ∈ C. In other words, C is closed under subspace addition. Because of this and the finite
dimensionality of IRn , it is possible to prove that C contains a unique largest element which contains every
other subspace in C. In fact, [C|A] turns out to be this unique largest subspace. In other words, [C|A] is
contained in kernel C, is A-invariant, and contains in every other subspace which is contained in kernel C
and is A-invariant. The reader may wish to prove that this is so.
12.4.2 Observability Reduction for Time-Invariant Systems

Although the observability reduction technique described earlier, also works in the time-invariant case, it is
possible to accomplish the same thing in a more algebraic manner without having to use the observability
Gramian. What we shall do is to explain how to construct from a given time-invariant linear system, a new
time-invariant linear system which is observable and which has the same transfer matrix as the system we
started with. To carry out this program we shall make use of the following fundamental algebraic fact.

Lemma 10 Let Mn×m and Np×m be given real matrices. There exists a solution Xp×n to the linear algebraic
equation
XM = N (12.22)
if and only if
kernel M ⊂ kernel N (12.23)

Proof: The solvability condition for the transposed equation M ′ X ′ = N ′ , is image M ′ ⊃ image N ′ . Taking
orthogonal complements, this condition becomes
(image M ′ )⊥ ⊂ (imageN ′ )⊥ (12.24)
But (image M ′ )⊥ = kernel M and (image N ′ )⊥ = kernel N . Substitution of these identities into (12.24)
yields (12.23) whic is the desired result.

We now turn to the problem of constructing from a given time-invariant linear system Σ = {A, B, C},
a new time-invariant linear system Σ̄ which is observable and which has the same transfer matrix as Σ.

As a first step, define Σ̄’s dimension n̄ = n − dim[C|A]. Next pick and full rank matrix Pn̄×n such that
kernel P = [C|A]. An easy way to do this is to define P ′ to be a basis matrix for < A′ |C ′ >. We claim that
kernel P ⊂ kernel C kernel P A ⊂ kernel P
The first containment is a consequence of the fact that [C|A] ⊂ kernel C. The second containment is a
consequence of the fact that [C|A] is A-invariant. In any event, these two containments plus Lemma 10
imply that the equations
P A = ĀP C = C̄P

have {unique} solutions Ān̄×n̄ and C̄p×n̄ respectively. Define B̄ = P B. In summary, we have defined a linear

system Σ̄ = {Ā, B̄, C̄} whose coefficient matrices are related to those of Σ = {A, B, C} by the equations
P A = ĀP B̄ = P B C = C̄P (12.25)
In view of Lemma 5, we see at once that Σ and Σ̄ have the same transfer matrix.

We claim that (C̄, Ā) is an observable pair. To establish this, we first note from the fact P A = ĀP , that
P Ai = Āi P, i ≥ 0. Hence CAi = C̄P Ai = C̄ Āi P, i ≥ 0. Therefore
   
C C̄
 CA   C̄ Ā 
 .. = .. P
 .   . 
CA(n−1) C̄ Ā(n−1)
Thus    
C̄ C
 C̄ Ā   CA 
rank 
 ..  ≥ rank 
  ..  = n̄

. .
C̄ Ā(n−1) CA(n−1)
But the matrix on the left has n̄ columns so its rank cannot be larger than n̄. Therefore
 

 C̄ Ā 
rank 
 ..  = n̄

.
C̄ Ā(n−1)

Hence by Lemma 9
 

 C̄ Ā 
rank 
 ..  = n̄

.
C̄ Ā(n̄−1)
so (C̄, Ā) is an observable pair. Therefore Σ̄ is an observable system with the same transfer matrix as Σ.

Suppose now that Σ 1s controllable. Our aim is to show that controllability is not lost under an ob-
servability reduction, just as before in the time-varying case. In other words, we want to verify that Σ̄ is
controllable. To accomplish this we again use the fact that P Ai = Āi P, i ≥ 0 together with the previous
definition of B̄, namely B̄ = P B, to write

[ B̄ ĀB̄ · · · Ā(n−1) B̄ ] = P [ B AB · · · A(n−1) B ]

Since Σ is controllable, [ B AB · · · A(n−1) B ] has linearly independent rows. So does P . Since the
product of two matrices which each have linearly independent rows, is a matrix with linearly independent
rows, [ B̄ ĀB̄ · · · Ā(n−1) B̄ ] must have linearly independent rows. This means that

rank [ B̄ ĀB̄ · · · Ā(n−1) B̄ ] = n̄

Hence by Lemma 4
rank [ B̄ ĀB̄ · · · Ā(n̄−1) B̄ ] = n̄
which estabilshes the controllability of Σ̄. Thus we see that reducing a controllable system to an observable
system in this manner results in a system which is controllable, just as in the time-varying case.

It is now clear that once we have we have a procedure for realizing a transfer matrix as a time-invariant
linear system, we will also have a procedure for constructing a realization wich is both controllable and
observable. First realize the transfer matrix as a time-invariant system.. Second reduce the realization to
one which is controllable. Third reduce the controllable realization to one which is observable. It can be
shown that the process also works if steps two and three are interchanged. We will return to this important
topic in the next chapter where we explain how to realize a transfer matrix.

12.4.3 Observable Decomposition

The preceding discussion was concerned solely with the problem of constructing from a given linear system,
a new observable linear system with the same transfer matrix as the original. Our aim now is to represent
the original system
ẋ = Ax + Bu y = Cx (12.26)
in a new coordinate system within which {Ā,B̄, C̄}
 appears naturally as a subsystem. Toward this end, let
∆ P
S be any (n − n̄) × n) matrix for which T = is nonsingular. Define
S
 ∆ ∆
Ã(n−n̄)×n̄ b(n−n̄)×(n−n̄) =
A SAT −1 b=
B SB
Recall that
P A = ĀP, B̄ = P B C = C̄P
From these expressions it follows that
   
Ā 0 B̄
T AT −1 = b TB = b CT −1 = [ C̄ 0]
à A B

Hence if we define z = T x, and partition z as


 
∆ z̄
z=
zb

where z̄ is a n̄-vector, then we can write

z̄˙ = Āz̄ + B̄u (12.27)


zḃ = bz + Bu
Ãz̄ + Ab b (12.28)
y = C̄ z̄ (12.29)

Note that y does not depend on zb and moreover that if zb(0) = 0, then the input-output relationship is
determined solely by the observable subsystem

z̄˙ = Āz̄ + B̄u y = C̄ z̄


b is is uniquely determined by (C, A) and is sometimes called the unobservable spectrum
The spectrum of A
of (C, A).
Chapter 13

Transfer Matrix Realizations

The transfer matrix of a m-input, p-output, n-dimensional, time-invariant linear system {A, B, C, D} is a
p × m matrix of proper rational functions defined by the formula


T (s) = C(sI − A)−1 B + D (13.1)

Any m-input, p-output linear system {A, B, C, D} of any dimension for which (13.1) holds, is a realization
of T (s). The aim of this chapter is to explain how to go about constructing realizations of any given matrix
of proper rational functions. We begin with transfer functions.

13.1 Realizing a Transfer Function

Suppose we are given a proper transfer function

β(s)
t(s) =
α(s)

with α(s) monic, which we wish to realize. The first step is to is to separate it into its strictly proper part
and what’s left: Define
β(s)
d = lim
s→∞ α(s)

Note that d must be finite because of properness. Moreover if we define

γ(s) = β(s) − d

then
β(s) γ(s)
= +d
α(s) α(s)
and
γ(s)
α(s)
is strictly proper. There are many ways to realize strictly proper rational functions. We shall discuss two
methods.

201
13.1.1 Realization Using the Control Canonical Form

Our aim is to realize the strictly proper transfer function

γ(s)
α(s)

For this, let (A, b) be the control canonical form determined by

α(s) = sn + an sn−1 + · · · a2 s + a1

That is

   
0 1 0 ··· 0 0
 0 0 1 ··· 0  0
 . .. .. .. ..  .
A=
 .
. . . . . 
 b= .
.
 0 0 0 ··· 1  0
−a1 −a2 −a3 · · · −an n×n 1 n×1

Note that because of the form of this pair,


 
1
 s 
(sI − A) 
 ..  = α(s)b

.
s(n−1)

Therefore
 
1
1  
s 

(sI − A)−1 b = .. (13.2)
α(s)  . 
s(n−1)
which is a formula well worth remembering.

Suppose
γ(s) = gn s(n−1) + · · · + g2 s + g1
Define

c = [ g1 g2 · · · gn ]
Then because of (13.2),
γ(s)
c(sI − A)−1 b =
α(s)
which means that {A, b, c, d} is a realization of the original transfer function

β(s)
α(s)

13.1.2 Realization Using Partial Fractions

Another way to realize a transfer function, which is fundamentally different that the way just described, is
to expand the transfer function in partial fractions, then realize each term in the expansion and then put
them together iinhe appropriate way. The disadvantage of this method compared with what was just done
is clear: the latter requires one to find all the roots of α(s), whereas the former does not. Despite this, the
partial fraction approach provides several insights into the realization process and for this reason will be
sketched in the sequel.

Suppose for simplicity, that α(s) has distinct n roots λ1 , . . . , λn . That is,

α(s) = (s − λ1 )(s − λ2 ) · · · (s − λn )

where λi 6= λj if i 6= j. Then the partial fraction expansion of the above transfer function would be

β(s) r1 rn
=d+ + ···+
α(s) (s − λ1 ) (s − λn )

where d is as determined before and ri is the residue


 
∆ β(s)
ri = (s − λi )
α(s) s=λi

The next step is to realize each of the transfer functions in the sum.. A simple one-dimensional realization
of the ith such transfer function is
∆ ∆ ∆
Ai = λi bi = 1 ci = ri (13.3)

Of course
∆ ∆ ∆
Ai = λi bi = ri ci = 1
will work just as well. In fact there are infinitely many ways to construct such a one-dimensional realization.
We leave it to the reader to verify that no matter which of these realizations we use for each i, if we define
∆ ∆ ∆
A = diagonal {A1 , A2 , . . . , An }n×n b = column {b1 , b2 , . . . , bn } c = [ c1 c2 · · · cn ]

then {A, b, c, d} realizes the given transfer function. Note that for any such realzation
 
λ1 0 0 0
 0 λ2 0 0 
A=
 ... .. .. 
0 . . 
0 ··· 0 λn

Note also that some of the λi and ri may be complex valued, in which case the preceding will be a complex
realization. If one starts with a real transfer function, complex realizations can always be avoided by first
taking care of complex terms in the original expansion. Suppose for example that rj and λj are non-real
numbers. Then because the transfer function under consideration has real coefficients, there must also be an
integer k such that λk = λ∗j and rk = rj∗ where ∗ denotes complex conjugate. If this is so then it is always
possible to write
rj rk (f1 s + f2 )
+ = 2 (13.4)
(s − λj ) (s − λk ) (s − 2σs + σ 2 + ω 2 )

where the fi , σ and ω are real numbers and λj = σ + ω −1. It can easily be seen that the real transfer
function in (13.4) is realized {for example} by the linear system
   
σ −ω 0
A= b= c = [ f1 σf2ω+f1 ]
ω σ 1

We leave it to the reader to work out how to realize a sum consisting of real transfer function of this type
and real transfer function whose denominators are of degree one. The reader may also wish to consider the
general case when α(s) has repeated roots.
13.2 Realizing a p×1 Transfer Matrix Using the Control Canonical
Form

Both of the preceding methods easliy generalize to the realization of a column vector of transfer functions.
We shall only discuss the extension of the first method.

Suppose we are given a p × 1 column vector T (s) of proper transfer functions. Without loss of generality
we can assume that all of the transfer functions in T (s) have the same monic common denominator α(s):

β 
1 (s)
α(s)
 
 β (s) 
 2 
 α(s) 
 
T (s) =  
 .. 
 . 
 
 
βp (s)
α(s)

Proceding along the same lines as before in the case of a single transfer function, we can write

T (s) = H(s) + d

where d is the constant p − vector



d = lim T (s)
s→∞

and H(s) is the strictly proper transfer matrix defined by


H(s) = T (s) − d

Suppose H(s) of the form


 γ1 (s)

α(s)
 
 
 γ2 (s) 
 α(s) 
 
H(s) =  
 .. 
 . 
 
 
γp (s)
α(s)

Let A and b be the control canonical form determined by α(s). For i ∈ {1, 2, . . . , p}, define ci so that

γi (s)
ci (sI − A)−1 b =
α(s)

much as was done before for a scalar transfer function. It follows that if we define
 
c1
∆  c2 
C= 
 ... 
cp

then {A, b, C} will realize H(s) and consequently {A, b, C, d} will realize T (s). Note that the realization is
controllable.
13.3 Realizing a p × m Transfer Matrix

Now suppose that T (s) is a p × m matrix. In other words

T (s) = [ T1 (s) T2 (s) · · · Tm (s) ]

where each Ti (s) is a p × 1 matrix of proper transfer functions. Using {for example} the method of the
preceding section, we can find, for each i ∈ {1, 2, . . . , m}, a ni -dimensional realization {Ai , bi , Ci , di } of
Ti (s). The reader should carefully verify that

A = block diagonal {A1 , A2 , . . . , Am }

B = block diagonal {b1 , b2 , . . . , bm }

C = [ C1 C2 · · · Cm ]

D = block diagonal {d1 , d2 , . . . dm }

is then a realization of T (s). In fact this realization is even controllable. We have proved the following.

Theorem 34 A matrix of rational functions is realizable by a finite-dimensional, time-invariant, linear


dynamical system if and only if each of its elements is a proper rational function.

We now know how

1. to realize any transfer matrix.


2. to reduce any time-invariant linear system to a controllable time-invariant linear system with the same
transfer matrix.
3. to reduce any time-invariant linear system to an observable time-invariant linear system with the same
transfer matrix.

Moreover we know that controllability is not lost if we reduce a time-invarinat controllable linear system to
an observable time-invariant linear system with the same transfer matrix. Thus the preceding three steps,
performed in sequence, will provide a controllable and observable time-invariant realization of any matrix of
proper rational functions. It can be shown that the modified procedure obtained by interchanging steps 2
and 3, also produces a controllable and observable time-invariant realization of any matrix of proper rational
functions. We summarize.

Theorem 35 Each matrix of proper rational functions admits a time-invaraint controllable and observable
realization.

13.4 Minimal Realizations



Let us agree to say that Σ = {A, B, C, D} is a minimal realization of a given transfer matrix T (s) if Σ realizes
T (s) and if the dimension of Σ’s state space is as least as small as the dimension of the state space of any
other time-invariant linear system which realizes T (s). Note that Σ is minimal in the sense defined earlier
in Chapter 3 because two time-invariant linear systems have the same transfer matrix just in case they have
the same weighting pattern. In view of Theorem 31, we see that a minimal realization of T (s) is the same
as a controllable and observable realization of Σ. Thus the three steps enumerated at the end of the last
section provide a procedure for constructing a minimal realization of any given transfer matrix.
13.5 Isomorphic Systems

Let us agree to say that two m-input, n-dimensional, p-output, time-invariant systems Σ = {A, B, C, D} and

Σ̄ = {Ā, B̄, C̄, D} are similar systems or isomorphic systems if there exists a nonsingular matrix Q such that

QAQ−1 = Ā QB = B̄ CQ−1 = C̄

System similarity is an equivalence relation on the class of all m-input, n-dimensional, p-output, time-
invariant systems. The transfer matrix of a linear system is invariant under system similarity – but it is not
necessarily a complete invariant. In other words, while two similar linear systems necessarily have the same
transfer matrix, it is not true in general that two systems with the same transfer matrix are similar, even if
both have the same dimension. On the other hand, two controllable and observable linear systems with the
same transfer matrix are in fact similar. This is a statement of the so-called “Isomorphism Theorem.”

Theorem 36 (Isomorphism) Two controllable and observable linear systems with the same transfer ma-
trix are isomorphic.

The theorem implies that two controllable and observable linear systems with the same transfer matrix are
“essentially the same” in that they differ from each other only by a change of state variables. The proof of
Theorem 36 depends on the following lemma which is useful in its own right.

Lemma 11 Let {Ā, B̄, C̄} and {A, B, C} be the coefficient matrix triples of two linear systems with the
same transfer matrix. Suppose that the former system is n̄-dimensional and controllable and that the latter
is n-dimensional and observable. Then

C̄ = CT T Ā = AT B = T B̄ (13.5)

where T the n × n̄ matrix



T = RR̄−1 ,

R̄−1 being a right inverse of the the controllability matrix of (Ā, B̄), and R = [ B AB ··· An̄−1 B ].

∆ ∆
Proof of Theorem 36: Let Σ̄ = {Ā, B̄, C̄} and Σ = {A, B, C} be two controllable and observable systems
with the same transfet matrix. Then both systems are minimal with the same state space dimension n̄ = n.

Let T be the n × n matrix T = RR̄−1 , where R̄−1 is a right inverse of the the controllability matrix of (Ā, B̄),
and R is the controllability matrix of (A, B). By Lemma 11,

C̄ = CT T Ā = AT B = T B̄ (13.6)

It is therefore enough to prove that T is nonsingular.

From the latter two equations in (13.6) there follows Ai B = T Āi B̄, i ≥ 0. Therefore R = T R̄. But
rank R = n, so rank T ≥ n. Since T is n × n, rank T = n so T is nonsingular.

Proof of Lemma 11: Equality of the transfer matrices of {Ā, B̄, C̄} and {A, B, C} implies that the transfer
matrices’ Markov matrices are equal. Thus

C̄ Āi B̄ = CAi B, i≥0

From the relations there follow

N̄ R̄ = N R C̄ R̄ = CR N̄ ĀR̄ = N AR N̄ B̄ = N B
where N is the observability matrix of (C, A) and
 

∆  C̄ Ā 
N̄ = 
 .. 

.
C̄ Ān−1

These equations respectively imply that

N −1 N̄ = RR̄−1 C̄ = CRR̄−1 N −1 N̄ Ā = ARR̄−1 B = N −1 N̄ B̄



where N −1 is a left inverse of N . Since T = RR̄−1 , it follows from the preceding that the lemma is true.
Chapter 14

Stability

Intutively speaking, a system is stable if nothing “blows up.” The aim of stability theory is to define stability
in precise mathematical terms to capture this idea. This turns out to be a bit of a challenge as there are
many different ways to define stability which are all useful. For example, one could say that a linear system
is stable if each bounded input produces a bounded output response. Or one could say a linear system is
stable if the state of the system remains bounded, no matter what the initial state, assuming the input is
zero. The former is an example of an “external” or input-output view of stability while the latter is an
“internal” view of stability. Under certain conditions the two views are equivalent. The aim of this chapter
is to define various notions of stability. We begin with the idea of internal stability.

14.1 Uniform Stability

Let Σ denote a linear system with state equation

ẋ = A(t)x + B(t)u

where A(t) and B(t) are bounded and piecewise continuous matrices defined on the time interval [0, ∞). We
say that Σ or A(t) is uniformly stable if there is a positive constant β such that

||Φ(t, τ )|| ≤ β, ∀t ≥ τ, ∀τ ≥ 0 (14.1)

Here Φ(t, τ ) is the state transition matrix of A(t) and || · || is any norm or induced norm on IRn×n for which
the submultiplicative property holds1 . It can be shown quite easily that uniform stability means that any
solution x(t) to ẋ = A(t)x is bounded in the sense that ||x(t)|| ≤ β||x(t0 )|| for all t ≥ t0 and all t0 ≥ 0. The
modifier “uniform” is used to emphasize the requirement that for uniform stability β must be independent
of t0 .

Uniform stability turns out to be less than one might hope for in at least three ways. First, even
though uniform stability ensures boundedness of all solutions to ẋ = A(t)x, it does not ensure that all such
solutions will tend to zero as t → ∞. Second, there are uniformly stable systems whose state responses to
bounded inputs grow without bound. Third, there are uniformly stable A(t)’s which, after arbitrarily small
perturbations of their elements, are no longer uniformly stable. Each of these situations are exemplified by
the scalar state equation ẋ = u. This system is uniformly stable because its state-transition matrix, namely

Φ(t, τ ) = 1 satisfies (14.1) with β = 1. Note first that the only solution to ẋ = 0 which tends to zero
1 For ∆
concreteness, one can think in terms of the infinty norm ||Mp×q || = maxi,j |mij |.

209
as t → ∞, is the zero solution itself! Note second that the bounded input u(t) = 1, t ≥ 0 produces the
unbounded state response x(t) = x(0) + t. Third note that if A is perturbed to A = ǫ where ǫ is any positive
number, then Φ(t, τ ) = eǫ(t−τ ) ; thus no matter how small ǫ is, A will no longer be uniformly stable because
||Φ(t, τ )|| will grow without bound as t → ∞.

The preceding suggest that what’s needed is a stronger notion of stability which precludes all of these
undesirable properties. We introduce such a notion in the next section.

14.2 Exponential Stability

Again let Σ denote a linear system with state equation

ẋ = A(t)x + B(t)u

where A(t) and B(t) are bounded and piecewise continuous matrices defined on the time interval [0, ∞). Let
us agree to call Σ or A(t) exponentially stable if there is are positive constants β and λ such that

||Φ(t, τ )|| ≤ βe−λ(t−τ ) , ∀t ≥ τ, ∀τ ≥ 0 (14.2)

It can be shown quite easily that exponential stability means that any solution x(t) to ẋ = A(t)x tends
to zero exponentially fast in the sense that ||x(t)|| ≤ β||x(t0 )||e−λ(t−t0 ) for all t ≥ t0 and all t0 ≥ 0. The
modifier “uniform” is sometimes used before the word exponential to emphasize the requirement that β and
λ must be independent of t0 . It should be clear that exponential stability implies uniform stability, but not
the converse.

At we have said, an exponentially stable A(t) is a matrix for which all solutions to ẋ = A(t)x tend to
zero exponentially fast. But what if we drop the requirement that the convergence be exponential? More
specifically, suppose we say that Σ or A(t) is uniformly asymptotically stable if for any positive number ǫ
there exists a time Tǫ such that

||Φ(t, t0 )|| ≤ ǫ ∀t ≥ Tǫ + t0 , ∀t0 ≥ 0

In other words, uniform asymptotically stability means that for any tolerance ǫ we might pick and no matter
what the value of t0 , we can be sure that ||x(t)|| will be at least as small as ǫ||x(t0 )|| provided t ≥ t0 + Tǫ .
This of course implies that x(t) is converging to zero at t → ∞. The modifier uniform in this case means
that Tǫ must be independent of t0 .

Let us note right away that exponential stability of A(t) implies uniform asymptotic stability of A(t).
What’s perhaps a little supprising is that the converse is also true.

Theorem 37 A bounded, piecewise-continuous matrix A(t) defined on [0, ∞) is uniformly asymptotically


stable if and only if it is exponentially stable.

The theorem implies that if all solutions to ẋ = A(t)x converge to zero, uniformly in t0 , then the convergence
must be exponential.

Proof of Theorem 37: It is enought to prove that uniform asymptotic stability implies exponential stability.
For this let ǫ be any positive number less than 1, and let T be any positive time such that

||Φ(t, τ )|| ≤ ǫ ∀t ≥ T + τ, ∀τ ≥ 0

Let
∆ 1
λ=− log ǫ
T
Because ǫ < 1, it must be true that λ > 0. Moreover e−λT = ǫ. Therefore

||Φ(T + τ, τ )|| ≤ e−λT ∀τ ≥ 0 (14.3)

Next let

a = sup ||A(t)||
t∈[0,∞)

Since A(t) is bounded, a ≤ ∞. By definition,

Φ̇(t, τ ) = A(t)Φ(t, τ ), Φ(τ, τ ) = I, ∀t ≥ τ, ∀τ ≥ 0

Thus Z t
Φ(t, τ ) = I + A(s)Φ(s, τ )ds, ∀t ≥ τ, ∀τ ≥ 0
τ
Therefore for all t ≥ τ and τ ≥ 0
Z t
||Φ(t, τ )|| = ||I + A(s)Φ(s, τ )ds||
τ
Z t
≤ ||I|| + || A(s)Φ(s, τ )ds||
τ
Z t
≤ 1+ ||A(s)Φ(s, τ )||ds
τ

Hence Z t
||Φ(t, τ )|| ≤ 1 + a||Φ(s, τ )||ds, ∀t ≥ τ, ∀τ ≥ 0
τ
Thus by the Bellman-Gronwall Lemma,

||Φ(t, τ )|| ≤ ea(t−τ ) , ∀t ≥ τ, ∀τ ≥ 0

Clearly
||Φ(t, τ )|| ≤ e{(a+λ)T −λ(t−τ )} , ∀T ≥ t − τ, ∀t − τ ≥ 0, ∀τ ≥ 0 (14.4)

Now fix t0 > 0 and t ≥ t0 . Let n be a nonnegative integer and t1 a number in [0, T ) such that
t − t0 = nT + t1 . Then by the composition rule for state-transition matices

Φ(t, t0 ) = Φ(t1 + nT + t0 , nT + t0 )Φ(nT + t0 , (n − 1)T + t0 ) · · · Φ(T + t0 , t0 )

Since the norm we are using is submitliplicative,

||Φ(t, t0 )|| ≤ ||Φ(t1 + nT + t0 , nT + t0 )||||Φ(nT + t0 , (n − 1)T + t0 )|| · · · ||Φ(T + t0 , t0 )||

From this, (14.3) and (14.4) there follows

||Φ(t, t0 )|| ≤ e{(a+λ)T −λ(t1 +nT +t0 −nT −t0 )} e−λ(nT +t0 −(n−1)T −t0 ) · · · e−λ(T +t0 −t0 )

Hence
||Φ(t, t0 )|| ≤ e{(a+λ)T −λ(t1 +nT −t0 )} = e{(a+λ)T −λ(t−t0 )}

Setting β = e(a+λ)T we thus arrive at the inequality

||Φ(t, t0 )|| ≤ βe−λ(t−t0 )

Since this holds for all t ≥ t0 and all t0 ≥ 0, A(t) is exponentially stable.
Example 64 The general solution to the differential equation
1
ẋ = − x
1+t
is
1 + t0
x(t) = x(t0 )
1+t
Clearly any such solution tends to zero as t → ∞ but not exponentially fast. Note that given any positive
number ǫ < 1, it is not possible to find another number Tǫ , not dependening on t0 , for which

||x(t)|| ≤ ǫ||x(t0 )||, ∀t ≥ t0 + Tǫ , ∀t0 > 0

Said differently, the system is not uniformly asymptotically stable even though x tends to zero.

14.3 Time-Invariant Systems

For time-invariant systems, one drops the modifier uniform and refers to uniformly stable systems simply as
stable systems. Time-invariant linear systems which are not stable are said to be unstable linear systems. For
time invariant case it is possible to completely chacterize both stability and exponential stability in terms of
the eigenvalues of A and there multiplicities. Let us begin with a brief review.

14.3.1 The Structure of the Matrix Exponential

Assume that A is a given constant n × n matrix, that J is its Jordan2 Normal Form and that T is a
nonsingular matrix such that
A = T JT −1
Then as we’ve already explained in Chapter 8 of the math notes,

etA = T etJ T −1

Now since J is a block diagonal matrix of the form


 
J1 0 ··· 0
 0 J2 ··· 0 
J =
 ... .. .. .. 
. . . 
0 0 · · · Jm

we can write  
etJ1 0 ··· 0
 0 etJ2 ··· 0 
etJ =
 ... .. .. .. 
. . . 
0 0 · · · etJm
Thus to characterize the kinds of time functions which can appear in etA , what we really need to do it to
characterize the kinds of time functions which appear in matrix exponentials of the form etJi where Ji is
a typical Jordan Block. Whatever these time functions are, they are all going to appear in etA in linear
combination with similar time functions from J’s other Jordan Blocks. This is because for any matrix B,
each entry of the matrix T BT −1 is a linear combination of the entries of the matrix B.
2 Named in honor of this course’s distinguished spokesperson.
Now it is quite clear that if
Ji = [ λi ]1×1
then etJi is simply the scalar exponential
etJi = etλi
Consider next the case when  
λi 1
Ji =
0 λi 2×2

Then  
tJi etλi tetλi
e =
0 etλi
This can easily be verified by checking to see that with etJi so defined,
d tJi
e = Ji etJi , and etJi t=0
=I
dt
Similarly, for  
λi 1 0
Ji =  0 λi 1
0 0 λi
one finds that  
t2 tλi
etλi tetλi 2e
etJi = 0 etλi te tλi 
tλi
0 0 e
More generally, if

 
λi 1 0 ··· 0 0
0 λi 1 ··· 0 0
 .. 
 . 0 
0 0 λi 0
Ji =  . .. .. .. . .. 
 . . .. 
 . . . . 
0 0 0 · · · λi 1
0 0 0 ··· 0 λi mi ×mi
then  
t2 tλi t(mi −2) tλi t(mi −1) tλi
etλi tetλi 2e ··· (mi −2)! e (mi −1)! e
 
 t(mi −3) tλi t(mi −2) tλi

 0 e tλi
te tλi
··· 
 (mi −3)! e (mi −2)! e 
 
 t(mi −4) tλi (mi −3) 
 0 0 etλi ··· t tλi 
 (mi −4)! e (mi −3)! e 
etJi =



 .. .. .. .. .. .. 
 . . . . . . 
 
 
 
 0 0 0 ··· etλi tetλi 
 
0 0 0 ··· 0 etλi mi ×mi
mi
Thus we see that the Jordan Block of the elementary divisor (s−λi ) contains within it, linear combinations
{actually just scalar multiples} of the time functions

etλi , tetλi , . . . , t(mi −1) etλi

We’ve arrived at the following characterization of the state transition matrix of A.


Theorem 38 Let A be an n × n matrix with elementary divisors

(s − λ1 )m1 , (s − λ2 )m2 , . . . , (s − λq )mq

Then each of the entries of etA is a linear combination of the time functions

etλi , tetλi , . . . , t(mi −1) etλi , i ∈ {1, 2, . . . , q}

Thus the qualitative behavior of etA and consequently all solutions to ẋ = Ax, is completely characterized
by A’s eigenvalues and the sizes of their Jordan blocks.

Real Matrices

Suppose A is real, that λ is one of its eigenvalues and that m is the size of its Jordan block. Then as we’ve
just noted, linear combinations of the time functions

eλt , teλt , . . . , tm−1 eλt

must appear in etA . Suppose λ is not real; i.e.,

λ = a + jω

where a and ω are real numbers and ω 6= 0. Then

λ∗ = a − jω

must also be an eigenvalue of A withthe same size Jordan block at λ. Because A is real, it can easily be
shown that the linear combinations of the time functions
∗ ∗ ∗
eλt , teλt , . . . , tm−1 eλt , eλ t , teλ t , . . . , tm−1 eλ t

which appear in etA can also be written as linear combinations of time functions of the forms

eat cos ωt, eat sin ωt, teat cos ωt, teat sin ωt, . . . , tm−1 eat cos ωt, tm−1 eat sin ωt

14.3.2 Asymptotic Behavior

Let A be either real or complex valued. On the basis of Theorem 38 we can draw the following conclusions.

1. If all of A’s eigenvalues have negative real parts, then ||etA || → 0 exponentially fast, as t → ∞3 . Such
A are said to be asymptotic stability matrices.

2. If at least one eigenvalue of A

(a) has a positive real part a, then


||etA ||
must grow without bound, as fast as eat , as t → ∞.
3 Note that for any finite integer i ≥ 0, and any positive number µ, ti e−µt → 0 as t → ∞.
(b) has a zero real part and a Jordan block of size m > 1, then

||etA ||

must grow without bound as fast tm−1 , as t → ∞.

In either case, such A are said to be unstable matrices.

3. If all of A’s eigenvalues have nonpositive real parts, and those with zero real parts all have Jordan
blocks of size 1, then
||etA ||
will remain finite as t → ∞ but will not approach 0. Such A are said to be stable matrices.

The key observation here is summarized as follows:

Theorem 39 A n × n matrix A is asymptotically stable; i.e.,

lim etA = 0
t→∞

if and only if all of A’s eigenvalues have negative real parts.

14.3.3 Discrete Systems

Just about everything we’ve discussed in the preceding section for the linear differential equation

ẋ = Ax

extends an a natural way to a linear recursion equation of the form

x(t + 1) = Ax(t), t = 0, 1, 2, . . .

For example, for the recursion equation we can write

x(t) = A(t−τ ) x(τ ), ∀t, τ ≥ 0

For this reason A(t−τ ) is called the {discrete-time} state transition matrix of A. Note that if J is A’s Jordan
normal form and A = T JT −1, then
At = T J t T −1 ,
so the asymptotic behavior of At is characterized by A’s eigenvalues, just as in the continuous time case
discussed above. We leave it to the reader to verify the following:

1. If all of A’s eigenvalues have magnitude less than 1, then ||At || → 0 exponentially fast, as t → ∞4 .
Such A are said to be discrete-time asymptotic stability matrices.

2. If at least one eigenvalue of A

(a) has magnitude a greater than 1, then


||At ||
must grow without bound, as fast as |a|t , as t → ∞.
4 Note that for any finite integer i ≥ 0, and any real number µ with magnitude less than 1, ti µt → 0 as t → ∞.
(b) has magnitude 1 and multiplicity m > 1, then

||At ||

must grow without bound as fast tm−1 , as t → ∞.

In either case, such A are said to be discrete-time unstable matrices.

3. If all of A’s eigenvalues have magnitudes no greater than 1, and those with magnitude 1 are all have
Jordan blocks of size 1, then
||At ||
will remain finite as t → ∞ but will not approach 0. Such A are said to be discrete-time stable matrices.

The discrete-time version of Theorem 39 is as follows:

Theorem 40 An n × n matrix A is discrete-time asymptotically stable; i.e.,

lim At = 0
t→∞

if and only if all of A’s eigenvalues have magnitudes less than 1.

14.4 Lyapunov Stability

The preceding discussion of stability and asymptotic stability for time-varying systems leaves unanswered
the question of how one might go about determining whether or not a given A(t) is uniformly stable or
asymptotically stable. In fact, for general time-varying linear systems, there is no constructive means for
doing this. Nonetheless, there are various ways to characterize uniform and asymptotic stability which prove
useful in certain applications. One of the most useful ideas developed so far is do to the late 19th century
engineer A. M. Lyapunov. Lyapunov’s idea applies not only to linear differential equations, but also to more
general nonlinear differential equations of the form

ẋ = f (x, t) (14.5)

where f satisfies f (0, t) = 0 and is at least locally Lipschitz in x and piecewise-continuous and bounded

t for every fixed x. Note that for the A(t)’s we’ve been considering, f (x, t) = A(t)x satisfies all of these
requirements.

Lyapunov’s idea is roughly as follows. As a first step one tries to find a continuously differentiable5 ,
non-negative, scalar-valued, function V (x, t) which is bounded in t on [0, ∞) for each fixed x and which
is radially unbounded6 in x for each fixed t. One further requires V to have the property that that along

each solution φ to (14.5), the total derivative the time function Vt = V (φ(t), t) is non-positive. For if this is
so, then Vt is monotone non-increasing, and therefore never larger than its initial value V (x0 , t0 ). Because
of this, and V ’s radial unboundedness, it would then follow that φ(t) is bounded where ever it exists and
consequently that φ(t) exists and is bounded on [t0 , ∞).

To recap, if one can find a continuously differentiable, non-negative, scalar valued, function V (x, t), which
is radially unbounded in x for each fixed t and bounded in t for each fixed x, and whose time derivative
5A function U (x, t) is continuously differentiable if its first derivatives with respect to x and t are all continuous in x and t.
6A nonnegative, scalar-valued function U (x) is radially unbounded if for each positive number r there is a positive number
r̄ such that U (x) > r whenever ||x|| > r̄. The definition implies that ||x|| is finite if U (x) is.
along any solution to (14.5) is non-positive, then each maximal solution φ to (14.5) must be bounded, and
therefore defined and bounded on [t0 , ∞). V ’s with these properties are typically called Lyapunov functions7 .

What makes Lyapunov’s idea so appealing, is that it is often possible to show that V̇t ≤ 0, without having
to explicitly compute a solution φ to (14.5). In particular, suppose V satisfies the condition

∂V (x, t) ∂V (x, t)
+ f (x, t) ≤ 0, ∀x ∈ IRn , ∀t ≥ 0 (14.6)
∂t ∂x
Then along any solution φ
   
∂V (x, t) ∂V (x, t) ∂V (x, t) ∂V (x)
V̇t = + φ̇ = + f (x, t) ≤0
∂t ∂x x=φ(t) ∂t ∂x x=φ(t)

In other words, to show that V̇t ≤ 0, its enough to check that (14.6) is true - and this can be done without
computing φ.

To exploit Lyapunov’s idea, it is thus necessary to construct a “candidate Lyapunov function” V in such
a way that that (14.6) holds. Unfortunately there is in general no systematic method for doing this. Indeed,
except for special cases such as f (x, t) = Ax with A constant, the construction of Lyapunov functions is very
much of an art.

14.4.1 Linear Time-Varing Differential Equations

For linear time-varing differential equations of the form

ẋ = A(t)x (14.7)

with A(t) piecewise-continuous and bounded on [0, ∞), one typically considers candidate Lyapunov functions
of the form
V (x, t) = x′ P (t)x (14.8)
where P (t) is a symmetric, positive definite, continuously differentiable, bounded matrix8 on [0, ∞). To
avoid difficulties in the limit at t → ∞ one typically demands that P (t) remain “positive definite in the
limit.” This is accomplished by stipulating that there should be a positive number µ such that

P (t) ≥ µI, ∀t ≥ 0 (14.9)

What this requirement means is that P (t)’s smallest eigenvalue, λmin (t), can never get be smaller than µ for
any value of t.

Suppose that such a P (t) can be found and that P (t) miraculously satisfies a matrix differential equation
of the form
Ṗ (t) + A′ (t)P + P (t)A(t) + Q(t) = 0, t ∈ [0, ∞) (14.10)
where Q(t) is a symmetric, bounded, piecewise-continous, positive semidefinite matrix on [0, ∞). We claim
that under these conditions, A(t) is uniformly stable. To understand why this is so, let us note that along a
solution x(t) to (14.7), V (x, t must satisfy

∆ d
V̇ (t) = V (x(t), t) = x′ (t)Ṗ (t)x(t) + x′ (t)A′ (t)P (t)x(t) + x′ (t)P (t)A(t)x(t)
dt
= x′ (t)(Ṗ (t) + A′ (t)P (t) + P (t)A(t))x(t)
7 The standard definition of a Lyapunov function also requires V to satisfy V (0, t) = 0, ∀t ≥ 0.
8 This would be a good time to review the sections in the Linear Algebra Notes which deal with norms, simultaneous
diagonalization, and quadratic forms.
Hence from (14.10)
V̇ = −x′ (t)Q(t)x(t) (14.11)
But Q(t) is positive semidefinite which means that V̇ ≤ 0, so V (x(t), t) is a nonincreasing function of t.
Thus V is bounded above by its initial value V (x(t0 ), t0 ). That is

V (x(t), t) ≤ V (x(t0 ), t0 ), ∀t ≥ t0

In view of (14.9),
1
||x(t)||2 ≤ V (x(t0 ), t0 ), ∀t ≥ t0
µ
∆ √
where || · || is the Euclidean norm ||z|| = z ′ z. To establish uniform stability we need finally to use the
assumption that P (t) is bounded on [0, ∞) Because P ’s eigenvalues depend continuously on its entries and
because P ’s entries are all bounded on [0, ∞), we can conclude that P ’s largest eigenvalue λmax (t) is also
bounded on [0, ∞). Let ν denote this bound. In other words λmax (t) ≤ ν, ∀t ∈ [0, ∞). This implies that

P (t) ≤ νI, ∀t ∈ [0, ∞) (14.12)


ν
and thus that V (x(t0 ), t0 ) ≤ ν||x(t0 )||2 , ∀t ≥ t0 . Therefore ||x(t)||2 ≤ µ ||x(t0 )||
2
or
r
ν
||x(t)|| ≤ ||x(t0 )||
µ

Since this holds for all t0 and neither µ nor ν depend on t0 A(t) must be uniformly stable as claimed.

Theorem 41 (Uniform Stability) Let A(t) be bounded and piecewise continuous on [0, ∞). Suppose there
is a positive number µ and bounded, symmetric matrics P (t) and Q(t) defined on [0, ∞) with P (t) continu-
ously differentiable and positive definite and Q(t) piecewise-continuous and positive semidefinite, such that
P (t) ≥ µI, ∀t ≥ 0 and
Ṗ (t) + A′ (t)P + P (t)A(t) + Q(t) = 0, t ∈ [0, ∞) (14.13)
Then A(t) is uniformly stable.

By making various stronger assumptions about Q(t), it is possible to say even more. We discuss two different
assumptions of this type.

Q(t) is Positive Definite on [0, ∞)

Suppose we assume in addition to what’s required in Theorem 41, that there is a positive number ρ such
that
Q(t) ≥ ρI, ∀t ≥ 0 (14.14)
Then using this followed by (14.9) we can write
ρ
x′ (t)Q(t)x(t) ≥ ρ||x(t)||2 ≥ V (x(t), t), ∀t ≥ 0
ν
From this and (14.11) there follows
d ρ
V (x(t), t) ≤ − V (x(t), , ∀t ≥ 0
dt ν
Thus  
d n ρt o ρ
t d ρ
e V (x(t), t) = e
ν ν V (x(t), t) + V (x(t), t) ≤ 0
dt dt ν
Clearly
ρ ρ
e ν t V (x(t), t)) ≤ e ν t0 V (x(t0 ), t0 ), ∀t ≥ t0
so
ρ
V (x(t), t)) ≤ e− ν (t−t0 ) V (x(t0 ), t0 ), ∀t ≥ t0
Using (14.9) and (14.12) we thus arrive at the inequality
µ − ρ (t−t0 )
||x(t)||2 ≤ e ν ||x(t0 )||2 , ∀t ≥ t0
ν
or r
µ − ρ (t−t0 )
||x(t)|| ≤ e 2ν ||x(t0 )||, ∀t ≥ t0
ν
Since µ, ν and ρ don’t depend on t0 , we can conclude that A(t) is exponentially stable. We summarize.

Corollary 10 (Exponential Stability) Let the hypotheses of Theorem 41 hold and suppose that there is
a positive number ρ such that Q(t) ≥ ρI, ∀t ≥ 0. Then A(t) is exponentially stable.

Q̇(t) Exists and is Bounded on [0, ∞)

Without assuming anything more about Q(t) than we did in the statement of Theorem 41 we can deduce
more by exploiting the fact from calculus that every bounded monotone nonincreasing function converges
in the limit. We have such a function, namely V (x(t), t). We’ve already established that V is monotone
nonincreasing and bounded above. Moreover, we know that V is bounded below by zero because P (t) is a
positive definite matrix. Thus we can conclude that the limit

V∞ = lim V (x(t), t))
t→∞

exists and is finite. This means that


Z t
lim V̇ (x(s), s)ds = V∞ − V (x(t0 ), t0 )
t→∞ t0

Hence from (14.11) we see that


Z t
lim x′ (s)Q(s)x(s)ds = V (x(t0 ), t0 ) − V∞ < ∞ (14.15)
t→∞ t0

We claim that if Q̇(t) exists and is bounded on [t0 , ∞) then (14.15) implies that

lim x′ (t)Q(t)x(t) = 0 (14.16)


t→∞

This equation can be used to deduce, to some extent, what happens to x(t) in the limit as t → ∞. For
example, if Q(t) = Q is a constant matrix, then the preceding implies that x → kernel Q. On the other hand
if x′ Q(t)x is the norm square of a signal y of the form y = C(t)x {i.e., Q(t) = C ′ (t)C(t)}, then this would
mean that for every initial state, the output of the unforced linear system

y = C(t)x ẋ = A(t)x

tends to zero as t → ∞.

Justification for (14.16) rests on the following lemma which is useful in its own right.
Lemma 12 ( Convergence ) Let µ : [0, ∞) → IR be a continuous function with a bounded, piecewise-
continuous derivative µ̇. If the limit
Z t
lim µ(τ )dτ (14.17)
t→∞ 0

exists and is bounded, then


lim µ(t) = 0 (14.18)
t→∞


Proof of Convergence Lemma: It is enough to show that for any positive number ǫ, the set Sǫ = {t :
|µ(t)| ≥ ǫ} is either empty or bounded from above. For if this is so, then for any number ǫ > 0, |µ(t)| must
be smaller than ǫ for any value of t greater than Sǫ ’s upper bound; and this in turn must imply that (14.18)
is true.

To proceed, pick ǫ > 0 and suppose that Sǫ is not bounded from above. Let b be any positive number
bounding µ̇; that is |µ̇(t)| < b, ∀t ≥ 0. The Cauchy convergence test together with the hypothesis that
the integral in (14.17) converges to a finite limit imply that a time t0 > 0 sufficiently large can be found for
which Z t
ǫ2
µ(s)ds < , ∀t > t0 (14.19)
t0 2b
Moreover, since Sǫ is supposed to be unbounded, t0 can be chosen large enough to be in Sǫ as well. Thus
|µ(t0 )| ≥ ǫ, so for t ≥ t0
 Z t 
sign{µ(t0 )}µ(t) = sign{µ(t0 )} µ(t0 ) + µ̇(s)ds
t0
Z t
≥ |µ(t0 )| − |µ̇(s)|ds
t0
Z t
≥ ǫ− bds
t0
= ǫ − b(t − t0 )

Therefore Z Z
t0 + ǫb t0 + ǫb
ǫ2
sign{µ(t0 )}µ(s)ds ≥ (ǫ − b(s − t0 ))ds =
t0 t0 2b
Clearly
Z t0 + ǫb
ǫ2
µ(s)ds ≥
t0 2b

which contradicts (14.19).

14.4.2 Time-Invariant Linear Differential Equations

A major problem with the results derived in the last section for time-varying linear systems is that they
don’t really provide constructive tests for deciding when a system is uniformly stable or exponentially stable.
This is because one has no effective means available for actually computing matrices P (t) and and Q(t) with
the required properties. Fortunately for time-invariant systems the situationis quite different. In the sequel
we will prove the following theorem which provides a constructive test for deciding whether or not a constant
matrix A is asymptotically stable.
Theorem 42 (Lyapunov Stability) Let A and Q be any real, constant n × n matrices.

• If A is an asymptotic stablity matrix, then the linear matrix equation

A′ P + P A + Q = 0 (14.20)

has exactly one solution P . Moreover, if Q is {symmetric} positive definite then so is P .


• If P and Q are positive definite matrices satisfying (14.20), then A is an asymptotic stability matrix.

The theorem provides an effective test for deciding whether or not a given matrix A is an asymptotic stability
matrix, without having to compute the matrix’s eigenvalues. The test consists of first picking any convenient
positive definite matrix for Q (say Q = I). Next one tries to solve (14.20), which is a linar algebraic equation.
A is an asymptotic stability matrix just in case (i) a solution P exists and (ii) P is positive definite. There
are a variety of ways to test a symmetric matrix for positive definiteness. See the Linear Algebra Notes for
this.

Proof of Theorem 42: Let us note right away that if Q and P are positive definite matrices satisfying
(14.20), then these matrices automatically satisfy the hypotheses of Corollary 10. From this it follows that
the statement following the second bullet of Theorem 42 is true.

To porve the statements following the first bullet, let A and Q be arbitrary real n × n matrices and define
Z t
∆ ′
X(t) = eµA QeµA dµ (14.21)
0


By changing the variable of integration from µ to τ = t − µ and then differentiating with respect to t we
obtain the linear matrix differential equation

Ẋ = A′ X + XA + Q (14.22)

Thus the matrix on the right side of (14.21) is the unique solution to (14.21) which starts in state X(0) = 0
at t = 0.

Now suppose that A is asymptotically stable. Then eµA QeµA is a matrix whose entries are all linear
combinations of decaying exponentials so the integral on the right side of (14.21) must converge to a finite
limit X∞ as t → ∞. This is the same thing as saying that the solution to (14.22) must converge to X∞ . In
the limit, Ẋ vanishes and we have
0 = A′ X∞ + X∞ A + Q
This means that for each real n × n matrix Q, there exists a solution to the equation

L(Y ) = −Q (14.23)

where L is the linear transformation L : Rn×n → IRn×n defined by



L(Y ) = A′ Y + Y A

Evidently L is an epimorphism {i.e., a surjective linear transformation}. In other words, image L = IRn×n .
Moreover L’s domain is IRn×n so

dim image L + dim kernel L = dim IRn×n

just like any linear function. It follows that dim kernel L = 0 and thus that kernel L = 0. This means that
there is only one Y which solves (14.23). In summary, if A is ayymptotically stable, then for each matrix Q,
the equation 0 = A′ Y + Y A + Q has exactly one solution Y and Y = X∞ .
Now suppose Q is positive definite. To complete the theorem’s proof, it is enough to show that X∞ {i.e.
P } is positive definite as well. But this is clear since for x 6= 0,
Z ∞  Z ∞ p

x′ X∞ x = x′ eµA QeµA dµ x = || QeµA x||2 dµ > 0
0 0

Therefore both statements following the first bullet of Theorem 42 are true.

14.5 Perturbed Systems

In this section we will develop a result which enables us to draw conclusions about the asymptotic stability
of matrices of the form A(t) + B(t) where A(t) is an asymptotic stability matrix and B(t) is a matrix whose
norm is “small” in a suitably defined sense. The result plays a central role in the analysis of slowly varying
time-varying linear systems and adaptive control systems. We begin with the idea of a nondestabilizing
signal.

14.5.1 Nondestabilizing Signals

Let us agree to say that a bounded, piecewise-continuous matrix B : [0, ∞) → IRn×m is nondestabilizing with
growth rate ρB ≥ 0 if (i) for each real number ρ > ρB there is a finite nonnegative constant bρ such that
Z t
||B(s)||ds ≤ bρ + ρ(t − τ ) t≥τ ≥0
τ

and (ii) ρB is the least nonnegative number with this property. Note that any piecewise-continuous bounded
matrix B : [0, ∞) is nondestabilizing with growth rate
ρB ≤ sup ||B(t)||
t∈[0,∞)

Examples of nondestabilizing signals with zero growth rates include all bounded, piecewise-continuous func-
tion B : [0, ∞) → IRn×m which tend to zero or which satisfy
Z ∞
||B(µ)||i dµ < ∞ (14.24)
0

for either i = 1 or i = 2. We will encounter other examples in the sequel.

The term “nondestabilizing” is prompted by the following basic result.

Theorem 43 (Nondestabilization) Let A : [0, ∞) → IRn×n and B : [0, ∞) → IRn×n be bounded,


piecewise-continuous matrices. Suppose that A is exponentially stable and let a and λ be positive constants
for which A’s state transition matrix φA (t, τ ) satisfies
||φA (t, τ )|| ≤ ae−λ(t−τ ) , ∀t ≥ τ ≥ 0
Suppose B is nondestabilizing with growth rate ρB and for each number ρ > ρB , let bρ be a constant for
which Z t
||B(s)||ds ≤ bρ + ρ(t − τ ), t ≥ τ ≥ 0 (14.25)
τ
Then for ρ > ρB , the state transition matrix of A + B satisfies
||ΦA+B (t, τ )|| ≤ ae(abρ −(λ−aρ)(t−τ )) , t ≥ τ ≥ 0
As an immediate consequence of this theorem we see that if B’s growth rate is sufficiently small, namely
λ
ρB <
a
then A + B will be exponentially stable. This of course will always be true if B’s growth rate is zero.

Proof of Theorem 43: Fix τ and define θ(t) = ΦA+B (t, τ ). Then

θ̇ = Aθ + Bθ
θ(τ ) = I

By viewing Bθ as a forcing function in the preceding, one may write the variation of constants formula
Z t
θ(t) = ΦA (t, τ ) + ΦA (t, s)B(s)θ(s)ds
τ

Therefore
Z t
||θ(t)|| = ||ΦA (t, τ ) + φA (t, s)B(s)θ(s)ds||
τ
Z t
≤ ||ΦA (t, τ )|| + ||ΦA (t, s)B(s)θ(s)||ds
τ
Z t
≤ ae−λ(t−τ ) + ae−λ(t−τ ) ||B(s)||||θ(s)||ds
τ

Hence Z t
λ(t−τ )
e ||θ(t)|| ≤ a + a||B(s)||{eλ(s−τ ) ||θ(s)||}ds
τ
By the Bellman-Gronwall Lemma,
Rt
eλ(t−τ ) ||θ(t)|| ≤ ae τ
a||B(s)||ds

Therefore Rt
||θ(t)|| ≤ ae(−λ(t−τ )+ τ
a||B(s)||ds)

Using (14.25) there follows


||θ(t)|| ≤ ae(abµ −(λ−aµ)(t−τ ))
which is the desired result.
Chapter 15

Feedback Control

Consider the n-dimensional linear system

ẋ = Ax + Bu (15.1)
y = Cx + Du (15.2)
s = C̄x (15.3)

which is now regarded as an internal model P of a physical process to be controlled. It is assumed that
A, B, C, C̄ and D are known approximately and that they can’t be altered. P is often called the process
model, u ∈ IRm the open-loop control input, y ∈ IRp the controlled output, and s ∈ IRq the sensor output1 .
It is not unusal for y and s to be one and the same, or more generaly for y to be readable for s in the sense
that for some matrix R, y = Rs.

Consider next the ‘closed-loop’ system exhibited diagramatically as follows.

r C uC u P y

Figure 15.1: Feedback Control System

Here r ∈ IRnr is a reference input vector which, depending on the particular situation, may or may not be
present. The function of C is to use sensor output s and r to generate a control signal uC which, when u is
set equal to uC , causes y to behave in some prescribed manner.

In these notes we consider controllers which can be modelled by linear, time-invariant dynamical systems.
Thus a linear controller C admits the description

ẋC = AC xC + BC s + EC r (15.4)
uC = CC xC + DC s + HC r (15.5)
1 All of the properties of feedback systems discussed in the sequel extend to processes with sensor outputs of the form

s = C̄x + D̄u. The less general output relation is adopted throughout for simplicity.

225
where AC , BC , CC , DC , EC , HC are constant matrices, xC ∈ IRnC is the state of C. Closed-loop control is
achieved by setting
u = uC (15.6)
so uC ∈ IRm .

Let us recall that time function differentiation is invariably avoided when implementing a controller
because of its adverse effect on noise. The preceding controller involves no differentiation and may be
constructed from standard analog or digital components. Relative to this, the preceding describes the most
general linear system for controlling a dynamical process.

15.1 State-Feedback

By a state-feedback control is meant a feedback law of the form

u = F x + Gr (15.7)
where F and G are constant matrices. If this law is applied to P the resulting closed-loop system PF,G is
described by the equations

ẋ = (A + BF )x + BGr (15.8)
y = (C + DF )x + DGr (15.9)
z = C̄x (15.10)

Algebraically the effect of state feedback is thus to transform a linear system P into another PF,G of the
same dimension and form. It is mainly because of the simplicity of this transformation that the study of
feedback control in state space terms has proved so useful. Although x usually can’t be measured directly,
thereby precluding implementation of a state-feedback law as a controller C such as described above, in due
course we shall justify our consideration of control laws of this type.

Observe that if G is nonsingular, then Im BG = B and < A + BF | B > is the controllable space of PF,G .
This being the case, the controllable space of PF,G turns out to be the same as the controllable space of P.

Proposition 12 For all F ,


< A + BF | B >=< A | B >

Proof: The equation


i
X i
X
(A + BF )j−1 B = Aj−1 B (15.11)
j=1 j=1

is valid for i = 1. Suppose (15.11) holds for i = k; then


k+1
X k
X
j−1
(A + BF ) B = (A + BF ) (A + BF )j−1 B + B
j=1 j=1
k
X
= (A + BF ) Aj−1 B + B (by hypothesis)
j=1
k
X
= A Aj−1 B + B
j=1
k+1
X
= Aj−1 B
j=1

By induction, (15.11) holds for all i ≥ 0 and in particular for i = n.

An immediate consequence of Proposition 12 is that (A + BF, B) is controllable if and only if (A, B)


is. A corresponding statement concerning the effect of state feedback on observability cannot be made; i.e.,
even if (C, A) is observable, (C + DF, A + BF ) may not be.

15.2 Spectrum Assignment

Consider the system PF,I to which the state feedback law u = F x + r has been applied. As already noted,
the spectrum of A + BF has a significant effect on the transient behavior of the system; e.g., the system is
asymptotically stable if and only if all eigenvalues of A + BF lie within the open left-half of the complex
plane. It is therefore important to understand the extent to which the spectrum of A + BF can be adjusted
with F .

Recall that the spectrum of A + BF is a symmetric set of complex numbers since the characteristic
polynomial of A + BF has real coefficients. In the sequel it will be shown that the symmetric spectrum of
A + BF can be freely assigned with suitable F just in case (A, B) is a controllable pair. This property of
controllable pairs depends upon the following preliminary results.

Proposition 13 Let (A, b) be a single-input, n dimensional, real controllable pair. For each symmetric set
of n complex numbers Λ there exists a real matrix F
spectrum (A + bf ) = Λ (15.12)

Proof: From the symmetric set Λ = {λi : i ∈ n} construct the polynomial


n
Y
ᾱ(s) = (s − λi ) = sn + ān sn−1 + ān−1 sn−2 + . . . + ā2 s + ā1
i=1

Let
α(s) = sn + an sn−1 + an−1 sn−2 + . . . + a2 s + a1
denote the characteristic polynomial of A and write T for the matrix which transforms (A, b) into control
canonical form. That is T AT −1 = AC and T b = bC where

   
0 1 0 ··· 0 0
 0 0 1 ··· 0  0
 . .. .. .. ..  .
AC = 
 .
. . . . . 
 bC =  .
.
 0 0 0 ··· 1  0
−a1 −a2 −a3 · · · −an n×n 1 n×1
Note that if we define the row vector

fC = [ a1 − ā1 a2 − ā2 · · · an − ān ]
then  
0 1 0 ··· 0
 0 0 1 ··· 0 
 . .. .. .. .. 
AC + bC fC = 
 .
. . . . . 

 0 0 0 ··· 1 
−ā1 −ā2 −ā3 · · · −ān
Thus ᾱ(s) is the characteristic polynomial of AC + bC fC which means that Λ is (AC + bC fC ) ’s spectrum.
Define

f = fC T
and note that T (A + bf )T −1 = AC + bC fC . Since A + bf and AC + bC fC are similar, both matrices must
have the same spectrum. Hence (15.12) is true.

Consider next the multi-input controllable pair (A, B). Spectrum assignment in this case reduces imme-
diately to the single-input case whenever there exists a vector b ∈ image B such that (A, b) is a controllable
pair. For when this is so, Proposition 13 can be used to construct a matrix f which assigns to A + bf a
prespecified symmetric spectrum. Since b ∈ image B, the equation b = Bg has a solution g; thus with
F = gf , there follows A + BF = A + bf so (A + BF )’s spectrum is as prescribed.

Unfortunately a vector b ∈ image B with the required property may not exist. For example, if A = 0
and B is the n × n identity then (A, B) is a controllable pair but there’s no b ∈ image I = for which (0, b) is
controllable. Indeed from our prior discussion, we know that such a vector b cannot exist if A is not cyclic.

To treat the general multi-input case, a slightly different approach will be taken. It will be shown that
if (A, B) is controllable, then there exists a vector b ∈ image B and a matrix F such that (A + BF, b) is
controllable. Thanks to M. Heymann, even more is true:

Lemma 13 (Heymann) Let (A, B) be an n-dimensional, m-input controllable pair and let b be any given
nonzero vector in image B . There exists a matrix F such that (A + BF, b) is a controllable pair.


Proof: Define B = image B. Since b 6= 0 there is a largest integer n1 , such that {b, Ab, . . . , An1 −1 b} is an
independent set. If n1 = n, set F = 0 since in this case (A, b) is controllable. If n1 < n, there exists a

vector b2 ∈ B such that b2 ∈ / R1 where R1 = span {b, Ab, . . . , An1 −1 b}. For if this were not so B would be
a subspace of R1 . Since AR1 ⊂ R1 , it would then follow that < A|B >= R1 (6= IRn ) which contradicts the
controllability hypothesis. Thus a vector b2 with the required properties exists. Let n2 be the largest integer
such that {b, Ab, . . . , An1 −1 b, b2 , Ab2 , . . . , An2 −1 b2 } is an independent set. Since (A, B) is controllable, this
process can be continued until for some integer k ≤ dim B,

{b, Ab, . . . , An1 −1 b, b2 , Ab2 , . . . , An2 −1 b2 , . . . , bk , . . . , Ank −1 bk }

is a basis for IRn .

Define
x1 = b (15.13)
and
xi = Axi−1 + b̄i−1 i = 2, 3, . . . , n (15.14)
where
b̄(n1 +n2 +...+nj )+1 = bj+1 j = 1, 2, . . . , k − 1
and all other b̄i = 0. It is easy to verify that {x1 , x2 , . . . , xn } is a basis for IRn . Compute vectors ui so that

Bui = b̄i i = 1, 2, . . . , n (15.15)

Finally, define
∆ −1
F = [ u1 u2 · · · un ] [ x1 x2 · · · xn ]
This means that
F xi = ui i ∈ {1, 2, . . . , n} (15.16)
It follows from (15.13)-(15.16) that

xi = Axi−1 + b̄i−1
= Axi−1 + Bui−1
= (A + BF )xi−1 i = 2, 3, . . . , n

Using (15.13), xi can therefore be written as

xi = (A + BF )i−1 b, i ∈ {1, 2, . . . , }

since the xi span IRn , the pair (A + BF, b) is controllable.

The main theorem on spectrum assignment is as follows.

Theorem 44 An n-dimensional matrix pair (A, B) is controllable if and only if for each symmetric set of
n complex numbers Λ, there exists a matrix F such that

spectrum (A + BF ) = Λ

We will only prove that controllability implies spectrum assignability. We leave the proof of the converse
implication to the reader as an exercise.

Proof: If (A, B) is controllable, there is {by Lemma 13} for any fixed nonzero vector b ∈ B, a matrix Fb
such that (A+BFb , b) is controllable. Proposition 15.12 thus provides a row vector f such that spectrum (A+

BFb + bf ) = Λ. Since b ∈ image B, the equation b = Bg has a solution g. With F = Fb + gf , there follows
A + BF = A + BFb + bf so F has the required property.

15.3 Observers

In the last section a procedure was developed for adjusting the closed-loop spectrum of a linear system with
a state-feedback law of the form
u = F x + Gv (15.17)
In practice direct measurement of system state is often not possible and so controllers such as the preceding
cannot typically be implemented. There is however an alternative controller which can be implemented
and which achieves similar results. The architecture of the control consists of a linear sub-system called an
“observer” whose role is to “asymptotically estimate” F x from measurement y; this estimate is then used
in place of F x in (15.17) to define the actual control signal u to be fedback to the process. In the sequel we
define and characterize observers.

15.3.1 Definition

Consider the unforced n-dimensional linear system

y = Cx, ẋ = Ax (15.18)

together with an additional output signal of the form

w = Lx (15.19)
where L is any given matrix with n columns. Our aim is to define a linear system of the general form

ż = Hz + Ky (15.20)
w
b = Mz + Ny (15.21)

in such a way so that along any trajectory of the combined linear system defined by (15.18)- (15.21), w b
converges to w in the limit as t → ∞. We call w
b an asymptotic estimate of w. Since (15.21) must hold even
when the estimate is exact, {i.e., when w
b = w}, we require the algebraic equations

Lx = M zx + N y y = Cx

to have a solution zx for each possible x ∈ IRn . In other words, for each such x we require there to be a
vector zx such that
(L − N C)x = M zx
Since for any given x the preceding is a linear algebraic equation, such a zx will exist if and only if

(L − M C)x ∈ image M

To ensure a solution no matter what the value of x we must therefore require that

image (L − N C) ⊂ image M

But this is exactly the condition for there to exist a solution V to the linear matrix equation

L − NC = MV (15.22)

Suppose that V is such a matrix and define the error signal

e=z−Vx (15.23)

where z and x are now the state vectors of (15.18) and (15.20) respectively. Then using (15.22), (15.18), and
b − w can be written as
(15.21) we see that the estimation error w

b−w
w = Mz + Ny − w
= M z + (N C − L)x
= Mz − MV x
= Me

b to asymptotically estimate w it is therfore enough to make sure that e converges to zero as t → ∞.


For w

Let us stipulate that the convergence of e is to be exponential by requiring e to satisfy a differential


equation of the form
ė = H̄e (15.24)
where H̄ a suitably defined constant, exponentially stable matrix. In order to determine what’s required for
(15.24) to hold, let us use the definition of e in (15.23) together with (15.22) and the (15.19)-(15.21) to write

ė = ż − V ẋ
= Hz + Ky − V Ax
= Hz + (KC − V A)x
= He + (HV + KC − V A)x

Therefore if we set H = H̄ and require HV + KC − V A = 0 to hold, then we will have satisfied the
requirement that (15.24) hold.
In the light of the preceding we now call the linear system {H, K, M, N } defined by (15.21) an L-observer
if for some matrix V the observer design equations

L = MV + NC (15.25)
V A = HV + KC (15.26)

both hold. Thus if {H, K, M, N } is an L observer, then the estimation error w


b = w will satisfy

b − w(t) = M eHt (z(0) − V x(0))


w(t)

for all initial states x(0) and z(0). There remains the problem of determining matrices H, K, M, N and V
for given C, A and L, in such a way that the observer design equations hold.

15.3.2 Full-State Observers

Just about the easiest solution to the observer design problem that one can think of, is the one for which
∆ ∆
M = L, N = 0, V = In×n and H = A − KC. Any observer of this type is called a full-state observer because
in this case z is an asymptotic estimate of x; i.e.

z(t) − x(t) = e(A−KC)t (z(0) − x(0))

Of course we we need to make sure that K is chosen to that A − KC is an asymptotically stable matrix.
One way to acomplish this, asuming (C, A) is an observable pair, is to exploit duality and use spectrum
assignment. In particular, we know that the spectrum of A-KC is the same as the spectrum of A′ −C ′ K ′ . We
also know that observability of (C, A) is equivalent to controllability of (A′ , C ′ ). Thus with an observability
assumption, the problem of choosing K to asymptotically stabilize A − KC is the same as the ptroblem of
choosing K ′ to asymptotically stabilize the A′ − C ′ K ′ with (A′ , C ′ ) controllable.

No matter howIn view of the one goes about defining K, the definitions of H, M, N, and V given above
show that a full-state observer is a system of equations of the form

ż = (A − KC)z + Ky w
b=z

These are also the equations of a time-invariant Kalman filter. In the time-invariant case, the main difference
between a a full-state observer and a Kalman filter is the way in which K is defined. In the case of an
observer, one typically apeals to spectrum assignment as we’ve just pointed out. In the Kalman filter case,
K is chosen in systematic manner to dealwith the trade-off between fast estimation, and reducing the effects
of measurement noise on the estimate if x. For the Kalman filtering problem to make sense, one must assume
that the sensor output to be processed is of the form y + n where n is an additive noise process with known
statistics. The way K is ultimately computed in this case is by solving what’s sometimes called a “matrix
Ricatti ” equation. Under suitable asumptions such as observability of (C, A),the K which results turns out
to stabilize A − KC. Kalman filtering is not restricted ti time-invariant systems or even to continuous time
systems. Kalman filtering has had a large impact in practice. The topic is typically addressed in detail in a
course on estimation theory.

Before concluding this section we point out that a full-state observer always has a dimension equal to
that of the process whose output w is to be estimated. Depending on A, C and L it is typically possible to
construct observers of lower dimension with the same capabilities. For example if L = C or if C = I one

need not not construct an observer at all. For in the first case w = y so we could define w
b = y while in the

second we could define wb = Ly since w = Ly . Even under less drastic assumptions, it is always possible to
generate an asymptotically correct estimate of w with an observer of dimension less than n.
15.3.3 Forced Systems

So far we have been concerned with the problem of asymptotically estimating w = Lx from measurement
of the output y = Cx of an unforced linear system with state equation ẋ = AX. In the case of a force
n-dimensional linear system such as

y = Cx, ẋ = Ax + Bu (15.27)

w = Lx can still be estimated asymptotically provided u can be sensed. Thus u cannot be an unmeasurable
input such as a “disturbance.” The way to go about estimating w in this case, is to first define observer
matrices H, K, M, N, and V so that the observer design equations

L = MV + NC (15.28)
V A = HV + KC (15.29)

hold, just as before, second to generate the estimate w


b using a modified observer of the form

ż = Hz + Ky + V Bu (15.30)
w
b = Mz + Ny (15.31)

We leave it to th reader to verify that if this is done, then with e = z − vx,

ė = He

and
b − w(t) = M eHt (z(0) − V x(0))
w(t)
just as before. these equations hold for all u. In essence what’s happening is that the influence of u on w is
being taken into account in just the right way by adding the term vBu to the differnetial equation for z.

15.4 Observer-Based Controllers

Consider the n-dimensional linear system

ẋ = Ax + Bu (15.32)
y = Cx + Du (15.33)
s = C̄x (15.34)

which we again regard as an internal model P of a physical process to be controlled. For any fixed F and G,
the state-feedback law
u = F x + Gr (15.35)
transforms P into a system of the form

ẋ = (A + BF )x + BGr (15.36)
y = (C + DF )x + DGr (15.37)
s = C̄x (15.38)

But this control cannot typically be implemented because F x may not be available as a sensor output. To
deal with this situation, suppose, instead of (15.35) we implement a control of the form

u=w
b + Gr (15.39)
ż = Hz + Ks + V Bu (15.40)
w
b = Mz + Ns (15.41)

where (15.40) and (15.41) is an F -observer designed on the basis on the equations

F = M V + N C̄ (15.42)
V A = HV + K C̄ (15.43)

Note that (15.39) - (15.41) define a controller C of the form

ż = (H + V BF )z + Ks + V BGr (15.44)
u = M z + N s + Gr (15.45)

We call C an observer-based controller. The closed-loop system with input r and output y consisting of P
and C is thus described by the system of equations

        
x ẋ A + BN C̄ BM x BG
y = [ C + DN C̄ DM ]   + DGr  =   +  r (15.46)
z ż K C̄ H + V BF z V BG

To understand the effect of C on P, let us first recall that

b − Fx = Me
w

and
ė = He (15.47)

where e = z − V x. Using the first of these equations and (15.39) it is possible to wtire

u = F x + Gr + M e

Application of this formula foru to the original open-loop equations defining P thus yields

ẋ = (A + BF )x + BM e + BGr (15.48)
y = (C + DF )x + DM e + DGr (15.49)

Combined with (15.47) we thus obtain the system


        
x ẋ A + BF BM x BG
y = [ C + DF DM ]    =   +  r (15.50)
e ė 0 H e 0

It is easy to verify that we could have gotton to these equations directly from (15.46) by simply making the
change of variables     
x I 0 x
=
e −V I z
In other words, the systems defined by (15.46) and (15.50) are similar and thus have the same spectrum
and the same transfer matrix. From this and the structure of (15.50) we can draw the following important
conclusions:

• Spectral Seperation Property: The spectrum of the closed-loop system consisting of P and the
observer-based controller C defined by (15.44) and (15.45) separates into two disjoint subsets, one the
spectrum of the observer matrix H and the other the spectrum of the system matrix A + BF which
would have resulted had we been able to apply the state-feedback law u = F x + Gr directly to P.
• Transfer Matrix Property: The transfer matrix from r to y of the closed-loop system consisting
of P and the observer-based controller C defined by (15.44) and (15.45) is exactlythe same as the
transfer matrix from r to y which would have resulted had we been able to apply the state-feedback
law u = F x + Gr directly to P.

The spectral is a direct consequence of the block triangular structure of the system matrix appearing in
(15.50).

The preceding provides a procedure for stabilizing any process model P which is controllable and which is
observable through its sensor output s. The spectrum can be freely assigned with F because of controllability

of (A, B) and, at least in the case when a full-state observer is used, the spectrum of H = A − K C̄ with K
because of observability of (C̄, A). As you might suspect, there is much, much more to the story than this.
Bibliography

[1] F. R. Gantmacher. The Theory of Matrices, volume 1 & 2. Chelsea Publishing, 1960.
[2] A. F. Filippov. Differential equations with discontinuous right-hand side. Amer. Math. Soc. Translations,
pages 199–231, 1964.

235
Index

A-invariant subspace, 109 closed orbit, 79


L-observer, 231 closed set, 53
Cl n , 32 closed set , 55
IRn , 32 Closed-loop control, 226
gcd, 5 codomain, 40
C,
l 3 cofactor, 9
C[s],
l 4 column span, 46
ith unit n- vector, 12 commutative, 143
n - vector, 8 companion form, 119
p - norm, 52 complete, 56, 141
IK, 3 completion, 110
IK[s], 4 complex conjugate, 145
IR, 3 composition, 41
IR[s], 4 composition rule, 78, 94
lcm, 5 congruence canonical form, 155
congruence transformation, 153
Fibonacci Sequence, 60 congruent, 153
invariant , 80 connected, 53
constant coefficient, 87
Abel, Jacobi, Liouville Formula, 96 contained, 32
Abelian group, 143 continuous-time linear dyanmical system, 165
adjoint, 145 continuously differentiable, 53, 216
algebraic multiplicity, 106, 124 control canonical form, 186
algebraic system, 3 controllability Gramian, 174
asymptotic estimate, 230 controllability index, 181
asymptotic stability matrices, 129, 214 controllability matrix, 179
attractor, 79 controllability reduction, 176
autonomous, 80, 87 controllable on [t0 , T ], 175
controllable space, 178
Banach Spaces, 56 controlled output, 225
basis, 37 converge, 55
basis matrices, 110 converge in the mean, 140
Bessel’s inequality, 139 convergent sequence, 55
block diagonal matrix, 111 convolution integral, 167
block triangular matrix, 111 convolution integrals, 97
bounded, 52, 67 convolution sum, 168
coordinates, 37
Cartesian Product, 29 coprime, 5
Cauchy Sequence, 55 cyclic, 117
Cauchy-Schwartz Inequality, 135 cyclic matrix, 118
characteristic equation, 106
characteristic polynomial, 106 dependent variables, 58
closed, 32 derivative, 54
closed interval, 53 diagonal matrix, 8

236
diamond notation, 59 fundamental matrix, 91
difference equation, 58
dimension, 40 Gauss Elimination, 22
dimension , 62 general linear group, 144
direct sum, 36 generator, 117
discrete dynamical system, 62 geometric multiplicity, 106, 124
discrete-time asymptotic stability matrices, 130, 215 globally Lipschitz continuous, 67
discrete-time linear dyanmical system, 165 golden mean, 61
discrete-time Markov sequence, 169 Gramian, 135
discrete-time realization, 170 greatest common divisor, 5
discrete-time stable matrices, 130, 216 group, 143
discrete-time state transition matrix , 168
discrete-time transfer matrix, 169 Hermetian, 147
discrete-time unstable matrices, 130, 216 Hermitian metric, 131
discrete-time weighting pattern, 170 Hilbert spaces, 141
divides, 4 homogeneous, 87
domain, 40
dot notation, 59 ideal, 4
dual, 196 identity, 143
dynamical system, 62, 77 identity matrix, 9
if and only if, 15
echelon form, 21 iff, 15
eigenspace, 105 image, 45, 46
eigenvalue, 105 impulse matrix realization, 171
eigenvector , 105 impulse response matrix, 167
elementary column operations, 27 indefinite matrices, 156
elementary divisors, 126 independent, 36
elementary row operation, 20 independent variables, 58
elementary row operations, 20 indeterminate, 4
elements, 7 indistinguishable, 187
endomorphism, 48 inf, 76
entries, 7 infimum, 76
epimorphism, 46 infinite dimensional, 36
equilibrium point, 78 infinity norm, 52
equivalence class, 29 initial value problem, 66
equivalence relation, 29 injective, 47
equivalent matrices, 28, 30 inner product, 131
estimation error, 230 input, 87, 165
Euclidean metric, 131 input , 165
Euclidean Norm, 52 input-output formula, 166
Euclidean space, 131 integral manifold, 80
Euler’s method, 82 intersection, 34
exponentially stable, 210 invaraint, 112
invariant factors, 120
field, 3 inverse, 143
finite dimensional, 36 invertible, 4
first integral, 80 irreducible, 4
forcing function, 87 isometric, 145
Fourier coefficients, 141 isomorphic, 48
Fourier series, 141 isomorphic systems, 206
full rank, 10 isomorphism, 48
full-state observer, 231
function, 40 Jordon Block, 125
Jordon Normal Form, 126 neighborhood, 53
nominal solution, 88
kernel, 47 nondestabilizing, 222
nonsingular, 15
Laplace Expansion, 9 norm, 51, 132
least common multiple, 5 normal, 146
left-half open interval, 53 normalize, 134
left-right equivalence canonical form, 29 normalized, 133, 134
left-right equivalence transformation, 30 normed vector space, 51
left-right equivalent matrices, 30
limit, 55 observability Gramian, 188
linear algebraic equations, 25 observability index, 197
linear combination, 36 observability matrix, 194
linear controller, 225 observability reduction, 190
linear matrix equations, 27 observable on [t0 , T ], 189
linear state differential equation, 87 observer design equations, 231
linear transformation, 41 observer-based, 233
linear vector spaces, 31 one-norm, 52
linearization, 87, 88 one-to-one, 47
linearly dependent, 36 onto, 46
linearly independent, 36 open ball, 53
Lipschitz Constant, 67 open interval, 53
locally Lipschitz Continuous, 67 open set, 53
lower triangular matrix, 8 open-loop control input, 225
Lyapunov Transformation, 99 order, 10, 58
ordinary differential equations, 58
main diagonal, 8 ordinary recursion equation, 58
Markov sequence, 167 orthogonal, 133, 144
Mathieu Equation, 98 orthogonal complement, 145
matrix, 7 orthogonal projection, 137
matrix addition, 11 orthogonal set, 134
matrix exponential, 92, 166 orthonormal, 134
matrix inverses, 15 orthonormal basis, 143
matrix inversion, 24 outer product, 132
matrix multiplication, 12 output, 165
matrix polynomial, 17
matrix representation, 43 pairwise orthogonal, 133
maximal interval of existence, 76 partitioned matrix, 7
maximal solution, 76 period, 79
metric space, 131 periodic , 79
minimal, 192 permutation group, 144
minimal polynomial, 114, 115 phase variable form., 63
minimal realization, 205 Picard Iteration, 69
minimum polynomial, 114 piecewise-continuous function, 54
minor, 10 positive definite, 156
modular distributative rule, 35 positive definite matrix, 156
monic, 4, 106 positive semi-definite function, 51
monomorphism, 47 positive semidefinite, 156
positive semidefinite matrix, 156
natural numbers, 165 positive-definite functions, 51
necessary and sufficient, 15 prime, 4
negative definite matrix, 156 principal axes, 154
negative semidefinite matrix, 156 principal leading minors, 157
principal subarray, 157 state equation, 166
process model, 225 state space, 77
proper, 32, 167 state space system, 62
pulse response matrix, 168 state transition matrix, 91, 129, 215
state variables, 62
quotient, 4 state-feedback control, 226
quotient set, 29 stationary, 87, 165
step size, 81
radially unbounded, 216 strictly proper, 167
rank, 10 sub-multiplicative property, 52
rational canonical form, 120 subarray, 9
reachable, 173 subset, 32
reachable space, 173 subspace, 32
readable, 225 sum, 33, 42
readout equation, 166 sup, 52
real bilinear form, 152 superposition rule, 41
real quadratic form, 152 supremum, 52, 67
realization, 170, 171, 201 surjective, 46
reduced echelon form, 22 symmetric, 106, 147, 153
reducible, 100 symmetric matrix, 8
reference input, 225 symmetry, 29
reflexivity, 29
relation, 29 time invariant, 87
remainder, 4 time-invariant, 165
representation, 37 time-varying, 165
resolvent, 116 trace, 95
restriction, 65, 110 trajectory, 78
right half-open interval, 53 transfer matrix , 167
right inverse, 15 transfer matrix realization, 171
rings, 4 transitivity, 29
Runge-Kutta method, 83 transpose, 10
triangle inequality, 51
sampling time, 169
scalar multiplication, 11, 31 uncontrollable spectrum, 184
scalars, 31 unforced, 87
semisimple, 107 uniform convergence, 56
sensor output, 225 uniformly asymptotically stable, 210
signature, 155 uniformly stable, 209
similar, 49, 103 union, 33
similar linear systems, 183 unit, 143
similar systems, 206 unit n - vectors, 36
similarity invariant, 112 unit vector, 8
similarity transformation, 49, 103 unitary group, 145
singular, 15 unitary matrix, 145
size, 7 unitary space, 131
span, 36 unobservable space, 194
spans, 36 unobservable space , 188
spectrum, 106 unobservable spectrum, 200
square matrix, 8 unstable linear systems, 212
square root, 157 unstable matrices, 129, 215
stable matrices, 129, 215 upper triangular matrix, 8
stable systems, 212
state, 77, 165 value, 40
variation of constants formula, 97
vector addition, 31
vectors, 31

weighting pattern, 170

zero matrix, 8
zero vector, 8, 31

You might also like