0% found this document useful (0 votes)
8 views784 pages

Index Theory in Mathematics and Physics

The document is a comprehensive text on Index Theory, detailing its applications in mathematics and physics. It covers a range of topics including Fredholm operators, differential equations, Sobolev spaces, and the Atiyah-Singer index formula, structured into four main parts. Each chapter provides in-depth discussions and methodologies relevant to the theoretical foundations and practical implications of index theory.

Uploaded by

Sanchita Sharma
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
8 views784 pages

Index Theory in Mathematics and Physics

The document is a comprehensive text on Index Theory, detailing its applications in mathematics and physics. It covers a range of topics including Fredholm operators, differential equations, Sobolev spaces, and the Atiyah-Singer index formula, structured into four main parts. Each chapter provides in-depth discussions and methodologies relevant to the theoretical foundations and practical implications of index theory.

Uploaded by

Sanchita Sharma
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Index Theory

with Applications to Mathematics and Physics

David Bleecker
Bernhelm Booß–Bavnbek
To David .
Contents

Synopsis xi
Preface xvi

Part I. Operators with Index and Homotopy Theory 1


Chapter 1. Fredholm Operators 2
1. Hierarchy of Mathematical Objects 2
2. The Concept of Fredholm Operator 3
3. Algebraic Properties. Operators of Finite Rank. The Snake Lemma 5
4. Operators of Finite Rank and the Fredholm Integral Equation 9
5. The Spectra of Bounded Linear Operators: Basic Concepts 10

Chapter 2. Analytic Methods. Compact Operators 12


1. Analytic Methods. The Adjoint Operator 12
2. Compact Operators 18
3. The Classical Integral Operators 25
4. The Fredholm Alternative and the Riesz Lemma 26
5. Sturm-Liouville Boundary Value Problems 28
6. Unbounded Operators 34
7. Trace Class and Hilbert-Schmidt Operators 53
Chapter 3. Fredholm Operator Topology 63
1. The Calkin Algebra 63
2. Perturbation Theory 65
3. Homotopy Invariance of the Index 68
4. Homotopies of Operator-Valued Functions 72
5. The Theorem of Kuiper 77
6. The Topology of F 81
7. The Construction of Index Bundles 82
8. The Theorem of Atiyah-Jänich 88
9. Determinant Line Bundles 91
10. Essential Unitary Equivalence and Spectral Invariants 108
Chapter 4. Wiener-Hopf Operators 120
1. The Reservoir of Examples of Fredholm Operators 120
2. Origin and Fundamental Significance of Wiener-Hopf Operators 121
3. The Characteristic Curve of a Wiener-Hopf Operator 122
4. Wiener-Hopf Operators and Harmonic Analysis 123
5. The Discrete Index Formula. The Case of Systems 125
6. The Continuous Analogue 129
vii
viii CONTENTS

Part II. Analysis on Manifolds 133

Chapter 5. Partial Differential Equations in Euclidean Space 134


1. Linear Partial Differential Equations 134
2. Elliptic Differential Equations 137
3. Where Do Elliptic Differential Operators Arise? 139
4. Boundary-Value Conditions 141
5. Main Problems of Analysis and the Index Problem 143
6. Numerical Aspects 143
7. Elementary Examples 144

Chapter 6. Differential Operators over Manifolds 156


1. Differentiable Manifolds — Foundations 157
2. Geometry of C∞ Mappings 160
3. Integration on Manifolds 166
4. Exterior Differential Forms and Exterior Differentiation 171
5. Covariant Differentiation, Connections and Parallelity 176
6. Differential Operators on Manifolds and Symbols 181
7. Manifolds with Boundary 190

Chapter 7. Sobolev Spaces (Crash Course) 193


1. Motivation 193
2. Definition 194
3. The Main Theorems on Sobolev Spaces 200
4. Case Studies 203

Chapter 8. Pseudo-Differential Operators 206


1. Motivation 206
2. Canonical Pseudo-Differential Operators 210
3. Principally Classical Pseudo-Differential Operators 214
4. Algebraic Properties and Symbolic Calculus 228
5. Normal (Global) Amplitudes 231

Chapter 9. Elliptic Operators over Closed Manifolds 237


1. Mapping Properties of Pseudo-Differential Operators 237
2. Elliptic Operators — Regularity and Fredholm Property 239
3. Topological Closure and Product Manifolds 242
4. The Topological Meaning of the Principal Symbol — A Simple Case
Involving Local Boundary Conditions 244

Part III. The Atiyah-Singer Index Formula 251

Chapter 10. Introduction to Topological K-Theory 252


1. Winding Numbers 252
2. The Topology of the General Linear Group 257
3. Elementary K-Theory 262
4. K-Theory with Compact Support 266
5. Proof of the Periodicity Theorem of R. Bott 269

Chapter 11. The Index Formula in the Euclidean Case 275


1. Index Formula and Bott Periodicity 275
CONTENTS ix

2. The Difference Bundle of an Elliptic Operator 276


3. The Index Theorem for Ellc (Rn ) 281

Chapter 12. The Index Theorem for Closed Manifolds 284


1. Pilot Study: The Index Formula for Trivial Embeddings 285
2. Proof of the Index Theorem for Nontrivial Normal Bundle 287
3. Comparison of the Proofs 301

Chapter 13. Classical Applications (Survey) 310


1. Cohomological Formulation of the Index Formula 311
2. The Case of Systems (Trivial Bundles) 316
3. Examples of Vanishing Index 317
4. Euler Characteristic and Signature 319
5. Vector Fields on Manifolds 325
6. Abelian Integrals and Riemann Surfaces 329
7. The Theorem of Hirzebruch-Riemann-Roch 333
8. The Index of Elliptic Boundary-Value Problems 337
9. Real Operators 356
10. The Lefschetz Fixed-Point Formula 357
11. Analysis on Symmetric Spaces: The G-equivariant Index Theorem 360
12. Further Applications 362

Part IV. Index Theory in Physics and the Local Index Theorem 363

Chapter 14. Physical Motivation and Overview 364


1. Classical Field Theory 365
2. Quantum Theory 373

Chapter 15. Geometric Preliminaries 394


1. Principal G-Bundles 394
2. Connections and Curvature 396
3. Equivariant Forms and Associated Bundles 400
4. Gauge Transformations 409
5. Curvature in Riemannian Geometry 414
6. Bochner-Weitzenböck Formulas 434
7. Characteristic Classes and Curvature Forms 443
8. Holonomy 455

Chapter 16. Gauge Theoretic Instantons 460


1. The Yang-Mills Functional 460
2. Instantons on Euclidean 4-Space 466
3. Linearization of the Moduli Space of Self-dual Connections 489
4. Manifold Structure for Moduli of Self-dual Connections 496

Chapter 17. The Local Index Theorem for Twisted Dirac Operators 513
1. Clifford Algebras and Spinors 513
2. Spin Structures and Twisted Dirac Operators 525
3. The Spinorial Heat Kernel 538
4. The Asymptotic Formula for the Heat Kernel 549
5. The Local Index Formula 576
x CONTENTS

6. The Index Theorem for Standard Geometric Operators 594

Chapter 18. Seiberg-Witten Theory 643


1. Background and Survey 643
2. Spinc Structures and the Seiberg-Witten Equations 655
3. Generic Regularity of the Moduli Spaces 663
4. Compactness of Moduli Spaces and the Definition of S-W Invariants 685

Appendix A. Fourier Series and Integrals - Fundamental Principles 705


1. Fourier Series 705
2. The Fourier Integral 707
Appendix B. Vector Bundles 712
1. Basic Definitions and First Examples 712
2. Homotopy Equivalence and Isomorphy 716
3. Clutching Construction and Suspension 718
Bibliography 723
Index of Notation 741

Index of Names/Authors 749


Subject Index 757
Synopsis

Preface. Target Audience and Prerequisites. Outline of History. Further


Reading. Questions of Style. Acknowledgments and Dedication.

Chapter 1. Fredholm Operators. Hierarchy of Mathematical Objects. The


Concept of a Bounded Fredholm Operator in Hilbert Space. Algebraic Properties.
Operators of Finite Rank. The Snake Lemma of Homological Algebra. Product
Formula. Operators of Finite Rank and the Fredholm Integral Equation. The
Spectra of Bounded Linear Operators (Terminology).

Chapter 2. Analytic Methods. Compact Operators. Adjoint and Self-


Adjoint Operators - Recalling Fischer-Riesz. Dual Characterization of Fredholm
Operators. Compact Operators: Spectral Decomposition, Why Compact Oper-
ators also are Called Completely Continuous, K as Two-Sided Ideal, Closure of
Finite-Rank Operators, and Invariant under ∗ . Classical Integral Operators. Fred-
holm Alternative and Riesz Lemma. Sturm-Liouville Boundary Value Problems.
Unbounded Operators: Comprehensive Study of Linear First Order Differential
Operators Over S 1 : Sobolev Space, Dirac Distribution, Normalized Integration
Operator as Parametrix, The Index Theorem on the Circle for Systems. Closed Op-
erators, Closed Extensions, Closed (not necessarily bounded) Fredholm Operators,
Composition Rule, Symmetric and Self–Adjoint Operators, Formally Self-Adjoint
and Essentially Self-Adjoint. Spectral Theory. Metrics on the Space of Closed
Operators. Trace Class and Hilbert-Schmidt Operators.

Chapter 3. Fredholm Operator Topology. Calkin Algebra and Atkinson’s


Theorem. Perturbation Theory: Homotopy Invariance of the Index, Homotopies
of Operator-Valued Functions, The Theorem of Kuiper. The Topology of F: The
Homotopy Type, Index Bundles, The Theorem of Atiyah-Jänich. Determinant
Line Bundles: The Quillen Determinant Line Bundle, Fredholm Determinants, The
Segal-Furutani Construction. Spectral Invariants: Essentially Unitary Equivalence,
What Is a Spectral Invariant? Eta Function, Zeta Function, Zeta Regularized
Determinant.

Chapter 4. Wiener-Hopf Operators. The Reservoir of Examples of Fred-


holm Operators. Origin and Fundamental Significance of Wiener-Hopf Operators.
The Characteristic Curve of a Wiener-Hopf Operator. Wiener-Hopf Operators
and Harmonic Analysis. The Discrete Index Formula. Noether’s Theorem for
the Hilbert Transform. The Case of Systems. The Continuous Analogue.

Chapter 5. Partial Differential Equations in Euclidean Space, Re-


visited. Review of Classical Linear Partial Differential Equations: Constant and
xi
xii SYNOPSIS

Variable Coefficients, Wave Equation, Heat Equation, Laplace Equation, Charac-


teristic Polynomial. Elliptic Differential Equations: Where Do Elliptic Differential
Operators Arise? Boundary-Value Conditions. Main Problems of Analysis and
the Index Problem. Calculations. Elementary Examples. The Noether(-Hellwig-
Vekua) Problem with Nonvanishing Index.

Chapter 6. Differential Operators over Manifolds. Motivation. Differ-


entiable Manifolds — Foundations: Tangent Space. Cotangent Space. Geometry
of C∞ Mappings: Embeddings, Immersions, Submersions, Embedding Theorems.
Integration on Manifolds: Hypersurfaces, Riemannian Manifolds, Geodesics, Ori-
entation. Exterior Differential Forms and Exterior Differentiation. Covariant Dif-
ferentiation, Connections and Parallelity: Connections on Vector Bundles, Parallel
Transport, Connections on the Tangent Bundle, Clifford Modules and Operators
of Dirac Type. Differential Operators on Manifolds and Symbols: Our Data, Sym-
bolic Calculus, Formal Adjoints. Elliptic Differential Operators. Definition and
Standard Examples. Manifolds with Boundary.

Chapter 7. Sobolev Spaces (Crash Course). Motivation. Equivalence of


Different Local Definitions. Various Isometries. Global, Coordinate-Free Definition.
Embedding Theorems: Dense Subspaces; Truncation and Mollification; Differential
Embedding; Rellich Compact Embedding. Sobolev Spaces Over Half Spaces. Trace
Theorem. Case Studies: Euclidean Space and Torus; Counterexamples.

Chapter 8. Pseudo-Differential Operators. Motivation: Fourier Inver-


sion; Symbolic Calculus; Quantization. Canonical and Principally Classical Pseudo-
Differential Operators. Pseudo-Locality; Singular Support. Standard Examples:
Differential Operators; Singular Integral Operators. Oscillatory Integrals. Kuran-
ishi Theorem. Change of Coordinates. Pseudo-Differential Operators on Mani-
folds. Graded ∗ -Algebra. Invariant Principal Symbol; Exact Sequence; Noncanon-
ical Op-Construction as Right Inverse. Coordinate-Free (Truly Global) Approach:
Bokobza-Haggiag-Fourier Transformation; Bokobza-Haggiag Amplitudes; Bokobza-
Haggiag Invertible Op-Construction; Approximation of Differential Operators.

Chapter 9. Elliptic Operators over Closed Manifolds. Continuity of


Pseudo-Differential Operators between Sobolev Spaces. Parametrices for Elliptic
Operators: Regularity and Fredholm Property. Topological Closures. Outer Tensor
Product on Product Manifolds. The Topological Meaning of the Principal Symbol
(Simple Case Involving Local Boundary Conditions).

Chapter 10. Introduction to Topological K-Theory. Winding Numbers.


One-Dimensional Index Theorem. Counter-Intuitive Dimension Two: Bending a
Plane. The Topology of the General Linear Group. The Grothendieck Ring of
Vector Bundles. K-Theory with Compact Support. Proof of the Bott Periodicity
Theorem.

Chapter 11. The Index Formula in the Euclidean Case. Index Formula
and Bott Periodicity: Three Integer Invariants. The Difference Bundle of an Elliptic
Operator: Operators Equal to the Identity at Infinity; Complexes of Vector Bundles
with Compact Support; Symbol Class in K-Theory with Compact Support. The
Index Theorem for Ellc (Rn ).
SYNOPSIS xiii

Chapter 12. The Index Theorem for Closed Manifolds. K-Theoretic


Proof of the Index Theorem by Embedding: Pilot Study — The Index Theorem
for Embeddings with Trivial Normal Bundle. Proof of the Index Theorem for Non-
Trivial Normal Bundle: The Difference Element Construction, Revisited; Symbol
Class; Thom Isomorphism of K-Theory; Definition of the Topological Index; Defi-
nition of the Analytic Index; Foundations of Equivariant K-Theory. Multiplicative
Property: Formulation; How it Fits into the Embedding Proof; Proof of the Mul-
tiplicative Property. Short Comparison of the Cobordism, the Embedding and the
Heat Equation Proof; Outlook to Spectral Theory, Asymmetry, and Inverse Prob-
lems.

Chapter 13. Classical Applications (Survey). General Appraisal. Coho-


mological Formulation of the Index Formula: Comparison of K-Theoretic and Co-
homological Thom Isomorphisms; Orientation Class; Chern Classes; Chern Charac-
ter; Todd Class; Chern Character Defect. The Case of Systems (Trivial Bundles).
Examples of Vanishing Index. Euler Number and Signature. Vector Fields on
Manifolds. Abelian Integrals and Riemann Surfaces. The Theorem of Riemann-
Roch-Hirzebruch. The Index of Elliptic Boundary-Value Problems. Real Opera-
tors. The Lefschetz Fixed-Point Formula. Analysis on Symmetric Spaces. Further
Applications.

Chapter 14. Physical Motivation and Overview. Mode of Reasoning in


Physics. String Theory and Quantum Gravity. The Experimental Side. Classical
Field Theory: Newton-Maxwell-Lorentz, Faraday 2-Form, Abstract Flat Minkowski
Space-Time, Relativistic Mass, Relativistic Kinetic Energy, Inertial System, Lorentz
Transformations and Poincaré Group, Relativistic Deviation from Flatness, Twin
Paradox, Variational Principles. Kaluza-Klein Theory: Simultaneous Geometriza-
tion of Electro-Magnetism and Gravity, Other Grand Unified Theories, String The-
ory. Quantum Theory: Photo-Electric Effect, Atomic Spectra, Quantizing Energy,
State Spaces of Systems of Particles, Basic Interpretive Assumptions. Heisenberg
Uncertainty Principle. Evolution with Time — The Schrödinger Picture. Nonrela-
tivistic Schrödinger Equation and Atomic Phenomena. Minimal Replacement and
Covariant Differentiation. Anti-Particles and Negative-Energy States. Unreason-
able Success of the Standard Model. Dirac Operator vs. Klein-Gordon Equation.
Feynman Diagrams.

Chapter 15. Geometric Preliminaries. Principal G-Bundles; Hopf Bun-


dle. Connections and Curvature: Connection 1-Form; Maurer-Cartan Form; Hori-
zontal Lift. Equivariant Forms and Associated Bundles: Associated Vector Bundles,
Equivariance, and Basic Forms; Horizontal Equivariant Forms; Covariant Differen-
tiation and the General Bianchi Identity; Inner Products, Hodge Star Operator,
and Formal Adjoints. Gauge Transformations: Distinguishing Gauge Transforma-
tions from Automorphisms; The Group of Gauge Transformations; The Action
of Gauge Transformations on Connections; Lie Algebra Analogy and Infinitesimal
Action. Curvature in Riemannian Geometry: The Bundle of Linear Frames; Con-
nections and Forms; Kozul Connection; The Orthonormal Frame Bundle; Metric
Connections; The Fundamental Lemma of Riemannian Geometry and the Levi-
Civita Connection; Local Coordinates and Christoffel Symbols; The Curvature of
the Levi-Civita Connection; First and Second Bianchi Identities; Ricci and Scalar
xiv SYNOPSIS

Curvature as Contractions, and Einstein’s Equation; All Possible Curvature Ten-


sors on Rn and the Kulkarni-Nomizu Product; Curvature Parts on 4-Manifolds and
Self-Duality. Bochner-Weitzenböck Formulas: Fibered Products; Contractions and
Components; The Connection Laplacian, the Hodge Laplacian, and the Bochner-
Weitzenböck Formula; Special Cases. Characteristic Classes and Curvature Forms:
Chern Classes as Curvature Forms; The Pfaffian; Pontryagin Classes; Other Char-
acteristic Classes Related to Index Theory; Multiplicative Classes; Todd Class
and L-Polynomials; Recalculating Characteristic Classes; Unifications on Almost-
Complex Manifolds. Holonomy.

Chapter 16. Gauge Theoretic Instantons. The Yang-Mills Functional.


Instantons on Euclidean 4-Space. Linearization of the Manifold of Moduli of Self-
dual Connections. Manifold Structure for Moduli of Self-dual Connections.

Chapter 17. The Local Index Theorem for Twisted Dirac Operators.
Clifford Algebras and Spinors: Clifford Algebra Basics; Spin Groups and Double
Cover; Spinor Representations; Supertrace. Spin Structures and Twisted Dirac Op-
erators: Čech Cohomology; Admittance of Spin Structures; Standard and Twisted
Dirac Operators; Chirality. The Spinorial Heat Kernel: Index, Spectral Asymme-
try and the Existence of the Heat Kernel; Solving the Spinorial Heat Equations;
Calculating Index and Supertrace; General Heat Kernels. The Asymptotic Formula
for the Heat Kernel: Why Asymptotic Expansion? The Radial Gauge; About the
Geometry of the Ball; Further Approximations. The Local Index Formula: Con-
tent and Meaning of the Local Index Formula; How the Curvature Terms Arise in
the Heat Asymptotics; The case m = 1 (surfaces); The case m = 2 (4-manifolds);
Proof of the Local Index Formula for Arbitrary Even Dimensions; Index Theorem
for Twisted Dirac Operators; A b Genus; Rokhlin’s Theorem. The Index Theorem
for Standard Geometric Operators: Index Theorem for Generalized Dirac Oper-
ators; Twisted Generalized Dirac Operators; The Hirzebruch Signature Formula;
The Gauss-Bonnet-Chern Formula; The Generalized Yang-Mills Index Theorem;
The Hirzebruch-Riemann-Roch Formula for Kähler Manifolds.

Chapter 18. Seiberg-Witten Theory. Background and Survey: Inter-


section Form and Homotopy Type of Compact Oriented Simply-Connected Four-
Manifolds; Unimodular Forms and Freedman’s Theorem; Existence of Differen-
tiable Structures; Donaldson’s Polynomial Invariants; Results Concerning Symplec-
tic Manifolds; Purely Geometric Applications. Spinc Structures and the Seiberg-
Witten Equations: Admittance of Spinc Structures, Spinc Dirac Operators; Moti-
vating and Defining the Unperturbed and the Perturbed Seiberg-Witten Equations.
Generic Regularity of the Moduli Spaces: Gauge Transformations; Moduli Space of
Solutions of the Perturbed S-W Equations; The Formal Dimension of the Moduli
Space; Seiberg-Witten Function; Quotient Manifolds; Manifold Structure for the
Parametrized Moduli Space; Generic Regularity; A Priori Bounds; Sobolev Esti-
mates; Compactness of Moduli Spaces; Definition of the S-W Invariant; Metric
Dependence of Connections and Dirac Operators; Oriented Cobordism; Fredholm
Transversality; Full Invariance of the S-W Invariant.

Appendix A. Fourier Series and Integrals — Fundamental Principles.


Fourier Series: The Fundamental Function Spaces on S 1 ; Density; Orthonormal
Basis; Fourier Coefficients; Plancherel’s Identity; Product and Convolution. The
SYNOPSIS xv

Fourier Integral: Different Integral Conventions; Duality Between Local and Global
— Point and Neighborhood — Multiplication and Differentiation — Bounded and
Continuous; Fourier Inversion Formula; Plancherel and Poisson Summation Formu-
lae; Parseval’s Equality; Higher Dimensional Fourier Integrals.
Appendix B. Vector Bundles. Basic Definitions and First Examples. Ho-
motopy Equivalence and Isomorphy. Clutching Construction and Suspension.
Bibliography. Key References. Classical and Recent Textbooks. References
to Technical Details; History; Perspectives.

Table 0.1. Suggested packages (selections, curricula) for upper-


undergraduate and graduate classes/seminars and teach-yourself

Aims Logical Order by Boxed Chapter Numbers

Index Theorem and App. B → 1.1-1.3, 2.1,2.2, 3.1-3.8, Thm. 5.11 → 6-7
Topolog. K-Theory → 8.5,9 → 10-12.2 → 13.1-13.5, 13.10,13.11 → 18.1

Index Theorem via App. A → 1.2,3.3 → 5.2, 6-7, 8.3 → 8.4,9.2 → 12.3
Heat Equation → 15 → 17

Gauge-Theoretic 14 → 5-6 → 9.2 → 12.3 → 13.8,13.11 → 15 → 16


Physics → 18

Spectral Geometry 1-4 → 6,15 → 12.3 → 13.4-13.11 → 17.4-17.6 → 18

Global and Micro- App. A,B → 1, 2, 3.1-3.5 → 3.10,4 → 5-6 → 14-15


Local Analysis → 7-9 → 10.1, 10.2 → 12.3,13 → 16,17.6,18
Preface

Target Audience and Prerequisites. The mathematical philosophy of index


theory and all its basic concepts, technicalities and applications are explained in
Parts I-III. Those are the easy parts. They are written for upper undergraduate
students or graduate students to bridge the gap between rule-based learning and
first steps towards independent research. They are also recommended as general
orientation to mathematics teachers and other senior mathematicians with different
background. All interested can pick up a single chapter as bedside reading.
In order to enjoy reading or even work through Parts I-III, we expect the
reader to be familiar with the concept of a smooth function and a complex separable
Hilbert space. Nothing more — but a will to acquire specialized topics in functional
analysis, algebraic topology, elliptic operator theory, global analysis, Riemannian
geometry, complex variables, and some other subjects. Catching so many different
concepts and fields can make the first three Parts a bit sophisticated for a busy
reader. Instead of ascending systematically from simple concepts to complex ones in
the classical Bourbaki style, we present a patch-work of definitions and results when
needed. In each chapter we present a couple of fully comprehensible, important,
deep mathematical stories. That, we hope, is sufficient to catch our four messages:

(1) Index theory is about regularization, more precisely, the index quantifies
the defect of an equation, an operator, or a geometric configuration from
being regular.
(2) Index theory is also about perturbation invariance, i.e., the index is a
meaningful quantity stable under certain deformations and apt to store
certain topological or geometric information.
(3) Most important for many mathematicians, the index interlinks quite di-
verse mathematical fields, each with its own very distinct research tradi-
tion.
(4) Index theory trains the student to recognize all the elementary topics of
linear algebra in finite dimensions in the sophisticated topics of infinite-
dimensional and nonlinear analysis and geometry.
Part IV is different. It is also self-contained. Choosing one or two chapters of
this Part IV of the book would make a suitable text for a graduate course in se-
lected topics of global analysis. All concepts will be explained fully and rigorously,
but much shorter than in the first Parts. This last Part is written for graduate stu-
dents, PhD students and other experienced learners, interested in low-dimensional
topology and gauge-theoretic particle physics. We try to explain the very place of
index theory in geometry and for revisiting quantum field theory. There are thou-
sands of other calculations, observations and experiments. But there is something
special about the actual and potential contributions of index theory. Index theory
xvi
PREFACE xvii

is about chirality (asymmetry) of zero modes in the spectrum and classifies connec-
tions (back ground fields) and a variety of other intrinsic properties in geometry
and physics. It is not just about some more calculations, some more numbers and
relations.
Outline of History. When first considering infinite-dimensional linear spaces,
there is the immediate realization that there are injective and surjective linear
endomorphisms which are not isomorphisms, and more generally the dimension of
the kernel minus that of the cokernel (i.e., the index) could be any integer. However,
in the classical theory of Fredholm integral operators which goes back at least to the
early 1900s (see [147]), one is dealing with compact perturbations of the identity
and the index is zero. Fritz Noether (in his study [324] of singular integral
operators and the oblique boundary problem for harmonic functions, published in
1920), was the first to encounter the phenomenon of a nonzero index for operators
naturally arising in analysis and to give a formula for the index in terms of a winding
number constructed from data defining the operator. Over some decades, this result
was expanded in various directions by G. Hellwig, I.N. Vekua and others (see
[425]), contrary to R. Courant’s and D. Hilbert’s expectation in [116] that
“linear problems of mathematical physics which are correctly posed behave like a
system of N linear algebraic equations in N unknowns”, i.e., they should satisfy the
Fredholm alternative and always yield vanishing index. Meanwhile, many working
mainly in abstract functional analysis were producing results, such as the stability
of the index of a Fredholm operator under perturbations by compact operators or
bounded operators of sufficiently small operator norm (e.g., first J.A. Dieudonné
[120], followed by F.V. Atkinson [49], B. Yood [449], I.Z. Gohberg and M.G.
Krein [178], etc.).
Around 1960, the time was ripe for I.M. Gelfand [159] to propose that the
index of an elliptic differential operator (with suitable boundary conditions in the
presence of a boundary) should be expressible in terms of the coefficients of highest
order part (i.e., the principal symbol) of the operator, since the lower order parts
provide only compact perturbations which do not change the index. Indeed, a con-
tinuous, ellipticity-preserving deformation of the symbol should not affect the index,
and so Gelfand noted that the index should only depend on a suitably defined
homotopy class of the principal symbol. The hope was that the index of an elliptic
operator could be computed by means of a formula involving only the topology of
the underlying domain (the manifold), the bundles involved, and the symbol of the
operator. In early 1962, M.F. Atiyah and I.M. Singer discovered the (elliptic)
Dirac operator in the context of Riemannian geometry and were busy working at
Oxford on a proof that the A-genus
b of a spin manifold is the index of this Dirac
operator. At that time, S. Smale happened to pass through Oxford and turned
their attention to Gelfand’s general program described in [159]. Drawing on the
foundational and case work of analysts (e.g., M.S. Agranovich, A.S. Dynin, L.
Nirenberg, R.T. Seeley and A.I. Volpert), particularly that involving pseudo-
differential operators, Atiyah and Singer could generalize Hirzebruch’s proof
of the Hirzebruch-Riemann-Roch theorem of 1954 (see [207]) and discovered and
proved the desired index formula at Harvard in the Fall of 1962. Moreover, the Rie-
mannian Dirac operator played a major role in establishing the general case. The
details of this original proof involving cobordism actually first appeared in [328].
A K-theoretic embedding proof was given in [44], the first in a series of five papers.
xviii PREFACE

This proof was more direct and susceptible to generalizations (to G-equivariant
elliptic operators in [42] and families of elliptic operators in [47]).
The proof of the Index Theorem in [44] was inspired by Grothendieck’s
proof and thorough generalization of the Hirzebruch-Riemann-Roch Theorem, ex-
plained in [86]. We shall present the approach in detail in Chapters 10-12 of this
book. The invariance of the index under homotopy implies that the index (say,
the analytic index) of an elliptic operator is stable under rather dramatic, but con-
tinuous, changes of its principal symbol while maintaining ellipticity. Using this
fact, one finds (after considerable effort) that the analytical index of an elliptic
operator transforms predictably under various global operations such as embed-
ding and extension. Using K-theory and Bott periodicity, a topological invariant
(say, the topological index ) with the same transformation properties under these
global operations is constructed from the symbol of the elliptic operator. One then
verifies that a general index function having these properties is unique, subject to
normalization. To deduce the Atiyah–Singer Index Theorem (i.e., analytic index
= topological index ), it then suffices to check that the two indices are the same in
the trivial case where the base manifold is just a single point. A particularly nice
exposition of this approach for twisted Dirac operators over even-dimensional man-
ifolds (avoiding many complications of the general case) is found in E. Guentner’s
article [194] following an argument of P. Baum.
Not long after the K-theoretical embedding proof (and its variants), there
emerged a fundamentally different means of proving the Atiyah–Singer Index The-
orem, namely the heat kernel method. This is worked out here (see Chapter 17 in
the important case of the chiral half D+ of a twisted Dirac operator D. In the
index theory of closed manifolds, one usually studies the index of a chiral half D+
instead of the total Dirac operator D, since D is symmetric for compatible connec-
tions and then index D = 0.) The heat kernel method had its origins in the late
1960s (e.g., in [291], inspired by [302] of 1949) and was pioneered in the works
[331], [166], [33]. In the final analysis, it is debatable as to whether this method
is really much shorter or better. This depends on the background and taste of the
beholder. Geometers and analysts (as opposed to topologists) are likely to find the
heat kernel method appealing. The method not only applies to geometric operators
which are expressible in terms of twisted Dirac operators, but also largely for more
general elliptic pseudo-differential operators, as R.B. Melrose has done in [292].
Moreover, the heat method gives the index of a “geometric” elliptic differential op-
erator naturally as the integral of a characteristic form (a polynomial of curvature
forms) which is expressed solely in terms of the geometry of the operator itself (e.g.,
curvatures of metric tensors and connections). One does not destroy the geometry
of the operator by using ellipticity-preserving deformations. Rather, in the heat
kernel approach, the invariance of the index under changes in the geometry of the
operator is a consequence of the index formula itself more than a means of proof.
However, considerable analysis and effort are needed to obtain the heat kernel for
2
e−tD and to establish its asymptotic expansion as t → 0+ . Also, it can be argued
that in some respects the K-theoretical embedding/cobordism methods are more
forceful and direct. Moreover, in [273], we are cautioned that the index theorem
for families (in its strong form) generally involves torsion elements in K-theory
that are not detectable by cohomological means, and hence are not computable
in terms of local densities produced by heat asymptotics. Nevertheless, when this
PREFACE xix

difficulty does not arise, the K-theoretical expression for the topological index may
be less appealing than the integral of a characteristic form, particularly for those
who already understand and appreciate the geometrical formulation of characteris-
tic classes. More importantly, the heat kernel approach exhibits the index as just
one of a whole sequence of spectral invariants appearing as coefficients of terms
of the asymptotic expansion (as t → 0+ ) of the trace of the relevant heat kernel.
(On p. 118, we guide the reader to the literature about these particular spectral
invariants and their meaning in modern physics. The required mathematics for
that will be developed in Section 17.4.) All disputes aside, the student who learns
both approaches and formulations to the index formula will be more accomplished
(and probably a good deal older).
Further Reading. What the coverage of topics in this book is concerned,
we hope our table of contents needs no elaboration, except to say that space limi-
tations prevented the inclusion of some important topics (e.g., the index theorem
for families; index theory for manifolds with boundary, other than the Atiyah-
Patodi-Singer Theorem; L2 -index theory and coarse geometry of noncompact man-
ifolds; R. Nest’s and B. Tsygan’s algebraic and operator theoretic index theory
of [317, 318]; P. Kronheimer’s and T. Mrowka’s visionary work on knot ho-
mology groups from instantons; lists of all calculated spectral invariants; aspects
of analytic number theory). However, we now provide some guidance for further
study. A fairly complete exposition, by Atiyah himself, of the history of index
theory from 1963 to 1984 is found in Volume 3 of [26] and duplicated in Volume
4. Volumes 3, 4 and 5 contain many unsurpassed articles written by Atiyah and
collaborators on index theory and its applications to gauge theory. In the infor-
mative — and charming [448], S.-T. Yau collected The founders of index theory:
reminiscences of and about Sir Michael Atiyah, Raoul Bott, Friedrich Hirzebruch,
and I. M. Singer. N. Hitchin’s short text [216] on the 2004 Abel Prize Laureates
describes the index theorem, where it came from, its different manifestations and a
collection of applications. It indicates how one can use the theorem as a tool in a
concrete fashion without necessarily retreating into the details of the proof. We all
owe a debt of gratitude to H. Schröder for the definitive guide to the literature on
index theory (and its roots and offshoots) through 1994 in Chapter 5 of the excellent
book [167] of P.B. Gilkey. We have benefited greatly not only from this book,
but also from the marvelous work [273] by H.B. Lawson and M.L. Michelsohn.
In that book, there are proofs of index formulas in various contexts, and numerous
beautiful applications illustrating the power of Dirac operators, Clifford algebras
and spinors in the geometrical analysis of manifolds, immersions, vector fields, and
much more. The classical book [389] of P. Shanahan is also a masterful, elegant
exposition of not only the standard index theorem, but also the G-index theorem
and its numerous applications. A fundamental source on index theory for certain
open manifolds and manifolds with boundary is the authoritative book [292] of
R.B. Melrose. In [366], Th. Schick reviews coarse index theory, in particular,
for complete partitioned manifolds. It has been introduced by J. Roe and pro-
vides a theory to use tools from C ∗ -algebras to get information about the geometry
of non-compact manifolds via index theory of Dirac type operators. [204] of N.
Higson and J. Roe gives a well-written presentation of the underlying ideas of
analytic K-homology and develops some of its applications. For a concrete calcula-
tion see also the concise [309, Section 7.4.2] and the Notes (forthcoming) [205]. See
xx PREFACE

also [137, 138, 139] of J. Eichhorn for heat kernel asymptotics on non-compact
manifolds and [316] of B.-W. Schulze and collaborators for index theory on sin-
gular spaces. [457] of W. Zhang gives an excellent introduction to various aspects
of Atiyah-Singer index theory via the Bismut/Witten-type deformations of elliptic
operators. Very close to our own view upon index theory is the plea [157] of M.
Furuta for reconsidering the index theorem, with emphasis on the localization
theorem. In the case of boundary-value problems for Dirac operators, we put quite
some care in the writing of our [83] jointly with K.P. Wojciechowski. The recent
book [155] of D. Fursaev and D. Vassilevich contains a detailed description
of main spectral functions and methods of their calculation with emphasis on heat
kernel asymptotics and their application in various branches of modern physics.
Following up on the classic [213] of F. Hirzebruch and D. Zagier on interrela-
tions between the index theorem and elementary number theory, the comprehensive
[376] of S. Scott covers the theory of traces and determinants on Banach algebras
of operators on vector bundles over closed manifolds, with emphasis on various al-
gebras of pseudo-differential operators. He gives a series of calculations that give
the flavor of the subject in tractable cases, and relates these calculations to Pois-
son and Selberg trace formulas. There is an impressive nonstandard proof of the
local Atiyah-Singer index theorem, using resolvent expansions in place of the usual
heat equation techniques. A wealth of radically new ideas of (partly yet unproven)
geometric use of instantons are given in [269] of P.B. Kronheimer and T.S.
Mrowka. Very inspiring is [226] of E.P. Hsu on stochastic analysis on manifolds.
It gives a reformulation of the heat equation proof of the index theorem in terms
of Wiener process asymptotics. Basically, that is what we should have after A.
Einstein’s famous 1905-discovery of the basis of heat conduction in diffusion. The
details are interesting, though, in particular because they open a window to discrete
analysis. A taste of the recent revival of D-branes and other exotic instantons in
string theory can be gained from [164] of H. Ghorbani, D. Musso and A. Lerda.
Indications can be found in the review [361] of F. Sannino about, how strongly
coupled theories of gauge theoretic physics result in perceiving a composite universe
and other new physics awaiting to be discovered. In the mathematically rigorous
and richly illustrated [355], N. Reshetikhin explains why and how topological
invariants by necessity appear in various quantizations of gauge theories.
The Question of Originality: Seeking a Balance between Mathemat-
ical Heritage and Innovation. Parts I-III and the two appendices teach what
mathematicians today consider general knowledge about the index theorem as one
of the great achievements of 20th century mathematics. But, actually, there are
two novelties included which even not all experts may be aware of: The first novelty
appears when rounding up our comprehensive presentation of the topology of the
space of Fredholm operators: we do not halt with the Atiyah-Jänich Theorem and
the construction of the index bundle, but also confront the student with a thorough
presentation of the various definitions of determinant line bundles. This is to re-
mind the student that index theory is not a more or less closed collection of results
but a philosophy of regularization, of deformation invariance and of visionary cross
connections within mathematics and between its various branches.
A second novelty in the first three Parts is the emphasis on global constructions,
e.g., in introducing and using the concept of pseudo-differential operators.
PREFACE xxi

Apart from these two innovations, the student can feel protected in the first
three Parts against any originality.
Basically, Part IV follows the same line. Happily we could also avoid excessive
originality in the chapters dealing with instantons and the Donaldson-Kronheimer-
Seiberg-Witten results about the geometry of moduli spaces of connections. There
we also summarize, refer, define, explain great lines and details like in the first three
Parts, though emphasizing variational aspects based on [59].
However, the core of Part IV is different. It consists of an original, full, quite
lengthy (in parts almost unbearably meticulous) proof of the Local Index Theorem
for twisted Dirac operators in Chapter 17 and its applications to standard geometric
operators. That long Chapter is thought as a new contribution to the ongoing search
for a deeper understanding of the index theorem and the “best” approach to it.
Clearly, a student looking for the most general formulation of the index theo-
rem and a proof apt for wide generalizations should concentrate on our Part III,
the so-called Embedding (or K-theoretic) Proof. However, a student wanting to
trace the germs of index calculations back in the geometry of the considered stan-
dard operators (all arising from various decompositions of the algebra of exterior
differential forms) should consult Section 17.5 with a full proof of the Local Index
Formula for twisted Dirac operators on spin manifolds (all terms will be explained)
and Section 17.6, where we derive the Index Theorem for Standard Geometric Op-
erators. These geometric index theorems are by far less general than Part III’s
embedding proof, but they are more geometric, and we hold, also more geomet-
ric than the usual heat equation proofs of the index theorem. Not striving for
greatest generality, we obtain index formulas for the standard elliptic geometric
operators and their twists. The standard elliptic geometric operators include the
signature operator d + δ : (1 + ∗) Ω∗ (M ) = Ω+ (M ) → Ω− (M ) = (1 − ∗) Ω∗ (M ),
the Euler-Dirac
√ operator d + δ : Ωev (M ) → Ωodd (M ), and the Dolbeault-Dirac
¯ ¯∗ −,ev

operator 2 ∂ + ∂ : Ω (M ) → Ω+,odd (M ) (all symbols will be defined). The
index formula obtained for the above operators yields the Hirzebruch Signature
Theorem, the Chern-Gauss-Bonnet Theorem, and the Hirzebruch-Riemann-Roch
Theorem, respectively. While these operators generally are not globally twisted
Dirac operators, locally they are expressible in terms of chiral halves of twisted
Dirac operators. That applies also to the Yang-Mills operator. Thus, even if the
underlying Riemannian manifold M (assumed to be oriented and of even dimen-
sion) does not admit a spin structure, we may still use the Local Index Theorem
for twisted Dirac operators to compute the index density and hence the index of
these operators. While it is possible to carry this out separately for each of the
geometric operators, basically all of these theorems are consequences of one single
index theorem for generalized Dirac operators on Clifford module bundles (all to
be defined). Using the Local Index Theorem for twisted Dirac operators, we prove
this index theorem first (our Theorem 17.59), and then we apply it to obtain the
geometric index theorems, yielding the general Atiyah-Singer Index Theorem for
practically all geometrically defined operators.
Style and Notations. To present the rich world of index theory, we have
chosen two different styles. We write all definitions, theorems, and proofs as concise
as possible to free the reader from dispensable side information. Where possible,
we begin the introduction of a new concept with a simple but generic example or
a review of the local theory, immediately followed by the corresponding global or
xxii PREFACE

general concept. That is one half of the book, so to speak the odd numbered pages.
The other half of the book consists of exercises (often with extended hints) and
historical reviews, motivations, perspectives, examples. We wrote those sections in
a more open web-like style. Important definitions, notions, concepts are in bold
face. Background information is in small between the signs I and J. In remarks
and notes, leading terms are in italics.
The reader will notice our bias towards elder literature when more recent ref-
erences would not add substantially more value. This is due not so much to the
age of the authors (both born before the middle of the last century) but rather
to the common pride of mathematicians belonging to a community where biblio-
graphic impact factors and research indices should rather be calculated in citations
after some decades of years than in numbers of recently appeared, cited and soon
forgotten publications.
There is also a distinction, due to Harald Bohr and disseminated by Børge
Jessen, between expansive and consolidating periods of each individual science.
While physics and biology had consolidating periods in the first half of the last
century and suffer now of the rapid change of ever new single and dispersed results,
mathematics has had and still has good decades of consolidation and of long-time
validity of key results. To the present authors, there is no reason to hide our
indifference to changing fashions.
Acknowledgments and Dedication. Due to circumstance, the responsible
author had to finish the book alone, though on the basis of extended drafts for
all chapters which had been worked out jointly with David Bleecker. First of
all, he thanks his wife Sussi Booß-Bavnbek for her encouragement, support and
love. The responsible author acknowledges the hints and help he received from I.
Avramidi (Socorro), G. Chen (Hangzhou), G. Esposito (Napoli), K. Furutani
(Tokyo), V.L. Hansen (Copenhagen), B. Himpel (Aarhus), P. Kirk (Blooming-
ton), T. Kori (Tokyo), M. Lesch, B. Sauer and B. Vertman (Bonn), H.J.
Munkholm (Odense), L. Nicolaescu (Notre Dame), M. Pflaum (Boulder), S.
Scott (London), R.T. Seeley (Newton), G. Su (Tianjin), K. Uchiyama (Tokyo)
and C. Zhu (Tianjin) in the critical phase of the final re-shuffling and updating. In
particular, he thanks Roskilde University for an extraordinary leave for that work;
the Università degli Studi di Napoli Federico II and the Istituto Nazionale di Fisica
Nucleare (INFN, Sezione di Napoli) for their hospitality during his stay there; and
P.C. Anagnostopoulos (Carlisle), B.J. Bianchini (Somerville), H. Larsen
and J. Larsen (Roskilde) and M. Lesch and B. Sauer (Bonn) for generously
providing useful LaTeX macros, hacks and environments and substantial help with
the final lay-out and promotion. A.E. Olsen (Roskilde) checked the bibliographic
data and assembled the Index of Names/Authors and Notations.
Both authors agreed to dedicate this book to their teachers, to the memory of
S.-S. Chern (thesis adviser of DB) on the occasion of his centenary in October 2011
and to the memory of F. Hirzebruch (thesis adviser of BBB) who appreciated
the announced dedication intended for his 85th birthday in October 2012, which
he did not live to celebrate. I take the liberty to change the dedication. This book
is dedicated to David.

Bernhelm Booß-Bavnbek, Roskilde (Denmark), December 2012


Part I

Operators with Index and


Homotopy Theory

If we do not succeed in solving a


mathematical problem, the reason
frequently is our failure to recog-
nize the more general standpoint
from which the problem before us
appears only as a single link in a
chain of related problems.

David Hilbert, 1900

1
CHAPTER 1

Fredholm Operators

Synopsis. Hierarchy of Mathematical Objects. The Concept of a Bounded Fredholm


Operator in Hilbert Space. The Index. Forward and Backward Shift Operators. Alge-
braic Properties. Operators of Finite Rank. The Snake Lemma of Homological Algebra.
Product Formula. Operators of Finite Rank and the Fredholm Integral Equation. The
Spectra of Bounded Linear Operators (Terminology and Basic Properties).

1. Hierarchy of Mathematical Objects


“In the hierarchy of branches of mathematics, certain points are recog-
nizable where there is a definite transition from one level of abstraction
to a higher level. The first level of mathematical abstraction leads us to
the concept of the individual numbers, as indicated for example by the
Arabic numerals, without as yet any undetermined symbol representing
some unspecified number. This is the stage of elementary arithmetic; in
algebra we use undetermined literal symbols, but consider only individ-
ual specified combinations of these symbols. The next stage is that of
analysis, and its fundamental notion is that of the arbitrary dependence
of one number on another or of several others – the function. Still more
sophisticated is that branch of mathematics in which the elementary
concept is that of the transformation of one function into another, or,
as it is also known, the operator.”
I Thus N. Wiener characterized the hierarchy of mathematical objects [443, p.1].
Very roughly we can say: Classical questions of analysis are aimed mainly at investiga-
tions within the third or fourth level. This is true for real and complex analysis, as well as
the functional analysis of differential operators with its focus on existence and uniqueness
theorems, regularity of solutions, asymptotic or boundary behavior which are of partic-
ular interest here. Thereby research progresses naturally to operators of more complex
composition and greater generality without usually changing the concerns in principle; the
work remains directed mainly towards qualitative results.
In contrast it was topologists, as Michael Atiyah variously noted, who turned sys-
tematically towards quantitative questions in their topological investigations of algebraic
manifolds, their determination of quantitative measures of qualitative behavior, the def-
inition of global topological invariants, the computation of intersection numbers and di-
mensions. In this way, they again broadly broke through the rigid separation of the
“hierarchical levels” and specifically investigated relations between these levels, mainly
of the second and third level (algebraic surface – set of zeros of an algebraic function)
with the first, but also of the fourth level (Laplace operators on Riemannian manifolds,
Cauchy-Riemann operators, Hodge theory) with the first.
This last direction, starting with the work of Wilhelm Blaschke and William
V. D. Hodge, continuing with Kunihiko Kodaira, Shiing-Shen Chern and Donald
Spencer, with Henri Cartan and Jean Pierre Serre, with Friedrich Hirzebruch,
Michael Atiyah and others, can perhaps be best described with the key word differential

2
1.2. THE CONCEPT OF FREDHOLM OPERATOR 3

topology or analysis on manifolds. Both its relation with and distinction from analysis
proper is that (from [19, p.57])
“Roughly speaking we might say that the analysts were dealing with
complicated operators and simple spaces (or were only asking simple
questions), while the algebraic geometers and topologists were only
dealing with simple operators but were studying rather general man-
ifolds and asking more refined questions.”
We can read, e.g. in [92, pp.278-283], elaborated in [94] and the literature given
in [326], to what degree the contrast between quantitative and qualitative questions and
methods must be considered a driving force in the development of mathematics beyond
the realm sketched above.
Actually, in the 1920’s already, mathematicians such as Fritz Noether and Torsten
Carleman had developed the purely functional analytic concept of the index of an op-
erator in connection with integral equations, and had determined its essential properties.
But “although its (the theory of Fredholm operators) construction did not require the
development of significantly different means, it developed very slowly and required the
efforts of very many mathematicians” [177, p.185]. And although Soviet mathematicians
such as Ilja N. Vekua had hit upon the index of elliptic differential equations at the
beginning of the 1950’s, we find no reference to these applications in the quoted principal
work on Fredholm operators. In 1960 Israel M. Gelfand published a programmatic
article asking for a systematic study of elliptic differential equations from this quantita-
tive point of view. He took as a starting point the theory of Fredholm operators with its
theorem on the homotopy invariance of the index (see below). Only after the subsequent
work of Michail S. Agranovich, Alexander S. Dynin, Aisik I. Volpert, and finally
of Michael Atiyah, Raoul Bott, Klaus Jänich and Isadore M. Singer, did it be-
come clear that the theory of Fredholm operators is indeed fundamental for numerous
quantitative computations, and a genuine link connecting the higher “hierarchical levels”
with the lowest one, the numbers. J

2. The Concept of Fredholm Operator


Let H be a (separable) complex Hilbert space, and let B denote the Banach
algebra (e.g., [332, p.128], [358, p.228], [365, p.201]) of bounded linear operators
T : H → H with the operator norm
kT k := sup {kT uk : kuk ≤ 1} < ∞,
where k · k denotes the norm in H induced by the inner product h · , · i.
Definition 1.1. An operator T ∈ B is called a Fredholm operator, if
Ker T := {u ∈ H : T u = 0} and Coker T := H/Im(T )
are finite-dimensional.
This means that the homogeneous equation T u = 0 has only finitely many
linearly independent solutions, and to solve T u = v, it is sufficient that v satisfies a
finite number of linear conditions (e.g., see Exercise 2.1b below). We write T ∈ F
and define the index of T by
index T := dim Ker T − dim Coker T.
The codimension of Im(T ) = T (H) = {T u : u ∈ H} is dim Coker T .
4 1. FREDHOLM OPERATORS

Remark 1.2. a) We can analogously define Fredholm operators T : H → H 0 ,


where H and H 0 are Hilbert spaces, Banach spaces, or general topological vector
spaces. In this case, we use the notation B(H, H 0 ) and F(H, H 0 ), corresponding
to B and F above. However, in order to counteract a proliferation of notation and
symbols in this section, we will deal with a single Hilbert space H and its operators
as far as possible. The general case H 6= H 0 does not require new arguments at this
point. Later, however, we shall apply the theory of Fredholm operators to elliptic
differential and pseudo-differential operators where a strict distinction between H
and H 0 (namely, the compact embedding of the domain H into H 0 , see Chapter 9)
becomes decisive.
b) All results are valid for Banach spaces and a large part for Fréchet spaces also.
For details see for example [347, p.182-318]. We will not use any of these but will
be able to restrict ourselves entirely to the theory of Hilbert space whose treatment
is in parts far simpler.
c) Motivated by analysis on symmetric spaces – with transformation group G –
operators have been studied whose index is not a number but an element of the
representation ring R(G) generated by the characters of finite-dimensional repre-
sentations of G [44, p.519f] or, still more generally, is a distribution on G [23,
p.9-17]. We will not treat this largely analogous theory, nor the generalization of
the Fredholm theory to the discrete situation of von Neumann algebras as it has
been carried out – with real-valued index in [90, 1968/1969].
Exercise 1.3. Let L2 (Z+ ) denote the space of sequences c = (c0 , c1 , c2 , . . .) of
complex numbers with square-summable absolute values; i.e.,

X 2
|cn | < ∞.
n=0
2
L (Z+ ) is a Hilbert space (see Chapter A below). Show that the shift operators
shift+ : (c0 , c1 , c2 , . . .) 7→ (0, c0 , c1 , c2 , . . .) and shift− : (c0 , c1 , c2 , . . .) 7→ (c1 , c2 , c3 , . . .)
are Fredholm operators with index(shift+ ) = −1 and index(shift− ) = +1.
[Warning: Just as we can regard L2 (Z+ ) as the limit of the finite dimensional vector
spaces Cm (as m → ∞), we can approximate shift+ by endomorphisms of Cm given,
relative to the standard basis, by the m × m matrix
0 ··· 0
 
0 0
. 
 1 . . . . . . . . . .. 

 
 0 ... ... ... 0  .
 
 
 . .
 .. .. ... ... 0 

0 ··· 0 1 0
Note that the kernel and cokernel of this endomorphism are one-dimensional, whence
the index is zero. (See Exercise 1.4 below.).
We have yet another situation, when we consider the Hilbert space L2 (Z) of se-
quences c = (. . . , c−2 , c−1 , c0 , c1 , c2 , . . .) with
X∞  
2 2 2
|c0 | + |cn | + |c−n | < ∞.
n=1
The corresponding shift operators are now bijective, and hence have index zero.]
1.3. ALGEBRAIC PROPERTIES. OPERATORS OF FINITE RANK. THE SNAKE LEMMA 5

3. Algebraic Properties. Operators of Finite Rank. The Snake Lemma


Exercise 1.4. For finite-dimensional vector spaces the notion of Fredholm
operator is empty, since then every linear map is a Fredholm operator. Moreover,
the index no longer depends on the explicit form of the map, but only on the
dimensions of the vector spaces between which it operates. More precisely, show
that every linear map T : H → H 0 where H and H 0 are finite-dimensional vector
spaces has index given by
index T = dim H − dim H 0 .

[Hint: One first recalls the vector-space isomorphism H/ Ker(T ) −→ Im(T ) and
then (since H and H 0 are finite-dimensional) obtains the well-known identity from
linear algebra
dim H − dim Ker T = dim H 0 − dim Coker T.]
[Warning: If we let the dimensions of H and H 0 go to ∞, we obtain only the formula
index T = ∞ − ∞. Thus, we need an additional theory to give this difference a
particular value.]
Exercise 1.5. For two Fredholm operators F : H → H and G : H 0 → H 0
consider the direct sum,
F ⊕ G : H ⊕ H 0 → H ⊕ H 0.
Show that F ⊕ G is a Fredholm operator with index(F ⊕ G) = index F + index G.
[Hint: First verify that Ker(F ⊕ G) = Ker F ⊕ Ker G, and the corresponding fact
for Im(F ⊕ G). Then show that
(H ⊕ H 0 ) /(Im F ⊕ Im G) ∼
= (H/ Im F ) ⊕(H 0 / Im G) .]
[Warning: When H = H 0 , we can consider the sum (not direct) F + G : H → H,
but in general this is not a Fredholm operator; e.g., we could set G := −F .]
Exercise 1.6. Show by algebraic means, that the composition G ◦ F of two
Fredholm operators F : H → H 0 and G : H 0 → H 00 is again a Fredholm operator.
[Hint: Which inequality holds between
dim Ker G ◦ F and dim Ker F + dim Ker G
and between
dim Coker G ◦ F and dim Coker F + dim Coker G ?]
[Warning: Why are these in general not equalities? Nevertheless, the chain rule
index G ◦ F = index F + index G can be proved, since the inequalities cancel out
when we form the difference. Sometimes in mathematics a result of interest appears
as an offshoot of the proof of a rather dull general statement. Such is the case when
we obtain the chain rule for differentiable functions by showing that the composite
of differentiable maps is again differentiable. Here, however, the proof of the index
formula requires extra work, and we must either utilize the topological structure by
functional analytic means (see Exercise 2.3, p.14) or refine the algebraic arguments.
The latter will be done next.]
6 1. FREDHOLM OPERATORS

We recall a basic idea of diagram chasing:


Definition 1.7. Let H1 , H2 , . . . be a sequence (possibly finite) of vector spaces
and let Tk : Hk → Hk+1 be a linear map for each k = 1, 2, . . . . We call
T
1 2 T
3 T
H1 −→ H2 −→ H3 −→ ···
an exact sequence, if for all k Im Tk = Ker Tk+1 .
In particular, the following is a list of equivalences and implications (where 0
denotes the 0-dimensional vector space)
1T
0 −→ H1 −→ H2 exact ⇔ Ker T1 = 0 (T1 injective),
T1
H1 −→ H2 −→ 0 exact ⇔ Im T1 = H2 (T1 surjective),
T1 ∼
0 −→ H1 −→ H2 −→ 0 exact ⇔ T1 : H1 −→ H2 is an isomorphism,

(
T
1 2 T T1 : H1 −→ T1 (H1 ) and
0 −→ H1 −→ H2 −→ H3 −→ 0 exact ⇔ ∼
T2 : T1H
(H1 ) −→ H3 .
2

In the last case, H2 ∼


= H1 ⊕ H3 , but not canonically so.
Remark 1.8. By definition, the sequence
1 T
0 −→ Ker T1 ,→ H1 −→ H2 −→ Coker T1 −→ 0
of linear maps between vector spaces is exact. If H1 , H2 are Hilbert spaces and T1
bounded, the sequence becomes an exact sequence of bounded mappings between
Hilbert spaces if and only if T1 has closed range (see also the Hint to Exercise 2.1).
The following theorem is a key device in Homological Algebra for various kinds
of decomposition and additivity theorems, see also Remark 1.11.
Theorem 1.9 (Snake Lemma). Assume that the following diagram of vector
spaces and linear maps is commutative (i.e., jF = F 0 i and qF 0 = F 00 p) with exact
horizontal sequences and Fredholm operators for vertical maps.
i p
0 → H1 → H10 → H100 → 0
↓F ↓ F0 ↓ F 00
j q
0 → H2 → H20 → H200 → 0.
Then we have index F − index F 0 + index F 00 = 0.
Proof. We do the proof in two parts.
1. Here we show that the following sequence is exact:
(1.1) 0 −→ Ker F −→ Ker F 0 −→ Ker F 00 −→
−→ Coker F −→ Coker F 0 −→ Coker F 00 −→ 0.
For this, we first explain how the individual maps are defined. By the commutativity
of the diagram, the maps Ker F → Ker F 0 and Ker F 0 → Ker F 00 are given by i and
p; and the maps Coker F → Coker F 0 and Coker F 0 → Coker F 00 are induced by j
and q in the natural way (please check).
Also Ker F 00 → Coker F is well defined: Let u00 ∈ Ker F 00 ; i.e., u00 ∈ H 00
and F 00 u00 = 0. Since p is surjective we can choose u0 ∈ H10 with pu0 = u00 . Then
F 0 u0 ∈ Ker q, since qF 0 u0 = F 00 pu0 = F 00 u00 = 0. By exactness, we have Ker q = Im j
1.3. ALGEBRAIC PROPERTIES. OPERATORS OF FINITE RANK. THE SNAKE LEMMA 7

and a unique (by the injectivity of j) element u ∈ H2 with ju = F 0 u0 . We map


u00 ∈ Ker F 00 to the class of u in H2 / Im F = Coker F . It remains to show that we
get the same class for another choice of u0 . Thus, let u e0 ∈ H 0 be such that pe u0 = u00
0 0
(p is in general not injective; hence, possibly u e 6= u ). As above, we have a u e ∈ H2
with je u = F 0 u0 . We must now find u0 ∈ H1 with F u0 = u − u e. Then we are
done. For this, note that j(u − u e) = ju − je u = F 0 u0 − F 0 u
e0 = F 0 (u0 − u
e0 ). Since
0 0 00 0 0
pu = pe u = u , we have u − u e ∈ Ker p. By exactness, we have u0 ∈ H1 with
iu0 = u0 − u e0 , whence F 0 iu0 = F 0 (u0 − u
e0 ). The left side is jF u0 and the right side
is j(u − ue) from above. Hence, F u0 = u − u e by the injectivity of j, as desired. We
introduced the map Ker F 00 → Coker F in such great detail in order to demonstrate
what is typical for diagram chasing: it is straightforward, largely independent of
tricks and ideas, readily reproduced, and hence somewhat monotonous. Therefore
we will forgo showing the exactness of (1.1) except at Ker F 0 , and leave the rest as
ı̃ p̃
an exercise. The exactness of Ker F → Ker F 0 → Ker F 00 , where ı̃ and p̃ are the
restrictions of i and p, means Im ı̃ = Ker p̃. Thus, we have two inclusions to show:
⊆: This is clear, since p ◦ i = 0 implies p̃ ◦ ı̃ = 0.
⊇: If u0 ∈ Ker p̃, then u0 ∈ Ker F 0 and pu0 = 0, and
by the exactness of H1 → H10 → H100 , we have u ∈ H1 with iu = u0 . It remains to
show that F u = 0, but this is clear, since jF u = F 0 iu = F 0 u0 = 0 and j is injective.
Note. Before we go to part 2 of the proof, we pause for a moment: It is interesting
that for H 0 = H ⊕ H 00 , i = j the inclusion, and p = q the projection, we recapture
the addition formula of Exercise 1.5. In this case, the exact sequence (1.1) then
breaks, as shown there, into two parts
0 → Ker F → Ker F 0 → Ker F 00 → 0 and
0 → Coker F → Coker F 0 → Coker F 00 → 0.
In the general case, however, we no longer have
dim Ker F 0 = dim Ker F + dim Ker F 00 and
dim Coker F 0 = dim Coker F + dim Coker F 00 ,
but instead we must consider the interaction (Ker F 00 → Coker F ). From a topo-
logical standpoint, the concept of the index of a Fredholm operator is a special case
of the general concept of the Euler characteristic χ(C) of a complex
Tk+1 k T Tk−1 Tk−2 Tk−3
C: −→ Ck −→ Ck−1 −→ Ck−2 −→ Ck−3 −→ · · ·
of vector spaces and linear maps (with Tk ◦ Tk+1 = 0) with finite Betti numbers
 
Ker Tk
bk := dim .
Im Tk+1
Here Ker Tk / Im Tk+1 is called the k-th homology space Hk (C) . Assuming that
all these numbers are finite, as well as the number of nonzero Betti numbers, we
define X
χ(C) := (−1)k bk
k
whence index F = χ(C) for
F
C : 0 −→ 0 −→ H −→ H −→ 0 −→ 0,
8 1. FREDHOLM OPERATORS

where C2 = C1 = H and Ci = 0 otherwise; thus, H2 (C) = Ker F and H1 (C) =


Coker F . This is the reason why we can follow in the proof of Theorem 1.9 the
well-known topological arguments (see [185, p.100f] or [140, p.52f] ) yielding the
addition (or pasting) theorem χ(C) − χ(C 0 ) + χ(C 0 /C) = 0. See also Section 13.4.
In particular, (1.1) is only a special case of the long exact homology sequence
→ Hk+1 (C 0 /C) → Hk (C) → Hk (C 0 )
(1.2) → Hk (C 0 /C) → Hk−1 (C) → Hk−1 (C 0 ) →,
[185, p.57f] or [140, p.125-128]. This more general description would have the
advantage that we only need to prove exactness of (1.2) at three adjacent places
with an argument independent of k rather than at six places as in our simplified
approach where we restricted ourselves to complexes of length two. But back to
our proof:
2. For each exact sequence
(1.3) 0 → A1 → A2 → · · · → Ar → 0,
of finite-dimensional vector spaces, we wish to derive the formula
r
X
(−1)k dim Ak = 0.
k=1

Notice first that for r sufficiently large (r > 3), the formula for the alternating sum
for (1.3) follows, once we know the formula holds for the exact (prove!) sequences
0 −→ A1 −→ A2 −→ Im(A1 → A2 ) −→ 0
and
0 −→ Im(A2 → A3 ) −→ A3 −→ · · · −→ Ar −→ 0.
Since these sequences have length less than r, the formula is proved by induction,
if we verify it for r = 1, 2, 3.
r = 1 : trivial, since then A1 ∼
= 0.
r = 2 : also clear, since then A1 ∼
= A2 .
r = 3 : clear, since 0 → A1 → A2 → A3 → 0 implies A3 ∼
= A2 /A1 ,
(1.4) whence dim A3 = dim A2 − dim A1 .
Thus, part 2 is finished and combining it with part 1, the Snake Lemma is proved.
Actually, we have proven much more, namely, whenever two of the three maps F,
F 0 , F 00 have a finite index, the third has finite index given by the snake formula. 
Exercise 1.10. Combine Theorem 1.9 and Exercise 1.6, to show that
(1.5) index G ◦ F = index F + index G.
[Hint: Consider the diagram
i p
0 −−−−→ H −−−−→ H ⊕ H 0 −−−−→ H 0 −−−−→ 0
  
  
yF yG◦F ⊕Id yG
j q
0 −−−−→ H 0 −−−−→ H 00 ⊕ H 0 −−−−→ H 00 −−−−→ 0,
where iu := (u, F u), jv := (Gv, v), p(u, v) := F u − v, and q(w, v) := w − Gv.]
1.4. OPERATORS OF FINITE RANK AND THE FREDHOLM INTEGRAL EQUATION 9

Remark 1.11. Alternative proofs of the product formula (1.5) can be found
in many places. They may appear shorter. Arguing via the Snake Lemma is more
lengthy, but it puts the product formula in the correct format of a topological
composition or gluing formula
(1.6) τ (Φ1 ∪ Φ2 ) = τ (Φ1 ) ∗ τ (Φ2 ) ∗ ε(Φ1 ∩ Φ2 ),
where we have in (1.5) the vanishing of the typical third term on the right side, the
error term ε(Φ1 ∩ Φ2 ), for τ := index; Φ1 , Φ2 ∈ F; ∪ := ◦; and ∗ := +.
A stunning impression of the intricacies of simple looking product formulas
may be gained by checking the proof of the corresponding product formula for the
index of closed (not necessarily bounded) densely defined Fredholm operators, see
Theorem 2.45, p.43f.

4. Operators of Finite Rank and the Fredholm Integral Equation


Exercise 1.12. Show that for any operator K : H → H of finite rank (i.e.,
dim K(H) < ∞), the sum Id +K is a Fredholm operator and index(Id +K) = 0.
Here, Id : H → H is the identity.
[Hint: Set h := Im K and recall Theorem 1.9 for the diagram

0 /h i /H p
/ H/h /0

(Id +K)h Id +K (Id +K)H/h


  
0 /h j
/H q
/ H/h / 0.

To see that the vertical maps are well defined, we only need (Id +K)(h) ⊆ h. The
commutativity of the diagram and the exactness of the rows are clear. Since one
can show Ker(Id +K) ⊆ h and dim Coker(Id +K) ≤ dim h, the Snake Formula gives
us the result once we show
(1.7) index(Id +K)h = 0
and
(1.8) index(Id +K)H/h = 0.
But (1.7) is clear from Exercise 1.4 and (1.8) is trivial because (Id +K)H/h = IdH/h .
Note that one could deduce that Id +K has finite index by using the observation
at the end of the proof of Theorem 1.9.]
Remark 1.13. One may be bothered by the way in which the proposed solution
produces the result so directly from the Snake Formula by means of a trick. As a
matter of fact, index(Id +K) can be computed in a pedestrian fashion by reduction
to a system of n linear equations with n unknowns where n := dim h. To do this
one verifies that every operator K of finite rank has the form
Xn
Ku = hu, ui i vi
i=1
with fixed u1 , ..., un , v1 , ..., vn ∈ H (note that every continuous linear functional is
of the form h·, u0 i). Whether this direct approach, as detailed for example in [365,
Theorem 4.9](see also [332, p.110f]), is in fact more transparent than the device
used with the Snake Formula, depends a little on the perspective. While in the
10 1. FREDHOLM OPERATORS

first approach the key point (namely the use of Exercise 1.4 for equation (1.7)) is
singled out and separated clearly in the remaining formal argument, we find in the
second more constructive approach rather a fusion of the nucleus with its packaging.
However, the use of the fairly nontrivial Riesz-Fischer Lemma is unnecessary in the
case where K is given in the desired explicit form, as in the following example.
Exercise 1.14. Consider the Fredholm integral equation of the second
kind Z b
u(x) + G(x, y)u(y) dy = h(x)
a
with degenerate (product-) weight function (or integral kernel )
n
X
G(x, y) = fi (x)gi (y)
i=1

with fixed a < b real and fi , gi square integrable on [a, b]. Prove the Fredholm
alternative: Either there is a unique solution u ∈ L2 [a, b] for every given right side
h ∈ L2 [a, b], or the homogeneous equation (h = 0) has a solution which does not
vanish identically. Moreover, the number of linearly independent solutions of the
homogeneous equation equals the number of linear conditions one needs to impose
on h in order that the inhomogeneous equation be solvable.
[Hint:
P Consider the operator Id +K on the Hilbert space L2 [a, b], where Ku =
hu, gi i fi , and apply Exercise 1.12. For the interpretation of the dimension of the
cokernel, see Exercise 2.1b below.]

5. The Spectra of Bounded Linear Operators: Basic Concepts


We close this chapter with a concept that belongs to the border region between
algebraic and analytic notions, the spectrum of a bounded linear operator. For
the corresponding definitions and elementary properties in the more general (and,
actually, different) case of not necessarily bounded linear operators we refer to
Definition 2.59 and Exercise 2.60, p.50.
Remark 1.15. The following comments correspond to the sets of Table 1.1.
For wanted arguments we refer to [332, Section 4.1]:
1. Res(T ) is open in C. We can argue as in Exercise 3.6, p.65 where one has to
prove that the group of units in any Banach algebra is open.
2. Spec(T ) is closed and bounded (i.e., compact) in C. Actually, one has Spec(T ) ⊂
{λ : |λ| ≤ kT k}.
3. Res(T ) ⊆ Fred(T ), and Fred(T ) is the union of at most a countable number of
open, connected components.
4. Spece (T ) ⊆ Spec(T ) , and Spece (T ) = Spec(π(T )), where π : B → B/K is the
projection.
5. Specp (T ) consists of isolated points of Spec(T ). Specp (T ) contains the limit
points of Spec(T ) which are Fredholm points.
6. Specc (T ) ⊆ Spece (T ).
7. Spec(T ) = Specp (T ) ∪ Specc (T ) ∪ Specr (T ), a union of disjoint sets.
1.5. THE SPECTRA OF BOUNDED LINEAR OPERATORS: BASIC CONCEPTS 11

Table 1.1. The spectra of bounded linear operators T : H → H

Symbol Name Definition


z ∈ C : (T − z Id)−1 ∈ B

1. Res(T ) resolvent set
2. Spec(T ) spectrum C\ Res(T )
3. Fred(T ) Fredholm points {z ∈ C : T − z Id ∈ F}
4. Spece (T ) essential spectrum C \ Fred(T )
point spectrum
5. Specp (T ) {z ∈ C : Ker(T − z Id) 6= {0}}
(eigenvalues)   
  Ker(T − z Id) = {0} , 
6. Specc (T ) continuous spectrum z∈C: Im(T − z Id) H, and
( Im(T − z Id) = H
  
( ))
Ker(T − z Id) = {0},
7. Specr (T ) residual spectrum z∈C:
codim Im(T − z Id) > 0

Example 1.16. (a) T := shift+ (see [234, 1970/1982, 5.3]) =⇒ Specp (T ) = ∅,


Specr (T ) = {z ∈ C : |z| < 1} and Specc (T ) = Spece (T ) = {z ∈ C : |z| = 1} .
(b) T compact and self-adjoint (see [234, 1970/1982, 6.2]) =⇒ Spec(T ) ⊆ R,
Specp (T ) = {λn : n = 1, 2, . . .}, Specr (T ) = {0} and Spece (T ) = {0}.
(c) T = Fourier transformation on L2 (R) (see [134, 1972, p.98]) =⇒ Spec(T ) =
Spece (T ) = Specp (T ) = {1, i, −1, −i}.
CHAPTER 2

Analytic Methods. Compact Operators

Synopsis. Adjoint and Self-Adjoint Operators — Recalling Fischer-Riesz. Dual


Characterization of Fredholm Operators. Compact Operators: Spectral Decomposition,
Why Compact Operators also are Called Completely Continuous, K as Two-Sided Ideal,
Closure of Finite-Rank Operators, and Invariant under ∗ . Classical Integral Operators.
Fredholm Alternative and Riesz Lemma. Sturm-Liouville Boundary Value Problems. Un-
bounded Operators: Comprehensive Study of Linear First Order Differential Operators
Over S 1 : Sobolev Space, Dirac Distribution, Normalized Integration Operator as Para-
metrix, The Index Theorem on the Circle for Systems. Closed Operators, Closed Ex-
tensions, Closed (not necessarily bounded) Fredholm Operators, Composition Rule, Sym-
metric and Self–Adjoint Operators, Formally Self-Adjoint and Essentially Self-Adjoint.
Spectral Theory. Metrics on the Space of Closed Operators. Trace Class and Hilbert-
Schmidt Operators.

1. Analytic Methods. The Adjoint Operator


I With Exercises 1.12 and 1.14, we have reached the limits of our so far purely al-
gebraic reasoning where we could reduce everything to the elementary theory of solutions
of n linear equations in n unknowns. In fact, the limit process n → ∞ marks the emer-
gence of functional analysis which went beyond the methods of linear algebra while being
motivated by its questions and results. This occurred mainly in the study of integral
equations.
In 1927, Ernst Hellinger and Otto Toeplitz stressed in their article Integral
equations and equations in infinitely many unknowns in the Enzyklopädie der mathema-
tischen Wissenschaften that “the essence of the theory of integral equations rests in the
analogy with analytic geometry and more generally in the passage from facts of alge-
bra to facts of analysis” [201, p.1343]. They showed in a concise historical survey how
the awareness of these connections progressed in the centuries since Daniel Bernoulli
investigated the oscillating string as a limit case of a system of n mass points:
“Orsus itaque sum has meditationes a corporibus duobus filo flexili in
data distantia cohaerentibus; postea tria consideravi moxque quatuor, et
tandem numerum eorum distantiasque qualescunque; cumque numerum
corporum infinitum facerem, vidi demum naturam oscillantis catenae
sive aequalis sive inaequalis crassitiei sed ubique perfecte flexilis.”1
The passage to the limit means for the Fredholm integral equation of Exercise 1.14
that more general nondegenerate weight functions G(x, y) are allowed which then can be
approximated by degenerate weights (e.g., polynomials). Such weights play a prominent

1
Petropol. Comm. 6 (1732/33, ed. 1738), 108-122. Our translation: “In these considerations
I started with two bodies at a fixed distance and connected by an elastic string; next I considered
three then four and finally an arbitrary number with arbitrary distances between them; but only
when I made the number of bodies infinite did I fully comprehend the nature of an oscillating
elastic chain of equal or varying thickness.”

12
2.1. ANALYTIC METHODS. THE ADJOINT OPERATOR 13

role in applications, especially when dealing with differential equations with boundary
conditions, as we will see below. J

For the theory of Fredholm operators developed here we must analogously


abandon the notion of an operator of finite rank and generalize it (to compact
operator, see below) whereby topological, i.e. continuity considerations, become
essential in connection with the limit process. This will bring out the full force of
the concept of Fredholm operator. We now turn to this topic.
Closed Image of Fredholm Operators. To begin with, the reader should
derive the topological closedness of the image of a Fredholm operator from the purely
algebraic property of finite codimension of the image.
Exercise 2.1. For F : H → H a Fredholm operator, prove:
a)Im F is closed.
b) There is an explicit criterion for deciding when an element of H lies in Im F .
Namely, let n = dim Coker F ; then there are u1 , ..., un ∈ H, such that for all w ∈ H,
we have
w ∈ Im F ⇔ hw, u1 i = · · · = hw, un i = 0.
[Hint for a): As a vector subspace of H, naturally Im F is closed under addition
and multiplication by complex numbers. However, here we are interested in topo-
logical closure, namely that in passing to limit points we do not leave Im F . This
has far-reaching consequences, since only closed subspaces inherit the completeness
property of the ambient Hilbert space, and hence have orthonormal bases (= com-
plete orthonormal vector systems, see e.g., [332, p.83] or [365, p.31 and Lemma
11.9]). The trivial fact that finite-dimensional subspaces are closed can be exploited:
Since Coker F = H/ Im F has finite dimension, we can find v1 , ..., vn ∈ H whose
classes in H/ Im F form a basis. The linear span h of v1 , ..., vn is then an algebraic
complement of Im F in H. Consider the map Φ : H⊕h → H with Φ(u, v) := F u+v.
Since Φ is linear, surjective, and (by the boundedness of F ) continuous, we have
that Φ is open (according to the open mapping principle; e.g., see [332, p.53]). It
follows that H \ F (H) = Φ(H ⊕ h \ H ⊕ {0}) is open.
[Hint for b): As a closed subspace of H, Im F is itself a Hilbert space possessing
a countable orthonormal basis w1 , w2 , . . .. Now set ui := vi − PP vi , where vi are

as above (i = 1, ..., n) and P : H → Im F is the projection P u := j=1 hu, wj i wj .
Then {u1 , ..., un } forms a basis for (Im F )⊥ , the orthogonal complement of Im F in
H.
For aesthetic reasons, one can orthonormalize u1 , ..., un by the Gram-Schmidt process
(i.e., without loss of generality, assume they are orthonormal). Then u1 , ..., un ,
w1 , w2 , w3 , . . . is a countable orthonormal basis for H.]
Remark 2.2. Be aware that in spite of the strength of the open-mapping
argument, it can not be applied to show that any subspace W of finite codimension
dim H/W < ∞ is closed. Of course, one could once again construct a bounded
surjective operator Φ : H ⊕ W → H, say by Φ(u, v) := u + v. But, in general,
H \ W can not be obtained as the image of an open subset of H ⊕ W by applying
Φ. Certainly H \ W 6= Φ(H ⊕ W \ H ⊕ {0}). Actually, the kernel Ker(f ) of any
unbounded linear functional provides a counterexample. It is a space of codimension
1, but it is not closed since closed Ker(f ) would imply continuity of f in 0 and hence
everywhere.
14 2. ANALYTIC METHODS. COMPACT OPERATORS

Exercise 2.3. Once more, prove the chain rule index G◦F = index F +index G
for Fredholm operators F : H → H 0 and G : H 0 → H 00 .
[Hint: In place of the purely algebraic argument in Exercise 1.10, use Exercise 2.1a
to first prove that the images are closed, and then use the technique of orthogonal
complements in Exercise 2.1b.]

How trivial or nontrivial is it to prove that operators have closed ranges? For
operators with finite rank and for surjective operators it is trivial, and for Fredholm
operators it was proved in Exercise 2.1a. Is it perhaps true that all bounded linear
operators have closed images? As the following counterexample explicitly shows,
the answer is no. Moreover, we will see below that all compact operators with
infinite-dimensional image are counterexamples.

Exercise 2.4. For a Hilbert space H with orthonormal system e1 , e2 , e3 , . . . ,


consider the contraction operator

X 1
Au := hu, ej i ej .
j=1
j

Show that Im A is not closed.


[Hint: Clearly A is linear and bounded (kAk =?), and furthermore, we have the
criterion for Im(A)

X ∞
X 2
v ∈ Im A ⇔ j hv, ej i ej ∈ H ⇔ j 2 |hv, ej i| < ∞.
j=1 j=1

It follows that for


∞ ∞
X 1 X 1
v0 := √ ej and vn := √ ej ; n = 1, 2, . . . .;
j=1
j j j=1
j jj 1/n

1/j a converges for a > 1


P
we get v0 ∈ H\ Im A and vn ∈ Im A. (The old trick:
(e.g., for a = 1 + 1/n), but diverges for a = 1.) To finally prove that the sequence
actually converges to v0 , observe that
∞ 2
2
X 1 j 1/n − 1
kv0 − vn k = .
j=1
j 2 j j 2/n

It is clear, that for each j


j 1/n − 1
→ 0 as n → ∞,
j 1/n
2
and then in particular hv0 − vn , ej i → 0. However, to show that kv0 − vn k con-
verges to 0Pas n → ∞, one must estimate more precisely. To do this, we exploit the

fact that j>j0 j −3 can be made smaller than any ε > 0 for j0 sufficiently large,
while on the other hand, choosing n sufficiently large (so large that (1 + ε)n ≥ j0 ),
we have
j 1/n − 1
< ε for j ≤ j0 .]
j 1/n
2.1. ANALYTIC METHODS. THE ADJOINT OPERATOR 15

The Adjoint Operator. We will draw further conclusions from the closure
of the image of a Fredholm operator and to do this we introduce adjoint operators.
The purpose is to eliminate the asymmetry between kernel and cokernel or, in other
words, between the theory of the homogeneous equation (questions of uniqueness
of solutions) and the theory of the inhomogeneous equation (questions of existence
of solutions). This is achieved by representing the cokernel of an operator as the
kernel of a suitable adjoint operator.
Projective geometry deals with a comparable problem via duality: one thinks of
space on the one hand as consisting of points, on the other as consisting of planes,
and depending on the point of view, a straight line is the join of two points or
the intersection of two planes. Analytic geometry passes from a matrix (aij ) to its
transpose (aji ) or (aji ) in the complex case to technically deal with dual statements.
We can do the same successfully for operators (= infinite matrices). The basic tool
is the following Representation Theorem with nice proofs in [332, Proposition 3.1.9]
or [365, Theorem 2.1]. Originally, it was proven independently by Frigyes Riesz
and Ernst Sigismund Fischer only for H := L2 ([a, b]).
Theorem 2.5 (E. Fischer, F. Riesz, 1907). Let H be a separable complex
Hilbert space. To each continuous linear mapping ϕ : H → C (called functional)
there exists a unique element u ∈ H such that ϕv = hv, ui for all v ∈ H.
Exercise 2.6. Show that on the space B(H) of bounded linear operators of a
Hilbert space H, there is a natural isometric (anti-linear) involution
∗ : B(H) → B(H)
which assigns to each T ∈ B(H) the adjoint operator T ∗ ∈ B(H) such that
hu, T ∗ vi = hT u, vi for all u, v ∈ H.
[Hint: It is clear that T ∗ v is well defined for each v ∈ H, since u 7→ hT u, vi is
a continuous linear functional on H; and so, by the preceding theorem, you can
express the functional through a unique element of H, which you may denote by
T ∗ v. The linearity of T ∗ is clear by construction. While proving the continuity
(i.e., boundedness) of T ∗ , show more precisely that kT ∗ k = kT k (i.e., that T 7→ T ∗
is an isometry).]
Just as easily, we have the involution property T ∗∗ = T , the composition rule
(T ◦ R)∗ = R∗ ◦ T ∗ , and conjugate-linearity (aT + bR)∗ = āT ∗ + b̄R∗ , where the bars
denote complex conjugation. Details can be found for example in [332, Theorem
3.2.3], [365, Sections 3.2 and 11.2]. Observe that for T ∈ B(H, H 0 ), where H and
H 0 may differ, the adjoint operator T ∗ is in B(H 0 , H).
Theorem 2.7. For F ∈ F, the u1 , ..., un in Exercise 2.1b form a basis of
Ker F ∗ , whence
Im F = (Ker F ∗ )⊥ and Coker(F ) = Ker(F ∗ ) .
Proof. First we note that u ∈ (Im F )⊥ exactly when
0 = hu, wi = hu, F vi = hF ∗ u, vi ,
for all w ∈ Im(F ) (i.e., for all v ∈ H); thus, (Im F )⊥ = Ker F ∗ . By again taking
orthogonal complements, we have (Ker F ∗ )⊥ = (Im F )⊥⊥ = Im F , since Im F is
closed by Exercise 2.1a. 
16 2. ANALYTIC METHODS. COMPACT OPERATORS

Observe that the above argument remains valid for any bounded linear operator
with closed range Such operators are also called normally-solvable operators. In
this case, we have the criterion that the equation F v = w is solvable exactly when
w⊥ Ker F ∗ . This is the basic
Lemma 2.8 (Polar Lemma). Let H, H 0 be Hilbert spaces and T ∈ B(H, H 0 ).
Then (Im T )⊥ = Ker T ∗ . Moreover, if Im T is closed, we have Im T = (Ker T ∗ )⊥ .
In the language of categories and functors (see e.g. [89, p.176]) this can be
reformulated in the following, a bit exaggerated way:
Theorem 2.9. The functor H 7→ H, H 0 7→ H 0 , B(H, H 0 ) 3 T 7→ T ∗ ∈
B(H 0 , H) is a contravariant functor on the category of (separable) Hilbert spaces
and bounded operators. It preserves the norm and is exact; i.e.,
T S R R∗ S∗ T∗
H −→ H 0 −→ H 00 −→ exact =⇒−→ H 00 −→ H 0 −→ H exact.
Proof. We show the exactness only at H 0 , i.e., Im S ∗ = Ker T ∗ . Since ST = 0,
we have T ∗ S ∗ = 0, hence Im S ∗ is contained in Ker T ∗ .
To show the opposite inclusion, we notice that Im T = Ker S is closed in H 0
and Im S = Ker R is closed in H 00 . So we have a decomposition of
(2.1) H 0 = Ker T ∗ ⊕ Ker S = (Im T )⊥ ⊕ Im T
and
(2.2) H 00 = Im S ⊕ (Im S)⊥
into pairs of mutually orthogonal closed subspaces. Notice also that
(2.3) S|Ker T ∗ : Ker T ∗ −→ Im S
is bounded, injective and surjective, hence its inverse is also bounded (though not
necessarily a Hilbert space isomorphism, i.e. not necessarily unitary).
Now let y ∈ H 0 with y ∈ Ker T ∗ , i.e., hy, y 0 i = 0 for all y 0 ∈ T (H). We consider
the mapping
[y] : H 00 −→ C given by y 0 + z 0 7→ hy 0 , yi,
where the splitting on the left side is according to (2.2) and the inner product on
the right side is taken in the Hilbert space H 0 . By construction, the mapping [y] is
linear and vanishes on the second factor of H 00 . On the first factor it is continuous
because of the homeomorphism of (2.3). Hence the functional [y] can be represented
by an element of the Hilbert space H 00 which we also will denote by [y]. So we have
(2.4) hy 0 + y 00 , S ∗ ([y])i = hSy 0 + Sy 00 , [y]i = hSy 0 , [y]i = hy 0 , yi = hy 0 + y 00 , yi
for all elements y 0 + y 00 ∈ H 0 with y 0 ∈ (Im T )⊥ and y 00 ∈ Im T according to the
decomposition (2.1). Note that the second and the third inner product in (2.4)
are taken in H 00 and the other inner products in H 0 . Equation (2.4) shows that
y = S ∗ ([y]). Thus Ker T ∗ ⊆ Im S ∗ , and Ker T ∗ = Im S ∗ , as desired. 
Theorem 2.10. A bounded linear operator F is a Fredholm operator, precisely
when Ker F and Ker F ∗ are finite-dimensional and Im F is closed. In this case,
index F = dim Ker F − dim Ker F ∗ .
Thus, in particular, index F = 0 in case F is self-adjoint (i.e., F ∗ = F ).
Proof. Use Theorem 2.7 and Exercise 2.1a. 
2.1. ANALYTIC METHODS. THE ADJOINT OPERATOR 17

Remark 2.11. a) Note that Ker F ∗ F = Ker F : “⊇” is clear; for “⊆”, take
u ∈ Ker F ∗ F , and then
hF ∗ F u, ui = hF u, F ui = 0,
and u ∈ Ker F . If Im F is closed (e.g., if F ∈ F), then we also have
Im F ∗ F = Im F ∗ .
Here “⊆” is clear. To prove “⊇”, consider F ∗ v for v ∈ H, and decompose v
into orthogonal components v = v 0 + v 00 with v 0 ∈ Im F and v 00 ∈ Ker F ∗ ; then
F ∗ v = F ∗ v 0 . In this way, we then have represented the kernel and cokernel of any
Fredholm operator F as the kernels of the self-adjoint operators F ∗ F and F F ∗ ,
respectively.
b) The contraction operator A of Exercise 2.4 provides an example of a bounded,
self-adjoint operator with Ker A (= Ker A∗ ) = {0} which is not a Fredholm opera-
tor.
Corollary 2.12. Let H, H 0 be Hilbert spaces and F : H → H 0 a bounded
Fredholm operator. Then F ∗ : H 0 → H is a Fredholm operator and we have
Ker F ∗ ∼
= Coker F, Coker F ∗ ∼
= Ker F, and index F ∗ = − index F.
Proof. Consider the exact sequence
F
0 −→ Ker F −→ H −→ H 0 −→ Coker F −→ 0.
Then by Theorem 2.9, the sequence
F∗
0 −→ Coker F −→ H 0 −→ H −→ Ker F −→ 0
is also exact, and the assertion follows. 

Positive operators. Another concept based on the scalar product is the no-
tion of positive operators.
Definition 2.13. If C ∈ B := B(H) with hCx, xi ≥ 0 for all x ∈ H, then C
is called a positive operator. We denote the (convex) set of such operators by
B+ .
Proposition 2.14. Any C ∈ B + is self-adjoint. Moreover, for any n ∈ N,
C ∈ B+ .
n

Proof. For all x ∈ H,


hCx, xi = hx, Cxi = hx, Cxi = hC ∗ x, xi =⇒ h(C − C ∗ ) x, xi = 0.
For D := C − C ∗ and all x, y ∈ H, we then have

hD(x + y) ,(x + y)i = 0 =⇒ hDx, yi + hDy, xi = 0
=⇒ hDx, yi = 0.
hD(x + iy) ,(x + iy)i = 0 =⇒ −i hDx, yi + i hDy, xi = 0
So D = 0 and C = C ∗ . We now have C 2 x, x = hCx, Cxi ≥ 0 and for n ≥ 3,
hC n x, xi = C n−2 (Cx) , Cx ,
whence the positivity of C n follows by induction. 
18 2. ANALYTIC METHODS. COMPACT OPERATORS

For real Hilbert spaces H, hCx, xi ≥ 0 does not imply C = C ∗ (e.g., hAx, xi =
0 ≥ 0 for any skew-symmetric A), but we assumed that H is complex.
Below, in Section 2.7 on trace class and Hilbert-Schmidt operators, we shall
prove the fundamental Square Root Lemma (Theorem 2.66, p.54ff) for all operators
belonging to B + by means of completely elementary arguments.

2. Compact Operators
So far we found that the space F of Fredholm operators is closed under compo-
sition and passage to adjoints and particularly that all operators of the form Id +T
belong to F when T is an operator of finite rank. We will increase this supply of
examples, in passing to compact operators by taking limits. This, however, does
not lead to Fredholm operators of nonzero index.
We begin with an exercise which emphasizes a simple topological property of
operators of finite rank, more generally characterizes the finite-dimensional sub-
spaces which are fundamental for the index concept, and prepares the introduction
of compact operators.

Exercise 2.15. a) Every operator with finite rank maps the unit ball (or any
bounded subset) of H to a relatively compact set.
b) If H is finite-dimensional, then the closed unit ball BH := {u ∈ H : kuk ≤ 1}
is compact.
c) If H is infinite-dimensional, then BH is noncompact.
[Hint for a) and b): Recall the theorem of Bernhard Bolzano and Karl Weier-
strass that says that every closed bounded subset of Rn (or Cn ) is compact.
For c): Every orthonormal system e1 , e2 , ... in H is a sequence in BH without a
convergent subsequence. For instance, how large is kei − ek k for i 6= k?]

Definition 2.16. We denote by K (or K(H)) the set of linear operators from
H to H which map the open unit ball (or more generally, each bounded subset of
H) to a relatively compact subset of H. Such operators are called compact (or
sometimes completely continuous) operators.

By Exercise 2.15a, the compact operators form the largest class of operators
that behave (in this respect) like finite rank operators, i.e., like the operators of
linear algebra which are defined via matrices.
The following theorem supports establishing the relative compactness of subsets
in function or mapping spaces, for instance in our proof of Rellich’s compact
embedding of Sobolev spaces, Theorem 7.15, p.201f. We state it and its common
reformulation without proof. For a clear (but a bit lengthy) proof we refer to [212,
Satz 3.10].

Theorem 2.17 (G. Ascoli, 1884; C. Arzelà, 1895). Let Y be a compact


topological space, (X, d) a metric space, and let C(Y, X) denote the metric space of
continuous mappings from Y to X, equipped with the uniform metric

d∞ (f, g) := sup {min{d(f (y), g(y)) , 1} : y ∈ Y } .


2.2. COMPACT OPERATORS 19

Then we have for every V ⊂ C(Y, X):

V is relatively compact ⇐⇒
(
(i) V is uniformly continuous, and
(ii) V (y) := {f (y) : f ∈ V } is relatively compact in X for all y ∈ Y .
Here “V is uniformly continuous” means the following: For all y ∈ Y and ε > 0
there exists a neighborhood Uy of y such that d(f (y 0 ), f (y)) < ε for all y 0 ∈ Uy and
all f ∈ V .
The Arzela-Ascoli Theorem is applied mostly in the following form.
Corollary 2.18. Let Y be a compact topological space and V ⊂ C(Y, C).
Then we have:
V is relatively compact ⇐⇒ V is uniformly continuous and bounded.

I Despite the risks inherent in pictures, we can perhaps best visualize compact op-
erators as “asymptotically” contracting maps which in the case of operators of finite rank
map the ball BH to a finite-dimensional disk, and in general to some sort of elliptical
spiral as in Figure 2.1.

operator of
finite rank

BH
H

general
compact
operator

Figure 2.1. Compact operators visualized as elliptical spirals

David Hilbert made this visualization precise in his spectral representation of a


compact operator K. Accordingly (in the normal case KK ∗ = K ∗ K) the value Ku can
be expanded into a series in eigenvectors u1 , u2 , . . . with the corresponding eigenvalues as
coefficients, i.e.,
X∞
Ku = λj hu, uj i uj ,
j=1
whereby the eigenvalues accumulate at 0 and the eigenvectors form an orthonormal system.
In the language of operator algebras, that means that K vanishes at infinity.
In direct analogy to the principal axis transformation of analytic geometry, we thus
obtain for the quadratic form defined by K the representation
X∞
hKu, ui = λj hu, uj i2 .
j=1

All proofs can be found in [234, 1970/1982, 6.2-6.4] or [332, Lemma 3.3.5 and Theorem
3.3.8], the historical background in [201, 16, 34 and 40] or more compactly in [247,
p.1064-1066]. J
20 2. ANALYTIC METHODS. COMPACT OPERATORS

We give an elementary proof of the spectral representation of a compact oper-


ator in the simplest case, namely when the operator is self-adjoint. Our proof is
inspired by [10, Section 55] and [176, Theorem 5.1]; see also [111, Section 7.3–4],
where the very same chain of arguments is played through for the special case of
self-adjoint Sturm-Liouville boundary problems.
Definition 2.19. We say that an operator is diagonalizable (or “discrete”)
if there is an orthonormal basis {ej : j = 1, 2, . . . } for H and a bounded set
{λj : j = 1, 2, . . . } in C such that

X
Tu = λj hu, ej iej
j=1

for every u ∈ H.
Note that the numbers hu, ej i are the coordinates for u in the basis {ej } and
that each λj is an eigenvalue for T corresponding to the eigenvector ej . So, the
matrix corresponding to T and the basis {ej } is the diagonal matrix
 
λ1

 λ2 .

..
.
We are going to prove the spectral decomposition for compact operators (first
proven by David Hilbert in 1904 for Fredholm integral equations and generalized
by his student Erhard Schmidt in 1905).
Theorem 2.20 (Hilbert-Schmidt Theorem, 1904). Every compact self-adjoint
operator K is diagonalizable.
In the proof we shall use a simple, well-known technical result (Lemma 2.22
below) which follows from the also well-known proposition:
Proposition 2.21. If a bounded operator T is self-adjoint, then
kT k = sup hT u, ui .
kuk=1

The proposition remains valid for normal operators, see for instance [332,
Proposition 3.2.25] or [358, Theorem 12.25]. The number on the right is also
called the numerical radius of T and is denoted by |||T ||| .
Proof. Let m denote the numerical radius of T . We deduce m ≤ kT k from
the Cauchy–Schwarz inequality
hT u, ui ≤ kT uk · kuk ≤ kT k for kuk = 1.
To prove m ≥ kT k we consider arbitrary u, v ∈ H and obtain
hT (u ± v), u ± vi = hT u, ui ± 2<hT u, vi + hT v, vi
(using the fact that T is self-adjoint), from which
4<hT u, vi = hT (u + v), u + vi − hT (u − v), u − vi ,
where <z denotes the real part of a complex number z.
To bring m into play, we recall
w w 2 2
hT w, wi = hT , i kwk ≤ m kwk for every w ∈ H.
kwk kwk
2.2. COMPACT OPERATORS 21

We then get the estimate


2 2 2 2
(2.5) 4<hT u, vi ≤ m ku + vk + ku − vk ≤ 2m kuk + kvk
with the right inequality deduced from the parallelogram law
2 2 2 2
ku + vk + ku − vk = 2 kuk + 2 kvk .
If we replace u by αu with α ∈ C and |α| = 1, the right side of (2.5) remains
unchanged and we obtain (for α := e−iθ when hT u, vi = |hT u, vi|eiθ for suitable
real θ, hence |hT u, vi| = <hT (e−iθ u), vi):
m 2 2
(2.6) hT u, vi ≤ kuk + kvk .
2
kuk
Suppose T u 6= 0. Then taking v := kT uk T u in (2.6) yields
2
kuk kT uk = hT u, vi ≤ m kuk .
Hence kT uk ≤ m kuk for all u ∈ H and so kT k ≤ m. 
Now, the key to the spectral representation is the following
Lemma 2.22. If K ∈ B(H) is compact and self-adjoint, then at least one of the
numbers kKk or −kKk is an eigenvalue of K.
Proof. The lemma is trivial if K = 0. Assume K 6= 0. It follows from Proposi-
tion 2.21 that there exists a sequence (un ) in H of unit vectors with hKun , un i → λ,
where λ = kKk or λ = −kKk.
To prove that λ is an eigenvalue of K, we first note that
2
0 ≤ kKun − λun k = kKun k2 − 2λhKun , un i + λ2
≤ 2λ2 − 2λhKun , un i −→ 0.
Thus
(2.7) Kun − λun −→ 0.
Since K is compact, there exists a subsequence (Kun0 ) of (Kun ) which converges
to some v ∈ H. Consequently, (2.7) implies that un0 → λ1 v, and by continuity of
K,
1
v = lim Kun0 = Kv.
λ
Hence Kv = λv and v 6= 0 since kvk = lim kλun0 k = |λ| = kKk. Thus λ is an
eigenvalue of K. 
Now we prove the theorem by repeated application of the preceding lemma.
Proof of Theorem 2.20. By Lemma 2.22, there exists an eigenvalue λ1 of
K and a corresponding unit eigenvector e1 with |λ1 | = kKk. Set H1 := H with
the closed subspace H2 := [e1 ]⊥ . Here
(2.8) [a, b, c . . . ] := Ca + Cb + Cc + . . .
denotes the linear span of vectors a, b, c . . . . Now, T -invariance of a subspace M
for a bounded operator T implies that M ⊥ is T ∗ -invariant. Since K is self-adjoint
it follows KH2 ⊆ H2 . Set K1 := K and K2 := K|H2 ∈ B(H2 ). Then K2 is
compact and self–adjoint.
22 2. ANALYTIC METHODS. COMPACT OPERATORS

If K2 6= 0, we repeat the previous argument. So, there exists an eigenvalue λ2


of K2 and a corresponding unit eigenvector e2 with
|λ2 | = kK2 k ≤ kK1 k = |λ1 |.
Clearly the pair (e1 , e2 ) is orthonormal. Now H3 := [e1 , e2 ]⊥ is a closed subspace
of H, H3 ⊂ H2 and KH3 ⊆ H3 . Setting K3 := K|H3 , the process continues. It
either stops when Kn = 0 or else we get a sequence (λn ) of eigenvalues of K and a
corresponding set {e1 , e2 . . . . } of mutually orthonormal eigenvectors such that
|λn+1 | = kKn+1 k ≤ kKn k = |λn |, n = 1, 2, . . . .
If (λn ) is an infinite sequence, then λn → 0. Indeed, assume this is not the
case. Since |λn | ≥ |λn+1 |, there exists an  > 0 such that |λn | ≥  for all n. Hence
for n 6= m,
2
kKen − Kem k2 = kλn en − λm em k = λ2n + λ2m > 2 .
But this is impossible since (Ken ) has a convergent subsequence due to the com-
pactness of K.
We are now ready to prove the diagonalization of K as asserted in the theorem.
Let u ∈ H be given. Pn−1
Case 1. Kn = 0 for some n: Since un := u − k=1 hu, ek iek is orthogonal to ej ,
j = 1, . . . , n − 1, the vector un belongs to Hn . Hence
Xn−1
0 = Kn un = Ku − λk hu, ek iek .
k=1
Case 2. Kn 6= 0 for all n: From what we have seen in case 1 (and the simple
kun k ≤ kuk),
Xn−1
kKu − λk hu, ek iek k = kKn un k ≤ kKn kkun k ≤ |λn |kun k ≤ |λn |kuk −→ 0,
k=1
which means that X∞
Ku = λk hu, ek iek . 
k=1

Corollary 2.23. Let K be a self-adjoint compact operator and ρ > 0 a real.


Then the operator K has only a finite number of linearly independent eigenvectors
such that the corresponding eigenvalues exceed ρ in modulus. In particular, only 0
can be an accumulation point of the eigenvalues, and each nonzero eigenvalue has
finite multiplicity.
In [10, Section 52] it is shown that the preceding assertion remains valid for
any compact operator; i.e., the assumption that K is self-adjoint is dispensable.

Note that an operator K


• is bounded (= continuous) if and only if for every u ∈ H
lim K(u + v) = K(u),
|v|−→0

• but is compact, if and only if, with respect to a complete orthogonal


system e1 , e2 , . . ., we have
lim hvi , ej i = 0 for each j =⇒ lim K(u + vi ) = K(u),
i−→∞ i→∞
2 P∞ 2
in which the hypothesis is weaker than kvi k = j=1 |hvi , ej i| → 0.
2.2. COMPACT OPERATORS 23

This describes the context in which Hilbert developed the idea of a compact
operator and why he called them completely continuous. For our purposes the
following result is sufficient.
Theorem 2.24. a) K is a “twosided (nontrivial) ideal” in the Banach algebra
B of bounded linear operators in a separable, infinite-dimensional Hilbert space H.
b) K is closed in B.
c) More precisely, K is the closure of the subset of finite-rank operators.
d) K is invariant under ∗; i.e., the adjoint of a compact operator is compact.
Proof. To (a): For bounded T and compact K, the operators T ◦ K and
K ◦ T are compact, by definition. We therefore have B ◦ K ⊆ K and K ◦ B ⊆ K.
Moreover, for λ ∈ C, we have λK compact. Now, let K and K 0 be compact. To
prove that K + K 0 is compact, we use the sequential criterion for compactness;
i.e., an operator is compact, if the image of a bounded sequence of points has a
convergent subsequence. Thus, let u1 , u2 , . . . be a bounded sequence in H. Then, by
a two-fold selection of subsequences, we can find ui1 , ui2 , ..., such that Kui1 , Kui2 , ...
and K 0 ui1 , K 0 ui2 , ...both converge in H, whence (K + K 0 ) ui1 ,(K + K 0 ) ui2 , ... also
converges. Finally, it is trivial that every compact operator is bounded, since the
image of the unit sphere is relatively compact and hence bounded; thus, K ⊆ B.
Since K includes the operators of finite rank (Exercise 2.15a) but not the identity
(Exercise 2.15c), K is a nontrivial ideal, and the assertion follows.
To (b): Let T ∈ B be an operator in the closure of K. In order to show that
each open cover (say, without loss of generality, by all of the balls of radius ε > 0
[130, 1966, p.298]) of the image T (BH ) of the closed unit ball of H has a finite
subcover, we use an “ε/3-proof” (as is usual in such situations): We choose K ∈ K
with kT − Kk < ε/3 and a finite open covering of K(BH ) by balls of radius ε/3,
with centers at Ku1 , , ..., Kum where u1 , ..., um ∈ BH . Then the ε-balls about
T u1 , ..., T um form the desired finite covering of T (BH ): For each u ∈ BH , there is
some i ∈ {1, ..., m}, such that kKu − Kui k < ε/3, and so
kT u − T ui k ≤ kT u − Kuk + kKu − Kui k + kKui − T ui k < ε.

To (c): We now ask which operators are limits in B of sequences of operators of


finite rank. Evidently, these limits typically lie outside of the space of finite rank
operators; e.g., see Exercise 2.28 below. Let e1 , e2 , ... be a complete orthonormal
system in H and let
Qn : H −→ [e1 , ..., en ]
denote the orthogonal projection from H to the linear span of the first n basis
elements. This truncation yields, for each T ∈ B, a sequence Q1 T, Q2 T, ... of
operators of finite rank which converges pointwise to T ; i.e., Qn T u → T u for each

u ∈ H. This does not mean that the sequence Q1 T, Q2 T, ... = (Qn T )1 converges
in B (i.e., in the operator norm) to T . Of course, not every bounded operator can
be the limit of a sequence of operators of finite rank (e.g., kQn Id − Idk = 1 for
all n); in fact, we have already proved in (b) that limits of compact operators (in
particular, finite-rank operators - see Exercise 2.15 a) must be compact. It remains

to prove that, for each K ∈ K, the sequence (Qn K)1 converges to K in B. For
this, we choose (for each ε > 0) a finite covering of K(BH ) by balls of radius ε/3
with centers at Ku1 , ..., Kum and some n ∈ N, so large that
kKui − Qn Kui k < ε/3 for i = 1, ..., m.
24 2. ANALYTIC METHODS. COMPACT OPERATORS


This is no problem, since (Qn )1 converges pointwise to the identity. For each u ∈
BH , we then have
kKu − Qn Kuk ≤ kKu − Kui k + kKui − Qn Kui k + kQn Kui − Qn Kuk
< ε/3 + ε/3 + ε/3
for some i ∈ {1, ..., m}. For the last term, note that kQn k = 1, whence
kQn Kui − Qn Kuk = kQn (Kui − Ku)k ≤ kKui − Kuk .
Thus, we have proven kK − Qn Kk < ε for all sufficiently large n, depending on ε.
under ∗, since
To (d): Obviously, the space of operators of finite rank is invariant P
n
every such operator T (as in Remark 1.13, p. 9)P is of the form T = i=1 h·, ui i vi ,
∗ n
where u1 , ..., vn ∈ H, and then (verify!) T = i=1 h·, vi i ui also has finite rank.
Now, we can reduce the general case K ∈ K to the finite-rank case. Namely,
approximate K by a sequence (Tn )∞ 1 of operators of finite rank whose adjoints
then approximate K ∗ , since (by Exercise 2.6),

kTn∗ − K ∗ k = (Tn − K) = kTn − Kk . 
Remark 2.25. The statements in (a), (b), and (d) apply also (admittedly with
somewhat different proofs; e.g., see [358, Theorems 4.18, 4.19] or [365, Section 4.3])
to the more general case of Banach spaces, but not statement (c). The search for a
counterexample began with a legendary treatise by Alexander Grothendieck
in [188] (1955) and led Per Enflo to success, published in [143] (1973), inci-
dentally supplemented by many nice examples of the correctness of (c) in spe-
cial cases, in particular for almost all well-known Banach spaces; see also [234,
1970/1982, 12.4], but — surprisingly not for B(H): in [405] (1981), it was proved
by Andrzej Szankowski that the Banach algebra B(H) of bounded operators
in complex separable Hilbert space is not approximative, i.e., there exists a com-
pact operator k : B(H) → B(H) which can not be approximated by a sequence
(fj )j=1,2,... : B(H) → B(H) of operators (on the operator space B(H)) of finite
range.
Remark 2.26. Since the closure of an ideal is again an ideal (as is trivially
proved) and since the operators of finite rank obviously form an ideal in B, we note
that (a) follows from (c) – admittedly, somewhat less directly than in the above
proof.
Remark 2.27. For practical needs the sequence (Qn K)∞ 1 stated in the proof
of (c) is a poor approximation of K by operators of finite rank. It presupposes the
knowledge of K on all of H and works with a completely arbitrary orthonormal sys-
tem. However K is frequently (see for example Exercises 2.28 and 2.29 below) given
in a form which suggests a special approximation or which points to a distinguished
orthonormal system, namely the eigenvectors of K. The spectral representation of
Theorem 2.20 is numerically relevant, because it implies that
Xn
Qn Ku = KQn u = λj hu, uj i uj
j=1
which does indeed permit a stepwise approximation.
Exercise 2.28. Show that the operator A in Exercise 2.4 is compact. Namely,
Pn 2
give an estimate for Au − j=1 1j hu, ej i ej which is independent of u, for kuk
< 1. Recall the Cauchy-Schwarz inequality |hu, ej i| ≤ kuk kej k.
2.3. THE CLASSICAL INTEGRAL OPERATORS 25

Exercise 2.29. Let [a, b] be a compact interval in R. Show that we obtain a


compact operator K on the Hilbert space L2 [a, b] for each square-integrable function
G on [a, b] × [a, b] via
Z b
Ku(x) := G(x, y)u(y) dy, x ∈ [a, b]
a
[Hint: Approximate the weight function G by step functions, and hence K by
operators with finite rank of the kind considered in Exercise 1.14 (integral operators
with degenerate weight). For details of the argument, see [176, Example II.14.3],
[365, Section 11.4], [243, Example V.2.19], [332, Propositon 3.4.16], and [352,
Theorem VI.23]. The last three references show that one can replace the interval
[a, b] by any locally-compact Hausdorff space with a fixed Radon integral and that
the integral operator is not only compact but Hilbert-Schmidt, i.e., with the
eigenvalues of K ∗ K going to zero fast enough to be summable (i.e., the trace
Tr(K ∗ K) < ∞).]
Exercise 2.30. For the map
L2 ([a, b] × [a, b]) −→ K(L2 [a, b])
G 7→ K ,
defined in Exercise 2.29, show that:
a) The map is linear and injective.
b) If G is the weight function of K, then
!1/2
Z b Z b
2
kKk ≤ |G| := |G(x, y)| dxdy .
a a

c) The adjoint operator K ∗ has the weight function G∗ (x, y) = G(y, x).
d) If K1 and K2 are given by weight functions G1 and G2 , then the operator K2 ◦
K1 belongs to the weight function
Z b
G(x, y) = G2 (x, z) G1 (z, y) dz.
a

[Hint: Linearity is clear. For injectivity, one naturally (Lebesgue integral!) need
RβRδ
only show that α γ G(x, y) dxdy = χ[α,β] , Kχ[γ,δ] , where χ[α,β] and χ[γ,δ] are
characteristic functions of subintervals [α, β] , [γ, δ] ⊆ [a, b]; then, we have G = 0
when K = 0. Details (and the generalization to the case of unbounded intervals)
are in [234, 1970/1982, 11.2]. For the proofs of (b), (c), and (d), one needs to use
the theorem of Fubini on iterated integrals; e.g., the details are in [234, 1970/1982,
11.2-11.3] or [332, Propositon 3.4.16].]

3. The Classical Integral Operators


I The type of operators considered here, i.e., those which are given by weight func-
tions square-integrable on the product space, are nowadays called Hilbert-Schmidt oper-
ators. For a rigorous treatment of abstract Hilbert-Schmidt operators, see Section 2.7
below. By introducing them, the Hungarian mathematician Frigyes (Friedrich) Riesz
(1907), not only generalized the theory of integral equations with continuous weight func-
tion created by Vito Volterra (Turin, 1896), Erik Ivar Fredholm (Stockholm, 1900)
and David Hilbert (Göttingen, 1904f), but drastically simplified it at the same time.
26 2. ANALYTIC METHODS. COMPACT OPERATORS

Indeed, proofs dealing with L2 -integration theory on Hilbert space compare very favorably
with the cumbersome work with uniform convergence in the Banach space of continuous
functions. J

Of course, there do exist important integral operators whose weight functions


are not square integrable: The best-known example is the Fourier transform (see
Appendix A below) Z ∞
fb(x) = e−ixy f (y) dy.
−∞
It maps L2 (R) bijectively onto itself, but the L2 -norm of the weight,
Z ∞Z ∞ Z ∞Z ∞
−ixy 2
e dxdy = 1dxdy
−∞ −∞ −∞ −∞
is strongly infinite. Somewhere between these two simplest principal types – the
compact Hilbert-Schmidt operators on the one hand and the invertible Fourier
transform on the other – lie the convolution operators, in particular the Wiener-
Hopf operators and other singular integral operators. Their weights do not belong
to L2 , but frequently they are at least componentwise integrable, for example
Z b
2
|G(x, y)| dy < ∞ for x ∈ R.
a

I All these operators are highly significant in kinematic as well as stochastic mod-
eling and in solving a multitude of physical, technical and economical problems: as in
the method of inverting differential operators into integral operators which goes back to
George Green and was developed on a large scale by David Hilbert; and as in the in-
direct treatment of collective and statistical phenomena or more generally in probabilistic
situations. We will return to a number of particularly interesting integral operators later
in this Part and in the following Part. A first survey is given in Table 2.1. J

4. The Fredholm Alternative and the Riesz Lemma


If F is a Fredholm operator on a separable Hilbert space H, then the statement
“index F = 0” can be expressed in familiar classical terminology as: Either the
equation F u = v has a unique solution u ∈ H for each v ∈ H, or the homogeneous
equation F u = 0 has a nontrivial solution. In the second case, there are at most
finitely many linearly independent solutions w1 , ..., wn of F w = 0 and just as many
linearly independent solutions u1 , ..., un of the adjoint homogeneous equation F ∗ u =
0; the inhomogeneous equation F u = v is solvable exactly when
hv, u1 i = · · · = hv, un i = 0.
This statement is called the Fredholm alternative (met before in Exercise 1.14);
its equivalence with “index F = 0” follows from Theorem 2.7. We saw already that
the Fredholm alternative holds for self-adjoint (F ∗ = F ), or more generally normal
(F ∗ F = F F ∗ ), Fredholm operators. Moreover, in analogy with finite-dimensional
linear algebra, it holds for operators of the form Id +T , when T is an operator of
finite rank. Even more important for many applications are the following operators
for which the Fredholm alternative holds:
Theorem 2.31 (F. Riesz, 1918). For each compact operator K, Id +K is a
Fredholm operator with vanishing index.
No. X G(x, y) u Key words Property Literature or Origin

1. [a, b] C 0 (X × X) C 0 (X) K : C 0 (X) → C 0 (X) com-


Fredholm integral Fredholm (1903)
pact

2. compact ⊂ Rn ” ” ” ”
Hilbert (1904-12)

3. [a, b] continuous for y ≤ x, and ” ”


Volterra integral Volterra (1896)
= 0 for y > x

4. ” bounded on X × X and in ”
Green’s [234, Sect. 8.2]
C 0 (X × X \ diag)
operator
−α
5. compact ⊂ Rn G(x, y) = g(x, y) |x − y| , ” pole singularity ” ”
Ku(x) := X G(x, y)u(y) dy

g cont., α ∈ (0, n)

6. [a, b] L2 (X × X) L2 (X) K : L2 (X) → L2 (X) com-


Hilbert-Schmidt Riesz (1907)
pact
operator

7. Rn ” ” ” Exercises 2.29, 2.30


[234, Sect. 11.3]

8. ” (2π)−n/2 e−i(x1 y1 +···xn yn ) ” isometry of L2 (X)


Fourier transform [134, Sect. 2.10] and our
Appendix A

9. R g(x, y) = k(x − y), ” see Appendix A.2 [134, Sect. 2.1]


convolution
k ∈ L1 (R)

10. ” ” see Section 4.6 [134, Sect. 3.6]


2.4. THE FREDHOLM ALTERNATIVE AND THE RIESZ LEMMA

R+
Wiener-Hopf
operator
Table 2.1.R Some fundamental integral operators of the form

−α
11. Rn G(x, y) = g(x, y) |x − y| , ” see Chapter 8
singular integral [294]
g cont., α ∈ (n, ∞)
operator
27
28 2. ANALYTIC METHODS. COMPACT OPERATORS

Note that our proof functions only in Hilbert space or in approximative Banach
space, see Remark 2.25. An alternative proof with wider applicability can be found
in [242, Theorem IV.5.26, p.238f].
Proof. Using Theorem 2.24c, approximate K by a sequence K1 , K2 ,... of
operators of finite rank and choose n with kK − Kn k < 1. ThenPId +K − Kn is

invertible. Indeed, let Q := K − Kn and consider the series k=0 Qk . Since
kQk < 1 and
XM XM XM k
Qk ≤ Qk ≤ kQk ,
k=N k=N k=N

the partial sums form a Cauchy sequence in the Banach algebra B. The series then
converges, and one has
X∞ X∞ 
(Id −Q) Qk = Qk (Id −Q) = Id .
k=0 k=0
Thus, we can write Id +K as a product:
Id +K = (Id +K − Kn )(Id +(Id +K − Kn )−1 Kn )
where the left factor is invertible and the right is Id + an operator of finite rank,
which we know (Exercise 1.12, p.9) is a Fredholm operator with index 0. By the
composition rule of Exercise 1.10 (p.8) or Exercise 2.3 (p.14), the statement is
proved. 
Remark 2.32. Note that in the above proof of the formula index(Id +K) = 0,
it was not needed that K is approximated by a sequence of operators of finite rank.
It was sufficient to have a crudely approximating operator K (with kKn − Kk <
1). Also, for the determination of the index, it was not necessary to compute
or know more precisely the inverse operator (Id +K − Kn )−1 . This method of
proof which reduces the general case to situations permitting explicit or at least
iterative solutions goes back to E. Schmidt and works only in Hilbert space. In
more general cases where the theorem still holds the proof starts by showing with
a compactness argument that Ker(Id +K) (and similarly Ker(Id +K ∗ )) is finite
dimensional. Then one needs a careful argument concerning the limit process in
order to show that Im(Id +K) is closed, and one gets only that Id +K ∈ F. Finally
the homotopy invariant of the index (see Theorem 3.11, p.68) implies the formula
index(Id +K) = 0.
Exercise 2.33. Formulate and prove the Fredholm alternative for the linear
Fredholm integral equation of the second kind
Z b
u(x) + G(x, y)u(y) dy = h(x),
a
2
where G ∈ L ([a, b] × [a, b]); see Exercise 1.14 (p.10), Theorem 2.7 (p.15), Exercise
2.29 (p.25), and Exercise 2.30 (p.25).

5. Sturm-Liouville Boundary Value Problems


I We will apply Theorem 2.31 and Exercise 2.33 to a classical boundary value prob-
lem in the theory of ordinary differential equations which is not very transparent. We
start with a dynamical system with finitely many degrees of freedom described by a sys-
tem of ordinary differential equations. Such ideally simple models are used in celestial
mechanics for the computation of planetary orbits, for investigations of the pendulum and
2.5. STURM-LIOUVILLE BOUNDARY VALUE PROBLEMS 29

gyroscope, for econometric simulation of economic processes, and for the treatment of
many other discrete oscillating systems (N -body problems). Many of these problems are
mathematically unsolved, but one knows at least that each set of initial values determines
a unique solution curve; i.e., given the differential equation (with half-way reasonable co-
efficients), the system is completely determined by its state at a single moment in time.
The key mathematical tools are the local existence and uniqueness theorems of Augustin
Cauchy, Emile Picard, Rudolf Lipschitz, Giuseppe Peano, and Ernst Lindelöf.
See, for example, [111, Chapters 1-2] or [356, p.164-170]. We have a different situation
with continuously distributed oscillating systems such as the oscillating string or flexi-
ble rod, electrical oscillations in wires, acoustical vibrations in tubes, heat conduction,
heat propagation and other diffusion processes, particularly the statistical treatment of
equilibria and motion. J

For many such processes, one has partial differential equations (see Part II
below) for instance of the form
∂2U 2
(2.9) ∂x2 = ρ ∂∂tU2 + F (x, t) , x ∈ [0, 1] .
Under suitable assumptions, (2.9) can be reduced to an ordinary differential equa-
tion by separating variables. More explicitly, setting
U (x, t) = u(x)ψ(t),
we obtain
(2.10) u00 + ru = f,
where r, f are given, and u is to be found. Hereby (2.10) usually inherits boundary
conditions, e.g.,
(2.11) u(0) = u(1) = 0
from (2.9). For details and generalizations, see [116, I, V.3], [61, p.25ff, 285ff,
351ff], [281, p.108-152] or [356, p.164-170].
We stick with this example which goes back to John Bernoulli’s brachis-
tochrone problem and more generally to the beginnings of the calculus of variations
and of geometric optics by Pierre de Fermat [247, Ch. 24]. For starters let
r = 0. Evidently the homogeneous differential equation associated with (2.10) (put
f = 0) has only the trivial solution u = 0 if the boundary conditions (2.11) are to be
satisfied. In this case there is a Green’s function (see e.g., [111, Theorem 7.2.2] or
[116, I, V.14-15]) which, for each (piecewise continuous) f , yields a solution of the
differential equation (2.10) with boundary conditions (2.11) given by the formula
Z 1
(2.12) u(x) = k(x, y)f (y) dy.
0

For the present case of (2.10) with r = 0 and (2.11),



x(y − 1), for x ≤ y,
k(x, y) =
y(x − 1), for x ≥ y
[116, I, V.15.1].
Now let f = 0 and let r be a positive real number. Then the general solution
of (2.10) has the form
√ √
u(x) = c1 ei rx
+ c2 e−i rx
;
30 2. ANALYTIC METHODS. COMPACT OPERATORS

i.e., each of the two conditions of (2.11) determines a one-dimensional family of


solutions where 
−1,√ for u(0) = 0,
c1 /c2 =
−e−2i r , for u(1) = 0.

The two families coincide, exactly when r is a multiple of π. If so, there is in
addition to the zero function another solution u0 of the homogeneous differential
equation
(2.13) u00 + ru = 0
which satisfies the boundary conditions (2.11). Thus we have here – in contrast
to the Uniqueness Theorem of Lipschitz – an ordinary differential equation of
second order with two conditions imposed (which, however, are not concentrated
at an initial point but distributed over two points) whose solution is not unique.
Another peculiarity of this case is that in contrast to the Existence Theorem of
Picard the inhomogeneous differential equation (2.10) subject to the conditions
(2.11) need not always have a solution. This happens for example if the driving
force f (the source term) equals the eigenfunction u0 , or more generally [116, I,
V.14.2] if
Z 1
f (x)u0 (x) dx 6= 0.
0

I This is the case of resonance which means that the system becomes unstable under
the influence of an exterior force. The lack of a solution does not mean that nothing
happens, but from the point of view of the user it is an indication that some critical
phenomenon might occur: A short in a wire, the collapse of a bridge, extreme concentration
of light beams which is technically utilized in a laser. Also, mathematically speaking, the
lack of a solution means only that no solution of the given or desired type exists, in
our example no bounded function which has a piecewise continuous second derivative.
Frequently, this is an indication that a reformulation or refinement of the question is
necessary. J

We note finally that in the other case, when the nontrivial solution sets with
u(0) = 0 and those with u(1) = 0 are disjoint, the equations (2.10) and (2.11)
always have a unique solution, and a Green function can be constructed which
carries the essential information of (2.10) and (2.11) and yields the solution for
each right hand side f in the integral form (2.12) [116, I, V.14.1]. Combining the
two cases we obtain a kind of Fredholm alternative:
Either the differential equation (2.10) together with the boundary conditions
(2.11) possesses a unique solution u for every given f , or else the homogeneous
equation (2.13) has a solution which does not vanish identically. In the second case
the equations (2.10) and (2.11) have a solution, if and only if the orthogonality
condition
Z 1
f (x)u(x) dx = 0
0
holds for each solution u of the homogeneous equation (2.13), where f is the right
hand side of (2.10).
The analogy with the Fredholm alternative for integral equations (Exercise 2.33,
p.28) is not accidental. When the solution is unique, the Green function makes
the connection via formula (2.12). But even when solutions do not necessarily
2.5. STURM-LIOUVILLE BOUNDARY VALUE PROBLEMS 31

exist or when they are not unique, then the classical theory manages to work with
generalized Green functions.
We will not deal with these questions in detail, but refer the reader to the
quoted literature. The fundamental methodological and, in our context, particu-
larly interesting point of view is perhaps best made precise as follows:
Exercise 2.34. Consider the differential equation
(2.14) u00 + pu0 + qu = f
on the interval [0, 1] with p, q, f ∈ C 0 [0, 1] and with the boundary conditions
(2.15) u(0) = a and u(1) = b.
Show that the integral equation
(2.16) v − Kv = g
is equivalent to the boundary-value problem (2.14), (2.15), if
Z 1
Kv(x) := G(x, y)v(y) dy, x ∈ [0, 1] ,
0

y(q(x)(1 − x) − p(x)), for y ≤ x,
G(x, y) :=
(1 − y)(q(x)x + p(x)), for y > x,
g := ph0 + qh − f, and
h := a(1 − x) + bx (hence, h0 = b − a).

[Hint: Show that every twice continuously differentiable solution of (2.14) and (2.15)
yields a solution v := u00 of (2.16), and conversely, every continuous solution v
of (2.16) gives a twice continuously differentiable solution of (2.14) and (2.15) by
means of
Z 1
u(x) := h(x) + k(x, y)v(y) dy, where
0

x(1 − y), for x ≤ y,
k(x, y) :=
y(1 − x), for x ≥ y.
Details may be found in [234, 1970/1982, 9.1] (see also [332, Section 3.4.18] for a
rigorous treatment of the general symmetric second-order differential equation with
a certain periodic self-adjoint boundary condition). For a first calculation and in
order to maintain continuity with the preliminary remarks, it is recommended that
one first try p = 0, q positive and constant, and set a = b = 0, see also p.67.]
Remark 2.35. With Exercise 2.33, the Fredholm alternative for the boundary-
value problem follows from the equivalence proved in Exercise 2.34. More precisely,
Id −K is a Fredholm operator on the Hilbert space L2 [0, 1], and index(Id −K) =
0. Thus, we have a Fredholm alternative relative to L2 [0, 1]. Actually, from the
Closed Graph Theorem and a regularity theorem (see Chapter 9 and Section 13.8;
incidentally, we see here that the Banach space theory is genuinely more difficult
than the Hilbert space theory), we have that each square integrable solution of
equation (2.14) is continuous, provided the right side is continuous. Thus, the
Fredholm alternative in C 0 [0, 1] holds: Either dim Ker(Id −K) = 0 and so (because
index(Id −K) = 0, and hence dim Coker(Id −K) = 0) the equation (2.14) has a
unique solution for each g ∈ C 0 [0, 1] (whence, the boundary-value problem (2.14),
32 2. ANALYTIC METHODS. COMPACT OPERATORS

(2.15) also has a unique solution for each f ∈ C 0 [0, 1] and fixed boundary values
a, b), or the homogeneous equation v − Kv = 0 has nontrivial solutions.
In the second case, the adjoint integral equation also has a nontrivial solution; i.e.,
there is a w ∈ L2 [0, 1] with
Z x
w(x) = (1 − x) (q(y)y + p(y))w(y) dy
0
Z 1
+x (q(y)(1 − y) − p(y))w(y) dy.
x

We can differentiate with respect to the upper and lower bounds, obtaining that w
is continuously differentiable and
Z x
0
w (x) = − (q(y)y + p(y))w(y) dy + (1 − x)(q(x)x + p(x))w(x)
0
Z 1
+ (q(y)(1 − y) − p(y))w(y) dy − x(q(x)(1 − x) − p(x))w(x)
x
Z x Z 1
= p(x))w(x) − ··· + ··· .
0 x

We bring p(x)w(x) to the left side, differentiate once more, and obtain
(2.17) (w0 − pw)0 + qw = 0.
From the integral equation for w, we have w(0) = w(1) = 0. Hence, every solution
w of the homogeneous adjoint integral equation is a solution of the formal adjoint
homogeneous differential equation (2.17) with the homogeneous boundary condi-
tions w(0) = w(1) = 0. By formal adjoint, we mean that for all u, w ∈ C 2 [0, 1]
with the homogeneous boundary condition, we have
Z 1 Z 1
u((w0 − pw)0 + qw) dx = (u00 + pu0 + qu) w dx
0 0

which one can verify through integration by parts. For brevity, we have taken all
functions to be real-valued.
In the second case of the Fredholm alternative, the problem (2.14) and (2.15) is
solvable exactly when (2.16) is solvable; i.e., when (2.17) has a nontrivial solution
w with hg, wi = 0, which means that in terms of f , we have
Z 1
f (x)w(x) dx = aw0 (0) − bw0 (1).
0

Details of this argument and similar treatments of other boundary-value problems


for ordinary differential equations of second order (Sturm-Liouville problems) can
be found in [111, Chapters 7 and 12] and [234, 1970/1982, 9.2].
Remark 2.36. If the boundary value problem (2.14) and (2.15) is equivalent
to the integral equation (2.16), what is the special nature of the presentation (2.16)
in comparison with (2.14) and (2.15)? We bring out three points:
1. The integral equation succeeds in combining two equations, the differential
equation and the boundary conditions, into one.
2. Let
L : C 2 [0, 1] −→ C 0 [0, 1]
2.5. STURM-LIOUVILLE BOUNDARY VALUE PROBLEMS 33

denote the differential operator defined by the left side of (2.14) and let
B : C 2 [0, 1] −→ C ⊕ C
denote the boundary operator defined by the left sides of (2.15). Then Exercise
2.34 says that the operators
L⊕B : C 2 [0, 1] −→ C 0 [0, 1]⊕C ⊕ C and
Id −K : L2 [0, 1] −→ L2 [0, 1]
are equivalent in the sense that Ker(L⊕B) ∼ = Ker(Id −K) and Coker(L⊕B) ∼ =
Coker(Id −K). Here Id −K is a bounded operator on a Hilbert space to itself,
while the functional analytic structure of L⊕B is much less clear. The equivalence
of the differential and the integral equation is a formal one, while the equivalence
of the C 0 /C 2 -theory and the L2 -theory is fairly elementary, but by no means ob-
vious: While nature poses its problems usually in the spaces C 2 or C 2(piecewise) ,
mathematicians decide freely in which spaces they want to solve these problems.
Hilbert spaces are used, not because of their intrinsic beauty, but because integral
equations on L2 can be treated more efficiently and more transparently than on C 0 .
The Regularity Theorem provides the justification for this procedure and shows at
the same time that the freedom of the mathematician is not arbitrary.
3. Numerically, the integral operator Id −K is dealt with by approximating the
compact operator K by operators of finite rank or by approximating the weight
function G(x, y) of K by degenerate weights of the form φ(x)ψ(y). Jacques-
Charles-Francois Sturm and Joseph Liouville first and successfully under-
took the systematic investigation of boundary value problems for ordinary differ-
ential equations of second order. It is quite characteristic that they also arrived
at their algebraic solution methods by an approximation principle: They investi-
gated related difference equations and then passed to the limit (Jour. de Math.
1 (1836), 106-186 and 373-444). The difference is that the approximation of the
integral equation in some cases (e.g., when G is continuous and nonnegative) can
be done very naturally by development into a series in eigenfunctions (analogous
to principal axis transformation of quadratic forms) so that the integral equation
becomes immediately clearer ([116, I, III.5.1]). In contrast, the approximation of a
differential equation by difference equations is done blindly so to speak. It requires
the ingenuity of a Sturm and Liouville (or nowadays extensive free computer
time) to regain the necessary information about the boundary value problem from
the discrete pieces. (Of course, the blind approximation always works, while for
many Sturm-Liouville problems no explicit eigenfunctions are known.)

I In the final analysis the three viewpoints arise from the duality between local and
global terms and operations. This duality pervades large parts of analysis (see the Fourier
Inversion Formula in Appendix A or the Index Formula itself): While differentiation of
a function is a purely local operation, the solution of a differential equation with initial
or boundary conditions always requires a certain global operation. This circumstance is
illustrated already by the Newtonian formula relating derivative and integral. It may also
explain why the local theory of (e.g., elliptic) differential equations is so difficult (one has
to do global theory anyway, namely in Rn ), and why at times a purposely global approach,
say starting with differential operators on closed manifolds, leads more quickly and easily
to fundamental local results. We resume this thought in Part II. J
34 2. ANALYTIC METHODS. COMPACT OPERATORS

6. Unbounded Operators
So far we have considered only bounded Fredholm operators, i.e., linear Fred-
holm operators from one separable Hilbert space H1 to another separable Hilbert
space H2 which are continuous and defined on all of H1 . Identifying H1 and H2
we ended up with the space of bounded Fredholm operators as a subspace of the
algebra of bounded operators B(H). That is the main line of presentation chosen
for this book. It is sufficient for establishing the Atiyah–Singer Index Theorem.
Therefore, a hurried reader may skip this section.
However, since our Definition 1.1 (p.3) of Fredholm operator is purely algebraic,
we can reformulate and generalize most of the functional analytical and topological
results of this book concerning bounded Fredholm operators to the unbounded case
(though still assuming the linearity of the operators).
There is good reason to make this generalization: Differential operators are
naturally defined on domains which are dense subspaces of the full L2 . However,
there is no reasonable way to extend them to endomorphisms, acting on the full L2
and with values in the same L2 –space. And on their domain, they are not bounded
relative to the L2 –norm. In particular for a closed (not necessarily bounded) oper-
ator T (see below) there are various ways to re-write or transform T as a bounded
operator, be it by Riesz transform (for self-adjoint T ), Cayley transform, or sim-
ply by equipping the domain Dom(T ) with the graph norm, to be recalled below
in (2.23), p.36. For index theory, these approaches are valuable to some extent,
as we shall show. However, they always distort the picture, and a treatment of
unbounded operator is in order.
Exercise 2.37. Let the unit circle S 1 be parametrized by the angle θ ∈ [0, 2π).
Consider the Hilbert space L2 (S 1 ) of all square Lebesgue integrable complex–valued
functions on S 1 with inner product
Z 2π
(2.18) hu, viL2 := u(θ)v(θ) dθ for u, v ∈ L2 (S 1 )
0
p
and norm kukL2 := hu, uiL2 (see Exercise A.1, p.705). On the dense sub-
space C 1 (S 1 ) of differentiable (periodic) functions with continuous derivative (l.c.,
Exercise A.1e), the differentiation d/dθ defines a linear operator
du
(2.19) T0 : C 1 (S 1 ) 3 u 7→ u0 = ∈ L2 (S 1 )

with Z 2π
0 1
Im(T0 ) = {v ∈ C (S ) : v(θ)dθ = 0}.
0
Show that T0 is not continuous as a mapping in L2 (S 1 ).
[Hint: ku0 kL2 can be arbitrarily large for kukL2 = 1.]
Of course, T0 = d/dθ is bounded if it is regarded as an operator from the
Banach space C 1 (S 1 ) to the Banach space C 0 (S 1 ). Many problems in analysis,
however, require exploiting Hilbert space structure for effective treatment. After
all, the state space of quantum mechanics is a Hilbert space; and the Spectral
Theorem, both for bounded and unbounded operators, is valid only in Hilbert space.
Admittedly, most results on Fredholm operators can also be obtained in Banach
space, but are prettier, more meaningful, and much simpler in Hilbert space.
2.6. UNBOUNDED OPERATORS 35

These problems lead to unbounded operators, but not irrevocably: Alterna-


tively, staying with our example, we can extend the differential operator T0 to the
first Sobolev space W 1 (S 1 ). It is the completion of the pre–Hilbert space C 1 (S 1 )
equipped with the inner product
Z 2π Z 2π
(2.20) hu, viW 1 := u(θ)v(θ) dθ + u0 (θ)v 0 (θ) dθ
0 0

and hence a Hilbert space. Then the extension (easily produced in Theorem 2.40
a)
(2.21) T : W 1 (S 1 ) −→ L2 (S 1 )
of T0 becomes a bounded Fredholm operator from the (whole) Hilbert space W 1 (S 1 )
to the (different) Hilbert space L2 (S 1 ). Thus, unbounded operators can be averted
in this sense.
In Part II of this book we shall follow this approach. It is quite effective for
the study of elliptic differential operators on closed manifolds, but not sufficient for
the study of boundary value problems. Different boundary conditions give different
extensions of the formal differential operator and different domains. Therefore,
varying boundary value problems for a fixed formal differential operator can often
best be treated in a shared framework, namely considering them all as densely
defined unbounded operators, operating by the same formal rules in the same basic
L2 space and distinguished only by their domains. The point is that two such
operators may be equal on a dense subspace, and yet be quite different. For a 1–
dimensional example see our discussion of Sturm–Liouville problems in Section 2.5.

Closed Operators. We emphasized that bounded operators in Hilbert space


share many properties with familiar matrix calculus in finite dimensions; that differ-
ential operators can be treated as bounded operators in suitably re-defined domains,
the so-called Sobolev spaces; and that those Sobolev spaces in our Part II will be
equipped with scalar products that make them separable Hilbert spaces. However,
a student would be mislead, if we over-emphasize the analogy with elementary
Euclidean linear algebra. To get a realistic feeling for the delicate aspects of dif-
ferential operators, of partial differential equations and global analysis, the reader
must become knowledgable of the basic concepts and fundamental results regarding
unbounded operators. That is the goal of the following short course in closed oper-
ators and not necessarily bounded Fredholm operators. The reader should work it
through, even though we can derive most of the key results of index theory without
referring to that world, to that language and theory.
Definition 2.38. Let H be a Hilbert space.
a) A densely defined operator (not necessarily bounded) in H is a linear mapping
T : Dom(T ) −→ H ,
where Dom(T ) – the domain of T – is a (linear) dense subspace of H.
b) If S and T are operators in H such that Dom(S) ⊆ Dom(T ) and Su = T u for
every u ∈ Dom(S), we say that T is an extension of S and write S ⊆ T .
c) For a densely defined operator T in H, we form the adjoint operator T ∗ in H
by letting Dom(T ∗ ) denote the subspace of elements u ∈ H for which the functional
v 7→ hT v, ui on Dom(T ) is bounded (= continuous). Since Dom(T ) is dense in H,
36 2. ANALYTIC METHODS. COMPACT OPERATORS

the functional extends by continuity to H, and thus there is a unique element T ∗ u


in H such that
(2.22) hv, T ∗ ui = hT v, ui for all v ∈ Dom(T ).

d) A closed operator in H is a densely defined operator whose graph


G(T ) := {(u, T u) : u ∈ Dom(T )}
is a closed subspace of H ⊕ H.
e) A densely defined operator T in H is closable if the (norm) closure of G(T ) in
H ⊕ H is the graph of an operator T . In that case T is a closed operator and it is
the minimal closed extension of T .
Exercise 2.39. a) Show that for every densely defined operator T in H we
have
(Im T )⊥ = Ker T ∗ ,
exactly as for bounded operators (see proof of Theorem 2.7).
b) Show that a closed operator, while not necessarily continuous, at least has some
‘decent’ ([332]) limit behavior: If (un ) is a sequence in Dom(T ) converging to some
u ∈ H and if (T un ) converges to some v ∈ H, then u ∈ Dom(T ) and T u = v.
Moreover, show that this limit behavior characterizes a closed operator.
c) Conclude that an everywhere defined closed operator is bounded.
d) More generally, let T be a closed operator in H with Dom(T ) = W ⊆ H. Show
that the inner product of the Hilbert (sub)space G(T ) ⊆ H ⊕ H induces an inner
product on W which makes W a Hilbert space. Conclude that T : W → H can be
considered as a bounded operator if W is equipped with the corresponding norm
induced by the graph of T .
e) Show that for every densely defined operator T in H the adjoint T ∗ is a closed
operator and we have an orthogonal decomposition
H ⊕ H = G(T ) ⊕ U G(T ∗ ),
where U is the unitary anti–involution on H ⊕ H given by U (w, v) = (−v, w).
Furthermore, T is closable iff T ∗ is densely defined, and in that case T = T ∗∗ .
f) Let H be a Hilbert space with u1 , u2 , . . . a complete orthonormal system (i.e., an
orthonormal basis for H). Let (λj )j=1,2... be an unbounded sequence of numbers.
Find a linear subspace D ⊂ H such that the multiplication operator Mid is closed,
given by the domain D and the operation
X X
D3u= cj uj 7→ Mid (u) := λj cj uj .

g) Prove: If T is a densely defined, closed operator in H, and T is injective with


dense range, then the same properties hold for T ∗ and for T −1 , and
(T ∗ )−1 = (T −1 )∗ .
[Hint: Deduce (a) from (2.22). For (c), apply the closed graph theorem. In (d), the
inner product in W underlying the graph norm
q
(2.23) kukT := k((u, T u)kG(T ) = kukH 2 + kT uk 2 , u ∈ W
H

induced in W by T is an immediate generalization of the inner product (2.20) of


the first Sobolev space. Note that the converse is not valid, i.e. not every bounded
2.6. UNBOUNDED OPERATORS 37

operator from such graph–norm equipped W to H can be considered as a closed


operator in H with domain W . E.g., if W ⊂ H is dense, then
Id |W = Id |H ,
so Id |W is not closed.
For (e), prove first G(T )⊥ = U G(T ∗ ). Since U is unitary, it follows that T ∗ is
closed. This gives the decomposition, but there is more to prove. Help can be
drawn from [332, Theorem 5.5]. For (f) recall that each element u in H has the
form X
u= dj uj ,
where the sum converges in the norm induced by the inner product. Taking inner
products with the uj ’s you see that the coordinates for u are determined by dj =
hu, uj i. Note also the Parseval identity (a generalization of Pythagoras’ theorem)
X
kuk2 = |dj |2 ,
obtained by computing hu, ui. See also Exercise 2.50c. A safety net for (g) is
provided in [332, Proposition 5.1.7].]
We shall dwell a little more on the example of the standard first order differ-
ential operator d/dθ over the circle S 1 . Our goal is to prove that the operator T
which acts like d/dθ on the domain W 1 (S 1 ) is a closed operator in L2 (S 1 ). There
is nothing surprising in the result; no interesting consequences are attached to it;
but it permits us to introduce and illustrate some of the most basic concepts of the
analysis of elliptic differential operators (to be developed below in Part II).

Recall the well-known fact that the functions


1
(2.24) ek (θ) = √ eikθ , k∈Z

form a complete orthonormal system (= orthonormal basis) for L2 (S 1 ) (see also
Appendix A below). It follows that the space C ∞ (S 1 ) is dense in the space L2 (S 1 ).
Clearly, the space C ∞ (S 1 ) can be considered as the space of smooth complex–valued
functions on the real line of period 2π.
Let u ∈ L2 (S 1 ). We shall use the notation
(2.25) b(k) := hu, ek i ,
u k∈Z
for the Fourier coefficients.
Recall the definition of the first Sobolev space W 1 (S 1 ) from above. Note that
the inner product defined in (2.20) is exactly the inner product induced by the
graph G(T ) ⊂ H ⊕ H. For u ∈ W 1 (S 1 ) we write
q
(2.26) kukW 1 := kukL22 + kT ukL22 ,
where T : W 1 (S 1 ) → L2 (S 1 ) extends d/dθ on C 1 (S 1 ), as in the following.
Theorem 2.40. a) The operator d/dθ on C 1 (S 1 ) has a unique continuous (i.e.,
bounded) extension T : W 1 (S 1 ) → L2 (S 1 ), where W 1 (S 1 ) is the first Sobolev space
W 1 (S 1 ) (see (2.20)).
b) The W 1 –norm and the Fourier coefficient norm
r
1 X
(2.27) kuk1 := √ b(k)2 (1 + k 2 )
u
2π k∈Z
38 2. ANALYTIC METHODS. COMPACT OPERATORS

are equivalent norms on W 1 (S 1 ).


c) The first Sobolev space W 1 (S 1 ) is contained in the Banach space of continuous
functions C 0 (S 1 ) and the inclusion is continuous.
d) For each θ ∈ S 1 , the Dirac distribution
(2.28) δθ : C 0 (S 1 ) 3 u 7→ u(θ) ∈ C
extends to a continuous mapping from W 1 (S 1 ) to C (i.e., a bounded operator of
rank 1 or continuous linear functional).
e) The normalized integration operator
Z θ
0 1 θ
(2.29) C (S ) 3 v 7→ v(s)ds − J (v) ∈ C 1 (S 1 )
0 2π
with Z 2π
J (v) := v(s) ds ,
0
extends to a bounded operator S : L2 (S 1 ) → W 1 (S 1 ). Moreover, up to a finite rank
operator, the operator S is a right and left inverse of T (such a quasi–inverse is
called a ‘parametrix’):
1
(2.30) T ◦ S = IdL2 − J and S ◦ T = IdW 1 −δ0 .

f ) As an operator in L2 (S 1 ), the differentiation operator T (with W 1 (S 1 ) ⊂


L2 (S 1 ) as its domain) is closed. Both its kernel and cokernel are one–dimensional,
so its index vanishes.
g) The range Im(T ) is closed in L2 (S 1 ).
Proof. of (a): For u ∈ C 1 (S 1 ) we have
kT ukL22 = hu0 , u0 iL2 ≤ hu, uiL2 + hu0 , u0 iL2 = kukW
2
1 .

So, T is continuous as a mapping from the dense subset C 1 (S 1 ) ⊂ W 1 (S 1 ) to


L2 (S 1 ), and so uniquely extends to W 1 (S 1 ) as a bounded operator.
of (b): We have
hek , ek iW 1 = 1 + k 2 ,
hence X
2
kukW 1 = hu, uiW 1 = u(k)|2 (1 + k 2 ).
|b
k∈Z
of (c): For θ ∈ S 1 and u ∈ C ∞ (S 1 ) we have
1 X 1 X
|u(θ)| = √ b(k)eikθ ≤ √
u |bu(k)|
2π k∈Z 2π k∈Z
1 X 1 1
=√ u(k)|(1 + k 2 ) 2 (1 + k 2 )− 2
|b
2π k∈Z
r r
1 X X
≤√ u(k)|2 (1 + k 2 )
|b (1 + k 2 )−1 ,
2π k∈Z k∈Z

where the last inequality is Schwarz’ inequality


p p
|ha, bi| ≤ ha, ai · hb, bi
in the Hilbert space L2 (S 1 ) (or, correspondingly in the Hilbert space `2 of square–
1 1
b(k)(1 + k 2 ) 2 ek and b := (1 + k 2 )− 2 ek .
P P
summable sequences) for a := u
2.6. UNBOUNDED OPERATORS 39

Clearly b ∈ L2 (S 1 ) since (1 + k 2 )−1 < ∞. To see a ∈ L2 (S 1 ), we apply the


P
preceding result (b) to u ∈ C ∞ (S 1 ) ⊂ W 1 (S 1 ). Applying (b) once again yields
sup |u(θ)| ≤ CkukW 1
θ∈S 1
where the constant C does not depend on u. We extend the estimate to the whole
W 1 (S 1 ) by density.
of (d): Clearly, δθ is continuous on C 0 (S 1 ) for each θ ∈ S 1 . Then the assertion
follows from (c).
of (e): Let v ∈ C 0 (S 1 ) and k ∈ Z, k 6= 0. Then
Z 2π Z θ ! Z 2π
1 J (v)
hSv, ek iL2 = √ v(s)ds e−ikθ dθ − θe−ikθ dθ
2π 0 0 (2π)3/2 0
Z 2π " Z ! #2π
θ
−1 1 −ikθ 1 1 −ikθ
=√ v(θ) e dθ + √ v(s)ds e
2π 0 −ik 2π 0 −ik
0
Z 2π  2π
J (v) 1 −ikθ J (v) 1 −ikθ
+ e dθ − θ e
(2π)3/2 0 −ik (2π)3/2 −ik 0
Z 2π
−i i
= √ v(θ)e−ikθ dθ + √ (J (v) − 0)
k 2π 0 k 2π
 2π
iJ (v) 1 −ikθ iJ (v)
+ 3/2
e − (2π − 0)
k(2π) −ik 0 k(2π)3/2
i
= − vb(k),
k
because the third term vanishes and the second and forth cancel each other. Hence
X 1
hSv, SviL2 = v (k)|2 2 + |hSv, e0 i|2 .
|b
k
k6=0

By partial integration and Schwarz’ inequality we obtain


Z 2π Z θ ! Z 2π
1 J (v)
|hSv, e0 i| = √ v(s)ds dθ − θdθ ≤ C0 kvkL2 .
2π 0 0 (2π)3/2 0
Clearly,
!
Z θ
θ 1
(2.31) T (Sv) (θ) = T v(s)ds − J (v) = v(θ) − J (v).
0 2π 2π
We estimate √
|J (v)| ≤ 2πkvkL2
and obtain
1
k(Sv)0 kL2 = kT SvkL2 ≤ kvkL2 + |J (v)| ≤ C1 kvkL2 .

So,
hSv, SviW 1 = kSvk2L2 + kT Svk2L2
 
X 1
= v (k)|2 2 + C02 kvk2L2 + C12 kvk2L2
|b
k6=0 k
≤ (1 + C02 + C12 )kvk2L2 .
40 2. ANALYTIC METHODS. COMPACT OPERATORS

Moreover, for u ∈ C 1 (S 1 ) we have


Z θ Z 2π
θ
(2.32) S(T u) (θ) = u0 (s)ds − u0 (s)ds = u(θ) − u(0).
0 2π 0
The assertion follows by a density argument.
of (f ): Let (un ) be a sequence in W 1 (S 1 ). We assume that (un ) and (T un ) converge
in L2 (S 1 ). We denote the limits by u, respectively w. To prove that T is closed
we have to show (by Exercise 2.39b) that u ∈ W 1 (S 1 ) and T u = w. From the
preceding (e) we obtain
k(un − um ) −(un (0) − um (0)) kW 1 = k(S ◦ T )(un − um )kW 1 ≤ CkT (un − um )kL2 .
This shows that the sequence (vn := un − un (0)) is convergent in W 1 (S 1 ) to a
limit v. So, the sequence(un − vn = un (0)) has a limit in L2 (S 1 ), hence in C, and
so also in W 1 (S 1 ). This proves that the sequence (un ) converges in W 1 (S 1 ) to u.
So, by (a), T un converges to T u and T u = w.
Clearly, Ker(T ) ⊇ {constants} and, by (2.32) Ker(T ) ⊆ Ker(ST ) = {constants}.
So,
Ker(T ) = {constants} = Coker(T ),
and hence index(T ) = 0.
of (g): Let (vn = T un ) be a sequence in Im T which converges in L2 (S 1 ) to a
v ∈ L2 (S 1 ). Without loss of generality we may assume that δ0 (un ) = 0 for all
n = 1, 2, . . . . According to (e), the operator S is bounded. So, the sequence
(un = S(T un ) + δ0 (un ) = S(T un )) converges in W 1 (S 1 ) to a u ∈ W 1 (S 1 ). Then
by the continuity of T we obtain v = T u. 
Remark 2.41. We gave a fully detailed proof of the preceding theorem because
the 1–dimensional case study provides a preview for the analytical part of this book:
each of the statements will be reproved (and the underlying concepts generalized) in
Part II by replacing S 1 by an arbitrary n–dimensional compact manifold without
boundary (i.e. closed manifold) and d/dθ by an arbitrary elliptic differential or
pseudo-differential operator of order m, acting on sections of a complex vector
bundle. More precisely, (a) will be generalized in Theorem 9.3 (p.237), (b) in
Exercise 7.2b (p.195), (c) in Theorem 7.13 (p.200), (d) in Theorem 7.14 (p.201),
(e) in Theorem 9.8 (p.239), (f) once again in Theorem 9.3 (p.237) and Step (iv) of
the cobordism proof in Section 3 (p. 303), and (g) in Theorem 9.10a (p.240).
In (e), we have
1 X
(T u)(θ) = √ (ik) eikθ u
b(k) and
2π k∈Z
 
1 X −i ikθ
(Sv)(θ) = √ vb(0) + e vb(k)
2π k6=0 k

for u ∈ Dom(T ) = W 1 (S 1 ) and v ∈ Dom(S) = L2 (S 1 ). This is a special way of


writing T and S as pseudo–differential operators with amplitudes
−i
t(θ, k) = ik and s(θ, k) = ,
k
respectively, see Chapter 8. We are only interested in the case k 6= 0, and note that
then t and s are invertible (namely 6= 0) and inverses of each other.
On arbitrary closed manifolds we will not have such simple global descriptions
of differential and pseudo–differential operators. However, the symbolic calculus
2.6. UNBOUNDED OPERATORS 41

we are going to develop below in Chapter 8 will provide us with globally defined
principal symbols (= the top order homogenous part of the amplitudes). We will
characterize an elliptic operator S by the invertibility of its principal symbol s and
construct a parametrix T from the inverse t of s (Section 9.2, p. 239ff).
A closer look at the proof of Assertion (c) reveals that the first Sobolev space
W 1 (S 1 ) coincides with the space of absolutely continuous functions on the real line
with period 2π and with first derivative belonging to L2 (S 1 ) (see also [352, p.257
and note p.305]).
Assertion (f) can also be proved in a different way, namely by showing that
T = R∗ with densely defined R. Actually, R = −T . Therefore the index of T must
vanish. This argument will be made precise below. The index must also vanish for
topological reasons because the dimension of S 1 is odd (see Result (a) in Section
13.3, p.317).
In continuation of Exercise 2.37, p.34, and the preceding Theorem 2.40, we close
this Section with an exercise of a system of r, r ∈ N, linear ordinary differential
equations of first order on S 1 = [0, R]/(0 ∼ R) with R ∈ R, R > 0 fixed. Like
before, one may take R = 2π and the angle θ as coordinate. Here, however, we
prefer general R and a real x as coordinate for better recalling elementary results for
systems of ordinary differential equations. In particular, we denote by C ∞ (S 1 , Cr )
the space of smooth r-vector valued functions on R of period R.
Exercise 2.42. Denote by gl(r, C), r ∈ N, the space of complex r × r matrices.
Let A : R → gl(r, C) be a smooth mapping with A(x + R) = A(x) for all x ∈ R.
Consider the operator (∇ + A)|S 1 : C ∞ (S 1 , Cr ) → C ∞ (S 1 , Cr ), where
d r times d
∇ := ⊕ ... ⊕ .
dx dx
a) Show that the operator (−∇ + A∗ )|S 1 is formally adjoint to (∇ + A)|S 1 . Here
A∗ (x) denotes the adjoint matrix of A(x), x ∈ R.
b) Denote by Φ : Cr → C ∞ ([0, ∞), Cr ) the fundamental solution for ∇ + A which
assigns a solution f ∈ C ∞ ([0, ∞), Cr ) to each initial value f (0) ∈ Cr , and define a
linear mapping φ := Φ|x=R : Cr → Cr by assigning f (0) 7→ f (R). Show that the
periodic solutions correspond to the fixpoints of φ.
c) Denote by Ψ the fundamental solution for −∇ + A∗ and define a corresponding
linear mapping ψ : Cr → Cr . Show that φ and −ψ −1 are adjoint.
d) Prove the Index Theorem on the Circle, namely
dim ker(∇ + A)|S 1 = dim ker(−∇ + A∗ )|S 1 , i.e., index(∇ + A)|S 1 = 0.

[Hint: To a) Define a Hermitian inner product on Cr by


Xr
hu, vi := ui vi for u = (u1 , . . . ), v = (v1 , . . . ).
i=1
Then deduce that the operators (∇ + A)|S 1 and (−∇ + A∗ )|S 1 are formally adjoint
from
Z R Z R
!
h(∇ + A)f, gidx − hf, (−∇ + A∗ )gidx = hf (R), g(R)i − hf (0), g(0)i = 0
0 0
for all f, g ∈ C ∞ (S 1 , Cr ).
To b) Clearly a periodic solution yields a fixpoint of φ. A delicate argument is
42 2. ANALYTIC METHODS. COMPACT OPERATORS

needed to show that each fixpoint yields a smooth periodic solution.


To c) To begin with, consider two functions f, g on R with values in Cr and two
points x0 < x1 . If (∇ + A)f = 0 and (−∇ + A∗ )g = 0, you get
hf (x1 ), g(x1 )i = hf (x0 ), g(x0 )i.
Put x0 = 0 and x1 = R and re-write
hf (R), g(R)i = hφ(f (0), g(R)i and hf (0), g(0)i = hf (0), ψ −1 g(R)i.
To d) If we do not require the periodicity condition, the solution uniquely exists
for a given initial value at a point. Hence the dimension of the space of periodic
solutions is at most r. Notice that, according to changes of A, the dimension of the
space of periodic solutions varies and takes values between 0 and r.]
Note. Alternatively, we could apply the heavy analysis tools of our Part II
to get the mentioned Index Theorem on the Circle. First we would determine the
principal symbol of (∇ + A)|S 1 , e.g., by the formula of Exercise 6.37, p. 183. Our
data are x0 ∈ R, ξ ∈ Tx∗0 S 1 , and e ∈ Cr . We represent ξ = df |x0 by a function f
vanishing at x0 like, e.g., f (x) = ξ(x−x0 ), and extend e to a vector-valued function
g, e.g., the constant function g := e1 on S 1 . Then we find
σ1 (∇ + A)(x0 , ξ)(e) = i(∇ + A)(f g)|x0
= i∇(ξ(x − x0 )e1)|x0 + iA(x)(ξ(x − x0 )e1)|x0
= iξe + iA(x0 )(0e) = iξe = (iξI r )e.
Here I r ∈ GL(r, C) denotes the identity. Whence, σ1 (∇ + A)(x0 , ξ) = iξI r is
invertible for ξ 6= 0, Next, we conclude that (∇ + A)|S 1 is an elliptic differential
operator on a closed manifold and has finite index which depends only on the
homotopy type of the principal symbol. All that will be explained in Part II.
Finally, we notice that the choice of A does not inflict the principal symbol. We
choose A = 0 constant and obtain index ∇ = 0. That suffices. As mentioned above
in a similar situation, this is in fine agreement with the basic topological insight that
any homogeneous polynomial elliptic symbol over an odd-dimensional manifold can
be deformed into the identity within the class of elliptic symbols.
Closed (not necessarily bounded) Fredholm Operators. It is easy to
generalize the concept of Fredholm operators to the unbounded case.
Definition 2.43. Let H be a complex separable Hilbert space. A linear (not
necessarily bounded) operator F with domain Dom(F ), null space Ker(F ), and
range Im(F ) is called Fredholm if the following conditions are satisfied.
(i) Dom(F ) is dense in H.
(ii) F is closed.
(iii) Both dim Ker(F ) and dim Coker(F ) are finite. The difference of the di-
mensions is called index(F ).
Then it follows that the range Im(F ) of F is a closed subspace of H (i.e., F
is normally-solvable) and that dim Ker(F ∗ ) is finite. So, a closed operator F is
characterized as a Fredholm operator by the same properties as in the bounded
case (see also Exercise 2.39a):
(iv) For any arbitrary v ∈ H the equation F u = v, v ∈ Dom(F ), is solvable
if and only if v is orthogonal to every solution w of F ∗ w = 0: differently
speaking, Im(F ) is closed.
2.6. UNBOUNDED OPERATORS 43

(v) Each of the equations F u = 0, F ∗ w = 0 has only finitely many linearly


independent solutions.
Condition (iv) can be replaced by
(iv’) kF uk ≥ Ckuk for all u ∈ Ker(F )⊥ ∩ Dom(F ) with a positive constant C.
We more generally show
Lemma 2.44. Let T : Dom(T ) → H be a closed (by our definition densely
defined) operator. Then T has closed range if and only if (iv’) is valid.
Proof. Let first condition (iv’) be satisfied. Let (fn )n=1,2,... be a sequence
that converges to f (shortly “fn → f ”) with fn ∈ Im(T ). Choose a sequence
(un ∈ Dom(T ) ∩ Ker(T )⊥ ) with T un = fn . We get un − um ∈ Ker(T )⊥ for all n, m,
and thus by (iv’)
1
kun − um k ≤ kfn − fm k −→ 0, n, m −→ 0.
C
Accordingly, un → u and T un → f which implies that u ∈ Dom(T ) and T u = f ,
i.e., f ∈ Im(T ), since T is closed. Thus Im(T ) is closed.
Now let Im(T ) be closed. Then the mapping
Dom(T ) ∩ Ker(T )⊥ −→ Im(T ) given by u 7→ T u
is surjective and injective for simple reasons. Moreover, both spaces are closed
subspaces of H, and thus can be interpreted as Hilbert spaces. The inverse of this
mapping exists and is a well-defined linear transformation of Im(T ) into Ker(T )⊥ ,
with domain Im(T ). Also this transformation is a closed linear operator of the
Hilbert space Im(T ) into the Hilbert space Ker(T )⊥ , as follows from the fact that
T is closed. Whence it must be a bounded linear transformation by standard
argument. However, this boundedness clearly amounts to the condition (iv’). 
Moreover, as in Theorem 2.10, F is a Fredholm operator if and only if F ∗ is a
Fredholm operator (proving the closedness of Im(F ∗ ) looks demanding, but follows
rather directly from Lemma 2.44), and (clearly) we have
index F = dim Ker F − dim Ker F ∗ = − index F ∗ .
In particular, index F = 0 in case F is self–adjoint (i.e., F ∗ = F , see below).
The composition of (not necessarily bounded) Fredholm operators yields again
a Fredholm operator. More precisely, we have the following composition rule.
Theorem 2.45 (I.Z. Gohberg and M.G. Krein, 1957). If F and G are (not
necessarily bounded) Fredholm operators then their product GF is densely defined
with
(2.33) Dom(GF ) = Dom(F ) ∩ F −1 Dom(G)
and is a Fredholm operator. Moreover,
(2.34) index GF = index F + index G .
Remark 2.46. The proof is considerably more involved than in the bounded
case. Why so? Can’t we consider the operators F, G as bounded Fredholm operators
in the graph norm of (2.23)? Yes. And then derive that the composition is Fredholm
and satisfies (2.34) from our previous product rule for bounded Fredholm operators
44 2. ANALYTIC METHODS. COMPACT OPERATORS

in Exercise 1.10? Yes, indeed. The delicate claim, however, is that the domain
Dom(GF ) of the composition is dense in H.
In the context of index theory, the product of closed unbounded Fredholm
operators shows up
(1) for elliptic differential operators over a closed (i.e., compact and without
boundary) manifold M and
(2) for elliptic differential operators over a compact manifold with smooth
boundary, subject to some regular (elliptic) boundary conditions.
The first case was touched upon in Exercise 2.37, p.34f and Theorem 2.40, p.37
with M = S 1 and is the subject of the main body of this monograph for arbitrary,
more general and more specific closed M . As we shall see in Part II, an elliptic
differential operator F of order k ≥ 1, acting between sections of vector bundles E1
and E2 over M can be considered both as a closed unbounded operator from L2 (E1 )
to L2 (E2 ) with domain Dom(F ) = W k (E1 ) ⊂ L2 (E1 ) or as an indexed family of
bounded operators Fr : W (E1 ) → W r−k (E2 ) for all real r, where W r (E1 )
r

denotes the rth Sobolev space (see Chapters 7 and 9). Combining F with a second
elliptic differential operator G, say of order m, with dom(G) = W m (E2 ) ⊂ L2 (E2 )
yields the domain
Dom(GF ) = Dom(F ) ∩ F −1 Dom(G)
= W k (E1 ) ∩ F −1 (W m (E2 )) = W k (E1 ) ∩ W m+k (E1 ) = W m+k (E1 ),
which is dense in L2 (E1 ) by definition. Here we exploited that the Sobolev spaces
can be defined via elliptic operators (along the lines of Exercise 7.2b, p. 195f).
Whence, in view of our set-up of Part II, there is nothing surprising in the preceding
theorem for elliptic operators on closed manifolds. Even in that case, however, the
result is not trivial.
The second case is much more intricate: We have to deal with a boundary
condition for the combined elliptic differential operator obtained by combining two
possibly radically different regular (elliptic) boundary conditions. Solely by classical
analysis arguments it might be difficult to prove the regularity of the combined
boundary condition.
Surprisingly, we can prove the density of Dom(GF ) ⊂ H by purely (but, ad-
mittedly, somewhat wired) functional analysis arguments, following [177] and [112,
Lemma 2.3 and Theorem 2.1].
We prepare the proof by a series of small lemmata.
Lemma 2.47. Let D be a dense subspace of H and Mn a closed subspace of
H of finite codimension n ∈ N. Then there exists a bounded idempotent (= not
necessarily orthogonal projection) P , such that
Im(P ) = Mn , dim Ker(P ) = n, Ker(P ) ⊂ D.
Recall from Remark 2.2, p.13, that subspaces of finite codimension are not
necessarily closed.
Proof. To begin with, choose a basis u1 , . . . , un of Mn⊥ . Then select v1 , . . . , vn ∈
D sufficiently close to the start basis (say kui − vi k < δ for all i = 1, . . . , n for suf-
ficiently small δ > 0), such that

(2.35) det hui , vj i i,j=1,...,n 6= 0.
2.6. UNBOUNDED OPERATORS 45

In particular, this implies linear independence of v1 , . . . , vn . Let Q denote the


space spanned by v1 , . . . , vn . Clearly Q ⊂ D. Then for every x ∈ H there exists
a unique decomposition
Pn x = w + v with w ∈ Mn and v ∈ Q. Indeed, we simply
set v := i=1 α i v i where the coefficients αi are uniquely determined by the
orthogonality relations
X n
x− αi vi , uj = 0, j = 1, . . . , n.
i=1
Note that the determinant of the preceding system of linear equations does not
vanish, by (2.35). Now define P x := w. This defines a projection operator
satisfying all claims stated above. In particular, P is bounded with
X
kP xk = kwk = kx − vk ≤ kxk + kvk = kxk + C |αi | ≤ (1 + C 0 )kxk. 
The claim of following lemma becomes wrong, if we drop the assumption of
finite codimension and closedness of M . Take, for instance, H := L2 ([a, b]), D a
dense subset of step functions, and M∞ := C 0 ([a, b]). For finite codimension of
M , the lemma seems obvious. However, to prove it rigorously we depend on the
preceding result.
Lemma 2.48. Under the assumptions of Lemma 2.47, the intersection Mn ∩ D
is dense in Mn .
Proof. For given u ∈ Mn and given approximation radius δ > 0, we set
−1
δ 0 := δ 1 + kId −P k and select an x ∈ D such that ku − xk ≤ δ 0 . As before,
we decompose x = w + v with w ∈ Mn and v ∈ Q ⊂ D, whence w = P x ∈ D, and
ku − wk ≤ ku − xk + k(Id −P )xk
= ku − xk + k(Id −P )(x − u)k ≤ 1 + kId −P k δ 0 = δ.


Now we can draw the decisive consequence for the composition of (not neces-
sarily bounded) Fredholm operators in Hilbert space.
Lemma 2.49. If F and G are (not necessarily bounded) Fredholm operators
then their product GF is densely defined and Fredholm.
Proof. 1. We first show that Dom(GF ) = {u ∈ Dom(F ) : F u ∈ Dom(G)}
(as defined in (2.33)) is dense in H. Since by assumption
dim Ker(F ∗ ) = dim Im(F )⊥ < ∞,
the space Im(F ) ∩ Dom(G) is dense in Im(F ) by the preceding lemma. Let now
u ∈ Dom(F ), then there exists u e ∈ Ker(F )⊥ ∩ Dom(F ) with F ue = F u, and we
have u − u e ∈ Ker(F ) ⊂ Dom(GF ). Since Im(F ) ∩ Dom(G) is dense in Im(F ), for
every δ 0 > 0 there exists a v ∈ Im(F ) ∩ Dom(G) such that kv − F uk < δ 0 . Let
v = F w with w ∈ Ker(F )⊥ , then w ∈ Dom(GF ). We get u e − w ∈ Ker(F )⊥ , and
0 0
thus keu − wk ≤ CkF u − vk ≤ Cδ . Set u := w + u − u e, and choose δ 0 such
0 0 0
that Cδ = δ, then we have ku − u k < δ and u ∈ Dom(GF ), which means that
Dom(GF ) is dense in Dom(F ). Since Dom(F ) is dense in H, we see that Dom(GF )
is dense in H.
2. Next we show that GF is closed. Suppose the sequence (un )n=1,2,... converges to
u (once again, shortly “un → u”) and GF un → v. We decompose F un = wn + zn
with wn ∈ Ker(G)⊥ and zn ∈ Ker(G). By the closedness of Ker(G)⊥ = Im(G∗ )
46 2. ANALYTIC METHODS. COMPACT OPERATORS

(see the reformulation (iv’) above), we have wn → w ∈ Ker(G)⊥ . Also either znk →
z ∈ Ker(G) for a suitable subsequence (znk ) or kzn k → ∞, since dim Ker(G) < ∞.
In the first case we get unk → u, F unk → w + z, GF unk → v, i.e., u ∈ Dom(GF )
and GF u = v, since F and G are closed operators. In the second case, let xn :=
un /kzn k, then
wn zn
xn −→ 0, F xn = + , GF xn −→ 0.
kzn k kzn k
But zn /kzn k must have a convergent subsequence, and we must get xnk → 0,
F xnk → w with kwk = 1, which is a contradiction, since F is closed. This proves
that GF is closed.
3. Now we show that GF satisfies condition (iv’) (before Lemma 2.44, p.43).
Indeed, let un ∈ Dom(GF ) ∩ Ker(GF )⊥ , kun k = 1, and GF un → 0. Then we again
write F un = wn +zn with wn ∈ Ker(G)⊥ and zn ∈ Ker(G). We get wn → 0 by (iv’)
for G, and again either znk → w or kzn k → ∞. In the first case we get GF unk → 0,
Gunk → w, and unk ∈ Ker(F )⊥ , i.e., unk → u, kuk = 1, u ∈ Dom(GF )∩Ker(GF )⊥ ,
and GF u = 0, a contradiction. In the second case, set again xn := un /kzn k. The
sequence (F xn ) must have a convergent subsequence, and thus we get
xnk −→ 0, F xnk −→ w, kwk = 1
a contradiction, because F is closed. This proves (iv’) for GF .
4. Finally, it is clear that
dim Ker(GF ) ≤ dim Ker(F ) + dim Ker(G) < ∞
and
codim Im(GF ) = dim Ker((GF )∗ ) = dim Ker(F ∗ G∗ )
≤ dim Ker(F ∗ ) + dim Ker(G∗ ) < ∞.
Here we apply that F ∗ G∗ is also closed with closed range (by the same arguments)
and we have (GF )∗ = F ∗ G∗ . 
Proof of Theorem 2.45. The preceding lemma yields the delicate result,
namely that GF is closed with closed range and that the in (2.33) defined Dom(GF )
is dense in H. We leave it to the reader to count the dimensions for the proof of
(2.34), respectively refer to [112, p.699] for the details of that counting. 
Exercise 2.50. Find out which of the following operators are (not necessarily
bounded) Fredholm operators in the sense of Definition 2.43.
a) (Bounded) Fredholm operators in the sense of Section 1.2?
b) Operators of finite rank?
c) The multiplication operator Mid of Exercise 2.39f for λj := j, j = 1, 2 . . . ?
d) The operator T which extends (to W 1 (S 1 )) the differentiation operator d/dθ on
C 1 (S 1 ) (Exercise 2.37)?
e) The Laplace operator ∆ on the unit disk in the plane (see Exercise 5.9, p. 145).
[Answer: (a) Yes. (b) Never. (c) Yes: for
nX∞ X∞ o
Dom(Mid ) := cj uj : j 2 |cj |2 < +∞
j=1 j=1

you obtain a densely defined closed operator which is injective and surjective. (d)
Yes, if H = L2 (S 1 ) and the domain of T is taken to be W 1 (S 1 ) ⊂ L2 (S 1 ) (as in
Theorem 2.40a). (e) Yes, if the domain is extended to the second Sobolev space
2.6. UNBOUNDED OPERATORS 47

and restricted by elliptic boundary conditions (see Exercise 5.9, p. 145). If the full
space of smooth functions on the disk is taken as the domain of the Laplacian
and no boundary conditions are imposed, the kernel of the Laplacian is infinite-
dimensional, consisting of all harmonic functions on the disk.]
Symmetric and Self–Adjoint Operators. As seen immediately after Def-
inition 2.43, the index of unbounded self–adjoint Fredholm operators vanishes like
in the bounded case. So, there is no immediate index problem. However, spectral
projections of self-adjoint Fredholm operators defined by the Spectral Theorem 2.61
(below on p.51) are of high interest in index theory, in particular for specifying el-
liptic boundary value problems. As a matter of fact, the most interesting operators
of geometry and gauge–theoretical physics are Laplacians or Dirac type operators
which are symmetric operators.
Index problems arise from symmetric operators in two ways:
(1) We may occasionally split a self–adjoint Fredholm operator P into the direct
sum of chiral nonsymmetric components
0 P−
 
P =
P+ 0
with P − = (P + )∗ and investigate the index of P + . Since
⊥
Ker(P ) = Ker(P + ) ⊕ Ker(P − ) and Ker(P − ) ∼
= Im(P + ) ,
the integer index(P + ) = n+ − n− gives the chiral asymmetry of Ker(P ); here
n± := dim Ker(P ± ). In Part III, Section 13.8, Example 13.11 and Results d and
f, we shall discuss a basic model of quantum chromodynamics and show how chiral
asymmetry appears for the symmetric euclidean Dirac operator in 4 dimensions
with —- more or less natural — boundary conditions imposed. This is also the way
one follows in most geometric and topological applications (e.g., when determining
the Euler characteristic or the signature of even–dimensional closed Riemannian
manifolds, see Part III, Section 13.4, p. 319ff).
(2) More generally, index problems arise from symmetric (i.e., formally self-
adjoint) elliptic operators on compact manifolds with boundary when non–self–
adjoint elliptic boundary conditions are imposed. In 2 dimensions, we shall give a
simple example for that in Part II with the Noether-Hellwig-Vekua Theorem 5.11
(pp.146ff). For more elaborated examples in odd dimensions see our Section 13.8,
Results d, f and g, mostly based on [83, Theorem 22.24].
For these perspectives and also for the completeness of our presentation, we
summarize the basic knowledge of symmetric and self–adjoint unbounded operators
and introduce to the corresponding Fredholm theory.
Definition 2.51. a) We say that a densely defined operator S in H is sym-
metric (or formally self–adjoint) if
hSu, vi = hu, Svi, u, v ∈ Dom(S).
b) Recall that a densely defined operator S in H is called self–adjoint if S ∗ = S
(in the sense of Definition 2.38c).
c) A densely defined symmetric operator S in H is called essentially self–adjoint
if its closure is self–adjoint, i.e. if S = S ∗ .
We give an interesting criterion for proving that a symmetric operator is essen-
tially self–adjoint. See also and similarly, [352, Theorem VIII.3, Corollary, p. 257]
48 2. ANALYTIC METHODS. COMPACT OPERATORS

and, differently, [305, Lemma 8.14] — Mizohata requires dense Im(T ) in H and
the existence of a positive constant a such that kT uk ≥ akuk for all u ∈ Dom(T ).
Lemma 2.52. Let T be a densely defined symmetric operator in a separable
complex Hilbert space H. We assume that both Im(T + i) and Im(T − i) are dense
in H. Then T is essentially self–adjoint.
Proof. First we show that Im(T ± i) = H. Let (un ) be a sequence in Dom(T )
and (T ± i)un converges to v0 . As with T , T is also symmetric, and we have
k(T ± i)uk2 = kuk2 + kT uk2 for all u ∈ Dom(T ± i) = Dom(T ),
since the mixed terms h±iu, T ui and hT u, ±iui cancel each other by symmetry.
Hence
kuk ≤ k(T ± i)uk ,
i.e., T ± i injective and (T ± i)−1 well defined on Im(T ± i) and bounded. We
conclude that (un ) converges to some u0 and T un converges too. Since T is closed,
u0 ∈ Dom(T ) and (T ± i)u0 = v0 . Thus Im(T ± i) is closed, so Im(T ± i) = H.
Now we show that T is self–adjoint. By definition, T is the minimal closed

extension, so T ⊆ T ∗ . Therefore it suffices to show that Dom(T ) ⊆ Dom(T ).

Then, let u ∈ Dom(T ). Since Im(T + i) = H, there is a w ∈ Dom(T ) so that
∗ ∗ ∗
(T + i)w = (T + i)u. But Dom(T ) ⊆ Dom(T ), so u − w ∈ Dom(T ) and

(T + i)(u − w) = 0 .

Since Im(T − i) = H, we have Ker(T + i) = {0} , so u = w ∈ Dom(T ). 
From the first part of the preceding proof we can distil
Corollary 2.53. Let T be a closed, injective operator. If T is bounded from
below by the identity, then T is a semi–Fredholm operator.
Here we used the following notation:
Definition 2.54. a) A densely defined closed operator T is called a semi–
Fredholm operator if and only if Im(T ) is closed and either Ker(T ) or Coker(T )
has finite dimension.
b) A densely defined, symmetric operator T is bounded from below by the
identity if we have
kuk2 ≤ hT u, ui for all u ∈ Dom(T ),
and consequently kuk ≤ kT uk.
We apply the lemma for H = L2 (S 1 ) and take for T the operator id/dθ with
domain C ∞ (S 1 ). Integration by parts shows that id/dθ is symmetric. For ek (θ) :=
√1 eikθ (k ∈ Z), we have (T ± i)(ek ) = (−k ± i) ek (or (T ± i)( 1 ek ) = ek ) and
2π −k±i
so Im(T ± i) is dense, since it contains the linear span of the complete orthonormal
system {ek : k ∈ Z}. Thus we have proved
Theorem 2.55. The operator id/dθ in L2 (S 1 ) with domain C ∞ (S 1 ) is sym-
metric and essentially self–adjoint.
We have shown in Theorem 2.40a that d/dθ on C 1 (S 1 ) extends to a continuous
operator T : W 1 (S 1 ) → L2 (S 1 ). Regarding T as an unbounded operator on L2 (S 1 )
with dense domain W 1 (S 1 ), we have already shown that T is closed (see Theorem
2.6. UNBOUNDED OPERATORS 49

2.40f). Indeed, T is the closure of d/dθ with domain C 1 (S 1 ) ⊂ L2 (S 1 ). Applying


the preceding theorem yields
Corollary 2.56. On the circle S 1 , the unbounded operator T on L2 (S 1 )
(which extends id/dθ on C 1 (S 1 )) with domain W 1 (S 1 ) is self–adjoint.
Remark 2.57. a) The arguments of the preceding proof can be generalized
to prove the well–known but seldom explicitly stated fact that each symmetric
elliptic differential operator on a closed Riemannian manifold (or, more generally,
on a complete Riemannian manifold) is essentially self–adjoint (see though [391,
Theorem 8.3], [60, Lemma 3.23] and, implicitly, [167, Lemma 1.6.3] or, differently
and carried through only for the Laplacian, [412, Proposition 8.2.4]).
b) Note the remarkable contrast with the case where the underlying manifold has
a boundary. Then there is a huge variety of domains to which a fixed symmetric
elliptic differential operators can be extended such that it becomes self–adjoint; and
there is a smaller, but still large variety where the extension becomes self–adjoint
and Fredholm. For the case of a 1–dimensional manifold (= the interval) see below
Exercise 2.58g–h. For the higher dimensional case we refer to [76, Section 3] and,
differently, [83, Chapter 20]. See also the review in Section 13.8 below (pp.337–356).
In [77, Proposition 7.15] it is shown that the set of all extensions of a given
symmetric elliptic differential operator T0 of first order over a compact smooth
Riemannian manifold X with boundary Y can be naturally identified with a dense
graded subspace β(T0 ) of the distribution space W −1/2 (Y ). It turns out that β(T0 )
carries a natural structure of a symplectic Hilbert space, i.e., a (real) Hilbert space
with a skew-symmetric, nondegenerate and bounded bilinear form ω, here induced
from the principal symbol of T0 over Σ in normal direction.2 Then all self–adjoint
extensions of T0 (in the underlying L2 –space) correspond to the Lagrangian sub-
spaces of β(T0 ) and the self–adjoint Fredholm extensions correspond to the La-
grangian subspaces which form a Fredholm pair with the canonical Lagrangian
subspace (the Cauchy data space). A Fredholm pair is a pair of closed subspaces
with finite-dimensional intersection and with sum of finite codimension (then the
difference of these two dimensions is called the index of the pair; see [83, Chapter
24] with a mild correction given in [78, Appendix] regarding the characterization
of Fredholm pairs with vanishing index).
c) A classical source to the systematic study of all self–adjoint extensions of a given
symmetric operator is the célèbre paper [319]. For supplementary studies, involv-
ing also the deficiency indices (i.e. the codimension of the range of T ± i) and
relations to index theory, we refer to [10, Section 78], [177] (summarized in [12]),
and, more recently, [276, Chapter 4] and [99].
Exercise 2.58. a) Show that a densely defined operator S in H is symmetric
if and only if S ⊆ S ∗ . So, in particular, each self–adjoint operator is symmetric.
b) Prove another criterion for S being symmetric, namely that hSu, ui ∈ R for
every u ∈ Dom(S).
c) Let S be a symmetric operator in H. Show that the two usual conditions (i.e.,

S = S and S = S ∗ ) for S being essentially self–adjoint are equivalent.
d) Let S be a densely defined operator in H which is essentially self–adjoint. Show
that then S is the only self–adjoint extension of S. Show that the converse is also
2The concept of symplectic manifolds is mentioned in Remark 6.9, p.165, and introduced
rigorously in Definition 18.23, p.652. They are one of the major objects of Seiberg-Witten Theory.
50 2. ANALYTIC METHODS. COMPACT OPERATORS

true, i.e., if S has one and only one self–adjoint extension, then S is essentially
self–adjoint.
e) Let S be a densely defined, symmetric operator in H. Show that S is self–adjoint
if and only if S ± i are both surjective operators.
f) Show that the operator Mid of Exercise 2.50c is self–adjoint.
g) Show that the operator id/dx in L2 ([0, 1]) with domain C0∞ ([0, 1]) (= smooth
functions with support in the interior (0, 1) of the interval) is symmetric but not
essentially self–adjoint.
h) Let P, Q be continuous real–valued functions on the interval [0, 1]. Let u 7→ u0
denote the differentiation. Show that the operator
u 7→ (P u0 )0 + Qu
in L2 ([0, 1]) with domain C0∞ ([0, 1]) is symmetric but not essentially self–adjoint.
i) Consider the multiplication operator Mid (u(x) 7→ xu(x)) and the differential
operators of (g) and (h) on L2 (R) with domain C0∞ (R) (= smooth functions with
compact support). Show that these operators are essentially self–adjoint.
[Hint: For (a) use (2.22). For (b) establish
3
X
4hu, vi = ik hu + ik v, u + ik vi
k=0

by straightforward calculation. For (c) recall that always S ⊆ T implies T ∗ ⊆ S ∗ ;


that every symmetric operator S is closable, because S is densely defined and
S ⊆ S ∗ ; and that S = S ∗∗ for each closable S. For (d) note first that every
symmetric operator is closable because S is densely defined and S ⊆ S ∗ , whence
S ⊆ S ⊆ S ∗ . Then use that S ⊆ R implies R∗ ⊆ S ∗ . A safety net for (e) is provided
e.g. in [332, Proposition 5.2.5], see also the second part of our proof of Lemma
2.52. (f) can be proved directly or by applying (e).
For (g) and (i) recall that the support of a function u ∈ C 0 (X) is by definition
the smallest closed subset of X outside which u vanishes identically: supp u :=
{z ∈ X : u(z) 6= 0}. Here X is a topological space, C 0 (X) denotes the set of com-
plex valued continuous functions on X, and L denotes the closure of L ⊂ X.
In (g) and (h) the symmetry follows from partial integration. Candidates for non–
uniquely determined self–adjoint extensions are provided in (g) by the unitarily
twisted periodic boundary conditions u(1) = eiϕ u(0) for all ϕ ∈ [0, 2π); and in
(h) e.g. by the Dirichlet boundary condition u(0) = u(1) = 0 and the Neumann
boundary condition u0 (0) = u0 (1) = 0. For(i) see also the extensive discussion in
[10, Sections 49, 77].]
Spectral Theory. The following definition extends the notion of resolvent set
and spectrum commonly defined for elements in B(H); cf. Table 1.1, p. 11.
Definition 2.59. a) For an operator T in H we define the resolvent set
Res(T ) as the set of all λ ∈ C for which the operator T − λ Id is bijective from
Dom(T ) onto H with bounded inverse.
b) The complement of the resolvent set is called the spectrum of T , and is denoted
by Spec(T ).
c) The function R(λ) := (T − λ Id)−1 defined on C \ Spec(T ) with values in B(H)
is called the resolvent function.
Exercise 2.60. a) Show that the resolvent set of any operator in H is open.
2.6. UNBOUNDED OPERATORS 51

b) Let T be a closed operator in H and λ0 ∈ C. Assume that the resolvent


R(λ0 ) exists and is compact. Show that then the spectrum of T consists entirely
of countably many isolated eigenvalues with finite multiplicities and without finite
accumulation point, and the resolvent R(λ) is compact for all λ ∈ Res(T ). If T is
self–adjoint, conclude that T is diagonalizable (= discrete), i.e., the eigenvectors
form a basis of Dom(T ).
c) Prove that the spectrum of a self–adjoint operator in H is a nonempty, closed
subset of R.
d) Prove the following spectral characterization of (not necessarily bounded) self–
adjoint Fredholm operators: A self–adjoint operator T has discrete spectrum of
finite multiplicity in a neighborhood of 0 ⇐⇒ Ker(T ) is finite–dimensional and
Im(T ) is closed.
e) Let S be a self–adjoint operator in H with compact resolvent. Show that S and
each bounded self–adjoint perturbation (i.e. operators of the form S + C where
C ∈ B(H) and C = C ∗ ) are Fredholm operators in the sense of Definition 2.38.
[Hint: For (a) cf. [332, Proposition 5.2.11] or [358, Exercise 13.17, p.365]. For (b)
cf. [242, Theorem III.6.29]; see also our Section 2.2 on compact operators (p. 18ff),
in particular Theorem 2.20. See also our discussion of the Green’s function for the
Sturm-Liouville problems in the preceding section. Note that the eigenfunctions
of the Green’s operator and of the Sturm-Liouville problem coincide, while the
eigenvalues of the Green’s operator are the reciprocals of the eigenvalues of the
Sturm-Liouville problem (in that case 0 is not an eigenvalue); see [111, p.194]. For
(c) cf. [242, Section V.3.5] or [332, Proposition 5.2.13]. For (d) cf. [133, Def.
XIII.6.1 and Thm. XIII.6.5]. Dunford and Schwartz show that if T is self–
adjoint, then λ is an isolated point of the spectrum of T if and only if Im(T − λ)
is closed. (Note that [133] define the essential spectrum differently). For (e) apply
(c) and (d).]

For a normal (not necessarily bounded) operator, there is a famous theorem


which expresses the operator as an integral of the coordinate function over the
operator’s spectrum with respect to a projection-valued measure. It supplements
the previously proven spectral decomposition for compact operators, Theorem 2.20,
p.20. We shall formulate it only for self-adjoint operators.

Theorem 2.61 (Spectral Theorem). Let T be a densely defined self-adjoint


operator in a complex separable Hilbert space H. Then there exists a uniquely
determined spectral measure E on the Borel subsets of R such that
R
(1) T = λ∈Spec(T ) λ dE(λ),
(2) E(M ) = 0 for all M ⊂ R with M ∩ Spec(T ) = ∅,
(3) E(M ) 6= 0 for all open M ⊂ R with M ∩ Spec(T ) 6= ∅.

Consequently, we can associate a well-defined operator f (T ) to T for each


function f that is integrable on Spec(T ). That result is trivial for a polynomial
f (with id(T ) = T as√ in (1) and 1(T ) = Id), but rather advanced for such simple
functions like f := · yielding an ultra-short proof of the Square Root Lemma. For
comparison, see our elementary but lengthy proof of that Lemma below on p.54f.
Different variants of the Spectral Theorem and a variety of different proofs are in
the literature. We recommend [332, Theorems 4.4.1 and 5.3.8] for the bounded
and for the not necessarily bounded case.
52 2. ANALYTIC METHODS. COMPACT OPERATORS

Metrics on the Space of Closed Operators. Recall that for a fixed separa-
ble complex Hilbert space H, we denoted by B(H) the algebra of bounded operators
from H to H. It is naturally equipped with the metric defined by the operator norm
kT − Sk.
We shall denote by C(H) the space of closed densely defined operators in H.
Clearly the operator norm does not make sense for unbounded operators. However,
for S, T ∈ C(H) the orthogonal projections PG(S) , PG(T ) onto the graphs of S, T
in H ⊕ H are bounded operators and
γ(S, T ) := kPG(T ) − PG(S) k
defines a metric for C(H), the projection metric.
It is also called the gap metric and it is (uniformly) equivalent with the metric
given by measuring the distance between the (closed) graphs. For details and the
proof of the following Lemma and Theorem, we refer to [112, Section 3].
Lemma 2.62. For T ∈ C(H) the orthogonal projection onto the graph of T in
H ⊕ H can be written (where RT := (I +T ∗ T )−1 ) as
RT T ∗ T ∗ RT ∗ T ∗ RT ∗
     
RT RT RT
PG(T ) = = = .
T RT T RT T ∗ T RT T T ∗ RT ∗ T RT I −RT ∗
Theorem 2.63 (H.O. Cordes, J.P. Labrousse, 1963). a) The space B(H)
of bounded operators on H is dense in the space C(H) of all closed operators in
H. The topology induced by the projection (∼= gap) metric on B(H) is equivalent to
that given by the operator norm.
b) Let CF(H) denote the space of closed (not necessarily bounded) Fredholm opera-
tors. Then the index is constant on the connected components of CF(H) and yields
a bijection between the integers and the connected components.
Exercise 2.64. Consider the multiplication operator Mid of Exercise 2.50c
and let Pj denote the orthogonal projection of H onto the linear span of the j-
th orthonormal basis element uj . Clearly the sequence (Pj ) does not converge in
B(H) in the operator norm. Show that, however, the sequence (Mid − 2jPj ) of
self–adjoint Fredholm operators converges in C(H) in the projection metric to Mid .
[Hint: On the subset of self–adjoint (not necessarily bounded) operators in the
space C(H), the projection metric is uniformly equivalent to the metric γ given by
γ(T1 , T2 ) := k(T1 + i)−1 − (T2 + i)−1 k ,
(cf. [79, Theorem 1.1]).]
Remark 2.65. The results by Heinz Cordes and Jean–Philippe Labrousse
may appear to be rather counter-intuitive. For (a), it is worth mentioning that the
operator-norm distance and the projection metric on the set of bounded operators
are equivalent, but not uniformly equivalent since the operator norm is complete,
while the projection metric is not complete on the set of bounded operators. Ac-
tually, this is the point of the first part of (a), see also the preceding exercise.
Assertion (b) says two things: (i) that the index is a homotopy invariant,
i.e. two Fredholm operators have the same index if they can be connected by
a continuous curve in CF(H); (ii) that two Fredholm operators having the same
index always can be connected by a continuous curve in CF(H). Both results are
also true in the category of bounded Fredholm operators. Actually, topologically
much farther reaching results for bounded Fredholm operators are shown in Chapter
2.7. TRACE CLASS AND HILBERT-SCHMIDT OPERATORS 53

3. For investigations of the topology of the subspace of self–adjoint (not necessarily


bounded) Fredholm operators we refer to [79] and [277].
The delicacy of Assertion (b) is partly due to the delicacy of varying domains.
However, if we fix a self–adjoint operator T with compact resolvent and dense
domain D ⊆ H and make D into a Hilbert space by the operator norm (along the
lines of Exercise 2.39d), one may investigate all closed Fredholm operators in H
with that domain D. A reasonable guess is that this space can be identified with
the full space of bounded Fredholm operators by identifying D with H, and that
this bijection is a homeomorphism.

7. Trace Class and Hilbert-Schmidt Operators


Here we give a rigorous definition and exposition of the fundamentals of trace
class and Hilbert-Schmidt operators on a general (complex, infinite-dimensional)
Hilbert space H. For an advanced reader, our presentation may seem a bit convo-
luted with all the small definitions, lemmata, propositions and theorems patched
together. We refer such a reader to [332, Section 3.4] where all we need is done
simply and directly. In this section, however, we prefer to confront our primary
readership with all the details of the involved calculations instead of hiding them
in general structural concepts and theorems.
We have already seen the important special case of Hilbert-Schmidt integral
operators with square-integrable kernels on function spaces (see Exercises 2.29 and
2.30, p. 25). We will also need trace class operators later in Section 3.9 when we give
a construction of the determinant line bundle over the space of Fredholm operators
with index
 0. A connection between determinants and traces is seen in the formula
det eA = eTr A for a A ∈ GL(N, C), or the related formula (see Proposition 3.44,
p. 96)
XN
Tr Λk A ,

(2.36) det(Id +A) =
k=0
where Λk A is the extension of A to the k-th exterior product Λk (Cn ) of Cn . For
those not familiar with exterior products, let λ1 (A) , . . . , λN (A) be the eigenvalues
of A, repeated according to algebraic (as opposed to geometric) multiplicity. In
view of Jordan canonical form, we have the formula (equivalent to (2.36)
YN XN  X 
(2.37) det(Id +A) = (1 + λj (A)) = λi1 (A) · · · λik (A)
j=1 k=0 hiik
P
where hii denotes the sum over all indices 0 < i1 < i2 < · · · < ik ≤ N ; this inner
k 
sum is in fact Tr Λk A . The k = 1 term is the trace of A, namely
XN
(2.38) Tr(A) = λi (A) ,
i=1
and in general the k-th term is the elementary symmetric
 polynomial of degree k
in the λi (A), which again is the same as Tr Λk A . The formula (2.36) extends
to operators on a Hilbert space H when A is trace class. Indeed, it is essentially
the Definition 3.45 (p. 98) of det(Id +A) that we adopt. The equivalence of (2.36)
and (2.37) for trace class operators A is true,
√ but not trivial (see [393]). Roughly,
A ∈ B(H) is trace class
P √ if the square-roots λn of the eigenvalues λn ≥ 0 of A∗ A
are summable (i.e., n λn < ∞). However, for this one would need to assume
that there exists a complete eigenbasis of A∗ A, and thus some assumptions on A
(e.g., compactness) would have to be made. We would rather have the compactness
54 2. ANALYTIC METHODS. COMPACT OPERATORS

of A emerge from our definition of trace class than assume compactness as part of
the definition. In the course leading to the definition we adopt, we begin with some
useful, basic definitions and results of independent interest.
The following Square Root Lemma is fundamental for establishing the Polar
Decomposition of bounded operators in Proposition 2.77, which plays a prominent
role in establishing the fact that the set I1 of trace class operators is closed under
addition (see below). It is a simple consequence of the Spectral Theorem 2.61,
p.51. Since we have not proven the Spectral Theorem, we give an elementary,
though elaborate (i.e., a bit lengthy) proof of the Square Root Lemma.
Theorem 2.66 (Square Root Lemma).√If C ∈ B + , then √ there is a unique
S ∈ B + , such that S 2 = C. Denoting S by C, the map C → C is continuous
in operator norm.

Proof. The power series for 1 − x about x = 0 is
X∞ X∞ (−1)n Yn−1
an xn = 1 + 1
− k xn .

(2.39) p(x) := 2
n=0 n=1 n! k=0

For n ≥ 1, the coefficients an are all negative. Thus,


X∞ X∞ √
|an xn | = lim 1 − 1 − x = 1.

|an | = lim
n=1 x→1− n=1 x→1−

The Weierstrass
√M -test then implies that p(x) converges uniformly and absolutely
for kxk ≤ 1 to 1 − x. Let
B1+ := B ∈ B + : kBk ≤ 1 .


Since X∞
n
A ∈ B1+ =⇒ kan An k ≤ |an | kAk ≤ |an | and |an | < ∞,
n=1
P∞
the Weierstrass M -test implies that n=0 an An converges uniformly on B1+ to a
continuous function p : B1+ → B, namely
X∞
p(A) := an An (for A ∈ B1+ ).
n=0
P∞
The convergence of n=0 kan An k also implies that the formal squaring and re-
arrangement of the series p(A) yields the expected result
2
(2.40) p(A) = Id −A (for A ∈ B1+ ).
Note that p(A) ≥ 0, since (using kAk ≤ 1)
X  X
k k Xk n 2 2
|an | An x, x ≤ |an | kAn xk kxk ≤ |an | kAk kxk ≤ kxk ,
n=1 n=1 n=1

which implies that


 Xk 
n
hp(A) x, xi = lim Id − |an | A x, x
k→∞ n=1
X 
2 k
n
= kxk − lim |an | A x, x ≥ 0.
k→∞ n=1

If C ∈ B1+ , we have Id −C ∈ B1+ , since


2 2
h(Id −C) x, xi = hx, xi − hCx, xi ≥ kxk − kCxk kxk ≥ kxk (1 − kCk) ≥ 0,
2.7. TRACE CLASS AND HILBERT-SCHMIDT OPERATORS 55

and then using Proposition 2.21 (p. 20), we get


kId −Ck = sup {h(Id −C) x, xi : x ∈ H, kxk = 1}
= sup {1 − hCx, xi : x ∈ H, kxk = 1} ≤ 1.
2 2
Thus in (2.40) we may take A = Id −C ∈ B1+ and (p(Id −C)) = p(A) = Id −A =
C. Hence, for C ∈ B1+ and kCk ≤ 1, a positive square root of C is p(Id −C). For
arbitrary C ∈ B + with C 6= 0, note that Id −C/ kCk ∈ B1+ and
  2
p C C
kCk p Id − = kCk = C.
kCk kCk
Then there is a positive square root S(C) of C ∈ B + , namely
(
p(Id−C) ,  if kCk ≤ 1,
S(C) := p C
kCk p Id − kCk , if kCk > 0.
√ √ p √ p √
Since αp(1 − x/α) = α 1 −(1 − x/α) = α x/α = x, for 0 ≤ x ≤ α ≤ 1,
the two formulas agree for 0 < kCk ≤ 1. The first p formula yields the continuity
of S on B1+ , in particular at C = 0. Since C 7→ kCk and C 7→ Id −C/ kCk are
continuous functions on B \ {0}, the second formula yields the continuity of S on
the rest of B + .
We now prove the uniqueness property. Suppose that R is another positive
square root of C. Then RC = R3 = R2 R = CR, so that R commutes with C and
hence with S(C) which (aside from the trivial case C = 0) can be written as a series

of powers of Id −C/ kCk. To show that R = S(C), we show Ker(S(C) − R) = {0}.
Since RS(C) = S(C) R, we have (S(C) + R)(S(C) − R) = 0 which implies that

(S(C) − R)(H) ⊆ Ker(S(C) + R) , and so (S(C) − R)(H) ⊆ Ker(S(C) + R) .


Since S(C) − R is self-adjoint, we then have
⊥ ∗
Ker(S(C) − R) = (S(C) − R) (H) = (S(C) − R)(H) ⊆ Ker(S(C) + R) .
Thus,

x ∈ Ker(S(C) − R) =⇒ hS(C) x, xi + hRx, xi = h(S(C) + R) x, xi = 0
=⇒ hS(C) x, xi = 0 and hRx, xi = 0,

since S(C) ≥ 0 and R ≥ 0. Now,


D E
2 2 2
hRx, xi = 0 =⇒ 0 = S(R) x, x = kS(R) xk =⇒ Rx = S(R) x = 0,

and similarly hS(C) x, xi = 0 ⇒ S(C) x = 0. Thus,


(S(C) − R)(x) = S(C) x − Rx = 0 − 0 = 0,

and so x ∈ Ker(S(C) − R) ∩ Ker(S(C) − R) = {0}. 

Corollary 2.67. Every A ∈ B is a C-linear combination of two self-adjoint


operators, and any self-adjoint operator is a C-linear combination of two unitary
operators.
56 2. ANALYTIC METHODS. COMPACT OPERATORS

Proof. Note that A + A∗ and i(A − A∗ ) are self-adjoint, and


A = 21 (A + A∗ ) − 2i (i(A − A∗ )) .

If B ∈ B is self-adjoint and kBk = 1, then Id −B 2 ≥ 0, and Id −B 2 makes sense.
Note that
 p   p 
(2.41) B = 21 B + i Id −B 2 + 21 B − i Id −B 2 ,

where B ± i Id −B 2 is unitary since
 p ∗  p 
B ± i Id −B 2 B ± i Id −B 2
 p  p 
= B ∓ i Id −B 2 B ± i Id −B 2 = B 2 + Id −B 2 = Id .

If 0 6= B ∈ B is self-adjoint, we can replace B by B/ kBk in (2.41) and multiply by


kBk. If B = 0, then B = 0 Id +0 Id. 

Definition 2.68. A ∈ B is trace class, written A ∈ I1 , if there √ is a complete


orthonormal system {e0 , e1 , . . .}, such that for the operator |A| := A∗ A, we have
X∞ X∞ 1 2
Tr |A| := h|A| ei , ei i = |A| 2 ei < ∞.
i=0 i=0
p
More generally, for p ∈ N, A ∈ Ip if |A| ∈ I1 . If A ∈ I2 , then A is called
Hilbert–Schmidt.
Remark 2.69. Note that Tr |A| is independent of the choice of {e0 , e1 , . . .}.
Indeed, if {f0 , f1 , . . .} is another complete orthonormal system, then
X∞ X∞ 1 2 X∞ X∞ D 1 E 2
h|A| ei , ei i = |A| ei =
2
|A| ei , fj
2
i=0 i=0 i=0 j=0
X∞ X∞  D 1 E2  X∞
= |A| 2 fj , ei = h|A| fi , fi i .
j=0 i=0 j=0

Remark 2.70. Note that p ≤ p0 ⇒ Ip ⊂ Ip0 , since


X∞ D p0 E X∞ 1 0 2 X∞  2
p0 −p)
|A| 2(
1 1
p p
|A| ei , ei = |A| 2 ei = |A| 2 ei
i=0 i=0 i=0
X∞  0
2 0 2 X∞ 2
|A| 2(
p −p)
= |A| 2(
1 1
p 1
p −p) 1
p
≤ |A| 2 ei |A| 2 ei
i=0 i=0
0 2
= |A| (
p −p)
1
p
2
Tr |A| .

Recall that K denotes the ideal of compact operators in B. Let Bf denote the
ideal of finite rank operators. We will eventually show (see Proposition 2.75) that
Bf ⊂ Ip ⊂ Ip0 ⊂ K, for 1 ≤ p ≤ p0 ,
but first we prove

Proposition 2.71. Any trace class operator is compact; i.e, I1 ⊂ K.


Proof. For any x ∈ H, we have
X∞  X∞
Ax = A hx, ei i ei = hx, ei i Aei .
i=0 i=0
2.7. TRACE CLASS AND HILBERT-SCHMIDT OPERATORS 57

Thus, we have (as always) that A is the pointwise limit of finite rank operators
X∞ X∞ Xn−1
A= h·, ei i Aei = e∗i ⊗ Aei = lim e∗i ⊗ Aei .
i=0 i=0 n→∞ i=0

To show that A is compact, it suffices to show convergence in norm; i.e.,


Xn−1 2 !
(2.42) lim sup{||(A − e∗i ⊗ Aei )(x) || : x ∈ H, kxk = 1} = 0.
n→∞ i=0

2
First, note that A ∈ I1 ⇒ |A| ∈ I1 , since
X∞ D 2 E X∞ X∞  1 1
2
2
|A| ei , ei = k|A| ei k = |A| 2 |A| 2 ei
i=0 i=0 i=0
X∞  1 1
2 1 2 X∞ 1 2
≤ |A| 2 |A| 2 ei = |A| 2 |A| 2 ei
i=0 i=0
1 2
= |A| 2
Tr |A| .

In the sup of (2.42), we may assume x ∈ Hn := span(e0 , . . . , en−1 ) , since
Xn−1
(A − e∗i ⊗ Aei )|Hn⊥ = 0.
i=0

Thus, it suffices to show that


n o
2
lim sup kAxk : x ∈ Hn , kxk = 1 = 0.
n→∞

2
Note that any x ∈ Hn with kxk = 1 may serve as fn in a complete orthonormal

extension {fi }i=0 of e0 , . . . , en−1 , and so we have
Xn−1 2 2
X∞ 2
X∞ 2

2

kAei k + kAxk ≤ kAfi k = k|A| fi k = Tr |A| .
i=0 i=0 i=0

Thus, as desired,
  Xn−1
2 2 2 2
x ∈ Hn , kxk = 1 =⇒ kAxk ≤ Tr |A| − kAei k =⇒
i=0
n
2 2
o
2
Xn−1 2
lim sup kAxk : x ∈ Hn , kxk = 1 = Tr |A| − lim kAei k = 0. 
n→∞ n→∞ i=0

If A ∈ I1 , then |A| ∈ I1 and hence |A| is compact, self-adjoint and positive.


By the Hilbert-Schmidt Theorem 2.20, there is a complete orthonormal system for

(Ker A) , say {ei : 0 ≤ i < N }, where N may be finite or ∞, such that |A| ei = µi ei
with µi > 0, µ0 ≥ µ1 ≥ µ2 ≥ . . . ≥ 0 (repeated according to multiplicity), and
limi→∞ µi = 0 if N = ∞. The positive eigenvalues µi of |A| are known as the
singular values of A. For x ∈ H,
X  X
N N XN
µi hx, ei i µ−1

Ax = A hx, ei i ei = hx, ei i Aei = i Aei
i=0 i=0 i=0
XN
= µi hx, ei i fi , where fi := µ−1
i Aei .
i=0

Thus, we have the so-called canonical expansion of A


XN
(2.43) A= µi h·, ei i fi ,
i=0
58 2. ANALYTIC METHODS. COMPACT OPERATORS

which (assuming N = ∞) converges in norm to the compact operator A, since


Xn n Xn o
A− µi h·, ei i fi = sup Ax − µi hx, ei i fi : x ∈ H, kxk = 1
n i=0 i=0
o

= sup kAxk : x ∈ span {e0 , . . . en } , kxk = 1
n o

= sup k(|A| x)k : x ∈ span {e0 , . . . en } , kxk = 1 = µn+1 → 0 as n → ∞.
We have
hfi , fj i = µ−1 −1 −1 −1 ∗
i Aei , µj Aej = µi µj hA Aei , ej i = δij , and

(AA∗ ) fi = (AA∗ ) µ−1 −1 ∗ −1 2 2 −1 2


 
i Aei = µi A(A Aei ) = µi A µi ei = µi µi Aei = µi fi .

It follows that {fi : 0 ≤ i < N + 1} is a complete orthonormal system for(Ker A∗ ) =
Im A, and the µi are also the singular values of A∗ , which (noting that fi =
µ−1 −1 ∗
i Aei ⇒ ei = µi A fi ) has the canonical expansion
XN
A∗ = µi h·, fi i ei .
i=0

The following definition then yields Tr A∗ = Tr A, for A ∈ I1 .


PN
Definition 2.72. For A ∈ I1 with canonical expansion A = i=0 µi h·, ei i fi ,
we define
XN XN  XN  X
N
Tr A := hAej , ej i = µi hej , ei i fi , ej = µj hfj , ej i .
j=0 j=0 i=0 j=0

This sum is absolutely convergent, since |µj hfj , ej i| ≤ µj kfj k kej k = µj and
XN XN
µj = h|A| ej , ej i = Tr |A| < ∞.
j=0 j=0
PN
Remark 2.73. If A ∈ I1 and A is self adjoint, then A = i=0 λi h·, ei i ei is
the canonical expansion of A, where {ei } is a complete orthonormal system for

(Ker A) with Aei = λi ei . In this case,
XN XN
Tr A = λj hej , ej i = λj ,
j=0 j=0

which converges absolutely and agrees with (2.38) when N < ∞.


Proposition 2.74. If {gi } is any complete orthonormal system for H and
A ∈ I1 , then X∞
hAgj , gj i = Tr A,
j=0
and the sum is absolutely convergent.
Proof. We first verify that the sum is absolutely convergent:
X∞ X∞ XN 
|hAgj , gj i| = µi hgj , ei i fi , gj
j=0 j=0 i=1
X∞ XN XN X∞
≤ µi |hgj , ei i hfi , gj i| = µi |hgj , ei i hfi , gj i|
j=0 i=1 i=1 j=0
XN X∞  1 X∞ 1
2 2 2 2
≤ µi |hgj , ei i| |hfi , gj i|
i=1 j=0 j=0
XN XN
= µi kei k kfi k = µi = Tr(|A|) = kAk1 .
i=1 i=1
2.7. TRACE CLASS AND HILBERT-SCHMIDT OPERATORS 59

P∞ PN
The absolute convergence of j=0 i=1 µi |hgj , ei i hfi , gj i| just shown allows the
interchange of the sums over i and j in the following:
X∞ X∞ XN 
hAgj , gj i = µi hgj , ei i fi , gj
j=0 j=0 i=1
X∞ XN
= µi hgj , ei i hfi , gj i
j=0 i=1
XN X∞ XN
= µi hgj , ei i hfi , gj i = µi hfi , ei i = Tr A. 
i=1 j=0 i=1

Proposition 2.75. For any p, p0 ∈ N with p ≤ p0 , we have


Bf ⊂ Ip ⊂ Ip0 ⊂ K.
Proof. Since Bf ⊂ Ip is clear, it suffices (by Remark 2.70) to prove that
p p
Ip ⊂ K for any p ∈ N. For A ∈ Ip , we have |A| ∈ I1 . Thus, |A| is compact.
PN p
Let A = i=0 µi h·, ei i ei be the canonical expansion of |A| . Then the canonical
PN 1/p 1/p
expansion of |A| is i=0 µi h·, ei i ei . Since µi → 0, we have µi → 0, and so
|A| is compact. This implies that A is compact. Indeed, since k|A| xk = kAxk, if
{|A| xn } has a convergent subsequence {|A| xni } for any bounded sequence {xn },
then {Axn } has the convergent subsequence {Axni } since it will also be a Cauchy
sequence: Axni − Axnj = |A| xni − |A| xnj → 0 as i, j → ∞. 
The set I1 is clearly closed under scalar multiplication, but less clearly under
addition, as the following exercise suggests.
Exercise 2.76. Show that there are 2×2 matrices A and B for which |A + B|
|A| + |B| (i.e. |A| + |B| − |A + B| is not positive).
To show that nevertheless I1 is closed under addition, it is convenient to first
introduce polar decomposition.
Proposition 2.77 (Polar Decomposition). Any A ∈ B can be uniquely ex-
pressed in the form A = U P , where
P ∈ B is self-adjoint
 and positive (i.e., hP x, xi ≥ 0 for all x ∈ H) with
⊥ ⊥
P (Ker A) ⊆ (Ker A) , and

U ∈ B, Ker U = Ker A, and hU (x) , U (y)i = hx, yi for x, y ∈ (Ker A) (i.e.,

U |(Ker A)⊥ is an isometry of (Ker A) onto U (H)).
Indeed,
√ √
 −1 
(2.44) ∗
P = A A and U = 0Ker A ⊕ A ◦ A A|(Ker A)⊥ ∗ .

Proof. We first establish uniqueness of the decomposition A = U P . Note


⊥ ⊥
that U ∗ U (Ker A) ⊆ (Ker A) , since

x ∈ (Ker A) , y ∈ Ker A =⇒ hU ∗ U (x) , yi = hU x, U yi = 0.
Also,

x, y ∈ (Ker A) =⇒ hU ∗ U x, yi = hU x, U yi = hx, yi =⇒ U ∗ U |(Ker A)⊥ = Id(Ker A)⊥ .

For x ∈ Ker A, (A∗ A) x = 0 = P 2 x, and for x ∈ (Ker A) we have


A Ax = (U P ) (U P ) x = P (U U )(P x) = P ∗ P x = P 2 x,
∗ ∗
60 2. ANALYTIC METHODS. COMPACT OPERATORS


since P x ∈ (Ker A) and U ∗ U |(Ker A)⊥ = Id (Ker A)⊥ . Thus, A∗ A = P 2 and so
√  −1
P = A∗ A by Theorem 2.66. Then A = U P ⇒ U |(Ker A)⊥ = A ◦ P |(Ker A)⊥ ,
and since U |Ker A = 0, we have the uniqueness. Defining P and U by (2.44), we
have A = U P , since
x ∈ Ker A ⇔ hP x, P xi = P 2 x, x = hA∗ Ax, xi = 0 =⇒ U P x = 0 = Ax, and
 −1 

x ∈ (Ker A) =⇒ U P x = A P |(Ker A)⊥ (P x) = Ax.

By definition, A∗ A is self-adjoint and positive. Also, since (A∗ A) |(Ker A)⊥ ∈
 

B (Ker A) is positive, by the uniqueness of positive square roots, we have
√ q
A∗ A = 0Ker A ⊕ (A∗ A) |(Ker A)⊥ ,
√ 


⊥ ⊥
whence A∗ A (Ker A) ⊆ (Ker A) . For x, y ∈ (Ker A) , we have
 √ −1 √ −1 

hU (x) , U (y)i = A ◦ A A|(Ker A)⊥ ∗
x, A ◦ A A|(Ker A)⊥ y
√ √
 −1 −1 
= A∗ A|(Ker A)⊥ A∗ A A∗ A|(Ker A)⊥ x, y
√ −1 √
 −1 
∗ ∗ ∗
= A A|(Ker A)⊥ A A|(Ker A)⊥ A Ax, y
 −1 
= (A∗ A) |(Ker A)⊥ A∗ Ax, y = hx, yi .
 
By definition, Ker A ⊆ Ker U and since we have just seen that Ker U |(Ker A)⊥ = 0,
we have Ker A = Ker U . 
Exercise 2.78. (a) Use polar decompositions of A, B and A + B to show that
for p = 1 and p = 2
1 1 1
p p p
(2.45) A, B ∈ Ip =⇒ (Tr |A + B| ) p ≤ (Tr |A| ) p +(Tr |B| ) p ,
and hence that A + B ∈ Ip . By Proposition 2.74 Tr : I1 → C is then linear. [Hint:
The case p = 2 is a bit less tricky. As a last resort, see [352, p. 208] for the case
p = 1.]
(b) Show that if A ∈ Ip for p ∈ N and B ∈ B, then A∗ , AB and BA are in Ip .
Moreover, for A ∈ I1 , prove that Tr(AB) = Tr(BA). [Hint: By Corollary 2.67 and
Part (a), we may assume that B is unitary.]
While (2.45) is valid for all p, here we only need it for p = 1 and 2. In general,
p 1
Ip is a normed linear space, the Schatten class, with norm kAkp := (Tr |A| ) p .
In particular, there is a norm (the trace norm) on I1 , given by
XN
(2.46) kAk1 := Tr(|A|) = µi ,
i=0
PN
where A = i=0 µi h·, ei i fi is the canonical expansion of A. We have
2
XN 2 2
kAxk = µ2i |hx, ei i| ≤ µ20 kxk ,
i=0
2.7. TRACE CLASS AND HILBERT-SCHMIDT OPERATORS 61

with equality for x = e0 . Thus,


XN
kAk = µ0 ≤ µi = kAk1 .
i=0
Note that Part (b) of Exercise 2.78 yields that Ip is an ideal of algebra B, whence
the use of the symbol “I”. When p = 2, we now show that kAk2 is in fact the norm
associated with an inner product. Note that Tr(D∗ C) exists for C, D ∈ I2 , since
4D∗ C = (C ∗ + D∗ )(C + D) + i(C ∗ − iD∗ )(C + iD)
−(C ∗ − D∗ )(C − D) − i(C ∗ + iD∗ )(C − iD)
X3 ∗  X3 2
= ik C + ik D C + ik D = ik C + ik D .
k=0 k=0
k 2 ∗
Then by Exercise 2.78, C + i D ∈ I1 and D C ∈ I1 . Thus, we have an inner
product h·, ·iTr on I2 given by
hC, DiTr := Tr(D∗ C) for C, D ∈ I2 ,
 1
2 2
p
with norm hC, CiTr = Tr |C| = kCk2 .

Proposition 2.79. For A ∈ I1 and B ∈ B, we have BA ∈ I1 (by Exercise


2.78) and
kBAk1 ≤ kBk kAk1 .
Proof. Let A = U |A| denote the polar decomposition of A ∈ I1 . Since
1
|A| ∈ I2 ,
2

 1 1
2
2 2
|Tr(BA)| = |Tr(BU |A|)| = Tr BU |A| 2 |A| 2
D 1 ∗ 1
E 2 D 1 1
E 2
= |A| 2 , BU |A| 2 = |A| 2 , BU |A| 2 .
Tr Tr
By the Cauchy-Schwarz inequality for h·, ·iTr ,
D 1 1
E 2 1 2 1 2 1 2
|A| 2 , BU |A| 2 ≤ |A| 2 BU |A| 2 = Tr(|A|) BU |A| 2
Tr 2 2 2
 1
∗ 1
  1 1

∗ ∗
= kAk1 Tr BU |A| 2
BU |A| 2
= kAk1 Tr |A| U B BU |A| .
2 2

For a complete orthonormal system e1 , e2 , . . . , we have


 1 1
 X∞ D 1 1
E
Tr |A| 2 U ∗ B ∗ BU |A| 2 = |A| 2 U ∗ B ∗ BU |A| 2 ei , ei
i=0
X∞ D 1 1
E X∞ D 1 1
E
∗ ∗
= U B BU |A| ei , |A| 2 ei ≤
2
kU ∗ B ∗ BU k |A| 2 ei , |A| 2 ei
i=0 i=0
 X∞ 
∗ ∗ 2
= kU B BU k h|A| ei , ei i = kAk1 kBk .
i=0
2 2 2
Thus, combining the above, |Tr(BA)| ≤ kAk1 kBk or |Tr(BA)| ≤ kAk1 kBk.
Letting BA = W |BA| denote the polar decomposition of BA, we then have
kBAk1 = Tr |BA| = |Tr(W ∗ BA)| ≤ kAk1 kW ∗ Bk = kAk1 kBk . 
Recall that Bf ⊂ B denotes the subspace of finite-rank operators. Relative to
the operator norm, the closure of Bf is the space K of compact
P∞ operators. Since
there are compact operators which are not trace class (e.g., n=1 n1 h·, en i en ), I1
is not a closed subspace of B in the operator norm. However, we have
62 2. ANALYTIC METHODS. COMPACT OPERATORS

Proposition 2.80. (I1 , k·k1 ) is a Banach space. The set Bf of finite-rank


operators is k·k1 -dense in I1 ; i.e., (I1 , k·k1 ) is the k·k1 -completion of Bf .
Proof. Let (An ) be a Cauchy sequence in (I1 , k·k1 ), then (An ) is a Cauchy
sequence of compact operators in (B, k·k) with limit A ∈ K. We need to show
A ∈ I1 with kA − An k1 → 0. Since (An ) is a Cauchy relative to k·k1 , the sequence
(kAn k1 ) is bounded (kAn k1 ≤ kAn − Am k1 + kAm k1 ≤ ε + kAm k1 for n ≥ m,
PN
where m is chosen sufficiently large). Let A = i=0 µi h·, ei i fi denote the canonical
PN
expansion of A. We first show kAk1 = i=0 µi < ∞. Since the case N < ∞ is
clear, let N = ∞. If kAn − Ak → 0, then kA∗n An − A∗ Ak → 0, since
kA∗ A − A∗n An k ≤ k(A∗ − A∗n ) A + A∗n (A − An )k
≤ kA∗ − A∗n k kAk + kA∗n k kA − An k ≤ (2 kAk + 1) kA − An k ,
for n sufficiently large. Using this and Theorem 2.66, we obtain
√ p
(2.47) kAn − Ak → 0 =⇒ k|A| − |An |k = A∗ A − A∗n An → 0.
Then for each (finite) m ∈ N,
Xm Xm Xm
h|A| ei , ei i = lim h|An | ei , ei i = lim h|An | ei , ei i ≤ kAn k1 .
i=1 i=1 n→∞ n→∞ i=1
Hence A ∈ I1 , since
Xm
kAk1 = lim h|A| ei , ei i ≤ sup {kAn k1 } < ∞.
m→∞ i=1 n∈N

By (2.47), we have (as p → ∞)


kA − Ap k → 0 ⇔ k(A − An ) −(Ap − An )k → 0 =⇒ |Ap − An | → |A − An | ,
and so
Xm Xm
h|A − An | ei , ei i = lim h|Ap − An | ei , ei i ≤ lim kAp − An k1
i=1 p→∞ i=1 p→∞
Xm
=⇒ kA − An k1 = lim h|A − An | ei , ei i ≤ lim kAp − An k1 .
m→∞ i=1 p→∞

Given any ε > 0, for n and p sufficiently large we have kAp − An k1 ≤ ε. Hence,
kA − An k1 ≤ ε for n sufficientlyP∞large (i.e., limn→∞ kA − An k1 = 0). As for the
density of Bf in I1 , let A = i=0 µi h·, ei i fi denote the canonical expansion of
A ∈ I1 − Bf . Then as n → ∞,
Xn X∞ X∞
A− µi h·, ei i fi = µi h·, ei i fi ≤ µi → 0. 
i=0 1 i=n+1 1 i=n+1
CHAPTER 3

Fredholm Operator Topology

Synopsis. Calkin Algebra and Atkinson’s Theorem. Perturbation Theory: Homo-


topy Invariance of the Index, Homotopies of Operator-Valued Functions, The Theorem
of Kuiper. The Topology of F: The Homotopy Type, Index Bundles, The Theorem of
Atiyah-Jänich. Determinant Line Bundles: The Quillen Determinant Line Bundle, Fred-
holm Determinants, The Segal-Furutani Construction. Spectral Invariants: Essentially
Unitary Equivalence, What Is a Spectral Invariant? Eta Function, Zeta Function, Zeta
Regularized Determinant.

1. The Calkin Algebra


I So far, we have introduced compact operators for purely practical reasons: Within
pure mathematics, they came from the search for a (closed) class of operators that exhibit
properties analogous to those of the operators of finite rank. In applied mathematics, they
enter through the theory of integral equations associated with the study of oscillations.
Actually, the compact operators have yet a deeper significance in the representation of
Fredholm operators. J

We recall our notation: H is a complex, separable Hilbert space; B denotes


the Banach algebra of bounded linear operators on H (in modern terminology, B is
even a C ∗ -algebra; see Exercise 2.6, where one needs to verify the additional axiom
2
kT ∗ T k = kT k ); K ⊆ B denotes the closed two-sided ideal of compact operators
(see Theorem 2.24); and F ⊆ B denotes the space of Fredholm operators. We begin
with a simple exercise.
Exercise 3.1 (J. W. Calkin [106], 1941). Show that the quotient space B/K,
consisting of equivalence classes π(T ) := {T − K : K ∈ K}, where T ∈ B, forms a
Banach algebra.
[Hint: Since K is a linear subspace, clearly B/K is a vector space. To prove that
B/K is an algebra, one must use the fact that K is a two-sided ideal. Then show
that since K is closed, B/K can be made into a Banach space by defining a norm
on B/K by
kπ(T )k := inf {kT − Kk : K ∈ K} = inf {kRk : R ∈ π(T )} .
It remains to show that
kπ(Id)k = 1 and kπ(T )π(S)k ≤ kπ(T )k kπ(S)k .
To prove the left equation, assume that there is a K ∈ K with kId −Kk < 1
and show that K is invertible (using the argument in the proof of Theorem 2.31
involving geometric series); this contradicts the compactness of K. To prove the
right inequality, apply the trick
inf kT S − Kk ≤ inf k(T − K1 )(S − K2 )k . ]
K∈K K1 ,K2 ∈K

63
64 3. FREDHOLM OPERATOR TOPOLOGY

Theorem 3.2 (F. V. Atkinson, 1951). If (B/K)× is the group of units (i.e.,
elements which are invertible with respect to multiplication) of B/K and π : B →
B/K is the natural projection, then we have
F = π −1 ((B/K)× ).
Exercise 3.3. Show that this theorem of Frederick Valentine Atkinson
can also be written as: An operator T ∈ B is a Fredholm operator exactly when
there are S ∈ B and K1 , K2 ∈ K, such that ST = Id +K1 and T S = Id +K2 . Such
an S is called a parametrix (or quasi-inverse) for T . One also says that T is
essentially invertible; i.e., invertible modulo K.
Exercise 3.4. Suppose that K1 and K2 in Exercise 3.3 are trace class (see
Section 2.7) and self-adjoint. Show that
(3.1) index T = Tr K1 − Tr K2 .
[Hint. Using K1 = ST − Id and K2 = T S − Id, show that T K1 = K2 T and
SK2 = K1 S. Using this, verify that if v (resp. w) is an eigenvector of K1 (resp.
K2 ) with eigenvalue λ, then T v (resp. Sw) is an eigenvector of K2 (resp. K1 ) with
eigenvalue λ. If Vλ (resp. Wλ ) is the eigenspace of K1 (resp. K2 ) for eigenvalue λ,
then check that ST |Vλ = (1 + λ) IdVλ and T S|Wλ = (1 + λ) IdWλ . Conclude that

T |Vλ : Vλ −→ Wλ for λ 6= −1. Also check that Ker T = V−1 and Coker T ∼ = W−1 =
Ker S. Verify that
X X
Tr K2 − Tr K1 = λ dim Wλ − λ dim Vλ ,
λ λ
and all but two desirable terms cancel in the difference of these absolutely conver-
gent sums.]
Proof of Theorem 3.2. For “⊆”, let F ∈ F. We show that π(F ) is invert-
ible. For this, consider the operator F ∗ F + P , where P : H → Ker F is orthogonal
projection. In Remark 2.11 (p.17), we have already shown that Ker F ∗ F = Ker F
and Im F ∗ F = Im F ∗ ; thus, F ∗ F + P is bijective and hence invertible in B. Since
P is compact (being of finite rank), it follows that π(F ∗ F ) = π(F ∗ )π(F ) is in-
vertible in B/K. Similarly, one shows with the help of the orthogonal projection
Q : H → Ker F ∗ that F F ∗ + Q in B and π(F )π(F ∗ ) in B/K are invertible. With a
left-inverse for π(F ∗ )π(F ) and a right-inverse for π(F )π(F ∗ ), it follows that π(F )
is invertible in B/K.
For “⊇”, let T ∈ B with π(T ) invertible in B/K; i.e., there is S ∈ B such that
T S and ST lie in π(Id). Now, π(Id) = {Id +K : K ∈ K)} consists of Fredholm
operators by Theorem 2.31 (indeed, of index zero, but that does not concern us
here). In particular, we then have that Ker ST and Coker T S are finite-dimensional.
Since
Ker T ⊆ Ker ST and Im T ⊇ Im T S,
it follows that T ∈ F. 
Remark 3.5. The trick in the first part of the above proof consists of first
considering F ∗ F and F F ∗ (whose invertibility modulo K is trivial) rather than
F , and only then drawing conclusions about π(F ). This has the advantage that
one need not explicitly exhibit the parametrix (i.e., inverse modulo K) for F . An
explicit, if somewhat cumbersome, proof of the theorem of Atkinson can be found
in [365, Theorems 5.4 and 5.5].
3.2. PERTURBATION THEORY 65

Exercise 3.6. Show that the set of Fredholm operators is open in the Banach
algebra of bounded linear operators on a fixed Hilbert space H. [Hint: Because of
the continuity of π (π is even contracting), it suffices to show that (B/K)× is open
in B/K. For this, show in general that the group of units A× in any Banach algebra
A is open; more precisely, show that about each a ∈ A× there is a ball of radius
1/ a−1 contained in A . For this, apply again the geometric series argument in
the proof of Theorem 2.31 or from Exercise 3.1.]
Exercise 3.7. Conclude from the theorem of Atkinson that the space of
Fredholm operators is closed under composition, the adjoint operation, and addition
of compact operators. Show that such a conclusion is not circular, since the earlier
proofs of the same results (e.g., Exercise 1.6 (p.5) and Theorem 2.10, (p.16)) were
not needed in the proof of Atkinson’s theorem.
Exercise 3.8. Illustrate Atkinson’s theorem with the shift+ operator (in
Exercise 1.3, p. 4) on L2 (Z+ ). In particular, show that the similarly defined shift−
is a parametrix (= an inverse modulo K) for shift+ . Which compact operators do
we get for (shift− ◦ shift+ ) − Id and for (shift+ ◦ shift− ) − Id?
Exercise 3.9. Using the theorem of Riesz (Theorem 2.31, p.26), show that
each parametrix G for a Fredholm operator F is itself a Fredholm operator, and we
have index G = − index F .
Exercise 3.10. From Exercise 3.7, we know already that F is closed under
addition of compact operators. Now show that the index is invariant:
index(F + K) = index F for all F ∈ F and K ∈ K.
[Hint: Show that each parametrix for F is also a parametrix for F + K, and apply
Exercise 3.9.]

2. Perturbation Theory
I The result of Exercise 3.10, which we obtained as an easy corollary of the Theorem
2.31 (p.26) of Riesz and of Theorem 3.2, is also due to Frederick Valentine Atkinson.
It represents a fundamental result of perturbation theory which asks how the properties of
a complicated system are related to those of an ideal system close by whose properties are
more easily computed or known. The idea comes from the variational calculus which asks
the opposite way, namely determining optimal shapes of curves and surfaces (e.g., minimiz-
ing some energy functionals) by comparison with less advantageous neighbors, formalized
in the famous Euler-Lagrange Equations and developed further in Morse Theory. In that
context, the basic idea of homotopy was expressed by the young Giuseppe Lodovico
(Luigi) Lagrangia (Lagrange) in his [271, Second letter to Euler, 12 August, 1755],
to us the birth certificate of deformation theory and differential topology:
“Differentiale ipsius y quatenus hic differentiatur, x manente, pro habendo
maximo, minimove formulae datae valore, ad distinctionem aliarum eius-
dem y differentiarum, quae in illa jam ingrediuntur, denotabo per δ; sic
et δdy est differentia ipsius dy, dum y crescunt quantitate δy; idem dic
generaliter de valore δF y [F y mihi est functio quaecumque (emphasized
by the authors) y].”1

1
Our translation: “I shall denote the (peculiar) derivative of y, which is here to differentiate
to obtain the largest or smallest value of a given formula while x remains unchanged, by δ — to
distinguish it from the other differentiations of that y which already enter that formula; in such a
66 3. FREDHOLM OPERATOR TOPOLOGY

Perturbation theory in a wider sense arose in celestial mechanics which tries to de-
termine the deviations of planetary orbits from the unperturbed Keplerian paths due to
the gravitational forces of other celestial bodies. While the methods used there point in
a different direction, it is the perturbation theory of Lord Rayleigh (concerned with
continuously extended oscillating systems) which leads frequently and typically to opera-
tors perturbed by the addition of a compact operator. This happens for example, when
in elasticity the passage is made from constant mass density to variable density. See for
example [116, I, V.13]. That these are as a rule compact perturbations, is due to the
fact that in the underlying partial and ordinary differential equations the terms of highest
order remain unchanged and only the coefficients of the derivatives of lower order are
modified. A theorem of Franz Rellich (see below Theorem 7.15, p. 201) explains why
this produces compact perturbations.
Quantum mechanics poses farther reaching perturbation problems which are in parts
mathematically unsolved. An example is the quantitative determination of energy levels
of complicated systems of quantum mechanics.
The oscillations and motions of quantum mechanical systems are largely determined
by the eigenvalues and eigenfunctions of the corresponding operators. Therefore, perturba-
tion theory usually amounts to applying approximation methods to solving the eigenvalue
problem of a complicated linear operator T + K which differs little from a simpler T with
a solved eigenproblem. Perturbation theory becomes spectral theory which studies the
different constituents of the spectrum of an operator. A reference is the comprehensive
exposition in [242].2
We will not pursue the physical applications any further here, since there is abundant
motivation for perturbation theory within mathematics. Consider for instance the above
mentioned calculus of variations of which local perturbations are an actual principle, or
geometric questions which ask how much a curve (asymptotically or in its shape) or a
surface, etc., changes if relevant parameters in their equations are modified. In particular,
we are interested in the degree to which our quantitative invariants dim Ker, dim Coker,
and index are independent of “small” perturbations. Here “small” does not exclusively
mean that the dimension of the image of the perturbing operator is small, as with operators
of finite rank and in a sense with compact operators, but may mean the perturbation is
small in operator norm. J

To get a feeling for the complications, we put together a list of results:


1. The group of invertible elements of a Banach algebra is open, by Exercise 3.6.
In particular, for each invertible, bounded, linear operator T , there is an ε ( :=

way δdy denotes the difference just of dy, when (all) the y increase by a value δy; likewise speak
generally of the value δF y [to me, F y is an arbitrary function (emphasized by the authors) of y].”
2
Kato’s perturbation theory is incomparably deeper than our investigation: While we con-
sider a single invariant, the index, Kato’s theory is concerned with countably many real parameters
associated with the power series expansion of the eigenvalues of a perturbed (symmetric) operator
T + cK where the parameters depend analytically on the perturbation. Just as one can classify
symmetric matrices in linear algebra
• -according to their rank
• -projectively, according to their index of inertia (Sylvester index) , and
• -orthogonally, according to their diagonal elements (after principal axis transformation)
we have in the perturbation theory of operators in Hilbert space several levels of stability: in-
dex/essential spectrum (see below)/perturbation parameters of the power series expansion. In
the crude mirror of finite-dimensional linear algebra, Kato’s theory is closest to the principal axis
transformation, while we restrict ourselves in index theory to consideration of the rank.
3.2. PERTURBATION THEORY 67

−1
T −1 ) such that for all S ∈ B with kSk < ε, we have:
(i) T + S ∈ F
(ii) index(T + S) = index T (= 0)
(iii) dim Ker(T + S) = dim Ker T (= 0)
(iv) dim Coker(T + S) = dim Coker T (= 0).
2. Further, by Exercise 3.10, for all T ∈ F and K ∈ K
(i) T + K ∈ F
(ii) index(T + K) = index T .
3. On the other hand, one can always find a perturbation of the identity
by a compact operator K such that dim Ker(Id −K) > 1080 making Ker(Id −K)
unimaginably large, since its dimension could not be matched by the atoms in a
universe of “only” 1011 galaxies. Namely, select an orthonormal basis for H and
define K as the orthogonal projection onto the linear span of the first 1080 + 1
basis elements. However, index(Id −K) = index Id (= 0) by the Riesz Theorem
(Theorem 2.31, p. 26).
4. In each arbitrarily small neighborhood of the zero operator there are Fred-
holm operators (namely, iterates of the shift operators multiplied by a small con-
stant ε) with any large or small index; e.g.,
index(0 + ε(shift+ )k ) = k.
Hence, in the neighborhood of 0, the index behaves (metaphorically) as a holomor-
phic function in the neighborhood of an essential singularity (Theorem of Felix
Casorati and Karl Weierstrass).
5. For the boundary-value problem
u00 + ru = 0, u(0) = u(1) = 0, r ∈ R, r > 0,
treated in Chapter 2 (see Exercise 2.34, p.31), or the equivalent problem
v − rKv = 0,
where Z x Z 1
Kv = (1 − x) yv(y) dy + x (1 − y)v(y)dy,
0 x
it was already shown that
1, for r = n2 π 2 and n ∈ N,

dim Ker(Id −rK) =
0, otherwise.
While dim Ker(Id −rK) is not perturbable if it is zero (this is also clear because
Id −rK is then invertible by the Riesz Theorem (Theorem 2.31, p. 26), whence Ex-
ercise 3.6 or the above result 1 applies), it is very prone to change when r =
n2 π 2 – however, only in one direction: the dimension can only decrease. In
other words, dim Ker(Id −rK) is upper semi-continuous; i.e., dim Ker(Id −rK) ≤
dim Ker(Id −r0 K) for all r sufficiently close to r0 .
6. Closed (not necessarily bounded) Fredholm operators with compact resol-
vent (typically elliptic differential operators of positive order on closed manifolds
or on compact manifolds with smooth boundary subject to suitable boundary con-
ditions) have either discrete spectrum or the whole set C as essential spectrum.
Nonvanishing index implies the second case, by Exercise 3.10, p.65.
68 3. FREDHOLM OPERATOR TOPOLOGY

The following theorem shows, for arbitrary small (in the operator norm sense)
perturbations what we already proved in Exercise 3.10 for compact perturbations:
Even though the dimensions of the kernel of an operator and of its adjoint are not
invariant under perturbations, the two jump by the same amount, so that their
difference (the index) remains constant. The perturbation-invariance of the index
is its most remarkable property. Together with the composition rule (Exercise 1.10,
p.8, or Exercise 2.3, p.14), it shows that the index has properties analogous to
homotopy invariants in algebraic topology such as the Euler characteristic χ(M ) of
a compact manifold M . Indeed, χ(M ) is in fact the index of a certain operator,
namely, d+δ from the space of even differential forms to the space of odd differential
forms on M , see Theorem 13.6b, Formula (13.12).

3. Homotopy Invariance of the Index


After these heuristic considerations, we now come to the aforementioned main
theorem.
Theorem 3.11 (J. Dieudonné, 1943). With regard to the operator-norm topol-
ogy, the mapping index : F → Z is locally constant.
Proof. By Exercise 3.6, we already know that there is a neighborhood in B
which is contained in F. We now amplify the argument used there: By Theorem 3.2
(p. 64), we first choose a parametrix G for F ; i.e., a G ∈ B such that
F G = Id +K1 and GF = Id +K2 ,
−1
where K1 , K2 ∈ K. We now show that for all T ∈ B with kT k < kGk , we have
F + T ∈ F. Recall the geometric series argument (see the hint for Exercise 3.6),
whereby the operators Id +T G and Id +GT are invertible, since kT Gk and kGT k
are less than 1. Thus (Id +GT )−1 G is a left inverse of F + T modulo K, since
(Id +GT )−1 G(F + T ) = (Id +GT )−1 (Id +K2 + GT )
(3.2) = Id +(Id +GT )−1 K2 .
Similarly, G(Id +T G)−1 is a right inverse of F + T modulo K. Thus, by Theorem
3.2, we have F + T ∈ F. Applying the composition rule (Exercise 2.3, p.14)) and
the Riesz Theorem (Theorem 2.31, p. 26), we easily obtain from (3.2) the index
formula
index (Id +GT )−1 + index G + index(F + T ) = 0.


Hence index(F + T ) = index F , since the index of an invertible operator vanishes,


and index G = − index F by Exercise 3.9. 

Remark 3.12. a) As mentioned in Theorem 2.63, the index remains locally


constant (and, in fact, distinguishes the connected components) also in the un-
bounded case, even for varying domains.
b) The following example shows that the index of Fredholm operators in a Fréchet
space is not a homotopy invariant.
Example 3.13. As explained in Appendix A, the functions z 7→ z k , k ∈ Z form
an orthonormal basis for the Hilbert space L2 (S 1 ) of square integrable complex
valued functions on the circle S 1 = {z : |z| = 1}. Let H+ denote the (closed)
3.3. HOMOTOPY INVARIANCE OF THE INDEX 69

subspace spanned by all z k with k nonnegative. Let P+ : L2 (S 1 ) → H+ denote the


orthogonal projection. Then each f ∈ C 0 (S 1 ) induces a bounded operator
Tf := P+ ◦ Mf |H+ : H+ −→ H+ ,
where Mf denotes multiplication by f . We shall see below in Exercise 4.3 that
Tf ∈ B(H+ ) depends continuously of f . Moreover, Theorem 4.4 implies that Tf is
a Fredholm operator if and only if f (z) 6= 0 for all z ∈ S 1 . In that case we have
index Tf = −W (f, 0), where W (f, 0) denotes the winding number of f around the
origin (e.g., W (z k , 0) = k for k ∈ Z).

(S 1 ) := P+ C ∞ (S 1 ) , each f ∈ C ∞ (S 1 ) correspondingly

On the Fréchet space C+
induces a continuous linear operator
∞ ∞
τf := π+ ◦ Mf |C+∞ (S 1 ) : C+ (S 1 ) −→ C+ (S 1 ),

where π+ denotes the restriction of P+ to C+ (S 1 ). Similarly as in the preceding
case, the operator τf depends continuously on f (in the respective Fréchet spaces).
However, the domain of τf is much smaller than the domain of Tf , and so are the
secondary sets. In particular, we now have that τf is Fredholm if and only if f has
not more than a finite number of zeros, each of finite order. If f has no zeros at
all, we still have index τf = index Tf = −W (f, 0). Now, τzk can continuously be
deformed into the identity within the space of Fredholm operators in our Fréchet
space, e.g., by ft (z) := z k − 2t, t ∈ [0, 1], since z k − 2 clearly is homotopic to a
constant function.
Exercise 3.14. Return to the space F of all bounded Fredholm operators in a
fixed Hilbert space. Show that the index is constant on the connected components
of F.
Exercise 3.15. Show that dim Ker : F → N ∪ {0} is upper semi-continuous;
i.e., dim Ker F does not expand suddenly when F is changed continuously. But it
may well shrink suddenly. As an example, consider a continuous path Ft , t ∈ [0, 1],
in F connecting an invertible operator F0 (with dim Ker F0 = 0) to a noninvertible
operator F1 (with dim Ker F1 > 0). More precisely, show that
dim Ker(F + T ) ≤ dim Ker F
for F ∈ F and kT k sufficiently small.

[Hint: Show that Ker(F + T ) ∩(Ker F ) = {0}. Details are found in [365, Proof
of Theorem 5.11].]
I As an aside, we mention that Dieudonné (in [120]) proved Theorem 3.11 only
implicitly without use of the index concept. The numerous interrelations of this theorem
can be seen from the fact that by now a number of quite diverse proofs exist. All proofs
have in common the reduction to the geometric series argument or the openness of the
group of units of a Banach algebra.
This idea is most apparent in [128, p.36f. and pp.133-148] where it is first shown
that A× /A× 0 is discrete where A is an arbitrary Banach algebra (with identity), A
×
its
group of units, and A× 0 is the connected component containing the identity. There is an
abstract index
i : A× −→ A× /A×0
defined in a natural way and whose continuity and hence local invariance is clear from
the definition. The main task consists in making the connection between this ideally
simple algebraic object and the real index. Less algebraic proofs can be found in [234,
70 3. FREDHOLM OPERATOR TOPOLOGY

1970/1982, 5.4], where the reduction to the openness of B× is achieved in a sequence of


explicit extensions and projections which are computed in detail. The trick of fixing one
dimension is carried out particularly elegantly in [20, p.104]. As in [234] and in contrast
to our proof above which is inspired by [109, 12/06] and [365, Theorem 5.11], the Atiyah
proof does not use the nontrivial theorems of Riesz and Atkinson and thus may be the
most transparent proof on the whole. We will render it next. J

Alternative Proof of Theorem 3.11. Let e0 , e1 , ... be a complete ortho-


normal system for the Hilbert space H. We take Hn to be the closure of the linear
span of the ei with i ≥ n, and we let Pn denote the orthogonal projection of H
onto Hn .
Step 1: Clearly Pn is self-adjoint and Pn ∈ F, since Ker Pn and Coker Pn are
finite-dimensional. Hence index Pn = 0, and for each F ∈ F, we then have
index Pn F = index Pn + index F = index F.

Step 2: Since dim Coker F < ∞ (F ∈ F), we have n0 such that e0 , e1 , ..., en0 −1
together with F (H) span H; in particular, for all n ≥ n0 ,
Pn F (H) = Hn and dim Coker Pn F = n.
(Incidentally, we see that dim Coker Pn F and also dim Ker Pn F can be made arbi-
trarily large with n.)
Step 3: Although the function dim Ker is only semi-continuous on F, we claim
that for G sufficiently near to F and n sufficiently large (as in Step 2)
dim Ker Pn G = dim Ker Pn F and dim Coker Pn G = dim Coker Pn F.
For G ∈ B and p : H → Ker Pn F the projection, consider the operator
Gb : H −→ Hn ⊕ Ker Pn F given by Gub := (Pn Gu, pu).

If G = F , then Fb is bijective, and hence has a bounded inverse by the Open


Mapping Principle. Identifying Hn ⊕ Ker Pn F with H, we then have Fb ∈ B × . By
the familiar argument in the hint to Exercise 3.6 (p. 65), there is a neighborhood of
Fb contained in B × (the units, i.e., the invertible operators belonging to the algebra
B). Since “G 7→ G” b is continuous, there is also a neighborhood V of F such
that for all G ∈ V the operator G b is an isomorphism. From the surjectivity of
G,
b it follows that Pn G(H) = Hn , whence dim Coker Pn G = n = dim Coker Pn F.
Moreover, Ker Pn G = G b −1 (Ker Pn F ), since by definition of G,
b a point u is mapped
to Ker Pn F by G exactly when the first component (i.e., Pn Gu) of Gu
b b vanishes.
Since Gb is an isomorphism, we then also have
dim Ker Pn G = dim G b −1 (Ker Pn F ) = dim Ker Pn F,
which establishes the above claim.
Summary: We have shown that for each F ∈ F, there is a natural number n and
η > 0 such that for all G ∈ B with kF − Gk < η, we have
index F = index Pn F = index Pn G = index G.
Here, one could replace “index” by “Ker” or “Coker” in the inner equality, but
not in the outer equalities. Moreover, Im(Pn F ) = Im(Pn G) = Hn and Ker Pn G =
b −1 (Ker Pn F ).
G 
3.3. HOMOTOPY INVARIANCE OF THE INDEX 71

Exercise 3.16. Formulate Theorem 3.11 for a continuous family of Fredholm


operators, by which we mean a continuous map G : X → F, where X is any topo-
logical space. More precisely, show that for all x0 ∈ X, there is a neighborhood U
and a natural number n, such that for all x ∈ U ,
Im Pn G(x) = Hn .
Then prove that the function
dim Ker Pn G : X −→ N ∪ {0}
is constant (say k) on U , and that there are k continuous functions
fi : X −→ H, i = 1, ..., k,
such that for all x ∈ U , the points f1 (x), ..., fk (x) form a basis of Ker(Pn G(x)).
[Hint: Imitate the preceding alternate proof, replacing G by G(x) and F by G(x0 ).
Let f1 (x0 ), ..., fk (x0 ) be a basis of Ker Pn G(x0 ) and set fi (x) := b −1 (fi (x0 )).]
G(x)
I The concept of a continuous family of operators comes from classical analysis
in the investigation of operators depending on a parameter or families of operators. In
the simplest examples, the parameter space X is the unit interval, all R, or a bordered
domain in a higher-dimensional Euclidean space (e.g., the domain of permissible control
variables). With somewhat more complex problems of analysis (e.g., as in the study of
elliptic boundary value problems), we are quickly forced to consider families with more
general parameter spaces: An elliptic differential equation defines a continuous family of
Fredholm operators, where the parameter space is the sphere bundle of covectors of the
underlying manifold restricted to the boundary; see Part III, Section 13.8. We will first
treat these questions not from the standpoint of applications, but rather out of natural
topological-geometric considerations, namely interest in deformation invariants 3. J

We assign to each compact parameter space X a group and to each continuous


family of Fredholm operators
G : X −→ F
we assign a group element, which is indeed invariant under deformation. This means
that another continuous family of Fredholm operators G0 : X → F is assigned to
the same group element, if G and G0 are homotopic. By this, we mean that G can
be continuously deformed to G0 (see Chapter 10); i.e., there is a continuous family
g of Fredholm operators parametrized by the product space X × I, where I = [0, 1],
g : X × I −→ F, such that g|X×{0} = G and g|X×{1} = G0 .
If X consists of a single point, then a continuous family of Fredholm operators
is just a single Fredholm operator, and the homotopy of G and G0 clearly means
3
Motivated by problems of optics (and also questions in astronomy, surveying, and architec-
ture), the projective geometry of the 17th century originated in the idea of searching for properties
of geometric figures which are invariant under transformations (central projection and cross sec-
tion). In addition to these linear transformations, the concept of a deformation (i.e., the continuous
change of a mathematical object) existed, for instance, when Johannes Kepler in 1604 noted
that if the plane is compactified, then ellipse, hyperbola, parabola and circle can be transformed
into one another by a continuous relocation of the foci (See [247, p.299]). But it was not until well
into the 19th and 20th centuries that deformation invariants were found for a greater variety of
mathematical objects. These include homology and cohomology theories, as presented axiomati-
cally in [140] for example, as well as the so-called K-theory, another branch of algebraic topology
which was developed by Michael F. Atiyah and Friedrich Hirzebruch and is specifically aimed
at the needs of analysis; see Part III below.
72 3. FREDHOLM OPERATOR TOPOLOGY

that G and G0 lie in the same (path) component of F. In this way, we can answer
fundamental questions concerning the nature of the connected components of F
(e.g., via approximation theory).
Before we study continuous families of Fredholm operators (i.e., the geometry
of F or the group of units (B/K)× by Theorem 3.11), we first turn to a simpler
problem, the geometric investigation of the group of units B × .

4. Homotopies of Operator-Valued Functions


I This section and the following sections of this chapter are central for understanding
index theory. They may be skipped in first reading and then later read in conjunction
with Part III. In particular, here we use some concepts from topology which will be made
precise only in Part III below. J

Exercise 3.17. Let X and Y be topological spaces and f, g and h continuous


maps from X to Y . Show that if f is homotopic to g and g is homotopic to h, then
f is homotopic to h.
We recall some definitions and introduce some notation.
1. Two continuous maps f and g from X to Y are called homotopic (written
f ∼ g), if one can continuously deform one into the other; i.e., there is a continuous
map F : X × I → Y (where I = [0, 1]) such that
F ◦ i0 = f and F ◦ i1 = g,
where it : X → X × {t} is the canonical inclusion (t ∈ I). We write Ft for F ◦ it ,
and can then roughly regard F as a 1-parameter family (over I) of continuous maps
from X to Y . We call F a homotopy of f to g.
2. By the transitivity (Exercise 3.17) and obvious symmetry and reflexivity of
the relation homotopic, it follows that the homotopy classes
f¯ := {g ∈ C(X, Y ) : g ∼ f }
where C(X, Y ) is the set of continuous functions from X to Y , and the homotopy
set
[X, Y ] := f¯ : f ∈ C(X, Y )

(3.3)
are well defined. Note that [point, Y ] corresponds to the set of pathwise connected
components of Y .
3. Two topological spaces X and Y are homeomorphic, if there is a bijective
map f : X → Y which is continuous in both directions.
4. Two topological spaces are homotopy equivalent, if there are continuous
maps f : X → Y and g : Y → X such that f ◦ g ∼ Id and g ◦ f ∼ Id. Clearly, the
real line R and the plane R2 are homotopy equivalent and have the same cardinality
(i.e., there is a bijection between them), but they are not homeomorphic, as we show
in Part III (or just note that the removal of a point disconnects R but not R2 .
5. Y is called a retract, retraction of X, if Y ⊆ X and there is a continuous
map f : X → Y with f |Y = Id. Such an f is called a retraction. If, in addition,
i ◦ f ∼ Id, where i : Y → X is inclusion, then Y is called a deformation retract
of X, and X and Y are homotopy equivalent. Each P ∈ X (or rather {P }) trivially
constitutes a retract of X (but the sphere, as the boundary of the solid ball, is not
a retract of the ball – see Part III below). If {P } is a deformation retract of X,
3.4. HOMOTOPIES OF OPERATOR-VALUED FUNCTIONS 73

then X is called contractible. The shape of X must then be starlike in a certain


sense.
Here we will study the homotopy type of operator spaces. Let B × (H) denote
the group of invertible operators on a Hilbert space H that we allow to be a finite-
dimensional complex vector space, say CN in which case B × (H) = GL(N, C).
Exercise 3.18. Investigate the group B × (H × H) (or GL(2N, C)) of invertible
operators on the product space H × H, which can be written as 2 × 2 (block)
matrices. For R, S ∈ B × (H) – or more generally for R, S : X → B × (H) continuous
with X a given topological space – show that
   
SR 0 R 0
∼ .
0 Id 0 S

[Hint: Consider the map F : X × [0, π/2] → B × (H × H), which is given by


    
cos t − sin t S 0 cos t sin t R 0
Ft := .
sin t cos t 0 Id − sin t cos t 0 Id
Here, we have for brevity written cos t and sin t for the operators cos t Id and sin t Id
∈ B(H). Show that the image of F really lies in B × (H × H) – not entirely in
B × (H) × B × (H) – and investigate F0 and Fπ/2 . Why can we use an interval [a, b]
(a 6= b) of R different from I in the definition of a homotopy?]

I The trigonometric functions which appeared in the preceding problem are typical of
homotopy investigations of linear spaces in which rotations and compressions or dilations
are the most important deformations. This considerably simplifies the explicit statement
of homotopies. Of course, it does not simplify the demonstration of the nonexistence of
a homotopy since this forces one to consider all homotopies, a task which in general is
solvable only with the crude means of algebraic topology; see Chapter 10 below. J

Recall from linear algebra the fact that the group GL(N, C) of invertible com-
plex N × N -matrices contains the compact subgroup U(N ), where U(N ) consists
of the unitary matrices of rank N ; i.e.,
U(N ) := A ∈ GL(N, C) : A∗ = A−1 ,


where A∗ is the conjugate transpose of A. Such A are matrices of C-linear trans-


formations of CN which preserve the usual Hermitian inner
 product in C
N
. For a
(complex) Hilbert space H, we have the group U(H) := T ∈ B(H) : T = T −1 .

Exercise 3.19. If R, S ∈U(H), show that the homotopy used in Exercise 3.18
does not leave U(H × H).
Exercise 3.20. Regard S 1 := {z : z ∈ C and |z| = 1} as a subset of C× {0}
(via z 7→ (z, 0)) and choose a ∈ S 1 (e.g., a = (1, 0)). Construct a continuous map
g : S 1 −→ U(2)
with the properties
(i) (g(z))(z) = a for all z ∈ S 1
(ii) g ∼ f , where f (z) = Id for all z ∈ S 1 .
[Warning: The exercise would be trivial and solvable without using the 2nd dimen-
sion (i.e., within U(1) rather than U(2)) if S 1 were contractible. In that case the
maps f : z 7→ 1 and g : z 7→ az −1 would be homotopic as maps from S 1 to S 1 (or
74 3. FREDHOLM OPERATOR TOPOLOGY

equivalently U(1)).]
[Hint: Reduce to Exercise 3.19 by setting
az −1
 
0
g(z) = . ]
0 za−1
We show that U(N ) and GL(N, C) are pathwise connected as follows. The
polar decomposition theorem of linear algebra states that every g ∈ GL(N, C) is
+
(uniquely) a product AP , where A ∈ U(N ) and P ∈ HN := the convex space of
positive-definite (and hence invertible), N × N Hermitian matrices. Thus, if U(N )
is path-connected, the multiplication map
+
U(N ) × HN −→ GL(N, C)
+
exhibits GL(N, C) as the continuous image of the path-connected space U(N )×HN ,
and so the path-connectedness of GL(N, C) follows from that of U(N ). To show
that U(N ) is connected, we may proceed as follows. If eN = (0, . . . , 0, 1) ∈ CN ,
then the map
f : U(N ) −→ S 2N −1 given by f (A) = AeN
is a continuous, open surjection with fibers f −1 (f (A)) = AU(N − 1), homeomorphic
to U(N − 1). Suppose that U(N ) is not connected. If U(N ) = V1 ∪ V2 where V1 and
V2 are nonvoid, open disjoint sets, then as S 2N −1 is connected, f (V1 ) ∩ f (V2 ) 6= φ,
say AeN ∈ f (V1 ) ∩ f (V2 ). Thus
AU(N − 1) = (V1 ∩ AU(N − 1)) ∪(V2 ∩ AU(N − 1)) ,
which implies that AU(N − 1) is not connected. Continuing, we arrive at the con-
tradiction that U(1) (a circle) is not connected. If u(N ) is the real vector space
of skew-Hermitian matrices, the exponential map exp : u(N ) → U(N ) has differ-
ential Id at IN and hence is a local homeomorphism about 0 ∈ u(N ). It follows
that U(N ) is locally path connected. Finally, a connected, locally path-connected
space is path-connected, since the path-components are then open and disjoint. In
Part III below, we further investigate the homotopy type of U(N ) which is only
partially known. In contrast, we can show for infinite dimensional H that U(H) is
contractible; see Remark 3.24 following Theorem 3.22 below.
Theorem 3.21. The group B × of invertible bounded linear operators on a
Hilbert space H is pathwise connected.
It is not true that the group of units of a Banach algebra is pathwise con-
nected. The group of units (B/K)× in the Calkin algebra is a counterexample; its
connected components – the connected components of Fredholm operators – are
mapped bijectively to Z by the index, as the following paragraph shows.
This theorem is usually proved by means of deeper results of spectral theory
(see, e.g., [128, 1972, p.134ff]). One first shows that every unitary operator U has a
spectral decomposition U = eiA = cos A + i sin A where A is a self-adjoint operator.
Then, by the Spectral Theorem 2.61,
t 7→ Ut := eitA , t ∈ I
is a continuous path in U(H) from Id to U . (If one is willing to use spectral theory,
this argument can replace the one above for the connectivity of U(N ).) One shows
further that each invertible operator R can be factored as R = U B where U is
3.4. HOMOTOPIES OF OPERATOR-VALUED FUNCTIONS 75


unitary, and B = R∗ R is self-adjoint, positive and invertible. Then one again
connects U with Id using Ut and B with Id with the path
t 7→ Bt := t Id +(1 − t)B, t ∈ [0, 1]
which does not go outside B × by the Spectral Theorem, since B is positive and
invertible. In this fashion t 7→ Ut Bt defines a path from R to Id.
Following an idea of Nicolaas Kuiper, here we provide a completely elemen-
tary proof of the theorem which perhaps is not as elegant as the proof outlined
above and which (as most elementary proofs) requires more calculation and per-
haps some more geometric imagination. The decisive advantage for us is that the
elementary proof generalizes effortlessly to a proof of Kuiper’s Theorem (Theorem
3.22) according to which [X, B × ] = 0, even if X does not consist of a single point as
in Theorem 3.21 but is an arbitrary compact topological space. While the content
of Theorem 3.21 remains unchanged in passing from CN to the infinite-dimensional
Hilbert space H, Theorem 3.22 exhibits a fundamental difference (see the Bott Pe-
riodicity Theorem in Chapter 10) between the linear algebra of finite-dimensional
vector spaces and the functional analysis of Hilbert space. This aspect we can bring
out clearly in the following proof of Theorem 3.21.
Proof of Theorem 3.21. Let R0 ∈ B × . We seek a continuous path in B ×
connecting R0 with Id. We proceed in two stages: In the first stage, we connect
R0 with an operator R2 which is the identity on a cleverly constructed infinite
dimensional subspace. In the second stage, we connect R2 with the identity of H.
Stage 1, step 1: We begin by recursively constructing a sequence of unit vectors
a1 , a2 , ... ∈ H and a sequence of 2-dimensional subspaces A1 , A2 , ... ⊆ H, such that
Ai ⊥Aj for i 6= j, and
ai ∈ Ai , R0 ai ∈ Ai for all i = 1, 2, . . . .
Start with any unit vector a1 ∈ H and a 2-dimensional subspace A1 which contains
a1 and R0 a1 . Then choose a unit vector
−1
a2 ∈ A⊥ A⊥

1 ∩ R0 1 ,

and let A2 be a 2-dimensional subspace with A2 ⊥A1 and containing a2 and R0 a2 ;


note that a2 ∈ R0−1 A⊥ a2 ∈ A⊥

1 ⇒ R 0 1 . Then proceed with
−1 −1
a3 ∈ A⊥ ⊥
A⊥ A⊥
 
1 ∩ A2 ∩ R0 1 ∩ R0 2 , etc.

The construction never breaks down, since the intersection of finitely many sub-
spaces of finite codimension in H (recall dim H = ∞) is never trivial.
Stage 1, step 2: Now we deform the operator R0 to R1 so that R1 ai is a unit
vector in the direction of R0 ai . Thus, define (for t ∈ I)
( L∞ ⊥
R0 u,  for u ∈ ( i=1 Ai ) ,
Rt u := 
(1 − t) + |R0tai | R0 u, for u ∈ Ai
Stage 1, Step 3: Deform the operator R1 to an operator R2 with the desired
property
R2 ai = ai for all i.
This will be done by constructing a suitable curve Tt ∈ U(H) (t ∈ [0, 1]) with
T0 = Id, Tt (Ai ) = Ai and T1 (R1 ai ) = ai . With R2 := T1 R1 , we then will have
L∞ ⊥
R2 ai = ai . Moreover, by construction, Tt will leave all vectors in ( i=1 Ai ) fixed.
76 3. FREDHOLM OPERATOR TOPOLOGY

The simple geometric construction of Tt (in which we have given up spatial intuition,
since a complex plane, of C-dimension 2, has R-dimension 4) can be reduced to
Exercise 3.20. Indeed, we map each Ai by an isometry αi onto C2 , in such a way
that the complex line {λR1 ai : λ ∈ C} is mapped to C × {0}. Then let gi : S 1 →
U(2) be a map with the properties (guaranteed by Exercise 3.20):
(i) (gi (z))(z) = (1, 0), for all z ∈ S 1 ⊂ C× {0}
(ii) There is a continuous map Fi : S 1 × I −→ U(2) with
Fi (·, 1) = gi and Fi (z, 0) = IdC2 for all z.
Let Bt ∈ U(2) (t ∈ [0, 1]) be a curve chosen so that B0 = Id and B1 αi (ai ) = (1, 0) ∈
S 1 ⊂ C× {0}. For t ∈ I, we now set
 L∞ ⊥
u, for u ∈ ( i=1 Ai ) ,
Tt u := −1 −1
αi Bt Fi (αi (R1 ai ) , t) αi u, for u ∈ Ai .
Then T0 = Id, and since αi (R1 ai ) ∈ S 1 ⊂ C× {0}, we have
T1 (R1 ai ) = αi−1 B1−1 Fi (αi R1 ai , 1) αi R1 ai
= αi−1 B1−1 (1, 0) = αi−1 (αi (ai )) = ai .
Thus,
t 7→ R1+t := Tt ◦ R1
is a continuous path in B from R1 to R2 with R2 |H 0 = Id, where H 0 ⊆ H denotes
×

the infinite-dimensional closed subspace spanned by a1 , a2 , ....

Stage 2, step 1: Relative to the decomposition H = H1 ⊕ H 0 , where H1 :=


(H ) denotes the orthogonal complement of H 0 in H, R2 has the form
0 ⊥

 
Q 0
,
∗ Id
where Q ∈ B × (H1 ) and the perturbation term ∗ can be deformed to zero by a
continuous path in B × (H)
   
Q 0 Q 0
R2+t = , t ∈ I; with R3 = .
(1 − t) ∗ Id 0 Id
Stage 2, step 2: By the classical argument (which one uses in set theory to
count Q or to demonstrate the equipotence of N and N × N), we can decompose H 0
into an infinite sum of Hilbert spaces H2 , H3 , .... Explicitly, let a1 , a2 , ... form an
orthonormal basis of H 0 . Decompose N into the infinite disjoint subsets
Nj := 2j−2 (2n − 1) : n ∈ N ) , j = 2, 3, 4, ....


and take Hj to
Lbe the closed subspace spanned by the ai with i ∈ Nj . In this way

we have H = j=1 Hj (recall H1 = (H 0 )⊥ ) and
 
Q 0 ··· 0
 .. 
 0 Id . 
R3 = Q ⊕ Id ⊕ Id ⊕ · · · := 
 .
.
 .. .. 
. 0 
0 ··· 0 Id
3.5. THE THEOREM OF KUIPER 77

If we now identify Hj and H1 for j 6= 1 (all infinite-dimensional, separable, Hilbert


spaces are trivially isomorphic), we can also write:
QQ−1 0 QQ−1 0
   
R3 = Q ⊕ ⊕ ⊕ ··· .
0 Id 0 Id
With the rotation of Exercise 3.18, we obtain a continuous path in B × (H1 ) ×
B × (H1 × H1 ) × · · · ⊆ B × (H) from R3 to an operator
R4 = Q ⊕ Q−1 ⊕ Q ⊕ Q−1 ⊕ Q ⊕ · · ·
 

= Q ⊕ Q−1 ⊕ Q ⊕ Q−1 ⊕ · · ·
 

and with one more rotation (this time in B × (H1 × H1 ) × B × (H1 × H1 ) × · · · ) to a


continuous path from R4 to R5 = Id ⊕ Id ⊕ · · · = IdH . 

5. The Theorem of Kuiper

I The idea of the last step of the preceding proof can be found in Albert Solomo-
novich Schwarz [371] and in Klaus Jänich [233]. It bares a secret which separates
fundamentally (from the topological standpoint) the linear functional analysis in Hilbert
space from the linear algebra of finite dimensional vector spaces: It is the possibility of
(figuratively speaking) escaping any squeeze by moving aside into a new dimension. If we
had decomposed H into only finitely many components H1 ⊕ · · · ⊕ Hm , we would have
gotten stuck in Hm , either with the homotopy from R3 to R4 or with the homotopy from
R4 to R5 (depending on whether m is even or odd). We meet a similar phenomenon when
investigating the geometry of unitary matrices where, e.g., the homotopy set [S i , U(N )]
(the well-known homotopy groups π i (U(N )) , where S i denotes the i -sphere of unit vectors
in Ri+1 ) is determined for 2N ≥ i + 1 by the Periodicity Theorem of Raoul Bott but
not known for all smaller values of N . Details are in Chapter 10.
Finally we wish to remark that generally in topology low-dimensional structures,
particularly 3- and 4-dimensional manifolds (recall the key words Vaughan Jones and
knot theory, Grigori Perelman’s proof of the Poincaré conjecture or Friedrich Hirze-
bruch’s work on the signature) are among the most difficult areas, while analogous ques-
tions for higher-dimensional objects were either solved decades ago or at least pose no
fundamentally new problems. This is the background which may help understand the
following basic theorem. J

Theorem 3.22 (N. Kuiper, 1964). For any compact, topological space X, the
homotopy set [X, B × (H)] consists of a single element, where B × (H) is the group of
bounded, invertible operators in the Hilbert space H.
Remark 3.23. Just like Theorem 3.21 (X = point), Theorem 3.22 holds for
nonseparable Hilbert spaces (see [229, No. 284/02f.] and for real Hilbert spaces (see
[270, p.19-30]). However, according to our earlier convention, we restrict ourselves
always to separable complex Hilbert space.
Remark 3.24. It is a corollary of Theorem 3.22 that B × (H) is contractible.
This would be completely trivial if Theorem 3.22 were valid for noncompact X,
for example for X = B × (H). But it is not this simple. Still there is a way (by
studying the nerve of an open cover of B × (H) as in Stage 0 of the following proof)
of reducing the question of contractibility to Theorem 3.22. See [270, p.27f.] and
[229, No. 284/01f.].
78 3. FREDHOLM OPERATOR TOPOLOGY

Remark 3.25. Contrary to the topological investigation of the general linear


group GL(N, C), there is no gain in restricting attention first to the group U(H)
of unitary operators (T T ∗ = T ∗ T = Id). In the classical case (see Chapter 10) the
advantage results from the compactness of U(N ) := U(CN ) but for dim H = ∞,
U(H) is not compact.
Proof. We adopt the proof of Theorem 3.21 with the necessary modifications
and several additional observations. In order to make a reasonable analogy between
a continuous family f : X → B × (H) of operators and a single operator R ∈ B × (H),
say R : {point} → B× (H), we must guarantee that Im f is at least contained in
a finite-dimensional subspace of B(H). Thus, before we consider the analogy, we
establish
Stage 0: Each continuous f0 : X → B × (H) with X compact is homotopic to a
continuous map f1 : X → B × (H) with f1 (X) ⊆ V , where V is a finite-dimensional
subspace of B(H).
To prove this, we use the openness of B × (H) (see Exercise 3.6, p. 3.6) and
place an open ball contained in B × (H) about each operator T ∈ f0 (X). This
gives us an open cover U of f0 (X). For safety reasons (see below), we replace each
ball U ∈ U by a ball U 0 with the same center, but with 1/3 the radius. Clearly,
U 0 := {U 0 : U ∈ U} is an open cover of f0 (X), since each operator in the image
of f0 is the center of some U 0 . As the continuous image of a compact set, f0 (X) is
also compact and it is covered by a finite subset of U 0 . Thus, we have f0 (X) ⊆ U∗ ,
SN
where U∗ = i=1 K(Ti , εi ) is the union of the finite set of open balls
K(Ti , εi ) := {T ∈ B(H) : kT − Ti k < εi } , i = 1, ..., N.
The balls are contained in B × (H) and are so small that K(Ti , 3εi ) is also contained
in B × (H).
Then we are essentially done with the verification of stage 0: Although U∗ is still
infinite-dimensional, it is visibly contractible to a simplicial complex with vertices
T1 , ..., TN , meaning a structure that lies entirely in the subspace of dimension ≤ N
of B(H) spanned by T1 , ..., TN , see Figure 3.1.
Since intuition (particularly for infinite-dimensional spaces) can be deceiving,
we will write out this argument precisely: For each t ∈ [0, 1] and T ∈ U∗ , we define
an operator
XN
gt (T ) := (1 − t)T + t φi (T )Ti ,
i=1
where φi (i = 1, ..., N ) is a partition of unity subordinate to U∗ ; i.e., φi : U∗ → R is
continuous and
(i) support φi ⊆ K(Ti , εi ); i.e., φi (T ) = 0 for kT − Ti k ≥ εi ,
(ii) 0 ≤ φi (T ) ≤ 1
PN
(iii) i=1 φi (T ) = 1 for T ∈ U∗
For example, set
ψi (T )
φi (T ) := PN for T ∈ U∗ , where
i=1 ψi (T )

εi − kT − Ti k , for T ∈ K(Ti , εi ) ,
ψi (T ) :=
0, otherwise.
In this way, g0 : U∗ → B × (H) is the inclusion and g1 : U∗ → B × (H) is a retraction of
U∗ onto a simplicial complex (composed of points, segments, triangles, tetrahedra,
3.5. THE THEOREM OF KUIPER 79

f0(X)
(subset of B £(H) µ B(H))

the finite covering U


*

T2 K(T2 ,"2 )
T1
K(T1 ,"1 )

T4
T5
T3
contraction onto T6
T7
simplicial complex
T2
T1

Figure 3.1. Two cases of visible contraction to a simplicial complex

and corresponding higher-dimensional simplices) with vertices T1 , ..., TN . To be sure


that the homotopy of g0 to g1 does not leave B × (H), we use an “ε/3-argument”
(see Figure 3.2): Let T ∈ U∗ . In considering gt (T ), we are (because of (i)) only
interested in those summation indices i for which T ∈ K(Ti , εi ). We compare
these balls and let K(Tm , εm ) be the one with the largest radius. By the triangle
inequality, each of these balls is contained in K(Tm , 3εm ). Thus, the convex hull of
T and these Ti (and hence gt (T ), 0 ≤ t ≤ 1) is contained in the ball K(Tm , 3εm ),
which by construction lies in B × (H).
With ft := gt f0 , t ∈ I, we obtain a homotopy in B × (H) of f0 to an f : X →
×
B (H) with the desired properties. This completes the essential work in carrying
over the proof of Theorem 3.21 to the more general situation of Theorem 3.22.
Indeed, the rest of the proof now is quite analogous; we no longer consider R0 as
a single element of B × (H), but rather as the intersection of the finite-dimensional
subspace V (in which f1 (X) lies) with B × (H). Then we need only show that {Id}
is a deformation retract of span(V, Id) ∩ B × (H), just as we showed in the proof of
Theorem 3.21 that R0 and Id are connected by a continuous path in B × (H), and
we are done. The proofs are identical in principle. In the particulars, we need the
following modifications:
Stage 1, step 1: We construct as above a sequence of unit vectors a1 , a2 , ... ∈
H and a sequence of pairwise orthogonal(n + 1)-dimensional subspaces A1 , A2 , ... ⊂
H such that ai ∈ Ai and Rai ∈ Ai for all R ∈ V and i = 1, 2, ...; here, n := dim V
and V is the vector space constructed in stage 0, except for containing in addition
the operator Id; i.e., replace V by span(V, Id).
80 3. FREDHOLM OPERATOR TOPOLOGY

3"m
Tm
"m
T T0
3"m
Ti

Figure 3.2. Keeping the homotopy of g0 to g1 in B × (H) by an


ε/3-argument

Stage 1, steps 2 and 3: Now we show that the canonical inclusion


γ0 : V ∩ B × (H) ,→ B × (H)
is homotopic to a map
γ2 : V ∩ B × (H) −→ B × (H)
with γ2 (R)(ai ) = ai for i = 1, 2, . . . and R ∈ V ∩ B × (H).
As before, in stage 1 step 2 (p. 75), we first go from γ0 to γ1 with
( L∞ ⊥
Ru, for u ∈ ( i=1 Ai ) ,
γ1 (R) u := Ru
|Rai | , for u ∈ Ai .

For step 3, one must again take up the rotation argument of Exercise 3.20 and
generalize it somewhat by induction: We regard
n o
2 2
S 2n−1 = z = (z1 , ..., zn ) ∈ Cn : |z1 | + ... + |zn | = 1)

as a subset of Cn × {0} ⊆ Cn+1 and construct, for each a ∈ S 2n−1 , a continuous


map g : S 2n−1 → U (n + 1) with the properties
(i) g(z)(z) = a for all z ∈ S 2n−1 and
(ii) g ∼ h, where h(z) = IdCn+1 for all z ∈ S 2n−1
(see also [229, No. 284/04]). We reduce our problem to this situation by mapping
each Ai by an isometry αi onto Cn+1 so that {T ai : T ∈ V } is mapped into Cn ×{0}.
Then, if Fi : S 2n−1 × I → U(n + 1) is the corresponding homotopy for a := αi (ai ),
then the following is a homotopy from γ1 to γ2 with the desired properties:
 L∞ ⊥
γ1 (T ) u, for u ∈ ( i=1 Ai ) ,
γ1+t (T ) u :=
αi−1 Fi (αi γ1 (T ) ai , t) αi γ1 (T ) u, for u ∈ Ai .
In particular, we have an infinite-dimensional H 0 (with orthonormal basis a1 , a2 , ...)
such that γ2 |H 0 = Id.
Stage 2: In the proof of Theorem 3.21, we have already implicitly shown that
B × ((H 0 )⊥ ) × IdH 0 , is a deformation retract of the group of all invertible operators
3.6. THE TOPOLOGY OF F 81

on H which are the identity on H 0 (step 1), and that B × ((H 0 )⊥ ) × IdH 0 , can be
contracted to {IdH } (step 2). The proof of Theorem 3.22 is now complete. 
Exercise 3.26. Prove the Theorem of Dixmier (and Douady) that shows
that U(H) is contractible in the strong operator topology. Recall or check, e.g., in
[332, Section 4.6, p.171 ]: The strong topology on B(H) is induced by the family
of seminorms of the form T 7→ kT xk for various x ∈ H. Since kT xk ≤ kT k kxk, we
immediately observe that the strong topology is weaker than the norm topology.
Kuiper’s theorem is about the norm topology and is much harder (although there
are now, depending on taste, more conceptual proofs than ours which is Kuiper’s
original one). The idea is very simple.
[Hint, following [122, Lemma 10.8.2]: First realize H as L2 [0, 1] and consider the
strongly continuous family Pt of orthogonal projections onto the subspaces L2 [0, t],
t ∈ [0, 1]. Of course Pt is given by multiplication by the characteristic function of
[0, t]. Now for t > √ 0, let Vt denote the obvious isometry of L2 [0, t] onto L2 [0, 1],
namely Vt (x)(s) = tx(ts) for s ∈ [0, 1] and x ∈ L2 [0, t], and let V0 = 0. Then
t 7→ Vt is strongly continuous. The contraction of U(H) to 1 is given by (u, t) 7→

(1 − Pt ) + Pt (Vt ) uVt Pt . At time t = 1, this is just the identity map u 7→ u on
U(H), while at time t = 0, this is the map u 7→ 1. Check that Vt∗ (y) (σ) = √1t y( σt )

for σ ∈ [0, t], and verify the unitarity of (1 − Pt ) + Pt (Vt ) uVt Pt for t ∈ (0, 1).]

6. The Topology of F
I Like the preceding one, this section is central for index theory but may be skipped
in first reading and read later in connection with the study of the topology of the general
linear group (Bott’s Periodicity Theorem, Chapter 10) and the topological interpretation
of elliptic boundary value problems (Sections 9.4 and 13.8). J

Our point of departure is the principal theorem on the homotopy invariance of


the index saying that the index is defined in a neighborhood of a Fredholm operator
and is constant there, and consequently on each connected component (Theorem
3.11, p.68). Regarding the topology of F, the space of Fredholm operators on a
given Hilbert space H, the following questions arise:
1. What is the number of connected components of F?
2. What can be said about the structure of the individual
connected components? How many holes are there of each type?
Conceptually imagine a serving of Swiss cheese. Not only do we note the numbers
of slices (Question 1) but also the kind of holes (Question 2), those that cross the
cheese like a channel and those which are enclosed like air bubbles.
We already have a partial answer for Question 1. There are at least Z con-
nected components, since the right-handed and left-handed shift, together with its
iterates (the natural powers), show that every integer can be the index of a Fred-
holm operator. We denote the path connected components of F by [point, F]. Then
the map index : [point, F] → Z is well-defined by Theorem 3.11, p.68, (homotopy
invariance of the index) and surjective. In fact, the map is injective also as we
will prove in this chapter. This answers Question 1 completely. In particular, it is
utterly impossible that F or the group of units (B/K)× of the quotient algebra is
contractible, in contrast to B × (Theorem 3.21, p.74). In addition, the individual
connected components of F differ as to homotopy type (Question 2) radically from
82 3. FREDHOLM OPERATOR TOPOLOGY

B × , which has an extremely simple structure according to the Theorem of Kuiper


(Theorem 3.22, p.77). In fact, we will show that the holes of F, its fissures, can be
arbitrarily complex in some sense, that F is a kind of model for all possible topolog-
ical structures (distinguishable via the functor K; see below). We can sketch this
aspect before starting with formal definitions and theorems, as follows: In algebraic
topology, for instance in the “K-theory” treated in Part III, one has methods which
assign to certain topological spaces (compact or triangulable spaces, differentiable
manifolds, etc.) certain algebraic objects (groups, rings, algebras, etc.) and to con-
tinuous (differentiable etc.) maps certain homomorphisms between the associated
algebraic objects. In this fashion certain topological-geometric phenomena, which
are very difficult to distinguish in concrete visualization, can be reduced to alge-
braic terms (made discrete, F. Waldhausen) for which there is a well-developed
formalism that is easier to comprehend. For example, it is immediate that there
is no epimorphism of the group Z onto the group Z ⊕ Z, while the nonexistence of
a retraction of the n-dimensional ball B n onto its boundary S n−1 is not immedi-
ately clear (see Theorem 10.26). The construction of the index bundle which will
be treated in this chapter is the basis of a further step from algebraic topology
to functional analysis or to elliptic topology (Atiyah). We will be able to explain
this term only in the following chapters which deal with the connection between
elliptic differential equations and Fredholm operators. In this way, deep geometric
questions in the proof of Bott’s Periodicity Theorem (generalization of the con-
cept winding number ) find an algebraic formalism in K-theory. Thereby the Bott
isomorphism K(X × R2 ) → K(X) is described by means of families of Fredholm
operators, and in this form can be understood more easily and elementarily by a
symmetric use of classical results of functional analysis (see Chapter 10). This is an
example where the interpretation in functional analysis makes the understanding
of geometric or algebraic situations easier or perhaps only possible.
Conversely, the topological and algebraic problems and methods serve analy-
sis: for example, when the index or the index bundle yield algebraic-topological
invariants for Fredholm operators or families of Fredholm operators which turn up
concretely in problems of analysis.

7. The Construction of Index Bundles


We now come to the construction of index bundles. Let T : X → F be a
continuous family of Fredholm operators in a Hilbert space H, where X is a compact
topological space. If X is connected, Theorem 3.11 (p.68) gives
index Tx = index Tx0 , for all x, x0 ∈ X.
In this way, we can assign to each T an integer, which is independent of possible
small continuous perturbations of T . For X connected, we therefore have a map
index : [X, F] −→ Z,
which is well defined on the homotopy set [X, F]; see (3.3), p.72. Actually, we can
extract much more information out of T .
Exercise 3.27. Show that for each continuous family T : X → F with constant
kernel dimension, i.e., dim Ker Tx = dim Ker Tx0 , for all x, x0 ∈ X, one can assign a
vector bundle Ker T over X, in a natural way. Here, a vector bundle E over X is
a continuous locally-trivial family of complex vector spaces Ex of finite dimension,
3.7. THE CONSTRUCTION OF INDEX BUNDLES 83

parameterized by the base space X. For the details of the definition, we refer to
the Appendix.
[Hint: Set Ker T := ∪x∈X {x} × Ker Tx , and give this the induced topology that it
inherits as a subset of X × H. Then show, as in the alternative proof (see p.70) of
Theorem 3.11, the property of local triviality; i.e., locally there is a basis of Ker Tx
which depends continuously on x. See also Exercise B.2c (p. 714) of the Appendix,
where X = S 1 .]
Exercise 3.28. Under the same assumptions as in Exercise 3.27, show that
there are naturally defined vector bundles Ker T ∗ and the quotient bundle Coker T .
Moreover, show that these bundles are isomorphic (see Appendix).
We denote the set of all isomorphism classes of vector bundles over X by
Vect(X). By Exercises 3.27 and 3.28, we have a map ι from the set C(X, F)
of continuous families of Fredholm operators with constant kernel dimension to
Vect(X)× Vect(X):
ι : C(X, F) −→ Vect(X) × Vect(X) given by
ι(T ) := ([Ker T ] , [Coker T ]) .
If X consists of a single point, then a vector bundle over X is simply a single vector
space, and Vect(X) can be identified with Z+ , since vector spaces are isomorphic
precisely when their dimensions are equal; symbolically, [·] = dim(·). In this case,
we then have the maps
ι δ
F −→ Z+ × Z+ −→ Z given by
ι δ
T 7→ ([Ker T ], [Coker T ]) 7→ [Ker T ] − [Coker T ],
where δ is the difference mapping and δ ◦ ι = index.
In the more general case where X is not a single point, we can formally write
such a difference [Ker T ] − [Coker T ], at least when the family T has constant
kernel dimension. Admittedly, this is meaningless for the moment: While one can
naturally introduce an addition ⊕ in Vect(X) by forming the direct sum pointwise
(see Exercise B.4, p. 715, of the Appendix), this only makes Vect(X) a semigroup.
However, one can go from the abelian semigroup Vect(X) to an abelian group (just
as from Z+ to Z), which we denote by K(X) and define as follows. An equivalence
relation on the product space Vect(X) × Vect(X) is defined by means of
(E, F ) ∼ (E ⊕ G, F ⊕ G) for G ∈ Vect(X),
and then
K(X) := (Vect(X) × Vect(X))/ ∼ .
The equivalence class of the pair (E, F ) is then written as E − F ; these we call
difference bundles.
The details of this construction, and its basic importance for algebraic topology,
is explained in Section 10.3. Here we only need to establish that the difference map
δ extends from Z+ × Z+ to Vect(X) × Vect(X) in a canonical way so that its values
form an abelian group K(X). (To be careful, X is always compact in this chapter,
but many of the steps carry over easily to more general cases.)
[Warning: In spite of the close analogy between the construction of K(X) and Z,
notice that (for X 6= point) the natural map
Vect(X) −→ K(X), given by E 7→ E − 0
84 3. FREDHOLM OPERATOR TOPOLOGY

need not be injective. In Chapter 10, we will get to know vector bundles E and F
over the two-dimensional sphere S which are not isomorphic, but when the same
vector bundle G is added to each of them, the results are isomorphic. (Visually, one
may note that the tangent bundle of S 2 and the real two-dimensional trivial bundle
over S 2 are not isomorphic, but the addition of a trivial one-dimensional bundle to
each yields isomorphic bundles; see the Appendix, Exercise B.13a, p. 719).]
With the help of the maps ι and δ just introduced, we can now deduce from
Exercise 3.27 and Exercise 3.28:
Exercise 3.29. To each continuous family T : X → F of Fredholm operators,
with constant kernel dimension, there can be assigned a difference bundle (index
bundle) index T := [Ker T ] − [Coker T ] ∈ K(X). Moreover, for a point space X

(T is then a single Fredholm operator and K(X) −→ Z), the concepts of index and
index bundle coincide.
[Hint: Set index := δ ◦ ι.]
Now we will show that the unrealistic condition that the kernel dimension
be constant can be dropped, and the index bundle is invariant under continuous
deformation just as the index (Theorem 3.11, p.68).
Theorem 3.30. Let X be compact. For each continuous family T : X → F of
Fredholm operators in a Hilbert space H, there is an index bundle
index T ∈ K(X)
assigned in a canonical way.
Proof. As in the alternative proof (see p.70) of Theorem 3.11, we first choose
an orthonormal system e0 , e1 , ... for H and consider the Fredholm operator Pn Tx
(which has the same index as Tx ), where Pn is again the projection of H onto the
closed subspace Hn spanned by en , en+1 , . . . . Since X is compact, we can choose n
such that
Im(Pn Tx ) = Hn for all x ∈ X, and then
dim Ker Pn Tx = dim Ker Pn Tx0 for all x, x0 ∈ X.
Indeed, for each y ∈ X, there is a natural number ny and a neighborhood Uy so
that for all x ∈ Uy , Tx will be sufficiently close to Ty to ensure (by means of the
alternative proof, p.70) that dim Ker Pn Tx = dim Ker Pn Ty for all n ≥ ny . Then
we pass to a finite subcover {Uy : y ∈ Y }, where Y is a finite subset of X, and
set n := max(ny : y ∈ Y ). Relative to the family Pn T , we can therefore (as in
Exercise 3.29) set
index T := index Pn T = [Ker Pn T ] − [Coker Pn T ]

= [Ker Pn T ] − [X ×(Hn ) ].
Thus, we have assigned an index bundle in K(X) to T . Indeed, we have expressed
the index bundle in normal form, in the sense that the bundle subtracted is trivial.
We must show that the definition is independent of the sufficiently large natural
number n. Without loss of generality, replace n by n + 1. Then, by construction,
[Coker Pn+1 T ] = [X × (Hn+1 )⊥ ] = [Coker Pn T ] ⊕ [X × Cen ].
To calculate [Ker Pn+1 T ], we note that for all x ∈ X,
Pn Tx |(Ker Pn Tx )⊥ : (Ker Pn Tx )⊥ −→ Hn
3.7. THE CONSTRUCTION OF INDEX BUNDLES 85

is bijective. Hence, by the closure of Hn and the open mapping principle, there is
a bounded inverse operator
Tex : Hn −→ (Ker Pn Tx )⊥ ⊆ H.
We set
vx := Tex (en ) ∈ (Ker Pn Tx )⊥ , so that
Pn Tx vx = en for all x ∈ X.
Then for all x ∈ X, we have
Ker Pn+1 Tx = Ker Pn Tx ⊕ Cvx and
[Ker Pn+1 Tx ] = [Ker Pn T ] ⊕ [{(x, zvx ) : x ∈ X, z ∈ C}] .

As with T , the family Te is also continuous (prove!). Thus, Te yields an isomorphism


of bundles
X × Cen −→ {(x, zvx ) : x ∈ X, z ∈ C} ,
(i.e., x 7→ vx is a continuous, nowhere zero section over X in X × H). Thus, the
index bundles index(Pn+1 T ) and index(Pn T ) are equal in K(X).
Finally, we must show the independence of the choice of orthonormal ba-
sis. For this problem we use the independence of n just shown. Let ee0 , ee1 , . . .
be another complete orthonormal system for H, and let Pem : H → H e m denote
the orthogonal projection onto the closed subspace H e m spanned by eem , eem+1 , . . ..
 
There is m e 0 ∈ N such that for all m ≥ m e 0 , we have Im Pem Tx = H e m . As Hn
and H e m both have finite codimension, so does Hn ∩ H e m . Since Ker Pn T is the
same if we replace en , en+1 , . . . by another orthonormal system for Hn , we may
assume that en , en+1 , . . . is adapted to Hn ∩ H e m in the sense that for some k, we
have that en+k , en+k+1 , . . . is a complete orthonormal system  for Hn ∩ Hm . Sim-
e

ilarly, we may assume that eem , eem+1 , . . . is chosen so that eem+ek , eem+ek+1 , . . . =
(en+k , en+k+1 , . . .) for some ek. Then (en+k , en+k+1 , . . .) is a common tail of both
sequences en , en+1 , . . . and eem , eem+1 , . . . . The orthogonal projection onto the span
of this common tail (i.e., onto Hn ∩ H e m ) is Pn+k = Pe e . Thus, by the above
m+k

index Pn T = index Pn+k T = index Pem+ek T = index Pem T. 


Remark 3.31. For simplicity, we have assumed that the continuous family
T : X → F of Fredholm operators is such that T is defined on the same Hilbert
space for all x ∈ X. In applications, we actually often have a different situation.
For example, in analysis the arising Hilbert spaces are function spaces with values
in a vector space Vx which can vary with the parameter x ∈ X. Then Tx is a
Fredholm operator on the Hilbert space H ⊗ Vx ; i.e., the Hilbert space on which
the continuous family of Fredholm operators operates is not constant. We will
encounter such examples mainly in Part III, when we interpret homomorphisms
of K-theory, in particular products of algebraic topology, with tools of analysis.
What can be done in such cases? In general, using the Theorem of Kuiper, one
can show that every Hilbert space bundle is trivial (i.e., isomorphic to a product
bundle X × H, where H is a single Hilbert space). This follows from [404, p.54f]
and the contractibility of the structure group B × , and more simply from [A, B × ] = 1
for all compact A (Theorem 3.22, p.77) following [404, p.148f] (e.g., in the case that
86 3. FREDHOLM OPERATOR TOPOLOGY

X is a compact, triangulable manifold). For the construction of index bundles, it


suffices to have only a local product structure which is already part of the definition
of a Hilbert bundle and this always holds (by the naturality of definitions) in our
applications.
Remark 3.32. In the proof of Theorem 3.30, we have strongly used the fact
that H is a Hilbert space. We treated the general case with variable Ker Tx by
reducing it, through composition with an orthogonal projection, to the simple spe-
cial case of constant kernel dimension. Instead, we can study the family T on the
quotient spaces H/ Ker Tx which produces the constant kernel dimension 0, or we
can make the operators Tx surjective by extending their domain of definition (to
H ⊕ Ker(Tx∗ ) say) which produces constant cokernel dimension 0. Details for these
two alternatives, which are already indicated in the usual proofs of the homotopy
invariance of the index (e.g., see [234, 1970/1982, 5.4] and our comments after
Theorem 3.11), can be found for the first alternative in [17, p.155-158] and for the
second alternative, for instance if H is a Fréchet space and T is an elliptic operator
(see below) in [47, p.122-127].
[Warning: In connection with Wiener-Hopf operators, we will later meet fami-
lies of Fredholm operators T : X → F for which it is quite possible that index T = 0
for all x ∈ X, while the global index bundle index T ∈ K(X) does not vanish. The
index bundle is simply a much sharper invariant than just an integer. Thus one
must be very careful if one wishes to infer properties of the index bundle from those
of the index. The next exercise may be comforting.]
Exercise 3.33. Show that for each continuous Fredholm family T : X → F, we

have index T ∗ = − index T , where T ∗ : X → F is the adjoint family (T ∗ )x := (Tx )
for x ∈ X.
Theorem 3.34. The construction of the index bundle in Theorem 3.30 depends
only on the homotopy class of the given family of Fredholm operators, and for X
compact yields a homomorphism of semigroups
index : [X, F] −→ K(X).
Proof. 1. We show first the homotopy-invariance of the index bundle. Let
T : X×I → F be a homotopy between the families of Fredholm operators T0 := T i0
and T1 := T i1 parametrized by X, where
it : X −→ X × {t} ,→ X × I, t ∈ I
are the mutually homotopic natural inclusions. The functoriality of index bundles
(Exercise 3.37) then yields
index T0 = i∗0 (index T ) = i∗1 (index T ) = index T1 ,
where the middle equality follows from the homotopy invariance of K(X); namely,
f ∗ = g ∗ for f ∼ g : Y → X (Theorem B.7, p. 716, Appendix).
2. For the homomorphism property of the index, we point out first that for two
Fredholm families S, T : X → F, we have a well defined product family T S : X → F
given by composition in F (i.e., (T S)x := Tx Sx ), and that [X, F] is then a
semigroup. On the other hand, K(X) has an addition defined via the direct sum of
vector bundles which makes it a group (Section 10.3). The index is a semigroup
3.7. THE CONSTRUCTION OF INDEX BUNDLES 87

homomorphism now, since


index T S = index T S + index Id = index(T S ⊕ Id)
= index(S ⊕ T ) = index S + index T.
In the second and fourth equalities, we have used the fact that (by the construction
in Theorem 3.30) the index bundle, for a family of Fredholm operators on a product
space H × H which can be written in the form
 
S 0
S ⊕ T := ,
0 T
coincides with the direct sum of the index bundles of the two diagonal elements.
Moreover, for the third equality, we have applied the usual trick of homotopy theory,
where T S⊕Id can be deformed into S⊕T in F(H×H). This is clear by the definition
of this homotopy in Exercise 3.18 (p. 73). Note that we do not need T, S ∈ B × (H)
since we are deforming in F(H × H), as opposed to B × (H × H). 
Remark 3.35. Note that in the proof of homotopy invariance, we did not need
to again investigate the topology of F, but rather everything followed from the
entirely different aspect of homotopy invariance of vector bundles. Why did it take
more work to prove the homotopy-invariance of the index of a single operator in
the proof of Theorem 3.11, p.68, while in the proof of Theorem 3.34 we did not use
this result, even though we can deduce it for X = {point}. The solution of this
paradox lies in the fact that the actual generalization of the homotopy invariance
of the index is already contained in the construction of index bundles in Theorem
3.30.
Exercise 3.36. Show that the index bundle of a continuous family of self-
adjoint Fredholm operators, T : X → F with Tx = Tx∗ for all x ∈ X is zero; i.e.,
index T = 0.
[Warning: One cannot deduce this from Exercise 3.33, since it is possible that K(X)
has a finite cyclic (torsion-) factor; e.g., possibly a + a = 0, but a 6= 0.]
[Hint: Consider the homotopy tT + i(1 − t) Id. Show that a self-adjoint operator A
has a real spectrum; i.e.,
A − z Id ∈ B × for z ∈ C − R :
Step 1. u ∈ Ker(A − z Id) implies zu = Au = A∗ u = z̄u, and so u = 0 since z 6= z̄.
Step 2. If v is in the orthogonal complement of Im(A−z Id), then hAu − zu, vi = 0
and so
hu, Avi = hAu, vi = hzu, vi = hu, z̄vi for all u ∈ H.
Thus, Av = z̄v, and v = 0 by step 1. Hence, Im(A − z Id) is dense in H.
Step 3. Let v ∈ H and let v1 , v2 , ... ∈ Im(A − z Id) be a sequence converging to v.
Then show that the sequence of unique preimages u1 , u2 , ... is a Cauchy sequence,
and set u := lim ui . Then note that Au − zu = v.]
We will now investigate the construction in the Theorem 3.30 more generally,
and show that it has all the properties that one might reasonably expect. We begin
with the following exercise.
Exercise 3.37. Show that our definition of index bundle is functorial: Let
f : Y → X be a continuous map (X and Y compact) and let T : X → F be a
continuous family of Fredholm operators. Then (see Exercise B.1d, p. 713, of the
88 3. FREDHOLM OPERATOR TOPOLOGY

Appendix), we have index T f = f ∗ (index T ).


[Hint: If e1 , e2 , ... is an orthonormal basis for H and the projection Pn is chosen
so that dim Ker Pn Tx is independent of x, then dim Ker Pn Tf(y) is also independent
of y ∈ Y . Thus, one does not need to choose another projection Pn to exhibit the
index bundle for the family T f : Y → F.]
Exercise 3.38. The proof of Theorem 3.34 yields (for X = point) a further
proof of the composition rule index T S = index S + index T for Fredholm operators.
We have already seen three proofs, namely Exercise 1.10 (p.8), Exercise 2.3 (p.14),
and Exercise 3.7 (p.65). Do these three proofs carry over without difficulty to the
case of families of Fredholm operators? Does one need a homotopy argument each
time? What is the real relationship between the four proofs?
Exercise 3.39. Show that the set
F0 := {T ∈ F : index T = 0}
is pathwise-connected.
[Hint: For T ∈ F0 choose an isomorphism
φ : Ker T −→ (Im T )⊥
(vector spaces of the same finite dimension!) and set

φ, on Ker T,
Φ := ⊥
0, on (Ker T ) .
By construction, we have T + Φ ∈ B × and T + tΦ ∈ F for t ∈ I (Why? What kind
of operator is Φ?). Now apply Theorem 3.21.]

8. The Theorem of Atiyah-Jänich


The sets Fi := {T ∈ F : index T = i} are bijectively (modulo compact op-
erators) mapped onto each other in a continuous fashion by shift operators (see
Exercise 1.3, p.4). Thus, Exercise 3.39 with the homotopy-invariance of the index
already proved in Theorem 3.11 (p. 68) gives us a bijection from the pathwise-
connected components of F to Z. Naturally, this result still does not say much
about the topology of F, and we do not want to carry it out in detail. Much more
informative is the following theorem which gives the result,
index : [{point} , F] −→ Z bijective,
as the special case X = {point}.
Theorem 3.40 (M. F. Atiyah, K. Jänich 1964). We have an isomorphism
index : [X, F] −→ K(X).
Proof. We show that the natural sequence of semigroups
index
[X, B × ] −→ [X, F] −→ K(X) −→ 0
is exact. From this result of M. Atiyah and K. Jänich along with the theorem of
N. Kuiper (Theorem 3.22, p.77), the present theorem follows. (For the concept of
exactness, see the material before Theorem 1.9, p. 6).
Step 1. The index bundle of a continuous family T : X → B × is trivially zero.
3.8. THE THEOREM OF ATIYAH-JÄNICH 89

Step 2. Now let T : X → F be a Fredholm family with vanishing index bundle.


In Exercise 3.39, we have stated what one must do in the case where X = {point};
i.e., T is a single Fredholm operator. Namely, one chooses an isomorphism
φ : Ker T −→ (Im T )⊥
and then 
φ, on Ker T,
Φ := ⊥
0, on (Ker T )
is an operator of finite rank with T + tΦ the desired homotopy in F to T + Φ ∈ B.
To generalize this to X 6= {point}, we now must deal with two difficulties:
S
l. x∈X Ker Tx is not always a vector bundle.
2. In general, there is no canonical choice of
φx : Ker Tx −→ (Im T )⊥ which depends continuously on x.
The first problem, we can quickly solve: We choose again an orthogonal projection
Pn : H → Hn , such that dim Ker Pn Tx is constant and

index T = [Ker Pn Tx ] − [X ×(Hn ) ] ∈ K(X).
Then “index T = 0” means
[Ker Pn T ] = [X × (Hn )⊥ ] in K(X),
which in turn means (see Section 10.3) that, for N sufficiently large, there is an
isomorphism of vector bundles

φ : (Ker Pn T ) ⊕ (X × CN ) −→ (X × (Hn )⊥ ) ⊕ (X × CN ).
This means that by augmenting with the trivial bundle X ×CN , we can also cure the
second difficulty in principle. Actually we can avoid this K-theoretical argument.
Indeed, we do not need to augment, if we take n to be large enough: As shown in
the proof of Theorem 3.30 (p.84), we have
Ker Pn+1 T ∼
= Ker Pn T ⊕ (X × C).
Thus, for m := n + N , the map φ can be regarded as an isomorphism
Ker Pm T −→ X × (Hm )⊥ .
In this way, the construction of φ in Exercise 3.39 carries over pointwise, whence
we obtain a homotopy of Fredholm families between
Pm T : X −→ F(H) and Pm T + Φ : X −→ B × (H).
Since the index of the orthogonal projection Pm : H → Hm vanishes (and hence
coincides with the index of Id), we can connect the constants Pm with Id in F(H)
(Exercise 3.39), and hence also connect Pm T with T in a further homotopy of
maps X → F(H). Each Fredholm family with vanishing index bundle is therefore
homotopic to a continuous family of invertible operators, and, in conjunction with
step 1, exactness in the middle of the short sequence is now proved.
Step 3: We must now show the surjectivity of index : [X, F] → K(X). By
the construction in Exercise B.12 (p. 719) of the Appendix (see also Section 10.3),
every element of K(X) can be written in the form [E] − [X × Ck ], where E is
a vector bundle over X and k ∈ Z+ . Hence we will be finished with the proof
if, for each vector bundle E, we can find a continuous family S of (surjective)
90 3. FREDHOLM OPERATOR TOPOLOGY

Fredholm operators on a suitable Hilbert space such that index S = [E]. Then, the
homomorphism of the index bundle construction (Theorem 3.30) yields
k
index shift+ S = [E] − [X × Ck ],
where shift (as in Exercise 1.3, p. 1.3) is the displacement to the right relative to
an orthonormal basis of the Hilbert space. Note that, since index shift+ = −1, the
constant family x → shift+ gives the index bundle −[X × C].
Thus, let E be a vector bundle over X. If X consists of a single point, then
we just complete an orthonormal basis of E (regarded as a subspace of an infinite-
dimensional separable Hilbert space) to an orthonormal basis of the whole Hilbert
space and set S := (shift− )dim E . Now, we consider the general case. By Exercise
10.11 (p. 264) of the Appendix, there is a vector bundle F over X and a finite-
dimensional (complex) vector space V such that E ⊕ F ∼ = X × V . Let π : V → E
denote the projection.
Let H be an arbitrary Hilbert space. We will see that it is simplest to consider
the desired operators Sx , x ∈ X, to be defined on the space Hom(V, H) of linear
transformations from V to H. We choose for the vector space V (which we can
identify with CN , N = dim V , via a basis) a Hermitian scalar product h·, ·i : V ×
V → C. For every pair (f, u) ∈ V × H, we have an element of Hom(V, H) defined
by
v −→ hf, vi u, v ∈ V,
which we will denote by f ⊗u; recall the isomorphism Hom(V, H) = V 0 ⊗H of linear
algebra, where V 0 (∼
= V ) is the vector space of linear maps from V to C. Then, we
have nXm o
Hom(V, H) = fi ⊗ ui : m ∈ N, fi ∈ V, ui ∈ H ,
i=1
where
zf ⊗ u = f ⊗ zu, z ∈ C and
(f + f ) ⊗ (u + u0 ) = f ⊗ u + f ⊗ u0 + f 0 ⊗ u + f 0 ⊗ u0 .
0

We define a scalar product


DXm Xm E Xm
fi ⊗ ui , fj0 ⊗ u0j := fi , fj0 ui , u0j ,
i=1 j=1 i,j=1

which makes Hom(V, H) a Hilbert space (since V is finite-dimensional, nothing


can go wrong; addition and scalar multiplication by complex numbers are defined
naturally). One easily checks that if f1 , ..., fN is an orthonormal basis for V and
e0 , e1 , e2 , . . . is an orthonormal basis for H, then
{fj ⊗ ei }j=1,...,N, i=0,1,2,...
is a orthonormal basis for Hom(V, H).
Each bounded linear operator T on H and endomorphism τ of V clearly define
a bounded linear operator τ ⊗ T on Hom(V, H) by means of
(τ ⊗ T )(f ⊗ u) := τ (f ) ⊗ T (u) .
For a fixed chosen basis e0 , e1 , e2 , . . . of H, we set
Sx := (πx ⊗ shift− ) + (IdV −πx ) ⊗ IdH .
3.9. DETERMINANT LINE BUNDLES 91

Then,

πx (f ) ⊗ ei−1 +(f − πx (f )) ⊗ ei , for i ≥ 1,
Sx (f ⊗ ei ) =
(f − πx (f )) ⊗ e0 , for i = 0,
and in particular for f ∈ Ex , we have Sx (f ⊗ e0 ) = 0. Hence, we have Im Sx =
Hom(V, H) and Ker Sx = {f ⊗ e0 : f ∈ Ex }. Thus, Ker Sx is isomorphic to Ex in
a natural way, and index S = [Ker S] − 0 = [E]. 

I The preceding theorems are significant on various levels: In the following chapters,
in dealing topologically with boundary value problems as well as in proving analytically
the Periodicity Theorem of the topology of linear groups, we will repeatedly use Theorems
3.30 and 3.34, i.e., the construction of the index bundle and its elementary properties, but
we will not use explicitly Theorem 3.40, our actual main theorem. Nevertheless, the last
theorem has fundamental significance for our topic, as it provides the reasons for the
theoretical relevance of the notion index bundle and explains why this concept proved
suitable to express deep relations in analysis as well as topology. J

9. Determinant Line Bundles

I A primary motivation for the study of determinant line bundles originated from
the desire of quantum physicists to produce a gauge-invariant volume element for the
purpose of computing (via functional integration) Greens functions for Dirac operators
coupled to gauge potentials. An obstruction to doing this is the nontriviality of the so-
called determinant bundle for a certain family of Fredholm operators, namely the family
of Dirac operators parametrized by the quotient space of gauge potentials modulo gauge
transformations. This obstruction signals the presence of so-called anomalies that arise
when physicists attempt to quantize the classical field theory. We refer to the comprehen-
sive [376] for a systematic presentation of determinants and traces on Banach algebras of
operators with emphasis on elliptic pseudo-differential operators and fascinating relations
to number theory. J

In this section we develop the notion of the determinant line bundle of a con-
tinuous family of Fredholm operators. Moreover, in the case of a family with index
0, we assist the reader by showing (see Exercise 3.53 and Corollary 3.54) that it
is the pull-back of a universal determinant line bundle that we construct over F0 ,
the space of Fredholm operators of index 0. The simplest way of describing this
line bundle (often referred to as the Quillen determinant line bundle, which arose
in [348]) is to declare the fiber above the point T ∈ F0 to be
ΛdT (ker T )∗ ⊗ ΛdT (ker T ∗ ), where dT := dim ker T = dim ker T ∗ .
However, the fact that these fibers may be pieced together to form a genuine line
bundle is not trivial since dT varies with T , and dT is an unbounded function of
T ∈ F0 . In doing this, we adopt a novel approach due primarily to Graeme
Segal (in [385]). Various ways of describing this universal bundle are summarized
in Theorem 3.55.

1. The Exterior Determinant. For compact X and a continuous family


T : X → F of Fredholm operators in a fixed Hilbert space H, we have seen that
there is an index bundle
index T ∈ K(X)
92 3. FREDHOLM OPERATOR TOPOLOGY

assigned in a canonical way. While the determinant of an operator on a Hilbert


space H exists only in restrictive circumstances, we now show that there is a well-
defined complex line bundle det T → X, known as the determinant line bundle of
T . For any α ∈ K(X), we will first define an isomorphism class det α ∈ K(X)
which is represented by a line bundle. For a continuous family T : X → F, we
then can (and do) simply define det T to be det(index T ), where index T ∈ K(X) is
given in the Atiyah–Jänich Theorem 3.40. We know that α = [E] − [F ] ∈ K(X) for
E, F ∈ Vect(X). For a finite-dimensional vector bundle V → X, it is convenient to

use the notation Λmax (V ) = Λdim V (V ). We claim that Λmax (E) ⊗ Λmax (F ) ∈
K(X) only depends on α, in the strong sense that if α = [E] − [F ] = [E 0 ] − [F 0 ],
then
∗ ∗
(3.4) Λmax (E) ⊗ Λmax (F ) ∼ = Λmax (E 0 ) ⊗ Λmax (F 0 )
∗ ∗
(E) ⊗ Λmax (F ) = Λmax (E 0 ) ⊗ Λmax (F 0 ) in K(X)). Indeed,
 max   
(not just Λ
using the fact that Λmax (V ⊕ W ) = Λmax (V ) ⊗ Λmax (W ), we have
[E] − [F ] = [E 0 ] − [F 0 ] ⇐⇒ E ⊕ F 0 ⊕ Ck = ∼ E 0 ⊕ F ⊕ Ck for some k
=⇒ Λmax (E ⊕ F 0 ) ∼
Λmax (E 0 ⊕ F )
=
∼ Λmax (E 0 ) ⊗ Λmax (F ) .
=⇒ Λmax (E) ⊗ Λmax (F 0 ) =
∗ ∗
Tensoring both sides with Λmax (E 0 ) ⊗ Λmax (E) and noting that for a line bundle
L, L ⊗ L∗ is isomorphic to the trivial bundle X × C, we then obtain (3.4). We can
now make the following
Definition 3.41. For X compact and α ∈ K(X), let E and F be vector
bundles over X with α = [E] − [F ].
a) We define the exterior determinant of α by

det α := Λmax (E) ⊗ Λmax (F ) ∈ K(X) .
 


b) By (3.4) the bundle Λmax (E) ⊗ Λmax (F ) is well defined (independent of the
choice of E and F ). By an abuse of notation, we also denote this isomorphism class
by det α.
c) For a continuous family T : X → F,
det T := det(index T ).
Note that while det T has been defined, this does not directly imply that

[
(3.5) Λmax (Ker(Tx )) ⊗ Λmax (Ker(Tx∗ ))
x∈X

can be given the structure of a (global) vector bundle over X. Doubts about this are
sometimes deflected by saying: while Ker(T ) and Ker(T ∗ ) are not defined globally
in general, dim Ker(Tx ) and dim Ker(Tx∗ ) jump by the same amount if x varies.
Actually, we know there exists a family R such that T + R is surjective. Then we
may set, as for the index bundle,
(3.6) det T := Λmax Ker(T + R),
and this is independent of R up to isomorphism. Hence, a busy reader may skip
our long proof of the following Proposition. However, it seems to us useful (and
comforting) to show that for some genuine vector bundles E and F over X with

index T = [E] − [F ], the fiber Λmax (Ex ) ⊗ Λmax (Fx ) of the manifestly well defined
3.9. DETERMINANT LINE BUNDLES 93

∗ ∗
bundle Λmax (E) ⊗Λmax (F ) → X is isomorphic to Λmax (Ker(Tx )) ⊗Λmax (Ker(Tx∗ ))
in a natural fashion. Indeed, as in the proof of Theorem 3.30 (p. 84), let Pn : H →

Hn = {e0 , . . . , en } be an orthogonal projection so that Pn Tx H = Hn for all x ∈ X.

We then have bundles E = Ker Pn T and F = Ker(Pn T ) = X × Hn⊥ , and a natural
isomorphism
∗ ∼ ∗
Λmax (Ker Pn Tx ) ⊗ Λmax Hn⊥ −→ Λmax (Ker(Tx )) ⊗ Λmax (Ker(Tx∗ ))


is supplied by taking G to be Tx in the following


Proposition 3.42. Suppose that G ∈ F and Pn G(H) = Hn . Then there is an
isomorphism (depending only on the choice Pn )
(3.7)
∗ ∼ ∗
ΨG,n : Λmax (Ker Pn G) ⊗ Λmax Hn⊥ −→ Λmax (Ker G) ⊗ Λmax (Ker G∗ ) .


Proof. First note that we have an isomorphism



G Ker(G)⊥ : (Ker G) −→ G(H) ,
e := G|

and hence
 

Ker Pn G = G−1 Hn⊥ = Ker(G) ⊕ Ker(G) ∩ G−1 Hn⊥
e −1 Hn⊥ ∩ G(H) .

(3.8) = Ker(G) ⊕ G

The orthogonal projection Qn : H → Hn⊥ is IdH −Pn . We are given that Pn G(H) =

Hn . This implies that Qn |G(H)⊥ : G(H) → Hn⊥ is injective, because
⊥ ⊥
v ∈ G(H) , Qn (v) = 0 =⇒ v ∈ Hn ∩ G(H)
=⇒ v = Pn (w) for some w ∈ G(H) , say w = G(u)
=⇒ G(u) = w = Pn (w) + Qn (w) = v + Qn (w)
=⇒ v = G(u) − Qn (w)
=⇒ hv, vi = hv, G(u) − Qn (w)i = hv, G(u)i − hv, Qn (w)i = 0,

since v ∈ G(H) and hv, Qn (w)i = hPn (w) , Qn (w)i = 0. We claim
 

Hn⊥ = Hn⊥ ∩ G(H) ⊕ Qn G(H) .

(3.9)
 

First note that Hn⊥ ∩ G(H) ∩ Qn G(H)

= {0}, since
 

u ∈ Hn⊥ ∩ G(H) ∩ Qn G(H)



=⇒ u = G(w) ∈ Hn⊥ for some w ∈ H, and u = Qn (v) for v ∈ G(H)
=⇒ hu, ui = hQn (v) , G(w)i = hv − Pn (v) , G(w)i
= hv, G(w)i − hPn (v) , G(w)i = 0 − hPn (v) , G(w)i = 0,

since v ∈ G(H) and G(w) ∈ Hn⊥ . The same proof yields the general fact that for
two subspaces V and W of an inner product space, the orthogonal projection of
V ⊥ onto W is orthogonal to V ∩ W . To obtain (3.9), it now suffices to show that
 

dim Hn⊥ ∩ G(H) ≥ dim Hn⊥ − dim Qn G(H) .

94 3. FREDHOLM OPERATOR TOPOLOGY

For this, note that

dim Hn⊥ ∩ G(H) ≥ dim Hn⊥ − codim(G(H))


 
 
⊥ ⊥
= dim Hn⊥ − dim G(H) = dim Hn⊥ − dim Qn G(H) ,
 

since we have shown that Qn |G(H)⊥ is injective. This also follows from

index G = index G + index Pn = index Pn G



=⇒ dim Ker G − dim G(H) = dim Ker Pn G − dim Hn⊥

=⇒ dim G(H) = dim Hn⊥ ∩ G(H) − dim Hn⊥ .


Let
 ∼
(3.10) G
e n := G| e −1 Hn⊥ ∩ G(H) −→ Hn⊥ ∩ G(H) ,
G (H ∩G(H)) : G
e e−1 ⊥
n

and  
⊥ ∼ ⊥
qG,n := Qn |G(H)⊥ : G(H) −→ Qn G(H) .
By (3.9) and (3.8), we have the isomorphisms
∼ ∗
e −1
Λmax (G ∗ max e −1
Hn⊥ ∩ G(H) )∗ −→ Λmax Hn⊥ ∩ G(H) ,

n ) : Λ (G
   ∗
−1 ∗ ⊥ ∼ ⊥
Λmax (qG,n ) : Λmax (Qn G(H) )∗ −→ Λmax G(H) , and
∗ ∼
Λmax Hn⊥ ∩ G(H) ⊗ Λmax Hn⊥ ∩ G(H) −→ C.
 

Then we obtain (3.7), as follows:



Λmax (Ker Pn G) ⊗ Λmax Hn⊥


e −1 H ⊥ ∩ G(H) )∗
= Λmax (Ker(G) ⊕ G

n
  

⊗ Λmax Hn⊥ ∩ G(H) ⊕ Qn G(H)


∗ e −1 Hn⊥ ∩ G(H) )∗
= Λmax (Ker G) ⊗ Λmax (G

  

⊗ Λmax Hn⊥ ∩ G(H) ⊗ Λmax Qn G(H)

.

e −1 ∗ max −1
Via Id ⊗Λmax (G n ) ⊗ Id ⊗Λ (qG,n ), this last tensor product is
∗ ∗

= Λmax (Ker G) ⊗ Λmax Hn⊥ ∩ G(H)
 

⊗ Λmax Hn⊥ ∩ G(H) ⊗ Λmax G(H)

 
∼ ∗ ⊥
= Λmax (Ker G) ⊗ Λmax G(H) .
In other words, for

  ∗
α ∈ Λmax (Ker G) , βn ∈ Λmax G e −1 Hn⊥ ∩ G(H) ,
  

bn ∈ Λmax Hn⊥ ∩ G(H) and an ∈ Λmax Qn G(H)

,

we have (α ⊗ βn ) ⊗(bn ⊗ an ) ∈ Λmax (Ker Pn G) ⊗ Λmax Hn⊥ , and we define


D E
ΨG,n ((α ⊗ βn ) ⊗(bn ⊗ an )) := e −1 )∗ (βn ) , bn α ⊗ Λmax (q −1 )(an ) .
Λmax (G n G,n
3.9. DETERMINANT LINE BUNDLES 95

Note that
Ψ−1
G,n (α ⊗ a) = (α ⊗ βn ) ⊗(bn ⊗ Λ
max
(qG,n )a) ,
D E
where βn and bn are chosen so that Λmax (G e −1
n ) ∗
(β n ) , bn = 1; i.e., bn is dual to
max e −1 ∗ max e −1
Λ (Gn ) (βn ), or equivalently, βn is the dual of Λ (Gn )(bn ). 
Thus, the set (3.5) can be given the structure of a genuine line bundle, namely
that which is induced by the bijection
∗ ∗
Λmax (Ker(Tx )) ⊗ Λmax (Ker(Tx∗ )) ←→ Λmax (Ker Pn Tx ) ⊗ Λmax Hn⊥ ,


and this line bundle structure is unique up to isomorphism. The set (3.5) then
serves as a standard representative of the isomorphism class det T .
2. The Quillen Determinant Line Bundle. What we have done so far is
sufficient for many purposes, but we will go on to construct a restricted version, the
so-called Quillen determinant line bundle q : Q → F and show that det T → X
is isomorphic to the pull-back via T of Q, at least in the case index T = 0.
Definition 3.43. The fiber of q : Q → F above F ∈ F, is simply
∗ ∗
q −1 (F ) := Λdim Ker F (Ker F ) ⊗ Λdim Ker F (Ker F ∗ ) ,
which is clearly a complex line.
However, as with the set (3.5), it is not immediately clear that there are suitable
local trivializations for
[
(3.11) Q := q −1 (F )
F ∈F

with continuous transition functions, because Ker F (or Ker F ∗ ) is not the fiber of
a vector bundle over F about points where F (or F ∗ ) is not surjective. It is true
that for any F ∈ F, there is some nF such that Pn F (H) = Hn⊥ for some suitable
neighborhood, say UF , of F in F. Moreover, one can construct a trivialization for
q|UF : q −1 (UF ) → UF . However, unlike the case of a family T : X → F with X
compact where we had fixed n for which Pn Tx (H) = Hn⊥ for all x ∈ X, note now
that nF is an unbounded function of F ∈ F (noncompact). At the very least, this
causes difficulties in defining the transition functions and exhibiting their continuity.
Instead of attempting this, we opt for an interesting, instructive alternative con-
struction, using transition functions that are determinants of the form det(Id +A)
where A is a trace class operator, defined below. This idea is based on informal
notes of Graeme Segal (see [385]), with extensions elaborated upon by Kenro
Furutani in [158], to whom we are indebted. We shall give this construction
only for Q|F0 , where F0 := {A ∈ F : index A = 0}. This may seem a drawback,
but there are various graceful ways to similarly construct Q over the components
of F with nonzero index. This may not be of crucial importance, since the one
convention is that if index T 6= 0, then det T = 0 if det T is defined at all (e.g.,
for T : Rn → Rm , we have index T = n − m and det T is undefined if n 6= m).
Another convention for index T = k > 0 is to add a zero operator 0 : Ck → H to T
yielding T ⊕ 0 : H ⊕ Ck → H ∈ F0 . Finally, we always can identify the connected
components of F (distinguished by the index), e.g., by shift operators after fixing
a basis for the underlying Hilbert space H. The bundle Q then constructed via
pull-back would, however, depend on the choice of the identifications.
96 3. FREDHOLM OPERATOR TOPOLOGY

For our readers, another potential difficulty in the Segal approach is that one
needs to know the functional analysis of trace class and Hilbert-Schmidt operators
(which is covered in Section 2.7), and Fredholm determinants which we consider
next.

3. Detour on the Classical Fredholm Determinant.


I Before continuing with the various determinant line bundles, we make a long de-
tailed detour into the classical subject of Fredholm determinants. In our context, we
shall explain Ivar Fredholm’s ideas mostly in an abstract and, one could blame us, a
superficial way. Our focus is the rigorous and geometric definition of a determinant line
bundle, generalizing the concept of the index bundle of topological K-theory. A wider and
historically more adequate discussion of the Fredholm determinant is given in [290], where
H.P. McKean presents three serious examples of what you can do with this machinery
in the study of Brownian motion; of the Kortwegde Vries equation (KdV) describing long
waves in shallow water; and of the distribution of eigenvalues of typical unitary matrices
and operators. J

For a complex vector space V with d = dim V finite, let Λk (V ) denote the
k-th exterior product of V (k = 0, . . . , d), where Λ0 (V ) = C. If T ∈ End(V ) (i.e.,
T : V → V is linear), then T induces Λk T ∈ End Λk (V ) , determined by

Λk T (v1 ∧ · · · ∧ vk ) = T v1 ∧ · · · ∧ T vk for k > 0,



(3.12)

and Λ0 T = IdC . Since dim Λd (V ) = 1, Λd T is multiplication by a scalar, and this


scalar is det T . Indeed det T can be defined in this way, and it is straightforward to
show that this agrees with the usual definition in terms of the matrix of T relative
to a basis (e.g., just take k = d and let v1 , · · · , vd be a basis in (3.12)). Note that
under the natural isomorphisms
 ∼ ∗ ∼
End Λd (V ) −→ Λd (V ) ⊗ Λd (V ) −→ C,

and Λd T ∈ End Λd (V ) corresponds to det T . There are immediate problems with
the notion det T when dim V = ∞, but det T can be defined when V is a Hilbert
space and T is sufficiently close to Id (e.g., T − Id is trace class), as we will show.
We begin with the following proposition whose proof can be shortened somewhat
if one presupposes that A can be put in Jordan canonical form.

Proposition 3.44. If V is a vector space with dim V = d < ∞, we have


Xd
Tr Λk A .

det(Id +A) =
k=0

Proof. We define the ε-tensor with the following sign convention:



 1, if j1 . . . jk is an even permutation of i1 . . . ik ,
(3.13) εji11...i
...jk
k
: = −1, if j1 . . . jk is an odd permutation of i1 . . . ik ,
0, otherwise,

and εi1 ...id : = ε1...d


i1 ...id .

In what follows, we use the multi-index notation

(i)k := (i1 , . . . , ik ) and (i)0d−k = (ik+1 , . . . , id ),


3.9. DETERMINANT LINE BUNDLES 97

where i1 , . . . , id run independently from 1 to d. We also use the abbreviated nota-


tions
∧(i)k v := vi1 ∧ · · · ∧ vik and ∧(i)0n−k v = vik+1 ∧ · · · ∧ vid , and
(j)
ε(i)kk = εji11...i
...jk
k
and ε(i)d = εi1 ...id = ε1...d
i1 ...id .

Assuming that v1 , . . . , vd is an orthonormal basis of V , we compute


Λd (Id +A) (v1 ∧ . . . ∧ vd ) = (v1 + Av1 ) ∧ . . . ∧(vd + Avd )


1 X Xd   
= ε(i)d ∧(i)k (Av) ∧ ∧(i)0d−k v
d! (i)d k=0
1 X Xd  1 X 
  
= ε(i)d ∧(i)k (Av) , ∧(j)k v ∧(j)k v ∧ ∧(i)0d−k v
d! (i)d k=0 k! (j)k
 X 
1 X X d 1   
= ε(i)d ∧(i)k (Av) , ∧(j)k v ∧(j)k v ∧ ∧(i)0d−k v
d! (i)d k=0 k! (j)k
 X  
1 X X d 1 (j)k 
= ε(i)d ∧(i)k (Av) , ∧(j)k v ε(i) ∧(i)d v
d! (i)d k=0 k! (j)k k
  
1 X X d 1 X 
= ε(i)d ∧(j)k (Av) , ∧(j)k v ∧(i)d v
d! (i)d k=0 k! (j)k
1 X Xd
Tr Λk A ∧(i)d v
 
= ε(i)d
d! (i)d k=0
X 
d
k

= Tr Λ A (v1 ∧ . . . ∧ vd ) . 
k=0

If H is a separable (complex) Hilbert space with dim H = ∞, and A is a suitable


operator on H, one is tempted to define
X∞
Tr Λk A .

(3.14) det(Id +A) =
k=0
This is certainly finite if A has finite rank, but the terms on the right side need not
exist, even if A is compact. If there is a complete orthonormal system {e0 , e1 , . . .}
with Aei = λi ei , λi ∈ C, then it is reasonable to define
Y∞
det(Id +A) = (1 + λi ) ,
i=0
Qn
provided the infinite product converges (i.e., thePpartial products i=0 (1 + λi )

converge). It is well known that this is the case if i=0 |λi | < ∞; i.e., if A is trace
class. However, we will proceed somewhat differently, in the spirit of Proposition
3.44 and (3.14).
Since H is a fixed separable (complex) Hilbert space with dim H = ∞, through-
out this section we use the notation
n o
B = B(H) := A ∈ End(H) : kAk = supkxk=1 {kAxk} < ∞ and
B × = B × (H) := A ∈ B : A−1 ∈ B ,


and recall from Section 2.7 that I1 denotes the Banach


√ space (and ideal in B) of
trace class operators, with norm kAk1 := Tr |A| = Tr A∗ A;
 see (2.46)) (p.60) and
Propositions 2.79 and 2.80 (p.62). Let Λk : B(H) → B Λk H denote the continuous
linear map determined by
Λk (A)(v1 ∧ · · · ∧ vk ) = Av1 ∧ · · · ∧ Avk .
98 3. FREDHOLM OPERATOR TOPOLOGY


Note that Λk (AB) = Λk (A) Λk (B), and Λk (A∗ ) = Λk (A) since
Λk (A)(v1 ∧ · · · ∧ vk ) , w1 ∧ · · · ∧ wk = hAv1 ∧ · · · ∧ Avk , w1 ∧ · · · ∧ wk i
1 X
= εi1 ···ik hAv1 , wi1 i · · · hAvk , wik i
k!
1 X
= εi1 ···ik hv1 , A∗ wi1 i · · · hvk , A∗ wik i
k! (i)k

= v1 ∧ · · · ∧ vk , Λk (A∗ )(w1 ∧ · · · ∧ wk ) .
Thus,
q q q

Λk (A) = Λk (A) Λk (A) = Λk (A∗ ) Λk (A) = Λk (A∗ A)
r   q
2
= Λk |A| = Λk (|A|) Λk (|A|) = Λk (|A|) .

Hence, the singular values of Λk (A) coincide with those of Λk (|A|), namely the
products µi1 · · · µik , where the µi are the singular values of A. Moreover,
 X 1 X
Λk (A) 1 = Tr Λk (|A|) = µi1 · · · µik ≤ µi1 · · · µik
i1 <···<ik k! i1 ,··· ,ik
1 k 1 k
= Tr(|A|) = (kAk1 ) ,
k! k!
and for any z ∈ C,
X∞ X∞ k
X∞  k
Tr Λk (A) z k ≤ Tr Λk (A) |z| ≤ Tr Λk (A) |z|
 
k=0 k=0 k=0
X∞ k
X∞ 1 k
= k
Λ (A) 1 |z| ≤ (|z| kAk1 ) ≤ e|z|kAk1 .
k=0 k=0 k!
P∞ 
Since k=0 Tr Λk (A) z k exists for any z (in particular z = 1), we may make the
following definition which agrees with the finite-dimensional case.
Definition 3.45. For A ∈ I1 , we define
X∞
Tr Λk (A) ,

det(Id +A) =
k=0
which is known as the Fredholm determinant of Id +A.
Proposition 3.46. For A, B ∈ I1 ,
|det(Id +A) − det(Id +B)| ≤ kA − Bk1 exp(kAk1 + kBk1 ) ,
whence det(Id +(·)) : I1 → C is a continuous function.
 
Proof. Since Tr is linear and Tr Λ0 (A) = Tr Λ0 (B) = Tr(IdC ) = 1,
X∞
Tr Λk (A) − Λk (B) .

|det(Id +A) − det(Id +B)| ≤
k=1
Observe that

Λk (A) − Λk (B) (v1 ∧ · · · ∧ vk ) , v1 ∧ · · · ∧ vk



 
X hBv1 ∧ · · · ∧ Bvp , v1 ∧ · · · ∧ vp i
k−1
=  · h(A − B) vp+1 , vp+1 i ,
p=0 · hAvp+2 ∧ · · · ∧ Avk , vp+2 ∧ · · · ∧ vk i
3.9. DETERMINANT LINE BUNDLES 99

since
Λk (A) − Λk (B) (v1 ∧ · · · ∧ vk )


= Av1 ∧ · · · ∧ Avk − Bv1 ∧ Av2 · · · ∧ Avk


+ Bv1 ∧ Av2 · · · ∧ Avk − Bv1 ∧ · · · ∧ Bvk
= (A − B) v1 ∧ Av2 ∧ · · · ∧ Avk + Bv1 ∧ Av2 ∧ · · · ∧ Avk
− Bv1 ∧ Bv2 ∧ Av3 · · · ∧ Ak + Bv1 ∧ Bv2 ∧ Av3 · · · ∧ Ak − Bv1 ∧ · · · ∧ Bvk
= (A − B) v1 ∧ Av2 ∧ · · · ∧ Avk + Bv1 ∧(A − B) v2 ∧ Av3 · · · ∧ Ak
+ Bv1 ∧ Bv2 ∧ Av3 · · · ∧ Ak − Bv1 ∧ · · · ∧ Bvk
k−1
X
= ··· = Bv1 ∧ · · · ∧ Bvp ∧(A − B) vp+1 ∧ Avp+2 ∧ · · · ∧ Avk .
p=0

1 k
Thus, using Λk (A) 1
= k! (kAk1 ) ,

Tr Λk (A) − Λk (B)


k−1
X p!(k − 1 − p)!
≤ |Tr(A − B)| Tr |Λp (B)| Tr Λk−1−p (A)
p=0
(k − 1)!
k−1 k−1
Tr |A − B| X p k−1−p kA − Bk1 X p k−1−p
≤ kBk1 kAk1 ≤ kBk1 kAk1
(k − 1)! p=0
(k − 1)! p=0
kA − Bk1 k−1
≤ (kBk1 + kAk1 ) ,
(k − 1)!
and so
X∞
Tr Λk (A) − Λk (B)

|det(Id +A) − det(Id +B)| ≤
k=1
X∞ 1 k−1
≤ kA − Bk1 (kBk1 + kAk1 )
k=1 (k − 1)!

≤ kA − Bk1 exp(kBk1 + kAk1 ) . 

Proposition 3.47. For A, B ∈ I1 , we have


det((Id +A)(Id +B)) = det(Id +A) det(Id +B) .
Proof. Note that det((Id +A)(Id +B)) exists, since(Id +A)(Id +B) = Id +A+
B + AB and A + B + AB ∈ I1 by Exercise 2.78 (p.60) and Proposition 2.79 (p.61).
PNA PNB
Let A = i=0 µi (A) h·, ei (A)i fi (A) and B = i=0 µi (B) h·, ei (B)i fi (B) denote
the canonical expansions (see (2.43), p.57) of A and B. Let
Xn
An := µi (A) h·, ei (A)i fi (A) if n < NA
i=0

and An = A if n ≥ NA < ∞. Define Bn similarly. Let


Vn = span {ei (A) , fi (A) , ei (B) , fi (B) : i ≤ n}
and note that An |Vn , Bn |Vn ∈ End(Vn ). Since dim Vn < ∞, we have
det((Id +An |Vn )(Id +Bn |Vn )) = det(Id +An |Vn ) det(Id +Bn |Vn ) .
100 3. FREDHOLM OPERATOR TOPOLOGY

As An → A, Bn → B, and An Bn → AB in (I1 , k·k1 ) by Proposition 2.79, we have

det((Id +A)(Id +B)) = det(Id +A + B + AB)


 
= det lim (Id +An + Bn + An Bn ) = lim det(Id +An + Bn + An Bn ) ,
n→∞ n→∞

since det(Id +(·)) is continuous on (I1 , k·k1 ) by Proposition 3.46. Since An = 0 on


Vn⊥ ,
X∞  X∞
Tr Λk (An ) = Tr Λk (An |Vn ) = det(1 + An |Vn ) ,

det(Id +An ) =
k=0 k=0

and similarly for An + Bn + An Bn . Thus,


det(Id +An + Bn + An Bn ) = det(Id +An |Vn + Bn |Vn + An |Vn Bn |Vn )
= det((Id +An |Vn )(Id +Bn |Vn ))
= det(Id +An ) det(Id +Bn ) , and

det((Id +A)(Id +B)) = lim det(Id +An + Bn + An Bn )


n→∞
= lim det(Id +An ) det(Id +Bn )
n→∞
= det(Id +A) det(Id +B) . 

Corollary 3.48. If A ∈ I1 , then Id +A is invertible ⇔ det(Id +A) 6= 0.


Proof. Note that
−1
(Id +A)(Id +B) = Id +A + B + AB = Id ⇔ B = −A(Id +A) .
−1
For B := −A(Id +A) , B ∈ I1 by Proposition 2.79. Then det(Id +A) 6= 0, since
det(Id +A) det(Id +B) = det((Id +A)(Id +B)) = det Id = 1.
If Id +A is not invertible, then −1 ∈ Spec(A) := {λ ∈ C : λ Id −A ∈ / B × }, and
we know that there is a unit eigenvector e with Ae = −e (since Id +A ∈ F0 ,
we have −1 ∈ / Spece (A) and, actually, −1 ∈ Specp (A)). Let P ∈ B denote the
orthogonal projection onto span(e) given by P (x) := hx, ei e. Note that AP = −P .

The orthogonal projection Q = I − P onto span(e) , obeys P + Q = Id, and
AQAP = −AQP = 0. Thus,
(Id +AQ)(Id +AP ) = Id +A(Q + P ) + AQAP Id = Id +A.
P∞
Since Id +AP = Id −P and det(Id −P ) = k=0 Tr Λk (−P ) = 1 − 1 = 0,
det(Id +A) = det(Id +AQ) det(Id +AP ) = 0. 

Remark 3.49. So, the zeros of the function A 7→ det(Id +A) arise exactly
where dim Ker A > 0. This gives a hint of the intimate relation between index
theory and the geometric study of determinants. It may as well legitimize our
construction of the determinant bundle via the index bundle. For much deeper
relations see, e.g., [315, Chapter X] where index theory yields an obstruction to
the existence of a gauge invariant determinant for classical Dirac operators coupled
to a background field (a connection).
3.9. DETERMINANT LINE BUNDLES 101

4. The Segal-Furutani Construction. We now begin the alternative con-


struction of the restriction q : Q|F0 → F0 based on work of G. Segal [385] with
contributions of K. Furutani [158]. We refer to [64, Appendix] for a short cocycle
definition of the Segal determinant bundle. For a fixed Hilbert space H, recall that
B × denotes the group of invertible elements of the ring B := B(H) and I1 denotes
the ideal of trace class operators.
Proposition 3.50. The space F0 of all Fredholm operators of index 0 on H
equals B × + I1 .
Proof. Let T ∈ F0 , and let σT : H → Ker T denote the orthogonal projection.
∼ ⊥
Since the index of T is 0, there is an isomorphism L : Ker T −→ (Im T ) . Then
L ◦ σT is a finite rank (and hence trace class) operator, and T + L ◦ σT ∈ B × .
Indeed, Ker(T + L ◦ σT ) = {0}, since

0 = (T + L ◦ σT )(v) = T (v) +(L ◦ σT )(v) ∈ Im T ⊕(Im T )
=⇒ T (v) = 0 and L(σT (v)) = 0 =⇒ v ∈ Ker T and v = σT (v) = 0.
Moreover, T + L ◦ σT is surjective, since

(T + L ◦ σT )(Ker T ) = (L ◦ σT )(Ker T ) = (Im T ) and
   
⊥ ⊥
(T + L ◦ σT ) (Ker T ) = T (Ker T ) = Im T.

Thus, T + L ◦ σT ∈ B × and T ∈ B× − L ◦ σT ⊂ B × + I1 . Conversely, any


element of B × + I1 is of the form C + K, where C ∈ B× and K ∈ I1 ⊂ K
by Proposition 2.75. Then C + K ∈ F by Exercise 3.7 (p.65) which makes use
of Theorem 3.2 (Atkinson), p.64. By Exercise 3.10 (p. 65), index(C + K) =
index C = 0, and so C + K ∈ F0 . 
For A ∈ I1 , let
T ∈ F0 : T + A ∈ B × = B × − A ⊂ B × + I1 .

UA :=
In other words, UA consists of all perturbations of invertible operators by −A ∈ I1 .
Then {UA : A ∈ I1 } is an open cover of F0 = B × +I1 in the topology of the operator
norm. To see that B × + I1 = ∪A∈I1 UA , note that if S ∈ B × + I1 , then S = G − A
for some A ∈ I1 and G ∈ B × . Thus, S + A = G ∈ B × and so S ∈ UA . Note that
for T0 ∈ F0 , we have T0 + LT0 ◦ σT0 ∈ B × and so T0 ∈ ULT0 ◦σT0 . For A, B ∈ I1 and
T ∈ UA ∩ UB , let
 
−1
(3.15) gAB (T ) := det (T + A) (T + B) ∈ C.

This is defined, since composing both sides of T + B = (T + A) +(B − A) on the


−1
left with (T + A) yields
−1 −1
(3.16) (T + A) (T + B) = Id +(T + A) (B − A) ∈ Id +I1 .
Then Definition 3.45 applies. Note that (3.16) also implies that gAB is continuous.
Indeed, UA ∩ UB → I1 , given by
−1
T 7→ (T + A) (B − A) ,
−1
is continuous since T 7→ (T + A) ∈ B is continuous in the operator norm topology,
and by Proposition 2.79 the composition multiplication B × I1 → I1 is (jointly)
102 3. FREDHOLM OPERATOR TOPOLOGY

continuous. Then Proposition 3.46 yields the continuity of


 
−1
T 7→ gAB (T ) = det Id +(T + A) (B − A) .
For A, B, C ∈ I1 and T ∈ UA ∩ UB ∩ UC , we have the cocycle condition
gAC (T ) = gAB (T ) gBC (T ) .
Indeed, using Proposition 3.47
   
−1 −1
gAB (T ) gBC (T ) = det (T + A) (T + B) det (T + B) (T + C)
   
−1 −1 −1
= det (T + A) (T + B)(T + B) (T + C) = det (T + A) (T + C) .

Definition 3.51 (Segal Determinant Line Bundle). Let π : S → F0 denote the


line bundle defined by {gAB }; i.e., S is the disjoint union of {UA × C : A ∈ I1 }, but
with the identifications
(T, zA ) ∈ UA × C ∼ (T, zB ) ∈ UB × C
 
−1
⇐⇒ zA = gAB (T ) zB = det (T + A) (T + B) zB .
In particular, for T ∈ B × + I1 , say T − A ∈ B × for A ∈ I1 , the fiber π −1 (T ) can
be written as
π −1 (T ) = [T, zA ] = {(T, zB ) ∈ UB × C : B ∈ I1 and zA = gAB (T ) zB } .
We now show that Q|F0 (with Q defined in (3.11), p.95) can be made into a
genuine line bundle by exhibiting identification of Q|F0 with S. Let T ∈ F0 . As
in Proposition 3.50, we have the orthogonal projection σT : H → Ker T and some

isomorphism L : Ker T −→ Ker T ∗ . Let e1 , . . . , ed be a basis for Ker T , and let

e∗1 , . . . , e∗d denote the dual basis for (Ker T ) . Define
∗
[
(3.17) φA : UA × C −→ Q|F0 := Λd (Ker T ) ⊗ Λd (Ker T ∗ ) by
T ∈UA
 
−1
φA (T, zA ) := zA det (T + L ◦ σT ) (T + A) ·
e∗1 ∧ · · · ∧ e∗d ⊗ L(e1 ) ∧ · · · ∧ L(ed ) , for T ∈ UA , zA ∈ C.
We now show that φA (T, zA ) is independent of the choice of basis e1 , . . . , ed ; later we
show that it is also independent of the choice of L. For a new basis F (e1 ) , . . . , F (ed )
where F ∈ GL(Ker T ), note that we have
(L ◦ F )(e1 ) ∧ · · · ∧(L ◦ F )(ed ) = (det F )(L(e1 ) ∧ · · · ∧ L(ed )) .
If F (ej ) = i Fji ei , then
P
X  X
∗ ∗ ∗
δij = F (ei ) (F (ej )) = F (ei ) Fjk ek = F (ei ) (ek ) Fjk .
k k
i
Since δij = k F −1 k Fjk , we have
P

i X i 

F (ei ) (ek ) = F −1 k = F −1 j e∗j (ek ) , and so
j
∗ −1 i ∗
X 
F (ei ) = F e .
j j
j
Thus,
∗ ∗
F (e1 ) ∧ · · · ∧ F (ed ) = det F −1 e∗1 ∧ · · · ∧ e∗d .

3.9. DETERMINANT LINE BUNDLES 103

Then φT is independent of the choice of basis e1 , . . . , ed , since


∗ ∗
F (e1 ) ∧ · · · ∧ F (ed ) ⊗(L ◦ F )(e1 ) ∧ · · · ∧(L ◦ F )(ed )
= det F −1 det F e∗1 ∧ · · · ∧ e∗d ⊗ L(e1 ) ∧ · · · ∧ L(ed )
 

= e∗1 ∧ · · · ∧ e∗d ⊗ L(e1 ) ∧ · · · ∧ L(ed ) .



We now show that φT is also independent of L. Suppose that L0 : Ker T −→ Ker T ∗
is another isomorphism. Then
L0 (e1 ) ∧ · · · ∧ L0 (ed ) = det L0 ◦ L−1 L(e1 ) ∧ · · · ∧ L(ed ) .


Thus,
 
−1
det (T + L0 ◦ σT ) (T + A) e∗1 ∧ · · · ∧ e∗d ⊗ L0 (e1 ) ∧ · · · ∧ L0 (ed )
 
−1
= det (T + L0 ◦ σT ) (T + A) det L0 ◦ L−1 ·


· e∗1 ∧ · · · ∧ e∗d ⊗ L(e1 ) ∧ · · · ∧ L(ed ) .


Hence, we must show that
   
−1 −1
det (T + L0 ◦ σT ) (T + A) det L0 ◦(τT ◦ L)
 
−1
= det (T + L ◦ σT ) (T + A) , or equivalently,
 
−1
det L0 ◦ L−1 = det (T + L0 ◦ σT ) ◦(T + L ◦ σT )

(3.18) .

For this, note that for v ∈ (Ker T ) ,
(T + L ◦ σT )(v) = T (v) = (T + L0 ◦ σT )(v) ,
and for v ∈ Ker T ,
(T + L ◦ σT )(v) = L(v) , while (T + L0 ◦ σT )(v) = L0 (v) .
Hence, (3.18) holds, since
−1
(T + L0 ◦ σT )(T + L ◦ σT ) = Id(Ker T )⊥ ⊕ L0 ◦ L−1 .


For T ∈ UA ∩ UB , we have
φA (T, zA ) = φB (T, zB ) ⇔ zA = gAB (T ) zB , since
 
−1
zA det (T + L ◦ σT ) (T + A)
 
−1
= gAB (T ) zB det (T + L ◦ σT ) (T + A)
   
−1 −1
= zB det (T + L ◦ σT ) (T + A) det (T + A) (T + B)
 
−1
= zB det (T + L ◦ σT ) (T + B) .
Thus, we have a well defined bijection
(3.19) Φ : S −→ Q|F0 given by Φ([T, zA ]) := φA (T ) .
So far we have not given Q|F0 a topology. At this point, the easiest way to give
Q|F0 a topology is to assert that Φ is a homeomorphism, since the line bundle S
has a topology given by the topologies on the UA × C, after taking the quotient by
the equivalence relation.
104 3. FREDHOLM OPERATOR TOPOLOGY

We summarize what has been done thus far.


Proposition 3.52. If Q|F0 is given the unique topology so that the bijection
Φ : S → Q|F0 in (3.19) is a homeomorphism, then Q|F0 inherits the structure of a
complex line bundle from S. Explicitly, the local trivializations of Q|F0 are given
by the maps
φA : UA × C −→ Q|UA ,
defined in (3.17), and the transition functions are gAB : UA ∩ UB → C, given by
 
−1
gAB (T ) := det (T + A) (T + B) .

Exercise 3.53. Use Proposition 3.42, p.93, and the fact that for any T ∈ F,
there is some nT such that PnT T (H) = Hn⊥T for some neighborhood of T in F,
in order to directly give Q|F0 a possibly different alternate line bundle structure.
Then show that this alternate structure is in fact equivalent to that induced by
Φ : S → Q|F0 .
Corollary 3.54. For a continuous family T = {Tx }x∈X : X → F0 , the
determinant bundle det T → X is the pull-back T ∗ (Q|F0 ) of the line bundle Q|F0 →
F0 . Here
det T := det(index T ).
Proof. For x ∈ X, we have
T ∗ (Q|F0 )x = (Q|F0 )Tx = Λmax (Ker(Tx ))∗ ⊗ Λmax (Ker(Tx∗ )) = (det T )x .
Thus, the fibers of T ∗ (Q|F0 ) coincide with those of det T . This does not yet prove
that det T = T ∗ (Q|F0 ) as line bundles, but this is a consequence of the preceding
Exercise 3.53 if we give Q|F0 its alternate structure. Without giving Exercise 3.53
away, let Pn : H → Hn := {e0 , . . . , en }⊥ be an orthogonal projection so that
Pn Tx H = Hn for all x ∈ X, and note that
∼ det Pn T = Λmax (Ker Pn T )∗ ⊗ Λmax (H ⊥ ) = (Pn T )∗ (Q|F ).
det T = 
n 0

Besides Q|F0 , we now show that there is another way to interpret the line bundle
S. This is also essentially due to Graeme Segal; see [385]. For T ∈ F0 = B × +I1 ,
let
FT := T + I1 and FT× := FT ∩ B × 6= ∅.
(Note that there appears to be a notational problem in using F0 and FT , but since
T ∈ F0 , we never have T = 0.) In particular, FId = Id +I1 is the set of operators
on which the Fredholm determinant (see Definition 3.45 is defined. For S ∈ FT× ,
define a continuous bijection
ΦS : FT −→ FId by ΦS (R) = S −1 R, where R ∈ FT .
Note that S −1 R ∈ FId , since

R = T + B, S = T + B 0 (where B, B 0 ∈ I1 )
−1 −1
=⇒ S −1 R = (T + B 0 ) (T + B) = Id +(T + B 0 ) (B − B 0 ) ∈ Id +I1 = FId .
Also, ΦS is clearly 1-1, and we see that ΦS is onto, as follows. For C = I +B 00 ∈ FId ,
we have (by Proposition 2.79
SC = S(I + B 00 ) = S + SB 00 ∈ FT× + I1 ⊆ B × + I1 .
3.9. DETERMINANT LINE BUNDLES 105

Since C = S −1 SC = ΦS (SC), ΦS is onto. On FId × C, we have an equivalence


relation
(R, z) ∼ (R0 , z 0 ) ⇔ z det R = z 0 det R0 .
Alternatively, there is a map
κ0 : FId × C −→ C given by κ0 (R, z) := z det R,
and the (huge) equivalence classes are just the preimages of points in C. We denote
the set of equivalence classes [R, z] by
n −1 o
EId := {[R, z] : (R, z) ∈ FId × C} = κ0 (w) : w ∈ C .

Since κ0 (Id, z) = z, there is a bijection


κ : EId −→ C given by κ([R, z]) = κ0 (R, z) = z det R,
and EId inherits a vector space structure from C via κ. For S ∈ FT× , we have a
map
κ0T,S : = κ0 ◦(ΦS × Id) : FT × C −→ C , or (for A ∈ I1 )
κ0T,S (T + A, z) = κ0 S −1 (T + A) , z = z det S −1 (T + A) ,
 

which defines an equivalence relation on FT × C. Note that for S 0 ∈ FT× , we have

κ0T,S (T + A, z) = z det S −1 (T + A) = z det S −1 S 0 S 0−1 (T + A)


 

= det S −1 S 0 z det S 0−1 (T + A) = det S −1 S 0 κ0T,S 0 (T + A, z) , and so


  

κ0T,S = det S −1 S 0 κ0T,S 0 .



(3.20)
Since det S −1 S 0 6= 0, κ0T,S 0 defines the same equivalence relation, say ∼T , on


FT × C. For T ∈ F0 = B × + I1 , let
ET : = (FT × C) / ∼T = {[R, z] : (R, z) ∈ FT × C} ,
E : = (T, [R, z]) : T ∈ B × + I1 and [R, z] ∈ ET , and let


p : E −→ B × + I1 be given by p(T, [R, z]) := T.


Note that for S ∈ FT× , ΦS : FT −→ FId induces
[ΦS ] : ET −→ EId given by [ΦS ]([T + A, z]) := [ΦS (T + A) , z] = S −1 (T + A) , z .
 

Then [ΦS ] : ET → EId defines a vector space structure on ET which is independent


of the choice of S ∈ FT× . However, the identification κT,S : ET → C, given by
κT,S ([R, z]) := κ([ΦS ] [R, z]) = κ([ΦS (R) , z]) = z det S −1 R ,


depends on S. Using (3.20), we have


κT,S = det S −1 S 0 κT,S 0 .


For A ∈ I1 , recall that UA := B × − A and that {UA : A ∈ I1 } is an open cover of


B × + I1 . Although κT,S : ET → C depends on S ∈ FT× , for each T ∈ UA , there is
a natural choice for S ∈ FT× , namely S = T + A. Then we may define
ψA : p−1 (UA ) −→ UA × C by (where [R, z] ∈ ET , T ∈ UA )
  
 −1
(3.21) ψA (T, [R, z]) := T, κT,(T +A) ([R, z]) = T, z det (T + A) R .
106 3. FREDHOLM OPERATOR TOPOLOGY

For B ∈ I1 , note that from κT,S = det S −1 S 0 κT,S 0 and (3.15), we have

 
−1
κT,(T +A) = det (T + A) (T + B) κT,(T +B) = gAB (T ) κT,(T +B) .

Thus, for zA = κT,(T +A) ([R, z]) and zB = κT,(T +B) ([R, z]), we have
−1

ψA ◦ ψB (T, zB ) = ψAA (T, [R, z]) = (T, zA ) , where
zA = κT,(T +A) ([R, z]) = gAB (T ) κT,(T +B) ([R, z]) = gAB (T ) zB .
Hence, the transition functions for the bundle p : E → B× + I1 relative to the
trivializations ΨA are the same as those for the bundle π : S → B× + I1 , and so the
ΨA induce an isomorphism

Ψ : E −→ S.
By Proposition 3.52, the bundle S|F0 is isomorphic to the Quillen bundle Q|F0 . In
summary we have
Theorem 3.55. There are three isomorphic complex line bundles over the com-
ponent F0 of Fredholm operators with index zero on a fixed Hilbert space H, namely
π : S −→ F0 , q : Q|F0 −→ F0 and p : E −→ F0 .
For A ∈ I1 , and T ∈ UA = B × − A, the point in q −1 (T ) corresponding to [T, zA ] ∈
π −1 (T ) is
 
−1
(3.22) zA det (T + L ◦ σT ) (T + A) e∗1 ∧ · · · ∧ e∗d ⊗ L(e1 ) ∧ · · · ∧ L(ed ) ,

which is independent of the choice of isomorphism L : Ker T −→ Ker T ∗ and the
choice of basis {e1 , . . . , ed } of Ker T ; here σT is the orthogonal projection of H onto
Ker T . The point in p−1 (T ) corresponding to [T, zA ] ∈ π −1 (T ) is
(T, [T + A, zA ]) ∈ p−1 (T ) ,
since (by (3.21))

ψA (T, [T + A, zA ]) = T, κT,(T +A) ([T + A, zA ])
  
−1
= T, zA det (T + A) (T + A) = (T, zA ) .

Remark 3.56. Recall that


n   o
−1
(3.23) [T + A, zA ] = (R, z) : zA = z det (T + A) R and R ∈ T + I1 .

Taking R = T + L ◦ σT ∈ (T + I1 ) ∩ B × in (3.23), we obtain


  
−1
(T, [T + A, zA ]) = T, [R, zA / det (T + A) R ]
= T, [R, zA det R−1 (T + A) ] .


The zA det R−1 (T + A) in this last expression is precisely the factor multiplying


e∗1 ∧· · ·∧e∗d ⊗L(e1 )∧· · ·∧L(ed ) in (3.22), but neither (3.23) nor (3.22) is determined
by this factor alone.
In the fiber p−1 (T ), there are standard elements
[T, 1] = (R, z) ∈ FT × C : z det S −1 R = det S −1 T for all S ∈ FT×

and
[T, 0] = (R, z) ∈ FT × C : z det S −1 R = 0 for all S ∈ FT× .

3.9. DETERMINANT LINE BUNDLES 107

Actually [T, 0] is just the zero element in p−1 (T ). For T ∈ F0 ,


[T, 1] = [T, 0] ⇐⇒ det S −1 T = 0 for all S ∈ FT× ⇐⇒ S −1 T ∈
/ B × for all S ∈ FT×
/ FT× (i.e., T is not invertible).
⇐⇒ T ∈
The canonical section σ : F0 → E is defined by σ(T ) = [T, 1]. We have just seen
that σ(T ) 6= 0 ⇔ T ∈ B × . On UA , we have (see (3.21))
  
−1
ψA (σ(T )) = ψA ([T, 1]) = T, det (T + A) T .
In other words, the local representative σA : UA → C of σ is given by
   
−1 −1
σA (T ) = det (T + A) T = det Id −(T + A) A .
 
−1
Setting zA := det (T + A) T in (3.22) and noting that
     
−1 −1 −1
det (T + L ◦ σT ) (T + A) det (T + A) T = det (T + A) T
(by Proposition 3.47, p.99), we obtain the Quillen version of the canonical section,
namely σq : F0 → Q|F0 given by
 
−1
σq (T ) = det (T + L ◦ σT ) T e∗1 ∧ · · · ∧ e∗d ⊗ L(e1 ) ∧ · · · ∧ L(ed ) ,
where again L : Ker T → Ker T ∗ is an arbitrary isomorphism and {e1 , . . . , ed } is an
arbitrary basis of Ker T .
Exercise 3.57. Show directly that for T ∈ F0 ,

0 ∈ Λd (Ker T ) ⊗ Λd (Ker T ∗ ) , / B× ,

if T ∈
σq (T ) =
1 ⊗ 1 ∈ C ⊗ C, if T ∈ B × .
From this, it would seem that σq is discontinuous. Is this really the case? One may
wish to read the relevant discussion and footnotes in [315, p.276].
[Hint: Show that the Quillen determinant bundle is the universal line bundle, be-
cause any complex line bundle on a compact space is an induced line bundle of the

Quillen determinant line bundle through the isomorphism K(X) −→ [X, F]. By
this fact, you may establish that the canonical section is not continuous.]
Remark 3.58. a) At the end of the next Section, we provide a brief glimpse
of the fascinating relations (discovered in [377]) between the rich structure of the
determinant line bundle (and the canonical section) to the corresponding object
in physics, namely the zeta-function regularized determinant of Dirac operators.
Presently, we will only note the following recent development. On the infinite-
dimensional manifold F0 , we have the Quillen-Segal determinant line bundle E.
This is a rather complicated object, say of degree one. On the space ΩFId of loops
{T : S 1 → FId }, i.e., T (θ) = Id +Aθ invertible for θ ∈ S 1 with Aθ ∈ I1 , we have
U (1)-valued functions, det(T ) (after normalization) which are simpler objects, say
of degree zero. Currently, a very active research field is the investigation of similar
transgressions between an object of degree k + 1 on a base and a corresponding
object of degree k on the free loop space of the base.
b) The reader will have noticed the extensive use we have made of the produc-
tive property valid for Fredholm determinants. Contrary to that, the ζ-function
regularized determinants of elliptic operators (see below pp.118f) do not obey a
precise multiplicative property. We refer to [253, 193] for a detailed study of the
108 3. FREDHOLM OPERATOR TOPOLOGY

so-called Multiplicative Anomaly. The two subjects, multiplicative anomaly and


determinant line bundle are intimately related to each other: roughly speaking, to
formalize the construction of taking the ratios of zeta-regularized determinants we
need the machinery of the determinant line bundle, see [69, Chapter 11].

10. Essential Unitary Equivalence and Spectral Invariants


I Our main focus in this book is on the index, which is stable under rather general
deformations of the operator in question. However, in this section we wish to broaden
the perspective somewhat by indicating some of the finer attributes of operators which
are captured by more sensitive quantities constructed from their spectra; i.e., spectral
invariants. J

The power and limitations of the homotopy theoretic technique in analysis,


specifically in the theory of Fredholm operators, are demonstrated in results of
Lawrence Brown, Lewis Coburn, Ronald Douglas, Peter Fillmore,
William Helton, Roger Howe and others. We will review them briefly; for
details see the collection [146, 1973].
1. Let B denote the Banach algebra of linear bounded operators on the Hilbert
space H with the closed ideal K of compact operators and the canonical projection
π : B → B/K. For S, T ∈ B, we define (in B/K) essential unitary equivalence or
unitary equivalence mod K, denoted by π(S) ≈ π(T ), by either of the following
equivalent ([146, 1973, p.77]) conditions:
(i) There is a unitary operator U ∈ B (i.e., U ∗ = U −1 ),
such that S − U T U ∗ ∈ K.
(ii) There is a unitary element v in B/K such that π(S) = vπ(T )v ∗ .
Now let S ∈ F; i.e., π(S) is an invertible element in B/K (see Theorem 3.2, p. 64).
What is the relationship in F between
• the topological relation S ∼ T (S and T are homotopic; i.e., S and T
can be connected by a continuous path in F, or equivalently: index S =
index T ) and
• the numerical-analytic relation π(S) ≈ π(T ); i.e., S and T are modulo K
unitarily equivalent?
Since B and even more U are connected (Theorem 3.21), the homotopy equivalence
follows trivially from the unitary equivalence modulo K. The converse is not imme-
diately clear. Rather, by looking, for example, at the homotopic operators Id and
− Id, it is apparent that Fredholm operators from the same path-connected compo-
nent of F may well differ by “more” than a compact operator. However, Brown,
Douglas and Fillmore showed in 1970: In the group of unitary elements of B/K
the relations “S ∼ T ” and “π(S) ≈ π(T )” coincide; i.e., essentially-unitary opera-
tors in Hilbert space can be joined by a continuous path in F (i.e., they have the
same index) if and only if they are unitarily equivalent modulo K. In still another
way:
By [146, 1973, p.71], the classes of unitarily equivalent unitary elements of
B/K form an infinite cyclic group with representatives
π(Id) or π((shift+ )n ) or π((shift− )n ), n ∈ N.
For all essentially-unitary operators T ∈ B (i.e., π(T ) is unitary in B/K), the
essential spectrum Spece (T ) = Spec π(T ) is contained in S 1 = {z ∈ C : |z| = 1)},
3.10. ESSENTIAL UNITARY EQUIVALENCE AND SPECTRAL INVARIANTS 109

since kπ(T )k = 1 (see Exercise 3.6, p.65). Since index(T − z Id) = 0 whenever
|z| > 1 , the homotopy invariance of the index implies that the inclusion Spece (T ) ⊆
S 1 is proper, only if index T = 0. Thus each class of unitarily equivalent unitary
elements in B/K is characterized by the index, where for fixed T the index of T −z Id
is a map on C − S 1 into Z with value 0 everywhere outside S 1 and constant value
n = index T inside S 1 as depicted in Figure 3.3. In geometric language: The index
is a complete unitary invariant for the unitary elements in B/K.

index(T¡zId) index(T¡zId) index(T¡zId)

¡2

Figure 3.3. Three classes of unitarily equivalent elements in B/K,


distinguished by the height n = 4, 1, −2 of the |z| < 1-towers

2. For normal operators (T ∈ N ⇔ T ∗ T − T T ∗ = 0), however, the essential


spectrum is the complete unitary invariant4, while the index is identically 0 in
the complement of the essential spectrum (see Remark 2.11, p. 17). The class of
operators which is the most natural next object of study after the normal and
Fredholm operators are (working modulo K at any rate) the essentially normal
operators, i.e., the operators T ∈ B for which π(T ) is normal, i.e., T T ∗ − T ∗ T ∈ K.
This includes in particular the compact perturbations of normal operators (nothing
new since their index vanishes) and the essentially-unitary operators. We have the
following theorem of Brown-Douglas-Fillmore [98, 1975]: Two essentially normal
operators S and T are unitarily equivalent modulo K, if and only if they have the
same essential spectrum X and if on every connected component of C \ X, we have
index(S − z Id) = index(T − z Id) [146, 1973, p.73-122].
The proof of this theorem with its dependence on delicate questions of topolog-
ical algebra and K-theory is by no means trivial. Let us consider the situation once
more: The essential spectrum X of an essentially normal operator T is a compact
subset of C. The expression index(T − z Id) defines a Z-valued function on C − X.
The connected components of C − X with nonvanishing index are (so to speak) the
obstructions to the normality of T .
3. While pathwise connectedness in F yields nothing but a decomposition
into the Z connected components (a classification by index T − z Id at the point

4I.D. Berg, Trans. Amer. Math. Soc. 160 (1971), 365-371, shows that every normal operator
can be diagonalized by a compact perturbation, and that S, T ∈ N are unitarily equivalent modulo
K (i.e., there is a unitary operator U with S −U T U ∗ ∈ K), if and only if Spece (T ) = Spece (S). For
self-adjoint S, T this result is due to John von Neumann, whose point of departure was a lemma
by Hermann Weyl (1909) saying that the accumulation points of the spectrum of a self-adjoint
operator remain unchanged under perturbation by a compact operator.
110 3. FREDHOLM OPERATOR TOPOLOGY

z = 0) we have in 2 above a finer classification by index towers on the connected


components of the complement of the essential spectrum as depicted in Figure 3.4.

index(T¡zId)

n1
n2

n3

n4

Figure 3.4. Index towers on the connected components of the


complement of the essential spectrum

We state some concrete consequences which form a transition to the next chap-
ter.
(i) An essentially normal operator T with index(T − z Id) = 0 for all z outside
the essential spectrum of T belongs to N + K, i.e., can be written as a sum of a
normal and a compact operator [146, p.118]
(ii) The family N + K is topologically closed (in the operator norm); thus
elements of the complement of N + K in B (even if their index vanishes) cannot be
approximated by a sequence in N + K [146, p.119]. Unfortunately, there is not yet
an elementary proof for this remarkable result.
(iii) It is another very interesting fact, that essentially normal operators whose
essential spectrum is described by the image of a simple, closed curve are, modulo K,
unitarily equivalent to Wiener-Hopf operators with the same characteristic curve.
Details are in the next chapter and in [146, p.73].

What Is a Spectral Invariant? Let A be a set of operators (possibly un-


bounded) on a Hilbert space, and let spec(A, A∗ ) := {(Spec(T ) , Spec(T ∗ )) : T ∈ A}.
Roughly speaking, a spectral invariant of A with values in a set C (typically Z, R
or C) is a function Φ : spec(A, A∗ ) → C which is preserved (i.e., invariant) under a
given set (often a group) G of transformations g : A → A; i.e., Φ(g(T )) = Φ(T ). We
say that a spectral invariant Φ0 is finer than Φ if Φ0 (S) = Φ0 (T ) ⇒ Φ(S) = Φ(T ),
but not conversely. One might expect that finer invariants are better, but they may
be more difficult to compute, and a courser one may solve the problem at hand.
The main spectral invariant that we have considered thus far is
index : spec(F, F ∗ ) −→ Z, given by index T = dim Ker T − dim Ker T ∗ .
This case brings out the point that spec(T ) and spec(T ∗ ) should include information
on the multiplicities of eigenvalues of T and T ∗ (e.g., if Ker T 6= {0} or Ker T ∗ 6=
3.10. ESSENTIAL UNITARY EQUIVALENCE AND SPECTRAL INVARIANTS 111

{0}, the multiplicity of the eigenvalue 0). For the index, the set G could be the
additive group K of compact operators K with K(T ) := T + K. Exercise 3.10
(p. 65) shows that the index is a spectral invariant of F under translations by K.
More generally, by the homotopy invariance of the index (Theorem 3.11, p.68), the
index is a spectral invariant of F under any set of transformations on F that map
each component of F into itself.
Note that any function Φ : spec(A, A∗ ) → C is a spectral invariant under a set
of transformations that leaves spec T and spec T ∗ invariant. For example, this is
the case if G := B × , acting via conjugation (i.e., g(T ) = gT g −1 ), since
gT g −1 − λ Id = g(T − λ Id) g −1 (for g ∈ B × )
shows that the resolvent set of gT g −1 is the same as that for T (see Definition
2.59). Of course, we may form a semi-direct product B × K and let it act on F
via (g, K) ·(T ) = gT g −1 + K. Clearly the index is still a spectral invariant on F
under this larger group. The result of Brown, Douglas and Fillmore in 1 above
that two essentially unitary operators have the same index if and only if they are
unitarily equivalent modulo K, can be interpreted as the statement that there is
no spectral invariant for essentially unitary operators which is finer than the index,
under the subgroup U K ⊂ B × K. However, the result in 2 says that this is
very far from the case when we enlarge the class of operators from the essentially
unitary operators to the essentially normal operators.
In addition to the index, many other spectral invariants have arisen, especially
for (unbounded) differential operators such as the Laplace and Dirac operators.
Referring to Theorem 2.40 (p. 37), consider the unbounded operator D0 =√−iT
in the Hilbert space L2 (S 1 ), where T |C 0 (S 1 ) denotes differentiation and i = −1.
Recall that T , and hence D0 , has dense domain W 1 (S 1 ) ⊂ L2 (S 1 ). Moreover, D0
is self-adjoint by Corollary 2.56, p. 49, and index D0 = 0. For ek (θ) := √12π eikθ ,
we have D0 ek = kek and the standard complete orthonormal system {ek }k∈Z of
L2 (S 1 ) consists of normalized eigenvectors of D0 with simple eigenvalues consti-
tuting Spec(D) = Z. The operator D0 is essentially the so-called Dirac operator
for the circle with the trivial spin structure. More generally, standard (first-order)
Dirac operators can be defined on spinor fields that live on oriented Riemannian n-
manifolds M with spin structures, and indeed certain “twisted” operators of Dirac
type do not require spin structures. We will consider them later in some detail. As
with the primordial example D0 on L2 (S 1 ), the spectra of operators of Dirac type,
say D, over compact spin manifolds have a discrete real spectrum of eigenvalues
(not necessarily simple) which is unbounded above and below. If the eigenvalues of
D are ordered so that |λ1 | ≤ |λ2 | ≤ · · · , then there is some constant C (depending
on M ) such that (see [167, Lemma 1.12.6, p.113])

|λk | ∼ Ck 1/n , where n = dim M .

The Eta Function. The eta function for D is a C-valued function of s ∈ C


defined, for <s sufficiently large, by
−s
X
ηD (s) := (sign λ) mλ |λ| ,
λ∈(spec D)\{0}

where mλ is the multiplicity of the eigenvalue λ. Note that ηD (s) is a measure of


the spectral asymmetry of spec D in the sense that if mλ = m−λ , then ηD (s) = 0.
112 3. FREDHOLM OPERATOR TOPOLOGY

This is the case for D0 :


X k −s
X∞ −s
ηD0 (s) = |k| = (1 − 1) |k| = 0.
k∈Z\{0} |k| k=1

It is known ([167, Lemma 1.13.1, p.114]) that Γ((s + 1) /2) ηD (s) extends to a mero-
morphic function (possibly 0) defined on C, all of whose poles (if any) are simple
and located at points of the form (n + 1 − k)/2, k ∈ N = {1, 2, 3, . . .}. Generally,
the reduced eta invariant of D (not to be confused with the eta function of D
or with ηD (0), also called the eta-invariant) is defined by
1
ηeD := 2 (ηD (0) + dim Ker(D)) mod Z.

The reduced eta invariant makes a natural appearance as a boundary term for
the Atiyah-Patodi-Singer index formula ([40]) for operators of Dirac type, with
certain boundary conditions, on manifolds of even dimension with boundary. In this
case the D in ηeD is an induced tangential Dirac operator on the odd -dimensional
boundary of the manifold.

Example 3.59. Although ηD0 (s) ≡ 0, consider Da := D0 − a for a ∈ R. If


a ∈ Z, then SpecDa = SpecD0 and ηDa (s) ≡ 0. Thus, assume that a ∈
/ Z, so that
0∈/ SpecDa = {k − a : k ∈ Z}. Then
X k−a −s
(3.24) ηDa (s) = |k − a| , for <s > 1.
k∈Z |k − a|

Note that ηDa (s) is periodic of period 1 in the variable a. Also,


X k+a −s
X −k + a −s
ηD−a (s) = |k + a| = |−k + a|
k∈Z |k + a| k∈Z |−k + a|
X −(k − a) −s
= |k − a| = −ηDa (s),
k∈Z |k − a|

whence ηDa (s) is odd in a as well. We will R∞ find that ηDa (0) is defined, but not by
the above sum in (3.24). Using Γ(x) := 0 tx−1 e−t dt, we have
Z ∞
2  λ −s
t(s−1)/2 λe−λ t dt = Γ 12 (s + 1) |λ| for <s > −1 and λ ∈ R \ {0} .
0 |λ|

Indeed, using the change of variable τ = λ2 t,


Z ∞ Z ∞
2 (s−1)/2 −τ −2
t(s−1)/2 λe−λ t dt = λ λ−2 τ e λ dτ
0 0
Z ∞ Z ∞
1−s (s−1)/2 −τ 1−s 12(s+1)−1 −τ
= λ−1 |λ| τ e dτ = λ−1 |λ| τ e dτ
0 0
 λ −s
= Γ 21 (s + 1) |λ| .
|λ|
Thus, with λ = k − a 6= 0,
Z ∞
2  k−a −s
t(s−1)/2 (k − a) e−(k−a) t dt = Γ 1
2 (s + 1) |k − a| ,
0 |k − a|
3.10. ESSENTIAL UNITARY EQUIVALENCE AND SPECTRAL INVARIANTS 113

and so for <s > 1,

1
X k−a −s
+ 1) ηDa (s) = Γ 12 (s + 1)

Γ 2 (s |k − a|
k∈Z |k − a|
Z ∞
X 2
= t(s−1)/2 (k − a) e−(k−a) t dt
k∈Z 0
Z ∞ X 2
= t(s−1)/2 (k − a) e−(k−a) t dt.
0 k∈Z

Since the final expression is analytic in s for <s > −1, it is the analytic continuation
of Γ 21 (s + 1) ηDa (s) which was originally defined only for <s > 1. We claim that
for t > 0,
X 2
 π 3/2 X∞ π 2 k2
(k − a) e−(k−a) t
= −2 ke− t sin(2πka) .
k∈Z t k=1

This is a consequence of the Poisson Summation Formula


X √ X
f (k) = 2π fb(2πk)
k∈Z k∈Z

which holds for rapidly decreasing functions f ∈ C↓∞ (R) (and under less stringent
conditions; see [61, p.445]). We apply this to

2 −iξ −ξ2 /4t −iaξ


f (x) = (x − a) e−(x−a) t , for which fb(ξ) = 3/2
e e .
(2t)

Then
X 2
(k − a) e−(k−a) t
k∈Z
√ X √ X −i2πk 2
= 2π fb(2πk) = 2π 3/2
e−(2πk) /4t −i2πak
e
k∈Z k∈Z (2t)
 3/2 X
2π ∞ 2 2
ke−π k /t e−i2πak − ei2πak

= −i
2t k=1
 π 3/2 X∞ 2 2
= −2 ke−π k /t sin(2πak) .
t k=1

Hence, for 1 < <s < 2

1
X k−a −s
+ 1) ηDa (s) = Γ 12 (s + 1)

Γ 2 (s |k − a|
k∈Z |k − a|
Z ∞ Z ∞
X 2 X 2
= t(s−1)/2 (k − a) e−(k−a) t dt = t(s−1)/2 (k − a) e−(k−a) t dt
k∈Z 0 0 k∈Z
Z ∞  π 3/2 X∞ 2 2
= −2 t(s−1)/2 ke−π k /t sin(2πak) dt
0 t k=1
X∞ Z ∞ 1 
3
s−2 −π 2 k2 /t
= −2π 2 k t 2 e dt sin(2πak) .
k=1 0
114 3. FREDHOLM OPERATOR TOPOLOGY

This last expression is analytic for <s < 2. Thus, it is the analytic continuation of
Γ 12 (s + 1) ηDa (s) for <s < 2. For s = 0, we get


√ X∞ Z ∞ 
3
−2 −π 2 k2 /t
πηDa (0) = −2π 2 k t e dt sin(2πak)
k=1 0
3
X∞ 1
= −2π 2 k sin(2πak)
k=1 π 2 k 2
2 X∞ 1
= −√ sin(2πak) .
π k=1 k
R 1/2 −2 2
P∞ 1
Since 4 0 (2x − 1) sin(2kπx) dx  =1 kπ , we have that − π k=1 k sin(2πxk) is the
Fourier sine series of 2x − 1 on 0, 2 . We then have
2 X∞ 1
ηDa (0) = − sin(2πak) = 2a − 1 (for 0 < a ≤ 21 ).
π k=1 k
Since ηD0 (s) = 0 and ηDa (s) is odd and periodic in a of period 1, ηDa (0) = 2a − 1
for 12 < a < 1 and ηD1 (0) = 0. Thus, ηDa (0) is the periodic extension (of period 1)
of 
0, for a = 0 or 1,
a 7→
2a − 1, for 0 < a < 1.
Thus ηDa (0) is a discontinuous function with a jump −2 as a crosses each integer.
The reduced eta invariant of Da is then
1 1
ηeDa = 2 (ηDa (0) + dim Ker(Da )) mod Z = 2 (2a − 1) mod Z
− 21 mod Z a + 12 mod Z
 
= a = for all a ∈ R.
It is easy to check the relation
d
(3.25) Da := −i − a = eixa D0 e−ixa for all a ∈ R.
dx
For integer a, (3.25) can be read as a special kind of unitary equivalence between the
operators Da and D0 with Ua := eixa unitary operator on L2 (S 1 ) and Ua∗ = U−a .
Note that Ua and Ua∗ keep the domains of the operators D0 and Da (namely the
first Sobolev space W 1 (S 1 ), defined in(2.20) on page 35) invariant. That explains
Spec Da = Spec D0 . If a ∈ R \ Z, the transformation Ua is still unitary with
Ua∗ = U−a , but it does not keep W 1 (S 1 ) ⊃ C 0 (S 1 ) invariant. Hence we obtain a
different spectral situation.
At this place, we shall not re-formulate (3.25) in the language of essential
unitary equivalence. We shall turn back to the example later when we discuss the
symbolic calculus in Part II, Chapter 7, pages 193ff.
The Zeta Function. In contrast to the eta function, the zeta function is
typically defined for certain unbounded operators P , such as Laplacians or squares
of Dirac operators, with discrete spectrum which is positive (or more generally
in an unbounded wedge containing the positive real axis). (However, see [377]
in which zeta functions for operators of Dirac type are defined.) In the context of
Laplacians for Riemannian manifolds the zeta function made an early appearance in

the seminal paper [302] of S. Minakshisundaram and Å. Pleijel. Let {λk }k=1 ,
denote the eigenvalues of P with positive real part, ordered so that 0 < <λ1 ≤
<λ2 ≤ · · · (repeated according to multiplicity). We define zeta function of P by
X∞
ζP (s) := λ−s
k .
k=1
3.10. ESSENTIAL UNITARY EQUIVALENCE AND SPECTRAL INVARIANTS 115

If P is a self-adjoint, elliptic differential operator of order d over a compact n-


manifold, this sum converges to an analytic function for <s > n/d, since |λk | ∼
Ck d/n (see again [167, Lemma 1.12.6, p.113]). It turns out that ζP (s) extends to
a meromorphic function (still called the zeta function of P and still denoted by
ζP (s)) on all of C. Assuming further that P is positive semi-definite, all of the
poles of ζP (s) are simple and they form a subset of {(n − k − 1) /d : k ∈ N}; see
[167, Theorem 1.12.5, p.112].
Example 3.60. For P = D02 = −d2 /dθ2 , we have eigenvalues k 2 , for k =
0, 1, 2, . . ., each of multiplicity 2. Thus, in this case
X∞
ζD02 (s) = 2 k −2s = 2ζ(2s) ,
k=1

where ζ is the well-known Riemann zeta function. Since ζ(z) is known to be analytic
except for a simple pole at z = 1 with residue 1, we have that ζD02 (s) is analytic
except for a simple pole at s = 1/2 = n/d with residue 1.
Closely related to the zeta function is the trace of the heat kernel for P,
namely
X∞
Tr e−tP := e−λk t .

k=nP

Here the sum is over all of the eigenvalues of P , say


λnP ≤ λnP +1 ≤ · · · ≤ λ1 ≤ λ2 ≤ · · · ,
not just the positive ones λ1 ≤ λ2 ≤ · · · . If π+ is the projection onto the closed
subspace spanned by the eigenspaces of P with positive eigenvalues λ1 ≤ λ2 ≤ · · · ,
then the so-called renormalized heat trace for P is
 X∞ −λ t  X0
(3.26) Tr e−tP π+ = e k = Tr e−tP − e−λk t .
k=1 k=nP

For τ = λt,
Z ∞ Z ∞ Z ∞
s−1 −λt s−1 −τ −1 −s
t e dt = (τ /λ) e λ dτ = λ τ s−1 e−τ dτ = λ−s Γ(s) .
0 0 0

Thus, for <s > n/d,


X∞ −1
X∞ Z ∞
ζP (s) = λ−s = Γ(s) ts−1 e−λk t dt
k=1 k k=1 0
Z ∞ X∞ Z ∞
−1 −1
ts−1 e−λk t dt = Γ(s) ts−1 Tr e−tP π+ dt.

(3.27) = Γ(s)
0 k=1 0
R∞
For a suitable function f (t), the function s 7→ 0 ts−1 f (t) dt is the so-called Mellin
transform of f (t). Thus, (3.27) says that ζP (s) is the Mellin transform of the
renormalized heat trace Tr e−tP π+ . The residues of the poles of Γ(s) ζP (s) are


spectral invariants in that they depend only on the spectrum of P . For any ε > 0,
Z ε Z ∞ 
s−1 −tP π+ s−1 −tP π+
 
Γ(s) ζP (s) = t Tr e dt + t Tr e dt .
0 ε

Since the second integral


 is analytic, the residues of Γ(s) ζP (s) only depend on the
behavior of Tr e−tP π+ for small t > 0. As we will do later, at least for certain
116 3. FREDHOLM OPERATOR TOPOLOGY

natural geometric operators P , it is possible to develop an asymptotic expansion


1 
(3.28) Tr e−tP ∼ n/d a0 (P ) + a1 (P ) t1/d + · · · + aN (P ) tN/d

t  
+ O t(N +1)/d as t → 0+ ,
where the ak (P ) are integrals of certain functions on M , which are expressible in
terms of the coefficients of P and their derivatives. Note that
X0 X∞  X0 m

e−λk t = 1
m! (−λ k ) tm .
k=nP m=0 k=nP

By (3.26), the coefficients of the asymptotic expansion (as t → 0+ ) of Tr e−tP π+




will have coefficients ãk (P ) which generally differ from ak (P ) when (k − n) /d ∈


{0, 1, 2, . . .}, namely
X0 m
1
ãk (P ) = ak (P ) − m! (−λk ) if m = (k − n) /d ∈ {0, 1, 2, . . .} .
k=nP
Note that for 0 ≤ k ≤ N,
Z 1 Z 1
1
ts−1 t(k−n)/d dt = ts−(n−k)/d−1 dt = for s > (n − N ) /d.
0 0 s −(n − k) /d
It follows that the residue of Γ(s) ζP (s) at the point s = (n − k) /d is ãk (P ); i.e.,
(3.29) Ress=(n−k)/d (Γ(s) ζP (s)) = ãk (P ) , k ∈ {0, 1, 2, . . .} .
It is well-known that about s = 0, Γ(s) has the
Pn initial
 Laurent expansion Γ(s) =
s−1 + γ + · · · , where γ = − limn→∞ log n − k=1 k1 = 0.5772 . . . is Euler’s con-
stant. Thus, ζP (s) is regular at s = 0, and
(3.30) ζP (0) = Ress=0 (Γ(s) ζP (s)) = ãn (P ) .
Note that ãn (P ) is the term in (3.28) which is t-independent. In the case that
P is the square of an operator D of Dirac type on a manifold of even dimension,
ãn (P ) is the index of the restriction, say D+ , of D to the space of positive spinor
fields with the space of negative spinor fields as the codomain (see Chapter 17 in
Part IV). Moreover, when P = D2 , ãn (P ) = index D+ is a topological invariant.
Sometimes (but not always!) ãn (P ) is a topological invariant even when P is not
of the form D2 , as in the following.
Example 3.61. There is a notion of a Laplace operator ∆ defined on the space
C 2 (M, R) of functions on manifold M with a Riemannian metric. In the case of a
smooth, compact surface M embedded in R3 and f ∈ C 2 (M, R), one can define ∆f
as the restriction to M of the ordinary Laplacian for R3 of the extension, say f¯,
of f to a neighborhood of obtained by constantly extending f along line segments
normal to M ; i.e., ∆f := ∆f¯ |M . Then P = −∆ has a discrete spectrum
0 = λ0 < λ1 ≤ λ2 ≤ . . .(see [167, Lemma 1.6.3]). As t → 0+ , we have the
asymptotic expansion
 X∞ −λ t
1 + Tr e−tP π+ = Tr e−tP = Tr et∆ =
 
e k
k=1
Z  Z   Z  
1
(3.31) ∼ dA + 31 K dA t + 15 1
K 2 dA t2 + · · · ,
4πt M M M
where K is the Gaussian curvature and dA is the element of area (see theR seminal
paper [291] of H. P. McKean, Jr. and I. M. Singer). Note that M dA is
3.10. ESSENTIAL UNITARY EQUIVALENCE AND SPECTRAL INVARIANTS 117

simply the area of M , which is then a spectral invariant of ∆; i.e., two surfaces
with the same spectrum for ∆ must have√thesame area. For an eigenfunction
uk with −∆uk = λk uk , w(p, t) := cos λk t uk (p) is clearly a solution of the
wave equation wtt = ∆w. Hence, the λk are proportional to the frequencies λk /2π
of possible fundamental harmonic tones emitted from the surface. In this sense,
spectral invariants of ∆ are quantities that can be heard, since they are determined
by the set of these tones. In particular, (3.31) impliesRthat the area of M can be
heard. We can also hear the total Gaussian curvature M K dA whose significance
is explained as follows. The Gaussian curvature at (0, 0, 0) of the surface z =
1 2 2
2 k1 x + k2 y is k1 k2 , which is negative for a hyperbolic paraboloid (saddle) and
positive for an elliptic paraboloid. At an arbitrary point p of a surface M in R3 ,
K is defined the same way by means of the best quadratic approximation to M
in a coordinate system centered at p and adapted toR M with the z-axis normal to
M at p. The Gauss-Bonnet Theorem asserts that M K dA = 2π(2 − 2g), where
the so-called genus g is the number of holes of M (e.g., g = 1 for a torus, and
g = 0 for a sphere). Intuitively, the more holes M has, the more negative Gaussian
curvature M has. Thus, a torus and a sphere not only look different, but they also
sound different, even if they have the same area. Incidentally, 2 − 2g is the Euler
characteristic
χ(M ) = #f aces − #edges + #vertices,
for a triangulation of M . At any rate, (3.31) tells us that
Z
(3.32) K dA = 2π(2 − 2g) = 2πχ(M )
M

can be heard. Moreover, M K 2 dA can be heard, as well as all of the higher order
R

terms in (3.31) which involve derivatives of K. These terms are computable, but
with efforts that soon exceed the rewards, especially in higher dimensions.
Example 3.62. Let 0 = λ0 < λ1 ≤ λ2 ≤ . . . . denote the eigenvalues of −∆
for a compact Riemannian manifold M as in Example 3.61. Let C0∞ (R) denote the
space of compactly supported, R-valued C ∞ functions on R. Let W : C0∞ (R) → R
denote the linear functional defined by
X∞ Z ∞ p
W (f ) := f (t) cos( λk t) dt, for f ∈ C↓∞ (R) .
k=0 −∞
P∞ √
Even though the sum k=0 cos( λk t) may not converge, one writes
X∞ p
W = cos( λk t) in the distributional sense,
k=0

since W is a distribution (a continuous linear functional on C0∞ (R) with the


topology of uniform convergence of each derivative on each compact subset of
R). The distribution
√  W has the interpretation as the trace of the wave ker-
nel Tr cos t −∆ as opposed to the heat kernel. We say that x ∈ R is in the
singular support of W (denoted by sing suppR W ) if there is no open interval
I about x and F ∈ C ∞ (I), such that W (f ) = I F (x)f (x)dx for all f ∈ C0∞ (R)
with f |R\I = 0. It is clear that 0 ∈ sing supp W . At least for generic M , it has
been proven (see [132] for much more) that (sing supp W ) \ {0} is the closure of the
length spectrum of M which is the set of multiples of lengths of smoothly closed
geodesics (curves whose sufficiently short subarcs are of minimal length between
118 3. FREDHOLM OPERATOR TOPOLOGY

their fixed endpoints) of M . In other words, at least generically, the closure of the
length spectrum of M is a spectral invariant.
Remark 3.63. From the above examples, one may have the impression that the
spectrum of ∆ for a compact Riemannian manifold contains so much information
that it might even determine M up to isometry. The first counterexample was
discovered by John Milnor [298] who found that the quotients of R16 by the
lattices E8 × E8 and E16 provide two flat tori which are isospectral (i.e., have
the same spectrum for ∆) but not isometric to each other. Since then a large
variety of counterexamples have been found, including one-parameter families of
nonisometric isospectral deformations. Moreover, nonisometric isospectral surfaces
were first found in 1992 (see [182]), which led to a negative answer to the query of
Mark Kac, “Can one hear the shape of a drum?” (see [237]).
For a survey of this and many other related topics, consult Chapter 9 of the
truly monumental book [54] of Marcel Berger. A short list of spectral invariants
derived from first coefficients of the asymptotic expansion of the heat kernel can
be found in [144]. Much longer lists are given in [168, 169]. Of special interest to
mathematicians are the monographs and reviews by Emilio Elizalde [142] and
Dmitri Vassilevich and collaborators [155, 421] which discuss main spectral
functions appearing in the context of modern physics, in particular quantum field
theory.
As will become clear from the asymptotic formula for the heat kernel (elabo-
rated below in Section 17.4, p. 549ff), in general, information about the whole spec-
trum can not be gained from the heat kernel asymptotics alone but requires insight
into the derivatives of the heat kernel and other tools. For the zeta-regularized de-
terminant, this is explained in the following section. For details see also our reviews
[60, Section 3.2] and [69]. Bauer et al. [51] give an interesting review of related
work on homogeneous spaces where new relations for the Hurwitz zeta-function
are obtained and representations and characters of the underlying symmetry group
enter into the calculations.
The Zeta Regularized Determinant. While a self-adjoint, elliptic differ-
ential operator P of order d over a compact n-manifold, with spectrum bounded
below, is far from possessing a Fredholm determinant in the sense of Definition
3.45 (p. 98), there is the so-called zeta regularized determinant of P defined and
motivated as follows. Note that for <s > n/d, we have
X∞  X∞ d −s log λ X∞
−s
ζP0 (s) = d
ds λk = ds e
k
= −λ−s
k log λk .
k=1 k=1 k=1
If P
one sets s = 0, then the right side becomes the formal undefined expression

− k=1 log λk which can formally be rewritten as other undefined expressions:
X∞ Y∞  
− log λk = − log λk = − log det P |π+ H ,
k=1 k=1

where π+ H is the projection of H onto the closure H+ of the span of theeigenspaces


of P with positive eigenvalues. Although as it stands, − log det
 P |H+ is a purely
formal meaningless expression, we can define − log det P |H+ to be ζP0 (0) , which
does exist since the meromorphic extension ζP (s) is regular at s = 0. Then the
zeta determinant of P |H+ is defined to be
0
detζ P |H+ := e−ζP(0) .

3.10. ESSENTIAL UNITARY EQUIVALENCE AND SPECTRAL INVARIANTS 119

This notion of determinant appeared in the 1971 paper [351] of D. Ray and I. M.
Singer. If 0 is not an eigenvalue of P , then it is natural to define
0
detζ (P ) := λnP · · · λ0 e−ζP(0) ,
where we recall that λnP ≤ · · · ≤ λ0 are the nonpositive eigenvalues of P . If 0 is
an eigenvalue of P , the consensus seems to be to eliminate it by restricting P to

the (Ker P ) . Of course, it would be nice to have a way of computing ζP0 (0). By
(3.27),
Z ∞
ts−1 Tr e−tP π+ dt = Γ(s) ζP (s) = s−1 + γ + · · · (ζP (0) + ζP0 (0) s + · · · )
 
0
= ζP (0) s−1 +(ζP0 (0) + γζP (0)) + · · · .
By (3.30), we then have the (rather intractable) formula
Z ∞ 
ζP0 (0) = −γãn (P ) + lim ts−1 Tr e−tP π+ dt − ãn (P ) s−1 .

(3.33)
s→0 0
Although ζP0 (0) is not locally computable, as with an (P ) (or ãn (P ) when P is
semi-definite), it is a sensitive spectral invariant with important applications not
only to quantum physics in relation to anomalies (see [315, Chapter X]), but it
has also been used in other contexts, e.g., to show the compactness of the space of
nonisometric compact surfaces with a given spectrum for ∆ (see [325]).
At first sight, Fredholm determinants and determinant line bundles of the pre-
vious section and the zeta-regularized determinants discussed here seem to have
little in common. It is a very remarkable development arising from the work of re-
searchers such as Jinsung Park, Simon Scott and Krzysztof Wojciechowski
(see e.g., [330], [375] and [377]) that there are relations between quotients of the
respective determinants. While we cannot go into the details here, perhaps the
reader can experience the flavor of such relations by simply looking at one of them,
say the following formula which is explained and proved in [375]:
∗ 
detζ (∆P1 ) detF S(P1 ) S(P1 )
= ∗ .
detζ (∆P2 ) detF S(P2 ) S(P2 )
Here (for i = 1 or 2), ∆Pi is essentially a Dirac Laplacian (i.e., DP∗ i DPi where DPi is
an operator of Dirac type on a manifold with boundary), Pi is a suitable boundary

condition, and S(Pi ) S(Pi ) is a boundary Laplacian, involving a generalized scat-
tering operator S(Pi ). Moreover, such formulas have interpretations in the context
of determinant line bundles over suitable spaces of boundary conditions.
It was remarked above that usually the finer a spectral invariant is, the more
difficult it is to compute. In order of increasing computational difficulty, we gen-
erally have: the index, the reduced eta invariant, the eta invariant, and the zeta-
determinant which seems to be the most delicate and informative of the four thus
far.
The thrill that spectral theory gives was formulated beautifully by Gerd
Grubb in the announcement of her Retirement Lecture, [191]: “It has been an
influential subject and an inspiration for my research through the times, giving me
the opportunity of 1) solving concrete questions related to geometry, 2) develop-
ing general theories and methods, in particular for boundary value problems, 3)
meeting with international researchers in the related fields.”
CHAPTER 4

Wiener-Hopf Operators

Synopsis. The Reservoir of Examples of Fredholm Operators. Origin and Funda-


mental Significance of Wiener-Hopf Operators. The Characteristic Curve of a Wiener-Hopf
Operator. Wiener-Hopf Operators and Harmonic Analysis. The Discrete Index Formula.
Noether’s Theorem for the Hilbert Transform. The Case of Systems. The Continuous
Analogue.

1. The Reservoir of Examples of Fredholm Operators


We already proved some deep theorems on Fredholm operators, but our supply
of examples is still very small, even trivial, as we only studied the following types
of Fredholm operators:
(1) The identity operator Id.
(2) The shift operator shift+ (with respect to an orthonormal basis); see Ex-
ample 1.3.
(3) The Riesz operators Id +K, where K is an operator with finite rank or,
more generally, a Hilbert-Schmidt integral operator of the form
Z
(Ku)(x) := G(x, y)u(y) dy
X
with square integrable weight function G; see Exercise 2.29.
(4) The differentiation operator on the first Sobolev space W 1 (S 1 ) and its
parametrix; see Theorem 2.40 and Exercise 2.42.
All the other Fredholm operators which appeared so far were elementary function-
analytic modifications of the above three basic types. For instance, the left-handed
shift is the adjoint of the right-handed shift, i.e., shift− = (shift+ )∗ . Further,
the (unitary) Fourier transformation (Exercise A.2e) F : L2 (R) → L2 (R) can be
written as a direct sum F = i Id ⊕(− Id) ⊕ (−i Id) ⊕ Id by decomposing L2 (R) into
a direct sum of four closed subspaces H1 , H2 , H3 , H4 , which are the eigenspaces
of F for the eigenvalues in . (A proof which explicitly exhibits the eigenfunctions,
the Hermite functions, can be found in [134, p.97-101].) Viewed in this fashion,
from the standpoint of our abstract operator theory on Hilbert space, the Fourier
transformation is nothing but a trivial modification of the identity.
We will now enlarge our supply of examples by a class of operators which is
connected to all four of the basic types (the relation to differentiation were disclosed
by Louis Boutet de Monvel in his profound study of local elliptic boundary
value problems in [87]): The Wiener-Hopf operators of the form Id +K, defined on
the Hilbert space L2 [0, ∞], where
Z ∞
Ku(x) := k(x − y)u(y) dy, for x ≥ 0 and k ∈ L1 (R).
0

120
4.2. ORIGIN AND FUNDAMENTAL SIGNIFICANCE OF WIENER-HOPF OPERATORS 121

We will give some background information before developing the mathematical


theory of these operators and the analogous discrete operators
X
S : L2 [Z+ ] −→ L2 [Z+ ], where (Su)n := fn−k uk ,
k≥0

whereby the fn are, for example, the Fourier coefficients of a continuous function
f ∈ C 0 (S 1 ) on the circle S 1 .

2. Origin and Fundamental Significance of Wiener-Hopf Operators


I Norbert Wiener wrote (1954) in his autobiography:
“However, the best of the work which he (Eberhard Hopf) and I un-
dertook together concerned a differential equation occurring in the study
of the radiation equilibrium of the stars. Inside a star there is a region
where electrons and atomic nuclei coexist with light quanta, the material
of which radiation is made. Outside the star we have radiation alone, or
at least radiation accompanied by a much more diluted form of matter.
The various types of particles which form light and matter exist in a
sort of balance with one another, which changes abruptly when we pass
beyond the surface of the star. It is easy to set up the equations for this
equilibrium, but it is not easy to find a general method for the solution
of these equations.
The equations for radiation equilibrium in the stars belong to a type
now known by Eberhard Hopf’s name and mine. They are closely re-
lated to other equations which arise when two different physical regimes
are joined across a sharp edge or a boundary, as for example in the
atomic bomb, which is essentially the model of a star in which the sur-
face of the bomb marks the change between an inner regime and an
outer regime; and, accordingly, various important problems concerning
the bomb receive their natural expression in Hopf-Wiener equations.
The question of the bursting size of the bomb turns out to be one of
these.
From my point of view, the most striking use of Hopf-Wiener equa-
tions is to be found where the boundary between the two regimes is in
time and not in space. One regime represents the state of the world up to
a given time and the other regime the state after that time. This is the
precisely appropriate tool for certain aspects of the theory of prediction,
in which a knowledge of the past is used to determine the future. There
are however many more general problems of instrumentation which can
be solved by the same technique operating in time. Among these is the
wave-filter problem, which consists in taking a message which has been
corrupted by a simultaneous noise and reconstructing the pure message
to the best of our ability.
Both prediction problems and filtering problems were of importance
in the last war and remain of importance in the new technology which
has followed it. Prediction problems came up in the control of anti-
aircraft fire, for an anti-aircraft gunner must shoot ahead of his plane
as does a duck shooter. Filter problems were of repeated use in radar
design, and both filter and prediction problems are important in the
modern statistical techniques of meteorology.” (N.W.: I am a Mathe-
matician, Victor Follancz Ltd., London, 1956.)
Here we cannot treat all three main areas of application of Wiener-Hopf operators
mentioned by Norbert Wiener, namely (i) the analysis of boundary-value problems,
122 4. WIENER-HOPF OPERATORS

(ii) filter problems in information theory, and (iii) time series analysis in statistics. We
have to concentrate on the aspect (i) (see Section 9.4 and Chapter 10, where we intend
to clarify the connection with topological-geometric questions). But it is useful for this
purpose to have an idea of the other applications, since it simplifies the transfer of the
methods in (ii) and (iii) to our area (i). J

3. The Characteristic Curve of a Wiener-Hopf Operator


From information sciences we are interested in the stance taken by electrical
engineers: The computation in C and the Fourier analysis of electric oscillations
with the classification of filters (or more generally control circuits) by the geometric
shape of the characteristic curve as depicted in Figure 4.1. Imagine a filter K acting
on an input signal u resulting in an output
Z ∞
Ku(x) = k(x − y)u(y) dy.
−∞

Figure 4.1. Scheme of a filter (left) and specifying a transmis-


sion region in the corresponding amplitude ratio curve (right)

Such linear, time independent and (if k(x) = 0 for x < 0) purely past-dependent
filters are good models for many devices of physics and technology. The information
scientist measures such channels of information by processing a pure sine wave
u(x) = eiωx through the filter
Z ∞ Z −∞
Ku(x) = k(x − y)eiωy dy = (with z = x − y) = − k(z)eiω(x−z) dz
−∞ ∞
Z ∞
iωx −iωz iωx b
=e k(z)e dz = e k(ω),
−∞

and sketching the characteristic values b k(ω) as a function of the phase ω or the
frequency 1/ω. Note that here and throughout the rest of this chapter, we define
R∞ √
k(ω) := −∞ k(z)e−iωz dz without the factor 1/ 2π which would only serve as a
b
distraction in the current context.
The amplitude ratio |b k(ω)| is only one measure for the linear distortion indicat-
ing its reinforcement or weakening. From it the transmission region [ω0 , ω1 ] may be
found via the condition |b k(ω)| ≥ κ. However, a true harmonic analysis is achieved
only if the phase shift, i.e., the argument of the complex number b k(ω), is taken into
account. The nonlinear distortion is given essentially by the shape of the curve
{bk(ω) : ω ∈ R}, see Figure 4.2. This characteristic curve, filter characteristic or
4.4. WIENER-HOPF OPERATORS AND HARMONIC ANALYSIS 123

periodogram coincides under certain conditions with the essential spectrum of the
operator K; see Theorem 4.14 (p. 129) below.

Figure 4.2. Left: characteristic curve of a filter without feedback.


Right: feedback is possible in a certain region. Note: the linear
distortion of both filters may be the same

For details, in particular for the relationship with the general theory of electric
circuits we refer to [134, p.170-176] and the literature quoted there.

4. Wiener-Hopf Operators and Harmonic Analysis


I We have just pointed out how complex analysis, with its varied geometric-topological
aspects, enters markedly into operator theory through information theory. Roughly, the
real reason is that the Fourier transform of a square-integrable function k which vanishes
identically on the left half-line is holomorphic on the upper half-plane C+ , see Figure 4.3
and [134, p.161f]. Formulated differently, the reason is that the position of singularities
in C of certain functions associated with dynamical systems carries information about the
asymptotic behavior of the oscillating system.

k^

C+

Figure 4.3. Fourier transform of a square-integrable function


which vanishes identically on the left half-line

The methods of complex analysis thus introduced are based on the idea (founded in
the notion of a holomorphic function expandable in a power series) of quantities which
vary smoothly and continuously and which are ultimately completely determined through
the knowledge of the function value and those of the derivatives at a single point. In
124 4. WIENER-HOPF OPERATORS

contrast, the statistical theory of time series analysis rests on the theory of real functions
and thus enters into functional analysis an experience of dealing mathematically (in the
framework of harmonic analysis) with curves which are pieced together from unrelated
parts.
With the terminology of the preceding section, we have (roughly) that every operator
on the past of u(x) which is linear and invariant under translation of the time origin can
be represented as a filter Z ∞
Ku(x) = k(y)u(x − y) dy
0
or as the limit of a sequence of such operators. If K is defined in this fashion as a linear
statistical prediction operator, for example, then the method of least squares yields an
optimality criterion of minimizing
Z ∞
|u(x + a) −(Ku) (x)|2 dx,
−∞

where a is a given prediction period and the function k which defines K is sought. However,
in a statistical theory no statements are made about single occurrence but only about large
numbers of such. Correspondingly, the prediction or extrapolation based on a single time
series u (the determination of k from a single u) does not make any sense. The optimality
criterion itself must be interpreted statistically, and the goodness of the operator must be
measured not by a single sample but by its average effect. Hence the stochastic processes
which appear are classified by their autocorrelation
Z T
1
ϕ(a) := lim u(x − a) u(x) dx, a ∈ R.
T →∞ 2T −T

When passing from u to the function ϕ, a certain part of the information content of the
time series u is isolated, while for the rest the specific features of u are ignored. For a
class of time series with known autocorrelation ϕ, the optimality criterion can be written
as a Wiener-Hopf equation
Z ∞
ϕ(x + a) − ϕ(x − y) k(y) dy = 0, x ≥ 0,
0
where ϕ and a are given and k is sought.
These methods have become standard fare in the statistical time series analysis
through the pioneering works [252, 1943] and [442], and can be found in any of the
textbooks on statistics and probability theory, frequently under the title Spectral theory
of stochastic processes. A survey with an abundance of examples from economics and
technology is provided by [230, 1967], which includes nonstationary processes also. The
details of these methods are not always interesting from our point of view (computation
of the index of Fredholm operators). Conversely, the computation of the index is as a
rule uninteresting for correlation theory, since the Wiener-Hopf operators which show up
usually have vanishing index, see [230, 1967, p.75]; but also note [458, 1970, p.147f]
who warns about the illusion of an easy computability of the optimal kernel functions and
points out the large computational effort necessary for the determination of the correlation
functions... in spite of their uniqueness and explicit solvability in principle. He suggests
adaptive algorithms as an alternative. These are associated with other types of Fredholm
operators, and the uniqueness of the solution is lost. In our context, we want to retain the
probabilistic method which roughly consists in forming averages by means of Lebesgue
integration, and in compressing and selecting information. The relevant information is
then that which (as autocorrelation and the prediction operator itself) yields statements
on the kind of connections and transitions between one curve segment (time series) and
the next (transition probabilities). This is exactly the same strategy that is practical in
algebraic topology which investigates, again roughly how geometric structures are com-
posed of simpler pieces (see Part III). On this background, the explanation takes shape of
4.5. THE DISCRETE INDEX FORMULA. THE CASE OF SYSTEMS 125

why the Wiener-Hopf operators, which originated in boundary value problems of analy-
sis and gained significance in probability theory, more recently turned out to be relevant
for the representation of operations in K-theory (see Section 10.5). It is simply because
they are (as all Fredholm operators) a functional analytic tool in the treatment of seams,
transitions, and relations. J

5. The Discrete Index Formula. The Case of Systems


In Appendix A, we become familiar with the Hilbert space L2 (S 1 ) of mea-
surable, square-integrable functions on the circle, and we cited the fact that the
functions z 7→ z n , n ∈ Z form an orthonormal basis for L2 (S 1 ).
Exercise 4.1. Let H± denote the subspaces of L2 (S 1 ) spanned by z k with
k ≥ 0, respectively with k < 0, and Hn the subspace of L2 (S 1 ) spanned by z k with
k ≥ n ∈ N. Show that the functions z 0 , z 1 , ..., z n−1 form a basis of the orthogonal
complement (Hn )⊥ of Hn in H0 .
Exercise 4.2. Let P denote the orthogonal projection L2 (S 1 ) → H+ and P 0
the orthogonal projection L2 (S 1 ) → H− , i.e., P 0 := I − P . Let f be a continuous
complex-valued function on the circle; i.e., f ∈ C 0 (S 1 ).
(a) Show that Tf := P Mf |H+ defines a bounded linear operator on the Hilbert
space H+ , where Mf denotes multiplication by f .
(b) Verify that for u ∈ H+ and n ∈ Z+
X∞
(Tf u)b(n) = fb(n − k)bu(k),
k=0

where fb(m) := hf, z m i is the m-th Fourier coefficient of f (see Appendix A). Tf
is the (discrete) Wiener-Hopf Operator assigned to f .
Exercise 4.3. Show that f 7→ Tf defines a continuous linear map
T : C 0 (S 1 ) −→ B(H+ ),

where the Banach algebra C 0 (S 1 ) has norm kf k := sup |f (z)| : z ∈ S 1 . [Hint:
kTf k ≤ kf k. Incidentally, is T a Banach algebra homomorphism; i.e., does it
respect the ring structure? See Step 2 in the first (extended) proof of Theorem 4.4
below.]
Theorem 4.4 (Discrete Gohberg-Krein Index Formula, 1956). If f ∈ C 0 (S 1 )
and f (z) 6= 0 for all z ∈ S 1 , then
(a) Tf : H+ → H+ is a Fredholm operator,
(b) index Tf = −W (f, 0). For the definition of winding number W (f, 0), see Sec-
tion 10.1.
Extended Proof. We begin with (a). Step 1: Let B := B(H+ ) denote
the Banach algebra of bounded linear operators on the Hilbert space H+ , and let
K ⊆ B denote the closed ideal of compact operators on H+ with π : B → B/K the
canonical projection onto the quotient algebra; see Chapter 2 also. From Exercise
4.3, it follows that π ◦ T : C 0 (S 1 ) → B/K is linear and continuous.
Step 2: Let C ∨ denote the subalgebra of C 0 (S 1 ) consisting of the continuous
functions representable by finite Fourier series. Let f, g ∈ C ∨ , say
Xn Xm
f (z) = fb(k)z k , g(z) = gb(k)z k ,
k=−n k=−m
126 4. WIENER-HOPF OPERATORS

for n, m ∈ N. In Appendix A, the Fourier coefficients of f g are already calculated:


X∞
fcg(j) = fb(j − k)bg (k), j ∈ Z,
k=−∞

where the sum is actually taken over only finitely many k. Thus we have (see also
Exercise 4.2b)
Tf Tg (z k ) = Tf g (z k ) for k ≥ m + n.
The operators Tf Tg and Tf g coincide on the subspace Hm+n of H+ . Since the
codimension of Hm+n in H+ is finite (= m + n), this means that Tf Tg − Tf g is an
operator of finite rank, and hence is compact. While T is not a homomorphism of
Banach algebras (give a counterexample with f := ... and g := ...), by passing
to the quotient algebra B/K, we have
πT (f g) = πT (f )πT (g).
Thus, π ◦ T is a homomorphism, when restricted to the subalgebra C ∨ .
Step 3: By the Approximation Theorem of Karl Weierstrass (see Chapter A or,
for a direct proof, [134, p.49]), each continuous function on a compact interval can
be uniformly approximated (i.e., in the sup-norm) by polynomials, and even more
so by rational functions. Thus, C ∨ is dense in C 0 (S 1 ). Since π ◦T is continuous, the
multiplicative property carries over; i.e., π ◦ T : C 0 (S 1 ) → B/K is a homomorphism
of Banach algebras.
Step 4: Since πT (1) = 1 (where the 1 on the left is the constant function z 7→ 1 and
the 1 on the right is the class {Id +K : K ∈ K}, it follows that π ◦ T takes invertible
functions into invertible elements of B/K. Hence, if f (z) 6= 0 for all z ∈ S 1 , then
π(Tf ) is invertible in B/K, and so Tf is a Fredholm operator by the Theorem of
Atkinson (Theorem 3.2, p. 64).
We now prove (b). We begin with the simplest case, the function f (z) = z m .
Relative to the canonical orthonormal basis of H+ consisting of the functions z n ,
n ≥ 0, the Wiener-Hopf operator Tzm assigned to f (Exercise 4.2b) has the form
of the one-sided shift operator (shift+ )m for m ≥ 0 and (shift− )|m| for m < 0. By
Exercise 1.3, we then have index Tzm = −m. From the continuity of T (Exercise
4.3) and the continuity (homotopy invariance or local constancy) of the index (see
Theorem 3.11, p. 68), it follows from (a) that index Tg = −m for any g ∈ C 0 (S 1 )
with values in C× = C \ {0} which can be connected to the function z m by a
continuous path of functions in C 0 (S 1 ) with values in C× . Now, the winding
number of the curve S 1 → C (defined by z m ) about the point 0 is m. Since
curves in C× are homotopic through curves in C× exactly when they have the same
winding number (see Section 10.1), we have index Tg = −W (g, 0), and the index
formula is proved. 

We are indebted to R.T. Seeley for the following outline of a much simpler
proof of the preceding theorem and a re-arrangement of Fritz Noether’s Theorem
(for more details see also [381] and our elaboration in Theorem 5.11). We use (with
the notations of Exercise 4.2):
Lemma 4.5. For f ∈ C 0 (S 1 ), the commutator [P, Mf ] is compact.
Proof. If f (t) = eikt then P Mf − Mf P has finite rank |k|. So, if f has a finite
Fourier expansion, the commutator is compact. The Lemma follows by Weierstrass
approximation. 
4.5. THE DISCRETE INDEX FORMULA. THE CASE OF SYSTEMS 127

Short Proof of Theorem 4.4. If |f (t)| > 0, then by Lemma 4.5


P (Mf )−1 P Mf = I + K and P Mf P (Mf )−1 = I + K 0
with K, K 0 compact operators, so P Mf is Fredholm.
To compute the index, if W (f, 0) = k, then, by homotopy invariance, we can assume
f (t) = eikt . Then
k ≥ 0 =⇒ dim Ker(P Mf ) = 0, dim Coker(P Mf ) = −k,
k < 0 =⇒ dim Ker(P Mf ) = −k, dim Coker(P Mf ) = 0. 
Definition 4.6. The (discrete) Hilbert transform on L2 (S 1 ) is H := P −
0
P = 2P − I.
Note that H2 = I, H∗ = H, and [H, Mf ] compact for f continuous.
Theorem 4.7 (F. Noether, 1920). If a and b are continuous: S 1 → C and
|a (t) − b2 (t)| > 0 for all t, then Ma + Mb H is Fredholm on L2 (S 1 ) and
2

(4.1) index(Ma + Mb H) = W (a − b, 0) − W (a + b, 0).


Proof. By Lemma 4.5, Ma +Mb H = P (Ma +Mb )P +P 0 (Ma −Mb )P 0 +K with
K compact. By Theorem 4.4, acting on H+ , P (Ma + Mb ) has index −W (a + b, 0).
A similar argument shows that, on H− , P 0 (Ma −Mb ) has index W (a−b, 0). Modulo
compact operators, Ma + Mb H is the direct sum of these two. 
Exercise 4.8. In the construction of the index bundle (Theorem 3.30), we
have seen that for each prescribed orthonormal basis e0 , e1 , e2 , . . . and Fredholm
operator S ∈ F(H+ ) , there is an n ∈ N such that
Pn S : H+ −→ Hn
is surjective, where Pn is the orthogonal projection of H+ onto the closed subspace
Hn spanned by the basis elements en , en+1 , en+2 , . . .. Now show that (in the case
of Wiener-Hopf operators) to each f ∈ C 0 (S 1 ) with f (S 1 ) ⊆ C \ {0}, one can ex-
plicitly give an n for which Pn Tf : H+ → Hn will be surjective.
[Hint: One naturally exploits the fact that we deal not with an arbitrary Hilbert
space, but rather with function spaces, where there is an additional structure: Ap-
proximate the function z 7→ 1/f (z) by a finite Fourier series
Xn
g(z) = gb(k) z k
k=−n

with n chosen large enough so that sup |f (z) g(z) − 1| : z ∈ S 1 < 1.]
Theorem 4.4 and Exercise 4.8 demand a detailed topological discussion, in
relation to Chapter 1 and in view of Part III. However the families of Wiener-Hopf
operators which we will encounter in the following are not of such elementary type.
So we need some generalizations.

The first generalization is apparent, if we interpret the Wiener-Hopf opera-


tor Tf as a prediction operator for a time series . . . , u−4 , u−3 , u−2 , u−1 , u0 of (say
geophysical) measurements,
P∞ as depicted in Figure 4.4.
Here, e.g., vn = k=0 fn−k u−k , with u−k = u b(k) the given time series, fn−k =
fb(n − k) the weighting and vn = (Tf u)b(n) the predicted time series.
From the standpoint of the statistician, it is now perfectly obvious (even if one
is interested in the weather in Frankfurt exclusively) that the inclusion of additional
128 4. WIENER-HOPF OPERATORS

Figure 4.4. Interpretation of a Wiener-Hopf operator as a pre-


diction operator

series of meteorological measurements (from Iceland or the Azores, say) can result
in more information than the most sophisticated evaluation of a single series of
data (of Frankfurt, for example) could provide. While for a single time series,
the weights fn−k are numbers, they must be matrices in the statistical analysis of
multiple time series. Hence, if we deal with an N -fold time series, the condition
f (z) 6= 0 which implies the Fredholm property must be replaced by det(f (z)) 6= 0,
where f (z) ∈ GL(N, C).
Exercise 4.9. Let H be a Hilbert space of complex-valued functions (e.g.,
H = L2 (S 1 ) or other examples in Appendix A). Show that the well-known notion
of tensor product from multilinear algebra for finite-dimensional vector spaces also
yields a sensible definition H ⊗CN . Convince yourself that H ⊗CN is again a Hilbert
space and (for the concrete examples) is related to the scalar-valued function space
H, in such a way that one can regard H ⊗ CN as being the corresponding function
space with values in CN .
[Hint: Compare the analogous considerations in the proof of Theorem 3.40 with
regard to the Hilbert space Hom(CN , H) isomorphic to H ⊗ CN . How does one
obtain a basis for H ⊗ CN from bases of H and CN ? Details of the algebraic
construction are in [356, 1970, p. 116f], and the peculiarities of infinite-dimensional
spaces (which are indeed no problem, when one factor of the tensor product is finite
dimensional) are found in [128, 1972, p.31 and 79f].]
Exercise 4.10. For a continuous map f : S 1 → GL(N, C), define the Wiener-
Hopf operator
Tf := P Mf |H+ ⊗ CN : H+ ⊗ CN −→ H+ ⊗ CN ,
where P : H ⊗ CN → H+ ⊗ CN is the projection, and Mf is multiplication by the
matrix function f . Show:
a) Tf is a Fredholm operator,
b) index Tf depends only on the homotopy class of f in the homotopy set
[S 1 , GL(N, C)].
[Hint: Repeat the arguments from Exercises 4.2a and 4.3, and Theorem 4.4. Be-
cause of (b), we can identify index Tf ∈ Z with the element [f ] in the fundamen-
tal group π1 (GL(N, C)) ∼= Z that f represents. Each continuous map of S 1 into
GL(N, C) is homotopic to a continuous map of S 1 into the space of invertible di-
agonal matrices of rank N . Therefore, set
[f ] := −W (det f, 0),
4.6. THE CONTINUOUS ANALOGUE 129

where det f (z) is the determinant of the matrix f (z). See also under Section 10.2.]
Exercise 4.11. In the next generalization, let X be a compact parameter
space. Assign to each continuous map f : S 1 × X → GL(N, C) a Fredholm family
Tf : X → F and also an index bundle index Tf ∈ K(X). Show that index Tf only
depends on the homotopy class of f .
[Hint: Note that f (z, x) is an invertible matrix that depends continuously on the
variables z and x. Apply Exercise 4.10, noting that we obtain Tf(·,x) ∈ F, for each
x ∈ X. Here F is the space of Fredholm operators on the Hilbert space H ⊗ CN .
Show that Tf(·,x) depends continuously on x, and then apply the construction from
Theorem 3.30, p.84.]
Exercise 4.12. For a further generalization let E be a complex vector bundle
over X of fiber dimension N . Figuratively speaking, one allows the vector space
CN to change from point to point. Given a function f (z, x) ∈ Iso(Ex , Ex ) which
depends continuously on z and x and therefore defines a family of automorphisms
of the vector bundle E, construct a family of Fredholm operators (in the variable
Hilbert space H ⊗ E ), and finally an index bundle index Tf ∈ K(X) that again
only depends on the homotopy class of f .
[Hint: See Theorem 3.30, Remark 3.31(p. 85), where we may take the base X to be
sufficiently nice (e.g., triangulable). Question: Do we really need the Theorem of
Kuiper in this Exercise (as in Remark 3.31) or can we proceed directly because of
the particular structure of the problem? See [20, p.115].]

6. The Continuous Analogue


In connection with local elliptic boundary-value problems (Section 9.4) and
topological investigations of the general linear group GL(N, C) (Chapter 10, the
Periodicity Theorem of Raoul Bott), we will return to the preceding construction.
For the moment, we will only consider the continuous analog of Theorem 4.4:
Exercise 4.13. Let L1 (R) denote the space of measurable, absolutely inte-
grable functions. Show that each ϕ ∈ L1 (R) defines a bounded linear operator
Kϕ : L2 (R+ ) → L2 (R+ ), via
Z ∞
(Kϕ u)(x) := ϕ(x − y)u(y) dy, x ∈ R+ .
0
[Hint: Regard L2 (R+ ) as a subspace of L2 (R), and then apply the results of Chap-
ter A on the convolution. For detailed estimates, see [415, 1937, p.90f].]
Theorem 4.14. Let ϕ ∈ L1 (R) with ϕ(t)
b + 1 6= 0 for all t ∈ R and let Kϕ be
as in Exercise 4.13. Then
Id +Kϕ : L2 (R) −→ L2 (R)
is a Fredholm operator, and we have
index(Id +Kϕ ) = W (ϕ
b + 1, 0),
where W (ϕb + 1, 0) is the winding number of the oriented curve t 7→ ϕ(t)
b + 1 (t ∈ R)
about the origin (see Section 10.1).
Remark 4.15. In this index formula, one always must be aware of the depen-
dence of the orientation in the definition of the winding number (for us, W (z, 0) = 1,
for z(t) = ei2πt , t ∈ [0, 1]) and the orientation in the Fourier transformation (for us,
130 4. WIENER-HOPF OPERATORS

R∞
ϕ(x)
b = −∞ e−ixy ϕ(y) dy). If one removes the minus sign in the exponent (e.g., as
does Mark Krein), then one obtains a minus sign in the index formula.
Remark 4.16. More exactly, for any ϕ ∈ L1 (R):
(i) Spece (Kϕ ) = {ϕ(t)
b : t ∈ R}
(ii) index(z Id −K ) = W (ϕ,
ϕ b z) for z ∈ Spece (Kϕ )
surjective for index z Id −Kϕ ≥ 0
(iii) z Id −Kϕ is
injective for index z Id −Kϕ ≤ 0.
Proofs for these results discovered by Mark Krein are found in [234, 1970/1982,
13.4], for example.
Remark 4.17. If we regard Id +Kϕ as a map of L1 (R), then under the as-
sumptions of Theorem 4.14, we have that Id +Kϕ is an isomorphism [443]. We
then have no index problem.
Proof. Instead of presenting a complete proof, we will comment on the very
different ways one can prove Theorem 4.14.
Approach 1: Reduce to Theorem 4.4 with the Cayley transformation κ(z) := z−iz+i ,
which maps the upper half-plane conformally onto the open unit disk, as depicted
in Figure 4.5.

·
C+
i

·(0)
0 1 ·(i)

·(1)

Figure 4.5. The Cayley transform κ : C+ → D2

For v ∈ L2 (S 1 ),
√ v(κ(x))
(U v)(x) :=2 , x∈R
x+i
defines an isometry from L2 (S 1 ) to L2 (R), which carries the Hilbert space
H+ (S 1 ) := v ∈ L2 (S 1 ) : vb(n) = 0 for n < 0


to the Hilbert space


u ∈ L2 (R) : u

(4.2) H+ (R) := b|(−∞,0) = 0 .
Equivalently, H+ (R) consists of the square-integrable functions on R which can be
analytically continued to the lower half-plane C− ; e.g., see [118, 1967, p.82-84].
Instead of working with the projection P : L2 (S 1 ) → H+ (S 1 ) (see Exercise
4.2), we utilize the corresponding projection Q : L2 (R) → H+ (R), where Q =
4.6. THE CONTINUOUS ANALOGUE 131

U P U −1 . To each continuous C-valued function f ∈ C 0 (R) of the form f = c +


b where c ∈ C and ϕ ∈ L1 (R), we assign a (continuous) Wiener-Hopf operator
ϕ,
Wf := Q(Mf ) |H+ (R) : H0 (R) → H0 (R), where Mf means multiplication by f . (In
contrast to such continuous Wiener-Hopf operators, one often refers to the discrete
Wiener-Hopf operators as Toeplitz operators). We now have defined three different
operators:
- the discrete Wiener-Hopf operator Tg , g ∈ C 0 (S 1 ),
- the convolution operator Kϕ , ϕ ∈ L1 (R), and
- the continuous Wiener-Hopf operator Wf for f = c + ϕ.
From the properties of U , it then follows [118, p.91] that
(i) Tg = U −1 Wf U , if f = g ◦ κ, and
cf (u) = fb ∗ u
(ii) W b, for all u ∈ H+ (R); i.e.,
R∞
W
cf (u) = cbu(x) + 0 ϕ(x − y)b u(y) dy, x ∈ R+ , c ∈ C and ϕ ∈ L1 (R) .
Denoting the Fourier transform by F : L2 (R) → L2 (R), we can write (ii) as
0
(ii ) F Wf = (c Id +Kϕ )F , where f = c + F ϕ.
Fact (i) expresses the unitary equivalence of discrete and continuous Wiener-Hopf
operators, which is trivial by the definition of Wf here. Fact (ii) requires some
caution with the Fourier transformation: In Appendix A, we deal only with the
harmonic analysis of periodic processes or of processes which (in some sense) abate
with increasing or decreasing time. However, in the kinematic and statistical analy-
sis of most natural, physical, technical, economic, etc. processes, the classical ma-
chinery is, in fact, not sufficient, since these processes oscillate about some mean
without being strictly periodic. The formally analogous Fourier analysis requires
functions on R which are identically 0 away from a point but are so strongly infi-
nite at this one point that the integral over all of R does not vanish. Physicists,
such as Paul Dirac, used this idea in their computations long before Norbert
Wiener rigorously proved the necessary generalizations of harmonic analysis [443],
and which Laurent Schwartz later placed on an even broader foundation with
his theory of distributions. In the sense of distributions the Fourier transformation
of the constant function 1 is just the Dirac distribution δ at the point 0. See [217,
p.21f] or [370, II, p.11]. Theorem 4.14 follows immediately from Theorem 4.4 with
(i) and (ii0 ).
Approach 2: When Allen Devinatz proved the unitary equivalence be-
tween discrete and continuous Wiener-Hopf operators, he showed more than is
actually necessary for the proof of Theorem 4.14. Alternatively, one can reduce
Theorem 4.14 to Theorem 4.4 in a pedestrian fashion via approximating f − 1 by
functions which are identically zero outside a bounded interval. If g is such a func-
tion, then g + 1 can be considered a continuous periodic function, i.e., an element of
C 0 (S 1 ). The approximation is done in such a way that g(z) + 1 6= 0 for all z ∈ S 1
and then Theorem 4.4 applies to g + 1.
We now proceed as in the passage from Fourier series to Fourier integrals (see
Appendix A) whereby the convergence questions must be considered very carefully.
For an indication of the computations involved, see e.g. [161, p.129-132].
Approach 3: One can avoid Theorem 4.4 and basically give a new proof
(e.g. [260, 1958, Theorem 9.2], [234, 1970/1982, 13.4], also for systems (i.e., ma-
trix valued f , [178, Theorem 4.1]. These proofs employ the famous factorization
132 4. WIENER-HOPF OPERATORS

method introduced by Eberhard Hopf and Norbert Wiener in their original


paper [444]. Its idea and essential content are presented very comprehensibly in
[442, p.153-157] (Norman Levinson’s heuristic addendum), [356, 1970, p.47-48]
(Friedrich Sommer’s survey paper on complex analysis) or [134, 1972, 176-184].
The last source contains an altogether good introduction to the function theoretic
properties of Hardy functions and the elements of our spaces H+ (S 1 ) and H+ (R).
Approach 4: One can use projection methods of a more general sort which
encompass the discrete as well as the continuous case. A detailed exposition is found
in [175, 1974]. Here, as in Approach 3, the stress is on finding explicit solutions
so that statements about the index enter more frequently in the opposite direction,
since “the applicability of one or another projection method to the Wiener-Hopf
integral equation is determined by the index” (l.c., p. 9l). Complete proofs of our
Theorem 4.14, in the manner of these projection methods, are in [346, 2.1.5, 2.4.1,
3.2]. Here Theorem 4.14 is not only proved for square integrable functions, but
at once for broad varieties of more general function spaces, as do the authors of
Approach 3. 
Finally
 we return once more to the discrete Wiener-Hopf operators whose to-
tality Tf : f ∈ C 0 (S 1 ) we will denote by T after adjoining the compact operators
on H+ .
Exercise 4.18. Show that the following is an exact sequence of Banach spaces
0 −→ K(H+ ) −→ T −→ C 0 (S 1 ) −→ 0,
where K(H+ ) denotes the space of compact operators on the Hilbert space H+ (S 1 ).
How are the arrows defined?
[Hint. One best begins with  the fact, established in the
 proof of Theorem 4.4,
that the commutator ideal Tϕ Tψ − Tϕψ : ϕ, ψ ∈ C 0 S 1 is contained in K(H+ ).
Show then
 that the quotient algebra T /K(H+ ) is mapped isometrically via T onto
C 0 S 1 , where the maps of the short sequence are defined using algebraic general-
ities. The proof is not entirely simple. One may consult [128, 1972, p.184]. Note
the similarity to the exact symbol sequence in the theory of partial differential
equations (see Part II below). Also compare it to the tensorial sequence in the case
of systems [128, 1972, p.202f].]
With these classes of Wiener-Hopf operators, we have greatly enlarged our
reservoir of examples. One can even show that, up to unitary equivalence modulo
unitary operators (see Section 3.10, p. 108), every essentially normal operator R on
a separable Hilbert space can be written in the form of a Wiener-Hopf operator.
More precisely:
1. If the essential spectrum of R has the form of a simple, closed curve (say
the image of the circle S 1 under an orientation-preserving, continuous, embedding
η : S 1 ,→ C) and index(R − z Id) = n for z interior to curve, then R is unitarily
equivalent to a compact perturbation of the multiplication operator Mη or the
discrete Wiener-Hopf operator Tη◦κ−n , where κ(z) := z [146, 1973, p.73].
2. Even when the essential spectrum cannot be parametrized so nicely, classes
of generalized Wiener-Hopf operators (namely on generalized Hardy spaces, where
the domain of holomorphy need not be the upper half-plane or the open unit disk,
but may be any bounded region of C) exist from which a model of R can be patched
together; see [146, 1973, p.122] and the original papers quoted there.
Part II

Analysis on Manifolds

But, we ask, will not the growth of


mathematical knowledge eventually
make it impossible for a single re-
searcher to embrace all parts of this
knowledge? In answer let me point
out how thoroughly, by the very na-
ture of the mathematical sciences,
any true progress brings with it the
discovery of more incisive tools and
simpler methods which at the same
time facilitate the understanding
of earlier theories and eliminate
older more awkward developments.
By acquiring these sharper tools
and simpler methods the individual
researcher succeeds more easily in
orienting himself in the different
branches of mathematics. In no
other science is this possible to the
same degree.

David Hilbert, 1900

133
CHAPTER 5

Partial Differential Equations in Euclidean Space

Synopsis. Review of Classical Linear Partial Differential Equations: Constant and


Variable Coefficients, Wave Equation, Heat Equation, Laplace Equation, Characteristic
Polynomial. Elliptic Differential Equations: Where Do Elliptic Differential Operators
Arise? Boundary-Value Conditions. Main Problems of Analysis and the Index Prob-
lem. Calculations. Elementary Examples. The Noether(-Hellwig-Vekua) Problem with
Nonvanishing Index.
I Elliptic operators on sections of complex vector bundles over manifolds provide a
primary source of Fredholm operators. In this Part, we explain how this happens. Before
that, we recall a bit of general knowledge about elementary geometric aspects of partial
differential equations in the plane or in n-dimensional Euclidean space. J

1. Linear Partial Differential Equations

I The theory of partial differential equations serves the characterization of motions


and equilibria with infinitesimal interactions and constitutes the mathematics of all quan-
tities varying in space and time (Norbert Wiener) which makes up a good part of
mathematical physics and of applied mathematics altogether.
We distinguish ordinary and partial differential equations. In ordinary differential
equations (ode), the unknown is a function or a system of functions which depend on a
single independent variable. In most applications, this variable is time. In partial differ-
ential equations (pde), the one or more unknown functions depend on several variables.
In applications, these variables are usually the coordinates of a point in space, but one of
them may be time. A differential equation expresses relations between measurable quanti-
ties and their changes in space and/or time (rates of change). In geometric language (see
Figure 5.1), solving an ordinary differential equation means finding a curve and solving a
partial differential equation means finding a family of curves or a surface or a manifold
of higher dimension, whereby the curvatures of the curves or surfaces must satisfy the
conditions expressed by the differential equation. J

In this Chapter, we restrict ourselves to the treatment of linear differential


equations of the form
Pu = f

where u and f are infinitely differentiable complex-valued functions on R and


X
P u(x) := aα (x)(Dα u)(x), x ∈ R.
α

134
5.1. LINEAR PARTIAL DIFFERENTIAL EQUATIONS 135

Figure 5.1. Finding a curve (ode task, left) and finding a family
of curves etc. (pde task, right)

n times
Here, α = (α1 , ..., αn ) ∈ Z+ × · · · × Z+ is a multi-index to specify the partial
derivative; e.g.
 2 2
(1,0,...,0) 1 ∂ (2,0,...,0) 1 ∂
D := , D := , and
i ∂x1 i ∂x21
 |a|
1 ∂ |α|
Dα := , where |α| := α1 + · · · + αn .
i ∂x1 · · · ∂xα
α1
n
n

Remark 5.1. It’s also convenient to carry the factor of i−|α| when integrat-
ing Hermitian inner products by parts. Then integration by parts can be done
symmetrically. For example, when n = 1,
Z b Z b
d d

dx f g = − f dx g + boundary terms, while
a a
Z b Z b
1 d
(Df ) g = f Dg + boundary terms for D = i dx .
a a

In this way we achieve that the differential operators Dα are formally self-adjoint
(see Exercise 6.42, p. 185) and yield better expressions under Fourier transformation
(note that under Fourier transform the operator 1i dx d
is converted into a simple
multiplication operator, see Exercise A.2, p. 708). A drawback of the factor is that
we have to re-define the Dα and the principal symbol (see (6.26), (6.30) below on
p.186ff) for real differential operators of odd order to stay in the real category.
But here we follow the notation of main stream analysis, which seems unaware of
this drawback, possibly out of a former neglect of operators of first order in their
community. Conversely, we emphasize that real differential operators of first order
(typically of Dirac type) are of interest in index theory, see our Section 13.9 and
the literature given there.

The coefficients aα are always taken to be infinitely differentiable; moreover,


aα = 0 for all but finitely many α. P is then called a differential operator of
order max {|α| : aα 6= 0}. We give the space C ∞ (Rn ) of infinitely differentiable
(complex-valued) functions on Rn the topology defined by the following family of
136 5. PARTIAL DIFFERENTIAL EQUATIONS IN EUCLIDEAN SPACE

semi-norms (for k ∈ Z+ , K ⊂ Rn compact):


X
kf kk,K := sup {|Dα f (x)| : x ∈ K} .
|α|≤k

Accordingly, a sequence f1 , f2 , ... of C ∞ functions converges to the constant function


0, if and only if the functions f and all their derivatives converge to 0 uniformly on
each compact subset of Rn .
Exercise 5.2. Show that a linear (for simplicity, assume scalar) differential
operator P is a continuous, linear, and local map P : C ∞ (Rn ) → C ∞ (Rn ); here
local means that
supp P f ⊆ supp f for all f ∈ C ∞ (Rn ), where
supp f := f −1 (R \ {0}) = the closure of {x ∈ Rn : f (x) 6= 0} .
Remark 5.3. Conversely, one can show that every continuous, linear, local
map P : C ∞ (Rn ) → C ∞ (Rn ) is a differential operator (if one allows the order to
be infinite and only finite on compact subsets). In fact, for each x ∈ Rn , the
map f 7→ (P f )(x) is a continuous linear form on C ∞ (Rn ) with one-point support
{x}, whence [370, I, Ch. III, Theorem XXXV] it is a finite linear combination of
derivatives (in the distributional sense) of the Dirac δ at x. As x varies in Rn , one
can piece together these distributions to obtain the desired differential operator
with C ∞ coefficients. The details are in [109, 1-03f.]. In 1960, Jaak Peetre
showed that one can drop the continuity assumption. One can find the completely
elementary proof, avoiding distribution theory, in [314, p.172-175].
Exercise 5.4 (Commutator Relation). Show that the linear differential oper-
ators with coefficients in C ∞ (Rn ) form a noncommutative algebra. Verify that the
commutator P Q − QP is a differential operator of order at most m + m0 − 1, if P
has order m and Q has order m0 .
Exercise 5.5. Show that the space of linear differential operators with constant
coefficients forms a commutative subalgebra which is isomorphic to the polynomial
algebra C[ξ1 , ..., ξn ] in the variables ξ1 , ..., ξn .
Exercise 5.6. Study the connection between the following partial differential
equations appearing most frequently in mathematical physics texts:
a) The wave equation
 
∂2u 2 ∂2u ∂2u ∂2u
∂t2 − c ∂x2 + ∂x
1
2 + ∂x 2
2
= f (x1 , x2 , x3 , t)
3

is the differential equation for the spreading of vibrations in a homogeneous medium,


where the right side vanishes if no force intervenes, and u denotes the displacement
(e.g., of a vibrating membrane). The constant c is identified with the propagation
speed of the wave.
b) The heat conduction equation or shortly heat equation in a homogeneous
isotropic body is
 2 
∂u ∂ u ∂2u ∂2u
∂t − α ∂x 2 + ∂x 2 + ∂x2 = f (x1 , x2 , x3 , t).
1 2 3

The coefficient α = k/(cp ρ) > 0 is called the thermal diffusivity. It is a material-


specific quantity depending on the thermal conductivity k, the mass density ρ, and
the specific heat capacity cp . Here, the right side vanishes when no sources or sinks
5.2. ELLIPTIC DIFFERENTIAL EQUATIONS 137

are present; u denotes the temperature. The heat equation describes the transfer
of heat energy by molecular and electron collisions within a substance (especially
a solid) due to a temperature gradient. It governs many other diffusion processes
and is to distinguish from the radically different equations describing the two other
heat transfer phenomena, convection and radiation.
c) The potential (or Poisson) equation for the potential of a static electric field
(for example) is
∂2u 2 2

∂x2
+ ∂∂xu2 + ∂∂xu2 = −4πf (x1 , x2 , x3 ),
1 2 3

where f is the given charge density and u denotes the potential whose negative
gradient is the electric field.

One can easily classify the (scalar) second order linear differential equations in
several independent variables. For the corresponding differential operator
X
P = aα Dα
|α|≤2

and a point x ∈ R, consider the characteristic form


X
aα (x) ξ1α1 · · · ξnαn
|α|=2
Pn
which is a quadratic form in ξ1 , ..., ξn since |a| = j=1 αj = 2. In analogy with the
classification of conic sections in affine geometry, P is called elliptic at the point
x, if the form is definite in the sense that
X
aα (x) ξ1α1 · · · ξnαn 6= 0 for (ξ1 , ..., ξn ) ∈ Rn \ {0} .
|α|=2

In this case, by a change of variable (ξi ) → (ηi ) (not necessarily orthogonal), one
can express the form (at x) as
± η12 + · · · + ηn2


We call P hyperbolic at x, if the characteristic form can be expressed as


η12 + · · · + ηn−1
2
− ηn2
by a change of variables; and P is parabolic at x, if we can express the form as
η12 + · · · + ηn−1
2
.
The wave equation is then hyperbolic (n = 4), the heat equation is parabolic
(n = 4), and the potential equation is elliptic (n = 3).

2. Elliptic Differential Equations


Roughly speaking the elliptic differential equations of second order differ from
the other classical types, in that there is no distinguished coordinate (e.g., time).
More precisely, if P (of order k) is not elliptic at x0 , then in general (i.e., except in
certain degenerate cases which can cause difficulties when using Hamilton-Jacobi
methods) there is a function f ∈ C ∞ (Rn ) with f (x0 ) = 0 and
X  α1  αn
∂f ∂f
aα (x0 ) ∂x (x 0 ) · · · ∂xn (x 0 ) = 0,
|α|=k 1
 
∂f ∂f
where the gradient ∂x 1
, . . . , ∂x n
(x0 ) ∈ R \ {0}; i.e., the directional derivatives
of f at x0 do not all vanish. By the Implicit Function Theorem, the set S :=
138 5. PARTIAL DIFFERENTIAL EQUATIONS IN EUCLIDEAN SPACE

{x : f (x) = 0)} is an (n − 1)-dimensional submanifold of Rn in a neighborhood of


x0 (see Figure 5.2).

Rn R

grad f f
0
x0

Figure 5.2. Distinguished coordinate and characteristic surface


at x0 for nonelliptic differential equation

The manifold S is called characteristic for the differential operator P at the


point x0 . Solutions which are otherwise smooth can have jumps of their second
derivatives only along these characteristic surfaces. (In physics the characteristic
surfaces are possible wave fronts.) Furthermore, one obtains from them certain
curves along which (separation of variables) the partial differential equation reduces
to a simpler differential equation of first order, the so-called transport equation. For
these reasons the study of characteristic surfaces is a central task in the theory of
nonelliptic differential equations.
However, we shall deal with elliptic differential operators which have no (real)
characteristic manifolds. The above may well be the reason why, in the theory of
elliptic differential equations, it is not initial value problems but boundary value
problems and problems on compact curved manifolds involving global questions
which are at the center of interest. Slightly exaggerated: Since elliptic operators
look locally the same in all directions (there are no characteristic manifolds, no
distinguished directions etc.), and since the local solvability presents no problems
(according to [217, Theorem 7.2.1] there are no local singularities which could
cause trouble globally). Also since there are sufficiently many local solutions (e.g.,
the large spaces of harmonic functions for the Laplace operator and holomorphic
functions for the Cauchy-Riemann operator), interesting global problems can be
formulated immediately and at times solved. We will come back to this philosophy
later.
We refer to [33] for the connection between elliptic equations and parabolic
initial value problems, which we will discuss more closely below in Section 17.3 for
twisted Dirac operators, see in particular Proposition 17.31, p.541. The solutions
of the parabolic heat equation, with arbitrary initial values, solve the potential
equation asymptotically. This fact is made the starting point for the heat equation
proof of the Atiyah-Singer Index Formula, explained in detail in [55, 167, 451].
Until now we have considered only single differential equations (i.e., scalar
differential operators). The treatment of simultaneous differential equations (where
the interacting unknowns cannot be decoupled) requires the concept of vectorial
5.3. WHERE DO ELLIPTIC DIFFERENTIAL OPERATORS ARISE? 139

differential operators. These are operators of the form P = α aα (x)Dα , where


P
(for each x ∈ Rn ), aα (x) is a linear map from a complex vector space V to a
complex vector space W . Relative to bases of V and W , one can regard the aα (x)
as matrices with complex entries. The differential equation P u = f , where u and
f are C ∞ vector-valued functions on Rn (u(x) ∈ V ∼ = CN and f (x) ∈ W ∼
= CM ),
can be regarded as a system of M differential equations in N unknown functions.
Then, we have
P : C ∞ (Rn , CN ) → C ∞ (Rn , CM ),
where C ∞ (Rn , CN ) denotes the C ∞ functions from Rn to CN .
Exercise 5.7. To what extent do the previous exercises carry over to vectorial
differential operators? A differential operator P of order k is said to be elliptic, if
(for all x ∈ Rn and (ξ1 , ..., ξn ) ∈ Rn \ {0}) the characteristic polynomial or the
principal part
X
aα (x) ξ1α1 · · · ξnαn
|α|=k
is an isomorphism from V to W ; in particular, dim V = dim W .
In the following paragraphs we will investigate the concept of ellipticity more
fully; in particular we will work out the geometric meaning of the principal part.
Here, we give only a few hints for why one is interested in elliptic differential
operators and what kinds of related questions come to the forefront.

3. Where Do Elliptic Differential Operators Arise?


Linear elliptic differential operators emerge in many different contexts.
(o) Linear elliptic differential equations of order 1 can not arise in Pmore than
two variables: An R-linear mapping Rn → C ∼ = R2 , (ξ1 , . . . , ξn ) 7→ αj ξj with
complex αj can not be injective for n > 2. That may explain why the main stream
of the partial differential equations community was late to show interest in the
study of geometrically defined differential operators of first order and Dirac type
(see our Parts III-IV): “First order?! That’s not a challenge.” Perhaps, they were
right regarding equations, but terribly wrong regarding systems.
(i) Modeling of equilibrium states of oscillating systems. A typical example
from mathematical physics is the Laplace equation of potential theory; see Exercise
5.6c above. More complicated problems require more complicated operators: Op-
erators with variable coefficients which only pointwise resemble Laplace operators
(e.g., when the material is not isotropic); operators of higher order; and opera-
tors on those function spaces, where the individual functions are not the concrete
distributions of a continuous quantity (e.g., temperature), but for instance, the
probability amplitudes (wave functions) describing a discrete quantum mechanical
system consisting of single electrons, atoms or molecules.
(ii) Investigation of classical operators on more complicated geometric surfaces.
In analogy with the Laplace operator ∆ = ∂ 2 /∂x21 + · · · + ∂ 2 /∂x2n (or rather the
positive semi-definite −∆), one can construct an operator on any Riemannian man-
ifold; see Chapter 6. Various properties of such operators depend on the form of
the manifold and serve to classify such surfaces and manifolds to some degree see
Chapter 13 below.
(iii) Probabilistic characterization of diffusion processes. In contrast to dis-
crete decay processes with transition probabilities, e.g. on lattices, we deal here
140 5. PARTIAL DIFFERENTIAL EQUATIONS IN EUCLIDEAN SPACE

with infinitesimal descriptions of flows and other processes, whereby the transition
probabilities are given in the form of vector fields. Depending on the model, the
random growth, the mean exit time (for problems with boundary), the expectation
of some other quantity, etc. appear as solutions of characteristic operators which
are associated with the Markov process via some infinitesimal consideration. Con-
ceptually, imagine a particle which performs a symmetrical random motion on the
lattice points of Zn by moving in equal time intervals one unit to one of the 2n
neighboring lattice points with transition probability 1/2n always, i.e., the transi-
tion probability is equidistributed and history independent. If f is a payoff function
defined on the lattice points, then the expectation of the payoff after one time unit
is given by the mean
1 Xn
P f (x) := (f (x + ek ) + f (x − ek )) ,
2n k=1
where the random motion placed the particle one unit ago at the point x ∈ Zn , and
e1 , ..., en are the canonical basis vectors of Rn . The linear operator P − Id is then
a discrete analogue of the operator 21 ∆ in that one can show that the statistical
operator P −Id yields half the Laplace operator, when the distances between lattice
points approach zero. The reason is the identity
Xn 1
(∆f )(x) = lim (f (x + hek ) − 2f (x) + f (x − hek )) ,
k=1 h→0 h2

which holds for sufficiently smooth functions. In this fashion, the Laplace operator
is linked with the Wiener process which models the random motion of very small
particles suspended in some fluid. The Wiener process is characterized probabilis-
tically by the fact that the random change x(t + s) − x(t) of a trajectory x possesses
a normal distribution, i.e., a particularly simple density function. Other probabil-
ity distributions yield different characteristic operators, but again elliptic ones if
the underlying random process is a diffusion process. A very elementary and clear
exposition can be found in [135]. Further details are in [241].
(iv) Branching of solutions of nonlinear differential equations. It should be
noted that physical, biological, or social systems rarely contain intrinsic justifica-
tions for the linearity assumption of mathematical models. The supposition that
the effect on a system under study is exactly proportional to the effect contradicts
the presence of friction and, more generally, the laws of thermodynamics. Lin-
ear models are therefore used exclusively for pragmatic reasons, “either in order
to facilitate computation or on account of the present imperfection of engineer-
ing techniques of realization” (of models) [442, p.12]. There are a multitude of
situations which unquestionably warrant the use of linear models, for example, in
the theory of elasticity of materials, whose deformations are nearly proportional to
the forces acting on them, or for many questions of stability theory and of control
theory, for which the underlying machinery has been made fairly linear by man.
On the other hand, some situations require nonlinear modeling, since the essential
phenomenon of branching of solutions cannot be described in any other way. (Some
examples from mechanics are the bending of a straight rod under a constant force,
the buckling of a flexible plate, the oscillations of a satellite in its orbital plane, and
the surface waves of a heavy fluid.)
These facts in no way render the study of linear models superfluous. Rather it
is true that very many nonlinear systems can be approximated by so-called implicit
operators which are linear, and in many cases also elliptic differential operators.
5.4. BOUNDARY-VALUE CONDITIONS 141

In these cases the index of the implicit linear elliptic differential operator plays an
important role for the derivation of the branching equation. The following example
illustrates why the theory of the branching of solutions of a nonlinear equation,
with an analytic variety as solution manifold, is a natural analogue of the Fredholm
theory with affine spaces as solution manifolds. Consider the nonlinear operator
(x, λ) 7→ T x − λx on H × R where H is a Hilbert space and T a (linear) compact
operator. The solution set {(x, λ) : T x − λx = 0} consists of the R-axis {0}×R and
the kernels Ker {T − λ Id} × {λ} of the operators T − λ Id, which are Fredholm for
λ 6= 0, depicted in Figure 5.3. Here the jumps of the kernel dimension of T − λ Id
(i.e., the eigenvalues of T ) are of special interest. See [422, Chs. VII/VIII, esp.
Sect. 27] and [231] for this rapidly developing theory.

¸i ¸j R

Figure 5.3. Analytic variety {(x, λ) : T x − λx = 0} as solution


manifold of the simple nonlinear operator (x, λ) 7→ T x − λx for
T ∈ K(H)

(v) Problems of optimization theory. Frequently elliptic differential equations


are solved by solving the associated variational problem, i.e., a problem of optimiza-
tion. Conversely, many complicated problems of optimization, particularly those
occurring in control theory, can be reduced to elliptic differential equations and in
this way made clearer and more accessible for particular questions. A comprehen-
sive exposition of this aspect can be found in [310].
(vi) Nonelliptic boundary value problems. Another area of applications is the
treatment of systems of nonelliptic differential equations which sometimes can be
represented as a family of elliptic differential operators in space coordinates para-
metrized by time. This is true for instance for the important type of parabolic
differential equations which describes a multitude of spacial growth and differenti-
ation processes. Here the connection between parabolic initial value problems and
families of elliptic operators is well researched (see above Section 5.2).

4. Boundary-Value Conditions
Notice that in (i), (iii) and (iv) boundary-value conditions play an essential
role, while in (ii) interesting and deep results can be found considering operators
on closed manifolds (see Chapter 6 below), thereby avoiding the analytic difficulties
of boundary-value problems. We will see below how closely connected boundary-
value problems are with problems on closed manifolds. In fact, in the geometric
142 5. PARTIAL DIFFERENTIAL EQUATIONS IN EUCLIDEAN SPACE

expressions of K-theory, every boundary-value problem on a region X with bound-


ary has a corresponding problem on the boundary ∂X of X and a problem on the
double X ∪∂X X of X (see Figure 5.4 and Section 13.8 below).
S
X @X X @X X

Circular diskB2 Circle S 1 Sphere S 2

Circular ring Two circles Torus T 2

Figure 5.4. Correspondence between manifolds with boundary,


boundaries, and closed doubles

Conversely, elliptic operators over a closed manifold reflect in this fashion how
complicated manifolds are built from macromolecules, the classical regions with
boundary of Euclidean space R.
Warning: Conceptually, the term boundary-value problems first brings to mind
the boundary-value problems of the theory of elasticity, where an oscillating mem-
brane is held fast along its border. This is mathematically the Dirichlet problem
u|∂X = 0, or more generally u|∂X = g, where g is a function on ∂X. But in many
applications, we deal with much more general types of boundary-value conditions.
Good examples for all that can occur on the boundary of a region are furnished by
the theory of diffusion processes described in (iii). We list just a few of the simplest
phenomena following [135, p.137-139], see also Figure 5.5:
(I) Backward jump of the particle upon reaching the boundary to a fixed
point x inside X, possibly according to a certain probability distribution π generally
depending on the boundary point y.
(II) Absorption: The particle stays for good at the boundary point first
reached.
(III) Extinction: The particle is annihilated upon first reaching the boundary.
(IV) Reflection: Symmetric reflection of the trajectory in the boundary.

For us, these different boundary-value problems only serve as a supply of con-
ceptual examples, and we will not pursue them further. But we wish to stress that
it is lastly the investigation and classification of the various boundary-value prob-
lems (just like the investigation and classification of various manifolds) that yield
the most interesting results. A simple but meaningful example is the Noether(-
Hellwig-Vekua) Theorem (Theorem 5.11, p. 146).
5.6. NUMERICAL ASPECTS 143

@X
X
y
x
x y
x0

@X

y0

Figure 5.5. Boundary occurrences in the theory of diffusion


processes: extinction (left) and reflection (right)

5. Main Problems of Analysis and the Index Problem


Let X be a region in Rn (or a C ∞ manifold; see below) and a differential
operator on X with C ∞ coefficients. Consider the equation P u = f , where u and
f are functions (not necessarily C ∞ ) on X. Somewhat vaguely, we can (following
[221]) formulate the following questions:
(i) Under which conditions on P and X can one obtain local or global existence
results?
(ii) Given X and P , how are the singularities of u and of f related? Lars
Hörmander shows in detail [loc. cit.] that these questions “are in fact so closely
related that they can be considered different forms of the same problem.” We
are interested in the index of elliptic problems that is in questions of type (i).
We will show in Part II that, for P u = f to have a solution at all, every elliptic
problem with suitable boundary-value conditions must possess an index, i.e., a finite
number of linearly independent solutions of the homogeneous equation (f = 0) and
a finite number of linear conditions for f . In Part III, we will introduce methods for
computing the index from the coefficients of P and from numerical invariants of the
structure of X, and conversely, for representing topological invariants of manifolds
as indices of elliptic operators.

6. Numerical Aspects
“Much of the modern work in partial differential equations looks highly
esoteric, and only a few years ago such work would have been considered
of no interest for applications, where one wants a solution expressed in
a workable form, say by a sufficiently simple formula. The advent of the
modern computing machines has changed this. If a problem involving
a differential equation is sufficiently understood theoretically, then, in
principle at least, a numerical solution can be obtained on a machine.
If the mathematics of the problem is not understood, then the biggest
machine and an unlimited number of machine-hours may fail to yield
a solution.” (COSRIMS Report of the National Science Foundation,
1969).
144 5. PARTIAL DIFFERENTIAL EQUATIONS IN EUCLIDEAN SPACE

I Sometimes the choice of numerical methods can cleverly be based on previous


knowledge of the index. While dealing with Wiener-Hopf operators in Chapter 4, we
pointed out such results of I. Z. Gohberg and I.A. Feldman (after Theorem 4.14, p.129).
Similarly, in the numerical treatment of nonlinear problems, different methods have been
recommended, depending on the index of the associated linear problem ([422], [231]).
V. Strassen and others uncovered the importance of the Theorem of Riemann-Roch (see
Section 13.7) and of other quantitative (index-) formulas of algebraic geometry for basic
questions of computational complexity (e.g., for the calculation of the computational steps
needed for inverting a matrix). Thus, there is an indirect relevance of index calculations
on computer oriented numerical mathematics in this setting as well. J

7. Elementary Examples

After these general remarks we will work out in detail some elementary exam-
ples.
Exercise 5.8. Investigate the (trivially elliptic) ordinary differential operator
on the unit interval I = [0, 1] with boundary ∂I = {0, 1}, defined by
P : C ∞ (I) × C ∞ (I) → C ∞ (I) × C ∞ (I),
(f, g) 7→ (f 0 , −g 0 ) ,
with three choices of boundary conditions C ∞ (I) × C ∞ (I) → C ∞ (∂I) ∼
=C×C
(i) R1 : (f, g) 7→ (f − g) |∂I
(ii) R2 : (f, g) 7→ f |∂I
(iii) R3 : (f, g) 7→ (f + g 0 )|∂I .
Determine the index of the operators (for i = 1, ..., 3)
(P, Ri ) : C ∞ (I) × C ∞ (I) → C ∞ (I) × C ∞ (I) × C ∞ (∂I).
[Hint: Clearly, dim Ker(P, Ri ) = 1. To determine the cokernel, one writes F, G ∈
C ∞ (I) and h = (h0 , h1 ) ∈ C × C, obtaining
Z t Z t
f (t) = F (τ )dτ + c1 , g(t) = − G(τ )dτ + c2
0 0

and two more equations for the boundary condition. The dimension of Coker(P, Ri )
is then the number of linearly independent conditions on F , G, and h which must
be imposed in order to eliminate the constants of integration. For each of the
three boundary conditions, check
R1 that there
R 1 is only one condition
R 1 on the triple
(F, G, h), namely h0 = h1 − 0 F (τ )dτ − 0 G(τ )dτ ; h0 = h1 − 0 F (τ )dτ ; resp.,
R1
h0 = h1 − 0 F (τ )dτ − G(0) + G(1). Conclude that the index vanishes in all three
cases.]
For a more comprehensive treatment of the existence and uniqueness of boundary-
value problems for ordinary differential equations (including systems), we refer to
[111] and [200, p.322-403]. Does the index always vanish?
We now consider the Laplace operator ∆ := ∂ 2 /∂x2 + ∂ 2 /∂y 2 , as a linear
elliptic differential operator from C ∞ (X) to C ∞ (X), where X is the unit disk
{z = x + iy : |z| ≤ 1} ⊂ C with boundary ∂X := {z ∈ C : |z| = 1}.
5.7. ELEMENTARY EXAMPLES 145

Exercise 5.9. For the boundary-value problem (named after Peter Gustav
Dirichlet) with boundary condition
R : C ∞ (X) → C ∞ (∂X), with R(u) = u|∂X ,
show that
a) Ker(∆, R) = {0} and
b) Im(∆, R)⊥ = {0} ,
where ⊥ is orthogonal complement in L2 (X) × L2 (∂X).1 In particular, it follows
that index(∆, R) = 0.
[Hint for (a): Ker(∆, R) consists of functions of the form u + iv, where u and v are
real-valued. Since the coefficients of the operators ∆ and R are real, we may assume
v = 0 without loss of generality. Thus, consider a real solution u with ∆u = 0 in
X and u = 0 on ∂X. Then (where ∇u := ( ∂u ∂u
∂x , ∂y ))
Z Z
2
(5.1) 0=− u∆u dxdy = |∇u| dxdy,
X X
whence ∇u = 0, noting that u is real. Thus, is constant, and indeed zero since
u = 0 on ∂X. The trick lies in the equality (5.1), an integration by parts which
is perhaps most simply derived from the integral theorem of George Gabriel
Stokes in the calculus of differential
R formsR (see Exercise 6.20, p. 172 and [356,
p.133f]). Stokes’ formula reads X dω = ∂X ω, where ω is a 1-form. We set
ω := u ∧ ∗du, where ∗ denotes the Hodge star operator (defined here via ∗du =
!
∗(ux dx + uy dy) = ux dy − uy dx, again, see Exercise 6.20) and obtain
2
dω = du ∧ ∗du + u ∧ d ∗ du = |∇u| dx ∧ dy + (u∆u) dx ∧ dy.
Using Stokes’ formula and u|∂X = 0, we have
Z Z Z Z Z
2
|∇u| dxdy + (u∆u) dxdy = dω = ω= u ∧ ∗du = 0.
X X X ∂X ∂X
From this and ∆u = 0, conclude that ∇u = 0 and u is constant.]
[Hint for (b): Choose L ∈ C ∞ (X) and l ∈ C ∞ (∂X) with (L, l) orthogonal to
Im(∆, R), whence (relative to the usual measures on X and ∂X)
Z Z
(5.2) (∆u) L + ul = 0 for all u ∈ C ∞ (X).
X ∂X
Use a 2-fold integration by parts (in the exterior calculus) to obtain
Z Z Z Z
(5.3) u∆L − (∆u) L = (u(d ∗ dL) − d(∗du) L) = (u ∗ dL − L ∗ du) .
X X X ∂X
First consider u with support supp(u) := the closure of {z ∈ X : u(z) 6= 0} con-
tained in the interior of X. Then
Z Z Z
u∆L = (∆u) L = − ul = 0,
X X ∂X

1Here, consider that the intersection of the orthogonal complement of Im(∆, R) relative to
the usual inner product in L2 (X) × L2 (∂X) with the space C ∞ (X) × C ∞ (∂X) is isomorphic to
Coker(∆, R). This is true, since the image of the natural Sobolev extension of (∆, R) is closed in
the L2 -norm, and its L2 -orthogonal complement is contained in C ∞ (X) × C ∞ (∂X).
146 5. PARTIAL DIFFERENTIAL EQUATIONS IN EUCLIDEAN SPACE

and so ∆L = 0. Now for u ∈ C ∞ (X) apply (5.2) and (5.3) to deduce that
Z Z Z
ul = − (∆u) L = (u ∗ dL − L ∗ du)
∂X X ∂X
Z     
= u x ∂L
∂x + y ∂L
∂y − L x ∂u
∂x + y ∂u
∂y .
∂X

Conclude that l = x ∂L
∂x + y ∂L
∂y and L|∂X = 0, and finally apply (a). Details are in
[217, p.264].]
Remark 5.10. The preceding result index(∆, R) = 0 (for Ru = u|∂X ) can
also be obtained by proving the symmetry of ∆ and that the L2 extension on the
domain defined by Ru = 0 is a self-adjoint Fredholm extension.

We now consider a C ∞ vector field ν : ∂X → C on the boundary ∂X =


{z : |z| = 1}. For u ∈ C ∞ (X), z ∈ ∂X, and ν(z) = α(z) + iβ(z), one defines
the directional derivative of the function u relative to the vector field ν at the point
z to be the number
∂u ∂u ∂u
∂ν (z) := α(z) ∂x (z) + β(z) ∂y (z) .
From the standpoint of differential geometry it is better either to denote the vector
∂ ∂
field by ∂ν or to write the directional derivative as simply as ν [u](z), since ∂x and
∂ ∂
∂y can be regarded as vector fields; see Chapter 6 below. The pair (∆, ∂ν ) defines
a linear operator

(∆, ∂ν ) : C ∞ (X) → C ∞ (X) ⊕ C ∞ (∂X) given by u 7→ (∆u, ∂u
∂ν ).

Theorem 5.11 (F. Noether, 1920, and, differently, G. Hellwig, I. N.


Vekua, both 1952). For p ∈ Z and ν(z) := z p as depicted in Figure 5.6, we

have that (∆, ∂ν ) is an operator with finite-dimensional kernel and cokernel, and

(5.4) index(∆, ∂ν ) = 2(1 − p) .

Remark 5.12. The preceding theorem of F. Noether (proved independently


and differently by G. Hellwig and I.N. Vekua) remains true, if we replace z p by
any nonvanishing vector field ν : ∂X → C \ {0} with winding number p as in Figure
5.7.
Moreover, in place of the disk, we can take X to be any simply-connected (i.e.,
without holes) domain in C with a smooth boundary ∂X; see Chapter 6 below. The
reason is the homotopy invariance of the index (see Theorem 3.11, p. 68) which holds
for elliptic differential operators on closed manifolds and on compact manifolds with
smooth boundary when admissible boundary conditions are imposed (see Section
13.8, p.337 below) .
Remark 5.13. One encounters the number 2(1 − p) also in the theory of Rie-
mann surfaces of genus p; e.g., as the Euler characteristic of a closed surface or,
deeper, in the theorem of Bernhard Riemann and Gustav Roch (see Section
13.7, p.333). This is no accident, but rather it is connected with the relation be-
tween elliptic boundary-value problems and elliptic operators on closed manifolds,
as mentioned above on p. 139. Specifically, there is a relation between the index of

(∆, ∂ν ) and the index of the Cauchy-Riemann operator for complex line bundles
over S 2 = P1 (C) with Chern number 1 − p (e.g., see Example 5.18 for a start).
5.7. ELEMENTARY EXAMPLES 147

Figure 5.6. The vector field ν : ∂X → C with winding number


p = 0, 1, 2

@X
º

p =2 X

Figure 5.7. Another vector field ν with winding number 2

Remark 5.14. Motivated by the method of replacing a differential equation


by difference equations, David Hilbert and Richard Courant expected “lin-
ear problems of mathematical physics which are correctly posed to behave like a
system of N linear algebraic equations in N unknowns... If for a correctly posed
problem in linear differential equations the corresponding homogeneous problem
possesses only the trivial solution zero, then a uniquely determined solution of the
general inhomogeneous system exists. However, if the homogeneous problem has
a nontrivial solution, the solvability of the nonhomogeneous system requires the
fulfillment of certain additional conditions.” This is the heuristic principle which
[116, II, p. 179/231] saw in the Fredholm Alternative (see Chapter 2). Hilbert, if
not Courant must have been aware of Noether’s 1920 counterexample in [324]
and its motivation in the hydrodynamics of fluids of weak friction. Then what
about [202] by Günter Hellwig (nicely explained in [197]) in the real setting
and [423, 424] by Ilya Nestorovich Vekua in complex setting? In hindsight,
148 5. PARTIAL DIFFERENTIAL EQUATIONS IN EUCLIDEAN SPACE

we may consider them rather as nice exercises than decisive breaks with the com-
mon perception about vanishing index. That was disproved by Noether long time
before.
We remark that in addition to these oblique-angle boundary-value problems, cou-
pled oscillation equations, as well as restrictions of boundary-value problems, even
with vanishing index, to suitable half-spaces, furnish further elementary examples
for index 6= 0. The simplest example of a system of first order differential operators
on the disk is provided in Exercise 5.18, p. 154 below. A world of more advanced,
and for differential geometry much more meaningful examples, is approached by
the Atiyah–Patodi–Singer Index Theorem, see Section 13.8 below.
Proof of Theorem 5.11. We follow [217, p.266f]. (Below on p.152f, we
shall give the outlines of a much simpler proof, based on Noether’s Index The-
orem regarding the discrete Hilbert transform, Theorem 4.7, p.127.) Since the

coefficients of the differential operators (∆, ∂ν ) are real, we may restrict ourselves

to real functions. Thus, u ∈ C (X) denotes a single real-valued function, rather
than a complex-valued function (i.e., a pair u1 + iu2 of real-valued functions u1 and
u2 ).

Ad Ker(∆, ∂ν ): It is well-known that Ker(∆) consists of real (or imaginary) parts
of holomorphic functions on X (e.g., see [9, p.175f]). Such functions are called
harmonic. Hence, u ∈ Ker(∆), exactly when u = <(f ) where f = u + iv is
holomorphic; i.e., the Cauchy-Riemann equation ∂f ∂ 1 ∂
∂ z̄ = 0 holds, where ∂ z̄ := 2 ( ∂x +

i ∂y ). Explicitly,
     
0 = ∂f∂ z̄ = 1
2 ∂x

+ i ∂
∂y (u + iv) = 1 ∂u
2 ∂x − ∂v
∂y + i ∂u
∂y + ∂v
∂x .

Every holomorphic (= complex differentiable) function f is twice complex differen-


tiable and its derivative is given by
 
∂f 1 ∂ ∂
∂z := 2 ∂x − i ∂y (u + iv)
   
= 21 ∂u
∂x + ∂v
∂y + i − ∂u
∂y + ∂v ∂u ∂u
∂x = ∂x − i ∂y .

In this way we have a holomorphic function φ := f 0 for each u ∈ Ker(∆). Since


  
p ∂u p ∂u p
∂u
∂ν = <(z ) ∂x + Im(z ) ∂y = <
∂u ∂u
∂x − i ∂y z = <(φ(z) z p ) ,

the boundary condition ∂u p p


∂ν = 0 (ν = z ) then means that the real part <(φ(z)z )
p
vanishes for |z| = 1. For p > 0, φ(z)z is holomorphic as well as φ, and hence for
φ(z) := ∂u ∂u
∂x − i ∂y , we have

with ν := z p , p ≥ 0

u ∈ Ker ∆, ∂ν
⇒ <(φ(z)z p ) ∈ Ker(∆, R) where R(·) := (·) |∂X .
Thus, we have associated the oblique-angle boundary-value problem for u with a
Dirichlet boundary-value problem for <(φ(z)z p ), which has only the trivial solution
by Exercise 5.9a. Since φ(z)z p is holomorphic with <(φ(z)z p ) = 0, the partial
derivatives of the imaginary part vanish, and so there is a constant C ∈ R such that
φ(z)z p = iC for all z ∈ X. If p > 0, then we have C = 0 (set z = 0). Henceφ = 0,

and (by the definition of φ) the function u is constant (i.e., dim Ker ∆, ∂ν = 1).


If p = 0, then φ(z) = iC; and so u(x, y) = −Cy + C, whence dim Ker ∆, ∂ν = 2
e
in this case.
5.7. ELEMENTARY EXAMPLES 149

We now come to the case p < 0, which curiously is not immediately reducible
to the case q > 0 where q := −p. One can try to look for a solution by simply
∂ ∂
turning ∂ν around to − ∂ν as illustrated in Figure 5.8.

@ (z )

z z { @ (z )

X X

Figure 5.8. Replacing ν by −ν does not change the winding number

However, this is futile since the winding numbers of ν and −ν about 0 are the
same. Besides, if νp (z) = z p , we do not have ∂ν∂−p = − ∂ν∂ p . In order to reduce
the boundary-value problem with p < 0 to the elementary Dirichlet problem, we
must now go through a more careful argument. Note that φ(z)z p can have a pole
at z = 0, whence <(φ(z)z p ) is not necessarily harmonic. We write the holomorphic
function φ(z) as a finite Taylor series
Xq
φ(z) = aj z j + g(z) z q+1 ,
j=0

where q := −p and g is holomorphic. We define a holomorphic function ψ by


Xq−1
ψ(z) := g(z) z + āj z q−j ,
j=0

with ψ(0) = 0. Then one can write


Xq−1
φ(z)z p = aq + aj z j−q − āj z q−j + ψ(z).

(5.5)
j=0
∂u
The boundary condition ∂ν = 0 implies <(φ(z)z p ) = 0 for |z| = 1. By (5.5),
we have 0 = <(φ(z)z ) = <(ψ(z) + aq ) for |z| = 1 since then z −1 = z̄. Since
p

ψ is holomorphic, we have again arrived at a Dirichlet boundary value problem;


this time for the function <(ψ(z) + aq ). From Exercise 5.9a, it follows again that
ψ(z) + aq is an imaginary constant, whence ψ(z) = ψ(0) = 0 and aq is purely
imaginary. We have
Xq−1
φ(z) = φ(z)z p z q = aq z q + aj z j − āj z 2q−j

j=0

for arbitrary a0 , a1 , ..., aq−1 ∈ C and aq ∈ iR. As a vector space over R, the set
n o
∂u ∂u ∂
∂x − i ∂y : u ∈ Ker ∆, ∂ν

has dimension 2q + 1; here we have restricted ourselves to real u, according to our


convention above. Since u is uniquely determined by φ up to an additive constant,
it follows that for ν = z p and p < 0,


dim Ker ∆, ∂ν = 2q + 2 = 2 − 2p.
150 5. PARTIAL DIFFERENTIAL EQUATIONS IN EUCLIDEAN SPACE



Ad Coker ∆, ∂ν : As Exercise 5.9b shows, the equation ∆u = F has a solution for
each F ∈ C ∞ (X). In view of this we can show
 C ∞ (X) × C ∞ (∂X) C ∞ (∂X)

Coker ∆, ∂ν = ∼
=
∂ ∂

Im ∆, ∂ν ∂ν (Ker ∆)

as follows. We assign to each representative pair (F, h) ∈ C ∞ (X) × C ∞ (∂X) the



class of h − ∂u∂ν ∈ C (∂X), where u is chosen so that ∆u = F . This map is
clearly well defined on the quotient space of pairs, and the inverse map is given

by h 7→ (0, h). Hence, we have found a representation for Coker(∆, ∂ν ) in terms
 ∂u
of the boundary functions ∂ν : u ∈ Ker ∆ , rather than the cumbersome pairs in


Im ∆, ∂ν . (This trick can always be applied for the boundary-value problems
(P, R), when the operator P is surjective.)
We therefore investigate the existence of solutions of the equation ∆u = 0 with

the inhomogeneous boundary condition ∂u ∂ν = h, where h is a given C function on
∂X. According to the trick introduced in the first part of our proof, it is equivalent
to asking for the existence of a holomorphic function φ with the boundary condition
<(φ(z)z p ) = h, |z| = 1, i.e., for a solution of a Dirichlet problem for <(φ(z)z p ).
By Exercise 5.9b, there is a unique harmonic function which restricts to h on the
boundary ∂X; hence, we have a (unique up to an additive imaginary constant)
holomorphic function θ with <θ(z) = h for |z| = 1.
In the case p ≤ 0, the boundary problem for φ is always solvable; namely, set
φ(z) := z −p θ(z). Hence, we have

= 0 for ν(z) = z p and p ≤ 0.

dim Coker ∆, ∂ν
For p > 0, we can construct a solution of the boundary-value problem for φ from
θ, if and only if there is a constant C ∈ R, such that (θ(z) − iC)/z p is holomorphic
(i.e., the holomorphic function θ(z) − iC has a zero of order at least p at z = 0).
Using the Cauchy Integral Formula, these conditions on the derivatives of θ at
z = 0 correspond to conditions on line integrals around ∂X. In this way, we have
2p − 1 linear (real) equations that h must satisfy in order that the boundary-value
problem has a solution. We summarize our results in Table 5.1 (ν(z) = z p ) and
Figure 5.9. 

Table 5.1. Dimensions of kernel and cokernel for varying p

∂ ∂ ∂
  
p dim Ker ∆, ∂ν dim Coker ∆, ∂ν index ∆, ∂ν
>0 1 2p − 1 2 − 2p
≤0 2 − 2p 0 2 − 2p

Warning 1: We already noted in the proof the peculiarity that the case p < 0
cannot simply be played back to the case p > 0. This is reflected here in the
asymmetry of the dimensions of kernel and cokernel and the index. It simply reflects
the fact that there are more rational functions with prescribed poles than there are
polynomials with corresponding zeros. See also Section 13.7, the Riemann-Roch
Theorem.
5.7. ELEMENTARY EXAMPLES 151

Figure 5.9. The dimensions of kernel and cokernel and the index
of the Laplacian with boundary condition given by ν(z) = z p for
varying p

Warning 2: In contrast to the Dirichlet Problem, which we could solve via


integration by parts (i.e., via Stokes’ Theorem), the above proof is function-theoretic
in nature and cannot be used in higher dimensions. This is no loss in our special
case, since the index of the oblique-angle boundary-value problem must vanish
anyhow in higher dimensions for topological reasons; see [217, p.265f] or Section
13.8 below. The actual mathematical challenge of the function-theoretic proof arises
less from the restriction dim X = 2 than from a certain arbitrariness, namely the
tricks and devices of the definitions of the auxiliary functions φ, ψ, θ, by means of
which the oblique-angle problem is reduced to the Dirichlet problem. Is there not
a canonical, straightforward general method for finding the index of a boundary
value problem? We will return to this question below (Section 13.8 ).
Warning 3: The theory of ordinary differential equations easily conveys the
impression that partial differential equations also possess a general solution in the
form of a functional relation between the unknown function (quantity) u, the in-
dependent variables x and some arbitrary constants or functions, and that every
particular solution is obtained by substituting certain constants or functions f, h,
etc. for the arbitrary constants and functions. (Corresponding to the higher degree
of freedom in partial differential equations, we deal not only with constants of in-
tegration but with arbitrary functions.) The preceding calculations, regarding the
boundary value problem of the Laplace operator, clearly indicate how limited this
notion is which was conceived in the 18th century on the basis of geometric intu-
ition and physical considerations. The classical recipe of first searching for general
152 5. PARTIAL DIFFERENTIAL EQUATIONS IN EUCLIDEAN SPACE

solutions and only at the end determining the arbitrary constants and functions
fails. For example, the specific form of boundary conditions must enter the analysis
to begin with.
We are indebted to R.T. Seeley for the following outline of a much simpler
proof of the Noether-Hellwig-Vekua Theorem 5.11. The arguments are based
on Noether’s Index Theorem for the (discrete) Hilbert transform and applied to
the elliptic boundary problem on the disk via boundary reduction in polar coordi-
nates. The arguments are close to Noether’s original paper [324] (see also [381]).
Moreover, it is neat to make the arguments this way - it gives a topological rationale
for 1 − p, rather than p or −p in the index formula (5.4).
We shall prove the following re-formulation of Theorem 5.11. As before (see,
e.g., p. 35) for greater precision we take the liberty to employ the terminology of
Sobolev spaces to be introduced rigorously below in Chapter 7.
Theorem 5.15 (Re-formulation of Theorem 5.11). Use (r, t) as polar coordi-
nates in the unit disk Ω. Consider the map
a 7→ T u := ∆u, (aur + but )(1, ·) , u ∈ W 2 (Ω), a, b real and of class C 1


with |a2 + b2 | > 0. Then T is Fredholm as a map into L2 (Ω) ⊕ W 1/2 (S 1 ) with
index T = −2W (a + ib, 0),
where W (a + ib, 0) denotes the winding number of the curve a + ib : S 1 → C× .
We use
Lemma 5.16. For any v ∈ W 2 (Ω) with ∆v = 0 we have on S 1 , vt = iHvr
and vr = −iHvt , where H denotes the (discrete) Hilbert transform of Definition
4.6 (p.127).
P∞
Proof. The results are immediate from the Ansatz v(r, t) = −∞ an r|n| eint .

Proof of Theorem 5.15, after the original [324], re-arranged. We
reduce the equation
(5.6) (∆u, aur + but ) = (f, g)
to an equation on the boundary. For f ∈ L2 (Ω), let Gf be the unique solution of
the Dirichlet problem
∆(Gf ) = f, Gf (1, t) = 0.
Then (5.6) ⇐⇒
(5.7) u = Gf + v, ∆v = 0, a(Gf )r + avr + 0 + bvt = g when r = 1.
By Lemma 5.16, the boundary equation is equivalent to
(5.8) (−iMa H + Mb ) vt = g − a(Gf )r ,
where Ma , Mb denote the multiplication operators.
Consider first the case2 that −ia + b = eikt , and set

X X∞
vt = cn eint , g − a(Gf )r = dm eimt .
−∞ −∞

2This case does not require identifying the target space, just to compute the index.
5.7. ELEMENTARY EXAMPLES 153

Then (5.8) reduces to



X −1−k
X ∞
X
(5.9) −i cm−k eimt + cm+k eimt = dm eimt .
m=k m=−∞ −∞

When k ≥ 0, this requires dm = 0 for −k − 1 < m ≤ k; we need dk = 0 to get a


solution for vt with mean value c0 = 0. Such a solution exists if and only if 2k + 1
conditions are satisfied, hence
k ≥ 0 =⇒ dim Coker(T ) = 2k + 1.
As for Ker(T ), when the right hand side of (5.9) is 0, then all cm = 0, so vt = 0, and
v is constant on the boundary, hence everywhere, and then so is u = Gf +v = 0+c.
Thus:
k≥0 =⇒ index(T ) = dim Ker(T ) − dim Coker(T ) = 1 − (2k + 1) = −2k.

When k < 0, the (5.9) always has a solution with mean value 0, so dim Coker(T ) =
0. And it has a nullspace of dimension −2k, so
k<0 =⇒ index(T ) = dim Ker(T ) − dim Coker(T ) = −2k − 0 = −2k.

For arbitrary a and b in C 1 with |a2 + b2 | > 0 for all t, we will show that T is
Fredholm.3 Then the index formula follows by homotopy. For u ∈ Ker(T ), u = v
in (5.7). By Noether’s Theorem 4.7 (p.127), −iMa H + Mb is Fredholm, so vt lies
in a space of finite dimension, hence v(1, ·) lies in a space of finite dimension. Also,
u = v is determined by the boundary values of v. So dim Ker(T ) < ∞.
Next, construct a right inverse to T , modulo compact operators. Define a
(Poisson type) operator D from W 1/2 (S 1 ) to W 2 (Ω) by
∆Dg = 0, Dg(1, t) = g(1, t)
1/2 1
and an operator C on W (S ) such that
(−iMa H + Mb ) = I + K, K compact.
Then define a right parametrix R by
Z • Z
 
R(f, g) := Gf + D C(g − Ma (Gf )r − C g − Ma (Gf )r dt;
0 S1

the outer integral is indefinite. We find, since (Gf )t = 0,


 
T R(f g) = δGF, Ma (Gf )r + g − Ma (Gf )r + K g − Ma (Gf )r
= (f, g) + 0, K 0 (f, g) ,


where K 0 is compact. Thus Im(T ), as a subspace of a closed space of finite codi-


mension, is itself closed with finite codimension. Since dim Ker(T ) is also finite, T
is Fredholm. 

3To be precise, this part of the proof, alas, requires identifying the target space as L2 (Ω) ⊕
W 1/2 (S 1 ) in order to have an operator with closed range. Also a proof is required that Ma H is
bounded on W 1/2 . All that follows at once from the set-up in Chapters 7 and 9.
154 5. PARTIAL DIFFERENTIAL EQUATIONS IN EUCLIDEAN SPACE


Exercise 5.17. Without using Theorem 5.11, show that Index(∆, ∂ν ) = 0 for

ν := z. This boundary-value problem, where ∂ν is the field normal to the boundary
∂X is named after Carl Neumann. From the topological viewpoint it is equivalent
(modulo constant functions) to the Dirichlet boundary-value problem defined by a
tangent vector field, see Figure 5.10.

Neumann Dirichlet

º(z) º(z)

z z

Figure 5.10. Topological equivalence of Neumann and Dirichlet


boundary condition

Exercise 5.18. Let X := {z = x + iy : |z| < 1} be the unit disk and define an
operator
T : C ∞ (X) × C ∞ (X) → C ∞ (X) ⊕ C ∞ (X) ⊕ C ∞ (∂X) by
T (u, v) := ∂u ∂v

∂ z̄ , ∂z , (u − v) |∂X ,

where C ∞ (X) := C ∞ (X, C) (i.e., we are back to complex-valued functions), ∂z ∂


=
1 ∂ ∂ ∂ 1 ∂ ∂
2 ( ∂x − i ∂y ) denotes the complex differentiation and ∂ z̄ = 2 ( ∂x + i ∂y ) denotes

the Cauchy-Riemann differential operator formally adjoint to ∂z . Prove that
index(T ) = 1.
[Hint: Show first that dim(Ker T ) = 1: Suppose that (u, v) ∈ Ker T . Then ∂u ∂ z̄ = 0
∂v
and ∂z = 0 in which case u is holomorphic and v is conjugate-holomorphic (i.e., v
is holomorphic). In particular, u and v are harmonic. Then since (u − v) |∂X = 0,
we have u = v on X. Why? Now u0 (z) = ∂u ∂v
∂z = ∂z = 0, and so u and v are the same
constant function. Then show Coker(T ) = {0}, or more precisely (Im T )⊥ = {0};
see the footnote to Exercise 5.9. For this, choose arbitrary f, g ∈ C ∞ (X) and
h ∈ C ∞ (∂X) and prove that f , g and h must identically vanish, if
Z Z
∂u ∂v
(u − v)h = 0 for all u, v ∈ C ∞ (X).

(5.10) ∂ z̄ f + ∂z g +
X ∂X
Note that for P, Q ∈ C ∞ (X)
 
∂Q ∂Q
d(P dz + Qdz) = ∂P∂ z̄ dz ∧ dz + ∂z dz ∧ dz = ∂P
∂ z̄ − ∂z dz ∧ dz
   
∂Q ∂Q
= ∂P ∂P
∂ z̄ − ∂z (dx − idy) ∧(dx + idy) = 2i ∂ z̄ − ∂z dx ∧ dy.

Thus you obtain the complex version of Stokes’ Theorem,


Z Z Z  
∂P ∂Q
P dz + Qdz = d(P dz + Qdz) = 2i ∂ z̄ − ∂z .
∂X X X
5.7. ELEMENTARY EXAMPLES 155

From this, you get


Z Z Z Z Z
1
∂u
∂ z̄ f = ∂
∂ z̄ (uf ) − u ∂f
∂ z̄ = uf dz − u ∂f
∂ z̄ and
2i ∂X
ZX ZX ZX X
−1
Z Z
∂v
∂z g =

∂z (vg) − v ∂g
∂z = vgdz̄ − v ∂g
∂z .
X X X 2i ∂X X
Hence,
Z Z  Z
 1
∂u ∂v
u ∂f ∂g

∂ z̄ f + ∂z g =− ∂ z̄ + v ∂z + (uf dz − vgdz̄) .
X X 2i ∂X
Assuming (5.10), you have
Z Z
∂u ∂v

0= ∂ z̄ f + ∂z g + (u − v)h
X ∂X
Z  Z Z
∂f ∂g
 1
=− u ∂ z̄ + v ∂z + (uf dz − vg dz̄) + (u − v)h.
X 2i ∂X ∂X
By considering u and v with compact support inside the open disk, you can de-
duce that ∂f ∂g
∂ z̄ = 0 and ∂z = 0 (i.e., f and g are analytic and conjugate analytic
respectively). Thus, (5.10) implies
Z Z
1
0= (uf dz − vg dz̄) + (u − v)h,
2i ∂X ∂X
for all u, v ∈ C ∞ (X). Choosing v = u, you have
Z
1
0= u(f dz − g dz̄) for all u ⇒ f dz = g dz̄ on ∂X
2i ∂X
⇒ f eiθ ieiθ dθ = −g eiθ ie−iθ dθ ⇒ f eiθ eiθ = −g eiθ e−iθ .
   

However, since f is analytic, the Fourier series of f eiθ eiθ has a nonzero  coefficient
for eimθ only when m > 0, and since g is conjugate analytic, g eiθ e−iθ has a
nonzero coefficient for eimθ only when m < 0. Thus, f = g = 0. Choosing v = −u,
(5.10) then yields
Z Z
0= (u − v)h = 2 uh for all u ∈ C ∞ (X) ⇒ h = 0.
∂X ∂X
Remark 5.19. In engineering one calls a system of separate differential equa-
tions
Pu = f
Qv = g,
which are related by a transfer condition R(u, v) = h, a coupling problem; when the
domains of u and v are different, but have a common boundary (or boundary part)
on which the transfer condition is defined, then we have a transmission problem;
e.g., see [68, p.7f]. Thus, we may think of T as an operator for a problem on the
spherical surface X ∪∂X X (see Exercise 6.50 below) with different behavior on the
upper and lower hemispheres, but with a fixed coupling along the equator.
CHAPTER 6

Differential Operators over Manifolds

Synopsis. Motivation. Differentiable Manifolds — Foundations: Implicit Function


Theorem, Tangent Space, Cotangent Space. Geometry of C∞ Mappings: Embeddings,
Immersions, Submersions, Embedding Theorems. Integration on Manifolds: Hypersur-
faces, Riemannian Manifolds, Geodesics, Orientation. Exterior Differential Forms and
Exterior Differentiation. Covariant Differentiation, Connections and Parallelity: Connec-
tions on Vector Bundles, Parallel Transport, Connections on the Tangent Bundle, Clifford
Modules and Operators of Dirac Type. Differential Operators on Manifolds and Symbols:
Our Data, Symbolic Calculus, Formal Adjoints. Elliptic Differential Operators: Definition
and Standard Examples. Manifolds with Boundary.

Motivation.
I For many decades now, workers in differential geometry and mathematical physics
have been increasingly concerned with differential operators (exterior differentiation, con-
nections, Laplacians, Dirac operators, etc.) associated to underlying Riemannian or space–
time manifolds. Of particular interest is the interplay between the spectral decomposition
of such operators and the geometry/topology of the underlying manifold. This has become
a large, diverse field involving index theory, the distribution of eigenvalues, zero sets of
eigenfunctions, Green functions, heat and wave kernels, families of elliptic operators and
their determinants, canonical sections, etc.. Moreover, Simon Donaldson’s analysis of
moduli of solutions of the nonlinear Yang–Mills equations and Seiberg–Witten theory have
led to profound insights into the classification of four–manifolds, which were not accessible
by techniques that are effective in higher dimensions. We shall touch upon many of these
topics, but we focus on index theory and its applications (in Parts III-IV). In this Chapter
and the following of this Part, we shall present an elementary introduction into the basic
notions, concepts, and tools of global analysis.
This chapter is rather tough going for beginners. A main point is to fix terminology
and remind the readers of basic prerequisites for reading our book. For a true learning of
and a good training with the foundational material on differentiable manifolds, the reader
will need some time and a good elementary textbook. We recommend to consult [199,
Chapter 4].
We begin with the concept of a closed manifold. It allows us to generalize and si-
multaneously drastically simplify the index problem by eliminating boundary conditions.
For example, the homogeneous Laplace equation ∆u = 0 on the disk has infinitely many
linearly independent solutions (e.g., (x + iy)n ), while the corresponding Laplace equation
on the sphere has a one-dimensional solution space consisting of the constant functions.
In this respect the notion of a differentiable manifold, does not make the mathematics
more complicated, but is a genuine first approximation to the difficult boundary value
problems in Euclidean space Rn .1
1
The development of mathematics shows again and again how, in the growth of knowledge,
the conceptual and non-conceptual form a unit, alternating, and fading into one another. A
most striking example is furnished by the famous four-color problem, which characteristically still
presents many puzzles in the plane, even after its computer aided solution, while the corresponding
156
6.1. DIFFERENTIABLE MANIFOLDS — FOUNDATIONS 157

But also from the point of view of immediate applications, the geometric concept of
a manifold played an important role. In fact, space-time problems defined initially and
canonically in Euclidean space frequently do not have unrestricted independent variables,
but these variables are restricted by side conditions to certain submanifolds of Euclidean
space. Examples are the constraints in mechanics; the path equations of electrodynamics
into which enter essentially the shape and surface of the conductor ; or the symmetry
conditions of elementary particle physics which replace the high dimensional Euclidean
state spaces by low dimensional state spaces in the form of manifolds. J

1. Differentiable Manifolds — Foundations


We begin with a compilation of the basic notions and elementary relations of the
concept of a differentiable manifold. As a general reference, we refer to [400, 97, 89]
. As emphasized above, for the foundational material on differentiable manifolds a
student could also consult [199, Chapter 4].
Exercise 6.1. Recall the following two classical theorems of differential calcu-
lus, which form the foundation of the concept of a differentiable manifold.
a. (Inverse Function Theorem). If f = (f1 ,..., fn ) is a C ∞ map from Rn
to Rn whose n × n Jacobian matrix (∂fi /∂xj )(p) has rank n at p ∈ Rn (i.e.,
its determinant is nonzero), then there is a neighborhood U of p in Rn which is
diffeomorphic by f to a neighborhood V of f (p) in Rn .
b. (Implicit Function Theorem). Let O ⊂ Rm be open and f = (f1 , ..., fn)
be a C ∞ map from O to Rn (m > n), whose m × n Jacobian matrix (∂fi /∂xj )(p)
at the point p = (p1 , ..., pm ) ∈ U has maximal rank n. Thus, for some permutation
of the coordinates xj , the first n × n submatrix has rank n, i.e., we have

det (∂fi /∂xj )(p) i,j=1,...,n 6= 0.
Then the isolevel set {x ∈ O : f (x) = f (p)} can be parametrized locally. More
precisely (see Figure 6.1), there is a differentiable map (implicit function) g =
(g1 , ..., gn ) defined in a neighborhood V of (pn+1 , ..., pm ) ∈ Rm−n with values in a
neighborhood W of (p1 , ..., pn ), such that W × V ⊂ O and for all x ∈ W × V we
have:
f (x) = f (p) ⇔ xi = gi (xn+1 , ..., xm ) for all i ∈ {1, ..., n} .
A topological manifold without boundary is a locally Euclidean, Hausdorff
topological space X. By locally Euclidean, we mean that for some n ∈ N, each point
of X has a neighborhood U which is homeomorphic via some function u : U → V
to an open subset V of Rn . The function u is called a chart for X. One might
concretely think of geography, where a curved and uneven piece of the earth’s
surface is mapped onto a flat piece of paper. A chart is also known as a local
coordinate system, when one wishes to stress the computational point of view.
A set A of charts, whose domains of definition form an open covering of X, is called
an atlas.
questions for closed manifolds have long been disposed of. “Most of the early attempts at solving
this problem were based on direct attack, and they not only failed, but did not even contribute any
useful mathematics.” Only “a new and highly indirect approach to the coloring problem based on
a generalization of Kirchhoff’s laws of circuit theory in a completely unforeseen direction”, proved
to be “successful in understanding a variety of combinatorial problems.” (Gian-Carlo Rota, The
Mathematical Sciences: A Report, 1969. Reprinted with the permission of the National Academy
of Sciences).
158 6. DIFFERENTIAL OPERATORS OVER MANIFOLDS

Rn

Rm

(p1 ,...,pn ) W p
W V

V
Rm{n
(pn+1 ,...,pm )

Figure 6.1. The Implicit Function Theorem, parametrizing lo-


cally (over V ) the isolevel set {x : f (x) = f (p)}

The number n, the local dimension, is constant on each connected component


of X. In our applications, n will not vary from component to component; so we
may speak of the dimension of the manifold X.
Now let X be an n-dimensional manifold and A an atlas for X. Of geometric
and analytic interest is the study of the coordinate changes u ◦ v −1 , for two charts
u, v ∈ A whose domains have nonvoid intersection. The change u ◦ v −1 is a contin-
uous function from one open subset of Rn to another. This is trivial, by definition.
However, in Rn one has a much richer structure, which permits to impose further
restrictions on the coordinate changes: The atlas A is called a C ∞ -atlas, if all
coordinate changes are C ∞ maps.
Exercise 6.2. Show: Each atlas A for X induces on an open subset W of X
an atlas A|W := {u|W : u ∈ A}. If A is a C ∞ -atlas, then so is A|W .
Exercise 6.3. Show that for each C ∞ -atlas A on an n-dimensional manifold
X, there is a commutative subalgebra (consisting of “C ∞ functions on X”) of the
algebra C 0 (X) of continuous complex-valued functions on X, namely
C ∞ (X) := φ ∈ C 0 (X) : φ ◦ u−1 is C ∞ on Im(u) for all u ∈ A .


Moreover, show that the usual properties hold; e.g.,


a) For φ ∈ C ∞ (X), we have φ|W ∈ C ∞ (W ), where W ⊆ X is open and C ∞ (W )
corresponds to the C ∞ atlas A|W .
b) Suppose φ is a fixed complex-valued function that is locally smooth (i.e., φ|W ∈
C ∞ (W ) for all W ∈ W, where W is an open covering of X). Then φ ∈ C ∞ (X).
c) For φ ∈ C ∞ (X) and ψ ∈ C ∞ (C), we have ψ ◦ φ ∈ C ∞ (X).
d) The constant functions on X are in C ∞ (X).
A C ∞ manifold is a topological manifold X with a “C ∞ -structure” C ∞ (X),
defined by a C ∞ -atlas A.
A C ∞ map from a C ∞ -manifold X to a C ∞ -manifold Y is a function f : X → Y
with φ ◦ f ∈ C ∞ (X) for all φ ∈ C ∞ (Y ). One can express this condition in terms
of local coordinates as follows: For each u in the atlas for X and v in the atlas for
Y , we have v ◦ f ◦ u−1 is C ∞ , as a map from an open subset of Rn to Rm , where
6.1. DIFFERENTIABLE MANIFOLDS — FOUNDATIONS 159

n = dim X and m = dim Y . We denote the set of all C ∞ maps from X to Y by


C ∞ (X, Y ).
The C ∞ manifolds X and Y are called diffeomorphic if there is a diffeo-
morphism from X to Y ; i.e., a bijective C ∞ map f ∈ C ∞ (X, Y ) whose inverse is
in C ∞ (Y, X). If X = Y , one also calls such a map an automorphism.
For technical reasons one frequently requires that a C ∞ manifold be paracom-
pact. This means that every covering of X by open subsets {Uj }j∈J possesses a
locally finite refinement {Vk }k∈K , in the sense that each Vk is an open subset of
some Uj(k) , and for each fixed x ∈ X there is a neighborhood Ox of x, such that the
set {k ∈ K : Ox ∩ Vk 6= ∅} is finite. Note that by defining Wj := ∪j(k)=j Vk (pos-
sibly void), we can obtain a locally finite refinement {Wj }j∈J of {Uj }j∈J without
changing the index set J. Unlike the stronger condition of compactness, we do not
require that K is finite. It is well-known that every metric space is paracompact
[130, p.186], and every paracompact space is normal [130, p.163].
Theorem 6.4. Let {Uj }j∈J be an open covering of a paracompact C ∞ n-
manifold X. Then there is a “C ∞ partition of unity” subordinate to {Uj }j∈J ,
namely, a family {ϕj ∈ C ∞ (X)}j∈J such that the following hold:
(i) ϕj ≥ 0,
(ii) We have supp ϕj := closure of {y ∈ X : ϕj (y) 6= 0} ⊆ Uj , and the family
{supp ϕj : j ∈ J} is locally finite; i.e., for each x ∈ X, there is a neigh-
borhood
P Vx such that {j ∈ J : Vx ∩ supp ϕj 6= ∅} is finite, and
(iii) j∈J ϕj (x) = 1 for all x ∈ X.

Proof. If we can produce a C ∞ partition of unity subordinate to a refinement


of {Uj }j∈J , then it is subordinate to {Uj }j∈J itself. We can produce a refinement of
{Uj }j∈J consisting of open subsets each of which are contained within the domain
of a coordinate chart which maps the open subset to a bounded subset of Rn .
Since X is paracompact, we may then assume (without loss of generality) that the
covering {Uj }j∈J is locally finite and for each j ∈ J, Uj is a subset of the domain
U
ej of a coordinate chart uj : U ej → Rn such that uj (Uj ) has compact closure in Rn .
Step 1. We select an open neighborhood Wx about each point x ∈ X, such that
the closure W x is contained in some Uj . There is a subset Y ⊆ X, such that
{Wy : y ∈ Y } is a locally finite covering of X. By defining

Vj := ∪ Wy : y ∈ Y, W y ⊆ Uj ,
we have a covering {Vj }j∈J of X. We also have V j ⊆ Uj by the local finiteness
of {Wy : y ∈ Y }. Indeed, suppose that x ∈ V j , then every neighborhood Ox of x
intersects some Wy with W y ⊆ Uj . But there is some Ox such that only finitely
many of such Wy , say Wy1 , . . . , Wym , intersect Ox . Since the union of finitely many
closed sets is closed,
x ∈ Wy1 ∪ . . . ∪ Wym ⊆ W y1 ∪ . . . ∪ W ym ⊆ Uj .
Step 2. For each j ∈ J, below we will construct a function ψj ∈ C ∞ (M, R), which
is positive on Vj ⊂ Uj and identically zero on X \ Uj . Then, as {Uj }j∈J is locally
finite, about each point of X there is a neighborhood on which all but a finite
number of the ψj are identically 0, and so ψ := j ψj is C ∞ and positive (since
P

∪j∈J Vj = X). Set ϕj := ψj /ψ and note that conditions (i), (ii) and (iii) hold. With
160 6. DIFFERENTIAL OPERATORS OVER MANIFOLDS

ej → Rn , we carry out the construction of


the help of the coordinate function uj : U
∞ n
ψj as follows. Define η ∈ C (R ) by
(  
−1
exp 1−|x| 2 , for |x| < 1,
η(x) :=
0, for |x| ≥ 1.
Since uj (V j ) is a closed subset of the compact subset uj (Uj ) ⊂ Rn , uj (V j ) is also
compact. Let δ > 0 be the distance from the compact set uj (V j ) to the closed
subset Rn − uj (Uj ). Cover uj (V j ) with finitely many open balls of radius less than
δ and centers a1 , . . . , am ∈ uj (V j ). Finally, for any p ∈ X, let
( P  
m u (p)−a
k=1 η j δ k , p ∈ Uj ,
ψj (p) :=
0, p ∈ X \ Uj .
x−ak

Note that ψj >  0 on Vj since η δ > 0 for |x − ak | < δ, and supp ψj ⊆ Uj since
Pm x−ak
k=1 η δ is 0 for x in a neighborhood of Rn − uj (Uj ). 
Remark 6.5. The preceding proof is typical for many nonconstructive, setthe-
oretic arguments in analysis. Actually, one almost always has canonically given
charts relative to which an explicit partition of unity can be provided.
Remark 6.6. The C ∞ partition of unity is an important tool which is used
to globally piece together locally given data in a smooth way. Although in some
applications we deal with analytic manifolds, generally we will stay in the category
of C ∞ manifolds, since there clearly is no analytic version of Theorem 6.4.

2. Geometry of C∞ Mappings
In elementary differential calculus, many geometrical questions concerning func-
tions (the location of extreme values, inflection points, etc.) can be answered by
investigating the derivatives (i.e., linear approximations) of the function. By means
of linear algebra, one can also study C ∞ mappings between manifolds. The essential
concepts for this are:

The directional derivative. Let x be a point of a C ∞ manifold X, ϕ ∈


C (X), and c : R → X a C ∞ map (a C ∞ curve) with c(0) = x. Then the direc-

tional derivative of the function ϕ in the direction of the curve c is defined to be


(ϕ ◦ c)0 (0), the derivative of ϕ ◦ c at 0 in the sense of elementary calculus. Two such
curves are equivalent, when the directional derivatives of each function relative to
the two curves are the same. We denote such an equivalence class by
0
c ∈ C ∞ (R, X) : e c) (0) = (ϕ ◦ c)0 (0) .

ċ(0) := e c(0) = x and ∀ϕ∈C ∞ (X) (ϕ ◦ e
Note that ċ(0) only depends on how c is defined near 0.

The tangent space. The set of directional derivatives, and hence the set of
equivalence classes of curves, forms a real vector space,
Tx X := (T X)x := {ċ(0) : c ∈ C ∞ (R, X), c(0) = x}
called the tangent space of X at x. Clearly, the multiplication of the directional
derivative c0 (0) by a real number λ is given by a λ-fold increase in the speed; i.e.,
˙
λċ(0) := c̃(0), where c̃(t) := c(λt) , for t ∈ R.
6.2. GEOMETRY OF C∞ MAPPINGS 161

Also, for two curves c1 and c2 , we can add ċ1 (0) and ċ2 (0) by setting
0
ċ1 (0) + ċ2 (0) := u−1 (u ◦ c1 + u ◦ c2 ) (0) ,
where u : U → Rn (n := dim X) is a chart with u(x) = 0 ∈ Rn and we arbitrarily
redefine c1 and c2 outside of a neighborhood of 0 ∈ R so that c1 (R) ∪ c2 (R) ⊂ U .
One can verify that these operations are well-defined and the axioms for a vector
space hold. Also one may check that dim(Tx X) = dim X. For this, one chooses
a C ∞ chart u : U → Rn , from the open neighborhood U of x to an open subset
of Rn . Then, for each positively directed coordinate line through u(x), there is
a corresponding C ∞ curve in X, as depicted in Figure 6.2. The corresponding

directional derivatives are denoted by ∂u 1
|x , . . . , ∂u∂ n |x and these form a basis for
Tx X.

Figure 6.2. Coordinate lines through u(x) and the corresponding


curves in X

The tangent bundle. The disjoint union T X := ∪x∈X Tx X of tangent spaces


has the structure of a real C ∞ vector bundle over X, namely the tangent bundle;
see Appendix, Exercise B.13, p. 719. A map s : X → T X is a section, if s(x) ∈ Tx X
for all x ∈ X. Let x ∈ X and let cx ∈ C ∞ (R, X) be a curve representing s(x); i.e.,
cx (0) = x and s(x) = ċx (0). Then the directional derivative of ϕ ∈ C ∞ (X) at x in
the direction s(x) is
0
s(x) [ϕ] := (cx ◦ ϕ) (0) .
If the function
s [ϕ] : X → R given by s [ϕ](x) := s(x) [ϕ]
is in C (X) for each ϕ ∈ C ∞ (X), then s : X → T X is a C ∞ section of T X or a

C ∞ vector field.
The differential. A C ∞ map f : X → Y determines a linear map (the dif-
ferential of f at x)
f∗x : Tx X → T Y given by f∗x (ċx (0)) = (f ◦˙ cx )(0) ,
f(x)

where ċx (0) ∈ Tx X. Sometimes we write f∗ |x instead of f∗x to clarify that the
differential f∗ is evaluated at x. Let u = (u1 , . . . , um ) : U → Rm and v =
(v1 , . . . , vn ) : V → Rn be coordinates about x and f (x) respectively, and let
(f1 , . . . , fn ) := v ◦ f ◦ u−1 : u(f −1 (V ) ∩ U ) → Rn .
162 6. DIFFERENTIAL OPERATORS OVER MANIFOLDS

∂ ∂
The matrix of f∗x with respect to the coordinate bases ( ∂u 1
,..., ∂um ) and
x x
( ∂v∂ 1 ,..., ∂
∂vn ) is given by
f(x) f(x)

∂ ∂fi
 
∂uj [fi ] = ∂uj (u(x)) ,
x
which is the n × m Jacobian matrix of (f1 , . . . , fn ). Note that f∗ : T X → T Y is a
bundle map (linear in the fibres, as explained in the Appendix p.712). Moreover, f
is called an immersion if f∗ is injective, an embedding if f and f∗ are injective
and f maps its domain X homeomorphically onto its image f (X) ⊂ Y , and a
submersion if f∗ is surjective at each x ∈ X.

Submanifolds. A subset Y ⊆ X is called a submanifold, if Y is a C ∞


manifold with the induced topology and the inclusion i : Y → X is a C ∞ embedding.
(One can also consider the image sets of immersions as submanifolds with self-
intersections, a concept that we will not pursue further.) For all y ∈ Y , Ty Y is a
linear subspace of Ty X in a natural way.
The following theorem illustrates some closely related representations of C ∞
manifolds.
Theorem 6.7. Let X and Y be C ∞ manifolds of dimensions m and n (m > n).
a) (Definition of manifolds through equations) Let q ∈ Y and f : X → Y be
a C ∞ map with f∗x surjective for all x ∈ f −1 {q}. Then f −1 {q} has the structure
of an (m − n)-dimensional submanifold of X, in a natural way.
b) (Representation of submanifolds of RN ) If Y is a submanifold of RN and
y ∈ Y , then (for a certain renumbering of the Euclidean coordinates x1 , ..., xN ),
the projection of Y to the n-dimensional subspace {(x1 , ..., xn , 0, . . . , 0)} ∼
= Rn is
n
a local coordinate system v : U → R for Y in some neighborhood U ⊆ Y of y.
Moreover, there is a neighborhood V of y in RN as in Figure 6.3, such that Y ∩ V
is the set of points satisfying the following system of equations for unique functions
gn+1 , . . . , gN ∈ C ∞ (v(U )):
xn+1 = gn+1 (x1 , ..., xn ) , . . . , xN = gN (x1 , ..., xn ) .
c) (Embedding Theorem) Every closed C ∞ manifold X can be embedded in RN
for N sufficiently large.

Remark 6.8. One might consider the special cases of part a) of Theorem 6.7
where the sphere is represented as a level set of the distance function, or where the
matrix manifold SL(n, R) is represented as a level set of the determinant function.
One can visualize (b) with Y = S 2 and N = 3. The proof of (a) and (b) follows
without difficulty from the Implicit Function Theorem (Exercise 6.1b). The com-
plete and elementary proofs for (b) and (c) can be found in [429, p.35-43] or [89,
p.91-92] where it is proved that every C ∞ n-manifold X can be embedded in R2n+1 .
Actually, H. Whitney proved that every C ∞ n-manifold X can be embedded in
R2n ; see [440]. Indeed, any closed smooth X can be embedded smoothly in R2n−1
if and only if the normal Stiefel-Whitney class wn−1 (X) = 0. For n 6= 4, this was
done in [198]. Decades later, the case n = 4 was finally settled by F. Fuquan in
[154] as a consequence of [63] by J. Boéchat and A. Haefliger and [124] by
S.K. Donaldson.
6.2. GEOMETRY OF C∞ MAPPINGS 163

Rn Y

y
u
V

RN

xN
xn+2
xn+1

Figure 6.3. Local representation of a submanifold Y (depicted


as a curve) by projection

We will not repeat the quoted proofs here, since we do not aim at minimizing the
dimension of the receiving Euclidean space. Instead of that we give an ultra-short
proof, here following [157, Section 2.2, pp.30-33]. This proof yields what we want,
namely an embedding into a finite-dimensional Euclidean space, but eventually of
quite high dimension. The basic idea goes back to work in algebraic geometry by
Kunihiko Kodaira, Fritz Hirzebruch and others, namely to fill large spaces
of functions or sections until an ample level is reached where one gets something
manageable or trivial.
Proof of Theorem 6.7c. We begin with an elementary set-theoretical ar-
gument, to illustrate the idea of filling: Any manifold X can be embedded in a huge
Euclidean space of highly infinite dimension, for suitable definition of the terms
topology, differential and embedding for infinite-dimensional manifolds. Indeed,
consider the natural mapping

ι : X −→ RC (X) given by x 7→ C ∞ (X) 3 f 7→ f (x) ∈ R ,


where RC (X) denotes the set of all mappings from the space C ∞ (X) of smooth
(here real-valued) functions to R, or, differently put, the direct product of copies of
R over all elements of C ∞ (X). This is a really huge Euclidean space. Each single
f ∈ C ∞ (X) may be perceived as a coordinate function, namely a reader or parser
of all x ∈ X. We give a formal argument for the fact that the preceding map ι is
an embedding. First we address the immersiveness. An analog

ι∗ : TX −→ T RC (X)
ċx (0) 7→ C (X) 3 f 7→ (f ◦˙ cx )(0) ∈ T R,

of the differential for ι is defined by taking the differential of each coordinate. Then
the mapping ι∗ maps any nonzero tangent vector v = ċx (0) to a nonvanishing

vector in T RC (X) . More precisely, for a nonzero tangent vector v, there is a
smooth function g so that the derivative of g in the direction of v is not zero. This
implies that the coordinate function, corresponding to g, of ι∗ (v) is nonzero.
164 6. DIFFERENTIAL OPERATORS OVER MANIFOLDS

Now, we check the injectivity of the mapping ι. For a pair x0 , x1 of distinct


points, there is a smooth function f so that the values f (x0 ) and f (x1 ) are distinct.

This implies that the points ι(x0 ), ι(x1 ) ∈ RC (X) take different values of the
coordinate function f , and we are done.
However, we want an embedding in a finite-dimensional Euclidean space. To do
that, we trim the full set C ∞ (X) of coordinate functions down to a finite number.
The set of the coordinate functions to be selected must be sufficiently ample to
yield still an embedding (the idea is very similar to the task of Exercise B.12).
Let’s begin anew. Denote by P(T X) the set of all tangent lines over all points
on X (as in our Exercise B.2 in the Appendix). The space P(T X) is a fiber bundle
over the compact base space X. (The concept of a fiber bundle embraces the concept
of vector bundles of our Appendix B, and the concept of principal G-bundles of our
Definition 15.1). Its fiber is a real projective space, which is compact. Hence, the
total space is also compact. For any tangent line ` at any point x, we can choose a
smooth function f` , whose derivative in the direction of ` is not zero. Then there
is an open neighborhood U` of ` in P(T X) such that all derivatives of f` for all
directions in U` are not zero. Since P(T X) is compact, we can cover it with finitely
many open subsets U`1 , U`2 , . . . , U`s . Then the mapping F := (f`1 , f`2 , . . . , f`s )
gives an immersion.
Unfortunately, the mapping F is not necessarily injective. However, since the
mapping F is an immersion, its restriction to a sufficiently small S neighborhood Ux
of an arbitrary x ∈ X is injective. We form the open subset x∈X Ux ⊂ X × X.
Its complement K is compact. By definition, for each pair y = (x0 , x1 ) ∈ K, the
points x0 and x1 are distinct. So there is an fy ∈ C ∞ (X) with fy (x1 ) 6= fy (x0 ).
Then there is an open neighborhood Vy of y in X × X such that fy (x01 ) 6= fy (x00 )
for all (x00 , x01 ) ∈ Vy . We select finitely many such open subsets Vy1 , Vy2 , . . . , Vyu
to cover the compact K. Now we set G := (fy1 , fy2 , . . . , fyu ). Then the mapping
(F, G) : X → Rs+u gives a desired embedding. 

The cotangent bundle. Let X be a C ∞ manifold with x ∈ X. In place of


the tangent space Tx X, one can consider its dual space Tx∗ X (other valid notations
∗ ∗
are(Tx X) , (T X)x , and(T ∗ X)x ) of linear maps from Tx X to R. An element of Tx∗ X
can be identified with the differential (see 4. above) at x of a real-valued function
ϕ ∈ C ∞ (X, R), namely ϕ∗x : Tx X → Tϕ(x) R ∼ = R. The notation (dϕ)x or dϕ|x is
also used for ϕ∗x . If u = (u1 , ..., un ) : U → X is a chart for X on a neighborhood
U of x, then the differentials (du1 )x , ..., (dun )x form a basis for Tx∗ X. One can also
give the disjoint union T ∗ X = ∪x∈X Tx∗ X a bundle structure, and indeed, T ∗ X is
exactly the dual bundle of T X; see Appendix, Exercise B.4, p. 715. For short, T ∗ X
is called the dual tangent bundle, covariant bundle, or most commonly, the
cotangent bundle.
Under a coordinate change v = κ ◦ u, the differentials change covariantly,
n
X
∂κi
dvi = ∂xj duj
j=1

while the tangent vectors transform contravariantly by means of the Jacobian of


κ−1 . From the standpoint of category theory, however, the tangent bundle is covari-
ant since f : X → Y yields a well-defined bundle map f∗ : T X → T Y whereas there
is a well defined bundle map fe: T ∗ X → T ∗ Y only when f is a diffeomorphism.
6.2. GEOMETRY OF C∞ MAPPINGS 165

Then we may define


 
(6.1) fe(αx ) (Z) := αx (f −1 )∗ (Z) for αx ∈ Tx∗ X and Z ∈ Tf(x) Y .

The situation for induced maps on sections is different, as we now explain. The
space of C ∞ sections of T X is denoted by C ∞ (T X) and such a section is known
as a vector field on X. A vector field on X generally does not push forward to
a well-defined vector field on Y unless f : X → Y is a diffeomorphism. Indeed, if
f is not onto, the purported push-forward will not be defined everywhere, while if
f is not 1-1, the purported push-forward may be ill-defined on f (X). The sections
in C ∞ (T ∗ X) (also denoted by Ω1 (X)) are known as 1-forms on X. For any f ∈
C ∞ (X, Y ) (not necessarily a diffeomorphism), there is a well-defined map

f ∗ : C ∞ (T ∗ Y ) → C ∞ (T ∗ X) given by

f (µ)(Zx ) := µf(x) (f∗ Zx ) for µ ∈ C ∞ (T ∗ Y ) and Zx ∈ Tx X,

which is known as the pull-back of 1-forms induced by f . Moreover, there is a


pull-back f ∗ : Ωk (Y ) → Ωk (X) of k-forms (see Appendix B and Exercise 6.20, p.
172) defined in the same way.
If X is a submanifold of Y with the embedding f : X ,→ Y , then although f is
not necessarily
 a diffeomorphism, we can still easily define f ∗ : T ∗ Y |f (X) → T ∗ X
∗ ∗
via f αf(x) (Zx ) := αf(x) (f∗ (Zx )) for αf(x) ∈ T Y |f (X) and Zx ∈ Tx X, and then
Ker f ∗ is a subbundle of T ∗ Y |f (X) known as the normal bundle of the embedding.
Put differently, the normal bundle consists of those covectors at points of f (X),
which annihilate all vectors tangent to f (X).

Remark 6.9. Cotangent bundles of manifolds arise naturally in abstract for-


mulations of classical mechanics and analytical mechanics, e.g., in the Hamiltonian
formulation of classical mechanics, which provides one of the major motivations
for the field: The set of all possible configurations of a system is modeled as a
manifold, and this manifold’s cotangent bundle describes the phase space of the
system. Locally, i.e., over a coordinate patch u(U ) ⊂ X for a chart u : U → X,
we have T ∗ X|u(U ) ∼ = R2n with the canonical (once the chart is chosen) coordi-
nates x1 , . . . , xn , du1 , . . . , dun , traditionally called the pairing of space and impulse
coordinates. For two such (x, µ) and (y, ν) we set

ω (x, µ), (y, ν) := h(x, µ), J(y, ν)i,
 
2n 0 In
where h·, ·i denotes the standard inner product in R and J := the
−In 0
usual skew-symmetric 2n × 2n block matrix. Then ω is a symplectic form (bilinear,
skew-symmetric and nondegenerate) for T ∗ X|u(U ) . Actually, the whole bundle T ∗ X
can be considered a symplectic manifold by defining (rather trivially) an exterior
nondegenerate skew-symmetric differential 2-form ω on it with dω = 0, see also our
Definition 18.23, p.652. A very readable introduction to local symplectic geometry
and Hamilton-Jacobi theory is given in [187, Chapter 5, pp.55-66, and Chapter
9, pp.97-106]. In the symbolic calculus of elliptic operators, the symplectic cone
T̊ ∗ X := T ∗ X \ X plays a fundamental role. It consists of the punctured cotangent
spaces.
166 6. DIFFERENTIAL OPERATORS OVER MANIFOLDS

Jets and Jet Bundles. A natural generalization of the cotangent bundle is


given by the concept of jets and jet bundles: While the cotangent bundle assembles
the differentials (the linear approximations) of a function, the jet bundles assemble
the higher derivatives (the Taylor expansions). We give a coordinate-free definition.
Definition 6.10. For a smooth manifold X and given x ∈ X, we denote by
Ix (X), or simply Ix , the ideal in C ∞ (X, C) of functions vanishing at x. If E → X
is a complex vector bundle over X we define
Zxk (E) := Ixk+1 · C ∞ (X; E), k = 0, 1, 2, . . . ,
and the complex vector space of k-jets of E at x, denoted by Jxk (E), is defined as
(6.2) Jxk (E) := C ∞ (X; E)/Zxk (E), k = 0, 1, 2, . . . .
The canonical linear map of C ∞ (X; E) into Jxk (E) is denoted by f 7→ jk (f )x and
jk (f )x is called the k-jet of f at x.
Remark 6.11. The cases k = 0, 1, 2 are of course well known for a function
f ∈ C ∞ (X), yielding the value, the differential and the Hessian of f at x for
j k (f )x ∈ Jxk (CX ). For general k, roughly speaking, jets are equivalence classes
of smooth maps between manifolds, which are represented by Taylor polynomials.
Probably the most intuitive description of a k-jet of f at x is the set of all partial
derivatives of f of order less than or equal to k.
Exercise 6.12. For any k ∈ Z+ and E → X define the k-jet bundle J k (E)
and prove the Jet Bundle Exact Sequence
0 −→ Lksym (T X, E) −→ J k (E) −→ J k−1 (E) −→ 0.
[Hint: You can work in local coordinates, or, equivalently, fix a local frame for
the tangent bundle T X and its dual frame for T ∗ X (for simplicity, assume that
X is Riemannian). Then show that the vector spaces {Jxk0 (E)}x0 ∈U (x,ε) in an ε-
neighborhood of x ∈ X can be trivialized nicely. Then define a space Lksym (V, W ) of
polynomials (symmetric real k-linear forms) on the vectors of a real vector space V
with values in a complex vector space W and define the symbol space Lksym (T X, E)
correspondingly. The projection J k (E) → J k−1 (E) is given pointwise by jxk (E) 7→
jxk−1 (E). To define the injection i : Lksym (T X, E) → J k (E) it suffices to give its
values pointwise for any v ∈ L(T Xx , R) = T ∗ Xx , respectively for the symmetric
form S k (v) and any e ∈ Ex . To do that represent v by a function ϕ ∈ C ∞ (X)
with ϕ(x) = 0 and dϕ|x = v and represent e by a section g ∈ C ∞ (X; E) with
1 k
g(x) = e. Then characterize i by the property i(S k (v) ⊗ e) := jk ( k! ϕ g)x . For a
very clear presentation see [328, pp.50-60]. Se also our coordinate-free description
of the principal symbol of a differential operator in Definition 6.36, p.183. It is
inspired by this exercise.]

3. Integration on Manifolds
Hypersurfaces. Suppose that X is an n-dimensional submanifold of Rn+1
(i.e., a hypersurface), and moreover assume that X is the boundary of a bounded
open subset of Rn+1 . From the notion of integration on Rn+1 where one has a
canonical volume element, we have a surface element on X, whence integration
over X is well-defined.
6.3. INTEGRATION ON MANIFOLDS 167

Riemannian Manifolds. In principle one can use the same recipe for a com-
pact C ∞ Riemannian manifold X.

Definition 6.13. A C ∞ manifold X is Riemannian, if it has been given a


Riemannian metric (tensor), namely for all x ∈ X the tangent space Tx X is
equipped with a fixed Euclidean metric tensor h··, ··i (positive, symmetric, nonde-
generate R-valued, bilinear form), such that for two C ∞ sections s1 and s2 of the
tangent bundle T X, the function hs1 , s2 i is in C ∞ (X).

It may be helpful to understand the concept of a Riemannian metric in local


1 n n
coordinates. Thus, let x ∈ X and u =  (u , . . . , u ) : U → R be coordinates with
∂ ∂
X ⊃ U 3 x, whence ∂u1 |x , . . . , ∂un |x a basis for Tx X. In these coordinates,  a met-
ric is represented by a positive definite, symmetric matrix gij (x) i,j=1,...,n , (i.e.,
gij (x) = gji (x) for all i, j and i,j gij (x)ξ i ξ j > 0 for all Rn 3 ξ = (ξ 1 , . . . , ξ n ) 6= 0),
P
where the coefficients depend smoothly on x.
Then the inner product of two tangent vectors A, PB ∈ ∂Tx X with coordinate
P j ∂
representations (a1 , . . . , an ) and (b1 , . . . , bn ) (i.e., A = i ai ∂u i and B = j b ∂uj )
becomes
X
hA, Bix = gij (x)ai bj .
1≤i,j≤n

∂ ∂
In particular, = gij (x). Similarly, the length of A is given by kAk :=
∂ui , ∂uj x
p
hA, Aix .
With the help of a C ∞ partition of unity (see Theorem 6.4, p. 159), one can
furnish every paracompact manifold with a Riemannian metric. Indeed, let {Uj }j∈J
be a locally finite covering of the n-manifold X by domains of coordinate charts
uj : Uj → Rn , say uj = u1j , . . . , unj , and let {ϕj ∈ C ∞ (X)}j∈J be a partition of


unity subordinate to {Uj }j∈J , then, for x ∈ X and A, B ∈ Tx X,


X  Xn 
d ukj (A) d ukj
 
gx (A, B) := ϕj (x) x x
(B)
j∈J with x ∈ Uj k=1

defines a Riemannian metric. Since there is no j-th term if x ∈/ Uj , the sum over J
is really finite on a neighborhood of each point.
On any submanifold X of the Euclidean space RN there is a natural Riemannian
metric induced by restricting the Euclidean inner product on RN to T X. Since we
have seen that any n-manifold X can be realized as a submanifold of RN for N
sufficiently large (Theorem 6.7c), we have another (less elementary) existence proof
for Riemannian metrics.

Geodesics. For a general metric space, a geodesic is defined as a curve which


realizes the shortest distance between any two sufficiently close points lying on it.
For Riemannian manifolds we can be more explicit. As in metric space, we ask
that geodesics are only locally the shortest distance between points. Additionally
we ask that they are parameterized with constant velocity, i.e., proportionally to arc
length. It is very fortunate that (locally) minimizing the energy will also minimize
the length — and give the wanted parametrization for free. The details of the
argument can be found in any textbook on Riemannian geometry, see, e.g., [235,
Section 1.4].
168 6. DIFFERENTIAL OPERATORS OVER MANIFOLDS

Let [t0 , t1 ] be a closed interval in R and c : [a, b] → X a smooth curve. The


length and the energy of c then are defined as
Z t1
1 t1
Z
L(c) := kċ(t)kdt and E(c) := kċ(t)k2 dt.
t0 2 t0
In physics, the massless term E(c) is usually called action of c where
P ic is considered

the orbit of a mass point. In local coordinates, we write ċ(t) = i γ (t) ∂u i |c(t) and

obtain
Z t1 X  12 1 t1 X
Z
i j
L(c) = gij (c(t))γ (t)γ (t) dt, E(c) = gij (c(t))γ i (t)γ j (t)dt.
t0 i,j
2 t 0 i,j

We also remark for later technical purposes that the length of a (continuous and)
piecewise smooth curve may be defined as the sum of the lengths of the smooth
pieces, and the same holds for the energy.
On a Riemannian manifold X, the distance dist(x, x0 ) between two points x, x0
can be defined as
inf{L(c) : c : [t0 , t1 ] → X piecewise smooth curve with c(t0 ) = x, c(t1 ) = x0 }.
If X is connected, it is also pathwise connected, i.e., any two points x, x0 ∈ X
can be connected by a path, actually by a piecewise smooth path. [Prove it by
decomposing X into the open (!) set Xx of all p ∈ X which can be connected with
x by a piecewise smooth path, and the open (!) complement X \ Xx consisting of
the union of all similarly defined sets Xq with q ∈
/ Xx . Since Xx is not empty, the
complement must be.] So, dist : X × X → [0, ∞) is well defined, and one checks
easily that it is a metric for X.
Definition 6.14. A smooth curve c : [t0 , t1 ] → X which is a critical point of
the energy functional is called a geodesic.
Recall that the Euler-Lagrange equations of a functional
Z t1
f t, c1 (t), . . . , cn (t), ċ1 (t), . . . , ċn (t) dt

I(c) :=
t0

for c = (c1 , . . . , cn ) : [t0 , t1 ] → Rn , are given by


d ∂f ∂f
− i = 0, i = 1, . . . , n.
dt ∂ ċi ∂c
Then the critical points of our energy functional E(c) are given by the system of n
second order differential equations
X
c̈i (t) + Γijk (c(t))ċj (t)ċk (t) = 0, i = 1, . . . , n
j,k

with
1 X i`
Γijk := g (gj`,k + gk`,j − gjk,` ),
2
`
where
X ∂
(g ij )i,j=1,...,n := (gij )−1 , (i.e., g i` g`j = δij ), and gj`,k := gj` .
∂uk
`

The expressions Γijk


are called Christoffel symbols. Christoffel symbols play
a prominent role in all concrete calculations with connections (a concept we shall
6.3. INTEGRATION ON MANIFOLDS 169

introduce on pp.176ff). In that context, they show up below in Section 6.5, Equation
(6.18), and become central in Section 15.5, in Equation (17.66) of Section 17.4,
p.558, and in analyzing Equations (18.107) in Section 18.4, p.697f in our Part IV,
beginning with Equations (15.47), (15.48) on p.421. The details of the preceding
deduction can be found in [235, Lemma 1.4.4] and many other places.
From the Local Existence and Uniqueness Theorem for systems of ordinary
differential equations we obtain
Proposition 6.15. Let X be a Riemannian manifold, x ∈ X, v ∈ Tx X. Then
there exist ε > 0 and precisely one geodesic c : [0, ε] → X (to be denoted by cv ) with
c(0) = x, ċ(0) = v.
One can show (and make precise) that, in addition, cv depends smoothly on x
and v. We then define
Definition 6.16. Let X be a Riemannian manifold, x ∈ X.
a) The mapping
expx : Vx → X
with Vx := {v ∈ Tx X : cv is defined on [0, 1]}
v 7→ cv (1)
is called the exponential map of X at x.
b) The point injectivity radius of x is

ρ(x) := sup ρ > 0 : expx is defined and injective on {v ∈ Tx X : kvk ≤ ρ} .

c) The injectivity radius of X is ρ(X) := inf{ρ(x) : x ∈ X}.


For example, the injectivity radius of the sphere X := S n is π, since the expo-
nential map of any point x maps the open ball of radius π in Tx X injectively onto
the complement of the antipodal point of x.
3. Orientation and Integrability. A C ∞ manifold X is oriented when
an atlas for X has been chosen such that the Jacobian matrix for each coordinate
change is positive (i.e., det(∂κi /∂xj ) > 0, for charts u and v, with u = κ ◦ v on
the intersection of their domains). More simply, without recourse to differential
calculus, we can also express this as follows. The bases of a finite-dimensional
vector space are divided into two orientation classes; two bases belong to the
same class, if their transformation matrix has positive determinant. By means of
an orientation of a C ∞ manifold, one may select two classes of bases from each
tangent space (positive and negative bases with regard to the fixed orientation) in
such a way that in a neighborhood of each point the choice is given by a continuous
or differentiable choice of basis.
The familiar Möbius band is an example of a nonorientable manifold. A sub-
manifold Y of an orientable manifold X (even of codimension only 1) is therefore
not always orientable. On the other hand, Y is automatically orientable if it is the
boundary of an open subset of an oriented manifold X. Then at each point y ∈ Y
one can define a basis of Ty Y to be positively oriented when a positive basis of
Tx X is obtained by adjoining an outward pointing vector. Since we will only be
interested in bounding manifolds with classically defined surface elements in most
of our applications, we state without proof that with the help of a suitable par-
tition of unity and local charts in which the metric can be expressed in terms of
curvilinear coordinates, the concept of integration of functions on Rn carries over
170 6. DIFFERENTIAL OPERATORS OVER MANIFOLDS

to the case of Riemannian manifolds. Orientability is required for the integration of


n-forms, because the sign of an n-form will make sense since it will not vary under
an orientation-preserving change of chart (see Exercise 6.20, p. 172). Actually, the
integration of R-valued functions (as opposed to n-forms) requires only a measure
or density (absolute value of an n-form) since the sign of such a function is already
unambiguous.
Definition 6.17. Let X be a smooth oriented paracompact Riemannian man-
ifold with dim X = n. We denote the metric tensor by g.
a) Let {Uj }j∈J be an open cover of X and {xj = (x1j , . . . , xnj ) : Uj → Rn } local,
positively oriented coordinates on Uj . Then each
p
νj,g := | det(gik )|dx1j ∧ · · · ∧ dxnj
defines a Lebesgue measure on each Uj and hence all together a Lebesgue measure
νg on X that is called the volume form of X.
b) Let π : E → X be a smoothp complex vector bundle over X with Hermitian
metric h·, ·ih . Set |e|h := he, eih for e ∈ E. An L2 -section of E is a Lebesgue
measurable map ψ : X → E (i.e., ψ −1 (U ) is Lebesgue measurable for any open
subset U ⊂ E) such that
(i) π ◦ ψ(x) = x for almost all x ∈ X except possibly a negligible set.
(ii) The function x 7→ |ψ(x)|h belongs to L2 (X, R).
The space of L2 -sections of E is denoted by L2 (E).
In the notation of Exercise 6.20 of the following section, we have νg = ∗0 (1) ∈
Ωn (X). That yields a coordinate-free definition
R of the volume form. It is worth
mentioning that the volume vol(X) := X νg ∈ [0, ∞) ∪ {∞} is well defined. We
leave it to the reader to check that L2 (E) is a Hilbert space with respect to the
scalar product
Z
(ϕ, ψ)0 := hϕ(x), ψ(x)iEx ,h νg ϕ, ψ ∈ L2 (E).
X

Remark 6.18. For some computations on manifolds it is impractical and con-


fusing to constantly revert to local coordinates. In such cases intrinsic coordinate-
invariant concepts of integration like the preceding definition of the volume form
by the linear star operator are welcome, and they require in part weaker hypothe-
ses. For the integration of n-forms, the existence of a volume element is essential,
or, more generally in modern terminology, the existence of a distinguished n-form
where n = dim X. See, e.g., [400, p.134/150].
Remark 6.19. For the time being, we will use only the above integration as-
pects of the Riemannian metric, and thus only scratch the surface of Riemannian
geometry which unfolds in the great classic theorems on parallel displacement (con-
nections), curvature and rigidity with their varied computations (see [54]). Our
Part IV deals extensively with such topics, in particular with connections on prin-
cipal G-bundles, rigorously defined in Section 15.2, p.396. Already in Section 6.5,
p.176ff, connections on vector bundles will be introduced and used for coordinate
free integration in our definition of Sobolev spaces and pseudo-differential oper-
ators in Chapters 7 and 8. We have noted (see (3.32), p. 117) that there is a
close relationship (the Gauss-Bonnet Theorem) between the integral of the Gauss-
ian curvature and the topological form of a surface, namely the genus or Euler
6.4. EXTERIOR DIFFERENTIAL FORMS AND EXTERIOR DIFFERENTIATION 171

characteristic. See also Chapters 12/13 below in Part III, and Part IV where the
higher-dimensional Gauss-Bonnet-Chern Theorem is proved using the local index
theorem (see Theorem 17.68, p. 612).

4. Exterior Differential Forms and Exterior Differentiation


I It is possible to construct (see Exercise 6.20 below), from the tangent bundles,
bundles of exterior differential forms by means of multilinear algebra. These are an
important tool in describing physical laws mainly in the areas of electromagnetism and
special relativity. This is the case when empirical relationships are to be expressed in
terms of an integral in such a way that the physicist or engineer can pursue qualitative
and quantitative changes resulting from modifications of the integrand or the domain
of integration. Such applications have stimulated further studies of exterior bundles in
differential topology; see our Chapters 13 (pp.310–362) and 15–18 (pp.394–704). J

We briefly summarize (details are found in [89, p.260f], [248, p.17f], and [356,
p.111-161] and the literature given there — for a quick guide to the content of this
and the following section we recommend the crash course [453, Chapter 1]; for an
extended elaboration see our Chapter 15, pp.394ff): For a real n-dimensional vector
space V , we form the vector space Λp (V ) of p-fold skew-symmetric tensors (or
p-vectors); these are the multilinear maps
p times
V ∗ × · · · × V ∗ → R, p ∈ N, V ∗ := L(V, R)
which change, under a permutation of the arguments, by a factor equal to the sign
of the permutation. One sets Λ0 (V ) := R and obtains Λ1 (V ) = V , Λn−1 (V ) ∼ = V,
Λn (V ) ∼
= R and Λp (V ) = {0} for p > n. For v ∈ Λp (V ) and w ∈ Λq (V ), we define
v ∧ w ∈ Λp+q (V ) by
1 X
(v ∧ w)(a1 , ..., ap+q ) := sgn(σ)(v ⊗ w) (aσ(1) , ..., aσ(p+q) )
p!q! σ

(sum over all permutations), which gives the exterior multiplication Pn Λp (V ) ×


Λq (V ) → Λp+q (V ). This multiplication makes Λ• (V ) := p=0 Λp
(V ) a graded
algebra, the exterior algebra of V .
If e1 , ..., en is a basis of V , then the np forms ei1 ∧ · · · ∧ eip with 1 ≤ i1 < · · · <


ip ≤ n yield a basis for Λp (V ). With this property, Λp (V ) is occasionally defined


(in order to avoid the suggestive but tedious definition via maps) as the space of
p-vectors: the space of formal linear combinations of the p-tuples of basis vectors
ei1 ∧ · · · ∧ eip with only the relation eσ(i1 ) ∧ · · · ∧ eσ(ip ) = sgn(σ)ei1 ∧ · · · ∧ eip .
A scalar product (= inner product) for V induces a scalar product h·, ·i for
Λp (V ) by

(6.3) hv1 ∧ · · · ∧ vp , w1 ∧ · · · ∧ wp i := det hvi , wj i .
Declaring an orthonormal basis e1 , ..., en of V to be positively oriented yields an
explicit isomorphism Λn (V ) ∼
= R via e1 ∧ · · · ∧ en 7→ 1, which only depends on the
chosen orientation and scalar product. The linear star operator
∗p : Λp (V ) → Λn−p (V ) is generated by
(6.4)
e i1 ∧ · · · ∧ e ip 7 → ej1 ∧ · · · ∧ ejn−p ,
where j1 . . . jn−p is selected such that ei1 , . . . eip , ej1 . . . ejn−p is a positive basis of
V.
172 6. DIFFERENTIAL OPERATORS OVER MANIFOLDS

Since the star operator is supposed to be linear, it is determined by its values


on some basis (6.4). It is characterized by the property u ∧ ∗p v = hu, vi e1 ∧ · · · ∧ en
for all u, v ∈ Λp (V ). In particular,
(6.5a) ∗0 (1) = e1 ∧ · · · ∧ en and
(6.5b) ∗n (e1 ∧ · · · ∧ en ) = 1,
if e1 , . . . , en is a positive basis. From the rules of multilinear algebra, it easily
follows that if A ∈ End(V ), and if f1 , . . . , fp ∈ V , then
∗p (Af1 ∧ · · · ∧ Afp ) = (det A) ∗p (f1 ∧ · · · ∧ fp ).
In particular, this implies that the star operator does not depend on the choice
of positive orthonormal basis in V , as any two such bases are related by a linear
transformation with determinant 1. For a negative basis instead of a positive one,
one gets a minus sign on the right hand sides of (6.4), (6.5a), (6.5b).
Exercise 6.20. Let X be a compact C ∞ manifold of dimension n with or
without boundary, with metric tensor g.
a) Show that the family of vector spaces Λp (Tx∗ X), x ∈ X, yields a real vector
bundle of fiber dimension np over X in a natural way. We denote this bundle by
Λp (T ∗ X) or shortly Λp (X). Correspondingly, define a bundle Λ• (T ∗ X) by sum-
mation. Check that g induces a smoothly varying inner product on the fibres of
Λp (T ∗ X) and Λ• (T ∗ X).
b) Customarily, one writes Ωp (X) := C ∞ (Λp (T ∗P X)), which is the space of exte-
n
rior differential p-forms on X and Ω• (X) = p=0 Ωp (X). For a C ∞ function
f , consider the differential df (see also Section 6.1) and show that the operator
d : Ω0 (X) → Ω1 (X) uniquely extends to a linear differential operator of first order
for each p = 0, . . . , n − 1
(6.6) d : Ωp (X) −→ Ωp+1 (X), such that d2 := d ◦ d = 0 and
p
d(α ∧ β) = (dα) ∧ β + (−1) α ∧ dβ for all α ∈ Ωp (X) and β ∈ Ωq (X).

c) Once we have an orientation of X, the definition of the linear star operator


carries over from the vector spaces Λp (Tx∗ ) to the vector bundles Λp (T ∗ X) and,
finally, to the vector spaces Ω• (X), then called Hodge star operator. Apply
(6.5a) to obtain
p
(6.7) ∗0 (1) = |det(gik )|dx1 ∧ · · · ∧ dxn
in local coordinates.
Prove the Hodge duality ∗p : Ωp (X) ∼= Ωn−p (X). Assume n even, and check whether
n
∗ := ⊕p=0 ∗p is an involution. If not, how should one modify ∗ to obtain an involu-
tion?
d) Prove that for a compact, oriented, n-dimensional, Riemannian manifold X with
boundary ∂X, we have Stokes’ Theorem
Z Z
dω = ω, for ω ∈ Ωp (X).
X ∂X

[Hint for a): In principle, use the same mechanism as in Exercise B.4. Note that for
charts u and w for X in a neighborhood of x ∈ X, we have the simple transformation
6.4. EXTERIOR DIFFERENTIAL FORMS AND EXTERIOR DIFFERENTIATION 173

rules, e.g., for a 1-form v ∈ Ω1 (X),


Xn Xn
v(x) = aj (x) duj x
= bi (x) dwi x
, where
j=1 i=1
−1 j

Xn ∂ u◦w
bi (x) := aj (x) (w(x)) .
j=1 ∂xi
For b): d is characterized by the Leibniz rule
d(v ∧ w) = dv ∧ w + (−1)p v ∧ dw, for v ∈ Ωp (X), w ∈ Ωq (X).
How is d written in local coordinates?
For c): Note that ∗n−p ∗p = (−1)p IdΩp (X) . That yields Hodge duality, but no
involution ∗. Assuming n even, say n = 2m, does not help. However, try
(6.8a) τp : = im+p(p−1) ∗p . Then
m+(2m−p)(2m−p−1) m+p(p−1)
(6.8b) τ2m−p ◦ τp = i i (−1)p ,
±
and so τ := ⊕2m 2
p=0 τp is an involution (τ = 1 = IdΩ• (X) ). The ±1 eigenspaces Ω
of τ are crucial for defining the Hirzebruch signature operator, see Sections 12.3
(p.302), 13.4 (p.321), and 17.6, Theorem 17.64 (p.605), where some of the preceding
tasks are executed, both in greater generality — and in more detail.
For d): See [196, p.182-187]. Incidentally, here one really needs the orientation.]
Remark 6.21. In algebraic terms, the exterior algebra Λ• (T ∗ X) is the universal
unital algebra generated by T ∗ X subject to the relations ξ ∧ ξ = 0 for ξ ∈ T ∗ X. To
sum up, the wedge product extends to the full bundle Λ• (T ∗ X) of exterior algebras
and to the space Ω• (X) of exterior differential forms on X. The mapping
w∧ : Ω• (X) → Ω• (X), ξ 7→ w ∧ ξ
is called (left) exterior multiplication.
There is also an interior multiplication for skew-symmetric tensors, exterior
algebras and exterior differential forms. Exterior multiplication adds an index if
possible and increases the degree of forms; interior multiplication cancels an index
and decreases the degree of forms. More precisely, let V be a real vector space
equipped with an inner product h·, ·i and dim V = n. Then there is a linear mapping
wx : Λp (V ) → Λp−1 (V ) for each w ∈ V ,
defined via
p
X
(6.9) wx(v1 ∧ · · · ∧ vp ) := hvj , wiv1 ∧ · · · ∧ vbj ∧ · · · ∧ vp ,
j=1

where vbj means that the factor vj is omitted.


Relative to the induced inner product on Λ• (V ), the mappings w∧ and wx
are adjoints (proved in (13.17), p.322, see also Section 17.1 and Proposition 17.11,
p.519). Like exterior multiplication, the mapping wx induces endomorphisms of
Λ• (T ∗ X) and Ω• (X), called the (left) interior multiplication.
The geometric meaning of interior multiplication is explained in Sections 6.6
(p.185) and 13.4 (p.321) for the principal symbol of the codifferential δ = d∗ ; in
Section 13.5 (p.329) for discussing the number of vector fields; and in Section 17.11
(p.519) for spinor representations.
174 6. DIFFERENTIAL OPERATORS OVER MANIFOLDS

The operator d : Ωp (X) → Ωp+1 (X) is known as the exterior derivative


operator or exterior differentiation. In local coordinates, x1 , . . . , xn defined
on a coordinate neighborhood U ⊆ M , a form α ∈ Ωp (X) can be written as
1 X
α= αi1 ···ip dxi1 ∧ · · · ∧ dxip ,
p!
1≤i1 ,...,ip ≤n

where the αi1 ···ip ∈ C (U ) are antisymmetric in the indices i1 , . . . , ip . On U ,
1 X
d αi1 ···ip ∧ dxi1 ∧ · · · ∧ dxip

dα =
p!
1≤i1 ,...,ip ≤n
1 X

αi1 ···ip dxi ∧ dxi1 ∧ · · · ∧ dxip .

= ∂xi
p!
1≤i1 ,...,ip ≤n

However, since d : Ω (X) → Ωp+1 (X) is uniquely determined by the coordinate-free


p

operation d : Ω0 (X) → Ω1 (X), one should be able to express d in a coordinate-free


manner. For this purpose (and because it is an important and basic notion), we
introduce Lie differentiation.
Let A ∈ C ∞ (T X) be a vector field on X. The theory of systems of first-order
ordinary differential equations guarantees that for each point p ∈ X, there is ε > 0
and a curve αp : (−ε, ε) → X such that α̇p (t) = Aαp(t) . This curve α is known as
the integral curve of A through p. Furthermore, ε can be chosen so that for all q
in some neighborhood U of p, the integral curve αq : (−ε, ε) → X exists, and there
is a well-defined C ∞ map
α : U ×(−ε, ε) → X given by α(q, t) := αq (t) ,
such that αt := α(·, t) : U → α(U, t) is a diffeomorphism for each t ∈ (−ε, ε).
Moreover, αt+s = αt ◦ αs whenever both sides are defined. In particular αt−1 = α−t
on α(U, t) ∩ U which is nonvoid for small t. Given a second vector field B ∈
C ∞ (T X), we have αt−1 ∗ Bαt(p) ∈ Tp X and so the curve
t 7→ αt−1 ∗ Bαt(p) = (α−t )∗ Bαt(p)
  

lies in the single vector space Tp X. It then makes sense to differentiate this curve
at t = 0 to obtain a vector in Tp X which is known as the Lie derivative of B with
respect to A at p, namely
d

(LA B)p := dt (α−t )∗ Bαt(p) t=0 .
It is also called the Lie bracket of A and B at p. We explain why: The
assignment p 7→ (LA B)p defines a vector field LA B ∈ C ∞ (T X). If A and B are
vector fields on Rn and p, δp ∈ Rn with |δp|  1, then (where ≈ denotes equality
modulo terms of first-order in t) we have
(α−t )∗ (δp) ≈ δp − t(dA)p (δp) and Bαt(p) ≈ Bp+tAp ≈ Bp + t(dB)p (Ap ) .
Thus,
 
(α−t )∗ Bαt(p) ≈ Bαt(p) − t(dA)p Bαt(p) ≈ Bp + t(dB)p (Ap ) − t(dA)p (Bp )
⇒ (LA B)p = (dB)p (Ap ) − (dA)p (Bp ) = Ap [B] − Bp [A] .
Hence, LA B = −LB A and viewing the vector fields A and B as differential opera-
tors on functions f , we have
(LA B)p (f ) = (df )p (Ap [B] − Bp [A]) = Ap [B [f ]] − Bp [A [f ]] ,
6.4. EXTERIOR DIFFERENTIAL FORMS AND EXTERIOR DIFFERENTIATION 175

and so as a differential operator LA B is the commutator of A and B, i.e.,


LA B = A ◦ B − B ◦ A = [A, B] .
In local coordinates x1 , . . . , xn , we have (automatically summing over repeated
indices)
LA B = LA B xi ∂xi = A B xi − B A xi
       
∂xi
j i j i

= A ∂xj B − B ∂xj A ∂xi .
Now we show how the exterior derivative can be expressed by the Lie derivative.
For a one-form ω, we then have the coordinate-free relation
(6.10) dω(A, B) = A [ω(B)] − B [ω(A)] − ω([A, B]) , since

A [ω(B)] − B [ω(A)] − ω([A, B])


= A ωj B j − B ωj Aj − ω Ai ∂xi B j ∂xj − B i ∂xi Aj ∂xj
      

= Ai B j ∂xi ωj + ωj ∂xi B j − B i Aj ∂xi ωj + ωj ∂xi Aj


 

− ωj Ai ∂xi B j + ωj B i ∂xi Aj
= ∂xi ωj Ai B j − B i Aj = ∂xi (ωj ) dxi ∧ dxj (A, B) = dω(A, B) .
 

Exercise 6.22. For ψ ∈ Ω2 (X) and vector fields A, B, C ∈ C ∞ (T X), show


that
dψ(A, B, C) = A [ψ(B, C)] + B [ψ(C, A)] + C [ψ(A, B)]
− ψ([A, B] , C) − ψ([C, A] , B) − ψ([B, C] , A)
= S(A [ψ(B, C)] − ψ([A, B] , C)) ,

where S denotes the sum over all cyclic permutations of (A, B, C). [Hint. Consider
the case ψ = ω ∧ ϕ, for ω, ϕ ∈ Ω1 (X), and use dψ = dω ∧ ϕ − dϕ ∧ ω. The general
case follows by linearity, since any 2-form is locally a sum of wedges of 1-forms.]

More generally, one can show by induction that for ψ ∈ Ωk (X)


k+1
X h  i
i+1
dψ(A1 , . . . , Ak+1 ) = (−1) Ai ψ A1 , . . . , A
ci , . . . , Ak+1
i=1
X  
i+j
(6.11) + (−1) ψ [Ai , Aj ] , A1 , . . . , A
ci , . . . , A
cj , . . . , Ak+1 ,
1≤i<j≤k+1

where Aci indicates that Ai is omitted. For a proof, see [248, p.36]. Note that the
extra numerical factor of 1/(k + 1) in the
 formula of [248, p.36] is ultimately due to
their convention that dx1 ∧ · · · ∧ dxn (∂x1 , . . . , ∂xn ) = 1/n! (see [248, p.7]), while
our convention is that dx1 ∧ · · · ∧ dxn (∂x1 , . . . , ∂xn ) = 1.
An extensive discussion of Lie brackets, and the general theory of Lie derivatives
can be found in [59, Chapter 0] and [251, Chapters I-II]. In particular, see the last
reference (Sections II.7.6-II.7.9, pp.63-66) for the place of interior multiplication
and for further relations between exterior differentiation and Lie derivation.
176 6. DIFFERENTIAL OPERATORS OVER MANIFOLDS

5. Covariant Differentiation, Connections and Parallelity

I Children of our motorized time are familiar with the concepts of speed and acceler-
ation and able to clearly distinguish between them: A car can move on a straight highway
with high speed, constant velocity and no acceleration; and it can move after a stoplight
with low velocity, but high acceleration. So much for small children. When they grow older
and have learned about the interpretation of force as a product of mass and acceleration
(Newton’s Second Law), most of them will fall back to the pre-Newtonian identification of
velocity and acceleration. Ask them to draw the trajectory of a thrown ball or rock! Most
will correctly draw a parabola which is a good approximation for a rock that is thrown
for short distances. But then ask them to mark the acting forces by directed arrows along
the trajectory! Most will draw tangent vectors of varying length (the impulses) instead of
the solely vertically acting constant gravitation. Some of the wise ones would explain the
apparent contradiction by referring to resulting force or to air resistance. The smartest
of all of them was the Greek philosopher and polymath Aristotle (384 - 322 BCE) who
derived a straight trajectory of finite length for the thrown stone until the impulse of the
initial throw was consumed, followed by vertical fall-down. To his student Alexander III
of Macedon, the later famous military leader and creator of one of the largest empires
of the ancient world, he explained the visible deviation of his theoretical trajectory from
observed orbits by the complexity of full reality, air resistance, wind influence, imperfect
shape of the thrown object etc.
A modern geometer may have two comments to that continuing confusion of concepts.
(I) Analyzing a single trajectory is not very challenging. Elementary calculus yields simple
definitions of the velocity ċ(t) ∈ R3 and the acceleration c̈(t) ∈ R3 of a sufficiently smooth
path c : [t0 , t1 ] → R3 at a point c(t) ∈ R3 for t ∈ [t0 , t1 ]. The student will see at once that
the vectors ċ(t) and c̈(t) have different directions, in general, and even can be perpendic-
ular to each other in natural parametrization, namely c̈(t) pointing to the center of the
curvature. Moreover, writing the equations of motion for the curve c with given initial
position, velocity and acceleration c(t0 ), ċ(t0 ), c̈(t0 ) in (x, y, z) coordinates gives a simple
one-dimensional problem. Only the vertical z-coordinate is relevant. One can neglect the
y-coordinate for a plane movement and the movement in x direction is not accelerated
in the absence of forces in that direction. In flat R3 , we can do without distinguishing
between the spaces of state (configuration), velocity (tangent) and acceleration (forces) as
long as we do correct calculations. At first, the geometer may wonder about the success
of classical mechanics with concepts that belong to different categories but are commonly
put in the same (Euclidean) space. Thinking about it, the geometer will explain the as-
tonishing correctness of sloppy and vague physics terminology by the flatness of Euclidean
space. There is nothing to worry about.
(II) In a second comment, the geometer would admit that there is a lot to worry
about. Recall Section 6.2, where it is natural to distinguish between the points of a
manifold X and equivalence classes of paths making its tangent bundle T X. We did
it in a coordinate-free manner, admitting non-Euclidean X, rigorously, canonically and
without special choices or ambiguities. Later we made choices, to put a Riemannian
structure, i.e., smoothly varying metrics on the tangent spaces. But the basic concept
of T X was canonical. Now, similarly, we might wish to define a second derivative —
canonically. That is impossible. Why? Consider a vector field s ∈ C ∞ (X; T X), i.e., a
section in the bundle T X → X. It specifies a direction, i.e., a tangent vector s(x) ∈ Tx X
in each point x ∈ X. We may consider the vector field s as a field of velocities. To get
something like a second derivative, an acceleration or a force we would specify a direction
v ∈ Tx X and take the limit

!? s(x + tv) − s(x)


(6.12) ∇v s|x : = lim .
t→0 t
6.5. COVARIANT DIFFERENTIATION, CONNECTIONS AND PARALLELITY 177

That looks familiar — except for two problems: We have to define a translation “+tv”
for small real t yielding a point x + tv in the neighborhood of x. That can be done using
a Riemannian metric for X (in contrast to the fact that the concept of tangent vector,
tangent space and tangent bundle was defined fully invariantly and without reference to
a metric). The second problem is more serious: There is no canonical way in T X of how
to compare two tangent vectors at different base points. We have to make choices. We
have to make parallel translations and to specify the ways to do them for a given bundle.
That is the concept of a connection. J

In this Part II, we use only a very simple concept of connection, first and most
general for real and complex vector bundles, and then more specifically for the
tangent bundle. We would like to emphasize that the concept of a connection has
many more ramifications than the few dry definition terms we give in this section.
It has become the central concept in modern low-dimensional geometry and, as
well, in gauge-theoretic quantum field theory and particle physics. That will be
explained in our Part IV.

Connections on Vector Bundles. Let X be a smooth manifold and E a


smooth real or complex vector bundle over X.
Definition 6.23. A connection (or covariant differentiation operator)
∇E on E is an R-linear first order differential operator
∇E : C ∞ (X; E) → C ∞ (X; T ∗ X ⊗ E)
satisfying the Leibniz rule
(6.13) ∇E (f s) = df ⊗ s + f ∇E s
for all functions f ∈ C ∞ (X, R) and sections s ∈ C ∞ (X; E).
Recall that the bundles T ∗ X ⊗ E and Hom(T X, E) are isomorphic (as real
vector bundles, see also Exercise B.4, p.715 for the complex category). Whence, we
can consider ∇E as a mapping
C ∞ (X; E) × C ∞ (X; T X) 3 (s, v) 7→ ∇Ev s := ∇E (s) (v) ∈ C ∞ (X; E)


with the following properties for v, w ∈ C ∞ (X; T X), s, r ∈ C ∞ (X; E), f ∈ C ∞ (X, R)
and c ∈ R:
(6.14) tensorial in v: ∇Ev+w s = ∇Ev s + ∇Ew s and ∇Efv s = f ∇Ev s;
(6.15) linear in s: ∇Ev (s + r) = ∇Ev s + ∇Ev r and ∇Ev cs = c∇Ev s;
product rule: ∇Ev (f s) = (df ) ∇Ev s + f ∇Ev s .
 
(6.16)

Exercise 6.24. a) Set X := Rn , E the (trivial) product bundle X × CN


(denoted by CN X in Exercise B.1a of the Appendix B, p.713) and show that the
attempted definition of (6.12) yields a connection.
b) Show that each complex or real vector bundle E over a smooth manifold X
admits a connection.
[Hint for (a): Check that all the properties of the preceding list are satisfied.
Hint for (b): Choose a locally finite covering of X by charts, and choose trivial-
izations of E; apply (a); assemble the local connections to a global operator; and
check the properties.]
178 6. DIFFERENTIAL OPERATORS OVER MANIFOLDS

Parallel Transport. Closely related to the formal definition of a connection


is the concept of parallel transport of geometric information.
Definition 6.25. Let ∇E be a connection for a vector bundle πE : E → X.
A section s ∈ C ∞ (X; E) is called parallel relative to ∇E along a smooth path
c : [t0 , t1 ] → X in X, if
(6.17) ∇Eċ(t) s(c(t)) = 0 for all t ∈ (t0 , t1 ).

Exercise 6.26. Let X, E, ∇E , c be as in the preceding definition. Define a


parallel translation
E
τc,t : Ec(t) −→ Ec(t0 ) for t ∈ (t0 , t1 ).
[Hint: Begin with local coordinates around x0 := c(t0 ). So, choose a coordinate
patch X ⊃ U 3 x0 and coordinates x1 , . . . , xn yielding coordinate vector fields
∂ ∂
∂x1 , . . . , ∂xn in T X|U . For simplicity, consider E as a real vector bundle of fiber di-
mension N . Through the identification E|U ∼ = U ×RN you obtain a basis s1 , . . . , sN
of sections of E|U . For the given connection ∇E , define the so-called Christoffel
symbols Γkij (j, k = 1, . . . , N, i = 1, . . . , n) by the condition
N
X
(6.18) ∇E∂ sj =: Γkij sk .
∂xi
k=1

See also our geometric interpretation of the Christoffel symbols in Equations (15.47)

and (15.48) of Part P IV below on pp.421ff. Let now s ∈ C (X; E). Locally, you
may write s(y) = k ak (y)sk (y). Putting s(t) := s(c(t)), you define a section of E

along c. Furthermore, let ċ(t) =: γ i ∂x
P
i . Then by (6.14)-(6.16) and (6.18), find
X X
∇Eċ(t) s(t) = ak (t)sk (c(t)) + γ i (t)ak (t) ∇E∂ sk (c(t))

∂xi
k i,k
X X
= ak (t)sk (c(t)) + γ i (t)ak (t)Γjik (c(t))sj (c(t)).
k i,j,k

Note that ∇Eċ(t) s(t) depends only on the values of s along the curve c, and not on
all the values of s in a neighborhood of the trajectory {c(t) : t ∈ [t0 , t1 ]}. Our
Equation (6.17) thus represents a linear system of first order ordinary differential
equations for the coefficients a1 (t), . . . , aN (t) of the section s(t) along c that you are
looking for. Therefore, for given initial values s(t0 ) ∈ Ec(t0 ) , you obtain a unique

=
solution of (6.17). This gives you an isomorphism Ec(t0 ) −→ Ec(t) for all t ∈ [t0 , t1 ].
Take the inverse as the wanted parallel translation.]
Thus, if x, x0 ∈ X, the fibers of E above x and x0 , Ex and Ex0 , respectively, can
be identified by choosing a curve c from x to x0 (x = c(0), x0 = c(1)) and moving
each s0 ∈ Ex along c to Ex0 by parallel translation. This identification depends
only on the choice of the curve c. Now assume that X is a compact manifold with
Riemannian metric g and injectivity radius ρ > 0 (see Definition 6.16, p. 169).
Then we have a geodesic with respect to the Riemannian metric g as canonical
curve, which is uniquely determined by the endpoints x, x0 , if dist(x, x0 ) < ρ.
Definition 6.27. Let X be a compact Riemannian manifold with injectivity
radius ρ > 0 and let E → X be a vector bundle equipped with a connection ∇E .
6.5. COVARIANT DIFFERENTIATION, CONNECTIONS AND PARALLELITY 179

For x, x0 ∈ X, with dist(x, x0 ) < ρ, we denote the parallel translation relative to


∇E along the unique geodesic from x0 to x with minimal length dist(x, x0 ) by
E
τx,x 0 : Ex0 −→ Ex .

Remark 6.28. a) In the preceding definition, we obtain parallel translation on


a Riemannian manifold from a connection. Conversely, we can regain a connection
from parallel translation by the recipe of (6.12) at the beginning of this section.
b) To explain the name connection, we refer to our Section 15.2, p.396ff, where
we develop the topic Connections and Curvature. In particular, we refer to Figure
15.1, p.396. More precisely (replacing P by E and M by X in the figure), consider
the tangent space Tp E at the point p ∈ E to the total space E of a vector bundle
π : E → X. Inside Tp E, there is a distinguished subspace, namely the tangent
space to the fiber Ex containing p (x = π(p)). This space is called the vertical
space Vp . However, there is no distinguished horizontal space Hp complementary
to Vp , i.e., satisfying Tp E = Vp ⊕ Hp . If we have a covariant derivative ∇E ,
however, we can parallely transport p for each v ∈ Tx X along a curve cv (t) with
cv (0) = x, ċv (0) = v. Thus, for each v, we obtain a curve pv (t) in E. The subspace
d
of Tp E spanned by all tangent vectors to E at p of the form dt pv (t)|t=0 then is a
suitable choice of a horizontal space Hp . In this manner, one obtains a rule how
the fibers in neighboring points are connected with each other.
Definition 6.29. Let π : E → X be a real or complex vector bundle on the
differentiable manifold X with Euclidean, respectively Hermitian bundle metric
h·, ·i. A connection ∇ on E is called metric (also Riemannian, respectively
Hermitian), if
(6.19) dhs, ri = h∇s, ri + hs, ∇ri for all s, r ∈ C ∞ (X; E).
A metric connection thus has to be compatible with the metric. To read
(6.19) correctly, you notice that hs, ri is a smooth function on X, whence dhs, ri ∈
C ∞ (T ∗ X). Applying the left side of (6.19) to a tangent vector field v ∈ C ∞ (T X)
yields a smooth function on X. Similarly, each of the two terms on the right
side of (6.19) yields a function when the vector field v is inserted into ∇s, ∇r ∈
Hom(T X, E).

Connections on the Tangent Bundle. Connections on the tangent bundle


T X are particularly important:
Definition 6.30. Let ∇ be a connection on the tangent bundle T X of a dif-
ferentiable manifold X.
a) A curve c : (t0 , t1 ) → X is called autoparallel or geodesic with respect to ∇,
if ∇ċ c ≡ 0, i.e., if the tangent field ċ of c is parallel along c.
b) The torsion tensor of ∇ is defined as
T (X, Y ) := ∇X Y − ∇Y X − [X, Y ] (X, Y ∈ C ∞ (T X)).

c) The connection ∇ is called torsion free, if T ≡ 0.


A classical result of Riemannian geometry (reformulated in greater generality
and proved in details below in Section 15.5, Proposition 15.39, pp.418ff) is now:
180 6. DIFFERENTIAL OPERATORS OVER MANIFOLDS

Theorem 6.31. On each Riemannian manifold X, there is precisely one con-


nection θ (on T X) that is metric (i.e., compatible in the sense of Definition 6.29)
and torsion free. It is determined by the formula
1
(6.20) hθX Y, Zi = XhY, Zi − ZhX, Y i + Y hZ, Xi
2 
− hX, [Y, Z]i + hZ, [X, Y ]i + hY, [Z, X]i .
The connection θ determined by (6.20) is called the Levi-Civita connection
of X. In the sequel, θ (or θg ) will always denote the Levi-Civita connection.
Clifford Modules and Operators of Dirac Type. There are many different
concepts of a Dirac operator in global analysis: classical and twisted Dirac opera-
tors on spin manifolds; operators of Dirac type with a square with scalar principal
symbol; generalized (or compatible) Dirac operators defined by arbitrary (or com-
patible) connections on bundles of Clifford modules over Riemannian manifolds;
full and split (odd-parity) Dirac operators; boundary Dirac operators; etc. The
concepts depend on various geometrical features like dimension parity, orientation
and chirality, almost complex structure, and suitable boundary. Each definition
has its own merits and range of application and we will return to them.
Let X be a compact smooth oriented manifold (with or without boundary)
with Riemannian metric g. Let dim X = n. Let S be a complex vector bundle over
X of Clifford modules; i.e., we have a representationc : C`(X) → Hom(S, S) with
(6.21) c(v)2 = −kvk2 IdSx for v ∈ T Xx and x ∈ X.
Recall that the Clifford bundle C`(X) consists of the Clifford algebras C`(T Xx , gx ),
x ∈ X, which are associative algebras with unit generated by T Xx and subject to
the relation v · w + w · v = −2gx (v, w). For details see Section 17.1 below on p.513ff.
We shall call c left Clifford multiplication and occasionally write
c : C ∞ (X; T X ⊗ S) −→ C ∞ (X; S).
We may assume that S is equipped with a Hermitian metric which makes Clifford
multiplication skew-adjoint, i.e., c(v)∗ = − c(v) for all v ∈ T Xx .
Definition 6.32. A connection ∇S : C ∞ (X; S) −→ C ∞ (X; T ∗ X ⊗ S) for S
will be called compatible with the Clifford module structure of S, if ∇S c = 0,
i.e., ∇S is a module derivation with
(6.22) (∇S c)(v)(s) = ∇S (c(v)s) − c(θg v)s − c(v)(∇S s) = 0 for all s ∈ C ∞ (X; S),
where θg denotes the Levi-Civita connection on X.
Patching locally constructed spin connections together proves
Theorem 6.33 (T. Branson, P. Gilkey [88]). There exist compatible con-
nections on S which extend the Riemannian connection on X to S.
Definition 6.34. Let D : C ∞ (X; S) −→ C ∞ (X; S) be a linear differential
operator of first order operating on smooth sections of a C`(X)-module S.
a) We call D an operator of Dirac type, if it can be written as
(6.23) D = c ◦j ◦ ∇S ,
where ∇S is a (not necessarily compatible) connection and

j : C ∞ (X; T ∗ X ⊗ S) −→ C ∞ (X; T X ⊗ S)
6.6. DIFFERENTIAL OPERATORS ON MANIFOLDS AND SYMBOLS 181

denotes the canonical identification. In terms of a local orthonormal frame v1 , . . . , vn


of T X we then have
Xn
(6.24) Ds|x = c(vν )(∇Svν s)|x .
ν=1

b) We call D a (compatible) Dirac operator, if it can be written as D = c ◦j◦∇S ,


where ∇S is a compatible connection.
Note. The Dolbeault complex (to be studied extensively in Parts III-IV) is
an example of a noncompatible Dirac operator.
As we shall see below in Section 13.8, Equation (13.31), all (total) Dirac oper-
ators are elliptic and formally self-adjoint with a Green’s formula
Z
(6.25) (Ds, s0 )0 − (s, Ds0 )0 = − J(y)hs|Y , s0 |Y i,
Y
where J(y) := c(n) denotes Clifford multiplication by the inward unit tangent
vector over the (possibly empty) boundary Y = ∂X and (·, ·)0 the scalar product
in L2 (X; S)).
For even n the splitting C`(X) = C`+ (X) ⊕ C`− (X) of the Clifford bundles
induces a corresponding splitting of S = S + ⊕ S − and a chiral decomposition
0 D−
 
D= .
D+ 0
The partial (chiral) Dirac operators D± are especially interesting in index the-
ory since they are also elliptic, but in general not self-adjoint and provide interesting
integer-valued invariants as their indices. In this book, we shall come back to Dirac
type operators incessantly.

6. Differential Operators on Manifolds and Symbols


We shall define linear differential operators on differentiable manifolds, acting
on sections of complex vector bundles.
Our Data. Let X be a C ∞ n-manifold. Let πE : E → X be a C ∞ complex
vector bundle over X of fiber-dimension N , i.e., a family of N -dimensional complex
vector spaces Ex , with the parameter x ranging over X, whose disjoint union carries
C ∞ structure in a natural way; see Appendix, Exercise B.13b, p. 719. We denote
the linear space of C ∞ sections of E by C ∞ (X; E) or shortly C ∞ (E) := {s : X →
E : πE ◦ s = idX }. Unless otherwise stated, we remain in the C ∞ category. For
example, let E denote the trivial product bundle X × CN , where we write CN X when
we wish to emphasize the bundle point of view. Then C ∞ (E) denotes the space of
CN -valued C ∞ functions on X. Strictly speaking, when N = 1, we should write
C ∞ (CX ) instead of C ∞ (X), but when the context is clear this is not necessary.
Since manifolds and vector bundles can be locally described in terms of coor-
dinate functions, the following definition makes sense:
Definition 6.35. Let πE : E → X, πF : F → X be two C ∞ complex vector
bundles over X of fiber-dimension N and M . A linear differential operator P
of integer order k ≥ 0 from E to F is a linear mapping P : C ∞ (E) → C ∞ (F ) that
satisfies the following conditions:
(i) For any section s ∈ C ∞ (E), the support of the section P s is contained in the
182 6. DIFFERENTIAL OPERATORS OVER MANIFOLDS

support of s.
(ii) Via local coordinates, the mapping P can be represented as a vectorial differ-
ential operator (see Exercise 5.7, p. 139) where derivatives of order ≤ k appear, but
not of order > k. More precisely, for all coordinate neighborhoods U ⊂ X and
trivializations τE : E|U ∼= U × CN and τF : F |U ∼= U × CM , the mapping P can be
locally expressed in the form
X
P [s](x) = τF−1 aα (x)Dα (τE ◦ s)|x , x ∈ U,

(6.26)
|α|≤k

where α ranges over all multi-indices (α1 , . . . , αn ) ∈ Z+ × . n. . × Z+ with |α| :=


α1 + · · · + αn , aα ∈ C ∞ (X; Hom(CN , CM )), and
∂ |α|
Dα := i−|α| ,
∂x1 · · · ∂xn αn
α1

where x1 , . . . , xn denote the chosen local coordinates on U .


Note (1). The splitting of our definition in two parts is for convenience only.
Actually, condition (i) makes it sufficient to check condition (ii) solely for s with
compact support contained in the coordinate patch U . That would be a weaker
condition. However, to argue along that line excludes analytic differential operators
in case they should be applied only on analytic sections. Mathematically, the
splitting is redundant: (i) can be deduced from (ii) as in Exercise 5.2, p.136. For
the deduction of (ii) from (i), see Remark 5.3, p.136.
Note (2). The reason for inclusion of the factor i−|α| is explained in Remark
5.1, p.135. There we also emphasize why this tradition, originating from mathe-
maticians working in analysis, is a bit unfortunate for topologists, geometers and
physicists when they are interested in differential operators of first order on real
vector bundles.
We write P ∈ Diff k (E, F ) for linear differential operators of order ≤ k. A real
differential operator of order k is defined similarly with C replaced by R.
Symbolic calculus. Fourier analysis makes it natural to replace the differ-
ential expressions by multiplication, e.g., to replace the differentiation Dα by the
monomial ξ α with ξ ∈ Rn . We can do that for all differential expressions in (6.26)
and obtain a polynomial in ξ at each x0 ∈ X. That polynomial depends heavily on
the choice of coordinates. To reduce that dependence, we can define the polyno-
mial in cotangent variables, i.e., choosing ξ ∈ T̊x∗0 X = Tx∗0 X \ {0}. The expression
defined in that way is called the total or complete symbol.
Workers in analysis often use the complete symbol. It is intimately related
to the concepts of quantization, see below. However, roughly speaking, it carries
too much information. Consequently, in general it is not defined independently
of the choice of coordinates. In the geometric tradition of, e.g., classifying conic
sections or more advanced curves and surfaces, it is obvious that one has to select
the relevant information. For index theory, that is the principal symbol, i.e., the
leading term of the total symbol. Fortunately, it will turn out that the definition of
the principal symbol does not depend on the choice of local coordinates. Whence,
it is a genuinely geometric object.
Let π : T̊ ∗ X → X denote the dotted cotangent bundle of X, i.e., the bundle
with the symplectic cone T̊ ∗ X := T ∗ X \ X as total space. Let π ∗ E → T̊ ∗ X
6.6. DIFFERENTIAL OPERATORS ON MANIFOLDS AND SYMBOLS 183

and π ∗ F → T̊ ∗ X denote the pull-backs of the vector bundles πE : E → X and


πF : F → X via π. We then have a vector bundle Hom(π ∗ E, π ∗ F ) → T̊ ∗ X. We
define a section σ(P ) of this bundle for each linear differential operator P , acting
between sections of E and F , i.e., σ ∈ C ∞ (Hom(π ∗ E, π ∗ F )). That is our view of
the principal symbol (similarly, e.g., [434, p.115-116]).
Definition 6.36. For a linear differential operator P : C ∞ (X; E) → C ∞ (X; F )
of order k ≥ 0 we define the symbol (also called the principal or the leading
symbol to emphasize that only the principal or leading terms of P enter into the
definition) via the formula
ik
P ϕk g x ,

(6.27) σ(P )(x, ξx )(e) =
k!
where x ∈ X; ξx ∈ Tx∗ X, ξx 6= 0; e ∈ Ex ; ϕ is a real-valued C ∞ function on X with
dϕx = ξx and ϕ(x) = 0; and g ∈ C ∞ (E) with g(x) = e. (Such choices are always
possible.)
For ξx ∈ Tx∗ X, note that the fiber (π ∗ E)ξx may be identified with Ex and we
will do so; we write the identification as (π ∗ E)ξx ∼ Ex .
Note. Most workers in analysis prefer to define the principal symbol in local
coordinates, i.e., taking the characteristic polynomial when the operator is given
by (6.26), see the following Exercise 6.37b. For the further treatment of elliptic
differential operators, this definition is somewhat opaque, since in that way the
principal symbol is defined in a piecewise unrelated fashion. Then, it may look
like a mystery that the locally given characteristic polynomials collectively define a
bundle homomorphism and that the principal symbol of a differential operator on
a manifold admits a geometric interpretation. Of course, basically it does not make
a big difference whether we introduce the principal symbol without coordinates or
with coordinates, as the following exercise shows. Even in our preferred coordinate-
free first way one has to use coordinates to show that (6.27) in Definition 6.36 is
well defined and yields a smooth section of the homomorphism bundle.
Exercise 6.37. a) Show that the point σ(P )(x, ξx )(e) ∈ Fx in Equation (6.27)
in the preceding Definition is well defined, i.e., it does not depend on the choices of
functions and sections representing cotangent vectors and fiber points. Moreover,
show that the definition yields a smooth section of the homomorphism bundle.
b) Choose a coordinate patch U ⊂ X and local coordinates x1 , . . . , xn on U . Write
ξx = ξ1 dx1 + · · · + ξn dxn with (ξ1 , . . . , ξn ) ∈ Cn , and write the product ξ1α1 · · · ξnαn
as ξ α . For e ∈ (π ∗ E)ξx ∼ Ex , define
 X  
σ(P )(ξx )(e) : = τF−1 x, aα (x) τE (x, e) ξ α ∈ Fx ∼ (π ∗ F )ξx
|α|=k
X  
= τF−1 x, a(α1 ,...,αn ) (x) τE (x, e) ξ1α1 · · · ξnαn ,
|α|=k

where τE , τF are local bundle trivializations like in (6.26). Making identifications


(π ∗ E)ξx ∼ Ex ∼ CN (via τE ) and (π ∗ F )ξx ∼ Fx ∼ CM (via τF ), we can write this
more transparently as
X
(6.28) σ(P )(ξx ) = a(α1 ,...,αn ) (x)ξ1α1 · · · ξnαn ∈ Hom(CN , CM ).
|α|=k
184 6. DIFFERENTIAL OPERATORS OVER MANIFOLDS

Show that in this way σ(P ) ∈ C ∞ Hom(π ∗ E, π ∗ F ) is well-defined, i.e., indepen-




dent of the choice of local coordinates and trivializations τE and τF .


[Hint: You can prove (b) by (a), i.e., writing the mapping in coordinate-free manner;
and you can prove (a) by (b), i.e., checking all choices and transformations. In both
cases, you have to do some calculations at some point. There are no free lunches.
The result is not obvious. Indeed, if we had summed over α with |α| = k − 1, the
resulting so-called subprincipal symbol is not well-defined.]
In this way, we have defined a linear map σ : Diff k (E, F ) → Smblk (E, F ),
where
(6.29) Smblk (E, F ) :=
{σ ∈ Hom(π ∗ E, π ∗ F ) : σ(x, λv) = λk σ(x, v) for all (x, v) ∈ T̊ ∗ X and λ > 0}.
Exercise 6.38. Show that the sequence of vector spaces
j σ
0 → Diff k−1 (E, F ) → Diff k (E, F ) → Smblk (E, F )
is exact, where j denotes the natural inclusion.
[Hint: This is clear from the representation in coordinates; see Exercise 6.37b.
Replacing Smblk (E, F ) by the subspace of polynomial symbols (see [328, p.63]),
one gets a surjective symbol map, and the exact sequence may be extended by zero
on the right.]
Exercise 6.39. Show that for P ∈ Diff k (E, F ) and Q ∈ Diff j (F, G), the
operator QP is in Diff k+j (E, G) with σ(QP ) = σ(Q) ◦ σ(P ).
[Hint: Carry out the proof using the chain rule first for the local vector-valued
differential operators (see p. 139), and then generalize.]
Formal Adjoints. In the following, we assume that the manifold X is com-
pact, oriented, closed (i.e., without boundary) and Riemannian, and that the vector
bundle E is equipped with a Hermitian metric; i.e., each fiber Ex has a nondegen-
erate, conjugate-symmetric bilinear form (··, ··)Ex which is C ∞ in the sense that
(e1 , e2 )E ∈ CR ∞ (X) for any two sections e1 , e2 ∈ C ∞ (E). Whence, we can form
the integral X (e1 , e2 )E , obtaining a Hermitian bilinear form on the vector space
C ∞ (E). Assume that the vector bundle F is also given a Hermitian metric.
Definition 6.40. Two operators P ∈ DiffRk (E, F ) and P R∗ ∈ Diff k (F, E) are
formally adjoint (or formal adjoints), if X (P e, f )F = X (e, P ∗ f )E for all
sections e ∈ C ∞ (E) and f ∈ C ∞ (F ).
Remark 6.41. a) The definition extends to a compact manifold X with smooth
boundary Σ = ∂X. In that case we require the symmetry condition for sections
e, f with support in the interior X̊ := X \ Σ of X.
b) Our definition is not in strict analogy with the theory of bounded operators
in Hilbert space (see Chapter 2), but with the theory of closed (not necessar-
ily bounded) operators in Hilbert space (here the space(es) of L2 -sections). See
our previous Definition 2.38c: For bounded operators the concepts of symmetry
and self-adjointness are equivalent; for closed not necessarily bounded operators in
Hilbert space they are not equivalent, in general. A given symmetric differential
operator with dense domain in L2 may have many different — or even no — self-
adjoint extensions. Happily, for elliptic operators (say differential and of order 1)
on closed manifolds we need not care too much about these distinctions: There is
6.6. DIFFERENTIAL OPERATORS ON MANIFOLDS AND SYMBOLS 185

no difference between minimal and maximal domains. They are always equal to
the first Sobolev space W 1 (X; E), and if the elliptic operators are symmetric with
that domain, they are closed and, in fact, self-adjoint.
Exercise 6.42. Show:
a) There is at most one formally adjoint differential operator P ∗ for a given P .

b) (P + Q)∗ = P ∗ + Q∗ , (P ◦ Q) = Q∗ ◦ P ∗ , and P ∗∗ = P .
Exercise 6.43. Show that for each P ∈ Diff k (E, F ) there is an adjoint differ-
ential operator P ∗ ∈ Diff k (F, E) such that σ(P ∗ ) = σ(P )∗ , where σ(P )∗ : π ∗ (F ) →
π ∗ (E) is the homomorphism pointwise adjoint to σ(P ).
[Hint: 1. Begin with the special case k = 0, where P ∈ Diff 0 (E, F ) is given by a
vector bundle homomorphism h : E → F (i.e., a family of linear maps h : Ex → Fx
parametrized smoothly by x ∈ X). Then P (e)(x) = hx (e(x)) and σ(P )(x, v) = hx ,
where e ∈ C ∞ (E) and v ∈ T̊x X). Let h∗x : Fx → Ex be the linear map adjoint to
hx relative to the Hermitian metrics on Ex and Fx . In this case, for f ∈ C ∞ (F )
(P ∗ f )(x) = h∗x (f (x)) and σ(P ∗ )(x, v) = h∗x .
Thus, the statement is proven for this trivial case.
2. For each χ ∈ C ∞ (T X) define an operator P ∈ Diff 1 (CX , CX ) by
P ϕ := 1i χ [ϕ] = 1i dϕ(χ) ,
where at any point x, χ [ϕ] (x) = dϕx (χ) is the derivative of ϕ in the direction of
χ|x . Then
σ(P )(x, v) = v(χ|x ) , where v ∈ T̊x X = Tx∗ X \ {0} .
Furthermore, by the Stokes Theorem in the classical Green form (e.g., see [196,
p.152 and 182-187] or the Cartan calculus in our Exercise 6.20, p. 172 below, which
we have already used in Exercise 5.9, p. 145 and which also applies here), we have
for all ϕ, ψ ∈ C ∞ (X)
Z  
χ [ϕ] ψ̄ + ϕdiv(ψχ) = i 1i χ[ϕ], ψ 0 − i ϕ, 1i div(ψχ) 0
 
0 =
X
where div(ψχ) ∈ C ∞ (X) denotes the divergence of the vector field ψχ and (·, ·)0
the scalar product in L2 (X, C). Thus, P ∗ ψ = 1i div(ψχ). One further checks that
(σ(P ∗ )(x, v))(zx ) = (σ(P )(x, v))(zx ) = v(χ|x ) zx , zx ∈ (CX )x .
Since v(χ|x ) is real and hence self-adjoint as a linear map from C to C, we have
σ(P ∗ ) = σ(P )∗ .
3. One may now show that every global differential operator can be constructed
from the two preceding types via sums and compositions (locally, this is entirely
trivial), and thus Exercise 6.43 reduces to Exercise 6.42.]
Remark 6.44. In contrast to Exercise 6.42, the solution of the preceding Exer-
cise 6.42 is not so trivial, even though we only applied Stokes’ theorem in the weak
form. Alternatively, one can first assign to each vector-valued differential operator
P : C ∞ (CU ) → C ∞ (CU ) given (over an open set U ⊂ Rn ) by
X
Pu = aα (x) Dα u
|α|≤k

the operator P ∗ : C ∞ (CU ) → C ∞ (CU ) given by


X
P ∗ v := Dα (a∗α v),
|α|≤k
186 6. DIFFERENTIAL OPERATORS OVER MANIFOLDS

where a∗α (x) is the adjoint (conjugate transpose) of the N × N matrix aα (x). Using
integration-by-parts, it then follows at once that
Z Z
(P u, v) = (u, P ∗ v) for all u, v ∈ C0∞ (CU ) , where
U U

C0 (CU ) := {w ∈ C ∞ (CU ) : supp w is a compact subset of U } ,
and where (·, ·) is the canonical Hermitian scalar product on CN . The major work
consists of globalizing this result; see [328, p.70-75], [314, p.181-183], or [434, p.117
f].
Elliptic Differential Operators. Definition and Standard Examples.
I Geometrically defined linear differential operators on closed manifolds (i.e., opera-
tors of Laplace type and operators of Dirac type, see below) are marked by two features,
namely the algebraic symmetry and regularity of their expression and the finite number
of linearly independent solutions. Geometers have always noticed these two features and
exploited them. They have been pleased with the ease and transparency of manipula-
tions, and were enthusiastic when recognizing geometric or topological invariants in the
dimensions of the solution spaces.
As seen from analysis, these two features are interrelated: algebraic regularity of the
principal symbol of a differential operator over a closed manifold implies that the dimen-
sion of the kernel of the operator is finite. For that, the key notion is the ellipticity of
the principal symbol. That notion will be explained now. In Chapter 9, we deduce the
regularity (=smoothness) of the solutions and the Fredholm properties from the elliptic-
ity. Then in Part III, we prove the Atiyah-Singer Index Theorem for elliptic operators
on closed manifolds. Roughly speaking, it gives a thorough explanation for the astonish-
ing and previously perceived of as somewhat mysterious interrelations between algebraic
symmetries of a formal expression (the principal symbol of a variety of geometric defined
elliptic operators) and geometric features of the underlying manifold. In Part IV, much
wider implications are drawn for low-dimensional topology and gauge-theoretic physics of
the same philosophy, namely exploiting symmetries and regularities of formal expressions
for sensing asymmetries and irregularities of related geometric or physical objects. J

Definition 6.45. Let P : C ∞ (E) → C ∞ (F ) be a differential operator with


principal symbol σ(P ) : T ∗ (X) → Hom(π ∗ E, π ∗ F ). If σ(P )(ξx ) is an isomorphism
for all nonzero ξx ∈ Tx∗ X, then P is called an elliptic differential operator.
Whence, a differential operator P is elliptic, if in each chart all of the locally-
defined vectorial differential operators are elliptic (see Exercise 5.7); i.e., for each
local representation of P over the chart domain U , the characteristic polynomial
of the principal part (associated to the terms of highest order, i.e., what we call
the principal symbol)
X
(6.30) pk (x, ξ) := aα (x) ξ α
|α|=k

is an invertible linear map for all x in U and all ξ ∈ Rn \ {0}.


Exercise 6.46. Show P elliptic ⇒ P ∗ elliptic.
We consider some standard examples. Even though C ∞ (E) and C ∞ (F ) are
not Hilbert spaces (and hence P is not Fredholm), here we take the index of P to
be dim Ker P − dim Ker P ∗ , where P ∗ denotes the formal adjoint of P . This is the
same as the usual index of a Fredholm extension of P to a suitable Sobolev space,
as is explained in Chapters 7 and 9 below.
6.6. DIFFERENTIAL OPERATORS ON MANIFOLDS AND SYMBOLS 187

Table 6.1. The principal symbol of elliptic standard operators

Operator P Domain σ(P )(x, ξ)


ODE system P := ∇ + A,
1 C ∞ (R, C) iξ Idr
A ∈ C ∞ (R, gl(r, C)), ∇ = dt d
Idr
Laplace operator
2 C ∞ (Rn , C) −(ξ12 + · · · + ξn2 )
P := ∆ = ∂x21 + · · · + ∂x2n
Cauchy-Riemann operator
3 C ∞ (R2 , C) 1
2 i(ξ1 + iξ2 )
P := ∂∂z̄ = 12 ( ∂x
∂ ∂
+ i ∂y )
4 Euler operator P := d + δ Ω• (X), Ωev (X) i(ξ∧ − ξx )
5 Dirac type operator P := D = c ◦∇S C ∞ (X; S) c(iξ)
Dirac Laplacian P := D2
6 C ∞ (X; S) −kξk2g(x) IdSx
Riemannian metric g

Exercise 6.47. Check the principal symbols of Table 6.1 and derive the ellip-
ticity of all the listed standard operators.
[Hint to (1): This operator (and its counterpart on S 1 for periodic matrices A) was
studied in Exercise 2.42, p.41f. Its principal symbol was calculated on p.42 (Note).
∂ ∞
To (2): Let ∂xi be a shorthand notation for ∂x i . Regard the domain C (Rn , C) as
n n
the space of sections of the trivial bundle R × C → R and write the spaces down
where the principal symbol of the Laplacian ∆ = ∂x21 + · · · + ∂x2n acts. Do it slowly
with the previous notations: Then you have
σ(∆) : T̊ ∗ Rn Hom π ∗ (Rn × C), π ∗ (Rn × C)

−→
(x, ξ1 dx1 + · · · + ξn dxn ) 7→ −(ξ12 + · · · + ξn2 ) ∈ End(π ∗ (Rn × C)x,ξ ) ,
regarded as multiplication on π ∗ (Rn × C)x,ξ ∼
= C by a real.
To (3): The Cauchy-Riemann operator is a first-order operator on the same space
of sections as in Example (2), but with a different symbol
∂ 1
σ( )(x, ξ1 dx1 + ξ2 dx2 ) = i(ξ1 + iξ2 ),
∂ z̄ 2
regarded as complex multiplication on π (R × C)x,ξ ∼
∗ n
= C. Similarly, you have the
2
complex differentiation operator ∂∂z̄ = 12 ( ∂x ∂ ∂
− i ∂y ). Note that ∆ = 4 ∂∂z̄∂z and at
ξ1 dx1 + ξ2 dx2 you confirm
1 1
σ(∆)(x, ξ1 dx1 + ξ2 dx2 ) = 4 i(ξ1 + iξ2 ) i(ξ1 − iξ2 ) = −(ξ12 + ξ22 ).
2 2
There is a type of exterior derivative on Ω0,0 (R2 , C) := C ∞ (R2 , C), namely the
Dolbeault operator
∂¯ : Ω0,0 (R2 , C) −→ Ω0,1 (R2 , C) := {hdz̄ : h ∈ C ∞ (R2 , C)},
∂f .
f 7→ ∂ z̄ dz̄

Here Ω0,1 (R2 , C) is called the space of complex forms of type (0, 1). For a compact
Riemann surface S one can define a strictly analogous operator ∂¯ : Ω0,0 (S, C) −→
Ω0,1 (S, C). Its index yields the classical Riemann-Roch Theorem. For any higher-
dimensional compact, complex manifold X of dimC X = m, there is a Dolbeault
operator complex ∂¯ : Ω0,k (X, C) → Ω0,k+1 (X, C), k = 0, 1, . . . m which can be rolled
188 6. DIFFERENTIAL OPERATORS OVER MANIFOLDS

up to give an elliptic operator


m
M m
M
∂¯ : Ω0,k (X, C) → Ω0,k (X, C).
k even k odd

The Hirzebruch-Riemann-Roch Theorem expresses index ∂, ¯ called the arithmetic


genus of X in terms of so-called Chern numbers. We shall explain all that below in
Section 13.7, p.333ff, and present a comprehensive generalization in Section 17.6,
see in particular p.615ff.
To (4): Recall from Section 6.4
• the definition of the bundle Λk (X) → X of complex exterior k-covectors
over the compact, orientable C ∞ Riemannian n-manifold X with metric
tensor g;
• the wedge product ∧, the interior product x and the (Hodge) star operator
∗ on Λ• (X);
• the space Ωk (X) := C ∞ (Λ(X)) of C ∞ sections of Λk (X), namely the
space of C-valued k-forms on X;
• the exterior derivative d : Ωk (X) → Ωk+1 (X) and the codifferential
(6.31a) δ : Ωk+1 (X) −→ Ωk (X) which is the formal adjoint of d; i.e.,
Z Z
(6.31b) (dα, β)0 = hdα, βig νg = hα, δβig νg = (α, δβ)0 ,
X X

where νg denotes the volume form, h·, ·ig the scalar product on the fibers of
the bundle Λ• (X) induced by g and (·, ·)0 the scalar product in L2 (X, Λ• (X));
and
• set again Ω• (X) = ⊕nk=0 Ωk (X).
Next, you bring δ in a manageable form. Try (here just for fun, but later highly
usable)
(6.32) δ = −(−1)nk ∗n−k d ∗k+1 : Ωk+1 (X) → Ωk (X).
Then you are ready to calculate the principal symbol of the first order operator
d + δ : Ω• (X) → Ω• (X). One way to do it would be the following exercise in the
manipulation of exterior forms (do it):
Choose a local system of coordinates x = (x1 , . . . , xn ) on X. Then the array
{dx , . . . , dxn } is the corresponding local frame for T ∗ X and {dxα := dxα1 ∧ · · · ∧
1

dxαk }α∈J a local frame for the bundle Λ• (X), where J denotes the set of all coor-
dinate selections α = (α1 , . . . , αk ) with k ≤ n and 1 ≤ α1P < · · · < αk ≤ n). For
given x ∈ X and e ∈ Λ• (X)x expand e = ω(x) with ω = α∈J fα dxα close to x
with smooth functions {fα }α∈J . Check that
n
X  XX ∂
dω = d fα dxα = f dxj ∧ dxα .
j α
j=1
∂x
α∈J α∈J

For ξ = ξ1 dx1 +· · ·+ξn dxn you will find σ(d)(x, ξ)(e) = iξ ∧e. Then σ(δ)(x, ξ)(e) =
−iξxe since δ = d∗ and the principal symbol of the adjoint operator is the adjoint
of the principal symbol (as you have shown in Exercise 6.43, p.185). Check that the
interior product ξx is the dual of the exterior product ξ∧ by applying the mappings
to an orthonormal local frame (we shall do it for you in our arguments for Theorem
6.6. DIFFERENTIAL OPERATORS ON MANIFOLDS AND SYMBOLS 189

13.6a (p.321). In this way you obtain

T̊ ∗ X Hom π ∗ Λ• (X), π ∗ Λ• (X)



σ(d + δ) : −→
(x, ξ) 7→ (π ∗ Λ• (X))(x,ξ) 3 e 7→ i(ξ ∧ e − ξxe).

Note that σ(d+δ)ξ is invertible (ξ 6= 0), and hence d+δ is elliptic. To see that, please
check σ(d + δ)ξ ◦ σ(d + δ)ξ = −kξk2 Idπ∗ Λ• (X)ξ = σ((d + δ)2 )ξ . You can find the
details below in Section 13.4, p.319f. The second order operator (d + δ)2 = dδ + δd
is called the Beltrami Laplacian. Since d + δ is formally self-adjoint, its index is
zero. By definition, elements of Ker(d + δ) are known as harmonic forms. Hodge
theory tells us that the algebra of harmonic forms, say H• (X), is isomorphic to the
cohomology algebra H ∗ (X; C) where wedge product of harmonic forms corresponds
to cup product in H ∗ (x; C) (see [186, p.43f and p.59-61] for a quick orientation).
To get something more interesting, you must restrict d + δ to the even or to the
odd forms. The restricted operators are still elliptic with symbols that are inverses
modulo a factor of −kξk2 . Their index is not necessarily zero. Indeed, restricting
to the even forms yields an operator with the Euler characteristic of X as index.
That will be explained in the mentioned Section 13.4.
To (5): The claim follows immediately from the definition of Dirac type operators
(see Definition 6.34, p.180 and the elementary properties of Clifford multiplication.
You can use either the global definition of (6.23) or the local definition of (6.24).
Ellipticity follows at once. Also it follows that the principal symbol is symmetric,
if the defining connection is metric (i.e., compatible with the metric structure of
the bundle, as explained in Definition 6.29, p.179).
Actually, the standard operators ∇ + A, ∂¯ and d + δ of (1), (3) and (4) are all
special instances of Dirac type operators. You could have obtained your results for
these operators by clever Clifford multiplication alone! Indeed, take for instance
S := Λ• (T ∗ X) and let c be the extension of c(ξ)(s) = ξ ∧ s − ξxs for ξ ∈ Tx∗ X ⊂
C`(Tx∗ X, gx ) and s ∈ Λ• (Tx∗ X) (here we identify T X and T ∗ X). The extension is
2
guaranteed by the fact that c(ξ) = −gx (ξ, ξ) Id.
To (6): Apply Exercise 6.39 and simple Clifford multiplication. There is a zoo of
Laplacians, depending on the assumptions about the underlying manifold X and
the bundle S and the connection ∇S . Special cases (the connection Laplacian and
the Hodge Laplacian) are discussed in our Section 15.6, p.434ff.]

I What is the meaning of the geometric description of the symbol of a differential,


and in particular elliptic operator? This is the basic question which will concern us from
now on. Recall from Remark 5.3 the geometric characterization of differential operators
as support compressing. Unfortunately, that property can not be quantified or graded.
Conversely, the symbol can express both quantitative and qualitative aspects. That as a
first vague answer.
As a second vague answer, imagine an electron microscope aimed at a point x of
our manifold. Under enlargement, both the neighborhood of x and the vector bundle
above it become linear, of course, while the hole in the cotangent bundle becomes large
like a sphere. With even greater magnification, we see, instead of a highly complicated
differential operator on infinite dimensional spaces of cross-sections (the object of study
of functional analysis), families of maps on spheres S n−1 into the general linear group
GL(N, C) (the object of study of linear algebra and of topology). The Atiyah-Singer
Index Formula is (crudely) a manifestation of this change of the levels of investigation.
190 6. DIFFERENTIAL OPERATORS OVER MANIFOLDS

The term symbol suggests the symbolic method of Oliver Heaviside, which owes its
power to the transition between the plane of operator theory and the plane of polyno-
mial algebra. For this reason it is characterized in [116, II, p.187/518] pointedly as the
separation of the algebraic part from the mathematical-conceptual part.
Initially, the relationship between the study of operators and the study of their sym-
bols was penetrated more deeply not by mathematicians, but by physicists with the ideas
of Erwin Schrödinger and Paul Adrien Maurice Dirac on the quantization of clas-
sical mechanical systems according to which classical mechanics deals with symbols and
quantum mechanics with operators (for quantum mechanical systems with spin, these are
differential operators for nontrivial vector bundles). Thus first one considers a problem of
classical physics (mechanics, electrodynamics), establishes the classical Hamiltonian func-
tion, changes to position and momentum coordinates, and obtains, e.g., for the harmonic
oscillator the function
1 2 k 2
h(x, p) = p + x .
2m 2
Interpretation: Consider on the real axis the motion of a particle of mass m with the
kinetic energy m 2
ẋ2 = 2m
1
p2 where x(t) is the location at time t and p = mẋ is the
momentum. Then one supposes that the particle moves in a force field whose potential
energy is k2 x2 . According to the quantization rule, we now choose a suitable Hilbert space,
as a rule an L2 (or Sobolev space as in Chapter 7 below), replace the position coordinate
d
x by multiplication by x and the momentum coordinate by the differential operator }i dx .
We then obtain (in the Schrödinger representation with } denoting the reduced Planck
constant) the operator H with
}2 d2 f k
Hf = − + x2 f,
2m dx2 2
with the 2-symbol
1 2
σ2 (H)(x, ξ) = ξ .
2m
The total symbol (which is not invariant and therefore not defined here) is
1 2 k 2
σ2,1,0 (H)(x, ξ) = ξ + x .
2m 2
So much for this simple example; see e.g. [203, II, p.257-293], who strongly advocates
this point of view. We shall return to this topic in our Part IV. J

7. Manifolds with Boundary


Instead of modeling a manifold locally on open subsets of E, we can also work
with charts that map open subsets of the topological space X homeomorphically
onto open subsets of the half-space Rn+ := {(x1 , ..., xn ) ∈ Rn : xn ≥ 0} such that
the coordinate changes are again C ∞ . In this way, we introduce the concept of
a C ∞ manifold with boundary, in the same way as we have done above for
(unbounded) manifolds. We call x ∈ X an interior point, if it has a neighborhood
which is mapped by a chart onto an open subset of Rn (i.e., contained in the
interior of Rn+ ). On the other hand, if there is a chart mapping x to a point on
the boundary of Rn+ , then x is called a boundary point of X; we write ∂X for
the set of all boundary points. As examples, the closed solid sphere or torus are
three-dimensional C ∞ manifolds with boundaries being the 2-sphere or 2-torus,
respectively.
Exercise 6.48. Let X be a C ∞ manifold with boundary.
a) Carry the concepts C ∞ (X), T X, T ∗ X, orientation, Riemannian metric, etc.,
over to this case.
6.7. MANIFOLDS WITH BOUNDARY 191

Y TyY

Figure 6.4. The X


inner normal y
field over the
boundary ∂X of º(y)
a Riemannian
manifold X

Y1 Y2
X1 X2
U1 U2
y f(y )

Figure 6.5. Gluing two manifolds X1 , X2 along their boundaries

b) Construct a C ∞ atlas for ∂X from a C ∞ atlas for X, showing ∂X is a C ∞


manifold of dimension n − 1, when dim X = n. Show that ∂∂X = ∅ and that ∂X
inherits Riemannian structure and orientation from those on X.
For short, we write Y = ∂X here. Since each C ∞ path in Y is a C ∞ path in X,
we have a canonical embedding of T Y into T X|Y . Over Y , we have the following
diagram of tangent and cotangent bundles
TY ∼
= T ∗Y
∩ ∩
(T X) |Y ∼= (T ∗ X) |Y .
The left inclusion is canonical; the other isomorphisms and the right inclusion
depend on the choice of a Riemannian metric.
Exercise 6.49. With the help of a Riemannian metric (·, ·), define a normal
field ν ∈ C ∞ (T X|Y ) such that (ν(y), w) = 0 and (ν(y), ν(y)) = 1 for all y ∈ Y and
w ∈ Ty Y , as depicted in Figure 6.4. Show that there are two such normal fields,
and characterize the inner one via the condition dϕ(ν(y)) ≥ 0 for all real-valued
ϕ ∈ C ∞ (X) which are positive except at y. Characterize the dual normal field
ν ∗ ∈ C ∞ ((T ∗ X)|Y ).

Exercise 6.50. Let X1 and X2 be C ∞ manifolds with boundaries, and let


f : Y1 → Y2 be a diffeomorphism of their boundaries. Show that one can construct
a C ∞ manifold X1 ∪f X2 in a canonical way by identifying the boundaries of X1
and X2 via f as in Figure 6.5.
[Hint: Form the disjoint union of X1 − Y1 , X2 − Y2 , and {(y, f (y)) : y ∈ Y1 }. It
is clear what a coordinate neighborhood of x ∈ Xi − Yi (i = 1, 2) should be. For
x = (y, f (y)), choose a neighborhood U1 of y in X1 , neighborhood U2 of f (y) in X2
with f (U1 ∩ Y1 ) = U2 ∩ Y2 . Then U1 and U2 form a neighborhood of (y, f (y)), see
192 6. DIFFERENTIAL OPERATORS OVER MANIFOLDS

X1 X2 u

Figure 6.6. Defining u ±v {1


U1 U2
an atlas for Rn+
X1 ∪f X2 from v
charts for X1 , X2

Figure 6.6. By this recipe, define an atlas for X1 ∪f X2 from the C ∞ atlases for X1
and X2 , such that X1 ∪f X2 becomes a topological manifold. It is not entirely easy
to prove that the coordinate changes are C ∞ . Without loss of generality, assume
X1 = X2 =: X and f = Id; in our applications, we always have this situation. In
order to avoid the difficulties with the corners originating from the way in which
the charts are joined, choose a Riemannian metric and extend the above-mentioned
normal vector field (Exercise 6.49) to a neighborhood of ∂X := Y . The integral
curves of the vector field provide a diffeomorphism (collar) of Y × [0, 1] with a
neighborhood of Y in X. The differentiable doubling of Y × [0, 1] along Y × {0} is
trivial; see also [97, 13.5-13.11].]

Remark 6.51. The definition of differential operators on manifolds with bound-


ary does not require any modification in relation to the case discussed above. The
following difference is essential, however: While every differential operator over a
closed (i.e., compact, without boundary) manifold has a formal adjoint by Exercise
6.43, one always has an extra term for manifolds with boundary
Z Z Z

(P e, f ) − (e, P f ) = (· · · )
X X ∂X
involving a differential operator over the boundary; e.g., see our discussion above on
the Sturm-Liouville boundary-value problem (Section 2.5), or more generally [328,
p.73-75] and [83, Proposition 3.4]. Thus, one of the advantages of the computa-
tions on compact manifolds without boundary is the existence of formal adjoints.
In passing to the boundary-value problems of interest in applications, Exercise 6.50
comes into play. More generally, in Section 13.8, we will assign to each boundary-
value problem over X associated operators over the two closed manifolds ∂X and
X ∪∂X X. Incidentally, these are manifolds which can be rather complicated topo-
logically, even in the classical case when X is a bounded domain in Rn . (See the
Heegaard diagrams, whereby each three-dimensional oriented manifold can be rep-
resented as such a doubling with boundary diffeomorphism, not necessarily the
identity [387, p.219]).
CHAPTER 7

Sobolev Spaces (Crash Course)

Synopsis. Motivation. Equivalence of Different Local Definitions. Various Isome-


tries. Global, Coordinate-Free Definition. Embedding Theorems: Dense Subspaces; Trun-
cation and Mollification; Differential Embedding; Rellich Compact Embedding. Sobolev
Spaces Over Half Spaces. Trace Theorem. Case Studies: Euclidean Space and Torus;
Counterexamples.

1. Motivation

I In this book, the concept of Sobolev spaces will be used to solve two different
problems. The first question is: How can we fit the analytic concept of (elliptic)
textbfdifferential operators into the framework of functional analysis? That question is
dealt with extensively in many textbooks of modern analysis. One of the answers to that
question is that Sobolev spaces permit a transition from Banach spaces (natural domains
of differential operators) to easy Hilbert spaces and provide a link to the theory of bounded
Fredholm operators in Hilbert space, as developed above in our Part I. We shall explain all
the needed definitions and results in the sequel. That is the easy question and we can be
short in dealing with it. The second question is much deeper: How can we fit the geometric
concept of a connection into the framework of functional analysis?. Recall from Section
6.5, Definition 6.34 and our Table 6.1 that all geometric standard operators are related to
connections, namely as operators of Dirac type, respectively, Dirac Laplacians (= squares).
We will show the fundamental role of connections in gauge theoretic physics and low-
dimensional topology in our Part IV. Then the key problem is to develop an all embracing
view of the space of all suitable connections in a given geometric or physics context. Such
a view is provided by the concept of a manifold. To establish the manifold character
of spaces of connections, we need linearization and parametrization tools. The most
important are infinite-dimensional analogues to the Implicit Function Theorem (IFT).
There we shall use Sobolev spaces once again. But then, our point will not be the rather
trivial aspect of Sobolev spaces, cultivated in this and the following sections and perhaps
overemphasized in analysis main stream literature (namely, the simple replacement of
Banach space theory by Hilbert space theory). In Sections 16.4, 18.3 and 18.4, our point
is rather the replacement of Fréchet spaces (where we have no IFT) by Banach spaces, see
Definition 16.19, p.497.
Like we have two different motivations for the introduction of Sobolev spaces, we
also have two different ways of doing it. As so often in global analysis, we have the
heritage from classical analysis with its proficiency in making calculations in coordinates.
In this section, we shall follow that tradition, making easier reading for a student — or a
teacher — who feels safer with coordinates, does not care so much about global geometric
meaning and accepts the arbitrary and tiresome coordinate shifts. In Section 16.4, p.497
we present an alternative, namely the natural (coordinate-free) Definition 16.19 of much
more delicate families of Sobolev Banach spaces. The global approach is mandatory in
Part IV but would also be quite appropriate in this section for readers who seek a really

193
194 7. SOBOLEV SPACES (CRASH COURSE)

simple, coordinate-free presentation and are not afraid of abstract geometric concepts, see
Remark 7.8 below. J

Let us leave the connections for later and return to our first question regard-
ing differential operators and their relations to Hilbert spaces. There is first of all
the L2 concept of a Lebesgue measurable square integrable function, which can be
transferred naturally to sections in a Hermitian vector bundle E on a Riemannian
manifold X: A section
Z u : X → E (not necessarily continuous) represents an ele-
ment of L2 (E) if hu, ui νg < ∞. Here h·, ·i is a Hermitian metric for the vector
X
bundle E, whence hu, ui is an R-valued function on X which is integrated with re-
spect to the volume element νg defined by the metric tensor g of X (see Section 6.3
above). In this way, L2 (E) becomes a Hilbert space with the usual identification
of sections that differ on a set of measure 0.
The traditional way of fitting a differential operator P ∈ Diff k (E, F ) into the
well-understood and powerful Hilbert space theory consists in considering P as a
map on L2 (E) to L2 (F ), restricted to functions with sufficient differentiability. We
proceeded this way above, when introducing the notion of formally adjoint operators
(Section 6.6). But even restricted to these subspaces, the differential operators are
not continuous in the norm topology of L2 . A simple example is the operator d/dt
1
which maps the sequence sin nt (converging to 0 in L2 (S 1 )) to the sequence cos nt
n
which does not converge in L2 (S 1 ). This circumstance leads to the extensive field
of classical mathematical research on unbounded linear operators. We gave a taste
in Section 2.6.
Thanks to Sergey Lvovich Sobolev, we now have a more potent tool for
the definition of Hilbert spaces with various differentiability properties, namely the
Sobolev spaces W s (X). More common notations for these spaces are H s (X) or
Ls2 (X). We prefer the W s of the Russian literature where the “W” reminds of weak
solutions, namely in distributional sense, to be explained below in Exercise 7.7a. In
a topology book like ours, H s (X) is reserved for cohomology. The notation Ls2 (X)
would be correct in emphasizing that we model the Sobolev spaces after the Hilbert
space L2 (X). However, it is a bit heavy and, moreover, we have so many spaces of
linear mappings, carrying an L. So we had better stick to W s (X).
These spaces have gained great importance in the theory of partial differential
equations, especially for existence questions where precise statements on the reg-
ularity of solutions are desired, but often cannot be expressed in the language of
C k -Banach spaces. Since introducing Sobolev spaces by coordinates is an estab-
lished part of now classical analysis, we may keep it short and refer to the abundant
textbook literature: the classic [6], [58, Chs. III/IV], [217, p.33-63], [280, p.1-118],
[314, p.184-200], [328, p.125-174], and [452, p.55 and 173f], and the more recent
[187, p.44ff] and [190, Sections 4.2, 6.2, 6.3 and 8.2].

2. Definition
Customary Definitions in Coordinates. In the following, we put together
some of the various customary equivalent definitions of Sobolev spaces. Here, we
restrict ourselves to the case of functions and where s ≥ 0. In the framework of
distribution theory, the spaces W s can be treated clearly and uniformly also for
s < 0; e.g., see the carefully written [187] and [190].
7.2. DEFINITION 195

The basic idea of Sobolev spaces is very simple: As explained above, we wish
a common functional analytic frame for linear differential operators. Let us try:
Definition 7.1. Let m ≥ 0 be an integer. We define the Sobolev space
W m (Rn ) as the intersection of the maximal domains of all elementary formally
self-adjoint differential operators
|α| ∂ α1 αn
Dα := (−i) ∂xα1

· · · ∂x α
, |α| := α1 + · · · + αn ≤ m,
n

of order ≤ m, i.e., to consist of all u ∈ L2 (Rn ) such that Dα u ∈ L2 (Rn ) for all
multiindices α with |α| ≤ m.
Recall that for u ∈ L2 (Rn ), we mean by Dα u ∈ L2 (Rn ) that there exists a
v ∈ L2 such that the distribution Dα u acts like v on all test functions w, i.e.,
Z

Dα u (w) := hu, (Dα ) wi0 = hu, Dα wi0 :=

u(x) Dα w(x) dx
Rn
Z
!
= v(x) w(x) dx = hv, wi0 for all w ∈ C0∞ (Rn ).
Rn
By Fourier analysis (Differentiation-Multiplication Conversion and Plancherel For-
mula, Exercises A.5b,d of Appendix A), this is equivalent to requiring ξ α u
b(ξ) ∈ L2
for |α| ≤ m, or, which is the same, (1 + |ξ|)m u
b ∈ L2 . The following exercise makes
you familiar with the arguments.
Exercise 7.2. Show that, for u ∈ C0∞ (Rn ), the two following norms | · |m and
k·km are equivalent (m ∈ N):
a)
X 1/2 Z
2 2
|u|m := |Dα u|0 , where |u|0 := hu, ui0 = u(x) u(x) dx.
|α|≤m Rn
 
∂2 ∂2
b) For ∆ = − ∂x21
+ ··· + ∂x2n ,
Z 1/2 Z 1/2
2 m 2 m
kukm := (1 + |ξ| ) |b
u(ξ)| dξ = h(1 + ∆) u, ui dx ,
Rn Rn
where the last equality is due to Exercise A.5(b,d), p. 711 in Appendix A.
[Hint: Start with the Fourier differentiation formula (see Appendix A), giving
Z X 
2 α 2 2
|u|m := (ξ ) |b u(ξ)| dξ.
Rn |α|≤m

Then prove that for some constant c,


 m X  m
2 2 2
1 + |ξ| ≤ (ξ α ) ≤ c 1 + |ξ|
|α|≤m

and deduce kukm ≤ |u|m ≤ c kukm .]
This leads to the following more general definition.
Exercise 7.3. For real nonnegative s, define the Sobolev space
n o
2
(Rn ) := u ∈ L2 (Rn ) : ξ 7→ (1 + |ξ| )s/2 u
b(ξ) ∈ L2 (Rn )
and show:
a) For all s ∈ N, the space W s (R) is the completion of C0∞ (Rn ) relative to the
196 7. SOBOLEV SPACES (CRASH COURSE)

(equivalent, by Exercise 7.2) s-norms |·|s and k·ks , and whence a Banach space.
b) By considering scalar products which induce the respective norms, W s (Rn )
becomes a Hilbert space.
c) The following inclusions are defined in a natural way, and are continuous and
dense
\∞
C0∞ (Rn ) ⊂ W ∞ := W s ⊂ . . . ⊂ W s+t ⊂ . . . ⊂ W s ⊂ . . . ⊂ W 0 := L2 (Rn ),
s=0
where W s := W s (Rn ) for short. See also Theorem 7.13, p. 200.
[Hint: For a: Investigate Cauchy sequences in C0∞ (Rn ) relative to |·|s .
For b: For natural s, this is clear by (a). For arbitrary real s, see the classical [217,
p.37 and 45f] or [280, p.35-37] or the more recent [187] and[190].
 s
2
For c: For the inclusions, note the monotonicity of 1 + |ξ| in s. For the proof
that C0∞ (Rn ) is dense in W s (Rn ), note that
• C0∞ (Rn ) is dense in the Schwartz space C↓∞ (Rn ) of rapidly decreasing
functions;
• the Fourier transform is an isometric isomorphism
s 
F : W s (Rn ) −→ L2 Rn , 1 + |ξ|2 dξ ; and
• F −1 C↓∞ (Rn ) = C↓∞ (Rn ).

s 
Since C↓∞ (Rn ) is clearly dense in L2 Rn , 1 + |ξ|2 dξ , it follows that C↓∞ (Rn ) is
dense in W s (Rn ), and you are done.
s
You may try another more direct proof of the density: Note that the space WK (Rn )
n s n
of functions with support in compact K ⊂ R is dense in W (R ). So, given
u ∈ WK s
(Rn ), how can you construct a sequence uν ∈ C0∞ (Rn ) with supp uν ⊂ K e
(another compact subset with K ⊂ Int K), e such that kuν − uks → 0 for ν → ∞?
Exploit the convolution (see Definition A.4c, p.710): You can choose compactly
supported standard
R test function ϕ ∈ C0∞ (Rn ), real valued with ϕ ≥ 0, supp ϕ ⊂
{|x| < 1} and Rn ϕ(x)dx = 1. Then put
ϕε (x) := ε−n ϕ(x/ε) and uε (x) := (ϕε ∗ u)(x) = hu(y), ϕε (x − y)i0 , ε > 0.

Note that supp ϕε ⊂ {|x| < ε} and ϕε (x)dx = 1. To show that uε ∈ C0 (Rn ) is an
R

easy exercise in the differentiability of integrals with parameters. The interesting


part is to prove that limε→0+ kuε − uks = 0. For that, work with the Fourier trans-
form: Since ϕ bε (ξ) = ϕ(εξ)
b and ϕ(0)
b = 1, the question is reduced to establishing
the relation
Z
− 1|2 |b
u(ξ)|2 ξ12s + · · · + ξn2s dξ = 0,

lim |ϕ(εξ)
b
ε→0+

which is evident from the dominated convergence theorem.]



Remark 7.4. a) The family ϕε 0<ε<1 (more precisely the family of linear

scalar operators Jε 0<ε<1 with Jε (·) := ϕε ∗ ·, mapping integrable functions to
smooth functions) is called a mollifying family because Jε smoothes out the as-
perities and we have Jε → Id in the appropriate norm.
b) There is a second basic technique for dealing with Sobolev spaces, namely trun-
cation. The essentials are contained in the following result: Let u ∈ W k (Rn ).
Consider for each R > 0 a smooth bump (cut-off) function χR ∈ C0∞ (Rn ) with
χR (x) ≡ 1 for |x| ≤ R, χR (x) ≡ 0 for |x| ≥ R + 1, but sufficiently moderate, say
7.2. DEFINITION 197

with differential |dχR (x)| ≤ 2 for all x ∈ Rn . Then χr · u ∈ W k (Rn ) for all R > 0
and, moreover,
Wk
χR · u −→ f, as R −→ ∞.
We leave the proof to the reader (or else see [321, Lemma 9.2.9]).
Exercise 7.5. Show that for s ∈ R, the formally self-adjoint operator (in fact,
a pseudo-differential operator, see the following chapter)
Z  s/2
2
(Λs u)(x) := eihx,ξi 1 + |ξ| u
b(ξ) d̄ξ,
Rn

with d̄x = (2π)−n/2 dx1 · · · dxn , defines an isomorphism (in particular, isometry)
Λs : W t+s (Rn ) → W t (Rn ) of Hilbert spaces for t ≥ 0 and t + s ≥ 0.
[Hint: Note that Parseval’s Formula (Appendix A, Exercise A.5d, p. 711) implies
the equality kuks = kΛs uk0 for u ∈ W s (Rn ), and that the family {Λs : s ∈ R}
forms a group since Λs ◦ Λr = Λs+r .]
1
Remark 7.6. A common abbreviation of the expression 1 + |ξ|2 2 is the sym-
bol hξi. Whence, in the notation of pseudo-differential operators of Chapter 8, we
can write Λs = Op(hξis ).
Global and Coordinate-Free Definitions. From the preceding presenta-
tion of Sobolev spaces in local Euclidean coordinates, the reader can catch the basic
idea, namely that Sobolev spaces are closures of spaces of differentiable functions
with regard to the L2 -norms of the highest derivative. However, it seems to us
that the power of the concept of Sobolev spaces becomes clearer in global and
coordinate-free presentation. We shall give several choices.
Exercise 7.7. Let X be a compact, oriented, C ∞ Riemannian n-manifold
(without boundary), and let E be a C ∞ Hermitian vector bundle over X of fiber
dimension N .
a) For a positive integer s, define the Sobolev space (to begin with, only the
underlying vector space)

(7.1) W s (E) := {u ∈ L2 (E) : for each P ∈ Diff s (E, E), there is v ∈ L2 (E)
such that hu, P wi0 = hv, wi0 for all w ∈ C ∞ (E)}.
Show that a section u ∈ L2 (E) liesPin W s (E), exactly when, for each local repre-
sentation of u in the form u(x) = ui (x)ei (x) relative to a local chart and local
basis e1 , ..., eN of E, we have ϕui ∈ W s (Rn ), for all C ∞ functions ϕ with support
in the domain of the chart.
b) Define W s (E) for s ∈ R+ , using this local recipe.
[Hint: For a: Definition 7.1 and Exercise 7.2a. Note that v is uniquely determined
by P . One says that v arises by weak application of the formal adjoint opera-
tor P ∗ on u (differentiation in the distributional sense, i.e., “P ∗ u = v” ⇐⇒
P ∗ u and v act identically on all test vector valued functions w ∈ C ∞ (E) with
hP ∗ u, wi0 = hu, P wi0 ).
For b: The crucial point is the independence of the set W s (E) of the choice of charts,
local trivializations, and the smoothing functions. Take care with the coordinate
changes: It is not entirely trivial that each diffeomorphism κ : U → V between open
subsets of Rn induces (via v 7→ v ◦ κ) an isomorphism WK s
(Rn ) → Wκs−1(K) (Rn ),
198 7. SOBOLEV SPACES (CRASH COURSE)

s
where K ⊂ V is compact and WK (Rn ) := {v ∈ W s (Rn ) : supp v ⊆ K}. An ele-
mentary proof for this is found in [217, p.57-59]. In our context it is simpler to
jump forward to Theorem 8.19, p. 223, where we shall show the invariance of the
space Lspc of (principally classical) pseudo-differential operators under a coordinate
change (but only for s ∈ Z+ ). However, the proof goes through smoothly for
s ∈ R+ . Then define W s (E) as in (a), where P ∈ Lspc (E, E). Instead of coordinate
invariance, which is self-evident, one must show, as in (a), that one obtains elements
of W s (Rn ) locally. Details are found in [218, p.169f] or [323, p.151f].]
Remark 7.8. (a) For a fixed choice of atlas, local trivializations of the bundle
E, and an appropriate C ∞ partition of unity, one obtains a norm and scalar prod-
uct which makes W s (E) a Hilbert space. Without such choices, we must do with
a Hilbertable space ([328]’s notation).
(b) Instead of arguing with all elements in Diff s (E, E) we can define W s (E) both as
a set and as a Hilbert space by specifying a single generating operator ΛE,s , ac-
tually a (principally classical) pseudo-differential operator (belonging to Lspc (E, E),
a space to be defined below in Section 8.3): Let {Uj }j∈J be a locally finite covering
of the underlying n-manifold X by domains of coordinate charts κj : Uj → Rn and
local trivializations τj : E|Uj → Uj × CN , and let {ϕj ∈ C ∞ (X)}j∈J be a partition
of unity subordinate to {Uj }j∈J . Let u ∈ L2 (E), i.e., X hu, uih νg < ∞ where
R

νg denotes the volume element for the Riemannian metric g on X and h·, ·ih the
Hermitian product on the vector bundle E, see Definition 6.17b, p.170. We set
X
ϕj · τj−1 ◦ ΛN
 
(7.2) ΛE,s u := s uj ◦ κj ,
j∈J
(
τj ◦ u ◦ κ−1
j , on κj (Uj ),
where uj := and ΛN
s := Λs ⊕ · · · ⊕ Λs with
0, on Rn \ κj (Uj ),
Λs : W s (Rn ) → L2 (Rn ) as defined in Exercise 7.5. The process used in Equation
(7.2) to construct the operator ΛE,s from the operators ΛN s given in local coordi-
nates is called gluing together. Similarly, we defined differential operators globally
via local coordinates in Definition 6.35, p.181 and shall define pseudo-differential
operators globally via local coordinates in Definition 8.13, p.215. We set
W s (E) := Λ−1 2

(7.3) E,s L (E) and hu, vis := hΛE,s u, ΛE,s vi0
for u, v ∈ W s (E). One checks that the Equations (7.1) and (7.3) yield the same
vector space W s (E). Contrary to the definition of the Sobolev space by (7.1), the
preceding definition based on the generating ΛE,s yields a scalar product at once
and makes W s (E) a Hilbert space. As before, however, the inner product is not
canonical but depends also here on the choice of coordinates etc, entering into the
definition of ΛE,s .
Warning: While we had Λ−s ◦ Λs = Id get on W s (Rn ) we now get an error term
(7.4) Rs := ΛE,−s ◦ ΛE,s − IdW s (E) ,
R
which is by definition an integral operator (Rs u)(x) = X Ks (x, x − y)u(y)dy with
smooth kernel Ks and therefore compact in B(W s (E), W s (E)) (to be proved like
in Exercise 2.29, p.25) and extendable to the whole L2 (E) and transforming it into
C ∞ (E).
(c) However, there are other choices to fix the scalar product: As with Exercise
7.2. DEFINITION 199

7.2b, on manifolds one may define a Laplace operator ∆ which is an elliptic, self-
adjoint, positive-definite second-order differential operator. For a natural s (and
also for real s, via the Spectral Theorem 2.61, p.51), we then may explicitly set
s
kuks := h(Id +∆) u, ui0 , u ∈ C ∞ (E).
(d) By [44, p.511] (the idea goes back to [284, p.134-197], see also [280, p.42]),
one can proceed in this way even further, if the vector bundle E is furnished
with a C ∞ connection ∇E (defined and discussed in Section 6.5, 176ff). Consider
∇E : C ∞ (E) → C ∞ (E ⊗ T ∗ X) as a differential operator. By composition with its
formally adjoint ∇∗ : C ∞ (E ⊗ T ∗ X) → C ∞ (E), we obtain a positive, semi-definite,
formally self-adjoint, Laplacian, namely ∆ := ∇∗ ◦ ∇.
Following up on this remark, we may replace our conventional introduction of
the Sobolev spaces via lengthy and, in principle, artificial coordinate transforma-
tions, by an alternative geometric (in particular, coordinate-free) definition, namely
by specifying one single generating differential operator ∇j .
Definition 7.9. We equip the bundle E → X with a Hermitian structure h
and a metric covariant differentiation operator ∇E : C ∞ (E) → C ∞ (T ∗ X ⊗ E). By
also employing a Riemannian metric g and Levi-Civita connection θg on X, we
obtain for any k = 0, 1, 2, . . . a connection in the tensor products (⊗k T ∗ X) ⊗ E
∇k,E : C ∞ (⊗k T ∗ X) ⊗ E → C ∞ (⊗k+1 T ∗ X) ⊗ E .
 

For each j = 1, 2, . . . we write shortly ∇j : C ∞ (E) → C ∞ (⊗j T ∗ X) ⊗ E for the




composition
∇E ∇1,E
∇j : C ∞ (E) −→ C ∞ (T ∗ X ⊗ E) −→ C ∞ (T ∗ X ⊗ T ∗ X ⊗ E)
∇2,E ∇j−1,E
−→ C ∞ (⊗j T ∗ X) ⊗ E .

−→ ···
a) For u, v ∈ C ∞ (E) and m > 0, we then set
Xm Z p
(u, v)m := h∇j u, ∇j viνg , and kukm := (u, u)m ,
j=0 X

where νg denotes the volume element for g, and the inner product h∇j u, ∇j vi is the
natural one constructed from the one induced by g on ⊗j T ∗ X and the Hermitian
structure on E.
b) The Sobolev space W m (E) is the completion of C ∞ (E) with the norm k·km .
Differently put, W m (E) is the space of sections u ∈ L2 (E) such that for all
j ∗

j = 1, . . . , m there exists vj ∈ L (⊗ T X) ⊗ E with ∇j u = vj weakly, i.e.,
2
Z Z
∗
hu, ∇j wi νg , for all w ∈ C ∞ (⊗j T ∗ X) ⊗ E .

hv, wi νg =
X X
We shall come back to this global type of definition in Section 16.4, Definition
16.19, p.497. For now, we leave it to the reader to check the accordance between
our two Definitions, i.e., the conventional definition by combining Exercises 7.3 and
7.7, and the preceding global Definition 7.9.
Remark 7.10. Not the underlying set of the Sobolev space W m (E) introduced
here, but its scalar product depends (as before, but differently) on several choices:
the metrics on X and E and the connection on E. It was pointed out in [321,
200 7. SOBOLEV SPACES (CRASH COURSE)

Theorem 9.2.24] that the identity map between two such versions of W m (E) is a
Banach space isomorphism, if we restrict ourselves to compatible connections (i.e.,
metric in the sense of Definition 6.29, p.179). When we expand our definition to
noncompact X this dependence is very dramatic and has to be taken into serious
consideration.
Remark 7.11. In the literature and in our Part IV, much more general Sobolev
spaces (Bessel potentials, etc.) are treated. For these, one starts with Lp theory
instead of the Hilbert spaces L2 , and works with weights other than our (1 +
2
|ξ| )s/2 . It is interesting that in “the study of classes of differential equations with
variable coefficients which are defined by conditions on their highest-order part”
(Hörmander) only those W -spaces play a role which are distinguished in a way
by their invariance on manifolds (translation invariance of L2 and diagonalizability
of the derivative by means of Fourier transformation). Indeed, in the present part of
our book, the L2 -modeled Sobolev spaces suffice. To describe the manifold structure
of moduli of self-dual connections, however, we shall, as mentioned before, define
wider Sobolev spaces in Section 16.4, in immediate generalization of the preceding
Definition 7.9.
Sobolev Spaces Over Half-Spaces. In many applications it is natural to
consider manifolds with boundary, modeled on half-spaces. For a comprehensive
treatment we refer to [83]. For now, the following Exercise may suffice.
Exercise 7.12. Let m be a positive integer.
a) For Rn+ := {x ∈ Rn : xn ≥ 0}, define (as in Exercise 7.7a) the space
W m (Rn+ ) := {u ∈ L2 (Rn+ ) : for each P ∈ Diff m (Rn+ ), there is
v ∈ L2 (Rn+ ) such that hu, P wi0 = hv, wi0 for all w ∈ C ∞ Rn+ }.


Prove that
W m (Rn+ ) = {u ∈ L2 (Rn+ ) : there is v ∈ W m (Rn ) with v|Rn+ = u}.
Show that W m (Rn+ ) is a Hilbert space.
b) Define the space W m (X) for a compact, orientable manifold X with boundary
via localization, and carry over Exercise 7.3c.
[Hint: For a: Set kukm := inf{kvkm : v ∈ W m (Rn ) and v|Rn+ = u}. Be careful with
restricting to the half-space: For m 6= 0, one must distinguish between W m (Rn+ )
and WRmn (Rn ), the space of W m functions with support in Rn+ ; see [217, p.51-54].
+
For b: The invariance under diffeomorphism is trivial here, since (without any loss
in the applications) we only consider whole numbers m. See also [217, p.60f].]

3. The Main Theorems on Sobolev Spaces


Here, we discuss briefly (partly without full proofs, for which we refer to the
literature) the three main theorems. The content of these results is illustrated in
Section 7.4, Case Studies.
We begin with a regularity theorem showing how one can pass from Hilbert
space results given in the language of Sobolev spaces to results in classical form.
Theorem 7.13 (S. L. Sobolev, 1938). Let X be a compact C ∞ manifold
(with or without boundary) and s > 0. Define the strength of the Sobolev space
7.3. THE MAIN THEOREMS ON SOBOLEV SPACES 201

W s (X) by
dim X
str(W s (X)) := s − .
2
Then W s (X) ,→ C k (X) for all k < str(W s (X)), and the embedding is continuous
with the Sobolev inequality
kukC k ≤ ε kukW s + Cε kukL2 for u ∈ W s (X) and k < str W s (X),
where ε > 0 can be made arbitrarily small, if Cε is sufficiently large.
Note . Whence, the strength of W s (X) is a measure for the regularity of
its elements: the bigger the strength, the more regular are the functions in that
space and thus it consists of fewer functions or, rather, elements or classes. More
precisely, an element u of W s (X) ⊂ L2 (X) is a class of functions which agree
almost everywhere. The theorem means this: In each class u ∈ W s (X), there is a
representative in C k (X), and each sequence of elements in W s (X) which converges
in the norm of W s (X) yields a sequence of representatives in C k (X) which converges
in the norm of the Banach space C k (X).
Proof. For the Euclidean case, we give a taste of the proof below in Theorem
7.16. Else see [58, p.167] and [328, p.159f] for X = T n := S 1 × · · · × S 1 , and
[328, p.169] for the transition to arbitrary X and to sections of vector bundles. For
the Sobolev inequality, see [6, 75-76/97f], where X is a codimension 0 submanifold
(with boundary) of Rn . 
The treatment of boundary value problems with Hilbert space methods is prob-
lematic, since in a fixed L2 space a function is only uniquely defined modulo its
values on sets of measure zero such as the boundary. Thus, the restriction of such a
function to the boundary is completely arbitrary. However, the following restriction
theorem, which also goes back to S.L. Sobolev, is helpful (e.g., for Y = ∂X and
m = 1).
Theorem 7.14. Let X be a compact, C ∞ manifold (possibly with boundary)
with a compact submanifold Y of codimension m, and let E be a vector bundle
over X. Then, for each integer s > m/2, the canonical restriction map C ∞ (E) →
m
C ∞ (E|Y ) extends to a continuous, linear, surjective map W s (E) → W s− 2 (E|Y ).
Proof. For the periodic case X = T n , Y = T n−1 and E the trivial line
bundle, we give a full proof below in Theorem 7.17. For the hyperplane problem
1
W (Rn+ ) → W s− 2 (Rn−1 ), see [217, p.54f] or [280, p.38]; for Y = ∂X, see [280,
p.44-48]; for X = T n and Y = T n−m and the general case, see [328, p.161 f]. 
Finally, the following lemma, named after Franz Rellich and proved by him
in a different formulation, brings in compact operators (see our Chapters 2ff), and
will furnish a further connection with the Fredholm theory of elliptic operators.
Theorem 7.15 (F. Rellich, 1930). If X is a compact C ∞ manifold (possi-
bly with boundary) and E is a complex vector bundle over X, then the inclusion
W m (E) ,→ W s (E) is compact for m > s ≥ 0.
There are many different proofs in the literature: See [6, p.144], when X is a
codimension 0 submanifold of Rn . For X = T n := S 1 × · · · × S 1 , see [58, p.169f] or
[328, p.158f], similarly the very clear presentation [190, Theorem 8.2], and [328,
p.168] for the general case. We shall give an explicit proof in the special scalar and
202 7. SOBOLEV SPACES (CRASH COURSE)

1
compact-supported Euclidean case of WK (Rn ) ,→ L2 (Rn ) (closely following [321,
Theorem 9.2.14]). One may wonder about a shorter, structural and more general
proof. Inspiration may be found in [391, Theorem 7.4]). See also our Remark 7.8b
with the compact error term Rs := ΛE,−s ΛE,s − IdW s (E) of (7.4) on p.198.

Explicit proof in scalar, Euclidean, compactly supported case. Let


R > 0 and let us show that the inclusion WB1 R (Rn ) ,→ L2 (Rn ) is compact. Let (uµ )
be a bounded sequence in W 1 (Rn ) supported in the ball BR = BR (0) := {|x| ≤ R}.
We have to show that the sequence contains a subsequence convergent in L2 . The
proof will be carried out in two steps.
Step 1. We will prove that for every 0 < δ < 1 the mollified sequence (see
Remark 7.4) (uµ,δ := ϕδ ∗ uµ ) admits a subsequence uniformly convergent on
BR+1 := BR+1 (0).
To prove this we will apply the Arzela-Ascoli Theorem (Theorem 2.17 and
Corollary 2.18, p.18f.). Whence, we shall show that
(i) ∃C=C(δ) ∀µ ∀|x|≤R+1 |uµ,δ (x)| < C,
(ii) ∀µ ∀|x|,|x0 |≤R+1 |uµ,δ (x) − uµ,δ (x0 )| < C|x − x0 |.
Indeed,
 
x−y
Z Z
−n −n
|(ϕδ ∗ uµ ) (x)| ≤ δ ϕ |uµ (y)| dy ≤ δ |uµ (y)| dy
|y−x| ≤ δ δ Bδ (x)
Schwarz
≤ C(n)δ −N kuµ kL2 (Rn ) · vol(Bδ )1/2 ≤ C(δ) kuµ kW 1 (Rn ) ,

by the continuous embedding of W 1 ,→ L2 , Exercise 7.3c.


Similarly,
Z
|uµ,δ (x) − uµ,δ (x0 )| ≤ |ϕδ (x − y) − ϕδ (x0 − y)| · |uµ (y)| dy
BR+1
Z
≤ C(δ) · |x − x0 | |uµ (y)| dy ≤ C(δ) · |x − x0 | kuµ kW 1 (Rn ) .
BR+1

Step 1 is completed.
Step 2. So, for each δ ∈ (0, 1) we have a uniformly convergent subsequence (vν,δ )
of the mollifier sequence (uµ,δ ). Using the diagonal procedure for δ = 1/ν → 0 and
ν → ∞, we pick for each ν ∈ N the function vν,1/ν and denote the corresponding
element of the original sequence (uµ ) by u0ν = uµ0 , i.e., the element uµ0 that yields
exactly vν,1/ν = ϕ1/ν ∗uµ0 by convolution with ϕ1/ν . Note that the sequence vν,1/ν
is uniformly convergent on BR by Step 1 and limν→∞ kvν,1/ν −u0ν kW 1 (Rn ) = 0 under
mollification.
We claim that the subsequence (u0ν ) ⊂ (uµ ) is convergent in L2 (BR ). Indeed,
for all natural ν and ρ

ku0ν −u0ρ kL2 (BR ) ≤ ku0ν −vν,1/ν kL2 (BR ) +kvν,1/ν −vρ,1/ρ kL2 (BR ) +kvρ,1/ρ −u0ρ kL2 (BR ) .

Each of the three terms on the right tends to 0 for ν → ∞ since the embedding
W 1 ,→ L2 is continuous (once again, Exercise 7.3c). Hence the subsequence (u0ν )
of our original bounded sequence (uµ ) is a Cauchy sequence in L2 (BR ) and thus it
converges. The compactness theorem is proved. 
7.4. CASE STUDIES 203

4. Case Studies
To illustrate the preceding theorems, we consider some simple special cases. To
begin with we prove a simple Euclidean version of Theorem 7.13.
Theorem 7.16. If s > n/2, then each u ∈ W s (Rn ) is bounded and continuous,
and the inclusion W s (Rn ) ,−→ C 0 (Rn ) is continuous.
Proof. By the Fourier Inversion Formula and the Integrable-Continuous Con-
b is in L1 (Rn ).
version (Exercise A.5, p.711 in Appendix A), it suffices to prove that u
Indeed, we get
Z Z  s/2  −s/2
2 2
(7.5) |b
u(ξ)| dξ ≤ |b
u(ξ)| 1 + |ξ| 1 + |ξ| dξ
Rn Rn
Z  s 1/2 Z  −s 1/2
2 2 2
≤ |b
u(ξ)| 1 + |ξ| dξ 1 + |ξ| dξ .
Rn Rn
1/2 1/2
Here we have used the Schwarz Inequality ha, bi ≤ ha, ai hb, bi ; note that the
first factor is finite by assumption and the latter factor is finite precisely for s > n/2.
By the Riemann-Lebesgue Lemma, we can even conclude that u(x) vanishes at
infinity. To prove the continuity of the inclusion, let u ∈ C0∞ (Rn ). We have for
x ∈ Rn , with our convention d̄ξ := (2π)−n/2 dξ,
Z Z
|u(x)| = eixξ u
b(ξ) d̄ξ ≤ |b
u(ξ)| d̄ξ.
Rn Rn
Estimate (7.5) and the definition in Exercise 7.2b then yield
sup |u(x)| ≤ K kukW s(Rn ) ,
x∈Rn
R 1/2
2
where the constant K (e.g., (2π)−n/2 Rn
(1 + |ξ| )−s dξ ) does not depend on
u. Since C0∞ (Rn ) s n
is dense in W (R ) (see Exercise 7.3c, p. 195), we are done. 
n n n
Theorem 7.17. We write the n-dimensional torus T in the form R /(2πZ ).
Then the restriction map C ∞ (T n ) → C ∞ (T n−1 ) (induced by the projection (y, θ) 7→
1
y of T n onto T n−1 ) extends to a continuous linear map W s (T n ) → W s− 2 (T n−1 ),
for s ≥ 1/2.
Proof. We follow [328, p.143-162]and build on Peter Lax’s remark in [274]
that, in the periodic case, certain technical difficulties vanish:
Step 1: C ∞ (T n ) consists of functions on Rn which are periodic of period 2π
in each variable. The functions
−n/2 2πihν,xi
eν (x) := (2π) e , ν ∈ Zn
form a complete orthonormal system for L(T n ); see the theory of Fourier series
(Appendix A). By definition (see Exercise 7.2b), we have u ∈ W s (T n ) exactly
when X
2 2 2
kuks := n
u(ν)| (1 + |ν| )s < ∞,
|b
ν∈Z
where the ν-th Fourier coefficient u
b(ν) is given by
Z
−n/2
u
b(ν) := (2π) u(x) e−2πihν,xi dx = hu, eν iL2 (T n ) , with
Tn
2
hν, xi := ν1 x1 + · · · + νn xn and |ν| = hν, νi .
204 7. SOBOLEV SPACES (CRASH COURSE)

b(ν)eν converges absolutely to u in the W s (T n ) topology.


P
Then the series ν u
P Step 2: For x = (y, θ) ∈ T
n−1
× S 1 and u ∈ C ∞ (T n ), we have u(y, θ) =
b(λ, µ)eλ (y)eµ (θ), where the sum is over all (λ, µ) ∈ Zn−1 × Z and the
(λ,µ) u
convergence on T n is uniform. Since eµ (0) = (2π)−1/2 , it follows that
X X
u(y, 0) = (2π)−1/2 n−1
e λ (y) u
b(λ, µ),
λ∈Z µ∈Z
n
where the series converges uniformly on T . Thus, we have
X
(u|T n−1 )b(λ) = (2π)−1/2 u
b(λ, µ).
µ∈Z

Step 3: We essentially follow [328, p.143f] (but avoid a minor error, the last
inequality on p. 143). For s ≥ 12 , b ≥ 1 and a : Z+ → R+ , we set
 2
−s/2  2
s/2
xµ = b−1/4 1 + µb and yµ = aµ b1/4 1 + µb .
P 2
x2µ yµ2 yields
P P
The Schwarz inequality µ∈Z xµ yµ ≤ µ∈Z µ∈Z
X 2  −s X  s
µ2 µ2
X
aµ ≤ b−1/2 1 + b a2µ b1/2 1 + b
µ∈Z µ∈Z µ∈Z
or X 2  −s X
1
X
µ2
s
aµ bs− 2 ≤ b−1/2 1 + b a2µ b + µ2 .
µ∈Z µ∈Z µ∈Z
By integral comparison,
 −s  Z ∞
−s
µ2 2
X
b−1/2 1 + b ≤ b−1/2 + 2 b−1/2 1 + xb dx
µ∈Z 0
∞ √
πΓ s − 21
Z 
−s
= b−1/2 + 2 1 + y2 dx = b−1/2 + .
0 Γ(s)

πΓ(s− 12 )
For Cs := 1 + Γ(s) we have (since b ≥ 1)
,
X 2 1
X s
aµ bs− 2 ≤ Cs a2µ b + µ2 .
µ∈Z µ∈Z
2
We set aµ := |b
u(λ, µ)| and b := 1 + |λ| and then obtain
X 2  s− 12 X  s
2 2 2
|b
u(λ, µ)| 1 + |λ| ≤ Cs u(λ, µ)| 1 + |λ| + µ2 .
|b
µ∈Z µ∈Z

Thus, with the above Step 2, we have


1
2
X 2 2
k(u|T n−1 )ks− 1 = n−1
|(u|T n−1 )b(λ)| (1 + |λ| )s− 2
2 λ∈Z
X X 2 1
2
= n−1
(2π)−1 b(λ, µ) (1 + |λ| )s− 2
u
λ∈Z µ∈Z
X X 2 1
−1 2
≤ (2π) n−1
u(λ, µ)| (1 + |λ| )s− 2
|b
λ∈Z µ∈Z
X X  s
2 2
≤ (2π)−1 Cs n−1
|b
u(λ, µ)| 1 + |λ| + µ 2
.
λ∈Z µ∈Z

Thus, p
ku|T n−1 ks− 1 ≤ Cs /2π kuks ,
2

and we are done, since C ∞ (T n ) is dense in W s (T n ). 


7.4. CASE STUDIES 205

The following case study gives insight into the possible loss of differentiability
under restrictions of Sobolev spaces:
Theorem 7.18. There is no continuous, linear map W s (Rn ) → W s (Rn−1 )
which extends the restriction map C0∞ (Rn ) → C0∞ (Rn−1 ) defined by u 7→ u(·, 0).
α
Proof. On the ball B n := {x ∈ Rn : |x| ≤ 1}, the function |x| is integrable,
if α > −n, since then in polar coordinates, we have
Z Z 1
α
|x| dx ≤ C rα+n−1 dr < ∞.
Bn 0
α
Now consider the function u(x) := |x| χ(x), where χ is a C ∞ function with com-
pact support and χ(x) = 1 for all x ∈ B n . For α = −1/2 and n = 2, we have
R1
u ∈ L2 (R2 ), but u(·, 0) ∈/ L2 (R), since 0 x−1 dx = ∞. Thus, the theorem is veri-

fied for s = 0, since there is a sequence {uν }ν=1 with uν ∈ C0∞ (R2 ) which converges

in W 0 (R2 ) (= L2 (R2 )) to u, but {uν (·, 0)}ν=1 is not a Cauchy sequence in W 0 (R)
(= L2 (R)). One can also easily construct counterexamples for s > 0, since by the
above argument it follows that the function u (defined there) lies in W 1 (Rn ) ex-
−1/4
actly when 2α > 2 − n (see Exercise 7.2a). For example, u(x) := |x| χ(x) is an
1 3 1 2 1
element of W (R ), but u(·, ·, 0) ∈ / W (R ) and u(·, 0, 0) ∈/ W (R). If a sequence
uν ∈ C0∞ (R3 ) converges to u ∈ W 1 (R3 ) and the restrictions uν (·, 0, 0) were to con-
verge in W 1 (R), then the limit in W 1 (R) would also be in C 0 (R) by Theorem 7.16;
this contradicts the form of u. 
Theorem 7.19. Without the assumption that X is compact, Theorem 7.15
above is false.
Proof. For each n ∈ N, one constructs un ∈ W 1 (R) with kun k1 < 3, as in
2 √ √ 2 √ 2
Figure 7.1. However, we have kun − u2n k0 ≥ 2n 1/ n − 1/ 2n = 2−1 ,

Figure 7.1. No Rellich Theorem for noncompact X := R

independent of n. Thus, the un lie in a bounded subset of W 1 (R), but there is no


subsequence convergent in W 0 (R) = L2 (R). 
CHAPTER 8

Pseudo-Differential Operators

Synopsis. Motivation: Fourier Inversion; Symbolic Calculus; Quantization. Canon-


ical and Principally Classical Pseudo-Differential Operators. Pseudo-Locality; Singular
Support. Standard Examples: Differential Operators; Singular Integral Operators. Oscil-
latory Integrals. Kuranishi Theorem. Change of Coordinates. Pseudo-Differential Opera-
tors on Manifolds. Graded ∗ -Algebra. Invariant Principal Symbol; Exact Sequence; Non-
canonical Op-Construction as Right Inverse. Coordinate-Free (Truly Global) Approach:
Bokobza-Haggiag Fourier Transformation; Bokobza-Haggiag Amplitudes; Bokobza-Haggiag
Invertible Op-Construction; Approximation of Differential Operators.

1. Motivation
I For better reading by a traditionally educated student, this chapter is based on a
local definition of pseudo-differential operators and then generalizes to operators acting on
sections of vector bundles over manifolds by charts and local trivializations of the bundles.
That approach has its merits since it has become the standard way of introducing pseudo-
differential operators and since it admits some easy elementary calculations. However, for
geometrically defined operators, the arbitrary character of the coordinate shifts does not
facilitate calculations and can even block for natural constructions (like the product of
pseudo-differential operators in nontrivial cases). J

Whence, a self-confident reader may skip the first four sections of this chapter
and advance directly to Section 8.5 where we give a coordinate-free description of
pseudo-differential operators. As we shall see in our Part III, global description
is more powerful for the investigation of the analytical index under embedding.
That said, the reader must be reminded that also our global description of pseudo-
differential operators is neither really invariant nor canonical : While it does not
depend on coordinates, it depends heavily on other choices, namely the choice of
metric structures on the manifold and the involved bundles and on the choice of
connections for those bundles.
We will now turn to a class of operators which, roughly speaking (details below),
are locally presentable in the form
Z
(P u)(x) := eihx,ξi p(x, ξ)b
u(ξ) d̄ξ,
Rn
where we use the convenient shorthand
−n/2
d̄ξ = (2π) dξ
n n/2
for Lebesgue measure on R divided by (2π) . Here p is called the amplitude of
the operator P , hx, ξi is its phase function, and
Z
u
b(ξ) := e−ihx,ξi u(x) d̄x
Rn

206
8.1. MOTIVATION 207

denotes the Fourier transform of u (see the crash course in Appendix A).
There are a number of reasons why these pseudo-differential operators have
commanded increasing attention since the appearance of the pioneering studies by
Solomon Grigoryevich Mikhlin on Singular Integral Equations (1948). We
mention the following overlapping aspects.
1. This class is large enough to contain in addition to the differential operators
(Exercise 8.5 below, p.212) the Green operators (see also Chapter 2) and other sin-
gular integral operators which play a role in solving partial differential equations.
In particular this class of pseudo-differential operators contains, with each elliptic
operator, its parametrix, i.e., a quasi-inverse modulo an operator of lower order. In
Theorem 9.8 (p.239) below we will incorporate this operator calculus into Hilbert
space theory, and in this fashion we will be able to derive easily the classical results
on elliptic operators (regularity theorems, finiteness of the index) using the ele-
mentary theory of Fredholm operators developed in Chapters 1-3. Thereby “some
of the techniques used in the case of differential operators appear here as general
properties of the class of integro-differential operators considered” (Seeley).
2. The class is small enough and close enough to the differential operators to
allow convenient computations. This standpoint is important, particularly because
the progress in functional analysis of the past decades permitted the definition of
more and more general and involved operator classes and phantom spaces (Thom),
while the exploration of their properties was too difficult and lagged behind. In
contrast, turning to pseudo-differential operators, for which an exact calculus was
developed, signalled “a trend in the theory of general partial differential equations
towards essentially constructive methods” (Hörmander).
3. A special aspect is the attempt to deal with differential operators with vari-
able coefficients, by means of pseudo-differential operators in first approximation, in
the same way differential operators with constant coefficients are treated by means
of the Fourier transform: For example, for f ∈ C ∞ (Rn ) with compact support and
n > 2, consider the inhomogeneous equation ∆u = f , where ∆ = ∂ 2 /∂x21 + · · · +
∂ 2 /∂x2n is the Laplace operator. With the Fourier transform (see the multiplication
rule in Exercise A.5b, Appendix A), we obtain −(ξ12 +· · ·+ξn2 )b u(ξ) = ∆u(ξ)
c = fb(ξ),
2 2
b(ξ) = −fb(ξ)/ |ξ| , as an L function (for n > 2), and further with the Fourier
i.e., u
inversion formula,
Z
−2
(Qf )(x) = u(x) = − eihx,ξi |ξ| fb(ξ)d̄ξ,
Rn

where Q is the inverse operator (fundamental solution) of ∆. In general, suppose


P is a differential operator with constant coefficients which can be written as a
polynomial P = p(D) where D = (−i∂/∂x1 , . . . , −i∂/∂xn ), and consider the inho-
mogeneous equation p(D)u = f , f ∈ C ∞ (Rn ) of compact support. We obtain in
the same way, at least formally, a solution u = Qf , where
Z
(Qf )(x) := eihx,ξi q(ξ)fb(ξ) d̄ξ,
Rn

and q(ξ) := p(ξ)−1 is the amplitude. In the process, a number of difficulties arise.
Indeed, Qf is in general not C ∞ , and possibly only a distribution, and the integral
must be interpreted, since the zeros of p can cause divergences. But these problems
can be resolved almost completely; e.g., see the following references of quite different
208 8. PSEUDO-DIFFERENTIAL OPERATORS

depth: [105], [171, 172, 173], [190, pp.108-110],[217, Chs. III and IV ], [223],
[411, Chapter 3].
Now, if (as in Chapter 5) U ⊆ Rn is open and
X
P = p(x, D) = aα (x) Dα , with aα ∈ C ∞ (U )
|α|≤k

is a differential operator with variable coefficients, then all these methods fail ini-
tially. But we can, “as a good physicist would” (Atiyah), formally invert the
operator P by freezing the coefficients at a point x0 ∈ U and considering P as a
perturbation of p(x0 , D) which is a differential operator with constant coefficients.
In this way we obtain as an approximate inverse of P a pseudo-differential operator
with the amplitude q(ξ) = p(x0 , ξ)−1 . In order to get a better approximate inverse
it is natural to slowly thaw the coefficients, i.e., to let the point x0 vary in U . This
yields the operator Z
(Qf )(x) := eihx,ξi q(x, ξ) fb(ξ) d̄ξ
Rn
−1
with the amplitude q(x, ξ) = p(x, ξ) (x ∈ U, f ∈ C ∞ (Rn ) with compact support).
This basic perturbation argument, which was supplied in the study of elliptic differ-
ential equations by the Italian mathematician Eugenio Elia Levi already in the
year 1907, thus finds its theoretical framework within the class of pseudo-differential
operators.
We remark (see Theorem 9.8, p. 239) that in the elliptic case an equally good
approximation is obtained by choosing as amplitude the inverse of the principal
part (symbol), i.e., the function pk (x, ξ)−1 which is homogeneous of degree −k in
ξ. There is a particularly simple calculus of such operators, since the asymptotic
expansions of p(x, ξ) and q(x, ξ) and the underlying iteration (usually necessary) is
avoidable here.
4. The theory of pseudo-differential operators allows a certain relaxing of cus-
tomary precision, a precision which is senseless, or at least exaggerated, in a number
of practical problems. Thus, in order to investigate regularity and solvability of the
differential equation P u = f , we do not need an actual inverse operator (funda-
mental solution), but (in the framework of Fredholm theory) it suffices to have a
parametrix, i.e., a quasi-inverse modulo certain elementary operators (see Chap-
ter 3). This has considerable computational advantages. In the case of differential
operators P = p(D) with constant coefficients (as in Item 3 above), we may take
the amplitude to be
q(ξ) := χ(ξ)p(ξ)−1 ,
where χ(ξ) is a fixed C ∞ bump function which is identically zero in a disk about
the origin and identically 1 for large ξ. In this fashion we obtain a well-defined
integral Z
(Qf )(x) := eihx,ξi q(ξ) fb(ξ) d̄ξ
Rn
and avoid the delicate convergence problems which would be met for the amplitude
p(ξ)−1 because of singularity at the zeros of p. While P Q = Id and QP = Id are
not valid, we still have
P Qf = f + Rf for f ∈ C ∞ (U ), where
Z
(Rf )(x) := r(x − ξ) f (ξ) d̄ξ.
Rn
8.1. MOTIVATION 209

and rb = χ − 1; so r ∈ C ∞ , and R is a smoothing operator. Indeed,


Z
(P (Qf ))(x) = eihx,ξi p(ξ)Qf
c (ξ) d̄ξ and
Rn
Z Z 
−ihy,ξi
Qf
c (ξ) = e e ihy,ηi
q(η) fb(η) d̄η d̄y = q(ξ)fb(ξ) = χ(ξ)p(ξ)−1 fb(ξ).
Rn Rn
Using the convolution formula of Exercise A.5e in Appendix A, we then have
Z
(P (Qf ))(x) = eihx,ξi p(ξ)χ(ξ)p(ξ)−1 fb(ξ)d̄y d̄ξ
Rn
Z
= eihx,ξi (1 + (χ(ξ) − 1)) fb(ξ) d̄ξ
R n
Z Z
= f (x) + eihx,ξi (χ(ξ) − 1) fb(ξ) d̄ξ = f (x) + r(x − ξ) f (ξ) d̄ξ.
Rn Rn
Since much is known about simple correction or residue operators such as R, a
parametrix Q serves just as well as a true fundamental solution for which R vanishes.
At any rate, fundamental solutions do not usually exist when passing to variable
coefficients in the perturbation method (sketched in Item 3) and when replacing
operators and their amplitudes by principal symbols (the terms of highest order in
the amplitudes). However, the simple computations modulo smoothing operators
and other operators of lower order can be used very efficiently and arise naturally
in the theory of pseudo-differential operators.
5. In Section 6.6 (p. 189) we pointed out, in connection with the symbolic
calculus, the basic significance of the change of levels in passing from operators to
the functions which characterize them in approximation. The pseudo-differential
operators form a class (and this is tied to the perturbation argument) whose op-
erators can at least in approximation (actually, precise in the global approach de-
scribed below in Section 8.5) be described by their amplitudes and symbols. The
latter are functions satisfying simple rules of computation resulting in a particularly
simple approximation theory for the corresponding operators; see for example the
composition rules in Theorem 8.27, p. 229. Surely, mathematicians such as Vito
Volterra, Erik Ivar Fredholm, David Hilbert and Friedrich (Frigyes)
Riesz had this goal in mind when they developed the theory of integral equations as
a means of dealing with differential operators. However, starting with the classical
representation Z
(Qf )(x) = K(x, z)f (z) dx,
it turns out that the formulation of the correct conditions for the weights K are by
far not as simple and natural as are those for the amplitudes and symbols in the
representation via Fourier transform. The same is true for the transformation and
composition rules (see Theorem 8.3, p. 210, and Theorem 8.14, p. 217).
6. From the topological standpoint, the following issues are particularly impor-
tant: In the passage between the levels of consideration, we prefer to operate with
symbols. They are more accessible by topological means than operators. In Part
III below, the larger class of pseudo-differential operators has a decisive advantage.
Indeed, the associated extension of the symbol space beyond the polynomial maps
to arbitrary C ∞ functions permits lifting a homotopy of the symbol of a differential
operator to a homotopy of the operator itself in the space of pseudo-differential
operators. This is generally impossible in the space of differential operators. Using
210 8. PSEUDO-DIFFERENTIAL OPERATORS

heavier topological machinery, this difficulty can be dealt with without the use of
pseudo-differential operators. However, the difficulties which occur are not to be
underestimated. For example, not much is known about the simplest question of
the existence of an elliptic system (in Rn ) of N differential equations of order k with
constant coefficients. For k = 1, this is the case exactly when the (N − 1)-sphere
S N −1 admits n − 1 linearly independent vector fields; hence, for N = n, (according
to a famous theorem of John Frank Adams) exactly for the values 2, 4, and
8. More about this is in [21]. In [157], M. Furuta presented a full proof of the
Atiyah-Singer Index Theorem without the use of pseudo-differential operators.
7. Finally, we point out that the class of pseudo-differential operators origi-
nally was developed only in connection with elliptic differential equations, and only
there (and with the closely related hypo-elliptic differential equations) the beautiful
properties listed above unfold fully. However, Lars Hörmander and other au-
thors succeeded, in a series of papers and monographs, in generalizing the concept
of a pseudo-differential operator in such a way that the theory Fourier integral op-
erators so created leads to new results also in heat transfer and wave operators, for
example. For this aspect, which we cannot pursue further, see [220, 225].

2. Canonical Pseudo-Differential Operators


We begin with the definition of the prototypes of our pseudo-differential oper-
ators in local form over the open subset U ⊆ Rn and acting on functions only:
Z
(8.1) (P u)(x) = eihx,ξi p(x, ξ) u
b(ξ) d̄ξ =: (Op(p)u) (x),
Rn
where x ∈ U , u ∈ C0∞ (U ) (i.e., u ∈ C ∞ (U ) and the support of u is compact).
Definition 8.1. a) The operator P = Op(p) is called a canonical pseudo-
differential operator of order k ∈ R, shortly P ∈ Lk (U ), if the amplitude p ∈
C ∞ (U ×Rn ) satisfies the following asymptotic conditions of growth as |ξ| → ∞: For
each compact subset K ⊂ U and multi-indices α = (α1 , ..., αn ), β = (β1 , ..., βn ) ∈
Zn+ , there is a C ∈ R such that for all x ∈ K and ξ ∈ Rn , we have
k−|α|
(8.2) Dxβ Dξα p(x, ξ) ≤ C(1 + |ξ|) .
(0,...,1,...0)
Recall that |α| := |α1 + · · · + αn | and that Dx := −i∂/∂xj where “1”
stands in the j-th place of the multi-index (0, . . . , 1, . . . 0).
b) The set of amplitudes satisfying (8.2) will be denoted by Sk (U × Rn ). We set
S• (U × Rn ) := k∈R Sk (U × Rn ) and L• (U ) := k∈R Lk (U ).
S S

Remark 8.2. a) Amplitudes satisfying (8.2) are often called symbols of Hörmander
type (1, 0). The definition extends to matrix valued amplitudes, needed below for
defining pseudo-differential operators acting on sections of vector bundles.
b) A global version of the estimates (8.2) is given below in (8.16).
c) The best constants in (8.2) provide a set of semi-norms which endow S• (U × Rn )
with the structure of a Fréchet algebra.
The estimate (8.2) plays a key role in the derivation of many useful properties
of pseudo-differential operators, as in the following:
Theorem 8.3. Each canonical pseudo-differential operator is a linear map from
C0∞ (U ) to C ∞ (U ).
8.2. CANONICAL PSEUDO-DIFFERENTIAL OPERATORS 211

b(ξ) is obviously a C ∞
Proof. For all ξ ∈ R, the integrand x 7→ eihx,ξi p(x, ξ) u
function. To show that the function
Z
x 7→ (Op(p)u) (x) = eihx,ξi p(x, ξ) u
b(ξ) d̄ξ
Rn

is also C , we must show that the integral converges sufficiently well so that the
order of integration and differentiation may be switched. More precisely, by the
Dominated Convergence Theorem of Henri Lebesgue, a function which is the
limit of a sequence of measurable functions, uniformly bounded by an (absolutely)
integrable function, is itself integrable and the limit and integral may be inter-
changed. To apply this to our situation, we must show that, for each x ∈ U and
each multi-index β, the function
 
ξ 7→ Dxβ eihx,ξi p(x, ξ) ub(ξ) (ξ ∈ Rn )
can be estimated by an integrable function. Since the support of u is compact, we
have (see Exercise A.5b, p. 711 of Appendix A) that
Z
α
ξ ub(ξ) = e−ihx,ξi Dα u(x) d̄x,
Rn
which goes to 0 as |ξ| → ∞. Hence the function ξ 7→ |ξ α ub(ξ)| is bounded for each
−1
multiindex α. Thus, ub(ξ) decreases faster than any power of |ξ| as |ξ| → ∞; i.e.,
n
for each N there is a constant C1 such that for all ξ ∈ R ,
u(ξ)| ≤ C1 (1 + |ξ|)−N .
|b
0
By (8.2), we have Dxβ p(x, ξ) ≤ C2 (1 + |ξ|)k for any β 0 ≤ β. Hence
 
|β|
Dxβ eihx,ξi p(x, ξ) ≤ C3 |ξ| (1 + |ξ|)k , and so
 
Dxβ eihx,ξi p(x, ξ) ub(ξ) ≤ C3 (1 + |ξ|)k+|β|−N ,
the right side being integrable for N sufficiently large. 
Remark 8.4. a) Recall roughly that quantization in quantum mechanics at-
tempts to convert functions of position and momentum (i.e., functions on T ∗ X)
into operators. One may think of Op(p) as a quantization of p. See also [187,
Exercise 3.1, p.36f]. There you find a sketch of how, e.g., the common commuta-
tor relations of Weyl quantization can be derived from properties of Op. Readers,
however, who have a geometric antenna and are truly interested in physics will be
bothered by the fact that our Op(p) is naturally defined only in Euclidean space. A
generalization of that quantization concept to manifolds depends on many choices.
The common way depends on the choice of charts. That does not lead very far, see
our Exercise 8.20, p.224 and our Chapter 9 or [187, Exercise 3.4] and [190, Sec-
tion 8.2]. Below in Section 8.5, we shall describe an alternative, thoroughly global
way to define Op(p) for suitable p over a closed manifold. However, that approach
will also depend on many choices (e.g., the choice of metric, connection and bump
function), as we will see. Apart from these choices, there are other choices one can
make. As [187, Chapter 11] indicates, one probably has to return to the visionary
notes by J.B. Keller, 1958, V.C. Maslov, 1972, and J. Leray, 1981 (precise
references are given in [187, p.130]) to find ideas for a quantization concept which
is physically realistic and geometrically meaningful.
b) The continuity of the operator Op(p) will be discussed later, see Theorem 9.3,
212 8. PSEUDO-DIFFERENTIAL OPERATORS

p.237.
c) We also postpone the discussion, whether an amplitude is determined from a
given pseudo-differential operator to Example 8.16, p.218.
Exercise 8.5. Show that the following standard operators define canonical
pseudo-differential operators modulo smoothing operators (i.e., pseudo-differential
operators, whose amplitudes have compact support in the second variable; see also
RemarkP 8.8, p. 213).
a) P = |α|≤k aα Dα , where aα ∈ C ∞ (U ).
b) (P u)(x) = Rn K(x, y)u(y)dy, where K ∈ C ∞ (U × U ) and the support of K(x, ·)
R

is compact for all x ∈ U . For example, the convolution u 7→ u ∗ ϕ with ϕ ∈ C0∞ (U )


(where K(x, y) = ϕ(x − y)) P is a canonical pseudo-differential operator.
c) The Riesz operator P = aα Rα , where aα ∈ C0∞ (U ), with aα = 0 for all but
finitely many multi-indices α, and
Z  α
ξ
(Rα u)(x) := eihx,ξi |ξ| u
b(ξ)d̄ξ.
Rn
(See the footnote for Exercise 8.24, p. 226 below.)
[Hint: For a: One applies the differential operator P to
Z
u(x) = eihx,ξi u
b(ξ) d̄ξ,
Rn
and then obtains a pseudo-differential operator with amplitude
X
p(x, ξ) = aα (x) ξ α , x ∈ U, ξ ∈ Rn .
|α|≤k

For b: Here also, we begin with the Fourier Inversion Formula. One obtains the
amplitude Z
p(x, ξ) = eihy,ξi K(x, y) dy.
Rn
For each fixed x, this is a multiple of the inverse Fourier transform of a function
with compact support, and hence p(x, ·) ∈ C↓∞ (Rn ), the space of rapidly decreasing
functions in C ∞ (Rn ). (Argue as in the proof of Theorem 8.3, where partial integra-
tion interchanges multiplication and differentiation.) Thus, the required conditions
on the amplitude hold Pfor each k ∈ Z.
α
For c: p(x, ξ) = χ(ξ) aα (x)(ξ/ |ξ|) where χ(ξ) is a cut-off function, i.e., χ(ξ) = 0
for |ξ| < ρ, ρ > 0, and χ(ξ) = 1 for |ξ| > ρ0 > ρ. The order of P is therefore k = 0,
and different choices of χ lead, modulo smoothing functions, to the same pseudo-
differential operator. One also calls Riesz operators singular integral operators,
since they can be alternatively represented by singular convolutions. For example,
if α = (1, 0, ..., 0), then one can (up to a constant factor, which we will ignore) write
Z
α −n−1
(aα R u)(x) = lim K(x, x − y)u(y) dy, where K(x, z) := aα (x)z1 |z| ,
ε→0 |x−y|>ε

whence the weight function K has a singularity at the point z = 0. For the connec-
tion between Riesz operators, Hilbert transformations, and Wiener-Hopf operators
in the case n = 1, see above our Chapter 4 and [345], where an algebra of pseudo-
multiplication operators in the half-space Rn+ is investigated. This algebra is formed
with the help of pseudo-differential operators and contains the Wiener-Hopf oper-
ators.]
8.2. CANONICAL PSEUDO-DIFFERENTIAL OPERATORS 213

Remark 8.6. While differential operators are local operators (see Exercise
5.2, p.136), a pseudo-differential operator can increase supports. For example, if
P is defined as convolution with ϕ ∈ C ∞ (U ) as in Exercise 8.5b, we can have
supp P u = supp u + supp ϕ. The translation operator with amplitude

p(x, ξ) := eihx0 ,ξi , x0 fixed,

which sends u(x) to u(x + x0 ) is not a canonical pseudo-differential operator. Ac-


tually, the asymptotic amplitude estimate guarantees pseudo-locality, a kind of
locality modulo operators of lower order, whereby (rather than the support) the
singular support (the closure of the set where a function is not C ∞ ) is not in-
creased. Details are found in [190, p.177], [224, Theorem 18.1.16], [323, p.151f] or
[328, p.260].

Exercise 8.7. Show that one can write every canonical pseudo-differential
operator P as an integral operator (for some λ ∈ R)
Z
λ
(P u)(x) = Kλ (x, x − y)(1 − ∆) u(y) d̄y, u ∈ C0∞ (U ),
U

where the weight function Kλ (x, z) is C ∞ for z 6= 0.


[Hint: Suppose the operator P has order k ∈ Z and amplitude p ∈ C ∞ (U × R).
The case k < −n is easily analyzed: Without loss of generality, suppose z 6= 0 and
show as in the proof of Theorem 8.3 (repeated partial integration) that K(x, z) :=
p(x, ξ) d̄ξ is actually C ∞ for z 6= 0, and λ = 0. K(x, z) is continuous even
R ihz,ξi
e
at z = 0, and differentiable there for k sufficiently negative. In the case k ≥ −n,
formally write the integral for K(x, z) as in the case k < −n where λ = 0. The
integral need not converge, since one no longer has the estimate

|p(x, ξ)| ≤ (1 + |ξ|)−n−1 .

Hence, insert the factor (= 1)


2 2
(1 + |ξ| )λ (1 + |ξ| )−λ
Pn
in the integrand. Then note that for the usual Laplace operator ∆z = − j=1 Dj2
(Dj = −i∂/∂zj ), we have
2 λ
eihz,ξi (1 + |ξ| )λ = (1 − ∆z ) eihz,ξi .

If λ > k + n, the case k ≥ −n reduces to the case λ = 0. Details are found in [323,
p.152].]

Remark 8.8. By Definition 8.1, a canonical pseudo-differential operator, whose


amplitude has compact support in the second variable, is of arbitrarily small order
(“k = −∞”), and so it can be presented as an integral operator with C ∞ weight
function (i.e., a smoothing operator).

Remark 8.9. Contrary to the classical notation for integral operators (where
the singularities of the weight function lie on the diagonal of U × U ), we write
K(x, x − y) instead of K(x, y) under the integral, and by this artifice obtain a
weight function K(x, z) that is singular only at z = 0.
214 8. PSEUDO-DIFFERENTIAL OPERATORS

3. Principally Classical Pseudo-Differential Operators


In order to adapt the theory of pseudo-differential operators to our problem
of treating elliptic differential equations, first on closed manifolds and then on
bordered domains, we must solve two problems.
Task 1 — Homogeneous Principal Symbol. Instead of the weight function
K or the amplitude p, we require the notion of principal symbol, a sort
P of homoge-
neous main part of the amplitude. For a differential operator P = |α|≤k aα Dα ,
the amplitude was the polynomial function (Exercise 8.5a)
X
p(x, ξ) = aα (x)ξ α , x ∈ U, ξ ∈ Rn .
|α|≤k

From this, the principal symbol (or shortly “symbol”) σ(P )(x, ξ) of P was taken
to be the homogeneous polynomial in ξ of order k obtained by taking the sum only
over the terms of highest order (|α| = k); see Chapter 5. This process of separation
does not carry over to the amplitude of an arbitrary canonical pseudo-differential
operator. Thus, we make the following four assumptions about the amplitude p of
a pseudo-differential operator of order k ∈ Z:
Assumptions 8.10. (i) For each compact subset K ⊂ U and multi-
indices α, β ∈ Zn+ , there is a C ∈ R such that for all x ∈ K and ξ ∈ Rn ,
we have
k−|α|
Dxβ Dξα p(x, ξ) ≤ C(1 + |ξ|) .
p(x,λξ)
(ii) The limit σk (p)(x, ξ) := limλ→∞ λk
exists for all x ∈ U and ξ ∈
Rn \ {0}. (
0, for |ξ| small,
(iii) For some cut-off function χ ∈ C ∞ (Rn ) with χ(ξ) =
1, |ξ| ≥ 1,
p(x, ξ)−χ(ξ)σk (p)(x, ξ) is the amplitude of a canonical pseudo-differential
operator of order k − 1.
(iv) p(x, ξ) has compact support in the variable x.
Note that (i ) is just a repetition of (8.2). For us, conditions (iii ) and (iv )
serve only a technical purpose, since we then obtain convergence of integrals and
estimates more easily (e.g., see the above hint to Exercise 8.7b). Actually, one
can forgo these conditions and, as in Theorem 8.14 (p. 217), go over to a Fourier
integral operator with a three-slot amplitude. For applications, we must drop these
further assumptions, and we do so for the additional reason that we define our
global pseudo-differential operators so that they possess amplitudes with compact
support only in their localized form (see below).
In contrast to the canonical pseudo-differential operators, whose amplitudes
only satisfy the estimate (i ), we now say that P is a (principally classical)
pseudo-differential operator (with compact support), if the amplitude p of P
meets all four conditions (i ) – (iv). We shall write P ∈ Lkpc (U ), where the acronym
“pc” stands for principally classical in accordance with one branch of modern
literature, see [73, Section 2.3].
Exercise 8.11. a) As mentioned in Remark 7.6, p.197, it is common to write
1/2
shortly hxi := 1 + |x|2 for x ∈ Rn . For k ∈ N, show that the amplitude
p(x, ξ) := hξi meets all four conditions (i ) – (iv), whence P := Op(p) ∈ Lkpc (Rn ).
k
8.3. PRINCIPALLY CLASSICAL PSEUDO-DIFFERENTIAL OPERATORS 215

b) Show that Lkpc (Rn ) $ Lk (Rn ).


[Hint: For a: Consider the expansion
X m X
ξ 2α ≤ 1 + |ξ|2 = Cm,α ξ 2α for m ∈ N,
|α|≤m |α|≤m

m!
where Cm,α = α!(m−|α|)! .
For b: Consider the amplitude p(x, ξ) := χ(ξ) with χ ∈ C↓∞ (Rn ), where C↓∞ (Rn )
denotes the Schwartz space of rapidly decreasing functions and show Op(χ) ∈
L0 (Rn ).]
Remark 8.12. a) Main stream deals with classical pseudo-differential oper-
ators (written P ∈ CLk (U )). That are operators generated by elements of the
subspace CSk (U × Rn ) ⊂ Sk (U × Rn ) consisting of classical (polyhomogeneous)
symbols. More precisely, an amplitude p ∈ Sk (U × Rn ) belongs to CSk (U × Rn ),
if it admits sequences pk−j ∈ C∞ (U × Rn ), j ∈ Z+ with
(8.3) pk−j (x, rξ) = rk−j pk−j (x, ξ), r ≥ 1, |ξ| ≥ 1,
such that
N
X −1
(8.4) p− pk−j ∈ Sm−N (U × Rn ) for all N ∈ Z+ .
j=0

P
The latter property is usually abbreviated p ∼ pk−j .
j=0
b) Clearly we have CLk (U ) ⊂ Lkpc (U ) ⊂ Lk (U ), more precisely:
(8.5) Lkpc (U ) = CLk (U ) + Lk−1 (U ).
For index theory of elliptic operators, it seems to us that the common restriction
to classical pseudo-differentialSoperators is not necessary. All we need can be done
within the wider space L•pc = Lkpc .
c) The preceding Exercise 8.11b shows that the full class L• of canonical pseudo-
differential operators is technically more convenient than the principally classical or
classical classes. However, canonical pseudo-differential operators carry the topo-
logical handicap that a principal symbol can not always be defined in a meaningful
way.
Task 2 — Manifolds and Coordinate Change. Our second task consists of
defining pseudo-differential operators on a paracompact C ∞ manifold X. We begin
with the scalar case. Thus, consider a linear map P : C0∞ (X) → C ∞ (X), where
C0∞ (X) again denotes the space of complex-valued C ∞ functions with compact
support. (We will consider operators on sections of vector bundles below in Exercise
8.21, p. 225). For each local coordinate system κ : U → Rn with U open in X, P
yields a local operator

−1 ∞ u ◦ κ, on U,
Pκ u := P (u ◦ κ) ◦ κ , u ∈ C0 (κ(U )), where u ◦ κ :=
0, on X \ U.
Definition 8.13. P : C0∞ (X) → C ∞ (X) is called a (principally classical)
pseudo-differential operator of order k on X, if Pκ is a (principally classi-
cal) pseudo-differential operator (with compact support) for all C ∞ charts κ with
216 8. PSEUDO-DIFFERENTIAL OPERATORS

relatively compact image. We write

P ∈ Lkpc (X).

The definition seems to be analogous to the introduction of differential opera-


tors on manifolds. Actually, the situation here is different and more complex, since
pseudo-differential operators do not need to be local (see the warning of Remark
8.6, p.213), while differential operators may actually be characterized by their local-
ity (i.e., supp P u ⊆ supp u); see Exercise 5.2, p. 136 and Remark 5.3. In particular,
we have the following problems:
1. How invariant is the definition of Lkpc (X)? Must one actually show that
the induced local operators are pseudo-differential operators for all charts, or is it
enough to check this for an atlas? This difficulty lies in the fact that the formation
of the local operators is not transitive; i.e., in general, one obtains two different
operators, if one first restricts a chart κ on U ⊆ X to an open subset U 0 ⊆ U
obtaining P(κ|U 0 ) and then considers the restriction P to C0∞ (κ(U 0 )). However, it
turns out that the difference is a smoothing operator of the simple form treated in
Exercise 8.5b.
2. With a differential operator P , the amplitude p(x, ξ) and the principal
symbol σ(P )(x, ξ) can be obtained intrinsically from the action of the operator,
without explicitly representing it in terms of local coordinates first. Namely, we
have (see Exercise 6.37, p. 183 above) in a local chart

ik  k

σ(P )(x, ξ)e = P (ϕ − ϕ(x)) u (x), with dϕx = ξ, u(x) = e,
k!
and trivially
 
p(x, ξ) = e−ihx,ξi P ψeihx,ξi (x) ,

where ψ ∈ C0∞ (X) with ψ = 1 in a neighborhood of x.


For pseudo-differential operators (in general, k is not positive) the first formula
does not make sense, and there is no known simple fully invariant formula for
the symbol of a pseudo-differential operator; the second formula holds only in an
approximate sense (e.g., see [323, p.152f]); the amplitude of a pseudo-differential
operator is not unique, but is only asymptotically determined by the operator.
Thus, the task of defining a global symbol (for pseudo-differential operators defined
on the whole manifold X) lies before us now. (Later, as announced before, we
shall give a genuinely global definition of the principal symbol and the total symbol
(amplitude), see Section 8.5 below.)
3. For this, we investigate the behavior of the local operators and their sym-
bols under a coordinate change, and determine the transformation rule in order to
obtain a global symbol. These calculations are somewhat lengthy, since under a
coordinate change, the phase hx, ξi and the amplitude p(x, ξ) cannot be directly
expressed in the form hy, ηi and q(y, η) in the new coordinates. By passing over to
an apparently larger operator class (the so-called Kuranishi Trick, see also [187,
p.34f]), one can drastically simplify these computations (Theorem 8.19, p. 223), as
well as the derivation of the composition rules, the formula for the symbol of the
adjoint operator (Theorem 8.27, p. 229), and the multiplicative properties under
tensor product (see [328, p.206-209] and [220, p.96]).
8.3. PRINCIPALLY CLASSICAL PSEUDO-DIFFERENTIAL OPERATORS 217

The Kuranishi Trick. The following theorem is our entrance ticket to the
micro-local analysis of pseudo-differential operators on manifolds. It is also of
independent interest.
Theorem 8.14 (M. Kuranishi, 1969). Let U ⊆ Rn be open and k ∈ Z. Let
Q be an operator of the form
(8.6) Z Z
eiϕ(x,y,ξ) q(x, y, ξ) u(y) d̄y d̄ξ, x ∈ U, u ∈ C0∞ (U ) =: Op(q)u (x),

(Qu)(x) =
Rn U
where the phase function ϕ is C ∞ and real-valued on U × U × Rn , and linear in
the variable ξ with
∂ϕ ∂ϕ
(8.7) ∂ξ1 (x, y, ξ) = ··· = ∂ξn (x, y, ξ) =0 for ξ 6= 0 ⇔ x = y,
and for each fixed x (resp. y) ϕ is without critical points (y, ξ) (resp. (x, ξ)). In
other words, for all (x, y, ξ) ∈ U × U × (Rn \ {0}),
 
(dξ ϕ)(x,y,ξ) = 0 ⇔ x = y, d(y,ξ) ϕ (x,y,ξ) 6= 0 and d(x,ξ) ϕ (x,y,ξ) 6= 0.
Moreover, we assume that the amplitude q ∈ C ∞ (U × U × Rn ) meets the follow-
ing conditions (analogous to the conditions (i)–(iv) on the amplitude of a pseudo-
differential operator, p. 214):
(i0 ) For each compact subset K ⊂ U and multi-indices α, β, γ ∈ Zn+ ,
there is a Cα,β,γ ∈ R such that for all x, y ∈ K and ξ ∈ Rn , we have
Dξα Dxβ Dyγ q(x, y, ξ) ≤ Cα,β,γ (1 + |ξ|)k−|α| .
q(x,y,λξ)
(ii0 ) σk (q)(x, y, ξ) := limλ→∞ λk
exists for ξ 6= 0 and
(x, y) ∈ U × U .
0 ∞ n 0, for |ξ| small,
(iii ) For some cut-off function χ ∈ C (R ) with χ(ξ) =
1, |ξ| ≥ 1,
q(x, x, ξ) − χ(ξ) σk (q)(x, x, ξ) is the amplitude of an element of Lk−1
pc (U ).
(iv 0 ) q(x, y, ξ) has compact support in the x and y variables.
Then Q can be written as a pseudo-differential operator (with compact support) of
order k, i.e., Q ∈ Lkpc (U ).
Note . Recall that we in this book deal mostly with the principal symbol
of pseudo-differential operators and write shortly “symbol” and σ(x, ξ) when we
mean “principal symbol” and “σk (x, ξ)”. In some places, however, we wish to mark
the order of the operator in the notation for the symbol. That is the case in the
preceding assumption (ii’).
Remark 8.15. a) These operators are special types of Fourier integral oper-
ators. The term is due to L. Hörmander who in a series of papers developed a
precise theory for them, which can be applied to the general theory of partial dif-
ferential equations. In doing so, he could resort to ideas of the Dutch mathemati-
cian and physicist Christian Huygens (1629-1695) and of the Russian mathe-
maticians Vladimir Igorevich Arnold, Yuriy Vladimirovich Egorov, and
Venyaminovich Clavdiy Maslov, who dealt with fundamentals of geometric
optics and the formalization of its more or less intuitive methods (aggregation prin-
ciple, quantization, etc., see also our Remark 8.4, p.211).
b) Below, in Step 1 of the proof of the preceding theorem, we shall address the
delicate convergence questions related to the integral in (8.6). Such integrals are
called oscillatory integrals. More precisely, let U ⊂ Rm and ϕ = ϕ(x, θ) ∈
218 8. PSEUDO-DIFFERENTIAL OPERATORS

C ∞ (U × (Rn \ {0})) be real valued (or, at least of nonnegative imaginary part)


with ϕ(x, λθ) = λϕ(x, θ) for λ > 0 and dϕ 6= 0 on all U × (Rn \ {0}). Let
p ∈ C ∞ (U × Rn ) satisfy the asymptotic estimate (8.2) introduced on p.210 for
fixed order k ∈ R. Then the oscillatory integral
Z
I(p, ϕ)(x) := eiϕ(x,θ) p(x, θ)dθ
Rn \{0}

belongs to C j (U ), if k is sufficiently negative, more precisely if k + j < −n. That is


not very deep. The interesting aspect of oscillatory integrals is that they also give
a meaning as distributions even if the order k of p is large. We shall explain that
below in Step 1 for our special case where we replace U by U × U , m by 2n, ϕ(x, θ)
by φ(x, y, θ) and p(x, θ) by q(x, y, θ). A general and more systematic treatment of
oscillatory integrals can be found in [187, Chapter 1] and [190, p.168].
c) Obviously, every pseudo-differential operator with amplitude p(x, ξ) can be writ-
ten as a Fourier integral operator with q(x, y, ξ) := p(x, ξ) and ϕ(x, y, ξ) := hx − y, ξi.

Before the proof of the preceding theorem, we shall emphasize that an ampli-
tude is not determined from a given pseudo-differential operator.

Example 8.16. Let U ⊂ Rn , φ(x, y, ξ) := hx − y, ξi, a ∈ C0∞ (U ) \ {0} and


1 ≤ j ≤ n. Then the amplitude
q(x, y, ξ) := ξj a(y) − a(x)ξj − Dxj a(x)
meets the conditions (i’)-(iv’) of Theorem 8.14 for k = 1, but Op(q) is just the zero
operator.

Remark 8.17. For pseudo-differential operators, the amplitude (also called


the total symbol or the dequantization) p(x, ξ) is neither uniquely determined from
Op(p) in general, since a perturbation by an infinitely smoothing amplitude can
generate the same operator. For the precise results we refer to [190, Proposition
7.8]. There is a vast literature on properly supported pseudo-differential op-
erators, which have the nice (and somewhat misleading) property that they have
uniquely determined amplitudes, see [187, Chapter 3], [190, Section 7.2], [224,
Section 18.1], [391, Section 3.1], [410, Section II.3]. The sad fact is the following:
If you build on coordinates and coordinate shifts, it seems that only the symbol (i.e.,
what is also called the principal symbol) has a geometric meaning. That is well and
easily defined for differential and pseudo-differential operators. It suppresses sub-
stantial parts of the underlying operator, but is sufficient for finding parametrices
and calculating the index of elliptic operators, as we shall show below in Chapter 9.
The good news is that there is a coordinate free definition of Fourier transformation
and pseudo-differential operators leading to a one-to-one quantization p 7→ Op(p),
see below Section 8.5. For the full embedding proof of the Atiyah-Singer Index
Theorem, mastering our coordinate free introduction of pseudo-differential oper-
ators and the one-to-one correspondence between operator and amplitude will be
decisive. The nongeometric constructions of the analysis main-stream do not suf-
fice. However, also our geometric construction below depends on choices (of metrics
and connections, as mentioned already in Remark 8.4a, p. 211). So, it supports a
powerful and transparent proof of the Index Theorem, but it does not offer a formal
solution to the mysteries of dequantization.
8.3. PRINCIPALLY CLASSICAL PSEUDO-DIFFERENTIAL OPERATORS 219

Proof of Theorem 8.14. Step 0: First we show the convergence of the


integral defining Qu(x). As the integral stands, it is only absolutely convergent
when the order k of Q is very negative. However, the following integral is absolutely
convergent for sufficiently large r ∈ N:
Z Z
r
eiϕ(x,y,ξ) t L (q(x, y, ξ) u(y)) d̄y d̄ξ, x ∈ U,
Rn U
t
where L denotes the formal adjoint of the operator
(dy ϕ) · dy + (dξ ϕ) · dξ
L := −i 2 2 2 : C ∞ (U × Rn ) −→ C ∞ (U × Rn ),
|dy ϕ| + |ξ| |dξ ϕ|
which (under the assumptions) is a well-defined differential operator such that
Leiϕ(x,·,·) = eiϕ(x,·,·) .
In this way, the original integral can be replaced by an absolutely-convergent in-
tegral via repeated integration by parts. (Note that owing to the linearity of ϕ in
2
ξ and the assumption d(y,ξ) ϕ 6= 0, the term |dy ϕ| in the denominator of L grows
2
like |ξ| for fixed x and y, while u(y) has compact support in y). Hence Qu(x) is
well-defined. Without difficulty it follows that Q is a linear map from C0∞ (U ) to
C ∞ (U ).
Step 1: We now show that on a neighborhood Ω of the diagonal in U × U , one
can find a C ∞ map ψ : Ω → GL(n, R) (see Figure 8.1) such that for all (x, y) ∈ Ω
and ξ ∈ Rn , we have
ϕ(x, y, ψ(x, y)ξ) = hx − y, ξi .
By the assumption that ϕ is linear in ξ, we can write ϕ in the form
Xn
∂ϕ
ϕ(x, y, ξ) = ϕj (x, y)ξj , where ϕj (x, y) := ∂ξ j
(x, y, ξ) .
j=1
We now show that the functional matrix

Figure 8.1. Finding ψ on a neighborhood Ω of the diagonal U × U

∂ϕ1 ∂ϕ1
···
 
∂x1 (x, y) ∂xn (x, y)
F (x, y) := 
 .. .. 
. . 
∂ϕn ∂ϕn
∂x1 (x, y) · · · ∂xn (x, y)
220 8. PSEUDO-DIFFERENTIAL OPERATORS

is invertible for x = y. However, for all y ∈ U , ϕ|U ×{y}×(Rn \{0}) has no critical
points, and so
Xn Xn
∂ϕ ∂ϕ
ξ 6= 0 ⇒ ∂xj (x, y, ξ) + ∂ξj (x, y, ξ) 6= 0.
j=1 j=1
Pn ∂ϕ
By (8.7) j=1 ∂ξj (x, y, ξ) = 0 ⇔ x = y. Thus, if x = y, then for any ξ 6= 0 there
is some j ∈ {1, ..., n} such that
Xn  Xn
∂ϕ ∂ϕ
0 6= ∂x (x, y, ξ) = ∂
∂xj (x, y, ξ) ξk = ξk ∂ϕ
∂xj (x, y) .
k
j k=1 ∂ξk j=1

Hence, the matrix F (x, x) has a trivial kernel and must be invertible. By assump-
tion ϕj (x, x) = 0. Thus we have the short Taylor expansion
Xn
ϕj (x, y) = ϕµj (x, y)(xµ − yµ ),
µ=1

where the functions ϕµj (x, y) are C ∞ near the diagonal of U × V . The matrix
ϕ(x, y) = (ϕµj (x, y)) is invertible in a neighborhood Ω of the diagonal, because
ϕ(x, x) = t F (x, x). Since
Xn
ϕ(x, y, ξ) = ϕj (x, y)ξj
j=1
Xn Xn
= (xµ − yµ ) ϕµj (x, y)ξj = hx − y, ϕ(x, y)ξi ,
µ=1 j=1

we have the desired property


ϕ(x, y, ψ(x, y) ξ) = hx − y, ξi ,
−1
where ψ(x, y) := ϕ(x, y) for (x, y) ∈ Ω. For later, we note that
 
2
ϕ(x, x) = t F (x, x) = ϕ00xξ (x, y, ξ) y=x := ∂x∂i ∂ξ
ϕ
j
(x, y, ξ) ,
y=x

whence in particular,
1
det ψ(x, x) = .
det ϕ00xξ (x, y, ξ)
y=x

Step 2: Now we eliminate the phase function ϕ. For this, we assume that for all
ξ,
supp q(·, ·, ξ) ⊆ Ω.
Then, for all x ∈ U , the integration domain in the formula for (Qu)(x) (see above)
is small enough so that the change of variable transformation ξ = ψ(x, y)θ can be
applied to obtain
Z Z
(8.8) (Qu)(x) = eihx−y,θi q(x, y, ψ(x, y)θ) |det ψ(x, y)| u(y) d̄y d̄θ.
U Rn
The new amplitude
(x, y, θ) 7→ a(x, y, θ) |det ψ(x, y)|
with a(x, y, θ) := q(x, y, ψ(x, y)θ) then automatically satisfies the conditions (ii0 ),
(iii0 ), and (iv 0 ). To check (i0 ), we must calculate: By the chain rule, we obtain the
formula (for z := (x, y) ∈ R2n )
 
Id 0
(∂z a(z, θ) , ∂θ a(z, θ)) = (∂z q(z, ψ(z)θ), ∂ξ q(z, ψ(z)θ)) ,
ψ 0 (z)θ ψ(z)
8.3. PRINCIPALLY CLASSICAL PSEUDO-DIFFERENTIAL OPERATORS 221

where ∂z denotes the partial derivatives with respect to the first 2n variables and
∂θ or ∂ξ denote those with respect to the last n variables. Hence, we have
∂a Xn ∂q
(z, θ) = (z, ψ(z) θ) ψ ij (z)
∂θi j=1 ∂ξj
k−1
≤ nCk (1 + |ψ(z)θ|) max ψ ij (z) ,
i,j,z

0
where Ck is a real number from assumption (i ) for q when z varies within a compact
domain K ⊆ Ω ⊆ U × U ⊆ R2n . Since we can find positive constants C1 and C2
with
C1 |θ| ≤ |ψ(z) θ| ≤ C2 |θ| , for all z ∈ K and θ ∈ Rn ,
we finally have the estimate
∂a ek (1 + |θ|)k−1 , where C
(z, θ) ≤ C ek ∈ R.
∂θi
Similarly one can obtain estimates for the higher derivatives, wherein the factor
|det ψ(x, y)| of the amplitude in (8.8) is irrelevant.
Step 3: Now, we consider the general case, where the support supp q(·, ·, ξ) is
not necessarily contained in Ω. In our applications of the theorem of Kuranishi
(see Theorem 8.19, p. 223) we are only concerned with a local argument; i.e., we
can manage with the case treated in step 2. We will therefore be brief in showing
that in general we may assume the first case without loss of generality. We choose
a nonnegative C ∞ function χ on U × U having support in Ω and being equal to 1 in
a neighborhood of the diagonal. Then Q can be written as the sum of two Fourier
integral operators, where one has the amplitude χq of the form in step 2, and the
other has the form
Z Z
(Ru)(x) = eiθ(x,y,ξ) r(x, y, ξ) u(y) d̄y d̄ξ,
U Rn

where r = (1 − χ)q is a C function vanishing in a neighborhood of the diagonal
in U × U . Just as in Exercise 8.7, it follows that R can be written in the form
Z
λ
(Ru)(x) = K(x, y)(1 − ∆) u(y) dy,
U

where the weight function K is C ∞ off the diagonal of U × U according to Ex-


ercise 8.7, and vanishes in a neighborhood of the diagonal by construction; K is
then C ∞ everywhere. By Exercise 8.5b, R can then be written as a canonical
pseudo-differential operator; the corresponding conditions (ii ), (iii ), and (iv ) are
met without difficulty.
Step 4: Without loss of generality, we may now assume that the operator Q
is given in the form
Z Z
(Qu)(x) = eihx−y,ξi q(x, y, ξ)u(y) d̄y d̄ξ
Rn Rn
Z Z 
ihx,ξi −ihy,ξi
= e e q(x, y, ξ)u(y)d̄y d̄ξ,
Rn Rn

where the braces enclose the Fourier transform of the product function q(x, y, ξ)u(y)
(extended by 0 values outside U ) or equivalently, by Appendix A, the convolution
222 8. PSEUDO-DIFFERENTIAL OPERATORS

of the Fourier transforms in the variable ξ


Z Z
−ihy,ξi
e q(x, y, ξ)u(y) d̄y = qb(x, ξ − η, ξ)b
u(η) d̄η,
Rn Rn

where qb(x, ·, ξ) is the Fourier transform of y 7→ q(x, y, ξ). Inserting the factor
eihx,ηi e−ihx,ηi and reversing the order of integration, we obtain
Z Z 
(Qu)(x) = eihx,ηi eihx,ξ−ηi qb(x, ξ − η, ξ)d̄ξ ub(η) d̄η
Rn Rn
Z
= eihx,ηi p(x, η) u
b(η) d̄η,
Rn

where (by a change of variables ζ = ξ − η)


Z
p(x, η) := eihx,ζi qb(x, ζ, ζ + η) d̄ζ.
Rn

We now show that p(x, η) is actually the amplitude of a pseudo-differential operator


(the support is trivially compact by construction, whence (iv ) already holds):
(i): Let α and β be multi-indices and let x range over a compact subset of Rn .
For the estimation of Dxβ Dηα p(x, η) , we first note that by the Fourier multiplication
rule (Appendix A, Exercise A.5b, p. 711), we have
Z
k−α
Dxβ Mθγ Dηα qb(x, θ, η) = e−ihy,ζi Dxβ Dyγ Dηα q(x, y, η) d̄y ≤ c(1 + |η|) ,
Rn

where γ is a further multi-index and Mθγ is multiplication by θγ = θ1γ1 · · · θnγn ; the


inequality follows from the assumptions (i0 ) and (iv 0 ) for q. Hence, for each positive
ν, we have
k−α −ν
(8.9) Dxβ Dηα qb(x, θ, η) ≤ c0 (1 + |η|) (1 + |θ|) .
Thus, by definition of p and by means of differentiation under the integral, we get
k−α
Dxβ Dηα p(x, η) ≤ c00 (1 + |η|) .
(ii): By the Mean-Value Theorem, we obtain for suitable ζ0 between 0 and ζ
Z  X 
p(x, η) = eihx,ζi qb(x, ζ, η) + eihx,ζi Dηα qb(x, ζ, η + ζ0 )ζ α d̄ζ
Rn |α|=1

= q(x, x, η) + a correction term E(x, η) .


We have already seen (in (8.9)) that
−ν
Dηα qb(x, ζ, η + ζ0 ) ≤ Cν (1 + |η + ζ0 |)k−1 (1 + |ζ|) ,
for arbitrarily large ν. Since |ζ0 | < |ζ|, we have
k−1 −ν+k−1
Dηα qb(x, ζ, η + ζ0 ) ≤ c0 (1 + |η|) (1 + |ζ|) .
k−1
Integrating with respect to ξ, we find |E(x, η)| ≤ C(1 + |η|) . Thus,
k−1
E(x, λη) C(1 + |λη|)
lim ≤ lim = 0, and so
λ→∞ λk λ→∞ λk

p(x, λη) q(x, x, λη)


σk (p)(x, η) = lim k
= lim = σk (q)(x, x, η).
λ→∞ λ λ→∞ λk
8.3. PRINCIPALLY CLASSICAL PSEUDO-DIFFERENTIAL OPERATORS 223

k−1
(iii ): It follows easily from the estimate E(x, η) ≤ C(1 + |η|) and the cor-
responding assumption (iii0 ) for q(x, x, η), that p(x, η) − χ(η) σk (p)(x, η) is the
amplitude of a canonical pseudo-differential operator of order k − 1. 
Remark 8.18. We note that, from the preceding constructive proof of Masa-
take Kuranishi, a simple formula for the symbol follows from steps 2 and 4:
σk (q)(x, x, ψ(x, x) η)
σk (p)(x, η) = ,
det ϕ00x,ξ (x, y, ξ)
y=x

where ψ(x, x) is the inverse of the functional matrix (also denoted F (x, x)) in step 1,
namely  
∂2ϕ
ϕ00x,ξ (x, y, ξ) y=x
= ∂xi ∂ξj (x, y, ξ) .
y=x
The term Ru in step 3 with amplitude r does not affect the symbol formula, since
r(x, x, η) = 0 for all x and η, whence σk (r)(x, x, η) = 0.
Coordinate Change and Pseudo-Differential Operators on Manifold.
Now we investigate the behavior of pseudo-differential operators under a coordinate
change, exploiting the Kuranishi Trick of transgressing to Fourier integral operators,
similarly in [187, p.34f]:
Theorem 8.19. Let κ : U → V be a diffeomorphism between relatively compact
open subsets of Rn . If P is a pseudo-differential operator (with compact support)
of order k ∈ Z on V , then the transported operator
Pκ (u) := P (u ◦ κ−1 ) ◦ κ, u ∈ C0∞ (U )
is a pseudo-differential operator (with compact support) of order k over U . If p and
q are amplitudes for P and Pκ resp., then their symbols are related by
0
σk (q)(x, ξ) = σk (p)(κ(x) , (t κ (x))−1 ξ), x ∈ U, ξ ∈ Rn \ {0} ,
where
∂κ1 ∂κn
···
 
∂x1 ∂x1
t 0
κ (x) = 
 .. .. 
. . 
∂κ1
· · · ∂κ
∂xn ∂xn
n

is the transpose of the functional matrix of κ at x.


Proof. Suppose that the operator P is of the form
Z
(P v)(y) = eihy,ηi p(y, η) vb(η) d̄η
n
ZR Z
= eihy−θ,ηi p(y, η) v(θ) d̄θ d̄η, v ∈ C0∞ (V ) , y ∈ V,
Rn Rn
−1
whence for v = u ◦ κ and y = κ(x), u ∈ C0∞ (U ) and x ∈ U :
Z Z
eihκ(x)−θ,ηi p(κ(x), η) u κ−1 (θ) d̄θd̄η

Pκ (u)(x) =
n n
ZR ZR
= eihκ(x)−κ(ξ),ηi p(κ(x), η) |det κ0 (ξ)| u(ξ) d̄ξd̄η
Rn Rn
by means of the change of variable κ(ξ) = θ. The transported operator is then a
Fourier integral operator with the phase function ϕ(x, ξ, η) := hκ(x) − κ(ξ), ηi and
224 8. PSEUDO-DIFFERENTIAL OPERATORS

amplitude q(x, ξ, η) := p(κ(x), η) |det κ0 (ξ)|. Since ϕ and q meet the hypotheses of
the Theorem of Kuranishi (Theorem 8.14, p. 217), we have
Z Z
Pκ (u)(x) = eihx−ξ,ηi p(κ(x), ψ(x, ξ) η) D(x, ξ) u(ξ) d̄ξd̄η,
Rn Rn
where
D(x, ξ) := |det κ0 (ξ)| |det ψ(x, ξ)|
and ψ(x, ξ) is the matrix-valued function constructed in step 1 of the proof of
Theorem 8.14; in particular,
 
∂κ 0
ψ(x, x)−1 = ϕ00x,η (x, ξ, η) ξ=x = ∂xνj = t κ (x)
and D(x, x) = 1.
By the Theorem of Kuranishi, Pκ is a pseudo-differential operator for which
we derived an explicit formula for the amplitude in the above proof. For the symbol,
we have
σk (q)(x, η) = σk (e
q )(x, x, η),
where
qe(x, ξ, η) = p(κ(x), ψ(x, ξ) η) D(x, ξ).
0
Since ψ(x, x)−1 = t κ (x) and D(x, x) = 1, we then obtain
 −1
0
σk (q)(x, η) = σk (p)(κ(x) , t κ (x) η). 

Exercise 8.20. Let X be a (paracompact) C ∞ n-manifold and k ∈ Z.


a) Show that the space Lkpc (X), defined by localization at the beginning of this
section just before Definition 8.13, coincides with the space of pseudo-differential
operators of order k on X, when X is a bounded open subset of Rn .
b) Define a canonical vector space structure on Lkpc (X).
c) Show Lkpc (X) ⊂ Lk+1pc (X).
d) Let κ : U → Rn , U open in X, be a local coordinate system for X, f ∈ C0∞ (U ),
and P ∈ Lkpc (X). Show that if f ∈ C ∞ (X) is identically 1 in a neighborhood of
κ−1 (x), then the formula
 
(x, ξ) 7→ qf (x, ξ) := e−ihx,ξi P (f (·) eihκ(·),ξi ) (κ−1 (x)); x ∈ κ(U ), ξ ∈ Rn ,
defines the amplitude of a pseudo-differential operator of order k on κ(U ), and that
if p is the amplitude for the localized operator Pκ , then
σk (qf ) (x, ξ) = σk (p)(x, ξ).

e) Set Smblk (X) := Smblk (CX , CX ); see our definition in Equation (6.29), p. 184,
in Chapter 6 above. Thus, s ∈ Smblk (X) ⇔ s : T̊ ∗ X → C with s(x, λv) = λk s(x, v)
for all x ∈ X and v ∈ Tx∗ X, v 6= 0. Show that the linear map
σk : Lkpc (X) → Smblk (X)
is well defined and coincides with the earlier definition (see Exercise 6.37, p.183)
on Diff k (X) ⊂ Lkpc (X).
[Hint: For a: Theorem 8.19.
For b: Proceed by using the vector space structure of C ∞ (X).
For c: Use amplitude estimates.
8.3. PRINCIPALLY CLASSICAL PSEUDO-DIFFERENTIAL OPERATORS 225

For d: Characterize Lkpc (X) within the space of linear operators from C0∞ (X) to
C ∞ (X) in the following way: P ∈ Lkpc (X) if and only if the associated qf is the
amplitude of a pseudo-differential operator of order k on an open subset of Rn for
all local coordinate systems and cut-off functions f . For details of the computation,
see [220, p.112] and [323].
For e: Recall that T̊ ∗ X denotes the symplectic cone T ∗ X \ X that consists of the
punctured cotangent spaces. It remains only to show that the locally well defined
symbol in d) transforms correctly under a coordinate change, so that it forms global
homomorphism from T̊ ∗ X × C to T̊ ∗ X × C (which is homogeneous of degree k in
the cotangent vectors. For this, check that the transformation rule in Theorem 8.19
can be written in the form
σ(q)(x, κ
e(η)) = σk (p)(κ(x) , η),
where η lies in T ∗ (Rn )κ(x) , which is the space of covectors at the point κ(x) canon-
ically identified with Rn , and κ e(η) is the pull back covector via κ; see (6.1), p.165,
in Appendix B; further details are found in [31, p.404-407].]
Exercise 8.21. Define the space Lkpc (E, F ) when E and F are complex vector
bundles over the C ∞ manifold X, and show the existence of a canonical linear map
σk : Lkpc (E, F ) → Smblk (E, F ).
[Hint: Represent an operator P : C0∞ (E) → C ∞ (F ) locally; i.e., choose a chart
κ : U → Rn , U ⊆ X open, and κ(U ) relatively compact, and trivializations E|U ∼=
U × CN and F |U ∼ = U × CM as a M × N matrix of pseudo-differential operators
(with compact support) of order k.]
Remark 8.22. a) There is a slight ambiguity in our definition of the symbol
space Smblk (E, F ). As explained in our defining Equation 6.29, p.184, we require
homogeneity
(8.10) σ(x, rξ) = rk σ(x, ξ) for x ∈ X, ξ ∈ Tx∗ X \ {0} and r > 0
for σ ∈ Smblk (E, F ). Our Assumption 8.10(ii), p.214 ensures (8.10). Clearly, homo-
geneity and smoothness at ξ = 0 contradict each other except for monomials. Our
convention is that Smblk (E, F ) denotes the space of homogeneous bundle homomor-
phisms of the lifted bundles π ∗ E, π ∗ F , where π : T̊ ∗ X → X and T̊ ∗ X = T ∗ X \ X,
i.e., we exclude ξ = 0. In various applications, however, symbols should be smooth
functions, thus the σ(x, ξ) should be smooth everywhere but homogeneous only in
the restricted sense:
(8.11) σ(x, rξ) = rk σ(x, ξ) for x ∈ X, |ξ| ≥ 1 and r ≥ 1
with a suitable Riemannian metric that yields the length of cotangent vectors.
b) In many places, we shall tacitly identify the homogeneous bundle mappings on
T ∗ X \ X by restriction with the smooth sections C ∞ S ∗ X, Hom(ρ∗ (E), ρ∗ (F )) .


Here S ∗ X denotes the sphere bundle of cotangent vectors (relative to a fixedRie-


mannian metric), ρ : S ∗ X → X the natural projection, and Hom ρ∗ (E), ρ∗ (F ) the
bundle of smooth bundle homomorphisms.
Remark 8.23. The definition of Lkpc (E, F ) through localizations is unsatisfac-
tory from a computational point of view. The choice of coordinates is awkward
with its avalanche of subscripts which are frequently unavoidable even in fairly
simple situations. The method is particularly unsatisfactory, when the operator (as
226 8. PSEUDO-DIFFERENTIAL OPERATORS

in Exercise 8.24 below) can be written in closed, global form explicitly and much
more clearly. [220, p.113f] contains the following idea for writing Lkpc (E, F ) by
means of the Kuranishi Theorem directly as a space of Fourier Integral Operators
with phase function ϕ : G → R and amplitude q : G → Hom(E, F ): Let G be a real
vector bundle of fiber dimension n over a neighborhood of the diagonal in X × X,
e.g., G = π ∗ (T ∗ X) where π is the projection π(x, y) := y, and
q(x, y, ξ) ∈ Hom(Ey , Fx ).
It turns out that one can formulate the necessary conditions on ϕ and q directly,
globally and with little difficulty: For example, ϕ is linear in the fibers and the
restriction of ϕ to a fiber has a critical point exactly when the fiber lies above a
point of the diagonal of X × X. Then (loc. cit.) Lkpc (E, F ) consists of all operators
that can be written as the sum of an operator with C ∞ kernel and one of the form
Z
−n
(P e)(x) := (2π) eiϕ(x,y,η) q(x, y, η)e(y) dydη,
T ∗X
where e ∈ C0∞ (E), dydη is the invariant volume element on the cotangent bundle
T ∗ X, and q an amplitude of order k which vanishes for (x, y) outside a small
neighborhood of the diagonal of X × X. By step 4 of the proof of Theorem 8.14, it
follows that
q(x, x, λη)
σk (P )(x, η) = lim , x ∈ X, η ∈ Tx∗ X \ {0} .
λ→∞ λk
We shall devote the whole Section 8.5, p.231ff to the details of a truly global con-
struction of a ‘Bokobza-Haggiag’ total symbol.
Singular Integral Operators. We show that the classical singular integral
operators fit nicely under our heading of principally classical pseudo-differential
operators.
Exercise 8.24. Show that the following singular integral operators are pseudo-
differential operators of order 0 over R or S 1 = R/2πZ 1 and determine their
amplitudes:
a) The Hilbert transform Q : C0∞ (R) → C ∞ (R), defined for u ∈ C0∞ (R) by
Z ∞
−1 −1
Z
u(y) u(y)
(Qu)(x) := (p.v.) dy := lim+ dy.
πi −∞ x − y πi ε→0 |x−y|>ε x − y
b) The projection operator P : C ∞ (S 1 ) → C ∞ (S 1 ), defined by
 imθ
imθ e , for m ≥ 0,
Pe :=
0, for m < 0.
c) The Toeplitz operator
gP + (Id −P ), for g ∈ C ∞ (S 1 ).
1More precisely: Write them as a sum of a pseudo-differential operator of the kind treated
so far and a smoothing operator. Many classical pseudo-differential operators Q are defined as
here via an amplitude q which is homogeneous in the second variable, but has a singularity at the
origin. Through multiplication by a C ∞ function χ which is identically 1 in a neighborhood of ∞,
we obtain a singularity-free amplitude qe(x, ξ) := χ(ξ)q(x, ξ), which defines a pseudo-differential
for Q
e in our (Hörmander’s) sense. Then Q e − Q has an amplitude with compact support and
consequently (with reasoning as in Remark 8.8, p.213) can be represented as an integral operator
with a C ∞ weight function.
8.3. PRINCIPALLY CLASSICAL PSEUDO-DIFFERENTIAL OPERATORS 227

[Hint: For a: Show that


Z ∞
(Qu)(x) = eixξ sign(ξ) u
b(ξ) d̄ξ,
−∞
as follows. We have
u(x − y)
Z Z
u(y)
dy = dy = (u ∗ gε )(x)
|x−y|>ε x−y |y|>ε y
Z ∞
= eixξ u
b(ξ) gbε (ξ) dξ, where
−∞
1

for |x| > ε,
x,
gε (x) :=
0, for |x| ≤ ε,
and for the last equality the well-known convolution formula is used. Thus, it is
natural to try to evaluate the improper integral
Z ∞ −iξt √
e
(p.v.) dt := 2π lim gbε (ξ).
−∞ t ε→0

Now distinguish cases according to the sign of ξ! One obtains


√ Z ∞
sin t
2π lim gbε (ξ) = −2i sign(ξ) lim dt = −πi sign(ξ) ,
ε→0 ε→0 |ξ|ε t
R∞ iz
since 0 sint t dt = π2 via contour integration of the function f (z) := ez , z = t + is,
along the curve shown in Figure 8.2. Compare also [134, p.93 and 150].

t
{R {" " R

Figure 8.2. Contour for integrating the function f

For b: Reduce to a) by means of the formula P u = 21 (u + Hu), where


Z
1 u(z)
(Hu)(eiθ ) := (p.v.) iθ
dz, u ∈ C ∞ (S 1 )
πi S 1 z − e
denotes the Cauchy-Hilbert transformation (on the circle) which is carried over to
the Hilbert transformation Q (on the line R) by means of the Cayley transformation;
see p. 130 in Chapter 4. A more direct way may be found in [44, p.525]. For this,
as in Exercise 8.20d form the expression
X∞ √
qf (x, ξ) := e−ixξ P (f (x)eixξ ) = 2π fb(n − ξ)eix(n−ξ)
n=0
√ X−∞
= f (x) − 2π fb(n − ξ)eix(n−ξ) ,
n=−1
where f ∈ C0∞ (R) with compact support in an interval of length < 2π, so that f
may be regarded as a function on the circle with support in a canonical coordinate
228 8. PSEUDO-DIFFERENTIAL OPERATORS

domain. P∞
Trick: For ξ < 0, estimate n=0 fb(n − ξ)eix(n−ξ) and its derivatives, showing that
−1
as ξ → −∞ they go to 0 faster than any power of |ξ| . Show that for ξ > 0,
P−∞ b
the sum n=−1 f (n − ξ)eix(n−ξ) has the corresponding property. By the hint for
Exercise 8.20d, one is done and obtains

1, for ξ > 0,
σ0 (P )(x, ξ) =
0, for ξ < 0.
For c: Reduce to b). Note that in the notation of Chapter 4
Tg , on C ∞ S 1  ∩ H0 ,
 
gP + (Id −P ) =
Id, on C ∞ S 1 ∩ H0⊥ (in L2 (S 1 )),
where Tg is the Wiener-Hopf operator induced by g with index Tg = −W (g, 0), if g
is nowhere zero on S 1 . In particular, by Exercise 1.5 (p. 5), index(gP + Id −P ) =
−W (g, 0).]

4. Algebraic Properties and Symbolic Calculus


Here we will show that all C ∞ symbols are obtained as symbols of pseudo-
differential operators. Moreover, we explain how one can calculate with symbols
instead of operators, using some simple rules. In a wording borrowed from algebraic
topology, the symbolic calculus is a functor from the category of infinite-dimensional
function spaces and systems of linear differential and pseudo-differential equations
to parameter dependent linear algebra in finite dimensions.
Theorem 8.25. Let E and F be complex vector bundles over a C ∞ manifold
X. Then there is an exact sequence
σ
(8.12) 0 −→ CLk−1 (E, F ) ,−→ Lkpc (E, F ) −→
k
Smblk (E, F ) −→ 0 ,
where σk (P ) denotes the principal (homogeneous leading) symbol of P ∈ Lkpc (E, F ).

Recall that Lk−1 (E, F ) (respectively CLk (E, F )) denote the space of (k − 1)th
order canonical pseudo-differential operators from sections of E to sections of F
(respectively kth order classical pseudo-differential operators) and that
(8.13) Lkpc (E, F ) = CLk (E, F ) + Lk−1 (E, F ),
like in Remark 8.12b, p.215.
Proof. Since CLk (E, F ) ∩ Lk−1 (E, F ) = CLk−1 (E, F ), the principal symbol
map σk : Lkpc (E, F ) → Smblk (E, F ) is well defined. Because of the decomposition
(8.13), it only remains to show that the symbol map is surjective.
Let s ∈ Smblk (E, F ). If π : X × X → X is the projection given by π(x, y) = y,
then the pull-back bundle
G := π ∗ (T ∗ X) → X × X
is a real vector bundle of fiber dimension n. In reference to Remark 8.23 after
Exercise 8.21 it suffices to give a phase function ϕ : G → R and an amplitude
a : G → Hom(E, F ) with the properties required by definition (see the conditions
(i0 )–(iv 0 ) in Theorem 8.14, p. 217) such that
a(x, y, η) = s(x, η), x ∈ X, η ∈ Tx∗ X \ {0} .
8.4. ALGEBRAIC PROPERTIES AND SYMBOLIC CALCULUS 229

For ϕ, we choose a real-valued function with ϕ(x, x, η) = 0 and



dϕ(x,x,η) = η ⊕ (−η) ⊕ 0 : T(x,x,η) G → R, where

T(x,y,η) G = Tx X ⊕ Ty X ⊕ Tη (T X) and η ∈ Tx∗ X.
∗ ∗ ∗ ∗ ∗

Such a map ϕ, which is also linear on the fiber and only possesses critical points
over the diagonal of X × X, is locally easy to construct relative  to a chart κ
about the point x. One simply sets ϕ(x, y, η) = κx − κy, κ−1 eη (see (6.1),
p. 165 for the definition of “˜”). A global construction may be carried out using
a partition of unity (see Theorem 6.4, p. 159) in a neighborhood of the diagonal.
For a : G → Hom(E, F ), we choose an arbitrary extension of a(x, x, η) := s(x, η)
to a neighborhood of the diagonal; then we can smoothly extend a by multiplying
by a C ∞ function (with support in the neighborhood) which is identically 1 in
a smaller neighborhood of the diagonal. An extension can be found, since the
diagonal is closed in X × X, and a can be regarded as a section of the lift of the
bundle Hom(E, F ) by means of the projection (x, y, η) 7→ (x, y). Compare with
step 1 of the proof of Theorem B.7 (p. 716) in Appendix B, in connection with
the Whitney Approximation Theorem; e.g., [314, p.34f] or [97, 14.8]. With this,
the proof (which strongly depends on the Theorem of Kuranishi, Theorem 8.14,
p.217) is done. A direct proof can be found in [434, p.134f]. 
Remark 8.26. We can fix a right inverse
Op : Smblk (E, F ) −→ Lkpc (E, F )
of σk , obtained by patching together the local Op-maps (8.1) via a fixed partition
of unity.
Theorem 8.27. The direct sum L•pc (E, F ) :=
L k
k Lpc (E, F ) forms a graded
algebra via composition, and is closed under the operation of taking formal adjoints.
For the symbols, we have the following calculation rules:
(a) If E, F , and G are complex vector bundles over the C ∞ manifold X, and P ∈
Lkpc (E, F ) and Q ∈ Ljpc (F, G), then Q ◦ P ∈ Lj+k
pc (E, G) and

σj+k (Q ◦ P ) (x, η) = σj (Q) (x, η) ◦ σk (P ) (x, η), x ∈ X, η ∈ Tx∗ X \ {0} .

(b) Let P ∈ Lkpc (E, F ), where the bundles E and F are equipped with Hermitian
metrics and the manifold X is Riemannian and oriented. Then, there is a unique
operator P ∗ ∈ Lkpc (F, E) with
Z Z
hP u, viF = hu, P ∗ viE , for all u ∈ C0∞ (E), v ∈ C0∞ (F ), and
X X
 ∗
σk (P ∗ ) (x, η) = σk (P ) (x, η) , for x ∈ X, η ∈ Tx∗ X \ {0} .

Proof. We begin with (b): The uniqueness of P ∗ is clear. For the proof of
existence, we need only to show that for each v ∈ C0∞ (F ) and each open coordinate
domain U ⊆ X (with E|U and F |U trivial) there is a PU∗ v ∈ C ∞ (E|U ) such that
Z Z
hP u, viF = hu, PU∗ viE for all u ∈ C0∞ (E|U ).
X X
Indeed, when such a PU∗ v exists, then it is uniquely determined, and so for a second
coordinate domain U 0
PU∗ v|U ∩U 0 = PU∗ 0 v|U ∩U 0 .
230 8. PSEUDO-DIFFERENTIAL OPERATORS

Then we have a global C ∞ section P ∗ v ∈ C ∞ (E) with


Z Z
hP u, viF = hu, P ∗ viE for all u ∈ C0∞ (E),
X X
which is constructed by covering X with  finitely many coordinate domains U and
writing u = j uj with uj ∈ C0∞ E|Uj by means of a C ∞ partition of unity.
P
Hence, let v and U be given. Without loss of generality, we assume that there
is a coordinate domain V which contains U as well as supp(v). (Otherwise one
covers supp(v) with finitely many coordinate domains and pieces v together from a
C ∞ partition of unity.) In local coordinates, relative to a C ∞ N -framing for E|V
and M -framing for F |V , one can write P in the form
Z
(P u)(x) = u(ξ) , v(x)i d̄ξ, u ∈ C0∞ (E|V ),
eihx,ξi hp(x, ξ)b
Rn
where p is an M × N matrix of amplitude functions with compact support. Hence,
Z Z
hP u, vi = eihx,ξi hp(x, ξ)bu(ξ) , v(x)iCM d̄ξ dx
Z Z Z
= eihx−y,ξi hp(x, ξ)u(y) , v(x)iCM d̄y d̄ξ dx
Z  Z Z 
ihx−y,ξi ∗ 0
= u(y) , e p (x, ξ)v(x) d x d̄ξ dy,
CN

where p (x, ξ) is the adjoint of p(x, ξ) (i.e., complex-conjugate transpose), an N ×M
matrix of complex numbers for each x ∈ V , ξ ∈ Rn , and the integrals here and
below are over Rn . Thus, we get hP u, vi = hu, P ∗ vi for
Z Z
(P ∗ v) (y) := eihx−y,ξi p∗ (x, ξ)v(x) d̄xd̄ξ = eiϕ(x,y,ξ) q(y, x, ξ)v(x) d0 xd̄ξ,

where q(y, x, ξ) := p∗ (x, ξ) and ϕ(x, y, ξ) = hx − y, ξi . By the Theorem of Kuran-


ishi, there now exists an amplitude pe so that
Z
(P ∗ v)(y) = eihy,ξi pe(y, ξ)b
v (ξ) d̄ξ.

Hence, we have found P ∗ v ∈ Lk (CN M


U , CU ) with the desired property. Also, we have

σ(P ∗ )(x, ξ) = lim λ−k q(x, x, λξ) = lim λ−k p∗ (x, λξ) = (σ(P )(x, ξ)) .
λ→∞ λ→∞

Furthermore, we remark that with this we obtain (for all u ∈ C∞ (CN V ) and v ∈
C∞ (CM
V ))
Z  Z 
ihy,ξi
hP u, vi = u(y) , e pe(y, ξ)b
v (ξ) d̄ξ dy
N
Z Z  C 
= e−ihy,ξi pe(y, ξ)∗ u(y) d̄y , vb(ξ) d̄ξ,
CM

where pe(y, ξ) is the adjoint matrix of the amplitude pe(y, ξ) of the operator P ∗ . By

Parseval’s Formula, Pcu(ξ) is exactly the expression in the curly braces. Hence,
we have the additional formula
Z

(8.14) P u(ξ) = e−ihy,ξi (e
c p(y, ξ)) u(y) d0 y.
8.5. NORMAL (GLOBAL) AMPLITUDES 231

For (a): This time we do not rely on the global representation of P and Q, but
rather on their definition by localizations. Without loss of generality, let X be an
open, relatively compact subset of Rn . Let p and q be the amplitudes of orders k
and j, belonging to P and Q, respectively. By (8.14), we obtain for
Z

(QP u)(x) = eihx−y,ξi q(x, ξ)(ep(y, ξ)) u(y) d̄y d̄ξ.

which is thus a pseudo-differential operator (of order k + j) by Theorem 8.14 (of


Kuranishi), p. 217, and

q(x, λη)(e p(y, λη))
σk+j (Q ◦ P )(x, η) = lim
λ→∞ λk+j

q(x, λη) (e
p(y, λη))
= lim lim
λ→∞ λj λ→∞ λk
= σj (Q)(x, η) ◦ σk (P )(x, η),
p) = σk (p)∗ by (b).
since σk (e 
Remark 8.28. The derivation of the formally adjoint operator seems trivial
here in comparison to the lengthy calculation for differential operators; see Exer-
cise 6.43, p.185. Actually, we have three entirely different problems: By definition,
it is trivial that every Fourier-integral operator P possesses a formal adjoint P ∗ . To
prove that P ∗ is a pseudo-differential operator if P is, we need more (namely, the
Theorem of Kuranishi or somewhat long-winded direct computations). It is possi-
ble, by the way, to prove the sharper result P ∈ Diff k ⇒ P ∗ ∈ Diff k in this fashion,
by analyzing carefully the various transformations in the Theorem of Kuranishi
with this goal in mind.

5. Normal (Global) Amplitudes


Motivation. Our aim is to provide a framework for a proof of the Index The-
orem, which, when compared to existing approaches, we believe is somewhat more
streamlined and globally expressed (i.e., free of local coordinates). This method
is based on defining pseudo-differential operators from sections of a vector bundle
E → X to sections of a bundle F → X in terms of a globally defined symbol which
is a section p ∈ C ∞ (Hom(π ∗ E, π ∗ F )) of the bundle of homomorphisms between
the lifts π ∗ E and π ∗ F to the cotangent bundle T ∗ X, where π : T ∗ X → X. The
definition of the operator, say Op(p) ∈ C ∞ (E, F ), associated with p entails the in-
troduction of a metric on X and connections on E and F . It could be argued that
(without a considerable background in modern differential geometry) this is not eas-
ier than using local coordinates and framings, but there are advantages. This global
approach to pseudo-differential operators is not new. It seems to have first appeared
in the paper of Juliane Bokobza-Haggiag [67]. Many subsequent developments
and applications have appeared steadily since then, in the work of Harold Widom
[441], Ezra Getzler [162], Stephen Fulling and G. Kennedy [153], Yuri
Safarov [359, 289], Markus Pflaum [338] and Theodore Voronov [426],
just to mention a few.
The application of the global approach to the index theorem has been mostly
in the context of the heat equation proof, rather than in the embedding proof.
Nevertheless, in preliminaries leading up to their treatment of the embedding proof
in their enlightening book, H. Blaine Lawson and Marie-Louise Michelsohn
232 8. PSEUDO-DIFFERENTIAL OPERATORS

[273, p. 188], point out the possible desirability of defining pseudo-differential op-
erators with a global symbol. In essence, here we are exploring this possibility.
We find that some of the difficulties are softened. In particular, the lifting of a
pseudo-differential operator to an invariant one in the proof of the twisted multi-
plication formula is made easier, since the global symbol can be lifted by means of
a connection. Moreover the thorny problem of forming suitable products of indi-
vidual pseudo-differential operators with identity operators over product manifolds
(see [273, p. 250f] or [44, p. 514f]) is alleviated by performing operations on the
globally-defined total product symbols. This will be made clearer below. One fun-
damental challenge inspired by this program is the task of constructing a global
symbol whose associated pseudo-differential operator is exactly the operator that
one may want, as opposed to an approximate operator with essentially the same
asymptotic principal symbol.
Remark 8.29. There are various exciting interrelations emerging between the
deformation-theoretical approach to Weyl quantization on Riemannian manifolds
and our global symbolic calculus. In particular, we refer to [337] and [426]. In-
terpreting a global symbolic calculus as deformation quantization then leads to
an additional proof of the Atiyah-Singer Index Theorem via the algebraic index
theorems in [317, 318].
Bokobza-Haggiag Fourier Transform. We work within the C ∞ category
unless stated otherwise. Let ρ be the injectivity radius of the compact manifold
X with Riemannian metric g and Levi-Civita connection ∇; i.e., for all x ∈ X,
the exponential map expx : Tx X → X relative to g is injective on the disk of
radius ρ about 0x ∈ Tx X. Let πE : E → X and πF : F → X be complex Her-
mitian vector bundles equipped with Hermitian ∇E : C ∞ (E) → C ∞ (T ∗ X ⊗ E)
and ∇F : C ∞ (F ) → C ∞ (T ∗ X ⊗ F ), where C ∞ (E) denotes the space of (smooth)
E
sections of πE : E → X. For x, y ∈ X, with d(x, y) < ρ, let τx,y : Ey → Ex denote
E
parallel translation relative to ∇ along the unique geodesic from y to x with mini-
mal length d(x, y). Here we emphasized in italics the terms which were introduced
more precisely in Sections 6.3 and 6.5.
Let ψ : [0, ∞) → [0, 1] be smooth, with ψ(r) = 1 for r ∈ [0, ρ/3] and ψ(r) = 0
for r ∈ [2ρ/3, ∞].
Definition 8.30. For π : T ∗ X → X and u ∈ C ∞ (E), we define the Bokobza-
Haggiag Fourier transform u∧ ∈ C ∞ (π ∗ E) (where x = π(ξ) and ξ ∈ T ∗ X)
by
Z
u∧ (ξ) := e−iξ(v) ψ(|v|)τx,exp
E
xv
[u(expx v)] d̄v ∈ Ex for ξ ∈ Tx∗ X,
Tx X
−n/2
where d̄v = (2π) dv and dv is the volume element on Tx X associated with gx .
For x, y ∈ X with d(x, y) < ρ, we have y = expx v for a unique v ∈ Tx X with
|v| = d(x, y), and we may define α ∈ C ∞ (X × X, [0, 1]) by

ψ(d(x, y)) = ψ(|v|), for d(x, y) < ρ ,
α(x, y) :=
0, for d(x, y) ≥ ρ.
Note that we can think of the function (in C ∞ (Tx X, Ex ))
E E
v 7→ ψ(|v|)τx,expxv
[u(expx v)] = τx,expxv
[α(x, expx v)u(expx v)] (v ∈ Tx X)
8.5. NORMAL (GLOBAL) AMPLITUDES 233

as a pull-back (of sorts), using τ E and expx : Tx X → X, of the bump function α(x, ·)
times u(·) in a neighborhood of x, and u∧ |Tx∗ X is the Fourier transform of this
“pull-back” of α(x, ·)u(·). The Bokobza-Haggiag inverse Fourier transform

(u∧ ) : T X → E of u∧ is given by
Z
∧ ∨
(u ) (v) := eiξ(v) u∧ (ξ) d̄ξ = ψ(|v|)τx,expx v [u(expx v)],
Tx∗ X
∨ ∨
where d̄ξ = (2π)−n/2 dξ. Since (u∧ ) (v) ∈ Ex , (u∧ ) is a section of the pull-back
of E to T X via π : T ∗ X → X. Moreover, we can recover u locally about x from

u∧ |Tx∗ X . In particular, for v = 0x ∈ Tx X, we have (u∧ ) (0x ) = u(x).
Pseudo-Differential Operators and Normal Amplitudes. For π : T ∗ X →
X and a section p ∈ C ∞ (Hom(π ∗ E, π ∗ F )) , (of Hom(π ∗ E, π ∗ F ) → T ∗ X) we define
an operator Op(p) : C ∞ (E) → C ∞ (F ) via
Z Z

Op(p)(u)x := e iξ(v)
p(ξ)(u (ξ)) d̄ξ = p(ξ)(u∧ (ξ)) d̄ξ
Tx∗ X Tx∗ X
v=0
Z Z 
= p(ξ) e−iξ(v) ψ(|v|)τx,exp
E
xv
[u(expx v)] d̄v d̄ξ
Tx∗ X Tx X
Z  
= p(ξ) e−iξ(v) ψ(|v|)τx,exp
E
x v [u(expx v)] d̄vd̄ξ
Tx X×Tx∗ X
Z  
(8.15) = e−iξ(v) p(ξ) τx,exp
E
x v [α(x, exp x v)u(expx v)] d̄vd̄ξ.
Tx X×Tx∗ X

Remark 8.31. a) In the spirit of Remark 8.4, p.211, one may think of Op(p)
as a quantization of p, but Op(p) continues to depend on many choices also in our
global setting (e.g., the choice of metric, connections, and α : X × X → [0, 1]).
b) Apart from these choices, there are other choices one can make, as discussed in
[426]. For example, if s ∈ [0, 1], let

Tx,expx sv : Texpx sv
X → Tx∗ X
denote parallel translation (with respect to the Levi-Civita connection) for T ∗ X
along the geodesic t 7→ expx tv in the reverse direction from expx sv to x. In [426]
(but with notation that differs from ours), an operator Op(p; s) (depending on s)
is associated to p via
Z
Op(p; s)(u)x = e−iξ(v) α(x, expx v)
Tx X×Tx∗ X
F E
τx,expx sv
p(Tx,expx sv (ξ))τexp x sv, expx v
u(expx v) d̄vd̄ξ.
When s = 0, we get
Z
Op(p; 0)(u)x = e−iξ(v) α(x, expx v)p(ξ)τx,
E
expx v (u(expx v)) d̄vd̄ξ
Tx X×Tx∗ X

which is precisely our Op(p). In interesting cases, the operators Op(p; s) for different
s differ by “lower order” operators which do not affect the index (if defined). Hence,
for simplicity, we only use s = 0. As stated in [426] the choice of s is related to
the choice of operator ordering of monomials in position and momentum variables
under quantization.
234 8. PSEUDO-DIFFERENTIAL OPERATORS

The connections ∇E and ∇F pull back via π : T ∗ X → X to connections on


the bundles π ∗ E → T ∗ X and π ∗ F → T ∗ X, which we continue to denote by
∇E and ∇F . The Levi-Civita connection for (X, g) determines a subbundle H
of T (T ∗ X) consisting of horizontal subspaces of T (T ∗ X), which is complementary
to the subbundle V of T (T ∗ X) consisting of vectors which are tangent to the fibers
of T ∗ X → X. There is a natural Riemannian metric, say g ∗ , on T ∗ X such that
V and H are orthogonal and g ∗ equals g on V and π ∗ g on H. Using ∇E and ∇F ,
along with the Levi-Civita connection for g ∗ , say ∇∗ , we may construct a covariant
derivative
e : C ∞ (Hom(π ∗ E, π ∗ F )) → C ∞ (T ∗ (T ∗ X) ⊗ Hom(π ∗ E, π ∗ F )) .

Since ∇∗ extends to ⊗k T ∗ (T ∗ X), we may “iterate” ∇
e to obtain
e k : C ∞ (Hom(π ∗ E, π ∗ F )) → C ∞ ⊗k T ∗ (T ∗ X) ⊗ Hom(π ∗ E, π ∗ F ) .


Now we are ready to give a geometric and truly global definition of a variant of
our spaces of principally classical pseudo-differential operators, introduced locally
in Section 8.3. Recall from Remark 6.28 (p.179) that a connection of a vector
bundle E splits the tangent space T E into two bundles, the bundle V of vertical
spaces and the bundle H of complementary horizontal spaces.
Definition 8.32. a) We say that p ∈ C ∞ (Hom(π ∗ E, π ∗ F )) is a Bokobza-
Haggiag amplitude (or total symbol) of order k ∈ R if for any H1 , . . . , HI ∈
C ∞ (H) with |H1 |, . . . , |HI | ≤ 1 and V1 , . . . , VJ ∈ C ∞ (V ), there are constants CIJ
(depending only on I, J and p), such that
   XJ k−J
I+J
(8.16) ∇
e p (H1 , . . . , HI , V1 , . . . , VJ ) ≤ CIJ 1 + |Vj | .
j=1

Moreover, we require that the k-th order asymptotic symbol (or principal sym-
bol) of p, namely
p(tξ)
(8.17) σk (p)(ξ) := lim (for ξ 6= 0)
tk
t→∞

exists, where the convergence is uniform on S(T ∗ X).


b) Then we call Op(p) a Bokobza-Haggiag2 pseudo-differential operator of or-
der k and write Op(p) ∈ LkBokobza (E, F ). We set L•Bokobza (E, F ) := LkBokobza (E, F ).
S
c) We denote the set of amplitudes of order k by Amplk (E, F ).
Clearly, for k 0 > k,
\−∞
Amplk0 (E, F ) ⊃ Amplk (E, F ) ⊃ Ampl−∞ (E, F ) := Amplk (E, F ).
k=0

For p ∈ Amplk (E, F ), we then have the operator, say Op(p) : C ∞ (E) → C ∞ (F ),
given by (8.15), which extends to a bounded operator Ops (p) : W s (E) → W s−k (F ),
where for any s ∈ R, W s (E) denotes the s-th Sobolev space of sections of E,
namely the completion of C ∞ (E) with respect to the norm k·ks defined by
Z  s
2 2 2
kuks := 1 + |ξ| |u∧ (ξ)| dξ.
T ∗X

2also called normal, see [338].


8.5. NORMAL (GLOBAL) AMPLITUDES 235

Recall that for k ∈ Z+ , and s > n/2 + k, there is a compact inclusion W s (E) ⊂
C k (E). For each s, the linear map
(8.18) Ops : Amplk (E, F ) → B(W s (E), W s−k (F ))
into the Banach space B(W s (E), W s−k (F )) of bounded linear transformations is
continuous (see the corresponding Exercise 9.6, p.239 or [273, p. 177f]). More-
over, for ϕ ∈ Ampl−∞ (E, F ), Ops (ϕ) is a compact operator for any s ∈ R, and
Ops (ϕ)(W s (E)) ⊂ C ∞ (F ); i.e., Ops (ϕ) is a smoothing operator.
Definition 8.33. We say that p ∈ Amplk (E, F ), and the corresponding oper-
ator Op(p), are elliptic if for some constant c > 0, p(ξ)−1 exists for |ξ| > c, and
for some constant K > 0
−k
p(ξ)−1 ≤ K(1 + |ξ|) for all ξ ∈ T ∗ X with |ξ| > c.
We set EllkBokobza (E, F ) := {p ∈ Amplk (E, F ) : p is elliptic} .
For p ∈ EllkBokobza (E, F ), there are q ∈ Ampl−k (E, F ), ϕE ∈ Ampl−∞ (E, E)
and ϕF ∈ Ampl−∞ (F, F ), such that
Ops−k (q) ◦ Ops (p) = IdW s (E) + Ops (ϕE ) and
Ops (p) ◦ Ops−k (q) = IdW s−k (F ) + Ops−k (ϕF ).
Since Ops (ϕE ) and Ops−k (ϕF ) are compact operators, it follows that Ops (p) is
Fredholm, and hence we may define
index(Ops (p)) := dim Ker(Ops (p)) − dim Coker(Ops (p)).
Note also that if Ops (p)u ∈ C ∞ (F ), then
u = Ops−k (q)(Ops (p)u) − Ops (ϕE )u ∈ C ∞ (E).
Thus, dim Ker(Ops (p)) < ∞, Ker(Ops (p)) ⊂ C ∞ (E), and Ker(Ops (p)) is indepen-
dent of s. As a consequence,
index(Ops (p)) = dim Ker(Ops (p)) − dim Coker(Ops (p))
is independent of s.
Approximation of Differential Operators. The reader should be aware of
a minor technical problem when dealing with our Bokobza-Haggiag amplitudes and
Bokobza-Haggiag pseudo-differential operators: There can be differential operators
which can not be generated by a Bokobza-Haggiag amplitude.
Let us have a closer look at the familiar case of an elliptic, linear differential
operator D : C ∞ (E) → C ∞ (F ) of given order k. As we have seen in the preceding
chapter, associated with D is its principal symbol σk (D) ∈ C ∞ (Hom(π ∗ E, π ∗ F )).
We have checked that σk (D) is independent of the choice of local coordinates and
observed that this would not be the case if lower-order terms were included. If σk (D)
is invertible outside of the zero section of T ∗ X, then D is said to be elliptic, which
we assume. If lower order terms were included and if we denoted this coordinate-
dependent, locally-defined “full symbol” by ploc (D)(ξ), then σk (D)(ξ) at ξ ∈ Tx∗ X
would be given by
ploc (D)(tξ)
lim
t→∞ tk
in comparison with (8.17). However, it is not clear that D is precisely Op(p) for
some globally defined p ∈ Amplk (E, F ). In the language of physicists, it is not clear
236 8. PSEUDO-DIFFERENTIAL OPERATORS

that D can be precisely dequantized. If such p exists, it would clearly depend on


choices of a Riemannian metric on X, connections for E and F and on the function
α : X × X → [0, 1] supported near the diagonal. However, in [67] and [441], it is
shown that given such choices, p can be found so that Op(p) and D differ by an
operator which is infinitely smoothing (and hence compact); i.e.,
D − Op(p) = Op(a) for a ∈ Ampl−∞ (E, F ).
By methods that are standard by now and to be summarized in the following chap-
ter, it follows that D has Fredholm Sobolev extensions Ds : W s (E) → W s−m (F )
for all s, with a common index, which is sometimes called the analytic index of D;
it is just the usual operator-theoretic index. It is simply denoted by index(D) and
if D∗ : C ∞ (F ) → C ∞ (E) denotes the formal L2 -adjoint of D, then
dim Ker(D) − dim Ker(D∗ ) = index(D)
= index(Op(p) + Op(a)) = index(Op(p)).
Thus, readers (including the authors) who are bothered by the fact that differential
operators may not be precisely dequantized, may take some solace in the fact that
elliptic differential operators may be approximated by a pseudo-differential operator
of the form Op(p), modulo smoothing operators which preserve the index.
Note. If D is a scalar differential operator, then there exists one and only one
amplitude p which is polynomial in each fiber such that D = Op(p).
We close this section with a few exercises.
Exercise 8.34. a) Show that L•Bokobza (E, F ) of Definition 8.32b is a graded

-algebra.
b) Show that the quantization Op : Amplk (E, F ) → LkBokobza (E, F ) is a bijection
for all k ∈ R.
c) Show that the spaces Amplk (E, F ) and LkBokobza (E, F ) are independent of the
choice of metrics and connections (contrary to the definition of Op).
d) Prove LkBokobza (E, F ) ⊂ Lkpc (E, F ).
e) Find a closed Riemannian manifold X, Hermitian bundles E, F → X with
fixed connections and k ∈ N such that Diff k (E, F ) 6⊂ LkBokobza (E, F ). Conclude
LkBokobza (E, F ) 6= Lkpc (E, F ). For that inequality, can E, F be trivial bundles? Can
you choose X = S 1 or, more generally, X = S n ?
[Hint: To a: Check composition and taking formal adjoints like in Section 8.4.
To b: By definition of LkBokobza (E, F ) and the linearity of Op it suffices to prove
the injectivity of Ops of (8.18) for s = 0. How can you exclude the existence of a
not identically vanishing p ∈ Amplk (E, F ) with Op(p) = 0 in spite of the Example
8.16, p.218?
To d: Fix the metric structures and connections. Then check the claim in coordi-
nates.
To e: See the references given at the beginning of this section, in particular [67],
[441] and [338].]
CHAPTER 9

Elliptic Operators over Closed Manifolds

Synopsis. Continuity of Pseudo-Differential Operators between Sobolev Spaces. Para-


metrices for Elliptic Operators: Regularity and Fredholm Property. Topological Closures.
Outer Tensor Product on Product Manifolds. The Topological Meaning of the Principal
Symbol (Simple Case Involving Local Boundary Conditions).
I In this chapter, we derive the basic mapping properties of pseudo-differential oper-
ators and show that out of formal properties (e.g., invertibility) of elliptic symbols, a series
of existence, regularity, and finiteness results for the associated differential and pseudo-
differential operators (as Fredholm operators) can be obtained. Moreover, we begin to
explore the outer tensor product that plays a decisive role in the K-theoretic proof of the
Index Theorem.J

1. Mapping Properties of Pseudo-Differential Operators


Convention 9.1. We do not strive for greatest generality. In what follows,
the manifold X is closed, i.e. compact, without boundary. We make this conven-
tion, in part for convenience (in order to make some proofs go easier), but also
because otherwise some of the following theorems would be meaningless or false;
see Exercise 9.23 below. Moreover, X continues to be oriented. For convenience
and to have an L2 -structure at our disposal, we assume that X is furnished with a
fixed Riemannian metric and that E and F are Hermitian vector bundles. Without
loss of generality, we will occasionally assume that the Hilbertable Sobolev spaces
are already furnished with a fixed norm or scalar product. The pseudo-differential
operators considered are principally classical. In particular, the principal symbol
is well defined by Assumption 8.10(ii). In a later Section 12.2, p.287ff, we shall
come back to the global, approach via Bokobza-Haggiag amplitudes and Bokobza-
Haggiag pseudo-differential operators.
Definition 9.2. A linear operator P : C0∞ (E) → C ∞ (F ) is called an operator
of order k ∈ Z if it extends to a continuous map Ps : W s (E) → W s−k (F ) for all
s ∈ R with s, s − k > 0. We denote by OPk (E, F ) the set of operators of order k.
Fixing norms on W s (E) and W s−k (F ), we set
(9.1) kP ks,s−k := sup{kP uks−k : u ∈ C ∞ (X; E) and kuks = 1}.
Formally, the following theorem says that the analytical order of a pseudo-
differential operator (determined by the asymptotic behavior its amplitude) coin-
cides with its order in the context of functional analysis (which is expressed by its
continuity relative to the norms of the Sobolev spaces). Actually, the claim is valid
for all pseudo-differential operators. For the proof no existence or homogeneity of
the principal symbol is required.
Theorem 9.3. For k ∈ Z, we have Lkpc (E, F ) ⊂ OPk (E, F ).
237
238 9. ELLIPTIC OPERATORS OVER CLOSED MANIFOLDS

Proof. We need to prove the estimate kP uks−k ≤ C kuks , u ∈ W s (E), where


C only depends on P ∈ Lkpc (E, F ) and s, and not on u. Since C0∞ (E) is dense in
W s (E), it suffices to prove the preceding estimate for u ∈ C0∞ (E).
Moreover, since the norms in W s (E) and W s−k (F ) are locally defined by Ex-
ercise 7.7 (p. 197), it suffices to show the inequality for u ∈ C0∞ (Rn ). By Exercise
7.5 (p. 197) we can further assume without loss of generality that k = 0 and s = 0.
Let Z
(P u)(x) = eihx,ξi p(x, ξ)b
u(ξ) d̄ξ, u ∈ C0∞ (Rn ),

be a pseudo-differential operator, whose amplitude p(x, ξ) vanishes for sufficiently


large x and satisfies the estimate (see 8.2, p. 210)
k−|α|
Dxβ Dξα p(x, ξ) ≤ C(1 + |ξ|) , (x, ξ) ∈ Rn × Rn ,
for all multi-indices α and β. Then the Fourier transform pb(·, ξ) of the function
x 7→ p(x, ξ) can be estimated by
(9.2) p(z, ξ)| ≤ CN (1 + |z|)−N for all N ∈ N,
|b
as was done for u
b in the proof of Theorem 8.3, p. 210. We have
Z Z Z
−ihη,xi ihx,ξi
P u(η) =
c e e p(x, ξ) d̄x u
b(ξ) d̄ξ = pb(η − ξ, ξ) u
b(ξ) d̄ξ.

By (9.2) and the Schwarz Inequality,


Z 2
|Pcu(η)|2 ≤ CN2
(1 + |η − ξ|)−N u
b(ξ) d̄ξ
Z 
2
≤ (2π)−n CN
2
(1 + |η − ξ|)−2N dξ kb
uk0 .

For N sufficiently large, integrating this with respect to η and using Parseval’s
Formula (Appendix A, Exercise A.5d), we obtain the desired result
2 2 2
kP uk0 = ||Pcu||20 ≤ C 0 kb
uk0 = C 0 kuk0 . 
Exercise 9.4. Interpret and prove:
P ∈ Lkpc (E, F ) =⇒ (P0 )∗ = (P ∗ )0 .
[Hint: Recall Theorem 8.27b, p.229.]
General functional analysis does not suffice to obtain the Index Theorem for
elliptic operators. The more refined structure of differential and pseudo-differential
operators is required. Apparently, the terms of highest order must be singled out.
Fortunately, they have a global meaning in spite of being defined locally. On the
basis of that double character the Index Theorem will be erected.
As a matter of fact, by introducing the spaces Lkpc (E, F ) of principally classical
pseudo-differential operators, we singled out the terms of highest order, the princi-
pal symbol. However, there is a slight ambiguity in defining the space Smblk (E, F )
of the induced principal symbols. Further above, in Equation (6.29), p.184 and
correspondingly in Exercise 8.20e, 224, we considered the space Smblk (E, F ) as a
subspace of smooth bundle homomorphisms from π ∗ E to π ∗ F over the base space
T̊ ∗ X, where π : T̊ ∗ X → X denotes the projection of the dotted cotangent bundle
T̊ ∗ X := T ∗ X \ X (the symplectic cone) onto X. Then the bundle homomorphisms
9.2. ELLIPTIC OPERATORS — REGULARITY AND FREDHOLM PROPERTY 239

belonging to Smblk (E, F ) are characterized by k-homogeneity in the cotangent


variable.
Equally well and topologically more manageable, we may identify the symbol
space Smblk (E, F ) with the whole space HomSX (E, F ) = C ∞ SX, Hom(π0∗ E, π0∗ F )
of all smooth bundle homomorphisms from π0∗ E to π0∗ F over the base space
SX := {ξ ∈ Tx∗ X : x ∈ X and |ξ|x = 1)} (the co-sphere bundle),
where π0 : SX → X denotes the base point map.
Exercise 9.5. Let σk : Lkpc (E, F ) → Smblk (E, F ) denote the well-defined (Ex-
ercise 8.20, p. 224), surjective (Theorem 8.25, p. 228) symbol map. Show Ker σk ⊆
OPk−1 (E, F ). [Hint: This is trivial by axiom (iii) (see Section 8.3, p. 214) and the
preceding Theorem 9.3.]
Exercise 9.6. Show that the short exact sequence
σ
Lkpc (E, F ) −→
k
Smblk (E, F ) −→ 0
splits; i.e., σk has a linear right inverse Opk : Smblk (E, F ) → Lkpc (E, F ) which
satisfies the continuity condition
kOp(ρ)ks,s−k ≤ C sup {|ρ(x, ξ)| : x ∈ X, |ξ| = 1} (s, s − k > 0),
where C does not depend on ρ ∈ Smblk (E, F ), and |ρ(x, ξ)| denotes the usual
matrix norm which may be defined via the Hermitian inner products on Ex and
Fx . [Hint: Go through the proof of Theorem 8.25 (p. 228) again.]
Exercise 9.7. Show conversely, that for all s, s − k > 0 and all P ∈ Lkpc (E, F ),
we have
sup {|σk (P )(x, ξ)| : x ∈ X, |ξ| = 1} ≤ kP ks,s−k .
[Hint: By a theorem of Israil Gohberg (see also [378, p.171]), one can find a
sequence {ϕν } of functions in C0∞ (Rn ) such that
(1) ϕν (x) = 0 for |x − x0 | > 1/ν ;
(2) kϕν k0 = 1 for all ν, and
(3) kP ϕν − σk (P )(x0 , ξ0 ) ϕν k0 −→ 0 as ν −→ ∞, where (x0 , ξ0 ) ∈ Rn ×
(Rn \ {0}) is any given point.
Details are found in [378, p.179] or [109, 22-05]. Caution: On Lkpc (E, F ) itself
there is a topology which is defined in a natural way by the condition that the map
P 7→ Ps be continuous for all s. However, then σk is not continuous; see [328,
p.175]. For other topologies on Lkpc (E, F ) we refer to [73, Section 2.3].]

2. Elliptic Operators — Regularity and Fredholm Property


As a generalization of our earlier definition for differential operators (see Sec-
tion 6.6 above), we call P ∈ Lkpc (E, F ) elliptic, if σk (P )(x, ξ) is an isomorphism
from Ex to Fx for all x ∈ X and ξ ∈ Tx∗ X, ξ 6= 0. We write P ∈ Ellk (E, F ).
Theorem 9.8 (Main result). For any P ∈ Ellk (E, F ), there is Q ∈ Ell−k (E, F ),
such that P Q − IdF ∈ OP−1 (F, F ) and QP − IdE ∈ OP−1 (E, E).
Remark 9.9. This existence theorem forms the foundation of our theory of
elliptic operators. Using the terminology introduced by David Hilbert, one calls
Q a parametrix for P , although a crude one: The classical parametrix (Green’s
function) inverts P , not only modulo operators of order −1, but also of order −∞
240 9. ELLIPTIC OPERATORS OVER CLOSED MANIFOLDS

(i.e., modulo so-called smoothing operators; see [219]). This can be obtained locally
by inverting the amplitude p of P = Op(p) and then setting Q := Op(p−1 ) and
globally by the corresponding operations in the Bokobza-Haggiag calculus explained
above in Section 8.5, pp.231ff.
Proof. Theorem 8.25 (p. 228) guarantees the existence of a Q ∈ L−k pc (F, E)
 −1
0
with σ−k (Q)(x, ξ) := σk (P )(x, ξ) , whence P Q ∈ Lpc (F, F ) by Theorem 8.27a
(p. 229) and σ0 (P Q − IdF ) = 0, and σ0 (P Q − IdF ) ∈ OP−1 (F, F ) by Exercise 9.5.

Theorem 9.10. Let P ∈ Ellk (E, F ) and s, s − k ≥ 0. Then we have:
a) Finiteness: The extension Ps : W s (E) → W s−k (F ) is a Fredholm oper-
ator with index independent of s.
b) Existence: P ∗ is elliptic and Coker Ps ∼ = Ker(P ∗ )s−k .
c) Regularity: Ker Ps = Ker P .
d) Homotopy-invariance: index P = index Ps depends only on the ho-
motopy class of σ(P ) := σk (P )|SX in IsoSX (E, F ). Here IsoSX (E, F ) de-
notes the space of C ∞ bundle isomorphisms π0∗ E ∼= π0∗ F , where π0 : SX →
X is the base-point map and SX denotes the co-sphere bundle as before;
IsoSX (E, F ) is equipped with a supremum norm as in Exercise 9.6.
Remark 9.11. In conjunction with c), existence says that the inhomogeneous
equation P u = f has a solution exactly when f ⊥ Ker P ∗ , and the solution is unique
if constrained to be orthogonal to Ker P in W 0 (E). By regularity, all classical
solutions (i.e., u ∈ C k (E)) of homogeneous elliptic differential equations P u = 0
(with C ∞ coefficients) lie in C ∞ (E). In the context of distribution theory (see
[223]), one obtains the sharper result that every weak solution (in the distribution
sense) is a strong solution (in the function sense); i.e., from the assumption u ∈
W 0 (E) and hu, P ∗ f i0 = 0 for all f ∈ C ∞ (F ), the conclusions u ∈ C ∞ (E) and
P u = 0 follow. Such regularity results, which were first proved in 1940 by Hermann
Weyl in the case of the Laplace operator P := ∆, are of special importance when
one is solving partial differential equations by variational methods (i.e., solving
through extremal conditions); see [217, p.96] and [280, p.214f].
Proof. For a: If Q ∈ L−k pc (F, E) is a parametrix (see Theorem 9.8) for P ,
then it follows from the Theorem of Franz Rellich (Theorem 7.15, p. 201) that
the composition
Qs−k Ps −Id
W s (E) −→ W s+1 (E) ,→ W s (E)
is a compact operator on W s (E), and correspondingly, Ps Qs−k − Id is a compact
operator on W s−k (F ). Thus, Ps : W s (E) → W s−k (F ) is a Fredholm operator by
Theorem 3.2, p. 64. (There, actually the proof was explicitly given only for endo-
morphisms, but this is no restriction for separable Hilbert spaces, since they are all
isomorphic.) By continuity considerations (Theorem 3.11, p. 68 and the preceding
Exercises 9.6 and 9.7, or easy norm comparison for Ps and Pt by means of Λs−r of
Exercise 7.5, p. 197) it follows that index Ps = index Pt .
For b: Without loss of generality (the Λ argument of Exercise 7.5, p. 197), let
k = s = 0. Then b) follows directly from Exercise 9.4 and Theorem 2.7.
For c: By definition, we have Ker Ps+1 ⊆ Ker Ps , since W s+1 (E) ⊆ W s (E). Con-
versely, by Theorem 9.8 there is a bounded operator K : W s (E) → W s+1 (E), such
9.2. ELLIPTIC OPERATORS — REGULARITY AND FREDHOLM PROPERTY 241

that Qs−k Ps u − Id u = Ku for all u ∈ W s (E), where Q is a parametrix for P .


Thus u ∈ W s+1 (E), if Ps u = 0. Hence, Ker Ps = Ker Ps+1 = · · · = Ker P , since
C ∞ = ∩W s ; see Exercise 7.3c (p. 195) and Theorem 7.13, p. 200).
For d: For Q ∈ Ellk (E, F ) and σk (Q) = σk (P ), we obtain index Q = index P
from Exercise 9.5 and the invariance of the index under perturbation by com-
pact operators (Exercise 3.10, p. 65). In general, by Exercise 9.6, each continuous
curve ρ : I → Smblk (E, F ) lifts to a corresponding continuous (in the operator
norm) curve r : I → Lkpc (E, F ) with σk ◦ r = ρ. Thus, if one can connect σk (Q)
with σk (P ) by a continuous curve in Iso∞ SX (E, F ), then P and Q can be con-
nected in Ellk (E, F ). Hence for all s, the Fredholm operators Qs and Ps lie in
the same component and have the same index by Theorem 3.11, p. 68. Finally, if
Q ∈ Ellj (E, F ) is an operator whose symbol σj (Q) coincides with σk (P ) on SX,
then σj (Q)−1 ◦ σk (P ) is the symbol of a self-adjoint operator R ∈ Ellk−j (E, E).
From σk (P ) = σj (Q) ◦ σk−j (R) = σk (QR), it follows by preceding arguments that
index P = index QR = index Q + index R = index Q,
since index R = 0 by (c) (see also the composition rule in Exercise 1.10, p. 8, and
the following Exercise 9.16a). 
The preceding results can be reformulated in the following more common esti-
mates.
Exercise 9.12 (Fundamental Elliptic Estimates). Let P ∈ Ellk (E, F ).
a) Regularity — Lifting Jack. If u ∈ L2 (E) and P u = v with v ∈ W s (F ) (e.g.,
v = 0, so u is a weak solution of P u = 0, i.e., in the distributional sense), then
u ∈ W s+k . Choosing norms in the Sobolev spaces you get kP uks < ∞ for all real
positive s, if P u ∈ C ∞ (F ). Conclude that u ∈ C ∞ (E), i.e., kuks+k < ∞ for all
positive s.
b) Variant of Gårding’s Inequality. For all real s ≥ 0 there exists a positive
constant Cs such that
 
(9.3) kuks+k ≤ Cs kP uks + kuks .

Definition 9.13. Let U ⊂ X be open and P ∈ Lkpc (CN N


U , CU ) with N ∈ N and
k ∈ R and P not necessarily elliptic. We call the operator P hypoelliptic, if for all
s ∈ R and u ∈ W s (U, CN
U ), P u smooth implies that u is smooth.

Remark 9.14. a) For other purely operator-theoretical definitions of hypoel-


lipticity see [391, Section 5] and [224, Chapter 22].
b) Elliptic regularity says that each elliptic operator is hypoelliptic. From an alge-
braic point of view, that property is deeply counter-intuitive for differential opera-
tors: It means that all derivatives of a section u will behave fine, if only a selection
of derivatives (those entering into the definition of P ) behave fine and the selection
is balanced, namely with elliptic principal symbol.
Convention 9.15. In the following, we write for short σ(P ) for the restriction
of σk (P ) to SX.
Exercise 9.16. Let E, F, G, H be Hermitian vector bundles over the closed,
oriented Riemannian manifold X; P ∈ Ellk (E, F ), Q ∈ Ellj (F, G), R ∈ Ellk0 (G, H).
Show that the following expressions are defined, and prove the formulas:
a) index P ∗ = − index P
242 9. ELLIPTIC OPERATORS OVER CLOSED MANIFOLDS

b) index QP = index P + index Q


c) index P ⊕ R = index P + index R
d) index P = 0, if σ(P )(x, ξ) depends only on x and not on ξ ∈ (SX)x
[Hint for d: A bundle isomorphism E → F and a multiplication operator Mψ ∈
Ell0 (E, F ) are defined via ψ(x) := σ(P )(x, ξ) for ξ ∈ (SX)x . Apply Theorem
9.10d.]
For later use we summarize our results in a slightly different form.
Theorem 9.17 (Elliptic Decomposition). Let X be a closed smooth oriented
Riemannian manifold and E, F Hermitian vector bundles over X. Let k ∈ Z and
P ∈ Lkpc (E, F ). We denote the formally adjoint operator by P ∗ and the Sobolev
extensions by
Ps+k : W s+k (E) −→ W s (F ) and (P ∗ )s+k : W s+k (F ) −→ W s (E).
We distinguish three cases.
a) If the principal symbol σk (P ) is either injective or surjective, we have a direct
sum decomposition in closed subspaces
W s+k (E) = Ker(Ps+k )⊕Im (P ∗ )s+2k and W s+k (F ) = Ker (P ∗ )s+k ⊕Im(Ps+2k ).
 

b) If the principal symbol σk (P ) is injective, then the operator P ∗ P is elliptic and


we have
Ker(Ps+k ) = Ker P = Ker(P ∗ ◦ P ) ⊂ C ∞ (E)
with only finitely many linearly independent sections in Ker P .
c) If the principal symbol σk (P ) is surjective, then the operator P P ∗ is elliptic and
we have
Ker (P ∗ )s+k = Ker P ∗ = Ker(P ◦ P ∗ ) ⊂ C ∞ (F )


with only finitely many linearly independent sections in Ker P ∗ .


Corollary 9.18. If P is elliptic, then both the spaces Ker P and Ker P ∗ are
finite-dimensional and consist of smooth sections. Moreover, the Sobolev extension
Ps+k : W s+k (E) → W s (F ) is a Fredholm operator with
index(Ps+k ) = dim Ker P − dim Ker P ∗ =: index P.

3. Topological Closure and Product Manifolds


Exercise 9.19. a) Form the closure Smblk (E, F ) of Smblk (E, F ) in the supre-
mum norm, and show that one then obtains all the continuous symbols. In particu-
lar, the space IsoSX (E, F ) of all continuous isomorphisms from π ∗ E to π ∗ F , where
π : SX → X is the projection, consists of the restrictions p|SX for p ∈ Smblk (E, F )
b) For each s, s − k ≥ 0, form the closure (in the operator norm) of the set of all
operators Ps with P ∈ Lkpc (E, F ), and show that σk can be continuously extended
to a surjective map on this space if the target of σk is enlarged to Smblk (E, F ).

We now write P ∈ Lkpc (E, F ), if P lies in the closure formed in Exercise 9.19, for
all s ≥ 0 (and s−k ≥ 0). One easily sees that our results up to now (in particular on
elliptic operators) remain valid in this larger class. The most important reason for
passing to the closure arises from the multiplicative behavior of pseudo-differential
operators on product manifolds. Here we give a first taste. We shall elaborate the
following concepts and results much further when
9.3. TOPOLOGICAL CLOSURE AND PRODUCT MANIFOLDS 243

• introducing the basic idea of Bott Periodicity in Equation (10.1), p.261;


• deriving the outer product of K-theory in Theorem 10.16, p.267; and
• proving the multiplicative property of the analytical index under embed-
ding in Section 12.2, p.287ff.
Exercise 9.20. Consider two closed, oriented, Riemannian manifolds X and
Y ; and Hermitian vector bundles E and F over X, and G and H over Y . Moreover,
let P ∈ Lkpc (E, F ) and Q ∈ Lkpc (G, H), k ∈ N.
a)Define a vector bundle E  G (and correspondingly F  H) on X × Y with fiber
Ex ⊗ Gy over (x, y).
b) Show that an operator P ⊗ IdG ∈ Lkpc (E  G, F  G) is defined by
(P ⊗ IdG ) (u ⊗ v) := P u ⊗ v, u ∈ C ∞ (E), v ∈ C ∞ (G),
which does not always lie in Lkpc (E  G, F  G).
c) Over the manifold X × Y , define the outer tensor product
P #Q : C ∞ (E  G) ⊕ C ∞ (F  H) → C ∞ (F  G) ⊕ C ∞ (E  H)
by the matrix
P ⊗ IdG − IdF ⊗Q∗
 
P #Q := .
IdE ⊗Q P ∗ ⊗ IdH
Prove the Multiplication Theorem: If P and Q are elliptic, then P #Q is elliptic,
and index(P #Q) = (index P )(index Q).
[Hint for b: For the sake of simplicity, assume that all bundles are trivial line
bundles (i.e., the case of functions). Then P ⊗ IdCY : C ∞ (X × Y ) → C ∞ (X × Y )
is the operator obtained when P acts on the first variable while the second is held
fixed. In local coordinates, for (x, y) ∈ X × Y and u ∈ C ∞ (X × Y ), we have:
Z Z
(P ⊗ IdCY ) u(x, y) = eihx−x̄,ξi p(x, ξ)u(x̄, y) d0 x̄ d̄ξ.

The amplitude pe of P ⊗IdCY is then given by pe(x, y, ξ, η) = p(x, ξ) (up to a constant


of integration). Show that the amplitude estimate
k−|α|
Dxβ Dξα p(x, ξ) ≤ C(1 + |ξ| + |η|)
can only hold for large |α| when Dxβ Dξα p(x, ξ) is identically zero (i.e., when p is a
polynomial in ξ). For the proof that
P ⊗ IdCY ∈ Lkpc (X × Y ) := Lkpc (CX×Y , CX×Y ),
explicitly construct a family {Rt : t ∈ (0, 1]} with
Rt ∈ L0pc (X × Y ) and (P ⊗ IdCY ) Rt ∈ Lkpc (X × Y ),
such that (P ⊗ IdCY ) Rt converges in the operator norm to P ⊗ IdCY as t ↓ 0, as
in [44, p.513-516]. Another proof can be found in [220, p.96f] and [224, Lemma
19.2.6 and Theorem 19.2.7], where the consideration of the difference variable z
in x and y (and ζ in ξ and η) is carried out in the framework of the theory of
Fourier integral operators, with its more flexible methods; see also the theorem of
Kuranishi (Theorem 8.14, p. 217).
For c: For the origin of the peculiar form of P #Q compare with golden rule of
tensoring chain complexes; see also Chapters 10 and 11. Details may be found
in [35, Section 10] , [378, p.190-193], [109, exp. 22], [44, p.526-529]. In Section
244 9. ELLIPTIC OPERATORS OVER CLOSED MANIFOLDS

12.2, pp.287ff, we shall prove the index formula for outer products in much greater
generality.]
Exercise 9.21. Let P ∈ Lkpc (E, F ) be an elliptic operator (i.e., assume that
σ(P ) ∈ IsoSX (E, F )). Show that index P = 0, if σ(P ) can be extended to an
isomorphism over all of BX := {ξ ∈ T ∗ X : |ξ| ≤ 1}.
[Hint: Show that a homotopy between the symbol of P and the symbol of the
multiplication operator (M u)(x) := σ(P )(x, 0)(u(x)) can be defined and apply
Theorem 9.10d and Exercise 9.16d.]
Exercise 9.22. Now, let E and F be trivial line bundles over the closed man-
ifold X. Show that index P = 0, if dim X > 2. [Hint: Reduce this to Exercise
9.21 by a suitable deformation of σ(P )(x, ξ); see [323, p.160f]. Compare also with
Section 13.3 below.]
Exercise 9.23. Carry the following theorems and exercises over to the case of
manifolds with boundary: Theorem 9.3, Exercise 9.4 (if the formal adjoint operator
is defined by (P u, v)0 = (u, P ∗ v)0 for all u and v with support contained in the
interior of X), Exercises 9.5–9.7, Theorem 9.8 (Why not Theorem 9.10?), Exercise
9.19, and Exercise 9.20b.

4. The Topological Meaning of the Principal Symbol — A Simple Case


Involving Local Boundary Conditions
I In this section, we shall explain the topological meaning of the principal symbol
by addressing a simple elliptic system of equations over a compact manifold with smooth
boundary, involving local boundary conditions.
To establish Fredholm properties and regularity results for elliptic operators, in this
chapter it was decisive that the operators acted on sections of vector bundles over closed
manifolds, i.e., manifolds that are compact and without boundary or any other singu-
larity. Historically, index theory began with the Noether(-Hellwig-Vekua) Theorem 5.11,
p.146, i.e., with a boundary value problem. In [102] of 1963, A.-P. Calderón proposed
to consider well-posed global, what we now call pseudo-differential boundary conditions
yielding Fredholm operators and regularity, the two essentials of the idea of ellipticity. With
hindsight, any rational discussion of elliptic boundary conditions should begin with that.
Clearly following his own nose, these ideas were worked out systematically by R.T. See-
ley in [379, 383] and found a spectacular application in the famous Atiyah-Patodi-Singer
Index Theorem for boundary conditions defined by spectral projections, see our Section
13.8, p.337ff in Part III. There we shall give a short summary of Seeley’s ideas and the
main subsequent results. For a comprehensive presentation we refer to the monograph
[83] and the more recent review [99].
In the preceding sections we saw that a topological object σ(P ) ∈ IsoSX (E, F ) is as-
sociated with each elliptic operator P : C ∞ (E) → C ∞ (F ), whereby index P depends only
on the homotopy type of σ(P ). Here SX is the covariant sphere bundle for a Riemannian
metric for X, and E and F are Hermitian vector bundles over the closed manifold X.
To show the beautiful relations between the local theory and global invariants and to
explain the topological meaning of the principal symbol, we switch to a simple boundary
value problem for a moment. J

A Simple Case Study. We consider, on a manifold X with boundary a local


boundary-value system (P, R) : C ∞ (E) → C ∞ (F ) ⊕ C ∞ (G), where E, F are vector
bundles over X, and G is a vector bundle over the boundary Y of X. The object
9.4. THE TOPOLOGICAL MEANING OF THE PRINCIPAL SYMBOL I 245

σ(P ) ∈ IsoSX (E, F ) is well defined, but it does not contain the necessary infor-
mation on the index of the boundary value problem (P, R) which may depend on
the specific choice of the boundary conditions R, by the Hellwig-Vekua Theorem
(Theorem 5.11, p.146). In the case of a differential operator P of first order we will
show, roughly, how the given suitable boundary conditions canonically determine
a continuation of σ(P ) beyond SX to the closed manifold SX ∪ BX|Y . We obtain
a topologically more significant object (conceptually: a closed line packs more topo-
logical information than an open one) whose homotopy type does in fact determine
index(P, R), as we will see in Section 13.8. Contrary to the technical explanations
of the algebraic meaning of ellipticity of boundary-value problems in the preceding
sections, we are concerned in the following case study with the geometric-topological
interpretation.
Theorem 9.24. An elliptic system P of N partial differential equations of
first-order over the compact, oriented, Riemannian manifold X together with a
system R of N/2 (suitable) boundary conditions over the boundary Y of X defines
(uniquely, up to homotopy) a continuous map of the closed manifold SX ∪(BX|Y )
into GL(2N, C) which coincides with σ(P ) ⊕ IdN on SX. Here, GL(2N, C) denotes
the group of complex, invertible 2N × 2N matrices and BX := {ξ ∈ T ∗ X : |ξ| ≤ 1}.
Proof. Verbal communication of I. M. Singer. (See also the elabora-
tions in [27, p.180-184], [109, 20/05-25/07] and [328, p.346-350]). Let
(P, R) : C ∞ (E) → C ∞ (F ) ⊕ C ∞ (G)
be a boundary-value system with P ∈ Diff 1 (E, F ) elliptic and E = F = CN X and
N/2
G = CY , N even. We shall specify the assumptions regarding the boundary
condition R below. Let ν ∈ Ty∗ X be the inner normal at the point y ∈ Y , see
Figure 9.1. Each covector ξ ∈ Ty∗ X can be written in the form zν + η with z ∈ R
and η ∈ Ty∗ Y , where Ty∗ Y can be taken to be a proper subspace of Ty∗ X by means
of the Riemannian metric (see Exercise 6.49, p. 191). We write σ(ξ) := σ1 (P )(y, ξ)
and obtain (since σ(ξ) is a homogeneous polynomial in ξ = (zν, η)) :
σ(zν + η) = σ(zν) + σ(η) = zσ(ν) + σ(η) : Ey → Fy (linear).
By a corresponding choice of basis for Fy , we may assume (without loss of general-
ity) that σ(ν) = Id, whence
(9.4) σ(zν + η) = z Id +σ(η).
Consider the space M+
η, consisting of the C ∞ functions h : R → E with σ(Dν +
d
η)h = Dh + σ(η)h = 0 which remain bounded as t → +∞; here D := −i dt . M+η
is spanned by the functions of the form h(t) = h0 eiλt , where λ is an eigenvalue of
the endomorphism σ(η) : Ey → Ey with Im λ > 0 and h0 ∈ Ey is an element of the
associated eigenspace. Correspondingly, define the space M− η . Thus, the spaces
±
Mη are naturally isomorphic to the sum of the (+) (resp. (−)) eigenspaces of σ(η).
For η 6= 0, the homomorphism λ Id +σ(η) = σ(λν + η) is regular for λ ∈ R by the
ellipticity of P ; i.e., σ(η) has no real eigenvalues and Ey can be represented as the
direct sum
(9.5) Ey ∼
= M+ −
η ⊕ Mη .
L
Now let G1 , ..., Gr be vector bundles over Y with Gj = G and
Rj : C ∞ (E) → C ∞ (G), j = 1, ..., r,
246 9. ELLIPTIC OPERATORS OVER CLOSED MANIFOLDS

be boundary conditions given by differential expressions such that associated initial


value map
(9.6) βη+ : M+
η → Gy

is an isomorphism for all y ∈ Y and η ∈ Ty∗ Y \ {0} . In the literature (see [83,
Remark 18.2c]), this is the condition of ellipticity for local boundary-value sys-
tems. We now show that σ(p)(y, ·) : (SX)y → Iso(Ey , Ey ) is (stably) homotopic
to a constant map, and that the homotopy is defined in a natural way by using
βη : η ∈ Ty∗ Y \ {0} . Thus, σ(P ) (more precisely σ(P ) ⊕ IdCN , see the Homotopy
 +

(SX )y
´
º
y
(BX )y

Figure 9.1. The cotangent ball BXy at y ∈ ∂X = Y with the


inner unit normal vector ν ∈ SXy and a unit vector η ∈ SYy

2 below) can be extended to (BX) |Y .


Homotopy 1: For η ∈ Ty∗ Y \ {0}, let πη± : Ey → M±
η ⊂ Ey be the projections
defined by (9.5) . By means of
(9.7) sσ(η) + (1 − s)(iπη+ − iπη− ), s ∈ I,
we obtain a homotopy of σ(η) to the map iπη+ − iπη− , see Figure 9.2; the geometric
meaning of this is that one may concentrate the eigenvalues of σ(η) on the eigen-

+i

{i

Figure 9.2. Contracting σ(η) to iπη+ − iπη−

values +i and −i which are independent of η, while the eigenspaces still depend on
− ±
η. By using h0 = h+ ±
0 + h0 ∈ Ey with h0 ∈ Mη one calculates that the eigenvalues
of the endomorphism defined in (9.7) always remain nonreal; thus, z Id + (9.7) is
nonsingular for z ∈ R. Hence, we have a homotopy in the space of elliptic symbols
(i.e., in (SX)|Y → GL(N, C) here) from σ to σ1 with σ1 (zν + η) := z Id +iπη+ −iπη− .
Homotopy 2: By means of e−iϕ iπη+ −eiϕ iπη− , ϕ ∈ [0, π/2], each σ1 (η) in Iso(Ey , Ey )
9.4. THE TOPOLOGICAL MEANING OF THE PRINCIPAL SYMBOL I 247

can be connected with the identity, but this deformation depends on the choice of
η and does not go through uniformly for all η ∈ Ty∗ Y \ {0}. GL(N, C) is too small
to implement the further homotopy. Hence, we enlarge σ1 by direct sum with
IdG ⊕ IdG to a map (SX)|Y → Iso(E ⊕ G ⊕ G, E ⊕ G ⊕ G), which we also denote
by σ1 . Because of the splitting Ey ∼
= M− + ∗
η ⊕ Mη , η ∈ Ty Y \ {0}, we have
 
−i 0 0 0
 0 i 0 0 
 0 0 1 0 ,
σ1 (η) =  

0 0 0 1

where the diagonal elements mean respectively the identities on M+
η or Mη or G
multiplied by the coefficients −i or +i or I. By the deformation
1 1
e−i 2 πs IdG ⊕ ei 2 πs IdG , s ∈ [0, 1] ,
we can uniformly deform σ1 (η) to
 
−i 0 0 0
 0 i 0 0  − +
 0 0 −i 0  on Mη ⊕ Mη ⊕ Gy ⊕ Gy .
σ2 (η) =  

0 0 0 i
Homotopy 3: Now it remains to deform σ2 (η) to a constant (σ2 (η) still depends
on the positions of the eigenspaces M± η ), in such a way that no real eigenvalues
appear and σ2 (zν + η) := z Id +σ2 (η) does not become singular for z ∈ R under
the deformation. This we achieve with the help of the boundary isomorphism
βη+ : M− ∼
η = Gy given by the ellipticity of the boundary-value problem in (9.6).
Indeed, there is a homotopy (which switches the second and third diagonal members
in σ2 (η)) of σ2 (η) to a constant map (on (SY )y )
   
1 0 0 0 −i 0 0 0
 −1
 0 0 βη+ 0  ◦ σ2 (η) =  0 −i 0 0 
 
η 7→ σ3 (η) := 
 0 β+  0
η 0 0  0 i 0 
0 0 0 1 0 0 0 i
which is multiplication by −i on Ey and by i on Gy ⊕ Gy ; the homotopy
 −1   
0 βη+ Id 0

βη+ 0 0 Id
follows from the Homotopy Lemma which says that
   
0 B Id 0

A 0 0 AB
in the space of automorphisms of the vector space V × W , if V and W are complex
vector spaces and A : V → W and B : W → V are linear with AB ∈ Iso(W, W ).
One proves the Homotopy Lemma by composing two homotopies: First connect
     
0 B i Id 0 (i sin ϕ) Id (cos ϕ) B
and via
A 0 0 iAB (cos ϕ) A (i sin ϕ) AB
ϕ ∈ [0, π/2], and then multiplication by e−iψ , ψ ∈ [0, π/2], provides the final
homotopy.
Homotopy 4: We have now deformed σ ⊕ Id to the constant map σ3 on (SY )y ;
248 9. ELLIPTIC OPERATORS OVER CLOSED MANIFOLDS

we extend this map to a map on all of (SX)y which is homotopic to σ1 ⊕ Id, by


means of the parametrization (see Figure 9.3)
(SX)y = {(cos θ) ν + (sin θ) η : η ∈ (SY )y and θ ∈ [0, π]} .
Namely, define on Ey ⊕ (Gy ⊕ Gy ) the automorphism

(SY )y

´
µ
º y
(SX )y

Figure 9.3. The parametrization of SXy by η ∈ (SY )y and θ ∈ [0, π]

 
cos θ − i sin θ 0
σ3 ((cos θ) ν + (sin θ) η) :=
0 cos θ + i sin θ
which by definition is homotopic to the constant map Id ⊕ Id. 
Generalizations (Heuristic). In the preceding proof, we have explicitly shown
with a sequence of homotopies how one can continuously extend the map σ ⊕
Id : SX → GL(2N, C) to a map σ e : SX ∪(BX)|Y → GL(2N, C) with σ e(y, 0) = Id2N
for all y ∈ Y :=boundary of X. In that proof, we have chosen formulations that
make the generalization for arbitrary bundles clear. To be sure, we must make
precise what we understand by stable homotopy; i.e., why we are content with an
extension of σ ⊕ Id2G : (SX)|Y → Iso(E ⊕ G ⊕ G, F ⊕ G ⊕ G) on (BX)|Y , even
though this object does not immediately extend over all of SX, since G is defined
only over Y . See also under Section 10.3. To bring about further generalizations
of the proof, we remark that the special form of the boundary conditions (which
were given by differential operators) played no role, since only the isomorphism
βη+ was needed; βη+ possibly could be defined through pseudo-differential boundary
conditions. The definition of the spaces M± η and the construction of the symbol
homotopies are made very easy by the polynomial form of σ(P ), i.e., its derivation
from a differential operator. This is why we have devoted so much space to this
case study. The construction of an extension of σ(P ) ⊕ IdN as an isomorphism over
(BX)|Y goes through more generally for elliptic pseudo-differential operators with
transmission properties which allow elliptic boundary problems; see Section 13.8,
p.337.
Remark 9.25. To each elliptic operator P ∈ Lkpc (E, F ) over the Riemannian
n-manifold X, one can assign a local index, namely the homotopy class of
σ(P )(x, ·) : (SX)x → Iso(Ex , Fx )
l l
S n−1 → GL(N, C)
where N is the fiber dimension of E. By the Bott Periodicity Theorem, whereby
(see Exercise 10.23, p. 271) for N sufficiently large (which we achieve here by adding
9.4. THE TOPOLOGICAL MEANING OF THE PRINCIPAL SYMBOL I 249

the identity) 
Z, for n even,
πn−1 (GL(N, C)) =
0, for n odd,
we obtain an integer deg(P ) for the local index; by continuity, it is independent of
the choice of x. If X is a manifold with boundary Y , then we can determine the
vector spaces M± + −
η and the integer µ(P ) := dim Mη − dim Mη , which is indepen-
∗ ∗
dent of the choice of η ∈ T̊ Y := T Y \ Y and does not automatically vanish for
n = 2.
If the operator P admits elliptic boundary conditions in the sense of the proof
of Theorem 9.24, then the condition µ(P ) = 0 follows, and by Theorem 9.24, the
condition deg(P ) = 0 holds. These two conditions are closely connected. By a
communication from M. F. Atiyah, in the special case N = 1, n = 2 (where
deg(P ) is the classical winding number of σ(P )(x, ·) in C \ {0} about the point 0),
the equation
(9.8) deg(P ) = ±µ(P )
holds. Also, in the general case, we have
(9.9) deg(P ) = ± deg M+
where deg M+ is the integer degree of a map fy : S n−3 → GL(M, C), y ∈ Y , which
is used to join together the trivial bundles over the upper and lower hemispheres
of S n−2 along the equator
n S n−3 (seeoAppendix B, Exercise B.9, p.718) to obtain
the bundle M+ := M+ : η ∈ T̊ ∗ Y over (SY )y ∼
y η y = S n−2 . Again, by continuity
arguments, it is clear that degree deg M+ does not depend on the choice of y in
fy . Equations (9.8) and (9.9) then represent a reformulation of the Bott Periodicity
Theorem. For this, see also [27, p.178], [109, 25-05], and in particular [328, p.351],
where a K-theoretic formulation of (9.8) and (9.9) is given.
As a result of Theorem 9.24, we have obtained (in deg(P ) 6= 0) a topological
obstruction to there being elliptic boundary conditions for the elliptic differential
operator P . In the classical theory E = F = CX , the obstruction can arise only
in the case n = 2, since for n ≥ 3 every homogeneous elliptic polynomial is of
even degree 2k and possesses an equal number k of zeros in the upper and lower
half-planes [217, p.246]. The situation is different for n = 2, where deg( ∂∂z̄ ) =
1. However, for all even n ≥ 4 there are elliptic differential operators P with
deg(P ) 6= 0; e.g., the Dirac operators (see [328, p.91f] or [68, p.26f]), defined using
m−1
the Clifford module R2 over X = R2m , have local index 1.
The example is comparable to the peculiarity shown by the Cauchy-Riemann
operator ∂∂z̄ in the classical theory for n = 2. Actually, there are many elliptic op-
erators which arise in Riemannian geometry, that have a nonvanishing local index,
and, for closed manifolds, characterize important topological invariants such as the
Euler number and signature by their global indices; see [68, p.30] and [41, p.46].
In the last work, as an expedient for the calculation involving the corresponding
topological invariants of manifolds with boundary, a nonlocal theory of boundary-
value problems is applied, with which one obtains a Fredholm theory in which the
above obstructions become irrelevant.
Exercise 9.26. Consider the disk X := {z ∈ C : |z| ≤ 1} with the circle Y :=
{z ∈ C : |z| = 1} as boundary. Go through the construction of Theorem 9.24 for
250 9. ELLIPTIC OPERATORS OVER CLOSED MANIFOLDS

the transmission operator (see Exercise 5.18)


 
∂u ∂v
T (u, v) := , , (u − v) |Y .
∂ z̄ ∂z
Remark 9.27. For the practical calculation of the index of an elliptic boundary-
value system, it suffices to reduce elliptic boundary-value systems of order k to
elliptic systems of order 1 on the symbol level. This deformation procedure on
the symbol level was introduced in [27, p.180f], and is entirely analogous to the
linearization of polynomial clutching functions, which plays a decisive role in one of
the proofs of the Bott Periodicity Theorem (see [28, p.241f]). It has the merit of
also carrying over to the case of arbitrary manifolds with boundary and nontrivial
bundles E, F , and G. This requires some standard tricks: One must identify the
bundles E and F in a neighborhood of the boundary Y of X, by means of the
isomorphism σ(P )(y, ν) : Ey → Fy , where ν is the inner normal at the boundary
point y, etc..
In such general cases it is possibly advantageous to go yet a step higher: One does
not specify a homotopy-theoretic extension of σ(P ) and IsoSX to an isomorphism
over SX ∪ BX|Y by means of the boundary isomorphism β, but rather one directly
constructs, from σ(P ) and β, a difference vector bundle in K(BX, SX ∪ BX|Y ), as
sketched in [328, p.346-351].
Beside these two ways of finding the correct topological object while preserving
the information about the index of an elliptic boundary-value problem, viz.
• linearization and continuation of the symbol by means of homotopies, and
• K-theoretic axiomatic characterization of a difference vector bundle,
there is a third way oriented more strongly towards functional analysis, viz.
• assignment of families of Wiener-Hopf operators to elliptic boundary-value
problems.
All three approaches depend not only on the Bott Periodicity Theorem (see
Chapter 10), but on essential elements of the different proofs. Conversely, the
Periodicity Theorem be derived via an investigation of a simple boundary-value
problem on the disk (see Exercise 10.24, p. 272).
In Theorem 4.4, we already provided the Index Theorem for Wiener-Hopf op-
erators. We will prove the Periodicity Theorem, following [20, p.116-120], as a
generalization of this theorem. Hence, the third approach via Hilbert space theory
is most convenient for us, and offers some formal advantages for the demonstration
of the analytical main theorems on the index of elliptic boundary-value problems.
Admittedly, this approach may be less conceptual than the first, and less elegant
than the second.
Part III

The Atiyah-Singer Index Formula

In perhaps most cases when we fail


to find the answer to a question,
the failure is caused by unsolved
or insufficiently solved simpler and
easier problems. Thus all depends
on finding the easier problem and
solving it with tools that are as
perfect as possible and with notions
that are capable of generalization.

David Hilbert, 1900

251
CHAPTER 10

Introduction to Topological K-Theory

Synopsis. Winding Numbers. 1-Dimensional Index Theorem. Counter-Intuitive


Dimension 2: Bending a Plane. The Topology of the General Linear Group. The
Grothendieck Ring of Vector Bundles. K-Theory with Compact Support. Proof of the
Bott Periodicity Theorem.
I It is the goal of this chapter to develop a larger portion of algebraic topology by
means of a theorem of Raoul Bott concerning the topology of the general linear group
GL(N, C), i.e., on the basis of linear algebra, rather than on the basis of the theory of
simplicial complexes and their homology and cohomology. There are several reasons for
doing so. First of all, it is a matter of taste and familiarity as to which approach one
prefers “codifying qualitative information in algebraic form” (Atiyah). In addition, there
are objective criteria such as simplicity, accessibility and transparency, which speak for
this path to algebraic topology. Finally, it turns out that this part of topology is most
relevant for the investigation of the index problem.
Before developing the necessary machinery, it seems advisable to explain some basic
facts on winding numbers and the topology of the general linear group GL(N, C). Note
that the group GL(N, C) moved to fore in Part II already in connection with the symbol
of an elliptic operator, and that the group Z of integers was in a certain sense the topic of
Part I, the Fredholm theory. In the following Part III, the concern is (roughly) the deeper
connection between the previous parts. We consider the Index Theorem of I.M. Singer
and M.F. Atiyah to be the correct and promising generalization of F. Noether’s (and
followers) theorems on the index of the (discrete) Hilbert transform and of the oblique
boundary value problem on the disk. See Chapter 4, Theorem 4.4 (p.125), Theorem
4.7 (p.127), Exercise 4.12 (p.129), and Theorem 4.14 (p.129), and Chapter 5, Theorems
5.11 (p.146) and 5.15 (p.152). To us, the geometric model of the Index Theorem is the
equivalence between local and global definitions of the winding number. J

1. Winding Numbers
I “How can numerical invariants be extracted from the raw material of geometry
and analysis?” (Morris Hirsch). A good example is the concept of winding number,
surely the best known item of algebraic topology: In his studies of celestial mechanics
the French physicist and mathematician Henri Poincaré turned to stability questions of
planetary orbits. Many of the related problems are not completely solved even today, e.g.,
the three body problem of describing all possible motions of three points which interact
via gravitation. See [306] for a review with references to ever new orbits discovered
theoretically and simulated on a computer. This problem, however, has been solved for
practical purposes through the last 50 years, as demonstrated by the unmanned soft
landing of the lunar module Luna 9 on February 3, 1966, and subsequent interplanetary
missions.
As a tool for the qualitative investigation of nonlinear (ordinary) differential equa-
tions, Poincaré introduced in 1881 the notion of the index I(P0 ) of a singular point P0

252
10.1. WINDING NUMBERS 253

for a system of two ordinary differential equations


ẋ = F (x, y), ẏ = G(x, y).
To do this, surround P0 by a closed curve C in the phase portrait (see the following
examples) and measure on it the angle of the rotation performed by the vector field
(F (x, y), G(x, y)) when (x, y) traverses C once counterclockwise. The angle is an integral
multiple of 2π, and this integer is I(P0 ). In case of a magnetic field one can actually see
I(P0 ) in the rotation of the needle, when a compass is moved along C. The sign of I(P0 )
says something about the geometry of the phase portrait close to an equilibrium position
P0 , but not necessarily something about stability (see [5, p.75-76] and Figure 10.1). We
will come back to this in Section 13.5.

P0

I(P0 ) = 1 =1 =1 =1

=1 = {1 = {2 =2

Figure 10.1. The index of a singular point of a system of two


ordinary differential equations, for eight qualitatively different sys-
tems

Poincaré returned to this topological argument in 1895, when he considered all


closed curves in an arbitrary space and classified them according to their deformation
properties.
The fundamental group was introduced in Poincaré’s work Analysis Situs (Oeuvres
6, pp.193–288), whose theme is purely topological-geometric-algebraic: an “analysis situs
in more than three dimensions.” Poincaré expected the abstract formalism to “do in
certain cases the service usually expected of the figures of geometry”. He mentioned three
areas of application: In addition to the Riemann-Picard problem of classifying algebraic
curves and the Klein-Jordan problem of determining all subgroups of finite order in an
arbitrary continuous group, he particularly stressed its relevance for analysis and physics:
“one easily recognizes that the generalized analysis situs would allow treating the equations
of higher order and specifically those of celestial mechanics” [the same way as H.P. had
done it before with simpler types of differential equations] “. . . I also believe that I did
not produce a useless work, when I wrote this treatise.” The complexity and limited
understanding of the topological problems did not however permit Poincaré to carry out
his program completely: “Each time I tried to limit myself I slipped into darkness.” J

Poincaré’s simplest result can be expressed in today’s terminology as follows


(see also [16, p.237-241]):
Theorem 10.1. Let f : S 1 → C× := C \ {0} be a continuous mapping of the
circle S 1 to the punctured plane of nonzero complex numbers C× . In other words,
254 10. INTRODUCTION TO TOPOLOGICAL K-THEORY

we have a closed path in the plane not passing through the origin. The following
hold:
i. The mapping f possesses a winding number which states how many
times the path rounds the origin; we write W (f, 0) or deg(f ).
ii. This degree is invariant under continuous deformations.
iii. The integer deg(f ) is the only such invariant, i.e., f can be deformed to
g, if and only if deg(f ) = deg(g).
iv. For each integer m, there is a mapping f with deg(f ) = m.
Arguments: Instead of a formal proof, we briefly assemble the different ways
of defining or computing deg(f ).
Geometrically: Replace f by g := f / |f |. This is a mapping from S 1 to S 1 .
Approximate g by a differentiable map h, and count (algebraically, i.e., with a sign
convention according to the derivative of h) the number of points in the preimage
of a point which is in general position. This method can also be characterized as
counting of the intersection numbers: Firstly, distort the curve in a suitable way so
that its arcs become visible and almost everywhere distinguishable (for instance,
by scaling with a factor 1 + χ(θ) where χ is a smooth bump function, constant
0 outside of [0, 2π], monotonously increasing on [0, π] and decreasing on (π, 2π]).
Then, draw an arbitrary ray emanating from the origin which does not pass through
an inflection or a self-intersection point of the path. Now count the intersections
of the path with the ray according to the rules of traffic of the right of way (H.
Weyl) — thus with a plus sign if the path has the right of way, and a minus sign
when the ray has the right of way (as in right driving countries, see Figure 10.2).
2

0 + +
+
{
+
2
+

Figure 10.2. The geometric count of the winding number of a


closed oriented curve in the punctured plane

Combinatorially (simplicial-homologically): We approximate f with a


piecewise linear path g, and then use combinatorial methods; i.e., we permit deleting
and adding of those edges of our polygonal path which are boundaries of 2-simplices.
These are triangles (or double edges with complementary orientation) whose interior
is completely contained in C× , and thus do not contain the origin (see Figure 10.3).
10.1. WINDING NUMBERS 255

0 0

1) 2)

0 0

3) 4)

Figure 10.3. The combinatorial (simplicial-homological) count of


the winding number by deforming (mild deformation — homotopy)
the curve into a piecewise linear figure in (2) and adding/removing
boundaries of 2-simplices (hard deformation — homology) in (3),
until two separate circles are obtained in (4)

HDifferential:
dg
We approximate f by a differentiable g, and then set deg(f ) :=
1
2πi 1
S g
. Here, we have regarded g as a map [0, 2π] → C× , and then the integral
R 2π g0(τ )
is defined as 0 g(τ ) dτ . From the Cauchy Integral Formula, it follows that the
integral is a multiple of 2πi, and hence deg(f ) is an integer.
Algebraic: Approximate f by a finite Fourier series
k
X
g(φ) = aν eiνφ , φ ∈ [0, 2π).
ν=−k

We regard g as a map S 1 → C× with S 1 = {z ∈ C : |z| = 1}. Consider now the


extension of g to the disk |z| < 1, where it is a finite Laurent series. We then obtain
a meromorphic function h, and set deg(f ) := N (h) − P (h) where N (h) and P (h)
denote the number of zeros and poles of h in |z| < 1.
Functional Analytic: Set deg(f ) := − index Tf , where Tf denotes the (dis-
crete) Wiener-Hopf operator, assigned to f , on the space L2+ (S 1 ) ⊂ L2 (S 1 ) spanned
by the functions z 0 , z 1 , z 2 , . . . . Then Tf is defined (for u ∈ L2+ (S 1 )) by
 P∞
ˆ
Td u(n) := k=0 f (n − k) û(k) , for n ≥ 0,
f
0, for n < 0,
256 10. INTRODUCTION TO TOPOLOGICAL K-THEORY

where fˆ(m) := hf, z m i denotes the m-th Fourier coefficient of f . For the details of
this, see Theorem 4.4 (p.125), and Theorem 4.14 (p.129) for the analogous repre-
sentation deg(f ) = index(I + Wφ ) via the (continuous) Wiener-Hopf operator
Z ∞
(Wφ u)(x) := φ(x − y)u(y) dy; x ∈ R+ , u ∈ L2 (R+ ),
0

where φ ∈ L (R) with φ̂ = f ◦κ−1 and κt := t−i


1
t+i denotes the Cayley transformation.
As a first example, one may recall the standard map a : S 1 → C× which is
given by a(z) = z, z ∈ C, |z| = 1. Here the equivalence of the various definitions is
clear. We omit the proof for complicated cases and refer to [9, p.151]. Compare also
Theorem 10.4 (p. 257) and Theorem 10.22 (p. 271) below and in another context
[206, p.120-131] or [97, p.161 f] .
Remark 10.2. In Theorem 10.1 all closed curves in the punctured plane are
compared and the essentially different ones separated. The different definitions
of winding number listed there reflect the main branches of topology with their
different techniques, goals and connections. Accordingly, depending on the point of
view chosen, very many generalizations of Theorem 10.1 to higher dimensions are
possible (see also Exercise 10.27, p. 274 and the pioneering work [261] by Leopold
Kronecker of 1869). As we shall see, the Index Theorem interrelates two of the
possible generalizations, a local one (coming from the integration over the symbol
form, yielding the topological index) and a global one (coming from the functional
analysis of Fredholm operators and yielding the analytic index).

R
R2

Figure 10.4. Left: There is just one way of bending the real line
into a closed manifold. Right: There are various possibilities of
bending the plane into closed manifolds

Remark 10.3. If one sticks with the classification of systems of ordinary dif-
ferential equations (which was the point of departure for Poincaré’s topological
papers) one would first try to distinguish the different possibilities of bending the
10.2. THE TOPOLOGY OF THE GENERAL LINEAR GROUP 257

real line into a closed curve in space or other higher dimensional spaces. In this fash-
ion H. Poincaré (but also see the Remark above) conceived (among other things)
the fundamental group π1 (X, x0 ) of a space X which arises from the homotopy
classification of closed paths S 1 → X which pass through the point x0 ∈ X. Here
only the embedding question (which depends on the structure of X) is of interest,
while the embedded images themselves of the compactified line are topologically
identical. The reason is that there is just one way of bending the real line into a
closed manifold, namely the form of a circle which may be traversed several times
and may wind so many times around one or the other hole, but still remains topo-
logically a circle. Roughly speaking, that means for ordinary differential equations
that global behavior can be revealed by local information in a finite number of test
points.
The situation is different, when we pass from ordinary to partial differential
equations. Here the classification essentially requires a differentiation between the
various possibilities of bending the plane or higher dimensional Euclidean spaces
into closed manifolds (see also [25]). Roughly speaking, that means that global
behavior for partial differential equations can not be revealed by punctual analysis
but will require some kind of integration of local information.
Now genuine global difficulties arise, since (as the diagrams of Figure 10.4
illustrate) there are already for the plane R2 different ways of bending it together.
For R2 it is still possible to survey completely the different forms which can be
classified according to the genus of the surface, i.e., the number of its handles (see
for instance [206, p.204 f]). The corresponding problem of the bending of R3 has
not been solved, although it is of special importance for the analysis of space-time
processes of the real world by means of partial differential equations. For example,
it took 100 years until the Poincaré conjecture (according to which every simply-
connected, three-dimensional, closed manifold is homeomorphic to the 3-sphere S 3 )
was confirmed in 2002/03 by Grigori Perelman in [333, 335, 334]. We shall not
comment on the proof in this monograph. 20 years earlier, the Poincaré conjecture
in four dimensions was established by Michael Freedman, see below Corollary
18.9 (p.648) in our Chapter 18 on Seiberg-Witten Theory.

The following generalization of Theorem 10.1 is due to Raoul Bott. It is a


true and fully understood achievement of topology, and in addition, touches on the
heart of the index prob1em for systems of elliptic differential equations.

2. The Topology of the General Linear Group


We consider continuous maps

f : S n−1 −→ GL(N, C), 2N ≥ n,

where S n−1 denotes the unit sphere in Rn and GL(N, C) denotes the general linear
group of all invertible linear maps from CN to CN .

Theorem 10.4 (R. Bott, 1958). If n is odd, each such f can be deformed to
a constant map. If n is even, one can define an integer deg(f ), such that f can be
deformed to another map g exactly when deg(f ) = deg(g). Moreover, there exist
maps having arbitrarily prescribed integer degree.
258 10. INTRODUCTION TO TOPOLOGICAL K-THEORY

Arguments. We will completely prove Theorem 10.4 in a different form below,


but emphasize now:
a) Theorem 10.1 is the special case of Theorem 10.4 for n = 2 and N = 1. For
n = 1 and N arbitrary, we recover the well-known fact that GL(N, C) is pathwise
connected; see the paragraph following Exercise 3.20, p. 73.
b) As we have formulated it here, the Bott Theorem is expressed in terms of
deformations (i.e., homotopies), a central theme of topology. In the formalism
of homotopy theory, the Bott Theorem says that for all n, N with 2N ≥ n the
homotopy groups πn−1 (GL(N, C)) (i.e., the group of homotopy classes of maps
S n−1 → GL(N, C)) are given as follows:

0 , if n is odd,
πn−1 (GL(N, C)) ∼
=
Z , if n is even.
In the second case, the isomorphism is given by

deg : πn−1 (GL(N, C)) −→ Z .

Using these concepts the classically expressed result of our Theorem 10.1 above
means that the first homotopy group (the fundamental group) of C× is isomorphic
to Z.
Theorem 10.4 yields an isomorphism of πn+1 (GL(N, C)) with πn−1 (GL(N, C)).
Therefore, it is also known as a periodicity theorem. Incidentally, there is a corre-
sponding theorem (with period 8) for GL(N, R). It has close connections with the
theory of real elliptic skew-adjoint operators; see Section 13.9 below.
c) As with Theorem 10.1 above, there are various ways in which the degree (for
even n) can be defined: First, a differential definition of deg(f ) is possible with
the help of a known, explicitly defined, invariant differential form ω on the C ∞
manifold GL(N, C). For the not entirely simple definition of this world constant
(Weltkonstante, F. Hirzebruch), we refer to [208, p.587 f]. One then sets
Z
deg(f ) := f ∗ (ω) ,
S n−1

where f ∗ (ω) denotes the pull-back form over S n−1 , and shows (!) that the invariant,
so defined, is an integer. As an alternative to this direct, somewhat computation-
ally cumbersome definition, one can also define deg(f ) geometrically, by means of
a stepwise reduction to the more intuitive, but topologically no less demanding,
notion of mapping degree of a continuous mapping of the (n − 1)-sphere into
itself. (In Theorem 10.1 the two concepts still coincided.) First one shows that,
without loss of generality, one can take 2N = n, since in the case 2N > n, f can
be deformed into a map of the form
" #
h(x) 0
g(x) = 0 IN- n ,
2

where h : Sn−1
→ GL(n/2, C) and IN- n denotes the (N − n2 ) × (N − n2 ) unit matrix.
2
All further constructions do not depend on the choice of g, since the preceding
factorization gives, more precisely,

πn−1 (GL(N, C)) ∼


= πn−1 (GL(n/2, C)) for N ≥ n/2.
10.2. THE TOPOLOGY OF THE GENERAL LINEAR GROUP 259

For those with enough background, we offer the following explanation of this: Form
the exact homotopy sequence
πi+1 (U(m) , U(n)) −→ πi (U(n)) −→ πi (U(m)) −→ πi (U(m) , U(n))
induced by U(n) ,→ U(m) for n ≤ m. Recall: U(n) := {A ∈ GL(n, C) : AA∗ = I}.
Form the exact sequence
πi+1 S 2n+1 −→ πi (U(n)) −→ πi (U(n + 1)) −→ πi S 2n+1
 

from the fibration U(n) → U(n + 1) → S 2n+1 . We have


πi+1 S 2n+1 ∼= πi S 2n+1 = 0,
 
for i < 2n;
namely, show with the Sard Theorem (after differentiable approximation) that on
account of different dimensions, not every map in such homotopy classes can be
surjective, whence one always has a point in the complement of the image in S 2n+1
from which one can contract. Thus, it follows that πi−1 (U(m)) ∼ = πi−1 (U(n)) for
1 ≤ i ≤ 2n ≤ 2m. Since the unitary group U(n) is a deformation retract of GL(n, C)
(prove!), the statement follows; see [404, 5.6 and 19.5].
Hence, let 2N = n. Then the first column of the matrix f defines a map
f1 : S n−1 → CN \ {0}. Since CN \ {0} = Rn \ {0}, we have a map
g := f1 / |f1 | ,
for which a degree (the mapping degree, the most natural generalization of Theorem
10.1) may be easily defined: One approximates g by a differentiable map h, and
chooses a point y ∈ S n−1 in general position. This means that the differential
h∗x (see Section 6.2 above) has maximal rank for all x ∈ h−1 (y); i.e., h∗x is an
isomorphism of tangent spaces. Then one counts the algebraic number of points in
h−1 (y) (i.e., the points x where h∗x reverses orientation are counted negatively).
Details for this are found in [206, p.121-131]. For other definitions of the mapping
degree of g, see [97, 14.9.6-10] (on intersection numbers – thus again a geometrical
definition, but in a more general setting), and [140, p.304 ff] (on the homology
of S n−1 ; i.e., fundamentally a combinatorial definition); furthermore, see Exercise
10.27, p. 274 (K-theoretical definition).
It turns out that the mapping degree of g
(I) contains the essential information on the qualitative behavior of f and
(II) is always divisible by (N − 1)!, so that we then define
(−1)N −1 deg(g)
deg(f ) := ,
(N − 1)!
where the sign (−1)N −1 is there for secondary technical reasons.
One can perhaps best visualize (I) (concerning the term visualize see d) below;
strictly speaking it is more a comprehension aid) as follows: Topologically, the linear
independence of the column vectors of a matrix means approximately the same as
their being perpendicular. Thus, if such is the case, each function S n−1 → GL(N, C)
is so rigid that it can be completely classified topologically by means of a single
column.
In fact (and this now concerns (II)) the maps S n−1 → GL(N, C) are so rigid
that in general, by far not every function S n−1 → CN \ {0} can appear as the first
column. What is the reason? Isn’t it possible to make an invertible matrix out
of any nonzero vector by adding orthogonal vectors? Yes and no: It is possible to
260 10. INTRODUCTION TO TOPOLOGICAL K-THEORY

do so at each single point, but not if the additions are to be made uniformly in
continuous dependence on the points of S n−1 . One may think of a sphere that has
no unit tangential vector field and thus eliminates the identity map id : S n−1 →
Rn \ {0} = CN \ {0} as a possible first column of an A : S n−1 → GL(N, C).
Otherwise assume (or deform to) A : S n−1 → U(N ) with A(z)(e1 ) = z for all
z ∈ S n−1 , where e1 denotes the first unit vector in Rn . Then, by A(z) unitary, the
vector field z 7→ A(z)b is continuous and orthogonal to z and hence a tangential
unit vector field for any choice of a unit vector b ∈ Rn orthogonal to e1 . Such a
sphere is the 2-sphere (see Exercise 10.28, p. 274), and we can not exclude that also
some of the higher-dimensional spheres share that property, namely not admitting
tangential unit vector fields.
For an evaluation of the difficulty of the divisibility theorem (II), see [207, Ap-
pendix p.182]. By Friedrich Hirzebruch the theorem was proven in a somewhat
different form “pretty much at the end of the study [on the Theorem of Riemann-
Roch] as a corollary of cobordism theory”. Hirzebruch added with hindsight,
that the divisibility theorem does not belong “at the end, but at the beginning”,
namely within “the Bott periodicity theory which is the basis of the newer proofs
of the Riemann-Roch Theorem”.
Note. With hindsight, in the language of Section 13.1, the facts are the fol-
lowing: The Chern character
e 2N ) → H 2N (S 2N ; Q) is given by
chN : K(S (−1)N −1 cN/(N −1)!,

since all other terms in the formula for chN cancel because of Bott periodicity
which here takes the form K(S e 2N ) ⊗ Q ∼ = H 2N (S 2N ; Q). If E denotes the vector
2N
bundle over S of complex dimension N , which is constructed by gluing with
f : S 2N −1 → GL(N, C), then (regarding the sign see above) deg f = chN ([E] −
[CNS 2N ]), where [E] denotes the class of E in K(S
2N
); furthermore the N th Chern

class cN (E) is mapped to deg(g) under the isomorphism H 2N (S 2N ; Z) −→ Z, where
2N −1 2N −1
g: S →S is defined by means of f as before.
d) What can be said so far on the substance of the Bott Periodicity Theorem? We
stay in the case n even, 2N ≥ n. Then, we have three statements:
(i) The mapping degree : πn−1 (GL(N, C)) → Z is a well defined homomor-
phism,
(ii) the map is surjective, and
(iii) the map is injective.
As we have seen, for (i) two approaches are available, a differential one and a
geometrical one. In this way, the degree was defined one time as an integral of
a differential form (hence it is, a priori, a real number), and the other time as a
quotient of two integers (a priori, as a rational number). Either way, statement (i)
causes no difficulties except for the integrability and divisibility theorems needed,
since the homotopy invariance of the degree is rather clear from the definition. One
might simply say: (i) concerns the definition of an integer homotopy invariant
— this is homology and comparatively simple. The additivity under catenation
of two mappings (the catenation suitably to be defined), i.e., the homomorphism
property follows directly from the definition(s).
Also (ii) is comparatively simple: Namely, one can form (as when tensoring
elliptic operators in Exercise 9.20, p. 243 and complexes in Remark 11.3, p.278) the
10.2. THE TOPOLOGY OF THE GENERAL LINEAR GROUP 261

outer tensor product of f : S n−1 → GL(N, C) and g : S m−1 → GL(M, C),


(10.1) f #g : S m+n−1 −→ GL(2M N, C), defined by
f (x) ⊗ IM − IN ⊗g ∗ (y)
 
(f #g)(x, y) := ,
IN ⊗g(y) f ∗ (x) ⊗ IM
where f and g are extended homogeneously to all of Rn \{0} , respectively Rm \{0}.
From the simple multiplication formula deg(f #g) = (deg f )(deg g) it then follows
that deg(ak ) = 1, where a : S 1 → GL(1, C) denotes the standard map z 7→ a(z) := z
of degree 1 and the map ak denotes its k-fold power:
k times
ak := a# · · · #a : S 2k−1 −→ GL(2k−1 , C).
Thus, product theory furnishes a generating map, meaning a generating element of
the infinite cyclic group of homotopy classes of maps S n−1 → GL(N, C) for n = 2k
of degree 1, whence the surjectivity of degree follows.
The statement (iii) is in contrast highly nontrivial: While in (i) and (ii) one
had to construct particular objects (a homotopy-invariant number and a suitable
homotopically nontrivial map), one must now show that mappings of the same
degree may be deformed into each other, in particular, that each map of degree 0
is homotopic to a constant map. Having defined the invariants, one wants to verify
their relevance and determine their value. This is exactly homotopy and not at all
intuitive. It can be seen from the fact that, although the homotopy classes of maps
S n−1 → GL(N, C) (for 2N < n) as well as those from spheres to spheres are in a
fashion closer to our space-time comprehension, they are largely unknown because
of their extreme complexity.
e) Without commenting on the very different proofs of the Periodicity Theorem
which exist to date, we would like to point out that all proofs proceed by induction
on n, more precisely by induction from n to n + 2. In the language of the outer
product theory presented above, it must be shown that f 7→ f #a is an isomorphism
of the homotopy group of dimension n − 1 with the homotopy group of dimension
n + 1.
f ) Interestingly, the algebraic definition of winding number (see Theorem 10.1) was
initially not susceptible to generalization to the present situation. However, the
discovery of the topological significance of elliptic boundary-value problems (see
above Theorem 9.24 and in retrospect and more generally [83], and below, Section
13.8) led Raoul Bott and Michael F. Atiyah to a new and elementary proof of
the Periodicity Theorem. Elementary in comparison with the original proof which
employed essentially the tools of the modern calculus of variations due to Marston
Morse; see, e.g. [297, p.124-132]. In a very deep sense, which we will explain below
in a particularly suitable form when presenting the proof, this proof generalizes and
unifies the algebraic and functional analytic definitions and furthermore “gets to
the heart of the problem” (Atiyah).
I The Lie group GL(N, C) and its real analogue play an important role everywhere
in mathematics. Accordingly the Periodicity Theorem and its off-spring K-theory (see
below) has had immediate applications to the index problem of elliptic harmonics (see
particularly Chapters 11-13 below) and has also proved to be a useful tool for a number
of deep geometric problems. An example is the computation of the number of linearly
independent vector fields on a sphere which the British mathematician John Frank
Adams carried out in 1962 with just these tools. A detailed exposition of this and some
other applications can be found in [228, Ch.15].
262 10. INTRODUCTION TO TOPOLOGICAL K-THEORY

The deep relation between the topology of the general linear group and the geometry
of differentiable manifolds, which becomes manifest in these successes, can be described
intuitively as follows: Among the simplest global topological invariants of a compact
oriented n-dimensional manifold X is the Euler characteristic χ(X). A famous theorem,
proven in 1895 by Henri Poincaré for n = 2, and in general by Heinz Hopf in 1925,
says that χ(X) can be found by means of a differentiable structure on X as the number
of singularities. Here a singularity is an isolated zero x of a tangent vector field v on X,
and the counting must be done with the proper multiplicity, namely the (local) index of
v at x (i.e., the degree of the mapping S n−1 → S n−1 ) which is given by v on the surface
of a ball about x. For a conceptually very plausible proof for n = 2 see [93, p.166-171],
see also Sections 13.4 and 13.5 below.
Many global topological invariants are known which are defined on X by means of
a classical (Riemannian or complex) structure, see in our Chapters 13 and 15–18 and
the classic [207]. These characteristic classes are without exception generalizations of
the Euler characteristic since “roughly, one considers the cycles where a given number
of vector fields become dependent”, as [19, p.59] remarks. The question of which way
a system of linearly independent vectors can become linearly dependent, forms the link
between topology and GL(N, C) (and GL(N, R)). J

3. Elementary K-Theory
The Ring of Vector Bundles. Let X be a compact topological space. From
our common knowledge summary in the Appendix, Definition B.6 (p.716) we recall
that we denote by Vect(X) the abelian semi-group of isomorphism classes of com-
plex vector bundles over X. If X consists of a single point, then Vect(X) ∼
= Z+ .
Now we generalize the construction which one uses to go from the semi-group Z+
to the group Z, in such a way that we can assign a group K(X) to the semigroup
Vect(X).
Theorem 10.5. Each abelian semi-group A (with zero element) yields in a
canonical way an abelian group GA := A × A/∼ and a semi-group homomorphism
ϕA : A → GA induced by a 7→ (a, 0). Here ∼ denotes the equivalence relation on
A × A defined by
(a1 , a2 ) ∼ (a01 , a02 ) ⇐⇒ ∃a, a0 ∈ A such that (a1 ⊕ a, a2 ⊕ a) = (a01 ⊕ a0 , a02 ⊕ a0 ) .
Proof. Let ∆ : A → A × A denote the diagonal homomorphism a 7→ (a, a) of
semi-groups. Then we set
GA := {(a1 , a2 ) + ∆(A) : ai ∈ A, i = 1, 2} ,
where
(a1 , a2 ) + ∆(A) := {(a1 ⊕ a, a2 ⊕ a) : a ∈ A} .
GA is a quotient semi-group in which there is an inverse for each element given by
−(a1 , a2 ) + ∆(A) = (a2 , a1 ) + ∆(A).
Thus, GA is a group. In this notation, the semi-group homomorphism is given by
ϕA (a) := (a, 0) + ∆(A). 
Remark 10.6. It is advisable to go through the definition of GA carefully as
we did in the proof, since the intuition one gains in the transition from Z+ to Z is
partly deceptive: Namely, a semi-group is not always embedded in a group. (The
10.3. ELEMENTARY K-THEORY 263

cancellation rule must already hold in the semi-group). The natural homomorphism
ϕA : A → GA is not necessarily injective; see Exercise 10.10 below.
Remark 10.7. In a certain sense, GA is the best possible group that can be
made from the semi-group A. More precisely, the following universal property
of ϕA holds: Every semi-group homomorphism h : A → H from A to an arbitrary
group H can be factored through ϕA in exactly one way; i.e., there is exactly one
group homomorphism h0 : GA → H such that the adjacent diagram

A
ϕA
/ GA

h0
h
! 
H
is commutative. The universal property is a generalization of the observation that
ϕA becomes an isomorphism, if A is already a group. That observation follows
from the functorial property of the assignment A 7→ (GA, ϕA ) on the category of
semi-groups. Alternatively, GA can be defined by generators and relations, namely
as a quotient group F A/RA, where F A denotes the free abelian group on A (which
consists of all finite linear combinations of elements of A with coefficients in Z),
and RA denotes the subgroup of F A generated by the subset
{1(a1 ⊕ a2 ) +(−1) a1 +(−1) a2 } .
The mapping ϕA : A → F A/RA is defined in the natural way and fulfills the
homomorphism condition ϕA (a1 ⊕ a2 ) = ϕA (a1 ) + ϕA (a2 ), since ϕA (a1 ⊕ a2 ) −
ϕA (a1 ) − ϕA (a2 ) is represented by 1(a1 ⊕ a2 ) +(−1) a1 +(−1) a2 . That ϕA is the
universal solution of the factorization problem is shown as follows: The unique-
ness of h0 is clear, since “h0 (ϕA (a)) = h(a) for a ∈ A” implies that h0 is already
given on a set of generators of F A/RA, whence h0 itself is uniquely defined. From
h0 (ϕA (a1 ⊕ a2 )) − h0 (ϕA (a1 )) − h0 (ϕA (a2 )) = 0, it follows that h0 is a group homo-
morphism. From the universal property, it follows easily that the two methods of
group construction are equivalent, and in particular that GA and F A/RA are iso-
morphic. In particular, GA may be thought of as the group of equivalence classes of
formal differences of elements of A, i.e., considering (a1 , a2 ) as a1 − a2 , see Exercise
10.9.
Now let A := Vect(X) (see our Definition B.6, p.716, in the Appendix), and X
compact.
Definition 10.8 (K-Theory). We denote the associated abelian group GA
by K(X). Then, for each vector bundle E over X, we obtain (by means of ϕA )
an element [E] ∈ K(X), and every element of K(X) can be written as a linear 
combination of such elements; see also Exercise 10.12a below. For the class CN
X
generated by the trivial vector bundle CN
X of complex fiber dimension N over the
basis X, we also simply write N .
Exercise 10.9. In the formalism developed here, describe the canonical ex-
tension of subtraction δ : Z+ × Z+ → Z to the difference-bundle construction
Vect(X) × Vect(X) → K(X). First show K(X) ∼ = Z, if X denotes a point.
Exercise 10.10. Show: The cancellation rule does not always hold in Vect(X).
[Hint: First illustrate with the two real bundles T S 2 and R2S 2 over the 2-sphere.
264 10. INTRODUCTION TO TOPOLOGICAL K-THEORY

These are not isomorphic (see Exercise 10.28, p. 274), but forming the direct sum
of each with the trivial line bundle RS 2 , we arrive at isomorphic bundles: with the
tangent bundle T S 2 consider the direct sum with the normal bundle N S 2 of the
canonical embedding of S 2 in R3 . In general, search for a nontrivial vector bundle
that becomes trivial when a trivial bundle is added to it. A detailed discussion
of special cancellation type rules can be found in [228, Ch.8]; for example, the
Uniqueness Theorem for Complex Vector Bundles says that trivial bundles over
manifolds of dimension n may be cancelled when the other summand has fiber
dimension ≥ n/2.]
Exercise 10.11. a) Show that each element of K(X) can be written in the
form [E] − N , where E ∈ Vect(X) and N ∈ N.
b) Show that two vector bundles E and F define the same element of K(X) (i.e.,
[E] = [F ]) exactly when E ⊕ CN N
X = F ⊕ CX , for some N .
c) One says that the bundles E and F are stably-equivalent when there are
natural numbers M and N such that
E ⊕ CN ∼
= F ⊕ CM .
X X

Show that the set I(X) of stable equivalence classes forms a group relative to the
operation of direct sum. [Hint for c): See Appendix, Exercise B.12, p. 719.]
Exercise 10.12. a) Show that, by means of the tensor product ⊗ for vector
bundles (see Appendix, Exercise B.4, p. 715), a multiplicative structure for K(X)
is furnished, making it a commutative ring with unit [CX ].
b) Show that each continuous map f : Y → X induces a ring homomorphism
f ∗ : K(X) → K(Y ) which only depends on the homotopy class of f .
c) Let i : Y → X denote the inclusion of a closed subset Y of X. Define relative
K-theory by setting K(X, Y ) := Ker(K(X/Y ) → K(Y /Y )), where X/Y denotes
the space obtained from X when Y is collapsed to a point {Y /Y }.
i. Let Y consist only of a single point x0 . Show that the group K(X, Y )
forms an ideal in K(X), and that K(X) splits into a direct sum
K(X) ∼ = K(X, x0 ) ⊕ K(x0 ) ∼ = K(X, x0 ) ⊕ Z = K(X)e ⊕ Z,
where K(X)
e := K(X, x0 ) is the essential part of K(X) and is isomorphic
to I(X).
ii. In general, define a natural map j ∗ : K(X, Y ) → K(X), and show that
j∗ i∗
the short sequence K(X, Y ) → K(X) → K(Y ) is exact.
[Hint for a): Use the universal property (Remark 10.7) to factorize Vect(X) ×
Vect(X) → K(X) through K(X) × K(X). For b): See Appendix, Theorem B.10,
p. 718; in particular, K(X) ∼ = K(Y ) when X and Y are homotopy equivalent.
For c): Work with the retraction r : X → x0 for the splitting in (i). See the
preceding Exercise 10.11 for I(X). First define j ∗ in (ii) more generally, when
j : (X 0 , Y 0 ) → (X, Y ) is a map of pairs of spaces (i.e., j : X 0 → X continuous with
j(Y 0 ) ⊆ Y ). Then set X 0 = X and Y 0 = Y . To check Im(j ∗ ) ⊂ Ker(i∗ ), factor

(Y, ∅)
j◦i
/ (X, Y )
O

$
(Y, Y )
10.3. ELEMENTARY K-THEORY 265

and note that K(Y, Y ) = 0. To prove the other direction, work with the difference
representation as in Exercise 10.11a.]
K-Theory and Functional Analysis. What is so special about K-theory
for the global analysis of elliptic operators? A first answer is given by the following
theorem.
Theorem 10.13. Let X be a compact space and let [X, F] denote the set of
homotopy classes of continuous maps T : X → F, where F denotes the space of
Fredholm operators in a Hilbert space H. The construction of index bundles (see
Section 3.7 above) induces a bijective map index : [X, F] → K(X) under which com-
position in F and addition in K(X) correspond, as do adjoints in F and negatives
in K(X).
Proof. See Theorem 3.40, p. 88. 
There is more about that relationship. In the proof of Theorem 10.5 and
the subsequent Remark 10.7, we learned two different constructions of the group
K(X) and of these the first is probably the most natural in connection with func-
tional analysis (as in Theorem 10.13). On the other hand, the second construction
immediately yields the universal property which historically motivated this for-
mal group construction in the papers of Claude Chevalley and Alexander
Grothendieck concerning algebraic geometry. It arose as a tool for the study of
problems involving functions that are additive on a semigroup with integral values.
This may also explain the relevance of K(X) for our index problem of el-
liptic operators: In Part II, we associated with each elliptic pseudo-differential
operator P : C ∞ (E) → C ∞ (F ) (E and F complex C ∞ bundles over the closed,
C ∞ Riemannian n-manifold X) its symbol σ(P ) ∈ IsoSX (E, F ), and we proved
that index P only depends on the homotopy type of σ(P ). Now let S 0 X :=
B + X ∪SX B − X denote the n-sphere bundle over X which arises by gluing two
copies B + X and B − X of the covariant unit-ball bundle
BX := {(x, ξ) : x ∈ X and ξ ∈ Tx∗ X with |ξ| ≤ 1}
along their common boundary SX. We lift E over B + X and F over B − X and glue
them (see Appendix, Exercise B.9, p. 718) over SX by means of σ(P ). This way we
σ(P )
obtain a vector bundle on S 0 X whose isomorphism class [E → F ] only depends on
the homotopy type of σ(P ). Conversely, the space of pseudo-differential operators
is so rich that σ(P ) has any desired homotopy type for suitable P . Because of the
special form of S 0 X, we therefore obtain all isomorphism classes of vector bundles
over S 0 X in this fashion. Thus the theory of elliptic equations yields a semigroup
homomorphism index : Vect(S 0 X) → Z which fits into the following diagram
Ell(X) / Vect(S 0 X) / K(S 0 X)

index index K−index


&  x
Z.
Here Ell(X) denotes the class of elliptic pseudo-differential operators on the closed
manifold X. Thus the universal property of K(S 0 X) guarantees the existence and
uniqueness of a K-index, which makes the diagram commutative. Being a group
homomorphism, the K-index is easier to analyze than the two index mappings,
especially as more is known about the group K(S 0 X).
266 10. INTRODUCTION TO TOPOLOGICAL K-THEORY

Remark 10.14. a) The commutativity of the preceding diagram is a direct


consequence of Theorem 10.13 and, after all, a rather simple consequence of the
given constructions and, most of all, Dieudonné’s discovery of the local constancy
(i.e., homotopy invariance, Theorem 3.11, p.68) of the index of Fredholm operators.
The preceding diagram does not give the essence of the Index Theorem. Beware,
the diagram leads only to the definition of the analytical index on suitable K-
groups. The analytical index is easy to define and almost impossible to calculate
directly. The very achievement by I.M. Singer and M.F. Atiyah was to find
a topological index, i.e., expressing the analytical index by an integral which is
explicitly calculable in many cases.
b) Presently, three different ways are known for associating with an elliptic operator
P a K-theoretic object via its symbol. The construction presented here which ends
σ(P )
up in K(B + X ∪SX B − X) is the most conceptual, since the class [E → F ] ∈
K(B + X ∪SX B − X) can be represented directly by a vector bundle on B + X ∪SX
B − X. The other constructions yield objects in K(BX, SX) or (equivalently, see the
following section) in K(T X) and basically amount to forming the difference class
σ(P )
[E → F ] − [F ]. The advantage of this nonconceptual construction of a difference
bundle (for details see Chapters 11 and 12 below) rests on the fact that the object
σ(P )
[E → F ] contains too much useless information on the specific form of the vector
bundles E and F which is completely irrelevant to the index problem. For example,
one can make E (or F ) trivial by adding a vector bundle V (Appendix, Exercise
B.12, p.719) while the index of P ⊕ IdV : C ∞ (E ⊕ V ) → C ∞ (E ⊕ V ) does not
change. More precisely: By evaluating the row-exact commutative diagram
o K(B + X) o K(B + X ∪SX B − X) o K(B + X ∪SX B − X, B − X) o
6

= ∼
=
  
K(X) / K(X) K(B + X, SX)
id

one finds that K(B X ∪SX B − X) ∼


+
= K(X) ⊕ K(BX, SX), whereby the first sum-
mand is irrelevant for the index problem and the second one is best dealt with in
the nonrelative form K(T X) in the framework of K-theory with compact support.
K-Theory and Cohomology. Exercise 10.11 says that K is a contravariant
functor on the category of compact topological spaces and continuous maps into
the category of commutative rings (with identity) and ring homomorphisms, and
this functor bears great resemblance (also in aspects not explained here) with the
cohomology functor H ∗ ; see [140, p.13 f]. The characteristic classes of vector
bundles (once again, see in our Chapters 13 and 15–18 and the classic [207]) yield
a variety of interesting operations K → H ∗ , and one can show that in fact K(X)⊗Z
R∼ = H even (X; R); see [17].

4. K-Theory with Compact Support


Until now, we have only defined K(X) for compact X. For locally compact X,
i∗ 
we now set K(X) := K(X + , +) = Ker K(X + ) → K(+) , where X + := X ∪ {+}
denotes the 1-point compactification of X (by the addition of the point +, where
i : {+} ,→ X + denotes the inclusion of this point). For compact X, this definition
adds nothing new to Definition 10.8. Alternatively, K(X) can be expressed in terms
10.4. K-THEORY WITH COMPACT SUPPORT 267

of complexes of vector bundles — see [18, p.489 ff]; then, the elements of K(Y \Z) =

K(Y, Z) can be taken to be equivalence classes of isomorphisms σ : E|Z −→ F |Z ,
where E and F are (complex) vector bundles over the compact set Y with closed
subset Z. Consider, e.g., the symbol of an elliptic operator over the manifold X
and set Y := BX and Z := SX, where Y \ Z is then diffeomorphic to the full
cotangent bundle T ∗ X.
Exercise 10.15. a) Verify that K(X) is a ring (without unit element) when
X is noncompact.
b) Show functoriality for proper maps f : Y → X; these are the maps which can
be continuously extended to Y + with values in X + .
[Hint for b): One may also define f to be proper exactly when f −1 (K) is compact
for all compact subsets K ⊆ X. From this comes the notion of K-Theory with
compact support — see also [140, p.5 and 269 ff]. In particular, each homeomor-

phism is proper, and we have K(X) −→ K(Y ) for homeomorphic X and Y , and
f ∗ = Id, if Y = X and f is homotopic (within the class of homeomorphisms) to
the identity. On the other hand, the mere homotopy type of X does not determine
K(X). (Example: K(R)  K(+)).]
Theorem 10.16. If X and Y are locally compact spaces, then (in addition to
the ring structures of K(X) and K(Y )) there is an outer product
 : K(X) ⊗ K(Y ) −→ K(X × Y ).
Remark 10.17. The outer product admits a particularly simple and natural
definition, if one adopts the above introduction of K(X) via complexes, and forms
the tensor product of complexes; see [44, p.490]. See also the closely related outer
tensor product for elliptic operators in Exercise 9.20c (p. 243) and for matrix-valued
functions in Equation (10.1) (Section 10.2, p.261 above).
Proof. Step 1) If X and Y are compact, then  is defined by forming the
vector bundle E  F over X × Y , where E is a vector bundle over X, F is over Y ,
and E  F has fiber Ex ⊗ Fy over (x, y).
Step 2) To carry this definition over to locally compact X and Y , we prove the
exactness of the short sequence
(10.2) 0 −→ K(X × Y ) −→ K(X + × Y + ) −→ K(X + ) ⊕ K(Y + ),
whereby K(X × Y ) is identified with the subgroup of K(X + × Y + ) which vanishes
on the axes X + and Y + . For this, we begin with the short exact sequence
j∗ i∗
(10.3) K(A, B) −→ K(A) −→ K(B)
from Exercise 10.12c (above) for compact topological spaces A and B with i : B ,→
A for the case where B is a retract of A; i.e., there is a continuous map r : A → B
which is the identity on B. Then ri = Id on B and i∗ r∗ = Id on K(B); thus,
we see that i∗ is surjective and r∗ is injective. Furthermore, we obtain (prove!) a
γ : K(A) → K(A, B) with γj ∗ = Id, whence j ∗ is injective. One says: The sequence
(10.3) splits (see [140, p.229f]), and we obtain a decomposition

(10.4) K(A) −→ K(A, B) ⊕ K(B).
To get (10.2), we apply (10.4) twice. First with
(10.5a) A := X + × Y + and B := X + × {+} ,
268 10. INTRODUCTION TO TOPOLOGICAL K-THEORY

next with
A := X + × Y + /X + and B := Y + .

(10.5b)
Since B is a retract of A in both cases, (10.3) splits and we obtain the formulas

(10.6) K(X + × Y + ) −→ K(X + ) ⊕ K(X + × Y + , X + )
and
K (X + × Y + )/X + ∼
= K(Y + ) ⊕ K (X + × Y + )/X + , Y + ,
 
(10.7)
from which the desired splitting of (10.2)
K(X + × Y + ) ∼= K(X + ) ⊕ K(Y + ) ⊕ K(X × Y )
follows because K (X + × Y + ) /X + , Y + ∼

= K((X ×Y )+ , +) = K(X ×Y ). One can
check that the splitting is compatible with the naturally defined arrows in (10.2).
Step 3) Now let x ∈ K(X) ⊆ K(X + ) and y ∈ K(Y ) ⊆ K(Y + ). Then x  y ∈
K(X + × Y + ) is well defined by Step 1). Actually, x  y can be regarded also as
an element of K(X × Y ) by (10.2), since i∗ (x  y) = 0, where i : X + ,→ X + × Y +
denotes the canonical inclusion (and correspondingly for Y + ,→ X + × Y + ). For
a proof of this, we write (Exercise 10.11a, p. 264) x = [E] − N and y = [F ] − M ,
where E ∈ Vect(X) and F ∈ Vect(Y ) , N, M ∈ Z+ ; we then have
x  y = [E  F ] − [N  F ] − [E  M ] + [N  M ] , and so

i (x  y) = [E ⊗ F+ ] − [E ⊗ M ] − [N ⊗ F+ ] + [N ⊗ M ] = 0,
since the fiber F+ ∼
= CM . (Beware: M denotes the trivial M -dimensional bundle
+
over Y in the first formula, but in the second, it denotes only the vector space
CM .) 
An important example of a locally compact space is furnished by Euclid-
ean space Rn , whose 1-point compactification is the n-sphere S n ; by definition
K(S n ) ∼
= K(Rn ) ⊕ Z holds, where the second summand, K({+}) = Z, is given
by the dimension of the vector bundle. This shows that K(Rn ) is actually the
interesting part of K(S n ).
Exercise 10.18. For an arbitrary paracompact X, go through the splitting
K(S n × X) ∼
= K(Rn × X) ⊕ K(X).
[Hint: Work with the sequence (10.3) from Step 2) of the preceding proof, where
+
B := X + is a retract of A := S n × X + /S n . Note that A = (S n × X) and
+
A/B = (Rn × X) , homeomorphically, see Figure 10.5.]
The importance of K-theory with compact support stems for one thing from the
fact that applications frequently involve noncompact, but locally compact, spaces
such as Euclidean space or tangent spaces. Of course it is possible, without undue
difficulties, to avoid noncompact spaces altogether (as with the passage from the
tangent bundle T X to the double ball bundle B + X ∪ B − X in Section 10.3 above)
— as artificial as this construction may appear. However, the splitting of Exercise
10.18 (also see the Remark 10.14b, p. 266) makes the locally compact formalism
genuinely simpler and perhaps conceptually clearer. In the following proof of the
Bott Periodicity Theorem which we adopt from [20], we will therefore stay in the
category of locally compact spaces: We already know that K(R0 ) = K({+}) ∼ =Z
and we obtain K(R1 ) = 0, since all complex vector bundles on the circle are trivial
(note that GL(N, C) is connected, see the gluing classification in the Appendix,
10.5. PROOF OF THE PERIODICITY THEOREM OF R. BOTT 269

S n £f+g X+

Figure 10.5. The 1-point compactification X + of a paracompact


X as a retract of (S n × X + )/S n ' (S n × X)+

Theorem B.10, p. 718). With some pains (as well as some projective geometry, see
Appendix, Exercise B.2a, p. 714, and [20, p.46f]), we could still compute K(R2 ) ∼=
Z. How does it go further?
The theorem which we will prove says that the sequence of these K-groups
continues. Thus, K(R3 ) = 0, K(R4 ) ∼
= Z, K(R5 ) ∼= 0, etc., and, more generally, for

each locally compact X, there is a natural isomorphism K(R2 × X) −→ K(X).

5. Proof of the Periodicity Theorem of R. Bott


In Chapter 4, we became acquainted with Wiener-Hopf operators, and as a
generalization of the discrete (and completely elementary) index theorem of Fritz
Noether, Israil Gohberg and Mark Krein (Theorem 4.4, p. 125), we gave
a construction (V, f ) 7→ F 7→ index F . In the pair (V, f ), V is a vector bundle
over a compact parameter space X, and f is an automorphism of π ∗ V , where
π : S 1 × X → X denotes the projection. For z ∈ S 1 and x ∈ X, f (z, x) is then
an automorphism of the fiber Vx , and it depends continuously on x and z. Now,
F : X → F denotes the associated family of Fredholm operators (after the Wiener-
Hopf recipe, formed on certain Hilbert spaces of half-space functions) and index
F ∈ K(X) denotes the index bundle of F , which is defined (in the special case
that the kernels of all of the Fredholm operators Fx have constant dimension) as
[Ker F ] − [Coker F ]. Moreover, we have seen that index F only depends on the
homotopy class of f .

Exercise 10.19. a) From the pair (V, f ), how can one construct a vector bundle
over S 2 × X that only depends on V and the homotopy class of f ?
b) Show that every E ∈ Vect(S 2 × X) can be obtained in this way.
[Hint for a): Decompose S 2 into the two hemispheres B + and B − with B + ∩
B − = S 1 , and form the bundle (π + )∗ V ∪f (π − )∗ V by means of the clutching
construction (see Appendix, Exercise B.9, p. 718), where π ± : B ± ×X → X denotes
the projection.
For b): Argue as in the proof of Theorem B.10 (p. 718) of the Appendix, where X
consists only of a single point. The parameter space plays only a subordinate role,
and so the proof actually carries over. It is convenient to normalize the map f that
one obtains so that f (1, x) is the identity on Vx .]
270 10. INTRODUCTION TO TOPOLOGICAL K-THEORY

Theorem 10.20. Let X be locally compact. Then a homomorphism α : K(R2 ×


X) → K(X) can be defined with the following properties:
i. α is functorial in X.
ii. If Y is another locally compact space, then we have the following multi-
plication rule, expressed by the commutative diagram:
t0 / K(R2 × X × Y )
K(R2 × X) ⊗ K(Y )
αX ⊗Id αX×Y
 
K(X) ⊗ K(Y )
t / K(X × Y ),

where t and t0 denote the outer tensor products  defined in Theorem


10.16.
iii. We have α(b) = 1, where b denotes the Bott class, that we define by
b := [E−1 ] − [E0 ] ∈ K(S 2 ).
Remark 10.21. Before we prove the Theorem, we have a few explanations:
a) We emphasize that in the following proof, the homomorphism α is defined by
the index of a Wiener-Hopf family of Fredholm operators; see also Exercise 10.24).
b) Recall that Em , m ∈ Z, denotes the line bundle defined over S 2 by means of
the clutching function, f (z) = z m ; see Appendix, Theorem B.10 (p. 718) or the
preceding Exercise 10.19a. Since E−1 and E0 have the same dimension (one), b
lies in K(S 2 , {+}) = K(R2 ). Thus, here we set X = {point}, and identify K(point)
with Z as in Exercise 10.9.
c) Note that b forms a basis in K(R2 ).
Proof. Step 1) We start with X compact. By Theorem 4.4 (p. 125) and
Exercise 4.12 (p. 129), in conjunction with the preceding Exercise 10.19, we can
(for each vector bundle E over S 2 × X) go through a construction E 7→ (V, f ) 7→
F 7→ Index F . In this way, a semi-group homomorphism Vect(S 2 × X) → K(X)
is defined, which can be extended (see Remark 10.7, p. 263) uniquely to a group
homomorphism α0 : K(S 2 ×X) → K(X). The restriction of α0 to K(R2 ×V ) (which
may be regarded as a subgroup of K(S 2 × X) by Exercise 10.18) then provides a
homomorphism α : K(R2 × X) → K(X). The functoriality of α means that, for
each element u ∈ K(R2 × X) and each continuous map g : X 0 → X (where X 0 is
another compact space), we have
∗ 
αX 0 (g × IdR2 ) u = g ∗ αX (u) ,
which is clear from the functorial nature of index bundles (Exercise 3.37, p. 87).
Step 2) We apply this definition for X = {point} and calculate α(b) = α0 [E−1 ] −
α0 [E0 ]. By construction, we have α0 [Em ] = −m (m ∈ Z), since by Theorem 4.4
(p. 125) we have the assignments
clutch Wiener-Hopf index
Em 7−→ (C, z m ) 7−→ Tzm 7−→ −m.
Hence α(b) = 1, and so (iii) is fulfilled.
Step 3) To prove the multiplication rule (ii) – first for X, Y compact – we must
consider the difference t(αX ⊗ Id)(u ⊗ v) − αX×Y (t(u ⊗ v)) for u ∈ K(R2 × X)
and v ∈ K(Y ); without loss of generality, we may assume v = 1 (the class of the
trivial line bundle CY over Y ), since all of the maps arising here are K(Y )-module
homomorphisms. By the functoriality of α, the difference π ∗ αX (u) − αX×Y (π ∗ u)
10.5. PROOF OF THE PERIODICITY THEOREM OF R. BOTT 271

vanishes, as needed. Here π : X × Y → X denotes the projection.


Step 4) The definition of α : K(R2 ×X) → K(X) for locally compact X carries over
just as in the proof of (ii) in the compact case via the one-point compactification
with the decomposition above in the proof of Theorem 10.16 and in Exercise 10.18.

Theorem 10.22 (Bott Periodicity Theorem). For each locally compact space
X, the homomorphism α : K(R2 × X) → K(X) is an isomorphism, whose inverse
β : K(X) → K(R2 × X) is given via outer multiplication x 7→ β(x) := b  x by the
Bott class b.
Proof. The proof follows by repeated use of the multiplicative property of α
expressed in Theorem 10.20(ii). In the following, for short we write only xy for the
outer product x  y or t(x ⊗ y).
!
αβ = Id: For the proof here, in (ii) substitute {point} for X and X for Y ; then
we have αβ(x) = (α(b)) (x) for each x ∈ K(X) by (ii), whence αβ = Id, since
α(b) = 1 by (iii).
! !
βα = Id: Let u ∈ K(R2 × X). We want to show that βαu := b(α(u)) = u, or
equivalently (if multiplying by b from the right) (a(u))b = ũ, where ũ := ρ∗ u
and ρ : X × R2 → R2 × X switches factors. From (ii) with R2 for Y , it follows
that (α(u))(b) equals α(ub). Now comes the trick, through which the proof of
βα = Id can be reduced to the rather banal fact αβ = Id already proven: On
K(R2 × X × R2 ), where the element ub lies, the map τ ∗ , which is the lifting along
the switching map
τ : R2 × X × R2 → R2 × X × R2 , where τ (a, b, c) := (c, b, a) ,
is the identity, since τ is homotopic to the identity on R2 × X × R2 through home-
omorphisms. (See the hint for Exercise 10.15b.) Namely, on R4 = R2 × R2 , τ is
given by the matrix  
0 0 1 0
 0 0 0 1 
 1 0 0 0  ,
 

0 1 0 0
which has determinant +1 and hence (one thinks of the transition to Jordan normal
form) lies in the same connected component of GL(4, R) as the identity. Hence, we
have
(α(u))b = α(ub) = α(τ ∗ (ub)) = α(bũ) = αβ ũ = ũ. 
Exercise 10.23. Show the following consequences of Theorem 10.22:

a. K(X × S 2 )−→ K(X) ⊗ K(S 2 ) for X compact.
Z , for n even,
b. K(Rn ) =
 0 , for n odd.
Z ⊕ Z , for n even,
c. K(S n ) =
Z, for n odd. 
Z , for n even,
d. N ≥ n/2 =⇒ πn−1 (GL(N, C)) =
0 , for n odd.
[Hint: While a), b) and c) follow directly from Theorem 10.22 with Exercise
10.18, the derivation of d) requires two further considerations. First, for a com-
pact manifold X of dimension n − 1, we have that for each E ∈ VectN (X) with
272 10. INTRODUCTION TO TOPOLOGICAL K-THEORY

−m
N ≥ m := [n/2 − 1], there is an F ∈ Vectm (X) such that E = F ⊕ CN X , i.e.,
each vector bundle over X is stably equivalent (see Exercise 10.11c, p. 264) to a
vector bundle of fiber dimension m. This is the basis theorem for vector bun-
dles. The uniqueness theorem then says that (under the same assumptions)
stably-equivalent vector bundles of fiber dimension N ≥ m + 1 are isomorphic. For
the proofs of these two lemmas (e.g., see [228, Ch.8]) one needs some homotopy
theory. The rest is trivial, since we can now represent the group I(X) of stable-
equivalence classes of bundles for N ≥ n/2 by Vect(X). In particular (see Exercise
∼ ∼
10.12c, p. 264), VectN (S n ) −→ I(S n ) −→ K(Rn ). The classical form of the peri-
odicity theorem now follows, since πn−1 (GL(N, C)) ∼= VectN (S n ) by Theorem B.10
(p. 718) of the Appendix.]
Exercise 10.24. As an alternative to Theorem 10.20, give a construction of
the isomorphism α : K(R2 × X) → K(X) by means of a family of elliptic boundary-
value problems.
[Hint: For X = {point} and f : S 1 → GL(N, C), consider (over the disk |z| < 1)
the transmission operator (see Exercise 5.18, p. 154)
 
∂u ∂v
Af : (u, v) 7→ , , f u|S 1 − v|S 1 ,
∂ z̄ ∂z
where u, v are N -tuples of complex-valued functions. By the same recipe, one can
also treat families of such boundary-value problems which are parametrized over a
space X; see [18, p.118-122].]
Remark 10.25. The connection with the construction of α via Wiener-Hopf
operators lies, roughly speaking, in the Poisson principle (i.e., in the Agranovich-
Dynin formula of [83, Chapter 21]), by which boundary-value problems can be
translated into problems on the boundary. Namely, extend (as in Exercise 8.24,
p. 226) the discrete Wiener-Hopf operator Tf to a pseudo-differential operator of
order 0 on the circle S 1 (via the identity on the basis elements of the form z m with
m < 0), which we still denote by Tf . Then, we have that
 
∂u ∂v
Sf : (u, v) 7→ , , Tf (zu|S 1 − v|S 1 )
∂ z̄ ∂z
is an elliptic problem with pseudo-differential boundary conditions; it has the same
kernel as the transmission problem Azf and isomorphic cokernel. On the other
hand, the operator Sf is formed directly by the composition of Az I with the prim-
itive boundary-value problem (u, v, w) → (u, v, Tf w). Here the letter I denotes
the constant function that assigns the unit matrix in GL(N, C) to each z ∈ S 1 .
While Exercise 5.18 says that index AI = 1, one finds index Az I = 0, whence
index Azf = index Sf = index Tf . Within a suitable algebra of boundary value
problems, then Sf induces the connection

   
∂ z̄ 0 0 I 0 0
∂ Sf
 0 ∂z 0 ↔ 0 I 0 
Mzf ◦ r − I ◦r 0 0 0 Tf
between a conventional elliptic system of partial differential equations of the first or-
der over the disk B 2 with ideally simple boundary-value conditions which are formed
via the restriction r : C ∞ (B 2 ) → C ∞ (S 1 ) and a trivial multiplication operator, and
a primitive boundary-value problem that consists of a (somewhat complex) elliptic
10.5. PROOF OF THE PERIODICITY THEOREM OF R. BOTT 273

g (y)

S n{1
f (y) x
f (x)

y
Bn

g (x)

Figure 10.6. Defining a retraction g : B n → S n−1 from a hypo-


thetical fixed-point free f : B n → B n

pseudo-differential operator only on the boundary. Actually, the change of con-


texts from differential boundary value problems in the plane to pseudo-differential
operators on the boundary line can be pushed further, and using a polynomial
approximation to f , a purely algebraic definition of the homomorphism α can be
given. In doing so, no Hilbert space theory and Fredholm operators are needed,
just as in the wholly elementary, but in parts quite tedious, proof [28]. A detailed
discussion of the advantages and disadvantages of the different methods for the
construction of α is in [18, p.131-136].
As a simple corollary of the periodicity theorem, we prove the classical fixed
point theorem of topology.
Theorem 10.26 (L.E.J. Brouwer, 1911). Each continuous map f of the
closed n-dimensional ball B n into itself has a fixed point.
Proof. If f (x) 6= x for all x ∈ B n , a continuous map g : B n → S n−1 is defined
by
g(x) := (1 − α(x))f (x) + α(x)x,
where α(x) ≥ 0 is chosen such that kg(x)k = 1; g is the identity on S n−1 , and
therefore a retraction of B n to S n−1 (i.e., g ◦ i = Id, where i : S n−1 ,→ B n denotes
the inclusion). See also Figure 10.6.
For odd n, we note that the composition
g∗ i∗
K(S n−1 ) −→ K(B n ) −→ K(S n−1 )
∼ Z⊕Z
is the identity, since g ◦i = Id. But it cannot be the identity, since K(S n−1 ) =
is not cyclic, whereas i∗ (K(B n )) is cyclic because K(B n ) ∼ = Z. For even n, note
that the corresponding composition of suspensions (see Appendix, Exercise B.9,
p. 718):
(Sg)∗ (Si)∗
K(S n ) −→ K(SB n ) −→ K(S n )
is the identity, and yet it cannot be the identity, since K(S n ) ∼
= Z ⊕ Z is not the
image of K(SB n ) ∼= Z. 
274 10. INTRODUCTION TO TOPOLOGICAL K-THEORY

Exercise 10.27. Define the mapping degree, deg f ∈ Z, of an arbitrary con-


tinuous map f : S n → S n , and show that homotopic maps have the same mapping
degree.
[Hint: Each homomorphism h : Z → Z clearly has a degree, namely d, such that
h(m) = dm for all m ∈ Z. Thus, for even n work with the subgroup K(Rn ) ∼ =Z
of K(S n ), which is carried into itself by f ∗ . For odd n, consider the suspension
Sf : S n+1 → S n+1 .]
Exercise 10.28. Show that the tangent bundle T (S n ) is nontrivial for even
n ≥ 2.
[Hint: More generally, each nowhere vanishing vector field v on S n (v ∈ C ∞ (T S n ))
provides a homotopy between the identity and the antipodal map ant : S n → S n ,
while deg Id 6= deg A for even n ≥ 2, contradicting Exercise 10.27.]
I The preceding examples show that important notions and results of classical alge-
braic topology can be developed as well, and perhaps more quickly and more easily, on the
basis of linear algebra via K-theory than they can by establishing, say, homology theory
by means of simplicial theory.
A number of sharper results, such as the well-known converse by Heinz Hopf of
Exercise 10.27 (equality of mapping degrees implies homotopy) in case n ≥ 2, cannot be
obtained by K-theoretic means but require deeper geometric considerations. As empha-
sized many times by Atiyah, the Periodicity Theorem is not only simpler than most major
theorems of classical algebraic topology, but also more relevant for the index problem and,
more generally, for many investigations of manifolds. Hereby, the philosophy is that the
usual algebraic topology destroys the structures too much while K-theory “comparable to
molecular biology” (Atiyah) searches for the essential macromolecules which make up the
manifold. It is then clear that for a comparison of the topology of the intricate manifold
with the building blocks one needs to first know the topology of the general linear group
which comprises the transition functions and, more generally, the Periodicity Theorem in
its K-theoretic form.
We shall apply the Periodicity Theorem in various ways for computing the index of
elliptic problems. Conversely, the index of special elliptic differential equations served to
prove the Periodicity Theorem; see especially Theorem 10.20 and Exercise 10.24. This
is no contradiction, no vicious circle, but an indication of how closely K-theory and
index theory of systems of linear elliptic equations are related both being linked by the
catchwords linear, finite dimensional and deformation invariant. J
CHAPTER 11

The Index Formula in the Euclidean Case

Synopsis. Index Formula and Bott Periodicity: Three Integer Invariants. The Dif-
ference Bundle of an Elliptic Operator: Operators Equal to the Identity at Infinity; Com-
plexes of Vector Bundles with Compact Support; Symbol Class in K-Theory with Compact
Support. The Index Theorem for Ellc (Rn ).

1. Index Formula and Bott Periodicity


Three Integer-Valued Invariants. In the preceding chapter, we used ana-
lytic tools (the Noether-Gohberg-Krein Index Theorem for Wiener-Hopf operators)
to prove the Bott Periodicity Theorem. We will now use it to derive an index the-
orem for elliptic integral operators in Rn . The basic idea is perhaps best described
via homotopy theory:
According to our Assumptions 8.10 (p.214) about the class Lkpc of (principally
classical) pseudo-differential operators, an elliptic pseudo-differential operator of
order k in Rn is given in the form
Z
(P u)(x) = (2π)−n eihx,ξi p(x, ξ)û(ξ) dξ,
Rn

where u is a C ∞ function on Rn with compact support and values in CN , and


the amplitude (see above Chapter 8) p is an N × N matrix-valued function with
principal symbol
p(x, λξ)
σ(P )(x, ξ) := lim ∈ GL(N, C), ξ 6= 0.
λ→∞ λk
Thus for fixed x, we have a continuous map
σ(P )(x, ·) : S n−1 −→ GL(N, C)
that has a well-defined degree for n even and N sufficiently large. This degree does
not depend on x, because of continuity (Rn is connected) and it was denoted by
deg(P ) in Remark 9.25, p. 248. In order to get interesting global problems, P is
usually combined with N k/2 boundary conditions to form an elliptic system in the
sense of [83, Chapter 18] or [190]. However, this is only possible in the stated
fashion if the local index deg(P ) vanishes. A sort of extremely simple boundary
condition — lacking the topological and analytical difficulty discussed in Section
9.4 — arises when we put k = 0 and p(x, ξ) = Id for x outside a compact subset K
of Rn .
Definition 11.1. The class of elliptic pseudo-differential operators of order 0
in Rn with this property of being equal to identity at infinity will be denoted by
Ellc (Rn ).
275
276 11. THE INDEX FORMULA IN THE EUCLIDEAN CASE

This class was first investigated by Robert T. Seeley. Obviously, see Exer-
cise 11.2 below, every P ∈ Ellc (Rn ) has a finite dimensional kernel and cokernel;
therefore index P is well-defined. If furthermore p(x, ξ) = Id for |x| ≥ r, r real,
then σ(P )(x, ξ) ∈ GL(N, C) for |x| + |ξ| ≥ r. In this way P defines a continuous
mapping S 2n−1 → GL(N, C), where S 2n−1 denotes the (2n − 1)-sphere of radius r
in R2n (the (x, ξ) -space). Since index P = index(P + Id) (Id the identity opera-
tor on functions), we may assume without loss of generality that N ≥ n. By the
homotopy theoretic form of the Bott Periodicity Theorem (Theorem 10.4, p. 257,
or Exercise 10.23d, p. 271), we have π2n−1 (GL(N, C)) ∼ = Z. Thus we have three
integer-valued invariants:

the analytic index, defined in the sense
index P
 of functional analysis,
the topological index, defined via homotopy
deg(σ(P )(·, ·))
 theory by the global behavior of σ(P ),
the local index, defined via homotopy for
deg P
even n by the pointwise behavior of σ(P )(x, ·).
We have deg(P ) = 0 (which is trivial) and index P = ± deg(σ(P )(·, ·)) (see Theorem
11.5, p. 281; be careful with the sign). The second formula is not trivial. Just as
with the Noether-(Gohberg-Krein) Index Formula for Wiener-Hopf operators on the
circle and the straight line [see Theorem 4.4 (p. 125), Exercise 4.10 (p. 128), and
Theorem 4.14 (p. 129)], its significance derives from the fact that on the left side
the analytic index, defined globally by the operator P , is an object of the analysis
of infinite-dimensional function spaces, while on the right side the topological index
is given by the symbol, i.e., by locally defined data of the linear algebra of finite
dimensional vector spaces (which are suitably integrated).
The proof of this index formula (see Theorem 11.5, p. 281) roughly rests on the
fact that Ellc (Rn ) is so rich that the analytic index (which as in Theorem 9.10d,
p. 240, only depends on the symbol and does not change under small deformations
of the symbol) can be considered an additive function on π2n−1 (GL(N, C)), and
thus as a multiple of the topological index. Comparing the topological and the
analytic index for the generators of the homotopy groups, we obtain equality.

2. The Difference Bundle of an Elliptic Operator


We will no longer pursue these homotopy-theoretic arguments, but rather we
carry out the details of the proofs in the more convenient formalism of K-theory,
which also permits an easier transition to the more general situation of the following
chapter.
Operators Equal to the Identity at Infinity. We begin with a definition,
packed in an exercise.
Exercise 11.2. Let X be a (not necessarily compact) oriented C ∞ Riemannian
manifold, and let Ellc (X) denote the class of elliptic pseudo-differential operators
of order 0 on X which are equal to the identity at infinity; i.e., for each
P ∈ Ellc (X), there is a compact subset K ⊆ X, such that P ϕ = ϕ for all C ∞
sections ϕ (in the domain of definition of P ) with support ϕ ∩ K = ∅, see Figure
11.1. The same condition should hold for the formal adjoint operator P ∗ .
a) Show that this definition coincides with the definition given in Section 11.1 for
X = Rn .
11.2. THE DIFFERENCE BUNDLE OF AN ELLIPTIC OPERATOR 277

K
supp ' X

Figure 11.1. An operator P is equal to the identity at infinity, if


P ϕ = ϕ for all sections ϕ with support near infinity, i.e., outside a
suitable compact subset

.
K = supp E

L{L0

Figure 11.2. Compact neighborhood L of K with K ⊂ L̊

b) Show that index P is well defined, depends only on σ(P ), and remains constant
under a C ∞ homotopy of the symbol within the space of elliptic symbols which are
the identity at infinity.
[Hint for b): Instead of repeating the proofs of Chapter 9, one can also reduce the
present case to the results of Chapter 9 directly. Indeed, one can embed K in a
bounded, compact, codimension-zero submanifold Y of X and then investigate the
doubled operator P̃ on the closed manifold X̃ := Y ∪∂Y Y . Show that index P̃ =
2 index P .]
Complexes of Vector Bundles with Compact Support. Much more
generally one can define, for locally compact Y , the group K(Y ) through com-
plexes of vector bundles with compact support. These are short sequences
α
0 → E 0 → E 1 → 0, where E 0 and E 1 are complex vector bundles over Y and
α is a vector bundle isomorphism outside a compact subset of Y . Two complexes
α β
E • = 0 → E 0 → E 1 → 0 and F • = 0 → F 0 → F 1 → 0 are called equivalent, if
γ
there is a complex G• = 0 → G0 → G1 → 0 with compact support over Y × I such
that E • = G• |Y ×{0} and F • = G• |Y ×{1} . The equivalence classes form a semigroup
C(Y ) with sub-semigroup C∅ (Y ) of elementary complexes with empty support
(i.e., the bundle maps over Y are isomorphisms). Then the sequence
d i∗
(11.1) 0 −→ C(Y )/C∅ (Y ) −→ K(Y + ) −→ K(+) −→ 0
is exact and splits; hence, C(Y )/C∅ (Y ) proves to be isomorphic to K(Y ).
278 11. THE INDEX FORMULA IN THE EUCLIDEAN CASE

The construction (which goes back to Michael Atiyah and Friedrich Hirze-
α
bruch) of difference bundles d(E • ) for complexes E • = 0 → E 0 → E 1 → 0 over
Y with compact support K goes roughly as follows: We choose a compact neigh-
borhood L of K, such that K is contained in the interior L̊ of L (see Figure 11.2).
In order to extend the complex E • to all of Y + , we replace it by an equivalent
complex whose bundles are trivial over L \ L̊:
α⊕Id
0 −→ E 0 |L ⊕ F −→ E 1 |L ⊕ F −→ 0 ,
where F ∈ Vect(L) is chosen (by means of Appendix, Exercise B.12, p. 719) so that
E 1 |L ⊕ F is trivial. Since α is an isomorphism on L \ L̊, we must have that E 0 |L ⊕ F
is trivial at least on L \ L̊. Let
 
τi : E i |L\L̊ ⊕ F |L\L̊ −→ L \ L̊ × CN , i = 0, 1
be trivializations with τ1 arbitrary and τ0 := τ1 ◦ (α ⊕ Id). Then the clutched
bundles (see Appendix, Theorem B.10, p. 718)
Gi := E i |L ⊕ F ∪τi ((Y + \ L̊) × CN ) ∈ Vect(Y + ), i = 0, 1,


are well defined and we set d(E • ) := [G0 ] − [G1 ] ∈ K(Y + ). Since the fiber dimen-
sions of G0 and G1 coincide, d(E • ) lies in K(Y ). Incidentally, one calculates easily
that d(E • ) = d(E • ⊕ H • ), if H • is an elementary complex, and that d(E • ) does
not depend on the choice of F .
Because of the universal property (see Remark 10.7, p. 263) of the functor K,
it suffices to define the splitting homomorphism of (11.1)
(11.2) e : K(Y + ) −→ C(Y )/C∅ (Y )
additively on Vect(Y + ). For E ∈ Vect(Y + ), one sets e(E) := E • |Y , where E •
β
denotes the complex 0 → E → p∗ i∗ E → 0, p : Y + → {+} denotes the retraction,
and β is an arbitrary extension of β+ := Id to an isomorphism on a neighborhood
of the point +. The support of E • is then compact and contained in Y , and hence
e(E) ∈ C(Y ). As an element of C(X)/C∅ (X), e(E) is independent of the choice of
the extension β. One sees immediately that eb ⊕ p∗ i∗ = Id, whence in particular
K(+) is the cokernel of d; and with suitable homotopies for the vector bundle
homomorphisms, we have ed = Id, whence d is injective.
Further details of this construction are found in [44, p.489ff] and in [384,
p.139-151], where complexes
α α αn
0 −→ E 0 −→
1
E 1 −→
2
· · · −→ E n −→ 0
of length n with ai+1 ◦ ai = 0 are considered, which are exact outside a compact
subset of Y . Incidentally, by means of tensor products of complexes of arbitrary
length (e.g., see [140, p.140 ff]) a ring structure on K(Y ) may be introduced in
a natural way.
α β
Remark 11.3. For complexes E • = 0 → E 0 → E 1 → 0 and F • = 0 → F 0 →
1
F → 0 of length 1, one obtains as an outer product the complex
φ  ψ
E •  F • := 0 −→ E 0  F 0 −→ E 1  F 0 ⊕ E 0  F 1 −→ E 1  F 1 −→ 0


of length 2, where (golden rule of multilinear algebra)


φ := α  Id + Id β and ψ := − Id β + α  Id .
11.2. THE DIFFERENCE BUNDLE OF AN ELLIPTIC OPERATOR 279

With the help of Hermitian metrics on the vector bundles one can rewrite E •  F •
as a complex of length 1
 θ
0 −→ E 0  F 0 ⊕ E 1  F 1 −→ E 1  F 0 ⊕ E 0  F 1 −→ 0, where
  

α  Id − Id β ∗
 
θ :=
Id β α∗  Id
and α∗ and β ∗ denote the adjoint homomorphisms. Details are in [17, p.93 f].
See also our previous use of outer products in Exercise 9.20c (p. 243) for elliptic
operators, in Equation (10.1) (Section 10.2, p. 261) for matrix-valued functions and
in Theorem 10.16 for K-groups.
If the locally compact space Y can be represented in the form Z \ A (where Z
is compact and A is closed in Z), then often in the literature for this special case
one sees
(11.3) K(Y ) = K(Z \ A) = K(Z, A) = CZ\A (Z)/C∅ (Z).
Hence, an element of K(Y ) is written as an equivalence class of an isomorphism
σ : E 0 |A → E 1 |A , where E 0 and E 1 are complex vector bundles over Z. In our
applications (where Y denotes the tangent bundle T X) one prefers K(BX, SX) to
K(T X), since BX and SX are compact for X compact. In fact, the construction
of the difference bundles in K-theory began with compact base space, although the
proof of (11.1) (see e.g., [17, p.88-94]) in its basic idea is not as simple as that for
the more general locally compact space.
The Symbol Class in K-Theory with Compact Support. We have packed
the link between index theory and K-theory in the following exercise.
Exercise 11.4. Show that the principal symbol σ(P ) of an elliptic operator
P ∈ Ellc (X) defines an element [σ(P )] ∈ K(T X) in a natural way, and that each
a ∈ K(T X) can be represented in this way; here, X is as in Exercise 11.2 and T X
denotes the tangent bundle of X, which can be identified with the cotangent bundle
T ∗ X by means of the Riemannian metric on X.
[Hint: Construct the difference bundle [σ(P )] ∈ K(T X) as in the preceding. In
addition, show  N
CB0 ∪σ(P ) CN

[σ(P )] = B∞ − [N ]
in the special case X = Rn , where T X + = (R2n )+ = S 2n = B0 ∪ B∞ with
B0 ∩ B∞ = S 2n−1 , and
n o
2 2
σ(P )(·, ·) : S 2n−1 = (x, ξ) : |x| + |ξ| = r2 −→ GL(N, C),
where r is so large that P ϕ = ϕ for all N -tuples ϕ of complex-valued functions
such that supp(ϕ) ∩ {x : |x| ≤ r} = ∅.
For the reverse direction set V := T X and represent a ∈ K(V ) (as with the splitting
φ
homomorphism (11.2)) by a complex 0 → F 0 → F 1 → 0, where the bundles F i ,
i = 0, 1 are restrictions to V of bundles of the same fiber dimension N over V + ; i.e.,

outside a compact subset L ⊆ V , we have isomorphisms τi : F i |(V \L) → (V \ L)×CN
−1
such that φ := (τ1 ) τ0 is a bundle isomorphism over V \ L.
If π : V → X denotes the base point map, then replace the bundles F i on V \ L
(where they are trivial) by π ∗ E i , where E i is the restriction of F i to the zero section
of V . On L this cannot be done in general. However, the following artifice (after
280 11. THE INDEX FORMULA IN THE EUCLIDEAN CASE

Y
v
¼(L)
x L X {
S½(V ) B½(V )jY
¼

Figure 11.3. Choice of Y ⊂ X and ρ > 0 such that L ⊂ Bρ (V )|Ȳ

homogeneous extension

homogeneous extension

trivial extension trivial extension

homogeneous extension

homogeneous extension

Figure 11.4. Deforming the representation 0 → F 0 → F 1 → 0


of an arbitrary a ∈ K(T X) to a symbol class

[44, p.492f]) is helpful: Choose an open, relatively compact subset Y of X that


includes π(L) and a real number ρ > 0, such that the compact set L is contained
in the ball bundle Bρ (V )|Ȳ (see Figure 11.3). Now show that Ȳ is a deformation
retract of Bρ (V )|Ȳ , and conclude (with Appendix, Theorem B.7, p. 716) that there
are isomorphisms, θi : F i |(Bρ (V )|Ȳ ) → π ∗ E i |(Bρ (V )|Ȳ ), which are extensions of
the trivialization isomorphisms given above over Ȳ \ Y . Thus one must show that
θi can be chosen so that the homomorphisms
θi (v) : Fvi −→ π ∗ E i v = Fxi , for π(v) = x ∈ Ȳ \ Y


−1
coincide with the composition (τi (x)) τ0 (v). Furthermore, if we require that θi
is the identity on the zero section then it is uniquely determined  up to homotopy.
Now define α := θ1 ◦ φ ◦ θ0−1 over ∂(Bρ (V )|Ȳ ) = (Sρ (V )|Ȳ ) ∪ Bρ (V )|(Ȳ \Y ) , and
on V |Ȳ (modulo the zero section) extend it to be homogeneous of degree 0, and
consider the given trivialization on π −1 (Y \ Ȳ ) (see Figure 11.4).
α
Finally, consider the complex 0 → π ∗ E 0 → π ∗ E 1 → 0 so obtained, where α is
homogeneous of degree 0, and is induced by an isomorphism E 0 → E 1 outside a
compact subset of the base X. Approximate it by a C ∞ mapping with the same
properties, where the E i can be taken to be C ∞ vector bundles (without loss of
generality — see Appendix, Exercise B.13b, 719). Incidentally, what simplifications
can be made for X = RN ?]
11.3. THE INDEX THEOREM FOR Ellc (Rn ) 281

3. The Index Theorem for Ellc (Rn )


Theorem 11.5. For all P ∈ Ellc (Rn ), we have the index formula
index P = (−1)n αn ([σ(P )]).

Here an : K(R2n ) → K(R0 ) = Z denotes the periodicity homomorphism produced
by iteration of
αX : K(R2 × X) −→ K(X) for X = R2(n−1) , R2(n−2) , . . .
(see Theorem 10.22, p. 271).
Proof. Step 1) For locally compact X, we have the following commutative
diagram:
[σ(·)]
Ellc (X) / K(T X)

index
% 
index
Z
Here the analytical index is defined on Ellc (X) by Exercise 11.2b. By Exercise
11.4 (surjectivity of the difference bundle construction [σ(·)]) it is well defined on
K(T X) and trivially additive (Exercise 1.5, p. 5). Hence, for X = Rn (where
K(T X) is isomorphic to Z by the Bott Periodicity Theorem above), the index is a
multiple of this isomorphism, whence
index P = Cn αn ([σ(P )]),
where the constant does not depend on P , but indeed may depend on n.
Step 2) We now want to show that Cn = (−1)n . For this we must find a P ∈
Ellc (R) with [σ(P )] = b  · · ·  b ∈ K(R2n ) and index P = (−1)n . Recall that
b ∈ K(R2 ) denotes the Bott class of Theorem 10.20, p. 270. The main problem
consists in finding a sufficiently simple operator P , so that one can compute its
analytical index. We already know that P cannot have constant coefficients, since P
is the identity at infinity; also, P is not a differential operator, since it has vanishing
order. There are various ways to solve this problem; see [16, p.243f], [21, p.110 ff],
and [222, p.141-146]. For us, it is most convenient to first show that we can restrict
ourselves to the case n = 1. As in Exercise 9.20c (p. 243), we have (with analogous
proof) the following multiplicative properties: P = Q#R, P ∈ Ellc (Rn ), Q ∈
Ellc (Rm ), R ∈ Ellc (Rk ), and m + k = n imply σ(P ) = σ(Q)#σ(R) and [σ(P )] =
[σ(Q)]  [σ(R)] and index P = (index Q)(index R). Since αn is multiplicative by
construction, we have Cn = (C1 )n .
Step 3) Thus, let n = 1 (i.e., X = R and T X = R2 = C = {x + iξ}. By definition,
the Bott class b is represented by the complex
(x+iξ)
0 / CC / CC / 0,

 
TX TX
where CC denotes the trivial vector bundle C × C → C of complex fiber dimension
1 over the basis C, with the notation of Definition 10.8. As it stands, the complex
still does not represent any pseudo-differential operator. As in Exercise 11.4, we
282 11. THE INDEX FORMULA IN THE EUCLIDEAN CASE

» {1
(x+i»)
1 1
1 e¼i(x{1)
1 1 1 1

{1 0 1 x x x x x
1 1 1 1 1 1
1 1

' ®
Deformation: dilation of one

Figure 11.5. Deforming the Bott clutching morphism ·(x + ξ)−1


to an elliptic symbol

can deform the bundle map φ (see Figure 11.5), which at the point (x, ξ) is defined
on the fiber C by
φ(x, ξ) : z 7→ z(x + iξ)−1 ,
to a map α with
α(x, ξ) = 1, for |x| ≥ 1,
α(x, λξ) = α(x, ξ), for λ > 0,
α(x, ξ) 6= 0, for ξ 6= 0.
For |x| ≤ 1, we explicitly set

eiπ(x−1) , for ξ > 0,
α(x, ξ) =
1, for ξ < 0.
Thus, after smoothing, we can represent α as the symbol of an elliptic pseudo-
differential operator T of order 0 on R, which is the identity outside of the interval
iπ(x−1)
[−1, 1] and inside it equals theP Toeplitz operator
P∞T̃ :=ν e P + (Id −P ) on the
1 ∼ ∞ ν
circle S = [−1, 1], where P : −∞ aν z →
7 0 aν z denotes the projection op-
erator. By construction, index T = index T̃ , and according to Exercise 8.24c, we
obtain index T̃ = W (eiπ(x−1) , 0) = −1. Thus, index(b) = −1, and C1 = −1 then
follows. 
Exercise 11.6. How can one directly prove C2 = 1, without using induction
from step 2 of the preceding proof?
[Hint: Consider, on the disk X := B 2 , the transmission operator
 
∂u ∂v
(11.4) A : (u, v) 7→ , , (u − v)|S 1
∂ z̄ ∂z
with index A = 1 (Exercise 5.18, p. 154), and construct an operator A0 (in a suitable
algebra Ell(X, ∂X) of elliptic boundary value problems, see, e.g., [83] or [190])
which is stably equivalent to A, and is equal to the identity in a neighborhood
of ∂X; use the deformation procedure of Exercise 9.26a (p. 249) or Theorem 9.24
(p. 245). Show index A0 = index A, whence index A00 = 1 if A00 ∈ Ellc (R2 ) denotes
the extension of A0 to all of R2 with A00 = Id outside B 2 .
Then it only remains to show that [σ(A00 )] is actually b  b. For this, represent b
as in Theorem 11.5 by the complex
ζ −1
0 −→ CC −→ CC −→ 0
and derive (using the recipe given in Remark 11.3, p. 278) the representation of
θ
b  b by the complex 0 → CT Z ⊕ CT Z → CT Z ⊕ CT Z → 0 over the tangent bundle
11.3. THE INDEX THEOREM FOR Ellc (Rn ) 283

T Z = C2 = {(z, ζ)} of the space Z = R2 = C = {z}. Then


" −1 #
z −1 − ζ̄
θ(z, ζ) := −1
ζ −1 (z̄)
can be deformed into σ(A0 )(z, ζ).]
I Theorem 11.5 is a beautiful result of a purposeful application of modern topolog-
ical methods to questions of analysis. Its theoretical ramifications are manifold, and we
mention briefly:
(i) Generalizations of the Noether-(Gohberg-Krein) Index Formula for Wiener-Hopf
operators on the circle or half-line (see above Chapter 4) to elliptic pseudo-differential
operators of order 0 on n-dimensional Euclidean space. See also [345] for a systematic
comparison of these two interesting operator classes.
(ii) Analytic definition of the degree of a (2n − 1)-dimensional homotopy class of
GL(N, C): In its homotopy theoretic form (see Section 11.1), Theorem 11.5 supplies an
explicit formula for the index, if one uses either one of the definitions of degree given
explicitly in Section 10.2 above. Conversely, one can use Theorem 11.5 to define the
degree (by analytic means) and generalize in this fashion the algebraic definition, which
in Section 10.1 was available only for n = 1, to the case n > 1. In [16, p.244], Atiyah
comments “the index of an operator is usually a less computable quantity than an integral,
say, the actual computation is for many theoretical goals not important, while the analytic
definition entails numerous theoretical advantages”. In that paper, a number of advantages
— a priori integrality, connections with analytic function theory and Lie groups — are
discussed in detail, as in our Chapters 13 and 15–18 below.
(iii) Applications to boundary-value problems. In Section 9.4, we learned methods
for trivializing the symbol of an elliptic boundary-value problem along the boundary.
Thereby (more precise discussion in [87, p.40]) one usually obtains an operator of order
0 which equals the identity near the boundary and therefore can be continued to an
operator of Ellc (Rn ), if the manifold with boundary is a region with boundary in Rn . As
sketched in the hint to Exercise 11.6, the index does not change under these manipulations.
Therefore, the index of a boundary-value problem can be computed as the case may be
via Theorem 11.5; however, the derivation of a closed formula (see Chapter 13 below)
requires a reduction to the Agranovich-Dynin Formula ([83, Chapter 21]) and thus to the
study of pseudo-differential operators also on the boundary which is a closed manifold.
(iv) In the following chapter we will derive from Theorem 11.5 an index formula for
elliptic operators on closed (= compact without boundary) manifolds, first provided the
latter can be embedded in Euclidean space trivially (i.e., with trivial normal bundle), then
for arbitrary embeddings. J
CHAPTER 12

The Index Theorem for Closed Manifolds

Synopsis. Proof of the Index Theorem by Embedding: Pilot Study — The Index
Theorem for Embeddings with Trivial Normal Bundle. Proof of the Index Theorem for
Non-Trivial Normal Bundle: The Difference Element Construction, Revisited; Symbol
Class; Thom Isomorphism of K-Theory; Definition of the Topological Index; Definition
of the Analytic Index; Foundations of Equivariant K-Theory. Multiplicative Property:
Formulation; How it Fits into the Embedding Proof; Proof of the Multiplicative Prop-
erty. Short Comparison of the Cobordism, the Embedding and the Heat Equation Proof;
Outlook to Spectral Theory, Asymmetry, and Inverse Problems.

I In this chapter, we shall show that the Atiyah-Singer Index Theorem can be de-
rived from Bott Periodicity by an embedding argument. In such a way, the proof of the
Index Theorem can be executed within K-theory. Below in Section 12.3 (p.301ff) we shall
compare the different basic ideas underlying the most prominent three proofs of the Index
Theorem, cobordism, embedding, and heat equation asymptotics. Roughly speaking, the
cobordism idea speaks immediately to everyone in analysis or topology who is interested in
manifolds with boundary as building blocks of the mathematical universe. For a compre-
hensive presentation we refer to [328], and for a wider context to [83]. The heat equation
idea appeals to differential geometers and yields fascinating interpretations of almost all
intermediate constructions in terms of geometry and particle physics. There are several
excellent textbooks prevalent, like [167] and [55]. For the special (but representative) case
of twisted Dirac operators, we shall explain that approach with all details in our Chapter
17 (pp.513–642).
To catch the idea of the K-theoretic embedding proof of the Index Theorem, it suffices
to read our pilot study for embeddings with trivial normal bundle in the following short
Section 12.1. To get the full proof, however, we must deal with some more delicate
concepts and somewhat lengthy proofs of technical details. That will be presented in the
long Section 12.2. Fairly complete and much shorter presentations can be found in the
original paper [44] and the marvelous classic textbooks [389], [224, Section 19.3], and
[273, Section III.13].
Then, why should the reader work through our long pedestrian presentation when
short ingenious - and also correct - presentations are available for the embedding-based
proof? We answer that question by recalling a story about Mark Kac giving a long
complicated calculation on the blackboard at Cornell with Richard Feynman in the
audience. To the extent that Kac’s calculation went more and more difficult, everybody in
the audience could notice Feynman’s pain listening to it until he raised himself, stepped to
the blackboard, took the chalk of Kac, cancelled one term against another term, cancelled
a third term against a fourth, turned to Kac with the words: “There it is, right, Mark?
Why are you wasting our time?” and returned to his seat. Whereupon Kac in his most
polite Polish English answered: “You see, Dick, I explain how to do the calculation when
you are not in the audience”. J

284
12.1. PILOT STUDY: THE INDEX FORMULA FOR TRIVIAL EMBEDDINGS 285

1. Pilot Study: The Index Formula for Trivial Embeddings


Let X be a closed (i.e., compact, without boundary), oriented, Riemannian man-
ifold of dimension n, which is trivially embedded (i.e., with trivial normal bundle) in
the Euclidean space Rn+m . Let E and F be Hermitian vector bundles over X and
P ∈ Ellk (E, F ), k ∈ Z. Then we have the following formula:

Theorem 12.1 (M. F. Atiyah, I. M. Singer 1963).

index P = (−1)n αn+m ext [σ(P )]  bm ,




where [σ(P )] ∈ K(T X) ∼ = K(T ∗ X) denotes the (principal) symbol class of P , b ∈ K(R2 )
the Bott class, ext : K(T N ) → K(R2(n+m) the natural extension explained in (12.2) below,
and αn+m : K(R2(n+m) ) → Z the iteration of the Bott isomorphism.

Remark 12.2. The preceding situation of trivial embedding arises in applications,


e.g., when X is a hypersurface, in particular the smooth boundary of a bounded domain in
Rn+1 . For the more general case of nontrivial embedding, for different modes of expressing
the topological index (the right side of the formula), and for a comparison of the various
proofs, see the sections below.

Proof. Step 1. First, we want to visualize the contents of the formula. The left
side is a well-defined (by Chapter 9) integer which depends only on the homotopy type of
the principal symbol σ [P ] ∈ IsoSX (E, F ), where SX denotes the cotangent sphere bundle.
However, how is the right side defined? The construction of [σ(P )] ∈ K(T X) was carried
out in Exercise 11.4 (p.279) for k = 0; the case k 6= 0 adds nothing new. (One can reduce
it to the case k = 0 directly via composition with Λ−k .)
Now, consider the Figure 12.1. Here N is a tubular neighborhood of X in Rn+m ;

Rn+m

embedding

X
X
N

Figure 12.1. An embedding of a closed n-dimensional manifold


X into Rn+m with trivial normal bundle and tubular neighborhood
diffeomorphic to X × Rm

i.e., a neighborhood of X which locally (and also globally, because of the triviality of the
embedding) has the form X × Rm .
The index formula then says that the analytic index map indexa , defined by the
left triangle of the following diagram, coincides with the topological index map, defined
286 12. THE INDEX THEOREM FOR CLOSED MANIFOLDS

by the right square. Differently put, the following diagram is commutative:


(−b)m
(12.1) Ellk (E, F )
[σ(·)]
/ K(T X) / K(T X × R2m ) = K(T N )

indexa indext ext


index

) Z o (−1)n+m αn+m
K(T Rn+m ) = K(R2(n+m) ).
Here,
(12.2) ext : K(T N ) −→ K(R2(n+m) )
is induced by the map (T Rn+m )+ → (T N )+ , which maps the complement of the open set
T N in T Rn+m to the point at infinity + ∈ (T N )+ .
Note. We prefer working with the finitely generated abelian groups K(T X), rather
than with single elliptic operators or classes of such: It is precisely the advantage of
topological methods that complicated objects of analysis, whose structure is only partially
explored, can be replaced purposely by simple quantities (in the case
Pn before us,2kby the
rank r of the group K(T X)). Incidentally, it is known that r = k=0 rank H (T X).
Therefore, justified by Exercise 11.4, we will henceforth consider the analytic index not
on Ellk (E, F ) but directly on K(T X) as indicated by the double arrow in the diagram.
Step 2. We choose a very direct way for the proof of the index formula by system-
atically replacing the horizontal homomorphisms (K-theoretic operations) by operations
associated with corresponding elliptic operators: Thus, for each elliptic operator P over X,
we construct an elliptic operator P 0 over N , which is the identity at infinity and satisfies
index P 0 = (−1)m index P and [σ(P 0 )] = [σ(P )]  bm .
We take P 0 to be P #T m , where T ∈ Ellc (R) denotes the standard operator with
index T = −1 given in the proof of Theorem 11.5, p. 281; further note that in Exercise
9.20 (p.243) the tensor product # was only defined for operators of order k > 0. Hence,
more precisely, we take P 0 ∈ Ellc (N ) to be an operator whose principal symbol on the
m
unit cosphere bundle S(X × Rm ) ⊂ T ∗ (X × Rm ) coincides with σ(P )#σ(T )#· · ·#σ(T );
e.g., for m = 1
− IdFx ⊗σ(T ∗ )(t, τ )
 
σ(P )(x, ξ)
(σ(P )#σ(T )) (x, t; ξ, τ ) = ∗ ,
IdEx ⊗σ(T )(t, τ ) σ(P )(x, ξ)
where (x, t; ξ, τ ) ∈ (X × R) × S(X × R). Since σ(T )(t, τ ) = 1 for t sufficiently large and
for τ negative, we can proceed like in Exercise 11.4 (p.279). Whence this principal symbol
can be conveniently deformed to a symbol σ 0 with σ 0 (x, t; ξ, τ ) = Id for t sufficiently large,
if we identify E 0 ⊕ F 0 and F 0 ⊕ E 0 by means of switching the summands. Here E 0 := p∗ E
0
where p : N = X × E → X denotes the projection, whence Ex,t = Ex , and F 0 is defined
similarly.
Step 3. Without loss of generality, we can take E 0 ⊕ F 0 to be a trivial bundle,
since otherwise we can form P 0 ⊕ IdG , where the bundle G over N is chosen so that
E 0 ⊕ F 0 ⊕ G is trivial. Hence, E 0 ⊕ F 0 is extended to a trivial bundle over all of Rn+m .
Then extend the operator P 0 to all of Rn+m , by the identity outside N , to an operator
P 00 ∈ Ellc (Rn+m ). Since P 0 is the identity near the boundary N̄ \ N of N it follows that
if P 00 u = 0 (on Rn+m ), then the support of u lies entirely in the interior of N , whence
Ker P 00 = Ker P 0 . By the same argument for the formal adjoint operators, it follows that
index P 00 = index P 0 ; [σ(P 00 )] = ext[σ(P 0 )] by construction.
Step 4. The formula given for index P then follows from the index formula in the
Euclidean case (Theorem 11.5, p.281). 
Remark 12.3. Instead of tensoring with the standard operator T ∈ Ellc (R), we can
also (for m = 2) tensor with the standard transmission operator
T ∈ Ell(B 2 , S 1 ),
12.2. PROOF OF THE INDEX THEOREM FOR NONTRIVIAL NORMAL BUNDLE 287

of Exercise 11.6, (11.4), p.282. That yields the construction of an elliptic boundary-value
problem over the bounded manifold N̄ , whose inner symbol can be deformed by the well-
known procedure (using the boundary symbols) such that it becomes the identity near the
boundary N̄ − N . The desired operators P 0 and P 00 are then provided. Correspondingly,
for arbitrary m, one can find a boundary-value problem whose index is 1 and whose symbol
induces the bundle b  · · ·  b (m-times): For this, one takes the differential operator d + δ
(in the exterior calculus of differential forms; see Exercise 6.20, p. 172) from the forms of
even order to those with odd order, with a suitable elliptic boundary-value problem in the
sense of [83].
Remark 12.4. The advantage of the construction of P 0 via boundary-value problems
is best exhibited, when the embedding of X in Rn+m is not trivial i.e., when N is no longer
X × Rm . In the next Section, following [18] and [21], we are going to employ devices of
equivariant K-theory (with transformation groups) for explicitly stating or axiomatically
characterizing the desired operator P 0 with the help of symmetry properties of standard
operators over the sphere. The use of equivariant K-theory can be avoided (also in the
case N 6= X × Rm of the following section) by passing to boundary-value problems and by
another method proposed by [221] using hypo-elliptic operators and stronger analytical
tools.
Remark 12.5. Just as we learned (in Sections 10.1 and 10.2) different ways for defin-
ing the degree, the index, computed here in K-theoretic terms, can be determined in
cohomological or integral form; see Section 13.1. This is possible without new or modified
proofs, but simply by routine exercises in algebraic topology, the transition from K-theory
to cohomology, whereby simply “one set of topological invariants is translated into an-
other”, so Atiyah and Singer. Which formula provides the “best answer” is largely a
matter of taste. It depends on which invariants are most familiar or can be computed
most easily.

2. Proof of the Index Theorem for Nontrivial Normal Bundle

I The standard, geometric, elliptic differential operators P : C ∞ (E) → C ∞ (F ) (e.g.,


see our Table 6.1, p.187) typically depend on the choice of a Riemannian metric on X,
and Hermitian metrics and connections on E and F . Such choices might be called the
geometric data defining the operator. As these geometric data vary smoothly, it can
be shown that the operator P k+m : W k+m (E) → W k (F ) varies continuously in the space
F W k+m (E), W k (F ) of bounded Fredholm operators and hence the integer index(P k+m )


is constant. Thus, we expect index P to be a topological attribute of X, E and P , such


as a combination of the Euler characteristic and/or other characteristic classes, which is
invariant under a change of the geometric data. Since P is defined in terms of geometric
data, one is tempted to try to compute the integer index P in terms of these data, as
we did in the case of embedding with trivial normal bundle. Also for non-trivial embed-
ding, the Atiyah-Singer Index Theorem says that this is not only possible, but in fact
index P is equal to a topological (integer) invariant known as the topological index of P
determined (as described in the next subsections) by a suitable homotopy class of the
principal symbol σ(P ). We denote the topological index of P by indext [σ(P )], see also
our previous diagram (12.1). Succinctly, the Atiyah-Singer Index Theorem (or Formula)
is indexa P = indext [σ(P )], where indexa P := dim Ker P −dim Coker P denotes the usual
index of P as an elliptic operator (possibly pseudo-differential), with the subscript “a”
standing for “analytic”.
In the next two subsections, we shall define indext [σ(P )]. (Later we shall describe
several equivalent definitions of it. The utility and appreciation of each definition de-
pends on one’s background.) Then we shall define indexa A, the analytic index for any
A ∈ K(T ∗ X), by lifting the index from elliptic operators to K-theoretical classes, and
288 12. THE INDEX THEOREM FOR CLOSED MANIFOLDS

formulate and prove the multiplicative property for it. That property is the key for the
embedding proof of the Index Theorem for non-trivial normal bundles. Our approach
requires a rudimentary knowledge of equivariant K-theory, provided further below. The
constructions will be done in the Bokobza-Haggiag (global) symbolic calculus explained
in Section 8.5. J

The Difference Element Construction, Revisited. First we elaborate the so-


called difference element construction, announced in Exercise 11.4 (p.279). Let (X, Y ) be
a compact pair where Y ⊂ X. Given bundles π0 : E0 → X and π1 : E1 → X and an

isomorphism σ : E0 |Y −→ E1 |Y , we construct an element χ(E0 , E1 ; σ) ∈ K(X, Y ). Let
X0 := X × {0} and X1 := X × {1} denote two copies of X, and let Z := X0 ∪Y X1 =
X0 ∪ X1 /[(y, 0) ∼ (y, 1)]; i.e., the disjoint union of X0 and X1 , but with (y, 0) and (y, 1)
identified for all y ∈ Y . We define a vector bundle F over Z by the clutching construction
(see Appendix, Exercise B.9, p.718) F := E0 ∪ E1 / [(e0 )y ∼ (σ(e0 ))y ] for all e0 ∈ E0,y ,
the fiber of E0 over the point y ∈ Y . The bundle F is the disjoint union of E0 over X0
and E1 over X1 but with the fibers over (y, 0) and (y, 1) identified via σy for all y ∈ Y .
We have a retraction
ρ : Z → X1 , given by ρ(x, ν) := (x, 1) for ν ∈ {0, 1}.
i j
From the sequence (X1 , ∅) −→ (Z, ∅) −→ (Z, X1 ), we obtain an exact sequence
j∗ i∗
0 −→ K(Z, X1 ) −→ K(Z) −→ K(X1 ) −→ 0,
and this sequence is split, with ρ∗ : K(X1 ) → K(Z) serving as a left inverse of i∗ . Let
F1 → Z denote the bundle ρ∗ (E1 ), namely the pull-back of E1 → X1 , via ρ. Note that
i∗ ([F ]−[F1 ]) = 0 since F |X1 = F1 |X1 . Thus, there is κ ∈ K(Z, X1 ) with j ∗ (κ) = [F ]−[F1 ].
Since Z/X1 ∼ = X/Y , K(Z, X1 ) ∼ = K(X, Y ), and hence κ corresponds to some element of
K(X, Y ), which by definition is χ(E0 , E1 ; σ) ∈ K(X, Y ). If Y = ∅, then we claim that
χ(E0 , E1 ; σ) = [E0 ] − [E1 ]. Indeed, for Y = ∅,
(F1 → Z) = ρ∗ (E1 → X1 ) = (E1 → X0 ) ∪ (E1 → X1 ), whereas
(F → Z) = (E0 → X0 ) ∪ (E1 → X1 ).
Consequently,
j ∗ (κ) = [F ] − [F1 ] = [(E0 → X0 ) ∪ (E1 → X1 )] − [(E1 → X0 ) ∪ (E1 → X1 )]
= [(E0 → X0 ) ∪ (0 → X1 )] − [(E1 → X0 ) ∪ (0 → X1 )]
= j ∗ ([E0 → X0 ] − [E1 → X0 ]),
and [E0 → X0 ] − [E1 → X0 ] ∈ K(Z, X1 ) = K(X0 ∪ X1 , X1 ) corresponds to [E0 ] − [E1 ] ∈
K(X) ∼= K(X0 ).
Definition 12.6 (Difference element). Given a compact pair (X, Y ), two complex
vector bundles π0 : E0 → X and π1 : E1 → X, and an isomorphism σ : E0 |Y → E1 |Y , in
the above notation the difference element χ(E0 , E1 ; σ) ∈ K(X, Y ) is defined by
χ(E0 , E1 ; σ) := (j ∗ )−1 ([F ] − [F1 ]) ∈ K(X0 ∪Y X1 , X1 ) ∼
= K(X, Y ),
where we have identified K(X0 ∪Y X1 , X1 ) with K(X, Y ).

Symbol Class, Thom Isomorphism, and the Topological Index of an Elliptic


Operator. Let X be a compact C ∞ n-manifold with cotangent bundle π : T ∗ X → X.
Relative to a Riemannian metric for X, we can consider a unit ball bundle
BX := {ξ ∈ T ∗ X : |ξ| ≤ 1},
with boundary the cotangent sphere bundle SX.
12.2. PROOF OF THE INDEX THEOREM FOR NONTRIVIAL NORMAL BUNDLE 289

The Symbol Class. If σ(P ) : T ∗ X → Hom(π ∗ E, π ∗ F ) is the principal symbol (or,


alternatively, the Bokobza-Haggiag amplitude in the sense of Section 8.5, pp.231ff) of an
elliptic operator P : C ∞ (E) → C ∞ (F ), then σ(P ) restricts to an isomorphism
σ(P )|SX : π ∗ E|SX → π ∗ F |SX .
We can then apply the above difference construction to obtain
χ(π ∗ E, π ∗ F ; σ(P )|SX ) ∈ K(BX, SX).
There is an isomorphism
K(T ∗ X) := K((T
e ∗
X)+ ) ∼
= K(BX/SX)
e =: K(BX, SX),
and so we may regard χ(π ∗ E, π ∗ F ; σ(P )|SX ) ∈ K(T ∗ X).
Definition 12.7. If P : C ∞ (E) → C ∞ (F ) is an elliptic operator with principal sym-
bol (or, as emphasized above, the Bokobza-Haggiag amplitude in the sense of Section 8.5,
pp.231ff) σ(P ) : T ∗ X → Hom(π ∗ E, π ∗ F ), then the symbol class of P is denoted and
defined by
[σ(P )] := χ(π ∗ E, π ∗ F ; σ(P )|SX ) ∈ K(T ∗ X).
Thom Isomorphism in K-Theory. After some preliminary work, we will eventually
produce an integer from [σ(P )] which will be the desired indext [σ(P )], the topological
index of P .
Let π : V → X be a complex vector bundle, where X is compact. Let Λi (V ) denote
the i-th exterior bundle of V over X. The pull-backs π ∗ Λi (V ) are then bundles over
V , say π i : π ∗ Λi (V ) → V . At each v ∈ V , we have a linear map αvi : π ∗ Λi (V ) v →
π ∗ Λi+1 (V ) v , given by exterior multiplication αvi (w) := v ∧ w. Since αvi+1 ◦ αvi = 0, we


have a complex over V , namely


α0 α1 αn−1
0 −→ π ∗ Λ0 (V ) −→ π ∗ Λ1 (V ) −→ · · · −→ π ∗ Λn (V ) −→ 0,
where n denotes the fiber dimension of V . If v 6= 0, we have Im αvi = Ker αvi+1 , so that
 

the complex is exact over V minus the zero section. Thus, the complex defines an element
λV ∈ K(V ) in the following way: Define bundles over V by
M ∗ k M ∗ k
π ∗ Λev (V ) := π Λ (V ) and π ∗ Λodd (V ) := π Λ (V ).
k even k odd

Relative to a Hermitian structure for V , let BV := {v ∈ V : |v| ≤ 1} and SV := {v ∈ V :


|v| = 1}. By summing the exterior multiplications and their adjoints (with opposite sign)
over the even forms we obtain an isomorphism over SV , namely

M i
αe − (αe )∗ |SV := α − (αi−1 )∗ |SV : π ∗ Λev (V )|SV −→ π ∗ Λodd (V )|SV .
 
i even

We mention that the adjoint of αvi ,


   
(αvi )∗ : π ∗ Λi+1 (V ) −→ π ∗ Λi (V )
v v

is given by interior multiplication (see (6.9), p.173 above in Section 6.4) by the Hermitian
dual v ∗ ∈ V ∗ . Applying the difference construction relative to the compact pair (BV, SV ),
we obtain the canonical difference element of the exterior algebra
λV := χ π ∗ Λev (V )|BV , π ∗ Λodd (V )|BV ; (αe − (αe )∗ )|SV ∈ K(BV, SV ) ∼

= K(V ).
The following result can be considered as a reformulation or generalization of Bott Pe-

riodicity (Theorem 10.22, p.271). It yields an isomorphism K(X) −→ K(X × R2 ) by
multiplication with the Bott class b that corresponds to the canonical exterior class λV
for trivial V = X × R2 . We shall not elaborate on that. For an independent proof of the
following result, see e.g., [273, Appendix C].
290 12. THE INDEX THEOREM FOR CLOSED MANIFOLDS

Theorem 12.8 (Thom Isomorphism in K-Theory). For a complex vector bundle π :


V → X, where X is compact, the homomorphism

Ψ : K(X) −→ K(V ), given by Ψ(a) := (π ∗ a)λV ,

is an isomorphism.

Note. The analogous result for noncompact X is proven in [240, Section IV.1] and
[205, Section 5.2].

Remark 12.9. To indicate the dependence of Ψ on π : V → X, we use the notation

ΨV →X : K(X) −→ K(V ) or simply ΨV .

Definition of the Topological Index. Of concern to us, is a special case of this iso-
morphism which arises as follows. Let X and Y be manifolds and f : X ,→ Y a smooth,
proper embedding. We have f∗ : T X ,→ T Y . While the normal bundle N of X in Y does
not have a complex structure, the normal bundle of T X in T Y does. Thus, the normal
bundle of T X in T Y is T N . This normal bundle is just the pull-back to T X of N ⊕ N .
Roughly, the fiber of T N → T X over v ∈ Tx X consists of pairs of factors thought of as
lying in manifold directions and fiber directions,
(12.3) (
0 u ∈ Nx normal vector to f (X) in Y at f (x) ∈ Y ,
(u, u ) ∈ Nx ×Nx , with
u0 normal vector to f∗ (Tx X) in Tf (x) Y at 0f (x) ∈ Tf (x) Y ;

thus u0 can also be regarded as in Nx under the identification of Tf (x) Y with T0f (x) (Tf (x) Y ).
The complex structure maps (u, v) to (v, −u). Thus, we have

ΨT N →T X : K(T X) −→ K(T N ).

Note that T N can be embedded into T Y as an open subset, and this embedding induces
an extension homomorphism h : K(T N ) → K(T Y ). The composition h ◦ ΨT N →T X
gives us a homomorphism

(12.4) f! := h ◦ ΨT N →T X : K(T X) −→ K(T Y ).

In the case where Y = Rn+m , we have T Y = R2(n+m) . If i : {0} → Rn+m denotes the

inclusion of the origin, then i! : K(T {0}) −→ K(R2(n+m) ), and plainly K(T {0}) ∼
= Z,
−1
since T {0} is just a point. Then i! ◦ f! is a homomorphism,

f i−1
!
indext : K(T X) −→ K(R2(n+m) ) −→
!
K(T {0}) ∼
= Z.

Exercise 12.10. Show that this is well defined (e.g., independent of the choice of f).
[Hint: see [273, p.244]].

Definition 12.11. For an elliptic operator P : C ∞ (E) → C ∞ (F ) with symbol class


[σ(P )] ∈ K(T ∗ X), the topological index of P is defined by

indext [σ(p)] := i−1


! f! [σ(P )].

Note . Here and in the following we identify the cotangent bundle T ∗ X with the
tangent bundle T X by fixing a Riemannian metric for X.

In Theorem 13.1 (p.313f) and Corollary 13.2 (p.315), we shall express the preceding
formula in cohomological terms.
12.2. PROOF OF THE INDEX THEOREM FOR NONTRIVIAL NORMAL BUNDLE 291

Definition of the Analytic Index. The analytic index of A ∈ K(T ∗ X) is defined as


α
follows. Recall that A can be regarded as an equivalence class of bundle maps V0 → V1 for

complex vector bundlesn V0 and V1 over T X. Moreover, itois required that the (nonregular)
support supp(α) := ξ ∈ T ∗ X | α(ξ) ∈ / Iso((V0 )ξ , (V1 )ξ ) be a compact subset of T ∗ X.
Using the compactness of X and the fact that the zero section of T ∗ X is clearly a retract of
the unit ball bundle B(T ∗ X), it is not difficult to show (e.g., see our Exercise 11.4, pp.279ff

or [44, p.492] or [273, p.246]) that we can represent any A ∈ K(T  X) by some Bokobza-
m ∞ ∗ ∗
Haggiag amplitude p ∈ EllBokobza (E, F ) ⊂ C Hom(π E, π F ) for some complex vector
bundles E and F over X, where p(ξ) is an isomorphism for |ξ| ≥ ε > 0 and m ∈ Z or,
even, m ∈ R can be chosen arbitrarily. Here “represent” means that the symbol class [p]
(explained in Definition 12.7) coincides with A.
Note . Here we switch to our preferred Bokobza-Haggiag global symbolic calculus
of Section 8.5, but the arguments run equally well for the conventional symbolic calculus
(where only the principal symbol is defined globally and a canonical representation of an
amplitude by a pseudo-differential operator is missing) of the other sections of Chapter 8
and Section 9.2, so far.
Definition 12.12. The analytic index of A ∈ K(T ∗ X) is defined by
indexa (A) := index(Op(p)) for p ∈ Ellm
Bokobza (E, F ) with [p] = A.

From Section 8.5 we recall that a Bokobza-Haggiag symbol is a section


p ∈ C ∞ (Hom(π ∗ E, π ∗ F )) defining Op(p) : C ∞ (E) → C ∞ (F )
by the following recipe of Equation 8.15 (p.233):
Z
eiξ(v) p(ξ) u∧ (ξ) d̄ξ

(12.5) Op(p)(u)x : =
Tx∗ X
v=0
Z  
= p(ξ) e−iξ(v) ψ(|v|)τx,exp
E
xv
[u(expx v)] d̄vd̄ξ,
Tx X×Tx∗ X

where, as usual, d̄ξ := (2π)−n/2 dξ and dξ is the volume element on Tx X associated


with the metric gx . Note that u∧ denotes the Bokobza-Haggiag Fourier transformation of
Definition 8.30 which is globally defined for fixed choices of metric structures for X, E, F ,
connections defining the parallel moves τ and a smooth bump function ψ : [0, ∞) → [0, 1]
within the injectivity radius of X.
Similarly to the earlier Exercise 12.10, we assign the following task:
Exercise 12.13. Check that index(Op(p)) is independent of the choice of m and
p ∈ EllmBokobza (E, F ) representing A.
[Hint. To p. Show that p0 ∈ Ell0Bokobza (E0 , F0 ) and p1 ∈ Ell0Bokobza (E1 , F1 ) both
represent A precisely when there are vector bundles E e and Fe over X × I , and for
π × Id : (T ∗ X) × I → X × I, a bundle map P : (π × Id)∗ E e → (π × Id)∗ Fe such that
0
P|(T ∗ X)×{t} ∈ EllBokobza (E|
e X×{t} , Fe|X×{t} ), and (for k = 0, 1) isomorphisms ηk and ϕk ,
such that we have a commutative diagram of bundle maps
P |(T ∗ X)×{k}
(π × Id)∗ E|
e (T ∗ X)×{k} / (π × Id)∗ Fe|(T ∗ X)×{k}

ηk ϕk

 pk ⊕Ink 
π ∗ Ek ⊕ Cnk / π ∗ F k ⊕ C nk ,
where Cnk denotes the trivial complex bundle over T ∗ X of fiber dimension nk (see also
[273, p. 247]). Using the invariance of the index under continuous deformation, you have
 
index Op P |(T ∗ X)×{0} = index Op P |(T ∗ X)×{1} .
292 12. THE INDEX THEOREM FOR CLOSED MANIFOLDS

Then using other standard properties of the index, you obtain


index (Op (p0 )) = index (Op (p0 ⊕ In0 ))
= index Op ϕ0 ◦ (p0 ⊕ In0 ) ◦ η0−1

 
= index Op P |(T ∗ X)×{0} = index Op P |(T ∗ X)×{1}
= index Op ϕ1 ◦ (p1 ⊕ In1 ) ◦ η1−1

= index (Op (p1 ⊕ In1 ))
= index (Op (p1 )) .

To m. For q (ξ) := (1 + |ξ|2 )1/2 IdE ∈ Ell1Bokobza (E, E), you have
Mq−m : Ellm 0
Bokobza (E, F ) −→ EllBokobza (E, F ), given by
−m
Mq−m (p) : = p ◦ q ∈ Ell0Bokobza (E, F ) for p ∈ Ellm
Bokobza (E, F ).

Moreover, since the operator Op(q −m ) is invertible, you have


index p ◦ q −m = index (p) .


Deduce that the definition of indexa (A) is independent of the choice of m.]

Formulation of the Multiplicative Property. In the embedding proof of the


Atiyah-Singer Index Theorem, the multiplicative property is used to reduce the index for-
mula for an elliptic pseudo-differential operator over a topologically complicated compact
manifold X to the case of a related pseudo-differential operator over an ordinary sphere
in which X can be embedded. For operators over spheres, the index formula can then
either be checked explicitly or further reduced (by means of Bott periodicity) to the case
of an operator over S 2 or S 1 . At the end of this paragraph (pp.294f), we will explain more
precisely how the multiplicative property fits into the general scheme of the embedding
proof.
Towards the Embedding Proof. Consider an embedding
0 0
(12.6) f : X ,→ Y of the compact manifold X into some manifold Y (say Rn or S n ).
From an elliptic pseudo-differential operator on X, we will construct an appropriate elliptic
pseudo-differential operator, with the same index, on a suitably compactified tubular
neighborhood, say S, of f (X) in Y . In other words, from an amplitude
∞ ∗ ∗ ∗ ∗
a ∈ Ellm
Bokobza (E, F ) ⊂ C (T X, Hom(πX E, πX F )), where πX : T X −→ X

with associated operator Op(a) : C ∞ (E) → C ∞ (F ), one needs to construct suitable com-
e → S and Fe → S and a symbol
plex vector bundles E
∞ ∗ ∗ e ∗ e
(12.7) c ∈ Ellm
Bokobza (E, F ) ⊂ C (T S, Hom(πS E, πS F )),
e e

with associated operator Op(c) : C ∞ (E) e → C ∞ (Fe); here πS : T ∗ S → S. The essen-


tial ingredient which is needed to produce c is an equivariant K-theory element b ∈
KO(m) (T ∗ S m ), where m = n0 − n and S m denotes the unit m-sphere. The choice of
b ∈ KO(m) (T ∗ S m ) which yields index Op(c) = index Op(a) is essentially the famous gen-
erating Bott element b (defined in Theorem 10.20, p. 270), but b will be arbitrary here.
Foundations of Equivariant K-Theory. We begin with a short review of relevant equi-
variant K-theory for those who desire it. The work of Graeme Segal [384] is an excellent,
authoritative exposition of the foundations of equivariant K-theory.
Let G be a group which acts to the left on X, via a ` : G × X → X. We write
g · x = `g (x) = `(g, x). Let π : E → X be a complex vector bundle over X and suppose
that there is a left action of G on E such that π(g·e) = g·π(e) and e 7→ g·e is linear on each
fiber Ex . Then π : E → X is called a G-vector bundle. As an example, if X is a manifold
and G acts on X smoothly, then the action on TC X := C ⊗ T X given by v 7→ d(`g )(v)
for v ∈ TC X makes TC X → X a G-vector bundle. More generally, Λk (TC X) → X is
12.2. PROOF OF THE INDEX THEOREM FOR NONTRIVIAL NORMAL BUNDLE 293

a G-vector bundle. A morphism from a G-vector bundle π1 : E1 → X to a G-vector


bundle π2 : E2 → X is a vector bundle morphism (linear on fibers) ϕ : E1 → E2 such that
ϕ(g · e) = g · ϕ(e). An isomorphism of G-vector bundles is a morphism which is bijective.
The direct sum of G-vector bundles is clearly a G-vector bundle and this operation induces
an abelian semi-group structure on the set of isomorphism classes of G-vector bundles.
We can then form the associated abelian group KG (X) via the Grothendieck construction
of Definition 10.8. Moreover, the tensor product of G-vector bundles yields a G-vector
bundle, and this induces a ring structure on KG (X). For a homogeneous space G/H

where H is a closed subgroup of G, there is a ring isomorphism KG (G/H) −→ R(H) :=
the representation ring of H. Recall that R(H) is the Grothendieck ring obtained from the
abelian semi-group of equivalence classes of representations of H with addition induced
by the direct sum. The tensor product of representations induces a multiplication on
R(H) making it a ring. More concisely, R(H) = KH ({point}). As with ordinary K-
theory, an element of KG (X) can also be described as equivalence classes of G-equivariant
morphisms E → F of G-bundles which are isomorphisms outside of a compact support
(i.e., morphisms with compact support).
We proceed with the construction of c ∈ Ellm
Bokobza (E, F ) in (12.7). For an elaboration
e e
of the used geometric terminology we refer to Chapter 15, see Table 12.1.

Table 12.1. References for applied geometric terminology

Principal G-bundles Definition 15.1, p.395


Representations and associated vector bundles Section 15.3, pp.400ff
Connection 1-forms Definition 15.5, p.398
Bundle of linear frames and subbundle of orthonormal frames Section 15.5, pp.414ff

Let πP : P → X denote the principal O(m)-bundle of orthonormal frames of the normal


bundle N → X for the embedding f : X ,→ Y , where dim X = n and dim Y = n0 . We
regard a frame p ∈ Px as a linear isometry p : Rm → Nx , where m = n0 − n and Nx is the
fiber of the normal bundle at x ∈ X. In terms of associated bundles, we have
N = P ×O(m) Rm = (P × Rm ) / O(m),
where O(m) acts on P × Rm via (p, v) · A := (p ◦ A, A−1 v). Note that O(m) also acts
on Rm+1 = Rm × R via A · (v, a) = (A(v), a), and the m-sphere S m ⊂ Rm+1 is invariant
under this action with two fixed points, the poles (0, ±1) ∈ S m . Let
S := P ×O(m) S m and let Q : P × S m −→ P ×O(m) S m = (P × S m ) / O(m)
denote the quotient map. We may regard πS : S → X as the m-sphere bundle over X
obtained by compactification of the normal bundle N via adjoining the section at infinity.
Choose a so(m)-valued connection 1-form ω on P ; there is actually a natural ω induced
by f : X → Y and a given Riemannian metric on Y . Then we have an O(m)-invariant
distribution H of horizontal subspaces (i.e., Hp = Ker ωp ) on P and hence on P × S m .
By the O(m)-invariance of H, Q∗ (H) is a well defined distribution on S. Moreover, since
πS∗ Q∗ (Hp ) = πP ∗ (Hp ) = Tπ(p) X, Q∗ (H) is complementary to the vertical distribution VS
of tangent spaces of the fibers of πS : S → X. We denote Q∗ (H) by HS . Thus, we have a
splitting
(12.8) T S = VS ⊕ HS = VS ⊕ Q∗ (H).
We also have T S = VeS∗ ⊕ H
∗ e S∗ , where
e S∗ := {α ∈ T ∗ S : α(VS ) = 0} and
H
VeS∗ := {β ∈ T ∗ S : β(HS ) = 0} ∼ ∗ m
= P ×O(m) T S .
294 12. THE INDEX THEOREM FOR CLOSED MANIFOLDS

In view of the splitting (12.8), there are identifications VeS∗ ∼ ∗


= VS∗ := (VS ) and He S∗ ∼
=
∗ ∗ m ∗ m
HS := (HS ) . Note that O(m) acts on the sphere S , and hence on T S via pull-back
of covectors. Thus, we may consider KO(m) (T ∗ S m ). The projection P × T ∗ S m → T ∗ S m
induces a map KO(m) (T ∗ S m ) → KO(m) (P × T ∗ S m ). Moreover, there is the general fact
that if G acts freely on X, then the projection Q : X → X/G induces an isomorphism

Q∗ : K(X/G) −→ KG (X) (see [384, p. 133]). Thus, we have
(Q∗ )−1
(12.9) KO(m) (T ∗ S m ) −→ KO(m) (P × T ∗ S m ) ∼ K(P ×O(m) T ∗ S m ) = K(VS∗ ).
−→

We define
(12.10) K(T ∗ X) ⊗ K(VS∗ ) −→ K(T ∗ S),
as follows. If E → T ∗ X and F → VS∗ are complex vector bundles, then for α0 ∈ T ∗ X and
β 0 ∈ VS∗ , we have unique α ∈ H
e S∗ and β ∈ VeS∗ such that α(v) = α0 (πS ) (v) for v in T S,

and β|VS = β 0 and β(HS ) = 0. Then Eα0 ⊗ Fβ 0 is the fiber of a bundle over T ∗ S at the
point α + β. Thus, we have K(T ∗ X) ⊗ K(VS∗ ) → K(T ∗ S) induced by [E] ⊗ [F ] 7→ [E ⊗ F ].
Using the homomorphisms (12.9) and (12.10), we then have
(12.11) K(T ∗ X) ⊗ KO(m) (T ∗ S m ) −→ K(T ∗ X) ⊗ K(VS∗ ) −→ K(T ∗ S).
For any representation ρ : O(m) → GL(Cq ), we have the associated vector bundle
P ×ρ Cq → X. Let R(O(m)) denote the representation ring of O(m). The assignment
ρ 7→ P ×ρ Cq extends to a ring homomorphism
R(O(m)) −→ K (X) ,
which is to say that K(X) is a R(O(m))-module. Moreover, recall that K (T ∗ X) is a
K (X)-module via u · v = (π ∗ u) v. Thus, ultimately K (T ∗ X) is an R(O(m))-module. We
are now in a position to state
The Multiplicative Property. For v ∈ KO(m) (T ∗ S m ) and u ∈ K(T ∗ X), we have u · v ∈
K(T ∗ S), via (12.11). Moreover, the multiplicative claim
!  
(12.12) indexa (u · v) = indexa indexO(m) v · u ,
where indexO(m) v · u ∈ K(T ∗ X) makes sense since indexO(m) v ∈ R(O(m)), and as


we have just noted, K(T ∗ X) is an R(O(m))-module. In particular, if indexO(m) v = 1 ∈


R(O(m)), then
indexa (u · v) = indexa u.
How the Multiplicative Property Fits into the Embedding Proof. So far, we have not
indicated how the multiplicative property fits into the embedding proof of the index for-
mula. Before proving (12.12), we give the K-theoretical version of the Index Theorem and
show how the proof of it rests on (12.12).
Theorem 12.14 (K-theoretical Version of the Index Theorem). For all u ∈ K(T ∗ X)
we have
(12.13) indexa (u) = indext (u),
where indexa (u) denotes the analytic index of Definition 12.12 (p.291) and indext (u) de-
notes the topological index of Definition 12.11 (p.290).
Recall from Definition 12.11 that we identify T ∗ X with T X by fixing a Riemannian
metric for X.
Proof. To prove the preceding index formula, one needs to show that indexa (u) =
indexa (f! u). Then
indexa (u) = indexa (f! u) = indexa ( i! i−1 (f! u)) = indexa (i! i−1
 
! ! f! u )

= indexa (i−1 −1
! f! u) = i! f! u = indext (u).
12.2. PROOF OF THE INDEX THEOREM FOR NONTRIVIAL NORMAL BUNDLE 295

Since f! = h ◦ ΨT N →T X is a composition of two maps (with h denoting the extension


homomorphism of (12.4), p.{290), the proof that indexa (u) = indexa (f! u) has two parts,
namely
1. indexa (ΨT N −→T X (u)) = indexa (u) and

2. indexa (ΨT N −→T X (u)) = indexa h (ΨT N −→T X (u)) .
Recall that T N is considered as a bundle over T X so that the K-theoretic Thom iso-
morphism ΨT N →T X : K(T X) → K(T N ) is well defined. A precise description of T N as
normal bundle of T X in T Y was given in (12.3).
Part 2 follows from the Excision Property and its proof is easier than part 1 (e.g., see
[273, p. 248 and p. 254]). Part 1 is a consequence of the multiplicative property. Indeed,
for πT∗ N : T N → T X,
ΨT N −→T X (u) = (πT∗ N u)λT N = u · i! 1,
where the last equality follows (in part) from the fact that the associated bundle P ×O(m)
T ∗Sm ∼= VS∗ is isomorphic to T N with one of its two summands compactified; note that

T N = π ∗ N ⊕ π ∗ N where π : T X → X. By various means (none very easy) it is known
that indexO(m) i! 1 = 1 ∈ R(O(m)); see [44, Proposition 4.4, p. 505] or incompletely in
[273, p. 253]. Thus,
indexa (ΨT N −→T X (u)) = indexa ((πT∗ N u)λT N ) = indexa (u · i! 1)
(12.12)  
= indexa indexO(m) i! 1 · u = indexa (u) . 

Proving the Multiplicative Property. Let u = [a] ∈ K(T X) and v = [b] ∈
KO(m) (T ∗ S m ) for first-order elliptic amplitudes
a ∈ Ell1Bokobza (E, F ) and b ∈ EllBokobza (E 0 , F 0 )
O(m),1

which means the following. For g ∈ O(m), let `g : S m → S m be given by `g x = gx.


The differential `g∗ : Tx S m → Tgx S m induces `∗g : Tgx ∗
S m → Tx∗ S m given by `∗g (ξgx )(Yx ) =
O(m),1
ξgx (`g∗ (Yx )) for Yx ∈ Tx S . Then b ∈ EllBokobza (E , F 0 ) means that, for π : T ∗ S m → S m ,
m 0

g ∈ O(m), e0 ∈ Ex0 and ξ ∈ Tx∗ S m , we require that b ∈ C ∞ (T ∗ S m , Hom (π ∗ E 0 , π ∗ F 0 ))


satisfy
ρF 0 (g) b(`∗g ξgx ) e0 = b (ξgx ) ρE 0 (g) e0 ∈ Fgx 0
 
,
where ρE 0 and ρF 0 are the given actions of O(m) on E 0 and F 0 . Note that ρE 0 (g) e0 ∈ Egx
0
,
and `∗g ξgx ∈ Tx∗ S m , since `∗g ξgx (Yx ) = ξgx (`g∗ Yx ). Associated with a and b, there are


pseudo-differential operators Op(a) : C ∞ (E) → C ∞ (F ) on X, and Op(b) : C ∞ (E 0 ) →


O(m),1
C ∞ (F 0 ) on S m . We show that the assumption that b ∈ EllBokobza (E 0 , F 0 ) together with
E0 F0
appropriate choice of connections ∇ and ∇ for E and F 0 implies that Op(b) is O(m)-
0

invariant, in the sense that


(12.14) Op(b) (ρE 0 (g)φ) = ρF 0 (g) Op(b) (φ) , for all φ ∈ C ∞ (E 0 ).
0 0
For this, we assume that ∇E and ∇F are compatible with the O(m)-actions in the
0
0 0
sense that for any curve γ : [c, d] → S m , and parallel translations τγE : Eγ(c) → Eγ(d) and
0
0 0
τγF : Fγ(c) → Fγ(d) , we have
0 0 0 0
ρE 0 (g) ◦ τγE = τ`Eg γ ◦ ρE 0 (g) and ρF 0 (g) ◦ τγF = τ`Fg γ ◦ ρF 0 (g);
i.e., there are commutative diagrams
0
Eγ(c) / Eγ(d)
0
and 0
Fγ(c) / Fγ(d)
0

ρE 0 (g) ρE 0 (g) ρF 0 (g) ρF 0 (g)


   
E`0 g γ(c) / E`0 F`0g γ(c) / F`0 .
g γ(d) g γ(d)
296 12. THE INDEX THEOREM FOR CLOSED MANIFOLDS

Then the invariance (12.14) of Op(b) is shown as follows


Op(b) (ρE 0 (g)(φ))x
Z  0 
= d̄v d̄ξ e−iξ(v) ψ(|v|)b(ξ) τx,exp
E
x v [ρ E 0 (g)(φ)(exp v)]
x
Tx X×Tx∗ X
Z
0
 
= d̄v d̄ξ e−iξ(v) ψ(|v|)b(ξ) ρE 0 (g)τgE−1 x,g−1 expx v [φ(g −1 expx v)]
Tx X×Tx∗ X
Z  0 
= d̄v d̄ξ e−iξ(v) ψ(|v|)ρF 0 (g)b(`∗g ξ) τgE−1 x,g−1 expx v [φ(g −1 expx v)]
Tx X×Tx∗ X
Z  0 
= ρF 0 (g) d̄v d̄ξ e−iξ(v) ψ(|v|)b(`∗g ξ) τgE−1 x,g−1 expx v [φ(g −1 expx v)]
Tx X×Tx∗ X
Z
 −iL∗ ξ(` −1 v)
d̄ `g−1 ∗ v d̄ `∗g ξ e

= ρF 0 (g) g g ∗

(Tg−1 x X)×(T ∗−1 X)


g x
 0 
ψ( `g−1 ∗ v )b(`∗g ξ) τgE−1 x,exp ` v [φ(expg−1 x `g−1 ∗ v)]
g −1 x g −1 ∗

Z  0 
= ρF 0 (g) v d̄ξe e−iξ(ev) ψ(|e
d̄e e τ E−1
v |)b(ξ) [φ(expg−1 x ve)]
e
g x,exp g −1 x
v
e
(Tg−1 x X)×(T ∗−1 X)
g x
 
= ρF 0 (g) Op(b) (φ)g−1 x .
Recall that πP : P → X is a principal O(m)-bundle over X, the bundle of orthonormal
frames of the normal bundle for the embedding f : X ,→ Y . There is a natural connec-
tion, say ω, on P which is inherited from the Levi-Civita connection on the orthonormal
frame bundle for Y . We have a ∈ Ell1Bokobza (E, F ) ⊂ C ∞ (T ∗ X, Hom (π ∗ E, π ∗ F )). For
πT ∗ P : T ∗ P → X, we wish to obtain a lift of a, namely
a ∈ C ∞ (T ∗ P, Hom (πT∗ ∗ P E, πT∗ ∗ P F )),
e
a Rg∗ ξp = e

which is O(m)-invariant in the sense that e a (ξp ). Note that ω gives us a
splitting Tp P = Hp ⊕ Vp and a corresponding splitting Tp∗ P = H e p∗ ⊕ Vep∗ , where
e p∗ := ξ ∈ Tp∗ P : ξ(Vp ) = 0 and Vep∗ := ξ ∈ Tp∗ P : ξ(Hp ) = 0 .
 
H
We have a pull-back πP∗ : T ∗ X → T ∗ P and note that πP∗ ξx ∈ H e p∗ for ξx ∈ Tx∗ X and x =
∗ ∗ ∼ e p . Any ξp ∈ Tp P decomposes uniquely as ξp := ηp + πP∗ ξx
∗ ∗
πP (p). Indeed, πP : Tx X −→ H
for some ηp ∈ Vep∗ and some ξx ∈ Tx∗ X. We simply define
 
ea (ξp ) := a (ξx ) = a πP∗−1 πHe ∗ (ξp ) .
p

Actually, for ξp ∈ e p∗ ,
H a (ξp ) is well-defined without the use of the connection, since
e
ξp ∈ e p∗
H =⇒ ξp = πP∗ ξx for a unique ξx =⇒ e
a (ξp ) = a (ξx ) .
a is O(m)-invariant, since the decomposition Tp∗ P = H
Note that e e p∗ ⊕ Vep∗ is invariant, i.e.,

Rg∗ Tp∗ P = Rg∗ (H e p∗ ) ⊕ Rg∗ (Vep∗ ) = H


e ∗ −1 ⊕ Ve ∗ −1

pg pg

by the Rg∗ -invariance of Hp and πP .


By means of the projection π1 : T ∗ (P × S m ) → X, we may pull back E → X and
F → X to bundles π1∗ E → T ∗ (P × S m ) and π1∗ F → T ∗ (P × S m ). Similarly, we simply
write π2∗ E 0 and π2∗ F 0 for the pull-backs of E 0 → S m and F 0 → S m to T ∗ (P × S m ) via
T ∗ (P × S m ) → S m . Let 1π2∗ E 0 denote the identity automorphism of π2∗ E 0 and let the
trivial extension of e a on T ∗ P to a function on T ∗ (P × S m ) be denoted by e a as well. We
then obtain
a ⊗ 1π2∗ E 0 ∈ C ∞ T ∗ (P × S m ), Hom π1∗ E ⊗ π2∗ E 0 , π1∗ F ⊗ π2∗ E 0 .

e
12.2. PROOF OF THE INDEX THEOREM FOR NONTRIVIAL NORMAL BUNDLE 297

This is but one of the four blocks in the matrix which will yield a representative of
[a] · [b] ∈ K(T ∗ S); see (12.15) below. However, there are difficulties with the required
uniform convergence near ξ = 0 on the sphere bundle |ξ|2 + |η|2 = 1 in the limit defining
the asymptotic symbol (see (8.17))

a ⊗ 1π2∗ E 0 (tξ, tη)
e
a ⊗ 1π2∗ E 0 )(ξ, η) = lim
σ1 (e
t→∞ t
a(tξ) ⊗ 1π2∗ E 0 a)(ξ) ⊗ 1π2∗ E 0 , ξ 6= 0,

e σ1 (e
= lim =
t→∞ t limt→∞ ae(0) t
⊗ 1π ∗ E0 ,
2
ξ = 0, η 6= 0,

σ1 (ea)(ξ) ⊗ 1π2 E 0 , ξ 6= 0,

=
0 ⊗ 1π2∗ E 0 = 0, ξ = 0, η 6= 0.

a ⊗ 1π2∗ E 0 (ξ, η) by ϕr0 (|ξ| , |η|), where the C ∞



This can be remedied by multiplying e
function ϕr0 : [0, ∞)2 → [0, 1] is chosen so that
(
1, for r ≤ r0 or tanr0
θ
≤ 1,
ϕr0 (r cos θ, r sin θ) = tan θ
h( r0 ), for r ≥ 2r0 ,

where the C ∞ function h : [0, ∞) → [0, 1] is chosen so that



1, for s ≤ 1,
h (s) =
0, for s ≥ 2.
Then

ϕr0 (|ξ| , |η|) e
a ⊗ 1π2∗ E 0 (ξ, η)
 q
|η|
|ξ|2 + |η|2 ≤ r0 or

 e a ⊗ 1π2∗ E 0 (ξ, η), for |ξ|
≤ r0 ,
= q
 h( |η| ) e 
a ⊗ 1 ∗ 0 (ξ, η), for |ξ|2 + |η|2 ≥ 2r0 .
r0 |ξ| π2 E

Note that the two formulas agree on the overlap region


 
|η|
q
(ξ, η) : ≤ r0 and |ξ|2 + |η|2 ≥ 2r0 ,
|ξ|
|η|  |η|
since h r0 |ξ| = 1 for |ξ| ≤ r0 . Then

ϕr0 (|tξ| , |tη|) e a ⊗ 1π2∗ E 0 (tξ, tη)
a ⊗ 1π2∗ E 0 )(ξ, η) = lim
σ1 (ϕr0 e
t→∞ t
ϕr0 (|tξ| , |tη|)ea(tξ) ⊗ 1π2∗ E 0
 
|tη| a(tξ)
= lim = lim h ⊗ 1π2∗ E 0
e
t→∞ t t→∞ r0 |tξ| t
   
|η| a(tξ) |η|
= lim h ⊗ 1π2∗ E 0 = h a)(ξ) ⊗ 1π2∗ E 0 .
σ1 (e
e
t→∞ r0 |ξ| t r0 |ξ|
 
The factor h r|η|
0 |ξ|
ensures uniform convergence on the sphere bundle |ξ|2 + |η|2 = 1 as
t → ∞. However, ϕr0 e a ⊗ 1π2∗ E 0 is not an isomorphism for |ξ|2 + |η|2 sufficiently large
because ϕr0 (0, |η|) = 0 if |η| > 2r0 . Thus, ϕr0 e a ⊗ 1π2∗ E 0 is not an elliptic symbol even if
restricted to the subbundle H e ∗ ⊕ T ∗ S m ⊂ T ∗ (P × S m ). This will be remedied when we
consider the full symbol e cr0 (ξ, η) in (12.15), which is elliptic on H e ∗ ⊕ T ∗ S m . Due to the
O(m)-equivariance of
a ⊗ 1π2∗ E 0 ∈ C ∞ (T ∗ (P × S m ), Hom(π1∗ E ⊗ π2∗ E 0 , π1∗ F ⊗ π2∗ E 0 )),
ϕr0 e

a ⊗ 1π2∗ E 0 down to some ϕr0 e
we can push ϕr0 e a ⊗ 1π2∗ E 0 Q in
−1 −1
C ∞ (T ∗ S, Hom((q ∗ ) π1∗ E ⊗ π2∗ E 0 , (q ∗ ) π1∗ F ⊗ π2∗ E 0 )),
 
298 12. THE INDEX THEOREM FOR CLOSED MANIFOLDS

where q ∗ denotes the precursor of



Q∗ : K(P ×O(m) T ∗ S m ) −→ KO(m) (P × T ∗ S m )
on the level of representative vector bundles. However, rather than considering
  
Op ϕr0 e a ⊗ 1π2∗ E 0 Q ,

a ⊗ 1π2∗ E 0 acting on the equivariant sections of π1∗ E ⊗



it is easier to work with Op ϕr0 e
π2∗ E 0 → P × S m which correspond to the sections of (q ∗ )−1 (π1∗ E ⊗ π2∗ E 0 ) → S. One
adjustment must be made when working over P × S m : When defining Op ϕr0 e a ⊗ 1π2∗ E 0
(or ultimately Op(ecr0 (ξ, η))) via a double integral as in (8.15), the integration is restricted
to the product
(Hp ⊕ Tf S m ) × (H e p∗ ⊕ Tf∗ S m ),

as opposed to integrating over all of T(p,f ) (P × S m ) × T(p,f m
) (P × S ), where
∗ m ∗ ∗ m e p ⊕ Tf S m .
T(p,f ) (P × S ) = Tp P ⊕ Tf S = Vep ⊕ H
e p ⊕ Tf S m ) = TQ(p,f ) S, and Ker Q∗ consists of tangent vectors to orbits
Note that Q∗ (H
of the O(m)-action on P × S m . Also,
e p ⊕ Tf S m = T(p,f ) (P × S m ),
Ker Q∗(p,f ) ⊕ H
but generally Ker Q∗(p,f ) " Vep .
Repeating the analogous construction (that we did for a ∈ Ell1Bokobza (E, F )) in the
case of the (pointwise) adjoint
a∗ ∈ Ell1Bokobza (F, E) ⊂ C ∞ T ∗ X, Hom (π ∗ F, π ∗ E) ,


we obtain
a∗ ⊗ 1π2∗ F 0 ∈ C ∞ (T ∗ (P × S m ), Hom(π1∗ F ⊗ π2∗ F 0 , π1∗ E ⊗ π2∗ F 0 )).
ϕr0 e
In a straightforward way, we also obtain lifts of
b ∈ EllBokobza (E 0 , F 0 ) ⊂ C ∞ T ∗ S m , HomO(m) (E 0 , F 0 ) and
O(m),1 

b∗ ∈ EllBokobza (F 0 , E 0 ) ⊂ C ∞ T ∗ S m , HomO(m) (F 0 , E 0 ) .
O(m),1 

to T ∗ (P × S m ) and form
b ∈ C ∞ (T ∗ (P × S m ), Hom(π1∗ E ⊗ π2∗ E 0 , π1∗ E ⊗ π2∗ F 0 )) and
ϕr0 1π1∗ E ⊗ e
b∗ ∈ C ∞ (T ∗ (P × S m ), Hom(π1∗ F ⊗ π2∗ F 0 , π1∗ F ⊗ π2∗ E 0 ).
ϕr0 1π1∗ F ⊗ e
We now define (note the switch from (|ξ| , |η|) to (|η| , |ξ|))
 
ϕr0 1π1∗ E ⊗ e
b (ξ, η) := ϕr0 (|η| , |ξ|)1π1∗ E ⊗ e
b(η) 6= ϕr0 (|ξ| , |η|)1π1∗ E ⊗ e
b(η),
since there is now a non-uniformity of convergence of the asymptotic symbol for small |η|,
as opposed to small |ξ|. For (ξ, η) ∈ H e ∗ ⊕ T ∗ S m ⊂ T ∗ (P × S m ) and for r0 > 0, we define

b∗
" #
ϕr0 (|ξ| , |η|)e
a ⊗ 1π2∗ E 0 − ϕr0 (|η| , |ξ|)1π1∗ F ⊗ e
(12.15) cr0 (ξ, η) :=
e .
ϕr0 (|η| , |ξ|)1π1∗ E ⊗ eb a∗ ⊗ 1π2∗ F 0
ϕr0 (|ξ| , |η|)e
Note that ecr0 (ξ, η) is homogeneous outside a ball bundle of fixed positive radius about the
zero section of H e ∗ ⊕ T ∗ S m . Although we have noted above that the individual entries,
such as ϕr0 (|ξ| , |η|)e
a ⊗ 1π2∗ E 0 , are not isomorphisms for large |η| when ξ = 0 (or in other
cases, for large |ξ| when η = 0), we will show that the entire transformation e cr0 (ξ, η) is
an isomorphism for |ξ|2 + |η|2 large, as follows. Note that
   
a∗ ⊗ 1π2∗ E 0
ϕr0 (|ξ| , |η|)e ϕr0 (|η| , |ξ|) 1π1∗ E ⊗ eb∗
cr0 (ξ, η))∗ := 
(e ,
 
 
−ϕr0 (|η| , |ξ|) 1π1∗ F ⊗ e b ϕr0 (|ξ| , |η|)e
a ⊗ 1π2∗ F 0
12.2. PROOF OF THE INDEX THEOREM FOR NONTRIVIAL NORMAL BUNDLE 299

cr0 (ξ, η))∗ e


and (e cr0 (ξ, η) is block diagonal with entries
 
a∗ e
ϕr0 (|ξ| , |η|)2 e b∗e
a ⊗ 1π2∗ E 0 + ϕr0 (|η| , |ξ|)2 1π1∗ E ⊗ e

b and
 
ϕr0 (|η| , |ξ|)2 1π1∗ F ⊗ eb∗ + ϕr0 (|ξ| , |η|)2 e a∗ ⊗ 1π2∗ F 0 .

(12.16) be ae

Note that for r0 sufficiently large, ϕr0 (|ξ| , |η|)2 and ϕr0 (|η| , |ξ|)2 are not simultaneously
0, since ϕr0 (|ξ| , |η|)2 = 0 holds only in a narrow cone-like wedge about the subspace
ξ = 0, truncated by removing a ball of radius r0 , and ϕr0 (|η| , |ξ|)2 = 0 only in a similar
region about the subspace η = 0. Thus, each of the entries in (12.16) are invertible
(indeed, positive) operators on π1∗ E ⊗ π2∗ E 0 and π1∗ F ⊗ π2∗ F 0 respectively for |ξ|2 + |η|2
sufficiently large, and then e cr0 (ξ, η) is also invertible for |ξ|2 + |η|2 sufficiently large. Since
ϕr0 (|ξ| , |η|) = ϕr0 (|η| , |ξ|) = 1 for |ξ|2 + |η|2 ≤ r02 , we know that for |ξ|2 + |η|2 ≤ r02 ,
cr0 (ξ, η) = e
e c(ξ, η) which (by definition) is the transformation e cr0 (ξ, η) without the ϕr0
factors. Thus, for r0 sufficiently large, the support of e cr0 is the same as that for e c, and
the push down of e cr0 to a function, called (e cr0 )Q on T ∗ (P ×O(m) S m ) is elliptic; i.e.,

−1
cr0 )Q ∈ Ell1Bokobza (q ∗ ) π1∗ E ⊗ π2∗ E 0 ⊕ π1∗ F ⊗ π2∗ F 0 ,
 
(e
 
−1
(q ∗ ) π1∗ F ⊗ π2∗ E 0 ⊕ π1∗ E ⊗ π2∗ F 0 .


O(m),1
cr0 )Q represents [a] · [b] for a ∈ Ell1Bokobza (E, F ) and b ∈ EllBokobza (E 0 , F 0 ),
To see that (e
one goes through the steps leading to the definition (12.11), bearing in mind that when
K-theory elements are defined in terms of compactly supported length-one complexes,
products formed from them (such as the one in (12.10)) are defined in terms of a length-
one complex between sums of tensor products; see also [44, p. 490 and p. 528]. Thus,
 
indexa ([a] · [b]) = index Op (e cr0 )Q .

As pointed out in [44, p. 513f], even though ecQ is not elliptic, it is a limit of the
cr0 )Q as r0 → ∞ in a strong enough sense that Ops (e
elliptic symbols (e cQ ) (for any s ∈ R)
is Fredholm and
 
index Op (e
cQ ) = index Ops (e
cQ ) = index Op (e cr0 )Q = indexa ([a] · [b]) .

We compute index Ops (e cQ ) as follows. Let π̄1 : P × S m → X and π̄2 : P × S m → S m


denote the obvious projections. As we have observed, instead of directly computing the
index of Op (e
cQ ), we can instead compute the index of the equivalent operator

π̄1∗ E ⊗ π̄2∗ E 0 ⊕ π̄1∗ F ⊗ π̄2∗ F 0
 
Op (e
c) : CO(m)

π̄1∗ F ⊗ π̄2∗ E 0 ⊕ π̄1∗ E ⊗ π̄2∗ F 0
 
−→ CO(m)
acting on equivariant (indicated by the subscript O(m)) sections defined on P ×S m . Using
b∗ )
" #
Op (e a) ⊗ 1π̄2∗ E 0 − 1π̄1∗ F ⊗ Op(e
Op(e c) =
1π̄1∗ E ⊗ Op(e b) Op(e a∗ ) ⊗ 1π̄2∗ F 0
and
a∗ ) ⊗ 1π̄2∗ E 0 b∗ )
" #

Op (e 1π̄1∗ E ⊗ Op(e
Op(e
c ) = ,
−1π̄1∗ F ⊗ Op(e
b) a) ⊗ 1π̄2∗ F 0
Op(e
c∗ ) Op(e
we get that Op(e c) is block diagonal with entries
a∗ ) Op (e
Op (e b∗ ) Op(e
a) ⊗ 1π̄2∗ E 0 + 1π̄1∗ E ⊗ Op(e b) and
1π̄1∗ F ⊗ Op(e b∗ ) + Op(e
b) Op(e a∗ ) ⊗ 1π̄2∗ F 0 .
a) Op(e
300 12. THE INDEX THEOREM FOR CLOSED MANIFOLDS

Thus,
Ker Op(e c∗ ) Op(e
c) = Ker (Op(e c))
 
a∗ ) Op(e b∗ ) Op(e

= Ker (Op(e a)) ⊗ 1π̄2∗ E 0 ∩ Ker(1π̄1∗ E ⊗ (Op(e b)))
 
a∗ ) ⊗ 1π̄2∗ F 0 ∩ Ker(1π̄1∗ F ⊗ Op(e b∗ ))

⊕ Ker Op(e a) Op(e b) Op(e
  
= Ker Op(e a) ⊗ 1π̄2∗ E 0 ∩ Ker(1π̄1∗ E ⊗ Op(e b))
 
⊕ Ker(Op(e a∗ ) ⊗ 1π̄2∗ F 0 ) ∩ Ker(1π̄1∗ F ⊗ Op(eb∗ )) ,

and
c∗ ) = Ker (Op(e
Ker Op(e c) Op(e c∗ ))
 
a∗ ) ⊗ 1π̄2∗ E 0 ∩ Ker(1π̄1∗ F ⊗ Op(e b∗ ) Op(e

= Ker Op(e a) Op(e b))
 
a∗ ) Op(e b∗ ))

⊕ Ker Op(e a) ⊗ 1π̄2∗ F 0 ∩ Ker(1π̄1∗ E ⊗ Op(e b) Op(e
 
a∗ ) ⊗ 1π̄2∗ E 0 ∩ Ker(1π̄1∗ F ⊗ Op(e

= Ker Op(e b))
 
b∗ )) .

⊕ Ker Op(e a) ⊗ 1π̄2∗ F 0 ∩ Ker(1π̄1∗ E ⊗ Op(e

Since Op(ea) ⊗ 1π̄2∗ E 0 commutes with 1π̄1∗ E ⊗ Op(e a) ⊗ 1π̄2∗ E 0 preserves


b), we have that Op(e
 
Ker 1π̄1 E ⊗ Op(b) , and
∗ e

  
a) ⊗ 1π̄2∗ E 0 ∩ Ker 1π̄1∗ E ⊗ Op(e
Ker Op(e b)
 

= Ker a) ⊗ 1π̄2∗ E 0 |Ker1
Op(e  .
π̄ ∗ E ⊗Op(b)
e
1

Similarly,
 
b∗ )
a∗ ) ⊗ 1π̄2∗ F 0 ∩ Ker 1π̄1∗ F ⊗ Op(e

Ker Op(e
 
= Ker (Op(ea∗ ) ⊗ 1π̄2∗ F 0 )|Ker1 e∗
 .
π̄ ∗ F ⊗Op(b )
1

Thus,
 

Ker Op(e
c) = Ker a) ⊗ 1π̄2∗ E 0 |Ker1
Op(e 
π̄ ∗ E ⊗Op(b)
e
1
 
a∗ ) ⊗ 1π̄2∗ F 0 )|Ker1
⊕ Ker (Op(e e∗
 ,
π̄ ∗ F ⊗Op(b )
1

and similarly
 
c∗ ) = Ker (Op(e
Ker Op(e a∗ ) ⊗ 1π̄2∗ E 0 )| Ker1 
π̄ ∗ F ⊗Op(b)
e
1
 

⊕ Ker a) ⊗ 1
Op(e ∗F 0
π̄2 |Ker1 ∗ ⊗Op(eb∗ ) .
π̄ E
1

We note that
   
b) = C ∞ π̄1∗ (E) ⊗ Ker Op(e
Ker 1π̄1∗ E ⊗ Op(e b) and
   
b∗ ) = C ∞ π̄1∗ (F ) ⊗ Ker Op(e
Ker 1π̄1∗ F ⊗ Op(e b∗ ) .
12.3. COMPARISON OF THE PROOFS 301


a) ⊗ 1π̄2∗ E 0 |Ker1
Thus, Op(e ∗ E ⊗Op(b)e
 is a differential operator on
π̄1
 

CO(m) π̄1∗ (E) ⊗ Ker Op(e
b) ;

i.e., a differential operator on the O(m)-invariant sections of π̄1∗ (E) ⊗ Ker Op(e
b), where
Ker Op(e b) is a finite-dimensional O(m)-module, and similarly for
a∗ ) ⊗ 1π̄2∗ F 0 )|Ker1
(Op(e e∗ ) .
∗ F ⊗Op(b

π̄1

Since Op(e
a) is an O(m)-invariant lift of Op(a), we have an isomorphism of O(m)-modules,
 
a) ⊗ 1π̄2∗ E 0 |Ker1 ∗ ⊗Op(eb) ∼

Ker Op(e = Ker (Op(a)) ⊗ KerO(m) Op(b),
π̄1 E

where the action is trivial on the Ker (Op(a)) factor. Similarly,


Ker((Op(ea∗ ) ⊗ 1π̄∗ F 0 )| ∼ ∗ ∗
e∗ ) = Ker (Op(a )) ⊗ KerO(m) Op(b ),
2 Ker(1π̄∗ F ⊗Op(b ))
1
∼ Ker (Op(a∗ )) ⊗ KerO(m) Op(b),
a∗ ) ⊗ 1π̄2∗ E 0 )| Ker(1 ∗ ⊗Op(eb)) ) =
Ker((Op(e
π̄1 F

a) ⊗ 1π̄2∗ F 0 |Ker(1 ∗ ⊗Op(eb∗ )) ) ∼ ∗



Ker( Op(e
π̄1 E
= Ker (Op(a)) ⊗ KerO(m) Op(b ).

Hence, as required,
index Op(e c)) − dim(Ker Op(e
c) = dim (Ker Op(e c∗ ))
   

 Ker Op(e a) ⊗ 1π̄2 E 0 |Ker 1 ∗ ⊗Op(eb)
∗  
π̄1 E

= dim    

 
⊕ Ker (Op(e a ) ⊗ 1π̄2∗ F 0 )|Ker 1 ∗ ⊗Op(eb∗ )

π̄1 F
   

Ker (Op(ea ) ⊗ 1 ∗ 0 ) 
π̄2 E | Ker 1 ∗ ⊗Op(e b)

π̄1 F
 
− dim 



 

⊕ Ker Op(e a) ⊗ 1π̄2∗ F 0 |Ker1 ∗ ⊗Op(eb∗ )
π̄ E
 1
= dim Ker (Op(a)) ⊗ KerO(m) Op(b)
− dim(Ker (Op(a∗ )) ⊗ KerO(m) Op(b))
+ dim(Ker (Op(a∗ )) ⊗ KerO(m) Op(b∗ ))
− dim(Ker (Op(a)) ⊗ KerO(m) Op(b∗ ))
= index [a] · KerO(m) Op(b) − KerO(m) Op(b∗ )

 
= index [a] · indexO(m) [b] = index u · indexO(m) v .

3. Comparison of the Proofs

I Isadore Singer (born 1924) and Michael Atiyah (born 1929) — we prefer the
order by age to the common lexicographic one — gave two more proofs of the Index
Formula, in addition to the embedding proof given above. These are the original cobordism
proof, presented in [328] in all details, and the newer heat equation proof, presented in
various books and proved for twisted Dirac operators here in Chapter 17 (pp.513–642) in
unusual detail. In this section, we cannot summarize them here, but we will comment
briefly.
All three of these proofs appear to be somewhat complicated. Several authors (among
others, [65], [103], [378], [380], [259] and [163]) tried to give simpler or more elementary
proofs for the Euclidean case and/or geometric operators. In the early judgment of [16,
p.245] “these different proofs differ only in the use and presentation of algebraic topology”
(instead of, and at times together with, the Bott Periodicity Theorem “older but not at
all elementary parts of topology” are employed) — “the analysis is essentially the same in
302 12. THE INDEX THEOREM FOR CLOSED MANIFOLDS

origin”. When M.F. Atiyah, R. Bott and V.K. Patodi found their radically different
approach via heat asymptotics, various true simplifications appeared. In particular, we
refer to the work of E. Getzler and P. Gilkey who have inspired our Chapter 17.
Moreover, in the framework of noncommutative geometry, A. Connes and collaborators
achieved the Index Theorem as a special case of natural localizations in operator algebras.
In this monograph, however, we shall not comment upon these approaches that belong to
a much wider context. J

The Cobordism Proof. That proof is sketched in [43] and worked out in detail
in [95], [109] and [328]. It was the first proof: It begins with a compact, oriented
Riemannian manifold (without boundary) of dimension 2l and defines d : Ωj → Ωj+1 and
δ : Ωj+1 → Ωj as the exterior (Cartan-) derivative of forms and its adjoint.1
These forms and derivatives of forms are explained in our Section 6.4 on Exterior
Differential Forms and Exterior Differentiation, P2l see j particularly Exercise 6.20, p. 172.
Recall Ωj := C ∞ Λj (T ∗ X) ⊗ C and Ω• := j=0 Ω . Then d + δ : Ω•
→ Ω• is a self-
adjoint, elliptic differential operator of first order whose square is the Laplace operator ∆
of Hodge theory (our Theorem 13.6a). If ∗ : Ωp → Ω2l−p denotes the Hodge star (duality)
operator of Exercise 6.20c, the formula of the same Exercise
(12.17) τ (v) := ip(p−1)+l ∗v, v ∈ Ωp
defines on Ω• an involution (i.e., τ ◦ τ = Id). If Ω± denote the ±1 eigenspaces of τ , we
define the signature operator (see Section 13.4 below and [45, p.575]) (d + δ)+ : Ω+ → Ω−
to be the restriction of d+δ to Ω+ . One can show (see below Theorem 13.6c) that (d + δ)+
is an elliptic operator and its index is the signature of the manifold X, often named after
Friedrich Hirzebruch.
For sufficiently many special manifolds (specifically for X = S 2l and X = Pl (C) :=
complex projective space of complex dimension l) one can now compute the signature (i.e.,
index(d + δ)+ ) using cohomology theory and derive an index formula for manifolds of even
dimension. Sufficiently many here means four things:
(i) By a deep result of cobordism theory in [413] by René Thom, every even-
dimensional manifold Y is in a certain sense cobordant to the special manifolds; in other
words, there is a bounded manifold Z whose boundary is built up from X and Y (see Fig-
ure 12.2). The concept of bordism is much coarser than homotopy. After first attempts
in 1895 by Henri Poincaré, cobordism was defined and successfully applied in 1938 in
[339] by the young Lev Pontryagin to relate cobordism of smooth manifolds to stable
homotopy of spheres.

X
Z Y

Figure 12.2. The manifolds X and Y are cobordant via Z

1
In [15], Atiyah gave his first public announcement of the index theorem on 16 July, 1962 at
the Bonn Arbeitstagung. There, the theorem was explained for (elliptic) Dirac operators on spin
manifolds. These operators were invented by Singer and Atiyah for that purpose and serve as a
model for differential operators associated with some geometrical structures since then.
12.3. COMPARISON OF THE PROOFS 303

(ii) Furthermore, René Thom proved the vanishing of the signature for bounding
manifolds.
Note. A modern formulation and proof of this (vanishing index) Cobordism Theorem
for operators of Dirac type is given in [83, Theorem 21.5] and generalized in [80, Section 6]
for any arbitrary linear formally self-adjoint (i.e., symmetric) elliptic differential operator
B over a compact manifold Z with smooth boundary ∂Z (= X ∪ (−Y ) in the actual
application) satisfying a weak inner unique continuation property: Then the induced
tangential operator ∂B over ∂Z splits naturally in block matrix form with index(∂B + ) =
0 for the induced lower left part operator ∂B + . As a matter of fact, a different, but
mathematically equivalent result was obtained by James Ralston already in 1970, but
ignored by topologists in 40 years. The result of [350] was simply that any such operator
B admits a regular (globally elliptic) boundary condition making B — subjected to that
boundary condition — self-adjoint and Fredholm. That yields a symplectic splitting of
∂B and the wanted vanishing of index(∂B + ).
Hence the index of the signature operator on an arbitrary 2l-dimensional manifold X
can be computed from the indices of the special signature operators [207, p.58].
(iii) For a Hermitian C ∞ -vector bundle E over X, let
 
ΩjE := C ∞ E ⊗ Λj (T ∗ X)

denote the space of j-forms with coefficients in E. By means of a covariant derivative ∇E


(inducing parallel translation along paths), one can define the operator
dE : ΩjE −→ Ωj+1
E via
E E
d (v ⊗ u) := ∇ v ∧ u + v ⊗ du, for u ∈ ΩjE and v ∈ C ∞ (E)
and its adjoint (dE )∗ , and one can establish the index formula for the generalized signature
+
operator dE + (dE )∗ for the vector bundle E which is defined by means of an involution
P2l
on ΩE := j=0 ΩjE .
This concludes the proof since every elliptic operator on a closed oriented smooth even-
dimensional oriented manifold X is equivalent in the sense of K-theory to a generalized
signature operator. More precisely:
Lemma 12.15 (Main Lemma of Global Analysis). Let X be a closed oriented smooth
even-dimensional oriented manifold. Then K(T X) is a ring over K(X) and the subgroup
K(X) · [σ((d + d∗ )+ )] generated by generalized signature operators via
  + 
E E∗
σ d +d = [E] · [σ((d + δ)+ )]

is so large, namely a subgroup of finite index, that practically all of K(T X) is generated.
Here practically means up to the image of K(X) in K(T X) and up to 2-torsion, where
the index must vanish as an additive function with values in Z.
At this place, the theory of pseudo-differential operators enters in order to achieve
that the symbols are arbitrary bundle isomorphisms over SX, and to allow reduction of
the index computation from Ell(X) to K(T X). Further, the Bott Periodicity Theorem is
used in somewhat generalized form in representing K(T X) approximately by K(X); see
[33, p.321f] and our explanations to the K-theoretic Thom Isomorphism in Theorem 12.8
(p.290).
(iv) The general index formula can be extended to an odd-dimensional manifold X,
using the multiplicative property of the index by tensoring with the standard operator
T with index 1 on S 1 and by passing to the even-dimensional manifold X × S 1 . If one
is not interested in the sign in the index formula, one can avoid the explicit definition
of T and simply pass to the squared (relative to the tensor product) operator on the
even-dimensional manifold X × X.
304 12. THE INDEX THEOREM FOR CLOSED MANIFOLDS

The Embedding Proof. It was given in Section 12.1, following [16], [18] and [21].
In this proof, the methods remain topological with the consideration of K(T X) instead of
the operator space Ell(X). The idea goes back to the proof of the Hirzebruch-Riemann-
Roch Theorem by Alexander Grothendieck explained in [86], see also Section 13.7
below. Roughly speaking, the basic difference between the Hirzebruch-inspired first
proof of the index theorem and the Grothendieck-inspired second proof is what they
consider elementary: spheres and projective spaces, as Hirzebruch did, following Thom,
or just points (or spheres and Euclidean spaces) as Grothendieck could do by replacing
the intricacies of cobordism by the more elementary concept of embedding and support it
by elaborate new structures. In the K-theoretic embedding proof of the index theorem,
one shows first, using the Bott Periodicity Theorem, that every elliptic operator on the
sphere or Euclidean space is equivalent, in sense of K-theory, to one of infinitely many
(more precisely only |Z|-many) standard operators. Then the case of an arbitrary elliptic
operator on arbitrary closed manifold is reduced to the standard case by embedding.
The advantage, as well as the weakness, of this proof lies in its perhaps some-
what forced directness. It succeeds on the one hand in eliminating cohomology and
cobordism theory completely, bringing out the functional analytic and topological pil-
lars (F. Noether’s Index Formula — usually ascribed to I. Gohberg and M. Krein —
and R. Bott’s Periodicity Theorem) plainly and in their most elementary form, and in
achieving through this simplicity of tools the greatest susceptibility to generalization (see
Section 13.11). On the other hand, under the imbedding (except for particularly smooth
ones, e.g. holomorphic embeddings of algebraic manifolds in a complex projective space)
the special structure of classical operators is completely destroyed. For example, the sig-
nature operator does not become another signature operator under the imbedding, and to
prove the Riemann-Roch Theorem for arbitrary compact complex manifolds (to mention
another problem defined by classical operators; see also Sections 13.7 and 17.6, pp.615ff
below), one has to leave this category, whereby many of the interesting and sometimes
open problems of modern differential topology become less transparent.

The Heat Equation Proof. In comparison with the first two proofs, which argue
more topologically, the heat equation proof of [33, 34] offers a completely different and,
initially, purely analytic approach to the index problem. The germinal idea goes back to
papers of Marcel Riesz on spectral theory of positive self-adjoint operators, and was
presented by M. F. Atiyah as early as 1966 at the International Congress of Mathemati-
cians in Moscow, and then published, also in connection with applications of the Index
Formula to fixed point problems in [31], [19] and in related form in [103], [380] and [152]
(see our Chapter 17, pp.513–642 for a comprehensive presentation):
1. From spectral analysis of the two nonnegative self-adjoint operators P ∗ P and P P ∗
to index(P ). One starts with an operator P ∈ Ellk (E, F ), k > 0, where E and F are
Hermitian C ∞ vector bundles on the n-dimensional, closed, oriented, Riemannian mani-
fold X. Then the operator P ∗ P is a nonnegative self-adjoint operator of order 2k with
a discrete spectrum (see Chapter 3 above) of nonnegative eigenvalues 0 ≤ λ1 ≤ λ2 ≤ · · ·
(the multiplicity may be larger than 1, hence “≤”), and the series
X∞
(12.18) θP ∗ P (t) := e−tλm
m=1

converges for all t > 0. By the way, for X = S 1 and P = −i dx d


, we obtain the theta
P∞ −tm2
function θ(t) = m=0 e of analytic number theory since the square integers are
exactly the eigenvalues of P ∗ P = ∆ = −d2 /dx2 .
Correspondingly, one forms the function θP P ∗ . The operators P ∗ P and P P ∗ have the
same nonzero eigenvalues, and only the eigenvalue 0 has in general different multiplicities,
namely dim Ker P and dim Ker P ∗ (see the Remark 2.11, p. 17). This way, one has a new
12.3. COMPARISON OF THE PROOFS 305

index formula
(12.19) index P = θP ∗ P (t) − θP P ∗ (t), t > 0.
Trivially, one may choose an arbitrary function φ on E with φ(0) = 1 instead of the
function m 7→ e−tm and thus obtain for each φ a further index formula
X X
index P = φ(λ) − φ(λ).
λ∈Spec(P ∗ P ) λ∈Spec(P P ∗ )

Now, the theta function is distinguished by permitting near t = 0 an asymptotic


development (for a precise definition of asymptotics, see Definitions 17.36 and 17.37 below
in Chapter 17)
X m/2k Z
(12.20) θP ∗ P (t) ∼ t µm (P ∗ P ) , (as t → 0+ ),
m≥−n X


where µm (P P ) is, for each m ∈ Z, a certain density on X which can be formed canonically
from the coefficients of the operator P ∗ P . Using (12.19) we get from (12.20) the explicit
integral representation
Z
(12.21) index P = µ0 (P ∗ P ) − µ0 (P P ∗ ) .
X
2. How the Heat Equation Enters. The convergence of the series in (12.18) has impli-
cations for the construction of solutions of the heat conduction equation (where ∆P := P ∗ P
is a generalized Laplace operator), as emphasized in [167, p.64-65]. Apparently that was
discovered and exploited first by the Swedish genius (and strange character) Torsten
Carleman in [108] when he found the poles of the ζ-function

X
ζ(s) := tr(∆−s ) = λ−s
m
m=1

of the Laplacian for a compact region X ⊂ R2 subject to the Dirichlet or Neumann


boundary condition from the asymptotic behavior of the heat kernel

X
K(x, y, t) = vm (x) vm (y) e−tλm , x, y ∈ X,
m=1

with {vm ; λm }m∈N denoting the normalized eigenfunctions and eigenvalues. The basic
ideas go even further back in history, namely to the expression of the fundamental solution
of the heat equation by sums of E.E. Levi, [279] (also called Hilbert-Levi’s parametrix
method ). In the late 1960s, apparently Takeshi Kotake was the first to connect the
asymptotic analysis of parabolic equations directly to index theory in [258, 259]. At that
time, asymptotics were well studied in probability theory, for instance in [420]. That
paper was used by McKean and Singer in their legendary [291] where help by Kotake
with the Levi sums is acknowledged. So much for the credits. Consider
∂u
(x, t) + ∆P u(x, t) = 0, x ∈ X, t ∈ [0, ∞)
∂t
with the initial condition u(·, 0) = u0 ∈ L2 (E). Here u is the unknown function on
X × [0, ∞) with values in the bundle E (the heat distribution). Then
t2 2 t3
Ht := e−t∆P = Id −t∆P + ∆P − ∆3P + · · · , t ≥ 0
2! 3!
is a well-defined family of bounded operators on the Hilbert space L2 (E) which satisfies
the heat equation
dHt
+ ∆P Ht = 0
dt
with initial value H0 = Id. Thus H yields for each initial distribution u0 the heat distri-
bution at time t via the formula u(·, t) = Ht u0 .
306 12. THE INDEX THEOREM FOR CLOSED MANIFOLDS

v m : m ∈ Z+

Since the eigenfunctions of ∆P form a complete orthogonal system
for L2 (E), the formula
X∞
tr e−t∆P = θ∆P (t) = e−tλm
m=1

is meaningful. The convergence of the series in (12.18) means that the evolution operators
Ht of the parabolic heat conduction equation belongs to the trace class for t > 0. By
means of the theory of pseudo-differential operators it follows more precisely that Ht is
a smoothing operator, i.e., an operator of order −∞ which is representable as an integral
operator
Z
(Ht v)(x) = Kt (x, y) v(y) ωy , v ∈ L2 (E), x ∈ X
X
with C ∞ weight function (kernel ) (x, y) 7→ Kt (x, y) ∈ L(Ey , Ex ) for t > 0 and volume
element ω.
3. Time Independence and Asymptotics. Then
Z
θ∆P (t) = tr Ht = µt , t > 0, where
X
X∞
µt (x) : = tr(Kt (x, x)) ωx = e−tλm |vm (x)|2 ωx
m=1

defines a density on X which at each x ∈ X is the pointwise trace of the operator Kt ,


times ωx . While µt can be expressed in terms of the coefficients of the operator P only
very indirectly (via the eigenvalues and eigenfunctions of ∆P = P ∗ P ), one has at each
point x ∈ X an asymptotic expansion
X m/2k
µt (x) ∼ t µm (x) for t → 0,
m≥−n

where the µm are purely local invariants of P ∗ P , which then implies (12.20).
4. An Intuitive (but Impractical) Solution. A proof of (12.20) with a recipe for the
computation of µm extracted from the theory of pseudo-differential operators is due to
[380]. It shows that the µm depend rationally on the coefficients of P and their derivatives
of orders ≤ n. A more intuitive and heuristic description of the µm for the special case
X = T n := Rn /(2πZn ), to which we paid special attention in our Sobolev case studies
(Chapter 7), can be found in [40, p.300f]. The idea of the proof goes back to the Indian
mathematician Subbaramiah Minakshisundaram and Carleman’s student Åke Plei-
jel, who in 1949 (long before an effective machinery for pseudo-differential operators was
established) computed the µm for the case P ∗ P = ∆, where ∆ denotes the invariantly
defined Laplace-Beltrami operator which depends only on the Riemannian metric on X,
i.e., P ∈ Ellk (E, F ) with k = 1 and E = CX . Precisely, S. Minakshisundaram and Å.
Pleijel (and later R. T. Seeley and T. Kotake, when generalizing their results) stud-
ied in place of P ∗ P the positive self-adjoint
P∞ operator  = Id +P ∗ P and, in place of the
−z
theta function, the zeta function ζ(z) := m=1 (λm ) summed over all (discrete positive)
eigenvalues of . As is well-known, the zeta function is well-defined for <(z) > dim X and
can be continued to a meromorphic function in the z-plane with finitely many real poles
of order 1 with behavior at poles known in principle. Among other things it is found that
z = 0 is not a pole and that the value ζ(0) can be expressed explicitly in terms of ∆P ; in
fact,
Z
(12.22) ζ(0) = ρ0 (∆P ) ,
X

where the right hand side is fairly complicated but can be computed in principle. On
the other hand, ζ(z) can be interpreted within spectral theory as trace(∆−z
P ), and finally
ζ(0) appears as the constant term in the asymptotic expansion of θ(t) as t → +∞. This
establishes the connection with the heat conduction approach. In particular, the measure
12.3. COMPARISON OF THE PROOFS 307

µ0 (∆P ) sought there is identical with the measure ρ0 (∆P ) in equation (12.22), and (12.20)
follows from (12.22) and similar computations of the residues of ζ at its poles.
5. From a General Algorithm to a Comprehensible Formula via Cancelation Proce-
dures. Thus, a general algorithm is available that is capable of producing the right hand
side of the index formula (12.21) in finitely many steps by means of a computer for ex-
ample. In contrast to the index formulas of the cobordism and imbedding proofs (into
which enter the derivatives of the coefficients of P up to order 2 only) this formula is in
the general case complicated numerically and algebraically mainly by the appearance of
derivatives up to order n (= dim X). While for algebraic curves of complex dimension 1
(= Riemannian surfaces, n = 2), the formula can be handled well computationally, the
general situation requires so much effort that M. F. Atiyah and R. Bott by their own
admission had initially “little hope for interpreting these integrals directly in terms of the
characteristic classes of E and X”, and therefore “the beautiful formula appeared to be
useless in this context.”
Only a series of subsequent papers on curvature tensors revealed that in the special
case when X is even-dimensional and P is the signature operator (d + δ)+ , all higher
order derivatives cancel out in Seeley’s formula for the measure µ0 (P ∗ P ), and only the
derivatives up to order 3 remain. Details of such computations appear first in [291], in the
case (not all that fortunate for this aspect) that P is the operator d+δ : Ωeven → Ωodd with
the Euler characteristic as index (see Section 13.4). This was generalized in 1971 by V.
K. Patodi — again by means of symmetry considerations — and extended to Riemann-
Roch operators (see Section 13.7) in particular. P. Gilkey succeeded shortly thereafter
in replacing Patodi’s complicated group theoretic cancellation procedure for the higher
derivatives by an axiomatic argument which was drastically simplified in [33] through
the use of stronger tools of Riemannian geometry. It says roughly that in each integrand
with the general qualitative properties of µ0 (P ∗ P ) the annoying higher derivatives can be
disregarded and µ0 (P ∗ P ) can be identified with the (normalized) Gaussian curvature.
In this fashion, a new purely analytic proof of the Hirzebruch Signature Theorem (the
index formula for classical operators) is achieved which implies the general index formula,
as in the cobordism proof (see above, items (iii) and (iv) in the Section 3), and with the
same topological arguments.
6. Heat Equation Asymptotics and Spectral Invariants. The significance of the heat
equation proof, which cannot be extended (just like the cobordism proof) to families of
elliptic operators and operators with group action, is at present difficult to estimate. Its
authors, who (as an aside) acknowledge that their “whole thinking on these questions has
been stimulated and influenced very strongly by the recent paper of Gelfand on Lie al-
gebra cohomology” [160], point out that their proof is hardly shorter than the imbedding
proof since it “uses more analysis, more differential geometry and no less topology. On
the other hand, it is more direct and explicit for the classical operators associated with
Riemannian structures: In particular, the local form of the Signature Theorem and its gen-
eralizations are of considerable interest in itself and should lead to further developments”
[33, p.281].
This prediction appears to materialize even beyond the realm of differential geometry:
The approach via the zeta function of the Laplace-Beltrami operator ∆ on X (whose values
yield real-valued invariants of the Riemannian metric ρ of X – spectral invariants – at any
point where the zeta function does not have a pole) has been extended to the systems
case, where the Laplace equation is replaced by the system of partial differential equations
of the total Laplace operator of Hodge theory which can be represented as the square of a
formally self-adjoint operator D (the Dirac operator). In analogy to the zeta function, one
considers for an operator D, which is not positive, the function η(z) := λ6=0 (sign λ) |λ|−z
P

where summation is over the eigenvalues of D with proper multiplicity. Again η(z) is a
holomorphic function for large |z|, which can be continued meromorphically to the whole
z-plane. Corresponding to the asymptotic expansion above in equation (12.20) for the
308 12. THE INDEX THEOREM FOR CLOSED MANIFOLDS

theta and zeta functions, one can look for an integral formula for η(0). Such a formula
Z
(12.23) η(0) = α(ρ̃) − integer
γ

is proven in [41, Thm. 4.14], when X can be obtained as the boundary of a 4-dimensional
manifold Y with Riemannian metric ρ̃, see Figure 12.3.

Figure 12.3. Typically, an integral formula for the spectral in-


variant η(0) of X is obtainable, if X is a boundary

Here ρ̃ is assumed to induce on X the metric ρ and to render a neighborhood of X


in Y isometric to X × [0, 1). The integrand α(ρ̃) is explicitly known (the first Hirzebruch
L-polynomial in Pontryagin forms of the Riemannian metric ρ̃), as well as the integral
correction term (the signature of Y defined by the topology of Y ). In this way, a
formula is obtained which relates the spectral invariant η(0), measuring the asym-
metry of the spectrum of D, with a differential geometric and a purely topological
invariant.
Perspectives.
I Keeping in mind that the Atiyah–Singer Index Theorem is given by local invari-
ants, and that the η–invariant (and R–torsion) are the next level, since their derivatives
are local, we meet a question repeatedly put forward by Gelfand: “what comes next?” To
this, Singer remarked in personal communication ([398]): “Just as η arises in boundary
value problems for smooth boundaries, I think the next level will come from corner contri-
butions when the boundary has corners.” Indeed, one is tempted to claim a hierarchy of
asymmetry invariants. At ground level, we have closed manifolds and the index of elliptic
operators. The index is a topological invariant which expresses the asymmetry between the
multiplicity of the zero-eigenvalues of A and A∗ . At the next level we have manifolds with
smooth boundary and elliptic operators with suitable (elliptic/regular) boundary condi-
tions. There, the index formula contains an additional, correction term, the η-invariant
of the induced operator over the boundary, see (12.23). Actually, the integer in (12.23) is
the index of the boundary value problem with spectral (Atiyah-Patodi-Singer) boundary
condition. Finally, when the boundary has corners, a third term, the Hörmander index of
symplectic analysis appears, see [428], elaborated further in [100, 101].
Surprisingly, each of the asymmetry (correction or error) terms has developed its own
life and surfaced with an independent meaning: the index in the integrality theorems dis-
cussed in this book; the η-invariant as the phase of the ζ-function regularized determinant
(elaborated in [377]); and the Hörmander and Maslov indices of symplectic analysis in
recent generalizations of classical Morse theory (see the historical account in [84]).
The formula (12.23) lies deeper than (12.20) or (12.21), where the integrand was
of local type. It has numerous relations to the cobordism as well as the embedding
(particularly via boundary-value problems) proof. It is not as esoteric as it may appear
to someone not so much interested in differential-topological problems, since it yields
(roughly) a direct geometric interpretation for the peculiarities of the distribution of the
12.3. COMPARISON OF THE PROOFS 309

eigenvalues of an operator D of Dirac type and of the Dirac Laplacian D2 : Classical results
in this direction by Victor W. Guillemin and others, say for example, that a Riemannian
manifold X is isometric with S n if the spectra of the Laplace operators coincide. One
cannot always expect that two manifolds with equal spectra of their Laplace operators are
isometric (the 16-dimensional tori yield a counterexample in [298]), but it appears that at
least extreme distributions of the eigenvalues (when they are not randomly distributed, but
lumped together near integers for example) carry with them extreme geometric situations
(in our example the closedness of the geodesics). According to an announcement in [397],
this answers in principle the classical question Can you hear the shape of the drum? (Mark
Kac) which had already motivated [291], see Figure 12.4. We discussed the present state
of knowledge of spectral geometry in Section 3.10 of our Part I. See also our review for
physicists in [74, Section 3].

?
Figure 12.4. One instant of spectral synthesis: Can you hear the
shape of a drum?

Further, a gate opened to the inverse problem of mathematical modeling of real phe-
nomena: Theoreticians frequently and on perforce only apply theory; i.e., they try, similar
to the axiomatic method within mathematics, to draw conclusions as far-reaching as pos-
sible about the concrete behavior from relatively modest assumptions about the existence
of certain laws. Conversely, the practitioner needs in general the inverse of the theory,
namely the exposure of regularities in the observations at hand. (Somewhat overstated,
the practitioner desires to fit a curve to given measurements, while the theoretician sees
his strength in a detailed discussion of the properties of a given curve.) In this sense,
the novelty consists in the attempt to estimate the parameters of a differential equation,
when information about special solutions (eigenfunctions and eigenvalues, for example) is
available.
According to a communication by Richard Bellman, the inverse problem in its most
general formulation goes back to Carl Gustav Jacob Jacobi (1804-1851). Today, the
Inverse Problem is studied in very different contexts, reaching from algebraic problems
[419, p.32] to structure identification and parameter estimation in control theory. A special
version of the Inverse Problem is the inversion of spectral analysis, spectral synthesis, which
goes far beyond the results of Riemannian geometry presented here [53]. J
CHAPTER 13

Classical Applications (Survey)

Synopsis. General Appraisal. Cohomological Formulation of the Index Formula:


Orientation Class; Chern Classes; Chern Character; Todd Class; Comparison of K-
theoretic and Cohomological Thom Isomorphisms; Chern Character Defect. The Case
of Systems (Trivial Bundles). Examples of Vanishing Index. Euler Number and Signa-
ture. Vector Fields on Manifolds. Abelian Integrals and Riemann Surfaces. The Theorem
of Riemann-Roch-Hirzebruch. The Index of Elliptic Boundary-Value Problems. Real Op-
erators. The Lefschetz Fixed-Point Formula. Analysis on Symmetric Spaces. Further
Applications.
I This last chapter of our Part III contains a quick survey on classical reformulations,
applications and generalizations of the Atiyah-Singer Index Formula. Further below, in
Chapters 15 and 17, we shall elaborate on details of the fundamental geometric operators
and reproof the corresponding — now classical — integrality theorems by specific asymp-
totics of the heat kernel, as announced in Equation (12.21), p.305. In the other chapters of
Part IV, the reader will find a full length exposition of the role of index theory in quantum
field theory and low-dimensional topology.
With the Atiyah-Singer Index Formula, we proved “one of the deepest and hardest
results of mathematics” which “is probably enmeshed more widely with topology and
analysis than any other single result” [213, p. VIII]. Up till now, in the first three parts
of this book, we are mainly interested in the varied sources and parts which flow together
or are put together in the Index Formula. But the formula itself is of great interest also.
It can express important and far-reaching ideas in various areas of application, through
numerous corollaries, specializations, reformulations and generalizations.
The appraisals of the Atiyah-Singer Index Formula in these applications are contra-
dictory: On the one hand, it permits us to attack complicated topological questions with
relatively simple analytic methods; on the other hand, this formula frequently serves only
“to derive a number of wholly elementary identities, which could have been proved much
more easily by direct means.” This is the judgment of Friedrich Hirzebruch and Don
Zagier about the relationship between the Index Formula and topics from elementary
number theory. Some of the following principal theorems, which at first could be proved
only in the framework of the Atiyah-Singer theory, have been proved directly in the mean-
time. This is true for the general Riemann-Roch Theorem (see Section 13.7 below) in
[416], its consequences for the classification of certain algebraic surfaces drawn by Kuni-
hiko Kodaira in [209], and for some of the vector field computations of M.F. Atiyah
and J. Dupont (see Section 13.5) in [257].
Other results like S. Donaldson’s, P. Kronheimer’s, T. Mrowka’s, N. Seiberg’s,
C.H. Taubes’ and E. Witten’s discoveries about low-dimensional smooth manifolds must
be considered as veritable achievements and in that sense true applications of index theory.
Therefore, we shall devote two Chapters (16 on Gauge Theoretic Instantons and 18 on
Seiberg-Witten Theory) to their presentation.
Much is unclear in the relationship between the Atiyah-Singer Index Theorem and its
applications. It is important and interesting that these many relationships exist, although

310
13.1. COHOMOLOGICAL FORMULATION OF THE INDEX FORMULA 311

considerable future efforts will be required to bare the real reasons for these relationships,
and, in the end, to better understand the unity of mathematics and the specifics and
interrelationships of its parts. Once again, we quote [213, [Link]]: “That a connection
exists, a number of people realized essentially at the same time.... And since neither we
nor anybody knows why there must be such a connection, this seemed like an ideal topic for
a book in order to confront other mathematicians with a puzzle for their embarrassment
or their entertainment — as the case may be.”
Whence, regarding contemporary mathematics, this book shall guide the reader to
three roles of index theory:
(i) Index theory explains and interrelates phenomena. It opens the eyes to puzzling
interconnections.
(ii) It trains young mathematicians in global analysis, i.e., distinguishing between
local and global features and interrelating them.
(iii) By advancing new territories, it formulates new hypotheses and proves new
results.
While most of the following areas of applications — and, as emphasized, additional
ones — will be presented in more detail in Part IV, this chapter is more of an overview of
the by now classical applications and a literature survey to them. J

1. Cohomological Formulation of the Index Formula


In Theorem 12.1 (p. 285) and Theorem 12.14 (p. 294), the Index Formula is
phrased in the language of K-theory: The right sides only involve vector bundles
and operations on vector bundles. The conversion to cohomological form is carried
out in [19, p.546-559]. While the individual calculations are somewhat complicated,
the underlying method (namely, the construction of a functor that assigns to each
vector bundle a characteristic cohomology class of the base X of E with coefficients
in Z, Q, or R) is rather clear. Here, we follow [301].

Orientation Class and Thom Isomorphism. Let E be a complex vector


bundle of fiber dimension N over a paracompact space X, with projection π : E →
X, and let E̊ denote the subspace of E obtained by removing the zero section. As
a real vector bundle of fiber dimension 2N , E is oriented, since all complex bases
e1 , ..., eN of the fiber Ex , x ∈ X, yield real bases e1 , ie1 , e2 , ie2 , . . . ,eN , ieN of
the same orientation. In the language of cohomology, an orientation for Ex is the
choice of a generating element Ox for H 2N (Ex , Ex \ {0} ; Z). Now, the following
observations go back to R. Thom:
Basic Observation 1.
(i) H i (E, E̊; Z) = 0 for all i < 2N .
(ii) The orientation of E defines a total orientation class U ∈ H 2N (E, E̊; Z)

 thecondition (jx ) U = Ox for all x ∈ X, where jx : (Ex , Ex \ {0}) ,→
by
E, E̊ is the embedding.
(iii) Via the cup product u 7→ π ∗ (u) ` U , u ∈ H i (X; Z), with the orientation
class U , an isomorphism

(13.1) ΦE : H ∗ (X; Z) −→ H ∗ (E, E̊; Z)
is defined which raises the dimension of the cohomology classes by 2N .
This is the Thom isomorphism that can be given more generally for all
312 13. CLASSICAL APPLICATIONS (SURVEY)

real, oriented vector bundles of arbitrary fiber dimension. For a compar-


ison with the Bott isomorphism of K-theory (our Theorem 10.22, p.271)
and the Thom isomorphism of K-theory (our Theorem 12.8, p.290), see
[19, p.546-559] and the technical details in the proof of Theorem 13.1 be-
low.
In the following, let Hc∗ (·) denote cohomology with compact supports defined
in analogy with K-theory with compact supports, see Section 10.4 (pp.266f) and
(11.3) (p.279 in our Section 11.2). In particular, we have Hci (E; Z) = Hci (E, E̊; Z).

Chern and Todd Classes. The N -th Chern class cN (E) of E is the class
Φ−1 (U ` U ) ∈ H 2N (X; Z); the total Chern class
c(E) = 1 + c1 (E) + ... + cN (E); ci (E) ∈ H 2i (X; Z)
is obtained (compare our geometric Definition 15.61, p.444) by the axiomatic con-
ditions of
• functoriality, i.e., f ∗ c(E) = c(f ∗ E) for f : Y → X, and
• homomorphy, i.e., c(E ⊕ F ) = c(E) ` c(F ).
Because of this homomorphism property, the Chern classes are not only defined on
Vect(X), but also on K(X), since c(E) = 1 if E is a trivial bundle.
We write
c(E) = (1 + y1 ) ` · · · ` (1 + yN ) with yi ∈ H 2 (X; Z),
where the −yi are the hypothetical zeros of the polynomial 1+c1 (E)t+...+cN (E)tN
so that (by Vieta’s formula) ck (E) is the k-th elementary symmetric polynomial in
the yi . By means of that formal factorization one obtains the Chern character
XN XN 1 XN 2
ch(E) := eyi = N + yi + yi + · · ·
i=1 i=1 2! i=1

with the understanding that when ck (E) is substituted for the k-th elementary
PN
symmetric polynomial in the yi , then i=1 eyi is expressed in terms of the elemen-
tary symmetric polynomials in the ck (E). Explicitly one computes (where we write
ck (E) simply as ck )
1 2
ch0 = dim E, ch1 = c1 , ch2 = 2 c1 − c2 , ch3 = 1
6 (3c3 − 3c2 c1 + c31 )
ch4 = 1
24 (−4c4 + 4c3 c1 + 2c22 − 4c2 c1 + c41 ), . . .
2

For two complex bundles E1 and E2 over X, we have


ch(E1 ⊕ E2 ) = ch(E1 ) + ch(E2 ) and ch(E1 ⊗ E2 ) = ch(E1 ) ch(E2 ).
Then E 7→ ch(E) induces a ring homomorphism ch : K(X) → H ev (X; Q), thereby
providing a natural transformation from K-theory to singular cohomology theory
with rational coefficients and compact support.
The Todd class of E → X is defined in the analogous way by
y1 yN
Td(E) := −y
· ... · ∈ H ev (X; Q).
1−e 1 1 − e−yN
One finds for the homogeneous parts (see Section 15.7, pp.443ff)
1 1 2 1
Td0 = 1, Td1 = 2 c1 , Td2 = 12 (c2 + c1 ), Td3 = 24 c2 c1
Td4 = 1
720 (−c4 + c3 c1 + 3c22 + 4c2 c1 − c41 ), . . .
13.1. COHOMOLOGICAL FORMULATION OF THE INDEX FORMULA 313

Incidentally, by means of Riemannian geometry one can express ch(E) and


Td(E) by the curvature matrix of the vector bundle E which one equips with a
Hermitian metric; e.g., see [45, p.551] and [40, p.310], or our Section 15.7 below.
Results. We can now prove
Theorem 13.1 (Cohomological Formulation of the Index Theorem). For an
elliptic operator P : C ∞ (E) → C ∞ (F ) on a closed manifold X of dimension n,
acting between smooth sections of complex vector bundles E → X and F → X,
with symbol class [σ(P )] ∈ K(T ∗ X), we have
index(P ) = (−1)n chT X [σ(P )] Td(T X ⊗R C) [T X].
 
(13.2)
Here
• [T X] ∈ H2n (T X; Q) denotes the fundamental cycle of the tangent bundle
T X which carries an orientation as an almost-complex manifold (namely,
divide the tangent space of T X into a horizontal = real and a vertical
= imaginary part); by fixing a Riemannian metric we identify T X and
T ∗ X;
• Td(T X ⊗ C) ∈ H ev (X; Q) denotes the Todd class of the complexification
T X ⊗R C with fiber dimension n;
• chT X [σ(P )] Td(T X ⊗R C) belongs to the ring Hcev (T X; Q) with
chT X : K(T X) −→ Hcev (T X; Q)
denoting the Chern character ring homomorphism; the right side of (13.2)
is evaluation of this element on [T X]; i.e., integration over T X.
Proof. From the K-theoretic Index Formula Theorem 12.14 (p. 294), we have
index(P ) = indext [σ(P )]. Whence to prove (13.2), it suffices to show that the right
side of (13.2) coincides with indext [σ(P )]. We could have done that before proving
the K-theoretic Index Theorem. All we need is the following careful comparison
of the Thom isomorphisms of K-theory (our Theorem 12.8, p.290) and of singular
homology theory (our Equation (13.1), p.311).
Step 1. Elementary Properties of the Thom Isomorphisms. We recall: For a
complex vector bundle π : E → X of fiber dimension k, we have Thom isomorphisms
ΨE : K(X) −→ K(E) and ΦE : H ∗ (X) −→ H ∗+2k (E; Z) ∼
c = H ∗+2k (E, E̊; Z).
By Theorem 12.8, ΨE : K(X) → K(E) is given by ΨE (u) = π ∗ (u)λE = π ∗ (u)ΨE (1)
for u ∈ K(X). If iE : X → E denotes the zero section, with induced i∗E : K(E) →
K(X), then
(i∗E ◦ ΨE )(u) = i∗E π ∗ (u)ΨE (1) = i∗E ΨE (1) u
 

= i∗E (λE )u = [Λev (E)] − [Λodd (E)] u.




For 1 ∈ H 0 (X), the Thom class is ΦE (1) ∈ H 2k (E), and for π ∗ : H ∗ (X) → H ∗ (E)
induced by π : E → X, we have ΦE (u) = π ∗ (u)ΦE (1). For the zero section iE : X →
E, we have i∗E : H ∗ (E) → H ∗ (X) as before, but now for cohomology, and for
u ∈ H j (X),
(i∗E ΦE )(u) = i∗E π ∗ (u)ΦE (1) = i∗E ΦE (1) · u,
 
(13.3)
where the pull-back i∗E ΦE (1) of the Thom class is called the Euler class of E (for
an explanation of that notation see below our Sections 13.4 on Euler numbers and
13.7 on the arithmetic genus), namely χ(E) := i∗E ΦE (1) ∈ H 2k (X; Z).
314 13. CLASSICAL APPLICATIONS (SURVEY)

Step 2. The Chern Character Defect. The diagram


Ψ
K(X) −−−E−→ K(E)
 
ch ch
y X y E
Φ
H ev (X) −−−E−→ H ev (E)
does not commute in general, and this leads
 to the introduction of the Chern
−1
character defect, I(E) := ΦE chE (ΨE (1)) . We may write
chE (ΨE (1)) = ΦE (1) · π ∗ x for some x ∈ H ∗ (X).
We then have
i∗E ΦE (1) · I(E) = i∗E ΦE (1) · Φ−1
  
χ(E)I(E) = E chE ΨE (1)
= i∗E ΦE (1) · Φ−1 ∗
= i∗E ΦE (1) · x
  
E ΦE (1) · π x
= i∗E ΦE (1) · i∗E π ∗ x = i∗E ΦE (1) · π ∗ x
 

= i∗E chE ΨE (1) = chX i∗E ΨE (1) = chX [Λev (E)] − [Λodd (E)] .
  

Qk
Thus, formally and with c(E) = j=1 (1 + xj ), we then have

chX [Λev odd
C (E)] − [ΛC (E)]
Φ−1

E chE (Ψ E (1)) = I(E) =
χ(E)
P P  P P 
2p≤k exp j1 <j2 ···<j2p xji − 2q+1≤k exp j1 <j2 ···<j2q+1 xji
= Qk
j=1 xj
k
Y 1 − exj
= .
j=1
xj

Recall that
k
Y xj
Td(E) = , and so
j=1
1 − e−xj
k k
Y −xj Y xj
Td(E) = xj
= (−1) k
= (−1)k I(E)−1 ,
j=1
1 − e j=1
1 − exj

i.e., I(E)−1 = (−1)k Td(E). For ξ ∈ K(X), we have


 
(Φ−1
E ◦ ch E ◦ Ψ E )(ξ) = Φ −1
E chE Ψ E (1) chX (ξ) = I(E)chX (ξ).

Indeed,

Φ−1 = Φ−1 ∗
= Φ−1 ∗
   
E chE ΨE (ξ) E chE ΨE (1) · π ξ E chE ΨE (1) chE (π ξ)
= Φ−1 = Φ−1
 ∗  
E chE ΨE (1) π chX (ξ) E chE ΨE (1) chX (ξ) = I(E)chX (ξ).

Step 3. Involving the Embedding f : X ,→ Rn+m . For i : {p0 } → Rn+m , we have


the normal bundle N0 = Rn+m p0 → {p0 } and T N0 = T Rn+m = Cn+m → T {p0 }.
Moreover, there are Thom isomorphisms
i! = Ψ0 : K(T {p0 }) −→ K(T Rn+m ) and Φ0 : Hc∗ (T {p0 }) −→ Hc∗ (T Rn+m ).
13.1. COHOMOLOGICAL FORMULATION OF THE INDEX FORMULA 315

Let u ∈ K(Cn+m ) ∼
= Z, say u = Ψ0 (ξ), for ξ ∈ K(T {p0 }) ∼
= Z. We have
n+m −1 −1
 
ch(u)[C ] = Φ0 ch(u) = Φ0 ch(Ψ0 (ξ))
= I Cn+m ch(ξ) = 1ξ = ξ = Ψ−1

p0 0 (u).

Hence, the right-most square in the following diagram commutes:


−1
Ψ h Ψ0
K(T X) −−−−→ K(T N ) −−−−→ K(T Rn+m = Cn+m ) ∼
= Z −−−−→ K(T {p0 })
   
   
ych ych ych ych
−1
Φ h Φ0
Hc∗ (T X) −−−−→ Hc∗ (T N ) −−−−→ Hc∗ (T Rn+m = Cn+m ) ∼
= Z −−−−→ Hc∗ (T {p0 }),
e

and under the identification with Z of the rings in the right-most square the homo-
morphisms are all just identities. The middle square also commutes, when h and
h are the extension homomorphisms. The left-most square does not commute in
e
general, since there is generally a nontrivial Chern character defect I(T N ) in the
relation
Φ−1 ch Ψ(ξ) = I(T N ) ch(ξ),

ξ ∈ K(T X).
In the midst of the computation of indext (P ) below, we use the result I(T N ) =
Qk xj
I(T X ⊗ C)−1 , which is seen as follows. From the relation I(E) = j=1 1−e xj , we
have I(E1 ⊗ E2 ) = I(E1 )I(E2 ). Hence, I(T N ) = I(T X ⊗ C)−1 follows from the
fact that as complex bundles over T X, we have
(T X ⊗ C) ⊕ T N ∼ = π ∗ (T X) ⊕ π ∗ (T X) ⊕ T N ∼= T (T X) ⊕ T N = T (T Rn+m )|T X ,


which is a trivial bundle. Without further interruption, we obtain


indext (P ) = (i! )−1 f! ([σ(P )]) = (Ψ0 )−1 (h ◦ Ψ)([σ(P )])
 

= ch (h ◦ Ψ)([σ(P )]) [Cn+m ] = e h ch Ψ([σ(P )]) [Cn+m ]


 

h∗ [Cn+m ] = ch Ψ([σ(P )]) [T N ]


 
= ch Ψ([σ(P )]) e
= Φ−1 ch Ψ([σ(P )]) [T X] = I(T N )ch([σ(P )]) [T X]
 

= I(T X ⊗ C)−1 ch([σ(P )]) [T X]




= (−1)n Td(TC X ⊗ C)chT X ([σ(P )]) [T X]




= (−1)n chT X ([σ(P )]) Td(TC X ⊗ C) [T X].




Corollary 13.2. If the underlying manifold X is oriented, the index formula
(13.2) simplifies to
Φ−1
n(n+1)/2
 
(13.4) index P = (−1) T X chT X ([σ(P )]) Td(T X ⊗ C) [X] ,

where ΦT X : H ∗ (X; Q) → H ∗ (T X, T˚X; Q) = Hc∗ (T X; Q) denotes the Thom iso-


morphism and [X] ∈ Hn (X; Q) is the fundamental cycle of the orientation of X.
Proof. For a compact oriented manifold X, we have the Thom isomorphism
ΦT X : H ∗ (X) → H ∗ (T X), which has the property
Φ−1
n(n−1)/2

(13.5) u[T X] = (−1) T X (u) [X]
1
for any u ∈ H ∗ (T X). The (−1) 2 n(n−1) is explained as follows: Let x1 . . . . , xn be
positively oriented local coordinates for X. Then the induced oriented local co-
ordinates for v1 ∂x1 + · · · + vn ∂xn ∈ T X are x1 , . . . , xn , v1 , . . . , vn . However, for
316 13. CLASSICAL APPLICATIONS (SURVEY)

the orientation that we chose for T X (even if X is not-orientable), the coordi-


nates x1 , v1 , . . . , xn , vn are positively oriented. Transforming x1 , . . . , xn , v1 , . . . , vn
to x1 , v1 , . . . , xn , vn requires (n − 1) + (n − 2) + · · · + 1 = n(n −1)/2 transpo-
sitions. We use (13.5) with u := (−1)n ch([σ(P )]) Td(T X ⊗ C) , noting that
n + n(n − 1)/2 = n(n + 1)/2, to obtain (13.4). 
Exercise 13.3. Assume oriented X and nonvanishing Euler class1 χ(X) 6= 0.
Show that then there is the simpler formula
 ch ([E] − [F ]) 
n(n+1)/2 X
(13.6) index P = (−1) chT X ([σ(P )]) Td(T X ⊗ C) [X] .
χ(X)
[Hint: Recall from our introduction of the Euler class in (13.3) the following: If
i : X → T X denotes the zero section, then
χ(T X) · Φ−1 = i∗ chT X [σ(P )] = chX i∗ [σ(P )]
  
T X chT X ([σ(P )])
= chX ([E] − [F ]) · chT X ([σ(P )])].
Remark 13.4. The calculation of the right sides in the index formulas (13.2),
(13.4) and (13.6) naturally involves only the evaluation of the highest dimensional
components of the cup product on the respective fundamental cycles.

2. The Case of Systems (Trivial Bundles)


In [19, p.600-602], a drastic simplification of the Index Formula is proved for
the case of trivial vector bundles (i.e., the elliptic operator P is applied to a sys-
tem of N complex-valued functions). The symbol of such an operator is then
a continuous map σ(P ) : SX → GL(N, C), where SX denotes the unit cotan-
gent sphere bundle of X. Looking at the induced cohomology homomorphism
(σ(P ))∗ : H ∗ (GL(N, C)) → H ∗ (SX) (coefficients arbitrary), it becomes evident
that in order to obtain a useful formula in this case, one must know something
about the cohomology of the Lie group GL(N, C). Since the unitary group U(N )
is a deformation retract of GL(N, C), we have H ∗ (GL(N, C)) = H ∗ (U(N )). More-
over, since we have a natural map ρN : U(N ) → U(N )/U(N − 1) = S 2N −1 , we may
obtain all essential information on H ∗ (GL(N, C)) from the well-known cohomology
of S 2N −1 . More precisely: Let ui ∈ H 2i−1 (S 2i−1 ) denote the natural generat-
ing element, where S 2i−1 is oriented as the boundary of the ball in Ci . Then set
∗ 2N −1
hii := (ρi ) (ui ) ∈ H 2i−1 (U(i)); e.g., hN
N ∈ H (U(N )). For i ≤ N , we obtain
N 2i−1
additional elements hi ∈ H (U(N )) by means of the normalization condition
j ∗ hN i
i = hi , where j : U(i) ,→ U(N ) denotes the canonical embedding. The hi
N

(i ≤ N ) form a system of generators for the algebra H (U(N )).

Results. For an elliptic system P of N pseudo-differential equations for N


complex-valued functions on a closed manifold X of dimension n, we have
a)
X ∗ N 
N i−1 (σ(P )) hi
(13.7) index P = (−1)n (−1) ` Td(X) [SX] .
i=1 (i − 1)!
Here [SX] denotes the fundamental cycle of the canonical orientation of SX and
Td(X) denotes the lift to SX of the Todd class of the complexification of T X; see

1Commonly one writes “χ(X)” instead of “χ(T X)” for the Euler class.
13.3. EXAMPLES OF VANISHING INDEX 317

Section 13.1 above.


b) If n ≤ 3 or if X is a hypersurface in Rn+1 , then Td(X) = 1, and hence
N +n−1 (σ(P ))∗ hN

 (−1)
 (n−1)! ,
n
for N > n,
index(P ) = mapping degree (ρ◦σ(P ))
 − (n−1)! , for N = n,
0, for N < n.

For the determination of the mapping degree of the composition ρ ◦ σ(P ) : SX →


S 2n−1 , see [97, 14.9.6-10] .
c) In the framework of Hodge theory (which provides a canonical isomorphism be-
tween the harmonic differential forms of degree p on a manifold and the p-th coho-
mology of the manifold with coefficients in C – see also Section 13.4 below), there are
explicit differential forms ωi ∈ Ω2i−1 (U(N )) (so-called bi-invariant forms), which

represent the generators hN i . If π : SX → X denotes the projection, Td ∈ Ω (X)
denotes the total differential form corresponding to the Todd class (involving the
PN i−1
curvature of the Riemannian manifold X) and ω := i=1 (−1) (i−1)!
ωi
∈ Ω2i−1 (U(N ))
denotes the total bi-invariant form, then one obtains the integral formula
Z

index P = (−1) n
σ(P ) ω ∧ π ∗ Td .
SX

3. Examples of Vanishing Index


In a series of special cases one can conclude that the index of an operator van-
ishes by using solely the format of the index formula in Theorem 12.1 (p. 285) or its
generalizations and alternative formulations in Sections 13.1 and 13.2 above, with-
out having to go through all of the somewhat complicated topological computations.
For some of these results (e.g., for a), and for the special case N = 1 and n > 2
in b), see Remark 13.5), one does not need the full index formula, but rather only
the simpler theorem of Exercise 11.4 (p. 279) that the index is a homomorphism
K(T X) → Z.
Results. Let X be a closed manifold of dimension n, E, F ∈ VectN (X) and
P ∈ Ellk (E, F ). Then we have index(P ) = 0 in the following cases
a) n odd and P a differential operator.
b) N < n and X a hypersurface in Rn+1 or n ≤ 3,
c) N = n/2 and the Euler number χ(X) 6= 0,
d) N = n/2 and n not divisible by four,
e) N < n/2,
where we shall assume that the vector bundles E and F are trivial in b)-e).

Arguments. To a): We show that a) follows very nicely from the general
cohomological index formula of Theorem 13.1 (p.313): Let ant : ξ 7→ −ξ denote the
antipodal map on the tangent bundle T X. Since σ(P ) at the point x is written in
terms of a matrix of homogeneous polynomials of the k-th degree with coefficients
in C and coordinates in Tx∗ X as variables, we have the symmetry condition
(13.8) σ(P )(ant(ξ)) = (−1)k σ(P )(ξ), ξ ∈ Tx X.
Here we have identified T X and T ∗ X by means of a Riemannian metric on X. Via
multiplication by eitπ , t ∈ [0, 1], one obtains a homotopy in IsoSX (E, F ) from σ(P )
318 13. CLASSICAL APPLICATIONS (SURVEY)

to −σ(P ), hence [σ(P )] and [−σ(P )] are equal in K(T X). We can then neglect the
sign in (13.8) and obtain
(13.9) ant∗ [σ(P )] = [σ(P )] ,
if P is a differential operator. We now apply the index formula Theorem 13.1 of
Section 13.1:
indext [σ(P )] = (−1)n {ch [σ(P )] ` Td(X)} [T X]
= (−1)n {ant∗ ch [σ(P )] ` Td(X)} (ant∗ [T X])
= (−1)n {ch [σ(P )] ` Td(X)} ((−1)n [T X])
= −index P, whence index P = 0.

Here we have used (13.9) in the third equality, to obtain ant∗ ch [σ(P )] = ch [σ(P )]
in H ∗ (T X; Q). Note also that ant inverts only the vertical part of the tangent
space T X, leaving the horizontal part unchanged: in localPcoordinates (x1 , ..., xn ),
with ξ represented by (x1 , ..., xn , ξ1 , ..., ξn ) where ξ = ξi dxi , we have ant(ξ)
represented by (x1 , ..., xn , −ξ1 , ..., −ξn ). Thus, the orientation of T X is reversed by
ant, precisely when n is odd.
Incidentally, with somewhat more topology, (see [45, p.600]) one can show
directly for odd n and P a differential operator that the mappings σ(P )(ant(·))
and (σ(P )(·))−1 are stably homotopic, whence [σ(P )] + [σ(P )] = 0 and [σ(P )] is
of finite order. Then index P ∈ Z must also be of finite order and hence zero,
since index : K(T X) → Z is a homomorphism. In this way, one obtains a) without
recourse to the explicit index formula.
To b): We went over b) in Result b) of Section 13.2.
To c)-e): The derivation of c)–e) from the Result a) of Section 13.2 can be found
in [45, p.602f].
Remark 13.5. One can also directly prove e) for the special case N = 1
and n > 2 without the full index theorem (see also Exercise 9.22, p. 244, and the
literature given there, where the same result is derived topologically in a pedestrian
way). For trivial line bundles the space of elliptic symbols can be expressed very
simply: Since GL(1, C) = C× can be contracted to the circle S 1 , the index is defined
on the set of homotopy classes [SX, S 1 ] = H 1 (SX; Z). Since (by Exercise 9.21,
p. 244) the index of an elliptic operator P is zero when its symbol σ(P ) depends
only on x (and not on ξ ∈ (SX)x ), it follows that the index vanishes on the image
of π ∗ in the following long exact cohomology sequence:

... / H 1 (BX) / H 1 (SX) / H 2 (BX, SX) / H 2 (BX) / ...


9

= index
 π∗

H 1 (X) Z

(We have omitted the coefficient ring Z from the cohomology groups here.) As we
already reported in our Basic Observation 1(i) of Section 13.1 above, by a classical
result of R. Thom, H 2 (BX, SX; Z) = 0 for n > 2, whence π ∗ is surjective. Thus,
we have proved that index P = 0 for each elliptic operator P defined on the space
of complex-valued functions on a manifold of dimension > 3. Compare with [21,
13.4. EULER CHARACTERISTIC AND SIGNATURE 319

p.103 f], where similar topological arguments are needed in certain cases for n = 2
and for nontrivial bundles.

4. Euler Characteristic and Signature


I We have seen already how (e.g., in our proof of the Bott Periodicity Theorem) an-
alytic methods are utilized in topology and, conversely, how the Index Formula expresses
the analytic index by topological means. This relationship between the topology of mani-
folds and the analysis of linear elliptic operators is further revealed by the fact that certain
invariants of closed manifolds can be realized as indices of classical geometric operators
which can be defined quite naturally on these manifolds. A detailed presentation is con-
tained in our Chapter 17, expanding the classic [287]. The case of compact manifolds
with boundary remained obscure for a long time, since the classical operators do not all
admit local elliptic boundary value systems in the way of Section 9.4. In [41] the break-
through came: The famous Atiyah-Patodi-Singer (APS) Index Theorem brought together
topological, differential and spectral invariants, associated to smooth compact manifolds
with boundary, and defined by elliptic operators, see also [68], [83] and our Section 9.4
below. J

The Euler Characteristic. The invariant which we consider here is the


Euler characteristic χ(X). We concentrate on a few technical issues and re-
fer to the marvelous review [211] for the fascinating history and wider aspects.
It is defined by the observation (a matter of solid geometry and probably known
long ago to Greek mathematicians) that, for every triangulation of a closed oriented
surface X, the alternating sum χ(X) = α0 − α1 + α2 of the number of vertices (α0 ),
edges (α1 ) and faces (α2 ) is the same and depends only on the number of handles,
the genus of the surface: χ(X) = 2 − 2g. Figure 13.1 shows the beginning of a
triangulation and the surface has genus 4.

Figure 13.1. Beginning triangulation of a surface X of genus g = 4

The Euler characteristic can be interpreted as the alternating sum χ(X) =


b0 − b1 + b2 of the number 0-, 1- and 2-dimensional holes, whereby the interior of
X consists of a 0-dimensional and g 1-dimensional holes and the exterior of further
g 1-dimensional holes and the 2-dimensional total space, thus b0 = b2 = 1 and
b1 = 2g.
320 13. CLASSICAL APPLICATIONS (SURVEY)

In the language of singular homology, bi is the rank of the i-th homology group
Hi (X; Z), the i-th Betti number, and in this form the definition of the Euler
characteristic can be extended to a topological manifold X of dimension n > 2. If
the αi are again defined by a triangulation of X, then we obtain
n n
χ(X) = α0 − α1 + · · · +(−1) α2 = b0 − b1 + · · · +(−1) bn .
The Euler characteristic is among the best understood topological invariants. For
example one knows (see e.g., [11, p.309 and 358], [387, p.246] and [185, p.99-103]):
(i) χ(X × Y ) = χ(X) · χ(Y ) ,
(ii) dim X odd =⇒ χ(X) = 0,
(iii) χ S 2m = 2 and χ(Pm (C)) = m + 1.
The Signature. The concepts of Euler characteristic and genus have experi-
enced various generalizations, inspired the construction of new invariants, and have
led to many surprising results. Contrary to the Euler characteristic, some of these
new invariants are still not fully understood. All that is Index Theory and will
be dealt with in the following sections and in Part IV. In this section we shall
only explain one single additional invariant, the signature. The signature of an
oriented topological manifold X of dimension 4q is a sharper invariant than the
Euler characteristic in several aspects (see Section 13.5). It is defined as follows.
Consider the real-valued quadratic form
(13.10) Q : H 2q (X; R) −→ R, defined by Q(a, a) := (a ` a) [X] ,
where a ` a denotes the cup product and [X] denotes the fundamental cycle of the
orientation of X. Then
(13.11) sig(X) := sig(Q) := p+ − p− ,
where p+ and p− denote the maximal dimension of subspaces of H 2q (X; R) on
which Q is positive- (respectively negative-) definite.
Q is nondegenerate, whence p+ + p− = b2q (recall bi = dim Hi (X; R) =
dim H i (X; R)). Furthermore, the Poincaré duality Hi (X; R) ∼
= H 4q−i (X; R) im-
plies χ(X) ≡ b2q mod 2, and we obtain the formula
(iv) dim X = 4q =⇒ χ(X) ≡ sig(X) mod 2.
Results. We now want to describe these two invariants analytically, when the
oriented, closed manifold X, with dim X = n even, is equipped with a differen-
tiable structure and a Riemannian metric. Referring to Exercise 6.20 (p. 172) and
our discussion of theL cobordism based proof of the Index Theorem on pp.302ff
n
above, let Ω• (X) := j=1 Ωj (X) denote the space of (complexified) exterior dif-
ferential forms, with exterior derivative d : Ω• (X) → Ω• (X) and its formal adjoint
δ : Ω• (X) → Ω• (X).
Theorem 13.6. a) The operator d+δ : Ω• (X) → Ω• (X) is an elliptic, formally
self-adjoint differential operator of order 1, whence index(d + δ) = 0. Occasionally,
it is called the DeRham-Dirac operator.
b) By restricting d + δ to the space Ωev (X) := j
L
jeven Ω (X) of even forms and
setting Ωodd (X) := j
L
jodd Ω (X), we obtain the Euler operator. It is an elliptic
differential operator of order 1
(13.12) (d + δ)ev : Ωev (X) → Ωodd (X) with index((d + δ)ev ) = χ(X).
13.4. EULER CHARACTERISTIC AND SIGNATURE 321

c) Let n ≡ 0 mod 4, say n = 4q, and let τ : Ω• (X) → Ω• (X) denote the involution
(defined by (12.17), p. 302) with the ±1-eigenspaces Ω± (X). By restricting d + δ to
the space Ω+ (X), we obtain the signature operator. It is an elliptic differential
operator of order 1
(13.13) (d + δ)+ : Ω+ (X) → Ω− (X) with index(d + δ)+ = sig(X).
 
d) As defined in Section 13.1, let U ∈ H n T X, T̊ X denote the orientation class
 
and Φ : H n (X) → H 2n T X, T̊ X the Thom isomorphism. Recall that there is a
characteristic class χ(T X) = Φ−1 (U ` U ) ∈ H n (X; Z), namely the Euler class of
the tangent bundle T X, for which χ(T
 X) [X] = χ(X). The class χ(T X) is repre-
sented by a certain n-form GB Ωθ , the Gauss-Bonnet
 form, which is defined
in terms of the curvature Ωθ of X; i.e., GB Ωθ is a multiple of the Pfaffian of Ωθ
(to be explained below in Section 15.6 in detail on pp.446ff ). Then,
Z
GB Ωθ ,

(13.14) χ(X) =
X
and this is known as the Chern-Gauss-Bonnet formula.
e) If n = 4q, the signature sig(X) can also be expressed as
Z
(13.15) sig(X) = Lq (p1 , . . . , pq ) [X] = Lq (e
p1 , . . . , peq ) ,
X
which is known as the Hirzebruch Signature Formula with Hirzebruch’s L
polynomial in Pontryagin’s characteristic classes.
Remark 13.7. The preceding theorem should not be considered an application
of the Index Theorem. Much more it is its origin and motivation: Strictly speaking
a)-c) only show that the Euler characteristic and the signature can be written as the
index of geometrically defined operators. Assertions d)-e) then recall two formulas
that we shall prove in our Part IV in unusual detail, applying a suitable version
of the Index Theorem. The formulas were known before the Index Theorem and
originally proved independently.
Arguments.
To a: To Exercise 6.47a (pp.187ff) we gave an extended hint how to show that
(13.16) σ(d)(x, ξ) = iξ∧ and σ(δ)(x, ξ) = −iξx , for (x, ξ) ∈ T̊ ∗ X,
and how to deduce ellipticity of d + δ. Recall from Remark 6.21 and Equation
(6.9) (p.173) that ξ∧ denotes exterior and ξx interior multiplication of forms with
the cotangent vector ξ ∈ Tx∗ X \ {0}. Two calculations were left to the reader and
shall be brought now in detail. For that, with the notation of Exercise 6.20 (p.172),
let ∗j : Λj (X) → Λn−j (X) denote the Hodge star operator, characterized by the
property (∗j α) ∧ β = hα, βi νg , where νg denotes the volume form on X induced
by a fixed Riemannian metric g. We use the same notation for the corresponding
Hodge star operator on forms.
(i) While it is easy to see that σ(d)(x, ξ) = iξ∧ , we need a little calculation to find
σ(δ)(x, ξ). By definition, δ = d∗ . Then σ(δ)(x, ξ) must be the adjoint of σ(d)(x, ξ).
All we have to do is to confirm that −iξx is the adjoint mapping to iξ∧ . Note
322 13. CLASSICAL APPLICATIONS (SURVEY)

that −iξx = −i ∗ ξ∧ ∗ which is better manageable, like our reformulation of the


codifferential δ in (6.32). Thus, we have for ω ∈ Ωj (X), α ∈ Ωj+1 (X)
j
hξ ∧ ω, αi νg = (ξ ∧ ω) ∧ ∗α = (−1) ω ∧(ξ ∧ ∗α)
 
j (n−j)j
(13.17) = (−1) ω ∧ (−1) ∗ ∗ (ξ ∧ ∗α)
nk
= (−1) hω, ∗(ξ ∧ ∗α)i νg = hω, ∗(ξ ∧ ∗α)i νg ,
since n is even. Here we used that
ω ∈ Ωj (X), α ∈ Ωj+1 (X) =⇒ ∗α ∈ Ωn−(j+1) (X) and ξ ∧ ∗α ∈ Ωn−j (X).

(ii) To show the ellipticity of d + δ, we recall that the square of d + δ is the (Hodge)
Laplace operator
∆ := (d + δ)2 = dδ + δd : Ω• (X) −→ Ω• (X),

which is of order 2 and homogeneous in the forms (i.e., ∆ Ωj (X) ⊆ Ωj (X)). The
ellipticity of d + δ, (as well as that of ∆) follows, once it is shown that
 2
σ(d + δ)(x, ξ) ◦ σ(d + δ)(x, ξ) ω = kξk ω,
−1 −2
since then σ(d + δ)(x, ξ) = kξk σ(d + δ)(x, ξ). We have

σ(d + δ)(x, ξ) ◦ σ(d + δ)(x, ξ) ω
 
= iξ ∧ iξ ∧ ω − i ∗ (ξ ∧ ∗ω) − i ∗ ξ ∧ ∗(iξ ∧ ω − i ∗ (ξ ∧ ∗ω))
 
= ξ ∧ ∗(ξ ∧ ∗ω) + ∗ ξ ∧ ∗ξ ∧ ω − ∗ ∗ (ξ ∧ ∗ω)
 
= ξ ∧ ∗(ξ ∧ ∗ω) + ∗ (ξ ∧ ∗ξ ∧ ω − ξ ∧ ∗ ∗ (ξ ∧ ∗ω))
 
= ξ ∧ ∗(ξ ∧ ∗ω) + ∗ ξ ∧ ∗(ξ ∧ ω) ± ξ ∧ ξ ∧ ∗ω
  2
= ξ ∧ ∗(ξ ∧ ∗ω) + ∗ ξ ∧ ∗(ξ ∧ ω) = kξk ω,
where it suffices to check the last equality for kξk = 1 in the case ω ∈ Λj ξ ⊥ ,


and in the case ω = ξ ∧ η where η ∈ Λj−1 ξ ⊥ , both of which are straightforward


(extend ξ to an oriented, orthonormal basis).
To b: Since d+δ is formally self-adjoint, its index is zero. By a simple orthogonality
argument, like in our Remark 2.11 (p.17), Ker(d+δ) = Ker ∆ follows. By definition,
elements of that space are known as harmonic forms. Hodge-de-Rham theory tells
us that the algebra of harmonic forms, say H• (X), is isomorphic to the cohomology
algebra H ∗ (X; C), where wedge product of harmonic forms corresponds to cup
product in H ∗ (X; C). In particular, for ∆j := ∆|Ωj (X) , we have the main theorem
of Hodge theory
(13.18) Ker (d + δ) |Ω (X) = Ker ∆j =: Hj (X) ∼ = H j (X; C), j = 0, . . . , n.

j

Moreover, there is a splitting of Ωj (X) into L2 orthogonal direct sums


Ωj (X) = Ker ∆j ⊕ Im ∆j−2 = Ker ∆j ⊕ Im dj−1 ⊕ Im δj+1 .
For the details of that Hodge-DeRham Decomposition, see Theorem 17.62, p. 603,
and Corollary 17.63. If we restrict d+δ to Ωev (X), we obtain a differential operator
(13.19) (d + δ)ev := (d + δ)|Ωev (X) : Ωev (X) −→ Ωodd (X) with formal adjoint
∗
(13.20) (d + δ)ev = (d + δ)odd := (d + δ)|Ωodd (X) .
13.4. EULER CHARACTERISTIC AND SIGNATURE 323

The restricted operators (d + δ)ev and (d + δ)odd are still elliptic with symbols that
are inverses modulo a factor of −kξk2 . The index of (d + δ)ev is not necessarily
zero. Indeed,
index((d + δ)even ) = dim Ker(d + δ)ev − dim Ker(d + δ)ev
X X
= dim Ker ∆j − dim Ker ∆j
j even j odd
X X
j
= dim H (X; C) − dim H j (X; C) = χ(X),
j even j odd

the Euler characteristic of X.


To c: To begin with, we explain the claimed decomposition Ω(X) = Ω+ (X) ⊕
Ω− (X). To be on the safe side, we argue on the level of the vector bundle Λj (X) →
X of complex exterior j-covectors over the compact, orientable C ∞ Riemannian
n-manifold X with metric tensor g. Until further notice, we assumeLn only that n
is even, say n = 2m. Since ∗2m−j ◦ ∗j = (−1)j , the total ∗ := j=0 ∗j is not an
involution. However, as shown in Exercise 6.20c and Equations (6.8a) and (6.8b),
L2m
setting τj := im+j(j−1) ∗j yields an involution τ := j=0 τj (τ 2 = 1 := IdΛ• (X) ).
Thus, Λ• (X) = Λ+ (X) ⊕ Λ− (X), where
Λ+ (X) := (1 + τ )Λ• (X) and Λ− (X) := (1 − τ )Λ• (X)
are the ±1 eigenbundles of τ (of positive fiber dimension); we set Ω± (X) :=
C ∞ (Λ± (X)). Using δ = −∗d∗ (for n even), one can check that (d+δ)τ = −τ (d+δ)
so that(d + δ) (Ω± (X)) = Ω∓ (X). Whence, the Hirzebruch signature operator
(d + δ)+ := (d + δ)|Ω+ (X) : Ω+ (X) → Ω− (X)
is well defined with formal adjoint (d + δ)− := (d + δ)|Ω− (X) . Since ∆ and τ
commute, it follows that
Mm
(13.21a) Ker(d + δ)+ = (1 + τ ) Ker ∆ = (1 + τj ) Hj and
j=0
m
∗ M
(13.21b) Ker(d + δ)− = Ker (d + δ)+ = (1 − τj ) Ker ∆ = (1 − τ ) Hj .
j=0

For j < m (strict), the maps


(1 ± τj ) : Hj (X) −→ Hj (X) ⊕ τj Hj (X) = Hj (X) ⊕ H2m−j (X)
are injections, and so for any j < m, (1 − τj ) Hj =∼ Ker ∆j ∼= (1 + τj ) Hj . Conse-
quently, in (13.21a) and (13.21b), the spaces (1 ± τj )Hj (X) have pairwise equal
dimensions and only when j = m do we have a contribution to the index:
index(d + δ)+ = dim Ker(d + d∗ )+ − dim Ker(d + δ)−
= dim((1 + τm ) Hm ) − dim((1 − τm ) Hm ) .
Note that
(
m+m(m−1) m2 ±i ∗m , for m odd,
τm = i ∗m = i ∗m =
∗m , for m even.
It follows that index(d + δ)+ = 0 for m odd (i.e., n ≡ 2 mod 4). Consider m
even, so that τm = ∗m . Let Hm (X)R denote the space of R-valued m-forms α
324 13. CLASSICAL APPLICATIONS (SURVEY)

with (d + δ)ga = 0. By the afore mentioned Hodge theory, we can translate the
cohomological quadratic (intersection) form of (13.10) into a quadratic form
Z Z
Q : H(X)R −→ R given by Q(α) := α∧α = hα, ∗m αig νg .
X X

Note that Q is positive-definite on (1+τm )Hm (X)R on which ∗m is Id and negative-


definite on (1 − τm )Hm (X)R . Thus, for n ≡ 0 mod 4, we have proved

index(d + δ)+ = sig(Q) = p+ − p− ,

i.e., (13.13) holds.


To d: The formula (13.14) is proved directly (i.e., without reference to U , Φ, or the
fact that Φ−1 (U ` U ) [X] = χ(T X) [X] = χ(X)) using the Local Index Theorem
in Part IV (Theorem 17.68, p. 612) applied to certain components of the operator
1
(d+δ)even . When n = 2, the formula reduces to the classical result χ(X) = 2π
R
X
K
of C. F. Gauss and O. Bonnet, where K denotes the Gaussian curvature of the
closed surface X.
To e: The Hirzebruch Signature Formula will be proved in Chapter 17 in full detail
(Theorem 17.64, pp.601–605). The proof will be based on the Index Theorem for
Generalized Dirac Operators (Theorem 17.59, proved via heat kernel asymptotics,
pp.595–600), and our determination of the Hirzebruch L-Polynomials (Equation
(15.111), explained and derived on pp.447–451).
Here we only sketch the basic concepts: The pj are the Pontryagin characteristic
j
classes which are defined in terms of Chern classes via pj := (−1) c2j (T X ⊗ C) and
the polynomial Lq (p1 , . . . , pq ) (in which the pj are multiplied via cup product) is
described as follows. We have an expansion

x1 xq X
··· = Lk (σ1 , . . . , σq ) ,
tanh x1 tanh xq =1

where σ1 , . . . , σq denote the elementary symmetric polynomials in x21 , . . . , x2q and


Lk (σ1 , . . . , σq ) is ultimately homogeneous of degree k in x21 , . . . , x2q . We then replace
σ1 , . . . , σq in Lq (σ1 , . . . , σq ) by p1 , . . . , pq to obtain Lq (p1 , . . . , pq ), see (15.111),
p. 451. In particular,
(13.22) 
 13 p1 ,
 for n = 4,
+ 1

index(d + δ) = sig(Q) = L(X) = 45 7p2 − p21 , for n = 8,
 1
 3
945 (62p3 − 13p2 p1 − 2p1 ), for n = 12.

One can represent pk by a 4k-form Rpek which involves the curvature tensor of X,
whence, as an example, sig X = 31 X pe1 for dim X = 4. The formula for pek in
terms of the curvature tensor of X is found in Section 15.7, specifically (15.102),
p. 447, where pej is denoted there by pk Ωθ to indicate its dependence on the
curvature form Ωθ of the Levi-Civita connection θ. In Lq (e
p1 , . . . , peq ) the forms pek
are multiplied via wedge product. For a statement and proof (using the Local Index
Theorem) of the twisted generalization of the Hirzebruch Signature Formula, see
Theorem 17.65, p. 606.
13.5. VECTOR FIELDS ON MANIFOLDS 325

Remark 13.8. We shall give three arguments for the significance of the signa-
ture and the signature formula.
a) The result (13.22) explains the once mysterious fact that p1 is divisible by 3.
This is just one in a series of divisibility results which we shall address in our Chap-
ters 17 and 18, like V. Rokhlin’s Corollary 17.56 and Theorem 18.14. See also the
divisibility results in the next section on vector fields and characteristic classes.
b) As explained above, the genus g or, equivalently, the Euler characteristic χ of a
closed oriented surface completely characterizes the homeomorphism type which for
surfaces agrees with the diffeomorphism type. In our Chapter 18 we shall explain
mutually complementary results by Simon Donaldson and Michael Freedman
which combined yield an analogous result, namely that two closed differentiable
simply-connected 4-manifolds X1 and X2 are homeomorphic if and only if
χ(X1 ) = χ(X2 ), sig(X1 ) = sig(X2 ), and type(Q1 ) = type(Q2 ),
where we shall distinguish between even and odd type of the intersection form.
c) Another beautiful proof of the significance of the signature comes from a simple
production of explicit examples of exotic spheres and at the same time the con-
struction of a topological manifold without smooth structure. Our Corollary 18.15
gives the existence of such manifolds. Now we summarize an explicit construction
that does the job. Roughly speaking (for details we refer to [211]), we construct a
manifold W (E8 ) of dimension 12 with boundary ∂ W (E8 ) =: Σ by gluing together
a copy of the disc bundle of the tangent bundle of S 6 for each edge of the graph
of E8 . Recall that E8 denotes the unique unimodular rank 8 even quadratic form
of signature 8, defined in (18.8), p.646. Such constructions are known under the
heading plumbing, a technique peculiarly mastered by John Milnor and Egbert
Brieskorn in the classic [295, 91], take also a quick glance at [1]. The Mayer-
Vietoris sequence implies that Σ is a homotopy sphere. Application of the Poincaré
conjecture for dimensions larger than 4, proved by Smale in [401], shows that Σ is
homeomorphic to S 11 . Then, as Hirzebruch and Kreck notice in [211], Σ is not

diffeomorphic to S 11 . Indeed, if there were a diffeomorphism f : Σ −→ S 11 , then
we could glue a 12-dimensional disc onto M along Σ and obtain a smooth manifold
M := W (E8 ) ∪f D12 , whose homology is trivial except in degree 0, 6, and 12. The
intersection form of this manifold is by construction the E8 -form, whose signature
is 8. Now we apply the signature theorem and obtain 8 = 62/945 (p3 (M ), [M ]), a
contradiction (note that the only potentially nontrivial Pontryagin class is p3 (M ),
an integral cohomology class). Thus Σ is an exotic sphere. If instead of a diffeo-
morphism we use a homeomorphism, we obtain a topological manifold M , which
by the same argument cannot admit a smooth structure.

5. Vector Fields on Manifolds


I Among the basic concepts of the analysis of a dynamical system, as well as of the
geometry of a differential manifold X, is the notion of a vector field. It is a C ∞ section v of
the tangent bundle (in classical terminology a variety of line elements), thus v ∈ C ∞ (T X).
Each v defines an ordinary differential equation ċ(t) = v(c(t)) for differentiable paths
(trajectories) c : R → X. According to the classical existence and uniqueness theorems,
the equation has a unique solution (at least defined on some open interval about 0) for any
given initial value c(0) = x ∈ X; for details see [97, p.74-87] . In every theory of flows the
singularities, i.e., the points belonging to the singular set Σv := {x ∈ X : v(x) = 0} are
of special interest, as they are the stagnation points of dynamics, the equilibrium points,
326 13. CLASSICAL APPLICATIONS (SURVEY)

Figure 13.2. Smooth vector field on X with singularity at x

or the so-called stationary solutions (the constant paths c(t) = x, see Figure 13.2). For
applications, one may think intuitively of examples from oceanography, of a magnetic field,
or of dynamics in economics. We saw above in Section 10.1 and Figure 10.1 that every
isolated singularity x of a vector field v defines locally a map S n−1 → S n−1 , n = dim X,
whose mapping degree we denote by Iv (x). The geometric interest in vector fields usually
stems from the question of parallelizability of the manifold, see Figure 13.3: Is it possible
to assign to a tangent vector v ∈ Tx X at a point x a tangent vector v 0 ∈ Tx0 X parallel to
v in a way which is independent of the path from x to x0 used in the process?

v
x
x0
v0

Figure 13.3. Parallel move of v ∈ Tx X to v 0 ∈ Tx0 X, here being


independent of the choice of the connecting path

This question is important for the physical notion of space and finally for the analysis
of motion since the concepts of force and acceleration depend on parallel displacement, as
indicated in our Section 6.5 on Connections and further developed in our Chapters 14 and
15 on Physical Motivation and Geometric Preliminaries. The question of parallelizability
is equivalent to the question of whether an n-dimensional manifold X possesses n vector
fields that are linearly independent at each point x ∈ X; i.e., whether the tangent bundle
T X is trivial. For example, by a famous theorem of J. F. Adams, S n−1 is parallelizable
if and only if Rn is a division algebra, i.e., for n = 2m and m = 1, 2, 4.
In the preceding Section 13.4, we saw that index and principal symbol of the Euler
operator (d + δ)ev and the signature operator (d + δ)+ are related to some characteristic
classes: the Euler class χ(X) and the Pontryagin classes {pj }. Such interrelations between
the separated worlds of analysis and topology are not so surprising. Both analysis and
topology have to do with vector fields. In this section, we shall explain that aspect. J
13.5. VECTOR FIELDS ON MANIFOLDS 327

Vector Fields, Linking Elliptic Operators and Characteristic Classes.


The geometrical idea and also the historical origin (with Eduard Stiefel, 1935,
and almost simultaneously and in a similar connection with Hassler Whitney)
of the topological invariants of Riemannian or complex manifolds (nowadays called
characteristic classes — see Sections 13.1–13.5 above and Section 15.7 below) lies
in the general investigation of r vector fields v1 , ..., vr on a manifold and their
singular set Σv1 ,...,vr consisting of those points x, where v1 (x) , ..., vr (x) are linearly
dependent. One knows, for example, that set Σv1 ,...,vr generically has dimension
r − 1, and that the cycle Σv1 ,...,vr defines a homology class (with suitable coefficient
group) which is a characteristic class of X and is independent of v1 , ..., vr ; see
also the large survey [414].
The index formula (see Section 13.1 above) says that the symbol of an elliptic
operator is a kind of characteristic class, and its index is a characteristic integer.
From this arose the program to illuminate the connection between elliptic operators
and vector fields on manifolds. Specifically, the existence of a certain number of
vector fields implies certain symmetry properties for classical operators and corre-
sponding results for their indices such as Euler number, signature and characteristic
numbers.
Results. Let X be a closed, oriented, Riemannian manifold.
a) Each vector field v generically has only finitely many zeros (i.e., by an arbitrarily
small
P perturbation a vector field can be put in this form), and we have χ(X) =
x∈Σv Iv (x), where Iv (x) denotes the local index of v at x defined above. Thus,
the right side of the formula consists of the finite weighted sum of zeros of v.
b) If dim X = 4q and X has a tangent field of 2-dimensional planes (i.e., an
oriented 2-dimensional subbundle of the tangent bundle), then χ(X) ≡ 0 mod 2
and sig(X) ≡ χ(X) mod 4.
c) If the 4q-dimensional manifold X has at least r vector fields that are independent
everywhere, then sig(X) ≡ 0 mod βr , where the values of βr are given in the
following table (βr+8 = 16βr ):
r 1 2 3 4 5 6 7 8
βr 2 4 8 16 16 16 16 32 .
Arguments.
To a): For dim X = 2, this is a classical result of H. Poincaré, for which one
can find a very clear sketch of the proof in [93, p.166-171]. The generalization for
dim X > 2 originated from H. Hopf, who also showed that χ(X) = 0, if and only
if there is a nowhere zero vector field on X.
To b) and c): We refer to [287], [22] and [37], where a series of related results are
proved by combining results of analysis, topology, and algebra, and in some cases
using the theory of real elliptic operators (see Section 13.9 below).
To illustrate the methods, we will only treat the first half of Hopf’s Theorem
(13.23) ∃v ∈ C ∞ (T X) with v(x) 6= 0 for all x ∈ X =⇒ χ(X) = 0.
See also the student blog [286, Vector Fields II] which explains b) in simple words.
One can obtain (13.23) from the explicit Atiyah-Singer Index Formula or the
Gauss-Bonnet Formula which says (see Theorem 13.6d) that
(13.24) χ(X) = χ(T X)[X], with χ(T X) := Φ−1 (U ` U ) = ṽ ∗ i∗ U,
where
328 13. CLASSICAL APPLICATIONS (SURVEY)

χ(X): Euler characteristic of the topological manifold X;


χ(T X): Euler form of the tangent bundle of X, now considered as differen-
tiable;
U ∈ H n (BX, SX; Z): orientation class of T X;

Φ : H j (X; Z) −→ H j+n (BX, SX; Z): Thom isomorphism;
i : (BX, ∅) ,→ (BX, SX): the trivial embedding; and
ṽ : X → BX: normalized vector field ṽ = v/ |v| with |v| =
6 0 by assumption.
Since iṽ(X) ⊂ SX, we have ṽ ∗ i∗ = 0. Whence, (13.24) proofs (13.23).
To get the same result, without recourse to the explicit index formula, one can
apply the general theory of elliptic operators (see Chapter 9 above) and standard
results of Hodge theory, reproduced in Theorem 13.6b,c. The key issue for this is
the fact that

σ(d + δ)(x, ξ)w = iξ × w for x ∈ X, ξ ∈ Tx∗ X, and w ∈ Λ• (Tx∗ X).

Here for V := Tx∗ X,

× : Λ• (V ) × Λ• (V ) −→ Λ• (V ) denotes the Clifford multiplication

given (in particular) for ξ ∈ Tx∗ X and w ∈ Λ• (Tx∗ X) by

ξ × w := ξ ∧ w − ξxw and w × ξ := ξ ∧ w − wyξ,

where x and y denote left and right interior multiplication. Note the difference
between Clifford multiplication ξ× and exterior multiplication ξ∧ . Note also ξxw =
∗(ξ ∧ ∗w) when n is even, so that indeed σ(d + δ)(x, ξ)w = iξ × w by (13.16), p. 321.
A general treatment of Clifford algebras is provided in Section 17.1.
By means of the Riemannian metric on X, a vector field v can be regarded as
a 1-form which yields a 0-th order differential operator Rv : Ω• (X) → Ω• (X) given
by Rv (u) := u × v, u ∈ Ω• (X) (i.e., via pointwise right Clifford multiplication by
2
v). Since (u × v) × v = u ×(v × v) = − |v| u, our assumption |v| > 0 implies that

(13.25) Rv : Ω• (X) automorphism with Ωev/odd (X) −→ Ωodd/ev (X).

Since

Rv ◦ σ(d + δ)(x, ξ) = w(iξ × w) × v(x)


= iξ ×(w × v(x)) = σ(d + δ) ◦ Rv (x, ξ),

σ(d + δ) commutes with σ(Rv ) = Rv . Let Rvodd := Rv |Ωodd(X) . Then by the Main
Theorem 8.25 of Symbolic Calculus,
−1
(13.26) Rvodd ◦ (d + δ)ev ◦ Rvodd − (d + δ)odd ∈ OP0 .

Using ((d + δ)ev ) = (d + d∗ )odd , we then have
(13.25)
 −1 
index(d + δ)ev = index Rvodd (d + δ)ev ◦ Rvodd
(13.26)
= index(d + δ)odd = − index(d + δ)ev ,

whence χ(X) = index(d + δ)ev = 0.


13.6. ABELIAN INTEGRALS AND RIEMANN SURFACES 329

I(x/a
) I(x)

x
x

Figure 13.4. Arc lengths, left ellipse, right lemniscate

6. Abelian Integrals and Riemann Surfaces


I Perhaps one of the first (in our topological sense) quantitative results of analysis is
contained in the major work of the Norwegian mathematician Niels Henrik Abel entitled
Mémoire sur une propriété générale d’une classe très étendue de fonctions transcendantes,
[2]. It was written in 1826, but published only posthumously in 1841. In it Abel takes up
a dispute of the numerical analysis of the 18th century about rectifiability, the possibility of
solving integrals by means of elementary functions: algebraic functions, circular functions,
logarithm and exponential functions.
Note . We leave it to the historians of mathematics to discern the different moti-
vations that drew rectifiability research. For the long run to solving algebraic equations
by radicals, such a study was conducted by Neslihan Saglanmak and Uffe Jankvist
in [360]. They point to the world view of cryptologists, atomists and alchemists (Viète,
Leibniz, Tschirnhaus) driving the research, while they refute needs of numerical analysis
as motivation for these investigations. Already in [437] A.N. Whitehead ridiculed the
perception of teachers of mathematics that learning the solution formula of second order
algebraic equations had anything to do with true practical applications. Similar atomistic
and alchemical ideas may be found behind the rectification program. A quote from Leib-
niz can be read either way: arguing for technology, astronomy, biology oriented research
in applications and/or arguing for a philosophical world view, remote from day-to-day
usefulness. He drew the attention of the mathematicians to constructions and curves
“quas natura ipsa simplici et expedito motu producere potest” (which nature itself can
produce by simple and complete motions), quoted after A. von Brill’ and M. Noether’s
comprehensive report [96, p.124]. See also [247, p.411f].
Since Jakob Bernoulli and Gottfried Wilhelm Leibniz, there was an interest
particularly in the integration of irrational functions, that turn up in many problems of
science and technology. The efforts (already of the 17th century) to rectify the ellipse (see
Figure 13.4), whose arc length is important for astronomy, lead to the computation of the
integral, for semi-axes a, b:
Z x
1 − k 2 t2 b2
I(x) = a p dt, k = 1 − 2 .
0 (1 − t2 )(1 − k2 t2 ) a
While investigating the deformation of an elastic rod under the influence of forces
acting at its ends, Jakob Bernoulli (1694) came across further irrational integrands. In
this connection he also introduced the lemniscate
n p p o
(± (x2 + x4 )/2, ± (x2 − x4 )/2) : 0 ≤ x ≤ 1 ,
whose arc length is given by Z x
dt
I(x) = √ .
0 1 − t4
330 13. CLASSICAL APPLICATIONS (SURVEY)

For this and the following, see [392, 1.1-1.2]. Leonhard Euler proved (1753) the addition
theorem for the lemniscate integral :
√ √
u 1 − v 4 + v 1 − u4
I(u) + I(v) = I(w) where w := .
1 + u2 v 2
It generalized a result of the Italian mathematician Giulio Carlo de’ Toschi di Fag-
nano (1714) on the doubling of the circumference of the lemniscate with compass and
ruler alone:
4u2 1 − u4

2
2I(u) = I(w) for w = .
(1 + u4 )2
Note. At that time, it was suspected already that the rectification even of elementary
curves was impossible in general. But only in 1835 did J. Liouville prove rigorously
that the elliptic integrals could not be solved elementarily, i.e., expressed by a finite
combination of algebraic, circular, logarithmic and exponential functions. As an aside,
it is now known that the problem of the elementary integration of arbitrary functions is
recursively undecidable.
Prior to di Fagnano, Johann Bernoulli (1698) accidentally discovered that the
difference of two arcs of the cubic parabola (y = x3 ) is integrable by elementary functions
and added to the problem of rectifying curves the new — and solvable — problem of finding
arcs of parabolas, ellipses, hyperbolas etc. whose sum or difference is an elementary
quantity – “just as arcs of a circle could be compared with one another through the
expressions for sin(α + β), sin 2α, etc.” [96, p.206].
A little later Euler succeeded in extending his addition theorem for the leminiscatic
integral to elliptic integrals of the first kind, i.e., in proving that
p p
u P (v) + v P (u)
(13.27) I(u) + I(v) = I(w) with w = , where
1 + u2 v 2
Z u
dt
I(u) := p and P (t) := 1 + at2 − t4 .
0 P (t)
Based on a comparison of elliptic arcs also due to Fagnano, Euler finally found another
generalization to elliptic integrals of higher kind. These are integrals of the form
Z u
r(t)
I(u) := p dt,
0 P (t)
where r is a rational function in one variable and P is a polynomial of third or fourth
degree with simple zeros. Here the addition theorem takes the form
(13.28) I(u) + I(v) = I(w) + W (u, v),
where w is, as above, an algebraic function of the arbitrary upper integration limits u and
v and W (u, v) = S1 (u, v) + log S2 (u, v) with rational functions S1 , S2 .
Euler already noticed that his methods cannot be used for the treatment of hyper -
elliptic integrals (with a polynomial P of fifth or higher degree). But only Abel found
the explanation for “the difficulties which Euler’s formulation necessarily encountered
when dealing with hyperelliptic integrals: The constant of integration (I(w) in (13.27) or
(13.28)), appearing in the transcendental equation, could not be replaced, as in the elliptic
case, by a single integral but only by two or more hyperelliptic integrals – a remarkable
circumstance which was in no way predictable... The question about the minimal number
of integrals which a given sum of integrals could be reduced, remained as the cardinal
question; it caused Abel to produce the elaborate and laborious counts which constitute
the main results of his great Paris paper and which brought him into the possession of the
notion of genus of an algebraic structure long before Riemann”, [96, p.211 f]. J
13.6. ABELIAN INTEGRALS AND RIEMANN SURFACES 331

Abelian Addition Theorem. Let R be a rational function and F a polyno-


mial, both in two variables. For a, x ∈ R (a fixed), consider the Abelian integral
Z x
I(x) := R(t, y)dt, where y satisfies the equation F (t, y) = 0.
a

For F (t, y) = y 2 − P (t) with P as above and R(t, y) = 1/y, we obtain an elliptic
integral of the first kind ; and for R(t, y) = r(t)/y with r a rational function, we get
an elliptic integral of a higher kind.
For a given value of t, there may be several corresponding solutions (roots) of
the equation F (t, y) = 0. Thus, one must specify which root will be substituted
for y in R(t, y). Hence, one selects an integration path γ : I → C in the plane
with <γ(0) = a, <γ(1) = x, and F (<γ, =γ) = 0. Regard I(x) as a line integral.
Then Abel proved in [2] that the sum of m arbitrary Abelian integrals I(xi ) (m
sufficiently large), with respective fixed integration paths (xi ∈ R, i = 1, ..., m)
can be written as the sum of only g Abelian integrals I(x̂j ) (j = 1, ..., g) and a
remainder term W (x1 , ..., xm ) which is the sum of a rational function S1 and the
log of another rational function S2 of the limits of integration x1 , ..., xm :
Xm Xg
(13.29) I(xi ) = I(x̂j ) + W (x1 , ..., xm ),
i=1 j=1

where the x̂j = x̂j (x1 , ..., xm ) are algebraic functions and the number g only depends
on the specific form of the polynomial F .
According to C.G.J. Jacobi, Abel attempted to solve two problems with his
addition theorem, the representability of an integral by closed expressions, and the
investigation of general properties of integrals of algebraic functions [96, p.205]. In
fact, for one thing, the addition theorem says something about the rectification
question: If g = 0 then I(x) = W (x) for m = 1, where W is constructed from
rational functions and the logarithm. Thus, in this case, the integral I(x) can be
solved by elementary functions. If g > 0 then in general at least g additional higher
transcendental functions are needed, namely the I(x̂j ). Most of all, the theorem
is an addition theorem just like the sum formulas for trigonometric functions, and
can be used in the back-up files of computers as interpolation formulas for realizing
standard functions of science and technology (see the good old [117]).
We quoted Abel’s theorem, in order to point out one of the earliest occurrences
of the fundamental invariant g which, as the index, has the dual character of being
both analytic and algebraic. Without entering a discussion of the function theoretic
aspects and geometric interpretations of Abel’s addition theorem, we will note
some results exhibiting the significance of the quantity g. These essentially are due
to Bernhard Riemann (except c), see the review [211].
Results.
a) Each polynomial F (t, y) defines a compact Riemann surface, an oriented sur-
face with a complex-analytic structure (i.e., a topological manifold with a distin-
guished atlas whose coordinate changes are holomorphic functions). Conversely,
one can define a complex-analytic structure on each compact, oriented topological
surface X such that X can be regarded as the Riemann surface of a polynomial
F (t, y).
b) Topologically, a compact Riemann surface X is characterized by its genus g,
i.e., the number of handles that must be fastened to the sphere in order to obtain
X, see Figure 13.5. Twice g is the number of closed curves needed to generate
332 13. CLASSICAL APPLICATIONS (SURVEY)

the first homology of X; i.e., there are γ1 , . . . , γ2g closed curves (namely, along the
lengths and girths of the handles, see Figure 13.6) such that P every closed curve γ
in X is homologous to a unique integral linear combination ni γi .

g=2

Figure 13.5. Compact Riemann surface of genus g = 2

Analytically (see the Abelian Addition Theorem of (13.29)) X has another in-
variant, the maximal number g1 of linearly independent holomorphic differential
forms α1 , ..., αg1 of degree 1Ron X. Actually, g = g1 ; i.e., the numerical complex-
ity of the Abelian integral R(t, y)dt with F (t, y) = 0 is equal to the genus of the
Riemann surface F . In particular, for the elliptic integrals, one obtains an elliptic
curve, namely the torus of genus 1 with two generating cycles γ1 and γ2 .
Riemann (1857) calls this quantity the Klassenzahl (class number). The term
genus originated with Alfred Clebsch (1864). For the connection with the classi-
cal notion of double periodicity in elliptic integrals, see [311, p.149-155], for example.

°2 °1

Figure 13.6. Two closed curves generate the homology of the torus
R
More generally, one can define a periodicity matrix ωij := γi αj for i =
1, ..., 2g, j = 1, . . . , g, which determines the complex-analytic structure of X by a
theorem of R. Torelli.
c) On each Riemann surface X of genus g there is a Cauchy-Riemann operator
∂¯ : f 7→ ∂f
∂ z̄ dz̄ which assigns to each complex-valued C

function f on X a complex
deferential form of degree 1. ∂ is an elliptic operator and index ∂¯ = 1 − g.
¯

Arguments.
To a) and b): All that follows from the classical theory of Riemann surfaces.
To c): Note that dim Ker ∂¯ = 1, since Ker ∂¯ consists of the global holomorphic
functions and these must be constant by the maximum principle. Coker ∂¯ = Ker ∂¯∗
13.7. THE THEOREM OF HIRZEBRUCH-RIEMANN-ROCH 333

consists of the (anti-) holomorphic differential forms which constitute a vector


space isomorphic to the space of holomorphic 1-forms.
One can also directly obtain c) from the Atiyah-Singer Index Formula, since
the Euler class χ(T X) is the only Chern class of a Riemann surface, whence for
some constant C

index ∂¯ = Cχ(T X)[X] = Cχ(X) = C(1 − 2g + 1) = C(2 − 2g).

It is then not difficult to derive that C = 1/2; see also the following Section 13.7,
Results a), c) and d), which include our present Result c) as a special case. See
also [311, p.132-141] where c) is proven, with reference to these generalizations,
in the form arithmetic genus = geometrical genus by elementary geometrical and
algebraic tools of classical projective geometry.

7. The Theorem of Hirzebruch-Riemann-Roch

I We introduce next a class of theorems for which the Atiyah-Singer Index Formula
yields new proofs or generalizations. We shall give the ideas here and postpone the precise
definitions and the fully detailed proofs to Section 17.6 with a 28 pages subsection solely
devoted to the Hirzebruch-Riemann-Roch Theorem (pp.615–642).
According to [96, p.280 f], who introduced the term Riemann-Roch Theorem, it deals
with the “counting of the constants of an algebraic function”, and more generally, with
establishing relations between the constants (number of singularities, degree, order, genus
etc.) of an algebraic curve, algebraic surface or complex manifold. Thus the history and
motivation of Riemann-Roch — detailed in [96] — are closely tied to the unsolved problem
of completely classifying algebraic varieties, see [209] and the recent review [211].

Note . a) For a function-theoretic interpretation of Riemann-Roch, see also [392,


4.6-4.7], where Riemann-Roch is used to prove Abel’s theorem and is expressly called
algebraic in contrast to the transcendental nature of Abel’s Theorem.
b) Alfred Clebsch, “who surely did not lack knowledge and versatility” wrote very
openly in a letter of August 1864 to Gustav Roch that he himself “understood very little
of Riemann’s treatise even after greatest efforts, and that Roch’s dissertation remained
for the most part incomprehensible to him.” (Quoted after [96, p.320].) We may add that
the historical process of correctly understanding the theorem apparently has not reached
its conclusion at least as far as algebraic functions are concerned.

We cannot convey the abundance of results in this subject. They are still too scattered
and the diversity of approaches too uncertain. We therefore restrict ourselves to a few
stages where, out of the complexity of the preceding problems, some unifying, very rich,
and consequential aspects developed little by little:
(i) Bernhard Riemann’s transcendental idea of the analysis on Riemann surfaces
(complex curves), i.e., his attempt to consider the totality of integrals of a fixed
algebraic function field (just the Abelian integrals).
(ii) The function theoretic treatment of the problems by Karl Weierstrass.
(iii) The interpretation from the point of view of differential geometry in the language
of Hodge theory due to Kunihiko Kodaira which was the basis for Friedrich
Hirzebruch’s generalization of Riemann-Roch to higher dimensional algebraic
varieties.
That makes the stages essential for us. J
334 13. CLASSICAL APPLICATIONS (SURVEY)

Dolbeault Operator and Arithmetic Genus. From Chapter 5 (proof of


the Hellwig-Vekua Theorem 5.11) and Chapter 6 (Table 6.1 and Exercise 6.47.3)
we recall the basic properties of the Cauchy-Riemann operator
 
∂ ∂ ∂
∂z = = 12 +i = 21 (∂x + i∂y ) : C ∞ (R2 , C) −→ C ∞ (R2 , C).
∂z ∂x ∂y
It is a first-order elliptic operator with principal symbol
∂ 1
σ( )(x, ξ1 dx1 + ξ2 dx2 ) = i(ξ1 + iξ2 ).
∂ z̄ 2
In the cited Exercise 6.47.3, we introduced the Dolbeault operator

∂¯ : Ω0,0 (R2 , C) −→ Ω0,1 (R2 , C), given by ∂(f¯ ) := ∂f dz̄,


∂ z̄
where Ω0,0 (R2 , C) := C ∞ (R2 , C) and Ω0,1 (R2 , C) := {hdz̄ : h ∈ C ∞ (R2 , C)}. We
mentioned its strictly analogous operator ∂¯ on compact Riemann surfaces.
More generally, one can consider the Dolbeault complex
∂¯ ∂¯ ∂¯
0 −→ Ω0,0 −→ Ω0,1 −→ · · · −→ Ω0,n −→ 0
for a Kähler manifold
P X of (complex) dimension n (i.e., a Pcomplex manifold with
Hermitian metric gik dzi dz̄k , whose associated 2-form α = gik dzi ∧dz̄k is closed,
i.e., dα = 0). Here Ω0,p denotes the space ofP complex exterior differential forms of
degree p, which can be written in the form (i) ai1 ···ip dz̄i1 ∧ · · · ∧ dz̄ip relative to
local coordinates z1 , . . . , zn . As in Exercise 6.20b (p. 172), ∂¯ denotes the exterior
derivative. If V is a holomorphic vector bundle over X, then via tensoring (as
above with the signature operator) one can construct the generalized Dolbeault
complex
∂¯ ∂¯ ∂¯
0 −→ Ω0,0 (V ) −→
V
Ω0,1 (V ) −→
V V
· · · −→ Ω0,n (V ) −→ 0,
where Ω0,0 (V ) = C ∞ (V ). Using the fact that V is holomorphic, one verifies that
∂¯V2 = 0, and thus one can define the Dolbeault cohomology groups
Ker ∂¯V |Ω0,q(V )

q
H (OV ) := ¯ .
∂V (Ω0,q−1 (V ))
Definition 13.9. a) In analogy with the above definition of the Euler charac-
teristic, one defines the holomorphic Euler characteristic of V as
Mn q
χhol (X, V ) := (−1) dim H q (OV ) .
q=0

b) If V is the trivial line bundle CX , then χhol (X) := χhol (X, CX ) is called the
arithmetic genus of X.
We refer to the charming review [211] by F. Hirzebruch and M. Kreck for
the fascinating history of different definitions of the arithmetic genus of a projective
smooth algebraic variety, mixed with informative personal recollections by FH, the
main actor in that field in the second half of the 20th century.
As with Hodge theory in the real category
H q (OV ) =∼ Hq (V ) := Ker ∂¯V + ∂¯∗ |Ω0,q(V ) = Ker  ¯ V |Ω0,q(V ) ,
  
V

¯ V := ∂¯V + ∂¯∗ 2 = ∂¯V ∂¯∗ + ∂¯∗ ∂¯V , so that



where  V V V
 
∗ even
M M
¯ ¯ Ω0,j (V ) −→ Ω0,j (V ) .

χhol (X, V ) = index ∂V + ∂V :
j even j odd
13.7. THE THEOREM OF HIRZEBRUCH-RIEMANN-ROCH 335

Results.
a) On a compact Riemann surface X of genus g, we consider meromorphic functions
w : X → C which have poles at the points xi ∈ X (i = 1, ..., r) of order at most
mi ∈ N and zeros at the points xj ∈ X (j = r + 1, ..., s) of order at least −mj ∈ N.
These form a complex vector space L(D) whose dimension l(D) is given by the
following formula
l(D) − l0 (D) = deg(D) − g + 1,
Ps Ps
where D := j=1 mj xj (formal sum), deg(D) := j=1 mj and l0 (D) denotes the

x2

x1 xs
X

Figure 13.7. Specifying a divisor on a Riemann surface X

dimension of the vector space L0 (D) of meromorphic differential 1-forms on X with


the corresponding behavior on the zeros and poles.
b) For deg(D) ≥ 2g − 1, the number l0 (D) vanishes, whence one obtains a proper
formula for l(D) in this case.
c) For each formal integral linear combination D of points of X, called a divisor, see
Figure 13.7, there is a holomorphic vector bundle {D} of complex fiber dimension 1
such that L(D) is isomorphic to the vector space Ker ∂¯{D} of holomorphic sections
of {D}, and L0 (D) is isomorphic to the vector space Coker ∂¯{D} ∼ = Ker ∂¯{D}

of
(anti-) holomorphic 1-forms with coefficients in {D}. Here,
∂¯{D} : Ω0,0 ({D}) −→ Ω0,1 ({D})
denotes the elliptic differential operator obtained from ∂¯ by tensoring, like in the
preceding definition of the generalized Dolbeault complex. With this construction,
we can rewrite Result a) in the form index ∂¯{D} = deg(D) − g + 1.
d) For a Kähler manifold X and a holomorphic vector bundle V over X, the
Hirzebruch-Riemann-Roch Theorem states that
(13.30) χhol (X, V ) = (ch(V ) ` Td(T X))[X].
Here ch(V ) denotes the Chern character of V and Td(T X) denotes the Todd class
of X; see Section 13.1 above.
If V is the trivial line bundle CX , then our (13.30) states that the arithmetic
genus is the same as the evaluation of the Todd class on the fundamental class,
called the Todd genus Td(T X)[X] which in turn can be expressed in terms of
p
Chern numbers. Indeed, let us denote by cpi the evaluation ci (T X) [X] of the
p-th power (with i · p = dimC X) of the Chern class ci (T X) on the fundamental
336 13. CLASSICAL APPLICATIONS (SURVEY)

class [X] ∈ H2n (X; Z) (cpi is called a Chern number). For the first four Todd
polynomials one then computes the table

dimC X 1 2 3 4
1 1 1 1
 
χhol (X) 2 c1 12 c2 + c21 24 c1 c2 720 −c4 + c3 c1 + 3c22 + 4c2 c21 − c41 .

Arguments. a) follows using c) and d) from the Atiyah-Singer Index For-


mula; a direct proof can be found in [435, p.117-119]. If D = 0 (so, in particular,
deg(D) = 0), then one recovers Results b) and c) of Section 13.6. Notice that
l0 (D) = l(KX − D), where KX is a divisor which canonically (up to a certain linear
equivalence, i.e., up to the divisor of a rational function) is defined by the zeros
and poles of an arbitrary nontrivial form on X. One has deg(KX ) = 2g − 2, For
details of the definition of the canonical divisor, and for an elementary proof of
l(D) − l(KX − D) = deg(D) − g + 1, see [311, p.104-107 and 145-147]; the duality
l0 (D) = l(KX − D) originated with Richard Dedekind and Heinrich Weber,
and is a special case of a very general, duality theorem of Jean-Pierre Serre
(see [207, 15.4]).
For b): A nontrivial meromorphic function on a compact Riemann surface has
as many poles as zeros (counting orders). Thus, l(D) = 0 if deg(D) < 0. Since
deg(KX ) = 2g − 2, it follows that deg(KX − D) < 0, for deg(D) ≥ 2g − 1.
For c): For the construction of the line bundle {D}, we cover X with a finite col-
lection of open sets {Uj }j∈J such that on each Uj there is defined a meromorphic
function f which has the zero and pole behavior that is the opposite of the portion
of D involving points of Uj ; so that fi /fj is a nowhere-zero holomorphic function
on Ui ∩ Uj . For example, one can choose Uj so small that at most one point (of the
finitely many points x1 , ..., xs found in D) lies in Uj , and then take fj to be a lo-
cally defined rational monomial which has a zero of order mk if mk > 0, and a pole
of order −mk if mk < 0. The clutching construction (Exercise B.9, p. 718, of the
Appendix) then yields the bundle {D} for which {fi } defines a global meromorphic
section f with zero behavior opposite D. Via g 7→ g/f , an isomorphism is defined
from the global holomorphic sections of {D} to the vector space of meromorphic
functions on X with the zero and pole behavior prescribed by the divisor D. Details
of this construction, which is rather typical for the topological approach to analytic
problems of function theory and algebraic geometry, can be found in [207, 15.2].
For d): The derivation from the Atiyah-Singer Index Formula can be found in Sec-
tion 17.6 (beginning p. 615); see also [328, p.324 ff] and [373], and in a somewhat
more general context (see Section 13.11 below) also in [45, p.563-565]. This transi-
tion from the classical Riemann-Roch Theorem to complex manifolds of arbitrary
dimension was initiated by Friedrich Hirzebruch (1953). His proof depended
essentially on the additional condition that X can be holomorphically embedded
in a complex projective space of a suitable dimension. By going back to the index
formula for elliptic operators, one can drop this restriction. Our proof in Section
17.6 is a rather differential geometric proof and ultimately depends on the heat
equation approach to the Local Index Theorem for Dirac operators. For the rather
easily calculable case of an elliptic curve (n = 1, g = 1), we refer to [33, p.311 f].
For a sketch of the general case see [loc. cit., p. 317 f].
13.8. THE INDEX OF ELLIPTIC BOUNDARY-VALUE PROBLEMS 337

Remark 13.10. The Hirzebruch-Riemann-Roch Theorem appears here as a


special application of the Atiyah-Singer Index Formula for elliptic operators of
which one, namely ∂¯V + ∂¯V∗ , was constructed taking advantage of the complex
structure. Even so, Result d) and already Result a) are the quintessential models
for the structure of the Index Formula: On the left side we have the difference of
globally defined quantities each of which can change even under small variations
of the initial data (as the dimension l(D) in a)), while on the right we have an
expression in terms of topological invariants of the problem. In a) it is the genus g
of the Riemannian surface X and the degree of the divisor D.
In the end, for complex manifolds, the general index formula and the Hirzebruch-
Riemann-Rochev Theorem (the special index formula for elliptic operators D̄V :=
∂¯V + ∂¯V∗ ) are equivalent: Let
ev
D̄ = ∂¯ + ∂¯∗ : Ω0,ev (Cn ) −→ Ω0,odd (Cn )

denote the Riemann-Roch operator for X = Cn . Then σ D̄ produces a generator
n−1
of the homotopy
 group π2n−1 (GL(N, C)), N = 2 , and all of K(T X) is generated
by the class σ D̄V modulo the image of K(X); see [33, p.321 f].
A survey of function theoretic and geometric applications (e.g., estimates for
the dimensions of systems of curves or differential forms and determination of Betti
numbers of complicated manifolds) can be found
• for the case of curves, (dimC X = 1) in [311, p.147 ff],
• for the classification of surfaces (dimC X = 2) in [250] including a multi-
tude of very concrete geometric facts which are derived directly from the
general Hirzebruch-Riemann-Roch Theorem (for line bundles, the preced-
ing Result d)), and
• for the general case in [373].
• For the general historic background of algebraic geometry see [121].

8. The Index of Elliptic Boundary-Value Problems

I We will now have a look at the index problem for boundary-value problems which
was solved in the twenties, forties and fifties of the last century for a number of special
cases in [324, 202], [425, p.316-330], [65, p.15-18] and the sources stated there (see also
our Theorem 5.11, p. 146). In the programmatic article of [159], it was called “description
of linear elliptic equations and their boundary problems in topological terms”, and this
initiated the search for the index formula for elliptic operators on closed manifolds [43].
From 1963 on, we witness the rapid development of elliptic topology, with varying forms
and proofs of the index formula and their widespread applications. During all this time,
manifolds without boundary remain in the center, although in 1964, one year after the
proof of the Atiyah-Singer Index Theorem for closed manifolds, [27] demonstrated the
topological significance of elliptic boundary conditions, and showed how, in principle, an
index formula for elliptic boundary value problems may be obtained by reducing it to the
unbordered case (Poisson principle). The essence of their message is reproduced by the
index calculation of a selected boundary value problem in our Section 9.4 (pp.244–250).
Moreover, connecting with the vanishing formulas of Section 13.3, in this fashion one can
find very quick proofs (and supplements) for the boundary-value problems with vanishing
index listed in [7]. Therefore, we may ask: Why then the slow development and the
relatively peripheral place of boundary value problems in index theory?
338 13. CLASSICAL APPLICATIONS (SURVEY)

Conceptual Challenges of Index Theory on Manifolds with Boundary. In


index theory, a topologist may expect a much more dominant place of manifolds with
boundary and boundary value problems: A topologist’s perception of space builds on local
charts, on triangulation, cutting and pasting, plumbing, surgery, connected sums. Why
then do we meet ignorance, reservations, time delay regarding a full scale development
and uptake of index theory on manifolds with boundary?
Immediately, we can point to the following explanations:
(i) For years an all too narrow class of boundary value problems has been prominent
in main-stream analysis.
(ii) It was difficult to resist the temptation of elegance and homotopy invariance.
(iii) One had to overcome the common (and sometimes ridiculed) perception of gen-
eralization by reduction to the well-understood case.
(iv) One should put aside the understandable fright of expected complexity.
Shortly, we shall explain what we mean by these traditions, temptations, difficulties,
and frights that had to be overcome.
To (i): For the Laplacian, for more than a century, main stream analysis had investigated
Dirichlet, Neumann and oblique angular boundary conditions (see also our Exercise 2.58h,
p.49f in dimension 1, and Figure 5.10, p.154 regarding dimension 2). In the early 1950s,
Zorya Yakolevna Šapiro (well-known also for her work in integral geometry together
with I.M. Gelfand) in [362] and Yaroslav Borisovich Lopatinskiı̆ in [282] (see also
the late obituary [181]) succeeded in finding the essential conditions for ellipticity of
local boundary conditions on the symbol level, i.e., necessary and sufficient conditions for
smoothness of the solutions and Fredholm property (finite index).
With hindsight it seems to us that for two decades their success blocked research in
geometrically perhaps more meaningful global boundary conditions (later to be defined by
pseudo-differential projections, see below). Certainly, Atiyah, Bott and Singer realized
very soon that, for instance, the signature operator D+ := (d + δ)+ does not admit local
elliptic boundary conditions in the sense of Šapiro and Lopatinskiı̆ (contrary to the
case studied in our Section 9.4) on a bordered manifold X. In spite of their important
role for closed manifolds in the various proofs and applications of the index formula, such
operators seemed to be out of the realm of index theory when passing to manifolds with
boundary.
Ironically, there were a complementary direction of research in the Soviet Union of
the 1950s: the systematic study of all closed extensions of various classes of operators in
Hilbert space. Main promoters were M.G. Kreı̆n, M.I. Višhik and M.Š. Birman, see the
review [12]. However, their research was focused on self-adjoint extensions, while striving
to widest generality and so also missing the point of index theory of geometrical operators
at that time.
To (ii): In 1971, L. Boutet de Monvel published a paper [87] of incredible elegance
and clarity where he assigned to local elliptic boundary value problems families of Wiener-
Hopf operators on the boundary and so were able to assemble all topological and analytical
ingredients to a clear presentation of the index problem for these types of boundary con-
ditions, ending with a beautiful and in many cases calculable index formula. Again with
hindsight we can see that his success was founded on ingeniously exploiting symbolic cal-
culus and homotopy invariance — and therein also found its limitations: Three years later,
N. Hitchin showed in [215] that, in modern terminology, under small variation of the
Riemannian metric the spectral projection of an operator of Dirac type can jump from
one connected component (of the Grassmannian of pseudo-differential projections with
the same principal symbol) into another component. Whence the index of global elliptic
(= well posed, see below) boundary problems for Dirac type operators is not homotopy in-
variant and can jump under slightest change of the defining connection, i.e., small change
of the coefficients. Still one year later, in the famous triple paper [41], M.F. Atiyah,
13.8. THE INDEX OF ELLIPTIC BOUNDARY-VALUE PROBLEMS 339

V.K. Patodi and I.M. Singer discovered that geometrically defined operators admit in
an extended sense certain globally elliptic boundary conditions. In the case of the before
cited signature operator D+ , they defined a boundary condition via a spectral projection
P≥ which canonically belongs to D+ and for which index D+ P≥ is well-defined and equal
to sig(X) up to dim Ker ∂, where ∂ denotes the tangential Dirac operator to D+ (see our
explanation of the signature deficiency below following (13.62), p.355). In particular, they
obtained the new formula
Z
sig(X) = Lq (p1 , . . . , pq ) − η(0)
X

(regarding η(0), recall the explanations around (12.23), p. 308) for the signature of a 4q-
dimensional compact, oriented, bordered Riemannian manifold. This was the end of an
index theory solely built upon homotopy invariance: On the left side of the preceding
formula we have a topological invariant. Indeed, sig(X) is homotopy invariant. On the
far right side we have a spectral invariant and further to the middle with the integral
a differential invariant. None of these two are homotopy invariant. Moreover, applying
the Atiyah-Patodi-Singer Index Theorem on odd-dimensional manifolds yields a kernel
dimension as index and so shows at once that geometrically meaningful boundary value
problems do not obey homotopy invariance (see below).
To (iii): For the transition from the index formula on closed manifolds to index problems
on manifolds with boundary, it is natural to look for the same or similar constructions as
those successful on closed manifolds. And they are there; there are three closed manifolds
immediately associated to a smooth compact manifold X with boundary: the closed
boundary Y := ∂X; the closed double X e := (−X) ∪Y X; and the closed ∂(BX) =
SX ∪SX|Y BX|Y , i.e., the closing of the cotangent sphere bundle SX over the boundary.
Reducing a boundary value problem to a problem over the boundary is a good classi-
cal method of numerical analysis, gathered under the heading boundary integral methods.
They are/were mandatory for reducing combinatorial complexity and for gaining numer-
ical stability. That approach is also viable for index theory, as we shall show below.
However, the induced index problem over the boundary is of very different nature, and
the recognition of geometric information is not easy.
Expanding a boundary value problem to a problem over the closed double is also an
option, by reading boundary conditions as transmission or coupling data. Also here it is
difficult to preserve geometric information.
Finally, reading boundary conditions as a recipe for the continuation of the principal
symbol over the closed sphere bundle works perfectly and preserves all available geometric
information in a perfect way, as shown in [87], though only for the small class of local
boundary value problems.
A radically new approach was needed for a truly geometric approach to the index
problem on manifolds with boundary. That approach was found in [41] by applying the
heat kernel asymptotics procedures, i.e., roughly speaking, by adding a separate dimen-
sion to obtain a parabolic problem and then to apply the Levi-Carleman techniques (or
the Duhamel method in [83]).
To (iv): Perhaps the most severe misapprehension would be to overestimate the technical
difficulties of obtaining valuable index formulas for manifolds with boundary. It is true
that, for instance for gluing two manifolds with boundary together to a partitioned man-
ifold, additivity formulas for Euler characteristic and signature are much easier to obtain
by classical topological methods like the Mayer-Vietoris sequence of cohomology than
by applying the Index Theorem for manifolds with boundary. That is possible as well
though, see [83, Chapter 25], and leads to a deeper result, namely determining the error
term in the general gluing situation, thus yielding the necessary and sufficient conditions
for precise additivity formulas and in that way characterizing Euler characteristic and
signature by their cutting-and-pasting invariance. Moreover, the physicists M. Ninomiya
340 13. CLASSICAL APPLICATIONS (SURVEY)

and C.-I. Tan have shown in [322] that the terms of the Index Theorem for the four-ball
can be easily found and give new insight in particle physics, see below. J

Corresponding to the geometric aims of this book, we don’t strive for greatest
generality but restrict our discussion of the index theory of boundary value problems
to Dirac type operators. That facilitates the presentation of the main ideas. For
more general elliptic operators the reader will find the necessary modifications (and
caveats) in [76, 80, 83, 99, 104, 190, 354, 379, 382, 383].

Elliptic Boundary Problems for Dirac Operators — Our Data. Let X


be a compact smooth oriented n-dimensional Riemannian manifold with boundary
Y , and let S → X be a bundle of Clifford modules with compatible Hermitian
structure and connection (covariant derivative) ∇S . Recall from Definition 6.34
that the total (elliptic) Dirac operator
D : C ∞ (X; S) −→ C ∞ (X; S)
is obtained by suitably composing a (compatible) connection ∇S : C ∞ (X; S) →
C ∞ (X; T ∗ X ⊗S) with the Clifford multiplication c : C ∞ (X; T M ⊗S) → C ∞ (X; S).
Strictly speaking, we deal with operators of Dirac type since we shall not assume
a spin structure.
As emphasized in Exercise 6.47.5, D is an elliptic differential operator. It is
formally self–adjoint and Green’s formula holds for all sections of S (often called
spinors) s and s0 :
Z
(13.31) (Ds, s0 ) − (s, Ds0 ) = − hJ(y)(s|Y ), s0 |Y i,
Y
where J := c(n) : S|Y → S|Y denotes the unitary bundle isomorphism given by
Clifford multiplication by the inward unit tangent vector. To prove (13.31) one
applies Stokes’ Theorem in a similar way as in Exercise 6.43. All details are written
down in [83, Proposition 3.4].
Until further notice, we assume that X is an even–dimensional manifold. Fol-
lowing physics terminology regarding Dirac’s original equation with n = 4 (14.30),
p.382, we denote by γ5 (the “γ5 matrix”) the global section of Hom(S, S) defined
locally by
γ5 := c(e1 ) . . . c(en ),
where {eµ } is any positively oriented orthonormal local frame of tangent vectors
and n the dimension of the manifold X. Since n is even, S splits into subbundles
S ± spanned by the eigensections of γ5 corresponding to the eigenvalue ±1, if n is di-
visible by 4, or ±i otherwise; over Y , the Clifford multiplication J switches between
S ± |Y and S ∓ |Y ; over the whole of X the Dirac operator splits correspondingly into
components
0 D−
 
D=
D+ 0
such that the right half Dirac operator D+ : C ∞ (X; S + ) → C ∞ (X; S − ) is for-
mally adjoint to the left half Dirac operator D− : C ∞ (X; S − ) → C ∞ (X; S + ).
To simplify the exposition we assume that the Riemannian metric and the
Hermitian structure are product near the boundary. Let us point out that the
results presented here remain true also for non–product structures. Admitting
non–product structures, however, makes the analysis more complicated especially
13.8. THE INDEX OF ELLIPTIC BOUNDARY-VALUE PROBLEMS 341

when one wants to discuss the ζ regularized determinant and related asymptotic
expansions.
Close to the boundary, say in a collared neighborhood N = [0, 1] × Y of Y in
X, the total Dirac operator splits into the following product form
0 −J −1
   
6∂ 0
(13.32) D|N = J (∂r + B)|N = ∂r + ,
J 0 0 −J 6 ∂ J −1 N

where r denotes the inward oriented normal (radial) coordinate in a collar neighbor-
hood of the boundary and 6 ∂ : C ∞ (Y ; S + ) → C ∞ (Y ; S + ) is the tangential Dirac
operator over the boundary. Notice that the radial coordinate does not enter in
J or B, and in fact
J 2 = −1 and J B = −BJ
as required by the formal self–adjointness of D.
We also discuss the operator D+ alone. It has the following form on the collar
N
(13.33) D+ |N = J (∂r + 6 ∂) |N .
The Dirac operator on an odd–dimensional manifold has the same form J (∂r + B)
on the collar. In this case the total operator D does not split, but the tangential
operator is a Dirac operator on an even–dimensional manifold and has therefore
the following form:
0 B−
 
(13.34) B= .
B+ 0
To give an example, we address the physics situation of quantum chromody-
namics, as we did in [82, Section 3].
Example 13.11. As manifold X we take a ‘volume’ V in R4 . We think of V
as a ball of large radius R. Actually, we are interested in the asymptotic situation
with R → ∞. As the bundle of Clifford modules we take
V × (S ⊗ C2 ) = S ⊗ C2 → V ,

(13.35)
the Clifford bundle of euclidean spinors with coefficients in the trivial bundle V ×C2
with Clifford action c(a) ⊗ 1. As the full Dirac operator D we take a twisted Dirac
operator defined by a connection A for V × C2 which is pure gauge on the boundary
of V .
To make the Example comprehensible, we have to explain the terms euclidean
spinors, connection, and pure gauge at the boundary (for details see [83, Chapter
5]).
Euclidean Dirac operator : As we shall discuss more closely below on pp.382ff in
relation to Dirac’s original (non-euclidean) equation, the (free) euclidean Dirac
operator
!

0 − ∂q
6D = ∂ : C ∞ (V ; S) −→ C ∞ (V ; S)
∂q 0
is canonically defined over R4 with
∂ ∂ ∂ ∂ ∂
(13.36) =i +j +k + ,
∂q ∂x1 ∂x2 ∂x3 ∂x4
342 13. CLASSICAL APPLICATIONS (SURVEY)

and
∂ ∂ ∂ ∂ ∂
(13.37) = −i −j −k + ,
∂q ∂x1 ∂x2 ∂x3 ∂x4
where the bundle S of euclidean spinors in 4 dimensions splits into a pair of quater-
nions S = V × (H ⊕ H) with Clifford multiplication c : C`4 → HomC (S, S) given
by the four complex 4 × 4 matrices
   
0 σµ 0 −1
c(eµ ) = γµ = for µ = 1, 2, 3 and c(e4 ) = γ4 =
σµ 0 1 0
with {σµ } denoting the Pauli matrices of (14.32) and {e1 , . . . e4 } a basis of R4 .
The connection defining the euclidean Dirac operator is just the standard connec-
tion d for S.
Twisted Dirac operator : Then any connection A for the trivial bundle V × C2
defines in a natural way a twisted Dirac operator 6 DA = 6 D ⊗A IdC2 . It is
characterized by the property
6 DA (s ⊗ f )(x) = (6 D(s) ⊗ f )(x)
whenever (Af )(x) = 0. It is a true (total) Dirac operator with regard to the induced
Clifford multiplication c(a) ⊗ 1 and the induced connection d ⊗ A.
Pure gauge at boundary: Now we must discuss the choice of the connection
A. From a physical point of view it does not suffice to consider only the trivial
choice,Pnamely the P standard connection d in V × C2 given by exterior differenti-
ation µ αµ eµ 7→ µ αµ ∂µ . Roughly speaking, the standard connection would
correspond to the description of two non–interacting fermions. To change that
we introduce a smooth family h of SU(2) matrices parametrized over ∂V ; this is
equivalent to introducing a smooth connection ∇ over the whole ball which is pure
gauge at the boundary with regard to h; or, equivalently, we introduce a vector–
valued field {Aµ } which is pure gauge at the boundary with regard to h providing
∂f
∇eµ f = ∂x µ
+ Aµ f − f Aµ for any f ∈ C ∞ (V ; V × C2 ) and eµ the µth basis vector
in R4 .
The geometrical idea behind demanding the interaction boundary
P term to be-
have as pure gauge is to get a non–trivial curvature form Ω∇ = µ<ν Fµν dxµ ∧ dxν
R 2
corresponding to an action Fµν dx < ∞ with
Fµν := [∇eµ , ∇eν ] = ∇eµ ∇eν − ∇ eν ∇eµ
= ∂µ Aν − ∂ν Aµ + [Aµ , Aν ] ∈ L2
and P Fµν (x) → 0 for |x| → ∞. This is equivalent to writing the connection ∇ =
d + µ Aµ with Aµ (x) → h∂µ h−1 for |x| → ∞.
The different mathematical descriptions are formalized in the following defini-
tion.
Definition 13.12. Let C(V ×C2 ) denote the affine space of smooth connections
for the bundle V × C2 . A connection A ∈ C(V × C2 ) is called pure gauge at the
boundary ∂V = S 3 , if in a collar [0, ε) × ∂V :
• A does not depend on the normal (radial) coordinate r; and
• there exists a smooth h : ∂V ∼ S 3 → SU(2) ∼ S 3 such that A = h◦d◦h−1
in the collar, where d denotes the standard connection given by exterior
differentiation.
13.8. THE INDEX OF ELLIPTIC BOUNDARY-VALUE PROBLEMS 343

˚ × C2 ) for the subspace of connections which are pure gauge at


We write C(V
the boundary. We shall return to this example below in Result f). For a thorough
presentation of gauge-theoretic physics see our Part IV.

Wanted Properties on Manifolds with Boundary. Now we shall discuss


the properties of a Dirac operator over a compact manifold with boundary. In the
beginning we shall not distinguish between the cases of even– and odd–dimensional
manifolds and whether we treat the total or the half Dirac operator. So, let
A ∈ {D, D+ } with A : C ∞ (X; E) → C ∞ (X; F ), E, F ∈ {S, S ± },
and product form, correspondingly with Γ ∈ {J , J}
(13.38) A = Γ(∂r + B) near the boundary Y .
Contrary to the case of a closed manifold, the space
(13.39) H(A, ∞) = {s ∈ C ∞ (X; E) : As = 0}
of harmonics of the operator A is an infinite–dimensional subspace of C ∞ (X; E).
There is also a question of regularity of the solutions. Let s denote a weak so-
lution of As = 0, which is an element of the space L2 (X; E) (or more generally
of W k (X; E) the k–th Sobolev space). On manifolds with boundary, it does not
follow that s is a smooth section of S. This leads us to the following general defi-
nition of the ellipticity in analogy to the basic properties of elliptic operators over
closed manifolds proved in Theorem 9.10a-c above (p.240): As with the Sturm-
Liouville problems in Section 2.5 in dimension 1 and within the general functional
analysis framework of closed extensions, as explained in Definition 2.38, differen-
tial operators have a maximal domain in L2 (X; E). To get a well-posed problem
(made precise in the following Definition 13.13), we have to restrict the domain.
One canonical candidate is to restrict the domain to the space C0∞ (X \ Y ; E) of
smooth sections with support in the interior of X \ Y or to its closure W01 (X; E)
in the first Sobolev space W 1 (X; E). Then we are back in Hilbert space analy-
sis and A|W01 (X;E) is the minimal closed extension of A|C0∞ (X\Y ;E) . But there are
many other reasonable extensions (also called realizations) AR to be chosen by
specifying a domain R ⊂ L2 (X; E).
Definition 13.13 (Well-posedness). Let A be a total or a half Dirac operator
on X and let AR be a closed extension of A in L2 (X, S) with domain R. We call
AR an elliptic (or well-posed) boundary problem for the operator A, if and
only if the following two conditions are satisfied:
(I) The extension AR : R → L2 of A is a Fredholm operator. In particular,
Im(AR ) is closed in L2 (i.e., AR is normally-solvable).
(II) The spaces Ker AR and Coker AR are (respectively in the case of the
cokernel: can be represented as) finite–dimensional subspaces of the spaces
of smooth sections.
Remark 13.14. Let H(A) denote the space of all L2 solutions of As = 0. Then
we may reformulate condition (II) in the definition as follows:
R ∩ H(A) ⊂ C ∞ and R∗ ∩ H(A∗ ) ⊂ C ∞ ,
where R∗ denotes the domain of the adjoint operator.
344 13. CLASSICAL APPLICATIONS (SURVEY)

It seems at first sight that the boundary Y does not appear in the definition,
but usually the domain R is defined by a condition posed on the sections on the
boundary: Let γ0 denote the restriction map γ0 (s)(y) := s(0, y). It gives a con-
tinuous map γ0 : W 1 (X; E) → L2 (Y ; E|Y ). There is a full proof of the sharper
Sobolev Trace Theorem in [83, Theorem 11.2], see also our Theorems 7.14 and 7.17
for precise formulation and the basic ideas. The condition which determines R is
typically given in the form:
R = {s ∈ W 1 (X; E) : T (γ0 (s)) = 0} ,
where T : L2 (Y ; E|Y ) → L2 (Y ; G) is a 0–th order pseudo–differential operator. Of
course T has to satisfy certain additional assumptions to guarantee fulfilment of
conditions (I) and (II) from the definition. We have to introduce the Calderón
projection in order to explain those conditions.
The Calderón projector P+ (A) was defined in [379, 382] in greater gen-
erality and as the decisive ingredient to A.P. Calderón’s program [102] of 1963.
For linear elliptic differential operators of first order, it is the (without loss of gener-
ality orthogonal) projection of L2 (Y ; E) onto the Cauchy data space, also called
Hardy space in Clifford analysis:
L2 (Y ;E|Y )
(13.40) H+ (A) := {s|Y : s ∈ C ∞ (X; E) and As = 0 in X \ Y } .
It is shown in [379, 382, 83] that P+ (A) is a pseudo-differential operator of
order zero and that the principal symbol p+ (y; ζ) of P+ (A) is equal to the orthog-
onal projection onto the direct sum of the eigenspaces of the automorphism b(y; ζ)
corresponding to the positive eigenvalues. Here b denotes the principal symbol of
the tangential operator 6 ∂. Now we are ready to formulate the conditions which the
operator T has to satisfy:
Definition 13.15 (Symbolic calculus). Let T ∈ L0pc (E|Y , E|Y ), i.e., let
T : C ∞ (Y ; E) −→ C ∞ (Y ; E)
be a principally classical pseudo–differential operator of order 0. We call T an
elliptic boundary condition for the operator A, if the following conditions are
satisfied:
(I’) For any real ρ, the extension T ρ : W ρ (Y ; E) → W ρ (Y ; E) of T has a closed
range.
(II’) Let σ(T ) denote the principal symbol of T . Then
Im(σ(T )) = Im(σ(T ) ◦ p+ ).
In particular the restriction σ(T )|Im(p+ ) : Im(p+ ) → Im(σ(T )) is an iso-
morphism of vector bundles.
Condition (I’) implies that NT , the orthogonal projection onto the kernel of
T , is a pseudo-differential operator (see [83, Proposition 18.11]). For the ease of
notation we shall denote the kernel of T by the same letter NT . Condition (II’)
implies that the couple (NT , H+ (A)) is a Fredholm pair of subspaces, i.e. a
pair of closed subspaces with finite–dimensional intersection and with closed sum of
finite codimension (then the difference of these two dimensions is called the index
of the pair; see [83]). It should be noted that apparently Bogdan Bojarski in
[66] was the first to bring the concept of Fredholm pairs into the index theory of
13.8. THE INDEX OF ELLIPTIC BOUNDARY-VALUE PROBLEMS 345

boundary value problems. It is common


 to call a pair of projections (P1 , P2 ) a
Fredholm pair, if Ker(P1 ), Im(P2 ) is a Fredholm pair.

Remark 13.16. a) We get a local elliptic condition in the sense of Šapiro and
Lopatinskiı̆ when the range of σ(T ) can be written as the lifting of the vector
bundle E|Y under the natural projection π : T̊ ∗ Y → Y .
b) Local boundary conditions continue to have their place in global analysis and
geometry, as we shall see below in the arguments to Result g regarding the cobor-
dism theorem. We also refer to G. Grubb and E. Schrohe who obtained remark-
able results on traces and quasi-traces on the Boutet de Monvel algebra in [193]
with applications to the study of the multiplicative anomaly of zeta-determinants
of elliptic operators.

Results.
a) Let T ∈ L0pc (E|Y , E|Y ). We denote by AT the extension of A with domain
{s ∈ W 1 (X; E) : s|Y ∈ NT }. Then the operator AT satisfies the conditions of
Definition 13.13 if and only if T satisfies the conditions of Definition 13.15.
b) Let T be as in Definition 13.15. Then the couple (NT , H+ (A)) is a Fredholm
pair of subspaces in L2 (Y ; E|Y ) with
(13.41)
index(NT , H+ (A)) = index {T ◦ P+ (A) : H+ (A) → Im(T )} = index AT .

c) The Calderón projection is closely related to another projection determined by


the tangential part B of A in the product form (13.38) in a collar neighborhood of
Y . Since A is of Dirac type and we assumed product form of all metric structures
near Y , B is a symmetric elliptic differential operator over Y . It has discrete real
eigenvalues and a complete system of L2 orthonormal eigensections. Let P≥ (B)
denote the spectral (Atiyah–Patodi–Singer) projection onto the subspace
L≥ (B) of L2 (Y ; E|Y ) spanned by the eigensections corresponding to the nonnega-
tive eigenvalues of B. It is a pseudo-differential operator and its principal symbol
p+ coincides with the principal symbol of the Calderón projection P+ (A).
d) We call the space of pseudo-differential projections with the same principal
symbol p+ the Grassmannian Grassp+ (= Grass(A)) and equip it with the op-
erator norm corresponding to L2 (Y ; E|Y ). Here projection means idempotent (i.e.
P = P 2 ). The Grassmannian has countably many connected components; two
projections P1 , P2 belong to the same component, if and only if the virtual codi-
mension

(13.42) i(P2 , P1 ) := index {P2 P1 : im P1 −→ im P2 }

of P2 in P1 vanishes; the higher homotopy groups of each connected component are


given by Bott periodicity.
e) Choosing the spectral projection P≥ (6 ∂) as boundary condition for the chi-
ral (half) operator D+ on an even-dimensional manifold X, we have the famous
Atiyah–Patodi–Singer Index Theorem which gives
Z
1
(13.43) index DP+≥ (∂/) = α(x) − (η∂/ (0) + dim Ker 6 ∂) .
X 2
346 13. CLASSICAL APPLICATIONS (SURVEY)

Here α(x) denotes the locally defined index density of D+ which expresses the
local chiral anomaly, and
Z ∞
X 1 z−1 2
(13.44) η∂/ (z) := sign λ|λ|−z = t 2 tr(6 ∂ e−t∂/ ) dt
Γ( z+1
2 ) 0
λ∈spec ∂
/\{0}

denotes the η–function of the tangential Dirac operator 6 ∂ (see our previous dis-
cussion on pp.111ff).
f ) For our four-dimensional Example 13.11, M. Ninomiya and C.I. Tan gave in
1985 in [322] the following application of the Atiyah-Patodi-Singer Index Theorem:
For a suitable metric we have
(13.45) index(6 D+
A )Ph = deg(h)
˚ × C2 ) with A = h ◦ d ◦ h−1 in a collar of ∂V and suitable smooth
for any A ∈ C(V
h : ∂V → SU(2) where Ph denotes the corresponding spectral Atiyah–Patodi–Singer
projection.
g) If n = dim X is odd, we have a variety of index problems and results for the
total Dirac operator D. Recall from  that we have D|N = J (∂r + B)|N in a
(13.34)
0 B−

collar N of Y = ∂X, with B = .
B+ 0
g1. The Atiyah–Patodi–Singer Index Formula in this case gives
index DP≥ (B) = − dim ker B + .
g2. Let dimX = 3, ∂X = Y connected and the genus g(Y ) ≥ 2. Let A
be a smooth flat connection on X × SU(2) which restricts to a product
connection B × Id in the collar neighborhood [0, 1] × Y × SU(2) of Y ×
SU(2) for an irreducible flat connection B on Y × SU(2). We define
the corresponding twisted signature operator 6 DA : W → W with
coefficients in su(2) by
6 DA (a, b) := (d∗A b, ∗dA a + da b) ,
 
where W := Ω0 (X) ⊗ su(2) ⊕ Ω0 (X) ⊗ su(2) , a a 0-form, b is a 1-form,
dA denotes the covariant derivative,
R and the adjoint is taken relative to
an L2 inner product (a, b)0 := − X tr (a ∧ ∗b) for a, b ∈ Ωp (X) ⊗ su(2).
Then we have
(13.46) index(6 DA )P≥ = 3 − 3g.
g3. The chiral projections π± : S|Y → (S|Y )± define two natural local elliptic
boundary conditions with
(13.47) index Dπ− − index Dπ+ = index B + = 0.
g4. Let D : C ∞ (X; S) → C ∞ (X; S) be a compatible Dirac operator over X,
m ∈ N and g : Y → H(m) a smooth family of Hermitian matrices. Then
we get a self–adjoint elliptic operator Dg,m acting like
 
mD 0
Dg,m :=
0 −mD
with a local elliptic boundary condition imposed by
  
s1 1 m m

dom Dg,m := ∈ W X; (S ⊗ C ) ⊕ (S ⊗ C ) : s2 |Y = (Γ ⊗ g)s1 |Y ,
s2
13.8. THE INDEX OF ELLIPTIC BOUNDARY-VALUE PROBLEMS 347

where mD := D ⊗ IdCm .
h) There are several formulas to describe the index jumps under change of the
boundary condition and/or the connection defining the Dirac operator. Here are
two of these index correction formulas:
h1. For P1 , P2 ∈ Grass(A) with A ∈ {D, D± } we have
(13.48) index AP2 = index AP1 + index (P2 ◦ P1 : Im(P1 ) → Im(P2 )) .
h2. Let {∇t , t ∈ [0, 1]} be a smooth homotopy of connections for a fixed
Hermitian Clifford modules bundle S over X. Let {Dt } be the induced
curve of Dirac type operators over X and {Bt } its corresponding family
of tangential operators. Then
(13.49) ind(D0 )P≥ (B0 ) − ind(D1 )P≥ (B1 ) = sf{Bt }t∈[0,1] ,
where the last expression is the spectral flow. Roughly speaking, the spec-
tral flow counts the net number of eigenvalues changing from the negative
real half axis to the non-negative one. For the concept of a spectral flow
see [71].
k) Let M be a closed odd-dimensional smooth Riemannian partitioned manifold
M = X− ∪Y X+ , where X− ∩ X+ = ∂X− = ∂X+ = Y
and Y a hypersurface. We assume that M \ Y does not have a closed connected
component (i.e., Y intersects any connected component of X− and X+ ). Let D
be an operator of Dirac type over M . Then the pair (H− (D), H+ (D)) of Cauchy
data spaces of D along Y , as defined above in (13.40), makes a Fredholm pair of
subspaces in L2 (Y ; S|Y ) and we have

(13.50) index D = index H− (D), H+ (D) .
Arguments. To a) Our Claim a) is modelled after Lopatinskiı̆’s result: As
mentioned before, he proved in [282], that for local boundary value problems a
condition expressed in symbolic calculus as in Remark 13.16 is necessary and suf-
ficient for smoothness of the solutions and finite index. For our wider classes of
boundary problems, it is not so difficult to deduce the Fredholm property from the
symbolic assumptions, see Claim b). To derive the regularity property from the
symbolic assumptions, one has to construct a kind of lifting jack like Gårding’s
Inequality, see above Exercise 9.12 and the details in [83, Theorem 19.6]. It is more
demanding to derive the symbolic calculus conditions from the normal solvability,
see [99].

To b) We show that index(NT , H+ (A)) = index{T P+ (A) : H+ (A) → Im(T )} and


refer to [83, Theorem 20.8] for the proof of the second equality.
Let us assume that z is an element of NT ∩ H+ (A). This implies that
(13.51) T z = 0 and P+ (A)z = z,
which shows that z is an element of the kernel of the operator T P+ (A). Let us
also observe that the second equality of (13.51) shows that there exists a uniquely
determined s, such that As = 0 and γ0 (s) = z. Therefore we have shown:
Ker AT = Ker T P+ (A) = NT ∩ H+ (A).
Now let us assume that w is an element of (NT + H+ (A))⊥ . It means that
w is perpendicular to H+ (A), hence P+ (A)w = 0 and that w is perpendicular to
348 13. CLASSICAL APPLICATIONS (SURVEY)

the kernel of T . Therefore there exists q such that w = T ∗ q, which provides the
identification of Ker P+ (A)T ∗ with the orthogonal complement of the sum and thus
proves the claim.

To c) Since P+ (A) and P≥ (B) have the same principal symbol, they differ only by a
compact operator, actually only by an infinitely smoothing operator as shown by S.
Scott in [374]. Instead of P≥ (B) one might consider weighted spectral projections
P≥a (B) and P>a (B) for arbitrary real a. They all are pseudo-differential projections
with the same principal symbol.
For the Cauchy-Riemann operator on the disc D2 = {|z| ≤ 1}, the Cauchy
data space is spanned by the eigenfunctions eikθ of the tangential operator ∂θ over
S 1 = [0, 2π]/{0, 2π} for nonnegative k. So, the Calderón projection and the Atiyah–
Patodi–Singer projection coincide in this case. One can generalize the preceding
example: For any smooth compact manifold X with boundary Y and any real
R ≥ 0, let X R denote the stretched manifold

X R := ([−R, 0] × Y ) ∪Y X.

Since we assume product structures with A = Γ(∂r + B) near Y , we have a well-


defined extension AR of A. L. Nicolaescu proved in [320] that the Calderón pro-
jection and the Atiyah–Patodi–Singer projection coincide up to a finite-dimensional
component in the adiabatic limit (R → +∞ in a suitable setting).
In [189], G. Grubb considered an arbitrary orthogonal pseudo-differential pro-
jection P of L2 (Y ; E) which makes a Fredholm pair with the Calderón projection
P+ (A). If P defines a self-adjoint boundary condition for A, then, she proved, there
exists a Dirac type operator B 0 over Y such that P = P> (B 0 ). Grubbs Lemma
shows that the Atiyah-Patodi-Singer boundary projection is the most general ad-
missible (self-adjoint) boundary condition, in the specified sense.

To d) For A = D, the symmetric total Dirac type operator, we have also an


important self–adjoint Grassmannian Grasssa (D) of smooth self–adjoint boundary
conditions of Atiyah–Patodi–Singer type. This is the subspace of Grass(D), which
consists of those smooth (orthogonal) projections Π, which satisfy the condition

−J ΠJ = Id −Π

with the anti-involution J of Clifford multiplication which appeared in the product


formula (13.32) near the boundary. Any element of Grasssa (D) defines a self–adjoint
elliptic boundary value problem for the operator D. With respect to the operator
norm, the topological space Grasssa (D) is connected and higher homotopy is given
by Bott periodicity.
In [82, Section 2] an intimate relation was found between the self-adjoint Grass-
mannian of the total Dirac operator and the Grassmannian of the chiral halves: The
image of the mapping
 
P 0
(13.52) Grass(D+ ) 3 P 7→ P # := ∈ Grasssa (D)
0 J(Id −P )J −1

provides us with the γ5 -Grassmannian Grasssa γ5 (D) of self–adjoint elliptic (well–


posed) boundary conditions for the total Dirac operator D which are all γ5 -invariant.
13.8. THE INDEX OF ELLIPTIC BOUNDARY-VALUE PROBLEMS 349

 
1 0
Indeed, in view of the chiral splitting γ5 takes the form γ5 = for dimen-
0 −1
sion of X divisible by 4 (otherwise we multiply by the imaginary unit i). Whence
γ5 P # γ5 = P # .
If we set (following [246])
1 J −1
 
1
(13.53) T := ,
2 J 1
we obtain a well-posed (i.e., elliptic) local boundary condition yielding a self-adjoint
Fredholm realization DT , namely satisfying
A) Definition 13.13 above (normal solvability) and
B) the Lagrangian (symmetry) condition Id −T = −J T J .
One decisive difference between the two boundary conditions P # and T defined in
equations (13.52) and (13.53) lies in the γ5 -symmetry: We have γ5 T γ5 = Id −T ,
i.e., T is not γ5 -invariant.
Our map P 7→ P # induces a natural identification
∼ ∼
(13.54) π0 (Grasssa +
γ5 (D)) ←− π0 (Grass(D )) −→ Z.

Moreover, for all P # ∈ Grasssa 2


γ5 (D) the L realization DP # has a discrete real spec-
trum. Each eigenvalue is of finite multiplicity and there are no finite accumulation
points. The spectrum is symmetric around the origin of the real axis (hence there
is no η function for the operator DP # on X, contrary to the relevance of the eta-
invariant η∂/ (0) of the tangential operator 6 ∂ on Y in the Atiyah-Patodi-Singer index
theorem (13.43)). The null space
Ker DP # := {s ∈ W 1 (X; S) : D(s) = 0 and P # (s|Y ) = 0}
consists solely of smooth spinors. It is finite–dimensional and splits naturally into

the direct sum Ker DP # = Ker DP+ ⊕ Ker DJ(Id −P )J −1 of a space of spinors of
#
positive chirality of dimension n+ (P ) and a space of spinors of negative chirality
of dimension n− (P # ) with
(13.55) index DP+ = n+ (P # ) − n− (P # ) ,
where P # and P are related through (13.52).
To sum up (see [82] for the elaboration of the physics meaning in quantum
chromodynamics, see also our Result f):
• Vanishing index DP+ with P ∈ Grass(D+ ) means that P # is a γ5 -invariant
well-posed self-adjoint boundary condition for the total Dirac operator
with global chiral symmetry:

n+ := dim Ker DP+ = n− := dim Ker DJ(Id −P )J −1

with dim Ker DP # = n− + n+ .


• Nonvanishing index DP+ means also γ5 -invariance of Ker DP # , but with
global chiral asymmetry.
All that follows at once from the general theory of global elliptic boundary
 the spectrum of DP # is λ 7→ −λ
problems for the Dirac operator. To prove that
s+
symmetric we consider an eigenspinor s = ∈ dom DP # with Ds = λs.
s−
350 13. CLASSICAL APPLICATIONS (SURVEY)

Because of the anti–diagonal form of D this means


D− s− = λs+ and D+ s+ = λs− .
 
s+
Then also belongs to dom DP # and is an eigenspinor of DP # with eigen-
−s−
value −λ, since
0 D− −D− s−
        
s+ −λs+ s+
= = = −λ
D+ 0 −s− D+ s+ λs− −s−
and, trivially,
J(Id −P )J −1 (s− |Y ) = 0 =⇒ J(Id −P )J −1 (−s− |Y ) = 0 .

To e) The proof of (13.43) in [41], and slightly differently in [83] (see also our
extended summary in [60, Section 2.4]), is based on the heat kernel method for
computing the index, but the process is less straightforward than in the closed case
because of the boundary condition. The appropriate heat kernel is constructed
by means of J.M.C. Duhamel’s method of [131]. As rediscovered in [129] and
elaborated in [83, Section 22.C], the Duhamel Principle allows one to study the
interior contribution and the boundary contribution separately and identify the
singularities caused by the boundary contribution: An exact kernel is obtained
from an approximate one by an iterative process initiated by writing the error as
the integral of a derivative of the convolution of the true and approximate kernel.
The initial approximate heat kernel is obtained by patching together two heat
kernels, denoted by Ec and Ed . Here, Ec is a heat kernel for a Dirac operator over
an infinite extension [0, ∞) × Y of the collared neighborhood N = [0, 1] × Y of Y in
X, for which the boundary condition P≥ (6 ∂) (ψ|Y ) = 0 is imposed. The heat kernel
−g+
Ed is the usual one (without boundary conditions) for e−tD D , where D ± are the
g g
chiral halves of the invertible Dirac operator, namely
   
D
g ± := D ± ∪ D ∓ : C ∞ X, e Sf± → C ∞ X, e Sf±

over the double X, e a closed manifold without boundary. Note that the proof
heavily relies on the existence of an invertible double (established in [83, Chap-
ter 9]). For that, a weak unique continuation property and the symmetry of the
tangential operator are exploited. Hence the proof does not generalize to arbi-
trary elliptic differential operators on compact manifolds with smooth boundary
and Atiyah-Patodi-Singer-type boundary conditions. For possible generalizations
— and obstructions, though, see [80].
In [41] the first term of the index theorem was calculated in cohomological
expressions, namely:
Z Z

α(x) = ch (S, ε) ∧ A
e (X, θ) ,
X X
where ch (S, ε) ∈ Ω∗ (X, R) is the total Chern character form of the complex vector
bundle S with compatible, unitary connection ε, and A e (X, θ) ∈ Ω∗ (X, R) is closely
related to the total A
b (X, θ) form relative to the Levi-Civita connection θ, namely
2k−m b
A (X, θ)4k = 2
e A (X, θ)4k .
One can not expect global chiral symmetry for the Atiyah–Patodi–Singer bound-
ary problem; in general, none of the expressions in formula (13.43) will vanish. For
13.8. THE INDEX OF ELLIPTIC BOUNDARY-VALUE PROBLEMS 351

sufficiently elementary
R operators and under additional assumptions some of the
three terms X α(x), η∂/ (0) and dim ker 6 ∂ will vanish (namely for local chiral sym-
metry of D+ , for symmetric spectrum of 6 ∂, and if 6 ∂ is invertible, respectively). It
also happens that fairly easy expressions are obtainable for non-vanishing terms
(see our Result f).

To f ) On the way to their result, Ninomiya and Tan made in [322] several
interesting observations: 1. A connection A for V × C2 is pure gauge at the
boundary if it can be written in the form
(13.56) A = d − (dh)h−1 in a collar of the boundary.
Moreover, if A is pure gauge at the boundary, then the tangential Dirac operator
B over ∂V corresponding to the partial (half) twisted Dirac operator 6 D+
A over V
takes the form
(13.57) B = 6 ∂ S 3 ⊗−(dh)h−1 Id = (Id ⊗h) (6 ∂ ⊗ IdC2 ) (Id ⊗h−1 )
with 6 ∂ ⊗ IdC2 = 6 ∂ ⊕ 6 ∂. Here 6 ∂ = 6 ∂ S 3 denotes the tangential operator over S 3
corresponding to the euclidean Dirac operator 6 D+ .
To prove (13.56), we find
Af = (hdh−1 )f = hd(h−1 f ) = h(h−1 df + d(h−1 )f ) = df − (dh)h−1 f .
To prove (13.57), we notice that the restriction of A to the boundary takes the form
−(dh)h−1 , therefore we get such a simple form for lifting 6 ∂ to the auxiliary bundle.
For details of the calculation see e.g. [329] and [83].
2. In the same article [322] Ninomiya and Tan pointed out that the Atiyah–
Patodi–Singer boundary condition is natural or physical in the following sense: Let
˚ × C2 ) with corresponding h : ∂V → SU(2). Then Ph := P≥ (6 ∂ ⊗h IdC2 ) ∈
A ∈ C(V
Grass(6 D+ ⊗A IdC2 ) and P (h) := (Ph )# ∈ Grasssa γ5 (6 D ⊗A IdC2 ). Consider the family
of operators {6 DA,P (h) }A∈C̊(V ×C2 ) which act like 6 D ⊗A IdC2 with domain

dom 6 DA,P (h)


    
1 2 Ph 0 s+
:= s ∈ W (V ; S ⊗ C ) : P (h)(s) = =0 .
0 J(Id −Ph )J −1 s−
It satisfies the following three fundamental conditions:
i. 6 DA,P (h) is self–adjoint;
ii. P (h) is γ5 -invariant; and
iii. the domain dom 6 DA,P (h) is gauge–invariant.
Before arguing for i-iii, we recall the meaning of iii: Let U : V → SU(2) denote
a gauge transformation
f (x) 7→ U (x)f (x)U −1 (x) ,
∂U
Aµ (x) 7→ U (x)Aµ (x)U −1 (x) − (x)U −1 (x) ,
∂xµ
then the connection A transforms as follows:
Aeµ f |x 7→ U (x)(Aeµ f |x )U −1 (x),
A 7→ U AU −1 , and
ΩA 7→ U ΩA U −1 − (dU )U −1 .
352 13. CLASSICAL APPLICATIONS (SURVEY)

This motivates the following definition:


Definition 13.17. A smooth family
C(V × C2 ) 3 A 7→ P (A) ∈ Grasssa ∼ +
γ5 (6 DA ) = Grass(6 DA )

is gauge–invariant, if we have
(13.58) P (A1 ) = U # P (A)(U # )−1
for all A, A1 ∈ C(V × C2 ) where A1 := U AU −1 with arbitrary U : V → SU(2) and
U # := Id ⊗C2 U .
Clearly, for any A, A1 ∈ C(V × C2 ) with A1 = U AU −1 we have pointwise (see
[329])
6 DA1 = 6 D ⊗A1 Id = U # (6 D ⊗A Id)(U # )−1 = U # 6 DA (U # )−1 .
The crucial point of property iii — gauge–invariance as defined in (13.58) — is
6 DA1 ,P (A1 ) = U # 6 DA,P (A) (U # )−1 ,
and, especially,
(13.59) dom(6 DA1 ,P (A1 ) ) = U # (dom 6 DA,P (A) ) ;
i.e., we require that the boundary condition transforms in a correct way under
variation of the background operator resp. of the connection.
Properties i and ii are obvious from the choices. The gauge invariance follows from
the corresponding transformation law for the tangential operator
(13.60) 6 ∂ ⊗h1 IdC2 = U # (6 ∂ ⊗h IdC2 )(U # )−1 ,
where the smooth families h, h1 := U |∂V h(U |∂V )−1 : ∂V → SU(2) correspond to the
connections A, A1 which are supposed to be pure gauge at the boundary. Equation
(13.60) implies that the eigenvalues do not change under gauge transformation and
that the eigenspaces transform like Eλ (6 ∂ h1 ) = U Eλ (6 ∂ h ). Hence P (h) satisfies i, ii,
and iii.
3. Now we prove (13.45). We apply the Atiyah–Patodi–Singer Index Formula for
˚ × C2 ) with corresponding h : S 3 → SU(2):
the operator 6 DA with A ∈ C(V
Z
1
index(6 D+ )
A Ph = α(x) − (ηB (0) + dim ker B) .
2
From (13.57) we have
ηB (0) = 2 η∂/S3 (0) and dim ker B = 2 dim ker 6 ∂ S 3 .
We find deg(h) for the value of the integral of the index density. This result is
actually independent of the choice of the metric. By Theorem 10.4 (Bott Periodicity
for h : S n−1 → GL(N, C) with n = 4 even, N = 2, 2N ≥ n, and SU(2) ⊂ GL(2, C)),
a multiple of ± deg(h) is the only integer-valued invariant we can get from h.
Then
(13.61) index(6 D+
A) = deg(h) − η∂/S3 (0) − dim ker 6 ∂ S 3 .
Ph

The two numbers on the right were found to vanish for the standard metric of R4 ,
slightly modified close to ∂V in a calculation done in [369] by A.M. Bincer and
J.R. Schmidt, see also [368]. In that metric the tangential Dirac operator on the
3–sphere 6 ∂ S 3 has a spectrum symmetric about λ = 0 and is invertible.
13.8. THE INDEX OF ELLIPTIC BOUNDARY-VALUE PROBLEMS 353

We also refer to a series of papers [254, 255, 256] on 4-dimensional spinor


analysis of T. Kori, which ensure the same result, namely a symmetric spectrum,
not containing zero, for a particular metric set–up coming from a metric over the 4–
ball which is product near the boundary. Actually, for Kori’s metric the Calderón
projector P+ and the Atiyah–Patodi–Singer projection P≥ coincide. Also from
[215] of N. Hitchin it follows directly that the tangential Dirac operator over the
3–sphere in standard metric is non–singular with symmetric spectrum.
4. Present wisdom in quantum chromodynamics relies on the relation between the
topological density of gauge field configurations and the so–called local chiral anom-
aly, namely, the appearance Rof a topological term in the conservation equation of
the chiral current, deg(h) = V α(x). Non-vanishing of deg(h) is usually regarded
as global chiral asymmetry in the sense of (13.55) (i.e., n+ 6= n− ), implying the
impossibility of a functional integral formulation of the ζ-regularized determinant
(see p.118 above) invariant under rigid , (i.e., global ) chiral transformations. For
γ5 -invariant boundary conditions and domains that are pure gauge at the bound-
ary, the message of the Atiyah–Patodi–Singer Index Theorem is clear: local chiral
anomaly (deg(h) 6= 0) implies global chiral asymmetry. However, one should keep
in mind that from a geometrical point of view there is no particular reason to
choose the Atiyah–Patodi–Singer (APS) boundary conditions among all the other
boundary problems which, as we shall see, equally satisfy i–iii; in other words, a
stronger requirement than covariance of the domain under gauge transformations
of the connection seems to be necessary in order to select that boundary condition.
There are other quite natural boundary conditions RA instead of PA which fulfil the
conditions i–iii (of p.351) and additionally provide global chiral symmetry, namely
iv. the vanishing of index 6 D+
A,RA = n+ (RA ) − n− (RA ).
There are various ways of obtaining global chiral symmetry by imposing el-
liptic, self–adjoint, γ5 -symmetric, and gauge–invariant boundary conditions in the
presence of local chiral anomaly (i.e. non–vanishing deg h for connections which
are pure gauge at the boundary). One way is the Calderón projector (introduced
on p.344). It removes all solutions such that kernel and cokernel become trivial.
This makes many calculations easy. But since the Calderón projector depends on
the gauge configuration also inside the region and not only on the boundary, this
must have consequences in, for example, the derivation of identities by variation of
the gauge field configuration in a subregion.
Instead of removing all solutions by imposing the Calderón projector one can
add further solutions to the original Dirac equation with APS boundary condition
until one gets global chiral symmetry. There are three ways to do that.
Let us begin with a given connection A in the auxiliary bundle V × C2 which is
pure gauge at the boundary. Hence it can be expressed in a collar of the boundary
by a mapping h : S 3 → SU(2) which has a degree (topological number) deg h. If
deg h = k is non–trivial, the dimensions n+ and n− of the zero modes (subject
to the APS boundary condition Ph ) do not coincide. Then, to get global chiral
symmetry we enlarge the solution spaces until n+ and n− become equal. More
precisely:

Alternative 13.18. The easiest, but physically hardly very meaningful way
of doing the equalization of the solution spaces is to take a second copy of the
coefficients bundle V × C2 and to choose a connection A0 which is pure gauge at
the boundary ∂V with a unitary mapping g of opposite degree −k.
354 13. CLASSICAL APPLICATIONS (SURVEY)

Then, instead of tensoring the original euclidean Dirac operator D+ solely with
the h–connection, we do two twistings: first with h, then with g. The resulting
twisted Dirac operator
D0+ := D+ ⊗A IdC2 ⊗A0 IdC2 = DA
+
⊗A0 IdC2
with coefficients in C2 ⊗ C2 = C4 admits again an APS boundary condition P 0
which is gauge invariant such that
n+ − n− = index D0+ P 0 = deg(h ⊗ g)
= deg hg = deg h + deg g = k − k = 0.
To get global chiral symmetry one can also apply a less trivial mirror process:
+
Alternative 13.19. Instead of twisting the global Dirac operator DA over the
full 4–ball V it suffices to twist the transversal (tangential) Dirac operator Bh with a
connection of opposite degree over the 3–sphere. We get a new operator Bh0 . Then
we apply the APS spectral projection Ph0 corresponding to the twisted operator
+
Bh0 to the original operator DA . It follows that Ph0 is an admissible boundary
+
value problem for DA . It belongs to the same Grassmannian as the standard
APS projection Ph and all the nice properties i–iii are guaranteed, but Ph0 belongs
to a different connected component. In fact, the index jumps by the winding
number yielding global chiral symmetry. An attractive feature of Alternative 13.19,
discussed in [307] (see also [308]), is that in fact the (non–free) operator DA is not
changed; only the boundary condition is changed.
Alternative 13.20. One more alternative is provided by a suitable spectral
cut (weighted spectral projection). For the problem of uniform choice of the spectral
cut, see [293, 73], and [85, Appendix].

To g1) On closed manifolds, symmetric (i.e., formally self-adjoint) differential op-


erators are essentially self-adjoint with vanishing index. And definitely, all differ-
ential operators, symmetric or not, have vanishing index on closed manifolds of
odd dimension (SectionExVanIndSect, p.317). So, the non-vanishing index of the
formally self-adjoint total Dirac operator D on an odd -dimensional manifold X is
caused by the boundary condition. As announced in Part I (pp.47f), it is one of
the most fascinating aspects of the index theory of Singer and Atiyah to follow
the emergence of asymmetry out of symmetry.
The proof sketched above for the Atiyah-Patodi-Singer Index Theorem is valid
in the odd-dimensional case as well and yields a formula that is simpler than in the
even-dimensional case:
Z
1
index DP≥ (B) = α(x) − (ηB (0) + dim Ker B) = − dim Ker B + ,
X 2
since
• the locally defined index density α(x) vanishes pointwise on odd-dimensional
manifolds (see (17.128) and Theorem 17.59);
• each eigenspinor ϕ of B with Bϕ = λϕ and λ 6= 0 is mirrored by an eigen-
spinor J ϕ of eigenvalue −λ, so the spectrum of B is symmetric around 0
and ηB (0) = 0;
• by the Cobordism Theorem (our Result g3), we have index B + = 0, so
dim Ker B = dim Ker B − + dim Ker B + = 2 dim Ker B + .
13.8. THE INDEX OF ELLIPTIC BOUNDARY-VALUE PROBLEMS 355

Note that for non–vanishing kernel of the tangential operator the Atiyah–Patodi–
Singer Problem for the total, symmetric Dirac operator is not self–adjoint, and
its index is not stable under small deformations. On the other hand, Green’s
formula (13.31) shows that in the case Ker(B + ) = {0} the operator DP≥ (B) is an
(unbounded) self–adjoint operator.
To g2) This is a particularly nice application of the Atiyah-Patodi-Singer Index
Formula and has gained some prominence in low-dimensional geometry and gauge
theory, see Taubes [407, Section 2, Proposition 4.9, and Lemma A.4] and Yoshida
[450, Sections 1-4]. For the definition of the applied geometric concepts we refer
to our Chapter 15 on geometric terminology (though we deviate here and use A
and B for connections as the physicists do). For the definition of the twisted Dirac
operator (which is only locally a true operator of Dirac type) we refer to our Chapter
17, where we discuss all technical details. For the derivation of (13.46) from the
APS Theorem (13.43) we refer to [83, pp.250–252].
To g3) The first equality in (13.47) is a trivial consequence of the chiral splitting.
But index Dπ± vanishes by Green’s formula so that we get index B + = 0. That is the
illustrious cobordism theorem, namely the vanishing of the index of any (half)
Dirac operator over a closed even–dimensional manifold Y that can be written as
the (half) tangential operator of a (total) Dirac operator over an odd–dimensional
manifold X with ∂X = Y . For generalizations and other proofs of the Cobordism
Theorem see our Note on p.303 in Section 12.3.
To g4) In quantum field theory, the operator Dg,m is known as the chiral bag
model. In mathematics it became prominent when Singer in [399] chose this
example to explain his view upon determinants of Dirac type operators, having
the ζ-function regularized determinant of the Dirac Laplacian as modulus and the
η-invariant of the operator itself as phase.
To h1) Treating elliptic boundary problems as projections yields a corresponding
variant of the Agranovich-Dynin formula for global boundary conditions of general-
ized Atiyah-Patodi-Singer type, i.e., belonging to the Grassmannian Grass(A). The
original Agranovich-Dynin Formula of [8] is for local elliptic boundary value
problems. Formula (13.48) is a consequence of (13.41) of Result b, see also [83,
Chapter 21]. The formula is a key result since it permits to switch easily from the
strict APS boundary condition P≥ to other possibly more appropriate boundary
conditions in the Grassmannian. One caveat arises from our discussion of Result f,
when we found for quantum chromodynamics that sticking to the true APS projec-
tion can be misleading and, for instance, applying a weighted APS projection more
natural. A second caveat arises from the fact that the signature of 4k-dimensional
manifolds with boundary is not the index of a true APS problem: From [41, 1975,
Section 4, Theorem 4.14 and Equation (17.9)] we extract the two following formulas:
Z Z
1
(13.62) sig X = L(x) − η∂/ (0) and index DP+≥ = L(x) − h − η∂/ev (0),
X 2 X

and hence the signature deficiency formula sig X = index DP+≥ + h, where h is
the multiplicity of the zero-eigenvalue of the tangential Dirac operator 6 ∂ to the
signature operator D+ . Then, by (13.48) we obtain sig X = index DP+Σ for the
projection PΣ onto any subspace Σ of Im P≥ of codimension h.
So, regarding the APS boundary condition one should keep in mind that, in
contrast to the Calderón projection, there is nothing canonical or natural about
356 13. CLASSICAL APPLICATIONS (SURVEY)

the spectral projection P≥ . It is just easy to define while the hard choice of the
cutting point — should it be 0 or another a ∈ R — can not always be avoided.
To h2) The clue of (13.49) is to trace the variation of the APS projection over
Y = ∂X under variation of the connection that defines the Dirac operator over X.
For that, the arguments of [41, p. 95] were worked out in [363, Theorem 1.4] and,
differently and in detail, in [278, Theorem 7.6]. It is also called the Spectral Flow
Theorem. The continuous dependence of P + (Bt ) on Bt (in the sense that P + (Bt )
has the same jumps as 1(−ε,ε) (Bt ), if ±ε 6∈ spec Bt ) is important in this theorem.
When Bt is self-adjoint — as is the case for the tangential operator of the signature
operator — it can be proved by standard techniques of functional analysis (cf. [83,
Chapter 17]). For the concept of a spectral flow see [71].
It is natural to consider a more general case. In [364], A. Savin, B.-W.
Schulze and B. Sternin gave a similar formula for the case that the tangential
family Bt is non-self-adjoint. It seems very satisfactory that [73] complemented
their result by proving the continuous dependence of the sectorial projection P + (Bt )
on Bt when Bt has no spectral points on the imaginary axis for all t ∈ [0, 1]. In her
comment to [73] in [192], G. Grubb gave a definition of the sectorial projection
from logarithms. That makes it a lot easier to derive the mentioned main result of
[73].
To k) Formula (13.50) is called the Bojarski Conjecture. It was suggested by B.
Bojarski in [66] in great generality and proved in [83, Chapter 24] for operators
of Dirac type. It relates the quantum quantity index with a classical quantity,
the Fredholm intersection index of the Cauchy data spaces from both sides of the
hypersurface Y . L. Nicolaescu found in [320] an even-dimensional analogue to
the Bojarski Conjecture, namely expressing the spectral flow of a curve of Dirac
type operators over M by the Maslov (intersection) index of the corresponding two
curves of Cauchy data spaces. For the concept of a spectral flow see [71]. That
spectral flow formula generalizes many predecessors, dating back to the Morse
Index Theorem, and has itself received various generalizations in recent years, see
the historical review in [85, Section 1].
By recovering the index of a given Dirac operator D over a partitioned manifold
from the Cauchy data spaces H± (D), the Bojarski Conjecture / Theorem can be
considered as a pioneering contribution to the Calderón Inverse Problem Pro-
gram. Roughly speaking, that program aims at recovering as much information as
possible about a geometric operator (Laplace or Dirac type) from its Cauchy data.
The program seems to be quite successful in two dimensions. In [195] C. Guil-
larmou and L. Tzou identified the connection of a Dirac operator on a Riemann
surface with boundary from the Cauchy data space up to natural gauge transforma-
tion. Many other mathematicians with quite different background (like P. Albin
and G. Uhlmann) work on the program.

9. Real Operators
Up to now, we have considered operators between spaces of sections of complex
vector bundles. One can also consider operators with real coefficients which operate
only on sections of real bundles. If P is such a real elliptic and skew-adjoint
operator, then trivially index P = 0. This is uninteresting, but now dim(Ker P ) is a
homotopy-invariant mod 2. The reason for this stems from the fact that the nonzero
eigenvalues of P all come in complex conjugate pairs (λ, λ̄); if one deforms the
13.10. THE LEFSCHETZ FIXED-POINT FORMULA 357

operator P so that λ goes to zero, then λ̄ goes to zero and consequently dim(Ker P )
increases by two. Already in 1959, R. Bott had discovered a real analogue to his
periodicity theorem (see Chapter 10 above) and proved that the homotopy groups
πi (GL(N, R)), for large N , are periodic in i with period 8, and for i ≡ 0 or i ≡
1 mod 8 are isomorphic to Z.

Results. a) In [46] the connection between these two analytic and topological
mod 2 invariants was determined, and the index theorem was carried over to the real
case, see also the elaboration in [157]. Actually, linear differential operators on real
vector bundles are decisive for M. Furuta’s geometric proof of the Index Theorem,
avoiding the use of pseudo-differential operators. Together, these considerations
provide a new and topologically much simpler proof of the real Bott Periodicity
Theorem. Two details are particularly noteworthy:
b) In order to connect the two invariants of the real theory, one must go outside of
the real theory, since the amplitude p of a real skew self-adjoint differential operator
P is defined via the Fourier transform, and is thus not real in general, but rather
complex with the condition p(x, −ξ) = p(x, ξ). Thus, the symbol of a real operator
does not immediately yield a suitable element of πi (GL(N, R)), but rather, must
be interpreted as a mapping f : S 2n−1 → GL(N, C) (as in Chapter 11) with the
condition f (−ξ) = f (ξ).
c) Another peculiarity lies in the fact that these mod 2 invariants (although having
only the values 0 or 1) in a certain sense are more complicated topologically, or in
any case, of a different type than the usual homology or cohomology classes (also
if one takes Z2 coefficients). Thus, in concrete situations they can provide decisive
additional information. This program was carried out for vector fields (see Section
13.5 above) in [22] and [37]. For a different approach see also the cited [157].

10. The Lefschetz Fixed-Point Formula


Let f : X → X be a continuous map with X compact, and let
Fix(f ) := {x ∈ X : f (x) = x}
denote the fixed-point set of f . Salomon Lefschetz (1926) introduced the for-
mula
X
(13.63) L(f ) = ν(x).
x∈Fix(f )

The Lefschetz number L(f ) on the left side of (13.63) is defined as the alternating
sum (−1)i tr(H i f ), where
P

H i f : H i (X; C) −→ H i (X; C)
denotes the cohomology endomorphism of the complex vector space H i (X; C) in-
duced by f . For details of the definition of the integer ν(x) on the right side of
(13.63) (which is ±1 for an isolated fixed point and equals 0 for a point where
f = Id in a neighborhood) and for the proof, see [11, p.531-542] or [185, p.222-
224]. Atiyah and Bott refined this beautiful formula in the mid 1960’s. Further-
more, the original formula is simplicially defined and in general hardly computable.
In other words, only the topology of X enters into the formula, while additional
structures are ignored. Atiyah and Bott removed this weakness by bringing the
additional structures into play.
358 13. CLASSICAL APPLICATIONS (SURVEY)

Table 13.1. Applications of the Atiyah-Bott-Lefschetz formula (ABL)

Situation X, P, f Weight ν(x) Meaning

1. X oriented, Rieman- L(f,


P P ) = L(f ) =
nian manifold, P := det(Id −f∗ (x)) ν(x) number of
d + δ|Ωev of Section | det(Id −f∗ (x))| fixed-points,
13.4, f : X → X = sign det(· · · ) = ±1 ABL ⊂ Lefschetz
smooth

2. X complex manifold, det(Id −f∗C (x))−1 , In particular, ABL


P := D of Section 13.7, f∗C (x) denotes the contains well-known
f : X → X holomor- complex differential theorems on algebraic
phic (C-linear approxi- functions for the case
mation to f at x); of Riemann surfaces:
note det(Id −f∗ (x)) = [208, pp. 595ff]
2
|det(Id −f∗C (x))|

3. X complex manifold, ABL has corollaries in


P := DV , V holomor- tr ϕ(x) algebraic geometry
phic vector bundle (see det(Id −f∗C (x))
Section 13.7 above),
f : X → X holomor-
phic, ϕ : f ∗ V → V
holomorphic
Yq
4. X oriented, Rieman- iq cot(rj,x ), ABL is a discrete ana-
j=1
nian manifold of di- where Tx X is decom- log of the Hirzebruch
mension 2q, P := D+ posed into a direct Signature Formula
signature operator of sum of 2-dimensional (f = Id). We obtain
Section 13.4, f : X → planes Vj,x on which theorems of this kind:
k
X orientation preserv- f∗ (x) is a rotation by If f is of order p with
ing isometry. rj,x . p a prime 6
= 2 (i.e.,
k
f (p ) = Id), then f
cannot have exactly
one fixed-point, since
then L(f, P ) ∈ Z, but
ν(x) ∈ C \ Z.

Results. Let X be a compact C ∞ manifold without boundary, and let P be


an elliptic differential operator. Let f : X → X be differentiable and commute with
P . Here we need to assume that f lifts to a bundle mapping, say f˜, so that f acts
on sections via (f · s) (x) = f˜(s(f −1 (x))). This is a special case of the equivariant
constructions of the next Section. Then f yields a well defined endomorphism of
the finite-dimensional vector spaces Ker P and Coker P . We define the Atiyah-
Bott-Lefschetz number as
(13.64) L(f, P ) := tr(f |Ker P ) − tr(f |Coker P ).
13.10. THE LEFSCHETZ FIXED-POINT FORMULA 359

Figure 13.8. Transversality of the fixed points of f : X → X

a) If f is the identity, then by definition L(f, P ) = index P , and one can apply the
Atiyah-Singer Index Formula.
b) In the other extreme case, where f has only isolated fixed points with multiplicity
±1, one obtains the formula
X
(13.65) L(f, P ) = ν(x).
x∈Fix(f )

On the right side, we sum over complex numbers ν(x) that depend only on the dif-
ferential f∗ (x) : Tx X → Tx X. The simplicity or transversality of the fixed-point
x means that the endomorphism Id −f∗ (x) is invertible, and the multiplicity ±1
is understood to be the sign of det(Id −f∗ (x)), see Figure 13.8.)In Table 13.1, there
are some applications of the Atiyah-Bott-Lefschetz formula (ABL) (13.65)
with the expressions for the respective values of ν(x).

AS
ABL

0
z

Figure 13.9. The Atiyah-Bott-Lefschetz Formula (ABL) as out-


grow of the Atiyah-Singer Index Theorem (AS)

For further details and the proof of (13.65), we refer to [29], [30], and [31].
Incidentally, the proof is essentially simpler than the proof of the index formula
(AS), and represents a weak version of the heat equation proof reviewed above on
pp.304ff and elaborated below in Chapter 17 in full detail: One considers the zeta
360 13. CLASSICAL APPLICATIONS (SURVEY)

function ζ(z) : z 7→ tr(f ∗ ◦ ∆−z ) whose value at z = 0 is easier to calculate (see


Figure 13.9), since it turns out that this zeta function is holomorphic not only for
<(z) > dim X, but also in all of C because of the assumptions on the fixed points;
see also the following section.

11. Analysis on Symmetric Spaces: The G-equivariant Index Theorem


Let G be a compact Lie group which acts smoothly to the left on a closed
(compact, without boundary) n-manifold
 X, via a C ∞ map ` : G × X → X. One
may think of G = {Id} or G = Id, g, g , ..., g r−1 , where g is a diffeomorphism of
2

order r (i.e., g r = Id). G can also be a more general finite group, or, as said, a
compact Lie group. We write g · x = `g (x) = `(g, x). Let πE : E → X be a C ∞
complex vector bundle over X and suppose that there is a left action of G on E
such that for all g ∈ G and e ∈ E, we have thatπ(g ·e) = g ·π(e) and e 7→ g ·e defines
a linear map Ex → Eg·x . Then πE : E → X is called a G-vector bundle. For
the wider concept of a G-principal bundle see below Definition 15.1 (p.395). For a
section u ∈ C ∞ (E) and a group element  g ∈ G, we have a section g · u ∈ C (E)

−1
defined by (g · u)(x) := g · u(g · x) for x ∈ X. Let πF : F → X be another
G-vector bundle.
Definition 13.21. An operator P : C ∞ (E) → C ∞ (F ) is a G-operator (or
G-invariant operator), if for all g ∈ G, we have P (g · u) = g · P (u).
If P is an elliptic (pseudo-) differential G-operator, then Ker P and Coker P
are preserved by the action of G, so that Ker P and Coker P are not only finite-
dimensional complex vector spaces, but can be regarded as G-modules. The homo-
morphism
G → Iso(Ker P ), given by g 7→ g|Ker P , g ∈ G,
is then a finite-dimensional representation of G, and g 7→ tr(g|Ker P ) is its charac-
ter. Corresponding considerations apply to Coker P .
We say that Ker P and Coker P are representation spaces for G. Recall that
R(G) denotes the Grothendieck ring obtained from the abelian semi-group of
equivalence classes of finite-dimensional representations of G with addition induced
by the direct sum. Tensor product of representations induces a multiplication on
R(G) making it a ring.
Definition 13.22. The index of an elliptic G-operator P : C ∞ (E) → C ∞ (F )
is defined by
indexG P := [Ker P ] − [Coker P ] ∈ R(G).
Moreover, for g ∈ G, we define (as in the Formula (13.64) of the Lefschetz
character in the preceding section) the virtual character
indexg P := tr(g : Ker P → Ker P ) − tr(g : Coker P → Coker P ).
It is an element of the representation ring R(G). If G = {Id}, then R(G) =
K(point) = Z and indexG (P ) = index P . In the general case, R(G) ∼ = K(∗) ∼
= ∗ is
a complicated object and correspondingly, indexG (P ) is a sharper (also homotopy-)
invariant than the integer index(P ).
To formulate the G-index Theorem, we need to define the topological in-
dex of an elliptic G-operator P in terms of its principal symbol, say σ(P ) ∈
C ∞ Hom(π ∗ E, π ∗ F ) . Note that the action of G on X induces an action on

13.11. THE G-EQUIVARIANT INDEX THEOREM 361

T ∗X ∼ = T X (via a G-invariant metric on X). Using the fact that P is a G-


operator, σ(P ) defines an element of [σ(P )] ∈ KG (T X). See also our short re-
view of equivariant K-theory on pp.292ff above. We proceed along the same
lines of the case where G is trivial. One first selects a G-equivariant embedding
f : X → Rn+m , where Rn+m is a representation space for G. The existence of
such an embedding is a consequence of the Peter-Weyl Theorem (see [327]). Al-
though it is more difficult to prove (especially if G is nonabelian), we have a Thom
isomorphism ΨT N →T X : KG (T X) → KG (T N ) and an extension homomorphism
h : KG (T N ) → KG (T Rn+m ). The composition h ◦ ΨT N →T X gives us a homomor-
phism
f! := h ◦ ΨT N →T X : KG (T X) −→ KG (T Rn+m ).
Moreover, for i : {0} → Rn+m , we have (i! )−1 : KG (T Rn+m ) → KG (0) = R(G).
Definition 13.23. For an elliptic G-operator P : C ∞ (E) → C ∞ (F ) with sym-
bol class [σ(P )] ∈ KG (T ∗ X), the topological G-index of P is defined by
indext,G (P ) := (i! )−1 ◦ f! ([σ(P )]) ∈ R(G).
For g ∈ G, the topological g-index of P is defined by

indext,g (P ) := tr (indext,G (P ))(g) .
Results. The proof of the Index Formula (Theorem 12.14) generalizes (see
[44]) without difficulty to yield the following
Theorem 13.24 (G-index Formula). For an elliptic G-operator P : C ∞ (E) →

C (F ) over a closed manifold X, we have indexG P = indext,G (P ).
As when G is trivial there is a cohomological form of indext,G (P ). This is
particularly easy to deduce when G acts trivially on X (i.e., all g act like the

identity). In that case, there is an isomorphism K(X) ⊗ R(G) −→ KG (X) induced
by tensoring bundles over X on which G acts trivially with product G-bundles
X × Vi , where Vi is an irreducible G-module (see [384]). Then we have
chG : KG (X) → H ∗ (X; C) ⊗ R(G) given by ch ⊗ Id on K(X) ⊗ R(G) ∼ = KG (X),
and we also have (using compact supports) chG : KG (T ∗ X) → H ∗ (T ∗ X; C)⊗R(G).
Theorem 13.25. For an elliptic differential G-operator P : C ∞ (E) → C ∞ (F ),
arising from a G-action which is trivial on X, the topological index of P is given
by
indexG P = indext,G (P ) = (−1)n chG ([σ(P )])Td(T X ⊗ C) [T X],


where the symbol class [σ(P )] is regarded as in KG (T ∗ X). Moreover, for g ∈ G,


indexg P = (−1)n tr(chG ([σ(P )])g)Td(T XC) [T X],


where tr denotes the trace of chG ([σ(P )])(g) in the R(G) factor of K(X) ⊗ R(G) ∼
=
KG (X), which results in an element of K(X).
Now suppose that the action of G on X is not trivial. For each g ∈ G, let X g :=
{x ∈ X : g · x = x} denote the set of fixed-points of g. Since there is a metric on X
such that G acts by isometries, it follows that X g is a union of finitely many compact
connected submanifolds of X, say X1 , . . . , Xkg , of possibly different dimensions
d1 , . . . , dkg . For k = 1, . . . kg , let ik : Xk → X denote the inclusion and let Nk → Xk
be the normal bundle of Xk in X. We have (ik )∗ : T Xk → T X and the normal
362 13. CLASSICAL APPLICATIONS (SURVEY)

bundle of T Xk in T X is T Nk → T Xk , which, recall, has a complex structure. We


wish to compute indexg (P ) as in Definition 13.22, and for this we may assume that
G is the cyclic group generated by g. Note that πk : T Nk → T Xk is a possibly
nontrivial G-bundle. Recall that we have the Thom element λT Nk ∈ KG (T Nk )
which provides the Thom isomorphism Ψk : KG (T Xk ) → KG (T Nk ) via Ψk (a) :=
πk∗ (a)λT Nk . Then the extension homomorphism hk : KG (T Nk ) → KG (T X) yields
(ik)! := hk ◦ Ψk : KG (T Xk ) → KG (T X).
Theorem 13.26 (Atiyah-Segal-Singer Fixed-Point Formula). In the above no-
tation, we have


d−k chg (ik∗ ) [σ(P )]
Xkg  
indexg (P ) = (−1) Td(T Xk ⊗ C) [T Xk ],
k=1 chg (λT Nk )
where the quotient has meaning in the context of the localization of the ring R(G)
at g since the trace of g in the representation defining λT Nk is nonzero.
For a more thorough discussion of the proof, see [42], [273, pp.259f] and [389,
pp.120f]. The references [31, 32], [273, pp.259f] and [389, pp.120f] also contain
the major instances of Theorem 13.26 obtained by using various standard elliptic
operators P , see Table 13.2.

Table 13.2. Major instances of the Atiyah-Segal-Singer fixed-


point formula

Elliptic Operator P Corresponding G-Theorem

d + δ : Ωev (X) → Ωodd (X) Lefschetz Fixed-Point Theorem

d + δ : Ω+ (X) → Ω− (X) G-Signature Theorem

∂ + ∂ : Ω0,ev (X) → Ω0,odd (X) Holomorphic Lefschetz Theorem

D± : C ∞ (Σ± (X)) → C ∞ (Σ∓ (X)) G-Spin Theorem

12. Further Applications


With the preceding overview, we have in no way encompassed all of the con-
nections, alternative formulations, generalizations and special applications of the
Atiyah-Singer Index Formula, but we have only touched those which have reached a
certain definitive form in their development. In the following Part IV (Chapters 16
and 18), we explain highlights of the ongoing work on applications in quantum field
theory and low-dimensional topology. In Chapter 14, we summarize the physical
motivation. In Chapter 15, we explain the required concepts of differential geometry
in detail. In Chapter 17 we harvest a full independent proof of the index theorem
for classical geometric operators as a side benefit of the geometric preliminaries
learned. For an insight into further perspectives and (mostly still) open problems,
we refer to the classic [394], [396], and [25] and the literature given there.
Part IV

Index Theory in Physics and the


Local Index Theorem

The beauty and profundity of the


geometry of fibre bundles were to
a large extent brought forth by
the (early) work of the man (S.-S.
Chern) who we are here to honor
today. I must admit, however, that
the appreciation of this beauty
came to physicists only in recent
years.

C.N. Yang, Chern Symposium 1979


[Berkeley, June], Eds. W.-Y. Hsiang
et al., Springer-Verlag, 1980, p. 252

363
CHAPTER 14

Physical Motivation and Overview

Synopsis. Mode of Reasoning in Physics. String Theory and Quantum Gravity.


The Experimental Side. Classical Field Theory: Newton-Maxwell-Lorentz, Faraday 2-
Form, Abstract Flat Minkowski Space-Time, Relativistic Mass, Relativistic Kinetic En-
ergy, Inertial System, Lorentz Transformations and Poincaré Group, Relativistic Deviation
from Flatness, Twin Paradox, Variational Principles. Kaluza-Klein Theory: Simultane-
ous Geometrization of Electro-Magnetism and Gravity, Other Grand Unified Theories,
String Theory. Quantum Theory: Photo-Electric Effect, Atomic Spectra, Quantizing En-
ergy, State Spaces of Systems of Particles, Basic Interpretive Assumptions. Heisenberg
Uncertainty Principle. Evolution with Time — The Schrödinger Picture. Nonrelativis-
tic Schrödinger Equation and Atomic Phenomena. Minimal Replacement and Covariant
Differentiation. Anti-Particles and Negative-Energy States. Unreasonable Success of Stan-
dard Model. Dirac Operator vs. Klein-Gordon Equation. Feynman Diagrams.

I Both the goals and methods of physics are different from those of mathematics.
Mathematicians have the rather nebulous goal of exploring and establishing that which is
logically possible and interesting, depending on the fashions of the time and the tastes of
the individual. Physicists have the sharper goals of discovering, explaining and predicting
actual phenomena in the physical world. The mode of reasoning in physics is rather
fuzzy by mathematical standards, as it is often partly based on conventional wisdom
and folklore rather than clear axioms. However, this reasoning is of great value if it
provides a satisfying explanation of experimental data and makes promising, testable
predictions. If physicists were forced to be mathematically rigorous every step of the way,
physics would not have advanced toward its goals nearly as much as it has. Although
the creative process in mathematics is generally fuzzy in its initial phases, a result is
not usually publishable until it has been proven within a quite definite framework of
commonly accepted logical standards, which goes far beyond the notion of reasonable
doubt in a court of law. Mathematicians are uneasy with many of the heuristic arguments
used by physicists sometimes involving manipulations of expressions which have not been
shown to exist (e.g., path integrals, infinite renormalizations, nonconvergent series, etc.).
On the other hand, physicists cannot be expected to have interest in mathematics that
seems unrelated to physical phenomena.
Since index theory was developed as a mathematical achievement, there has emerged
a prominent group of theoretical physicists who appear to be somewhat unconventional,
namely the string theorists (and we may include ‘supersymmetrists’ and ‘quantum gravi-
tationists’ as well, see the anthology [75]). A taste of the recent revival of D-branes and
other exotic instantons in string theory can be gained from [164] of H. Ghorbani, D.
Musso and A. Lerda. Indications can be found in the review [361] of F. Sannino about,
how strongly coupled theories of gauge theoretic physics result in perceiving a composite
universe and other new physics awaiting to be discovered. It is clear that most string
theorists believe that what they are doing is of physical relevance, but as yet no direct
experimental confirmation has emerged. What they have certainly uncovered is a truly

364
14.1. CLASSICAL FIELD THEORY 365

awesome body of mathematics that has had a big positive impact on purely mathemati-
cal research in neighboring areas. Many mathematicians envy the mathematical insights
that string theorists have had. Indeed, Seiberg-Witten Theory is but a small portion of
the mathematics inspired largely by the insights of Edward Witten, one of the leading
string theorists. However, one characteristic of a physical theory that conventional physi-
cists deem essential is that the theory be testable by experiment. Currently, the physical
refutation of string theory seems just as remote as its confirmation. Based on this, the
conventional physicist can argue that string theory (or quantum gravity) is not really a
physical theory at all. This is not because it is false, but because it is not falsifiable, in
the sense that it seems unlikely that it can be proven or disproved experimentally in the
foreseeable future. José Gracia-Bondı́a [183, p.6], e.g., emphasizes that masses and
energies on our planet are much too small to make a difference for possible falsifications
of common ideas of string theory and quantum gravity. However,when Giampiero Es-
posito in [145, Section 8.1] addresses the experimental side of quantum gravity, one of
his points is the immense capacity of modern computer supported and partly space based
astronomy, which gives access to data involving previously unimaginable large masses and
energies. See also the recent Nobel citations in physics for a nontechnical view on the new
observational capacities. Moreover, as noted by Bryce DeWitt in [119, p.417], string
theory provides (in some cases) substantially simplified schemes and diagrams for basic
calculations. His example is the replacement of four different Feynman diagrams by a
single one in string theory: a thing “that, from a nonspecialists point of view, make it
look rather pretty”.
If it were suddenly found that string theory has no physical relevance, most likely only
a handful of string theorists would remain, namely those who really consider themselves
to be primarily mathematicians.
There are such mathematicians (misnamed mathematical physicists) who are inter-
ested in strict mathematics that seems to have physical relevance or is motivated by
physical considerations. In the overview that follows, it is hoped that the reader may gain
some understanding of why many concepts in this book (e.g., elliptic operators, complex
vector bundles, pseudo-differential operators, Hilbert spaces, distributions, etc.) may have
great physical relevance and how in large part they were initially motivated by physical
considerations. Of course, mathematics being motivated by physics is not a new phenom-
enon, but rather an old one. There was hardly any distinction between mathematics and
theoretical physics before the 1900s. The dubious mid-twentieth century goal of attaining
ivory purity in mathematics, devoid of any hint of lowly physical application, seems to
have been largely temporary, although many practitioners remain.
The reader is not expected to understand every detail in the following lengthy (yet
necessarily incomplete and historically vague) overview of quantum field theories in mod-
ern physics. However, she or he may take whatever is digestible, realizing that this material
is neither a prerequisite nor a substitute for the more precise (if drier) mathematics of the
chapters that follow. Those who are unfamiliar with relativity or quantum physics are
likely to discover that the logical possibilities of the world of physics can be every bit as
beautiful and strange as those encountered in far-reaching mathematical diversions. J

1. Classical Field Theory


In classical (as opposed to quantum) physics, particles are viewed as point-like
objects that move along paths which are solution curves of systems of ordinary
differential equations determined by a force field. For example, there is Newton’s
equation mr00 (t) = F(r(t)) , where F : R3 → R3 is a given force field. More
generally, the force may also depend on the velocity r0 (t) as well as r(t). Indeed, an
electromagnetic (E-M) field consists of a pair of vector fields E and B (which may
366 14. PHYSICAL MOTIVATION AND OVERVIEW

be time-dependent). A test particle of charge e moves according to the Lorentz


force law
d e
(14.1) (mr0 (t)) = eE(r(t) , t) + r0 (t) × B(r(t) , t) .
dt c
The situation is complicated not only by the fact that a real (not test) particle
contributes to E and B, but also by the fact that E and B satisfy a system of
partial differential equations, namely Maxwell’s equations

1 ∂B
(1) ∇ × E + =0 (2) ∇ · B = 0
c ∂t
1 ∂E 1
(3) ∇ · E = ρ (4) ∇ × B − = J.
c ∂t c
Here c denotes the speed of light, ρ : R3 → R is proportional to the charge density of
a continuous medium of charged particles, and J is essentially the current density of
the medium (i.e., J =ρv, where v is the velocity vector field of the medium). Thus,
E and B are influenced by each other, as well as the by the motions of the charged
medium that they are supposed to influence via the Lorentz force law (14.1).
Maxwell’s equations can be simplified conceptually by considering the following
2-form, called the E-M field strength or Faraday 2-form

(14.2) F := cE1 dx ∧ dt + cE2 dy ∧ dt + cE3 dz ∧ dt


+ B1 dy ∧ dz + B2 dz ∧ dx + B3 dx ∧ dy.

Exercise 14.1. (a) Check that Maxwell’s equations (1) and (2) are equivalent
to dF = 0.
(b) Defining the source 1-form j by

j = ρdt − c−2 (J1 dx + J2 dy + J3 dz) ,

verify that the Maxwell equations (3) and (4) say that δF = j, where δ is the
codifferential (the formal adjoint of d) relative to the Lorentz-Minkowski metric
c2 dt2 − dx2 − dy 2 − dz 2 .

This exercise implies that Maxwell’s equations can be immediately generalized


to arbitrary 4-manifolds with Lorentz metric tensors (i.e., space-times). Hence,
Maxwell’s equations fit quite naturally into general relativity, although historically,
relativity was built around Maxwell’s equations. The Lorentz force law (14.1) also
can be written invariantly on a space-time M . Indeed, the world line of a test
particle of rest mass m0 and charge e is a curve s 7→ γ(s) ∈ M which obeys the
equation
D 0 e #
(14.3) m0 γ (s) = F (γ 0 (s) , ·) ,
ds c
where γ 0 denotes the tangent vector field of γ, ds D
denotes covariant differentiation
along γ (i.e., ∇γ 0 ) and the sharp “#” on the right side indicates that the covector
F (γ 0 (s) , ·) has been converted to a vector by raising indices using the metric. In
flat Minkowski space, ds D
(γ 0 (s)) is simply γ 00 (s), but s is not necessarily the time
coordinate. Rather think of s as the arc length as an inherent parametrization.
14.1. CLASSICAL FIELD THEORY 367

Equation (14.3) implies that the length of γ 0 (s) is constant. Indeed,


 
0 2 D 0 0 eD #
E
1 d
2 m0 ds |γ (s)| = m0 γ (s) , γ (s) = F (γ 0 (s) , ·) , γ 0 (s)
ds c
e
(14.4) = F (γ 0 (s) , γ 0 (s)) = 0,
c
since F is anti-symmetric. Note that (14.3) is not scale invariant; i.e., if γ is
replaced by γa where γa (s) := γ(as), then γa is not necessarily a solution of (14.3)
for a 6= 1. In Minkowski space with the metric c2 dt2 − dx2 − dy 2 − dz 2 , equation
(14.3) splits into spatial and temporal components which are empirically correct
2
only when |γ 0 (s)| = c2 . Indeed,
γ(s) = (t(s) , x(s) , y(s) , z(s)) =: (t(s) , r(s))
!
2
0 2 2 0 2 0 2 2 0 2 −2 r0 (s)
=⇒ |γ (s)| = c t (s) − |r (s)| = c t (s) 1−c
t0 (s)
!
2
2 0 2 |v(t(s))|
= c t (s) 1− ,
c2

where v(t) = d
dt r(s(t)) assuming that t0 (s) > 0, so that t = t(s) can be inverted.
Then
!− 1
2 2
2 |v(t(s))|
(14.5) |γ 0 (s)| = c2 ⇐⇒ t0 (s) = 1− =: β(s) ,
c2
and γ 0 (s) = (β(s) , β(s) v(t(s))).
Exercise 14.2. Check that (14.3) splits into the pair of equations
 v 
d d
m0 βc2 = eE · v.

(a) dt (m0 βv) = e E + × B (b) dt
c
Note that (a) is the Lorentz force law (14.1), where m = m0 β is the so-called
relativistic mass (m ≈ m0 for |v|  c, and m → +∞ as |v| ↑ c). The right side
of (b) is the rate at which the E-M field does work on the particle; note that B
does no work since (v × B) · v = 0. Thus, mc2 = m0 βc2 on the left side of (b) must
be the energy E of the particle (i.e., E = mc2 ). Note that m0 c2 is the rest energy
and the relativistic kinetic energy is
 
2 4
mc2 − m0 c2 = m0 c2 (β − 1) = 21 m0 |v| + c2 O (|v| /c) as |v| /c → 0.

Abstract Minkowski space-time consists of a four-dimensional vector space


(or more precisely, affine space) M with scalar product h·, ·i of signature (+, −, −, −).
By translation and the usual identifications, h·, ·i determines a scalar product on
the tangent space at each point of M . A coordinate system (t, r) := (t, x, y, z) on
M is called an inertial system if h·, ·i = c2 dt2 − dr2 := c2 dt2 − dx2 − dy 2 − dz 2 . If
(t̄, r̄) := (t̄, x̄, ȳ, z̄) is another inertial system, then there is a linear transformation
L : R4 → R4 and a point (t0 , r0 ) ∈ R4 , such that
(t̄, r̄) = L((t, r)) + (t0 , r0 ).
The fact that c2 dt̄2 −dr̄2 = h·, ·i = c2 dt2 −dr2 places a restriction on L. It preserves
the scalar product with diagonal matrix I1,3 having diagonal entries 1, −1, −1, −1,
T
in the sense that [L] I1,3 [L] = I1,3 , where [L] is the matrix of L relative to the
368 14. PHYSICAL MOTIVATION AND OVERVIEW


standard basis of R4 . Such L ∈ GL R4 are known as Lorentz transformations
and comprise the Lorentz group O(1, 3). The Lorentz transformations together
with the translations of R4 , generate the Poincaré group. Even if inertial sys-
tems (t̄, r̄) and (t, r) are based at the same point O ∈ M, we do not necessarily
have the equality t̄ = t of time coordinate functions (i.e., coordinate time has no
absolute meaning). On the other hand, if γ : (a, b) → M is a smooth curve, the
2
condition |γ 0 (s)| := hγ 0 (s) , γ 0 (s)i = c2 does have invariant meaning. The correct
interpretation is that s represents the time on a clock carried by the particle with
world line γ. Equation (14.5) tells us that t0 (s) = β ≥ 1 meaning that coordinate
time in the inertial system (t, r) generally runs faster than the proper time of a
particle which is moving relative to this inertial system. Assuming that the earth
does not deviate from the t-axis of some inertial system in a nearly Minkowskian
space-time, a high velocity space traveler with world line γ will find his earth-bound
twin is older when he returns. The earth-bound twin ages according to his proper
time which coincides with coordinate time t, while the space traveler ages according
2
to his proper time, namely s, where |γ 0 (s)| = c2 . Thus, assuming that their clocks
are synchronized just before departure at t = s = 0, we then have
 − 1
0 2 2 2
t (s) = 1 − |v| /c > 1 =⇒ t > s.

One might argue that by symmetry, we should also have that s > t, but the situation
is not symmetric, since the world line of the space traveler is not close to the time
axis of an inertial system, because his acceleration is considerable (i.e., his world
line is not nearly straight). Thus, while we have the so-called twin paradox, there
is no contradiction. The result has been confirmed by experiment using particles
instead of humans.
Things are complicated by the fact the metric tensor gµν on a realistic space-
time is not flat like that of Minkowski space. The deviation from flatness is due
to the presence of E-M fields (i.e., radiation), particles (neutral and charged) and
gravity waves that can propagate in space-time even in the absence of radiation
and matter. One measure of curvature is the symmetric Ricci curvature tensor
Rµν (where µ, ν = 0, 1, 2, 3) to be defined later (see 15.61). The vanishing of the
Ricci tensor is necessary (but not sufficient) in order that a space-time be locally
isometric to Minkowski space. The scalar curvature is the trace S = g µν Rµν ,
where we automatically sum over repeated indices on different levels (the Einstein
convention). The Einstein field equation (10 scalar equations) is
−8πK
(14.6) Rµν − 12 Sgµν = Tµν ,
c2
where K is the universal gravitational constant and Tµν is the symmetric stress-
energy-momentum tensor which is formed in a canonical way from the E-M field and
a continuous approximation of the energy-momentum density of particle-like matter
(cosmologists sometimes take these particles to be entire galaxies). In essence,
the Einstein field equation (14.6) tells us how the nongravitational stress-energy-
momentum Tµν of radiation and matter influences the curvature of space-time.
2
Neutral particles move along geodesics γ(s) of space-time such that |γ 0 (s)| = c2 .
The apparent curvature of such geodesics when projected onto what we perceive as
space, is due to gravity which is just the geometry of space-time.
14.1. CLASSICAL FIELD THEORY 369

The vanishing of the Ricci tensor (and hence the scalar curvature) does not
imply that space-time is locally flat. Indeed, the full curvature tensor has ten ad-
ditional components that constitute the Weyl conformal curvature tensor (defined
in Section 15.5). Thus, it is quite possible to have a curved space-time which is
devoid of matter and radiation (i.e., Tµν = 0) which satisfies the so-called empty
space equation Rµν − 21 Sgµν = 0. This equation can be formulated in terms of a
variational principle. Indeed, let D be a compact domain in a space-time M with
metric tensor g. Let L be the functional that assigns to each metric tensor g 0 , the
quantity
Z
(14.7) LD (g 0 ) := S(g 0 ) µg0 ,
D
where S(g ) denotes the scalar curvature of g 0 and µg0 is its volume element.
0

Remark 14.3. When we write about general relativity, strictly speaking, there
should be an additional term in the preceding Einstein action functional in form
of a surface integral of the trace K of the extrinsic curvature over the boundary
∂D of D to cancel the second derivatives coming from the scalar curvature. This
is a standard procedure now. If we do not add it, then the variational problem
is not well defined. This term was introduced in [165]. Moreover, it may be
necessary to add the cosmological constant Λ. It seems that nowadays everybody
believes
R that Λ R6= 0. That is, the action of general relativity in (14.7) should be
D
(S − 2Λ) + 2 ∂D K. (We are indebted to I. Avramidi for this remark.)
It is found (e.g., see [59, p. 125] for a coordinate-free proof) that g is a critical
point of L within the space of those g 0 agreeing with g on the boundary of D, if and
only if g satisfies the empty space equation Rµν − 12 Sgµν = 0 in D. In order to obtain
the full equation (14.6) including Tµν , it is necessary to add additional terms to L
for each type of nongravitational particle or field that resides in space-time. These
terms are known as actions or Lagrangians (L itself is the purely gravitational
Lagrangian). The action over D for an E-M field F = Fµν dxµ ∧ dxν (see (14.2))
2
for a fixed metric g is proportional to − 21 D |F |g µg , where
R

2
|F |g := 12 g µν g ρσ Fµρ Fνσ
(i.e., the standard Lorentz-invariant norm-square relative to the metric tensor g).
In the special case of Minkowski space with flat metric g, writing F in terms of
inertial coordinates as in (14.2), we have
 
2 2 2
− 12 |F |g = 12 |E| − |B| .
2 2
Under a change of inertial coordinate system on Minkowski space |E| and |B|
2 2
can change, but |E| − |B| is invariant. The 2-form F is a covariant object, but
(as with coordinate time) E and B separately have no absolute significance. The
combined Lagrangian over D is
Z Z
2
(14.8) LD (g, F ) := S(g) µg − k2 |F |g µg .
D D
For an arbitrary covariant symmetric 2-tensor h, we have the following for the
partial directional derivative of L(g, F ) at g in the direction h
Z  
σ 2
d
dt L(g + th, F ) t=0
= −R µν + 1
2 Sg µν + kF µ Fνσ − k
4 |F | g gµν hµν µg .
D
370 14. PHYSICAL MOTIVATION AND OVERVIEW

Thus, g is a critical point for L(g, F ) in the sense that this directional derivative is
0 for all h, when
 
2
(14.9) Rµν − 12 Sgµν = k Fµ σ Fνσ − 14 |F |g gµν .
For suitable k, depending on the choice of units, this is the Einstein field equation
(14.6) for a space-time with E-M radiation. The right side is indeed propor-
tional to the accepted stress-energy-momentum tensor for the E-M field F .
We can also get Maxwell’s equation δF = 0 from a variational principle;
note
R that j = 0 in the absence of sources which we assume here. Indeed, consider
1 2
2 D |F |g µg as a functional of F , instead of g. We assume that F and its variations
satisfy the other Maxwell equation dF = 0. Assuming that D is simply-connected,
we can write any
R variation of F 0 of F as dA0 for some 1-form A0 . The associated
1 2
variation of 2 D |F |g µg is the directional derivative
 Z  Z Z Z
0 2 0 0
d
dt 2
1
|F + tF |g µg = hF , F i µg = hdA , F i µg = hA0 , δF i µg ,
D t=0 D D D
0
assuming that A vanishes on the boundary of D. This variation is 0 for all such
A0 exactly when δF = 0 in D. In summary, the vanishing of the first variations
of L(g, F ) with respect to g and F are Einstein’s equation (14.9) and Maxwell’s
equation δF = 0, respectively.
The Maxwell equation dF = 0 implies that F can be written locally as F =
−dA, where A is a 1-form known to physicists as the 4-vector potential, and
the minus sign stems from the fact that in mechanics forces generally act in the
direction opposite the gradient of the potential energy. Such an A (satisfying F =
−dA) exists on any simply-connected domain where dF = 0 and is called a gauge
potential for F . However, A is not unique, since for any function ϕ ∈ C ∞ (M ),
F = −dA = −d(A + dϕ), whence A + dϕ also serves as a gauge potential for F .
The transformation A 7→ A+dϕ is called a gauge transformation. In terms of A,
the equation δF = 0 becomes the wave equation −δdA = 0; in Minkowski space,
−δd = c−2 ∂t2 − ∂x2 − ∂y2 − ∂z2 . It is convenient to regard A as more fundamental than
F, since dF = 0 follows immediately from F = −dA, and from the wave equation
δdA = 0, we see that singularities of A propagate with speed c. However, the fact
that A is not uniquely determined by F is somewhat of a drawback. To alleviate
this, the so-called Lorentz condition δA = 0 is sometimes imposed on A, but
there are generally plenty of functions ϕ for which δ(A + dϕ) = 0, namely solutions
ϕ of the scalar wave equation δdϕ = 0.
Geometrical Unification for the Simplest Kaluza-Klein Theory. The
E-M field strength F is not built into the metric gµν , and Einstein spent many
years trying to incorporate F into the geometry of space-time, thereby obtaining
a unified field theory. Actually, in [239] and [244], it was shown that E-M and
gravity could be geometrized simultaneously by forming a 5-dimensional manifold
by attaching circles to the points of space-time. In this remarkable Kaluza-Klein
theory, the illusion that the universe has only four space-time dimensions is not
necessarily due to the smallness of the circles (although the theory predicts that
they are very small), but rather it is due to the perfect homogeneity in the circular
direction. Off hand, it is difficult to believe that anything useful can come from
adding an unobservable dimension, but indeed all grand unified theories (GUTs)
add at least 24 dimensions in order to unify all known nongravitational forces. The
14.1. CLASSICAL FIELD THEORY 371

strings in string theory live in a 10 or 26 dimensional space-time (depending on


whether supersymmetry is incorporated or not), not even counting the gauge di-
mensions. For the remainder of this section we describe the geometrical unification
that occurs for the simplest Kaluza-Klein theory.
In modern terminology (made precise in Section 15.1, the original Kaluza-Klein
theory introduces a fiber bundle π : P → M over space-time M , whose fibers
π −1 (x) are circles. The pull-back π ∗ g of the metric g on M is degenerate on the
fibers. Thus, we need additional structures to complete π ∗ g to a Lorentzian metric
on P . While we do not assume that each fiber π −1 (x) is explicitly identified with
a standard circle, we do suppose that there is a notion of what it means for a point
p ∈ P to be rotated through an angle θ along its fiber, say p 7→ Rθ (p). We let ∂θ
be the vector field given by
d
(∂θ )p = dθ Rθ (p) θ=0 .
Of course, there is a standard metric on the fibers which gives ∂θ fixed length
throughout P , but we need to specify what vectors in T P should be orthogonal to
the fibers. This is conveniently accomplished by introducing a 1-form à on P , such
that Ã(∂θ ) = 1 and Rθ∗ à = Ã, so that the distribution of kernels of à is preserved
by Rθ . A nondegenerate metric g̃ on P is then given (for X, Y ∈ Tp P ) by

g̃(X, Y ) = (π ∗ g)(X, Y ) + k Ã(X) Ã(Y ) = g(π∗ X, π∗ Y ) + k Ã(X) Ã(Y ) ,

where k > 0 is some constant to be determined (i.e., g̃ = π ∗ g + k à ⊗ Ã). The


subspace of Tp P which is orthogonal to the fiber through p (i.e., orthogonal to
(∂θ )p ) is then the kernel of Ãp . Note that the rotation Rθ : P → P is an isometry
of (P, g̃) since
     

Rθ∗ g̃ = Rθ∗ π ∗ g + kRθ∗ Ã ⊗ Ã = (π ◦ Rθ ) g + k Rθ∗ Ã ⊗ Rθ∗ Ã
= π ∗ g + k à ⊗ Ã,
where we have used the fact that π ◦ Rθ = π (i.e., Rθ preserves fibers  setwise),
and Rθ∗ à = Ã. Note that à can be recovered from g̃ by taking Ker Ãp to be the
subspace of Tp P which is g̃-orthogonal to ∂θ , and Ãp (∂θ ) = 1.
We now come to the main reasons for introducing this 5-dimensional cylin-
drical universe (P, g̃) with its associated 1-form à derived from g̃ and the circle
action Rθ . The first very striking fact (e.g., see [59]) is that if γ is a geodesic in
P relative to g̃, with Ã(γ 0 ) 6= 0, then the projection γ̄ := π ◦ γ of γ onto M is
the path of a charged particle subject to an E-M field. The Faraday 2-form F of
this E-M field is the unique 2-form on M such that π ∗ F = −dÃ. The existence of
such F is in part a consequence of the invariance of A under pull back by Rθ . The
charge/mass ratio of the particle is proportional to Ã(γ 0 ), or equivalently g̃(γ 0 , ∂θ ),
i.e., essentially the vertical component of γ 0 relative to g̃. If Ã(γ 0 ) = 0, then γ̄ is
the path of neutral particle in (M, g), i.e., a space-time geodesic. The fact that ∂θ
is a vector field generated by a 1-parameter group of isometries Rθ implies that the
charge/mass ratio g̃(γ 0 (s) , ∂θ ) is constant, independent of s, as it should be. The
equation π ∗ F = −dà suggests that à is related to potential 1-forms of F . Indeed,
suppose that there is an open set U ⊆ M and a map σ : U → P , such that π ◦ σ = I
(i.e., σ is a local section of the bundle π : P → M ). Then Aσ := σ ∗ Ã is locally a
372 14. PHYSICAL MOTIVATION AND OVERVIEW

potential 1-form of F , since on U we have


 

−dAσ = −d σ ∗ à = −σ ∗ dà = σ ∗ π ∗ F = (π ◦ σ) F = F.

Exercise 14.4. Suppose that σ 0 : U → P is another local section, say σ 0 (x) =


Rϕ(x) (σ(x)), for some function ϕ ∈ C ∞ (U ). Verify that
Aσ0 := σ 0∗ Ã = σ ∗ Ã + dϕ = Aσ + dϕ.
Thus, Aσ0 is related to Aσ by a gauge transformation.
Remark 14.5. Nowadays many refer to the transformation p 7→ Rϕ(π(p)) (p) of
π −1 (U ) as a gauge transformation, and the transformation Aσ 7→ Aσ0 = Aσ + dϕ is
induced by it. Moreover, differential geometers refer to the invariant 1-form à on
P as a connection 1-form. We will develop these notions much more systematically
in the next chapter.
The crucial point here is that E-M forces and gravity are simultaneously en-
coded in the metric g̃ on P , thereby achieving a geometrical unification of these
forces. It is interesting to observe that the addition of an extra dimension to space-
time to geometrize E-M is very much in the same spirit that time was adjoined to
space in order to geometrize gravity in general relativity.
Even by itself, the fact that the geodesics of (P, g̃) project to paths of charged
particles would be sufficient to take the 5-dimensional Kaluza-Klein theory seriously,
but there is yet another surprise. The scalar curvature S(g̃) of the metric g̃ is
constant on each fiber and thus projects to a well-defined function on M, still
denoted by S(g̃). However, S(g̃) is not just the scalar curvature S(g) of g, but
rather it is given (e.g., see [59]) by
2
S(g̃) = S(g) − 21 k |F |g .
Consequently, the combined Lagrangian of gravity and E-M (see (14.8)) can then
be written simply as
Z Z Z
2
LD (g, F ) = S(g) µg − k2 |F |g µg = S(g̃) µg .
D D D

Thus, the scalar curvature S(g̃) of P yields the combined Lagrangian. Roughly put,
the Einstein field equation in a nonempty universe with an E-M field (but no matter)
is obtained from an empty bundle universe, in the sense that the E-M stress-
energy-momentum source is encoded in the geometry of the metric g̃. All of this
admits suitable generalization to the case where the fibers are not just circles, but
rather general Lie groups (typically SU(N ) , SO(N ) or products of these in physical
applications). The 1-forms on these higher dimensional bundles are Lie-algebra-
valued connection 1-forms which physicists call gauge potentials when they
are pulled down to M via a local section. The corresponding field strengths
(known as curvatures to differential geometers) are no longer R-valued 2-forms
such as the Faraday F , but rather they have values in certain vector bundles over M .
There is the rather obvious hope that these field strengths describe the other forces.
For example there are the weak forces that cause, among other events, the decay
of the neutron; and the strong forces that are indirectly responsible for holding
the nucleus together and directly responsible for binding quarks together inside
individual hadrons such as the proton, neutron, pions, etc.. However, one must be
14.2. QUANTUM THEORY 373

wary about extrapolating classical field theory (which is all we have discussed up
to this point) to such small systems which are governed by quantum theory.

2. Quantum Theory
I The following account of quantum field theory (QFT ) is quite condensed. The
interested reader will find many quite recent monographs on QFT in libraries and book
stores with many details. We particularly recommend [431, 432, 433], but also [13, 50,
390, 454, 455, 456] .
Just as Newtonian mechanics breaks down for systems moving at high speeds near
that of light, classical field theory does not describe systems of atomic dimensions or
smaller very well. The classical picture told us that there are diffuse and wave-like back-
ground fields such as the electromagnetic (E-M) field F and the metric tensor g of general
relativity, and in sharp distinction to these there were point-like particles that move in
trajectories determined by these fields, as well as influencing them. However, the reader
has no doubt heard that under certain conditions, light (E-M radiation) produces results
that are better understood by assuming that it is made of a stream of particles, known as
photons. Notably, when light falls on certain metallic surfaces in a vacuum electrons are
emitted from the atoms at a rate which can typically be billions of times larger than the
rate that is calculated under the assumption that each atom absorbs all of the energy it
receives from the continuous E-M wave that contacts it. The most natural explanation of
this photoelectric effect, is that the E-M wave is not continuous, but rather it is made of
chunks (quanta) that have sufficient energy (depending on the wave length) to immedi-
ately dislodge the electrons from the atoms they come in contact with. It was Einstein
who was awarded a Nobel Prize in 1922 in part for his explanation of the photo-electric
effect in terms of the quantum theory of light. However, he veered away from the dramatic
developments in quantum mechanics, preferring to work on unifying the classical field the-
ories of E-M and gravity without adding an extra dimension as in the Kaluza-Klein theory.
He did not succeed. Just as electro-magnetic fields exhibit particle-like properties, it was
also discovered that particles (e.g., electrons) exhibit wave-like properties. In an experi-
ment where electrons are fired at a double-slit they collectively make a diffraction pattern
of impacts on a screen behind the slit. Thus, the sharp particle versus wave dichotomy in
classical physics must admit some fuzziness. J

Energy Levels and Eigenvalues. Quantum mechanics and quantum field


theory grew out of the attempt to describe this state of affairs and to make predic-
tions as accurately as possible. One of the most perplexing phenomena confronting
the founders of quantum mechanics was that of atomic spectra. A spectrograph
reveals that atoms emit and absorb light at fairly discrete wave-lengths or energies.
The classical planetary model for the hydrogen atom has the electron circling the
proton under the inverse square Coulomb law. It predicts that the electron will
radiate E-M energy at a continuously increasing rate and will actually spiral into
the proton as it gives up its energy in a short time, rendering the atom unstable.
Although it first seems speculative, one might hypothesize that the various energy
levels of the atom are actually eigenvalues of some differential operator, just as the
d2
frequencies of a vibrating string are the eigenvalues of a constant multiple of dx 2

acting on the space of functions vanishing at the ends. If the eigenvalues of the op-
erator are to represent energies, the operator should have the physical dimensions
of energy. The Coulomb potential energy (due to the charge of the nucleus) of an
electron at distance r to the nucleus of a hydrogenic atom (or ion) with Z protons is
−Ze2 /r, where e is the proportional to the charge of the electron, depending on the
374 14. PHYSICAL MOTIVATION AND OVERVIEW

system of units. Moreover, the simplest rotationally invariant differential operator


on R3 is the Laplacian ∆ = ∂x2 +∂y2 +∂z2 . The most obvious operator, having dimen-
sions of energy, formed from −Ze2 /r and ∆ is α∆−Ze2 /r where α is a constant with
dimensions of energy times length2 . The point spectrum (assuming α < 0) of this
operator (densely defined on L2 R3 ) is found to be Z 2 e4 /(4αn2 ), n = 1, 2, 3, . . . .
For hydrogenic atoms, there is a choice of α such that this spectrum is consis-
tent with the observed energy spectrum. (Note that the observed spectral lines are
at energies which are actually differences of the above eigenvalues, as the electron
jumps between the possible energy levels.) In order to describe α, let m be the
mass of the electron and let M be the much larger mass of the nucleus. We define
the reduced mass of the system to be µ = m(1 + m/M ) ≈ m. The experimentally
~2
suitable value for α is found to be − 2µ h
, where ~ = 2π ≈ 6.6256 × 10−27 erg. sec.,
and h is Planck’s constant introduced by Max Planck around 1900 in connection
~2
with black body radiation. One can easily check that 2µ has units of energy times
2
length2 so that − 2µ
~
∆ has units of energy. Hence we have the striking coincidence
that the point spectrum of the operator
~2
(14.10) Ê := −∆ − Ze2 /r

coincides to good approximation with the energy levels
Z 2 e4 Z 2 e4 µZ 2 e4
(14.11) = − 2 = −
4αn2 ~
4 2µ n2 2~2 n2
of a hydrogenic atom with Z protons.
Quantization Procedures. A procedure, by which one replaces a classical
observable A (e.g., energy, momentum, position) by an operator  whose spectrum
ranges over all possible experimentally observed values of the observable, is known
as a quantization. A systematic review of various quantization procedures is given
in [355]. Thus, we have roughly succeeded in quantizing the energy, say E, of an
electron in a Coulomb potential, by replacing this energy by the operator Ê (14.10).
Note that
~2 ~2 2 1  2 2 2

∂x + ∂y2 + ∂z2 =

− ∆=− (±i~∂x ) +(±i~∂y ) +(±i~∂z )
2µ 2µ 2µ
1

resembles the classical expression 2µ p2x + p2y + p2z for the kinetic energy of an
object of mass m and momentum p = px i + py j + pz k. This suggests that the
quantization of the classical observable p should be (where the minus sign is con-
ventional)
p̂ = p̂x i + p̂y j + p̂z k := (−i~∂x ) i+(−i~∂y ) j+(−i~∂z ) k,
2
and − 2µ
~
∆ is the quantization of the kinetic energy of an object of mass µ. Our
atomic example then suggests that the quantization of a classical potential en-
ergy function V (r) should be the multiplication operator V̂ given by V̂ (ψ)(r) :=
V (r) ψ(r) , for ψ : R3 → C. In particular, for the coordinate function x, we should
have x̂(ψ)(r) = xψ(r), etc.. We have been very vague about the domains (spaces
of functions ψ : R3 → C) of these operators. In practice, physicists feel comfortable
with this vagueness, as long as they can make physical sense of their results. For
example, if a ∈ R3 , the function ψ(r) = eia·r/~ is a simultaneous eigenfunction
14.2. QUANTUM THEORY 375

of the momentum operators  p̂x , p̂y , p̂z in the sense that p̂x (ψ) = a1 ψ, etc., but
this ψ is not in L2 R3 , C . Also, the Dirac delta distribution δ(r − a) might be
regarded as an eigenfunction for the position operators x̂, ŷ, ẑ. Perhaps the most
appropriate domain would be the space of tempered distributions (i.e., continu-
ous linear functionals on the Schwartz space of rapidly decreasing functions); at
least this would be big enough to encompass the above examples. At any rate, the
functions in the domains of these operators are known as states of the particle
(e.g., the electron, in the above atomic example). We mention that the states de-
scribing systems of particles are essentially tensor products (sometimes symmetric
and sometimes skew-symmetric) of the individual particles. States which differ by
a constant, complex factor are identified (considered physically indistinguishable),
making the space of states an infinite-dimensional complex projective space. The
following is a basic interpretive assumption of quantum mechanics, which gives it
some physical sense.
Postulate. Let [ψ] be a state with representative ψ of norm kψk = 1 in a
Hilbert space H (typically L2 R3 , C in the single particle case). Suppose that a
quantized observable (i.e., a self-adjoint operator) A has a eigenvalue λ. Then the
probability that the observable is measured to be λ when the particle (or system)
is in the state [ψ] is the norm-square of the projection of ψ onto the eigenspace of
λ. More generally, if the self-adjoint operator R ∞A has a spectral resolution (i.e., a
projection-valued measure P on R with A = −∞ λ dPλ ), then the probability that
the observable is measured
R  to be in some interval I when the particle (or system)
is in state [ψ] is I
dP λ (ψ) , ψ .
If the self-adjoint, quantized observable A has a complete set of eigenvectors
un , n = 1, 2, 3, . . . , with Aun = λn un , then according to the Postulate, the expec-
tation of measurements of this observable for the state [ψ] (kψk = 1) is simply


X ∞
X ∞
X
2
λn |hψ, un i| = λn hψ, un i hψ, un i = hψ, Aun i hψ, un i
n=1 n=1 n=1

* +
X
(14.12) = Aψ, hψ, un i un = hAψ, ψi .
n=1

We get the same end result in the general case where A has a spectral resolution.
2
As a consequence of (14.12) in the single particle case, we show that r 7→ |ψ(r)|
2 R 2 3
(where kψk = R3 |ψ(r)| d r = 1) is the probability density for the position of the
particle in state [ψ]. Indeed, for a domain D ⊆ R3 , let χD : R3 → {0, 1} be the
characteristic function of D. Classically, this observable is 1 if the particle is in
D and 0 otherwise. As for functions on R3 in  general, the quantization of χD
is the multiplication operator χ̂D on L2 R3 , C given by χ̂D (u)(r) = χD (r) u(r).
According to (14.12), the quantum mechanical expectation of this observable for
the state [ψ] is
Z Z
2
hχ̂D ψ, ψi = hχD (r) ψ(r) , ψ(r)i d3 r = |ψ(r)| d3 r.
R3 D

2
This shows that |ψ| is the probability density for the position of the particle. Our
aim now is to show that the probability density for the momentum of the particle
376 14. PHYSICAL MOTIVATION AND OVERVIEW

2
is the function p 7→ ψ̃(p) , where
Z
−3/2
ψ̃(p) := (2π~) ψ(r) e−ip·r/~ d3 r,
R3

which is essentially the Fourier transform of ψ. By the Fourier Inversion Theorem,


Z
−3/2
ψ(r) = (2π~) ψ̃(p) eip·r/~ d3 p,
R3

and formally
Z
−3/2
p̂x (ψ)(r) = −i~∂x ψ(r) = (2π~) px ψ̃(p) eip·r/~ d3 p.
R3

More generally, for suitable functions f (p) (momentum-dependent observables), it


is natural to let
Z
−3/2
fˆ(ψ)(r) : = (2π~) f (p) ψ̃(p) eip·r/~ d3 p.
R3

For a domain D in momentum space and associated characteristic function χD (p),


we then have
Z
−3/2
χ̂D (ψ)(r) = (2π~) χD (p) ψ̃(p) eip·r/~ d3 p.
R3

The quantum mechanical expectation that the particle has momentum in D


is then
D E Z 2
hχ̂D (ψ) , ψi = χD ψ̃, ψ̃ = ψ̃(p) d3 p,
D
2
where we have used Parseval’s equality. This identifies ψ̃ as the probability
h i
density for the momentum of the particle in the state ψ̃ . Note that pseudo-
differential operators are essentially quantizations of functions of momentum.

Positions and Momenta. In classical physics, particles move along trajecto-


ries and have well-defined positions and momenta at all times. In quantum mechan-
ics, position and momentum cannot both be simultaneously determined with arbi-
trarily high precision. This is a consequence of the (Heisenberg) Uncertainty
Principle which can be deduced as follows. The uncertainty of an observable A for
a particle in state [ψ] (with kψk = 1) is the standard deviation of the measurements
of the observable, or equivalently, the root of the expectation of measurements of
2
the observable (A − hAψ, ψi) , namely
D E1/2
2
∆ψ A := (A − hAψ, ψi) ψ, ψ = k(A − hAψ, ψi) ψk .

As shown below, the Cauchy-Schwarz inequality and a little algebra reveal that for
observables A and B we have the Uncertainty Principle
1
∆ ψ A · ∆ψ B ≥ 2 |([A, B] ψ, ψ)| ,
14.2. QUANTUM THEORY 377

where [A, B] = AB − BA. Indeed,


 2 1/2  2 1/2
∆ψ A · ∆ψ B = A −(Aψ, ψ) ψ, ψ B −(Bψ, ψ) ψ, ψ
= k(A −(Aψ, ψ)) ψk k(B −(Bψ, ψ)) ψk

≥ (A −(Aψ, ψ)) ψ,(B −(Bψ, ψ)) ψ

≥ = (A −(Aψ, ψ)) ψ,(B −(Bψ, ψ)) ψ

(A −(Aψ, ψ)) ψ,(B −(Bψ, ψ)) ψ
= 21 
− (B −(Bψ, ψ)) ψ,(A −(Aψ, ψ)) ψ
1 (A −(Aψ, ψ))(B −(Bψ, ψ))
= 
2 −((B −(Bψ, ψ))(A −(Aψ, ψ))) ψ, ψ
1 1
(14.13) = 2 |((AB − BA) ψ, ψ)| = 2 |([A, B] ψ, ψ)| .
Consequently, for two noncommuting quantized observables, there is a possible
obstruction to obtaining arbitrarily low uncertainties in the measurements of both.
For example, [x̂, p̂x ](ψ) = x(−i~∂x ψ) −(−i~∂x (xψ)) = i~ψ (i.e., [x̂, p̂x ] = i~I), and
so
∆ψ x̂ · ∆ψ p̂x ≥ 21 |(i~ψ, ψ)| = ~2 |(ψ, ψ)| = ~2 .
Thus in quantum mechanics, the position and momentum of a particle in the same
direction cannot both be determined with arbitrarily high precision. For example,
initial conditions for Newton’s equation (or its relativistic analogs) for a particle
cannot be exactly specified, and the philosophy of determinism loses its grip on
reality. However, classical mechanics works well in ordinary circumstances due to
the smallness of ~.
 
Exercise 14.6. Show that if a function ψ ∈ C 1 R3 ∩ L2 R3 is a state for
which ∆ψ x̂ · ∆ψ p̂x , ∆ψ ŷ · ∆ψ p̂y and ∆ψ ẑ · ∆ψ p̂z all have the minimal value ~2 , then
2
|ψ| is a normal (Gaussian) distribution in each variable. For simplicity, you may
wish to assume (at first) that (x̂ψ, ψ) = (p̂x ψ, ψ) = 0 (and similarly for y and z).
To proceed, consider when the inequalities in 14.13 are equalities.
Time Evolution. So far we have said nothing about how states and/or ob-
servables evolve with time in quantum mechanics. While there are a number of
standard ways of introducing time evolution, here we will proceed in a somewhat
unusual manner, by drawing upon special relativity for motivation. Suppose that
γ(s) = x0 (s) , x1 (s) , x2 (s) , x3 (s) = (ct(s) , r(s))


is the trajectory of a particle of rest mass m0 in Minkowski space-time, parametrized


2 2
in the standard way so that hγ 0 (s) , γ 0 (s)i = c2 t0 (s) − kr0 (s)k = c2 . The energy-
momentum 4-vector of the particle is
p(s) : = m0 γ 0 (s) = m0 (ct0 (s) , r0 (s))
(14.14) = m0 (cβ(s) , β(s) v(t(s))) = (mc, mv) = (E/c, p) .
Since the quantization of p is p̂ =−i~∇, it is fitting that (on the basis of covariance)
more generally we should have p̂µ = i~g µν ∂ν . As g ii = −1 for i = 1, 2, 3, we then
have p̂ = − i~∇, whereas for µ = 0, we obtain Ê/c = p̂0 = i~g 00 ∂0 = +i~∂0 =
+i~ 1c ∂t (plus!). In other words, one ought to have Ê = i~∂t . However, most
physicists do not think of i~∂t as being the quantization Ê of the classical energy E,
378 14. PHYSICAL MOTIVATION AND OVERVIEW

since Ê is generally determined by other means. For example, in the nonrelativistic


1 2
setting where E = 2m kpk + V (r) for some potential V , we have seen that Ê =
2
− 2m
~
∆ + V̂ . In the so-called Schrödinger picture, the relation Ê = i~∂t is regarded
as defining the time evolution of states [ψ], in the sense that
(14.15) i~∂t ψ = Êψ.
Actually, we should be more precise here. The E in p = (E/c, p) is the total energy
of the particle, which is close to the sum of kinetic energy, potential energy, and
rest energy
1 2
(14.16) E≈ kpk + V (r) + m0 c2 ,
2m0
where we assume that V (r) is shifted to zero by adding a constant when p = 0, so
that E = m0 c2 when the particle is at rest. Thus, according to (14.15)
~2
(14.17) i~∂t ψ = Êψ = − ∆ψ + V (r) ψ + m0 c2 ψ.
2m0
This is not quite the usual Schrödinger equation because of the term m0 c2 ψ. How-
2
ever, if in (14.17) we make the replacement ψ(r, t) → e−im0 c t/~ ψ(r, t) (which at
each time replaces ψ by an equivalent state), we then obtain the official time-
dependent Schrödinger equation
~2
(14.18) i~∂t ψ = − ∆ψ + V (r) ψ.
2m0
An obvious defect of (14.18) is its noninvariance under Lorentz transformations
which arises from the approximation (14.16). However, it does have the virtue
2
that ψ(·, 0) 7→ ψ(·, t) is a unitary transformation, so that for each time t, |ψ(·, t)|
may be regarded as a probability density if kψ(·, 0)k = 1. This  unitary property

2
1
can be seen from the fact that the infinitesimal generator i~ − 2m~
0
∆ + V̂ is a
skew-hermitian operator because of the factor of i.
2
Exercise 14.7. Show that R3 |ψ(r, t)| d3 r is constant by formally differenti-
R

ating under the integral, using (14.18). You may assume that ψ(·, t) and its spacial
derivatives decay rapidly enough as |r| → ∞ to neglect boundary terms when inte-
grating by parts.
Another nice consequence of (14.18) is that (under suitable decay assumptions
on ψ(r, t) and its derivatives as krk → ∞) the expectation of the position vector of
the particle in state [ψ], namely
Z
2
R(t) := |ψ(r, t)| r d3 r,
R3
obeys not only
Z
1 1 1 ∗
(14.19) R0 (t) = P(t) := hp̂ψ, ψi = −i~∇ψ(r, t) ψ(r, t) d3 r,
m0 m0 m 0 R3
where P(t) is the expectation of the momentum, but also Newton’s equation
Z
00 2
(14.20) m0 R (t) = − |ψ(r, t)| ∇V (r) d3 r,
R3
where we note that the right side is the expectation of the force on the particle in
state [ψ].
14.2. QUANTUM THEORY 379

Exercise 14.8. Formally derive equations (14.19) and (14.20). Again, assume
that ψ and its derivatives suitably decay so that the boundary terms produced
when integrating by parts can be discarded.
One reason why the nonrelativistic (14.18) and its many-particle generalizations
are so successful in dealing with atomic phenomena is that electrons in atoms
travel at speeds of only around c/100, according to a simple approximate classical
calculation.
Perhaps it would have been better to insist on Lorentz invariance from the
2 2
beginning, thereby replacing the relation (E/c) − kpk = m20 c2 by its quan-
tized analog, namely the Klein-Gordon equation for the C-valued function ψ
on Minkowski space
−~2 c−2 ∂t2 − ∆ ψ = m20 c2 ψ.

(14.21)
However, note that there is no vestige of a potential in this equation. If we are
to introduce electromagnetism in some way, we should do it in a Lorentz invariant
way. If we write (14.21) in the covariant form
(14.22) −~2 g µν ∂µ ∂ν ψ = m20 c2 ψ,
then the most obvious way of introducing E-M is simply to add a multiple of the
gauge potential A = (Aµ ) to the operator ∂ = (∂µ ). However, in order that the
resulting equation be invariant under gauge transformations A → A0 := A + dϕ,
where ϕ : M → R, we need to subject ψ to a gauge transformation. In order to keep
2
|ψ| gauge-invariant, we might try ψ → ψ 0 := eiϕ ψ. Indeed, we have the identity
(∂µ − i(Aµ + ∂µ ϕ)) eiϕ ψ = eiϕ (∂µ − iAµ ) ψ,
 0
∂µ − iA0µ ψ 0 = ∂µ − iA0µ ψ .

or
Thus, we obtain the desired invariance

− ~2 g µν (∂µ − iAµ )(∂ν − iAν ) ψ = m20 c2 ψ


⇐⇒ −~2 g µν ∂µ − iA0µ (∂ν − iA0ν ) ψ 0 = m20 c2 ψ 0 .


In order for the units to work out, Aµ must be replaced by (const.) · Aµ having the
−1 e
same dimensions as ∂µ , namely (length) . The natural choice is c~ Aµ where e is
the charge. Thus, one may incorporate E-M into the Klein-Gordon equation by
ie
replacing ∂µ by ∂µ − c~ Aµ , obtaining
  
ie ie
(14.23) −~2 g µν ∂µ − Aµ ∂ν − Aν ψ = m20 c2 ψ.
c~ c~
Physicists call this minimal replacement, while differential geometers recognize
that this amounts to replacing ordinary derivatives by covariant derivatives.
Note that ψ changes under a change of gauge. Thus, rather than taking ψ to be a
C-valued function on space-time, this wave function is more properly regarded as an
equivariant C-valued function on the Kaluza-Klein circle bundle P (or equivalently
as a section of the associated complex line bundle). Covariant differentiation is then
forced upon us, since ordinary differentiation of sections of a vector bundle makes
2 2 2
no invariant sense. Observe that eiϕ ψ = |ψ| , whence |ψ| is gauge invariant.
2
However there is a problem with interpreting |ψ(t, ·)|R as a probability, because,
2
even with A = 0, it does not follow from (14.23) that R3 |ψ(t, r)| d3 r is constant,
380 14. PHYSICAL MOTIVATION AND OVERVIEW

as with solutions of Schrödinger’s equation (14.18). Instead, one finds that the real
quantity
Z    
ie ∗
2= ψ(t, r) ∂0 + A0 (t, r) ψ(t, r) d3 r
R3 c~
Z      
ie ie
= i ψ ∂0 − A0 ψ − ψ ∂0 + A0 ψ ∗ d3 r

R3 c~ c~
Z
2e 2
(14.24) = i(ψ ∗ ∂0 ψ − ψ∂0 ψ ∗ ) − A0 |ψ| d3 r
R3 c~
is conserved (i.e., independent of t). However, the integrand, say ρ(t, r), is not
necessarily of fixed sign everywhere even if A0 = 0, and thus does not represent
a probability density. The usual interpretation is that ρ(t, r) is proportional to
a charge probability density, but this is odd because ψ is supposedly the state
of a single particle of a definite charge e. This difficulty foreshadows the fact
that relativistic quantum theories are generally multi-particle theories in which
the number of particles is not fixed in the presence of an external potential, and
anti-particles of charge −e are naturally built in at the outset. The presence of
anti-particles manifests itself in negative-energy states ψ for which i~∂t ψ = Eψ
with E < 0.
I The full story is known as quantum field theory, as opposed to relativistic
quantum mechanics. For physicists, quantum field theory (particularly, quantum electro-
dynamics (QED)) is enormously successful, since very accurate, verifiable predictions are
made. However, for mathematicians, conventional quantum field theory leads to an unsat-
isfactory state of affairs. Indeed, no mathematical quantum field theory (which satisfies
a reasonable set of axioms) with realistic interacting particles has ever been constructed
in four space-time dimensions. From a mathematical perspective, the great tragedy is
that conventional physicists are getting fantastic answers, with little concern that a solid
theoretical foundation has yet to be found. J

Back to the Klein-Gordon Equation. Returning to (14.23), it is possible


to make contact with nonrelativistic quantum mechanics in a limiting sense, as
follows. In order to relate (14.23) to Schrödinger’s equation (14.18), for a solution
2
ψ of (14.23), we define ψs := eim0 c t/~ ψ. Note that
 2
 2
i~∂t ψs = i~∂t eim0 c t/~ ψ = eim0 c t/~ i~∂t ψ − m0 c2 ψ .

(14.25)

Thus, the transformation ψ → ψs has the effect of removing the rest energy m0 c2
from ψ. For a solution ψ of (14.23), we expect that ψs will approximately satisfy
(14.18) for some choice of V . We will show this under the assumptions that Ai = 0
for i = 1, 2, 3, A0 is time-independent (∂t A0 = 0) and terms without factors of c2
are negligible in comparison with those that do. Using (14.25), we then have
2 2
i~∂t ψ = e−im0 c t/~ i~∂t ψs + m0 c2 ψs ≈ e−im0 c t/~ m0 c2 ψs

 2 
−~2 ∂t2 ψ = i~∂t (i~∂t ψ) = i~∂t e−im0 c t/~ i~∂t ψs + m0 c2 ψs
2
 2 
= e−im0 c t/~ −~2 ∂t2 ψs + 2m0 c2 i~∂t ψs + m0 c2 ψs
2
 2 
(14.26) ≈ e−im0 c t/~ 2m0 c2 i~∂t ψs + m0 c2 ψs .
14.2. QUANTUM THEORY 381

Then under the above assumptions, (14.23) yields


  
2 2 2 −1 ie −1 ie
m0 c ψ = −~ c ∂t − A0 c ∂t − A0 ψ + ~2 ∆ψ
c~ c~
e2
≈ −c−2 ~2 ∂t2 ψ + 2c−2 eA0 i~∂t ψ + 2 A20 ψ + ~2 ∆ψ
  c
−2 −im0 c2 t/~
2 
≈ c e 2m0 c i~∂t ψs + m0 c2 ψs
2

 2
 e2
(14.27) + 2c−2 eA0 e−im0 c t/~ m0 c2 ψs + 2 A20 ψ + ~2 ∆ψ
c
or
−~2 e2 −~2
 
2 eA0
i~∂t ψs ≈ ∆ψs − eA0 ψs + A ψs = ∆ψs − eA0 1 − ψs .
2m0 2m0 c2 0 2m0 2m0 c2
Thus, assuming that the electrostatic potential energy eA0 is small compared
eA0
with the rest energy m0 c2 so that 2m 0c
2 is negligible, we approximately have

Schrödinger’s equation (14.18) with potential V = −eA0 . Incidentally, we men-


tion that expressing the density in the conserved quantity (14.24) in terms of
2
ψs = eim0 c t/~ ψ (using the first equation in (14.26)) yields
2e 2
i(ψ ∗ ∂0 ψ − ψ∂0 ψ ∗ ) − A0 |ψ|
c~
1 2

i~(ψs∗ ∂t ψs − ψs ∂t ψs∗ ) + 2 m0 c2 − eA0 |ψs |

=
c~    
2m0 c i~ ∗ ∗ eA0 2
(14.28) = (ψ ∂t ψs − ψs ∂t ψs ) + 1 − |ψs | .
~ 2m0 c2 s m0 c2
Thus, the density in (14.24) is for eA0  m0 c2 , etc., is approximately 2m0 c/~ times
2
the usual Schrödinger probability density |ψs | .
For the hydrogenic atom where A0 = Ze/r and A = 0, the relevant exact
product solutions ψ(t, r) = e−iEt/~ R(r) Yl,m (θ, ϕ) of (14.23) can be found (see
[367, p. 470]) and the corresponding energy levels to order 4 in the parameter
γ := Ze2 /(~c) are given by
γ2 γ4
   
n 3
En,l = m0 c2 1 − 2 − 4 6

(14.29) − + O γ ,
2n 2n l + 12 4
where n = 1, 2, 3, . . . is the total quantum number and l = 0, 1, . . . , n − 1 is the
azimuthal quantum number of the state. The first term m0 c2 is the rest energy
and the second term is
2
2
2
2 γ 2 Ze /(~c) m0 e4 Z 2
−m0 c = −m 0 c = −
2n4 2n4 2~2 n4
which coincides with (14.11) with m0 = µ. The third term is a relativistic correction
that predicts that there is a small spread (fine structure) in the energy levels for a
fixed n, since the different values l = 0, 1, . . . , n − 1 yield different energies (i.e., the
degeneracy is broken). However, the predicted spread is larger than the observed
spread. The problem is that, while the Klein-Gordon equation 14.23 is the most
obvious relativistic wave equation, it is not the correct one for electrons. Indeed,
there is a first-order relativistic equation for a multi-component (spinorial) wave
function that works much better, namely the Dirac equation which we consider
next.
382 14. PHYSICAL MOTIVATION AND OVERVIEW

Dirac’s Equation. The search for a first-order relativistic wave equation was
partly motivated by the fact that Schrödinger’s equation only involves a first deriva-
tive with respect to t, and the evolution of states
 is simply given by a one-parameter
group of unitary transformations on L2 R3 generated by the skew-Hermitian op-
erator −iÊ/~ formed from the quantized energy Ê. As was eventually discovered,
the problem with the Klein-Gordon equation really is not with the second-order
time derivative per se, but Dirac’s search for a first-order relativistic equation led
to the correct equation for electron wave functions. What follows is a rough out-
line of his reasoning. Consider a first-order differential operator with constant (but
possibly complex matrix) coefficients γ µ , say
A = γ µ ∂µ = γ 0 ∂0 + γ 1 ∂1 + γ 2 ∂2 + γ 3 ∂3 ∂0 := c−1 ∂t .

(14.30)
If A is relativistically (or Lorentz) invariant, then so is
2
A2 = (γ µ ∂µ ) = 21 (γ µ γ ν + γ ν γ µ ) ∂µ ∂ν .
2 2 2 2
For gµν dxµ dxν = dx0 − dx1 − dx2 − dx3 , we have g µν = gµν , and constant
multiples of the operator g µν ∂µ ∂ν are Lorentz invariant. Thus, it is reasonable to
impose the condition (for some scalar or matrix K 6= 0) that
µ ν
(14.31) 1
2 (γ γ + γ ν γ µ ) = Kg µν .
If the γ µ and K are assumed to be complex scalars then there are no such γ µ ,
since these scalars would be nonzero and would anticommute. (Recall (g µν ) =
diag(1, −1, −1, −1) for the Lorentzian metric.) With some perseverance, one can
prove that if the γ µ are n × n matrices and K = In (the n × n identity), then the
least n for which there are solutions to (14.31) is n = 4. One standard solution is
   
0 I2 0 j 0 σj
γ = , γ = ,
0 −I2 −σj 0
where
     
0 1 0 −i 1 0
(14.32) σ1 = , σ2 = , σ3 =
1 0 i 0 0 −1
are the so-called Pauli matrices. A source of headaches is the fact that the γ µ
are not unique, since we can replace γ µ by Bγ µ B −1 for any invertible 4 × 4 matrix
B. At any rate, Dirac’s equation for a C4 -valued (4-component) wave function
ψ is
(14.33) i~γ µ ∂µ ψ = m0 cψ
where m0 is the rest mass of the particle associated with ψ. The electromagnetic
gauge potential 1-form A = Aµ dxµ is again naturally introduced via minimal re-
placement:  
µ ie
i~γ ∂µ − Aµ ψ = m0 cψ.
c~
For (Aµ ) = (Ze/r, 0), this equation can be separated (see [367, p. 486]) and
one finds that there are solutions of the form ψ(t, r) = e−iEt/~ Ψ(r), where Ψ(r)
decays suitably as r → ∞ and E > 0 are the energy levels. They are indexed by
n = 1, 2, 3, . . . and |k| = 1, 2, . . . , n and are given (to fourth order in γ := Ze2 /(~c))
by
γ2 γ4
   
2 n 3 6

En,|k| = m0 c 1 − 2 + 4 − +O γ .
2n 2n |k| 4
14.2. QUANTUM THEORY 383

The fine structure exhibited by the third term agrees much better with observa-
tions than that predicted by the Klein-Gordon equation (see (14.29)). Although
discrepancies with the observed spectrum still exist (e.g., the Lamb shift), they
are accounted for within the more accurate context of quantum electrodynamics
(QED) which is a quantum field theory. In this theory, the wave function ψ and
the E-M gauge potential A are replaced by distributions with values that are op-
erators in a multi-particle Hilbert space of states. Although the mathematics of
QED is shady, the formalities involved give rise to recipes for computing physi-
cal quantities in terms of formal power series in the dimensionless fine structure
constant α := e2 /(c~) ≈ 1/137. Such series are called renormalized perturbations
series. The coefficients of such series are computed by summing up integrals as-
sociated with so-called Feynman diagrams. For a rigorous — and comprehensible
— introduction see [355]. One trouble is that integrals associated with Feynman
diagrams that have loops are infinite. By various procedures known as renormal-
ization techniques, finite values for the coefficients of the perturbations series
are extracted. Field theories for which this is the case (e.g., QED) are known as
renormalizable. The various renormalization techniques all lead to the same values
for the coefficients, which is reassuring. For QED, the terms of these series become
smaller at least initially. Eventually, the coefficients become incalculable due to
the huge number and complexity of the Feynman diagrams. The consensus among
those who have studied these series in some detail is that the coefficients eventually
increase rapidly enough so that the series do not converge. Contrary to popular
misconceptions (even held by good physicists) a formal power series in α does not
necessarily converge, even if α ≈ 1/137. However, as with asymptotic series, be-
fore divergence sets in, one obtains amazing accuracy compared with experimental
results. For example, we have the following values for the magnetic moment of
the electron:
e~
Experiment: 2me c (1.00115965241 ± 20)
e~
QED: 2me c (1.00115965238 ± 26) .

Thus, in spite of the profound mathematical problems with QED (e.g., its very ex-
istence as a mathematical theory, beyond computational recipes), QED is hailed as
one of the most successful physical theories from the perspective of most physicists.
Returning to the Dirac equation (14.33), we have not yet indicated the sense
in which it is Lorentz invariant. For the Klein-Gordon equation (14.22), Lorentz-
invariance means that if ψ is a solution and L ∈ O(1, 3) is a Lorentz transformation,
then ψ ◦ L is also a solution. For the Dirac equation, one considers the universal
double cover C : SL(2, C) → O0 (1, 3) of the identity component O0 (1, 3) of O(1, 3).
There is a representation r : SL(2, C) → GL(4, C), such that if ψ is a solution
of Dirac’s equation i~γ µ ∂µ ψ = m0 cψ, and A ∈ SL(2, C) then A−1 ψ ◦ r(A) is
also a solution. This is the meaning of Lorentz-invariance for the Dirac equation.
Since the representation r is the sum of two irreducible spin- 21 representations, the
Dirac equation not only takes into account the spin of the electron, but it forecasts
the existence of the positron. In the terminology of modern differential geometry
(introduced in the next Chapter), the Dirac wave function is a section of a complex
4-dimensional vector bundle (the Dirac bispinor bundle) which is associated to a
double cover (by spinor frames) of the bundle of space-time oriented orthonormal
Lorentz frames. If the E-M gauge potential is included, then gauge invariance
384 14. PHYSICAL MOTIVATION AND OVERVIEW

dictates that the Dirac bispinor bundle is associated with the fibered product of
the spinor frame bundle and the Kaluza-Klein U(1) (circle) bundle.
Of course, electrons and photons are just part of the total picture. In place of
the U(1) circle bundle used in the original Kaluza-Klein theory to introduce E-M,
one uses principal bundles P with a larger Lie group to incorporate other
nongravitational forces. The wave functions of the fundamental particles of matter
are sections of various vector bundles that are associated (via group representa-
tions) to the so-called fibered product of the bundle P with the bundle of spinor
frames over space-time. The known fundamental particles of matter include the 6
flavors of quarks (u (up), d (down), s (strange), c (charm), t (top), b (bottom))
each in three colors (R (red), G (green), B (blue)) together with the leptons (e−
(electron), µ− (muon), τ − (tau)) together with their associated neutrinos (νe (elec-
tron neutrino), νµ (muon neutrino), τ − (tau neutrino)). These are organized into
three generations (see the table below) which are essentially identical except in the
masses of the corresponding particles, higher generation particles generally being
heavier than their lower generation counterparts; e.g., the masses of e− , µ− , τ − are
approximately .511, 105.66, and 1784 MeV, respectively.
Generation 1 2 3
charge + 32

u(R,G,B) c(R,G,B) t(R,G,B)
quarks 2
 charge − 3 d(R,G,B) s(R,G,B) b(R,G,B)
charge − 1 e− µ− τ−
leptons
charge 0 νe νµ ντ
We mention that for each quark q(R,G,B) , there is an anti-quark q̄(C,M,Y ) of the
opposite complementary color (C, M, Y ) = (cyan, magenta, yellow). In addition
to gravity, the known forces are as follows. There is the strong force of QCD
(quantum chromodynamics) acting between the quarks inside strongly interacting
particles (hadrons) such as neutrons, protons and pions, which is mediated by 8
gluons (one for each vector in a basis for su(3) :=Lie algebra of SU(3)). The col-
ors of the quarks can be regarded as charges that respond to the strong force in
the sense that electrically charged particles respond to E-M fields. The fact that
quarks are confined within hadrons seems to be related to the fact that SU(3) is
nonabelian, which causes gluons to interact with each other. Unlike the coulomb
force, strong forces between quarks weaken as the separation distance decreases
to zero, but strong forces strengthen dramatically when the distance increases to
the diameter of a hadron. Individual leptons (e.g., electrons, neutrinos) appear
in the open (unconfined) since, being colorless, they are unaffected by gluons. In
mathematical terms, SU(3) acts trivially on the lepton sector of the relevant repre-
sentation. Corresponding to 4 generators of the Lie algebra of SU(2)× U(1), there is
the weak force mediated by the Z and ±W vector bosons and the E-M force due
to photons. Incidentally, U(1) of E-M is not simply the U (1) factor in SU(2)× U(1).
In terms of the Pauli matrices σk of (14.32), a set of standard generators for the
complexified Lie algebra su(2) ⊗ C are the matrices 12 σ3 and σ ± := 12 (σ1 ± iσ2 ).
The generator for electric charge in the Lie algebra su(2) ⊕ u(1) is a linear combi-
nation of 2i σ3 ∈ su(2) and the generator i ∈ u(1), while the Z boson is associated
with an independent linear combination of these generators. The ±W bosons are
associated with σ ± . The weak force acts on leptons as well as quarks. Among
other things, it is responsible for the decay of an isolated neutron into a proton,
electron and antineutrino in about 15 minutes on average. It was primarily for
14.2. QUANTUM THEORY 385

their work in exhibiting that the electro-weak unification was feasible in the con-
text of a spontaneously broken SU(2)× U(1) gauge theory (the GWS-theory) that
Sheldon Glashow, Steven Weinberg and Abdus Salam were awarded the
1979 Nobel Prize for Physics. The Z and ±W vector bosons were detected by
experimentalists in 1983. Unlike true gauge bosons such as the photon and gluons,
the Z and ±W are massive, due to the fact that the SU(2) × U(1) gauge symmetry
is broken, leaving the U(1) of E-M as the surviving gauge group. The standard ex-
planation for how the symmetry was broken is known as the Higgs mechanism.
According to a press release of CERN of July 4, 2012, the Higgs particles in the
theory have finally been found, see also [110]. There is another peculiarity of the
weak force in that it acts only the so-called left halves of quarks and leptons. To
understand this a little better, recall that a bispinor field for a particle is locally
C4 -valued. Two of these components correspond to the particle and two correspond
to the antiparticle. Of the two components for the particle, one is the left-handed
component and one is the right-handed component, and these are called the chiral
halves of the particle; the antiparticle part also has chiral halves. Mathematically,
the weak force associated with “su(2)-like” broken generators in the complexified
su(2) ⊕ u(1) act in the usual way (as SU(2) acts on C2 ) on certain chiral doublets,
such as e− 0 0
L , veL and (uL , dL ). In a less complicated world, dL might simply be
the left-handed chiral half dL of the down quark, but in this world, d0L is a linear
combination of dL , sL , and bL , where the coefficients form the first row of the so-
called Cabibbo-Kobayashi-Maskawa 3×3 matrix. The other rows of the CKM
matrix are determined by the weak SU(2) doublets (cL , s0L ) and (tL , b0L ), when s0L
and b0L are written as linear combinations of dL , sL , and bL . Incidentally, in a less
complicated world where d0L = dL , etc. or where there are fewer than 3 generations,
we might not exist. Indeed, then certain time-asymmetric weak reactions (e.g., K 0
meson decay) would not occur (see [336, p. 725]. It has been speculated that in the
absence of such reactions, certain quark-nonconservation processes in grand unified
theories might cause so much matter and anti-matter annihilation that there would
be too few quarks left to make enough nucleons (see [115, pp. 176f]). At any
rate, to account for all of these known nongravitational forces, the gauge group of
the principal bundle, say P , over space time M (before symmetry-breaking) must
include SU(3) × SU(2) × U(1), a far cry from the U(1) group for the circle bundle
of the original Kaluza-Klein theory. As a point of historical interest, Oscar Klein
was the first to introduce SU(2) gauge fields (commonly known as Yang-Mills fields)
and he even anticipated their use in modeling weak interactions in [245], 30 years
before the GWS model was developed.

In the Search for Symmetry. Many have attempted to incorporate the


group SU(3) × SU(2) × U(1) into a larger simple group (e.g., SU(5), SO(10), etc.),
thereby obtaining a grand unified theory (GUT). The SO(10) (or more precisely,
Spin(10)) GUT is particularly tidy, since the two fundamental spinor representation
of Spin(10) are 16-dimensional and the left chiral halves of an individual generation
of fundamental particles of matter fits perfectly (and correctly) into one of these
16-dimensional representations, while the right chiral halves fit into the other. For
the first generation of (say, left) chiral halves, 12 of the 16 dimensions are accounted
for by the 4 left chiral halves of the up and down quarks and antiquarks replicated
in three colors, 2 dimensions come from the left chiral halves for the electron and
positron, and the remaining 2 dimensions are occupied by the left chiral halves for
386 14. PHYSICAL MOTIVATION AND OVERVIEW

the electron neutrino and antineutrino. Incidentally, for years many believed that
neutrinos were massless and not bispinorial (i.e., having just 2-component wave
functions), but recent experiments strongly suggest (if not prove) that neutrinos
have a small mass which necessitates the existence of left-handed and right-handed
neutrinos and antineutrinos, instead of just left-handed neutrinos and right-handed
antineutrinos. It should be emphasized that not only do all of the left chiral halves
of the fundamental particles fit by virtue of dimension count into a fundamental
representation of Spin(10), but under the usual inclusions
su(3) × su(2) × u(1) ⊂ su(5) ⊂ so(10) ∼
= spin(10)
of Lie algebras the left chiral halves all respond correctly to the various forces
under the spinor representation. In the Lie algebras of possible grand unification
groups, there are generators which are not in su(3) × su(2) × u(1) and hence do not
correspond to standard known forces. The forces associated with these generators
are thought to be involved in processes that convert quarks to leptons, which for
example can lead to proton decay, a process which has yet to be detected. The
force of gravity is not encoded in a grand unification group, but rather it is gauged
in a different sense by the Lorentz group (or its cover SL(2, C)) for the bundle of
frames (or spinor frames) for space-time itself. In this way, gravity seems to resist
attempts to unify it with other forces, and no universally convincing method of
quantizing it has been forthcoming. The best hope seems to reside in string theory.
It must be stressed that unlike the Schrödinger wave function ψs which speci-
fies the quantum state of a single particle in nonrelativistic quantum mechanics and
2
whose modulus square |ψs | is the position probability density, the Dirac bispinor
wave functions ψ (or sections) for leptons and quarks, do not admit such an easy
interpretation. One source of confusion is that these wave functions are not re-
garded as quantum fields even though in a certain nonrelativistic limit, two of the
components of the Dirac wave function ψ can be identified with the two compo-
nents of the Pauli wave function that satisfies the so-called Pauli equation which
is a Schrödinger equation for single particles with spin 12 . In view of this, Dirac
bispinor wave functions ψ are referred to as first quantized wave functions, while
the process that converts such ψ to operator-valued distributions is known as sec-
ond quantization. However, as some correctly point out, an object should only
be quantized once. Thus, one should regard Dirac bispinor wave functions ψ as
classical states which have yet to be quantized. However, there are difficulties with
interpreting such ψ as classical states. For example, there are two pointwise scalar
products for Dirac bispinors ψ, one is simply ψ ∗ ψ, where ψ ∗ is conjugate trans-
pose of the C4 -valued, while the other is ψ̄ψ := ψ ∗ γ 0 ψ where ψ̄ := ψ ∗ γ 0 is the
so-called dual bispinor. Since ψ ∗ ψ ≥ 0 and its integral over R3 is time-independent
(as a consequence of Dirac’s equation γ µ ∂µ ψ = mcψ), one might be tempted to
think of it as a probability density. However, when ψ is quantized (i.e., turned
into a suitable operator-valued distribution), the expectation values of ψ ∗ ψ have
the interpretation as the charge density of a collection of positively and negatively
charged particles and antiparticles, and so the positivity of ψ ∗ ψ does not survive
quantization. Moreover, the prequantum indefinite scalar product ψ̄ψ has posi-
tive expectation values after quantization and is interpreted as an energy operator.
Generally, many classical fields cannot be given a reasonable physical significance
until they are quantized. For example, although the forces of QCD are mediated by
mass 0 gluons (corresponding to photons in QED) and thus might be expected to
14.2. QUANTUM THEORY 387

have a long range (decaying as r−2 , rather than exponentially), no unconfined, long
range effects of gluons are evident, unlike the case of E-M fields of photons. Since
gluons seem to be confined to very small regions inside hadrons, it would appear
very speculative to treat them as classical wave-like fields. By the photoelectric ef-
fect, we know that E-M does not behave much like a wave even at the vastly greater
dimensions of an atom. It is nevertheless believed that the classical solutions of the
field equations for nonabelian gauge fields (particularly, in 4-dimensional Euclidean
space, with positive-definite metric) do yield at least a first-order approximation to
certain quantum effects for such fields, especially with regard to tunneling phenom-
ena. Although we cannot go into the details of how this works in quantum field
theory, there is a similar situation in quantum mechanics, which can be understood
through the following discussion.

Ground States and Classical Action. Consider the Schrödinger operator


~2
H := H0 + V where H0 := − 2m ∆ and V : R3 → R is a suitable potential (e.g.,
V− := min(V, 0) ∈ L2 R3 + L∞ R3 and V+ := max(V, 0) ∈ L2loc R3 ). Then
  

H0 + V is an unbounded essentially self-adjoint operator on the dense domain


C0∞ R3 of L2 R3 , see [353, p. 185]. One then obtains a stronglycontinuous


1-parameter group of unitary transformations exp(−itH/~) of L2 R3 , see [352,


p. 265]. For f in the Schwartz space of rapidly decreasing functions on R3 , the
solution of the problem

~2
i~∂t ψ = − ∆ψ + V ψ, ψ(x, 0) = f (x)
2m
is given by ψ(x, t) = [exp(−itH/~) f ](x). Even though H0 and V do not commute
in general, there is a formula due to T. Kato and H. F. Trotter (see [243] and
[417]) that yields
      k !
i −it −it
exp − tH = lim exp V exp H0 .
~ k→∞ ~k ~k

It is well-known (and not hard to prove) that


−3/2 Z !
   2
i 2πi~t im |x − y|
exp − tH0 [f ](x) = exp f (y) d3 y,
~ m R3 2~t

whence
   
−it −it
exp V exp H0 [f ](x)
~k ~k
−3/2Z  
2
m |x−x0 |
 it
2πi~t 2 (t/k)2 −V(x)
f (x0 ) d3 x0 .
~k
= e
mk R3

For x0 , x1 , . . . , xk ∈ R3 , let
k
!
2
X m |xj − xj−1 | t
At (x0 , x1 , . . . , xk ) := 2 − V (xj ) .
j=1
2 (t/k) k
388 14. PHYSICAL MOTIVATION AND OVERVIEW

Taking xk = x, the Kato-Trotter formula then yields


ψ(x, t) = [exp(−itH/~) f ](x)
     k
−it −it
= lim exp V exp H0 f (x)
k→∞ ~k ~k
 −3k/2Z
2πi~t i
= lim e ~ At(x0 ,...,xk ) f (x0 ) d3 x0 · · · d3 xk−1 .
k→∞ mk R3k

Let γ : [0, t] → R3 be a path with γ(jt/k) = xj for j = 0, . . . , k − 1 and


 
k
γ(s) = (xj+1 − xj ) s − j + xj for s ∈ [jt/k,(j + 1) t/k] .
t
Then the classical action for the path γ is
Z t
m 0 2
A(γ) : = |γ (s)| − V (γ(s)) ds
0 2
k−1
X Z (j+1)t/k m  |xj − xj−1 | 2
= − V (γ(s)) ds
j=0 jt/k
2 t/k
k  2 !
X m |xj − xj−1 | t
(14.34) ≈ − V (xj ) = At (x0 , x1 , . . . , xk ) .
j=1
2 t/k k

Now, integrating with respect to x1 , . . . , xk−1 is like integrating over all polygonal
paths that have k−1 segments, starting at x0 at time 0 and ending at some arbitrary
point x at time t. Illustrative examples are given in [355]. As k → ∞, the variety
of such paths is sufficiently great so that we (as R. P. Feynman) are tempted to
write !
Z Z
i
A(γ)
ψ(x, t) = e ~ dγ f (x0 ) d3 x0 ,
R3 Pt(x0 ,x)

where Pt (x0 , x) is the space of continuous paths γ : [0, t] → R3 with γ(0) = x0 and
γ(t) = x, A(γ) is the action defined in (14.34), and dγ is some kind of measure on
Pt (x0 , x). In other words, the kernel for the operator exp − ~i tH is formally
Z
i
K(x0 , x, t) = e ~ A(γ) dγ.
Pt(x0 ,x)
i
The integrand e ~ At(γ) oscillates the least about paths for which A(γ) is stationary
among paths in Pt (x0 , x), namely classical paths that are solutions of Newton’s
equation mγ 00 = −∇V. Hence we expect K(x0 , x, t) to be most greatly influenced
by classical paths (typically only one) with γ(0) = x0 and γ(t) = x. As ~ → 0, this
effect becomes more pronounced, and presumably we obtain classical mechanics in
the limit. Note that formally
   2
2 i
|K(x0 , x, t)| = exp − tH δ(x0 ) , δ(x)
~
is the probability density that at time t the particle will be found at x, given that
it was at x0 at time t. While the path integral is a suggestive formalism for the
14.2. QUANTUM THEORY 389

rigorous Kato-Trotter limit, M. Kac (see [236]) noticed that if one replaces the
time variable t by a pure imaginary parameter −iτ , then we obtain
k
!
2
i 1 X m |xj − xj−1 | τ
A−iτ (x0 , x1 , . . . , xk ) = − 2 + V (xj )
~ ~ j=1 2 (τ /k) k
Z τ
1 m 0 2 1
= − |γ (s)| + V (γ(s)) ds =: − AE (γ) ,
~ 0 2 ~
where γ : [0, τ ] → R3 is a polygonal path through x0 , x1 , . . . , xk and AE (γ) is
the so-called Euclidean action. Then M. Kac was able to express K(x0 , x, −iτ )
rigorously as a path integral in terms of conditional Wiener measure Wxτ0 ,x on the
set Pτ (x0 , x) of continuous paths γ : [0, τ ] → R3 with γ(0) = x0 and γ(τ ) = x. The
Feynman-Kac Formula (see [174, Theorem 3.2.3]) is then
!
1 τ /2
Z Z
(14.35) K(x0 , x, −iτ ) = exp − V (γ(s)) ds dWxτ0 ,x (γ) .
~ −τ /2
Rτ 0 2
Note that the kinetic part 0 m 2 |γ (s)| ds of the Euclidean action AE (γ) has been
absorbed into the measure Wxτ0 ,x , and (since it is convenient for some purposes)
the paths γ are reparametrized symmetrically using [−τ /2, τ /2] instead of [0, τ ].
It is also of interest that the subset of paths γ with Hölder exponent larger than
1 1
2 (which includes the piecewise C paths) has Wiener measure 0 and hence this
subset does not contribute to the integral. However, one still expects that the
greatest contributions to K(x0 , x, −iτ ) come from fluctuations about a path which
minimizes the Euclidean action AE (γ). Such paths are solutions of the Euler-
Lagrange equation (with conditions γ(0) = x0 and γ(τ ) = x)
mγ 00 (s) = ∇V (γ(s))
which differs from Newton’s equation mγ 00 (s) = −∇V (γ(s)) by a minus sign. We
now interpret K(x0 , x, −iτ ), at least formally. For simplicity, suppose that H =
~2
− 2m ∆ + V where V (r) increases rapidly enough as krk → ∞ so that there is
a complete orthonormal set of eigenfunctions u0 , u1 , u2 , . . . of H with eigenvalues
(energies) arranged in increasing order, say E0 ≤ E1 ≤ E2 ≤ . . . (degeneracy
allowed). The kernel of exp − ~i tH is given by



i
X
K(x0 , x, t) = e− ~ tEn un (x) un (x0 ).
n=0
Indeed,

Z !
− ~i tEn
X
ψ(x, t) := e un (x) un (x0 ) f (x0 ) d3 x0
R3 n=0
solves Schrödinger’s equation at least formally and ψ(x, 0) is the eigenfunction
expansion of f (x). Replacing t by −iτ , we obtain

τ
X
K(x0 , x, −iτ ) = e− ~ En un (x) un (x0 ).
n=0

Letting τ → ∞, we formally obtain (for some k > 0)


τ
X
e ~ E0 K(x0 , x, −iτ ) = un (x) un (x0 ) + O e−kt ,


En =E0
390 14. PHYSICAL MOTIVATION AND OVERVIEW

where the sum is only over lowest energy states. If x0 is a minimum for V ,
then the constant path γ : [0, τ ] → {x0 } clearly minimizes the Euclidean action
Rτ m 0 2
0 2
|γ (s)| + V (γ(s)) ds for loops at x0 . Hence, such paths are likely to make
τ P∞ 2
K(x0 , x0 , −iτ ) ≈ e− ~ E0 En =E0 |un (x0 )| larger at minima for V than at other
x0 . Indeed, we expect the position probability densities for ground energy states
to be concentrated about the minima for V . Classical intuition leads us to suspect
that if there are N absolute minima for V , then there ought to be N independent
eigenfunctions, each peaked at a different minimum; i.e., that there is an N -fold de-
generacy in the lowest energy level (i.e., En = E0 for n = 0, 1, . . . , N − 1). However,
there is a very general result stating that if V is continuous and bounded below and
~2
H = − 2m ∆ + V is essentially self-adjoint, then the ground state is nondegenerate
and is represented by a real, positive function. Indeed, there is an elegant proof of
this in [174, Corollary 3.3.4] based in part on the Feynman-Kac Formula (14.35).

The Concept of an Instanton. We will examine a simple example in dimen-


sion 1 (i.e., x ∈ R) to illustrate this and to introduce the concept of an instanton.

Figure 14.1. Quartic potential and its negative (dashed)

2
Let V (x) = 12 x2 − x20 , as shown in Figure 14.1. Note that the Schrödinger
~2
operator H = − 2m ∆+V commutes with the parity operator P given by(P f )(x) :=
f (−x), and hence the eigenspaces of H split into even functions and odd functions
(+1 and −1 eigenspaces of P ). Classical intuition falsely suggests that there are two
normalized E0 -energy eigenfunctions, say ψ peaked at x0 and P ψ peaked at −x √0
2 2
(kψk = kP ψk = 1). √ Suppose that this is the case, and let ψ+ = (ψ + P ψ) / 2
and ψ− = (ψ − P ψ) / 2 be the associated even and odd states. Then for any
x ∈ R, as τ → ∞,
τ
e ~ E0 K(x, −x, −iτ ) ∼ ψ+ (−x) ψ+ (x) + ψ− (−x) ψ− (x)
2 2
(14.36) = |ψ+ (x)| − |ψ− (x)| .
From (14.35), K(x0 , x, −iτ ) is the path integral of a positive function, and hence
τ
e ~ E0 K(x, −x, −iτ ) ≥ 0 so that
2 2
(14.37) |ψ+ (x)| − |ψ− (x)| ≥ 0.
14.2. QUANTUM THEORY 391

If this inequality were strict even at a single point, then we would get the contra-
diction
2 2
(14.38) 0 < kψ+ (x)k − kψ− (x)k = 1 − 1 = 0.
We now argue (as physicists might) that the inequality (14.37) is strict near x =
x0 . Consider the Euclidean potential −V whose graph is shown dashed in Fig-
ure 14.1. There is a classical solution γ∞ : (−∞, ∞) → (−x0 , x0 ) of the Euclid-
ean equation of motion mγ 00 (s) = ∇V (γ(s)) for a particle, with total energy
1 0 2
2 mγ∞ (s) − V (γ∞ (s)) = 0, that moves from the top of the left hill to the top
of the right hill. Physicists call such trajectories (and their analogs in quantum
Rτ 0 2
field theory) instantons. The Euclidean action AE (γ) = 0 m 2 |γ (s)| + V (γ(s)) ds
of the instanton is finite, since making the change of variable x = γ∞ (s), we have
Z ∞ Z ∞
1 0 2
2 mγ ∞ (s) + V (γ ∞ (s)) ds = 2V (γ∞ (s)) ds
−∞ −∞
Z x0 Z x0
−1 1
= 2V (x) dxds dx = 2V (x) p dx
−x0 −x0 2V (x) /m
Z x0 p
(14.39) = 2mV (x) dx < ∞.
−x0

The instanton γ∞ minimizes the action functional AE (γ) among suitable competing
paths from −x0 to −x0 . Let γτ := γ∞ |[−τ /2,τ /2] . We have for all τ > 0,
Z τ /2 Z x0 p
AE (γτ ) = 2V (γτ (s)) ds < 2mV (x) dx < ∞.
−τ /2 −x0
τ
Since the kinetic part of AE (γτ ) is implicit in dW−x 0 ,x0
, by a formal applica-
tion of Laplace’s method, we expect that  there is a finite, nonzero contribution

1 τ /2
R
to K(−x0 , x0 , −iτ ) proportional to exp − ~ −τ /2 AE (γτ ) ds ; i.e., for some con-
stant C > 0,
!
1 τ /2
Z Z
K(x0 , −x0 , −iτ ) = exp − V (γ(s)) ds dWxτ0 ,−x0 (γ)
~ −τ /2
!
1 τ /2 1 ∞
Z  Z 
≥ C exp − AE (γτ ) ds ≥ C exp − AE (γ∞ ) ds
~ −τ /2 ~ −∞
1 x0 p
 Z 
= C exp − 2mV (x) dx > 0.
~ −x0
Since E0 ≥ 0, as τ → ∞,
2 2 τ
|ψ+ (x0 )| − |ψ (x0 )| ∼ e ~ E0 K(−x0 , x0 , −iτ )
1 x0 p
   Z 
1
≥ C exp − AE (γ∞ ) ≥ C exp − 2mV (x) dx > 0.
~ ~ −x0
Hence, we arrive at the contradiction (14.38). In the above heuristic argument, the
contradiction was produced by the instanton γ∞ , and hence physicists are led to
attribute the nondegeneracy of the ground state to the presence of this instanton.
From our previous discussion, we know that in fact there is a unique ground state
represented by a positive wave function ψ0 . Since H commutes with the parity
operator P , we know that P ψ0 also is a positive representative of the ground state
392 14. PHYSICAL MOTIVATION AND OVERVIEW

and hence P ψ0 = ψ0 (i.e., ψ0 is even). The evenness of the ground state is also a
2
consequence of (14.36). Regardless of whether ψ0 is even or odd, we have that |ψ0 |
is even and hence a particle in the state ψ0 has the same probability of appearing
in an interval about x0 as it does in the reflected interval about −x0 , even for
measurements made in rapid succession. This is so even though the height of the
potential barrier between −x0 and x0 forbids travel of a classical particle of energy
E0 between these two points. In other words, we have the phenomenon of quantum
tunneling. There is also a quantitative link between the instanton and quantum
tunneling. It turns out that by WKB methods, the transmission amplitude for
a particle to penetrate the potential barrier from x0 to −x0 is proportional to
exp − ~1 AE (γ∞ ) (see [238, pp.545-554]).


Much of this discussion on the role of instantons in removing degeneracy and


tunneling caries over, at least metaphorically, to quantum field theory. Since we
will be primarily concerned with nonabelian gauge fields (e.g., where the gauge
group is SU(2), isomorphic to the unit quaternions S 3 ), we confine ourselves to a
very brief account of what instantons are in context of the quantum field theory of
pure Yang-Mills fields and how they correspond to the ones we have discussed in
relation to quantum mechanics. The configuration space in the quantum mechanics
of a single particle is simply R3 , but for pure Yang-Mills fields, the configuration
space is essentially the space Ω1 R3 , su(2) of all smooth 1-forms on R3 (with values
in the Lie algebra su(2) of traceless skew-Hermitian matrices), modulo the action
by the group of gauge transformations. Of course, any SU(2)-bundle over R3 is just
a product, R3 × SU(2). Hence,  a gauge transformation amounts1 to3a map φ : R →
3
∞ 3
SU(2);i.e., φ ∈ C R , SU(2) . The right action of φ on A ∈ Ω R , su(2) is given
by
A · φ := φ−1 (Aφ + dφ) .
Note that
−1
A ·(φη) = (φη) (A(φη) + d(φη)) = η −1 φ−1 (A(φη) +(dφ) η + φdη)
= η −1 φ−1 Aφ + φ−1 dφ η + dη
 

= η −1 φ−1 (Aφ + dφ) η + dη = (A · φ) · η.


 

We require that φ(r) → Id ∈ SU(2) as krk → ∞, and φ yields a map φ0 : S 3 →


SU(2) ∼= S 3 which is classified up to homotopy by its degree. A gauge transforma-
tion φ is called homotopically
 trivial if φ0 has degree 0. Moreover, it is required of
1 3
A ∈ Ω R , su(2) that the field strength

FA := dA + A ∧ A ∈ Ω2 R3 , su(2)


2 2
be square integrable; i.e., R3 |FA | < ∞, where |b| := 12 Tr(b∗ b) for b ∈ su(2).
R

More precisely, for


 Z 
2
Ω1 R3 , su(2) finite := A ∈ Ω1 R3 , su(2) :
 
|FA | < ∞
R3

the configuration space is



Ω1 R3 , su(2) finite
C := ,
C ∞ (R3 , SU(2))0
14.2. QUANTUM THEORY 393

where C ∞ R3 , SU(2) 0 is the group of homotopically trivial gauge transformations.




The nonlinear Yang-Mills functional V : C → [0, ∞) , defined by


Z
2
(14.40) V([A]) := 21 |FA | ,
R3
plays the role of the potential energy function V in quantum mechanical setting of a
single particle. Note that V is well defined, since one can verify that FA·φ = φ−1 FA φ
2 2 2
and |FA·φ | = φ−1 FA φ = |FA | using the fact that for b ∈ su(2) and C ∈ SU(2)
2
 ∗ 
C −1 bC = 12 Tr C −1 bC C −1 bC
2
= 12 Tr C −1 b∗ CC −1 bC = 21 Tr C −1 b∗ bC = |b| .
  
(14.41)
Suppose that φi : R3 → SU(2) (i = 1, 2) are inequivalent (i.e., deg(φ01 ) 6= deg(φ02 )).
For Ai := φ−1 1 3
i dφi ∈ Ω R , su(2) , we have Ai = 0 · φi and so FAi = F0 = 0. There
0
is no φ with deg(φ ) = 0 such that A1 · φ = A2 . Indeed,
A1 · φ = A2 =⇒ φ−1 −1
1 dφ1 · φ = φ2 dφ2 =⇒ (0 · φ1 ) · φ = 0 · φ2

=⇒ 0 · φ1 φφ−1 = 0 =⇒ d φ1 φφ−1
 
2 2 =0
(14.42) =⇒ φ1 φφ2−1 = Id =⇒ φ1 φ = φ2 =⇒ deg(φ01 ) = deg(φ02 ) .
Thus, [A1 ] and [A2 ] are distinct, absolute minima  the potential V, and there are
for
infinitely many distinct minima of the form φ−1 dφ , one for each possible value of


deg(φ0 ).
In analogy with the single particle setting, an instanton is a certain curve
connecting minimum [A1 ] to minimum [A2 ] in the configuration space C parame-
trized by τ ∈ (−∞, ∞).  Such a curve can be regarded as the class of a point,
say A ∈ Ω1 R4 , su(2) where we mod out by suitable gauge transformations in
C ∞ R4 , SU(2) 0 . Again by analogy, we want to minimize the Euclidean action of
this curve among competitors running from [A1 ] to [A2 ]. This Euclidean action is
R 2 2
naturally taken to be R4 |FA | , where FA := dA+A∧A as before, and |FA | is com-
4
puted using the Euclidean metric on R , as opposed to the Minkowski metric. Note
that this is analogous to replacing t by −iτ in the single particle setting. One major
goal of the following chapters is to use the Atiyah-Singer  Index Theorem
 −1 to prove
that the set of instantons connecting [A1 ] = φ−1

1 dφ 1 to [A 2 ] = φ2 dφ2 form
an (8k − 3)-dimensional manifold, where k = |deg φ1 − deg φ2 |. It is important to
know this dimension in order to estimate its full effect with regard to vacuum tun-
neling between [A1 ] and [A2 ]. The interested reader may find some insight into this
very tricky business in [56] and [114]. It should also be noted that the Lagrangians
(actions) of other fields (e.g., Dirac bispinor fields for fundamental particles) must
be added to the self-actions of pure gauge potentials. The Index Theorem is also
essential for estimating the effects of these other fields (e.g., Euclidean fermionic
lowest energy modes) on Green’s functions of quantum field theory (see [372]).
CHAPTER 15

Geometric Preliminaries

Synopsis. Principal G-Bundles; Hopf Bundle. Connections and Curvature: Connec-


tion 1-Form; Maurer-Cartan Form; Horizontal Lift. Equivariant Forms and Associated
Bundles: Associated Vector Bundles, Equivariance, and Basic Forms; Horizontal Equi-
variant Forms; Covariant Differentiation and the General Bianchi Identity; Inner Prod-
ucts, Hodge Star Operator, and Formal Adjoints. Gauge Transformations: Distinguishing
Gauge Transformations from Automorphisms; The Group of Gauge Transformations; The
Action of Gauge Transformations on Connections; Lie Algebra Analogy and Infinitesimal
Action. Curvature in Riemannian Geometry: The Bundle of Linear Frames; Connections
and Forms; Kozul Connection; The Orthonormal Frame Bundle; Metric Connections; The
Fundamental Lemma of Riemannian Geometry and the Levi-Civita Connection; Local
Coordinates and Christoffel Symbols; The Curvature of the Levi-Civita Connection; First
and Second Bianchi Identities; Ricci and Scalar Curvature as Contractions, and Einstein’s
Equation; All Possible Curvature Tensors on Rn and the Kulkarni-Nomizu Product; Cur-
vature Parts on 4-Manifolds and Self-Duality. Bochner-Weitzenböck Formulas: Fibered
Products; Contractions and Components; The Connection Laplacian, the Hodge Lapla-
cian, and the Bochner-Weitzenböck Formula; Special Cases. Characteristic Classes and
Curvature Forms: Chern Classes as Curvature Forms; The Pfaffian; Pontryagin Classes;
Other Characteristic Classes Related to Index Theory; Multiplicative Classes; Todd Class
and L-Polynomials; Recalculating Characteristic Classes; Unifications on Almost-Complex
Manifolds. Holonomy.
I Here we provide fundamental definitions and results concerning the geometry and
topology of fiber bundles that is essential to understanding gauge theories. In most journal
articles on the subject, it is assumed that the reader knows this material or can dig it out
from various sources. To cut down on the frustration, we develop the following topics,
assuming a more modest background: J

15.1. Principal G-bundles


15.2. Connections and Curvature
15.3. Equivariant Forms and Associated Bundles
15.4. Gauge Transformations
15.5. Curvature in Riemannian Geometry
15.6. Bochner-Weitzenböck Formulas
15.7. Characteristic Classes and Curvature Forms
15.8. Holonomy

1. Principal G-Bundles
A Lie group is simply a group which is a smooth (C ∞ ) manifold for which
the map (g1 , g2 ) 7→ g1 g2−1 is a C ∞ map from G × G to G. Let P be a manifold
on which a Lie group G acts freely and smoothly on the right. Thus, there is a
394
15.1. PRINCIPAL G-BUNDLES 395

smooth map P × G → P which we denote by (p, g) 7→ pg such that (pg1 ) g2 =


p(g1 g2 ), and if pg = p for some (p, g) ∈ P × G, then g is the group identity e.
For our purposes, one may assume that G is a matrix group, such as SU(N ) ,
SO(N ) , etc.. We assume that the quotient space M := P/G can be made into
a manifold such that the projection π : P → M is smooth. Also, we assume that
P is locally trivial. This means that each x ∈ M has a neighborhood U , such
that there is a diffeomorphism T : π −1 (U ) → U × G with T (pg) = (π(p) , s(pg))
where s : π −1 (U ) → G satisfies s(pg) = s(p) g for all p ∈ π −1 (U ) and g ∈ G.
In other words, P is locally equivalent to a product with the standard action.
For comparison with the somewhat complementary concept of vector bundles see
Appendix B (pp.712 ff).
Definition 15.1. If the conditions of the preceding paragraph hold, then we
say that π : P → M is a principal G-bundle with total space P , base space
M , structure group G, and projection π.
The diffeomorphism T : π −1 (U ) → U × G is known as a local trivialization. In
the case where U can be taken to be all of M (i.e., T : P = π −1 (M ) → M × G),
we say that principal G-bundle π : P → M is (globally) trivial. Let U be an open
subset of M , and let σ : U → P be a map such that π ◦ σ = id, where id denotes
the identity map on U . Then σ is called a local section. There is a one-to-one
correspondence between local sections and local trivializations. Indeed, given σ,
define T : π −1 (U ) → U × G by T (σ(x) g) = (x, g). Note that T is well-defined since
any p ∈ π −1 (U ) can be written uniquely as σ(x) g since G acts freely and transitively
on the fiber π −1 (π(p)). Conversely, given a local trivialization T , the equation
T (σ(x) g) = (x, g) serves to define σ (i.e., σ(x) = T −1 (x, g) g −1 = T −1 (x, e)). It
follows that a principal G-bundle is trivial precisely when it has a global section
(i.e., local section σ : U → P with U = M ). The corresponding statement for sphere
bundles is false. For example, the unit tangent bundle S(K) of a Klein bottle K is
a circle bundle which is not globally a product, since otherwise one could define a
frame field on K even though K is nonorientable. Nevertheless, K does have a unit
tangent vector field. This also shows that S(K) cannot be made into a principal
S 1 bundle, where S 1 is regarded as the group U (1) := eiθ : θ ∈ R . Indeed, an
orientation is precisely what is necessary in order to define a free S 1 action on
the unit tangent bundle S(M ) of a surface M (with metric). Thus, for orientable
surfaces M , S(M ) is trivial exactly when there is a unit tangent vector field.
Exercise 15.2. Recall that
SU(2) : = {A ∈ GL(2, C) : A∗ A = I2 , det A = 1}
  
a −b
: a, b ∈ C, |a| + |b| = 1 ∼
2 2
= = S3.
b a
Let G ∼= S 1 denote the subgroup of elements with b = 0 (i.e., a = eiθ ). Show that
the left coset space SU(2) /G = AS 1 : A ∈ SU(2) can be identified with S 2 so
that the quotient map Q : SU(2) → SU(2) /G may be regarded as a principal S 1 -
bundle π : SU(2) → S 2 (which is known as a Hopf bundle). Show that this bundle
is nontrivial. [Hint. Identify R3 with the space of traceless Hermitian matrices via
the Pauli map  
z x − iy
σ : r = (x, y, z) 7→ .
x + iy −z
396 15. GEOMETRIC PRELIMINARIES

For A ∈ SU(2), show that c(A)(r) := σ −1 (Aσ(r) A∗ ) defines a homomorphism


2
c : SU(2) → SO(3). Let π : SU(2) →  S be given by π(A) := c(A) e3 where

e3 := (0, 0, 1) and check that π Ae = π(A).]

2. Connections and Curvature


In Section 6.5 (pp.176 ff), we defined connections on vector bundles. Connec-
tions (or gauge fields) on principal G-bundles can be defined in various equivalent
ways. From a conceptual standpoint, the following definition is closest to our pre-
vious one and perhaps best:
Definition 15.3. a) A connection on a principal G-bundle π : P → M
smoothly assigns to each p ∈ P a subspace Hp of the tangent space Tp P , such
that π∗p : Hp → Tπ(p) M is an isomorphism and Rg∗ (Hp ) = Hpg , where Rg : P → P
is defined by Rg (p) = pg.
(b) We shall denote the space of connections on P by C(P ).
The so-called horizontal subspace Hp is complementary to the vertical subspace
Vp := Ker(π∗p ) which is the tangent space of the fiber π −1 (π(p)) at p. Thus, a
connection serves to select a horizontal complement to each vertical subspace in a
smooth G-invariant fashion, see Figure 15.1.

Hp
Vp
Tp P

P
¼ p

¼ {1(x)

M
x =¼(p) ¼

Figure 15.1. Selecting a horizontal complement Hp to each ver-


tical subspace Vp of the tangent space Tp P at p ∈ P

One way of defining Hp would be to let it be the subspaces of Tp P annihilated


by a differential 1-form of maximal rank with values in a fixed vector space having
15.2. CONNECTIONS AND CURVATURE 397

the same dimension as Vp , namely dim P − dim M = dim G. The natural choice for
this vector space is the Lie algebra of G which we denote by g. The Lie algebra of G
is the tangent space Te G of G at the identity e ∈ G, and for A, B ∈ g, there is a Lie
bracket [A, B] ∈ g. While we will not go into the definition of [A, B] for general Lie
groups, in the case matrix Lie groups G ⊆ GL(N, C) it is easy to describe. Indeed,
as GL(N, C) is an open subset of the linear space gl(N, C) of all N × N complex
matrices, the tangent space of TI G at the identity matrix I may be identified with
a subspace of gl(N, C), and for the A, B ∈ g = TI G ⊆ gl(N, C), the Lie bracket
[A, B] is just the commutator AB − BA. In all of what follows, we will assume that
G is a matrix Lie group. For a matrix A ∈ gl(N, C), we define

X 1 k
exp(A) := A .
k!
k=0

Then we show that g is the set, say s, of all A ∈ gl(N, C) such that exp(tA) ∈ G
for all t ∈ R. Since g := TI G, it is clear that s ⊆ g. To show g ⊆ s, suppose that
A ∈ g, and let A e denote the vector field on G defined by A eg := Lg∗ (A), where
0 0
Lg : G → G is given by Lg (g ) = gg . Since G is a matrix group where tangent
vectors are considered to reside in gl(N, C) and Lg is a linear transformation of
gl(N, C), A eg = Lg∗ (A) is simply gA. Now, t 7→ exp(tA) is the solution curve at I
of the vector field Ae on G, since
∞ ∞
!
X 1 k k X 1
d
dt (exp(tA)) = d
dt t A = tk−1 Ak
k! (k − 1)!
k=0 k=1

!
X 1 k−1 k−1
= t A A = exp(tA) A = A eexp(tA) .
(k − 1)!
k=1

In particular, exp(tA) ∈ G for all t ∈ R and hence g ⊆ s.


Definition 15.4. a) Note that for g ∈ G, the map Adg : G → G , given by
Adg (h) := ghg −1 fixes the identity I. Thus its derivative (Adg∗ )I at I is a linear
transformation of TI G or g. For A ∈ g, we have (at t = 0)
d
(Adg∗ )I (A) = dt Adg (exp(tA))
−1
d d
exp tgAg −1 = gAg −1 .
 
= dt g exp(tA) g = dt

b) We denote (Adg∗ )I (A) by adg (A) and the homomorphism ad : G → GL(g) ,


given by g 7→ adg is known as the adjoint representation of G.
Observe that for A, B ∈ g,
adexp(tA) (B) = exp(tA) B exp(−tA) = B + t(AB − BA) + O t2 ,


and so
d

dt adexp(tA) (B) t=0
= AB − BA = [A, B] .
Thus, the derivative ad∗I : g → End(g) of the map ad : G → GL(g) at I is given
by (ad∗I (A))(B) = [A, B]. It is convenient to denote ad∗I by ad, and so ad : g →
End(g) is given by
(ad(A))(B) = (ad∗I (A))(B) = [A, B] .
398 15. GEOMETRIC PRELIMINARIES

For a principal G-bundle π : P → M and any A ∈ g, there is a vector field A∗


(known as a fundamental vertical vector field ) defined at p ∈ P by
A∗p := d
dt (p exp(tA)) t=0 .
Definition 15.5. A connection 1-form for the principal G-bundle π : P → M
is a g-valued 1-form ω on P , such that for each A ∈ g, g ∈ G, p ∈ P and X ∈ Tp M ,
we have two conditions satisfied:
(C1 ) ω(A∗ ) = A and
(15.1)
(C2 ) ωpg (Rg∗ X) = adg−1 (ωp (X)) = g −1 ωp (X) g.
The Definitions 15.3 and 15.5 are related as follows. Given ω as in definition
15.5, define
Hp := {X ∈ Tp M : ωp (X) = 0} .
Then condition (C1 ) in (15.1) insures that π∗p : Hp → Tπ(p) M is an isomorphism,
and condition (C2 ) guarantees that Rg∗ (Hp ) = Hpg . Note that adg−1 in (C2 ) is
needed so that it is consistent with (C1 ), since
Rg∗ (A∗ ) = dtd d
pgg −1 exp(tA) g

(p exp(tA) g) = dt
d
 ∗
= dt pgAdg−1 (exp(tA)) = adg−1 (A) pg
and (C1 ) imply that
 ∗ 
ωpg (Rg∗ (A∗ )) = ωpg adg−1 (A) pg = adg−1 (A) = adg−1 (ωp (A∗ )) .

Of course, when G is abelian, adg−1 is the identity and (C2 ) says that ωpg is invariant
under Rg (i.e., Rg∗ ω = ω when G is abelian).

Exercise 15.6. Let G be a (closed) Lie subgroup of a matrix Lie Group G.


It can be verified that G/G naturally has the structure of a manifold and that
π : G → G/G is a principal G-bundle. Let g denote the Lie algebra of G. The

Maurer-Cartan form for G is the g-valued 1-form ω ∈ Ω1 G, g on G e given at
g by ω g (Lg∗ A) = ω g (gA) = A. Suppose that g = g ⊕ m where adg (m) = m for
all g ∈ G (i.e., m is an adG -invariant subspace of g). Let πg : g → g denote the
projection onto g along m. 
(a) Check that the form ω := πg ◦ ω ∈ Ω1 G, g is a connection 1-form for
π : G → G/G. Why do we need adg (m) = m?
(b) Return to Exercise 15.2 and use the preceding construction to explicitly find
a natural connection 1-form ω ∈ Ω1 (SU(2) , g) for the Hopf bundle SU(2) →
SU(2) /G.
Given a connection 1-form ω, we can decompose any X ∈ Tp P into horizontal

and vertical parts, X H and X V respectively, where ωp X H = 0, π∗p X V = 0,
and X = X H + X V . If φ is a k-form on P with values in a vector space W , then
we may define a new k-form φH on P by
φH (X1 , . . . , Xk ) := φ X1H , . . . , XkH .


We define the covariant derivative Dω φ of φ relative to ω to be the W -valued


(k + 1)-form
H
(15.2) Dω φ := (dφ) .
15.2. CONNECTIONS AND CURVATURE 399

The curvature Ωω of ω is simply the covariant derivative of ω relative to itself


H
(15.3) Ωω := Dω ω = (dω) .
Proposition 15.7. We have
(15.4) Ωω = Dω ω = dω + ω ∧ ω,
(ω ∧ ω)(X, Y ) := ω(X) ω(Y ) − ω(Y ) ω(X) .
Proof. Both sides of (15.4) agree on any pair (X, Y ) of vectors where one of
X V or Y V is 0, since then ω(X) = 0 or ω(Y ) = 0. Thus, it suffices to show that
both sides agree on a pair (A∗ , B ∗ ) of fundamental vertical fields. In this case,
(Dω ω)(A∗ , B ∗ ) = 0, and
dω(A∗ , B ∗ ) + (ω ∧ ω)(A∗ , B ∗ )
= A∗ [ω(B ∗ )] − B ∗ [ω(A∗ )] − ω([A∗ , B ∗ ]) + (ω ∧ ω)(A∗ , B ∗ )
= −ω([A∗ , B ∗ ]) + ω(A∗ ) ω(B ∗ ) − ω(B ∗ ) ω(A∗ )
= −ω([A∗ , B ∗ ]) + [A, B] .
Thus, it suffices to check that

(15.5) [A∗ , B ∗ ] = [A, B] .
We have (evaluating derivatives with respect to s and t at 0)
 
[A∗ , B ∗ ]p = dt
d
Rexp(−tA)∗ Bp∗ exp(tA)
d d

= dt Rexp(−tA)∗ ds p exp(tA) exp(sB)
d d
= dt ds (p exp(tA) exp(sB) exp(−tA))
d d
= ds dt p exp(tA) exp(sB) exp(−tA)
d d ∗
= ds pAdI∗ (A)(sB) = ds (p exp(s [A, B])) = [A, B]p ,
verifying (15.5). 
We mention that for general Lie groups (15.4) is written as
(15.6) Dω ω = dω + 1
2 [ω, ω] ,
where
1
2[ω, ω](X, Y ) := 12 ([ω(X) , ω(Y )] − [ω(Y ) , ω(X)])
is defined purely in terms of Lie brackets instead of commutators of matrices.
Exercise 15.8. We use the notation in Exercise 15.6. For A ∈ g, let A
e ∈


C T G denote the vector field on G , given by Ag := Lg∗ A = gA. Similarly, for
e
B ∈ g, we define
h B. i
e
(a) Show that A, B
e ^
e = [A, B].
(b) Use formula 6.10 (p. 175) to show that
  h i
dω A,e Be = −ω A, e = − [A, B] .
e B

(c) Conclude that the
 curvature
 Ωω ∈ Ω2 G, g of the connection ω in Exercise
15.6, is given by Ωω A,
e B e = πg ([A, B]).
(d) Consider the special case G = SU(2) and G = {exp(itσ(e3 )) : t ∈ R} , as in
400 15. GEOMETRIC PRELIMINARIES

that − 2i σ(e1 ) , − 2i σ(e2 ) , − 2i σ(e3 ) is a basis for su(2) and



Example 15.2. Check
− 2i σ(v) , − 2i σ(w) = − 2i σ(v × w) for
 3
 v,i w ∈R . iConclude that if m ⊆ su(2) in
Exercise 15.8 (b) is chosen to be span − 2 σ(e1 ) , − 2 σ(e2 ) , then at any g ∈ SU(2)
we have
Ωgω g − 2i σ(v) , g − 2i σ(w) = (e3 ·(v × w)) 2i σ(e3 ) .
 

The following notion of horizontal lift will be crucial in many key computations.
Definition 15.9. For principal G-bundle π : P → M with connection ω and a
vector field Y on M the vector field X on P , such that ω(X) = 0 and π∗ (X) = Y
is called the horizontal lift of Y .
Remark 15.10. Note that X is unique since π∗ : Hp → Tπ(p) M is an iso-
morphism. Moreover, for any g ∈ G, note that Rg∗ (X) satisfies π∗ (Rg∗ (X)) =
(π ◦ Rg )∗ (X) = π∗ (X) = Y and ω(Rg∗ (X)) = 0. Thus, Rg∗ (X) is also a horizontal
lift of Y , and by uniqueness Rg∗ (X) = X (i.e., horizontal lifts are Rg∗ -invariant).
In particular, for a fundamental vertical vector field A∗ (A ∈ g) and a horizontal
lift X, we have (at any p ∈ P )
[A∗ , X]p = dt
d d

(15.7) Rexp(−tA)∗ Xp exp(tA) t=0 = dt (Xp ) t=0 = 0.

3. Equivariant Forms and Associated Bundles


Associated Vector Bundles, Equivariance, and Basic Forms. Let
r : G → GL(W ) be a representation (i.e., a homomorphism), where GL(W ) denotes
the general linear group of a vector space W . For a principal G-bundle  π : P → M ,
there is a right action of G on P × W , given by (p, w) g = pg, r g −1 w . Let [p, w]
denote the orbit {(p, w) g : g ∈ G} of (p, w), and let P ×G W denote the quotient
space
P ×W
P ×G W := = {[p, w] : (p, w) ∈ P × W } .
G
It is not difficult to verify that πW : P ×G W → M, where πW ([p, w]) := π(p) , is
a vector bundle which is known as an associated vector bundle of P via r. Note
−1
that for p ∈ P and x = π(p), any two points in the fiber πW (x) have unique
representatives of the form (p, w1 ) and (p, w2 ), and it is easily verified that [p, w1 ] +
−1
[p, w2 ] := [p, w1 + w2 ] is a well-defined addition. The fibers πW (x) are isomorphic
to W , but not canonically so, since the isomorphism [p, w] 7→ w depends on the
choice of p ∈ π −1 (x).
Let s : M → P ×G W be a section, and define f : P → W by the equation
s(π(p)) = [p, f (p)]. Since
[p, f (p)] = s(π(p)) = s(π(pg)) = [pg, f (pg)] = [p, r(g) f (pg)] ,
f has the equivariance property
−1
f (pg) = r(g) f (p) .
Thus, we have an isomorphism (s 7→ f )
0
n o
−1
C ∞ (P ×G W ) ∼= Ω (P, W ) := f ∈ C ∞ (P, W ) : f (pg) = r(g) f (p) .
0
Ω (P, W ) is called the space of W -valued equivariant functions (0-forms) on P .
Whether one works with sections or equivariant functions, is largely a matter of
taste or convenience. We will use equivariant forms as well.
15.3. EQUIVARIANT FORMS AND ASSOCIATED BUNDLES 401

Definition 15.11. Let π : P → M be a principal G-bundle and let r : G →


GL(W ) be a representation. We denote the space of C ∞ , W -valued k-forms on
P by equivariantΩk (P, W ). For k > 0, the space of horizontal, equivariant,
k
W -valued k-forms on P , denoted by Ω (P, W ), consists of all α ∈ Ωk (P, W ), such
that for all vector fields X1 , . . . , Xk and g ∈ G, we have
(H) α(X1 , . . . , Xk ) = 0 if π∗ (Xi ) = 0 for some i ∈ {1, . . . , k} and
−1 −1
(E) α(Rg∗ X1 , . . . , Rg∗ Xk ) = r(g) α(X1 , . . . , Xk ) (i.e., Rg∗ α = r(g) α).

Remark 15.12. For the representation ad : G → GL(g), we may consider


k 1
Ω (P, g). However, a connection 1-form ω is not in Ω (P, g), since ω(A∗ ) = A
in violation of Condition (H) in Definition 15.11. Nevertheless, if ω 0 is another
1
connection, then ω − ω 0 ∈ Ω (P, g), since (ω − ω 0 )(A∗ ) = 0. In other words, the
1
space C(P ) of connections on P is an affine space based on Ω (P, g).

Exercise 15.13. (a) In the notation of the preceding Remark, while ω ∈ /


1 2
Ω (P, g), show that the curvature 2-form Ωω ∈ Ω (P, g), relative to the repre-
sentation ad : G → GL(g). [Hint. One may use (15.3) and (15.4) in conjunction
with condition(C2 ) in (15.1).]
(b) If G is abelian, deduce that Ωω = π ∗ Ωω ω 2
0 for a unique form Ω0 ∈ Ω (M, g),

where π is pull-back on forms induced by π : P → M .
(c) If ω is as in Exercise 15.8 (d) and π : SU(2) → S 2 (see also Exercise 15.2,
i
p. 395), show that the form Ωω 0 in (b) is 2 σ(e3) ν, where ν denotes the area 2-
2 i
form of S . [Hint. Show that πg∗ g − 2 σ(v) = c(g)(v × e3 ) and note that
νx (a, b) = x ·(a × b) for x ∈ S 2 and a, b ∈ Tx S 2 ⊂ R3 .]
0
Remark 15.14 (Basic Forms). Let s ∈ Ω (P, W ), η ∈ Ωk (M, R) and β = π ∗ η.
Define s ⊗ β ∈ Ωk (P, W ) by
(s ⊗ β)p (X1 , . . . , Xk ) := β(X1 , . . . , Xk ) s(p) = η(π∗ X1 , . . . , π∗ Xk ) s(p) .

Then s ⊗ β clearly meets Condition (H), and it meets Condition (E), since
(s ⊗ β)pg (Rg∗ X1 , . . . , Rg∗ Xk ) = η(π∗ (Rg∗ X1 ) , . . . , π∗ (Rg∗ Xk )) s(pg)
−1 −1
= η(π∗ X1 , . . . , π∗ Xk ) r(g) s(p) = r(g) (η(π∗ X1 , . . . , π∗ Xk ) s(p))
−1
= r(g) (s ⊗ β)p (X1 , . . . , Xk ) .
k
Thus, s ⊗ β ∈ Ω (P, W ), and we call such forms basic. Although not every α ∈
k k P
Ω (P, W ) is basic, any α ∈ Ω (P, W ) can be written as a finite sum i si ⊗ βi of
k
basic forms. Hence many facts concerning forms in Ω (P, W ) can be verified first
for basic forms, and then extended by linearity.

Covariant Differentiation and the General Bianchi Identity. Equivari-


ance is preserved by covariant differentiation, namely
k k+1
Dω : Ω (P, W ) → Ω (P, W ) .
For this, note that Dω α = dαH satisfies condition (H) in Definition 15.11, and since
 H
the distribution of horizontal subspaces is Rg∗ invariant, Rg∗ X H = Rg∗ (X) and
402 15. GEOMETRIC PRELIMINARIES

H
so Rg∗ β = Rg∗ β H for any form β on P . Then Dω α meets condition (E), since


  H H   H
H −1
Rg∗ (Dω α) = Rg∗ (dα) = Rg∗ dα = d Rg∗ α = d r(g) α
 H
−1 −1 H −1
= r(g) dα = r(g) (dα) = r(g) Dω α.
Moreover, there is a very convenient formula given in the following
k
Proposition 15.15. For a representation r : G → GL(W ) and α ∈ Ω (P, W ),
we have
(15.8) Dω α = dα + r0 (ω) ∧ α,
where r0 denotes the Lie algebra representation (i.e., the derivative of r : G →
GL(W ) at I) and (where σ runs over all permutations of {1, . . . , k})
1 X σ
(r0 (ω) ∧ α)(X1 , . . . , Xk+1 ) := (−1) r0 (ω(Xσ1 )) α Xσ2 , . . . , Xσk+1 .

k! σ

Proof. To verify (15.8), we need to show that


(15.9) Dω α(X1 , . . . , Xk+1 ) = dα(X1 , . . . , Xk+1 ) + (r0 (ω) ∧ α)(X1 , . . . , Xk+1 )
when each Xi is a fundamental vertical field or horizontal. If all of the Xi are
horizontal, then both sides of (15.9) agree, since ω(Xi ) = 0. By (6.11), p. 175,
k+1
X h  i
i+1
(15.10) dα(X1 , . . . , Xk+1 ) = (−1) Xi α X1 , . . . , X
ci , . . . , Xk+1
i=1
k+1
X  
i+j
+ (−1) α [Xi , Xj ] , X1 , . . . , X
ci , . . . , X
cj , . . . , Xk+1 .
1≤i<j≤k+1

Note that for fundamental vertical fields A∗ and B ∗ , we have [A∗ , B ∗ ] = [A, B] .
Thus, both sides of (15.9) are zero when two or more of the Xi are vertical. In the
remaining case, one of the Xi is vertical (say X1 = A∗ ) and the rest are horizontal.
Then, the left side of (15.9) is 0 and it remains to verify the right side is 0. For
this we may assume that the horizontal X2 , . . . , Xk+1 are horizontal lifts of vector
fields Y2 , . . . , Yk+1 on M . By (15.7), [A∗ , X2 ] = · · · = [A∗ , Xk+1 ] = 0 and the right
side of (15.9) is (at p ∈ P )
A∗ [α(X2 , . . . , Xk+1 )] + r0 (ω(A∗ )) α(X2 , . . . , Xk )
d

= dt αp exp tA R(exp tA)∗ X2 , . . . , R(exp tA)∗ Xk+1 t=0
+ r0 (A) α(X2 , . . . , Xk )
d ∗ 0

= dt Rexp tA α(X2 , . . . , Xk+1 ) t=0 + r (A) α(X2 , . . . , Xk )
d 0
= dt r(exp(−tA)) αp (X2 , . . . , Xk+1 ) t=0 + r (A) α(X2 , . . . , Xk )
= r0 (−A) α(X2 , . . . , Xk ) + r0 (A) α(X2 , . . . , Xk ) = 0. 
k
The space Ω (P, W ) can be identified with the space Ωk (P ×G W ) of k-forms
with values in the vector bundle P ×G W as follows. If Y1 , . . . , Yk are vector fields on
k
M , with horizontal lifts Ye1 , . . . , Yek , then it is easy to check that for α ∈ Ω (P, W ),
we have   0
α Ye1 , . . . , Yek ∈ Ω (P, W ) ∼ = C ∞ (P ×G W ) .
15.3. EQUIVARIANT FORMS AND ASSOCIATED BUNDLES 403

Then we define a form αM ∈ Ωk (P ×G W ) at x = π(p) by


h  i
αM (Y1 , . . . , Yk ) = p, α(p) Ye1 , . . . , Yek .
k
Conversely, the same equation can be used to define α ∈ Ω (P, W ) for a given
αM ∈ Ωk (P ×G W ). It is easy to see that the correspondence
k
(15.11) Ω (P, W ) ∼
= Ωk (P ×G W ) (α ↔ αM )
is actually independent of the choice of connection ω. Moreover, via this correspon-
k k+1
dence Dω : Ω (P, W ) → Ω (P, W ) provides us with a corresponding operator
(also denoted Dω ), say
Dω : Ωk (P ×G W ) → Ωk+1 (P ×G W ) .
Remark 15.16. Since it would be cumbersome to adhere to the notation αM ,
k
we shall simply use the same symbol α, whether we regard α as in Ω (P, W ) or as
k
in Ω (P ×G W ). The context will either be clear, irrelevant, or made explicit.
Of fundamental importance is
Proposition 15.17 (General Bianchi Identity). If ω is any connection on a
2
principal G-bundle π : P → M and Ωω ∈ Ω (P, g) denotes the curvature of ω, then
(15.12) Dω (Ωω ) = 0.
Proof. We compute
Dω (Ωω ) = dΩω + ad(ω) ∧ Ωω = dΩω + ω ∧ Ωω − Ωω ∧ ω
= d(dω + ω ∧ ω) + ω ∧(dω + ω ∧ ω) − (dω + ω ∧ ω) ∧ ω
= d(ω ∧ ω) + ω ∧ dω − (dω) ∧ ω
= ((dω) ∧ ω − ω ∧ dω) + ω ∧ dω − dω ∧ ω = 0. 

Note that Dω (Dω ω) = Dω (Ωω ) = 0. However, unlike ordinary exterior dif-


ferentiation d : Ωk (M, R) → Ωk+1 (M, R) which satisfies d2 = 0, the composition
k k+2
Dω ◦ Dω : Ω (P, W ) → Ω (P, W ) is not zero in general, and the curvature Ωω is
the obstruction in the following sense.
Proposition 15.18. If ω is any connection on a principal G-bundle π : P →
2 k
M , Ωω ∈ Ω (P, g) denotes the curvature of ω, and α ∈ Ω (P, W ), we have
(15.13) Dω (Dω α) = r0 (Ωω ) ∧ α.
Proof. We compute
Dω (Dω α) = Dω (dα + r0 (ω) ∧ α) = d(dα + r0 (ω) ∧ α) + r0 (ω) ∧(dα + r0 (ω) ∧ α)
= d(r0 (ω) ∧ α) + r0 (ω) ∧ dα + r0 (ω) ∧(r0 (ω) ∧ α)
= (d(r0 (ω))) ∧ α − r0 (ω) ∧ dα + r0 (ω) ∧ dα + (r0 (ω) ∧ r0 (ω)) ∧ α
= (r0 (dω) + r0 (ω) ∧ r0 (ω)) ∧ α
= r0 (dω + ω ∧ ω) ∧ α = r0 (Ωω ) ∧ α. 
404 15. GEOMETRIC PRELIMINARIES

Inner Products, Hodge Star Operator, and Formal Adjoints. In order


to define a formal adjoint to Dω , we need to introduce some inner products. Let h
be a Riemannian metric on M , and suppose that K is an inner product on W for
which r : G → GL(W ) is orthogonal ; i.e., for all g ∈ G,
r(g) ⊆ O(W ) := {A ∈ GL(W ) : K(Aw1 , Aw2 ) = K(w1 , w2 )} .
Such a K always exists if G is compact, by an averaging argument. Then we define
−1
an inner product on each fiber πW (x) of the associated vector bundle πW : P ×G
W → M by
h[p, w1 ] , [p, w2 ]ix := K(w1 , w2 ) ,
which is independent of the choice of representatives (p, wi ) ∈ [p, w1 ] by the orthog-
onality of r : G → GL(W ). This pointwise inner product gives us a pairing
h·, ·i : C ∞ (P ×G W ) × C ∞ (P ×G W ) → C ∞ (M, R)
by simply defining hs, ti(x) := hs(x) , t(x)ix for s, t ∈ C ∞ (P ×G W ). We can show
that
(15.14) d hs, ti = hDω s, ti + hs, Dω ti ∈ Ω1 (M, R) ,
where the inner product on the right is just between the values in P ×G W . First
note that for A ∈ g, the product rule for differentiation yields (at u = 0)
0= d
du K(r(exp uA) w1 , r(exp uA) w2 ) = K(r0 (A) w1 , w2 ) + K(w1 , r0 (A) w2 ) .
0
Then regarding s, t ∈ Ω (P, W ) and using (15.8), we get
d(K(s, t)) = K(ds, t) + K(s, dt)
= K(ds, t) + K(s, dt) + K(r0 (ω) s, t) + K(s, r0 (ω) t)
= K(ds + r0 (ω) s, t) + K(s, dt + r0 (ω) t)
= K(Dω s, t) + K(s, Dω t) .
This equality of right-invariant R-valued 1-forms on P yields (15.14). Using the
Riemannian metric h, there is a pairing
(15.15) h·, ·i : Ωk (P ×G W ) × Ωk (P ×G W ) → C ∞ (M, R) ,
such that for s, t ∈ C ∞ (P ×G W ) and β, γ ∈ Ωk (M, R)
hs ⊗ β, t ⊗ γi = hs, ti h(β, γ) ,
where h(β, γ) is given locally by
1 i1 ···ik 1 i1 j1
(15.16) h(β, γ) := β γi1 ···ik := h · · · hik jk βj1 ···jk γi1 ···ik .
k! k!
Here, βj1 ···jk = β(∂j1 , . . . , ∂jk ) for local coordinate vector fields ∂1 , . . . , ∂n and the
hij are the entries of the inverse of the matrix [hij ], where hij := h(∂i , ∂j ). In
(15.16) and elsewhere we adopt the Einstein summation convention where repeated
indices on different levels are assumed to be summed from 1 to n = dim M . Rather
than introducing sections with compact support, let us assume that M is compact.
Then we have the inner product
(15.17) (·, ·) : Ωk (P ×G W ) × Ωk (P ×G W ) → R,
given by Z
(α1 , α2 ) := hα1 , α2 i |νh | ,
M
15.3. EQUIVARIANT FORMS AND ASSOCIATED BUNDLES 405

where |νh | denotes the density on M relative to h, given locally in coordinates


x1 , · · · , xn by
1/2
|νh | = |det(hij )| dx1 · · · dxn .
For α ∈ Ωk (P ×G W ), we define
2 2
kαk := (α, α) ∈ R and |α| := hα, αi ∈ C ∞ (M, R) .
Suitable modifications can be made to handle the case where W is complex, with
Hermitian scalar product K and r : G → GL(W ) is a unitary representation.
To explicitly construct a formal adjoint of Dω on Ωk (P ×G W ), we introduce
the (Hodge) star operator (for k ∈ {1, . . . , n = dim M }
∗ : Ωm (M, R) → Ωn−m (M, R) for m ∈ {1, . . . , n = dim M } .
In order to define ∗, we need to assume that M is oriented
 with volume form given
locally in an oriented coordinate system x1 , · · · , xn by
1/2
νh := |det [hij ]| dx1 ∧ · · · ∧ dxn .
Recall our previous definition of the linear ∗-operator in Equation (6.4) (p.171) that
was discussed in Exercise 6.20 (p.172). In agreement with this, the operator ∗ is
defined to be the unique linear map, such that for all α, β ∈ Ωm (M, R)
(15.18) α ∧ ∗β = hα, βi νh .
Note that ∗ can be defined pointwise. There is also a local formula (proven in [59,
p. 5])
1 1/2
(15.19) (∗β)jm+1 ···jn = |det [hij ]| β j1 ···jm εj1 ···jm jm+1 ···jn ,
m!
where ε is antisymmetric in its indices with ε12···n = 1 and νh (∂1 , . . . , ∂n ) > 0. In
[59, p. 5-6] it is also shown that for β ∈ Ωm (M, R),
m(n−m)
∗2 β := ∗(∗β) = sign(det [hij ])(−1) β.
For a Riemannian (positive definite) metric h, we have sign(det [hij ]) = 1, whereas
for h Lorentzian sign(det [hij ]) = −1. In particular, for n = dim M = 4, note that
∗2 = Id on Ω2 (M, R) for Riemannian h, while ∗2 = − Id for Lorentzian h. Thus,
for Riemannian h, one has a decomposition
Ω2 (M, R) = Ω2+ (M, R) ⊕ Ω2− (M, R) ,
where
Ω2± (M, R) := β ∈ Ω2 (M, R) : ∗β = ±β .


Forms in Ω2+ (M, R) are called self-dual, while forms in Ω2− (M, R) are called anti-
self-dual. For a Lorentzian 4-manifold, there is a similar notion, but only after
complexification where we can decompose Ω2 (M, C) into the ±i eigenspaces of ∗.
Unless otherwise stated, we assume that h is Riemannian. Of course, we can extend
the notion of star operator to the spaces Ωm (P ×G W ) ∼
= C ∞ (P ×G W )⊗Ωm (M, R)
via ∗(s ⊗ β) = s ⊗ ∗β. Moreover, for dim M = 4, we still have a decomposition
Ω2 (P ×G W ) = Ω2+ (P ×G W ) ⊕ Ω2− (P ×G W )
into self-dual and anti-self-dual 2-forms. Note that if the orientation of M is re-
versed, then according to (15.18), ∗ changes sign, and Ω2+ (·) and Ω2− (·) are inter-
changed. Also, for dim M = 4, ∗ : Ω2 (·) → Ω2 (·) is invariant under a conformal
change of metric. Indeed, if h is replaced by λh for a positive λ ∈ C ∞ (M, R), in the
406 15. GEOMETRIC PRELIMINARIES

1/2
local formula (15.19) |det(hij )| gains a factor of λ2 , while β j1 j2 gains a factor of
λ from the raising of the two indices (since hij becomes λ−1 hij ).
−2
0
For s, s0 ∈ C ∞ (P ×G W ), β ∈ Ωm (M, R) and β 0 ∈ Ωm (M, R), the following
definition is convenient
K((s ⊗ β) ∧(s0 ⊗ β 0 )) := K(s, s0 ) β ∧ β 0 .
0
Then for α ∈ Ωm (P ×G W ) and α0 ∈ Ωm (P ×G W ),
0
K(α ∧ α0 ) ∈ Ωm+m (M, R)
is defined by linearity. We have
m
(15.20) d(K(α ∧ ∗α0 )) = K(Dω α ∧ α0 ) + (−1) K(α ∧ Dω α0 ) ,
since (using (15.14))
dK((s ⊗ β) ∧(s0 ⊗ β 0 )) = d(K(s, s0 )) ∧ β ∧ β 0
m
+ K(s, s0 ) dβ ∧ β 0 + K(s, s0 )(−1) β ∧ dβ 0
= K(Dω s, s0 ) ∧ β ∧ β 0 + K(s, Dω s0 ) ∧ β ∧ β 0
m
+ K(s, s0 ) dβ ∧ β 0 + K(s, s0 )(−1) β ∧ dβ 0
= K((Dω s ∧ β + s ⊗ dβ) ∧(s0 ⊗ β 0 ))
m
+ (−1) K((s ⊗ β) ∧(Dω s0 ∧ β 0 + s ⊗ dβ 0 ))
= K(Dω (s ⊗ β) ∧(s0 ⊗ β 0 ))
m
+ (−1) K((s ⊗ β) ∧ Dω (s0 ⊗ β 0 )) .
Moreover, if m0 = m, we have
K(α ∧ ∗α0 ) = hα, α0 i vh .
Proposition 15.19. The formal adjoint of
Dω : Ωm (P ×G W ) −→ Ωm+1 (P ×G W )
on a compact, oriented, Riemannian n-manifold M is the covariant codifferential
δ ω : Ωm+1 (P ×G W ) −→ Ωm (P ×G W ) ,
given by
nm
(15.21) δ ω := − (−1) ∗ Dω ∗
In other words, for α ∈ Ωm (P ×G W ) and α0 ∈ Ωm+1 (P ×G W ), we have
(Dω α, α0 ) = (α, δ ω α0 ) .
Proof.R If we show that dγ = (hDω α, α0 i − hα, δ ω α0 i) vh , then (Dω α, α0 ) −
(α, δ ω α0 ) = M dγ = 0 by Stoke’s Theorem. Using (15.20), we compute
m
dγ = d(K(α ∧ ∗α0 )) = K(Dω α ∧ ∗α0 ) + (−1) K(α ∧ Dω (∗α0 ))
 
m (n−m)m 2
= K(Dω α ∧ ∗α0 ) + (−1) K α ∧(−1) ∗ Dω (∗α0 )
nm
= K(Dω α ∧ ∗α0 ) + K(α ∧ ∗((−1) ∗ Dω (∗α0 )))
= K(Dω α ∧ ∗α0 ) − K(α ∧ ∗(δ ω α0 )) = (hDω α, α0 i − hα, δ ω α0 i) vh ,
as required. 
15.3. EQUIVARIANT FORMS AND ASSOCIATED BUNDLES 407

To obtain formulas for Dω and δ ω in local coordinates, let σ : U → P be a


local section on a coordinate neighborhood U , and let α ∈ Ωm (P ×G W ). For
−1
each x ∈ U , we have an isomorphism (P ×G W )x = πW (x) → W , given by
[σ(x) , w] 7→ w. This yields an isomorphism Ω ((P ×G W )| U ) ∼
m
= Ωm (U, W ) which
k
we denote by α 7→ α e. One can easily check that if α ∈ Ω (P, W ) is the equivariant
form corresponding to α, then α e = σ ∗ α. By (15.8), we have Dω α = dα + r0 (ω) ∧ α
and
D]ω α = σ ∗ (D ω α) = σ ∗ (dα + r 0 (ω) ∧ α)

= d(σ ∗ α) + r0 (σ ∗ ω) ∧ σ ∗ α = de α + r0 (σ ∗ ω) ∧ α
e.
1 m

In the local coordinates x , . . . , x on U we may write
1 X
α
e= ei1 ...im dxi1 ∧ · · · ∧ dxim ,
α
m!
where it is assumed that αei1 ...im is antisymmetric in i1 , . . . , im . Then
(15.22)
  Xn    
k+1
D
] ωα = (−1) ∂jk α ej1 ...jbk ...jm+1 + r0 (σ ∗ ω(∂jk )) α
ej1 ...jbk ...jm+1 ,
j1 ...jm+1
k=1

where jbk means that jk is omitted. Using ∗f


α = ∗e
α, we have
n(m−1) g
− (−1) δω α = ∗D
^ ω ∗ α = ∗D
^ ω ∗α

(15.23) ∗α) + r0 (σ ∗ ω) ∧ ∗f
= ∗(d(f α) + r0 (σ ∗ ω) ∧ ∗e
α) = ∗(d(∗e α) .
While it is possible to get this formula with (15.19), in order to find the components
(δ ω α)i1 ...im−1 , it is easier to compute the formal adjoint of the operator α e 7→ D
] ω α,

using (15.22) and integration by parts assuming that α is compactly supported in


U . The final result is
   
ωα −1/2 1/2 ij1 ...jm−1
δg = − |h| hi j · · · hi
1 1 j ∂i |h|
m−1 m−1
α
e
i1 ...im−1

(15.24) − r0 ((σ ∗ ω)(∂i )) α


eii1 ...im−1 .
Associated principal bundles induced by Lie group homomorphisms.
The construction of vector bundles associated to a given principal G-bundle π : P →
M via a representation r : G → GL(W ) can be generalized to the case where r is
replaced by left action of G on any manifold F , say r : G → C ∞ (F, F ). Of particular
use to us in applications will be the case where G acts on a second group G0 via
g · g 0 := γ(g) g 0 where γ : G → G0 is a homomorphism. The following proposition
shows that in this case the associated bundle is a principal G0 -bundle P 0 → M .
Also, a connection for π : P → M gives rise to a connection for π : P 0 → M .
Moreover, an equivariant map between vector spaces V and V 0 in representations
of G and G0 gives rise to a vector bundle morphism between the associated vector
bundles P ×G V and P 0 ×G0 V 0 .
Proposition 15.20. For Lie groups G and G0 , let π : P → M be a principal
G-bundle and let γ : G → G0 be a homomorphism. Then there is a canonically
constructed principal G0 -bundle π 0 : P 0 → M . Moreover, there is a canonical map
Γ : P → P 0 which is γ-equivariant, in the sense that Γ(pg) = Γ(p) γ(g) (i.e., Γ◦Rg =
Rγ(g) ◦ Γ). For any p0 ∈ P 0 , Γ−1 (p0 ) is an orbit of the action of Ker γ on P , so that
Γ is an embedding only if Ker γ = 0.
408 15. GEOMETRIC PRELIMINARIES

Proof. We define the principal G0 -bundle π 0 : P 0 → M as follows. There is a


right action R : (P × G0 ) × G → P × G0 of G on P0 × G0 , given by
R(g)(p, g 0 ) = (p, g 0 ) · g := pg, γ g −1 g 0 .
 

Let the orbit of (p, g 0 ) be [p, g 0 ] := {(p, g 0 ) · g : g ∈ G} and let


P × G0
P 0 := P ×G G0 := = {[p, g 0 ] : p ∈ P, g ∈ G0 } .
G
Define an action of G0 on P by [p, g 0 ] · h0 := [p, g 0 h0 ] for h0 ∈ G0 . Since
pg, h g −1 g 0 h0 = [p, g 0 h0 ] ,
  

this action is well defined. The action is free, since


[p, g 0 h0 ] = [p, g 0 ] =⇒ (p, g 0 h0 ) = pg, γ g −1 g 0 for some g ∈ G
 

=⇒ g = e and g 0 h0 = γ g −1 g 0 = g 0 =⇒ h0 = e0 ,


where e ∈ G and e0 ∈ G0 are the identities. Then π 0 : P 0 → M (with π 0 ([p, g 0 ]) =


π(p)) is a principal G0 -bundle. Let Γ : P → P 0 be given by Γ(p) := [p, e0 ]. We check
that Γ is γ-equivariant:
Γ(pg) = [pg, e0 ] = pgg −1 , γ(g) e0 = [p, e0 ] γ(g) = Γ(p) γ(g) .
 

Note that Γ−1 (p0 ) is an orbit of the action of Ker γ on P : For p1 , p2 ∈ P ,


Γ(p1 ) = Γ(p2 ) = p ⇐⇒ [p1 , e0 ] = [p2 , e0 ]
⇐⇒ ∃g ∈ G, s.t. p1 g, γ g −1 e0 = (p2 , e0 )
 

⇐⇒ p2 = p1 g and γ g −1 = e0 (i.e., g ∈ Ker γ).





Remark 15.21. If Pe0 is a another principal G0 -bundle and Γ e : P → Pe0 is γ-


0 0 0 e g 0 . This
equivariant, then P is isomorphic to P via the bijection Γ(p) g ←→ Γ(p)
e
0
bijection is well defined (and then clearly G -equivariant). Indeed, for π(p1 ) =
π(p2 ), we have p2 = p1 g for some g ∈ G, and
Γ(p1 ) g10 = Γ(p2 ) g20 ⇐⇒ Γ(p1 ) g10 = Γ(p1 g) g20 = Γ(p1 ) γ(g) g20
⇐⇒ g10 = γ(g) g20 ⇐⇒ Γ(p
e 1 ) g 0 = Γ(p
1
e 1 ) γ(g) g 0 = Γ(p
2
e 2 ) g0 .
2

With regard to induced connections, we have


Proposition 15.22. Let γ 0 : g → g0 denote the Lie algebra map for γ : G → G0 .
For any connection 1-form ω ∈ Ω1 (P, g) on P , there is a uniqueconnection
 1-form
0
ω 0 ∈ Ω1 (P 0 , g0 ), such that Γ∗ ω 0 = γ 0 ◦ ω. Moreover, we have Γ∗ Ωω = γ 0 ◦ Ωω .

Proof. Since Γ◦Rg = Rγ(g) ◦Γ for all g ∈ G, the G-invariant distribution H of


ω-horizontal subspaces is mapped by Γ∗ : P → P 0 to a well-defined γ(G)-invariant
distribution of horizontal subspaces, say Γ(H) on Γ(P ). Via the various Rg0 ∗ for
g 0 ∈ G0 , Γ(H) uniquely extends to a G0 -invariant horizontal distribution, say H0 ,
on all of P 0 . Then H0 determines a connection 1-form ω 0 on P 0 . By definition of
ω 0 , we have Γ∗ ω 0 = 0 on the horizontal subspaces of ω. Thus, Γ∗ ω 0 = γ 0 ◦ ω on H.
As for vertical vectors, if A ∈ g, we have
Γ∗ A∗p = dt d d

Γ(p exp(tA)) = dt Γ(p) γ(exp(tA))
0 ∗
= d
dt (Γ(p) exp(tγ (A))) = γ 0 (A)Γ(p) , and so
15.4. GAUGE TRANSFORMATIONS 409

∗
(Γ∗ ω 0 )(A∗ ) = ω 0 (Γ∗ (A∗ )) = ω 0 γ 0 (A) = γ 0 (A) = (γ 0 ◦ ω)(A∗ ) .
Hence, Γ∗ ω 0 = γ 0 ◦ ω, and uniqueness is clear. Moreover,
 0
Γ∗ Ωω = Γ∗ dω 0 + 12 [ω 0 , ω 0 ] = dΓ∗ ω 0 + 21 [Γ∗ ω 0 , Γ∗ ω 0 ]


= d(γ 0 ◦ ω) + 12 [γ 0 ◦ ω, γ 0 ◦ ω] = γ 0 ◦ dω + 21 [ω, ω]


= γ 0 ◦ Ωω . 
Another useful fact concerns associated bundles.
Proposition 15.23. Suppose that we have representations
r : G → GL(V ) and r0 : G0 → GL(V 0 )
and a linear map φ : V → V 0 which is equivariant, in the sense that
φ(r(g)(v)) = r0 (γ(g))(φ(v)) .
Then there is a vector bundle morphism
φ× : P ×G V → P 0 ×G0 V 0 , given by φ× ([p, v]) = [Γ(p) , φ(v)] .
Proof. Note that φ× is well-defined, since
[p1 , v1 ] = [p2 , v2 ]
p2 g, r g −1 (v2 ) for some g ∈ G
 
⇒ (p1 , v1 ) =
⇒ ([p1 , e0 ], φ(v1 )) = ([p2 g, e0 ], φ r g −1 v2 ) for some g ∈ G
 
 
−1
⇒ ([p1 , e0 ], φ(v1 )) = ([p2 , e0 ]γ(g) , r0 γ(g) (φ(v2 ))) for some g ∈ G
⇒ [[p1 , e0 ], φ(v1 )] = [[p2 , e0 ], φ(v2 )]
⇒ [Γ(p1 ) , φ(v1 )] = [Γ(p2 ) , φ(v2 )] . 
Remark 15.24. It is easy to see that φ× is injective (or surjective) if and only
if φ is injective (or surjective). In particular, if φ is an isomorphism, then so is
φ× : P ×G V → P ×G0 V 0 .

4. Gauge Transformations
While physicists speak of gauge transformations of particle fields and gauge
potentials, each of these is induced by gauge transformations of a principal bundle
defined as follows.
Distinguishing Gauge Transformations from Automorphisms.
Definition 15.25. A gauge transformation of a principal G-bundle π : P →
M is a diffeomorphism F : P → P , such that for all p ∈ P and g ∈ G,
(G1 ) F (pg) = F (p) g and
(G2 ) π(F (p)) = π(p) .
We denote the group of gauge transformations by GA(P ).
Remark 15.26. Condition (G2 ) implies that the fibers are mapped into them-
selves. If (G2 ) were dropped, then (G1 ) and the fact that F is a diffeomorphism
imply that there is a diffeomorphism f : M → M , such that π(F (p)) = f (π(p)). In
this more general case (i.e., if (G2 ) is dropped), F is called an automorphism of
P . We denote the group of automorphisms of P by Aut(P ).
410 15. GEOMETRIC PRELIMINARIES

Aut(P ) acts to the left on the space C(P ) of connection 1-forms on P , as


k
well as the spaces Ω (P, W ) of horizontal, equivariant k-forms on P (relative to a
representation r : G → GL(W )) via pull-back:
(
−1 ∗
 C(P )
F · α := F α, α ∈ k
Ω (P, W ) .
To prove that F · α ∈ C(P ) for α ∈ C(P ), note that F −1 preserves the fundamental
vertical fields A∗ for A ∈ g, since F −1 (p exp tA) = F −1 (p) exp tA. Thus,
 ∗ 
F −1 α (A∗ ) = α F∗−1 (A∗ ) = α(A∗ ) = A.


∗
(i.e., Condition (C1 ) of (15.1) is met by F −1 α). Also, since F −1 ◦Rg = Rg ◦F −1 ,
 ∗  ∗ ∗  ∗ 
(15.25) Rg∗ F −1 α = F −1 Rg∗ (α) = F −1 g −1 αg = g −1 F −1 α g,


k
so that condition (C2 ) of (15.1) is met. It is also easy to prove that F ·α ∈ Ω (P, W )
k
for α ∈ Ω (P, W ). Indeed, for any fundamental vertical field A∗ ,
∗
F · α(A∗ , · · · ) = F −1 α(A∗ , · · · ) = α F −1 ∗ (A∗ ) , · · · = 0,
 

since we have observed that F −1 ∗ (A∗ ) is vertical. Also Rg∗ (F · α) = r g −1 (F · α)


 
k
by the same sort of computation as (15.25). Since Ω (P, W ) ∼ = Ωk (P ×G W ) (see
k
(15.11)), GA(P ) also acts on Ω (P ×G W ). Recall (see (15.15)) that if the repre-
sentation r : G → GL(W ) is orthogonal relative to some inner product k on W and
there is a Riemannian metric h on M , then there is a pairing
h·, ·i : Ωm (P ×G W ) × Ωm (P ×G W ) → C ∞ (M, R) .
It is straightforward to check that this pairing is invariant under GA(P ), and it
follows that GA(P ) acts by isometries on the pre-Hilbert spaces Ωm (P ×G W ); for
k
this, it may be easier to work with the related pairing on Ω (P, W ).
k
Proposition 15.27. For F ∈ GA(P ) and α ∈ Ω (P, W ), we have
F ·(Dω α) = DF ·ω (F · α) .
Proof. Since wedge and d commute with pull-back,
∗ ∗
F ·(Dω α) = F −1 (Dω α) = F −1 (dα + r(ω) ∧ α)
∗ ∗
= F −1 (dα) + F −1 (r(ω) ∧ α)
 ∗    ∗  ∗ 
= d F −1 α + r F −1 ω ∧ F −1 α = DF ·ω (F · α) . 

The Group of Gauge Transformations, Revisited. There is another way


of looking at GA(P ). Let
C(P, G) := f ∈ C ∞ (P, G) : f (pg) = g −1 f (p) g = Adg−1 f (p) .

(15.26)
Since we have assumed that G is a matrix Lie group (i.e., G ⊆ GL(N, C)), the
adjoint action of G on itself (i.e., g · g0 = Adg (g0 ) = gg0 g −1 ) can be regarded as a
representation Ad : G → GL(gl(N, C)). Hence
0
C(P, G) ⊆ Ω (P, gl(N, C)) ∼
= Ω0 (P ×G gl(N, C)) .
15.4. GAUGE TRANSFORMATIONS 411

Note that C(P, G) has an induced group operation given, for f1 , f2 ∈ C(P, G), by
(f1 f2 )(p) = f1 (p) f2 (p).
Proposition 15.28. There is an isomorphism Φ : C(P, G) → GA(P, G) of
groups given by
−1
Φ(f )(p) := pf (p) ,
Proof. Note that Φ(f ) ∈ GA(P, G), since
−1 −1 −1
Φ(f )(pg) = pgf (pg) = pg g −1 f (p) g = pf (p) g = Φ(f )(p) g,
and Φ is a homomorphism, since
−1
Φ(f1 f2 )(p) = p(f1 (p) f2 (p))
 
−1 −1
= p f2 (p) f1 (p) = (Φ(f1 ) ◦ Φ(f2 ))(p) .
−1
For F ∈ GA(P, G), define Ψ(F ) ∈ C ∞ (P, G) by F (p) = pΨ(F )(p) . Then Ψ(F ) ∈
C(P, G), since
−1 −1
pgΨ(F )(pg) = F (pg) = F (p) g = pΨ(F )(p) g
−1 −1
⇒ gΨ(F )(pg) = Ψ(F )(p) g
−1
⇒ Ψ(F )(pg) = g Ψ(F )(p) g.
−1 −1
Note that Ψ is the inverse of Φ, since pΨ(Φ(f ))(p) = Φ(f )(p) = pf (p) and
−1
Φ(Ψ(F ))(p) = pΨ(F )(p) = F (p). 

The Action of Gauge Transformations on Connections. After the pre-


ceding algebra, now the differential geometry begins.
Proposition 15.29. Let f ∈ C(P, G) . For ω ∈ C(P ) and X ∈ Tp P , we have
−1
(Φ(f ) · ω)(X) = f (p) f −1 ∗p (X) + f (p) ω(X) f (p) .

(15.27)
m
For α ∈ Ω (P, W ) and X1 , . . . , Xk ∈ Tp P, we have
(15.28) (Φ(f ) · α)p (X1 , . . . , Xm ) = r(f (p))(αp (X1 , . . . , Xm )) .

Proof. Let γ : R →P be a curve with γ 0 (0) = X ∈ Tp P, and let f ∈ C(P, G).


At t = 0, we have
d d d d
Φ(f )∗ (X) = dt Φ(f )(γ(t)) = dt γ(t) f (γ(t)) = dt pf (γ(t)) + dt γ(t) f (p)
d −1 d
= dt pf (p) f (p) f (γ(t)) + dt Rf(p) (γ(t))
 ∗
−1
= f (p) f∗p (X) + Rf(p)∗ (X) .
pf(p)
−1
We have used the fact that t 7→ f (p) f (γ(t)) is a curve through I ∈ G with tangent
−1 d −1
vector f (p) f∗p (X) ∈ TI G = g, and hence dt pf (p) f (p) f (γ(t)) coincides with
−1 ∗ −1
the fundamental vertical field f (p) dfp (X) at pf (p). Since Φ(f ) = Φ f −1 ,


we also have
   ∗
−1
(X) = f (p) f −1 ∗p (X)

Φ(f ) −1
+ Rf(p)−1 ∗ (X) .
∗ pf(p)
412 15. GEOMETRIC PRELIMINARIES

Thus,
 ∗    
−1 −1
(Φ(f ) · ω)(X) = Φ(f ) ω (X) = ω Φ(f ) (X)

 ∗ 
−1

= ω f (p) f ∗p
(X) + R −1 (X)
f(p) ∗
pf(p)−1

= f (p) f −1 ∗p (X) + Rf(p) ∗



−1 ω(X)

−1
 −1
= f (p) f ∗p
(X) + f (p) ω(X) f (p) .
Using the fact that α vanishes on vertical vectors, we also have
 
−1∗
(Φ(f ) · α)p (X1 , . . . , Xm ) = Φ(f ) α (X1 , . . . , Xm )
p
    
−1 −1
= α Φ(f ) X1 , . . . , Φ(f ) Xm
∗ ∗
 
= α Rf(p)−1 ∗ (X1 ) , . . . , Rf(p)−1 ∗ Xm
   −1
∗ −1
= Rf(p) −1 α (X1 , . . . , Xm ) = r f (p) (αp (X1 , . . . , Xm ))
= r(f (p))(αp (X1 , . . . , Xm )) ,
yielding (15.28). 
Corollary 15.30. For ω ∈ C(P ) and f ∈ C(P, G) ,
ΩΦ(f )·ω = f Ωω f −1 .
Proof. Using Proposition 15.7 (p. 399) and (15.28) where the representation
r is ad : G → GL(g),
−1∗ −1∗ −1∗ −1∗
ΩΦ(f )·ω = ΩΦ(f ) ω
= dΦ(f ) ω + Φ(f ) ω ∧ Φ(f ) ω
−1∗ −1∗ −1∗
= Φ(f ) dω + Φ(f ) (ω ∧ ω) = Φ(f ) Ωω
= Φ(f ) · Ωω = ad(f ) Ωω = f Ωω f −1 . 
Corollary 15.31. The pairing
h·, ·i : Ωk (P ×G W ) × Ωk (P ×G W ) → C ∞ (M, R) ,
k
of (15.15) is preserved under the action of GA(P ) on Ωk (P ×G W ) ∼
= Ω (P, W ) in
the sense that for α, α0 ∈ Ωk (P ×G W ),
(15.29) hF · α, F · α0 i = hα, α0 i .
0
Proof. In the special case k = 0, with s, s0 ∈ Ω (P, W ) ∼
= Ω0 (P ×G W ), and
letting F = Φ(f ), we have
hF · s, F · s0 i = hΦ(f ) · s, Φ(f ) · s0 i
= K(r(f )(s) , r(f )(s0 )) = K(s, s0 ) = hs, s0 i ,
since r : G → GL(W ) is orthogonal relative to K. For s, s0 ∈ Ω0 (P ×G W ) and
β, β 0 ∈ Ωk (P ×G W ), for the basic forms s ⊗ β and s0 ⊗ β 0 , we have
hF ·(s ⊗ β) , F ·(s0 ⊗ β 0 )i = h(F · s) ⊗ β,(F · s0 ) ⊗ β 0 i
= hF · s, F · s0 i h(β, β 0 ) = hs, s0 i h(β, β 0 ) = hs ⊗ β, s0 ⊗ β 0 i .
For arbitrary α, α0 ∈ Ωk (P ×G W ) , (15.29) follows by linearity. 
15.4. GAUGE TRANSFORMATIONS 413

k+1
Corollary 15.32. For F ∈ GA(P ) and β ∈ Ω (P, W ), we have
F ·(δ ω β) = δ F ·ω (F · β) .
Proof. In view of Corollary 15.31, the global inner product (·, ·) of (15.17) is
k
also preserved by the action of GA(P ). Thus, for all α ∈ Ω (P, W ), we have

F · α, δ F ·ω (F · β) = DF ·ω (F · α) , F · β
 

= (F · Dω α, F · β) = (Dω α, β) = (α, δ ω β) = (F · α, F · δ ω β) ,
and it follows that F ·(δ ω β) = δ F ·ω (F · β). 
Lie Algebra Analogy and Infinitesimal Action. Let
0
C(P, g) := Ω (P, g) = s ∈ C ∞ (P, g) : s(pg) = adg−1 s(p) .


There is a map Exp : C(P, g) → C(P, G), defined for s ∈ C(P, g) by



X 1 m
(15.30) Exp(s)(p) := exp(s(p)) = s(p) .
m=0
m!
m
Note that s(p) ∈ g ⊆ gl(N, C) so that s(p) makes sense, and Exp(s) ∈ C(P, G)
since
Exp(s)(pg) = pg exp(s(pg)) = pg exp g −1 s(p) g


= pgg −1 exp(s(p)) g = p exp(s(p)) g = Exp(s)(p) g.


Thus, we also have
Φ ◦ Exp : C(P, g) → GA(P, G) .
Just as exp : g → G is a local diffeomorphism on a neighborhood of 0 ∈ g to a
neighborhood of I ∈ G, relative to the C k topology (k ≥ 0), Φ◦Exp is a continuous
bijection of a neighborhood of 0 ∈ C(P, g) to a neighborhood of Id ∈ GA(P, G).
One might think of C(P, g) as the Lie algebra of the Lie group GA(P, G). Then
k
corresponding to the group representation of GA(P, G) on Ω (P, W ), we have a Lie
k
algebra representation given, for s ∈ C(P, g) and α ∈ Ω (P, W ), by
 
d
(s · α)p (X1 , . . . , Xm ) := dt (Φ(Exp(ts)) · α)p (X1 , . . . , Xm )
t=0
d
= dt (r(Exp(ts)(p)) α(X1 , . . . , Xm )) t=0
= d
dt (r(exp(ts(p))) α(X1 , . . . , Xm )) t=0 = r0 (s(p))(α(X1 , . . . , Xm )) .
There is also an infinitesimal version of the action of C(P, g) on C(P ) defined by
d
s · ω := dt (Φ(Exp(ts)) · ω) t=0
.
Proposition 15.33. For s ∈ C(P, g) and ω ∈ C(P ),
1
s · ω = − (ds + [ω, s]) = −Dω s ∈ Ω (P, g) .
Proof. For X ∈ Tp P ,
d
s·ω = dt (Φ(Exp(ts)) · ω)(X) t=0
   
d −1 −1
= dt Exp(ts)(p) Exp(ts) (X) + Exp(ts) ω(X) Exp(ts)
∗ t=0
= −ds(X) − [ω(X) , s] .
414 15. GEOMETRIC PRELIMINARIES

We computed the derivative of the first term as follows. For γ : R →P a curve with
γ 0 (0) = X ∈ Tp P , we have (at t = u = 0)
   
d −1
dt Exp(ts)(p) Exp(ts) (X)

d d

= dt exp(ts(p)) du exp(−ts(γ(u)))
d 0
= dt (exp(ts(p))(−tds(γ (0))))
d
= dt (exp(ts(p))(−tds(X)))
d d
= dt (exp(ts(p)))(0) + I dt (−tds(X)) = −ds(X) . 

Exercise 15.34. Let π : P → M be a principal G-bundle. Suppose that Uα is


open in M and Tα : π−1 (Uα ) → Uα × G is a local trivialization of P , so that
Tα (p) = (π(p) , sα (p)) , where sα (pg) = sα (p) g.
Similarly let Tβ = π × sβ : π−1 (Uβ ) → Uβ × G be another local trivialization with
Uα ∩ Uα 6= ∅. Define a local section σα : Uα → π −1 (Uα ) by σα (x) = Tα−1 (x, e),
where e denotes the identity of G, and similarly
 define
 σβ : Uβ → π −1 (Uβ ).
(a) Show that Tα (σα (x) g) = (x, g) and Tα ◦ Tβ−1 : (Uα ∩ Uβ ) × G ←- is given by
 
Tα ◦ Tβ−1 (x, g 0 ) = (x, gαβ (x) g 0 ) ,

where gαβ : Uα ∩ Uβ → G is a well-defined function (known as a transition func-


tion for P ) given by
−1
gαβ (π(p)) := sα (p) sβ (p) .
Conclude that Tα ◦ Tβ−1 is a gauge transformation of the trivial principal G-bundle
(Uα ∩ Uβ ) × G → G.
(b) Let ω ∈ Λ1 (P, g) be a connection 1-form on P . Show that
σβ (x) = σα (x) gαβ (x) ,
−1 ∗ −1
σβ ∗ ω = gαβ (σα ω) gαβ + gαβ dgαβ , and
−1 ∗ ω
σβ ∗ Ωω = gαβ (σα Ω ) gαβ .

5. Curvature in Riemannian Geometry


The Bundle of Linear Frames. Let M be a C ∞ n-manifold. A linear frame
of M at x ∈ M is an isomorphism u : Rn → Tx M . Note that if e1 , . . . , en denotes
the standard basis of Rn , then u(e1 ) , . . . , u(en ) is a basis of Tx M . Let LMx denote
the set of all linear frames at x, let LM = ∪x∈M LMx , and let π : LM → M
denote the map u 7→ x. For A ∈ GL(n, R) and u ∈ LMx , we have u ◦ A ∈ LMx ,
with u ◦ A = u only if A = I. Thus, GL(n,R) acts freely to the right on LM .
Indeed, LM can be made into a C ∞ n + n2 -manifold, such that π : LM → M
is a principal GL(n, R)-bundle, known as the bundle of linear frames for M . The
reader is probably familiar with tensor representations of GL(n, R). For example,
∗ ∗
let S 2 (Rn ) ⊂ (Rn ) ⊗(Rn ) denote the space of symmetric, bilinear  forms on Rn . We
have the tensor representation r : GL(n, R) → GL S 2 (Rn ) given, for s ∈ S 2 (Rn )
and A ∈ GL(n, R), by
(r(A)(s))(v1 , v2 ) := s A−1 v1 , A−1 v2 for v1 , v2 ∈ Rn .

(15.31)
15.5. CURVATURE IN RIEMANNIAN GEOMETRY 415

Using the abbreviation Gn = GL(n, R) , the associated bundle LM ×Gn S 2 (Rn ) →


M is the bundle of symmetric, bilinear forms on the tangent spaces of M . For
[u, s] ∈ LM ×Gn S 2 (Rn ) and X, Y ∈ Tx M , note that [u, s](X, Y ) := s(u(X) , u(Y ))
defines a bilinear form on Tx M . In the case of the defining representation
r : GL(n, R) → GL(n, R) , given by the identity map A 7→ A,
we have the isomorphism
(15.32) LM ×Gn Rn ∼
= T M , given by [u, v] 7→ u(v) ,
for u ∈ LM and v ∈ Rn .
Definition 15.35. We define the canonical 1-form on LM to be the element
1
ϕ ∈ Ω (LM, Rn ), given for u ∈ LM and X ∈ Tu LM by
ϕu (X) := u−1 (π∗ (X)) .
1
Connections and Forms. We know (see (15.11)) that Ω (LM, Rn ) is iso-
morphic to Ω1 (LM ×Gn Rn ) = Ω1 (T M ), the space of 1-forms on M with values in
T M , or in other words, the space of endomorphisms of T M . The canonical 1-form
ϕ corresponds to the identity endomorphism. Indeed, for u ∈ LM and X ∈ Tu LM ,
we have
ϕM (π∗ X) := [u, ϕu (X)] = u, u−1 (π∗ (X)) = u u−1 (π∗ (X)) = π∗ (X) ,
  

where we have used the identification (15.32).


A connection 1-form ω for the principal GL(n, R)-bundle π : LM → M is called
a linear connection for M . The torsion Θ of ω is the covariant derivative of the
canonical 1-form with respect to ω, namely
2
(15.33) Θ := Dω ϕ ∈ Ω (LM, Rn ) ∼
= Ω2 (M, T M ) ,
which (as indicated) can be regarded as a 2-form on M with values in T M . A linear
connection gives us a way of differentiating a vector field, say Y on M , with respect
to another vector
  field0 X, as follows. LetY denote
e
 the1 ω-horizontal lift of Y to
n
LM . Then ϕ Y ∈ Ω (LM, R ), and D ϕ Y
e ω e ∈ Ω (LM, Rn ) ∼ = Ω1 (M, T M ).
  
Thus, we may regard Dω ϕ Ye as a 1-form with values in T M . Evaluating this
1-form on X gives us a vector field, commonly denoted by ∇X Y . Explicitly, for
any u ∈ LM , we have
           
(∇X Y )π(u) = u Dω ϕ Ye X
e = u d ϕ Ye Xe
u u
  
(15.34) =u X eu [ϕ Ye ] ∈ Tπ(u) M.

Exercise 15.36. (a) Show that for vector fields X, Y ∈ C ∞ (T M ) and f ∈



C (M ) ,
(K1 ) ∇f X Y = f ∇X Y and
(K2 ) ∇X (f Y ) = df (X) Y + f ∇X Y.
(b) A Kozul connection ∇ for M is defined to be a map C ∞ (T M ) × C ∞ (T M ) →
C ∞ (T M ) written as (X, Y ) 7→ ∇X Y such that (K1 ) and (K2 ) hold. Show that a
Kozul connection arises from a unique linear connection for M .
416 15. GEOMETRIC PRELIMINARIES

(c) Let M be a submanifold of Rn+k . For X, Y ∈ C ∞ (T M ) , say X = (X1 , . . . , Xn+k )


and Y = (Y1 , . . . , Yn+k ) , where Xi , Yi ∈ C ∞ (M ), define
X [Y ] := (X [Y1 ] , . . . , X [Yn+k ]) ,
d
where X [Yi ]x := dt Yi (γ(t)) t=0 and γ is a curve in M with γ 0 (0) = Xx . Verify that
the following defines a Kozul connection for M :
(∇X Y )x := Px (X [Y ]) for any x ∈ M,
where Px is orthogonal projection of Rn+k onto Tx M .
The torsion Θ of ω is 0, if and only if ∇X Y − ∇Y X − [X, Y ] = 0 for all vector
fields X and Y . Indeed, for u ∈ LM , we have
     
u Θ X,e Ye = u (Dω ϕ) X, e Ye
         h i
=u X e ϕ Ye − u Ye ϕ X e − u ϕ X,e Ye
h i
(15.35) = ∇X Y − ∇Y X − π∗ X, e Ye = ∇X Y − ∇Y X − [X, Y ] .

We can similarly express the curvature Ωω in terms of ∇. First note that


    h  i h  i h i
Ωω X,e Ye = dω X, e Ye = X e ω Ye − Ye ω X e − ω X, e Ye
h i h iV  ∗
= −ω X, e Ye =⇒ X, e Ye = −Ωω u X,
e Ye .
u

Then recall that for A ∈ gl(n, Rn ) and A∗ the fundamental vertical vector field on
LM , we have
h  i h  i      
A∗ ϕ Z e = d ϕu exp(tA) Z e = d exp(tA)−1 ϕu Ze = −Aϕu Ze .
dt dt

Thus,
∇X ∇Y Z − ∇Y ∇X Z − ∇[X,Y ] Z
 h h  ii  h h  ii h iH h  i
=u X eu Ye ϕ Ze − u Yeu X e ϕ Ze − u X,e Ye ϕ Z
e
h i h iH  h  i h iV h  i
=u e Ye − X,
X, e Ye ϕ Ze = u X,e Ye ϕ Z
e

(15.36)
  ∗ h  i     
= u −Ωω
u X, Y
e e ϕ Ze = u Ωω X e =: Ωω (X, Y )(Z) ,
eu , Yeu ϕu Z

where the final equality defines what it means to regard the curvature form
2
Ωω ∈ Ω (LM, gl(n, Rn ))
as being in Ω2 (End T M ).

The Orthonormal Frame Bundle. Now, suppose that h is a Riemannian


metric on M . Then u ∈ LMx is called an orthonormal frame if u : Rn → Tx M is
an isometry where Rn has its standard inner product (i.e., h(u(v) , u(w)) = v · w =
v 1 w1 + · · · + v n wn ). The set F M of all orthonormal frames at all points of M is the
total space of a principal O(n)-bundle π : F M → M called the orthonormal frame
bundle of M relative to h. Note that F M is a submanifold of LM .
15.5. CURVATURE IN RIEMANNIAN GEOMETRY 417

Definition 15.37. If ω is a linear connection for M whose horizontal subspaces


(those subspaces annihilated by ω) at points of F M are contained in tangent spaces
of F M , then ω is called a metric connection (relative to h).
Note that ω|F M is automatically a connection 1-form for the principal O(n)-
bundle F M → M , but the horizontal subspace of ω at u ∈ F M is not necessarily
Ker(ω|Tu F M ); i.e., ω is not necessarily metric. There is another useful characteriza-
tion of metric connections described as follows. The metric h on M corresponds to
0  
some H ∈ Ω LM, S 2 (Rn ) , where r : GL(n, R) → GL S 2 (Rn ) denotes the tensor
representation for the space S 2 (Rn ) of symmetric bilinear forms, defined in (15.31).
Indeed for u ∈ LM and v1 , v2 ∈ Rn ,
(15.37) H(u)(v1 , v2 ) := h(u(v1 ) , u(v2 )) .
For u ∈ F M , H(u)(v1 , v2 ) = v1 · v2 , so that H|F M is constant, namely the usual
dot product. In fact, if ι denotes the usual dot product, F M = H −1 (ι). Let ω be
a connection on 1-form on LM . According to (15.8),
Dω H = dH + r0 (ω)(H) .
Proposition 15.38. The following are equivalent
1. ω is a metric connection,
2. Dω H = 0 on LM,
3. X [h(Y, Z)] = h(∇X Y, Z) + h(Y, ∇X Z) , for all vector fields X, Y and Z.
Proof. Since H|F M is constant, for X ∈ Tu F M , we have dH(X) = 0 and
((Dω H)(X))(v1 , v2 ) = dH(X)(v1 , v2 ) + (r0 (ω(X))(H))(v1 , v2 )
= − (ω(X) v1 ) · v2 − v1 ·(ω(X) v2 ) .
Since the left side is 0 for X vertical and the right side is 0 for X horizontal, both
sides are zero for all X ∈ Tu F M , namely
(15.38) (Dω H)(X) = 0 and (ω(X) v1 ) · v2 + v1 ·(ω(X) v2 ) = 0.
Suppose that ω is metric. Then for any u ∈ F M , all horizontal Y ∈ Tu LM are
in Tu F M , and hence (Dω H)(Y ) = 0 by (15.38). Since GL(n, R) acts transitively
via Rg∗ on the set of all horizontal subspaces of LM at points of a given fiber, we
then have that (Dω H)(Y ) = 0 for all horizontal Y ∈ Tu LM for all u ∈ LM . Since
(Dω H)(Y ) = 0 for all vertical Y , we then have ω metric ⇒ Dω H = 0 on LM .
Now, suppose that Dω H = 0 on LM , and ω(Y ) = 0 for Y ∈ Tu LM . To show that
ω is metric, we need to prove that Y ∈ Tu F M . However,
0 = (Dω H)(Y ) = dH(Y ) + (r0 (ω(Y ))(H)) = dH(Y ) .
Thus, as dH has maximal rank at u ∈ F M (Exercise), Y ∈ Tu H −1 (ι) = Tu F M.


For (3) note that for x = π(u) ,


h     i     
Xx [h(Y, Z)] = Xeu H ϕ Ye , ϕ Ze =Xeu [H] ϕ Ye , ϕ Ze
 h  i      h  i
+ Hu X eu ϕ Ye , ϕ Ze + Hu ϕ Ye , X eu ϕ Ze
    
= (Dω H)u ϕ Ye , ϕ Ze + hx (∇X Y, Z) + hx (Y, ∇X Z) . 
418 15. GEOMETRIC PRELIMINARIES

The Fundamental Lemma of Riemannian Geometry. After many pages


with notations, definitions and re-formulations, we can present a true achievement:

Proposition 15.39 (Fundamental Lemma of Riemannian Geometry). For each


Riemannian metric h on a manifold M , there is a unique metric linear connection
1-form ω on LM relative to h with torsion 0 (i.e., Θ = Dω ϕ = 0).
Definition 15.40. The connection ω in this proposition restricts to an so(n)-
valued connection θ := ω|F M which we call the Levi-Civita connection for h,
generalizing our previous definition in Section 6.5, Equation (6.20) (p.180). The
existence and uniqueness proof below is also valid for nondegenerate, indefinite
metrics h. One just replaces the standard dot product on Rn by a standard, non-
degenerate, indefinite scalar product.
Proof. By (15.8), for the canonical 1-form ϕ, we have
Dω ϕ = dϕ + ω ∧ ϕ,
or more precisely, for all u ∈ LM and W1 , W2 ∈ Tu LM ,
Dω ϕ(W1 , W2 ) = dϕ(W1 , W2 ) + ω(W1 ) ϕ(W2 ) − ω(W2 ) ϕ(W1 )
where ω(W1 ) ∈ gl(n, R) is operating on ϕ(W2 ) ∈ Rn . A vector field Z on M
0
corresponds to some Z 0 ∈ Ω (LM, Rn ) via Z 0 (u) = u−1 (Z). Let Z e denote the
unique vector field on LM such that π∗ Z = Z and Z denotes the horizontal
e e
 lift of
Z relative to some arbitrary, fixed connection ω0 on LM (i.e., ω0 Z e = 0). From
the fact that horizontal subspaces are sent to horizontalsubspaces
 by Rg∗ , it follows
0
that Z is invariant under Rg∗ . Moreover, we have ϕ Z = Z . Indeed, at each
e e
u ∈ LM,
Z 0 (u) = u−1 Zπ(u) = u−1 (π∗ (Z

eu )) = ϕu (Z
eu ).
Let X and Y be vector fields on M . Then for an arbitrary connection 1-form ω on
LM
Dω ϕ(X,
e Ye ) = dϕ(X,
e Ye ) + ω(X)ϕ(
e Ye ) − ω(Ye )ϕ(X)
e

= dϕ(X, e 0 − ω(Ye )Y 0 ,
e Ye ) + ω(X)Y
0
or for all vector fields Z on M with corresponding Z 0 ∈ Ω (LM, Rn ),
e Ye ), Z 0 )
H(Dω ϕ(X,
(15.39) e Ye ), Z 0 ) + H(ω(X)Y
= H(dϕ(X, e 0 , Z 0 ) − H(ω(Ye )X 0 , Z 0 ).
We also have
0
(15.40) (Dω H) (Z)(X
e , Y 0 ) = dH(Z)(X
e 0
, Y 0 ) + r0 (ω(Z))(H)(X
e 0
, Y 0)
0
= dH(Z)(X
e , Y 0 ) − H(ω(Z)X
e 0 , Y 0 ) − H(X 0 , ω(Z)Y
e 0 ).

From (15.39) and (15.40), we see that if Dω ϕ = 0 and Dω H = 0, then


e Ye ), Z 0 ) = H(ω(Ye )X 0 , Z 0 ) − H(ω(X)Y
H(dϕ(X, e 0, Z 0)
and
0
dH(Z)(X
e , Y 0 ) = H(ω(Z)X
e 0 , Y 0 ) + H(X 0 , ω(Z)Y
e 0 ).
15.5. CURVATURE IN RIEMANNIAN GEOMETRY 419

We want to solve for H(ω(X)Y e 0 , Z 0 ). This is accomplished by forming the following


mysterious combination which one can derive from Young diagrams for 3-tensors
(see [436]), or (with a little luck) by trial and error.
0 0 0 0 0 0
 , Z ) + dH(Y )(X , Z ) − dH(Z)(X , Y )
dH(X)(Y
e e e

0 0 e Ye ), X 0 )
− H(dϕ(X, e Ye ), Z ) + H(dϕ(Z, e X),
e Y ) + H(dϕ(Z,
   
H(ω(X)Y e 0 , Z 0 ) + H(Y 0 , ω(X)Z e 0)
+H(ω(Ye )X 0 , Z 0 ) + H(X 0 , ω(Ye )Z 0 ) 
   
  

 −H(ω( e 0 , Y 0 ) − H(X 0 , ω(Z)Y
Z)X e 0) 

=  
0 0 0 0
 
 H(ω(Y )X , Z ) − H(ω(X)Y , Z )
e e 

− 0 0 0 0
+H(ω( X)Z , Y ) − H(ω( Z)X , Y )
  e e  
   
+H(ω(Ye )Z 0 , X 0 ) − H(ω(Z)Y e 0, X 0)
0 0
= 2H(ω(X)Y
e , Z ).
Thus, we have

(15.41) e 0 , Z 0 ) = dH(X)(Y
2H(ω(X)Y e 0
, Z 0 ) + dH(Ye )(X 0 , Z 0 ) − dH(Z)(X
e 0
, Y 0)
 
− H(dϕ(X, e Ye ), Z 0 ) + H(dϕ(Z, e Y 0 ) + H(dϕ(Z,
e X), e Ye ), X 0 ) .

Hence, if Dω ϕ = 0 and Dω H = 0, then ω is uniquely determined by (15.41).


Conversely, suppose that we define ω by (15.41) and by the requirement that
ω(A∗ ) = A. Then using (15.41), it is straightforward to check that not only do
we have that ω is torsion-free, namely
e Ye ), Z 0 ) = H(dϕ(X,
H(Dω ϕ(X, e Ye ), Z 0 ) + H(ω(X)Y
e 0 , Z 0 ) − H(ω(Ye )X 0 , Z 0 ) = 0,
but also that ω is a metric connection, namely
   
0
(Dω H) (Z)(X
e , Y 0 ) = dH(Z)(X
e 0
, Y 0 ) − H ω(Z)X
e 0 , Y 0 − H X 0 , ω(Z)Y
e 0 = 0.

Of course one should also check that ω defined by (15.41) and ω(A∗ ) = A, also
has the property Rg∗ ω = g −1 ωg. This is automatic on vertical vectors, but we need
Rg∗ ω (X
eu ) = g −1 ω(X

eu )g

for all vector fields X on M and g ∈ GL(n, R), or equivalently,


(15.42)    
H(ug) Rg∗ ω (X eu )Y 0 (ug) , Z 0 (ug) = H(ug) g −1 ω(X
eu )gY 0 (ug) , Z 0 (ug)


for all vector fields X, Y, Z on M and g ∈ GL(n, R). For this we use
   
H(ug) Rg∗ ω (X eu )Y 0 (ug) , Z 0 (ug) = H(ug) ω(Rg∗ Xeu )Y 0 (ug) , Z 0 (ug)

 
(15.43) eug )Y 0 (ug) , Z 0 (ug)
= H(ug) ω(X
and
 
H(ug) g −1 ω(X
eu )gY 0 (ug) , Z 0 (ug)
 
= r g −1 · H(u) g −1 ω(X eu )gY 0 (ug) , Z 0 (ug)

   
(15.44) = H(u) ω(X eu )gY 0 (ug) , gZ 0 (ug) = H (u) ω(X eu )Y 0 (u) , Z 0 (u) ,
420 15. GEOMETRIC PRELIMINARIES

and then apply (15.41) to the final expressions in (15.43) and (15.44). The result
(15.42) then follows by showing that the R-valued functions of the form
 
0
dH(X)(Y
e , Z 0 ) and H dϕ(X, e Ye ), Z 0

are Rg invariant. As can be easily checked, this follows from the fact that H, ϕ
and X 0 are equivariant, and Rg∗ (X)
e =X e (and similarly for Y and Z). Thus, the
1-form ω defined by (15.41) is a connection, and is the unique torsionless metric
connection on LM relative to h. 
Levi-Civita Connection
 in Local Coordinates and Christoffel Sym-
bols. Let x1 , . . . , xn be a system of local coordinates in a neighborhood U of M .
Then the coordinate vector fields ∂k := ∂x∂ k yield a section σ : U → LM of the
frame bundle given, for v ∈ Rn and x ∈ U , by σ(x)(v) =
P k
v (∂k )x . The image
σ∗ (Tx M ) is a subspace of Tσ(x) LM which is a complement of the vertical subspace
0
of Tσ(x) LM . As A ∈ GL(n, R) varies, the subspaces Hσ(x)A := RA∗ Tσ(x) LM then
define a connection on LM |π−1(U ) . This may locally serve to define the fixed con-
Pn k
nection ω0 in the above proof. For a vector field V (x) = k=1 v(x) ∂k , note that
Veσ(x)A := RA∗ (σ∗x (Vx )) is the ω0 -horizontal lift of V , and
 
0 −1 −1
Vσ(x)A : = ϕ Ve = (σ(x) A) (π∗ (RA∗ (σ∗x (Vx )))) = (σ(x) A) (Vx )
−1
= A−1 σ(x) (Vx ) = A−1 v(x) .
We have H(σ(x) A)(v, w) = h(σ(x) A(v) , σ(x) A(v)), and in particular
hij (x) := hx (∂i , ∂j ) = h(σ(x)(ei ) , σ(x)(ej )) = H(σ(x))(ei , ej ) or
(σ ∗ H)(ei , ej ) = σ ∗ (H(ei , ej )) = hij .
fkσ(x) = σ∗x (∂k ) and (∂ 0 ) −1
We have ∂ k σ(x) = ϕ(σ∗x (∂k )) = σ(x) (∂k ) = ek . Thus, the
result (see (15.41))
2H(ω(eei )e0j , e0k ) = dH(eei ) e0j , e0k + dH(eej )(e0i , e0k ) − dH(eek ) e0i , e0j
 
 
(15.45) − H(dϕ(eei , eej ), e0k ) + H(dϕ(eek , eei ), Y 0 ) + H(dϕ(eek , eej ), e0i )

becomes (on the image σ(U ))



(15.46) 2H ω(σ∗x (∂i ))ej , ek = dH(∂ei )(ej , ek ) + dH(∂ej )(ei , ek ) − dH(∂ fk )(ei , ej )
 
− Hσ(x) (dϕ(∂ei , ∂ej ), ek ) + H(dϕ(∂fk , ∂ei ), Y 0 ) + H(dϕ(∂
fk , ∂ej ), ei ) .

On the image σ(U ) again,


 h i h i h i 
H(dϕ(∂ei , ∂ej ), ek ) = H ∂ei ϕ(∂ej ) − ∂ej ϕ(∂ei ) − ϕ ∂ei , ∂ej , ek

= H ∂ei [ej ] − ∂ej [ei ] − ϕ(0) , ek = 0,
and
 
(dH)σ(x) (∂ei )(ej , ek ) = d H|σ(U ) (ej , ek ) σ∗x (∂i )
= σ ∗ d H|σ(U ) (∂i )(ej , ek )


= d σ ∗ H|σ(U ) (ej , ek ) (∂i ) = ∂i (hjk ) .


 
15.5. CURVATURE IN RIEMANNIAN GEOMETRY 421

k
Thus, letting (ωi )j be defined by
X k
(σ ∗ ω) (∂i )(ej ) = ω(σ∗ (∂i ))(ej ) = ω(eei )(ej ) = (ωi )j ek ,
k

we have
X l  X l 
2Hσ(x) (ω(∂ei )ej , ek ) = 2Hσ(x) (ωi )j el , ek = 2Hσ(x) (ωi )j el , ek
k k
X l
=2 hkl (ωi )j .
l
k
Hence, using the classical notation for the Christoffel symbols Γkij := (ωi )j , we get
X X l
2 hkl Γlij = 2 hkl (ωi )j = ∂i [hjk ] + ∂j [hik ] − ∂k [hij ] or
l l
l
(15.47) Γlij = (ωi )j = 1 lk
2 h (∂i [hjk ] + ∂j [hik ] − ∂k [hij ]) ,

which is the classical formula. Note that Γlij is symmetric in i and j, and hence the
Levi-Civita connection is often called a symmetric connection. We now show that
n
X
(15.48) ∇∂i ∂j = Γlij ∂l .
l=1
    
Since ϕσ(x) ∂ei = ei is constant, we have d ϕ ∂ei = 0, and so
              
Hσ(x) Dω ϕ ∂ei ∂ej , ek = Hσ(x) d ϕ ∂ei + ω ∂ej ϕ ∂ei , ek
    X X
l
= Hσ(x) ω ∂ej ei , ek = hkl (ωj )i = hkl Γlji .
l l

Then for each k,


!
X X X
hx Γlji ∂l , ∂k = Γlji hx (∂l , ∂k ) = hkl Γlji
l l l

    
ω
= Hσ(x) D ϕ ∂ei ∂ej , ek
      
= hx σ(x) Dω ϕ ∂ei ∂ej , σ(x)(ek ) = hx (∇∂i ∂j , ∂k ) ,

from which (15.48) follows. Let h denote the matrix whose entries are hij . There
is a key identity that we will use later, namely (using the summation convention)
n  √
1 X 
hij ∇∂i ∂j = − ∂i hli det h ∂l or

l=1
det h
1 h √ i
(15.49) hij Γlij = − √ ∂i hli det h .
det h
This is based on the identity
n
X
∂k [det h] = (det h) hij ∂k [hij ] or hij ∂k [hij ] = ∂k [log det h] ,
i,j=1
422 15. GEOMETRIC PRELIMINARIES

which is shown as follows. Let hj denote the j-th column of h and write the
determinant as a multilinear function of its columns, say det h = det(h1 , . . . , hn ).
Then (where we use Cramer’s rule for the third equality)
n
X
∂k [det h] = det(h1 , . . . , hj−1 , ∂k [hj ] , hj+1 , . . . , hn )
j=1
n
X 1
= (det h) det(h1 , . . . , hj−1 , ∂k (hj ) , hj+1 , . . . , hn )
j=1
det h
n
X j
= (det h) h−1 ∂k (hj )
j=1
n
X n
X
= (det h) hji ∂k (hij ) = (det h) hij ∂k (hij ) .
i,j=1 i,j=1

To obtain (15.49), we then compute (where a sum over i and j is implicit)

hij Γlij = hij Γlij = 21 hij hlk (∂i [hjk ] + ∂j [hik ] − ∂k [hij ])
= hij hlk ∂i [hjk ] − 21 hlk ∂k [hij ]


= hij ∂i hlk hjk − ∂i hlk hjk − 12 hlk ∂k [hij ]


    

= −∂i hlk hij hjk − 12 hlk hij ∂k [hij ] = −∂i hli − 21 hlk hij ∂k [hij ]
   
 √ 
= −∂i hli − 12 hlk ∂k (log det h) = −∂i hli − hlk ∂k log det h
   

−1 √ h√ i
det h∂i hli + hlk ∂k
 
=√ det h
det h
−1 h √ i
=√ ∂i hli det h .
det h

The Curvature of the Levi-Civita Connection. As in (15.4) and (15.6),


the curvature of the Levi-Civita connection θ = ω|F M is
2
Ωθ := Dθ θ = dθ + θ ∧ θ = dθ + 21 [θ, θ] ∈ Ω (F M, so(n)) ,

where so(n) = o(n) := A ∈ gl(n, R) : AT = −A denotes the Lie algebra of O(n)
(or SO(n)). Let us denote the bundle of skew-symmetric (relative to h) endomor-
phisms of the tangent spaces Tx M by so(T M ). Then so(T M ) = F M ×O(n) so(n)
and
2
Ω (F M, so(n)) ∼
= Ω2 M, F M ×O(n) so(n) = Ω2 (M, so(T M ))


Thus, we may regard Ωθ as belonging to Ω2 (M, so(T M )), and for vectors X, Y ∈
∗
Tx M , Ωθ (X, Y ) ∈ so(Tx M ), we define R ∈ C ∞ M, ⊗4 (T M ) by

R(W, Z, X, Y ) := h Ωθ (X, Y )(Z) , W .



(15.50)
Note that R(W, Z, X, Y ) is antisymmetric in (X, Y ) and in (W, Z). We can relate
the curvature tensor to the Gaussian curvature of surfaces as follows. Let X, Y ∈
Tx M and let S denote the surface composed of geodesic segments issuing from x
with initial tangent vectors in the 2-plane Π = span(X, Y ) ⊆ Tx M . Then the
15.5. CURVATURE IN RIEMANNIAN GEOMETRY 423

Gaussian curvature of S (with the induced metric) at x is given by


R(X, Y, X, Y )
(15.51) K(Π) := 2,
h(X, X) h(Y, Y ) − h(X, Y )
which is independent of the choice of the basis {X, Y } and is known as the sectional
curvature of Π.
Exercise 15.41. Here you show that the curvature tensor of the unit n-sphere
S n (with metric tensor induced from Rn+1 ) at the point x ∈ S n is given by
(15.52) R(W, Z, X, Y ) = hX, Zi hY, W i − hY, Zi hX, W i ,
where W, Z, X, Y ∈ Tx S n = x⊥ ⊂ Rn+1 , and h·, ·i denotes the usual dot product.
This implies that all the sectional curvatures K(Π) of S n are equal to 1.
(a) Let en+1 = (0, . . . , 0, 1) ∈ S n . Show that the subgroup
{g ∈ O(n + 1) : gen+1 = en+1 }
can be identified with O(n) and that the map π : O(n + 1) → S n , given by π(g) =
gen+1 induces a diffeomorphism O(n+1) ∼ n n
O(n) = S . Hence, π : O(n + 1) → S is the
special case of the bundle π : G → G/G in Exercise 15.6, where G = O(n + 1) and
G = O(n).
(b) Show that for x ∈ S n and g ∈ π −1 (x) , we have g|Rn : Rn ∼ = x⊥ = Tx S n is an
isometry. This shows that π : O(n + 1) → S n may be regarded as the orthonormal
frame bundle π : F (S n ) → S n .
n
(c) Noting that any vector in Tg (F (S ))T = Tg (O(n + 1)) is of the form gA for some
A ∈ o(n + 1) = A ∈ gl(n + 1, R) : A = −A , show that the canonical 1-form
1
ϕ ∈ Ω (F (S n ) , Rn ) is given by
ϕ(gA) = Aen+1
Note that A ∈ o(n + 1) ⇒ Aen+1 · en+1 = 0 ⇒ Aen+1 ∈ Rn .
(d) Show that
o(n + 1) = o(n) ⊕ m, where
m := {A ∈ o(n + 1) : Aen+1 · en+1 = 0} ∼
= Rn ,
and where the isomorphism m ∼
= Rn is defined by A 7→ Aen+1 , with inverse α : Rn →
m , given by
α(x)(y) := hen+1 , yi x − hx, yi en+1 for x, y ∈ Rn .
Also verify that adg m = gmg −1 = m, and for all x, y ∈ Rn
[α(x) , α(y)](en+1 ) = 0 so that [α(x) , α(y)] ∈ o(n) , while
[α(x) , α(y)](z) = hx, zi y − hy, zi x for x, y, z ∈ Rn .

(e) According to Exercise 15.6 (p. 398) with G = O(n + 1) and G = O(n), the
o(n)-valued 1-form
ω := πo(n) ◦ ω ∈ Ω1 (O(n + 1) , o(n)) (where ω(gA) := A)
is a connection 1-form for π : O(n + 1) → O(n + 1) / O(n) or π : F (S n ) → S n .
Show that ω is the Levi-Civita connection for S n (i.e., Dω ϕ = 0). For this it is best
to evaluate Dω ϕ on the pair horizontal vector fields g 7→ gα(v) and g 7→ gα(w) for
some v, w ∈ Rn .
424 15. GEOMETRIC PRELIMINARIES

(f) For X ∈ Tx S n = x⊥ ⊆ Rn+1 , show that the horizontal  lift Xg of X at g ∈


e
O(n + 1) = F S n is gα g −1 X . Then, using the formula Ωω A,

e B e = πg ([A, B]) of
Exercise 15.8 (p. 399) and Part (d) above, show that for X, Y, Z ∈ Tx S n ,
   
R(X, Y )(Z) := gΩω g
e Ye ϕg Ze = hX, Zi Y − hY, Zi X,
X,

so that (15.52) holds.


Remark 15.42. Note that under a dilation of the metric, say h 7→ ch for some
c > 0, the Levi-Civita connection ∇ does not change, as is apparent from (15.48)
and (15.47). Hence by virtue of (15.36), the curvature Ωω ∈ Ω2 (End T M ) does
not change However, the curvature tensor R(W, Z, X, Y ) := h Ωθ (X, Y )(Z) , W
changes by a factor of c, because of the involvement of the metric. Moreover, due
to the factor of c2 in the denominator of (15.51), the sectional curvatures then
change by a factor of c−1 . Thus, a sphere of radius r (whose metric tensor is c = r2
times that of the unit sphere) sectional curvatures all equal to r−2 , which is the
Gaussian curvature of all its great 2-spheres.
First Bianchi Identity. Proposition 15.44 below implies (by linearity) that
R is uniquely determined at a point x by its sectional curvatures at x, but first
we need to derive the so-called First Bianchi Identity. Using the fact that θ is
torsion-free and (15.13), we obtain
0 = Dθ Θ = Dθ Dθ ϕ = Ωθ ∧ ϕ,


Regarding Ωθ ∈ Ω2 (M, so(T M )) and ϕ = Id ∈ Ω1 (M, End(T M )), we then have


0 = Ωθ ∧ ϕ (X, Y, Z) = 13 Ωθ (X, Y )(Z) + Ωθ (Z, X)(Y ) + Ωθ (Y, Z)(X)
 

Taking the h inner product with W yields the First Bianchi Identity, namely
(15.53) R(W, Z, X, Y ) + R(W, Y, Z, X) + R(W, X, Y, Z) = 0.
Exercise 15.43. From the antisymmetry of R(W, Z, X, Y ) in (W, Z) and in
(X, Y ) and the First Bianchi Identity identity (15.53), obtain the identity
(15.54) R(W, Z, X, Y ) = R(X, Y, W, Z) .
[Hint. Add the equations (each a First Bianchi Identity)
R(W, Z, X, Y ) + R(W, Y, Z, X) + R(W, X, Y, Z) = 0,
R(Y, W, Z, X) + R(Y, X, W, Z) + R(Y, Z, X, W ) = 0,
−R(X, Y, W, Z) − R(X, Z, Y, W ) − R(X, W, Z, Y ) = 0, and
−R(Z, X, Y, W ) − R(Z, W, X, Y ) − R(Z, Y, W, X) = 0.
and use the above antisymmetry.]
Proposition 15.44. If R(X, Y, X, Y ) = 0 for all X, Y ∈ Tx M , then R = 0
at x.
Proof. We know that R(W, X, Y, Z) is antisymmetric in (W, X) and in (Y, Z).
Hence, if we knew that R(W, X, Y, Z) is antisymmetric in any other pair, say (W, Y ),
then it would be antisymmetric in all pairs; e.g., for (X, Y ),
R(W, X, Y, Z) = −R(X, W, Y, Z) = R(Y, W, X, Z) = −R(W, Y, X, Z) .
15.5. CURVATURE IN RIEMANNIAN GEOMETRY 425

The First Bianchi Identity would then yield the desired result

0 = R(W, X, Y, Z) + R(W, Y, Z, X) + R(W, Z, X, Y )


= R(W, X, Y, Z) − R(W, X, Z, Y ) − R(W, X, Z, Y )
= R(W, X, Y, Z) + R(W, X, Y, Z) + R(W, X, Y, Z)
= 3R(W, X, Y, Z) .

Thus, it remains to prove R(Y, X, W, Z) = −R(W, X, Y, Z). By assumption,

0 = R(W, X + Z, W, X + Z)
= R(W, X, W, X) + R(W, X, W, Z) + R(W, Z, W, X) + R(W, Z, W, Z)
= R(W, X, W, Z) + R(W, Z, W, X)
= 2R(W, X, W, Z) by (15.54).

Then, as required,

0 = R(Y + W, X, Y + W, Z)
= R(Y, X, Y, Z) + R(W, X, Y, Z) + R(Y, X, W, Z) + R(W, X, W, Z)
= R(W, X, Y, Z) + R(Y, X, W, Z) . 

The Second Bianchi Identity. The general Bianchi Identity of (15.12),


yields Dθ Ωθ = 0 which in the context of Levi-Civita connections is called the
Second Bianchi Identity. In order to write this identity in terms of R, it is con-
venient to introduce the notion of the standard horizontal vector field w on F M
associated to w ∈ Rn

Definition 15.45. Given w ∈ Rn , the standard horizontal vector field w


on F M associated to w assigns to each u ∈ F M the unique vector wu ∈ Tu F M ,
such that θ(wu ) = 0 and ϕ(wu ) = w.

A standard horizontal vector field w is not Rg∗ invariant, since


 
(15.55) ϕ(Rg∗ wu ) = g −1 ϕ(wu ) = g −1 w ⇒ Rg∗ wu = g −1 w ,
ug

and so w is not a horizontal lift. One can also define w on LM relative to a given
linear connection ω on LM .
For x, y, z ∈ Rn , we let x, y, z denote the associated standard horizontal vector
fields. Now Dθ ϕ = 0 implies (using the fact that ϕ(y) = y is constant) that

Dθ ϕ (x, y) = x [ϕ(y)] − y [ϕ(x)] − ϕ([x, y]) = ϕ([x, y]) ,



(15.56) 0 =

In other words, the Lie bracket [x, y] is a vertical vector field. Indeed, from

Ωθ (x, y) = Dθ θ (x, y) = dθ(x, y)




(15.57) = x [θ(y)] − y [θ(x)] − θ([x, y]) = −θ([x, y]) ,


426 15. GEOMETRIC PRELIMINARIES

we see that the vertical part of [x, y] is −Ωθ (x, y). Using (15.56), we have
0 = 12 Dθ Ωθ (x, y, z) = 12 dΩθ (x, y, z)


= x Ωθ (y, z) + y Ωθ (z, x) + z Ωθ (x, y)


     

+ Ωθ ([x, y] , z) + Ωθ ([z, x] , y) + Ωθ ([y, z] , x)


= x Ω (y, z) + y Ωθ (z, x) + z Ωθ (x, y)
 θ     

= d Ωθ (y, z) (x) + d Ωθ (z, x) (y) + d Ωθ (x, y) (z) .


  
(15.58)
∗
We can identify R ∈ C ∞ M, ⊗4 (T M ) with R ∈ Ω0 F M, ⊗4 Rn∗ defined, for


u ∈ F M , by
R(u)(w, z, x, y) = Ωθu (x, y)(z) · w.


Since
Dθ R (v)(w, z, x, y) = ((dR)(v))(w, z, x, y) = d Ωθ (x, y) (v)(z) · w,
  

we see from (15.58) that


Dθ R (x)(w, v, y, z) + Dθ R (y)(w, v, z, x) + Dθ R (z)(w, v, x, y) = 0.
  
(15.59)
In terms of R this identity (the Second Bianchi Identity) is typically written as
(15.60) (∇X R)(W, V, Y, Z) + (∇Y R)(W, V, Z, X) + (∇Z R)(W, V, X, Y ) = 0.
Exercise 15.46. A C ∞ Riemannian manifold M with metric h has constant
sectional curvature at a point x if all of the 2-planes Π in Tx M have the same
sectional curvature (automatic if dim M = 2.) Use (15.60) to show that if dim M ≥
3, M is connected, and M has constant sectional curvature K(x) at each x ∈ M ,
then in fact K(x) is independent of x. This is known as Schur’s Theorem.
[Hint. First use Proposition 15.44 to deduce that at x ∈ M , you have
R(W, Z, X, Y ) = K(x)(h(X, Z) h(Y, W ) − h(Y, Z) h(X, W )) .
Then apply (15.60),and choose W = Z and V = Y with X, Y, Z orthogonal.]
Ricci and Scalar Curvature as Contractions, and Einstein’s Equation.
Recall from tensor algebra, that we speak of contraction when we set two indices
equal to each other and sum over that index.
Definition 15.47. (a) The Ricci curvature of h is the contraction (trace) of
R in the first and third slots, namely for an orthonormal basis E1 , . . . , En of Tx M
and X, Y ∈ Tx M,
n
X n
X
(15.61) Ric(X, Y ) := R(Ei , X, Ei , Y ) = R(Ei , Y, Ei , X) = Ric(Y, X) ,
i=1 i=1

where we have used (15.54) to deduce that Ric is a symmetric 2-tensor.


(b) The scalar curvature S is the contraction of Ric, namely
n
X
S= Ric(Ei , Ei ) .
i=1

Exercise 15.48. The Einstein tensor of h is defined to be Ric − 21 Sh. By


contracting (15.60) in the pairs (W, Y ) and (V, Z). Show that
0 = div Ric − 12 Sh := ∇Ei Ric − 12 Sh (Ei , ·) ∈ Ω1 (M ) .
 
15.5. CURVATURE IN RIEMANNIAN GEOMETRY 427

When h has signature (3, 1) (e.g, when (M, h) is a space-time), the Einstein equa-
tion of general relativity is Ric − 12 h = −8πK
c2 T (see (14.6)), where T denotes the
symmetric stress-energy-momentum tensor which is known to be divergence-free
by conservation of energy and momentum. The left side Ric − 12 h of Einstein’s
equation is the most obvious geometric candidate for a divergence-free symmetric
tensor.
All Possible Curvature Tensors on Rn and the Kulkarni-Nomizu Prod-
uct. We will need to study the abstract space R(Rn ) of all possible curvature ten-
sors on Rn , defined as follows. Let e1 , . . . , en denote the standard basis of Rn . For

R ∈ ⊗4 (Rn ) , let Rhijk := R(eh , ei , ej , ek ). We define R(Rn ) to be the set of all

R ∈ ⊗4 (Rn ) , such that
(A) Rhijk = −Rihjk = Rihkj , and
(15.62)
(B) Rhijk + Rhkij + Rhjki = 0.

Thus, R(Rn ) consists of those R ∈ ⊗4 (Rn ) antisymmetric in the first pair and
the last pair of indices, and which satisfy the First Bianchi Identity. We have seen
that Rhijk = Rjkhi then follows.
 Hence, R(Rn ) can be regarded as the subspace of
the vector space S Λ (R ) of symmetric linear endomorphisms of Λ2 (Rn ) which
2 n

satisfy Condition (B) in (15.62). Since Condition (B) is automatic in S Λ2 (Rn )


except when h, i, j,and k are distinct,
dim(R(Rn )) = dim S Λ2 (Rn ) − n4
 
 
= 21 n(n−1) n(n−1)
+ 1 − n4 = 12
1 2
n n2 − 1 .
 
2 2

Let S(Rn ) denote the space of symmetric bilinear forms on Rn . There is a linear
map, which we call the Ricci map,
Xn
r : R(Rn ) −→ S(Rn ) , given by r(R) := Rhihk ,
h=1
and a scalar map
n
X n
X
s : R(Rn ) → R given by s(R) := Tr(R) := r(R)ii = Rhihi .
i=1 h,i=1
 

Note that Ker r ⊆ Ker s, so that Ker s = Ker s ∩(Ker r) ⊕ Ker r. Thus, we have
a decomposition
 
⊥ ⊥
(15.63) R(Rn ) = R1 ⊕ R2 ⊕ R3 := (Ker s) ⊕ Ker s ∩(Ker r) ⊕ Ker r.
The tensor representation O(n) → GL(R(Rn )) is given for g ∈O(n) by
(g · R)(v1 , . . . , v4 ) := R g −1 v1 , . . . , g −1 v4 .


and the representation O(n) → GL(S(Rn )) is defined similarly. Since the map r
is O(n)-invariant and s is O(n)-equivariant (i.e., g · s(R) = s(g · R) , for g ∈O(n)),
(15.63) is a decomposition of R(Rn ) into subspaces which are O(n)-invariant. The
subspace R1 , R2 and R3 are actually irreducible, since R(Rn ) is irreducible as a
GL(n, Rn )-module corresponding to the Young symmetrizer diagram
1 3
(15.64)
2 4
428 15. GEOMETRIC PRELIMINARIES

and r and s are the only independent contractions in R(Rn ) (see [436, 153ff]). This

symmetrizer takes a tensor in ⊗4 (Rn ) and symmetrizes it in the indices in positions
1 and 3 (the top row of the diagram), and in positions 2 and 4 (the bottom row of
the diagram). Then the result is antisymmetrized in the indices in positions 1 and
2 (the first column of the diagram) and in positions 3 and 4 (the second column of
the diagram). In order to determine the R1 , R2 , R3 components of R ∈ R(Rn ),
we introduce a bilinear, symmetric map (the Kulkarni-Nomizu product, up to
a constant factor)
∨ : S 2 (Rn ) × S 2 (Rn ) → R(Rn ) ,
say ∨(P, Q) := P ∨ Q, where
1
(15.65) (P ∨ Q)hijk := 2 (Phj Qik − Pij Qhk + Pik Qhj − Phk Qij ) .

Note that this is simply twice the result of applying the diagram (15.64) to P ⊗ Q.
Exercise 15.49. Check that P ∨ Q ∈ R(Rn ); i.e., verify conditions (A) and
(B) in (15.62).
Moreover, if I denotes the usual dot product on Rn (i.e., Iij = δij =Kronecker
delta), then
(I ∨ Q)hijk = 12 (δhj Qik − δij Qhk + δik Qhj − δhk Qij ) ,
r(I ∨ Q)ik = 21 (nQik − Qik + δik Tr(Q) − Qik )
= 21 ((n − 2) Qik + Tr(Q) δik ) .
Hence,
1
(15.66) r(I ∨ Q) = 2 ((n − 2) Q + Tr(Q) I) ,
(15.67) s(I ∨ Q) = (n − 1) Tr(Q) ,
and in particular,
1
(15.68) r(I ∨ I) = 2 ((n − 2) I + nI) = (n − 1) I, and
(15.69) s(I ∨ I) = (n − 1) n.

Proposition 15.50. The adjoint r∗ of r : R(Rn ) → S(Rn ) is 21 I∨ : S(Rn ) →


R(Rn ) ; i.e., r∗ (Q) = 2I ∨ Q, in the sense that for R ∈ R(Rn ) and Q ∈ S(Rn ),
hr(R) , Qi = R, 21 I ∨ Q .
For n > 2, I∨ : S(Rn ) → R1 ⊕ R2 is an isomorphism of vector spaces. More
precisely, if
S0 (Rn ) := {P ∈ S(Rn ) : Tr(P ) = 0} ,
then

I∨ : S0 (Rn ) ⊕ S0 (Rn ) → R1 ⊕ R2
respects the summands, and for P, Q ∈ S0 (Rn ),
(15.70) hI ∨ P, I ∨ Qi = (n − 2) hP, Qi .
15.5. CURVATURE IN RIEMANNIAN GEOMETRY 429

Proof. For R ∈ R(Rn ) and Q ∈ S(Rn ), we have (summing over repeated


indices),
hR, I ∨ Qi = Rhijk (I ∨ Q)hijk
= 12 Rhijk (δhj Qik − δij Qhk + δik Qhj − δhk Qij )
 
= 12 r(R)ik Qik + r(R)hk Qhk + r(R)hj Qhj + r(R)ij Qij
= 2r(R)ik Qik = h2r(R) , Qi .

This shows that r∗ (Q) = 21 I ∨ Q. Also, hR, I ∨ Ii = h2r(R) , Ii = 2s(R) , which


⊥ ⊥
shows that Ker(s) is spanned by I ∨ I. Since the image of r∗ is (Ker r) , the
n n ⊥
mapping I∨ : S(R ) → R(R ) maps onto (Ker r) = R1 ⊕ R2 . Moreover, for
P, Q ∈ S(Rn ),
hI ∨ P, I ∨ Qi = h2r(I ∨ P ) , Qi = h((n − 2) P + Tr(P ) I) , Qi
(15.71) = (n − 2) hP, Qi + Tr(P ) Tr(Q) .
In particular, (15.70) holds for P, Q ∈ S0 (Rn ), and
2 2 2
(15.72) |I ∨ Q| = (n − 2) |Q| + Tr(Q) ,
which shows that I∨ : S(Rn ) → R(Rn ) is injective for n 6= 2. Moreover, we have
(I∨)(S0 (Rn )) = R2 , since
I ∨ Q ∈ R2 ⇔ 0 = s(I ∨ Q) = Tr(r(I ∨ Q)) = (n − 1) Tr(Q) .
 
⊥ ⊥
Also, as S0 (Rn ) is spanned by I, (I∨) S0 (Rn ) = R1 by (15.68). 

For R ∈ R(Rn ), suppose that according to the decomposition (15.63)


R = R1 + R2 + R3 .
Then R1 is the projection of R onto Ker(s), namely
hR,I∨Ii s(R)
R1 := kI∨Ik2
(I ∨ I) = n(n−1) (I ∨ I) ,
2 2
 to compute kI ∨ Ik = (n1 − 2) n + n = 2n(n − 1). For
where (15.71) is used
2
n = 2, dim R R = 1, and hence R = R1 = 2 s(R)(I ∨ I), and R2 = R3 = 0.
Thus, we now assume n > 2. A candidate for R2 is obtained as follows. Note that
⊥ ⊥
I ∨ r(R) = 2r∗ (r(R)) ∈ (Ker r) , and subtracting off the (Ker s) component, we

get an element of Ker s ∩(Ker r) = R2 , namely (n − 2) hP, Qi + Tr(P ) Tr(Q)
 
I ∨ r(R) − hI∨r(R),I∨Ii
|I∨I|2 (I ∨ I) = I ∨ r(R) − (n−2)s(R)+ns(R)
n(n−1) I
1

= I ∨ r(R) − n s(R) I .
The operator A : R(Rn ) → R2 , given by
A(R) := I ∨ r(R) − n1 s(R) I



is not quite a projection onto Ker s ∩(Ker r) , but rather note that for R ∈ R2 ,
say R = I ∨ Q with Tr(Q) = 0, we have (using (15.66) and (15.67))
A(R) = A(I ∨ Q) = I ∨ r(I ∨ Q) − n1 s(I ∨ Q) I = 12 (n − 2)(I ∨ Q) .

430 15. GEOMETRIC PRELIMINARIES

2
Hence for n > 2, n−2 A|R2 is the identity and A|(R1 ⊕R3 ) = 0, since clearly A|R3 =
A|Ker r = 0 and
A(I ∨ I) = I ∨ r(I ∨ I) − n1 s(I ∨ I) I


= I ∨ (n − 1) I − n1 (n − 1) nI = 0.


2
Thus, n−2 A : R(Rn ) → R2 is an orthogonal projection and
2 2
∨ r(R) − n1 s(R) I .

R2 = n−2 A(R) = n−2 I

In summary, we have
Proposition 15.51. For n = 2, R = R1 = s(R)(I ∨ I) , while for n > 2,
R = R1 + R2 + R3 , where
s(R) ⊥
R1 = ∨ I) ∈ R1 = (Ker s) ,
n(n−1) (I
2 ⊥
I ∨ r(R) − n1 s(R) I ∈ R2 = Ker s ∩(Ker r) ,

R2 = n−2 and
R3 = R − R1 − R2 ∈ Ker r.
Remark 15.52. The parts R1 , R2 and R3 have names:
R1 is the constant curvature part of R,
R2 is the traceless Ricci part of R, and
(15.73) R3 (usually denoted W ) is the Weyl part of R.

s(R)
(a) R1 gets its name as follows. For R of the form n(n−1) (I ∨ I) , we have (from
n
(15.65)) that for independent vectors X, Y ∈ R with Π := span(X, Y ),
s(R)
R(X, Y, X, Y ) = ∨ I)(X, Y, X, Y )
n(n−1) (I
 
s(R) 2
= n(n−1) (X · X)(Y · Y ) − (X · Y ) ,

so that the sectional curvature (see (15.51, p. 423)) of Π, namely


R(X, Y, X, Y ) s(R)
K(Π) = 2 2 2 = ,
kXk kY k − (X · Y ) n(n − 1)
is independent of the plane Π (i.e., K is a constant function on the set of planes).
2
I ∨ r(R) − n1 s(R) I is determined by the traceless Ricci

(b) Note that R2 = n−2
tensor r(R) − n1 s(R) I, called so since Tr r(R) − n1 s(R) I = s(R) − s(R) = 0.


(c) The Weyl part of the curvature tensor of a Riemannian manifold M with met-
ric h is known as the Weyl conformal curvature tensor (or simply Weyl tensor )
of (M, h) , and is denoted by W . We should mention that each Tx M can be iso-
metrically identified with Rn . Such an identification is unique up to O(n) and the
various spaces R1 , R2 , R3 are O(n)-invariant, so that the split R = R1 + R2 + R3
can be made invariantly. Indeed, for the curvature tensor field R, we can replace I
in (15.73) by the metric h to define R1 , R2 and R3 . Thus,
W = R3 = R − R1 − R2
s(R) 2
∨ r(R) − n1 s(R) h .

=R− n(n−1) (h ∨ h) − n−2 h
15.5. CURVATURE IN RIEMANNIAN GEOMETRY 431

Let W # denote the (1, 3)-tensor obtained from W by raising the first index,
i
namely W # jkl = him Wmjkl . Then W # has the property that it is invariant
under a conformal change of metric, say h 7→ e2σ h for some σ ∈ C ∞ (M ). If
#
(Rσ ) σ denotes the (1, 3) version of the Riemann curvature tensor of e2σ h, then a
somewhat lengthy computation in [141] yields
#σ 2 # #
(15.74) R# − (Rσ ) = |dσ| (h ∨ h) + 2(e
σ ∨ h) ,
where σ e := ∇(dσ) − dσ ⊗ dσ is a symmetric 2-tensor. The 2-tensor ∇(dσ) is the
covariant derivative of dσ with respect to the Levi-Civita connection θ of h, and
it is known as the Hessian of σ. The symmetry of ∇(dσ) is due to fact that θ
is torsion-free. Since the right side of (15.74) has Weyl part 0, the (1, 3)-version
of the Weyl tensor is unchanged. Also in [141] it is shown that if W = 0 and
dim(M ) ≥ 4, then about each point x ∈ M , there is a neighborhood U and a
function σ ∈ C ∞ (U ), such that the curvature tensor Rσ of e2σ h is 0 (i.e., (M, h) is
conformally flat). The result (15.74) shows that W = 0 is necessary in order that
(M, h) be conformally flat. If dim(M ) = 3, then W = 0 is automatic, since
1 2 2
3 3 − 1 = dim R R3 = dim R1 R3 ⊕ R2 R3 .
   
6=
12
Thus, for dim(M ) = 3, R is determined by the Ricci tensor. However for dim(M ) =
3, W = 0 does not imply that (M, h) is conformally flat. For dim(M ) = 2, W = 0
is again automatic, but conformal flatness does not follow from W = 0. Instead,
one proves conformal flatness (i.e., the existence of isothermal parameters) by other
means.

Curvature Parts on 4-Manifolds and Self-Duality. We will focus on ori-


ented Riemannian 4-manifolds (M, h). First however, in any dimension, due to
the antisymmetry of R(W,  Z, X, Y ) in (X, Y ) and in (W, Z), we can view R as a
section of End Λ2 (T M ∗ ) , or an operator sending 2-forms to 2-forms. This opera-
tor is called the curvature operator. In terms of a local frame field, the curvature
operator applied to a 2-form α yields the 2-form R(α)
b defined by
1 kl
(15.75) R(α)
b
ij := 2 Rijkl α ,

where αkl = hkp hlq αpq . Note that for a metric of constant sectional curvature 1, R
b
is the identity. The symmetry R(W, Z, X, Y ) = R(X, Y, W, Z) implies that R b is a
symmetric endomorphisms of Λ2 (T M ∗ ), since
h(R(α)
b , β) = 41 Rijkl αkl β ij = 41 Rklij αkl β ij = h(α, R(β)).
b

For oriented Riemannian 4-manifolds, we have another automorphism of Λ2 (T M ∗ ),


namely the Hodge star ∗ : Λ2 (T M ∗ ) → Λ4−2 (T M ∗ ). While the decomposition,
R(Rn ) = R1 ⊕ R2 ⊕ R3 of (15.63) consists of O(n)-irreducible subspaces, we now
show that, for n = 4, R3 is  Recall that the Hodge star ∗
 not SO(4)-irreducible.
satisfies ∗2 = Id on Λ2 R4∗ , and we have Λ2 R4∗ = Λ2+ ⊕ Λ2− , where
Λ2+ := (1 + ∗) Λ2 R4∗ and Λ2− = (1 − ∗) Λ2 R4∗
 

are the self-dual and anti-self-dual subspaces (±1 eigenspaces of ∗). Relative to the
standard basis e1 , e2 , e3 , e4 of R4∗ , a basis of Λ2± is
e2 ∧ e3 ± e1 ∧ e4 , e3 ∧ e1 ± e2 ∧ e4 , e1 ∧ e2 ± e3 ∧ e4 .
432 15. GEOMETRIC PRELIMINARIES


We can write Rb and ∗ in block form relative to the decomposition Λ2 R4∗ =
2 2
Λ+ ⊕ Λ− , say
   
A B I 0
(15.76) R=
b , ∗= ,
BT C 0 −I

where A = AT and C = C T since R b is symmetric. We know that dim R R4 =
1 2 2
12 4 4 − 1 = 20, but the dimension of the space of all 6 × 6 symmetric matrices is
6(6 + 1) /2 = 21. The
 discrepancy is due to the fact that ∗ is orthogonal
 to R R4
in the space S Λ2 of all symmetric
 endomorphisms of Λ2 R4∗ due to the Bianchi
4
Identity, namely for R ∈ R R

(15.77) 0 = εijkl (Rijkl + Riljk + Riklj ) = 3εijkl Rijkl = 3 h∗, Ri ,

where we recall from (15.19) that (∗α)ij = 21 εijkl αkl . Thus,

R R4 ∼ = S 0 Λ2 := ∗⊥ := S ∈ S Λ2 : hS, ∗i = 0 .
   
(15.78)
By (15.76) and (15.77), we have 0 = h∗, Ri = Tr(A) − Tr(C), while Tr(A) +
Tr(C) = Tr(R) = 12 Rijij = 21 s(R). Hence, Tr(A) = Tr(C) = 41 s(R), and defining
Ae := A − 1 Tr(A) I and C
e := C − 1 Tr(C) I, we have
3 3
       
s(R) I 0 0 B A
e 0 0 0
(15.79) R= + + + e .
12 0 I BT 0 0 0 0 C

As s(R) , B, and (traceless symmetric) A e and Ce vary independently, each of the



summands in (15.79) varies over a subspace of R R4 , say C1 , C2 , C3 , and C4 , from
left to right. Note that
 2

C1 = α Id ∈ S Λ
 :α∈R ,
C2 = R ∈ S Λ2  : R ◦ ∗ = − ∗ ◦R ,
C3 = R ∈ S Λ2  : h∗, Ri = 0, R ◦ ∗ = ∗ ◦ R = R and
C4 = R ∈ S Λ2 : h∗, Ri = 0, R ◦ ∗ = ∗ ◦ R = −R .
Thus, as ∗ is SO(4)-invariant, each of C
i is an SO(4)-invariant subspace for the
tensor representation SO(4) → O S Λ2 . For Proposition 15.54 below, we will
use the following

Lemma 15.53. Let r : O(n) → O(V ) be an irreducible representation. Then


the restriction r|SO(n) is either irreducible, or V is the direct sum of two irreducible
r|SO(n) -invariant subspaces of equal dimension.

Proof. Suppose that r|SO(n) is not irreducible, and let V 0 be a proper, irre-
ducible r|SO(n) -invariant subspace of V . Let A ∈ O(n) with det A = −1. Then
r(A)(V 0 ) is r|SO(n) -invariant, since for any B ∈ SO(n) , we have C := A−1 BA ∈
SO(n), and
r(B)(r(A)(V 0 )) = r(BA)(V 0 ) = r(AC)(V 0 ) = r(A)(r(C) V 0 ) ⊆ r(A)(V 0 ) .
Since V 0 is irreducible, either r(A)(V 0 ) ∩ V 0 = V 0 or r(A)(V 0 ) ∩ V 0 = {0}. If
r(A)(V 0 ) ∩ V 0 = V 0 , then V 0 is a proper r-invariant subspace of V , contrary to
assumption. If r(A)(V 0 ) ∩ V 0 = {0}, then V 0 + r(A)(V 0 ) is a direct sum. Moreover,
15.5. CURVATURE IN RIEMANNIAN GEOMETRY 433

V 0 + r(A)(V 0 ) is also r-invariant, since any C ∈ O(n) is of the form BA−1 and AB 0
for some B, B 0 ∈ SO(n), and we have

r(C)(V 0 + r(A)(V 0 )) = r(C)(V 0 ) + r(C) r(A)(V 0 )


= r(AB 0 )(V 0 ) + r BA−1 r(A)(V 0 )


= r(A)(V 0 ) + r(B)(V 0 ) = V 0 + r(A)(V 0 ) .

Thus, V = V 0 + r(A)(V 0 ) is a direct sum of two r|SO(n) -invariant subspaces of


equal dimension. It remains to show that r(A)(V 0 ) is r|SO(n) -irreducible. If W is
a proper, irreducible r|SO(n) -invariant subspace of r(A)(V 0 ), then (contrary to the
−1
r|SO(n) -irreducibility of V 0 ) r(A) (W ) is a proper, r|SO(n) -invariant subspace of
V , since for any B ∈ SO(n) , BA−1 = A−1 B 0 for some B 0 ∈ SO(n) , and so
0

 
−1 −1 −1
r(B) r(A) (W ) = r(A) (r(B 0 )(W )) = r(A) (W ) . 

Proposition 15.54. With respect to the decomposition (see (15.63))


 
⊥ ⊥
R R4 = R1 ⊕ R2 ⊕ R3 := (Ker s) ⊕ Ker s ∩(Ker r) ⊕ Ker r,

(15.80)

∼ S 0 Λ2 of (15.78), we have
 
under the identification R R4 =

C1 ∼= R1 = {constant curvature parts} ,


C2 ∼= R2 = {traceless Ricci parts} , and
C3 + C4 ∼
= R3 = {Weyl parts} ,
in the terminology of (15.73).

Proof.
 We know that the Ri are irreducible,
 O(4)-invariant subspaces of
R R4 , with dim(R1 ) = 1, dim(R2 ) = dim S0 R4 = 4(4 + 1) /2 − 1 = 9, and

dim(R3 ) = dim R R4 − 10 = 42 42 − 1 /12 − 10 = 10.


 

By Lemma 15.53, the odd-dimensional O(4)-irreducible subspaces R1 and R2 re-


main irreducible under SO(4). Since the isomorphism R R4 ∼ = S 0 Λ2 is SO(4)-
 

equivariant and there are four SO(4)-invariant summands Ci , R3 must split into two
irreducible, SO(4)-invariant summands, each of dimension 5. These are necessarily
C3 and C4 , and then clearly C1 ∼
= R1 and C2 ∼= R2 . 

We can decompose the Riemann curvature tensor R of an oriented, Riemannian


4-manifold (M, h) as

R = R1 + R2 + W = R1 + R2 + W + + W − ,

where W + and W − correspond respectively to the C3 and C4 components of the


Weyl tensor W .

Definition 15.55. If W − = 0, then (M, h) is called self-dual, and if W − = 0,


then (M, h) is called anti-self-dual.

The terminology is fitting, since considering ∗ and W ± ∈ C ∞ (End(Λ(T M ∗ ))),


we have ∗ ◦ W ± = W ± ◦ ∗ = ±W ± .
434 15. GEOMETRIC PRELIMINARIES

6. Bochner-Weitzenböck Formulas
Let π : P → M be a principal G-bundle and let ω be a connection1-form on
k
P . Suppose that ρ : G → GL(W ) is a representation, and let Ω (P, W ) denote the
space of horizontal equivariant forms (see Definition 15.11, p. 401). In this section,
we assume that M is compact, oriented Riemannian n-manifold with metric h. Let
θ denote the Levi-Civita connection (see Definition 15.40, p. 418) on the principal
SO(n)-bundle πF : F M → M of oriented, orthonormal frames.
Fibered Products. Let
P ×f F M := {(p, u) ∈ P × F M : π(p) = πF (u)} .
The group G × SO(n) acts freely on P ×f F M via (p, u)(g1 , g2 ) = (pg1 , ug2 ), and
π ×f πF : P ×f F M → M
is readily verified to be a principal G × SO(n)-bundle, called the fibered product
of P and F M . The subscript f in ×f stands for fibered (note that P ×f F M 6=
P × F M ). Observe that
π1 : P ×f F M → P and π2 : P ×f F M → F M,
given by π1 (p, u) := p and π2 (p, u) := u are principal bundles with groups SO(n)
and G respectively. Note that π1∗ ω is a g-valued 1-form on P ×f F M , while π2∗ θ is a
so(n)-valued 1-form on P ×f F M . The direct sum π1∗ ω ⊕ π2∗ θ is a g ⊕ so(n)-valued
1-form on P ×f F M . It is not hard to verify that π1∗ ω ⊕ π2∗ θ is a connection 1-form
for π ×f πF : P ×f F M → M . To avoid cumbersome expressions, let us adopt the
notation
(15.81) ω ⊕ θ := π1∗ ω ⊕ π2∗ θ.
If Rn∗ denotes the dual space of Rn , then we define the space of tensors contravariant
of degree r and covariant of degree s by
r s
T r,s := Rn ⊗ · · · ⊗ Rn ⊗ Rn∗ ⊗ · · · ⊗ Rn∗ ,
and the tensor representation tr,s : SO(n) → GL(T r,s ) is given by
(15.82) tr,s (A)(v1 ⊗ · · · ⊗ vr ⊗ η1 ⊗ · · · ⊗ ηs )
:= Av1 ⊗ · · · ⊗ Avr ⊗ η1 ◦ A−1 ⊗ · · · ⊗ ηs ◦ A−1 .
 

Then we have a representation ρ ⊗ tr,s : G × SO(n) → GL(W ⊗ T r,s ), and we may


k
consider the spaces Ω (P ×f F M, W ⊗ T r,s ) of horizontal, equivariant W ⊗ T r,s -
valued k-forms on P ×f F M . We have the usual isomorphism (see (15.11), p. 403)
k
(15.83) Ω (P ×f F M, W ⊗ T r,s ) ∼= Ωk M,(P ×f F M ) ×G×SO(n) (W ⊗ T r,s )


between horizontal, equivariant k-forms and k-forms with values in the associated
vector bundle. More significantly for this section is the fact that either of the spaces
in (15.83) can be identified with the subspace of elements of
0
 
Ω P ×f F M, W ⊗ T r,s+k ∼
r,s+k
= C ∞ M,(P ×G W ) ⊗ T (M )


which are antisymmetric in the last k slots. This is accomplished via standard
horizontal vector fields (see Definition 15.45, p. 425). Recall that for w ∈ Rn , the
standard horizontal vector field w is defined on F M . However, we can take the
horizontal lift of w to P ×f F M , relative to the connection π1∗ ω on π2 : P ×f F M →
15.6. BOCHNER-WEITZENBÖCK FORMULAS 435

F M, in order to obtain a vector field w e on P ×f F M . Let us just call we the


standard horizontal vector field on P ×f F M associated to w ∈ Rn , and to avoid
notational complication use the same notation w as we do for w on F M . Note that
for (g, A) ∈ G×O(n), we have

(15.84) (R)(g,A)∗ w(p,u) = (A−1 w)(pg,u◦A) ,

(cf. (15.55), p. 425). Thus, regarding G and O(n) as subgroups of G×O(n), note
that w on P ×f F M is G-invariant, and although not SO(n)-invariant, it transforms
nicely.
k 0
For α ∈ Ω (P ×f F M, W ⊗ T r,s ), we define α0 ∈ Ω P ×f F M, W ⊗ T r,s+k


by
α0 (p, u)(η1 , . . . , ηr , v1 , . . . , vs+k ) := α(p,u) (η1 , . . . , ηr , v1 , . . . , vs )(vs+1 , . . . , vs+k ) ,
where η1 , . . . , ηr ∈ Rn∗ , v1 , . . . , vs+k ∈ Rn , and vs+1 , . . . , vs+k denote the standard
horizontal vector fields on P ×f F M associated with vs+1 , . . . , vs+k . Note that
0
α0 ∈ Ω P ×f F M, W ⊗ T r,s+k , since (using (15.84))


α0 ((p, u)(g, A))(η1 , . . . , ηr , v1 , . . . , vs+k )


= α(pg,u◦A) (η1 , . . . , ηr , v1 , . . . , vs )(vs+1 , . . . , vs+k )
 
= α(p,u) (η1 , . . . , ηr , v1 , . . . , vs ) R(g,A)∗ Av s+1 , . . . , R(g,A)∗ Av s+k
 


= R(g,A) α (η1 , . . . , ηr , v1 , . . . , vs ) Av s+1 , . . . , Av s+k
(p,u)
 
r,s −1 
= (ρ ⊗ t ) (g, A) α(p,u) (η1 , . . . , ηr , v1 , . . . , vs ) Av s+1 , . . . , Av s+k
= ρ g −1 α(p,u) (η1 ◦ A, . . . , ηr ◦ A, Av1 , . . . , Avs ) Av s+1 , . . . , Av s+k
  

= ρ g −1 (α0 (p, u)(η1 ◦ A, . . . , ηr ◦ A, Av1 , . . . , Avs+k ))



 −1

= ρ ⊗ tr,s+k (g, A) (α0 (p, u))(η1 , . . . , ηr , v1 , . . . , vs+k ) .

In addition to the covariant exterior differentiation operator


k k+1
(15.85) Dω⊕θ : Ω (P ×f F M, W ⊗ T r,s ) −→ Ω (P ×f F M, W ⊗ T r,s ) ,
we also have
0
∇ω⊕θ := Dω⊕θ : Ω P ×f F M, W ⊗ T r,s+k

(15.86)
1
−→ Ω P ×f F M, W ⊗ T r,s+k .


Even though one can regard


k 0
Ω (P ×f F M, W ⊗ T r,s ) ⊆ Ω P ×f F M, W ⊗ T r,s+k


and
k+1 0
(P ×f F M, W ⊗ T r,s ) ⊆ Ω P ×f F M, W ⊗ T r,s+k+1


1

= Ω P ×f F M, W ⊗ T r,s+k ,


the operator in (15.85) is not a restriction of the operator in (15.86) if k > 0.


Hence, we have introduced a different notation, namely ∇ω⊕θ in (15.86). We also
436 15. GEOMETRIC PRELIMINARIES

add that it is customary to denote the evaluation of the 1-form ∇ω⊕θ α on a vector
0
X by ∇ω⊕θ r,s+k

X α, namely for α ∈ Ω P ×f F M, W ⊗ T and X ∈ T (P ×f F M ),
∇ω⊕θ ω⊕θ

X α := ∇ α (X) .
Although not equal, Dω⊕θ α and ∇ω⊕θ α are nevertheless related by
Proposition 15.56. For
k 0
α ∈ Ω (P ×f F M, W ⊗ T r,s ) ⊆ Ω P ×f F M, W ⊗ T r,s+k


and x1 , . . . , xk+1 standard horizontal vector fields on P ×f F M associated with


x1 , . . . , xk+1 ∈ Rn , we have
k+1
X i+1
Dω⊕θ α (x1 , . . . , xk+1 ) = ∇xω⊕θ
 
(15.87) (−1) i
α (x1 , . . . , xbi , . . . , xk+1 ) .
i=1

Proof. Using the fact that [xi , xj ] is vertical (see (15.56)), we compute
Dω⊕θ α (x1 , . . . , xk+1 ) = (dα)(x1 , . . . , xk+1 )


k+1
X h  i
i+1
= (−1) xi α x1 , . . . , xbi , . . . , xk+1
i=1
k+1
X  
i+j
+ (−1) α [xi , xj ] , x1 , . . . , xbi , . . . , xbj , . . . , xk+1
1≤i<j≤k+1
k+1
X h  i
i+1
= (−1) xi α x1 , . . . , xbi , . . . , xk+1
i=1
k+1
X   
i+1
= (−1) d α x1 , . . . , xbi , . . . , xk+1 (xi )
i=1
k+1
X i+1
= (−1) d(α(x1 , . . . , xbi , . . . , xk+1 ))(xi )
i=1
k+1
X i+1
∇xω⊕θ

= (−1) i
α (x1 , . . . , xbi , . . . , xk+1 ) . 
i=1

Contractions and Components. Contractions of tensor fields are most ef-


ficiently displayed in terms of components. While one usually thinks that this
involves the choice of some coordinate system or ad hoc choice of basis, one of the
advantages of working on the frame bundle (or more generally P ×f F M ) is that
one can use the standard horizontal vector fields e1 , . . . , en corresponding to the
standard basis e1 , . . . , en of Rn and standard dual basis e1 , · · · , en ∈ Rn∗ . Thus,
k
there are standard W -valued components of α ∈ Ω (P ×f F M, W ⊗ T r,s ) defined
at each (p, u) ∈ P ×f F M by
αji11···i i1 ir

···js ;q1 ···qk (p, u) := α(p,u) (eq1 , · · · , eqk ) e , · · · , e , ej1 , · · · , ejs ∈ W.
r

(It might not be necessary to emphasize that we drop the semicolon in the α-
1
expression on the left side, if r = s = 0.) Let ϕ ∈ Ω (F M, Rn ) denote the canonical
1-form. Note that the components ϕ1 , . . . , ϕn of ϕ are R-valued forms vanishing
on vertical vectors. Using π2 : P ×f F M → F M , we can pull back the form ϕ to
15.6. BOCHNER-WEITZENBÖCK FORMULAS 437

m
a form π2∗ (ϕ) ∈ Ω (P ×f F M, Rn ), which we continue to denote by ϕ and whose
k
components are still denoted by ϕ1 , . . . , ϕn . For α ∈ Ω (P ×f F M, W ⊗ T r,s ), we
can write
1 X i1 ···ir
α ei1 , · · · , eir , ej1 , · · · , ejs = ϕq1 ∧ · · · ∧ ϕqk .

α
k! q ,...,q j1 ···js ;q1 ···qk
1 k

0 r,s
Also, for α ∈ Ω (P ×f F M, W ⊗ T ), we use the notation
 i1 ···ir i1 ···ir
αji11···i r
···js |q := ∇ ω⊕θ
e q
α = ∇ω⊕θ α j1 ···js ;q .
j1 ···js
H
Note that for any f ∈ C ∞ (P ×f F M, W ) , (df ) (ei ) = df (ei ) = ei [f ], from which
it follows that
αji11···i i1 ···ir
  i1 ···ir 
(15.88) ···js |q = d αj1 ···js (ei ) = eq αj1 ···js .
r

In terms of this notation, (15.87) can be rewritten as


k
i1 ···ir X l
(15.89) Dω⊕θ α j1 ···js ;q0 ···qk
= (−1) αji11···i
···js ;q0 ···qbl ···qk |ql .
r

l=0
0 
Moreover, if h denotes the Riemannian metric on M , and H ∈ Ω P ×f F M, T 0,2
is defined by
H(p, u)(v1 , v2 ) := h(u(v1 ) , u(v2 )) ,
n
then as u : R → Tπ(u) M is an isometry,
H(p, u)(ei , ej ) = h(u(ei ) , u(ej )) = δij .
Thus, the components of H are constant functions. Thus, in view of (15.88), indices
can be raised (or lowered) before or after applying ∇ω⊕θ or Dω⊕θ producing the
same result. We have already found (see Proposition 15.19, p. 406) the formal
nm
adjoint of Dω : Ωm (P ×G W ) → Ωm+1 (P ×G W ) to be δ ω := − (−1) ∗ Dω ∗, but
for the purpose of stating and proving the B-W formulas, it is convenient to have a
m
lifted version of δ ω defined directly on Ω (P ×f F M, W ) instead of Ωm (P ×G W ).
Since there is a composition of isomorphisms
Ψ : Ωm (M, P ×G W ) = ∼ Ωm (P, W ) =
∼ Ωm (P ×f F M, W ) ,
all we really need is the lifted version ∗, say
(15.90) ∗ := Ψ−1 ◦ ∗ ◦ Ψ,
in which case the lifted version of δ ω is given by
nm m+1 m
(15.91) δ ω = − (−1) ∗ Dω ∗ : Ω (P ×f F M, W ) → Ω (P ×f F M, W ) .
For a (ω ⊕ θ)-horizontal subspace H ⊆ T(p,u) (P ×f F M ), it is not hard to see that
∗ is simply the usual star operator acting on the restrictions of horizontal forms to
H, where H is given the metric and orientation which make
(15.92) (π ×f πF )∗ : H → Tπ(p) M
an orientation-preserving isometry. In terms of the components αq1 ···qm of α ∈
m n−m
Ω (P, W ), one can verify that ∗α ∈ Ω (P, W ) is given by
1 q1 ···qm
(∗α)qm+1 ···qn = α εq1 ···qm qm+1 ···qn .
m!
438 15. GEOMETRIC PRELIMINARIES

where εq1 ···qn is antisymmetric in its indices with ε1···n = 1.

Proposition 15.57. (a) For

m m+1
Dω : Ω (P ×f F M, W ) → Ω (P ×f F M, W )

m
and α ∈ Ω (P ×f F M, W ), we have

m
X l
(Dω α)q0 ···qm = Dω⊕θ α

(15.93) q0 ···qm
= (−1) αq0 ···qbl ···qm |ql .
l=0

(b) For

m+1 m
δω : Ω (P ×f F M, W ) → Ω (P ×f F M, W ) ,

m+1
given by (15.91) and β ∈ Ω (P ×f F M, W ), we have

m
X
nm
(15.94) (δ ω β)q1 ···qm = (− (−1) ∗ Dω (∗β))q1 ···qm = − β iq1 ···qm |i .
i=1

Proof. The equation (15.93) is just a restatement of (15.87) in the special


case T r,s = T 0,0 = R. For (15.94) we compute as follows. From

1 q ···q
(∗β)qm+2 ···qn = βq ···q ε 1 m+1qm+2 ···qn
(m + 1)! 1 m+1

we obtain

n
X l−m−1
Dω (∗β)rm+1 ···rn = (−1) (∗β)rm+1 ···rbl ···rn |rl
l=m+1
n
1 X l−m−1

q ···q

= (−1) βq1 ···qm+1 |rl ε 1 m+1 rm+1 ···rbl ···rn .
(m + 1)!
l=m+1

In the following, we use the identity

i ···i
εi1 ···ip kp+1 ···kn εj1 ···jp kp+1 ···kn = (n − p)!δj11 ···jpp ,
15.6. BOCHNER-WEITZENBÖCK FORMULAS 439

i ···i
where the generalized Kronecker delta δj11 ···jpp := +1 (or −1), depending on whether
j1 · · · jp is an even (or odd) permutation of i1 · · · ip , and 0 otherwise. Then
(n − m − 1)!(m + 1)!∗Dω (∗β)r1 ···rm
rm+1 ···rn
= Dω (∗β) εrm+1 ···rn r1 ···rm
n
!
X l−m−1 q1 ···qm+1 rm+1 ···rbl ···rn
= (−1) βq1 ···qm+1 |rl ε εrm+1 ···rn r1 ···rm
l=m+1
n
!
X m(n−m) q1 ···qm+1 rm+1 ···rbl ···rn
= (−1) βq1 ···qm+1 |rl ε εr1 ···rm rl rm+1 ···rbl ···rn
l=m+1
n
|r q ···q
mn−m
X
= (n − m − 1)!(−1) βq1 ···qm+1 l δr11···rm
m+1
rl
l=m+1
n
|rl
mn−m
X
= (n − m − 1)!(m + 1)!(−1) βr1 ···rm rl
l=m+1
n
X
mn
= (n − m − 1)!(m + 1)!(−1) β rlr1 ···rm |rl ,
l=m+1
nm
(−1)
and multiplication by − (n−m−1)!(m+1)! yields (15.94). 

The Bochner-Weitzenböck Formula. If we regard


k k 0
Ω (P, W ) ∼
= Ω (P ×f F M, W ) ⊆ Ω P ×f F M, W ⊗ T 0,k ,


then we have
k 1
∇ω⊕θ : Ω (P, W ) −→ Ω P ×f F M, W ⊗ T 0,k .

∗
We denote the formal adjoint of this map by ∇ω⊕θ . We then have the so-called
connection Laplacian
∗ k k
∇ω⊕θ ∇ω⊕θ : Ω (P, W ) −→ Ω (P, W ) .
0 
Since ∇ω⊕θ is the same as Dω⊕θ on Ω P ×f F M, W ⊗ T 0,k , we can use Propo-
sition 15.57 (with m = 0) to obtain
Corollary 15.58. For
k k
α ∈ Ω (P, W ) ∼= Ω (P ×f F M, W ) ∼
= Ωk (M, P ×G W ) ,
we have  ∗ 
∇ω⊕θ ∇ω⊕θ α = −αi1 ···ik |j j
i1 ···ik

We also have the Hodge Laplacian


k k
∆ω := δ ω Dω + Dω δ ω : Ω (P ×f F M, W ) → Ω (P ×f F M, W )
which does not depend on the Levi-Civita connection and only depends on the
metric h on M via the Hodge star operator.
The next result is the Bochner-Weitzenböck formula. It relates the two Lapla-

cians ∇ω⊕θ ∇ω⊕θ and ∆ω . First, there is some notation. Let
∗ ω
Ωω
ij = (π1 Ω )(ei , ej ) ,
440 15. GEOMETRIC PRELIMINARIES

where π1 : P ×f F M → P and Ωω denotes the curvature form of ω. Let

Rijkl := ei · Ωθ (ek , el )(ej ) .




These are the components of the Riemann curvature tensor R regarded as in


0  0 
Ω F M, T 0,4 (Rn ) or Ω P ×f F M, T 0,4 (Rn ) . The components of the Ricci cur-
vature are then
Rjl := δ ik Rijkl = Rijil .
k
Theorem 15.59 (Bochner-Weitzenböck Formula). For α ∈ Ω (P, W ), we have
 ∗ 
(15.95) (∆ω α)i1 ···ik = ∇ω⊕θ ∇ω⊕θ α
i1 ···ik
k k
X l   X l
+ (−1) ρ0 Ωω
il j αji − (−1) Ril j αji
1 ···il ···ik 1 ···il ···ik
b b
l=1 l=1
k
X l
− (−1) Rjmim il j αji
1 ···im−1 jm im+1 ···il ···ik .
b
l6=m=1

Proof. Using Proposition 15.57, we have


j
(δ ω Dω α)i1 ···ik = − (Dω α) i1 ···ik |j
k
X l
= −αi1 ···ik |jj − (−1) αji and
1 ···il ···ik |il j
b
l=1
k
X l+1
(Dω δ ω α)i1 ···ik = (−1) (δ ω α)i1 ···ibl ···ik |il
l=1
k
X l+1
=− (−1) αji .
1 ···il ···ik |jil
b
l=1

Thus,
k
X  
l
(∆ω α)i1 ···ik = −αi1 ···ik |jj + (−1) αji − αji .
1 ···il ···ik |jil 1 ···il ···ik |il j
b b
l=1
 ∗ 
The first term on the right is −αi1 ···ik |jj = ∇ω⊕θ ∇ω⊕θ α . Regarding
i1 ···ik
0 
α ∈ Ω P ×f F M, W ⊗ T 0,k and using (15.93) and (15.13), we have
j
αji − αji = Dω⊕θ Dω⊕θ α i1 ···ib ···i ;i j
1 ···il ···ik |jil 1 ···il ···ik |il j l k l
b b
 0
j
= ρ ⊗ t0,k Ωω⊕θ (α)
 
.
i1 ···ibl ···ik ;il j

Recall that for w ∈ W , β ∈ Λk (Rn∗ ), g ∈ G and B ∈ SO(n), we have


j1 jk
ρ ⊗ t0,k (g, A) (w ⊗ β) i B −1 . . . B −1
  
= βj1 ···jk ρ(g)(w) .
1 ···ik i1 ik
15.6. BOCHNER-WEITZENBÖCK FORMULAS 441

Thus, for A ∈ g and C ∈ so(n), we have


 0   0  
ρ ⊗ t0,k (A, C)(w ⊗ β) = ρ0 ⊗ I + I ⊗ t0,k (A, C)(w ⊗ β)
i1 ···ik i ···i
 1 k
0 j1 jk
= ρ (A)(w) βj1 ···jk − C i1 βj1 ···jk + . . . + C ik βj1 ···jk w.
Hence,
 0    
ρ ⊗ t0,k Ωω⊕θ (α) = ρ0 Ω ω
il j αi0 i1 ···ibl ···ik
i0 i1 ···ibl ···ik ;il j
X
− Rj0i0 il j αj0 i1 ···ibl ···ik − Rjmim il j αi0 i1 ···im−1 jm im+1 ···ibl ···ik ,
m∈{1,...,b
l,...,k}

where in the sum, we do not mean to imply that m < l, only that m 6= l. Thus,
raising i0 and contracting i0 with j, we obtain
 0  j  j 
ρ ⊗ t0,k Ωω⊕θ (α) = ρ 0 Ωωi l j α i ···ib ···i
i1 ···ibl ···ik ;il j 1 l k
 X 
j0 jm j
− R il αj0 i1 ···ibl ···ik + R im il j α i ···i .
1 m−1 jm im+1 ···il ···ik
b
m∈{1,...,l,...,k}
b

Finally we obtain
k
 ∗  X l
 
(∆ω α)i1 ···ik = ∇ω⊕θ ∇ω⊕θ α + (−1) αji − αji
i1 ···ik 1 ···il ···ik |jil 1 ···il ···ik |il j
b b
l=1
k
 ∗  X l
 0  j
= ∇ω⊕θ ∇ω⊕θ α + (−1) ρ ⊗ t0,k Ωω⊕θ (α)
i1 ···ik i1 ···ibl ···ik ;il j
l=1
 ∗ 
= ∇ω⊕θ ∇ω⊕θ α
i1 ···ik
  j  
k 0 ω j0
X l ρ Ω il j α i1 ···ibl ···ik
− R α
il j0 i1 ···il ···ik
+ (−1)  P
b

jm j
l=1
− m∈{1,...,l,...,k}
b R α
im il j i1 ···im−1 jm im+1 ···ibl ···ik

k
 ∗  X l  j 
= ∇ω⊕θ ∇ω⊕θ α + (−1) ρ0 Ωω il j α i ···ib ···i
i1 ···ik 1 l k
l=1
k
X k
X
l l
− (−1) Ril j αji − (−1) Rjmim il j αji ,
1 ···il ···ik 1 ···im−1 jm im+1 ···il ···ik
b b
l=1 l6=m=1

as required. 

Special Cases. We consider some special cases for the degree k of the form
k
α ∈ Ω (P, W ) and the dimension n of the underlying manifold M .
1
1. (k = 1) For 1-forms α ∈ Ω (P, W ) ∼ = Ω1 (M, P ×G W ), the last term in
(15.95) is absent. Taking the L inner product of ∆ω α with α, we then have
2

2
(∆ω α, α) = ∇ω⊕θ α+ (Ric(α) , α) − (ρ0 (Ωω ) α, α) .
It follows that if Ric − ρ0 (Ωω ) ∈End Ω1 (M, P ×G W ) is pointwise nonnegative,


but not zero everywhere, then ∆ω α = 0 ⇒ α = 0. In particular, if G is trivial, we


442 15. GEOMETRIC PRELIMINARIES

obtain S. Bochner’s result (see [62]) that a compact, Riemannian manifold with
positive Ricci curvature admits no nonzero harmonic 1-form.
2 2
2. (k = 2) For α ∈ Ω (P, W ) ∼ = Ω (P, W ) ∼= Ω2 (M, P ×G W ) , we get

 
(∆ω α)i1 i2 = ∇ω⊕θ ∇ω⊕θ α

i1 i2
  j   j 
0 ω 0 ω
− ρ Ωi1 j α i2 − ρ Ωi2 j α i1
 
+ Ri1 j αji2 − Ri2 j αji1 + Rmi2 i1 j αjm − Rmi1 i2 j αjm .


In terms of the product ∨ (see (15.65), p. 428) and using the hat “∧” for converting
curvature type tensors to operators on 2-forms (see (15.75), p. 431), we have

Ri1 j αji2 − Ri2 j αji1 = 2 (h ∨ Ric) (α) i1 i2 .


Using the first Bianchi identity (15.53) and known symmetries of R,


Rmi2 i1 j αjm − Rmi1 i2 j αjm = Rmi2 i1 j αjm + Rmi1 ji2 αjm
= −Rmji2 i1 αjm = Rmji1 i2 αjm = −Ri1 i2 jm αjm = −2R(α)
b
i1 i2 ,

where R(α) denotes the image of α under the curvature operator R (see (15.75),
p. 431). Moreover, it is convenient to define [ρ0 (Ωω ) , α] via
 j   j 
(15.96) [ρ0 (Ωω ) , α]i1 i2 := ρ0 Ωω
i1 j α i2 − ρ0 Ωωi2 j α i1 .
2
Then (for α ∈ Ω (P, W )), we can write
∗ ∧
∆ω α = ∇ω⊕θ ∇ω⊕θ α − [ρ0 (Ωω ) , α] + 2(h ∨ Ric) (α) − 2R(α)
b .
When n = dim M = 3, note that by Proposition 15.51 (p. 15.51) and the fact
that the Weyl tensor is zero for n = 3,
R = 61 S(h ∨ h) + 2h ∨ Ric − 31 Sh = h ∨ 2Ric − 12 Sh .
 

Thus for n = 3,
∗ ∧
∆ω α = ∇ω⊕θ ∇ω⊕θ α − [ρ0 (Ωω ) , α] + 2(h ∨ Ric) (α) − 2R(α)
b
ω⊕θ ∗ ∧
∇ω⊕θ α − [ρ0 (Ωω ) , α] + 2(h ∨ Ric) (α)

= ∇
∧
− 2 h ∨ 2Ric − 21 Sh (α)
ω⊕θ ∗ ω⊕θ 0 ∧
α − [ρ (Ωω ) , α] − 2(h ∨ Ric) (α) + Sα.

= ∇ ∇
If n = 4, we have
1
∨ h) + h ∨ Ric − 41 Sh + W,

R= 12 S(h
and so
∗ ∧
∆ω α = ∇ω⊕θ ∇ω⊕θ α − [ρ0 (Ωω ) , α] + 2(h ∨ Ric) (α) − 2R(α)
b
∗ ∧
= ∇ω⊕θ ∇ω⊕θ α − [ρ0 (Ωω ) , α] + 2(h ∨ Ric) (α)
1
∧
S(h ∨ h) + h ∨ Ric − 14 Sh + W (α)

− 2 12
∗
= ∇ω⊕θ ∇ω⊕θ α − [ρ0 (Ωω ) , α] + 31 Sα + W (α) .
We have decompositions
α = α+ + α− , Ωω = Ωω+ + Ωω− and W = W + + W −
15.7. CHARACTERISTIC CLASSES AND CURVATURE FORMS 443

into self-dual and


 anti-self-dual parts. Under the SO(4)-equivariant isomorphism
so(4) ∼= Λ2 R4∗ , given by lowering an index, the irreducible, SO(4)-invariant
subspaces Λ+ and Λ− , correspond to irreducible, subspaces, say so+ and so− which
are SO(4)-invariant with respect to the adjoint action. Thus,
(15.97) ∼ so+ ⊕ so−
so(4) =
with [so(4) , so± ] ⊆ so± , and so [so− , so+ ] ⊆ so− ∩ so+ = {0} and [so± , so± ] ⊆ so± .
It then follows from (15.96) that [ρ0 (Ωω− ) , α+ ] = [ρ0 (Ωω+ ) , α− ] = 0, where the fact
that α is W -valued is irrelevant to the argument. Thus,
∗
∆ω α = ∇ω⊕θ ∇ω⊕θ α − ρ0 Ωω+ , α+ − ρ0 Ωω− , α−
     

+ 13 Sα + W + α+ + W − α− .
 

The following direct consequence will be useful to us later.


2
Proposition 15.60. When α ∈ Ω (P, W ) is anti-self-dual (α+ = 0), Ωω is
self-dual (Ωω− = 0) and h is a self-dual metric (i.e., W eyl− = 0), we have
∗
∆ω α = ∇ω⊕θ ∇ω⊕θ α + 31 Sα.
Thus,
2
(∆ω α, α) = ∇ω⊕θ α + 13 (Sα, α) .
If S ≥ 0 and ∆ω α = 0, then either S = 0 and ∇ω⊕θ α = 0, or S 6= 0 and α = 0.

7. Characteristic Classes and Curvature Forms


Before working through this Section, we recommend the reader to browse
through our Part III (pp.251ff) to register where the various characteristic classes
appear naturally in classical Index Theory.
Let E → M be a C ∞ complex, Hermitian vector bundle of complex dimension
m, with Hermitian inner product h·, ·i. For x ∈ M, a (unitary) frame of Ex is a
linear, isometry u : Cm → Ex (i.e., hu(z) , u(w)i = z1 w1 + · · · + zm wm ). If u is a
frame and A ∈ U (m), then uA := u ◦ A is a frame. The set U (E) consisting of all
frames at all points of M can be made into a C ∞ manifold, such that π : U (E) → M
is a principal U (m)-bundle, namely the bundle of unitary frames of E. Without
difficulty, one can prove that U (E) ×U(m) Cm ∼ = E, via [u, w] 7→ u(w), where the
representation is the inclusion map U (m) → GL(m, C).
Chern Classes as Curvature Forms. Let ω be a connection 1-form on
U (E), which is a 1-form with values in the Lie algebra
u(m) = {A ∈ gl(m, C) : A∗ = −A}
of U (m), and with the required properties of Definition 15.5, p. 398. We will express
the Chern classes of E → M in terms of the curvature Ωω of ω.
We begin by defining functions sk : gl(m, C) → C by means of
m
X
det(A + tI) = sk (A) tm−k .
k=0

Note that sk (A) is a homogeneous polynomial of degree k in the entries of A, namely


1 X j1 ···jk i1
sk (A) = δi1 ···ik a j1 · · · aikjk ,
k!
(i),(j)
444 15. GEOMETRIC PRELIMINARIES


where A = aij , (i) := (i1 , . . . , ik ) ranges over all sequences of k distinct elements
···jk
of {1, . . . , m}, and δij11···ik
= +1 (resp. −1), depending on whether (i) is an even
···jk
(resp. odd) permutation of (j), and δij11···i k
= 0 if {i1 , . . . , ik } 6= {j1 , . . . , jk }. The
sk are invariant under the adjoint action of GL(m, C) on gl(m, C), in the sense that
sk BAB −1 = sk (A) for all A ∈ gl(m, C) and B ∈ GL(m, C), since


det BAB −1 + tI = det B(A + tI) B −1 = det(A + tI) .


 

2
The curvature Ωω ∈ Ω (U (E) , u(m)) can be regarded as a matrix of C-valued
 j
2-forms, say Ωω = Ωij , such that Ωij = −Ω i , and we define
1 X j1 ···jk i1
sk (Ωω ) := δi1 ···ik Ω j1 ∧ · · · ∧ Ωikjk ∈ Ω2k (U (E) , C) .
k!
(i),(j)

From the fact that Rg∗ Ωω = adg−1 Ωω for all g ∈ U (m) and the invariance of sk
under the adjoint action, it follows that sk (Ωω ) is invariant under Rg∗ . Since Ωω
also vanishes on vertical vectors, we know that there is a 2k-form, say σk (Ωω ) ∈
Ω2k (M, C), on M such that sk (Ωω ) = π ∗ σk (Ωω ). Note that
H H
π ∗ (dσk (Ωω )) = d(π ∗ σk (Ωω )) = d(π ∗ σk (Ωω )) = d(sk (Ωω ))
k
1 X X j1 ···jk i1 i
= δi1 ···ik Ω j1 ∧ · · · ∧ (dΩ pjp )H ∧ · · · ∧ Ωikjk = 0,
k! p=1
(i),(j)

i i
since (Ωi j )H = Ωi j and (dΩ pjp )H = (Dω Ωω ) pjp = 0 by the Bianchi identity.
Thus, σk (Ωω ) is closed and determines a de Rham cohomology class [σk (Ωω )] ∈
H 2k (M ; C).
Definition 15.61. (a) The k-th Chern class of the complex, Hermitian vector
bundle E is defined as
i k
[σk (Ωω )] = σk 2π
i
Ωω .
  
ck (E) := 2π
(b) The form
 
i ω
ck (E, ω) := σk Ω ∈ Ω2k (M, C)

is the k-th Chern form of the complex Hermitian vector bundle E → M for the
connection 1-form ω on U (E).
i i

The factor i in 2π Ωω ensures that σk 2π Ωω ∈ Ω2k (M, R), since
iq iq j
iΩ jq = −iΩ jq = iΩ qiq ⇒ σk (iΩω ) = σk iΩω = σk (iΩω ) ,


2k i ω
and
 i ω
 ∈ H2k (M ; R). The factor of 2π in 2π Ω is a normalization implying
so ck (E)
σk 2π Ω ∈ H (M ; Z). A full proof of this would carry us too far afield. How-
ever, we will show that [σk (Ωω )], and hence ck (E), is independent of the choice of
ω. Indeed, let
1
α := ω1 − ω0 ∈ Ω (U (E) , u(m)) and ωt := ω0 + tα, 0 ≤ t ≤ 1.
Then ωt is a connection with curvature Ωωt = dωt + 21 [ωt , ωt ]. Let sek denote the

k-linear symmetric form such that sk (A) = sek A, . k. ., A . Note that the invariance
15.7. CHARACTERISTIC CLASSES AND CURVATURE FORMS 445

of sk and hence sek under the adjoint action yields (at u = 0)


d

0= du sek exp(uB) A1 exp(−uB) , . k. ., exp(uB) Ak exp(−uB)
= sek ([B, A1 ] , A2 , . . . , Ak ) + sek (A1 , [B, A2 ] , A3 , . . . , Ak )
+ . . . + sek (A1 , A2 , A3 , . . . , [B, Ak ]) .
Using this and the Bianchi identity dΩωt + [ωt , Ωωt ] = 0, we have
1 d
sk (Ωωt ) = sek dt
d
(Ωωt ) , Ωωt , k−1. . . , Ωωt

k dt
= sek dα + [ωt , α] , Ωωt , k−1 . . . , Ωωt


= sek dα, Ωωt , k−1


. . . , Ωωt + sek [ωt , α] , Ωωt , k−1
. . . , Ωωt
 

= sek dα, Ωωt , k−1


. . . , Ωωt − (k − 1) sek α, [ωt , Ωωt ] , Ωωt , k−2 . . . , Ωωt
 

= sek dα, Ωωt , k−1


. . . , Ωωt + (k − 1) sek α, dΩωt , Ωωt , k−2 . . . , Ωωt
 

= d sek α, Ωωt , k−1


. . . , Ωωt .


Thus,
Z 1 
ω1 ω0 ωt ωt
sk (Ω ) − sk (Ω ) = kd sek (α, Ω , . . . , Ω ) dt .
0
ωt ωt 2k
Since, sek (α, Ω , . . . , Ω ) ∈ Ω (U (E) , C) is invariant under Rg∗ ,
Z 1
sek (α, Ωωt , . . . , Ωωt ) dt = π ∗ β,
0

for some unique form β ∈ Ω2k (M, C), and σk (Ωω1 ) − σk (Ωω0 ) = dβ. Hence,
[σk (Ωω0 )] = [σk (Ωω1 )], as required.
Remark 15.62. Suppose that U(E) is reducible to an SU(m)-bundle, i.e., we
have a subprincipal SU(m)-bundle U(E)0 → M . Then we show that c1 (E) = 0.
Let ω0 be an arbitrary connection 1-form on U(E)0 . We can extend the distribution
of horizontal subspaces for ω0 on U(E)0 to all of U(E) by the requiring that the
distribution be Rg∗ -invariant for all g ∈ U(m). Let ω denote the resulting connec-
tion on U(E). We know that ω0 and Ωω0 are su(m)-valued. While ω is u(m)-valued
and has values outside su(m), we can show that Ωω is su(m)-valued. Indeed,
H H
Ωω |U(E)0 = (dω) |U(E)0 = (dω0 ) = Ωω0 .
Thus, Ωω has values in su(m) on ω-horizontal subspaces at points of U(E)0 . Since
su(m) is invariant under the adjoint action of U(m) and Rg∗ Ωω = adg−1 Ωω , we
know that Ωω is su(m)-valued throughout U(E). Since s1 (A) = Tr(A) = 0 for
A ∈ su(m), we have c1 (E) = 0, when U(E) is reducible to an SU(m)-bundle.
Once the Chern classes ck (E) are determined, the Chern character ch(E) ∈
H ∗ (M ; Q) may be defined in terms of the ck (E). Alternatively, we can get ch(E)
directly as follows. For A ∈ u(m) ,
   X ∞  k !
i k 1 i
Tr exp t A = rk (A) t , for rk (A) := Tr A .
2π k! 2π
k=0

As with the sk , the rk are invariant under the adjoint action of U(m) on u(m).
Hence the horizontal form rk (Ωω ) is Rg∗ -invariant and rk (Ωω ) = π ∗ (ρk (Ωω )) for
446 15. GEOMETRIC PRELIMINARIES

some closed form ρk (Ωω ) ∈ H 2k (M ; R) (actually H 2k (M ; Q) , where Q denotes the


field of rational numbers) whose class [ρk (Ωω )] is independent of ω. Then


X M ]
[ 12 dim M
ω
(15.98) ch(E) := [ρk (Ω )] ∈ H 2k (M ; Q) .
k=0 k=0

Note that 8π 2 r2 (A) = − Tr A 2
for A ∈ u(m), and so

−1
(15.99) ch(E)2 = [ρ2 (Ωω )] = [Tr(Ωω ∧ Ωω )] ∈ H 4 (M ; Q) ,
8π 2
2
where we have regarded Ωω ∈ Ω2 (M, End(E)) ⊆ Ω (U (E) , u(m)) and composition
of endomorphisms is implicit in the wedge Ωω ∧ Ωω .

The Pfaffian. Other characteristic classes can be represented by forms. The


Euler class of an oriented Riemannian 2m-manifold is represented by the Gauss-
Bonnet form
1 X
GB Ωθ εi1 ···i2m Ωθi1 i2 ∧ · · · ∧ Ωθi2m−1 i2m

(15.100) :=
22m π m m!
(i)

where Ωθ denotes the curvature form of any connection θ (not necessarily Levi-
Civita) on the principal SO(2m)-bundle πF : F M → M of oriented orthonormal
frames. In (15.100), we may regard Ωθ ∈ Ω2 (M, End(T M )) and the components
Ωθij are relative to a locally defined orthonormal frame field. Alternatively, (15.100)
may be regarded as the unique form on M which when pulled back to F M via
πF∗ denotes the form given by the same formula, but where Ωθij = Ωθ (ei , ej ), the
2
components of Ωθ ∈ Ω (F M, so(2m)) relative to the standard, horizontal fields
e1 , . . . , e2m . The form (15.100) arises from the homogenous polynomial of degree
m on so(2m) known as the Pfaffian, defined for A ∈ so(2m), by
m
(−1) X
(15.101) Pf(A) := εi1 ···i2m Ai1 i2 · · · Ai2m−1 i2m .
2m m!
(i)

1
θ
 θ

Thus, GB Ω = Pf 2π Ω . The Pfaffian is invariant under the adjoint action of
SO(2m) (but not O(2m)). Indeed, for B ∈ O(2m), we have

Pf BAB −1 = Pf BAB T
 
m
(−1) X 
= m εi1 ···i2m (Bi1 j1 Aj1 j2 Bi2 j2 ) · · · Bi2m−1 j2m−1 Aj2m−1 j2m Bi2m j2m
2 m!
(i)
m
(−1) X
= m εi1 ···i2m Bi1 j1 Bi2 j2 · · · Bi2m−1 j2m−1 Bi2m j2m Aj1 j2 · · · Aj2m−1 j2m
2 m!
(i)

= det(B) Pf(A) .

The Gauss Bonnet Theorem, which is a special case of the Index Theorem, asserts
that the integral of the Gauss-Bonnet form over (compact) M is χ(M ).
15.7. CHARACTERISTIC CLASSES AND CURVATURE FORMS 447

The Pontryagin Classes.


I We will also encounter the Pontryagin classes pk (M ) ∈ H 4k (M ; Z) of M . They
were shortly introduced by Lev Semenovich Pontryagin in 1942 in [340] and elaborated
by him in 1947 in his highly influential treatise [341] with English translation in [342].
They have a special status in the history of modern mathematics:
• They were the first characteristic classes clearly perceived as global invariants
which measure the deviation of a local product structure from a global product
structure.
• They were defined systematically and offered at once a unifying geometric con-
cept in algebraic topology, differential geometry and algebraic geometry.
• They were designed to investigate composed and partitioned manifolds and de-
formation problems by distinguishing whether a manifold is a boundary or not
(Pontryagin classes vanish on a manifold that is a boundary).
• They moved research about low-dimensional (in particular, 4-dimensional) topol-
ogy into the center of attention and showed up (via the Hirzebruch Signature
Formula) in practically all results concerning 4q-dimensional manifolds since
then.
• Being defined in terms of smooth structures but applied to topological problems,
they laid the ground for the modern research on existence and uniqueness of
smooth structures.
• They continue to raise hard problems around their homotopy invariance and
their wider ramifications.
Pontryagin recalls in [344, Chapter 2] (similarly in [343]) that he constructed and
applied the classes (then called cycles) already before the war, years before [340]. He was
led to their construction, he wrote, when he in 1936 led the foundation to the study of the
“interrelations between the investigation of higher homotopy groups of spheres and the
theory of smooth manifolds”. A short account of the mathematical adventures connected
to these classes can be found in [303]. J

Today, these classes may be defined in terms of the Chern classes of the com-
plexified tangent bundle TC M := C ⊗ T M, which can be regarded as the associated
bundle F M ×SO(n) Cn where the representation SO(n) → U(n) is just inclusion.
Note that F M is a principal subbundle of the unitary frame bundle U (TC M ) of
TC M, where the Hermitian metric H on TC M is given in terms of the complex
bilinear extension hC of the Riemannian metric h via

H(X, Y ) = hC X, Ȳ for X, Y ∈ TC M.
A connection θ on F M determines a unique connection 1-form
θc ∈ Ω1 (U (TC M ) , u(n)) ,
such that θ = θc |F M . By definition, the Pontryagin class pk (M ) ∈ H 4k (M ; Z) is
1 θc
 4k
represented by the unique 4k-form σ2k 2π Ω ∈ Ω (M ), such that
(15.102)
2k
pk Ωθ : = πc ∗ σ2k 1 θc 1 θc i
Ωθc
   
2π Ω = s2k 2π Ω
= (−i) s2k 2π
k 1 X j ···j
i θc
Ωθi1cj1 ∧ · · · ∧ Ωθi2k

= (−1) s2k 2π Ω = 2k
δi11···i2k
2k c
j2k ,
(2π) (2k)! (i),(j)
i
 
for πc : U (TC M ) → M . In other words, since c2k (TC M ) = σ2k 2π Ωθc , we have
k
(15.103) pk (M ) = (−1) c2k (TC M ) ∈ H 4k (M ; Z) .
448 15. GEOMETRIC PRELIMINARIES

  
Since Ωθ = Ωθc |F M , it follows that σ2k Ωθc = σ2k Ωθ , where σ2k Ωθ is the
unique 4k-form, such that for π : F M → M ,
1 X j ···j
π ∗ σ2k Ωθ = s2k 2π 1
Ωθ = Ωθi1 j1 ∧ · · · ∧ Ωθi2k j2k
 
2k
δi11···i2k
2k

(2π) (2k)! (i),(j)


θ T
Note that
1 θ
 Ω has values in so(n). For A ∈ so(n) (so that A = −A), we have
sk 2π Ω = 0 for k odd, since
n
X
sk (A) tn−k = det(A + tI) = det AT + tI = det(−A + tI)


k=0
n
X n
X
n n n−k k
= (−1) det(A − tI) = (−1) sk (A)(−t) = sk (A)(−1) tn−k .
k=0 k=0

By the same argument, when E is a real Riemannian vector bundle we have


ck (C ⊗ E) = 0 for k odd. For more information on this approach to character-
istic classes, see [249] and [301].

Other Characteristic Classes Related to Index Theory. There are other


characteristic classes that will arise in index theorems. Some of these are defined
and manipulated more efficiently through the use of power series as follows. First
note that for A ∈ gl(ν, C) with eigenvalues {λ1 , . . . , λν }, there is B ∈ GL(m, C)
such that BAB −1 is upper triangular with diagonal entries λ1 , . . . , λν . Then
ν
X ν
Y ν
X
sk (A) tν−k = det(A + tI) = (λj + t) = σk (λ1 , . . . , λν ) tν−k ,
k=0 j=1 k=0

where σ0 := 1, σ1 := i=1 λi , and generally
X
σk (λ1 , . . . , λν ) := λi1 · · · λik
1≤i1 <···<ik ≤ν

denotes the elementary symmetric polynomial of degree k in x1 , . . . , xν . We have


seen that each sk (A) (and hence each σk ) together with a Hermitian vector bundle
E → M with a connection gives rise to a characteristic form or class, namely the
i
k-th Chern form σk 2π Ωω or class ck (E). Nowany polynomial in the σk gives rise
i
to a corresponding polynomial in the σk 2π Ωω where the multiplication is wedge
product, or a polynomial in the ck (E) where the multiplication is cup product. By
the Fundamental Theorem of Symmetric Polynomials, any symmetric polynomial in
(λ1 , . . . , λν ) can be expressed uniquely as a polynomial in σ1 , . . . , σν . Hence, from
a symmetric polynomial in (λ1 , . . . , λν ), we can obtain new characteristic forms or
classes, which are (however) ultimately polynomials in Chern forms or classes.

Multiplicative Classes. One way to manufacture symmetric polynomials


forming a symmetric function (e.g., the product) of a given power series P∞in each
λk , and considering the Taylor polynomials, goes as follows. Let b(x) = n=0 bn xn
be a formal power series in the single variable x with bn ∈ C. We may form the
product
(15.104)

X X∞ ∞
X
b(x1 ) . . . b(xν ) = βk1 ,...,kν xk11 · · · xkνν = Bk (σ1 , . . . , σk ) ,
k=0 k1 +···+kν =k k=0
15.7. CHARACTERISTIC CLASSES AND CURVATURE FORMS 449

where σj := σj (x1 , . . . , xν ) for j = 1, . . . , ν. By the Fundamental Theorem of


Elementary Polynomials, the Bk (σ1 , . . . , σk ) are uniquely determined by b. The
sequence {Bk (σ1 , . . . , σk )} is known as the multiplicative sequence determined by
b(x); the ideas here are due to F. Hirzebruch (see [207]). Given A ∈ gl(ν, C),
and taking {x1 , . . . , xν } = {iλ1 , . . . , iλν } where {λ1 , . . . , λν } denotes the set of
eigenvalues of A, we obtain a function
A 7→ Bk (A) := Bk (σ1 (iλ1 , . . . , iλν ) , . . . , σk (iλ1 , . . . , iλν ))
which (like the set of eigenvalues of A) is invariant under the adjoint action. Note
that if A ∈ u(ν), then theiλj are real. If each σj in Bk (σ1 , . . . , σk ) is replaced by
i
the Chern form σk 2π Ωω where Ωω denotes the curvature form of a connection
1-form ω on the unitary frame bundle U (E) for a Hermitian vector bundle E → M ,
then we obtain a closed 2k-form
i
Ωω , . . . , σk 2π
i
Ωω ∈ Ω2k (M, C)
 
Bk (E, ω) := Bk σ1 2π
which determines a cohomology class, say Bk (E) ∈ H 2k (M ; C). Hence, for each
formal power series b(x) = n=0 bn xn , there is an associated total form and total
P
class, namely
B(E, ω) := B0 (E, ω) + B1 (E, ω) + · · · ∈ Ω• (M, C) , and
B(E) := B0 (E) + B1 (E) + · · · ∈ H ∗ (M ; C) .
Since there seems to be no official notation to denote the assignments
(b(x) , E, ω) 7→ B(E, ω) ∈ Ω• (M, C) and (b(x) , E) 7→ B(E) ∈ H ∗ (M ; C)
of a characteristic form (or class) to a formal power series and bundle with connec-
tion, let
MF(b(x) , E, ω) := B(E, ω) and MC(b(x) , E) := B(E) ,
where MF stands for multiplicative form and MC stands for multiplicative class.
We use MFk (b(x) , E, ω) and MCk (b(x) , E) for the homogeneous parts:
X
MF(b(x) , E, ω) = MFk (b(x) , E, ω) and
X k
MC(b(x) , E) = MCk (b(x) , E) .
k
For two formal power series b1 (x) and b2 (x), we have
(15.105) MC(b1 (x) b2 (x) , E) = MC(b1 (x) , E) MC(b2 (x) , E) ,
using b1 (x1 ) b2 (x1 ) · · · b1 (xν ) b2 (xν ) = b1 (x1 ) . . . b1 (xν ) b2 (x1 ) . . . b2 (xν ).
Given another Hermitian vector bundle E 0 → M of dimension ν 0 , we may form
the direct sum E ⊕ E 0 → M . Connections ω and ω 0 on U (E) and U (E 0 ) yield a
0 0
connection ω ⊕ ω 0 on U (E ⊕ E 0 ) , with curvature form Ωω⊕ω = Ωω ⊕ Ωω . Note
that for x001 . . . , x00ν , x00ν+1 , . . . , x00ν+ν 0 = (x1 . . . , xν , x01 , . . . , x0ν 0 ), we have



X
Bk00 (σ100 , . . . , σk0000 ) = b(x001 ) . . . b(x00ν ) b x00ν+1 . . . b x00ν+ν 0
 

k00 =0
= b(x1 ) . . . b(xν ) b(x01 ) . . . b(x0ν 0 )

X ∞
X
= Bk (σ1 , . . . , σk ) Bk0 (σ10 , . . . , σk0 ) .
k=0 k0 =0
450 15. GEOMETRIC PRELIMINARIES

Consequently, in the ring H ∗ (M ; C) we have


(15.106) MC(b(x) , E ⊕ E 0 ) = MC(b(x) , E) MC(b(x) , E 0 ) .
Since the conjugate bundle E is the associated bundle U (E) ×U(ν) Cν relative
to the conjugate representation, the curvature form for the conjugate bundle E is
i ω
− 2π Ω . If A ∈ u(ν), and {x1 , . . . , xν } = {iλ1 , . . . , iλν } where {λ1 , . . . , λν } denotes
the set of eigenvalues of A, then the eigenvalues of A are {−λ1 , . . . , −λν } and
{−iλ1 , . . . , −iλν } = {−x1 , . . . , −xν }. Thus,

(15.107) MC b(x) , E = MC(b(−x) , E) .
Note that
the total Chern class of E = c(E) = MC(1 + x, E) .
Todd Class and L-Polynomials. A class which arises in the Hirzebruch-
Riemann-Roch Theorem is the total Todd class of E defined by
 
x
Td(E) := MC , E ∈ H ∗ (M ; Q) ,
1 − e−x
where Q denotes the rationals. For a compact, complex manifold M , Td(T M ) [M ]
is the Todd genus of M , which is a kind of holomorphic Euler characteristic of
M . One can compute Td(E) in terms of Chern classes. Indeed, using the Algebra
SymmetricPolynomials in Mathematica, we get (where ck = ck (E))
T d0 (E) = 1, T d1 (E) = 12 c1 , T d2 (E) = 12
1
c2 + c21 ,


1 1
−c4 + c3 c1 + 3c22 + 4c2 c21 − c41 , . . . .

T d3 (E) = 24 c1 c2 , T d4 (E) = 720
In treating the case where E is the complexification of a real, even-dimensional
Riemannian bundle F → M (i.e., E = C⊗F ), we proceed as follows. If A ∈ so(ν, R),
where ν = 2µ is even, then the λj are not only pure imaginary, but they come in
conjugate pairs, so that
(15.108) (x1 , . . . , xν ) = (iλ1 , . . . , iλν ) = (y1 , −y1 , . . . , yµ , −yµ ) for yl ∈ R.

Hence it seems more appropriate to express σj (x1 , . . . , xν ) in terms of σl y12 , . . . , yµ2 .
To this end, note that
Xν Yν Yµ
σj (x1 , . . . , xν ) = (1 + xk ) = (1 + yl )(1 − yl )
j=1 k=1 l=1
Yµ X µ l
1 − yl2 = (−1) σl y12 , . . . , yµ2 .
 
=
l=1 l=1
Hence if (15.108) holds, then

0, for j odd,
σj (x1 , . . . , xν ) = l 2 2

(−1) σl y1 , . . . , yµ , for j = 2l even.
k
Since pk (M ) = (−1) c2k (TC M ) (15.103), it makes sense to define the Pontryagin
classes of F by
k
pk (F ) := (−1) c2k (C ⊗ F ) ∈ H 4k (M ; Z)
This implies that for b(x) even, MC(b(x) , C ⊗ F ) can be obtained by writing
2 2
b(x1 ) . . . b(xν ) = b(y1 ) b(−y1 ) . . . b(yµ ) b(−yµ ) = b(y1 ) · · · b(yµ )

X
ek σ1 y12 , . . . , yµ2 , . . . , σk y12 , . . . , yµ2 ,
 
= B
k=0
15.7. CHARACTERISTIC CLASSES AND CURVATURE FORMS 451


and then replacing σj y12 , . . . , yµ2 by the j-th Pontryagin class pj (F ) ∈ H 4j (M ; Z).
Thus, for real, even-dimensional, Riemannian bundles F and b(x) even, we use the
more direct notation
 
2
MC b(y) , F := MC(b(x) , C ⊗ F ) .

There are various special cases arising in index theorems, which we now consider.
For a real, 2µ-dimensional Riemannian bundle F → M , we have the total A b class
of F defined by
 
y/2
(15.109) A(F ) := MC
b ,F ,
sinh(y/2)
which occurs in the index formula for the Dirac operator and its twists. We have

Yν/2 yj /2 X
bk σ1 y 2 , . . . , y 2 , . . . , σk y 2 , . . . , y 2 ,
 
= A 1 µ 1 µ
j=1 sinh(yj /2)
k=0

where (with the aid of Mathematica if desired)


2
Ab0 = 1, A b1 (σ1 ) = −1 σ1 , Ab2 (σ1 , σ2 ) = −4σ2 + 7σ1
24 5760
3
−16σ3 + 44σ 2 σ 1 − 31σ 1
Ab3 (σ1 , σ2 , σ3 ) =
967680
2 2 4
−192σ 4 + 512σ3 σ1 + 208σ2 − 904σ2 σ1 + 381σ1
(15.110) A b4 (σ1 , . . . , σ3 ) = ,....
464486400
To obtain Abk (F ), replace each σj by pj (F ). In connection with the Hirzebruch-
Signature Theorem, we have
 
y
the total Hirzebruch L class of F = L(F ) := MC , F , and
tanh y

1 7σ2 − σ12
L0 = 1, L1 (σ1 ) = σ1 , L2 (σ1 , σ2 ) = ,
3 45
62σ3 − 13σ2 σ1 − 2σ13
L3 (σ1 , σ2 , σ3 ) = ,
945
381σ4 − 71σ3 σ1 − 19σ22 + 22σ2 σ12 − 3σ14
(15.111) L4 (σ1 , . . . , σ4 ) = ,....
14175
Suppose that F → M is the realification of a complex bundle FC → M (i.e.,
we just restrict scalar multiplication for F√ C to real scalars). If J : F → F denotes
the map given by scalar multiplication by −1, then there is C-linear extension of
J, say JC : C ⊗ F → C ⊗ F . Since JC2 = − Id, we have C ⊗ F = F 1,0 ⊕ F 0,1 , where
F 1,0 := {V − iJV : V ∈ F } denotes the +i eigenbundle of the JC and F 0,1 :=
{V + iJV : V ∈ F } denotes the −i eigenbundle of JC . Note that FC ∼ = F 1,0 via
V 7→ V − iJV , and F C ∼ =F 0,1
via V 7→ V + iJV . Since
2
x2 x2

x/2
= 2 = x/2 x/2  
sinh(x/2) ex/2 − e−x/2 e e − e−x/2 e−x/2 ex/2 − e−x/2
x2 x −x
= = ,
(ex − 1)(1 − e−x ) 1 − e−x 1 − ex
452 15. GEOMETRIC PRELIMINARIES

it follows from (15.105), (15.107) and (15.106) that


b )2 = Td(FC ) Td FC = Td FC ⊕ FC = Td(C ⊗ F ) .
 
A(F

Recalculating Characteristic Classes. Some characteristic classes are not


expressible in terms of MC(b(x) , E) for some formal power series b(x), but they
can be described in a similar way. The Chern character ch(E) is not MC(ex , E).
However, adding (instead of multiplying as in (15.104)) one obtains

X
ex1 + · · · + exν = chk (σ1 , . . . , σk )
k=0

If each σj in chk (σ1 , . . . , σk ) is replaced by the Chern class cj (E), then we obtain
(15.112) ch(E) = ch0 (E) + ch1 (E) + · · · ∈ H ∗ (M ; R) .
We compute
ch0 (E) = dim E, ch1 (E) = c1 , ch2 (E) = 21 c21 − c2 ,
ch3 (E) = 16 3c3 − 3c2 c1 + c31 ,


1
−4c4 + 4c3 c1 + 2c22 − 4c2 c21 + c41 , . . . .

ch4 (E) = 24
While we do not have ch(E ⊕ E 0 ) = ch(E) ch(E 0 ) as in (15.106),
ch(E ⊕ E 0 ) = ch(E) + ch(E 0 ) and ch(E ⊗ E 0 ) = ch(E) ch(E 0 ) .
0 0
In terms of curvature forms, the first relation is clear from Ωω⊕ω = Ωω ⊕ Ωω , while
0 ω
the
 second  follows from the fact that the curvature form for E ⊗ E is (Ω ⊗ Id) ⊕
0
Id ⊗Ωω together with
Xν Xν 0 0 Xν Xν 0 0
exk +xk0 = exk exk 0 .
k=1 k0 =1 k=1 k0 =1
More generally, one could consider elementary polynomials σk (b(x1 ) , . . . , b(xν )) for
k other than 1 or ν, although finding uses for such might be a challenge.
The Euler class of a real, oriented Riemannian bundle F of dimension 2µ is
not generally expressible in terms of Pontryagin classes of C ⊗ F . We proceed as
follows. If A ∈ so(2µ, R), then A is SO(2µ)-similar to a matrix of the form
µ  
M 0 −yk
.
yk 0
k=1

The eigenvalues λj of A come in pure imaginary conjugate pairs ±iyk , so that


(x1 , . . . , xν ) := (iλ1 , . . . , iλν ) = (y1 , −y1 , . . . , yµ , −yµ ) for yl ∈ R, and
µ
(−1) X
Pf(A) : = εi1 ···i2µ Ai1 i2 · · · Ai2µ−1 i2µ
2µ µ!
(i)
µ µ
= (−1) A12 · · · A2µ−1,2µ = (−1) (−y1 ) · · ·(−yµ ) = y1 · · · yµ .
Note that y1 · · · yµ is not a symmetric polynomial in y12 , · · · , yµ2 . However, if ω is a
connection on the bundle π : SO(F ) → M of oriented frames of F and Ωω denotes
the curvature, then
µ
π ∗ χ(F, ω) = (−1) Pf 2π1
Ωω

(15.113)
15.7. CHARACTERISTIC CLASSES AND CURVATURE FORMS 453

for a unique closed, 2µ-form χ(F, ω) ∈ Ω2µ (M, R), the Euler form of E relative to
ω, which by definition represents the Euler class
(15.114) χ(F ) := [χ(F, ω)] ∈ H 2µ (M ; R) .

In the case F = T M and ω = θ, note for the Gauss-Bonnet form GB Ωθ , defined
in (15.100), that
(15.115)
m
π ∗ GB Ωθ = (−1) Pf 2π 1
Ωθ = π ∗ χ(T M, θ) =⇒ GB Ωθ = χ(T M, θ) .
  

Remark 15.63. We prefer to write χ(T M ) as GB(T M ). There is a good reason


for this. The Gauss-Bonnet-Chern Theorem states that the Euler characteristic
χ(M ) of M, defined as the alternating sum of numbers of faces (or Betti numbers)
of M , is given by Z
GB Ωθ = GB(T M ) [M ] .

χ(M ) =
M
On the other hand, for a generic characteristic class, say C(T M ) ∈ H ∗ (M ), fre-
quently one defines C(M ) to be C(T M ) [M ]. The Gauss-Bonnet-Chern Theorem is
not simply a definition, and yet that is exactly what it looks like if one writes it as
χ(M ) = χ(T M ) [M ]. Thus, we prefer to use GB(T M ) in place of χ(T M ), although
admittedly changing established notation is a losing battle, no matter how noble
the cause.
Observe that A ∈ so(2µ) is the realification of B = diag(iy1 , . . . , iyµ ) ∈ su(µ),
and
µ µ
(15.116) det(iB) = (−1) y1 · · · yµ = (−1) Pf(A) .

Unifications on Almost-Complex Manifolds. For a manifold M , J ∈


End(T M ) is an almost-complex structure if J 2 = −I. If J exists, then T M
becomes a complex vector bundle by defining (a + ib) X = aX + bJX and M
is an almost-complex manifold. The formula (15.116) will now be used to
show that for a compact, almost-complex, manifold M with dimR M = 2m, we
have cm (T M ) = GB(T M ) = χ(T M ). If h0 is any Riemannian metric on M ,
then h(X, Y ) := h0 (X, Y ) + h0 (JX, JY ) is compatible with the complex structure,
say J, on T M (i.e., h(JX, JY ) = h(X, Y )). We then have a Hermitian metric
hX, Y i := h(X, Y ) + ih(X, JY ) on T M regarded as a complex vector space. Note
that
hJX, Y i = h(JX, Y ) + ih(JX, JY ) = h(JX, Y ) + ih(X, Y )
= −h(X, JY ) + ih(X, Y ) = i(h(X, Y ) + ih(X, JY )) = i hX, Y i
hX, JY i = h(X, JY ) + ih X, J 2 Y = h(X, JY ) − ih(X, Y )


= −i(h(X, Y ) + ih(X, JY )) = −i hX, Y i .


The unitary frame bundle U M := U (T M ) (relative to h·, ·i) is then a subbundle
of F M . We may regard F M as the principal bundle U M ×U(m) O(2m) associated
to U M via the inclusion U(m) → O(2m); see 15.20, p.407. If the Levi-Civita
connection for h on F M restricts to a connection, say ω, on U M (i.e., the horizontal
subspaces at points in U M are contained in T (U M )), then M is Kähler (by one
definition). However, if this is not the case, then we can still uniquely extend any
connection ω on U M to a connection, say θ, on F M (see Proposition 15.22, p.408).
454 15. GEOMETRIC PRELIMINARIES

If we regard ω as u(m)-valued, then Ωθ |U M is just the realification of the curvature


Ωω . Now,
1 X
GB Ωθ = 2m m εi1 ···i2m Ωθi1 i2 ∧ · · · ∧ Ωθi2m−1 i2m

2 π m! (i)
 m X
1 1
= m εi1 ···i2m Ωθi1 i2 ∧ · · · ∧ Ωθi2m−1 i2m
2 m! 2π (i)
m 1
Ωθ , and

= (−1) Pf 2π
   m X
i ω 1 i ···jm i i
σm Ω = δij11···i (Ωω ) 1j1 ∧ · · · ∧(Ωω ) mjm
2π m! 2π m
(i),(j)
i ω

= det 2π Ω .
Hence, in view of (15.116) and restricting to U M , we have
M
GB Ωθ = (−1) Pf 2π 1
Ωθ = det 2π i
Ωω = σ m i ω
   
2π Ω , and

(15.117) cm (T M ) = GB(T M ) = χ(T M ) .


It often happens that a Hermitian bundle E arises an associated bundle E =
P ×G W , relative to some unitary representation r : G → U (W ), for some principal
G-bundle P → M which is not necessarily the unitary frame bundle U (E). Since
r : G → U (W ) is a homomorphism of the Lie group for P to U (W ), there is an
associated principal U (W )-bundle (see Proposition 15.20, p. 407)
P × U (W )
P 0 = P ×G U (W ) := = {[p, g 0 ] : p ∈ P, g 0 ∈ U (W )} ,
G
and an r-equivariant map Γ : P → P 0 . There is also a map
F : P 0 → U (E) , given by
F ([p, g 0 ])(w) = [p, g 0 (w)] , for p ∈ P, g 0 ∈ U (W ) , w ∈ W.
Note that F is well-defined and equivariant, since
F pg, r g −1 g 0 (w) = pg, r g −1 g 0 (w) = [p, g 0 (w)] , and for h0 ∈ U (W ) ,
     

F ([p, g 0 ] h0 )(w) = F ([p, g 0 ◦ h0 ])(w) = [p, g 0 (h0 (w))]


= F ([p, g 0 ])(h0 (w)) = (F ([p, g 0 ]) ◦ h0 )(w) .
As F is also bijective, it is an isomorphism of principal bundles. Thus we have a
morphism
R := F ◦ Γ : P → U (E)
and R(P ) may be regarded as a principal r(G)-subbundle of the U (W )-bundle
U (E). Using Proposition 15.22 (p.15.22), if ω is a connection on P , then there is a
unique connection ω 0 on P 0 such that ω = Γ∗ ω 0 . For the connection ωE := F −1∗ ω 0
on U (E), we then have
(15.118) r0 ◦ ω = R−1∗ ωE and r0 ◦ Ωω = R−1∗ ΩωE .
In the notation of (15.104), let
Bk (P, ω, r0 ) := Bk σ1 i 0
◦ Ωω , . . . , σ k i 0
◦ Ωω
 
2π r 2π r , and
0 0 0 •
MF(b(x) , P, ω, r ) := B0 (ω, r ) + B1 (ω0 , r ) + · · · ∈ Ω (M, C) .
15.8. HOLONOMY 455

From r0 ◦ Ωω = R−1∗ ΩωE , we get Bk (P, ω, r0 ) = Bk (E, ωE ). Thus,


MF(b(x) , E, ωE ) = MF(b(x) , P, ω, r0 ) and
MC(b(x) , E) = MC(b(x) , P, r0 ) := class of MF(b(x) , P, ω, r0 ) .
With similar notation, one also has
c(E, ωE ) = c(P, ω, r0 ) and c(E) = c(P, r0 ) , and similarly
(15.119) ch(E, ωE ) = ch(P, ω, r0 ) and ch(E) = ch(P, r0 ) .
Consider the special case G = U(ν) = U(Cν ) and r : U(ν) → U(W ). If
E0 := P ×U(ν) Cν , then we can often find MC(b(x) , E) = MC(b(x) , ω, r0 ) or
other characteristic of E classes (such as ch(E)) in terms of the Chern classes
ck (E0 ). If possible, one just expresses the elementary symmetric polynomials in
the eigenvalues of r0 (A) in terms of those for A ∈ u(ν). For example, consider the
exterior product bundle E = Λp (E0 ), where r : U(ν) → U(Λp (Cν )). If A ∈ u(ν)
has eigenvectors e1 , . . . , eν with eigenvalues λ1 , . . . , λν , then an eigenbasis of r0 (A)
is 
ei1 ∧ · · · ∧ eip : 1 ≤ i1 < · · · < ip ≤ n ,
and the associated eigenvalues of r0 (A) are λi1 + · · · + λip . For {x1 , . . . , xν } =
{iλ1 , . . . , iλν }, the j-th elementary symmetric polynomial in the xi1 + · · · + xip
(1 ≤ i1 < · · · < ip ≤ n) is the coefficient of tj in the expansion of
Y  
1 + xi1 + · · · + xip t ,
(i)p

where the multi-index (i)p ranges over {1 ≤ i1 < · · · < ip ≤ n}. These coefficients
can in turn be expressed as polynomials in the σk (x1 , . . . , xν ). The Chern classes
of Λp (E0 ) are then the same polynomials in the ck (E0 ). Since
Yν Xν X
(1 + exk ) =

p
exp xi1 + · · · + xip ,
k=1 p=1 (i)
• 2j
the Chern character chj (Λ (E0 )) ∈ H (M ; Q) can be found by expanding the
product on the left and writing the symmetric, homogeneous j-th degree part of
the power series as a polynomial in the σk (x1 , . . . , xν ) , regarded as ck (E0 ).
When E0 = C ⊗ F0 for some real, Riemannian bundle F0 of dimension 2µ, then
ch(Λ• (C ⊗ F0 )) may be obtained by expanding
Yµ  Yµ    
(1 + eyk ) 1 + e−yk = eyk /2 e−yk /2 + eyk /2 e−yk /2 eyk /2 + e−yk /2
k=1 k=1

(15.120) = 4 cosh2 (yk /2)
k=1

in terms of σk y12 , . . . , yµ2 which are then replaced by the Pontryagin classes pk (F0 ).
In other words,
ch(Λ• (C ⊗ F0 )) = MC 4 cosh2 (y/2) , F0 .

(15.121)

8. Holonomy
In this section, we assume that M is connected. Let ω be a connection 1-form
on the principal G-bundle π : P → M . Fix a point p0 ∈ P , and let P0 denote the
set of all points p ∈ P which can be joined to p0 by a smooth horizontal curve
γ : [a, b] → P , say γ(a) = p0 , γ(b) = p and ω(γ 0 (t)) = 0 for t ∈ (a, b). The
holonomy group of ω with reference point p0 is Hol(ω, p0 ) := {g ∈ G : p0 g ∈ P0 }.
456 15. GEOMETRIC PRELIMINARIES

It can be proven (see [248, 83-85]) that P0 is an immersed submanifold of P , and


π|P0 : P0 → M is a principal Hol(ω, p0 )-bundle, which is known as the holonomy
bundle of ω through p0 . If Hol(ω, p0 ) is a proper subgroup of G, then ω is said to
be reducible to Hol(ω, p0 ). If Hol(ω, p0 ) = G, then ω is irreducible.
The isotropy subgroup at ω for the action of the group GA(P ) of gauge trans-
formations on the space C(P ) of connection 1-forms on P will be denoted by
Iω := {F ∈ GA(P ) : F · ω = ω}. Under the isomorphism Φ : C(P, G) → GA(P )
of Proposition 15.28 (p. 411), we can identify Iω with the subgroup Φ−1 (Iω ) of
C(P, G).
Proposition 15.64. The homomorphism Iω → G , given by Φ(f ) 7→ f (p0 )
maps Iω isomorphically onto the centralizer of Hol(ω, p0 ) in G, namely
Z(Hol(ω, p0 )) := {g ∈ G : gg0 = g0 g for all g0 ∈ Hol(ω, p0 )} .
Proof. Let Φ(f ) ∈ Iω and g0 ∈ Hol(ω, p0 ). To prove that f (p0 ) ∈ Z(Hol(ω, p0 )),
we need to show that f (p0 ) g0 = g0 f (p0 ). Let γ be a horizontal curve joining p0
to p0 g0 . Then γ · f (p0 ) is a horizontal curve joining p0 f (p0 ) to p0 g0 f (p0 ). Since

Φ(f ) ω = ω, Φ(f ) ◦ γ is a horizontal curve joining p0 f (p0 ) to Φ(f )(p0 g0 ) =
Φ(f )(p0 ) g0 = p0 f (p0 ) g0 . Since
π(γ(t) · f (p0 )) = π(Φ(f )(γ(t))) ,
the curves γ ·f (p0 ) and Φ(f )◦γ are horizontal lifts of the same curve in M , and they
have the same initial point, namely p0 f (p0 ). By the uniqueness of horizontal lifts
with the same starting point (see [248, 69]), the endpoints p0 g0 f (p0 ) and p0 f (p0 ) g0
must agree, whence g0 f (p0 ) = f (p0 ) g0 (i.e., f (p0 ) ∈ Z(Hol(ω, p0 ))). To see that
Iω → G is injective, we use (15.27) on p. 411, namely

−1
(Φ(f ) · ω)(X) = f (p) f −1

(15.122) ∗p
(X) + f (p) ω(X) f (p) ,

for X ∈ Tp P . If Φ(f ) ∈  Iω , then Φ(f ) · ω = ω and (15.122) implies that if X


is horizontal, then f −1 ∗p (X) = 0. Then f −1 (and hence f ) is constant on all
horizontal curves and in particular f is the constant f (p0 ) on P0 . Since P0 meets
each fiber of P and f is equivariant, f is uniquely determined by f (p0 ). Thus, Iω →
G is injective. To prove that Iω → Z(Hol(ω, p0 )) is onto, let g 0 ∈ Z(Hol(ω, p0 )) ,
and let f (q0 ) = g 0 for all q0 ∈ P0 . For arbitrary p ∈ P , there is some g ∈ G such
that pg ∈ P0 , and we define
f (p) := gf (pg) g −1 = g −1 g 0 g.
To show that f is well-defined, suppose that ph ∈ P0 . Then gh−1 ∈ Hol(ω, p0 ) and
g 0 ∈ Z(Hol(ω, p0 )) ⇒ g 0 gh−1 = gh−1 g 0 ⇒ g −1 g 0 g = h−1 g 0 h.
 

Note that by the definition of P0 , the horizontal subspace Hq0 of ω at any q0 ∈ P0


is contained in Tq0 P0 . Since f |P0 is constant, for each X ∈ Hq0 , we have
−1
(Φ(f ) · ω)(X) = f (q0 ) ω(X) f (q0 ) = 0.
Thus, the horizontal subspace of Φ(f ) · ω at any q0 ∈ P0 coincides with Hq0 . Since
P0 meets each fiber of P and horizontal subspaces are Rg∗ -invariant, the horizontal
subspaces of Φ(f ) · ω and ω coincide at all points of P . Thus, Φ(f ) ∈ Iω . 
15.8. HOLONOMY 457

2
Proposition 15.65. The curvature Ωω ∈ Ω (P, g) at any point of the holonomy
bundle P0 of ω (with reference point p0 ) has values in the Lie algebra g0 of the
holonomy group Hol(ω, p0 ).
Proof. Note that ω|P0 clearly has values in g0 . Thus, (dω) |P0 = d(ω|P0 ) has
values in g0 . Since the horizontal subspace
 of ω at any q0 ∈ P0 is contained in
Tq0 P0 , we have Ωω (X, Y ) = dω X H , Y H ∈ g0 for all X, Y ∈ Tq0 P . 

Proposition 15.66. Let π : P → M be a principal U(1)-bundle with a con-


nection 1-form ω, and  let ρ : V → P be a local trivialization of P . Let D denote
the closed unit disk reit : r ∈ [0, 1] , t ∈ R ⊂ C, and let f : D → V denote the
restriction of a smooth immersion of a larger open disk. Then there is a unique
function h : [0, 2π] → R, such that γ e(t) := ρ(γ(t)) eih(t) defines an ω-horizontal lift
2
of γ(t) and h(0) = 0. If Ωω = dω ∈ Ω (P, iR) denotes the curvature 2-form of ω,
then the element eih(2π) of the holonomy group of ω at p0 = ρ(f (1)) determined by
γ
e is given by
 Z 
ih(2π) ∗ ω
(15.123) e = exp − (ρ ◦ f ) Ω .
D

Proof. For γ : [0, 2π] → V , given by γ(t) = f eit , a curve γ e : [0, 2π] → P
with π ◦ γ e = γ and γ e(0) = ρ(γ(0)) has the form γ e(t) = ρ(γ(t)) eih(t) for some
h : [0, 2π] → R with h(0) = 0. The curve γ e is a horizontal lift of γ if and only if
ω(eγ 0 (t)) = 0 for all t ∈ [0, 2π]. Note that
     
e0 (t0 ) = dt
γ d
ρ(γ(t)) eih(t) = dt
d
ρ(γ(t)) eih(t0 ) + dtd
ρ(γ(t0 )) eih(t)

= Reih(t0 ) ∗ (ρ∗ (γ 0 (t0 ))) + (ih0 (t0 ))ρ(γ(t0 )) , and so
 

γ 0 (t0 )) = ω(Reih(t0 ) ∗ (ρ∗ (γ 0 (t0 )))) + ω ih0 (t0 )ρ(γ(t0 ))
ω(e
= adeih(t0 ) ω(ρ∗ (γ 0 (t0 ))) + ih0 (t0 )
= ω(ρ∗ (γ 0 (t0 ))) + ih0 (t0 ) .
Thus, γ γ 0 (t)) = 0) if and only if
e is a horizontal lift of γ (i.e., ω(e
h0 (t) = iω(ρ∗ (γ 0 (t))) = i(ρ∗ ω)(γ 0 (t)) ; i.e.,
Z t Z t
∗ 0
(ρ∗ ω) f∗ dτ
d iτ

h(t) = h(t) − h(0) = i (ρ ω)(γ (τ )) dτ = i e dτ
0 0
Z t Z t
∗ 
(ρ∗ ω) f∗ ieiτ dτ = i (ρ ◦ f ) ω ieiτ dτ.
 
=i
0 0

We then have (15.123), since using Stokes’ Theorem,


Z Z Z
∗ ∗  ∗
(ρ ◦ f ) Ωω = d (ρ ◦ f ) ω = (ρ ◦ f ) ω
D D ∂D
Z 2π

(ρ ◦ f ) ω ieiτ dτ = −ih(2π) .
 
= 
0

2
Remark 15.67. Note that since U (1) is abelian, we have that Ωω ∈ Ω (P, iR)
is right-invariant as well as horizontal, and so there is a unique F ω ∈ Ω2 (M, R)
458 15. GEOMETRIC PRELIMINARIES

such that π ∗ (F ω ) = iΩω . Thus,


Z Z

h(2π) = i (ρ ◦ f ) Ωω = f ∗F ω ,
D D

note that D f ∗ F ω is defined even if f (D) V . Moreover, the element say gγ of


R

the holonomy group of ω at p0 = ρ(f (1)) determined by a horizontal lift γ


e of γ is
also defined even if f (D) V . Thus, it makes sense to ask whether
 Z 
∗ ω
gγ = exp i f F .
D

This is the case, and it can be proven by first establishing a version of Proposition
15.66 for domains with corners and then partitioning D into such domains each of
which is mapped by F into an open set over which P is trivial. Incidentally, if we
define AV ∈ Ω1 (V, R) by ρ∗ (ω) = iAV , then AV does depend on ρ, and it has the
interpretation of being an electromagnetic gauge potential. Then F ω = ρ∗ (iΩω ) =
ρ∗ (idω) = d(iρ∗ (ω)) = −dAV , is the electromagnetic field, regarded as a 2-form.

Let πE : E → M be a Hermitian line bundle over a closed surface M (i.e., M


is compact, without boundary, and dimR M = 2) and let π : P → M denote the
principal U(1)-bundle of unitary frames with a connection 1-form ω. Suppose that
ψ ∈ C ∞ (E) with finite zero set Z := {z ∈ M : ψ(z) = 0} = {z1 , z2 , . . . , zn }. Con-
sider the closed unit disk D = reit : r ∈ [0, 1] , t ∈ R ⊂ C and for k ∈ {1, . . . , m},
let fk : D ,→ M denote the restriction of a smooth embedding of a larger open disk,
such that P is trivial over fj (D). Assume that fj (D) ∩ fk (D) = ∅ for j 6= k. Then
for M1 := M −∪nj=1 fj (D), we have a section ψ0 := ψ/ |ψ| and a section ρ : M1 → P ,
where ρ(x) : C → Ex denotes the frame given by ρ(x)(z) := zψ0 (x) ∈ Ex . Since P
is trivial over fj (D), we have a local section ρj : fj (D) → P . Let gj : U(1) → U(1)
be defined by

(ρ ◦ fj ) eit = ρ fj eit = (ρj ◦ fj ) eit gj eit =: ((ρj ◦ fj ) · gj ) eit .


    

The degree of the zero zj of ψ, denoted deg(ψ; zj ), is defined to be the degree (or
winding number) of gj . There is  a function gej : R → R (unique up to an additive
it ie
gj(t)
multiple of 2π), such that gj e = e and 2π deg(ψ; zj ) = gej (2π) − gej (0). Since
d gj(t) 0
ie
gj(t)
 ie
dt e = e ie
gj (t), we then have
Z 2π Z 2π −1
gej0 (t) dt = −igj eit d
eit dt.

(15.124) 2π deg(ψ; zj ) = dt gj
0 0

Theorem 15.68. As above, let πE : E → M be a Hermitian line bundle over


a compact, orientable 2-manifold M , let ω be a connection on the U(1)-bundle
of P of unitary frames of E, and let ψ ∈ C ∞ (E) have finitely many zeros. For
F ω ∈ Ω2 (M, R) determined by π ∗ (F ω ) = iΩω , we have
Z
1 Xn
(15.125) c1 (E) [M ] = Fω = deg(ψ; zj ) .
2π M j=1
15.8. HOLONOMY 459

i
σ (Ωω ) which is
R
RProof. Recall from Definition 15.61 that c1 (E) [M ] := 2π M 1
1 ω ∗
2π M
F . Using the notation above and A := ρ (−iω), we compute
Z Z Z Z
ω
F = −dA = − A= ρ∗ (iω)
M1 M1 ∂M1 ∂M1
Xn Z Xn Z ∗

=− ρ (iω) = − (ρ ◦ fj ) (iω)
j=1 ∂fj(D) j=1 ∂D
Xn Z

=− ((ρj ◦ fj ) · gj ) (iω)
j=1 ∂D
Xn Z 2π
iω ((ρj ◦ fj ) · gj )∗ ieit

(15.126) =− dt.
j=1 0
We have
((ρj ◦ fj ) · gj )∗ ieit0 = dt d
(ρj ◦ fj ) eit gj eit t=t0
  

d
(ρj ◦ fj ) eit t=t0 gj eit0
 
= dt
   
d
(ρj ◦ fj ) eit0 gj eit0 gj ei(t−t0 )

+ dt
t=t0
it0
  it0 −1
 ∗
gj∗ ieit0

= Rgj(eit0 )∗ (ρj ◦ fj )∗ ie + gj ie .
((ρj ◦fj )·gj )(eit0 )

Since ω is Rg -invariant and ω(B ∗ ) = B for B ∈ u(1) = iR, we then have


−1
iω ((ρj ◦ fj ) · gj )∗ ieit = iω (ρj ◦ fj )∗ ieit0 + igj ieit0 gj∗ ieit0 .
  

Thus, using (15.124) and (15.126),


Z Xn Z Xn

(15.127) Fω = − (ρj ◦ fj ) iω + 2π ord(ψ; zj ) .
M1 j=1 ∂D j=1

For r ∈ (0, 1] , let Dr 


= rD and let fj,r : D → M be given by fj,r (z) := fj (rz).

= r fj∗ reit ieit ,
 
Since ω (ρj ◦ fj )∗ ieit is bounded on D \ {0} and fj,r eit
ie it

we have
Z Z Z
∗ ∗
ρ∗j ω = rfj∗ ρ∗j ω = O(r) .
 
(ρj ◦ fj,r ) ω = fj,r
∂D ∂D ∂Dr

Thus, using fj,r in place of fj in (15.127) and letting r → 0+ , we obtain


Z Xn
F ω = 2π deg(ψ; zj ) . 
M1 j=1

Remark 15.69. In Theorem 15.68 if the image ψ(M ) ⊂ E intersects the image
0(M ) ∼
Pn
= M of the zero section 0 ∈ C ∞ (E) transversally, then j=1 deg(ψ; zj ) is
the intersection number of the surfaces 0(M ) and ψ(M ) in E (i.e., the algebraic
number of signed intersections, where the sign is ±1, depending on whether the
combined orientation of surfaces at an intersection point agrees with that of E).
In particular, let M be a compact, oriented, embedded surface in an oriented 4-
manifold X and let N M denote the normal bundle with the orientation induced by
those of X and M . Then the intersection number of M and the exponential of a
small section of C ∞ (N M ) transverse to the zero section is the self-intersection
number of M in X, and by Theorem 15.68 this is c1 (N M ) [M ].
CHAPTER 16

Gauge Theoretic Instantons

Synopsis. The Yang-Mills Functional. Instantons on Euclidean 4-Space. Lineariza-


tion of the Manifold of Moduli of Self-dual Connections. Manifold Structure for Moduli
of Self-dual Connections.
I In the preceding Chapter 15, we developed the geometric preliminaries that are
required for reading the present chapter. A gentle introduction to the very basics of
this subject can also be found in [312]. More than 30 new textbooks (in English) on
gauge theory have been published during the past 10-15 years. Apparently, that reflects
the intense activity of this discipline in quantum field theory and particle physics. For all
those who want to immerse themselves in this particular area of theoretical physics, a wide
selection of early concepts and results is given in [357]. The awakening of more recent
physics applications, including string theory and anomalies in gauge theory, is described in
the classic textbook [313]. It is written for physicists, but includes rigorous mathematical
definitions and arguments.
In the recent [266, 269] the next 20 years of theory are laid out: Peter Kron-
heimer’s and Tomasz Mrowka’s analysis of moduli spaces of singular instantons (in-
cluding index calculations in terms of representation theory) is impressive. In terms of big
results, [263, 264], are important. Simon Donaldson himself has moved to Calabi-Yau
gauge theory, e.g., [127]. J

1. The Yang-Mills Functional


Notation. Regarding notations, a reader will notice a certain clash of culture
between differential geometers on the one side and topologists and physicists on the
other: Differential geometers are proud of the origin and long history in their field
of concepts which only recently have come to play a role in topology and physics.
They love to stick to traditional notation, contrary to topologists and physicists
who prefer the freedom of making their own notation, see our Translation Table
16.1.
ad-Invariant Inner Products. In order to define the Yang-Mills functional
on the space C(P ) of connections on a principal G-bundle π : P → M , we need
an inner product K on the Lie algebra g of G. This inner product needs to
be invariant under the adjoint action ad : G → GL(g) of G on g (introduced in
Definition 15.4, p.397), namely
!
K(A, B) = K(adg (A) , adg (B)) = K gAg −1 , gBg −1


for all g ∈ G and A, B ∈ g, where we continue to assume that G is a matrix group,


so that adg (A) = gAg −1 .
Remark 16.1. a) When G is O(n) or SO(n), then
g = so(n) = A ∈ GL(n, R) : AT = −A ,


460
16.1. THE YANG-MILLS FUNCTIONAL 461

Table 16.1. Standard topology and gauge theory notation vs. our notation

Standard in Introduced,
Concept Our notation
topology and GT where:
Space of connections Def. 15.3,
A C(P )
(affine config. space) p.396
Def. 15.25,
Gauge group G GA(P )
p.409
Eq. 16.31,
Quotient space B M = C(P )/ GA(P )
p.489
Space of moduli of Eq. 16.32,
M M+
self-dual connections p.489

and we can use the clearly ad-invariant


K(A, B) := − Tr(AB) = Tr AB T .

 Pn
Note that K(A, A) = Tr AAT = i,j=1 A2ij > 0 for A 6= 0.
b) When G is U(n), g = u(n) = {A ∈ GL(n, C) : A∗ = −A}, and one can take
K(A, B) := − Re(Tr(AB)) = Re(Tr(AB ∗ )) .
Pn Pn 2
Note that K(A, A) = Re(Tr(AA∗ )) = i,j=1 Aij Aij = i,j=1 |Aij | > 0 for A 6= 0.
If these are the only cases of interest, the reader may skip the next two paragraphs.
For a compact Lie group G, such an ad-invariant K can be produced via in-
tegration as follows. Select an arbitrary nonzero v0 ∈ Λm (g, R). A volume form
v ∈ Ωm (G, R) is defined at any g 0 ∈ G by
 
vg0 (A1 , . . . , Am ) := v0 Rg−1 −1
0 ∗ (A1 ) , . . . , Rg 0 ∗ (Am ) .

Note that v is right-invariant in the sense that Rg∗ v = v. Indeed, noting that

Rg0 g = Rg ◦ Rg0 =⇒ Rg0 g∗ = Rg∗ ◦ Rg0 ∗ =⇒ Rg−1 −1 −1


0 g∗ = Rg 0 ∗ ◦ Rg∗ ,

we have
Rg∗ v

g0
(A1 , . . . , Am ) = vg0 g (Rg∗ A1 , . . . , Rg∗ Am )
 
= v0 Rg−1 −1
0 g∗ Rg∗ A1 , . . . , Rg 0 g∗ Rg∗ Am
 
= v0 Rg−1 −1 −1 −1
0 ∗ Rg∗ Rg∗ (A1 ) , . . . , Rg 0 ∗ Rg∗ Rg∗ (Am )
 
= v0 Rg−1 −1
0 ∗ (A1 ) , . . . , Rg 0 ∗ (Am ) = vg0 (A1 , . . . , Am ) .

For A, B ∈ g and an arbitrary inner product K0 on g, consider the form α(A,B) ∈


Ωm (G, R), defined by
α(A,B) (g 0 ) := K0 (adg0 (A) , adg0 (B)) vg0 .
Then let Z
K(A, B) := α(A,B) .
G
462 16. GAUGE THEORETIC INSTANTONS

We have α(adg A,adg B) = Rg∗ α(A,B) , since




α(adg A,adg B) (g1 ) = K0 (adg1 (adg A) , adg1 (adg B)) vg1


= K0 (adg1 g A, adg1 g (B)) Rg∗ (v)(g1 )
= Rg∗ α(A,B) (g1 ) .


The ad-invariance of K then follows, since


Z Z
Rg∗ α(A,B)

K(adg A, adg B) = α(adg A,adg B) =
ZG G

= α(A,B) = K(A, B) .
G
There are other ways of producing ad-invariant inner products K on g. Re-
call from Definition 15.4 that ad : g → End(g) denotes the derivative of ad at the
identity, and is given by ad(A)(B) = [A, B]. The Killing form β is a symmetric
bilinear form on g defined for A, B ∈ g by
β(A, B) := Tr(ad(A) ◦ ad(B)) .
−1

Note that ad gAg = adg ◦ ad(A) ◦ adg−1 , since
ad gAg −1 (B) = gAg −1 , B = gAg −1 B − BgAg −1
  

= gA g −1 Bg g −1 − g g −1 Bg Ag −1 = adg ad(A) adg−1 (B) .


  

Then β is invariant under the adjoint action, since


β gAg −1 , gBg −1 = Tr ad gAg −1 ◦ ad gBg −1
  

= Tr adg ◦ ad(A) ◦ adg−1 ◦ adg ◦ ad(B) ◦ adg−1

= Tr adg ◦ ad(A) ◦ ad(B) ◦ adg−1 = Tr(ad(A) ◦ ad(B)) .
Of course, β need not be definite; e.g., β = 0 for abelian groups such as U (1).
However, for compact, semi-simple G, Cartan’s Criterion guarantees that −β is
positive definite, and hence would serve as an ad-invariant inner product K. If
ad : G → GL(g) is irreducible, then any two ad-invariant inner products, say K1
and K2 , on g must agree up to a multiplicative constant, since any eigenspace of
K1 relative to K2 would be invariant.

Yang-Mills Functional and Equation. In what follows, we will assume that


G is compact, connected and semi-simple, in which case we can and do take K on
g to be −β.
Definition 16.2. Let π : P → M be a principal G-bundle, where M has a
Riemannian metric h, g has an ad-invariant inner product K, and M and G are
compact. Let C(P ) denote the set of connection 1-forms on P . The Yang-Mills
functional
YM : C(P ) → R+ := [0, ∞)
is defined by Z
ω 2 2
YM(ω) := 1
2 kΩ k = 1
2 |Ωω | vh ,
M
where
2
Ωω = dω + ω ∧ ω = dω + 1
2 [ω, ω] ∈ Ω (P, g) ∼
= Ω2 (M, P ×G g)
16.1. THE YANG-MILLS FUNCTIONAL 463

2
denotes the curvature of the connection ω ∈ C(P ), and |Ωω | = hΩω , Ωω i in terms
of the pairing (15.15) on p. 404.
Recall from Remark 15.12, p. 401, that C(P ) has the structure of an affine
1
space based on the vector space Ω (P, g) which can then be regarded as a formal
1
tangent space of C(P ) at any ω ∈ C(P ). For τ ∈ Ω (P, g) and t ∈ R, note that
ωt := ω + tτ ∈ C(P ). At t = 0,
d ωt d
dωt + 21 [ωt , ωt ] = dτ + 12 ([τ, ω] + [ω, τ ])

dt Ω = dt
= dτ + [ω, τ ] = dτ + ad(ω) ∧ τ = Dω τ.
Thus at t = 0, we have
ω ω
d
dt YM(ωt ) = d 1
dt 2 (Ω , Ω ) = (Dω τ, Ωω ) = (τ, δ ω Ωω ) ,
where
n+1
δ ω = (−1) ∗ Dω ∗ : Ω2 (M, P ×G g) → Ω1 (M, P ×G g)
denotes the formal adjoint of
Dω : Ω1 (M, P ×G g) → Ω2 (M, P ×G g) ,
namely δ ω is the covariant codifferential (see Proposition 15.19, p. 406). Since all
inner products involved are positive definite, all the directional derivatives
d
dt YM(ω + tτ ) t=0

will be zero precisely when


δ ω Ωω = 0.
This is called the (source-free) Yang-Mills equation. Formally, it characterizes
the critical points (connections) for the Yang-Mills functional.
We mention that YM : C(P ) → R+ is invariant under the action on C(P ) of the
group GA(P ) of gauge transformations (introduced
∗ in Definition 15.25 of Section
15.4), since for F ∈ GA(P ) , F · ω = F −1 ω yields
∗
ΩF ·ω = d(F · ω) + (F · ω) ∧(F · ω) = F −1 (dω + ω ∧ ω)
∗
= F −1 Ωω = F · Ωω ,
2 2
and then |F · Ωω | = |Ωω | by Corollary 15.31.

Extremal Properties of Self-Dual and Anti-Self-Dual Curvatures and


Chern Character. If dim M = 4 and M is oriented, then we have the Hodge ∗
operator on Ω2 (M, P ×G g), and it makes sense to speak of Ωω as being self-dual
(∗Ωω = Ωω ) or anti-self-dual (∗Ωω = −Ωω ). For these, the Yang-Mills equation
always holds, since
δ ω Ωω = − ∗ Dω ∗ Ωω = − ∗ Dω ∗ ±Ωω = ∓ ∗(Dω Ωω ) = 0,
by the Bianchi identity Dω Ωω = 0 (see Proposition 15.17, p. 403).
Proposition 16.3. In the notation of Definition 16.2, if Ωω is self-dual or
anti-self-dual, then ω is an absolute minimum of the functional YM.
464 16. GAUGE THEORETIC INSTANTONS

Proof. Let gC = C ⊗ g denote the complexification of g. The inner product


K on g extends to a Hermitian inner product KC on gC via
KC (z ⊗ A, w ⊗ B) = zwKC (A, B) for A, B ∈ g and z, w ∈ C.
Then the orthogonal (relative to K) representation ad : G → SO(g) extends to
adC : G → SU(gC ), and we may form the associated complex vector bundle E :=
P ×G gC which has a Hermitian structure inherited from KC . Let U (E) → M
denote the bundle of unitary frames. Each p ∈ P gives rise to a unitary mapping
pb: g →Eπ(p) , simply via
pb(A) = [p, A] .
In this way we have an embedding P ,→ U (E) given by p 7→ pb, which makes P a
principal subbundle of U (E). A connection ω on P uniquely extends to a connection
ωE on U (E) in such a way that the horizontal subspaces of ωE at points of P are
those of ω. The second Chern character class of E is (see 15.99, p. 446)
−1
ch(E)2 = [Tr(ΩωE ∧ ΩωE )] ∈ H 4 (M ; Q) .
8π 2
k
Note that ΩωE |P ∈ Ω (P, u(gC )) ∼
= Ω2 (M, End(E)) is related to Ωω via
ΩωE |P = ad(Ωω ) .
Hence, for Xi ∈ Tp P,
Tr(ΩωE (X1 , X2 ) ◦ ΩωE (X3 , X4 )) = Tr(ad(Ωω (X1 , X2 )) ◦ ad(Ωω (X1 , X2 )))
= −K(Ωω (X1 , X2 ) , Ωω (X1 , X2 )) .
Thus, we have
Tr(ΩωE ∧ ΩωE )(X1 , X2 , X3 , X4 )
1 X σ
= (−1) Tr(ΩωE (Xσ1 , Xσ2 ) ◦ ΩωE (Xσ3 , Xσ4 ))
2!2! σ
X σ
= − 41 (−1) K(Ωω (Xσ1 , Xσ2 ) , Ωω (Xσ3 , Xσ4 ))
σ
= −K(Ωω ∧ Ωω )(X1 , X2 , X3 , X4 )
= −K(Ωω ∧ ∗ ∗ Ωω )(X1 , X2 , X3 , X4 )
= −K(Ωω , ∗Ωω ) π ∗ (νh )(X1 , X2 , X3 , X4 ) .
Hence,
−1
Z Z
1
ch(E)2 [M ] = Tr(ΩωE ∧ ΩωE ) = hΩω , ∗Ωω i νh .
8π 2 M 8π 2 M

Incidentally, since adC : G → SU(gC ) we have c1 (E) = 0, in which case


 
2
ch(E)2 = 12 c1 (E) − 2c2 (E) = −c2 (E) .

Writing Ωω = Ωω+ + Ωω− , we then have


Z Z Z
2 ω ω ω+ 2 2
8π ch(E)2 [M ] = hΩ , ∗Ω i νh = Ω νh − Ωω− νh .
M M M
16.1. THE YANG-MILLS FUNCTIONAL 465

Thus,
Z Z Z
2 2 2
YM(ω) = 1
2 |Ωω | vh = 1
2 Ωω+ νh + 1
2 Ωω− νh
M M M
Z Z 
2 2
= 1
2 Ωω+ νh + 1
2 Ωω+ νh − 8π 2 ch(E)2 [M ]
M M
Z
ω+ 2
(16.1) = Ω νh − 4π 2 ch(E)2 [M ] ,
M
and similarly
Z Z
ω 2 2
(16.2) YM(ω) = 1
2 |Ω | vh = Ωω− νh + 4π 2 ch(E)2 [M ] .
M M
Thus, if Ωω self-dual or anti-self-dual, then ω will furnish an absolute minimum for
YM. 
Taking the difference of (16.1) and (16.2) yields
Corollary 16.4. In the notation of the proof of Proposition 16.3,
Z Z
2 ω+ 2 2
(16.3) 8π ch(E)2 [M ] = Ω νh − Ωω− νh .
M M
Thus if ch(E)2 [M ] > 0, then there is no connection ω ∈ C(P ) with anti-self-dual
curvature Ωω (i.e., Ωω+ = 0). If ch(E)2 [M ] < 0,then there is no ω ∈ C(P ) with
self-dual curvature Ωω (i.e., Ωω− = 0). If ch(E)2 [M ] = 0, then Ωω+ = 0 ⇒ Ωω =
0, and Ωω− = 0 ⇒ Ωω = 0 (i.e., all ω ∈ C(P ) with self-dual or anti-self-dual
curvature are flat).
By a convenient abuse of terminology, connections ω for which Ωω is self-dual
(resp. anti-self-dual) are known as self-dual (resp. anti-self-dual) connections.
Such connections are also known as instantons, particularly when the base M
is S 4 . This deserves some explanation. We saw in Section 14.14.2 (p. 391-393)
that, generally speaking, an instanton is a solution of a Euclidean action principle,
which minimizes the Euclidean action among all suitable paths joining two minima
for a potential function. When the general concept is applied to the configuration
space of connections (gauge potentials) modulo gauge transformations, an instanton
1 ω 2
is a minimum for the Yang-Mills functional 2 R4 |Ω | vh for Euclidean R4 with
R

the standard metric h, where certain asymptotic conditions at ∞ are imposed on


ω which make the integral converge. Moreover, the functional is defined on the
quotient space of connections modulo the group of gauge transformations defined
on R4 which tend to the identity at infinity. Recall that the star operator on 2-forms
in dimension 4 is conformally invariant, so that self-duality and anti-self-duality are
preserved under conformal changes of metric. Moreover, the invariance of ∗ implies
that the functional YM is invariant under conformal changes, since
Z Z
ω 2
1
2 |Ω | v h = 1
2 K(Ωω ∧ ∗Ωω ) .
R4 R4
2
Alternatively, note that replacing h by λh, |Ωω | and vh acquire factors of λ−2 and
λ2 respectively. Since R4 is conformally equivalent to S 4 − {∞} via stereographic
projection, we can work over S 4 − {∞}. It turns out that via a crucial theorem
in [418], the physically reasonable asymptotic conditions imposed on ω and the
gauge transformations at ∞ are precisely those that enable one to extend (over
466 16. GAUGE THEORETIC INSTANTONS

the point ∞ ∈ S 4 ) a self-dual connection ω for the bundle R4 × G → R4 to a


self-dual connection ω for a principal G-bundle π : P → S 4 . The extension of the
bundle R4 × G → R4 is generally nontrivial, depending  on the homotopy class of
gauge transformations over the slice S 3 = R3 ∪ {∞} × {0}. For G = SU(2) this

homotopy class is determined by the degree of φ0 : S 3 → SU(2) −→ S 3 , discussed
at the end of Section 14.2. Of course, one can work over compact, orientable
Riemannian 4-manifolds M other than S 4 .

2. Instantons on Euclidean 4-Space


In the sections that follow, we will construct manifolds of gauge-equivalence
classes of self-dual connections (instantons) and compute their dimensions. How-
ever, these results require the existence of such connections, or else they might
just be sophisticated statements about the empty set. Here our primary goal is to
produce self-dual connections for principal SU(2)-bundles
 4 P → S 4 with arbitrary
2
nonpositive Chern number −k := c2 P ×SU(2) C [S ]. Our construction will be
motivated from the standpoint of the Riemannian geometry of conformally related
metrics on S 4 or R4 , but first we give a brief history.

A Brief History of Instantons in Physics and Mathematics. The first


such instanton that came from the physics community was the BPST instanton,
named after A. A. Belavin, A. M. Polyakov, A. S. Schwarz and Y. S. Tyup-
kin (see [52]). The standard BPST instanton was given on R4 , in the sense that it
was presented as an su(2)-valued 1-form on R4 , which can be regarded as the pull-
back of a connection 1-form, say ω0 , on the trivial principal SU(2)-bundle R4 ×SU(2)
by the global section x 7→ (x, I). It is not difficult to extend ω0 to a connection
1-form ω on nontrivial principal SU(2)-bundle P → S 4 , namely the quaternionic
Hopf bundle S 7 → S 4 . Then ω turns out to be the well-known universal connection
on this bundle. Conformal transformations of S 4 act on the space of (anti-)self-dual
connections on S 4 . While the conformal transformations which are isometries of
S 4 preserve ω, the 5-parameter family of boosts acts effectively on ω to produce
a 5-parameter family of instantons. Here, a boost is a conformal transformation
of a sphere which contracts toward one point of the sphere and dilates about the
antipode. The standard example is the function z 7→ αz on the extended complex
plane (Riemann sphere). For S 4 , four parameters suffice to locate the pair of an-
tipodal points and the dilation factor is the fifth parameter. The Chern number
of the bundle S 7 → S 4 is −1, and so the 5-parameter BPST family of instantons
pertains to this case. Instantons for arbitrary negative Chern number −k were
exhibited in [446], and 5k-parameter families of such instantons were produced
by t’Hooft (unpublished), in [445], and in [113]. In [232] R. Jackiw, C. Nohl
and C. Rebbi used the fact that these families could be augmented to conformally
invariant families, to enlarge the number of parameters to 5k + 4. In [38], M. F.
Atiyah, N. J. Hitchin and I. M. Singer applied the index theorem to find that
the maximal number of effective parameters is 8k − 3 (k ≥ 1), a result which was
also derived in [372]. Not long after, the analysis was applied to arbitrary princi-
pal G-bundles over S 4 (G simple and compact); see [39] and [57]. The problem of
actually constructing the most general 8k − 3 family of solutions was solved in [48]
using the techniques originating with the twistor theory of Roger Penrose and
some algebraic geometry. A construction involving “only” linear algebra was finally
16.2. INSTANTONS ON EUCLIDEAN 4-SPACE 467

developed by M.F. Atiyah, V.G. Drinfeld, N.J. Hitchin and Y.I. Manin in
[36]. We will describe this construction later in this section. The interested reader
may augment this brief history by consulting the excellent surveys [136] and [283].
The Jackiw-Nohl-Rebbi (5k + 4)-Parameter Family of Instantons. In
what follows, we will derive the Jackiw-Nohl-Rebbi (5k + 4)-parameter family of
instantons in a very natural way from the viewpoint of Riemannian  geometry. Re-
call from (15.97) that corresponding to the decomposition Λ2 R4 ∼ = Λ +
⊕ Λ −
into
self-dual and anti-self-dual 2-forms, there is a Lie algebra decomposition so(4) ∼=
so+ ⊕ so− which we can make explicit as follows. Under the index-lowering isomor-
 ∼
phism Λ2 R4 −→ so(4), we have e2 ∧e3 ±e1 ∧e4 , e3 ∧e1 ±e2 ∧e4 , and e1 ∧e2 ±e3 ∧e4
corresponding to
     
0 0 0 ±1 0 0 −1 0 0 1 0 0
 0 0 1 0    0 0
 0 ±1   ,  −1 0 0
 0 
(16.4a)   0 −1 0 0  ,  1 0
,
0 0   0 0 0 ±1 
∓1 0 0 0 0 ∓1 0 0 0 0 ∓1 0
respectively. Choosing + in ± (and − in ∓), these are the self-dual ’t Hooft
matrices η1 , η2 and η3 , which form a basis of so+ . Choosing − in ± (and + in ∓),
we have the anti-self-dual ’t Hooft matrices η 1 , η 2 and η 3 which form a basis
of so− . One readily verifies that [ηa , ηb ] = −2εabc ηc and [η a , η b ] = −2εabc η c . For
the Hermitian Pauli matrices
     
0 1 0 −i 1 0
σ1 := , σ2 := , σ3 := ,
1 0 i 0 0 −1
we have iσa ∈ su(2) with [iσa , iσb ] = −2εabc iσc . Thus, ηa 7→ iσa defines an

isomorphism so+ −→ su(2), while η a 7→ iσa yields so− ∼ = su(2).
Recall from Definition 15.40 that the Levi-Civita connection θ for a Riemannian
4-manifold M is an so(4)-valued connection 1-form on the bundle of F M of ortho-
∼ ∼
normal frames. Relative to the isomorphisms so(4) −→ so+ ⊕so− −→ su(2)⊕su(2),
+ −
the connection θ splits into two su(2)-valued forms θ and θ . Suppose that M is
simply R4 with some metric tensor h. Then there is a global section σ : R4 → F R4
(e.g., apply the Gram-Schmidt procedure to the standard coordinate fields ∂i :=
∂ + −
∂xi , i = 1, · · · , n). We may pull back θ and θ to obtain the su(2)-valued forms
± ∗ ± 4
A := σ θ on R . Let
 ±
F ± := σ ∗ Ωθ = dA± + 12 A± , A±
 

denote the field strengths. It is not difficult to find conditions on the metric h such
that ∗F + = F + or ∗F − = F − . Indeed,  F + ⊕ F − is the decomposition of the
2
b ∈ End Ω (M ) according to the decomposition of its values
curvature operator R
in Ω2 (M ) = Ω+ (M ) ⊕ Ω− (M ). Writing
   
A B I 0
Rb= , ∗ = ,
BT C 0 −I
as in (15.76), we have F + = A B while F − = B T C . Then
   
 
+
  I 0  
∗F = A B = A −B , and
0 −I
 

 T  I 0
= B T −C .
 
(16.5) ∗F = B C
0 −I
468 16. GAUGE THEORETIC INSTANTONS

Hence,
F − is self-dual ⇐⇒ C = 0,
F + is anti-self-dual ⇐⇒ A = 0, and
F + is self-dual ⇐⇒ B = 0 ⇐⇒ F − is anti-self-dual.
From (15.79) we have
s(R) e and C = s(R) + C
A= +A e
12 12
and the Weyl tensor (as a curvature type operator) is
" #
Ae 0
W = e .
0 C
Thus, if W = 0 (i.e., h is conformally flat) and s(R) = 0, then A = 0 and C = 0.
Hence, in this case F − is self-dual and F + is anti-self-dual. Recall that B = 0 if
the traceless Ricci tensor Rij − 41 s(R) h = 0, in which case h is an Einstein metric.
It is not easy to produce Einstein metrics, but it is trivial to produce conformally
flat metrics, namely take
 2 
2
(16.6) h = f 2 ds2 := f 2 dx1 + · · ·(dxn ) for 0 < f ∈ C ∞ (Rn ) .

Levi-Civita Connection and Curvature Tensor for Flat Metric. The


next proposition provides the Levi-Civita connection and the curvature tensor of
such h for arbitrary n. We also need s(R) = 0 to conclude that F − is self-dual and
F + is anti-self-dual. In the case n = 4, this proposition says that
s(R) = 0 ⇐⇒ ∆f = 0
(i.e., f is harmonic).
Proposition 16.5. Let F Rn → Rn denote the frame bundle of Rn with the
metric h = f 2 ds2 , and let τ : Rn → F Rn denote the global section
1
τ (x) = (∂1 , · · · , ∂n ) ,
f (x)
where ∂i = ∂/∂xi are the standard coordinate vector fields. If θ ∈ C(F Rn ) ⊆
Ω1 (F Rn , so(4)) denotes the Levi-Civita connection of h, then τ ∗ θ ∈ Ω1 (Rn , so(4))
is given by
i 1
(τ ∗ θ) j = ∂j f dxi − ∂i f dxj .

f
2
Let Ωθ ∈ Ω (Rn , so(4)) denote the curvature of θ. With
h X h X
τ ∗ Ωθ i = 12 τ ∗ Ωθ ijk dxj ∧ dxk = 21 Rh ijk dxj ∧ dxk ,

the curvature tensor R ∈ C ∞ Rn , T 0,4 Rn with Rhijk := f 2 Rh ijk is given in terms




of the K-N product ∨ (see (15.65)) by


2
R = −2f ∇2 f ∨ I + 4(df ⊗ df ) ∨ I − |df | I ∨ I
 
2
(16.7) = −2f ∇2 f + 4(df ⊗ df ) − |df | I ∨ I,
16.2. INSTANTONS ON EUCLIDEAN 4-SPACE 469

 2 Pn 2
where ∇2 f ij = ∂i ∂j f, (df ⊗ df )ij = ∂i f ∂j f, and |df | = i=1 (∂i f ) . The scalar
curvature s(R) of h is
 
2
s(R) = f −4 (n − 1) −2f ∆f + (4 − n) |df | ,

which is −6f −3 ∆f when n = 4.



Proof. Regarding τ (x) as an isometry Rn −→ Tx Rn we have τ (x)(ei ) =
−1
f (x) (∂i )x where ei denotes the i-th standard unit vector in Rn . Recall that the
1
canonical 1-form ϕ ∈ Ω (F Rn , Rn ) is given by ϕu (X) = u−1 π∗ (X). Thus,
       
−1 −1 −1 −1
(τ ∗ ϕ)x f (x) (∂i )x = ϕ τ∗ f (x) (∂i )x = τ (x) π∗ τ∗ f (x) (∂i )x
 
−1 −1
= τ (x) f (x) (∂i )x = ei .

Hence,

φ1 , . . . , φn := φ := τ ∗ ϕ ∈ Ω1 (Rn , Rn )


is the dual coframe field for f −1 ∂1 , . . . , f −1 ∂n in the sense that




   j
−1 j
φj f −1 ∂i = (τ ∗ ϕ) f (x) (∂i )x = (ei ) = δij .

In other words, φi = f dxi . The Levi-Civita connection 1-form θ ∈ C(F Rn ) ⊆


Ω1 (F Rn , so(n)) on F Rn is uniquely determined by the condition that the torsion
vanishes; i.e., 0 = Dθ ϕ = dϕ + θ ∧ ϕ. We have

0 = τ ∗ Dθ ϕ = τ ∗ (dϕ + θ ∧ ϕ) = d(τ ∗ ϕ) + τ ∗ (θ) ∧(τ ∗ ϕ) = dφ + τ ∗ (θ) ∧ φ.




Writing τ ∗ (θ) as a skew-symmetric matrix (θij ) of 1-forms θij , this becomes dφi =
Pn
− j=1 θij ∧ φj . On the other hand,

n
X
dφi = d f dxi = df ∧ dxi = ∂j f dxj ∧ dxi

j=1
n
X n
X
=− ∂j f dxi ∧ dxj = − f −1 ∂j f dxi ∧ φj
j=1 j=1
n
X
f −1 ∂j f dxi − ∂i f dxj ∧ φj .

=−
j=1

Thus, by the uniqueness of θ, we have

θij = f −1 ∂j f dxi − ∂i f dxj .



470 16. GAUGE THEORETIC INSTANTONS

2
Since the curvature Ωθ ∈ Ω (F Rn , so(n)) is given by Ωθ = dθ + θ ∧ θ,
n
h X
τ ∗ Ωθ i
= dθhi + θhp ∧ θpi
p=1
−1
∂i f dx − ∂h f dxi
h

=d f
n
X
+ f −2 ∂p f dxh − ∂h f dxp ∧ ∂i f dxp − ∂p f dxi
 
p=1
−1
∧ ∂i f dxh − ∂h f dxi + f −1 d ∂i f dxh − ∂h f dxi
  
=d f
n
X
+ f −2 ∂p f dxh − ∂h f dxp ∧ ∂i f dxp − ∂p f dxi
 
p=1
n
X
= −f −2 ∂q f ∂i f dxq ∧ dxh − ∂q f ∂h f dxq ∧ dxi

q=1
n
X
+ f −1 ∂q ∂i f dxq ∧ dxh − ∂q ∂h f dxq ∧ dxi

q=1
Xn  
2
+ f −2 ∂p f ∂i f dxh ∧ dxp + ∂h f ∂p f dxp ∧ dxi − (∂p f ) dxh ∧ dxi
p=1
n
X
= f −1 ∂q ∂i f dxq ∧ dxh − ∂q ∂h f dxq ∧ dxi

q=1
Xn  
−2 2
+f 2∂p f ∂i f dxh ∧ dxp + 2∂h f ∂p f dxp ∧ dxi − (∂p f ) dxh ∧ dxi .
p=1

Thus, using the fact dxq ∧ dxh (∂j , ∂k ) = δjq δkh − δkq δjh ,


h
Rhijk = f 2 Rhijk = τ ∗ Ωθ i
(∂j , ∂k )
n
X
∂q ∂i f δjq δkh − δkq δjh − ∂q ∂h f δjq δki − δkq δji
 
=f
q=1
n
X
2∂p f ∂i f δjh δkp − δkh δjp + 2∂h f ∂p f δjp δki − δkp δji
 
+
p=1
Xn
2
δjh δki − δkh δji

− (∂p f )
p=1

= f ∂j ∂i f δkh − ∂k ∂i f δjh − ∂j ∂h f δki − ∂k ∂h f δji




n
X
∂k f ∂i f δjh − ∂j f ∂i f δkh + ∂h f ∂j f δki − ∂h f ∂k f δji
 
+2
p=1
n
X 2
− (∂p f ) ) δjh δki − δkh δji .

+(
p=1
16.2. INSTANTONS ON EUCLIDEAN 4-SPACE 471

In terms of the Kulkarni-Nomizu product (see (15.65))


(Q ∨ I)hijk = 21 (Qik δhj − Qhk δij + Qhj δik − Qij δhk ) ,
we can write
2
R = −2f ∇2 f ∨ I + 4(df ⊗ df ) ∨ I − |df | I ∨ I
 
2
(16.8) = −2f ∇2 f + 4(df ⊗ df ) − |df | I ∨ I,
 2 Pn 2
where ∇2 f ij = ∂i ∂j f, (df ⊗ df )ij = ∂i f ∂j f, and |df | = i=1 (∂i f ) . The scalar
curvature of h is then
 
2
s(R) = f −4 (n − 1) Tr −2f ∇2 f + 4(df ⊗ df ) − |df | I
 
2
= f −4 (n − 1) −2f ∆f + (4 − n) |df | ,

where the factor of f −4 comes from raising two indices using h (i.e., s(R) = Rhi hi ).

− ∗ − 1 4


We can express
∗ θ−
 − 4
 A := τ θ ∈ Ω R , su(2) and field strength
the connection
F := τ Ω ∈ Ω R , su(2) using  the ’t Hooft and Pauli matrices as follows.
The projection of τ ∗ θ ∈ Ω1 R4 , so(4) onto Ω1 R4 , so− is
 
3
X X4
ij
1
4
 (τ ∗ θ) (η a )ij  η a ,
a=1 i,j=1
P4
where the 14 comes from the fact that i,j=1 (η a )ij (η a )ij = 4 for each a. Under the

isomorphism so− −→ su(2) given by η a 7→ iσa ,
 
3 4
1 X X kj
A− =  (τ ∗ θ) (η a )kj  iσa
4 a=1
k,j=1
 
3 4
1 X X
f −1 ∂j f dxk − ∂k f dxj (η a )kj  iσa

= 
4 a=1
k,j=1
3 4
1X X
(16.9) = (η a )kj ∂j (log f ) dxk (iσa ) .
2 a=1
k,j=1

4

There is an alternate
 expression in terms of vector notation. First, for u, u :=
1 2 3 4
u , u , u , u , we have (using the Einstein summation convention)

uk v j η akj σa = uk v j εabc δkb δjc − δka δj4 − δja δk4 σa




= εabc δkb δjc uk v j − δka δj4 uk v j − δja δk4 uk v j σa




= εabc ub v c − v 4 ua − v a u4 σa


= u × v + u4 v − v 4 u · σ,

(16.10)
where
w · σ := w1 σ1 + w2 σ2 + w3 σ3 .
472 16. GAUGE THEORETIC INSTANTONS

Basic Facts about Quaternions. Readers who have worked with the algebra
H of quaternions
u = u4 + u1 i + u2 j + u3 k,
where i2 = j2 = k2 = −1, ij = −ji = k, etc., will sense their involvement in (16.10).
As ensuing developments are more easily expressed with quaternions, we recall some
basic facts. By definition

Im(u) : = u1 i + u2 j + u3 k = u·(i, j, k) and


u : = u4 − u1 i − u2 j − u3 k = u4 − Im(u) = u4 − u·(i, j, k) .

One computes that the quaternion product uv of u ∈ H with v ∈ H is given by

uv = u · v − u × v + u4 v − v 4 u ·(i, j, k) ,


where u · v denotes the ordinary dot product of u and v as vectors in R4 . Note that
uu = u · u and

vu = v · u − v × u + v 4 u − u4 v ·(i, j, k)


= u · v + u × v + u4 v − v 4 u ·(i, j, k) = uv.



There is an algebra isomorphism H −→ RI + isu(2) given by

u4 + u1 i + u2 j + u3 k ←→ u4 I − u1 iσ1 − u2 iσ2 − u2 iσ3 = u4 I − u·iσ.

Under this isomorphism, we have

uk v j η akj (iσ)a = u × v + u4 v − v 4 u · iσ


←→ − u × v + u4 v − v 4 u ·(i, j, k) = Im(uv) .


Hence, we obtain
3 4
1X X
A− = (η a )kj ∂j (log f ) dxk (iσa )
2 a=1
k,j=1
 
(16.11) ←→ 12 Im dx∂(log f ) = − 12 Im(∂(log f ) dx) ,

where the quaternion differential dx is

dx = dx4 + dx1 i + dx2 j + dx3 k,

and the quaternion-valued function ∂(log f ) is given by

∂(log f ) := ∂4 (log f ) + ∇(log f ) ·(i, j, k)


:= ∂4 (log f ) + ∂1 (log f ) i + ∂2 (log f ) j + ∂3 (log f ) k.

Alternatively, in nonquaternionic notation with u = dx = dx, dx4 and v =
∂ log f = (∇(log f ) , ∂4 (log f )), we have
i
A− = dx × ∇(log f ) + dx4 ∇(log f ) − ∂4 (log f ) dx · σ,

(16.12)
2
16.2. INSTANTONS ON EUCLIDEAN 4-SPACE 473

which is considerably less tidy than (16.11). Before considering the field strength
F − , we mention that a direct computation yields
− 12 dx ∧ dx = dx2 ∧ dx3 + dx1 ∧ dx4 i 


+ dx3 ∧ dx1 + dx2 ∧ dx4  j


+ dx1 ∧ dx2 + dx3 ∧ dx4 k and
(16.13)
− 12 dx ∧ dx = dx2 ∧ dx3 − dx1 ∧ dx4 i 


+ dx3 ∧ dx1 − dx2 ∧ dx4  j


+ dx1 ∧ dx2 − dx3 ∧ dx4 k.
Thus, dx ∧ dx is self-dual and dx ∧ dx is anti-self-dual.

Computing the Field Strength F − . According to (16.8),


 
3 4
− i X X
Fjk =  (η a )hi (Pf ∨ I)hijk  σa ,
4 a=1
h,i=1

where
2
(16.14) Pf := −2f ∇2 f + 4(df ⊗ df ) − |df | I.
We now consider some obvious choices for the harmonic dilation factor f . Of course
for f = 1, we obtain the standard flat metric ds2 with R = 0, A± = 0 and F ± = 0.
For f (x) = λ2 r−2 (where 0 < λ ∈ R and r := |x| > 0) one also finds that R = 0,
either by direct computation of Pf , or by verifying that f 2 ds2 is the pull-back of
ds2 under the inversion x 7→ λ2 r−2 x of R4 in the sphere r = λ. If we add them,
taking
(16.15) f = 1 + λ2 r−2 = 1 + f0 (x) ,
then
2
Pf = −2(1 + f0 ) ∇2 f0 + 4(df0 ⊗ df0 ) − |df0 | I
= −2∇2 f0 = −4λ2 r−6 4x ⊗ x − r2 I ,


since
∇2 f0 = ∂j ∂i f0 = −2λ2 ∂j r−4 xi = −2λ2 ∂j r−4 xi − 2λ2 r−4 δij
  
ij
−3
= 4λ2 r2 ∂j r2 xi − 2λ2 r−4 δij = 2λ2 r−6 4xj xi − r2 δij .
 
(16.16)

Thus, for f = 1 + λ2 r−2 , we have a nonzero field strength. Moreover, in this case
 λ2 ∂i r−2

2 −2
∂i (log f ) = ∂i log 1 + λ r =
1 + λ2 r−2
2 −4 i 2
−2λ r x λ xi
= = −2 ,
1 + λ2 r−2 r2 r2 + λ2
or in quaternion notation,
λ2 x λ2 x
∂(log f ) = −2 2 2 2
and ∂(log f ) = −2 .
r r +λ r r + λ2
2 2
474 16. GAUGE THEORETIC INSTANTONS

Hence,

λ2
  
  x
A− (x) ←→ 12 Im dx∂(log f ) = 12 Im dx −2 2 2
r r + λ2
  2   2 
λ x λ xdx
= −Im dx 2 2 = Im
r r + λ2 r2 r2 + λ2
λ2 1
= 2 2 Im(xdx) ∈ H,
r r + λ2
or alternatively,

λ2 1
A− (x) = − (dx) ×x + dx4 x − x4 dx · iσ

2 2
r r +λ 2

λ2 1
x×dx + x4 dx − dx4 x · iσ.

= 2 2 2
r r +λ

Note that f = 1 + λ2 r−2 and A− are singular at x = 0. However, we will find that a
gauge transformation can be applied to A− to yield a local connection form which
extends smoothly over x = 0. For now, we will compute the value YM(A− ) of the
Yang-Mills functional at A− .

Proposition 16.6. We have


Z
− 2
1
F− νh = 16π 2 .

YM A = 2 h
R4

Proof. To avoid confusion in what follows, we denote norms computed with


2
h = 1 + λ2 r−2 ds2 by | · |h and norms computed with ds2 by | · |e . Note that due
to conformal invariance of the YM functional,
Z Z
− 2 2
F h νh = F − e d4 x.
R4 R4

As stated earlier, we are also using the metric K = −β (minus the Killing form on
su(2)). To compute this, we use
3
X
ad(iσb )(iσc ) = [iσb , iσc ] = −2 εbcd iσd and
d=1
3
X 3
X
ad(iσa ) ad(iσb )(iσc ) = −2 εdbc ad(iσa )(iσd ) = −4 εdbc εdae (iσe )
d=1 d,e=1

to deduce that
3
X
K(iσa , iσb ) = − Tr(ad(iσa ) ad(iσb )) = 4 εdbc εdac = 8δab .
c,d=1


The inner product on so− under the isomorphism so− −→ su(2), given by η a 7→ iσa ,
ij
is then twice the contraction inner product (η a )ij (η b ) = 4δab . Since A = C = 0 in
16.2. INSTANTONS ON EUCLIDEAN 4-SPACE 475

(16.5), F + and F − have equal magnitude. Thus, we find


4 4
2 X X
F− e 1
Fij− , Fij− 1
K Fij− , Fij− + K Fij+ , Fij+
  
= 2 K = 4
i,j=1 i,j=1
4 4
X 2 1 −4
X 2 2
= 1
42 Rhkij = 2f (Rhkij ) = 21 f −4 |R|e
h,k,i,j=1 h,k,i,j=1
2
(16.17) = 12 f −4 |Pf ∨ I|e .
Using (15.72) with n = 4 and (16.16), we have
2 2 2 2 2
|Pf ∨ I|e = f rm−e |Pf |e + Tr(Pf ) = 2 |Pf |e = 2 −4λ2 r−6 4x ⊗ x − r2 I e
2
= 32λ4 r−12 4x ⊗ x − r2 I e = 32λ4 r−12 4xi xj − r2 δij 4xi xj − r2 δ ij
 

= 32λ4 r−12 16xi xj xi xj − 8r2 xi xj δ ij + r4 δij δ ij




= 32λ4 r−12 16r4 − 8r4 + 4r4 = 384λ4 r−8 .




Thus,
2 2
F− 1 −4
|Pf ∨ I|e = 12 f −4 384λ4 r−8

e
= 2f
−4 192λ4
1 + λ2 r−2 192λ4 r−8 =

(16.18) = 4,
(r2 + λ2 )
and

192λ4
Z Z
2
F− d4 x = Vol S 3 3

e 4r dr
R4 0 (r2 + λ2 )
Z ∞
r3
= 2π 2 · 192λ 4
4 dr.
0 (r2 + λ2 )
Using the substitution y = r2 + λ2 (with dy = 2rdr), we compute
Z ∞ Z ∞ Z ∞
r3 1 r2 1 y − λ2
4 dr = 2 4 dy = 2 dy
0 (r2 + λ2 ) 2
λ2 (r + λ )
2
λ2 y4
Z ∞
∞ 1
= 21
y −3 − λ2 y −4 dy = − 41 y −2 + 16 λ2 y −3 λ2
= .
λ2 12λ4
Hence,
2π 2 · 192λ4
Z Z
2 2
F− h νh = F− d4 x = = 32π 2 . 
R4 R4
e 12λ4
Remark 16.7. If we had only integrated over the ball r ≤ λ, we would have
obtained half this value, since
∞ 1 1 1
− 41 y −2 + 61 λ2 y −3 2λ2 = = .
24λ4 2 12λ4
2
Thus, the total field strength R4 |F − |h νh of the instanton is 32π 2 , and as half of
R

this is in the ball r ≤ λ, we say that the size of the instanton is λ.



Let π0 : P0 = R4 \ {0} × SU(2) → R4 \ {0} denote the trivial principal SU(2)-
bundle, and let ρ : R4 \ {0} denote the section ρ(x) := (x, I). There is a unique
connection 1-form ω0 ∈ C(P ), such that ρ∗ ω0 = A− . Since we have found that
YM(A− ) < ∞, the results in [418] imply that the P0 can be extended over 0 and
476 16. GAUGE THEORETIC INSTANTONS

∞ to a bundle P → S 4 in such a way that ω0 extends to a smooth connection


1-form ω on P . However, it is instructive to do this explicitly.
We can always extend π0 : P0 → R4 \ {0} trivially to π : R4 × SU(2) → R4
with π(x, B) = x, but ω0 will not extend smoothly. Another possibility is to take
a trivial bundle π1 : P1 → R4 but identify it nontrivially with π0 : P0 → R4 \ {0}
over R4 \ {0}. In other words, identify (x, B) ∈ P1 with (x, g(x) B) ∈ P0 for some
function g : R4 \ {0} → SU(2). This identification mapping Q : P1 → P0 given by
Q(x, B) := (x, g(x) B) = (x, B) B −1 g(x) B ,


defines a gauge transformation of P0 , since for C ∈ SU(2) , we have


Q((x, B) C) = Q(x, BC) = (x, g(x) BC) = Q(x, B) C.
The associated function q ∈ C(P0 , SU(2)) of Proposition 15.28 is given by
q(x, B) = B −1 g(x) B.
Under the identification mapping Q, the connection form ω0 on P0 , when viewed
on P1 |R4 \{0} , is just Q∗ ω0 . Our task is then to find g : R4 \ {0} → SU(2) such that
Q∗ ω0 extends smoothly to all of P1 . According to (15.27) in Proposition 15.29,
Q∗ ω0 = q −1 dq + q −1 ω0 q.
Let ρ1 : R4 \ {0} → P1 denote the section ρ1 (x) = (x, I). If ρe := Q ◦ ρ1 , then using
the fact q(ρ1 (x)) = q(x, I) = I −1 g(x) I = g(x) , we have
ρe∗ ω0 = ρ∗1 (Q∗ ω0 ) = ρ∗1 q −1 dq + q −1 ω0 q


−1 −1
(16.19) = (q ◦ ρ1 ) d(q ◦ ρ1 ) + (q ◦ ρ1 ) ρ∗1 ω0 (q ◦ ρ) = g −1 dg + g −1 A− g.
We claim that there is g : R4 \ {0} → SU(2), such that ρe∗ ω0 = ρ∗1 (Q∗ ω0 ) extends
smoothly at 0 ∈ R4 . Indeed, take
1 x
(16.20) g(x) := (x4 I − ix · σ) ←→ .
r r
We compute
   
−1 x  x  x dx xrdr x dx x(x · dx)
g dg = d = − 3 = −
r r r r r r r r3
1 1
= 2 (xdx − (x · dx)) = 2 Im(xdx) ,
r r
and similarly
1
gd g −1 = 2 Im(xdx) .

r
Thus,
λ2 1 λ2
A− = 2 2 gd g −1 .

2
Im(xdx) = 2 2
r r +λ r+ λ
Hence, using the fact 0 = dI = d g −1 g = d g −1 g + g −1 dg, we have


λ2
 
ρe∗ ω0 = g −1 dg + g −1 A− g = g −1 dg + g −1 2 −1

gd g g
r + λ2
λ2 λ2
= g −1 dg + 2 d g −1 g = g −1 dg − 2 g −1 dg

r +λ 2 r + λ2
r2 −1 r2 1 1
= 2 2
g dg = 2 2 2
Im(xdx) = 2 Im(xdx) ,
r +λ r +λ r r + λ2
16.2. INSTANTONS ON EUCLIDEAN 4-SPACE 477

which is smooth at 0 ∈ R4 . We can show that A− itself (without gauge transfor-



mation) extends smoothly at ∞ as follows. Let J : R4 \ {0} −→ R4 \ {0} denote the
1 x −
inversion map J(x) = x = r2 . For A to extend smoothly at ∞, we need to  check
that J ∗ A− extends smoothly across 0. We have J ∗ g = g −1 (and J ∗ g −1 = g),
since   x
x 2 x −1
g(J(x)) = g 2 = rx = = g(x) .
r r 2 r
Then J ∗ A− extends smoothly over x = 0 by the computation
λ2 λ2
 
J ∗ A− = J ∗ 2 −1
J ∗ (g) dJ ∗ g −1
 
2
gd g = ∗ 2 2
r +λ J (r ) + λ
1 1 1 1
= −2 2
g −1 dg = −2 2 2
Im(xdx) = Im(xdx) .
r +λ r +λ r 1 + r2 λ2

The function g : R4 \ {0} → SU(2) −→ S 3 is used to identify (clutch) P0 and P1
over R4 \ {0} to obtain π : P → S 4 . Since g restricts to the (degree 1) identity map
S 3 → S 3 on S 3 ⊆ R4 \ {0}, we suspect that the Chern number c2 (E 0 ) [S 4 ] of the
0 2
associated bundle E := P ×SU(2) C is ±1, where the representation r : SU(2) →
GL C is just inclusion. We now verify that c2 (E 0 ) [S 4 ] = −1, using the result
2
2
|F − |h νh = 32π 2 of Proposition 16.6.
R
R4

Proposition 16.8. Let π : P → M be an arbitrary principal SU(2)-bundle over


a compact 4-manifold M . If E 0 := P ×SU(2) C2 and E := P ×SU(2) su(2)C , then
(16.21) 4c2 (E 0 ) = c2 (E) = −ch2 (E) .
For the bundle π : P → S 4 defined above by clutching P0 and P1 over R4 \ {0}, we
have c2 (E 0 ) [S 4 ] = −1.
Proof. Note that any C ∈End C2 ∼

= C2 ⊗ C2∗ can be written uniquely  in
the form C = zB + wI, where B ∈ su(2) and z, w ∈ C, so that End C2 =
su(2)C ⊕ CI and thesummands are SU(2)-invariant subspaces of the representation
ρ : SU(2) →End C2 given by
ρ(B)(ξ) = B ◦ ξ ◦ B −1 .
Hence,
C2 ⊗ C2 ∼ = C2 ⊗ C2∗ = End C2 = su(2)C ⊕ CI.


Here we have used the fact that r : SU(2) → GL C2 and r∗ : SU(2) → GL C2∗ are
 

equivalent. This follows from the fact that the irreducible representations of SU(2)
are determined by dimension 2s+1 (s = spin = 0, 21 , 1, · · · ), but a direct equivalence
can be exhibited as follows. There is a standard skew-symmetric bilinear form ε on
C2 given by
 
v1 w1
ε((v1 , v2 ) ,(w1 , w2 )) := v1 w2 − v2 w1 = det .
v2 w2
Define C2 → C2∗ by v 7→ ε(v, ·). Then for any B ∈ SU(2) ,
B(v) 7→ ε(B(v) , ·) = ε B −1 B(v) , B −1 (·) = det B −1 ε v, B −1 (·)
  

= ε v, B −1 (·) ,

478 16. GAUGE THEORETIC INSTANTONS

−1
since det B −1 = det(B) = 1, and so v 7→ ε(v, ·) is an equivalence. In terms of


the bundles

E 0 := P ×SU(2) C2 and E = P ×SU(2) su(2)C ,

we have (where 1C denotes the trivial line bundle over S 4 )

E ⊕ 1C ∼
= E0 ⊗ E0.

We have observed (see Remark 15.62) that for associated SU(2)-bundles V , ch1 (V ) =
2
c1 (V ) = 0 and ch2 (V ) = 21 c1 (V ) − c2 (V ) = −c2 (V ). Hence,

ch(E) ⊕ ch(1C ) = 4 − 2c2 (E) and


2 2
ch(E 0 ⊗ E 0 ) = ch(E 0 ) = (2 − c2 (E 0 )) = 4 − 4c2 (E 0 )

imply that 4c2 (E 0 ) = c2 (E) = −ch2 (E). In the case of the clutched bundle π : P →
S 4 , we have found that
Z Z Z Z
ω+ 2 ω− 2 ω+ 2 2
8π ch(E)2 S 4 =
2
F−
 
Ω νh − Ω νh = Ω νS 4 = h
νh
S4 S4 S4 R4
Z
2
= F− e
νe = 32π 2 .
R4

Hence we have

ch(E)2 S 4 = 4 and 4c2 (E 0 ) S 4 = −ch2 (E) S 4 = −4,


     

and so c2 (E 0 ) S 4 = −1.
 


Alternative Computation of the Field Strength. Before going  on to find


instantons for bundles π : P → S 4 with arbitrary negative c2 (E 0 ) S 4 , we will

compute the field strength F − = τ ∗ Ωθ rather easily by using the gauge equivalent
local connection form Ae := ρe∗ ω0 = 2 2 Im(xdx) instead of A− .
1
r +λ

Proposition 16.9. The local field strength Fe := ρe∗ Ωω0 = dA e∧A


e+A e is given
by

λ2
(16.22) Fe(x) = 2 dx ∧ dx,
(r2 + λ2 )


where x ∈ H ∼
= R4 . Moreover, F − = τ ∗ Ωθ is given by

−1 λ2 x x
(16.23) F − (x) = g(x) Fe(x) g(x) = 2 (dx ∧ dx) .
(r2 + λ2 ) r r
16.2. INSTANTONS ON EUCLIDEAN 4-SPACE 479

Proof. We have

F 0 = dA0 + A0 ∧ A0
 
1 1 1
=d 2 2
Im(xdx) + 2 2
Im(xdx) ∧ 2 Im(xdx)
r +λ r +λ r + λ2
 
1 − Im(d(xx) ∧ xdx)
=
(r2 + λ2 )
2 + r2 + λ2 Im(dx ∧ dx) + Im(xdx) ∧ Im(xdx)
 
1 − Im((d(x)x + xdx) ∧ xdx)
=
(r2 + λ2 )
2 + r2 + λ2 Im(dx ∧ dx) + Im(xdx) ∧ Im(xdx)
−r2 Im(dx ∧ dx) − Im(xdx ∧ xdx)
 
1
=
(r2 + λ2 )
2 + r2 + λ2 Im(dx ∧ dx) + Im(xdx) ∧ Im(xdx)
1 2

=
2 2 2 λ Im(dx ∧ dx) + Im(xdx) ∧ Im(xdx) − Im(xdx ∧ xdx) .
(r + λ )

To obtain (16.22), it remains to show that

(16.24) Im(xdx ∧ xdx) = Im(xdx) ∧ Im(xdx) .

For this, we have

xdx ∧ xdx = (Re(xdx) + Im(xdx)) ∧(Re(xdx) + Im(xdx))


= Re(xdx) ∧ Re(xdx) + Im(xdx) ∧ Im(xdx)
+ Im(xdx) ∧ Re(xdx) + Re(xdx) ∧ Im(xdx)
(16.25) = Re(xdx) ∧ Re(xdx) + Im(xdx) ∧ Im(xdx) ,

where Im(xdx) ∧ Re(xdx) + Re(xdx) ∧ Im(xdx) = 0, since real and pure imaginary
quaternions commute, while 1-forms anticommute. Note that

Im(xdx) ∧ Im(xdx) = −Im(xdx) ∧ Im(xdx) = − (− Im(xdx)) ∧(− Im(xdx))


= − Im(xdx) ∧ Im(xdx) .

So the fact that Im(xdx) ∧ Im(xdx) is pure imaginary, and taking the imaginary
part of (16.25), yields (16.24). We have (16.23), either by direct computation using
(16.19), or by using the general result (15.28). 

Remark 16.10. Recall (see (16.13)) that

− 12 dx ∧ dx = dx2 ∧ dx3 + dx1 ∧ dx4 i + · · ·.




Since the 2-forms dx2 ∧ dx3 + dx1 ∧ dx4 ,. . . are each of norm-square 2 and i,
j, and k have norm-square 8 with the Killing metric on S 3 ∼
= SU(2), we have that
2
|dx ∧ dx| = 4(3 ·(2 · 8)) = 192. Thus,

2 2 192λ4
F− e
= |F 0 |e = 4,
(r2 + λ2 )

in agreement with (16.18).


480 16. GAUGE THEORETIC INSTANTONS

Finding Instantons for Negative Chern Number. Suppose that we have


found a principal SU(2)-bundle
π : P → S 4 with c2 (E 0 ) S 4 = −k
 

(recall E 0 := P ×SU(2) C2 ). Then (see (16.3)) for any connection 1-form ω on P ,


(16.26) Z Z 
1 ω− 2 ω+ 2
−4k = 4c2 (E 0 ) S 4 = −ch2 (E) S 4 =
   
Ω νh − Ω ν h .
8π 2 S 4 S4

Thus, we can only have self-dual connections (Ωω− = 0) if k ≥ 0, and only anti-
self-dual connections (Ωω+ = 0) if k ≤ 0. If k = 0 (e.g., P = S 4 × SU(2)), then all
such (anti-)self-dual connections are flat, namely Ωω = 0. We produced self-dual
connections in the case k = 1, by considering the so− (∼= su(2))-valued component
θ− of the Levi-Civita connection θ for the metric h = f 2 ds2 on R4 by taking
f = 1 + λ2 r−2 . If we had considered θ+ for such h, then we would have found
anti-self-dual connections for the case k = −1. It is reasonable (and correct) that
for arbitrary k ≥ 1, we should consider the harmonic function
k
X λ2i
f (x) := 1 + 2
i=1 |x − xi |

for nonzero λ1 , . . . , λk ∈ R and distinct points x1 , . . . , xk ∈ R4 . In order to show


that the self-dual connection θ− with
(16.27) A− = τ ∗ θ− = − 21 Im(∂(log f ) dx)
yields a self-dual connection on some principal SU(2)-bundle π : P → S 4 with
0 4

c2 (E ) S = −k, we proceed as follows. From (16.17),
2 2 2
(16.28) F− e
= 12 f −4 kPf ∨ Ike = f −4 kPf ke .
−2
Letting fi (x) = λ2i |x − xi | , we have f = 1 + i fi (1 ≤ i ≤ k) and
P

2
Pf = −2f ∇2 f + 4(df ⊗ df ) − |df | I
 X   X  X X  X 2
= −2 1 + f i ∇2 1 + fi + 4 dfi ⊗ dfj − dfi I
i j i j i
X X
2 2

= −2 ∇ fi + −2fj ∇ fi + 4dfi ⊗ dfj − hdfi , dfj i I
i i,j
X X
∇2 f i + −2fj ∇2 fi + 4dfi ⊗ dfj − hdfi , dfj i I ,

= −2
i i6=j

where in the last equality we have used the fact that fi2 ds2 is flat (so that Pfi = 0)
to deduce that the terms with i = j vanish. Setting ri (x) := |x − xi |, we have
fi = λ2i ri−2 ,
 −1  −2
dfi = λ2i d ri−2 = λ2i d ri2 = −λ2i ri2 d ri2
 

−2
= −λ2i ri2 2(x − xi ) · dx, and
∇ fi pq = ∂p ∂q fi = ∂p ∂q λ2i ri−2 = λ2i ∂p ∂q ri−2
2
  

= 2λ2i ri−6 4(xp − xpi )(xq − xqi ) − ri2 δpq , and so



16.2. INSTANTONS ON EUCLIDEAN 4-SPACE 481

−2 2
2
|dfi | = −λ2i ri2 2(x − xi ) · dx = 4λ4i ri−6 and
2
∇ 2 fi = 48λ4 ri−8 .

For j 6= i, the quantities fj , |dfj | and ∇2 fj are bounded about xi . Thus as ri → 0,


2 2
kPf ke = 4 ∇2 fj + O ri−6 = 192λ4i ri−8 + O ri−6 .
 

Then
 −4
2 2
X 2
F− e
= f −4 |Pf |e = 1 + λ2i ri−2 |Pf |e
i
 X −4
= 1 + λ2i ri−2 + λ2j rj−2 192λ4i ri−8 + O ri−6

i6=j
 X −4
2
= 192λ4i ri2 + λ2i + λ2j (ri /rj ) + O ri2 ,

i6=j

which is bounded near xi . Also, as r := |x| → ∞, we have


2 2
X 2
F − e = f −4 |Pf |e = 4 ∇2 fi + O r−10 ≤ Cr−8 ,

i
2
for some constant C which tends to 0 as(λ1 , . . . , λk ) → 0. Thus, |F − |e is integrable.
By K. Uhlenbeck’s Theorem (or by directly using clutching functions gi as in (16.20)
about the xi ), one can deduce that A− , defined by (16.27) on R4 −{x1 , . . . , xk }, lifts
to a smooth connection 1-form, say ω, for some principal SU(2)-bundle π : P → S4.
0 0
4

By (16.26), we can determine the Chern number c2 (E ) S , where E := P ×SU(2)
C2 . Indeed,
−1 −1
Z Z
ω+ 2 2
c2 (E 0 ) S 4 = F − e νe .
 
Ω ν h =
32π 2 S 4 32π 2 R4
This integral is independent of the nonzero λ1 , . . . , λk ∈ R and the distinct positions
points x1 , . . . , xk ∈ R4 . We can evaluate the integral by first choosing x1 , . . . , xk
far enough apart so that if Bi denotes the Euclidean ball of radius 1 about xi , then
ri /rj ≤ ε in Bi for all j 6= i. Then
Z Z
2 −4
F − e νe = 192λ4i ri2 + λ2i νe + O(ε) .
Bi Bi
2
|F − |e
R
Since R4 −∪i Bi
νe → 0 as (λ1 , . . . , λk ) → 0, and
1
r3
Z Z
−4
lim 192λ4i ri2 + λ2i νe = 2π 2 · 192 lim λ4 4 dr
λi →0 Bi λ→0 0 (r2 + λ2 )
1
= 2π 2 · 192 lim λ4 − 41 y −2 + 16 λ2 y −3

λ→0 λ2
1
= 2π 2 · 192 · = 32π 2 ,
12
we obtain, as ε → 0 and (λ1 , . . . , λk ) → 0,
−1
Z
0 2
F − e νe → −k.
 4
c2 (E ) S =
32π 2 R4
But c2 (E 0 ) S 4 is independent of ε and λ1 , . . . , λk . Thus, c2 (E 0 ) S 4 = −k.
   
482 16. GAUGE THEORETIC INSTANTONS

We can extend the k + 4k = 5k-parameter family of harmonic functions


k
X λ2i
1+ 2
i=1 |x − xi |
to the (5k + 5)-parameter family
k
X λ2i
2.
i=0 |x − xi |
The first family is a limiting case of the second if we take λ0 = |x0 | and let x0 → ∞.
Moreover, the analysis of A− and F − about the points xi is the same as before so
that the singularities of A− are removable and A− lifts and extends to a smooth
connection on a principal SU(2)-bundle π : P → S 4 . The first family is a limiting
case ofthe second if we take λ0 = |x0 | and let x0 → ∞. Thus, the Chern number
c2 (E 0 ) S 4 for the associated bundle E 0 = P ×SU(2) C2 for functions of the second
family will also be −k. Not all of the parameters of the 5k + 5 are effective,
since multiplying f by a constant does not change A− by (16.9) or (16.12). Thus,
there are at most 5k + 4 effective parameters. Actually, there are 5k + 4 effective
parameters only for k ≥ 3. For k = 1, there are only 5 effective parameters
corresponding to the position and size of the BPST instanton, while for k = 2
there are 13 effective parameters (see [232]). In the next section, we prove that the
correct number of parameters is 8k − 3 for all k. There was really no guarantee that
the procedure we used (i.e., generating connections from conformally-flat metrics
f 2 ds2 with vanishing scalar curvature) would give us all of the instantons. There is
a procedure (the ADHM procedure) developed by M. F. Atiyah, V. G. Drinfeld,
N. J. Hitchin, and Y. I. Manin in [36] for capturing all instantons, which we
describe below. However, we mention that the singular metrics
k
!2
2 2
X λ2i
(16.29) f ds = 2 ds2
i=0 |x − xi |

are interesting in their own right. We have noticed that for f (x) = λ2 r−2 (i.e.,
the case k = 1) one finds that f 2 ds2 is the pull-back of ds2 under the inversion
x 7→ λ2 r−2 x of R4 in the sphere r = λ. In other words, when a ball about the
origin is given in the metric f 2 ds2 it becomes isometric to the exterior of a ball of
R4 with the flat metric. While f 2 ds2 in (16.29) is not perfectly flat in a ball about
an xi , the influence of the other terms is comparatively slight, so that the metric is
asymptotically flat as one approaches xi . Thus, the manifold R4 −{x1 , . . . , xk } with
metric f 2 ds2 joins together k asymptotically flat regions, and serves to motivate
the study of gravitational instantons.

The Atiyah-Drinfeld-Hitchin-Manin (8k − 3)-Parameter Family of In-


stantons. It is not difficult to describe the instantons constructed by the ADHM
procedure. The difficulty lies in proving that all instantons are captured by the
procedure. We will content ourselves with a suitably motivated description. Re-
call that for a harmonic function f on R4 , a su(2)-valued connection 1-form with
self-dual curvature is defined by
Af = − 21 Im(∂(log f ) dx) ,
16.2. INSTANTONS ON EUCLIDEAN 4-SPACE 483

where we have identified su(2) with the pure imaginary quaternions. When
k
X λ2j
f (x) = 1 + 2,
j=1 |x − xj |
 
∂f
Af (x) = − 12 Im(∂(log f ) dx) = − 12 Im dx
f
 
k
1 X 2 x − xj
(16.30) = Im λ d(x − xj ) .
f (x) j=1 j |x − xj |4

For a quaternion variable x, we apply the convenient identity


!    !
x 1 1
Im 4 dx = Im d ,
|x| x x
which is derived from the computation
    ! ! ! ! !
1 1 x x x x x 1 1
d = 2 d 2 = 2d 2 = 2 xd 2 + 2 dx
x x |x| |x| |x| |x| |x| |x| |x|
!
x 1 x 1 
−2
 x
= 2 xd 2 + 2 2 dx = d |x| + 4 dx,
|x| |x| |x| |x| |x|
 
−2
using the fact that d |x| is real. Thus, we may rewrite (16.30) as
 
Pk  λj   λj 
j=1 x−xj d x−xj
Af (x) = Im Pk  λj   λj   .
 
1 + j=1 x−xj x−xj

We can generalize this (admittedly with some hindsight) as follows. Although the
λj were taken to be real, we now let them be quaternions. For a column matrix
T
λ = [λ1 , . . . , λk ] ∈ Hk and B a k × k matrix of quaternions, define u : H → Hk by
−1
u(x) := (B − xIk ) λ.
For

T
h i
u∗ := u(x) = u1 (x), . . . , uk (x) , 

  
Pk u∗ du
u∗ du := let Aλ,B (x) := Im .
j=1 uj (x)duj , and  1+|u(x)|2
2 Pk 2
|u(x)| u∗ u = j=1 |uj (x)| ,

:= 

Note that Af (x) in (16.30) is the special case, where the B is the diagonal k × k
matrix with diagonal entries x1 , . . . , xk ∈ R4 ∼
= H.
Assumption 16.11 (ADHM Conditions). In order that Aλ,B define a connec-
tion with self-dual curvature and removable singularities, we need to assume that
T
the following two conditions hold (where B ∗ = B ) :
(I) B is symmetric and λλ∗ + BB ∗ is a real k × k matrix.
(II) For any x ∈ H and ξ ∈ Hk , we have
(B − xIk ) ξ = 0 and λ∗ ξ = 0 =⇒ ξ = 0.

484 16. GAUGE THEORETIC INSTANTONS

In the special case that B is diagonal, then (I) is obviously met. If, in addition,
we assume that all λi are nonzero and B has distinct diagonal entries, then (II) is
met. Indeed, with these assumptions,
(B − x) ξ = 0 and λT ξ = 0
=⇒ for each i, (bii − x) ξi = 0 and λT ξ = 0
=⇒ for each i, ξi = 0 or x = bii and λT ξ = 0
=⇒ ξ = 0, or there is at most one i, with ξi 6= 0 and λi ξi = λT ξ = 0
=⇒ ξ = 0.

Theorem 16.12 (Instantons Described). If (λ, B) satisfies Conditions (I) and


(II), then Aλ,B is an instanton (i.e., the curvature Fλ,B of Aλ,B is self-dual). Also,
Aλ,B lifts and extends to a connection on a principal SU(2)-bundle π : P → S 4 .
Moreover, if λ0 = T λq and B 0 = T BT −1 for some unit q ∈ H (i.e., q ∈ S 3 ) and
some T ∈ SO(k) , then (λ0 , B 0 ) also satisfies Conditions (I) and (II), and
Aλ0 ,B 0 = q −1 Aλ,B q,
so that Aλ0 ,B 0 is equivalent (via a gauge transformation) to Aλ,B .
Proof. First we introduce the following notation:
q
2
σ(x) := 1 + |u(x)| and
   
1 1 1 1
U (x) := = −1 .
σ(x) u(x) σ(x) (B − xIk ) λ
Note that
  
1
U ∗ dU = σ −1 d σ −1
u∗
 
1
u
    
0  1
= σ −1 1 u∗ σ −1 + d σ −1
 
du u
 
1
= σ −2 u∗ du + σ −1 d σ −1 1 u∗
  
u
= σ −2 u∗ du + σd σ −1 .


Since σd σ −1 is real, we then have




!
u∗ du
Aλ,B (x) := Im 2 = Im(U ∗ dU ) .
1 + |u(x)|
Moreover,
Fλ,B = dAλ,B + Aλ,B ∧ Aλ,B
= d Im(U ∗ dU ) + Im(U ∗ dU ) ∧ Im(U ∗ dU )
= Im(d(U ∗ dU )) + Im((U ∗ dU ) ∧(U ∗ dU )) ,
since d commutes with the algebraic projection Im, and
Im(U ∗ dU ) ∧ Im(U ∗ dU ) = Im(U ∗ dU ) ∧ Im(U ∗ dU )
16.2. INSTANTONS ON EUCLIDEAN 4-SPACE 485

by the computation
U ∗ dU ∧ U ∗ dU = (Re(U ∗ dU ) + Im(U ∗ dU )) ∧(Re(U ∗ dU ) + Im(U ∗ dU ))
= Im(U ∗ dU ) ∧ Im(U ∗ dU ) + Re(U ∗ dU ) ∧ Re(U ∗ dU )
+ Im(U ∗ dU ) ∧ Re(U ∗ dU ) + Re(U ∗ dU ) ∧ ImU ∗ dU
= Re(U ∗ dU ) ∧ Re(U ∗ dU ) + Im(U ∗ dU ) ∧ Im(U ∗ dU ) ,
where the cross terms cancel, since 1-forms anti-commute, while real and pure
imaginary quaternions commute. Thus,
Fλ,B = Im(d(U ∗ dU ) + (U ∗ dU ) ∧(U ∗ dU ))
= Im(dU ∗ ∧ dU + (U ∗ dU ) ∧(U ∗ dU )) = Im(FU ) ,
where
FU := dU ∗ ∧ dU + (U ∗ dU ) ∧(U ∗ dU ) .
Let Mm,n (H) denote the set of m × n matrices with quaternion entries. While
U ∗ U = 1, we have
P := U U ∗ ∈ Mk+1,k+1 (H) .
Note that P 2 = U U ∗ U U ∗ = P and P : Hk+1 → Hk+1 is a projection onto
span(U ) := {U q : q ∈ H} .
We have
FU = U ∗ (dP ∧ dP ) U
by the following computation (where we use d(U ∗ U ) = d(1) = 0)
U ∗ (dP ∧ dP ) U = U ∗ (d(U U ∗ ) ∧ d(U U ∗ )) U
= U ∗ (((dU ) U ∗ + U dU ∗ ) ∧((dU ) U ∗ + U dU ∗ )) U
= (U ∗ (dU ) U ∗ + dU ∗ ) ∧(dU + U (dU ∗ ) U )
= dU ∗ ∧ dU + U ∗ dU ∧ U ∗ dU
+ ((dU ∗ ) U + U ∗ (dU )) ∧(dU ∗ ) U
= dU ∗ ∧ dU + U ∗ dU ∧ U ∗ dU + d(U ∗ U ) ∧(dU ∗ ) U = FU .
Hence, the self-duality of FU will follow, once the self-duality of dP ∧ dP is shown.
To demonstrate that dP ∧ dP is self-dual, it turns out to be easier to work with the
complementary projection Q := I − P . Note that
Q := I − P =⇒ dP ∧ dP = d(I − P ) ∧ d(I − P ) = dQ ∧ dQ.
In order to find an appropriate formula for Q, we first seek v ∈ Mk+1,k (H), whose
columns span a subspace, say U ⊥ , of Mk+1 (H) which is orthogonal to U ; i.e.,
U ∗ v = 0. Writing  
v1
v= ∈ Mk+1,k (H) ,
v2
where v1 ∈ M1,k (H) and v2 ∈ Mk,k (H), we want
 

 v1

= v1 + u∗ v2

0 = σU v = 1 u
v2
 ∗
−1 ∗−1
= v1 + (B − xIk ) λ v2 = v1 + λ∗ (B − xIk ) v2 .
486 16. GAUGE THEORETIC INSTANTONS


This is achieved by taking v2 = (B − xIk ) and v1 = −λ∗ . Thus, we take
−λ∗
 
v= ∗ ∈ Mk+1,k (H) .
(B − xIk )
Note that Condition II says precisely that v has rank k. Then v ∗ v = Mk,k (H)
is invertible, and the orthogonal projection Q of Hk+1 onto U ⊥ is given by the
quaternionic version of the usual formula, namely
−1
Q = v(v ∗ v) v∗ .
Indeed,
−1 −1
QU = v(v ∗ v) (v ∗ U ) = 0 and Qv = v(v ∗ v) v ∗ v = v.
Now,
U ∗ (dP ∧ dP ) U = U ∗ d(1 − Q) ∧ d(1 − Q) U = U ∗ dQ ∧(dQ) U.
Since U ∗ v = 0, we have
    
−1 −1 −1
U ∗ dQ = U ∗ d v(v ∗ v) v ∗ = U ∗ (dv)(v ∗ v) v ∗ + vd (v ∗ v) v ∗
 
−1 −1 −1
= U ∗ (dv)(v ∗ v) v ∗ + U ∗ vd (v ∗ v) v ∗ = U ∗ (dv)(v ∗ v) v ∗ .

Similarly, since v ∗ U = (U ∗ v) = 0, we have
 
−1 −1
(dQ) U = d v(v ∗ v) v ∗ U = v(v ∗ v) (dv ∗ ) U.

Thus,
−1 −1
FU = U ∗ dQ ∧(dQ) U = U ∗ (dv)(v ∗ v) v ∗ ∧ v(v ∗ v) (dv ∗ ) U
−1
= U ∗ (dv)(v ∗ v) ∧(dv ∗ ) U.
Note that
−λ∗
     
0 0
dv = d ∗ = = dx and
(B − xIk ) − (dx) Ik −Ik
dv ∗ = d [−λ, B − xIk ] = dx [0, −Ik ] .
Hence,
−1
FU = U ∗ dQ ∧(dQ) U = U ∗ (dv)(v ∗ v) ∧(dv ∗ ) U
 
0 −1
= U∗ dx(v ∗ v) ∧ dx [0, −Ik ] U
−Ik
   
0 −1 1
= σ −1 1 u∗ dx(v ∗ v) ∧ dx [0, −Ik ] σ −1
 
−Ik u
−1
= σ −2 u∗ (dx)(v ∗ v) ∧(dx) u .
By (16.13), dx ∧ dx is a self-dual, quaternion-valued form. The intervening factor
−1 −1
(v ∗ v) in (dx)(v ∗ v) ∧(dx) prevents us from asserting that FU is self-dual. On
∗ −1 −1
the other hand, if (v v) ∈ Mk,k (H) has real entries, then (v ∗ v) commutes with
dx, and
 
−1
FU = σ −2 u∗ (v ∗ v) dx ∧ dx u
16.2. INSTANTONS ON EUCLIDEAN 4-SPACE 487

is indeed self-dual. It is precisely Condition (I) which guarantees that v ∗ v (and


−1
hence (v ∗ v) ) is real. To show this, first note that
−λ∗
 
∗ ∗
= λλ∗ + (B − xIk )(B − xIk )
 
v v = −λ B − xIk ∗
(B − xIk )
2
= λλ∗ + BB ∗ − (Bx + xB ∗ ) + |x| Ik .
This is real for all x ∈ H ⇔ λλ∗ + BB ∗ is real and Bx + xB ∗ is real for all x ∈ H.
We show that Bx+xB ∗ is real for all x ∈ H ⇔ B symmetric. We have B symmetric
⇔ B T = B ⇔ B ∗ = B. Assume that Bx + xB ∗ is real for all x ∈ H. Then
xB ∗ = (Bx) = xB,
and in particular for x = 1, B ∗ = B (i.e., B T = B). Conversely, for B symmetric,
we have B ∗ = B and so for all x ∈ H,
Bx + xB ∗ = Bx + xB = xB + Bx = (Bx + xB ∗ ),


namely Bx + xB ∗ is real.
It is easy to check that FU is pure imaginary. Thus,
 
−1
Fλ,B = Im(FU ) = FU = σ −2 u∗ (v ∗ v) dx ∧ dx u.
It is also easy to verify that as |x| → ∞
 
2 −8
|Fλ,B | = O |x| ,
2 
so that |Fλ,B | ∈ L2 R4 , and Uhlenbeck’s Theorem then implies that Aλ,B lifts
and extends to a connection on a principal SU(2)-bundle π : P → S 4 .
The verification of the fact that (λ0 , B 0 ) also satisfies Conditions (I) and (II) is
routine. Note that B 0 is symmetric, since
T
B 0T = RBR−1 = R−1T B T RT = RB T R−1 = B 0 .
Moreover B 0 B 0∗ + λ0 λ0∗ is real, since
∗ ∗
B 0 B 0∗ + λ0 λ0∗ = RBR−1 RBR−1 + (Rλq)(Rλq)
= RBR−1 R B R + (Rλq)(q ∗ λ∗ R∗ )
−1∗ ∗ ∗


= RBR−1 RB ∗ RT + Rλλ∗ RT


= RBB ∗ RT + Rλλ∗ RT = R(BB ∗ + λλ∗ ) RT .


Finally,
(B 0 − xIk ) ξ = 0 and λ0∗ ξ = 0

=⇒ RBR−1 − xIk ξ = 0 and (Rλq) ξ = 0


=⇒ R(B − xIk ) R−1 ξ = 0 and q ∗ λ∗ RT ξ = 0


=⇒ (B − xIk ) R−1 ξ = 0 and λ∗ R−1 ξ = 0


=⇒ R−1 ξ = 0 =⇒ ξ = 0.
−1
Note that for uλ,B (x) := (B − xIk ) λ, we have
−1 −1
uλ0 ,B 0 (x) = (B 0 − xIk )
λ0 = T BT −1 − xIk T λq
−1 −1
= T (B − xIk ) T −1

T λq = (B − xIk ) λq = uλ,B (x) q,
488 16. GAUGE THEORETIC INSTANTONS

and so
! !

(uq) d(uq) qu∗ (du) q
Aλ0 ,B 0 (x) = Im 2 = Im 2
1 + |uq(x)| 1 + |u(x)|
!
u∗ (du)
= q Im 2 q = qAλ,B (x) q. 
1 + |u(x)|

Remark 16.13. The principal SU(2)-bundle π : P → S 4 in Theorem 16.12 has

c2 (E 0 ) [S 4 ] = −k, for k in (16.11) and E 0 = P ×SU(2) C2 .


We have shown this for the special case, where the B is the diagonal k × k matrix
with distinct diagonal entries. If the set of (λ, B) satisfying Conditions (I) and (II)
of (16.11) is connected, then c2 (E 0 ) [S 4 ] = −k for all such (λ, B) by continuity. As
emphsized above, a much more difficult task would be to show that every instanton
arises in this fashion, as stated in Theorem 16.14 below. Using methods from
algebraic geometry and twistor theory, this was accomplished in [36]. An excellent,
expanded presentation is found in [24], in which Theorem 16.14 appears on p. 26.
The reader will notice some differences between our presentation and Atiyah’s,
since Atiyah writes his quaternions as x = x1 + x2 i+x3 j+x4 k, whereas we have
written x = x4 + x1 i+x2 j+x3 k. This causes a reversal of orientation. It also (for
our own good) made us feel obligated to go through the entire construction down
to the last detail in the proof of Theorem 16.12.

Theorem 16.14 (Instantons Captured). Every instanton on R4 , which  extends


to an instanton on a principal SU(2)-bundle π : P → S 4 with c2 P × C2 [S 4 ] = −k,
is of the form Aλ,B for (λ, B) satisfying Conditions (I) and (II). Moreover, Aλ,B
is equivalent (via a gauge transformation) to Aλ0 ,B 0 ⇔ λ0 = T λq and B 0 = T BT −1
for some unit q ∈ H (i.e., q ∈ S 3 ) and some T ∈ SO(k).

Given Theorem 16.14, we can count the number of independent parameters


in the space of instantons modulo gauge transformations. For a fixed k, the real
dimension of the space of (λ, B) for which B is symmetric, is

4k + 4 · 21 k(k + 1) = 2k 2 + 6k.
∗ T
Since (B ∗ B + λ∗ λ) = B ∗ B + λ∗ λ, we have (B ∗ B + λ∗ λ) = B ∗ B + λ∗ λ, and so
T
(Im(B ∗ B + λ∗ λ)) = Im B ∗ B + λ∗ λ = − Im(B ∗ B + λ∗ λ) .


Thus, Im(B ∗ B + λ∗ λ) is skew-symmetric and setting it equal to 0 eliminates at most


3· 21 k(k − 1) dimensions. Quotienting by the group S 3 ×SO(k) of gauge equivalences
among the Aλ,B further reduces the dimension by at most 3 + 12 k(k − 1). Thus,
the dimension of the space of Aλ,B satisfying conditions (I) and (II) is at least

2k 2 + 6k − 3 · 12 k(k − 1) − 3 + 21 k(k − 1) = 8k − 3.


In the next section, we apply the index theorem to prove that the space of in-
stantons modulo gauge transformations is a manifold of dimension 8k − 3. Thus,
the preceding intuitive dimension count is actually correct. In other words, the
conditions imposed are actually all independent, in case there was any doubt.
16.3. LINEARIZATION OF THE MODULI SPACE OF SELF-DUAL CONNECTIONS 489

3. Linearization of the Moduli Space of Self-dual Connections


In Section 16.1 we observed that YM : C(P ) → R+ is invariant under the action
on C(P ) of the
 group GA(P ) of gauge transformations, since for F ∈ GA(P ) ,
−1 ∗
F ·ω = F ω yields
F ·ω
∗
Ω = d(F · ω) + (F · ω) ∧(F · ω) = F −1 (dω + ω ∧ ω)
∗
= F −1 Ωω = F · Ωω ,
2 2
and then |F · Ωω | = |Ωω | by Corollary 15.31. In particular, the set of critical
points of YM is preserved by this action, as well as the (possibly empty) set of
(anti-) self-dual connections (absolute minima of YM, if they exist). Thus, the
quotient space
(16.31) M := C(P ) / GA(P )
is a natural object for study. In particular, one would like to know the extent to
which we may regard M as a manifold, say modeled on some infinite dimensional
Fréchet space. Also, if C(P )+ denotes the space of self-dual connections (i.e., with
self-dual curvature), it would be of interest to compute the dimension of the space
of moduli of self-dual connections
+
(16.32) M+ := C(P ) / GA(P ),
at least where it is a submanifold of M. In the heuristic discussion which follows,
we will talk rather loosely about infinite dimensional manifolds, and their tangent
spaces and normal spaces. However, this discussion will motivate a precise theorem
with a precise proof. In Section 16.4, we will provide the framework within which
it makes sense to speak of M+ as being a submanifold of M.
Heuristic discussion: For ω ∈ C(P ), the orbit of the action of GA(P ) on
C(P ) is
GA(P ) · ω := {F · ω : F ∈ GA(P )} = {Φ(f ) · ω : f ∈ C(P, G)}
in the notation of Proposition 15.28, p. 411. Formally, the Lie algebra of C(P, G)
is C(P, g) and the infinitesimal action of C(P, g) on C(P ) is given (see Proposition
15.33, p. 413), for s ∈ C(P, g) and ω ∈ C(P ), by
1
s·ω = d
dt (Φ(Exp(ts)) · ω) t=0
= − (ds + [ω, s]) = −Dω s ∈ Ω (P, g) .
Thus the formal tangent space at ω ∈ C(P ) of the orbit GA(P ) · ω is
n 1
o
Tω (GA(P ) · ω) := Dω s ∈ Ω (P, g) : s ∈ C(P, g) .
1
Note that Tω (GA(P ) · ω) ⊆ Ω (P, g), which is the vector space on which the affine
1
space C(P ) is modeled (see Remark 15.12, p. 401). For τ ∈ Ω (P, g) and s ∈ C(P, g),
we have
(τ, Dω s) = (δ ω τ, s) ,
where δ is covariant codifferential, the formal adjoint of Dω (see Proposition 15.19,
ω

p. 406). Thus, the formal normal space to GA(P ) · ω at ω is


n 1
o
(16.33) Nω (GA(P ) · ω) := τ ∈ Ω (P, g) : δ ω τ = 0 .
The formal slice of the action is
(16.34) Sω := {ω + τ : τ ∈ Nω (GA(P ) · ω)} = {ω + τ : δ ω τ = 0} .
490 16. GAUGE THEORETIC INSTANTONS

Intuitively, we expect that every connection ω 0 in a suitably small neighborhood of


ω will be gauge-equivalent to a connection in Sω ; i.e., there is some F ∈ GA(P ),
such that F · ω 0 ∈ Sω .
Recall from Section 15.8 that the isotropy subgroup of ω is
Iω := {F ∈ GA(P ) : F · ω = ω} .
Note that Iω leaves S setwise fixed, since, for τ ∈ Nω (GA(P ) · ω),
F ∈ Iω =⇒ F · (ω + τ ) = ω + F · τ and δ ω (F · τ ) = F · δ F ·ω τ = F · δ ω τ = 0,
where we have used Corollary 15.32 on p. 413 Unless Iω acts trivially on Sω , there
will be gauge-equivalent connections in Sω , and so Sω will not parametrize M :=
C(P ) / GA(P ) in a 1-1 fashion even locally about ω. Note that if g ∈ Z(G) := the
center of G, then Rg : P → P is in GA(P ), and for any ω 0 ∈ C(P ) we have Rg∗ ω 0 = ω 0
by (15.27), p. 411. A condition which implies that Iω consists of only these central
gauge transformations is that ω be irreducible; this is immediate from Proposition
15.64, p. 456, which asserts that Iω is isomorphic to the centralizer Z(Hol(ω, p0 )) (in
G) of the holonomy group Hol(ω, p0 ). The center of a semi-simple, compact G must
be finite. In this case, for irreducible ω, there can be no nontrivial one-parameter
subgroup in Iω .
Proposition 16.15. Let α ∈ C(P, g). Then Φ(Exp(tα)) ∈ Iω for all real t, if
and only if Dω α = 0. Thus, the following are equivalent (where Hol(ω, p) denotes
the holonomy group of ω with arbitrary reference point p ∈ P ):
 1

(i) Ker Dω : C(P, g) → Ω (P, g) = {0}.
(ii) Z(Hol(ω, p)) = {g ∈ G : gg0 = g0 g for all g0 ∈ Hol(ω, p)} is discrete.
Proof. Using (15.33) on p. 413, for any real t0 , we have
d d
dt (Φ(Exp(tα)) · ω) t=t0
= dt (Φ(Exp((t0 + t) α)) · ω) t=0
d
= dt (Φ(Exp(t0 α) Exp(tα)) · ω) t=0
d
= dt ((Φ(Exp(t0 α)) · Φ(Exp(tα))) · ω) t=0
d
= dt (Φ(Exp(t0 α)) ·(Φ(Exp(tα)) · ω)) t=0
  ∗ 
d −1
= dt Φ Exp(t0 α) (Φ(Exp(tα)) · ω)
t=0
∗ d

= Φ(Exp(−t0 α)) dt Φ(Exp(tα)) ·ω t=0
∗ ω
= Φ(Exp(−t0 α)) (−D α) .
Since Φ(Exp(0 · α)) · ω = ω, Φ(Exp(tα)) · ω = ω for all t ⇔ Dω α = 0. 
Definition 16.16. Let π : P → M principal G-bundle, where G is a compact,
semi-simple Lie group. A connection 1-form ω on P is called weakly irreducible
if (i), or equivalently (ii), holds in Proposition 16.15.
If ω ∈ C(P )+ is a weakly-irreducible, self-dual connection, then
0 0
ω 0 ∈ C(P )+ ∩ Sω ⇐⇒ ∗Ωω = Ωω and δ ω (ω 0 − ω) = 0.
Writing τ = ω 0 − ω (or ω 0 = ω + τ ), we have
0
Ωω = dω 0 + 1
2 [ω 0 , ω 0 ] = dω + 1
2 [ω, ω] + dτ + ad(ω) ∧ τ + 1
2 [τ, τ ]
ω ω 1
=Ω +D τ + 2 [τ, τ ] .
16.3. LINEARIZATION OF THE MODULI SPACE OF SELF-DUAL CONNECTIONS 491

0
1

Thus, ∗Ωω = Ωω ⇔ (1 − ∗) Dω τ + 2[τ, τ ] = 0. Hence,
(A) δ ω τ = 0 and

(16.35) ω 0 ∈ C(P )+ ∩ Sω ⇐⇒ 1

(B) (1 − ∗) Dω τ + 2 [τ, τ ] = 0.
Condition (A) is linear and the linearization of the quadratic Condition (B) is
(1 − ∗) Dω τ = 0. Hence, the formal tangent space of C(P )+ ∩ Sω at ω is
n 1
o
Tω C(P )+ ∩ Sω := τ ∈ Ω (P, g) : δ ω τ = 0 and (1 − ∗) Dω τ = 0 .

(16.36)

Spinor-Free Dimension Calculation. The preceding motivates the follow-


ing key result, which specifies the dimension of the hypothetical manifold C(P )+ ∩Sω
near a weakly-irreducible self-dual connection ω for a suitable principal G-bundle
π : P → M . The original proof in [39] makes use of the Dirac operator on spinor
fields and the Ab genus, along with the Index Theorem. The proof here is spinor-free,
and is essentially Hodge-theoretic. The Index Theorem is used mainly to handle the
twist in the bundle P ×G gC . A reader may be surprised how long our Rproof of the
Formal Dimension Theorem is. Indeed, once we have that index T0 = M ch(E) `
(something involving Riemannian curvature) one is tempted to conclude (by look-
ing at homogeneous parts) that the answer is 8k − 3 (namely, the index of the
untwisted operator T0 ) and then compute the index of the twisted operator T us-
ing the Hodge Theorem. That sounds easy. However, when checking the details, it
seems to us that one has to go through the same lengthy calculations as in the proof
here. A radically different excision/homotopy proof idea may be more promising,
when we first reduce to the trivial bundle and then deform to the trivial connection.
Someone should try.
Theorem 16.17 (Formal Dimension Theorem). Let G be a Lie group and as-
sume G semi-simple and compact. Let ω ∈ C(P )+ be a weakly-irreducible, self-
dual connection for a principal G-bundle π : P → M over a self-dual, compact,
connected, oriented Riemannian 4-manifold M with nonnegative scalar curvature
S 6= 0. For the formal tangent space Tω (C(P )+ ∩ Sω ) defined by (16.36), we have
(16.37) dim Tω C(P )+ ∩ Sω = 2ch(P ×G gC ) [M ] − 12 dim(G)(χ(M ) − sig(M )) ,


where P ×G gC → M denotes the vector bundle associated to π : P → M via the


complex adjoint representation adC : G → GL(gC ) , gC := C ⊗ g, χ(M ) denotes the
Euler characteristic of M , and sig(M ) denotes the signature of M .
Proof. Let E := P ×G gC , and let Ωk (E) denote the space of E-valued k-forms
on M . By (15.11), p. 403, we may (and often do) make the identification
∼ k
Ωk (E) −→ Ω (P, gC ) .
Let Ω2− (E) denote the space of anti-self-dual E-valued 2-forms on M . Define
T : Ω1 (E) → Ω0 (E) ⊕ Ω2− (E) , for τ ∈ Ωk (E) , by
T (τ ) := δ ω τ, 12 (1 − ∗) Dω τ .

(16.38)
Note that Ker T = Tω (C(P )+ ∩ Sω ). Of course, Ω0 (E) ⊕ Ω2− (E) can be regarded
as the space of sections of the single bundle E ⊗ Λ0 (T ∗ M ) ⊕ Λ2− (T ∗ M ) , and so
T is a differential operator (of order 1) mapping
C ∞ E ⊗ Λ1 (T ∗ M ) −→ C ∞ E ⊗ Λ0 (T ∗ M ) ⊕ Λ2− (T ∗ M ) .
 
492 16. GAUGE THEORETIC INSTANTONS

Since
dim Λ0 (T ∗ M ) = 4 = 1 + 3 = dim Λ0 (T ∗ M ) ⊕ Λ2− (T ∗ M ) ,


we know that T will be elliptic if its symbol σ(T ) is injective. From the local
formulas for Dω and δ ω (see (15.22) and (15.24), p. 407), it follows that σ(T ) =
IdE ⊗σ(T0 ), where σ(T0 ) denotes the symbol of the untwisted (i.e., coefficients no
longer in E) operator T0 : Ω1 (M ) → Ω0 (M ) ⊕ Ω2− (M ) given by
T0 (γ) := δγ, 21 (1 − ∗) dγ .


Let S(M ) ⊆ T ∗ M denote the unit cosphere bundle. Then


σ(T0 ) : S(M ) → Hom T ∗ M, Λ0 (T ∗ M ) ⊕ Λ2− (T ∗ M )


is given, for ξ ∈ S(M ) and η ∈ Tx∗ M , by


σξ (η) = − hξ, ηi , 12 (1 − ∗) ξ ∧ η .


Now, σ(T0 )ξ (η) = 0 ⇒ hξ, ηi = 0 and ∗(ξ ∧ η) = ξ ∧ η. If ∗(ξ ∧ η) = ξ ∧ η and νx


denotes the volume element at x, then
2
|ξ ∧ η| νx = ξ ∧ η ∧ ∗(ξ ∧ η) = ξ ∧ η ∧ ξ ∧ η = 0,
and so η = cξ for some constant c. However, then 0 = hξ, ηi = c hξ, ξi = c, in which
case η = 0. Thus, σ(T0 )ξ is injective for all ξ ∈ S(M ), and T is elliptic.
Since Ker T = Tω (C(P )+ ∩ Sω ), our goal is to compute dim Ker T . We show
that Ker T ∗ = 0, and then go on to compute
dim Ker T = dim Ker T − dim Ker T ∗ = index T
via the index formula. To obtain a formula for T ∗ , we compute (where α ∈ Ω0 (E),
β ∈ Ω2− (E), and τ ∈ Ω1 (E))
(T ∗ (α, β) , τ ) = ((α, β) , T (τ )) = (α, β) , δ ω τ, 12 (1 − ∗) Dω τ


= (α, δ ω τ ) + β, 12 (1 − ∗) Dω τ


(16.39) = (Dω α, τ ) + (β, Dω τ ) = (Dω α + δ ω β, τ ) ,

where we note that 21 (1 − ∗) is an orthogonal projection onto Ω2− (E) in which β
resides, so that β, 21 (1 − ∗) Dω τ = (β, Dω τ ). Thus,
T ∗ (α, β) = Dω α + δ ω β.
We now prove that Ker T ∗ = 0. Suppose that T ∗ (α, β) = 0, so that Dω α = −δ ω β.

Since ω ∈ C(P )+ (i.e., Ωω ∈ Ω2+ (E)) and β ∈ Ω2− (E) ⊆ Ω2+ (E) , we have
2
kDω αk = (−δ ω β, Dω α) = − (β, Dω Dω α) = − (β, [Ωω , α]) = 0.
Thus, T ∗ (α, β) = 0 ⇒ −δ ω β = Dω α = 0. Since we have assumed that ω is
weakly-irreducible, Dω α = 0 ⇒ α = 0. As β ∈ Ω2− (E),
δ ω β = 0 =⇒ Dω β = −Dω ∗ β = ∗(∗Dω ∗ β) = − ∗ δ ω β = 0
=⇒ ∆ω β = δ ω Dω β + Dω δ ω β = 0 =⇒ β = 0,
by Proposition 15.60, p. 443, which applies under our assumptions on ω and M .
Hence, Ker T ∗ = 0, and dim Ker T = index T .
16.3. LINEARIZATION OF THE MODULI SPACE OF SELF-DUAL CONNECTIONS 493

To compute index T , we use the cohomological index formula (13.4) of Part III,
p.315:
index T = Φ−1 ch(σ(T )) ` Td(T M ⊗ C) [M ]


= Φ−1 ch(IdE ⊗σ(T0 )) ` Td(T M ⊗ C) [M ]




= ch(E) ` Φ−1 (ch(σ(T0 ))) ` Td(T M ⊗ C) [M ] ,




where we recall that σ(T ) = IdE ⊗σ(T0 ), Φ denotes the Thom isomorphism, and
Td(T M ⊗ C) denotes the Todd class. The unitary frame bundle U (E) of E =
P ×G gC is reducible to an SU(N ) bundle (N = dim g), since the orthogonal (relative
to K) representation ad : G → SO(g) extends to adC : G → SU(gC ) which serves to
define E = P ×G gC . By the Remark 15.62, p. 445, we then have ch1 (E) = c1 (E) =
0, and then ch(E) = dim g + ch(E)2 . Note that
Φ−1 (ch(σ(T0 ))) ` Td(T M ⊗ C) 2 [M ] = index T0 .


We will save the verification that


Φ−1 (ch(σ(T0 ))) ` Td(T M ⊗ C) 0 = 2


for last, but assuming that this is correct, we have


(16.40) index T = 2ch2 (E) [M ] + (dim g) index T0 .
We now show that
(16.41) index T0 = − 21 (χ(M ) − sig(M )) .
By Hodge theory (see Equation (13.18) of Section 13.4 above, and Theorem 17.62,
p. 603, and Corollary 17.63 below for more details) the cohomology space H k (M ; R)
can be identified with the space of harmonic k-forms:
β ∈ H k (M ; R) ⇐⇒ ∆β = 0 ⇐⇒ dβ = 0 and δβ = 0.
Let bk := dim H k (M ; R) denote the k-th Betti number. It is easy to check that
∗ commutes with ∆. Thus H 2 (M ; R) is preserved by ∗ and splits into the ±1
eigenspaces of ∗, say H±2
(M ; R) of dimensions b± + −
2 , with b2 = b2 + b2 . For α ∈
2 2
H+ (M ; R) and β ∈ H− (M ; R), we have
2 2
α ∧ α = α ∧ ∗α = |α| ν, β ∧ β = −β ∧ ∗β = − |β| ν,
and
α ∧ β = −α ∧ ∗β = − hα, βi ν = − hβ, αi ν = −β ∧ ∗α = −β ∧ α =⇒ α ∧ β = 0,
− ∼
since 2-forms commute. Thus, sig(M ) = b+ 1 3
2 −b2 . Since ∗ : H (M ; R) −→ H (M ; R) ,
we have b3 = b1 , and

− 21 (χ(M ) − sig(M )) = − 12 b0 − b1 + b2 − b3 + b4 − b+

2 − b2

= − 12 2 − 2b1 + 2b− −
 
2 = b1 − 1 + b2 .

Thus, it suffices to prove that dim Ker T0 = b1 and dim Ker T0∗ = 1 + b−2 . We claim
1

Ker T0 = H 1 (M ; R). Since T0 (γ) = δγ, 2 (1 − ∗) dγ , it is clear that H 1 (M ; R) ⊆
Ker T0 . If γ ∈ Ker T0 , then δγ = 0 and (1 − ∗) dγ = 0. Hence,
0 = δ((1 − ∗) dγ) = δdγ − δ ∗ dγ = δdγ + ∗d ∗ ∗dγ = δdγ + ∗d2 γ = δdγ,
and
δdγ = 0 =⇒ 0 = (δdγ, γ) = (dγ, dγ) =⇒ dγ = 0.
494 16. GAUGE THEORETIC INSTANTONS

Thus, Ker T0 = H 1 (M ; R), and dim Ker T0 = b1 . We now prove that


Ker T0∗ = H 0 (M ; R) ⊕ H−
2
(M ; R) .
By a computation strictly analogous to (16.39), T0∗ (α, β) = dα + δβ. Thus, clearly
H 0 (M ; R) ⊕ H−
2
(M ; R) ⊆ Ker T0∗ . If (α, β) ∈ Ker T0∗ , then dα + δβ = 0 and
0 = δ(dα + δβ) = δdα + δ 2 β = δdα.
Hence, 0 = (δdα, α) = (dα, dα) and dα = 0, so that α ∈ H 0 (M ; R). Then
dα + δβ = 0 =⇒ δβ = 0,
and since β ∈ Ω2− (M, R),
0 = ∗δβ = − ∗ ∗d ∗ β = d ∗ β = −dβ.
Thus, β ∈ 2
H− and Ker T0∗ = H 0 (M ; R) ⊕ H−
(M ; R), 2
(M ; R).
−1
We now prove that the degree 0 part of Φ (ch(σ(T0 ))) ` Td(T M ⊗ C) is 2.
The degree 0 part of Td(T M ⊗ C) is 1 (directly from the definition in Section 13.1),
and so we need to show that the degree 0 part of Φ−1 (ch(σ(T0 ))) is 2. We use the
fact that the Thom class U ∈ H 4 (BM, SM ) and the Euler class χ(T M ) ∈ H 4 (M )
are related by
π ∗ χ(T M ) = i∗ U, where π : BM → M and i : (BM, ∅) → (BM, SM ).
For V := Λ1 (T ∗ M )C and F := Λ0 (T ∗ M )C ⊕ Λ2− (T ∗ M )C , we prove
(16.42) χ(T M ) ` Φ−1 (ch(σ(T0 )))0 = ch(V )2 − ch(F )2 .
We have the commutative diagram (where the coefficients are rational)

`χ(T M )
H 0 (M ) / H 4 (M )
∗,0 ∗,4
πB Φ πB
∗,0 ∗,4
πB,S πB,S
w  (  '
i∗,0 ∗,4
H 0 (BM ) o / H 4 (BM, SM ) i / H 4 (BM ).
`U
H 0 ([Link] )

Since Φ−1 lowers the degree by 4, Φ−1 ch(σ(T0 )) 0 = Φ−1 ch(σ(T0 ))2 . Thus,
 

∗,4
χ(T M ) ` Φ−1 (ch(σ(T0 )))0

πB
∗,4 ∗,0
Φ−1 (ch(σ(T0 )))0

= πB (χ(T M )) ` πB
∗,4 ∗,0
Φ−1 (ch(σ(T0 ))2 )

= πB (χ(T M )) ` πB
∗,0
= i∗,4 U ` i∗,0 πB,S Φ−1 (ch(σ(T0 ))2 )

 
∗,0
= i∗,4 U ` πB,S Φ−1 (ch(σ(T0 ))2 ) = i∗,4 ΦΦ−1 ch(σ(T0 ))2


∗,4
= i∗,4 (ch(σ(T0 ))2 ) = ch(πB
∗ ∗
V )2 − ch(πB F ) 2 = πB (ch(V )2 − ch(F )2 ) .
∗,4
Since M is homotopy equivalent to BM , we know that πB is an isomorphism, and
the identity (16.42) is proven.
To show Φ−1 (ch(σ(T0 )))0 = 2, it suffices to prove that
(16.43) ch(V )2 − ch(F )2 = 2χ(T M ) when χ(T M ) 6= 0.
This identity is conveniently proved by representing the characteristic classes in
terms of forms, as in Section 15.7. The Hermitian vector bundles, V = Λ1 (T ∗ M )C
and F = Λ0 (T ∗ M )C ⊕ Λ2− (T ∗ M )C , are complexifications of Riemannian bundles.
16.3. LINEARIZATION OF THE MODULI SPACE OF SELF-DUAL CONNECTIONS 495

Hence the principal unitary frame bundles of V and F reduce to principal orthogonal
frame bundles, say O(V ) and O(F ). When the curvature forms of the unitary
frame bundles are restricted to O(V ) and O(F ), they have values in spaces of
skew-symmetric real matrices (i.e., Lie algebras of orthogonal groups). Hence,
in verifying (16.43), it suffices to work with real skew-symmetric matrices when
checking the corresponding identity on the level of invariant polynomials. The
representation of SO(4) associated with T M is just the identity Id : SO(4) → SO(4).
T
The dual representation is B 7→ B −1 , which is also Id for B ∈ SO(4), hence the
representation of SO(4) associated with Λ1 (T ∗ M ) is also Id. The representation
for Λ0 (T ∗ M ) ⊕ Λ2− (T ∗ M ) is the direct sum of the trivial representation and the

representation SO(4) → GL(Λ2− (Rn∗ )) −→ GL(so− ). On the Lie algebra level, the
second representation can be expressed in terms of the ’t Hooft matrices η a ∈ so−
(a = 1, 2, 3) given in (16.4a), p. 467. For A ∈ SO(4), we compute
3
X 3
X 3
X
1 1 1
[A, η b ] = 4 (A · η c ) [η c , η b ] = 4 (A · η c )(−2εcba η a ) = 2 (A · η c ) εabc η a .
c=1 a,c=1 a,c=1
P3
Defining the matrix S(A) ∈ so(3) by [A, η b ] = a=1 S(A)ab η a , we have
3
X
1
S(A)ab = 2 (A · η c ) εabc .
c=1

Thus,
S(A)12 = 12 A · η 3 = A12 − A34 , S(A)23 = 1
2A · η 1 = A23 − A14 , and
1
S(A)31 = 2A · η 2 = A31 − A24 .
Hence for A ∈ so(4),
1 2
 1  2

2 tr −A − 2 tr −S(A)
X 2 2 2 2
= (Aij ) − (A12 − A34 ) − (A23 − A14 ) − (A31 − A24 )
i<j
X
1
= 2(A12 A34 + A23 A14 + A31 A24 ) = 4 εi1 i2 i3 i4 Ai1 i2 Ai3 i4 .
(i)

Thus,
  X
2
1
tr −A2 − 1 1

8π 2 8π 2 tr −S(A) = 2 32π 2 εi1 i2 i3 i4 Ai1 i2 Ai3 i4 ,
(i)

which yields (16.43), upon replacing A by the curvature form Ωθ of the Levi-Civita
connection θ (or actually any connection) on F M . 

Self-dual Connections on Principal SU(2)-Bundles are Irreducible. If


G = SU(2), by (16.26), p. 480, we have
c2 P ×SU(2) C2 S 4 = −k ⇐⇒ ch2 P ×SU(2) su(2) S 4 = 4k.
   
(16.44)
Also, for M = S 4 , we have χ(M ) = 2 and sig(M ) = 0. Thus, (16.37) becomes
(16.45) below. Moreover for k ≥ 1, in the proof of the following, we show that any
ω ∈ C(P )+ is irreducible.
496 16. GAUGE THEORETIC INSTANTONS

Corollary 16.18. Let P → S 4 be a principal SU(2)-bundle with


c2 P ×SU(2) C2 S 4 = −k, for k ≥ 1.
 

If ω is a self-dual connection on P (i.e., ω ∈ C(P )+ ), then ω is irreducible and


dim Tω C(P )+ ∩ Sω = 8k − 3,

(16.45)
in the notation of (16.36) for the formal tangent space of C(P )+ ∩ Sω at ω.
Proof. Once ω is shown to be irreducible, (16.45) follows from (16.37) and
(16.44). Suppose that the holonomy group Hol(ω, p) of such ω is a strict subgroup
of SU(2). The Lie algebra g0 of Hol(ω, p) has dimension at most 1, since two
independent elements of su(2) have a bracket outside of their span, and g0 cannot
be su(2) since then SU(2) = exp(g0 ) ⊆ Hol(ω, p). By Proposition 15.65, p. 457, the
restriction Ωω |P0 has values in g0 . Thus, dim g0 6= 0, since
Ωω |P0 = 0 =⇒ Ωω = 0 =⇒ −k = c2 P ×SU(2) C2 = 0.


Hence, dim g0 = 1, and the connected component of Hol(ω, p) containing Id is a


circle. Thus, ad : Hol(ω, p) → GL(g0 ) is trivial. Then, for g ∈ Hol(ω, p), we have
Rg∗ (Ωω |P0 ) = adg−1 Ωω |P0 = Ωω |P0 (i.e., Ωω |P0 is invariant under Hol(ω, p)). Hence,
Ωω |P0 = π ∗ α for a unique g0 -valued 2-form α on M . For a local section σ : S 4 → P0 ,
we have (by the Bianchi identity dΩω + [ω, Ωω ] = 0)
dα = d(σ ∗ π ∗ α) = σ ∗ d(Ωω |P0 ) = σ ∗ (dΩω |P0 ) = σ ∗ (− [ω, Ωω ] |P0 ) = 0,
since ω|P0 and Ωω |P0 have values in g0 which has a trivial bracket. Since ∗Ωω = Ωω ,
it follows that ∗α = α and δα = − ∗ d ∗ α = − ∗ dα = 0, so that α is a g0 -valued,
harmonic 2-form on S 4 . However, any such form on S 4 is zero, since H 2 (S 4 ) = 0.
Finally, by the transitivity of the SU(2) action on P and the equivariance of Ωω ,
α = 0 =⇒ Ωω |P0 = π ∗ α = 0 =⇒ Ωω = 0 =⇒ k = 0. 

4. Manifold Structure for Moduli of Self-dual Connections


Our Setting. Let C(P )+ denote the set of weakly-irreducible, self-dual connec-
tions on a principal G-bundle P over a compact, oriented, Riemannian 4-manifold
M . We only assume that G is semi-simple, C(P )+ 6= ∅, and that there are no
∆ω -harmonic forms in Ω2− (P ×G g) (e.g., M is self-dual with nonnegative scalar
curvature S 6= 0; see Proposition 15.60, p. 443). Let ω ∈ C(P )+ and recall from
(16.34)
n 1
o
(16.46) Sω := ω + τ : τ ∈ Ω (P, P ×G g) and δ ω τ = 0 .
In this section we will prove that (in a suitable sense) C(P )+ ∩ Sω is a submanifold
of the affine space C(P ) in a neighborhood of ω, and this submanifold has dimension
dim Tω C(P )+ ∩ Sω := 2ch(P ×G gC ) [M ] − 21 dim(G)(χ(M ) − sig(M )) ,


in the notation of Theorem 16.17. We introduce the space C(P )+ m ⊆ C(P )


+
of
+
mildly-irreducible self-dual connections, and show that C(P )m / GA(P ) can be made
into a Hausdorff topological space, such that each class [ω] ∈ C(P )+m / GA(P ) pos-
sesses a neighborhood U which is homeomorphic to a neighborhood of ω in Sω .
These homeomorphisms are shown to constitute an atlas which makes the quotient

space C(P )+m / GA(P ) a C manifold. In doing all of this, we need to state some
16.4. MANIFOLD STRUCTURE FOR MODULI OF SELF-DUAL CONNECTIONS 497

basic results about Lp -Sobolev Banach spaces of sections of vector bundles. For
p = 2, precise statements and proofs were given in our crash course, Chapter 7
(pp.193-205). For general p, proofs or references to proofs can be found in the
excellent survey articles [107] and [329], elaborated in the very useful textbook
[321]. We shall need these spaces in mildly nonlinear problems where it may be
advantageous to use several values of p at the same time (see the Sobolev Multi-
plication Theorem 16.24 below). Roughly speaking, if the nonlinearity involves a
power of a given section u ∈ W p,∗ (CM ), we shall use that then ua ∈ W p/a,∗ (CM ),
where the first Sobolev space is modeled over Lp (M ) and the second over Lp/a (M ).
In the following review of these results, M is a compact Riemannian n-manifold
with metric h, where n is not necessarily 4 until further notice.
Basic Results about Lp -Sobolev Banach Spaces of Sections of Vector
Bundles. Let E → M be a C ∞ Hermitian vector bundle, say E = P ×G W , for
a principal G-bundle π : P → M , and unitary representation ρ : G → U(W ). Let
θ denote the Levi-Civita connection on F M and let ω0 be a connection 1-form on
P . In the notation of Section 15.6, there is a connection ω0 ⊕ θ on P ×f F M (see
(15.81), p. 434), and a covariant derivative operator
∇ω0 ⊕θ : C ∞ (E ⊗ T r,s (M )) → C ∞ E ⊗ T r,s+1 (M ) .


For this, see (15.86), p. 435, with k = 1, and note that


0
C ∞ (E ⊗ T r,s (M )) ∼
= Ω (P ×f F M, W ⊗ T r,s ) , while
1
C ∞ E ⊗ T r,s+1 (M ) ∼ = Ω (P ×f F M, W ⊗ T r,s ) .


To shorten notation, we write ∇ω0 ⊕θ simply as ∇.


 For a nonnegative integer k and
p ∈ [1, ∞), and u ∈ C ∞ (E) = C ∞ E ⊗ T 0,0 (M ) , we set
  p1
X k Z
(16.47) kukp,k :=  |∇j u|p νh  ,
j=0 M

where νh denotes the volume element of the Riemannian metric h, and



 u, for j = 0,
∇j u :=

j times
 ∇ ◦ · · · ◦ ∇ u ∈ C ∞ E ⊗ T 0,j (M ) , for j > 0.


In (16.47), the numerical value |∇j u| of the j-fold composition is computed using
the Hermitian structure on E and the metric h on M .
Definition 16.19. The Sobolev space W p,k (E) is the completion of C ∞ (E)
with the norm k·kp,k .
Remark 16.20. The Banach space W p,k (E) coincides with the subspace of
L (E) := W p,0 (E) consisting of sections that have weak derivatives of orders ≤ k
p

in Lp . More precisely, u ∈ W p,k (E), if for each j ≤ k, there is vj ∈ Lp (E⊗T 0,j (M )),
such that for all w ∈ C ∞ (E ⊗ T 0,j (M )), we have
Z Z D E
j
hvj , wi νh = u,(∇∗ ) w νh ,
M M
where ∇∗ denotes the formal adjoint of ∇. We say that ∇j u = vj in the weak
(or distributional) sense.
498 16. GAUGE THEORETIC INSTANTONS

We have the following standard results (see [107], [329] and [321]). For Propo-
sitions 16.21-16.25 compare the corresponding L2 -modeled results in our Chapter
7 and note the new nonlinear feature of Proposition 16.24. For Theorems 16.26-
16.28 compare the corresponding Euclidean results in our Exercise 6.1 of Section
6.1 (p.157).

Proposition 16.21 (Sobolev Extension). Let D : C ∞ (E) → C ∞ (F ) be a linear


differential operator of order m (E, F Hermitian vector bundles over M ). For each
k ≥ 0, D has a unique continuous extension

Dp,k+m : W p,k+m (E) → W p,k (E) with kDp,k+m (α)kp,k ≤ K kαkp,k+m ,

for some K > 0, independent of α ∈ W p,k+m (E).

Proposition 16.22 (Compact Rellich Inclusion). Let C m (E) denote the Ba-
nach space of m-times (strongly) differentiable sections of E with norm
m
X
sup ∇j u x .

kukC m :=
j=0 x∈M

Let n = dim M , k ∈ {0, 1, 2, . . .} and p ∈ [1, ∞). For 0 ≤ m < k − np , we have a


compact (i.e., completely continuous) inclusion

(16.48) W p,k (E) ⊆ C m (E) .


n
For k > m and k − p > m − nq , we also have a compact inclusion

(16.49) W p,k (E) ⊆ W q,m (E).

Proposition 16.23 (Fundamental Elliptic Estimate). Assume that the differ-


ential operator D : C ∞ (E) → C ∞ (F ) is elliptic of order m, with formal adjoint
D∗ : C ∞ (F ) → C ∞ (E). Suppose that for some u ∈ Lp (E), we have v ∈ W p,k (F )
such that Z Z
hv, wi νh = hu, D∗ wi νh ,
M M

for all w ∈ C ∞ (F ) (i.e., Du := v exists weakly in W p,k (F )). Then u ∈ W p,k+m (E).
Moreover, for each k ≥ 0, there is a constant Ck > 0 independent of u, such that
 
(16.50) kukp,k+m ≤ Ck kDukp,k + kukp .

If Du ∈ C ∞ (F ), then for all k ≥ 0, kDukp,k < ∞, and we have kukp,k+m < ∞,


in which case u ∈ C ∞ (E) by (16.48). In particular, weak solutions of Du = 0 are
C ∞.

Proposition 16.24 (Sobolev Multiplication). Let Q : E1 × E2 → E3 be a


smooth, bilinear map of Riemannian (or Hermitian) vector bundles over M (com-
pact with dim(M ) = n). Then subject to the conditions below, the induced map
Q : C ∞ (E1 ) × C ∞ (E2 ) → C ∞ (E3 ) extends uniquely to a bounded bilinear map
Q : Lp1 ,k1 (E1 ) × Lp2 ,k2 (E2 ) → Lp3 ,k3 (E3 ), i.e.,

Q(x1 , x2 ) p3 ,k3
≤ C kx1 kp1 ,k1 kx2 kp2 ,k2
16.4. MANIFOLD STRUCTURE FOR MODULI OF SELF-DUAL CONNECTIONS 499

for C depending only on Q, the connections ∇E1 , ∇E2 , ∇E3 , and the Riemannian
metric on M . The conditions are that k3 ≤ min {k1 , k2 } and
    n o
 k1 − n + k2 − n if max k − n
, k − n
n p1 p2 o 1 p1 2 p2 o < 0
k3 − pn3 ≤ n
n n
 min k1 − , k2 −
p1 p2 if max k1 − p1 , k2 − pn2 > 0.
n

n o
If max k1 − pn1 , k2 − pn2 = 0, then it suffices that
n o    
k3 − pn3 < min k1 − pn1 , k2 − pn2 = k1 − pn1 + k2 − pn2 ,
where the inequality is strict. When p1 = p2 = p3 = p and k3 = k2 = j ≤ k = k1 >
n/p, we obtain the boundedness of
Q : W p,k (E1 ) × W p,j (E2 ) → W p,j (E3 ).
(Note that by (16.48) we have W p,k (E1 ) ⊆ C 0 (E1 ) for k > n/p, whence the induced
map Q may be defined pointwise.)
Proposition 16.25 (Elliptic Decomposition). Let D : C ∞ (E) → C ∞ (F ) be a
differential operator of order m with a symbol which is injective or surjective, and
let D∗ denote the formal adjoint of D. If D∗p,k+m : Lp,k+m (F ) → Lp,k (E) denotes
the Sobolev extension of D∗ , then we have the following direct sum decompositions
into closed subspaces for k ≥ 0 and p ≥ 2n/(2m + n) (e.g., p ≥ 2),
W p,k+m (E) = Ker(Dp,k+m ) ⊕ Im(D∗p,k+2m ) and
W p,k+m (F ) = Ker(D∗p,k+m ) ⊕ Im(Dp,k+2m ).
If the symbol of D is injective, then D∗ ◦ D is elliptic and
Ker(Dp,k+m ) = Ker(D) = Ker(D∗ ◦ D) ⊆ C ∞ (E)
is finite-dimensional. If the symbol of D is surjective, then D ◦ D∗ is elliptic and
Ker(D∗p,k+m ) = Ker(D∗ ) = Ker(D ◦ D∗ ) ⊆ C ∞ (F )
is finite-dimensional. In particular, if D is elliptic (i.e., with injective and surjective
symbol), then both Ker(D) and Ker(D∗ ) are finite-dimensional.
Given the estimate (16.50), this last result is not difficult to prove (e.g., see
[107]). Proposition 16.25 is indispensable in verifying the hypotheses of the follow-
ing implicit function theorems for Banach spaces in applications where the differ-
ential of the map is an elliptic operator.
Theorem 16.26 (Implicit Function Theorem I). Let B1 and B2 be Banach
spaces and let F : B1 → B2 be C k (1 ≤ k ≤ ∞). Assume that F∗x : Tx B1 → TF(x) B2
is a surjection and Tx B1 = Ker(F∗x ) ⊕ H, where H is closed. Then F −1 (F (x))
is a C k submanifold of B1 in a neighborhood of x, and its tangent space at x is
Ker(F∗x ).
We will also need the following version, whose proof is also in [272].
Theorem 16.27 (Implicit Function Theorem II). Let B1 , B2 and B3 be Banach
spaces and let F : B1 × B2 → B3 be a C k map (1 ≤ k ≤ ∞) with F (x1 , x2 ) =
x3 . Suppose that the partial derivative of F in the B2 -direction at (x1 , x2 ) (i.e.,
(D2 F )(x1 ,x2 ) : B2 → B3 ) . Then there are neighborhoods U1 of x1 and U2 of x2 ,
such that there is a unique C k map G : U1 → U2 whose graph {(x, G(x)) : x ∈ U1 }
500 16. GAUGE THEORETIC INSTANTONS

is F −1 (x3 ) ∩ U1 × U2 . Indeed, there is a neighborhood U3 of x3 and a unique C k


function H : U1 × U3 → U2 such that for all z ∈ U3 we have F (x, H(x, z)) = z.
When B1 = {0}, we obtain a special case of Theorem 16.27 which is usually estab-
lished beforehand, namely
Theorem 16.28 (Inverse Function Theorem). F : B2 → B3 be a C k map (1 ≤
k ≤ ∞) with F (x2 ) = x3 . If (DF )x2 : B2 → B3 is a bicontinuous isomorphism,
then there are neighborhoods U1 of x1 and U3 of x3 and a unique C k function
H : U3 → U2 , such that for all z ∈ U3 we have F (H(z)) = z.
Generalized Connection 1-Forms and Manifold Structure. In the rest
of this section, the assumptions of the opening paragraph are in force (e.g., dim M =
k
4). We set E := P ×G g and make the identification of the space Ω (P, g) (of
equivariant forms on P ) with Ωk (E). For any (smooth) ω ∈ C(P ), recall that
1
C(P ) = ω + Ω (P, g). Thus, we can identify C(P ) with the set of equivalence
1
classes of pairs in C(P ) × Ω (P, g), where
(16.51) (ω1 , α1 ) ≡ (ω2 , α2 ) :⇐⇒ ω1 + α1 = ω2 + α2 ⇐⇒ ω1 − ω2 = α2 − α1 .
1 ∼
Since Ω (P, g) −→ C ∞ (E ⊗ Λ1 (M )), say α ←→ α e, we can also regard C(P ) as the
set of equivalence classes of pairs in C(P ) × C ∞ (E ⊗ Λ1 (M )), where

(ω1 , α e2 ) :⇐⇒ (ω^


e1 ) ≡ (ω2 , α 1 − ω2 ) = α
e2 − α
e1 .

We define C(P )p,k as the set of equivalence classes of C(P )×W p,k E ⊗ Λ1 (M ) ,
where the equivalence relation is still defined by (16.51), and ω1 and ω2 are still
smooth, but α e1 and α e2 are not necessarily smooth elements of W p,k E ⊗ Λ1 (M )
(although α e2 − αe1 is smooth). One may wish to think of C(P )p,k as the set of all
generalized connection 1-forms  on P which differ from a smooth connection by an
element of W p,k E ⊗ Λ1 (M ) . We have only given a more precise meaning to this
thought. In what follows, for simplicity we will not adhere to the notation α e, but
simply use α. Note that each ω ∈ C(P ) defines a bijection
φω : C(P )p,k → W p,k E ⊗ Λ1 (M ) , where φω ([(ω, α)]) = α.


For ω1 , ω2 ∈ C(P ), φω1 ◦ φ−1 p,k



ω2 is a translation of W E ⊗ Λ1 (M ) , since

φω1 ◦ φ−1

ω2 (α) = φω1 ([(ω2 , α)])
= φω1 ([(ω1 , α + (ω2 − ω1 ))]) = α + (ω2 − ω1 ) .

serve as a collection of charts making C(P )p,k a C ∞ Banach



Thus, the φω ω∈C(P )
manifold.
Theorem 16.29 (Self-dual Slices). Let ω ∈ C(P )+ and assume that there are
no nonzero ∆ω -harmonic forms in Ω2− (E) (E = P ×G g). Let Sω be defined by
(16.46). For 2 ≤ k ∈ Z and p ∈ [1, ∞) with pk > 4, we have that C(P )+ ∩ Sω is a
C ∞ submanifold of C(P )p,k+1 .
Proof. Recall from (16.35) that
ω + τ ∈ C(P )+ ∩ Sω ⇐⇒ 0 = Q(τ ) := δ ω τ,(1 − ∗) Dω τ + 1

2 [τ, τ ] .
16.4. MANIFOLD STRUCTURE FOR MODULI OF SELF-DUAL CONNECTIONS 501

Thus, C(P )+ ∩ Sω = Q−1 (0) and we wish to apply Implicit Function Theorem
I (say, IFT I) to a Sobolev extension of Q. To this end, we first show that
Q : C ∞ E ⊗ Λ1 (M ) → C ∞ E ⊗ Λ0 (M ) ⊕ Λ2− (M ) has a C ∞ extension to
Qp,k+1 : W p,k+1 E ⊗ Λ1 (M ) → W p,k E ⊗ Λ0 (M ) ⊕ Λ2− (M ) ,
 
(16.52)
provided that p(k + 1) > 4; the stronger inequality pk > 4 will be used later. Note
that Q is a sum of the first-order elliptic differential operator τ 7→ (δ ω τ,(1 − ∗) Dω τ )
and the quadratic map τ 7→(1 − ∗) 12 [τ, τ ]. The differential operator has an exten-
sion to W p,k+1 E ⊗ Λ1 (M ) by Proposition 16.21. To show that the quadratic map
also extends, first note that, since p(k + 1) > 4, Proposition 16.24 yields a bounded
bilinear extension
W p,k+1 E ⊗ Λ1 (M ) × W p,k+1 E ⊗ Λ1 (M ) −→ W p,k+1 E ⊗ Λ2 (M )
  

of the bilinear function (τ1 , τ2 ) 7→ [τ1 , τ2 ]. Since τ 7→ (τ, τ ) is clearly bounded and
linear and (1 − ∗) is a differential operator of order 0 so that Proposition 16.21
applies, the composition
τ 7→ (τ, τ ) 7→ (1 − ∗) 12 [τ, τ ]
defines a bounded, quadratic map
W p,k+1 E ⊗ Λ1 (M ) → W p,k+1 E ⊗ Λ2− (M ) ⊆ W p,k E ⊗ Λ2− (M ) .
  

Thus, Q has the extension Qp,k+1 in (16.52). The differential(Qp,k+1 )∗0 of Qp,k+1
at τ = 0 is the linear part given for τ 0 ∈ W p,k+1 E ⊗ Λ1 (M ) by
 
(Qp,k+1 )∗0 (τ 0 ) := (δ ω )p,k+1 τ 0 ,(1 − ∗)(Dω )p,k+1 τ 0 .
This is just the Sobolev extension Tp,k+1 of the elliptic operator
T (τ 0 ) := δ ω τ 0 , 12 (1 − ∗) Dω τ 0


used in the proof of Theorem 16.17, p. 491. There we showed that Ker T ∗ = {0},
under the assumptions that ω is weakly-irreducible and the space of ∆ω -harmonic
forms in Ω2− (E) is trivial. With the goal of applying IFT I to Qp,k+1 , we wish to
use Proposition 16.25 with D = T ∗ to deduce that (Qp,k+1 )∗0 is onto, but we need
to show that T ∗ is elliptic first. Since T ∗ (α, β) = Dω α + δ ω β (see (16.39), p. 492),
the symbol of T ∗ is given by
σ(T ∗ )ξ (α, β) = αξ − β(ξ # , ·),
where ξ # is defined by ξ = h(ξ # , ·). If ξ 6= 0 and σ(T ∗ )ξ (α, β) = 0, then
2
0 = αξ(ξ # ) − β(ξ # , ξ # ) = α |ξ| =⇒ α = 0,
#
and so β(ξ , ·) = 0. This means that β is a sum of bicovectors in the 3-dimensional
orthogonal complement of ξ, in which case
2
0 = β ∧ β = −β ∧ ∗β = − |β| νh ,
since β ∈ Λ2− (T M ∗ ). Thus, σ(T ∗ )ξ is injective, and an isomorphism for dimensional
reasons. Hence T ∗ is elliptic, and (Qp,k+1 )∗ is onto by Proposition 16.25 with
D = T ∗ . Using the ellipticity of T , we have the splitting

W p,k+1 E ⊗ Λ1 (M ) = Ker T ⊕ Im Tp,k+2
 

of the domain of (Qp,k+1 )∗0 . Then we may then finally apply IFT I to deduce that
−1 
in a neighborhood U of ω, (Qp,k+1 ) (0) is a submanifold of W p,k+1 E ⊗ Λ1 (M )
502 16. GAUGE THEORETIC INSTANTONS

−1
of dim Ker T . Since C(P )+ ∩ U ⊆ (Qp,k+1 ) (0) ∩ U is clear, it remains to show
−1
that (Qp,k+1 ) (0) ∩ U ⊆ C(P )+ ∩ U , where U is possibly replaced by a smaller
neighborhood of ω. For τ ∈ W p,k+1 E ⊗ Λ1 (M ) with Qp,k+1 (τ ) = 0, we have
Tp,k+1 (τ ) = − 0,(1 − ∗) 21 [τ, τ ] ∈ W p,k+1 E ⊗ Λ1 (M ) ,
 

and then using (16.50),


   
1
kτ kp,k+2 ≤ C kTp,k+1 (τ )kp,k+1 + kτ kp = C 2 [τ, τ ] p,k+1
+ kτ kp
 
2
≤ C C 0 kτ kp,k+1 + kτ kp .
Replacing k by k + 1, k + 2,. . ., we have kτ kp,j < ∞ for all j. Thus, τ ∈
C ∞ E ⊗ Λ1 (M ) by (16.48) and ω + τ is a C ∞ self-dual connection. It remains to


show that U can be chosen so that the C ∞ elements of U are weakly-irreducible.


Definition 16.30. Let OPp,1 denote the Banach space of bounded
 (continuous
and linear) transformations from W p,k+1 (E) to W p,k E ⊗ Λ1 (M ) .
For τ ∈ C ∞ E ⊗ Λ1 (M ) , by Proposition 16.21, (Dω+τ )p,k+1 ∈ OPp,1 . Indeed,


for α ∈ W p,k+1 (E), we have


Dω+τ p,k+1 (α) = (Dω )p,k+1 (α) + [τ, α]

p,k p,k
ω
≤ (D )p,k+1 (α) + k[τ, α]kp,k+1
p,k
≤ C1 kαkp,k+1 + C2 kτ kp,k+1 kαkp,k+1
 
≤ C1 + C2 kτ kp,k+1 kαkp,k+1 ,
by Propositions 16.21 and 16.24. Indeed, this shows that the function
τ 7→ Dω+τ p,k+1 ∈ OPp,1


extends to a continuous affine map


Φ : W p,k+1 E ⊗ Λ1 (M ) → OPp,1 .


We know that Ker Dω : C ∞ (E) → C ∞ E ⊗ Λ1 (M ) = {0}, since we have as-




sumed that ω is weakly-irreducible. Moreover, Ker(Dω ) = Ker(δ ω Dω ) and δ ω Dω


is elliptic. Suppose that (Dω )p,k+1 (α) = 0, for α ∈ W p,k+1 (E). By Proposition
16.22, α ∈ C 1 (E) for 1 < k + 1 − p4 or kp > 4. Then (δ ω Dω )(α) = 0 weakly, since
for α ∈ C 1 (E) and β ∈ C ∞ (E), we can perform integration by parts:
Z Z Z

α,(δ ω Dω ) β νh = hα, δ ω Dω βi νh = hDω α, Dω βi νh = 0,
M M M

where the last equality holds since 0 = (Dω )p,k+1 (α) = Dω α for α ∈ W p,k+1 (E) ⊆
C 1 (E). Thus, for α ∈ W p,k+1 (E) with pk > 4, we may apply Proposition 16.23 to
obtain
(Dω )p,k+1 (α) = 0 =⇒ (δ ω Dω )(α) = 0 weakly
=⇒ α ∈ C ∞ (E) and (δ ω Dω )(α) = 0
2
=⇒ α ∈ C ∞ (E) and kDω (α)k = (δ ω Dω α, α) = 0
=⇒ α ∈ Ker(Dω ) =⇒ α = 0,
16.4. MANIFOLD STRUCTURE FOR MODULI OF SELF-DUAL CONNECTIONS 503

by the assumption that ω is weakly-irreducible. Thus, Ker(Dω )p,k+1 = {0}, and


so Φ(0) is in the open set, say OP×p,1 , of injective elements of OPp,1 . Since Φ is
continuous, there is some neighborhood U0 of 0 in W p,k+1 (E) such that Φ(τ ) ∈
OP×p,1 for τ ∈ U0 . Thus replacing U with U ∩ U0 , we are assured that the C

elements of U are weakly-irreducible. 

Theorem 16.31 (Local Slices). Let ω ∈ C(P ) be weakly-irreducible. For 1 ≤


p < ∞ and k ∈ Z with pk > 4, there are positive constants C1 and C2 , such
that for every τ ∈ W p,k+1 E ⊗ Λ1 (M ) with kτ kp,k+1 < C1 , there is a unique
σ(τ ) ∈ W p,k+2 (E) with kσ(τ )kp,k+2 < C2 and
(16.53) (δ ω )p,k+1 (Φ(Exp(σ(τ ))) ·(ω + τ ) − ω) = 0.
Here, the function (τ, s) 7→ Φ(Exp(s)) ·(ω + τ ) − ω, which has meaning for smooth 
pairs (τ, s), is proved to analytically extend to a function on W p,k+1 E ⊗ Λ1 (M ) ×
W p,k+2 (E) with values in W p,k+1 E ⊗ Λ1 (M ) , so that (16.53) makes sense. In
other words, Φ(Exp(σ(τ ))) ·(ω + τ ) ∈ Sωp,k+1 , where
n o
(16.54) Sωp,k+1 := ω + τ 0 ∈ C(P )p,k+1 : (δ ω )p,k+1 (τ 0 ) = 0 .

Moreover, τ 7→ σ(τ ) is a C ∞ function.


Proof. Let
Q : C ∞ E ⊗ Λ1 (M ) × C ∞ (E) → C ∞ (E)

(16.55)
be defined by
Q(τ, s) := δ ω (Φ(Exp(s)) ·(ω + τ ) − ω) .

As in (15.30), p. 413, let f = Exp(s) ∈ C(P, G) ⊆ C P, GL CN , where we
continue to assume that G is a matrix group. According to (15.27), p. 411, we have
Φ(f ) · ω = f d f −1 + f ωf −1 , and so


Φ(Exp(s)) ·(ω + τ ) = Exp(s) d(Exp(−s)) + Exp(s)(ω + τ ) Exp(−s) .


For A, B ∈ g, we will use the following identity:

X 1 i
Exp(A) B Exp(−A) = (adA ) (B) ,
i=0
i!

where adA (B) = [A, B]. This follows from


1 i
X 1
(adA ) B = An BAm , i ≥ 0,
i! m+n=i
n!m!

which can be proved by straightforward induction. Similarly, we have


1 i
X 1 n m
(ads ) (ds) = − s d(s ) ,
(i + 1)! m+n=i+1
n!m!

for s ∈ C(P, g) ∼
= C ∞ (E), and this yields

X 1 i
Exp(s) d(Exp(−s)) = − (ads ) (ds) , and
i=0
(i + 1)!
504 16. GAUGE THEORETIC INSTANTONS

Φ(Exp(s)) · (ω + τ ) = Exp(s) d(Exp(−s)) + Exp(s)(ω + τ ) Exp(−s)


∞ ∞
X 1 i
X 1 i
= − (ads ) (ds) + (ads ) (ω + τ )
i=0
(i + 1)! i=0
i!
∞ ∞ ∞
X 1 i
X 1 i−1
X 1 i
= − (ads ) (ds) + ω − (ads ) ([ω, s]) + (ads ) (τ )
i=0
(i + 1)! i=1
i! i=0
i!
∞ ∞
X 1 i
X 1 i
= ω− (ads ) (ds + [ω, s]) + (ads ) (τ )
i=0
(i + 1)! i=0
i!
∞ ∞
X 1 i
X 1 i
= ω− (ads ) (Dω s) + (ads ) (τ ) .
i=0
(i + 1)! i=0
i!
Hence for p(k + 1) > 4, by repeated use of Proposition 16.24,
∞ ∞
X 1 i
X 1 i
(τ, s) 7→ Φ(Exp(s)) ·(ω + τ ) − ω = − (ads ) (Dω s) + (ads ) (τ ) ,
i=0
(i + 1)! i=0
i!
extends to a C ∞ (indeed, analytic) map
W p,k+1 E ⊗ Λ1 (M ) × W p,k+2 (E) → W p,k+1 E ⊗ Λ1 (M ) .
 

Composing this map with (δ ω )p,k+1 gives us an extension of Q in (16.55), say


Qp,k+1 : W p,k+1 E ⊗ Λ1 (M ) × W p,k+2 (E) → W p,k E ⊗ Λ1 (M ) .
 

The derivative of Qp,k+1 at (0, 0) is


 
(Qp,k+1 )∗(0,0) (τ 0 , s0 ) = (δ ω )p,k+1 − (Dω )p,k+2 (s0 ) + τ 0 .
As we wish to apply Implicit Function Theorem II (IFT II) to Qp,k+1 , we note
that the partial derivative D2 (Qp,k+1 )(0,0) of Qp,k+1 in the W p,k+2 (E) direction
is the Sobolev extension (−∆ω )p,k+2 of the formally self-adjoint elliptic operator
−∆ω = −δ ω Dω . By Proposition 16.25,
W p,k (E) = Ker(−∆ω ) ⊕(−∆ω )p,k+2 W p,k+2 (E)


= (−∆ω )p,k+2 W p,k+2 (E) ,




since ω weakly-irreducible implies that Ker(−∆ω ) = Ker Dω = 0. Thus, our


D2 (Qp,k+1 )(0,0) = (−∆ω )p,k+2 is onto with trivial kernel, and hence is a bicon-
tinuous isomorphism by the Open Mapping Theorem. Thus, the IFT II applies to
give us the constants C1 and C2 , and the C ∞ function σ. 
(Small) Gauge Equivalence Close to Weakly-Irreducible Connections.
Theorem 16.31 provides us with a local slice for the action of a Sobolev extension
of GA(P ) on the space C(P )p,k+1 . In other words, every connection ω 0 ∈ C(P )p,k+1
in a suitably small neighborhood of ω ∈ C(P ) is gauge-equivalent to a connec-
tion in Sωp,k+1 via some generalized gauge transformation Exp s, where kskp,k+2
p,k+2
small. Of course, one would like to precisely define a Sobolev extension GA(P )
of GA(P ). Since GA(P ) (or equivalently C(P, G)) is not the space of sections
of a vector bundle, some elaboration is needed.  We have assumed
 that G is a
matrix group, say a Lie subgroup of GL CN . Now GL CN (and hence G) is
16.4. MANIFOLD STRUCTURE FOR MODULI OF SELF-DUAL CONNECTIONS 505


contained in the vector space gl CN of all linear endomorphisms of CN . Thus,
0 
C(P, G) (and hence GA(P )) can identified
 with a subset of Ω P, gl CN , where
the representation G → GL gl CN is the adjoint representation (i.e., g · A =
0
gAg −1 , g ∈ G and A ∈ gl CN ). Now Ω P,gl CN ∼
= C ∞ P ×G gl CN
  

which has a Sobolev extension W p,k P ×G gl CN . Thus, it makes sense to define


p,k 
C(P, G) to be the closure in W p,k P ×G gl CN of the subset corresponding
0 
to C(P, G) ⊆ Ω P, gl CN  . Since (by Proposition 16.22) we have a continuous
inclusion W p,k P ×G gl CN ⊆ C 0 P ×G gl CN for pk > 4 , it follows that
p,k
C(P, G) consists of continuous, ad-equivariant, G-valued functions on P . Note
p,k p,k
that C(P, G) may then be identified with a certain set, say GA(P ) , of con-
tinuous gauge transformations (i.e., continuous (as opposed to C ∞ ), equivariant,
p,k
fiber-preserving homeomorphisms of P ). A proof that (for pk > dim M ) GA(P )
p,k
is actually a Lie group, modeled on the Banach space W (E) in such a way that
p,k
Exp : W p,k (E) → GA(P ) is a local diffeomorphism, can be found in [304] for
the case p = 2, but their proof works for pk > dim M as well. Although as a map
between Banach spaces, σ in Theorem 16.31 is a C ∞ function of τ , we still need to
prove that σ(τ ) is C ∞ if τ is C ∞ as the next result states.

Theorem 16.32 (Gauge Regularity). In Theorem 16.31, σ(τ ) is C ∞ if τ is


C ∞ , provided C1 is chosen small enough.

Proof. From the proof of Theorem 16.31, we know that s := σ(τ ) ∈ W p,k+2 (E)
obeys the equation

0 = (δ ω )p,k+1 (Φ(Exp(s)) ·(ω + τ ) − ω)


= (δ ω )p,k+1 (Exp(s) d(Exp(−s)) + Exp(s)(ω + τ ) Exp(−s) − ω)
∞ ∞
!
ω
X 1 i

ω
 X 1 i
= (δ )p,k+1 − (ads ) (D )p,k+2 (s) + (ads ) (τ )
i=0
(i + 1)! i=0
i!
P∞   !
ω 1 i ω
− (D ) (s) − (ads ) (D ) (s)
= (δ ω )p,k+1 p,k+2
P∞ i
i=1 (i+1)! p,k+2
.
+ i=0 i!1 (ads ) (τ )

Thus,
  
(∆ω )p,k+2 (s) = (δ ω )p,k+1 (Dω )p,k+2 (s) = (δ ω )p,k+1 R s,(Dω )p,k+2 (s) , τ ,

where
∞ ∞

ω
 X 1 i

ω
 X 1 i
R s,(D )p,k+2 (s) , τ = − (ads ) (D )p,k+2 (s) + (ads ) (τ ) .
i=1
(i + 1)! i=0
i!

Note that

   X 1 i
 
(δ ω )p,k+1 R s,(Dω )p,k+2 (s) , τ = − (ads ) (∆ω )p,k+2 (s)
i=1
(i + 1)!
 
+ P s,(δ )p,k+1 (s) ,(Dω )p,k+2 (s) , τ,(δ ω )p,k+1 (τ ) ,
ω
506 16. GAUGE THEORETIC INSTANTONS

where P is a power series in its arguments and contains  terms which


 result when
ω i i ω
(δ )p,k+1 is passed through the factor (ads ) in (ads ) (D )p,k+2 (s) . Thus,


!
X 1 i
1+ (ads ) (∆ω )p,k+2 (s)
i=1
(i + 1)!
 
= P s,(δ ω )p,k+1 (s) ,(Dω )p,k+2 (s) , τ,(δ ω )p,k+1 (τ ) .
For arbitrarily small C2 in Theorem 16.31, the constant C1 exists. By (16.48), we
may then assume that C2 is small enough so that kskp,k+2 < C2 ⇒ kskC 0 is small
enough so that the endomorphism

!
X 1 i
Ψ(s) := 1 + (ads ) : W p,k (E) → W p,k (E)
i=1
(i + 1)!
can be inverted. Thus,
 
−1
(∆ω )p,k+2 (s) = Ψ(s) P s,(δ ω )p,k+1 (s) ,(Dω )p,k+2 (s) , τ,(δ ω )p,k+1 (τ ) .
−1
Since the Wp,k norm of Ψ(s) is finite, as well as the Wp,k norms of all of the argu-
ments of Q (recall that τ ∈ C ∞ E ⊗ Λ1 (M ) ), it follows that (∆ω )p,k+2 (s)

<
p,k
∞. Hence by (16.50), s ∈ W p,k+4 (E), and repeating the argument with k replaced
by k + 2, etc., yields s ∈ C ∞ (E). 
Obstructions to (Large) Gauge Equivalences. We have proved that every
C ∞ connection in a sufficiently small neighborhood (in C(P )p,k+1 , pk > 4) of a
weakly-irreducible ω ∈ C(P ) is gauge equivalent, via a unique small (i.e., close to Id
in GA(P )p,k+2 ) gauge transformation which is necessarily smooth, to a connection
in Sω . We have yet to introduce hypotheses which will enable us to prove that no
two smooth connections near ω in Sω are gauge equivalent by a possibly large gauge
transformation. For this we need a stronger condition on ω than weak-irreducibility.
If we insist that ω be irreducible, then at some later point, we would have to face
the problem of proving that every connection near an irreducible connection is
irreducible. While this is quite believable, there seems to be no simple proof. To
get around the difficulty, we introduce a milder irreducibility condition which is still
stronger than weak-irreducibility. To this end, note that the adjoint representation
ad : G → GL(g) induces a representation r = ad ⊗ ad∗ : G → GL(End g) given by
r(g)(h) = adg ◦ h ◦ adg−1 . There is an inner product k ⊗ k ∗ on End g induced from
k (minus the Killing form) on g, and r is orthogonal with respect to k ⊗ k ∗ . Now
End g = H1 ⊕ H1⊥ , where H1 denotes the invariant subspace
(16.56) H1 := {h ∈ End(g) : r(g) (h) = h, for all g ∈ G} .
Proposition 16.33. Suppose that G is connected and semi-simple. If adg0 ∈
H1 for some g 0 ∈ G, then g 0 ∈ Z(G) := the center of G.
Proof. For all g ∈ G, we have adg0 = r(g)(adg0 ) = adg ◦ adg0 ◦ adg−1 or
adgg0 g−1 g0−1 = Id. We claim that gg 0 g −1 g 0−1 ∈ Z(G). Indeed, since G is connected,
any g 00 ∈ G can be written as exp A. Using the fact that exp ◦adg = Adg ◦ exp for
all g ∈ G, we then have
Adgg0 g−1 g0−1 (g 00 ) = Adgg0 g−1 g0−1 (exp A) = exp adgg0 g−1 g0−1 A = exp A = g 00 .

16.4. MANIFOLD STRUCTURE FOR MODULI OF SELF-DUAL CONNECTIONS 507

Thus, gg 0 g −1 g 0−1 ∈ Z(G). Since G is semi-simple, Z(G) is discrete. It follows that


gg 0 g −1 g 0−1 is the identity of G, because g (and hence gg 0 g −1 g 0−1 ) can be connected
to the identity by a path in G. Thus, gg 0 = g 0 g and g 0 ∈ Z(G) . 
Definition 16.34. We call ω ∈ C(P ) mildly-irreducible if there is no nonzero
element of H1⊥ , defined in (16.56), which is fixed by all transformations r(g0 ) as g0
ranges over the holonomy group Hol(ω, p) of ω.
Clearly, if ω is irreducible (i.e., Hol(ω, p) = G), then it is mildly-irreducible
(i.e., Z(Hol(ω, p)) is discrete). Moreover, we have:
Proposition 16.35. Let π : P → M be a principal G-bundle, where G is con-
nected and semi-simple. Let ω ∈ C(P ). For E := P ×G g, let End(E) = End1 (E) ⊕

End1 (E) denote the orthogonal decomposition arising from End g = H1 ⊕ H1⊥ in
(16.56). We have ω mildly-irreducible, if and only if
  
⊥ ⊥
Ker Dω : C ∞ (End1 (E) ) → Ω1 End1 (E) = 0.
Moreover, if ω is mildly-irreducible, then
1. Z(Hol(ω, p)) = Z(G).
2. The isotropy group Iω of ω is {Rg : g ∈ Z(G)}.
3. ω is weakly-irreducible (see Definition 16.16, p. 490).

Proof. Suppose that Dω α = 0 for α ∈ C ∞ (End1 (E) ), where ω is mildly-
irreducible. Then α (regarded as in C(P, H1⊥ )) is constant on the holonomy bundle
P0 of ω through some p0 ∈ P (see Section 15.15.8). Since α is equivariant as well, we
have (for g0 ∈ Hol(ω, p) = holonomy group of ω) r(g0−1 )α(p0 ) = α(p0 g0 ) = α(p0 ),
whence α(p0 ) ∈ H1⊥ is fixed by all members of r(Hol(ω, p)). Thus, α(p0 ) = 0 since
ω is mildly-irreducible, and so α = 0. Conversely, if ω is not mildly-irreducible,
then there is 0 6= h ∈ H1⊥ invariant under r(Hol(ω, p)). Let α denote the constant
function h on P0 , and note that α extends by equivariance to a nonzero element of
C(P, H1⊥ ) in Ker Dω .
For (1), suppose that gc ∈ Z(Hol(ω, p)). Then adgc ∈ End(g) is invariant under
all elements of r(Hol(ω, p)), since for all A ∈ g and g0 ∈ Hol(ω, p),
(r(g0 )(adgc ))(A) = adg0 ◦ adgc ◦ adg−1 (A)
 0
= g0 gc g0−1 Ag0 gc −1 g0−1 = gc Agc −1 = adgc (A) .
The same is true of the projection of adgc onto H1⊥ . Hence, ω mildly-irreducible
implies that adgc ∈ H1 . Then gc ∈ Z(G) by Proposition 16.33. For (2) note that
any z ∈ Z(G), Rz : P → P is a gauge transformation which acts trivially on C(P ),
since Rz · ω = Rz∗−1 ω = adz ω = ω for all ω ∈ C(P ). By Proposition 15.64, p. 456,
the homomorphism Iω → G given by Φ(f ) 7→ f (p0 ) maps the isotropy group Iω
isomorphically onto the centralizer Z(Hol(ω, p)). Thus, for ω mildly-irreducible, it
follows from Z(Hol(ω, p)) = Z(G) that Iω = {Rz : z ∈ Z(G)}. For (3), note that
since Z(Hol(ω, p)) = Z(G) is discrete by (1), ω is weakly-irreducible by Proposition
16.15, p. 490. 
In distinction to Proposition 16.31, the following result essentially states that
near a mildly-irreducible connection ω, the set
n 1
o
Sω = ω + τ : τ ∈ Ω (P, P ×G g) and δ ω τ = 0
508 16. GAUGE THEORETIC INSTANTONS

serves as a global slice. In other words, no two distinct connections in Sω near a


mildly-irreducible ω are gauge equivalent by a possibly large gauge transformation.
Some key ideas we use in the proof were inspired by [39].
Theorem 16.36 (Global Slices). Let ω ∈ C(P ) be mildly-irreducible. Then
there is a constant C > 0, such that if kω − ω 0 kp,k+1 ≤ C for ω 0 ∈ C(P ), then ω 0 is
mildly-irreducible. Also, if pk > 4, ω1 , ω2 ∈ Sω with kω − ωi kp,k+1 ≤ C (i = 1, 2),
p,k+2
⊆ W p,k+2 P ×G End CN with ω2 = F −1∗ ω1 , then

and there is F ∈ GA(P )
F ∈ Iω = {Rz : z ∈ Z(G)}, and hence ω2 = ω1 .
Proof. By Proposition 16.35, we know that ω 0 ∈ C(P ) is mildly-irreducible,
if and only if
 0  
⊥ ⊥
Ker Dω : C ∞ (End1 (E) ) → Ω1 End1 (E) = 0.
0
Using Proposition 16.24, note that ω 0 7→ Dω extends to a C ∞ map
⊥ ⊥
C(P )p,k+1 → L(W p,k+1 (End1 (E) ), W p,k (End1 (E) ⊗ Λ1 (M )),
where L(V1 , V2 ) := bounded linear maps from V1 to V2 . Since the subset of injective
bounded linear maps is open, we see that the mildly-irreducible connections form
an open subset of C(P )p,k+1 . This proves the first assertion of Theorem 16.36.
Let τ = ω2 − ω1 ∈ W p,k+1 E ⊗ Λ1 (M ) , and let
p,k+2
⊆ W p,k+2 P ×G gl CN

f ∈ C(P, G)
p,k+2
correspond to F −1 ∈ GA(P ) (i.e., f = Φ F −1 or F (p) = pf (p) for all p ∈ P ).


Note that 2 < (k + 2) − 4/p for pk > 4, whence f is C 2 by (16.48), p. 498. By


(15.27), p. 411, we have
ω2 = F −1∗ ω1 ⇐⇒ ω2 = f −1 df + f −1 ω1 f
= f −1 (df + ω1 f ) = f −1 (df + ω1 f − f ω1 ) + ω1
(16.57) ⇐⇒ τ = ω2 − ω1 = f −1 (df + [ω1 , f ]) = f −1 (Dω1 f ) .
The idea is to prove that if C is sufficiently small (so that kτ kp,k+1 will also be
small), then F will be close to some Rz (z ∈ Z(G)) or equivalently z −1 f will be
close enough to the constant map (to the identity Id of G) so that we can apply
Theorem 16.31 to conclude (since
 ω1 , ω2 ∈ Sω ) that Rz−1 ◦ F = Id (or F = Rz ).
Since Ker Dω1 |End1(E)⊥ = 0, we know that the positive self-adjoint elliptic
operator    
⊥ ⊥
∆ω1 = δ ω1 Dω1 : C ∞ End1 (E) → C ∞ End1 (E)
 

has a positive smallest eigenvalue λ. Thus, for η ∈ C ∞ End1 (E) ,
2 2
(16.58) kDω1 ηk = (∆ω1 η, η) ≥ λ kηk .
Recall from (16.57) that τ = f −1 (Dω1 f ) ∈

W p,k+1 E ⊗ Λ1 (M ) , where f ∈
p,k+2
C(P, G) . Via the representation ad : G →O(g) ⊂ End(g), we may asso-
p,k+2
ciate to each f ∈ C(P, G) , some fEnd ∈ W p,k+2 (End(E)). Note also that
ad : g → End(g) induces a map
W p,k+1 E ⊗ Λ1 (M ) → W p,k+1 End(E) ⊗ Λ1 (M ) .
 
16.4. MANIFOLD STRUCTURE FOR MODULI OF SELF-DUAL CONNECTIONS 509

Applying this to both sides of the equation τ = f −1 (Dω1 f ), we obtain an equation,


say
−1
τEnd = f −1 (Dω1 f ) End = fEnd (Dω1 fEnd ) ,


where both sides are in W p,k+1 End(E) ⊗ Λ1 (M ) . According to

End(E) = End1 (E) ⊕ End1 (E) ,
we can orthogonally decompose
f −1 (Dω1 f ) End = f −1 (Dω1 f ) End1 + f −1 (Dω1 f ) End⊥ .
  
1

Then, pointwise,
2 2
2 2
f −1 (Dω1 f ) = |(Dω1 f )End | ≥ (Dω1 f )End⊥

|τEnd | = End
.
1

Upon integrating this, (16.58) yields


2 2
2
(16.59a) kτEnd k ≥ (Dω1 f )End⊥ ≥ λ fEnd⊥
1
.
1

From Proposition 16.22, we know that there is a constant K so that


(16.60a) max(|τEnd |) ≤ K kτEnd kp,k+1 .
 
2 D E
ω1
Since d fEnd⊥
1
= 2 f ⊥
End1 , D fEnd1 ,

 
2  
ω1
d fEnd⊥
1
≤ 2 D f End⊥
1
fEnd⊥
1

= 2 |τEnd | fEnd⊥
1
≤ K kτEnd kp,k+1 fEnd⊥
1
.
 
2
Hence, for kτ kp,k+1 sufficiently small, we can assure that d fEnd⊥1
<  on

M. This implies that fEnd⊥


1
can be made arbitrarily close to a constant function
on M as kτ kp,k+1 → 0. In fact, by (16.59a) and (16.60a), we know this constant
function must be zero. Consequently, fEnd = fEnd1 +fEnd⊥
1
can be made arbitrarily
0
C close to the subbundle End(E)1 . By Proposition 16.33 and the definition of
p,k+2
fEnd , it follows that f ∈ C(P, G) can be made arbitrarily C 0 close to the
constant map P → {z} for some z ∈ Z(G) for kτ kp,k+1 sufficiently small. Then
p,k+2
zf −1 ∈ C(P, G) is C 0 close to the constant map with value I ∈ G. Thus,
−1 p,k+2
zf = Exp s for some s ∈ C(P, G) , and
  −1  −1
(Exp s) · ω1 = zf −1 · ω1 = zf −1 d zf −1 + zf −1 ω1 zf −1
 

= f −1 df + f −1 ω1 f = F · ω1 = ω2 .
By the same computation as in the proof of Theorem 16.31, the equation
(Exp s) · ω1 − ω1 = ω2 − ω1 = τ
can be expanded to yield

!
X 1
1+ adns Dω1 s = τ.
n = 1
(n + 1)!
510 16. GAUGE THEORETIC INSTANTONS

Since max |s| can be made arbitrarily small for kτ kp,k+1 sufficiently small. Thus,
for sufficiently small kτ kp,k+1 , |ads | will be small enough so that the series can be
inverted, say

!−1 ∞
X 1 n
X
1+ ads = 1+ cn adns
n=1
(n + 1)!
P∞ n=1
for constants cn . We then have Dω1 s = (1 + n=1 cn adns )τ , and by Proposition
16.24
X∞
n
(16.61) kDω1 skp,j ≤ (1 + cn K n kskp,j ) kτ kp,k+1 ,
n=1
for each j ∈ {0, . . . , k + 1} and some constant K > 0, independent of s and τ . As
a consequence of the definition of Sobolev spaces, we have
 
(16.62) kskp,j+1 ≤ C kDω1 skp,j + kskp,0 .
As kτ kp,k+1 → 0, we have kskp,0 → 0 (since then max |s| → 0 and M has finite
volume). By (16.61), as kτ kp,k+1 → 0, we then have kDω1 skp,0 → 0, in which case
kskp,1 → 0 by (16.62). We then see inductively that if kskp,j → 0 as kτ kp,k+1 → 0,
then kskp,j+1 → 0 as kτ kp,k+1 → 0 for j ∈ {0, . . . , k + 1}. Thus, kskp,k+2 → 0 as
kτ kp,k+1 → 0. Hence for kτ kp,k+1 sufficiently small, we will have kskp,k+2 small
enough so that if Exp s · ω1 = ω2 , then Exp s = Id by Theorem 16.31 since ω1 and
ω2 are in Sω . 
Use of Local Slices to Introduce an Atlas on M and to Determine
+
dim M+ . Let C(P )m denote the set (assumed nonvoid) of mildly irreducible self-
dual C ∞ connections on P . We now show how to use the local slices Sω , ω ∈
+ +
C(P )m to introduce a C ∞ atlas on M+ := C(P )m / GA(P ) in such a way that the
topology induced on M+ by the atlas is Hausdorff and dim(M+ ) = 2ch(E)[M ] −
dim G(χ(M ) − sig(M )) found in Theorem 16.17, p. 491.
+ +
For each ω ∈ C(P )m , we choose a neighborhood of ω, say Uω ⊆ Sω ∩ C(P )m ,
such that the function Uω → M+ given by
ω 0 7→ [ω 0 ] := {F · ω 0 : F ∈ GA(P )}
is injective; Uω exists by Theorem 16.36. Let [Uω ] denote the image of Uω in M+ ,
and let φω : [Uω ] → Uω denote the inverse. We know from Theorem 16.29, that we
p,k+1
can (and do) choose Uω small enough so that Uω is a C ∞ submanifold of C(P ) .
To show that the φω constitute an atlas, we need to prove that
φω ◦ φ−1
ω 0 : φω 0 ([Uω ] ∩ [Uω 0 ]) → φω ([Uω ] ∩ [Uω 0 ])
is C ∞ if [Uω ] ∩ [Uω0 ] is nonempty. Select ω 00 ∈ Uω0 , such that [ω 00 ] ∈ [Uω ]. Then,
there is F ∈ GA(P ), such that F · ω 00 ∈ Uω . The map ω̃ 7→ F · ω̃ is a C ∞
p,k+1
diffeomorphism of C(P ) , as one can verify using Proposition 15.29, p. 411, and
p,k+1
Proposition 16.24. Hence F · Uω0 is a C ∞ submanifold of C(P ) containing
00 00
F · ω ∈ Uω . A neighborhood of F · ω in F · Uω0 will be contained in the ball
of radius C1 about ω in Theorem 16.31, provided we choose the Uω small enough.
Then Theorem 16.31 provides us with a C ∞ map ω + τ 7→ Exp(σ(τ ))(ω + τ ) which
will carry this neighborhood of F · ω 00 smoothly into Uω , proving that φω ◦ φ−1 ω 0 is
C ∞ at the arbitrary ω 00 ∈ φω0 ([Uω ] ∩ [Uω0 ]). Note that Theorem 16.32 is needed to
ensure that if ω + τ is C ∞ , then Exp(σ(τ )) given by Theorem 16.31 is in GA(P )
16.4. MANIFOLD STRUCTURE FOR MODULI OF SELF-DUAL CONNECTIONS 511

(i.e., Exp(σ(τ )) is C ∞ ). The topology on M+ is the smallest topology that makes


the maps φω continuous. To show that M+ is Hausdorff, select ω and ω 0 with
[ω] 6= [ω 0 ], and set τ = ω 0 − ω ∈ Ω1 (E). We argue as in [39] by first noting that
[ω] 6= [ω 0 ] means that for all f ∈ C(P, G) ∼
= GA(P )
0 6= f −1 · ω − ω 0 = f −1 df + f −1 ωf − ω + ω − ω 0
(16.63) = f −1 (df + ωf − f ω) + ω − ω 0 = f −1 Dω f − τ.
As in the proof of Theorem 16.36, we can consider the fEnd ∈ C ∞ (End E) (E :=
P ×G g) associated with f ∈ C(P, G); and τEnd ∈ Ω1 (End E) associated with
τ ∈ Ω1 (E). If we can show that
−1
(16.64) fEnd Dω (fEnd ) − τEnd 2,0
>  > 0,

as f varies over C(P, G), then f −1 · ω − ω 0 2,0 will be bounded away from 0. For
pk > 4 (indeed p(k + 1) > 2), by (16.49) we would then have f −1 · ω − ω 0 p,k+1
bounded away from 0. Thus, the orbits of ω and ω 0 will be bounded away from each
other in C(P )p,k+1 , and we will have that M+ is Hausdorff. In order to establish
(16.64), we introduce
Dω − τEnd : C ∞ (End E) → Ω1 (End E) given by
(Dω − τEnd )(fEnd ) := Dω (fEnd ) − fEnd τEnd , and

∆ω,τ := (Dω − τEnd ) (Dω − τEnd ) : C ∞ (End E) → C ∞ (End E).
Now, since fEnd is an isometry of End E, we have
−1
fEnd Dω (fEnd ) − τEnd 2,0
= kDω (fEnd ) − fEnd τEnd k2,0

= k(Dω − τEnd )(fEnd )k2,0 ≥ λ fEnd 2,0
,

where fEnd denotes the projection of fEnd onto the orthogonal complement of K :=
Ker(∆ ) and λ denotes the smallest positive eigenvalue of ∆ω,τ . We need to show
ω,τ

that fEnd 2,0
is bounded away from 0. Suppose, on the contrary, that there is a

sequence fn ∈ GA(P ), such that k(fn )End 2,0 k → 0. Since (fn )End is an isometry,

k(fn )End k2,0 is constant, and so (fn )End − (fn )End is a bounded sequence in K, with
dim K < ∞ since ∆ω,τ is elliptic. Thus, by passing to a subsequence, we may

assume that (fn )End − (fn )End converges. Calling the limit h∞ ∈ K, we have
  2 2
2 ⊥ ⊥
kh∞ − (fn )End k2,0 = h∞ − (fn )End − (fn )End + (fn )End → 0,
2,0 2,0
2
whence (fn )End → h∞ in L (End E). We remark that h∞ = (f∞ )End for some
f∞ ∈ GA(P ). To see this, note that ad : G → O(g) has discrete kernel Z(G).
Thus, if (f )End = (f 0 )End , then f 0 = zf for some z ∈ Z(G). Hence, fn may be
replaced by zn fn which converges in L2 to some f∞ ∈ GA(P ) for which (f∞ )End =
−1 ω
limn→∞ (zn fn )End = limn→∞ (fn )End = h∞ . Thus, Dω f∞ −f∞ τ = 0 or f∞ D f∞ −
τ = 0 in violation of (16.63). In summary, we have proved the following.
Theorem 16.37 (Main Theorem). Let π : P → M be a principal G-bundle
with G compact and semi-simple over a compact, oriented, self-dual, Riemannian
4-manifold with scalar curvature S ≥ 0 and S 6= 0. Then the space C(P )+
m / GA(P )
512 16. GAUGE THEORETIC INSTANTONS

of moduli of mildly-irreducible self-dual connections has (if nonempty) the structure


of a Hausdorff C ∞ manifold of dimension
1
2ch(P ×G gC ) [M ] − 2 dim(G)(χ(M ) − sig(M )) .
Remark 16.38. a) If
n o
+ +
C0 (P )m := ω ∈ C(P )m : Ker ∆ω : Ω2− (P ×G gC ) ←- = {0}


+
is nonempty, then the same conclusion holds for C0 (P )m / GA(P ), and we may drop
the self-duality and positive scalar curvature assumptions on M . It is likely a proof
+ +
that C0 (P )m = C(P )m for a generic class of metrics on M can be constructed along
the lines found [148].
b) For this chapter, we read, distilled, reworked and verified the literature with
one single goal: a faithful and comprehensible description of Donaldson’s truly
unprecedented application of index theory. Correspondingly, for the result that the
moduli space is a manifold, we restrict ourselves to the theory until about 1982 and
assume self-duality and positive scalar curvature. Also today, these assumptions
seem natural in the physics context (and for the ADHM construction of Theorem
16.14). For topologists, this may not be really on point. A modern reader should not
be confused about that: Taubes’ proof that all definite 4-manifolds admit instan-
tons, (and later versions where there is an obstruction bundle to gluing instantons)
together with Uhlenbeck’s generic metrics theorem are the analytic results that
make the theory topologically useful. Actually, we can omit the generic metrics the-
orem by perturbing the equations by some compact perturbation like Donaldson
initially did, and like it is done in Seiberg-Witten theory, see our Theorem 18.32
(p.662), to ensure the map to 2-forms has 0 as a regular value. The interested reader
might look at how Kronheimer and Mrowka deal with genericity in their recent
papers [269, 267, 268]. (We are indebted to P. Kirk for these considerations.)
With more space, we would have liked to describe the general 4-manifold situation,
and, perhaps even how manifolds with boundary enter the scene.
CHAPTER 17

The Local Index Theorem for Twisted Dirac


Operators

Synopsis. Clifford Algebras and Spinors: Clifford Algebra Basics; Spin Groups and
Double Cover; Spinor Representations; Supertrace. Spin Structures and Twisted Dirac
Operators: Čech Cohomology; Admittance of Spin Structures; Standard and Twisted
Dirac Operators; Chirality. The Spinorial Heat Kernel: Index, Spectral Asymmetry and
the Existence of the Heat Kernel; Solving the Spinorial Heat Equations; Calculating In-
dex and Supertrace; General Heat Kernels. The Asymptotic Formula for the Heat Kernel:
Why Asymptotic Expansion? The Radial Gauge; About the Geometry of the Ball; Fur-
ther Approximations. The Local Index Formula: Content and Meaning of the Local Index
Formula; How the Curvature Terms Arise in the Heat Asymptotics; The case m = 1 (sur-
faces); The case m = 2 (4-manifolds); Proof of the Local Index Formula for Arbitrary Even
Dimensions; Index Theorem for Twisted Dirac Operators; A b Genus; Rokhlin’s Theorem.
The Index Theorem for Standard Geometric Operators: Index Theorem for Generalized
Dirac Operators; Twisted Generalized Dirac Operators; The Hirzebruch Signature For-
mula; The Chern-Gauss-Bonnet Formula; The Generalized Yang-Mills Index Theorem;
The Hirzebruch-Riemann-Roch Formula for Kähler Manifolds.
I One of our main goals in this chapter, will be to show that the classical geometric
operators such as the signature operator, the de Rham operator, the Dolbeault operator
and even the Yang-Mills operator can all be locally expressed in terms of twisted Dirac
operators. The index of any of these operators (and their twists) can then be obtained
from the Local Index Theorem for twisted Dirac operators which is proved in unusual
detail. This theorem supplies a globally defined n-form on M , whose integral is the index
of an operator which is perhaps only locally of the form of a twisted Dirac operator, as with
the classical geometric operators. This n-form (or index density) is expressed in terms of
curvature forms of characteristic classes. The Index Theorem thus obtained then becomes
a formula that relates a global invariant quantity, namely the index of an operator, to the
integral of a local quantity involving curvature. This is in the spirit of the Gauss-Bonnet
Theorem which is a special case. J

1. Clifford Algebras and Spinors

I The concept of spinors and spin groups is loaded with comprehensive physical and
mathematical meaning. One is tempted to attribute a general mysterious significance to
spinors. Perhaps rightly so. For the learner, however, we shall give a rather formal and,
hopefully, pleasantly unexciting rigorous introduction. J

Clifford Algebra Basics. Let V be a real vector space with a symmetric


positive-definite inner product h·, ·i and induced norm k·k. Think of V as the
513
514 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

tangent space of a Riemannian manifold in a given point endowed with the corre-
sponding Euclidean metric.
Definition 17.1. The Clifford algebra C`(V ) is the real algebra generated
by V and R with the relation
vw + wv = −2 hv, wi , for all v, w ∈ V.
Note that the inner (Euclidean) product h·, ·i of V induces an inner product for
C`(V ) by the determinant like in Equation (6.3), p.171. The product of v and w in
2
C`(V ) is denoted by the plain juxtaposition vw. Also, v 2 := vv = − hv, vi = − kvk ,
and vw = −wv if hv, wi = 0. In the following, we take {e1 , . . . , en } to be an
orthonormal basis of V .
Example 17.2. If dim V = 1, then e21 := e1 e1 = −1, and C`(V ) is isomorphic
to the algebra C of complex numbers, via α0 + α1 e1 7→ α0 + iα1 , for α0 , α1 ∈ R.
Example 17.3. If dim V = 2, then it is easy to check that
α0 + α1 e1 + α2 e2 + α3 e1 e2 7→ α0 + α1 i + α2 j + α3 k
defines an isomorphism C`(V ) with the algebra H of quaternions. Note that
2
(e1 e2 ) = e1 e2 e1 e2 = −e1 e1 e2 e2 = −1, and (e1 e2 ) e1 = −e21 e2 = e2 , etc.
Example 17.4. For dim V = 3, one can check that there is an isomorphism

C`(V ) −→ H ⊕ H determined by
e1 7→ (−i, i) , e2 7→ (−j, j) , e3 7→ (−k, k) .

Example 17.5. If dim V = 4, we have an isomorphism, C`(V ) −→ H(2) := the
algebra of 2 × 2 quaternionic matrices, determined by
       
0 i 0 j 0 k 0 −1
(17.1) e1 7→ , e2 7→ , e3 7→ , e4 7→ .
i 0 j 0 k 0 1 0
For dim V = n, we write C`(V ) =: C`n . In [273], it is shown that there is the
following table of algebra isomorphisms
n 0 1 2 3 4 5 6 7 8
(17.2)
C`n ∼
= R C H H ⊕ H H(2) C(4) R(8) R(8) ⊕ R(8) R(16)
Here R(k), C(k) and H(k) denote the algebras of k × k matrices with entries in
R, C and H respectively. Moreover, it is also proven that there is periodicity

relation C`n+8 −→ C`8 ⊗ C`n = R(16) ⊗ C`n , so that this table can be extended
indefinitely. The case of nondegenerate indefinite inner products with signature
(r, s) is also handled in [273]; we have only considered (n, 0).
Let Λ• (V ) = ⊕nk=1 Λk (V ) denote the exterior algebra of V . While Λ• (V ) is not
isomorphic to C`(V ) as an algebra, there is a linear isomorphism of vector spaces

(17.3) L : Λ• (V ) −→ C`(V ) determined by
L(ei1 ∧ · · · ∧ eik ) := ei1 · · · eik , (i1 < · · · < ik ) ,
where we continue to let {e1 , . . . , en } be an orthonormal basis of V . It can be
shown that L is O(n)-equivariant and independent of the choice of orthonormal
basis. Moreover via L, the natural inner product on Λ• (V ) gives us an inner
product and norm on C`(V ). There is an exponential map
exp : C`(V ) −→ PC`(V )
∞ 1 k ,
x 7→ k=0 k! x
17.1. CLIFFORD ALGEBRAS AND SPINORS 515

k
which converges, since xk ≤ ck kxk for some constant c depending on n but
not on x. Indeed, for x, y ∈ C`(V ), each of the 2n components of xy (relative
n
q {e1 · · · ek : i1 < · · · < ik }) can be no larger than 2 kxk kyk, and so
to the basis
2
kxyk ≤ 2n (2n kxk kyk) = 23n/2 kxk kyk.
We define the bracket (or commutator)
 of any x, y ∈ C`(V ) , by [x, y] =
xy − yx. The linear subspace L Λ2 (V ) is closed under bracket, since
[ei ej , eh ek ] = ei ej eh ek − eh ek ei ej = ei ej eh ek + eh (ei ek + 2δik ) ej
= ei ej eh ek + eh ei ek ej + 2δik eh ej
= ei ej eh ek −(ei eh + 2δih ) ek ej + 2δik eh ej
= ei ej eh ek − ei eh ek ej − 2δih ek ej + 2δik eh ej
= ei ej eh ek + ei eh (ej ek + 2δjk ) − 2δih ek ej + 2δik eh ej
= ei ej eh ek + ei eh ej ek + 2δjk ei eh − 2δih ek ej + 2δik eh ej
= ei ej eh ek − ei (ej eh + 2δjh ) ek + 2δjk ei eh − 2δih ek ej + 2δik eh ej
= −2δjh ei ek + 2δjk ei eh − 2δih ek ej + 2δik eh ej .
For A = (aij ) , B = (bij ) ∈ so(n) (i.e., the Lie algebra of antisymmetric n × n
matrices), we have (where we implicitly sum over all indices)
[aij ei ej , bhk eh ek ] = aij bhk [ei ej , eh ek ]
= −2aij bhk δjh ei ek + 2δjk aij bhk ei eh − 2δih aij bhk ek ej + 2δik aij bhk eh ej
= −2(AB)ik ei ek − 2(AB)ih ei eh + 2(BA)kj ek ej + 2(BA)hj eh ej
= −4 [A, B]ij ei ej .
Thus, h i
X X X
− 41 aij ei ej , − 41 bhk eh ek = − 14 [A, B]ij ei ej ,
i,j h,k i,j
which implies that

X
c0 : L Λ2 (V ) given by c0 − 41
 
(17.4) −→ so(n), aij ei ej := A
i,j
 ∼
is an isomorphism L Λ2 (V ) −→ so(n) of Lie algebras.

Spin Groups and Double Cover. We begin with a geometric and explicit
construction of the spin groups.
Definition 17.6. a) We define the spin group,
Spin(n) := exp L Λ2 (V ) ,


where V is any n-dimensional Euclidean space. By definition, Spin(n) inherites the


structure of a Lie group.
b) We denote the corresponding Lie algebra by spin(n) and note
spin(n) = L Λ2 (V ) = so(n).


Before showing that


c : Spin(n) → SO(n), given by c(exp x) := exp(c0 (x))
is a well defined double cover, we consider some examples.
516 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS


Example 17.7. For C`2 , L Λ2 R2 = {te1 e2 : t ∈ R} , and

X 1 k
exp(te1 e2 ) = (te1 e2 )
k!
k=1
2 3
= 1 + te1 e2 + 12 t2 (e1 e2 ) + 61 t3 (e1 e2 ) + · · ·
= 1 − 21 t2 + · · · + t − 61 t3 + · · · e1 e2 = cos(t) + sin(t) e1 e2 .


Thus, Spin(2) = {cos(t) + sin(t) e1 e2 : t ∈ R} .



Example 17.8. For C`3 , L Λ2 R3 = {a1 e2 e3 + a2 e3 e1 + a3 e1 e2 : a1 , a2 , a2 ∈ R}.
If a : = a1 e2 e3 + a2 e3 e1 + a3 e1 e2 , then
2
a2 = (a1 e2 e3 + a2 e3 e1 + a3 e1 e2 )
2 2 2
= a21 (e2 e3 ) + a22 (e3 e1 ) + a23 (e1 e2 ) + a1 a2 (e2 e3 e3 e1 + e3 e1 e2 e3 ) + · · ·
2
= − a21 + a22 + a23 = − kak .


sin t
Thus, with t = 1 − 16 t2 + · · · (analytic), it follows that

sin(kak)
exp(a) = cos(kak) + a − 16 a3 + · · · = cos(kak) +

a, and
kak
 X3 
Spin(3) = α0 + α1 e2 e3 + α2 e3 e1 + α3 e1 e2 : αk2 = 1 ,
k=0

which may be regarded as the 3-sphere of unit quaternions.

Example 17.9. To exhibit Spin(4), it is convenient to utilize the duality de-


composition Λ2 R4 = Λ+ ⊕ Λ− . Then L Λ2 R4 = L(Λ+ ) ⊕ L(Λ− ) and we


set

a := a1 12 (e2 e3 + e1 e4 ) + a2 12 (e3 e1 + e2 e4 ) + a3 21 (e1 e2 + e3 e4 ) ∈ L Λ+ ,




b := b1 12 (e2 e3 − e1 e4 ) + b2 21 (e3 e1 − e2 e4 ) + b3 21 (e1 e2 − e3 e4 ) ∈ L Λ− .





Under the isomorphism F : C`4 −→ H(2), determined by (17.1), it is easy to check
that
 
a1 i+a2 j + a3 k 0
F (a + b) = ,
0 b1 i+b2 j + b3 k
and so
 
exp(a1 i+a2 j + a3 k) 0
F (exp(a + b)) =
0 exp(b1 i+b2 j + b3 k)
" sin(kak)
#
cos(kak) + kak a 0
= sin(kbk) .
0 cos(kbk) + kbk b


Thus, we have F : Spin(4) −→ S 3 × S 3 . To delineate Spin(4) itself, first note that
for v4 = e1 e2 e3 e4 ,
 
−1 0
F (v4 ) = F (e1 e2 e3 e4 ) = .
0 1
17.1. CLIFFORD ALGEBRAS AND SPINORS 517

Hence, Spin(4) consists of all elements of C`4 of the form

exp(a + b) = F −1 (F (exp(a + b)))


sin(kak) sin(kbk)
= 12 (1 − v4 ) cos kak + a+ 21 (1 + v4 ) cos kbk + b
kak kbk
sin(kak)
= 12 (cos kak + cos kbk) + a
kak
sin(kbk)
+ b + 12 (cos kbk − cos kak) v4 .
kbk

Note that here exp(a) exp(b) = exp(a + b) = exp(a) + exp(b)!

The vector representation or double cover c : Spin(n) → SO(n)is defined


by means of the next result. Here dim(V ) = n, and the spaces L Λ1 (V ) ⊂ C`(V )
are identified.

Proposition 17.10. For g ∈ Spin(n) and v ∈ V = L Λ1 (V ) , let

(17.5) c(g)(v) := gvg −1 ∈ C`(V ) .



Then c(g)(v) ∈ L Λ1 (V ) = V . Also, c(g) ∈ SO(n) := SO(V ) and

(17.6) c : Spin(n) → SO(n).

is a double covering homomorphism (universal for n ≥ 3). Moreover, for c0 defined


as in (17.4), we have

(17.7) c(exp a) = exp(c0 (a)) ,

and so c0 is the Lie algebra homomorphism for c.


 
Proof. We write g = exp(a) = exp − 14 i,j aij ei ej for a ∈ L Λ2 (V ) . To
P 

show that gvg −1 ∈ V , it suffices to verify that

(17.8) exp(ta) v exp(−ta) = exp(tA) v,

where A = c0 (a) ∈ so(n). Since each side of (17.8) is a C`(V )-valued power series
in t with infinite radius of convergence, we need only check that all derivatives of
both sides agree at t = 0; i.e.,

dk
(17.9) dtk
(exp(ta) v exp(−ta)) = Ak (v) , k = 0, 1, 2, . . . .
t=0

In verifying this, we will use the identity

ei ej ek − ek ei ej = ei (−ek ej − 2δkj ) − ek ei ej = −ei ek ej − 2δkj ei − ek ei ej


= −(−ek ei − 2δki ) ej − 2δkj ei − ek ei ej
= 2(δki ej − δkj ei ) .
518 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

At t = 0,
X
d
dt (exp(ta) v exp(−ta)) = [a, v] = − 14 aij [ei ej , v]
i,j
X X
= − 41 aij [ei ej , vk ek ] = − 14 aij vk (ei ej ek − ek ei ej )
i,j,k i,j,k
X X
= − 41 aij vk 2(δki ej − δkj ei ) = − 12 (aij vk δki ej − aij vk δkj ei )
i,j,k i,j,k
X X
= − 21 (aij vi ej − aij vj ei ) = aij vj ei = A(v) .
i,j i,j

At arbitrary t,
d d
dt (exp(ta) v exp(−ta)) = du (exp((t + u) a) v exp(−(t + u) a))|u=0
d
= exp(ta) du (exp(ua) v exp(−ua))|u=0 exp(−ta)
= exp(ta) A(v) exp(−ta) .

Hence,
dk
dtk
(exp(ta) v exp(−ta)) = exp(ta) Ak (v) exp(−ta) ,

and evaluating both sides at t = 0 yields (17.9).


Now, c(g) ∈ SO(n), since
2 2
− kc(g) vk = − gvg −1 = gvg −1 gvg −1
 

2 2
= gvvg −1 = − kvk g −1 g = − kvk .

Thus, c(Spin(n)) ⊆ SO(n). Note that (17.7) does in fact hold by (17.8) with
t = 1. Since Spin(n) and SO(n) are connected and c0 : spin(n) → so(n) is an
isomorphism, it follows that c(Spin(n)) is the connected component of I ∈ SO(n),
namely SO(n) itself, and c : Spin(n) → SO(n) is a covering homomorphism. Since
= Z2 for n > 2, it follows from covering space theory that π1 (Spin(n)) ∼
π1 (SO(n)) ∼ =
π1 (SO(n)) / Ker c. Then for n > 2, Spin(n) is the universal, (simply-connected)
covering space of SO(n) if ±1 ∈ Ker c. Certainly, 1 ∈ Ker c, and if −1 ∈ Spin(n),
−1
then −1 ∈ Ker c, since c(−1)(v) = −1v(−1) = v. Thus, it remains to check that
−1 ∈ Spin(n), but this is immediate from exp(te1 e2 ) = cos(t) + sin(t) e1 e2 with
t = π. For n = 2, c : Spin(2) → SO(2) is still a double cover, since

c(exp(te1 e2 ))(e1 ) = exp(te1 e2 ) e1 exp(−te1 e2 )


= (cos(t) + sin(t) e1 e2 ) e1 (cos(t) − sin(t) e1 e2 )
= cos2 (t) e1 − cos(t) sin(t) e1 e1 e2
+ sin(t) cos(t) e1 e2 e1 − sin2 (t) e1 e2 e1 e1 e2
cos2 (t) − sin2 (t) e1 + 2 cos(t) sin(t) e2

=
= cos(2t) e1 + sin(2t) e2 .

However, π1 (SO(2)) ∼
= Z so that the covering is not universal for n = 2. 
17.1. CLIFFORD ALGEBRAS AND SPINORS 519

Spinor Representations. Besides the vector representation c : Spin(n) →


SO(n), there are fundamental spinor representations, which we will describe. Since
the index of an elliptic differential operator on a closed, odd-dimensional manifold
is always 0, for simplicity we assume that n is even, say n = 2m. Then there is a
unique (up to equivalence) irreducible representation (homomorphism of algebras
over R)
ρ : C`2m → End(Σ2m ) ,
where End(Σ2m ) denotes the algebra of C-linear endomorphisms of some complex
vector space Σ2m , the elements of which are called spinors. Here irreducible means
that Σ2m has no proper subspace which is invariant under all operators in ρ(C`2m ).
In the following we give an explicit construction of Σ2m and ρ.
Let h·, ·i denote the standard Hermitian inner product on Cm given by
Xm
hz, wi = zk wk .
k=1
m n 2m
We identify C with R = R , and for w ∈ Cm , we have the C-linear ext function
w∧ : Λk (Cm ) −→ Λk+1 (Cm ) ,
given by α 7→ w∧α for α ∈ Λk (Cm ). Moreover, like in Remark 6.21, p.173, there is
a C-linear int function
wx : Λk (Cm ) → Λk−1 (Cm ) , defined via
k
X j+1
wx(v1 ∧ · · · ∧ vk ) := (−1) hvj , wi v1 ∧ · · · ∧ vbj ∧ · · · ∧ vk ,
j=1

where vbj means that the factor vj is omitted. All that was discussed in Section
6.4, p.171 in the real case. While wx is C-linear, the function Cm → End(Λ• (Cm ))
given by w 7→ wx is R-linear (but C-conjugate linear). Similarly to the real case
explained in Equation (6.3), p.171, we define a Hermitian inner product h·, ·i on
Λk (Cm ) induced by that on Cm , such that {ei1 ∧ · · · ∧ eik : i1 < · · · < ik } is an
orthonormal basis for Λk (Cm ) if e1 , . . . , em is an orthonormal basis for Cm . Relative
to this inner product, wx and w∧ are adjoints, since for all v1 , · · · , vk , u1 , · · · , uk−1 ∈
Cm we have
hv1 ∧ · · · ∧ vk , w∧ u1 ∧ · · · ∧ uk−1 i
k
X j+1
= (−1) hvj , wi hv1 ∧ · · · ∧ vbj ∧ · · · ∧ vk , u1 ∧ · · · ∧ uk−1 i
j=1
X 
k j+1
= (−1) hvj , wi v1 ∧ · · · ∧ vbj ∧ · · · ∧ vk , u1 ∧ · · · ∧ uk−1
j=1

= hwx(v1 ∧ · · · ∧ vk ) , u1 ∧ · · · ∧ uk−1 i .
Proposition 17.11. Let ρ1 : Cm → End(Λ• (Cm )) be given by
ρ1 (w)(α) := (w∧ − wx)(α) = w∧ α − wxα.
Then ρ1 uniquely extends to an R-linear homomorphism
ρ : C`2m → End(Λ• (Cm ))
of algebras over R.
520 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

Proof. Using w∧ w∧ α = 0 and wx(wxα) = 0, we obtain


(ρ1 (w) ◦ ρ1 (w))(α) = w∧ (w∧ α − wxα) − wx(w∧ α − wxα)
= w∧ w∧ α − w∧ (wxα) − wx(w∧ α) + wx(wxα)
= −w∧ (wxα) − wx(w∧ α)
(17.10) = − hw, wi α,
where in the last equality we used
wx(w∧ (v1 ∧ · · · ∧ vk ))
Xk j+1
= hw, wi v1 ∧ · · · ∧ vk − (−1) hvj , wi w∧ v1 ∧ · · · ∧ vbj ∧ · · · ∧ vk
j=1
= hw, wi v1 ∧ · · · ∧ vk − w∧ (wx(v1 ∧ · · · ∧ vk )) .
Generalizing (17.10), it follows that
ρ1 (w1 ) ◦ ρ1 (w2 ) + ρ1 (w2 ) ◦ ρ1 (w1 )
= −(hw1 , w2 i + hw2 , w1 i) Id = −2< hw1 , w2 i Id .
Since < hw1 , w2 i is the standard inner product on R2m ∼
= Cm , ρ1 extends uniquely
to an R-linear homomorphism
ρ : C`2m → End(Λ• (Cm ))
of algebras over R. 
Let Cl2m := C ⊗R C`2m denote the complex Clifford algebra. One of our goals
is to prove that the complex linear extension of ρ, say
ρC : Cl2m → End(Λ• (Cm )) ,
is an isomorphism of algebras over C. For this and other reasons, it is convenient
to introduce more notation. Let (f1 , · · · , fm ) be an orthonormal basis of Cm , then
(17.11) (e1 , · · · , e2m ) := (f1 , if1 , · · · , fm , ifm )
is an oriented, orthonormal basis of R2m , and the complex volume element, →
Clifford analysis is
(17.12) ωC := im e1 · · · e2m ∈ Cl2m ;
this is independent of the choice of oriented orthonormal basis of R2m . We have
m
ωC2 = (−1) e1 · · · e2m e1 · · · e2m
m 2m (2m−1)+(2m−2)+···+1
= (−1) (−1) (−1)
m+2m(2m−1)/2 m+m(2m−1) 2m2
= (−1) = (−1) = (−1) = 1, and so
2
ρC ωC2

(17.13) ρC (ωC ) = = ρC (1) = Id .
Since ρ1 (w) = w∧ − wx is the difference between an operator and its adjoint, ρ1 (w)
is skew-adjoint. We define an array of skew-adjoint operators
γj := ρC (ej ) = ρ(ej ) , for j ∈ {1, . . . , 2m} .
In harmony with the physics literature, we set γn+1 = γ2m+1 := γ1 · · · γn , so that
ρC (ωC ) = im γ2m+1 . The case n = 4 abounds in physics, as does γ5 = γ1 γ2 γ3 γ4 , but
due to the indefiniteness of the metric in relativity, γ42 = −γi2 for i = 1, 2, 3. We
will remain in the Euclidean category. Using the fact that the γj are skew-adjoint,
17.1. CLIFFORD ALGEBRAS AND SPINORS 521

it is straightforward (but tedious) to show directly that ρC (ωC ) is self-adjoint. But,


we can easily get this from the fact that ρC (ωC ) is clearly either self-adjoint or
skew-adjoint, and, unlike ρC (ωC ), the square of a skew-adjoint transformation has
2
nonpositive eigenvalues (squares of pure imaginaries). Since ρC (ωC ) = Id, the
eigenvalues of ρC (ωC ) are ±1. As ρC (ωC ) is self-adjoint, the eigenspaces of ρC (ωC ),
say
• − •
(17.14) Σ+ m m
2m := (ρC (ωC ) + Id) Λ (C ) and Σ2m := (ρC (ωC ) − Id) Λ (C ) ,

are orthogonal. Using the notation (17.11),


ρ(e1 e2 ) = (e1∧ − e1 x) ◦(e2∧ − e2 x)
 
= f1∧ − f1 x ◦ if1∧ − (if1 ) x
 
= f∧ − f1 x ◦ if1∧ + i(f1 x)
= i(f1∧ ) ◦(f1 x) − i(f1 x) ◦(f1∧ )

= i (f1∧ ) ◦(f1 x) − (f1 x) ◦(f1∧ ) ,
and in general

ρ(e2j−1 e2j ) = i (fj∧ ) ◦(fj x) −(fj x) ◦(fj∧ ) .
Thus, the action of ρ(e2j−1 e2j ) on fj1 ∧ fj2 ∧ · · · ∧ fjl is given by
ρ(e2j−1 e2j )(fj1 ∧ fj2 ∧ · · · ∧ fjl )

i(fj1 ∧ fj2 ∧ · · · ∧ fjl ) , if jk = j for some k,
(17.15) =
−i(fj1 ∧ fj2 ∧ · · · ∧ fjl ) , if jk =
6 j for all k.

Hence, when restricted to Λl (Cm ), ρC (ωC ) is


m−l l
im ρ(e1 · · · e2m ) |Λl(Cm ) = im il (−i) Id = (−1) Id , and so
M
Σ+ ev m
2m = Λ (C ) := Λl (Cm ) , while
l even
M
Σ−
2m = Λ
odd
(Cm ) := Λl (Cm ) .
l odd

Using this (or the easy fact that for j ∈ {1, . . . , 2m} , γ2m+1 ◦ γj = −γj ◦ γ2m+1 ), we

have γj Σ± ∓ ± ∓

2m = Σ2m , and indeed γj : Σ2m −→ Σ2m with inverse −γj . Moreover,
±
for j, k ∈ {1, . . . , 2m} , the spaces
 Σ2m± are each invariant under the compositions
γj ◦ γk , and so ρ(spin(2m)) Σ± 2m ⊂ Σ2m . For j 6= k,

(γj ◦ γk ) = γk∗ ◦ γj∗ = −γk ◦ −γj = −γj ◦ γk and
Tr(γj ◦ γk ) = − Tr(γk ◦ γj ) = − Tr(γj ◦ γk ) .
Thus, the elements of ρ(spin(2m)) are skew-adjoint and traceless, and so we see
ρ(Spin(2m)) ⊂ SU(Λ• (Cm )). In summary, ρ : Spin(2m) → SU(Λ• (Cm )) is the
orthogonal direct sum of two special unitary half-spinor or chiral representations
ρ± : Spin(2m) → SU Σ±

(17.16) 2m .

Definition 17.12. Let


Σ2m := Λ• (Cm ) and π ± := 12 (ρC (ωC ) ± Id) : Σ2m → Σ±
2m .
522 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

The supertrace of an endomorphism A ∈ End(Σ2m ) is


     
Str(A) := Tr π + ◦ A|Σ+ − Tr π − ◦ A|Σ−
2m 2m

= Tr(A ◦ ρC (ωC )) = im Tr(A ◦ γ2m+1 ) .


The following result will be crucial in evaluating the local index density of the
twisted Dirac operator.
Proposition 17.13. For k ∈ {1, . . . , 2m} with j1 , j2 , . . . , jk distinct, we have
Tr(γj1 γj2 · · · γjk ) = 0, and
Str(γj1 γj2 · · · γjk ) = im Tr(γj1 γj2 · · · γjk γ2m+1 )

0, if k < 2m,
= m
(−2i) εj1 ...j2m , if k = 2m.
The 2n endomorphisms consisting of Id and those γj1 γj2 · · · γjk with j1 < j2 < · · · <
jk , k ∈ {1, . . . , 2m}, form a basis of End(Σ2m ).
Note that we write the short γi γj instead of the long γi ◦ γj . Recall from (3.13)
(p.96) in our Section 3.9 on determinants that our sign convention for the ε-tensor
εj(1)...j(n) is +1, if the mapping j : {1, . . . n} → {1, . . . , n} is an even permutation;
it is -1, if j is an odd permutation; and 0 else.
Proof. Since γj Σ± ∓

2m ⊆ Σ2m , we have Tr(γj ) = 0. More generally, for k odd
and j1 , j2 , . . . , jk distinct, we have γj1 γj2 · · · γjk Σ± ∓
2m ⊆ Σ2m and Tr(γj1 γj2 · · · γjk ) =
0. For k even and j1 , j2 , . . . , jk distinct, we have
k−1
γj1 γj2 · · · γjk = (−1) γj2 · · · γjk γj1 = −(γj2 · · · γjk ) γj1 , and so
Tr(γj1 γj2 · · · γjk ) = − Tr((γj2 · · · γjk ) γj1 ) = − Tr(γj1 (γj2 · · · γjk )) = 0.
If k < 2m, j1 , j2 , . . . , jk are distinct, and the complementary set of indices is
{h1 , . . . , h2m−k } := {1, . . . , 2m} − {j1 , j2 , . . . , jk }, then
Str(γj1 γj2 · · · γjk ) = im Tr(γj1 γj2 · · · γjk γ2m+1 )
= ±im Tr γh1 γh2 · · · γh2m−k = 0.


If k = 2m and j1 , j2 , . . . , jk are distinct, then


Str(γj1 γj2 · · · γj2m ) = εj1 ...j2m Str(γ1 γ2 · · · γ2m )
   2 
2
= εj1 ...j2m im Tr (γ2m+1 ) = εj1 ...j2m im Tr i−m ρC (ωC )
m m m
= εj1 ...j2m im (−1) Tr(Id) = εj1 ...j2m (−i) 2m = εj1 ...j2m (−2i) .
2
Since dim End(Σ2m ) = (dim Σ2m ) = 22m , the 2n endomorphisms consisting of Id
and the γj1 γj2 · · · γjk with j1 < j2 < · · · < jk will form a basis of End(Σ2m ), if they
are shown to be linearly independent. For this, let
X
c0 I + cj1 ...jk γj1 γj2 · · · γjk = 0,
j1 <j2 <···<jk

for c0 , cj1 ...jk ∈ C and let {h1 , . . . , h2m−k } := {1, . . . , 2m} \ {l1 , l2 , . . . , lk } for some
l1 < l2 < · · · < lk . Then
  X 
0 = Str γh1 γh2 · · · γh2m−k c0 I + cj1 ...jk γj1 γj2 · · · γjk
j1 <j2 <···<jk
 m
= cl1 ...lk Str γh1 γh2 · · · γh2m−k γl1 γl2 · · · γlk = ±(−2i) cl1 ...lk ,
17.1. CLIFFORD ALGEBRAS AND SPINORS 523

with the convention that cl1 ...lk = c0 if k = 0. 


Corollary 17.14. The representation ρC : Cl2m → End(Σ2m ) is irreducible
and an isomorphism of complex algebras. Moreover, the representations
ρ± : Spin(2m) −→ End Σ±

2m

are irreducible and inequivalent.


Proof. The last statement of Proposition 17.13 implies that ρC : Cl2m →
End(Σ2m ) is an isomorphism, so that in particular ρC (Cl2m ) = End(Σ2m ). As
End(Σ2m ) acts transitively on the set of subspaces of Σ2m of a given dimension,
End(Σ2m ) leaves no proper subspace of Σ2m invariant, and hence ρC is irreducible.
Let
End0 (Σ2m ) := A ∈ End(Σ2m ) : A Σ± ±
 
2m ⊆ Σ2m and
1 ± ∓
 
End (Σ2m ) := A ∈ End(Σ2m ) : A Σ2m ⊆ Σ2m .
 ∼  ∼
The linear isomorphism L : Λ• R2m −→ C`2m extends to LC : Λ• C2m −→ Cl2m .
We have
Cl2m = LC Λev C2m ⊕ LC Λodd C2m ,
 

ρC L Λev C2m ⊆ End0 (Σ2m ) , and




ρC L Λodd C2m ⊆ End1 (Σ2m ) .




As ρC is an isomorphism, these last two inclusions are equalities. In particular, the


restrictions
π ± ◦ ρC : L Λev C2m → End Σ±
 
2m
ev 2m

are irreducible
 representations of the subalgebra
 L Λ C ⊂ Cl2m . Since
ev 2m 2 2m
L Λ C is generated by L Λ R = spin(2m), there is also no proper

subspace of Σ+ 2m or Σ 2m which is invariant under
 spin(2m) or under Spin(2m) =
exp(spin(2m)). Thus, ρ± : Spin(2m) → SU Σ± 2m are irreducible representations of
Spin(2m). Using the computation in Example 17.7, each of e1 e2 , . . . , e2m−1 e2m are
in Spin(2m). Hence e1 e2 · · · e2m−1 e2m ∈ Spin(2m). Since ρ(e1 e2 · · · e2m−1 e2m ) =
γ2m+1 = ±i−m on Σ± + −
2m , the representations π ◦ ρ and π ◦ ρ are inequivalent. 

Proposition 17.15. Let R : Cl2m → End(V ) be a finite-dimensional repre-


LN
sentation. Then V = k=1 Wk where W1 , . . . , WN of V are invariant subspaces,
such that Rk : Cl2m → End(Wk ) defined by Rk (α) = R(α) |Wk is equivalent to
ρC : Cl2m → End(Λ• (Cm )) for all k = {1, . . . , N }. In particular, all irreducible
representations of Cl2m are equivalent to ρC . Moreover, let
Hom0 (Σ2m , V ) := {F ∈ Hom(Σ2m , V ) : F (ρC (α)(w)) = R(α)(F (w))}
denote the subspace of Hom(Σ2m , V ) of Cl2m -equivariant linear maps; note that
Cl2m acts trivially on Hom0 (Σ2m , V ). There is then an isomorphism of Cl2m -
modules

Φ : Hom0 (Σ2m , V ) ⊗ Σ2m −→ V, given by Φ(φ ⊗ ψ) := φ(ψ) .
Proof. Let e1 , . . . , e2m be an oriented, orthonormal basis for R2m . Let σj =
ie2j−1 e2j ∈ Cl2m for j ∈ {1, . . . , m}. Note that [e2j−1 e2j , e2k−1 e2k ] = 0 for
all j, k ∈ {1, . . . , m}, and so [σj , σk ] = 0 and [R(σj ) , R(σk )] = 0. Since σj2 =
524 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

ie2j−1 e2j ie2j−1 e2j = e2j−1 e2j−1 e2j e2j = 1, the eigenvalues of R(σj ) are ±1. Thus,
there are simultaneous eigenspaces of the R(σj ), j ∈ {1, . . . , m} , indexed by
m
ε = (ε1 , . . . , εm ) ∈ Zm
2 := {1, −1} × · · · × {1, −1} , namely
V (ε) := {v ∈ V : R(σj )(v) = εj v} .
L
We have V = ε∈Zm V (ε). Let
2

m
Y Y
α(ε) = 21 (1 − ε1 ) e2 · · · 12 (1 − εm ) e2m = 1
2 (1 − εj ) e2j = e2j .
{j: εj =−1}
j=1

Then

α(ε) σk , if εk = 1,
σk α(ε) = ie2k−1 e2k α(ε) = = εk α(ε) σk .
−α(ε) σk , if εk = −1,
For 1m := (1, . . . , 1) ∈ Zm
2 , we claim

R(α(ε)) : V (1m ) → V (ε)


is a well-defined isomorphism. Indeed, for v ∈ V (1m ), we have R(α(ε)) v ∈ V (ε),
since
R(σk )(R(α(ε)) v) = R(σk α(ε)) v = R(εk α(ε) σk ) v
= εk R(α(ε)) R(σk ) v = εk (R(α(ε)) v) ,
2 2
and α(ε) = ±1 ⇒ R(α(ε)) = ± Id. Thus,
M
V (ε) = 2m dim V (1m ) .

dim V = dim m ε∈Z2

Let {v1 , . . . , vN } be a basis for V (1m ) and for k ∈ {1, . . . , N }, let


Wk = span {R(α(ε)) vk : ε ∈ Zm
2 }.

We claim that Wk is a Cl2m -module. For this first note that Cl2m is generated by
{α(ε) : ε ∈ Zm
2 } ∪ {σj : j ∈ {1, . . . , m}} ,

since for any j ∈ {1, . . . , m}, the elements e2j ∈ {α(ε) : ε ∈ Zm 2 }, σj e2j = −e2j−1 ,
and {e1 , . . . , e2m } generate Cl2m . Thus, it suffices to show that R(α(ε))(Wk ) ⊂ Wk
for all ε ∈ Zm 2 , and R(σj )(Wk ) ⊂ Wk . To this end,

ε, ε0 ∈ Zm 0 00
2 =⇒ α(ε ) α(ε) = ±α(ε ) , for some ε00 ∈ Zm
2
=⇒ R(α(ε0 )) R(α(ε)) vk = ±R(ε00 ) vk
=⇒ R(α(ε))(Wk ) ⊂ Wk , for all ε ∈ Zm
2 ;

moreover
vk ∈ V (1m ) ⇒ R(σj ) R(α(ε)) vk = ±R(α(ε)) R(σj ) vk = ±R(α(ε)) vk
=⇒ R(σj )(Wk ) ⊂ Wk .
LN
Thus, V = k=1 Wk . Note that dim Wk ≤ 2m , since {R(α(ε)) vk : ε ∈ Zm 2 } spans
Wk . Since Wk is a Cl2m -module, the same proof as that of Proposition 17.13
yields that the 22m endomorphisms consisting of Id and those R(ej1 · · · ejk ) with
1 ≤ j1 < j2 < · · · < jk ≤ 2m, k ∈ {1, . . . , 2m}, form a linearly independent subset
17.2. SPIN STRUCTURES AND TWISTED DIRAC OPERATORS 525

2
of End(Wk ). Since dim(End(Wk )) = (dim Wk ) = 22m , we obtain End(Wk ) =
{R(α) |Wk : α ∈ Cl2m }. Indeed

Cl2m −→ End(Wk ) via α 7→ R(α) |Wk ,
and hence each Wk is an irreducible Cl2m -module and Wk is isomorphic to the
specific module Λ• (Cm ). Note that Φ : Hom0 (Σ2m , V ) ⊗ Σ2m → V is indeed a
morphism, since
 
−1
Φ(α ·(φ ⊗ ψ)) = Φ R(α) ◦ φ ◦ ρC (α) ⊗ ρC (α)(ψ)
 
−1
= R(α) ◦ φ ◦ ρC (α) (ρC (α)(ψ)) = R(α)(φ(ψ)) .

Since Σ2m is irreducible, Hom0 (Σ2m , Σ2m ) = C Id. Since V ∼


LN
= k=1 Σ2m , we then
LN
have (where πk : k=1 Σ2m → Σ2m denotes the projection)
 MN  M
N
Hom0 (Σ2m , V ) ∼
= Hom0 Σ2m , Σ2m = Cπk .
k=1 k=1
m
and hence dim(Hom0 (Σ2m , V ) ⊗ Σ2m ) = N · 2 = dim(V ). For zk ∈ C,
M 
N
Φ zk πk ⊗ ψk = (z1 ψ1 , . . . , zN ψN ) ,
k=1

and Φ is then onto, and an isomorphism for dimensional reasons. 


Remark 17.16. Alternatively, it is known (see [436]) that, up to equivalence,
the only irreducible representation of the algebra End(W ) for any complex or real
vector space W is the defining representation, namely Id : End(W ) → End(W ).
Since Cl2m ∼ = End(Σ2m ), it follows that, up to equivalence, ρC : Cl2m → End(Σ2m )
is the only irreducible representation of Cl2m . Note that ρC restricts to ρ : C`2m →
End(Σ2m ). For dimensional reasons, ρ is not an isomorphism of real algebras.
However, it is clearly an irreducible complex representation of C`2m , since any
invariant subspace for ρ would also be invariant for ρC . Additional considerations
found in [273] imply that ρ : C`2m → End(Σ2m ) is the unique (up to isomorphism)
irreducible real representation of C`2m , but we will not be using this fact.

2. Spin Structures and Twisted Dirac Operators


Let M be a compact, oriented Riemannian n-manifold. Until further notice,
we do not assume that n is even. Let F M denote the principal SO(n)-bundle of
oriented, orthonormal frames.
Spin Structures and Čech Cohomology. A spin structure for M consists
of a principal Spin(n)-bundle P → M and a map C : P → F M which is equivariant
in the sense that C (pg) = C (p) c(g), where c : Spin(n) → SO(n) denotes the double
cover of (17.5), namely c(g)(v) = gvg −1 . Spin structures do not always exist, but

given a coordinate ball U ⊆ M and a local trivialization T : F M |U −→ U × SO(n),
there is the obvious local spin structure

T ◦(Id ×c) : U × Spin(n) → U × SO(n) −→ F M |U .
Our immediate goal is to establish the meaning and sketch the proof of Proposition
17.20 below which exhibits the obstruction to finding a (global) spin structure for
M . Suppose that U = {Uα : α ∈ J} is an open covering of M , such that each
intersection Uα1 ∩ · · · ∩ Uαk of finitely many of the Uα is contractible (e.g., take
526 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

each Uα to be a convex, normal coordinate ball). Such a covering is known as a



Leray covering. There is then a local trivialization Tα : F M |Uα −→ Uα × SO(n),
say
Tα (u) = (π(u) , sα (u)) ,
where sα (ug) = sα (u) g for all g ∈ SO(n). Since
−1 −1 −1 −1
sα (ug) sβ (ug) = sα (u) g(sβ (u) g) = sα (u) gg −1 sβ (u) = sα (u) sβ (u) ,
there is a well defined transition function
−1
gαβ : Uα ∩ Uβ → SO(n) given by gαβ (π(u)) := sα (u) sβ (u) .
−1
Note that gβα = gαβ and, more generally, we have the so-called cocycle condition
gαβ gβγ gγα = I.
Conversely, if we are given functions gαβ : Uα ∩ Uβ → SO(n), satisfying the cocycle
condition (together with gαα ≡ I), then a principal SO(n)-bundle over M can be
constructed as the set of equivalence classes for the relation on the disjoint union
of the Uα × SO(n), where we declare that (x, Aα ) ∈ Uα × SO(n) is equivalent
to (x, Aβ ) ∈ Uβ × SO(n) if Aα = gαβ (x) Aβ for x ∈ Uα ∩ Uβ . Since, gαα ≡ I,
−1
gβα = gαβ , and gαβ gβγ gγα = I, this is an equivalence relation. As each of the
0
Uα ∩Uβ is contractible, we can find a lift gαβ : Uα ∩Uβ → Spin(n) of gαβ : Uα ∩Uβ →
0
SO(n) with c ◦ gαβ = gαβ , where c : Spin(n) → SO(n) denotes the double covering
0
(vector representation). We can (and do) choose the collection of gαβ so that
0 0 −1
gβα = (gαβ ) . Now, on Uα ∩ Uβ ∩ Uγ , we have
0 0 0 0 0 0

c gαβ gβγ gγα = gαβ gβγ gγα = I =⇒ gαβ gβγ gγα = ±1 ∈ Z2 ⊆ Spin(n) ,
yielding the neutral element of multiplication in the field Z2 . Let
0 0 0 0
Eαβγ := gαβ gβγ gγα : Uα ∩ Uβ ∩ Uγ → Z2 .
0 0
Note that Eαβγ is symmetric (as well as antisymmetric, as Eαβγ ∈ Z2 ) in α, β, γ,
since
0 0 0 0 0 0
gαβ gβγ gγα = ±1 =⇒ gαβ gβγ = ±gαγ
0 0 0 0 0 0 0 0
 
=⇒ Eαβγ = gαβ gβγ gγα = gγα gαβ gβγ = Eγαβ , and
0 0
 −1 0
 −1 0
 −1 0
 −1 0 0 0 0
Eαβγ = Eαβγ = gγα gβγ gαβ = gαγ gγβ gβα = Eαγβ .
The collection E 0 = Eαβγ 0

is an example of a Čech 2-cochain (with values in
the field Z2 ) relative to the cover U, the group of which we denote by C 2 (U; Z2 ).
The coboundary of E 0 is the Čech 3-cochain in C 3 (U; Z2 ) defined by
−1 −1
(δE 0 )αβγδ := Eβγδ 0
· Eαγδ0 0
· Eαβδ 0
· Eαβγ .

Exercise 17.17. Find a natural definition of a Čech 1-cochain h = hαβ and
define its coboundary.
The following shows that δE 0 = 1:
(δE 0 )αβγδ = Eβγδ
0 0 0 0 0 0
 0 0

Eαγδ Eαβδ Eαβγ = Eδβγ Eδγα Eβδα Eβαγ
0 0 0 0
 0 0 0 0 
= gδβ gβγ gγα gαδ gβδ gδα gαγ gγβ
0 0 0 0
 0 0 0 0 
= gδβ gβγ gγα gαδ gδα gαγ gγβ gβδ = 1,
17.2. SPIN STRUCTURES AND TWISTED DIRAC OPERATORS 527

meaning that E 0 is a Čech 2-cocycle, the group of which is denoted by Z 2 (U; Z2 ).


00 0
Suppose that we choose a different set of lifts of the gαβ , say gαβ = hαβ gαβ for
1 00 00 00 00
h = {hαβ } ∈ C (U; Z2 ). Then for Eαβγ := gαβ gβγ gγα ,
00 0
−1 00 00 00
 0 0 0 −1
Eαβγ Eαβγ = gαβ gβγ gγα gαβ gβγ gγα
00 00 00
 0
 −1 0 −1 0 −1
= gαβ gβγ gγα gγα gβγ gαβ
= hαβ hβγ hγα = hβγ hγα hαβ =: (δh)αβγ ,
00 0
so that E and E differ by δh. The group of 2-coboundaries is
B 2 (U; Z2 ) := δh : h ∈ C 1 (U; Z2 ) ,


and it is easy to check that for any h ∈ C 1 (U; Z2 ), δδh = 1, so that B 2 (U; Z2 ) ⊂
Z 2 (U; Z2 ). Hence, we have shown that E 0 and E 00 determine the same Čech co-
homology class
Z 2 (U; Z2 )
[E 0 ] = [E 00 ] ∈ H 2 (U; Z2 ) := 2 .
B (U; Z2 )
Of course H k (U; Z2 ) can be defined for k = 0, 1, 2, . . .. It can be shown that for
Leray coverings U, H k (U; Z2 ) is naturally isomorphic to the usual (say, singular)
cohomology group H k (M ; Z2 ), with Z2 -coefficients.
Definition 17.18. Let E 0 ∈ Z 2 (U; Z2 ) denote the Čech 2-cocycle given by
0 0 0 0
Eαβγ := gαβ gβγ gγα : Uα ∩ Uβ ∩ Uγ → Z2 ,
0
where the gαβ : Uα ∩ Uβ → Spin(n) are lifts of the transition functions gαβ : Uα ∩
Uβ → SO(n) for the oriented frame bundle F M of a compact, oriented Riemannian
n-manifold relative to a Leray covering U = {Uα : α ∈ J}. The class w2 (M ) :=
[E 0 ] ∈ H 2 (M ; Z2 ) is known as the second Stiefel-Whitney class of M . More
generally, by using transition functions, any equivalence class of a principal SO(k)-
bundle P → M (where k is not necessarily the dimension of M ) can be identified
with some [P ] ∈ H 1 (M, SO(k)) and a function
w2 : H 1 (M, SO(k)) → H 2 (M ; Z2 )
may be defined in the same way as w2 (M ) was defined in the case of F M . Thus,
w2 (M ) is w2 ([P ]) in the special case P = F M , but for convenience we write w2 (M )
instead of w2 ([F M ]).
Remark 17.19. Given a principal SO(k1 )-bundle P1 → M and a principal
SO(k2 )-bundle P2 → M , one can define a principal SO(k1 ) × SO(k2 ) bundle P1 ×
P2 → M (fibered product) which determines an SO(k1 + k2 )-bundle P → M by
means of the injection SO(k1 ) × SO(k2 ) → SO(k1 + k2 ) as in Proposition 15.20,
p.407. Note that transition functions for P → M can be taken to be products
of transition functions for P1 → M (with values in SO(k1 ) × Id) and transition
functions for P2 → M (with values in Id × SO(k2 )). Since such products commute,
it is clear from our construction that w2 ([P ]) = w2 ([P1 ]) w2 ([P2 ]), or regarding
H 2 (M ; Z2 ) as an additive group (as it is usually the case) we have
(17.17) w2 ([P ]) = w2 ([P1 ]) + w2 ([P2 ]) .
This formula may seem wrong to those already familiar with Stiefel-Whitney classes,
since generally there is also a cup product term w1 ([P1 ]) ` w1 ([P2 ]) on the right.
However, w1 ([P ]) is trivial for SO(k)-bundles, which is sufficient for our purposes.
528 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

Admittance of Spin Structures. The following proposition exhibits the


obstruction to finding a (global) spin structure to a compact, oriented Riemannian
manifold M .
Proposition 17.20. Let M be a compact, oriented Riemannian n-manifold.
Then M admits a spin structure if and only if w2 (M ) = 0 (i.e., w2 (M ) is the
identity of H 2 (M ; Z2 )). In this case H 1 (M ; Z2 ) acts freely on the set of inequivalent
spin structures.
Proof. If w2 (M ) = 0, then E 0 ∈ B 2 (U; Z2 ) and so E 0 is the Čech coboundary
of a Čech 1-cochain (see Exercise 17.17), say F = {Fαβ }; i.e.,
0 −1
Eαβγ = (δF )αβγ = Fβγ (Fαγ ) Fαβ = Fβγ Fγα Fαβ = Fαβ Fβγ Fγα .
We construct a spin structure C : P → F M as follows. Let geαβ 0 0
:= Fαβ gαβ ∈
0
Spin(n), and note that while the gαβ did not necessarily satisfy the cocyle condition,
0
the geαβ do:
0 0 0 0 0 0
geαβ geβγ geγα = Fαβ gαβ Fβγ gβγ Fγα gγα
0 0 0 0
2
= Fαβ Fβγ Fγα gαβ gβγ gγα = Eαβγ = 1.
Thus, a principal Spin(n)-bundle P → M can be constructed from the transition
0 0 0 0
functions geαβ . Since c(e
gαβ ) = c(Fαβ gαβ ) = c(±gαβ ) = gαβ , where c : Spin(n) →
SO(n) denotes the double cover, we obtain a spin structure C : P 0 → F M . Con-
versely, given a spin structure C : P 0 → F M , with transition functions geαβ 0
, there
1 0 0 0
is F ∈ C (U; Z2 ) with geαβ = Fαβ gαβ . However, a choice F , differing from F , with
δF 0 = δF = E 0 might lead to different spin structure P 0 → M , defined via tran-
00 0 0
sition functions geαβ := Fαβ gαβ . While F and F 0 are not cocycles, their difference
0 −1 0 −1
= E 0 E 0−1 = 1. If F 0 F −1 is a coboundary, say

F F is a cocycle, since δ F F
0 −1 0
F F = δk for some k ∈ C (U; Z2 ), then
0
Fβα Fαβ = kβ kα−1 or Fαβ = kα−1 Fαβ
0
kβ .
This condition yields a well defined principal bundle isomorphism φ : P → P 0 de-
termined locally via the maps φα : Uα × Spin(n) → Uα × Spin(n) defined by
φα (x, aα ) = (x, kα aα ) .
To see that the φα yield a well defined φ : P → P 0 , note that (where ≡0 denotes the
equivalence relation used in defining P 0 from the geαβ
00
)
(x, kα aα ) ≡0 (x, kβ aβ ) ⇔ kα aα = geαβ
00
kβ aβ
0 0
⇔ kα aα = Fαβ gαβ kβ aβ ⇔ aα = kα−1 Fαβ
0 0
kβ gαβ aβ
0 0
⇔ aα = Fαβ gαβ aβ ⇔ aα = geαβ aβ ⇔ (x, aα ) ≡ (x, aβ ) .
That φ : P → P 0 is Spin(n)-equivariant and φ ◦ C = C 0 ◦ φ for the coverings
C : P → F M and C 0 : P 0 → F M , follows from these obvious properties for the φα
(recall kα ∈ Z2 = {1, −1}, so that c(aα ) = c(kα aα )). Thus, if F 0 F −1 = δk (i.e.,
F 0 and F differ by a coboundary), then F and F 0 define equivalent spin structures
and the converse also holds. Note that
00 0 0 0 −1 0 0 −1 0
geαβ = F 0 F −1 αβ geαβ
0

geαβ = Fαβ gαβ = Fαβ Fαβ Fαβ gαβ = Fαβ Fαβ .
Hence, for a given F , there is a one-to-one correspondence between H 1 (M ; Z2 ) and
the set S(M ) of inequivalent spin structures, induced by F 0 F −1 7→ F 0 F −1 ge0 .
17.2. SPIN STRUCTURES AND TWISTED DIRAC OPERATORS 529

Alternatively put, there is a free action of H 1 (M ; Z2 ) on S(M ), given simply by


multiplication, namely (z · ge0 )αβ := zαβ geαβ
0
for z ∈ H 1 (M ; Z2 ). 

Proposition 17.21. Let P be a principal U(1)-bundle over a connected, ori-


ented 2-manifold M . Since U(1) ∼ = SO(2) , we may regard P as a principal SO(2)-
bundle. Thus, P possesses a Stiefel-Whitney class w2 (P ) ∈ H 2 (M ; Z2 ) as well as
a Chern class c1 (P ) ∈ H 2 (M ; Z). We have
w2 (P ) [M ] ≡ c1 (P ) [M ] mod 2.
In the special case P = F M , we have (see (15.117), p.454)
c1 (F M ) [M ] = χ [M ] = 2 − 2 genus(M )
which is even and so w2 (M ) := w2 (F M ) = 0.

Proof. Let ω ∈ Λ1 (P, u(1)) be a connection 1-form for P as a U(1)-bundle,


and let U = {Uα : α ∈ J} be a Leray covering by smoothly embedded disks so that
we have trivializing local sections σα : Uα → P |Uα . For a U(1)-bundle, the formula
(see Exercise 15.34, p. 414) relating σα∗ ω to σβ∗ ω simplifies:
−1 ∗ −1 −1
σβ ∗ ω = gαβ (σα ω) gαβ + gαβ dgαβ = σα∗ ω + gαβ dgαβ or
∗ −1
σβ ω − σα∗ ω = gαβ dgαβ = d(log gαβ ) .

Let iAα := σα∗ ω. The local curvature forms Ωα = dσα∗ ω = idAα and Ωβ = dσβ∗ ω =
idAβ agree on the overlaps Uα ∩ Uβ to yield a well-defined 2-form F ∈ Ω2 (M, R)
1
given locally by F := −dAα , and 2π F represents c1 (P ) ∈ H 2 (M ; Z) in de Rham
cohomology. The isomorphism
Z 2 (U; R)
H 2 (M ; R) → H 2 (U; R) :=
B 2 (U; R)

from de Rham cohomology to Čech cohomology is obtained via the intermediate


isomorphisms

2 H 0 U, Z 2 ∼  ∼
H (M ; R) = −→ H 1 U, Z 1 −→ H 2 (U; R) ,
dH 0 (U, A1 )
which we describe as follows. For a Leray cover U of M and for 0 ≤ j, k ≤ 2, let
C j U, Ak denote the  group of Čech j-cochains c which assign to an ordered (j + 1)-

tuple Uα0 , . . . , Uαj , with Uαi ∈ U, a k-form cα0... αj ∈ Ωk Uα0 ∩ . . . ∩ Uαj , R .
Similarly, let
C j U, Z k := c ∈ C j U, Ak : dcα0 ...αj = 0 for all (α0 , . . . , αj ) and
  

c ∈ C j U, Ak : cα0 ...αj = dbα0... αj for some
 
j k

C U, B := .
bα0 ...αj ∈ Ωk−1 Uα0 ∩ . . . ∩ Uαj , R
The short exact sequences of cochain complexes
 
∼ i  d
0 → C ∗ U, Z 0 −→ R → C ∗ U, A0 → C ∗ U, Z 1 → 0 and


 i  d
0 → C ∗ U, Z 1 → C ∗ U, A1 → C ∗ U, Z 2 → 0

530 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

give rise (in the standard way) to long exact sequences of cohomology groups
0 → H 0 U, Z 0 → H 0 U, A0 → H 0 U, Z 1 → H 1 U, Z 0 → H 1 U, A0
    
 δ1 2
→ H 1 U, Z 1 → H U, Z 0 → H 2 U, A0 · · · and
 

(17.18)
 δ0 1
0 → H 0 U, Z 1 → H 0 U, A1 → H 0 U, Z 2 → H U, Z 1 → H 1 U, A1 · · ·
   

Since the Aj are fine sheaves (admitting


 partitions of unity), we have (see [434,
Theorem 3.11, p.56]) H i U, Aj = 0 for i > 0 and j ≥ 0. Thus,
 δ
2 ∼ H 0 U, Z 2 0
∼ δ
H 1 U, Z 1 =1 H 2 U, Z 0 = HČech
2

Hde Rham (M ; R) −→ −→ (U; R) .
dH 0 (U, A1 )
The first isomorphism is obtained by noting that

c ∈ C 0 U, A2 : dcα0 = 0 and
 
H 0 U, Z 2 = Z 0 U, Z 2 =
 
,
0 = (δc)α0 α1 = cα1 − cα0 for all α0 , α1

and so any c ∈ H 0 U, Z 2 gives rise to a globally defined, closed 2-form on M .
Similarly H 0 U, A1 can  be identified with the space of globally defined 1-forms
on M , and dH 0 U, A1 ∼ = the space of exact 2-forms on M . The isomorphism δ0
2
in (17.18) is given as follows. Let [F ] ∈ Hde Rham (M ; R) for a closed 2-form F ∈
2
Ω (M, R). We have  that F |U α
is exact, say F |Uα = −dAα for some Aα ∈ Ω1 (Uα )
0 1
or A ∈ C U, A . We have (δA)αβ = Aβ − Aα . Now
 
d (δA)αβ = dωβ − dωα = F |(Uα ∩Uβ ) − F |(Uα ∩Uβ ) = 0,
 
so that δA  ∈ C 1 U, Z 1 and δ(δA) = 0 so that δA ∈ Z 1 U, Z 1 and [δA] ∈
H 1 U, Z 1 . Note that 0 1
 [δA] is not necessarily 0, since A ∈ C U, A is not nec-
0 1
essarily in C U, Z . At any rate, δ0 in (17.18) is given by δ0 [F ] := [−δA]. We
1 1
now define δ1 in
2 1
 (17.18). For [B] ∈ H U,1Z we 0
 have dBα0 α1 = 0 and 0 =
δB ∈ C U, Z , and so there is Cα0 α1 ∈ C U, A , such that dCα0 α1 = Bα0 α1 .
Moreover,
(δC)α0 α1 α2 = Cα1 α2 − Cα0 α2 + Cα0 α1

d (δC)α0 α1 α2 = dCα1 α2 − dCα0 α2 + dCα0 α1 = Bα1 α2 − Bα0 α2 + Bα0 α1
= (δB)α0 α1 α2 = 0.
 
Thus, δC ∈ C 2 U, Z 0 and δ(δC) = 0 so that δC ∈ Z 2 U, Z 0 and δ1 ([B]) :=
[δC] ∈ H 2 U, Z 0 . We wish to compute
 1 
δ1 δ0 c1 (P ) = δ1 δ0 2π F .
Recall that F |Uα = −dAα , and
−1
iAβ − iAα = σβ ∗ ω − σα∗ ω = gαβ dgαβ = d log gαβ ,
where the simple-connectedness of Uα ∩ Uβ yields a well-defined function
log gαβ ∈ Ω0 (Uα ∩ Uβ , iR) (unique up to additive multiples of 2πi),
such that exp(log gαβ ) = gαβ . Then we have
∈ H 2 (U; R) .
 1   1   1

δ1 δ0 c1 (P ) = δ1 δ0 2π F = δ1 − 2π δA = δ 2πi log g
17.2. SPIN STRUCTURES AND TWISTED DIRAC OPERATORS 531

In fact δ1 δ0 c1 (P ) ∈ H 2 (U; Z), since gα1 α2 gα−1 g


0 α2 α0 α1
= 1 implies
1 1

δ 2πi log g α0 α1 α2 = 2πi (log gα1 α2 − log gα0 α2 + log gα0 α1 )
1
log gα1 α2 gα−1

= 2πi g
0 α2 α0 α1
∈ Z.
 1 
Thus, δ 2πi log g ∈ H 2 (U; Z) is the Čech cohomological version of c1 (P ) ∈
H 2 (M ; Z). The Stiefel-Whitney class w2 (P ) ∈ H 2 (U; Z2 ) is [E0 ] = [δg 0 ], where
0 √ 1 1
we may take gαβ = gαβ := exp 4πi log gαβ . Then δ 2πi log g α0 α1 α2 is an odd
0
integer on Uα0 ∩ Uα1 ∩ Uα2 ⇔ (δg )α0 α1 α2 is −1, and it follows that the mod 2 re-
log g is [δg 0 ]. Thus, w2 (P ) ∈ H 2 (U; Z2 ) is the mod 2 reduction
 1 
duction of δ 2πi
of c1 (P ) ∈ H 2 (M ; Z), and w2 (P ) [M ] ≡ c1 (P ) [M ] mod 2. 
Standard and Twisted Dirac Operators and Chirality. Until further
notice, we again assume that n = dim M is even, n = 2m, and moreover that
C
there is a spin structure P → F M → M . We may then form the Hermitian
positive and negative spinor bundles Σ± (M ) := P ×Spin(n) Σ± 2m associated to the
half-spinor representations ρ± : Spin(n) → SU Σ±

2m of (17.16), p. 521. Recall that
ρ+ ⊕ ρ− : Spin(n) → SU(Σ2m ) is the restriction of ρ : C`2m → End(Σ2m ). For
v ∈ R2m ⊂ C`2m , we have ρ(v) : Σ± ∓
2m → Σ2m , with
−1
ρ(c(a) v) = ρ ava−1 = ρ(a) ρ(v) ρ(a)

for a ∈ Spin(n) .
Thus, v 7→ ρ(v) induces a well defined vector bundle morphism T M → End(Σ(M ))
where Σ(M ) := Σ+ (M ) ⊕ Σ− (M ), or equivalently a so-called Clifford multipli-
cation
c : T M ⊗ Σ(M ) → Σ(M ) .
± ∓
with c(T M ⊗ Σ (M )) ⊆ Σ (M ). The Riemannian metric on M gives an identi-
fication of T ∗ M with T M and hence we may regard c : T ∗ M ⊗ Σ(M ) → Σ(M ),
which induces a map (still denoted by c) on the level of sections
c : Ω1 (M ) ⊗ C ∞ (Σ(M )) → C ∞ (Σ(M )) .
The Levi-Civita connection 1-form θ ∈ Ω1 (F M, so(n)) on F M was introduced
in Definition 15.40 (p. 418). It pulls back via C : P → F M to a form C ∗ (θ) ∈

Ω1 (P, so(n)), which when composed with the isomorphism c0−1 : so(n) −→ spin(n),
0−1 ∗ 1
gives us a form ω = c (C (θ)) ∈ Ω (P, spin(n)). It follows from the equivariance
of C : P → F M with respect to c : Spin(n) → SO(n) (i.e., C (pg) = C (p) c(g)), that
ω is a connection 1-form for P → M . Thus, we have a covariant differentiation
operator associated with ω, say
∇Σ : C ∞ (Σ(M )) → Ω1 (M ) ⊗ C ∞ (Σ(M )) .
Definition 17.22. For an oriented Riemannian manifold M with spin struc-
ture, the (standard) Dirac operator (also known as the Atiyah-Singer operator
in [273]), is the composition
DΣ := c ◦∇Σ : C ∞ (Σ(M )) → C ∞ (Σ(M )) .
For many applications, it will be necessary to use twisted Dirac operators which
are defined as follows. Let E → M be a Hermitian vector bundle, and let U (E) →
M denote the principal bundle of unitary frames of E. We equip U (E) with any
connection 1-form, say ε ∈ C(U (E)), and then we have an associated covariant
differentiation operator ∇E : C ∞ (E) → Ω1 (M )⊗C ∞ (E). For the bundle E ⊗Σ(M )
532 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

(which is an associated bundle of the fibered product U (E) ×f P ), we have a


covariant differentiation operator
∇ := ∇E ⊗ 1 + 1 ⊗ ∇Σ : C ∞ (E ⊗ Σ(M )) → Ω1 (M ) ⊗ C ∞ (E ⊗ Σ(M )) .
This corresponds to the connection 1-form on U (E) ×f P which is the direct sum
of the pull-backs of the connection 1-forms on U (E) and P to a U (E) ×f P .
Definition 17.23. For an oriented Riemannian manifold M with spin struc-
ture, and a Hermitian vector bundle E → M with unitary connection ε ∈ C(U (E)),
the twisted Dirac operator D associated with (E, ε) is
(17.19) D := (1 ⊗ c) ◦ ∇ : C ∞ (E ⊗ Σ(M )) → C ∞ (E ⊗ Σ(M )) .
Note that ∇(C ∞ (E ⊗ Σ± (M ))) ⊆ Ω1 (M ) ⊗ C ∞ (E ⊗ Σ± (M )), while
(1 ⊗ c) Ω1 (M ) ⊗ C ∞ E ⊗ Σ± (M ) ⊆ C ∞ E ⊗ Σ∓ (M ) .
 

Thus, we obtain a chiral decomposition


D = D+ ⊕ D− , where
D± : C ∞ E ⊗ Σ± (M ) → C ∞ E ⊗ Σ∓ (M ) .
 

The symbol of the first-order differential operator D is computed as follows. For


φ ∈ C ∞ (M ) with φ(x) = 0 and ψ ∈ C ∞ (E ⊗ Σ(M )) , we have at x
(1 ⊗ c) ◦ ∇(φψ) = (1 ⊗ c) ◦((dφ) ψ + φ∇ψ) = (1 ⊗ c) ◦(dφ) ψ
= (1 ⊗ c(dφ)) ψ.
Thus, the symbol σ(D) : Tx M ∗ → End(Σ(M )) at the covector ξx ∈ Tx M ∗ is given
by
σ(D)(ξx ) = 1 ⊗ c(ξx ) : Σx → Σx .
For ξx 6= 0, σ(D)(ξx ) is an isomorphism, since
2 2
σ(D)(ξx ) ◦ σ(D)(ξx ) = 1 ⊗ c(ξx ) = − |ξx | Id .
Thus, D is an elliptic operator. Moreover, since σ(D+ ) and σ(D− ) are restrictions
of σ(D), it follows that D+ and D− are elliptic. In what follows, we show that D
is formally self-adjoint, and D+ and D− are formal adjoints of each other.
At times it is best to express Dψ in terms of a local orthonormal frame field
E1 , . . . , En on M . If ϕ1 , . . . , ϕn is the dual coframe, then, for any vector field X,
we have
X X X
(∇ψ)(X) = (∇ψ) ( ϕj (X) Ej ) = ϕj (X)(∇ψ)(Ej ) = ϕj (X) ∇Ej ψ;
j j j

i.e.,
X
∇ψ = ϕj ⊗ ∇Ej ψ.
j

Thus, with “·” denoting Clifford multiplication (1 ⊗ c) and using T ∗ M ∼


= T M , we
have
X X
(17.20) Dψ = ϕj · ∇Ej ψ = Ej · ∇Ej ψ.
j j
17.2. SPIN STRUCTURES AND TWISTED DIRAC OPERATORS 533

Note that Dψ is independent of the choice of local orthonormal frame field. More-
over, in terms of local coordinates, say y1 , . . . , yn with associated coordinate fields
∂i := ∂/∂y i , we have
X
(17.21) Dψ = hij ∂i · ∇∂j ψ,
i,j
ij
where h are the entries of the inverse of the matrix [hij ] = [h(∂i , ∂j )]. In the
proof of Proposition 17.25 below, we take advantage of the fact that we can always
choose E1 , . . . , En in a neighborhood U of a point x ∈ M so that ∇Ej Ek = 0 at x,
where ∇ denotes the covariant derivative for the Levi-Civita connection θ. More
generally we have
Proposition 17.24. Let σ : U → F M be a local section of the frame bundle
of a Riemannian manifold M , where U ⊆ M is open. Let E1 , . . . , En denote the
orthonormal frame field on U given, at y ∈ U , by (Ej )y := σ(y)(ej ), where ej
denotes the j-th standard unit vector in Rn . If θ ∈ Ω1 (F M, so(n)) is a connection
1-form on F M , with associated covariant differentiation operator ∇, then we have
Xn n
X
(17.22) ∇Ej Ek = (σ ∗ θ)ik (Ej ) Ei = θik (σ∗y (Ej )) Ei .
i=1 i=1

For x ∈ M and u ∈ πF−1 (x),


let N be an embedded submanifold of F M with
Tu N = Hu := Ker θu . There is a neighborhood V of x, such that there is a unique
section σ : V → F M with σ(x) = u, σ(V ) ⊆ N and σ∗x (Tx M ) = Hu . In this case,
∇Ej Ek = 0 at x by (17.22).
Proof. For a vector field Y on M , the corresponding equivariant function
0
Ω (F M, Rn ) is ϕ(Ye ), where Ye denotes the horizontal lift of Y relative to the connec-
1
tion θ and ϕ ∈ Ω (F M, Rn ) denotes the canonical 1-form (ϕu (Y ) := u−1 (πF ∗ Y )).
For horizontal lifts Ye , W
f of vector fields Y and W on M , we have
 
θ
ϕ(∇^Y W ) = D (ϕ(W )) Y
f e = Ye [ϕ(W f )].

Let Ee1 , . . . , E
en denote the horizontal lifts of E1 , . . . , En . Note that ϕ(E
ek ) is con-
stant on σ(U ), since
ek ) = σ(y)−1 (πF ∗ (E
ϕσ(y) (E ek )) = σ(y)−1 ((Ek ) ) = ek ,
y

and so d(ϕ(E ej = σ∗x (Ej ) and Dθ = d + θ, we have


ek ))(σ∗y (Ej )) = 0. Then, as E
H

Dθ (ϕ(E ej ) = Dθ (ϕ(E
ek ))(E ek ))(σ∗y (Ej ))

= d(ϕ(E
ek ))(σ∗y (Ej )) + θ(σ∗y (Ej )) (ϕ(E
ek ))

= θ(σ∗y (Ej )) (ϕ(E


ek )).
Hence, it follows that
 
−1
∇^
Ej Ek = (ϕ|H ) Dθ (ϕ(E
ek ))(E
ej )
n n
−1
X X
= θik (σ∗x (Ej ))(ϕ|H ) (ei ) = θik (σ∗y (Ej )) E
ei
i=1 i=1
and (17.22) follows by applying (πF )∗ to both sides. The remaining assertions
follow directly from the implicit function theorem. 
534 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

For ψ, ψ 0 ∈ C ∞ (E ⊗ Σ(M )) and ν the volume form for (M, h), let
Z
0
(ψ, ψ ) := hψ, ψ 0 i ν.
M

Proposition 17.25. The twisted Dirac operator D of (17.19) is formally self-


adjoint, namely
(17.23) (Dψ1 , ψ2 ) = (ψ1 , Dψ2 ) .
Moreover, D is the formal adjoint of D− ; i.e., for ψ1+ ∈ C ∞ (E ⊗ Σ+ (M )) and
+

ψ2− ∈ C ∞ (E ⊗ Σ− (M )) , we have
D+ ψ1+ , ψ2− = ψ1+ , D− ψ2− .
 
(17.24)
Proof. Let ψ1 , ψ2 ∈ C ∞ (E ⊗ Σ(M )). Since ∇ is the covariant differentiation
for a connection on U (E) ×f P with group U(N ) × Spin(n) which preserves the
Hermitian structure on CN ⊗ Σn , it follows that
(17.25) d hψ1 , ψ2 i = h∇ψ1 , ψ2 i + hψ1 , ∇ψ2 i .
Moreover, for ρ : C`n → End(Σn ), a ∈ Spin(n), v ∈ Rn , and σ ∈ Σn , we have
 
−1
ρ(a)(ρ(v) σ) = ρ(a) ◦ ρ(v) ◦ ρ(a) ◦ ρ(a) (σ)
= ρ ava−1 (ρ(a) σ) = ρ(c(a) v)(ρ(a) σ) .

(17.26)
Replacing a with exp(ta0 ) for a0 ∈ spin(n) and differentiating (17.26) with respect
to t at t = 0, we get
(17.27) ρ0 (a0 )(ρ(v) σ) = ρ(c0 (a0 ) v)(σ) + ρ(v)(ρ0 (a0 ) σ) .
From this it follows that, for ϕ ∈ Ω1 (M ) and ψ ∈ C ∞ (E ⊗ Σ(M )),
∇(ϕ · ψ) = (∇θ ϕ) · ψ + ϕ · ∇ψ.
For a local orthonormal frame field E1 , . . . , En , chosen so that ∇Ej Ej = 0 at some
fixed x ∈ M , using (17.25) and (17.27), we have (at x)
X X
hDψ1 , ψ2 i = Ej · ∇Ej ψ1 , ψ2 = − ∇Ej ψ1 , Ej · ψ2
j j
X 
= − Ej [hψ1 , Ej · ψ2 i] − ψ1 , ∇Ej (Ej · ψ2 )
j
X  
= −Ej [hψ1 , Ej · ψ2 i] + ψ1 , ∇Ej Ej · ψ2 + Ej · ∇Ej ψ2
j
X
= −Ej [hψ1 , Ej · ψ2 i] + ψ1 , Ej · ∇Ej ψ2
j
X
= −Ej [hψ1 , Ej · ψ2 i] + hψ1 , Dψ2 i .
j

Consider the 1-form α given by


α(Y ) = hψ1 , Y · ψ2 i .
We claim that at x,
X X
δα := ∗d(∗α) = − Ej [α(Ej )] = − Ej [hψ1 , Ej · ψ2 i] .
j j

The j-th component of the related α e ∈ Ω0 F M, T 0,1
on F M is given at u ∈ F M by
α
ej = αe(u)(ej ) = α(πF ∗ (ej )), where ej denotes the j-th standard horizontal vector
field at u. Recall that (πF )∗ : H → Tπ(p) M denotes the orientation-preserving
17.2. SPIN STRUCTURES AND TWISTED DIRAC OPERATORS 535

isometry on the horizontal subspace H ⊂ T(p,u) (P ×f F M ), defined in (15.92)


(p.437 of Section 15.6. According to (15.94) on p. 438,
X
δα = − ej [e
αj ] ,
j

which is constant on each fiber of F M . For any y ∈ U , we have ej (σ(y)) = E


ej (σ(y)),
since both sides are horizontal and
   
ϕ E ej (σ(y)) = σ(y)−1 πF ∗ E ej (σ(y))
−1
= σ(y) (Ej (y)) = ej = ϕ(ej (σ(y))) ,
1
where ϕ ∈ Ω (F M, Rn ) denotes the canonical 1-form. For any y ∈ U ,
(σ ∗ α
ej )(y) = α
ej (σ(y)) = α(πF ∗ ej (σ(y)))
 
= α πF ∗ E
ej (σ(y)) = α(Ej (y)) , and so

Ej [α(Ej )] = d(α(Ej ))(Ej ) = d(σ ∗ α


ej )(Ej ) = σ ∗ (de
αj )(Ej )
= de
αj (σ∗ (Ej )) = σ∗ (Ej ) [e
αj ] .
However at u = σ(x), we have σ∗ (Ej ) [e
αj ] = ej [e
αj ]. Thus, at x
X X
δα = − αj ] = −
ej [e Ej [α(Ej )] .
j j

Then (17.23) follows, since


X
hDψ1 , ψ2 i ν − hψ1 , Dψ2 i ν = − Ej hψ1 , Ej · ψ2 i ν
j
= (δα) ν = ∗(δα) = − ∗ ∗d(∗α) = −d(∗α)
Z
=⇒ (Dψ1 , ψ2 ) −(Dψ2 , ψ1 ) = (hDψ1 , ψ2 i − hψ1 , Dψ2 i) ν
M
Z
=− d(∗α) = 0.
M

Applying (17.23) when ψ1 has values in E ⊗ Σ+ (M ) and ψ2 has values in the


orthogonal bundle E ⊗ Σ− (M ), we obtain (17.24). 

The index of the self-adjoint operator D is 0, but the index of D+ (or D− ) is


not necessarily 0. Using the fact that D+ and D− are adjoints, we have
index D+ = dim Ker D+ − dim Coker D+
  

= dim Ker D+ − dim Ker D−


 

= dim Ker D− ◦ D+ − dim Ker D+ ◦ D− .


 

Moreover, since D2 = (D− ◦ D+ ) ⊕(D+ ◦ D− ), it is convenient to define


2
D+ := D2 |C ∞(E⊗Σ+(M )) = D− ◦ D+
(17.28) 2
D− := D2 |C ∞(E⊗Σ−(M )) = D+ ◦ D− .

A detailed study of D2 is desirable, since the above yields


index D+ = dim Ker D+ 2 2
  
− dim Ker D− .
536 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

To develop a suitable formula for D2 , we compute (at x ∈ M where ∇Ej Ek = 0)


X X X
D2 ψ = D(Dψ) = D( Ej · ∇Ej ψ) = E j · ∇ Ej ( Ek · ∇Ek ψ)
j j k
X X
= Ej · ∇Ej (Ek · ∇Ek ψ)
j k
X 
= Ej · ∇Ej Ek · ∇Ek ψ + Ek · ∇Ej (∇Ek ψ)
j,k
X
= Ej · Ek · ∇Ej (∇Ek ψ)
j,k
X X
= Ej · Ej · ∇Ej (∇Ek ψ) + Ej · Ek · ∇Ej (∇Ek ψ)
j j6=k
X X
1
 
=− ∇ Ej ∇ Ej ψ + 2 Ej · Ek · ∇Ej (∇Ek ψ) − ∇Ek ∇Ej ψ .
j j6=k

There is an invariant second derivative ∇2 ψ ∈ C ∞ M, E ⊗ Σ(M ) ⊗ T 0,2 (M )




defined by
∇2X,Y ψ = ∇X (∇Y ψ) − ∇∇X Y ψ.
2
The value of ∇(X,Y ) ψ at x ∈ M depends on Xx but is independent of how Xx is
extended, this is also true for Yx , but less obviously so. Indeed,
∇2X,Y ψ − ∇2Y,X ψ = ∇X (∇Y ψ) − ∇Y (∇X ψ) − ∇∇X Y −∇Y X ψ
= ∇X (∇Y ψ) − ∇Y (∇X ψ) − ∇[X,Y ] ψ
 2 
= Dε⊕ω ψ (X, Y ) = Ωε⊕ω (X, Y ) ψ,

and we know that Ωε⊕ω (X, Y ) is independent of extensions. The trace of ∇2 ψ is


the connection Laplacian, which — since ∇Ej Ek = 0 at x, is given at x by
X
∇2Ej ,Ej ψ = ∇Ej ∇Ej ψ − ∇∇Ej Ej ψ = ∇Ej ∇Ej ψ .
 
(17.29) ∆ψ :=
j

Remark 17.26. The connection Laplacian ∆ coincides with


−δ ε⊕ω ◦ Dε⊕ω : Ω0 (M, E ⊗ Σ(M )) → Ω0 (M, E ⊗ Σ(M )) ,
where Dε⊕ω : Ω0 (M, E ⊗ Σ(M )) → Ω1 (M, E ⊗ Σ(M )) denotes the covariant deriv-
ative and δ ε⊕ω is its formal adjoint, namely the covariant codifferential given in
Proposition 15.19.

We have ∇Ej (∇Ek ψ) − ∇Ek ∇Ej ψ = Ωε⊕ω (Ej , Ek ) ψ, and so
X
D2 ψ = −∆ψ + 21 Ej · Ek · Ωε⊕ω (Ej , Ek ) ψ
j6=k
X  X 
= −∆ψ + 1
2 Ej · Ek · (Ωε (Ej , Ek ) ⊗ Id)(ψ) − 1
4 Rhijk Eh · Ei · ψ
h,i
j6=k
X X
= −∆ψ + 1
2 Ωεjk Ej · Ek ψ − 1
8 Rhijk Eh · Ei · Ej · Ek · ψ,
j,k h,i,j,k
17.2. SPIN STRUCTURES AND TWISTED DIRAC OPERATORS 537

where Ωεjk = Ωε (Ej , Ek ) ⊗ Id ∈ Ω2 (M, End(E ⊗ Σ(M ))). We show that


X
(17.30) − 18 Rhijk Ej · Ek · Eh · Ei ψ = 14 Sψ,
h,i,j,k
P
where S = h,i Rhihi denotes the scalar curvature of M . For i, j, k distinct Ei ·
Ej · Ek is invariant under cyclic permutation. Hence, by the first Bianchi identity,
X X
Rhijk Ei Ej Ek = 13 (Rhijk + Rhkij + Rhjki ) Ei Ej Ek = 0.
i,j,k distinct i,j,k distinct

Thus, in view of the symmetries of Rhijk , we may assume that in the sum on the
left of (17.30), no three indices are distinct. Then (17.30) follows from
X X
Rhijk Ej Ek Eh Ei = 2 Rhihi Eh Ei Eh Ei
h,i,j,k h,i
X X
= −2 Rhihi Eh Eh Ei Ei = −2 Rhihi = −2S.
h,i h,i

In summary, we have
Proposition 17.27. Let Ωε ∈ Ω2 (M, End(E)) denote the curvature of the
connection ε on U (E) and let S denote the scalar curvature of M . For ψ ∈
C ∞ (E ⊗ Σ(M )), and an orthonormal frame E1 , . . . , En at x ∈ M , let
X
(Rε ψ)(x) := 12 Ωεjk Ej · Ek · ψ(x) ,
j,k

where Ωεjk ε 2
= Ω (Ej , Ek ) ⊗ Id ∈ Ω (M, End(E ⊗ Σ(M ))). We have
(17.31) D2 ψ = −∆ψ + Rε ψ + 14 Sψ.
Corollary 17.28. Let M be a compact Riemannian manifold with spin struc-
ture and let E be a Hermitian vector bundle over M . If the symmetric transforma-
tion
Rε + 14 S (x) ∈ End(Ex ⊗ Σx (M ))

(17.32)
is nonnegative semi-definite at each x ∈ M , then Dψ = 0 ⇒ ∇ψ = 0; i.e., all
harmonic twisted spinors on M are parallel. If M is connected and (17.32) is
nonnegative semi-definite at each x ∈ M and positive definite at some x0 ∈ M ,
then Ker D = 0, i.e., there are no nonzero twisted harmonic spinors on M . In
particular (taking E = 0), if S ≥ 0 and S 6= 0, then there are no nonzero harmonic
spinors on M .
Proof. By Proposition 17.25 and Remark 17.26,
2
kDψk = D2 ψ, ψ = −∆ψ + Rε ψ + 41 Sψ, ψ
 

= δ ε⊕ω ◦ Dε⊕ω ψ, ψ + Rε + 14 S ψ, ψ
  

2
= k∇ψk + Rε + 14 S ψ, ψ ≥ 0,
 
(17.33)
where the inequality is strict if ∇ψ 6= 0. Thus, Dψ = 0 ⇒ ∇ψ = 0. Now
 
2
∇ψ = 0 =⇒ d |ψ| = h∇ψ, ψi + hψ, ∇ψi = 0
2
=⇒ |ψ| constant =⇒ ψ = 0 or ψ(x) 6= 0 for all x ∈ M .
538 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

1
ε

Hence, if Dψ = 0 for ψ 6= 0, then ψ(x 0 ) 6= 0 and assuming that R + 4 S (x 0 ) is
positive definite, we have Rε + 14 S ψ, ψ > 0 at x0 and so Rε + 1

 4 S ψ, ψ > 0,
which contradicts (17.33). Thus, Dψ = 0 ⇒ ψ = 0, if Rε + 14 S (x0 ) is positive
definite for some x0 . 

3. The Spinorial Heat Kernel


Recall from (17.28) that we have a pair of self-adjoint elliptic operators
2
D+ := D2 |C ∞(E⊗Σ+(M )) = D− ◦ D+ ,
2
D− := D2 |C ∞(E⊗Σ−(M )) = D+ ◦ D− .
Index, Spectral Asymmetry and the Existence of the Heat Kernel.
For λ ∈ C, let
2
:= ψ ∈ C ∞ E ⊗ Σ± (M ) : D±
2
  
Vλ D± ψ = λψ .
From the general theory of formally self-adjoint, elliptic operators on compact man-
ifolds, we know that
2 2
  
Spec D± = λ ∈ C : Vλ D± 6= {0}
2
consists of the eigenvalues of D± and is a discrete subset of [0, ∞), the eigenspaces
±
Vλ D± are finite-dimensional, and an L2 (E
2
 ⊗ Σ (M ))-complete orthonormal
 set

2 + 2 2
of vectors can be selected
 from the V λ D± . Note that D V λ D+ ⊆ Vλ D− ,
2
since for ψ ∈ Vλ D+
2
D+ ψ = D+ ◦ D− D+ ψ = D+ D− ◦ D+ (ψ)
    
D−
= D+ D+ 2
(ψ) = D+ (λψ) = λD+ (ψ) ,


and similarly D− Vλ D− 2
 2

⊆ Vλ D+ . For λ 6= 0,
D± |Vλ(D2 ) : Vλ D± 2 2
 
→ Vλ D∓
±

is an isomorphism, since it has inverse λ1 D∓ . Thus the set of nonzero eigenvalues


2 2
(and their multiplicities) of D+ coincides with that of D− . However, in general
2 2 2 2
= index D+ 6= 0.
    
dim V0 D+ − dim V0 D− = dim Ker D+ − dim Ker D−
2
 2

Since dim Vλ D+ − dim Vλ D− = 0 for λ 6= 0, obviously
+ 2 2
  
index D = dim V0 D+ − dim V0 D−
X
e−tλ dim Vλ D+2 2
 
= − dim Vλ D− .
λ∈Spec(D+ 2
)
The point is that the sum can be expressed as the integral of the supertrace of
the heat kernel for the spinorial heat equation ∂ψ 2
∂t = −D ψ, from which the local
index theorem for D+ will eventually follow. However, first we need to establish
the existence of the heat kernel.
2
Let the positive eigenvalues of D± be placed in an ascending sequence 0 < λ1 ≤
λ2 ≤ λ3 ≤ . . . where each eigenvalue is repeated according to its multiplicity. Let
u± ± 2
1 , u2 , . . . be an L -orthonormal sequence in C ∞ (E ⊗ Σ+ (M )) with D± 2

j =
± ± + +
2
 2
λj uj (i.e., uj ∈ Vλj D± ). We let u01 , . . . , u0 + be an L -orthonormal basis of
n
Ker D+ 2
= Ker D+ , and u− − 2
01 , . . . , u0n− be an L -orthonormal basis of Ker D− =
2

±
Ker D− . We can pull back the bundle E ⊗ Σ (M ) via either of the projections
17.3. THE SPINORIAL HEAT KERNEL 539

M × M ×(0, ∞) → M given by π1 (x, y, t) := x and π2 (x, y, t) := y and take the


tensor product of the results to form a bundle
K± := π1∗ E ⊗ Σ± (M ) ⊗ π2∗ E ⊗ Σ± (M ) → M × M ×(0, ∞) .
 

Proposition 17.29. For t > t0 > 0, the series k 0± , defined by



X
k 0± (x, y, t) := e−λj t u± ±
j (x) ⊗ uj (y) ,
j=1

converges uniformly in C (K |M ×M ×(t0 ,∞) ) for all q ≥ 0. Hence k 0± ∈ C ∞ (K± ),


q ±

and (for t > 0)



∂ ± X
(17.34) k (x, y, t) = − λj e−λj t u± ± 2 ±
j (x) ⊗ uj (y) = −D± k (x, y, t) .
∂t j=1

Proof. Recall from the Sobolev Embedding Theorem (equation (16.48), p. 498),
that there are constants ck > 0 such that
n
u± ±
j C q ≤ ck uj 2,k for 0 ≤ q < k − .
2
Thus for
n
k(q) := q + + τ,
2
where τ > 0 so that q < k − n2 , we have
2
e−λj t u± ±
j ⊗ uj Cq
≤ e−λj t u±
j Cq

j Cq
≤ e−λj t c2k(q) u±
j 2,k(q)
.
Moreover, by the Fundamental Elliptic Estimate (Proposition 16.23, p. 498), there
are constants Ck independent of j, such that
 
u±j 2,k+2 ≤ Ck D±
2 ±
u j 2,k + u ±
j 2
 
= Ck λj uj 2,k + uj 2 ≤ Ck (λj + 1) u±
± ±
j 2,k .

As u±
j 2,0
0
= 1, iteration yields constants Ck(q) , such that
[k(q)/2] [k(q)/2]

j 2,k(q)
0
≤ Ck(q) (λj + 1) u±
j 2,0
0
= Ck(q) (λj + 1) and
[k(q)/2]
(17.35) u±
j Cq
≤ ck(q) u±
j 2,k(q)
0
≤ ck(q) Ck(q) (λj + 1) .
Combining the above estimates, we then have (for j sufficiently large such that
λj ≥ 1)
2
e−λj t u± ±
j ⊗ uj Cq
≤ e−λj t c2k(q) u±
j 2,k(q)
k(q) k(q)
≤ e−λj t c2k(q) Ck(q)
02
(λj + 1) ≤ e−λj t c2k(q) Ck(q)
02
(2λj )
00 k(q)
≤ Ck(q) e−λj t λj 00
, where Ck(q) 02
:= 2k(q) c2k(q) Ck(q) .
In order to apply the Weierstrass M-test to deduce the (uniform) convergence of
P∞ −λj t ± P∞ −λj t k(q)
j=1 e uj ⊗ u± q
j in C (K|M ×M ×(t,∞) ), we need to show that j=1 e λj
converges; note that the inclusion of the arbitrary positive parameter τ in the
definition of k(q) is designed to handle the uniform C q convergence in t as well. It
is easy to check that
k
e−x/2 xk ≤ Dk := e−k (2k) ,
540 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

by noting that
   1
d
dx e−x/2 xk = xk−1 − 12 x + k e− 2 x = 0 ⇐= x = 2k.

Thus,

k(q) k(q)
e−λj t/2 (λj t) ≤ Dk(q) , e−λj t λj ≤ Dk(q) t−k(q) e−λj t/2 , and so
∞ ∞
k(q)
X X
e−λj t λj ≤ Dk(q) t−k(q) e−λj t/2 .
j=1 j=1

If we can show that λj ≥ Cj α for some positive constants C and α, then

∞ ∞
X X α
e−λj t/2 ≤ e−Cj t/2
< ∞,
j=1 j=1

by comparison with an integral of the form (for α,β ∈ (0, ∞))


Z ∞  
α 1 −1 1
e−βx dx = β αΓ < ∞.
0 α α
α
Thus, itnremains to show that o λj ≥ Cj forPsome positive constants C and α. Let
j j
Vj± := ± ±
= i=1 ai u± ± n
P
i=1 ai ui : ai ∈ C . For u i ∈ Vn and k > 2 , we have
(using Proposition 16.23, p. 498)
[k/2]
u± ≤ ck u± 2


C0 2,k
≤ ck Ck D± +1
2,0
[k/2] ±
≤ ck Ck (λj + 1) u 2,0
X 1
j 2
[k/2] 2
= Ck00 (λj + 1) |ai | ,
i=1

where we have used the characterization


[k/2]
2
D± +1 (u± )
[k/2] 2,0
(λj + 1) = sup .
u± ∈Vj± ku± k2,0

Hence, for any x ∈ M and ai ∈ C,

X 1
Xj j 2
[k/2] 2
ai u±
i (x) ≤ Ck00 (λj + 1) |ai | .
i=1 i=1

Let F1 (x) , . . . , FN (x) (N = dim E · dim Σ±


n ) be an orthonormal basis for Ex ⊗

Σ± (M )x , and make the specific choice ai := u± i (x) , Fh (x) for fixed h ∈ {1, . . . , N }.
Then we are one step closer to find the wanted positive constants C and α such
17.3. THE SPINORIAL HEAT KERNEL 541

that λj ≥ Cj α . Indeed, not summing over h, we obtain


Xj 2 Xj

i (x) , Fh (x) = ai u±
i (x) , Fh (x) |Fh (x)|
i=1 i=1
Xj Xj
= ai u±
i (x) , Fh (x) Fh (x) ≤ ai u±
i (x)
i=1 i=1
X 1
j 2
[k/2] 2
≤ Ck00 (λj + 1) |ai |
i=1
X 1
j 2 2
[k/2]
= Ck00 (λj + 1) u±
i (x) , Fh (x) .
i=1

P 2
1
j 2
Dividing this by i=1 u±
i (x) , Fh (x) and squaring, we have
Xj 2 k

j (x) , Fh (x) ≤ Ck002 (λj + 1) .
i=1

Now summing over h, we get


j XN Xj
X 2 2 k

i (x) = u±
i (x) , Fh (x) ≤ N Ck002 (λj + 1) .
h=1 i=1
i=1

Integrating over M , we obtain


j
X 2 k
j= u±
i 2,0
≤ N Ck002 (λj + 1) Vol(M ) .
i=1

Thus, for k > n/2 and j sufficiently large, we have the desired result
  k1   k1
j 1 1
(17.36) λj ≥ − 1 ≥ 2 jk. 
Ck002 Vol(M ) Ck002 Vol(M )

Solving the Spinorial Heat Equations. We specify and justify our termi-
nology.
Definition 17.30. The positive and negative twisted spinorial heat ker-
2
nels (or the heat kernels for D± ) k ± ∈ C ∞ (K± ) are given by
±
n
X ∞
X
±
k (x, y, t) := u± ±
0i (x) ⊗ u0i (y) + e−λj t u± ±
j (x) ⊗ uj (y) .
i=1 j=1

The total twisted spinorial heat kernel (or the heat kernel for D2 ) is
k = k + , k − ∈ C ∞ K+ ⊕ C ∞ K− ∼ = C ∞ K+ ⊕ K− ⊆ C ∞ (K) ,
   

(17.37) where K := π1∗ (E ⊗ Σ(M )) ⊗ π2∗ (E ⊗ Σ(M )) .


The terminology is justified in view of the following
Proposition 17.31. Let ψ0± ∈ C ∞ (E ⊗ Σ± (M )) and let
Z
±
ψ (x, t) = k ± (x, y, t) , ψ0± (y) y ν(y) ,
M
542 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

where ν denotes the volume form of M . Then for t > 0, ψ ± solves the heat equation
with initial spinor ψ0± :
∂ψ ±
= −D± 2 ±
ψ and lim ψ ± (·, t) = ψ0± in C q for all q ≥ 0.
∂t t→0+

Moreover, for ψ0 ∈ C ∞ (E ⊗ Σ(M )) and


Z
ψ(x, t) := hk(x, y, t) , ψ0 i ν(y),
M
∂ψ
we have = −D2 ψ and limt→0+ ψ(·, t) = ψ0 (·) in C q for all q ≥ 0.
∂t
Proof. Since k ± is C ∞ , we may differentiate under the integral and use (17.34)
∂ψ ±
to deduce (note t > 0) that = −D± 2
ψ. We now show that limt→0+ ψ ± (·, t) =
∂t
ψ0± in C q for all q ≥ 0. First note that
Z
±
ψ (x, t) = k ± (x, y, t) , ψ0± (y) y ν(y)
M
Z * n± ∞
+
X X
= u±
0i (x) ⊗ u±
0i (y) + e−λj t u±
j (x) ⊗ u± ±
j (y) , ψ0 (y) ν(y) .
M i=1 j=1

Using the uniform convergence of the series to interchange sum and integral we
obtain
n± Z
X 
±
ψ (x, t) = u±0i (y) , ψ0
±
(y) ν(y) u±
0i (x)
i=1 M

X Z 
+ e −λj t
u± ±
j (y) , ψ0 (y) ν(y) u±
j (x)
j=1 M


X ∞
X
u± ±
 ±
e−λj t u± ±
 ±
= 0i , ψ0 u0i (x) + j , ψ0 uj (x) .
i=1 j=1

It remains to prove that in the C norm, the limit as t → 0+ may be taken under
q

the infinite sum. This is permitted if



X
u± ±
 ±
j , ψ0 uj (x) Cq
< ∞.
j=1

Since ψ0± is C ∞ , for any l = 0, 1, . . . , we have


2 l ±
ψ0 ∈ C ∞ E ⊗ Σ± (M ) ⊆ L2 E ⊗ Σ± (M ) , and
  

   
2 l ± 2 l ±
u± ±
= λlj u± ±
  
j , D± ψ0 = D ± uj , ψ0 j , ψ0 2,0 .
2,0 2,0
Thus, for any l = 0, 1, . . . ,
∞ 2 2
2 l ±
X
λ2l u± ±
 
j j , ψ0 2,0
= D± ψ0 < ∞, and so
j=1

u± ±
≤ Kl λ−l

j , ψ0 2,0 j , for some Kl > 0.
17.3. THE SPINORIAL HEAT KERNEL 543

[k(q)/2]
Using u±
j Cq
0
≤ ck(q) Ck(q) (λj + 1) (see 17.35), we have

X ∞
X
u± ±
 ±
j , ψ0 uj Cq
≤ Kl λ−l
j u±
j Cq
j=1 j=1
X∞
[k(q)/2]
≤ 0
Kl ck(q) Ck(q) λ−l
j (λj + 1) < ∞,
j=1

since we can choose l arbitrarily large and we have shown (see (17.36)) that λj ≥
Cj α for some positive constants C and α. Thus, the sum
±
n
X ∞
 ± X
u± ±
u± ±
 ±
0i , ψ0 u0i + j , ψ0 uj ,
i=1 j=1

which is known to converge in L (E ⊗ Σ (M )) to ψ0± , in fact converges to ψ0± in


2 ±

C q (E ⊗ Σ± (M )). By dominated convergence,



X ∞
X
u± ±
u± u± ±
−λj t
  ±
e j , ψ0 j Cq ≤ j , ψ0 uj Cq
< ∞,
j=1 j=1

implies that we have the following limit in C q :


±
n
X ∞
 ± X
u± ±
lim+ e−λj t u± ±
±
 ±
lim+ ψ (·, t) = 0i , ψ0 u0i + j , ψ0 uj
t→0 t→0
i=1 j=1
±
n
X ∞
X
u± ±
u± u± ±
  ± ±
= 0i , ψ0 0i + j , ψ0 uj = ψ0 .
i=1 j=1

The analogous assertions for ψ with ψ0 are proved in the same way. 
Calculating Index and Supertrace. Note that for x ∈ M , the Hermitian
inner product h , ix on (E ⊗ Σ(M ))x gives us a conjugate-linear bijective map ψ 7→

ψ ∗ (·) := h · , ψix from (E ⊗ Σ(M ))x to its dual (E ⊗ Σ(M ))x . Thus, for t > 0, we
may regard  
k(x, y, t) ∈ Hom (E ⊗ Σ(M ))x ,(E ⊗ Σ(M ))y ,
and similarly for k ± (x, y, t). For any finite dimensional Hermitian vector space
(V, h·, ·i) with orthonormal basis e1 , . . . , eN , we have (for v ∈ V )
XN XN
Tr(v ∗ ⊗ v) = h(v ∗ ⊗ v)(ei ) , ei i = hv ∗ (ei ) v, ei i
i=1 i=1
XN XN XN 2 2
= hhei , vi v, ei i = hei , vi hv, ei i = |hei , vi| = |v| .
i=1 i=1 i=1
In particular, k ± (x, x, t) ∈ End((E ⊗ Σ± (M ))x ) and
±
n ∞
X 2 X 2
±
u± e−λj t u±

Tr k (x, x, t) = 0i (x) + j (x) .
i=1 j=1

Since this series converges uniformly and = u± u±


j 2,0 = 1, we have
0i 2,0
Z ∞
X
Tr k ± (x, x, t) νx = n± + e−λj t < ∞.

M j=1
544 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

2
For t > 0, we define the bounded linear operator e−tD± ∈ End L2 (E ⊗ Σ± (M ))


by
n± ∞
2  X  ± X
u± ±
e−λj t u± ±
−tD± ±
 ±
e ψ = 0i , ψ0 u0i + j , ψ0 uj .
i=1 j=1
2
−tD±
Note that e is of trace class, since
  ∞  Z
2
−tD± ±
X
−λj t
 
Tr e = n + e = Tr k ± (x, x, t) νx < ∞.
j=1 M

Now, we have

X

+ 2 2 +
e−λj t − e−λj t
   
index D = dim V0 D+ − dim V0 D− =n −n +
j=1

X ∞
X
= n+ + e−λj t − n− + −λj t

e
j=1 j=1
Z Z
Tr k + (x, x, t) νx − Tr k − (x, x, t) νx
 
=
ZM M

Tr k + (x, x, t) − Tr k − (x, x, t) νx .
 
(17.38) =
M
2
, we also have the operator e−tD ∈ End L2 (E ⊗ Σ(M )) of
2 2 2

Since D = D+⊕ D−
trace-class. Its trace is given by
 2
 Z Z
Tr e−tD = Tr k + (x, x, t) + Tr k − (x, x, t) νx .
 
Tr(k(x, x, t)) νx =
M M
The supertrace of k(x, x, t) is defined by
Str(k(x, x, t)) := Tr k + (x, x, t) − Tr k − (x, x, t) ,
 

and in view of (17.38), we have


Z
+

(17.39) index D = Str(k(x, x, t)) νx .
M
The left side is independent of t and so the right side is also independent of t.
The main task now is to determine the behavior of Str(k(x, x, t)) as t → 0+ . We
suspect that for each x ∈ M , as t → 0+ , k(x, x, t) and Str(k(x, x, t)) are influenced
primarily by the geometry (e.g., curvature of M and E) near x, since the heat
sources of points far from x are not felt very strongly at x for small t. In the
sections to follow, we show that
*  12 +
iΩθ /4π
 ε  
iΩ /2π
(17.40) lim Str(k(x, x, t)) = Tr e ∧ det , νx ,
t→0+ sinh(iΩθ /4π)
where the meaning of the right side is explained in the following digression.
The curvature form of the connection ε on E is denoted by Ωε , while Ωθ denotes
the curvature form of the Levi-Civita connection θ for F M . We have (recall 2m =
dim M )
∞  k m  k
iΩε /2π
X 1 i ε k ε
X 1 i k
(17.41) e := Ω ∧ ··· ∧ Ω = Ωε ∧ · · · ∧ Ωε ,
k! 2π k! 2π
k=0 k=0
17.3. THE SPINORIAL HEAT KERNEL 545

k
 k

where Ωε ∧ · · · ∧ Ωε ∈ Ω2k (End(E)). Also Tr Ωε ∧ · · · ∧ Ωε ∈ Ω2k (M ) and
 ε  M m
Tr eiΩ /2π ∈ Ω2k (M ) .
k=1

This (by one of manyLm equivalent definitions) is a representative of the total Chern
character ch(E) ∈ k=1 H 2k (M ; Q). Now Ωθ ∈ Ω2 (End(T M )) has values in the
skew-symmetric endomorphisms of T M . A skew-symmetric endomorphism of R2m ,
say B ∈ so(n), has pure imaginary eigenvalues ±irk , where rk ∈ R (1 ≤ k ≤ m).
z/2
Thus, iB has real eigenvalues ±rk . Now is a power series in z with radius
sinh(z/2)
isB/2
of convergence 2π. Thus, is defined for s sufficiently small and has
sinh(isB/2)
rk s/2
eigenvalues each repeated twice. Hence
sinh(rk s/2)
  m  2
isB/2 Y rk s/2
det = and
sinh(isB/2) sinh(rk s/2)
k=1
 1 m
isB/2 2 Y rk s/2
det = .
sinh(isB/2) sinh(rk s/2)
k=1

The last product is a power series in s of the form


m ∞
Y rk s/2 X
ak r12 , . . . , rm
2
 2k
(17.42) = s ,
sinh(rk s/2) j=0
k=1

where the coefficient ak r12 , . . . , rm 2
is a homogeneous, symmetric polynomial in
r12 , . . . , rm
2
of degree k. One can always express any such a symmetric polynomial
as a polynomial in the elementary symmetric polynomials σ1 , . . . , σm in r12 , . . . , rm
2
,
where
Xm Xm Xm
σ1 := ri2 , σ2 := ri2 rj2 , σ3 := ri2 rj2 rk2 , . . . .
i=1 i<j i<j<k

These in turn may be expressed in terms of SO(n)-invariant polynomials in the


entries of B ∈ so(n) via
m
Y m
Y
λ2 + rj2

det(λI − B) = (λ + irj )(λ − irj ) =
j=1 j=1
Xm
σk r12 , . . . , rm
2
 2(m−k)
= λ .
k=1

On the other hand,


 
m
X 1 X ···j2k
det(λI − B) =  sgnji11···i B ij11 · · · B ij2k2k  λ2(m−k) , and so
(2k)! 2k
k=1 (i),(j)
1 X ···j2k
σk r12 , . . . , rm
2
sgnji11···i B ij11 · · · B ij2k2k ,

=
(2k)! 2k
(i),(j)
546 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

where (i) = (i1 , · · · , i2k ) is an ordered 2k-tuple of distinct elements of {1, . . . , 2m}
···j2k
and (j) is a permutation of (i) with sign sgnij11···i 2k
. If we replace B ij with the
1
 i
2-form 2π Ωθ j relative to an orthonormal frame field, we obtain the Pontryagin
forms
1 X ···j2k θ
p k Ωθ = sgnji11···i Ωi1 j1 ∧ · · · ∧ Ωθi2k j2k ,

2k 2k
(2π) (2k)! (i),(j)

which represent the Pontryagin classes of the SO(n) bundle F M. Note that pk Ωθ
is independent of the choice of framing by the ad-invariance of the polynomials σk .
(The concept of ad-invariant metrics was discussed at the beginning  of Section 16.1,
p.460). Getting back to (17.40), if we express the ak r12 , . . . , rm2
as polynomials,
say Ak (σ1 , . . . , σk ), in the σj (j ≤ k), we can ultimately write

 1 ∞
isB/2 2 X
det = Ak (σ1 , . . . , σk ) s2k .
sinh(isB/2)
k=0

1 θ
Formally replacing B by 2π Ω , we finally have motivated the definition

 21 ∞
iΩθ /4π
 X
Ak p1 Ωθ , . . . , pk Ωθ ,
 
det :=
sinh(iΩθ /4π)
k=0

θ

where the pj Ω are multiplied via wedge product when evaluating the forms
Ak p1 Ωθ , . . . , pk Ωθ ; the order of multiplication
 doesnot matter since pj Ω
θ
θ θ
is of even degree 4j. Also, since Ak p1 Ω , . . . , pk Ω is a 4k-form, there are
only a finite number of nonzero terms in the infinite sum. Abbreviating pj Ωθ
simply by pj , one finds

 12
iΩθ /4π

1 1
7p21 − 4p2

det =1− p1 +
sinh(iΩθ /4π) 24 5760
1
31p31 − 44p1 p2 + 16p3 + · · · .

(17.43) −
967 680

This (by one definition) represents the total A-class


b of M , denoted by
m
M
(17.44) A(M
b )∈ H 2k (M ; Q) ,
k=1

where actually A(M


b ) has nonzero components only in H 2k (M ; Q) when k is even
(or 2k ≡ 0 mod 4). In (17.40) the multi-degree forms (17.41) and (17.43) have been
wedged, and the top (2m-degree) component (relative to the volume form) has been
harvested. In terms of characteristic classes, (17.39) and (17.40) yield
  
(17.45) index D+ = ch(E) ` A(M b ) [M ] ,

which is the index formula for the twisted Dirac operator D+ .


17.3. THE SPINORIAL HEAT KERNEL 547

General Heat Kernels. In the next section, we give a different construction


of the total twisted spinorial heat kernel, which yields its asymptotic expansion as
t → 0+ . In order to deduce that this construction actually yields the same result
as (17.37), we introduce the definition of a general heat kernel for D2 . Propositions
17.29 and 17.31 imply that the heat kernel (17.37) satisfies the criteria in this
definition. We prove below that any two general heat kernels are the same. Hence,
if a differently constructed function also fits this definition, it must agree with the
heat kernel (17.37).
We set

H := π1∗ (E ⊗ Σ(M )) ⊗ π2∗ (E ⊗ Σ(M )) → M × M ×(0, ∞) ,
where π1 (x, y, t) := x and π2 (x, y, t) := y. We fix a Riemannian metric h on M and
denote the induced volume form by νh .
Definition 17.32. A general heat kernel for D2 is a section κ ∈ C 0 (H),
where κ(x, y, t) ∈ Hom (E ⊗ Σ± (M ))x ,(E ⊗ Σ± (M ))y is C 2 in x, C 1 in t, and
∂t + Dx2 κ(x, y, t) = 0. Moreover, for all ψ ∈ C 0 (E ⊗ Σ± (M )), we require

Z
(17.46) lim κ(y, x, t) ψ(y) νh (y) = ψ(x) .
t→0+ M

Remark 17.33. Here and elsewhere, the phrase “C 2 in x” means not only that
κ(x, y, t) is C 2 in x for fixed t, but that the second partials with respect to local
coordinates for x are jointly continuous in (x, y, t). Also “C 1 in t” means that
∂t κ(x, y, t) is jointly continuous in (x, y, t).
Lemma 17.34. Let π : M × [0, ∞) → M be given by π(x, t) = x, and let ψ1 ,
ψ2 ∈ C 0 (π ∗ (E
 ⊗ Σ(M ))) be C 2 in x and C 1 in t. Suppose that ∂t + D2 ψ1 = 0
and ∂t + D2 ψ2 = 0. Then for any t > 0, we have
Z
d
hψ1 (x, t − s) , ψ2 (x, s)ix νh (x) = 0 for s ∈ (0, t) .
ds M
R
In other words, M hψ1 (x, t − s) , ψ2 (x, s)ix νh (x) is independent of s in [0, t].
Proof. For L := ∂t +D2 (here we emphasize that D2 solely contains derivatives
in the space directions by writing Dx2 ),
0 = h(Lψ1 )(x, t − s) , ψ2 (x, s)i − hψ1 (x, t − s) ,(Lψ2 )(x, s)i
= h− (∂s ψ1 )(x, t − s) , ψ2 (x, s)i − hψ1 (x, t − s) , (∂s ψ2 )(x, s)i
+ Dx2 ψ1 (x, t − s) , ψ2 (x, s) − ψ1 (x, t − s) , Dx2 ψ2 (x, s)
d
= ds (hψ1 (x, t − s) , ψ2 (x, s)i)
+ Dx2 ψ1 (x, t − s) , ψ2 (x, s) − ψ1 (x, t − s) , Dx2 ψ2 (x, s) .
Since D2 is formally self-adjoint, integrating over M , we then have
Z
d
hψ1 (x, t − s) , ψ2 (x, s)ix νh (x)
ds M
Z
d
= ds hψ1 (x, t − s) , ψ2 (x, s)ix νh (x) = 0,
M

since the integrand hψ1 (x, t − s) , ψ2 (x, s)ix is C 1 in t. 


548 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

Proposition 17.35. If κ is a general heat kernel for D2 , then κ(y, x, t) =



κ(x, y, t) . Moreover, if κ1 and κ2 are general heat kernels for D2 , then κ1 = κ2 .

Proof. For any fixed α ∈ (E ⊗ Σ± (M ))x and β ∈ (E ⊗ Σ± (M ))y , we set


∗ ∗
ψ1 (z, t) := κ1 (z, x, t) (α) and ψ2 (z, t) := κ2 (z, y, t) (β) for (z, t) ∈ M × (0, ∞).
Then

∂t + Dz2 ψ1 = ∂t + Dz2 κ1 (z, x, t) (α)
  
∗
= ∂t + Dz2 κ1 (z, x, t) (α) = 0,



and similarly ∂t + Dz2 ψ2 = 0. According to (17.46),
 Z 
hψ1 (y, t) , βi = lim+ κ2 (z, y, s) ψ1 (z, t − s) νh (z) , β
s→0 M y
Z
= lim+ hκ2 (z, y, s) ψ1 (z, t − s) , βiy νh (z)
s→0
ZM

= lim+ ψ1 (z, t − s) , κ2 (z, y, s) (β) z
νh (z)
s→0 M
Z
= lim hψ1 (z, t − s) , ψ2 (z, s)iz νh (z) , and
s→0+ M

 Z 
hα, ψ2 (x, t)i = α, lim− κ1 (z, x, t − s) ψ2 (z, s) νh (z)
s→t M x
Z
= lim− hα, κ1 (z, x, t − s) ψ2 (z, s)ix νh (z)
s→t
ZM

= lim− κ1 (z, x, t − s) (α) , ψ2 (z, s) z νh (z)
s→t
ZM
= lim hψ1 (z, t − s) , ψ2 (z, s)iz νh (z) .
s→t− M

Hence, using Lemma 17.34,



κ1 (y, x, t) (α) , β = hψ1 (y, t) , βi
Z
= lim+ hψ1 (z, t − s) , ψ2 (z, s)iz νh (z)
s→0
ZM
= lim hψ1 (z, t − s) , ψ2 (z, s)iz νh (z)
s→t− M

= hα, ψ2 (x, t)i = α, κ2 (x, y, t) (β) ,

from which we have



κ2 (x, y, t) = κ1 (y, x, t) .

In the case κ1 = κ2 = κ, we have κ(x, y, t) = κ(y, x, t) , and so

κ1 (x, y, t) = κ1 (y, x, t) = κ2 (x, y, t) . 
17.4. THE ASYMPTOTIC FORMULA FOR THE HEAT KERNEL 549

4. The Asymptotic Formula for the Heat Kernel


“Exact answers are great when we can find them; there’s something
very satisfying about complete knowledge. But there’s also a time when
approximations are in order. If we run into a sum or a recurrence whose
solution doesn’t have a closed form (as far as we can tell), we still would
like to know something about the answer; we don’t have to insist on all
or nothing. And even if we do have a closed form, our knowledge might
be imperfect, since we might not know how to compare it with other
closed forms.” (R. Graham, D. Knuth, and O. Patashnik, [184,
p. 425])

I Asymptotic expansions for the heat kernel have many different applications and
can be obtained in various ways. In the following, we present one way. It is very close to
classical analysis, but takes 25 pages. Once over that hurdle, it leads immediately to the
Local Index Formula. J

Defining Asymptotic Expansions Rigorously. In mathematics and math-


ematical physics, the terms asymptotic solution and asymptotic expansion are heav-
ily loaded with meaning. Often they are used in talk, informally, and on risk of
ambiguities. Before inviting the reader to that informed community we shall recall
a few rigorous definitions, mostly following [356, p. 27].
Definition 17.36. (a) (Paul Bachmann, 1894). Let f and g be holomorphic
functions in domain G ⊂ C and z0 a point of the boundary of G. We write
f (z) = O(g(z)) (“f (z) is Big Oh of g(z)”) for z → z0
if there exist C, δ ∈ R, C, δ > 0, such that |f (z)| ≤ C|g(z)| whenever |z − z0 | ≤ δ.
We write
f (z) = o(g(z)) (“f (z) is Small Oh of g(z)”) for z → z0
if there exist δ ∈ R, δ > 0, to each ε ∈ R, ε > 0, such that |f (z)| ≤ ε|g(z)| whenever
|z − z0 | ≤ δ.
(b) (Asymptotic Expansion 1). Let w0 , w1 , w2 , . . . be a sequence of holomorphic
functions in a domain G ⊂ C satisfying the conditions
(17.47) wn (z) 6≡ 0 and wn+1 (z) = o(wn (z)) for z → z0
for every n = 0, 1, 2, . . . , where z0 is a point
P∞ of the boundary of G. Let f be a
holomorphic function in G. A formal series ν=0 aν wν (z) is called an asymptotic
expansion of f for z → z0 with respect to the sequence w0 , w1 , w2 , . . . , if
n
X
(17.48) f (z) = aν wν (z) + o(wn (z)) for z → z0 .
ν=0

In this case we write



X
(17.49) f (z) ∼ aν wν (z) for z → z0 .
ν=0
550 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

Note. There are some caveats:


• While f (z) = O(g(z)) for z → z0 defines an equivalence relation f ∼ g,
the definition of f (z) = o((g(z)) for z → z0 does not. Clearly it is reflexive
only for functions identical 0 near z0 . So there is good reason to prefer
the explicit formulation of asymptotic relations. In critical cases, we shall
not succumb to the temptation of the concept of equivalence classes which
otherwise should
P∞ be preferred for aesthetic reason.
• The series ν=0 aν wν (z) is not necessarily convergent, but its partial
sums give a good approximation of f (z) near z0 . It is easy to see that
the coefficients aν in (17.48) are uniquely determined by f , if such an
expansion exists. On the other hand different functions may very well
have the same asymptotic expansion for z → z0 .
• Suppose that the domain G extends to infinity, then it is natural to con-
sider the sequence wν (z) = 1/z ν , ν = 0, 1, 2, . . . , for z → ∞. Similarly, we
shall meet sequences of the form wν (z) = z ν or wν (z) = z −ν exp(−ct−1 )
when the domain G extends to 0.
• Depending on the context, it is common to consider only functions on
the real line with t → ∞ or t → 0 in dynamical systems (like our heat
expansion); respectively, µ → ∞ or µ → 0 in ultra-violet and ultra-red
cut-off of wave lengths in quantum field theory; or functions on the natural
grid N with n → ∞ in discrete computer sciences.
• In Definition 17.37 below, we shall modify Definition 17.36 for functions
of several variables.

Why Asymptotic Expansion of the Heat Kernel? The well-known heat


kernel (or fundamental solution) for the ordinary heat conduction equation ut = ∆u
in Euclidean space Rn , is given by
 
−n/2 2
(17.50) e(x, y, t) = (4πt) exp − |x − y| /4t .

(Like in many places of Part II, we write ut for ∂t u.) Note that the signs on the right
side of the heat conduction equation and the spinorial heat equation in front of the
essentially negative standard Laplacian ∆, respectively, in front of the essentially
positive connection Laplacian D2 are opposite.
Since e(x, y, t) only depends on r = |x − y| and t, it is convenient to write
−n/2
exp −r2 /4t .

(17.51) e(x, y, t) = E(r, t) := (4πt)

We do not expect such a simple expression for the heat kernel k = (k + , k − ) of


Definition 17.30 on p. 541. However, we will show that for x, y ∈ M (of even
dimension n = 2m) with r = d(x, y) := Riemannian distance from x to y sufficiently
small, we have an asymptotic expansion as t → 0+ for k(x, y, t) of the form
XQ
k(x, y, t) ∼ HQ (x, y, t) := E(r, t) hj (x, y) tj ,
j=0

for any fixed integer Q > m + 4, where


 
hj (x, y) ∈ Hom (E ⊗ Σ(M ))x ,(E ⊗ Σ(M ))y , j ∈ {0, 1, . . . , Q} .
17.4. THE ASYMPTOTIC FORMULA FOR THE HEAT KERNEL 551

Definition 17.37 (Asymptotic Expansion 2). For d(x, y) =: r and t sufficiently


small, the meaning of k(x, y, t) ∼ HQ (x, y, t) as t → 0+ is that
|k(x, y, t) − HQ (x, y, t)| ≤ CQ E(r, t) tQ+1 ≤ CQ tQ−m+1 ,
where CQ is a constant, independent of (x, y, t).
We then will have
Q Q
−m −m
X X
(17.52) k(x, x, t) ∼ (4πt) hj (x, x) tj = (4π) hj (x, x) tj−m .
j=0 j=0

This may seem a bit odd, since we know from (17.39) that for any t > 0,
Z
Str(k(x, x, t)) νh (x) = index D+ ,

M

which is constant, independent of t. As t → 0+ , we deduce from (17.52) that


(
for j ∈ {0, 1, . . . , m − 1}, while
Z
0,
Str(hj (x, x)) νh (x) = m +
M (4π) index(D ) , for j = m.
Thus, to prove the Local Index Formula in the next chapter, it will suffice to show
that
*  12 +
iΩθ /4π
 ε  
−m iΩ /2π
(4π) Str(hm (x, x)) = Tr e ∧ det , νh (x) .
sinh(iΩθ /4π)

While this may not be the intellectual equivalent of climbing Mount Everest, it is
not for the faint of heart. We will find some shortcuts, and we encourage the reader
to find other approaches to the summit. In what follows, we develop machinery
to obtain a local formula for the heat kernel of D2 about a point x ∈ M using a
local section (known as the radial gauge) of the frame bundle, which is as simple
as possible. We will then use this to construct what ought to be the asymptotic
expansion for the heat kernel of the spinorial heat equation ψt = −D2 ψ. The
proof that this is indeed the asymptotic expansion for the heat kernel will then be
carried out. This will be the foundation for the proof of the Local Index Formula
in the next chapter.

The Radial Gauge. In this paragraph, we construct the radial gauge. For
any v ∈ Rn , we denote the standard horizontal vector field on F M (with respect
to the Levi-Civita connection 1-form θ, defined on p. 418) by v. Recall that v is
determined by the conditions
θ(v) = 0 and v = ϕ(v u ) = u−1 (πF ∗ (v u )) , for all u ∈ F M,
1
where ϕ ∈ Ω (P, Rn ) denotes the canonical 1-form. Let C : P → F M be the
spin structure for M and let πU(E) : U (E) → M denote the unitary frame bun-
dle of E with connection ε. We then have the fibered product bundle πU(E) ×f
(πF ◦ C ) : U (E) ×f P → M with connection ε ⊕ θ,e where θe := c0−1 (C ∗ θ). We
1
may pull back the canonical form ϕ to a form in Ω (U (E) ×f P, Rn ) via the map
U (E) ×f P → P → F M . For simplicity, we will denote this pullback of ϕ by the
552 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

same symbol ϕ. Moreover, we can define the notion of the standard horizontal
vector field v on U (E) ×f P for v ∈ Rn by means of the conditions
 
ε ⊕ θe (v) = 0 and v = ϕ(v) .
For fixed x ∈ M , select a frame ux ∈ (F M )x and some frame u0x ∈ P with C (u0x ) =
ux ; u0x is sometimes called a spinor frame at x. Also select a (unitary) frame
wx0 ∈ U (E)x . Then let wx := (wx0 , u0x ) ∈ (U (E) ×f P )x . For v ∈ Rn , let ηv : R →
U (E) ×f P denote the integral curve of v with initial point ηv (0) = wx . Thus, for
πf : U (E) ×f P → M ,
ηv0 (t) = v ηv(t) and ϕ(ηv0 (t)) = ϕ v ηv(t) = u−1 πf ∗ v ηv(t) = v.
 

Define ζ : Rn → U (E) ×f P by ζ(v) = ηv (1). Note that ζ(tv) = ηtv (1) = ηv (t) and
ζ(0) = ηv (0) = wx . By the implicit function theorem, there is a ball B(0, r0 ) ⊆ Rn
(with r0 > 0) about 0 ∈ Rn , such that ζ|B(0,r0 ) defines a smoothly embedded
submanifold of U (E) ×f P through wx and πf ◦ ζ|B(0,r0 ) : B(0, r0 ) → M is a diffeo-
morphism onto its image, say B. There is a unique local section σ : B → U (E) ×f P
whose image is ζ(B(0, r0 )). This local section σ is known as the radial gauge
for the choice (wx0 , u0x ) ∈ (U (E) ×f P )x ; it depends on the connection ε ⊕ θ.
e Since
U (E) ×f P ⊂ U (E) × P, the map σ has two component local sections; i.e.,
σ = σ1 × σ2 , where σ1 : B → U (E) and σ2 : B → P .
Using the spin structure C : P → F M , we also have a local section
(17.53) u := C ◦ σ2 : B → F M
of the frame bundle (i.e., a frame field). For y ∈ B and the standard basis vectors
ek ∈ Rn , we set
(17.54) Ek (y) := ((C ◦ σ2 )(y))(ek ) .
Proposition 17.38. For v ∈ Rn , the curve ξv := πf ◦ ηv is a geodesic relative
to the metric h on M with ξv0 (0) = ux (v). If ξv (1) ∈ B, the map
−1
φ := πf ◦ ζ|B(0,r0 ) : B → B(0, r0 )
assigns to the point ξv (1) ∈ B, the components relative to the frame ux of ξv0 (0),
namely
(17.55) φ(ξv (1)) = u−1 0
x (ξv (0)) = v.
In other words, φ : B → B(0, r0 ) is a normal coordinate system.
Proof. We claim that the Ek are parallel along each of the curves ξv := πf ◦ηv
in the sense that ∇θξv0 Ek = 0. Let u = C ◦ σ2 as in (17.53) and let E ek (u(y)) ∈
Tu(y) F M denote the θ-horizontal lift of Ek (y). We have that ϕη (t) (Ek ) = ek , since
e
v

ek ) = u(ξv (t))−1 (π∗ (E


ϕu(ξv(t)) (E ek )) = u(ξv (t))−1 (Ek (ξv (t)))
−1
= u(ξv (t)) (u(ξv (t))(ek )) = ek .
Then using the definition ∇θX Y := π∗ (ϕ−1 (X[ϕ( e Ye )])),
  h i
−1
∇θξv0 (t) Ek = π∗ ϕu(ξ v(t))
e0 (t)
ξ v u(ξv(t)) ϕ u(ξ v(·)) ( E
ek (ηv (·)))
  
−1
= π∗ ϕu(ξv(t)) ξev0 (t)u(ξv(t)) [ek ] = 0.
17.4. THE ASYMPTOTIC FORMULA FOR THE HEAT KERNEL 553

In particular, the curve ξv is a geodesic (for the metric h on M ), with ξv0 (0) = ux (v),
since its tangent vector field ξv0 (t) is a linear combination (with constant coefficients)
of the parallel vector fields Ek ; explicitly,
0
ξv0 (t) = (πf ◦ ηv ) (t) = πf ∗ (ηv0 (t)) = πf ∗ v ηv(t) = πF ∗ v u(ξv(t))
 

= u(ξv (t))(v) (where we used v = ϕ(v u ) = u−1 (πF ∗ (v u )) )


 Xn  Xn
= u(ξv (t)) vk e k = vk u(ξv (t))(ek )
k=1 k=1
Xn
(17.56) = vk (Ek )ξv(t) .
k=1

We get (17.55) from


ζ(v) = ηv (1) =⇒ (πf ◦ ζ)(v) = (πf ◦ ηv )(1) = ξv (1)
=⇒ φ(ξv (1)) = v = u−1 0
x (ξv (0)) .

In other words, φ is the restriction of the exponential map from Tx M to M . These


are called the normal coordinates about x relative to the frame ux . 

We denote the Cartesian coordinates of points in B(0, r0 ) ⊆ Rn by y 1 , . . . , y n ,
2 2
and we let r2 = y 1 + · · · +(y n ) . In what follows, we identify points y ∈ B with
their coordinate points in B(0, r0 ). In particular the original point x, about which
all of these constructions started, is 0 ∈ B(0, r0 ), and the curve ξv is simply given
by
ξv (t) = tv = t(v1 , . . . , vn ) .
We use the notation ∂j := ∂/∂y j for the standard coordinate vector fields. By
(17.56) we have u(ξv (t))(v) = ξv0 (t), and so
Ek (0) = ux (ek ) = u(ξek (0))(ek ) = ξe0 k (0) = d
dt (tek ) t=0 = ∂k |y=0 .
Thus, the orthonormal vector fields E1 (y) , . . . , En (y) coincide with ∂1 , . . . , ∂n at
y = 0, but not necessarily elsewhere in B(0, r0 ) unless the metric h is flat, as we
shall see. In general, the functions
hij (y) := hy (∂i , ∂j )
are not the constant functions δij = hy (Ei , Ej ), and we will eventually need to
determine them to second order about y = 0.
The radial vector field ∂r on B − {x} (or B(0, r0 ) \ {0}) is defined by
1 Xn 1 1 ξy0 (1)
(∂r )y := y j (∂j )y = d
dt (y + ty) t=0
= d
dt (1 + t) y t=0
= .
r j=1 r r r
More generally, for t > 0, we have
1 0
(∂r )ξy(t) = ξ (t) ,
|y| y
since for |y| > 0,
1 Xn j
Xn ξy (t)j
(∂r )ξy(t) = ξy (t) (∂j )ξy(t) = (∂j )ξy(t)
|ξy (t)| j=1 j=1 |ξy (t)|
Xn tyj 1 Xn 1 d 1 0
= (∂j )ξy(t) = y j (∂j )ξy(t) = (ty) = ξ (t) .
j=1 |ty| |y| j=1 |y| dt |y| y
554 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

If σ : U → U (E) ×f P is a radial gauge for the connection ε ⊕ θ, e then


 
σ ∗ ε ⊕ θe (∂r ) = 0, since for y 6= 0,
        1 
∗ 0
σ ε ⊕ θ (∂r )y = ε ⊕ θ (σ∗ ∂r ) = ε ⊕ θ σ∗
e e e ξ (1)
r y
1   1  
ε ⊕ θe σ∗ ξy0 (1) = ε ⊕ θe ηy0 (1) = 0.

(17.57) =
r r
This fact simplifies covariant differentiation of twisted spinor fields in the radial
direction.

About the Geometry of the Ball. We continue with local simplifications.


We begin with a pair of basic results that will also be very useful.

Lemma 17.39 (Gauss’s Lemma). Let t ∈ R and let u, v ∈ Rn be nonzero


orthogonal vectors relative
Pn to the standard
Pn inner product. Then at any point tu ∈
B(0, r0 ), the vectors i=1 ui ∂i and j=1 v j ∂j are orthogonal relative to h. In other
words, for tu ∈ B(0, r0 ),
Xn Xn Xn
δij ui v j = hij (0) ui v j = 0 =⇒ hij (tu) ui v j = 0.
i,j=1 i,j=1 i,j=1

Proof. We may assume that |u| = |v| = 1. Define V : R × [0, r0 ) → B(0, r0 )


by
V (s, t) = t(u cos s + v sin s) .
Then for any s, the curve t 7→ V (s, t) is a unit speed geodesic. Indeed: Let ∂t V =
V∗ (∂t ) and ∂s V = V∗ (∂s ). These are well-defined vector fields on span{u, v} ∩
B(0, r0 ) \ {0} with [∂t V, ∂s V ] = V∗ ([∂t , ∂s ]) = 0. Since t 7→ V (s, t) is a geodesic for
any fixed s, we have ∇∂t V ∂t V = 0. Thus, as the Levi-Civita connection is a metric
connection,
∂t V [h(∂t V, ∂t V )] = 2h(∇∂t V ∂t V, ∂t V ) = 0, so that
hV(s,t) (∂t V, ∂t V ) = hV(s,0) (∂t V, ∂t V ) = 1.
Since the Levi-Civita connection is also torsion-free, we then have
∂t V [h(∂t V, ∂s V )] = h(∇∂t V ∂t V, ∂s V ) + h(∂t V, ∇∂t V ∂s V )
= h(∂t V, ∇∂t V ∂s V ) = h(∂t V, ∇∂s V ∂t V + [∂t V, ∂s V ])
= h(∂t V, ∇∂s V ∂t V ) = 21 ∂s V [h(∂t V, ∂t V )] = 12 ∂s V [1] = 0.
Pn i
Writing i=1 (cos s) u ∂i +(sin s) v i ∂i sin s simply as u cos s + v sin s, we have
d

0 = ∂t V [h(∂t V, ∂s V )] = dt hV(s,t) (u cos s
+ v sin s) , t(−u sin s + v cos s)

⇒ hV(s,t) (u cos s + v sin s) , t(−u sin s + v cos s)
= hV(s,0) (u cos s + v sin s, 0) = 0.

In particular, for all t ∈ [0, r0 ), we have 0 = hV(0,t) (u, tv) = thut (u, v), and so
hut (u, v) = 0, as required. 
17.4. THE ASYMPTOTIC FORMULA FOR THE HEAT KERNEL 555


Proposition 17.40. If y = y 1 , . . . , y n is a normal coordinate system on
B(0, r0 ) about x ∈ M relative to the Riemannian metric h, ∂1 , . . . , ∂n are the
coordinate vector fields, and hij (y) := hy (∂i , ∂j ) for y ∈ B(0, r0 ), then
Xn
(17.58) hij (y) y j = y i .
j=1

If Rikjl (0) denotes the component R(∂i , ∂k , ∂j , ∂l ) of the Riemann curvature tensor
for the Levi-Civita connection at y = 0 (define the components Rij (0) correspond-
ingly like in (15.75), p. 431) then
Xn 3
hij (y) = δij − 31 Rikjl (0) y k y l + O |y| ,
k,l=1
Xn 3
hij (y) = δij + 13 Rikjl (0) y k y l + O |y| ,
k,l=1
α
Xn 3
α
(17.59) h := (det h) = 1 − 13 α Rkl (0) y k y l + O |y| .
k,l=1

Proof. Since t 7→ ty is a geodesic with speed |y|, we have


2
Xn  Xn
|y| = hty y i ∂i , y j ∂j = hij (ty) y i y j for all t.
i,j=1 i,j=1

By Gauss’s Lemma, we have


n
X
hij (y) y i v j = 0 for any v ∈ Rn with hy, vi = 0.
i,j=1

We may write any w ∈ Rn as w = αy + v for some α ∈ R and such a v. Then


n
X n
X n
X 2
hij (y) y i wj = α hij (y) y i y j + hij (y) y i v j = α |y|
i,j=1 i,j=1 i,j=1
Xn
= hy, αy + vi = hy, wi = y j wj ,
j=1

from which (17.58) is immediate.


To obtain (17.59), we execute a long series of very simple operations. First we
differentiate (17.58):
Xn 
δik = ∂k y i = ∂k hij (y) y j

j=1
Xn
∂k (hij (y)) y j + hij (y) ∂k y j ,

=
j=1

Using the convenient notation hij,k (y) = ∂k (hij (y)), we then have
Xn
δik = hik (y) + hij,k (y) y j .
j=1

Setting y = 0, we obtain hik (0) = δik . Differentiating again,


 Xn 
0 = ∂l hik (y) + hij,k (y) y j
j=1
Xn
hij,kl (y) y j + hij,k (y) δlj

= hik,l (y) +
j=1
Xn
(17.60) = hik,l (y) + hil,k (y) + hij,kl (y) y j ,
j=1
556 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

and setting y = 0, we obtain


hil,k (0) + hik,l (0) = 0.
This implies hil,k (0) = 0, since
hil,k (0) = hli,k (0) = −hlk,i (0) = −hkl,i (0)
= hki,l (0) = hik,l (0) = −hil,k (0) .
Thus, by simple Taylor approximation
Xn 3
hij (y) = δij + 21 hij,kl (0) y k y l + O |y| .
k,l=1

Note that
Xn
δip = hij (y) hjp (y)
j=1
Xn  Xn 3

= δij + 21 hij,kl (0) y k y l + O |y| hjp (y)
j=1 k,l=1
Xn 3
= hip (y) + 21 hij,kl (0) y k y l hjp (y) + O |y|
j,k,l=1
Xn 3
= hip (y) + 21 hij,kl (0) hjp (0) y k y l + O |y|
j,k,l=1
Xn 3
ip ip
(17.61) =⇒ h (y) = δ − 2 1
hip,kl (0) y k y l + O |y| .
k,l=1

Differentiation of (17.60) yields


 Xn 
0 = ∂p hij,kl (y) y j + hil,k (y) + hik,l (y)
j=1
Xn
= hij,klp (y) y j + hip,kl (y) + hil,kp (y) + hik,lp (y) , and so
j=1
(17.62) hip,kl + hil,pk + hik,lp = 0 at y = 0.
This implies that hij,kl (0) = hkl,ij (0), since at y = 0,
2hij,kl = hij,kl + hji,kl = −(hik,lj + hil,jk ) −(hjk,li + hjl,ik )
= −(hki,lj + hkj,il ) −(hli,jk + hlj,ki )
(17.63) = hkl,ji + hlk,ij = 2hkl,ij .
We will use this to show that
Xn Xn
k l
1
2 hij,kl (0) y y = − 1
3 Rikjl (0) y k y l .
k,l=1 k,l=1

We also need to write Rikjl (0) in terms of the derivatives of the hij at 0. If
ω ∈ Ω1 (B, GL(n)) denotes the pull-back of the Levi-Civita connection 1-form on
LM via the y-coordinate frame field, then
Rikjl = (dω + ω ∧ ω)ik (∂j , ∂l ) .
Here LM denotes the bundle of linear frames for M , introduced in Section 15.5,
p.414. With the notations and summation conventions of (15.47) (p. 421) we obtain
l
(17.64) ω(∂i )j = Γlij := 21 hlk (∂i [hjk ] + ∂j [hik ] − ∂k [hij ]) ,
which is 0 at y = 0, since we have shown hil,k (0) = 0. Thus, ω ∧ ω = 0 at y = 0
and h i h i
i i
Rikjl (0) = Rikjl (0) = dωki (∂j , ∂l ) = ∂j ω(∂l )k − ∂l ω(∂j )k .

17.4. THE ASYMPTOTIC FORMULA FOR THE HEAT KERNEL 557

2
Since hlk = δ lk + O |y| by (17.61), (17.64) yields
l 2
ω(∂i )j = 21 (hjl,i + hil,j − hij,l ) + O |y| .
Thus (where henceforth all terms are evaluated at y = 0),
h i h i
i i
Rikjl = ∂j ω(∂l )k − ∂l ω(∂j )k
= ∂j 12 (hki,l + hli,k − hlk,i ) − ∂l 21 (hki,j + hji,k − hjk,i )
   

= 12 (hki,lj + hli,kj − hlk,ij ) − 21 (hki,jl + hji,kl − hjk,il )


= 21 (hli,kj − hlk,ij ) − 21 (hji,kl − hjk,il )
= 12 (hli,kj − hji,kl + hjk,il − hlk,ij )
= hli,kj − hji,kl .
using (17.63). We show
hij,lk = − 31 (Rikjl + Riljk ) .
Indeed,
Rikjl + Riljk = (hli,kj − hji,kl ) +(hki,lj − hji,lk )
= −2hji,kl + hli,kj + hki,lj = −2hji,kl + hil,jk + hik,lj
= −2hji,kl − hij,kl = −3hji,kl .
Hence,
Xn Xn
1
2 hij,kl y k y l = − 21 31 (Rikjl + Riljk ) y k y l
k,l=1 k,l=1
Xn
= − 31 Rikjl y k y l .
k,l=1
Now
 
3 
det h = det δij − 31 Rikjl (0) y k y l + O |y|
3
= det(δij ) − Tr 31 Rikjl (0) y k y l + O |y|


3
= 1 − 13 Rikil (0) y k y l + O |y| and so
α 3
(det h) = 1 − 13 αRikil (0) y k y l + O |y| . 
1
The Levi-Civita connection θ ∈ Ω (F M, so(n)) is a form on F M and it is the
restriction of the linear connection ω ∈ Ω1 (LM, gl(n)) on LM . For
u := C ◦ σ2 : B → F M,
the framing E1 := u(e1 ) , . . . , En := u(en ) , is a local section of F M , while the
coordinate framing ∂1 , . . . , ∂n is a local section, say I : B → LM . We have the
pull-backs u∗ θ ∈ Ω1 (B, so(n)) and I ∗ ω ∈ Ω1 (B, gl(n)). Moreover, if ∇θ denotes the
covariant derivative operator for the Levi-Civita connection, we have
Xn k
∇ω∂i ∂j = ((I ∗ ω)(∂i )) j ∂k and
k=1
Xn k
(17.65) ∇θEi Ej = ((u∗ θ)(Ei )) j Ek .
k=1
We verify the second of the equations (17.65). The proof of the first equation is
completely analogous, and we have shown that
Xn
∇ω∂i ∂j = Γkij ∂j ,
k=1
558 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

where the Γkij are the Christoffel symbols given by (15.47), namely
k
((I ∗ ω)(∂i )) j = Γkij = 12 hlk (∂i [hjk ] + ∂j [hik ] − ∂k [hij ]) .
By definition (see (15.34)), for vector fields X and Y ,
∇θX Y y := u(y) Dθ (ϕ(Ye ))u(y) (X)
 
e .
Using
 
u∗ ϕ(E ej ) = u(y)−1 π∗ (E
ej ) = ϕu(y) (E ej )
y
−1 −1
= u(y) π∗ (u∗y (Ej )) = u(y) (Ej ) = ej ,
we then have
∇θEi Ej = u(y) Dθ (ϕ(E
 
ej ))u(y) (E
ei )
y

= u(y) Dθ (ϕ(E

ej ))u(y) (u∗y (Ei ))

= u(y) d(ϕ(E ej ))u(y) (u∗y (Ei )) + θu(y) (u∗y (Ei )) ϕu(y) (E
ej )
   
= u(y) u∗ d(ϕ(E ej ))y (Ei ) + u(y) (u∗ θ) (Ei ) ej
y
   
∗ ∗
= u(y) d(u (ϕ(E ej )))y (Ei ) + u(y) (u θ) (Ei ) ej
y

= u(y) d(ej )y (Ei ) + u(y) (u∗ θ)y (Ei ) ej


 
 
= u(y) (u∗ θ)y (Ei ) ej
n
X k 
= u(y) (u∗ θ)y (Ei ) j ek
k=1
Xn k
= (u∗ θ)(Ei ) j
(Ek )y .
k=1
Proposition 17.41. We use the notation of Proposition 17.40. Let I ∗ ω ∈
1
Ω (B, GL(n)) denote the pull-back of the Levi-Civita connection 1-form on LM via
the y-coordinate frame field ∂1 , . . . , ∂n , and let ∇ denote the covariant derivative
for the Levi-Civita connection so that
Xn n
X
k
∇∂i ∂j = ((I ∗ ω)(∂i )) j ∂k = Γkij (y) ∂k .
k=1
k=1
k
Then the Christoffel symbols := Γkij (y)
obey ω(∂i ) j
Xn 2
(17.66) Γkij (y) = 31 (Ripjk (0) + Rikjp (0)) y p + O |y| .
p=1

Moreover, if Ej (y) := u(ej ) := (C ◦ σ2 )y (ej ) (j = 1, . . . n) is the radial framing


(which equals ∂1 , . . . , ∂n at y = 0 and is parallel along the geodesics through y = 0),
then
Xn  Xn 
3
(17.67) Ej (y) = δjk + 16 Riklj (0) y i y l ∂k + O |y| ,
k=1 i,l=1
and
Xn  k
∇Ei Ej = (u∗ θ)y (Ei ) Ek
k=1 j
Xn 2
1 l
(17.68) = 2 Rkjli (0) y Ek + O |y| .
k,l=1
17.4. THE ASYMPTOTIC FORMULA FOR THE HEAT KERNEL 559

Proof. Using the summation convention and (17.62), we have

Γkij = 12 hlk (∂i [hjl ] + ∂j [hil ] − ∂l [hij ])(y)


= 12 hlk (hjl,i + hil,j − hij,l )(y)
2
= 21 δ lk (hjl,ip + hil,jp − hij,lp ) y p + O |y|
2
= 12 (hjk,ip + hik,jp − hij,kp ) y p + O |y|
2
= 21 (−hkp,ij − hij,kp ) y p + O |y|
2 2
= −hkp,ij y p + O |y| = −hij,kp y p + O |y|
2
= 31 (Ripjk (0) + Rikjp (0)) y p + O |y| ,

and so (17.66) holds.


Since Ej = ∂j at y = 0, there are functions bkjl , such that

Ej (y) = δjk + bkjl (y) y l ∂k .




As the Ej are parallel in the ∂r -direction, for all y ∈ B(0, r0 ) we have

0 = r∇∂r Ej (y) = r∇∂r ∂j + bkjl (y) y l ∂k




= ∇yi ∂i ∂j + bkjl (y) y l ∂k




= y i ∇∂i ∂j + bkjl (y) y l ∂k




= y i ∇∂i ∂j + ∇∂i bkjl (y) y l ∂k




= Γkij (y) y i ∂k + bkjl (y) y i y l ∇∂i ∂k + ∂i bkjl (y) y l y i ∂k




= Γkij (y) y i ∂k + bkjl (y) y i y l Γpik (y) ∂p + ∂i bkjl (y) y i y l + bkji (y) y i ∂k
 
 
= bkji (y) y i + Γkij (y) y i + ∂i bkjl (y) y i y l + bqjl (y) Γkiq (y) y i y l ∂k .


 2
Since Γkij (y) = O |y| , all terms except possibly bkji (y) y i are O |y| . As the entire
expression vanishes to all orders, we must have bkji (0) = 0, and so
2
bkji (y) = bkji,l (0) y l + O |y| .

Considering the second-order part and using (17.66), we get


k k
1
yi yl

0 = 3 (Riljk (0) + Rikjl (0)) + bji,l (0) + bjl,i (0)
− 31 Riklj (0) + bkji,l (0) + bkjl,i (0) y i y l , and

= so
bkji,l (0) + bkjl,i (0) = 1
6 (Riklj (0) + Rlkij (0)) .

Hence,
 
3
Ej (y) = δjk + bkjl (y) y l ∂k = δjk + bkjl,i (0) y i y l + O |y|

∂k
 
3
= δjk + 21 bkjl,i (0) + bkji,l (0) y i y l + O |y|

∂k
 
3
= δjk + 16 Riklj y i y l + O |y|

∂k ,
560 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

2
and we have (17.67). As for (17.68), we have, modulo O |y| ,
∇Ei Ej = ∇∂i Ej
 Xn 
= ∇∂i ∂j + ∇∂i 16 Rpklj (0) y p y l ∂k
k,p,l=1
 Xn 
= ∇∂i ∂j + ∂i 6 1
Rpklj (0) y p y l ∂k
k,p,l=1
Xn Xn
= ∇∂i ∂j + 6 1
Riklj (0) y l ∂k + 16 Rpkij (0) y p ∂k
k,l=1 k,p=1
Xn
= ∇∂i ∂j + 61 Riklj (0) + Rlkij (0) y l ∂k

k,l=1
Xn Xn
Γkij (y) ∂k + 16 Riklj (0) + Rlkij (0) y l ∂k

=
k=1 k,l=1
Xn Xn
1
Riljk (0) + Rikjl (0) y l ∂k

= 3
k=1 l=1
Xn
+ 16 Riklj (0) + Rlkij (0) y l ∂k

k,l=1
Xn   l
1
 1
= R iljk (0) + Rikjl (0) + R iklj (0) + R lkij (0) y ∂k
k,l=1 3 6
Xn
1
Riljk (0) + 16 Rikjl (0) + 16 Rlkij (0) y l ∂k

=
k,l=1 3
Xn
1
Rkjli (0) − 16 Rkijl (0) − 16 Rklij (0) y l ∂k

= 3
k,l=1
Xn   l
1 1
= R kjli (0) − R kjli (0) + R kijl (0) + R klij (0) y ∂k
k,l=1 2 6
Xn
= 12 Rkjli (0) y l Ek . 
k,l=1

A section ψ ∈ C ∞ (E ⊗ Σ(M )) can be identified with an equivariant function


0 
ψe ∈ Ω U (E) ×f P, CN ⊗ Σ2m . The local section σ : B → U (E) ×f P allows us
form the pullback σ ∗ ψe = ψe ◦σ ∈ C ∞ B, CN ⊗ Σ2m and pulling back once more via


φ−1 : B(0, r0 ) → B, we have ψe ◦ σ ◦ φ−1 ∈ C ∞ B(0, r0 ) , CN ⊗ Σ2m . This allows




us to locally treat the twisted spinor field ψ ∈ C ∞ (E ⊗ Σ(M )) as a function on


B(0, r0 ) with values in the single vector space CN ⊗ Σ2m . In particular, expressions
such as ∂i ψ, which are encountered in the following results then make sense. For
y ∈ B(0, r0 ), let h(y) denote the matrix whose entries are hij (y) = hy (∂i , ∂j ). It is
customary and convenient to use the abusive notation
p p q
h(y) := det h(y) = det[hy (∂i , ∂j )],

so that h dy 1 ∧ · · · ∧ dy n is the volume element for h on B(0, r0 ). To save vertical
space, we set
−1 1
h /2 := √ .
h
Moreover, we adhere to the summation convention whereby sums are automatically
taken over repeated indices if they are not summed explicitly. For better reading,
these dummy indices will typically be set in mutually opposite position, i.e., as
subscript and superscript.
Recall that in terms of a local orthonormal framing E1 , . . . , En we have
X X
Dψ = ϕj · ∇Ej ψ = Ej · ∇Ej ψ.
j j
17.4. THE ASYMPTOTIC FORMULA FOR THE HEAT KERNEL 561

This is independent of the local orthonormal framing. However, if we take


E1 , . . . , En to be the radial framing and identify ψ and Dψ with ψe ◦ σ ◦ φ−1 and
g ◦ σ ◦ φ−1 ∈ C ∞ B(0, r0 ) , CN ⊗ Σ2m , then

X  
g ◦ σ ◦ φ−1 = γ j d ψe ◦ σ ◦ φ−1 (Ej ) + σ ∗ ε ⊕ θe (Ej ) ψe ◦ σ ◦ φ−1 .
 

j

−1
Of course it would be cumbersome to maintain the notation  ψ ◦ σ ◦ φ , and so we
e
−1 ∞ N
will simply denote ψe ◦ σ ◦ φ ∈ C B(0, r0 ) , C ⊗ Σ2m by ψ, in which case
X
γ j Ej [ψ] + γ j σ ∗ ε ⊕ θe (Ej ) ψ

Dψ =
j
X
γ j Ej [ψ] + γ j (σ ∗ ε)(Ej ) ψ + γ j σ ∗ θe (Ej ) ψ.

=
j
PN P2m
Note that ψ can be written locally as a sum ψ(y) = p=1 q=1 ap,q (y) vp ⊗ φq ,
where ap,q ∈ C ∞ (B(0, r0 ) , C) , vp ∈ CN and φq ∈ Σ2m . Also, (σ ∗ ε)y (∂j ) ∈
End CN acts only on vp in vp ⊗ φq , while σ ∗ θe y (∂j ) ∈ End(Σ2m ), as well as
 

γ j , act only on φq in vp ⊗ φq .
At y = 0, Ej = ∂j and σ ∗ ε(Ej ) ⊕ θ(E

e j ) = 0. Hence we simply have
X
(Dψ)(0) = γ j (∂j ψ)(0) .
j

We will need the lead-order terms of (σ ∗ ε)y (Ej ) and σ ∗ θe y (Ej ) as functions of y.


Since (σ ∗ ε)0 (∂j ) = 0 in the radial gauge, we have


Xn 2
Xn  2

(σ ∗ ε)y (∂j ) = εij y i + O |y| or (σ ∗ ε)y = εij y i + O |y| dy j
i=1 i=1

for some εij ∈ End CN . At y = 0, the pull-back of the curvature is
(σ ∗ Ωε )0 = d(σ ∗ ε)0 +(σ ∗ ε)0 ∧(σ ∗ ε)0 = d(σ ∗ ε)0
Xn  Xn
= d εij y i dy j = εij dy i ∧ dy j .
i=1 i,j=1
∗ ε
Denoting the pull-back σ Ω by F , we then have
εij = Fij := (σ ∗ Ωε )0 (∂i , ∂j ) .
2
Since Ej = ∂j + O |y| by (17.67),
Xn 2
(17.69) (σ ∗ ε)y (Ej ) = (σ ∗ ε)y (∂j ) = Fij y i + O |y| .
i=1
Using (17.4) and (17.68), we have
σ ∗ θe y (Ej ) = σ2∗ c−1 C ∗ θ y (Ej ) = c−1 ((σ2∗ ◦ C ∗ ) θ)y (Ej )
 

∗ 
= c−1 (C ◦ σ2 ) θ y (Ej ) = c−1 (u∗ θ)y (Ej )
Xn  
1 k l ∗
=− γ γ (u θ) y (E j )
k,l=1 4 kl
Xn Xn
1 k l1
=− γ γ 2 Rklij (0) y i
k,l=1 4 i=1
Xn
(17.70) = − 81 γ k γ l Rklij (0) y i .
i,k,l=1
562 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

Proposition 17.42. For ψ ∈ C ∞ B(0, r0 ) , CN ⊗ Σ2m , we have




 √ 
∆ψ = h− /2 ∇∂j hij h ∇∂i ψ .
1

Proof. Using (15.49), p. 421, we have

∆ψ = hij ∇2∂i ,∂j ψ = hij ∇∂j (∇∂i ψ) − hij ∇ψ ∇∂j ∂i




= hij ∇∂j (∇∂i ψ) − ∇ψ hij ∇∂j ∂i



  √  
= hij ∂j (∇∂i ψ) + hij ε ⊕ θe (∂j )(∇∂i ψ) + ∇ψ h− /2 ∂j hij h ∂i
1

 √ 
= hij ∂j (∇∂i ψ) + hij ε ⊕ θe (∂j )(∇∂i ψ) + h− /2 ∂j hij h ∇∂i ψ
1

 √  √  
= h− /2 hij h ∂j (∇∂i ψ) + ∂j hij h ∇∂i ψ
1


+ h− /2 hij h ε ⊕ θe (∂j )(∇∂i ψ)
1

 √  √
= h− /2 ∂j hij h ∇∂i ψ + h− /2 hij h ε ⊕ θe (∂j )(∇∂i ψ)
1 1

 √ 
= h− /2 ∇∂j hij h ∇∂i ψ .
1


We will generalize this a bit (see Proposition 17.44 below), but first we need

Proposition 17.43. Let B(0, r0 ) ⊂ Rn as before and let F : B(0, r0 ) → R be


of the form F (y) = f (kyk) = f (r). Then

 
n−1
∆F := hij ∇2∂i ,∂j (F ) = f 00 (r) + h− /2 ∂r [ h] + f 0 (r) .
1
(17.71)
r

Proof. By the same computation as in the proof of


 Proposition 17.42, we
√  √
have ∆F = h−1/2 ∇∂j hij h ∇∂i F = h−1/2 ∂j hij h ∂i F . Then

√ √ 0 yi
   
−1/2 ij −1/2 ij
h ∂j h h ∂i F
=h ∂j h h f (r)
r
√ f 0 (r)
 
−1/2 j
=h ∂j y h (using (17.40))
r
√ f 0 (r)  √ f 0 (r)
   
−1/2 j −1/2 j
=h y ∂j h +h ∂j y h
r r

 0
f (r) nf 0 (r)
0

f (r)  
= y j ∂j + h− /2 y j ∂j
1
h +
r r r

 00 0
f 0 (r) nf 0 (r)

rf (r) − f (r) yj −1/2 j
 
= yj + h y ∂ j h +
r2 r r r
0
f (r)  √  0
nf (r)
= f 00 (r) − + h− /2 ∂r h f 0 (r) +
1

r r

 
n − 1
= f 00 (r) + h− /2 ∂r [ h] + f 0 (r) .
1

r
17.4. THE ASYMPTOTIC FORMULA FOR THE HEAT KERNEL 563

Proposition 17.44. Let F : B (0, r0 ) → R be of the form F (y) = f (kyk) =


f (r). Then for ψ ∈ C ∞ B (0, r0 ) , CN ⊗ Σ2m as above, we have, where r =
|y| , y ∈ B (0, r0 ),
∆ (F ψ) (y) = (∆F ) ψ + 2hij (∇∂i F ) ∇∂j ψ + F ∆ψ

 
n−1
= f 00 (r) ψ (y) + h− /2 ∂r [ h] + f 0 (r) ψ (y)
1

r
(17.72) + 2f 0 (r) ∂r ψ + f (r) ∆ψ.
Proof. By repeated use of the product rule,
 √ 
∆ (F ψ) = h− /2 ∇∂j hij h ∇∂i (F ψ)
1

 √ √ 
= h− /2 ∇∂j hij h (∇∂i F ) ψ + hij h F ∇∂i ψ
1

= (∆F ) ψ + 2hij (∇∂i F ) ∇∂j ψ + F ∆ψ.


In view of Proposition 17.43, it remains to compute
hij (y) (∇∂i F ) ∇∂j ψ = hij (y) (∂r f ∂i r) ∇∂j ψ = hij (y) ∂r f y i /r ∇∂j ψ


∂r f
= hij (y) y i ∇∂j ψ = ∂r (f ) ∇ yj ∂ ψ = ∂r f ∇∂r ψ,
r r j

where we have used Proposition 17.40. Note that (17.58) says that y is an eigen-
vector of the matrix [hij (y)], and hence y is also eigenvector of the inverse matrix
[hij (y)]. Finally note that ∇∂r ψ = ∂r ψ + σ ∗ ε ⊕ θe (∂r ) ψ = ∂r ψ in the radial gauge


by (17.57). 

Let ∆e= ∂12 +· · ·+∂n2 denote the usual Laplace operator in Rn with coordinates
y 1 , . . . , y n and ∂i := ∂/∂y i . Setting x = 0 in (17.51), p. 550, the fundamental
solution of ut = ∆e u in Euclidean n-space is
−n/2
exp − 14 r2 /t , for t > 0.

(17.73) E (r, t) := (4πt)
2 2
where r2 = y 1 + . . . + (y n ) . Using Proposition 17.43,
n−1
(17.74) ∂t E = ∆e E = ∂r2 E + ∂r E.
r
If ∆ denotes the Laplace operator on B for the metric h, then by (17.71),
√ √
 
n−1
∆E = ∂r2 E + h− /2 ∂r [ h] + ∂r E = ∆e E + h− /2 ∂r [ h]∂r E.
1 1
(17.75)
r
For 0 ≤ Q ∈ Z, let ΨQ ∈ C ∞ B × (0, ∞) , CN ⊗ Σ2m be of the form


XQ
ΨQ (y, t) := E (r, t) Uk (y) tk ,
k=0

where Uk ∈ C ∞ B, CN ⊗ Σ2m . If U0 (0) ∈ CN ⊗ Σ2m is arbitrarily specified, we




seek a formula for Uk (y), k = 0, . . . , Q, such that


D2 + ∂t ΨQ (y, t) = E(r, t) tQ D2 (UQ )(y) ,

(17.76)
where the square of the Dirac operator D is given by
X
D2 ψ = −∆ψ + 21 Ωεjk Ej · Ek · ψ + 14 Sψ,
j,k
564 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

as computed in Proposition 17.27,


 p. 537. It is convenient to define the 0-th order
operator F on C ∞ B, CN ⊗ Σ2m via
X
F [ψ] := Ωεjk Ej · Ek · ψ, so that
1
2 j,k

D2 = −∆ψ + F + 14 S [ψ] .


The desired formula for the Uk (y) involves the operator A on C ∞ B, CN ⊗ Σ2m


given (where h1/4 := ( h)1/2 and h−1/4 := (h−1/2 )1/2 ) by
h i h i
A [ψ] := −h1/4 D2 h−1/4 ψ = h1/4 ∆ h−1/4 ψ − F + 41 S [ψ] .


For s ∈ [0, 1], let


As [ψ](y) := A [ψ](sy) .

Proposition 17.45. Let U0 (0) ∈ CN ⊗ Σ2m , and let V0 ∈ C ∞ B, CN ⊗ Σ2m




denote the constant function V0 (y) ≡ U0 (0). Then the Uk (y) which satisfy (17.76)
are given by
−1/4
Uk (y) = h(y) Vk (y) , where
Z Y
k−1 i 
(17.77) Vk (y) = (si ) Ask−1 ◦ · · · ◦ As0 [V0 ] (y) ds0 . . . dsk−1 ,
Ik i=0

and where I k = {(s0 , . . . sk ) : si ∈ [0, 1] , i ∈ {0, . . . , k − 1} } .


PQ
Proof. Let Σ(y, t) := k=0 Uk (y) tk . Using (17.72) and (17.75), we compute

∂t + D2 [ΨQ (y, t)] = ∂t − ∆ + F + 41 S [E(r, t) Σ(y, t)]


 

= ∂t [E(r, t)] Σ(y, t) + E(r, t) ∂t Σ(y, t)


− ∆ [E(r, t) Σ(y, t)] + F + 14 S E(r, t) Σ(y, t)


= ∂t [E(r, t)] Σ(y, t) + E(r, t) ∂t Σ(y, t)


− ∆(E(r, t)) Σ(y, t) − 2∂r (E(r, t)) ∂r Σ(y, t) − E(r, t) ∆(Σ(y, t))
+ F + 14 S E(r, t) Σ(y, t)

h√ i
= (∂t − ∆e ) [E(r, t)] Σ(y, t) − h− /2 ∂r
1
h ∂r [E(r, t)] Σ(y, t)
− 2∂r (E(r, t)) ∂r Σ(y, t) + E(r, t)(∂t Σ(y, t) − ∆Σ(y, t))
+ F + 14 S E(r, t) Σ(y, t)

 h√ i 
= −∂r [E(r, t)] h− /2 ∂r
1
h Σ(y, t) + 2∂r Σ(y, t)
+ E(r, t) ∂t Σ(y, t) − ∆Σ(y, t) + F + 14 S Σ(y, t) .
 

Since
−n/2
∂r exp − 14 r2 /t
 
∂r [E(r, t)] = (4πt)
−n/2
exp − 14 r2 /t ∂r − 14 r2 /t = − 2t
r
  
= (4πt) E(r, t) ,
17.4. THE ASYMPTOTIC FORMULA FOR THE HEAT KERNEL 565

we have
−1
∂t + D2 [ΨQ (y, t)]

E(r, t)

r −1/2
= 2t h ∂r ( h)Σ(y, t) + 2∂r Σ(y, t)
+ ∂t Σ(y, t) − ∆Σ(y, t) + F + 14 S Σ(y, t)


Q  r −1/2

∂r ( h)Uk (y) + rt ∂r Uk (y)

X h
= 2t tk
+ kt Uk (y) − ∆Uk (y) + F + 41 S Uk (y)
k=0
Q
 √  !
X k + 2r h−1/2 ∂r ( h) Uk (y) + r∂r Uk (y)
= tk−1
1

k=0 + −∆ + F + 4 S [Uk−1 (y)]
+ −∆ + F + 14 S [UQ (y)] tQ ,


where U−1 (y) := 0. Thus, to achieve (17.76), we need


 r 1 √ 
r∂r Uk (y) + k + h− /2 ∂r ( h) Uk (y) = ∆ − F − 41 S [Uk−1 (y)] .

2
For fixed 0 6= y0 ∈ B, we define
uk (s) := Uk (sy0 ) for s ∈ [0, 1] .
At sy0 , we have r = s ky0 k and so
∂ 1
r∂r = s ky0 k ∂s = s ky0 k ∂s = s∂s .
∂r ky0 k
Thus, we obtain the first-order, linear ordinary differential equation
 i
duk s −1/2 d hp
h(sy0 ) uk (s) = ∆ − F − 14 S [uk−1 (sy0 )] .

s + k+ h
ds 2 ds
The integrating factor is
Z  i
k 1 −1/2 d hp 1/4
exp + 2h h(sy0 ) ds = sk h(sy0 ) .
s ds
Hence, for k > 0, we obtain
d k 1/4

1/4
s h(sy0 ) uk (s) = sk−1 h(sy0 ) ∆ − F − 41 S [uk−1 (sy0 )] and

ds Z s
1/4 1/4
sk h(sy0 ) uk (s) = sek−1 h(e ∆ − F − 14 S [uk−1 (e

sy0 ) sy0 )] de
s,
0

where we note that both sides are 0 for s = 0 when k > 0. Setting s = 1, replacing
y0 by y, we have Uk (y) = uk (1) and (for k > 0),
Z 1
−1/4 1/4
sk−1 h(sy) ∆ − F − 14 S [Uk−1 ](sy) ds.

Uk (y) = h(y)
0

When k = 0,
d 1/4

−1/4
h(sy0 ) u0 (s) = 0 =⇒ u0 (s) = h(sy0 ) u0 (0)
ds
−1/4
=⇒ U0 (y) = h(y) U0 (0) .
566 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

−1/4
Since Uk (y) = h(y) Vk (y) , we have V0 (y) = U0 (0) and
Z 1
1/4 h i
Vk (y) = sk−1 h(sy) ∆ − F − 14 S h−1/4 Vk−1 (sy) ds
0
Z 1 Z 1
k−1
= sk−1 As [Vk−1 ](y) ds = (sk−1 ) Ask−1 [Vk−1 ](y) dsk−1
0 0
Z 1 Z 1 
k−1 k−2
= (sk−1 ) Ask−1 (sk−2 ) Ask−2 [Vk−2 ] dsk−2 (y) dsk−1
0 0
Z 1 Z 1
k−1
k−2 
= (sk−1 )
(sk−2 ) Ask−1 ◦ Ask−2 [Vk−2 ](y) dsk−2 dsk−1
Z0 0Y 
k−1 i 
= (si ) Ask−1 ◦ Ask−2 ◦ · · · ◦ As0 [V0 ] (y) ds0 . . . dsk−1 ,
Ik i=0

where we have used the fact that the As are linear differential operators and all
functions are C ∞ in order to bring the Asj inside the si -integrals for i 6= j. 
Note that U0 (0) ∈ CN ⊗ Σ2m may be arbitrarily specified, and once U0 (0) is
 are uniquely determined via (17.77). Next we define hk (y) ∈
chosen, the Um (y)
End CN ⊗ Σ2m by
(17.78) hk (y)(U0 (0)) := Uk (y)

(in particular, h0 (0) = Id ∈ End CN ⊗ Σ2m ), and
Q
X
hk (y) tk ∈ C ∞ B, End CN ⊗ Σ2m

(17.79) HQ (0, y, t) := E(r, t) .
k=0

We may regard the preceding expression for HQ (0, y, t) as yielding


 
HQ (x, y, t) ∈ Hom Ex ⊗ Σ(M )x , Ey ⊗ Σ(M )y ,
where x ∈ M is the point about which we have chosen normal coordinates. For y
sufficiently close to x, we set
Q
X
(17.80) HQ (x, y, t) := E(d(x, y), t) hk (x, y) tk .
k=0

Further Approximations. To obtain a globally defined version of HQ we


proceed as follows. For r > 0, let
δr (M × M ) := {(x, y) ∈ M × M : d(x, y) < r} .
For sufficiently small r, say 0 < r < rM = injectivity radius of M , the collection
{HQ (x, y, t)} yields a section of the bundle H|δr(M ×M )×[0,∞) , where (as in Section
17.3)

H := π1∗ (E ⊗ Σ(M )) ⊗(π2∗ (E ⊗ Σ(M )))
(17.81) ∼ K := π ∗ (E ⊗ Σ(M )) ⊗ π ∗ (E ⊗ Σ(M )) ,
= 1 2

and πi : M × M × [0, ∞) → M (i = 1, 2) are the projections. For some constants


r1 , r2 with 0 < r1 < r2 < rM , there is a C ∞ function ρ : [0, ∞) → [0, 1] with
ρ|[0,r1 ] = 1 and ρ|[r2 ,∞) = 0. Then we have ϕ ∈ C ∞ (M × M, R) defined by
ϕ(x, y) := ρ(d(x, y)) .
17.4. THE ASYMPTOTIC FORMULA FOR THE HEAT KERNEL 567

with ϕ|δr1(M ×M ) = 1 and ϕ = 0 on the complement of δr2 (M × M ). Then ϕHQ


extends by zero values to a section GQ ∈ C ∞ (H), namely,

ϕ(x, y) HQ (x, y, t) , for (x, y) ∈ δr2 (M × M ) , t > 0,
(17.82) GQ (x, y, t) :=
0, for (x, y) ∈
/ δr2 (M × M ) , t > 0.

From GQ (x, y, t) we will give a different construction of the heat kernel k(x, y, t)
and obtain its asymptotic expansion as t → 0+ . However, before the proof of the
validity of the construction, it is best to provide some motivation for it, as follows.
One expects that at least for small t, GQ (x, y, t) is a good approximation for
k(x, y, t), since the effect at small time t of a heat source at x should not be ap-
preciably felt at a distantpoint y. Nevertheless, unlike the true heat kernel, we do
not expect that ∂t + Dx2 GQ (x, y, t) = 0 exactly. Let

KQ,0 (x, y, t) := ∂t + Dx2 GQ (x, y, t) .




In order to begin the correction process, we consider the problem

∂t + Dx2 ψ(x, y, t) = −KQ,0 (x, y, t) , with ψ(x, y, 0) = 0.




The solution of this problem should be given approximately by


Z tZ
ψ(x, y, t) ≈ − GQ (x, z, s) KQ,0 (z, y, t − s) ν(z) ds
0 M
Z tZ
(17.83) =− GQ (x, z, t − s) KQ,0 (z, y, s) ν(z) ds.
0 M

To compactify the expressions to follow, it is convenient to introduce

Definition 17.46. For f1 , f2 ∈ C 0 (H) , the convolution f1 ∗ f2 ∈ C 0 (H) is


defined by
Z tZ
(17.84) (f1 ∗ f2 )(x, y, t) := f1 (x, z, s) ◦ f2 (z, y, t − s) νh (z) ds.
0 M

By a change of variable, we also have the alternate expression


Z tZ
(f1 ∗ f2 )(x, y, t) = f1 (x, z, s) ◦ f2 (z, y, t − s) νh (z) ds
0 M
Z 0 Z
= f1 (x, z, t − σ) ◦ f2 (z, y, σ) νh (z) (−dσ)
t M
Z tZ
(17.85) = f1 (x, z, t − s) ◦ f2 (z, y, s) νh (z) ds.
0 M

It is convenient to record here that when f1 and f2 are C 1 ,

∂xi ((f1 ∗ f2 )(x, y, t)) = (∂xi f1 ∗ f2 )(x, y, t) , and



(17.86) ∂yi ((f1 ∗ f2 )(x, y, t)) = f1 ∗ ∂yi f2 (x, y, t) .
568 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

Moreover using (17.84) and (17.85),


Z
∂t ((f1 ∗ f2 )(x, y, t)) = f1 (x, z, t) ◦ f2 (z, y, 0) νh (z)
M
Z tZ
+ f1 (x, z, s) ◦ ∂t f2 (z, y, t − s) νh (z) ds, and
Z 0 M
∂t ((f1 ∗ f2 )(x, y, t)) = f1 (x, z, 0) ◦ f2 (z, y, t) νh (z)
M
Z tZ
+ ∂t f1 (x, z, t − s) ◦ f2 (z, y, s) νh (z) ds.
0 M

Thus, when f1 (·, ·, 0) ≡ 0 and f2 (·, ·, 0) ≡ 0, we have


(17.87) ∂t ((f1 ∗ f2 )(x, y, t)) = (f1 ∗ ∂t (f2 ))(x, y, t) = (∂t (f1 ) ∗ f2 )(x, y, t) .
In terms of convolution, (17.83) can be written as
ψ(x, y, t) ≈ −(GQ ∗ KQ,0 )(x, y, t) .
At least formally, we have
∂t + Dx2 ψ(x, y, t) = − ∂t + Dx2 (GQ ∗ KQ,0 )(x, y, t)
 

 t
Z Z
= − ∂t + Dx2 GQ (x, z, t − s) ◦ KQ,0 (z, y, s) v(z) ds
0 M
Z tZ
= −∂t GQ (x, z, t − s) ◦ KQ,0 (z, y, s) v(z) ds
0 M
Z tZ
− Dx2 GQ (x, z, t − s) ◦ KQ,0 (z, y, s) v(z) ds
0 M
Z
= − lim− GQ (x, z, t − s) ◦ KQ,0 (z, y, s) v(z) ds
s→t M
Z tZ
∂t + Dx2 GQ (x, z, t − s) ◦ KQ,0 (z, y, s) v(z) ds


0 M
Z tZ
= −KQ,0 (x, y, t) − KQ,0 (x, z, t − s) ◦ KQ,0 (z, y, s) v(z) ds
0 M
= −KQ,0 (x, y, t) −(KQ,0 ∗ KQ,0 )(x, y, t) .

Then one expects that, to greater accuracy than ∂t + Dx2 GQ (x, y, t) ≈ 0, we have
∂t + Dx2 (GQ − GQ ∗ KQ,0 )(x, y, t) ≈ 0.


We compute
∂t + Dx2 (GQ − GQ ∗ KQ,0 )(x, y, t)


= KQ,0 (x, y, t) + ∂t + Dx2 (ψ(x, y, t))




= KQ,0 (x, y, t) − KQ,0 (x, y, t) −(KQ,0 ∗ KQ,0 )(x, y, t)


= −(KQ,0 ∗ KQ,0 )(x, y, t) .
Then we consider the problem
∂t + Dx2 ψ(x, y, t) = (KQ,0 ∗ KQ,0 )(x, y, t) ,

17.4. THE ASYMPTOTIC FORMULA FOR THE HEAT KERNEL 569

whose solution should be given approximately by


ψ(x, y, t) ≈ (GQ − GQ ∗ KQ,0 ) ∗(KQ,0 ∗ KQ,0 )(x, y, t)
= GQ ∗(KQ,0 ∗ KQ,0 )(x, y, t) − GQ ∗ KQ,0 ∗(KQ,0 ∗ KQ,0 )(x, y, t) .
How to get forward from here? Adding this correction to GQ − GQ ∗ KQ,0 , presum-
ably gives a more accurate approximation to the fundamental solution, namely
GQ − GQ ∗ KQ,0 + GQ ∗(KQ,0 ∗ KQ,0 ) − GQ ∗ KQ,0 ∗(KQ,0 ∗ KQ,0 ) .
Repeating, we conjecture that the exact fundamental solution is

X k k factors
GQ + GQ ∗ (−1) KQ,0 ∗ ··· ∗ KQ,0 .
k=1

To prepare the rigorous demonstration of this, we first establish some boundedness


properties of the operator whose kernel is GQ .
Proposition 17.47. For ψ ∈ C l (E ⊗ Σ(M )) (l ≥ 0) and
Z
(GQ ψ)(y, t) := GQ (x, y, t) ψ(x) νh (x) ,
M

there is a constant cl independent of t, such that


(17.88) k(GQ ψ)(·, t)kC l ≤ cl kψkC l .
Also,
(17.89) lim (GQ ψ)(y, t) = ψ(y) .
t→0+

Moreover for K ∈ C l (H) (H as in (17.81)), we have


(17.90) kGQ ∗ KkC l(T ) ≤ cl,T kKkC l(T ) ,

where C l (T ) denotes the C l norm for restrictions of sections in C l (H) to M × M ×


[0, T ], where T may be chosen arbitrarily large, and cl,T is a constant depending on
l and T .
√ √
Proof. Making the change of variable w = (x − y) / t, x = y + w t,
Z
(GQ ψ)(y, t) = ϕ(x, y) HQ (x, y, t) ψ(x) νh (x)
M
Z XQ
= ϕ(x, y) E(r, t) hk (x, y) ψ(x) tk νh (x)
M k=0
Z XQ
−n/2 − 41 |y−x|2 /t
= ϕ(x, y)(4πt) e hk (x, y) ψ(x) tk νh (x)
M k=0
Z  √ 1 2
ρ |w| t e√− 4 |w|
XQ 
−n/2 k
= (4π) t  √  νh (w) .
k=0 Rn ·hk y + w t, y ψ y + w t
Thus, as the hk are C ∞ , for any fixed t > 0,
k(GQ ψ)(·, t)kC l ≤ cl kψkC l
for some constant cl , and we have (17.88). For (17.89), note that each integrand
in the above final expression for (GQ ψ)(y, t) is bounded by an integrable function,
570 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

1 2
namely a constant multiple of e− 4 |w| . Thus, we may apply the Lebesgue Domi-
nated Convergence Theorem to obtain (17.89):
lim (GQ ψ)(y, t)
t→0+
Q  √ 1 2
ρ |w| t e√− 4 |w|
Z 
−n/2
X
= (4π) lim+ tk  √ νh (w)
t→0 Rn ·hk y + w t, y ψ y + w t
k=0
Z  √ 1 2  √   √
−n/2
= (4π) lim ρ |w| t e− 4 |w| h0 y + w t, y ψ y + w t νh (w)
n t→0+
ZR
−n/2 1 2
= (4π) e− 4 |w| h0 (y, y) ψ(y) νe (w) = h0 (y, y) ψ(y) = ψ(y) .
Rn
√ √
To see (17.90), we make the change of variables w = (z − x) / s, z = x + w s
in computing
Z tZ
(GQ ∗ K)(x, y, t) = ϕ(x, z) HQ (x, z, s) ◦ K(z, y, t − s) νh (z) ds
0 Rn
Z tZ XQ
−m − 41 |z−x|2 /s
= ϕ(x, z)(4πs) e hk (x, z) ◦ K(z, y, t − s) sk νh (z) ds
0 Rn k=0

(17.91)
Z tZ √ −m − 1 w2
!
ϕ(x, x + w s)(4π) e 4
= PQ √ √ νh (w) ds.
0 Rn · k=0 hk (x, x + w s) ◦ K(x + w s, y, t − s) sk
Moreover, we have

(GQ ∗ K)(x, y, t)
∂t
√ −m
exp − 41 w2
Z   
ϕ(x, x + w s)(4π)
= PQ √ √ νh (w)
Rn · k=0 hk (x, x + w s) ◦ K(x + w s, y, t − s) sk
Z tZ  √ −m
exp − 41 w2
 
ϕ(x, x + w s)(4π)
+ PQ √ ∂ √
0 Rn · k=0 hk (x, x + w s) ◦ ∂t K(x + w s, y, t − s) sk
(17.92) · νh (w) ds.
Since the hk are fixed C ∞ sections, the forms of (17.91) and (17.92) imply that the
derivatives of GQ ∗ K of order ≤ l on M × M × [0, T ] have bounds (depending on
T ) in terms of those of K, so that k(GQ ∗ K)kC l(T ) ≤ cl,T kKkC l(T ) . 

We set
KQ,0 (x, y, t) := ∂t + Dy2 GQ (x, y, t) and

Z tZ
KQ,j (x, y, t) := KQ,0 (x, z, s) ◦ KQ,j−1 (z, y, t − s) νh (z) ds
0 M
= (KQ,0 ∗ KQ,j−1 )(x, y, t) for j ≥ 1.
The next result implies that for Q sufficiently large,

X j+1
(17.93) κQ := GQ + (−1) GQ ∗ KQ,j
j=0

exists and is a general heat kernel for D2 ; in particular, κQ is independent of Q.


17.4. THE ASYMPTOTIC FORMULA FOR THE HEAT KERNEL 571

Theorem 17.48. For any integer k ≥ 2, if we choose Q > m+2k and  T > t0 >
k
0, the series (17.93)
 defining κQ (x, y, t) converges in C H|M ×M ×[t 0 ,T ] , κQ (x, y, t)
satisfies ∂t + Dx2 κQ = 0, and
Z
(17.94) lim κQ (y, x, t) ψ(y) νh (y) = ψ(x) ,
t→0+ M
0 ±
for all ψ ∈ C (E ⊗ Σ (M )) (i.e., κQ is a general heat kernel in the sense of
Definition 17.32, p. 547). Moreover,
(17.95) |κQ (x, y, t) − GQ (x, y, t)| ≤ CtQ−m+1 ,
for some constant C independent of (x, y, t) ∈ M × M ×(0, T ).
Proof. Once we prove that for k ≥ 2 and Q > m + 2k + 1 the series

X j+1
KQ (x, y, t) := (−1) KQ,j (x, y, t)
j=0

converges in C k H|M ×M ×[0,T ] , then according to Proposition 17.47,

X ∞
X
j+1 j+1
GQ + (−1) GQ ∗ KQ,j = GQ + GQ ∗ (−1) KQ,j
j=0 k=0
= GQ + GQ ∗ KQ ,

where the convergence of the first infinite sum is in C k H|M ×M ×[0,T ] . Thus, noting
that although GQ is not C k at t = 0,

X j+1
GQ ∗ KQ,j ∈ C k H|M ×M ×(0,∞)

κQ := GQ + (−1)
j=0

will exist and the convergence
 will be in C k H|M ×M ×[t0 ,T ] if 0 < t0 < T < ∞.
Then for k ≥ 2, ∂t + Dx2 κQ can be computed via term-by-term differentiation:
∂t + Dx2 κQ (x, y, t)


= ∂t + Dx2 GQ (x, y, t)



 t
X Z Z
j+1 2
+ (−1) ∂t + Dx GQ (x, z, s) ◦ KQ,j (z, y, t − s) νh (z) ds
j=0 0 M

X Z
j+1
= KQ,0 (x, y, t) + (−1) lim− GQ (x, z, t − s) ◦ KQ,j (z, y, s) νh (z)
s→t M
j=0

X Z tZ
j+1
+ (−1) KQ,0 (x, z, s) ◦ KQ,j (z, y, t − s) νh (z) ds
j=1 0 M

X ∞
X
j+1 j+1
= KQ,0 (x, y, t) + (−1) KQ,j (x, y, t) + (−1) KQ,j+1 (x, y, t) = 0.
j=0 k=0

To prove that the series for KQ (x, y, t) converges in the C k H|M ×M ×[t0 ,T ] , we will
estimate the terms of this sum and their derivatives so that the Weierstrass M -test
can be applied. Recall that
KQ,0 (x, y, t) := ∂t + Dx2 GQ (x, y, t) , where

572 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS


ϕ(x, y) HQ (x, y, t) , for (x, y) ∈ δr2 (M × M ) , t > 0,
GQ (x, y, t) :=
0, for (x, y) ∈
/ δr2 (M × M ) , t > 0,
and ϕ(x, y) := ρ(d(x, y)). Note that GQ (x, y, t) and all of its local derivatives
vanish for d(x, y) ≥ r2 , and so we assume that d(x, y) < r2 in what follows. For
d(x, y) < r1 , we have ϕ(x, y) = ρ(d(x, y)) = 1, and
KQ,0 (x, y, t) = ∂t + Dx2 GQ (x, y, t) = ∂t + Dx2 (ϕ(x, y) HQ (x, y, t))
 

= ∂t + Dx2 HQ (x, y, t) = E(d(x, y) , t) tQ D2 (UQ )(x, y) .




For r := d(x, y) ∈ [r1 , r2 ) ,


KQ,0 (x, y, t) = ∂t + Dx2 GQ (x, y, t) = ∂t + Dx2 (ϕ(x, y) HQ (x, y, t))
 

= ∂t + Dx2 HQ (x, y, t) + ∂t + Dx2 ((ϕ(x, y) − 1) HQ (x, y, t))


 

= E(r, t) tQ D2 (UQ )(y) + ∂t + Dx2 ((ϕ(x, y) − 1) HQ (x, y, t)) ,




2
and because of the factor t−m e−d(x,y) /4t
in HQ (x, y, t),
2
|KQ,0 (x, y, t)| ≤ E(r, t) t Q
D2 (UQ )(y) + C0 t−m−2 e−r1 /t
2
for some (x, y, t)-independent constant C0 . Since e−r1 /t is O(tk ) for all k > 0,
putting this together with the result for d(x, y) < r1 , we have
2
|KQ,0 (x, y, t)| ≤ CQ e−r /4t Q−m
t ≤ CQ tQ−m ,
for all (x, y, t) ∈ M × M × [0, ∞), for some (x, y, t)-independent constant CQ . Using
2
the same reasoning, and noting that differentiation of e−d(x,y) /4t with respect to
i i −1
local coordinates x or y introduces factors of t , while applying ∂t introduces
factors of t−2 , it is straightforward to see that
s
(17.96) (∂t ) ∂xi1 · · · ∂xip ∂yj1 · · · ∂yjq KQ,0 (x, y, t) ≤ CQ (p + q + s) tQ−m−p−q−2s ,
for some constant CQ (N ) depending monotonically on N = 0, 1, 2, . . . , say with
CQ (0) = CQ . We will use this below. Recall
Z tZ
|KQ,1 (x, y, t)| = KQ,0 (x, z, s) ◦ KQ,0 (z, y, t − s) νh (z) ds .
0 M
Then
Z tZ
|KQ,1 (x, y, t)| ≤ |KQ,0 (x, z, s) ◦ KQ,0 (z, y, t − s)| νh (z) ds
0 M
Z tZ
≤ |KQ,0 (x, z, s)| |KQ,0 (z, y, t − s)| νh (z) ds
0 M
Z tZ
Q−m
≤ CQ 2 sQ−m (t − s) νh (z) ds
0 M
Z t
Q−m
= V (M ) CQ 2 sQ−m (t − s) ds
0
2
Γ(Q − m + 1)
= V (M ) CQ 2 t2Q−2m+1 .
Γ(2Q − 2m + 2)
For the last equality, recall that for a, b ≥ 0,
Z 1
b Γ(a + 1) Γ(b + 1)
xa (1 − x) dx = .
0 Γ(a + b + 2)
17.4. THE ASYMPTOTIC FORMULA FOR THE HEAT KERNEL 573

Thus,
Z t Z t  a  
b s s b
sa (t − s) ds = ta tb 1 − ds
0 0 t t
Z t  a 
a+b s s b  s 
=t 1− ds substituting x = , t dx = ds
0 t t t
Z 1 Z 1
b b
= ta+b xa (1 − x) t dx = ta+b+1 xa (1 − x) dx
0 0

a+b+1 Γ(a+ 1) Γ(b + 1)


=t .
Γ(a + b + 2)
More generally, if f1 , f2 ∈ C 0 (H) with
|f1 (x, y, t)| ≤ C1 tp1 and |f2 (x, y, t)| ≤ C2 tp2 (for p1 , p2 ≥ 0),

Z tZ
|(f1 ∗ f2 )(x, y, t)| = f1 (x, z, s) ◦ f2 (z, y, t − s) νh (z) ds
0 M
Z t
p2
≤ C1 C2 V (M ) sp1 (t − s) ds
0
Γ(p1 + 1) Γ(p2 + 1)
= C1 C2 V (M ) tp1 +p2 +1 .
Γ(p1 + p2 + 2)
Applying this, we get the following fact that will also be needed,
 
Γ(p2 + 1) Γ(p3 + 1)
(17.97) |(f1 ∗ f2 ∗ f3 )(x, y, t)| ≤ C1 C2 C3 V (M )
Γ(p2 + p3 + 2)
Γ(p1 + 1) Γ(p2 + p3 + 2)
· V (M ) tp1 +p2 +p3 +2
Γ(p1 + p2 + p3 + 3)
2 Γ(p1 + 1) Γ(p2 + 1) Γ(p3 + 1) p1 +p2 +p3 +2
= C1 C2 C3 V (M ) t .
Γ(p1 + p2 + p3 + 3)
Consequently,
∗3 f1 (x, y, t) := |(f1 ∗ f1 ∗ f1 )(x, y, t)|


3
2 Γ(p1 + 1) 3p1 +2
≤ C13 V (M ) t ,
Γ(3(p1 + 1))
and by induction, we see that
 k factors

∗k f1 (x, y, t) := f1 ∗ · · · ∗ f1 (x, y, t)


k
k−1 Γ(p1 + 1) (p1 +1)k−1
≤ C1k V (M ) t .
Γ((p1 + 1) k)
Thus, for j = 0, 1, 2, . . . ,
|KQ,j (x, y, t)| = ∗j+1 KQ,0 (x, y, t)
j+1
j+1 j Γ(Q − m + 1)
≤ CQ V (M ) t(Q−m+1)(j+1)−1 .
Γ((Q − m + 1)(j + 1))
574 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

We have assumed that Q > m + 2k and k ≥ 2. Consequently, in the preceding


expression all the powers (Q − m + 1)(j + 1) − 1 of t are positive. Hence, the M -
P∞ j+1
test will apply to give uniform convergence of the series j=0 (−1) KQ,j (x, y, t)
on M × M × [0, T ] if for every constant R ≥ 0

X Rj
< ∞,
j=0
Γ((Q − m + 1)(j + 1))

Rj
P∞
but as Q > m, this holds by comparison with j=0 Γ(j+1) = eR .
k

We now show the C H|M ×M ×[0,T ] -convergence of

X j+1
KQ (x, y, t) = (−1) KQ,j (x, y, t) ,
j=0

for Q > m + 2k. Through repeated use of (17.86) and (17.87), we have for j ≥ 3
s
(∂t ) ∂xi1 · · · ∂xip ∂yj1 · · · ∂yjq KQ,j (x, y, t)
 
s  
= (∂t ) ∂xi1 · · · ∂xip KQ,0 ∗ KQ,j−3 ∗ ∂yj1 · · · ∂yjq KQ,0 (x, y, t) .

For N = p + q + s ≤ k, we have
s
|(∂t ) ∂xi1 · · · ∂xip KQ,0 | ≤ CQ (N ) tQ−m−p−2s ,
j−2 j−3
|KQ,j−3 | ≤ CQ (N ) V (M )
j−2
Γ(Q − m + 1)
· t(Q−m+1)(j−2)−1 , and
Γ((Q − m + 1)(j − 2))
(17.98) ∂yj1 · · · ∂yjq KQ,0 ≤ CQ (N ) tQ−m−q ,
where all the powers of t are positive, since Q > m + 2k ≥ m + 2(p + q + s). Using
(17.86), (17.87), (17.97), and (17.98), we have
s
| (∂t ) ∂xi1 · · · ∂xip ∂yj1 · · · ∂yjq KQ,j (x, y, t) |
 
s  
= (∂t ) ∂xi1 · · · ∂xip KQ,0 ∗ KQ,j−3 ∗ ∂yj1 · · · ∂yjq KQ,0 (x, y, t)
j−2
j Γ(Q − m + 1)
j−1
≤ CQ (N ) V (M ) · t(Q−m+1)j−(q+2s+p)−1
Γ((Q − m + 1)(j − 2))
Γ(Q − m − p − 2s + 1) Γ((Q − m + 1)(j − 2)) Γ(Q − m − q + 1)
·
Γ(Q − m − p − 2s +(Q − m + 1)(j − 2) − 1 + Q − m − q + 3)
j j−1
≤ CQ (N ) V (M ) · t(Q−m+1)j−(q+2s+p)−1
j−2
Γ(Q − m − p − 2s + 1) Γ(Q − m + 1) Γ(Q − m − q + 1)
· .
Γ((Q − m + 1) j −(p + 2s + q))
Since Γ((Q − m + 1) j −(p + 2s + q)) ≥ Cj!
Pfor some constant C, the series of these
terms can be estimated from above, using Rj /j! < ∞ as in the case k = 0. Thus,
the M -test can be used to obtain the uniform convergence (for Q > m + 2k) on
M × M × [0, T ] of

X j+1 s
(−1) (∂t ) ∂xi1 · · · ∂xip ∂yj1 · · · ∂yjq KQ,j (x, y, t) ,
j=0
17.4. THE ASYMPTOTIC FORMULA FOR THE HEAT KERNEL 575

 P∞ j+1
and hence the C k H|M ×M ×[0,T ] -convergence of KQ = j=0 (−1) KQ,j (x, y, t).
P∞ j+1
We now prove (17.95). From κQ = GQ + j=0 (−1) GQ ∗ KQ,j , we have

|κQ (x, y, t) − GQ (x, y, t)| = |(GQ ∗ KQ )(x, y, t)| .

Moreover for some constant C,


X∞
|KQ (x, y, t)| ≤ |KQ,j (x, y, t)|
j=0
j+1
X∞ j+1 j Γ(Q − m + 1)
≤ CQ V (M ) t(Q−m+1)(j+1)−1
j=0 Γ((Q − m + 1)(j + 1))
j+1
X∞ j+1 j Γ(Q − m + 1)
= tQ−m CQ V (M ) t(Q−m+1)j
j=0 Γ((Q − m + 1)(j + 1))
Q−m
(17.99) ≤ Ct .

With K = KQ in (17.91), we obtain

(GQ ∗ KQ )(x, y, t)
Z tZ √ −m − 1 w2
!
ϕ(x, x + w s)(4π) e 4
= PQ √ √ νh (w) ds.
0 M · k=0 hk (x, x + w s) ◦ KQ (x + w s, y, t − s) sk

Then using (17.99), for some constants C 0 and C 00 , we have

|κQ (x, y, t) − GQ (x, y, t)| = |(GQ ∗ KQ )(x, y, t)|


Q
Z tX
0 Q−m k
≤ C (t − s) s ds
0 k=0
Q
X
= C 0 tQ−m+k+1
k=0
≤ C 00 tQ−m+1 as t → 0+ ,

which yields (17.95). Finally, note that


Z
κQ (y, x, t) ψ(y) vy − ψ(x)
M
Z Z
≤ (κQ (y, x, t) − GQ (y, x, t)) ψ(y) vy + GQ (y, x, t) ψ(y) vy − ψ(x)
M M
Z Z
≤ CtQ−m+1 |ψ(y)| vy + GQ (y, x, t) ψ(y) vy − ψ(x) .
M M

Thus, (17.94) follows from (17.89) in Proposition 17.47, p. 569. 


576 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

5. The Local Index Formula


Content and Meaning of the Local Index Formula. In the previous
section, we showed that for x, y ∈ M (of even dimension n = 2m) with r = d(x, y),
denoting the Riemannian distance from x to y, sufficiently small, the heat kernel
k(x, y, t) of D2 has an asymptotic expansion as t → 0+ of the form
XQ
k(x, y, t) ∼ HQ (x, y, t) := E(r, t) hj (x, y) tj ,
j=0

for any fixed integer Q > m + 4, where


 
hj (x, y) ∈ Hom (E ⊗ Σ(M ))x ,(E ⊗ Σ(M ))y , j ∈ {0, 1, . . . , Q} .

In particular, with y = x, we have


Q Q
−m −m
X X
j
(17.100) k(x, x, t) ∼ (4πt) hj (x, x) t = (4π) hj (x, x) tj−m .
j=0 j=0

From (17.39) we know that for any t > 0,


Z
Str(k(x, x, t)) νh (x) = index D+ ,

M

which is constant, independent of t. As t → 0+ , we deduce from (17.100) that


Z
Str(hj (x, x)) νh (x) = 0 for j ∈ {0, 1, . . . , m − 1} , while
M
Z
−m
Str(hm (x, x)) νh (x) = index D+ .

(17.101) (4π)
M

We have a somewhat cumbersome formula for hm (x, x), and it is already clear
from the formula that hm (x, x) is determined by the metric h and the connection
ε on U (E) in a neighborhood of the point x (indeed by finitely many derivatives
of h and ε at x). One might regard the gist of the Index Formula for twisted
Dirac operator as exhibiting the global quantity index(D+ ) as the integral of a
form which may be locally computed. From this perspective, (17.101) does the
job. The direct computation of hm (x, x) is actually not very difficult for m = 1
and 2 (i.e., for dim(M ) = 2 or 4), and we will carry it out explicitly. However,
for large values of m it is rather cumbersome and one would like a more tractable
formula for hm (x, x). It is also desirable to express hm (x, x) in terms of curvature
forms, thereby showing that hm (x, x) only depends on the 2-jet of the metric and
the 1-jet of ε. (For a rigorous definition of k-jets we refer back to Section 6.2 with
Definition 6.10, p.166.) Moreover, since index(D+ ) is insensitive to perturbations
in h and ε, one expects that hm (x, x) νh can be expressed in terms of curvature
forms of characteristic classes for M and E. The Local Index Formula below does
this. Moreover, since the Local Index Formula is a purely local result, it may be
applied to obtain the (global) Index Formula for elliptic operators which are only
locally of the form of twisted Dirac operator D+ . Indeed, if A is such an operator
(possibly on a nonspin manifold) and k is the heat kernel for A∗ A ⊕ AA∗ , then
from the spectral resolution of A, we still have
Z
Str(hm (x, x)) νh (x) = index(A) ,
M
17.5. THE LOCAL INDEX FORMULA 577

where the supertrace Str is defined in the natural way. Now, the crucial observation
is that the Local Index Formula allows us to compute Str(hm (x, x)) once A is
represented locally as a twisted Dirac operator. We will see a number of examples
of such A in the next section. In fact, it is not easy to find any first-order elliptic
operators of geometrical significance which are not locally twisted Dirac operators,
or 0-th order perturbations thereof. Our goal in this section is to prove

Theorem 17.49 (The Local Index Formula). Let dim M = 2m and x ∈ M .


Let B be some neighborhood about x and D : C ∞ (B, E(B) ⊗ Σ(B)) ←- a twisted
Dirac operator. Let D be determined by a connection ε on E(B), a metric h on B
with Levi-Civita connection θ, and a spin structure over F B. Let hm (x, x) be as in
the diagonal asymptotic expansion (17.100) of the heat kernel of D2 . If Ωε denotes
the curvature of ε and Ωθ denotes the curvature of θ, then

−m
D ε
 iΩθ /4π 1/2 E
(17.102) (4π) Str(hm (x, x)) = Tr(eiΩ /2π ) ∧ det , νh (x) ,
sinh(iΩθ /4π)

where the meaning of the right side was explained in the paragraph following (17.40),
p. 544.

How the Curvature Terms  Arise in the Heat Asymptotics. Using nor-
mal coordinates y 1 , . . . , y 2m ∈ B(r0 , 0) about x ∈ M and the radial gauge, and
selecting V0 ∈ CN ⊗ Σ2m , we have (see (17.80), (17.78) and Proposition 17.45,
p. 564),
(17.103) Z Y
m−1 i
  h i
hm (x, x)(V0 ) = (si ) Asm−1 ◦ · · · ◦ As0 Ve0 (0) ds0 . . . dsm−1 ,
Im i=0

where Ve0 ∈ C ∞ B(r0 , 0) , CN ⊗ Σ



2m is the constant extension of V0 . Recall that
∞ N

for ψ ∈ C B(r0 , 0) , C ⊗ Σ2m , we have

As [ψ](y) := A [ψ](sy) , where


h i
A [ψ] := h1/4 ∆ h−1/4 ψ − F + 14 S [ψ] .


While the right side of (17.103) may seem unwieldy, there is substantial simpli-
h i
fication due to facts that Asm−1 ◦ · · · ◦ As0 V0 (y) is evaluated at y = 0 in
e
h i
(17.103) and that only those terms of Asm−1 ◦ · · · ◦ As0 Ve0 (0) which involve
γn+1 := γ1 · · · γ2m will survive when the supertrace Str(hm (x, x)) is taken (see
Proposition 17.13, p. 522). To see how we may take advantage of these facts, we
578 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

need to expand
h i h i
−1 −1
A [ψ] = −h /4 D2 h /4 ψ = h /4 ∆ h /4 ψ − F + 14 S [ψ]
1 1


−1
√ −1
= h /4 ∇∂j hij h∇∂i (h /4 ψ) − F + 14 S [ψ]
 
 √  !
−1/4 ∂j hij h∇∂i h−1/4 ψ 1

=h √ − F + 4 S [ψ]
e j ) ∇∂ h−1/4 ψ

+ hij h(ε ⊕ θ)(∂ i
  √  
−1/4
∂j hij h ∂i (h−1/4 ψ) + (ε ⊕ θ)(∂
e i ) h−1/4 ψ
=h  √   
e (∂j ) ∂i (h−1/4 ψ) + (ε ⊕ θ)(∂
+ hij h(ε ⊕ θ) e i ) h−1/4 ψ

− F + 14 S [ψ]

 √   
−1 −1 −1
= h /4 ∂j hij h∂i (h /4 ψ) + h /4 ∂j h /4 hij (ε ⊕ θ)(∂
1
e i) ψ
1
e j ) ∂i (h−1/4 ψ) + hij (ε ⊕ θ)(∂
+ hij h /4 (ε ⊕ θ)(∂ e j ) (ε ⊕ θ)(∂
e i) ψ
− 21 Ωεjk ⊗ γ j γ k ψ − 14 Sψ.


In order to exhibit the parts of A which are of pure order 0, 1 and 2, we expand
further:
 
−1 −1
A [ψ] = hij ∂j ∂i ψ + h /4 ∂j (hij h /4 ) + h /4 hji ∂j (h /4 ) ∂i ψ
1 1

 √ 
−1 −1
+ h /4 ∂j hij h∂i (h /4 ) ψ
 
e i ) ∂j ψ + h−1/4 ∂j h1/4 hij (ε ⊕ θ)(∂
+ hij (ε ⊕ θ)(∂ e i) ψ
−1/4
+ hij (ε ⊕ θ)(∂
e j ) ∂i ψ + hij h /4 (ε ⊕ θ)(∂
1
e j ) ∂i (h )ψ
+ hij (ε ⊕ θ)(∂ 1
Ωεjk ⊗ γ j γ k
ψ − 14 Sψ

e j ) (ε ⊕ θ)(∂
e i) ψ −
2
= hij ∂j ∂i ψ
 
−1 −1
+ h /4 ∂j (hij h /4 ) + h /4 hji ∂j (h /4 ) ∂i ψ
1 1

+ hij (ε ⊕ θ)(∂
e i ) ∂j ψ + hij (ε ⊕ θ)(∂
e j ) ∂i ψ
 
−1
+ h /4 ∂j h /4 hij (ε ⊕ θ)(∂
1
e j ) ∂i (h−1/4 )ψ
e i ) ψ + hij h1/4 (ε ⊕ θ)(∂

+ hij (ε ⊕ θ)(∂ e i ) ψ − 1 Ωε ⊗ γ j γ k ψ − 1 Sψ.



e j ) (ε ⊕ θ)(∂
2 jk 4
Finally,
A [ψ] = hij ∂j ∂i ψ
 
−1 −1
+ h /4 ∂j (hij h /4 ) + h /4 hji ∂j (h /4 ) + 2hji (ε ⊕ θ)(∂
1 1
e j ) ∂i ψ
   
h−1/4 ∂j h1/4 hij (ε ⊕ θ)(∂ e j ) ∂i (h−1/4 )
e i ) + hij h1/4 (ε ⊕ θ)(∂
(17.104) +    ψ.
+ hij (ε ⊕ θ)(∂ e i ) − 1 Ωε ⊗ γ j γ k − 1 S
e j ) (ε ⊕ θ)(∂
2 jk 4

We wish to alter the operator A in such a way that the alteration does not affect
Str(hm (0, 0)), but the altered operator is much simpler. Note that in (17.104), the
differential operator A has three parts of orders 2, 1, and 0 which contain 1, 3, and 5
terms each. Correspondingly, there are (1 + 3 + 5)m = 9m terms in the composition
Asm−1 ◦ · · · ◦ As0 . However, any of these terms of Asm−1 ◦ · · · ◦ As0 which involve
17.5. THE LOCAL INDEX FORMULA 579

fewer than 2m gamma matrices γ i will not contribute to Str(hm (0, 0)). The term
of A which produces the most (four) γ i is a subterm of hij (ε ⊕ θ)(∂ e j ) (ε ⊕ θ)(∂
e i ),
namely
(17.105)
 
hij θ(∂ e i ) ψ = hij Rklpj (0) 1 γ k γ l y p Rk0 l0 qi (0) 1 γ k0 γ l0 y q ψ + O |y|3 .

e j ) θ(∂
8 8

Although this subterm contributes four γ , it introduces two factors of y i which


i

must be differentiated by terms in subsequent factors of Asm−1 ◦ · · · ◦ As0 in order


to contribute to Str(hm (0, 0)). We say that the degree of contribution (degc ) of
the subterm (17.105) to Str(hm (0, 0)) is 4 − 2 or 2. Note that hij ∂j ∂i ψ has no γ i ,
but it contains two differentiations which can serve to eliminate two factors of yi in
previous terms. Hence, although hij ∂j ∂i contains no γ i , we take degc hij ∂j ∂i ψ =
2. In general, we have
Definition 17.50. For B(0, r0 ) = {y ∈ Rn : |y| ≤ r0 }, let
T : C ∞ B(0, r0 ) , CN ⊗ Σ2m → C ∞ B(0, r0 ) , CN ⊗ Σ2m
 

be an operator of the form


X ···ik
T (ψ)(y) = fji11···j p
(y) γ j1 · · · γ jp (∂i1 · · · ∂ik ψ)(y) ,
(i),(j)
···ik
where fji11···j p
∈ C ∞ (B(0, r0 ) , R) and (i) ranges over all multi-indices (i1 , . . . , ik ) ∈
k
× {1, . . . , n}. Then the degree of contribution of T is
degc (T ) := p + k − r, where
n o
···ik d
r := min d : fji11···j p
(y) = O |y| ;
(i),(j)
i.e., r is smallest of the degrees of the lead parts of the Taylor expansions of the
···ik
coefficients fji11···jp
. For an integer q ≥ 0, we use Oc (q) to denote an operator with
degc ≤ q.
Referring to (17.104), the only term in
     
−1 −1
h /4 ∂j hij h /4 + h /4 hji ∂j h /4 + 2hji (ε ⊕ θ)(∂
1 1
e j ) ∂i

which has an degc of at least 2 is


2hji θ(∂
e j ) ∂i .
In (17.104), the sum of the terms of
    
h−1/4 ∂j h1/4 hij (ε ⊕ θ)(∂ e j ) ∂i h−1/4
e i ) + hij h1/4 (ε ⊕ θ)(∂
   
+ hij (ε ⊕ θ)(∂ e i ) − 1 Ωε ⊗ γ j γ k − 1 S
e j ) (ε ⊕ θ)(∂
2 jk 4

which have degc ≥ 2 is


 
hij ∂j θ(∂
e i ) + hij θ(∂ 1
Ωεjk ⊗ γ j γ k .

e i) −
e j ) θ(∂
2

Since there are only m factors in Asm−1 ◦ · · · ◦ As0 , the maximal number of γ i in
factors of the terms of Asm−1 ◦· · ·◦As0 will be 2m. To achieve this maximal number
it is necessary (but not sufficient) that such a term be a composition of m factors
each with the maximal degc , namely 2. Thus, we can alter A without changing
Str(hm (0, 0)) by retaining only those terms of A with degc = 2. Moreover, only
580 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

the lead part of the Taylor expansion in y of such terms need be retained. In the
next lemma, we collect the relevant expansions that we have already found for the
metric and the connections in terms of normal coordinates and the radial gauge, at
least to the order we need.

Lemma 17.51. We have


2 2
εy (∂j ) = 12 Fij y i + O |y| = 12 Ωεij (0) y i + O |y| ,
2
θey (∂j ) = 18 Rklji (0) γ k γ l y i + O |y| ,
3
hij (y) = δij − 31 Rikjl (0) y k y l + O |y| ,
3
hij (y) = δij + 31 Rikjl (0) y k y l + O |y| ,
α 3
hα := (det h) = 1 − 13 αRkl (0) y k y l + O |y| ,
X
1 ε j k 1
Ωεij (0) ⊗ γ i γ j + O |y| , and

2 Ωjk γ γ ψ = 2 i,j
i ij 2
γ = h γj = γi + O |y| .

We define

θe1 (∂j ) := 1 k l i
8 Rklji (0) γ γ y and
X X
0 i j
(17.106) F := 1
2 Fij ⊗ γ γ = 1
2 Ωεij (0) ⊗ γ i γ j .
i,j i,j

Thus, using
 
1 k l j

∂i θ(∂
e i ) = ∂i
8 Rklij (0) γ γ y

= 18 Rklij (0) γ k γ l δij = 18 Rklii (0) γ k γ l = 0,

a permissible alteration of A is
X   
A0 := ∂i2 + 2θ(∂
e i ) ∂i + ∂i θ(∂ e i )2 − F 0
e i ) + θ(∂
i
X 
= ∂i2 + 2θ(∂ e i )2 − F 0 .
e i ) ∂i + θ(∂
i

We now show that even though


   X 
e i ) ∂i = deg 1
degc 2θ(∂ c 4 Rklij (0) γ k γ l y j ∂i = 2,
j

the term 2θ(∂ e i ) ∂i may be dropped without changing Str(hm (0, 0)). Note that if
this 2θ(∂i ) ∂i is on the far left in a term of A0sm−1 ◦ · · · ◦ A0s0 , then the contribution
e
of this term to A0sm−1 ◦ · · · ◦ A0s0 [V0 ](0) will be zero, because of the factor of y j in
2θ(∂
e i ) ∂i which is set to 0. Thus, it suffices to prove that 2θ(∂ e i ) ∂i commutes with
2
∂i2 , θ(∂
e i ) and F 0 , modulo terms which do not affect Str(hm (0, 0)). In other words,
we need to check that the commutators
(17.107)
hX i hX i hX i
X X
e i ) ∂i , F 0 ,
θ(∂ θ(∂
e i ) ∂i , ∂ 2 , and θ(∂
e i ) ∂i , e j )2
θ(∂
j
i i j i j
17.5. THE LOCAL INDEX FORMULA 581

each have degc < 4. We have


hX i hX X i
e i ) ∂i , F 0
θ(∂ = e i ) ∂i , 1
θ(∂ F jk ⊗ γ j k
γ
i i 2 jk
X h i
1
= 2 e i ) ∂i , Fjk ⊗ γ j γ k
θ(∂
i,j,k
X
1
Rpliq (0) γ p γ l y q ∂i , Fjk ⊗ γ j γ k
1 
= 2
i,j,k,p 8
X
1 q
 p l j k
= 12 8 y Rpliq (0) Fjk ⊗ γ γ , γ γ ∂i .
i,j,k

  hP i
Since degc γ p γ l , γ j γ k ≤ 2, we have degc i
e i ) ∂i , F 0 ≤ 2. Moreover,
θ(∂

hX X i
θ(∂
e i ) ∂i , ∂j2
i j
X  
= θ(∂i ) ∂i ∂j2 − ∂j2 θ(∂
e e i ) ∂i
i,j
X    
= e i ) ∂i ∂ 2 − ∂j ∂j θ(∂
θ(∂ e i ) ∂i + θ(∂
e i ) ∂j ∂i
j
i,j
X      
= e i ) ∂i ∂ 2 − ∂ 2 θ(∂
θ(∂ e i ) ∂i + 2∂j θ(∂ e i ) ∂j ∂i + θ(∂ e i ) ∂ 2 ∂i
j j j
i,j
X   X
−2∂j − 18 Rklji (0) γ k γ l y j ∂j ∂i

= −2∂j θ(∂ e i ) ∂j ∂i =
i,j i,j
X
= 14 Rklji (0) γ k γ l ∂j ∂i = 0,
i,j

sinceh
Rklji is antisymmetric i in i and j, while ∂j ∂i is symmetric. We show
P e P e 2
degc i θ(∂i ) ∂i , j θ(∂j ) ≤ 2 as follows:

hX X i
θ(∂
e i ) ∂i , e j )2
θ(∂
i j
X  
= θ(∂ e j )2 − θ(∂
e i ) ∂i ◦ θ(∂ e j )2 θ(∂
e i ) ∂i
i,j
X  
= θ(∂ e j )2 + θ(∂
e i ) ∂i θ(∂ e j )2 ∂i − θ(∂
e i ) θ(∂ e j )2 θ(∂
e i ) ∂i
i,j
X   h i
= θ(∂ e j )2 + θ(∂
e i ) ∂i θ(∂ e j )2 ∂i
e i ) , θ(∂
i,j
X  
= θ(∂ e j )2 + Oc (4 − 3 + 1)
e i ) ∂i θ(∂
i,j
X  
= θ(∂ e j )2 + Oc (2) .
e i ) ∂i θ(∂
i,j

The first term appears to have degc 4 which is too high, but we now show that it
actually has degc 2. Using

e j ) = Rrsqj γ r γ s y q and
−8θ(∂
2 0 0 0
82 θ(∂
e j ) = Rrsqj Rr0 s0 q0 j γ r γ s γ r γ s y q y q ,
582 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

we have
 
− 83 θ(∂
e i ) ∂i θ(∂ e j )2
 0 0 0

= Rklpi γ k γ l y p ∂i Rrsqj Rr0 s0 q0 j γ r γ s γ r γ s y q y q
0 0
 0

= Rklpi γ k γ l y p Rrsqj Rr0 s0 q0 j γ r γ s γ r γ s δiq y q + δiq0 y q
 0 0
 0
 0 0

= Rklpi Rrsij Rr0 s0 q0 j γ k γ l γ r γ s γ r γ s y p y q + Rklpi Rrsqj Rr0 s0 ij γ k γ l γ r γ s γ r γ s y p y q
 0 0
  0 0

= Rklpi Rrsij Rr0 s0 qj γ k γ l γ r γ s γ r γ s y p y q + Rklpi Rrsqj Rr0 s0 ij γ k γ l γ r γ s γ r γ s y p y q
 0 0

= Rklpi Rr0 s0 ij Rrsqj γ k γ l γ r γ s γ r γ s y p y q
 0 0

+ Rklpi Rrsqj Rr0 s0 ij γ k γ l γ r γ s γ r γ s y p y q + Oc (4 − 2)
 0 0

= 2Rklpi Rrsqj Rr0 s0 ij γ k γ l γ r γ s γ r γ s y p y q + Oc (2) .

The first term which appears to be Oc (4) is actually Oc (2), since we have by trivially
interchanging the indices and using the symmetry of R:
 0 0

Rklpi Rrsqj Rr0 s0 ij γ k γ l γ r γ s γ r γ s y p y q
 0 0

= Rklqi Rrspj Rr0 s0 ij γ k γ l γ r γ s γ r γ s y p y q
 0 0

= −Rklqi Rrspj Rr0 s0 ji γ k γ l γ r γ s γ r γ s y p y q
 0 0

= −Rklqj Rrspi Rr0 s0 ij γ k γ l γ r γ s γ r γ s y p y q
 0 0

= −Rrspi Rklqj Rr0 s0 ij γ k γ l γ r γ s γ r γ s y p y q
 0 0

= −Rklpi Rrsqj Rr0 s0 ij γ r γ s γ k γ l γ r γ s y p y q
 0 0

= −Rklpi Rrsqj Rr0 s0 ij γ k γ l γ r γ s γ r γ s y p y q + Oc (4 − 2) ,
 
where we have used γ r γ s , γ k γ l = Oc (2) in the last equality.
Thus, we have shown that all the commutators in (17.107) have degc ≤ 2 < 4.
In summary, we have

Proposition 17.52. Let


X
θe1 (∂j ) := 18 Rklji (0) γ k γ l y i ,
k,l,i
X X
F 0 := 12 Fij ⊗ γ i γ j = 12 Ωεij (0) ⊗ γ i γ j , and
i,j i,j
X 
(17.108) A0 := e i )2 − F 0 .
∂i2 + θ(∂
i

For V0 ∈ CN ⊗ Σ2m , define


(17.109) Z Y
m−1 i
  h i
0
hm (0, 0)(V0 ) := (si ) A0sm−1 ◦ · · · ◦ A0s0 Ve0 (0) ds0 . . . dsm−1 ,
Ik i=0
17.5. THE LOCAL INDEX FORMULA 583

where Ve0 ∈ C ∞ B(r0 , 0) , CN ⊗ Σ2m is the constant extension of V0 ∈ CN ⊗ Σ2m .




Then
Str(hm (0, 0)) = Str h0m (0, 0) .

(17.110)
In other words, in the computation of Str(hm (0, 0)) given by (17.103), we may
replace A by A0 .
At this point, we have that
Z
−m
index D+ = (4π)

Str(hm (x, x)) νh (x) ,
M
where the form Str(hm (x, x)) νh (x) only depends on νh (x), the curvature of the
Levi-Civita connection of h at x, and the curvature of the connection ε on U (E)
at x. Since the index(D+ ) is invariant under perturbations of the metric h and the
connection ε, we suspect that in fact Str(hm (x, x)) νh (x) can be expressed in terms
of curvature forms of characteristic classes, whose integrals are also invariant. It
remains to do this by establishing (17.102). Before doing this in general, we verify
(17.102) in the cases m = 1 and m = 2 (i.e., for surfaces and 4-manifolds). For
readers who have no use for the Local Index Theorem beyond dimension 4, this is
sufficient. For m = 1 and m = 2, we do not need to use Mehler’s Formula but
rather we proceed directly using (17.109), (17.110) and
A0s [ψ](y) := A0 [ψ](sy)
X X X
∂i2 [ψ] (sy) + s2 e i )2 ψ(sy) − 1
Fjk ⊗ γ j γ k ψ(sy) .
 
= θ(∂ 2 j,k
i i

The case m = 1 (surfaces).


Verification of the Local Index Formula for m = 1. We have
Z 1 Z 1
0
h01 (0, 0)(V0 ) = (s0 ) A0s0 [V0 ](0) ds0 = A0s0 [V0 ](0) ds0
0 0
Z 1 Z 1
= A0 [V0 ](s0 0) ds0 =A0 [V0 ](0) ds0 = A0 [V0 ](0)
0 0
X 
e i )2 [V0 ](0) − 1
= ∂i2 [V0 ](0) + θ(∂ Fjk ⊗ γ j γ k (V0 )
2 j,k
X
1 j k

= −2 Fjk ⊗ γ γ (V0 ) .
j,k
m
For m = 1, Str(γ2m+1 ) = (−2i) yields Str(γ1 γ2 ) = −2i. Hence,
 
−1 −1
X
(4π) Str(h1 (0, 0)) = (4π) Str − 12 Fjk ⊗ γ j γ k
j,k
−1 −1 1
= (4π) Tr(−F12 ) Str(γ1 γ2 ) = −(4π) Tr(Ωε12 (0))(−2i)
i
= Tr Ωε12 (0) .

Thus, we have (17.102) in the case m = 1, since
*  12 +
iΩθ /4π
 ε  
iΩ /2π
Tr e ∧ det , νh (0)
sinh(iΩθ /4π)
i
= hTr(iΩε (0) /2π) , νh (0)i = Tr Ωε12 (0) . 

584 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

The case m = 2 (4-manifolds). For m = 2, we also have an elementary


proof of the Local Index Formula.

Verification of the Local Index Formula (17.102) for m = 2. To be-


gin with we shall prove the following equation:

X 0 0
(17.111) h02 (0, 0)(V0 ) = 1 1
6 82 2 Rklpi (0) Rk0 l0 pi (0) γ k γ l γ k γ l [V0 ]
i,p,k,l,k0 ,l0
 
j 0 k0 j k
X
1 ε ε
+ 8 Ω 0
j k 0 (0) ◦ Ω jk (0) ⊗ γ γ γ γ [V0 ] .
j,k,j 0 ,k 0

Indeed,

Z 1 Z 1
1 0
h02 (0, 0)(V0 ) = (s1 ) (s0 ) A0s1 A0s0 [V0 ](0) ds0 ds1
0 0
Z 1 Z 1
1
= (s1 ) A0 A0s0 [V0 ](s1 0) ds0 ds1
0 0
Z 1 Z 1
1
(s1 ) A0 A0s0 [V0 ] (0) ds0 ds1
 
=
0 0
Z 1
1
A0 A0s0 [V0 ] (0) ds0
 
= 2
0
Z 1 h X X i
= 1
A0 s20 e i )2 V0 −
θ(∂ 1
Fjk ⊗ γ j γ k V0 (0) ds0
2 i 2 j,k
0
Z 1 hX
i h X i
= 1
s2 0
A e i )2 V0 (0) ds0 − 1 A0 1
θ(∂ Fjk ⊗ γ j k
γ V 0 (0)
2 0 2 2
0 i j,k
hX i h X i
1 0 e i )2 V0 (0) − 1 A0 1 Fjk ⊗ γ j γ k V0 (0)

= 6A θ(∂ 2 2
i j,k
h i  
j 0 k0
X 2 e 2
X
1 1 1 1 j k

= 6 (∂j ) θ(∂ i ) V0 + 2 · 2 F j 0 k0 ⊗ γ γ
2 Fjk ⊗ γ γ V0
i,j j,k,j 0 ,k0
X h i X  0 0

2 e 2
= 1
6 (∂j ) θ(∂ i ) V0 + 8
1
0 0
Fj 0 k0 ◦ Fjk ⊗ γ j γ k γ j γ k [V0 ] .
i,j j,k,j ,k

Now we observe

X h i
1 2 e i )2 V0
6 (∂j ) θ(∂
i,j
X 2  2
= 1
6 (∂j ) Rklip (0) 81 γ k γ l y p [V0 ]
i,j,p,k,l
 
k l k 0 l0 p q
X 2
1 1
= 6 82 (∂j ) Rklip (0) Rk 0 l0 iq (0) γ γ γ γ y y [V0 ]
i,j,p,q,k,l,k0 ,l0
X 0 0
= 61 812 2 0 0
Rklip (0) Rk0 l0 ip (0) γ k γ l γ k γ l [V0 ] .
i,p,k,l,k ,l

Thus (17.111) is proved.


17.5. THE LOCAL INDEX FORMULA 585

 0 0

m 2
Since Str(γ2m+1 ) = (−2i) ⇒ Str γ k γ l γ k γ l = (−2i) εklk0 l0 when m = 2,
we get the simple re-arrangement
 
k l k 0 l0
X
Str(h2 (0, 0)) = 13 812 Tr(IdE ) R klip (0) R k 0 l0 ip (0) Str γ γ γ γ
i,p,k,l,k0 ,l0
X  0 0 
1
Tr Ωεj0 k0 (0) ◦ Ωεjk (0) Str γ j γ k γ j γ k

+8 0 0
j,k,j ,k
X 2
1 1
= 3 82 dim E · Rklip (0) Rk0 l0 ip (0)(−2i) εklk0 l0
i,p,k,l,k0 ,l0
X 2
+ 81 Tr Ωεj0 k0 (0) ◦ Ωεjk (0) (−2i) εj 0 k0 jk

0
j,k,j ,k 0
X
= 31 842 dim E · Ripkl (0) Rpik0 l0 (0) εklk0 l0
i,p,k,l,k0 ,l0
X
− 21 Tr Ωεj0 k0 (0) ◦ Ωεjk (0) εj 0 k0 jk .

0
j,k,j ,k 0

We rewrite this as follows. In R4 , for 2-forms α = 12 k,l αkl dy k ∧ dy l and β =


P
0 0
1 k l
P
k,l βk l dy ∧ dy ,
0 0
2
X
αkl βk0 l0 εklk0 l0 = 4 hα ∧ β, νh i .
k,l,k0 ,l0

Thus, we have
X
Tr Ωεj0 k0 (0) ◦ Ωεjk (0) εj 0 k0 jk = 4 hTr(Ωε ∧ Ωε ) , νh (0)i .

j,k,j 0 ,k0
and
Str(h2 (0, 0))
X
= 1 4
3 82 dim E · Rpikl (0) Ripk0 l0 (0) εklk0 l0 − 2 hTr(Ωε ∧ Ωε ) , νh (0)i .
i,p,k,l,k0 ,l0
We can write the first sum as follows. Since
X
Ωθpi = 21 Rpikl dy k ∧ dy l , we have
k,l
X
Ωθ ∧ Ωθ Ωθjp ∧ Ωθpi
 
ji
=
p
X 0 0
= 1
4 Rjpkl Rpik0 l0 dy k ∧ dy l ∧ dy k ∧ dy l
p,k,l,k0 ,l0
X
1
= 4 (Rjpkl Rpik0 l0 εklk0 l0 ) ν, and
p,k,l,k0 ,l0
X X
Tr Ωθ ∧ Ω θ
Ωθ ∧ Ωθ ii = 14
 
= 0 0
(Ripkl Rpik0 l0 εklk0 l0 ) ν.
i i,p,k,l,k ,l
Thus,
Str(h2 (0, 0))
1 4
Tr Ωθ ∧ Ωθ , νh (0) − 2 hTr(Ωε ∧ Ωε ) , νh (0)i

= 3 82 dim E · 4
1
Ωθ ∧ Ωθ − 2 Tr(Ωε ∧ Ωε ) , νh (0) .

= 12 dim E · Tr
From the definition (15.102) of the Pontryagin forms we get
1 X
p1 Ωθ = δ j1 j2 Ωθi1 j1 ∧ Ωθi2 j2

2
(2π) 2! i1 ,i2 ,j1 ,j2 i1 i2

1 X 1
Ωθi1 i2 ∧ Ωθi2 i1 = − 2 Tr Ωθ ∧ Ωθ .

=− 2
8π i1 ,i2 8π
586 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

In dimension 4, we have
 12
iΩθ /4π

1
= 1 − p1 Ωθ , and

det θ
sinh(iΩ /4π) 24
 2
 ε 
iΩ /2π i ε i
Tr e = Tr IE + Tr(Ω ) + 2 1
Tr(Ωε ∧ Ωε )
2π 2π
i 1
= dim E + Tr(Ωε ) − 2 Tr(Ωε ∧ Ωε ) .
2π 8π
Thus, as required,
*  12 +
iΩθ /4π
 ε  
iΩ /2π
Tr e ∧ det , νh (0)
sinh(iΩθ /4π)
 
1 θ
 1 ε ε
= − (dim E) p1 Ω − 2 Tr(Ω ∧ Ω ) , νh (0)
24 8π
   
1 −1 θ θ
 1 ε ε
= − (dim E) Tr Ω ∧ Ω − 2 Tr(Ω ∧ Ω ) , νh (0)
24 8π 2 8π
1 1
dim E · Tr Ωθ ∧ Ωθ − 2 Tr(Ωε ∧ Ωε ) , νh (0)

=
16π 2 12
1
= 2 Str(h2 (0, 0)) . 
(4π)
Proof of the Local Index Formula for Arbitrary Even Dimensions.
Our starting point for the derivation of the Local Index Formula (17.102) for ar-
bitrary even dimensions is the determination of the kernel, say ea (x, y, t), for the
generalized 1-dimensional heat equation
ut = uyy − a2 y 2 u, u(y, t) ∈ R, (y, t) ∈ R×(0, ∞) ,
where a ∈ R is a given constant. If a = 0, we have the familiar result
1 − 1 (y−x)2
e0 (x, y, t) = √ e 4t .
4πt
For a 6= 0, we derive Mehler’s formula
(17.112) !
1 1 2 2
 
ea (x, y, t) = q exp − sinh(2at) cosh(2at) x + y − 2xy .
4π sinh(2at) 4 2a
2a

Note that we recover e0 (x, y, t) as a → 0. We know that ea (x, y, t) should be


symmetric in x and y, but not necessarily a function of y −x, since ut = uyy −a2 y 2 u
is not translation invariant. The simplest suitable form is
−1
ea (x, y, t) = f (t) 2 exp 41 g(t) x2 + y 2 + h(t) xy .
 

We have
1

(ea )y (x, y, t) = ea (x, y, t) 2 g(t) y + h(t) x
(ea )y (x, y, t) 12 g(t) y + h(t) x + ea 21 g(t)

(ea )yy (x, y, t) =
 2 
= ea (x, y, t) 12 g(t) y + h(t) x + 21 g(t) , and
17.5. THE LOCAL INDEX FORMULA 587

 
e−1
a (ea )t −(ea )yy + a2 y 2 ea (x, y, t)
f 0 (t) 1 0
= − 21 + 4 g (t) x2 + y 2 + h0 (t) xy

f (t)
 
2 2
− 41 g(t) y 2 + g(t) h(t) xy + h(t) x2 + 21 g(t) + a2 y 2
 0 
1 f (t)
 
2
= −2 + g(t) + 41 g 0 (t) − g(t) + 4a2 y 2
f (t)
 
2
+ 4 g (t) − h(t) x2 +(h0 (t) − g(t) h(t)) xy.
1 0

Equating to zero the coefficients of this quadratic polynomial in (x, y), we get

g 0 = g 2 − 4a2 =⇒ g(t) = −2a coth(2at + C)


a
q
2
p
h(t) = 21 g 0 (t) = 12 (2a) csch2 (2at + C) =
sinh(2at + C)
f0
 Z 
0
= −g =⇒ f (t) = C exp 2a coth(2at + C) dt
f
= C 0 exp(ln sinh(2at + C)) = C 0 sinh(2at + C) .

Luckily, we also have h0 (t) = g(t) h(t). Hence,

ea (x, y, t)
!
1 1 2 2
 
=p exp − sinh(2at) cosh(2at + C) x + y − 2xy
C 0 sinh(2at + C) 4 2a

The behavior is correct as a and t approach 0 if and only if C = 0 and C 0 = 2π/a,


and we have (17.112).
Now suppose that B is a real, skew-symmetric n × n matrix where n = 2m is
even. We will find the heat kernel for
(17.113)
2
ut (y, t) = ∆u − |By| u = ∆u + B 2 y, y u, u(y, t) ∈ R, (y, t) ∈ Rn ×(0, ∞) ,

where ∆u = ∂12 u + · · · + ∂n2 u. The eigenvalues of B are of the form ±ir1 , . . . , ±irm ,
and for some R ∈ SO(n),

R−1 B 2 R = − diag r12 , r12 , . . . , rm


2 2

, rm .

Let y 0 := R−1 y and u0 (y 0 , t) := u(y, t) = u(Ry 0 , t). If ut = ∆u + B 2 y, y u, then

u0t (y 0 , t) = ut (Ry 0 , t) = ∆u(Ry 0 , t) + B 2 Ry 0 , Ry 0 u(Ry 0 , t)


= ∆0 u0 (y 0 , t) + R−1 B 2 Ry 0 , y 0 u0 (y 0 , t)
Xm
= ∆0 u0 (y 0 , t) − 02 02
 0 0
(17.114) rj2 y2j−1 + y2j u (y , t) .
j=1
588 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

Let v1 = v1 (y1 , t) with v1t = ∂12 v1 − a21 v1 , and v2 = v2 (y2 , t) with v2t = ∂22 v2 − a22 v2 .
Then for v(y1 , y2 , t) := v1 (y1 , t) v2 (y2 , t), we have
vt (y1 , y2 , t) = v1t (y1 , t) v2 (y2 , t) + v1 (y1 , t) v2t (y2 , t)
= ∂12 v1 − a21 v1 v2 (y2 , t) + v1 (y1 , t) ∂22 v2 − a22 v2
 

= ∂12 v1 (y1 , t) v2 (y2 , t) − a21 v1 (y1 , t) v2 (y2 , t)




+ v1 (y1 , t) ∂22 v2 − v1 (y1 , t) a22 v2 (y2 , t)




= ∂12 + ∂22 v(y1 , y2 , t) − a21 + a22 v(y1 , y2 , t) .


 

Thus, the heat kernel for (17.114) is


m
Y
e0 (x0 , y 0 , t) := erj x02j−1 , y2j−1
0
, t erj x02j , y2j
0
 
,t .
j=1

The heat kernel for (17.113) is then


eB (x, y, t) = e0 (x0 , y 0 , t) = e0 R−1 x, R−1 y, t , and


m
−m
Y 2rj
eB (0, 0, t) = e0 R−1 0, R−1 0, t = (4π)

(17.115) .
j=1
sinh(2rj t)

Let AB : C ∞ (Rn ) → C ∞ (Rn ) denote the operator defined by


2
AB [f ](y) := ∆u(y) − |By| f (y) , and let
AB B
s [f ](y) := A [f ](sy) .

−n/2 
With E(r, t) := (4πt) exp −r2 /4t , just as before (only easier) we have an
asymptotic expansion (as t → 0+ )
−1
XQ
E(|y| , t) eB (0, y, t) ∼ hB j
j (0, y) t , where
j=0
Z Y
j−1 i
  
hB
j (0, y) = (si ) AB B
sj−1 ◦ · · · ◦ As0 [1] (y) ds0 . . . dsj−1 .
Im i=0

In particular,
−1
XQ
E(0, t) eB (0, 0, t) ∼ hB j
j (0, 0) t ,
j=0

and in view of (17.115) and n = 2m,


m m
−1 n/2 −m
Y 2rj Y 2rj t
E(0, t) eB (0, 0, t) = (4πt) (4π) = ,
j=1
sinh(2rj t) j=1 sinh(2rj t)

which is analytic at t = 0. Thus for t sufficiently small,


m
Y 2rj t X∞
= hB j
j (0, 0) t .
j=1
sinh(2r j t) j=0

Since the eigenvalues of iB are ±r1 , . . . , ±rm , we have


m   12
Y 2rj t 2tiB
= det .
j=1
sinh(2rj t) sinh(2tiB)
17.5. THE LOCAL INDEX FORMULA 589

Recall (see (17.42)) that


m ∞ ∞
Y rk s/2 X  2k X
= ak r12 , . . . , rm
2
s = Ak (σ1 , . . . , σk ) s2k ,
sinh(rk s/2)
k=1 k=0 k=0


where the coefficient ak r12 , . . . , rm
2
is a homogeneous,
 symmetric polynomial in
r12 , . . . , rm
2
of degree k. We write ak r12 , . . . , rm
2
as a polynomial
 Ak (σ1 , . . . , σk ) in
the elementary symmetric polynomials σk = σk r12 , . . . , rm 2
. With s/2 = 2t so that
s = 4t,
  12 m
2tiB Y 2rj t
det =
sinh(2tiB) j=1
sinh(2r j t)

X ∞
X
2k
= Ak (σ1 , . . . , σk )(4t) = 42k Ak (σ1 , . . . , σk ) t2k .
k=0 k=0

Thus, hB
j (0, 0) = 0 for j odd, and with

Ak (B) := Ak (σ1 , . . . , σk ) , where


1 X j1 ···j2l i1
σl := δi1 ···i2l B j1 · · · B ij2l2l ,
(2l)!
(i),(j)

we have

42k Ak (B) = hB
2k (0, 0)
Z Y
2k−1 i
 
= (si ) AB s2k−1 ◦ · · · ◦ A B
s0 [1](0) ds0 . . . ds2k−1 .
I 2k i=0

Let C ∈ R. If u satisfies (17.113), and v(y, t) := eCt u(y, t), then v satisfies

2
(17.116) vt (y, t) = ∆v − |By| v + Cv.

Thus, the kernel for the heat equation (17.116) is

eB,C (x, y, t) := eCt eB (x, y, t) ,

since this solution has the correct behavior as t → 0+ . With


2
AB,C [f ](y) := ∆f (y) − |By| f (y) + Cf (y) ,
AB,C
s [f ](y) := AB,C [f ](sy) , and

Aj (B, C) := hB,C
j (0, 0)
Z Y
j−1 i
 
:= (si ) AB,C B,C
sj−1 ◦ · · · ◦ As0 [1](0) ds0 . . . dsj−1 ,
Ij i=0
590 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

we then have
X∞ B,C −1
hl (0, 0) tl = E(|y| , t) eB,C (0, 0, t)
l=0
m
Y 2rj t X∞ X∞
= eCt = eCt hB j
j (0, 0) t = e
Ct
42k Ak (B) t2k
j=1
sinh(2r j t) j=0 k=0

X∞ C j tj X∞ X∞ 1 j 2k
= Ak (B) t2k = C 4 Ak (B) tj+2k
j=0 j! k=0 j,k=0 j!
X∞  X 1 j 2k

= C 4 Ak (B) tl , or
l=0 j+2k=l j!

(17.117)
Z Y
l−1 i
  X 1 j 2k
(si ) AB,C B,C
sl−1 ◦ · · · ◦ As0 [1](0) ds0 . . . dsl−1 = C 4 Ak (B) .
Il i=0 j+2k=l j!
Observe that both sides are ultimately homogeneous polynomials of degree l in the
variables C and Bij . Suppose that we replace these variables by operators C
b and
B
bij in End(W ) for a finite dimensional vector space W , specifically

W := CN ⊗ Σ2m ,
X
b := − 1
C 2 Ωεpq (0) ⊗ γ p γ q , and
p,q
X
bpj := 1 Rklpj (0) IdCN ⊗iγ k γ l .

B 8 k,l

b in AB,C yields an operator AB, : C ∞ (Rn , W ) ←-,


bC
Replacing B by Bb and C by C b

namely
bpj 0 y j 0 ψ + Cψ
X
AB,C (ψ) := ∆ψ − bpj y j B
B
b b b
p,j,j 0
X  X  X 
k 0 l0 j 0
X
= ∂i2 + 1
8 Rklpj (0) γ k γ l y j 18 Rk 0 l0 pj 0 (0) γ γ y
i p k,l,j k0 ,l0 ,j 0
X
1 ε i j
−2 Ωij (0) ⊗ γ γ ,
i,j

which is the same as the operator A0 in (17.108); note the presence of the factor

of i = −1 in the definition of B bpj which accounts for a crucial sign change in
B,
bC 0
order that A = A . The right side of (17.117) still makes sense as an element
b

of End(W ) if multiplication is taken to be composition. But, since generally op-


erators do not commute, one could get a different result unless some ordering in
compositions of the factors C b and B bij in monomials is specified. With operator
replacement, the left side may also be interpreted as in End(W ), if for a given
w ∈ W , we define
Z Y 
l−1 i
 
hB,
bC B,
bC B,
bC
(0, 0) := (s ) A ◦ · · · ◦ A [1](0) ds . . . ds l−1 (w)
b b b
l i sl−1 s0 0
Il i=0
Z Y 
l−1 i
 
(si ) AB,
bC B,
bC
(17.118) = ◦ · · · ◦ A [ w](0) ds . . . dsl−1 ,
b b
sl−1 s0 e 0
Il i=0

where w e is the constant function in C ∞ (Rn , W ) with value w; note that the 1 in
[1] of (17.118) may be regarded as the map w 7→ w. e There is already a definite
17.5. THE LOCAL INDEX FORMULA 591

ordering of the factors in the various terms of the expansion of


Z Y
l−1 i
 
(si ) AsB,l−1
bC
◦ · · · ◦ AsB,0 C [1](0) ds0 . . . dsl−1 .
b b b

Il i=0

Since we know that the two sides of (17.117) agree as polynomials in commuting
variables C and Bij , if we replace these variables by operators, there is a corre-
sponding reordering of operators within the terms of the right side so that both
sides are the same element of End(W ). In other words,
Z Y
l−1 i
 
(si ) AB,
bC B,
bC
sl−1 ◦ · · · ◦ As0 [1](0) ds0 . . . dsl−1
b b

Il i=0
X 42k b j Ak (B)),
(17.119) = R(C b
j+2k=l j!
where R(C b j Ak (B))
b is a reordering of the operators C b and B bij in the monomial
j
terms of C Ak (B) so that (17.119) holds. Now,
b b
 X   X 
b◦B bpj = − 1 ε q r 1 k l
C 2 Ω qr (0) ⊗ γ γ ◦ 8 R klpj (0) IdC N ⊗iγ γ
p,r k,l
X
1 ε q r k l
= − 16 Rklpj (0) Ωqr (0) ⊗ iγ γ γ γ
q,r,k,l
X
1
= − 16 Rklpj (0) Ωεqr (0) ⊗ iγ k γ l γ q γ r + Oc (2)
q,r,k,l
bpj ◦ C
=B b + Oc (2) .
Thus, for j + 2k = m = n/2,
   
b j Ak (B))
Str R(C b = Str Cb j R(Ak (B))
b ,

where R(Ak (B))


b is Ak (B)b with a possible reordering within the monomial terms
of Ak (B). However, these monomial terms are products of the operators
b
X
bpj = 1
B 8 Rklpj (0) iγ k γ l ,
k,l

and since the Rklij (0) are just scalars, such operators commute modulo terms of
degc 2. Hence, for j + 2k = m = n/2,
     
Str R(C b j Ak (B))
b = Str Cb j R(Ak (B))
b = Str Cb j Ak (B))
b .
Thus, while a reordering R is needed to make (17.119) correct, when l = m = n/2,
we have
42k
  X 
B,
bC j
Str hm (0, 0) = Str R(C Ak (B))
b b b
j+2k=m j!

42k b j
X 
(17.120) = Str C Ak (B) .
b
j+2k=m j!

With ν := dy 1 ∧ · · · ∧ dy n and n = 2m,


m
Str γ j1 γ j2 · · · γ j2k−1 γ j2k = (−2i) dy j1 ∧ dy j2 ∧ · · · ∧ dy j2k−1 ∧ dy j2k , ν
  

= −2idy j1 ∧ dy j2 ∧ · · · ∧ −2idy j2k−1 ∧ dy j2k , ν ,


 

where both sides are 0 if k < m. Thus, in (17.120) we may replace


X X
Cb = −1
2 Ωεpq ⊗ γ p γ q by 2iΩε = − 21 Ωεpq (−2idy p ∧ dy q ) ,
p,q p,q
592 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

and
X X
1
Rklpj iγ k γ l 1 θ
Rklpj i −2idy k ∧ dy l
1

B
bpj =
8 by 2 Ωpj = 8
k,l k,l

provided we take the inner product, with ν, of the trace (in End CN ) of the result.
The following computation then completes the proof of the Local Index Theorem
(i.e., Theorem 17.49, p. 577):
−m −m B,C
(4π) Str(hm (0, 0)) = (4π) Str(hm (0, 0))
b b

 
−m
X 1 b j 2k
= Str (4π) (C) 4 Ak (B) b
j+2k=m j!
* +
X 1 j 2i ε 42k 1 θ
= ∧ ( Ω )∧ A ( Ω ), ν
2k k 2
j+2k=m j! 4π (4π)
X 
∞ 1  X∞
j i ε 1 θ
= Tr ∧ ( 2π Ω ∧ Ak ( 2π Ω ), ν
j=0 j! k=0
*  θ
 21 +
 ε  iΩ /4π
= Tr eiΩ /2π ∧ det ,ν .
sinh(iΩθ /4π)

Index Theorem for Twisted Dirac Operators and the A b genus. As a


consequence of the Local Index Theorem for twisted Dirac Operator we have the
integrated version itself, namely
Theorem 17.53 (Index Theorem for Twisted Dirac Operators). Let M be a
compact, oriented Riemannian 2m-manifold with a spin structure, and let E be a
Hermitian vector bundle over M with unitary connection ε ∈ C(U (E)). Then the
index of the Dirac operator
D+ : C ∞ E ⊗ Σ+ (M ) → C ∞ E ⊗ Σ− (M )
 

is given by  
index D+ = ch(E) ` A(M
b ) [M ] ,
where ch(E) ∈ H ∗ (M ; Q) denotes the Chern character of E (see (15.98) or (15.112))
and A(M
b ) ∈ H ∗ (M ; Q) denotes the total A
b class of M (see (15.109) with F = T M
and A(M ) := A(T M ) or the equivalent formulation (17.44)).
b b
Proof. We have already proven this implicitly in the discussion preceding
(17.45) on p. 546, but there we assumed the Local Index Theorem (p. 577). As this
was done many pages ago, it is fitting to recall the computation, especially now
that we have finally established the Local Index Theorem:
Z
+ (17.38)
Tr k + (x, x, t) − Tr k − (x, x, t) νx
  
index D =
M
Z Z
= Str(k(x, x, t)) νx = lim Str(k(x, x, t)) νx
+
M M t→0
 12
iΩθ /4π
Z D    E
iΩε /2π
= Tr e ∧ det , νx νx by Theorem 17.49
M sinh(iΩθ /4π)
 
= ch(E) ` A(M
b ) [M ] ,
where the final equality was explained on p. 546. 
17.5. THE LOCAL INDEX FORMULA 593

Corollary 17.54. Let M be a compact, oriented Riemannian 2m-manifold


with a spin structure and positive scalar curvature, then A(M
b ) := A(M
b ) [M ] = 0.
Proof. Take E = 0. By Corollary 17.28, p. 537, Ker D+ = Ker D− = {0} so
b ) = index D+ = 0.
that A(M 

As an application, especially relevant to the study of intersection forms of


smooth 4-manifolds and Seiberg-Witten theory, we have
Theorem 17.55. If M is a compact, orientable C ∞ 4k-manifold with a spin
structure, then A(M
b ) := A(M
b ) [M ] is an integer. Moreover, if k is odd, then A(M
b )
is even.
Proof. By Theorem 17.53 with E trivial, A(M b ) = A(Mb ) [M ] = index D+
and so A(M
b ) is an integer. If k = 2a + 1 is odd, then according to the table (17.2)

and the periodicity C`n+8 −→ C`8 ⊗ C`n = R(16) ⊗ C`n , for the real C`4k we have
∼ ∼ ∼ ∼
C`4k −→ C`8a+4 −→ (⊗a R(16)) ⊗ C`4 −→ R 24a ⊗ H(2) −→ H 24a+1 .
 

Thus, C`4k can be identified with the algebra of 24a+1 × 24a+1 matrices with entries
4a+1
in H (the quaternions) which act via R-linear transformations to the left on H2 .
4a+1
Of course, there is right action of H on H2 via H-scalar multiplication, and this
 4a+1
right action commutes with the left action of H 24a+1 on H2 . If we define a
24a+1
complex on structure on H via right multiplication by a pure imaginary unit
4a+1
quaternion (e.g., −k), then H2 becomes a complex vector space of dimension
 4a+1
24a+2 . The left action of H 24a+1 on H2 is C-linear, and the complexified
∼ 4a+1
4a+2

Cl4k is the full algebra (= C 2 ) of C-linear transformations of H2 . The
4a+1
complex spinor space Σ4k is then H2 with its complex structure. We still have
the R-linear right action of H on H4a+1 , and although this not C-linear, this right
action commutes with left multiplication by the volume element ωC = i2k e1 · · · e4k =
k 4a+1

(−1) e1 · · · e4k , since ωC is in the real C`4k = H 2 . Thus, the spinor spaces
Σ±4k := (1 ± ω C ) Σ4k are invariant under the R-linear right action of H on Σ4k ∼ =
4a+1 ±
H2 . In other words Σ4k are right H-modules over R. The spinor bundles
Σ± ±
4k (M ) := P ×Spin(4k) Σ4k are also right H-modules since the right H-action on
Σ± ± ±

4k commutes with the spinor representations ρ : Spin(4k) → SU Σ 4k which are
just restrictions of Clifford multiplication to Spin(4k) ⊂ C`4k . For the same reason,
the H-action on Σ± 4k (M ) commutes with

Clifford multiplication c : Ω1 (M ) ⊗ C ∞ (Σ(M )) → C ∞ (Σ(M )) and


Σ ∞ 1 ∞
covariant differentiation ∇ : C (Σ(M )) → Ω (M ) ⊗ C (Σ(M )) ,

and hence with the Dirac operator D = c ◦∇Σ . Thus, the spaces
Ker D± : C ∞ Σ± (M ) → C ∞ Σ∓ (M )
 

are right H-modules. The real dimension of any H-module V is a multiple of 4 (see
Exercise 17.57 below). Hence,

A(M
b ) = A(M
b ) [M ] = index D+ = dimC Ker D+ − dimC Ker D− is even. 
594 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

Corollary 17.56 (V. Rokhlin’s Theorem, 1952). The signature sig(M ) of


a compact, orientable C ∞ 4-manifold M with a spin structure is a multiple of 16.
Proof via Index Theory. The Hirzebruch Signature Theorem (to be proven
independently in Section 17.6; see Theorem 17.64, p. 605) states that sig(M ) =
1 1
3 p1 (M ). Since A(M ) = − 24 p1 (M ) by 15.110, we obtain sig(M ) = −8A(M ) ,
b b
which is a multiple of 16 by Theorem 17.55. 
Exercise 17.57. Let φ : H → EndR (V ) be a real right representation (i.e.,
φ(qq 0 ) = φ(q 0 ) φ(q)) with dim V < ∞. By completing parts (a), (b) and (c), show
that V = V1 ⊕· · ·⊕Vh where the Vi are invariant subspaces on which φ is equivalent
to the right action of H on itself.
a. Let r : H → EndR (H) be given by r(q)(q 0 ) = q 0 q. Show that r is irreducible;
i.e., r(W ) ⊆ W for some subspace W ⊆ H implies that W is {0} or H.
[Hint. H is a division algebra.]
b. For any nonzero v ∈ V , show that φv : H → End(φ(H) v) given by φv (q) =
φ(q) |φ(H)v is equivalent to r. [Hint. For fv : H → φ(H) v given by fv (q) =
φ(q) v, show first that Ker fv is r-invariant.]
c. Show that V = φ(H) v1 ⊕ · · · ⊕ φ(H) vh for some v1 , . . . , vh ∈ V .

6. The Index Theorem for Standard Geometric Operators


I Our goal here is to obtain index formulas for the standard elliptic geometric op-
erators and their twists. As listed in the table of standard elliptic operators in Table 6.1
(p.187) and summarized in our first survey of the applications of the Index Theorem in
Chapter 13 (pp.310ff), the standard elliptic geometric operators include
the signature operator d + δ : (1 + ∗) Ω• (M ) → (1 − ∗) Ω• (M ) ,
the Euler-Dirac operator d + δ : Ωev (M ) → Ωodd (M ) , and

the Dolbeault-Dirac operator 2 ∂¯ + ∂¯∗ : Ω0,ev (M ) → Ω0,odd (M ) .


Here, d denotes the exterior derivative, δ denotes the exterior coderivative (i.e., the formal
adjoint d∗ of d), ∗ denotes the Hodge star operator, and ∂¯ and its adjoint ∂¯∗ denote the
complex analogs of d and d∗ on complex manifolds, which, along with Ω0,∗ (M ), will be
defined. The index formula obtained for the above operators yields, as sketched in Chapter
13 (pp.310ff),
the Hirzebruch Signature Theorem, sig(M ) = L(M )[M ], Theorem 17.64 (p.605),
Z
the Chern-Gauss-Bonnet Theorem, χ(M ) = GB(Ωθ ), Theorem 17.68 (p.612), and
M

the Hirzebruch-Riemann-Roch Theorem, Index(∂ + ∂ ) = Td(T M )[M ], Theorem 17.77
(p.637),
respectively. While these operators generally are not globally twisted Dirac operators,
locally they are expressible in these terms. Thus, even if the underlying Riemannian
manifold M (still assumed to be oriented and of even dimension n = 2m) does not admit
a spin structure, we may still use the Local Index Theorem for twisted Dirac operators to
compute the index density and hence the index of these operators. While it is possible to
carry this out separately for each of the geometric operators, basically all of these theorems
are consequences of an index theorem for generalized Dirac operators on Clifford module
bundles (defined below). Using the Local Index Theorem for twisted Dirac operators,
we prove this index theorem first, and then we apply it to obtain the geometric index
theorems. J
17.6. THE INDEX THEOREM FOR STANDARD GEOMETRIC OPERATORS 595

An Index Theorem for Generalized Dirac Operators. For a compact,


oriented manifold M of even dimension n = 2m with Riemannian metric h, let
Cl(Tx M ) denote theS complexified Clifford algebra of Tx M with inner product hx .
Then Cl(T M ) := x∈M Cl(Tx M ) is the total space of the so-called complex Clif-
ford bundle Cl(T M ) → M . This bundle of algebras is defined whether M admits
a spin structure or not. As a complex vector bundle, it is canonically isomorphic
to Λ• (M, C), but the algebra structure is different. We can also describe Cl(T M )
as follows. Let c : Spin(n) → SO(n) denote the covering
 given by c(g)(v) = gvg −1 ,
n n 1 n
where v ∈ R and we identify R with L Λ (R ) ⊂ Cln , where L denotes the
linear isomorphism of (17.3), p.514. Define the representation
r : SO(n) → End(Cln ) by
r(c(g))(α) = gαg −1 for g ∈ Spin(n) .
This is simply the extension to Cln of the defining representation of SO(n) on

Rn −→ L Λ1 (Rn ) ⊂ Cln . Relative to r, Cl(T M ) is then the associated bundle
F M ×SO(n) Cln . Note that Clifford multiplication is SO(n)-equivariant, in the sense
that for A ∈ SO(n), r(A) ∈ End(Cln ) is an algebra automorphism of Cln . Thus,
regarding Cl(T M ) as F M ×SO(n) Cln , a well-defined algebra structure on each fiber
is given by [p, α1 ] [p, α2 ] := [p, α1 α2 ], since
pA, r A−1 (α1 ) r A−1 (α2 ) = pA, r A−1 (α1 α2 ) = [p, α1 α2 ] .
      

Let Q : Cln → End(W ) be an algebra representation where W is a complex


vector space with dim(W ) < ∞. There is a Hermitian inner product h·, ·i on W ,
such that for any unit vector v ∈ S n ⊂ Rn , we have Q(v) ∈ U (W ) := unitary group
0
for W with h·, ·i. Indeed, for any Hermitian inner product h·, ·i on W , let
Z
0
hw1 , w2 i := hQ(v) w1 , Q(v) w2 i dv.
Sn
2 ∗
Using Q(v) = −I for v ∈ S n , we deduce that Q(v) = −Q(v) from
D E
2
hQ(v) w1 , w2 i = Q(v) w1 , Q(v) w2 = h−w1 , Q(v) w2 i .

Definition 17.58. Let q : Spin(n) → U (W ) be a representation such that


(17.121) q(g) ◦ Q(α) ◦ q g −1 = Q(r(c(g)) α) for all g ∈ Spin(n) , α ∈ Cln .


If M admits a spin structure C : P → F M , the representation q provides us with


an associated bundle
W (M ) := P ×Spin(n) W.
In the case q : Spin(n) → U (W ) is of the form q0 ◦ c, where c : Spin(n) → SO(n)
denotes the double cover and q0 : SO(n) → U (W ) is a representation, then we may
form a bundle
W (M ) := F M ×SO(n) W ,
whether M admits a spin structure or not. In either case, we call W (M ) a Clifford
module bundle.
We consider some examples. Note that Cl(T M ) ∼ = Λ• (M, C) is a Clifford
module bundle. Indeed, let W = Cln and Q : Cln → End(Cln ) be left multiplication
596 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

of Cln on itself, and let q : Spin(n) → U(Cln ) be given by q(g)(β) = gβg −1 . Then
as required
q(g) ◦ Q(α) ◦ q g −1 (β) = g α g −1 βg g −1
 

= gαg −1 β = Q gαg −1 β = Q(r(c(g)) α)(β) .




If M admits a spin structure, the primary example is the spin bundle Σ(M ). Here
W = Σn , and Q = ρC : Cln → End(Σn ) denotes the unique irreducible representa-
tion, which restricts to q = ρ : Spin(n) → End(Σn ). Then
q(g) ◦ Q(α) ◦ q g −1 (β) = ρ(g) ◦ ρC (α) ◦ ρ g −1 (β)
 

= ρC gαg −1 (β) = ρC (r(c(g)) α)(β) = Q(r(c(g)) α)(β) .




Note that when M admits a spin structure, P → F M , one can also consider the
bundle P ×Spin(n) W relative to the representation Q|Spin(n) : Spin(n) → U (W ),
but since it may not be the case that q = Q|Spin(n) , this bundle is not necessarily
isomorphic to W (M ) obtained from q. For example when Q : Cln → End(Cln ) is
left multiplication, the bundle P ×Spin(n) W relative to Q|Spin(n) is not Cl(T M ), but
rather it is designated by ClSpin (T M ) which is actually isomorphic to a direct sum
of 2m copies of Σ(M ) , m = 21 dim M .
To show that a Clifford module bundle W (M ) is indeed a bundle of Cl(T M )x -
modules, we let
c : Cl(T M ) ⊗ W (M ) → W (M ) be defined via
c([p, α] ⊗ [p, w]) := [p, Q(α) w] .
which is a well-defined Clifford multiplication. Indeed, using (17.121),
h i h i h   i
−1 −1 −1
c pA, r(A) α ⊗ pA, q(A) w = pA, Q r A−1 α q(A) w
h i 
−1
= pA, q A−1 Q(α) q(A) q(A) (w) = pA, q A−1 Q(α)(w)
  

= [p, Q(α)(w)] = c([p, α] ⊗ [p, w]) .



Moreover, for α ∈ L Λ1 (Rn ) ⊂ Cln , we have
h i
2 2
c([p, α] ⊗ c([p, α] ⊗ [p, w])) = p, Q(α) w = − |[p, α]| [p, w] ,

so that c does in fact make W (M )x a Cl(T M )x -module. As in the special case of


Σ(M ), there is a decomposition W (M ) = W + (M ) ⊕ W − (M ) defined as follows. If
E1 , . . . , En is a positively oriented orthonormal frame at x ∈ M , let
(17.122) ωC (x) := im E1 · · · En ∈ Cl(T M )x .
2
It is easy to check that this is independent of the choice of frame and ωC (x) = 1
(see (17.13), p. 520). Then
2
µC (x) := c(ωC (x)) ∈ End(W (M )x ) and µC (x) = Id, and
(17.123) W ± (M ) := (I ± µC )W (M ) =⇒ W (M ) = W + (M ) ⊕ W − (M ) .
For any 0 6= X ∈ Tx M, we may take E1 = X/ |X|. Since c(E1 ) is an isomorphism
of Wx (M ) which anticommutes with µC (x), we have

(17.124) c(X) = |X| c(E1 ) : W ± (M ) −→ W ∓ (M ) .
17.6. THE INDEX THEOREM FOR STANDARD GEOMETRIC OPERATORS 597

Since W (M ) is an associated bundle of F M or the Spin(n)-bundle P of a spin


structure P → F M , the Levi-Civita connection θ on F M (or its lift to P ) yields a
covariant derivative operator

∇W : C ∞ (W (M )) → Ω1 (W (M )) −→ C ∞ Ω1 (M ) ⊗ W (M ) .


Moreover, c : Cl(T M ) ⊗ W (M ) → W (M ) induces map on sections,


c : C ∞ (Cl(T M )) ⊗ C ∞ (W (M )) → C ∞ (W (M )) (same notation).
For α ∈ C ∞ (Cl(T M )) and ψ ∈ C ∞ (W (M )), we have
∇W (c(α ⊗ ψ)) = c(∇α ⊗ ψ) + c α ⊗ ∇W ψ ,

(17.125)
where ∇ denotes the usual Levi-Civita covariant derivative on C ∞ (Cl(T M )) ∼
=
Ω• (M, C). Indeed, using the infinitesimal version of (17.121), namely
q 0 (A) Q(α) − Q(α) q(A0 ) = Q(r0 (A) α) ,
0
it follows that on F M where α⊗ψ ∈ Ω (Cl2m ⊗ W ) := equivariant Cl2m ⊗W -valued
functions on F M ,
Dθ (Q(α ⊗ ψ)) = d(Q(α ⊗ ψ)) + q 0 (θ) Q(α ⊗ ψ)
= Q(dα ⊗ ψ) + Q(α ⊗ dψ) + Q((r0 (θ) α) ⊗ ψ + α ⊗ q 0 (θ) ψ)
= Q(dα ⊗ ψ) + Q(r0 (θ)(α) ⊗ ψ) + Q(α ⊗ dψ) + Q(α ⊗ q 0 (θ) ψ)
= Q Dθ (α) ⊗ ψ + Q α ⊗ Dθ ψ .
 

Via the Riemannian metric, Λ1 (Tx∗ M ) ∼


= Λ1 (Tx M ) ⊂ Cl(T M ) , and so
Ω1 (M ) ⊗ C ∞ (W (M )) ⊂ C ∞ (Cl(T M )) ⊗ C ∞ (W (M ))
Thus, we have an operator DW := c ◦∇W ∈ End(C ∞ (W (M ))),
∇W c
DW : C ∞ (W (M )) → Ω1 (M ) ⊗ C ∞ (W (M )) → C ∞ (W (M )) ,
which we call a generalized Dirac operator. In view of (17.124), we also have
DW ± : C ∞ W ± (M ) → C ∞ W ∓ (M ) ,
 

as with Dirac operators. We now show that DW is locally a twisted Dirac operator.
Indeed, we show that DW is globally a twisted Dirac operator when M admits a spin
structure P → F M , from which the local statement follows since locally M admits
spin structures. Recall (see Proposition 17.15, p. 523) that for any Cl2m -module
W , we have a Cl2m -equivariant linear isomorphism

Φ : Hom0 (Σ2m , W ) ⊗ Σ2m −→ W via Φ(φ ⊗ ψ) := φ(ψ) ,
where Hom0 (Σ2m , W ) consists of the Cl2m -equivariant linear maps Σ2m → W . If
M admits a spin structure P → F M , then Φ yields an isomorphism

ΦM : W (M ) −→ E0 ⊗ Σ(M ) = End0 (Σ(M ) , W (M )) ⊗ Σ(M ) , where
E0 := End0 (Σ(M ) , W (M )) = P ×Spin(n) End0 (Σ2m , W ) .
Since φ ∈ Hom0 (Σ2m , W ) , we have
Φ(φ ⊗ ρ(ωC ) ψ) = φ(ρ(ωC ) ψ) = c(ωC ) φ(ψ) = µC φ(ψ) , and so
 
± ±
(17.126) ΦM W (M ) = E0 ⊗ Σ(M ) .
598 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

As both C ∞ (Σ(M )) and C ∞ (W (M )) have covariant differentiation operators aris-


∗
ing ultimately from the Levi-Civita connection, C ∞ W (M ) ⊗ Σ(M ) also has
∗ ∗
such an operator, say ∇W ⊗Σ . Now E0 is a subbundle of W (M ) ⊗ Σ(M ) , and

∇W ⊗Σ (C ∞ (E0 )) ⊂ Ω1 (E0 ) as a consequence of (17.125). Indeed,
φ ∈ C ∞ (E0 ) =⇒ cW (v) φ(σ) = φ(cΣ (v) σ)
=⇒ ∇X (cW (v) φ(σ)) = ∇X (φ(cΣ (v) σ))
 
⊗Σ∗
=⇒ cW (∇X v)(φ(σ)) + cW (v) ∇W X φ (σ) + cW (v)(φ(∇X σ))
 
⊗Σ∗
= ∇W X φ (cΣ (v) σ) + φ(cΣ (∇X v) σ) + φ(cΣ (v) ∇X σ)
   
⊗Σ∗ W ⊗Σ∗
=⇒ cW (v) ∇WX φ (σ) = ∇ X φ (cΣ (v) σ)
=⇒ ∇X φ ∈ C ∞ (E0 ) .

For ∇E0 := ∇W ⊗Σ |C ∞(E0 ) , we then have the twisted Dirac operator

DE0 := (1 ⊗ cΣ ) ◦ ∇ : C ∞ (E0 ⊗ Σ(M )) → C ∞ (E0 ⊗ Σ(M )) , where


∇ := ∇E0 ⊗ 1 + 1 ⊗ ∇Σ : C ∞ (E0 ⊗ Σ(M )) → Ω1 (M ) ⊗ C ∞ (E0 ⊗ Σ(M )) .

For ΦM : C ∞ (W (M )) −→ C ∞ (E0 ⊗ Σ(M )) corresponding to W (M ) ∼
= E0 ⊗Σ(M ),
we then have
DW = Φ−1
M ◦D
E0
◦ ΦM .
Since ΦM is canonical, this shows that DW is canonically equivalent to twisted
Dirac operator in the case M has a spin structure, and in the general case DW is
locally such. If M admits a spin structure, then the Index Theorem for twisted
Dirac operators yields

index DW = (ch(E0 ) ` A(M



b )) [M ] .
However, we will see that

(17.127) ch(E0 ) ` A(M


b ) = ch(W (M )) ` A(M
e ),

where A(M
e ) denotes the total characteristic class defined by
 
y/2
(17.128) A(M
e ) := MC sinh(y) ,TM ,
 
y/2
which is almost the same as A(M
b ) = MC sinh(y/2) , T M . A key observation is
that the right side of (17.127) is defined even if M does not admit a spin structure.
Shortly we will see that the standard total forms that represent the total classes
ch(E0 ) ` A(M
b ) and ch(W (M )) ` A(M e ) are identical. The local index density
W
for D is computed using a local spin structure about a point. However, the
result of that local computation is identical to the restriction of a globally-defined
form which represents ch(W (M )) ` A(M e ). In essence, the local spin structure
is a computational aid, while the index density itself can be expressed without
reference to spin structures. Thus, once (17.127) is shown on the level of forms, we
have an index formula for the generalized Dirac operator DW + : C ∞ (W + (M )) →
C ∞ (W − (M )), namely
17.6. THE INDEX THEOREM FOR STANDARD GEOMETRIC OPERATORS 599

Theorem 17.59 (Index Theorem for Generalized Dirac Operators). For a Clif-
ford module bundle W (M ) over an oriented, closed (i.e., compact and without
boundary) Riemannian manifold M , we have
  
(17.129) index DW + = ch(W (M )) ` A(M e ) [M ] ,

where A(M
e ) is defined in (17.128).
Proof. For a normal coordinate ball V ⊂ M , let C : PSpin → F M |V be a
spin structure. The Levi-Civita connection θ on F M lifts to a unique connec-
tion, say θe := c0−1 (C ∗ (θ)) ∈ Ω1 (P, spin(n)). Note that Σ(V ) = P ×Spin(n) Σn .
Let R : PSpin → U(Σ(V )) denote the morphism determined by the representation
ρ : Spin(n) → U(Σ(V )), as in (15.118), p. 454. To compute the Chern forms for
Σ(V ) relative to the connection ω on U(Σ(V )) , such ρ0 ◦ θe = R∗ ω, we make use of
the result (see (15.119), p.455)
 
e ρ0
cj (Σ(V ) , ω) = cj Σ(V ) , θ,

Hence we need to determine the eigenvalues of ρ0 (b) for b ∈ spin(n). Any b ∈


spin(n), can be written in the form
Xm
b = bλ := 21 λj e2j−1 e2j
j=1

for some oriented, orthonormal basis e1 , . . . , en of Rn . Recall that Proposition 17.11


(p.519) provides us with an explicit representation
ρ : C`2m → End(Λ• (Cm )) .
If (f1 , · · · , fm ) is an orthonormal basis of Cm , then
(e1 , · · · , e2m ) := (f1 , if1 , · · · , fm , ifm )
is an oriented orthonormal basis of Rn . By (17.15), p.521, we have

ρ(e2j−1 e2j ) fj1 ∧ fj2 ∧ · · · ∧ fjp
 
i fj1 ∧ fj2 ∧ · · · ∧ fjp , if jk = j for some k,
=
−i fj1 ∧ fj2 ∧ · · · ∧ fjp , if jk 6= j for all k.
 P 
m
Hence, fj1 ∧fj2 ∧· · ·∧fjp is an eigenvector of ρ 12 j=1 λj e2j−1 e2j with eigenvalue
X X 
i
2 λj − λj .
j∈{j1 ,...,jp } j ∈{j
/ 1 ,...,jp }

Thus, in order to find the total Chern character form ch(Σ(V ) , ω), we compute
1 P
 
Xm X P
λj − j ∈ λj
j∈{j1 ,...,jp } / {j 1 ,...,jp }
p
e2
p=1 (j)
Ym   Ym
eλk /2 + e−iλk /2 = 2 cosh 12 λk .

=
k=1 k=1
 
and replace each σk λ21 , . . . , λ2k by pk Ωθ . Hence chk (Σ(V ) ,ω), which is de-
fined on V , coincides with polynomial in Pontryagin forms pj Ωθ that are defined
throughout M . Note also that since W (M ) := F M ×SO(n) W is associated to F M,
600 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

ch(W ) is expressible in terms of theglobally-defined curvature form Ωθ . Thus, the



same is true for ch W (M ) ⊗ Σ(V ) and ch(E0 ). Since
1 1
2 λi  2λi
Ym Ym Ym
1

= 2 cosh 2 λi
k=1 sinh 1 λi 1 1

2
k=1 2 sinh 2 λicosh 2 λi
k=1
1
2 λi
Ym Ym
1

= 2 cosh 2 λi ,
k=1 sinh(λi ) k=1

we have the following equality of forms

A(V,
b θ) = A(V,
e θ) ∧ ch(Σ(V ) , θ) .
The desired equality (17.127) on the level of forms then follows from
 
ch(E0 |V , θ) ∧ A(V,
b θ) = ch(E0 |V , θ) ∧ A(V,
e θ) ∧ ch(Σ(V ) , θ)

= (ch(E0 |V , θ) ∧ ch(Σ(V ) , θ)) ∧ A(V,


e θ)
(17.130) = ch(E0 |V ⊗ Σ(V ) , θ) ∧ A(V,
e θ) = ch(W (M ) , θ) ∧ A(V,
e θ) .

In view of the preliminary discussion, we have shown (17.129). 

Twisted Generalized Dirac Operators. Recall that W (M ) is an associated


bundle of F M or a covering P . Thus, Theorem 17.59 does not yet include the case
of arbitrary twisting, say by a Hermitian vector bundle E → M with a covariant
derivative ∇E arising from a unitary connection 1-form ε on U(E). To handle this,
we proceed as follows. If we let Cl(T M ) act trivially on E, we have
IdE ⊗ c : C ∞ (Cl(T M )) ⊗ C ∞ (E ⊗ W (M )) → C ∞ (E ⊗ W (M ))
and can define a twisted generalized Dirac operator

∇E⊗W
DE,W : C ∞ (E ⊗ W (M )) −→ Ω1 (M ) ⊗ C ∞ (E ⊗ W (M ))
IdE ⊗ c
−→ C ∞ (E ⊗ W (M )) .
Moreover, we have a decomposition
   
+ −
E ⊗ W (M ) = E ⊗ W (M ) ⊕ E ⊗ W (M ) and an operator
   
+ −
DE,W + : C ∞ E ⊗ W (M ) −→ C ∞ E ⊗ W (M ) .

Theorem 17.60 (Index Theorem for Twisted Generalized Dirac Operators).


For a Clifford module bundle W (M ) over an oriented, compact Riemannian man-
ifold M and a Hermitian vector bundle E → M with a covariant derivative ∇E
arising from a unitary connection 1-form ε on U(E), we have
   
index DE,W + = ch(E) ` ch(W (M )) ` A(M e ) [M ]
 
(17.131) = ch(E ⊗ W (M )) ` A(M
e ) [M ] ,

where A(M
e ) is defined in (17.128).
17.6. THE INDEX THEOREM FOR STANDARD GEOMETRIC OPERATORS 601

Proof. We proceed as in the proof of Theorem 17.59, noting that now DE,W is
locally a twisted Dirac operator just as DW was, except that there is an additional
twist by E. In other words, while the generalized Dirac operator DW is locally the
twisted Dirac operator DE0 , the twisted generalized Dirac operator DE,W is locally
the twisted Dirac operator DE⊗E0 . The computation (17.130) generalizes to
ch(E ⊗ E0 |V , ε ⊗ θ) ∧ A(V,
b θ)
 
= ch(E ⊗ E0 |V , ε ⊗ θ) ∧ A(V,
e θ) ∧ ch(Σ(V ) , θ)

= ch(E|V , ε) ∧(ch(E0 |V , θ) ∧ ch(Σ(V ) , θ)) ∧ A(V,


e θ)
= ch(E|V , ε) ∧ ch(E0 |V ⊗ Σ(V ) , θ) ` A(V,
e θ)
(17.132) = ch(E|V , ε) ∧ ch(W (M ) , θ) ∧ A(V,
e θ) ,
which implies (17.131) as before. 

The Hirzebruch Signature Formula. The standard irreducible representa-


tion of Cl(n) is the C-linear extension ρC : Cl2m → End(Λ• (Cm )) of ρ : C`2m →
End(Λ• (Cm )) which is described as follows. Let ρ1 : Cm → End(Λ• (Cm )) be given
by
ρ1 (w)(α) := (w∧ − wx)(α) = w∧ α − wxα.
We think of Cm as R2m . By Proposition 17.11, p. 519, ρ1 uniquely extends to an
R-linear homomorphism
ρ : C`2m → End(Λ• (Cm )) ,
of algebras over R. Then ρC : Cl2m → End(Λ• (Cm )) is the C-linear extension
 of ρ.
There is also a highly reducible representation Q : Cl2m → End Λ• C2m that is
determined (where we regard C2m ⊂ Cl2m ) by
Q1 : C2m → End Λ• C2m ,


which is the same as ρ1 : Cm → End(Λ• (Cm )) with m replaced by 2m. However in


the case of Q, we think of C2m as being the complexification 
of R2m instead of R2m
m • 2m
as the realification of C . Let q0 : SO(2m) → End Λ C denote the complex
extension of the usual representation SO(2m) → End Λ• R2m . Since ∧ and x


are SO(2m)-invariant operations, for A ∈ SO(2m) and v ∈ C2m , we have


q0 (A) ◦ Q1 (v) ◦ q0 A−1 = Q1 (A(v)) = Q1 (r(A) v) .


It follows that for A ∈ SO(2m), and α ∈ Cl2m , we have


q0 (A) ◦ Q(α) ◦ q0 A−1 = Q(r(A) α) , and


q(g) Q(α) q g −1 = Q(r(c(g)) α) for all g ∈ Spin(n) and α ∈ Cln ,




where q = q0 ◦ c. Hence, Λ• (TC M ) := F M ×SO(2m) Λ• C2m is a Clifford module



+ −
bundle. We now determine Λ• (TC M ) and Λ• (TC M ) . For e1 , . . . , en an oriented,
n m
orthonormal basis of R and ωC = i e1 · · · en , we show that
Q(ωC )(ei1 ∧ · · · ∧ eik ) = im+k(k−1) ∗(ei1 ∧ · · · ∧ eik ) ,
where ∗ denotes the Hodge star operator on Λ• C2m . Let


∗(ei1 ∧ · · · ∧ eik ) = ej1 ∧ · · · ∧ ejn−k .


602 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

Since ei1 · · · eik ej1 · · · ejn−k = e1 · · · en in Cln , we have


Q(ωC )(ei1 ∧ · · · ∧ eik ) = Q im ei1 · · · eik ej1 · · · ejn−k (ei1 ∧ · · · ∧ eik )


k(n−k)
= im (−1)

Q ej1 · · · ejn−k ei1 · · · eik (ei1 ∧ · · · ∧ eik )
k(2m−k)
= im (−1)

Q ej1 · · · ejn−k Q(ei1 · · · eik )(ei1 ∧ · · · ∧ eik )
−k2 k(k−1)/2
= im (−1)

Q ej1 · · · ejn−k (−1) Q(eik · · · ei1 )(ei1 ∧ · · · ∧ eik )
m −k+k(k−1)/2  k
= i (−1) Q ej1 · · · ejn−k (−1) 1
k(k−1)/2
= im (−1) ej1 ∧ · · · ∧ ejn−k = im+k(k−1) ∗(ei1 ∧ · · · ∧ eik ) .
Hence if ∗k := ∗|Λk(C2m ) ,
Mn
(17.133) Q(ωC ) = im+k(k−1) ∗k .
k=0
Usually one defines the star operator on forms, sections of exterior products of
T ∗ M , instead of sections of exterior products of T M , but the Riemannian metric
identifies T ∗ M with T M , and so we have a choice of leaving the above alone or
dualizing it, replacing ρ1 : Cm → End(Λ• (Cm )) by
∗ 
ρ∗1 : Cm → End Λ• (Cm ) defined by
ρ∗1 (w)(α) := (w[ ∧ − w[ x)(α) = w[∧ α − w[ xα.
The advantage of dualizing is that we will find that then the Dirac operator DW
is the familiar sum d + δ of the exterior derivative d and its adjoint (the exterior
∞ •
codifferential δ), instead of the less familiar operators
 on C (Λ (T M )). Thus,

we go ahead and dualize, in which case W = Λ• C2m

. Referring to (17.122),
(17.123) and (17.133), we then have that µC (x) := c(ωC (x)) ∈ End(W (M )x ) is
given by
n
M n
M
µC (x) = τ (x) := τk := im+k(k−1) ∗k at x ∈ M , and
k=0 k=0
∗
W ± (M ) = Λ± (M ) := (1 ± τ ) Λ• (TC M ) .
2
We may directly check that τ (x) = Id. Indeed,
  
im+(2m−k)(2m−k−1) ∗2m−k im+k(k−1) ∗k
= im+(2m−k)(2m−k−1)+m+k(k−1) (∗2m−k ∗k )
2
+4m2 −4mk k(2m−k) k2 −k2
= i2k (−1) Id = (−1) (−1) Id = Id .
Note that DW = d + δ, since
DW ϕ = c E k (∇Ek α) = ϕk ∧(∇Ek α) − E k x(∇Ek α) = (d + δ) ϕ.


Definition 17.61. a) In view of this and the theorem and corollary below,
d + δ is sometimes called the DeRham-Dirac operator.
b) A form ϕ ∈ Ωk (M ) := Ωk (M, C) is called harmonic if (d + δ) ϕ = 0. We let
Hk (M ) := ϕ ∈ Ωk (M, C) : (d + δ) ϕ = 0


denote the space of harmonic forms.


17.6. THE INDEX THEOREM FOR STANDARD GEOMETRIC OPERATORS 603

We show the coincidence of our definition of harmonic forms with the common
language of denoting functions as harmonic
 when they belong to the kernel of the
Laplacian. Since DW is elliptic, Ker DW is finite-dimensional. Moreover, we have
Theorem 17.62 (Hodge-DeRham Decomposition). For 0 ≤ k ≤ n, there is an
orthogonal decomposition
Ωq (M ) = Hq (M ) ⊕ d Ωq−1 (M ) ⊕ δ Ωq+1 (M )
 

(17.134) = Hq (M ) ⊕ dδ(Ωq (M )) ⊕ δd(Ωq (M )) .


Proof. The operator
2
∆ := dδ + δd = (d + δ) : Ωq (M ) → Ωq (M )
is elliptic, and hence we have the orthogonal decomposition
Ωq (M ) = Ker ∆ ⊕ ∆(Ωq (M )) .
Moreover, Ker ∆ = Hq (M ). Indeed,
α ∈ Hq (M ) =⇒ (d + δ) α = 0
=⇒ dα = 0 and δα = 0 =⇒ (dδ + δd) α = 0,
and conversely if α ∈ Ker ∆, then
2 2
0 = ((dδ + δd) α, α) = kδαk + kdαk =⇒ dα + δα = 0.
It remains to prove that we have an orthogonal decomposition
∆(Ωq (M )) = d(Ωq (M )) ⊕ δ Ωq+1 (M ) .

(17.135)
The summands are orthogonal since
β ∈ Ωq−1 (M ) and γ ∈ Ωq+1 (M ) =⇒ (dβ, δγ) = d2 β, γ = 0.


Moreover, for any α ∈ Ωq (M ), we have


∆α = dδα + δdα ∈ d Ωq−1 (M ) ⊕ δ Ωq+1 (M ) ,
 
(17.136)

and so ∆(Ωq (M )) ⊆ d Ωq−1 (M ) ⊕ δ Ωq+1 (M ) .


 
 
For the reverse inclusion, note that d Ωq−1 (M ) and δ Ωq−1 (M ) are both in

Hq (M ) , since
α ∈ Hq (M ) =⇒ (dβ, α) = (β, δα) = 0 = (γ, dα) = (δγ, α) ,

and so d Ωq−1 (M ) ⊕ δ Ωq+1 (M ) ⊆ Hq (M ) = ∆(Ωq (M )) .
 

Note that (17.136) and (17.135) then give the second equality of (17.134). 

Corollary 17.63. Suppose that dγ = 0 for some γ ∈ Ωq (M ). There is a


unique α ∈ Hq (M ), such that for some β ∈ Ωq−1 (M ), γ = α + dβ. In other words,
every cohomology class in the de Rham cohomology space

q Ker d : Ωq (M ) → Ωq+1 (M )
H (M ) :=
dΩq−1 (M )
has a unique harmonic representative.
604 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

Proof. Theorem 17.62 yields a unique α ∈ Hq (M ) such that


γ = α + dβ + δβ 0
for some β ∈ Ωq−1 (M ) and β 0 ∈ Ωq+1 (M ). Now
0 = dγ = dα + d2 β + dδβ 0 = dδβ 0
2
=⇒ (dδβ 0 , β 0 ) = 0 =⇒ kδβ 0 k = 0 =⇒ δβ 0 = 0. 
Now we are ready to formulate and prove the Hirzebruch Signature Formula in
full detail. It was announced in our Section 13.4, Theorem 13.6e. To fix the notation
and to keep this section rather self contained, we recall the basic constructions
sketched there.
If Hm (M ; R) (∼= H m (M ; R)) denotes the space of R-valued harmonic forms on
M , we have a bilinear form
Z
B : Hm (M ; R) × Hm (M ; R) −→ R given by B(α, β) := α∧β.
M
If m is odd, B is antisymmetric.
R If m Ris even (i.e. dim M = 2m ≡ 0 mod 4), B is
symmetric. Let (α, β) := M α ∧ ∗β = M hα, βi νM . Since
2 m(2m−m) m
(∗m ) = (−1) Id = (−1) Id,
we have
Z Z
m
B(α, β) = α∧β = α ∧ ∗m (∗m β) = (α,(−1) ∗m β) .
M M
2
Note that τm = im+m(m−1) ∗m = im ∗m . Thus,
m even =⇒ τm = ∗m =⇒ B(α, β) = (α, ∗m β) = (α, τm β) .
Observe that α ∈ Hk (M ; R) ⇔ ∗α ∈ Hn−k (M ; R), since (using δ = ± ∗ d∗)
dα = 0 and δα = 0 ⇐⇒ dα = 0 and d(∗α) = 0
⇐⇒ d(∗(∗α)) = 0 and d(∗α) = 0 ⇐⇒ δ(∗α) = 0 and d(∗α) = 0.
Alternatively, we know on general grounds that τ = c(ωC ) and d + δ = DW anti-
commute. At any rate,
k∗ k τ
Hk (M ; R) ∼
= Hn−k (M ; R) and Hk (M ; C) ∼
= Hn−k (M ; C)
Moreover, B(α, β) = (α, ∗m β) = (α, τm β) implies that
sig(M ) := signature of B
= dimR ((1 + ∗) Hm (M ; R)) − dimR ((1 − ∗) Hm (M ; R))
= dimC ((1 + τ ) Hm (M ; C)) − dimC ((1 − τ ) Hm (M ; C)) ;
i.e., sig(M ) is the difference between the dimensions of the spaces of self-dual and
anti-self-dual harmonic forms. In the present case, the generalized Dirac operator
+
DW + = (d + δ) is known as the signature operator, because
 
+
index (d + δ) : C ∞ Λ+ (M ) → C ∞ Λ− (M )


+ −
= dim Ker(d + δ) − dim Ker(d + δ)
= dim((1 + τ ) H∗ (M ; C)) − dim((1 − τ ) H∗ (M ; C))
= dim((1 + τ ) Hm (M ; C)) − dim((1 − τ ) Hm (M ; C)) = sig(M ) .
17.6. THE INDEX THEOREM FOR STANDARD GEOMETRIC OPERATORS 605

The last line follows from the fact that for k 6= m,


1 ± τk : Hk (M ; C) → Hk (M ; C) ⊕ H2m−k (M ; C)
is clearly injective, and so for k 6= m,
dim(1 + τ ) Hk (M ; C) = dim Hk (M ; C) = dim(1 − τ ) Hk (M ; C) .
Now Theorem 17.59 yields
  
sig(M ) = index DW,+ = ch(W (M )) ` A(M e ) [M ]
 
= ch(Λ• (T ∗ M )) ` A(M
e ) [M ] .

Using (15.120), p.455, and (17.128), we obtain


∗ 
ch Λ• (TC M ) ` A(M
e )
 
2
 y/2
= MC 4 cosh (y/2) , T M ` MC ,TM
sinh y
 
y/2
= MC 4 cosh2 (y/2) ,TM
sinh y
 
2 y/2
= MC 4 cosh (y/2) ,TM
2 sinh(y/2) cosh(y/2)
 
y
(17.137) = MC ,TM .
tanh(y/2)
Let degm denote the m-th degree part of a power series in y1 , . . . , ym . Then
Y  Y 
m yk m yk /2
degm = 2m degm
k=1 tanh(yk /2) k=1 tanh(yk /2)
Y  Y 
m yk m yk
= 2m 2−m degm = degm .
k=1 tanh yk k=1 tanh yk

y 
Since L(M ) := L(T M ) = MC , T M , we then obtain
tanh y
Theorem 17.64 (Hirzebruch Signature Theorem). Let M be a compact, ori-
ented Riemannian 2m-manifold, where m is even. Then
sig(M ) = L(M ) [M ] =: L(M ) = L-genus of M .
In terms of the Pontryagin classes pk = pk (T M ) , we have
(
1/3 p [M ] , 2m = 4,
1
sig(M ) = 
1/45 7p − p2 [M ] , 2m = 8,
2 1

in particular, and one may extend this using additional Hirzebruch L-polynomials
of (15.111).
There is a twisted version of this theorem which we describe as follows. Let
E → M be a Hermitian vector bundle with a covariant derivative ∇E arising
from a unitary connection 1-form ε on U(E). Recall from Section 15.3, that there
is an exterior covariant derivative operator Dε : Ωk (M, E) → Ωk+1 (M, E) which
generalizes d : Ωk (M ) → Ωk+1 (M ), but Dε ◦ Dε 6= 0 in general. Moreover, Dε has a
formal adjoint δ ε : Ωk+1 (M, E) → Ωk (M, E). Using the local formulas (15.22) and
606 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

 ∗ 
(15.23), we see that for W = Λ• C2m , the twisted generalized Dirac operator
is
E
(17.138) DE,W = (d + δ) := Dε + δ ε : Ω• (M, E) → Ω∗ (M, E) .
which is called the twisted DeRham-Dirac operator. The twisted version of the
Hirzebruch
 Signature
 Theorem has the additional twist that perhaps unexpectedly,
E,+
index (d + δ) 6= (ch(E) ` L(M )) [M ] in general.

Theorem 17.65 (Twisted Hirzebruch Signature Theorem). Let M be a com-


pact, oriented Riemannian 2m-manifold, where m is even, and let E → M be a
Hermitian vector bundle with a covariant derivative ∇E arising from a unitary con-
nection 1-form ε on U(E). Then, with Ω± (M, E) := C ∞ (E ⊗ Λ± (M ))), the index
of the twisted signature operator
E,+
(d + δ) = DE,W + : Ω+ (M, E) → Ω− (M, E)
is given by
   
y

E,+
index (d + δ) = ch(E) ` MC ,TM [M ]
tanh(y/2)
  
y
= ch2 (E) ` MC ,TM [M ] = (ch2 (E) ` L(M )) [M ] ,
tanh y
Lm
where ch2 (E) := j=0 2j chj (E) .

Proof. Using Theorem 17.60 and (17.137),


    
E,+ ∗ 
index (d + δ) = ch(E) ` ch Λ• (TC M ) ` A(M
e ) [M ]
  
y
= ch(E) ` MC ,TM [M ] .
tanh y/2
Now, recalling the definition of the polynomial components Lk of the total Hirze-
bruch L class we find
   
y m y/2
MC , T M = 2 MC ,TM
tanh y/2 tanh y/2
Mm
= 2m 2−2k Lk (y) .
k=1

Thus,
 
y
ch(E) ` MC , T M [M ]
tanh y/2
Mm Mm 
= chj (E) ` 2m 2−2k Lk (y) [M ]
j=1 k=1
M 
= 2m−2k chj (E) ` Lk (y) [M ]
j+2k=m
M 
= 2j chj (E) ` Lk (y) [M ]
j+2k=m
= (ch2 (E) ` L(M )) [M ] . 
17.6. THE INDEX THEOREM FOR STANDARD GEOMETRIC OPERATORS 607

Almost Complex Structures. Recall that J ∈ End(T M ) is an almost


complex structure of a manifold M , if J 2 = −I. If J exists, then T M be-
comes a complex vector bundle by defining (a + ib) X = aX + bJX. For a complex
structure to exist, dim M must be even.
Exercise 17.66. Use Theorem 17.65 to show that if the sphere S n (n = 2m > 0)
admits an almost complex structure, then m = 1, 2 or 3:
(a) Show that the total Pontryagin class p(S n ) is 1 ∈ H 0 (S n ). [Hint. If N is the
(trivial) normal bundle of S n ⊆ Rn+1 , then T S n ⊕ N is trivial.]
(b) In the identity (x + x1 ) · · ·(x + xm ) = xm + σ1 xm−1 + · · · + σm , where σi de-
notes the i-th elementary symmetric polynomial in the xi , successively substitute
x = x1 , . . . , x = xm and add the results to get Newton’s formula
m
X m
X
m m−1
(17.139) 0 = (−1) xm
i +(−1) σ1 xm−1
i + · · · + mσm .
i=1 i=1

(c) Assuming that S 2m admits


 an almost complex structure, use (17.139) and the
fact cm T S 2m = χ S 2m (see (15.117) p. 454) to deduce that
  2m  (−1)m+1 2m+1
2m chm T S 2m S = .
(m − 1)!

(d) Finally show that (c) and Theorem 17.65 imply that m must be 1, 2, or 3.
Remark 17.67. Of course S 2 has a complex (and hence almost-complex) struc-
ture. S 4 does not have a complex structure, since otherwise,
 
1 = 1 − p1 T S 4 = c C ⊗ T S 4 = c T S 4 ⊕ T S 4
 
  2
= c T S 4 ` c T S 4 = 1 + c2 T S 4 = 1 + 2c2 T S 4 ,
 

 
but c2 T S 4 = χ S 2m 6= 0; the same argument works for S 4k , using
k
1 +(−1) pk T S 4 = c C ⊗ T S 4k .
 

We show that S 6 has a complex structure as follows. Recall that the Cayley numbers
O := H × H, with multiplication
pq = (p1 , p2 )(q1 , q2 ) = (p1 q1 − q2 p2 , q2 p1 + p2 q1 ) ,
are neither associative nor commutative, but (1, 0) is a multiplicative identity. Let
p := (p1 , −p2 ) and <(p) := 12 (p + p) = <(p1 ). Then
hp, qi := <(pq) = <(p1 q1 + q2 p2 ) and
2 2 2
|p| = hp, pi = <(p1 p1 + p2 p2 ) = |p1 | + |p2 | .
Thus, if we identify p and q with vectors in R8 , then hp, qi is just the usual dot
product, and
Σ6 := {p ∈ O : |p| = 1 and hp,(1, 0)i = 0}
is a 6-sphere. Note that
p ∈ Σ6 =⇒ p2 = (p1 p1 − p2 p2 , p2 p1 + p2 p1 )
2
= (−p1 p1 − p2 p2 , p2 p1 − p2 p1 ) = − |p| (1, 0) = −(1, 0) .
608 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

One can check that |pq| = |p| |q| and that (in spite of nonassociativity), p(pq) = p2 q.
These are rather involved computations, where, e.g., one eventually needs to use
<(xy) = <(yx) for x, y ∈ H, and
p2 q1 p1 + p2 q1 p1 −(p2 p1 q1 + p2 p1 q1 ) = (p2 q1 )(p1 + p1 ) − p2 (p1 + p1 ) q1 = 0.
For p ∈ Σ6 , we have hpq, pq 0 i = hq, q 0 i for all q, q 0 ∈ O, since |pq| = |p| |q| = |q| .
Suppose that hq,(1, 0)i = 0 and hq, pi = 0, so that q ∈ Tp Σ6 . Then
hpq,(1, 0)i = hpq, −ppi = − hq, pi = 0 and
hpq, pi = hpq, p(1, 0)i = hq,(1, 0)i = 0.
Thus, we have a well-defined linear map Jp : Tp Σ6 → Tp Σ6 given by Jp q = pq.
Moreover, Jp defines an almost complex structure since
2
Jp2 (q) = p(pq) = p2 q = − |p| q = −q.

The Chern-Gauss-Bonnet Formula. As noted before, the representation


Q : Cl2m → End Λ• C2m ,


determined by Q(w) = w∧ α − wxα for w ∈ C2m , is highly reducible. Indeed for


purely dimensional reasons, Q is the sum of 2m copies of the unique irreducible
• m
representation ρC : Cl2m →
 End(Λ (C )). We now describe a coarser (if m > 1)
decomposition of Λ• C2m into just two Cl2m -modules, which leads to the Chern-
Gauss-Bonnet Formula. Let
Λ± := (1 ± ωC (Q))(Λ• (Cm )) and Λev(odd),± := Λev(odd) C2m ∩ Λ± .


In view of the fact that for w ∈ C2m ,


 
Q(w) Λ± = Λ∓ and Q(w) Λev(odd) C2m = Λodd(ev) C2m
 
 
=⇒ Q(w) Λev(odd),± = Λodd(ev),∓ ,

the Cl2m -module Λ• C2m splits into two submodules:




Λ• C2m = Λev,+ ⊕ Λodd,− ⊕ Λodd,+ ⊕ Λev,− =: W e ⊕ W o .


  

Consequently, we have two generalized Dirac operators


e
DW = d + δ ∈ End Ωev,+ (M ) ⊕ Ωodd,− (M )

and
o
DW = d + δ ∈ End Ωodd,+ (M ) ⊕ Ωev,− (M ) ,

(17.140)
where Ωev(odd),± (M ) := C ∞ Λev(odd),± (M ) . Note that


e
DW +
: Ωev,+ (M ) → Ωodd,− (M ) and
o
DW +
: Ωodd,+ (M ) → Ωev,− (M ) , with
 e
  
index DW + = ch Λev,+ (M ) ⊕ Λodd,− (M ) ` A(M

e ) [M ] and
 o
  
index DW + = ch Λodd,+ (M ) ⊕ Λev,− (M ) ` A(M

e ) [M ] .
The signature operator can be written as
+ e o
(d + δ) = DW +
⊕ DW +
: Ω+ (M ) → Ω− (M ) .
17.6. THE INDEX THEOREM FOR STANDARD GEOMETRIC OPERATORS 609

The Euler operator is


χ
(17.141) (d + δ) := d + δ : Ωev (M ) → Ωodd (M ) ,
whose index (according to Hodge Theory) is χ(M ). Indeed,
Mm  Mm 
χ(M ) := dim H 2k (M ) − dim H 2k−1 (M )
Mmk=0  Mmk=1 
2k
= dim H (M ) − dim H2k−1 (M )
k=0 k=1
χ χ ∗
= dim(Ker(d + δ) ) − dim Ker((d + δ) )
χ
(17.142) = index((d + δ) ) .
e o χ
In terms of DW + and DW + , we have that (d + δ) is
e
 o
∗
DW + ⊕ DW + : Ωev,+ (M ) ⊕ Ωev,− (M ) → Ωodd,− (M ) ⊕ Ωodd,+ (M ) .
Hence,
 e
  o
∗ 
χ
index((d + δ) ) = index DW ,+ + index DW ,+
 e
  o

= index DW ,+ − index DW ,+
 
= ch Λev,+ (M ) ⊕ Λodd,− (M ) ` A(M

e ) [M ]
 
− ch Λodd,+ (M ) ⊕ Λev,− (M ) ` A(M

e ) [M ]
ch(Λev,+ (M )) − ch(Λev,− (M )) 
 
(17.143) = ` A(M
e ) [M ] .
+ ch Λodd,− (M ) − ch Λodd,+ (M )

The individual results for the ch Λev(odd),± (M ) are not all that simple or inter-
esting (as far as we know), but the two combinations
ch Λev,+ (M ) − ch Λev,− (M ) and ch Λodd,− (M ) − ch Λodd,+ (M )
   

admit dramatic simplification. Indeed, we will find that when m := 12 dim M is odd
the first of these is 0 and the second, when cupped with A(M e ), is the Euler class.
When m is even, the opposite is the case. Thus, whether m is even or odd, we will
obtain the Chern-Gauss-Bonnet Theorem.
∗ 
Let Λ : SO(n) → U Λ• (Cn ) denote the representation determined by
∗ ∗ ∗
Λ(A)(α) = AT α = α ◦ AT for α ∈ (Cn ) = Λ1 (Cn ) , A ∈ SO(n) .


The various ch Λev(odd),± (M ) can be computed by finding the eigenvalues of
0 • n ∗ 0 • n ∗
 
Λ (B) ∈ End Λ (C ) for B ∈ so(n), where Λ : so(n) → u Λ (C ) is the Lie
n
algebra representation. There is an
 orthonormal basis e 1 , . . . , e n of R for which B
Lm 0 −yk
is of the form k=1 . One purpose of the rather lengthy digression in
yk 0
the next paragraph is to show that
       
ev,+ ev,− odd,− odd,+
ch Λ0 (B) − ch Λ0 (B) + ch Λ0 (B) − ch Λ0 (B)
Ym
= (−2 sinh yk ) ,
k=1
from which the Chern-Gauss-Bonnet Theorem will follow. The results in the di-
gression are also used in proving the Hirzebruch-Riemann-Roch Theorem.
610 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS


Digression. Let ϕ1 , . . . , ϕn denote the basis of (Rn ) dual to e1 , . . . , en . Then
Λ0 (B)(ϕ1 + iϕ2 )(e1 ) = (ϕ1 + iϕ2 ) B T e1 = (ϕ1 + iϕ2 )(−Be1 )


= (ϕ1 + iϕ2 )(−y1 e2 ) = −iy1


0
Λ (B)(ϕ1 + iϕ2 )(e2 ) = (ϕ1 + iϕ2 )(−Be2 ) = (ϕ1 + iϕ2 )(y1 e1 ) = y1
Λ0 (B)(ϕ1 + iϕ2 ) = −iy1 ϕ1 + y1 ϕ2 = −iy1 (ϕ1 + iϕ2 ) .
Thus,
Λ0 (B)(ϕ1 + iϕ2 ) = −iy1 (ϕ1 + iϕ2 ) and Λ0 (B)(ϕ1 − iϕ2 ) = iy1 (ϕ1 − iϕ2 ) .
For j = 1, . . . , m, let ξj = ϕ2j−1 + iϕ2j and ξ¯j = ϕ2j−1 − iϕ2j . For a multi-index
(i)k := (i1 , . . . , ih ) where 1 ≤ i1 < · · · < ih ≤ m, we let
ξ(i) := ξi ∧ · · · ∧ ξi
h 1 and ξ¯(i) = ξ¯i ∧ · · · ∧ ξ¯i .
h h 1 h

Note that ξ(i)m = ξ1 ∧ · · · ∧ ξm and ξ¯(i)m = ξ¯1 ∧ · · · ∧ ξ¯m . The 2m · 2m = 2n forms


∗
ξ(i)h ∧ ξ¯(j)k make a basis of eigenvectors of Λ0 (B) on Λ• (Cn ) since
Λ0 (B) ξ(i)h ∧ ξ¯(j)k = −i((yi1 + · · · + yih ) −(yj1 + · · · + yjk )) ξ(i)h ∧ ξ¯(j)k .


However, we need a basis in Λev(odd),± . For this, we will determine ∗ ξ(i)h ∧ ξ¯(j)k


where ∗ denotes the complex-linear extension of the usual star operator on Λ• (Rn ).
If b(α, β) is the symmetric, bilinear (not Hermitian) extension of the usual inner
product on Λ• (Rn ), we have
α ∧ ∗β = b(α, β) ϕ1 ∧ · · · ∧ ϕn .
Note that
2 2
b(ξj , ξj ) = b(ϕ2j−1 + iϕ2j , ϕ2j−1 + iϕ2j ) = |ϕ2j−1 | − |ϕ2j | = 0, and
2 2
b ξj , ξ¯j = b(ϕ2j−1 + iϕ2j , ϕ2j−1 − iϕ2j ) = |ϕ2j−1 | + |ϕ2j | = 2.


Hence,
2h+k , if (i0 )h0 = (i)h , (j 0 )k0 = (j)k ,

b ξ¯(i0 )h0 ∧ ξ(j 0 )k0 , ξ(i)h ∧ ξ¯(j)k =

0, otherwise.
Moreover, using ξ1 ∧ ξ¯1 = (ϕ1 + iϕ2 ) ∧(ϕ1 − iϕ2 ) = −2iϕ1 ∧ ϕ2 , we get
m
(−2i) ϕ1 ∧ · · · ∧ ϕn = ξ1 ∧ ξ¯1 ∧ · · · ∧ ξm ∧ ξ¯m
m(m−1)/2
= (−1) ξ1 ∧ · · · ∧ ξm ∧ ξ¯1 ∧ · · · ∧ ξ¯m
= im(m−1) ξ1 ∧ · · · ∧ ξm ∧ ξ¯1 ∧ · · · ∧ ξ¯m , or
−m m(m−1)
νn := ϕ1 ∧ · · · ∧ ϕn = (−2i) i ξ1 ∧ · · · ∧ ξm ∧ ξ¯1 ∧ · · · ∧ ξ¯m
2 2
(17.144) = 2−m im ξ1 ∧ · · · ∧ ξm ∧ ξ¯1 ∧ · · · ∧ ξ¯m = 2−m im ξ(i)m ∧ ξ¯(i)m .
From
ξ¯(i0 )h0 ∧ ξ(j 0 )k0 ∧ ∗ ξ(i)h ∧ ξ¯(j)k = b ξ¯(i0 )h0 ∧ ξ(j 0 )k0 , ξ(i)h ∧ ξ¯(j)k νn ,
  

we deduce that for some scalar C = C((i)h ,(j)k ) ∈ C,


∗ ξ(i) ∧ ξ¯(j) = Cξ(j c ) ∧ ξ¯(ic )

(17.145) h k m−k m−h
,
where (j c )m−k is the multi-index complementary (j)k in the sense that
{j1 , . . . , jk } ∪ j1c , . . . , jm−k
c

= {1, . . . , m} .
17.6. THE INDEX THEOREM FOR STANDARD GEOMETRIC OPERATORS 611

We find below that


h 2
(17.146) C = C((i)h ,(j)k ) = 2h+k−m (−1) im ε(ic )m−h(i)h ε(j)k(j c )m−k .
Indeed if (i0 )h0 6= (i)h or (j 0 )k0 6= (j)k , then
ξ¯(i0 )h0 ∧ ξ(j 0 )k0 ∧ ξ(j c )m−k ∧ ξ¯(ic )m−h = b ξ¯(i0 )h0 ∧ ξ(j 0 )k0 , ξ(i)h ∧ ξ¯(j)k νn = 0,
 

while if (i0 )h0 = (i)h and (j 0 )k0 = (j)k , then


 ∼
ξ¯(i)h ∧ ξ(j)k −→ ∧Cξ(j c )m−k ∧ ξ¯(ic )m−h
hm
= C(−1) ξ(j)k ∧ ξ(j c )m−k ∧ ξ¯(i)h ∧ ξ¯(ic )m−h
hm
= C(−1) ε(j)k(j c )m−k ε(i)h(ic )m−h ξ(j)m ∧ ξ¯(i)m
hm
= C(−1) ε(i)h(ic )m−h ε(j)k(j c )m−k ξ(j)m ∧ ξ¯(i)m
hm h(m−h)
= C(−1) (−1) ε(ic )m−h(i)h ε(j)k(j c )m−k ξ(j)m ∧ ξ¯(i)m
h
= C(−1) ε(ic )m−h(i)h ε(j)k(j c )m−k ξ(j)m ∧ ξ¯(i)m
h 1
= C(−1) ε(ic )m−h(i)h ε(j)k(j c )m−k −m m2 νn
2 i
¯ ¯ h+k

= b ξ(i)h ∧ ξ(j)k , ξ(i)h ∧ ξ(j)k νn = 2 νn ,
which yields the value of C in (17.146). We have
τh+k ξ(i)h ∧ ξ¯(j)k = im+(h+k)(h+k−1) ∗ ξ(i)h ∧ ξ¯(j)k
 

= im+(h+k)(h+k−1) C((i) ,(j) ) ξ(j c )


h k∧ ξ¯(ic )
m−k m−h

h 2
= i m+(h+k)(h+k−1) h+k−m
2 (−1) im ε(ic )m−h(i)h ε(j)k(j c )m−k ξ(j c )m−k ∧ ξ¯(ic )m−h
= im(m+1) i(h+k)(h+k−1)+2h 2h+k−m ε(ic )m−h(i)h ε(j)k(j c )m−k ξ(j c )m−k ∧ ξ¯(ic )m−h
1
m(m+1)+ 21(h−k)(h−k+1)
= (−1) 2 2h+k−m ε(ic )m−h(i)h ε(j)k(j c )m−k ξ(j c )m−k ∧ ξ¯(ic )m−h ,
where we have used
(h + k)(h + k − 1) + 2h = ((h − k)(h − k + 1)) + 4hk.
In the case h + k = m, we have
1 1 1 1
m(m + 1) + (h − k)(h − k + 1) = m(m + 1) + (2h − m)(2h − m + 1)
2 2 2 2
= m2 + 2h2 − 2mh + h ≡ m + h mod 2, and so
m+h
h + k = m =⇒ τm ξ(i)h ∧ ξ¯(j)k = (−1) ε(ic )m−h(i)h ε(j)k(j c )m−k ξ(j c )m−k ∧ ξ¯(ic )m−h .


Note that if h + k = m, then (1 ± τm ) ξ(i) ∧ ξ¯(j) are both eigenvectors of τm



h k
(with the same eigenvalue for Λ0 (B)), unless one is 0. To determine when
(1 ± τm ) ξ(i)h ∧ ξ¯(j)k = 0,


note that (17.145) yields


τm ξ(i)h ∧ ξ¯(j)k = ±ξ(i)h ∧ ξ¯(j)k ⇔ (j)k = (ic )m−h (or (i)h = (j c )m−k ).


 2
In this case, ε(ic )m−h(i)h ε(j)k(j c )m−k = ε(ic )m−h(i)h = 1, and so
 
m+h
(17.147) τm ξ(i)h ∧ ξ¯(ic )m−h = (−1) ξ(i)h ∧ ξ¯(ic )m−h , for 0 ≤ h ≤ m,
612 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

while (1 ± τm ) ξ(i)h ∧ ξ¯(j)k 6= 0 for (i)h 6= (j c )m−k . Consequently, if (j)k 6= (ic )m−h ,

∗
the eigenvector (1 + τh+k ) ξ(i)h ∧ ξ¯(j)k of Λ0 (B) |Λev(odd),+ (Cm ) corresponds to an

∗
eigenvector (1 − τh+k ) ξ(i)h ∧ ξ¯(j)k of Λ0 (B) |Λev(odd),− (Cm ) with the same eigen-


value. In the remaining case (j)k = (ic )m−h , we have


    
Λ0 (B) ξ(i)h ∧ ξ¯(ic )m−h = −i (yi1 + · · · + yih ) − yic1 + · · · + yicm−h .

First Consequences. Thus taking (17.147) into account, we have


       
ev,+ ev,− odd,− odd,+
ch Λ0 (B) − ch Λ0 (B) + ch Λ0 (B) − ch Λ0 (B)
Xm X m+h
   
= (−1) exp − (yi1 + · · · + yih ) − yic1 + · · · + yicm−h
h=0 (i)h
m
Xm X h
  
= (−1) (−1) exp yic1 + · · · + yicm−h −(yi1 + · · · + yih )
h=0 (i)h
m
Ym  Ym
= (−1) e − e−yk =
yk
(−2 sinh yk ) , and so
k=1 k=1

     
ev,+ ev,−
ch Λ0 (B) − ch Λ0 (B)
(17.148)  
odd,−
 
odd,+
  A(B)
e
+ ch Λ0 (B) − ch Λ0 (B)
Ym Ym yk /2 m m
= (−2 sinh yk ) = (−1) y1 . . . ym = (−1) Pf(B) ,
k=1 k=1 sinh yk
where Pf(B) denotes the Pfaffian defined in (15.101), p.446.
Theorem 17.68 (Chern-Gauss-Bonnet Theorem). Let M be a compact, ori-
entable, Riemannian manifold of even dimension n = 2m, and let Hk (M ) denote
the space of harmonic k-forms on M . Then
n
X
χ k
(−1) dim Hk (M )

χ(M ) = index((d + δ) ) =
k=0
Z
GB Ωθ ,

= GB(T M ) [M ] =
M
θ
 n
where GB Ω ∈ Ω (M ) denotes the Gauss-Bonnet form, determined by
m
π ∗ GB Ωθ = (−1) Pf 2π 1
Ωθ
 

1 X
= 2m m εi1 ···i2m Ωθi1 i2 ∧ · · · ∧ Ωθi2m−1 i2m ,
2 π m! (i)

where Ωθ denotes the curvature of the Levi-Civita connection θ (or indeed any
connection) on the bundle π : F M → M of oriented, orthonormal frames.
Proof. Using the preceding (17.142), (17.148) and our previous (15.115),
p.453,
Xn k χ
(−1) dim Hk (M ) = index((d + δ) )

k=0
ch(Λev,+ (M )) − ch(Λev,− (M )) 
 
= ` A(M
e ) [M ]
+ ch Λodd,− (M ) − ch Λodd,+ (M )
Z
GB Ωθ .

= GB(T M ) [M ] = 
M
17.6. THE INDEX THEOREM FOR STANDARD GEOMETRIC OPERATORS 613

There is also a twisted version of the Chern-Gauss-Bonnet Theorem, which has


the additional twist (or rather untwist) that the index of the twisted Euler operator
is only affected by the dimension of the twisting bundle E, rather than any actual
twisting (i.e., nontriviality) of E. We define the twisted Euler operator
E,χ E
(d + δ) : Ωev (M, E) → Ωodd (M, E) := (d + δ) |Ωev(M,E)
E E
as the restriction (d + δ) |Ωev(M,E) of the twisted DeRham-Dirac operator (d + δ)
in (17.138).

Theorem 17.69 (Twisted Chern-Gauss-Bonnet Theorem). Let M be a com-


pact, oriented Riemannian 2m-manifold, and let E → M be a Hermitian vector
bundle with a covariant derivative ∇E arising from a unitary connection 1-form ε
on U(E). Then the index of the twisted Euler operator is given by
 
E,χ
index (d + δ) = dim E · GB [T M ] [M ] = dim E · χ(M ) .

Proof. We compute
   e
  o
∗ 
E,χ
index (d + δ) = index DW ,E+ + index DW ,E+
 e
  o

= index DW ,E+ − index DW ,E+
 
= ch(E) ` ch Λev,+ (M ) ⊕ Λodd,− (M ) ` A(M

e ) [M ]
 
− ch(E) ` ch Λodd,+ (M ) ⊕ Λev,− (M ) ` A(M

e ) [M ]
ch(Λev,+ (M ))− ch(Λev,− (M )) 
 
= ch(E) ` ` A(M
e ) [M ]
+ch Λodd,− (M ) − ch Λodd,+ (M )
= (ch(E) ` GB(T M )) [M ] = (ch0 (E) ` GB(T M )) [M ]
= dim E · GB(T M ) [M ] . 

The Generalized Yang-Mills Index Theorem. Recall that in (17.140) we


e o
introduced the two generalized Dirac operators DW and DW and we used their
twisted versions
e E
DW ,E = (d + δ) ∈ End Ωev,+ (E) ⊕ Ωodd,− (E)

and
o E
DW ,E = (d + δ) ∈ End Ωodd,+ (E) ⊕ Ωev,− (E)

(17.149)

in the proof of the Twisted Gauss-Bonnet Theorem. For reasons that will be ex-
e o
plained below, we call DW ,E and DW ,E Yang-Mills-Dirac operators.

Theorem 17.70 (The Yang-Mills-Dirac Index Theorem). Let M be a compact,


oriented Riemannian 2m-manifold, where m is even, and let E → M be a Hermitian
vector bundle with a covariant derivative ∇E arising from a unitary connection 1-
form ε on U(E). Then
 e

index DW ,E+ = 12 (ch2 (E) ` L(M )) [M ] + 21 dim E · χ(M ) and
 o

index DW ,E+ = 12 (ch2 (E) ` L(M )) [M ] − 21 dim E · χ(M ) .
614 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

E,+ e o
Proof. Using (d + δ) = DW ,+ ⊕ DW ,+ and the above proof of Theorem
17.69, we have
   e
  o

E,+
index (d + δ) = index DW ,+ + index DW ,+ and
   e
  o

E,χ
index (d + δ) = index DW ,+ − index DW ,+ .

The results follow from adding and subtracting, since


 
E,+
index (d + δ) = (ch2 (E) ` L(M )) [M ] and
 
E,χ
index (d + δ) = dim E · χ(M ) ,

by the Twisted Hirzebruch Signature Theorem (Theorem 17.65, p.606) and Twisted
Chern-Gauss-Bonnet Theorem (17.69). 

We now explain the Yang-Mills-Dirac nomenclature. In the proof of the Formal


Dimension Theorem (Theorem 16.17, p.491), we computed the index of the operator

T : Ω1 (E) → Ω0 (E) ⊕ Ω2− (E) , given byT (α) := δ ω α, 12 (1 − ∗) Dω α ,




where E = P ×G gC for some principal G-bundle over a compact Riemannian 4-


manifold M and where gC denotes the complexification of the Lie algebra of G.
The kernel of T can be regarded as the formal dimension of tangent space (at ω) of
the manifold of moduli of connections on P with self-dual curvature. The operator
T bears a strong resemblance to the operator
o
DW ,E+
: Ωodd,+ (E) → Ωev,− (E) ,
E
which is the restriction of (d + δ) := Dω + δ ω ∈ End(Ω• (E)). Indeed, for

π± := 12 (1 ± τ ) : Ω• (E) → Ω•,± (E) ,


o
we have π− ◦ T = DW ,E
◦ π+ . We have isomorphisms

π+ |Ω1(E) : Ω1 (E) −→ Ωodd,+ (E) ⊂ Ω1 (E) ⊕ Ω3 (E) , and

π− |Ω0(E)⊕Ω2−(E) : Ω0 (E) ⊕ Ω2− (E) −→ Ωev,− (E) .

Thus, using ch2 (E) = dim E + 2 ch1 (E) + 4 ch2 (E), it follows that
 −1 o 
index(T ) = index π− |Ω0(E) ◦ DW ,E ◦ π− |Ω0(E) ⊕ Ω2− (E)
 o

= index DW ,E+ = 21 (ch2 (E) ` L(M )) [M ] − 21 dim E · χ(M )
= 21 ((dim E + 2 ch1 (E) + 4 ch2 (E)) ` L(M )) [M ] − 1
2 dim E · χ(M )
= 12 (4 ch2 (E) [M ] + dim E · L(M ) [M ]) − 1
2 dim E · χ(M )
1
(17.150) = 2 ch2 (E) [M ] − 2 dim E ·(χ(M ) − sig(M )) ,

in agreement with the computations in the proof of the Formal Dimension Theorem
(Theorem 16.17, p.491).
17.6. THE INDEX THEOREM FOR STANDARD GEOMETRIC OPERATORS 615

The Hirzebruch-Riemann-Roch Formula. Now we shall widen our hori-


zon.
Complex Structures. Let M be a complex manifold with
n = dim M = 2 dimC M = 2m.
In other words, M is a smooth n-manifold, and there is a covering {U } of M and
a collection {ϕU } of coordinate charts ϕU : U → Cm , such that
ϕV ◦ ϕ−1
U : ϕU (U ∩ V ) → ϕV (U ∩ V )

is holomorphic (i.e., ϕV ◦ ϕ−1


 m
U ∗ : TϕU(x) C → TϕV (p) Cm is complex linear for
each x ∈ U ∩ V ). The tangent spaces Tx M then possess a well-defined√map Jx ∈
Tx M (with Jx2 = − Idx ) which corresponds to multiplication by i = −1 under
∼ −1
(ϕU )∗ : Tx M → TϕU(x) Cm −→ Cm (i.e., Jx (X) = (ϕU )∗ (iϕU (X))). The bundle
automorphism J ∈ End(T M ) is known as the complex structure of the complex
manifold M . While it is tempting to explicitly make Tx M a complex vector space
by defining iX to be JX for X ∈ Tx M , this ultimately leads to profound confusion,
since it is customary (and of great utility) to consider the complexification of the
real vector space Tx M , namely
(TC M )x := C ⊗ Tx M = {V + iW : V, W ∈ Tx M }
of complex dimension 2m. The problem is that multiplication by i in TC M is
not the same as the complex linear extension of J to TC M . In particular while J
preserves the real subspace Tx M ⊂ (TC M )x , multiplication by i does not. Thus, to
avoid confusion, we use Jx instead of i for the complex structure on Tx M . Let
ϕU (x) = z 1 (x) , . . . , z m (x) = x1 (x) + iy 1 (x) , . . . , xm (x) + iy m (x) .
 

The Dolbeault Complex. We define C-valued, R-linear functionals dz j and dz̄ j


on Tx M via
dz j := dxj + idy j : Tx M → C and
dz̄ j := dxj − idy j : Tx M → C.

Any R-linear functional on Tx M , such as dz k or dz̄ k , extends uniquely to a C-linear


functional on (TC M )x , and we use the same symbols to denote these extensions;

i.e., dz j , dz̄ j ∈ (TC M )x . The local complex vector fields (local sections of TC M )
∂zk := 21 ∂xk − i∂yk , ∂z̄k := 12 ∂xk + i∂yk
 


are dual to dz j and dz̄ j ∈ (TC M ) , in the sense that
dz j (∂zk ) = 12 dxj + idy j ∂xk − i∂yk = δkj ,
 

dz̄ j (∂z̄k ) = 12 dxj − idy j ∂xk + i∂yk = δkj ,


 

dz̄ j (∂zk ) = dz j (∂z̄k ) = 0.


There is complex-linear extension of Jx to (TC M )x . We denote this extension by
the same symbol Jx . Since Jx2 = − Id, the eigenvalues of Jx are i and −i, and the
eigenspaces of Jx ∈ End((TC M )x ) are
Tx1,0 M := {V − iJV : V ∈ Tx M } and Tx0,1 M := {V + iJV : V ∈ Tx M } ,
616 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

respectively. Note that {∂z1, . . . , ∂zm } and {∂z̄1 , . . . , ∂z̄m } are local framings of
C ∞ Tx1,0 M and C ∞ Tx0,1 M , respectively. We set

Λp,0 (TC M ∗ )x := the vector space of all antisymmetric multi-complex-linear


p
functionals defined on Tx1,0 M × · · · × Tx1,0 M.

The Λp,0 (TC M ∗ )x are the fibers of a complex vector bundle Λp,0 (TC M ∗ ) → M . Let
Ωp,0 (M ) denote the space of C ∞ sections of Λp,0 (TC M ∗ ); i.e.,

Ωp,0 (M ) := C ∞ Λp,0 (TC M ∗ ) .




On a coordinate neighborhood U , such a section is of the form


1 X
fj1 ···jp dz j1 ∧ · · · ∧ dz jp ,
p!
(j)

where the fj1 ···jp ∈ C ∞ (U, C) are antisymmetric in j1 · · · jp . Similarly, we may


define the fibers Λ0,q(TC M ∗ )x , the bundle Λ0,q (TC M ∗ ) → M and the space Ω0,q (M )
:= C ∞ Λ0,q (TC M ∗ ) of sections which locally are of the form
1 X
fk1 ···kq dz̄ k1 ∧ · · · ∧ dz̄ kq .
q!
(k)

More generally, one has the bundle Λp,q (TC M ∗ ). The space of its sections is denoted
by Ωp,q (M ) := C ∞ (Λp,q (TC M ∗ )). The sections are locally of the form
1 X
fj1 ···jp ;k1 ···kq dz j1 ∧ · · · ∧ dz jp ∧ dz̄ k1 · · · ∧ dz̄ kq
p!q!
(j)(k)

and are called forms of bidegree (p, q). By writing dz j = dxj + idy j and dz̄ j =
dxj − idy j , we can regard such
 forms as ordinary forms p+q
 in Ω (M, C). Conversely,
j 1 j j j 1 j j
writing dx = 2 dz + dz̄ and dy = 2i dz − dz̄ , we see that

M
Λl (T M ∗ , C) : = C ⊗ Λl (T M ∗ , R) −→ Λp,q (TC M ∗ ) and
p+q=l

X
Ωl (M, C) : = C ⊗ Ωl (M, R) −→ Ωp,q (M ) .
p+q=l

There is an operator ∂¯ : Ω0,q (M ) → Ω0,q+1 (M ) given locally by


 X 
¯ 1 k1 kq
∂ fk1 ···kq dz̄ ∧ · · · ∧ dz̄
q! (k)
1 X Xm
∂z̄k0 fk1 ···kq dz̄ k0 ∧ dz̄ k1 ∧ · · · ∧ dz̄ kq .

:=
q! (k) k0 =1

More generally, one analogously defines ∂¯ : Ωp,q (M ) → Ωp,q+1 (M ), as well as


∂ : Ωp,q (M ) → Ωp+1,q (M ). While we have given these operators locally, they are
independent of local holomorphic coordinates. The operator ∂ + ∂¯ is the restriction
of the usual exterior derivative on Ωp,q (M ) ⊂ Ωp+q (M, C), namely

∂ + ∂¯ = d|Ωp,q(M ) : Ωp,q (M ) → Ωp+1,q (M ) + Ωp,q+1 (M ) .


17.6. THE INDEX THEOREM FOR STANDARD GEOMETRIC OPERATORS 617

This is a consequence of the fact that, for f ∈ C ∞ (M, C), locally we have
m
X
∂ + ∂¯ f = ∂zk (f ) dz j + ∂z̄k (f ) dz̄ j


k=1
Xm
1
dxk + idy k + 1
dxk − idy k
   
= 2 ∂xk f − i∂yk f 2 ∂xk f + i∂yk f
j=1
m
X
= ∂xk (f ) dxk + ∂yk (f ) dy k = df.
j=1

We have ∂ 2 = 0, ∂ ∂¯ + ∂∂
¯ = 0 and ∂¯2 = 0, since
2
0 = d2 = ∂ + ∂¯ = ∂ 2 ⊕ ∂ ∂¯ + ∂∂
¯ ⊕ ∂¯2 .


In particular, since ∂¯2 = 0, we have a chain complex


∂¯ ∂¯ ∂¯
0 −→ Ω0,0 (M ) −→ Ω0,1 (M ) −→ · · · −→ Ω0,m (M )
and the Dolbeault cohomology spaces
¯ Ω0,q(M )

0,q Ker ∂|
H (M ) := ¯ 0,q−1 .
∂(Ω (M ))
In order to define harmonic representatives of Dolbeault cohomology classes,
we need a Hermitian metric to define an adjoint
∂¯∗ : Ω0,q+1 (M ) → Ω0,q (M ) ,
for ∂¯ : Ω0,q (M ) → Ω0,q+1 (M ). Suppose that we are given a Riemannian metric h
on M , so that for all x ∈ M and V, W ∈ Tx M, we have
h(JV, W ) = −h(V, JW ) , or equivalently h(JV, JW ) = h(V, W ) .
Such a metric can always be found by setting h(V, W ) = h0 (JV, JW ) + h0 (V, W )
for an arbitrary metric h0 . Then hx uniquely extends to a complex bilinear form
hC on C ⊗R Tx M, so that for V1 , V2 , W1 , W2 ∈ Tx M
hC (V1 + iV2 , W1 + iW2 )
:= h(V1 , W1 ) − h(V2 , W2 ) + i(h(V1 , W2 ) + h(V2 , W1 )) .
We have
hC (V ± iJV, W ± iJW ) = h(V, W ) − h(JV, JW ) ± i(h(V, JW ) + h(JV, W )) = 0.
Thus, the restrictions of hC to T 1,0 M and T 0,1 M are 0. For
(hC )j k̄ := hC (∂zj , ∂z̄k ) = hC (∂z̄k , ∂zj ) =: (hC )k̄j ,

we have (hC )j̄k = (hC )j k̄ and (where ⊗s denotes the symmetric tensor product)
X
hC = (hC )j k̄ dz j ⊗ dz̄ k +(hC )k̄j dz̄ k ⊗ dz j
j,k
X X
(hC )j k̄ 21 dz j ⊗ dz̄ k + dz̄ k ⊗ dz j = 2 (hC )j k̄ dz j ⊗s dz̄ k .

=2
j,k j,k
618 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

Chirality and the ∗ Map, Revisited. We shall elaborate a few of the intriguing
relations between real and complex theory.

Definition 17.71. The Kähler 2-form κ ∈ Ω2 (M, C) on a complex manifold


with Riemannian metric h with h(JV, JW ) = h(V, W ) is given by

κ(X, Y ) := hC (JX, Y ) .

Note that κ ∈ Ω1,1 (M ), since locally

κ(∂zj , ∂z̄k ) = hC (J∂zj , ∂z̄k ) o X


= hC (i∂zj , ∂z̄k ) = i(hC )j k̄ , and =⇒ κ = i (hC )j k̄ dz j ∧ dz̄ k .
j,k
κ(∂z̄j , ∂z̄k ) = κ(∂zj , ∂zk ) = 0

If ∂x1 , ∂y1 , . . . , ∂xm , ∂ym is orthonormal at x ∈ M , then

(hC )j k̄ = hC (∂zj , ∂z̄k ) = hC 21 ∂xj − i∂yj , 12 ∂xk + i∂yk


 

= 14 h(∂xj , ∂xk ) + 14 h ∂yj , ∂yk = 12 δjk .




Thus, for example, (hC )kk̄ = 21 h(∂xk , ∂xk ) = 12 h ∂yk , ∂yk = 12 , which is why we


use the notation (hC )j k̄ , instead of simply hj k̄ . Perhaps the bar over one of the
indices suffices to avoid any confusion in practice, but there remain oddities, such
as 0 = (hC )11 6= h11 = h(∂x1 , ∂x1 ) = 1, which would yield contradictions if we were
to denote (hC )11 simply by h11 .
If ∂x1 , ∂y1 , . . . , ∂xm , ∂ym is orthonormal at x ∈ M , then at x
Xm Xm
κ=i (hC )j k̄ dz j ∧ dz̄ k = i (hC )kk̄ dz k ∧ dz̄ k
j,k=1 k=1
i Xm i Xm
dz k ∧ dz̄ k = dxk + idy k ∧ dxk − idy k
 
=
2 k=1 2 k=1
Xm
k k
= dx ∧ dy , and
k=1

Xm  Xm 
∧m κ = dxk1 ∧ dy k1 ∧ · · · ∧ dxkm ∧ dy km
k1 =1 km =1
1 1 m m
= m! dx ∧ dy ∧ · · · ∧ dx ∧ dy .

Moreover, according to (17.144), we have

(17.151) νh = dx1 ∧ dy 1 ∧ · · · ∧ dxm ∧ dy m


2
= 2−m im dz 1 ∧ · · · ∧ dz m ∧ dz̄ 1 ∧ · · · ∧ dz̄ m .

If ∂x1 , ∂y1 , . . . , ∂xm , ∂ym is not necessarily orthonormal, we still have
m
X 
νh = m!1
∧m κ = im! ∧m (hC )j k̄ dz j ∧ dz̄ k .
j,k

Define a Hermitian metric Hx (complex linear in first slot and conjugate linear
in the second slot) on C ⊗R Tx M by

(17.152) H(V, W ) := (hC ) V, W .
17.6. THE INDEX THEOREM FOR STANDARD GEOMETRIC OPERATORS 619

Note that H|Tx1,0 M and H|Tx0,1 M are nondegenerate and relative to H, Tx1,0 M ⊥Tx0,1 M .
Indeed, for V, W ∈ Tx M, we have
H(V ± iJV, W ± iJW ) = H(V, W ) + H(V, ±iJW ) + H(±iJV, W ) + H(iJV, iJW )
= H(V, W ) + H(JV, JW ) = 2h(V, W ) and
H(V ± iJV, W ∓ iJW ) = H(V, W ) + H(V, ∓iJW ) + H(±iJV, W ) − H(iJV, iJW )
= ±i(h(V, JW ) + h(JV, W )) = 0.
We set
Hjk := H(∂zj , ∂zk ) = (hC )(∂zj , ∂z̄k ) = (hC )j k̄ ,
Hj k̄ := H(∂zj , ∂z̄k ) = (hC )(∂zj , ∂zk ) = (hC )jk = 0,
Hk̄j := H(∂z̄k , ∂zj ) = (hC )(∂z̄k , ∂z̄j ) = (hC )k̄j̄ = 0, and
Hj̄ k̄ := H(∂z̄j , ∂z̄k ) = (hC )(∂z̄j , ∂zk ) = (hC )j̄k .

In particular, if ∂x1 , ∂y1 , . . . , ∂xm , ∂ym is orthonormal at x ∈ M , then
Hjk = H(∂zj , ∂zk ) = (hC )j k̄ = 21 δjk = (hC )j̄k = H(∂z̄j , ∂z̄k ) = Hj̄ k̄ ,
√ √ √ √ 
whence 2∂z1 , . . . , 2∂zm , 2∂z̄1 , . . . , 2∂z̄m is orthonormal for H.
There is a conjugate-linear bundle isomorphism

[ : (TC M )x → (TC M )x
given, for V, W ∈ (TC M )x , by
[x (V )(W ) := H(W, V ) .
(T M ∗ ) and [ T 0,1 M = Λ0,1 (T M ∗ ). Indeed,
1,0
 1,0

We have [ T M =Λ
Xm Xm
[(∂zk ) = Hlk dz l and [(∂z̄k ) = Hl̄k̄ dz̄ l , since
l=1 l=1
Xm
[(∂zk )(∂zj ) = H(∂zj , ∂zk ) = Hjk = Hlk dz l (∂zj ) ,
l=1
[(∂zk )(∂z̄j ) = H(∂z̄j , ∂zk ) = Hj̄k = 0,
Xm
[(∂z̄k )(∂z̄j ) = H(∂z̄j , ∂z̄k ) = Hj̄ k̄ = Hl̄k̄ dz̄ l (∂z̄j ) .
l̄=1
[(∂z̄k )(∂zj ) = H(∂zj , ∂z̄k ) = Hj k̄ = 0.

In particular, if ∂x1 , ∂y1 , . . . , ∂xm , ∂ym is orthonormal for h at x ∈ M , then
[(∂zk ) = 21 dz k and [(∂z̄k ) = 12 dz̄ k . There is a conjugate-linear inverse to [, de-
noted by

# : (TC M )x → (TC M )x .

Now H on (TC M )x induces a Hermitian inner product (still denoted by H) on


∗ ∗
Λ1 (TC M )x , given for α, β ∈ Λ1 (TC M )x , by
H(α, β) = H(#β, #α) ,
where the switch in the order is due to the conjugate linearity of #. We then have
∗
an induced Hermitian inner product on Λk (TC M )x for 1 ≤ k ≤ m, where
{ei1 ∧ · · · ∧ eik : i1 < · · · < ik }
∗
is an orthonormal basis for Λk (TC M )x if e1 , . . . , em is an orthonormal basis of
∗  ∗
Λ1 (TC M )x . By restriction, we have Hermitian inner products on Λp,q (TC M )x
as well.
620 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

Let (·, ·) denote the Hermitian L2 inner product on Ω0,q (M ) induced by H,


namely Z p
(α, β) := H(α(x), β(x)) νh (x), with kαk := (α, α).
M
We may then speak of the formal adjoint
∂¯∗ : Ωp,q+1 (M ) → Ωp,q (M ) of ∂¯ : Ωp,q (M ) → Ωp,q+1 (M ) ,
having the property
¯ β = α, ∂¯∗ β .
 
∂α,
In order to exhibit a formula for ∂¯∗ , we have a C-linear star operator
∗ : Ωk (M, C) → Ωn−k (M, C) determined by
k
α ∧ ∗β := hC (α, β) νh for α, β ∈ Ω (M, C) .
Some authors (e.g., [186]) use H instead of hC in the definition of ∗, which makes
∗ conjugate-linear, while we, as others (e.g., [179]), use a C-linear star operator.
Thus, here

α ∧ ∗β̄ := gC α, β̄ νh = H(α, β) νh .
k
Since dimR M = 2m is even, ∗2 |Ωk(M,C) = (−1) Id and the formal adjoint of
d : Ωk (M, R) → Ωk+1 (M, R) is δ := −∗d∗. By the C-linearity of ∗ and the fact
¯ δ has the C-linear extension
that d is the C-linear extension of ∂ + ∂,
δ = −∗d∗ = −∗ ∂ + ∂¯ ∗ = (−∗∂∗) ⊕ −∗∂∗ ¯ .
 

By (17.145), we have ∗Ωp,q (M ) = Ωm−q,m−p (M ), and so


 
¯ Ωp+1,q (M ) ⊆ −∗∂¯ Ωm−q,m−(p+1) (M )
 
−∗∂∗
⊆ (−∗) Ωm−q,m−p (M ) ⊆ Ωp,q (M ) , while


(−∗∂∗) Ωp+1,q (M ) ⊆ (−∗∂) Ωm−q,m−p−1 (M )


 

⊆ (−∗) Ωm−q+1,m−p−1 (M ) ⊆ Ωp+1,q−1 (M ) .




Thus, for α ∈ Ωp,q (M ) and β ∈ Ωp+1,q (M ), we have


¯ β = H(∂α, β) ,

H(dα, β) = H ∂α + ∂α,
¯ ¯
 
H(α, δβ) = H α, −∗∂∗β − ∗∂∗β = H α, −∗∂∗β , and so
Z Z Z

H(∂α, β) νh = H(dα, β) νh = hC dα, β̄ νh
M
ZM ZM
 
= hC α, δ β̄ νh = hC α, δβ νh
ZM Z M

¯ β νh .
 
= H(α, δβ) νh = H α, −∗∂∗
M M
Similarly, Z Z
¯ β νh =

H ∂α, H(α,(−∗∂∗) β) νh .
M M
Thus, the formal adjoints of ∂ and ∂¯ are given by
¯ and ∂¯∗ = − ∗ ∂∗.
∂ ∗ = −∗∂∗
17.6. THE INDEX THEOREM FOR STANDARD GEOMETRIC OPERATORS 621

Hodge Decomposition and Harmonic Representation. We define the space of


harmonic (0, q)-forms by
H0,q (M ) := α ∈ Ω0,q (M ) : ∂¯ + ∂¯∗ α = 0 .
 

Theorem 17.72 (Hodge Decomposition). There is an orthogonal decomposition


Ω0,q (M ) = H0,q (M ) ⊕ ∂¯ Ω0,q−1 (M ) ⊕ ∂¯∗ Ω0,q+1 (M )
 

= H0,q (M ) ⊕ ∂¯∂¯∗ Ω0,q (M ) ⊕ ∂¯∗ ∂¯ Ω0,q (M ) .


   
(17.153)
Proof. The operator
2
∆∂¯ := ∂¯∂¯∗ + ∂¯∗ ∂¯ = ∂¯ + ∂¯∗ : Ω0,q (M ) → Ω0,q (M )


is elliptic, and hence we have the orthogonal decomposition


Ω0,q (M ) = Ker ∆∂¯ ⊕ ∆∂¯ Ω0,q (M ) .


Moreover, Ker ∆∂¯ = H0,q (M ). Indeed,


α ∈ H0,q (M ) =⇒ ∂¯ + ∂¯∗ α = 0


¯ = 0 and ∂¯∗ α = 0 =⇒ ∂¯∂¯∗ + ∂¯∗ ∂¯ α = 0,



=⇒ ∂α
and conversely if α ∈ Ker ∆∂¯, then
2 2
0 = ∂¯∂¯∗ + ∂¯∗ ∂¯ α, α = ∂α
¯ ¯ ¯ + ∂¯∗ α = 0.
 
+ ∂α =⇒ ∂α
It remains to prove that we have an orthogonal decomposition
∆∂¯ Ω0,q (M ) = ∂¯ Ω0,q−1 (M ) ⊕ ∂¯∗ Ω0,q+1 (M ) .
  
(17.154)
The summands are orthogonal since
¯ ∂¯∗ γ = ∂¯2 β, γ = 0.
β ∈ Ω0,q−1 (M ) and γ ∈ Ω0,q+1 (M ) =⇒ ∂β,
 

Moreover, for any α ∈ Ω0,q (M ), we have


∆∂¯α = ∂¯ ∂¯∗ α + ∂¯∗ ∂α
¯ ∈ ∂¯ Ω0,q−1 (M ) ⊕ ∂¯∗ Ω0,q+1 (M ) ,
   
(17.155)
and so ∆∂¯ Ω0,q (M ) ⊆ ∂¯ Ω0,q−1 (M ) ⊕ ∂¯∗ Ω0,q+1 (M ) .
  

For the reverse inclusion, note that ∂¯ Ω0,q−1 (M ) and ∂¯∗ Ω0,q−1 (M ) are both in
 

H0,q (M ) , since
α ∈ H0,q (M ) =⇒ ∂β,¯ α = β, ∂¯∗ α = 0 = γ, ∂α ¯ = ∂¯∗ γ, α ,
   


and so ∂¯ Ω0,q−1 (M ) ⊕ ∂¯∗ Ω0,q+1 (M ) ⊆ H0,q (M ) = ∆∂¯ Ω0,q (M ) .
  

Note that (17.155) and (17.154) then give the second equality of (17.153). 
Corollary 17.73. Suppose that ∂γ ¯ = 0 for some γ ∈ Ω0,q (M ). There is a
unique α ∈ H (M ), such that for some β ∈ Ω0,q−1 (M ), γ = α + ∂β.
0,q ¯ In other
0,q
words, every cohomology class in H (M ) has a unique harmonic representative.
Proof. Theorem 17.72 yields a unique α ∈ H0,q (M ) such that
¯ + ∂¯∗ β 0
γ = α + ∂β
for some β ∈ Ω0,q−1 (M ) and β 0 ∈ Ω0,q+1 (M ). Now
¯ = ∂α
0 = ∂γ ¯ + ∂¯2 β + ∂¯∂¯∗ β 0 = ∂¯∂¯∗ β 0
2
=⇒ ∂¯∂¯∗ β 0 , β 0 = 0 =⇒ ∂¯∗ β 0 = 0 =⇒ ∂¯∗ β 0 = 0.


622 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

Kähler Manifolds. Suppose that ∇ denotes the covariant derivative for the
Levi-Civita connection of the Riemannian metric h. We have seen that it is always
possible to choose coordinates about a point x ∈ M such that the coordinate
vector fields have vanishing ∇-derivatives at x. However, it is not necessarily  the
1 1 m m
case that such  coordinates can be chosen
 of the form x , y , . . . , x , y , where
z 1 , . . . , z m = x1 + iy 1 , . . . , xm + iy m is a complex-analytic coordinate chart. If
for each x ∈ M such coordinates can be found, then the complex manifold M with
Riemannian metric h is a Kähler manifold. While one can take this to be the
definition of a Kähler manifold, usually one of the other equivalent conditions in
the following theorem is taken to be the definition.
Theorem 17.74. Let M be a complex manifold with complex structure J, and
Riemannian metric h, with Levi-Civita covariant derivative ∇. The following are
equivalent.

I. About each x ∈ M , there is a complex chart x1 + iy 1 , . . . , xm + iy m ,
such that ∇(∂xi ) = ∇ ∂yi = 0 at x.
II. ∇J = 0 (i.e., (∇J)(X) = J(∇X) − ∇(J(X)) = 0.
III. The Kähler 2-form κ ∈ Ω2 (M, R), given by κ(X, Y ) := h(X, JY ), is
closed, i.e., dκ = 0.
Proof. Assuming I, at any point x we have II, since

(∇J)(∂xk ) = J(∇∂xk ) − ∇(J(∂xk )) = 0 − ∇ ∂yk = 0 and
  
(∇J) ∂yk = J ∇∂yk − ∇ J ∂yk = 0 + ∇(∂xk ) = 0
with the induced Hermitian metric H on C ⊗R T M defined by (17.152), and the
Levi-Civita covariant derivative ∇. We now show that II ⇒ III. For vector fields
X, Y, Z, we have

3(dκ)(X, Y, Z) = X [κ(Y, Z)] + Y [κ(Z, X)] + Z [κ(X, Y )]


− κ([X, Y ] , Z) − κ([Y, Z] , X) − κ([Z, X] , Y ) .
If X, Y, Z are coordinate vector fields, then (since ∇ is torsion-free)
∇X Y − ∇Y X = [X, Y ] = 0, and so

3(dκ)(X, Y, Z) = X [κ(Y, Z)] + Y [κ(Z, X)] + Z [κ(X, Y )]


= X [h(Y, JZ)] + Y [h(Z, JX)] + Z [h(X, JY )]
= h(∇X Y, JZ) + h(Y, ∇X (JZ)) + h(∇Y Z, JX) + h(Z, ∇Y (JX))
+ h(∇Z X, JY ) + h(X, ∇Z (JY ))
= h(∇X Y, JZ) − h(∇X Z, JY ) + h(∇Y Z, JX) − h(∇Y X, JZ)
+ h(∇Z X, JY ) − h(∇Z Y, JX)
= h(∇X Y − ∇Y X, JZ) + h(∇Z X − ∇X Z, JY ) + h(∇Y Z − ∇Z Y, JX) = 0.
To show that III ⇒ I, we note that by a C-linear change of complex coordinates
about x, we may assume (since h is positive definite) that
X
hC = 2 (hC )j k̄ dz j ⊗s dz̄ k
j,k
X  X  
2
aj k̄l z l + aj k̄l̄ z̄ l + O kzk dz j ⊗s dz̄ k ,

= δjk +
j,k,l l
17.6. THE INDEX THEOREM FOR STANDARD GEOMETRIC OPERATORS 623

2 Pm
where kzk = j=1 z j z̄ j and

(hC )j k̄ = (hC )j̄k = (hC )kj̄


X X X
aj k̄l z̄ l + aj k̄l̄ z l = akj̄l z l + akj̄ l̄ z̄ l

=⇒ aj k̄l z l + aj k̄l̄ z̄ l =
l l l
(17.156) =⇒ aj k̄l̄ = akj̄l and aj k̄l̄ = akj̄l .
Since
κ(∂zj , ∂z̄k ) = hC (∂zj , J∂z̄k ) = hC (∂zj , −i∂z̄k ) = −ihC (∂zj , ∂z̄k ) = −i(hC )j k̄ ,
and similarly κ(∂zj , ∂zk ) = κ(∂z̄j , ∂z̄k ) = 0, we have
Xm Xm
κ = 21 κ(∂zj , ∂z̄k ) dz j ∧ dz̄ k = − 2i (hC )j k̄ dz j ∧ dz̄ k .
j,k=1 j,k=1

Thus,
Xm  
0 = dκ = − 2i ∂zl (hC )j k̄ dz l + ∂z̄l (hC )j k̄ dz̄ l ∧ dz j ∧ dz̄ k
j,k=1
Xm
= − 2i aj k̄l dz l + aj k̄l̄ dz̄ l ∧ dz j ∧ dz̄ k

j,k=1
=⇒ aj k̄l = alk̄j and aj k̄l̄ = aj l̄k̄ .
For bkpq = bkqp ∈ C, let
Xm Xm
z k = wk + 12 bkpq wp wq and dz k = dwk + bkpq wq dwp .
p,q=1 p,q=1
2 2
Then, modulo O kwk or O kzk terms,
X
δjk + aj k̄l z l + aj k̄l̄ z̄ l dz j ⊗s dz̄ k

hC =
j,k
X  X 
= δjk + aj k̄l wl + aj k̄l̄ w̄l
j,k l
 Xm   Xm 0 0

j
· dw + bjpq wq dwp ⊗s dw̄k + bkp0 q0 w̄q dw̄p
p,q=1 p,q=1
X  X 
l l
= δjk + aj k̄l w + aj k̄l̄ w̄
j,k l
Pm
dwj ⊗s dw̄k + p,q=1 bjpq wq dwp ⊗s dw̄k
 
· m 0 0
+ p,q=1 bkp0 q0 w̄q dwj ⊗s dw̄p
P
X  X 
aj k̄l wl + aj k̄l̄ w̄l dwj ⊗s dw̄k

= δjk +
j,k l
X Xm
+ δjk bjpq wq dwp ⊗s dw̄k
j,k p,q=1
X Xm 0 0
+ δjk bk0 0 w̄q dwj ⊗s dw̄p
j,k p,q=1 p q
X  X 
= δjk + aj k̄l wl + aj k̄l̄ w̄l dwj ⊗s dw̄k
j,k l
Xm Xm
+ bkjl wl dwj ⊗s dw̄k + bj w̄l dwj ⊗s dw̄k
j,k,l=1 j,k,l=1 kl
X  X   
aj k̄l + bkjl wl + aj k̄l̄ + bjkl w̄l dwj ⊗s dw̄k .

= δjk +
j,k l

Thus, we should (if possible) choose

bkjl = −aj k̄l and bjkl = −aj k̄l̄ (i.e., bkjl = −akj̄ l̄ ).
624 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

We can do this provided that


1. aj k̄l = alk̄j , so that bkjl is symmetrical in j and l, and
2. aj k̄l = akj̄ l̄ , so that bkjl = −aj k̄l =⇒ bjkl = −aj k̄l̄ .
But this is the case, since we have seen that, by (17.156),
dκ = 0 =⇒ aj k̄l = alk̄j and akj̄ l̄ = ak̄jl = aj k̄l . 

Useful Formulae for Complex Partial Derivatives. Until further notice, we as-
sume that M (with Riemannian metric  h) is a Kähler manifold with complex co-
ordinates x1 + iy 1 , . . . , xm + iy m about  some point x ∈ M such that the par-
tial derivatives ∂x1 , ∂y1 , . . . , ∂xm , ∂ym are orthonormal at x relative to hx , and
∇(∂xi ) = ∇ ∂yi = 0 at x. We find a formula for ∂¯∗ : Ω0,q (M ) → Ω0,q (M ) which

will allow us to exhibit 2 ∂¯ + ∂¯∗ locally as a twisted Dirac operator. In view of


Theorem 17.74, for α ∈ Ω0,q (M ), we have


Xm
¯ =
∂α dz̄ j ∧ ∇∂z̄j α at x ∈ M ,
j=1
  
since ∇∂z̄j dz̄ k (∂z̄l ) = ∂z̄j dz̄ k (∂z̄l ) − dz̄ k ∇∂z̄j ∂z̄l = 0 at x implies that

∇∂z̄j αk1 ···kq dz̄ k1 ∧ · · · ∧ dz̄ kq = ∂z̄j αk1 ···kq dz̄ k1 ∧ · · · ∧ dz̄ kq at x.
 

We derive a formula for ∂¯∗ , namely for α ∈ Ω0,q (M ),


Xm
∂¯∗ α = − dz̄ j x∇∂z̄j α at x.

(17.157)
j=1

We use the result (17.145), namely

(17.158) ∗ ξ(i) ∧ ξ¯(j)



h k
h 2
= 2h+k−m (−1) im ε(ic )m−h(i)h ε(j)k(j c )m−k ξ(j c )m−k ∧ ξ¯(ic )m−h ,
where (j)k denotes the multi-index (j1 , . . . , jk ) with 1 ≤ j1 < · · · < jk ≤ m, and
(j c )m−k denotes the multi-index complementary to (j)k in the sense that
{j1 , . . . , jk } ∪ j1c , . . . , jm−k
c

= {1, . . . , m}
and 1 ≤ j1c < · · · < jm−k
c
< m. Note that (17.158) implies
    2 c c

∗ dz h ∗ dz̄ (j)k = ∗ dz q ∧ 2k−m im ε(j)k(j c )m−k dz (j )m−k ∧ dz̄ (i )m
2
 c c

= 2k−m im ε(j)k(j c )m−k ∗ dz q ∧ dz (j )m−k ∧ dz̄ (i )m ,
 
which is zero unless h is jp for some jp in (j)k . In the following, jbp , j denotes
k−1
(j)k , but with jp removed, and (jp , j c )m−k+1 denotes the multi-index (j c )m−k with
jp inserted in the correct position which produces an increasing order. Note that
k−p jp j c
(17.159) ε(jbp ,j ) (jp ,j c ) = (−1) δ(jp ,j c ) ε(j)k(j c )m−k ,
k−1 m−k+1

since jp is moved through k − p indices in (j)k (j c )m−k to produce


 
jbp , j jp (j c )m−k+1 ,
k−1
17.6. THE INDEX THEOREM FOR STANDARD GEOMETRIC OPERATORS 625

j jc
and jp is moved through δ(jpp ,j c ) indices in jp (j c )m−k+1 to produce (jp , j c )m−k+1 .
We use (17.159) in the following computation
  
∗ dz jp ∗ dz̄ (j)k
2
 c c

= 2k−m im ε(j)k(j c )m−k ∗ dz jp ∧ dz (j )m−k ∧ dz̄ (i )m
j jc
2
 c c

= 2k−m im ε(j)k(j c )m−k δ(jpp ,j c ) ∗ dz (jp ,j )m−k+1 ∧ dz̄ (i )m
2 j jc
= 2k−m im ε(j)k(j c )m−k δ(jpp ,j c ) 2(m−k+1)+m−m

dz̄ (
m−k+1 m2 jbp ,j )
· (−1) i ε(jbp ,j ) (jp ,j c ) k−1
k−1 m−k+1

m2 m−k+1 j jc (jbp ,j )k−1


= 2(−1) (−1) ε(j)k(j c )m−k δ(jpp ,j c ) ε(jbp ,j ) (jp ,j c )m−k+1 dz̄
k−1

k−1 j jc (jbp ,j )k−1


= 2(−1) ε(j)k(j c )m−k δ(jpp ,j c ) ε(jbp ,j ) (jp ,j c )m−k+1 dz̄
k−1

dz̄ ( p )k−1 = 2(−1) dz̄ ( p )k−1 = dz̄ jp x dz̄ (j)k .


k−1 k−p jb ,j p−1 jb ,j
= 2(−1) (−1)

It then follows that for β ∈ Ω1,0 (M ) and α ∈ Ω0,q (M ) ,

∗(β∧ ∗ α) = β̄ x α.
1 k1
∧ · · · ∧ dz̄ kq and formally substituting ∂ =
P
Thus, with α = q! (k) αk1 ···kq dz̄
k0
P
k0 ∂z k0 dz ∧ for β∧ , we have

∂¯∗ α = −∗(∂ ∗ α) = −∂¯ x α


X  1 X
dz̄ k0 ∂z̄k0 x αk1 ···kq dz̄ k1 ∧ · · · ∧ dz̄ kq

=−
k0 q! (k)
1 X
∂z̄k0 αk1 ···kq dz̄ k0 x dz̄ k1 ∧ · · · ∧ dz̄ kq .
  
=−
q! k0 ,(k)

At x ∈ M (where the ∂z̄k are parallel), we then have (17.157).


An Inspiring Fake Proof of the Hirzebruch-Riemann-Roch Theorem. Ideally,
we would
√ like to∗show that Ω0,∗ (M ) is a Clifford module bundle W (M ) and that
¯ ¯
W

D = 2 ∂ + ∂ . We could then obtain the Hirzebruch-Riemann-Roch Theorem
as a consequence of the Index Theorem for Generalized Dirac Operators (Theorem
17.59, p.599). However, there are technical details of great consequence which stand
in the way of doing this, as we now explain.
Note that
1,0
Cm → (C ⊗ Cm ) ⊂ C ⊗ Cm given by
√ √ 1 1
v 7→ 2 v 0,1 := 2 (v + iJv) = √ (v + iJv)
2 2
is an isometry. Indeed, if h·, ·im denotes the standard Hermitian inner product on
Cm ,
√ √ √ √
H(1/ 2 (v + iJv) , 1/ 2 (v + iJv)) = h1/ 2 (v + iJv) , 1/ 2 (v − iJv)i
1
= 2 (hv, vi + hJv, Jvi) = hv, vi .
626 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

∗
We take W := Λ0,∗ (C ⊗ Cm ) ∼ = Λ• (Cm ) and let Q : Cln → End(W ) be deter-
mined by
1
Q1 (v) α = 2 2 [ v 0,1 ∧ α − [ v 0,1 xα ,
  
∗
for v ∈ Rn = Cm . Let q : U(m) → U(W ) = Λ0,∗ (Cm ) denote the usual represen-
∗ 
tation induced by q1 : U(m) → U Λ0,1 (Cm ) which we describe as follows. For
A ∈ U(m) ⊂ SO(2m), we have the C-linear extension Ac ∈ U(2m) which preserves
the summands of the decomposition
1,0 0,1
C ⊗R Cm = (C ⊗R Cm ) ⊕(C ⊗R Cm ) .
b ∈ U (Cm )∗ given by A(ϕ) = ϕ ◦ A−1 , and its C-linear extension

We also have A b
2m
∗  m ∗ 
bc ∈ U C
A = U (C ⊗R C ) which preserves the summands of
∗ ∗1,0 ∗0,1
(C ⊗R Cm ) = (C ⊗R Cm ) ⊕(C ⊗R Cm ) .
For A ∈ U(m) , we take
 
bc )0,1 ∈ End (C ⊗R Cm )∗0,1 = End Λ0,1 (C ⊗R Cm )∗ ,

q1 (A) = (A
∗
and let q(A) denote the natural extension of q1 (A) to all of Λ0,∗ (C ⊗R Cm ) . For
A ∈ U(m) and v ∈ Cm , we have
q(A) ◦ Q(v) ◦ q A−1 = Q(Av) .


∗ ∗
Indeed for α ∈ W = Λ0,∗ (Cm ) = Λ0,∗ (C ⊗R Cm ) ∼ = Cln ,
q(A) ◦ Q(v) ◦ q A−1 α = q(A) Q(v) q A−1 α
  
1
= 2 2 q(A) [ v 0,1 ∧ q A−1 (α) − [ v 0,1 x q A−1 (α)
     
1
= 2 2 q(A) [ v 0,1 ∧(α) − q(A) [ v 0,1 xα
 

1
   
0,1 0,1
= 2 2 [ (Av) ∧ α − [ (Av) xα
= Q(Av)(α) .
Hence for A ∈ U(m) and α ∈ Cln , we have
q(A) Q(α) q A−1 = Q(r(A) α) .


In view of this, there are well-defined


c : Cl(T M ) ⊗ W (M ) → W (M ) ,
c : Cl(T M ) ⊗ Λ0,∗ (TC M ∗ ) → Λ0,∗ (TC M ∗ ) .
Moreover, c : Cl(T M ) ⊗ W (M ) → W (M ) induces maps on sections,
c : C ∞ (Cl(T M )) ⊗ C ∞ (W (M )) → C ∞ (W (M )) ,
c : C ∞ (Cl(T M )) ⊗ Ω0,∗ (M ) → Ω0,∗ (M ) .
For α ∈ C ∞ (Cl(T M )) and ψ ∈ C ∞ Ω0,∗ (M ) , firstly, we have


∇0,1 (c(α ⊗ ψ)) = c(∇α ⊗ ψ) + c α ⊗ ∇0,1 ψ ,



(17.160)
where
∇0,1
∇0,1 : Ω0,∗ (M ) → Ω1 (M ) ⊗ Ω0,∗ (M )
17.6. THE INDEX THEOREM FOR STANDARD GEOMETRIC OPERATORS 627

is just the complex extension of the Levi-Civita covariant derivative which has the
property that
ψ ∈ Ω0,k (M ) =⇒ ∇X ψ ∈ Ω0,k (M ) ,
and, secondly, the generalized Dirac operator
∇0,1 c
DW : Ω0,∗ (M ) → Ω1 (M ) ⊗ Ω0,∗ (M ) → Ω0,∗ (M ) ,

which is found to be 2 ∂¯ + ∂¯∗ . However, the problem is that


∗
W (M ) = Λ0,∗ (TC M ∗ ) = U M ×U(m) Λ0,∗ (C ⊗R Cm )
is an associated bundle of U(M ) instead of F M or P for a spin structure P → F M ,
as is required in the proof of Theorem 17.59, p. 599; otherwise, it is not clear that
DW is locally a twisted Dirac operator. Let us assume, for the moment, that this
does not matter. Then Theorem 17.59 would yield
  
index ∂¯ + ∂¯∗ = index DW + = ch Λ0,∗ (TC M ∗ ) ` A(M
 
e ) [M ] .

Now, identifying ck (T M ) with σk (y1 , . . . , ym ), we have


 Xm X Ym
ch Λ0,∗ (TC M ∗ ) = eyj1 +···+yjh = (1 + eyk ) , and
h=0 (j)h k=1

Ym Ym yk /2
ch Λ0,∗ (TC M ∗ ) ` A(M (1 + eyk )

e )=
k=1 k=1 sinh(yk )
Ym yk /2
Ym yk
= (1 + eyk ) sinh(y ) =
k=1 k k=1 1 − e−yk

= Td T 1,0 (M ) = Td(T M ) , and so



(17.161)
index ∂¯ + ∂¯∗ = Td(T M ) [M ] .


This is correct and indeed it is the Hirzebruch-Riemann-Roch formula, but since


W (M ) = Λ0,∗ (TC M ∗ ) is associated to U(M ) instead of F M or a spin structure
bundle P , Theorem 17.59 does not apply. This suggests that the notion of Clifford
module bundle can be generalized in such a way that Theorem 17.59 still holds. In
fact this is the case, and we will examine the case at hand to determine a suitable
generalization.
To fix the fact that Λ0,∗ (TC M ∗ ) is associated to U(M ), instead of F M or P ,
ideally one would like to find a homomorphism
η : U(m) → SO(2m) ,
a representation r : SO(2m) → GL(V ), and an equivariant isomorphism
∗
φ : Λ0,∗ (C ⊗ Cm ) → V.
Then we would have Λ0,∗ (TC M ∗ ) ∼= F M ×SO(2m) V , exhibiting Λ0,∗ (TC M ∗ ) as a
bundle associated with F M . It is natural to take η to be the inclusion and φ
∗
to be the identity. However, Λ0,∗ (C ⊗ Cm ) is not an invariant subspace of the

representation SO(2m) → Λ• (C ⊗ Cm ) , but rather of its restriction to U(m).
While we will not prove the fact that there is no suitable choice r and φ, at least
the most obvious attempt fails. The somewhat involved resolution of this difficulty
is doubly worthwhile, since it also provides is a nice way of motivating Seiberg-
Witten theory. In a nutshell, we will exploit the fact that while
∗ 
Λ0,∗ : U(m) → End Λ0,∗ (C ⊗ Cm )
628 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

does not extend to a representation of SO(2m), we have linear isomorphisms


∗ ∼ ∼
Λ0,∗ (C ⊗ Cm ) −→ Λ• (Cm ) −→ Σ2m ,
and there is the spinor representation ρ : Spin(2m) → End(Σ2m ). However on the
Lie algebra level, the restriction of ρ0 : spin(2m) → End(Σ2m ) to u(m) ⊂ s0(2m) ∼ =
0 ∗ 
spin(2m) does not coincide with Λ0,∗ : u(m) → End Λ0,∗ (C ⊗ Cm ) . Indeed,
there is a crucial 1-dimensional twist which is needed to make a correct identifica-
tion. We proceed with the details.
A Crucial 1-Dimensional Twist.
√ Since we will be considering C ⊗ Cm , we√use
J to denote multiplication by −1 in Cm and i to denote multiplication by −1
∗ 0,1
in C ⊗ Cm . Note that the map [0,1 : Cm → (C ⊗ Cm ) given by [0,1 (v) =
1
2− 2 [(v + iJv) is a C-linear isometry, since it is clearly an R-linear isometry and
1 1 1
[0,1 (Jv) = 2− 2 [(Jv − iv) = 2− 2 [(−i(v − iJv)) = i2− 2 [(v − iJv) = i[0,1 (v) .
Thus, [0,1 extends to a C-linear isometry
∼ ∗
Λ [0,1 : Λ• (Cm ) −→ Λ0,∗ (C ⊗ Cm ) .


Recall that Λ• (Cm ) provides a representation space for an irreducible Cl2m -module
ρC : Cl2m → End(Λ• (Cm )) .
Here ρC denotes the C-linear extension of
ρ : C`2m → End(Λ• (Cm )) ,

which was described as follows. Let R2m −→ Cm denote the usual isomorphism
(x1 , y1 , . . . , xm , ym ) = (x1 + iy1 , . . . , xm + iym ) .
m
For w ∈ C , define

ρ1 : R2m −→ Cm → End(Λ• (Cm )) by
ρ1 (w)(α) := (w∧ − wx)(α) = w∧ α − wxα.
We verified that ρ1 extends to ρ : C`2m → End(Λ• (Cm )). In this way one may view
the space Σ2m of spinors concretely as Λ• (Cm ). There is then the representation
ρ|Spin(n) : Spin(n) → SU(Σ2m ) .

m
 2m → End(Λ (C )) tα
Since ρ : C` is an algebra homomorphism, for α ∈ spin(n) =
L Λ (C ) ⊂ C`2m , we have ρ(e ) = etρ(α) , and so
2 m

0  
d
ρ etα |t=0 = dt
d
etρ(α) |t=0 = ρ(α) .

ρ|Spin(n) (α) = dt
Thus,
0
ρ|Spin(n) = ρ|spin(n) = ρ|L(Λ2(Cm )) .
Recall that we also have the 2-fold covering c : Spin(n) → SO(n). Since there is a
natural injection
ι : U(m) → SO(2m) ,
at first one might guess that on the level of Lie algebras, the action of u(m) ⊂
so(2m) ∼ = spin(2m) on Λ• (Cm ) given by
0
ρ|Spin(n) : u(m) → End(Λ• (Cm ))
0
coincides with the usual (Λ• ) : u(m) → End(Λ• (Cm )). Recall that Λ• : U(m) →
End(Λ• (Cm )) denotes the usual representation induced by the defining (identity)
17.6. THE INDEX THEOREM FOR STANDARD GEOMETRIC OPERATORS 629

representation U(m) → U(Cm ). Explicitly, for A ∈ U(m) and 0 ≤ l ≤ m, we have


Λl (A) ∈ End Λl (Cm ) given by

Λl (A)(fj1 ∧ fj2 ∧ · · · ∧ fjl ) := Afj1 ∧ Afj2 ∧ · · · ∧ Afjl ,


0 0
and this extends to Λ• : U(m) → End(Λ• (Cm )). Although ρ|Spin(n) 6= (Λ• ) , the

two representations are related in a way which will ultimately exhibit 2 ∂¯ + ∂¯∗


at least locally as a twisted generalized Dirac operator, thereby


0 justifying the com-
0
putation (17.161) via forms. We now describe how ρ|Spin(n) and (Λ• ) are related.
For λ = (λ1 , . . . , λm ) ∈ Rm , let Bλ ∈ u(m) denote the diagonal matrix

Bλ := diag(iλ1 , . . . , iλm ) .
0
The endomorphism (Λ• ) (Bλ ) ∈ End(Λ• (Cm )) leaves each of the Λl (Cm ) (l =
0 
1, · · · , m) fixed. Explicitly, Λl (Bλ ) ∈ End Λl (Cm ) (induced by Bλ ∈ End(Cm ))
is given simply by

Λl (Bλ )(fj1 ∧ fj2 ∧ · · · ∧ fjl ) = i(λj1 + · · · + λjl )(fj1 ∧ fj2 ∧ · · · ∧ fjl ) .


0
We wish to compare (Λ• ) (Bλ ) with
0
ρ|Spin(n) c0−1 (ι(Bλ )) ∈ End(Σ2m ) = End(Λ• (Cm )) .


Recall that for A = aij ∈ so(n), the Lie algebra isomorphism c0 : spin(2m) =

 ∼
L Λ2 R2m −→ so(n) for the covering c : Spin(2m) → SO(2m) is given by
X
c0 − 41

aij ei ej := A.
i,j

For ι : U(m) → SO(n) , the associated Lie algebra map ι0 : u(m) → so(n) applied to
Bλ is given by

ι0 (Bλ )(e2j−1 ) = λj e2j and ι0 (Bλ )(e2j ) = −λj e2j−1 , j = 1, . . . , m, so that


ι0 (Bλ )2j,2j−1 = λj and ι0 (Bλ )2j−1,2j = −λj .

Then
 Xn   Xm 
ι0 (Bλ ) = c0 − 41 ι(Bλ )ij ei ej = c0 − 21 ι(Bλ )2j−1,2j e2j−1 e2j
i,j=1 j=1
 Xm 
0 1
= c 2 λj e2j−1 e2j .
j=1

We have shown (see (17.15), p. 521) that the action of ρ(e2j−1 e2j ) on fj1 ∧ fj2 ∧
· · · ∧ fjl is given by

ρ(e2j−1 e2j )(fj1 ∧ fj2 ∧ · · · ∧ fjl )



i(fj1 ∧ fj2 ∧ · · · ∧ fjl ) , if jk = j for some k,
=
−i(fj1 ∧ fj2 ∧ · · · ∧ fjl ) , if jk 6= j for all k,
X 
l j
= 2i δjk (fj1 ∧ fj2 ∧ · · · ∧ fjl ) − i(fj1 ∧ fj2 ∧ · · · ∧ fjl ) , and so
k=1
630 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

0
ρ|Spin(n) c0−1 (ι0 (Bλ )) (fj1 ∧ fj2 ∧ · · · ∧ fjl )


 X m 
= ρ 21 λj e2j−1 e2j (fj1 ∧ fj2 ∧ · · · ∧ fjl )
j
Xm Xl  Xm 
=i λj δjjk (fj1 ∧ fj2 ∧ · · · ∧ fjl ) − i 21 λj (fj1 ∧ fj2 ∧ · · · ∧ fjl )
j=1 k=1 j=1
0
 Xm 
= Λl (Bλ )(fj1 ∧ fj2 ∧ · · · ∧ fjl ) − i 12

λj (fj1 ∧ fj2 ∧ · · · ∧ fjl )
j=1
 0  Xm 
= Λl (Bλ ) − i 21 λj (fj1 ∧ fj2 ∧ · · · ∧ fjl ) .
j=1

Since this holds for arbitrary l,


 Xm 
0
ρ c0−1 (ι0 (Bλ )) = − 21 i λj +(Λ• ) (Bλ )

j=1
0
= − 12 Tr(Bλ ) Id +(Λ• ) (Bλ ) , or
0
(Λ• ) (Bλ ) = 21 Tr(Bλ ) Id +ρ c0−1 (ι0 (Bλ )) .

(17.162)
0
Although (Λ• ) (Bλ ) and ρ c0−1 (ι0 (Bλ )) are different, this formula provides a def-

• 0
inite comparison.
More specifically, while (Λ ) does not have values in ρ(spin(2m))
2 2m
=ρ L Λ R , it does have values in the slightly larger space ρC (iR ⊕ spin(2m)).
Note that iR ⊕ spin(2m) is the Lie algebra of
U(1) Spin(2m) := {zg : g ∈ Spin(2m) , z ∈ C, |z| = 1} ⊂ Cl(2m) .
This suggests (as we will show) that possibly there is a homomorphism
0
j : U(m) → U(1) Spin(2m) such that ρC ◦ j 0 = (Λ• ) : u(m) → End(Λ• (Cm )) .
If so, then by (17.162),
j 0 (Bλ ) = ρ−1 •0
C ◦ Λ (Bλ ) =
1
2 Tr(Bλ ) + c0−1 (ι0 (Bλ )) ⊆ iR ⊕ spin(2m) .
From this, it can be conjectured that for A ∈ U(m) with an orthonormal basis of
eigenvectors v1 , . . . , vm and eigenvalues eiθ1 , . . . , eiθm , one can define
 Xm  Ym
j(A) = exp 2i cos 12 θk + sin 12 θk vk Jvk .
 
θk
k=1 k=1

However, there is a problem (not impossibly difficult) with showing that this yields
a well defined, continuous homomorphism, independent of the choice of the vk .
Alternatively, we obtain j as follows. Note that the homomorphism
(17.163) χ : U(1) × Spin(2m) → U(1) Spin(2m) given by χ(z, σ) := zσ
is onto with kernel {±(1, 1)}, since
zg = 1 =⇒ g = z −1 =⇒ g ∈ R ∩ U(1) =⇒ g = ±1 =⇒ (z, g) = ±(1, 1) .
Thus, we are motivated to define
U(1) × Spin(n)
(17.164) Spinc (n) := U(1) Spin(2m) ∼
= .
{±(1, 1)}
Often Spinc (n) is defined to be the group on the right. This has the advantage
of making it clear that Spinc (n) is not U(1) × Spin(n), and our definition exhibits
17.6. THE INDEX THEOREM FOR STANDARD GEOMETRIC OPERATORS 631

Spinc (n) as a subgroup of Cl(2m), with Lie algebra iR ⊕ spin(2m). There is a


double-covering homomorphism (with kernel {1, −1})
Sq × c : Spinc (n) → U(1) × SO(n) given by
(Sq × c)(zg) := z 2 , c(g) ,

(17.165)
 
2 
which is well-defined since (−z) , c(−g) = z 2 , c(g) .

Proposition 17.75. There exists a lifting homomorphism j : U(m) → Spinc (n),


such that det × ι = j ◦(Sq × c); i.e., the diagram
(17.166)
Spinc (n)
ρc
/ U(Σ2m ) = U(Λ• (Cm ))
5
j
Sq×c

U(m)
det × ι
/ U(1) × SO(n)

commutes. Moreover,
ρC ◦ j = Λ• : U(m) → U(Λ• (Cm )) .
Proof. The existence of a continuous map j : U(m) → Spinc (n) with det ×ι =
j ◦(Sq × c) will be deduced from covering space theory by showing
(det × ι)# (π1 (U(m) , I)) ⊆ (Sq × c)# (π1 (Spinc (n) , 1)) .
We have
π1 (U(m) , I) ∼
= Z, π1 (Spinc (n) , 1) ∼
= Z, and

π1 (U(1) × SO(n) ,(1, I)) = Z × Z2 .
The generator, say gU(m) , of π1 (U(m) , I) is represented by U(1) → U(m) , given by
eiθ 7→ diag eiθ , 1, . . . , 1 , and the generator gSO(n) of π1 (SO(n) , I) is represented
by  
iθ cos θ − sin θ
e 7→ ⊕ I2n−2 .
sin θ cos θ
By means of gU(m) and  gSO(n) , we identify
 π1 (U(1) × SO(n) ,(1, I)) = Z × Z2 . We
have (det × ι)# gU(m) = gU(1) , gSO(n) , since
   

 iθ cos θ − sin θ
(det × ι) diag e , 1, . . . , 1 = e , ⊕ I2n−2 ,
sin θ cos θ
The Group of Covering Transformations. The group G of covering transfor-
mations of the covering Sq × c : Spinc (n) → U(1) × SO(n) consists of Id and the
map zg 7→ −zg. The image, say N , of π1 (Spinc (n) , 1) under the monomorphism
(Sq × c)# is then a subgroup of π1 (U(1) × SO(n) ,(1, I)) = Z × Z2 . Note that N is
normal since Z × Z2 is abelian. From standard covering space theory we then have
Z2 ∼
=G∼ = (Z × Z2 ) /N . There are several subgroups of Z × Z2 of index 2, namely
Z× {0}, 2Z × Z2 , and
(17.167) N = {(p, p mod 2) : p ∈ Z}

which is generated by (1, 1) = gU(1) , gSO(n) . To prove (17.167), it suffices to
observe that
eiθ 7→ eiθ/2 cos 21 θ + sin 12 θ e1 e2 , 0 ≤ θ ≤ 2π
  
632 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

is a loop in Spinc (n) projecting via Sq × c to a representative of (1, 1). This loop
then represents the generator, say gc , of π1 (Spinc (n) , 1) ∼
=N ∼= Z. Now
  
h
iθ/2 1
 1
 i
iθ θ θ
(Sq × c) e , cos 2 θ + sin 2 θ e1 e2 = e , c cos + sin e1 e2
2 2
   
cos θ − sin θ
= eiθ , ⊕ I2n−2 = (det × ι) diag eiθ , 1, . . . , 1

sin θ cos θ

=⇒ (det × ι)# gU(m) = (Sq × c)# (gc )
=⇒ (det × ι)# (π1 (U(m))) = (Sq × c)# (π1 (Spinc (n))) .
Covering space theory then provides us with a topological lifting j in the diagram
(17.166) which sends I ∈ U(m) to 1 ∈ Spinc (n). We prove that j is a homomor-
phism as follows. Since det × ι is a homomorphism and Sq × c is a two-fold cover,
for all A1 , A2 ∈ U(m), we have
−1 −1
j(A1 A2 ) j(A2 ) j(A1 ) = ±1,
but the left side is a continuous function of (A1 , A2 ) which is 1 at (A1 , A2 ) =
−1 −1
(I, I) , whence j(A1 A2 ) j(A2 ) j(A1 ) = 1 for all (A1 , A2 ) in the connected space
U(m) × U(m). Of course, the lift j of the C ∞ map det × ι is smooth, since the
covering Sq × c is a local diffeomorphism. That ρC ◦ j = Λ• : U(m) → U(Λ• (Cm ))
can be deduced as follows. Since
 
0 0−1 0
(ρC ◦ j) (Bλ ) = ρC ◦(Sq × c) ◦(det × ι) (Bλ )
0−1
= ρC ◦(Sq × c) (Tr(Bλ ) , ι0 (Bλ ))
= ρC 21 Tr(Bλ ) + c0−1 (ι0 (Bλ ))


0
= 21 Tr(Bλ ) + ρ c0−1 (ι0 (Bλ )) = (Λ• ) (Bλ ) ,


0
by (17.162), (ρC ◦ j) and Λ•0 agree on the (Cartan) subalgebra t := {Bλ : λ ∈ Rm } ⊂
u(m) which is the Lie algebra of the (maximal) torus
T := diag eiλ1 , . . . , eiλm : λ ∈ Rm ⊂ U(m) .
 

Thus, Λ• |T = (ρC ◦ j) |T . Those sufficiently familiar with the theory of representa-


tions of compact Lie groups will then conclude that Λ• = ρC ◦ j, but for those less
familiar with the theory, we offer the following. Since for every A ∈ U(m) there is
B ∈ U(m) such that BAB −1 ∈ T , we have
 
−1
Tr(Λ• (A)) = Tr Λ• (B) Λ• (A) Λ• (B) = Tr Λ• BAB −1


= Tr (ρC ◦ j) BAB −1 = Tr((ρC ◦ j)(A)) .




Hence the characters Tr ◦(ρC ◦ j) and Tr ◦Λ• of the representations ρC ◦j and Λ• are
the same, which implies (see 4.10.2, p. 107 of [430]) that ρC ◦j and Λ• are equivalent.
Thus ρC ◦j and Λ• differ by a constant multiple on each of the irreducible subspaces
Λl (Cm ). But the constant multipliers are all 1, since ρ◦j and Λ• agree on T . Hence,
Λ• = ρC ◦ j; i.e., we have
ρC
Spinc (n) −→ U(Λ• (Cm )) = U(Σ2m )
(17.168) ↑j ↓ Id 
Λ• • m
U(m) −→ U(Λ (C )) .
17.6. THE INDEX THEOREM FOR STANDARD GEOMETRIC OPERATORS 633

The Canonical Line Bundle. Let M be a complex manifold of complex dimen-


sion m with a Riemannian metric h which makes M a Kähler manifold. The homo-
morphism j : U(m) → Spinc (n) yields a Spinc (n)-bundle PSpinc(n) → M, namely
PSpinc(n) := U(M ) ×U(m) Spinc (n) , and a bundle morphism
fj : U(M ) → PSpin(n) (namely, u 7→ [u, 1] ),
which is equivariant in the sense that fj (uA) = fj (u) j(A). The bundle Λ0,n (TC M ∗ )
of complex dimension 1 is known as the canonical line bundle for M . Let PU(1)
denote the unitary frame bundle for Λn,0 (TC M ). Then
∼ Λn,0 (TC M ) = PU(1) ×U(1) C.
Λ0,n (TC M ∗ ) =
There is an equivariant bundle morphism ∧n : U(M ) → PU(1) given by ∧n (u) 7→
u(e1 )∧· · ·∧u(em ), where e1 , . . . , em denotes the standard basis for Cm and u : Cm →
T
√M is a unitary frame. Since we will find that it is only necessary to exhibit
2 ∂ + ∂¯ locally as a twisted Dirac operator, until further notice, we now hypo-
thesize (as is always locally the case)
Assumptions 17.76. There are
0
1. a spin structure PSpin(n) → F M and principal U(1)-bundle PU(1) , and
0
2. an equivariant bundle map πχ : PU(1) ×f PSpin(n) → PSpinc(n) , such that
πχ ((p0 , p) g) = πχ ((p0 , p)) χ(g) for all g ∈ U(1) × Spin(n) .
We then have a commutative diagram of bundle morphisms
0
PU(1) ×f PSpin(n)
πχ

fj 
U(M ) / PSpinc(n)
fdet × ι
∧n πSq×c
 ) 
PU(1)
Id ×In
/ PU(1) ×f F M

equivariant with respect to the following diagram of group homomorphisms:


U(1) × Spin(n)
χ

U(m)
j
/ Spinc (n)
det × ι
det Sq×c
 ) 
U(1)
Id ×I
/ U(1) × SO(n) .

Since ((Sq × c) ◦ χ)(z, g) = (Sq × c)(zg) = z 2 , c(g) , under the composition πSq×c ◦
πχ , a point (p0 , p) ∈ PU(1)
0
×f PSpin(n) is sent to a point (q 0 , q) ∈ PU(1) ×f F M where
0
q is independent of the choice of p. Thus, πSq×c ◦ πχ determines a bundle map, say
0
πs : PU(1) → PU(1) .
Moreover, since πs (p0 z) = πs (p0 ) z 2 , πs is a double cover. Since πs and C ⊗ C ∼
=C
are equivariant relative to the homomorphism z 7→ z 2 of U(1), it follows that
634 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

0
L := PU(1) ×U(1) C is a line bundle, such that
L⊗L∼
= PU(1) ×U(1) C = Λn,0 (TC M ) ∼
= Λ0,n (TC M ∗ ) .
In other words, under the assumption (17.76), L is a square-root of the canonical
line bundle for M . In general, L exists locally. As usual, let
Σ± (M ) := PSpin(n) ×Spin(n) Σ± + −
2m and Σ(M ) := Σ (M ) ⊕ Σ (M ) .
We claim that there is a natural isomorphism

(17.169) L ⊗ Σ(M ) −→ Λ0,∗ (TC M ∗ ) .
Let µ : U(1) → U(C) be given by µ(z) w = zw. We then have the representation
µ ⊗ ρ : U(1) × Spin(n) → U(C ⊗R Σ2m ) .
There is also an isomorphism φ : C ⊗R Σ2m → Σ2m given by φ(z ⊗ σ) := zσ. Note
that φ is equivariant in the sense that for each (ζ, g) ∈ U(1) × Spin(n) ,
(µ⊗ρ)(ζ,g)
C ⊗R Σ2m → C ⊗R Σ2m
φ↓ ↓φ
ρC(χ(ζ,g))
Σ2m → Σ2m
commutes. Indeed, for z ⊗ σ ∈ C ⊗R Σ2m ,
ρC (χ(ζ, g))(φ(z ⊗ σ)) = ρC (ζg)(zσ) = ζρ(g)(zσ) = ζzρ(g)(σ)
= φ(ζz ⊗ ρ(g)(σ)) = φ((ι ⊗ ρ)(ζ, g) z ⊗ σ) .
Consequently, the diagram (where Φ(T ) := φ ◦ T ◦ φ−1 )
µ⊗ρ
U(1) × Spin(n) → U(C ⊗R Σ2m )
χ↓ ↓Φ
ρC
Spinc (n) → U(Σ2m ) ,
↑j ↓ Id
Λ•
U(m) → U(Λ• (Cm ))
commutes and extends (17.168). Since Λ• = ρC ◦j : U(m) → U(Σ2m ) = U(Λ• (Cm )),
it follows (see Proposition 15.23, p.409) that
(17.170) U(M ) ×U(m) Λ• (Cm ) ∼= PSpinc(n) ×Spinc(n) Σ2m .
Moreover, since the diagrams
(µ⊗ρ)(ζ,g) µ⊗ρ
C ⊗R Σ2m → C ⊗R Σ2m U(1) × Spin(n) → U(C ⊗R Σ2m )
φ↓ ↓φ and χ↓ ↓Φ
ρC(χ(ζ,g)) c ρC
Σ2m → Σ2m Spin (n) → U(Σ2m )
commute, the isomorphism φ : C ⊗ Σ2m → Σ2m is equivariant. Then (see Proposi-
tion 15.23, p.409)
L ⊗ Σ(M ) = PU(1)×Spin(n) ×U(1)×Spin(n) (C ⊗ Σ2m )
∼ ∼
−→ PSpinc(n) ×Spinc(n) Σ2m −→ U(M ) ×U(m) Λ• (Cm )

(17.171) −→ Λ0,∗ (TC M ∗ ) ,
where we have used (17.170). This is the desired natural isomorphism (17.169).
Note that via (17.171), Λ0,∗ (TC M ∗ ) may be regarded as an associated bundle of
PU(1)×Spin(n) . However, Λ0,∗ (TC M ∗ ) is not associated to F M or PSpin(n) and hence is
17.6. THE INDEX THEOREM FOR STANDARD GEOMETRIC OPERATORS 635

not a Clifford module bundle in the strict sense of Definition (17.58). Nevertheless,
under the Assumption (17.76), Λ0,∗ (TC M ∗ ) is a twisted Clifford module bundle,
since Σ(M ) is a Clifford module bundle and Λ0,∗ (TC M ∗ ) is obtained by twisting
Σ(M ) with L. To prove the Hirzebruch-Riemann-Roch Theorem, √ it
 remains to
check the twisted Dirac operator DL,Σ2m corresponds to 2 ∂¯ + ∂¯∗ . More pre-
cisely, the above isomorphism in (17.171), say

φM : L ⊗ Σ(M ) −→ Λ0,∗ (TC M ∗ ) ,
provides us with a linear isomorphism
Γ(φM ) : C ∞ (L ⊗ Σ(M )) → Ω0,∗ (M ) given by
Γ(φM ) [ψ](y) = φM (ψ(y)) .

Exhibiting 2 ∂¯ + ∂¯∗ as a twisted Dirac operator, at least locally. Of course


the isomorphism φM of vector bundles also induces a corresponding isomorphism,


still denoted by Γ(φM ), between the spaces of exterior forms with values in these
vector bundles. Now we need to show that, √
under Γ(φM  ), the twisted Dirac operator
DL,Σ2m on C ∞ (L ⊗ Σ(M )) corresponds to 2 ∂¯ + ∂¯∗ on Ω0,∗ (M ). In other words,
 √
Γ(φM ) DL,Σ2m ψ = 2 ∂¯ + ∂¯∗ (Γ(φM ) [ψ]) .
 

Recall that
DL,Σ2m := (1 ⊗ c) ◦ ∇ : C ∞ (L ⊗ Σ(M )) → C ∞ (L ⊗ Σ(M )) ,
while

2 ∂¯ + ∂¯∗ := Ω0,∗ (M ) → Ω0,∗ (M ) .


We first need to show that Γ(φM ) respects covariant differentiation


Γ(φM ) [∇ψ] = ∇(Γ(φM ) [ψ]) ,
but this is a consequence of the facts that φ is equivariant and that the connection
forms on PU(1)×Spin(n) and PSpinc(n) are both pull-backs via
πχ πSq×c
PU(1)×Spin(n) −→ PSpinc(n) −→ PU(1)×SO(n)
of the connection 1-form θ0,n ⊕ θ on PU(1)×SO(n) , where θ denotes the Levi-Civita
connection 1-form on F M and θ0,n denotes the connection induced by θ on the
unitary frame bundle for Λ0,n (TC M ∗ ), namely PU(1) . We show

Γ(φM ) [Dψ] = 2 ∂¯ + ∂¯∗ (Γ(φM ) [ψ])

(17.172)
as follows. It is enough to show that (17.172) holds at an arbitrary point x ∈
M . We continue to let x1 + iy 1 , . . . , xm + iy m be complex coordinates about
x ∈ M such that ∂x1 , ∂y1 , . . . , ∂xm , ∂ym are orthonormal at x relative to hx , and
∇(∂xi ) = ∇ ∂yi = 0 at x. It is convenient to replace our model Λ• (Cm ) for Σ2m

by Λ0,∗ (C ⊗R Cm ) which is unitarily equivalent to Λ• (Cm ) via the C-linear map
∼ ∗
Λ• (Cm ) −→ Λ0,∗ (C ⊗R Cm ) induced by
√ ∼ ∗
2 [ ◦ π 0,1 : Cm −→ Λ0,1 (C ⊗R Cm ) , i.e.,
√ ∗
Λ1 (Cm ) 3 v 7→ 2 [ 12 (v + iJv) ∈ Λ0,1 (C ⊗R Cm )

636 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS


Note that [ ◦ π 0,1 (v) = π 0,1 (v) = √12 |v|, so that
 
2 [ ◦ π 0,1 (v) = |v|. Clif-
∗
ford multiplication of v ∈ R2m ∼
= Cm on Λ0,∗ (C ⊗R Cm ) is then

2 [ ◦ π 0,1 v∧ − [ ◦ π 0,1 vx .
  

Any section of ψ ∈ C ∞ (L ⊗ Σ(M )) is locally of the form


m X
−1
X
dz̄ j1 ∧ · · · ∧ dz̄ jk

ψ= ψj1 ...jk (φM )
k=1 (j)k

and Γ(φM ) [ψ] ∈ Ω0,∗ (M, C) is given by


m X
X
Γ(φM ) [ψ] = ψj1 ...jk dz̄ j1 ∧ · · · ∧ dz̄ jk , ψj1 ...jk ∈ C ∞ (M, C) .
k=1 (j)k

Let
dz 1 ∧ · · · ∧ dz m dz̄ 1 ∧ · · · ∧ dz̄ m
µC = and µ̄C = .
|dz 1 ∧ · · · ∧ dz m | |dz̄ 1 ∧ · · · ∧ dz̄ m |
Then µ̄C is a local frame for the canonical line bundle Λ0,n (TC M ∗ ); i.e., µ̄C is a
√ 0
local section of PU(1) and we let µ̄C denote the corresponding local section of PU(1)
so that
−1  √
dz̄ j1 ∧ · · · ∧ dz̄ jk = µ̄C ⊗ dz̄ j1 ∧ · · · ∧ dz̄ jk .

(φM )
Since dz k and dz̄ k are parallel at x, µC and µ̄C are also parallel at x. Moreover,

[ ◦ π 0,1 (∂xk ) = [ 12 ∂xk + i∂yk = [(∂z̄k ) = 12 dz̄ k and


 

[ ◦ π 0,1 ∂yk = [ 12 ∂yk − i∂xk = [(−i∂z̄k ) = i[(∂z̄k ) = 2i dz̄ k .


  

Thus at x,
m X 
X −1 
DL,Σ2m ψ = DL,Σ2m ψj1 ...jk (φM ) dz̄ j1 ∧ · · · ∧ dz̄ jk
k=1 (j)k
m X
X √
DL,Σ2m ψj1 ...jk µ̄C ⊗ dz̄ j1 ∧ · · · ∧ dz̄ jk

= and
k=1 (j)k


√1 D L,Σ2m ψj1 ...jk µ̄C ⊗ dz̄ j1 ∧ · · · ∧ dz̄ jk

2
     √
= [ ∂x0,1 l ∧ − [ ∂x0,1 l x ∂xl (ψj1 ...jk ) µ̄C ⊗ dz̄ j1 ∧ · · · ∧ dz̄ jk
     √
+ [ ∂y0,1 l ∧ − [ ∂y0,1 l x ∂yl (ψj1 ...jk ) µ̄C ⊗ dz̄ j1 ∧ · · · ∧ dz̄ jk

= 21 dz̄ l ∧ − 12 dz̄ l x ∂xl (ψj1 ...jk ) µ̄C ⊗ dz̄ j1 ∧ · · · ∧ dz̄ jk


+ 2i dz̄ l ∧ − 2i dz̄ l x ∂yl (ψj1 ...jk ) µ̄C ⊗ dz̄ j1 ∧ · · · ∧ dz̄ jk


= dz̄ l ∧ − dz̄ l x 12 ∂xl + i∂yl (ψj1 ...jk ) µ̄C ⊗ dz̄ j1 ∧ · · · ∧ dz̄ jk
 
√
= ∂z̄l (ψj1 ...jk ) dz̄ l ∧ − dz̄ l x µ̄C ⊗ dz̄ j1 ∧ · · · ∧ dz̄ jk .
17.6. THE INDEX THEOREM FOR STANDARD GEOMETRIC OPERATORS 637

Hence,

√1 Γ(φM ) DL,Σ2m ψj1 ...jk µ̄C ⊗ dz̄ j1 ∧ · · · ∧ dz̄ jk
 
2

= ∂z̄l (ψj1 ...jk ) dz̄ l ∧ −dz̄ l x Γ(φ) µ̄C ⊗ dz̄ j1 ∧ · · · ∧ dz̄ jk
 

= ∂z̄l (ψj1 ...jk ) dz̄ l ∧ −dz̄ l x dz̄ j1 ∧ · · · ∧ dz̄ jk




= ∂¯ + ∂¯∗ ψj1 ...jk dz̄ j1 ∧ · · · ∧ dz̄ jk


 

= ∂¯ + ∂¯∗ Γ(φM ) ψj1 ...jk µ̄C ⊗ dz̄ j1 ∧ · · · ∧ dz̄ jk ,
  

and (17.172) holds by linearity. Since


M
Σ+ ev m
2m = Λ (C ) := Λl (Cm ) , while
l even
M
Σ−
2m = Λ
odd
(Cm ) := Λl (Cm ) ,
l odd

we have
φM φM
L ⊗ Σ+ (M ) ∼= Λ0,ev (TC M ∗ ) and L ⊗ Σ− (M ) ∼ = Λ0,odd (TC M ∗ ) , and
−1

Γ(φM ) ◦ DL,Σ2m ◦ Γ(φM ) = 2 ∂¯ + ∂¯∗ : Ω0,ev (M ) → Ω0,odd (M ) .



Rewriting 2 ∂¯ + ∂¯∗ Yields the Hirzebruch-Riemann-Roch Theorem. Hence


under Assumption (17.76), we have exhibited 2 ∂¯ + ∂¯∗ as a twisted Dirac op-


erator, and locally so if (17.76) fails globally. Thus, on a suitable√neighborhood B


about any point (e.g., a normal ball), the local index density for 2 ∂¯ + ∂¯∗ coin-

0
cides with that for DL,Σ2m on B, where PU(1) , PSpin(n) , L and Σ(B) may only exist
over B. Let θ denote the Levi-Civita connection 1-form on F M , let θe denote the
connection 1-form c0−1 (C ∗ (θ|F B )) on PSpin(n) , let θ0,m denote the connection form
on PU(1) = U Λ0,m (TC M ∗ ) induced by θ|U(M ) , and let θe0,m = 12 πs∗ θ0,m denote the

0 0
related connection on PU(1) where πs : PU(1) → PU(1) denotes the double cover. The

¯ ¯∗

local index density form for 2 ∂ + ∂ |Ω0,ev(M ) is the same as that for DL,Σ2m on
B, namely
     
ch L, θe0,m ∧ A(M,
b θ) = ch L, θe0,m ∧ ch PSpin(n) , θ, e ρ0 ∧ A(M,
e θ)
    
0
= ch PU(1) , θe0,m , µ0 ∧ ch PSpin(n) , θ, e ρ0 ∧ A(M,
e θ)
 
0
= ch PU(1) × PSpin(n) , θe0,m ⊕ θ,(µe ⊗ ρ)0 ∧ A(M, e θ)

= ch U(M ) , θ|U(M ) , Λ0,∗0 ∧ A(M,



e θ)
0,∗ ∗ 
(17.173) = ch Λ (TC M ) ∧ A(M,
e θ) = Td(T M, θ) .
Thus, we obtain
Theorem 17.77 (Hirzebruch-Riemann-Roch Theorem). Let M be a compact
Kähler manifold. Then
index ∂¯ + ∂¯∗ : Ω0,ev (M ) → Ω0,odd (M ) = Td(T M ) [M ] .
 

A Twisted Version of the H-R-R Theorem. There is also a twisted version


of the Hirzebruch-Riemann-Roch Theorem. Let E → M be a Hermitian vector
bundle with a covariant derivative ∇E arising from a unitary connection 1-form ε
638 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

on U(E). We have the spaces Ωp,q (M, E) := C ∞ (E ⊗ Λp,q (TC M ∗ )) of E-valued


forms of bidegree (p, q), which are locally of the form (in multi-index notation)
1 X
φ= f(j)p ;(k)q ⊗ dz (j)p ∧ dz̄ (k)q ,
p!q!
(j),(k)

where f(j)p ;(k)q is a local section of E. Moreover, there are global operators

∂ E : Ωp,q (M, E) → Ωp+1,q (M, E) and ∂¯E : Ωp,q (M, E) → Ωp,q+1 (M, E)
determined locally by
1 X  
(j)p
∂E φ = ∇E
∂z h f h
(j)p ;(k)q dz ∧ dz ∧ dz̄ (k)q , and
p!q!
h,(j),(k)
1 X  
(17.174) ∂¯E φ = ∇E
∂z̄h f h
(j)p ;(k)q dz̄ ∧ dz
(j)p
∧ dz̄ (k)q .
p!q!
h,(j),(k)

Note that just as ∂¯ + ∂¯∗ was shown to be locally a twisted Dirac operator (twisted
by L), ∂¯E + ∂¯E∗ is also locally a twisted Dirac operator (twisted by E ⊗ L).
Theorem 17.78 (Twisted Hirzebruch-Riemann-Roch Theorem). Let M be a
compact Kähler manifold and let E → M be a Hermitian vector bundle. Then
index ∂¯E + ∂¯E∗ : Ω0,ev (M, E) → Ω0,odd (M, E) = (ch(E) ` Td(M )) [M ] .


Proof. Using the local computation (17.173), we get


 
ch E ⊗ L, (ε, θe0,m ) ∧ A(M,
b θ)
   
= ch(E, ε) ∧ ch L, θe0,m ∧ A(M,
b θ)
= ch(E, ε) ∧ Td(T M, θ) . 

Getting the Holomorphic Hirzebruch-Riemann-Roch Theorem. There is a per-


haps more familiar version of Theorem 17.78 which applies when E → M is holo-
morphic. We say that the complex vector bundle of complex fiber dimension N
πE : E → M over the complex manifold M is holomorphic if E has the structure of
a complex manifold, and about each point x ∈ M there is a neighborhood B and a
biholomorphic vector bundle map
−1
φB : πE (B) → B × CN ,
which we call a holomorphic trivialization. If E has complex fiber dimension
N , then the C-linear frames of E (linear maps u from CN to Ex for x ∈ M ) form a
principal GL(N, C)-bundle, say π eE : GL(E) → M , and if E has a Hermitian struc-
ture, U(E) → M is a principal subbundle. Although U(E) does not generally have
a complex structure, we show that GL(E) does. Each holomorphic trivialization
φB determines a section uφ : B → GL(E) given (for z ∈ CN ) by
uφ (y)(z) = φ−1
B (y, z) .

An arbitrary frame u ∈ GL(E) at y ∈ B is of the form uφ (y) ◦ Aφ (u) for some


Aφ (u) ∈ GL(N, C). We then obtain a map
φ̃B : GL(E) |B → B × GL(N, C) given by φ̃B (u) = (y, Aφ (u))
17.6. THE INDEX THEOREM FOR STANDARD GEOMETRIC OPERATORS 639

−1
If φ0B 0 : πE (B 0 ) → B 0 × CN is another holomorphic trivialization, then
uφ (y) ◦ Aφ (u) = u = uφ0 (y) ◦ Aφ0 (u) , and so
   
−1
φ̃0B 0 ◦ φ̃−1
B (y, Aφ (u)) = φ̃ 0
B 0 (u) = y, uφ 0 (y) ◦ u φ (y) ◦ A φ (u) .

Now φ̃0B 0 ◦ φ̃−1 0 0


B : (B ∩ B ) × GL(N, C) → (B ∩ B ) × GL(N, C) is holomorphic, since
0
for (y, C) ∈ (B ∩ B ) × GL(N, C),
   
−1
φ̃0B 0 ◦ φ̃−1
B (y, C) = y, u φ0 (y) ◦ uφ (y) ◦ C
 −1 
(17.175) = y, φ0−1 B 0 (y, ·) ◦ φ−1
B (y, ·) ◦ C

is a holomorphic function of (y, C). We now define an operator


∂¯E : Ωp,q (M, E) → Ωp,q+1 (M, E)
that generalizes ∂¯ : Ωp,q (M ) → Ωp,q+1 (M ) which is the special case E = M × C. If
f1 , . . . , fN denotes the standard basis for CN and sj (x) := φ−1 B (x, fj ), then locally,
∂¯E is given by
X 
¯
∂E αk,(i)(j) sk ⊗ dz (i)p
∧ dz̄ (j)q
k,(i)p ,(j)q
X
∂¯ αk,(i)(j) sk ⊗ dz (i) ∧ dz̄ (j) .

(17.176) :=
k,(i),(j)

This is well-defined (independent of local holomorphic trivialization and coordi-


nates), since ∂¯ kills holomorphic functions. By definition, the kernel of ∂¯E consists
of the holomorphic sections of E. Note that ∂¯E ◦ ∂¯E = 0 and so we have an exact
sequence
∂¯ ∂¯ ∂¯ ∂¯
Ω0,0 (M, E) →
E E
··· → Ω0,q (M, E) →
E
Ω0,q+1 (M, E) →
E
··· ,
along with cohomology groups
Ker ∂¯E |Ω0,q(M,E)

q
H (OE ) := ¯ .
∂E (Ω0,q−1 (M, E))
Since ∂¯E ◦ ∂¯E = 0, ∂¯E is like a covariant derivative operator for a connection 1-
form on GL(E) with zero curvature. Indeed, on B × GL(N, C) there is a standard
flat connection 1-form given in terms of coordinates (y, C) by C −1 dC. On B 0 ×
GL(N, C), with coordinates (y, C 0 ), we have C 0−1 dC 0 . In view of (17.175),
−1
C 0 = φ0−1
B 0 (y, ·) ◦ φ−1
B (y, ·) ◦ C and so
−1 −1
 
dC 0 = d φ0−1 −1
◦ C 0 + φ0−1 ◦ φ−1

B 0 (y, ·) ◦ φ B (y, ·) B 0 (y, ·) B (y, ·) ◦ dC.

Because of the first term, it is not true that C 0−1 dC 0 = C −1 dC. However, the (0, 1)-
−1
component of the first term is 0 since φ0−1 B 0 (y, ·) ◦ φ−1
B (y, ·) is a holomorphic
function of y, and so
¯ 0 = C −1 ∂C
C 0−1 ∂C ¯ on (B ∩ B 0 ) × GL(N, C) .
¯ 0 = 0 and ∂C
In particular, the equations ∂C ¯
0,1 −1 0
 =0,1 0 define the same horizontal distri-
bution of subspaces of T eE (B ∩ B ) ⊆ T (GL(E)), of C-dimension m, and
π
so we obtain a well-defined horizontal distribution of T 0,1 (GL(E)). A genuine con-
nection on GL(E) determines a horizontal distribution of real subspaces T (GL(E))
of R-dimension 2m, and hence a horizontal distribution TC (GL(E)) of C-dimension
640 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

¯ is the well-defined (0, 1)-component of an ill-defined flat


2m. In essence, C −1 ∂C
connection C dC, and ∂¯E is the well-defined (0, 1)-component of the ill-defined
−1
¯
covariant derivative for C −1 dC. To avoid additional notation, we still use C −1 ∂C
−1 ¯
to denote the global (0, 1)-form given locally as C ∂C. Proposition 17.82 below
asserts that the Hermitian metric on E determines a standard well-defined (1, 0)-
component ω 1,0 whose sum with C −1 ∂C¯ is a complex connection 1-form
¯ ∈ Ω1 (GL(E) , gl (n, C)) ,
ω = ω 1,0 ⊕ C −1 ∂C C C

where glC (n, C) = C ⊗R gl(n, C). By complex connection 1-form, we mean that
although ω is not necessarily the complex extension of a (real) connection 1-
form in Ω1 (GL(E) , gl(n, C)), at each point u ∈ GL(E) the form ω is defined on
TC GL(E)u := C ⊗R Tu GL(E) and satisfies the usual relations
d
(u exp tA) |t=0 = A and Rg∗ ω = g −1 ωg,

ωu dt
for A ∈ gl(N, C) and g ∈ GL(N, C).
Definition 17.79. Let ω be a complex connection 1-form on GL(E) for a holo-
morphic vector bundle E. Then ω is compatible with the complex structure
¯
on E if ω 0,1 = C −1 ∂C.
Let Herm(N ) denote the space of all Hermitian N × N matrices. We denote
by r : GL(N, C) → End(Herm(N )) the representation given by
T
r(g)(η) = g −1 ηḡ −1 , for g ∈ GL(N, C) and η ∈ Herm(N ) .
Note that r(g) is just pull-back by g −1 , since
T T  
X T (r(g)(η)) Ȳ = X T g −1 ηḡ −1 Ȳ = g −1 X η g −1 Y .

Definition 17.80. If E has a Hermitian structure H, then ω is a Hermitian


connection for H if Dω H e = 0, where
e ∈ Ω̄0 (GL(E) , Herm(N )) = f ∈ C ∞ (GL(E) , Herm(N )) : f (ug) = r g −1 f (u) .
 
H
denotes the equivariant function corresponding to H.
Remark 17.81. If ∇E : C ∞ (M, E) → Ω1 (M, E) is the covariant derivative
arising from ω, then ω is compatible with the complex structure on E iff
1,0
π 0,1 ◦ ∇E = ∂¯E or ∇E = ∂¯E + ∇E .
Moreover, ω is a Hermitian connection for H if
d(H(V, W )) = H ∇E V, W + H V, ∇E W .
 

The connection in the following Proposition is variously known as the Chern


connection, the (1, 0)-connection, the metric connection, and the canonical con-
nection for the Hermitian, holomorphic vector bundle E.
Proposition 17.82. Let M be a complex manifold and let E → M be a holo-
morphic vector bundle with Hermitian metric H. Then there is a unique complex
connection 1-form ω on GL(E) which is both compatible with the complex structure
on E and a Hermitian connection for H.
17.6. THE INDEX THEOREM FOR STANDARD GEOMETRIC OPERATORS 641

Proof. First we prove uniqueness. Suppose that ω exists. Then


(17.177) 0 = Dω He = dHe + r0 (ω) H e − ωH
e = dH e − Hω.
e
Note also that
= r g −1 (H(u)) = g T H(u)

H(ug)
e e e ḡ
(17.178) e ∗ ) = AT H(u)
=⇒ dH(A e + H(u)
e Ā.
u

Using the complex structure on GL(E), ω decomposes as ω = ω 1,0 ⊕ ω 0,1 . Observe


that ω 1,0 and ω 0,1 have values in C⊗R gl(n, C), and any A ∈ gl(n, C) splits as A1,0 +
A0,1 , where A1,0 = A0,1 ∈ C ⊗R gl(n, C). Since ω is compatible with the complex
structure on E, ω 0,1 = C −1 ∂C¯ in terms of the coordinates (y, C) in B × GL(N, C).
Then ω 1,0 is uniquely determined by taking the (1, 0)-part of (17.177), namely
e − ω 1,0 T H e 0,1 =⇒ ω 1,0 T H
 
0 = ∂H e − Hω e = ∂H e 0,1
e − Hω
  −1 T
e T − ω 0,1 T H
 
=⇒ ω 1,0 = ∂ H e 0,1 H
e − Hω e =He −1T ∂ H eT .

T T −1
¯

Since ω 0,1 = C −1 ∂C = ∂ C̄ T C̄ T , we get
e T − ∂ C̄ T C̄ T −1 H
 
ω 1,0 e −1T ∂ H eT .
 
(17.179) =H
Thus, ω is unique. We mention that taking the (0, 1)-part of (17.177) and using
H
e =H e T gives the same result. Now if ω 1,0 is defined by (17.179), then the sum
e T − ∂ C̄ T C̄ T −1 H
 
e −1T ∂ H ¯
e T ⊕ C −1 ∂C
 
ω := H
e −1T ∂ C̄ T C̄ T −1 H e T − ∂ C̄ T C̄ T −1 H
   
¯
e T ⊕ C −1 ∂C
  
=H
e −1T C̄ T ∂ C̄ T −1 H
 
¯
e T ⊕ C −1 ∂C

=H

is a complex connection on GL(E). Indeed, we have Rg∗ ω = g −1 ωg, since


−1 ¯ ¯ g, and
= g −1 C −1 ∂C

(Cg) ∂(Cg)
 
  −1T  −1  T  T  T −1   T
r g −1 H e ∂ r g H
e − ∂ Cg Cg r g −1 H e
 −1T   T   −1  T 
= gT H e ḡ ∂ gT H e ḡ − ∂ Cg T Cg T gT H
e ḡ

e T g − ḡ T ∂ C T C T −1 ḡ T −1 ḡ T H
       
= g −1 H e −1T ḡ −1T ḡ T ∂ H eTg

e T − ḡ T ∂ C̄ T C̄ T −1 H
 
= g −1 H
e −1T ∂ H e T g.
 

To show ω(A∗ ) = A for A ∈ gl(n, C), note that (17.178) implies


e T A1,0 + ĀT 1,0 H
e T (A∗u ) = H e T , and then

∂H

e T (A∗ ) − ∂ C̄ T (A∗ ) C̄ T −1 H
 
εe(A∗ ) = H
e −1T ∂ H ¯
e T ⊕ C −1 ∂C(A ∗
 
)
  
e −1T H e T A1,0 + ĀT 1,0 He T − ĀT 1,0 H e T ⊕ A0,1
 
=H
= A1,0 ⊕ A0,1 = A. 
642 17. THE LOCAL INDEX THEOREM FOR TWISTED DIRAC OPERATORS

If E is a holomorphic Hermitian bundle over a Kähler manifold and ∇E is the


covariant derivative of a complex connection 1-form on GL(E) which is a Hermitian
connection and compatible with the complex structure on E, then ∂¯E defined in
(17.174) agrees with ∂¯E defined in (17.176). Hence, by a Hodge theoretic proof
strictly analogous to that of Theorem 17.72, when E is Hermitian and holomorphic
we have
H q (OE ) ∼
= Hq (E) := Ker ∂¯E + ∂¯E ∗ |Ω0,q(M,E) .
 

The following theorem is then immediate.


Theorem 17.83 (Holomorphic Hirzebruch-Riemann-Roch Theorem). Let M
be a compact Kähler manifold and let E → M be a Hermitian holomorphic vector
bundle. Then the holomorphic Euler characteristic of E is given by
Xm q
χhol (E) := (−1) dim H q (OE )
q=0

= index ∂¯E + ∂¯E∗ : Ω0,ev (M, E) → Ω0,odd (M, E)




= (ch(E) ` Td(M )) [M ] .
In particular, χhol (E), which seems to depend on the holomorphic structure of E,
actually only depends on the topology of E.
CHAPTER 18

Seiberg-Witten Theory

Synopsis. Background and Survey: Intersection Form and Homotopy Type of Com-
pact Oriented Simply-Connected Four-Manifolds; Unimodular Forms and Freedman’s
Theorem; Existence of Differentiable Structures; Donaldson’s Polynomial Invariants; Re-
sults Concerning Symplectic Manifolds; Purely Geometric Applications. Spinc Structures
and the Seiberg-Witten Equations: Admittance of Spinc Structures, Spinc Dirac Opera-
tors; Motivating and Defining the Unperturbed and the Perturbed Seiberg-Witten Equa-
tions. Generic Regularity of the Moduli Spaces: Gauge Transformations; Moduli Space of
Solutions of the Perturbed S-W Equations; The Formal Dimension of the Moduli Space;
Seiberg-Witten Function; Quotient Manifolds; Manifold Structure for the Parametrized
Moduli Space; Generic Regularity; A Priori Bounds; Sobolev Estimates; Compactness of
Moduli Spaces; Definition of the S-W Invariant; Metric Dependence of Connections and
Dirac Operators; Oriented Cobordism; Fredholm Transversality; Full Invariance of the
S-W Invariant.

1. Background and Survey

I Since the early 1980s, S.K. Donaldson and others have been proving results about
smooth 4-manifolds using moduli spaces of instantons (self-dual SU(2)-connections modulo
gauge transformations). The approach inspired by N. Seiberg and E. Witten (see the
original [386] and the more recent [265] which is the modern point of view in gauge theory)
uses monopole moduli spaces (twisted spinor fields paired with abelian U(1)-connections,
modulo gauge transformations). Seiberg-Witten theory not only provides simpler proofs
of most of Donaldson’s results, but has generated many new results, notably Taubes’
work concerning symplectic manifolds (e.g., see [408] and [409]). Like in our Chapter 16
on Gauge Theoretic Instantons, our goal is not to review the most recent results of the
continuously evolving theory but rather to deliver a hopefully digestible presentation of the
main results: We first introduce some notation, and review the major results. Afterwards,
we give sometimes sketches and sometimes details of the proofs. J

Intersection Form and Homotopy Type of Compact Oriented Simply-


Connected Four-Manifolds. For a compact, simply-connected (and connected),
oriented 4-manifold X, let [X] ∈ H 4 (X; Z) denote the fundamental class determined
by the orientation. For the standard definitions and results from algebraic topology
in what follows, we refer the reader to [185] and [403]. Let
QX : H 2 (X; Z) × H 2 (X; Z) −→ Z
denote the symmetric bilinear form (the so-called intersection form of X), given
by
(18.1) QX (α, β) := [X] a (α ` β) = hα ` β, [X]i ,
643
644 18. SEIBERG-WITTEN THEORY

where ` and a are the cup and cap products, and h·, ·i denotes the pairing between
homology and cohomology, namely the Kronecker product
h·, ·i : H q (X; Z) × Hq (X; Z) −→ Z.
Note that if the orientation of X is switched, then QX changes to −QX . It is known
(see [185, p.180]) that α 7→ hα, ·i defines a surjection
h : H q (X; Z) −→ Hom(Hq (X; Z); Z) ,
but in general there may be a kernel. The Universal Coefficient Theorem (see [185,
p.189]) yields a split exact sequence
h
0 −→ Ext(Hq−1 (X; Z); Z) −→ H q (X; Z) −→ Hom(Hq (X; Z); Z) −→ 0.
Since we assumed that X is simply-connected, H1 (X; Z) = 0, and so

(18.2) h : H 2 (X; Z) −→ Hom(H2 (X; Z), Z),
which is a free, abelian Z-module (i.e., a lattice). In the smooth category,
Z
X
Q (α, β) = α̃ ∧ β̃,
X

where α̃ and β̃ are closed 2-forms representing (nontorsion) integral classes α and
β, and the integral is defined via the given orientation on X. Poincaré duality
states that we have an isomorphism

(18.3) P := [X] a (·) : H i (X; Z) −→ H4−i (X; Z).
Thus, H2 (X; Z) ∼
= H 2 (X; Z) is also a lattice. Hence, we may alternatively use
QX : H2 (X; Z) × H2 (X; Z) −→ Z.
The term intersection form follows from the fact that if two classes in H2 (X; Z)
are represented by a compact, oriented embedded surfaces which intersect trans-
versely, then QX on this pair is the intersection number of the two surfaces (i.e.,
the algebraic number of signed intersections, where the sign is ±1, depending on
whether the combined orientation of surfaces at an intersection point agrees with
that of X).
Remark 18.1. a) In the preceding definition of the intersection form we re-
stricted ourselves to simply-connected manifolds. However, for a not-necessarily
simply connected four-manifold one observes that the cup product descends to the
quotient H 2 (X; Z)/ torsion. Poincaré duality and the preceding Universal Coeffi-
cient Theorem then says it is unimodular on this quotient. Consequently, many
of this chapter’s concepts and results can be re-formulated and proven in greater
generality.
b) To give an example for the range of possible generalizations we mention that
the following Proposition 18.2 is true for any X, simply connected or not, and
the conclusion is stronger: any class in H 2 (X; Z) is represented by an embedded
(and connected) surface; there is no need to take a homological sum. The proof is
the same as ours, except it starts with the fact that the Atiyah-Hirzebruch spec-
tral sequence shows that the map from oriented 2-dimensional bordism to X is an
isomorphism. This gives a map of a surface (instead of a sphere as in the simply-
connected case in the following proof) which one can de-singularize. One can also
give a completely elementary proof, that every class in H 2 (X; Z) is represented by
a continuous map of a surface using singular homology and triangles, just as in the
18.1. BACKGROUND AND SURVEY 645

classic proof that yields the first homology group from making the first homotopy
group abelian.

Proposition 18.2. If X is a compact, simply-connected, oriented 4-manifold,


then any class in H2 (X; Z) is a homological sum of compact, oriented embedded
surfaces.

Proof (Sketch). Since X is simply-connected, the Hurewicz Isomorphism



Theorem (see [403, p.398]) yields an isomorphism π2 (X, x0 ) −→ H2 (X; Z). This
implies that the generators of H2 (X; Z) can be represented by maps S 2 → X. Such
a map can be approximated (within the same homotopy class) by an immersion f
with transverse self-intersections. We can find a small  4-ball Bi about each self-
2
intersection i (say i ∈ {1, . . . , g}), such that f S
 ∩ Bi consists of two disks Di1
and Di2 intersecting at a point. Note that f S 2 ∩ ∂Bi is the union of the disjoint
circles ∂D1 and ∂D2 which can be joined by an embedded tube Ti ⊆ Bi (e.g.,
(1 − u) ei2πt , uei2πt , t, u ∈ [0, 1]). Replacing Di1 ∪ Di2 by Ti for each  i, we obtain
2
an embedded surface Σ. Note that homology classes of Σ and f S in H2 (X; Z)
are the same in H2 (X; Z), since these classes trivially determine the same element
in H2 (X \ ∪i B; Z) which is isomorphic to H2 (X; Z) because attaching 4-balls to
X \ ∪i B does not alter H2 (see [185, p.118]). Moreover, since Σ is obtained by
attaching g handles to S 2 \ ∪gi=1 (Di1 ∪ Di2 ), Σ is an oriented surface of genus g. 

Exercise 18.3. The connected sum M1 #M2 of two n-manifolds M1 and M2


is the manifold obtained by removing a ball from each and gluing the exposed spher-
ical boundaries together. Show that for two compact, simply-connected, oriented
4-manifolds X1 and X2 , we have

(18.4) QX1 #X2 ∼


= QX1 ⊕ QX2 .
If needed, see [300, p.103-105].

Since X is simply-connected,

H 1 (X; Z) ∼
= Hom(H1 (X; Z), Z)⊕ Tor(H0 (X; Z) = {0}, and
P P
{0} = H 1 (X; Z) ∼
= H3 (X; Z) and H 3 (X; Z) ∼
= H1 (X; Z) = {0}.

Thus, QX provides all of the homological/cohomological information about X.


Indeed, Milnor (in [296]) announced that, as a consequence of a theorem of J.H.C.
Whitehead in [439], the oriented homotopy type of a simply-connected, compact,
oriented four-manifold X is completely determined by QX . A proof is given in
[300, p.103-5].

Unimodular Forms and Freedman’s Theorem. The form QX is nonde-


generate over Z, meaning that

w 7→ QX (w, ·) defines an isomorphism H2 (X; Z) −→ Hom(H2 (X; Z), Z).

This isomorphism is just


P −1 h
−1 ∼ ∼
h◦P : H2 (X; Z) −→ H 2 (X; Z) −→ Hom(H2 (X; Z), Z).
646 18. SEIBERG-WITTEN THEORY

Indeed, we can write w = [X] a α for a unique α = P −1 (w) ∈ H 2 (X; Z). Then for
z := [X] a β ∈ H2 (X; Z),
QX (w, z) = QX ([X] a α, [X] a β) = QX (α, β) = [X] a (α ` β)
= [X] a (β ` α) = ([X] a β) a α
= z a α = hα, zi = h ◦ P −1 (w) (z) .
  

Thus, the matrix of QX relative to a basis (i.e., a set of independent generators)


of H2 (X; Z) not only has integer entries, but also has determinant ±1 (i.e., is
unimodular ). Indeed, if {e1 , ..., en } is a basis of H2 (X; Z) and {δ1 , ..., δn } denotes
the dual basis, thenPnondegeneracy implies that there are α1 , ..., αn ∈ H2 (X; Z),
n
such that (for αi = k=1 αik ek , αik ∈ Z)
n
X
(18.5) δi = QX (αi , ·) or δij = δi (ej ) = QX (αi , ej ) = αik [QX ]kj .
k=1
Thus, [αik ] is the inverse matrix of [QX ] with integer entries, and det(QX ) det([αik ])
=1⇒ det(QX ) = ±1.
In [149], Michael Freedman showed that any unimodular symmetric bilin-
ear form over Zn (0 < n ∈ Z) is isomorphic to QX for some topological, simply-
connected, compact, oriented 4-manifold X. Moreover, up to homeomorphism there
are at most two such X having a given intersection form. In order to state this
result more precisely, we introduce some terminology.
A unimodular symmetric bilinear form Q on a lattice, say Zn (n is the rank
of Q), is called even if Q(α, α) is even for all α ∈ Zn . Otherwise, Q is called odd.
Note Q is even iff all of the diagonal elements of a matrix representing Q are even,
since
Xn n
X n
X
Q(α, α) = αi2 Q(ei , ei ) + 2 αi αj Q(ei , ej ) ≡ αi2 Q(ei , ei ) mod 2.
i=1 i<j i=1

The signature sig(Q) of Q is the number of positive minus the number of negative
entries when Q is diagonalized over R. If sig(Q) = ±rank(Q), then Q is definite
and otherwise Q is indefinite.
The forms which are indefinite are easy to describe (see [300]):
1. If Q is indefinite and odd of rank n, then, relative to some basis, the matrix [Q]
of Q is given by the standard forms
p q
(18.6) [Q] = diag(1, ..., 1, −1, ..., −1), p + q = n.
2. If Q is indefinite and even, then relative to some basis, [Q] is block-diagonal,
with
(18.7) [Q] = H ⊕ ... ⊕ H ⊕ ±(E8 ⊕ ... ⊕ E8 ) , (at least one H summand) where
 
2 1 0 0 0 0 0 0
 1 2 1 0 0 0 0 0 
 
 0 1 2 1 0 0 0 0 
   
0 1  0 0 1 2 1 0 0 0 
(18.8) H := and E8 :=  .
1 0  0
 0 0 1 2 1 0 1  
 0 0 0 0 1 2 1 0 
 
 0 0 0 0 0 1 2 0 
0 0 0 0 1 0 0 2
18.1. BACKGROUND AND SURVEY 647

Remark 18.4. We show that (18.6) is the intersection form for the connected
sum
     2 q 
(18.9) #p CP2 # #q CP2 := CP2 #...#CP
p 2
# CP #...#CP2 .

Here CP2 is CP2 but with the orientation


 reversed. Recall (see [185, p.119]) that
H2 CP2 ∼ = Z with generator α := CP1 1 , where


CP1 1 := {[1, z1 , 0]} ⊂ CP2 .




To compute Q(α, α), we note that CP1 1 meets the homologous




CP1 2 := {[1, 0, z2 ]} ⊂ CP2 (with CP1 2 = α = CP1 1 )


      

at the single point [1, 0, 0] with intersection number +1, whereas in CP2 the inter-
section number is −1.
Remark 18.5. Over the reals, H (where “H” stands for hyperbolic) is equiva-
lent to diag(1, −1), but H is not equivalent to diag(1, −1) over Z, since diag(1, −1)
is odd but H is even. We can see that H is the intersection  form for S2 × 2
 S ,
2 2 2 2
as follows. Each of the generators S × {y} or {x} × S in H2 S × S has
self-intersection number 0, since
(for y 6= y 0 ), S 2 × {y} = S 2 × {y 0 } , but S 2 × {y} ∩ S 2 × {y 0 } = ∅.
   

The intersection number of S 2 × {y} with {x} × S 2 is 1; here S 2 × S 2 inherits its


orientation from the S 2 factors. Note that S 2 × S 2 is not homotopy equivalent to
CP(2)#CP(2), since H  diag(1, −1).
Remark 18.6. Note that the number of E8 summands is 18 of the signature of
the even, indefinite form Q in (18.7). The form E8 gets its name from the fact that
it is isomorphic to the Cartan matrix for e8 , one of the exceptional Lie algebras.
Let e1 , . . . , en denote the standard basis vectors of Rn . In general, for any integer
k > 0, there is a lattice E4k in R4k generated by the vectors
1
ei ± ej , −ei ± ej (i, j ∈ {1, . . . , 4k} , i 6= j) , and 2 (±e1 ± · · · ± e4k ) ,
where the numbers of “+ signs” and “− signs” in the last expression are equal
mod 2. In the case k = 2, there are 4 82 + 27 = 240 of these vectors and these can
be identified with the roots of e8 (note dim e8 = 8 + 240 = 248); see [227, p.65].
Moreover, there is a transitive group of isometries on this set (see [300, p.28]). By
definition, the form E4k is the restriction to E4k of the usual dot product on R4k .
To see that this definition is consistent with the form E8 defined in (18.8), one may
do the following exercise.
Exercise 18.7. Find generators f1 , . . . , f8 of the lattice E8 , such that the matrix
[fi · fj ] is the one in (18.8). [Hint. One can choose
f1 := e1 − e2 , f2 := e3 − e2 and f8 := 12 ((e1 + · · · + e6 ) − (e7 + e8 )) .
Find f3 , . . . , f7 .]
If Q is definite and even, the rank of Q is a multiple of 8. When the rank n = 8,
up to isomorphism there is only E8 . When n = 16, there are E8 ⊕ E8 and E16 .
When n = 24, we have 24, including E8 ⊕ E8 ⊕ E8 , E8 ⊕ E16 , and the so-called
Leech lattice. However, the possibilities increase very rapidly. When n = 32, there
are over 10 million different classes. When n = 40, there are over 1051 different
648 18. SEIBERG-WITTEN THEORY

classes. There are also definite, nondiagonalizable odd Q; e.g, the direct sum of
a definite, even form and diag(1, ..., 1). Moreover, there are indecomposable, odd,
definite Q. These results are stated in [300], but the primary source is [388].
Theorem 18.8 (M. Freedman, [149]; see also [150]). Any unimodular qua-
dratic form Q over Zn (n arbitrary) is isomorphic to QX for some topological,
simply-connected, compact, oriented 4-manifold X. Moreover, up to homeomor-
phism, X is unique if Q is even. If Q is odd there are two such X. For one of
these X × S 1 is smoothable, and for the other X × S 1 is not smoothable (i.e., its
Kirby-Siebenmann invariant is nonzero).
Corollary 18.9 (Poincaré’s Conjecture in Dimension 4). If a topological
manifold X is homotopy equivalent to S 4 (e.g., X is simply-connected and has
trivial even intersection form QX = 0), then X is homeomorphic to S 4 .
Note that the famous proof of Poincaré’s original conjecture in dimension 3
was first achieved in 2003 by Grigori Perelman by rather different methods, see
Remark 10.3, p.256.
Corollary 18.10. There are over 10 million topologically distinct, simply-
connected topological 4-manifolds with even, definite intersection forms of rank 32.
Existence of Differentiable Structures. By the following result of Don-
aldson, none of the manifolds in Corollary 18.10 admits a differentiable structure.
Theorem 18.11 (S.K. Donaldson, [123]). If X is a compact, smooth (i.e.,
C ∞ ), simply-connected 4-manifold with a definite intersection form QX , then QX
is a standard diagonalizable form (i.e., according to Theorem 18.8, topologically X
must be a connected sum of CP2 s or CP2 s).
Donaldson also proved some results about the possible indefinite, even forms
of smooth, compact, simply connected, 4-manifolds.
Theorem 18.12 (Donaldson, [125]). If the form
(18.10) (⊕M H) ⊕(±(⊕N E8 ))
is realized by a smooth, compact, simply-connected 4-manifold and N > 0, then we
must have M ≥ 3.
A K3 surface (named after E.E. Kummer, K. Kodaira and E. Kähler) is a
smooth, compact, simply-connected, complex surface X (dimC X = 2) with trivial
canonical bundle Λ2,0 (TC M ∗ ) (see p.616 and 633). An example is the complex
variety z14 + z24 + z34 + z44 = 0 in CP3 . It is a deep fact that all K3 surfaces are
diffeomorphic. However, in the holomorphic category the family of K3 surfaces is
20-dimensional [186, p. 593]. For a K3 surface, the intersection form is known to
be
QK3 = (⊕3 H) ⊕(− ⊕2 E8 ) .
By Proposition 17.20 (p. 528) we know that a compact, orientable manifold
X admits a spin structure if and only if w2 (X) = 0. There is also the following
characterization in terms of the intersection form QX .
Proposition 18.13. Let X be a compact, simply-connected 4-manifold. Then
X has a spin structure (i.e., w2 (X) = 0) ⇔ QX is even.
18.1. BACKGROUND AND SURVEY 649

Proof. Since X is simply-connected, H 2 (X; Z) is free and so reduction mod


2 r∗ : H 2 (X; Z) → H 2 (X; Z2 ) is onto. We first establish a special case of Wu’s
formula
(18.11) QX X 2
2 (w2 (X) , r∗ (α)) = Q2 (r∗ (α) , r∗ (α)) for all α ∈ H (X; Z) ,

where QX 2 2
2 : H (X; Z2 ) × H (X; Z2 ) → Z2 is simply given by

QX X
2 (r∗ (α) , r∗ (α)) := Q (α, α) mod 2 ∈ Z2 .

By Proposition 18.2, we may assume that the Poincaré dual of α is [Σg ] ∈ H2 (X; Z)
for some compact orientable surface Σg embedded in X. Using (17.17), p. 527,
QX

2 (w2 (X) , r∗ (α)) = w2 T X|Σg , r∗ ([Σg ])
= hw2 (T Σg ⊕ N Σg ) , r∗ ([Σg ])i
= hw2 (T Σg ) , r∗ ([Σg ])i + hw2 (N Σg ) , r∗ ([Σg ])i .
By Proposition 17.21 (p.529), this is
= (hc1 (T Σg ) , [Σg ]i + hc1 (N Σg ) , [Σg ]i) mod 2 = hc1 (N Σg ) , [Σg ]i mod 2
= QX ([Σg ] , [Σg ]) mod 2 = QX (α, α) mod 2 = QX
2 (r∗ (α) , r∗ (α)) .

The term hc1 (N Σg ) , [Σg ]i is the self-intersection number QX ([Σg ] , [Σg ]) of Σg .


That follows from Remark 15.69 (p.459). Thus,
QX
2 (w2 (X) , r∗ (α)) = hc1 (N Σg ) , [Σg ]i mod 2 = QX ([Σg ] , [Σg ]) mod 2

= QX (α, α) mod 2 = QX
2 (r∗ (α) , r∗ (α)) ,

and so we have (18.11). 

Recall that X admits a spin structure if w2 (X) = 0. Thus, Rokhlin’s Theorem


(see Corollary 17.56, p. 594) may be restated as
Theorem 18.14 (Rokhlin’s Theorem). The signature sig(X) of a simply con-
nected, smooth, compact 4-manifold X with QX even must be a multiple of 16.
Thus, N in (18.10) must be even.
Corollary 18.15. The topological manifold with intersection form E8 guar-
anteed to exist by Freedman’s Theorem 18.8 does not admit a smooth structure.
In view of Theorems 18.11, 18.12 and 18.14,
Theorem 18.16. Among compact, simply connected, smooth 4-manifolds X
with QX even (i.e., those which are spin), K3 surfaces (including those with oppo-
site orientation) have simplest possible QX with an E8 summand, namely
(⊕3 H) ⊕(± ⊕2 E8 ) .
Proof. By Theorem 18.11, if QX is even, then it must be indefinite, and hence
it must be of the form (⊕M H) ⊕(± ⊕N (E8 )). By Theorem 18.12, we have M ≥ 3,
and by Theorem 18.14, N ≥ 2. The K3 surface and its opposite realize the lower
bounds. 

Note that for X = #p (K3) # #q S 2 , we have
QX = (⊕3p H) ⊕(− ⊕2p E8 ) ⊕(⊕q H) = (⊕3p+q H) ⊕(− ⊕2p E8 )
650 18. SEIBERG-WITTEN THEORY

For this X and all known examples of smooth, compact, simply connected 4-
manifolds with even, indefinite QX as in (18.10), one has M ≥ 32 N . For such
QX , we have b2 (X) = 2M + 8N and |sig(X)| = 8N . Thus,
11
b2 (X) ≥ 8 |sig(X)| ⇔ 8(2M + 8N ) ≥ 11 · 8N
⇔ 16M ≥ (88 − 64) N = 24N ⇔ M ≥ 32 N.
The conjecture that M ≥ 32 N is the same as
Conjecture 1 (The 11 8 -conjecture). Let X be a simply connected, smooth,
compact 4-manifold with QX even. Then we have
11
b2 (X) ≥ 8 |sig(X)| ,
where b2 (X) denotes the second Betti number (the rank of QX ) and sig(X) is the
signature of QX .
While the 11
8 -conjecture may still be open, M. Furuta [156] used Seiberg-
Witten Theory to prove a strict bound of 10
8 , namely
10
b2 (X) = 2M + 8N ≥ 10N + 2 = 8 |sig(X)| + 2.
For X a smooth, simply-connected, compact orientable 4-manifold with self-
dual Betti number b2+ (X) > 1 and odd, Donaldson also found some invariants
(18.12) qdk̄ (X) : ×dk̄ /2 H2 (X; Z) −→ Z
which are symmetric polynomials (Donaldson’s polynomial invariants) of de-
gree dk̄ , k̄ = 1, 2, 3... . Here
(18.13) dk̄ = 8k̄ − 3(1 − b1 (X) + b2+ (X)),
which reduces to the even integer 8k̄ − 3(b2+ (X) + 1) under the above assumptions
on X. The number dk̄ in (18.13) is the virtual dimension of the moduli space
Mk̄ of anti-self-dual
 connections for a principal SU(2)-bundle P → X with k̄ :=
c2 P ×SU(2) C2 [X].
Remark 18.17. Donaldson (and almost everyone else now) have found it
convenient to work with moduli spaces of anti-self-dual (ASD) connections, instead
of the self-dual connections which we have dealt with in this book. However, one
can translate between the two by  changing the orientation of X. In particular,
we defined k := −c2 P ×SU(2) C2 [X], and replacing the orientation [X] by − [X],
Donaldson’s k (which we have denoted by k̄) is obtained. Moreover, b2+ (X) be-
comes b2− (X) under [X] → − [X].
As we will now explain, we implicitly obtained the analogous formula for vir-
tual dimension dk of the moduli space Mk of self-dual connections (where k :=
−c2 P ×SU(2) C2 [X]), namely
dk = 8k − 3 1 − b1 (X) + b2− (X) .


Indeed, this is the index of the operator T : Ω1 (E) → Ω0 (E) ⊕ Ω2− (E), where
E := P ×SU(2) su(2), and
T (sig) := δ ω sig, 21 (1 − ∗) Dω sig , for sig ∈ Ωk (E) .


We found (see (16.40) and (16.41) p.493, or (17.150) p.614) that


1
index(T ) = 2ch2 (E) [M ] − 2 dim E ·(χ(M ) − sig(M )) .
18.1. BACKGROUND AND SURVEY 651

To show that this is dk , recall from Proposition 16.8 (p.477) that for E 0 := P ×SU(2)
C2 and E := P ×SU(2) su(2)C , we have −4k = 4c2 (E 0 ) = c2 (E) = −ch2 (E), and so
index(T ) = 8k − 32 (χ(M ) − sig(M ))
2 − 2b1 (X) + b2+ (X) + b2− (X) − b2+ (X) − b2− (X)
3

= 8k − 2
= 8k − 3 1 − b1 (X) + b2− (X) .


Donaldson’s polynomial invariants can be used to distinguish smooth 4-manifolds


having the same intersection form. For example, we have
Proposition 18.18 (Donaldson, Kronheimer, [126], p.27). For any simply
connected, complex surface S with b2+ (S) > 3, there is a smooth 4-manifold which
is homotopy equivalent to S, but not diffeomorphic to S, nor to any complex surface.
In dimension 4, there is a radical difference between the smooth category and
the topological category. Unlike in the topological category, the function
[X]smooth 7→ [QX ]
is far from being 1-1 or onto. Moreover, differential topology in dimension 4 is
unlike that in other dimensions. For one thing, there are uncountably many exotic
R4 ’s (homeomorphic to R4 but not diffeomorphic to R4 or to each other). That
was proved by Taubes [406] based on ideas of Gompf, [180]. Some years before,
a smart combination of Donaldson’s and Freedman’s results had led to the
dramatic discovery of the fact that R4 admits more than one smooth structure. In
[349], the surprise is nicely described which arose from finding that “there exists a
compact subset in R4 which is not contained in a compact contractible submanifold
smooth in the strange structure”. There is no exotic Rn for n 6= 4. In dimensions
greater than 4, there can be only a finite number of nondiffeomorphic compact
manifolds which have the same homotopy type and Pontryagin classes. However,
by generalizing examples of Donaldson, Robert Friedman and John Morgan
proved
Theorem 18.19 (Friedman and Morgan,  [151]).
 There are infinitely many
2 2
nondiffeomorphic smooth structures on CP # #q CP for any q ≥ 9.

Recall that two C ∞ n-manifolds M0 and M1 are h-cobordant if there is a



C (n + 1)-manifold N , with ∂N = M0 ∪ M1 , such that M0 and M1 are strong
deformation retracts of N . The triple (N, M0 , M1 ) is called an h-cobordism.
Theorem 18.20 (h-Cobordism Theorem, [401]; see also [299]). Let the triple
(N, M0 , M1 ) be an h-cobordism with N (and hence M0 and M1 ) simply-connected.
If n = dim(Mi ) ≥ 5, then N is diffeomorphic to M0 × [0, 1] , and consequently M0
and M1 are diffeomorphic.
We can use Theorem 18.19 to show that the analog of Smale’s h-cobordism
theorem fails dramatically in dimension 4 in the smooth category. In contrast,
Friedman showed that the h-cobordism theorem holds in the topological category
in dimension 4. First, there is the result of C.T.C. Wall, [427]:
Theorem 18.21 (Wall 1964). Two simply-connected, smooth, compact 4-
manifolds with isomorphic intersection forms are h-cobordant.
652 18. SEIBERG-WITTEN THEORY

Thus, there is an h-cobordism N between CP2 #(#9 CP2 ) and any of its infin-
itely many exotic versions, say X. However, since CP2 #(#9 CP2 ) and X are not
diffeomorphic, N cannot be diffeomorphic to a product.
Let X be a smooth, compact, oriented 4-manifold (not necessarily simply-
connected). For any L ∈ H 2 (X; Z) (which can be interpreted as a complex line
bundle) we will later define a Seiberg-Witten invariant SW (L) ∈ Z, provided
b2+ (X) ≥ 2. Thus,
(18.14) SW : H 2 (X; Z) −→ Z.
Note . More precisely, as we shall explain below on pp.692-693, Definition
(18.101), the Seiberg-Witten invariants depend a priori on a Spinc structure, not
on a class in H 2 (X; Z). The map Spinc (X) → H 2 (X) taking s to c1 (s) is neither
injective nor surjective in general. Its image is the space of characteristic elements
(lifts of w2 ). Its kernel can be non-zero when H1 (X; Z2 ) is non-zero. So, liter-
ally, we can’t really write SW : H 2 (X; Z) → Z. Until further explanation, (18.14)
should be read as an affine statement: if we fix a Spinc structure, others are in 1-1
correspondence with H 2 (X; Z).
Moreover, we will see that (as with the Donaldson polynomials) the Seiberg-
Witten invariants are all 0, unless b1 (X) + b2+ (X) is odd. Reportedly (see [265]),
many (if not all) of the previous instanton results of Donaldson and his coworkers
(e.g., Kronheimer, Friedman, Morgan, Mrowka) can be proved more easily
and pushed further using Seiberg-Witten invariants rather than using Donaldson’s
invariants. Moreover, Taubes has proven a number of new results concerning
symplectic manifolds, which we describe following a few definitions.
Remark 18.22. There are two famous results that have instanton but appar-
ently no Seiberg-Witten or Heegård-Floer proofs: one is Taubes’ above mentioned
theorem that there exist uncountably many C ∞ structures on R4 , and the other is
Furuta’s theorem that the 3-dimensional homology cobordism group is infinitely
generated.
Results Concerning Symplectic Manifolds.
Definition 18.23. A symplectic form or symplectic structure on a smooth
manifold X is a closed 2-form ω on X which is nondegenerate in the sense that
∼ ∗
ω[ (V ) := ω(·, V ) defines an isomorphism Tx X −→ (Tx X) at each x ∈ X. The pair
(X, ω) is called a symplectic manifold. A diffeomorphism f : X1 → X2 between
two symplectic manifolds (X1 , ω1 ) and (X2 , ω2 ) such that f ∗ ω2 = ω1 is called a
symplectomorphism. If such symplectomorphism exists, then (X1 , ω1 ) and
(X2 , ω2 ) are called symplectomorphic.
For a symplectic manifold (X, ω), one can choose (uniquely up to homotopy)
an almost complex structure J which is compatible with ω in the sense that
hJ (V, W ) := ω(V, JW ) is a Riemannian metric. Indeed, let Met(M ) ⊂ T 0,2 (M )
denote the convex set of metric tensors and let J (M, ω) be the set of complex
structures compatible with ω. See also our discussion of almost complex structures
above on p.607. Following [288, p.61], we will define a canonical surjective function
R : Met(X) → J (X, ω). Then not only is J (X, ω) nonvoid, but also given J0 , J1 ∈
J (X, ω) with J0 = R(g0 ) and J1 = R(g1 ), we have a homotopy
J(t) := R((1 − t) g0 + tg1 ) , connecting J0 with J1 in J (X, ω).
18.1. BACKGROUND AND SURVEY 653

For g ∈ Met(X), we define R(g) ∈ J (X, ω) as follows. Let A ∈ Aut(T X) be


determined by g(AV, W ) = ω(V, W ). Then the adjoint A∗ of A relative to g is
defined by
g(V, A∗ W ) = g(AV, W ) , and A∗ = −A, (i.e., A is g-skew-symmetric) since
g(V, A∗ W ) = g(AV, W ) = ω(V, W ) = −ω(W, V ) = −g(AW, V ) = g(V, −AW ) .
Now −A2 = A∗ A ∈ Aut(T X) is g-symmetric and positive-definite, and hence it has
a unique g-symmetric, positive-definite square root, say S ∈ Aut(T X), such that
S 2 = A∗ A = −A2 . Let R(g) := J := S −1 A. Then J 2 = S −2 A2 = −A−2 A2 = − Id.
Moreover,
ω(V, JW ) = g(AV, JW ) = g AV, S −1 AW


is symmetric in V and W , and ω(V, JV ) = g AV, S −1 AV > 0 for V 6= 0, since




A is invertible and S −1 is positive-definite. Thus, kJ (V, W ) := g AV, S −1 AV =




ω(V, JW ) is a positive-definite metric, and so R(g) ∈ J (X, ω).


For a compatible J ∈ J (X, ω) , the associated canonical class
KX := −c1 (TJ X) = c1 (ΛJ2,0 (X))
then only depends only on ω. In the case dim X = 4, ω is self-dual relative to the
metric g and orientation ω ∧ ω.
Theorem 18.24 (Taubes, [409]). If (X, ω) is a compact, symplectic 4-manifold
with b2+ (X) ≥ 2, then we have
SW (KX ) = ±1 and (KX ` [ω]) [X] ≥ 0.
If SW (k) 6= 0 for any other class k ∈ H 2 (X; Z), then
|k ` [ω]| ≤ KX ` [ω],
with equality only if k = ±KX .
Corollary 18.25. If (X, ω) is a compact, symplectic 4-manifold, then b1 (X)+
2+
b (X) is odd; otherwise, SW (KX ) = 0, contrary to Theorem 18.24.
Remark 18.26. The preceding Corollary, being a nice application of the highly
intricate Theorem 18.24, can be proved independently: Since X is symplectic, it
admits an almost complex structure. Then it is algebraic topology to show that
a closed four-manifold admits an almost complex structure iff sig(X) + χ(X) ≡ 0
mod 4. This is equivalent to b1 + b2+ odd. Similarly, in Corollary 18.28, the
existence of an almost complex structure for #3 CP2 can be derived immediately
from sig(X) + χ(X) = 3 + 5 ≡ 0 mod 8. Proving the following Connected Sums
Corollary 18.27 is a nice exercise, but clearly more demanding than the two other
corollaries.
Corollary 18.27. If X = X1 #X2 , where neither X1 nor X2 have negative-
definite intersection forms (i.e., b2+ (X1 ) and b2+ (X2 ) are strictly positive), then
X does not have a symplectic structure compatible with the given orientation. For
example, when n ≥ 2 and m ≥ 0,
     
(18.15) #n CP2 # #m CP2 = CP2 # #n−1 CP2 # #m CP2
admits no symplectic structure.
Proof. It is known that such X have Seiberg-Witten invariant 0. 
654 18. SEIBERG-WITTEN THEORY

The following proves a conjecture of Mark Gotay.


Corollary 18.28. While #3 CP2 has an almost complex structure, it has no
symplectic form.
Proof. 1. There are many ways to prove that the connected sum #3 CP2
admits an almost complex structure. Here is one way: In order for c ∈ H 2 (X; Z) to
be the first Chern class of an almost complex structure on the compact, orientable
4-manifold X, it is necessary and sufficient that c = w2 (X) mod 2, and QX (c, c) =
2χ(X) + 3 sig(X). For X = #3 CP2 ,
∼ ∼
H 2 (X; Z) −→ ⊕3 Z, QX −→ diag (1, 1, 1) , w2 (X) ←→ (1, 1, 1) and
2χ(X) + 3 sig(X) = 2 · 5 + 3 · 3 = 19.
If c ←→ (k1 , k2 , k3 ) , then
c = w2 (X) mod 2 ⇐⇒ k1 , k2 , and k3 are odd, and
X
X
Q (c, c) = 2χ(X) + 3σ(X) ⇐⇒ ki2 = 19.
Thus, (k1 , k2 , k3 ) = (1, 3, 3) will do.
2. Of course, #3 CP2 has no symplectic form by Corollary 18.27. 
Let K ∈ H 2 (CP2 ; Z) be the Chern class of the canonical line bundle Λ2,0 (CP2 )
of CP2 . We have
K = −c1 (T CP2 ) = −3[ξ] ˆ = −3[κ/π],
where ξˆ denotes the dual of the tautological bundle ξ := (p, v) ∈ CP2 × C3 : v ∈ p


and κ denotes the Kähler form of CP2 with the standard Fubini-Study metric. Note
that
(K ` [κ]) CP2 = (−3[κ/π] ` [κ]) CP2 = −3/π < 0.
   

More generally, using Seiberg-Witten theory Clifford Taubes proved


Theorem 18.29 (Taubes, [409]). For any symplectic structure CP2 , ω , let K


denote the Chern class of the canonical line bundle of an almost complex structure J
2
compatible with ω (i.e.,
 ω(X, Y ) = h(JX, Y ) for a Riemannian metric h on CP ).
2
Then (K ` [ω]) CP < 0.
Purely Geometric Applications. The Donaldson invariants and (more sim-
ply) SW invariants can also be used to prove the Thom Conjecture in considerable
generality and specifically the classical case of CP2 (see [262]), namely
Theorem 18.30 (Thom Conjecture). In a compact Kähler surface X, any
complex curve C (i.e., a holomorphic immersion of a Riemann surface) has the
smallest genus among all surfaces immersed in X which are homologous to C.
For a purely geometrical application, we have the result
Theorem 18.31 (Witten, [447]). No compact 4-manifold X with b1 (X) > 0
and with a nonzero SW -invariant can have a metric with positive scalar curvature.
We also have the following result of C. LeBrun
Theorem 18.32 (LeBrun, [275]). Any compact, Einstein 4-manifold X with
a nonzero SW -invariant satisfies χ(X) ≥ 3 sig(X).
18.2. SPINc STRUCTURES AND THE SEIBERG-WITTEN EQUATIONS 655

Remark 18.33. A result of Hitchin ([214, 1974]) says that an Einstein 4-


manifold X satisfies χ(X) ≥ 32 |sig(X)| , with equality implying that X is flat or
is covered by a K3 surface (with χ = 24 and sig = −16). Thus, Theorem 18.32
provides a stronger result when sig(X) > 0. In the case of a complex surface,
χ(X) ≥ 3 sig(X) is the same as c2 ≥ c1 2 − 2c2 or the Miyaoka-Yau inequality
3c2 (X) ≥ c1 (X)2 .

2. Spinc Structures and the Seiberg-Witten Equations


Admittance of Spinc Structures. Recall from (17.164), p.630, that
U(1) × Spin(n)
Spinc (n) := U(1) Spin(n) ∼
= .
{±(1, 1)}
Moreover, recall that there is a homomorphism
rc : Spinc (n) −→ U(1) × SO(n) , given by
rc (zg) := z 2 , c(g) ,

(18.16)
with kernel {[1, I], [1, −I]} ∼
= Z2 , where c : Spin(n) → SO(n) denotes the 2-fold
cover.
Definition 18.34. A Spinc structure for an oriented Riemannian n-manifold
X consists of a principal Spinc (n)-bundle π c : PSpinc (n) → X, a principal U(1)-
bundle π1 : PU(1) → X, and an rc -equivariant bundle map
(18.17) πrc : PSpinc (n) −→ PU(1) × F X,
where F X → X denotes the principal SO(n)-bundle of oriented orthonormal frames.
Here, PU(1) × F X denotes the fibered (as opposed to Cartesian) product consist-
ing of pairs (p1 , p2 ) with π1 (p1 ) = πF (p2 ). Naturally, PU(1) × F X is a principal
U(1) × SO(n)-bundle. The Chern class in H 2 (X; Z) of the line bundle associated
to PU(1) → X is called the canonical class of the Spinc (n)-structure.
By Proposition 17.20 (p.528) an oriented, Riemannian manifold X has a spin
structure if and only if w2 (X) = 0 ∈ H 2 (X; Z2 ), where w2 (X) denotes the sec-
ond Stiefel-Whitney class defined in Definition 17.18, p.527. The condition under
which X admits a Spinc (n)-structure is much weaker. Indeed, we will show that
any oriented, Riemannian 4-manifold X has a Spinc (n)-structure. To describe the
condition, observe that the short exact sequence
m r
0 −→ Z −→ Z −→ Z/2Z = Z2 −→ 0 (where m denotes multiplication by 2)
induces a long exact sequence
(18.18)
m r δ
... −→ H i (X; Z) −→

H i (X; Z) −→

H i (X; Z2 ) −→ H i+1 (X; Z) −→ ...,
where δ denotes the Bockstein homomorphism of homological algebra (see [403]).
Theorem 18.35. An oriented Riemannian n-manifold X has a Spinc structure
if and only if the Stiefel-Whitney class w2 (X) ∈ H 2 (X; Z2 ) is the mod 2 reduction
r∗ ([w̃]) (see (18.18)) of an integral class [w̃] ∈ H 2 (X; Z). Then [w̃] is the canonical
class of some Spinc (n)-structure for X.
656 18. SEIBERG-WITTEN THEORY

Proof. Up to smooth equivalence, complex line bundles (or equivalently prin-


cipal U(1)-bundles) over X may be regarded (via their transition functions) as
elements of H 1 (X; U(1)) (C̆ech cohomology with coefficients in the sheaf of germs
of U(1)-valued functions); see the beginning of Section .17.2 where the construction
was carried out for the frame bundle. Associated with the exact sequence
(18.19) 0 −→ Z −→ R −→ U(1) −→ 0,
we have the long exact C̆ech cohomology sequence
(18.20)
c1
... −→ H 1 (X; R)
e −→ H 1 (X; U(1)) −→ H 2 (X; Z) −→ H 2 (X; R)
e −→ ...,

The homomorphism c1 : H 1 (X; U(1)) → H 2 (X; Z) is defined in the same way as the
homomorphism w2 : H 1 (X; SO(n)) → H 2 (X; Z2 ) was defined in Definition 17.18,
p.527, where the exact sequence in that case was 0 → Z2 → Spin(n) → SO(n) →
0. By one definition, the first Chern class of a line bundle in H 1 (X; U(1)) is
just its image under c1 in (18.20). In (18.20), Re is the sheaf of germs of C ∞ R-
valued functions on X, as opposed to the constant sheaf R whose C̆ech cohomology
coincides with the usual de Rham cohomology H ∗ (X; R). It is not hard to show

(via partitions of unity) that H i (X; R)
e = 0 for i > 0. Thus, c1 : H 1 (X; U(1)) −→
H 2 (X; Z); i.e., up to isomorphism a principal U(1)-bundle are determined by its
first Chern class. Associated with the sequence
i c
0 −→ Z2 −→ Spin(n) −→ SO(n) −→ 0,
we have exact C̆ech cohomology sequence
(18.21)
i∗ c∗ w2
... −→ H 1 (X; Z2 ) −→ H 1 (X; Spin(n)) −→ H 1 (X; SO(n)) −→ H 2 (X; Z2 ),
where for nonabelian groups G, H 1 (X; G) is a pointed set instead of a group. By
Definition 17.18, p.527, w2 assigns to each equivalence class of SO(n)-bundles, its
second Stiefel-Whitney class. Thus, for a principal SO(n)-bundle ξ, w2 (ξ) is the
obstruction to finding a spin structure covering ξ, and w2 ([F X]) (or w2 (X)) the is
obstruction in the special case of the bundle F X of oriented orthonormal frames.
For the sequence
i rc
0 −→ Z2 −→ Spinc (n) −→ U(1) × SO(n) −→ 0,
there is an exact sequence
i rc
... −→ H 1 (X; Z2 ) −→

H 1 (X; Spinc (n)) −→

c̃ +w
(18.22) −→ H 1 (X; U(1)) ⊕ H 1 (X; SO(n)) 1−→2 H 2 (X; Z2 ),
c ∗ r
where c̃1 denotes the composition H 1 (X; U(1)) → 1
H 2 (X; Z) → H 2 (X; Z2 ). Sup-
−1
pose r∗ ([w̃]) = w2 (X) for some [w̃] ∈ H (X; Z). Then c1 ([w̃]) ∈ H 1 (X; U(1))
2

defines (up to equivalence) a principal U(1)-bundle PU(1) . Thus, the obstruction to


obtaining a Spinc structure over PU (1) × F X with canonical class [w̃] is
(c̃1 + w2 ) c−1 −1
 
1 ([w̃]) ⊕ [F X] = c̃1 c1 ([w̃]) + w2 ([F X])
(18.23) = r∗ ([w̃]) + w2 (X) = w2 (X) + w2 (X) = 0,
since w2 (X) is in the Z2 -module H 2 (X; Z2 ).
18.2. SPINc STRUCTURES AND THE SEIBERG-WITTEN EQUATIONS 657

Conversely, if there is a Spinc structure over PU (1) × F X with canonical class


[w̃] ∈ H 2 (X; Z), then
c̃1 c−1

1 ([w̃]) + w2 ([F X]) = 0,
in which case
w2 ([F X]) = c̃1 c−1 = r∗ ◦ c1 ◦ c−1
 
1 ([w̃]) 1 [w̃] = r∗ ([w̃]) . 
Remark 18.36. By the exact sequence (18.18), any two integral classes [w̃] ∈
r∗−1 (w2 (X)) differ by an element of
 
r∗
Ker H 2 (X; Z) −→ H 2 (X; Z2 ) ∼= m∗ H 2 (X; Z) = 2H 2 (X; Z).

Thus, 2H 2 (X; Z) acts on the set r∗−1 (w2 (X)) of choices for [w̃]. Once a choice of the
canonical class [w̃] is made, any
 two choices of isomorphism classes [PSpin
c (n) ] differ
1 c
by the image of i∗ H (X; Z2 ) in (18.22). Assuming that X has a Spin structure,
the set of Spinc structures on X is then parametrized by 2H 2 (X; Z)⊕i∗ H 1 (X; Z2 ) .


We thank Bob Little for providing the references [210] and [285] for the fol-
lowing result.
Theorem 18.37. The Stiefel-Whitney class w2 (X) ∈ H 2 (X; Z2 ) of a compact,
orientable Riemannian 4-manifold X is in fact the mod 2 reduction of some in-
tegral class [w̃] ∈ H 2 (X; Z). Consequently, any compact, orientable Riemannian
4-manifold admits a Spinc -structure.

Proof. We need to show that w2 (X) ∈ r∗ H 2 (X; Z) , for r∗ : H 2 (X; Z) →
H 2 (X; Z2 ) as in (18.18). Let T 2 (X) denote the torsion subgroup of H 2 (X; Z).
Suppose that we can show that
⊥
r∗ H 2 (X; Z) = r∗ T 2 (X) , where

(18.24)
⊥
r∗ T 2 (X) := z ∈ H 2 (X; Z2 ) : z ` r∗ (t) = 0, ∀t ∈ T 2 (X) .


Then it would suffice to show that w2 (X) ` r∗ (t) = 0 for all t ∈ T 2 (X). According
to Wu’s formula
w2 (X) ` r∗ (t) = r∗ (t) ` r∗ (t) = r∗ (t ` t) = 0,
where the last equality follows since t ` t is in the torsion subgroup of H 4 (X; Z) ∼
= Z,
whence t ` t = 0. Thus it remains to show (18.24). Now
⊥
r∗ H 2 (X; Z) ⊆ r∗ T 2 (X) ,


since r∗ (h) ` r∗ (t) = r∗ (h ` t) = 0 because h ` t is a torsion element of the


torsionless group H 4 (X; Z) ∼
= Z. Thus, it suffices to prove that dim r∗ H 2 (X; Z) =
2
⊥
dim r∗ T (X) as vector spaces over Z2 . Since
` : H 2 (X; Z2 ) × H 2 (X; Z2 ) −→ H 4 (X; Z2 ) ∼
= Z2
 ⊥ 
is nondegenerate, dim r∗ T 2 (X) = dim H 2 (X; Z2 ) − dim r∗ T 2 (X) . Let bi =
rank H i (X; Z) and let ci denote the number of cyclic summands of H i (X; Z) of order
equal to a power of 2 in the primary decomposition. We have dim r∗ T 2 (X) = c2 ,
and dim r∗ H 2 (X; Z) = b2 + c2 . Moreover,
 
m∗
δ H 2 (X; Z2 ) = Ker H 3 (X; Z) −→ H 3 (X; Z) ∼

= ⊕c3 Z2 .
658 18. SEIBERG-WITTEN THEORY

Thus, from the exact sequence


m r δ m
... → H 2 (X; Z) →∗ H 2 (X; Z) →

H 2 (X; Z2 ) → H 3 (X; Z) →∗ H 3 (X; Z) → ...,
we obtain
dim H 2 (X; Z2 ) = dim r∗ H 2 (X; Z) + dim δ H 2 (X; Z2 ) = b2 + c2 + c3 .
 

Hence,
⊥
dim r∗ T 2 (X) = dim H 2 (X; Z2 )− dim r∗ T 2 (X)


= b2 + c2 + c3 − c2 = b2 + c3 .

Now, dim r∗ H 2 (X; Z) = b2 + c2 . Thus, it remains to show that c2 = c3 . Using
the Universal Coefficient Theorem ([185, p.194] or [403, Corollary 4, p.244]),
(18.25) H 3 (X; Z) ∼
= F3 (X) ⊕ T 2 (X),
where T i (X) denotes the torsion subgroup of Hi (X; Z) and Fi (X) = Hi (X; Z)/T i (X).
By Poincaré duality (which we can use, since X is orientable),
∼ ∼
(18.26) H 2 (X; Z) −→ H2 (X; Z) −→ F2 (X) ⊕ T 2 (X).
Thus, the torsion subgroup of H 3 (X; Z) is the same as that of H 2 (X; Z), and hence
c3 = c2 . 

In the case of compact, almost complex 2m-manifolds X, w2 (X) is the mod


2 reduction of the integral class c1 (Λm,0 (X)) = −c1 (T X), and hence X admits a
Spinc -structure with canonical class c1 (Λm,0 (X)). Since it might be the case that
H 1 (X; Z2 ) 6= 0, this does not uniquely determine a Spinc structure for X. However,
we will now show that there is a standard Spinc structure for X, using the existence
of the lift h : U(m) → Spinc (2m) of det ×ι : U(m) → U(1) × SO(2m).
Proposition 18.38. Any Riemannian 2m-manifold X with almost complex
structure J, with a metric h compatible with J (i.e., h(JV, JV ) = h(V, V )) admits a
standard Spinc -structure, with canonical class being the Chern class of the canonical
line bundle Λm,0 (X).
Proof. Let PU(m) denote the unitary frame bundle of X. The homomorphism
j : U(m) → Spinc (2m) of Proposition 17.75 (p.631) provides a principal Spinc (2m)-
bundle PSpinc → X and j-equivariant map
PU(m) −→ PU(m) ×j Spinc (2m) =: PSpinc
(see Proposition 15.20, p.407). The homomorphism det ×ι : U(m) → U(1) ×
SO(2m) yields a principal bundle PU(1)×SO(2m) → X which can be identified with
PU(1) × F X, where F X denotes the oriented orthonormal frame of X and PU(1)
denotes the frame bundle for the canonical line bundle Λm,0 (X); for this one may
use Remark 15.21 (p. 408) and note that a frame in PU(m) determines a frame
in PU(1) and a (oriented) frame in F X. The homomorphism rc : Spinc (2m) →
U(1) × SO(2m) (that satisfies rc ◦ j = det ×ι) then yields an rc -equivariant map

PSpinc −→ PU(1)×SO(2m) −→ PU(1) × F X,
which provides the desired standard Spinc -structure. 
18.2. SPINc STRUCTURES AND THE SEIBERG-WITTEN EQUATIONS 659

π c
Spinc Dirac Operators. For a Spinc structure PSpinc (n) →
r
PU(1) × F X we
c
have a Spin (n)-principal bundle
π c
π c : PSpinc (n) −→
r
PU(1) × F X −→ X.
Moreover, there is a representation
ρC
Spinc (n) −→ U(Σ2m )
which is just the restriction of ρC : Cl(n) → End(Σ2m ) to Spinc (n) ⊂ Cl(n). As
in (17.14), p. 17.14, there are also the Spinc (n)-invariant eigenspaces Σ±
2m of νC :=
im e1 · · · e2m . Thus, we may form the associated bundles
±
Σ± + −
c (X) := PSpinc (n) ×ρC Σ2m −→ X and Σc (X) := Σc (X) ⊕ Σc (X) .

Let ω be a connection on PU(1) and let θ denote the Levi-Civita connection on


c
F X. Then ω ⊕ θ is a connection on PU(1) × F X and (ω ⊕ θ) := πr∗c (ω ⊕ θ) may
be regarded as a connection on PSpinc (n) under the isomorphism of Lie algebras

spinc (n) −→ u(1) ⊕ so(n) induced by the double-covering rc : Spinc (n) → U(1) ×
c
SO(n). Hence, there is a covariant differentiation operator for (ω ⊕ θ) , say
c
∇(ω⊕θ) : C ∞ (Σc2m (X)) −→ C ∞ Λ1 (X) ⊗ Σc2m (X) .


Moreover, there is a well-defined Clifford multiplication


c
Λ1 (X) ⊗ Σc (X) −→ Σc (X) with c Λ1 (X) ⊗ Σ± ∓

c (X) ⊂ Σc (X) ,

induced by the representation ρC : Cl(n) → End(Σ2m ). Thus, we have a Spinc -


Dirac operator
c
c ◦∇(ω⊕θ)
Dc : C ∞ (Σc (X)) −→ C ∞ (Σc (X)) .
For index computations, as well as for a better understanding of this operator, we
will show that Dc is locally a twisted Dirac operator. The twist is by a locally
defined square root of the canonical line bundle
L := PU(1) ×µ C −→ X,
where µ : U(1) → U(C) is simply given by µ(z) w = zw; note that c1 (L) is (by
πr c
definition) the canonical class of the Spinc structure PSpinc (n) → PU(1) × F X. Let
BX be a coordinate ball of X. Since PU(1) and F X are trivial over BX , there
0
is a µ2 -equivariant double cover PU(1) → PU(1) |BX and a trivial spin structure
PSpin(n) → F X|BX . Define a 2-fold cover
χ : U(1) × Spin(2m) −→ U(1) Spin(2m) = Spinc (2m) by χ(z, σ) := zσ.
Since PSpinc (n) |BX is also trivial, we have a χ-equivariant covering
0 πχ
PU(1) × PSpin(n) −→ PSpinc (n) |BX
and a 4-fold rc ◦ χ-equivariant covering
0 πχ πr c 
πrc ◦ πχ : PU(1) × PSpin(n) −→ PSpinc (n) |BX −→ PU(1) × F X |BX

Let L0 (BX ) := PU(1)
0
×µ C. By Proposition 15.23 (p. 409) L0 (BX )⊗L0 (BX ) −→ L|BX
naturally, since φ : C ⊗ C → C (given by φ(w1 ⊗ w2 ) := w1 w2 ) satisfies
φ (µ ⊗ µ)(z) (w1 ⊗ w2 ) = φ(zw1 ⊗ zw2 ) = z 2 w1 w2 = µ µ2 (z) (φ(w ⊗ w0 )) .
  
660 18. SEIBERG-WITTEN THEORY

Also using Proposition 15.23 again, we have



φ× : L0 (BX ) ⊗ Σ2m (BX ) −→ Σc2m (X) |BX .
Indeed in this case, let φ : C ⊗ Σ2m → Σ2m be given by φ(w ⊗ ψ) = wψ, and note
that for (z, σ) ∈ U(1) × Spin(2m), we have χ-equivariance in the sense
 
φ (z, σ) ·(w ⊗ ψ) = φ zw ⊗ (σ · ψ) = zw(σ · ψ)
= zσ ·(wψ) = χ(z, σ) · wψ = χ(z, σ) · φ(w ⊗ ψ) .
Because of the naturality of the isomorphism φ× , the twisted Dirac operator
0
DL (BX ) ∈ End(C ∞ (L0 (BX ) ⊗ Σ2m (BX )))
corresponds to Dc ∈ End(C ∞ (Σc2m (X) |BX )), in the sense that
0
DL (BX ) (ψ) = φ−1

× Dc (φ× (ψ)) .

Proposition 18.39. For a Spinc -structure πrc : PSpinc (n) → PU(1) × F X, let
Ω ∈ Ω2 (X, iR) denote the curvature of the connection ω on PU(1) , and let S be the
ω

scalar curvature of X. For ψ ∈ C ∞ (Σc (X)), and an orthonormal frame E1 , . . . , En


at x ∈ X, let Rωx ∈ Cl(T M )x be defined by
X
Rω 1
x = 2 Ωω ω ω
jk Ej Ek , and let (R ψ)(x) := Rx · ψ(x) .
j,k

Then, we have
(18.27) Dc2 ψ = −∆ψ + 21 Rω ψ + 14 Sψ.
Proof. Since Dc is locally a twisted Dirac operator, this (18.27) is a conse-
quence of Proposition 17.27 (p. 537). Note that we have the factor 12 in 21 Rω ψ since
the local Dirac operator is twisted by L0 = PU(1)
0
×µ C, rather than

L = PU(1) ×µ C −→ L0 ⊗ L0 ,
0
and the curvature Ωω ∈ Ω2 (X, iR) of the lift ω 0 := 12 π1∗ (ω) of ω to PU(1)
0
is 21 Ωω .
0
Note since π1 : PU(1) → PU(1) |BX is µ2 -equivariant, π1∗ (A∗ ) = 2A∗ for A ∈ u(1) so
that (as required of a connection)
ω 0 (A∗ ) = 12 π1∗ (ω)(A∗ ) = 12 ω(π1∗ (A∗ )) = 12 ω(2A∗ ) = A. 
Motivating the Seiberg-Witten Equations. One way to motivate the
Seiberg-Witten equations without delving into supersymmetric QCD ([386]) is to
try to generalize to 4-manifolds with Spinc structures, the fact that spin manifolds
with positive scalar curvature have no nonzero harmonic spinors (see Corollary
17.28, p. 537). Mimicking the computation in the proof of Corollary 17.28, we have
for ψ ∈ C ∞ (Σc (X))
2
Dc2 ψ = Dc2 ψ, ψ = −∆ψ + 12 Rω ψ + 14 Sψ, ψ
 

c 2
(18.28) = ∇(ω⊕θ) ψ + 12 (Rω ψ, ψ) + 14 (Sψ, ψ) .
Note that
D X E
hRω ψ, ψi = 1
2 Ωω
jk Ej Ek · ψ(x) , ψ(x)
j,k
X
= 1
2 Ωω
jk hEj Ek · ψ(x) , ψ(x)i .
j,k
18.2. SPINc STRUCTURES AND THE SEIBERG-WITTEN EQUATIONS 661

Let Ωω = Ω+ + Ω− be the decomposition of Ωω into its self-dual and anti-self-dual


parts, and let Rω = R+ +R− be the corresponding decomposition of Rω . We claim
that
(Rω ψ, ψ) = R+ ψ + , ψ + + R− ψ − , ψ − ,
 

where ψ = ψ + + ψ − is the decomposition of


ψ ∈ C ∞ (Σc (X)) = C ∞ Σ+ ∞
Σ−
 
c (X) ⊕ C c (X) .

Indeed, for νC := i2 e1 e2 e3 e4 = −e1 e2 e3 e4 , we have νC R± = ±R± , since νC2 = Id


and
νC e2 e3 = −e1 e2 e3 e4 e2 e3 = −e1 e2 e3 e2 e3 e4 = e1 e4 ,
νC e3 e1 = −e1 e2 e3 e4 e3 e1 = e2 e3 e4 e3 = e2 e4 , and
νC e1 e2 = −e1 e2 e3 e4 e1 e2 = −e1 e2 e1 e2 e3 e4 = e3 e4 .

Thus, νC R± ψ + = ±R± ψ + ⇒ R± ψP +
∈ C ∞ (Σ±
c (X)). However, we know that
± + ∞ 1
R ψ ∈ C (Σc (X)), since Rx = 2 j,k Ωω
+
jk j k is in the even part of Cl(T M )x .
E E
Thus, R ψ = 0 and R ψ ∈ C (Σc (X)). Similarly, R+ ψ − = 0 and R− ψ − ∈
− + + + ∞ +

C ∞ (Σ−
c (X)). Thus, (18.28) becomes

2 c 2
Dc2 ψ = ∇(ω⊕θ) ψ 1
R+ ψ + , ψ + + 1
R− ψ − , ψ − + 41 (Sψ, ψ) .
 
+ 2 2

In particular,
2 c 2
Dc2 ψ + = ∇(ω⊕θ) ψ + 1
R+ ψ + , ψ + + 1
Sψ + , ψ + .
 
(18.29) + 2 4

Note that at x ∈ X,
D X E
Rω (x) ψ + (x) , ψ + (x) = 1
2 Ωω+
jk Ej E k · ψ +
(x) , ψ +
(x)
j,k
X
= 1
2 Ωω+ + +
jk Ej Ek · ψ (x) , ψ (x)
j,k
Ωω+ ψ + (x) ,

(18.30) = x ,q

where q(ψ + (x)) is the self-dual part of


X
Q ψ + (x) := 12 Ej Ek · ψ + (x) , ψ + (x) ϕj ∧ ϕk ∈ Λ2x (X, iR) .

(18.31)
j,k

Note that Q(ψ + (x)) has values in iR, since


Ej Ek · ψ + (x) , ψ + (x) ϕj ∧ ϕk = ψ + (x) , Ek Ej ψ + (x) ϕj ∧ ϕk
= − ψ + (x) , Ek Ej ψ + (x) ϕk ∧ ϕj = −hEk Ej ψ + (x) , ψ + (x)iϕk ∧ ϕj .

The Unperturbed and the Perturbed S-W Equations. Since q appears


throughout S-W theory, it is fitting to make a formal definition.
Definition 18.40. The quadratic map q : C ∞ (Σ+ c (X)) → Ω
2+
(X, iR) is de-
fined by
q ψ + (x) = 21 (1 + ∗) Q ψ + (x) ,
 

where Q(ψ + (x)) is given by (18.31). See also Remarks 18.42 and 18.44 below.
662 18. SEIBERG-WITTEN THEORY

From (18.29) and (18.30), we obtain


2 c 2
Dc2 ψ + = ∇(ω⊕θ) ψ + 1
Ωω+ , q ψ + 1
Sψ + , ψ + ,
 
+ 2 + 4

and perhaps the simplest way to guarantee that


Dc ψ + = 0 and 0 6= scalar curvature S ≥ 0 ⇒ ψ + = 0,
2
is to assume that Ωω+ = q(ψ + ) so that 12 hΩω+ , q(ψ + )i = 12 |q(ψ + )| ≥ 0. We write
Dc as Dcω and Ω+ as Ωω+ , to indicate their dependence on ω.
Definition 18.41. (a) Based on the above motivation, the (unperturbed) S-W
equations (S-W for Seiberg-Witten) for the pair (ω, ψ + ) are simply
Dcω ψ + = 0 and Ωω+ = q ψ + .


An immediate consequence of our derivation is that when 0 6= S ≥ 0, (ω, ψ + ) is


a solution of the S-W equations if and only if ψ + = 0 and Ωω+ = 0 (i.e., ω has
anti-self-dual curvature).
(b) Since the space of solutions of the S-W equations is not always nice, it is
generally necessary to study the perturbed S-W equations
Dcω ψ + = 0 and Ωω+ = q ψ + + η,

(18.32)
where η ∈ Ω2+ (X, iR) is a given self-dual 2-form which supplies the perturbation.
Remark 18.42. The quadratic map q : C ∞ (Σ+ c (X)) → Ω
2+
(X, iR) is induced
+ 2+
by a pointwise map qx : Σc (X)x → Λx (X,iR). This qx is derived from a purely
algebraic quadratic map q0 : Σ+2 → iΛ
2+
R4 , given by
 X +
q0 (ψ) = 12 ej ek ψ + , ψ + ej ∧ ek
j,k
= e2 e3 ψ + , ψ + e2 ∧ e3 + e1 e4 ψ + , ψ + e1 ∧ e4 + · · ·
1 1 + +
= 2 (e2 e3 + e1 e4 ) + 2 (e2 e3 − e1 e4 ) ψ , ψ e2 ∧ e3
+ 2 (e2 e3 + e1 e4 ) − 2 (e2 e3 − e1 e4 ) ψ , ψ + e1 ∧ e4
1 1 +
+ ···
1 + +
= 2 (e2 e3 + e1 e4 ) ψ , ψ (e2 ∧ e3 + e1 ∧ e4 ) + · · ·
= 1
2 (e2 e3 + e1 e4 ) · ψ + , ψ + (e2 ∧ e3 + e1 ∧ e4 )
+ 1
2 (e3 e1 + e2 e4 ) · ψ + , ψ + (e3 ∧ e1 + e2 ∧ e4 )
+ 1
2 (e1 e2 + e3 e4 ) · ψ + , ψ + (e1 ∧ e2 + e3 ∧ e4 ) .

Exercise 18.43. We can make q0 (ψ) more explicit by identifying Σ+ 2 with C


2

and (in view of Example 17.9, p. 516) by using the fact that Clifford multiplication
by 12 (e2 e3 + e1 e4 ) corresponds to quaternionic multiplication by i on H, or by iσ1
on C2 , where σ1 denotes the first Pauli matrix (see 14.32, p. 382), etc.. Verify that

(a) q0 (ψ) = i ψ2 ψ 1 + ψ1 ψ 2 (e2 ∧ e3 + e1 ∧ e4 )

+ ψ2 ψ 1 − ψ1 ψ 2 (e3 ∧ e1 + e2 ∧ e4 )
+ ψ1 ψ 1 − ψ2 ψ 2 (e1 ∧ e2 + e3 ∧ e4 ) ∈ iΛ2+ R4 .
 

1 2 4
(b) 2 |q0 (ψ)| = |ψ| .
18.3. GENERIC REGULARITY OF THE MODULI SPACES 663

Remark 18.44. Note that q0 (ψ) = qe0 (ψ, ψ) where the real bilinear form qe0 is
given by

qe0 (ψ, ξ) = i ψ2 ξ 1 + ψ1 ξ 2 (e2 ∧ e3 + e1 ∧ e4 )

+ ψ2 ξ 1 − ψ1 ξ 2 (e3 ∧ e1 + e2 ∧ e4 )
+ i ψ1 ξ 1 − ψ2 ξ 2 (e1 ∧ e2 + e3 ∧ e4 ) ∈ C ⊗ Λ2+ R4 ,
 

or equivalently
qe0 (ψ, ξ) = 21 (q0 (ψ + ξ) − q0 (ψ) − q0 (ξ))
+ 2i (q0 (ψ + iξ) − q0 (ψ) − q0 (ξ)) .

Also observe that qe0 (·, ·) is linear in the first slot, conjugate
 linear in the second
slot, and qe0 (ξ, ψ) = −e q0 (ψ, ξ) so that qe0 (ψ, ψ) ∈ iΛ2+ R4 . Similar statements hold
for the form
qe: C ∞ Σ+ ∞
Σ+ −→ Ω2+ (X, C)
 
(18.33) c (X) × C c (X)

associated with q : C ∞ (Σ+


c (X)) → Ω
2+
(X, iR).

3. Generic Regularity of the Moduli Spaces


Gauge Transformations and the Moduli Space of Solutions  of the
Perturbed S-W Equations. A certain  subgroup GA 1 P c
Spin (n) of the group
of gauge transformations GA PSpinc(n) acts on the space of solutions (ω, ψ + ) of
the perturbed S-W equations. In this context, the quotient space is known as the
moduli space. We will show that for a generic choice of the perturbation η, the
moduli space is a compact manifold. The Seiberg-Witten invariant is obtained by
integrating a certain form over this manifold. Much of the effort involved is directed
toward showing that this Seiberg-Witten invariant is well-defined, independent of
a suitable choice of perturbation η and Riemannian metric for X, so that it only
depends on the differentiable structure on X and the choice of Spinc structure.
Note . Our notation may appear to be a bit nonstandard. Readers trained
in Donaldson’s pioneering work may expect the gauge groups to be denoted by
G and the space of connections by A , possibly with some Sobolev decoration, see
our Translation Table 16.1 (p.461). Younger readers will not bother and, hopefully,
enjoy our rather classical style of notation.
 
The subgroup GA1 PSpinc(n) of the group GA PSpinc(n) of gauge transforma-
tions of PSpinc(n) → X is defined as follows. For s ∈ C ∞ (X, U(1)), note that
U(1) × Spin(n)
[s(x) , 1] = {±(s(x) , 1)} ∈ = Spinc (n)
{±(1, 1)}
is in the center of Spinc (n). Let Fs ∈ C ∞ PSpinc(n) , PSpinc(n) be given by


Fs (p) := p · [s(π c (p)) , 1] ,


 
where π c : PSpinc(n) → X, is in GA PSpinc(n) . Then Fs ∈ GA PSpinc(n) , since
(18.34) Fs (pg) = pg · [s(π c (p)) , 1] = p · [s(π c (p)) , 1] g = Fs (p)g,
664 18. SEIBERG-WITTEN THEORY

for any g ∈ Spinc (n). ∞


 Thus, via s 7→ Fs , we may regard C (X, U(1)) as a subgroup,
say GA1 PSpinc(n) , of GA PSpinc(n) . Note that for πrc : PSpinc(n) → PU(1) × F X,
x = π c (p) with p ∈ PSpinc(n) , we have
 
2
πrc (Fs (p)) = πrc (p · [s(x) , 1])) = πrc (p) · rc ([s(x) , 1]) = πrc (p) · s(x) , Id .
 2
Thus, Fs ∈ GA1 PSpinc(n) induces the gauge transformation p1 7→ p1 s(x) on
PU(1) . Hence, using Proposition 15.29 (p. 411) with fs ∈ C PU(1) , U(1) , given by
2 2
fs (p1 ) = p1 s(π1 (p1 )) = p1 ((π1∗ s)(p1 )) = p1 π1∗ s2 (p1 )


(where π1 : PU(1) → X), we have that C ∞ (X, U(1)) acts on the set C PU(1) of


connections ω on PU(1) via


−1
+ fs ωfs−1 = fs fs∗
−1
+ ω = π1∗ s2 d π1∗ s−2 + ω
 
s · ω = fs fs∗
= π1∗ s2 ds−2 + ω = π1∗ −s2 2s−3 ds + ω = ω + π1∗ −2s−1 ds .
  

Moreover, Proposition 15.29 implies that C ∞ (X, U(1)) ultimately acts on the space
C ∞ (X, Σ+
c (X)) via the simple rule

(s · ψ)(x) = s(x) ψ(x) .


We show that if (ω, ψ) solves the S-W equations, then
s · (ω, ψ) := (s · ω, s · ψ) = (ω + π1∗ −2s−1 ds , sψ)

(18.35)
(for all s ∈ C ∞ (X, U(1))) is also a solution. By Corollary 15.30
 (p. 412) and the
fact that the group U(1) of PU(1) is abelian, Ωs·ω = π1∗ s2 Ωω π1∗ s−2 = Ωω .
Alternatively,
Ωs·ω = d(s · ω) = d ω + π1∗ −2s−1 ds = dω − 2π1∗ d s−1 ds
 

= dω − 2π1∗ −s−2 ds ∧ ds + s−1 d2 s = Ωω ,




since ds∧ds = 0 and d2 s = 0. Also, since the inner product on Σ+


c (X) is Hermitian,
we have
2
q(s · ψ) = q(sψ) = ss̄q(ψ) = |s| q(ψ) = q(ψ) .
Hence, as scalar multiplication does not affect duality, we have
(18.36) Ωω+ = q(ψ) + η ⇒ Ω(s·ω)+ = Ωω+ = q(ψ) + η = q(s · ψ) + η.
Using Proposition 15.27 (p. 410) and the linearity of Clifford multiplication, we also
have
(18.37) Dcs·ω (s · ψ) = s · Dcω ψ, so that Dcω ψ = 0 ⇒ Dcs·ω (s · ψ) = 0.
We denote the orbit of (ω, ψ) ∈ C PU (1) × C ∞ (Σ+ ∞

c (X)) under C (X, U(1))
by
[ω, ψ] := {s · (ω, ψ) : s ∈ C ∞ (X, U(1))} .
Definition 18.45. For a fixed self-dual η ∈ Ω2+ (X, iR), the set
Mη := [ω, ψ] : Dcω ψ + = 0 and Ωω+ = q ψ + + η
 

of orbits of solutions (ω, ψ) of the perturbed S-W equations is known as the moduli
space for the Spinc structure (PSpinc(n) , PU(1) ).
18.3. GENERIC REGULARITY OF THE MODULI SPACES 665

The Formal Dimension of the Moduli Space. Mη is not always a manifold


in a natural way. However, we will show that there are generic perturbations η such
that Mη (if nonempty) is naturally a finite-dimensional manifold at points [(ω, ψ)]
where ψ 6= 0.
We compute the dimension of Mη formally as follows. Any two connections
on PU(1) differ by a one-form which is the pull-back of a one-form on the base X.
Thus, C PU(1) is an affine space with associated vector space

1  π∼1
Ω PU (1) , iR −→ Ω1 (X, iR) ,

which then formally serves as the tangent space  Tω∞ C PU (1) . The formal tangent
space of the orbit of a point (ω, ψ) ∈ C PU (1) × C (Σ+ c (X)) under the action of
C ∞ (X, U(1)) is the subspace of
T(ω,ψ) C PU (1) × C ∞ Σ+ := Ω1 PU (1) , iR ⊕ C ∞ Σ+
   
c (X) c (X)

consisting of elements of the form (where θ ∈ C ∞ (X, R))


d itθ d
ω + π1∗ −2e−itθ d eitθ , eitθ ψ
  
e · (ω, ψ) =
dt t=0 dt t=0
1 ∞ +

= (−2idθ, iθψ) ∈ Ω (X, iR) × C Σc (X) ,
1 
where we have (and will continue) to identify Ω PU (1) , iR and Ω1 (X, iR). Now,
Ω1 PU (1) , iR ⊕ C ∞ (Σ+

c (X)) has an inner product, given by
Z
h(iα, ξ) ,(iα0 , ξ 0 )i = h(α, α0 ) + <(hξ, ξ 0 i) νh ,
X
0
where h(α, α ) denotes the usual inner product on covectors and νh denotes the
volume form, each induced by the Riemannian metric h on X. A vector (iα, ξ) is
normal to the orbit if we have (where δ = − ∗ d∗ denotes the formal L2 -adjoint of
d)
0 = h(iα, ξ) ,(−2idθ, iθψ)i = hα, −2dθi + 21 (hξ, iθψi + hiθψ, ξi)
= −2(δα) − 2i (hξ, ψi − hψ, ξi) , θ for all θ ∈ C ∞ (X, R), or

(18.38) iδα + 41 (hψ, ξi − hξ, ψi) = 0.


Consider the so-called Seiberg-Witten function
Φ× : C PU(1) × C ∞ Σ+ −→ Ω2+ (X, iR) ⊕ C ∞ Σ+
  
(18.39) c (X) c (X) ,

given by
(18.40) Φ× (ω, ψ) = (Ωω+ − q(ψ) − η, Dcω ψ).
−1
The set of solutions of the S-W equations is Φ× (0, 0). We formally compute
the differential
Φ×
∗(ω,ψ) (iα, ξ) =
d
dt (Ω
(ω+tiα)+
− q(ψ + tξ) , Dcω+tiα (ψ + tξ))
t=0
+
= (i(dα) − qe(ξ, ψ) − qe(ψ, ξ) , Dcω ξ + 21 iα · ψ),
where α · ψ is shorthand for c(α# ⊗ ψ) (i.e., Clifford multiplication of ψ by the
vector field α# dual to the one-form α).
666 18. SEIBERG-WITTEN THEORY

Since the tangent space of the moduli space M at [ω, ψ] can be formally iden-
tified with the intersection of Ker Φ×
∗(ω,ψ) with the normal space of the orbit of the
action of C ∞ (X, U(1)) through (ω, ψ), we see that formally the tangent space of
M at [(ω, ψ)] is the kernel of the real operator

(18.41) B : Ω1 (X, iR)⊕C ∞ Σ+ ∞ 2+


(X, iR)⊕C ∞ Σ−
 
c (X) −→ C (X, R)⊕Ω c (X) ,

given by
 
B(iα, ξ) := iδα + 41 (hψ, ξi − hξ, ψi) , Φ×
∗(ω,ψ) (iα, ξ) .

We will now compute the index B, which provides a lower bound on dim(Ker B).
In general, dim(Coker B) 6= 0, but later we will show that for a generic choice of
−1
self-dual 2-form η, Coker(B) = 0 at any point in Φ× η (0, 0) for which ψ 6= 0.
Hence the index of B will turn out to be the dimension of a (perturbed) moduli
space.
One consequence of the Atiyah-Singer Index Theorem is that the index of an
elliptic operator is determined by its highest order part. The top order part of
differential operator B is the first-order differential operator
 
+
δ(iα) , d(iα) , Dcω ξ = δ ⊕ d+ ⊕ Dcω (iα, ξ).

(18.42) B1 (iα, ξ) :=

Each of the operators

δ ⊕ d+ : Ω1 (X, iR) → Ω0 (X, iR) ⊕ Ω2+ (X, iR) and


Dcω : C ∞ Σ+ ∞
Σ−
 
c (X) → C c (X)

is elliptic. Indeed, δ ⊕ d+ is elliptic, since we have previously shown that δ ⊕ d−


is elliptic (see the proof of Theorem 16.17, p. 491, where T0 := δ ⊕ d− ) and δ ⊕ d+
results from δ⊕d− by changing the orientation of X, which does not affect ellipticity.

The formal adjoint of (δ ⊕ d+ ) is d ⊕ δ|Ω2+(X,iR) , and
 ∗ 
index δ ⊕ d+ = dim Ker δ ⊕ d+ − dim Ker δ ⊕ d+ = b1 − 1 + b+
  
2
− −
= 12 (b1 + b3 ) − 1 + 21 b+
 1 + 
2 − b2 + 2 b2 + b2
= − 12 (1 − b1 + b2 − b3 + 1) − 1
2 sig(X)
= − 12 (χ(X) + sig(X)) .

Now Dcω is locally a Dirac operator twisted by L0 . Thus, we may apply the Local
Index Formula for twisted Dirac operators (Theorem 17.49, p. 577). Since locally
L0 ⊗ L0 = L, on the level of forms we have (where ω 0 := 21 π1∗ (ω))

c1 (L0 , ω 0 ) = 2c1 (L, ω) , and


2
ch(L0 , ω 0 ) = 1 + c1 (L0 , ω 0 ) + 21 c1 (L0 , ω 0 )
2
= 1 + 12 c1 (L, ω) + 81 c1 (L, ω) .
18.3. GENERIC REGULARITY OF THE MODULI SPACES 667

Then the Local Index Formula yields


 
2
index Dcω = 1 + 12 c1 (L) + 18 c1 (L) ` A(X)
b [X]
 
2
= 1 + 12 c1 (L) + 18 c1 (L) ` 1 − 241
p1 (T X) [X]
 
2
= 18 c1 (L) − 241
p1 (T X) [X]
 
2 2
= 18 c1 (L) [X] − 24
1
3 sig(X) = 18 c1 (L) [X] − sig(X) ,

where p1 (T X) [X] = 3 sig(X) by the Hirzebruch Signature Theorem (Theorem


17.64, p. 605). Since index B will turn out to be the dimension of a (perturbed)
moduli space, we denote index B by d(X, L). In summary, the index of the real
operator B is then given by

d(X, L) : = index B = index B1 = index δ ⊕ d+ + indexR Dcω




= b1 − 1 + b+ ω

2 + indexR Dc
 
2
= − 12 (χ(X) + sig(X)) + 2 · 81 c1 (L) [X] − sig(X)
= 41 c1 (L)2 − 2χ(X) − 3 sig(X) .

(18.43)

Remark 18.46. Note that as indexR (Dcω ) = 2· indexC (Dcω ) is even,

indexR (B) ≡ index δ ⊕ d+ mod 2.




Since index(δ ⊕ d+ ) = b1 − b2+ − 1, indexR (B) is even if and only if b1 + b2+ is odd
(i.e., indexR (B) and b1 + b2+ have opposite parity).

The Moduli Space as a Manifold. Having computed the formal dimension


d(X, L) of moduli space, one would like to verify that under suitable conditions it
is a manifold. This requires the introduction of function spaces and hard analysis.
−1
The set of solutions of the S-W equations is Φ× (0, 0). In this chapter, we have
assumed thus far that all objects are in the C ∞ category. Spaces of C ∞ sections can
be made into Fréchet spaces, but without extra conditions, the implicit function
theorem fails for maps between Fréchet spaces. Thus, even if we could prove that
× −1
Φ×

∗(ω,ψ) is onto, one could not deduce that Φ (0, 0) is a manifold. Instead,
we enlarge the spaces involved to suitable Sobolev spaces which are Hilbert (or
Banach) spaces. In that category, one has inverse and implicit function theorems,
as we have seen in Sections 77.3 and/or 16.4. Moreover, we will need

Theorem 18.47 (Unique Continuation, see [14]). Let E → X be a vector bun-


dle over a connected manifold X and let D : C ∞ (E) ←- be a second-order, elliptic
differential operator whose symbol is a multiple of the identity at each nonzero cov-
ector (i.e., a scalar symbol). If Du = 0 and u = 0 on some open set, then u = 0 on
X.

For an elementary proof of the weak unique continuation property for perturbed
Dirac type operators see [83, Theorem 8.2], elaborated in [70, 81] and [60, Theorem
1.33].
668 18. SEIBERG-WITTEN THEORY

The appropriate Sobolev extension of the S-W function (18.39) is (for suffi-
ciently large k)
2,k+1 2,k+1
Φ× PU (1) × W 2,k+1 Σ+
 
:C c (X)

−→ W 2,k Λ2+ (X, iR) ⊕ W 2,k Σ−


 
c (X) .
4
Note that we need k + 1 > 2 (i.e., k ≥ 2) in order that q defines a bounded bilinear
form on
W 2,k+1 Σ+ 2,k+1
Σ+ 2,k+1
Λ2+ (X, iR)
  
c (X) × W c (X) → W

according to Proposition
 16.24, p. 498. We remark that while a connection 1-
form ω ∈ C PU(1) is not a 1-form defined on X, any two connection 1-forms ω1
and ω0 on PU(1) differ by a 1-form ω1 − ω0 that uniquely projects to a 1-form
 by ω1 − ω0 . Thus,
on X which we also denote  given a fixed arbitrary choice of a
connection ω0 ∈ C PU(1) , the space C PU(1) can be identified with Ω1 (X, iR), via
ω1 ↔ ω1 − ω0 . Then
C 2,k+1 PU(1) := ω0 + W 2,k+1 Λ1 (X, iR) ,
 

which is really independent of the choice of ω0 .


Assume that k ≥ 2. Let
∗
W 2,k+1 Σ+ : = W 2,k+1 Σ+

c (X) c (X) − {0} and
2,k+1 2,k+1
∗
Σ+ × W 2,k Λ2+ (X, iR) .
 
(18.44) CWP k : = C PU(1) × W c (X)

The tangent space of the Banach manifold CWP k at (ω, ψ, η) is


∗
T(ω,ψ,η) CWP k := W 2,k+1 Λ1 (X, iR) × W 2,k+1 Σ+ × W 2,k Λ2+ (X, iR) .
 
c (X)

Define
F : CWP k −→ W 2,k Λ2+ (X, iR) ⊕ W 2,k Σ−
 
c (X) by
2,k+1
(18.45) F (ω, ψ, η) := Φ× (ω, ψ) = (Ωω+ − q(ψ) − η, Dcω ψ).
Theorem 18.48. If for some triple (ω, ψ, η) ∈ CWP k we have F (ω, ψ, η) = 0,
then the differential
F∗ : T(ω,ψ,η) CWP k −→ W 2,k Λ2+ (X, iR) ⊕ W 2,k Σ−
 
c (X)

at (ω, ψ, η) is given by
 
+
F∗ (ω 0 , ψ 0 , η 0 ) = (dω 0 ) − qe(ψ 0 , ψ) − qe(ψ, ψ 0 ) − η 0 , Dcω ψ 0 + 12 ω 0 · ψ ,
and F∗ is onto.
Proof. Note that for η 0 ∈ W 2,k Λ2+ (X, iR) , F∗ (0, 0, −η 0 ) = (η 0 , 0), so that


Im(F∗ ) ⊇ W 2,k Λ2+ (X, iR) ⊕ 0.




It suffices to show that for any ξ 0 ∈ W 2,k (Σ− 0 2,k+1



c (X)), there is some ω ∈ W Λ1 (X, iR)
and ψ 0 ∈ W 2,k+1 (Σ+ c (X)), such that

ξ 0 = Dcω ψ 0 + 21 ω 0 · ψ,
since then
 
+
F∗ ω 0 , ψ 0 ,(dω 0 ) − qe(ψ 0 , ψ) − qe(ψ, ψ 0 ) − η 0 = (η 0 , ξ 0 ).
18.3. GENERIC REGULARITY OF THE MODULI SPACES 669

Thus, it remains to show that the map


D : W 2,k+1 Λ1 (X, iR) × W 2,k+1 Σ+ −→ W 2,k Σ−
  
c (X) c (X)
given by
(18.46) D(ω 0 , ψ 0 ) = Dcω ψ 0 + 21 ω 0 · ψ
is onto. Since Dcω is elliptic, the operator D is also elliptic. Thus, according to
Proposition 16.25 (applied to the elliptic formal adjoint D∗ of D), D is onto if
Ker(D∗ ) is 0. We determine D∗ as follows. For ξ ∈ C ∞ (Σ− c (X)) , we have (where
(·, ·) denotes the L2 inner product on C ∞ (Σ−
c (X)))
Z
ω 0 1 0 0 ω
ω 0 , i Im 12 c(·)ψ, ξ g νg ,
 
Dc ψ + 2 ω · ψ, ξ = (ψ , Dc ξ) +
X
using the fact that ω 0 is iR-valued. Thus,
D∗ ξ = i Im 1
⊕ Dcω ξ

2 c(·)ψ, ξ
= 4i (hc(·)ψ, ξi − hξ, c(·)ψi) ⊕ Dcω ξ.
Hence
D∗ ξ = 0 ⇔ Dcω ξ = 0 and Im hc(·)ψ, ξi = 0.
We have assumed that ψ 6= 0 and Dc ψ = 0. If Dcω ξ = 0 and ξ 6= 0, then by
ω

the Theorem 18.47 (p. 667), neither ψ nor ξ can vanish on an open set. (Note
that we are actually applying Theorem 18.47 with D = Dcω ∗ Dcω which is a second-
order operator with scalar symbol, and Dcω ξ = 0 ⇔ Dcω∗ Dcω ξ = 0.) Hence there
must be a point x ∈ X where both ψ and ξ are nonzero. At x the covector
Im 12 c(·)ψ, ξ is nonzero. Indeed, first recall that Clifford multiplication takes

R4 onto R· SU(Σ+ 2
2 , Σ2 ). Since SU(2) acts transitively on the unit sphere in C , we
know there is some vector V ∈ Tx (X), so that c(V )ψ(x) = iξ(x) and
1 1 1
Im 2 c(V )ψ(x), ξ(x) = Im 2 iξ(x), ξ(x) = 2 hξ(x), ξ(x)i > 0.
Thus, necessarily ξ = 0, and D in (18.46) is onto as required. 
Theorem 18.48 and Theorem 16.26 (Implicit Function Theorem I, p. 499) yield
Corollary 18.49. The parametrized solution space
SWP k := F −1 (0, 0)
= (ω, ψ, η) ∈ CWP k : Ωω+ − q(ψ) − η = 0, Dcω ψ = 0

(18.47)
(if nonvoid) is a Hilbert submanifold of CWP k with tangent space at (ω, ψ, η) being
the kernel of F∗ .
We need to indicate the sense in which the group C ∞ (X, U (1)) of gauge trans-
formations can be enlarged to a C ∞ Banach manifold W 2,k+2 (X, U (1)) (where
k ≥ 1) and group for which the group operation (r, s) 7→ rs−1 is C ∞ . We have
(where the last inclusion is compact by Proposition 16.22, p.498, since 2(k + 2) > 4)
W 2,k+2 (X, U(1)) := s ∈ W 2,k+2 (X, C) : ss̄ = 1 (a.e.)


ι
⊆ W 2,k+2 (X, C) ⊆ C0 (X, C),
and so W 2,k+2 (X, U(1)) inherits a topological metric from W 2,k+2 (X, C). Moreover,
since C0 (X, U(1)) is closed in C0 (X, C),
W 2,k+2 (X, U(1)) = ι−1 (C0 (X, U(1))) is closed in W 2,k+2 (X, C).
670 18. SEIBERG-WITTEN THEORY

Hence, W 2,k+2 (X, U(1)) is a complete metric space. For any s ∈ W 2,k+2 (X, U(1)),
we define a function
Exps : W 2,k+2 (X, iR) −→ W 2,k+2 (X, U(1)) by Exps (iθ)(x) = eiθ(x) s(x).
For iθ ∈ W 2,k+2 (X, iR), the fact that Exps (iθ) ∈ W 2,k+2 (X, U(1)) follows from
Proposition 16.24 (p. 498) and
Proposition 18.50. If f : E → F is any C ∞ fiber-preserving map (not neces-
sarily linear on fibers), then left-composition by f defines a C ∞ map from W p,k (E)
to W p,k (F ) if k − np > 0.
Proof. See [329]. 
Proposition 18.51. For k ≥ 1, W 2,k+2 (X, U(1)) is a C ∞ Banach manifold
and W 2,k+2 (X, U(1))2 3 (t, s) 7→ ts−1 = ts̄ ∈ W 2,k+2 (X, U(1)) is C ∞ .
Proof. Let d(z1 , z2 ) denote the angular distance between z1 and z2 in the unit
circle U(1). For ρ ∈ [0, π], define
Uρ (s) := t ∈ W 2,k+2 (X, U(1)) : d(t(x), s(x)) < ρ for all x ∈ X , and


Oρ := iθ ∈ W 2,k+2 (X, iR) : |iθ(x)| < ρ for all x ∈ X .




Since the inclusion W 2,k+2 (X, iR) ⊆ C 0 (X, iR) is continuous, Oρ is open. More-
over,
Exps |Oρ : Oρ −→ Uρ (s)
is bijective. We denote the inverse by Logs : Uρ (s) → Oρ , and it is given by
 
Logs (t)(x) = i arg t(x)s(x) ,

where arg(eiθ ) = θ for |θ| < π. For s ∈ W 2,k+2 (X, U(1)), we consider the set of
charts Logs . To define a differentiable structure on W 2,k+2 (X, U(1)), we need to
show that
Logs ◦ Logt−1 : Logt (Uρ (s) ∩ Uρ (t)) −→ Logs (Uρ (s) ∩ Uρ (t))
is C ∞ . Note that if ρ < π2 , then
 
Logs ◦ Logt−1 (iθ) (x) = Logs Logt−1 (iθ) (x) = i arg Logt−1 (iθ) (x)s(x)
 
   
= i arg Expt (iθ) (x)s(x) = i arg eiθ(x) t(x)s(x)
 
= i θ(x) + i arg t(x)s(x) = i θ(x) + Logs (t)(x).

Thus, Logs ◦ Logt−1 is a translation of W 2,k+2 (X, iR) by Logs (t) ∈ W 2,k+2 (X, iR),
which is C ∞ . Hence, W 2,k+2 (X, U(1)) is a Banach manifold. The mapping from
W 2,k+2 (X, U (1))2 to W 2,k+2 (X, U(1)) given by (t, s) 7→ ts−1 = ts̄ is C ∞ , since in
terms of charts about t0 , s0 and t0 s̄0 it is given by
 
i arg t t0 , i arg(s s0 ) = (Logt0 (t), Logs0 (s))
  
7→ Logt0 s̄0 (ts̄) = i arg t s̄ t0 s̄0 = i arg t t0 s̄ s0 = i arg t t0 − i arg(s s0 )
(i.e., (u, v) 7→ u−v). One could also show that W 2,k+2 (X, U(1)) is a closed (Hilbert)
submanifold and subgroup of W 2,k+2 (X, C∗ ). Since W 2,k+2 (X, C∗ ) is covered by
a single chart in which the group operations are certainly C ∞ (using Propositions
18.3. GENERIC REGULARITY OF THE MODULI SPACES 671

16.24 and 18.50), the restrictions of the group operations to W 2,k+2 (X, U(1)) are
smooth. Thus, W 2,k+2 (X, U(1)) is a closed Lie subgroup of W 2,k+2 (X, C∗ ). 
For k ≥ 2, W 2,k+2 (X, U(1)) acts smoothly and freely on CWP k via
s ·(ω, ψ, η) = ω + π1∗ −2s−1 ds , sψ, η .
 
(18.48)
Note that s−1 ds ∈ W 2,k+1 (X, Λ1 (X, iR)) by Proposition 16.24 (p. 498), since s−1 ∈
W 2,k+2 (X, U(1)) and ds ∈ W 2,k+1 (X, Λ1 (X, C))) and 2(k + 1) > 4; also s−1 ds is
iR-valued, since
s−1 ds = s−1 ds = sd s−1 = s(−s−2 ds) = −s−1 ds).


For k ≥ 3, s−1 ds is C 1 since k + 1 − 24 ≥ 2 > 1 and Proposition 16.22 (p. 498)


then applies. Thus, for k ≥ 3 we can compute d s−1 ds = −s−2 ds ∧ ds = 0 in the
usual way. The computations leading to (18.36) and (18.37) then apply to show
that SWP k ⊆ CWP k is invariant under the action of W 2,k+2 (X, U(1)) for k ≥ 3.
Quotient Manifolds. We wish to construct quotient manifolds
CWP k /W 2,k+2 (X, U(1)) and SWP k /W 2,k+2 (X, U(1)).
In doing this, we will need to repeatedly use a family of elliptic operators parame-
trized by (ω, ψ, η) ∈ CWP k for k ≥ 2. Using (18.41) for the motivation (and noting
that η is just carried along for the ride for now), this family is
B(ω,ψ,η),k : W 2,k+1 Λ1 (X, iR) ⊕ W 2,k+1 Σ+
 
c (X)

−→ W 2,k (X, iR) ⊕ W 2,k Λ2+ (X, iR) ⊕ W 2,k Σ−


 
c (X) ,
given by
δω 0 − 41 (hψ 0 , ψi − hψ, ψ 0 i)
 

(18.49) B(ω,ψ,η),k (ω 0 , ψ 0 ) :=  dω 0+ − qe(ψ 0 , ψ) − qe(ψ, ψ 0 )  .


Dω ψ 0 + 21 ω 0 · ψ
Proposition 16.24 (p. 498) is used in showing that this is well-defined. For the time
being, both (ω, ψ, η) ∈ CWP k and k ≥ 2 will be fixed, and we write B(ω,ψ,η),k
simply as B. It is convenient to introduce the following notation.
Notation 18.52. Let

U1 := W 2,k+1 Λ1 (X, iR) , U2 := W 2,k+1 (Σ+
c (X)) ,

V2 := W 2,k Λ2+ (X, iR) , V3 := W 2,k (Σ−



V1 := W 2,k (X, iR) , c (X)) .
Then we write
B 1 , B 2 , B 3 : U1 ⊕ U2 −→ V1 ⊕ V2 ⊕ V3 , where

B =
B 1 : U1 ⊕ U2 −→ V1 with B 1 (ω 0 , ψ 0 ) = δω 0 − 41 (hψ 0 , ψi − hψ, ψ 0 i) ,
B 2 : U1 ⊕ U2 −→ V2 with B 2 (ω 0 , ψ 0 ) = dω 0+ − qe(ψ 0 , ψ) − qe(ψ, ψ 0 ) , and
B 3 : U1 ⊕ U2 −→ V3 with B 3 (ω 0 , ψ 0 ) = Dcω ψ 0 + 12 ω 0 · ψ.
Since B is elliptic, the symbols of B 1 , B 2 , and B 3 are surjective but not in-
jective. We need to compute the appropriate Sobolev extensions of the L2 formal
adjoints of B 1 , B 2 , and B 3 .
∗
B 1  : W 2,k+2 (X, iR) −→ U1 ⊕ U2 ,

B 2  : W 2,k+2 Λ2+ (X, iR) −→ U1 ⊕ U2 , and

B 3 : W 2,k+2 (Σ− c (X)) −→ U1 ⊕ U2 .
672 18. SEIBERG-WITTEN THEORY

The final results are:


∗
(iθ) =  idk+2 θ, − 2i θψ ,

B1
∗
B 2 (γ) = δ k+2 γ, Sψ∗ (γ) , and
∗  
k+2
B 3 (ξ) = Tψ∗ (ξ) ,(Dcω ) ξ ,

where Sψ∗ and Tψ∗ are 0-th order operators, described below. Since iR is a real
vector space with real inner product ha, bi := ab, the spaces Ωj (X, iR) are also
real vector spaces with real inner products, and so the formal adjoints need to
be computed relative to real L2 inner products. Hence, although Σ± c (X) has
a Hermitian inner product, say hψ1 , ψ2 i , we must use the ∗ real inner product
<(hψ1 , ψ2 i) = 12 (hψ1 , ψ2 i + hψ2 , ψ1 i) . We first compute B 1 :
Z Z
iθ, B 1 (ω 0 , ψ 0 ) νg = iθ, δω 0 + 41 (hψ, ψ 0 i − hψ 0 , ψi) νg
X
ZX Z
= hiθ, δω 0 i νg − 1 0 0
4 (hiθψ, ψ i + hψ , iθψi) νg
X X
Z Z
0 1 0
= hidθ, ω i νg − 2 <(hiθψ, ψ i) νg
X X
Z
idθ, − 2i θψ ,(ω 0 , ψ 0 ) νg .

=
X

2 ∗

For B , we first note that there is a 0-th order differential operator

Sψ : C ∞ Σ+ 2+
given by Sψ (ψ 0 ) = qe(ψ 0 , ψ) + qe(ψ, ψ 0 ).

c (X) → Ω (X, iR)

The adjoint of Sψ , say Sψ∗ : Ω2+ (X, iR) → C ∞ (Σ+


c (X)) , has the property that for
2+
γ ∈ Ω (X, iR),

hγ, qe(ψ 0 , ψ) + qe(ψ, ψ 0 )i = hγ, Sψ (ψ 0 )i = < Sψ∗ (γ) , ψ 0 .

While one can find an explicit expression for Sψ∗ (γ), the important point is that

Sψ∗ : Ω2+ (X, iR) →C ∞ (Σ+


c (X))

is again a 0-th order differential operator. The formal adjoint of B 2 is computed


via
Z Z
0 0
2
γ, B (ω , ψ ) νg = hγ, dω 0 − qe(ψ 0 , ψ) − qe(ψ, ψ 0 )i νg
X X
Z Z
= hδγ, ω 0 i νg − hγ, qe(ψ 0 , ψ) + qe(ψ, ψ 0 )i νg
ZX ZX
0
= hδγ, ω i νg − < Sψ∗ (γ) , ψ 0 νg
X X
Z
δγ, −Sψ∗ (γ) ,(ω 0 , ψ 0 ) νg .

(18.50) =
X
18.3. GENERIC REGULARITY OF THE MODULI SPACES 673

∗
For B 3 , let ξ ∈ C ∞ (Σ−
c (X)), and note that
Z Z
0 0
3
< ξ, B (ω , ψ ) νg = < ξ, Dcω ψ 0 + 12 ω 0 · ψ νg
X X
Z
= < ξ, 12 ω 0 · ψ + < hDcω ξ, ψ 0 i νg
ZX
= Tψ∗ (ξ) , ω 0 + < hDcω ξ, ψ 0 i νg
X
Z
Tψ∗ (ξ) , Dcω ξ ,(ω 0 , ψ 0 ) νg .

(18.51) =
X
Here Tψ∗ denotes the 0-th order operator which is the adjoint of the 0-th order
operator
Tψ : Ω1 (X, iR) → C ∞ (Σ−
c (X)), given by Tψ (ω 0 ) := 21 ω 0 · ψ.
Theorem 18.53. For k ≥ 2, the quotient space of moduli
MCWP k := CWP k /W 2,k+2 (X, U(1))
has the structure of a Hausdorff C ∞ Hilbert manifold.
Proof. Step 1. We first will produce a local slice of the action of W 2,k+2 (X, U(1))
on CWP k . Previously (see (18.38)) we found that at a point (ω, ψ, η) ∈ CWP k , a
vector (ω 0 , ψ 0 , η 0 ) ∈ T(ω,ψ,η) (CWP k ) is formally L2 -orthogonal to the orbit
W 2,k+2 (X, U(1)) · (ω, ψ, η)
iff
δω 0 − 14 (hψ 0 , ψi − hψ, ψ 0 i) = 0,
0 0 1

i.e., iff (ω , ψ ) ∈ Ker B . Note that

− 2B 1∗ ⊕ 0 : W 2,k+2 (X, iR) −→ U1 ⊕ U2 ⊕ V2 ,


given by −2Bk1∗ ⊕ 0 (iθ) := (−2idθ, iθψ, 0) ,


is the differential at 1 ∈ W 2,k+2 (X, U(1)) of the map


W 2,k+2 (X, U(1)) −→ CWP k , given by
s 7→ s ·(ω, ψ, η) := ω − 2π1∗ s−1 ds , sψ, η .
 

Since the symbol of B 1∗ is injective, we have the decomposition (see Proposition


16.25, p. 499)
U1 ⊕ U2 = Ker B 1 ⊕ Im B 1∗
 
(18.52)
into closed subspaces which are L2 -orthogonal. For a fixed (ω, ψ, η), we define a
C ∞ map
J = J(ω,ψ,η) : Ker(B 1 ) × V2 × W 2,k+2 (X, U (1)) −→ CWP k by


J(x, s) : = s ·((ω, ψ, η) + x) ; i.e.,


J (ω , ψ 0 , η 0 ), s = s · (ω, ψ, η) + (ω 0 , ψ 0 , η 0 )
0
 

= bigl(ω + ω 0 − 2π1∗ s−1 ds , s(ψ + ψ 0 ) , η + η 0 .


 

At (0, 0, 1) ∈ Ker(B 1 ) × V2 × W 2,k+2 (X, U (1)), the differential



J∗(0,0,1) : Ker B 1 ⊕ V2 ⊕ V1 −→ T(ω,ψ,η) (CWP k ) −→ U1 ⊕ U2 ⊕ V2

674 18. SEIBERG-WITTEN THEORY

is given by
J∗(0,0,1) ((ω 0 , ψ 0 ) , η 0 , iθ) = (ω 0 − 2idθ, iθψ + ψ 0 , η 0 )
= (ω 0 , ψ 0 , η 0 ) + (−2idθ, iθψ, 0) = (ω 0 , ψ 0 , η 0 ) − 2B 1∗ (iθ) .

Thus, on Ker B 1 ⊕ V2 ⊕ 0, J∗(0,0,1) is the inclusion
Ker B 1 ⊕ V2 ⊕ 0 ⊆ U1 ⊕ U2 ⊕ V2 ,


and on 0 ⊕ 0 ⊕ V1 ,
J∗(0,0,1) = −2B 1∗ : V1 → U1 ⊕ U2 ⊕ 0.
Thus using (18.52), we have that  J∗(0,0,1) is onto. The kernel of J∗(0,0,1) is trivial.
1 1∗
Indeed, since Ker B ⊥ Im B , we have
(18.53) (ω 0 , ψ 0 , η 0 ) + B 1∗ (iθ) = 0 ⇒ (ω 0 , ψ 0 , η 0 ) = 0 and B 1∗ (iθ) = 0,
and iθ = 0, since Ker(B 1∗ ) is trivial (recall ψ 6= 0). Thus, by the Inverse Function

Theorem, Theorem 16.28 (p. 500), there are open neighborhoods U of 0 ∈ Ker B 1 ×
V2 and V of 1 ∈ W 2,k+2 (X, U(1)), such that
J|(U ×V ) : U × V → J(U × V )
is a diffeomorphism onto a neighborhood J(U × V ) of (ω, ψ, η) ∈ CWP k . Thus,
points of CWP k near (ω, ψ, η) are uniquely of the form
J(x, s) = s ·((ω, ψ, η) + x) for x ∈ Ker(B 1 ) × V2
(i.e., for x in the normal space to the orbit of (ω, ψ, η)). In other words, U is a local
slice of the action. On our way to a global slice, let
Q : CWP k −→ CWP k / W 2,k+2 (X, U(1))
denote the quotient function and define a diffeomorphism
j(ω,ψ,η) : U −→ J(ω,ψ,η) (U × {1}) by
j(ω,ψ,η) (x) := J(ω,ψ,η) (x, 1) = (ω, ψ, η) + x.
Step 2. We now show that the mapping
ϕ(ω,ψ,η) : U −→ CWP k / W 2,k+2 (X, U(1)) ,
given by
(18.54) ϕ(ω,ψ,η) (K) := [(ω, ψ, η) + K] = Q ◦ j(ω,ψ,η) (K)
is 1-1 if U is chosen small enough. Suppose that for K1 = (ω1 , ψ1 , η1 ) and K2 =
(ω2 , ψ2 , η2 ) ∈ U, we have ϕ(ω,ψ,η) (K1 ) = ϕ(ω,ψ,η) (K2 ). Then
s ·((ω, ψ, η) + K1 ) = (ω, ψ, η) + K2
for some s ∈ W 2,k+2 (X, U(1)) , i.e.,
s ·((ω, ψ, η) + (ω1 , ψ1 , η1 ))
= ω − 2π1∗ s−1 ds + ω1 , sψ + sψ1 , η + η1
 

(18.55) = (ω + ω2 , ψ + ψ2 , η + η2 ) .
Thus, ω − 2π1∗ s−1 ds + ω1 = ω + ω2 , sψ + sψ1 = ψ + ψ2 , and η1 = η2 . Hence,


2π1∗ s−1 ds = ω1 − ω2 and (s − 1) ψ = (ψ2 − ψ1 ) − (s − 1) ψ1 .



18.3. GENERIC REGULARITY OF THE MODULI SPACES 675

Now, let
n o
Uε := (ω 0 , ψ 0 , η 0 ) ∈ U : kψ 0 k2,k+1 ≤ ε and kω 0 k2,k+1 ≤ ε ⊆ U.

For K1 , K2 ∈ Uε , we have
(18.56) 2s−1 ds 2,k+1
= kω1 − ω2 k2,k+1 ≤ 2ε,
and for some constant C,
k(s − 1) ψk2,k+1 ≤ kψ2 − ψ1 k + k(s − 1) ψ1 k2,k+1
 
≤ 2ε + C ks − 1k2,k+1 kψ1 k2,k+1 ≤ 2 + C ks − 1k2,k+1 ε.

Proposition 16.22 and (18.56) yield


s−1 ds C0
≤ C 0 s−1 ds 2,k+1
≤ C 0 ε,
for some constant C 0 and k ≥ 2. Thus, for ε sufficiently small, we deduce that s
is arbitrarily C0 -close to some constant, say s0 , with ks − s0 kC 0 ≤ ε0 . We show
s0 = 1. If s0 6= 1, then for ε0 < 21 |1 − s0 |,
1
|s(x) − 1| > 2 |1 − s0 | > 0 for all x ∈ X.
00
For k ≥ 2, there is a also a constant C such that
 
(18.57) k(s − 1) ψkC 0 ≤ C 00 k(s − 1) ψk2,k+1 ≤ C 00 2 + C k(s − 1)k2,k+1 ε

can be made arbitrarily small with ε. However, since ψ 6= 0, there is some x0 with
ψ(x0 ) 6= 0 and
1
|(s(x0 ) − 1) ψ(x0 )| = |(s(x0 ) − 1)| |ψ(x0 )| ≥ 2 |1 − s0 | |ψ(x0 )| ,
which contradicts (18.57). Thus, s0 = 1, and (18.56) then implies that s ∈ V for
sufficiently small ε. Since J|Uε ×V is 1-1 and s ∈ V , (18.55) then yields
J(K1 , s) = J(K2 , 1) ⇒ (K1 , s) = (K2 , 1) ⇒ K1 = K2 .
Thus, by replacing U by Uε for sufficiently small ε, the mapping
ϕ(ω,ψ,η) : U −→ ϕ(ω,ψ,η) (U ) = Q ◦ j(ω,ψ,η) (U ) ⊆ CWP k / W 2,k+2 (X, U(1))
is 1-1.
Step 3. We wish to show that the
(18.58) ϕ−1
(ω,ψ,η) : ϕ(ω,ψ,η) (U ) −→ U

form a collection of coordinate charts for CWP k / W 2,k+2 (X, U(1)) with C ∞ coor-
dinate transitions. Suppose that we have two such charts, say
ϕ−1 −1
(ω1 ,ψ1 ,η1 ) : ϕ(ω,ψ,η) (U1 ) −→ U1 and ϕ(ω2 ,ψ2 ,η2 ) : ϕ(ω,ψ,η) (U2 ) −→ U2
with [(ω0 , ψ0 , η0 )] ∈ ϕ(ω1 ,ψ1 ,η1 ) (U1 ) ∩ ϕ(ω2 ,ψ2 ,η2 ) (U2 ) .
We may assume (ω0 , ψ0 , η0 ) ∈ j(ω1 ,ψ1 ,η1 ) (U1 ). There is s0 ∈ W 2,k+2 (X, U(1)), such
that (ω0 , ψ0 , η0 ) · s0 ∈ j(ω2 ,ψ2 ,η2 ) (U2 ) . We show that there is a neighborhood U10 of
(ω0 , ψ0 , η0 ) in j(ω1 ,ψ1 ,η1 ) (U1 ) and some s = eiθ ∈ C ∞ U10 , W 2,k+2 (X, U(1)) , such
that for all x ∈ U10
Rs0 s (x) := x · s0 · s(x) ∈ j(ω2 ,ψ2 ,η2 ) (U2 ) ,
676 18. SEIBERG-WITTEN THEORY

where Rs0 s : U10 → CWP k is C ∞ . The following computation will then yield
−1
ϕ−1
(ω2 ,ψ2 ,η2 ) ◦ ϕ(ω1 ,ψ1 ,η1 ) = j(ω2 ,ψ2 ,η2 ) ◦ Rs0 s ◦ j(ω1 ,ψ1 ,η1 ) on U10 ,
which is a composition of C ∞ mappings, as required: For x = j(ω1 ,ψ1 ,η1 ) (y) ∈ U10
−1
(i.e., y ∈ j(ω 1 ,ψ1 ,η1 )
(U10 ) ⊆ U1 ),
   −1 
ϕ−1
(ω2 ,ψ2 ,η2 ) ◦ ϕ (ω1 ,ψ ,η
1 1 ) (y) = Q ◦ j (ω2 ,ψ ,η
2 2 ) ◦ Q ◦ j (ω1 ,ψ ,η
1 1 ) (y)
−1 −1
= Q ◦ j(ω2 ,ψ2 ,η2 ) ([x]) = Q ◦ j(ω2 ,ψ2 ,η2 ) ([x · s0 · s(x)])
−1 −1
= Q ◦ j(ω2 ,ψ2 ,η2 ) ([Rs0 s (x)]) = j(ω2 ,ψ2 ,η2 ) (Rs0 s (x))
 −1 
= j(ω2 ,ψ2 ,η2 ) ◦ Rs0 s ◦ j(ω1 ,ψ1 ,η1 ) (y) .

We seek θ ∈ C ∞ U10 , W 2,k+2 (X, R) such that, for




s = eiθ ∈ C ∞ U10 , W 2,k+2 (X, U(1)) ,




we have
1
B(ω2 ,ψ2 ,η2 )
((ω, ψ, η) · s0 · s(ω, ψ, η) − (ω2 , ψ2 , η2 )) = 0
for all (ω, ψ, η) in a neighborhood of (ω0 , ψ0 , η0 ) in j(ω1 ,ψ1 ,η1 ) (U1 ) . Define
F : j(ω1 ,ψ1 ,η1 ) (U1 ) × W 2,k+2 (X, R) −→ W 2,k (X, R) by

1
(ω, ψ, η) · s0 · eiθ − (ω2 , ψ2 , η2 )

F ((ω, ψ, η) , θ) = B(ω2 ,ψ2 ,η2 )
1
ω − ω2 − 2s−1 iθ

(18.59) = B(ω2 ,ψ2 ,η2 ) 0 ds0 − 2idθ, e s0 ψ − ψ2 , η − η2 .

Using the fact that (ω0 , ψ0 , η0 ) · s0 ∈ J(ω2 ,ψ2 ,η2 ) (U2 ), we have F ((ω0 , ψ0 , η0 ) , 0) = 0.
The existence of θ will follow from the Implicit Function Theorem (Theorem 16.27,
p. 499), provided that at ((ω0 , ψ0 , η0 ) , 0)
∂θ F : W 2,k+2 (X, R) −→ W 2,k (X, R)
is onto. However, using
1
B(ω2 ,ψ2 ,η2 )
(ω 0 , ψ 0 , η 0 ) = δω 0 + 41 (hψ, ψ 0 i − hψ 0 , ψi) , we have

∂θ F (θ0 ) = B(ω
1
2idθ0 , −iθ0 s−1

2 ,ψ2 ,η2 ) 0 ψ0 , 0

= 2iδdθ0 + ψ2 , −iθ0 s−1


1 0 −1

4 0 ψ0 − −iθ s0 ψ0 , ψ2

= 2iδdθ0 + iθ0 ψ2 , s−1


1 0 −1

4 0 ψ0 + iθ s0 ψ0 , ψ2

= 2i δdθ0 + 81 θ0 s−1 −1

0 ψ0 , ψ2 + ψ2 , s0 ψ0 .
Since the elliptic operator
θ0 7→ δdθ0 + 81 θ0 s−1 −1

0 ψ0 , ψ2 + ψ2 , s0 ψ0

is self-adjoint, it suffices to prove that its kernel is trivial. If ∂θ F (θ0 ) = 0, then


Z
δdθ0 + 81 θ0 s−1 −1
 0
0= 0 ψ0 , ψ2 + ψ2 , s0 ψ0 , θ νg
ZX
2
|dθ0 | + 18 s−1 −1
 02
(18.60) = 0 ψ0 , ψ2 + ψ2 , s0 ψ0 θ νg .
X
18.3. GENERIC REGULARITY OF THE MODULI SPACES 677

In the case s−1 0


0 ψ0 = ψ2 , this last equation holds only if θ = 0. However, in general
we must ensure that the neighborhoods
  U associated with points (ω, ψ, η) are chosen
small enough so that for all ω e , ψ, ηe ∈ j(ω,ψ,η) (U ), the operator
e
D E D E
∆ψe : θ 7→ δdθ + 41 ψ, e ψ + ψ, ψe θ
 
is invertible. This is certainly true for ω e , ψ,
e ηe = (ω, ψ, η) . The invertible op-
erators from W 2,k+2 (X, R) to W2,k (X, R) form an open set in the Banach space
Hom W 2,k+2 (X, R) , W 2,k (X, R) of continuous linear maps. By Proposition 16.24
(p. 498), we have
  D E D E
∆ψe − ∆ψ θ = 41 ψe − ψ, ψ + ψ, ψe − ψ θ
2,k 2,k

≤ C kψk2,k ψe − ψ kθk2,k
2,k

≤ C kψk2,k ψe − ψ kθk2,k+2 .
2,k+1

Thus, ∆ψe − ∆ψ ≤ C kψk2,k ψe − ψ which shows that ∆ψe will be invertible


2,k+1
for ψe sufficiently W 2,k+1 -close (in fact W -close) to ψ. Thus, the neighborhoods U
2,k

associated with points (ω, ψ, η) can be chosen small enough so that the coordinate
transition functions will be smooth.
Step 4. It remains to show that the topology of CWP k / W 2,k+2 (X, U(1))
induced by the charts is Hausdorff. Since all spaces involved are second countable,
it suffices to show that if a sequence [xn ] in CWP k / W 2,k+2 (X, U(1)) converges to
both [y] and [z], then [y] = [z]. If [xn ] converges to both [y] and [z], then there are
sequences rn and tn ∈ W 2,k+2 (X, U(1)) such that rn · xn → y and tn · xn → z. Let
yn = rn · xn , then yn → y and for sn := rn−1 tn , we have sn · yn = tn · xn → z.
Thus, it suffices to prove that if yn → y and sn · yn → z, then there is a convergent
subsequence, say sni → s, in which case sni · yn → y · s, and so z = y · s (since
CWP k is a metric space) and [y] = [z]. In other words, we need to show that the
action of W 2,k+2 (X, U(1)) on CWP k is proper. To show this, let
(18.61) (ωn , ψn , ηn ) −→ (ω, ψ, η) ∈ CWP k ,
and suppose that
sn ·(ωn , ψn , ηn ) −→ (β, φ, ζ) for sn ∈ W 2,k+2 (X, U(1)) .
Then
ωn − 2s−1 2,k+1

n dsn = sn · ωn =: βn −→ β ∈ C PU(1)
⇒ −2s−1 2,k+1
Λ1 (X, iR) .

n dsn = βn − ωn −→ β − ω ∈ W
4
Since (k + 2) − 2 > 0 for k > 0,
2
sn ∈ W 2,k+2 (X, U(1)) ⇒ sn ∈ C 0 (X, U(1)) and |sn (x)| = 1.
Thus, sn is bounded in W p,0 (X, U(1)) for all p ≥ 1. Also,
2s−1 2,k+1
Λ1 (X, iR) ⊆ W 2,2 Λ1 (X, iR) .
 
n dsn = βn − ωn ∈ W
Thus, by Proposition 16.24 (p. 498) with
   
k3 − pn3 = 0 − 45 < − 46 = 0 − 64 + 2 − 42 = k1 − n n
 
p1 + k2 − p2 ,
678 18. SEIBERG-WITTEN THEORY

we have that
dsn = 12 sn (βn − ωn ) ∈ W 5,0 Λ1 (X, iR) , and


(18.62) kdsn k5,0 ≤ C ksn k6,0 kβn − ωn k2,2 .


 
Now βn − ωn is bounded in W 2,k+1 Λ1 (X, iR) and hence in W 2,q Λ1 (X, iR)
for q ≤ k + 1 (e.g., for q ≤ 3). Thus, as both sn and dsn are bounded in
W 5,0 (X, U(1)) , we have that  sn bounded in W (X,
5,1
U(1)) . Since βn − ωn is
bounded in W 2,3
X, Λ (iR) , it is also bounded in W 5,1 (X, U(1)) , since 1 − 54 ≤
1

2 − 43 . Hence, by (18.62) and the Proposition 16.24 (p. 498) with ki = 1, pi = 5,


ki − p4i = 15 > 0, i = 1, 2, 3,

kdsn k5,1 ≤ C ksn k5,1 kBn − ωn k5,1 .

Thus, kdsn k5,1 is bounded, and hence ksn k5,2 is bounded. Now, as k ≥ 2,

W 2,k+1 Λ1 (X, iR) ⊆ W 2,3 Λ1 (X, iR) ⊆ W 3,2 Λ1 (X, iR) ,


  

since 3 − 24 > 2 − 43 . Hence, as βn − ωn ∈ W 2,k+1 Λ1 (X, iR) is bounded in




W 3,2 Λ1 (X, iR) , we have


kdsn k3,2 ≤ C ksn k5,2 kβn − ωn k3,2 .

Thus, ksn k3,3 is bounded, and consequently ksn k2,3 is bounded. Inductively, sup-
pose that ksn k2,j+1 is bounded for some 2 ≤ j ≤ k. Since j + 1 − 42 = j − 1 > 0,
Proposition 16.24 (p. 498) yields
kdsn k2,j+1 ≤ C ksn k2,j+1 kBn − ωn k2,j+1 ,

so that ksn k2,j+2 is bounded. Hence, by induction, ksn k2,k+2 is bounded. By


Theorem 7.15 (p. 201), there is a subsequence, say s0n , of sn which converges in
W 2,k+1 (X, U(1)). Now, in W 2,k+1 (X, U(1)),
ds0n − ds0m = 21 s0n (βn0 − ωn0 ) − 12 s0m (βm
0 0
− ωm )
= 21 (s0n − s0m )(βn0 − ωn0 ) + 12 s0m (βn0 − ωn0 ) − 21 s0m (βm
0 0
− ωm )
= 12 (s0n − s0m )(βn0 − ωn0 ) + 12 s0m ((βn0 − ωn0 ) − (βm
0 0
− ωm )) .

Thus, by Proposition 16.24 (p. 498),


kds0n − ds0m k2,k+1 ≤ C1 k(s0n − s0m )k2,k+1 kβn0 − ωn0 k2,k+1
+ C2 ks0m k2,k+1 k(βn0 − ωn0 ) − (βm
0 0
− ωm )k2,k+1 .

Hence, ds0n converges in W 2,k+1 (X, C), and so s0n converges in W 2,k+2 (X, U(1)) ,
since s0n is known to converge in W 2,0 (X, U(1)). 

Manifold Structure for the Parametrized Moduli Space. We now turn


our attention to defining the manifold structure for the set
MSWP k := SWP k /W 2,k+2 (X, U (1))
called the parametrized moduli space of solutions of the perturbed SW equa-
tions, where the parameter is the perturbation η ranging over W 2,k (X, Λ2+ (X)).
18.3. GENERIC REGULARITY OF THE MODULI SPACES 679

Theorem 18.54. For k ≥ 3, the parametrized moduli space MSWP k (if non-
empty) is a closed Hilbert submanifold of MCWP k . At a point (ω, ψ, η) ∈ SWP k ,
the differential Q∗(ω,ψ,η) of the projection Q : CWP k → MCWP k restricts to an
isomorphism
Q∗(ω,ψ,η) : T(ω,ψ,η) (SWP k ) ∩ Ker B 1 −→ T[(ω,ψ,η)] (MSWP k ) .


In other words, the tangent space at [(ω, ψ, η)] ∈ MSWP k is the projection of the
orthogonal complement of the tangentspace of the orbit W 2,k+2 (X, U(1)) ·(ω, ψ, η)
in T(ω,ψ,η) (F −1 (0, 0)) = Ker F∗(ω,ψ,η) , namely
T[(ω,ψ,η)] (MSWP k )

−→ Ker F∗(ω,ψ,η) ∩ (ω 0 , ψ 0 , η 0 ) : δω 0 − 12 (hψ 0 , ψi − hψ, ψ 0 i) = 0
 

+
= (ω 0 , ψ 0 , η 0 ) : (dω 0 ) − qe(ψ 0 , ψ) − qe(ψ, ψ 0 ) − η 0 = 0 and


Dcω ψ 0 + 21 ω 0 · ψ = 0, δω 0 − 21 (hψ 0 , ψi − hψ, ψ 0 i) = 0


= (ω 0 , ψ 0 , η 0 ) : B 1 (ω 0 , ψ 0 ) = 0, B 2 (ω 0 , ψ 0 ) − η 0 = 0, B 3 (ω 0 , ψ 0 ) = 0 .

(18.63)
Proof. Recall that SWP k was shown to be a submanifold of CWP k which
is invariant under W 2,k+2 (X, U(1)) for k ≥ 3. We need to exhibit MSWP k as a
submanifold of MCWP k . This will be accomplished by showing, via the Implicit
Function Theorem, that MSWP k is a submanifold when viewed in a coordinate
neighborhood of MCWP k . Recall (see (18.45)) that
F : CWP k −→ W 2,k Λ2+ (X, iR) ⊕ W 2,k Σ−
 
c (X) is given by

F (ω, ψ, η) := (Ωω+ − q(ψ) − η, Dcω ψ).


A standard coordinate chart about [ω, ψ, η] ∈ MCWP k is
−1
ϕ(ω,ψ,η) : ϕ(ω,ψ,η) (U ) −→ U

(see (18.54) and (18.58)), where


ϕ(ω,ψ,η) := Q ◦ j(ω,ψ,η) : U −→ MSWP k ,

U ⊆ Ker(B 1 ) × V2 , and Q ◦ j(ω,ψ,η) (x) = [x + (ω, ψ, η)] . Since (ω, ψ, η) will
be fixed, we drop the subscript (ω, ψ, η) . It suffices to show that in terms of the
coordinate chart about (ω, ψ, η) ∈ SWP k , we have that ϕ−1 (MSWP k ∩ ϕ(U )) is
a submanifold of U near 0 ∈ U . We have
−1
ϕ−1 (MSWP k ∩ ϕ(U )) = (Q ◦ j) (MSWP k ∩ ϕ(U ))
−1
= j −1 ◦ Q|j(U ) (MSWP k ∩ ϕ(U )) = j −1 (SWP k ∩ j(U ))
−1
= j −1 F −1 (0, 0) ∩ j(U ) = (F ◦ j) (0, 0) .


By the Implicit Function Theorem I (Theorem 16.26, p. 499), we need to show that
(F ◦ j)∗0 is onto. Now

(F ◦ j)∗0 : Ker(B 1 ) ⊕ W 2,k Λ2+ (X, iR) −→ V2 ⊕ V3 ,



(18.64)
is given by
(F ◦ j)∗0 (ω 0 , ψ 0 , η 0 ) = F∗(ω,ψ,η) (ω 0 , ψ 0 , η 0 ) = B 2 (ω 0 , ψ 0 ) − η 0 , B 3 (ω 0 , ψ 0 ) .

680 18. SEIBERG-WITTEN THEORY

We assume that (ω, ψ, η) ∈ SWP k , in which case (F ◦ j)(0) = (0, 0) . Since Im(F ◦ j)∗0 ⊇
V2 ⊕0, it suffices to prove that B 3 Ker(B 1 ) = V3 . The proof of Theorem  18.48 con-
tains a proof that B 3 : U1 ⊕ U2 → V3 is onto. Since U 1 ⊕ U2 = Ker B 3
⊕ Im B 3∗
,
3 3∗
we know that B Im B = V3 . Thus, it suffices to show that
3∗
⊆ Ker B 1 or Ker B 3 ⊇ Im B 1∗ .
   
Im B
Thus, we need B 3 ◦ B 1∗ = 0. To this end
B 3 B 1∗ (iθ) = B 3 idθ, − 2i θψ = Dcω − 2i θψ + 2i dθ · ψ
  

= − 2i θDcω ψ − 2i dθ · ψ + 2i dθ · ψ = − 2i θDcω ψ = 0,
since we have assumed (ω, ψ, η) ∈ SWP k . Finally note that we have the identifica-
tion

T(ω,ψ,η) MSWP k −→ Ker((F ◦ j)∗0 )
= (ω 0 , ψ 0 , η 0 ) : B 1 (ω 0 , ψ 0 ) = 0, B 2 (ω 0 , ψ 0 ) − η 0 = 0, B 3 (ω 0 , ψ 0 ) = 0 ,

(18.65)
as required. 
We have a smooth projection mapping
p : MSWP k −→ W 2,k Λ2+ (X, iR) , where p([ω, ψ, η]) = η.

(18.66)
Let
(18.67)
SW k (η) := p−1 (η) = [ω, ψ, η] ∈ MSWP k : Ωω+ − q(ψ) − η = 0, Dcω ψ = 0 .


Theorem 18.55. Near a point (ω, ψ, η) ∈ SWP k where p[ω,ψ,η]∗ is onto (i.e.,
η is an achieved regular value of p), we have that SW k (η) := p−1 (η) is a smooth
manifold of dimension
d(X, L) := b1 − 1 + b2+ + 14 c1 (L)2 − sig(X) .
 

Proof. By the Implicit Function Theorem I (Theorem 16.26, p. 499), the set
SW k (η) is a smooth submanifold of MSWP k . Moreover, by (18.65) and (18.66),
+
(18.68) p∗[ω,ψ,η] (ω 0 , ψ 0 , η 0 ) = η 0 = B 2 (ω 0 , ψ 0 ) = (dω 0 ) − qe(ψ 0 , ψ) − qe(ψ, ψ 0 ) .
Thus,
T[ω,ψ,η] p−1 (η) = ker p∗[ω,ψ,η] : T[ω,ψ,η] MSWP k −→ W 2,k Λ2+ (X, iR)
 
n
∼ +
−→ (ω 0 , ψ 0 , 0) ∈ T(ω,ψ,η) CWP k : (dω 0 ) − qe(ψ 0 , ψ) − qe(ψ, ψ 0 ) = 0,
Dω ψ 0 + 21 ω 0 · ψ = 0, and δω 0 − 1
2 hψ, ψ 0 i = 0
which can be identified with the kernel of the elliptic operator
B : W 2,k+1 Λ1 (X, iR) ⊕ W 2,k+1 Σ+
 
c (X)

−→ W 2,k (X, iR) ⊕ W 2,k Λ2+ (X, iR) ⊕ W 2,k Σ−


 
c (X)

of (18.49). We need to show that if p[ω,ψ,η]∗ is onto, then B is onto, so that


index(B) = dim Ker(B). Assuming that p[ω,ψ,η]∗ in (18.68) is onto, the map
+
(18.69) (ω 0 , ψ 0 ) 7→ (dω 0 ) − qe(ψ 0 , ψ) − qe(ψ, ψ 0 )
must be onto, even for (ω 0 , ψ 0 ) constrained by
B 3 (ω 0 , ψ 0 ) := Dω ψ 0 + 21 ω 0 · ψ = 0
18.3. GENERIC REGULARITY OF THE MODULI SPACES 681


(see (18.65)). Thus, the projection of the image of B to W 2,k Λ2+ (X, iR) ⊕
W 2,k (Σ−
c (X)) contains W
2,k
Λ2+ (X, iR) ⊕ 0. We already know from the proof of
Theorem 18.48 that for F (ω, ψ, η) = 0 and ψ 6= 0, the map (where (ψ 0 , ω 0 ) is now
unrestricted)
(ψ 0 , ω 0 ) 7→ Dω ψ 0 + 12 ω 0 · ψ
is onto. Thus, the projection of the image of B to W 2,k Λ2+ (X) ⊕W 2,k (Σ−

c (X)) is
2,k
onto. It remains to show that the image B contains W (M, iR) ⊕ 0 ⊕ 0. For an ar-
bitrary function θ ∈ W 2,k+2 (X, R) , we have (ω, ψ, η) · eitθ = ω + 2itdθ, e−itθ ψ, η .


Since F (ω, ψ, η) = 0 ⇒ F (ω, ψ, η) · eitθ = 0, we know that


d −itθ
 
dt ω + 2itdθ, e ψ, η t=0 = (2idθ, −iθψ, 0) ∈ ker F∗(ω,ψ,η) .
Thus,
δ(2dθ) + 2i (h−iθψ, ψi − hψ, −iθψi) ,
 
B(2idθ, −iθψ) =
F∗(ω,ψ,η) (2idθ, −iθψ, 0)
= 2δ(dθ) + 2 (hθψ, ψi + hψ, θψi) , 0, 0 = 2 δdθ + 14 hψ, ψi θ , 0, 0 .
1
  

Now δdθ = −∆θ, where ∆ denotes the usual Laplace operator on C ∞ (X, R), ex-
tended to W 2,k+2 (X, R) . Since the operator −∆ + 14 hψ, ψi is elliptic and formally
self-adjoint, to show that
θ 7→ 2i δdθ + 41 hψ, ψi θ


is onto W 2,k (M, iR) it suffices to show that the kernel of −∆ + 41 hψ, ψi is 0. How-
ever, −∆θ + 14 hψ, ψi θ = 0 implies
Z Z
2
−∆θ + 14 hψ, ψi θ θ νg = kdθk + 41 hψ, ψi θ2 νg .

0=
X X
Thus, θ = 0 since ψ 6= 0. Hence, B is onto, and its index is the dimension of its
kernel. This index was found (see (18.43)) to be b1 −(1+b2+ )+ 14 c1 (L)2 − sig(X) .



Generic Regularity. We would like to use the Sard-Smale Theorem  (see
[402]) stated below to conclude that the set of all η ∈ W 2,k Λ2+ (X) , for which
p−1 (η) is either void or a manifold of dimension d(X, L), is residual. Recall that
a residual set is a subset containing the intersection of a countable collection of
open, dense subsets. A residual set of a complete metric space is dense by the Baire
Category Theorem.
Theorem 18.56 (Sard-Smale Theorem). Let f : B1 → B2 be a C r Fredholm
mapping between Banach manifolds (i.e., f∗x is Fredholm for each x ∈ B1 ), and let
(18.70) C := {f (x) : x ∈ B1 and f∗x is not onto}
(i.e., C denotes the set of critical values of f ). If r > {max(0, index(f∗x )) : x ∈ B1 }
(in particular, if f is C ∞ ) and B1 is separable, then B2 \ f (C) is residual.
Remark 18.57. Often B2 \ f (C) is called the set of regular values, but if
f is not onto, B2 \ f (C) will contain points which are not values of f at all (e.g.,
consider a constant map f ).

In order to apply Theorem 18.56 to p : MSWP k → W 2,k Λ2+ (X, iR) , we
need
682 18. SEIBERG-WITTEN THEORY

Theorem 18.58. For each (ω, ψ, η) ∈ SWP k ,


p∗[ω,ψ,η] : T[ω,ψ,η] MSWP k −→ W 2,k Λ2+ (X, iR)


is Fredholm. Moreover, the index of p∗[ω,ψ,η] is the same as the index of


B(ω,ψ,η),k : W 2,k+1 Λ1 (X, iR) ⊕ W 2,k+1 Σ+
 
c (X)

−→ W 2,k (X, iR) ⊕ W 2,k Λ2+ (X, iR) ⊕ W 2,k Σ−


 
c (X) .

Proof. We will use some of the facts we have shown about the related operator
B(ω,ψ,η),k . We use Notation 18.52 (p. 671), namely

U1 := W 2,k+1 Λ1 (X, iR) , U2 := W 2,k+1 (Σ+ c (X)) ,

Λ (X, iR) , V3 := W 2,k (Σ−


2,k 2,k 2+

V1 := W (X, iR) , V2 := W c (X)) ,

and
B 1 , B 2 , B 3 : U1 ⊕ U2 −→ V1 ⊕ V2 ⊕ V3 ,

B := where
0 0 0 0
1
B : U1 ⊕ U2 −→ V1 1
with B (ω , ψ ) := δω − 1
4 (hψ , ψi − hψ, ψ 0 i) ,
B 2 : U1 ⊕ U2 −→ V2 with B (ω , ψ ) := dω − qe(ψ , ψ) − qe(ψ, ψ 0 ) , and
2 0 0 0+ 0

B 3 : U1 ⊕ U2 −→ V3 with B 3 (ω 0 , ψ 0 ) := Dcω ψ 0 + 21 ω 0 · ψ.
From (18.63), we have
T[ω,ψ,η] MSWP k
n
∼ +
−→ (ω 0 , ψ 0 , η 0 ) ∈ T(ω,ψ,η) CWP k : (dω 0 ) − qe(ψ 0 , ψ) − qe(ψ, ψ 0 ) = η 0 ,
Dcω ψ 0 + 21 ω 0 · ψ = 0, δω 0 − 21 (hψ 0 , ψi − hψ, ψ 0 i) = 0
(ω 0 , ψ 0 , η 0 ) ∈ T(ω,ψ,η) CWP k :
 
=
B 1 (ω 0 , ψ 0 ) = 0, B 2 (ω 0 , ψ 0 ) = η 0 , B 3 (ω 0 , ψ 0 ) = 0

−→ Ker B 1 ⊕ B 3 .


Moreover,
+
p∗[ω,ψ,η] (ω 0 , ψ 0 , η 0 ) = η 0 = (dω 0 ) − qe(ψ 0 , ψ) − qe(ψ, ψ 0 ) ,
or under the final isomorphism in (18.65),
p∗[ω,ψ,η] = B 2 |Ker(B 1 ⊕B 3 ) : Ker B 1 ⊕ B 3

−→ V2 .
In order to show that p∗[ω,ψ,η] is Fredholm, we need the following
 ∼ 
1. Ker p∗[ω,ψ,η] −→ Ker B 2 |Ker(B 1 ⊕B 
3) is finite-dimensional, and
2. Im p∗[ω,ψ,η] = B 2 Ker B 1 ⊕ B 3 is closed and of finite codimension in
V2 .
Now (1.) holds, since
Ker p∗[ω,ψ,η] = Ker B 2 |ker(B 1 ⊕B 3 ) = Ker B 1 ⊕ B 2 ⊕ B 3 = Ker(B) ,
  

and B is elliptic.
To prove (2.), we first show that it suffices to prove that

(18.71) B(U1 ⊕ U2 ) ∩(V1 ⊕ 0V2 ⊕ V3 ) = 0.
Note that under the obvious imbedding V2 → 0V1 ⊕ V2 ⊕ 0V3 ,
 ∼
B 2 ker B 1 ⊕ B 3 −→ B(U1 ⊕ U2 ) ∩(0V1 ⊕ V2 ⊕ 0V3 ) .
18.3. GENERIC REGULARITY OF THE MODULI SPACES 683


Thus, since B(U1 ⊕ U2 ) is closed in V1 ⊕ V2 ⊕ V3 , B 2 ker B 1 ⊕ B 3 is closed in
V2 . Assuming (18.71), we have a 1-1 projection
⊥ ⊥ ∼
π : B(U1 ⊕ U2 ) −→ (V1 ⊕ 0V2 ⊕ V3 ) = 0V1 ⊕ V2 ⊕ 0V3 −→ V2 .
Note that from the yet to be proven fact (18.71), we also obtain

V1 ⊕ V2 ⊕ V3 = B(U1 ⊕ U2 ) + (V1 ⊕ 0V2 ⊕ V3 )
= B(U1 ⊕ U2 ) + 0V1 ⊕ V2 ⊕ 0V3 .

Thus, B 1 ⊕ B 3 : U1 ⊕ U2 → V1 ⊕ V3 is onto. We will prove that


  ⊥

Im(π) = π B(U1 ⊕ U2 ) = B 2 Ker B 1 ⊕ B 3 ,

from which we will obtain


⊥ ∼ ⊥
(18.72) π : B(U1 ⊕ U2 ) −→ B 2 Ker B 1 ⊕ B 3 .
⊥ ⊥
Then, since B is Fredholm, dim B(U1 ⊕ U2 ) < ∞, and dim B 2 Ker B 1 ⊕ B 3
< ∞. Since

0V1 ⊕ B 2 Ker B 1 ⊕ B 3 ⊕ 0V3




= B(U1 ⊕ U2 ) ∩(0V1 ⊕ V2 ⊕ 0V3 ) ⊆ B(U1 ⊕ U2 ) ,


  ⊥

we have π B(U1 ⊕ U2 ) ⊆ B 2 Ker B 1 ⊕ B 3 .
⊥
To show the reverse inclusion, let
 w ∈ B 2 Ker B 1 ⊕ B 3 ⊆ V2 . For arbitrary
v ∈ V1 ⊕ V3 , we have v = B 1 ⊕ B 3 (u) for some u ∈ U1 ⊕ U2 , since B 1 ⊕ B 3 : U1 ⊕
U2 → V1 ⊕ V3 is onto. Define
L : V1 ⊕ V3 −→ R by L(v) = − B 2 (u) , w .
⊥
L is well defined since w ∈ B 2 Ker B 1 ⊕ B 3 . Thus, there is some w0 ∈ V1 ⊕ V3
such that
w0 , B 1 ⊕ B 3 (u) = hw0 , vi = L(v) = − B 2 (u) , w


for all u ∈ u ∈ U1 ⊕ U2 . Then


hw0 + w, B(u)i = w0 , B 1 ⊕ B 3 (u) + B 2 (u) , w = 0,


⊥ ⊥
so that w0 + w ∈ B(U1 ⊕ U2 ) . Since w0 ∈ V1 ⊕ V3 and w∈ B 2 Ker B 1⊕ B 3 ,
⊥ ⊥
we have π(w0 + w) = w. Thus, B 2 Ker B 1 ⊕ B 3

⊆ π B(U1 ⊕ U2 ) , and we
have the isomorphism 18.72. Thus, from the yet to be proven (18.71), we obtain

index p∗[ω,ψ,η] = dim Ker p∗[ω,ψ,η] − dim Coker p∗[ω,ψ,η]
 ⊥ 
= dim Ker B − dim B 2 Ker B 1 ⊕ B 3
 

= dim Ker B − dim B(U1 ⊕ U2 )
= dim Ker B − dim Coker B = index B.

Hence, it only remains to show B(U1 ⊕ U2 ) ∩(V1 ⊕ 0V2 ⊕ V3 ) = 0. Suppose that
(f, 0, ξ) ∈ V1 ⊕ 0V2 ⊕ V3 and for all (ω 0 , ψ 0 ) ∈ U1 ⊕ U2 , we have (where the zero-th
684 18. SEIBERG-WITTEN THEORY

order operator Tψ and its adjoint Tψ∗ were defined near 18.51)

0 = hB(ω 0 , ψ 0 ) ,(f, 0, ξ)i = B 1 (ω 0 , ψ 0 ) , f + B 3 (ω 0 , ψ 0 ) , ξ


D ∗ E D ∗ E
= (ω 0 , ψ 0 ) , B 1 f + (ω 0 , ψ 0 ) , B 3 ξ
= (ω 0 , ψ 0 ) , i df, − 2i f ψ + (ω 0 , ψ 0 ) , Tψ∗ ξ, Dcω ξ
 

(18.73) = ω 0 , i df + Tψ∗ ξ + ψ 0 , Dcω ξ − 2i f ψ .

Then

(18.74) i df + Tψ∗ ξ = 0 and Dcω ξ − 2i f ψ = 0.

Since Dcω ψ = 0,

0 = Dcω Dcω ξ − 2i Dcω (f ψ) = Dcω Dcω ξ − 2i df · ψ − 2i f Dcω ψ


= Dcω Dcω ξ + 12 Tψ∗ ξ · ψ.

Thus,
Z Z
Dcω Dcω ξ + 21 Tψ∗ ξ · ψ, ξ v = hDcω Dcω ξ, ξi + 1
Tψ Tψ∗ ξ , ξ v

0= 2
ZX X

= hDcω ξ, Dcω ξi + 1
2 Tψ∗ ξ, Tψ∗ ξ v,
X

and so
0 = Tψ∗ ξ = −idf and i
2fψ = Dcω ξ = 0.

Hence, f is constant, and 2i f ψ = 0 then implies f = 0, since ψ 6= 0. Also, at any


point p where ψ(p) 6= 0, the map Λ1p (X, iR) → Wp− (X) , given by α 7→ α·ψ(p) is 1-1
since α·(α · ψ(p)) = ψ(p) 6= 0, and it is onto since dimR Λ1p (X, iR) = dimR Wp− (X) =
4. Hence Tψ and Tψ∗ are isomorphisms at any point where ψ(p) 6= 0. Hence 0 =
Tψ∗ ξ ⇒ ξψ = 0, but then ξ = 0 on the nonvoid open set where ψ 6= 0. Since Dcω ξ = 0,
unique continuation (Theorem 18.47) then yields ξ = 0. Thus, (f, 0, ξ) = (0, 0, 0)
and we have (18.71), as required. 

Theorem 18.59. For a η in a residual subset of W 2,k Λ2+ (X, iR) , either
p−1 (η) (i.e., SW k (η)) is empty
 or p−1 (η) is a submanifold of MSWP k of dimen-
1
1 2+

sion d(X, L) = b − 1 + b + 4 c1 (L)2 − sig(X) .

Proof. According to Theorem 18.58,

p : MSWP k → W 2,k Λ2+ (X, iR)




has a Fredholm derivative throughout


 MSWP k . Thus, it follows from Theorem
18.56 that W 2,k Λ2+ (X, iR) \ p(C) is residual,
 where C denotes the critical set
of p. For any point η ∈ W 2,k Λ2+ (X, iR) \ p(C) , we have that p∗ is onto at each
point in p−1 (η) . If p−1 (η) is not empty, then p−1 (η) is a submanifold of MSWP k
of dimension d(X, L) by Theorem 18.55. 
18.4. COMPACTNESS OF MODULI SPACES AND S-W INVARIANTS 685

4. Compactness of Moduli Spaces and the Definition of S-W Invariants


Here we establish that moduli spaces of solutions of generically perturbed S-W
equations are compact, and proceed to show that the S-W invariants are defined.
A Priori Bounds. First we examine the consequences of the S-W equations
in conjunction with the Spinc Bochner-Weitzenböck formula (18.27), p. 660.
Proposition 18.60. If (ω, ψ) ∈ C PU(1) × C ∞ (Σ+

c (X)) is a solution of the
perturbed S-W equations Dcω ψ = 0, Ωω+ = q(ψ) + η, for some η ∈ Ω2+ (X, iR), then
we have
2
0 = Dc2 ψ, ψ = − h∆ψ, ψi + 1
2 hRω ψ, ψi + 41 S |ψ|
1 4 1 2
(18.75) = − h∆ψ, ψi + 2 |ψ| + 2 hη, q(ψ)i + 41 S |ψ| .
at each point of X. Integrating this, yields
Z
2 4 2
(18.76) |∇ψ| + 12 |ψ| + 21 hη, q(ψ)i + 14 S |ψ| νg = 0.
X

Proof. Using (18.30) on p. 18.30, Ωω+ = q(ψ)+η, and Exercise 18.43 (p. 662),
we have
hRω ψ, ψi = Ωω+ , q(ψ) = hq(ψ) + η, q(ψ)i
2 4
= |q(ψ)| + hη, q(ψ)i = |ψ| + hη, q(ψ)i ,
from which (18.75) follows from Dcω ψ = 0 and Proposition 18.39, p. 660. Note that
(18.76) follows upon integration, since ∆ := −∇∗ ∇. 
Proposition 18.61. For (ω, ψ, η) as in Proposition 18.60, we have

√ 
2
(18.77) |ψ(x)| ≤ max 2 |η(y)| − 21 S(y), 0 .
y∈X

Hence, there is an a priori bound on kψkC 0 . In particular, if η = 0 and S ≥ 0,


then we must have ψ = 0 and Ωω+ = q(ψ) + η = 0.
Proof. Let E1 , . . . , E4 be locally-defined orthonormal frame field, parallel at
a point x ∈ X (i.e., ∇Ei Ej = 0 at x). Then
  X4
2
∆ |ψ| = Ei [Ei [hψ, ψi]]
i=1
4
X
= h∇Ei ∇Ei ψ(x), ψ(x)i + 2 h∇Ei ψ(x), ∇Ei ψ(x)i + hψ(x), ∇Ei ∇Ei ψ(x)i
i=1
2
= h−∇∗ ∇ψ(x), ψ(x)i + 2 |∇ψ(x)| + hψ(x), −∇∗ ∇ψ(x)i .
Hence,  
2 2
∆ |ψ| = 2 |∇ψ| − 2< h∇∗ ∇ψ, ψi .
 
2 2
If |ψ| achieves its maximum at x0 ,then ∆ |ψ| (x0 ) ≤ 0, and so at x0
 
2 2
−2< h∆ψ, ψi = 2< h∇∗ ∇ψ, ψi = 2 |∇ψ| − ∆ |ψ| ≥ 0.
From (18.75), we have
1 4 2
2 |ψ| + 12 < hη, q(ψ)i + 41 S |ψ| = < h∆ψ, ψi .
686 18. SEIBERG-WITTEN THEORY

√ 2
< hη, q(ψ)i ≤ |η| |q(ψ)| ≤ 2 |η| |ψ| ,
2 4
using 12 |q(ψ)| = |ψ| from Exercise 18.43 b. Thus, at x0 ,
1

2
√ 1

2 4
√ 2 2
2 |ψ| − 2 |η| + 2 S |ψ| = 12 |ψ| − 12 2 |η| |ψ| + 14 S |ψ|
1 4 2
≤ 2 |ψ| + 12 < hη, q(ψ)i + 41 S |ψ| = < h∆ψ, ψi ≤ 0.
Hence, either ψ(x0 ) = 0 (and so ψ = 0), or for all x ∈ X,
√  √ 
2 2
|ψ(x)| ≤ |ψ(x0 )| ≤ 2 |η(x0 )| − 21 S(x0 ) ≤ max 2 |η(y)| − 21 S(y) . 
y∈X

Corollary 18.62. For (ω, ψ, η) as in Proposition 18.60, we have


√ 2
Ωω+ = |q(ψ) + η| ≤ 2 |ψ| + |η|
 √ 
(18.78) ≤ max 2 |η(y)| − 12 2S(y), 0 + max(|η(y)|) .
y∈X y∈X

For a fixed oriented, Riemannian manifold (X, g) and form η, there is an upper
bound on the number of Spinc -structures for which the S-W moduli space has a
nonnegative formal dimension.
2 4
Proof. Using 21 |q(ψ)| = |ψ| from Exercise 18.43 b and Proposition 18.61,
(18.78) is evident. For a Spinc structure PSpinc → PU(1)×SO(4) → X, let L denote
the complex line bundle PU(1) ×U(1) C. Recall that the formal dimension of the
moduli space is
 
1 2
(18.79) 4 c1 (L) [X] − 2χ(X) − 3 sig(X) ,
2
where c1 (L) [X] denotes the evaluation of the cup-square of the first Chern class
c1 (L) on [X]. If this is assumed to be nonnegative, then
2
(18.80) 2χ(X) + 3 sig(X) ≤ c1 (L) [X].
For any connection ω on PU(1) with curvature form Ωω ∈ Ω2 (X, iR), we have
Z Z
2 ω ω −1
i i
Ωω + + Ωω− ∧ Ωω + + Ωω−
 
c1 (L) [X] = 2π Ω ∧ 2π Ω = 4π 2
X X
Z
−1
Ωω + + Ωω− ∧ ∗Ωω + − ∗Ωω−
 
= 4π 2

ZX
−1
= 4π2 Ωω+ ∧ ∗Ωω+ − Ωω+ ∧ ∗Ωω− + Ωω− ∧ ∗Ωω+ − Ωω− ∧ ∗Ωω−
ZX  
2 2
1
= 4π2 Ωω+ − Ωω+ , Ωω− + Ωω− , Ωω+ − Ωω− νg
X 
2 2
(18.81) = 4π1 2 Ωω+ − Ωω− .
2
By (18.78), we have an upper bound on kΩω+ k . Then (18.81) together with the
2
lower bound (18.80) on c1 (L) [X] gives us an upper bound on FA− . Hence, there
2
is an upper bound on kΩω k . By Hodge theory (see Theorem 17.63, p. 603), the
i
closed form 2π Ωω has a unique representative, say β ω , in the lattice of integral
i
harmonic 2-forms. Since kβ ω k ≤ 2π Ωω , we have an upper bound on kβ ω k and
there are only a finite number of such β ω in the lattice within a ball. Thus, there
i

are only finitely many possibilities for the canonical class c1 (L) = 2π Ωω . For each
18.4. COMPACTNESS OF MODULI SPACES AND S-W INVARIANTS 687

of these possibilities, there are only #H 1 (X; Z2 ) < ∞ distinct Spinc structures (see
18.36, 657). 

Sobolev Estimates.
Lemma 18.63. Let ω0 be a C ∞ connection on PU(1) → X. For any k ≥ 0, let
ω ∈ C 2,k+1 (PU(1) ). Then there is s ∈ W 2,k+2 (X, U(1)), such that s · ω := ω0 + α,
where α ∈ W 2,k+1 (Λ1 (X)), δα = 0, and
2 2
(18.82) kαk2,k+1 ≤ C Ωω+ 2,k
+ K,
where C and K are independent of α.
Proof. For some α0 ∈ W 2,k+1 (Λ1 (X)) we have ω = ω0 + α0 . For eiθ ∈
2,k+2
W (X, U(1)),
eiθ · ω = ω0 + α0 − 2idθ.
Thus, we first find θ ∈ W 2,k+2 (X, R), such that δ(α0 − 2idθ) = 0 or
(18.83) δdθ = − 2i δα0 .
Since δd = −∆ is a formally self-adjoint elliptic operator, this can be solved for θ
as long as δα0 is L2 orthogonal to ker(∆) which consists of the constant functions.
However, hδα0 , ci2,0 = hα0 , dci2,0 = 0, for any constant function c. Hence, we can
solve for θ and take α = α0 + 2idθ to obtain δα = 0. However, further modifications
are necessary to produce α satisfying 18.82. Thus, let α1 = α0 + 2idθ. Now,

·ω
(18.84) Ωω = Ωe = Ωω0 +α1 = Ωω0 + dα1 .
By the Hodge Decomposition Theorem (Theorem 17.62, p. 603), we can uniquely
write α1 = h + β, where h is harmonic (i.e., dh = 0 and δh = 0) and β ∈
W 2,k (Λ1 (X)) is orthogonal to the subspace of harmonic forms. Note that 0 =
δα1 = δ(h + β) = δβ. We have
(18.85) Ωω+ = Ωω0 + + dα1+ = Ωω0 + + dh+ + dβ + = Ωω0 + + dβ + .
The operator
(18.86) δ ⊕ d+ : Ω1 (X) −→ Ω0 (X) ⊕ Ω2+ (X)
is elliptic and β⊥ Ker(δ ⊕ d) = Ker(δ ⊕ d+ ) , since
+
(dγ) = 0 =⇒ 0 = δ((1 + ∗) dγ) = δdγ + δ ∗ dγ
= δdγ − ∗d ∗ ∗dγ = δdγ − ∗d2 γ = δdγ
=⇒ 0 = (δdγ, γ) = (dγ, dγ) =⇒ dγ = 0.
Thus, there is a constant C (independent of β) such that
 
2 2 2 2 2
kβk2,k+1 ≤ C kδβk2,k + dβ + 2,k = C dβ + 2,k = C Ωω+ − Ωω0 + 2,k
2 2 2
≤ C Ωω+ 2,k
+ C Ωω0 + 2,k
= C Fω+ 2,k
+ K0 .
2
However, α1 = h+β and we cannot deduce that khk2,k+1 ≤ K 00 , and hence we need
a further gauge transformation. The group H 1 (X; Z) ⊂ H 1 (X; R) can be regarded
as a lattice in the b1 -dimensional vector space of harmonic 1-forms. For a harmonic
688 18. SEIBERG-WITTEN THEORY

1-form ξ ∈ H1 (X; Z) ,we have a well-defined function s0 ∈ C ∞ (X, U(1)) , given by
R
s0 (x) := exp 2πi γ ξ where γ is a path joining a fixed x0 to x in X. We have

eiθ s0 · ω = s0 · eiθ · ω = ω0 + α1 − 2s0 −1 ds0 = ω0 + β + h − 4πiξ.


 

2
If d is the k · k2,k+1 -diameter of a fundamental cell of the lattice H 1 (X; Z) , then
we can choose ξ ∈ H 1 (X; Z) so that kh − 4πiξk ≤ 2πd. Then for s := eiθ s0 and
α := β + h − 4πiξ, we have s · ω := ω0 + α, where
2 2 2 2
kαk2,k+1 = kβ + h − 4πiξk2,k+1 ≤ kβk2,k+1 + kh − 4πiξk2,k+1
2
≤ C Fω+ 2,k
+ K 0 + 2πd

and C and K 0 + 2πd are independent of α. 

Theorem 18.64. For k ≥ 4, let (ω 0 , ψ 0 ) ∈ C 2,k+1 (PU(1) ) × W 2,k+1 (Σ+


c (X)) be
a solution of
0 0
(18.87) Ωω + − q(ψ 0 ) = η and Dcω ψ 0 = 0,
where η ∈ iΩ2+ (X) is a C ∞ form and an achieved regular value of the projection
p : MSWP k+1 → W 2,k (Λ2+ (X, iR)). Let ω0 be a fixed choice of a C ∞ connection.
In accordance with Lemma 18.63, for some s ∈ W 2,k+2 (X, U(1)), we have
s · ω 0 = ω0 + α, for α ∈ W 2,k+1 (iΛ1 (X)) with δα = 0.
Let (ω, ψ) := s · (ω 0 , ψ 0 ). We assume that the harmonic component of 2πi
1
α lies in a
fixed fundamental domain of the lattice of integral harmonic 1-forms (see the proof
of Lemma 18.63). Then there are constants C(k 0 ) depending only on g, η, ω0 and
k 0 ≥ 3, such that
(18.88) kαk2,k0 + kψk2,k0 ≤ C(k 0 ),
where kψk2,k0 is computed using the Levi-Civita connection for X and the connec-
tion ω0 on PU(1) . In particular, α and ψ are C ∞ , as is ω = s · ω 0 = ω0 + α.
Proof. Since W 2,5 (Σ+ 2 +
c (X)) ⊆ C (Σc (X)), the argument of Proposition 18.61
applies to yield a uniform bound on kψkC 0 . Here and elsewhere, uniform bound
means a bound which only depends on η and the metric g on X (in particular, not
on ψ or ω). By (18.76), we have
Z Z
2 2 4
(18.89) |∇ω ψ| ≤ 1 1 1
4 |S| |ψ| + 2 |ψ| + 2 |hη, q(ψ)i| νg .
X X

However, this does not yet give us a uniform bound on kψk2,1 , since we are using
ω0 (not ω) to define kψk2,1 . Now

(18.90) ∇ω ψ = ∇ω0 ψ + 12 αψ or ∇ω0 ψ = ∇ω ψ − 21 αψ.


Thus, we need a uniform bound on kαψk2,0 to produce a uniform bound on kψk2,1 .
For this it suffices to produce a uniform bound on kαk2,0 , since there is a uniform
bound on kψkC 0 . We can in fact produce a uniform bound on kαk2,2 as follows.
By Lemma 18.63, we have
2 2
(18.91) kαk2,k+1 ≤ C Ωω+ 2,k
+ K,
18.4. COMPACTNESS OF MODULI SPACES AND S-W INVARIANTS 689

2
whence it suffices to obtain a uniform bound on kΩω+ k2,1 to get a uniform bound
2
on kαk2,2 . Since
δ + d : Ω2 (X) −→ Ω1 (X) ⊕ Ω3 (X)
has injective symbol, the Sobolev extension
2,k+1
: W 2,k+1 Λ2 (X) −→ W 2,k Λ1 (X) ⊕ W 2,k Λ3 (X)
  
(δ + d)
has a finite-dimensional kernel H2 (X) consisting of C ∞ harmonic 2-forms and
W 2,k+1 Λ2 (X) = H2 (X) ⊕ H2 (X)⊥ .


2,k+1
Since (δ + d) |H2 (X)⊥ has a continuous inverse on its image, there is a constant

C, such that for all β ∈ W 2,k+1 Λ2 (X) ,
2,k+1
β⊥ 2,k+1
≤ C (δ + d) β ,
2,k

where β ⊥ denotes the orthogonal projection of into H2 (X)⊥ . Since δ = ± ∗ d∗, if


β is self-dual, then δβ = ± ∗ dβ, and so kδβk2,k = kdβk2,k . Thus,

β ⊥ 2,k+1 ≤ 2C d2,k+1 β 2,k .
Applying this with β = Ωω+ , we get
⊥ √
(18.92) Ωω+ ≤ 2C d2,k+1 Ωω+ 2,k
.
2,k+1

Since all norms are equivalent on the finite-dimensional kernel H2 (X), the fact that
there is a uniform C 0 bound on Ωω+ = q(ψ) + η gives us a uniform W 2,k bound on
the harmonic part of Ωω+ . To get a uniform bound on Ωω+ itself, first note (where
θ denotes the Levi-Civita connection) that
(18.93) ∇θ Ωω+ = qe(∇ω ψ, ψ) + qe(ψ, ∇ω ψ) + ∇θ η.
While we are in the process of getting a uniform bound on kψk2,1 , we already know
that kψkC 0 and k∇ω ψk2,0 are uniformly bounded by (18.77) and (18.89). Thus,
∇θ Ωω+ 2,0 is uniformly bounded by (18.93). This yields a uniform bound on

d2,k+1 Ωω+ 2,0
in (18.92), and hence on (Ωω+ ) . Thus, kΩω+ k2,1 is uniformly
2,1
2
bounded. By (18.91), kαk2,2 is then uniformly bounded, and kψk2,1 is uniformly
bounded via (18.90). Now we wish to show that kψk2,3 is uniformly bounded. Note
that
0 = Dcω ψ = Dcω0 ψ + 12 α · ψ ⇒ Dcω0 ψ = − 21 α · ψ.
Since kαk2,2 and kψkp2 ,0 are bounded for each p2 ≥ 1, Proposition 16.24 (p. 498)
implies that kα · ψkp3 ,k3 is uniformly bounded, provided
 
k3 − p43 ≤ min 2 − 24 , 0 − p42 = − p42 .
In particular, taking p2 ≥ 4, kα · ψk4,0 is uniformly bounded, and hence kDcω0 ψk4,0
is uniformly bounded. Since Dcω0 is an elliptic operator, we not only have a contin-
uous map
Dcω0 : W 4,1 (Σ+
c (X)) −→ W
4,0
(Σ−
c (X)),
but also by Proposition 16.23 (498),
 
kψk4,1 ≤ C kDcω0 ψk4,0 + kψk4,0
690 18. SEIBERG-WITTEN THEORY

for some constant C depending only on Dcω0 . Thus, there is a uniform bound on
kψk4,1 . Since kαk2,2 and kψk4,1 are bounded, Proposition 16.24 (p. 16.24) implies
that kα · ψkp3 ,k3 is uniformly bounded, provided k3 − p43 < min 2 − 42 , 1 − 44 = 0.


Thus, with k3 = 1 and p3 = 3, kDcω0 ψk3,1 = kα · ψk3,1 is uniformly bounded. Using


 
kψk3,2 ≤ C kDcω0 ψk3,1 + kψk3,1 ,

we deduce that kψk3,2 is uniformly bounded. Then the cited Proposition 16.24
implies that kDcω0 ψk2,2 = kα · ψk2,2 is uniformly bounded, since kαk2,2 and kψk3,2
are uniformly bounded (note that max 2 − 24 , 2 − 43 = 23 > 0). Using

 
2,3
kψk2,3 ≤ C (Dcω0 ) ψ + kψk2,2 ,
2,2

kψk2,3 is uniformly bounded. Moreover, from Ω+ +


ω = q(ψ)+η, we get that kΩω k3,2 is
uniformly bounded since Proposition 16.24 implies that kq(ψ)k3,2 is bounded (here
k − np = 2 − 43 > 0). By (18.82), we then have a uniform bound on kαk2,4 .
In summary, thus far we have shown kαk2,k0 + kψk2,k0 ≤ C(k 0 ) for k 0 = 2 and
k = 3. Assume that we have this result for some k 0 ≥ 3. Once again, we apply
0

Proposition 16.24 and get kDcω0 ψk2,k0 = kα · ψk2,k0 is uniformly bounded since
k 0 − 24 > 0. Then,
 
kψk2,k0 +1 ≤ C kDcω0 ψk2,k0 + kψk2,k0

gives us a uniform bound on kψk2,k0 +1 . From Ωω+ = q(ψ) + η and the fact that
kq(ψ)k2,k0 is uniformly bounded, we get a uniform bound for kΩω+ k2,k0 . Finally,
(18.91) gives us a uniform bound on kαk2,k0 +1 . 

Corollary 18.65. For k ≥ 4, let (ωn0 , ψn0 ) ∈ C 2,k+1 (PU(1) ) × W 2,k+1 (Σ+
c (X))
be a sequence of solutions of the S-W equations
0 0
(18.94) Ωω + − q(ψ) = η and Dcω ψ 0 = 0,
where η ∈ iΩ2+ (X) is a C ∞ form and an achieved regular value of the projection
p : MSWP k+1 → W 2,k (iΛ2+ (X)). The sequence (ωn , ψn ) := sn ·(ωn0 , ψn0 ) of C ∞ solu-
tions of the S-W equations produced in Theorem 18.64 has a subsequence which is
convergent in the C ∞ topology to a C ∞ solution (ω, ψ) of the S-W equations (18.94).
Proof. The inclusion of W 2,k spaces in the corresponding C k−3 spaces is
compact for k ≥ 3. This means that a sequence which is bounded in W 2,k has
a subsequence convergent in C k−3 . Because of the bound (18.88) for k ≥ 3, we
can choose a subsequence of (An , ψn ) converging in C 0 . Then we can choose a
subsequence of the subsequence converging in C 1 . Continuing, we obtain a sequence
of subsequences, and the diagonal subsequence converges in C k for all k, and hence
in C ∞ to a C ∞ solution. 

Compactness of Moduli Spaces. Corollary 18.65 does not imply that the
solution (ω, ψ) to which (ωn , ψn ) converges is in p−1
k (η), since we might have ψ = 0,
and then (ω, 0) ∈/ CWP k ; see (18.44), p. 668). In order to prove that
 the manifold
p−1
k (η) is compact, we need to avoid using those η ∈ W
2,k
Λ2+ (X) for which there
are reducible solutions (i.e., solutions for which ψ = 0). If b+ 2 > 0, we now show
18.4. COMPACTNESS OF MODULI SPACES AND S-W INVARIANTS 691

how such η can be avoided. If (ω, ψ) is a reducible solution of the S-W equations,
then
(18.95) Ωω+ = q(ψ) + η = q(0) + η = η.
We can always write
Ωω = H(Ωω ) + dα ,
where H(Ωω ) is the harmonic part of Ωω and α∈ W 2,k+1 1
 (Λ (X, iR)). The coho-
i
mology class of Ω is determined by PU(1) , since 2π Ω = c1 (PU(1) ). Thus, H(Ωω )
ω ω

is uniquely determined
 by the metric and PU(1) and is independent of  the choice
of ω ∈ C 2,k+1 PU(1) . For a given metric g and PU(1) , we set γg PU(1) := H(Ωω ) .
Thus,
+ + +
(H(Ωω ) + dα) = H(Ωω ) + dα+ = γg PU (1) + dα+ .
Let
H+ : W 2,k+1 (Λ2 (X, iR)) −→ H2+ (X; iR)g
denote the L2 -orthogonal projection (given by Hodge theory) onto the self-dual har-

monic 2-forms relative to g. Then (18.95) will not hold for any ω ∈ C 2,k+1 PU(1) ,
as long as
+
(18.96) H+ (η) 6= γg PU(1) ,
+
since then H+ (Ωω ) = γg PU(1) 6= H+ (η) . The affine subspace
n + o
Ag PU(1) := η ∈ W 2,k+1 (Λ2+ (X, iR)) : H+ (η) = γg PU (1)

(18.97)

is closed and of codimension b+


2 in W
2,k+1
(Λ2+ (X, iR)).
Theorem 18.66 (Compactness of SW k (η)). For any fixed metric g and Spinc
structure for X, let pk : MSWP k → W 2,k (Λ2 (X, iR)+ ) denote the projection where
k ≥ 4. Assume that b+ 2 > 0. For generic η ∈ W
2,k
(Λ2 (X, iR)+ ) (i.e., for η in a
−1
residual subset) and SW k (η) := pk (η), either SW k (η) = ∅ or SW k (η) is a compact
C ∞ manifold of dimension
d(X, L) = b1 − 1 + b2+ + 41 c1 (L)2 − sig(X) ,
 
(18.98)
where L denotes the canonical line bundle for the Spinc -structure. In the case
η ∈ Ω2+ (X, iR) := C ∞ (Λ2+ (X, iR)) is an achieved regular value of
pk : MSWP k → W 2,k (Λ2 (X, iR)+ ),
we let SW ∞ (η) denote the set of C ∞ solutions of the S-W equations modulo
C ∞ (X, U(1)). Then the natural map
fk : SW ∞ (η) → SW k (η),
given by fk ([ω, ψ]) = [ω, ψ]k is a bijection.
Proof. By Theorem 18.59 (p. 684), the only new issue in the first claim is
compactness. By intersecting the  residual subset of Theorem 18.59 with the open,
dense complement of Ag PU(1) (using b+ 2 > 0), we obtain residual  subset, say T ,
of W 2,k (Λ2 (X, iR)+ ) such for η ∈ T , there are no ω ∈ C 2,k+1 PU(1) with Ωω+ = η,
and hence no solutions of the perturbed S-W equations with ψ = 0. Thus, for η ∈ T ,
the limit (ω, ψ) in Corollary 18.65 is not reducible and hence it is in p−1 k (η) which
is then compact. We now prove that fk : SW ∞ (η) → SW k (η) is a bijection. If
[ω 0 , ψ 0 ]k ∈ SW k (η), then (by Theorem 18.64), [ω 0 , ψ 0 ]k = [ω, ψ]k for a C ∞ solution
692 18. SEIBERG-WITTEN THEORY

(ω, ψ) of the S-W equations perturbed by η. Thus, fk is onto. To show that fk is


1-1, we argue as follows. If [ω1 , ψ1 ]k = [ω2 , ψ2 ]k for some C ∞ ω1 , ψ1 , ω2 , ψ2 , then
(ω2 , ψ2 ) = s ·(ω1 , ψ1 ) for some s ∈ W 2,k+2 (X, U(1)) , i.e.
(18.99) 2s−1 ds = ω2 − ω1 , and ψ2 = sψ1 .
Since locally s = eiθ , we have 2idθ = 2s−1 ds = ω2 − ω1 is C ∞ and hence θ and s
are C ∞ . Hence [ω1 , ψ1 ] = [ω2 , ψ2 ] in SW ∞ (η). 
Remark 18.67. Using the compactness of SW k (η0 ) we obtain: If pk has an
achieved regular value say η0 , then all η ∈ W 2,k (Λ2 (X, iR)+ ) which are sufficiently
close to η0 will be achieved regular values. Since C ∞ (Λ2+ (X, iR)) is dense in
W 2,k (Λ2 (X, iR)+ ), there will be C ∞ achieved regular values for which the bound
(18.88) applies.
Definition of the S-W Invariant. We wish to show how the moduli space
Mk+1 (η) is used to define an integer, known as the S-W-invariant for a given Spinc
structure for X. It turns out that the S-W-invariant is independent of a suitable
choice of η and of the Riemannian metric on X, provided that b+ 2 ≥ 2. We first
give the definition and establish the independence later.
Let η ∈ Ω2+ (X, iR)+ = C ∞ (Λ2+ (X, iR)) be a regular achieved value of
pk : MSWP k → W 2,k (Λ2 (X, iR)+ )
so that Theorem 18.66 applies, and hence SW ∞ (η) is a compact, smooth manifold
of dimension d(X, L). We let Sol(η) denote the set of all C ∞ solutions of the S-W
equations (perturbed by η) so that
SW ∞ (η) = Sol(η) /C ∞ (X, U(1)) .
Choose x0 ∈ X and consider the subgroup say G0 ⊆ C ∞ (X, U(1)) of those s, such
that s(x0 ) = 1. We then have a mapping
(18.100) π : Sol(η) /G0 −→ SW(η) := SW ∞ (η)
with π([(ω, ψ)]0 ) = [(ω, ψ)] , where [(ω, ψ)]0 denotes the G0 -orbit of (ω, ψ) ∈ Sol(η) .
We claim that (18.100) is a principal U(1)-bundle. For [(ω, ψ)]0 ∈ SW k /G0 , let
U(1) act on Sol(η) /G0 from the right via
[(ω, ψ)]0 · eiθ0 = e−iθ0 ·(ω, ψ) 0 = ω, e−iθ0 ψ 0 .
   

To show that the action is free, suppose that in SW k /G0 we have


ω, e−iθ0 ψ 0 = [(ω, ψ)]0 · eiθ0 = [(ω, ψ)]0 .
 

Then there is s ∈ G0 , such that ω = −2s−1 ds + ω (i.e., s is constant and hence 1,


since s ∈ G0 ) and e−iθ0 ψ = s−1 ψ = ψ. Thus, since ψ 6= 0, e−iθ0 = 1, and hence
the U(1) action is free. The fibers of π are the orbits of this U(1)-action. Indeed,
if π([(ω1 , ψ1 )]0 ) = π([(ω2 , ψ2 )]0 ) , then [(ω1 , ψ1 )] = [(ω2 , ψ2 )] . Hence (ω2 , ψ2 ) = s ·
(ω1 , ψ1 ) for some s ∈ C ∞ (X, U(1)). We can write
−1 −1
s = s s(x0 ) s(x0 ) = s0 s(x0 ) , where s0 := s s(x0 ) ∈ G0 .
Then
[(ω2 , ψ2 )]0 = [s ·(ω1 , ψ1 )]0 = [s0 s(x0 ) ·(ω1 , ψ1 )]0
−1 −1
= [s0 ·(ω1 , ψ1 )]0 · s(x0 ) = [(ω1 , ψ1 )]0 · s(x0 ) .
18.4. COMPACTNESS OF MODULI SPACES AND S-W INVARIANTS 693

Thus, U(1) acts transitively on the fibers of π, and it is clear that the U(1)-action
preserves the fibers. The local triviality of 18.100 may be shown ultimately using
constructions as in the proofs of Theorem 18.53 (p. 673) and Theorem 18.54 (p. 679),
and so (18.100) defines a principal U(1)-bundle. Let c1 (η) ∈ H 2 (SW(η) ; Z) denote
the first Chern class of (18.100). Note that d(X, L) in (18.98) is even iff b1 + b+
2 is
odd. Once we give SW(η) an orientation, if d(X, L) is even, then the evaluation of
the 21 d(X, L)-fold cup product

d(X,L)/2 d(X,L)/2
c1 (η) := c1 (η) ` ··· ` c1 (η)

on the fundamental class [SW(η)] ∈ H d(X,L)/2 (SW(η) ; Z) yields the Seiberg-


Witten invariant
d(X,L)/2
(18.101) SW([PSpinc ]) := c1 (η) [SW(η)] ∈ Z.

The orientation for SW(η) is defined as follows. The tangent space of SW(η)
is (see Theorem 18.58, p. 682)

T[ω,ψ,η] p−1 (η)



 ∼
= Ker p∗[ω,ψ,η] : T[ω,ψ,η] MSWP k −→ W 2,k Λ2+ (X, iR) −→ Ker B(ω,ψ,η),k
 0 0 0 
 (ω , ψ , η ) ∈ T(ω,ψ,η)
 CWP k :
+


(dω 0 ) − qe(ψ 0 , ψ) − qe(ψ, ψ 0 )
  

−→  ,

 0 = B(ω 0 , ψ 0 , η 0 ) :=  Dcω ψ 0 + 12 ω 0 · ψ 
δω 0 − 21 hψ, ψ 0 i
 

where

B = B(ω,ψ,η),k : W 2,k+1 Λ1 (X, iR) ⊕ W 2,k+1 Σ+


 
(18.102) c (X)

−→ W 2,k (X, iR) ⊕ W 2,k Λ2+ (X, iR) ⊕ W 2,k Σ−


 
c (X) .

Thus, we need some way to orient Ker(B). Associated with any Fredholm operator
F : V → W, between Hilbert spaces, there is the 1-dimensional space
(18.103)

det F := Λtop (Ker F ) ⊗ Λtop (Ker F ∗ ) := Λdim(ker F ) (Ker F ) ⊗ Λdim(ker F ) (Ker F ∗ ).

Topological ramifications of this construction are discussed following Definition


3.41 (pp.92ff). In the event that F is onto (i.e., Ker F ∗ = {0}), orienting det F is
equivalent to orienting Ker F . Let
δω 0 − 1−t 0
 
2 hψ, ψ i
+
(18.104) B(t)(ω 0 , ψ 0 , η 0 ) :=  (dω 0 ) − (1 − t)(eq (ψ 0 , ψ) + qe(ψ, ψ 0 ))  .
ω 0 1−t 0
Dc ψ + 2 ω · ψ = 0

This is a family of elliptic operators over the unit interval [0, 1] and it can be
shown that {det B(t) : t ∈ [0, 1]} is a continuous line bundle over [0, 1] . As [0, 1] is
contractible, this line bundle is trivial and an orientation for det B(1) determines
an orientation for det B(0) (i.e., for ker B, since B is onto, by the assumption that
η is an achieved regular value of p; see the proof of Theorem 18.55, p. 680). Now

B(1) = δ, d+ , Dcω ,

694 18. SEIBERG-WITTEN THEORY

and we have (see the computations following (18.42), p. 666)


Ker B(1) = H1 (X, iR) ⊕ Ker Dcω and
Ker B(1)∗ = iR ⊕ H2+ (X, iR) ⊕ Ker(Dcω∗ ) .
Thus,
det B(1) = Λtop H1 (X, iR) ⊕ ker Dcω ⊗ Λtop iR ⊕ H2+ (X, iR) ⊕ Ker(Dcω∗ ) .
 

As ker Dcω and ker Dcω∗ are complex spaces, they have a natural orientation. More-
over, iR has the natural orientation i. Thus, we need only to choose a fixed orienta-

tion for H 1 (X, iR) ∼
= H1 (X, iR) and H 2+ (X, iR) −→ H2+ (X, iR) to determine an
orientation for det B(1). We assume that this has been done. Hence, the orientation
class [SW(η)] in (18.101) is determined.

Metric Dependence of Connections and Dirac Operators. It remains


to show that SW(PSpinc ) is independent of the metric g on X and the perturbing
form η. Let F X(g) denote the bundle of oriented orthonormal frames for to the
metric g. We first point out that when the metric is on X is changed from g1 to g2 ,
the bundle of oriented orthonormal frames on X changes from F X(g1 ) to F X(g2 ).
These are seen to be isomorphic as follows.
Let
Aut(T X) ⊂ End(T X) = T 1,1 X
denote the bundle of invertible linear transformations of the tangent spaces of X.
First we show that there is a canonical automorphism ϕ ∈ C ∞ (Aut(T X)) of the
tangent bundle, such that
g1 (Y, Z) = (g2 · ϕ)(Y, Z) := g2 (ϕ(Y ) , ϕ(Z))
for all Y, Z ∈ Tp X and p ∈ X. More precisely, we prove
Proposition 18.68. For Riemannian metrics g1 and g2 on X, there is a unique
positive, g1 -symmetric ϕ ∈ C ∞ (Aut(T X)) such that
g1 = g2 · ϕ and g1 (ϕ(Y ) , Z) = g1 (Y, ϕ(Z)) .
Proof. At each point p ∈ X, αp ∈ GL(Tp X) exists, such that g2 · αp = g1 at
p. Just let αp take g1 -orthonormal basis of Tp X to a g2 -orthonormal basis of Tp X.
If βp ∈ GL(Tp X) also satisfies g2 · βp = g1 , then βp = αp ◦ A for some orthogonal
(relative to g1 ) A ∈ Og1 (Tp X). Indeed,
g1 · A = g1 · αp−1 ◦ βp = g1 · αp−1 · βp = g2 · βp = g1 .
 

Using T to denote transpose relative to g1 , we have


T
βp ◦ βpT = (αp ◦ A) ◦(αp ◦ A) = αp ◦ A ◦ AT ◦ αpT = αp ◦ αpT .
Thus, the positive g1 -symmetric map αp ◦ αpT is independent of the choice of αp ,
q
and it has a unique, g1 -symmetric, positive square-root αp ◦ αpT . We take ϕp ∈
q
Aut(Tp X) to be αp ◦ αpT , and ϕ ∈ C ∞ (Aut(T X)) is defined. 

Notation 18.69. Let Sg+1 X → X denote the bundle of positive g1 -symmetric


operators.
18.4. COMPACTNESS OF MODULI SPACES AND S-W INVARIANTS 695

Thus, we have a canonical ϕ ∈ C ∞ Sg+1 X ⊆ C ∞ (Aut(T X)) satisfying g2 · ϕ = g1 .




We define
Lϕ : F X(g1 ) −→ F X(g2 )
by Lϕ (u) := ϕp ◦ u where F X(g1 ) 3 u : R4 → Tp X is an isometry (i.e., an oriented
orthonormal frame relative to g1 ). Note that Lϕ is equivariant (Lϕ (u ◦ A) = ϕ ◦
u ◦ A = Lϕ (u) ◦ A for A ∈ SO(4)). Consequently, the Levi-Civita connection for
F X(g2 ) pulls back to a connection (but not necessarily the Levi-Civita connection)
for F X(g1 ). We can extend Lϕ to
I × Lϕ : PU(1) × F X(g1 ) −→ PU(1) × F X(g2 ).
c
Then pull back by I × L−1 ϕ of Spin -bundles πr c (g1 ) : PSpinc (g1 ) → PU(1) × F X(g1 )
gives us a canonical one-to-one correspondence between the Spinc structures for
(X, g1 ) and those for (X, g2 ) :
Jϕ !
PSpinc (g1 ) −→ PSpinc (g2 ) := I × L−1 ϕ PSpinc (g1 )
(18.105) ↓ πrc (g1 ) ↓ πrc (g2 )
I×Lϕ
PU(1) × F X(g1 ) −→ PU(1) × F X(g2 ), where
!
I × L−1
ϕ PSpinc (g1 )
p, L−1
  
:= ((p, u) , pe) ∈ PU(1) × F X(g2 ) × PSpinc (g1 ) : πrc (g1 )(e
p) = ϕ u .
That is, the fiber of PSpinc (g2 ) over (p, u) ∈ PU(1) × F X(g2 ) is given by
 ! 
I × L−1
ϕ PSpinc (g1 ) = {(p, u)} × PSpinc (g1 )(p,L−1
ϕ u)
.
(p,u)
c
The action of σ ∈ Spin (4) on PSpinc (g2 ) is given by
((p, u) , pe) · σ = ((p, u) · rc (σ) , pe · σ) .
The map πrc (g2 ) : PSpinc (g2 ) → PU(1) × F X(g2 ) is defined by
πrc (g2 )((p, u) , pe) := (p, u) ,
c
and note that πrc (g2 ) is r -equivariant, since
πrc (g2 )(((p, u) , pe) · σ) = πrc (g2 )(((p, u) · rc (σ) , pe · σ))
= (p, u) · rc (σ) = πrc (g2 )((p, u) , pe) · rc (σ) .
p) = p, L−1

Also, for pe ∈ PSpinc (g1 ) with πrc (g1 )(e ϕ u , we have

p) := ((p, u) , pe) ∈ PSpinc (g2 ) and


Jϕ (e
(πrc (g2 ) ◦ Jϕ )(e
p) = πrc (g2 )((p, u) , pe) = (p, u)
= (I × Lϕ ) p, L−1

ϕ u = ((I × Lϕ ) ◦ πrc (g1 ))(e
p) .
Connections and equivariant functions (e.g., twisted spinor fields) on PSpinc (g2 )
pullback via Jϕ to connections and equivariant functions on PSpinc (g1 ). Thus,
the connections and twisted spinor fields on PSpinc (g2 ) associated with the metric
g2 can all be identified with connections and twisted spinor fields on PSpinc (g1 )
associated with a fixed metric g1 . However, the pull-back of the Dirac operator
and the star operator for g2 are not the same as the usual Dirac operatorand star
operator for g1 , and we must elaborate on this. Note that ϕ ∈ C ∞ Sg+1 X induces
ϕ∗ ∈ C ∞ (Aut(Λ∗ (X))), where
(ϕ∗ α)(Y1 , ..., Yk ) = α(ϕ(Y1 ) , ..., ϕ(Yk )) .
696 18. SEIBERG-WITTEN THEORY

The star operator ∗g2 of g2 = g1 · ϕ−1 is given by


∗g2 = ϕ−1∗ ◦ ∗g1 ◦ ϕ∗ .
Observe that
η ∈ Ω2+g1 (X, iR) ⇔ ϕ−1∗ η ∈ Ω2+g2 (X, iR), since
ϕ−1∗ η = ϕ−1∗ ◦ ∗g1 ◦ ϕ∗ ϕ−1∗ η = ϕ−1∗ (∗g1 η) .
  
∗g2
Note that
Ωω+g2 = ϕ−1∗ η ⇔ 1 −1∗
◦ ∗g1 ◦ ϕ∗ Ωω = ϕ−1∗ η

2 1+ϕ
1 ∗ ∗ ω
⇔ 2 (ϕ + ∗g1 ◦ ϕ ) Ω = η
1 ∗ ω
⇔ 2 (1 + ∗g1 ) ϕ Ω = η
∗ ω +g1
⇔ (ϕ Ω ) = η.
+
Thus, in the case ψ = 0, the S-W equation Ωω g2 = ϕ−1∗ η relative to g2 is equivalent
+
to (ϕ∗ Ωω ) g1 = η. The other S-W equation Dcω ψ = 0 also depends on the metric.
Let θ2 ∈ C(PSO(4) (g2 )) denote the Levi-Civita connection of g2 and let ω ∈ C(PU(1) ).
Then L∗ϕ θ2 ∈ C(PSO(4) (g1 )), but θ1 6= L∗ϕ θ2 in general. Due to the commutativity
of (18.105),
 
−1 ∗ −1 ∗
(rc ) Prc (g1 ) ω ⊕ L∗ϕ θ2 = Jϕ∗ (rc )

Prc (g2 ) (ω ⊕ θ2 ) .

If Σc,g (X) = Σ+
c,g (X)⊕Σc,g (X) denotes the bundle of virtual twisted spinors relative
to g, then
Jϕ ×Spinc IΣ4

Σc,g1 (X) = P Spinc (g1 ) × Spinc Σ4 −→ PSpinc (g2 ) ×Spinc Σ4 = Σc,g2 (X) ,
where
(Jϕ ×Spinc IΣ4 )([p, w]) := [Jϕ (p) , w] .

If ψ ∈ C (Σc,g2 (X)) corresponds to the equivariant Σ4 -valued function ψe ∈
0 −1
Ω (PSpinc (g2 ), Σ4 ), then (Jϕ ×Spinc IΣ4 ) (ψ) ∈ C ∞ (Σg1 ,c (X)) corresponds to the
pull-back
0
Jϕ∗ ψe := ψe ◦ Jϕ ∈ Ω (PSpinc (g1 ), Σ4 ).

Consequently, we will denote the isomorphism C ∞ (Σc,g2 (X)) −→ C ∞ (Σc,g1 (X))
−1
induced by (Jϕ ×Spinc IΣ4 ) simply by

Jϕ∗ : C ∞ (Σc,g2 (X)) −→ C ∞ (Σc,g1 (X)).
In other words, there is a well-defined notion (i.e., Jϕ∗ ) of pull-back of twisted
spinors, induced by ϕ ∈ C ∞ Sg+1 X . We let ∇(ω,θ2 ) denote covariant differentia-


Prc (g2 ) (ω ⊕ θ2 ) , and ∇(ω,Lϕ θ2 )
−1 ∗
tion on C ∞ (Σc,g2 (X)) for the connection (rc )

−1 ∗
denote covariant differentiation on C ∞ (Σc,g1 (X)) for (rc ) Prc (g1 ) ω ⊕ L∗ϕ θ2 .


One can check that for ψ ∈ C ∞ (Σc,g2 (X)),


(ω,L∗ θ2 ) ∗  (ω,θ )
∇Y ϕ Jϕ ψ = Jϕ∗ ∇ϕ(Y )2 ψ.

In other words, regarding ∇(ω,Lϕ θ2 ) Jϕ∗ ψ ∈ C ∞ Λ1 (X) ⊗ Σc,g2 (X) as a Σg (X)-


valued 1-form,
∗  
∇(ω,Lϕ θ2 ) Jϕ∗ ψ = ϕ∗ ⊗ Jϕ∗ ∇(ω,θ2 ) ψ .

18.4. COMPACTNESS OF MODULI SPACES AND S-W INVARIANTS 697

We denote Clifford multiplication relative to g by


cg : Λ∗ (X) ⊗ Σg,c (X) −→ Σg,c (X) .
We have
cg1 ϕ∗ ⊗ Jϕ∗ (α ⊗ ψ) = cg1 (ϕ∗ α) ⊗ Jϕ∗ ψ = Jϕ∗ (cg2 (α ⊗ ψ)) .
  

Thus,
  
Jϕ∗ D(ω,θ2 ) ψ = Jϕ∗ cg2 ∇(ω,θ2 ) (ψ)

  
= cg1 ϕ∗ ⊗ Jϕ∗ ∇(ω,θ2 ) (ψ)
 ∗ 
= cg1 ∇(ω,Lϕ θ2 ) Jϕ∗ ψ .
Letting
(ω,L∗ θ2 ) ∗
Dc ϕ := cg1 ◦∇(ω,Lϕ θ2 ) ,
we then have
(ω,L∗ θ2 ) ∗ 
 
Jϕ∗ Dc(ω,θ2 ) ψ = Dc ϕ Jϕ ψ
or
(ω,L∗ θ2 )
Dc ϕ = Jϕ∗ ◦ Dc(ω,θ2 ) ◦ Jϕ∗−1 : C ∞ (Σg1 (X)) −→ C ∞ (Σg1 (X)) .
Thus,
(ω,L∗ θ2 )
Dc(ω,θ2 ) Jϕ∗−1 ψ = 0 ⇔ Dc ϕ ψ = 0.
We can also view
cg : Λ∗ (X) ⊗ Σc,g (X) −→ Σc,g (X) as cg : Λ∗ (X) −→ End(Σc,g (X)) .
Recall (see (18.33), p. 663) that there is a bilinear map
qeg : C ∞ Σ+ ∞
Σ+ −→ Ω2+g (X, C) .
 
c,g (X) × C c,g (X)
We have
 
qeg1 Jϕ∗ ψ, Jϕ∗ ζ = (ϕ∗ qeg2 )(ψ, ζ) or qeg1 (ψ, ζ) = (ϕ∗ qeg2 ) Jϕ∗−1 ψ, Jϕ∗−1 ζ .


If qg denotes the quadratic map corresponding to qeg , we have


 
Ωω+g2 − qg2 Jϕ∗−1 ψ = ϕ−1∗ η
  
⇔ ϕ∗ Ωω+g2 − qg2 Jϕ∗−1 ψ =η
 
+
⇔ (ϕ∗ Ωω ) g1 − ϕ∗ qg2 Jϕ∗−1 ψ = η
+g1
(18.106) ⇔ (ϕ∗ Ωω ) − qg1 (ψ) = η.
In summary, we have the following equivalence between the S-W equations for the
metric g2 := g1 · ϕ−1 and the ϕ-perturbed S-W equations involving spinor fields
which are associated with the fixed Spinc structure
PSpinc (g1 ) −→ PU(1) × F X(g1 )
for the metric g1 :
(18.107)
(ω,θ2 )
Dc Jϕ∗−1 ψ = 0 (ω,L∗ θ2 )
  ⇐⇒ Dc ϕ ψ = 0
Ωω+g2 − qg2 Jϕ∗−1 ψ = ϕ−1∗ η +
(ϕ∗ Ωω ) g1 − qg1 (ψ) = η
698 18. SEIBERG-WITTEN THEORY

We call the latter equations, the g1 -based S-W equations perturbed by ϕ ∈


C ∞ (Aut(T X)) and η ∈ Ω2,+g1 (X, iR). In Section 18.3, the metric g was fixed
and the Seiberg-Witten
 equations were only perturbed by η ∈ Ω2,+g (X, iR) (or
2,k 2,+g
W Λ (X, iR) . Here, we may still regard the metric g1 as being fixed and
C ∞ . However, there is an additional space of perturbations, namely C ∞ Sg+1 X ⊂


C ∞ (End(T X)) or an appropriate Sobolev enlargement, say


0 0
W 2,k Sg+1 X ⊂ W 2,k (End(T X)) ,


where k 0 is sufficiently large depending on k. One complication which arises is that


(in local coordinates) the coefficients of the Dirac operator D(ω,L∗ θ2 ) are not neces-
ϕ
0
sarily C ∞ if ϕ ∈ W 2,k Sg+1 X . Since the Levi-Civita connection for a metric may


be locally expressed (via Christoffel symbols) in terms of the metric coefficients


0
and their first derivatives, the coefficients of D(ω,L∗ θ2 ) will be in W 2,k −1 ⊆ C m
ϕ
for m < k 0 − 3. For elliptic operators with sufficiently regular coefficients, there
are versions (see [107]) of the Fundamental Elliptic Estimate (Proposition 16.23,
p. 16.23, see also Exercise 9.12, p.241), Elliptic Decomposition (Proposition 16.25,
p.499, see also Theorem 9.17, p.242) and Unique Continuation (Theorem 18.47)
with the same conclusions. We always choose k 0 large enough so that these con-
clusions hold. We mention that the Sobolev norms in these theorems are still all
defined in terms of the fixed C ∞ metric g1 , as opposed to g1 · ϕ−1 . We need to
address how the theoremsof Section 18.3 are modified when the additional pertur-
0
bation space W 2,k Sg+1 X is included.
We set
0
CWPS k := CWP k × W 2,k Sg+1 X) ,


where the “S” stands for symmetric. We have the extended map (see (18.45))
F S : CWPS k −→ W 2,k Λ2+ (X, iR) ⊕ W 2,k Σ−
 
c (X) ,

given by
 
∗ ω +g1 (ω,L∗ϕ θ2 )
F S(ω, ψ, η, ϕ) := (ϕ Ω ) − qg1 (ψ) − η, Dc ψ .

If F S(ω, ψ, η, ϕ) = 0, then the differential F S∗(ω,ψ,η,ϕ) is onto. Indeed, the proof of


Theorem 18.48 shows that even the partial derivative of F S(A, ψ, η, ϕ) with respect
to (A, ψ, η) is onto, where ϕ is held fixed. Essentially all we need to do is apply
Theorem 18.48 (p. 668) in the case where the underlying metric is g1 · ϕ−1 . Even
though g1 · ϕ−1 need not be C ∞ , the same argument goes through if k 0 is chosen
large enough. Thus,
−1
SWPS k := (F S) (0, 0)
is a Hilbert submanifold of CWPS k by Theorem 16.26 (Implicit Function Theorem
I, p. 499). The group W 2,k+2 (X, U(1))
 of gauge transformations acts trivially on the
0
extra parameter space W 2,k Sg+1 X . Consequently, Theorem 18.53 (p. 673) implies
that

MCWPS k := CWPS k /W 2,k+2 (X, U(1)) −→ CWP k /W 2,k+2 (X, U(1)) × Sk
is a Hausdorff C ∞ Hilbert manifold. Since
Jϕ∗−1 (s · ψ) = Jϕ∗−1 s−1 ψ = s−1 Jϕ∗−1 (ψ) ,

18.4. COMPACTNESS OF MODULI SPACES AND S-W INVARIANTS 699

the group W 2,k+2 (X, U(1)) leaves the space SWPS k of parametrized g1 -based so-
lutions of equations (18.107) invariant. We may then form the quotient
MSWPS k := SWPS k /W 2,k+2 (X, U(1)),
and prove that this is a closed Hilbert submanifold of MCWPS k as in Theorem
18.54, p. 679. There is a projection map
0
ps : MSWPS k −→ W 2,k Λ2+ (X, iR) × W 2,k Sg+1 X) ,
 

given by
ps([(A, ψ, η, ϕ)]) = (η, ϕ) ,
and the g1 -based moduli space SW g1 ,k (η, ϕ) at a fixed (η, ϕ) is defined by
−1
SW g1 ,k (η, ϕ) := (ps) (η, ϕ) .

Taking Regard of the Underlying Metric by Oriented Cobordism.


For the moduli spaces SW k (η) (defined in (18.67), p. 680), the underlying metric
g was fixed and hence it was not included in the notation. If we do include it, we
write SW k (η, g). Then we have

SW g1 ,k (η, ϕ) −→ SW k (ϕ−1∗ η, g1 · ϕ−1 ).
In the definition SW k (η) := p−1
k (η) (see Theorem 18.66) the right side depends on
the choice of η, but it could also depend on the chosen metric on X. Thus, in place
of (18.101) we should really write
SW([PSpinc ], η, g) := c1 (η, g)d(L)/2 [SW(η, g)],
and show that this is independent of the choice of a suitable pair (η, g), in the case
b+
2 ≥ 2. Here suitable means that η is in the residual subset of Theorem 18.66,
where the underlying metric is g. In terms of the g1 -based framework, we are to
show that
SW g1 ([PSpinc ], η, ϕ) := c1 (η, ϕ)d(L)/2 [SW g1 (η, ϕ)]
is independent of a suitable pair (η, ϕ), namely a pair for which (ϕ−1∗ η, g1 · ϕ−1 ) is
+
suitable. One necessary condition for the suitability of (η, ϕ) is that (ϕ∗ Ωω ) g1 6= η
for all connections ω, so that there will be no solutions of the S-Wg1 equations for
which ψ = 0.
 0
Definition 18.70. We call the set of such pairs in W 2,k Λ2+ (X, iR) ×W 2,k (Sg1 X)
meeting this condition FRk , since the condition guarantees that W 2,k+2 (X, U (1))
acts freely on CWPS k , and we call such (η, ϕ) an FRk pair.
As in the case where the metric is fixed, we will show that the complement of
FRk lies in a submanifold of codimension b+ 2 . Then as a consequence of the Fred-
holm Transversality Theorem below, for b+ 2 ≥ 2, two FRk pairs (η0 , ϕ0 ) and (η1 , ϕ1 )
can be joined by a path say (η(t) , ϕ(t)) of FRk pairs. We will eventually show
that generically such a path yields an oriented cobordism between SW g1 (η0 , ϕ0 )
and SW g1 (η1 , ϕ1 ). The Chern classes c1 (η, ϕ) ∈ H 2 (SW g1 ([PSpinc ], η, ϕ); Z) are
restrictions of the Chern class, say c1 ∈ H 2 (MCWPS; Z) , of the U(1) bundle
CWPS/G0 −→ MCWPS := CWPS/C ∞ (X, U (1)).
700 18. SEIBERG-WITTEN THEORY

d(L)/2
Thus, the evaluation of c1 on the two cobordant (and hence homologous) sub-
manifolds SW g1 (η0 , ϕ0 ) and SW g1 (η1 , ϕ1 ) yields the same integer. Hence,
SW g1 ([PSpinc ], η0 , ϕ0 ) := c1 (η0 , ϕ0 )d(L)/2 [SW g1 (η0 , ϕ0 )]
d(L)/2 d(L)/2
= c1 [SW g1 (η0 , ϕ0 )] = c1 [SW g1 (η1 , ϕ1 )]
= c1 (η1 , ϕ1 )d(L)/2 [SW g1 (η1 , ϕ1 )] =: SW g1 ([PSpinc ], η1 , ϕ1 ),
or equivalently,
SW([PSpinc ], ϕ−1∗
0 η, g1 · ϕ0 ) = SW([PSpinc ], ϕ−1∗
1 η, g1 · ϕ1 ).
Since any suitable pair (η0 , g0 ) can be written in the form ϕ−1∗

0 η, g1 · ϕ0 for some
suitable (η, ϕ0 ), the independence of SW([PSpinc ], η, g) on (η, g) will then be estab-
lished.
We now show that the set
 0
(η, ϕ) ∈ W 2,k Λ2+ (X, iR) × W 2,k (Pg1 (T X)) :
 
c
(18.108) (FRk ) := +
(ϕ∗ FA ) g1 = η for some A ∈ C 2,k+1 (PU (1) )
of unsuitable pairs is contained in a submanifold of codimension b+ 2 . Recall that
for any metric g on X, the g-harmonic representative Hg (Ωω ) = γg (PU(1) ) of [Ωω ] =
c
−2πi c1 (PU(1) ) ∈ 2πi H 2 (X; Z) is independent of the choice of ω. If (η, ϕ) ∈ (FRk ) ,
+
then (ϕ∗ Ωω ) g1 = η for some ω, and
  
+ +
0 = (ϕ∗ Ωω ) g1 − η ⇒ 0 = Hg2 ϕ−1∗ (ϕ∗ Ωω ) g1 − η
+
= Hg2 Ωω+g2 − ϕ−1∗ η = Hg2 (Ωω ) g2 − Hg2 ϕ−1∗ η
 

= γg2 (PU(1) )+g2 − Hg2 ϕ−1∗ η




or
Hg2 ϕ−1∗ η = γg2 (PU(1) )+g2 ∈ Hg2+

2
(X, iR).
Thus, it suffices to prove that the map
0
J : W 2,k Λ2+ (X, iR) × W 2,k (Sg1 X) −→ Hg2+

2
(X, iR), given by
−1∗

J(η, ϕ) = Hg2 ϕ η ,
has a surjective differential at (η, ϕ), since J −1 γg2 (PU (1) )+g2 will then be a sub-

c
(X, iR) = b+

manifold of codimension dim Hg2+ 2 2 containing (FRk ) . We have
 +g +g
J(η,ϕ)∗ (η 0 , ϕ0 ) = Hg2 ϕ−1∗ ϕ0∗ ϕ−1∗ (η) 2 + Hg2 ϕ−1∗ η 0 2
 ∗ +g2 +g
= Hg2 ϕ−1 ◦ ϕ0 ◦ ϕ−1 (η) + Hg2 ϕ−1∗ η 0 2 .

(X, iR), we have ϕ∗ ξ ∈ W 2,k Λ2+ (X, iR) and J(η,ϕ)∗ (ϕ∗ ξ, 0) =

Given any ξ ∈ Hg2+ 2
Hg2 ϕ−1∗ ϕ∗ ξ = Hg2 (ξ) = ξ, as required.


In order to produce an oriented cobordism between SW g1 (η0 , ϕ0 ) and SW g1 (η1 , ϕ1 ),


we need to recall some definitions and results from transversality theory.
Definition 18.71 (Transversality). Given C ∞ maps f1 : M1 → N and f2 : M2 →
N for Banach manifolds M1 , M2 and N , we say that f1 and f2 are transverse to
each other if for all (x1 , x2 ) ∈ M1 ×M2 with f1 (x1 ) = f2 (x2 ), we have f1∗ (Tx1 M1 )+
f2∗ (Tx2 M2 ) = Ty N , where y = f1 (x1 ) = f2 (x2 ). We write f1 t f2 .
18.4. COMPACTNESS OF MODULI SPACES AND S-W INVARIANTS 701

We can show that if f1 and f2 are transverse to each other, then (with ∆N :=
diag(N ))
−1
(f1 × f2 ) (∆N ) := {(x1 , x2 ) ∈ M1 × M2 : f1 (x1 ) = f2 (x2 )}
is a submanifold of M1 × M2 . Indeed, in terms of a coordinate ball ϕ : U → B about
y = f1 (x1 ) = f2 (x2 ) ∈ U (where B is a ball about 0 in some Banach space), we
have
 
−1
(f1 × f2 ) ∆ϕ−1( 1 B )
2
−1 1  −1 1 
 
(x1 , x2 ) ∈ (ϕ ◦ f1 ) 2 B ×(ϕ ◦ f2 ) 2B :
=
(ϕ ◦ f1 ) (x1 ) − (ϕ ◦ f2 ) (x2 ) = 0
−1
= (ϕ ◦ f1 − ϕ ◦ f2 ) (0) ,
where
−1 1  −1 1 
ϕ ◦ f1 − ϕ ◦ f2 : (ϕ ◦ f1 ) 2B ×(ϕ ◦ f2 ) 2B −→ B
is given by
(ϕ ◦ f1 − ϕ ◦ f2 )(x1 , x2 ) := (ϕ ◦ f1 )(x1 ) − (ϕ ◦ f2 )(x2 ) .
Since f1 and f2 are transverse, (ϕ ◦ f1 − ϕ ◦ f2 )∗ is onto at each point of the set
−1
(ϕ ◦ f1 − ϕ ◦ f2 ) (0). Thus, by the Implicit Function Theorem I (Theorem 16.26,
−1
p. 499), (f1 × f2 ) (∆N ) is a submanifold of M1 × M2 .
We also have
Theorem 18.72 (Fredholm Transversality). Let f1 : M1 → N and f2 : M2 → N
be C ∞ maps for Banach manifolds M1 , M2 and N with dim M2 < ∞ and f2
Fredholm. Then there is a map f˜2 : M2 → N arbitrarily close to f2 in the topology
of C ∞ convergence on compact sets), such that f1 t f˜2 . Moreover, if f1 t f2 for
points in a closed subset C of M2 , then we may assume that f˜2 = f2 on C.
One consequence of the FTT (Fredholm Transversality Theorem) is that any
two points, say p and q, in the complement of a codimension 2 submanifold M1
of a pathwise connected Banach manifold N can be joined by a curve which lies
outside of M1 . Indeed, let M2 := [0, 1], let f2 : [0, 1] → N be a curve joining p to
q, and let f1 : M1 → N denote the inclusion. By the FTT, f2 can be perturbed
to f˜2 : [0, 1] → N (still joining p to q), with f1 t f˜2 . It must be the case that
f˜2 ([0, 1]) ∩ M1 = ∅, since f1∗ (Tx1 M1 ) + f2∗ (Tx2 M2 ) = Ty N is impossible as M1 has
codimension 2 and dim f2∗ (Tx2 M2 ) ≤ 1. As a corollary, we have
Theorem 18.73. For b+ 2 ≥ 2, two FRk pairs (η0 , ϕ0 ) and (η1 , ϕ1 ) can be joined
by a path say (η(t) , ϕ(t)) of FRk pairs.
The Full Invariance of the S-W Invariant.
0
f1 = ps : MSWPS k −→ W 2,k+1 Λ2,+ (X, iR) × W 2,k (Sg1 X)


and
0
f2 : [0, 1] −→ W 2,k+1 Λ2,+ (X, iR) × W 2,k (Sg1 X)


is a curve joining two FRk pairs (η0 , ϕ0 ) and (η1 , ϕ1 ) which are regular values of
ps. Since (η0 , ϕ0 ) and (η1 , ϕ1 ) are regular values of ps,we have that ps t f2 at the
endpoints 0 and 1 of [0, 1]. By the FTT, f2 can be perturbed to f˜2 (still joining
702 18. SEIBERG-WITTEN THEORY

 −1
(η0 , ϕ0 ) and (η1 , ϕ1 )), so that ps t f˜2 . Then f1 × f˜2 (∆N ) is a submanifold of
MSWPS k × [0, 1]. Moreover,
 −1
C := f1 × f˜2 (∆N )
n o
= ([ω, ψ, η, ϕ], t) ∈ MSWPS k × [0, 1] : (η, ϕ) = ps([A, ψ, η, ϕ]) = f˜2 (t)
= ∪0≤t≤1 SW g1 ,k (f˜2 (t)) × {t} .
We have that
 −1     
∂ f1 × f2 ˜ (∆N ) = SW g1 ,k (f˜2 (0)) × {0} ∪ SW g1 ,k (f˜2 (1)) × {1}

= (SW g1 ,k (η0 , ϕ0 ) × {0}) ∪(SW g1 ,k (η1 , ϕ1 ) × {1}) .


We still must show that C is compact and orientable. Once this is done, we proceed
as follows. The mapping π : C → MSWPS k , given by ([ω, ψ, η, ϕ], t) 7→ [ω, ψ, η, ϕ]
exhibits the equality of homology classes (we drop the index k)

[SW g1 (η0 , ϕ0 )] = [SW g1 (η1 , ϕ1 )] ∈ H d(L) (MCWPS; Z) .


The invariance of the S-W invariant then follows:
SW g1 ·ϕ0 ([PSpinc ], η0 ) = SW g1 ([PSpinc ], η0 , ϕ0 ) = c1 (η0 , ϕ0 )d(L)/2 [SW g1 (η0 , ϕ0 )]
d(L)/2 d(L)/2
= c1 [SW g1 (η0 , ϕ0 )] = c1 [SW g1 (η1 , ϕ1 )]
d(L)/2
= c1 (η1 , ϕ1 ) [SW g1 (η1 , ϕ1 )] = SW g1 ([PSpinc ], η1 , ϕ1 )
= SW g1 ·ϕ1 ([PSpinc ], η1 ).

We now show that C is compact. The compactness results (see Corollary


18.65 and Theorem 18.66) for the individual moduli spaces depend on the estimate
(18.88), namely
(18.109) kαk2,k0 + kψk2,k0 ≤ C(k 0 ),

where the constants C(k 0 ) only depend on X, η, ω0 and k 0 ≥ 3. There, X had a


fixed Riemannian metric g and the dependence of C(k 0 ) on X implies a possible
0
dependence of C(k 0 ) on g (or equivalently, on ϕ ∈ W 2,k (Sg1 X)). However, in re-
tracing the steps involved in producing the constant C(k 0 ),  one2,k
finds as long as η
00
2,k+1 2,+
and ϕ vary within a compact subset of W Λ (X, iR) ×W (Sg1 X) for suf-
ficiently large k and k 00 , there is an upper bound on the constants C(k 0 ) for a fixed
k 0 . For example, the relevant inequalities come from applying the Fundamental El-
liptic Estimate (Proposition 16.23, p. 498) and Sobolev Multiplication (Proposition
16.24, p. 498). In all cases the constants involved depend on the Sobolev norms
(relative to g1 ) of η and ϕ of sufficiently large order. Thus, we can find constants
C(k 0 ) such that 18.109 (or a g1 -based version) holds for (η, ϕ) ∈ f˜2 ([0, 1]). The
compactness of C then follows.
We can produce an orientation on C as follows. Let z := (z1 , 1) := ([A, ψ, η, ϕ], t)
∈ C. The tangent space of C at z is
n o

∈ Tz1 (MSWPS k ) × R : (ps)∗ (v) = af˜20 (t) .

Tz C = v, a ∂t
18.4. COMPACTNESS OF MODULI SPACES AND S-W INVARIANTS 703

We wish to identify Tz C as the kernel of a Fredholm map which onto, and proceed
as in the discussion following (18.103). Let
 
+
F S(ω, ψ, η, ϕ) := (ϕ∗ Ωω ) g1 − qg1 (ψ) − η, D(ω,L∗ θ2 ) ψ ,
ϕ

so that for z1 = [ω, ψ, η, ϕ] , F S(ω, ψ, η, ϕ) = 0.


Now (ω 0 , ψ 0 , η 0 , ϕ0 ) ∈ Tz1 (MSWPS k ) means that

F S∗(ω,ψ,η,ϕ) (ω 0 , ψ 0 , η 0 , ϕ0 ) = 0, and B 1 (ω 0 , ψ 0 , η 0 , ϕ0 ) + L1 (ϕ0 ) = 0,

where
B 1 (ω 0 , ψ 0 , η 0 , ϕ0 ) := δω 0 − 41 (hψ 0 , ψi − hψ, ψ 0 i)
and L1 = L1 (ω, ψ, η, ϕ) is a zero-th order operator. We also have
 2 0 0
B (ω , ψ ) − η 0 + L2 (ϕ0 ) ,

0 0 0 0
F S∗(ω,ψ,η,ϕ) (ω , ψ , η , ϕ ) = ,
B 3 (ω 0 , ψ 0 ) + L3 (ϕ0 )
where
+
B 2 (ω 0 , ψ 0 ) : = (ϕ∗ dω 0 ) − qeg1 (ψ 0 , ψ) − qeg1 (ψ, ψ 0 )
(18.110) B 3 (ω 0 , ψ 0 ) : = D(ω,L∗ θ2 ) ψ 0 + 21 ω 0 · ψ
ϕ

and Li = Li (ω, ψ, η, ϕ) is a zero-th order operator (i = 2, 3). We have

v := (ω 0 , ψ 0 , η 0 , ϕ0 ) , a ∂t

∈ Tz C(ω 0 , ψ 0 , η 0 , ϕ0 )

 1 0 0
 B (ω , ψ ) + L1 (ϕ0 ) = 0,
⇐⇒ F S∗(ω,ψ,η,ϕ) (ω 0 , ψ 0 , η 0 , ϕ0 ) = 0, and
(η , ϕ ) = (ps)∗ (ω 0 , ψ 0 , η 0 , ϕ0 ) = af˜20 (t) ,
 0 0
 1 0 0
 B (ω , ψ ) + L1 (ϕ0 ) = 0,
⇐⇒ B 2 (ω 0 , ψ 0 ) − η 0 + L2 (ϕ0 ) = 0,
B (ω , ψ ) + L3 (ϕ0 ) = 0, and (η 0 , ϕ0 ) − af˜20 (t) = 0,
 3 0 0

H (v) := B 1 (ω 0 , ψ 0 ) + L 0
 1

 1 (ϕ ) = 0, 
H (v) := B 2 (ω 0 , ψ 0 ) − 21 η 0 + af˜22 0
(t) + L2 (ϕ0 ) = 0,
 2

⇐⇒

 H 3 (v) := B 3 (ω 0 , ψ 0 ) + L3 (ϕ0 ) = 0, and
H (v) := (η 0 , ϕ0 ) − af˜20 (t) = 0.

 4

0
Since f1 = ps and f˜2 are transverse, any ξ ∈ L2,k Λ2,+ (X, iR) ⊕ L2,k (Sg1 X) is of


the form ξ = (η 0 , ϕ0 )+f˜2∗ c ∂t ∂


for some c ∈ R and (ω 0 , ψ 0 , η 0 , ϕ0 ) ∈ Tq1 (MSWPS k ) .


Hence, the Fredholm map H = H 1 , H 2 , H 3 , H 4 is onto and its kernel is the space
Tz1 (MSWPS k ). If F S(ω, ψ, η, ϕ) = 0, then the differential F S∗(ω,ψ,η,ϕ) is onto.
For s ∈ [0, 1], define the deformation Hs by

Hs1 (v) : = B 1 (s)(ω 0 , ψ 0 ) + (1 − s) L1 (ϕ0 ) ,


 
Hs2 (v) : = B 2 (s)(ω 0 , ψ 0 ) − 12 (1 − s) η 0 + af˜22
0
(t) + (1 − s) L2 (ϕ0 ) ,
Hs3 (v) : = B 3 (s)(ω 0 , ψ 0 ) + (1 − s) L3 (ϕ0 ) , and
H 4 (v) : = (η 0 , ϕ0 ) − (1 − s) af˜0 (t) ,
s 2
704 18. SEIBERG-WITTEN THEORY

where B i (s) denotes the deformation defined in (18.104). When s = 0, we arrive


at the operator with block matrix
 1 
B (1) 0 0
 B 2 (1) 0 0 
 B (1) 0 0  ,
 3 

0 Id 0
 0
where the boldface entries are defined on L2,k Λ2,+ (X, iR) ⊕ L2,k (Sg1 (X)). The
kernel of H1 is ker B 1 (1) ⊕ B2 (1) ⊕ B 3 (1) ⊕(0, 0) ⊕ R, and the cokernel of H1 is
coker B 1 (1) ⊕ B 2 (1) ⊕ B 3 (1) ⊕(0, 0) ⊕ 0. Each of these has a natural orientation
(see the discussion following (18.104)).
APPENDIX A

Fourier Series and Integrals - Fundamental


Principles

Synopsis. Fourier Series: The Fundamental Function Spaces on S 1 ; Density; Ortho-


normal Basis; Fourier Coefficients; Plancherel’s Identity; Product and Convolution. The
Fourier Integral: Different Integral Conventions; Duality Between Local and Global —
Point and Neighborhood — Multiplication and Differentiation — Bounded and Continu-
ous; Fourier Inversion Formula; Plancherel and Poisson Summation Formulae; Parseval’s
Equality; Higher Dimensional Fourier Integrals.
I The following fundamentals and elementary facts are standard mathematical knowl-
edge today, and can be found in a great number of text books in analysis. As a general
reference, we mention [134]. J

1. Fourier Series
We use the notation S 1 := {z ∈ C : |z| = 1}, and define the function spaces
C-valued functions on S 1 with

0 1 the Banach space of continuous
C (S ) :=  1
 the norm kf k∞ := sup |f (z)| : z ∈ S ,
the Banach space of RC-valued, integrable functions on S 1 with
L1 (S 1 ) :=
 the norm kf k1 := S 1 |f | , 1
the Hilbert space of square-integrable C-valued functions
p on S
L2 (S 1 ) := R
with inner product hf, gi := S 1 f ḡ and norm kf k2 := hf, f i.
Warning: Functions in L1 (S 1 ) or L2 (S 1 ) are identified if they agree outside a set
of measure zero. In particular, f is identified with the zero function, if f is zero
almost everywhere; i.e., f is nonzero only on a set of measure zero. In this way, we
have kf k = 0 precisely when f = 0. Thus, strictly speaking, the elements of L1 (S 1 )
or L2 (S 1 ) are not functions, but rather equivalence classes of functions. While
this is true, in practice it is much simpler and generally harmless to disregard this
fine distinction, and we will do this in what follows. Moreover, it is convenient
to regard L1 (S 1 ) as L1 ([0, 1]), and L2 (S 1 ) as L2 ([0, 1]), and we will do so often
without comment. Then
Z 1 Z 1
kf k1 := |f (x)| dx and hf, gi := f (x)g(x) dx.
0 0

Perhaps f (0) 6= f (1), but this does not matter in L or L2 since {0, 1} is of measure
1

0. However, C 0 (S 1 ) and C 0 ([0, 1]) are not naturally identified.


Exercise A.1. Show that L2 (S 1 ) with h·, ·i is indeed a Hilbert space. You
have to show:
705
706 A. FOURIER SERIES AND INTEGRALS - FUNDAMENTAL PRINCIPLES

a) h·, ·i : L2 (S 1 ) × L2 (S 1 ) → C is well defined. [Hint: The pointwise estimate


2 2
2|f (z)g(z)| = 2 |f (z)| |g(z)| ≤ |f (z)| + |g(z)|
shows that f ḡ ∈ L1 (S 1 ) for f, g ∈ L2 (S 1 ).]
b) h·, ·i is sesquilinear, namely hf, gi is C-linear in f , and hf, gi = hg, f i (i.e., hf, gi
is conjugate linear in g). Also h·, ·i is positive; i.e., hf, f i ≥ 0 and hf, f i = 0 only
for f = 0. [All of this is trivial.]
c) L2 (S 1 ) is a complex vector space. [Hint: For closure under addition, prove H.
Minkowski’s inequality kf + gk2 ≤ kf k2 + kgk2 (the Triangle Inequality).]
d) L2 (S 1 ) is complete. [Hint: In order to prove that a Cauchy sequence {fn } in
L2 (S 1 ) (i.e., kfn − fm k2 → 0 as n, m → ∞) possesses a limit f ∈ L2 (S 1 ) (i.e.,
kfn − f k2 → 0 as n → ∞), one applies the fundamental convergence theorems
which distinguish the Lebesgue integral from the Riemann integral. The rather
technical proof can be found in [134, p.16-20].]
e) L2 (S 1 ) is separable. [Hint: Show that the family of piecewise constant functions,
having rational real and imaginary parts and with jumps at finitely many rational
points, is dense in L2 (S 1 ).]
Approximation. With the help of the smoothing functions of the kind
(  
−1
c exp (x − a) (x − b)−1 , for a < x < b,
(A.1) g(x) :=
0, for x ≤ a or x ≥ b,
it follows that C ∞ (S 1 ) is dense in L2 (S 1 ).
Convolution. In L1 (S 1 ), there is a commutative and associative product
(known as convolution) given by
Z 1
(f ∗ g)(x) := f (x − y)g(y) dy, x ∈ [0, 1] ,
0
where we assume that f is extended periodically of period 1 so that f (x − y) makes
sense when x − y ∈/ [0, 1]. This makes L1 (S 1 ) an algebra (without identity). By
applying the theorem of G. Fubini on iterated integrals, we obtain
kf ∗ gk1 ≤ kf k1 kgk1 .
Moreover, one can show that, relative to ∗, L2 (S 1 ) is an ideal in L1 (S 1 ), hence f ∗ g
is in L2 (S) whenever one of the factors lies in L2 (S). See [134, p.41].
Orthonormal systems. The family {z n : n ∈ Z}, where z n : S 1 → C denotes
the function that assigns to each z ∈ S 1 the value z n , is a complete orthonormal
system in L2 (S 1 ). Regarding L2 (S 1 ) as L2 ([0, 1]), the corresponding functions have
the form e2πinx .
Fourier series. This orthonormal system is complete; i.e., its linear span is
dense in L2 (S 1 ). Because of this, each function f ∈ L2 (S 1 ) can be expanded in a
Fourier series
X∞ Xk
f = fb(n)z n , i.e., lim k f − fb(n)z n k2 = 0,
n=−∞ k→∞ n=−k

with the Fourier coefficients


Z 1
(A.2) fb(n) := hf, z n i = f (x)e−2πinx dx.
0
A.2. THE FOURIER INTEGRAL 707

Note that f equals its infinite Fourier series, in the sense that the partial sums
n 2 1
P
|n|≤k f (n)z converge to f in the L (S )-norm as k → ∞, but not necessarily
b
pointwise. The function
n o∞
f 7→ fb(n) = . . . , fb(−1) , fb(0) , fb(1) , . . .
n=−∞

is an isomorphism from L2 (S 1 ) to the space L2 (Z) of absolute square-summable


sequences of complex numbers. The isomorphism is an isometry, namely
2
X∞
kf k2 = |fb(n) |2 (Plancherel’s Identity).
n=−∞

Details are in [134]. The Fourier coefficients are also defined for f ∈ L1 (S 1 ), and
by Fubini’s Theorem, it then follows [134, p.42] that

∗ g (n) = fb(n) gb(n).


f[
Incidentally, the algebra A of those sequences which appear as the Fourier coeffi-
cients of integrable functions has been barely investigated. Riemann’s theory of
trigonometric series, which were not necessarily the Fourier series of an integrable
function, led to the modern theory of generalized functions. It seems, however, that
the verdict of [134, p.43] has not lost its validity: “The best information available
to date [1972] indicates that A has no decent description at all.”
On the other hand, when the product f g (in the sense of pointwise multipli-
cation) is integrable (e.g., when f, g ∈ L2 (S), by Exercise A.1a), then the Fourier
coefficients satisfy
X∞
fcg (n) = (fb ∗ gb)(n) := fb(n − k) gb(k).
k=−∞

For the proof, we do not need Fubini’s Theorem as above, but rather we insert the
Fourier series of g in the formula for fcg (n), and then use the usual limit theorems
for the Lebesgue integral, to interchange the integral and sum.

2. The Fourier Integral


One can proceed from the standard representation of functions on a circle as
functions of period 1 on the real line, to the more general case of period T , and
then let T go to ∞. This leads to the concept of the Fourier transform. Let L1 (R)
denote the Banach space of integrable functions with
Z ∞
kf k1 := |f (x)| dx < ∞.
−∞

Different Integral Conventions. In [134, p.86f.], the Fourier transform of


f ∈ L1 (R) is defined as
Z ∞
(A.3) fDM (ξ) :=
b f (x)e−i2πξx dx,
−∞
R1
which is a natural extension of (A.2), namely fb(n) := 0 f (x)e−i2πnx dx.
In the previous edition [72, p.82] of this book, the Fourier transform was given as
Z ∞
fbB (ξ) := f (x)e−iξx dx = fbDM (ξ/2π).
−∞
708 A. FOURIER SERIES AND INTEGRALS - FUNDAMENTAL PRINCIPLES

This agrees with the main stream in analysis and information theory. In [358,
p.167] (and [61, p.423]), we find the definition
Z ∞
1 1 1
fR (ξ) := √
b f (x)e−iξx dx = √ fbB (ξ) = √ fbDM (ξ/2π).
2π −∞ 2π 2π
In Remark A.3 (below, p.710), there are cogent reasons for adopting any one of the
above definitions. Since our emphasis in this book is on the application of Fourier
transforms to differential equations and we wish the Fourier transform to be an
L2 isometry, we use fbR as Remark A.3 suggests. Thus, we define the Fourier
transform of f ∈ L1 (R) via
Z ∞
1
(A.4) f (ξ) :=
b √ f (x)e−iξx dx.
2π −∞

Since it is annoying to have to include the factor 1/ 2π, we adopt the notation
√ Z ∞
d̄x := dx/ 2π, so that f (ξ) =b f (x)e−iξx d̄x.
−∞

Convolution, Multiplication, Differentiation, and Inversion. We con-


sider the Fourier transformation on the spaces C↓∞ (R), L1 (R) and L2 (R). Here,
C↓∞ (R) denotes the space of rapidly decreasing C ∞ functions on R (with complex
values). Rapidly decreasing means that these functions and all their derivatives tend
to 0 at infinity, even when they are multiplied by arbitrary polynomials. L1 (R) is
an algebra (without identity) under convolution
Z ∞
(f ∗ g)(x) := f (x − y)g(y)dy,
−∞

and we have [134, p.87f.]:


kf ∗ gk1 ≤ kf k1 kgk1 .
L2 (R) denotes the Hilbert space (proved as in Exercise A.1) of square integrable
functions with
Z ∞ p
hf, gi := f (x)g(x) dx and kf k2 := hf, f i.
−∞

By the argument of Exercise A.1a, it follows that C↓∞ (R) is dense in L1 (R), as well
as in L2 (R). Naturally one cannot expect, as with functions on the (compact) circle
S 1 , that C ∞ (R) or C 0 (R) will be contained in the Lebesgue spaces.
A further complication arises since we no longer have L2 ⊆ L1 , or even L1 ⊆ L2 .
For example, if f (x) = 1 for |x| < 1 and f (x) = 0 for |x| ≥ 1, then
−2/3 −2/3
f (x) |x| ∈ L1 (R) \ L2 (R) and (1 − f (x)) |x| ∈ L2 (R) \ L1 (R).
Exercise A.2. The following statements are easy to prove:
(a) For each f ∈ L1 (R) and each ξ ∈ R, fb(ξ) is well defined.
(b) If f ∈ C 1 (R), f 0 ∈ L1 (R) and ixf (x) stands for the function x 7→ ixf (x) which
is assumed to be in L1 (R), then we have
(A.5) fb0 (ξ) = iξ fb(ξ) and

(A.6) (fb)0 (ξ) = −ixf


\ (x) (ξ).
A.2. THE FOURIER INTEGRAL 709

The proof of each is via integration by parts. Using the following notation for the
various operations above
1
(Df ) (x) := f 0 (x), (M f ) (x) := xf (x), and Ff := fb,
i
we have
FDf = M Ff and DFf = −FM f.
More generally, by induction, we have (for f ∈ C↓∞ (R) and p, q ∈ N)
q
M p Dq Ff = (−1) FM q Dp f.
This Differentiation-Multiplication Conversion Formula is of fundamental
importance for the treatment of differential operators (with constant coefficients)
which are converted into simple multipliers. See Chapter 8 on pseudo-differential
operators. From the topological viewpoint these formulas are most remarkable,
because they express a duality between local and global properties: Thus (A.5)
relates the smoothness of f with the rate of decay (asymptotic behavior) of fb, and
(A.6) relates the smoothness of fb with the decay of f . In fact, fb is differentiable
(a local property) when f decreases so fast that the Fourier integral of −ixf (x)
converges. This local–global duality is also a feature of the index formula for elliptic
operators, and we will deal with it further in that context.
(c) The Fourier Inversion Formula
Z ∞
1
f (x) = √ fb(ξ)eiξx dξ
2π −∞
holds for f ∈ C↓∞ (R). In direct analogy with the role of Fourier coefficients in
Fourier series, fb(ξ) is the density of the frequency ξ in the harmonic decomposition
of f .
[Hint: Prove the formula first for functions with compact support (i.e., vanishing
outside a compact subset of R). In this case there are no difficulties with the limit
process which reduces to functions of period T and then let T go to infinity. Use
smoothing functions as defined in (A.1) in order to approximate rapidly decreasing
C ∞ functions by functions of compact support. The L1 (R) estimates needed next
are somewhat tricky, but can be looked up in [134, p.89f]. A shorter direct proof
can be found in [217, p.18f].]
(d) As a corollary to the proof of (c), one obtains the Plancherel formula
kfbk2 = kf k2
and that
F : C↓∞ (R) −→ C↓∞ (R)
is linear and bijective, where F again denotes the Fourier transformation. By the
Fourier Inversion Formula, we obtain the inverse transformation
F −1 f (x) = (Ff ) (−x).


(e) Extend F from C↓∞ (R) to L2 (R)! [Hint: Approximate f ∈ L2 (R) in the
L2 (R)-norm by a sequence {fn } with fn ∈ C↓∞ (R). Using the additivity of F and
the Plancherel Identity, show that {fbn } is a Cauchy sequence in L2 (R), whence
fb := lim fbn defines an element of the Hilbert space L2 (R). Finally check that fb
indeed depends only on f and not on the choice of the sequence. In this way, one
obtains an isomorphism from L2 (R) to L2 (R), which we denote by F again.]
710 A. FOURIER SERIES AND INTEGRALS - FUNDAMENTAL PRINCIPLES

(f) The spaces C↓∞ (R) and L2 (R) share the property that they are mapped into
themselves by F. This is not true for L1 (R). Still, one can easily show [134, p.102]
that for f ∈ L1 (R),
(i) fb ∈ C 0 (R) ,
(ii) lim|ξ|→∞ fb(ξ) = 0, and

(iii) f[∗ g (ξ) = 2π fb(ξ) gb(ξ).
Remark A.3. Here we consider the merits of the three definitions
Z ∞
fbDM (ξ) : = f (x)e−i2πξx dx,
−∞
Z ∞
fbB (ξ) : = f (x)e−iξx dx = fbDM (ξ/2π), and
−∞
Z ∞ Z ∞
1 1
fbR (ξ) : = f (x)e−iξx d̄x = √ f (x)e−iξx dx = √ fbDM (ξ/2π),
−∞ 2π −∞ 2π
found in [134, p.87], [72, p.82], and [358, p.167]. The main advantage of fbB is that

there is no √ 2π or 2π. However, in terms of fB , the Plancherel formula becomes
b
kfB k2 = 2π kf k2 so that f 7→ fB is not an isometry, which is a drawback. We
b b
do have k fbDM k2 = kf k2 . Moreover, fbDM is good for expressing the remarkable
Poisson summation formula
X∞ 1 X∞
(A.7) f (kL) = fbDM (k/L) ,
k=−∞ L k=−∞
which holds for f ∈ C↓∞ (R) and any L > 0. This relates the sum of f over the
lattice {kL : k ∈ Z} to the sum of fb over the reciprocal lattice {k/L : k ∈ Z}. Using
fbR or fbB , this becomes

X∞ 2π X∞ 1 X∞
f (kL) = fbR (2πk/L) = fbB (2πk/L)
k=−∞ L k=−∞ L k=−∞
which is less aesthetic and harder to recall. The convolution theorem gives(f ∗ g)bB =
fb gb and (f ∗ g)bDM = fbDM gbDM , both of which look better than (f ∗ g)bR =
√B B
2π fbR gbR . So far, fbDM seems to be the best choice. However,
(f 0 )bDM (ξ) = 2πiξ fbDM (ξ),
and the excess baggage of the 2π makes fbDM a bit cumbersome for applications
to differential equations. Thus, we have adopted fbR , but not passionately and√not
exclusively; indeed in almost all of Chapter 4, we use fbB since the factor 1/ 2π
only serves as a needless distraction.
Higher Dimensional Fourier Integrals. By the theorem of Guido Fubini,
the closed linear span of the n-fold products
f1 (x1 ) · · · fn (xn )
of functions in L (R) is in L (Rn ) (Prove!). Thus, the preceding concepts and
2 2

results carry over directly to the case of several variables:


Definition A.4. (a) As above, define the spaces C↓∞ (Rn ), L1 (Rn ), and L2 (Rn ).
(b) Investigate the Fourier transform
Z
f (ξ) :=
b f (x)e−ihx,ξi d̄x with d̄x = (2π)−n/2 dx1 · · · dxn ,
Rn
A.2. THE FOURIER INTEGRAL 711

where x = (x1 , ..., xn ), ξ = (ξ1 , ..., ξn ), and hx, ξi = x1 ξ1 + ... + xn ξn .


(c) Investigate the convolution
L1 (Rn ) × L1 (Rn ) −→ 1
R L (R )
n

(f, g) 7→ (f ∗ g)(x) := Rn f (x − y)g(y)dy, x ∈ Rn .


We leave it to the reader to check that f ∗ g is well defined, i.e., for almost all
x ∈ Rn , the integrand y 7→ f (x − y)g(y) belongs to L1 .
Exercise A.5. Show:
(a) The Fourier Inversion Formula
Z
f (x) = fb(ξ)eihx,ξi d̄ξ, where f ∈ C↓∞ (Rn ).
Rn
(b) The Differentiation-Multiplication Conversion for multi-indices p, q,
∂ |q|
M p Dq F = (−1)|q| FDp M q , with Dq := (−1)|q| ,
∂xq11 · · · ∂xqnn
where Ff := fb for f ∈ C↓∞ (Rn ) and |q| := q1 + · · · + qn .
(c) The Integrable-Continuous Conversion
f ∈ L1 (Rn ) =⇒ fb ∈ C 0 (Rn ).

(d) The Plancherel Formula kfbk2 = kf k2 and more general Parseval’s Equal-
ity Z Z
f (x)g(x) dx = fb(ξ)b
g (ξ) dξ, for f, g ∈ L2 (Rn ).
Rn Rn

(e) The Convolution Theorems


 
fcg = (2π)n/2 fb ∗ gb and fbgb = (2π)−n/2 f[
∗g,

if f, g, fb, gb ∈ L1 (Rn ).
Details are in [134, pp.132f] or [217, pp.17–19].
APPENDIX B

Vector Bundles

Synopsis. Basic Definitions and First Examples. Homotopy Equivalence and Iso-
morphy. Clutching Construction and Suspension.

1. Basic Definitions and First Examples


Let X be a topological space. A family of vector spaces over X is a topo-
logical space E together with
(i) a continuous surjective map p : E → X and
(ii) a vector space structure of finite dimension in each Ex := p−1 (x),
which carries the topology induced by E.
By vector spaces, we mean complex vector spaces, unless explicitly indicated
otherwise. The mapping p is called the projection; E is called the total space
of the family; X is the parameter space or base space of the family; for x ∈ X,
Ex is the fiber over x. A section of a family p : E → X is a continuous map
s : X → E such that (p ◦ s)(x) = x for all x ∈ X (see Figure B.1).

s(x)

Ex

X s

Figure B.1. A section s of a family p : E → X of vector spaces


over the base space X

A homomorphism (bundle map) from one family p : E → X to another


q : F → X is a continuous map φ : E → F such that
(i) q ◦ φ = p and
712
B.1. BASIC DEFINITIONS AND FIRST EXAMPLES 713

(ii) φx : Ex → Fx is a linear map for each x ∈ X. We write φ ∈ Hom(E, F ).


We say that such a φ is an isomorphism when φ is bijective and φ−1 is continuous.
E and F are called isomorphic when there is an isomorphism between them. We
write φ ∈ Iso(E, F ) and E ∼= F.
Exercise B.1. a) Let V be a finite-dimensional vector space; e.g., V = CN .
Show that a family of vector spaces over X is obtained by taking E := X × V with
p : E → X being the projection onto the first factor. This is the product family VX
with fiber V .
b) If F is a family which is isomorphic to a product family, then one calls F trivial.
Show that a trivial family of finite dimensional (real) vector spaces is obtained, if
E := (x, y, −λy, λx) : x, y, λ ∈ R and x2 + y 2 = 1 and


p(x, y, ·, ·) := (x, y).


In general, prove that a bundle F is isomorphic to a product family VX , if and only
if one can find N sections si : X → P such that s1 (x), ..., sN (x) forms a basis for
Fx for each x ∈ X; here N = dim V .
c) Let Y be a subset of X and E a family of vector spaces over X with projection
p. Show that the induced map p−1 (Y ) → Y is a family over Y . We call this the
restriction of E to Y , and write E|Y for this family.
d) More generally: Let Y be an arbitrary topological space and f : Y → X a
continuous map. Define the induced family f ∗ (p) : f ∗ (E) → Y as follows: Take
f ∗ (E) to be the subspace of Y × E consisting of points (y, e) with f (y) = p(e); the
projection and the vector space structure on the fibres are self-evident (see Figure
B.2). Show: For each further map g : Z → Y , there is a natural isomorphism

=
(f g)∗ (E) → g ∗ f ∗ (E) which one obtains by mapping each point of the form (z, e)
with z ∈ Z and e ∈ E to the point (z, g(z), e). If f : Y → X denotes the inclusion,

=
then there is an isomorphism E|Y → f ∗ (F ) given by mapping e ∈ E|Y to the point
(p(e), e).

f f(y)

f(Y )
f *(E )y Y
Ef(y)

Figure B.2. The induced bundle (pull-back) f ∗ E of E by a map f

A family of vector spaces is called locally trivial, if each x ∈ X possesses a


neighborhood U such that E|U is trivial. A locally trivial family is called a vector
bundle; trivial families are called trivial bundles. If f : Y → X and E is a vector
714 B. VECTOR BUNDLES

bundle over X, clearly f ∗ (E) is a vector bundle over Y which we call the induced
bundle (lifted bundle, pull-back).
Note: If E is a vector bundle over X, then x 7→ dim(Ex ) is a locally constant
function on X, and hence constant on each connected component of X. If dim(Ex )
is constant on all of X, then one says that E has (fiber-)dimension equal to the
common fiber dimension dim(Ex ). If X is a manifold, then the real dimension of
E (regarded as a topological space) equals dim(X) + 2 dim(Ex ). Vector bundles of
fiber dimension 1 are also called line bundles.
Since a vector bundle is locally trivial, each section can be written locally as
a vector-valued function on the base space. For a vector bundle E, we denote the
space of sections of E by C 0 (E); C 0 (E) is a vector space in a natural way via
pointwise addition, etc.
Exercise B.2. a) Let V be a (complex) vector space and PV its associated
projective space of all one-dimensional linear subspaces of V . We can write PV =
(V \ {0})/ ∼, where ∼ denotes the equivalence relation v ∼ w ⇔ λv = w for some
λ ∈ C. We define HV ⊆ PV × V as the set of all (x, v) such that x ∈ PV , v ∈ V ,
and v belongs to the complex line x. Show that HV is a vector bundle in a natural
way. (The construction goes back to Heinz Hopf.)
b) Go through the corresponding construction in the (more intuitive) category of
real vector bundles when V is real (e.g., Rn ), and show that HV is a real subbundle
of PV × V of fiber dimension 1, and that HV is nontrivial if dim V ≥ 2.
c) For the moment, we remain in the real category and consider the following family
(parametrized by θ ∈ [0, 2π]) of integro-differential equations for C ∞ functions f
on the unit interval which satisfy the boundary condition f (0) = f (1) :
Z 1
df
cos θ f (x) + sin θ = cos θ f (x) dx, θ ∈ [0, π] ,
dx 0
Z 1 Z 1
cos θ f (x) + sin θ Lf (x)) = cos θ f (x) dx + sin θ Lf (x) dx, θ ∈ [π, 2π] .
0 0
Here, L : C 0 (S 1 ) → C 0 (S 1 ) is a fixed operator with L2 = − Id. Show that the
solutions of the family of equations form a real vector bundle over the circle S 1 :=
R/2πZ that is nontrivial and isomorphic to the bundle HR2 . (See Figure B.3.
Actually, every real line bundle over S 1 is either trivial or isomorphic to HR2 , see
also [97, 3.23.9].)

Figure B.3. The simplest nontrivial bundle


B.1. BASIC DEFINITIONS AND FIRST EXAMPLES 715

[Hint for b): In contrast to the complex numbers, −1 cannot be deformed into 1
without going through 0. Thus, a real bundle is nontrivial, if it remains connected
after the zero section is removed.
For c): First show that for 0 ≤ θ ≤ π, the solutions are the constant functions c1,
and for π ≤ θ ≤ 2π the solutions are the functions c cos θ 1−sin θ L(1) , c ∈ R. With
the topology of S × C 0 (I) or of S 1 × R (since every solution can be written in the
form c1 + c2 L(1)), construct a family of 1-dimensional (real) vector spaces over S 1
and show the local triviality. For this use the initial value map f 7→ f (0). Note that
this map can vanish for a solution f 6= 0 at a parameter θ0 , namely if cos θ0 = sin θ0
L(1) (0). How can one proceed in a neighborhood of θ0 ? Distinguish the cases where
L(1) (0) is positive, negative or zero. Incidentally, howPdoes one obtain an L with

L2 = − Id? Start with the Fourier series f (x) = a0 + ν=1 (aν sin νx + bν cos νx),
and replace aν by bν+1 and bν by −aν−1 ; see also [395].]

Remark B.3. a) and b) describe the origin of the bundle concept in analytic
and projective geometry. Part c) is characteristic for many functional analytic
situations with jumps, where the passage from one side to another (from one solution
curve to another of the same equation) cannot be understood within the given space
but requires an extension of the system (e.g., by parametrization). A basic model
for such a process is present in the geometry of number fields (see Hint for b)). Many
classical results of analysis – especially concerning the dependence of the solutions
of a functional equation on the variation of its coefficients and on the zeros and
poles of its solutions – can be aptly formulated in the language of vector bundles.
Conversely, the theorem of Grothendieck, for example, that every holomorphic
vector bundle E on the Riemannian number sphere S 2 = P(C2 ) can be represented
as a Whitney sum E1 ⊕ · · · ⊕ En of line bundles (Am. J. Math. 79 (1957), 121-138)
was known to analysts at the beginning of the century: See G. Birkhoff, Math.
Ann. 54 (1913), 122-139, where Grothendieck’s theorem appears as a theorem
about matrices of analytic functions. Birkhoff was led to this theorem through
his investigation of the singular points of ordinary differential equations; further see
D. Hilbert, Gött. Nachr. (1905), 307-358, who gave a proof of Grothendieck’s
theorem for N = 2 in his Fundamentals of a General Theory of Integral Equations
in connection with the Riemannian Problem (Contributed by M. Schneider).

Exercise B.4. Show that the usual operations for vector spaces in linear alge-
bra also make sense for vector bundles. In particular, for vector bundles E and F
over the same base, investigate the direct sum E ⊕F , the tensor product E ⊗F ,
the homomorphism bundle Hom(E, F ), the isomorphism bundle Iso(E, F ),
and the dual bundle E ∗ := Hom(E, CX ). Show that the bundles E ∗ ⊗ F and
Hom(E, F ) are isomorphic. Also, carry over the concepts of subspace and quotient
space from linear algebra to the corresponding concepts of a subbundle F of E
and a quotient bundle E/F .
[Hint: Make use of the fact that the corresponding operations in the structure
group GL(N, C) are continuous! Example: recall E ∗ = Hom(E, CX ) and intro-
duce a topology on E ∗ |U which makes U × CN → E ∗ |U a homeomorphism, where
φ : E|U → U × CN is a local trivialization for E over the open subset U ⊆ X. Let
ψ : E|V → V × CN be another trivialization. Do φ and ψ define the same topology
on E ∗ |U ∩V ? Does the continuity of (ψ ◦ φ−1 )∗ : (U ∩ V ) × CN → (U ∩ V ) × CN
follow from that of (U ∩ V ) × CN → (U ∩ V ) × CN ? For this, write the two chart
716 B. VECTOR BUNDLES


changes in the form U ∩ V → GL(N, C) and prove (!) that GL(N, C) → GL(N, C)
is continuous.]
Remark B.5. There is also the powerful concept of an outer tensor product
E  F as a vector bundle on X × Y , if E is a vector bundle on X and F is a vector
bundle on Y , see our Sections 9.3 and 12.2, pp.242ff and pp.287ff.

2. Homotopy Equivalence and Isomorphy


Definition B.6. We denote the set of isomorphism classes of vector bundles
over X by Vect(X), and let VectN (X) denote the subset of Vect(X) consisting of
the classes of bundles of fiber dimension N .
Note that Vect(X) is an abelian semi-group under the operation ⊕. In Vect(X),
there is a naturally distinguished element, namely the class of the trivial bundle of
dimension N . A vector bundle over a point is a vector space, and hence Vect(X) can
be identified with the semigroup Z+ of nonnegative integers, in this case. However,
in the general ease, when there are nontrivial bundles (see Exercise B.2 above), the
isomorphism classes of vector bundles are not determined by their dimensions.
Two continuous mappings f, h : X → Y are homotopic, if there is a continuous
map F : X × I → Y (I := [0, 1]) such that F0 := F (·, 0) = f and F1 := F (·, 1) = h.
The map f : X → Y is a homotopy-equivalence, if there is a continuous map
g : Y → X such that g ◦ f ∼ IdX and f ◦ g ∼ IdY (“∼” means homotopic). X
and Y are then called homotopy equivalent. The set of homotopy classes of
maps X → Y is denoted by [X, Y ]. X is called contractible, if X is homotopy
equivalent to a point.
Theorem B.7. (i) If f : X → Y is a homotopy equivalence, then the transfor-
mation f ∗ : Vect(Y ) → Vect(X) of Exercise B.1d is bijective. (Assume X and Y
compact.)
(ii) If X is contractible, then every bundle over X is trivial, and Vect(X) is iso-
morphic to the nonnegative integers.
Proof. (ii) follows easily from (i), since VectN (P ) consists only of the isomor-
phism class of the trivial bundle of dimension N , in the case where P is a point.
(i) follows from the fact that F0∗ E ∼ = F1∗ E, if F : X × I → Y is a homotopy of f
and E is a vector bundle over Y . We give a proof of this in three steps:
Step 1: Let H be a vector bundle over X × I, fix a τ ∈ I, and let s ∈
C 0 (H|X×{τ } ). We show that s can be extended to a section S ∈ C 0 (H) with
S(·, τ ) = s. Since a section of a vector bundle can be regarded locally as a graph
of a continuous vector-valued function, one can locally apply the Tietze Extension
Theorem [130]: We can find a neighborhood U about each (x, τ ), and a section
t ∈ C 0 (H|U ) such that t and s coincide on (X × {τ }) ∩ U (see Figure B.4).
By the compactness of X, we can obtain a finite system {Uj }, {tj } with X ×
{τ } ⊂ ∪Uj . If {φj } is a C 0 partition of unity subordinate to {Uj } (see also Theorem
6.4, p. 159), then we set

tj φj , on Uj ,
sj :=
0, on (X × I) \ Uj .
0
sj ∈ C 0 (H) is a well-defined extension of s.
P
By construction sj ∈ C (H); hence,
Step 2: From Step 1, we conclude that two vector bundles G and G0 over X × I
which are isomorphic over X × {τ }, are also isomorphic in a neighborhood of
B.2. HOMOTOPY EQUIVALENCE AND ISOMORPHY 717

(x,¿) x
U

X
X£f¿g

X£ I

Figure B.4. Extending a section s locally over a collar segment U

X ×{τ }: Each s ∈ Iso(G|X×{τ } , G0 |X×{τ } ) can be regarded as a section of H|X×{τ } ,


where H denotes the bundle Hom(G, G0 ) of linear maps from fibres of G to cor-
responding fibres of G0 . Let S ∈ C 0 (H) be an extension of s. Then the set
W := {z ∈ X × I : Sz ∈ Hom(G, G0 ) is bijective} is open in X × I (by the classical
zero determinant argument) and contains all of X × {τ } by construction. Since
the inverse map of GL(N, C) is continuous, it follows that the mapping z → Sz−1 is
continuous, and hence a bundle isomorphism is defined on W .
Step 3. We now set G := F ∗ E and G0 := p∗ (Fτ )∗ E, where Fτ (x) := F (x, τ ) and
p : X × I → X denotes the projection.

I
p
X X£I X

F F¿
Y
Figure B.5. Why the isomorphism classes of (Fτ )∗ E do not de-
pend on τ

By Exercise B.1d, G and G0 are isomorphic over X × {τ }, and by Step 2,


they are also isomorphic in a whole neighborhood, which we can take to be a strip
X×δ(τ ), by the compactness of X. For all ρ ∈ δ(τ ), we then have (Fρ )∗ E ∼
= (Fτ )∗ E.
Since the unit interval I is compact and connected, we obtain that the isomorphism
classes of (Fτ )∗ E do not depend on τ (see Figure B.5). 
Remark B.8. The statements proved in the first two steps of the proof also
apply to more general situations and are occasionally formulated as independent
718 B. VECTOR BUNDLES

theorems; e.g., see [28, p.233f] or [17, p.16]. One can also express the result of our
proof somewhat more generally (we write X × Z instead of X × I): Each vector
bundle E over the topological space X × Z (X and Z compact) can be regarded
as a continuous family of vector bundles E over X, where the parameter z is in Z,
and the isomorphism classes of E in Vect(X) are locally constant.

3. Clutching Construction and Suspension


Vector bundles are often given via a clutching or gluing construction: Let
X = X1 ∪ X2 and A = X1 ∩ X2 , where all spaces are compact. Let Ei be a vector
bundle over Xi and φ : E1 |A → E2 |A an isomorphism. Then we define the vector
bundle E1 ∪φ E2 over X as follows. As a topological space E1 ∪φ E2 is the quotient
space of the disjoint union E1 + E2 under the equivalence relation which identifies
e1 ∈ E1 |A with φ(e1 ) ∈ E2 |A . If we regard X as the corresponding quotient space
of X1 + X2 , then we obtain a natural projection p : E1 ∪φ E2 → X, and p−1 (x) has
a natural vector space structure for each x ∈ X.
Exercise B.9. a) Show that E1 ∪φ E2 is a vector bundle.
b) Show that the isomorphism class of E1 ∪φ E2 depends solely on the homotopy
class of the isomorphism φ : E1 |A → E2 |A .
[Hint for a): It remains only to show that E1 ∪φ E2 is locally trivial. Outside of A,
this is clear. In order to extend a trivialization of E1 on a neighborhood V1 ⊆ X of
a point a ∈ A to a trivialization of E2 over V2 with a ∈ V2 ⊆ X2 , argue as in Step 2
of Theorem B.7. See also [28, p.235] or [17, p.21]. For b): Reduce to Theorem B.7.
See also the given sources.]
We now give VectN (X) a homotopy-theoretic interpretation, when X can be
represented as the suspension S(Y ) of another space Y .1 Here the suspension
S(Y ) denotes the union of two cones over Y . Thus, we write S(Y ) := C + (Y ) ∪
C − (Y ), where C + (Y ) := Y × [0, 12 ]/Y × {0} and C − (Y ) := Y × [ 12 , 1]/Y × {1} (see
Figure B.6). Then Y = C + (Y ) ∩ C − (Y ). We note that the suspension S(S n ) of
the n-sphere is homeomorphic to the (n + 1)-sphere S n+1 .
Theorem B.10. The clutching of trivial bundles over C + (Y ) and C − (Y ) de-

=
fines a natural isomorphism [Y, GL(N, C)] → VectN (S(Y )).
Proof. (i) Each l : Y → GL(N, C) yields a bundle over SY via the clutching
of the N -dimensional trivial bundles over the two cones, and homotopic maps l0
and l1 yield isomorphic bundles; see the proof of Theorem B.7(i). (ii) Conversely,
we have the composition
VectN (SY ) −→ VectN (C − Y ) ⊕ VectN (C + Y ) −→ [Y, GL(N, C)].
The left arrow is given by the restrictions of the bundle, where one obtains triv-
ial bundles since C ± Y are contractible (see Theorem B.7(ii)). If α± are such
trivializations, then the right arrow is defined by taking the homotopy class of
−1
(α+ |Y )(α− |Y ) : Y → GL(N, C), which actually only depends on the homotopy
±
classes of α and hence only on the isomorphism class in VectN (SY ). (iii) By
construction the functions given in (i) and (ii) are inverse to each other. 

1With the concept of Grassmann manifolds, one can give a homotopy-theoretic definition of
VectN (X) for arbitrary X; see [17, pp.24–30].
B.3. CLUTCHING CONSTRUCTION AND SUSPENSION 719

C +(Y )
S(Y )
C {(Y )

Figure B.6. The suspension SY of Y

Exercise B.11. Show HC2 ∼ = CB0 ∪a CB∞ , where HC2 denotes the complex line
bundle over PC2 = C ∪ {∞} = S 2 = B0 ∪ B∞ defined in Exercise B.2a; here (z0 , z1 )
are homogeneous coordinates for PC2 with (0, 1) = ∞, and z = z1 /z0 denotes the
coordinate for C, B0 := {z ∈ C : |z| ≤ 1}, and B∞ := {z ∈ C : |z| ≥ 1} ∪ {∞} (the
two canonical hemispheres of S 2 ). Finally, a : z 7→ z, denotes the standard map
S 1 → C× := {C \ {0}} = GL(1, C).
Exercise B.12. Show that for each bundle E, there is a bundle F such that
E ⊕F is trivial. [Hint: Show first with the help of a finite open cover of the compact
parameter space X and a suitable partition of unity that C 0 (E) contains an ample
subspace; i.e., a subspace V ⊆ C 0 (E) such that each point of E is in the image of
a section s ∈ V . If dim V = N , then we have an epimorphism φ : X × CN → E,
and consequently there is an isomorphism E ⊕ F → X × CN , where F is the
kernel bundle of φ; see also [17, p.26 f] or the related technique in the proof of our
Embedding Theorem 6.7c, p.164.]
Exercise B.13. Let X be a topological space which in addition possesses the
structure of a C ∞ manifold of dimension n (see Chapter 6).
a) Show that the tangent bundle T X is a (real) vector bundle over X; do the same
for the normal bundle N X, when X is a submanifold of a Riemannian manifold Y .
b) When may one call a continuous vector bundle over X, whose total space is a
C ∞ manifold, a C ∞ vector bundle? Show that each (continuous) vector bundle
over X is isomorphic to a C ∞ vector bundle.
[Hint for a): First investigate the case X = S 1 and show that T S 1 is isomorphic
to the real line bundle defined in Exercise B.1b. In the general case, this direct
method is also possible, since one can (see Theorem 6.7c, p. 162) embed X into a
high dimensional Euclidean space, and thus realize T X as a (real) subbundle of a
higher dimensional trivial bundle. It is easier (especially since T X is in general not
trivial; e.g., for X = S 2 , see Exercise 10.28, p. 274) to carry out a local analysis,
where each chart u from the C ∞ atlas yields a local trivialization of T X via the
differential forms (du1 , ..., dun ); see Chapter 6. See also the discussion below.
For b): For the definition of C ∞ vector bundles, see also [97, Ch.3], and for the
720 B. VECTOR BUNDLES

topological equivalence of the categories of C 0 and C ∞ vector bundles, see the


Whitney Approximation Theorem in [97, p.66]. Details of the argument are in
[206, p.101], where it is shown that one can make E itself into a C ∞ vector bundle.]
Remark B.14. From Exercise B.13b, it follows (without loss of generality) that
we only need to investigate C ∞ vector bundles, if the base is C ∞ . Exercise B.12
does not say that we always encounter trivial bundles (compare with the analogous
- but deeper embedding theorem for manifolds, Theorem 6.7c, p. 162): Roughly, the
carrying along of additional irrelevant parameters is not only redundant, but can
also produce so much noise that this noise destroys the structure of the problem
or renders it unrecognizable. This is the case with the index problem for elliptic
operators, whose solution consists exactly in distinguishing certain vector bundles
generated by the symbol of the operator; see Part III.
The idea of a vector bundle originates in the analysis of non-Euclidean mani-
folds. Otherwise, according to Theorem B.7(ii), there are only trivial bundles, since
Euclidean space is contractible. As an example, let c : I → X be a differentiable
path in a manifold X. Classical mechanics considers the velocity vector ċ(t) for
t ∈ I. For the physicist it was always clear (by reasons of physics) how c(t) is
multiplied by a scalar, how ċ1 (t1 ) and ċ2 (t2 ) are added or how the equality of ċ1 (t1 )
and ċ2 (t2 ) is checked when c1 , c2 : I → X are two paths in X with c1 (t1 ) = c2 (t2 ).
From the point of view of physics no confusion between a velocity vector and a
position vector was conceivable, but the earliest mathematical abstractions could
not express the difference: If the position vectors c(t) are represented as triples of
real numbers (c1 (t), c2 (t), c3 (t)), then the velocity vector is (ċ1 (t), ċ2 (t), ċ3 (t)) where
ċi (t) denotes the derivative of ċi at t. Thus, c(t) and ċ(t) are both elements of the
single vector space R3 . This low level of abstraction was fully sufficient as long as
X was Euclidean space (actually an affine space – but choose a base point and call
it 0). In this case there is indeed a natural interpretation of the velocity vector
ċ(t) at the space point c(t) as a velocity vector e c˙ (t) at the point ec(t) = 0. In fact,
consider the translated path e c : I → X with e c(τ ) := c(τ ) − c(t), τ ∈ I. This means
that the tangent spaces at the various points of X can be identified canonically (via
the retraction r : X → {0}) with the tangent space at the point 0. In the language
of vector bundles, we could say T X ∼ = r∗ (T X|{0} ). Here T X denotes the totality of
all velocity vectors, the fiber Tx X is R3 , the restricted bundle T X|{0} is precisely
the space {0} × R3 ∼ = R3 , and the induced bundle r∗ ({0} × R3 ) is the trivial bundle
3
X ×R .
Let us next consider the case that X is a submanifold of a Euclidean space
Y , say the 2-sphere in 3-space. Every velocity vector ċ(t) can be considered as an
element of the tangent space of Y at the point c(t) and hence an element of the
tangent space of Y at 0 (due to the translation described above), i.e., as an n-tuple
of derivatives of the coordinate functions of c. Already by reason of dimensions it
is clear that, in general, the tangent space Tx X of all velocity vectors X of paths
through x cannot be identified with the full tangent space T0 Y but only with some
subspace, and it may be different for each x.
Example B.15. X = S 1 and Y = R2 . If we follow the translation with a
rotation through the angle ϕ, then we have identified all tangent spaces Tx S 1 with
the same subspace of T0 R2 , namely with the tangent space T(1,0) S 1 (see Figure
B.7). We write T S 1 = q ∗ (T S 1 |{(1,0)} ), where q : S 1 → (1, 0) denotes the retraction
B.3. CLUTCHING CONSTRUCTION AND SUSPENSION 721

x = e i' = (cos ', sin ')

'
({1,0) (1,0)

Tx (S 1)

Figure B.7. Identifying Tx S 1 with T(1,0) S 1

and T S 1 |{(1,0)} can be identified with {(1, 0)} × R. Consequently T S 1 ∼= S 1 × R,


1
i.e., the tangent bundle of S is trivial.
In analytic terms this circumstance is expressed by saying that there exists a
nowhere vanishing tangential vector field on S 1 ; i.e., at each point a nonvanishing
velocity vector can be chosen and in a continuous fashion. For example, choose
at x = (cos θ, sin θ) the unit velocity vector ċ(φ/2π), where c : I → X is given by
d d
c(t) = (cos 2πt, sin 2πt). Usually we write dθ x
instead of ċ(φ/2π). Then dθ is the
1 ∼ 1
nowhere vanishing vector field and it defines an isomorphism T S = S × R. The
isomorphism is, for all x ∈ S 1 , given by the map Tx S 1 → R which assigns to a
d
velocity vector the value λ, if it is equal to λ · dθ x
.
Example B.16. X = S 2 and Y = R3 . Here the situation is different. There
is no canonical way of identifying all tangent spaces. In other words, the tangent
bundle T S 2 is nontrivial. If not there would exist two tangential vector fields on
S 2 which are linearly independent at each point of S 2 . But on S 2 there is not even
a single vector field that vanishes nowhere (Exercise 10.28, p. 274). This example
shows that it might be useful (indeed necessary) to distinguish tangent spaces at
different points. As one goes on to consider manifolds without an explicitly given
imbedding in Euclidean spaces, as in physics with the theory of relativity, the notion
of a bundle becomes indispensable.
I The modern concept of a bundle evolved from the topology and geometry of mani-
folds as practiced by Heinz Hopf and others since the 1920’s. In the 1950’s the notion was
precisely formulated, and the classification of bundles and their systematic employment in
deep problems of geometry and analysis started. For the theory of characteristic classes,
see e.g., the concise [170] or the more elaborated [207, p.49f], [249, Chapter XII], and
our Sections 13.5 and 15.7. Crudely expressed, the success of these methods derives from
their utilizing given or manufactured classical structures (such as the tangent bundles or
bundles of differential forms) on the manifolds under discussion to the greatest extent
possible and thus shifting the plane of study from manifolds — which are conceptually
simpler but harder to understand — to vector bundles which are more easily analyzed.
Put differently: For topological manifolds and other triangulable spaces, one depends
at first on the combinatorial methods of the analysis of cell decompositions. For differen-
tiable manifolds only the group of diffeomorphisms is a priori available for investigation.
Vector bundles offer more opportunity for manipulations because of their richer structure.
Principally, large parts of linear algebra can be used directly. Incidentally, these play
an important role also for the other approaches, albeit under the surface. For example
(see above Exercise B.4) one can perform linear constructions with vector bundles such as
722 B. VECTOR BUNDLES

forming direct sums and quotients, which is impossible to do with manifolds. Clutching
functions which can be used to build complicated manifolds from simpler ones (see e.g.,
Exercise 6.50, p. 191) are typically diffeomorphisms that become linear only in the first
derivative (functional matrix or Jacobian). In contrast, Theorem B.7 shows how much (via
linear clutching functions) the topology of vector bundles can be reduced to the geometry
of the matrix spaces of linear algebra.
In the early 1960’s linear algebra had matured enough with the Periodicity Theorem
(for the (stable) homotopy groups of invertible matrices) discovered by Raoul Bott just
before. Michael Francis Atiyah and Friedrich Hirzebruch extracted from these
methods an abstract formalism – K-theory – which they developed as a generalized co-
homology theory, using stability classes of vector bundles; see Chapter 10 and our view
upon characteristic classes and vector fields in Section 13.5 (pp.325ff). J
Bibliography

[1] Manifold Atlas Project, [Link] ongoing, The mathematicians


behind the project with M. Kreck as Managing Editor, focus on constructions and invari-
ants, general theory and open problems. They also plan to build up historical information.
[2] N. H. Abel, ‘Mémoire sur une propriété générale d’une classe très étendue de fonctions
transcendantes’. Mémoires présentés par divers savants à l’Académie Royale des Sciences de
l’Institut de France VII (1841), 176–264, reprinted in [3, I, pp.145–211], English translation
in [4].
[3] — Complete works of Niels Henrik Abel. Vol. 1 and 2. Edited by L. Sylow and S. Lie.
(Œuvres complètes de Niels Henrik Abel. Vol. 1 and 2.) Reprint of the new edition published
1881 by Grøndahl and Son.. Cambridge: Cambridge University Press, 2012 (French).
[4] — Abel on analysis. Kendrick Press, Heber City, UT, 2007, Papers of N. H. Abel on abelian
and elliptic functions and the theory of series, Translated from the French, with commentary
and notes, by Philip Horowitz.
[5] R. Abraham and J. E. Marsden, Foundations of mechanics. Benjamin/Cummings Pub-
lishing Co. Inc. Advanced Book Program, Reading, Mass., 1978, Second edition, revised and
enlarged, With the assistance of Tudor Raţiu and Richard Cushman.
[6] R. A. Adams, Sobolev spaces. Academic Press [A subsidiary of Harcourt Brace Jovanovich,
Publishers], New York-London, 1975, Pure and Applied Mathematics, Vol. 65.
[7] M. S. Agranovič, ‘Elliptic singular integro-differential operators’. Uspehi Mat. Nauk 20/5
(125) (1965), 3–120.
[8] M. S. Agranovič and A. S. Dynin, ‘General boundary value problems for elliptic systems in
an n-dimensional domain’. Dokl. Akad. Nauk. SSSR 146 (1962), 511–514, Russian, English
translation Soviet Math. Dokl. 3 (1962/63), 1323–1327.
[9] L. V. Ahlfors, Complex analysis. An introduction to the theory of analytic functions of
one complex variable. McGraw-Hill Book Company, Inc., New York-Toronto-London, 1953.
[10] N. I. Akhiezer and I. M. Glazman, Theory of linear operators in Hilbert space. Dover
Publications Inc., New York, 1993, Translated from the Russian and with a preface by
Merlynd Nestell, Reprint of the 1961 and 1963 translations, Two volumes bound as one.
[11] P. Alexandroff and H. Hopf, Topologie. I. Springer-Verlag, Berlin, 1974, Berichtigter
Reprint, Die Grundlehren der mathematischen Wissenschaften, Band 45.
[12] A. Alonso and B. Simon, ‘The Birman-Kreı̆n-Vishik theory of selfadjoint extensions of
semibounded operators’. J. Operator Theory 4/2 (1980), 251–270, addenda in J. Operator
Theory 6 (1981), 407.
[13] L. Alvarez-Gaumé and M. A. Vázquez-Mozo, An invitation to quantum field theory, vol.
839. Springer-Verlag, Heidelberg, 2012.
[14] N. Aronszajn, ‘A unique continuation theorem for solutions of elliptic partial differential
equations or inequalities of second order’. J. Math. Pures Appl. (9) 36 (1957), 235–249.
[15] M. F. Atiyah, Harmonic spinors and elliptic operators, Arbeitstagung Lecture, Bonn,
mimeographed, Notes taken by S. Lang, 16 July, 1962.
[16] — ‘Algebraic topology and elliptic operators’. Comm. Pure Appl. Math. 20 (1967), 237–249,
reprinted in [26, Vol. 3, pp.57–71].
[17] — K-theory, Lecture notes by D. W. Anderson. W. A. Benjamin, Inc., New York-
Amsterdam, 1967.
[18] — ‘Bott periodicity and the index of elliptic operators’. Quart. J. Math. Oxford Ser. (2)
19 (1968), 113–140, reprinted in [26, Vol. 2, pp.605–633].

723
724 BIBLIOGRAPHY

[19] — ‘Global aspects of the theory of elliptic differential operators’. In: Proc. Internat. Congr.
Math. (Moscow, 1966). Izdat. “Mir”, Moscow, 1968, pp. 57–64, reprinted in [26, Vol. 3,
pp.73–81].
[20] — ‘Algebraic topology and operators in Hilbert space’. In: Lectures in Modern Analysis
and Applications. I. Springer, Berlin, 1969, pp. 101–121, reprinted in [26, Vol. 2/48] and
free available at [Link]
[21] — ‘Topology of elliptic operators’. In: Global Analysis (Proc. Sympos. Pure Math., Vol.
XVI, Berkeley, Calif., 1968). Amer. Math. Soc., Providence, R.I., 1970, pp. 101–119,
reprinted in [26, Vol. 3, pp.387–407].
[22] — Vector fields on manifolds, Arbeitsgemeinschaft für Forschung des Landes Nordrhein-
Westfalen, Heft 200. Westdeutscher Verlag, Köln, 1970, reprinted in [26, Vol. 2, pp.727–749].
[23] — Elliptic operators and compact groups, Lecture Notes in Mathematics, Vol. 401. Springer-
Verlag, Berlin, 1974, reprinted in [26, Vol. 3, pp.499–593].
[24] — ‘Classical groups and classical differential operators on manifolds’. In: Differential oper-
ators on manifolds (Centro Internaz. Mat. Estivo (C.I.M.E.), III Ciclo, Varenna, 1975).
Cremonese, Rome, 1975, pp. 5–48, reprinted in [26, Vol. 4, pp.341–385].
[25] — ‘Trends in pure mathematics.’. In: Proc. of the 3rd Internat. Congress on Mathematical
Education (Karlsruhe), 1976, pp. 61–74, reprinted in [26, Vol. 1, pp.261–276].
[26] — Collected works. Vol. 1-5, Oxford Science Publications. The Clarendon Press Oxford
University Press, New York, 1988.
[27] M. F. Atiyah and R. Bott, ‘The index problem for manifolds with boundary’. In: Dif-
ferential Analysis, Bombay Colloq., 1964. Oxford Univ. Press, London, 1964, pp. 175–186,
reprinted in [26, Vol. 3, pp.25–37].
[28] — ‘On the periodicity theorem for complex vector bundles’. Acta Math. 112 (1964), 229–247,
reprinted in [26, Vol. 2, pp.337–357].
[29] — Notes on the Lefschetz fixed point theorem for elliptic complexes (Russ. transl. Moscow
1966), Notes, Harvard University, mimeographed, 1965.
[30] — ‘A Lefschetz fixed point formula for elliptic differential operators’. Bull. Amer. Math.
Soc. 72 (1966), 245–250, reprinted in [26, Vol. 3, pp.83–89].
[31] — ‘A Lefschetz fixed point formula for elliptic complexes. I’. Ann. of Math. (2) 86 (1967),
374–407, reprinted in [26, Vol. 3, pp.91–125].
[32] — ‘A Lefschetz fixed point formula for elliptic complexes. II. Applications’. Ann. of Math.
(2) 88 (1968), 451–491, reprinted in [26, Vol. 3, pp.127–169].
[33] M. F. Atiyah, R. Bott and V. K. Patodi, ‘On the heat equation and the index theorem’.
Invent. Math. 19 (1973), 279–330, reprinted in [26, Vol. 4, pp.11–63].
[34] — ‘Errata to: “On the heat equation and the index theorem” (Invent. Math. 19 (1973),
279–330)’. Invent. Math. 28 (1975), 277–280, reprinted in [26, Vol. 4, pp.65–70].
[35] M. F. Atiyah, R. Bott and A. Shapiro, ‘Clifford modules’. Topology 3/suppl. 1 (1964),
3–38, reprinted in [26, Vol. 2, pp.299–336].
[36] M. F. Atiyah, V. G. Drinfeld, N. J. Hitchin and Y. I. Manin, ‘Construction of instan-
tons’. Phys. Lett. A 65/3 (1978), 185–187, reprinted in [26, Vol. 5, pp.21–25].
[37] M. F. Atiyah and J. L. Dupont, ‘Vector fields with finite singularities’. Acta Math. 128
(1972), 1–40, reprinted in [26, Vol. 2, pp.761–802].
[38] M. F. Atiyah, N. J. Hitchin and I. M. Singer, ‘Deformations of instantons’. Proc. Nat.
Acad. Sci. U.S.A. 74/7 (1977), 2662–2663, reprinted in [26, Vol. 5, pp.7–9].
[39] — ‘Self-duality in four-dimensional Riemannian geometry’. Proc. Roy. Soc. London Ser. A
362/1711 (1978), 425–461, reprinted in [26, Vol. 5, pp.27–65].
[40] M. F. Atiyah, V. K. Patodi and I. M. Singer, ‘Spectral asymmetry and Riemannian
geometry’. Bull. London Math. Soc. 5 (1973), 229–234, reprinted in [26, Vol. 4, pp.71–79].
[41] — ‘Spectral asymmetry and Riemannian geometry. I, II and III’. Math. Proc. Cambridge
Philos. Soc. 77, 78 and 79 (1975, 1975 and 1976), 43–69, 405–432 and 71–99, reprinted in
[26, Vol. 4, pp.81–169].
[42] M. F. Atiyah and G. B. Segal, ‘The index of elliptic operators. II’. Ann. of Math. (2) 87
(1968), 531–545, reprinted in [26, Vol. 3, pp.223–237].
[43] M. F. Atiyah and I. M. Singer, ‘The index of elliptic operators on compact manifolds’.
Bull. Amer. Math. Soc. 69 (1963), 422–433, reprinted in [26, Vol. 3, pp.11–24].
[44] — ‘The index of elliptic operators. I’. Ann. of Math. (2) 87 (1968), 484–530, reprinted in
[26, Vol. 3, pp.171–219].
BIBLIOGRAPHY 725

[45] — ‘The index of elliptic operators. III’. Ann. of Math. (2) 87 (1968), 546–604, reprinted in
[26, Vol. 3, pp.239–299].
[46] — ‘Index theory for skew-adjoint Fredholm operators’. Inst. Hautes Études Sci. Publ. Math.
37/1 (1969), 5–26, reprinted in [26, Vol. 3, pp.349–372].
[47] — ‘The index of elliptic operators. IV’. Ann. of Math. (2) 93 (1971), 119–138, reprinted in
[26, Vol. 3, pp.301–322].
[48] M. F. Atiyah and R. S. Ward, ‘Instantons and algebraic geometry’. Comm. Math. Phys.
55/2 (1977), 117–124, reprinted in [26, Vol. 5, pp.11–20].
[49] F. V. Atkinson, ‘The normal solubility of linear equations in normed spaces’. Mat. Sbornik
N.S. 28(70) (1951), 3–14.
[50] T. Banks, Modern quantum field theory. A concise introduction. Cambridge University
Press, Cambridge, 2008.
[51] W. Bauer, K. Furutani and C. Iwasaki, ‘Spectral analysis and geometry of sub-Laplacian
and related Grushin-type operators’. In: Partial differential equations and spectral theory,
Oper. Theory Adv. Appl., vol. 211. Birkhäuser/Springer Basel AG, Basel, 2011, pp. 183–290.
[52] A. A. Belavin, A. M. Polyakov, A. S. Schwarz and Y. S. Tyupkin, ‘Pseudoparticle
solutions of the Yang-Mills equations’. Phys. Lett. B 59/1 (1975), 85–87.
[53] J. J. Benedetto, Spectral synthesis. B. G. Teubner, Stuttgart, 1975, Mathematische
Leitfäden.
[54] M. Berger, A panoramic view of Riemannian geometry. Springer-Verlag, Berlin, 2003.
[55] N. Berline, E. Getzler and M. Vergne, Heat kernels and Dirac operators, Grundlehren
der Mathematischen Wissenschaften [Fundamental Principles of Mathematical Sciences],
vol. 298. Springer-Verlag, Berlin, 1992.
[56] C. W. Bernard, ‘Physical effects of instantons’. In: Geometrical and topological methods
in gauge theories (Proc. Canad. Math. Soc. Summer Res. Inst., McGill Univ., Montreal,
Que., 1979), Lecture Notes in Phys., vol. 129. Springer, Berlin, 1980, pp. 1–13.
[57] C. W. Bernard, N. H. Christ, A. H. Guth and E. J. Weinberg, ‘Pseudoparticle para-
meters for arbitrary gauge groups’. Phys. Rev. D (3) 16/10 (1977), 2967–2977.
[58] L. Bers, F. John and M. Schechter, Partial differential equations, Lectures in Applied
Mathematics, Vol. III. Interscience Publishers John Wiley & Sons, Inc. New York-London-
Sydney, 1964.
[59] D. D. Bleecker, Gauge theory and variational principles, Global Analysis Pure and Ap-
plied Series A, vol. 1. Addison-Wesley Publishing Co., Reading, Mass., 1981.
[60] D. D. Bleecker and B. Booß-Bavnbek, ‘Spectral invariants of operators of Dirac type on
partitioned manifolds’. In: Aspects of boundary problems in analysis and geometry, Oper.
Theory Adv. Appl., vol. 151. Birkhäuser, Basel, 2004, pp. 1–130. arXiv:math/0304214[math.
AP].
[61] D. D. Bleecker and G. Csordas, Basic partial differential equations. International Press,
Cambridge, MA, 1996.
[62] S. Bochner, ‘Vector fields and Ricci curvature’. Bull. Amer. Math. Soc. 52 (1946), 776–797.
[63] J. Boéchat and A. Haefliger, ‘Plongements différentiables des variétés orientées de di-
mension 4 dans R7 ’. In: Essays on Topology and Related Topics (Mémoires dédiés à Georges
de Rham). Springer, New York, 1970, pp. 156–166.
[64] M. Bohn, On Rho invariants of fiber bundles, 2009. arXiv:0907.3530v1[[Link]].
[65] B. Bojarski, ‘On the index problem for systems of singular integral equations’. Bull. Acad.
Polon. Sci. Sér. Sci. Math. Astronom. Phys. 11 (1963), 653–655.
[66] — ‘The abstract linear conjugation problem and Fredholm pairs of subspaces’. In: In Memo-
riam I.N. Vekua, Tbilisi Univ., Tbilisi, 1979, pp. 45–60, Russian.
[67] J. Bokobza-Haggiag, ‘Opérateurs pseudo-différentiels sur une variété différentiable’. Ann.
Inst. Fourier (Grenoble) 19/1 (1969), 125–177, x.
[68] B. Booß, Elliptische Topologie von Transmissionsproblemen, Bonn. Math. Schr., vol. 58,
1972, Inaugural-Dissertation zur Erlangung des Doktorgrades der Hohen Mathem.-Naturw.
Fakultät der Rheinischen Friedrich-Wilhelms-Universität zu Bonn.
[69] B. Booß-Bavnbek, The determinant of elliptic boundary problems for Dirac opera-
tors. Reader on the Scott–Wojciechowski Theorem, [Link]
[Link], 1999.
726 BIBLIOGRAPHY

[70] — ‘Unique continuation property for Dirac operators, revisited’. In: Geometry and topology:
Aarhus (1998), Contemp. Math., vol. 258. Amer. Math. Soc., Providence, RI, 2000, pp. 21–
32.
[71] — ‘Basic functional analysis puzzles of spectral flow’. J. Aust. Math. Soc. 90/2 (2011),
145–154. arXiv:1010.6084[[Link]].
[72] B. Booß-Bavnbek and D. D. Bleecker, Topology and analysis, Universitext. Springer-
Verlag, New York, 1985, The Atiyah-Singer index formula and gauge-theoretic physics,
Translated from the German by D. Bleecker and A. Mader.
[73] B. Booß-Bavnbek, G. Chen, M. Lesch and C. Zhu, ‘Perturbation of sectorial projections
of elliptic pseudo-differential operators’. J. Pseudo-Differ. Oper. Appl. 3 (2012), 49–79,
10.1007/s11868-011-0042-5. arXiv:1101.0067v4[[Link]].
[74] B. Booß-Bavnbek, G. Esposito and M. Lesch, ‘Quantum gravity: unification of princi-
ples and interactions, and promises of spectral geometry’. SIGMA Symmetry Integrability
Geom. Methods Appl. 3 (2007), Paper 098, 29 pages. arXiv:0708.1705[hep-th].
[75] B. Booß-Bavnbek, G. Esposito and M. Lesch (eds.), New paths towards quantum gravity,
With contributions by J. Ambjørn, J. Jurkiewicz and R. Loll; I.G. Avramidi; B. Booß–
Bavnbek; P. Bouwknegt; J.M. Gracia-Bondı́a; N. Reshetikhin; H. Zessin. Lect. Notes Phys.
Vol. 807. Springer-Verlag, Berlin, 2010.
[76] B. Booß-Bavnbek and K. Furutani, ‘The Maslov index: a functional analytical definition
and the spectral flow formula’. Tokyo J. Math. 21/1 (1998), 1–34.
[77] — ‘Symplectic functional analysis and spectral invariants’. In: Geometric aspects of par-
tial differential equations (Roskilde, 1998), Contemp. Math., vol. 242. Amer. Math. Soc.,
Providence, RI, 1999, pp. 53–83.
[78] B. Booß-Bavnbek, K. Furutani and N. Otsuki, ‘Criss–cross reduction of the Maslov
index and a proof of the Yoshida–Nicolaescu Theorem’. Tokyo J. Math. 24 (2001), 113–128.
[79] B. Booß-Bavnbek, M. Lesch and J. Phillips, ‘Unbounded Fredholm operators and spec-
tral flow’. Canad. J. Math. 57/2 (2005), 225–250. arXiv:math/0108014v3[[Link]].
[80] B. Booß-Bavnbek, M. Lesch and C. Zhu, ‘The Calderón projection: new definition and
applications’. J. Geom. Phys. 59/7 (2009), 784–826. arXiv:0803.4160v1[[Link]].
[81] B. Booss-Bavnbek, M. Marcolli and B.-L. Wang, ‘Weak UCP and perturbed monopole
equations’. Internat. J. Math. 13/9 (2002), 987–1008. arXiv:math/0203171v1[[Link]].
[82] B. Booß-Bavnbek, G. Morchio, F. Strocchi and K. Wojciechowski, ‘Grassmannian
and chiral anomaly’. J. Geom. Phys. 22 ( 1997), 219–244.
[83] B. Booß-Bavnbek and K. P. Wojciechowski, Elliptic boundary problems for Dirac op-
erators, Mathematics: Theory & Applications. Birkhäuser Boston Inc., Boston, MA, 1993.
[84] B. Booß-Bavnbek and C. Zhu, ‘General spectral flow formula for fixed maximal domain’.
Cent. Eur. J. Math. 3/3 (2005), 558–577 (electronic). arXiv:math/0504125v2[[Link]].
[85] — ‘The Maslov index in weak symplectic functional analysis’. Ann. Global. Anal. Geom.
(2013), online first, 38 pages. arXiv:1301.7248[[Link]].
[86] A. Borel and J.-P. Serre, ‘Le théorème de Riemann-Roch’. Bull. Soc. Math. France 86
(1958), 97–136, Exposition of A. Grothendieck’s proof and thorough generalization of the
Hirzebruch-Riemann-Roch-Theorem.
[87] L. Boutet de Monvel, ‘Boundary problems for pseudo-differential operators’. Acta Math.
126/1-2 (1971), 11–51.
[88] T. P. Branson and P. B. Gilkey, ‘Residues of the eta function for an operator of Dirac
type’. J. Funct. Anal. 108/1 (1992), 47–87.
[89] G. E. Bredon, Topology and geometry, Graduate Texts in Mathematics, vol. 139. Springer-
Verlag, New York, 1993.
[90] M. Breuer, ‘Fredholm theories in von Neumann algebras. I’. Math. Ann. 178 (1968), 243–
254.
[91] E. Brieskorn, ‘Beispiele zur Differentialtopologie von Singularitäten’. Invent. Math. 2
(1966), 1–14.
[92] — ‘Über die Dialektik in der Mathematik’. In: Mathematiker über die Mathematik.
Springer-Verlag, Berlin, 1974, pp. 220–286, in the collection of essays [326].
[93] — ‘The development of geometry and topology, Notes of introductory lectures given at the
University of La Habana in 1973’. Mat. z. Berufspraxis Math. Heft 17 (1976), 109–203.
[94] — ‘Gibt es eine Wiedergeburt der Qualität in der Mathematik?’. In: Wissenschaft zwischen
Qualitas und Quantitas. Birkhäuser, Basel, 2003, pp. 243–410.
BIBLIOGRAPHY 727

[95] E. Brieskorn et al., Die Atiyah-Singer-Indexformel, Seminarvorträge, Bonn, mimeo-


graphed, 1963.
[96] A. Brill and M. Noether, ‘Die Entwicklung der Theorie der algebraischen Funktionen in
älterer und neuerer Zeit.’. Jber. Deutsch. Math.-Verein. 3 (1892/93), 107–556.
[97] T. Bröcker and K. Jänich, Introduction to differential topology. Cambridge University
Press, Cambridge, 1982, Translated from the German by C. B. Thomas and M. J. Thomas.
[98] L. G. Brown, R. G. Douglas and P. A. Fillmore, ‘Unitary equivalence modulo the com-
pact operators and extensions of C ∗ -algebras’. In: Proceedings of a Conference on Operator
Theory (Dalhousie Univ., Halifax, N.S., 1973) (Berlin). Springer, 1973, pp. 58–128. Lecture
Notes in Math., Vol. 345.
[99] J. Brüning and M. Lesch, ‘On boundary value problems for Dirac type operators. I.
Regularity and self-adjointness’. J. Funct. Anal. 185/1 (2001), 1–62.
[100] U. Bunke, ‘On the gluing problem for the η-invariant’. J. Differential Geom. 41/2 (1995),
397–448.
[101] — ‘Index theory, eta forms, and Deligne cohomology’. Mem. Amer. Math. Soc. 198/928
(2009), vi+120.
[102] A. P. Calderón, ‘Boundary value problems for elliptic equations’. In: Outlines Joint Sym-
pos. Partial Differential Equations (Novosibirsk, 1963). Acad. Sci. USSR Siberian Branch,
Moscow, 1963, pp. 303–304.
[103] — ‘The analytic calculation of the index of elliptic equations’. Proc. Nat. Acad. Sci. U.S.A.
57 (1967), 1193–1194.
[104] — Lecture notes on pseudo-differential operators and elliptic boundary value problems.
I. Consejo Nacional de Investigaciones Cientificas y Tecnicas Instituto Argentino de
Matemática, Buenos Aires, 1976, Cursos de Matemática, No. 1. [Courses in Mathematics,
No. 1].
[105] A. P. Calderón and A. Zygmund, ‘On singular integrals’. Amer. J. Math. 78 (1956),
289–309.
[106] J. W. Calkin, ‘Two-sided ideals and congruences in the ring of bounded operators in Hilbert
space’. Ann. of Math. (2) 42 (1941), 839–873.
[107] M. Cantor, ‘Elliptic operators and the decomposition of tensor fields’. Bull. Amer. Math.
Soc. (N.S.) 5/3 (1981), 235–262.
[108] T. Carleman, ‘Propriétés asymptotiques des fonctions fondamentales des membranes vi-
brantes’. In: 8. Skand. Mat.-Kongr. (Stockholm, 1934), 1935, pp. 34–44.
[109] H. Cartan and Schwartz L. et al., Séminaire Henri Cartan, 16e année: 1963/64,
dirigée par Henri Cartan et Laurent Schwartz. Théorème d’Atiyah-Singer sur l’indice d’un
opérateur différential elliptique. Fasc. 1/2, Exposés 1/2 à 15/25. Secrétariat mathématique,
Paris, 1965.
[110] A. Cho, ‘Higgs Boson Makes Its Debut After Decades-Long Search’. Science 337 (2012),
141–143.
[111] E. A. Coddington and N. Levinson, Theory of ordinary differential equations. McGraw-
Hill Book Company, Inc., New York-Toronto-London, 1955.
[112] H. O. Cordes and J.-P. Labrousse, ‘The invariance of the index in the metric space of
closed operators’. J. Math. Mech. 12 (1963), 693–719.
[113] E. Corrigan and D. B. Fairlie, ‘Scalar field theory and exact solutions to a classical SU(2)
gauge theory’. Phys. Lett. B 67/1 (1977), 69–71.
[114] E. Corrigan and P. Goddard, ‘Some aspects of instantons’. In: Geometrical and topolog-
ical methods in gauge theories (Proc. Canad. Math. Soc. Summer Res. Inst., McGill Univ.,
Montreal, Que., 1979), Lecture Notes in Phys., vol. 129. Springer, Berlin, 1980, pp. 14–44.
[115] G. D. Coughlan and J. E. Dodd, The ideas of particle physics: an introduction for
scientists, second ed.. Cambridge University Press, New York, 1991, (illustrated, reprinted,
revised).
[116] R. Courant and D. Hilbert, Methods of mathematical physics. Vol. I and II. Interscience
Publishers, Inc., New York, N.Y., 1953 and 1962, German original of 1924 and 1931.
[117] P. J. Davis and P. Rabinowitz, Methods of numerical integration. Dover Publications Inc.,
Mineola, NY, 2007, Corrected reprint of the second (1984) edition.
[118] A. Devinatz, ‘On Wiener-Hopf operators’. In: Functional Analysis (Proc. Conf., Irvine,
Calif., 1966). Academic Press, London, 1967, pp. 81–118.
728 BIBLIOGRAPHY

[119] B. DeWitt, ‘Quantum gravity: yesterday and today’. Gen. Relativity Gravitation 41/2
(2009), 413–419.
[120] J. A. Dieudonné, ‘Sur les homomorphismes d’espaces normés’. Bull. Sci. Math. (2) 67
(1943), 72–84.
[121] — Course on algebraic geometry. 1: An outline of the history and development of algebraic
geometry. (Cours de géométrie algébrique. I: Aperçu historique sur le développement de la
géométrie algébrique.). Paris: Presses Universitaires de France, 974 (French).
[122] J. Dixmier, C ∗ -algebras. North-Holland Publishing Co., Amsterdam, 1977, Translated from
the French by Francis Jellett, North-Holland Mathematical Library, Vol. 15.
[123] S. K. Donaldson, ‘An application of gauge theory to four-dimensional topology’. J. Dif-
ferential Geom. 18/2 (1983), 279–315.
[124] — ‘The orientation of Yang-Mills moduli spaces and 4-manifold topology’. J. Differential
Geom. 26/3 (1987), 397–428.
[125] — ‘Polynomial invariants for smooth four-manifolds’. Topology 29/3 (1990), 257–315.
[126] S. K. Donaldson and P. B. Kronheimer, The geometry of four-manifolds, Oxford Mathe-
matical Monographs. The Clarendon Press Oxford University Press, New York, 1990, Oxford
Science Publications.
[127] S. K. Donaldson and E. Segal, ‘Gauge theory in higher dimensions, II’. In: Surveys in
differential geometry. Volume XVI. Geometry of special holonomy and related topics, Surv.
Differ. Geom., vol. 16. Int. Press, Somerville, MA, 2011, pp. 1–41. arXiv:0902.3239v1[math.
DG].
[128] R. G. Douglas, Banach algebra techniques in operator theory. Academic Press, New York,
1972, Pure and Applied Mathematics, Vol. 49.
[129] R. G. Douglas and K. P. Wojciechowski, ‘Adiabatic limits of the η–invariants. The
odd–dimensional Atiyah–Patodi–Singer problem’. Comm. Math. Phys. 142 (1991), 139–168.
[130] J. Dugundji, Topology. Allyn and Bacon Inc., Boston, Mass., 1966.
[131] J. Duhamel, ‘Mémoire sur la méthode générale relative au mouvement de la chaleur dans
les corps solides plongés dans les milieux dont la température varie avec le temps’. J. Éc.
polyt. Paris 14/22 (1833), 20.
[132] J. J. Duistermaat and V. W. Guillemin, ‘The spectrum of positive elliptic operators and
periodic bicharacteristics’. Invent. Math. 29/1 (1975), 39–79.
[133] N. Dunford and J. T. Schwartz, Linear operators. Part I–III, Wiley Classics Library.
John Wiley & Sons Inc., New York, 1988, General theory, With the assistance of William
G. Bade and Robert G. Bartle, Reprint of the 1958, 1963 and 1971 originals, A Wiley-
Interscience Publication.
[134] H. Dym and H. P. McKean, Fourier series and integrals. Academic Press, New York, 1972,
Probability and Mathematical Statistics, No. 14.
[135] E. B. Dynkin and A. A. Juschkewitsch, Sätze und Aufgaben über Markoffsche Prozesse,
Aus dem Russischen übersetzt von K. Schürger. Vorwort zur deutschen Ausgabe von K.
Krickeberg. Heidelberger Taschenbücher, Band 51. Springer-Verlag, Berlin, 1969.
[136] T. Eguchi, P. B. Gilkey and A. J. Hanson, ‘Gravitation, gauge theories and differential
geometry’. Phys. Rep. 66/6 (1980), 213–393.
[137] J. Eichhorn, ‘Index theory for generalized Dirac operators on open manifolds’. In: C ∗ -
algebras and elliptic theory, Trends Math.. Birkhäuser, Basel, 2006, pp. 73–128.
[138] — Global analysis on open manifolds. Nova Science Publishers Inc., New York, 2007.
[139] — Relative index theory, determinants and torsion for open manifolds. World Scientific
Publishing Co. Pte. Ltd., Hackensack, NJ, 2009.
[140] S. Eilenberg and N. Steenrod, Foundations of algebraic topology. Princeton University
Press, Princeton, New Jersey, 1952.
[141] L. P. Eisenhart, Riemannian geometry, Princeton Landmarks in Mathematics. Princeton
University Press, Princeton, NJ, 1997, Eighth printing, Princeton Paperbacks.
[142] E. Elizalde, Ten physical applications of spectral zeta functions, Lecture Notes in Physics.
New Series m: Monographs, vol. 35. Springer-Verlag, Berlin, 1995.
[143] P. Enflo, ‘A counterexample to the approximation problem in Banach spaces’. Acta Math.
130 (1973), 309–317.
[144] G. Esposito, Dirac operators and spectral geometry, Cambridge Lecture Notes in Physics,
vol. 12. Cambridge University Press, Cambridge, 1998.
[145] — An introduction to quantum gravity, 2011. arXiv:1108.3269v1[hep-th].
BIBLIOGRAPHY 729

[146] P. A. Fillmore (ed.), Proceedings of a Conference on Operator Theory (Dalhousie Univer-


sity, Halifax, Nova Scotia, April 13th and 14th, 1973), Lecture Notes in Mathematics, Vol.
345. Springer-Verlag, Berlin, 1973.
[147] E. I. Fredholm, ‘Sur une classe d’équations fonctionnelles’. Acta Math. 27/1 (1903), 365–
390.
[148] D. S. Freed and K. K. Uhlenbeck, Instantons and four-manifolds, second ed., Mathe-
matical Sciences Research Institute Publications, vol. 1. Springer-Verlag, New York, 1991.
[149] M. H. Freedman, ‘The topology of four-dimensional manifolds’. J. Differential Geom. 17/3
(1982), 357–453.
[150] M. H. Freedman and F. Quinn, Topology of 4-manifolds, Princeton Mathematical Series,
vol. 39. Princeton University Press, Princeton, NJ, 1990.
[151] R. Friedman and J. W. Morgan, ‘Algebraic surfaces and 4-manifolds: some conjectures
and speculations’. Bull. Amer. Math. Soc. (N.S.) 18/1 (1988), 1–19.
[152] D. Fujiwara, ‘On an analytic index-formula for elliptic operators’. Proc. Japan Acad. 44
(1968), 147–150.
[153] S. A. Fulling and G. Kennedy, ‘The resolvent parametrix of the general elliptic linear
differential operator: a closed form for the intrinsic symbol’. Trans. Amer. Math. Soc. 310/2
(1988), 583–617.
[154] F. Fuquan, ‘Embedding four manifolds in R7 ’. Topology 33/3 (1994), 447–454.
[155] D. Fursaev and D. Vassilevich, Operators, geometry and quanta, Theoretical and Math-
ematical Physics. Springer, Dordrecht, 2011, Methods of spectral geometry in quantum field
theory.
[156] M. Furuta, ‘Monopole equation and the 11 8
-conjecture’. Math. Res. Lett. 8/3 (2001), 279–
291.
[157] — Index theorem. 1, Translations of Mathematical Monographs, vol. 235. American Mathe-
matical Society, Providence, RI, 2007, Translated from the 1999 Japanese original by Kaoru
Ono, Iwanami Series in Modern Mathematics. Index theorem. 2 (in Japanese) is published
in 2002.
[158] K. Furutani, ‘On the Quillen determinant’. J. Geom. Phys. 49/3-4 (2004), 366–375.
[159] I. M. Gelfand, ‘On elliptic equations’. Russian Math. Surveys 15/3 (1960), 113–123.
[160] — ‘The cohomology of infinite dimensional Lie algebras: some questions of integral geom-
etry’. In: Actes du Congrès International des Mathématiciens (Nice, 1970), Tome 1.
Gauthier-Villars, Paris, 1971, pp. 95–111.
[161] I. M. Gelfand, D. A. Raikow and G. E. Schilow, Kommutative normierte Alge-
bren, Übersetzung und wissenschaftliche Redaktion von Helmut Boseck. Mathematische
Forschungsberichte, XIII. VEB Deutscher Verlag der Wissenschaften, Berlin, 1964.
[162] E. Getzler, ‘Pseudodifferential operators on supermanifolds and the Atiyah-Singer index
theorem’. Comm. Math. Phys. 92/2 (1983), 163–178.
[163] — ‘A short proof of the local Atiyah-Singer index theorem’. Topology 25/1 (1986), 111–117.
[164] H. Ghorbani, D. Musso and A. Lerda, Stringy instanton effects in N = 2 gauge theories,
2010. arXiv:1012.1122v1[hep-th].
[165] G. W. Gibbons and S. W. Hawking, ‘Action integrals and partition functions in quantum
gravity’. [Link]. D15 (1977), 2752–2756.
[166] P. B. Gilkey, ‘Curvature and the eigenvalues of the Laplacian for elliptic complexes’. Ad-
vances in Math. 10 (1973), 344–382.
[167] — Invariance theory, the heat equation, and the Atiyah-Singer index theorem, second ed.,
Studies in Advanced Mathematics. CRC Press, Boca Raton, FL, 1995.
[168] — Asymptotic formulae in spectral geometry, Studies in Advanced Mathematics. Chapman
& Hall/CRC, Boca Raton, FL, 2004.
[169] — ‘The spectral geometry of operators of Dirac and Laplace type’. In: Handbook of global
analysis. Elsevier Sci. B. V., Amsterdam, 2008, pp. 289–326, 1212.
[170] P. B. Gilkey, R. Ivanova and S. Nikčević, ‘Characteristic classes’. In:
J.-P. Francoise, G.L. Naber and Tsou S.T. (eds.), Encyclopedia of Mathematical Physics,
Vol. 3 (Oxford). Elsevier, 2006, pp. 448–496.
[171] G. Giraud, ‘Équations à intégrales principales; étude suivie d’une application’. Ann. Sci.
École Norm. Sup. (3) 51 (1934), 251–372.
[172] — ‘Sur certaines opérations du type elliptique.’. C. R. Acad. Sci., Paris 200 (1935), 1651–
1653.
730 BIBLIOGRAPHY

[173] — ‘Sur une classe générale d’équations à intégrales principales.’. C. R. Acad. Sci., Paris
202 (1936), 2124–2127.
[174] J. Glimm and A. Jaffe, Quantum physics, second ed.. Springer-Verlag, New York, 1987,
A functional integral point of view.
[175] I. Z. Gohberg and I. A. Feldman, Faltungsgleichungen und Projektionsverfahren zu ihrer
Lösung. Birkhäuser Verlag, Basel, 1974, Übersetzung aus dem Russischen von Reinhard
Lehmann und Jürgen Leiterer, Mathematische Reihe, Band 49.
[176] I. Z. Gohberg and S. Goldberg, Basic operator theory. Birkhäuser Boston, Mass., 1981.
[177] I. Z. Gohberg and M. G. Krein, ‘The basic propositions on defect numbers, root numbers
and indices of linear operators’. Amer. Math. Soc. Transl. (2) 13 (1960), 185–264.
[178] — ‘Systems of integral equations on a half line with kernels depending on the difference of
arguments’. Amer. Math. Soc. Transl. (2) 14 (1960), 217–287.
[179] S. I. Goldberg, Curvature and homology. Dover Publications Inc., Mineola, NY, 1998,
Revised reprint of the 1970 edition.
[180] R. E. Gompf, ‘An infinite set of exotic R4 ’s’. J. Differential Geom. 21/2 (1985), 283–300.
[181] M. L. Gorbachuk et al., ‘Yaroslav Borisovich Lopatins0 kiı̆ (November 9, 1906–March 10,
1981)’. Ukraı̈n. Mat. Zh. 58/11 (2006), 1443–1445.
[182] C. Gordon, D. Webb and S. Wolpert, ‘Isospectral plane domains and surfaces via Rie-
mannian orbifolds’. Invent. Math. 110/1 (1992), 1–22.
[183] J. M. Gracia-Bondı́a, ‘Notes on quantum gravity and noncommutative geometry’. In: New
paths towards quantum gravity (B. Booß-Bavnbek, G. Esposito and M. Lesch, eds.), Lecture
Notes in Physics, vol. 807. Springer Berlin / Heidelberg, 2010, pp. 3–58, 10.1007/978-3-642-
11897-5 1. arXiv:1005.1174v1[hep-th].
[184] R. L. Graham, D. E. Knuth and O. Patashnik, Concrete mathematics. Addison-Wesley
Publishing Company Advanced Book Program, Reading, MA, 1989, A foundation for com-
puter science.
[185] M. J. Greenberg, Lectures on algebraic topology. W. A. Benjamin, Inc., New York-
Amsterdam, 1967.
[186] P. Griffiths and J. Harris, Principles of algebraic geometry. Wiley-Interscience [John
Wiley & Sons], New York, 1978, Pure and Applied Mathematics.
[187] A. Grigis and J. Sjöstrand, Microlocal analysis for differential operators, London Mathe-
matical Society Lecture Note Series, vol. 196. Cambridge University Press, Cambridge, 1994,
An introduction.
[188] A. Grothendieck, ‘Produits tensoriels topologiques et espaces nucléaires’. Mem. Amer.
Math. Soc. 1955/16 (1955), 140.
[189] G. Grubb, ‘Spectral boundary conditions for generalizations of Laplace and Dirac operators’.
Comm. Math. Phys. 242 (2003), 243–280.
[190] — Distributions and operators, Graduate Texts in Mathematics, vol. 252. Springer, New
York, 2009.
[191] — Encounters with spectral theory, unpublished, 2012, Retirement Lecture, given 2 Novem-
ber, 2012, at Copenhagen University.
[192] — ‘The sectorial projection defined from logarithms’. Math. Scand. 111/1 (2012), 118–126.
[193] G. Grubb and E. Schrohe, ‘Traces and quasi-traces on the Boutet de Monvel algebra’.
Ann. Inst. Fourier (Grenoble) 54/5 (2004), 1641–1696, xvii, xxii.
[194] E. Guentner, ‘K-homology and the index theorem’. In: Index theory and operator algebras
(Boulder, CO, 1991), Contemp. Math., vol. 148. Amer. Math. Soc., Providence, RI, 1993,
pp. 47–66.
[195] C. Guillarmou and L. Tzou, ‘Identification of a connection from Cauchy data on a Rie-
mann surface with boundary’. Geom. Funct. Anal. 21/2 (2011), 393–418.
[196] V. W. Guillemin and A. Pollack, Differential topology. Prentice-Hall Inc., Englewood
Cliffs, N.J., 1974.
[197] W. Haack, ‘Randwertprobleme höherer Charakteristik für ein System von zwei elliptischen
Differentialgleichungen’. Math. Nachr. 8 (1952), 123–132.
[198] A. Haefliger and M. W. Hirsch, ‘On the existence and classification of differentiable
embeddings’. Topology 2 (1963), 129–135.
[199] V. L. Hansen, Fundamental concepts in modern analysis. World Scientific Publishing Co.
Inc., River Edge, NJ, 1999.
BIBLIOGRAPHY 731

[200] P. Hartman, Ordinary differential equations, Classics in Applied Mathematics, vol. 38.
Society for Industrial and Applied Mathematics (SIAM), Philadelphia, PA, 2002, Corrected
reprint of the second (1982) edition [Birkhäuser, Boston, MA; MR0658490 (83e:34002)],
With a foreword by Peter Bates.
[201] E. Hellinger and O. Toeplitz, Integralgleichungen und Gleichungen mit unendlichvielen
Unbekannten. Chelsea Publishing Company, New York, N. Y., 1953, Reprinted from the
1928 edition [Teubner, Leipzig, “Sonderausgabe des 1927 erschienenen Artikels II C 13 aus
der Enzyklopädie der Mathematischen Wissenschaften”, pp. 1335–1616].
[202] G. Hellwig, ‘Das Randwertproblem eines linearen elliptischen Systems’. Math. Z. 56 (1952),
388–408.
[203] R. Hermann, Vector bundles in mathematical physics. Vols. I, II. W. A. Benjamin, Inc.,
New York, 1970.
[204] N. Higson and J. Roe, Analytic K-homology, Oxford Mathematical Monographs. Oxford
University Press, Oxford, 2000, Oxford Science Publications.
[205] — Operator K-Theory and the Atiyah-Singer Index Theorem, 2004, In preparation for
Princeton University Press.
[206] M. W. Hirsch, Differential topology. Springer-Verlag, New York, 1976, Graduate Texts in
Mathematics, No. 33.
[207] F. Hirzebruch, Topological methods in algebraic geometry, Third enlarged edition. New
appendix and translation from the second German edition by R. L. E. Schwarzenberger, with
an additional section by A. Borel. Die Grundlehren der Mathematischen Wissenschaften,
Band 131. Springer-Verlag New York, Inc., New York, 1966.
[208] — ‘Elliptische Differentialoperatoren auf Mannigfaltigkeiten’. In: Arbeitsgemeinschaft für
Forschung des Landes Nordrhein-Westfalen, Natur–, Ingenieur- und Gesellschaftswis-
senschaften, Heft 33. Westdeutscher Verlag, Köln, 1966, pp. 583–608.
[209] — ‘Hilbert modular surfaces’. Enseignement Math. (2) 19 (1973), 183–281.
[210] F. Hirzebruch and H. Hopf, ‘Felder von Flächenelementen in 4-dimensionalen Mannig-
faltigkeiten’. Math. Ann. 136 (1958), 156–172.
[211] F. Hirzebruch and M. Kreck, ‘On the concept of genus in topology and complex analysis’.
Notices Amer. Math. Soc. 56/6 (2009), 713–719.
[212] F. Hirzebruch and W. Scharlau, Einführung in die Funktionalanalysis. Bibliographisches
Institut, Mannheim, 1971, B. I.-Hochschultaschenbücher, No. 296*.
[213] F. Hirzebruch and D. Zagier, The Atiyah-Singer theorem and elementary number theory.
Publish or Perish Inc., Boston, Mass., 1974, Mathematics Lecture Series, No. 3.
[214] N. J. Hitchin, ‘Compact four-dimensional Einstein manifolds’. J. Differential Geometry 9
(1974), 435–441.
[215] — ‘Harmonic spinors’. Advances in Math. 14 (1974), 1–55.
[216] — ‘The Atiyah-Singer index theorem’. In: H. Holden, R. Piene (eds.), The Abel Prize,
2003–2007. Springer-Verlag, Berlin, 2010, pp. 115–150, The first five years, With 1 DVD.
[217] L. Hörmander, Linear partial differential operators, Die Grundlehren der mathematischen
Wissenschaften, Bd. 116. Academic Press Inc., Publishers, New York, 1963.
[218] — ‘Pseudo-differential operators and hypoelliptic equations’. In: Singular integrals (Proc.
Sympos. Pure Math., Vol. X, Chicago, Ill., 1966). Amer. Math. Soc., Providence, R.I., 1967,
pp. 138–183.
[219] — ‘The calculus of Fourier integral operators’. In: Prospects in mathematics (Proc. Sym-
pos., Princeton Univ., Princeton, N.J., 1970). Princeton Univ. Press, Princeton, N.J., 1971,
pp. 33–57. Ann. of Math. Studies, No. 70.
[220] — ‘Fourier integral operators. I’. Acta Math. 127/1-2 (1971), 79–183.
[221] — ‘On the existence and the regularity of solutions of linear pseudo-differential equations’.
Enseignement Math. (2) 17 (1971), 99–163.
[222] — ‘On the index of pseudodifferential operators’. In: Elliptische Differentialgleichungen,
Band II. Akademie-Verlag, Berlin, 1971, pp. 127–146. Schriftenreihe Inst. Math. Deutsch.
Akad. Wissensch. Berlin, Reihe A, Heft 8.
[223] — The analysis of linear partial differential operators. II, Grundlehren der Mathematis-
chen Wissenschaften [Fundamental Principles of Mathematical Sciences], vol. 257. Springer-
Verlag, Berlin, 1983, Differential operators with constant coefficients.
732 BIBLIOGRAPHY

[224] — The analysis of linear partial differential operators. III, Grundlehren der Mathematis-
chen Wissenschaften [Fundamental Principles of Mathematical Sciences], vol. 274. Springer-
Verlag, Berlin, 1985, Pseudodifferential operators.
[225] — The analysis of linear partial differential operators. IV, Grundlehren der Mathematis-
chen Wissenschaften [Fundamental Principles of Mathematical Sciences], vol. 275. Springer-
Verlag, Berlin, 1985, Fourier integral operators.
[226] E. P. Hsu, Stochastic analysis on manifolds, Graduate Studies in Mathematics, vol. 38.
American Mathematical Society, Providence, RI, 2002.
[227] J. E. Humphreys, Introduction to Lie algebras and representation theory. 3rd printing, rev.,
Graduate Texts in Mathematics. Springer-Verlag, New York, 1980.
[228] D. Husemoller, Fibre bundles, second ed.. Springer-Verlag, New York, 1975, Graduate
Texts in Mathematics, No. 20.
[229] L. Illusie, ‘Contractibilité du groupe linéaire des espaces de Hilbert de dimension infinie
(d’après N. Kuiper)’. In: Séminaire Bourbaki, Vol. 9. Soc. Math. France, Paris, 1995,
pp. Exp. No. 284, 105–113.
[230] A. G. Ivachnenko and V. G. Lapa, Cybernetics and forcasting techniques. American El-
sevier, New York, 1967, translation from Russian.
[231] J. Ize, ‘Bifurcation theory for Fredholm operators’. Mem. Amer. Math. Soc. 7/174 (1976),
viii+128.
[232] R. Jackiw, C. Nohl and C. Rebbi, ‘Conformal properties of pseudoparticle configurations’.
[Link]. D15 (1977), 1642.
[233] K. Jänich, ‘Vektorraumbündel und der Raum der Fredholm-Operatoren’. Math. Ann. 161
(1965), 129–142.
[234] K. Jörgens, Lineare Integraloperatoren. B. G. Teubner Stuttgart, 1970, Mathematische
Leitfäden. Eng. transl. Pitman, San Francisco, 1982.
[235] J. Jost, Riemannian geometry and geometric analysis, sixth ed., Universitext. Springer,
Heidelberg, 2011.
[236] M. Kac, ‘On some connections between probability theory and differential and integral
equations’. In: Proceedings of the Second Berkeley Symposium on Mathematical Statistics
and Probability, 1950 (Berkeley and Los Angeles). University of California Press, 1951,
pp. 189–215.
[237] — ‘Can one hear the shape of a drum?’. Amer. Math. Monthly 73/4, part II (1966), 1–23.
[238] M. Kaku, Quantum field theory. The Clarendon Press Oxford University Press, New York,
1993, A modern introduction.
[239] T. Kaluza, ‘Zum Unitätsproblem der Physik.’. Berl. Ber. 1921 (1921), 966–972 (German).
[240] M. Karoubi, K-theory. Springer-Verlag, Berlin, 1978, An introduction, Grundlehren der
Mathematischen Wissenschaften, Band 226.
[241] N. Karoui and H. Reinhard, ‘Processus de diffusion dans Rn ’. In: Séminaire de Prob-
abilités, VII (Univ. Strasbourg, année universitaire 1971–1972). Springer, Berlin, 1973,
pp. 95–117. Lecture Notes in Math., Vol. 321.
[242] T. Kato, Perturbation theory for linear operators, Die Grundlehren der mathematischen
Wissenschaften, Band 132. Springer-Verlag New York, Inc., New York, 1966.
[243] — ‘Trotter’s product formula for an arbitrary pair of self-adjoint contraction semigroups’. In:
Topics in functional analysis (essays dedicated to M. G. Kreı̆n on the occasion of his 70th
birthday), Adv. in Math. Suppl. Stud., vol. 3. Academic Press, New York, 1978, pp. 185–195.
[244] O. Klein, ‘Quantentheorie und fünfdimensionale Relativitätstheorie’. Z. Phys. 37 (1926),
895–906, 10.1007/BF01397481.
[245] — ‘On the theory of charged fields’. Surveys High [Link]. 5 (1986), 269–285, Reprinted
from New Theories in Physics, Proc. of a Conf. held in Warsaw, 1938, Institute for Intel-
lectual Collaboration, Paris.
[246] S. Klimek and K. P. Wojciechowski, ‘Adiabatic cobordism theorems for analytic torsion
and η–invariant’. J. Funct. Anal. 136 (1996), 269–293.
[247] M. Kline, Mathematical thought from ancient to modern times. Oxford University Press,
New York, 1972.
[248] S. Kobayashi and K. Nomizu, Foundations of differential geometry. Vol I. Interscience
Publishers, a division of John Wiley & Sons, New York-Lond on, 1963.
[249] — Foundations of differential geometry. Vol. II. Interscience Publishers John Wiley & Sons,
Inc., New York-London-Sydney, 1969.
BIBLIOGRAPHY 733

[250] K. Kodaira, ‘On the structure of compact complex analytic surfaces. I’. Amer. J. Math. 86
(1964), 751–798.
[251] I. Kolář, P. W. Michor and J. Slovák, Natural operations in differential geometry.
Springer-Verlag, Berlin, 1993, free copy from [Link]
html.
[252] A. Kolmogoroff, ‘Interpolation und Extrapolation von stationären zufälligen Folgen’. Bull.
Acad. Sci. URSS Sér. Math. [Izvestia Akad. Nauk. SSSR] 5 (1941), 3–14.
[253] M. Kontsevich and S. Vishik, ‘Geometry of determinants of elliptic operators’. In: Func-
tional analysis on the eve of the 21st century, Vol. 1 (New Brunswick, NJ, 1993), Progr.
Math., vol. 131. Birkhäuser Boston, Boston, MA, 1995, pp. 173–197.
[254] T. Kori, ‘Index of the Dirac operator on S 4 and the infinite-dimensional Grassmannian on
S 3 ’. Japan. J. Math. (N.S.) 22/1 (1996), 1–36.
[255] — ‘Chiral anomaly and Grassmannian boundary conditions’. In: Geometric aspects of par-
tial differential equations (Roskilde, 1998), Contemp. Math., vol. 242. Amer. Math. Soc.,
Providence, RI, 1999, pp. 35–42.
[256] — ‘Spinor analysis on C2 and on conformally flat 4-manifolds’. Japan. J. Math. (N.S.) 28/1
(2002), 1–30.
[257] U. Koschorke, ‘Framefields and nondegenerate singularities’. Bull. Amer. Math. Soc. 81
(1975), 157–160.
[258] T. Kotake, ‘The fixed point theorem of Atiyah-Bott via parabolic operators’. Comm. Pure
Appl. Math. 22 (1969), 789–806.
[259] — ‘An analytic proof of the classical Riemann-Roch theorem’. In: Global Analysis (Proc.
Sympos. Pure Math., Vol. XVI, Berkeley, Calif., 1968). Amer. Math. Soc., Providence,
R.I., 1970, pp. 137–146.
[260] M. G. Krein, ‘Integral equations on the half-line with a kernel depending on the difference
of the arguments’. Uspehi Mat. Nauk 13/5 (83) (1958), 3–120, Amer. Math. Soc. Transl.
(2) 22 (1962), 163-288.
[261] L. Kronecker, ‘On systems of functions of several variables. (Über Systeme von Funktionen
mehrerer Variabeln)’. Berl. Monatsber. 1869 (1869), 159–193, 688–698 (German).
[262] P. B. Kronheimer and T. S. Mrowka, ‘The genus of embedded surfaces in the projective
plane’. Math. Res. Lett. 1/6 (1994), 797–808.
[263] — Dehn surgery, the fundamental group and SU(2), 2003. arXiv:math/0312322v1[math.
GT].
[264] — ‘Witten’s conjecture and property P’. Geom. Topol. 8 (2004), 295–310 (electronic).
arXiv:math/0311489v5[[Link]].
[265] — Monopoles and three-manifolds, New Mathematical Monographs, vol. 10. Cambridge
University Press, Cambridge, 2007.
[266] — ‘Instanton Floer homology and the Alexander polynomial’. Algebr. Geom. Topol. 10/3
(2010), 1715–1738.
[267] — Khovanov homology is an unknot-detector, 2010. arXiv:1005.4346[[Link]].
[268] — Filtrations on instanton homology, 2011. arXiv:1110.1290[[Link]].
[269] — ‘Knot homology groups from instantons’. J. Topol. 4/4 (2011), 835–918. arXiv:0806.
1053v2[[Link]].
[270] N. H. Kuiper, ‘The homotopy type of the unitary group of Hilbert space’. Topology 3 (1965),
19–30.
[271] J. L. Lagrange, ‘Lagrange’s Letter to Euler of 12 August, 1755 (in Latin)’. In:
J.A. Serret and G. Darboux (eds.), Œuvres de Lagrange 14. Gauthier-Villars, Paris, 1892,
pp. 366–375.
[272] S. Lang, Differential manifolds. Addison-Wesley Publishing Co., Inc., Reading, Mass.-
London-Don Mills, Ont., 1972.
[273] H. B. Lawson, Jr. and M.-L. Michelsohn, Spin geometry, Princeton Mathematical Series,
vol. 38. Princeton University Press, Princeton, NJ, 1989.
[274] P. D. Lax, ‘On Cauchy’s problem for hyperbolic equations and the differentiability of solu-
tions of elliptic equations’. Comm. Pure Appl. Math. 8 (1955), 615–633.
[275] C. LeBrun, ‘Einstein metrics and Mostow rigidity’. Math. Res. Lett. 2/1 (1995), 1–8.
[276] M. Lesch, Operators of Fuchs type, conical singularities, and asymptotic methods, Teubner-
Texte zur Mathematik [Teubner Texts in Mathematics], vol. 136. B. G. Teubner Verlagsge-
sellschaft mbH, Stuttgart, 1997.
734 BIBLIOGRAPHY

[277] — ‘The uniqueness of the spectral flow on spaces of unbounded self-adjoint Fredholm op-
erators’. In: Spectral geometry of manifolds with boundary and decomposition of manifolds,
Contemp. Math., vol. 366. Amer. Math. Soc., Providence, RI, 2005, pp. 193–224.
[278] M. Lesch, H. Moscovici and M. J. Pflaum, ‘Connes-Chern character for manifolds
with boundary and eta cochains’. Mem. Amer. Math. Soc. 220/1036 (2012). arXiv:0912.
0194[[Link]].
[279] E. E. Levi, ‘Sulle equazioni lineari totalmente ellittiche alle derivate parziali’. Palermo
Rend. 24 (1907), 275–317 (Italian).
[280] J.-L. Lions and E. Magenes, Problèmes aux limites non homogènes et applications. Vol.
1, Travaux et Recherches Mathématiques, No. 17. Dunod, Paris, 1968.
[281] J. D. Logan, Applied partial differential equations, Undergraduate Texts in Mathematics.
Springer-Verlag, New York, 1998.
[282] Y. B. Lopatinskiı̆, ‘On a method of reducing boundary problems for a system of differential
equations of elliptic type to regular integral equations’. Ukrain. Mat. Ž. 5 (1953), 123–151.
[283] J. Madore, ‘Geometric methods in classical field theory’. Phys. Rep. 75/3 (1981), 125–204.
[284] E. Magenes, ‘Spazi d’interpolazione ed equazioni a derivate parziali’. In: Atti del Settimo
Congresso dell’ Unione Matematica I (Genova, 1963). Edizioni Cremonese, Rome, 1965,
pp. 134–197.
[285] W. S. Massey, ‘On the Stiefel-Whitney classes of a manifold. II’. Proc. Amer. Math. Soc.
13 (1962), 938–942.
[286] A. Mathew, Climbing Mount Bourbaki: thoughts on mathematics, [Link]
[Link]/, ongoing, well done student blog on a variety of subjects with focus on
Atiyah’s and Deligne’s work.
[287] K. H. Mayer, ‘Elliptische Differentialoperatoren und Ganzzahligkeitssätze für charakteris-
tische Zahlen’. Topology 4 (1965), 295–313.
[288] D. McDuff and D. Salamon, Introduction to symplectic topology, Oxford Mathematical
Monographs. The Clarendon Press Oxford University Press, New York, 1995, Oxford Science
Publications.
[289] P. McKeag and Y. Safarov, ‘Pseudodifferential operators on manifolds: a coordinate-free
approach’. In: Partial differential equations and spectral theory, Oper. Theory Adv. Appl.,
vol. 211. Birkhäuser/Springer Basel AG, Basel, 2011, pp. 321–341.
[290] H. P. McKean, ‘Fredholm determinants’. Cent. Eur. J. Math. 9/2 (2011), 205–243.
[291] H. P. McKean and I. M. Singer, ‘Curvature and the eigenvalues of the Laplacian’. J.
Differential Geom. 1/1 (1967), 43–69.
[292] R. B. Melrose, The Atiyah-Patodi-Singer index theorem, Research Notes in Mathematics,
vol. 4. A K Peters Ltd., Wellesley, MA, 1993.
[293] R. Melrose and P. Piazza, ‘Families of Dirac operators, boundaries and the b–calculus’.
J. Differential Geom. 46 (1997), 99–167.
[294] S. G. Mikhlin, Multidimensional singular integrals and integral equations, Translated from
the Russian by W. J. A. Whyte. Translation edited by I. N. Sneddon. Pergamon Press,
Oxford, 1965.
[295] J. Milnor, ‘Differentiable structures on spheres’. Amer. J. Math. 81 (1959), 962–972.
[296] — ‘On simply connected 4-manifolds’. In: Symposium internacional de topologı́a algebraica
International symposi um on algebraic topology. Universidad Nacional Autónoma de México
and UNESCO, Mexico City, 1958, pp. 122–128.
[297] — Morse theory, Based on lecture notes by M. Spivak and R. Wells. Annals of Mathematics
Studies, No. 51. Princeton University Press, Princeton, N.J., 1963.
[298] — ‘Eigenvalues of the Laplace operator on certain manifolds’. Proc. Nat. Acad. Sci. U.S.A.
51 (1964), 542.
[299] — Lectures on the h-cobordism theorem, Notes by L. Siebenmann and J. Sondow. Princeton
University Press, Princeton, N.J., 1965.
[300] J. Milnor and D. Husemoller, Symmetric bilinear forms. Springer-Verlag, New York,
1973, Ergebnisse der Mathematik und ihrer Grenzgebiete, Band 73.
[301] J. Milnor and J. D. Stasheff, Characteristic classes. Princeton University Press, Prince-
ton, N. J., 1974, Annals of Mathematics Studies, No. 76.
[302] S. Minakshisundaram and Å. Pleijel, ‘Some properties of the eigenfunctions of the
Laplace-operator on Riemannian manifolds’. Canadian J. Math. 1 (1949), 242–256.
BIBLIOGRAPHY 735

[303] A. S. Mishchenko, ‘The Hirzebruch formula: 45 years of history and the current state’.
Algebra i Analiz 12/4 (2000), 16–35, Russian, translation in St. Petersburg Math. J. 12/4
(2001), 519–533.
[304] P. K. Mitter and C.-M. Viallet, ‘On the bundle of connections and the gauge orbit
manifold in Yang-Mills theory’. Comm. Math. Phys. 79/4 (1981), 457–472.
[305] S. Mizohata, The theory of partial differential equations. Cambridge University Press, New
York, 1973, Translated from the Japanese by Katsumi Miyahara.
[306] R. Montgomery, ‘A new solution to the three-body problem’. Notices Amer. Math. Soc.
48/5 (2001), 471–481.
[307] G. Morchio and F. Strocchi, ‘Boundary terms, long range effects, and chiral symme-
try breaking’. In: Fields and Particles, Proceedings Schladming, Austria, Mitter, H., and
Schweifer, W. (eds.). Springer–Verlag, Berlin–Heidelberg–New York, 1990, pp. 171–214.
[308] — ‘Chiral symmetry breaking and θ vacuum structure of QCD’. Ann. Physics 324/10 (2009),
2236–2254.
[309] H. Moriyoshi and T. Natsume, Operator algebras and geometry, Translations of Mathe-
matical Monographs, vol. 237. American Mathematical Society, Providence, RI, 2008, Trans-
lated from the 2001 Japanese original by the authors.
[310] C. B. Morrey, Jr., Multiple integrals in the calculus of variations, Die Grundlehren der
mathematischen Wissenschaften, Band 130. Springer-Verlag New York, Inc., New York,
1966.
[311] D. Mumford, Algebraic geometry. I, Classics in Mathematics. Springer-Verlag, Berlin, 1995,
Complex projective varieties, Reprint of the 1976 edition.
[312] G. L. Naber, Topology, geometry, and gauge fields, second ed., Texts in Applied Mathe-
matics, vol. 25. Springer, New York, 2011, Foundations.
[313] M. Nakahara, Geometry, topology and physics, second ed., Graduate Student Series in
Physics. Institute of Physics, Bristol, 2003.
[314] R. Narasimhan, Analysis on real and complex manifolds, second ed.. Masson & Cie, Éditeur,
Paris, 1973, Advanced Studies in Pure Mathematics, Vol. 1.
[315] C. Nash, Differential topology and quantum field theory. Academic Press Ltd., London,
1991.
[316] V. E. Nazaı̆kinskiı̆, A. Y. Savin, B. Y. Sternin and B.-W. Schulze, ‘On the index of
elliptic operators on manifolds with edges’. Mat. Sb. 196/9 (2005), 23–58 (Russian), English
translation in Sb. Math. 196 (2005), no. 9-10, 1271-1305.
[317] R. Nest and B. Tsygan, ‘Algebraic index theorem’. Comm. Math. Phys. 172/2 (1995),
223–262.
[318] — ‘Formal versus analytic index theorems’. Internat. Math. Res. Notices 1996/11 (1996),
557–564.
[319] J. v. Neumann, ‘Allgemeine Eigenwerttheorie Hermitescher Funktionaloperatoren’. Math.
Ann. 102/1 (1930), 49–131.
[320] L. I. Nicolaescu, ‘The Maslov index, the spectral flow, and decomposition of manifolds’.
Duke Math. J. 80 (1995), 485–533.
[321] — Lectures on the geometry of manifolds. World Scientific Publishing Co. Inc., River Edge,
NJ, 1996.
[322] M. Ninomiya and C.-I. Tan, ‘Axial anomaly and index theorem for manifolds with bound-
ary’. Nucl. Phys. B257 (1985), 199–225.
[323] L. Nirenberg, ‘Pseudo-differential operators’. In: Global Analysis (Proc. Sympos. Pure
Math., Vol. XVI, Berkeley, Calif., 1968). Amer. Math. Soc., Providence, R.I., 1970, pp. 149–
167.
[324] F. Noether, ‘Über eine Klasse singulärer Integralgleichungen’. Math. Ann. 82/1-2 (1920),
42–63.
[325] B. Osgood, R. Phillips and P. Sarnak, ‘Extremals of determinants of Laplacians’. J.
Funct. Anal. 80/1 (1988), 148–211.
[326] Otte, M. et al. (ed.), Mathematiker über die Mathematik. Springer-Verlag, Berlin, 1974,
Wissenschaft und Öffentlichkeit.
[327] R. S. Palais, ‘Imbedding of compact, differentiable transformation groups in orthogonal
representations’. J. Math. Mech. 6 (1957), 673–678.
736 BIBLIOGRAPHY

[328] — Seminar on the Atiyah-Singer index theorem, With contributions by M. F. Atiyah, A.


Borel, E. E. Floyd, R.T. Seeley, W. Shih and R. Solovay. Annals of Mathematics Studies,
No. 57. Princeton University Press, Princeton, N.J., 1965.
[329] — Foundations of global non-linear analysis. W. A. Benjamin, Inc., New York-Amsterdam,
1968.
[330] J. Park and K. P. Wojciechowski, ‘Scattering theory and adiabatic decomposition of the
ζ-determinant of the Dirac Laplacian’. Math. Res. Lett. 9/1 (2002), 17–25.
[331] V. K. Patodi, ‘Curvature and the eigenforms of the Laplace operator’. J. Differential Geom-
etry 5 (1971), 233–249.
[332] G. K. Pedersen, Analysis now, Graduate Texts in Mathematics, vol. 118. Springer-Verlag,
New York, 1989.
[333] G. Perelman, The entropy formula for the Ricci flow and its geometric applications, 2002.
arXiv:[Link]/0211159[[Link]].
[334] — Finite extinction time for the solutions to the Ricci flow on certain three-manifolds,
2003. arXiv:[Link]/0307245[[Link]].
[335] — Ricci flow with surgery on three-manifolds, 2003. arXiv:[Link]/0303109[[Link]].
[336] M. E. Peskin and D. V. Schroeder, An introduction to quantum field theory. Addison-
Wesley Publishing Company Advanced Book Program, Reading, MA, 1995, Edited and with
a foreword by David Pines.
[337] M. J. Pflaum, ‘A deformation-theoretical approach to Weyl quantization on Riemannian
manifolds’. Lett. Math. Phys. 45/4 (1998), 277–294.
[338] — ‘The normal symbol on Riemannian manifolds’. New York J. Math. 4 (1998), 97–125
(electronic).
[339] L. S. Pontryagin, ‘A classification of continuous transformations of a complex into a sphere.
I, II.’. C. R. Acad. Sci. URSS (2) 19 (1938), 147–149, 361–363.
[340] — ‘Characteristic cycles on manifolds’. C. R. (Doklady) Acad. Sci. URSS (N.S.) 35 (1942),
34–37.
[341] — ‘Characteristic cycles on differentiable manifolds’. Mat. Sbornik N. S. 21(63) (1947),
233–284, Russian, English translation in [342].
[342] — ‘Characteristic cycles on differentiable manifolds’. Amer. Math. Soc. Translation 1950/32
(1950), 1–72, English translation of [341].
[343] — ‘A brief biographical sketch of L. S. Pontrjagin written by himself’. Uspekhi Mat. Nauk
33/6(204) (1978), 7–21, Russian, English translation in Russian Math. Surveys 33/6 (1978),
7–24.
[344] — Autobiography, posthumous, [Link] 1998, Russian,
298 pages.
[345] S. Prössdorf, ‘Über eine Algebra von Pseudodifferentialoperatoren im Halbraum’. Math.
Nachr. 52 (1972), 113–139.
[346] — Einige Klassen singulärer Gleichungen. Birkhäuser Verlag, Basel, 1974, Mathematische
Reihe, Band 46, Engl. translation North-Holland /Elsevier, New York, 1978.
[347] D. Przeworska-Rolewicz and S. Rolewicz, Equations in linear spaces. PWN—Polish
Scientific Publishers, Warsaw, 1968, Translated from the Polish by Julian Musielak, Mono-
grafie Matematyczne, Tom 47.
[348] D. Quillen, ‘Determinants of Cauchy-Riemann operators on Riemann surfaces’. Funkt-
sional. Anal. i Prilozhen. 19/1 (1985), 37–41, 96.
[349] F. Quinn, ‘Smooth structures on 4-manifolds’. In: Four-manifold theory (Durham, N.H.,
1982), Contemp. Math., vol. 35. Amer. Math. Soc., Providence, RI, 1984, pp. 473–479.
[350] J. V. Ralston, ‘Deficiency indices of symmetric operators with elliptic boundary condi-
tions’. Comm. Pure Appl. Math. 23 (1970), 221–232.
[351] D. B. Ray and I. M. Singer, ‘R-torsion and the Laplacian on Riemannian manifolds’.
Advances in Math. 7 (1971), 145–210.
[352] M. Reed and B. Simon, Methods of modern mathematical physics. I. Functional analysis.
Academic Press, New York, 1972.
[353] — Methods of modern mathematical physics. II. Fourier analysis, self-adjointness. Acad-
emic Press [Harcourt Brace Jovanovich Publishers], New York, 1975.
[354] S. Rempel and B.-W. Schulze, Index theory of elliptic boundary problems. North Oxford
Academic Publishing Co. Ltd., London, 1985, Reprint of the 1982 edition.
BIBLIOGRAPHY 737

[355] N. Reshetikhin, ‘Lectures on quantization of gauge systems’. In: New paths towards
quantum gravity (Extended lecture notes of the summer school, Holbæk, Denmark, May
12–16, 2008). Springer, Berlin, 2010, pp. 125–190. Lecture Notes in Phys., Vol. 807.
arXiv:1008.1411[math-ph].
[356] E. Roubine (ed.), Mathematics applied to physics. Springer-Verlag New York Inc., New
York, 1970, UNESCO, Paris.
[357] V. Rubakov, Classical theory of gauge fields. Princeton University Press, Princeton, NJ,
2002, Translated from the 1999 Russian original by Stephen S. Wilson.
[358] W. Rudin, Functional analysis. McGraw-Hill Book Co., New York, 1973, McGraw-Hill
Series in Higher Mathematics.
[359] Y. Safarov, ‘Pseudodifferential operators and linear connections’. Proc. London Math. Soc.
(3) 74/2 (1997), 379–416.
[360] N. Saglanmak and U. T. Jankvist, ‘What did they seek and what did they find? Combi-
natoric solutions for algebraic equations - from Cardano to Cauchy. I’. Normat 53/2 (2005),
54–72 (Danish), with English summary.
[361] F. Sannino, ‘Conformal dynamics for TeV physics and cosmology’. Acta Phys. Polon. B 40
(2009), 3533–3743. arXiv:0911.0931[hep-ph].
[362] Z. Y. Šapiro, ‘On general boundary problems for equations of elliptic type’. Izvestiya Akad.
Nauk SSSR. Ser. Mat. 17 (1953), 539–562.
[363] A. Y. Savin and B. Y. Sternin, ‘The index defect in the theory of nonlocal problems and
the η-invariant’. Mat. Sb. 195/9 (2004), 85–126.
[364] A. Y. Savin, B. Y. Sternin and B.-W. Schulze, ‘On invariant index formulas for spectral
boundary value problems’. Differ. Uravn. (translation in Differential Equations 35 (1999),
no. 5, 709–718) 35/5 (1999), 705–714, 720.
[365] M. Schechter, Principles of functional analysis, second ed., Graduate Studies in Mathe-
matics, vol. 36. American Mathematical Society, Providence, RI, 2002.
[366] T. Schick, Lectures on coarse index theory, Cortona, [Link]
Cortona2012/program/[Link], 2012.
[367] L. I. Schiff, Quantum mechanics. 2nd ed.. McGraw-Hill, 1968.
[368] J. Schmidt, ‘Chiral anomaly for the Dirac operator in an instanton background via the
Atiyah–Patodi–Singer theorem’. Phys. Rev. D (3) 36 (1987), 2539–2544.
[369] J. Schmidt and A. Bincer, ‘Chiral asymmetry and the Atiyah–Patodi–Singer index for the
Dirac operator on a four–dimensional ball’. Phys. Rev. D (3) 35 (1987), 3995–4000.
[370] L. Schwartz, Théorie des distributions. Tome I/Tome II, Actualités Sci. Ind., no.
1091/1122 = Publ. Inst. Math. Univ. Strasbourg 9/10. Hermann & Cie., Paris, 1950.
[371] A. S. Schwarz, ‘On the homotopic topology of Banach spaces’. Dokl. Akad. Nauk SSSR
154 (1964), 61–63, English translation in Sov. Math. Dokl. 5 (1964), 57-59.
[372] — ‘Instantons and fermions in the field of an instanton’. Comm. Math. Phys. 64 (1979),
233–268.
[373] R. L. E. Schwarzenberger, ‘Appendix One’. In: F. Hirzebruch, Topological Methods in
Algebraic Geometry (Berlin-Heidelberg-New York). Springer-Verlag, 1966, pp. 159–201.
[374] S. Scott, ‘Determinants of Dirac boundary value problems over odd–dimensional mani-
folds’. Comm. Math. Phys. 173 (1995), 43–76.
[375] — ‘Zeta determinants on manifolds with boundary’. J. Funct. Anal. 192/1 (2002), 112–185.
[376] — Traces and determinants of pseudodifferential operators, Oxford Mathematical Mono-
graphs. Oxford University Press, Oxford, 2010.
[377] S. Scott and K. P. Wojciechowski, ‘The ζ-determinant and Quillen determinant for a
Dirac operator on a manifold with boundary’. Geom. Funct. Anal. 10/5 (2000), 1202–1236.
[378] R. T. Seeley, ‘Integro-differential operators on vector bundles’. Trans. Amer. Math. Soc.
117 (1965), 167–204.
[379] — ‘Singular integrals and boundary value problems.’. Am. J. Math. 88 (1966), 781–809.
[380] — ‘Complex powers of an elliptic operator’. In: Singular Integrals (Proc. Sympos. Pure
Math., Chicago, Ill., 1966). Amer. Math. Soc., Providence, R.I., 1967, pp. 288–307.
[381] — ‘Elliptic singular integral equations’. In: Singular Integrals (Proc. Sympos. Pure Math.,
Chicago, Ill., 1966). Amer. Math. Soc., Providence, R.I., 1967, pp. 308–315.
[382] — ‘The resolvent of an elliptic boundary problem’. Amer. J. Math. 91 (1969), 889–920.
[383] — ‘Topics in pseudo-differential operators’. In: Pseudo-Diff. Operators (C.I.M.E., Stresa,
1968). Edizioni Cremonese, Rome, 1969, pp. 167–305.
738 BIBLIOGRAPHY

[384] G. Segal, ‘Equivariant K-theory’. Inst. Hautes Études Sci. Publ. Math. 34/1 (), 129–151.
[385] — ‘The definition of conformal field theory’. In: Topology, geometry and quantum field
theory, London Math. Soc. Lecture Note Ser., vol. 308. Cambridge Univ. Press, Cambridge,
2004, pp. 421–577.
[386] N. Seiberg and E. Witten, ‘Monopoles, duality and chiral symmetry breaking in N = 2
supersymmetric QCD’. Nuclear Phys. B 431/3 (1994), 484–550.
[387] H. Seifert and W. Threlfall, Seifert and Threlfall: a textbook of topology, Pure and
Applied Mathematics, vol. 89. Academic Press Inc. [Harcourt Brace Jovanovich Publishers],
New York, 1980, Translated from the German edition of 1934 by Michael A. Goldman, With
a preface by Joan S. Birman, With “Topology of 3-dimensional fibered spaces” by Seifert,
Translated from the German by Wolfgang Heil.
[388] J.-P. Serre, A course in arithmetic. Springer-Verlag, New York, 1973, Translated from the
French, Graduate Texts in Mathematics, No. 7.
[389] P. Shanahan, The Atiyah-Singer index theorem, Lecture Notes in Mathematics, vol. 638.
Springer, Berlin, 1978, An introduction.
[390] M. Shifman, Advanced topics in quantum field theory. Cambridge University Press, Cam-
bridge, 2012, A lecture course.
[391] M. A. Shubin, Pseudodifferential operators and spectral theory, Springer Series in Soviet
Mathematics. Springer-Verlag, Berlin, 1987, Translated from the Russian by Stig I. Ander-
sson.
[392] C. L. Siegel, Topics in complex function theory. Vol. I–III, Translated from the original
German by A. Shenitzer and D. Solitar. Interscience Tracts in Pure and Applied Mathemat-
ics, No. 25. Wiley-Interscience A Division of John Wiley & Sons, New York-London-Sydney,
1969, 1971 and 1973.
[393] B. Simon, ‘Notes on infinite determinants of Hilbert space operators’. Advances in Math.
24/3 (1977), 244–273.
[394] I. M. Singer, ‘Elliptic operators on manifolds’. In: Pseudo-Diff. Operators (C.I.M.E.,
Stresa, 1968). Edizioni Cremonese, Rome, 1968, pp. 333–375.
[395] — Operator theory and K-theory, 1970, Lecture presented in the Arbeitsgemeinschaft für
Forschung des Landes Nordrhein-Westfalen, unpublished.
[396] — ‘Future extensions of index theory and elliptic operators’. In: Prospects in mathematics
(Proc. Sympos., Princeton Univ., Princeton, N.J., 1970). Princeton Univ. Press, Princeton,
N.J., 1971, pp. 171–185. Ann. of Math. Studies, No. 70.
[397] — ‘Eigenvalues of the Laplacian and invariants of manifolds’. In: Proceedings of the In-
ternational Congress of Mathematicians (Vancouver, B. C., 1974), Vol. 1. Canad. Math.
Congress, Montreal, Que., 1975, pp. 187–200.
[398] — Personal communication, 1999, letter, unpublished.
[399] — ‘The η–invariant and the index’. In: Yau, S.–T. (ed.), Mathematical Aspects of String
Theory. World Scientific Press, Singapore, 1988, pp. 239–258.
[400] I. M. Singer and J. A. Thorpe, Lecture notes on elementary topology and geometry. Scott,
Foresman and Co., Glenview, Ill., 1967.
[401] S. Smale, ‘Generalized Poincaré’s conjecture in dimensions greater than four’. Ann. of Math.
(2) 74 (1961), 391–406.
[402] — ‘An infinite dimensional version of Sard’s theorem’. Amer. J. Math. 87 (1965), 861–866.
[403] E. H. Spanier, Algebraic topology. McGraw-Hill Book Co., New York, 1966.
[404] N. Steenrod, The topology of fibre bundles, Princeton Mathematical Series, vol. 14. Prince-
ton University Press, Princeton, N. J., 1951.
[405] A. Szankowski, ‘B(H) does not have the approximation property’. Acta Math. 147/1-2
(1981), 89–108.
[406] C. H. Taubes, ‘Gauge theory on asymptotically periodic 4-manifolds’. J. Differential Geom.
25/3 (1987), 363–430.
[407] — ‘Casson’s invariant and gauge theory’. J. Differential Geom. 31/2 (1990), 547–599.
[408] — ‘The Seiberg-Witten invariants and symplectic forms’. Math. Res. Lett. 1/6 (1994), 809–
822.
[409] — ‘More constraints on symplectic forms from Seiberg-Witten invariants’. Math. Res. Lett.
2/1 (1995), 9–13.
[410] M. E. Taylor, Pseudodifferential operators, Princeton Mathematical Series, vol. 34. Prince-
ton University Press, Princeton, N.J., 1981.
BIBLIOGRAPHY 739

[411] — Partial differential equations. I, Applied Mathematical Sciences, vol. 115. Springer-
Verlag, New York, 1996, Basic theory.
[412] — Partial differential equations. II, Applied Mathematical Sciences, vol. 116. Springer-
Verlag, New York, 1996, Qualitative studies of linear equations.
[413] R. Thom, ‘Quelques propriétés globales des variétés différentiables’. Comment. Math. Helv.
28 (1954), 17–86.
[414] E. Thomas, ‘Vector fields on manifolds’. Bull. Amer. Math. Soc. 75 (1969), 643–683.
[415] E. C. Titchmarsh, Introduction to the theory of Fourier integrals, third ed.. Chelsea Pub-
lishing Co., New York, 1986.
[416] D. Toledo and Y. L. L. Tong, ‘A parametrix for ∂ and Riemann-Roch in Čech theory’.
Topology 15/4 (1976), 273–301.
[417] H. F. Trotter, ‘Approximation of semi-groups of operators’. Pacific J. Math. 8 (1958),
887–919.
[418] K. K. Uhlenbeck, ‘Removable singularities in Yang-Mills fields’. Comm. Math. Phys. 83/1
(1982), 11–29.
[419] S. M. Ulam, A collection of mathematical problems, Interscience Tracts in Pure and Applied
Mathematics, no. 8. Interscience Publishers, New York-London, 1960.
[420] S. R. S. Varadhan, ‘Diffusion processes in a small time interval’. Comm. Pure Appl. Math.
20 (1967), 659–685.
[421] D. V. Vassilevich, ‘Spectral problems from quantum field theory’. In: Spectral geometry of
manifolds with boundary and decomposition of manifolds, Contemp. Math., vol. 366. Amer.
Math. Soc., Providence, RI, 2005, pp. 3–21.
[422] M. M. Vaynberg and V. A. Trenogin, Theory of branching of solutions of non-linear
equations. Noordhoff International Publishing, Leyden, 1974, Translated from the Russian
by Israel Program for Scientific Translations.
[423] I. N. Vekua, ‘Systems of differential equations of the first order of elliptic type and boundary
value problems, with an application to the theory of shells’. Mat. Sbornik N. S. 31(73) (1952),
217–314 (Russian), German translation in [424].
[424] — Systeme von Differentialgleichungen erster Ordnung vom elliptischen Typus und
Randwertaufgaben; mit einer Anwendung in der Theorie der Schalen, Mathematische
Forschungsberichte, II. VEB Deutscher Verlag der Wissenschaften, Berlin, 1956, German
translation of [423].
[425] — Generalized analytic functions. Pergamon Press, London, 1962.
[426] T. Voronov, ‘Quantization of forms on the cotangent bundle’. Comm. Math. Phys. 205/2
(1999), 315–336.
[427] C. T. C. Wall, ‘On simply-connected 4-manifolds’. J. London Math. Soc. 39 (1964), 141–
149.
[428] — ‘Non-additivity of the signature’. Invent. Math. 7 (1969), 269–274.
[429] A. H. Wallace, Differential topology: First steps. W. A. Benjamin, Inc., New York-
Amsterdam, 1968.
[430] N. R. Wallach, Harmonic analysis on homogeneous spaces. Marcel Dekker Inc., New York,
1973, Pure and Applied Mathematics, No. 19.
[431] S. Weinberg, The quantum theory of fields. Vol. I. Cambridge University Press, Cambridge,
2005, Foundations.
[432] — The quantum theory of fields. Vol. II. Cambridge University Press, Cambridge, 2005,
Modern applications.
[433] — The quantum theory of fields. Vol. III. Cambridge University Press, Cambridge, 2005,
Supersymmetry.
[434] R. O. Wells, Jr., Differential analysis on complex manifolds, second ed., Graduate Texts
in Mathematics, vol. 65. Springer-Verlag, New York, 1980.
[435] H. Weyl, The concept of a Riemann surface, Translated from the third German edition
by Gerald R. MacLane. ADIWES International Series in Mathematics. Addison-Wesley
Publishing Co., Inc., Reading, Mass.-London, 1964.
[436] — The classical groups, Princeton Landmarks in Mathematics. Princeton University Press,
Princeton, NJ, 1997, Their invariants and representations, Fifteenth printing, Princeton
Paperbacks.
[437] A. N. Whitehead, ‘The aims of education. A plea for reform’. Math. Gazette 8 (1916),
191–203, reprinted in [438].
740 BIBLIOGRAPHY

[438] — The aims of education and other essays. Macmillan, New York, 1929, 1985 paperback
reprint, Free Press.
[439] J. H. C. Whitehead, ‘On simply connected, 4-dimensional polyhedra’. Comment. Math.
Helv. 22 (1949), 48–92.
[440] H. Whitney, ‘The self-intersections of a smooth n-manifold in 2n-space’. Ann. of Math. (2)
45 (1944), 220–246.
[441] H. Widom, ‘A complete symbolic calculus for pseudodifferential operators’. Bull. Sci. Math.
(2) 104/1 (1980), 19–63.
[442] N. Wiener, Extrapolation, interpolation, and smoothing of stationary time series. With
engineering applications. The Technology Press of the Massachusetts Institute of Technology,
Cambridge, Mass, 1949.
[443] — The Fourier integral and certain of its applications, Cambridge Mathematical Library.
Cambridge University Press, Cambridge, 1988, Reprint of the 1933 edition, With a foreword
by Jean-Pierre Kahane.
[444] N. Wiener and E. Hopf, ‘Über eine Klasse singulärer Integralgleichungen.’. Sitzungsber.
Preuß. Akad. Wiss., Phys.-Math. Kl. 1931/30-32 (1931), 696–706 (German).
[445] F. Wilczek, Some problems in gauge field theories, 1977, Print-77-1029 (IAS, Princeton).
[446] E. Witten, ‘Some Exact Multi - Instanton Solutions of Classical Yang-Mills Theory’.
[Link]. 38 (1977), 121.
[447] — ‘Monopoles and four-manifolds’. Math. Res. Lett. 1/6 (1994), 769–796.
[448] S.-T. Yau (ed.), The founders of index theory: reminiscences of and about Sir Michael
Atiyah, Raoul Bott, Friedrich Hirzebruch, and I. M. Singer, second ed.. International Press,
Somerville, MA, 2009.
[449] B. Yood, ‘Properties of linear transformations preserved under addition of a completely
continuous transformation’. Duke Math. J. 18 (1951), 599–612.
[450] T. Yoshida, ‘Floer homology and splittings of manifolds’. Ann. of Math. (2) 134/2 (1991),
277–323.
[451] — Index theorem for Dirac operators. Kyoritsu Shuppan Publishers, 1998 (Japanese).
[452] K. Yosida, Functional analysis, fourth ed.. Springer-Verlag, New York, 1974, Die
Grundlehren der mathematischen Wissenschaften, Band 123.
[453] Y. Yu, The index theorem and the heat equation method, Nankai Tracts in Mathematics,
vol. 2. World Scientific Publishing Co. Inc., River Edge, NJ, 2001.
[454] E. Zeidler, Quantum field theory. I. Basics in mathematics and physics. Springer-Verlag,
Berlin, 2006, A bridge between mathematicians and physicists.
[455] — Quantum field theory. II. Quantum electrodynamics. Springer-Verlag, Berlin, 2009, A
bridge between mathematicians and physicists.
[456] — Quantum field theory. III. Gauge theory. Springer, Heidelberg, 2011, A bridge between
mathematicians and physicists.
[457] W. Zhang, Lectures on Chern-Weil theory and Witten deformations, Nankai Tracts in
Mathematics, vol. 4. World Scientific Publishing Co. Inc., River Edge, NJ, 2001.
[458] J. S. Zypkin, Adaption und Lernen in kybernetischen Systemen.. Oldenbourg Blg.,
München-Wien, 1970, Translated from the Russian.
Index of Notation

∆ Laplace operator Σ(M ), Σ± (M ) the Hermitian spinor


∆P generalized Laplace operator, 305 bundles, 531
Beltrami Laplacian, 189 Σ±c (X), Σc (X) bundles associated to Spin
c

connection (covariant) Laplacian on structure, 659


Ωk (W ), 439 Σ2m vector space of spinors, 521
connection Laplacian, 199, 536 Σ±2m vector space of half-spinors, 521
Dirac Laplacian, 187 Σc,g (X), Σ± c,g bundles of virtual twisted
Euclidean, 144 spinors relative to g, 696
Hodge Laplacian, 439 Σv1 ,...,vr singular set of array of vector
Γ(s) Gamma function, 112, 204 fields, 325
Γijk Christoffel symbols Θ torsion of a connection 1-form, 415

for connection, 178, 421 α : K(R2 × X) −→ K(X) Bott
for Riemannian metric, 168 isomorphism, 270
Λ• (V ) graded exterior algebra of V , 171, β(T ) symplectic space of all extensions of
514 closed symmetric T , 49
Λ• (T ∗ X) total bundle of exterior forms, β(A, B) Killing form, 462
172 χ Euler characteristic / class of a
Λp (T ∗ X) vector bundle of exterior forms, surface, 117
172 complex, 7
Λp (V ) vector space of p-fold complex vector bundle, 313
skew-symmetric tensors, 171 real Riemannian bundle, 453
Λs , ΛE,s generating operator for Sobolev topological manifold, 262, 320, 609
spaces, 197, 198 χ : U(1) × Spin(2m) → Spinc (2m) 2-fold
Ωk (P, W ) smooth equivariant W -valued cover, 659
k-forms on P , 401 χhol holomorphic Euler characteristic
Ωp,q (M ) forms of bidegree (p, q), 616 of a complex manifold = arithmetic
Ω0,k space of complex forms of type genus, 334
(0,k), 187 of a holomorphic vector bundle, 334, 642
k
Ω (P, W ) smooth horizontal equivariant δ
W -valued k-forms on P , 401 Bockstein homomorphism of homological
Ωω curvature of connection ω, 399 algebra, 655
Ω• (X) total space of exterior differential codifferential operator on forms, 188
forms, 172 difference bundle construction, 263
Ωp (X) space of exterior differential Dirac distribution, 38, 136
p-forms, 172 ∂X boundary of a manifold X, 143
Ω2± (M, R) spaces of (anti-) self-dual forms, δ ω covariant codifferential, 406
405 ε-tensor, 96
Φ Thom isomorphism of singular ηi , η i ’t Hooft matrices, 467
cohomology, 311, 313 ηD (s) eta function of Dirac type operator
Φ× Seiberg-Witten function, 665 D, 111
Φ isomorphism from horizontal equivariant ηeD reduced eta invariant, 112
forms to the gauge group, 411 γ gap (projection) metric for C(H), 52
Ψ Thom isomorphism of K-theory, 290, 313 γ5 global section, γ5 -matrix, 340
Ψ inverse of Φ, 411 κ general heat kernel, 547
, Σ(M ), Σ± (M ) the Hermitian spinor κ(z) Cayley transformation, 130
bundles, 531 κ ∈ Ω1,1 (M ) Kähler 2-form, 618

741
742 INDEX OF NOTATION

λ size of instanton, 475 |A| symmetric (absolute) factor of operator


λV canonical difference element of Λ• (V ) A, 56
for complex vector bundle V → X, 289 A atlas of charts, 157
µ(P ) boundary index of an elliptic A× group of units of Banach algebra A, 64,
operator, 249 65

ν, ∂ν , n field normal to the boundary, 154, A space of connections (affine
181, 191 configurations space), see C(P ), 461
νC normalized γ5 -matrix, 659 A(M
e ) generalized total characteristic class,
νg volume form of X with metric tensor g, 598
170, 405 A∗ fundamental vertical vector field of
[ω, ω] bracket of g-valued 1-form, 399 element A in Lie algebra, 398
ω connection 1-form on P , 398 ad adjoint representation of G on g, 397
ω symplectic form on smooth manifold, 652 Adg adjoint action of G on G, 397
ωC complex volume element, 520 ad derivative of ad at identity, 397
φ× associated vector bundle isomorphism, Ab A roof
409 A(F
b ) total A b class of real Riemannian
bundle F → M , 451
ϕ(a) autocorrelation function, 124
A(M
b ) := A(T
b M ) total A b class of
π : P → M principal G-bundle, 395
oriented Riemannian manifold M
π|P0 : P0 → M holonomy bundle, 456
with spin structure, 546
πk (X, x0 ) homotopy group
Af transmission (coupling) operator, 272
π1 (X, x0 ) fundamental group, 77, 128
Amplk (E, F ) space of amplitudes of order
πrc rc -equivariant bundle map defining a k, 234
Spinc (n) structure, 655 ant antipodal map
πE vector bundle base point projection, 181 on sphere, 274
ρ, ρ± , ρC spinor representations, 519–521 on tangent bundle, 317
ρ(X) injectivity radius of Riemannian A± pull-backs of θ± to R4 , 467
manifold X, 169 AR closed extension (realization) of
[σ(P )] symbol class, 279, 289 operator with domain R, 343
σ : B → U (E) ×f P radial gauge for fixed Aut(P ) group of all automorphisms of P ,
x ∈ M , 552 409
σi Pauli matrices, 467
σ(P ) principal symbol bundle B real elliptic operator associated to Spinc
homomorphism of structure, 666
P ∈ Lkpc (E, F ), 214, 228 B 2 (U ; Z2 ) group of Čech 2-coboundaries
differential operator, 184 with values in Z2 relative to the cover
σ(P )(x, ξx ) principal symbol U , 527
of P ∈ Lpc • (E, F ), 214, 217, 228 B, B± , B, 6 ∂ tangential Dirac operators, 341
of Bokobza-Haggiag amplitude, 234 B quotient space (space of moduli of all
of differential operator, 183, 216 connections), see M(P ), 461
σk (λ1 , . . . , λν ) elementary symmetric b Bott class, 270
polynomial of degree k in x1 , . . . , xν , b equivariant K-theory element, 292
448 bk Betti number, 7, 320
B(H) Banach algebra of bounded operators
τ involution on forms, 173
E parallel translation on a Hilbert space H, 3
τc,t
B+ convex set of positive operators, 17
along path c, 178
E b+ self-dual Betti number, 653
τx,x 0 parallel translation
Bf ideal of finite rank operators, 56
within injectivity radius, 179 BH closed unit ball in Hilbert space, 18
τE , τF local bundle trivialization, 183
(θij ) local Levi-Civita connection forms, C 2 (U ; Z2 ) group of Čech 2-cochains with
469 values in Z2 relative to the cover U ,
θ Levi-Civita connection, 180, 418 526
θ± decomposition of Levi-Civita C↓∞ (R) Schwartz space of rapidly
connection, 467 decreasing functions, 113
ϕ canonical 1-form on LM , 415 C× complex units, 253
ζP (s) zeta function of semi-positive elliptic Cj , j = 1, . . . , 4 SO(4)-irreducible
operator P , 114 decomposition for n = 4, 432
INDEX OF NOTATION 743

C : P → F M spin structure for Riemannian C(Y ) semi-group of equivalence classes of


manifold M , 525 complexes of vector bundles with
ċ(0) tangent vector of a path c, 160 compact support, 277
c (left) Clifford multiplication, 180, 340, C(Y, X) space of continuous mappings
531, 659 from Y to X, 18
c(E) total Chern class of complex vector
bundle E, 312
Dω⊕θ exterior derivative on
c : Spin(n) → SO(n) double covering
P ×G W -valued k-forms, 435
homomorphism, 517
D Dirac operator
ck (E) Chern class of complex vector bundle
D± (partial) chiral Dirac operators, 181,
E, 312, 444
532
C(P, G) smooth equivariant G-valued
Dc Spinc -Dirac operator, 659
functions on P , 410 (ω,L∗
ϕ θ2 )
C 0 (S 1 ) Banach space of continuous Dc lifted Spinc -Dirac operator,
C-valued functions on S 1 , 705 697
c1 : H 1 (X; U(1)) → H 2 (X; Z) Čech 6 D (free) Euclidean Dirac operator, 341
cohomology isomorphism, 656 6 DA twisted Euclidean Dirac operator,
C ∞ (Rn ) complex valued C ∞ function on 342
Rn , 135 operator of Dirac type, 180

C (T X) space of smooth vector fields, 165 standard Dirac operator for spin
C(H) space of closed densely defined structure, 531
operators in H, 52 twisted Dirac operator, 532
ch(E) Chern character, 260, 312, 445 DW , DW ± Dirac operator
CF(H) space of closed (not necessarily generalized Dirac operators, 597
odd
bounded) Fredholm operators, 52 DW ,E , D W ev ,E Dirac operator
C ∞ (Rn , CN ) C ∞ functions from Rn to Yang-Mills Dirac operators, 613
CN , 139 odd ev

DW , DW Dirac operator
C (X; E) linear space of smooth sections
generalized Euler operators, 608
of E, = C ∞ (E), 181
D divisor on Riemann surface, 335
Cl(T X) complex Clifford bundle, 595
(d + δ)+ signature operator, 321, 604
C`(X) Clifford bundle on Riemannian
(d + δ)E,+ twisted signature operator, 346,
manifold, 180
606
C`± (X) chiral Clifford bundles, 181
(d + δ)ev , (d + δ)χ Euler operator, 320, 609
cg Clifford multiplication relative to g, 697 ∂
Cauchy-Riemann operator, 187
C`(T Xx , gx ) Clifford algebras of tangent ∂z
vectors, 180 ∂ Dolbeault operator, 187, 334
C`(V ) Clifford algebra of real vector space ∂ E generalized Dolbeault operator, 334,
V , 514 639
Cl2m complex Clifford algebra, 520 ∂x21 + · · · + ∂x2n Euclidean Laplace
C`n Clifford algebra of real n space, 514 operator, 187
C0∞ smooth sections with compact degc degree of contribution, 579
support, 186 d exterior differentiation, 172
Coker T cokernel of the operator T , 3 d + δ deRham-Dirac operator, 320, 602
C(P ) space of C ∞ connections on P , 396 dk , dk virtual dimensions of moduli spaces
C(P )+ Mk , Mk of self-dual and anti-self-dual
space of self-dual connections, 489 connections for classes of principal
space of self-dual weakly irreducible SU(2)-bundles, 650
connections over fixed 4-manifold, Dα (symmetrized) partial derivative of
496 multiindex α, 135, 182
C(P )p,k Sobolev space of connections, 500 Dω covariant derivative relative to ω, 398
C(P )+m mildly-irreducible self-dual d(E • ), χ(E0 , E1 ; σ) difference bundle for
connections, 496, 507 complex E • , 278, 288
CP2 complex projective space of complex deg degree
dimension 2, 647 deg(P ) local index of an elliptic
CWP k Banach manifold of connections, operator, 249
parametrized by Sobolev space, 668 deg M+ boundary clutching degree of an
CN X trivial product bundle of complex fiber
elliptic operator, 249
dimension N over base space X, 181 deg(D) degree of divisor D, 335
744 INDEX OF NOTATION

deg(f ) winding number, mapping degree, Exps of Sobolev Banach manifolds at


254, 257–259 s ∈ W 2,k+2 (X, C), 670
det determinant ext natural extension in topological
π : S → F0 Segal determinant line K-theory, 285
bundle, 102
ζ-regularized determinant, 119
F correction operator between usual and
det α exterior determinant of α ∈ K(X),
connection Laplacian, 564
92
fb Fourier coefficient, integral, transform,
det F determinant line of Fredholm
706, 710
operator F , 693
f! : K(T X) → K(T N ) induced
det T exterior determinant of family of
homomorphism of smooth proper
Fredholm operators, 92
embedding f : X ,→ Y with normal
Fredholm determinant, 97, 98
bundle N , 290
q : Q → F Quillen determinant line
f∗ |x differential of f : X → Y ∈ C ∞ at
bundle, 95
x ∈ X, 161
Diff k (E, F ) space of linear differential
F (H) space of Fredholm operators on a
operators of order ≤ k, 182
Hilbert space H, 3
d∞ uniform metric, 18
F0 set of Fredholm operators of index 0, 88
dist distance, 168
f ∗ lifting (pull-back)
div divergence of a vector field, 185
of differential forms, 165
Dom(T ) domain of (not necessarily
of vector bundles, 713
bounded) operator T , 35
F · α action of GA(P ) on C(P ) or
k
Ω (P, W ), 410
E P ×G gc and/or all-purpose bundle, 491
Fi (X) = Hi (X; Z)/T i (X) essential
E spectral measure, 51
(torsion-free) homology, 658
E̊ dotted vector bundle, i.e., after removing
π : F M → M orthonormal frame bundle,
the zero-section, 182
416
E(r, t) fundamental solution of standard
F ± field strengths of A± , 467
heat equation in Euclidean n-space,
(FRk )c set of unsuitable pairs, 700
550, 563
FT affine Fredholm version of I1 , 104
E(c) energy functional, 168
F X(g) bundle of oriented orthonormal
E8 unimodular rank 8 even quadratic form
frames for the metric g, 694
of signature 8 (exceptional form), 325,
646
E (exceptional) lattice in R4k , 647 G Lie group of matrices, 394
e8 exceptional Lie algebra, 647 G(T ) graph of an operator T , 36
O(m),k
EllBokobza (E, F ) space of O(m)-invariant G gauge group, see GA(P ), 461
elliptic amplitudes of order k, 295 g Lie algebra of G, 397
EllkBokobza (E, F ) space of elliptic GA Grothendieck group of abelian
amplitudes of order k, 235 semi-group, 262
Ell(X) class of elliptic pseudo-differential GA1 PSpinc (n) subgroup acting on the
operators on closed manifold X, 265 solution space of perturbed S-W
Ellc (Rn ) class of elliptic pseudo-differential equations, 663
operators of order 0 in Rn being = Id GA(P ) group of C ∞ gauge transformations
at ∞, 275 of P , 409
Ellk (E, F ) space of elliptic principally GB(Ωθ ) Gauss-Bonnet form of a connection
classical pseudo-differential operators θ with curvature form Ω, 321, 446
of order k, 239 gC complexification of g, 464
Em m-times twisted line bundle defined g = (gij (x)) metric tensor, 167
over S 2 , 270 GL(N, C) Lie group of invertible complex
End(V ) algebra of C-linear endomorphisms N × N matrices, 42
of complex vector space V , 519 gl(N, C) Lie algebra of complex N × N
exp exponential map matrices, 41
exp(A) of matrix A in Lie algebra, 397 Grasssa (D) self–adjoint Grassmannian, 348
exp : C`(V ) → C`(V ), 514 Grassp+ Grassmannian of
expx of Riemannian manifold X at x, 169 pseudo-differential projections with the
Exp : C(P, g) → C(P, G), 413 same principal symbol, 345
INDEX OF NOTATION 745

g(S) genus (Klassenzahl) of surface S = index


numerical complexity of Abelian indexG P (analytic) index of an elliptic
integral, 319, 332 G-operator, 360
indexg P (analytic) virtual character, 360
H hyperbolic intersection form, 646 indext,G P topological G-index, 361
H k (U ; Z2 ) Čech cohomology group with indext,g P topological g-index, 361
values in Z2 relative to the cover U , indexa analytic index, 286, 291
527 indext topological index, 286, 290
H q (M ) deRham cohomology space, 603 index and index bundle of a continuous
H q (OE ) generalized Dolbeault cohomology family of Fredholm operators, 82, 84
groups, 334, 639 index of a Fredholm operator, 3
H 0,q (M ) Dolbeault cohomology space, 617 i(P2 , P1 ) virtual codimension of a Fredholm
H ∗ (X; R) cohomology functor with pair of projections, 345
coefficients in ring R, 266 I(X) group of stable equivalence classes of
HQ (x, y, t), GQ (x, y, t) approximative heat bundles, 264
kernels, 566, 567
Hc∗ (·) cohomology with compact supports, J (or G) Green’s function on boundary, 181
312 Jε mollifier, 196
H± discrete Hardy spaces, 125 J integration operator, 38
H algebra of quaternions, 514 J (M, ω) set of complex structures
H(Ωω ) harmonic part of Ωω , 691 compatible with symplectic form ω,
Hk (M ) space of harmonic k-forms, 602 652
H+ L2 -orthogonal projection onto the J k (E) k-jet bundle, 166
self-dual harmonic 2-forms relative to jk (f )x k-jet of section f at x, 166
g, 691 Jxk (E) k-jets of complex vector bundle E at
} reduced Planck constant, 190 x, 166
H (discrete) Hilbert transform, 127
K(x, y) weight (kernel) of integral
H+ (A) Cauchy data space, 344
operators, 212
Hk (C) homology space of a complex C, 7
+ K(A, B) ad-invariant inner product on g,
HN convex space of positive-definite 461
Hermitian matrices, 74 K(Π) Gaussian curvature, 423
Hol(ω, p0 ) holonomy group, 455 KQ,0 (x, y, t) heat kernel error term, 567
Hom(E, F ), Iso(E, F ) bundles of vector K ideal of compact operators, 18
bundle homomorphisms and k, k± twisted spinorial heat kernels, 541
isomorphisms, 713, 715 KC complexified Killing type form, 464
Hom0 (Σ2m , V ) space of Cl2m -equivariant Ker T kernel of an operator T , 3
linear maps, 523 KG (X) equivariant K-group, 293
Hp , Vp horizontal, vertical subspace of Tp P KO(m) (T ∗ S m ) equivariant K-group, 294
at p ∈ P , 396 Kϕ convolution operator, 129
HV Hopf bundle, 714 K(X) K-group of topological space X, 83,
263, 266
I(P0 ) index of a singular point, 252 K(X, Y ) relative K-group, 264, 267
I(p, ϕ)(x) oscillatory integral, 218 KS canonical divisor of Riemann surface S,
Iω gauge isotropy subgroup at ω, 456 336
= imaginary part, 331 KX canonical class of compatible almost
I1 ideal of trace class operators, 56 complex structure, 653
I2 ideal of Hilbert–Schmidt operators, 56
Ip Schatten class, 56 L canonical line bundle on Spinc -manifold,
Id identity operator, 9 659
I(E) Chern character defect, 314 L• (E, F ) space of canonical
Im(T ) image of an operator T , 3 pseudo-differential operators, 210
In n × n unit matrix, 165 L•Bokobza (E, F ) space of Bokobza-Haggiag
index pseudo-differential operators, 234
index(M, N ) index of a Fredholm pair L•pc (E, F ) space of principally classical
(M, N ) of closed subspaces, 344 pseudo-differential operators, 214, 216,
indext topological index, 287 224, 225, 229
indexO(m) v O(m)-character of Lϕ equivariant transformation of oriented
equivariant K-class v, 294 orthonormal frames, 695
746 INDEX OF NOTATION

Lq (p1 , . . . , pq ) Hirzebruch L-polynomials, Euclidean construction, 210


324 patched global construction, 229
L(F ) total Hirzebruch L class, 451 OPp,1 Banach space of bounded operators
L linear isomorphism Λ• (V ) → C`(V ), 514 W p,k+1 (E) → W p,k E ⊗ Λ1 (X) , 502


L(c) length functional, 168


L1 (S 1 ) Banach space of C-valued P Poincaré duality, 644
integrable functions on S 1 , 705 P ×f F M fibered product, 434
L2 (E) = L2 (X; E) Hilbert space of P ± chiral splitting of an operator P , 47
Lebesgue measurable square integrable PV projective space of complex vector
sections in a Hermitian vector bundle space V , 714
E on a Riemannian manifold X, 194 p(x, ξ) total symbol/amplitude
L2 (S 1 ) Hilbert space of square-integrable (dequantization), 218
C-valued functions on S 1 , 705 P(T X) set of all tangent lines over all
L2 (X) Hilbert space of complex valued points on a manifold X, 164
square integrable functions on X, 10 P≥ (B) spectral (Atiyah–Patodi–Singer)
L2 (Z) = `2 Hilbert space of square projection, 345
summable series, 4 P+ (A) Calderón projection, 344
Lksym (V, W ) space of symmetric k-linear P ∗ formally adjoint operator, 184
forms on V with values in W , 166 Pf(A) Pfaffian of A ∈ so(2m), 446
LA B = [A, B] Lie bracket of vector fields P ×G W associated vector bundle, 400
A, B, 174 pk Pontryagin
L(f ) Lefschetz number, 357 pk (Ωθ ) forms, 546
L(f, P ) Atiyah-Bott-Lefschetz number, 358 pk (M ) class, 324, 447
Lg , `g left action of g, 292, 360, 397
[Q] matrix of quadratic form Q, 646
LM → M bundle of linear frames, 414  
Lp (E) Lp sections of E → M , 497 q : C ∞ Σ+c (X) → Ω
2+ (X, iR) quadratic

map, 661
Mη moduli space for Spinc structure, 664 qdk̄ (X) Donaldson’s polynomial invariants,
[M ] fundamental cycle of oriented manifold 650
M , 313 QX intersection form on H 2 (X; Z), 643
M space of moduli of connections, 489
M+ space of moduli of self-dual R(W, Z, X, Y ) curvature tensor, 422
connections, 489 R vector space of curvature tensors, 427
moduli of C(P )+ m , 510 Rε transformation of twisted spinors, 537
M± eigenspaces of boundary symbol, 245 Rω , R± correction forms for
M space of moduli of self-dual connections, Spinc -Dirac-Laplacian, 660, 661
see M+ (P ), 461 r : R(Rn ) −→ S(Rn ) Ricci map, 427
m mass of a particle, 190 rc = Sq × c : Spinc (n) → U(1) × SO(n)
m : Z → Z multiplication by 2, 655 double-covering homomorphism, 655
MC multiplicative class assigned to a R1 , R2 , R3 O(n)-irreducible
characteristic class, 449 decomposition of R, 428
MCWP k quotient space of moduli, 673 Rα Riesz (singular integral) operator, 212
MCWPS k MCWPS , MSWPS k r∗ : H 2 (X; Z) → H 2 (X; Z2 ) homology
Hausdorff C ∞ Hilbert (quotient) reduction mod 2, 649
manifolds, 698 < real part, 20
Met(M ) convex set of metric tensors, 652 Res (T ) resolvent set of the operator T , 10,
MF multiplicative form, 449 50
Mid spectral multiplication operator, 36 R(λ) resolvent function, 50
M1 #M2 connected sum, 645 Rg right action of g on P , 396
R(H) representation (Grothendieck) ring of
O Cayley numbers, 607 group H, 293, 360
O, o Big Oh, Small Oh, 549 Ric(X, Y ) Ricci tensor, 426
OE generalized Dolbeault complex, 639
OPk (E, F ) space of operators of order k, S scalar curvature, 426
237 S(Y ) suspension of topological space Y ,
Op(p) pseudo-differential operator 718
• principally classical
Spc
(quantization) of amplitude p
Bokobza-Haggiag (global) construction, symbols/amplitudes, 214
231, 233, 234 Sω slice of action through ω, 489
INDEX OF NOTATION 747

S symmetric bilinear forms, 427 T operator, Ω1 (E) → Ω0 (E) ⊕ Ω2− (E), 491
S(M ) set of inequivalent spin structures on T i (X) torsion subgroup of H i (X; Z), 657
manifold M , 528 T algebra of discrete Wiener-Hopf
S• (U × Rn ) symbols of Hörmander type operators, 132
(1, 0), 210 T (X, Y ) torsion tensor, 179
s · ω infinitesimal action of C(P, g) on C(P ), T 2 2-dimensional torus, 142
413 T ∗ X differential, = ϕ∗x , 164
s : R(Rn ) → R scalar map, 427 Td(E) Todd class of a complex vector
Sg+1 X bundle of positive g1 -symmetric bundle E, 312, 450
operators, 694 Tf discrete Wiener-Hopf operator, 69, 125,
sgn(σ) sign of permutation σ, 171 255
shift± shift operators, 4 T n n-dimensional torus, 201
sig signature Tr A, tr A trace of A ∈ I1 , 58
of a quadratic form Q, 320, 646 Tr R trace of curvature tensor R, 427
of a topological manifold X, 320 T r,s P ×G W -valued tensors, 434
sk homogeneous polynomial of degree k, Tx X tangent space of X at x, 160
443
U(H) group of unitary operators, 73
sk (Ωω ) Chern form, 444
U(N ) group of unitary N × N matrices, 73
Smblk (E, F ) space of k-homogeneous
bundle homomorphisms on T̊ ∗ X, 184, Vλ eigenspace, 538
225 V ∗ dual vector space, 171
SO(n) special orthogonal group, 396 Vect(X) abelian semi-group of isomorphism
so(n) Lie algebra of antisymmetric n × n classes of complex vector bundles over
matrices, 515 X, 83, 262, 716
Sol(η) set of all C ∞ solutions of the S-W vol(X) volume of Riemannian manifold X,
equations perturbed by η, 692 170
span = [. . . ] linear span of vectors, 21, 400
Spec spectrum, 10, 50 W Sobolev spaces
Specc continuous, 10 W k modeled on L2 , 35
Spece essential, 10 W s (E) bundle sections, 197
Specp point (discrete), 10 W s (Rn ) Euclidean, 195
Specr residual spectrum, 10 W m (Rn + ) over half spaces, 200
s compactly supported, 198
spec(A, A∗ ) domain for spectral invariance WK
of operators belonging to a set A, 110 W p,k modeled on Lp
Spin(n) spin group, 515 W p,k (E) bundle sections, 497
Spinc W (M ) Clifford module bundle, 595
Spinc structure for an oriented W (f, 0) winding number
Riemannian n-manifold, 655 see deg(f ), 69, 125
Spinc (n) nth Spinc -group, 655 W ± (anti-) self-dual part of R3 , 433
spin(n) spin algebra, 515 w2 (M ) second Stiefel-Whitney class of M ,
Str supertrace 527
of heat kernel k, 544 Wf continuous Wiener-Hopf operator, 131,
of spinor endomorphism A, 522 256
str(W s (X)) strength of a Sobolev space,
X + 1-point compactification of locally
201
compact X, 266
SU(n) special unitary group, 395
X̊ interior of a manifold with boundary, 184
supp support, 136
[X, Y ] homotopy set, 72
SW : H 2 (X; Z) −→ Z Seiberg-Witten
invariants (relative to one fixed Spinc YM(ω) Yang-Mills functional, 462
structure),
  652
SW PSpinc Seiberg-Witten invariant, Z 2 (U ; Z2 ) group of Čech 2-cocycles with
693 values in Z2 relative to the cover U ,
SW ∞ (η) set of C ∞ solutions of the S-W 527
equations modulo C ∞ (X, U(1)), 691 Zn n-dimensional lattice, 140
SW k (η) smooth submanifold of MSWPk , (·, ·) L2 -inner product on Ωk (P ×G W ), 404
680 (·, ·)0 inner product in L2 (X; E), 170
SWP k parametrized solution space, 669 < ·, · > inner product in Hilbert space, 3
748 INDEX OF NOTATION

< ·, · >h Hermitian metric on vector ∨ : S 2 (Rn ) × S 2 (Rn ) → R(Rn ) vee


bundle, 170 (Kulkarni-Nomizu) product, 428
[·, ·] bracket (commutator), 515 ∧ exterior multiplication (wedge product),
[. . . ] = span(. . . ) linear span of vectors, 21, 171
400 ·∧ Bokobza-Haggiag (global) Fourier
∗ star transform of section u on Riemannian
convolution, 706, 708, 711 X in Hermitian bundle with fixed
Hodge star operator on forms, 172, 405 connection and fixed bump function,
star operator on exterior algebra, 172 232
taking the adjoint operator, 15 ·x (left) interior multiplication, 173, 519
 outer tensor product, 243, 261, 267, 268, ·∧ (left) exterior multiplication, 173, 519
278, 279
∪f gluing
two manifolds by a diffeomorphism of
their boundaries, 191
two vector bundles by an isomorphism
over the intersection of their bases,
718
|α| degree of multiindex α, 135
k·k norm
L2 -norm on C ∞ (P ×G W ), 405
|·|e norm relative ds2 , 474
|·|h norm relative h = f 2 ds2 , 474
k·k1 trace norm on I1 , 60
k·kk Sobolev norm, 37
k·kk,K semi norm on C ∞ (Rn ), 136
k·kp,k Lp Sobolev norm, 497
operator norm, 3
Sobolev norm, 195
h·, ·i C ∞ (M )-valued pairing on
C ∞ (P ×G W ), 404
N
A·, · B pairing
H q (X; Z) × Hq (X; Z) → Z, 644
ˆ Fourier transform
Fourier integral, 26, 708
Fourier series, 706
˚dotted
for vector bundles, 182
interior of manifold with boundary, 184
∇ nabla operator
∇E connection (covariant differentiation
operator) on vector bundle E, 177
∇2 invariant second derivative, 536
∇ω⊕θ covariant derivative on T r,s (W ),
435
∇u gradient of function u, 145
⊕ direct sum, 5
⊗ tensor products, 166
⊥ orthogonal, 13
t transversality, 700
∼ asymptotic expansion, 549
∼ homotopy of maps, homotopy
equivalence, 72, 716
` cup product, 311
× Clifford multiplication
see also c, 328
(·∧ )∨ Bokobza-Haggiag inverse Fourier
transform, 233
Index of Names/Authors

Cursive numbers refer to the bibliography.

Most given names and personal data are from Indexes of Biographies
of selected mathematicians at
[Link]

Abel, Niels Henrik (1802–1829), 329–333, 511, 512, 718, 719, 722, 723–725, 734,
723 736
Abraham, Ralph H. (*1936), 253, 723 Atkinson, Frederick Valentine (1916–2002),
Adams, John Frank (1930–1989), 210, 261, xvii, 64, 65, 70, 101, 725
326 Avramidi, Ivan, 369, 726
Adams, Robert A., 194, 201, 723
Agranovich, Mikhail Semenovich (*1930), Bachmann, Paul Gustav Heinrich
xvii, 3, 337, 355, 723 (1837–1920), 549
Ahlfors, Lars Valerian (1907–1996), 148, Banks, Tom, 373, 725
256, 723 Bauer, Wolfram, 118, 725
Akhiezer, Naum Il’ich (1901–1980), 20, 22, Baum, Paul Frank (*1936), xviii
49, 50, 723 Belavin, Alexander ”Sasha” Abramovich
Albin, Pierre, 356 (*1942), 466, 725
Bellman, Richard Ernest (1920–1984), 309
Alexander III of Macedon (356–323 BCE),
Benedetto, John J., 309, 725
176
Berg, I. David, 109
Alexandrov, Pavel Sergeyevich
Berger, Marcel, 118, 170, 725
(1896–1982), 320, 357, 723
Berline, Nicole, 138, 284, 725
Alonso, Alberto, 49, 338, 723
Bernard, Claude W., 393, 466, 725
Alvarez-Gaumé, Louis, 373, 723
Bernoulli, Daniel (1700–1782), 12
Ambjørn, Jan, 726 Bernoulli, Jacob (1654–1705), 329
Aristotle (384–322 BCE), 176 Bernoulli, Johann (1667–1748), 29, 330
Arnold, Vladimir Igorevich (1937–2010), Bers, Lipman (1914–1993), 194, 201, 725
217 Bincer, Adam M. 352, 737
Aronszajn, Nachman (1907–1980), 667, 723 Birkhoff, George David (1884–1944), 715
Arzelà, Cesare (1847–1912), 18 Birman, Mikhail Šlemovic (1928–2009), 338
Ascoli, Giulio (1843–1896), 18 Bismut, Jean-Michel (*1948), xx
Atiyah, Michael Francis, Sir (*1929), Blaschke, Wilhelm Johann Eugen
xvii–xix, 2–4, 70, 71, 82, 86, 88, 112, (1885–1962), 2
129, 138, 199, 208, 210, 225, 227, 232, Bleecker, David D. (*1948), xxi, xxii, 29,
243, 245, 249, 250, 252, 253, 257, 261, 49, 113, 118, 175, 369, 371, 372, 405,
262, 266–269, 272–274, 278–281, 667, 707, 708, 710, 725, 726
283–285, 287, 291, 295, 299, 301–304, Boéchat, Jacques, 162, 725
306–308, 310–313, 316, 318, 319, 327, Bochner, Salomon (1899–1982), 439, 442,
336–339, 345, 349, 350, 354–357, 359, 725
361, 362, 466, 467, 482, 488, 491, 508, Bohn, Michael, 101, 725

749
750 INDEX OF NAMES/AUTHORS

Bohr, Harald August (1887–1951), xxii Connes, Alain (*1947), 302


Bojarski, Bogdan (*1931), 301, 337, 344, Cordes, Heinz Otto (*1925), 44, 46, 52, 727
356, 725 Corrigan, Francis Edward ”Ed” (*1946),
Bokobza-Haggiag, Juliane, 231, 236, 725 393, 466, 727
Bolzano, Bernard Placidus Johann Coughlan, Guy D., 385, 727
Nepomuk (1781–1848), 18 Courant, Richard (1888–1972), xvii, 29, 30,
Bonnet, Pierre Ossian (1819–1892), 321, 33, 66, 147, 190, 727
324 Csordas, George, 29, 113, 708, 725
Booß–Bavnbek, Bernhelm (*1941), xx, xxii,
47, 49, 52, 53, 108, 118, 155, 192, 200, Davis, Philip J. (*1923), 331, 727
214, 239, 244, 246, 249, 261, 272, 275, Dedekind, Julius, Wilhelm Richard
282–284, 287, 303, 308, 319, 339–341, (1831–1916), 336
344, 347–351, 354–356, 364, 667, 707, Deligne, Pierre René, Viscount (*1944), 734
710, 725, 726, 730 Devinatz, Allen (1922–2008), 130, 131, 727
Borel, Armand (1923–2003), xviii, 304, 726, DeWitt, Bryce Seligman (1923–2004), 365,
731, 736 728
Bott, Raoul (1923–2005), xviii, xix, 3, 77, Fagnano dei Toschi, Giulio Carlo di
129, 138, 225, 243, 245, 249, 250, 252, (1682–1766), 330
257, 261, 269, 273, 302–304, 307, Dieudonné, Jean Alexandre Eugène
336–338, 357, 359, 362, 718, 722, 724 (1906–1992), xvii, 68, 69, 266, 337, 728
Boutet de Monvel, Louis (*1941), 120, 283, Dirac, Paul Adrien Maurice (1902–1984),
338, 339, 726 131, 190, 382
Dirichlet, Johann Peter Gustav Lejeune
Bouwknegt, Peter, 726
(1805–1859), 145
Branson, Thomas P. (1953–2006), 180, 726
Dixmier, Jacques (*1924), 81, 728
Bredon, Glen E., 16, 157, 162, 171, 726
Dodd, James Edmund (*1952), 385, 727
Breuer, Manfred (1939–2011), 4, 726
Dolbeault, Pierre (*1924), 334
Brieskorn, Egbert Valentin (*1936), 3, 262,
Donaldson, Simon Kirwan (*1957), 156,
302, 325, 327, 726, 727
162, 310, 325, 460, 512, 643, 648,
Brill, Alexander Wilhelm von (1842–1935),
650–652, 663, 728
329–331, 333, 727
Douady, Adrien (1935–2006), 81
Brouwer, Luitzen Egbertus Jan
Douglas, Ronald George (*1938), 69, 74,
(1881–1966), 273
108, 109, 111, 128, 132, 350, 727, 728
Brown, Lawrence G., 108, 109, 111, 727
Drinfeld, Vladimir Gershonovich (*1954),
Bröcker, Theodor (*1938), 157, 192, 229,
467, 482, 488, 512, 724
256, 259, 317, 325, 714, 719, 720, 727
Dugundji, James (1919–1985), 23, 159, 716,
Brüning, Jochen (*1947), 49, 244, 340, 727
728
Bunke, Ulrich (*1963), 308, 727
Duhamel, Jean Marie Constant
(1797-1872), 339, 350, 728
Calderón, Alberto Pedro (1920–1998), 208,
Duistermaat, Johannes Jisse (Hans)
244, 301, 304, 340, 344, 727
(1942–2010), 117, 728
Calkin, John Wilson, 63, 727
Dunford, Nelson (1906–1986), 51, 728
Cantor, Murray, 497–499, 698, 727
Dupont, Johan Louis, 310, 327, 357, 724
Carleman, Torsten (1892–1949), 3, 305,
Dym, Harry, 11, 27, 120, 123, 126, 132,
306, 339, 727
227, 705–711, 728
Cartan, Henri Paul (1904–2008), 2, 70, 136,
Dynin, Alexander S., xvii, 3, 355, 723
239, 243, 245, 249, 302, 727
Dynkin, Eugene Borisovich (*1924), 140,
Casorati, Felice (1835–1890), 67
142, 728
Cauchy, Augustin-Louis (1789–1857), 29
Chen, Guoyuan, 214, 239, 354, 356, 726 Egorov, Yuri V. (*1933), 217
Chern, Shiing-shen (1911–2004), xxii, 2, Eguchi, Tohru, 467, 728
321, 363 Eichhorn, Jürgen (*1942), xx, 728
Chevalley, Claude (1909–1984), 265 Eilenberg, Samuel (1913–1998), 8, 71, 259,
Cho, Adrian, 385, 727 266, 267, 278, 728
Christ, Norman Howard (*1943), 466, 725 Einstein, Albert (1879–1955), xx, 370, 373
Clebsch, Rudolf Friedrich Alfred Eisenhart, Luther Pfahler (1876–1965),
(1833–1872), 332, 333 431, 728
Coburn, Lewis A., 108 Elizalde, Emilio, 118, 728
Coddington, E.A., 20, 29, 32, 51, 144, 727 Enflo, Per H. (*1944), 24, 728
INDEX OF NAMES/AUTHORS 751

Esposito, Giampiero (*1962), 118, 364, 365, Griffiths, Phillip Augustus (*1938), 189,
726, 728, 730 620, 648, 730
Euler, Leonhard (1707–1783), 319, 330 Grigis, Alain, 165, 194, 196, 211, 216, 218,
223, 730
Fairlie, David B., 466, 727 Grothendieck, Alexander (*1928), xviii, 24,
Feldman, Israel A., 132, 144, 730 265, 304, 715, 726, 730
Fermat, Pierre de (1601–1665), 29 Grubb, Gerd (*1939), 107, 119, 194, 196,
Feynman, Richard Phillips (1918–1988), 201, 208, 211, 213, 218, 275, 282, 340,
284, 388 345, 348, 356, 730
Fillmore, Peter A., 108–111, 132, 727, 729 Guentner, Erik, xviii, 730
Fischer, Ernst Sigismund (1875–1954), 15 Guillarmou, Colin, 356, 730
Floyd, Edwin E. (1924–1990), 736 Guillemin, Victor W. (*1937), 117, 173,
Fredholm, Erik Ivar (1866–1927), xvii, 25, 185, 309, 728, 730
27, 96, 209, 729 Guth, Alan Harvey (*1947), 466, 725
Freed, Daniel S. (*1959), 512, 729
Gårding, Lars (*1919), 241, 347
Freedman, Micheal Hartley (*1951), 257,
325, 646, 648, 651, 729
Haack, Wolfgang Siegfried (1902–1994),
Friedman, Robert, 651, 652, 729
147, 730
Fubini, Guido (1879–1943), 25, 706, 710
Haefliger, André (*1929), 162, 725, 730
Fujiwara, Daisuke, 304, 729
Hansen, Vagn Lundsgaard (*1940), 156,
Fulling, Stephen A., 231, 729
157, 730
Fuquan, Fang, 162, 729
Hanson, Andrew J., 467, 728
Fursaev, Dmitri, xx, 118, 729
Harris, Joseph (Joe) Daniel (*1951), 189,
Furuta, Mikio, xx, 163, 210, 357, 650, 652,
620, 648, 730
729
Hartman, Philip (*1915), 144, 731
Furutani, Kenro (*1949), 49, 95, 101, 118,
Hawking, Stephen William (*1942), 369,
340, 725, 726, 729
729
Gauss, Johann Carl Friedrich (1777–1855), Heaviside, Oliver (1850–1925), 190
321, 324 Hellinger, Ernst David (1883–1950), 12, 19,
Gelfand, Israil Moiseevic (1913–2009), xvii, 731
3, 131, 307, 308, 337, 338, 729 Hellwig, Günter (1926–2004), xvii, 47, 146,
Getzler, Ezra (*1962), 138, 231, 284, 301, 147, 152, 337, 731
302, 725, 729 Helton, J. William, 108
Ghorbani, Hossein, xx, 364, 729 Hermann, Robert (*1931), 190, 731
Gibbons, G.W., 369, 729 Higson, Nigel, xix, 290, 731
Gilkey, Peter Belden (*1946), xviii, xix, 49, Hilbert, David (1862–1943), xvii, 1, 19, 20,
111, 112, 115, 116, 118, 138, 180, 284, 23, 25–27, 29, 30, 33, 66, 133, 147, 190,
302, 305, 307, 467, 721, 726, 728, 729 209, 239, 251, 305, 715, 727
Giraud, Georges (1889–1943), 208, 729, 730 Hirsch, Morris William (*1933), 162, 252,
Glashow, Sheldon Lee (*1932), 385 256, 257, 259, 720, 730, 731
Glazman, Izrail Markovic, 20, 22, 49, 50, Hirzebruch, Friedrich Ernst Peter
723 (1927–2012), xvii, xix, xx, xxii, 2, 18,
Glimm, James Gilbert (*1934), 389, 390, 71, 77, 163, 258, 260, 262, 266, 278,
730 302–304, 310, 311, 321, 325, 331, 333,
Goddard, Peter (*1945), 393, 727 334, 336, 449, 657, 721, 722, 731
Gohberg, Israel (1928–2009), xvii, 3, 20, 43, Hitchin, Nigel (*1946), xix, 338, 466, 467,
44, 49, 131, 132, 144, 239, 269, 304, 482, 488, 491, 508, 511, 512, 655, 724,
730 731
Goldberg, Seymour, 20, 620, 730 Hodge, William Vallance Douglas
Gompf, Robert E., 651, 730 (1903–1975), 2, 322
Gorbachuk, Myroslav L., 338, 730 Holden, Helge, 731
Gordon, Carolyn S., 118, 730 Hopf, Eberhard Friedrich Ferdinand
Gotay, Mark J., 654 (1902–1983), 121, 132, 740
Gracia-Bondı́a, José M., 365, 726, 730 Hopf, Heinz (1894–1971), 262, 274, 320,
Graham, Ronald Lewis (*1935), 549, 730 327, 357, 657, 714, 721, 723, 731
Green, George (1793–1841), 26 Horowitz, Philip, 723
Greenberg, Marvin Jay, 8, 320, 357, Howe, Roger Evans (*1945), 108
643–645, 647, 658, 730 Hsiang, Wu-Yi (*1937), 363
752 INDEX OF NAMES/AUTHORS

Hsu, Elton P., xx, 732 Knuth, Donald Ervin (*1938), 549, 730
Humphreys, James Edward (*1939), 647, Kobayashi, Shoshichi (1932–2012), 171,
732 175, 448, 456, 721, 732
Husemöller, Dale Harper (Husemoller), Kodaira, Kunihiko (1915–1997), 2, 163,
(*1933), 645–648, 732, 734 310, 333, 337, 648, 733
Huygens, Christiaan (1629–1695), 217 Kolmogorov, Andrey Nikolaevich
Hörmander, Lars (1931-2012), 131, 138, (1903–1987), 124, 733
143, 146, 148, 151, 194, 196, 198, 200, Kolář, Ivan, 175, 733
201, 207, 208, 210, 213, 216–218, 225, Kontsevich, Maxim (*1964), 107, 733
226, 240, 241, 243, 249, 308, 709, 711, Kori, Tosiaki (*1941), 353, 733
731, 732 Koschorke, Ulrich (*1941), 310, 733
Kotake, Takeshi, 301, 305, 306, 733
Illusie, Luc (*1940), 77, 80, 732
Kreck, Matthias (*1947), 325, 331, 333,
Ivakhnenko, Alekseǐ Grigorèvich
334, 723, 731
(1913–2007), 124, 732
Kreĭn, Mark Grigorievich (1907–1989), xvii,
Ivanova, Raina, 721, 729
3, 43, 44, 49, 130, 131, 269, 304, 338,
Iwasaki, Chisato, 118, 725
730, 733
Ize, Jorge, 141, 144, 732
Krickeberg, Klaus (*1929), 140, 142, 728
Jackiw, Roman W. (*1939), 466, 482, 732 Kronecker, Leopold (1823–1891), 256, 733
Jacobi, Carl Gustav Jacob (1804–1851), Kronheimer, Peter Benedict (*1963), xix,
309, 331 xx, 310, 460, 512, 643, 651, 652, 654,
Jaffe, Arthur (*1937), 389, 390, 730 728, 733
Jahnke, Hans Niels (*1948), 3, 735 Kuiper, Nicolaas Hendrik (1920–1994), 75,
Jankvist, Uffe Thomas, 329, 737 77, 81, 82, 85, 88, 129, 733
Jessen, Børge Christian (1907–1993), xxii Kummer, Ernst Eduard (1810–1893), 648
John, Fritz (1910–1994), 194, 201, 725 Kuranishi, Masatake, 217, 221, 223, 224,
Jones, Vaughan Frederick Randal (*1952), 229–231, 243
77 Kähler, Erich (1906–2000), 334, 335, 453,
Jost, Jürgen (*1956), 167, 169, 732 618, 622, 624, 633, 648, 654
Jurkiewicz, Jerzy, 726
Yushkevic (Juschkewitsch), Alexander A., Labrousse, Jean-Philippe, 44, 46, 52, 727
140, 142, 728 Lagrange, Joseph-Louis (1736–1813), 65,
Jänich, Klaus Werner (*1940), 3, 77, 88, 733
157, 192, 229, 256, 259, 317, 325, 714, Lang, Serge (1927–2005), 499, 723, 733
719, 720, 727, 732
Lapa, Valentin Grigorievic, 124, 732
Jörgens, Konrad (1926–1974), 11, 19, 24,
Lawson, H. Blaine, Jr., xviii, xix, 231, 232,
25, 27, 31, 32, 70, 86, 130, 131, 732
235, 284, 289–291, 295, 362, 514, 525,
Kac, Mark (1914–1984), 118, 284, 309, 389, 531, 733
732 Lax, Peter David (*1926), 203, 733
Kahane, Jean-Pierre (*1926), 2, 130, 131, Lebesgue, Henri Léon (1875–1941), 211
740 LeBrun, Claude, 654, 733
Kaku, Michio, 392, 732 Lefschetz, Solomon (1884–1972), 357
Kaluza, Theodor Franz Eduard Leibniz, Gottfried Willhelm von
(1855–1954), 370, 732 (1646–1716), 329
Karoubi, Max, 290, 732 Leray, Jean (1906–1998), 211
Karoui, Nicole El (*1944), 140, 732 Lerda, Alberto, xx, 364, 729
Kato, Tosio (1917–1999), 25, 28, 51, 66, Lesch, Matthias (*1961), 49, 52, 53, 214,
387, 732 239, 244, 303, 340, 354, 356, 364, 726,
Keller, Joseph Bishop (*1923), 211 727, 730, 733, 734
Kennedy, Gerard, 231, 729 Levi, Eugenio Elia (1883–1917), 208, 305,
Kepler, Johannes (1571–1630), 71 339, 734
Kirk, Paul, 512 Levinson, Norman (1912–1975), 20, 29, 32,
Klein, Oskar Benjamin (1894–1977), 370, 51, 132, 144, 727
385, 732 Lindelöf, Ernst Leonard (1870–1946), 29
Klimek, Slawomir, 349, 732 Lions, Jacques-Louis (1928–2001), 194, 196,
Kline, Morris (1908–1992), 19, 29, 71, 329, 199, 201, 240, 734
732 Liouville, Joseph (1809–1882), 33, 330
INDEX OF NAMES/AUTHORS 753

Lipschitz, Rudolf Otto Sigismund Mrowka, Tomasz (*1961), xix, xx, 310, 460,
(1832–1903), 29, 30 512, 643, 652, 654, 733
Logan, J. David, 29, 734 Mumford, David Bryant (*1937), 332, 333,
Loll, Renate, 726 336, 337, 735
Lopatinskiı̆, Yaroslav Borisovich Musso, Daniele, xx, 364, 729
(1906–1981), 338, 345, 347, 734
Naber, Gregory L., 460, 735
Nakahara, Mikio, 460, 735
Madore, John, 467, 734
Narasimhan, Raghavan (*1937), 136, 186,
Magenes, Enrico (1923–2010), 194, 196,
194, 229, 735
199, 201, 240, 734
Nash, Charles, 100, 107, 119, 735
Manin, Yuri Ivanovich (*1937), 467, 482,
Natsume, Toshikazu, xix, 735
488, 512, 724
Nazaı̆kinskiı̆, Vladimir E., xx, 735
Marcolli, Matilde (*1969), 667, 726
Nest, Ryszard, xix, 232, 735
Marsden, Jerrold Eldon (1942–2010), 253,
Neumann, John von (1903–1957), 49, 109,
723
735
Maslov, Viktor Pavlovic (*1930), 211, 217,
Neumann, Carl Gottfried (1832–1925), 154
308
Nicolaescu, Liviu, 197, 200, 202, 348, 356,
Massey, William S., 657, 734
497, 498, 735
Mathew, Akhil, 327, 734
Nikčević, Stana, 721, 729
Mayer, Karl Heinz, 319, 327, 734
Ninomiya, Masao, 339, 346, 351, 735
Mayer, Walther (1887–1948), 339
Nirenberg, Louis (*1925), xvii, 198, 213,
McDuff, Margaret Dusa Waddington
216, 225, 244, 735
(*1945), 652, 734
Noether, Fritz Alexander Ernst
McKeag, Peter, 734 (1884–1941), xvii, 3, 47, 126, 127,
McKean, Henry P. (*1930), xviii, 11, 27, 146–148, 152, 153, 252, 269, 304, 337,
96, 116, 120, 123, 126, 132, 227, 305, 735
307, 309, 705–711, 728, 734 Noether, Max (1844–1921), 329–331, 333,
Mehler, Gustav Ferdinand (1835–1895), 586 727
Melrose, Richard Burt (*1949), xviii, xix, Nohl, Craig R., 466, 482, 732
354, 734 Nomizu, Katsumi (1924–2008), 171, 175,
Michelsohn, Marie-Louise, xviii, xix, 231, 448, 456, 721, 732
232, 235, 284, 289–291, 295, 362, 514,
525, 531, 733 Osgood, Brad G., 119, 735
Michor, Peter Wolfram (*1949), 175, 733 Otsuki, Nobukazu (*1941), 49, 726
Mies, Thomas, 3, 735 Otte, Michael (*1938), 3, 735
Mikhlin, Solomon Grigoryevich
(1908–1990), 27, 207, 734 Palais, Richard Sheldon (*1931), xvii, 166,
Milnor, John Willard (*1931), 118, 261, 184, 186, 192, 194, 198, 201, 203, 204,
309, 311, 325, 448, 645–648, 651, 734 213, 216, 239, 245, 249, 250, 284, 301,
Minakshisundaram, Subbaramiah 302, 336, 351, 497, 498, 670, 735, 736
(1913–1968), xviii, 114, 306, 734 Park, Jinsung, 119, 736
Minkowski, Hermann (1864–1909), 706 Parseval des Chênes, Marc-Antoine
Mishchenko, Alexandr S., 447, 735 (1755–1836), 197, 230, 238
Mitter, Peter K., 505, 735 Patashnik, Oren, 549, 730
Miyaoka, Yoichi, 655 Patodi, Vijay Kumar (1945–1976), xviii,
Mizohata, Sigeru, 48, 735 112, 138, 249, 302–304, 306–308, 313,
Montgomery, Richard, 252, 735 319, 336, 337, 339, 345, 349, 350, 355,
Morchio, Giovanni, 341, 348, 349, 354, 726, 356, 724, 736
735 Peano, Giuseppe (1858–1932), 29
Morgan, John Willard (*1946), 651, 652, Pedersen, Gert Kjærgård (1940–2004), 3, 9,
729 10, 13, 15, 19, 20, 25, 31, 36, 37, 50,
Moriyoshi, Hitoshi, xix, 735 51, 53, 81, 736
Morrey, Charles Bradfield (1907–1984), Peetre, Jaak (*1935), 136
141, 735 Penrose, Roger (*1931), 466
Morse, Harold Calvin Marston Perelman, Grigori Yakovlevich (*1966), 77,
(1892–1977), 261 257, 648, 736
Moscovici, Henri, 356, 734 Peskin, Michael E., 385, 736
754 INDEX OF NAMES/AUTHORS

Pflaum, Markus J., 231, 232, 234, 236, 356, Salamon, Dietmar Arno (*1953), 652, 734
734, 736 Sannino, Francesco (*1968), xx, 364, 737
Phillips, John, 52, 53, 726 Šapiro, Zorya Yakovlevna (*1914), 338,
Phillips, Ralph Saul (1913–1998), 119, 735 345, 737
Piazza, Paolo, 354, 734 Sarnak, Peter Clive (*1953), 119, 735
Picard, Charles Émile (1856–1941), 29, 30 Savin, Anton Yu., xx, 356, 735, 737
Piene, Ragni, 731 Scharlau, Winfried (*1940), 18, 731
Planck, Max Karl Ernst Ludwig Schechter, Martin (*1930), 3, 9, 13, 15, 24,
(1858–1947), 374 25, 64, 69, 70, 194, 201, 725, 737
Pleijel, Åke (1913–1989), xviii, 114, 306, Schick, Thomas, xix, 737
734 Schiff, Leonard Isaac (1915–1971), 381,
Poincaré, Jules Henri (1854–1912), 252, 382, 737
253, 256, 257, 262, 302, 327, 648 Shilov, Georgi E., 131, 729
Pollack, Alan, 173, 185, 730 Schmidt, Erhard (1876–1959), 20, 28
Polyakov, Alexandr Markovic (*1945), 466, Schmidt, Jeffrey R., 352, 737
725
Schneider, Michael (1942–1997), 715
Pontryagin, Lev Semenovic (1908–1988),
Schrödinger, Erwin Rudolf Josef Alexander
302, 321, 447, 736
(1887–1961), 190, 378
Przeworska-Rolewicz, Danuta, 4, 736
Schroeder, Daniel V., 385, 736
Prößdorf, Siegfried (1939–1998), 132, 212,
Schrohe, Elmar (*1956), 107, 345, 730
283, 736
Schröder, Herbert, xix
Quillen, Daniel Grey (1940–2011), 91, 736 Schubring, Gert, 3, 735
Quinn, Frank Stringfellow, 651, 729, 736 Schulze, Bert-Wolfgang (*1944), xx, 340,
356, 735–737
Rabinowitz, Philip (1926–2006), 331, 727 Schwartz, Jacob T. (1930–2009), 51, 728
Raikov, Dmitrii Abramovich (1905–1980), Schwartz, Laurent (1915–2002), 51, 70, 131,
131, 729 136, 239, 243, 245, 249, 302, 727, 737
Ralston, James V., 303, 736 Schwarz, Albert S. (*1934), 38, 39, 77, 393,
Ray, Daniel Burrill, 119, 736 466, 725, 737
Rayleigh, Lord, John William Strutt Schwarzenberger, Rolph L.E., 336, 337, 737
(1842–1919), 66 Scott, Simon G., xx, 91, 107, 114, 119, 308,
Rebbi, Claudio (*1943), 466, 482, 732 348, 737
Reed, Michael, 25, 41, 47, 60, 387, 736 Seeley, Robert Thomas (*1932), xvii, 126,
Reinhard, Hervé, 140, 732 152, 207, 239, 243, 244, 276, 301, 304,
Rellich, Franz (1906–1955), 18, 66, 201, 240 306, 340, 344, 736, 737
Rempel, Stephan, 340, 736
Segal, Ed, 460, 728
Reshetikhin, Nicolai Yuryevic (*1958), xx,
Segal, Graeme Bryce (*1941), xviii, 91, 95,
374, 383, 388, 726, 737
101, 104, 278, 292, 294, 361, 362, 724,
Rham, Georges de (1903–1990), 322
738
Riemann, Georg Friedrich Bernhard
Seiberg, Nathan ”Nati” (*1956), 310, 643,
(1826–1866), 146, 330–333, 707
660, 738
Riesz, Frigyes (1880–1956), 15, 25–27, 65,
Seifert, Karl Johannes Herbert
70, 209
(1907–1996), 192, 320, 738
Riesz, Marcel (1886–1969), 304
Serre, Jean-Pierre (*1926), xviii, 2, 304,
Roch, Gustav (1839–1866), 146, 333
336, 648, 726, 738
Roe, John, xix, 290, 731
Shanahan, Patrick Daniel, xix, 284, 362,
Rokhlin, Vladimir Abramovich
738
(1919–1984), 325, 594
Rolewicz, Stefan, 4, 736 Shapiro, Arnold Samuel, 243, 724
Rota, Gian-Carlo (1932–1999), 157 Shifman, Mikhail Arkadyevic (*1949), 373,
Roubine, Élie, 737 738
Rubakov, Valery A., 460, 737 Shih, Weishu, 736
Rudin, Walter (1921–2010), 3, 20, 24, 51, Shubin, Mikhail (*1944), 49, 202, 218, 241,
708, 710, 737 738
Siegel, Carl Ludwig (1896–1991), 330, 333,
Safarov, Yuri, 231, 734, 737 738
Saglanmak, Neslihan, 329, 737 Simon, Barry (*1946), 25, 41, 47, 49, 53,
Salam, Abdus (1926–1996), 385 60, 338, 387, 723, 736, 738
INDEX OF NAMES/AUTHORS 755

Singer, Isadore Manuel (*1924), xvii–xix, 3, Uhlenbeck, Karen Keskulla (*1942), 465,
4, 86, 112, 116, 119, 157, 170, 199, 227, 475, 512, 729, 739
232, 243, 245, 249, 252, 266, 267, 278, Uhlmann, Gunther, 356
280, 284, 285, 287, 291, 295, 299, 301, Ulam, Stanislaw Marcin (1909–1984), 309,
302, 305–309, 313, 318, 319, 336–339, 739
345, 349, 350, 354–357, 361, 362, 466, UNESCO, 29, 128, 132, 145, 171
491, 508, 511, 715, 724, 725, 734, 736,
738 Varadhan, Srinivasa R. S., 305, 739
Sjöstrand, Johannes, 165, 194, 196, 211, Vassilevich, Dmitri, xx, 118, 729, 739
216, 218, 223, 730 Vaı̂nberg, M. M., 141, 144, 739
Slovák, Jan, 175, 733 Vekua, Ilia (1907–1977), xvii, 3, 47, 146,
Smale, Stephen (*1930), xvii, 325, 651, 681, 147, 152, 337, 739
738 Vergne, Michèle (*1943), 138, 284, 725
Sobolev, Sergie Lvovich (1908–1989), 194, Viallet, Claude-M., 505, 735
200, 201 Vietoris, Leopold (1891–2002), 339
Solovay, Robert Martin (*1938), 736 Vishik, Mark Iosifovic (1921– 2012), 338
Sommer, Friedrich, 132, 549 Vishik, Simeon, 107, 733
Spanier, Ernst Paul (*1920), 643, 645, 655, Volpert, Aizik I., xvii, 3
658, 738 Volterra, Vito (1860–1940), 25, 27, 209
Spencer, Donald Clayton (1912–2001), 2 Voronov, Theodore Th., 231–233, 739
Stasheff, James Dillon (Jim) (*1936), 311, Vázquez-Mozo, Miguel Á., 373, 723
448, 734
Steenrod, Norman Earl (1910–1971), 8, 71, Waldhausen, Friedhelm (*1938), 82
85, 259, 266, 267, 278, 728, 738 Wall, Charles Terence Clegg (*1936), 308,
Sternin, Boris Yu., xx, 356, 735, 737 651, 739
Stiefel, Eduard L. (1909–1978), 327 Wallace, Andrew H., 162, 739
Stokes, George Gabriel (1819–1903), 145 Wallach, Nolan R., 632, 739
Strassen, Volker (*1936), 144 Wang, Bai-Ling, 667, 726
Strocchi, Franco, 341, 348, 349, 354, 726, Ward, Richard S., 466, 725
735 Webb, David L., 118, 730
Sturm, Jacques-Charles-François Weber, Heinrich Martin (1842–1913), 336
(1803–1855), 33 Weierstrass, Karl Theodor Wilhelm
Szankowski, Andrzej, 24, 738 (1815–1897), 18, 67, 126, 333
Weinberg, Erick J. (*1947), 466, 725
Tan, Chung-I, 340, 346, 351, 735 Weinberg, Steven (*1933), 373, 385, 739
Taubes, Clifford Henry (*1954), 310, 355, Weitzenböck, Roland (1885–1955), 439
512, 643, 651–654, 738 Wells, Raymond O., Jr. , 183, 186, 229,
Taylor, Michael Eugene (*1946), 49, 208, 530, 739
218, 738, 739 Weyl, Hermann Klaus Hugo (1885–1955),
Thom, René (1923–2002), 302–304, 311, 109, 240, 254, 336, 419, 428, 525, 739
318, 739 Whitehead, Alfred North (1861–1947), 329,
Thomas, Paul Emery (1927–2005), 327, 739 739, 740
Thorpe, John A., 157, 170, 738 Whitehead, John Henry Constantine
Threlfall, Wiliam Richard Maximilian (1904–1960), 645, 740
Hugo (1888–1949), 192, 320, 738 Whitney, Hassler (1907–1989), 162, 327,
Titchmarch, Edward Charles (1899–1963), 740
129, 739 Widom, Harold (*1932), 231, 236, 740
Toeplitz, Otto (1881–1940), 12, 19, 731 Wiener, Norbert (1894–1964), 2, 121, 124,
Toledo, Domingo, 310, 739 130–132, 134, 140, 740
Tong, Yue Lin L., 310, 739 Wilczek, Frank Anthony (*1951), 466, 740
Torelli, Ruggiero (1884–1915), 332 Witten, Edward (*1951), xx, 310, 365, 466,
Trenogin, Vladilen A., 141, 144, 739 643, 654, 660, 738, 740
Trotter, Hale F., 387, 739 Wojciechowski, Krzysztof P. (1953–2008),
Tsygan, Boris, xix, 232, 735 xx, 47, 49, 107, 114, 119, 192, 200, 244,
Tsypkin, Yakov Zalmanovic (1919–1997), 246, 261, 272, 275, 282–284, 287, 303,
124, 740 308, 319, 339–341, 344, 347–351, 355,
Tyupkin, Yu. S., 466, 725 356, 667, 726, 728, 732, 736, 737
Tzou, Leo, 356, 730 Wolpert, Scott A., 118, 730
756 INDEX OF NAMES/AUTHORS

Yang, Chen-Ning Franklin (*1922), 363


Yau, Shing-Tung (*1949), xix, 655, 738,
740
Yood, Bertram, xvii, 740
Yoshida, Tomoyoshi, 138, 355, 740
Yosida, Kõsaku (1909–1990), 194, 740
Yu, Yanlin, 171, 740

Zagier, Don Bernard (*1951), xx, 310, 311,


731
Zeidler, Eberhard, 373, 740
Zessin, Hans, 726
Zhang, Weiping, xx, 740
Zhu, Chaofeng (*1973), 214, 239, 303, 308,
340, 354, 356, 726
Zygmund, Antoni (1900–1992), 208, 727
Subject Index

Terms and topics are not in alphabetical order but grouped according
to teaching goals. → refers, or auto-refers, to a separate main entry.

Abelian integrals, 331 full proof for twisted generalized Dirac


Abstract Minkowski space-time, 367 operators via heat kernel
Actions (Lagrangians), 369 asymptotics, 513
(Anti-) self-dual connections, 489 heat-equation proof, sketch, 304
= instantons, 465 K-theoretical version, 294
Formal Dimension Theorem, 491 Atkinson’s Theorem, 64
forms, 405
global slices, 508 Banach manifolds
instantons’ capturing, 488 critical value of Fredholm mapping, 681
instantons’ description, 484 Fredholm mapping, 681
local regularity, 505 Fredholm Theorem, 682
Fredholm transversality theorem, 701
local slices, 503
regular value of Fredholm mapping, 681
manifold structure and dimension of
residual set theorem, 684
space of moduli of mildly-irreducible
Sard-Smale Theorem, 681
self-dual connections, 511
transverse pair of mappings, 700
mildly irreducible, 496, 507, 510
Bott Periodicity Theorem
self-dual slices, 500
classical, 257, 272
size of instanton, 475
K-theoretical, 271
space of moduli, 461, 489
periodicity homomorphism, 281
weakly irreducible, 496
real, 357
Arzelà-Ascoli Theorem, 18
Boundary value problems
Asymmetry
boundary integral method, 339
eta function, 111
chiral bag model, 355
index (chiral asymmetry), 3, 349
elliptic (well-posed), 244, 343, 344
Asymptotic expansions
local elliptic boundary condition, 142,
for formal series, 549 146, 244, 246, 338, 345–347
Landau-Bachmann notation O, o, 549 Poisson principle, 272
Atiyah-Singer (AS) Index Theorem
cobordism proof, sketch, 302 Cabibbo-Kobayashi-Maskawa 3 × 3 matrix,
cohomological formulation, 313 385
embedding (K-theoretic) proof, sketch, Calderón projector, 344
304 Calderón’s inverse problem, 356
Euclidean, for operators equal to identity Calkin algebra, 63
at infinity, 281 Canonical
for generalized Dirac operator, 599 1-form on bundle of linear frames, 415
for trivial embeddings, 285 class of Spinc (n)-structure →, 655
for twisted Dirac operator, 546, 592 class of almost complex structure, 653
for twisted generalized Dirac operator, difference element of exterior algebra, →
600 K-theory, 289
full K-theoretical proof, 287 expansion, 57, 58

757
758 SUBJECT INDEX

line bundle, 633 divisor on Riemann surface, 335


Cayley numbers, 607 Dolbeault
Characteristic classes, 266, 327 cohomology groups, 334
Chern cohomology groups, generalized, 639
character, 312, 445 cohomology spaces, 617
character defect, 314 operator → Dirac operators, 187
class, 312, 444 elliptic curve, 332
number, 336 holomorphic Euler characteristic, 334,
Euler characteristic →, 319 642
Euler class, 313, 321, 453 K3 surface, 648
functor, 311 Kähler
generalized total characteristic class, 598 2-form, 618
multiplicative form, 449 Kähler
orientation class, 311 manifold, 334
Pontryagin Theorem, 622
class, 324, 447 Riemann surface, 331
form, 546 symplectic structures →, 652
singular set of array of vector fields, 325, Connections, 177
327 (anti-) self-dual forms →, 405
Stiefel-Whitney class, 527 basic forms, 401
symbolic calculus →, 443 canonical 1-form on bundle of linear
Todd class, 312 frames, 415
Todd genus, 335 Chern connection, 640
total Ab class, 451 Christoffel symbols, 178, 421
total Chern class, 312 compatible (with Clifford module
total form, 449 structure), 180
total Hirzebruch L class, 451 Branson-Gilkey Theorem , 180
total Todd class, 450 compatible with complex structure, 640
Chiral connection 1-form, 398
asymmetry, 349 exponential map, 413
bag model, 346, 355 formal normal space, 489
chiral switch, 596 formal tangent space, 489, 491
decomposition, 340, 532 Gauss-Bonnet form, 446
Dolbeault and Hodge star, 620 geodesic, 179
representation, 521 Hermitian, 232, 640
Classical action, 388 infinitesimal action, 413
Classical paths, 388 Levi-Civita connection →, 180
Clifford analysis Maurer-Cartan form, 398
(left) Clifford multiplication, 180, 328, metric (Riemannian, Hermitian), 179
531, 659 module derivation, 180
bracket (commutator), 515 on principal G-bundle, 396
Clifford algebra, 514 parallel sections, 178
Clifford module, 180 parallel translation, 178, 232
Clifford module bundle, 595 projection of moduli space
complex Clifford algebra, 520 uniform estimates, 688
complex Clifford bundle, 595 quotient space of moduli, 673
complex volume element, 520 quotient space of moduli of all
exponential map, 514 connections, 461
Closed Graph Theorem, 31 Sobolev space of connections, 500
Codimension, 3 Sobolev-parametrized space
Cokernel, 3 Banach manifold, 668
Complex manifolds, 615 differential, 668
almost-complex structure, 313, 453 space of connections, 396, 461
arithmetic genus, 188, 334 space of moduli of connections, 489
canonical divisor of Riemann surface, 336 torsion free, 179
Cauchy-Riemann operator, 332 torsion tensor, 179, 415
complex structure, 615 weakly irreducible, 490
complex volume element, 520 Curvatures, 372
SUBJECT INDEX 759

Chern form, 444 elliptic →, 137


contractions of tensor fields, 436 Euler operator, 187
curvature tensor, 422 exterior differentiation, 174, 435
decomposition formally adjoint, 184, 405
(anti-) self-dual part, 433 formally self-adjoint (symmetric), 47
constant curvature part, 430 heat (conduction) equation, 136
traceless Ricci part, 430 hyperbolic equation, 137
Weyl part, 430 hypo-elliptic, 210
decomposition on 4-manifold, 433 Laplacian, → Laplace operators, 187
Euler form, 453 linear, 181
first Bianchi identity, 424 on S 1 , 38
Gaussian curvature, 423 order, 135
general Bianchi identity, 403 parabolic equation, 137
of connection, 399 real, 182
Ricci curvature, 426 vectorial, 138
Ricci tensor, 426 wave equation, 136
scalar curvature, 426, 654 wave fronts, 138
second Bianchi identity, 425, 426 Diffusion processes
trace of curvature tensor, 427 and heat asymptotics, xx, 305, 549
vector space of curvature tensors, 427 boundary conditions, 142
5-Dimensional cylindrical universe, 371
Degree (mapping degree, winding number), Dirac operators, 307
69 (free) euclidean Dirac operator, 341
alternative definitions, 258 (generalized) Dirac operator, 597
for n-sphere, 258 (partial) chiral Dirac operator, 532
for mapping the n-sphere into GL(N, C), (standard) Atiyah-Singer operator, 302,
259 531
from sphere to sphere, 274 twisted generalized Dirac operator, 600
homotopy of GL(N, C), 283 Cauchy-Riemann operator, 187, 332, 334
Ninomiya-Tan model, 353 chiral decomposition, 181
of boundary condition, 147 compatible, 181
of closed planar curve, 254 DeRham-Dirac operator, 320, 602
of zero of vector field, 458 Dolbeault complex, 334
Determinants Dolbeault operator, 187, 334
canonical section of determinant line euclidean spinors, 341
bundle, 107 Euler operator, 187, 320, 609
exterior determinant, 92, 693 gauge transformation, 351
Fredholm determinant, 98 gauge-invariant family, 352
multiplicative anomaly, 108 generalized Dolbeault complex, 334
Quillen determinant line bundle, 95 generalized Dolbeault operator, 639
Segal determinant line bundle, 102 generalized Euler operator, 608
zeta determinant, 118 global chiral symmetry, 349
Diagram chasing, 6 Green’s formula, 340
exact sequence, 6 local chiral symmetry, 351
Snake Lemma, 6 of Dirac type, 180, 340
Differential forms, → Vector analysis, 171 partial (chiral) Dirac operator, 181, 340
Differential operators product form in collar of boundary, 341
(principal) symbol, 183 pure gauge at boundary, 342
approximation by Bokobza-Haggiag real, 356
pseudo-differential operators, 235 Riemann-Roch operator, 337
Cauchy-Riemann operator, 187 signature operator, 321, 323, 604
characteristic manifold, 138 Spinc -Dirac operator, 659
characteristic polynomial, 183 spinors, 340
characteristic polynomial (principal tangential operator over boundary, 341
part), 139 total (elliptic) Dirac operator, 340
codifferential, 188, 406 twisted DeRham-Dirac operator, 606
covariant differentiation, 177, 435, 659 twisted Dirac operator, 342, 532
Dirac, → Dirac operators, 180 twisted signature operator, 346, 606
760 SUBJECT INDEX

Yang-Mills Dirac operator, 613 convolution, 706, 708, 711


Dirac’s equation, 382 convolution theorems, 711
Distribution differentiation-multiplication conversion,
Dirac δ-distribution, 38 709, 711
singular support, 117 Fourier coefficient, 37, 706
wave kernel, 117 Fourier integral, 708, 710
Domain of an operator, 35 Fourier inversion formula, 709, 711
Donaldson’s polynomial invariants, 650 Fourier series, 706
Donaldson’s Theorem for smooth generalized function, 707
4-manifolds integrable-continuous conversion, 711
with definite intersection forms, 648 Parseval identity, 37, 711
with even, indefinite intersection forms, Plancherel formula, 709, 711
648 Poisson summation formula, 710
rapidly decreasing smooth functions, 708
E-M field strength, 366
Fredholm alternative, 10, 26
Einstein field equation, 368
for Sturm-Liouville boundary-value
Eleven eighth’s conjecture, 650
problems, 31
Elliptic
Riesz Lemma for compact perturbations
Cauchy data space, 344
of identity, 26
classical Poisson equation, 137
decomposition, 242 Gauge transformations, 409
differential operators, definition, 139, 186 action on connections, 411
fundamental elliptic estimates, 241 formal slice of action, 489
Gårding’s Inequality, 241 gauge group, 409, 461
main result, 239
gauge potential, 370, 372
regularity, 33, 241, 347
of particle fields, 370
summary of main results, 240
Generalized Green function, 31
topological meaning (case), 245
Genus
topology, 337
Klassenzahl (former terminology), 319
unique continuation property, 667
numerical complexity, 332
well-posed, 343
of an algebraic structure, 330
Elliptic integrals, 330
of Riemann surface, 117, 257, 331
genus of elliptic curve, → genus, 332
Thom’s Genus Conjecture, 654
of higher kind, 331
Geometrical unification, 370
of the first kind, 331
Gohberg-Krein Product Lemma for
Empty bundle universe, 372
unbounded Fredholm operators, 43
Empty space equation, 369
Graph of an operator, 36
Euler characteristic
Green’s formula, 181
Euler operator, 320
Groups
of a complex, 7
adjoint action, 397
of a Riemann surface, 117
adjoint representation, 397
of topological manifold, 319, 609
character, 360
Faraday 2-form, 366 cocycle condition, 526
Feynman-Kac Formula, 389 degree →, 274
Fibration elementary symmetric polynomial, 448
of unitary group, 259 free abelian group, 263
Field strengths (curvatures), 372 Grothendieck ring, 360
Fischer-Riesz Representation Theorem, 15 Lie group of matrices, → Lie algebras,
Fixed points 394
Atiyah-Bott-Lefschetz number, 358 Pfaffian polynomial, 446
Brouwer Theorem, 273 representation, 400
Lefschetz number, 357 respresentation space, 360
multiplicity, 359 special orthogonal group, 396
transversal, 359 special unitary group, 395
Fourier analysis splitting exact sequence, 267
Bokobza-Haggiag subgroup of finite index, 303
Fourier transform, 232, 291 t’Hooft matrices, 467
inverse Fourier transform, 233 universal property, 263, 265
SUBJECT INDEX 761

h-cobordism homotopy classes, 72


Theorem of Milnor, 651 homotopy equivalent spaces, 72, 716
Hardy spaces homotopy set, 72
discrete, 125 invariants, 68, 81
in Clifford analysis, 344 retract, retraction, 72, 267
Heat equation Hydrogenic atoms, 374
classical heat distribution, 305, 550, 563
classical, in homogeneous isotropic body, Implicit Function Theorem
136 classical, 157
spinorial, 551 for Banach spaces, 499
trace of the heat kernel, 115 Inverse Function Theorem, 157
Heat kernel asymptotics, 571 Inverse Function Theorem for Banach
approximative heat kernel, 566, 567 spaces, 500
asymptotic expansion, 550 Index
convolution of kernels, 567 analytic index, 276, 281, 285, 286
general heat kernel for Dirac Laplacian, analytic index, definition, 291
547 homotopy invariance, 81
heat kernel error term, 567 index bundle, 82, 84, 269
Hilbert-Levi’s parametrix method, 305 index density, 346
Mehler’s formula, 586 local index, 276
supertrace, 544 of a singular point, 252
theta function, 304 of an elliptic G-invariant operator, 360
twisted spinorial heat kernel of bounded Fredholm operator, 3
positive and negative, 541 of Fredholm pair, 344
total, 541 topological G-index, 361
(Heisenberg) Uncertainty Principle, 376 topological g-index, 361
Higgs mechanism, 385 topological index, 276, 285
Hilbert transform, 127 topological index, definition, 290
Hilbert-Schmidt Diagonalization Theorem, virtual character, 360
20 virtual codimension of Fredholm pair of
Hodge theory projections, 345
Hodge-DeRham decomposition, 322 Index Theorems
main theorem, 322 Agranovich-Dynin index correction
Holonomy, 455 formula, 347, 355
gauge isotropy subgroup, 456 Atiyah-Bott-Lefschetz (ABL) Formula,
holonomy bundle, 456 359
holonomy group, 455 Atiyah-Patodi-Singer (APS), 345
Homology, cohomology, 255 Atiyah-Segal-Singer Fixed-Point
Čech Formula, 362
cochains and coboundaries, 526 Atiyah-Singer (AS) Index Theorem →,
cohomology class, 527 294
cohomology group with values in Z2 , Bojarski Conjecture, 356
527 Chern-Gauss-Bonnet Theorem, 321, 612
Bockstein homomorphism, 655 Cobordism Theorem, 303, 355
deRham cohomology space, 603 Cordes-Labrousse Theorem for
homology space of complex, 7 unbounded Fredholm operators, 52
of n-sphere, 259 Dieudonné Theorem of local constancy,
Poincaré duality, 644 68
simplicial complexes, 252 Discrete (Noether)-Gohberg-Krein Index
Thom isomorphism, 311 Formula, 125
total class, 449 for continuous Wiener-Hopf operator, 131
with compact support, 312 for convolution operator on real line, 129
Homotopy, 255 for Fredholm mappings between Banach
Bott Periodicity Theorem →, 258 manifolds, 682
contractible, 73, 716 G-index Formula, 361
deformation retract, 72 Hirzebruch Signature Theorem, 321, 605
fundamental group, 253, 257, 258 Hirzebruch-Riemann-Roch Theorem,
homotopic maps, 72, 716 335, 637
762 SUBJECT INDEX

Holomorphic Hirzebruch-Riemann-Roch symbol class, 289


Theorem, 642 Thom Isomorphism, 290
Local Index Formula, 577 torsion, 87
Noether’s Index Formula, 127 with compact support, 266, 267
Noether(-Hellwig-Vekua), Theorem of, Kuiper’s Theorem, 77
146, 152
on the circle, 41 Laplace operators
signature deficiency formula, 355 Bochner-Weitzenböck Formulas, 434
spectral flow theorem, 356 classical potential (Poisson, Laplace)
Twisted Chern-Gauss-Bonnet Theorem, equation, 137
613 connection Laplacian, 189, 199, 439, 536
Twisted Hirzebruch Signature Theorem, Dirac Laplacian, 187, 535
606 Euclidean, 187
Twisted Hirzebruch-Riemann-Roch Hodge Laplacian, 189, 307, 439
Theorem, 638 Laplace-Beltrami, 307
Lattice, 644
winding number model, 253
Levi-Civita connection, 418
Yang-Mills-Dirac Index Theorem, 613
Christoffel symbols, 558
Inertial system, 367
decomposition, 467
Instantons, → (Anti-) self-dual
field strength, 467
connections, 393
for Bokobza-Haggiag Fourier transform,
Integral operators
232
convolution, 27
fundamental lemma of Riemannian
Fourier integral operator, 226
geometry, 418
Fourier transform, 27
in local coordinates, 420, 469
Fredholm integral equation of the second
pull-back, 467
kind, 10, 27, 28
Lie algebras, 397
Green’s operator, 27
ad-invariant inner product, 461
Hilbert transform, 127
complexification, 464
Hilbert-Schmidt operator, 27
complexified Killing type form, 464
integral kernel, 10
exponential map, 397
Mellin transform, 115
inner product, 460
normalized integration operator, 38
Lie-algebra-valued connection 1-form,
oscillatory integrals, 217
372
singular integral operator, 26, 27, 226
of antisymmetric matrices, 515
Volterra integral, 27
Linear frames, bundle, 414
Wiener-Hopf operator, 27
Linear span, 21
Inverse problems, 309
Lorentz
condition, 370
Kirby-Siebenmann invariant, 648
force law, 366
Klein-Gordon equation, 379
transformations, 368
K-theory
Lp -sections, 497
α-homomorphism (Bott isomorphism),
270 Magnetic moment of the electron, 383
analytic index, multiplicative property, 4-Manifolds
294 curvature parts and self-duality, 431
Atiyah-Jänich Theorem, 88 Freedman’s Theorem, 648
Bott class, 270, 285 instantons →, 465
Bott periodicity theorem →, 271 LeBrun’s Theorem on Einstein
canonical difference element of exterior 4-manifolds, 654
algebra, 289 Poincaré’s conjecture in dimension 4, 648
difference bundle, 83, 263, 278, 279, 288 Rokhlin’s Theorem, 649
equivariant K-theory, 287 Seiberg-Witten invariant, 652
extension homomorphism, 290 self-dual weakly irreducible connections,
K-group, 263 496
outer tensor product, 267 smooth structure
relative K-theory, 264 not admitting, 649
ring structure, 278 with E8 summand, 649
splitting homomorphism, 279 smooth structures
SUBJECT INDEX 763

Donaldson and Kronheimer Theorem, semi–Fredholm operator, 48


651 shift operator, 4
Friedman and Morgan Theorem, 651 symmetric (formally self-adjoint), 47
symplectic Toeplitz operator, 226, 282
Taubes’ Theorem, 653 trace class, 54, 56
Wall’s Cobordism Theorem, 651 unbounded Fredholm operators, 42
Witten’s Theorem, excluding positive Wiener-Hopf operator, 26
scalar curvature, 654 Wiener-Hopf operator, continuous, 129
Maxwell’s equations, 366, 370 Wiener-Hopf operator, discrete
Minimal replacement (covariant (Toeplitz), 125
derivatives), 379 Orthonormal frame bundle, 416
Miyaoka-Yau inequality, 655
Moduli space, 664 Parametrix (quasi-inverse), 64, 65
Momentum space, 376 Pauli matrices, 382, 467
Morse Index Theorem, 356 Perturbation invariance, 68
Poincaré group, 368
Noncommutative geometry, 302 Postulate of quantum mechanics, 375
Norms Principal G-bundles, 371, 395
Fourier coefficient norm, 37 all-purpose bundle, 491
gap metric, 52 associated vector bundle, 400
graph norm, 36 automorphism, 409
L2 -norm, 26 fibered product, 434
numerical radius, 20 gauge transformation →, 409
operator norm, 3
Hopf bundle, 395
Sobolev norm, 37, 195, 497
horizontal equivariant forms, 401
trace norm, 60
horizontal lift of vector field, 400
Open Mapping Principle, 13 transition functions, 414
Operators with larger Lie group, 384
adjoint operator, 15, 35 Projections
bounded from below, 48 (spectral) APS projection, 345
chiral (nonsymmetric) components, 47 Calderón projector, 344
closable, 36 Fredholm pair of subspaces, 344
closed operator, 36 Grassmannian, 345
compact (completely continuous), 18 idempotent, 345
convolution operator, 26 of family of vector spaces, 712
densely defined (not necessarily projection metric, 52
bounded), 35 sectorial projection, 356
diagonizable (discrete), 20 self-adjoint Grassmannian, 348
Dirac operator →, 180 smooth Seiberg-Witten projection
essential unitary equivalence, 108 mapping, 680
essentially invertible, 64 Pseudo-differential operators
essentially normal, 109 amplitude, 206, 275
essentially self–adjoint, 47 Bokobza-Haggiag
essentially-unitary, 108 amplitude, 234
extension of unbounded operator, 35 pseudo-differential operator, 234
formal adjoint, 32 canonical, 210
Fredholm operator, 3 classical, 215
Hilbert-Schmidt operator, 25, 56 doubled operator on closed double, 277
integral operators →, 26 elliptic, 235, 239
minimal closed extension, 36 equal to identity at infinity, 275, 276
normal, 109 Fourier integral operator, 210, 226
normally-solvable, 16, 42, 343 G-invariant, 298, 360
of finite rank, 9 hypo-elliptic, 241, 287
operator of order k, 237 Kuranishi trick, 216, 223
positive operator, 17 local index, 248, 275
realization, 343 outer tensor product, 243
Schatten class, 60 phase function, 206
self-adjoint, 47 principal symbol, 275
764 SUBJECT INDEX

principally classical, 214, 215 Hilbert manifold of solutions, 679


properly supported, 218 parametrized moduli space of
pseudo-locality, 213 solutions, 678
quantization, 236 S-W equations, unperturbed, 662
singular support, 213 S-W invariant
smoothing operator, 213, 226, 235 definition, 693
Weyl quantization, → Quantization, 211 full invariance, 701
Purely gravitational Lagrangian, 369 S-W-equations, wider perturbed, 698
Signals and filters, 122
Quantization Signature
quantization procedures, 374 definition, 320
symbolic calculus →, 182 signature formula, 308, 605
Quantum field theory, 380 signature operator → Dirac operators,
energy levels and eigenvalues, 373 321
Quantum mechanical expectation, 376 Singular support
Quarks, 384 of distribution, 117
Quaternions, 514 Smooth manifolds, 158
automorphism, 159
Relativistic kinetic energy, 367
connection, 177
Renormalization techniques, 383
definition of manifolds through
Rest energy, 367
equations, 162
Riemannian manifolds, 167
diffeomorphic, diffeomorphism, 159
4-manifolds →, 431
differential of smooth map, 161
Christoffel symbols, 168
directional derivative of a function, 160
distance between points, 168
energy of a curve, 168 dotted cotangent bundle, 182
Euler-Lagrange equations, 168 dual tangent (covariant or cotangent)
exponential map (spray), 169 bundle, 164
fundamental lemma, 180, 418 embedding with not necessarily trivial
geodesic, 117, 168, 232 normal bundle, 290
injectivity radius, 169, 232 embedding with trivial normal bundle,
isospectral, 118 285
length of a curve, 168 embedding, Embedding Theorem, 162
length spectrum, 118 hypersurface, 162, 166
Levi-Civita connection →, 232 immersion, 162
metric, metric tensor, 167 integral curve of vector field, 174
normal coordinates about point relative jets and jet bundles, 166
to frame, 553 Lie bracket, 174
normal field to the boundary, 191 Lie derivative, 174
partitioned, 347 mollifying family, 196
point injectivity radius, 169 normal bundle of embedding, 165, 290
radial gauge, 552 orientation, 169
volume form, 170, 405 pull-back (lifting) of differential forms,
165
Schrödinger equation Sard Theorem, 259
formal solution, 389 smooth maps, 158
Schrödinger operator, 387 smooth structures
Schrödinger picture, 378 Donaldson and Kronheimer Theorem,
Schrödinger probability density, 381 651
time-dependent (relativistic), 378 Friedman and Morgan Theorem, 651
Seiberg-Witten (SW) function, 665 submanifold, 162
compactness of moduli of solutions, 691 submersion, 162
joining of FR-pairs, theorem, 701 symplectic cone, 165, 225
MSWP k moduli of solutions, tangent bundle, 161
parametrized, 678 tangent space, 160
parameter pair, freely acting (FR), 699 vector field, 161, 165, 325
quadratic map, 661 with boundary, 190, 337
Regular Value Theorem, 680 Sobolev spaces, 197, 234
S-W equations, perturbed, 662 generating operator, 198
SUBJECT INDEX 765

modeled on L2 , 195 vector representation (double c, 517


modeled on Lp , 497 vector space of half-spinors, 521
on S 1 , 35, 152 vector space of spinors, 521
on the disk, 152 Spinc structures, 655
operator of order k, 237, 502 and the Seiberg-Witten equations, 655
Rellich Compact Inclusion Theorem, 201 canonical line bundle, 659
restriction problems, 205 existence theorem, 655, 657
Sobolev Embedding Theorem, 200 moduli space for Spinc structures =
Sobolev Inequality, 201 orbits of solutions of the perturbed
Sobolev Restriction Theorem (torus S-W equations →, 664
case), 203 Stiefel-Whitney class mod 2, 655
strength, 200 States, 375
truncation, 196 Strong force of quantum chromodynamics
Source 1-form, 366 (QCD), 384
Space-time with E-M radiation, 370 Support, 50
Spectral invariants, 108, 110 Symbolic calculus, 275
characteristic curve, 110, 122 (principal) symbol class, 285
eta function, 111, 346 asymptotic condition of growth, 210
eta invariant, 112 Bokobza-Haggiag
reduced eta invariant, 112 amplitude, 234
spectral flow, 347 Bokobza-Haggiag symbol, 291
trace of the heat kernel, 115 characteristic form, 137
zeta determinant, 118 characteristic polynomial, 139, 183
zeta function, 114, 307 dequantization and quantization →, 190
Spectrum, 50 elliptic principal (polynomial) symbol,
continuous spectrum, 10 186
energy levels and eigenvalues →, 373 leading symbol of differential operator,
essential spectrum, 10, 108 183
isospectral, 118 principal symbol, 234
length spectrum, 117 subprincipal symbol, 184
Mark Kac’s query, 117, 118, 309 symbol space of symmetric k-linear
of a bounded operator, 10 forms, 166
point (discrete) spectrum, 10 topological meaning of principal symbol
renormalized heat trace, 115 (case), 245
residual spectrum, 10 total symbol (amplitude,
resolvent function, 50 dequantization), → quantization,
resolvent set, 50 218
singular values, 57 Symmetric bilinear form
spectral representation, 24 curvature tensor →, 427
Spectral Theorem, 51 definite, 646
symmetric spectrum, 351 E8 , 325, 646, 649
Spin structures even, odd, 646
Ab class, 546 indefinite, 646
for Riemannian manifold, 525 intersection form, 643
half (chiral) spinor representation, 521 Killing form, 462
Hermitian spin bundle, 531 Kulkarni-Nomizu product, 428
Rokhlin’s Generalized Theorem, 593 rank, 646
Rokhlin’s Theorem, 594 signature, 646
set of inequivalent spin structures on unimodular, 646
manifold M , 528 Symplectic structures
spin algebra, 515 Gotay’s conjecture, 654
spin group, 515 symplectic form, 652
spinor bundles, positive and negative, symplectic invariant
531 Hörmander index, 308
spinor frame, 552 Maslov index, 308
spinor representation, 519, 520 symplectic manifold, 652
spinors, 519 almost complex structure →, 652
supertrace, 522 symplectormorphism, 652
766 SUBJECT INDEX

Taubes’ Theorems, 654 bi-invariant forms, 317, 616


codifferential, 188
Thom isomorphism complex forms, 187
equivariant, 361 equivariant forms, 400, 401
in K-theory, 290 ext function, → Exterior multiplication,
of singular cohomology, 311 519
Time series analysis, 124 exterior algebra, 171, 514
Topological invariants exterior differential forms on manifold,
Betti number, 7, 319 172
characteristic classes →, 266 Gauss-Bonnet form, 321
Euler characteristic →, 7 harmonic forms, 317, 602
genus →, 257 Hodge star operator, 172, 405
intersection number, 259, 459 Hodge-DeRham decomposition, 603, 621
self-intersection number, 459 horizontal equivariant forms, 401
signature →, 308 int function, → Interior multiplication,
winding number, → Degree, 69 519
Topological manifolds, 157 invariant form, Weltkonstante, →
atlas of charts, 157 Canonical difference element of
boundary point, 190 exterior algebra, 258
chart, 157 jets, jet bundle exact sequence, 166
coordinate system, 157 skew-symmetric tensor, 171
dimension, 158 star operator, 171
interior points, 190 Vector bundles, 82, 713
Leray covering, 526 base (parameter) space, 712
plumbing, 325 basis theorem on stable equivalence, 272
triangulation, 319 clutching (gluing), 718
tubular neighborhood, 285 complexes with compact support, 277
Topological vector space direct sum, 715
Hilbertable, 198 dual bundle, 715
Topology elementary complexes, 277
bending Rn into closed manifold, 257 family of vector spaces, 712
cobordant, 302 fiber, fiber-dimension, 712, 714
connected sum, 645 G-vector bundle, 360
h-cobordism, 651 holomorphic trivialization, 638
homeomorphic, 72 homomorphism bundle, 712, 715
knot theory, 77 isomorphism bundle, 715
locally finite refinement of covering, 159 line bundle, 714
paracompact, 159 locally trivial family of vector spaces, 713
partition of unity, 78 outer tensor product, 278
Poincaré conjecture, 77, 257 pull-back, 714
proper maps, 267 quotient bundle, 715
residual set, 681 restriction, 713
starlike, 73 section, 712
strong topology, 81 stably-equivalent, 264
suspension, 273, 718 subbundle, 715
Transmission problems, 272, 282 tensor product, 715
total space, 712
Unified field theories trivial bundle, 713
Kaluza-Klein theory, 370 uniqueness theorem, 272
Unique continuation property, 303, 350, 667
Wave equation, 370
4-Vector potential, 370 wave kernel, 117
Vector analysis Weak force, 384
(anti-) holomorphic differential forms, Winding number, → Degree, 254
333
(left) exterior multiplication, 171, 173, Yang-Mills functional, 393, 462
328 Yang-Mills Dirac operator, 613
(left) interior multiplication, 173 Yang-Mills equation, 463
(right) exterior differentiation, 174

Common questions

Powered by AI

Gauge theory and its mathematical structures contribute to understanding the physical universe by providing a robust framework for describing fundamental forces and particles' interactions. The theory uses principal bundles, connections, and curvature, capturing the essentials of concepts like electromagnetism, the weak and strong nuclear forces. Gauge transformations account for the local symmetries inherent in physical laws, creating a mathematically enriched foundation that supports the construction of quantum field theories. This alignment between abstract mathematical architecture and physical phenomena fosters profound insights into unifying forces, driving the exploration of concepts like string theory and quantum gravity .

The concept of the principal symbol in pseudo-differential operator theory arises as a key element that captures the leading order behavior of differential operators. It acts as an invariant that provides insights into the differentiability properties of operators and their analytical indices. The principal symbol is crucial for classifying operators, particularly in distinguishing elliptic operators, and it is integral in the microlocal analysis framework, where it aids in understanding the propagation of singularities in solutions. Thus, it is an indispensable tool in the generalization of differential operators to broader contexts, significantly impacting the resolution of PDEs .

Fredholm operators are integral in functional analysis due to their role in solving linear integral equations that resemble finite-dimensional linear systems. A Fredholm operator, by definition, has a finite-dimensional kernel and cokernel, and a closed range. The Fredholm index, which is the difference dim(Ker F) - dim(Coker F), remains invariant under compact perturbations, reflecting a certain topological stability . Furthermore, Fredholm operators are foundational for the development of index theorems in mathematics, connecting linear algebra, geometry, and analysis .

The Atiyah-Singer Index Theorem is monumental in mathematics and physics as it provides a profound connection between analysis, geometry, and topology. It computes the analytical index of an elliptic differential operator by relating it to topological invariants. This theorem has had far-reaching implications, introducing new tools and concepts such as topological K-theory and shaping modern areas like gauge theory, string theory, and the study of anomalies in quantum field theory. The theorem's versatility in diverse mathematical contexts highlights its foundational role in advancing the understanding of geometric structures and their analytical properties .

In operator theory, bounded linear operators are evaluated based on properties like continuity and range closedness. A critical aspect of such operators, especially within the context of Fredholm operators, is that they possess a closed range and finite-dimensional kernel and cokernel. For a bounded operator to qualify as Fredholm, these conditions ensure that the analytical framework is maintained, facilitating the application of index theory to solve related problems. The closed range condition ensures the solvability of associated equations, reinforcing the stability of these operators under perturbations within the functional analytic framework .

Wiener-Hopf operators are significant in harmonic analysis and index theory as they serve as a classic example of Fredholm operators. These operators are instrumental in solving integral equations on half-lines, leveraging Fourier transform techniques. In index theory, they illustrate the discrete index formula and its continuous analogue, providing tangible instances of index computations through singular integral equations. Their intertwining with harmonic analysis exemplifies the detailed relationship between functional analytic concepts and practical problem solving in engineering and physics .

K-theory with compact support is employed in algebraic topology to analyze vector bundles over topological spaces that are not necessarily compact, using the notion of compactified spaces. It involves extending the insights gained from topological K-theory to noncompact scenarios by considering vector bundles over compactified one-point spaces, providing a framework for analyzing local versus global topological properties. This approach helps define homomorphisms that preserve topological invariants, thus making it possible to track the influences of local changes on global bundles, which is crucial for many K-theory applications in mathematics and physics .

Homotopy invariance is critical in the theory of elliptic differential operators as it ensures that the index of an operator is preserved under homotopic deformations. This property is pivotal because it reflects the topological nature of the index, making it invariant under continuous deformations of the manifold or the operator itself, thus linking differential geometry with topological concepts. This invariance is applicable to elliptic differential operators on closed manifolds and compact manifolds with smooth boundaries, where it preserves the analytical indices across homotopic classes of operators and manifolds .

The Bochner-Weitzenböck formulas play a pivotal role in analyzing differential operators on manifolds by linking curvature terms to analytical invariants. These formulas decompose Laplace-type operators into geometric components, revealing how curvature influences function properties on manifolds. This linkage supports proving results concerning eigenvalues, regularity, and the index of elliptic operators. By articulating relationships between geometry and analysis, the Bochner-Weitzenböck formulas underscore the manifold's topology and geometry's impact on solving differential equations, serving essential functions in global analysis and index theory applications .

Sobolev spaces are essential in the study of partial differential equations (PDEs) on manifolds as they provide the necessary framework to handle the irregularities of solutions. They extend the concept of differentiation to weaker conditions, which allows solutions that may not be classically differentiable to be captured within this functional space. Sobolev spaces incorporate both the manifold's geometric properties and the function's analytical properties, thus playing an indispensable role in proving existence, uniqueness, and regularity of solutions to PDEs .

You might also like