Electromagnetism as a Gauge Theory
Electromagnetism as a Gauge Theory
1 Introduction
This article is part of a series on physics for mathematicians. My ultimate goal for this series
is to build up to description of the fundamental forces of nature (or at least what we currently
believe them to be) that a certain type of mathematician can find understandable, aesthetically
pleasing, and geometrically natural.
We’ll start by exploring the theory of electromagnetism from four different perspectives:
first as it’s usually presented in a physics class, then rewritten to make the symmetries of special
relativity manifest, then in the framework of Lagrangian mechanics, and finally in terms of a
connection on a principal 𝑈 ( 1) -bundle. This last description, in addition to looking quite a
bit more aesthetically pleasing and less arbitrary than the first, also places electromagnetism
into a class of field theories called gauge theories, which also includes (after quantizing) almost
all of the interactions appearing in the Standard Model of particle physics. In the final section
we’ll briefly describe how this more general theory works, though we won’t touch on any of the
quantum aspects at all.
The prerequisites for this piece are unfortunately a bit steeper than some of the earlier
articles in the series. We’ll depend heavily on the theory of connections on 𝐺 -bundles; there is
an earlier article in this series going over this material. I’m also going to assume some exposure
to Lagrangian mechanics (the material in the first article in this series should be enough) and to
special relativity, though I also include a very brief review of both when they become relevant.
Finally, it might be helpful if you’ve seen Maxwell’s equations in a physics class at some point in
the past, although it’s not at all a requirement.
The mathematical objects used to embed physics in geometry can get basically arbitrarily
complicated, and in the interest of concreteness I’ve stopped well short of maximum generality
in this article; for example, spacetime will always be R 4 with the usual flat metric from special
relativity. As a result, we’ll miss out on a lot of gorgeous geometry and topology whose develop-
ment was intimately connected to the ideas presented here. Some of this can be found in the
sources I list below.
Some sources I found helpful while preparing this article include:
• Gauge Fields, Knots and Gravity by John Baez and Javier P. Muniain. This book starts
“further back” in the chain of prerequisites than I do but covers a lot of the same material
in a thorough and pedagogically skilled way, and I recommend it.
• The Geometry of Physics: An Introduction by Theodore Frankel. This book is large and a
little bit unwieldy, but it’s a good reference for most of the material it covers.
Section 2 Tensorializing Electromagnetism 2
• It’s possible to go much deeper than I do here, and the sort of mathematician who gets
excited about ∞-categories can find a lot to dig into in this area. I don’t have references
for this perspective that are as good as the ones just listed, but you can start with the nLab
pages on gauge theory and fields in physics and follow the references.
• Lagrangian mechanics and the calculus of variations figure very prominently in the ap-
proach I’ve chosen, and probably the most natural way to present this material is as
differential calculus on a jet bundle and an object called the variational bicomplex. One
place to learn this is from this textbook by Ian M. Anderson.
I’m grateful to Yuval Wigderson for many helpful comments on an earlier draft of this article.
2 Tensorializing Electromagnetism
We’ll start with an extended discussion of classical electromagnetism; it will serve as a sort of
prototype for the more general gauge theories we’ll eventually land on. We’ll start by recasting
electromagnetism so that the symmetries of the theory are more apparent than they are in the
usual presentation in physics classes. Partly this is in service of the sections that follow, but
I also think it’s aesthetically pleasing in its own right, and every math student with at least a
passing interest in physics should see it at least once.
This material is standard, so we will be brief. A much more detailed presentation can be
found in Baez and Muniain, which I recommend.
reach another particle and act on it according to the Lorentz force law. Maxwell’s equations
describe the dynamics of E and B, and each equation is commonly associated with a name:
This appears to contradict the fact that constant-velocity motion is a symmetry of physics,
that is, that the laws of physics are preserved by the map
(𝑡 , 𝑥, 𝑦 , 𝑧) ↦→ (𝑡 , 𝑥 − 𝑣𝑡 , 𝑦 , 𝑧),
or versions of this conjugated by a rotation. For a while, most physicists concluded that Maxwell’s
theory must only be valid in one coordinate system; this was “explained” by the idea that
the electric and magnetic fields were disturbances in some physical substance and that the
privileged coordinate system was the one in which this substance was at rest.
But, as you may remember if you’ve studied special relativity, this turned out not to be
consistent with experiment. The theory that won the day says instead that electromagnetism is
in a sense preserved by constant-velocity motion, but that the map above isn’t the right formula
for it. We get our desired symmetry if we instead use the Lorentz transformation:
𝑣𝑥
(𝑡 , 𝑥, 𝑦 , 𝑧) ↦→ 𝛾 𝑡 − 2 , 𝛾 (𝑥 − 𝑣𝑡 ), 𝑦 , 𝑧 ,
𝑐
where 𝛾 = ( 1 − 𝑣 2 /𝑐 2 ) − 1/2 , which preserves the line (𝑡 , 𝑐𝑡 , 0, 0) and therefore the speed 𝑐 . It
looks similar to the “wrong” coordinate change only when 𝑣 /𝑐 is small.
Special relativity is essentially just the statement that the laws of physics are preserved by
Lorentz transformations. Relativistic physics naturally takes place in Minkowski space, which
is R 4 with the metric
⟨−, −⟩ = 𝑐 2 𝑑𝑡 2 − 𝑑𝑥 2 − 𝑑𝑦 2 − 𝑑𝑧 2 .
The group of linear isometries of Minkowski space which preserve both the orientation and the
positive time direction is called the restricted Lorentz group 𝑆𝑂 + ( 1, 3) , and it is generated by
Lorentz transformations and spatial rotations.
We will be using this metric throughout the text to convert between vectors and covectors.
We’ll use the “musical isomorphism” notation: if 𝑣 is a vector and 𝛼 is a covector, then 𝑣 ♭ and 𝛼 ♯
are defined by the relations
𝑣 ♭ (𝑤 ) = ⟨𝑣 , 𝑤 ⟩
⟨𝛼 ♯ , 𝑤 ⟩ = 𝛼 (𝑤 )
for any vector 𝑤 . Physicists refer to these operations as “lowering an index” or “raising an index”
respectively. If you write a covector or a vector in coordinates, both of these operations have the
effect of negating all the spatial components.
The path of a particle is then a function 𝑥 : R → R 4 for which ⟨𝑑𝑥/𝑑𝑠 , 𝑑𝑥/𝑑𝑠 ⟩ ≥ 0 and
𝑑𝑥/𝑑𝑠 points in the forward time direction. We have equality here if and only if the particle is
travelling at the speed of light. (The distinction between this notation and the 𝑥 coordinate
function should hopefully be clear; we will not refer to the latter very much.) If this inequality is
always strict, the length
12
1
∫ 𝑠1
𝑑𝑥 𝑑𝑥
, 𝑑𝑠
𝑐 𝑠0 𝑑𝑠 𝑑𝑠
of some section of a path is called the proper time, written 𝜏 ; you should think of it as the
amount of time that would be recorded on a clock travelling alongside the particle. It’s usually
helpful to use 𝜏 as the parameter for 𝑥 , which amounts to insisting that ⟨𝑑𝑥/𝑑𝑠 , 𝑑𝑥/𝑑𝑠 ⟩ = 𝑐 2 ;
since only the image of the path is physically relevant, nothing is lost by doing this. (The picture
is more complicated in the case of a massless particle, for which ⟨𝑑𝑥/𝑑𝑠 , 𝑑𝑥/𝑑𝑠 ⟩ = 0 everywhere
and therefore 𝜏 is unsuitable as a parameter. For simplicity, we’ll restrict our attention to the
massive case so this issue will never come up.)
Section 2 Tensorializing Electromagnetism 5
Many concepts from nonrelativistic physics have natural generalizations which arise by
replacing 𝑡 derivatives with 𝜏 derivatives. For example, the 4-velocity of the particle is the
vector 𝑢 = 𝑑𝑥/𝑑𝜏 . Analogously, the 4-acceleration and 4-force are 𝑑 2 𝑥/𝑑𝜏 2 and 𝑚𝑑 2 𝑥/𝑑𝜏 2
respectively.
The relativistic analogue of momentum, called the energy-momentum, is defined as 𝑝 =
𝑚 (𝑑𝑥/𝑑𝜏) = 𝑚𝑢 . Note that we are using the convention that “mass” — our 𝑚 — is a coordinate-
independent notion and “energy” is the time component of energy-momentum; some sources
call these “rest mass” and “relativistic mass” respectively, but I think this is confusing.
For a charged particle, we similarly define the charge-current as 𝑗 = 𝑞𝑢 . (There is a slight
inconsistency between these naming conventions: unlike for energy-momentum, the charge is
1 1/2
𝑐 ⟨𝑗 , 𝑗 ⟩ , not the time component of charge-current.) Exactly as for mass, some sources have a
confusing distinction between “rest charge” and “relativistic charge,” where the latter refers to
the time component of the charge-current, but we’ll again reserve the word “charge” for the
coordinate-independent notion.
Continuous charge distributions are represented by a vector field called the charge-current
density 𝐽 = (𝜌, J) . In particular, unlike in the previous paragraph, we do want to identify
the charge density with the time component of a vector, which is not preserved by Lorentz
transformations. Indeed, “charge per unit volume” is not a Lorentz-invariant notion, since
different coordinate systems will disagree about the volume of given region of space.
From now on, we are going to start using units in which 𝑐 = 1.
(Indeed, the cross products that have been appearing all over this discussion indicate that if
we’re going to describe this in terms of tensor fields, a second exterior power ought to show up!)
With 𝐹 in hand we are free to forget about the special coordinate change rule for E and B;
it follows directly from the fact that the electric and magnetic fields are the coordinates of a
2-form on spacetime. When we apply our favorite Lorentz transformation, we get that:
𝑑 2𝑥
𝑚 = (𝜄𝑗 𝐹 ) ♯ .
𝑑𝜏 2
(If 𝛼 is a 2-form and 𝑣 is a vector, recall that 𝜄𝑣 𝛼 denotes the interior product, the 1-form defined
by (𝜄𝑣 𝛼) (𝑤 ) = 𝛼 (𝑣 , 𝑤 ) .)
The four Maxwell equations can be divided into two groups: the ones that don’t refer to
charge and current, and the ones that do. The first group — the absence of magnetic charges
and Faraday’s law — are together equivalent to the single equation
𝑑𝐹 = 0.
We’ll write the other two — Gauss’s and Ampère’s Laws — in terms of the operator 𝑑 ∗ , which
we’ll call the exterior divergence. (It’s also often written as 𝛿 , but we are reserving that symbol
for variations in Lagrangian mechanics.) Since this object might be unfamiliar to some readers
we’ll digress a bit to describe it.
On any 𝑛 -manifold 𝑀 , an inner product ⟨−, −⟩ on a tangent space 𝑇𝑥 𝑀 induces one on
each ∧𝑘 𝑇𝑥∗ 𝑀 , which we’ll also write as ⟨−, −⟩ . If 𝑣 1 , . . . , 𝑣𝑛 is an orthonormal basis of 𝑇𝑥 𝑀 and
𝛼 1 , . . . , 𝛼𝑛 is the dual basis of 𝑇𝑥∗ 𝑀 , then the pure wedges of the form 𝛼𝑖1 ∧ · · · ∧ 𝛼𝑖𝑘 form an
orthonormal basis of ∧𝑘 𝑇𝑥∗ 𝑀 under the induced inner product.
If 𝑀 is oriented, 𝛼 and 𝛽 are 𝑘 -forms, and one of them is compactly supported, we’ll define
∫
(𝛼, 𝛽) = ⟨𝛼, 𝛽⟩ 𝑑 𝑛 𝑥,
𝑀
where 𝑑 𝑛 𝑥 is the volume form on 𝑀 . Note that this new (−, −) product takes entire 𝑘 -forms as
its arguments and produces a single number, whereas ⟨−, −⟩ is an inner product in each fiber
Section 3 A Lagrangian for Electromagnetism 7
separately. We then define 𝑑 ∗ to be the adjoint to 𝑑 under the (−, −) product, that is, if 𝛼 is a
𝑘 -form, we define 𝑑 ∗ 𝛼 by requiring
(𝑑 ∗ 𝛼, 𝛽) = (𝛼, 𝑑 𝛽)
𝑑 ∗ (𝑣 ♭ ) = ★𝑑 ★ ( 𝑓𝑡 𝑑𝑡 − 𝑓𝑥 𝑑𝑥 − 𝑓𝑦 𝑑𝑦 − 𝑓 𝑧 𝑑𝑧)
= ★𝑑 ( 𝑓𝑡 𝑑𝑥 ∧ 𝑑𝑦 ∧ 𝑑𝑧 − 𝑓𝑥 𝑑𝑡 ∧ 𝑑𝑦 ∧ 𝑑𝑧 + 𝑓 𝑦 𝑑𝑡 ∧ 𝑑𝑥 ∧ 𝑑𝑧 − 𝑓 𝑧 𝑑𝑡 ∧ 𝑑𝑥 ∧ 𝑑𝑦 )
= (𝜕𝑡 𝑓𝑡 + 𝜕𝑥 𝑓𝑥 + 𝜕𝑦 𝑓𝑦 + 𝜕𝑧 𝑓 𝑧 ) · ★(𝑑𝑡 ∧ 𝑑𝑥 ∧ 𝑑𝑦 ∧ 𝑑𝑧)
= 𝜕𝑡 𝑓𝑡 + 𝜕𝑥 𝑓𝑥 + 𝜕𝑦 𝑓 𝑦 + 𝜕𝑧 𝑓 𝑧
= div 𝑓 .
Note that the expression for the divergence has plus signs everywhere despite the minus signs
in the definition of the metric. This is because the coefficients on the spatial components
are negated twice: once by the conversion of 𝑣 to a 1-form and once by the Hodge star. It’s
often useful to think of 𝑑 ∗ of a 𝑘 -form as representing a sort of divergence even when 𝑘 > 1; I
encourage you, for example, to repeat this computation with a section of ∧2𝑇 𝑀 and convince
yourself that it resembles a divergence.
At any rate, with this in hand, we can write the remaining Maxwell equations as
𝑑 ∗𝐹 = 𝐽 ♭.
In the nonrelativistic case we mentioned the charge-current conservation law 𝑑 𝜌/𝑑𝑡 = − div J.
This is equivalent to div 𝐽 = 0, which in fact follows directly by applying 𝑑 ∗ to the above equation.
It’s helpful to think of 𝑑 ∗ 𝐹 = 𝐽 ♭ as the natural Lorentz-invariant extension of Gauss’s Law,
which arises as the 𝑡 component of this equation: Gauss’s Law tells us that charges are sources
for the electric field, but if we are set on treating charge as the time component of the charge-
current vector and we want Lorentz invariance, we are forced to conclude that Ampère’s Law
holds as well.
theory in terms of Lagrangian mechanics. We do this for a few reasons: because it will make the
generalization more straightforward, because it puts our theory into the same framework as the
rest of classical physics, and because it will be necessary to have done this when, in a future
article, we build the quantum version of this story.
We’ll start by briefly recalling how nonrelativistic Lagrangian mechanics works. (I’m assum-
ing that the reader has seen this before; what follows is probably not sufficient to learn it for the
first time!) We have a particle moving in R 3 along a path 𝑥 : R → R 3 . Lagrangian mechanics
posits that to every physical situation we might want to model we can associate a Lagrangian
¤ 𝑡 ) , and that the trajectories that are allowed by the laws of physics are the ones that give
𝐿 (𝑥, 𝑥,
critical points of the action
∫
𝑑𝑥
𝑆 [𝑥] = 𝐿 𝑥 (𝑡 ), (𝑡 ), 𝑡 𝑑𝑡 .
𝑑𝑡
(The square bracket notation is often used by physicists to emphasize that the argument is a
function.)
It’s usually not possible to take this completely literally; after all, that integral is almost never
finite. To formalize this condition we consider variations of 𝑥 with compact support, that is,
homotopies ℎ : (−𝜖, 𝜖) × R → R 3 for which ℎ 0 (𝑡 ) = 𝑥 (𝑡 ) for all 𝑡 , and ℎ𝑢 (𝑡 ) = 𝑥 (𝑡 ) for all 𝑡
outside some compact interval. (Here we’re following the common practice of writing the first
argument to ℎ as a subscript.) We then say 𝑥 is a critical point of 𝑆 if, for any such ℎ ,
𝑑
𝑆 [ℎ𝑢 ] = 0.
𝑑𝑢 𝑢=0
While 𝑆 [𝑥] is probably not finite, the difference 𝑆 [ℎ𝑢 ] − 𝑆 [𝑥] will be, since ℎ𝑢 and 𝑥 are equal
outside a compact interval, and this is all that’s needed to make sense of the derivative appearing
in this equation. The original action integral can be thought of as just a formal tool for producing
this equation. One can show that the derivative vanishes for all ℎ if and only if 𝑥 satisfies the
Euler-Lagrange equations
𝑑 𝜕𝐿 𝜕𝐿
= .
𝑑𝑡 𝜕𝑥¤𝑖 𝜕𝑥𝑖
A very important special case comes from considering a particle of mass 𝑚 moving under
the influence of a conservative force, that is, a force F which depends only on the position of the
particle and for which F = − grad 𝑉 for some real-valued function 𝑉 . (We call 𝑉 a potential.)
The correct equations of motion in this case arise from the action
1
∫
𝑆 [𝑥] = 𝑚 𝑥¤ (𝑡 ) 2 − 𝑉 (𝑥 (𝑡 )) 𝑑𝑡 .
2
Our action will have three terms, two of which are∫direct analogues of the two terms in
the nonrelativistic example above. The analogue of the 21 𝑚 𝑥¤ 2 𝑑𝑡 term is straightforward: we
simply replace the velocity with the 4-velocity and set
1
∫
𝑆𝐾 [𝑥] = 𝑚 ¤ 𝑥⟩𝑑𝜏.
⟨𝑥, ¤
2
(The K is for “kinetic.”) Let’s see what happens to 𝑆𝐾 when we vary 𝑥 . We’ll follow the common
physics convention of using 𝛿 to denote (𝑑/𝑑𝑢)|𝑢=0 , so that for example the tangent vector
(𝑑/𝑑𝑢)ℎ𝑢 (𝜏)|𝑢=0 ∈ 𝑇𝑥 (𝜏 ) R 4 is written 𝛿 𝑥 (𝜏) , or just 𝛿 𝑥 . We have
∫ ∫
𝑑 (𝛿 𝑥)
𝛿𝑆𝐾 [𝑥] = 𝑚 , 𝑥¤ 𝑑𝜏 = −𝑚 ⟨𝛿 𝑥, 𝑥⟩𝑑𝜏.
¥
𝑑𝜏
In the first equality, we used the fact that 𝛿 (𝑑𝑥/𝑑𝜏) = 𝑑 (𝛿 𝑥)/𝑑𝜏 , which is just the commutativity
of partial derivatives. In the second, we integrated by parts and used the fact that our variation
vanishes outside of a compact interval to conclude that the boundary term is zero.
Our setting looks a bit different. Where this expression has the 1-form 𝑑𝑉 , we need the
2-form 𝐹 , and we also need it to be paired with the charge-current vector 𝑗 = 𝑞 𝑥¤ . We can take
this as a hint about the form of the term we’re looking for: we should try to find a 1-form 𝐴 for
which 𝑑𝐴 = 𝐹 and pair it with 𝑗 . Such a 1-form is called an electromagnetic potential, and
luckily we know from Maxwell’s equations that 𝑑𝐹 = 0, so, at least locally, it always exists. We
therefore set ∫ ∫
𝑆 int [𝑥] = − 𝐴 (𝑗 ) 𝑑𝜏 = −𝑞 𝑥 ∗ (𝐴).
R
(The “int” is short for “interaction,” since this term describes the interaction between the particle
and the field.)
We can then compute
∫
𝑑
𝛿𝑆 int [𝑥] = − 𝑑𝐴 (𝛿 𝑥, 𝑗 ) + 𝑞 𝐴 (𝛿 𝑥) 𝑑𝜏 ;
𝑑𝜏
this follows from plugging the tangent vectors 𝛿 𝑥 = 𝑑/𝑑𝑢 and 𝑗 = 𝑞 (𝑑/𝑑𝜏) into the definition
of the exterior derivative 𝑑𝐴 . The second term vanishes since it’s the integral of the derivative of
a quantity which is zero outside a compact interval, so we end up with simply
∫ ∫ ∫
− 𝑑𝐴 (𝛿 𝑥, 𝑗 ) 𝑑𝜏 = (𝜄𝑗 𝐹 ) (𝛿 𝑥) 𝑑𝜏 = ⟨𝛿 𝑥, (𝜄𝑗 𝐹 ) ♯ ⟩ 𝑑𝜏.
Section 3 A Lagrangian for Electromagnetism 10
And this is exactly what we wanted: if our action is given by 𝑆𝐾 + 𝑆 int , then 𝑥 is a critical point
if and only if 𝑚 𝑥¥ = (𝜄 𝑗 𝐹 ) ♯ , which is the Lorentz force law.
and moving the 𝜏 integral to the inside. When we do this, we can indeed conclude that 𝑑 ∗ 𝐹 = 𝐽 ♭ ;
I’ll leave the details of the computation as an exercise.
3.4 Summary
All together, our action is:
1 1
∫ ∫ ∫
𝑆 [𝑥, 𝐴] = 𝑚 ¤ 𝑥⟩𝑑𝜏
⟨𝑥, ¤ − 𝐴 (𝑗 ) 𝑑𝜏 − ⟨𝐹 , 𝐹 ⟩ 𝑑 4 𝑥.
2 2
Section 4 Electromagnetism as a Gauge Theory 11
A choice of 𝑥 and 𝐴 gives a critical point of 𝑆 if and only if they satisfy the Lorentz force law
and Maxwell’s equations. One interesting feature of our derivation is the role of the interaction
term −𝐴 (𝑗 ) . This single term is responsible both for the force exerted on the particle by the field
and for the fact that the charge-current acts as a source for the field. It’s often useful to think of
the Hamiltonian/Lagrangian picture of mechanics as “automatically” incorporating Newton’s
Third Law — the one about equal and opposite reactions — and our situation can be seen as
an example: when we write a Lagrangian in which the field acts on a particle we find that the
particle also acts on the field.
We could also have extracted the equations of motion directly from the Euler-Lagrange
equations; they can be applied to the path in the form described above, and there is an analogous
“field-theoretic version” which pertains to things like the variation of 𝐴 . We may talk about how
to apply this machinery more systematically in a future companion piece to this article.
𝑠 ′∗ (𝜔) = 𝑔 (𝑠 ∗ 𝜔)𝑔 −1 + 𝑑 𝑔 · 𝑔 −1 .
(This is Exercise 1 in Section 3 of the connections article. That section uses 𝐴 to refer to what we
are about to start calling 𝑖 𝐴 .)
If we take 𝐺 = 𝑈 ( 1) , so that 𝔤 = 𝔲( 1) = 𝑖 R, then the ambiguity in the choice of electromag-
netic potential takes exactly this form: writing 𝑠 ∗ 𝜔 = 𝑖 𝐴 and 𝑔 (𝑥) = 𝑒 𝑖 𝜙 (𝑥 ) , the formula above
becomes
𝑠 ′∗ (𝜔) = 𝑖 𝐴 + 𝑑 (𝑒 𝑖 𝜙 ) · 𝑒 −𝑖 𝜙 = 𝑖 (𝐴 + 𝑑𝜙).
Section 4 Electromagnetism as a Gauge Theory 12
Ω(𝑣 1 , 𝑣2 ) = 𝑑𝜔 (𝑣 1 , 𝑣2 ) + [𝜔 (𝑣 1 ), 𝜔 (𝑣 2 )].
When the group is abelian, this formula simplifies in two ways: the second term vanishes, and
Ω can be pulled back to give a well-defined 𝔤-valued 2-form on 𝑀 . (In general if 𝑠 and 𝑠 ′ are
two different sections then 𝑠 ∗ Ω and 𝑠 ′∗ Ω differ by conjugating by a 𝐺 -valued function.) For any
section 𝑠 ,
𝑠 ∗ Ω = 𝑠 ∗ (𝑑𝜔) = 𝑑 (𝑠 ∗ 𝜔) = 𝑖 · 𝑑𝐴 = 𝑖 𝐹 ,
so we conclude that the field strength is the curvature of the electromagnetic potential!
Representing the potential with a connection certainly makes the choices more geometri-
cally natural, but this does not mean that we’ve made the electromagentic potential unique!
Any automorphism of the bundle will take our chosen connection to a different connection
with the same curvature. In fact, the fact that the laws of physics don’t care which trivialization
we used means they must also be preserved by automorphisms of this form. This is a good
example of the distinction between “passive” and “active” symmetries; there is an analogous
situation in ordinary Newtonian mechanics: the fact that the laws of physics don’t care where
we put the origin in our coordinate system (passive) means that the laws of physics must also
be preserved by translations (active).
The automorphisms just discussed are called gauge transformations, and the fact that
they preserve the laws of physics is called gauge symmetry. Physicists refer to a choice of
trivialization as choosing a gauge; it’s often helpful when solving certain physical problems to
fix a gauge in some clever way that simplifies the computation, which amounts to imposing
some condition on 𝐴 , just as a clever choice of coordinate system might make it easier to solve
some problem in Newtonian mechanics. (The term “gauge symmetry” can be applied more
generally to any symmetry which can be specified locally in spacetime, but in this article we’ll
only be concerned with this special case.)
The reader may be wondering why we work with the Lie group 𝑈 ( 1) rather than R. For the
classical field theories on R 4 considered in this article, as far as I know there is no reason to
prefer one over the other, but when it comes time to quantize this theory or extend it to cover
more topologically interesting spacetimes, the difference will become relevant, and so we might
as well make the choice now that will still serve us then.
Write 𝑥 ∗ : R → 𝑃 for any horizontal lift of 𝑥 back up to 𝑃 , say the one that passes through 𝑥¯ ( 0) .
If we write
𝑥¯ (𝜏) = 𝑥 ∗ (𝜏) · 𝑖 𝛼 (𝜏)
for some function 𝛼 : R → R, then 𝛼¤ = 1𝑖 𝜔 (𝑥)
¤̄ . Think of 𝛼 as the displacement within the fiber
from where the particle would have been if it were parallel transported along 𝑥 .
This can perhaps serve as motivation for the following definition: we define a metric on 𝑃
according to the rule
1 1
⟨𝑣 , 𝑤 ⟩𝑃 = ⟨𝜋 ∗𝑣 , 𝜋 ∗𝑤 ⟩ +
𝜔 (𝑣 ) 𝜔 (𝑤 ) .
𝑖 𝑖
In other words, we use the metric from the base for horizontal vectors, the unique 𝑈 ( 1) -invariant
metric of total length 2𝜋 for the vertical vectors, and make horizontal and vertical vectors
orthogonal to each other.
Our new action is then:
1 1
∫ ∫
¯ 𝜔] = 𝑚
𝑆 [𝑥, ¤̄ ¤̄
⟨𝑥, 𝑥⟩𝑃 𝑑𝜏 + ⟨𝐹 , 𝐹 ⟩𝑑 4 𝑥.
2 2
(The claim is not that this is the same as our old action, just that it produces the same equations
of motion! The old action, in fact, doesn’t respect gauge symmetry, so it would be no good here.)
To write this action we are taking advantage of the fact that, since the group is abelian, 𝐹 = 1𝑖 𝑠 ∗ Ω
is independent of the choice of section 𝑠 .
The most natural way to set up the calculus of variations is as differential calculus on a jet
bundle, but I am deliberately avoiding introducing this level of complexity in this article. This
formalism would enable us to vary the connection directly and write the variation 𝛿𝑆 in way
that’s manifestly independent of choices. We will instead pick a section 𝑠 of 𝑃 and use it to
write all the quantities appearing in the action as functions on R 4 as in the previous section; the
resulting equations of motion won’t depend on 𝑠 . As before, we’ll write 𝑖 𝐴 = 𝑠 ∗ 𝜔 . We get
∫ h i ∫
𝛿𝑆 = 𝑚 ¤ 𝑑𝜏 − ⟨𝛿 𝐴, 𝑑 ∗ 𝐹 ⟩𝑑 4 𝑥.
−⟨𝛿 𝑥, 𝑥¥ − 𝛼¤ (𝜄𝑥¤ 𝐹 ) ♯ ⟩ + 𝛿 𝛼 · 𝛼¥ + (𝛿 𝐴) ( 𝛼¤ 𝑥)
particle causes all sorts of mathematical headaches. I have stuck with the particle so far because
I think the resulting geometric picture with the Lorentz force law and the geodesics is more
concrete. But the math is quite a bit nicer if the matter takes the form of a field as well, and it
will be this “everything is fields” version of the theory that we will eventually want to quantize.
(□ + 𝑚 2 )𝜙 = 0,
where
𝜕2 𝜕2 𝜕2 𝜕2
□ = 𝑑 ∗𝑑 = 2
− 2− 2− 2
𝜕𝑡 𝜕𝑥 𝜕𝑦 𝜕𝑧
is the d’Alembertian operator. This should be thought of as an equation of motion for the field:
if you know the values of 𝜙 and 𝜕𝜙/𝜕𝑡 on a time slice, this equation tells you how to evolve
them forward or backward in time. Our goal will be to build a theory of Klein-Gordon fields that
interact with electromagnetism.
This equation has a solution of the form
𝜙 (𝑥) = 𝑒 𝑖 ⟨𝑝,𝑥 ⟩
for any vector 𝑝 for which ⟨𝑝, 𝑝⟩ = 𝑚 2 . These are called plane wave solutions. This property
means that either 𝑝 or −𝑝 — depending on the sign of the 𝑡 component — is a 4-momentum for
a particle of mass 𝑚 . It’s useful to think of these solutions as being like a “massive version” of
the light waves we got as vacuum solutions to Maxwell’s equation. (Another difference is that
𝜙 is a scalar while 𝐴 is a vector.) A general solution can written as an integral over plane wave
solutions.
The Klein-Gordon equation can be extracted from the action
∫
𝑆 [𝜙] = ⟨𝑑𝜙, 𝑑𝜙⟩ − 𝑚 2 𝜙 𝜙¯ 𝑑 4 𝑥.
Note again the similarity to the Maxwell theory; the differences are the complex conjugates
(which are there to make the action take real values), the presence of the mass term 𝑚 2 𝜙 𝜙¯ , and
the fact that 𝜙 is a scalar. Note that if we separate 𝜙 into real and imaginary parts, we can rewrite
this action as a sum of two similar expressions, one for each part, reflecting the fact that asking
𝜙 to satisfy the Klein-Gordon equation is equivalent to asking for its real and imaginary parts to
both do so separately, not interacting with each other at all.
Section 4 Electromagnetism as a Gauge Theory 15
wouldn’t be especially interesting; solutions of the resulting theory would just be Klein-Gordon
fields together with vacuum solutions to Maxwell’s equations, evolving separately with no
interaction. There’s one feature of our action that will turn out to be the key to adding a more
interesting interaction to the theory: the fact that 𝑆 [𝜙] is preserved by the action of 𝑈 ( 1) on C.
Since we’ve seen that the electromagnetic potential can be naturally represented as a connection
on a principal 𝑈 ( 1) -bundle 𝑃 , this suggests a way to produce the interaction we want. We’ll let
𝐸 be the associated vector bundle to 𝑃 arising from the action of 𝑈 ( 1) on C, and then “upgrade”
our field 𝜙 from a complex-valued function to a section of 𝐸 .
Our electromagnetic potential induces a connection on 𝐸 , which we can immediately put
to use. Once 𝜙 is a section of 𝐸 , the “𝑑𝜙 ” appearing the action is no longer a well-defined
mathematical object, but we can replace it with the covariant derivative ∇𝜙 . (A covariant
derivative without a subscript like this denotes the 𝐸 -valued 1-form (𝑣 ↦→ ∇𝑣 𝜙) .) It’s common
to multiply the action of 𝔲( 1) on C by a constant 𝑞 when building 𝐸 and its induced connection;
this constant plays an analogous role to the charge of the particle, controlling the strength of the
interaction with the electromagnetic field. So our action becomes (before and after choosing a
trivialization):
∫
𝑆 [𝜙, 𝜔] = ⟨∇𝜙, ∇𝜙⟩ − 𝑚 2 𝜙 𝜙¯ + ⟨𝐹 , 𝐹 ⟩ 𝑑 4 𝑥
∫
= ¯ − 𝑚 2 𝜙 𝜙¯ + ⟨𝐹 , 𝐹 ⟩ 𝑑 4 𝑥.
⟨𝑑𝜙 + 𝑖𝑞𝐴𝜙, 𝑑𝜙 − 𝑖𝑞𝐴 𝜙⟩
(While it doesn’t show up explicitly, 𝜔 is present in the definition of both ∇ and 𝐹 , and 𝑞 is
implicit in the definition of ∇.) Note that even though we can’t canonically identify the fibers
of 𝐸 with C, expressions like 𝜙 𝜙¯ are still well-defined because 𝐸 is a 𝑈 ( 1) -bundle and 𝑈 ( 1)
respects the Hermitian metric on C. This all works out precisely because our original action was
written in terms of quantities that were preserved by the 𝑈 ( 1) action.
I encourage you to check that, using this action, the equation of motion for the Klein-Gordon
field becomes
(∇∗ ∇ + 𝑚 2 )𝜙 = 0,
where ∇∗ is defined analogously to 𝑑 ∗ in a way whose details I am leaving for you to fill in. For
the electromagnetic field we get
♭
𝑑 ∗ 𝐹 = 𝑖𝑞 𝜙¯ ∇𝜙 − 𝜙 ∇𝜙 .
The quantity on the right side of this last equation can therefore be called the charge-current
density of the Klein-Gordon field and written 𝐽 ; just as when we discussed Maxwell’s equations,
applying 𝑑 ∗ to both sides produces a conservation law for this quantity. Note that both 𝜙 and
the connection appear in both of these equations, so the time evolution of each depends on the
other. We say that we’ve coupled the Klein-Gordon field to electromagnetism, and 𝑞 is called
the coupling constant.
While we’ve only worked out this one example, its essential features give us a sort of recipe
for coupling a field theory to electromagnetism, and this recipe is widely applicable: find a 𝑈 ( 1)
symmetry of a field theory with values in some vector space, build a vector bundle out of this
action and induce a connection on it using the electromagnetic potential, and finally allow your
fields to take values in this bundle, replacing any ordinary derivatives with covariant derivatives.
Section 5 Yang-Mills Theory 16
5 Yang-Mills Theory
Over the course of this article, we’ve built up a description of electromagnetism in which the
potential takes the form of a connection on a principal 𝑈 ( 1) -bundle. This is worth doing partly
just for the nice geometric picture that it produces, but there is a deeper reason to present
electromagnetism in this way: it’s this picture that directly generalizes to the other interactions
in the Standard Model of particle physics.
The generalization is quite simple: we replace 𝑈 ( 1) with an arbitrary compact Lie group 𝐺 .
The resulting field theories are called Yang-Mills theories. The weak interaction corresponds to
the choice 𝐺 = 𝑆𝑈 ( 2) , and the strong interaction to 𝐺 = 𝑆𝑈 ( 3) . These groups are nonabelian,
which affects several aspects of the resulting field theory. One of them, unfortunately, is that
quantum effects become important enough that the classical version of the theory is no longer
a good physical model for anything in the real world, which severely limits the number of useful
things we can say here. Still, in this brief final section we’ll discuss a few of the changes we have to
make — and that do carry over to the quantum theory — when we move from electromagnetism
to a nonabelian Yang-Mills theory. We’ll mostly stick to the “single charged particle” version of
the theory for simplicity.
assign it a role analogous to electromagnetic charge. But new complications arise when we
try to describe the motion of the particle directly in terms of the path in R 4 (rather than in 𝑃 ).
Given any geodesic 𝑥¯ (𝜏) on 𝑃 and any ℎ ∈ 𝐺 , the path 𝑦¯ (𝜏) = 𝑥¯ (𝜏) · ℎ is also a geodesic and
projects down to the same path on the base, but 𝜔 ( 𝑦¤̄ ) = Ad ℎ − 1 · 𝜔 (𝑥) ¤̄ .
This means that, if we’d like to write the laws of motion just in terms of the projected path
𝑥 (𝜏) = 𝜋 (𝑥¯ (𝜏)) in R 4 , we can’t assign our particle a “charge” in 𝔤 in a well-defined way. The
charge is instead naturally a section of the vector bundle Ad 𝑃 := 𝑃 ×𝐺 𝔤, where 𝔤 carries the
adjoint action of 𝐺 . (Indeed, (𝑥, ¤̄ and (𝑥¯ · ℎ, Ad ℎ −1 · 𝜔 (𝑥))
¯ 𝜔 (𝑥)) ¤̄ are the same point in Ad 𝑃 by
definition.) We are therefore free to define 𝑞 (𝜏) = 𝑚𝜔 (𝑥 (𝜏)) as long as we think of this as a
¤̄
point of Ad 𝑃 lying above 𝑥 (𝜏) .
It no longer makes sense to say that 𝑞 is a constant. After all, it lives in different fibers of Ad 𝑃
at different times. But we do have the next best thing: I encourage you to check that 𝑞 is parallel
transported along 𝑥 under the connection on Ad 𝑃 induced by 𝜔 .
The curvature Ω satisfies 𝑅 𝑔∗ Ω = Ad 𝑔 − 1 · Ω and this, together with the fact that it vanishes on
vertical vectors, means we’re free to regard it as an Ad 𝑃 -valued 2-form on R 4 . So, even though
neither 𝑞 nor Ω can be naturally identified with an element of 𝔤, they take values in the same
bundle, and this is all we need for something like the Lorentz force law to make sense. The
equation of motion for the particle that arises from our action is
𝑚 𝑥¥ = 𝜅 (𝑞, (𝜄𝑥¤ Ω) ♯ ).
𝑞¤ = [𝑞, 𝐴 (𝑥)]
¤
𝑚 𝑥¥ = 𝜅 (𝑞, (𝜄𝑥¤ 𝐹 ) ♯ )
In electromagnetism, picking a gauge only really mattered for writing down the action; the
equations of motion themselves only referred to the gauge-invariant quantities 𝑞 and 𝐹 . This
is no longer true in the nonabelian case! If, as is helpful for many computations, we want to
identify all these objects with functions landing in a vector space rather than sections of some
bundle, we can’t forget about the gauge symmetry even for the equations of motion: 𝑞 and 𝐹
now depend on the gauge, and the equations of motion also involve 𝐴 directly.
𝐷𝛼 (𝑋 1 , . . . , 𝑋𝑘 +1 ) = 𝑑𝛼 (𝑋 1𝐻 , . . . , 𝑋𝑘𝐻+1 ).
The right generalization of the equation 𝑑𝐹 = 0 from electromagnetism — the one which
depends only on the existence of the potential 𝐴 and not on anything else about the action — is
the second Bianchi identity, which says that 𝐷Ω = 0.
To write the analogue of the other Maxwell equation, by analogy with 𝑑 ∗ , we define an
operator 𝐷 ∗ by the rule ∫ ∫
⟨𝐷 ∗ 𝛼, 𝛽⟩𝜅 = ⟨𝛼, 𝐷 𝛽⟩𝜅 .
Section 5 Yang-Mills Theory 18
If we define the charge-current as the Ad 𝑃 -valued vector 𝑞 (𝜏) 𝑥¤ (𝜏) and form a charge-current
density 𝐽 with delta functions in the usual way, then the other equation arising from our action
is
𝐷 ∗Ω = 𝐽 ♭.
This is called the Yang-Mills equation.
If we pick a gauge, the equation becomes
𝑑 ∗ 𝐹 + [𝐴, 𝐹 ] = 𝐽 ♭ ,
where the not especially good notation [𝐴, 𝐹 ] refers to the 𝔤-valued 1-form formed by first using
the metric to contract 𝐴 with 𝐹 , forming the 𝔤 ⊗ 𝔤-valued 1-form 𝜄𝐴 ♯ 𝐹 , and then applying the Lie
bracket. (This is a situation where there is some virtue to the parade of indices that physicists
use to work with these objects.) In particular, once again 𝐴 appears in the equations of motion,
not just 𝐹 .
The “charged field” story from earlier also works in this more general setting. If we start with
a field theory that takes values in some vector space 𝑉 with a representation of 𝐺 , then just as
before we can move it to the vector bundle 𝑃 ×𝐺 𝑉 by replacing all ordinary derivatives with
covariant derivatives. Most of the details are unchanged.
One difference in the nonabelian case worth highlighting is the status of the charge-current
density 𝐽 . There will still be an equation of the form 𝐷 ∗ Ω = 𝐽 ♭ , and we can use this as the
definition of 𝐽 . This means that 𝐽 is an “Ad 𝑃 -valued vector field,” that is, a section of 𝑇 R 4 ⊗ Ad 𝑃 .
We can still extract a sort of conservation law from the Yang-Mills equation: applying 𝐷 ∗ to both
sides gives us that 𝐷 ∗ (𝐽 ♭ ) = 0. This follows the fact that 𝐷 ∗ 𝐷 ∗ Ω = 0, which I encourage you to
check. (Note that this relation is special to Ω; it is not the case that (𝐷 ∗ ) 2 = 0 in general!)
A very important difference between the Yang-Mills and Maxwell equations is that even
when 𝐽 = 0, the Yang-Mills equation isn’t linear, so in particular we can’t solve it by adding
together simple solutions like the light waves in electromagnetism. This nonlinearity is a major
source of headaches, especially for the quantum version of the theory — for example, it means
that gluons, the strong-force analogue of photons, interact with each other as well as with
charged particles — and there are a large number of both mathematical and physical open
problems surrounding it.