Symmetry and Relativity in Physics
Symmetry and Relativity in Physics
University of Oxford
Third Year, Part B2
Caroline Terquem
Department of Physics
[Link]@[Link]
1 Symmetries 7
1.1 Newton’s law of motion and continuous symmetries . . . . . . . . . . . . . . 8
1.1.1 Translation in space . . . . . . . . . . . . . . . . . . . . . . . . . . . 9
1.1.2 Rotation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9
1.1.3 Translation in time . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10
1.2 Lagrangian and Hamiltonian formalisms . . . . . . . . . . . . . . . . . . . . 10
1.2.1 Generalised coordinates . . . . . . . . . . . . . . . . . . . . . . . . . 11
1.2.2 D’Alembert’s principle . . . . . . . . . . . . . . . . . . . . . . . . . . 11
1.2.3 Euler-Lagrange equations . . . . . . . . . . . . . . . . . . . . . . . . 13
1.2.4 The principle of least action . . . . . . . . . . . . . . . . . . . . . . . 16
1.2.5 Lagrangian for an electromagnetic force . . . . . . . . . . . . . . . . 19
1.2.6 Field theory for continuous systems . . . . . . . . . . . . . . . . . . 20
1.3 Conservation laws and Noether’s theorem . . . . . . . . . . . . . . . . . . . 24
1.3.1 Cyclic coordinates . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24
1.3.2 Noether’s theorem . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25
1.3.3 Applications . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27
3
3 Fundamental concepts in special relativity 51
3.1 Deriving the Lorentz transformation for a boost . . . . . . . . . . . . . . . . 52
3.2 Spacetime diagram . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 53
3.3 Relativistic effects on measurements . . . . . . . . . . . . . . . . . . . . . . 54
3.3.1 Proper length and length contraction . . . . . . . . . . . . . . . . . . 54
3.3.2 Proper time and time dilation . . . . . . . . . . . . . . . . . . . . . . 55
3.4 The invariance of the spacetime interval . . . . . . . . . . . . . . . . . . . . 56
3.4.1 Classification of worldlines . . . . . . . . . . . . . . . . . . . . . . . . 56
3.4.2 Causality in relativity . . . . . . . . . . . . . . . . . . . . . . . . . . 57
3.5 Relativistic velocity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 58
3.5.1 Velocity composition . . . . . . . . . . . . . . . . . . . . . . . . . . . 58
3.5.2 Velocity transformation . . . . . . . . . . . . . . . . . . . . . . . . . 58
3.6 Four-vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 59
3.6.1 Overview and definition . . . . . . . . . . . . . . . . . . . . . . . . . 59
3.6.2 Minkowski metric: length and scalar product . . . . . . . . . . . . . 60
3.6.3 Index notation, covariant and contravariant four-vectors . . . . . . . 62
3.6.4 Covariant and contravariant derivatives . . . . . . . . . . . . . . . . 66
3.6.5 Transformations of four-vectors in index notation . . . . . . . . . . . 69
4
5 Special relativity and electromagnetism 91
5.1 Tensors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 91
5.1.1 Preliminary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 91
5.1.2 Definition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 92
5.1.3 Operations on tensors . . . . . . . . . . . . . . . . . . . . . . . . . . 93
5.1.4 Symmetric and antisymmetric tensors . . . . . . . . . . . . . . . . . 96
5.2 Invariance of the electric charge . . . . . . . . . . . . . . . . . . . . . . . . . 96
5.3 Four-vectors for the current and potential . . . . . . . . . . . . . . . . . . . 96
5.3.1 Definition of the four-current . . . . . . . . . . . . . . . . . . . . . . 96
5.3.2 Conservation law . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 97
5.3.3 Lorentz gauge . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 97
5.3.4 Definition of the four-potential . . . . . . . . . . . . . . . . . . . . . 99
5.4 Electromagnetic tensor . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 100
5.4.1 Definition of the electromagnetic tensor . . . . . . . . . . . . . . . . 101
5.4.2 Maxwell’s equations in terms of the electromagnetic tensor . . . . . 102
5.5 Lorentz transformation of the electromagnetic field . . . . . . . . . . . . . . 102
5.6 The Lorentz force . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 103
5.7 Lagrangian formulation of electromagnetism . . . . . . . . . . . . . . . . . . 104
5.7.1 Covariant relativistic Lagrangian for a particle . . . . . . . . . . . . 104
5.7.2 Euler-Lagrange equations . . . . . . . . . . . . . . . . . . . . . . . . 106
5.7.3 Lagrangian for the electromagnetic field . . . . . . . . . . . . . . . . 108
5.8 Energy conservation for fields in vacuum . . . . . . . . . . . . . . . . . . . . 108
5.8.1 Noether’s theorem for the electromagnetic field . . . . . . . . . . . . 108
5.8.2 Canonical energy-momentum tensor . . . . . . . . . . . . . . . . . . 109
5.8.3 Symmetric energy-momentum tensor . . . . . . . . . . . . . . . . . . 111
5
These notes borrow from the following books (with the chapters of the lecture notes for
which the books are most relevant indicated in parentheses):
A. Zee, Group Theory in a Nutshell for Physicists, Princeton University Press, 2016
(chapter 2)
L. Susskind and A. Friedman, Special Relativity and Classical Field Theory, The
Theoretical Minimum, Penguin Books, 2017 (chapters 3, 4 and 5)
A. M. Steane, Relativity Made Relatively Easy, Oxford University Press, 2012 (chap-
ters 3 and 4)
These notes are intended to supplement the course, but they should not be seen as a
substitute for textbooks. It is highly recommended to regularly consult textbooks, as they
offer much more detailed explanations, along with numerous examples and problems.
6
Chapter 1
Symmetries
Symmetries underlie all fundamental physical laws. They play a crucial role in shaping
the theory of quantum mechanics, and lead to the theory of relativity once the invariance
of the speed of light is assumed.
The most straightforward way to define a symmetry is to use Feynman’s rewording of
Weyl1 : a thing is symmetrical if there is something we can do to it so that after we have
done it, it looks the same as it did before2 .
In classical physics, a symmetry is a transformation that leaves the equation of motion
unchanged, meaning it transforms one possible motion into another while leaving all mea-
surable quantities (observables) unaffected. Symmetries therefore restrict the solutions
of the equations of motion, and can be used to put constraints on these solutions. For
example, consider the rotation around the axis of angular momentum of a planet orbiting
a star: rotating the entire orbit (which may be elliptical) results in another valid orbit
with the same properties.
In quantum mechanics, a symmetry not only transforms allowed states into other allowed
states but also enables the formation of new states through superpositions of these trans-
formed states. These superpositions still respect the symmetry, offering a broader set of
solutions while maintaining the underlying symmetry. For example, if a system is rota-
tionally invariant, applying a rotation to a state |ψ > yields a new state R|ψ >. The
superposition of these two states, |ψ > +R|ψ >, is itself a valid quantum state with the
same symmetry properties. This contrasts with classical mechanics, where there is no way
to combine two trajectories to produce another valid trajectory.
Another important symmetry in quantum mechanics involves the exchange of particles.
Bosons are characterised by the fact that exchanging two bosons leaves the wavefunction
unchanged, whereas fermions exhibit the opposite behaviour — exchanging two fermions
causes the wavefunction to change sign.
1
Hermann Weyl, Symmetry, 1952, Princeton University Press, Princeton, NJ
2
Richard P. Feynman, Robert B. Leighton, Matthew Sands, 1963, The Feynman Lectures on Physics –
Volume 1, Chap. 52 Symmetry in Physical Laws, Addison-Wesley
7
These symmetries are examples of global symmetries, meaning the transformations do not
depend on position in space.
There is another type of invariance which is called gauge symmetry: it is related to the
fact that fundamental laws can be described in different ways. Different descriptions are
related by a transformation which is a function of position. The transformations associated
with gauge symmetries do not change the physical state of the system; rather, they give a
different but precisely equivalent mathematical description of the same physical situation,
which depends on the coordinates. The introductory review The role of symmetry in
fundamental physics by David J. Gross3 is recommended for more details.
A symmetry is either continuous or discrete. Continuous symmetries involve transfor-
mations that can be described using one or more continuous parameters, resulting in an
infinite number of possible transformations. Examples include rotations around a point or
an axis, where the continuous parameter is the rotation angle, and translations, where the
continuous parameter is the distance by which objects or fields are shifted. Lorentz sym-
metries, that include both rotations and boosts (changes in velocity) are also continuous
symmetries. According to Noether’s theorem, every continuous symmetry corresponds to
a conserved quantity.
Discrete symmetries involve transformations that cannot be described by any contin-
uous parameter, resulting in a finite or countably infinite number of possible transforma-
tions. Examples include parity symmetry, such as reflection through the origin (r → −r),
time reversal (t → −t) and charge conjugation (transformation that changes particles
into their antiparticles). Crystal symmetries are also discrete because they are defined
by specific mathematical operations that map the crystal onto itself. For instance, if the
periodicity of a crystal lattice is R, only translations that shift the crystal by an integer
multiple of R will leave the crystal unchanged. Unlike continuous symmetries, discrete
symmetries are not associated with conservation laws. Nevertheless, they can still impose
constraints and selection rules on physical processes. For example, the requirement of
parity symmetry in quantum systems can prohibit certain transitions.
In what follows, we will focus on continuous symmetries.
In this section, we illustrate how continuous symmetries imply conservation laws starting
from Newton’s law of motion. We will show that invariance under translation, rotation,
and time translation corresponds respectively to conservation of momentum, angular mo-
mentum, and energy.
Consider an isolated system of two masses m1 and m2 with position vectors r1 (t) and
3
Proceedings of the National Academy of Sciences of the USA, 1996, p. 14256–14259
8
r2 (t) which exert the forces f1 and f2 on each other. The equations of motion are:
where the dots represent time derivatives. We assume that the forces are conservative and
note U the interacting potential, such that fi = −∇ri U , where i = 1, 2 and the subscript
denotes the coordinates with respect to which the derivatives are taken.
Since t0 is arbitrary, this shows that U depends solely on the relative position r = r1 − r2
of the particles. Consequently, f1 = −f2 (weak form of the law of action and reaction),
and the forces themselves depend only on r. The equations of motion then yield:
d
(m1 ṙ1 + m2 ṙ2 ) = 0. (1.2)
dt
Invariance under translation therefore implies conservation of momentum.
1.1.2 Rotation
We now assume that the system is invariant under any rotation around a specific axis.
This implies that the force f1 applied to the rotated r is equal to the rotation of f1 (r) (and
similarly for f2 ). This can only be satisfied if the magnitude |f1 | depends solely on the
modulus r of r, and not on its direction (otherwise, rotating r could alter the magnitude
of the force). Using spherical coordinates, the potential U (r) can be written as U (r, θ, ϕ),
with the time dependence implicit. If ∇U depends only on r, this implies that U can be
expressed as U = g(r) + Cθ, where g is a function of r alone and C is a constant. There
cannot be any additional dependence on ϕ, because such dependence would introduce a
θ-dependence in the ϕ-component of the gradient.
This results in the force f1 = fr (r) r̂ − (C/r) θ̂, where a hat denotes a unit vector. This
suggests that, in principle, the force could have a tangential component. However, such a
component would imply that work is done as the particle moves around a circle, violating
9
rotational invariance. Therefore, C = 0, implying that U depends only on r, and the force
is directed along r. This conclusion means that internal forces are central, which is the
strong form of the law of action and reaction.
In summary, rotational invariance implies that the force is central. It follows that:
d
(m1 r1 ×ṙ1 + m2 r2 ×ṙ2 ) = r1 ×f1 + r2 ×f2 = (r1 − r2 ) ×f1 = 0. (1.3)
dt
Here we assume that the system is invariant under a time translation that shifts t back-
wards by some arbitrary interval τ . This implies fi (r1 , r2 , t − τ ) = fi (r1 , r2 , t). Therefore,
the change in U associated with a time translation can only be a function of t, and possibly
τ , which we denote as g: U (r1 , r2 , t − τ ) = U (r1 , r2 , t) + g(t, τ ). The function g can be
chosen arbitrarily because it does not affect the accelerations. Therefore, we take g = 0.
The equality above is satisfied in particular at time t = τ , so U (r1 , r2 , 0) = U (r1 , r2 , τ ) .
Since τ is arbitrary, this implies that U does not explicitly depend on time (it only depends
on time through r1 and r2 ). It follows that:
dU dK
= ṙ1 · ∇r1 U + ṙ2 · ∇r2 U = −ṙ1 · f1 − ṙ2 · f2 = − , (1.4)
dt dt
Invariance under time translation therefore implies conservation of the total energy K+U .
The examples above illustrate how continuous symmetries lead to conservation laws. How-
ever, they also show that Newton’s law of motion is not the most effective approach for
deriving general principles about these laws. For this purpose, we use the Lagrangian and
Hamiltonian formalisms, which were introduced in the first-year Mechanics course4 . In
that course, the Euler-Lagrange equations were derived from the principle of least action.
However, this does not explain why the Lagrangian is equal to the kinetic energy minus
the potential energy for conservative forces. Noether’s theorem was also introduced, but
only for the limited case of transformations of the generalised coordinates, with time held
constant. Here, we will derive the Euler-Lagrange equations from d’Alembert’s principle,
which offers the advantage of enabling us to calculate the Lagrangian. Additionally, we
will derive Noether’s theorem for the more general case where transformations of both the
generalised coordinates and time are considered.
4
See also Prof. David Marshall’s S7/B7 Classical Mechanics Lecture Notes
10
1.2.1 Generalised coordinates
We consider a system composed of N particles, each with mass mi and position vector
ri (t). Constraints may exist, implying that the ri are not all independent.
One type of constraint, known as holonomic, can be expressed as algebraic equations
involving the coordinates of the particles and, possibly, time. For example, a particle
restricted to move along a circular path of radius R must satisfy x2 + y 2 = R2 . Another
example is the rolling without slipping condition for a cylinder of radius R along a table
in one dimension. The relevant coordinates are x, the distance along the table, and θ, the
angle of rotation of the cylinder. The rolling without slipping condition is ẋ = Rθ̇, which
integrates into x − Rθ = 0. In such cases, we can introduce a set of generalised coordinates
qi (t), with i ranging from 1 to some integer less than 3N , by eliminating the constraints.
These coordinates, which do not necessarily have the units of length, are generally not
simply related to the position vectors, but they can be convenient to use even in the
absence of constraints. We refer to q̇i = dqi /dt as generalised velocities, although they do
not necessarily have the units of conventional velocity.
Other types of constraints are referred to as nonholonomic. An example is the rolling
of a disc with radius R on a two-dimensional table without slipping, where the plane of
the disc stays vertical, but the disc can freely rotate about a vertical axis. The relevant
coordinates are x and y, which represent the position of the disc’s centre on the table,
and θ, the rotation angle of the disc. The rolling without slipping condition gives the
velocity-dependent constraints ẋ = Rθ̇ cos α and ẏ = Rθ̇ sin α, where α is related to the
direction of motion. These conditions cannot be integrated into a relation involving only
the coordinates, and thus do not reduce the number of generalised coordinates. Another
case is motion restricted to a certain region of space, which imposes a condition on the
coordinates without reducing their number. For example, a particle confined to move
inside a sphere of radius R must satisfy x2 + y 2 + z 2 < R2 .
Although the rolling without slipping condition in two dimensions does not reduce
the number of generalised coordinates, it makes them dependent on each other through
their time derivatives. For example, if x and θ are known, y can be calculated through
ẏ 2 = R2 θ̇2 − ẋ2 . The generalised coordinates are only independent of each other if all the
constraints are holonomic.
Constraints are often enforced by forces of constraint, such as the friction force in the
case of rolling without slipping, or forces that keep particles confined within a certain
region of space.
Forces of constraint restrict motion to certain directions or surfaces, ensuring that the
configuration of the system remains consistent with those constraints. For example, a
normal force prevents penetration through a surface, tension in a rod prevents elongation
or compression and, in rolling motion without slipping, friction prevents sliding at the
contact point. If slipping were allowed, friction could do work, introducing non-ideal
11
constraints beyond this analysis.
We now introduce the concept of virtual displacements, δri , which are infinitesimal,
hypothetical changes in the positions ri of the particles of the system. These displacements
are ‘frozen in time’, meaning they represent possible instantaneous changes in configuration
that are consistent with the constraints of the system, without considering the actual
motion or dynamics.
For example, for a bead sliding on a rigid ring, a virtual displacement lies along the
circumference of the ring, respecting the constraint that the bead cannot leave it. The
normal force from the ring is perpendicular to this displacement and therefore performs
no work. This principle extends to more complex systems, where the geometry of the
constraints determines the possible directions of virtual displacements.
In general, however, constraint forces need not be perpendicular to the motion or to the
virtual displacement of each particle individually. For example, in an Atwood machine, the
tension in the string does virtual work on each mass separately, but the two contributions
cancel because the virtual displacements are related by the constraint. The hallmark of
an ideal constraint is that the total virtual work of all constraint forces on the system
vanishes:
X
Fc,i · δri = 0. (1.5)
i
This is the defining property of ideal constraints used in d’Alembert’s principle: although
constraint forces may exchange energy between parts of the system, they do no net virtual
work on the system as a whole.
The figure below5 illustrates the concept of a virtual displacement for a pendulum
whose pivot moves in time, along the trajectory x0 (t). The rod is rigid, so a virtual
displacement δr at time t0 moves the mass along the circular arc defined by the position
of the pivot at that time. In contrast, the actual displacement dr occurs while the pivot
itself moves, so the trajectory of the mass is not exactly circular.
5
From X. Jaén, J. Salud, C. Serra, J. Calaf. M. Khoury, 2023, Fundamental Mechanics, Newtonian
Mechanics for Engineering, Iniciativa Digital Politècnica, Universitat Politécnica de Catalunya, Barcelona
12
Virtual displacements describe all possible instantaneous changes consistent with the cons-
traints, independent of the dynamics of the system. This makes them particularly conve-
nient in the formulation of d’Alembert’s principle, which can then be expressed without
prior knowledge of the actual motion.
Each particle i is subject to both forces of constraint and additional forces, not related
to the constraints, denoted by Fnc,i . From Newton’s law of motion: Fc,i +Fnc,i −mi r̈i = 0.
Taking the dot product with δri , summing over all particles i and applying equation (1.5)
leads to the conclusion that, for any set of virtual displacements δri :
N
X
(Fnc,i − mi r̈i ) · δri = 0. (1.6)
i=1
13
where the generalised force Fk is defined as:
X ∂ri
Fk = Fnc,i · (1.9)
∂qk
i
While Fk may not have the units of a force, Fk δqk does have the units of work.
We now consider the term involving acceleration in equation (1.6):
X X ∂ri X d ∂ri
d ∂ri
mi r̈i · δri = mi r̈i · δqk = mi ṙi · − mi ṙi · δqk (1.10)
∂qk dt ∂qk dt ∂qk
i i,k i,k
We have:
!
d ∂ri X ∂ ∂ri
∂ ∂ri
∂ X ∂ri ∂ri
= q̇l + = q̇l + (1.11)
dt ∂qk ∂ql ∂qk ∂t ∂qk ∂qk ∂ql ∂t
l l
| {z }
ṙi
In deriving the last equality, we used ∂ q̇l /∂qk = 0. When we perform a virtual displace-
ment, we take a configuration of the system characterized by the generalised coordinates
qk and velocities q̇l = dql /dt, and change qk by δqk instantaneously. Since there is no
time evolution, the generalised velocities q̇l are not affected by the changes in qk . We are
therefore effectively considering that q̇l is independent of qk during the displacement, and
thus ∂ q̇l /∂qk = 0, even if along a real trajectory q̇l generally depends on the qk .
where we retain the time derivative since we are calculating the velocity, not a change
due to a virtual displacement. The velocity ṙi is then a linear function of the generalized
velocities q̇k , which leads to the relation:
∂ri ∂ ṙi
= (1.13)
∂qk ∂ q̇k
Substituting this expression and equation (1.11) into equation (1.10) gives:
X X d ∂ ṙi
∂ ṙi
mi r̈i · δri = mi ṙi · − mi ṙi · δqk (1.14)
dt ∂ q̇k ∂qk
i i,k
14
Substituting equations (1.8) and (1.16) into d’Alembert’s principle (1.6) then yields:
X d ∂T ∂T
− − Fk δqk = 0 (1.17)
dt ∂ q̇k ∂qk
k
Equation (1.17) holds for any system, but if the constraints are holonomic (so the gen-
eralized coordinates are independent), each coefficient of δqk must vanish individually,
yielding:
d ∂T ∂T
− = Fk , for all k (1.18)
dt ∂ q̇k ∂qk
These equations apply whether the forces Fnc,i can be derived from a potential or not. If
the forces derive from a scalar potential V (q1 , . . . , qN , t), we have Fnc,i = −∇ri V , where
the subscript indicates that the derivatives are taken with respect to the coordinates of
particle i. Then, equation (1.9) can be rewritten as:
X ∂ri ∂V
Fk = − ∇ri V · =− (1.19)
∂qk ∂qk
i
Since V depends only on the coordinates and possibly on time, but not on the velocities, it
does not involve the q̇k . Therefore, we can write ∂T /∂ q̇k = ∂ (T − V ) /∂ q̇k . Substituting
this expression, along with equation (1.19), into equation (1.18) yields the Euler-Lagrange
equations:
d ∂L ∂L
− = 0, for all k, (1.20)
dt ∂ q̇k ∂qk
which is simply Newton’s law of motion. This equivalence is expected, since d’Alembert’s
principle, and hence the Euler-Lagrange equations, originate directly from Newton’s laws.
15
The significance, however, lies in the generality of the Lagrangian formulation, which repro-
duces Newtonian mechanics in simple Cartesian cases but extends seamlessly to systems
with constraints or curvilinear coordinates, where Newton’s laws would be cumbersome to
apply. Furthermore, these equations arise from a principle far more general than Newton’s
law, which we explore in the next subsection.
6
For an in-depth discussion of this principle, see Richard P. Feynman, Robert B. Leighton, Matthew
Sands, 1963, The Feynman Lectures on Physics – Volume 2, Chap. 19 The Principle of Least Action,
Addison-Wesley.
16
another set of values for qi ) by following the trajectory that makes the action stationary
(typically a minimum). The action, denoted by S, is defined as:
ˆ tB
S= L [qi (t) , q̇i (t) , t] dt, (1.23)
tA
where L is the Lagrangian and tA and tB are the times when the system is at points A
and B, respectively. Hamilton’s principle applies to systems where non-constraint forces
are either conservative or derivable from a generalised potential that may depend on the
velocities.
We assume that the constraints are holonomic, so the generalised coordinates qi are
independent. If the path qi (t) makes S stationary, introducing an arbitrary infinitesimal
variation δqi (t) at all times between tA and tB , while keeping qi fixed at these endpoints,
will not change S. In other words, δS = 0.
It is important to note that these variations in qi (t) span the entire time interval between
tA and tB , and affect the entire trajectory of the system. This differs from the virtual
displacements discussed earlier, which were instantaneous changes at a fixed time. Here,
a variation δqi (t) in qi (t) induces a corresponding variation δ q̇i in q̇i (t). By definition,
the generalised velocity along the displaced trajectory qi + δqi is given by q̇i + δ q̇i =
d (qi + δqi ) /dt, which implies that δ q̇i = d (δqi ) /dt.
Note also that these variations are not constrained by the system’s dynamics. Since the
exact trajectory is unknown, we examine all nearby paths to determine which path makes
the action S stationary.
ˆ tB ˆ tB
δS = L (qi + δqi , q̇i + δ q̇i , t) dt − L (qi , q̇i , t) dt. (1.24)
tA tA
where the derivatives are evaluated at (qi , q̇i , t). Importantly, ∂L/∂qi accounts only for
the explicit dependence of L on qi , without considering the induced change in q̇i due to
variations in qi (and similarly for ∂L/∂ q̇i ). In other words, when we vary the action S,
we treat δqi and δ q̇i as independent variables, which is standard practice in variational
calculus. This approach views the trajectory as a mathematical construct characterised
by coordinates and velocities at each point, allowing us to vary these quantities indepen-
dently to explore other possible trajectories.
17
The next step is to acknowledge the relationship between δqi and δ q̇i by integrating by
parts the term in equation (1.25) involving δ q̇i :
ˆ ˆ ˆ tB
tB tB
∂L tB
∂L ∂L d (δqi ) d ∂L
δ q̇i dt = dt = δqi − δqi dt. (1.26)
tA ∂ q̇i tA ∂ q̇i dt ∂ q̇i tA tA dt ∂ q̇i
The boundary term in the square brackets is zero since δqi = 0 at both t = tA and tB .
Substituting this into equation (1.25), we obtain:
ˆ tB X ∂L
d ∂L
δS = − δqi (t)dt, (1.27)
tA ∂qi dt ∂ q̇i
i
where we now indicate explicitly the time dependence of δqi . Since the qi are independent
variables, the variations δqi are also independent. Therefore, for δS = 0 to hold for any
arbitrary variations δqi , each term within the integrand must vanish separately7 , resulting
in the Euler-Lagrange equations (1.20).
We have shown that Hamilton’s principle leads to the Euler-Lagrange equations, and it
is evident that the reverse is also true. These equations are quite general. In their standard
form, as given by equations (1.20), they are valid when the constraints are holonomic and
the non-constraint forces can be derived from a potential or generalised potential.
However, the applicability of the Euler-Lagrange equations extends beyond such idealised
cases. Even when the forces acting on a system are not derivable from a potential, such
as friction or other dissipative forces, the Euler-Lagrange equations in the form of equa-
tions (1.18) still hold. In these situations, the generalised forces Fk can be included
explicitly to account for non-conservative effects, allowing the equations to remain valid
without requiring a potential function.
Moreover, Hamilton’s principle, and consequently the Euler-Lagrange equations, can be
extended to handle certain types of nonholonomic constraints, though this is generally not
very practical.
Deriving the Euler-Lagrange equations from d’Alembert’s principle showed that, in
classical mechanics, the Lagrangian takes the form T − V . However, as shown in this
section, these equations can also be derived directly from Hamilton’s principle without
specifying a particular form of the Lagrangian. Hamilton’s principle is more fundamental
and provides a unified framework for deriving the equations of motion for a wide variety
of physical systems, encompassing both classical and quantum mechanics, as well as rela-
tivity. It also extends to fields in addition to particles.
7
This follows from the fundamental lemma of the calculus of variations, which states that if
ˆ b
f (x)η(x)dx = 0
a
for all arbitrary functions η(x) with continuous first and second derivatives on [a, b], then f (x) = 0 on
[a, b].
18
The Euler-Lagrange equations therefore remain valid in the context of rel-
ativity, provided the Lagrangian is appropriately constructed to satisfy the postulates of
relativity.
The Lagrangian defined by the principle of least action is not uniquely determined.
Indeed, if L is a function that renders the action stationary, then:
df
L̃ = L + , (1.28)
dt
where f is a function of qi , q̇i and t, will also make the action stationary. To see this,
define:
ˆ tB ˆ tB ˆ tB
df
S̃ = L̃ [qi (t), q̇i (t), t] dt = L [qi (t), q̇i (t), t] dt + dt = S + f (tB ) − f (tA ) .
tA tA tA dt
(1.29)
Variations δqi (t) that hold qi fixed at tA and tB yield δ S̃ = δS, since the additional
boundary term contributes nothing to the variation. Therefore, a path qi (t) for which
δS = 0 also satisfies δ S̃ = 0, implying that both L and L̃ are valid Lagrangians.
f = q (E + ṙ×B) . (1.30)
Denoting ϕ as the electric scalar potential and A as the magnetic vector potential, we
have:
∂A
E = −∇ϕ − and B = ∇×A. (1.31)
∂t
This leads to the expression for the force:
∂A
f = q −∇ϕ − + ṙ× (∇×A) . (1.32)
∂t
Using the vector identity ṙ× (∇×A) = ∇ (ṙ · A) − (ṙ · ∇) A, we observe that:
∂A ∂A dxi ∂A dA
+ (ṙ · ∇) A = + = .
∂t ∂t dt ∂xi dt
This is the total time derivative of A along the particle’s trajectory. Equation (1.32) can
then be rewritten as:
dA
f = q −∇ (ϕ − ṙ · A) − .
dt
Since the electromagnetic field, and therefore ϕ and A, are independent of the particle’s
velocity, the x-component of the force can be expressed as:
∂ d ∂
fx = q − (ϕ − ṙ · A) + (ϕ − ẋAx ) ,
∂x dt ∂ ẋ
19
where ∂ϕ/∂ ẋ = ∂Ax /∂ ẋ = 0, and similarly for the y and z-components. Using Cartesian
coordinates, with L = T − V and T = m ẋ2 + ẏ 2 + ż 2 /2, the Euler-Lagrange equa-
V = q (ϕ − ṙ · A) . (1.33)
Thus, the Lagrangian for a single charged particle in an electromagnetic field is given by:
1
L = mṙ2 − q (ϕ − ṙ · A) . (1.34)
2
Consider a one-dimensional crystal composed of N atoms, each with the same mass m,
arranged along the x-axis. At equilibrium, the i-th atom is at position xi , and we denote
its displacement from equilibrium by ϕi (t). We assume that each atom interacts only
with its nearest neighbours, so the total interaction potential, Utotal , is the sum of the
interaction potentials between successive atoms, which depend on the distance between
them:
N
X −1
Utotal = U (xi+1 + ϕi+1 − xi − ϕi ) .
i=1
We allow for an external potential V that depends on the displacement of the atoms.
The Lagrangian for this system is then given by:
N
X 1
L= mϕ̇2i − V (ϕi ) − Utotal .
2
i=1
Let a denote the equilibrium separation between atoms, so that xi = ia. For small
displacements, we expand the interaction potential in a Taylor series:
1
U (xi+1 + ϕi+1 − xi − ϕi ) = U (a + ϕi+1 − ϕi ) ≃ k (ϕi+1 − ϕi )2 ,
2
20
where we have used the fact that U has a minimum at a, and we set this minimum to zero.
Here, k is the second derivative of the potential evaluated at a, and is constant along the
chain. The Lagrangian then becomes:
N −1
NX
X 1 1
L= mϕ̇2i − V (ϕi ) − k (ϕi+1 − ϕi )2 . (1.35)
2 2
i=1 i=1
∂ϕ
ϕi+1 − ϕi = ϕ (xi + a) − ϕ (xi ) ≃ a (xi ) .
∂x
where the derivative is evaluated at the position of the i–th atom. This can be interpreted
as the potential energy of a string segment of length a, stretched by a small displacement
ϕ(xi ), suggesting that ka = T serves as the effective tension in the string.
In the continuous limit N → +∞, where a = dx, we replace m with ρa, where ρ is the
mass per unit length. The discrete sums in equation (1.35) then become integrals over the
length of the chain:
ˆ L
" 2 2 #
1 ∂ϕ 1 ∂ϕ
L= ρdx − T dx − dxV (ϕ) , (1.36)
0 2 ∂t 2 ∂x
where L = N a is the total length of the crystal, and V is the potential density defined by:
N
X ˆ L
V (ϕi ) = dxV (ϕ) .
i=1 0
This derivation can be naturally extended to three dimensions, leading to the Lagrangian:
ˆ
3 ∂ϕ
L= d x L ϕ, ∇ϕ, , (1.37)
∂t
where the integral is over the volume of the system and L is the Lagrangian density. It is
given by:
2
1 ∂ϕ 1
L= ρ − T (∇ϕ)2 − V (ϕ) , (1.38)
2 ∂t 2
with V now representing the potential density in a three-dimensional system.
21
Euler-Lagrange equations for fields
We now impose the condition that the action S is stationary, with S given by:
ˆ tB ˆ
dt d3 x L ϕ, ∇ϕ, ∂t ϕ ,
S= (1.39)
tA
where ∂t = ∂/∂t. Following a similar approach as in the particle case, but noting that the
Lagrangian now depends on ϕ, ∇ϕ and ∂t ϕ, we consider variations of the field ϕ. If ϕ (r, t)
is the field that makes S stationary, then any small variation δϕ, applied throughout the
volume and at all times between tA and tB , while keeping ϕ fixed at the boundaries and
at the times tA and tB , will leave S unchanged. This means δS = 0, with:
ˆ tB ˆ
3
δS = dt d x L ϕ + δϕ, ∇ϕ + δ (∇ϕ) , ∂t ϕ + δ (∂t ϕ) − L ϕ, ∇ϕ, ∂t ϕ . (1.40)
tA
where i = 1 . . . 3 and ∂1,2,3 ≡ ∂x,y,z . Here, δ(∂t ϕ) and δ(∂i ϕ) represent the changes in
∂t ϕ and ∂i ϕ resulting from the variations in ϕ. Using the relations δ(∂t ϕ) = ∂t (δϕ) and
δ(∂i ϕ) = ∂i (δϕ) and applying integration by parts, we obtain:
ˆ tB ˆ tB
∂L ∂L
dt δ(∂t ϕ) = −dt ∂t δϕ, (1.42)
tA ∂(∂t ϕ) tA ∂(∂t ϕ)
ˆ ˆ
∂L ∂L
dxi δ(∂i ϕ) = − dxi ∂i δϕ. (1.43)
∂(∂i ϕ) ∂(∂i ϕ)
The boundary terms vanish because δϕ is zero at times tA and tB as well as on the surface
enclosing the system.
Note on partial derivatives: The term ∂t (. . .) on the right-hand side of equation (1.42)
involves taking the derivative with respect to time while holding x, y and z constant, but
it accounts for the time dependence of the field and its derivatives in L via the chain
rule. Thus, it is not a total time derivative, since it does not account for explicit time
dependence in the coordinates, but it is also not strictly a partial derivative, as it includes
implicit time dependence through the field. If we were to ‘unpack’ ϕ and make L an explicit
function of x, y, z and t, instead of treating it as a function of the field and its derivatives,
then integration by parts above would yield the conventional partial time derivative. The
same reasoning applies to ∂x . Throughout this course, we will see many instances of such
‘partial’ derivatives, and the context will typically clarify which derivatives are being used.
22
Since δS = 0 must hold for any arbitrary variation δϕ, the expression in brackets within
the integrand must vanish, yielding the Euler-Lagrange equation for the field:
∂L ∂L ∂L
∂t + ∂i − = 0. (1.44)
∂(∂t ϕ) ∂(∂i ϕ) ∂ϕ
In this case, only one equation arises because there is a single field being varied. In
more general cases, the Lagrangian density may depend on multiple fields or on a vector
field with three spatial components, each varied independently, leading to a set of Euler-
Lagrange equations.
For the one-dimensional crystal discussed earlier, the Lagrangian density is given by
equation (1.38), which yields:
∂L ∂L
= ρ∂t ϕ, = −T ∂i ϕ.
∂(∂t ϕ) ∂(∂i ϕ)
∂2ϕ
ρ − T ∇2 ϕ = 0,
∂t2
which is a wave equation describing the propagation of a displacement ϕ within the crystal,
p
with a wave speed T /ρ.
Previously, starting from the Lagrangian for a single particle, we derived a Lagrangian for
a system of particles whose motion is described by a field ϕ (r, t), with the principle of
least action leading to a wave equation for this field. This approach can be generalised to
apply to fields in general, whether or not particles are present.
Consider, for example, the electromagnetic field (E, B) in the absence of charges and
currents. Our aim is to find a Lagrangian density such that applying the principle of least
action yields Maxwell’s equations. In Problem Set 1, we will verify that the Lagrangian
density is given by:
ϵ0 1 2
L = E2 − B . (1.45)
2 2µ0
By expressing the electric and magnetic fields in terms of the scalar potential ϕ and the
vector potential A, we can rewrite L as a function of ϕ, A and their derivatives. Applying
the principle of least action to this Lagrangian density results in Maxwell’s equations.
23
1.3 Conservation laws and Noether’s theorem
We now return to the topic of symmetries.
x-axis, then the potential V cannot depend on x. From equation (1.20), this implies that
∂L/∂ ẋ, which is equal to the x-component of the momentum px = mẋ, is constant.
Now, consider motion in a plane using polar coordinates q1 = r and q2 = θ. The
2 2 2
Lagrangian becomes L = m ṙ + r θ̇ /2 − V (r, θ). If the system is invariant under
rotation about the vertical axis through the origin, then V cannot depend on θ. From
equation (1.20), this implies that ∂L/∂ θ̇, which is equal to the z-component of the angular
momentum Jz = mr2 θ̇, is constant.
These two examples show that the Euler-Lagrange equations allow us to easily recover
results that were more laboriously derived using Newton’s law of motion in section 1.1.1.
More generally, if the Lagrangian of a system does not depend on the coordinate qk ,
then the system is invariant, or symmetric, under a transformation that changes qk . Such
a coordinate is referred to as a cyclic coordinate. In this case, equation (1.20) implies that:
∂L
pk ≡ , (1.46)
∂ q̇k
is a constant. We refer to pk as the conjugate momentum or canonical momentum, although
it does not necessarily have the units of momentum.
8
For a detailed discussion, see David Griffiths, Introduction to Electrodynamics, section 8.2
24
say, x, the conserved quantity is px = mẋ+qAx summed over the particles, not just the sum
of mẋ. The additional terms qAx represent how the charges contribute to the momentum
within the magnetic fields. For each particle, Newton’s second law still applies and gives
f = dpmec /dt, but the conserved quantity in the absence of external forces is not the sum
of the mechanical momenta because the magnetic fields carry additional momentum.
To restore the concept of action and reaction, we need to reinterpret it in terms of mo-
mentum exchange rather than direct forces or consider the forces as being mediated by
the field. When a charged particle moves, it generates a changing magnetic field, which
can exert forces on other charged particles. Conversely, changes in the motion of other
charged particles affect the magnetic field, which in turn affects the original particle. This
entire interaction is mediated by the field. For the system as a whole (particles + fields),
Newton’s third law holds. However, a particle does not directly exert a force on another
particle. Instead, particles transfer momentum to the fields, altering the field configura-
tion and momentum, and the fields transfer momentum back to the particles, exerting
forces on them. This process is more accurately described within the framework of special
relativity, which will be discussed later in these notes.
t → t′ = t + ϵτ (qk , t) , (1.47)
qi (t) → qi′ (t′ ) = qi (t) + ϵηi (qk , t) , (1.48)
where τ and ηi may depend on all the qk and time t. The dependence of t′ on the
coordinates qk is particularly important for applications in relativity. We will also need
the change in qi at fixed time, which we denote η̄i :
25
From equation (1.48), we have qi′ (t′ ) ≡ qi′ (t + ϵτ ) = qi (t) + ϵηi (qk , t). From equation (1.49),
we have qi′ (t + ϵτ ) = qi (t + ϵτ ) + ϵη̄i (qk , t + ϵτ ). Comparing these two expressions, we
obtain, to first order in ϵ:
η̄i (qk , t) = ηi (qk , t) − τ q̇i (t). (1.50)
For a time translation, ϵ represents the time interval by which t is shifted, with τ = 1
and ηi = 0 when the spatial coordinates remain unchanged. In this case, qi′ (t′ ) = qi (t) and
qi′ (t) = qi (t) − ϵq̇i , which can also be expressed as qi′ (t) = qi (t − ϵ). This means that a time
translation by ϵ replaces each coordinate qi (t) with qi (t − ϵ).
For rotations, ϵ corresponds to the angle of rotation θ, while for spatial translations, it
represents the distance by which the positions are shifted. In both these cases, τ = 0.
ˆ tA +δtA ˆ tB ˆ tB +δtB
S′ = − L qi′ (t), q̇i′ (t), t dt + L qi′ (t), q̇i′ (t), t dt + L qi′ (t), q̇i′ (t), t dt.
tA tA tB
(1.53)
Since δtA is an infinitesimal quantity, the first integral can be approximated by evaluating
the integrand at tA and multiplying by δtA . The same approximation applies to the third
integral. This yields:
ˆ tB
′ ′ ′ ′ ′
S − S = L qi (tB ), q̇i (tB ), tB δtB − L qi (tA ), q̇i (tA ), tA δtA + δL dt, (1.54)
tA
where:
X
∂L ∂L
δL = L qi′ (t), q̇i′ (t), t
− L qi (t), q̇i (t), t = ′
qi (t) − qi (t) ′
+ q̇i (t) − q̇i (t) .
∂qi ∂ q̇i
i
(1.55)
Here, the coordinates qi (t) undergoing the transformation define a trajectory that satisfies
the Euler-Lagrange equations (1.20). Using these and equation (1.49), we can rewrite δL
as:
X d ∂L ∂L
X
d ∂L
δL = ϵη̄i + ϵη̄˙ i = ϵη̄i , (1.56)
dt ∂ q̇i ∂ q̇i dt ∂ q̇i
i i
26
where the derivatives of L are evaluated at qi (t), q̇i (t), t .
For the first two terms on the right-hand side of equation (1.54), we can drop the primes
to first order in ϵ. This is because these expressions are multiplied by the infinitesimal
quantity δt = ϵτ , and the difference between the primed and unprimed quantities is also
first order in ϵ. We then combine these terms as follows:
tB
L qi′ (tB ), q̇i′ (tB ), tB δtB − L qi′ (tA ), q̇i′ (tA ), tA δtA = L qi (t), q̇i (t), t δt
t
ˆ
A
tB
d
= L qi (t), q̇i (t), t δt dt.
tA dt
ˆ tB
!
d X ∂L
S′ − S = ϵη̄i + L qi (t), q̇i (t), t δt dt. (1.57)
tA dt ∂ q̇i
i
For the condition S ′ − S = 0 to be satisfied for arbitrary times tA and tB , the integrand
must be zero. This implies that the function inside the time derivative must remain
constant over time:
X ∂L
η̄i + Lτ = const. (1.58)
∂ q̇i
i
!
X ∂L X ∂L
q̇i − L τ − ηi = const, (1.59)
∂ q̇i ∂ q̇i
i i
which is a form of Noether’s theorem. Since this theorem is derived without assuming
a specific form for the Lagrangian, it remains applicable in relativity, provided the
coordinate transformations preserve the action.
Noether’s theorem applies not only to mechanics but also to field theory, where the
symmetries are related to fields rather than just coordinates. This concept will be further
explored in Problem Set 1.
1.3.3 Applications
Translation in space
27
Now, consider N point masses with position vectors ri = (xi , yi , zi ). We define q3i−2 =
xi , q3i−1 = yi and q3i = zi for i = 1 . . . N . We assume that the system is invariant under
translation by any vector c. This corresponds to τ = 0, η3i−2 = cx , η3i−1 = cy and η3i = cz
for all i. Noether’s theorem then yields:
X ∂L X X X
ηi = cx mi ẋi + cy mi ẏi + cz mi żi = const.
∂ q̇i
i i i i
Since this must be satisfied for any values of the independent quantities cx , cy and cz , each
sum must be constant:
X X X
mi ẋi = Px , mi ẏi = Py , mi żi = Pz ,
i i i
where P = (Px , Py , Pz ) is the total momentum. Therefore, invariance under any transla-
tion implies that the total momentum is conserved.
Rotation
In polar coordinates, with q1 = r, q2 = θ, the Lagrangian is given by L = m ṙ2 + r2 θ̇2 /2−
V (r, θ). Consider a rotation around the vertical axis passing through the origin, which
corresponds to τ = 0, η1 = 0 and η2 = 1. If the potential is invariant under this rotation,
the action remains unchanged. Noether’s theorem then yields ∂L/∂ θ̇ = mr2 θ̇ = const,
confirming that the angular momentum about the z-axis is conserved.
Translation in time
For a time translation, τ = 1 and ηi = 0 for all i. If the system is invariant under time
translations, then Noether’s theorem yields:
X ∂L
H≡ q̇i − L = const. (1.60)
∂ q̇i
i
This quantity is known as the Hamiltonian. For a particle with a Lagrangian of the form
L = T − V , the Hamiltonian is given by H = T + V , representing the total energy.
Noether’s theorem then states that this energy is conserved when the system in invariant
under time translation.
This condition is equivalent to stating that the Lagrangian cannot explicitly depend
on time, meaning ∂L/∂t = 0, as can be shown by writing the rate of change of H:
dH X ∂L
d ∂L
dL
= q̈i + q̇i − . (1.61)
dt ∂ q̇i dt ∂ q̇i dt
i
28
Substituting this expression into equation (1.61) yields:
dH X d ∂L ∂L ∂L
= − q̇i − . (1.63)
dt dt ∂ q̇i ∂qi ∂t
i
The terms in the brackets vanish because L satisfies the Euler-Lagrange equations (1.20).
This yields:
dH ∂L
=− . (1.64)
dt ∂t
Therefore, if the Lagrangian has no explicit time dependence (∂L/∂t = 0), the Hamilto-
nian H is conserved.
tonian relate to the particle alone. The term qA in the canonical momentum represents
an exchange of momentum between the particle and the field, whereas the energy density
above pertains to the field alone. For a complete description of the total system, including
the energy of the electromagnetic fields, one must construct a Hamiltonian that includes
both the particle and the fields. This will be discussed in a later chapter.
29
30
Chapter 2
If two transformations preserve the invariance of a geometrical figure, then applying one
transformation followed by the other will also preserve the invariance of the figure. Simi-
larly, two transformations that leave the laws of physics unchanged can be performed
successively, resulting in another transformation that maintains the same invariance. Fur-
thermore, every transformation has an inverse that undoes its effect. More generally, a set
of transformations associated with symmetries forms a group, which is the focus of this
chapter1 .
2.1.1 Definition
(a) Closure: ∀ g1 , g2 ∈ G, g1 · g2 ∈ G,
1
For further reading, refer to the chapter Groups and Representations in Prof. Andre Lukas’ second-year
Mathematical Methods Lecture Notes
31
g2−1 · g1−1 · (g1 · g2 ) = e, and (g1 · g2 ) ·
Associativity implies that ∀ g1 , g2 ∈ G,
g2−1 · g1−1 = e, which means that g2−1 · g1−1 = (g1 · g2 )−1 .
2.1.2 Examples
Groups are of central importance in mathematics and are not always related to transfor-
mations. Here are a few examples:
• Group (Z, +) of all positive and negative integers under the operation of addition.
This is an infinite countable set, meaning that while the number of elements is
infinite, there exists a bijection between G and the set of natural numbers 1, 2, 3, . . ..
• Group (R∗ , ·) of non-zero real numbers under the operation of multiplication. This
is an infinite uncountable set.
In this chapter, we will focus on groups whose elements are transformations, which are
particularly significant in physics. Among these are the groups of translations T(n) and the
groups of rotations R(n) in n-dimensions. The internal operation for these groups is the
composition of transformations, meaning performing one transformation after another.
Known as the orthogonal group, it consists of all n×n orthogonal matrices. These matrices
M are real and satisfy M −1 = M ⊤ . We verify that they form a group:
(a) Closure: if M1 and M2 are orthogonal, then (M1 M2 )−1 = M2−1 M1−1 = M2⊤ M1⊤ =
(M1 M2 )⊤ , demonstrating that M1 M2 is also orthogonal.
32
(b) Associativity: matrix multiplication is associative in general, so it is associative for
orthogonal matrices in particular.
(c) Identity element: the identity matrix I is orthogonal and therefore an element of the
group.
Known as the special orthogonal group, it is the subgroup of O(n) which contains all n × n
orthogonal matrices with determinant equal to 1. Once again, we verify that they form a
group:
(c) Identity element: the identity matrix I is orthogonal and has a determinant of 1.
33
(d) Inverse element: if M is orthogonal, we have already shown that M −1 is also or-
thogonal. Additionally, det (M ) det M −1 = det M M −1 = det (I) = 1, implying
the determinant of M −1 .
In the case n = 2, SO(2) is the subgroup of O(2) for which ϵ = +1. As shown above,
the matrices in SO(2) represent rotations in Cartesian coordinates.
2.2 Representations
We mentioned earlier that the matrices in O(2) and SO(2) represent certain transforma-
tions in Cartesian coordinates. This leads us to the topic of representations.
This ensures that the representation preserves the group structure, meaning the matrices
M (g) defined in this way form a group under matrix multiplication. We can demonstrate
this as follows:
(a) Closure: if M (g) and M (g ′ ) are the matrices associated with the elements g ∈ G and
g ′ ∈ G, then the relation above shows that the product M (g) M (g ′ ) is associated
with the element g · g ′ ∈ G.
(c) Identity element: if e is the identity element of G, then equation (2.4) implies that
M (e) M (g) = M (e · g) = M (g). Similarly, M (g) M (e) = M (g · e) = M (g).
Therefore, M (e) is the identity element of the group formed by the matrices M (g).
Additionally, since these matrices are invertible, M (e) M (g) = M (g) implies that
M (e) = M (g) [M (g)]−1 = I. Thus, the matrix associated with e is the identity
matrix.
34
(d) Inverse element: equation (2.4) gives M (g) M g −1 = M g · g −1 = M (e) = I.
2.2.2 Examples
Group of real numbers under addition
Consider the group (R, +) of real numbers under the operation of addition. Any real
number a can be represented by the matrix:
!
1 0
D (a) = .
a 1
We verify that D(a)D(b) = D(a+b), ensuring that equation (2.4) is satisfied. The identity
element of (R, +) is 0 with D(0) = I. Therefore, the matrices D(a) are a representation
of the group (R, +). We will revisit these matrices when discussing the group of Galilean
transformations.
Group of rotations
We have previously shown that any matrix in SO(2) represents a rotation in two dimen-
sions. Conversely, the group R(2) of rotations in two dimensions can be represented by the
special orthogonal group SO(2), meaning each element of R(2) corresponds to a unique
matrix in SO(2). We denote these matrices as R (θ):
!
cos θ − sin θ
R (θ) = . (2.5)
sin θ cos θ
35
The angle α between two vectors r1 and r2 can be defined using the scalar product:
cos α = r⊤1 r2 / (∥r1 ∥∥r2 ∥). Since R conserves the norm, it will also conserve the an-
gle between the two vectors if it conserves the scalar product r⊤ ′⊤ ′
1 r2 . We have r1 r2 =
(Rr1 )⊤ Rr2 = r⊤ ⊤ ⊤ ′ ′
1 R Rr2 = r1 r2 . Therefore, the angle between r1 and r2 is the same as
the angle between r1 and r2 .
We have demonstrated that matrices in SO(2) leave the norm and scalar product
invariant, consistent with SO(2) being a representation of rotations in two dimensions.
Conversely, if we require matrices representing rotations to conserve the scalar product
(and therefore the norm), they must satisfy R⊤ R = I and thus belong to the O(2) group.
Moreover, as shown in section 2.1.3, only the matrices in O(2) with a determinant of 1
represent rotations, thus the representation of the group of rotations in two dimensions is
SO(2).
This shows that rotations can be defined either directly by the transformation R (θ),
or by specifying that they leave the norm and scalar product invariant, which implies
that their representation must be orthogonal matrices. To distinguish rotations from
reflections, we must additionally require that these matrices have a determinant of 1.
Consider a rotation by an infinitesimal angle around the origin. Without loss of generality,
this can be represented by a matrix R = I + X where X is first-order in the infinitesimal
angle. We define rotations as transformations that leave the scalar product (and therefore
the norm) invariant. As shown above, this condition implies R⊤ R = I. To first order in
the infinitesimal angle, this yields I + X ⊤ + X = I, and therefore X ⊤ = −X, meaning
36
that X is antisymmetric. This implies that X is proportional to the following matrix:
!
0 −1
J= , (2.6)
1 0
and since X is first-order in the infinitesimal angle, the factor, which we denote as δθ, is
proportional to this angle (we will show a posteriori that δθ is indeed the infinitesimal
angle). Therefore, to first order in δθ:
!
1 −δθ
R = I + δθJ = . (2.7)
δθ 1
!
x
Applying this transformation to a vector r = in Cartesian coordinates results in a
y
new vector r′ where x′ = x − δθy and y ′ = y + δθx. This recovers the geometric definition
of a rotation, confirming that δθ is indeed the rotation angle. From now on, we denote
the matrix as R (δθ).
Lie algebra is based on the idea that a rotation by an angle θ is equivalent to per-
forming N rotations by infinitesimal angles δθ = θ/N . In the limit as N → ∞, these two
approaches are equivalent. The matrix representing the rotation by an angle θ is then
given by2 :
N
N θ
R (θ) = lim [R (δθ)] = lim I+ J = eθJ , (2.8)
N →+∞ N →+∞ N
where the exponential is defined by:
+∞
X (θJ)k
eθJ = . (2.9)
k!
k=0
The matrix J is called the generator of the rotation group, and the set of generators for
a group is known as the Lie algebra. In this specific case, the generator is an element of
the group, though this is not generally true for all groups.
Since J 2 = −I, we have J 2n = (−1)n I and J 2n+1 = JJ 2n = (−1)n J. Therefore, we can
separate the contributions from the odd and even powers in the exponential series (2.9),
yielding:
+∞ +∞
X θ2k k
X θ2k+1
R (θ) = (−1) I + (−1)k J. (2.10)
(2k)! (2k + 1)!
k=0 k=0
| {z } | {z }
cos θ sin θ
Inserting the values of I and J, this gives:
!
cos θ − sin θ
R (θ) = . (2.11)
sin θ cos θ
2
We use the relation ex = limN →+∞ (1 + x/N )N . To prove this, we proceed as follows:
ln (1 + x/N ) ln (1 + ϵ)
lim (1 + x/N )N = lim exp x = exp x lim .
N →+∞ N →+∞ x/N ϵ→0 ϵ
Using L’Hôpital’s rule, limx→a (f (x)/g(x)) = limx→a (f ′ (x)/g ′ (x)), we obtain limN →+∞ (1 + x/N )N = ex .
37
Thus, using only the requirement that certain quantities remain invariant under the trans-
formation, we recover the matrix obtained earlier through geometric arguments. Although
we only imposed the condition that the matrix be orthogonal, without explicitly requiring
a determinant constraint, we still obtain matrices with a determinant of 1, which belong
to SO(2). This is because, by assuming that for infinitesimal transformations R was in-
finitesimally close to I, we excluded reflections and thus ruled out matrices that satisfy
the orthogonality condition but have a determinant of -1.
The Lie algebra of SO(2) is one-dimensional, meaning there is only one generator. This
is because only one parameter is needed to describe a rotation around a fixed point in a
plane. In general, the number of generators corresponds to the number of independent
parameters required to describe all the transformations within the group. We now examine
rotations around an arbitrary axis in three dimensions.
As before, we define a rotation in three dimensions as a transformation that preserves
the scalar product, implying it is represented by a matrix R that satisfies R⊤ R = I. For
an infinitesimal rotation by an angle δθ around a unit vector n̂, R can be written as
R(δθ, n̂) = I + X, with X ⊤ = −X. In two dimensions, any antisymmetric matrix is a
multiple of J. In three dimensions, X can be expressed as a linear combination of the
following three matrices:
0 0 0 0 0 1 0 −1 0
Jx = 0 0 −1 , Jy = 0 0 0 , Jz = 1 0 0 . (2.12)
0 1 0 −1 0 0 0 0 0
Therefore, X = δθx Jx + δθy Jy + δθz Jz , where the matrix I + δθx Jx corresponds to a ro-
tation by an angle δθx around the x-axis, and similarly for the y- and z-terms.
To relate δθx , δθy and δθz to δθ and n̂, we express the rotation matrix R(δθ, n̂) = I + X
as:
1 −δθz δθy
R(δθ, n̂) = δθz 1 −δθx . (2.13)
−δθy δθx 1
We then show that, for this matrix to represent a rotation by an angle δθ around the
unit vector n̂, its components must satisfy δθi = ni δθ for i = x, y or z. To see this, we
note first that n̂ remains invariant under the rotation, meaning Rn̂ = n̂. This implies
that the vector (δθx , δθy , δθz ) must be parallel to n̂, with a proportionality coefficient we
denote as ϵ. Next, examining the form of R, we observe that a vector u transforms as
Ru = u + ϵn̂×u. For a vector u perpendicular to n̂, the norm of Ru − u should be uδθ
38
for an infinitesimal rotation by an angle δθ. This requirement implies ϵ = δθ.
The matrices Jx , Jy and Jz serve as the generators of the SO(3) group and form its
Lie algebra. Since there are three independent parameters required to describe a rotation
in three dimensions, there are three corresponding generators. Note that the determinant
of each of these matrices is 0, so they do not belong to SO(3) themselves.
Following the reasoning from the previous section, we can decompose a three-dimen-
sional rotation by an angle θ around the vector n̂ into a sequence of infinitesimal rotations
by angles δθ = θ/N around n̂, with N → +∞. Each of these infinitesimal rotations is
associated with the matrix R(δθ, n̂), which can be written as:
θ θ θ
R(δθ, n̂) = I + δθnx Jx + δθny Jy + δθnz Jz = I + nx Jx + ny Jy + nz Jz .
N N N
To determine the coefficients of the matrix R (θ, n̂), we expand the exponential as
follows:
+∞
X (θn̂ · J)k
eθn̂·J = . (2.16)
k!
k=0
Using the expressions for the matrices Ji given in equations (2.12), we obtain:
0 −nz ny
n̂ · J = nz 0 −nx . (2.17)
−ny nx 0
Given that n2x + n2y + n2z = 1, it follows that (n̂ · J)3 = −n̂ · J (note that this does not imply
that the square of the matrix equals minus the identity, as the matrix is not invertible).
Therefore, for k ≥ 1, we have (n̂ · J)2k = (−1)k+1 (n̂ · J)2 and (n̂ · J)2k−1 = (−1)k−1 n̂ · J.
39
By separating the contributions from the odd and even powers in the exponential
series (2.16), we obtain:
+∞ +∞
X θ2k k 2
X θ2k−1
R (θ, n̂) = I − (−1) (n̂ · J) + (−1)k−1 n̂ · J, (2.18)
(2k)! (2k − 1)!
k=1 k=1
2.3.3 Commutators
The group of rotations is non-commutative, meaning that performing one rotation followed
by another is generally not the same as performing the second rotation followed by the
first. Consider two infinitesimal transformations M = I + X and M ′ = I + X ′ (not
necessarily rotations). These transformations commute if M M ′ = M ′ M . To first order in
X and X ′ , this implies XX ′ = X ′ X. We define the commutator as the matrix:
[X, X ′ ] = XX ′ − X ′ X. (2.20)
eA eB = eC , (2.21)
where
1 1
C = A + B + [A, B] + ([A, [A, B]] − [B, [A, B]]) + higher order terms.
2 12
This formula is not limited to Lie algebra generators, but it requires the series expansion
involving commutators to converge, which may not always happen for matrices other than
generators.
40
2.4 Galilean and Lorentz transformations in mechanics
Consider three observers moving with respect to each other. The laws of physics as seen
by observer j are a transformation T (i → j) of the laws seen by observer i. Transforming
from observer 1 to observer 2, and then from observer 2 to observer 3, must yield the same
result as transforming directly from observer 1 to observer 3. In other words:
T (2 → 3)T (1 → 2) = T (1 → 3).
This indicates that the transformations T form a group. Consequently, we can generate a
representation of this group using Lie algebra to derive the transformations.
Consider a frame (S) with origin O where a point mass is located at position x on the
x-axis at time t. Now, consider another frame (S′ ) that moves with velocity v towards
decreasing x, with its origin O′ coinciding with O at t = 0. We synchronise the time
origin in (S′ ) with that in (S). This type of transformation, where only the velocity of the
reference frame changes, is called a boost. A Galilean transformation is characterised
! by
t
t′ = t and x′ = x + vt. If we define a spacetime position vector in (S) as r = , then the
x
Galilean transformation from frame (S) to frame (S′ ) can be represented by the matrix:
!
1 0
D (v) = , (2.22)
v 1
which acts on the vector r, since D (v) r = r′ . The property D(v)D(v ′ ) = D(v +v ′ ) reflects
the fact that transforming to a frame moving at velocity v, and then to another frame
moving at velocity v ′ relative to (S′ ), is equivalent to moving directly to a frame moving
at velocity v + v ′ relative to (S). This is the well-known composition of the velocities.
If we consider a transformation to a frame moving with an infinitesimal velocity δv,
then D = I + δvX, where: !
0 0
X= , (2.23)
1 0
is the generator of the group of Galilean transformations in one dimension. Since X 2 = 0:
+∞
X (vX)k
D(v) = evX = = I + vX,
k!
k=0
as expected.
41
Boost and translations in one dimension
Now, consider the case where, in addition to a boost, there is a translation in space and
time of the coordinate system. Specifically, let frame (S′ ) have its origin O′ at x = x0
at t = 0, with the origin of time in frame (S′ ) corresponding to t = t0 in frame (S). The
Galilean transformation then results in x′ = x + vt + x0 and t′ = t − t0 .
We observe that while the boost results only in a homogeneous term, the translations
introduce inhomogeneous
terms. To manage these, we extend the usual spacetime position
1
vector to s = t . This extension allows us to incorporate translations into linear
x
transformations via matrix operations. Specifically, the Galilean transformation from
frame (S) to frame (S′ ) can be represented by the following matrix:
1 0 0
M (v, t0 , x0 ) = −t0 1 0 , (2.24)
x0 v 1
which acts on the vector s. Transforming from frame (S) to frame (S′ ) and then to frame
(S′′) transforms s into s′′ such that:
1 1 0 0 1 0 0 1 1 0 0 1
′′ ′ ′
t = −t0 1 0 −t0 1 0 t = −t0 − t0 1 0 t ,
x ′′ ′
x0 v 1 ′ x0 v 1 x x0 + x0 − v t 0 v + v ′
′ ′ 1 x
where v ′ , x′0 and t′0 are related to the transformation from frame (S′ ) to frame (S′′). The
3 × 3 matrix on the right-hand side represents the transformation from frame (S) to frame
(S′′), characterised by the following parameters:
It is clear that the group is not commutative, meaning the order in which the transforma-
tions are performed matters.
We can construct the Lie algebra by writing the matrices associated with each of the
infinitesimal transformations3 :
with:
0 0 0 0 0 0 0 0 0
XK = 0 0 0 , XH = −1 0 0 , XP = 0 0 0 . (2.25)
0 1 0 0 0 0 1 0 0
3
The subscript ‘K’ for the generator associated with a boost signifies kinematic transformations, as
boosts are related to changes in velocity. The subscript ‘P’ is used for the generators of spatial translations
because they are generated by the momentum operator. Finally, the subscript ‘H’ refers to the Hamiltonian,
which is the generator of time evolution.
42
For an arbitrary infinitesimal transformation, we have:
Since XK2 = X 2 = X 2 = 0 and the product of any two of these matrices is zero, except
H P
for XK XH = −XP , it can be shown that:
So far, we have only considered a single spatial dimension. However, in three dimensions,
the coordinate axes of (S′ ) could be rotated relative to those of (S), the velocity of (S′ )
relative to (S) could have any direction, and the origin of (S′ ) could be shifted by a vector
(x0 , y0 , z0 ). Consequently, the spacetime position vector and the transformation matrix
take the following form:
1 1 0 0 0 0
t −t0 1 0 0 0
s = x , M = x0 vx
,
(2.27)
y y0 v y R
z z0 v z
where R is the rotation matrix. There are now 10 parameters describing a transformation:
3 for rotations, 3 for boosts, 3 for spatial translations and 1 for the time translation.
Consequently, there are 10 generators in the Lie algebra, each represented by a 5 × 5
matrix. We can derive these generators by considering an infinitesimal transformation
associated with each parameter in turn.
A translation of time by an infinitesimal amount δt0 results in a matrix transformation
M = I + δt0 XH , where XH is a 5-dimensional extension of the 3 × 3 matrix given in the
previous section:
0 0 0 0 0
−1 0 0 0 0
XH = 0 0
.
(2.28)
0 0 0
0 0
43
Similarly, spatial translations are associated with the generators XP x , XP y and XP z , while
boosts are associated with the generators XKx , XKy and XKz . All of these generators
can be derived in the same manner as XP and XK were calculated above. For example:
0 0 0 0 0 0 0 0 0 0
0 0 0 0 0 0 0 0 0 0
XP x = 1 0
, XKx = 0 1
.
(2.29)
0 0 0 0 0 0
0 0 0 0
For the rotations, the generators XJi , with i = x, y, z, are those calculated in section 2.3.2
for the SO(3) group, but extended to 5 × 5 matrices:
0 0 0 0 0
0 0 0 0 0
XJi = 0 0
.
(2.30)
0 0 Ji
0 0
General transformation matrices can be derived from these 10 generators using the same
method applied in the simpler cases discussed above.
We can gain insights into the nature of the Galilean transformations by examining
their commutation relations. If two generators commute, it means that the associated
transformations do not interfere with each other. Conversely, if they do not commute,
their effects are interdependent.
• [XJi , XJj ] = ±XJk if i ̸= j: this indicates that performing rotations around the i
and j-axes results in a net effect equivalent to a rotation around the k-axis,
• [XJi , XKj ] = ±XKk if i ̸= j: rotations mix boost directions, reflecting the fact that
boosts transform as vectors under rotations,
• [XKi , XKj ] = 0: successive boosts in different directions do not interfere with each
other, consistent with Galilean relativity,
• [XKi , XH ] = −XPi : boosts and time translations are related to spatial translations,
as seen from the transformation matrix (2.24), where vt produces a spatial transla-
tion.
44
The fact that these commutators all result in elements of the Galilean group confirms that
the transformations indeed form a group: the product of any two elements (whether they
are rotations, boosts, translations, or a combination) is another element of the Galilean
group, thereby satisfying the closure property.
larly, in (S′ ), where the position of the wavefront advances by ∆r′ during the time interval
∆t′ , we have ∆s’2 = −c2 ∆t′2 + ∆x′2 + ∆y ′2 + ∆z ′2 = 0. Therefore, the Lorentz transfor-
mation from (t, x, y, z) to (t′ , x′ , y ′ , z ′ ) leaves ∆s2 invariant for electromagnetic waves in
vacuum. This result is actually much more general, as will be discussed in chapter 3, and
any ∆s2 is invariant under a Lorentz transformation.
We note :
ct −1 0 0 0
x 0 1 0 0
s=
y,
and g =
0
, (2.31)
0 1 0
z 0 0 0 1
where the spacetime position vector s has ct as its first coordinate instead of just t, for
unit consistency, and where g is known as the Minkowski metric. The space defined by the
45
vectors s is called Minkowski space. The quantity that remains invariant under the Lorentz
transformation, ∆s2 , can be expressed as the product ∆s⊤ g∆s. Letting L represent the
Lorentz transformation matrix, we have ∆s′ = L∆s. Thus, ∆s′⊤ g∆s′ = (L∆s)⊤ gL∆s =
∆s⊤ L⊤ gL∆s. Invariance requires this expression to be equal to ∆s⊤ g∆s. Since this
condition must hold for any ∆s, it implies the following requirement:
L⊤ gL = g. (2.32)
Just as we derived the matrices R of SO(3) by solving R⊤ R = I, we now seek the matrices
L that satisfy equation L⊤ gL = g. Taking the determinant on both sides of this equation
yields det(L)2 = 1, so det(L) = ±1. As noted in section 2.3.1, the Lie algebra, which
generates L from infinitesimal transformations, only produce matrices with a determi-
nant of 1. These transformations are smoothly connected to the identity transformation.
Matrices with a determinant of −1 represent transformations that are not smoothly con-
nected to the identity and cannot be generated by the Lie algebra. Therefore, infinitesimal
transformations yield only matrices in SO(1, 3), not those in O(1, 3). The latter can be
obtained by multiplying the matrices in SO(1, 3) by g or −g, similar to how the matrices
of O(2) are derived from those of SO(2) by multiplying them by the matrix P , as defined
in equation (2.3).
We now verify that the set of matrices satisfying equation (2.32) with a determinant
of 1 forms a group:
(a) Closure: consider two matrices L1 and L2 that satisfy equation (2.32). Using the
associativity of matrix multiplication, we find (L1 L2 )⊤ g (L1 L2 ) = L⊤ ⊤
2 L1 gL1 L2 =
L⊤
2 gL2 = g, showing that L1 L2 also satisfies equation (2.32). Additionally, if
det (L1 ) = det (L2 ) = 1, then det (L1 L2 ) = det (L1 ) det (L2 ) = 1.
46
(b) Associativity: matrix multiplication is associative in general, and thus it remains
associative for matrices that satisfy equation (2.32) in particular.
(c) Identity element: the identity matrix I satisfies I ⊤ gI = g and its determinant is 1.
−1
(d) Inverse element: if L satisfies L⊤ gL = g, multiplying both sides by L⊤ on
−1 ⊤
−1 −1 −1
⊤ −1
the left and by L on the right gives g = L gL = L gL , showing
−1 −1
that L also satisfies the condition. Additionally, since det L = 1/det (L), if
−1
det (L) = 1, then det L = 1.
Generators
The generators of these transformations, denoted as XJi for i = x, y, z, are those derived
in section 2.3.2 for the SO(3) group, but extended to 4 × 4 matrices:
0 0 0 0
0
XJi = . (2.34)
0
Ji
0
47
so Kx must take the form:
a b 0 0
d e 0 0
Kx = . (2.35)
0 0 0 0
0 0 0 0
Similarly, the generators for boosts along the y and z-axes are:
0 0 1 0 0 0 0 1
0 0 0 0 , and XKz = 0 0 0 0
XKy = 1
. (2.37)
0 0 0 0 0
0 0
0 0 0 0 1 0 0 0
General transformations
Having established all the generators of the Lorentz group, we can now express any trans-
formation matrix L, incorporating both rotations and boosts, as follows:
L = eθx Jx +θy Jy +θz Jz −ζx XKx −ζy XKy −ζz XKz . (2.38)
Here (ζx , ζy , ζz ) is known as the boost vector. This means that a boost by ζx in the x-
direction is achieved through a sequence of infinitesimal boosts ϵx , and similarly for the y
and z-directions. The choice of the minus sign will be explained later.
In the case of the Galilean transformations, knowing the transformation in advance
allowed us to directly equate the boost vector with the velocity vector. However, Galilean
transformations can result in velocities that exceed the speed of light, as velocities simply
add when transitioning through successive moving frames. Consequently, for Lorentz
transformations, the boost vector cannot be directly equated with the velocity vector.
The relationship between the boost vector and velocity will be derived below.
Commutation relations
As is evident from the form of the generators, rotations only affect the spatial coordinates,
while boosts mix spatial and time coordinates. The rotations alone form the SO(3) group.
However, boosts alone do not constitute a group. This is illustrated by the commutator
[XKx , XKy ] = −XJz , which shows that performing successive boosts in different directions
results in a rotation.
More generally, the commutation relations are as follows:
48
• [XJi , XJj ] = ±XJk if i ̸= j, as seen in the Galilean group,
Similar to the Galilean group, the fact that these commutators yield elements within
the Lorentz group confirms that the transformations collectively form a group: the com-
bination of any two elements (whether rotations, boosts, or their combinations) results in
another element of the Lorentz group, thus satisfying the closure property.
Boosts
Now, assuming there is no rotation and only a finite boost ζx along the x-direction, we
denote the transformation matrix as Λx instead of L. Following the same method used in
section 2.3.1 for the SO(2) group, we obtain:
+∞
−ζx XKx
X (−ζx XKx )k
Λx = e =I+ , (2.39)
k!
k=1
we find that, for k ≥ 1, (XKx )2k = I2×2 and (XKx )2k−1 = XKx .
Therefore, we can separate the contributions from the odd and even powers in the
exponential series (2.39), which yields:
+∞ +∞
X (−ζx )2k X (−ζx )2k+1
Λx = I + I2×2 + XKx . (2.41)
(2k)! (2k + 1)!
k=1 k=0
| {z } | {z }
cosh (−ζx ) − 1 sinh (−ζx )
Substituting the values of I2×2 and XKx , we get:
cosh ζx − sinh ζx 0 0
− sinh ζx cosh ζx 0 0
Λx = . (2.42)
0 0 1 0
0 0 0 1
49
The same calculation for a finite boost in the y- or z-directions yields:
cosh ζy 0 − sinh ζy 0 cosh ζz 0 0 − sinh ζz
0 1 0 0 0 1 0 0
Λy = , Λz = . (2.43)
− sinh ζy 0 cosh ζy 0
0 0 1 0
0 0 0 1 − sinh ζz 0 0 cosh ζz
A more general matrix Λ can be obtained for a boost in any direction by calculating
Λ = e−ζx XKx −ζy XKy −ζz XKz .
The final step in obtaining the Lorentz transformation is to relate the boost to the
velocity. Since the transformation involves only a boost and no translation, the origin O′
of the coordinate system (S′ ) coincides with the origin O of the coordinate system (S) at
time t = 0. When considering a boost in the x-direction, this means that the trajectory
of O′ is x = vx t in frame (S), while it is x′ = 0 in frame (S′ ). Now, using s′ = Λx s, we
obtain x′ = (cosh ζx ) x − (sinh ζx ) ct, which, as stated, must be 0 when x = vx t. This
yields tanh ζx = vx /c. It now becomes clear why we chose the boost to have a minus sign
in equation (2.38): this way, the boost has the same sign as the velocity component. From
here on, we denote βx = vx /c. Therefore:
1 1 βx
cosh ζx = p =p , and sinh ζx = cosh ζx tanh ζx = p .
2 1 − βx2 1 − βx2
1 − tanh ζx
Dropping the subscript x, we can thus rewrite the transformation matrix as:
γ −βγ 0 0
−βγ γ 0 0
, with γ = p 1 v
Λ= , and β = . (2.44)
0 0 1 0
1 − β2 c
0 0 0 1
50
Chapter 3
In the previous chapter, we derived the Lorentz transformation using the conservation of
the spacetime interval and group theory formalism. This allowed us to obtain a compre-
hensive form of the transformation, including rotations and boosts in any direction. In
this chapter, we will start with a brief review of a simple derivation of the transforma-
tion for a boost along one axis, as this emphasises the fundamental principles underlying
the Lorentz transformation. Following this, we will discuss some core aspects of special
relativity before introducing the important concept of the 4-vector.
Events are described within a reference frame (or frame, for short), which is a coor-
dinate system that incorporates both space and time. A specific frame can be visualised
as a collection of metre sticks to measure distances and a network of clocks to track time
at every point in space. These clocks must be synchronised, meaning they all display the
same time at a given moment and advance at the same rate when observed from within
the frame. A particularly important type of frame is the inertial reference frame, where a
particle not subject to any force moves at a constant speed in a straight line (or remains
at rest if initially stationary).
The laws of physics are the same in all inertial reference frames.
This principle was recognised early on by Galileo and Newton. For Newton, time was
universal, the same for all observers in all frames. This implied that a light ray moving
with velocity c in one frame would be perceived as moving with velocity c + v by an ob-
server travelling at velocity v in the opposite direction of the ray within the same frame.
However, by the end of the 19th century, precise measurements showed that the speed of
light remained constant for observers in different inertial frames.
The speed of light in a vacuum is the same in all inertial reference frames.
51
This implies that time cannot be universal, and that if the clocks in two different frames
are synchronised at one moment, they will not remain synchronised if the frames move
relative to each other. Space and time coordinates are inseparable, and their combination
is referred to as spacetime, or Minkowski space. A point with coordinate (t, x, y, z) is
called an event (not an actual occurrence, but simply a point in spacetime, representing
something that could happen). A path traced in spacetime is called a worldline.
where γ is a function of v. If (S′ ) were moving with velocity −v (towards decreasing x), we
would have x′ = (x + vt) γ(−v). In the first case, the point O at some time t0 measured
in (S) is viewed in (S′ ) as having coordinate x′ = −vt0 γ(v), whereas in the second case,
it is viewed as having coordinate x′ = vt0 γ(−v). Apart from the change of sign of x′ , the
two situations are identical, so that we must have γ(−v) = γ(v).
Returning to the case where (S′ ) moves with velocity v, this is equivalent to (S) moving
with velocity −v relative to (S′ ). Thus, the transformation from (t′ , x′ ) to x is the same
as given by equation (3.1) but with the sign of v reversed. This yields:
x = x′ + vt′ γ(v).
(3.2)
52
Equations (3.1) and (3.4), with γ given by equation (3.3), together constitute the Lorentz
transformation. Since the frames move relative to each other along the x-axis, y ′ = y and
z ′ = z.
By defining β = v/c, and using the coordinate ct instead of t, we can express the
Lorentz transformation in a compact form:
x′ = γ (x − βct) , (3.6)
y ′ = y, (3.7)
z ′ = z, (3.8)
Since frame (S) moves with velocity −v relative to frame (S′ ), we can also express the
transformation from (S′ ) to (S) by reversing the sign of the velocity. This results in:
ct = γ ct′ + βx′ ,
(3.9)
x = γ x′ + βct′ ,
(3.10)
′
y=y, (3.11)
′
z=z, (3.12)
where β = v/c as defined above. This implies that Λ−1 is obtained from Λ by replacing β
with −β wherever it appears.
To conveniently represent events and worldlines in a frame (S), we use a diagram where
the horizontal axis represents x (considering only one spatial coordinate) and the vertical
axis represents time, typically as ct. Each event is indicated by a coordinate pair (ct, x),
specifying its position in spacetime. An object’s worldline at rest in frame (S) is depicted
as a vertical line.
53
In this diagram, the worldline of a
light ray, x = ct, is shown as a line
inclined at 45 degrees to both the x-
and ct-axes (red line in the diagram).
Next, we aim to plot the x′ and ct′
axes of frame (S′ ). The equation of
the x′ -axis in (S′ ) is ct′ = 0. In frame
(S), according to equation (3.5), this
becomes ct = βx. Similarly, the equa-
tion for the ct′ -axis, which is x′ = 0 in
(S′ ), becomes x = βct in (S).
Therefore, in the x-ct diagram, the x′ and ct′ axes are symmetrical with respect to the
x = ct line. They are represented as the blue lines in the diagram, where ϕ is defined as
ϕ = tan−1 β.
Proper length refers to the length an object has in a frame (S) in which it is at rest.
Suppose this object is a rod lying along the x-axis, between the coordinates x1 and x2
(with x2 > x1 ). By definition, L0 = x2 − x1 is the proper length. Now, consider a frame
(S′ ) moving with velocity v along the x-axis. The length of the rod in (S′ ) is defined as the
distance L between its endpoints when their positions are measured simultaneously in (S′ ).
Let t′ be the time at which the measurement is made in (S′ ), with the endpoints located
at x′1 and x′2 such that x′2 − x′1 = L. Using equation (3.10), we get x2 = γ (x′2 + βct′ ) and
x1 = γ (x′1 + βct′ ). Substracting these two equations yields:
L = L0 /γ. (3.13)
Thus, the length in a frame moving relative to the rest frame of the object along its length
is always shorter than the proper length. Note that if frame (S′ ) were moving along the y
or z-axes, the length of the object in (S′ ) would be the same as that in (S).
54
To understand the significance of this
result, let us represent the endpoints
of the rod as measured in both frames
on the diagram. In (S), assume the
rod lies between points O and A at
time t = 0. Now, assume we make the
measurement in (S′ ) at t′ = 0. Ac-
cording to equations (3.5) and (3.6),
the endpoint which is at O in (S) at
t = 0 is at x′ = 0 in (S′ ) at t′ = 0.
The other endpoint must lie along the x′ axis since we are considering events that are
simultaneous in (S′ ) at t′ = 0. The position of this point, denoted as A′ , is given by
equation (3.10) as x′2 = x2 /γ. Therefore, in frame (S), the proper length of the rod is
L0 = OA ≡ x2 , while in frame (S′ ) the length of the rod is L = OA′ ≡ x′2 . The lengths
are measured between different points, which is why they differ.
p
We can position A′ precisely on the x′ -axis by noting that X = x′2 cos ϕ = x′2 / 1 + tan2 ϕ =
p p p p
x′2 / 1 + β 2 . With x′2 = x2 /γ = x2 1 − β 2 , this yields X = x2 1 − β 2 / 1 + β 2 .
Proper time is the time interval measured by a clock that is stationary relative to the
events being timed. In other words, it is the time elapsed between two events in the frame
where both events occur at the same location. Instead of referring to events, we can also
consider a particle moving along a worldline in a spacetime diagram, travelling from one
event to another. If we move with the particle along its trajectory, our clock measures
the particle’s proper time. This can also be described by saying that, in the rest frame
of the particle, where it moves vertically (indicating a constant spatial location), the time
measured is the proper time. We denote the proper time as τ .
Suppose two events with the same coordinate x in (S) occur at times t1 and t2 > t1 . By
definition, t2 − t1 = τ is the proper time. In frame (S′ ), these events occur at times t′1 and
t′2 , with the time interval between them being ∆t′ = t′2 − t′1 . According to equation (3.5),
we have t′1 = γ (t1 − βx/c) and t′2 = γ (t2 − βx/c). Substracting these equations gives:
Thus, the time interval in a frame moving relative to the rest frame of the events is always
longer than the proper time.
55
3.4 The invariance of the spacetime interval
In section 2.4.2 of chapter 2, we began by asserting the invariance of the spacetime interval:
where ∆x, ∆y, ∆z and ∆t describe the separations in space and time between two events in
the four-dimensional spacetime. We then derived the Lorentz transformation directly from
this invariant using Lie algebra. Here, we have derived the Lorentz transformation using
the two postulates of relativity. It is straightforward to verify that this transformation,
given by equations (3.5)-(3.8), ensures −c2 ∆t′2 + ∆x′2 + ∆y ′2 + ∆z ′2 = −c2 ∆t2 + ∆x2 +
∆y 2 + ∆z 2 , consistent with the spacetime interval being invariant.
Timelike worldlines
1/2
For massive particles travelling at speeds v < c, dr ≡ dx2 + dy 2 + dz 2 = vdt and
2
therefore ds < 0. This is known as a timelike interval. In the rest frame of the particle,
where the time is the proper time:
56
objects. This can also be understood by taking the limit γ → +∞ in equations (3.13)
and (3.14), which show that, in a frame moving at a velocity close to c, lengths contract
to nearly zero while time intervals stretch to infinity. Thus, a photon does not experience
either space or time.
Since proper time is the time measured by a clock moving with the object, and since
dτ = 0 for photons, it follows that photons do not have a rest frame. Unlike massive
particles, there is no inertial frame in which a photon is at rest.
Spacelike worldlines
For hypothetical particles travelling faster than light (known as tachyons, although they
are believed not to exist), we have ds2 > 0. This is known as a spacelike interval. In a
one-dimensional spacetime diagram, such a worldline would lie below that of a light ray.
Note that there is no frame in which two events separated by a spacelike interval occur
at the same location. If such a frame existed, ds2 would be equal to −dτ 2 , which is not
possible since ds2 > 0.
Consider an event (ct, x, y, z) in the spacetime of an inertial frame. Can this event influence
another event with coordinates [c (t + ∆t) , x + ∆x, y + ∆y, z + ∆z]? In other words, can
a signal sent from (x, y, z) at time t reach the location (x + ∆x, y + ∆y, z + ∆z) before
1/2
time t + ∆t? This is only possible if ∆x2 + ∆y 2 + ∆z 2 ≤ c∆t, which means ∆s2 ≤ 0.
Therefore, there cannot be any causal relationship between events separated by spacelike
intervals.
This diagram illustrates the concepts
discussed above. The origin O repre-
sents an event in the present (t = 0).
It can influence any event within the
so-called future light cone and can
receive information from any event
within the past light cone. Communi-
cation between events located on the
boundaries of the cones must occur
at the speed of light. Conversely, the
event at O cannot influence or be in-
fluenced by events outside of the light
cones.
Similar cones can be drawn from any event in spacetime, not just from O.
57
3.5 Relativistic velocity
3.5.1 Velocity composition
cosh ζ ′ − sinh ζ ′ 0 0 cosh ζ − sinh ζ 0 0
− sinh ζ ′ cosh ζ ′
′′
0 0 − sinh ζ cosh ζ 0 0
Λ = . (3.17)
0 0 1 0
0 0 1 0
0 0 0 1 0 0 0 1
with ζ ′′ = ζ + ζ ′ . Hence, the transformation from (S) to (S′′) is a boost with velocity
v ′′ along the x-axis such that v ′′/c = tanh ζ ′′ = tanh (ζ + ζ ′ ) . This can be expressed as
v ′′/c = (tanh ζ + tanh ζ ′ ) / (1 + tanh ζ tanh ζ ′ ), leading to the relativistic velocity compo-
sition formula:
v ′′ v/c + v ′ /c
= . (3.19)
c 1 + vv ′ /c2
While this formula can also be derived using the transformation matrix given by equa-
tion (2.44), the calculation is more involved.
We now examine how the velocity components of a particle change when observed from
different inertial frames. Let u be the velocity of the particle in frame (S), and u′ the
58
velocity in frame (S′ ), which moves with velocity v along the x-axis relative to (S). By
definition:
dx dx′
ux = , u′x = ′ ,
dt dt
and similarly for the y and z-components.
From equations (3.5)-(3.8), we have:
Dividing these last three equations by the first and using β = v/c, we get the relativistic
velocity transformation1 :
ux − v uy uz
u′x = , u′y = , u′z = . (3.20)
1 − ux v/c2 γ (1 − ux v/c2 ) γ (1 − ux v/c2 )
3.6 Four-vectors
In this chapter, we have so far focused on Lorentz transformations that are specifically
boosts. The following discussion, however, applies to any general Lorentz transformation.
To maintain consistency and simplicity, we will however continue to denote the transfor-
mation matrix as Λ, avoiding the introduction of new notation.
Since time and space are intrinsically linked through Minskowski spacetime, it is natural
to define a vector where time and the three spatial coordinates are treated equally. This
is called a 4-vector, and it is represented as a column vector. For example, the position
4-vector is expressed as:
x0
1
x
Xµ =
x2 (3.21)
x3
1
To make the first equation easier to remember, it can be expressed as:
vP/S + vS/S′
vP/S′ = ,
1 + vP/S vS/S′ /c2
where vP/S = ux and vP/S′ = u′x are the velocities of particle P in frame (S) and (S′ ), respectively, and
vS/S′ = −v is the velocity of frame (S) relative to frame (S′ ).
59
Here, x0 = ct represents the time component (scaled by c for unit consistency), while
x1 = x, x2 = y and x3 = z are the spatial components. The convention is to use Greek
letters as superscripts to denote the components of 4-vectors, with µ ranging from 0 to
3, and integers i = 1, 2, 3 to denote the spatial components of the standard Euclidean
3-vectors. We will also use capital letters to denote 4-vectors. Thus, X µ refers to the
4-vector as a whole, with no specific value assigned to µ, while xµ refers to the individual
components of the 4-vector, with µ taking values from 0 to 3.
By definition, the position 4-vector transforms according to the Lorentz transforma-
tion. This means that if X µ , as defined above, is a position 4-vector in frame (S), then
in frame (S′ ), which moves with velocity v along the x-axis relative to (S), the position
4-vector is X ′µ = ΛX µ , where Λ is the transformation matrix given by equation (2.44).
As we will see later in this course, various physical quantities transform between frames
in the same way as the spacetime position vector. Consequently, these quantities can be
combined into a 4-vector that, like the position 4-vector, follows the Lorentz transformation
rules. In fact, this defines a 4-vector: any set of spacetime components that transform
according to the Lorentz transformation. Formally, this means that if Aµ is a 4-vector in
frame (S), then in frame (S′ ) it transforms as:
This implies that if Aµ in frame (S) represents some physical quantity, then its Lorentz-
transformed 4-vector A′µ represents the same physical quantity in frame (S′ ).
Previously we noted that, under a Lorentz transformation, the invariant quantity is not
the commonly defined squared length of the spacetime interval, c2 ∆t2 + ∆x2 + ∆y 2 + ∆z 2 ,
but rather ∆s2 = −c2 ∆t2 + ∆x2 + ∆y 2 + ∆z 2 . In fact, ∆s2 is the square of the length
of the spacetime interval in hyperbolic space. This concept is formally expressed using a
metric, which is represented by a matrix.
A metric is a mathematical construct that defines the distance between points in a
given space. It is represented by a matrix that encodes information about the shape,
angles, and volumes within the space:
• In a flat (Euclidean) space, time and space are independent, and the metric is rep-
resented by the 3 × 3 identity matrix, which yields the commonly defined length of
a vector (corresponding to the familiar Pythagorean distance formula). This is valid
in the limit of weak gravitational fields and low velocities.
• In general relativity, the metric determines how distances and angles are measured
in a spacetime that is curved due to strong gravitational fields. Examples of curved
spacetimes include the Schwarzschild metric around a non-rotating spherical mass
and the Kerr metric around a rotating mass. The theory provides accurate de-
60
scriptions of phenomena at velocities close to the speed of light, and extends to
non-inertial (accelerating) frames.
The Minkowski metric, used in special relativity, has components that reflect the insepa-
rability of space and time and ensures the preservation of the interval ds2 under Lorentz
transformations. We have encountered this matrix in chapter 2 but we recall it here for
easier reference:
−1 0 0 0
0 1 0 0
g=
0
: Minkowski metric. (3.23)
0 1 0
0 0 0 1
In Euclidean space, the squared length of a 3-vector is given by ∥r∥2 = r⊤ Ir, where I
is the identity matrix, serving as the metric. In Minkowski space, the squared length of a
4-vector Aµ is defined analogously, but with the Minkowski metric g:
Given:
a0
1
a µ ⊤
Aµ = 0 1 2 3
a2 , and (A ) = a , a , a , a ,
a3
Based on this definition of the length, we have ∆s2 = (∆X µ )2 . This quantity remains
invariant for any interval between two events, including the interval between the origin O
(where both t and the spatial coordinates are zero) and a point with arbitrary coordinates.
Consequently, (X µ )2 is invariant for any position 4-vector.
Using this invariant, we can recover the relationship between the matrices g and Λ that
we derived in chapter 2 (see eq. [2.32]) and used to derive the general expression for Λ. We
recall this result here for easy reference. If we transform into a frame (S′ ), the invariance
of the square of the length means that (X ′µ )2 = (X µ )2 , that is to say (X ′µ )⊤ gX ′µ =
61
(X µ )⊤ gX µ . Using X ′µ = ΛX µ , this yields (X µ )⊤ Λ⊤ gΛX µ = (X µ )⊤ gX µ . Since this
must be satisfied for any position 4-vector X µ , it follows that :
Λ⊤ gΛ = g. (3.26)
Now, consider an arbitrary 4-vector Aµ . In frame (S′ ), it is given by A′µ = ΛAµ . This
yields (A′µ )2 ≡ (A′µ )⊤ gA′µ = (ΛAµ )⊤ gΛAµ = (Aµ )⊤ Λ⊤ gΛAµ = (Aµ )⊤ gAµ ≡ (Aµ )2 .
Thus, we conclude that:
Aµ · B µ ≡ (Aµ )⊤ gB µ . (3.27)
Aµ · B µ = −a0 b0 + a1 b1 + a2 b2 + a3 b3 , (3.28)
While matrix notation is convenient for deriving the Lorentz transformation via Lie al-
gebra, index notation, which explicitly handles matrix indices, is generally more suitable
for working with the equations of special (and general) relativity. This approach will be
discussed in the following section.
Aµ · B µ = gµν aµ bν , (3.29)
where gµν are the coefficients of the matrix g. As usual, the first index µ indicates a row,
and the second index ν indicates a column. The above expression utilises the Einstein
summation convention for repeated indices, which means it is a shorthand for:
3 X
X 3
Aµ · B µ = gµν aµ bν .
µ=0 ν=0
62
Covariant and contravariant four-vectors
To simplify the notation further, for any column 4-vector Aµ , we define the corresponding
row 4-vector denoted Aµ , with components aµ given by:
aµ = gµν aν . (3.30)
The expression (3.30) provides the relationship between the components of Aµ and those
of Aµ . However, we often write Aµ = gµν Aν , mixing index and vector notation, as a
shorthand for equation (3.30). This practice allows us to use the generic name of the
4-vector as a proxy notation for its components, bypassing the need to explicitly define
those components.
Just as we derived Aµ from Aµ , we can reverse the process. Let g µν represent the
coefficients of the matrix g −1 :
where δνµ is the Kronecker symbol (equal to 1 if µ = ν, and 0 otherwise). Even though
g −1 = g, we use distinct notations for the coefficients to enhance clarity. Specifically,
g 00 = g00 = −1, while all other non-zero components are equal to 1. Using equation (3.30),
we obtain g µν aν = g µν gνξ aξ = δξµ aξ . Therefore:
aµ = g µν aν . (3.33)
This expression can also be derived using equation (3.31), which gives Aµ = g −1 (Aµ )⊤ .
Once again, the notation Aµ = g µν Aν is often used as a shorthand for equation (3.33).
The column 4-vector Aµ is known as a contravariant vector, while the row 4-vector Aµ
is known as a covariant vector. The operations described by equations (3.30) and (3.33)
are commonly referred to as ‘lowering’ and ‘raising’ indices, respectively.
63
Transformation of a covariant four-vector
The Minkowski metric satisfies g ⊤ = g, but the result we derive here is general and
applies to any metric, so we will not rely on this property. When transforming to frame
(S′ ), the expression above becomes A′µ = (A′µ )⊤ g ⊤ = (ΛAµ )⊤ g ⊤ = (Aµ )⊤ Λ⊤ g ⊤ . Taking
the transpose of equation (3.26) yields g ⊤ = Λ⊤ g ⊤ Λ, which implies Λ⊤ g ⊤ = g ⊤ Λ−1 .
Substituting this into the expression for A′µ , we obtain A′µ = (Aµ )⊤ g ⊤ Λ−1 , leading to:
The square length of Aµ is defined by equation (3.24). Using equation (3.31), we have
⊤
gAµ = (Aµ )⊤ and (Aµ )⊤ = Aµ g −1 . Substituting these into equation (3.24), we obtain:
⊤
(Aµ )2 = Aµ g −1 (Aµ )⊤ .
Since this expression is now in terms of Aµ , it provides a definition for (Aµ )2 , which is
thus equal to (Aµ )2 :
⊤ ⊤
(Aµ )2 ≡ Aµ g −1 (Aµ )⊤ = Aµ g −1 (Aµ )⊤ . (3.35)
Using equation (3.30), we can replace gµν bν with bµ in equation (3.29). This yields the
following expression for the scalar product: Aµ · B µ = Aµ Bµ . Alternatively, this can be
derived by writing Aµ · B µ = (Aµ )⊤ gB µ = (Aµ )⊤ gg −1 (Bµ )⊤ = (Bµ Aµ )⊤ = Bµ Aµ , where
the final equality follows from the fact that the transpose of a scalar is the scalar itself.
64
Since the Minkowski metric is symmetric, meaning gµν = gνµ , we can also replace gµν aµ
with aν instead in equation (3.29), which yields Aµ · B µ = Aµ B µ .
3
X 3
X
µ µ µ µ ν
A · B = A B µ = Aµ B = a bν = aν bν . (3.36)
ν=0 ν=0
The first term serves as a formal notation for the scalar product, the second and third
terms represent the standard product of a row vector by a column vector, and the fourth
and fifth terms are expressed in terms of the individual components of these vectors.
Note that the product of a column vector with another column vector, such as Aµ B µ ,
is not defined according to the matrix and vector multiplication rules we have used so far.
The same is true for Aµ Bµ . These products will be dealt with using tensor operations.
Naturally, the above expression remains valid when B µ = Aµ , thus Aµ · Aµ = Aµ Aµ .
The process of summing over a pair of one upper (contravariant) and one lower (covari-
ant) index is called index contraction. Performing index contraction on a 4-vector results
in a scalar.
Conversely, if Aµ is a covariant 4-vector and Aµ B µ remains invariant under
transformation, then B µ is also a 4-vector and is contravariant. Indeed, the
invariance of the product requires that A′µ · B ′µ = Aµ · B µ , which can be expressed as
A′µ B ′µ = Aµ B µ . Using equation (3.34), this leads to Aµ Λ−1 B ′µ = Aµ B µ . Multiplying
both sides on the left by Λ (Aµ )−1 yields B ′µ = ΛB µ . This demonstrates that B µ is a
contravariant 4-vector. Similarly, if B µ is a contravariant 4-vector and Aµ B µ is invariant,
then Aµ is also a 4-vector and is covariant.
Until now, we have carefully distinguished between row and column 4-vectors. Howe-
ver, except when using matrix notation to multiply 4-vectors by matrices, this distinction
is not essential, just as it is unnecessary to specify whether a spatial vector in 3-dimensional
Euclidean space is a row or a column vector. Therefore, from this point forwards, we
will represent all covariant and contravariant 4-vectors as row vectors, under-
standing that Aµ B µ denotes the sum aµ bµ of the products of their components, analogous
to the scalar product of two spatial vectors. By using indices to refer to the components
of vectors and matrices, we eliminate the need for explicit matrix notation. This approach
will also prove to be much more effective for handling tensors.
65
3.6.4 Covariant and contravariant derivatives
In this section, we extend the various derivative operators defined in Euclidean space to
Minkowski spacetime.
Gradient
frame (S). The change in f between two infinitesimally separated points in spacetime is
given by the differential:
∂f
dfS = dxµ . (3.37)
∂xµ
In frame (S′ ), f is a function of x′0 , x′1 , x′2 , x′3 , and its differential is:
∂f
dfS′ = dx′µ . (3.38)
∂x′µ
Since x′µ is itself a function of the coordinates in (S), we have:
∂x′µ ν
dx′µ = dx , (3.39)
∂xν
which yields:
∂f ∂x′µ ν ∂f
dfS′ = dx = dxν . (3.40)
∂x′µ ∂xν ∂xν
The second equality follows from the chain rule. Therefore dfS′ = dfS , indicating that
the differential of a scalar function is invariant.
Equation (3.37) expresses the fact that the scalar product of the contravariant 4-
vector dX µ , whose components are dxµ , with the vector whose components are ∂f /∂xµ ,
is invariant. Thus, this vector is a covariant 4-vector, denoted ∂µ f . More generally,
we define the gradient operator as a covariant 4-vector:
∂ ∂ ∂ ∂ 1∂ ∂ ∂ ∂
∂µ = 0
, 1, 2, 3 = , , , . (3.41)
∂x ∂x ∂x ∂x c ∂t ∂x ∂y ∂z
Introducing the standard gradient operator in 3-dimensional space, ∇, this can be written
in a more compact form as:
1∂
∂µ = ,∇ . (3.42)
c ∂t
66
More generally, we can write:
∂ ′µ = ∂µ Λ−1 . (3.43)
µ 1∂
∂ = − ,∇ . (3.45)
c ∂t
∂ ′µ = Λ∂ µ . (3.46)
Although ∂µ and ∂ µ are operators, they can be manipulated as if they were ordinary
4-vectors. This allows us to use index contraction and write ∂µ = gµν ∂ ν , or ∂ µ = g µν ∂ν ,
where the 4-vector name serves as a proxy notation for its components, as discussed earlier.
The results above hold for any Lorentz transformation. In the specific case of a boost,
∂ ′µ f
= Λ∂ µ can be explicitly written as:
1 ∂ 1∂
− γ −βγ 0 0 −
c ∂t′ c ∂t
∂ ∂
∂x′ −βγ γ 0 0
= ∂x .
(3.47)
∂ ∂
∂y ′ 0 0 1 0
∂y
∂ ∂
0 0 0 1
∂z ′ ∂z
Divergence
So far, we have considered cases where the 4-vector gradient operator acts on a scalar
function, yielding a 4-vector. This can be expressed by stating that the product of a
gradient 4-vector and a scalar function results in a 4-vector. Now, we consider the scalar
product of the gradient 4-vector with another 4-vector, which results in a scalar. We
67
start with a contravariant 4-vector Aµ , whose components aµ depend on the spacetime
coordinates. The divergence of this 4-vector is defined as:
1 ∂a0
∂ µ · Aµ = ∂µ Aµ = + ∇ · A. (3.48)
c ∂t
In the Minkowski metric, Aµ has the same spatial components as Aµ , leading to:
1 ∂a0
∂ µ · Aµ = ∂ µ Aµ = − + ∇ · A. (3.49)
c ∂t
The divergence remains invariant as it represents the scalar product of two 4-vectors,
ensuring that conservation laws hold in all inertial frames.
D’Alembertian
Now, consider the effect of taking the divergence of the 4-vector ∂ µ f . We denote the
resulting operator (known as d’Alembertian, or wave operator) by the symbol □, and we
have:
1 ∂2f ∂2f ∂2f ∂2f
□f = ∂ µ · ∂ µ f = ∂µ ∂ µ f = − 2 2 + + + . (3.50)
c ∂t ∂x2 ∂y 2 ∂z 2
1 ∂2
□ = ∂ µ · ∂ µ = ∂µ ∂ µ = ∂ µ ∂µ = − + ∇2 . (3.51)
c2 ∂t2
68
the wave equation retains its form in any inertial frame. Specifically, if a scalar
function f (xµ ) satisfies the wave equation in frame (S), then f (x′µ ) will also satisfy the
wave equation in frame (S′ ) when expressed in terms of the coordinates in (S′ ), with
derivatives taken with respect to x′µ .
We have introduced index notation for 4-vectors and the metric matrix g, but not yet for
the Lorentz transformation matrix. This section addresses that. Consider a contravariant
4-vector Aµ . In matrix notation, it transforms as A′µ = ΛAµ . In previous courses, we
referred to the coefficients of a matrix M as Mij , where i indicates a row and j indicates a
column. However, using only lower indices to label the coefficients of Λ would imply that
both indices are covariant. While this might not pose a problem for 4-vectors, it can lead
to confusion when dealing with tensors, as we will see later in the course.
Thus, in index notation, the relationship between A′µ and Aµ is written as:
Here, Aν serves as a proxy notation for its components aν . Thus, the right-hand side is
the sum of the products of the coefficients of Λ and the components of Aν . Since Aν is a
column 4-vector, consistency with matrix notation requires that the contravariant (upper)
index µ in Λµν corresponds to a row, while the covariant (lower) index ν corresponds to a
column2 . Staggering the indices clarifies the notation.
The matrix Λ has dimension 4 × 4, but by contracting the lower index (multiplying by Aν
and summing over ν), we obtain a 4-vector.
With this notation, we can express the product of the matrices gΛ as gλκ Λκν . Conse-
quently, equation (3.26) becomes:
2
For the metric matrix g, it is unnecessary to distinguish between upper and lower positions for the
indices, as the matrix is symmetric, and therefore the roles of the indices are interchangeable. Therefore,
we use the notations gµν or g µν , rather than g µν .
69
Λλµ gλκ Λκν = gµν . (3.54)
By multiplying both sides of the second equation by g ρµ and gνπ (and summing over µ
and ν), and using equation (3.32), we obtain:
This notation extends the concept of lowering and raising indices from 4-vectors to ma-
trices.
ν
From equation (3.34), we also have A′µ = Aν Λ−1 µ . Comparing this with equation (3.55)
ν
implies Λµν = Λ−1 µ . Using the definition of Λµν from equation (3.55), we can verify
that Λµν Λµζ = gµκ Λκλ g λν Λµζ = gζλ g λν = δζν , where we have used equation (3.54).
We summarise these results:
ν
Λµν = Λ−1 µ
or equivalently Λµν Λµζ = δζν . (3.57)
∂x′µ
≡ ∂ν x′µ = Λµν , (3.58)
∂xν
meaning that the Lorentz transformation matrix Λ is the Jacobian matrix of the
transformation from the coordinates xν to x′µ . We can thus define a 4-vector Aµ as any
set of spacetime components that transform according to:
∂x′µ ν
A′µ = A = Aν ∂ν x′µ . (3.59)
∂xν
70
Equation (3.59) provides the most general definition of a 4-vector, as it encompasses
all possible transformations from one set of coordinates xµ to another set x′µ . The (linear)
Lorentz transformation is merely a special case applicable to special relativity. In contrast,
general relativity allows for transformations that are not necessarily linear and can involve
any smooth mapping between coordinate systems.
µ
The inverse transformation gives dX µ = Λ−1 ν
dX ′ν = Λν µ dX ′ν . This implies that:
∂xµ
≡ ∂ ′ν xµ = Λν µ . (3.60)
∂x′ν
From equation (3.55), we have A′µ = Λµν Aν . Therefore, equation (3.60) provides the
transformation of a covariant vector as:
∂xν
A′µ = Aν = Aν ∂ ′µ xν . (3.61)
∂x′µ
71
72
Chapter 4
In this chapter, the spacetime position 4-vector is represented as X µ = (ct, r), where
r = (x, y, z) is the usual spatial 3-vector. We also use the notation X µ = x0 , x1 , x2 , x3
when convenient.
4.1 Four-velocity
We begin by constructing a 4-vector for velocity, extending the concept from 3-dimensional
space to spacetime.
dX µ
Uµ = , (4.1)
dτ
where τ is the proper time along the worldline. Since dτ is invariant and dX µ is a 4-vector,
U µ is indeed a 4-vector. Explicitly, this is written as:
dt dx dy dz dt dt
Uµ = c , , , = c ,u .
dτ dτ dτ dτ dτ dτ
73
As discussed in section 3.4.1, a massive particle moves along a timelike worldline, corre-
p
sponding to ds2 = −c2 dt2 + dr2 = −c2 dτ 2 . This gives: dt/dτ = dt/ dt2 − dr2 /c2 =
p
1/ 1 − u2 /c2 , which can be written as:
dt
= γu , (4.2)
dτ
p
where γu = 1/ 1 − u2 /c2 . Therefore:
U µ = (γu c, γu u) . (4.3)
Since U µ is a 4-vector, its squared length is an invariant. In the rest frame of the particle,
where u = 0, we have U µ = (c, 0), resulting in:
(U µ )2 = −c2 . (4.4)
Consider a particle with velocity u in frame (S) and velocity u′ in frame (S′ ), where (S′ )
moves with velocity v along the x-axis relative to (S). The 4-velocity in (S′ ) is given by
U ′µ = ΛU µ , which can be explicitly written as:
γu′ c γv −βγv 0 0 γu c
γu′ u′x −βγv
γv 0 0 γu ux
γ ′ u′ = 0
, (4.5)
u y 0 1 0
γu uy
γu′ uz′ 0 0 0 1 γu uz
where β = v/c and the subscript ‘v’ on γ indicates that it is related to the velocity of
frame (S′ ). This transformation yields the following equations:
γu′ c = γv γu (c − βux ) ,
γu′ u′x = γv γu (−βc + ux ) ,
γu′ u′y = γu uy ,
γu′ u′z = γu uz .
Dividing the last three equations by the first one, we can recover the components of u′
calculated in section 3.5.2.
74
4.2 Four-momentum
In his seminal 1905 paper On the Electrodynamics of Moving Bodies, Einstein derived the
expressions for relativistic momentum, p = γmv, and energy, E = γmc2 , by analyzing
how momentum and energy must transform to be conserved across all inertial frames while
remaining consistent with the constancy of the speed of light. In this section, we will use
the principle of least action to construct a relativistic Lagrangian and recover these results.
The determination of the appropriate Lagrangian for a physical system is guided by sym-
metry principles, among other factors. In classical mechanics, the action should be in-
variant under Galilean transformations, while in special relativity, it should be invariant
under Lorentz transformations, ensuring the consistency of physical laws across all inertial
frames. If the equations of motion are already known, we can work backwards to identify a
corresponding action that reproduces these equations. Alternatively, proposing an action
based on symmetry principles, physical intuition, or other considerations can lead to new
equations of motion that may predict previously unknown phenomena.
In special relativity, we already know that the equation of motion is dp/dt = f , with
p = γmv, which serves as a benchmark for constructing a relativistic Lagrangian. Once
such a Lagrangian is established, it can be used to further develop the theory of special
relativity, particularly in the context of electromagnetism.
To extend the concept of action to the relativistic case, we must ensure that it remains
stationary along a worldline and is independent under changes of reference frame. This
means that if the action is stationary in one frame, it must also be stationary in all
frames. For a single particle, this condition is satisfied if the Lagrangian is proportional
to the spacetime interval ds along the worldline, which is itself proportional to the proper
time interval dτ . Thus, we express the action as:
ˆ τB
S∝ dτ.
τA
For a particle moving with velocity u, and using equation (4.2), this becomes S ∝
´ tB p ´t
1 − u2 /c2 dt. In the limit u/c ≪ 1, this integral reduces to tAB 1 − u2 /(2c2 ) dt, to
tA
the first non-zero order in u/c. In classical mechanics, in the absence of external forces,
the Lagrangian is given by the kinetic energy mu2 /2. This expression should match the
75
relativistic Lagrangian in the limit u/c ≪ 1, up to an additive constant (since adding
a constant to the Lagrangian does not alter the equations of motion). Therefore, the
proportionality coefficient in the relativistic Lagrangian is −mc2 .
If the particle interacts with an external force, an additional term must be included
in the Lagrangian to account for it. We denote this term as Lint and assume it depends
on r and t, but not on ṙ. The more general case, where Lint may also depend on ṙ, will
be discussed later in the course. This term is constructed so that that it does not alter
the invariance of the action, ensuring it remains stationary in all frames, and it yields the
correct equations of motion through Euler-Lagrange equations.
The relativistic Lagrangian for one particle is then:
r
2 ẋ2 + ẏ 2 + ż 2
L = −mc 1− + Lint (r, t) . (4.6)
c2
Consider a system of N particles. If they do not interact with each other, the Lagrangian
for the system is simply the sum of the individual Lagrangians. When mutual interactions
are present, they are incorporated into Lint such that the Lagrangian becomes:
s
N
X ẋ2 + ẏk2 + żk2
L= −mk c2 1 − k + Lint (r1 , r2 , . . . , rN , t) (4.7)
c2
k=1
Here, the subscript k labels the k-th particle, and does not indicate a covariant coordinate.
Therefore, rk represents the position vector of the k-th particle.
In chapter 1, we noted that the Euler-Lagrange equations are applicable regardless of the
explicit form of the Lagrangian, and therefore remain valid for the relativistic Lagrangian.
These equations (1.20) state that, for each particle k:
d ∂L ∂L d ∂L ∂L d ∂L ∂L
= , = , = . (4.8)
dt ∂ ẋk ∂xk dt ∂ ẏk ∂yk dt ∂ żk ∂zk
In the absence of internal or external forces, the right-hand side of these equations is
zero, implying that the quantities ∂L/∂ ẋk , ∂L/∂ ẏk and ∂L/∂ żk are conserved. These
conserved quantities are the components of momentum. Therefore, for each particle, and
now dropping the subscript k, we express the spatial momentum vector as p = (px , py , pz )
with px = ∂L/∂ ẋ = γu mux , and similarly for the other components, where we have used
L given by equation (4.6). This gives:
p = γu mu. (4.9)
76
In the limit u/c ≪ 1, this simplifies to the familiar p = mu.
In section 1.3.2, we also noted that Noether’s theorem remains valid for the relativistic
Lagrangian. This implies that, when the system is invariant under spatial translation,
which is the case when there are no external forces acting on the system, the total
momentum is conserved.
Energy
Furthermore, when the Lagrangian does not explicitly depend on time, there exists a
conserved quantity, the energy E, given by equation (1.60). For a single particle, this
equation is:
∂L ∂L ∂L
E= ẋ + ẏ + ż − L = const. (4.10)
∂ ẋ ∂ ẏ ∂ ż
Assuming Lint = 0 in the expression (4.6) for the Lagrangian, this is equal to p · u +
mc2 /γu = γu m u2 + c2 /γu2 , which simplifies to:
E = γu mc2 . (4.11)
This represents the total energy of a single particle in the absence of external forces and
is commonly referred to as the relativistic energy. In the limit u/c ≪ 1, this reduces
to:
1
E = mu2 + mc2 .
2
The constant term mc2 is the rest energy of the particle. It represents the energy that
the particle possesses due to its mass alone, independent of its motion. This term is
significant in relativity, illustrating the concept that mass and energy are equivalent. It
does not appear in classical mechanics, where energy is purely a function of motion and
position, and does not account for the intrinsic energy of mass.
Since mc2 is the energy of a particle when it is at rest, it is a fixed value and thus
Lorentz invariant. Therefore, the mass m is an intrinsic property of the particle.
By subtracting the rest energy from the total energy E, we obtain the kinetic energy
(in the absence of external forces and therefore of other forms of energy, such as potential
energy):
K = (γu − 1) mc2 .
For a system of N particles, equation (4.10) still applies, provided the first three terms
are summed over all the particles. If there are no external nor internal forces (Lint = 0),
the conserved quantity is the total energy:
N
X
Etotal = γuk mk c2 . (4.12)
k=1
77
4.2.3 Four-momentum and energy-momentum relation
µ µ E
P = mU = ,p . (4.13)
c
Note that E here includes only the rest and kinetic energies. The 4-momentum is de-
fined this way even when additional energy, such as potential energy from external forces,
is present, meaning that potential energy does not directly contribute to the 4-momentum.
E 2 = m2 c4 + p2 c2 . (4.15)
This expression holds for any particle with mass m. It can be easily checked that it is
consistent with p = γu mu and E = γu mc2 .
v
E ′ = γv (E − vpx ) , p′x = γv − 2 E + px , p′y = py , p′z = pz . (4.16)
c
Massless particles travel with speed u = c, leading to a situation where the momentum
γu mu involves multiplying an infinite quantity (γu ) by zero (m). In this case, we define the
momentum in terms of energy rather than velocity. The energy-momentum relation (4.15)
shows that E 2 → p2 c2 as m → 0. This expression remains valid for m = 0, leading to:
E
p= . (4.17)
c
78
Thus, the 4-momentum of a massless particles is given by:
µ E E
P = , n̂ , (4.18)
c c
where n̂ is the unit vector in the direction of propagation of the particle. This implies:
(P µ )2 = 0. (4.19)
The energy of a photon is given by E = hν, where ν is the frequency. For a photon
moving along the x-axis, the 4-momentum in frame (S) is P µ = (hν/c, hν/c, 0, 0). In a
frame (S′ ) moving with velocity v relative to (S), and using equation (4.16) along with
β = v/c, we find that hν ′ /c = γv (1 − β) hν/c, which simplifies to:
s
1−β
ν′ = ν. (4.20)
1+β
This equation represents the Doppler shift, which can also be derived using the concepts
of wavelength contraction and time (period) dilatation.
Pµ
µ E p
K = = , . (4.21)
ℏ ℏc ℏ
ω
Kµ = ,k , (4.22)
c
where ω = 2πν is the angular frequency, and the wavevector is related to the wavelength
by k = 2π/λ.
79
Since K µ is a 4-vector, its squared length (K µ )2 = −ω 2 /c2 + k 2 , where k 2 = k · k, is
invariant. For massless particles, as discussed in the previous subsection, this invariant
is zero. This leads to the condition ω/k = c, which shows that massless particles are
described by waves propagating at the speed of light.
For massive particles, in the rest frame where p = 0 and E = mc2 , the invariant
calculated from equation (4.21) is (mc/ℏ)2 . This gives the following dispersion relation:
2 2
2 2 2 mc
ω =c k + . (4.23)
ℏ
Consider the spacetime position 4-vector X µ = (ct, r). The scalar product X µ · K µ =
k · r − ωt is invariant. This quantity represents the phase of the wave associated with the
particle.
The wave 4-vector was initially introduced by defining it in terms of the 4-momentum
of a particle. Now, consider a wave not associated with a particle, such as a sound wave,
with the form:
f (r, t) = A cos (k · r − ωt) . (4.24)
Can we still describe this wave using a 4-vector K µ , as defined in equation (4.22)? This
would imply that the phase ϕ = k · r − ωt of the wave is invariant under Lorentz trans-
formations. We observe that ∂ µ ϕ is a 4-vector, and it is equal to (ω/c, k) = K µ . This
confirms that K µ in this context is indeed a 4-vector, and the invariance of X µ · K µ here
indicates that the phase of the wave is invariant.
4.3 Four-force
The Euler-Lagrange equations (4.8) show that, when a system of N particles is subject to
external and/or internal forces, the momentum of the k-th particle changes according to:
dpk
= fk , (4.25)
dt
where fk = ∇rk Lint is the force acting on the k-th particle. The subscript on the gradient
operator indicates the coordinates with respect to which the derivatives are calculated.
In the same way that we cannot create a velocity 4-vector by dividing dX µ by dt,
even though u = dr/dt, we also cannot create a force 4-vector by diving dP µ by dt, even
though the force is given by dp/dt. Instead, we define the force 4-vector, or 4-force, F µ ,
as:
dP µ
Fµ = . (4.26)
dτ
Using dt/dτ = γu , this yields F µ = γu (dP µ /dt), which can be written as:
µ 1 dE
F = γu , γu f , (4.27)
c dt
80
where f is the force acting on the particle (we now omit the subscript k). These two
equations establish the relativistic form of the equation of motion, valid in any inertial
frame.
The scalar product of the two 4-vectors U µ and F µ is invariant and equal to its value
in the particle’s rest frame. Therefore:
µ µ 2 dE dm
U · F = γu − + f · u = −c2 , (4.28)
dt dτ
where we have used E = mc2 and t = τ in the rest frame.
When a particle interacts with a field, its mass can change. This principle underlies
the Higgs mechanism, where particles acquire mass through their interaction with the
Higgs field. During this process, the particle’s energy does not follow E = mc2 , as energy
is exchanged through complex quantum field interactions. However, once the process
completes and the particle acquires its rest mass, E = mc2 applies in the rest frame, and
the mass remains constant, so dm/dτ = 0. Equation (4.28) then implies:
dE
= f · u, (4.29)
dt
where E represents the relativistic energy, which includes both kinetic energy and rest
mass energy. Since the rest mass is constant, dE/dt is the rate of change of kinetic energy.
Just as in classical mechanics, this is equal to the work done by the force per unit time.
Since F µ is a 4-vector, it transforms according to Lorentz transformations, allowing us
to calculate the force in different inertial frames
4.4.1 Four–acceleration
Given that P µ = mU µ , equation (4.26) can be rewritten as:
dU µ
Fµ = m ,
dτ
which leads us to define the 4-acceleration:
dU µ
Aµ = . (4.30)
dτ
81
We have:
dγu 1 d (u · u) u·a
= γu3 2 = γu3 2 , (4.32)
dt 2c dt c
where a = du/dt is the standard spatial acceleration vector.
Therefore, the 4-acceleration can be expressed as:
u·a u·a
Aµ = γu2 γu2 , γu2 2 u + a . (4.33)
c c
Because f = dp/dt and p = γu mu, there is no simple relation between the force f and the
acceleration a.
When a particle accelerates (with acceleration that may vary over time), its rest frame
can be defined as the inertial frame in which the particle is momentarily at rest at any
given instant. This frame, called the instantaneous rest frame, changes instantaneously
as the particle’s velocity changes. In this frame, the particle experiences an acceleration
known as proper acceleration, which is the physical acceleration felt by the particle and
measurable by an accelerometer moving with it. This can be shown by expressing the
4-force in this frame as:
µ dm
F = c , f0 , (4.34)
dτ
where f0 represents the force in the rest frame, and we use dE/dτ = c2 dm/dτ . Since the
particle’s velocity is zero in its instantaneous rest frame, equation (4.33) implies that the
4-acceleration in this frame is:
Aµ = (0, a0 ) , (4.35)
where a0 is, by definition, the proper acceleration. Therefore, F µ = mAµ yields a0 = f0 /m.
This result also implies that dm/dτ = 0, meaning the mass remains constant in the
particle’s rest frame, as discussed in section 4.3.
Since Aµ is a 4-vector, its squared length is an invariant, leading to the relation:
(Aµ )2 = a20 .
In the rest frame, U µ has no spatial component while Aµ only has a spatial component.
Consequently, their scalar product is zero. Since this result is invariant, it implies that
the 4-velocity and 4-acceleration are always orthogonal, leading to:
U µ · Aµ = 0. (4.36)
82
4.4.3 Proper acceleration and rapidity
Consider a particle initially at rest in an inertial frame (S). In the particle’s instantaneous
rest frame, time is measured as the proper time τ , with τ = 0 marking the instant the
particle begins to accelerate with a proper acceleration a0 . At a later time τ , the particle
is moving with velocity u(τ ) in frame (S), and its instantaneous rest frame, denoted as
S(τ ), moves relative to (S) with the same speed.
At time τ + dτ , the particle’s velocity in frame S(τ ) is a0 dτ . The particle’s velocity
u (τ + dτ ) in frame (S) can be found using the velocity transformation given by equa-
tion (3.20):
a0 dτ + u(τ )
u (τ + dτ ) = .
1 + a0 dτ u(τ )/c2
Here, we assume that a0 and u(τ ) are aligned, so that u (τ + dτ ) remains parallel to the
boost velocity. However, this result could be extended to more general cases. Defining
du = u (τ + dτ ) − u(τ ), and neglecting the second-order term in dudτ , we obtain:
du
= a0 dτ.
1 − u2 /c2
The left-hand side simplifies to c d tanh−1 (u/c) , where tanh−1 (u/c) is the rapidity ζ,
ˆ τ
1
ζ(τ ) = a0 dτ ′ , (4.37)
c 0
where we have used ζ = 0 at τ = 0. In the non-relativistic limit u/c ≪ 1, ζ ≃ u/c and the
´τ
equation simplifies to u(τ ) = 0 a0 dτ ′ , as expected.
83
N
Etotal
Pkµ
X
µ
P = = , ptotal , (4.38)
c
k=1
is conserved if evaluated long before and long after the collision. Here the subscript k
refers to the k-th particle.
This sum is also a 4-vector. Indeed, each Pkµ is a 4-vector, transforming to a frame
(S′ ) as Pk′µ = ΛPkµ , where the coefficients of Λ depend only on the velocity between frames
(S) and (S′ ). If all Pkµ are evaluated simultaneously in frame (S), the transformation
might result in the Pk′µ not being simultaneous in frame (S′ ). However, if the 4-vectors
are evaluated when the particles are far apart, they are time-independent, as each particle
conserves its energy and momentum. Thus, P ′µ = ΛP µ , indicating that P µ is itself a
4-vector, and therefore its squared length is invariant. Calculating it in the centre of
momentum frame, where ptotal = 0, we obtain (P µ )2 = −Etotal 2 /c2 = − k (γuk mk )2 c2 ,
P
If we denote the energy, momentum and mass of the system (comprising all particles) as
E, p and m respectively, and use subscripts ‘before’ and ‘after’ to indicate these quantities
before or after the collision, we have the following results:
• Classical collisions:
• Relativistic collisions:
– Mass:
84
4.5.2 Centre of momentum frame
Let P µ = (Etotal /c, ptotal ) represent the 4-momentum of the entire system of particles in
frame (S), and let P ∗µ = (E ∗ /c, 0) represent the 4-momentum in the centre of momentum
(CM) frame. Here, E ∗ is the energy of the system in the CM frame. The CM frame moves
with velocity vcm relative to frame (S), and we choose the x-axis to be in the direction of
that velocity. We denote γcm as the Lorentz factor associated with the velocity vcm and
Λcm as the transformation matrix from frame (S) to the CM frame.
Using Λcm P µ = P ∗µ , we have γcm −vcm Etotal /c2 + ptotal,x = 0 and ptotal,y = ptotal,z =
0. This implies that the x-axis, that is, the direction of vcm , is along ptotal , and:
ptotal c2
vcm = . (4.39)
Etotal
In the diagram, we have assumed that the proton moving forwards in the CM frame also
moves upwards. However, if it were to move downwards instead, the results presented
here would remain unchanged. Hereafter, we denote quantities in the CM frame with a
superscript ‘∗’ and quantities after collision with a prime. Velocities are denoted u, except
for the centre of mass velocity, which is vcm .
85
In classical mechanics, the conservation of momentum in frame (S) implies (p′1 + p′2 )2 =
p′2 ′2 ′ ′ 2 ′2 ′2 2
1 + p2 + 2p1 · p2 = p1 , and the conservation of energy implies v1 + v2 = v1 . This leads
to p′1 · p′2 = 0, indicating that the protons scatter at right angle to each other.
We will now show that this is no longer the case in relativity. Denoting the mass of
the proton as mp , we have the following relation in the CM frame:
mp u∗1 mp u∗2
p∗1 = p∗2 ⇒ p = p ⇒ u∗1 = u∗2 .
1 − u∗2
1 /c2 1 − u ∗2 /c2
2
Since u2 = 0, the velocity transformation from frame (S) to the CM frame yields u∗2 =
−vCM . Therefore, u∗1 = u∗2 = vcm .
Similarly, the fact that the protons have equal and opposite momenta after the collision
implies u′∗ ′∗
1 = u2 . Conservation of energy in the CM frame then implies that the γ factors
of the protons must be the same before and after the collision. Consequently, the velo-
cities after the collision are the same as before, meaning u′∗ ′∗
1 = u2 = vcm , which implies
p′∗ = p∗ = γcm mp vcm .
E∗ ∗
P1′∗µ = ∗ ∗ ∗
, p cos θ , p sin θ , 0 ,
c
E1′ ′
P1′µ = ′
, p cos θ1 , p1 sin θ1 , 0 .
c 1
We have P1′µ = (Λcm )−1 P1′∗µ , where (Λcm )−1 is the transformation associated with the
velocity −vcm . This yields:
p′1 cos θ1 = γcm vcm E ∗ /c2 + p∗ cos θ∗ and p′1 sin θ1 = p∗ sin θ∗ .
sin θ∗ sin θ∗
tan θ1 = . Similarly, tan θ 2 = .
γcm (1 + cos θ∗ ) γcm (1 − cos θ∗ )
tan θ1 + tan θ2 2
tan (θ1 + θ2 ) = = 2 .
1 − tan θ1 tan θ2 (γcm − 1) sin θ∗
86
4.5.4 Emission and absorption
Spontaneous emission
Spontaneous emission is a process in which an excited atom (or molecule) loses energy
by emitting a photon and transitioning to a lower energy state. In the rest frame of
the atom before emission, the 4-momentum of the atom is P µ = (M c, 0), where M c2
is the rest energy of the excited atom. After emission, the 4-momentum of the atom is
P ′µ = (E ′ /c, p′ ), and that of the photon is Pγ′µ = Eγ′ /c, p′γ . Let M ′ c2 denote the rest
M 2 − M ′2 c2
′ E0
Eγ = = 1− E0 .
2M 2M c2
Radioactive decay
Radioactive decay is a natural process by which an unstable atomic nucleus loses energy by
emitting radiation, which can be in the form of alpha particles, beta particles, or gamma
rays:
• α-decay: The parent nucleus emits an α particle and decays into a lighter nucleus,
In both types of β-decay within a nucleus, the electron or positron and the (anti)neutrino
are emitted, resulting in the transformation of the nucleus into a different element.
• γ-decay: An unstable nucleus releases energy in the form of γ-rays, usually following
α or β-decay when the daughter nucleus is produced in an excited state.
Let us consider β + -decay. If a free proton could undergo this process, it would be rep-
resented as: p −→ n + e+ + νe . In the rest frame of the proton, conservation of energy
87
implies: mp c2 = mn c2 + me c2 + kinetic energy + Eνe , where mn and me are the masses of
the neutron and positron, respectively, and Eνe is the energy of the neutrino, whose mass
is negligible. This requires mp c2 > mn c2 + me c2 . However, this condition is not satisfied,
as mp c2 = 938.27 MeV and mn c2 = 939.57 Mev.
The situation is different when the proton is within a nucleus because of the binding energy.
This is illustrated by the reaction: (A, Z) −→ (A, Z − 1) + e+ + νe , where (A,Z) denotes
a nucleus with atomic number A and charge number Z. Energy conservation now requires
the rest mass of (A,Z) to be greater than the sum of the rest masses of (A,Z-1) and the
positron. This condition is satisfied, for example, for A = 11 and Z = 6, corresponding to
Carbon-11, which can decay into Boron-11.
Absorption
Absorption of a particle as it collides with another particle at rest can result in either the
two particles sticking to each other, or the emission of different particles. Here are a few
examples of such processes:
• Excitation: Absorption of a photon by an atom, raising the energy state of the atom.
When two protons collide with sufficient energy, a pion can be produced:
p + p −→ p + p + π 0 .
In the rest frame of one of the protons before the collision, the 4-momenta are P1µ =
(E1 /c, p1 ) and P2µ = (mp c, 0), where E1 is the energy of the moving proton. Conservation
of energy requires the kinetic energy of the moving proton to be equal to mπ c2 plus the
kinetic energy of the particles after collision, where mπ is the mass of the pion. The
minimum value of E1 required to produce a pion is called the threshold energy. This
corresponds to all the particles being at rest in the CM frame after the collision. The
total 4-momentum in the CM frame is then P ∗µ = (2mp c + mπ c, 0). The squared length
2
of the total 4-momentum is invariant, meaning1 (P1µ + P2µ ) = (P ∗µ )2 , which implies
1
We could define the total momentum before collision as P µ = (mp c + E1 /c, p1 ) and square it, instead
of squaring P1µ + P2µ . However, this would introduce p21 , as (P µ )2 = − (mp c + E1 /c)2 + p21 . Using
p21 c2 = E12 − m2p c4 , we obtain (P µ )2 = −2m2p c2 − 2E1 mp . This result can be derived more directly
by using (P µ )2 = (P1µ + P2µ )2 = (P1µ )2 + (P2µ )2 + 2P1µ · P2µ , which avoids introducing p1 .
88
2 2
(P1µ ) + (P2µ ) + 2P1µ · P2µ = (P ∗µ )2 . This yields −2m2p c2 − 2E1 mp = − (2mp c + mπ c)2 ,
which simplifies to:
(2mp + mπ ) c2 − 2mp c2
E1 = .
2mp
In general, when a particle collides with its antiparticle, they annihilate, producing other
particles, typically photons, neutrinos or lighter particles, depending on the available en-
ergy. The number and types of particles produced depend on the specific interaction and
the energy of the colliding particles. Annihilation leading to photons is a clear example of
the conversion of mass into energy, as the entire rest mass of the particle and antiparticle
is transformed into the energy of the resulting particles.
An example of annihilation producing lighter particles is the collision of a proton with
an antiproton, which results in the production of pions (π + , π − , π 0 ).
When an electron and positron collide at low energies, they annihilate and produce
photons: e− + e+ −→ γ + γ. In the centre of momentum frame of the particles before
the collision, the total momentum is zero. Therefore, after the collision, momentum con-
servation requires the two photons to have equal and opposite momenta. Additionally,
conservation of energy implies that each photon has an energy at least equal to the rest
mass of the electron (or positron, as their rest masses are identical).
The 4-momentum vectors for the photon before and after the collision are:
h h h h h
Pγµ = , , 0, 0 , Pγ′µ = , cos θ, sin θ, 0 .
λ λ λ′ λ′ λ′
89
The 4-momentum vector for the electron before the collision is Peµ = (mp c, 0), and we
denote the 4-momentum after the collision as Pe′µ .
Conservation of 4-momentum gives:
2 2
Pγµ + Peµ − Pγ′µ = Pe′µ .
By isolating Pe′µ , we avoid involving the unknown momentum p′e in the calculation. Using
2 2
the relations (Pγµ ) = Pγ′µ = 0 and (Peµ ) = Pe′µ , we obtain:
2 2
h hh hh h
Pγµ · Peµ − Pγµ Pγ′µ − Peµ · Pγ′µ = − mp c + ′
− ′
cos θ + ′ mp c = 0.
λ λλ λλ λ
Finally, solving this expression yields:
h
λ′ − λ = (1 − cos θ) . (4.40)
mp c
For forward scattering (θ = 0), there is no shift in wavelength (i.e., λ′ = λ), while for
backward scattering (θ = π), the shift is maximized.
90
Chapter 5
In the second-year Electromagnetism course, it was demonstrated that electric and mag-
netic fields transform into each other when moving between reference frames, indicating
that they cannot be treated as separate entities. In this chapter, we will discuss how these
fields combine into a tensor. We begin with an introduction to tensors.
5.1 Tensors
As mentioned earlier, one of the postulates of relativity is that the laws of physics are the
same in all inertial frames. Consequently, we aim to express these physical laws in a way
that remains invariant under changes in coordinates or reference frames. This is referred
to as the covariant formulation of physical laws.
This concept was illustrated in chapter 4, where the laws of dynamics were expressed
using 4-vectors, which transform according to well-defined rules between reference frames.
For example, the law F µ = mAµ in frame (S) retains its form F ′µ = mA′µ in a frame
(S′ ), moving with velocity v relative to (S), because both 4-vectors transform according
to F ′µ = Λµν F ν and A′µ = Λµν Aν .
More generally, tensors are mathematical objects that transform according to rules
that ensure the covariant nature of the physical laws in which they appear.
5.1.1 Preliminary
Consider an ensemble of particles, all with the same mass m and velocity u in frame (S).
In the rest frame (S0 ) of the particles, where there are n0 particles per unit volume, each
particle has an energy of mc2 , resulting in an energy density of n0 mc2 = ρ0 c2 , where
ρ0 ≡ n0 m is the mass density.
Frame (S) moves relative to frame (S0 ) with velocity −u. Due to length contraction along
the direction of this velocity, a unit volume in the rest frame corresponds to a volume
of 1/γu in frame (S). Therefore, the number of particles per unit volume in frame (S) is
91
γu n0 . In this frame, each particle has an energy γu mc2 , so the energy density becomes
2
γu n0 × γu mc2 = γu2 ρ0 c2 . Notably, this equals ρ0 u0 , where u0 is the first coordinate of
the 4-velocity:
U µ = u0 , u1 , u2 , u3 = (γu c, γu ux , γu uy , γu uz ) .
By multiplying the components of U µ with each other and then by ρ0 , we obtain the 16
terms ρ0 uµ uν , denoted as tµν :
where i, j = 1 . . . 3. As discussed earlier, t00 represents the energy density of the ensemble
of particles. Writing t0i = γu n0 × γu mc × ui , we see that this represents the flux of energy
in the i-direction, divided by c. Finally, tij = γu n0 × γu mui × uj represents the flux of the
i–th component of momentum in the j-direction.
These quantities are all physically meaningful, so it is of interest to see how they transform
from one frame to another. We denote the set of numbers tµν as T µν , that is to say
T µν = ρ0 U µ U ν . Using the transformation properties of the 4-velocity, we have:
This is a specific case of tensors, and we now define them more generally.
5.1.2 Definition
µ ...µ
A tensor of type (m, n) is denoted Tν11...νnm and is defined as an object that transforms
under a Lorentz transformation according to:
µ ...µ κ ...κ
T ′ ν11...νnm = Λµ1 κ1 . . . Λµm κm Λν1 λ1 . . . Λνn λn Tλ11...λnm (5.1)
Using equations (3.58) and (3.60), we can alternatively define a tensor as an object that
transforms according to:
92
This more general definition applies to any transformation between the sets of coordinates
xµ and x′µ .
Note that upper and lower indices are sometimes staggered for clarity, though there is
no strict convention regarding which should be positioned to the right. The type of the
tensor refers to the number m of contravariant indices and the number n of covariant in-
dices. The rank, or order, is given by m + n. A tensor T µ of type (1, 0) is a contravariant
vector (rank 1), whereas a tensor Tν of type (0, 1) is a covariant vector (rank 1). Tensors
with both covariant and contravariant indices are referred to as mixed tensors. A scalar is
conventionally described as a tensor of rank 0. In the following sections, we discuss higher
rank tensors.
We now verify that we can raise and lower indices in the same manner as with 4-vectors.
For simplicity, we consider a covariant tensor of order 2, Tαβ , though this process can be
generalised to any mixed type tensor of any order. The tensor transforms as:
′
Tαβ = Λαδ Λβ ϵ Tδϵ = gαα1 Λα1 α2 g α2 δ gββ1 Λβ1 β2 g β2 ϵ Tδϵ ,
where we have used equation (3.55). Next, we multiply both sides by g ζα and g ηβ (and
sum over α and β):
′
g ζα g ηβ Tαβ = g ζα gαα1 Λα1 α2 g ηβ gββ1 Λβ1 β2 g α2 δ g β2 ϵ Tδϵ
ζ η
= δα1 Λα1 α2 δβ1 Λβ1 β2 g α2 δ g β2 ϵ Tδϵ
= Λζ α2 Ληβ2 g α2 δ g β2 ϵ Tδϵ
Multiplication
Aµ = a0 , a1 , a2 , a3 and B ν = b0 , b1 , b2 , b3 ,
93
we can generate 42 = 16 numbers by multiplying each component of Aµ with each com-
ponent of B µ . This set of numbers, aµ bν , forms a tensor, denoted as:
T µν = Aµ B ν , (5.3)
and aµ bν are the components of the tensor. It is straightforward to verify that T µν satisfies
the definition [5.1] of a tensor since both Aµ and B ν transform as 4-vectors:
Addition
Only tensors of the same type (m, n) (and therefore the same rank m + n) can be added
together. This means they must have the same number of covariant and contravariant
indices. The resulting tensor has the same type and rank as the original tensors, and is
formed by adding the corresponding components of the tensors being summed.
Contraction
We have already encountered index contraction in chapter 3 while dealing with 4-vectors.
Now, we extend this concept to more general tensors. Contracting a tensor involves
summing over one upper index and one lower index within the same tensor or across
different tensors.
µ ...µ
Contracting a single tensor Tν11...νnm over indices µi and νj is achieved by setting µi =
νj , renaming it as, say, a, and summing over all possible values of a. Using Einstein’s
summation convention, this is written as:
µ ...µ aµ ...µ
m
Tν11...νj−1
i−1 i+1
a νj+1 ...νn
µ ...µm µ ...µ
The result is a new tensor Uν11...νj−1
i−1 i+1
νj+1 ...νn which is of order n + m − 2. We can verify
that this is indeed a tensor by examining its transformation properties:
94
Using the property of the Lorentz transformation matrix given by equation (3.57), which
λ λ
gives Λaκi Λa j = δκij , we simplify this expression to:
We can also contract different tensors by summing over matching upper and lower
indices, as demonstrated in the following example:
µν µν α
Uβλ = Tαβ Rλ .
µν µν α ′ ′ ′′ µ′ ν ′ α′′
U ′ βλ = T ′ αβ R′ λ = Λµµ′ Λν ν ′ Λαα Λβ β Λαα′′ Λλλ Tα′ β ′ Rλ′′
′ α′
Using the relation Λαα′′ Λαα = δα′′ (eq. [3.57]), we obtain:
µν ′ ′′ µ′ ν ′ α′ ′ ′′ µ′ ν ′
U ′ βλ = Λµµ′ Λν ν ′ Λββ Λλλ Tα′ β ′ Rλ′′ = Λµµ′ Λν ν ′ Λββ Λλλ Uβ ′ λ′′
µν
demonstrating that Uβλ is indeed a tensor. This approach can be generalised to the
contraction of any number of indices involving any number of mixed tensors.
If m = n, a full contraction can be performed by summing over all possible pairs of
matching indices, as illustrated in the following example:
µν αβ
Tαβ Rµν .
This operation reduces the rank of the resulting tensor to zero, yielding a scalar that is
invariant. This invariance can be demonstrated with a simple case:
T ′µν Rµν
′
= Λµα Λν β Λµγ Λν δ T αβ Rγδ .
γ δ
Using the properties Λµα Λµγ = δα and Λν β Λν δ = δβ , we obtain T ′µν Rµν
′ = T αβ R . This
αβ
demonstrates that the scalar remains invariant under the transformation.
95
5.1.4 Symmetric and antisymmetric tensors
A tensor is said to be symmetric if it remains unchanged when any two indices are in-
terchanged. Specifically, a rank-2 tensor T µν is symmetric if it satisfies the condition
T µν = T νµ .
Conversely, a tensor is said to be antisymmetric if its components change sign when
any two indices are interchanged. For a rank-2 tensor Aµν , this property is expressed as
Aµν = −Aνµ , which implies that all the diagonal elements of the tensor are zero.
Here, the integral is taken over the surface Σ enclosing the volume V, with n̂ representing
the surface normal. The surface Σ is stationary in frame (S), and the electric field E is
measured at a given time t and position (x, y, z) in (S) through the force on a test charge
at rest in (S). Experimentally, it has been shown that this integral does not depend on
the specific choice of the surface Σ enclosing the volume V, irrespective of whether Q is in
motion or not. In other words, Gauss’s law holds true even when the charge carriers are
moving.
Moreover, it has been experimentally verified that the integral in equation (5.4) remains
the same in any inertial frame. This means that Σ can be at rest in (S) or any other
inertial frame (S′ ), as long as it encloses the volume containing the charges and the force
is measured in the frame where the surface is at rest, the charge Q defined by equation (5.4)
remains unchanged. This demonstrates the invariance of charge.
Consider an ensemble of charges all moving with the same average velocity u in a frame
(S). In the rest frame (S0 ) of the particles, the charge density is ρ0 . Frame (S) moves
relative to frame (S0 ) with velocity −u. Due to length contraction in the direction of this
velocity, a unit volume in the rest frame corresponds to a volume of 1/γu in frame (S).
Additionally, the charge of each particle remains invariant. Therefore, the charge density
in (S) is ρ = γu ρ0 (following the same reasoning as in section 5.1.1).
The current density in frame (S) associated with these charges is j = ρu = γu ρ0 u. This is
the spatial component of ρ0 U µ , whre U µ = (γu c, γu u) is the 4-velocity of a particle. This
96
naturally leads to the definition of the current 4-vector, or 4-current, J µ as:
Since J µ is a 4-vector, in a frame (S′ ) moving with velocity v relative to (S), it transforms
into J ′µ = Λµν J ν . This implies that the charge density ρ′ and current density j′ in (S′ )
are given by:
ρ′ c = Λ0 ν J ν and ji′ = Λi ν J ν .
We align the x-axis with the velocity u. Therefore, frame (S) moves relative to frame
(S0 ) with velocity −ux̂, where x̂ is a unit vector in the x-direction. Now, consider frame
(S′ ) moving with velocity vx̂ relative to (S), and with velocity wx̂ relative to (S0 ). The
transformation of the 4-current from (S) to (S′ ) yields:
v vu
ρ′ c = γv ρc − jx = γv γu ρ0 c 1 − 2 , (5.6)
c c
where we have used ρ = γu ρ0 and jx = γu ρ0 u for the second equality. We can also obtain
ρ′ by noting that, since in frame (S′ ) the charges move with velocity −wx̂, the charge
density in that frame is ρ′ = γw ρ0 . We now verify that this expression for ρ′ is consistent
with equation (5.6). To express γu γv in terms of γw , we use the invariance of the scalar
product. Consider a particle at rest in frame (S0 ). Its 4-vectors in (S0 ) and (S) are
U0µ = (c, 0) and U µ = (γu c, γu ux̂), respectively. Now consider a particle at rest in (S′ ).
Its 4-vectors in (S0 ) and (S) are V0µ = (γw c, γw wx̂) and V µ = (γv c, γv vx̂), respectively.
The invariance of the scalar product, U0µ V0µ = U µ Vµ , yields −γw c2 = γu γv −c2 + uv .
97
Gauge transformation
The magnetic vector potential A is not uniquely defined, as different potentials can pro-
duce the same magnetic field. If both A and à correspond to the same field, then
B = ∇×A = ∇×Ã, which implies ∇×(à − A) = 0. The vector à − A is curl-free
and can thus be written as the gradient of a scalar function Γ, so à − A = ∇Γ. Therefore:
à = A + ∇Γ. (5.11)
∂ Ã ∂A
+ ∇ϕ̃ = + ∇ϕ,
∂t ∂t
which can also be written as:
∂
∇ϕ̃ = ∇ϕ − (∇Γ) .
∂t
This is satisfied with:
ϕ̃ = ϕ − ∂Γ/∂t. (5.12)
Such a transformation from (ϕ, A) to ϕ̃, Ã is called a gauge transformation.
Although the curl of the vector potential is specified by equation (5.10), its divergence
can be chosen freely. Specifically, ∇ · Ã = ∇ · A + ∇2 Γ. By choosing Γ appropriately,
given A, the vector potential à can be made to have any desired divergence.
In the presence of electric charges with density ρ and electric currents with density j,
Maxwell’s equations are given by:
ρ
∇·E= , (5.13)
ϵ0
∇ · B = 0, (5.14)
∂B
∇×E = − , (5.15)
∂t
∂E
∇×B = µ0 j + µ0 ϵ0 . (5.16)
∂t
With E and B given by equations (5.9) and (5.10), Maxwell’s equations (5.14) and (5.15)
are satisfied. Next, we substitute equation (5.9) into Gauss’s law (5.13):
∂ ρ
∇2 ϕ + (∇ · A) = − . (5.17)
∂t ϵ0
98
Using the identity ∇× (∇×A) = ∇ (∇ · A) − ∇2 A, the above equation can be written
as:
∂2A
2 ∂ϕ
∇ A − µ0 ϵ0 2 − ∇ ∇ · A + µ0 ϵ0 = −µ0 j. (5.18)
∂t ∂t
As previously noted, the divergence of the vector potential A can be chosen freely. We
adopt the Lorentz gauge, defined by the condition:
∂ϕ
∇ · A = −µ0 ϵ0 . (5.19)
∂t
Substituting this condition into equations (5.17) and (5.18), we can recast Maxwell’s
equations into inhomogeneous wave equations for ϕ and A :
1 ∂2ϕ ρ
∇2 ϕ − 2 2
=− , (5.20)
c ∂t ϵ0
1 ∂ 2A
∇2 A − 2 2 = −µ0 j, (5.21)
c ∂t
Given the definitions of the 4-current (5.5) and the d’Alembertian operator (3.51), we can
introduce:
µ ϕ
A = ,A , (5.22)
c
ϕ′ /c = Λ0 ν Aν and A′ i = Λi ν Aν .
Here, ϕ′ refers to the same functional form as ϕ but expressed in terms of the coordinates
in (S′ ). Similarly, A′ is expressed in terms of the coordinates in (S′ ), with its direction
also altered by the transformation.
99
The Lorentz gauge (5.19) can also be expressed in terms of the 4-potential as:
∂µ Aµ = 0. (5.24)
The invariance of the d’Alembertian operator □ ensures that equation (5.23), which en-
capsulates all of Maxwell’s equations, is in a covariant form. In other words, if Aµ and
J µ transform to A′µ and J ′µ in another inertial frame, then □A′µ = −µ0 J ′µ still holds.
This guarantees that the wave equation for the 4-potential maintains the same form across
different inertial frames, thereby preserving the consistency of physical laws in relativity.
Similarly, the invariance of the scalar product ensures that the gauge condition (5.24) is
also in a covariant form.
100
where we have used the definition (3.45) of the contravariant 4-gradient ∂ µ . We express
the derivatives using ∂ µ rather than ∂µ to maintain a consistent pattern of signs in the
expressions for the components of E and B.
The expressions above lead us to define the so-called Maxwell electromagnetic tensor:
F µν = ∂ µ Aν − ∂ ν Aµ , (5.31)
F 00 F 01 F 02 F 03 0 Ex /c Ey /c Ez /c
10
F 11 F 12 F 13
= −Ex /c −By
F 0 Bz
F µν
=
F 20
. (5.32)
F 21 F 22 F 23
−E /c −B
y z 0 B
x
F 30 F 31 F 32 F 33 −Ez /c By −Bx 0
We can also define the covariant form of the electromagnetic tensor as:
For the Minkowski metric, which has only diagonal components, this becomes:
0 −Ex /c −Ey /c −Ez /c
Ex /c 0 Bz −B y
Fµν = gµµ gνν F µν
=
E /c −B
. (5.34)
y z 0 Bx
Ez /c By −Bx 0
The contravariant tensor can be reconstructed from the covariant one by multiplying
both sides of equation ( 5.33) by g αµ g βν and summing over µ and ν, resulting in:
g αµ g βν Fµν = F αβ . (5.35)
101
5.4.2 Maxwell’s equations in terms of the electromagnetic tensor
We can now express Maxwell’s equation (5.23) using the electromagnetic tensor:
□Aµ ≡ ∂ν ∂ ν Aµ = ∂ν (∂ µ Aν − F µν ) = ∂ µ ∂ν Aν − ∂ν F µν = −∂ν F µν ,
∂ν F µν = µ0 J µ . (5.36)
Although equation (5.23) explicitly represents only the inhomogeneous Maxwell’s equa-
tions (5.13) and (5.16), the use of the 4-potential naturally ensures that the homogeneous
equations (5.14) and (5.15) are also satisfied. This is because these equations are au-
tomatically fulfilled when the electric and magnetic fields are expressed in terms of the
potentials.
When we replace equation (5.23) with equation (5.36), we still retain the inhomoge-
neous equations, but the homogeneous ones are no longer guaranteed, as the 4-potential
is no longer used. Let us first verify that equation (5.36) does indeed capture the inho-
mogeneous equations. For µ = 0, this equation gives ∂ν F 0ν = ∇ · E/c = µ0 ρc, which
corresponds to Gauss’s law ∇ · E = ρ/ϵ0 . For µ = i, with i = 1 . . . 3, we obtain:
1 ∂F i0 ∂F i1 ∂F i2 ∂F i3
+ + + = µ0 ji .
c ∂t ∂x ∂y ∂z
By substituting the components of the electromagnetic field tensor, it can be readily shown
that these three equations yield ∇×B = µ0 j + µ0 ϵ0 (∂E/∂t). Thus, equation (5.36) is
equivalent to the two inhomogeneous Maxwell’s equations (5.13) and (5.16). To obtain the
complete set of Maxwell’s equations, we need to include the two homogeneous equations.
It is straightforward to verify that these can be expressed as:
F ′µν = Λµκ Λν λ F κλ .
102
This implies that the electric and magnetic fields in (S′ ), denoted by E′ and B′ , are related
to those in (S) through the following transformations:
Ex′ Ex Ex Ex
= F ′01 = γ 2 − β2γ2 = , (5.38)
c c c c
Ey′ Ey γ
= F ′02 =γ + βγBz = (Ey − vBz ) , (5.39)
c c c
Ez′ Ez γ
= F ′03 =γ + βγBy = (Ez + vBy ) , (5.40)
c c c
Bx′ = F ′23 = Bx , (5.41)
Ez v
By′ = F ′31 = βγ + γBy = γ By + 2 Ez , (5.42)
c c
Ey v
Bz′ = F ′12 = −βγ + γBz = γ Bz − 2 Ey . (5.43)
c c
where the subscripts ∥ and ⊥ denote the components parallel and perpendicular to the
velocity v, respectively.
These field transformations are derived using the properties of the electromagnetic
tensor. They can also be understood by considering the transformation of charge densities
and currents between reference frames. Although this alternative approach is more labo-
rious, it provides deeper insights into the physical nature of how fields transform. This
detailed derivation is presented in Appendix A.
f = q (E + u×B) . (5.46)
This expression leads to the following equation for the rate of change of the particle’s
momentum:
dp
= q (E + u×B) . (5.47)
dt
Since the mass of the particle is constant, the rate of change of the particle’s relativistic
energy is given by:
dE
= f · u = qE · u. (5.48)
dt
103
As discussed in section 4.3, these equations can be encapsulated in the relativistic form of
the equation of motion:
dP µ
µ µ γu dE
F = , with F = , γu f . (5.49)
dτ c dt
F µ = qF µν Uν , (5.50)
Equation (5.50) yields a covariant form of the relativistic equation of motion for a particle
with charge q in an electromagnetic field:
dP µ
= qF µν Uν . (5.51)
dτ
104
This yields:
p
−g µν dXν dXµ = cdτ. (5.52)
Using this relation, the action for a free particle can be written as:
ˆ B ˆ B
2
p
S= −mc dτ = −mc −g µν dXν dXµ . (5.53)
A A
The worldline of the particle can be parametrised by the proper time τ , or by any other
parameter θ that monotonically increases with τ . Expressing the differentials in terms of
θ, we have:
dXµ
dXµ = dθ.
dθ
Substituting this into the action, we obtain:
ˆ θB
r
dXν dXµ ′
S= −mc −g µν dθ , (5.54)
θA dθ′ dθ′
where θA and θB are the values of θ at points A and B on the worldline.
Next, we introduce the interaction with the electromagnetic field, which is expressed
similarly in both classical and relativistic mechanics. The generalised potential accounting
for the electromagnetic force in the classical limit is given by equation (1.33), and it remains
valid here. This implies that the term to be added to the action is:
ˆ tB
−q (ϕ − ṙ · A) dt. (5.56)
tA
105
where the total time derivative is computed along the particle’s trajectory. This confirms
that gauge symmetry is satisfied.
Using the potential and dispacement 4-vectors, equation (5.56) can be rewritten as:
ˆ tB ˆ B
dXµ
qAµ dt = qAµ dXµ .
tA dt A
To make this consistent with the previous expression (5.54) for the action, we parametrise
the worldline by the same parameter θ, yielding the total action:
ˆ θB
r !
dXν dXµ dXµ
S= −mc −g µν + qAµ dθ′ . (5.57)
θA dθ′ dθ′ dθ′
r
dXν dXµ dXµ
L = −mc −g µν + qAµ . (5.58)
dθ dθ dθ
As discussed in chapter 3, the notation Xµ has been used to represent both the 4-vector
and its components. To make the analysis clearer, we now distinguish the position 4-vector
Xµ from its individual components xα and, similarly, the potential 4-vector Aµ from its
components aα , where α = 0 . . . 3.
The Lagrangian L depends on the variables xα and their derivatives dxα /dθ, similar to
the Lagrangian we discussed in chapter 1 for classical mechanics, which depended on the
generalised coordinates qi and their time derivatives. As before, applying the principle of
least action yields the Euler-Lagrange equations in the same form as equations (1.20):
d ∂L − ∂L = 0,
for α = 0 . . . 3. (5.59)
dθ dxα ∂xα
∂
dθ
Next, we verify that these equations lead to the equation of motion (5.51). Writing the
Lagrangian in terms of the variables xα and dxα /dθ was essential to obtain the Euler-
Lagrange equations (5.59). To proceed with the unpacking of these equations, we need to
revert to the proper time.
106
We have:
αν dxν + g µα dxµ dxα
∂L mc g
= r dθ dθ + qaα = m dθ + qaα , (5.60)
dxα 2 dx ν dx µ
dτ
∂ −g µν
dθ dθ dθ dθ
dxα dxα dτ
= ,
dθ dτ dθ
we obtain the first term on the left-hand side of the Euler-equations (5.59):
dxα d2 xα daα dτ
d ∂L = d
α
m + qa = m 2 +q . (5.61)
dθ dxα dθ dτ dτ dτ dθ
∂
dθ
The second term on the left-hand side of the Euler-Lagrange equations is:
d2 xα ∂aµ ∂aα
dxµ
m 2 =q − . (5.64)
dτ ∂xα ∂xµ dτ
dX α dU α dP α
d
m =m = . (5.65)
dτ dτ dτ dτ
This confirm that equation (5.64), derived from the Euler-Lagrange equations, is identical
to the covariant equation of motion (5.51).
107
5.7.3 Lagrangian for the electromagnetic field
In chapter 1, we discussed that the Lagrangian density for the electromagnetic field in a
specific inertial frame, in the absence of charges and currents, is given by:
ϵ0 2 1 2
L= E − B . (5.66)
2 2µ0
In problem set 1, it was shown that the Euler-Lagrange equations for this Lagrangian,
when expressed in terms of the potentials ϕ and A, do indeed yield Maxwell’s equations.
Since Maxwell’s equations maintain their form in all inertial frames under Lorentz trans-
formations, the Lagrangian density can be written as in equation (5.66) in any frame,
with E and B representing the electric and magnetic fields in that frame. However, this
Lagrangian is not in covariant form, meaning transformations between frames requires
transforming the fields. To express it in covariant form, we observe that:
E2
µν 2
F Fµν =2 B − 2 , (5.67)
c
which leads to:
1 µν
L=− F Fµν . (5.68)
4µ0
Since F µν is a tensor, the scalar F µν Fµν is invariant, as shown in section (5.1.3), satisfying
the requirement for the Lagrangian.
In the absence of currents and charges, the system remains invariant under both time and
spatial translations, which do not alter the spacetime interval and therefore are consistent
with Lorentz transformations. In Problem Set 1, we derived Noether’s theorem for systems
invariant under time translation. For a Lagrangian density L that depends on a field ϕ
and its derivatives, this theorem is expressed as:
∂L ∂L
∂t ∂t ϕ − L + ∂i ∂t ϕ = 0, (5.69)
∂ (∂t ϕ) ∂ (∂i ϕ)
where summation over i = 1 . . . 3 is implied. Using 4-vector notation, this can be rewritten
as:
108
∂L µ
∂µ ∂0 ϕ − δ0 L = 0, (5.70)
∂ (∂µ ϕ)
This result extends to cases where the Lagrangian depends on N fields ϕk . In such cases,
ϕ in the expression above is replaced by ϕk , with the first term in parentheses summed
over k.
For the electromagnetic field, the Lagrangian is given by equation (5.68) and depends
on the the fields Aρ , which are the components of the 4-potential, as expressed in equa-
tion (5.31). Thus, ϕk is replaced by Aρ , yielding:
∂L ρ µ
∂µ ∂ν A − δν L = 0. (5.72)
∂ (∂µ Aρ )
We now define the canonical stress-energy tensor, also referred to as the canonical energy-
momentum tensor, as:
∂L
T µλ = − ∂λ Aρ + δλµ L, (5.73)
∂ (∂µ Aρ )
where the choice of sign for T µλ will be justified later. The fact that this is a mixed
tensor will be demonstrated below. Multiplying by g νλ and summing over λ gives the
contravariant tensor:
∂L
T µν = − ∂ ν Aρ + g νµ L. (5.74)
∂ (∂µ Aρ )
∂µ T µν = 0. (5.75)
109
The term canonical reflects the construction of the tensor from the Lagrangian, while the
names ‘stress-energy’ or ‘energy-momentum tensor’ will be clarified below. The fact that
T µν is a tensor follows from the invariance of L and the covariant transformation properties
of all components in its definition, such as ∂ ν Aρ and g νµ , under Lorentz transformations.
Using Fαβ = gακ gβλ F κλ and F αβ from equation (5.31), the Lagrangian density (5.68)
can be rewritten as:
1
κ λ λ κ
α β β α
L = − gακ gβλ ∂ A − ∂ A ∂ A −∂ A , (5.76)
4µ0 | {z }| {z }
Fαβ F αβ
1
= − gβλ ∂α Aλ − gακ ∂β Aκ g αγ ∂γ Aβ − g βϵ ∂ϵ Aα . (5.77)
4µ0 | {z }| {z }
Fαβ F αβ
This yields:
∂L 1
= − gβρ F µβ − gαρ F αµ + g αµ Fαρ − g βµ Fρβ (5.78)
∂ (∂µ Aρ ) 4µ0
1
= − (gαρ F µα + g αµ Fαρ ) , (5.79)
2µ0
where we used the antisymmetry of F µα and Fµα to obtain the second equality. Writing:
∂L 1
ρ
= − g αµ Fαρ . (5.80)
∂ (∂µ A ) µ0
1 αµ
T µν = g Fαρ ∂ ν Aρ + g νµ L. (5.81)
µ0
Using F0ρ from equation (5.34) and L from equation (5.66), we obtain:
110
1 ϵ0 1 2 ϵ0 2 1 2
T 00 = − 2
E · ∂t A − E 2 + B = E + B + ϵ0 ∇ · (ϕE) , (5.82)
µ0 c 2 2µ0 2 2µ0
µν 1 αµ ρν 1 1 αµ
T =− g Fαρ F + g νµ F αβ Fαβ + g Fαρ ∂ ρ Aν . (5.83)
µ0 4 µ0
1 αµ 1 µβ 1
Qµν ≡ g Fαρ ∂ ρ Aν = F ∂ β Aν = ∂β F µβ Aν , (5.84)
µ0 µ0 µ0
where we used equation (5.36) which implies that, in vacuum, ∂β F µβ = 0. Taking the
divergence of Qµν gives:
1 1
∂µ Qµν = ∂µ ∂β F µβ Aν = ∂β ∂µ F βµ Aν , (5.85)
µ0 µ0
where the second equality follows from swapping the indices µ and β. Since F βµ = −F µβ
and partial derivatives commute, we conclude that ∂µ Qµν = 0.
111
Consequently, we can define the energy-momentum tensor as T µν = T µν + Qµν ,
with T µν and Qµν defined above. This results in:
1 1
T µν = − g αµ Fαρ F ρν + g νµ F αβ Fαβ . (5.86)
µ0 4
This tensor is symmetric, gauge-invariant, traceless and satisfies the conservation law:
∂µ T µν = 0. (5.87)
ϵ0 E2 B2
E ≡ T 00 = + , (5.88)
2 2µ0
Si 1
≡ T 0i = (E×B)i , (5.89)
c µ0 c
ϵ0 E2 B2
ij 1
−σij ≡ T = −ϵ0 Ei Ej − Bi Bj + δij + . (5.90)
µ0 2 2µ0
∂E
+ ∇ · S = 0. (5.91)
∂t
Integrating over the entire volume of the system yields the Hamiltonian of the system:
ˆ
H ≡ Edv = const. (5.92)
Equation (5.91) expresses conservation of energy, where E represents the energy density
and S is the Poynting vector, the flux of energy.
Integrating over the entire volume of the system gives the conjugate momentum of the
system: ˆ
S
P≡ dv = const. (5.94)
c2
Equation (5.93) expresses conservation of momentum, where S/c2 represents the momen-
tum density of the field and σij is a stress (force per unit area).
112
This can be further understood by noting that photons have momenta equal to their energy
divided by c. Thus, if we associate the energy density E with photons, the momentum
density becomes g = E/c. Writing momentum conservation for a volume element then
yields the continuity equation:
∂g
+ ∇ · (cg) = 0, (5.95)
∂t
since photons travel at velocity c. Comparing this equation to equation (5.91) gives:
S
= g. (5.96)
c2
where S is the surface enclosing the volume V and nj is the j-component of the unit
normal vector to the surface element ds, directed outwards. The term on the left-hand
side represents the rate of change of momentum within the volume and, consequently, is
equal to the force exerted on the volume element by its surroundings. Thus, −σij is the
i-component of the force per unit surface (stress) exerted by the electromag-
netic field within the volume on a surface element whose normal vector points
in the j-direction. This is a symmetric tensor, satisfying σij = σji .
In vacuum, the electromagnetic field exerts no forces, so the right-hand side of equa-
tion (5.97) vanishes. In the presence of charges, the momentum of the charges must be
included alongside that of the field on the left-hand side.
E cg
µν
T = . (5.98)
cg
− σij
113
114
Chapter 6
Electromagnetic waves are generally associated with a nonzero Poynting vector, which
describes the transport of energy by the wave. The power passing through the surface of
a sphere at infinity is equal to the flux of the Poynting vector through that surface. If this
flux is zero, no energy is carried away to infinity. We define radiation as the energy flux
that propagates to infinity.
When the electric and magnetic fields decay as the inverse square of the distance, the
Poynting vector falls off as the inverse fourth power of the distance. In such cases, the
power passing through a sphere at infinity is zero, and no radiation is produced.
Systems that radiate energy include accelerating charges, which we will examine in this
chapter. Accelerating charges, such as those in oscillating antennas, produce electro-
magnetic waves that carry information over long distances, which is the basis for radio,
television, and satellite communications. Electromagnetic radiation also enables us to
see the stars, communicate via satellites, and receive information from the early universe
through the cosmic microwave background, among many other phenomena.
In this chapter, we determine the electromagnetic field produced by moving charges, and
calculate the energy they radiate.
Before solving the wave equation satisfied by the potential 4-vector for general charge and
current densities, we first examine the electromagnetic field produced by a charge moving
with uniform velocity. This specific case provides valuable insight into how relativity
influences the structure and orientation of the field lines.
Consider a particle with charge q moving with uniform velocity v in an inertial frame (S).
In its rest frame (S′ ), where the charge is stationary at the origin, the electromagnetic
field at a point P with position vector r′ is given by:
115
qr′
E′ = , B′ = 0, (6.1)
4πϵ0 r′3
where primed quantities denote the values in (S′ ). In frame (S), we align the x-axis with
v. Using equations (5.44) and (5.45), the electromagnetic field in (S) can be expressed as:
v×E′
E∥ = E′∥ , E⊥ = γE′⊥ , B=γ , (6.2)
c2
where the subscripts ∥ and ⊥ indicate the components parallel and perpendicular to the
velocity v, respectively. Aligning the z-axis with v, the field in frame (S) is evaluated at
r = (x, y, z), which is related to the coordinates in frame (S′ ) by:
z ′ = γ (z − vt) , x′ = x, y ′ = y.
We find:
v2
r′2 = γ 2 (z − vt)2 + x2 + y 2 = γ 2 D2 cos2 ψ + D2 sin2 ψ = γ 2 D2 1 − 2 sin2 ψ . (6.3)
c
Here, r′ is the distance from P to the charge in (S′ ), while D is the corresponding distance
in (S). The difference between these two quantities arises from length contraction when
transforming from the rest frame (S′ ) to (S).
Equations (6.2) can be expressed as E = Ex x̂+Ey ŷ +Ez ẑ = γEx′ x̂+γEy′ ŷ +Ez′ ẑ, resulting
in:
q ′ ′ ′
qγD
E= γx + γy + z = , (6.4)
4πϵ0 r′3 4πϵ0 r′3
where coordinate transformations have been applied.
Substituting the expression for r′ from equation (6.3) into this result gives:
1 qD
E (r, t) = 3/2
. (6.5)
γ 2 1 − (v 2 /c2 ) sin2 ψ
4πϵ0 D3
The electric field points away from the instantaneous position of the charge.
This is non-intuitive because the signal received at P was emitted when the particle was
116
at a different (retarded) position. In the limit v/c ≪ 1, this reduces to the familiar field
of a stationary point charge.
The figure below illustrates the electric field lines for different values of γ (Credit: A. K.
Singal, 2020, J. Phys. Commun., 4):
v×E 1 µ0 qv×D
B= = 3/2 . (6.6)
2
c2 γ 2 1 − (v 2 /c2 ) sin ψ 4πD3
In the limit v/c ≪ 1, this simplifies to the Biot-Savart’s law, where the current is replaced
by qv.
The electric and magnetic fields given by equations (6.5) and (6.6) are solutions to
Maxwell’s equations for charge and current densities given by:
where the specific case of uniform velocity, v(t) = dr/dt = const, was examined. We
now consider more general charge and current densities, and seek solutions to Maxwell’s
equations in the form of equation (5.23), which is restated here for convenience:
1 ∂ 2 Aµ
□Aµ ≡ − + ∇2 Aµ = −µ0 J µ , (6.8)
c2 ∂t2
where the 4-potential is Aµ = (ϕ/c, A), the 4-current is J µ = (ρc, j), and the Lorentz
gauge is assumed. All functions depend on r and t.
117
This equation identifies G as the potential generated by a point charge that exists only
instantaneously at r = rc at time t = tc , effectively ‘flashed’ at that moment. This
implies G (r, rc , t, tc ) = G (r − rc , t − tc ). Since equation (6.8) is linear, its solutions can
be constructed by summing over all positions and times:
ˆ +∞ ˆ +∞
Aµ (r, t) = dtc d3 rc µ0 J µ (rc , tc ) G (r − rc , t − tc ) . (6.10)
−∞ −∞
Solving for G thus determines Aµ . We next introduce the Fourier transform of G, denoted
as G̃, which satisfies:
ˆ +∞
1
G (R, θ) = G̃ (R, ω) e−iωθ dω, (6.11)
2π −∞
´ +∞
where we used −∞ e−iωθ dω = 2πδ(θ) and defined k = ω/c. Since this holds for all t, and
thus for all θ, it follows1 :
k 2 + ∇2 G̃ (R, ω) = −δ (R) .
(6.13)
´ +∞
1
This relies on the fact that, if −∞ f (r, ω) e−iωθ dω = 0 for all θ, then f (r, ω) = 0 for all ω. To see
this, multiply the integral by eiω0 θ , where ω0 is arbitrary, and integrate over θ:
ˆ +∞ ˆ +∞ ˆ +∞
dω f (r, ω) dθ e−i(ω−ω0 )θ = dω 2πf (r, ω) δ (ω − ω0 ) = 2πf (r, ω0 ) = 0.
−∞ −∞ −∞
118
e±ikR
G̃(R, ω) = . (6.15)
4πR
ˆ +∞ ˆ +∞
1 1 δ (θ ∓ R/c)
G(R, θ) = G̃(R, ω)e−iωθ dω = e−iω(θ∓R/c) dω = . (6.16)
2π −∞ 8π 2 R −∞ 4πR
This shows that G is a sum of spherical waves, whose phase is given by φ(R, t) = θ ∓R/c =
t − tc ∓ R/c. The potential originates at R = 0 at time t = tc and propagates out-
wards. To represent a physically meaningful forward-propagating wave, it must satisfy
φ(R + λ, t + T ) = φ(R, t), where T = λ/c is the period, selecting φ(R, t) = θ − R/c.
ˆ +∞
µ µ0 J µ (rc , t − R/c)
A (r, t) = d3 rc , (6.18)
4π −∞ R
where R = |r − rc (tc )|, with rc (tc ) being the position of the charge at time tc = t −
|r − rc (tc )| /c. This expression shows that the field at point r and time t is influenced by
the charge at location rc at the earlier time t − |r − rc | /c, where |r − rc | /c is the time
required for the field to propagate from rc to r. For this reason, the potential given in
equation (6.18) is referred to as the retarded potential.
It is important to note that the retarded potential depends on the velocities of the
charges through the current density j, but not on their acceleration. Consequently, the
potential is identical for a charge moving with uniform velocity v and for an
accelerated charge that has the same velocity v at a specific instant.
119
The case of a charge moving with uniform velocity v in frame (S) was discussed in sec-
tion 6.1.1, where the electric and magnetic fields were calculated directly. Here, we focus
on the potentials instead, as these are the quantities that remain the same for both uni-
formly moving and accelerated charges. In the rest frame (S′ ) of this charge, where it is
at the (fixed) location r′c , the scalar and vector potentials are given by:
q
ϕ′ r′ = A′ r′ = 0.
, (6.19)
4πϵ0 |r′ − r′c |
Note that these are not the potentials in the instantaneous rest frame if the charge is
accelerating, as we will discuss later. Since Aµ is a 4-vector, the transformation to frame
(S) gives:
ϕ ϕ′ γq
(r, t) = γ = , (6.20)
c c 4πϵ0 c |r′ − r′c |
v ϕ′ γqv
A∥ (r, t) = γ = , (6.21)
c c 4πϵ0 c2 |r′ − r′c |
A⊥ (r, t) = 0, (6.22)
• Event 1: the charge is at r′c (a fixed location) and emits a field (or photon) at time
t′c ,
• Event 2: the observer is at r′ and receives the field (or photon) at time t′ = t′c +
|r′ − r′c | /c .
The 4-vector interval between these two events in (S′ ) is R′µ = [c (t′ − t′c ) , r′ − r′c ], and
since the interval is light-like we have (R′µ )2 = 0.
• Event 2: the observer is at r and receives the field (or photon) at time t = tc +
|r − rc (tc )| /c, noting that the charge has moved from rc by the time the observer
receives the field.
The 4-vector interval between these two events in (S) is Rµ = [c (t − tc ) , r − rc (tc )] and,
again, since the interval is light-like, we have (Rµ )2 = 0.
To calculate |r′ − r′c (tc )|, we instead calculate c (t′ − t′c ), as these two expressions are
equal. Transforming from frame (S) to frame (S′ ), we obtain:
h v i
c t′ − t′c = γ c (t − tc ) − · (r − rc (tc )) .
(6.23)
c
120
With U µ = (γc, γv) representing the 4-velocity of the charge, the right-hand side of the
expression can be identified as −U µ Rµ /c. Substituting this result into equations (6.20),
(6.21) and (6.22) gives:
ϕ 1 γcq
(r, t) = − , (6.24)
c 4πϵ0 c U µ Rµ
1 γqv
A (r, t) = − ,, (6.25)
4πϵ0 c U µ Rµ
1 qU ν
Aν (r, t) = − , (6.26)
4πϵ0 c U µ Rµ
We now rewrite the potentials in a more explicit form in terms of the coordinates. Defining
R ≡ r − rc (tc ), as introduced in section 6.1.2, equation (6.23) becomes:
v b
c t′ − t′c = γ 1 − · R
R, (6.27)
c
where we have used c (t − tc ) = R and R b is a unit vector in the direction of r − rc (tc ). The
left-hand side, c (t − tc ), is equal to |r′ − r′c (tc )|. Substituting this into equations (6.20),
′ ′
q 1 q β
ϕ (r, t) = , A (r, t) = , (6.28)
4πϵ0 1−β·R
b R 4πϵ0 c 1 − β · R
b R
The Liénard-Wiechert potentials (6.26) and (6.28) were derived for a charge with
uniform velocity v. However, as noted earlier, they remain valid for an accelerated charge
that has the velocity v at time tc , even though the potentials given by equation (6.19),
which we used as the starting point for the calculation and represent the potentials in the
rest frame of a charge moving with uniform velocity, are not equal to the potentials of the
accelerated charge in its instantaneous rest frame, as will be seen in the next section.
121
∂A
E = −∇ϕ − , B = ∇×A, (6.29)
∂t
where the derivatives are with respect to the coordinates r and t in frame (S). Since tc
depends on both r and t, the calculation is quite laborious. However, having already
derived the fields for a uniformly moving charge in section 6.1.1, we now calculate only
the terms involving the acceleration. The total fields are obtained by adding the two
contributions. The component of the electric field that depends on acceleration is denoted
Erad , known as the radiation field, as it contributes to the radiated power, as will be
shown in the next section.
Using the potentials derived earlier, we have:
q R · ∂i β
b ∂t βi β i R · ∂t β
b
Erad,i = − 2 − − 2
4πϵ0 R 1−β·R b 1−β·R b c 1−β·R b c
h i
−q cR b · ∂i β + ∂t βi − ∂t βi β · R
b + βi R b · ∂t β
= 2 (6.30)
4πϵ0 1 − β · R b cR
The position of the charge at time tc is rc , with velocity and acceleration given by:
drc d2 rc
v(tc ) = βc = , a(tc ) = .
dtc dt2c
Writing:
R |r − rc (tc )| 1h i1/2
tc = t − =t− =t− (x − xc )2 + (y − yc )2 + (z − zc )2 , (6.31)
c c c
with r = (x, y, z) and rc = (xc , yc , zc ), we calculate:
Ri β·R −R
bi
∂i tc = − + ∂i tc , which yields ∂i tc = . (6.32)
Rc R 1−β·R
b c
Similarly,
β·R 1
∂t tc = 1 + ∂t tc , which yields ∂t tc = . (6.33)
R 1−β·R
b
To this radiation field, we add the component depending only on velocity, given by
equation (6.5). This near-zone field, denoted as Enf , varies as 1/R2 and dominates over
122
the radiation field close to the charge. Rewriting D and ψ in equation (6.5) in terms of
R and β, remembering that rc (t) − rc (tc ) = (t − tc ) v = vR/c for a charge in uniform
motion, yields the total electric field E = Erad + Enf :
h i
R − β R×
q b b Rb − β ×β̇
E (r, t) = + . (6.35)
3
4πϵ0 1 − β · R
b γ 2 R2 cR
Here β ≡ v(tc )/c and R = |r − rc (tc )|, with tc = t − |r − rc (tc )| /c. The magnetic field is:
R×E
b
B (r, t) = . (6.36)
c
In the instantaneous rest frame of the particle, where β = 0 at that instant, the electric
field does not reduce to Coulomb’s law. In other words, if a charge is moving with velocity
v(tc ), the potentials in the instantaneous rest frame at a given time t0 are not
the same as those of a charge moving with uniform velocity v0 = v(t0 ) in its
rest frame.
In contrast, equations (6.35) and (6.36) show that when a charge is accelerating, both E
and B include terms proportional to 1/R. This implies that an accelerated charge emits
radiation. We now proceed to calculate the radiated power.
Larmor’s formula
As mentioned earlier, at distances far from the charge, the radiation field dominates over
the near-zone field contributions.
Specifically, from equation (6.35), this condition is sat-
isfied when R ≫ c/ γ 2 β̇ . This regime is the focus of our discussion here.
123
Retaining only the radiation field and considering the limit β ≪ 1, equation (6.34) sim-
plifies to:
q R×
b R×
b β̇
Erad (r, t) = . (6.37)
4πϵ0 c R
Using spherical coordinates (R, θ, φ), with the z-axis aligned along the acceleration of the
particle at time t, and denoting unit vectors with a hat, this can be rewritten as:
qa sin θ
Erad (r, t) = θ̂, (6.38)
4πϵ0 c2 R
where a = cβ̇ = v̇ is the acceleration. The corresponding magnetic radiation field, given
by Brad = R×E
b rad /c (eq. [6.36]) becomes:
qa sin θ
Brad (r, t) = φ̂, (6.39)
4πϵ0 c3 R
yielding the Poynting vector:
1 1 2 b q 2 a2 sin2 θ b
S (r, t) = Erad ×Brad = Erad R = R. (6.40)
µ0 µ0 c 16π 2 ϵ0 c3 R2
Here, R = |r − rc (tc )|, where rc (tc ) is the position of the charge at the retarded time tc
and tc = t − |r − rc (tc )| /c.
The radiation field can also be derived by analysing the motion of the charge and
determining how far the information about its velocity, influenced by the acceleration,
propagates within a finite amount of time. As this approach offers additional insight into
the structure of the radiation field, we have included it in Appendix A.
We now apply energy conservation to relate the flux of the Poynting vector to the
energy emitted by the charge.
124
Now, consider an observer positioned at the surface element dΣ. Energy conservation, as
expressed by equation (5.91) and integrated over the spherical cone, gives S· RdΣb = dU/dt,
where dU represents the energy emitted into the spherical cone during the time interval
dt, as measured by the observer. For a non-relativistic charge, equation (6.33) yields
∂t tc ≃ 1, which implies that dt is approximately equal to the time interval dtc during
which the radiation is emitted. This can also be seen by noting that rc (tc ) ≃ rc (t) for non-
relativistic motion, and therefore tc ≃ t − |r − rc (t)| /c, which leads to dtc ≃ dt. Thus, the
power received by the observer in the infinitesimal solid angle dΩ, given by dP = dU/dt,
is equal to the power radiated by the charge into that solid angle. This is expressed as:
q 2 a2 sin2 θ dΩ
dP = . (6.41)
4πϵ0 c3 4π
To calculate the total power Prad radiated through the sphere of radius R, we integrate
dP over the surface area dΣ = R2 sin θdθdφ, or equivalently over the solid angle dΩ =
sin θdθdφ:
ˆ π
q 2 a2
Prad = sin3 θdθ.
8πϵ0 c3 0
q 2 a2
Prad = , (6.42)
6πϵ0 c3
Notably, Prad is independent of R and t. This means that the power crossing a spherical
surface of radius R at time t is equal to the power crossing another spherical surface of a
different radius at a different time. This result reflects the principle of energy conservation.
125
Before extending this result to relativistic charges, we first apply it to the cases of an
oscillating dipole and an antenna.
The radiated power is therefore given by equation (6.42), with the acceleration a =
−dω 2 sin ωt. Averaging over an oscillation period, this results in:
ω 4 p20
⟨Prad ⟩ = . (6.43)
12πϵ0 c3
Here, it is assumed that the motion is non-relativistic,meaning v ∼ dω ≪ c, which implies
d ≪ λ, where λ = 2πc/ω is the wavelength of the electric field, which oscillates at the
same frequency as the charge’s motion.
When an electromagnetic wave interacts with a charged particle, it exerts a force that
causes the particle to oscillate like a dipole. Consequently, the particle radiates energy in
directions other than that of the incident wave, leading to the scattering of the elec-
tromagnetic wave. If the particle’s motion is non-relativistic, the emitted radiation has
the same frequency as the incident wave. This phenomenon is known as is Thomson
scattering.
In the atmosphere, electrons bound to atomic nuclei are set into oscillations by the the
electric field of electromagnetic waves originating from the Sun. Since the size of atoms
is much smaller than the wavelength of visible light, the radiation can be analysed using
126
the formula derived earlier for an oscillating dipole. The scattering of light by particles
significantly smaller than the wavelength of the radiation is known as Rayleigh’s scat-
tering. In this case, the radiated power is proportional to ω 4 ∝ 1/λ4 , meaning that blue
light is scattered approximately four times more effectively than red light.
As a result, light observed away from the direction of the incident beam is enriched in
high-frequency (blue) components compared to the spectral distribution of the incident
sunlight, while the transmitted beam becomes progressively redder. This explains why
the sky appears blue.
At sunset, light travels through up to 40 times more air in the atmosphere than at midday,
leading to significant scattering of blue light out of the observer’s line of sight. Conse-
quently, the remaining light is dominated by red wavelengths, making sunsets appear red.
Sunrises are generally less red than sunsets because daytime turbulence stirs more parti-
cles into the atmosphere, enhancing scattering and making the effect more pronounced at
sunset.
Let ρ(z, t) denote the volume charge density within the antenna. The charge conservation
equation gives ∂ρ/∂t = −∇ · j, where j = (I/s)ẑ is the current density and s is the cross-
sectional area of the wire. This yields ∂λ/∂t = −∂I/∂z, where λ(z, t) = ρ(z, t)s is the
charge per unit length in the antenna. Substituting the expression for I, we find:
2I0
λ(z, t) = ± sin ωt,
ωd
where the upper sign applies for z > 0 and the lower sign for z < 0.
127
The charge contained in an infinitesimal length dz at position z is dq = λ(z, t)dz. The
dipole moment of this charge distribution is:
ˆ d/2 ˆ 0 ˆ d/2
!
2I0 I0 d
p(t) = ẑ zdq = sin ωt − zdz + zdz ẑ = sin ωt ẑ.
−d/2 ωd −d/2 0 2ω
I02
⟨Prad ⟩ = (kd)2 . (6.44)
48πϵ0 c
Note that antennas designed to emit radio waves typically have lengths comparable to the
wavelength of the electromagnetic field they produce. As a result, the radiated power for
such waves cannot be calculated using the above formulae.
In principle, the radiated power Prad can be calculated directly by constructing the Poynt-
ing vector from the electric and magnetic fields given by equations (6.35) and (6.36),
retaining only the radiation components. However, it is simpler to extend the result ob-
tained for a non-relativistic charge, noting that Prad is Lorentz invariant.
To demonstrate this, let us consider the charge in its instantaneous rest frame (St )
at time t. Denote dU ′ the energy radiated by the particle during a time interval dt′ , as
measured in (St ). During this interval, the charge’s velocity increases by dv ′ . Since dt′ is
infinitesimally small, dv ′ ≪ c, meaning the charge is non-relativistic in (St ). Consequently,
the radiated power, dU ′ /dt′ , is given by equation (6.42).
In frame (S), where the observer is at rest, let dt and dU denote the corresponding time
interval and emitted energy. Since dt′ represents the time interval during which the parti-
cle emits energy, the relationship between dt and dt′ is cdt = γ (cdt′ + (v/c) · ds′ ) , where
γ is the Lorentz factor associated with the velocity v(t) and ds′ is the displacement of the
particle in (St ) during dt′ . Since the particle’s velocity increases from zero to dv ′ during
dt′ , it follows that ds′ < dv ′ dt′ . To first order in dt′ , we then have dt = γdt′ .
To calculate dU , consider the field as being composed of photons. A photon emitted dur-
ing dt′ in the direction θ′ has a 4-momentum (E ′ /c, p′ (θ′ )), with E ′ /c = p′ (θ′ ). In frame
128
(S), this photon’s energy transforms to γ (E ′ /c + (v/c) · p′ (θ′ )). Because the power radi-
ated in (St ) is proportional to sin2 θ′ , for every photon emitted in the direction θ′ , there
is another photon emitted in the opposite direction, −θ′ , with p′ (−θ′ ) = −p′ (θ′ ). The
energy of these two photons in (S) is then γ (2E ′ /c). Summing over all photons emitted
during dt′ , the total emitted energy is dU = γdU ′ .
Consequently, dU/dt = dU ′ /dt′ , confirming that the radiated power Prad is Lorentz
invariant. Since Prad is given by equation (6.42) in the instantaneous rest
frame, it retains this value in any frame, regardless of whether the charge is
relativistic or not.
In the instantaneous rest frame of the charge, the 4-acceleration is A′µ = (0, a′ ), where a′
is the proper acceleration. In this frame, the radiated power is given by equation (6.42)
as Prad = q 2 A′µ A′µ / 6πϵ0 c3 .
Since A′µ A′µ is a Lorentz invariant, this expression is covariant. Thus, in any frame where
the 4-acceleration of the charge is Aµ , the radiated power can be expressed in covariant
form as:
q2
Prad = Aµ Aµ . (6.45)
6πϵ0 c3
v · a 2 ·a 2 v · a 2 a 2 v 2
2v
µ 4 4 6 2
A Aµ = γ −γ + γ v+a =γ a + − 2 . (6.46)
c c2 c c
where a = dv/dtc is the spatial acceleration. The last two terms in the brackets combine
to yield − (v×a)2 /c2 , so the radiated power Prad can be expressed as:
" 2 #
q2
v×a
Prad = γ6 a2 − . (6.47)
6πϵ0 c3 c
129
130
Appendix A
131
The information that the particle begins to decelerate propagates outwards from x = 0
at the speed c. By time t = T , this information has reached a distance cT from x = 0.
Consequently, only observers within a sphere of radius cT , centered at x = 0, are aware
that the particle has decelerated. Observers beyond this sphere (region I in the figure) still
perceive the electric field of a charge moving at a constant velocity v0 . If the charge had
not decelerated, it would have reached x = v0 T at time t = T . Observers in region I thus
observe the electric field produced by a charge located at x = v0 T , moving with velocity
v0 . This is because the particle had been travelling at constant velocity since t = −∞,
resulting in the whole space filled with field lines that follow the motion of the charge.
As shown in section 6.1.1, the electric field of a charge moving at constant velocity points
radially away from the instantaneous position of the charge. For instance, at point C in
the figure, the electric field is directed along the line segment CD.
The information that the particle has come to rest propagates outwards from x = v0 τ /2
at the speed c. By time t = T , this information has travelled a distance c(T − τ ) from
x = v0 τ /2. Consequently, for observers within a sphere of radius c(T − τ ) centered at
x = v0 τ /2 (region II in the figure), the electric field corresponds to that of a stationary
charge located at x = v0 τ /2. At point B in the figure, the electric field points along the
line segment AB and is given by EB = q/[4πϵ0 c2 (T − τ )2 ]. Point B is chosen such that
the angle θ between AB and the x–axis is the same as that between CD and the x–axis.
Since v0 ≪ c and τ ≪ T , the position x = v0 τ /2 ≪ cT . This implies that the separation
between the points at x = 0 and at x = v0 τ /2 is negligible compared to the other distances
in the problem. Thus, EB ≃ q/(4πϵ0 R2 ), where R = cT .
We now proceed to calculate the electric field in the transition region, which corresponds
to the spherical shell of thickness cτ .
Consequently, since the two cones (C1 ) and (C2 ) have the same opening angle, they carry
the same amount of flux originating from the charge q.
132
This observation implies that the flux of
E through the surface (Σ), bounded by
the bases of (C1 ) and (C2 ), is zero. For
this to be true, the line segments AB and
CD must belong to the same field line,
connected by the segment BC.
From the geometry of the problem, it follows that Eθ /Er = v0 T sin θ/(cτ ), where Er
and Eθ are the radial and transverse components of the electric field within the shell,
respectively. Since the perpendicular component of E must remain continuous, we have
Er = EB , which yields:
qv0 sin θ
Eθ = ,
4πϵ0 c2 τ R
qa sin θ
Eθ = . (A.1)
4πϵ0 c2 R
We arrive at the important result that Eθ scales as 1/R, unlike the field of a static charge,
which scales as 1/R2 . Since Er ∝ 1/R2 , it follows that as R → ∞, or equivalently T → ∞,
E → Eθ , meaning the electric field becomes purely transverse at large distances.
Accompanying this electric field is a magnetic field given by Bφ = Eθ /c, which is perpen-
dicular to both R̂, the unit vector radial vector, and E. Since Eθ Bφ ∝ 1/R2 , the Poynting
vector integrated over the surface of a sphere of radius R remains finite, indicating that
energy is radiated away. For this reason, Eθ and Bφ are referred to as the radiation
fields, denoted by Erad and Brad . In contrast, the components of the fields that decay
faster than 1/R dominate near the charge and are known as near–zone fields.
133