0% found this document useful (0 votes)
2 views55 pages

Differential Geometry Notes

The document presents lecture notes on differential and Riemannian geometry specifically tailored for understanding General Relativity. It aims to provide a mathematical toolbox for readers, focusing on key concepts and interpretations without delving deeply into physics. The notes are designed for varying levels of understanding and include practical examples to clarify abstract concepts.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views55 pages

Differential Geometry Notes

The document presents lecture notes on differential and Riemannian geometry specifically tailored for understanding General Relativity. It aims to provide a mathematical toolbox for readers, focusing on key concepts and interpretations without delving deeply into physics. The notes are designed for varying levels of understanding and include practical examples to clarify abstract concepts.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Dierential and Riemannian

Geometry for General Relativity:


A Useful Toolbox
by an aspiring (but still amateur) geometrist

DenisWerth
International Center for Fundamental Physics, 2020-2021
Preface
We have been doing physics for several years now. Since then, physics was mostly Fourier
transforms, solving dierential equations and linear algebra. The question behind most
of the physical problems was usually How to diagonalize the Hamiltonian? This is just
because we did a lot of Quantum Physics. We were (and still are) used to the Dirac
notations i.e. the bras and the kets, operators, commutators, etc. Even when it comes
to contracting indices and using the metric tensor, Special Relativity gave us some ease.
With General Relativity, the paradigm has changed, and so the mathematical formalism.
Everything is continuous and formula are impossible to remember1 . A lot of questions are
raised. What is the tangent vector tangent to? What does the connection connect? Does
a one-form act on a function? How to physically interpret the Riemann tensor? What
is the link between the Lie derivative, Killing vectors and symmetries? Does a change of
coordinates act at the same point? If you are just like me and all these questions prevent
you from sleeping, these lecture notes are made for you.

What are these Lecture Notes? This handout presents the mathematics of Gen-
eral Relativity. It aims at oering the reader a mathematical toolbox and ease behind
the concepts of Riemannian geometry. We will introduce the useful basic concepts of
dierential and Riemannian geometry, discuss their interpretations and give useful tips to
easily apply these abstract concepts to physics. I will try to write these lecture notes so
that you can have dierent reading levels. For example, you can entirely read the handout
or just the essential at the end of each section. Several examples in violet also illustrate
the abstract concepts throughout the script. We will not prove most of the results to
keep these lecture notes relatively short. Finally, these lecture notes are mainly based on
David Tong's lectures2 and the heavy "Big Black Book"3 .

What These Lecture Notes are Not? Personally in General Relativity, I know
how to apply formula, compute with several indices, integrate, change frames, etc. But
I feel that I do not deeply understand the mathematical objects I am using. To help me
understand the mathematics of General Relativity and to make things clear in my mind,
I started to write these lecture notes. And I learned a lot! Obviously, the aim was not
to write another lecture notes on General Relativity because many already exist on the
web. Instead, I wanted to focus on mathematics.
First, I need to say that this is not a physics lecture. We will not discuss Einstein
equations, present some solutions, linearize the theory to recover gravitational waves,
etc... It means basically that we will not make a clear and deep link between geometry
and gravity. But wait... Gravity is geometry! This rather simple statement gives me the
opportunity to focus on mathematics and geometry instead of physics, hoping that physics
will become more intuitive afterward. The rest is no longer my responsibility because (i)
I am just a beginner i.e. I know pretty much nothing about General Relativity, and (ii)
we have a wonderful teacher.

Finally, I would like to thank our teachers Marios Petropoulos and Philippe Grand-
clément for making me understand that I do not fully understand General Relativity.
1 Can you write the torsion tensor in terms of components without looking at the formula?
2 David Tong, General Relativity , University of Cambridge Part III Mathematical Tripos
3 Charles W. Misner, Kip S. Thorne and John Archibald Wheeler, Gravitation

2
Contents
1 Dierential Geometry 5
1.1 Dierential Manifolds . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5
1.2 Vectors Redened into Tangent Vectors . . . . . . . . . . . . . . . . . . . . 6
1.3 Integral Curves . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8
1.4 Transformation Law for Tangent Vectors . . . . . . . . . . . . . . . . . . . 10
1.5 One-Forms . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12
1.6 Tensors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13
1.7 Operations on Tensor Fields . . . . . . . . . . . . . . . . . . . . . . . . . . 14
1.8 Commutators and Lie Derivative . . . . . . . . . . . . . . . . . . . . . . . 16
2 Riemannian Geometry 19
2.1 Abstract and Component Geometry . . . . . . . . . . . . . . . . . . . . . . 19
2.2 Metric . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20
2.3 Covariant Derivative and Connection . . . . . . . . . . . . . . . . . . . . . 22
2.3.1 Denition of the Connection . . . . . . . . . . . . . . . . . . . . . . 23
2.3.2 Levi-Civita Connection and Christoel Symbols . . . . . . . . . . . 25
2.3.3 Useful Properties of the Levi-Civita Connection . . . . . . . . . . . 26
2.4 Torsion and Curvature . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28
2.4.1 Abstract Denition and Components . . . . . . . . . . . . . . . . . 28
2.4.2 Identities and Properties of the Riemann Tensor . . . . . . . . . . . 31
2.5 Ricci and Einstein Tensors . . . . . . . . . . . . . . . . . . . . . . . . . . . 33
2.6 Parallel Transport . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 34
2.7 Geodesics and auto-parallels . . . . . . . . . . . . . . . . . . . . . . . . . . 36
2.7.1 Ane and Non-ane Parametrisations . . . . . . . . . . . . . . . . 37
2.7.2 Geodesics as Extremal Line Elements . . . . . . . . . . . . . . . . . 38
2.7.3 Computing Christoel Symbols from the Action . . . . . . . . . . . 40
3 Advanced Topics 43
3.1 Symmetries . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 43
3.1.1 Killing Vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 43
3.1.2 Conserved Charges . . . . . . . . . . . . . . . . . . . . . . . . . . . 44
3.1.3 Useful Identity Relating Curvature and Killing Vectors . . . . . . . 46
3.1.4 Maximal Symmetry and Constant Curvature . . . . . . . . . . . . . 47
3.2 Variational Calculus . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 48
3.2.1 Covariant Volume Element . . . . . . . . . . . . . . . . . . . . . . . 48
3.2.2 Divergence Theorem . . . . . . . . . . . . . . . . . . . . . . . . . . 49
3.2.3 Einstein-Hilbert Action . . . . . . . . . . . . . . . . . . . . . . . . . 51
3.2.4 Action Variation . . . . . . . . . . . . . . . . . . . . . . . . . . . . 52
The Essential in a Scheme 55

3
4
Chapter 1
Dierential Geometry
The aim of this section is to fully understand the objects we manipulate in General
Relativity. Most of the denitions are usually very formal and one loses intuition. Here,
we give a meaningful sense of how dierent objects should be seen or at least imagined.
Even if our discussion of dierential geometry is not particularly rigorous, we will be
careful about building up the mathematical objects in the right logical order.

1.1 Dierential Manifolds


The rst basic notion that we introduce is the concept of manifold. As most physicists
are satised with a more fuzzy and intuitive denition, we will not discuss some elaborate
framework. Instead, we should think of a manifold as a curved D-dimensional space so
that if you zoom in to any patch, the manifold looks like RD . Viewed more globally, the
manifold may have interesting curvature or topology.

Figure 1.1. Example of a dierential manifold in two dimensions which locally looks
like R2 .

If one wants to describe physics with manifolds, one needs to make manifolds dif-
ferential. In simple words, a dierential manifold M means that we can continuously

5
(smoothly) go from one point p ∈ M of the manifold to another point q ∈ M, and that
M locally looks like RD . An illustration is provided above.

Example. Some simple examples in mathematics include Euclidean space RD , the sphere SD and the
torus TD = S1 × ... × S1 . In statistical physics, the phase space of N particles is a 6N -dimensional
manifold.

The advantage of locally identify a dierential manifold M to RD is that we can


now import our knowledge of how to do maths on RD . For example, we now how to
parametrize a curve or to dierentiate functions on RD .

A dierential manifold is a smooth curved D-dimensional space that is locally Eu-


clidean.

1.2 Vectors Redened into Tangent Vectors


We know what vectors are from undergraduate physics, and we know what dierential
operators are. But we are not used to equating the two. Here, we give the tools to un-
derstand in what sense a vector should now be seen as a tangent vector i.e. a dierential
operator.

In classical mechanics working in a at space, we describe the position of a particle as


a vector x ∈ R3 dened as an "arrow" drawn from some origin to a point. This notion
does not generalize to other manifolds. For example, a line connecting two points on a
sphere is not a vector. In general, there is no way to think of a point p ∈ M as a vector.

To rst understand tangent vectors at a given point p ∈ M, we would like to give


a coordinate independent (formal) denition of dierentiation, and then describe it us-
ing coordinates. Let us denote C ∞ (M) the set of all smooth functions over a manifold M.

Denition. A tangent vector Xp is an object that dierentiates functions at a point


p ∈ M. Specically, Xp : C ∞ (M) → R satisfying

(i) Xp (f + g) = Xp (f ) + Xp (g) for all f, g ∈ C ∞ (M).

(ii) Xp (f ) = 0 when f is a constant function.

(iii) Xp (f g) = f (p)Xp (g) + Xp (f )g(p) for all f, g ∈ C ∞ (M).

The last property is of course the Leibniz rule. Intuitively, tangent vectors tell us how
things change in a given direction. They do this by dierentiating. It is easy to check
that the objects


∂µ = , (1.1)
p ∂xµ p

6
which act on functions obeys all the requirements of a tangent vector1 . However, they are
much more than simple tangent vectors.

Indeed, the set of all tangent vectors at a point p ∈ M forms an D-dimensional vector
space called the tangent space Tp (M). The tangent vectors ∂µ provide a basis for
p
Tp (M). This means that we can rewrite any tangent vector as

Xp = X µ ∂ µ , (1.2)
p

with X µ the components of the tangent vector in this basis. Note that to be fully rigorous
regarding the notation, one needs to write ∂µ instead of ∂µ because this object is a tangent
vector.

What is it tangent to? So far, we have not really explained where the name "tangent
vector" comes from. Consider a smooth curve in M that passes through the point p ∈ M.
We describe the curve by some coordinates xµ (λ) where λ parameterizes the curve such
that λ = 0 at p ∈ M. Before we learned any dierential geometry, we would say that the
tangent vector to the curve at λ = 0 is
dxµ (λ)
Xµ = . (1.3)
dλ λ=0

But we can take these to be the components of the tangent vector Xp which we dene
as
dxµ (λ)
Xp = ∂µ . (1.4)
dλ λ=0 λ=0

The tangent vector tells us how fast any function f ∈ C ∞ (M) changes as we move
along the curve. This gives also meaning to the term "tangent space" for Tp (M). It is
literally the space of all possible tangents to curves passing through the point p ∈ M. An
illustration of a two dimensional manifold embedded in R3 is shown below. Note that the
tangent spaces Tp (M) and Tq (M) at dierent points p 6= q are dierent. There is no way
to compare vectors in Tp (M) and vectors in Tq (M).

So far we have only dened tangent vectors at a point p. It is useful to consider objects
in which there is a choice of tangent vector for every point p ∈ M. We call objects that
vary over space elds. A vector eld X is dened to be a smooth assignment of a
tangent vector Xp to each point p ∈ M. Given a coordinate basis, we can expand any
vector eld as

X = X µ ∂µ , (1.5)
where the X µ are now smooth functions on M. This tangent vector eld X , given a
certain curve parameterized by xµ (λ), acts on a function as a dierential operator
dxµ ∂f df (λ)
X(f ) = X µ ∂µ f = µ
= . (1.6)
dλ ∂x dλ
1 Note that the index µ is a subscript rather than superscript that we use for the coordinates xµ .

7
Figure 1.2. Illustration of a tangent vector space Tp (M) at the point p ∈ M with a
tangent vector X = X ∂µ in a two dimensional manifold.
µ

It means that a tangent vector acting on a function tells us about how this function
vary along the curve.

At any point p ∈ M on the manifold, we dene a tangent vector Xp : C ∞ (M) → R


that satises properties of dierentiation (linearity, Leibniz rule, ...). The set of
objects ∂µ forms a basis of the tangent space Tp (M). Thus the tangent vector
p
can be decomposed as

Xp = X µ ∂µ . (1.7)
p

Allowing the components X µ to vary on M, we dene a vector eld X : C ∞ (M) →


C ∞ (M) as

X = X µ ∂µ . (1.8)
The action of a tangent vector eld X , given a certain curve parameterized by
xµ (λ), on a function f is

df (λ)
X(f ) = X µ ∂µ f = , (1.9)

which is a number when evaluated at a point p ∈ M.

1.3 Integral Curves


The last section was formal but essential to understand that a tangent vector is a dier-
ential object. Fortunately, there is a slightly dierent way of thinking about vector elds
on a manifold. This point of view is the one I use when I think of tangent vectors.

8
We have seen previously that the components of a tangent vector are

dxµ (λ)
Xµ = . (1.10)

It means that to compute a tangent vector Xp at a given point p ∈ M, one must


specify a curve passing through p. As the curve is dened by xµ (λ), one must dierenti-
ate each coordinate with respect to λ to recover a tangent vector.

Example. Let us consider the curve C : {θ = π/2, ϕ = λ} on the two dimensional sphere S2 that we
parameterize with θ and ϕ. This curve clearly is the equator. To explicitly compute a tangent vector at
some point, one needs to specify a basis. We choose {∂θ , ∂ϕ }. Then our vector X is written

 dθ   
0
X= dλ
dϕ = = ∂ϕ . (1.11)

1

The previous example shows that one can explicitly write a tangent vector if a specic
curve is given. Indeed, a tangent vector at a given point can be oriented in dierent ways.

Alternatively, given a vector eld X = X µ ∂µ , we can integrate the dierential equation


(1.10) subject to some initial conditions. The generated streamlines are called integral
curves. The integral lines are then simply describe by

xµ (λ) = xµ (0) + λX µ + O(λ2 ), (1.12)

that illustrates how the tangent vector components give the variation along a curve.

Example. • Consider again the sphere S2 in polar coordinates with a vector X = ∂ϕ . The integral
curves solve the equation (1.10), which are

dθ dϕ
=0 and = 1, (1.13)
dλ dλ

that has the solution θ = θ0 and ϕ = ϕ0 + λ. The integral lines are shown below.
• Consider the vector eld on R2 with Cartesian coordinates X = (1, x2 ) = ∂x + x2 ∂y . The equation for
the integral curves is now

dx dy
=1 and = x2 , (1.14)
dλ dλ

that has the solution x(λ) = x0 + λ and y(λ) = y0 + 31 (x0 + λ)3 . The associated ow lines are shown
below.

9
(a) X = X µ ∂µ = ∂ϕ . (b) X = X µ ∂µ = ∂x + x2 ∂y .

Figure 1.3. Some integral curves corresponding to the two examples given above.

Given a curve C parameterized by xµ (λ), one can compute the components of a


tangent vector X µ using dierentiationa
dxµ (λ)
.Xµ = (1.15)

Alternatively, given a tangent vector X = X µ ∂µ , one can integrate (1.15) and
obtain integral curvesb parameterized by xµ (λ). A tangent vector tells us about
variation along a curve, as shown by

xµ (λ) = xµ (0) + λX µ + O(λ2 ). (1.16)


A tangent vector must be seen as a directional derivative operator along a curve.
a Note that this is possible if one rst chooses a basis and a set of coordinates.
b Note that in practice, only very simple tangent vectors can be integrated.

1.4 Transformation Law for Tangent Vectors


An especially useful basis in the tangent space at a point2 p ∈ M is induced by any
coordinate system


eµ = = ∂µ , (1.17)
∂xµ
that must be seen as a directional derivative along a curve. A transformation from one
basis to another in the same tangent space Tp (M) at the event p ∈ M is produced by a
matrix
2 It is more appropriate to call a point an event.

10
∂ ∂ ∂xµ
e0ν = = = eµ Jνµ . (1.18)
∂x0ν ∂xµ ∂x0ν
The matrix J is the Jacobian of the transformation xµ → x0µ . One must see this
transformation as the chain rule. The components of a tangent vector must transform
by the inverse matrix

X 0ν = (J −1 )νµ X µ . (1.19)
Note that we have (J −1 )νµ Jτµ = δτν . This inverse transformation law guarantees com-
patibility between the expressions X = X 0ν e0ν and X = X µ eµ

X = X 0ν e0ν = ((J −1 )ντ X τ )(eµ Jνµ ) = X τ ((J −1 )ντ Jνµ )eµ = X τ δτµ eµ = X µ eµ . (1.20)

Example. Let us consider the two dimensional at plane3 . One can use polar coordinates xµ : {r, θ} or
Cartesian coordinates x0µ : {x = r cos θ, y = r sin θ} to dene an event, a curve, etc. The Jacobian of the
xµ → x0µ transformation is dened as
 ∂x ∂x
   
cos θ −r sin θ 1 r cos θ r sin θ
J= ∂r
∂y
∂θ
∂y , J −1 = . (1.21)
∂r ∂θ
sin θ r cos θ r − sin θ cos θ

Now if we consider a tangent vector X = (0, 1) = ∂θ in polar coordinates, one can rewrite4 it in Cartesian
coordinates

∂θ = Jθµ ∂µ = Jθx ∂x + Jθy ∂y = −r sin θ∂x + r cos θ∂y = −y∂x + x∂y . (1.22)
Our tangent vector X is written X = (−y, x) in Cartesian coordinates. Notice that in Cartesian co-
ordinates, it is not obvious that our tangent vector gives rise to integral curves rotating around the origin.

The change of basis and coordinates xµ → x0µ of a tangent vector X is made by


the use of the Jacobianab
∂xµ ∂x0µ
Jνµ = , (J −1 )µν = satisfying Jµν Jτµ = δτν . (1.23)
∂x0ν ∂xν
The basis tangent vectors eµ = ∂µ transform with J

e0ν = eµ Jνµ . (1.24)


The tangent vector coordinates X µ , depending on the choice of coordinates, trans-
form with J −1

X 0µ = (J −1 )µν X ν . (1.25)
a Note that the equations give the components of J and J −1 .
b We remember this transformation as dierentiating the new coordinates with respect to the
old coordinates.

3 We eliminate the origin because it is a singular point for polar coordinates.


4 Note that because we transform a basis tangent vector, we use the Jacobian and not its inverse.

11
1.5 One-Forms
We have seen that tangent vectors at an event p ∈ M live in the tangent space Tp (M).
Since Tp (M) is a vector space (one can choose {∂µ } as a basis), there exists a dual vec-
tor space to Tp (M) denoted Tp? (M). This space is called the cotangent space and its
elements are linear functions from Tp (M) to R.

An element ω ∈ Tp? (M) of the cotangent space is called a one-form. A one-form acts
on a tangent vector to give a number. This procedure must be seen as a scalar product5

ω ∈ Tp? (M) : Tp (M) → R


(1.26)
X → hω|Xi

The simplest example of a one-form is the dierential df of a function f . The action


of a tangent vector X on f being X(f ) = X µ ∂µ f , the action of df on X is

hdf |Xi = X(f ) = X µ ∂µ f. (1.27)

Noting that df is expressed in terms of the coordinates as df = ∂µ f dxµ , it is natural


to regard {dxµ } as a basis for Tp? (M). Moreover, we have

∂xµ
µ
hdx |∂ν i = ν
= δνµ . (1.28)
∂x

An arbitrary one-form ω can then be written

ω = ωµ dxµ , (1.29)

where ωµ are the components of ω . Given a tangent vector X , one can act on this tangent
vector with the one-form6

hω|Xi = hωµ dxµ |X ν ∂µ i = ωµ X ν hdxµ |∂µ i = ωµ X ν δνµ = ωµ X µ . (1.30)

From dx0ν = ∂x0ν


∂xµ
dxµ , we nd the coordinate transformation of the one-form compo-
nents

∂x0µ
ων0 = ωµ (1.31)
∂xν

5 We draw an analogy with quantum mechanics when we write hφ|ψi with |ψi ∈ H and hφ| ∈ H? , the
Hilbert space and its dual space respectively.
6 Note that the scalar product is dened between a tangent vector and a one-form and not between
two tangent vectors or two one-forms.

12
A one-form ω is an element of the dual vector space Tp? (M) that acts on a tangent
vector X to give a numbera hω|Xi = ωµ X µ . Any one-form ω can be decomposedb
on the basis {dxµ } as

ω = ωµ dxµ , (1.32)
where ωµ are the components of ω . The basis elements satisfy
∂xµ
µ
hdx |∂ν i = ν
= δνµ . (1.33)
∂x
a This action denes an inner product.
b Remember that one-form components have low indices whereas tangent vector components
have upper indices.

1.6 Tensors
A tensor of type (q, r) is a multilinear object which maps q elements of Tp? (M) and r
elements of Tp (M) to a number. Such an object, called a tensor of rank q + r, is written
in terms of the bases described earlier as

T = T µ1 ...µq ν1 ...νr ∂µ1 ... ∂µq dxν1 ... dxνr . (1.34)


Note that we deliberately write the string of lower indices after the upper indices, and
the tensor product between basis elements is not written. In some sense this is unneces-
sary, and we do not lose any information by writing Tνµ11...ν r . Nonetheless, we will see later
...µq

that it is a useful habit to get into.

Example. A tangent vector is a tensor of type7 (1, 0) and a one-form is a tensor of type (0, 1).

Let Xi = Xiµ ∂µ with 1 ≤ i ≤ r and ωj = ωj µ dxµ with 1 ≤ j ≤ q be i tangent vectors


and j one-forms. The action of a tensor T on them yields a number

T (ω1 , ..., ωq ; X1 , ..., Xr ) = T µ1 ...µq ν1 ...νr ωµ1 ... ωµq X ν1 ... X νr . (1.35)
As with vector elds and one-forms, we can ask how the components of a tensor
transform. We know that tangent vector basis elements ∂µ and one-form basis elements
dxµ transform as

∂ν0 = Jνµ ∂µ and dx0ν = (J −1 )νµ dxµ . (1.36)


Then, it is clear that the lower components of a tensor transform by multiplying by J
and the upper components by multiplying by J −1 . So, for example, a rank (1, 2) tensor
transforms as

∂xµ ∂x0τ ∂x0λ σ


T 0µρν = (J −1 )µσ Jρτ Jνλ T σ τ λ = T τ λ, (1.37)
∂x0σ ∂xρ ∂xν
7 Using the fact that Tp?? (M) = Tp (M).

13
for a xµ → x0µ coordinate transformation. Note that we usually do not write the prime
because there is no ambiguity regarding the change of coordinates. This transformation
is just the chain rule and acts at two dierent points.

Similarly to tangent vectors elds, we can dene tensor elds by making each com-
ponents of the tensor a function that can be smoothly evaluated on the manifold. Note
that on a manifold of dimension D, a tensor T of type (q, r) has Dq+r components. For
a tensor eld, each of these is a function over M.

In General Relativity, every equation satises the principle of general covariance


i.e. the equation preserves its form under a general coordinate transformation. As one

can dene a tensor independently of the basis that we chose, a tensorial equation is valid
in any coordinate system i.e. in any frame. Note that, in particular, a tensor is zero (at
a point) in one coordinate system if and only if the tensor is zero (at the same point) in
another coordinate system.

A tensor T of type (q, r) acts on q one-formes and r tangent vectors to give a


number

T (ω1 , ..., ωq ; X1 , ..., Xr ) = T µ1 ...µq ν1 ...νr ωµ1 ... ωµq X ν1 ... X νr . (1.38)
Under a coordinate transformation, the lower components of a tensor transform by
multiplying by J and the upper components by multiplying by J −1 . Intuitively,
this is the chain rule. For example,
∂xµ ∂x0τ ∂x0λ σ
T 0µρν = T τ λ. (1.39)
∂x0σ ∂xρ ∂xν
The principle of general covariance tells us that every physical equation should be
written in a tensorial form such that it is valid in every coordinate system.

1.7 Operations on Tensor Fields


There are a number of operations that we can do on tensor elds. The basis algebraic
operations are the following.

Linear combination. Given two (q, r) tensors A, B and two scalars α, β , their linear
combination T = αA + βB is also a (q, r) tensor. In terms of components, this reads

T µ1 ...µq ν1 ...νr = αAµ1 ...µq ν1 ...νr + βB µ1 ...µq ν1 ...νr . (1.40)


Tensor product. Given a (q, r) tensor A and a (q0 , r0 ) tensor B , their tensor product
T = A ⊗ B is a (q + q 0 , r + r0 ) tensor. In terms of components, it reads
µ1 ...µq+q0 µq+1 ...µq+q0
T ν1 ...νr+r0 = Aµ1 ...µq ν1 ...νr B νr+1 ...νr+r0 . (1.41)
Contraction. Given a (q, r) tensor, one can associate to it a (q − 1, r − 1) tensor via
contraction of one upper and one lower index

14
T µ1 ...µq ν1 ...νr → T µ1 ...µq−1
ν1 ...νr−1 = T
µ1 ...µq−1 λ
ν1 ...νr−1 λ . (1.42)
Note that contraction over dierent pairs of indices will in general give rise to dierent
tensors. For example, Tνµ
µ
and Tµν
µ
will be dierent in general.

Symmetrisation and anti-Symmetrisation. Given a (0, 2) tensor T = Tµν dxµ dxν ,


one can decompose it into its symmetric and anti-symmetric parts as

1
S(X, Y ) = [T (X, Y ) + T (Y , X)]
T (X, Y ) = S(X, Y ) + A(X, Y ) with 2 (1.43)
1
A(X, Y ) = [T (X, Y ) − T (Y , X)]
2
In index notation, this becomes
1 1
Sµν = (Tµν + Tνµ ) and Aµν = (Tµν − Tνµ ), (1.44)
2 2
which is just like taking the symmetric and anti-symmetric part of a matrix. These
operations being frequently used, we introduce some new notation. We dene
1 1
T(µν) = (Tµν + Tνµ ) and T[µν] = (Tµν − Tνµ ). (1.45)
2 2
Note that a product of a symmetric and an anti-symmetric tensor is always zero.
Indeed, let us consider S a symmetric tensor (Sµν = Sνµ ) and A an anti-symmetric
tensor (Aµν = −Aνµ ). Then,

S µν Aµν = −S νµ Aνµ = −S µν Aµν , (1.46)


where in the last step, we rewrite the indices because these are dummy indices. We nally
obtain S µν Aµν = 0.
Example. • These operations generalise to other tensors. For a (3, 1) tensor symmetric under the two
rst upper indices, one obtains
1 µνρ
T (µν)ρ σ = (T σ + T νµρ σ ). (1.47)
2
• Similarly, for a totally symmetric (0, 3) tensor, one obtains

1
T(µνλ) = (Tµνλ + Tµλν + Tνµλ + Tνλµ + Tλµν + Tλνµ ), (1.48)
3!
where we considered all permutations of three elements (hence the 3!). For a totally anti-symmetric
tensor, one obtains the same but with a minus sign when the permutation cannot be obtain from the
ordered indices with a cyclic permutation
1
T[µνλ] = (Tµνλ − Tµλν − Tνµλ + Tνλµ + Tλµν − Tλνµ ). (1.49)
3!
• If a (0, 2) tensor T is totally anti-symmetric, then Tµν = −Tνµ .

• The product of a totally symmetric tensor with a totally anti-symmetric tensor is zero. Indeed, let us
consider Aµν symmetric and Bµν anti-symmetric. Then one obtains

Aµν Bµν = −Aνµ Bνµ = −Aµν Bµν , (1.50)

15
because µ and ν are dummy variables.

How to prove that an object is a tensor? To prove that an object is a tensor, one
has to verify that the object (its components) transforms as a tensor. Schematically, if
one obtains

T 0 = (...) T + junk, (1.51)


then the object T is not a tensor. Note that when one wants to manipulate tensors, it is
usually easier to deal with its components.

A linear combination of tensors is a same type tensor.

A tensor product between a (q, r) tensor and a (q 0 , r0 ) tensor gives a (q + q 0 , r + r0 )


tensor.

Given a (q, r) tensor, one can associate to it a (q − 1, r − 1) tensor via contraction


of one upper and one lower indexa .

Given a tensor, one can dene the symmetric and the anti-symmetric part of the
tensor
1 1
T(µν) = (Tµν + Tνµ ) and T[µν] = (Tµν − Tνµ ). (1.52)
2 2
These denitions are generalized to other tensors of higher rank. A product of a
symmetric tensor with an anti-symmetric tensor is zero.

An object T that transforms as

T 0 = (...) T + junk, (1.53)


is usually not a tensor.
a Note that Tνµ
µ
and Tµν
µ
will be dierent in general.

1.8 Commutators and Lie Derivative


Given two tangent vector elds X and Y , we cannot multiply them together to get a
new tangent vector eld. Roughly speaking, this is because the product XY is a second
order dierential operator rather a rst order operator. This reveals itself in a failure of
Leibnizarity for the object XY ,

XY (f g) = X(f Y (g) + gY (f )) = X(f )Y (g) + f X(Y (g)) + gX(Y (f )) + X(g)Y (f ),


(1.54)
which is not the same as f XY (g) + gXY (f ).

16
However, one can build a new tangent vector eld by taking the commutator [X, Y ],
which acts on function f as

[X, Y ](f ) = X(Y (f )) − Y (X(f )). (1.55)


This is also known as the Lie bracket. Evaluated in a coordinate basis, the commu-
tator is given by
ν ν
 
µ ∂Y µ ∂X ∂f
[X, Y ](f ) = X µ
−Y µ
. (1.56)
∂x ∂x ∂xν
It is not dicult to check that the commutator obeys the Jacobi identity

[X, [Y , Z]] + [Y , [Z, X]] + [Z, [X, Y ]] = 0, (1.57)


which one can remember as the sum of the cyclic permutations of X , Y and Z being
zero. This ensures that the set of all tangent vector elds on a manifold M has the
mathematical structure of a Lie algebra.

So far, we know how to dierentiate a function f . This requires us to introduce a


tangent vector eld X , and the new function X(f ) can be viewed as the derivative of f
in the direction of X . But is it possible to dierentiate a tangent vector eld? Formally
speaking, we need to subtract two tangent vector living on two dierent tangent spaces.
This procedure is done using the Lie derivative denoted LX . Without any construction,
we give the denition of the Lie derivative.

The Lie derivative of a function f along X is


∂f
LX f = X(f ) = X µ . (1.58)
∂xµ
In other words, acting on functions with the Lie derivative coincides with the action
of the tangent vector eld.

The Lie derivative of a tangent vector eld Y along X is the commutator

LX Y = [X, Y ]. (1.59)
Note the analogy with quantum mechanics and for example the angular momentum.
Written in terms of components, it is
∂Y µ ν ∂X
µ
(LX Y )µ = LX Y µ = X ν ν
− Y ν
= X ν ∂ν Y µ − Y ν ∂ν X µ . (1.60)
∂x ∂x
We usually write (LX Y )µ the components of the new tangent vector as LX Y µ . Using
the Jacobi identity, one can show that

LX LY Z − LY LX Z = L[X,Y ] Z. (1.61)
The denition of the Lie derivative is extended to one-forms and tensors. One just has
to take into account all possible congurations to permute a certain number of objects
and remember that a normal sum "lower index with an upper index" creates a plus sign
and "an upper index with a lower index" creates a minus sign. For a (1, 1) tensor T , the
Lie derivative along X reads

17
LX T µν = X τ ∂τ T µν + T µτ ∂ν X τ − T τν ∂τ X µ . (1.62)
Example. • In the two dimensional plane in Cartesian coordinates, let us consider f (x, y) = x2 − sin y
and X = sin x∂y − y 2 ∂x . Then, the Lie derivative of f along X is

LX f = (sin x∂y − y 2 ∂x )(x2 − sin y) = − sin x cos y − 2xy 2 . (1.63)


• In the two dimensional plane in Cartesian coordinates, let us consider ω = y 2 dx + x2 dy and X =
∂x + xy∂y . Then, the Lie derivative components of ω along X is

LX ω0 = X τ ∂τ ω0 + ωτ ∂0 X τ = (∂x + xy∂y )y 2 + y 2 ∂x 1 + x2 ∂x xy = 2xy 2 + x2 y


(1.64)
LX ω1 = X τ ∂τ ω1 + ωτ ∂1 X τ = 2x + x3 .
Combining the two results, we can write

LX ω = (2xy 2 + x2 y)dx + (2x + x3 )dy. (1.65)

Given two tangent vector elds X and Y , one can dene the commutator [X, Y ] =
XY − Y X that can be written in terms of components in the following way
ν ν
 
µ ∂Y µ ∂X
ν
[X, Y ] = X −Y . (1.66)
∂xµ ∂xµ
The commutator satises the Jacobi identity

[X, [Y , Z]] + [Y , [Z, X]] + [Z, [X, Y ]] = 0, (1.67)


The commutator is also a tangent vector and so acts on functions to give a number.

The Lie derivative of a function f along X is


∂f
LX f = X(f ) = X µ
. (1.68)
∂xµ
The Lie derivative of a tangent vector eld Y along X is the commutator

LX Y = [X, Y ]. (1.69)
The Lie derivative satises

LX LY Z − LY LX Z = L[X,Y ] Z, (1.70)
and can be extended to one-forms and tensors. For a (1, 1) tensor T , the Lie
derivative along X readsa

LX T µν = X τ ∂τ T µν + T µτ ∂ν X τ − T τν ∂τ X µ . (1.71)
a We should remember that a ()µ ()µ gives a "+" and ()µ ()µ gives a "-".

18
Chapter 2
Riemannian Geometry

2.1 Abstract and Component Geometry


At this stage, you might wonder why things are dened in a very abstract way when in
practice we only use components to explicitly write expressions. Let us introduce two
dierent aspects of geometry which in fact are useful in dierent situations.

Abstract dierential geometry treats a tangent vector as existing in its own right,
without necessity to give its breakdown into components

X = X µ ∂µ = X 0 ∂0 + X 1 ∂1 + ..., (2.1)
just as one is accustomed nowadays in electromagnetism to treat the electric eld E ,
without having to write out its components. The abstract approach is useful when one
wants to derive results in a simple way. For example in electromagnetism, considering
the gradient operator ∇ instead of its components ∂x , ∂y , etc is the quickest and simplest
mathematical scheme one knows to derive general results.

Dierential geometry as expressed in the language of components is convenient or


necessary or both when one is dealing even at the level of elementary algebra with the
most simple applications of Relativity, from the expression of the Friedmann Universe
to the curvature around a static black hole. In practice, one uses dierential geometry
as expressed in terms of components. But never forget that the mathematics of General
Relativity can be seen as abstract objects existing in their own right. The last picture is
suitable for generalizing General Relativity for example when working in quantum gravity.

Example. Let us take the example of the Lie derivative of a tangent vector Y along X . One can either
dene it by

LX Y = [X, Y ], (2.2)
or by
ν ν
 
µ ∂Y µ ∂X
LX Y µ
= X −Y . (2.3)
∂xµ ∂xµ

However note that the Jacobi identity for the Lie derivative is easier to derive using abstract notation
than in terms of components.

19
2.2 Metric
We have yet to meet the star of the show. There is one object that we can place on a
manifold whose importance dwarfs all others, at least when it comes to understanding
gravity. This is the metric. The existence of a metric brings a whole host of new concepts
to the table which, collectively, are called Riemannian geometry.

We all know that the metric is a way to measure distances between points on a man-
ifold. It does, indeed, provide this service but it is not its initial purpose. Instead, the
metric is an inner product on each tangent vector space Tp (M).

Denition. A metric g is a (0, 2) tensor eld that is:

(i) Symmetric g(X, Y ) = g(Y , X).

(ii) Non-degenerate i.e. if for any p ∈ M, g(X, Y )|p = 0 for all Y ∈ Tp (M) then
Xp = 0.

With a choice of coordinate, one can write the metric as

g = gµν dxµ dxν . (2.4)


The object g is often written as a line element ds2 and this expression is abbreviated
as

ds2 = gµν dxµ dxν . (2.5)


Note that the metric can be seen as a matrix gµν . If so, this matrix is symmetric
gµν = gνµ with real coecients1 . Using the spectral theorem, we know that the metric
can be diagonalized i.e. we can always pick a basis eµ of Tp (M) so that the matrix gµν
is diagonal. The non-degeneracy condition ensures that none of these diagonal elements
vanish. Some are positive, some are negative. The number of negative entries is called the
signature of the metric. In classical General Relativity, one always uses (−, +, +, +, ...)
like the Minkowsky metric. There is some mild logic behind this particular convention.
When thinking about geometry, the choice (−, +, +, +, ...) is preferable as it ensures that
length distances are positive. When thinking about quantum eld theory, the choice
(+, −, −, −, ...) is preferable as it ensures that frequencies and energies are positive.

In particular, the metric is used to dene the norm of a tangent vector X

||X||2 = X µ Xµ = gµν X µ X ν . (2.6)


Here, one can import notions from Special Relativity regarding the terminology of
the norm of a vector depending on its sign. Indeed for a tangent vector X describing
1 Real coecients come from the fact that the metric represents a physical object: the gravitational
eld.

20
an object's velocity along a curve xµ (λ) parameterized by λ, depending on the sign of
X µ Xµ = gµν X µ X ν , the tangent vector can be spacelike2 , null3 or timelike4

X µ Xµ > 0 spacelike
X µ Xµ = 0 null (2.7)
X µ Xµ < 0 timelike

Figure 2.1. Illustration of the light cone and the three types of tangent vectors.

We can use the metric to determine the length of a curve. Given a parametrisation
x (λ) of curve with tangent vector X µ = dx , its length between two points a and b is
µ
µ

given by
ˆ b
(2.8)
p
`a→b = dλ −gµν X µ X ν .
a
Note that for timelike tangent vectors describing most of the physical trajectories, we
have gµν X µ X ν < 0 so the minus sign ensures that we have a positive quantity inside the
square root.

Raising and Lowering of Indices. Given a tensor, one can raise and lower indices using
the metric. These operations can of course be combined in various ways. For example,
given a tangent vector eld X µ , we can associate its dual covector eld Xµ

Xµ = gµν X ν , (2.9)
and likwise for covectors

X µ = g µν Xν , (2.10)
2 Thesevectors describe the velocity of objects travelling faster than the speed of light i.e. tachyons.
3 Thesevectors describe the velocity of massless objects travelling at the speed of light i.e. mainly
photons (maybe also gravitons?).
4 These vectors describe the velocity of massive objects/particles travelling slower than the speed of
light.

21
Note that there are dierent ways of lowering the indices, and they will in general give
rise to dierent tensors. It is therefore important to keep track of this in the notation.
For example gµν T ντ = Tµτ and not Tτ µ . This is the reason why the lower indices of a
tensor are written more to the right that the upper indices.

Finally note that this notation of raising and lowering indices with the metric is con-
sistent with denoting the inverse metric by raised indices because

g µν = g µτ g νρ gτ ρ , (2.11)
and raising one index of the metric gives the Kronecker tensor,

g µλ gλν = gνµ = δνµ . (2.12)

The metric g is a (0, 2) tensor that enables us to dene a line element

ds2 = g = gµν dxµ dxν . (2.13)


The metric is symmetric gµν = gνµ and can always be diagonalized. Its inverse is
g µν so that g µλ gλν = δνµ . It is also used to dene the norm of a tangent vector X

||X||2 = X µ Xµ = gµν X µ X ν . (2.14)


When regarding a tangent vector eld X as a particle/object's velocity along a
trajectory/curve xµ (λ) parametrized with λ, its norm is said to be spacelike, null or
timelike depending on its sign. Along such a curve, the metric is used to determine
its length between two points a and b
ˆ b
(2.15)
p
`a→b = dλ −gµν X µ X ν .
a

2.3 Covariant Derivative and Connection


We have already met one version of dierentiation. A tangent vector eld X is, at heart,
a dierential operator and provides a way to dierentiate a function f . We write this
simply as X(f ). As we saw previously, dierentiating higher tensor elds is a little more
tricky because it requires us to subtract tensor elds at dierent points. Yet tensors eval-
uated at dierent points live in dierent tangent vector spaces, and it only makes sense to
subtract these objects if we can rst nd a way to map one tangent vector space into the
other. To do this, we used the ow generated by a tangent vector X as a way to perform
this mapping, resulting in the idea of the Lie derivative LX .

There is a dierent way to take derivatives, one which ultimately will prove more
useful. The derivative is again associated to a tangent vector eld X . However, this time
we introduce a dierent object, known as connection to map the tangent vector spaces
at one point to another tangent vector space at another. The result is an object, distinct
from the Lie derivative, called the covariant derivative.

22
2.3.1 Denition of the Connection
First, we give an abstract denition of the covariant derivative, also called connection.
This denition will be useful to understand that there are many such connections and so
covariant derivatives. Ultimately, we will chose one connection, the Levi-Civita connec-
tion, to perform computations in the physical spacetime.

Denition. A connection is a map from one tangent vector space Tp (M) to another
Tq (M). We usually write this map as ∇(X, Y ) = ∇X Y and the object ∇X is called the
covariant derivative. Note that the connection takes as inputs two tangent vectors. It
satises the following properties for all tangent vectors elds X, Y and Z ,
(i) ∇X (Y + Z) = ∇X Y + ∇X Z .
(ii) ∇f X+gY Z = f ∇X Z + g∇Y Z for all functions f and g .
(iii) ∇X (f Y ) = f ∇X Y + (∇X f )Y where we dene ∇X f = X(f ).

The covariant derivative endows the manifold M with more structure. To elucidate
this, we can evaluate the connection in a basis ∂µ of the tangent vector elds. We can
always express this as

∇∂ρ ∂ν = Γµρν ∂µ , (2.16)


with Γµρν the components of the connection. It is no coincidence that these are denoted
by the same greek letter that we use for the Christoel symbols. However, for now, you
should not conate the two; we will see the relation between them when we introduce the
Levi-Civita connection.

Mapping two dierent tangent vector spaces is what allows the connection to act as a
derivative. In what follows, we will use the notation

∇ µ = ∇ ∂µ . (2.17)
This makes the covariant derivative ∇µ look similar to a partial derivative. Using the
properties of the connection, we can write a general derivative of a tangent vector eld as
∇X Y = ∇X (Y µ ∂µ ) = X(Y µ )∂µ + Y µ ∇X ∂µ
(2.18)
= X ν ∂µ Y µ ∂µ + X ν Y µ ∇ν ∂µ = X ν ∂ν Y µ + Γµνρ Y ρ ∂µ .


The fact that we can strip of the overall factor of X ν means that it makes sense to
write the components of the covariant derivative as

(∇ν X)µ = ∇ν X µ = ∂ν Y µ + Γµνρ Y ρ , (2.19)


where we use the sloppy and sometimes confusing notation (∇ν X)µ = ∇ν X µ . Note
that the covariant derivatives coincides with the Lie derivative on functions LX f =
∇X f = X(f ). It also coincides with the old-fashioned partial derivative ∇µ f = ∂µ f .
However, its action on tangent vector elds diers. In particular, the Lie derivative
LX Y = [X, Y ] depends on both X and the rst derivative of X while, as we have seen
above, the covariant derivative depends only on X . This is the property that allows us

23
to write ∇X = X ν ∇ν and think ∇µ as an operator in its own right. In contrast, there is
no way to write "LX = X µ Lµ ".

The Connection is Not a Tensor. We can see this immediately from the denition 5

∇(X, f Y ) = ∇X (f Y ) = f ∇X Y + Y X(f ). This is not linear in the second argument,


which is one of the requirement of a tensor.

However, let us illustrate this using components. We ask what the connection looks
like in a dierent basis

∂ ∂xν ∂
∂µ0 = = = Jµν ∂µ . (2.20)
∂x0µ ∂x0µ ∂xν
In the basis {∂µ0 }, the connection is written ∇∂ρ0 ∂ν0 = Γ0µ
ρν ∂µ . Substituting in the
0

transformation (2.20), we have

Γ0µ 0 λ σ λ σ λ τ σ λ
ρν ∂µ = ∇Jρσ ∂σ (Jν ∂λ ) = Jρ ∇∂σ (Jν ∂λ ) = Jρ Jν Γσλ ∂τ + Jσρ ∂λ ∂σ Jν . (2.21)

We can write this as

Γ0µ 0
(2.22)
σ λ τ σ τ
 σ λ τ σ τ
 −1 µ 0
ρν ∂µ = Jρ Jν Γσλ + Jρ ∂σ Jν ∂τ = Jρ Jν Γσλ + Jρ ∂σ Jν (J )τ ∂µ .

Stripping o the basis tangent vectors ∂µ0 , we see that the components of the connection
transform as

Γ0µ
ρν = (J
−1 µ σ λ τ
)τ Jρ Jν Γσλ + (J −1 )µτ Jρσ ∂σ Jντ . (2.23)

The rst term coincides with the transformation of a tensor. But the second term,
which is independent of Γ, instead depends on ∂J , is novel. This is the characteristic
transformation property of a connection.

Dierentiating other Tensors. One can use the Leibnizarity of the covariant derivative
to extend its action to any tensor eld. Here, we give some examples. For a one-form ω ,
we obtain

∇µ ων = ∂µ ων − Γρµν ωρ . (2.24)

For a (1, 1) tensor T , one obtains

∇µ Tρν = ∂µ Tρν + Γνµτ Tρτ − Γτµρ Tτν . (2.25)

The pattern is clear; for every upper index we get a +ΓT term while for every lower
index we get a −ΓT term.

5 Wesee here in practice the power of using abstract dierential geometry rather than in terms of
components.

24
The connection ∇(X, Y ) = ∇X Y is a map from one tangent vector space X ∈
Tp (M) to another Y ∈ Tq (M). The connection is not a tensor. Evaluated in a
basis {∂µ }, it reads

∇∂µ ∂ν = ∇µ ∂ν = Γσµν ∂σ , (2.26)


where Γµρν are the components of the connectiona . In terms of components, the
connection (or the covariant derivative) readsb

∇µ X ν = ∂µ X ν + Γνµσ X σ . (2.27)
The covariant derivative coincides with the Lie derivative for scalar functions
∇X f = X µ ∇µ f = LX f = X(f ) = X µ ∂µ f . Using the Leibnitz rule, one can
extend the action of the covariant derivative to any tensor eldsc

∇µ ων = ∂µ ων − Γρµν ωρ . (2.28)

∇µ Tρν = ∂µ Tρν + Γνµτ Tρτ − Γτµρ Tτν . (2.29)


a At this stage, the Γ's are distinct from the Christoel symbols and so one can dene many
dierent connections.
b Note that we always use the sometimes confusing notation ∇ X ν = (∇ X)ν = (∇ X)ν .
µ µ ∂µ
c For every upper index we get a +ΓT term while for every lower index we get a −ΓT term.

2.3.2 Levi-Civita Connection and Christoel Symbols


So far, note that the connection ∇ has been dened without the use of the metric and that
one can dene many dierent connections, depending on the choice for the Γ's. However,
something nice happens if we have both a connection and a metric. This nice property is
called the fundamental theorem of Riemannian geometry.

Theorem. There exists a unique connection that is compatible with the metric g in the
sense that

∇X g = 0 for all tangent vector elds X. (2.30)


This connection is said to be metric compatible. In other words, a connection is
metric compatible if the covariant derivative of the metric with respect to this connection
is everywhere zero. We can demonstrate both the existence and uniqueness of a metric
compatible connection by deriving a manifestly unique expression for the connection co-
ecients in terms of the metric. To accomplish this, we expand out the equation of the
metric compatibility for the three dierent permutations6 of the indices
λ λ

∇σ gµν = ∂σ gµν − Γσµ gλν − Γσν gµλ = 0

∇µ gνσ = ∂µ gνσ − Γλµν gλσ − Γλµσ gνλ = 0 (2.31)

∇ν gσµ = ∂ν gσν − Γλνσ gλµ − Γλνµ gσλ = 0.

6 The metric is symmetric so we have three permutations instead of six.

25
We subtract the second and third of these from the rst, then multiply by the inverse
of the metric to obtain7
1
Γσµν = g σρ (∂µ gνρ + ∂ν gµρ − ∂ρ gµν ) . (2.32)
2
This is one of the most important formulas in this subject; commit it to memory.
This connection we have derived from the metric is the one on which conventional Gen-
eral Relativity is based. It is known as the Levi-Civita connection. The associated
connection coecients are called Christoel symbols. We then found a link between
two independent objects: the metric and the connection. Both concepts can be related if
one assumes a particular connection: the Levi-Civita connection.

The fundamental theorem of Riemannian geometry states the existence of a unique


connection that is compatiblea with the metric in the sense that

∇σ gµν = 0. (2.33)
This connection is called the Levi-Civita connection and its components are the
Christoel symbols
1
Γσµν = g σρ (∂µ gνρ + ∂ν gµρ − ∂ρ gµν ) . (2.34)
2
The metric and the connection, two a priori independent objects, can be linked by
choosing a specic connection: the Levi-Civita one, on which conventional General
Relativity is based.
a Note that we need also the torsion-free property to dene a unique metric compatible connec-
tion.

2.3.3 Useful Properties of the Levi-Civita Connection


After dening the Levi-Civita connection, one can rewrite the denition of the Lie deriva-
tive by using the covariant derivative. Indeed, the Lie derivative of Y along X is
LX Y µ = X ν ∂ν Y µ − Y ν ∂ν X µ
= X ν ∇ν Y µ − Y ν ∇ν X µ − (Γµνσ X ν Y σ − Γµνσ X σ Y ν ) (2.35)
= X ν ∇ν Y µ − Y ν ∇ν X µ ,
because using the Levi-Civita connection8 , we have Γσµν = Γσνµ . This shows that the Lie
derivative LX Y µ is also a tangent vector eld that transforms as a tensor. Note, however,
that the Lie derivative, in contrast to the covariant derivative, is dened without reference
to any metric.

Using the Levi-Civita connection, it is easy to show that the inverse metric also has
zero covariant derivative
7 Notethat we obtain the result if we consider the connection coecients being symmetric with respect
to both lower indices. This is a property of a torsion-free manifold. This notion will be introduce later
when we dene the torsion tensor.
8 Meaning we also consider the torsion-free property.

26
∇σ gµν = 0 and ∇σ g µν = 0. (2.36)

This nice property allows us to raise and lower indices inside the covariant derivative
by passing the metric through the covariant derivative. For example, if Xµ is a obtained
by lowering an index of the tangent vector X µ by Xµ = gµν X ν , then

∇σ Xµ = ∇σ (gµν X ν ) = gµν ∇σ X ν , (2.37)


where we used both the Leibnitz rule and the fact that ∇σ gµν = 0. For a higher order
tensor, it reads

∇σ Tνµ = g µτ ∇σ Tτ ν . (2.38)

Covariant derivatives commute on scalars. This is of course a familiar property of the


ordinary partial derivative, but it is also true for the second covariant derivatives of a
scalar and is a consequence of the symmetry of the Christoel symbols in the second and
third indices Γσµν = Γσνµ and is also known as the no torsion property of the covariant
derivative. Namely, we have

∇µ ∇ν f − ∇ν ∇µ f = ∇µ ∂ν f − ∇ν ∂µ f
(2.39)
= ∂µ ∂ν f − Γλµν ∂λ f − ∂ν ∂µ f − Γλνµ ∂λ f = 0,

because the usual partial derivatives commute. Note that the second covariant derivatives
on higher rank tensors do not commute. We will come back to this in our discussion of
the curvature tensor later on.

Another useful property is when the covariant derivative lies in between a dened
constant norm. Let us consider a tangent vector X with constant norm9 Xµ X µ = .
Then,

X µ ∇ν Xµ = X µ ∂ν Xµ − Γσνµ X µ Xσ
= ∂ν (Xµ X µ ) − Xµ ∂ν X µ − Γσνµ X µ Xσ
(2.40)
= −Xµ ∂ν X µ − Xσ (∇ν X σ − ∂ν X σ )
= −Xµ ∇ν X µ .

Using the fact that one can raise and lower indices inside the covariant derivative with
the metric and using the previous result, one obtains

X µ ∇ν Xµ = −gµα g µβ X α ∇ν Xβ = −δαβ X α ∇ν Xβ = −X β ∇ν Xβ , (2.41)


which leads to X µ ∇ν Xµ = 0.

9 Forexample, this is the case for velocity tangent vectors parametrized with the proper time. In this
case we have  = 0, ±1 depending on whether the object/particle is massive or not.

27
The Lie derivative can be expressed in term of the covariant derivativea

LX Y µ = X ν ∇ ν Y µ − Y ν ∇ ν X µ . (2.42)
One can raise and lower indices inside the covariant derivative by passing the metric
through the covariant derivativeb . For example,
∇σ Xµ = gµν ∇σ X ν
(2.43)
∇σ Tνµ = g µτ ∇σ Tτ ν .
Covariant derivatives commutec on scalars f

∇µ ∇ν f − ∇ν ∇µ f = 0, (2.44)
but it is not the case for higher rank tensors. For tangent vectors X with constant
norm Xµ X µ = , one has

X µ ∇ν Xµ = 0. (2.45)
a This property comes from Γσµν = Γσνµ .
b This property comes from ∇σ gµν = 0 and ∇σ g µν = 0.
c This property comes from the fact that usual partial derivatives commute.

2.4 Torsion and Curvature


We saw previously that one can dene a connection. In practice, we use the Levi-Civita
connection whose components can be expressed in terms of the metric. However, the
dened connection is not a tensor. Fortunately, one can use the connection to construct
two tensors.

2.4.1 Abstract Denition and Components


The rst tensor that one can construct is the torsion. It is a (1, 2) tensor T and so takes
as inputs a one-form ω and two tangent vectors X and Y . It reads

T (ω; X, Y ) = ω(∇X Y − ∇Y X − [X, Y ]). (2.46)


The second tensor that one can construct is the Riemann tensor, also called cur-
vature. It is a (1, 3) tensor R and so takes as inputs a one-form ω and three tangent
vectors X, Y and Z . It reads

R(ω; X, Y , Z) = ω(∇X ∇Y Z − ∇Y ∇X Z − ∇[X,Y ] Z). (2.47)


Alternatively, one could think of torsion T as a map between two tangent vector spaces

T (X, Y ) = ∇X Y − ∇Y X − [X, Y ]. (2.48)


Similarly, the Riemann tensor R can be viewed as a map between two tangent vector
spaces that gives a dierential operator acting on a tangent vector space

28
R(X, Y ) = ∇X ∇Y − ∇Y ∇X − ∇[X,Y ] . (2.49)
It is not obvious from the denition that these objects are tensors. To actually demon-
strate that T and R are tensors, we need to show that they are linear in all arguments.
This is a good example of using the abstract notation to derive expressions. Linearity
in ω is straightforward. For the others, there are some small calculations to do. For
example, we must show that T (ω; f X, Y ) = f T (ω; X, Y ) where f is a scalar function.
To see this, we just run through the denition of the various objects

T (ω; f X, Y ) = ω(∇f X Y − ∇Y (f X) − [f X, Y ]). (2.50)


We then use ∇f X Y = f ∇X Y , ∇Y (f X) = f ∇Y X + Y (f )X and [f X, Y ] =
f [X, Y ] − Y (f )X . The two Y (f )X terms cancel, leaving us with

T (ω; f X, Y ) = f ω(∇X Y − ∇Y X − [X, Y ]) = f T (ω; X, Y ). (2.51)


Similarly, for the Riemann tensor we have

R(ω; f X, Y , Z) = ω(∇f X ∇Y Z − ∇Y ∇f X Z − ∇[f X,Y ] Z)


= ω(f ∇X ∇Y Z − ∇Y (f ∇X Z) − ∇f [X,Y ] Z + ∇Y (f )X Z)
= ω(f ∇X ∇Y Z − f ∇Y ∇X Z − Y (f )∇X Z − f ∇[X,Y ] Z + ∇Y (f )X Z)
= f R(ω; X, Y , Z).
(2.52)
Linearity in Y follows from the linearity in X because both tangent vectors play a
similar role. But we still need to check linearity in Z ,

R(ω; X, Y , f Z) = ω(∇X ∇Y (f Z) − ∇Y ∇X (f Z) − ∇[X,Y ] (f Z))


= ω(∇X (f ∇Y Z + Y (f )Z) − ∇Y (f ∇X Z + X(f )Z)
− f ∇[X,Y ] Z − [X, Y ](f )Z)
= ω(f ∇X ∇Y + X(f )∇Y Z + Y (f )∇X Z + X(Y (f ))Z (2.53)
− f ∇Y ∇X Z − Y (f )∇X Z − X(f )∇Y Z − Y (X(f ))Z
− f ∇[X,Y ] Z − [X, Y ](f )Z)
= f R(ω; X, Y , Z).

Thus, both torsion and curvature dene new tensors on our manifold.

The next step is to evaluate these tensors in a coordinate basis {∂µ } with the dual
basis {dxµ }. The components of the torsion are
ρ
Tµν = T (dxρ ; ∂µ , ∂ν )
= dxρ (∇µ ∂ν − ∂ν ∂µ − [∂µ , ∂ν ])
(2.54)
= dxρ (Γσµν ∂σ − Γσνµ ∂σ )
= Γρµν − Γρνµ ,
where we used the fact that dxρ ∂σ = δσρ and [∂µ , ∂ν ] = 0. We learn that, even though Γρµν
is not a tensor, the anti-symmetric part Γρµν − Γρνµ does form a tensor. Clearly, the torsion
is anti-symmetric in the lower two indices

29
ρ
Tµν ρ
= Tνµ . (2.55)
Connections which are symmetric in the lower indices, so Γρµν = Γρνµ have zero torsion.
Such connections are said to be torsion-free. This is the case for the Levi-Civita con-
nection10 .

The components of the Riemann tensor are given by11


σ
Rρµν = dxσ (∇µ ∇ν ∂ρ − ∇ν ∇µ ∂ρ − ∇[∂µ ,∂ν ] ∂ρ )
= dxσ (∇µ ∇ν ∂ρ − ∇ν ∇µ ∂ρ )
= dxσ (∇µ (Γλνρ ∂λ ) − ∇ν (Γλµρ ∂λ )) (2.56)
= dx ((∂µ Γλνρ )∂λ + Γλνρ Γτµλ ∂τ
σ
− (∂ν Γλµρ )∂λ − Γλµρ Γτνλ ∂τ )
= ∂µ Γσνρ − ∂ν Γσµρ + Γλνρ Γσµλ − Γλµρ Γσνλ .

Clearly, the Riemann tensor is anti-symmetric in its last two indices

σ
Rρµν σ
= −Rρνµ . (2.57)
Note that it would be quite unpleasant to have to verify the tensorial nature of this
expression by explicitly checking its behaviour under coordinate transformations.

Using the connection, one can dene the torsion, a (1, 2) tensor T , by

T (ω; X, Y ) = ω(∇X Y − ∇Y X − [X, Y ]), (2.58)


and the Riemann tensor, a (1, 3) tensor R, by

R(ω; X, Y , Z) = ω(∇X ∇Y Z − ∇Y ∇X Z − ∇[X,Y ] Z). (2.59)


Their components are obtained by evaluating T and R in a coordinate basis {∂µ }
with the dual basis {dxµ }. One obtains
ρ
Tµν = Γρµν − Γρνµ
σ
(2.60)
Rρµν = ∂µ Γσνρ − ∂ν Γσµρ + Γλνρ Γσµλ − Γλµρ Γσνλ .
Connections with vanishing torsion are called torsion-free connections. These con-
nections verify

Γρµν = Γρνµ . (2.61)


The Levi-Civita connection is torsion-free.

10 Actually, the torsion-free property is one requirement to derive the Levi-Civita connection components
in terms of the metric. The fundamental theorem of Riemannian geometry states the existence of a unique,
torsion-free, connection that is metric compatible.
11 Note the slightly counter intuitive, but standard ordering of the indices.

30
2.4.2 Identities and Properties of the Riemann Tensor
There is a closely related calculation in which both the torsion and te Riemann tensors
appear. We look at the commutator of covariant derivatives acting on tangent vector
elds. One can check that

[∇µ , ∇ν ]Z σ = ∇µ ∇ν Z σ − ∇ν ∇µ Z σ = Rρµν
σ
Z ρ − Tµν
ρ
∇ρ Z σ . (2.62)
This expression is known as the Ricci identity. Note that in practice, we will always
work with the Levi-Civita connection so that the torsion is zero. One then obtains

[∇µ , ∇ν ]Z σ = ∇µ ∇ν Z σ − ∇ν ∇µ Z σ = Rρµν
σ
Z ρ. (2.63)
It is perhaps surprising that the commutator [∇µ , ∇ν ], which appears to be a dier-
ential operator, has an action on vector elds which (in the absence of torsion, at any
rate) is a simple multiplicative transformation. The Riemann tensor measures that part
of the commutator of covariant derivatives which is proportional to the vector eld, while
the torsion tensor measures the part which is proportional to the covariant derivative of
the vector eld; the second derivative does not enter at all. One must remember that the
Riemann tensor measures the failure of the covariant derivative to commute.

The extension of the above formula to any higher order tensors follows the usual
pattern, with one Riemann tensor contracted with every index. For example,

[∇µ , ∇ν ]T αβ = ∇µ ∇ν T αβ − ∇ν ∇µ T αβ = Rσµν
α
T σβ + Rσµν
β
T ασ . (2.64)
There are also a number of symmetric properties satised by the Riemann tensor when
we use the Levi-Civita connection. We will just state but not prove the following results.

If we lower an index on the Riemann tensor and write Rσρµν = gσλ Rρµν
λ
, then the
resulting tensor obeys the following symmetry identities
Rσρµν = −Rσρνµ
Rσρµν = −Rρσµν (2.65)
Rσρµν = Rµνσρ .
The Riemann tensor also satises the following cyclic permutation relation

Rσρµν + Rσµνρ + Rσνρµ = 0 (2.66)


We can now count how many independent components the Riemann tensor really has
in dimension 4. The anti-symmetric property of the last two indices implies that the sec-
ond pair of indices can only take (3 × 4)/2 = 6 independent values. The anti-symmetric
property of the rst two indices implies the same. The symmetric property implies that
the Riemann tensor can be seen as a symmetric 6 × 6 matrix and thus has (6 × 7)/2 = 21
components. The cyclic property adds an additional constraint so that the Riemann ten-
sor has 20 independent components. Note that in dimension D, the Riemann tensor has
D2 (D2 − 1)/12 independent components.

Finally, the Riemann tensor also satises the so-called Bianchi identity

31
∇λ Rσρµν + ∇σ Rρλµν + ∇ρ Rλσµν = 0. (2.67)
Note that for a general connection there would be additional terms involving the
torsion tensor. This cyclic identity is closely related to the Jacobi identity for the covariant
derivative

[[∇µ , ∇ν ], ∇ρ ] + [[∇ν , ∇ρ ], ∇µ ] + [[∇ρ , ∇µ ], ∇ν ] = 0. (2.68)

Figure 2.2. Illustration of the Ricci identity with the Levi-Civita connection.

The Riemann tensor measures the failure the failure of the covariant derivative to
commutea

[∇µ , ∇ν ]Z σ = ∇µ ∇ν Z σ − ∇ν ∇µ Z σ = Rρµν
σ
Z ρ. (2.69)
The above formula is extended to any tensor. For example,

[∇µ , ∇ν ]T αβ = ∇µ ∇ν T αβ − ∇ν ∇µ T αβ = Rσµν
α
T σβ + Rσµν
β
T ασ . (2.70)
The Riemann tensor veries the following properties
Rσρµν = −Rσρνµ
Rσρµν = −Rρσµν
(2.71)
Rσρµν = Rµνσρ
Rσρµν + Rσµνρ + Rσνρµ = 0.
and the Bianchi identity

∇λ Rσρµν + ∇σ Rρλµν + ∇ρ Rλσµν = 0. (2.72)


In dimension D, the Riemann tensor has D2 (D2 − 1)/12 independent components.
a This formula is valid when using a torsion-free connection.

32
2.5 Ricci and Einstein Tensors
There are a number of further tensors that we can build from the Riemann tensor and that
are especially important in General Relativity. First, it is frequently useful to consider
contractions of the Riemann tensor. Even without the metric, we can form a contracted
Riemann tensor called the Ricci tensor
σ
Rµν = Rµσν . (2.73)
The Ricci tensor associated with the Levi-Civita connection, which is always used
in practice, is symmetric. It inherits its symmetry from the Riemann tensor. We write
Rµν = g σρ Rσµρν = g ρσ Rρνσµ , giving

Rµν = Rνµ . (2.74)


One can go one step further with contractions and create a function R over the man-
ifold. This is the Ricci scalar, also called scalar curvature,

R = Rµµ = g µν Rµν . (2.75)


This scalar can be seen as the trace of the Ricci tensor. The Bianchi identity has a
nice implication for the Ricci tensor. If we write the Bianchi identity

∇λ Rσρµν + ∇σ Rρλµν + ∇ρ Rλσµν = 0, (2.76)


and multiply by g µλ g ρν , one obtains

∇µ Rµσ − ∇σ R + ∇ν Rνσ = 0, (2.77)


which means that ∇µ Rµν = 21 ∇ν R. This motivates us to introduce the Einstein tensor
1
Gµν = Rµν − Rgµν , (2.78)
2
which has the property that it is covariantly constant, or divergent less, meaning

∇µ Gµν = 0. (2.79)
The Riemann tensor contains all the information about the curvature of the manifold.
However, note that vanishing Christoel symbols does not mean the manifold is at. This
is because the connection is not a tensor. But the Riemann tensor is a tensor and so if
it vanishes in one coordinate system, it must vanish in all of them. Given some horrible
coordinate system, with non-vanishing Christoel symbols, we can always compute the
corresponding Riemann tensor and thus the scalar curvature R to see if the manifold is
actually at.

The elementary and universal applicable method for computing the components of the
Riemann tensor and thus the Ricci tensor and nally the curvature starts from the metric
components in a coordinate basis, and proceeds by the following scheme:
σ µ
Rµσν Rµ
(2.80)
Γ∼∂g R∼∂Γ+ΓΓ
gµν −−−→ Γσµν −−−−−−→ Rρµν
σ
−−−→ Rµν −→ R.
We see here that if one knows the metric, one knows everything about the manifold.
Note that the Christoel symbols can be obtained in a straightforward by varying the

33
(a) Positive curvature R > 0. (b) Negative curvature R < 0.

Figure 2.3. Surfaces with dierent curvatures.

action. We will see this in another section.

One obtains the Ricci tensor by contracting the Riemann tensor


σ
Rµν = Rµσν . (2.81)
The Ricci scalar, also called scalar curvature, is dened as the trace of the Ricci
tensor

R = Rµµ = g µν Rµν . (2.82)


One can then dene the Einstein tensor
1
Gµν = Rµν − Rgµν , (2.83)
2
which is covariantly constant, or divergent less, meaning

∇µ Gµν = 0. (2.84)
Starting from the metric, one can have a complete description of the manifold by
following this scheme
σ µ
Rµσν Rµ
(2.85)
Γ∼∂g R∼∂Γ+ΓΓ
gµν −−−→ Γσµν −−−−−−→ Rρµν
σ
−−−→ Rµν −→ R.

2.6 Parallel Transport


Although we have now met a number of properties of the connection and especially the
Levi-Civita connection, we have not yet fully explained its name. How does it connect
dierent tangent vector spaces?

34
The answer is that the connection connects tangent vector spaces at two dierent
points of the the manifold by mean of a map called parallel transport. As we stressed
earlier, such a map is necessary to dene dierentiation. Note that it does not make any
sense to ask if two tangent vectors are parallel in a curved space. However, given a metric
and a curve connecting these two points, one can compare the two by dragging one along
the curve to the other using the covariant derivative.

Take a tangent vector eld X and consider some associated integral curves with co-
ordinates xµ (λ), such that
dxµ
Xµ = . (2.86)

We say that a tensor eld T is parallely transported along the dened curve if

∇X T = 0. (2.87)
To illustrate this, consider the parallel transport of a second tangent vector eld Y .
In terms of the components, the last condition reads

X µ ∇µ Y ν = X µ (∂µ Y ν + Γνµρ Y ρ ). (2.88)


If we now evaluate this on the curve thinking of X µ ∂µ Y ν = dY ν

, one obtains
dY ν
+ Γνµρ X µ Y ρ = 0. (2.89)

These are a set of coupled, ordinary dierential equations. Given an initial condition
x (λ0 ) corresponding to a point p ∈ M, these equations can be solved to nd a unique
µ

tangent vector at each point along the curve.

Parallel transport is path dependent. It depends on both the connection and the
underlying path wich, in this case, is characterised by the tangent vector eld X . To
parallely transport an object between two points, one must rst dene the parametrized
curve.
Example. On the two dimensional sphere in spherical coordinates (θ, ϕ), let us parallely transport the
vector Y = ∂ϕ = (0, 1) between the point p = (θ = α, ϕ = ϕ0 ) and q = (θ = α, ϕ = ϕ0 + δ) that is
along a curve parallel to the equator. First we need to dene a parametrized curve C . Here we take
C = {θ = α, ϕ = ϕ0 + δ λ} where λ is the parameter. We then dene the tangent vector corresponding
to this integral curve
dxµ
X = X µ ∂µ = ∂µ = 0 × ∂θ + δ∂ϕ = (0, δ), (2.90)

so that

X µ ∇µ Y ν = 0 leads to ∇ϕ Y µ = 0. (2.91)
Using the denition of the covariant derivative in terms of the components, the last equation is a set
of coupled rst order dierential equations
∂ϕ Y θ + Γθϕϕ Y ϕ = 0
(
(2.92)
∂ϕ Y ϕ + Γϕ θ
ϕθ Y = 0.

35
In practice, these equations are very hard to solve. Here, we know the Christoel symbols and the
fact that θ is constant. Solving these equations (with constant θ = α) where the constants of integration
are found with the initial tangent vector leads to
(
Y θ = sin α sin(cos α δ)
(2.93)
Y ϕ = cos(cos α δ).
Note that after a round trip (δ = 2π ), we do not recover the same tangent vector except along the
equator θ = α = π/2.

Figure 2.4. Illustration of the above example of parallel transport on a sphere.

The connection maps to tangent vector spaces at two dierent points of the manifold
by the parallel transport. A tensor T is parallely transported along an integral curve
xµ (λ) dened by the tangent vector eld X which satises X µ = dx if
µ

∇X T = 0. (2.94)
For a tangent vector eld Y , it reads in terms of components
dY ν
+ Γνµρ X µ Y ρ = 0. (2.95)

These are a set of coupled, ordinary dierential equations that can be solved exactly
only in a very few cases.

2.7 Geodesics and auto-parallels


A geodesic is a curve tangent to a tangent vector eld X that obeys

∇X X = 0. (2.96)
Along a curve xµ (λ) parametrized by λ, we can write the above equation in terms of
components

36
d2 xµ ν
µ dx dx
ρ
+ Γ νρ = 0. (2.97)
dλ2 dλ dλ
This is precisely the geodesic equation. We can characterise geodesics by the property
that their tangent vectors are parallely transported (do not change) along the curve. For
this reason geodesics are also known as auto-parallels.

Note that for the Levi-Civita connection, we have ∇X g = 0. This ensures that for
any tangent vector eld Y parallely transported along a geodesic X , we have
d
g(X, Y ) = 0, (2.98)

which tells us that Y makes the same angle with the tangent vector X along each point
of the geodesic.

A geodesic, also called auto-parallel curve, is a curve that is always tangent to its
tangent vector X , namely

∇X X = 0, (2.99)
which in terms of components reads
d2 xµ ν
µ dx dx
ρ
+ Γνρ = 0. (2.100)
dλ2 dλ dλ
These are coupled dierential equations that can be solved to nd the curve
parametrized by xµ (λ).

2.7.1 Ane and Non-ane Parametrisations


To understand the signicance of how one parametrises geodesics, observe that the geodesic
equation (2.97) is not reparametrisation invariant. Indeed, if one change of parametrisa-
tion λ → λ̃, then

dxµ dλ̃ dxµ


= , (2.101)
dλ dλ dλ̃
and therefore the geodesic equation can be written
2
d2 xµ dxν dxρ d2 λ̃ dxµ


+ Γµνρ =− . (2.102)
dλ̃2 dλ̃ dλ̃ dλ̃ dλ2 dλ̃

Thus the geodesic equation retains its form only under ane changes λ̃ = aλ + b
so that ddλλ̃2 = 0. Parameters that make the right-hand side of (2.102) vanish are called
2

ane parameters and are related to each other by ane transformations. The geodesic
equation (2.97) written in this form is said to be anely parameterised.

Conversely, if we nd a curve xµ (λ̃) that satises

37
d2 xµ dxν dxρ dxµ
+ Γµνρ = C(λ̃) , (2.103)
dλ̃2 dλ̃ dλ̃ dλ̃
for some function C(λ̃), we can deduce that this curve is the trajectory of a geodesic, but
that it is simply not parametrised by an ane parameter. Denoting f (λ̃) = λ, an ane
parameter λ is determined by




C(λ̃) = − i.e. = exp ds C(s) , (2.104)
f˙2 dλ̃
where f˙ = ddλλ̃ . So for C(λ̃) = 0, we have an ane transformation between λ and λ̃. Note
that the proper time τ is an ane parameter for massive particle trajectories.

In general, an auto-parallel curve in terms of components is such that


d2 xµ ν
µ dx dx
ρ
dxµ
+ Γ νρ = C(λ) , (2.105)
dλ2 dλ dλ dλ
where C(λ) is a function of the parameter λ. However, one can always chose a
parameter λ for which C(λ) = 0. Such parameters are called ane parameters.
Ane parameters are related to each other by an ane transformation λ → λ̃ =
aλ + b. For example, proper timea τ for massive particles is an ane parameter.
a Dened by ds2 = −c2 d2 τ .

2.7.2 Geodesics as Extremal Line Elements


If we assume the dierential manifold M to be the four dimensional spacetime, all objects
follow geodesics. In other words, objects follow paths that extremizes their line element12 .
Let us take the example of a timelike13 particle trajectory parametrized by X µ = dx .
µ

Here, X µ is associated with the particle's velocity. This is why we usually denote this
vector V or U . The elementary trajectory length is

(2.106)
p
ds = −gµν dxµ dxν .
We then introduce the following action
ˆ ˆ
(2.107)
p
S= dλ L = dλ −gµν X µ X ν .

We know that extremizing the action leads to the Euler-Lagrange equations


dxµ
 
∂L d ∂L
µ
= with Xµ = . (2.108)
∂x dλ ∂X µ dλ
Let us now apply the Euler-Lagrange to the previous dened Lagrangian L =
p
−gµν dX µ dX ν =
q
. Recalling that the metric gµν depends on xµ , one obtains
µ dxν
−gµν dx
dλ dλ

12 This is the principle of least action.


13 Such that Xµ X µ = gµν X µ X ν < 0.

38
∂L 1 ∂gµν µ ν
σ
=− X X
∂x 2L ∂xσ (2.109)
∂L 1
σ
= − gσν X ν ,
∂X L

so that, using the Leibnitz rule and the chain rule d



= X α ∂α , one has

dX ν
   
d ∂L d 1 1 ∂gσν 1
= gσν Xν
+ X α α X ν + gσν . (2.110)
dλ ∂X σ dλ L L ∂x L dλ

Writing down the Euler-Lagrange equation with all the terms and rearranging the
expressions leads to

d 2 xν
 µ ν
dxν

1 dx dx
gσν + ∂ µ gσν − ∂σ gµν = C(λ)gσν , (2.111)
dλ2 2 dλ dλ dλ

where we have dened C(λ) = L1 dLdλ


. A standard trick here is to see that the rst term
inside the parenthesis can be symmetrized

1
∂µ gσν = (∂µ gσν + ∂ν gσµ ) . (2.112)
2

Multiplying the above equation by the inverse metric g ασ and using the denition of
the Christoel symbols, one nally obtains

d 2 xα µ
α dx dx
ν
dxα
+ Γ µν = C(λ) . (2.113)
dλ2 dλ dλ dλ

We note that this is the geodesic equation using a non-ane parameter. However, if
one executes the same steps with the following Lagrangian

L = −gµν X µ X ν , (2.114)

the resulting Euler-Lagrange equations leads to an anely parametrized geodesic

d2 xα µ
α dx dx
ν
+ Γ µν = 0. (2.115)
dλ2 dλ dλ

39
The Euler-Lagrange equation associated with a Lagrangian L(xµ , dx ) is
µ

dxµ
 
∂L d ∂L
= with Xµ =
= ẋµ . (2.116)
∂xµ dλ ∂X µ dλ
q
Extremizing the element line with the Lagrangian L = −gµν dx leads to a
µ dxν
dλ dλ
non-anely parametrized geodesic
d 2 xα µ
α dx dx
ν
dxα
+ Γ µν = C(λ) , (2.117)
dλ2 dλ dλ dλ
where C(λ) = . However, one can simplify the derivation by considering
1 dL
L dλ
. Extremizing the action with this Lagrangian leads to an anely
µ dxν
L= −gµν dx
dλ dλ
parametrized geodesic
d2 xα µ
α dx dx
ν
+ Γ µν = 0. (2.118)
dλ2 dλ dλ
One can always choose an ane parameter (for example proper time for massive
particles) and apply the Euler-Lagrange equations without the square root.

2.7.3 Computing Christoel Symbols from the Action


If the answer to a problem or the result of a computation is not simple, then there is no
simple way to obtain it. But when a long computation gives a short answer, then one
looks for a better method. It is exactly the case when computing Christoel symbols from
the metric. It usually involves much wasted eort. Indeed in most cases, one computes
many Γ's that turn out to be zero. Here, we present the "geodesic Lagrangian" method
that provides an economical way to tabulate the Γs.

One normally thinks that the connection coecients Γσµν must be known before one
can write the geodesic equation. However, when one anely parametrizes the geodesic
equation, it is simple to apply the Euler-Lagrange equation and explicitly nd the geodesic
equation. Thus, one can use this fact to compute the Christoel symbols.

To do this, one must explicitly write down the Lagrangian that leads to an ane
parametrised geodesic
dxµ dxν
L = −gµν = −gµν ẋµ ẋν , (2.119)
dλ dλ
and use the Euler-Lagrange equation. One then identies the resulting equations with
the anely parametrized geodesic equations and reads out the Christoel symbols.
Example. Let us consider the two dimensional sphere parametrized with (θ, ϕ). We want to compute
the Christoel symbols without using the formula involving the metric. Instead, we write the anely
parametrized Lagrangian
 
L = −gµν ẋµ ẋν = − θ̇2 + sin2 θϕ̇2 . (2.120)
The Euler-Lagrange for x0 = θ leads to

40
∂L ∂L d ∂L
= −2 sin θ cos θ ϕ̇2 , = −2θ̇ and = −2θ̈, (2.121)
∂θ ∂ θ̇ dλ ∂ θ̇
so that

θ̈ − sin θ cos θ ϕ̇2 = 0. (2.122)


The Euler-Lagrange for x1 = ϕ leads to
∂L ∂L d ∂L
= 0, = −2 sin2 θ ϕ̇ and = −4 sin θ cos θθ̇ϕ̇ − 2 sin2 θ ϕ̈, (2.123)
∂ϕ ∂ ϕ̇ dλ ∂ ϕ̇
so that
cos θ
ϕ̈ + 2 θ̇ϕ̇ = 0. (2.124)
sin θ
These two equations enable us to read out the Christoel symbols
cos θ
Γθϕϕ = − sin θ cos θ and Γϕ θ
θϕ = Γϕθ = . (2.125)
sin θ

One can computes with less eort the connection components Γσµν by writing down
the Lagrangian
dxµ dxν
L = −gµν = −gµν ẋµ ẋν , (2.126)
dλ dλ
and using the Euler-Lagrange equation. One then identies the resulting equations
with the anely parametrized geodesic equations and reads out the Christoel
symbols.

41
42
Chapter 3
Advanced Topics

3.1 Symmetries
We all know that symmetries are very important in physics because in most cases they
simplify the problem. We also know that symmetries are very important because a sym-
metry implies that something is conserved. This is of course the N÷ther's theorem. Here,
we discuss the symmetries of the spacetime metric. Intuitively, the notion of symmetry
is clear. If you hold up a round sphere, it looks the same no matter what way you rotate
it. We want a way to state this mathematically.

3.1.1 Killing Vectors


Let us rst recall the expression of the Lie derivative along ξ of a (0, 2) tensor T written
in terms of the covariant derivative

Lξ Tµν = ξ σ ∇σ Tµν + Tµσ ∇ν ξ σ + Tσν ∇µ ξ σ . (3.1)


This formula becomes particularly simple if applied to the metric because ξ σ ∇σ gµν = 0
when using the Levi-Civita connection. After lowering the indices, one obtains

Lξ gµν = ∇µ ξν + ∇ν ξµ . (3.2)
To mathematically dene a symmetry, we need the concept of a family of integral
curves1 , also called a ow. A ow can then be identied with a tangent vector eld ξ
which points along the tangent vector to the ow at each point p ∈ M
dxµ
ξµ = . (3.3)

This ow is said to be an isometry if the metric looks the same at each point along
a given ow line. Mathematically, this means that an isometry satises

Lξ g = 0, (3.4)
which, according to (3.2) can be written

∇µ ξν + ∇ν ξµ = 0. (3.5)
1 The concept of integral curves was dened in Section 1.3.

43
This is the Killing equation and any tangent vector eld ξ satisfying this equation
is known as a Killing vector.

In practice, it is not always easy to nd Killing vectors except in some peculiar situa-
tions. Indeed, let us explicitly write that the Lie derivative along some tangent vector ξ
og the metric vanishes

ξ α ∂α gµν + gµα ∂ν ξ α + gαν ∂µ ξ α = 0. (3.6)


If we now try to solve this equation with a constant tangent vector ξ and assuming
that the metric does not depend on a coordinate xn i.e. ∂g
∂xn
µν
= 0, then the above equation
boils down to

ξ α ∂α gµν = 0. (3.7)
We then see that the tangent vector ξ = (0, ..., 1, ..., 0) where the 1 is at the nth posi-
tion is a Killing vector.

Moreover, it is easy to show that if ξ and χ are Killing vectors, then any linear com-
bination with non-varying constants aξ + bχ and [ξ, χ] are also Killing vectors.
Example. • For a two dimensional sphere described in (θ, ϕ) coordinates, the metric reads2

ds2 = dθ2 + sin2 θ dϕ2 . (3.8)


The metric does not explicitly depend on ϕ so ξ = ∂ϕ = (0, 1) is a Killing vector.

• For the Schwarzschild metris written in the usual spherical coordinates (t, r, θ, ϕ)
   −1
2M 2M
2
ds = − 1 − 2
dt + 1 − dr2 + r2 dΩ2 , (3.9)
r r
one sees that it does not depend on t nor ϕ so that ξ = ∂t = (1, 0, 0, 0) and χ = ∂ϕ = (0, 0, 0, 1) are
Killing vectors.

A Killing vector is a tangent vector ξ that satises the Killing equation

Lξ g = 0 i.e. ∇µ ξν + ∇ν ξµ = 0. (3.10)
If the metric does not dependent on a coordinate xn , then

ξ µ = (0, ..., 0, 1
|{z} , 0, ..., 0), (3.11)
nth position

is a Killing vector.

3.1.2 Conserved Charges


We are used to the fact that symmetries lead to conserved quantities. In the present
context, the concept of "symmetries" is replaced by "symmetries of the metric", and we
therefore expect conserved charges associated with the presence of Killing vectors. Here
2 We set the radius R = 1.

44
are two important examples.

Let ξ µ be a Killing vector and xµ (λ) be a geodesic associated to the tangent vector
X µ , then the quantity

ξµ X µ , (3.12)
is a conserved quantity along the geodesic. Indeed,

X ν ∇ν (ξµ X µ ) = X ν (X µ ∇ν ξµ + ξµ ∇ν X µ )
µ ν
= X
| {zX } ∇ν ξµ +ξµ X ν ∇ν X µ
symmetric
| {z } | {z } (3.13)
anti-symmetric =0 geodesic

= 0.
Let ξ µ be a Killing vector and T µν the covariantly conserved symmetric energy-
momentum tensor, ∇µ T µν = 0, then the current

ξν T µν , (3.14)
is covariantly conserved. Indeed,

∇µ (ξν T µν ) = |{z}
T µν ∇µ ξν +ξν ∇µ T µν = 0. (3.15)
| {z } | {z }
symmetric
anti-symmetric =0

Hence, as we now have a conserved current, we can associate with it a conserved


charge in the way discussed above. The argument evidently relies on the fact that T µν is
symmetric and covariantly conserved i.e. divergent-less.
Example. As an example of application, let us derive the energy conservation equation for the Schwarzschild
metric where we assume3 θ = π/2
   −1
2M 2M
2
ds = − 1 − 2
dt + 1 − dr2 + r2 dϕ2 . (3.16)
r r
We know from te previous section that ξtµ = ∂t and ξϕµ = ∂ϕ are Killing vectors. Let us consider a tangent
vector eld4 U such that Uµ U µ =  where  = 0, ±1. We then know that −E = ξµt U µ and L = ξµϕ U µ are
conserved quantities5 . Denoting U µ = (U t , U r , 0, U ϕ ), one obtains
gµν ξtµ U ν = gtt U t
(3.17)
gµν ξϕµ U ν = gϕϕ U ϕ ,
which leads to
E L
Ut = and Uϕ = . (3.18)
1 − 2M
r
r2

Writing Uα U α = gtt (U t )2 + grr (U r )2 + gϕϕ (U ϕ )2 =  leads to


L2
  
2M
r 2
(U ) + − 1− = E2. (3.19)
r2 r

3 We assume planar trajectories.


4 Namely the velocity of an object/particle.
5 The total energy and the angular momentum respectively measured by observers at innity.

45
Let ξ µ be a Killing vector and xµ (λ) be a geodesic associated to the tangent vector
X µ i.e. X ν ∇ν X µ = 0, then the charge

ξµ X µ , (3.20)
is a conserved quantity along the geodesic.

Let ξ µ be a Killing vector and T µν a covariantly conserved symmetric tensora ,


∇µ T µν = 0, then the current

ξν T µν , (3.21)
is covariantly conserved i.e. divergent-less.
a For example the energy-momentum tensor.

3.1.3 Useful Identity Relating Curvature and Killing


Vectors
In Riemannian geometry the rich interplay between symmetries and geometry is reected
in relations between the curvature tensor and Killing vectors of a metric.

Let ξ µ be a Killing vector, then it must obey


σ
∇µ ∇ν ξρ = Rµνρ ξσ , (3.22)
where σ
Rµνρ is the Riemann tensor. Indeed, when applying three times the Riemann tensor
identity

∇µ ∇ν ξ σ − ∇ν ∇µ ξ σ = Rρµν
σ
ξρ, (3.23)
and the Killing equation

∇µ ξν + ∇ν ξµ = 0, (3.24)
one obtains
σ
∇µ ∇ν ξρ = ∇ν ∇µ ξρ − Rρµν ξσ
σ
= −∇ν ∇ρ ξµ − Rρµν ξσ
σ σ
= −∇ρ ∇ν ξµ − Rρµν ξσ + Rµνρ ξσ
σ σ
(3.25)
= ∇ρ ∇µ ξν − Rρµν ξσ + Rµνρ ξσ
σ σ σ
= ∇µ ∇ρ ξν − Rρµν ξσ + Rµνρ ξσ − Rνρµ ξσ
σ σ σ
= −∇µ ∇ν ξρ − Rρµν ξσ + Rµνρ ξσ − Rνρµ ξσ
so that
σ
∇µ ∇ν ξρ = Rµνρ ξσ , (3.26)
where we used the Bianchi identity.

46
Let ξ µ be a Killing vector, then it must obey
σ
∇µ ∇ν ξρ = Rµνρ ξσ , (3.27)
where Rµνρ
σ
is the Riemann tensor.

3.1.4 Maximal Symmetry and Constant Curvature


In order to understand how to dene and characterise maximally symmetric spaces, we
will need to obtain some more information about how Killing vectors can be classied.
Our starting point is, as in the previous section, the identity reproduced here with the
explicit x-dependence included for present purposes

σ
∇µ ∇ν ξρ (x) = Rµνρ (x)ξσ (x). (3.28)

In particular, this shows that the second derivatives of the Killing vector at a point x0
are again expressed in terms of the value of the Killing vector itself at that point. This
means that, remarkably, a Killing vector eld ξµ (x) is completely and uniquely determined
everywhere by the values of ξµ (x0 ) and ∇µ ξν (x0 ) at a single point x0 . Since, in an D-
dimensional spacetime there can be at most D linearly independent vectors (ξµ (x0 )) at
a point, and at most D(D − 1)/2 independent anti-symmetric matrices (∇µ ξν (x0 )), we
reach the conclusion that an D-dimensional spacetime can have at most

D(D − 1) D(D + 1)
D+ = , (3.29)
2 2
independent Killing vectors. A spacetime with this maximal number of Killing vectors is
called maximally symmetric.

Example. The D-dimensional Minkowski spacetime is maximally symmetric. Note that D(D + 1)/2
for D = 4 agrees with the dimension of the Poincaré group6 , the group of transformations that leave
the Minkowski metric invariant. We can also cite the de Sitter and the anti-de Sitter spacetimes being
maximally symmetric.

One can also show that maximally symmetric spacetimes have constant curvature
R = constant. For such spacetimes, one can show that the Riemann tensor takes the
following form

R
Rµνρσ = (gµρ gνσ − gµσ gνρ ). (3.30)
D(D − 1)

63 rotations, 3 boosts, 3 translations and parity.

47
A D-dimensional spacetime is maximally symmetric is it has
D(D + 1)
, (3.31)
2
independent Killing vectors. Those spacetimes have constant curvature R =
constant and the Riemann tensor can be written
R
Rµνρσ = (gµρ gνσ − gµσ gνρ ). (3.32)
D(D − 1)

3.2 Variational Calculus


Although we have met variational calculus when deriving the geodesic equation from
the Euler-Lagrange equation, we will now introduce variational calculus of tensors. Ulti-
mately, the aim is to be familiar with how to derive the Einstein equation starting from
the Einstein-Hilbert action. This way of deriving the Einstein equation is more suitable
for beyond classical General Relativity extensions because all our fundamental theories of
physics are described by action principles.

3.2.1 Covariant Volume Element


In this section we will address the issue of generally covariant integration in a spacetime
equipped with
´ Da metric. It is immediately apparent that the integral of a scalar f (x) over
spacetime d x f (x) is not generally covariant because one has to introduce the Jacobian
of the coordinate transformation which in general is not equals to one.

Let us recall the standard tensorial transformation behaviour of the metric under
coordinate transformations x → x0 ,
∂xµ ∂xν
0
gαβ (x0 ) = gµν (x). (3.33)
∂x0α ∂x0β
it follows that the absolute value of the metric determinant |g| = |det g| does not
transform like a scalar but instead transforms as
2
∂x0
  
∂x −2
g = det
0
g = det g. (3.34)
∂x0 ∂x

In particular its square rootg transforms as
 0
∂x −1 √
g 0 = det (3.35)
p
g.
∂x

Therfore the combined expression dD x g is invariant under general coordinate trans-
formation

(3.36)
p
dD x0 g 0 = dD x g,
and can be used to dene integrals of scalars f (x) in a generally covariant way

48
ˆ ˆ

(3.37)
p
D 0 0
d x g f (x ) = dD x g f (x).
0

Note that this is also frequently the quickest way to determine the volume element in
non- Cartesian coordinates in Euclidean space.
Example. Let us derive the volume element of the three dimensional euclidean space in spherical
coordinates (r, θ, ϕ) denoted by {x0µ }. The usual Cartesian coordinates are denoted {xµ }. Instead of
laboriously determining the Jacobi matrix for the coordinate transformation, and then calculating its
determinant, all one needs to know is the metric

ds2 = dr2 + r2 dθ2 + r2 sin2 θdϕ2 . (3.38)


Computing the determinant of the metric is straightforward because it is diagonal g = r sin θ and
4 2

therefore we have

d3 x = g d3 x0 = r2 sin θ dr dθ dϕ, (3.39)
which is of course the standard result.


The covariant volume element for integration is dD x g which is invariant under
general coordinate transformation x → x0

(3.40)
p
dD x0 g 0 = dD x g,
and can be used to dene integrals of scalars f (x) in a generally covariant way
ˆ ˆ

(3.41)
p
D 0 0
d x g f (x ) = dD x g f (x).
0

3.2.2 Divergence Theorem


Gauss's theorem, also known as the divergence theorem, states that if you integrate
a total derivative, you get a boundary term. There is a particular version of this theorem
in curved space. Before dive in, let us rst state the following result.

Lemma. The contraction of the Christoel symbols can be written as


1 √
Γµµν = √ ∂ν g, (3.42)
g
where g = det g . Note that for Lorentzian manifolds, the metric determinant is negative
and so one needs to perform the following replacement det g → det |g|. Indeed, from the
explicit relation between the Christoel symbols and the metric, one obtains
1 1   1
Γµµν = g µρ ∂ν gµρ = Tr g −1 ∂ν g = Tr [∂ν log g] . (3.43)
2 2 2
However, there is a useful identity for the log of any diagonalisable matrix A. They
obey

Tr [log A] = log det A. (3.44)

49
This is clearly true for a diagonal matrix, since the determinant is the product of
eigenvalues while the trace is the sum. But both trace and determinant are invariant
under conjugation, so this is also true for diagonalisable matrices. Applying it to our
metric formula above, we have

1 1 1 1 1
Γµµν = Tr [∂ν log g] = ∂ν log det g = ∂ν detg = √ ∂ν detg, (3.45)
p
2 2 2 detg detg
which is the claimed result. With this at hand, we can now prove the following theorem.

Theorem. Consider a region of a D-dimensional manifold M with boundary ∂M. Let


nµ be an outward-pointing, unit vector orthogonal to ∂M. Then, for any tangent vector
eld X µ on M, we have
ˆ ˆ
D √ √
d x g ∇µ X = µ
dD−1 x γ nµ X µ , (3.46)
M ∂M
where γij is the pull-back of the metric to ∂M and γ = det γ .

Proof. Using the lemma above, the integrand is

√ √ √ 1 √ √
g ∇µ X µ = g(∂µ X µ + Γµµν X ν ) = g(∂µ X µ + X ν √ ∂ν g) = ∂ν ( g X µ ), (3.47)
g

so that the integral is


ˆ ˆ
√ √
D
d x g ∇µ X = µ
dD x ∂ν ( g X µ ), (3.48)
M M
which is now an integral of an ordinary partial derivative so we can apply the usual
divergence, also called Stokes' theorem, that we are familiar with. It remains only to
evaluate what is happening at the boundary ∂M. For this, it is useful to pick coordinates
so that the boundary ∂M is a surface of constant xn . Furthermore, we will restrict to
metrics of the form7
 
γij 0
gµν = . (3.49)
0 1
Then by our usual result of integration, we have
ˆ ˆ
√ √
d x ∂ν ( g X µ ) =
D
dD−1 x γ X n . (3.50)
M ∂M
The unit normal vector nµ is given by nµ = (0, 0, ..., 1), which satises gµν nµ nν = 1 as
it should. We then have nµ = gµν nν = (0, 0, ..., 1) so wa can write
ˆ ˆ
√ √
D
d x ∂ν ( g X µ ) = dD−1 x γ nµ X µ , (3.51)
M ∂M
which is the result we need. As the nal expression is a covariant quantity, it is true in
general.

7 We construct a metric γij for the boundary hypersurface.

50
The contraction of the Christoel symbols can be written
1 √
Γµµν = √ ∂ν g, (3.52)
g
wherea g = det g .

Consider a region of a D-dimensional manifold M with boundary ∂M. Let nµ be


an outward-pointing, unit vector orthogonal to ∂M. Then, for any tangent vector
eld X µ on M, we have
ˆ ˆ
√ √
D
d x g ∇µ X = µ
dD−1 x γ nµ X µ , (3.53)
M ∂M
where γij is the pull-back of the metric to ∂M and γ = det γ .
a Note that for Lorentzian manifolds, the metric determinant is negative and so one needs to
perform the following replacement det g → det |g|.

3.2.3 Einstein-Hilbert Action


Gravity is also a fundamental theory that can be described by action principles. The
straight-jacket of dierential geometry places enormous restrictions on the kind of actions
that we can write down. These restrictions ensure that the action is something intrinsic
to the metric itself, rather than depending on our choice of coordinate.

In General Relativity, the gravitational eld is identied with a metric gµν on a four
dimensional Lorentzian manifold that we call spacetime. To build the Einstein-Hilbert
action that governs the metric dynamics, we know, from a previous section, that we need
the covariant volume element to integrate over a manifold. Furthermore, given that we
only have the metric to play with, the simplest scalar function is the Ricci scalar R. This
motivates us to consider the wonderful concise action
ˆ

S= d4 x −g R, (3.54)

which is the famous Einstein-Hilbert action8 . Note that the minus sign under the square
root arises because we are in a Lorentzian spacetime. As a quick sanity check, recall that
the Ricci tensor takes the schematic form R ∼ ∂Γ + ΓΓ while the Levi-Civita connection
itself is Γ ∼ ∂g . This means that the Einstein-Hilbert action is second order in derivatives,
just like other actions we consider in physics.

Note that written in this way, the Einstein-Hilbert action is suitable to consider that
classical General Relativity is just an approximated theory of a more general theory with
the following action
ˆ

S= d4 x −g (R + σ2 R2 + σ3 R3 + ...), (3.55)

8 Note that a dimensional analysis requires that the physical action must be multiplied by c3 /(16πG).

51
where σi are constants. For example truncating the series expansion up to second order
i.e. keeping only the ∼ R term gives rise to the so-called Starobinsky potential which
2

can describe ination during an early stage of the Universe9 .

In General Relativity, the gravitational eld is identied with a metric gµν on a four
dimensional Lorentzian manifold that we call spacetime. The metric dynamics is
governed by the Einstein-Hilbert action
ˆ

S= d4 x −g R. (3.56)

3.2.4 Action Variation


In this section we would like to determine the Euler-Lagrange equation arising from the
Einstein-Hilbert action. We do this in the usual way, by starting with some xed metric
gµν (x) and seeing how the action changes when we shift gµν (x) → gµν (x) + δgµν (x). We
then try to write the integrand of δS as proportional to δg µν .

Writing the Ricci scalar as R = g µν Rµν , the Einstein-Hilbert action changes as


ˆ
√ √ √
d4 x [δ −g]g µν Rµν + −g[δg µν ]Rµν + −g g µν [δRµν ] . (3.57)

δS =

The aim is then to derive the variation of the inverse metric, the covariant volume
element and the Ricci tensor.

It turns out that it is slightly easier to think of the variation in terms of the inverse
metric δg µν . This is equivalent to the variation of the metric δgµν , the two are related by
gρµ g µν = δρν so that

(δgρµ )g µν + gρµ δg µν = 0 i.e. δg µν = −g µρ g νσ δgρσ . (3.58)


The middle term in (3.57) is already proportional to δg µν . We now deal with the rst
and third terms in turn.

First, we use the standard trick log detA = Tr[log A] for any diagonalisable matrix A
to write
1
δ(detA) = Tr[A−1 δA]. (3.59)
detA
Applying this result to the metric, we have
√ 1 1 1√
δ −g = √ (−g)g µν δgµν = −gg µν δgµν . (3.60)
2 −g 2
Using g µν δgµν = −gµν δg µν , one obtains the variation of the covariant volume element
9 However,this potential has been proved wrong by the latest Planck satellite data of the Cosmic
Microwave Background.

52
1√ √
δ −g gµν δg µν .
−g = − (3.61)
2
We now claim that that the nal term g µν δRµν is a total derivative (boundary term)
and can be written as10

g µν δRµν = ∇µ X µ with X µ = g ρν δΓµρν − g µν δΓρνρ . (3.62)


using all the previous results, the variation of the action can then be written
ˆ

  
1
δS = 4
dx µν µ
−g Rµν − R gµν δg + ∇µ X . (3.63)
2
This nal term is a total derivative and by the divergence theorem, stated in a previous
section, we ignore it. Requiring that the action is extremised δS = 0, one obtains the
metric equation of motion11
1
Gµν = Rµν − R gµν = 0. (3.64)
2

Varying the Einstein-Hilbert action reads

ˆ
√ √ √
d4 x [δ −g]g µν Rµν + −g[δg µν ]Rµν + −g g µν [δRµν ] . (3.65)

δS =

Varying the covariant volume element and the Ricci tensor reads
√ 1√
δ −g = − −g gµν δg µν
2 (3.66)
g µν δRµν = ∇µ X µ with X µ = g ρν δΓµρν − g µν δΓρνρ .
The last variation turns out to be a boundary term. Using all the previous results,
the variation of the action is written
ˆ

  
1
δS = 4
dx µν µ
−g Rµν − R gµν δg + ∇µ X . (3.67)
2
Ignoring the boundary term and requiring the action to be extremised δS = 0, one
obtains the Einstein eld equation in vacuum
1
Gµν = Rµν − R gµν = 0. (3.68)
2

Note that they simplify somewhat; if we contract the last equation with g µν , we nd
that R = 0. Substituting this back in, the vacuum Einstein equation is simply the
requirement that the metric is Ricci at

Rµν = 0. (3.69)
10 We do not prove this result because it is very long. Note that one can easily derive this result in
normal coordinates but these are not introduced in these lecture notes. For interested people, I found a
Youtube video that explicitly makes this derivation.
11 In vaccum.

53
Note that we happily discarded the boundary term, a standard practice whenever we
invoke the variational principle. It turns out that there are some situations in General
Relativity where we should not be quite so cavalier. In such circumstances, one can be
more careful by invoking the so-called Gibbons-Hawking boundary term.

54
The Essential in a Scheme

Dierential Manifold M

Dening tangent vectors X = X µ ∂µ ,


one-forms ω = ωµ dxµ , tensors T = Tνµ ∂µ dxν ,
Lie derivative L

Metric gµν Connexion and covariant


derivative ∇X T

Geodesics
´ by extremizing
Auto-parallel curves and
p
S= dλ −gµν X µ X ν .
parallel transport

Spacetime symmetries
Killing vectors ξ µ
Fundamental theorem
of Riemannian geormetry

Conserved quantities

Levi-Civita connexion
(torsion-free, ∇X g = 0)

´ √
Einstein-Hilbert action S = d4 x −g R

Einstein eld equations


Black holes, gravitational waves, etc...

55

You might also like