Lecture Notes Geometry 20
Lecture Notes Geometry 20
1 Introduction
It is indeed very satisfying for a mathematician to define an affine space as being a set acted on by a vector
space (and this is what I do here) but this formal approach, although elegant, must not hide the phenomenological
aspect of elementary geometry, its own aesthetics: yes, Thales theorem expresses the fact that a projection is an
affine mapping, no, you do not need to orient the plane before defining oriented angles. . . but this will prevent
neither the Euler circle from being tangent to the incircle and excircles, nor the Simson lines from enveloping a
three-cusped hypocycloid! -Michelle Audin.
In these lecture notes we aim to summarize some of the main points of the first chapters of the book Geometry
by Michelle Audin: We will also treat some elementary notions of differential geometry after. All our vector
spaces will be over R for simplicity.
2 Affine space
A natural setting for discussing elementary geometry of lines and planes is affine space. We are mostly con-
cerned with properties of parallelism and intersection, not yet with distances. The approach is firmly rooted in
linear algebra as we will see affine space is just a vector space where one refuses to say where the origin is as
geometrically there is no natural choice of origin.
Definition 2.1. (Audin 1.1)
Θ
A set E is called an affine space if there exist a vector space E and a map E × E −
→ E. We use the notation
−−→
Θ(A, B) = AB. The map should satisfy two conditions:
ΘA −−→
1. For all A ∈ E the map B 7−−→ AB is a bijection from E to E.
−→ −−→ −−→
2. For all A, B, C ∈ E we have AC + CB = AB.
We say E is directed by vector space E or E underlies E. Also sometimes the image of ΘA is named EA : it
is E viewed as a vector space with the origin at A.
−−→
For two points A, B ∈ E we may express ΘA ◦ Θ−1 −1
B : E → E as follows: ΘA ◦ ΘB v = AB + v. This is because
−
−→ −→
Θ−1
B v = C for a unique C ∈ E such that v = BC. But then ΘA C = AC and the second property of affine space
−→ −−→ −−→ −→ −−→
says AC = AB + BC. In other words ΘA ◦ Θ−1 B v = AC = AB + v.
Going in the opposite direction we can promote any vector space V to an associated affine space V as follows.
Define V = V as a set and define Θ : V × V → V by Θ(A, B) = B − A. This provides many examples of affine
spaces but just like with vector spaces there is in some sense only one for each dimension. We call it affine
n-space. To get even more concrete we could do the above procedure to V = Rn . However being so concrete is
not always the easiest approach to difficult geometric problems. Often it is not the coordinates that count but
rather the relationships between them and these are observed more readily in an affine setting.
Definition 2.2. (Audin p.10)
A subset F ⊂ E of an affine space E is called an affine subspace if it is empty or for some A ∈ F we have
ΘA (F) is a linear subspace of E.
1
Definition 2.4. (Audin p.12)
Points A0 . . . Ak ∈ E are affine independent if hA0 , . . . Ak i has dimension k. When k = dim E we say
(A0 , . . . , Ak ) is an affine frame of E.
Affine independence is closely related to linear independence as becomes clear when we choose one of the
points as an origin. More precisely,
−−−→
Lemma 2.1. The points A0 . . . Ak ∈ E are affine independent if and only if the vectors A0 Ai for i = 1, . . . k
are linearly independent.
Proof. F = hA0 , . . . Ak i, is an affine subspace so ΘA0 : E → E sends F to a linear subspace F = ΘA0 (F). Affine
independence means that dim F = k. On the other hand, by definition of ΘA0 the space F is spanned by the k
−−−→
vectors A0 Ai for i = 1, . . . k. Therefore dim F = k implies these are all linearly independent and vice versa.
In particular the above lemma asserts the relation between affine frames in E and bases of E. This is of
great importance for applying linear algebra to understand affine space. Of course the choice of origin A0 in
the above was arbitrary, we could just as well have chosen some other point Ai .
Definition 2.5. (Audin p.12)
Two affine subspaces are said to be parallel if they have the same direction (underlying vector space).
More precisely this means that affine subspaces F, G ⊂ E are parallel iff for every A ∈ F and B ∈ G we have
ΘA (F) = ΘB (G) as linear subspaces of E.
Definition 2.6. (Audin p.14)
φ
A map E −
→ F between two affine spaces is called an affine map if for some O ∈ E there is a linear map
f −−−−−−−→ −−→ →
−
E−→ F such that for all M ∈ E we have φ(O)φ(M ) = f (OM ). Sometimes we use the notation f = φ .
Sometimes it is useful to reformulate this condition in terms of the bijection ΘO : E → E and Θφ(O) : F → F
as follows: f ◦ ΘO = Θφ(O) ◦ φ. These compositions are conveniently visualized in a (commutative) diagram:
φ
E −−−−→ F
Θ Θ
y O y φ(O)
f
E −−−−→ F
So after choosing an origin O and the compatible choice of origin φ(O) we find that φ is represented by the
linear map f .
Examples of affine maps:
−−→
1) For u ∈ E the translation tu : E → E is an affine map defined by tu (A) = B with AB = u.
2) For O ∈ E and λ ∈ R the (central) dilation with center O and ratio λ, denoted hO,λ : E → E, is the affine
−−−−−−−→ −−→
map defined by OhO,λ (M ) = λOM .
3) Another example is the projection πF ,L onto affine subspace F ⊂ E in the direction of linear subspace L ⊂ E.
This makes sense only if L + F = E and L ∩ F = 0 where F is the direction of F. In that case πF ,L (A) = B
−−→
where B ∈ F is the unique point such that AB ∈ L. Point B is indeed unique because if C was another such
−−→ −→ −→
point then AB − AC would be in F ∩ L. Point B exists because for any O ∈ F we can write OA = ` + v for
−−→ f
some v ∈ F, ` ∈ L and set B to be such that OB = v. The corresponding linear map is the projection ` + v 7− →v
−→ −−−−−−−−−−−−→
for any ` ∈ L, v ∈ F . Indeed f (OA) = πF ,L (O)πF ,L (A).
2
Proof. Given affine frame A0 , . . . An and O = A0 Lemma 2.1 tells us that ΘO sends A1 , . . . An to a basis
→
−
v1 , . . . vn of E. Using the same notation as above the linear map f = φ completely determines φ. Also
−1
f (vi ) = Θφ(O) ◦ φ ◦ ΘO (vi ) = Θφ(O) (φ(Ai )) shows that f in turn is completely determined by the images of the
frame Ai under φ. Here we used that a linear map is determined by what it does on the basis vi .
−
−→
AB −−→ −−→
For four collinear points A, B, C 6= D the ratio −
−→ means the scalar λ such that AB = λCD.
CD
3
Exercises
1. Can you give an example of an affine space that is not a vector space?
2. Prove that for any points B, C of affine subspace F ⊂ E we have ΘB (F) = ΘC (F).
3. Prove that through any two points of an affine space passes a unique line.
4. Let A be an m × n matrix and b ∈ Rm . Define F = {x ∈ Rn : Ax = b}. Prove that F is an affine subspace
of Rn . What is its direction? When is it empty? Express the dimension of F in terms of the rank of A.
5. Given an affine map φ : E → F between affine spaces and M an affine subspace of F. Prove that φ−1 (M)
is an affine subspace of E.
6. Prove that an affine map is completely determined by the image of an affine frame.
7. Is Pappus theorem valid if the lines D, D0 are two distinct lines in three-dimensional affine space?
8. The set of all invertible affine maps from E to itself is denoted GA(E).
(a) Prove that the composition of two affine maps φ and ψ is again an affine map, and that the linear map
−−−→ → −→
−
associated to ψ ◦φ is the composition of the linear maps associated to ψ and φ, i.e. that ψ ◦ φ = ψ φ .
(b) Let φ ∈ GA(E), and let h(O, λ) be a central dilation. Compute φ ◦ h(O, λ) ◦ φ−1 .
(c) Prove that two dilations with the same center commute.
(d) Compute h(B, λ0 ) ◦ h(A, λ).
(e) Explain whether or not the set of all dilations is a subgroup of GA(E).
−−→ −−→
(f) Prove AB = −BA for any A, B ∈ E.
9. In this exercise we aim to prove Menelaus’ theorem in a couple of steps, see also Audin Exercise I.37.
In an affine plane consider three distinct non-collinear points A, B, C and points A0 on line BC, B 0 on CA
−−→ −− → −−→
A0 B B0 C C0A
and C 0 on AB all distinct from A, B, C. Points A0 , B 0 , C 0 are collinear if and only if −−
0
→ · −−0→ · −−
0
→ = 1.
AC B A C B
(c) The composition of the three maps above sends B to itself, why must it be of the form hB,λ for some
λ ∈ R?
(d) The first two maps fix the line C 0 B 0 .
(e) The third map fixes C 0 B 0 if and only if A0 is on that line.
(f) Conclude that if the composition of the three is the identity then A0 , B 0 , C 0 are collinear.
−−→ −−→
(g) The point B is not on A0 C 0 because then A0 B and C 0 B are proportional and hence the sides of the
triangle BC and AB would not be independent.
(h) Conversely if A0 , B 0 , C 0 are collinear then show the composition of the three is the identity because
it fixes both the line C 0 B 0 and B (not on that line).
10. (a) Let V1 and V2 be linear subspaces of a vector space V . Under which conditions is V1 ∪ V2 a linear
subspace of V ? Prove your claim.
(b) Let F1 and F2 be affine subspaces of an affine space E, directed by E. Under which condition is
F1 ∪ F2 an affine subspace of E? Prove your claim.
11. Let A be a point on D be a line in an affine plane P with underlying vector space P . For a one-dimensional
linear subspace L ⊂ P not equal to ΘA (D) we consider the mapping π : P → P defined as follows. π(M )
−−−→
is the point M 0 defined by M 0 ∈ D and M M 0 ∈ L.
−−−→
(a) Why can there be only one point M 0 with properties M 0 ∈ D and M M 0 ∈ L?
(b) Prove that π is an affine map.
(c) What is the linear map →−
π : P → P associated to π?
(d) Why is π not defined if L is in the same direction as D?
12. All affine n-spaces are ’the same’.
(a) Prove that any two vector spaces of dimension n are isomorphic (i.e.) there is a linear bijection from
one to the other that has a linear inverse.
(b) Suppose E and F are two n-dimensional affine spaces. Pick B ∈ E and C ∈ F and use ΘB , ΘC and
the previous part to construct an affine bijection between E and F with affine inverse.
4
3 Euclidean space in general
By Euclidean space we mean affine space that also has a notion of dot-product (aka scalar or inner)-product
h·, ·i. This allows us to discuss angles and distances as Euclid did but still using vector spaces as our foundation.
Definition 3.1. (Audin p.44)
A Euclidean vector space is a vector space with a choice of scalar product. A Euclidean affine space E is an
−−→
affine space directed by a Euclidean vector space. The distance between two points A, B is d(A, B) = |AB|.
In a Euclidean vector space E the notion of perpendicular is very important. For any subset S ⊂ E we
define S ⊥ = {x ∈ E : ∀s ∈ S, hs, xi = 0}. When S is a subspace of E then E = S ⊕ S ⊥ . Recall that A = B ⊕ C
means every element a of vector space A can uniquely be written as a = b + c for some b in subspace B ⊂ A
and c in subspace C ⊂ A.
5
Theorem 3.2. (Audin thm 2.2, p.49)
Suppose E is an n dimensional Euclidean affine space. Any element of Isom(E) can be written as the composition
of at most n + 1 reflections.
Proof. If the affine isometry ψ has a fixed point O then there is a linear isometry L = ΘO ◦ ψ ◦ Θ−1O which by
Theorem 3.1 is the composition of at most n reflections through hyperplanes containing O. This means that the
same is true for ψ because we may conjugate each of these reflections by ΘO . More precisely if L = sH1 ◦ . . . sHj
then ψ = Θ−1 −1 −1 −1 −1
O ◦ L ◦ ΘO = ΘO ◦ sH1 ◦ ΘO ◦ ΘO ◦ · · · ◦ ΘO ◦ sHj ◦ ΘO and each ΘO ◦ sH1 ◦ ΘO is areflection in
an affine hyperplane through O.
In case ψ does not fix any point then choose A ∈ E and set A0 = ψA. Using Lemma 3.1 we see that σH ◦ ψ
fixes A so we conclude ψ is the composition of at most n + 1 reflections.
Linear reflections reverse orientation in the sense that they have determinant −1. Since the determinant of
the product is the product of the determinants this means that linear isometries have determinant ±1. The
linear isometries with determinant 1 are known as the rigid motions. The same terminology is used for the
affine isometries. We use the notation Isom+ (E) for the rigid affine motions and O+ (E) or SO(E) for the linear
rigid motions.
Choosing an orthonormal basis of Euclidean n-dimensional space E we may identify O(E) with a subset
O(n) ⊂ Mn inside the set Mn of the n × n matrices. Recall the columns of the matrix of a linear map
are the images of the base vectors and a linear isometry preserves the inner product. Therefore the columns
of the matrix with respect to an orthonormal basis must again form an orthonormal basis. In other words
2
O(n) = {A ∈ Mn : AAt = I}, where At is the transpose of matrix A. Identifying Mn with Rn we see that
A 7→ AAt is a continuous map and O(n) is the inverse image of a point so O(n) must be a closed subset of Mn
with respect to this topology. Also O(n) is bounded since all the columns must be unit vectors so all in all we
concluded that O(n) is a compact subset of Mn .
More precisely the determinant shows that O(n) decomposes neatly into two equal parts, the matrices with
determinant 1 called O+ (n) and those with determinant −1.
The two dimensional case is fundamental so we examine O+ (2) in more detail. It will be shown to be
identified with the unit circle U in the complex plane.
Lemma 3.2. (Audin prop 3.4, p. 53)
The group O+ (2) is isomorphic and homeomorphic to the complex units U ⊂ C.
a c
Proof. Matrix A = belongs to O(2) if and only if the columns form an orthonormal basis. So
b d
a2 + b2 = 1 and c2 + d2 = 1 and ac + bd = 0. This means that c = −b and d = a for some ∈ {−1, 1}. In
+ a −b
fact = det A. The map ϕ : O (2) → U given by ϕ = a + ib is a homomorphism of groups and it
b a
is also a continous bijection with continuous inverse as the reader can check by explicit calculations, finishing
the proof.
We thus find a surjective map R → U → O+ (2) sending θ ∈ R to eθi . The image in O+ (2) is called the
rotation with angle θ. In section 4 we will have more to say about angles and plane geometry.
A 0-simplex is a point, a 1-simplex is simply an interval and a 2-simplex is a triangle. 3-simplices are known
as tetrahedra. The face of simplex [A0 , . . . An ] opposite to Ai is the (n − 1)-simplex [A0 , . . . Âi . . . An ] where the
hat means we removed Ai from the sequence. The faces of a 2-simplex are simply its sides viewed as segments.
The numbers βi appearing in the definition of simplex are known as barycentric coordinates.
Definition 3.4. Numbers β0 , . . . βn ∈ R are barycentric coordinates of point B with respect to an affine
Pn −−→
frame A0 , . . . An if i=0 βi BAi = 0.
Pn
Barycentric coordinates are not unique but they are unique when we add the P
condition that i=0 βi = 1.
n
This condition can always be met because affine independence of the Ai implies i=0 βi 6= 0. The n-simplex
n+1
can be defined as those points with barycentric coordinates in [0, 1] .
6
What is the middle of a simplex? There are several competing notions and we will survey a few of them:
centroid, circumcenter and orthocenter. For 2-simplices (triangles) we will investigate how they relate to each
other.
Looking at the definition of simplex we defined the centroid G of the simplex as follows. G ∈ E is the
1 −−→
Pn
unique point such that i=0 n+1 GAi = 0.
The faces of the simplex also have centroids called Gi and the lines Gi Ai all meet in G. To see why
( this is the
1−t
n j 6= i
case, notice that the segment Ai Gi is precisely the set of points with barycentric coordinates βj =
t j=i
1 1 −−→ −−→
for t ∈ [0, 1]. Since G has barycentric coordinates βk = n+1 this is the case t = n+1 . In fact nGGi = Ai G
(exercise).
A competing notion of the middle of a simplex is known as the circumcenter O. It is the center of
the unique sphere that passes through all vertices of the simplex. First we should define the (n − 1)-sphere
−
→
S n−1 (c, r) with center c ∈ E and radius r ≥ 0 in affine n-space E to be S n−1 (c, r) = {A ∈ E : |cA| = r}. The
following lemma shows that the circumcenter is uniquely defined by the simplex.
Lemma 3.3. For an n-simplex [A0 , . . . , An ] in affine Euclidean n-space E there exists a unique (n − 1)-sphere
S n−1 (O, r) such that ∀i : Ai ∈ S n−1 (O, r).
Proof. We argue by induction on the dimension n. The base case n = 0 is clear so now assume the statement is
true in all dimensions < n. The face of our simplex opposite to An spans an affine hyperplane Hn in which there
is a unique (n − 1)-sphere with center On passes through all points Ai with i < n. By Lemma 3.1 there exists
a unique hyperplane F in E equidistant from A0 and An and also a line L orthogonal to Hn passing through
On . Since An is affine independent from the other Ai there must be a (unique) intersection point O = L ∩ F.
−−→
We claim that S n−1 (O, |OAn |) passes through all points Ai of the simplex. Indeed orthogonal projection on
−−→ −−→ −−→ −−→
Hn shows |OA0 | = |OAi | with i < n and by construction |OAn | = |OA0 |.
Finally the altitudes of a simplex are the affine lines Hi through Ai orthogonal to the faces opposite Ai .
Interestingly the altitudes do not necessarily meet in a single point when n > 2. For example the 3-simplex
[(1, 0, 0), (0, 0, 1), (1, 1, 0), (0, 0, 0)] has non-intersecting altitudes H2 and H1 as is easily seen by inscribing this
tetrahedron inside a regular cube.
For n = 2 the altitudes do meet in a single point called the orthocenter H of the triangle A = [A0 , A1 , A2 ].
To see why this is the case consider the triangle B = [B0 , B1 , B2 ] defined by hG,−2 (Ai ) = Bi where G is the
centroid of A. Notice that Ai Aj is parallel to Bi Bj by Lemma 2.4. Also when i, j, k are distinct Ai is the
−−−→ −−→ −−→
midpoint of the segment Bj Bk . This is because the midpoint of segment Aj Ak is Gi and Bj Ai = GAi − GBj =
−−→ −−→ −−−→ −−−→ −−−→
−2GGi + 2GAj = 2Gi Aj and likewise Bk Ai = 2Gi Ak . Finally the altitudes of A are precisely the perpendicular
bisectors of the sides of B. The latter meet in the circumcenter of B so this point is the orthocenter of A.
Looking carefully at what we just proved we find the following theorem of Euler:
Theorem 3.3. (Euler line)
The circumcenter, centroid and orthocenter O, G, H of triangle B = [B0 , B1 , B2 ] lie on a line and more precisely
−−→ −−→
−2GO = GH. Also on this line is center N of the circle through the midpoints Gi of the sides of our triangle.
Moreover, N is the midpoint of the segment HO.
Proof. As before we use two triangles A, B, where B = hG,−2 A. By construction they have the same centroid
G and Ai = Gi . Moreover, we showed the circumcenter O = OB of B is the orthocenter HA of A. By definition
this means the orthocenter H = HB of B is hG,−2 HA = hG,−2 O. We conclude that O, H, G lie on a line and
−−→ −−→ −−→
moreover GH = hG,−2 GO = −2GO.
Notice N is the circumcenter OA of A and so hG,−2 OA = OB = O. Therefore O, G, N, H lie on the same
−−→ −−→ −−→ −−→ −−→ −−→ −−→ −−→ −−→ −−→ −−→
line and GO = hG,−2 GN = −2GN . This means that N O = N G + GO = 32 GO = 2GO + GN = HG + GN =
−−→
HN .
By similar methods much more can be said about the circle through the midpoints of the sides of our triangle.
It is the famous nine-point circle.
Theorem 3.4. (Nine point circle)
Given triangle B = [B0 , B1 , B2 ], there is a circle that passes through the following nine special points: the
midpoints Gi of the sides and intersections Hi of the altitude Hi with the opposite side and the midpoints of the
segments Bi H.
Proof. Keeping the previous notation we call our triangle B and define the nine-point circle C as the circle
through the Gi = Ai . There must be a unique such circle by Lemma 3.3, also it has center N . We now need to
show that C contains both the Hi and the midpoints Mi of the segments Bi H.
7
First hG,− 12 ◦ hH,2 = hN,−1 because the left hand side has fixed point N by the previous theorem. Also
looking at the left hand side we see that hN,−1 (Mi ) = Ai . Since hN,−1 is an isometry we have |N Mi | = |N Ai |
so the Mi are on C.
Second Hi and Gi are both on line Bj Bk for i, j, k distinct. To show that |N Hi | = |N Gi | we argue as follows.
The lines HHi and OGi are parallel because both perpendicular to Bj Bk . Since by the previous theorem N is
the midpoint of segment HO, Lemma 2.4 shows N Ni is also perpendicular to Bj Bk where Ni is the midpoint
of segment Hi Gi . It follows by a reflection in line N Ni that |N Hi | = |N Gi | as desired.
There are in fact many more interesting points on the nine-point circle. For example Feuerbach found there
are precisely four circles that are tangent to all three sides of our triangle and the nine-point circle is tangent
to all four of those!
Exercises
1. Is it true that for any subset S in a Euclidean vector space E we have (S ⊥ )⊥ ?
2. Imagine a line L in affine Euclidean plane E and a vector u ∈ L in the direction (underlying vector
subspace) of L. The affine map γL,u : E → E defined by γL,u = tu ◦ σL is called a glide reflection.
(a) Prove that the glide reflection γL,u is an affine isometry.
(b) Why can γL,u not be written as the composition of two or less reflections? (Hint: fixed points)
(c) Write γL,u explicitly as a composition of three reflections.
3. Verify that sF ◦sF is the identity. Give an example of linear subspaces F, G ⊂ E such that sF ◦sG 6= sG ◦sF .
2
4. If we choose a basis O(E) can be identified with a subset of Rn where n = dim(E). Show that if f is the
composition of an even number of reflections then there is a continuous function γ : [0, 1] → O(E) with
γ(0) = idE and γ(1) = f .
5. Imagine a 2-simplex [A0 , A1 , A2 ] with centroid G and Gi the centroid of the face opposite to Ai . Using
−−→ −−→ −−→ −−−→ −−−→ −−→ −−→
GA0 + GA1 + GA2 = 0 = G0 A1 + G0 A2 , prove that 2GG0 = A0 G.
6. Prove that for any nonzero λ ∈ R and O ∈ E the dilation hO,λ sends any n-simplex in E to another
n-simplex in E.
7. By a rotation around point A in affine Euclidean plane E we mean any composition of two affine reflections
in lines passing through A. For A, B ∈ E let ρA be a rotation around A and ρB a rotation around B 6= A.
(a) What is the determinant of the linear map −
ρ→
A : E → E associated to ρA ?
−−−−−→ −
→ −
→
(b) Prove that ρA ◦ ρB = ρA ◦ ρB .
(c) What is the determinant of −
ρ−−◦−−
Aρ→? B
(d) Is it true that ρA ◦ ρB must also be a rotation around some point C ∈ E? Prove or provide a
counterexample.
Definition 4.1. A rotation is an element of O(E) that can be written as the composition of two reflections.
In fact the set O+ (E) of linear isometries of positive determinant coincides with the set of all rotations. It
forms a commutative subgroup isomorphic to O+ (2) or the unit circle in the complex plane, see Lemma 3.2.
Notice that the identity is also a rotation according to this definition. As we mentioned before, reflections
reverse orientation while rotations preserve it. Orientation is an important and often overlooked concept so we
will provide a definition.
8
Definition 4.2. (Orientation)
Recall a basis for a vector space is an ordered sequence of independent vectors that span V . An orientation of
a vector space V is a choice of an equivalence class of bases of V . Here two bases of V are said to be equivalent
if the linear map sending one to the other has positive determinant.
For V = R2 the standard choice of orientation is to take the equivalence class of the standard basis (e1 , e2 ).
It often appears as saying that the positive direction of rotation is counter-clockwise. In R3 orientation is often
referred to as the ’right hand rule’ but we will deal with space later.
As promised reflections do reverse orientation in the following sense. The reflection s = sv⊥ is easy to
describe in the basis v, v 0 where
v is orthogonal
to v 0 . By definition s(v) = −v while s(v 0 ) = v 0 so the matrix
−1 0 1
with respect to this basis is so the determinant is −1. To connect with the above definition we
0 1
would say that the bases (v, v 0 ) and (s(v), s(v 0 )) belong to different equivalence classes, i.e. orientations.
Since there are only two orientations for E and a single reflection reverses the orientation, all rotations
preserve orientation. This is why they play such an important role.
Closely tied to the concept of rotation is the notion of angle. But what exactly is an angle? This is not as
straightforward as it seems. In fact we will define four related notions of angle:
What corresponds to the ’usual’ notion of angles as real numbers modulo 2π is the measure of oriented angle.
However it is NOT well defined in our Euclidean E vector space. The common definition of an angle θ between
hu,vi
vectors u, v using the inner product as cos θ = |u||v| is also incomplete. The right hand side is symmetric in u, v
so without additional assumptions it cannot possibly distinguish between θ and −θ. Do we go from u to v or
vice versa and which way do we go?
The issue is that there is no natural choice for orientation of E. Of course we can simply choose an arbitrary
orientation but this muddles the theory and clutters the proofs. Much like affine space arises from our refusal
to make an arbitrary choice of origin.
4.1 Angles
In this subsection we define four notions of angle the most important of which is called the oriented angle
between two vectors. It is firmly rooted in our understanding of rotations.
Lemma 4.1. (Audin prop 1.1, p.65)
For any (ordered) pair of unit vectors in (u, u0 ) ∈ E there is a unique rotation that sends u to u0 .
Proof. Suppose u, u0 are unit vectors and choose a unit vector v such that u, v form an orthonormal basis. Then
u0 = au + bv for some a, b ∈ R with a2 + b2 =1. With respect to the basis u, v any orthonormal linear map
a −b
sending u to u0 and determinant 1 has matrix , see also Lemma 3.2.
b a
Our oriented angles now come out by considering pairs of vectors, modulo rotations.
Definition 4.3. (Audin p.66)
Define  to be the set of pairs of unit vectors (u, v) in E. Two pairs (p1 , p2 ), (q1 , q2 ) ∈  are said to be
equivalent if there exists a rotation that sends both p1 to q1 and p2 to q2 . The equivalence classes are are called
oriented angles and the set of oriented angles is denoted A. The class of (u, v) is denoted2 ∠(u, v) and is
called the oriented angle between u and v. The flat angle is ∠(v, −v) for any unit vector v.
0 0
This definition is extended to any pair of non-zero vectors u0 , v 0 where we set ∠(u0 , v 0 ) = ∠( |uu0 | , |vv0 | ).
To clarify the relation between angles and rotations further we define a map from oriented angles to rotations
O+ (E) as follows.
Lemma 4.2. (Audin lemma 1.2, p.66)
The map Φ̂ : Â → O+ (E) is defined by setting Φ̂(u, v) to be the unique rotation sending u to v. It defines a
bijection Φ : A → O+ (E) by sending the class of ∠(u, v) to Φ̂(u, v).
1 Remember to write down the matrix of a linear transformation one just lists the images of the base vectors as columns of this
matrix.
2 Audin uses just (u, v) to denote both the pair and the oriented angle.
9
Proof. To show Φ̂ is surjective choose any f ∈ O+ (E) and unit vector u ∈ E. We have Φ̂(u, f (u)) = f since
both sides send u to f (u) and there is only one such rotation by Lemma 4.1.
Next we need to check that Φ is well-defined by showing that if (p1 , p2 ) is equivalent in the sense of Definition
4.3 to (q1 , q2 ) then Φ̂ maps them to the same rotation in O+ (E). By Lemma 4.1 there is a unique rotation
g taking p1 to q1 and by assumption it also takes p2 to q2 . Suppose f = Φ̂(p1 , p2 ) so f (p1 ) = p2 then plane
rotations commute so f (q1 ) = f ◦ g(p1 ) = g ◦ f (p1 ) = g(p2 ) = q2 , showing f = Φ̂(q1 , q2 ).
Finally we should show that if Φ̂(p1 , p2 ) = Φ̂(q1 , q2 ) = f then there is a single rotation that sends both p1 to
q1 and p2 to q2 . By Lemma 4.1 there is a unique rotation g sending p1 to q1 . By assumption f ◦ g ◦ f −1 sends
p2 to p1 to q1 to q2 . Since the rotations form a commutative group f ◦ g ◦ f −1 = g.
The map Φ is bijective so with any rotation we may identify an oriented angle in A. We showed there is a
unique rotation that is the image of ∠(u, v) so it rotates unit vector u to v. We call this the rotation over angle
∠(u, v) and to any rotation there corresponds such an angle.
Moreover since the map Φ is a bijection we may use it to ’add’ angles by composing their images in the
group O+ (E). More precisely we set
Unpacking the definitions this says that the sum of the angles is an angle ∠(a, b) where b = f (a) and the
rotation f is the result of first rotating u to v and then w to z. At the cost of changing z we may always choose
v = w and then this kind of addition is just placing the angles next to each other (draw a picture!). This gives
the important special case
∠(u, v) + ∠(v, z) = ∠(u, z)
Now let us turn to measuring angles with respect to an orientation. This depends on the following lemma.
Lemma 4.3. Suppose (p1 , p2 ) and (q1 , q2 ) are two orthonormal bases of E in the same orientation O and
r ∈ O+ (E) is a rotation. The matrix of r with repsect to (p1 , p2 ) is the same as that with respect to (q1 , q2 ).
The orientation O thus determines an isomorphism FO : O+ (E) → O+ (2).
Proof. The linear map ρ sending basis p to basis q must be in O+ (E) since both bases are orthonormal and in
the same orientation. Now rotations commute so if rp1 = ap1 + bp2 and rp2 = cp1 + dp2 then also rq1 = rρp1 =
ρrp1 = ρ(ap1 + bp2 ) = aq1 + bq2 and similarly rq2 = rρp2 = ρrp2 = ρ(cp1 + dp2 ) = cq1 + dq2 .
The isomorphism FO comes from choosing any basis in O. This is an isomorphism by Lemma 3.2.
Since the matrix of a rotation only depends on the orientation of the basis we can use it to define our measure
of angle.
It should be clear from this definition that the measure of the angle really does depend on the choice of
orientation. Also the definition agrees with cos µ = hu, vi where µ = µ∠(u, v) for unit vectors u, v. The measure
of the flat angle is π mod 2π. Since FO is an isomorphism we have µ(a) + µ(b) = µ(a + b) mod 2π as implied
by the word measure.
Returning to the oriented angles ∠(u, v) we next define the angle between two lines and the geometric angle
between two vectors. Neither depends on a choice of orientation.
Definition 4.5. (Angle between lines and geometric angle)
The angle between two lines spanned by u and v in E is defined to be ∠(u, v) mod ∠(u, −u)
The geometric angle between two vectors u, v ∈ E is the equivalence class of ∠(u, v) under the equivalence
relation where we identify ∠(x, y) with ∠(y, x) for any x, y ∈ E.
One should check that the definition of the angle between lines does not depend on the choice of spanning
vectors u and v. To this end recall that ∠(u, −u) represents the flat angle and flipping the sign of u or v amounts
to adding or subtracting such a flat angle.
The point of defining the notion of geometric angle is that this is what is actually invariant under all
isometries. By definition a rotation ρ ∈ O+ (E) sends preserves oriented angles in the sense that ∠(u, v) =
∠(ρu, ρv). However a reflection s does not.
Lemma 4.4. (Audin prop 1.10, p.70 )
For any unit vectors u, v ∈ E and a reflection s we have ∠(su, sv) = ∠(v, u).
10
Proof. The reflection s0 = s(u−v)⊥ sends u to v and v to u. The composition s ◦ s0 is a rotation so ∠(v, u) =
∠(s ◦ s0 (v), s ◦ s0 (u)) = ∠(su, sv).
If follows from this discussion that geometric angles are preserved by all isometries, both rotations and
reflections.
So far we considered angles in a two-dimesional Euclidean vector space E. These can be carried over to the
affine Euclidean plane E as follows.
Definition 4.6. (Angles in the affine Euclidean plane)
For points O, A, B ∈ E with O 6= A and O 6= B, the oriented angle between the segments (1-simplices) [O, A]
−→ −−→
and [O, B] is ∠(OAB), defined as ∠OAB = ∠(OA, OB).
For a pair of lines L1 , L2 intersecting at O the angle between them is the angle between the lines ΘO (Li ) in E.
The geometric angle and the measure of the angle ∠OAB are defined as explained above.
which is the flat angle. The second statement follows from ∠(x, y) = −∠(y, x).
Notice the cyclic permutation of the order of the vertices A, B, C in the angle sum: ABC, BCA, CAB.
In the presence of an orientation this becomes the familiar theorem about the angle sum in a triangle. If
A, B, C are distinct points in an oriented affine Euclidean plane then θ∠ABC + θ∠BCA + θ∠CAB = π mod 2π.
This follows from the previous statement since the measure of the sum is the sum of the measures modulo 2π.
More interestingly we have the classic proposition about the angles in the circumference.
Theorem 4.1. (Audin prop 1.17, p.74)
If A, B, C are distinct points in an affine Euclidean plane and O is the circumcenter of the triangle [A, B, C]
then
∠OAB = 2∠CAB
Proof. O is the center of the circle through A, C so it is fixed by the reflection in the perpendicular bisector
H of segment [A, C]. This is the affine line of all points of equal distance to both A and C. Since σH reverses
orientation and permutes A and C we have ∠COA = ∠ACO. Lemma 4.5 tells us that ∠COA + ∠ACO +
∠OAC = 2∠COA + ∠OAC is flat. In the same way 2∠CBO + ∠OCB is also flat. Adding two flat angles gives
the angle 0 so
0 = 2∠COA + ∠OAC + 2∠CBO + ∠OCB = 2∠CBA + ∠OAB
Since ∠(x, y) = −∠(y, x) we are done.
Returning briefly to our discussion of centers of triangles, for each pair of lines there is a pair of bisectors of
the two angles defined by the lines. If the lines are directed by unit vectors u, v then these bisectors are directed
by u + v and u − v. They are perpendicular and cut the angles in two as reflection in either of the bisectors
permutes the original two lines.
One can prove that given a triangle the six bisectors of the lines generated by the sides of the triangle meet
in triples to define four points. These points are the centers of four circles tangent to the extended sides of the
triangle. One of them is inside the circle and is called the inscribed circle. As was mentioned earlier these circles
are all tangent to the nine-point circle, making it effectively a thirteen-point circle. Curiously, if instead we
trisect each angle then the six trisectors inside the triangle meet in adjacent pairs to define a regular triangle.
This is known as Morley’s theorem.
Finally let us mention an interesting map that is not an isometry but as we will see later it reverses angles
as if it were an ordinary reflection. In some sense inversions are reflections in a circular mirror.
Definition 4.7. (Audin def 4.1, p.84)
For λ ∈ R and O a point in affine Euclidean plane E, consider the inversion map IO,λ : E − {O} → E − {O}
−−−→ λ −−→
defined by IO,λ (M ) = M 0 where M 0 is the point on line OM such that OM 0 = |OM |2 OM .
11
5 Three dimensional Euclidean geometry
In this section we briefly explore three dimensional Euclidean geometry. Our main focus is the classification of
the regular polyhedra, also known as the Platonic solids. Traditionally they are a centerpiece of mathematics,
for example Euclid’s elements concludes with constructions of the polyhedra and an argument that there can
only be five such. Even though the Platonic solids do not seem to have much significance in themselves, the
tools needed to define, classify and construct them nicely illustrate this kind of geometry.
A new feature in dimension three is that in addition to usual angles between planes one now also has a
notion of ’solid angle’ or a piece of a sphere cut out by a cone. As such we are naturally led to consider the
geometry of the sphere itself.
12
Definition 5.3. (Audin Def 4.1 p.122)
A polyhedron in P ⊂ E is the convex hull of finitely many points in E. It is said to be non-degenerate if the
affine subspace generated by these points is E. A point p ∈ P is called interior if there exists a ball around p
that is also contained in P . The set of all interior points is called P o .
In other words a non-degenerate polyhedron always contains an affine frame and hence the simplex spanned
by it. It also means that there exist interior points. For non-degenerate polyhedra we will give a careful
definition of the vertices, edges and faces. This is not as straightforward as one might think.
Definition 5.4. Define the vertices of a polyhedron P to be the intersection of all T ⊂ E whose convex hull is
P . Any subset of n of the vertices determines a hyperplane H. We say H ∩ P is a face of P if P o ∩ H = ∅.
In case the intersection of two distinct faces contains two vertices we call it an edge. The sets of vertices, edges
and faces are generally denoted by V, E, F .
Even though a polyhedron P may be the convex hull of ten points, it may have fewer vertices, for example
take the eight corners of a regular cube and place two more points inside. The convex hull will still be that
cube and the eight corners will be the vertices. It is true that if P is the convex hull of set S then V is a subset
of S (exercise!).
Our terminology is mostly adapted for dimension three. In two dimensions the sides of a triangle should
now be called faces and there are no edges. In three dimensions the edges are intervals connecting two vertices
(exercise!) but in higher dimensions they will have higher dimension too and more should be said to describe
the combinatorial structure of the polyhedron completely.
Now that we defined faces properly we can say what we mean by a regular polyhedron. Again the definition
we give is mostly tailored to dimension three, in higher dimensions we can ask even more regularity.
Definition 5.5. A non-degenerate convex polyhedron in Euclidean affine space is called regular if all its faces
are isometric to a single regular polygon and if for each pair of vertices a, b there is an isometry taking a to b
that sends any edge adjacent to a to an edge adjacent to b.
Simple examples of regular polyhedra in any dimension are the orthoplexes which are the convex hull of the
unit vectors ±ei in Rn (viewed as affine n space). Dually the n-cubes [−1, 1]n are also regular.
These two examples illustrate a general phenomenon of duality of convex polyhedra. Recall that for triangles
we defined the centroid to be the center of gravity. This notion works for any finite subset S ⊂ E. It is just the
average of the points in S. The centroid of a polyhedron is the centroid of its vertices.
Definition 5.6. The dual of a convex polyhedron P is the convex hull of the centroids of the faces of P .
Definition 5.7. A great circle on S 2 is the intersection of S 2 with a two dimensional linear subspace of E. A
half space defined by vector v ∈ E is the set Hv = {x ∈ E|hv, xi ≥ 0}. A convex spherical polygon is the
intersection of S 2 and finitely many half spaces.
The most important special cases of convex spherical n-gons are: the case n = 1 which is known as a
hemisphere. The case n = 2 is called a time zone (lune) and the case n = 3 is called a spherical triangle.
We think of the great circles as the ’straight lines’ (geodesics). Later in the course we will see that indeed
among all paths that run on the surface of the sphere the shortest path between two points is always an arc of
a great circle. By arc of great circle we just mean a connected part. Unlike Euclidean straight lines, any great
circles intersects and not just in one point but actually in two points. This pair of points is always antipodal
(opposite). There is a fundamental isometry called the antipodal isometry a : E → E defined by a(x) = −x
and we say it sends points x on the sphere to their antipode −x.
We make the assumption that there exists an area function Area : P → R defined on the set P of all P ⊂ S 2
bounded by finitely many segments of great circles4 . We assume Area satisfies the following properties:
1. Area(S 2 ) = 4π
2. If P, Q ∈ P and P ∩ Q consists of segments of great circles then Area(P ∪ Q) = Area(P ) + Area(Q).
paradox.
13
The angle between two great circles is the geometric angle between the two corresponding planes.
Theorem 5.1. (Audin Prop 3.1, p. 121)
The sum of the angles of a spherical triangle is the area of the triangle plus π.
Proof. (Nicer than Audin!). Extend the sides of our spherical triangle ABC to three great circles. Each pair
of great circles will meet in an additional point antipodal to the original, say A0 is antipodal to A and B 0 the
antipode of B and C 0 the antipode of C. The antipodal isometry sends triangle ABC to triangle A0 B 0 C 0 which
means they have equal area. For each vertex of our triangle two of the three great circles intersect and these
two circles bound two time zones ending at the vertex and its antipode. One of these time zones contains ABC
−−→
the other contains A0 B 0 C 0 . The time zones are related by a π rotation with axis AA0 . The measure of the
geometric angle of both the time zones at A is called α ∈ [0, π), the angle at B is called β and the one at C is
called γ.
The total of six time zones covers the sphere completely with a triple overlap precisely at ABC and A0 B 0 C 0 .
Since the area of a time zone with angle µ is 2µ and the area of the sphere is 4π finishes the proof:
V −E+F =2
Proof. Without loss of generality we may assume that all faces of our polyhedron Q ⊂ E are triangular. This is
because we can always divide a face with more corners by adding a diagonal edge. This increases both E and
F by one while not changing the number of vertices. Hence V − E + F is unchanged. By non-degeneracy we
may choose a point O in the interior of our polyhedron Q. Applying the map ΘO : E → E where we wrote the
underlying vector space as E to avoid conflict of notation we now consider a polyhedron P = ΘO (Q) containing
the origin in its interior.
x
The radial projection map ρ : E − {0} → S 2 defined by ρ(x) = |x| sends the vertices, edges and faces of P to
2
a tiling of the sphere S by triangles. More precisely the images of the edges end at the images of the vertices
and their complement on the sphere is a union of spherical triangles.
We calculate in two ways the sum of the (measures of the geometric) angles of all the triangles. Collecting
the angles around each vertex we see that they add up to 2π at each vertex P so we get 2πV in total. On the
other hand we can add the angle sum of each of the triangular faces to get faces f Area(f ) + π = 4π + πF
since the F faces cover S 2 precisely.
Putting it all together we found
2πV = 4π + πF
Since all faces are assumed to be triangles we have 3F = 2E so − F2 = −E + F . Dividing the above by 2π we
find 2 = V − F2 = V − E + F .
Euler’s polyhedron formula restricts the possibilities for regular polyhedra severely.
Lemma 5.1. If in a convex polyhedron all faces have s vertices and around every vertex r edges meet then
{r, s} ⊂ {{3, 3}, {3, 4}, {3, 5}}
Proof. Any edge belongs to two faces so sF = 2E and every edge has two end points so rV = 2E. Solving for
E in Euler’s formula we find 2 = V − E + F = 2r E − E + 2s E so dividing by 2E we find
1 1 1 1
+ = +
E 2 r s
The number of edges is positive so 1r + 1s > 12 . Since r, s are supposed to be integers ≥ 3 there are only finitely
many possibilities to satisfy this equation. Five to be precise as one can see by listing all possibilities.
So far we have shown that if regular polyhedra exist in 3-space then the possible values of s, r are as above.
However we do not yet know if for each of these allowed values there exists such a polyhedron. This is the
content of the final book of Euclid’s Elements:
14
Theorem 5.3. (the Platonic solids)
Up to isometries and dilations there exist precisely five regular polyhedra in Euclidean affine 3-space. They are:
1. The tetrahedron (r, s) = (3, 3)
The existence of tetrahedron, cube and octahedron is not hard to establish explicitly. The construction of
the dodecahedron and icosahedron is more surprising.
The strategy is to start with a cube and place roofs on each of the faces of a cube in a symmetric fashion.
The roof has a square horizontal base and a top line segment parallel to a pair of sides of the cube. The
remaining faces are determined by connecting the vertices of the square to the endpoints of the top segment.
The five edges not part of the square should have equal length. The two dihedral angles that the faces make
with the square base should add up to π/2. The top angles of the triangles should be 2π/5.
The vertices of a dodecahedron are (±1, ±1, ±1), (0, ±τ, ±τ −1 ), (±τ −1 , 0, ±τ ), (±τ, ±τ −1 , 0).
Higher dimensional analogues of the Platonic solids exist in any dimension. We have already met the
simplices and the generalization of the cube is the hypercube [−1, 1]n . Its dual is the analogue of the octahedron
and also exists in any dimension. However dodecahedron and icosahedron are unique to three dimensions. In
dimensions five and up the three regular solids: simplex, hypercube and its dual the orthoplex are the only ones.
In dimension four there are three more that go by the name of 120-cell 600-cell (analogues of dodecahedron
and icosahedron) and something new called the 24-cell. One way to understand their existence is as lifts of the
symmetries of the threedimensional Platonic solids. This uses the fact that SO(3) is doubly covered by the
3-sphere S 3 .
5.5 Quaternions
No discussion of three-dimensions is complete without mention of the quaterions. Much like complex numbers
are very effective in discussing the plane the four-dimensional space of quaternions illuminates the isometry
group of three space. Following Hamilton we boldly construct these new numbers:
Definition 5.8. The quaternions H is a 4-dimensional vector space with basis 1, i, j, k and bilinear, associative
product H × H → H determined by Hamilton’s relations
i2 = j2 = k2 = −1 ij = k jk = i ki = j
qv qu = −v · u + qv×u
15
6 Differential geometry
In the second half of this course we explore what remains of geometry when the space is curved and distances
are distorted. Such spaces will be introduced as Riemannian5 charts6 . Our main concern will be to formulate
what we mean by a ’straight line’ in such a curved space. The curves taking the place of straight lines are
known as geodesics and we will show that they exist at least locally as the curve minimizing the distance from
point A to point B. In doing so we will need a fair amount of techniques from differential calculus including
the existence of solutions to ordinary differential equations and a little variational calculus.
6.1 Derivative
Definition 6.1. A C 1 function is a function whose partial derivatives exist and are continuous. Likewise C 2
means the partial derivatives of the partial derivatives are continuous and so on.
f Pm
Imagine a C 1 function P − → Rm defined on an open subset P ⊂ Rn . Set f = i=1 fi ei for some functions
f 0 (p)
fi : P → R. The derivative of f at p ∈ P is the linear transformation Rn −−−→ Rm whose matrix with respect
∂fi
to the standard bases is ( ∂xj
(p))i=1...m,j=1...n .
In case P is not open we say f is differentiable if there exists a differentiable function on an open set containing
P that coincides with f on P .
The matrix of the derivative with respect to the standard bases is also known as the Jacobian matrix.
∂fi
Its (i, j)-th entry is the j-th partial derivative ∂xj
(p) of the i-th component of f . Even though it is often
enough to work in the standard bases it is important to notice that our derivative is the linear transformation
P ∂fi
corresponding to the Jacobian matrix, not the Jacobian itself. In other words f 0 (p)ej = ∂j f (p) = i ∂x j
(p).
In geometry it is often useful to be able to change to a basis adapted to the situation. Also, phrasing things in
terms of linear transformations makes the formulas easier to manage.
It is traditional to call a function X : P → Rn for P ⊂ Rn a vector field on P . We will take for granted
the following properties of the derivative:
Lemma 6.1. (Properties of the derivative)
f g
1. (Chain rule) For C 1 functions P − → R we have (g ◦ f )0 (p) = g 0 (f (p)) ◦ f 0 (p).
→Q−
f,g
2. For C 1 functions P −−→ Q and α, β ∈ R the function αf + βg is also C 1 .
f,g
3. For C 1 functions P −−→ R the product f g is also C 1
f
→ Q is C 1 if and only if its coefficient functions fi are C 1 where f (p) =
P
4. The function P − i fi (p)ei .
f
Definition 6.2. A C 1 function Rn ⊃ P −
→ Q ⊂ Rn is said to be a diffeomorphism if it is a bijection whose
1
inverse is also C .
Diffeomorphisms are often important as symmetries or coordinate changes because by definition we do not
lose any information applying them to some object. The inverse function theorem provides lots of theoretical
f g
examples of diffeomorphism. One way to state the theorem is that if a C 1 function P −
→Q− → R has invertible
derivative f 0 (p) for some p ∈ P then there are open neighbourhoods p ∈ A ⊂ P and f (p) ∈ B ⊂ Q such that
the restriction of f to A is a diffeomorphism with image B.
f
A geometric way to think about the derivative of a map P −→ Q between open subsets P ⊂ Rn and Q ⊂ Rm
f 0 (p)
is the following. At every point p ∈ P the derivative at p is a linear map Rn −−−→ Rm and to visualize f 0 (p)v for
some v ∈ Rn we use the chain rule. To make the vector v more geometric we interpret it as the velocity vector
of some curve γ : (−, ) → P for some small > 0. So γ(0) = p and γ̇(0) = v. We cannot apply f directly to v
but we can now consider what happens to γ as we apply f . This produces a new curve β = f ◦ γ with image in
Q. The velocity vector w of this new curve β at f (p) = β(0) is a natural geometric way to transport the vector
v at p to a vector w at f (p). In formulas
In the third step we used the chain rule. We leave it as an exercise to the reader that this construction did not
depend on the particular choice of γ.
5 [Link] wrote his PhD thesis on this topic. His advisor was C. Gauss.
6 Usually
one considers Riemannian manifolds (see the course in the 3rd year) instead of charts but for our purposes charts are
sufficient.
16
Another good way to think about the derivative f 0 (p)v is as a directional derivative ∂v f (p):
f (p + hv) − f (p)
∂v f (p) = lim
h→0 h
The chain rule tells us that ∂v f (p) = f 0 (p)v. This is left as an exercise to the reader.
17
Since all geometric properties derive from the metric, isometries can be thought of as those transformations
that preserve shape. As the name suggests the Euclidean linear isometries O(E) from a Euclidean vector space to
itself are examples of the above Riemannian notion of isometry. In this case we have (P, g) = (Q, h) = (Rn , gE ),
the standard inner product in every point.
Other examples of isometries arise when we describe the same object in Euclidean space using two different
Riemannian charts. For example let us make a completely different chart describing (part of) the sphere in R3 .
2
−1)
The stereographic map σ : R2 → S 2 ⊂ R3 is defined as σ(p) = (2p,|p| |p|2 +1 where p = (p1 , p2 ). It has inverse
−1 x 2 2
σ (x, z) = 1−z where x = (x1 , x2 ) ∈ R defined on all of S exept the north pole where z = 1.
The stereographic map σ has injective derivative at every point and is a C 2 function. Therefore we can
use it to turn P = R2 into a Riemannian chart (P, g) with pull-back metric g = σ ∗ gE . Recall we also had the
geographic Riemannian chart (G, ϕ∗ gE ) of the S 2 . The map σ −1 ◦ ϕ when restricted to the correct domain is
an example of an isometry between Riemannian charts.
Another simple example is an isometry of t : H2 → H2 defined by t(x, y) = (x + 1, y). Since t0 (p) is the
identity at any point p we get t∗ ghyp = ghyp . There are many more isometries of H2 and we will get back to
them later.
d
∂i L(γ(t), γ 0 (t), t) = (∂n+i L(γ(t), γ 0 (t), t))
dt
The proof of this theorem depends on a simple lemma
Lemma 6.2. Suppose M is a C 2 function on M : [0, 1] → Rn . If M is such that for all C 2 functions
h : [0, 1] → Rn with h(0) = h(1) = 0 we have
Z 1
M (t)h(t)dt = 0
0
To apply the chain rule in a clean way we introduce W : (−v, v) × R → P × Rn+1 defined by W (, t) =
(V (, t), V 0 (, t), t). Then the derivative with respect to inside the integral is equal to (L ◦ W )0 (0, t)e1 =
18
L0 (W (0, t))W 0 (0, t)e1 by the chain rule7 . Now W 0 (0, t)e1 = h(t)ei + h0 (t)en+i so our integrand from (1) is
(L ◦ W )0 (0, t)e1 = L0 (W (0, t))(h(t)ei + h0 (t)en+i ) = (∂i L)(γ(t), γ 0 (t), t)hi (t) + (∂i+n L)(γ(t), γ 0 (t), t)h0i (t)
Partial integration gets rid of the h0 in the second term (using h(0) = h(1) = 0) and turns our integral into
Z 1 Z 1
0 = H 0 (0) = ∂i L(γ(t), γ 0 (t), t)hi (t) − ∂t (∂i+n L)(γ(t), γ 0 (t), t) hi (t)dt =
M (t)h(t)dt = 0
0 0
Since this must hold for arbitrary choices of h Lemma 6.2 shows M = 0 finishing the proof.
6.4 Geodesics
Our main application of the calculus of variations is to minimize the length of curves in a Riemannian chart.
Such curves will be shown to be solutions to a system of ordinary differential equations known as the geodesic
equation (4).
Definition 6.7. Any C 2 curve in a Riemannian chart that solves the geodesic equation (4) is called a geodesic
(curve).
For technical reasons
R1 it is easier not to directly find curves that minimize the length but rather minimize
the energy S(γ) = 0 |γ 0 (t)|2 dt. By reparametrizing we can assume our curve γ minimizing the length has unit
speed, so |γ̇(t)| = 1. Since 12 = 1 we see γ must also minimize the easier S which uses the squared norm.
R1
For example take a curve in flat 2-space γ : [0, 1] → R2 . We aim to minimize S(γ) = 0 |γ 0 (t)|2 dt so
L(x, y, t) = y12 + y22 .
The Euler-Lagrange equations we need to solve are: γ 00 (t) = 0 since
d d
(∂2+i L(γ(t), γ 0 (t), t) = (2γi0 (t)) = 2γi00 (t) = 0 = ∂i L(γ(t), γ 0 (t), t)
dt dt
It follows that γ(t) = at + b for some a, b ∈ R2 .
Moving on to the general case, consider a curve γ : [0, 1] → Rn . For the argument P it is important to allow
n
all curves, not just those in the Riemannian chart (P, g) with P ⊂ Rn . Writing γ(t) = i=1 γi ei we find
X X
|γ 0 (t)|2 = g(γ(t))(γi0 (t)ei , γj0 (t)ej ) = gij (γ(t))γi0 (t)γj0 (t)
i,j i,j
R1 R1P Pn
In this case we seek to minimize S(γ) = 0 |γ 0 (t)|2 dt = 0 i,j gij (γ(t))γi0 (t)γj0 (t)dt so L(x, y, t) = i,j=1 gij (x)yi yj
is a function from R2n+1 to R.
The Euler-Lagrange equations now read for any fixed k ∈ {1, . . . n}:
d
∂k L(γ(t), γ 0 (t), t) = (∂n+k L(γ(t), γ 0 (t), t)) (2)
dt
Pn ∂gij
The left hand side is i,j=1 ( ∂xk )yi yj . For the right hand side we first compute
n
X n
X n
X
∂n+k L(x, y, t)) = gkj (x)yj + gik (x)yi = 2 gik (x)yi
j=1 i=1 i=1
P ∂gik 0 0
P ∂gjk 0 0
Using symmetry to interchange the dummy indices i, j we get i,j ∂xj (γ(t))γi (t)γj (t) = i,j ∂xi (γ(t))γi (t)γj (t)
Putting it all together we found a differential equation for γ:
n n
X X ∂gjk ∂gik ∂gij
0=2 gks (γ(t))γs00 (t) + (γ(t)) + (γ(t)) − (γ(t)) γi0 (t)γj0 (t) (3)
s=1 i,j=1
∂xi ∂xj ∂xk
7 Here e1 refers to the first coordinate in (−v, v) × [0, 1], which is . Also W 0 (0, t)e1 = ∂ W (, t)|=0 .
19
It is convenient to isolate the second derivative term γs00 (t). This can be achieved since the matrix (gij (γ(t)))
describing the metric is positive definite and symmetric. It follows that it has only positive eigenvalues and hence
positive determinant. To see this pick an eigenvector v with eigenvalue λ and write v T (gij (γ(t)))v = v T λv =
2 −1 P −1
λ|v| > 0. It thus has an inverse (grk (γ(t))) so k grk (γ(t))gks (γ(t)) = δrs (Kronecker delta). Applying this
to both sides of our differential equations we get the geodesic equation. For fixed r ∈ {1, . . . n} we have:
n
X
0 = γr00 (t) + Γrij (γ(t))γi0 (t)γj0 (t) (4)
i,j=1
For any value of i, j, r ∈ {1, . . . n} the function Γrij : P → R is known as the Christoffel symbols and is defined
as:
n
1 X −1 ∂gjk ∂gik ∂gij
Γrij (p) = grk (p) (p) + (p) − (p) (5)
2 ∂xi ∂xj ∂xk
k=1
For example the eight Christoffel symbols for the sphere with geographic coordinates are Γ1ij = 0 except
Γ122 = − cos µ sin µ. Γ211 = Γ222 = 0 and Γ212 = Γ221 = cot µ = cos µ
sin µ . So the geodesic equations for γ = (γ1 , γ2 ) is
γ̈1 − cos µ sin µγ̇2 γ̇2 = 0 = γ̈2 + 2 cot µγ̇1 γ̇2 . While a little complicated at first sight it should be clear that the
meridians γ(t) = (t, c) for some constant c ∈ R satisfy this equation.
Returning to the general case: the fundamental theorem on the existence and uniqueness of ordinary differ-
ential equations assures us that geodesics always exist. Intuitively the next theorem states that in a Riemannian
chart one can start walking ’straight’ in any direction at any point of the chart.
Theorem 6.2. (Existence and uniqueness of geodesics)
In a Riemannian chart (P, g) any p ∈ P ⊂ Rn and v ∈ Rn determinesPn a unique C 2 curve γ : [C, D] → P
00 0 0
such that γ(0) = p and γ̇(0) = v and for all m we have γm (t) + i,j=1 Γm
ij γi (t)γj (t). The domain [−D, D] is
not quite unique but if we have another curve γ̃ with the same properties and domain [A, B] then γ = γ̃ on
[A, B] ∩ [C, D].
We arrived at geodesics by attempting to minimize the distance between two points but what we found is
actually more like the curves that are ’straight’. Depending on global issues travelling a straight line may or
may not actually minimize the distance travelled.
The key to the proof is the relation between the Christoffel symbols and the second order partial derivatives:
Lemma 6.3.
n
X 1
h∂i ∂j φ(p), ∂k φ(p)i = Γrij (p)gkr (p) = (∂i gjk + ∂j gki − ∂k gij )
r=1
2
Proof. Dropping the evaluation at p for brevity the right hand side can be expanded using (5):
1X −1 1X 1
gkr gr` (∂i gj` + ∂j g`i − ∂` gij ) = δk` (∂i gj` + ∂j g`i − ∂` gij ) = (∂i gjk + ∂j gki − ∂k gij )
2 2 2
r,` `
20
P −1
Here we used the defining equation for the inverse of g, namely r gkr gr` = δk` .
Writing the dot product with a · and using the product rule for the dot product we have: ∂i gjk = ∂i (∂j φ ·
∂k φ) = (∂i ∂j φ) · ∂k φ + ∂j φ · (∂i ∂k φ). write this equation three times while cyclically permuting the indices i, j, k:
∀k : hβ̈(t), ∂k φ(γ(t))i = 0
Pn
By the chain rule β̇(t) = (φ ◦ γ)0 (t) = φ0 (γ(t))γ̇(t) = i=1 ∂i φ(γ(t))γ̇i (t). Differentiating once more using both
product rule and chain rule we find
n
X n
X
β̈(t) = ∂j ∂i φ(γ(t))γ̇i (t)γ̇j (t) + ∂i φ(γ(t))γ̈i (t)
i,j=1 i=1
Taking the inner product with ∂k φ(γ(t)) on both sides and abbreviating p = γ(t) we can make use of our lemma
as follows
n
X n
X
hβ̈(t), ∂k φ(γ(t))i = h∂j ∂i φ(p), ∂k φ(p)iγ̇i (t)γ̇j (t) + h∂s φ(p), ∂k φ(p)iγ̈i (t) =
i,j=1 s=1
n n
X 1 X
(∂i gjk + ∂j gki − ∂k gij )γ̇i (t)γ̇j (t) + gsk (p)γ̈s (t) = 0
i,j=1
2 s=1
Up to a factor 2 the final line is precisely equation (3) which is equivalent to the geodesic equation.
Lemma 6.3 actually gives us more intuition about the Christoffel symbols. They express the projection of
the second derivative in terms of the basis ∂k φ(p) of the tangent space Tφ(p) . More precisely, for any i, j we
Pn
have ∂i ∂j φ(p) = k=1 Γkij (p)∂k φ(p) + Nij (p) for some vectors Nij (p) perpendicular to Tφ(p) . If we denote by
πq the orthogonal projection of Rm onto the linear subspace Tq then we may reformulate our lemma as
n
X
πφ(p) ∂i ∂j φ(p) = Γkij (p)∂k φ(p) (9)
k=1
21
For example in P = R2 with X(x, y) =(2x + y 2 , x) 3 2
and Y (x, y) = (3y + x , y ) we get ∂X (Y )(x, y) =
2
3x 3 2x+y 2 3 2 2
∂(2x+y2 ,x) Y (x, y) = Y 0 (x, y)(2x + y 2 , x) = = 6x +3x y +3x
x 2xy . More generally if X =
0 2y
Pn Pn P
i=1 Xi ei and Y = j=1 Yj ej then (∂X Y )(p) = i,j Xi (p)(∂i Yj )(p)ej . Another way to think about (∂X Y )(p)
d
is that we find an integral curve γ for X so γ(0) = p and γ̇(0) = X(p). Then (∂X Y )(p) = dt Y (γ(t)).
Our new operation on vector fields can be used to give an important criterion for space being ’flat’.
Lemma 6.4. For any three vector fields X, Y, Z on Rn we have
Proof. To see why we simply expand the definitions, for brevity we drop the evaluation at p from our notation:
X X
∂X (∂Y Z) = Xk ∂k (Yi (∂i Zj ))ej = Xk ∂k Yi + Yi ∂k ∂i Zj ej
i,j,k i,j,k
As we saw in the geodesic equation the proper generalization of ∂X Y to an arbitrary Riemannian chart is
P known as the Levi-Civita connection ∇X Y . It turns out to be related to
called the covariant derivative also
the Christoffel symbols by ∇ei ej = ij Γkij ek . We will see that in this notation the geodesic equation would be
∇γ̇ γ̇ = 0.
2. ∇f X+Z (Y ) = f ∇X Y + ∇Z Y
3. g(p)(∇X Y (p), Z(p)) + g(p)(Y (p), ∇X Z(p)) = ∂X(p) G(p), where G(s) = g(s)(Y (s), Z(s))
4. ∇X Y − ∇Y X = [X, Y ]
In (Rn , gE ) theP
formula ∇X Y = ∂X PYn defines an LC-connection. In checking the first two axioms the hardest
n
bit is ∇X (f Y ) = i=1 Xi ∂i (f Y ) = i=1 Xi (∂i f )Y + Xi f ∂i Y = (∂X f )Y + f ∇X Y . The third axiom is true
because the partial derivative satisfies a product rule with respect to the standard dot product. The last axiom
was our definition of the commutator so there is nothing to check.
While a little abstract this definition is encoding exactly what we were doing before. In the theorem below
we will see the Christoffel symbol comes out as the coefficients of ∇ei ej so ∇ei ej = i,j,k Γkij ek . Assuming this
P
for a moment the geodesic equation can be rewritten as ∇γ̇ γ̇ = 0 withPnthe caveat that we need to extend γ̇ to a
vector field V such that V (γ(t)) = γ̇(t) for all t. Indeed using V = j=1 Vj ej the properties of the connection
give us
X X X d X X
∇γ̇ γ̇ = γ̇i ∇ei Vj ej = γ̇i (∂i Vj )ej + γ̇i Vj ∇ei ej = Vk (γ(t))ek + Γkij γ̇i Vj ek = γ̈k (t)+ Γkij γ̇i Vj ek
i,j i,j j
dt
i,j,k i,j,k
What is important is that the Riemannian metric g determines the LC-connection uniquely, much like our
discussion in the previous subsection.
Theorem 6.4. (Fundamental lemma of Riemannian geometry)
unique Levi-Civita connection ∇ on Riemannian chart (P, g).
There exists aP
Also ∇ei ej = i,j,k Γkij ek .
Proof. For simplicity we write the Riemannian metric as h·, ·i and do not explicitly write its dependence on the
point. Condition 3) is then abbreviated as
22
We start by proving uniqueness of the LC-connection by proving it is determined entirely by g much like
the proof of Lemma 6.3. For vector fields X, Y, Z we write Equation (10) three times cycling around X, Y, Z:
Subtracting the third equation from the sum of the first two we can solve for h∇X Y, Zi to get
1
h∇X Y, Zi = ∂X hY, Zi + ∂Y hZ, Xi + ∂Z hX, Y i − hY, [X, Z]i − hZ, [Y, X]i + hX, [Z, Y ]i (11)
2
Uniqueness follows from this equation since if we had two LC-connections ∇, ∇˜ then for all Z we would find
˜ ˜
h∇X Y, Zi = h∇X Y, Zi which is only possible if ∇X Y = ∇X Y . Existence follows from the same equation (11)
because taking X, Y, Z = ei , ej , ek we find
n
1 X
h∇ei ej , ek i = ∂i gjk + ∂j gki − ∂k gij = Γrij (p)gkr (p)
2 r=1
using [ea , eb ] = 0 and the definition of the Christoffel symbols (5). Using the inverse g −1 we thus find ∇ei ej =
P n k
k=1 Γij ek as desired. This proves that an LC-connection exists because by properties 1) and 2) we have
X X X X
∇X Y = Xi ∇ei (Yj ej ) = Xi ∂i Yj ej + Xi Yj ∇ei ej = Xi ∂i Yj ej + Xi Yj Γkij
i,j i,j i,j k
When our Riemannian metric comes from pulling back the Euclidean metric along a suitable map φ : P → Rm
as in the previous subsection then the LC-connection is easier to understand. It is basically using the LC-
connection in (Rm , gE ) given by Definition 6.8 and then projecting the result onto the tangent space in Rm .
More precisely set g = φ∗ gE to be a metric on P ⊂ Rn and φ : P → Rm as in the previous subsection. We
claim that the LC-connection in that case is given by φ0 ∇X Y = π∇X̃ Ỹ where π(p) : Rm → Rm is the linear
map that projects orthogonally onto the tangent space Image(φ0 (p)) and X̃ and Ỹ are vector fields on Rm such
that X̃ ◦ φ = φ0 X and likewise Ỹ ◦ φ = φ0 Y . Also recall that by φ0 X we mean the vector field on Image(φ)
defined by φ0 X(φ(p)) = φ0 (p)X(p). This means that on the image of φ we get
0
∇X̃ Ỹ (φ(p)) = Ỹ (φ(p)) X̃(φ(p)) = (φ0 Y )0 (p)X(p) = ∂X (φ0 Y )(p)
Pn
When X = ei and Y = ej we get φ0 Y = ∂j φ so that by equation (9) φ0 ∇ei ej = π∂i ∂j φ = k=1 Γkij (p)∂k φ Since
Pn
φ0 ek = ∂k φ we conclude ∇ei ej = π∂i ∂j φ = k=1 Γkij (p)ek .
Lemma 6.5. (Isometry invariance)
f
For any isometry P − → Q between Riemannian charts we have f 0 (∇X Y ) = ∇f 0 X f 0 Y . Also γ is a geodesic if
and only if f ◦ γ is a geodesic.
Proof. Define a map ∆ : Vec(P ) × Vec(P ) → Vec(P ) by ∆X Y = (f −1 )0 ∇f 0 X f 0 Y . If we can show that ∆ is an
LC-connection on P then by uniqueness ∇ = ∆ so ∇X Y = (f −1 )0 ∇f 0 X f 0 Y proving f 0 (∇X Y ) = ∇f 0 X f 0 Y .
Next, if γ is a geodesic in P and β = f ◦ γ then ∇γ̇ γ̇ = 0. Applying f 0 to both sides yields 0 = f 0 ∇γ̇ γ̇ =
∇f 0 γ̇ f 0 γ̇ = ∇β̇ β̇, showing β is a geodesic too. Conversely if β is a geodesic then so is γ because f −1 is an
isometry too.
We we have seen that for the LC-connection on (Rn , gE ) we had ∇X ∇Y Z − ∇Y ∇X Z − ∇[X,Y ] Z = 0 For
other Riemannian charts this is not at all the case.
Definition 6.10. (Riemannian Curvature)
The Riemann curvature of n-dimensional Riemannian chart (P, g) is the map R : V ec(P )×V ec(P )×V ec(P ) →
V ec(P ) defined by
R(X, Y )Z = ∇X ∇Y Z − ∇Y ∇X Z − ∇[X,Y ] Z
It follows that curvature is invariant under isometries in the following sense
Lemma 6.6. (Isometry invariance of curvature)
f
→ Q between Riemannian charts we have f 0 (R(X, Y )Z) = R(f 0 X, f 0 Y )f 0 Z
For any isometry P −
23
Proof. Using Lemma 6.5
Concretely the coefficients of the Riemann curvature can be computed in terms of Christoffel symbols.
Lemma 6.7. The curvature is linear in X, Y, Z and for C 2 -functions
P ` f : P → R we have f R(X, Y )Z =
R(X, Y )f Z = R(X, f Y )Z = R(f X, Y )Z. If we set R(ei , ej )ek = ` Ri,j,k e` then
X
`
Ri,j,k = ∂i Γ`jk − ∂j Γ`ik + Γrjk Γ`ir − Γrik Γ`jr
r
Proof. Linearity in X, Y, Z follows directly from the properties of the LC-connection. The property with function
f is left to the reader as an exercise.
Notice [ei , ej ] = 0 so R(ei , ej )ek = ∇ei ∇ej ek − ∇ej ∇ei ek . Now
X X X X X
∇ei ∇ej e k = ∇ei Γrjk er = ∂i Γrjk er + Γrjk Γsir es = ∂i Γ`jk + Γrjk Γ`ir e`
r r r,s ` r
24