0% found this document useful (0 votes)
19 views191 pages

Tensor Algebra for Mechanics Students

This textbook is designed for graduate students and young researchers in mechanics, focusing on tensor algebra and analysis with applications in differential geometry. It aims to provide a modern introduction to the mathematical tools necessary for continuum mechanics, covering essential topics such as second and fourth-rank tensors, differential geometry of curves and surfaces, and curvilinear coordinates. The book includes over a hundred exercises to support learning, with a clear and concise presentation tailored for the scientific community in mechanics.

Uploaded by

Aryavart Diploma
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
19 views191 pages

Tensor Algebra for Mechanics Students

This textbook is designed for graduate students and young researchers in mechanics, focusing on tensor algebra and analysis with applications in differential geometry. It aims to provide a modern introduction to the mathematical tools necessary for continuum mechanics, covering essential topics such as second and fourth-rank tensors, differential geometry of curves and surfaces, and curvilinear coordinates. The book includes over a hundred exercises to support learning, with a clear and concise presentation tailored for the scientific community in mechanics.

Uploaded by

Aryavart Diploma
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

A short course on

Tensor algebra and analysis for


arXiv:2111.14633v9 [[Link]] 10 Jan 2025

engineers
with applications to differential geometry

Paolo Vannucci
[Link]@[Link]
To Carla, Bianca and Alessandro

ii
Preface

This textbook, addressed to graduate students and young researchers in mechanics, has
been developed from the class notes of different courses in continuum mechanics that I have
been delivering for several years as part of the Master’s program in MMM - Mathematical
Methods for Mechanics, at the University Paris-Saclay.

Far from being exhaustive, as any primer text, the intention of this book is to introduce
students in mechanics and engineering to the mathematical language and tools that are
necessary for a modern approach to continuum theoretical and applied mechanics. The
presentation of the matter is hence tailored for this scientific community and necessar-
ily it is different, in terms of language and objectives, from that normally proposed to
students of other disciplines, such as physics, especially general relativity, or pure math-
ematics.

What has motivated me to write this textbook is the idea of collecting in a single, in-
troductory book a set of results and tools useful for studies in mechanics and presenting
them in a modern, succinct way. Almost all the results and theorems are proved and
the reader is guided along a tour that starts from vectors and ends at the differential
geometry of surfaces, passing through the algebra of tensors of second and fourth orders,
the differential geometry of curves, the tensor analysis for fields and deformations and the
use of curvilinear coordinates.

Some topics are specially treated, such as rotations, the algebra of fourth order tensors,
fundamental for the mechanics of modern materials, or the properties of differential op-
erators. Some other topics are intentionally omitted because they are less important to
continuum mechanics or too advanced for an introductory text. Though, in some modern
texts, tensors are directly presented in the most general setting of curvilinear coordinates,
I preferred here to choose a more traditional approach, introducing first tensors in Carte-
sian coordinates, normally used for classical problems. Then, an entire chapter is devoted
to the passage to curvilinear coordinates and to the formalism of co- and contravariant
components.

The tensor theory and results are specially applied to introduce some subjects concerning
differential geometry of curves and surfaces. Also in this case, the presentation is mainly
intended for applications to continuum mechanics and, in particular, in view of courses
on slender beams or thin shells. All the presentation of the topics of differential geometry
is extensively based on tensor algebra and analysis.

iii
More than a hundred exercises are proposed to the reader, many of them completing the
theoretical part through new results and proofs. All the exercises are entirely developed
and solved at the end of the book, in order to provide the reader with thorough support
for his learning.
In Chapter 1, vectors and points are introduced and also, with a small anticipation of some
results of the second Chapter, also applied vectors are visited. Chapter 2 is completely
devoted to the algebra of second-rank tensors and the succeeding Chapter 3 to that of
fourth-rank tensors. Intentionally, these are the only two types of tensors introduced
in the book: They are the most important types of tensors in mechanics, by them we
can represent deformation, stress and the constitutive laws. I preferred not to introduce
tensors in an absolutely general way but to go directly to the most important tensors for
applications in mechanics; for the same reason, the algebra of other tensors, namely of
third-rank tensors, is not presented in this primer text.
The analysis of tensors is first done using differential geometry of curves, in Chapter 4,
for differentiation and integration with respect to only one variable, then introducing the
differential operators for fields and deformations, in Chapter 5.
Then, a generalization of second-rank tensor algebra and analysis in the sense of the use
of curvilinear coordinates is presented in Chapter 6, where the notion of metric tensor,
co- and contravariant components and Christoffel’s symbols are introduced.
Finally, Chapter 7 is entirely devoted to an introduction to the differential geometry of
surfaces. Classical topics such as the first and second fundamental forms of a surface, the
different types of curvatures, the Gauss-Weingarten equations or the concepts of minimal
surfaces, geodesics and the Gauss-Codazzi conditions are presented, with all these topics
being of a great interest in mechanics.
I tried to write a coherent, almost self-contained manual of mathematical tools for gradu-
ate students in mechanics, with the hope of helping young students progress in their stud-
ies. The exposition is as simple as possible, sober, sometimes minimalist. I intentionally
avoided burdening the language and the text with nonessential details and considerations,
but I have always tried to grasp the essence of a result and its usefulness.
It is my most sincere hope that the reader who dares to persevere through the pages of
this book will find a benefit for his studies in continuum mechanics. This is, eventually,
the goal of this primer text.

Versailles, June 16, 2022

iv
Acknowledgments

I am indebted to many persons for the topics of this book. Professor E. G. G. Virga,
University of Pavia, introduced me, the first, to tensor algebra, during my PhD at the
University of Pisa, many years ago.
Then, I have had the privilege of collaborating with Professor G. Verchery, at the Univer-
sity of Burgundy, who introduced me to the representation methods based upon tensor
invariants. This has been very useful for developing some results in the algebra of fourth
rank tensors.
I wish also to thank Professor P. M. Mariano, University of Florence, who always pushed
me to go forward and to consider problems of modern mechanics; many of the discussions
I had with him have been very important to me.
I have also had many interesting and useful discussions with Professor J. Lerbet, Uni-
versity of Evry, and with Doctor C. Fourcade, of Renault S.A.; I wish to thank them
sincerely.
I am also grateful to Prof. A. Frediani, who has been for me an example of honesty in
teaching and science, and I cannot forget Prof. P. Villaggio, my PhD director, who passed
by some years ago: his teaching and personality leaved an indelebile trace in all my life
of scientist.
Finally, I wish to thank my wife, Carla, my daughter, Bianca, and my son, Alessan-
dro. Without them, nothing would have been possible; because of them, many things
happen.

v
vi
List of symbols

:= : definition symbol
|: such that
∃!: exists and is unique
R: set of real numbers
E: ordinary 3D Euclidean space
V: vector space of translations, associated with E
Lin(V): vector space of second-rank tensors
Lin(V): vector space of fourth-rank tensors
x, y, z etc.: scalars (elements of R)
p, q, r etc.: points (elements of E)
u, v, w etc.: vectors (elements of V)
u = p − q etc.: vector difference of two points of E
L, M, N etc.: second-rank tensors (elements of Lin(V))
L, M, N etc.: fourth-rank tensors (elements of Lin(V))
u · v etc.: scalar product of vectors
L · M etc.: scalar product of second-rank tensors
u × v etc.: cross product of vectors
u ⊗ v etc.: dyad (second-rank tensor) of vectors
u = |u|: norm of a vector
L = |L|: norm of a tensor
|p − q|, |u − v|, |L − M|: distance of points, vectors or tensors
det L: determinant of a second-rank tensor L
L⊤ : transpose of a second-rank tensor L
L−1 : inverse of a second-rank tensor L
L−⊤ = (L⊤ )−1 = (L−1 )⊤
Sym(V): subspace of Lin(V) of symmetric second-rank tensors
Skw(V): subspace of Lin(V) of skew second-rank tensors
Sph(V): subspace of Lin(V) of spherical second-rank tensors
Dev(V): subspace of Lin(V) of deviatoric second-rank tensors
Orth(V): subspace of Lin(V) of orthogonal second-rank tensors

vii
Orth(V)+ : subspace of Lin(V) of rotation second-rank tensors
S: unit sphere of V: S = {u ∈ V| |u| = 1}
δij : Kronecker’s delta
ℜ: real part of a complex quantity
ℑ: imaginary part of a complex quantity
L = A ⊠ B: conjugation product of two second-rank tensors
Ssph : spherical projector
Ddev : deviatoric projector
Ttrp : transpose projector
Ssym : symmetry projector
Wskw : antisymmetry projector
I: identity of Lin(V)
Is : restriction of I to symmetric tensors of Lin(V)

viii
Contents

Preface iii

Acknowledgments v

List of symbols vii

1 Points and vectors 1


1.1 Points and vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1
1.2 Scalar product, distance, orthogonality . . . . . . . . . . . . . . . . . . . . 3
1.3 Basis of V, expression of the scalar product . . . . . . . . . . . . . . . . . . 4
1.4 Applied vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
1.5 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11

2 Second rank tensors 13


2.1 Second-rank tensors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13
2.2 Dyads, tensor components . . . . . . . . . . . . . . . . . . . . . . . . . . . 14
2.3 Tensor product . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15
2.4 Transpose, symmetric and skew tensors . . . . . . . . . . . . . . . . . . . . 15
2.5 Trace, scalar product of tensors . . . . . . . . . . . . . . . . . . . . . . . . 17
2.6 Spherical and deviatoric parts . . . . . . . . . . . . . . . . . . . . . . . . . 18
2.7 Determinant, inverse of a tensor . . . . . . . . . . . . . . . . . . . . . . . . 19
2.8 Eigenvalues and eigenvectors of a tensor . . . . . . . . . . . . . . . . . . . 23
2.9 Skew tensors and cross product . . . . . . . . . . . . . . . . . . . . . . . . 25
2.10 Orientation of a basis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30
2.11 Rotations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30
2.12 Reflexions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 41
2.13 Polar decomposition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42
2.14 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 44

3 Fourth rank tensors 47


3.1 Fourth-rank tensors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 47
3.2 Dyads, tensor components . . . . . . . . . . . . . . . . . . . . . . . . . . . 48
3.3 Conjugation product, transpose, symmetries . . . . . . . . . . . . . . . . . 49
3.4 Trace and scalar product of fourth-rank tensors . . . . . . . . . . . . . . . 51
3.5 Projectors, identities . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 52
3.6 Orthogonal conjugator . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 54

ix
3.7 Rotations and symmetries . . . . . . . . . . . . . . . . . . . . . . . . . . . 56
3.8 The Kelvin formalism . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 57
3.9 The polar formalism for plane tensors . . . . . . . . . . . . . . . . . . . . . 59
3.10 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 60

4 Tensor analysis: curves 63


4.1 Curves of points, vectors and tensors . . . . . . . . . . . . . . . . . . . . . 63
4.2 Differention of curves . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 64
4.3 Integral of a curve of vectors and length of a curve . . . . . . . . . . . . . . 67
4.4 The Frenet-Serret basis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 70
4.5 Curvature of a curve . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 72
4.6 The Frenet-Serret formulae . . . . . . . . . . . . . . . . . . . . . . . . . . . 74
4.7 The torsion of a curve . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 75
4.8 Osculating sphere and circle . . . . . . . . . . . . . . . . . . . . . . . . . . 77
4.9 Evolute, involute and envelopes of plane curves . . . . . . . . . . . . . . . 78
4.10 The theorem of Bonnet . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 80
4.11 Canonic equations of a curve . . . . . . . . . . . . . . . . . . . . . . . . . . 81
4.12 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 82

5 Tensor analysis: fields 87


5.1 Scalar, vector and tensor fields . . . . . . . . . . . . . . . . . . . . . . . . . 87
5.2 Differentiation of fields, differential operators . . . . . . . . . . . . . . . . . 87
5.3 Properties of the differential operators . . . . . . . . . . . . . . . . . . . . 89
5.4 Theorems on fields . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 94
5.5 Differential operators in Cartesian coordinates . . . . . . . . . . . . . . . . 97
5.6 Differential operators in cylindrical coordinates . . . . . . . . . . . . . . . . 98
5.7 Differential operators in spherical coordinates . . . . . . . . . . . . . . . . 102
5.8 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 104

6 Curvilinear coordinates 107


6.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 107
6.2 Curvilinear coordinates, metric tensor . . . . . . . . . . . . . . . . . . . . . 107
6.3 Co- and contra-variant components . . . . . . . . . . . . . . . . . . . . . . 110
6.4 Spatial derivatives of fields in curvilinear coordinates . . . . . . . . . . . . 116
6.5 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 119

7 Surfaces in E 121
7.1 Surfaces in E, coordinate lines, tangent planes . . . . . . . . . . . . . . . . 121
7.2 Surfaces of revolution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 123
7.3 Ruled surfaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 125
7.4 First fundamental form of a surface . . . . . . . . . . . . . . . . . . . . . . 126
7.5 Second fundamental form of a surface . . . . . . . . . . . . . . . . . . . . . 127
7.6 Curvatures of a surface . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 130
7.7 The theorem of Rodrigues . . . . . . . . . . . . . . . . . . . . . . . . . . . 132
7.8 Classification of the points of a surface . . . . . . . . . . . . . . . . . . . . 134
7.9 Developable surfaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 135

x
7.10 Points of a surface of revolution . . . . . . . . . . . . . . . . . . . . . . . . 137
7.11 Lines of curvature, conjugated directions, asymptotic directions . . . . . . 138
7.12 The Dupin’s conical curves . . . . . . . . . . . . . . . . . . . . . . . . . . . 139
7.13 The Gauss-Weingarten equations . . . . . . . . . . . . . . . . . . . . . . . 140
7.14 The Theorema Egregium . . . . . . . . . . . . . . . . . . . . . . . . . . . . 142
7.15 Minimal surfaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 143
7.16 Geodesics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 145
7.17 The Gauss-Codazzi compatibility conditions . . . . . . . . . . . . . . . . . 149
7.18 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 154

Suggested texts 157

Solutions of the exercises 159

xi
xii
Chapitre 1
Chapter E1LEMENTS D’ALGEBRE TENSORIELLE

Points and vectors


1.1 ESPACE EUCLIDIEN
Les événements de la mécanique classique se placent dans l'espace euclidien à trois dimensions,
que nous définissons ainsi: on dit que E est un espace euclidien tridimensionnel s’il existe un
1.1vectoriel
espace PointsV, quiandlui est vectors
associé, de dimension trois, dans lequel il est défini un produit
scalaire, et tel que:
We consider in the following a point space E, whose elements are points p. In classical
• mechanics,
les élémentsEvisdetoV,bequi sont deswith
identified vecteurs, sont des transformations
the Euclidean de Espace,
three-dimensional en lui-même:
wherein events
are intended to v ∈ V, v : E → E ;
be set. On E, we admit the existence of an operation, the difference of
any couple of its elements:
• la somme de deux éléments de V est définie q − p,comme
p, q ∈ E.
(u + v)( p ) = u( v( p)) ∀ u, v ∈ V et ∀ p ∈ E ;
We associate to E a vector space V whose dimension is dimV = 3 and whose elements are
• vectors
∀ p et qv∈representing
E , ∃! v ∈ V translations
: q = v(p). over E:
Pour mieux comprendre tout cela,∀p,
il faut
q ∈ d’abord
E, ∃! v introduire
∈ V| q − pdeux
= [Link] assez importants.

Any element v ∈ V is hence a transformation over E that, using the previous definition,
1.2 POINTS ET VECTEURS
can be written as :
Nous choisissons une fois pour toutes un espace euclidien E; ses éléments sont appelés points. E
doit être identifié avec l'espace
∀v ∈ordinaire
V, v : E où
→ nous
E| q vivons.
= v(p) → q = p + v.
L'espace vectoriel V sera appelé espace des translations de E et les éléments de V seront appelés
We remark that the result of the application of the translation v depends upon the
translations.
argument p:
Analysons donc les propriétés énoncées
q = pci-dessus;
+ v ̸= p1on
+vcommence
= q1 , avec la dernière. Ecrire q = v(p)
signifie que v est une transformation de E en lui-même, c’est à dire, on part d’un point de E pour
whose geometric meaning is depicted in Fig. 1.1. Unlike difference, the sum of two points
arriver encore en un point de E, et que cette transformation est intégralement déterminée par la
is not defined and is meaningless.
valeur prise sur un point de E. Graphiquement:
q
q’1
v
v

p
p’1

Figure 1.1: Same translationFigure 1.1 to two different points.


applied
Remarque: le même vecteur peut opérer différentes transformations, en fonction du point
d’application: q = v(p), mais aussi q’= v(p’). 1
Nous utiliserons, à la place de l'écriture q = v(p), une écriture qui a un sens géométrique plus direct:
q= p + v.
Elle définit la somme d’un point et d’un vecteur comme un point. De la relation ci-dessus on tire
We define the sum of two vectors u and v as the vector w such that
(u + v)(p) = u(v(p)) = w(p)
This means that, if
q = v(p) = p + v,
then
r = u(q) = q + u = w(p),
see Fig. 1.2, which shows that the above definition actually coincides with the parallelo-
gram rule and that Chapitre 1
u + v = v + u,
as obvious, for the sum over a vector space commutes. It is evident that the sum of more
than two vectorsv= q −[Link] iteratively, summing up a vector at a time to the sum of
can
La
thesomme de deux
previous points, ainsi que la différence d’un point avec un vecteur, ne sont pas définies.
vectors.
On
Therevient maintenant
null vector à la deuxième
o is defined as the propriété:
difference of any two coincident points:
(u + v)( p) = u( v ( p))o :=∀p u,
− vp∈∀p
V ∈ ∀ p∈E .
et E;
Soit
o is qunique
= v(p),and
ou q=
thep only
+ v, etvector
soit r=such
u(q),that
ou r= q + u. Alors,
r= p + v + u = p + w,
v + o = v ∀v ∈ V.
où w est le vecteur formé par la somme de u et de v. Graphiquement tout cela correspond à la
fameuse règle du parallélogramme:
q
u
v
r
w
p
v
u

Figure 1.2
Figure 1.2: Sum of two vectors: the parallelogram rule.
Remarque: par les propriétés générales d’un espace vectoriel, ou plus simplement
géométriquement, à l'aide de la figure ci-dessus, on a:
In fact,
v + u = u + v.
∀p ∈ E, v + o = v + p − p → p + v + o = p + v ⇐⇒ v + o = v.
En particulier, faire u + v équivaut à faire le chemin pointillé indiqué sur la figure 1.2.
1
Le vecteurcombination
A linear nul o est défini comme vlai isdifférence
of n vectors defined asdethe
deux points
vector coïncidents. Le vecteur nul est
unique, et il est le seul vecteur tel que
w := ki vi , ki ∈ R, i = 1, ..., n.
v + o= v ∀ v ∈ V.
The n + 1 vectors w, vi , i = 1, ..., n, are said to be linear independent if there does not
Ces
existdeux
a setpropriétés du kvecteur
of n scalars nul sont
i such that très facilement
the above equation démontrables avec
is satisfied and arelasaid
propriété que l’on a
to be linear
expliquée
dependentci-dessus.
in the opposite case.
Un 1vecteur
We adoptw here
tel que
and in the following the Einstein notation for summations: All the times when an
index is repeated in a monomial,
n P
then the summation with respect to that index, called the dummy index,
is understood, e.g.,wki=
∑ i ∈ R , i = 1, 2 ,..., n ,
vi = k ui k, i vi .k We
i i
then say that the index i is saturated. If a repeated index is
underlined, then it is noti =a1 dummy index, i.e. there is no summation.

est dit être une combinaison linéaire des n vecteurs ui, où les scalaires ki sont les coefficients de la
2
combinaison. Si, pour un w et pour les n ui donnés il n’existe aucun ensemble de ki tel que la
relation ci-dessus soit satisfaite, alors les n+1 vecteurs w et ui sont dits linéairement indépendants ;
cela signifie qu’il n’est pas possible d’exprimer w comme somme des ui, où, ce qui est la même
chose, que la combinaison linéaire des n+1 vecteurs w et ui peut avoir comme résultat le vecteur nul
si et seulement si tous les coefficients de la combinaison sont des zéros. Dans le cas contraire les
1.2 Scalar product, distance, orthogonality
A scalar product on a vector space is a positive definite, symmetric, bilinear form. A form
ω is a function
ω : V × V → R,
i.e. ω operates on a couple of vectors to give a real number, a scalar. We will indicate
the scalar product of two vectors u and v as2

ω(u, v) = u · v.

The properties of bilinearity prescribe that, ∀u, v ∈ V and ∀α, β ∈ R,

u · (αv + βw) = αu · v + βu · w,
(αu + βv) · w = αu · w + βv · w,

while symmetry implies that

u · v = v · u ∀u, v ∈ V.

Finally, the positive definiteness means that

v · v > 0 ∀v ∈ V, v · v = 0 ⇐⇒ v = o.

Any two vectors u, v ∈ V are said to be orthogonal ⇐⇒

u · v = 0.

Thanks to the properties of the scalar product, we can define the Euclidean norm of a
vector v as the nonnegative scalar, denoted equivalently by v or |v|,

v = |v| := v · v

Theorem 1. The norm of a vector has the following properties: ∀u, v ∈ V, k ∈ R,

|u · v| ≤ u v (Schwarz′ s inequality);
|u + v| ≤ u + v (Minkowski′ s triangular inequality);
|kv| = |k|v.

Proof. Schwarz’s inequality: It is sufficient to prove that

(u · v)2 ≤ u · u v · v.

Let x = v · v and y = −u · v. Then, by the positive definiteness of the scalar product, we


get
(xu + yv) · (xu + yv) ≥ 0,
2
The scalar product ω(u, v) is also indicated as < u, v >.

3
which implies that

x2 u · u + 2xyu · v + y 2 v · v = (v · v)2 u · u − 2v · v(u · v)2 + v · v(u · v)2 ≥ 0;

supposing v ̸= o (otherwise, the proof is trivial), we get the thesis on dividing by v · v.


Minkowski’s inequality: Because the two members of the inequality to be proved are
nonnegative, it is sufficient to prove that

(u + v) · (u + v) ≤ (u + v)2 = u2 + 2uv + v 2 .

This can be proved easily:

(u + v) · (u + v) = u · u + 2u · v + v · v = u2 + 2u · v + v 2
≤ u2 + 2|u · v| + v 2 ≤ u2 + 2uv + v 2 ,

in which the last operation follows from the Schwarz’s inequality.


The proof of the third property is immediate, it is sufficient to use the same definition of
norm.

We define distance between any two points p and q ∈ E the scalar

d(p, q) := |p − q| = |q − p|.

Similarly, the distance between two any vectors u and v ∈ V is defined as

d(u, v) := |u − v| = |v − u|.

Two points or two vectors are coincident if and only if their distance is null.
The unit sphere S of V is defined as the set of all the vectors whose norm is one:

S := {v ∈ V| v = 1}.

1.3 Basis of V, expression of the scalar product


There is a general way to define a basis for a vector space of any kind. Here, we limit
the introduction of the concept of basis to the case of V only, of interest in classical
mechanics. Generally speaking, a basis B of V is any set of three linearly independent
vectors ei , i = 1, 2, 3 of V:
B = {e1 , e2 , e3 }.
The introduction of a basis for V is useful for representing vectors. In fact, once a basis B
is fixed, any vector v ∈ V can be represented as a linear combination of the vectors of the
basis, where the coefficients vi of the linear combination are the Cartesian components of
v:
v = vi ei = v1 e1 + v2 e2 + v3 e3 .

4
Though the choice of the elements of a basis is completely arbitrary, the only condition
being their linear independency, we will use in the following only orthonormal bases, which
are bases composed of mutually orthogonal vectors of S, i.e. satisfying

ei · ej = δij ,

where the symbol δij is the so-called Kronecker’s delta:


(
1 if i = j,
δij =
0 if i ̸= j.

The use of orthonormal bases has great advantages; namely, it allows us to give a very
simple rule for the calculation of the scalar product:

u · v = ui ei · vj ej = ui vj δij = ui vi = u1 v1 + u2 v2 + u3 v3 .

In particular, it is
v · ei = vk ek · ei = vk δik = vi , i = 1, 2, 3.
So, the Cartesian components of a vector are the projection of the vector on the three
vectors of the basis B; such quantities are the director cosines of v in the basis B. In fact,
if θ is the angle formed by two vectors u and v, then

u · v = u v cos θ.

This relation is used to define the angle between two vectors,


u·v
θ = arccos ,
uv
which can be proved easily: Given two vectors u and v, we look for c ∈ R such that the
vector u − cv is orthogonal to v:
u·v u·v
(u − cv) · v = 0 ⇐⇒ c = = 2 .
v·v v
Now, if u is inclined of θ on v, its projection uv on the direction of v is

uv = u cos θ,

and, by construction (see Fig. 1.3), it is also

uv = c v.

So,
u u u·v u·v
c= cos θ → cos θ = 2 ⇒ cos θ = .
v v v uv

We remark that while the scalar product, being an intrinsic operation, does not change
for a change of basis, the components vi of a vector are not intrinsic quantities, but they

5
u-cv u

𝜃
uv v

Figure 1.3: Angle between two vectors.

are basis-dependent: A change of the basis makes the components change. The way this
change is done will be introduced in Section 2.11.
A frame R for E is composed of a point o ∈ E, the origin, and a basis B of V:

R := {o, B} = {o; e1 , e2 , e3 }.

The use of a frame for E is useful for determining the position of a point p, which can be
done through its Cartesian coordinates xi , defined as the components in B of the vector
p − o:
xi := (p − o) · ei , i = 1, 2, 3.
Of course, the coordinates xi of a point p ∈ E depend upon the choice of o and B.

1.4 Applied vectors


We introduce now a set of definitions, concepts and results that are widely used in physics,
especially in mechanics. For that, we need to anticipate some results that are introduced
in the next Chapter, namely that of cross product, in Section 2.9, and of complementary
projector, in Exercise 2, Chapter 2. This slight deviation from the good rule of consistent
progression in stating the results is justified by the fact that, actually, the matter exposed
hereafter is still that of vectors. The readers can, of course, come back to the topics of
this section once they have studied Chapter 2.
We call applied vector vp a vector v associated to a point p ∈ E. In physics the concept
of applied vector3 is often employed, for example, to represent forces4 . We define the
resultant of a system of n applied vectors vip as the vector
n
X
R := vip .
i=1

We define moment of an applied vector vp about a point o, called the center of the moment,
the vector
Mo := (p − o) × vp ,
3
In the literature, applied vectors are also called bound vectors.
4
The fact that in classical mechanics forces can be represented by vectors is actually a fundamental
postulate of physics. Forces are vectors that cannot be considered belonging to the translation space V;
nevertheless, the definitions and results found earlier are also valid for vectors ∈
/ V.

6
and resultant moment of a system of n applied vectors vip about a point o the vector
n
X
Mro := (pi − o) × vip .
i=1

We remark that R, Mo and Mro are not applied vectors.


If u ∈ S| u × vp = o, then

b := |(I − u ⊗ u)(p − o)| = |p − o| sin θ

is called the moment arm of vp with respect to the center o. It measures the distance of
o from the line of action, i.e. the line passing through p and parallel to vp , cf. Fig. 1.4.

o u

b
𝜃
vp
p

Figure 1.4: Moment arm of an applied vector.

Theorem 2. (Transport of the moment). If Mo1 is the moment of an applied vector


vp about a center o1 , the moment Mo2 of vp about to another center o2 is

Mo2 = Mo1 + (o1 − o2 ) × vp .

Proof. Referring to Fig. 1.5,

Mo2 = (p − o2 ) × vp = (p − o1 + o1 − o2 ) × vp
= (p − o1 ) × vp + (o1 − o2 ) × vp
= Mo1 + (o1 − o2 ) × vp .

o2

vp

p
o1
Figure 1.5: Scheme for the transposition of the moment.

7
A consequence of this theorem is that Mo1 = Mo2 ⇐⇒ vp × (o1 − o2 ), i.e. if vp and
o1 − o2 are parallel. It follows from this that the moment of an applied vector does not
change when calculated about the points of a straight line parallel to the vector itself or,
more importantly, if vp is translated along its line of action.
The above theorem can be extended to the resultant moment of a system of applied
vectors to give (the proof is quite similar)

Mro2 = Mro1 + (o1 − o2 ) × R. (1.1)

Also in this case, the resultant moment does not change when R and o1 − o2 are parallel
vectors, but not exclusively, as another possibility is that R = o: For the systems of
applied vectors with null resultant, the resultant moment is invariant with respect to the
center of the moment.
An interesting relation can be found if the two members of the last equation are projected
onto R, which gives
Mro1 · R = Mro2 · R : (1.2)
The projection of the resultant moment onto the direction of R does not depend upon
the center of the moment.
A particularly important case of system with null resultant is that of a couple, which is
composed by two opposite vectors v, −v, applied to two points p and q:

vp = −vq .

Of course, by definition, R = o for any couple and, as a consequence, the resultant


moment Mr of a couple, called the moment of the couple and simply denoted by M, is
independent of the center of the moment (that is why the index denoting the center of
the moment is omitted): Referring to Fig. 1.6,

vp
p u
𝜃
q
o
bc vq

Figure 1.6: Scheme of a couple.

M = (p − o) × vp + (q − o) × vq = (p − o) × v − (q − o) × v
= ((p − o) − (q − o)) × v = (p − q) × v.

If u ∈ S| u × v = o, then

bc := |(I − u ⊗ u)(p − q)| = |p − q| sin θ

8
is the couple arm. We then have

M = |(p − q) × v| = |p − q|v sin θ = bc v.

The central axis A of a system of n applied vectors with R ̸= o is the axis such that

Mra × R = o ∀a ∈ A.

Theorem 3. (Existence and uniqueness of the central axis). The central axis of
a system of n vectors exists and is unique.

Proof. Existence: We need at least a point a ∈ E| Mra = kR, k ∈ R ⇒ Mra × R = o.


From Eq. (1.1), ∀o ∈ E we get

Mra × R = Mro × R + ((o − a) × R) × R = Mro × R − R2 (o − a) + (R · (o − a))R.

Then, if we take for o − a the vector


Mro × R
o−a= ,
R2
it is evident that we get
Mra × R = o.
Hence, the point
Mro × R
a=o− ∈ A.
R2
So, because Mr does not change when calculated with respect to the points of an axis
parallel to R, A is the axis passing through a and parallel to R. Its equation is
Mro × R
p=a+t R=o− + t R, t ∈ R.
R2

Uniqueness: Suppose another axis  ≠ A exists, which is necessarily parallel to A. If


q ∈ Â, again using Eq. (1.1) we get

Mrq × R = Mra × R + ((a − q) × R) × R.

In this equation the left-hand side and the first term on the right-hand side are null by
the definition of central axis. Because (a − q) × R is perpendicular to R and R ̸= o by
hypothesis, the left-hand side is null if and only if a = q ⇒ Â = A.

The central axis has another remarkable property:

Theorem 4. (Property of minimum of the central axis). The points of the central
axis minimize the resultant moment.

9
Proof. When Mr is calculated about a point a ∈ A, it is parallel to R, which is not the
/ A. In this last case, hence, Mr has also a component orthogonal
case for any point q ∈
to R. Then, by virtue of the invariance of the projection of Mr onto R, Eq. (1.2), Mr
gets its minimum value when calculated about the points of A.

Let us now consider the case of systems for which

Mro · R = 0 ∀o ∈ E.

This is namely the case of systems of coplanar or parallel vectors (cf. Exercise 7). Because
in this case, for the points a ∈ A, it must be at the same time Mra ·R = 0 and Mra ×R = o,
the only possibility is that
Mra = o ∀a ∈ A,

i.e., in these cases A is the axis of points that make the resultant moment vanish.
Two systems of applied vectors are equivalent if they have the same resultant R and the
same resultant moment Mro about any center o ∈ E. The equivalence does not depend
upon the center o. In fact, by Eq. (1.1), if two systems have the same R and the same
Mro1 , with o1 a given point, then also Mro2 will be the same, ∀o2 ∈ E.

Theorem 5. (Reduction of a system of applied vectors). A system of applied


vectors is always equivalent to the system composed by the resultant R applied at a point
o and by a couple with moment M = Mro , with o any point of E.

Proof. By construction, R is the same for the two systems; moreover, for the equivalent
system (resultant plus couple) it is

M + (o − o) × R = M.

So, if the couple has a moment M = Mro , the two systems are equivalent.

In practice, this theorems affirms that it is always possible to reduce a system of n applied
vectors to only an applied vector equal to R and to a couple or, if one of the two vectors
composing the couple is applied to the same point of R, to two applied vectors. It is
worth noting that the equivalence of two systems is preserved if a vector is translated
along its line of action, because in such a case R and Mro do not change.
Finally, a system of n applied vectors is said to be equilibrated if

R = o, Mro = o ∀o ∈ E.

We note that, because R = o, the center o can be any point of E.

10
1.5 Exercises
1. Prove that the null vector is unique.
2. Prove that the null vector is orthogonal to any vector.
3. Prove that the norm of the null vector is zero.
4. Prove that
u · v = 0 ⇐⇒ |u − v| = |u + v| ∀u, v ∈ V.

5. Prove the linear forms representation theorem: Let ψ : V → R be a linear function.


Then, ∃! u ∈ V such that
ψ(v) = u · v ∀v ∈ V.

6. Consider a point p and two noncollinear vectors u, v ∈ S at p. Show that a vector


w is the bisector of the angle formed by u and v if and only if w · u = w · v.
7. Show that in the case of systems composed of coplanar or parallel applied vectors
with R ̸= o, Mro · R = 0 ∀o ∈ E.
8. Prove that any system of applied vectors with R = o is equivalent to a couple.
9. Prove that a system of applied vectors all passing through a point p is equivalent to
R applied to p.
10. Prove that if for a system of applied vectors Mro = o, then the system is equivalent
to R applied to o. Then, show that if o ∈ A, this is the case of coplanar or parallel
vectors.
11. Prove that a system of applied vectors is equilibrated if and only if any equivalent
system is equilibrated.
12. Prove that two applied vectors form an equilibrated system if and only if they are
two opposite vectors applied to the same point.
13. Prove that a system of applied vectors is equilibrated if all the vectors pass through
the same point and R = o.

11
12
Chapter 2

Second rank tensors

2.1 Second-rank tensors


A second-rank tensor L is any linear application from V to V:

L : V → V | L(αi ui ) = αi Lui ∀αi ∈ R, ui ∈ V, i = 1, ..., n.

Though here V indicates the vector space of translations over E, the definition of tensor1
is more general and, in particular, V can be any vector space.
Defining the sum of two tensors as

(L1 + L2 )u = L1 u + L2 u ∀u ∈ V, (2.1)

the product of a scalar by a tensor as

(αL)u = α(Lu) ∀α ∈ R, u ∈ V

and the null tensor O as the unique tensor such that

Ou = o ∀u ∈ V,

then the set of all the tensors L that operate on V forms a vector space, denoted by
Lin(V). We define the identity tensor I as the unique tensor such that

Iu = u ∀u ∈ V.

Different operations can be defined for the second-rank tensors. We consider all of them
in the following sections.
1
We consider, for the time being, only second-rank tensors, that constitute a very important set of
operators in classical and continuum mechanics. In the following, we also introduce fourth-rank tensors.

13
2.2 Dyads, tensor components
For any couple of vectors u and v, the dyad2 u ⊗ v is the tensor defined by

(u ⊗ v)w := v · w u ∀w ∈ V.

The application defined above is actually a tensor because of the bilinearity of the scalar
product. The introduction of dyads allows us to express any tensor as a linear combination
of dyads. In fact, it can be proved that if B = {e1 , e2 , e3 } is a basis of V, then the set of
nine dyads
B 2 = {ei ⊗ ej , i, j = 1, 2, 3},

is a basis of Lin(V), so that dim(Lin(V)) = 9. This implies that any tensor L ∈ Lin(V)
can be expressed as
L = Lij ei ⊗ ej , i, j = 1, 2, 3,

where the Lij s are the nine Cartesian components of L with respect to B 2 . The Lij s can
be calculated easily:

ei · Lej = ei · Lhk eh ⊗ ek ej = Lhk ei · eh ek · ej = Lhk δih δjk = Lij .

The above expression is sometimes called the canonical decomposition of a tensor. The
components of a dyad can be computed as follows:

(u ⊗ v)ij = ei · (u ⊗ v) ej = u · ei v · ej = ui vj . (2.2)

The components of a vector v resulting from the application of a tensor L on a vector u,


can now be calculated:

v = Lu = Lij (ei ⊗ ej )(uk ek ) = Lij uk δjk ei = Lij uj ei → vi = Lij uj . (2.3)

Depending upon two indices, any second-rank tensor L can be represented by a matrix,
whose entries are the Cartesian components of L in the basis B:
 
L11 L12 L13
L =  L21 L22 L23  .
L31 L32 L33

Because any u ∈ V, depending upon only one index, can be represented by a column
vector, Eq. (2.3) represents actually the classical operation of the multiplication of a 3 × 3
matrix by a 3 × 1 vector.
2
In some texts, the dyad is also called the tensor product; we prefer to use the term dyad because
the expression tensor product can be ambiguous, as it is used to denote the product of two tensors, see
Section 2.3.

14
2.3 Tensor product
The tensor product of L1 and L2 ∈ Lin(V) is defined by

(L1 L2 )v := L1 (L2 v) ∀v ∈ V.

By linearity and Eq. (2.1), ∀L, L1 , L2 ∈ Lin(V), u ∈ V, we get

[L(L1 + L2 )]v = L[(L1 + L2 )v] = L(L1 v + L2 v)


= LL1 v + LL2 v = (LL1 + LL2 )v →
L(L1 + L2 ) = LL1 + LL2 .

We remark that the tensor product is not symmetric:

L1 L2 ̸= L2 L1 ;

however, by the same definition of the identity tensor and of tensor product,

IL = LI = L ∀L ∈ Lin(V).

The Cartesian components of a tensor L = AB can be calculated using Eq. (2.3):

Lij = ei · (AB)ej = ei · A(Bej ) = ei · A(Bhk (ej )k eh ) = Bhk δjk ei · Aeh


= Bhk δjk ei · (Apq (eh )q ep ) = Apq Bhk δjk δqh δip = Aih Bhj .

The above result simply corresponds to the row-column multiplication of two matrices.
Using that, the following two identities can be readily shown:

(a ⊗ b)(c ⊗ d) = b · c(a ⊗ d) ∀a, b, c, d ∈ V,


(2.4)
A(a ⊗ b) = (Aa) ⊗ b ∀a, b ∈ V, A ∈ Lin(V).

Finally, the symbol L2 is normally used to denote, in short, the product LL, ∀L ∈
Lin(V).

2.4 Transpose, symmetric and skew tensors


For any tensor L ∈ Lin(V), there exists just one tensor L⊤ , called the transpose of L,
such that
u · Lv = v · L⊤ u ∀u, v ∈ V. (2.5)
The transpose of the transpose of L is L:

u · Lv = v · L⊤ u = u · (L⊤ )⊤ v ⇒ (L⊤ )⊤ = L.

The Cartesian components of L⊤ are obtained by swapping the indices of the components
of L:
L⊤ ⊤ ⊤ ⊤
ij = ei · L ej = ej · (L ) ei = ej · Lei = Lji .

15
It is immediate to show that
(A + B)⊤ = A⊤ + B⊤ ∀A, B ∈ Lin(V),
while
u · (AB)v = Bv · A⊤ u = v · B⊤ A⊤ u ⇒ (AB)⊤ = B⊤ A⊤ .
Moreover,
u · (a ⊗ b)v = a · u b · v = v · (b ⊗ a)u ⇒ (a ⊗ b)⊤ = b ⊗ a. (2.6)

A tensor L is symmetric ⇐⇒
L = L⊤ .
In such a case, because Lij = L⊤
ij , we have

Lij = Lji .
A symmetric tensor is hence represented, in a given basis, by a symmetric matrix and has
just six independent Cartesian components. Applying Eq. (2.5) to I, it is immediately
recognized that the identity tensor is symmetric: I = I⊤ .
A tensor L is antisymmetric or skew ⇐⇒
L = −L⊤ .
In this a case, because Lij = −L⊤ ij , we have (no summation on the index i, see footnote
1, Chapter 1)
Lij = −Lji ⇒ Lii = 0 ∀i = 1, 2, 3.
A skew tensor is hence represented, in a given basis, by an antisymmetric matrix whose
components on the diagonal are identically null in any basis; finally, a skew tensor only
depends upon three independent Cartesian components.
If we denote by Sym(V) the set of all the symmetric tensors and by Skw(V) that of all
the skew tensors, then it is evident that, ∀α, β, λ, µ ∈ R,
Sym(V) ∩ Skw(V) = O,
αA + βB ∈ Sym(V) ∀A, B ∈ Sym(V),
λL + µM ∈ Skw(V) ∀L, M ∈ Skw(V),
so Sym(V) and Skw(V) are vector subspaces of Lin(V) with dim(Sym(V)) = 6, while
dim(Skw(V)) = 3.
Any tensor L can be decomposed into the sum of a symmetric, Ls , and an antisymmetric,
La , tensor:
L = Ls + La ,
with
L + L⊤
Ls = ∈ Sym(V)
2
and
L − L⊤
La = ∈ Skw(V),
2
so that, finally,
Lin(V) = Sym(V) ⊕ Skw(V).

16
2.5 Trace, scalar product of tensors
There exists one and only one linear form
tr : Lin(V) → R,
called the trace, such that
tr(a ⊗ b) = a · b ∀a, b ∈ V.
For its same definition, that has been given without making use of any basis of V, the
trace of a tensor is a tensor invariant, i.e. a quantity, extracted from a tensor, that does
not depend upon the basis.
Linearity implies that
tr(αA + βB) = αtrA + βtrB ∀α, β ∈ R, A, B ∈ Lin(V).
It is just linearity to give the rule for calculating the trace of a tensor L:
trL = tr(Lij ei ⊗ ej ) = Lij tr(ei ⊗ ej ) = Lij ei · ej = Lij δij = Lii . (2.7)
A tensor is hence an operator whose sum of the components on the diagonal,
trL = L11 + L22 + L33 ,
is constant, regardless of the basis.
Following the same procedure above, it is readily seen that
trL⊤ = trL,
which implies, by linearity, that
trL = 0 ∀L ∈ Skw(V). (2.8)

The scalar product of tensors A and B is the positive definite, symmetric bilinear form
defined by
A · B = tr(A⊤ B).
This definition implies that, ∀L, M, N ∈ Lin(V), α, β ∈ R,
L · (αM + βN) = αL · M + βL · N,
(αL + βM) · N = αL · N + βM · N,
L · M = M · L,
L · L > 0 ∀L ∈ Lin(V), L · L = 0 ⇐⇒ L = O.
These properties give the rule for computing the scalar product of two tensors A and
B:
A · B = Aij (ei ⊗ ej ) · Bhk (eh ⊗ ek ) = Aij Bhk (ei ⊗ ej ) · (eh ⊗ ek )
= Aij Bhk tr[(ei ⊗ ej )⊤ (eh ⊗ ek )] = Aij Bhk tr[(ej ⊗ ei )(eh ⊗ ek )]
= Aij Bhk tr[ei · eh (ej ⊗ ek )] = Aij Bhk ei · eh ej · ek
= Aij Bhk δih δjk = Aij Bij .

17
As in the case of vectors, the scalar product of two tensors is equal to the sum of the
products of the corresponding components. In a similar manner, or using Eq. (2.4)1 , it
is easily shown that, ∀a, b, c, d ∈ V,

(a ⊗ b) · (c ⊗ d) = a · c b · d = ai bj ci dj ,

while by the same definition of the tensor scalar product,

trL = I · L ∀L ∈ Lin(V).

Similar to vectors, we define Euclidean norm of a tensor L the nonnegative scalar, denoted
either by L or |L|, √ p p
L = |L| = L · L = tr(L⊤ L) = Lij Lij ,
and the distance d(L, M) of two tensors L and M the norm of the tensor difference:

d(L, M) := |L − M| = |M − L|.

2.6 Spherical and deviatoric parts


Let L ∈ Sym(V); the spherical part of L is defined by

1
Lsph := trL I,
3
and the deviatoric part by
Ldev := L − Lsph ,
so that
L = Lsph + Ldev .
We remark that
1
trLsph = trL trI = trL ⇒ trLdev = 0,
3
i.e. the deviatoric part is a traceless tensor. Let A, B ∈ Lin(V); then

1 1
Asph · Bdev = trA I · Bdev = trA trBdev = 0, (2.9)
3 3
i.e. any spherical tensor is orthogonal to any deviatoric tensor.
The sets
 
sph sph 1
Sph(V) := A ∈ Lin(V)| A = trAI ∀A ∈ Lin(V) ,
3
Dev(V) := A ∈ Lin(V)| Adev = A − Asph ∀A ∈ Lin(V)
 dev

form two subspaces of Lin(V); the proof is left to the reader. For what is proved above,
Sph(V) and Dev(V) are two mutually orthogonal subspaces of Lin(V).

18
2.7 Determinant, inverse of a tensor
The reader is probably familiar with the concept of determinant of a matrix. We show
here that the determinant of a second-rank tensor can be defined intrinsically and that it
corresponds with the determinant of the matrix that represents it in any basis of V. For
this purpose, we first need to introduce a mapping:

ω :V ×V ×V →R

is a skew trilinear form if ω(u, v, ·), ω(u, ·, v) and ω(·, u, v) are linear forms on V and
if
ω(u, v, w) = −ω(v, u, w) = −ω(u, w, v) = −ω(w, v, u) ∀u, v, w ∈ V. (2.10)

Using this definition, we can state the following

Theorem 6. Three vectors are linearly independent if and only if every skew trilinear
form on them is not null.

Proof. In fact, let u = αv + βw, then for any skew trilinear form ω,

ω(u, v, w) = ω(αv + βw, v, w) = αω(v, v, w) + βω(w, v, w) = 0

because of Eq. (2.10) applied to the permutation of the positions of the two v and the
two w.

It is evident that the set of all the skew trilinear forms is a vector space, that we denote
by Ω, whose null element is the null form ω0 :

ω0 (u, v, w) = 0 ∀u, v, w ∈ V.

For a given ω(u, v, w) ∈ Ω, any L ∈ Lin(V) induces another form ωL (u, v, w) ∈ Ω,


defined as
ωL (u, v, w) = ω(Lu, Lv, Lw) ∀u, v, w ∈ V.
A key point3 for the following developments is that dim Ω = 1.
This means that ∀ω1 , ω2 ̸= ω0 ∈ Ω, ∃λ ∈ R such that

ω2 (u, v, w) = λω1 (u, v, w) ∀u, v, w ∈ V.

So, ∀L ∈ Lin(V), there must exist λL ∈ R such that

ω(Lu, Lv, Lw) = ωL (u, v, w) = λL ω(u, v, w) ∀u, v, w ∈ V. (2.11)


3
The proof of this statement is rather involved and outside of our scope; the interested reader is
referred to the classical textbook by Halmos on linear algebra, Section 31 (see the bibiography). The
theory of the determinants is developed in Section 53.

19
The scalar4 λL is the determinant of L and in the following it will be denoted as det L.
The determinant of a tensor L is an intrinsic quantity of L, i.e. it does not depend upon
the particular form ω, nor on the basis of V. In fact, we have never introduced, so far,
a basis for defining det L, hence it cannot depend upon the choice of a basis for V, i.e.
det L is a tensor invariant.
Then, if ω a and ω b ∈ Ω, because dim Ω = 1, there exists k ∈ R, k ̸= 0 such that
ω b (u, v, w) = k ω a (u, v, w) ∀u, v, w ∈ V ⇒
ω b (Lu, Lv, Lw) = k ω a (Lu, Lv, Lw) →
ωLb (u, v, w) = k ωLa (u, v, w).
Moreover, by Eq. (2.11) we get
ω a (Lu, Lv, Lw) = ωLa (u, v, w) = λaL ω a (u, v, w),
ω b (Lu, Lv, Lw) = ωLb (u, v, w) = λbL ω b (u, v, w),
so that
λbL k ω a (u, v, w) = λbL ω b (u, v, w) = ωLb (u, v, w) =
k ωLa (u, v, w) = λaL k ω a (u, v, w) ⇐⇒ λaL = λbL ,
which proves that det L does not depend upon the skew trilinear form, but only upon
L.
The definition given for det L allows us to prove some important properties. First of
all,
det O = 0;
in fact, ∀ω ∈ Ω,
det O ω(u, v, w) = ω(Ou, Ov, Ow) = ω(o, o, o) = 0 ∀u, v, w ∈ V
because ω operates on three identical, i.e. linearly dependent, vectors. Moreover, if L = I,
then
det I ω(u, v, w) = ω(Iu, Iv, Iw) = ω(u, v, w)
if and only if
det I = 1. (2.12)
A third property is that ∀a, b ∈ V,
det(a ⊗ b) = 0. (2.13)
In fact, if L = a ⊗ b, then
det L ω(u, v, w) = ω(Lu, Lv, Lw) = ω((b · u)a, (b · v)a, (b · w)a) = 0
because the three vectors on which ω ∈ Ω operates are linearly dependent; because u, v
and w are arbitrarily chosen, this implies Eq. (2.13).
An important result is the
4
More precisely, det L is the function that associates a scalar with each tensor (Halmos, Section 53).
We can, however, for the sake of practice, identify det L with the scalar associated with L, without
consequences for our purposes.

20
Theorem 7. (Theorem of Binet). ∀A, B ∈ Lin(V)

det(AB) = det A det B. (2.14)

Proof. ∀ω ∈ Ω and ∀u, v, w ∈ V,


λAB ω(u, v, w) = ω(ABu, ABv, ABw) = ω(A(Bu), A(Bv), A(Bw)) =
λA ω(Bu, Bv, Bw) = λA λB ω(u, v, w) ⇐⇒ λAB = λA λB ,
which proves the theorem.

A tensor L is called singular if det L = 0, otherwise it is non-singular.


Considering Eq. (2.11), with some effort but without major difficulties, one can see that,
if in a basis B of V it is L = Lij ei ⊗ ej , then
X
det L = ϵπ(1),π(2),π(3) L1,π(1) L2,π(2) L3,π(3) ,
π∈P3

where P3 is the set of all the permutations π of {1, 2, 3} and the ϵi,j,k s are the components
of the Ricci’s alternator5 :

 1 if {i, j, k} is an even permutation of {1, 2, 3},
ϵi,j,k := 0 if {i, j, k} is not a permutation of {1, 2, 3}
−1 if {i, j, k} is an odd permutation of {1, 2, 3}.

The above rule for det L coincides with that for calculating the determinant of the matrix
whose entries are the Lij s. This shows that, once chosen a basis B for V, det L coincides
with the determinant of the matrix representing it in B, and finally that
det L = L11 L22 L33 + L12 L23 L31 + L13 L32 L21
(2.15)
− L11 L23 L32 − L22 L13 L31 − L33 L12 L21 .

This result shows immediately that ∀L ∈ Lin(V), and regardless of B, we have

det L⊤ = det L. (2.16)

Using Eq. (2.15), it is not difficult to show that, ∀α ∈ R,

det(I + αL) = 1 + αI1 + α2 I2 + α3 I3 , (2.17)

where I1 , I2 and I3 are the three principal invariants of L:


tr2 L − trL2
I1 = trL, I2 = , I3 = det L. (2.18)
2
5
We recall that a permutation of an ordered set of n objects is even if it can be obtained as the product
of an even number of transpositions, i.e. exchange of places, of any couple of its elements, it is odd if the
number of transpositions is odd. For the set {1, 2, 3} the even permutations are {1, 2, 3}, {3, 1, 2}, {2, 3, 1},
while the odd ones are {2, 1, 3}, {1, 3, 2}, {3, 2, 1}; any triplet having at least a repeated number is not a
permutation.

21
A tensor L ∈ Lin(V) is said to be invertible if there exists a tensor L−1 ∈ Lin(V), called
the inverse of L, such that
LL−1 = L−1 L = I. (2.19)
If L is invertible, then L−1 is unique. By the above definition, if L is invertible, then

u1 = Lu ⇒ u = L−1 u1 .

Theorem 8. Any invertible tensor maps triples of linearly independent vectors into triples
of still linearly independent vectors.

Proof. Let L be an invertible tensor and u1 = Lu, v1 = Lv, w1 = Lw, where u, v, w are
three linearly independent vectors. Let us suppose that there exist h, k ∈ R such that

u1 = hv1 + kw1 .

Then, because L is invertible,

L−1 u1 = L−1 (hv1 + kw1 ) = hL−1 v1 + kL−1 w1 = hv + kw,

which goes against the hypothesis. Consequently, u1 , v1 and w1 are linearly independent.

This result, along with the definition of determinant, Eq. (2.11), and Theorem 6, proves
the

Theorem 9. (Invertibility theorem). L ∈ Lin(V) is invertible ⇐⇒ det L ̸= 0.


Using the theorem of Binet, 7, along with Eqs. (2.12) and (2.19), we get

1
det L−1 = .
det L

Equation (2.19) applied to L−1 , along with the uniqueness of the inverse, gives immedi-
ately that
(L−1 )−1 = L,
while
B−1 A−1 = B−1 A−1 AB(AB)−1 = (AB)−1 .
The operations of transpose and inversion commute:

L⊤ (L⊤ )−1 = I = L−1 L = I⊤ = (L−1 L)⊤ = L⊤ (L−1 )⊤ ⇒


(L−1 )⊤ = (L⊤ )−1 := L−⊤ .

22
2.8 Eigenvalues and eigenvectors of a tensor
If there exists a λ ∈ R and a v ∈ V, except the null vector, such that

Lv = λv, (2.20)

then λ is an eigenvalue and v an eigenvector, relatif to λ, of L. It is immediate to observe


that, thanks to linearity, any eigenvector v of L is determined to within a multiplier, i.e.
that kv is an eigenvector of L too ∀k ∈ R. Often, the multiplier k is fixed in such a way
that |v| = 1.
To determine the eigenvalues and eigenvectors of a tensor, we rewrite Eq. (2.20) as

(L − λI)v = o. (2.21)

The condition for this homogeneous system having a non null solution is

det(L − λI) = 0;

this is the so-called characteristic or Laplace’s equation. In the case of a second-rank


tensor over V, the Laplace’s equation is an algebraic equation of degree three with real
coefficients. The roots of the Laplace’s equation are the eigenvalues of L; because the
components of L, and hence the coefficients of the characteristic equation, are all real,
then the eigenvalues of L are all real or one real and two complex conjugate.
For any eigenvalue λi , i = 1, 2, 3, of L, the corresponding eigenvectors vi can be found
solving Eq. (2.21), once set λ = λi .
The proper space of L relatif to λ is the subspace of V composed of all the vectors that
satisfy Eq. (2.21). The multiplicity of λ is the dimension of its proper space, while the
spectrum of L is the set composed by all of its eigenvalues, each one with its multiplic-
ity.
L⊤ has the same eigenvalues of L, because the Laplace’s equation is the same in both the
cases:
det(L⊤ − λI) = det(L⊤ − λI⊤ ) = det(L − λI)⊤ = det(L − λI).
However, this is not the case for the eigenvectors, that are generally different, as a nu-
merical example can show.
Developing the Laplace’s equation, it is easy to show that it can be written as

det(L − λI) = −λ3 + I1 λ2 − I2 λ + I3 = 0,

which is merely an application of Eq. (2.17). If we denote L3 = LLL, using Eq. (2.18)
one can prove the Cayley-Hamilton theorem:

Theorem 10. (Cayley-Hamilton theorem). ∀L ∈ Lin(V),

L3 − I1 L2 + I2 L − I3 I = O.

23
A quadratic form defined by L is any form ω : V × V → R of the type

ω = v · Lv;

if ω > 0 ∀v ∈ V, ω = 0 ⇐⇒ v = o, then ω and L are said to be positive definite. The


eigenvalues of a positive definite tensor are positive. In fact, if λ is an eigenvalue of L,
which is positive definite, and v its eigenvector, then

v · Lv = v · λv = λv2 > 0 ⇐⇒ λ > 0.

Let v1 and v2 be two eigenvectors of a symmetric tensor L relative to the eigenvalues λ1


and λ2 , respectively, with λ1 ̸= λ2 . Then

λ1 v1 · v2 = Lv1 · v2 = Lv2 · v1 = λ2 v2 · v1 ⇐⇒ v1 · v2 = 0.

Actually, symmetric tensors have a particular importance, specified by the spectral theo-
rem:

Theorem 11. (Spectral theorem). The eigenvectors of a symmetric tensor form a


basis of V.
This theorem6 is of paramount importance in linear algebra: It proves that the eigenvalues
of a symmetric tensor L are real valued and, remembering the definition of eigenvalues
and eigenvectors, Eq. (2.20), that there exists a basis BN = {u1 , u2 , u3 } of V composed of
eigenvectors of L, i.e. by vectors that are mutually orthogonal and that remain mutually
orthogonal once transformed by L. Such a basis is called the normal basis.
If λi , i = 1, 2, 3, are the eigenvalues of L, then the components of L in BN are

Lij = ui · Luj = ui · λj uj = λj δij

so finally in BN we have
L = λi ei ⊗ ei ,
i.e. L is diagonal and is completely represented by its eigenvalues. In addition, it is easy
to check that

I1 = λ1 + λ2 + λ3 , I2 = λ1 λ2 + λ2 λ3 + λ3 λ1 , I3 = λ1 λ2 λ3 .

A tensor with a unique eigenvalue λ of multiplicity three is said to be spherical; in such


a case, any basis of V is BN and
L = λI.
Eigenvalues and eigenvectors have also another important property: Let us consider the
quadratic form ω := v · Lv, ∀v ∈ S, defined by a symmetric tensor L. We look for
the directions v ∈ S whereupon ω is stationary. Then, we have to solve the constrained
problem
∇v (v · Lv) = o, v ∈ S.
6
The proof of the spectral theorem is omitted here; the interested reader can find a proof of it in the
classical text by Halmos, p. 155, see the suggested texts.

24
Using the Lagrange’s multiplier technique, we solve the equivalent problem

∇(v,λ) (v · Lv − λ(v2 − 1)) = 0,

which restitutes the equation


Lv = λv
and the constraint |v| = 1. The above equation is exactly the one defining the eigenvalue
problem for L: The stationary values (i.e. the maximum and minimum) of ω hence
corresponds to two eigenvalues of L and the directions v, whereupon stationarity is get,
coincide with the respective eigenvectors.
Two tensors A and B are said to be coaxial if they have the same normal basis BN , i.e. if
they share the same eigenvectors. Let u be an eigenvector of A, relative to the eigenvalue
λA , and of B, relatif to λB . Then,

ABu = AλB u = λB Au = λA λB u = λA Bu = BλA u = BAu,

which shows, on the one hand, that also Bu is an eigenvector of A, relative to the same
eigenvalue λA ; in the same way, of course, Au is an eigenvector of B relative to λB . In
other words, this shows that B leaves unchanged any proper space of A and vice versa.
On the other hand, we see that, at least for what concerns the eigenvectors, two tensors
commute if and only if they are coaxial. Because any vector can be written as a linear
combination of the vectors of BN , and for the linearity of tensors, we finally have proved
the commutation theorem:

Theorem 12. (Commutation theorem). Two tensors commute if and only if they are
coaxial.

2.9 Skew tensors and cross product


Because dim(V) = dim(Skw(V)) = 3, an isomorphism can be established between V and
Skw(V), i.e. between vectors and skew tensors. We establish hence a way to associate
in a unique way a vector to any skew tensor and inversely. For this purpose, we first
introduce the following theorem:

Theorem 13. The spectrum of any tensor W ∈ Skw(V) is {0} and the dimension of its
proper space is 1.

Proof. This theorem states that zero is the only real eigenvalue of any skew tensor and
that its multiplicity is 1. In fact, let w be an eigenvector of W relative to the eigenvalue
λ. Then

λ2 w2 = Ww · Ww = w · W⊤ Ww = −w · WWw
= −w · W(λw) = −λw · Ww = −λ2 w2 ⇐⇒ λ = 0.

25
Then, if W ̸= O its rank is necessarily 2, because det W = 0 ∀W ∈ Skw(V); hence, the
equation
Ww = o (2.22)
has ∞1 solutions, i.e. the multiplicity of λ is 1, which proves the theorem.

The last equation also shows the way the isomorphism is constructed: In fact, using Eq.
(2.22) it is easy to check that if w = (a, b, c), then
 
0 −c b
w = (a, b, c) ⇐⇒ W =  c 0 −a  . (2.23)
−b a 0

The proper space of W is called the axis of W and it is indicated by A(W):

A(W) := {u ∈ V| Wu = o}.

The consequence of what shown above is that dim A(W) = 1. With regard to Eq. (2.23),
one can easily check that the equation
1
u·u= W·W (2.24)
2
is satisfied only by w and by its opposite −w. Because both these vectors belong to
A(W), choosing one of them corresponds to choose an orientation for E, see the next
section. We always make our choice according to Eq. (2.23), which fixes once and for
all the isomorphism between V and Skw(V) that makes correspond any vector w with
one and only one axial tensor W and vice-versa, any skew tensor W with a unique axial
vector w.
It is worth noting that the above isomorphism between the vector spaces V and Skw(V)
implies that to any linear combination of vectors a and b corresponds an equal linear
combination of the corresponding axial tensors Wa and Wb and vice-versa, i.e. ∀a, b ∈
R
w = αa + βb ⇐⇒ W = αWa + βWb , (2.25)
where W is the axial tensor of w. Such a property is immediately checked using Eq.
(2.23).
We define cross product of two vectors a and b the vector

a × b = Wa b,

where Wa is the axial tensor of a. If a = (a1 , a2 , a3 ) and b = (b1 , b2 , b3 ), then by Eq.


(2.23) we get
a × b = (a2 b3 − a3 b2 , a3 b1 − a1 b3 , a1 b2 − a2 b1 ).
It is immediate to check that such a result can also be obtained using Ricci’s alterna-
tor
a × b = ϵijk aj bk ei , (2.26)

26
or even computing the symbolic determinant
 
e1 e2 e3
a × b = det  a1 a2 a3  .
b1 b2 b 3

The cross product is bilinear: ∀a, b, u ∈ V, α, β ∈ R,


(αa + βb) × u = αa × u + βb × u,
u × (αa + βb) = αu × a + βu × b.
In fact, the first equation above is a consequence of Eq. (2.25), while the second one is a
simple application to axial tensors of the same definition of tensor.
Three important results concerning the cross product are stated by the following theo-
rems.

Theorem 14. (Condition of parallelism). Two vectors a and b are parallel, i.e.
b = ka, k ∈ R, ⇐⇒
a × b = o.

Proof. This property is actually a consequence of the fact that any eigenvalue of a tensor
is determined to within a multiplier:
a × b = Wa b = o ⇐⇒ b = ka, k ∈ R,
for Theorem 13.

Theorem 15. (Orthogonality property).


a × b · a = a × b · b = 0. (2.27)

Proof.
a × b · a = Wa b · a = b · Wa⊤ a = −b · Wa a = −b · o = 0,
a×b · b = Wa b · b = b · Wa⊤ b = −b · Wa b ⇐⇒ a × b · b = 0.

Theorem 16. a × b is the axial vector of the tensor (b ⊗ a − a ⊗ b).

Proof. First of all, by Eq. (2.6) we see that


(b ⊗ a − a ⊗ b) ∈ Skew(V).
Then,
(b ⊗ a − a ⊗ b)(a × b) = a · a × b b − b · a × b a = 0
for Theorem 15.

27
Theorem 16 allows us to show another important result about cross product: the anti-
symmetry of the cross product:

Theorem 17. (Antisymmetry of the cross product). The cross product is antisym-
metric:
a × b = −b × a ∀a, b ∈ V. (2.28)

Proof. Let W1 = (b ⊗ a − a ⊗ b) be the axial tensor of a × b and W2 = (−a ⊗ b + b ⊗ a)


that of −b × a. Evidently, W1 = W2 which implies Eq. (2.28) for the isomorphism
between V and Lin(V).

This property and, again, Theorem 16 lets us derive the formula for the double cross
product:

u × (v × w) = −(v × w) × u = −(w ⊗ v − v ⊗ w)u = u · w v − u · v w. (2.29)

Another interesting result concerns the mixed product:

u × v · w = Wu v · w = −v · Wu w = −v · u × w = w × u · v, (2.30)

and similarly
u × v · w = v × w · u.
Using this last result, we can obtain a formula for the norm of a cross product; if a = a ea
and b = b eb , with ea , eb ∈ S, are two vectors forming the angle θ, then

(a × b) · (a × b) = a × b · (a × b) = (a × b) × a · b = −a × (a × b) · b =
(−a · b a + a2 b) · b = b · (a2 I − a ⊗ a)b = a2 b · (I − ea ⊗ ea )b = (2.31)
a2 b2 eb · (I − ea ⊗ ea )eb = a2 b2 (1 − cos2 θ) = a2 b2 sin2 θ → |a × b| = ab sin θ.
v ∣u⨉v∣
So, the norm of a cross product can be interpreted, geometrically, as the area of the
parallelogram spanned by the u two vectors. As a consequence, the absolute value of the
mixed product (2.30) measures the volume of the prism delimited by three non coplanar
vectors, cf. Fig. 2.1.

∣detL∣∣u⨉v⋅w∣
∣u⨉v⋅w∣ Lw
w L
v Lv
∣u⨉v∣ ∣L*(u⨉v)∣
u Lu

Figure 2.1: Geometrical meaning of the cross and mixed products before (left) and after
(right) the application of a tensor L on the vectors u, v, w.

28
Because the cross product is antisymmetric and the scalar one is symmetric, it is easy to
check that the form
β(u, v, w) = u × v · w
is a skew trilinear form. Then, Eq. (2.11), we get
Lu × Lv · Lw = det L u × v · w. (2.32)
Following the interpretation given above for the absolute value of the mixed product,
we can conclude that | det L| can be interpreted as a coefficient of volume expansion7 ,
cf. again Fig. 2.1. A geometrical interpretation can then be given to the case of a non
invertible tensor, i.e. of det L = 0: It crushes a prism into a flat region (the three original
vectors become coplanar, i.e. linearly dependent).
The adjugate of L is the tensor
L∗ := (det L)L−⊤ .
From Eq. (2.32) we get hence
det L u × v · w = Lu × Lv · Lw = L⊤ (Lu × Lv) · w ∀w ⇒
Lu × Lv = L∗ (u × v).

It is useful, for further development, to calculate the powers of W:


W2 = WW = −W⊤ (−W⊤ ) = (WW)⊤ = (W2 )⊤ , (2.33)
i.e., W2 is symmetric. Moreover, if we take w ∈ S, which is always possible because
eigenvectors are determined to within an arbitrary multiplier,
W2 u = WWu = w × (w × u) = w · uw − w · wu
(2.34)
= −(I − w ⊗ w)u ⇒ W2 = −(I − w ⊗ w);
We remark that W2 u gives the opposite of the projection of any vector u ∈ V onto the
direction orthogonal to w, see Exercise 2.
Applying recursively the previous results,
W3 = WW2 = −W(I − w ⊗ w) = −W + (Ww) ⊗ w= −W
W4 = WW3 = −W2
(2.35)
W5 = WW4 = −W3
etc.
An important property of any couple axial tensor W−axial vector w ∈ S is
1
WW = − |W|2 (I − w ⊗ w), (2.36)
2
while Eq. (2.24) can be generalized to any two axial couples w1 , W1 and w2 , W2 :
1
w1 · w2 = W 1 · W 2 .
2
The proof of these two last properties is rather easy and left to the reader.
7
This result is classical and fundamental for the analysis of deformation in continuum mechanics.

29
2.10 Orientation of a basis
It is immediate to observe that a basis B = {e1 , e2 , e3 } can be oriented in two opposite
ways8 : For example, once two unit mutually orthogonal vectors e1 and e2 are chosen,
there are two opposite unit vectors perpendicular to both e1 and e2 that can be chosen
to form B.
We say that B is positively oriented or right-handed if
e1 × e2 · e3 = 1,
while B is negatively oriented or left-handed if
e1 × e2 · e3 = −1.
Schematically, a right-handed basis is represented in Fig. 2.2, where a left-handed basis
is represented too with a dashed e3 .

Figure 2.2: Right- and left-handed bases.

With a right-handed basis, by definition, the axial tensors of the three vectors of the basis
are
We1 = e3 ⊗ e2 − e2 ⊗ e3 ,
We2 = e1 ⊗ e3 − e3 ⊗ e1 ,
We3 = e2 ⊗ e1 − e1 ⊗ e2 .

2.11 Rotations
In the previous chapter, we have seen that the elements of V represent translations over
E. A rotation, i.e. a rigid rotation of the space, is an operation that transforms any two
vectors u, v ∈ V into two other vectors û, v̂ ∈ V in such a way that
u = û, v = v̂, u · v = û · v̂, (2.37)
i.e., a rotation is a transformation that preserves norms and angles. Because a rotation
is a transformation from V to V, rotations are tensors, so we can write
v̂ = Rv,
8
It is evident that this is true also for one- and two-dimensional vector spaces.

30
with R the rotation tensor or simply rotation.
Conditions (2.37) impose some restrictions on R:

û · v̂ = Ru · Rv = u · R⊤ Rv = u · v ⇐⇒ R⊤ R = I = RR⊤ .

A tensor that preserves the angles belongs to Orth(V), the subspace of orthogonal tensors;
we leave to the reader the proof that actually Orth(V) is actually a subspace of Lin(V).
Replacing in the above equation v with u shows immediately that an orthogonal tensor
also preserves the norms. By the uniqueness of the inverse, we see that

R ∈ Orth(V) ⇐⇒ R−1 = R⊤ .

The above condition is not sufficient to characterize a rotation; in fact, a rotation must
transform a right-handed basis into another right-handed basis, i.e. it must preserve the
orientation of the space. This means that it must be

ê1 × ê2 · ê3 = Re1 × Re2 · Re3 = e1 × e2 · e3 .

By Eq. (2.32), we get hence the condition9

det R(e1 × e2 · e3 ) = e1 × e2 · e3 ⇐⇒ det R = 1.

The tensors of Orth(V) that have a determinant equal to 1 form the subspace of proper
rotations or simply rotations, indicated by Orth(V)+ or also by SO(3). Only tensors of
Orth(V)+ represent rigid rotations of E 10 .

Theorem 18. Each tensor R ∈ Orth(V) has the eigenvalue ±1, with +1 for rotations.

Proof. Let u be an eigenvector of R ∈ Orth(V) corresponding to the eigenvalue λ. Be-


cause R preserves the norm, we have

Ru · Ru = λ2 u2 = u2 → λ2 = 1.

We must now prove that there exists at least one real eigenvector λ. To this end, we
consider the characteristic equation

f (λ) = λ3 + k1 λ2 + k2 λ + k3 = 0,

whose coefficients ki are real-valued, because R has real-valued components. It is imme-


diate to recognize that
lim f (λ) = ±∞.
λ→±∞

9
From the condition R⊤ R = I and through Eq. (2.16) and the theorem of Binet, we recognize
immediately that det R = ±1 ∀R ∈ Orth(V).
10
A tensor S ∈ Orth(V) such that det S = −1 represents a transformation that changes the orientation
of the space, like mirror symmetries do, see Section 2.12.

31
So, because f (λ) is a real-valued continuous function, actually a polynomial of λ, there
exists at least one λ1 ∈ R such that

f (λ1 ) = 0.

In addition, we already know that ∀R ∈ Orth(V), det R = ±1 and that, if λi , i = 1, 2, 3


are the eigenvalues of R, then det R = λ1 λ2 λ3 . Hence, the following two are the possible
cases:
i. λ1 ∈ R and λ2 , λ3 ∈ C, with λ3 = λ2 , the complex conjugate of λ2 ;
ii. λi ∈ R ∀i = 1, 2, 3.
Let us consider the case of R ∈ Orth(V)+ , i.e. a (proper) rotation → det R = 1. Then,
in the first case above,

det R = λ1 λ2 λ2 = λ1 [ℜ2 (λ2 ) + ℑ2 (λ2 )].

But
ℜ2 (λ2 ) + ℑ2 (λ2 ) = 1
because it is the square of the modulus of the complex eigenvalue λ2 . So in this case

det R = 1 ⇐⇒ λ1 = 1.

In the second case, λi ∈ R ∀i = 1, 2, 3, either λ1 > 0, λ2 , λ3 < 0, or all of them are


positive. Because the modulus of each eigenvalue must be equal to 1, either λ1 = 1 or
λi = 1 ∀i = 1, 2, 3 (in this case, R = I).
Following the same steps, one can easily show that ∀S ∈ Orth(V) with det S = −1, there
exists at least one real eigenvalue λ1 = −1.

Generally, a rotation tensor rotates the basis B = {e1 , e2 , e3 } into the basis B̂ = {ê1 , ê2 , ê3 }:

Rei = êi ∀i = 1, 2, 3 ⇒ Rij = ei · Rej = ei · êj . (2.38)

This result actually means that the jth column of R is formed by the components in the
basis B of the vector êj of B̂. Because the two bases are orthonormal, such components
are the director cosines of the axes of B̂ with respect to B.
Geometrically, any rotation is characterized by an axis of rotation w, |w| = 1, and by an
amplitude φ, i.e. the angle through which the space is rotated about w. By definition, w
is the (only) vector that is left unchanged by R, i.e.

Rw = w,

or, in other words, it is the eigenvector corresponding to the eigenvalue +1.


The question is then: How can a rotation tensor R be expressed by means of its geomet-
rical parameters, w and φ? To this end we have a fundamental theorem:

32
Theorem 19. (Euler’s rotation representation theorem). ∀R ∈ Orth(V)+ ,

R = I + sin φW + (1 − cos φ)W2 (2.39)

with φ the rotation’s amplitude and W the axial tensor of the rotation axis w.

Proof. We observe preliminarily that

Rw = Iw + sin φWw + (1 − cos φ)WWw = Iw = w (2.40)

i.e. that Eq. (2.39) actually defines a transformation that leaves unchanged the axis w,
like a rotation about w must do, and that +1 is an eigenvalue of R.
We need now to prove that Eq. (2.39) actually represents a rotation tensor, i.e. we must
prove that
RR⊤ = I, det R = 1.
Through Eq. (2.35) we get

RR⊤ = (I + sin φW + (1 − cos φ)W2 )(I + sin φW + (1 − cos φ)W2 )⊤


= (I + sin φW + (1 − cos φ)W2 )(I − sin φW + (1 − cos φ)W2 )
= I + 2(1 − cos φ)W2 − sin2 φW2 + (1 − cos φ)2 W4
= I + 2(1 − cos φ)W2 − sin2 φW2 − (1 − cos φ)2 W2 = I.

Then, through Eq. (2.34), we obtain

R = I + sin φW + (1 − cos φ)W2


= I + sin φW − (1 − cos φ)(I − w ⊗ w) (2.41)
= cos φI + sin φW + (1 − cos φ)w ⊗ w.

To go on, we need to express W and w ⊗ w; if w = (w1 , w2 , w3 ), then by Eq. (2.23) we


have  
0 −w3 w2
W =  w3 0 −w1 
−w2 w1 0
and by Eq. (2.2),  
w12 w1 w2 w1 w3
w ⊗ w =  w1 w2 w22 w2 w3  ,
w1 w3 w2 w3 w32
which on injecting into Eq. (2.41) gives
 
cos φ+(1−cos φ)w12 −w3sin φ+w1 w2 (1−cos φ) w2 sin φ+w1 w3 (1−cos φ)
R=  w3 sin φ+w1 w2 (1−cos φ) cos φ+(1−cos φ)w22 −w1 sin φ+w2 w3 (1−cos φ) .
−w2 sin φ+w1 w3 (1−cos φ) w1 sin φ+w2 w3 (1−cos φ) cos φ+(1−cos φ)w32
(2.42)

33
This formula gives R as a function exclusively of w and φ, the geometrical elements of
the rotation. Then
det R = (w2 + (1 − w2 ) cos φ)(cos2 φ + w2 sin2 φ)
and because w = 1, det R = 1, which proves that Eq. (2.39) actually represents a rotation.
We eventually need to prove that Eq. (2.39) represents the rotation about w of amplitude
φ. To this end, we choose an orthonormal basis B = {e1 , e2 , e3 } of V such that w = e3 ,
i.e. we analyze the particular case of a rotation of amplitude φ about e3 . This is always
possible thanks to the arbitrariness of the basis of V. In such a case, Eq. (2.38) gives
 
cos φ − sin φ 0
R =  sin φ cos φ 0  . (2.43)
0 0 1
Moreover,
   
0 −1 0 0 0 0
W =  1 0 0 , w ⊗ w =  0 0 0 ,
0 0 0 0 0 1
 
−1 0 0
W2 = −(I − w ⊗ w) =  0 −1 0 .
0 0 0

Hence
   
1 0 0 0 −1 0
I + sin φW + (1 − cos φ)W2 =  0 1 0  + sin φ  1 0 0  +
0 0 1 0 0 0
    (2.44)
−1 0 0 cos φ − sin φ 0
+ (1 − cos φ)  0 −1 0  =  sin φ cos φ 0  = R.
0 0 0 0 0 1

Equation (2.39) gives another result: To obtain the inverse of R it is sufficient to change
the sign of φ. In fact, because W ∈ Skw(V) and through Eq. (2.33)
R−1 = R⊤ = (I + sin φW + (1 − cos φ)W2 )⊤ = I + sin φW⊤ + (1 − cos φ)(W2 )⊤
= I − sin φW + (1 − cos φ)W2 = I + sin(−φ)W + (1 − cos(−φ))W2 .

The knowledge of the inverse of a rotation also allows us to perform the operation of
change of basis, i.e. to determine the components of a vector or of a tensor in a basis
B̂ = {ê1 , ê2 , ê3 } rotated with respect to an original basis B = {e1 , e2 , e3 } by a rotation R
(in the following equations, the symbol ˆ indicates a quantity specified in the basis B̂).
Considering that
ei = R−1 êi = R⊤ êi = Rhk
⊤ ⊤
(êh ⊗ êk )êi = Rhk δki êh

34
we get, for a vector u,

u = ui ei = Rki ui êk
i.e.

ûk = Rki ui → û = R⊤ u.
We remark that, because R⊤ = R−1 , the operation of change of basis is just the opposite
of the rotation of the space (and actually, we have seen that it is sufficient to take the
opposite of φ in Eq. (2.39) to get R−1 ).
For a second-rank tensor L we get
⊤ ⊤ ⊤ ⊤
L = Lij ei ⊗ ej = Lij Rmi êm ⊗ Rnj ên = Rmi Rnj Lij êm ⊗ ên ,

i.e.
⊤ ⊤ ⊤
L̂mn = Rmi Rnj Lij = Rmi Lij Rjn → L̂ = R⊤ LR.

We remark something that is typical of tensors: the components of a r-rank tensor in a


rotated basis B̂ depend upon the rth powers of the directors cosines of the axes of B̂, i.e.
on the rth powers of the components Rij of R.
If a rotation tensor is known through its Cartesian components in a given basis B, it is
easy to calculate its geometrical elements: The rotation axis w is the eigenvector of R
corresponding to the eigenvalue 1, so it is found solving the equation

Rw = w

and then normalizing it, while the rotation amplitude φ can be found using (2.39) along
with (2.34): Because the trace of a tensor is an invariant, we get
trR − 1
trR = 3 + (1 − cos φ)tr(−I + w ⊗ w) = 1 + 2 cos φ → φ = arccos .
2
It is interesting to consider the geometrical meaning of Eq. (2.39). For this purpose, we
apply Eq. (2.39) to a vector u, see Fig. 2.3,

Ru = (I + sin φW + (1 − cos φ)W2 )u


= u + sin φw × u + (1 − cos φ)w × (w × u)
The rotated vector Ru is the sum of three vectors; in particular, sin φWu is always
orthogonal to u, w and (1 − cos φ)W2 u. If u · w = 0, then (1 − cos φ)W2 u is also parallel
to u, see the sketch on the right in Fig. 2.3.
Let us consider now a composition of rotations. In particular, let us imagine that a vector
u is rotated first by R1 around w1 through φ1 , then by R2 around w2 through φ2 . So,
first, the vector u becomes the vector

u1 = R1 u.

Then, the vector u1 is rotated about w2 through φ2 to become

u12 = R2 u1 = R2 R1 u.

35
w (1-cos𝜑) W2u (1-cos𝜑) W2u
sin𝜑 Wu

rotation axis
Ru Ru sin𝜑 Wu
u
𝜑
𝜑
u

Figure 2.3: Rotation of a vector.

Let us now suppose that we change the order of the rotations: R2 first and then R1 . The
final result will be the vector
u21 = R1 R2 u. (2.45)
Because the tensor product is not symmetric (i.e., it has not the commutativity property),
generally11
u12 ̸= u21 .
In other words, the order of the rotations matters: Changing the order of the rotations
leads to a different final result. An example is shown in Fig. 2.4.

z z
rotation of 90°
about axis y
rotation of 90°
about axis z o y o y
z
x x

o y z z
x

rotation of 90° o y o y
about axis y rotation of 90°
x x
about axis z

Figure 2.4: Non-commutativity of the rotations.

This is a fundamental difference between rotations and displacements, that commute, see
Fig. 1.2, because the composition of displacements is ruled by the sum of vectors:

w =u+v =v+u (2.46)


11
We have seen in Theorem 12, that two tensors commute ⇐⇒ they are coaxial, i.e. if they have
the same eigenvectors. Because the rotation axis is always a real eigenvector of a rotation tensor, if two
tensors operate a rotation about different axes, they are not coaxial. Hence, the rotation tensors about
different axes never commute.

36
This difference, which is a major point in physics, comes from the difference in the oper-
ators: vectors for the displacements and tensors for the rotations.
Any rotation can be specified by the knowledge of three parameters. This can be easily
seen from Eq. (2.39): the parameters are the three components of w, that are not
independent because q
w = |w| = w12 + w22 + w32 = 1
and by the amplitude angle φ. The choice of the parameters by which to express a rotation
is not unique. Besides the use of the Cartesian components of w and φ, cf. Eq. (2.42),
other choices are possible, let us see three of them:
i. Physical angles: The rotation axis w is given through its spherical coordinates ψ,
the longitude, 0 ≤ ψ < 2π, and θ, the colatitude, 0 ≤ θ ≤ π, see Fig. 2.5, the third
parameter being the rotation amplitude φ. Then

𝜃
w
o
y
𝜓
x

Figure 2.5: Physical angles.

w2
w = (sin θ cos ψ, sin θ sin ψ, cos θ) → θ = arccos w3 , ψ = arctan ,
w1

and, Eq. (2.42),


 
cψ 2 sθ2 + cφ(cθ2 + sψ 2 sθ2 ) sψcψsθ2 (1 − cφ) − cθsφ cψsθcθ(1 − cφ) + sψsθsφ
R =  sψcψsθ2 (1 − cφ) + cθsφ sψ 2 sθ2 + cφ(cθ2 + cψ 2 sθ2 ) sψsθcθ(1 − cφ) − cψsθsφ  ,
cψsθcθ(1 − cφ) − sψsθsφ sψsθcθ(1 − cφ) + cψsθsφ cθ2 + cφ(cψ 2 sθ2 + sψ 2 sθ2 )

where cψ = cos ψ, sψ = sin ψ, cθ = cos θ, sθ = sin θ, cφ = cos φ and sφ = sin φ. We


remark that all the components of R so expressed depend upon the first powers of
the circular functions of φ. Hence, for what is said above, with this representation of
the rotations, the components of a rotated r-rank tensor depend upon the rth power
of the circular functions of φ, i.e. of the physical rotation, but not of ψ nor of θ.
ii. Euler’s angles: In this case, the three parameters are the amplitude of three particular
rotations into which the rotation is decomposed. Such parameters are the angles ψ,
the precession, θ, the nutation, and φ, the proper rotation, see Fig. 2.6 These three
rotations are represented in Fig. 2.7. The first one, of amplitude ψ, is made about z
to carry the axis x onto the knots line xN , the line perpendicular to both the axes z

37
and ẑ, and y onto y; by Eq. (2.38), in the frame {x, y, z} it is
 
cos ψ − sin ψ 0
Rψ =  sin ψ cos ψ 0  .
0 0 1

The second one, of amplitude θ, is made about xN to carry z onto ẑ; in the frame
{xN , y, z}, it is  
1 0 0
Rθ =  0 cos θ − sin θ  ,
0 sin θ cos θ
while in the frame {x, y, z},

Roθ = (R−1 ⊤ −1 ⊤
ψ ) Rθ Rψ = Rψ Rθ Rψ .

The last rotation, of amplitude φ, is made about ẑ to carry xN onto x̂ and y onto ŷ;
in the frame {xN , y, ẑ}, it is
 
cos φ − sin φ 0
Rφ = sin φ cos φ
 0 ,
0 0 1

while in {x, y, z},

Roφ = (R−1 ⊤ −1 ⊤ −1 −1 ⊤ ⊤
ψ ) (Rθ ) Rφ Rθ Rψ = Rψ Rθ Rφ Rθ Rψ .

Any vector u is transformed, by the global rotation, into the vector

û = Ru.

But we can also write


û = Roφ u,

z
z

𝜃
y

y
𝜑
𝜓 x

x
xN
Figure 2.6: Euler’s angles.

38
y y


z
z


o y o
y
𝜃 y

𝜓 xN o y 𝜑 x


x xN

Figure 2.7: Euler’s rotations, as seen from the respective axes of rotation.

where u is the vector transformed by the rotation Roθ ,

u = Roθ u,

and u is the vector transformed by the rotation Rψ :

u = Rψ u.

Finally,
û = Ru = Roφ Roθ Rψ u → R = Roφ Roθ Rψ ,
i.e. the global rotation tensor is obtained composing, in the opposite order of execu-
tion of the rotations, the three tensors all expressed in the original basis. However,

R = Roφ Roθ Rψ = Rψ Rθ Rφ R⊤ ⊤ ⊤
θ Rψ Rψ Rθ Rψ Rψ = Rψ Rθ Rφ ,

i.e., the global rotation tensor is also equal to the composition of the three rotations,
in the order of execution, if the three rotations are expressed in their own particular
bases. This result is general, not bounded to the Euler’s rotations nor to three
rotations.
Performing the tensor multiplications, we get
 
cos ψ cos φ − sin ψ sin φ cos θ − cos ψ sin φ − sin ψ cos φ cos θ sin ψ sin θ
R =  sin ψ cos φ + cos ψ sin φ cos θ − sin ψ sin φ + cos ψ cos φ cos θ − cos ψ sin θ  .
sin φ sin θ cos φ sin θ cos θ

The components of a vector u in the basis B̂ are then given by

û = R⊤ u = R⊤ ⊤ ⊤
φ Rθ Rψ u,

and those of a second-rank tensor

L̂ = R⊤ LR = R⊤ ⊤ ⊤
φ Rθ Rψ LRψ Rθ Rφ .

iii. Coordinate angles: In this case, the rotation R is decomposed into three successive
rotations α, β, γ, respectively about the axes x, y and z of each rotation, i.e.

R = Rα Rβ Rγ

39
with
     
1 0 0 cos β 0 − sin β cos γ − sin γ 0
Rα =  0 cos α − sin α  , Rβ =  0 1 0  , Rγ =  sin γ cos γ 0  ,
0 sin α cos α sin β 0 cos β 0 0 1

so finally
 
cos β cos γ − cos β sin γ − sin β
R =  cos α sin γ − sin α sin β cos γ cos α cos γ + sin α sin β sin γ − sin α cos β  .
sin α sin γ + cos α sin β cos γ sin α cos γ − cos α sin β sin γ cos α cos β

Let us now consider the case of small rotations, i.e. |φ| → 0. In such a case,

sin φ ≃ φ, 1 − cos φ ≃ 0

so that the Euler’s theorem, Eq. (2.39), becomes

R ≃ I + φW,

i.e. in the small rotations approximation, any vector u is transformed into

Ru ≃ (I + φW)u = u + φw × u, (2.47)

i.e. by a skew tensor and not by a rotation tensor. The term (1 − cos φ)W2 u has disap-
peared, as it is a higher-order infinitesimal quantity, and the term φw × u is orthogonal
to u. Because φ → 0, the arc is approximated by its tangent, the vector φw × u, see
Fig. 2.8. Applying to Eq. (2.47) the procedure already seen for the composition of finite

w
𝜑 w×u
rotation axis

Ru u
Ru
𝜑 w×u
𝜑
𝜑
u

Figure 2.8: Small rotations.

amplitude rotations, we get

u1 = R1 u = (I + φ1 W1 )u = u + φ1 w1 × u,
u21 = R2 u1 = (I + φ2 W2 )u1 = u1 + φ2 w2 × u1
= u + φ1 w1 × u + φ2 w2 × u
+ φ1 φ2 w2 × (w1 × u).

40
If the order of the rotations is changed, the last term becomes φ1 φ2 w1 × (w2 × u), which
is, in general, different from φ1 φ2 w2 × (w1 × u): Strictly speaking, also small rotations do
not commute12 . However, for small rotations, φ1 φ2 is negligible with respect to φ1 and
φ2 : In this approximation, small rotations commute. We remark that the approximation
(2.47) gives, for the displacements, a law that is quite similar to that of the velocities of
the points of a rigid body:
v = v0 + ω × (p − o)
This is quite natural, because

, ω=
dt
i.e. a small amplitude rotation can be seen as the rotation made with finite angular
velocity ω in a small time interval dt.

2.12 Reflexions
Let us consider now tensors S ∈ Orth(V) that are not a rotation, i.e. such that det S = −1.
Let us call S an improper rotation. A particular improper rotation, whose all eigenvalues
are equal to -1, is the inversion or reflexion tensor:
SI = −I.
The effect of SI is to transform any basis B into the basis −B, i.e. with all the basis
vectors changed in orientation (or, equivalently, to change the sign of all the components
of a vector). In other words, SI changes the orientation of the space. This is also the
effect of any other improper rotation S, that can be decomposed into a proper rotation
R followed by the reflexion SI 13 :
S = SI R. (2.48)
Let n ∈ S; then
SR = I − 2n ⊗ n (2.49)
is the tensor that operates the transformation of symmetry with respect to a plane or-
thogonal to n. In fact
SR n = −n, SR m = m ∀m ∈ V| m · n = 0.
SR is an improper rotation; in fact, by Eq. (2.4),
(I − 2n ⊗ n)(I − 2n ⊗ n)⊤ = (I − 2n ⊗ n)(I − 2n ⊗ n)
= I − 2n ⊗ n − 2n ⊗ n + 4(n ⊗ n)(n ⊗ n) = I,
while by the same definition of trace and through Eqs. (2.13) and (2.17),
tr2 (n ⊗ n) − tr(n ⊗ n)(n ⊗ n)
det(I − 2n ⊗ n) = 1 − 2tr(n ⊗ n) + 4 − 8 det(n ⊗ n) = −1.
2
12
This can happen for some vectors, all the times that w1 · u = w2 · u, like for the case of a vector u
orthogonal to both w1 and w2 ; however, this is no more than a curiosity, it has no importance in practice.
13
The application of Binet’s theorem shows immediately that det S = −1, while SI R(SI R)⊤ =
SI RR⊤ S⊤ ⊤
I = −I(−I) = I: The decomposition in Eq. (2.48) actually gives an improper rotation.

41
Let S = SI R be an improper rotation. Then
⊤
(Su) × (Sv) = (SI Ru) × (SI Rv) = det(SI R) (SI R)−1 (u × v)


= det SI det R(R−1 S−1 ⊤ −1 ⊤


I ) (u × v) = −(−R I) (u × v) = R(u × v).

The transformation by S of any vector u gives

Su = SI Ru = −Ru,

i.e. it changes the orientation of the rotated vector; this is not the case when the same
improper rotations transforms the vectors of a cross product: The rotated vector, result
of the cross product, does not change of orientation, i.e. the cross product is insensitive
to a reflexion. That is why, strictly speaking, the result of a cross product is not a vector,
but a pseudo-vector: It behaves like vectors apart for the reflexions. For the same reason,
a scalar result of a mixed product (scalar plus cross product of three vectors) is called a
pseudo-scalar because in this case, the scalar result of the mixed product changes of sign
under a reflexion, which can be checked easily.

2.13 Polar decomposition

Theorem 20. (Square root theorem). Consider L ∈ Sym(V) and positive definite.
Then, there exists a unique tensor U ∈ Sym(V) and positive definite such that

L = U2 .

Proof. Existence: Consider L, U, V ∈ Sym(V) and positive definite, and

L = ωi ei ⊗ ei

a spectral decomposition of L, ωi > 0 ∀i. Define U as



U = ωi ei ⊗ ei ;

then, by Eq. (2.4)1 , we get


U2 = L.
Uniqueness: Suppose that also
V2 = L
and √
let e be an eigenvector of L corresponding to the (positive) eigenvalue ω. Then, if
λ = ω,
o = (U2 − ωI)e = (U + λI)(U − λI)e,
and once we set
v = (U − λI)e,
we get
Uv = −λv ⇒ v = o ⇒ Ue = λe

42
because U is positive definite and −λ cannot be an eigenvalue of U because λ > 0. In a
similar way,
Ve = λe ⇒ Ue = Ve
for every eigenvector e of L. Because, based on the spectral theorem, it exists a basis of
eigenvectors of L, U = V.

We symbolically write that √


U= L.

For any F ∈ Lin(V), both FF⊤ and F⊤ F clearly ∈ Sym(V). If in addition det F > 0,
then
u · F⊤ Fu = (Fu) · (Fu) ≥ 0
with the zero value obtained ⇐⇒ Fu = o or, what is equivalent, because det F > 0 ⇒ F
is invertible, ⇐⇒ u = o. As a consequence, F⊤ F is positive definite. In a similar way,
it can be proved that FF⊤ is also positive definite.
A particular tensor decomposition14 is given by the

Theorem 21. (Polar decomposition theorem). ∀F ∈ Lin(V)| det F > 0 exist, and
are uniquely determined, two positive definite tensors U, V ∈ Sym(V) and a rotation R
such that
F = RU = VR.

Proof. Uniqueness: Let F = RU be a right polar decomposition of F; because R ∈


Orth(V)+ and U ∈ Sym(V),

F⊤ F = UR⊤ RU = U2 ⇒ U = F⊤ F.

By the square-root theorem, tensor U is unique, and because

R = FU−1 ,

R is unique too.
Now, let F = VR be a left polar decomposition of F; by the same procedure, we get

FF⊤ = V2 → V = FF⊤ ,

so V is unique, and also,


R = V−1 F.
Existence: let √
U= F⊤ F
so U ∈ Sym(V) and it is positive definite, and let

R = FU−1 .
14
This decomposition is fundamental to the theory of deformation of continuum bodies.

43
To prove that F = RU is a right polar decomposition, we just have to show that R ∈
Orth(V)+ . Since det F > 0 and det U > 0 (the latter because all the eigenvalues of U are
strictly positive), by the theorem of Binet, also det R > 0. Then,

R⊤ R = (FU−1 )⊤ (FU−1 ) = U−1 F⊤ FU−1


= U−1 U2 U−1 = I ⇒ R ∈ Orth(V)+ .

Now, let
V = RUR⊤ ,
then V ∈ Sym(V) and is positive definite, see Exercise 22, and

VR = RUR⊤ R = RU = F,

which completes the proof.

2.14 Exercises
1. Prove that
Lo = o ∀L ∈ Lin(V).

2. Prove that, if a straight line r has the direction of u ∈ S, then the tensor giving the
projection of a vector v ∈ V on r is u ⊗ u (the orthogonal projector), while the one
giving the projection on a direction orthogonal to r is I − u ⊗ u (the complementary
projector), see Fig. 2.9.
3. For any α ∈ R, a, b ∈ V and A, B ∈ Lin(V), prove that

(αA)⊤ = αA⊤ , (A + B)⊤ = A⊤ + B⊤ , (a ⊗ b)A = a ⊗ (A⊤ b).

4. Prove that
L + O = L ∀L ∈ Lin(V).

5. Prove that
trI = 3, trO = 0.

r
u

(I-u⊗u)v
(u⊗u)v

Figure 2.9: Projected vectors.

44
6. Prove that, ∀A, B ∈ Lin(V),

tr(AB) = tr(BA).

7. Prove that, ∀L, M, N ∈ Lin(V),

L⊤ · M⊤ = L · M, LM · N = L · NM⊤ = M · L⊤ N.

8. Prove the assertions in Eq. (2.4).


9. Prove that any form defined by a tensor L can be written as a scalar product of
tensors:
v · Lw = L · v ⊗ w ∀v, w ∈ V, L ∈ Lin(V).

10. Prove that Sym(V) and Skw(V) are orthogonal, i.e. prove that

A · B = 0 ∀A ∈ Sym(V), B ∈ Skw(V).

11. For any L ∈ Lin(V), prove that, if A ∈ Sym(V), then

A · L = A · Ls ,

while if B ∈ Skw(V), then


B · L = B · La .

12. Let A, B, C, D ∈ Lin(V); prove that

A · (BCD) = (B⊤ A) · (CD) = (AD⊤ ) · (BC).

13. Prove that L · W = 0 ∀W ∈ Skw(V) ⇐⇒ L ∈ Sym(V).


14. Express by components the second principal invariant I2 of a tensor L.
15. Prove that, if a = (a1 , a2 , a3 ), b = (b1 , b2 , b3 ), c = (c1 , c2 , c3 ), then
 
a1 a2 a3
a × b · c = det  b1 b2 b3  .
c1 c2 c3

16. Prove the uniqueness of the inverse tensor.


17. Show, using the Cartesian components, that all the dyads are singular.
18. Prove that if L is invertible and α ∈ R − {0} then

(αL)−1 = α−1 L−1 .

19. Prove that if W is the axial tensor of w, then


1
WW = − |W|2 (I − w ⊗ w).
2
45
20. Prove that for any two axial couples w1 , W1 and w2 , W2 , we have:
1
w1 · w2 = W1 · W2 .
2

21. Prove that ∀u, v ∈ V, u × v = o ⇐⇒ u ⊗ v ∈ Sym(V).


22. Let L ∈ Sym(V) and positive definite and R ∈ Orth(V)+ , then prove that RLR⊤ ∈
Sym(V) and that it is positive definite.
23. Prove that the spectrum of Lsph is composed of only
1
λsph = trL,
3
and that any u ∈ S is an eigenvector.
24. Prove that the eigenvalues λdev of Ldev are given by

λdev = λ − λsph ,

where λ is an eigenvalue of L.

46
Chapter 3

Fourth rank tensors

3.1 Fourth-rank tensors


A fourth-rank tensor L is any linear application from Lin(V) to Lin(V):

L : Lin(V) → Lin(V) | L(αi Ai ) = αi LAi ∀αi ∈ R, Ai ∈ Lin(V), i = 1, ..., n.

Defining the sum of two fourth-rank tensors as

(L1 + L2 )A = L1 A + L2 A ∀A ∈ Lin(V),

the product of a scalar by a fourth-rank tensor as

(αL)A = α(LA) ∀α ∈ R, A ∈ Lin(V)

and the null fourth-rank tensor O as the unique tensor such that

OA = O ∀A ∈ Lin(V),

then the set of all the tensors L that operate on Lin(V) forms a vector space, denoted by
Lin(V). We define the fourth-rank identity tensor I as the unique tensor such that

IA = A ∀A ∈ Lin(V).

It is apparent that the algebra of fourth-rank tensors is similar to that of second-rank


tensors and, in fact, several operations with fourth-rank tensors can be introduced in
almost the same way, in some sense shifting from V to Lin(V) the operations. However,
the algebra of fourth-rank tensors is richer than that of the second-rank ones and some
care must be taken.
In the following sections, we consider some of the operations that can be done with
fourth-rank tensors.

47
3.2 Dyads, tensor components
For any couple of tensors A and B ∈ Lin(V), the (tensor) dyad A ⊗ B is the fourth-rank
tensor defined by
(A ⊗ B)L := B · L A ∀L ∈ Lin(V).
The application defined above is actually a fourth-rank tensor because of the bilinearity
of the scalar product of second-rank tensors. Applying this rule to the nine dyads of the
basis B 2 = {ei ⊗ej , i, j = 1, 2, 3} of Lin(V) leads to the introduction of the 81 fourth-rank
tensors
ei ⊗ ej ⊗ ek ⊗ el := (ei ⊗ ej ) ⊗ (ek ⊗ el )
that form a basis, B 4 = {ei ⊗ ej ⊗ ek ⊗ el , i, j = 1, 2, 3}, for Lin(V). We remark hence
that dim(Lin(V)) = 81. A useful result is that
(ei ⊗ ej ⊗ ek ⊗ el )(ep ⊗ eq ) = (ek ⊗ el ) · (ep ⊗ eq )(ei ⊗ ej ) = δkp δlq (ei ⊗ ej ). (3.1)
Any fourth-rank tensor can be expressed as a linear combination (the canonical decom-
position):
L = Lijkl ei ⊗ ej ⊗ ek ⊗ el , i, j = 1, 2, 3,
where the Lijkl s are the 81 Cartesian components of L with respect to B 4 . The Lijkl s are
defined by the operation
(ei ⊗ ej ) · L(ek ⊗ el ) = (ei · ej ) · (Lpqrs ep ⊗ eq ⊗ er ⊗ es )(ek ⊗ el )
= (ei ⊗ ej ) · (Lpqrs δrk δsl ep ⊗ eq )
= Lpqrs δrk δsl δip δjq = Lijkl .
The components of a tensor dyad can be computed without any difficulty:
A ⊗ B = (Aij ei ⊗ ej ) ⊗ (Bkl ek ⊗ el ) = Aij Bkl ei ⊗ ej ⊗ ek ⊗ el ⇒
(A ⊗ B)ijkl = Aij Bkl ,
so that, in particular,
((a ⊗ b) ⊗ (c ⊗ d))ijkl = ai bj ck dl .
Concerning the identity of Lin(V),
Iijkl = (ei ⊗ el ) · I(ek ⊗ el ) = (ei ⊗ ej ) · (ek ⊗ el ) = ei · ek ej · el = δik δjl ⇒
I = δik δjl (ei ⊗ el ⊗ ek ⊗ el ).
The components of A ∈ Lin(V), resulting from the application of L ∈ Lin(V) on B ∈
Lin(V), can now be easily calculated:
A = LB = Lijkl (ei ⊗ ej ⊗ ek ⊗ el )(Bpq ep ⊗ eq )
= Lijkl Bpq δkp δlq (ei ⊗ ej ) (3.2)
= Lijkl Bkl (ei ⊗ ej ) ⇒ Aij = Lijkl Bkl .
Moreover,
L(A ⊗ B)C = L((A ⊗ B)C) = L(B · CA) = B · C LA = ((LA) ⊗ B)C ⇒
L(A ⊗ B) = (LA) ⊗ B.

48
Using this result and Eq. (3.1), we can determine the components of a product of fourth-
rank tensors:
AB = Aijkl (ei ⊗ ej ⊗ ek ⊗ el )Bpqrs (ep ⊗ eq ⊗ er ⊗ es )
= Aijkl Bpqrs (ei ⊗ ej ⊗ ek ⊗ el )(ep ⊗ eq ) ⊗ (er ⊗ es )
= Aijkl Bpqrs [(ei ⊗ ej ⊗ ek ⊗ el )(ep ⊗ eq )] ⊗ (er ⊗ es ) (3.3)
= Aijkl Bpqrs [δkp δlq (ei ⊗ ej )] ⊗ (er ⊗ es )
= Aijkl Bklrs (ei ⊗ ej ⊗ er ⊗ es ) ⇒ (AB)ijrs = Aijkl Bklrs .

Depending upon four indices, a fourth-rank tensor L cannot be represented by a matrix;


however, we will see in Section 3.8 that a matrix representation of a fourth-rank tensor is
still possible and that it is currently used in some cases, e.g. in elasticity.

3.3 Conjugation product, transpose, symmetries


For any two tensors A, B ∈ Lin(V) we call conjugation product the tensor A⊠B ∈ Lin(V)
defined by the operation

(A ⊠ B)L := ALB⊤ ∀L ∈ Lin(V).

As a consequence, for the dyadic tensors of B 2 ,

(ei ⊗ ej ) ⊠ (ek ⊗ el ) = ei ⊗ ek ⊗ ej ⊗ el , (3.4)

so that
(A ⊠ B)ijkl = Aik Bjl .
Moreover, by the uniqueness of the identity I, ∀A ∈ Lin(V),

(I ⊠ I)A = IAI⊤ = A ⇒ I = I ⊠ I.

The transpose of a fourth-rank tensor L is the unique tensor L⊤ such that

A · (LB) = B · (L⊤ A) ∀A, B ∈ Lin(V).

By this definition, setting A = ei ⊗ ej , B = ek ⊗ el gives

(L⊤ )ijkl = Lklij .

A consequence is that

A · (LB) = B · (L⊤ A) = A · (L⊤ )⊤ B ⇒ (L⊤ )⊤ = L.

Moreover,

M · (A ⊗ B)⊤ L = L · (A ⊗ B)M
= L · AM · B = M · (BA · L)
= M · (B ⊗ A)L ⇒ (A ⊗ B)⊤ = B ⊗ A,

49
while, cf. Exercise 7, Chapter 2,

M · (A ⊠ B)⊤ L = L · (A ⊠ B)M
= L · AMB⊤ = A⊤ L · MB⊤ = M⊤ A⊤ L · B⊤
= (M⊤ A⊤ L)⊤ · (B⊤ )⊤ = L⊤ AM · B = AM · LB
= M · A⊤ LB = M · (A⊤ ⊠ B⊤ )L ⇒
(A ⊠ B)⊤ = A⊤ ⊠ B⊤ .

The property
(AB)⊤ = B⊤ A⊤
can be proved in the same manner used for the analogous property of the second-rank
tensors.
A tensor L ∈ Lin(V) is symmetric ⇐⇒ L = L⊤ . It is then evident that

L = L⊤ ⇒ Lijkl = Lklij ,

relations that are known as major symmetries. There are 36 major symmetries on the
whole, so that a symmetric fourth-rank tensor has 45 independent components. More-
over,

A ⊠ B = (A ⊠ B)⊤ = A⊤ ⊠ B⊤ ⇐⇒ A = A⊤ , B = B⊤ ,
A ⊗ B = (A ⊗ B)⊤ = B ⊗ A ⇐⇒ B = λA, λ ∈ R.

Let us now consider the case of a L ∈ Lin(V) such that

LA = (LA)⊤ ∀A ∈ Lin(V).

Then, by Eq. (3.2),


Lijkl = Ljikl ,
relations that are called left minor symmetries: a tensor L having the left minor symme-
tries has values in Sym(V). On the whole, the number of left minor symmetries is 27.
Finally, consider the case of a L ∈ Lin(V) such that

LA = L(A⊤ ) ∀A ∈ Lin(V);

then, again by Eq. (3.2), we get


Lijkl = Lijlk ,
which are relations called minor right-symmetries, whose total number is also 27. It is
immediate to recognize that if L has the minor right-symmetries, then

LW = O ∀W ∈ Skw(V).

We say that a tensor has the minor symmetries if it has both the right and left minor
symmetries; the total number of minor symmetries is 45, because, as can be easily checked,

50
some of the left and right minor symmetries are the same, so finally a tensor with the
minor symmetries has 36 independent components.
If L ∈ Lin(V) has the major and minor symmetries, then the number of independent
symmetry relations is actually 60 (some minor and major symmetries coincide), so in
such a case L depends upon 21 independent components only. This is the case, for
instance, of the classical elasticity tensor.
Finally, the 6 Cauchy-Poisson symmetries1 are those of the type

Lijkl = Likjl .

A tensor having the major, minor and Cauchy-Poisson symmetries is said to be completely
symmetric, i.e. swapping any couple of indices gives an identical component. In that case,
the number of independent components is only 15.

3.4 Trace and scalar product of fourth-rank tensors


We can introduce the scalar product between fourth-rank tensors in the same way we did
for second-rank tensors. We first introduce the concept of trace for fourth-rank tensors,
denoted by tr4 , once again using the dyad (here, the tensor dyad):

tr4 A ⊗ B := A · B.

The easy proof that tr4 : Lin(V) → R is a linear form is based upon the properties of the
scalar product of second-rank tensors and it is left to the reader. An immediate result is
that
tr4 A ⊗ B = Aij Bij ,
Then, using the canonical decomposition, we have that

tr4 L = tr4 (Lijkl (ei ⊗ ej ) ⊗ (ek ⊗ el )) = Lijkl (ei ⊗ ej ) · (ek ⊗ el ) = Lijkl δik δjl = Lijij

and that

tr4 L⊤ = tr4 (Lklij (ei ⊗ ej ) ⊗ (ek ⊗ el )) = Lklij (ei ⊗ ej ) · (ek ⊗ el ) = Lklij δik δjl = Lijij = tr4 L.

Then, we define the scalar product of fourth-rank tensors as

A · B := tr4 (A⊤ B).

By the properties of tr4 , the scalar product is a positive definite symmetric bilinear
form:
αA · βB = tr4 (αA⊤ βB) = αβtr4 (A⊤ B) = αβA · B,
A · B = tr4 (A⊤ B) = tr4 (A⊤ B)⊤ = tr4 (B⊤ A) = B · A,
A · A = tr4 (A⊤ A) = (A⊤ A)ijij = Aklij Aklij > 0 ∀A ∈ Lin(V), A · A = 0 ⇐⇒ A = O.
1
The Cauchy-Poisson symmetries have played an important role in a celebrated diatribe of the XIXth
century in elasticity, that between the so-called rari- and muti-constant theories.

51
By components

A · B = tr4 ((Aklij ei ⊗ ej ⊗ ek ⊗ el )(Bpqrs ep ⊗ eq ⊗ er ⊗ es ))


= tr4 (Aklij Bpqrs δkp δlq (ei ⊗ ej ) ⊗ (er ⊗ es ))
= Aklij Bpqrs δkp δlq (ei ⊗ ej ) · (er ⊗ es ) = Aklij Bpqrs δkp δlq δir δjs = Aklij Bklij .

The rule for computing the scalar product is hence always the same, as was already seen
for vectors and second-rank tensors: All the indexes are to be saturated.
In complete analogy with vectors and second-rank tensors, we say that A is orthogonal to
B ⇐⇒
A·B=0
and we define the norm of L as
√ p p
|L| := L · L = tr4 L⊤ L = Lijkl Lijkl .

3.5 Projectors, identities


For the spherical part of any A ∈ Sym(V) we can write
1 1 1
Asph := trA I = I · A I = (I ⊗ I)A = Ssph A,
3 3 3
where
1
Ssph := I ⊗ I
3
is the spherical projector, i.e. the fourth-rank tensor that extracts from any A ∈ Lin(V)
its spherical part. Moreover,

Adev := A − Asph = IA − Ssph A = Ddev A,

where
Ddev := I − Ssph
is the deviatoric projector, i.e. the fourth-rank tensor that extracts from any A ∈ Lin(V)
its deviatoric part. It is worth noting that

I = Ssph + Ddev .

Moreover, about the components of Ssph ,

sph 1 1
Sijkl = (ei ⊗ ej ) · (I ⊗ I)(ek ⊗ el ) = I · (ei ⊗ ej ) I · (ek ⊗ el )
3 3
1 1 1
= tr(ei ⊗ ej )tr(ek ⊗ el ) = δij δkl → Ssph = δij δkl (ei ⊗ ej ⊗ ek ⊗ el ).
3 3 3
We remark that
Ssph = (Ssph )⊤ .

52
We introduce now the tensor Is , restriction of I to A ∈ Sym(V). It can be introduced as
follows: ∀A ∈ Sym(V)
1
A = (A + A⊤ ),
2
and
1 1
A = IA = (IA + IA⊤ ) = (Iijkl Akl + Iijkl Alk )(ei ⊗ ej ⊗ ek ⊗ el );
2 2

because A = A , there is insensitivity to the swap of indexes k and l, so
1 1
A = (Iijkl Akl + Iijlk Alk )(ei ⊗ ej ⊗ ek ⊗ el ) = (δik δjl + δil δjk )Akl (ei ⊗ ej ⊗ ek ⊗ el ).
2 2
Then, if we admit the interchangeability of indexes k and l, i.e. if we postulate the
existence of the minor right-symmetries for I, then I = Is , with
1
Is = (δik δjl + δil δjk )(ei ⊗ ej ⊗ ek ⊗ el ).
2
It is apparent that
s s
Iijkl = Iklij ,
i.e. Is = (Is )⊤ , but also that

s 1 s
Iijkl = (δil δjk + δik δjl ) = Ijikl ,
2
i.e., Is has also the minor left-symmetries; in other words, Is has the major and minor
symmetries, like an elasticity tensor, while this is not the case for I. In fact

Iijkl = Ijilk = δik δjl ̸= δil δjk = Ijikl = Iijlk .

Because Ssph and Ddev operate on Sym(V), it is immediate to recognize that it is also

Ddev = Is − Ssph ⇒ Is = Ssph + Ddev .

It is worth noting that

(Ddev )⊤ = (Is − Ssph )⊤ = (Is )⊤ − (Ssph )⊤ = Is − Ssph = Ddev .

We can now determine the components of Ddev :

dev s sph 1 1
Dijkl = Iijkl − Sijkl = (δik δjl + δil δjk ) − δij δkl →
 2  3
1 1
Ddev = (δik δjl + δil δjk ) − δij δkl (ei ⊗ ej ⊗ ek ⊗ el ).
2 3

We remark that the result (2.9) implies that Ssph and Ddev are orthogonal projectors, i.e.
they project the same A ∈ Sym(V) into two orthogonal subspaces of V, Sph(V) and
Dev(V).
The tensor Ttrp ∈ Lin(V) defined by the operation

Ttrp A = A⊤ ∀A ∈ Lin(V),

53
is the transposition projector, whose components are
trp
Tijkl = (ei ⊗ ej ) · Ttrp (ek ⊗ el ) = (ei ⊗ ej ) · (el ⊗ ek ) = δil δjk .

The following operation defines the symmetry projector Ssym ∈ Lin(V):

1
Ssym A = (A + A⊤ ) ∀A ∈ Lin(V),
2

while the antisymmetry projector Wskw ∈ Lin(V) is defined by

1
Wskw A = (A − A⊤ ) ∀A ∈ Lin(V).
2

Also Ssym and Wskw are orthogonal projectors, because they project the same A ∈ Lin(V)
into two orthogonal subspaces of Lin(V): Sym(V) and Skw(V), see Exercise 10, Chapter
2.
We prove now two properties of the projectors: ∀A ∈ Lin(V),

1 1
(Ssym + Wskw )A = (A + A⊤ ) + (A − A⊤ ) = A = IA ⇒ Ssym + Wskw = I. (3.5)
2 2
Then,

1 1
(Ssym −Wskw )A = (A+A⊤ )− (A−A⊤ ) = A⊤ = Ttrp A ⇒ Ssym −Wskw = Ttrp . (3.6)
2 2

3.6 Orthogonal conjugator


For any U ∈ Orth(V), we define its orthogonal conjugator U ∈ Lin(V) as

U := U ⊠ U.

Theorem 22. (Orthogonality of U). The orthogonal conjugator is an orthogonal tensor


of Lin(V), i.e. it preserves the scalar product between tensors:

UA · UB = A · B ∀A, B ∈ Lin(V).

Proof. By the assertion in Exercise 12 of Chapter 2 and because U ∈ Orth(V), we have

UA · UB = (U ⊠ U)A · (U ⊠ U)B = UAU⊤ · UBU⊤


= U⊤ UAU⊤ · BU⊤ = AU⊤ · BU⊤ = AU⊤ U · B = A · B.

54
Just as for tensors of Orth(V), we also have

UU⊤ = U⊤ U = I.

In fact, see the assertion in Exercise 4:

UU⊤ = (U ⊠ U)(U⊤ ⊠ U⊤ ) = UU⊤ ⊠ UU⊤ = I ⊠ I = I. (3.7)

The orthogonal conjugators also have some properties in relation with projectors:

Theorem 23. Ssph is unaffected by any orthogonal conjugator, while Ddev commutes with
any orthogonal conjugator.

Proof. For any L ∈ Sym(V) and U ∈ Orth(V),


 
1 1 1
US sph
L = (U ⊠ U) I ⊗ I L = (trL)(U ⊠ U)I = (trL)UIU⊤
3 3 3
1 1 1
= (trL)I = I · L I = (I ⊗ I)L = Ssph L.
3 3 3

Moreover,
 
1 1 1
S sph
UL = I ⊗ I (U ⊠ U)L = (I ⊗ I)(ULU⊤ ) = (I · ULU⊤ )I
3 3 3
1 1 1 1 1
= tr(ULU⊤ )I = tr(U⊤ UL)I = (trL)I = I · LI = (I ⊗ I)L = Ssph L.
3 3 3 3 3

Thus, we have proved that


Ssph U = USsph = Ssph ,

i.e. that the spherical projector Ssph is unaffected by any orthogonal conjugator. Further-
more

Ddev UL = (Is − Ssph )UL = Is UL − Ssph UL = UL − Ssph L = (U − Ssph )L

and

UDdev L = U(Is − Ssph )L = UIs L − USsph L = UL − Ssph L = (U − Ssph )L,

so that
Ddev U = UDdev .

55
3.7 Rotations and symmetries
We ponder now how to rotate a fourth-rank tensor, i.e., what are the components of

L = Lijkl ei ⊗ ej ⊗ ek ⊗ el

in a basis B ′ = {e′1 , e′2 , e′3 } obtained rotating the basis B = {e1 , e2 , e3 } by the rotation
R = Rij ei ⊗ ej , R ∈ Orth(V)+ . The procedure is exactly the same already followed for
vectors and second-rank tensors:
⊤ ′ ⊤ ′ ⊤ ′ ⊤ ′
L = Lijkl ei ⊗ ej ⊗ ek ⊗ el = Lijkl Rpi ep ⊗ Rqj eq ⊗ Rrk er ⊗ Rsl es
⊤ ⊤ ⊤ ⊤
= Rpi Rqj Rrk Rsl Lijkl e′p ⊗ e′q ⊗ e′r ⊗ e′s ,
i.e.
⊤ ⊤ ⊤ ⊤
L′pqrs = Rpi Rqj Rrk Rsl Lijkl .
We see clearly that the components of L in the basis B ′ are a linear combination of those
in B, the coefficients of the linear combination being fourth powers of the director cosines,
the Rij s. The introduction of the orthogonal conjugator2 of the rotation R,

R = R ⊠ R,

allows us to give a compact expression for the rotation of second- and fourth-rank tensors
(for completeness we recall also that of a vector w);

w′ = R⊤ w,
L′ = R⊤ LR = (R⊤ ⊠ R⊤ )L = R⊤ L,
L′ = (R⊤ ⊠ R⊤ )L(R ⊠ R) = R⊤ LR.
Checking the above relations with the orthogonal conjugator R is left to the reader. It is
worth noting that, actually, these transformations are valid not only for R ∈ Orth(V)+ ,
but more generally for any U ∈ Orth(V), i.e. also for symmetries.
If U denotes the tensor of change of basis under any orthogonal transformation, i.e. if we
put U = R⊤ for the rotations, then the above relations become
w′ = Uw,
L′ = ULU⊤ = (U ⊠ U)L = UL, (3.8)
L′ = (U ⊠ U)L(U ⊠ U)⊤ = ULU⊤ .

Finally, we say that L ∈ Lin(V) or L ∈ Lin(V) is invariant under an orthogonal transfor-


mation U if
ULU⊤ = L, ULU⊤ = L;
right multiplying both terms by U or by U and through Eq. (3.7), we get that L or L are
invariant under U ⇐⇒
UL = LU, UL = LU,
2
Here the symbol R indicates the orthogonal conjugator of R, not the set of real numbers.

56
i.e. ⇐⇒ L and U, or L and U commute. This relation allows, for example, the analysis
of material symmetries in anisotropic elasticity.

If a tensor is invariant under any orthogonal transformation, i.e. if the previous equations
hold true ∀U ∈ Orth(V), then the tensor is said to be isotropic. A general result3 is
that a fourth-rank tensor L is isotropic ⇐⇒ there exist two scalar functions λ, µ such
that
LA = 2µA + λtrA I ∀A ∈ Sym(V).

The reader is referred to the book of Gurtin (see the references) for the proof of this result
and for a deeper insight into isotropic functions.

3.8 The Kelvin formalism


As already mentioned, though fourth-rank tensors cannot be organized in and represented
by a matrix, nevertheless a matrix formalism for these operators exists. Such formalism
is due to Kelvin4 and it is strictly related to the theory of elasticity, i.e. it concerns the
Cauchy’s stress tensor σ, the strain tensor ε and the elasticity tensor E. The relation
between σ and ε is given by the celebrated (generalized) Hooke’s law:

σ = Eε.

Both σ, ε ∈ Sym(V) while E = E⊤ and it has also the minor symmetries, so E has just
21 independent components5 . In the Kelvin formalism, the six independent components
of σ and ε are organized into column vectors and renumbered as follows
   

 σ1 = σ11 
 
 ε1 = ε11 




 σ2 = σ22 






 ε2 = ε22 



 σ3 =√σ33   ε3 =√ε33 
{σ} = , {ε} = .

 σ4 = √2σ23 
 
 ε4 = √2ε23 

σ5 = √2σ31 ε5 = √2ε31

 
 
 


 
 
 


σ6 = 2σ12
 
ε6 = 2ε12

The elasticity tensor E is reduced to a 6 × 6 matrix [E] as a consequence of the minor


symmetries induced by the symmetry of σ and ε; this matrix is symmetric because

3
Actually, this is quite a famous result in classical elasticity, the Lamé’s equation, defining an isotropic
elastic material.
4
W. Thomson (Lord Kelvin): Elements of a mathematical theory of elasticity. Philos. Trans. R.
Soc., 146, 481-498, 1856. Later, Voigt (W. Voigt: Lehrbuch der Kristallphysik. B. G. Taubner, Leipzig,
1910) gave another, similar matrix formalism for tensors, more widely known than the Kelvin one, but
less effective.
5
Actually, the Kelvin formalism can also be extended without major difficulties to tensors that do not
possess all the symmetries.

57
E = E⊤ :
 √ √ √ 
E11 = E1111 E12 = E1122 E13 = E1133 E14 = √2E1123 E15 = √2E1131 E16 = √2E1112

 E12 = E1122 E22 = E2222 E23 = E2233 E24 = √2E2223 E25 = √2E2231 E26 = √2E2212 

E13 =√E1133 E23 =√E2233 E33 =√E3333 E34 = 2E3323 E35 = 2E3331 E36 = 2E3312
 
[E] =  .
 
 E14 = √2E1123 E24 = √2E2223 E34 = √2E3323 E44 = 2E2323 E45 = 2E2331 E46 = 2E2312 
 
 E15 = √2E1131 E25 = √2E2231 E35 = √2E3331 E45 = 2E2331 E55 = 2E3131 E56 = 2E3112 
E16 = 2E1112 E26 = 2E2212 E36 = 2E3312 E46 = 2E2312 E56 = 2E3112 E66 = 2E1212
In this way, the matrix product
{σ} = [E]{ε} (3.9)
is equivalent to the tensor form of the Hooke’s law and all the operations can be done
by the aid of classical matrix algebra6 , e.g. the computation of the inverse of E, the
compliance tensor.
An important operation is the expression of tensor U in Eq. (3.8) in the Kelvin formalism;
some tedious but straightforward passages give the result:
2 2 2
√ √ √
U11 U12 U13 √2U12 U13 √2U13 U11 √2U11 U12
 
2
U21 2
U22 2
U23

2 2 2
√2U22 U23 √2U23 U21 √2U21 U22 
U U U 2U U 2U U 2U31 U32
 
√ 31 √ 32 √ 33 32 33 33 31
[U ] = 
 
√2U21 U31 √2U22 U32 √2U23 U33 U23 U32 + U22 U33 U33 U21 + U31 U23 U31 U22 + U32 U21

 
√2U31 U11 √2U32 U12 √2U33 U13 U32 U13 + U33 U12 U31 U13 + U33 U11 U31 U12 + U32 U11
 
2U11 U21 2U12 U22 2U13 U23 U12 U23 + U13 U22 U11 U23 + U13 U21 U11 U22 + U12 U21
With some work, it can be checked that
[U ][U ]⊤ = [U ]⊤ [U ] = [I],
i.e. that [U ] is an orthogonal matrix in R6 . Of course,
[R] = [U ]⊤
is the matrix that in the Kelvin formalism represents the tensor R = U⊤ . The change of
basis for σ and ε are hence done through the relations
{σ ′ } = [U ]{σ}, {ε′ } = [U ]{ε},
which applied to Eq. (3.9) give
{σ} = [E]{ε} → [U ]⊤ {σ ′ } = [E][U ]⊤ {ε′ } → {σ ′ } = [U ][E][U ]⊤ {ε′ }
i.e. in the basis B ′
{σ ′ } = [E ′ ]{ε′ },
where
[E ′ ] = [U ][E][U ]⊤ = [R]⊤ [E][R]
is the matrix representing E in B ′ in the Kelvin formalism. Though it is possible to give
the expression of the components of [E ′ ], they are so long that they are omitted here.
6
Mehrabadi and Cowin have shown that the Kelvin formalism transforms second- and fourth-rank
tensors on R3 into vectors and second-rank tensors on R6 (M. M. Mehrabadi, S. C. Cowin: Eigentensors
of linear anisotropic elastic materials. Q. J. Mech. Appl. Math., 43, 15-41, 1990).

58
3.9 The polar formalism for plane tensors
The Cartesian representation of tensors makes use of quantities that are basis-dependent,
and the change of basis implies algebraic transformations rather complicate. The question
of representing tensors using other quantities than Cartesian components is hence of
importance. In particular, it should be interesting to represent a tensor making use of
only invariants of the tensor itself and of angles, the simplest geometrical way to determine
a direction.

In the case of plane tensors this has been done by Verchery7 who introduced the so-called
polar formalism. This is basically a mathematical technique to find the invariants of a
tensor of any rank. Here, we give just a short insight into the polar formalism of fourth-
rank tensors of the elastic type, i.e. having the minor and major symmetries, omitting
the proof of the results8 .

The Cartesian components of a plane fourth-rank elasticity-type tensor T in a frame


rotated through an angle θ can be expressed as

T1111 = T0 + 2T1 + R0 cos 4(Φ0 − θ) + 4R1 cos 2(Φ1 − θ),


T1112 = R0 sin 4(Φ0 − θ) + 2R1 sin 2(Φ1 − θ),
T1122 = −T0 + 2T1 − R0 cos 4(Φ0 − θ),
T1212 = T0 − R0 cos 4(Φ0 − θ),
T1222 = −R0 sin 4(Φ0 − θ) + 2R1 sin 2(Φ1 − θ),
T2222 = T0 + 2T1 + R0 cos 4(Φ0 − θ) − 4R1 cos 2(Φ1 − θ).

In the above equations, T0 , T1 , R0 , R1 are tensor invariants, with all of them non negative,
while Φ0 and Φ1 are angles whose difference, Φ0 −Φ1 , is also a tensor invariant, so fixing one
of the two polar angles corresponds to fixing a frame. In particular, the tensor invariants
have a direct physical meaning (e.g., for the elasticity tensor, they are linked to material
symmetries and to strain energy decomposition). We remark also that the change of frame
is extremely simple in the polar formalism: It is sufficient to subtract the angle θ formed
by the new frame from the two polar angles.

The Cartesian expression of the polar invariants can be found inverting the previous

7
G. Verchery: Les invariants des tenseurs d’ordre 4 du type de l’élasticité, Proc. Colloque EU-
ROMECH 115, 1979.
8
A detailed presentation of the method can be found in P. Vannucci: Anisotropic elasticity, Springer,
2018.

59
expressions:

1
T0 = (T1111 − 2T1122 + 4T1212 + T2222 ),
8
1
T1 = (T1111 + 2T1122 + T2222 ),
8
1p
R0 = (T1111 − 2T1122 − 4T1212 + T2222 )2 + 16(T1112 − T1222 )2 ,
8
1p
R1 = (T1111 − T2222 )2 + 4(T1112 + T1222 )2 ,
8
4(T1112 − T1222 )
tan 4Φ0 = ,
T1111 − 2T1122 − 4T1212 + T2222
2(T1112 + T1222 )
tan 2Φ1 = .
T1111 − T2222

3.10 Exercises
1. Prove Eq. (3.4).
2. Prove that
(AB)⊤ = B⊤ A⊤ .

3. Prove that
A ⊗ BL = A ⊗ L⊤ B.

4. Prove that
(A ⊠ B)(C ⊠ D) = AC ⊠ BD.

5. Prove Eq. (3.3) using the result of the previous exercise.


6. Prove that
(A ⊗ B)(C ⊠ D) = A ⊗ ((C⊤ ⊠ D⊤ )B).

7. Prove that
(A ⊠ B)(C ⊗ D) = ((A ⊠ B)C) ⊗ D.

8. Let p ∈ S and P = p ⊗ p, then prove that

P ⊠ P = P ⊗ P.

9. Prove that, ∀A ∈ Lin(V),


IA = AI = A.

10. Show that


(A ⊗ B) · (C ⊗ D) = A · C B · D.

60
11. Show that
I I
Ssph = ⊗ .
|I| |I|

12. Show that


dim(Sph(V)) = 1, dim(Dev(V)) = 5.

13. Show the following properties of Ssph and Ddev :

Ssph Ssph = Ssph ,


Ddev Ddev = Ddev ,
Ssph Ddev = Ddev Ssph = O.

14. Prove the results in Eqs. (3.5) and (3.6) using the components.
15. Show that
Ssph · Ssph = 1,
Ddev · Ddev = 5,
Ssph · Ddev = 0.

16. Make explicit the orthogonal conjugator SR of the tensor SR in Eq. (2.49).
17. Using the polar formalism, it can be proved that the material symmetries conditions
in plane elasticity are all condensed into the equation

R0 R1 sin 4(Φ0 − Φ1 ) = 0;

determine the different types of possibles elastic symmetries.

61
62
Chapter 4

Tensor analysis: curves

4.1 Curves of points, vectors and tensors


The scalar products in V, Lin(V) and Lin(V) allow us to define a norm, the Euclidean
norm, so they automatically endow these spaces with a metric, i.e. we are able to measure
and calculate a distance between two elements of such a space and in E. This allows us
to generalise the concepts of continuity and differentiability already known in R, whose
definition intrinsically makes use of a distance between real quantities.

Let πn = {pn ∈ E, n ∈ N} be a sequence of points in E. We say that πn converges to p ∈ E


if
lim d(pn − p) = 0.
n→∞

A similar definition can be given for sequences of vectors or tensors of any rank. Through
this definition of convergence we can now make the concepts of continuity and of curve
precise.

Let [a, b] be an interval of R; the function

p = p(t) : [a, b] → E

is continuous at t ∈ [a, b] if for each sequence {tn ∈ [a, b], n ∈ N} that converges to t, the
sequence πn defined by pn = p(tn ) ∀n ∈ N converges to p(t) ∈ E. The function p = p(t)
is a curve in E ⇐⇒ it is continuous ∀t ∈ [a, b]. In the same way we can define curves of
vectors and tensors:
v = v(t) : [a, b] → V,
L = L(t) : [a, b] → Lin(V),
L = L(t) : [a, b] → Lin(V).

Mathematically, a curve is a function that lets correspond to a real value t (the parameter)
in a given interval, an element of a space: E, V, Lin(V) or L(V).

63
4.2 Differention of curves
Let v = v(t) : [a, b] → V be a curve of vectors and g = g(t) : [a, b] → R a scalar function.
We say that v is of the order o with respect to g in t0 ⇐⇒

|v(t)|
lim = 0,
t→t0 |g(t)|

and we write
v(t) = o(g(t)) for t → t0 .
A similar definition can be given for a curve of tensors of any rank. We then say that the
curve v is differentiable in t0 ∈]a, b[ ⇐⇒ ∃v′ ∈ V such that

v(t) − v(t0 ) = (t − t0 )v′ (t0 ) + o(t − t0 ).

We call v′ (t0 ) the derivative 1 of v at t0 . Applying the definition of derivative to v′ we


define the second derivative v′′ of v and recursively all the derivatives of higher orders. We
say that v is of class Cn if it is continuous with its derivatives up to the order n; if n ≥ 1,
v is said to be smooth. A curve v(t) of class Cn is said to be regular if v′ ̸= o ∀t. Similar
definitions can be given for curves in E, Lin(V) and Lin(V), thus defining derivatives of
points and tensors. We remark that the derivative of a curve in E, defined as a difference
of points, is a curve in V (we say, in short, that the derivative of a point is a vector).
For what concerns tensors, the derivative of a tensor of rank r is a tensor of the same
rank.
Let u, v be curves in V, L, M curves in Lin(V), L, M curves in Lin(V) and α a scalar
function, all of them defined and at least of class C1 on [a, b]. The same definition of
derivative of a curve gives the following results, whose proof is let to the reader:

(u + v)′ = u′ + v′ ,
(αv)′ = α′ v + αv′ ,
(u · v)′ = u′ · v + u · v′ ,
(u × v)′ = u′ × v + u × v′ ,
(u ⊗ v)′ = u′ ⊗ v + u ⊗ v′ ,
(L + M)′ = L′ + M′ ,
(αL)′ = α′ L + αL′ ,
(Lv)′ = L′ v + Lv′ ,
(LM)′ = L′ M + LM′ ,
(L · M)′ = L′ · M + L · M′ ,

1 dv
The derivative is also written as , v,t or also as v̇, with the last symbol usually reserved, in
dt
physics, to the case where t is the time. For the sake of brevity, we omit to indicate the derivative of v
at t0 as v′ (t0 ), writing simply v′ .

64
(L ⊗ M)′ = L′ ⊗ M + L ⊗ M′ ,
(L ⊠ M)′ = L′ ⊠ M + L ⊠ M′ ,
(L + M)′ = L′ + M′ ,
(αL)′ = α′ L + αL′ ,
(LL)′ = L′ L + LL′ ,
(LM)′ = L′ M + LM′ ,
(L · M)′ = L′ · M + L · M′ .
We remark that the derivative of any kind of product is made according to the usual rule
of the derivative of a product of functions.
Let R = {o; B} be a reference frame of the euclidean space E, composed of an origin o
and a basis B = {e1 , e2 , e3 } of V, ei · ej = δij ∀i, j = 1, 2, 3 and let us consider a point
p(t) = (p1 (t), p2 (t), p3 (t)). If the three coordinates pi (t) are three continuous functions over
the interval [t1 , t2 ] ∈ R, then, by the definition given above, the mapping p(t) : [t1 , t2 ] → E
is a curve in E and the equation

 p1 = p1 (t)
p(t) = (p1 (t), p2 (t), p3 (t)) → p2 = p2 (t)
p3 = p3 (t)

is the parametric point equation of the curve: To each value of t ∈ [t1 , t2 ] it corresponds a
point of the curve in E, see Fig. 4.1.
Chapitre 1

e3

p(t)=(p1(t), p2(t), p3(t)) p(t)


r(t)

o
t1 t t2 R e2

e1

Figure 4.1: MappingFigure 1.7 of points.


of a curve
L(t)= Lij(t) ei ej, i, j= 1, 2, 3,
Theestvector
une courbe tensorielle.
function Souvent,
r(t) = p(t)−o en mécanique,
is the le paramètre
position vector in R;
of point pt est le temps de déroulement d’un
the equation
certain événement; nous verrons dans les chapitres suivants plusieurs exemples de courbes de

r1 = r1 (t)
points, vecteurs et tenseurs dont le paramètre est le temps, et leur
 signification mécanique.
r(t) = ri (t)ei = r1 (t)e1 + r2 (t)e2 + r3 (t)e3 → r2 = r2 (t)
Il faut aussi remarquer qu’une courbe peut avoir plusieurs représentations

r3 = r3 (t)paramétriques : en fait, si
t est le paramètre choisi pour représenter une courbe, par exemple une courbe de points, l’équation
is the parametric vector equation of the curve: To each value of t ∈ [t1 , t2 ] there corresponds
p (t ) p (t )
a vector of V that determines a point of the curve in E through the operation p(t) =
o+ décrit
r(t).la même courbe, étant lié à t par le changement de paramètre
(t ) .
65

1.22 DERIVEE D’UNE COURBE


Considérons une courbe de points p= p(t); on définit dérivée en t= to de la courbe p(t) par rapport au
paramètre t la limite
Similarly, if the components Lij (t) are continuous functions of a parameter t, the mapping
L(t) : [t1 , t2 ] → Lin(V) defined by
L(t) = Lij (t)ei ⊗ ej , i, j = 1, 2, 3,
is a curve of tensors. In the same way we can give a curve of fourth-rank tensors L(t) :
[t1 , t2 ] → Lin(V) by
L(t) = Lijkl (t)ei ⊗ ej ⊗ ek ⊗ el , i, j, k, l = 1, 2, 3.

It is worth noting that the choice of the parameter is not unique: The equation p = p[τ (t)]
still represents the same curve p = p(t), through the change of parameter τ = τ (t).
The definition given above for the derivative of a curve of points p = p(t) in t = t0 is
equivalent to the following one2 (probably more familiar to the reader)
Chapitre 1
dp(t) p(t0 + ε) − p(t0 )
= lim ,
dt ε→0 ε
dp dr dL
p, r, L. dp(t)
dt 4.2, where
represented in Fig. dt itdtis apparent that r′ (t) = is a vector.
dt
e3

p(to)
r’(to)
p(to+ )
r(to)
r(to+ )
o
e2

e1
Figure 1.8
Figure 4.2: Derivative of a curve.
Si on applique les opérations de limite aux composantes, on reconnaît immédiatement que
dp dr dL
An important case is pthat
i ( t ) e iof ri (t ) eiv(t)
, a vector , Lij (t ) enorm
whose i e j ,v(t) is constant ∀t:
dt dt dt
2 ′ ′ ′ ′ ′
(v d’une
c’est-à-dire que la dérivée ) = (v · v) a=comme
courbe v · v composantes
+ v · v = 2v les ·dérivées
v = 0 :des composantes de (4.1)
la
courbe donnée. Sur la base de cette considération, c’est facile de comprendre les formules
the derivative
suivantes, of such aaux
qui généralisent vector is les
courbes orthogonal to it ∀t.d’une
règles de dérivation Thefonction
contrary is variable
d’une also true, as is
réelle:
immediately apparent.
(u v) u v ;
Finally, using the( above
v) v and
rules v , assuming
(t) : R Rthat ; the reference frame R is independent of
t, we get easily that
(u v) u v u v ;
( u v ) u pv′ (t)u= vp;′ (t) ei ,
i
v[ (t )] v [ v(′t(t)
)] =(t ),v ′ (t)(t e
) : ,R R;
i i
( u v ) u v′ u v ′; (4.2)
L (t) = Lij (t) ei ⊗ ej ,
( L u) L u L′ u ;
L (t) = L′ (t) ei ⊗ ej ⊗ ek ⊗ el ,
( L M ) L M L M . ijkl
2
UnThis is true also for the derivatives of vector or tensor curves.
cas particulier, et important dans les applications, est celui d’un vecteur variable mais constant
en module ; dans ce cas la dérivée est toujours orthogonale au vecteur donné. En fait, soit v= v(t),
avec v (t ) v R . Cherchons la dérivée de la norme 66 au carré, qui est sans doute nulle parce que la
norme est constante par hypothèse :
(v 2 ) ( v v) v v v v 2v v 0,
donc les deux vecteurs sont orthogonaux ; on constate immédiatement que le contraire est vrai
i.e., the derivative of a curve of points, vectors or tensors is simply calculated differen-
tiating the coordinates of the components. Using this result, it is immediate to prove
that
(L⊤ )′ = L′⊤ ,
(L⊤ )′ = L′⊤ ,
while for any invertible tensor L it is (we state the following results without proof3 .)

(L−1 )′ = −L−1 L′ L−1 ,



(det L)′ = det L tr(L′ L−1 ) = det L L⊤ · L−1 = det L L′ · L−⊤ .

Let Q(t) : R → Orth(V)+ a differentiable function. We call spin tensor the tensor S(t)
defined as
S(t) := Q′ (t)Q⊤ (t).
Then, we have the following4

Theorem 24. (Characterization of the spin tensor). S(t) ∈ Skw(V) ∀t ∈ R.

Proof. As Q(t) ∈ Orth(V)+ ∀t, then


′ ′
QQ⊤ = I ⇒ (QQ⊤ )′ = Q′ Q⊤ + QQ⊤ = I′ = O ⇒ QQ⊤ = −Q′ Q⊤

so

S⊤ = (Q′ Q⊤ )⊤ = QQ⊤ = −Q′ Q⊤ = −S.

4.3 Integral of a curve of vectors and length of a


curve
We define integral of a curve of vectors r(t) between a and b ∈ [t1 , t2 ] the curve that is
obtained by integrating each component of the curve:
Z b Z b
r(t) dt = ri (t) dt ei .
a a

If the curve is regular, we can generalize the second fundamental theorem of the integral
calculus Z t
r(t) = r(a) + r′ (t∗ ) dt∗ .
a
Because
r(t) = p(t) − o, r′ (t) = (p(t) − o)′ = p′ (t),
3
The interested reader can find these proofs in the text by Gurtin, see the suggested texts.
4
The spin tensor and the following result are of importance in kinematics: If t is time and Q(t) ∈
OrthV + , then the axial vector of S(t) is ω(t), the angular velocity.

67
r (t ) r ( a ) r (t * )dt * .
a

Si on considère que
r (t ) p(t ) o,
r ( a ) p ( a ) o,
r (t ) ( p (t ) o ) p (t ),
we also get
l’équation ci-dessus peut être réécrite comme Z t
p(t)
t = p(a)
* *
+ p′ (t∗ ) dt∗ .
p (t ) p( a ) p (t )dt . a
a
The integral of a vector function is the generalization of the vector sum, see Fig. 4.3.
L’intégrale d’une fonction vectorielle est, d’une certaine façon, la généralisation de la somme
vectorielle, voir la figure 1.9.
t
e3 r (t * )dt*
a

p(a)
p(t)
r(a)
r(t)
o
e2

e1
Figure 1.9
Figure 4.3: Integral of a vector curve.
Une façon simple d’établir la position d’un point p(t) sur une courbe donnée, est celle de fixer un
point quelconque po sur la courbe, et de mesurer la longueur de l’arc de courbe compris entre
Letpor(t)
=p(to:) et
[a,p(t)
b] →; cette
E belongueur est appelée
a regular curve,abscisse curviligneofs(t),
σ a partition [a,etb]on
ofpeut
the démontrer
type a =que
t <t < 0 1
t t
... < tn = b, and s(t ) *
( p(t ) o) dt * *
r (t ) dt . *
to to
σmax = max |ti − ti−1 |.
i=1,...,n
La longueur totale d’une courbe r= r(t), avec t [a, b], sera
The length ℓσ of the polygonal line whose vertices are the points r(ti ) is hence:
b
l r (t) dt . n
a X
ℓσ = |r(ti ) − r(ti−1 )|.
De la formule de s(t), on a
i=1
2 2 2
We define length ofdsther curve r(t) dr2 dr3number
dr1 the (positive)
(t ) 0
dt dt dt dt
ℓ := sup ℓσ .
et donc s(t) est une fonction croissante avec t ; deσla formule précédente on tire la longueur d’un arc

- 22 -
Theorem 25. Let r(t) : [a, b] ⇒ E be a regular curve, then
Z b
ℓ= |r′ (t)|dt.
a

Proof. By the fundamental theorem of calculus,


Z ti
r(ti ) − r(ti−1 ) = r′ (t)dt,
ti−1

so that, using Minkowski’s inequality,


Z ti Z ti

|r(ti ) − r(ti−1 )| = r (t)dt ≤ |r′ (t)|dt,
ti−1 ti−1

68
whence Z b
ℓ≤ |r′ (t)|dt. (4.3)
a
Because r′ (t) is continuous on [a, b], ∀ε > 0 ∃δ > 0 such that |t−t| < δ ⇒ |r′ (t)−r′ (t)| < ε.
Let t ∈ [ti−1 , ti ] and σmax < δ, which is always possible by the choice of the partition σ;
again by the Minkowski’s inequality,
|r′ (t)| ≤ |r′ (t) − r′ (ti )| + |r′ (ti )| < ε + |r′ (ti )|,
whence
Z ti Z ti Z ti
′ ′
|r (t)|dt < |r (ti )|dt + ε(ti − ti−1 ) = r′ (ti )dt + ε(ti − ti−1 )
ti−1 ti−1 ti−1
Z ti Z ti
≤ r′ (t)dt + (r′ (ti ) − r′ (t))dt + ε(ti − ti−1 )
ti−1 ti−1

≤ |r(ti ) − r(ti−1 )| + 2ε(ti − ti−1 ).


Summing up over all the intervals [ti−1 , ti ] we get
Z b
|r′ (t)|dt ≤ ℓσ + 2ε(b − a) ≤ ℓ + 2ε(b − a),
a

and because ε is arbitrary, Z b


|r′ (t)|dt ≤ ℓ,
a
which by Eq. (4.3) implies the thesis.

Let t = f (τ ) : [c, d] → [a, b] be a bijective function that operates the change of parameter
from t to τ . If rt (t) : [a, b] → V is a parametric equation of a curve, rτ : [c, d] → V is a
re-parameterization of the same curve. We then have the following

Theorem 26. The length of a curve does not depend upon its parameterization.

Proof. Let rt (t) : [a, b] → E be a regular curve and t = f (τ ) : [c, d] → [a, b] be a change of
parameter; then dt = f ′ (τ )dτ and
Z b Z d Z d
′ ′ ′
ℓ= |rt (t)|dt = |rt (f (τ ))f (τ )|dτ = |r′τ (τ )|dτ.
a c c

A simple way to determine a point p(t) on a curve is to fix a point p0 on the curve and to
measure the length s(t) of the arc of curve between p0 = p(t = 0) and p(t). This length
s(t) is called curvilinear abscissa5 :
Z t Z t
′ ∗ ∗
s(t) = |r (t )|dt = |(p(t∗ ) − o)′ |dt∗ . (4.4)
0 0
5
The curvilinear abscissa is also called arc-length or natural parameter.

69
From Eq. (4.4) we get
ds
= |r′ (t)| > 0,
dt
so that s(t) is an increasing function of t and the length of an infinitesimal arc is
q
ds = dr12 + dr22 + dr32 .

For a plane curve y = f (x), we can always put t = x, which gives the parametric
equation
p(t) = (t, f (t)),
or in vector form
r(t) = t e1 + f (t) e2 ,
from which we obtain
ds p
= |r′ (t)| = |p′ (t)| = 1 + f ′2 (t), (4.5)
dt
which gives the length of a plane curve between t = x0 and t = x as a function of the
abscissa x: Z xp
s(x) = 1 + f ′2 (t)dt.
x0

4.4 The Frenet-Serret basis


We define the tangent vector τ (t) to a regular curve p = p(t) as the vector
p′ (t)
τ (t) := .
|p′ (t)|
By the definition of the derivative, this unit vector is always oriented as the increasing
values of t; hence, the straight line tangent to the curve in p0 = p(t0 ) has the equation
q(t̄) = p(t0 ) + t̄ τ (t0 ).
If the curvilinear abscissa s is chosen as parameter for the curve, through the change of
parameter s = s(t) we get
p′ (t) p′ [s(t)] 1 dp(s) ds(t) dp(s)
τ (t) = ′
= ′
= ′ = → τ (s) = p′ (s). (4.6)
|p (t)| |p [s(t)]| s (t) ds dt ds
So, if the parameter of the curve is s, the derivative of the curve is τ , i.e. it is automatically
a unit vector. The above equation, in addition, shows that the change of parameter does
not change the direction of the tangent, because it is only a scalar, the derivative of the
parameter’s change, that multiplies the vector. Nevertheless, generally speaking, a change
of parameter can change the orientation of the curve.
Because the norm of τ is constant, its derivative is a vector orthogonal to τ , see Eq. (4.1).
That is why we call principal normal vector to a curve the unit vector
τ ′ (t)
ν(t) := . (4.7)
|τ ′ (t)|

70
ν is defined only on the points of the curve where τ ′ ̸= o, which implies that ν is not
defined on the points of a straight line. This simply means that there is not, among the
infinite unit normal vectors to a straight line, a normal with special properties, a principal
one, linked to τ in a unique way.
Unlike τ , whose orientation changes with the choice of the parameter, ν is an intrinsic
local characteristic of the curve: It is not affected by the choice of the parameter. In
fact, by its same definition, ν does not depend upon the reference frame; then, because
the direction of τ is also independent upon the parameter’s choice, the only factor that
could affect ν is the orientation of the curve, which depends upon the parameter. But
a change in the orientation affects, in (4.7), both τ and the sign of the increment dt, so
that τ ′ (t) = dτ /dt does not change, nor does ν, which is hence an intrinsic property of
the curve.
The vector
β(t) := τ (t) × ν(t)
is called the binormal vector; by construction, it is orthogonal to τ and ν and it is a unit
vector. In addition, it is evident that
Chapitre 1
τ × ν · β = 1,
Les trois vecteurs introduits ci-dessus sont évidemment orthogonaux deux à deux et de norme
set {τ; ,enν,outre
so theunitaire β} forms a positively
c’est évident que oriented othonormal basis that can be defined at any

regular point of a curve with , τ ̸= o. Such a basis is called the Frenet-Serret local basis,
local in the sense that it changes with the position along the curve. The plane (τ , ν) is
et donc {τ, ν, β} est une base orthonormée directe, nommée trièdre de Frenet, définie en chaque
the osculating
point de la plane, the
courbe, et quiplane
change(ν, β)lathe
avec normal
position, figureplane
1.10 ; and
c’est the
pourplane (β,
cela que ceτtrièdre
) theest
rectifying
plane,appelé
see aussi
Fig. trièdre
4.4. local. plan τ−ν s’appelle
TheLeosculating planeplanis osculateur,
particularly ν−β plan normal
le planimportant: If etwe
le plan
consider a
β−τ plan rectifiant.
β

normal plane

ν$ β
rectifying plane β ν τ
ν$
τ$
osculating plane

Figure 1.10
Figure 4.4: The Frenet-Serret basis.
Le plan osculateur est particulièrement important : si on considère un plan qui passe par trois points
quelconques, non alignés, de la courbe, ce plan tend vers le plan osculateur lorsque ces trois points
plane sepassing through
rapprochent threetout
l’un à l’autre nonaligned
en restant surpoints of En
la courbe. theeffet,
curve, when
on peut thesequepoints
démontrer le plan become
closerosculateur en unstill
and closer, pointremaining
donné de la on
courbe
the est le planthe
curve, qui plane
se rapproche
tendsmieux
to theà laosculating
courbe au plane:
voisinage de ce
The osculating [Link]
plane Si la
a courbe
point est
of plane, le plan
a curve is osculateur
hence the est le plan qui
plane contient
that betterla courbe.
approaches the
curve On peutthe
near aussipoint.
démontrer
A que le vecteur
plane curvenormal ν est toujours
is entirely dirigé du
contained incoté
theduosculating
plan rectifiantplane,
dans which
lequel se trouve la courbe, voire, pour les courbes planes, ν est toujours dirigé vers la concavité de
is fixed.
la courbe.
The principal normal ν is always oriented toward the part of the space, with respect
to the1.25 COURBUREplane,
rectifying D’UNE Cwhere
OURBE the curve is; in particular, for a plane curve, ν is always

Il est important, dans plusieurs cas, de pouvoir évaluer de combien une courbe s’éloigne d’une ligne
droite au voisinage d’un point. Pour cela, on calcule 71 le vecteur tangent en deux points proches l’un
de l’autre, l’un à l’abscisse curviligne s, et l’autre à s+ε, et on mesure l’angle χ(s, ε) qu’ils forment,
voir la figure 1.11. On définit alors courbure de la courbe en s la limite

.
directed toward the concavity of the curve. To show that, it is sufficient to prove that the
vector p(t + ε) − p(t) forms with ν an angle ψ ≤ π/2, i.e. that (p(t + ε) − p(t)) · ν ≥ 0.
In fact,

1
p(t + ε) − p(t) = ε p′ (t) + ε2 p′′ (t) + o(ε2 ) ⇒
2
1 2 ′′
(p(t + ε) − p(t)) · ν = ε p (t) · ν + o(ε2 ),
2
but
p′′ (t) · ν = (τ ′ |p′ | + τ |p′ |′ ) · ν = (|τ ′ ||p′ |ν + τ |p′ |′ ) · ν = |τ ′ ||p′ |,
so that, to within infinitesimal quantities of order o(ε2 ), we obtain
Chapitre 1
1
(p(t + ε) − p(t)) · ν = ε2 |τ ′ ||p′ | ≥ 0.
2
1.25 COURBURE D’UNE COURBE
Il 4.5
est important, dans plusieursof
Curvature cas,adecurve
pouvoir évaluer de combien une courbe s’éloigne d’une ligne
droite au voisinage d’un point. Pour cela, on calcule le vecteur tangent en deux points proches l’un
deItl’autre, l’un à l’abscisse
is important, curviligne
in several s, ettol’autre
situations, à s+ how
evaluate , et on mesure
much l’angle
a curve (s, away
moves ) qu’ils forment,
from a
voir la figure 1.11. On définit alors courbure de la courbe en s la limite
straight line, in the neighborhood of a point. To do that, we calculate the angle formed
by the tangents at two close( s,points,
) determined by the curvilinear abscissae s and s + ε,
c ( s ) lim
and we measure the angle .
0 χ(s, ε) that they form, see Fig. 4.5.

e3 p(s)
p(s+ ) (s+ ) (s+ )
r(s) (s+ )
(s) (s) v(s+ )
r(s+ )
o
e2

e1
Figure 1.11
Figure 4.5: Curvature of a curve.
La courbure est donc un scalaire positif qui mesure la rapidité de variation de direction de la courbe
par unité de parcours sur la courbe même ; c’est évident que pour une ligne droite la courbure est
toujours
We thennulle.
define curvature of the curve in p = p(s) as the limit
Démontrons que la courbure est liée à la dérivée seconde de la courbe :
χ(s, ε)
c(s) = lim .
( s, ) ( s, ) ε
sin ε→0 2 ( s, )
c( s ) lim lim lim sin
0 0 0 2
The curvature is hence a non-negative scalar that measures the rapidity of variation in
v ( s, ) (s ) ( s)
the direction of the curvelim per unitlim length of the curve (that ( s ) is p (why
s ) . c(s) is defined as a
0 0
function of the curvilinear abscissa); by its same definition, the curvature is an intrinsic
property of the curve, i.e. independent of the parameter’s choice. For a straight line, the
D’ailleurs
curvature is everywhere identically null.
d [ s (t )] d ds d d 1 d
p (t ) ,
dt ds dt ds 72 ds p ( t ) dt

et donc
1 d (t )
c( s ) (s) ,
p (t ) dt p (t )
The curvature is linked to the second derivative of the curve; referring to Fig. 4.5, it
is
χ(s, ε) sin χ(s, ε) 2 χ(s, ε)
c(s) = lim = lim = lim sin
ε→0 ε ε→0 ε ε→0 ε 2
v(s, ε) τ (s + ε) − τ (s)
= lim = lim = |τ ′ (s)| = |p′′ (s)|.
ε→0 ε ε→0 ε

Another formula for the calculation of c(s) can be obtained if we consider that

dτ [s(t)] dτ ds dτ ′ dτ 1 dτ
= = |p (t)| → = ′ ,
dt ds dt ds ds |p (t)| dt

so that
1 dτ |τ ′ (t)|
c(s) = |τ ′ (s)| = = . (4.8)
|p′ (t)| dt |p′ (t)|

A better formula can be obtained using the complementary projector onto τ , i.e. the
tensor I − τ ⊗ τ , introduced in Exercise 2.2:

p′′ · p′
p′′ |p′ | − p′
dτ 1 dτ 1 d p′ (t) 1 |p′ |
= ′ = ′ ′
= ′
ds |p (t)| dt |p (t)| dt |p (t)| |p | |p′ |2
′′ ′′ ′′
p −τ p ·τ p
= ′ 2
= (I − τ ⊗ τ ) ′ 2 .
|p | |p |

Consequently,
dτ (s) 1
c(s) = = ′ 2 |(I − τ ⊗ τ )p′′ |.
ds |p |
Now, we use Eq. (2.36) with w = τ ; denoting by Wτ the axial tensor of τ , then

1
Wτ Wτ = − |Wτ |2 (I − τ ⊗ τ ),
2
whence
Wτ Wτ
I − τ ⊗ τ = −2 = −Wτ Wτ ,
|Wτ |2
because if τ = (τ1 , τ2 , τ3 ), then
   
0 −τ3 τ2 0 −τ3 τ2
|Wτ |2 = Wτ · Wτ =  τ3 0 −τ1  ·  τ3 0 −τ1 
−τ2 τ1 0 −τ2 τ1 0
= 2(τ12 + τ22 + τ32 ) = 2.

So, because Wτ ∈ Skw(V),


Wτ u = τ × u ∀u ∈ V.

73
Finally, using Eq. (2.29), the orthogonality property of cross product, Eq. (2.27) and Eq.
(2.31), we get

|(I − τ ⊗ τ )p′′ | = | − Wτ Wτ p′′ | = | − Wτ (τ × p′′ )| = | − τ × (τ × p′′ )|


|p′ × p′′ |
= |τ × (τ × p′′ )| = |τ × p′′ | = ,
|p′ |
so that, finally,
|p′ × p′′ |
c= . (4.9)
|p′ |3

Applying this last formula to a plane curve p(t) = (x(t), y(t)), we get

|x′ y ′′ − x′′ y ′ |
c= 3 ,
(x′2 + y ′2 ) 2

and if the curve is given in the form y = y(x), so that the parameter t = x, then we
obtain
|y ′′ |
c= 3 .
(1 + y ′2 ) 2

This last formula shows that if |y ′ | ≪ 1, then

c ≃ |y ′′ |.

This result is fundamental in the linearized (infinitesimal) theory of beams, plates and
shells.

4.6 The Frenet-Serret formulae


From Eq. (4.7) for t = s and Eq. (4.8), we get

=cν (4.10)
ds
which is the first Frenet-Serret Formula, giving the variation in τ per unit length of the
curve. Such a variation is a vector whose norm is the curvature and the direction is that
of ν. We remark that, because t = s, by Eq. (4.6) it is also

p′′ (s) = c(s)ν(s). (4.11)

Let us now consider the variation in β per unit length of the curve; because β is a unit
vector, we have

· β = 0,
ds
and
d(β · τ ) dβ dτ
β·τ =0 ⇒ = ·τ +β· = 0.
ds ds ds
74
Through Eq. (4.10) and because β · ν = 0 we get


· τ = −c β · ν = 0,
ds


so that is necessarily parallel to ν. We then set
ds


= ϑν,
ds

which is the second Frenet-Serret formula. The scalar ϑ(s) is called the torsion of the
curve in p = p(s). So, we see that the variation in β per unit length is a vector parallel
to ν and proportional to the torsion of the curve.
We can now find the variation in ν per unit length of the curve:

dν d(β × τ ) dβ dτ
= = ×τ +β× = ϑ ν × τ + c β × ν,
ds ds ds ds

so finally

= −c τ − ϑ β,
ds
which is the third Frenet-Serret formula: The variation in ν per unit length of the curve
is a vector of the rectifying plane.
The three formulae of Frenet-Serret (discovered independently by J. F. Frenet in 1847
and by J. A. Serret in 1851) can be condensed in the symbolic matrix product
 ′    
 τ  0 c 0  τ 
ν′ =  −c 0 −ϑ  ν .
 ′ 
β 0 ϑ 0 β
 

The matrix in the equation above is called the matrix of Cartan, and it is skew.

4.7 The torsion of a curve


We have introduced the torsion of a curve in the previous section, with the second formula
of Frenet-Serret. The torsion measures the deviation of a curve from flatness: If a curve is
planar, it belongs to the osculating plane and β, which is perpendicular to the osculating
pane, is hence a constant vector. So, its derivative is null, and by the Frenet-Serret second
formula, ϑ = 0.
Conversely, if ϑ = 0 everywhere, β is a constant vector and hence the osculating plane
does not change and the curve is planar. So we have that a curve is planar if and only if
the torsion is null ∀p(s).

75
Using the Frenet-Serret formulae in the expression of p′′′ (s), we get a formula for the
torsion:
dp ds
p′ (t) = |p′ |τ = = s′ τ ⇒ |p′ | = s′ →
ds dt

p′′ (t) = s′′ τ + s′ τ ′ = s′′ τ + s′2 = s′′ τ + c s′2 ν →
ds
p′′′ (t) = s′′′ τ + s′′ τ ′ + (c s′2 )′ ν + c s′2 ν ′
dτ dν
= s′′′ τ + s′′ s′ + (c s′2 )′ ν + c s′3
ds ds
= s′′′ τ + s′′ s′ cν + (c s′2 )′ ν − c s′3 (cτ + ϑβ)
= (s′′′ − c2 s′3 )τ + (s′′ s′ c + c′ s′2 + 2c s′ s′′ )ν − c s′3 ϑβ,

whence, through Eq. (4.9),

p′ × p′′ · p′′′ = s′ τ × (s′′ τ + c s′2 ν) · [(s′′′ − c2 s′3 )τ


+ (s′′ s′ c + c′ s′2 + 2c s′ s′′ )ν − c s′3 ϑβ]
|p′ × p′′ |2 ′ 6
= −c2 s′6 ϑ = −c2 |p′ |6 ϑ = − |p | ϑ,
|p′ |6

so that, finally,
p′ × p′′ · p′′′
ϑ=− ′ .
|p × p′′ |2
We remark that, while the curvature is linked to the second derivative of the curve, the
torsion is also a function of the third derivative.
Unlike curvature, which is intrinsically positive, the torsion can be negative. In fact, again
using the Frenet-Serret formulae,
1 1
p(s + ε) − p(s) = ε p′ + ε2 p′′ + ε3 p′′′ + o(ε3 )
2 6
1 2 1 3
= ετ + ε cν + ε (cν)′ + o(ε3 )
2 6
1 2 1
= ετ + ε cν + ε3 (c′ ν − c2 τ − c ϑβ) + o(ε3 )
2 6
1
⇒ (p(s + ε) − p(s)) · β = − ε3 c ϑ + o(ε3 ).
6

The above dot product determines whether the point p(s + ε) is located, with respect to
the osculating plane, on the side of β or on the opposite one, see Fig. 4.6: If following
the curve for increasing values of s, ε > 0, the point passes into the semi-space of β from
the opposite one, because 1/6 c ε3 > 0, it will be ϑ < 0, while in the opposite case it will
be ϑ > 0.
This result is intrinsic, i.e. it does not depend upon the choice of the parameter, hence of
the positive orientation of the curve; in fact, ν is intrinsic, but changing the orientation
of the curve, τ , and hence β, change in orientation.

76
La torsion, contrairement à la courbure qui est toujours positive, peut être négative. En particulier,
une fois établi un sens de parcours sur la courbe, c’est-à-dire une fois choisie une abscisse
curviligne, on peut démontrer que si, en suivant ce sens, la courbe sort du plan osculateur du côté de
β, alors la torsion est négative, elle est positive dans le cas contraire, voire figure 1.12. Ce résultat
est invariant: on peut démontrer que le signe de la torsion est une caractéristique intrinsèque de la
courbe, et ne dépend pas du paramétrage choisi.

cos α >0 : cos α <0 :


p(s+ε)
β β
$

ν τ$ s
α ν $τ
α
p(s) osculating p(s+ε)
plane p(s)
osculating
plane
s

Figure 4.6: Torsion of a curve.


Figure 1.12
4.8 Osculating
Ainsi que pour la courbure, onsphere
a une formuleand circle
de calcul de la torsion :

The osculating sphere6 to a curve at a point p is a sphere to which the curve tends to
- 28 -
adhere in the neighborhood of p. Mathematically, if qs is the center of the sphere relative
to the point p(s), then

|p(s + ε) − qs |2 = |p(s) − qs |2 + o(ε3 ).

Using this definition, discarding the terms of order o(ε3 ) and using the Frenet-Serret
formulae, we get:
1 1
|p(s + ε) − qs |2 = |p(s) − qs + εp′ + ε2 p′′ + ε3 p′′′ + o(ε3 )|2
2 6
1 2 1
= |p(s) − qs + ετ + ε c ν + ε3 (cν)′ + o(ε3 )|2
2 6
= |p(s) − qs | + 2ε(p(s) − qs ) · τ + ε2 + ε2 c(p(s) − qs ) · ν
2

1
+ ε3 (p(s) − qs ) · (c′ ν − c2 τ − c ϑβ) + o(ε3 ),
3
which gives
(p(s) − qs ) · τ = 0,
1
(p(s) − qs ) · ν = − = −ρ,
c
c′ ρ′
(p(s) − qs ) · β = − 2 = ,
cϑ ϑ
and finally
ρ′
β,qs = p + ρ ν − (4.12)
ϑ
so the center of the sphere belongs to the normal plane; the sphere is not defined for a
plane curve. The quantity ρ is the radius of curvature of the curve, defined as
1
ρ= .
c
6
The word osculating comes from the latin word osculo, which means to kiss; an osculating sphere or
circle or plane is a geometric object that is very close to the curve, as close as two lovers are in a kiss.

77
The radius of the osculating sphere is
s  ′ 2
ρ
ρs = |p − qs | = ρ2 + .
ϑ

The intersection between the osculating sphere and the osculating plane at a same point
p is the osculating circle. This circle has the property of sharing the same tangent in p
with the curve and its radius is the radius of curvature, ρ. From Eq. (4.12) we get the
position of the osculating circle center q:

q = p + ρ ν. (4.13)

An example can be seen in Fig. 4.7, where the osculating plane, circle and sphere are
shown for a point p of a conical helix.

𝜷
osculating shpere

𝜌s qs
osculating circle

p 𝜌 q 𝝂
𝝉
osculating plane

Figure 4.7: Osculating plane, circle and sphere for a point p of a conical helix.

The osculating circle is a diametral circle of the osculating sphere only when q = qs , so if
and only if
ρ′ c′
= − 2 = 0,
ϑ cϑ
i.e. when the curvature is constant.

4.9 Evolute, involute and envelopes of plane curves


For any plane curve γ(s), the center of the osculating circle q describes a curve δ(σ) that
is called the evolute of γ(s) (s and σ are curvilinear abscissae). A point q of the evolute
is then given by Eq. (4.13). We call involute of a curve γ(s) a curve µ(σ) whose evolute
is γ(s). We call envelope of a family of plane curves φ(s, κ), κ ∈ R being a parameter, a
curve that is tangent, in each of its points, to the curve of φ(s, κ) passing through that
point.

78
Let us consider the evolute δ(σ) of a curve γ(s); the tangent to δ(σ) is the vector, cf. Eq.
(4.13),
dq dq ds
τδ = = .
dσ ds dσ
But, cf. again Eq. (4.13) and the Frenet-Serret formulae,

dq dp dρ dν dρ dρ
= + ν +ρ = τ + ν − ρ c τ = ν,
ds ds ds ds ds ds
so
dq dρ ds
τδ = = ν.
dσ ds dσ
Because
dq
= |τ δ | = |ν| = 1,

then
dρ ds dρ dσ
=1 ⇒ =
ds dσ ds ds
and
τ δ = ν.
The evolute, δ(σ), of γ(s) is hence the envelope of its principal normals ν(s).
This result helps us in finding the equation of the involute µ(σ) of a curve γ(s); let
p = p(s) be a point of γ(s); then, if b ∈ µ(σ), it must be that

(b − p) · ν = 0

where ν is the principal normal to γ(s) in p, because γ(s) is the evolute of µ(σ), which
implies, for the last result, that τ = ν µ , with τ the tangent to γ(s) in p and ν µ the
principal normal to µ(σ) in b, see Fig. 4.8.
Therefore,
b(s) − p(s) = f (s)τ (s) → b(s) = p(s) + f (s)τ (s),
with f = f (s) a scalar function of s; to remark that b = b(s), i.e. the arc-length s of γ(s)
is the parameter also for µ(s), but in general σ ̸= s. Upon differentiation, we get

b′ (s) = (1 + f ′ (s))τ (s) + f (s)c(s)ν(s).

Then, because b′ (s) = |b′ (s)|τ µ is orthogonal to ν µ = τ , it is parallel to ν, so it must be


that
1 + f ′ (s) = 0 ⇒ f (s) = a − s, a ∈ R.
Finally, the equation of the involute µ(s) to γ(s) is

b(s) = p(s) + (a − s)τ (s),

and we remark that the involute is not unique.

79
Figure 4.8: Evolute, δ, and involutes for a = 0, denoted by µ, and a = 1, dashed, of a
catenary γ.

4.10 The theorem of Bonnet


The curvature, c(s), and the torsion, ϑ(s), are the only differential parameters that com-
pletely describe a curve. In other words, given two functions c(s) and ϑ(s), then a curve
exists with such a curvature and torsion (we remark that there are no conditions bounding
these parameters). This is proved in the following

Theorem 27. (Bonnet’s theorem). Given two scalar functions c(s) ∈C1 and ϑ(s) ∈C0 ,
there always exists a unique curve γ ∈C3 whose curvilinear abscissa is s, curvature c(s)
and torsion ϑ(s).

Proof. Let  
τ
e= ν 
β
be the column vector whose elements are the vectors of the Frenet-Serret basis. Then
de(s)
= C(s)e(s) (4.14)
ds
with  
0 c(s) 0
C(s) =  −c(s) 0 −ϑ(s) 
0 ϑ(s) 0

80
the matrix of Cartan. Adding the initial condition
 
e1
e(0) =  e2 
e3

we have a Cauchy problem for the basis e(0). As known, such a problem admits a unique
solution, i.e. we can associate to c(s) and ϑ(s) a family of bases e(s) (that are orthonormal
because if one of them were not so, the Cartan’s matrix should not be skew). Call τ (s)
the first vector of the basis e(s) and define the function
Z s
p(s) := p0 + τ (s∗ )ds∗ ;
0

p(s) is the curve looked for (it depends upon an arbitrary point p0 , i.e. upon an inessential
rigid displacement). In fact, because |τ | = 1, s is the curvilinear abscissa of the curve.
Then, it is sufficient to write the Frenet-Serret equations identifying them with the system
(4.14).

4.11 Canonic equations of a curve


We call the canonic equations of a curve at a point p0 the equations of the curve referred
to the Frenet-Serret basis in p0 . For this purpose, we expand the curve in a Taylor series
of initial point p0 :
1 1
p(s) = p0 + s p′ (0) + s2 p′′ (0) + s3 p′′′ (0) + o(s3 ).
2 6
In the Frenet-Serret basis,
dcν
p′ (0) = τ (0), p′′ (0) = c(0)ν(0), p′′′ (0) = = c′ (0)ν(0)−c2 (0)τ (0)−c(0)ϑ(0)β(0),
ds s=0

so
1 1
p(s) = p0 + s τ (0) + s2 c(0)ν(0) + s3 (−c2 (0)τ (0) + c′ (0)ν(0) − c(0)ϑ(0)β(0)) + o(s3 ).
2 6
The coordinates of a point p(s) close to p0 , in the basis {τ (0), ν(0), β(0)}, are hence
1
p1 (s) = s − c2 (0)s3 + o(s3 ),
6
1 1
p2 (s) = c(0)s2 + c′ (0)s3 + o(s3 ),
2 6
1
p3 (s) = − c(0)ϑ(0)s3 + o(s3 ).
6
The projections of the curve onto the planes of the Frenet-Serret basis hence have, close
to p0 (i.e. retaining the first non null term in the expressions above), the following
equations:

81
• On the osculating plane (
p1 (s) = s,
1
p2 (s) = c(0)s2 ,
2
or, eliminating s,
1
p2 = c(0)p21 ,
2
which is the equation of a parabola.
• On the rectifying plane (
p1 (s) = s,
1
p3 (s) = − c(0)ϑ(0)s3 ,
6
or, eliminating s,
1
p3 = − c(0)ϑ(0)p31 ,
6
which is the equation of a cubic parabola.
• On the normal plane 
 p2 (s) = 1 c(0)s2 ,

2
1
 p3 (s) = − c(0)ϑ(0)s3 ,

6
or, eliminating s,
2 ϑ2 (0) 3
p23 = p,
9 c(0) 2
which is the equation of a semicubic parabola, with a cusp at the origin, hence a
singular point, though the curve p(s) is regular.

4.12 Exercises
1. Using the same definition of derivative of a curve, prove the relations in Sect. (4.2).
2. Prove the relations in Eq. (4.2).
3. The curve whose polar equation is

r = a θ, a ∈ R,

is an Archimedes’ spiral, Fig. 4.9 a). Find its curvature c(θ) and its length ℓ(θ),
and prove that any straight line passing through the origin is divided by the spiral
in segments of constant length 2π a (that is why the Archimede’s spiral is used to
record disks).
4. The curve whose polar equation is

r = a ebθ , a, b ∈ R,

82
y y

x x

a) b)

Figure 4.9: The Archimedes’, a), and logarithmic, b), spirals.

is the logarithmic spiral. Prove that the origin is an asymptotic point of the curve,
find its curvature c(θ) and its length ℓ(θ), and show that the length of the segments
in which a straight line by the origin is divided by two consecutive intersections with
the spiral varies as a geometrical progression. Then, prove its equiangular property:
The angle α between p(θ) − o and τ (θ) is constant. Finally, show that the evolute
of the logarithmic spiral is a logarithmic spiral itself (an hence that its involute is
still a logarithmic spiral, that’s why Jc. Bernoulli coined for this curve the Latin
sentence eadem mutata resurgo.)
5. The curve whose parametric equation is
p(θ) = a(cos θ + θ sin θ)e1 + a(sin θ − θ cos θ)e2
with θ the angle formed by p(θ) − o with the axis x1 , is the involute of the circle,
Fig. 4.10. Find its curvature c(θ) and its length ℓ(θ), and prove that its evolute is
exactly the circle of center o and radius a (that is why the involute of the circle is
used to profile gears).
y

𝞶 p

q
𝜃
x
a

Figure 4.10: The involute of the circle and its evolute, the circle.

6. The curve whose parametric equation is


p(θ) = a cos ωθe1 + a sin ωθe2 + bωθe3

83
is a circular helix, i.e. a helix that winds on a circular cylinder of radius a, Fig.
4.11. Show that the angle φ formed by the helix and any generatrix of the cylinder
is constant (a property that defines a helix in the general case). Then, find its length
ℓ(θ), its curvature c(θ), torsion ϑ(θ), and the pitch d, i.e., the distance between two
successive intersections of the helix with a generatrix of the cylinder. Prove then the
Bertrand’s theorem: A curve is a cylindrical helix if and only if the ratio c/ϑ = const.
Finally, prove that for the above circular helix there are two constants A and B such
that
p′ × p′′ = Au(θ) + Be3 ,
with
u = sin ωθe1 − cos ωθe2 ;
then, find A and B.

d
𝜑
p 𝞽
o y

Figure 4.11: The circular helix.

7. Find the equation of the cycloid, i.e. of the curve that is the trace of a point of
a circle of radius r rolling without slipping on a horizontal axis, see Fig. 4.12.
Calculate the length of the cycloid for a complete round of the circle, determine its
curvature, and show that the evolute of the cycloid is the cycloid itself (Huygens,
1659).
8. The planar curve whose parametric equation is

p(t) = te1 + cosh te2

is the catenary (Jc. Bernoulli, 1690; Jn. Bernoulli, Leibniz, Huygens, 1691). It is
the equilibrium curve of a heavy, perfectly flexible, and inextensible cable. Calculate
the curvature of the catenary and the equation of its evolute and of its involutes
(see Fig. 4.8).
9. The planar curve whose parametric equation is
 
t
p(t) = cos t + ln tan e1 + sin te2
2

84
y

p
𝜃
𝞶
o x
q

Figure 4.12: The cycloid and its evolute.

is the tractrix (Perrault, 1670; Newton, 1676; Huygens, 1693). This is the curve
along which an object moves, under the influence of friction, when pulled on a
horizontal plane by a line segment attached to a tractor that moves at a right angle
to the initial line between the object and the puller at an infinitesimal speed, see
Fig. 4.13. Show that the length of the tangent to the tractrix between the points
on the tractrix itself and the axis x is constant ∀t, calculate the length of the curve
between t1 and t2 , calculate the curvature of the tractrix, and finally show that its
evolute is the catenary.

y q

𝞶
p
𝞽
o g x

Figure 4.13: The tractrix and its evolute.

10. For the curve whose cylindrical equation is


(
r = 1,
z = sin θ

find the highest curvature and determine whether or not it is planar.


11. Let p = p(t) be the path of a moving particle of masse m, with t being the time.
Define the velocity and the acceleration of p as, respectively, the first and the second
derivative of p with respect to t. Decompose these two vectors in the Frenet-Serret
basis and interpret physically the result. Recalling the second Newton’s principle of
mechanics, what about the forces on p?

85
86
Chapter 5

Tensor analysis: fields

5.1 Scalar, vector and tensor fields


Let Ω ⊂ E and f : Ω → V. We say that f is continuous at p ∈ Ω ⇐⇒ ∀ sequence
πn = {pn ∈ Ω, n ∈ N} that converges to p ∈ E, the sequence {vn = f (pn ), n ∈ N}
converges to f (p) in V. The function f (p) : Ω → V is a vector field on Ω if it is continuous
at each p ∈ Ω. In the same way we can define a scalar field φ(p) : Ω → R and a tensor
field, L(p) : Ω → Lin(V).
A deformation is any continuous and bijective function f (p) : Ω → E, i.e. any transforma-
tion of a region Ω ⊂ E into another region of E; bijectivity imposes that to any point p ∈ Ω
corresponds one and only one point in the transformed region, and vice-versa, which is
the mathematical condition expressing the physical constraint of mass conservation.
Finally, the basic difference between fields/deformations and curves, is that a field or a
deformation is defined over a subset of E, not of R. In practice, this implies that the
components of the field/deformation are functions of three variables, the coordinates xi
of a point p ∈ Ω.

5.2 Differentiation of fields, differential operators


Let ψ(p) be a scalar or vector or tensor field or also a deformation; we define directional
derivative of ψ(p) in the direction of e ∈ S the limit

dψ(p) ψ(p + αe) − ψ(p)


:= lim , α ∈ R.
de α→0 α

The directional derivative measures the rate of variation of ψ(p) in the direction of e. In
the particular case of e = ei , i = 1, 2, 3, i.e. of the directions of the basis {e1 , e2 , e3 } of
V, then
dψ(p) ψ(p + αei ) − ψ(p)
= lim
dei α→0 α

87
is the partial derivative of ψ with respect to xi ; e.g., if i = 1, then
dψ(p) ψ(x1 + α, x2 , x3 ) − ψ(x1 , x2 , x3 )
= lim .
de1 α→0 α
∂ψ
The partial derivative with respect to xi is usually indicated as or also as ψ,i .
∂xi
Let v(p) : Ω → V; we say that v is differentiable in p0 ∈ Ω ⇐⇒ ∃ gradv ∈ Lin(V) such
that
v(p0 + u) = v(p) + gradv(p) u + o(u)
when u → o. If v is differentiable ∀p ∈ Ω, gradv defines a tensor field on Ω called the
gradient of v. It is also possible to define higher order differential operators, using higher
order tensors, but this will not be done here. If v is continuous with gradv ∀p ∈ Ω, then
v is of class C1 (smooth).
Let v be a vector field of class C1 on Ω. Then, the divergence of v is the scalar field
defined by
divv := tr(gradv),
while curlv is the unique vector field that satisfies the relation

(gradv − gradv⊤ )u = (curlv) × u ∀u ∈ V.

The divergence of a tensor field L is the unique vector field divL that satisfies

(divL) · u = div(L⊤ u) ∀u = const. ∈ V.

Let φ(p) : Ω → R be a scalar field over Ω. Similar to the case of vector fields, we say that
φ is differentiable at p0 ∈ Ω ⇐⇒ ∃ gradφ ∈ V such that

φ(p + u) = φ(p) + gradφ(p) · u + o(u)

when u → o. If φ is differentiable ∀p ∈ Ω, gradφ defines a vector field on Ω called the


gradient of φ. If gradφ is differentiable, its gradient is the tensor gradII φ called second
gradient or Hessian. It is immediate to show that under continuity assumption,

gradII φ = (gradII φ)⊤ .

A level set of a scalar field φ(p) is the set SL such that

φ(p) = const. ∀p ∈ SL .

Considering hence two points p and p + u of the same SL , then by the definition of
differentiability of φ(p) itself, we see that gradφ is a vector that is orthogonal to SL at p.
The curves of E that are tangent to gradφ ∀p ∈ Ω are the streamlines of φ; they have
the property to be orthogonal to any SL of φ ∀p ∈ Ω.
gradφ allows to calculate the directional derivative of φ along any direction n ∈ S as

= gradφ · n.
dn
88
The highest variation of φ is hence in the direction of gradφ, and |gradφ| is the value of
this variation; we remark also that gradφ is a vector directed as the increasing values of
φ.
Similarly, for a vector field v the directional derivative along any direction n ∈ S can be
computed as
dv
= gradv n.
dn

Let ψ be a scalar of vector field of class C2 at least. Then, the laplacian ∆ψ of ψ is


defined by
∆ψ := div(gradψ).
By the linearity of the trace, and hence of the divergence, we see easily that the laplacian of
a vector field is the vector field whose components are the laplacian of each corresponding
component of the field. A field is said to be harmonic on Ω if its laplacian is null ∀p ∈
Ω.
The definitions given above for differentiable field, gradient and class C1 can be repeated
verbatim for a deformation f (p) : Ω → E.

5.3 Properties of the differential operators


The differential operators, gradient, divergence, curl and laplacian, have some interesting
properties, that are useful for calculations; they are introduced in this section.

Theorem 28. (Gradient of products). Let φ, ψ be scalar and u, v, w be vector fields,


with all of them differentiable. Then:

i) grad(φψ) = φ gradψ + ψ gradφ,


ii) grad(φv) = φ gradv + v ⊗ gradφ, (5.1)
iii) grad(v · w) = (gradw)⊤ v + (gradv)⊤ w.

Proof. The proof is based upon the definition of gradient itself1 :


i) (φψ)(p + u) = φψ + grad(φψ) · u + o(u),
but also
(φψ)(p + u) = φ(p + u)ψ(p + u) = (φ + gradφ · u + o(u))(ψ + gradψ · u + o(u))
= φψ + φ gradψ · u + ψ gradφ · u + o(u)
= φψ + (φ gradψ + ψ gradφ) · u + o(u),

so by comparison
grad(φψ) = φgradψ + ψgradφ.
1
For the sake of brevity, we omit to indicate the point p, e.g. we simply write φ for φ(p), gradφ for
gradφ(p) etc.

89
ii) in the same way
(φv)(p + u) = φv + grad(φv)u + o(u) = φv + grad(φv)u + o(u),
but also
(φv)(p + u) = φ(p + u)v(p + u) = (φ + gradφ · u + o(u))(v + gradv u + o(u))
= φv + φgradv u + gradφ · u v + o(u)
= φv + (φ gradv + v ⊗ gradv)u + o(u),
so comparing the two results we get
grad(φv) = φ gradv + v ⊗ gradv.
iii) in the same way
(v · w)(p + u) = v · w + grad(v · w) · u + o(u),
but also
(v · w)(p + u) = v(p + u) · w(p + u) = (v + gradv u + o(u)) · (w + gradw u + o(u))
= v · w + v · (gradw u) + (gradv u) · w + o(u)
= v · w + ((gradw)⊤ v + (gradv)⊤ w) · u + o(u),
whence, by comparison of the two results,
grad(v · w) = (gradw)⊤ v + (gradv)⊤ w.

Another important result2 , relating the gradient and the curl of a vector field, is the
following theorem.

Theorem 29. If v is a differentiable vector field, then


1
(gradv)v = (curlv) × v + gradv2 .
2
Proof.
(curlv) × v = (gradv − (gradv)⊤ )v = (gradv)v − (gradv)⊤ v
1
= (gradv)v − ((gradv)⊤ v + (gradv)⊤ v),
2
and by property iii) of the previous theorem,
(gradv)⊤ v + (gradv)⊤ v = grad(v · v) = gradv2 ,
so that
1
(curlv) × v = (gradv)v − gradv2 ,
2
whence we obtain the thesis.
2
This result is fundamental to fluid mechanics, as it allows us to get an interesting form of the Navier-
Stokes equations.

90
The proof of the following properties of the gradient are left to the reader as an exer-
cise:
grad(v · w) = (gradw)v + (gradv)w + v × curlw + w × curlv,
grad(u · v w) = (u · v)gradw + (w ⊗ u)gradv + (w ⊗ v)gradu, (5.2)
gradv · gradv⊤ = div((gradv)v − (divv)v) + (divv)2 .

Theorem 30. (Divergence of products). Let φ, u, v, w, L be differentiable scalar,


vector or tensor fields. Then:
i) div(φv) = φdivv + v · gradφ,
ii) div(v ⊗ w) = vdivw + (gradv)w,
iii) div(φL) = φdivL + Lgradφ,
iv) div(L⊤ v) = L · gradv + v · divL,
v) div(v × w) = w · curlv − v · curlw.

Proof. i) Using the definition of divergence and property ii) of Theorem 28, we get
div(φv) = tr(grad(φv)) = tr(φ gradv + v ⊗ gradφ)
= φ tr(gradv) + tr(v ⊗ gradφ) = φ divv + v · gradφ.
ii) By the definition of divergence of a tensor, ∀a = const. ∈ V, and using the previous
property along with property iii) of Theorem 28:
div(v ⊗ w) · a = div((v ⊗ w)⊤ a) = div(w ⊗ v a) = div(v · a w)
= v · a divw + w · grad(a · v)
= divw v · a + w · (gradv)⊤ a + w · (grada)⊤ v
= (vdivw + gradv w) · a.
iii) By the definition of divergence of a tensor, ∀a = const. ∈ V, and using the property
i) along with the iii) of Theorem 28:
div(φL) · a = div((φL)⊤ a) = div(φL⊤ a) = φdiv(L⊤ a) + L⊤ a · gradφ
= φdivL · a + a · Lgradφ = (φdivL + Lgradφ) · a.
iv) By the definition of divergence of a tensor and using the previous property:
div(L⊤ v) = div(L⊤ vj ej ) = div((vj L⊤ )ej ) = div(vj L⊤ )⊤ · ej
= div(vj L) · ej = vj divL · ej + L gradvj · ej
= v · divL + (Lpq ep ⊗ eq (gradvj )m em ) · ej
= v · divL + Lpq vj,m δqm δjp = v · divL + L · gradv.

v) This property can be proved making use of the expression of the cross product with
the Ricci’s alternator, given in Eq. (2.26):
curlv = ϵijk vk,j ei

91
and
v × w = ϵijk vj wk ei ,
whence

div(v × w) = div(ϵijk vj wk ei ) = ϵijk (vj wk ),i = ϵijk vj,i wk + ϵijk wk,i vj .

Moreover,

w · curlv = wm em · ϵpqr vr,q ep = ϵpqr vr,q wm δpm = ϵpqr vr,q wp = ϵqrp vr,q wp

and
v · curlw = vm em · ϵpqr wr,q ep = ϵpqr wr,q vm δpm = ϵpqr wr,q vp = −ϵqpr wr,q vp ,
so finally, comparing the last three results (all the subscripts are dummy indexes, so their
denomination is inessential),

div(v × w) = w · curlv − v · curlw.

The divergence has also the following properties

div(gradv⊤ ) = grad(divv),
div((gradv)v) = gradv · gradv⊤ + v · grad(divv), (5.3)
div(φLv) = φL⊤ · gradv + φv · divL⊤ + Lv · gradφ,

whose proof is a good exercise for the reader.


The relations of gradient and divergence with the curl are given by the following theo-
rem.

Theorem 31. Let φ and v be scalar and vector fields of class C2 ; then

i) div(curlv) = 0,
ii) curl(gradφ) = o.

Proof. i) Using again the Ricci’s alternator to represent the cross product,

div(curlv) = div(ϵijk vk,j ei ) = ϵijk vk,j divei + ϵijk vk,ji = ϵijk vk,ji
= v3,21 + v1,32 + v2,13 − v2,31 − v3,12 − v1,23 = 0.

ii) In a similar manner,

curl(gradφ) = ϵijk φ,kj ei = φ,32 + φ,13 + φ,21 − φ,23 − φ,31 − φ,12 = 0.

92
The following theorem gives an interesting relation between the curl of a vector and the
divergence of its axial tensor.

Theorem 32. (Curl of an axial vector). Let w be a differentiable vector field and W
its axial tensor field. Then,
curlw = −divW.

Proof. Using properties iv) and v) of Theorem 30 and because W = −W⊤ , ∀a = const. ∈
V we get

div(w × a) = a · curlw − w · curla = a · curlw,


div(Wa) = div(−W⊤ a) = −W · grada − a · divW = −a · divW.

Now, because ∀a, w × a = Wa ⇒ div(w × a) = div(Wa), we get the thesis.

The way the curl of a curl3 is computed is given by the following theorem.

Theorem 33. (Curl of a curl). Let vbe a vector field of class ≥C2 . Then,

curl(curlv) = grad(divv) − ∆v.

Proof. Using properties iv) and v) of Theorem 30, along with the first of Eq. (5.3),
∀a = const. ∈ V we get

div((curlv) × a) = a · curl(curlv) − curlv · curla = a · curl(curlv),

and by the definition of curl and laplacian

div((curlv) × a) = div((gradv − (gradv)⊤ )a) = div(gradv a) − div((gradv)⊤ a)


= (gradv)⊤ · grada + a · div(gradv)⊤ − gradv · grada − a · div(gradv)
= a · (div(gradv)⊤ − div(gradv)) = a · (grad(divv) − ∆v),

whence, by comparison,
curl(curlv) = grad(divv) − ∆v.

The proof of the following properties of the curl are can be done using the above results
and it is a good exercise:

curl(φv) = φcurlv + gradφ × v,


(5.4)
curl(v × w) = (gradv)w − (gradw)v + vdivw − wdivv.

Finally, we have a theorem also for the laplacian of a product.


3
This relation is useful in fluid mechanics, for writing the vorticity equation.

93
Theorem 34. (Laplacian of products). Let φ, ψ, u, v be scalar and vector fields of
class ≥C2 . Then:
i) ∆(φψ) = 2gradφ · gradψ + φ∆ψ + ψ∆φ,
ii) ∆(v · w) = 2gradv · gradw + v · ∆w + w · ∆v.

Proof. i) Using properties i) of Theorems 28 and 30, we get

∆(φψ) = div(grad(φψ)) = div(φ gradψ + ψ gradφ)


= div(φ gradψ) + div(ψ gradφ)
= φ div(gradψ) + gradψ · gradφ + ψ div(gradφ) + gradφ · gradψ
= 2gradφ · gradψ + φ ∆ψ + ψ ∆φ.

ii) Using properties iii) of Theorem 28 and iv) of Theorem 30, we obtain

∆(v · w) = div(grad(v · w)) = div((gradw)⊤ v + (gradv)⊤ w)


= div((gradw)⊤ v) + div((gradv)⊤ w)
= gradw · gradv + v · div(gradw) + gradv · gradw + w · div(gradv)
= 2gradv · gradw + v · ∆w + w · ∆v.

5.4 Theorems on fields


We recall here some classical theorems on fields and operators.

Theorem 35. (Harmonic fields). If v(p) is a vector field of class ≥ C2 such that

divv = 0, curlv = o,

then v is harmonic: ∆v = o.

Proof. By the definition of curl,

curlv = o ⇒ gradv − (gradv)⊤ = o ⇒ div(gradv − (gradv)⊤ ) = o,

and through Eq. (5.3)1 , the definition of laplacian and because by hypothesis divv = 0,
we have
div(gradv − (gradv)⊤ ) = ∆v − grad(divv) = ∆v.

We state now without proof a lemma4 that, basically, allows us to transform a volume
integral on a domain Ω to a surface integral on the boundary surface ∂Ω.
4
In the following, ∂Ω indicates the boundary of Ω.

94
Theorem 36. (Divergence lemma). Let v(p) be a vector field of class ≥ C1 on a
regular region Ω ⊂ E. Then,
Z Z
v ⊗ n dA = gradv dV.
∂Ω Ω

This lemma is fundamental for proving the three forms of the Gauss theorem, which is of
the paramount importance in many fields of mathematical physics.

Theorem 37. (Divergence or Gauss theorem). Let φ, v, L be, respectively, a scalar,


vector and tensor field of class ≥ C1 on a regular region Ω ⊂ E . Then:
Z Z
i) φn dA = gradφ dV,
Z ∂Ω ΩZ
ii) v · n dA = divv dV,
Z ∂Ω Z Ω
iii) Ln dA = divL dV.
∂Ω Ω

Proof. i) ∀a = const. ∈ V, by the lemma of divergence,


Z Z Z
grad(φa)dV = φa ⊗ n dA = a ⊗ φn dA,
Ω ∂Ω ∂Ω

but also, by ii) of Theorem 28,


Z Z Z
grad(φa)dV = (φ grada + a ⊗ gradφ)dV = a ⊗ gradφ dV,
Ω Ω Ω

whence, by comparison, Z Z
φn dA = gradφ dV.
∂Ω Ω
ii) Again by the divergence lemma,
Z Z Z Z
tr gradv dV = tr v ⊗ n dA = tr(v ⊗ n)dA = v · n dA,
Ω ∂Ω ∂Ω ∂Ω

but also Z Z Z
tr gradv dV = tr(gradv)dV = divv dV,
Ω Ω Ω
so, by comparison, Z Z
v · n dA = divv dV.
∂Ω Ω
iii) ∀a = const. ∈ V, by the lemma of divergence, property iv) of Theorem 30 and ii) just
proved,
Z Z Z Z
⊤ ⊤
div(L a)dV = (L a) · n dA = a · Ln dA = a · Ln dA,
Ω ∂Ω ∂Ω ∂Ω

95
but also
Z Z Z

div(L a)dV = (divL) · a + L · grada dV = a · divL dV,
Ω Ω Ω

so, once more by comparison,


Z Z
Ln dA = divL dV.
∂Ω Ω

The following identities follow directly from the Gauss theorem:


Z Z
v · Ln dA = (v · divL + L · gradv)dV,
Z ∂Ω ΩZ

(Ln) ⊗ v dA = ((divL) ⊗ v + L(gradv⊤ ))dV, (5.5)


Z∂Ω ZΩ
(w · n)v dA = (vdivw + (gradv)w)dV.
∂Ω Ω

A direct consequence of the Gauss theorem is the following result.

Theorem 38. (Flux theorem). Let v(p) be a vector field of class ≥ C1 on an open
subset R of E. Then,
Z
divv = 0 ⇐⇒ v · n dA = 0 ∀Ω ⊂ R.
∂Ω

Proof. It immediately follows from the ii) of the Gauss theorem.

Another consequence of the Gauss Theorem is the next theorem.

Theorem 39. (Curl theorem). Let v(p) be a vector field of class ≥ C1 on a regular
region Ω ⊂ E; then Z Z
n × v dA = curlv dV.
∂Ω Ω

Proof. If V is the axial tensor of v, by Theorem 32 and iii) of the Gauss theorem,
Z Z Z Z Z
n × v dA = − v × n dA = − Vn dA = − divV dV = curlv dV.
∂Ω ∂Ω ∂Ω Ω Ω

The following classical theorems on fields are recalled here without proof.

96
Theorem 40. (Potential theorem). Let v(p) be a vector field of class ≥ C1 on a
simply connected region Ω ⊂ E. Then,

curlv = o ⇐⇒ v = gradφ

with φ(p) the potential, a scalar field of class ≥ C2 .

Theorem 41. (Stokes theorem). Let v(p) be a vector field of class ≥ C1 on a regular
region Ω ⊂ E, Σ an open surface whose support is the closed line γ and n ∈ S the normal
to Σ, see Fig. 5.1. Then,
I Z
v · dℓ = curlv · n dA.
γ Σ

The parametric equation of γ must be chosen in such a way that

p′ (t1 ) × p′ (t2 ) · n > 0 ∀t2 > t1 .

𝛴 n

Figure 5.1: Scheme for the Stokes theorem.

Theorem 42. (Green’s formula). Let φ(p), ψ(p) be two scalar fields of class ≥ C2 on
a regular region Ω ⊂ E, with n ∈ S the normal to ∂Ω. Then,
Z   Z
dφ dψ
ψ −φ dA = (ψ ∆φ − φ ∆ψ)dV.
∂Ω dn dn Ω

5.5 Differential operators in Cartesian coordinates


The Cartesian expression of the differential operators can be found without difficulty
by applying the properties of such operators shown previously and considering that the

97
vectors of a Cartesian basis are fixed. The final result is5 ,

gradf = f,i ei ,
gradv = vi,j ei ⊗ ej ,
divv = vi,i ,
divL = Lij,j ei , (5.6)
∆f = f,ii ,
∆v = ∆vi ei = vi,jj ei ,
curlv = (v3,2 − v2,3 )e1 + (v1,3 − v3,1 )e2 + (v2,1 − v1,2 )e3 .

The so-called operator nabla ∇,


∂· ∂· ∂· ∂·
∇ := ei = e1 + e2 + e3 (5.7)
∂xi ∂x1 ∂x2 ∂x3
is often used to indicate the differential operators:

gradf = ∇f,
divv = ∇ · v,
curlv = ∇ × v,
∆f = ∇2 f.

5.6 Differential operators in cylindrical coordinates


The cylindrical coordinates ρ, θ, z of a point p, whose Cartesian coordinates in the (fixed)
frame {o; e1 , e2 , e3 } are p = (x1 , x2 , x3 ), are shown in Fig. 5.2. They are related together
by
q
ρ = x21 + x22 ,
x2 (5.8)
θ = arctan ,
x1
z = x3 ,

or conversely

x1 = ρ cos θ,
x2 = ρ sin θ, (5.9)
x3 = z.

To notice that ρ ≥ 0 and that the anomaly θ is bounded by 0 ≤ θ < 2π. A vector
p − o = xi ei , in the cylindrical basis is expressed has

p − o = ρeρ + zez
5
In what follows, and also in the following sections, f, v, L are, respectively, scalar, vector and tensor
fields.

98
e3

ez
p e𝜽
o e𝝆

𝜽
z e2
e1 𝝆

Figure 5.2: Cylindrical coordinates.

and the rotation tensor transforming the Cartesian basis {e1 , e2 , e3 } into the cylindrical
one, {eρ , eθ , ez } is  
cos θ − sin θ 0
Q =  sin θ cos θ 0 ,
0 0 1
so the relations between the vectors of the Cartesian and the cylindrical bases are

eρ = cos θe1 + sin θe2 ,


eθ = − sin θe1 + cos θe2 ,
ez = e3 ,

and viceversa
e1 = cos θeρ − sin θeθ ,
e2 = sin θeρ + cos θeθ , (5.10)
e3 = ez .

The question is: How can we express the differential operators in the (moving) frame
{p; eρ , eθ , ez }? To this end, we can proceed as follows: From Eq. (5.8)
x1 x2


 f,1 = f,ρ − f,θ 2 ,

 ρ ρ
∂ρ ∂θ ∂z 
x x
f,i = f,ρ + f,θ + f,z → f,2 = f,ρ + f,θ 1 ,
2 (5.11)
∂xi ∂xi ∂xi 
 ρ ρ 2


f = f .
,3 ,z

So, by Eqs. (5.6)1 and (5.10),


   
x1 x2 x2 x1
gradf = fi ei = f,ρ − f,θ 2 (cos θeρ −sin θeθ )+ f,ρ + f,θ 2 (sin θeρ +cos θeθ )+f,z ez .
ρ ρ ρ ρ

Finally, by Eq. (5.9) and through some standard operations, we obtain


1
gradf = f,ρ eρ + f,θ eθ + f,z ez .
ρ

99
The gradient of a vector field v can be obtained in a similar way. If we denote by vCart the
vector v expressed by its Cartesian components (v1 , v2 , v3 ) and by vcyl the same vector
expressed through the cylindrical ones, {vρ , vθ , vz }, then, cf. Section 2.11,

v1 = vρ cos θ − vθ sin θ,

vCart = Qvcyl → v2 = vρ sin θ + vθ cos θ, (5.12)

v = v .
3 z

Applying Eq. (5.11) to these components, we get

x1 x2
vi,1 = vi,ρ − vi,θ 2 ,
ρ ρ
x2 x1 (5.13)
vi,2 = vi,ρ + vi,θ 2 ,
ρ ρ
vi,3 = vi,z .

Injecting these expressions into Eq. (5.6)2 for the vi,j s and the (5.10) for the ei s, gives
finally6

1
gradv = vρ,ρ (eρ ⊗ eρ ) + (vρ,θ − vθ )(eρ ⊗ eθ ) + vρ,z (eρ ⊗ ez )
ρ
1
+ vθ,ρ (eθ ⊗ eρ ) + (vθ,θ + vρ )(eθ ⊗ eθ ) + vθ,z (eθ ⊗ ez )
ρ
1
+ vz,ρ (ez ⊗ eρ ) + vz,θ (ez ⊗ eθ ) + vz,z (ez ⊗ ez ),
ρ

or, in matrix form,


1
 
 vρ,ρ (vρ,θ − vθ ) vρ,z
 ρ 

 1 
 vθ,ρ
gradv =  (vθ,θ + vρ ) vθ,z ,
 ρ 

 1 
vz,ρ vz,θ vz,z
ρ

By the definition of divergence, we get immediately

1
divv = vρ,ρ + (vθ,θ + vρ ) + vz,z . (5.14)
ρ

Now, from Eq. (5.6), we see that divL is the vector whose components are the divergence
of the rows of the matrix representing L. So, we need first to calculate the Cartesian

6
Though straightforward, the details of the calculations for this formula, as for the following ones, are
particularly long and tedious, and for this reason they are omitted here; however, they are a very good
exercise for the reader.

100
components of L as functions of the cylindrical ones, cf. Section 2.11:


 L11 = − sin θ(Lρθ cos θ − Lθθ sin θ) + cos θ(Lρρ cos θ − Lθρ sin θ),


L12 = cos θ(Lρθ cos θ − Lθθ sin θ) + sin θ(Lρρ cos θ − Lθρ sin θ),









 L13 = Lρz cos θ − Lθz sin θ,

 L21 = − sin θ(Lθθ cos θ + Lρθ sin θ) + cos θ(Lθρ cos θ + Lρρ sin θ),




LCart = QLcyl Q⊤ → L22 = cos θ(Lθθ cos θ + Lρθ sin θ) + sin θ(Lθρ cos θ + Lρρ sin θ),


 L23 = Lθz cos θ + Lρz sin θ,




L31 = Lzρ cos θ − Lzθ sin θ,






L32 = Lzθ cos θ + Lzρ sin θ,






 L =L .
33 zz

Then, applying Eqs. (5.10) and (5.14) in Eq. (5.6)3 for the vectors vi = (Li1 , Li2 , Li3 ), i =
1, 2, 3, we get, through long but standard passages and after putting θ = 0 in order to
obtain the components of divL in the basis {eρ , eθ , ez },
 
1
divL = ((ρLρρ ),ρ + Lρθ,θ − Lθθ ) + Lρz,z eρ
ρ
 
1
+ Lθρ,ρ + (Lθθ,θ + Lρθ + Lθρ ) + Lθz,z eθ
ρ
 
1
+ ((ρLzρ ),ρ + Lzθ,θ ) + Lzz,z ez .
ρ
To obtain ∆f = f,ii , we need to apply twice Eq. (5.11), which gives
 
x1 x2 x1 ρ − x1 ρ,1 x2 2x2 ρρ,1
f,11 = f,ρ − f,θ 2 = f,ρ1 + f,ρ − f ,θ1 + f ,θ
ρ ρ ,1 ρ ρ2 ρ2 ρ4
ρ2 − x21
   
x1 x2 x1 x1 x2 x2 2x1 x2
= f,ρρ − f,ρθ 2 + f,ρ 3
− f,ρθ − f,θθ 2 2
+ f,θ 4
ρ ρ ρ ρ ρ ρ ρ ρ
2 2
sin θ cos θ sin θ sin θ sin θ cos θ
= f,ρρ cos2 θ − 2f,ρθ + f,ρ + f,θθ 2 + 2f,θ ,
ρ ρ ρ ρ2
 
x2 x1 x2 ρ − x2 ρ,2 x1 2x1 ρρ,2
f,22 = f,ρ + f,θ 2 = f,ρ2 + f,ρ 2
+ f,θ2 2 − f,θ
ρ ρ ,2 ρ ρ ρ ρ4
ρ2 − x22
   
x2 x 1 x2 x2 x1 x1 2x1 x2
= f,ρρ + f,ρθ 2 + f,ρ 3
+ f,ρθ + f,θθ 2 2
− f,θ 4
ρ ρ ρ ρ ρ ρ ρ ρ
2 2
sin θ cos θ cos θ cos θ sin θ cos θ
= f,ρρ sin2 θ + 2f,ρθ + f,ρ + f,θθ 2 − 2f,θ ,
ρ ρ ρ ρ2
f,33 = f,zz .
Then, adding these terms together, we finally have
1 1
∆f = (ρf,ρ ),ρ + 2 f,θθ + f,zz .
ρ ρ

101
The laplacian ∆v of a vector field v, Eq. (5.6)6 , can be obtained by following the same
steps for each one of the components in Eq. (5.12), which gives
 
1 1 1
∆v = (ρvρ,ρ ),ρ + 2 vρ,θθ + vρ,zz − 2 (vρ + 2vθ,θ ) eρ
ρ ρ ρ
 
1 1 1
+ (ρvθ,ρ ),ρ + 2 vθ,θθ + vθ,zz − 2 (vθ − 2vρ,θ ) eθ
ρ ρ ρ
 
1 1
+ (ρvz,ρ ),ρ + 2 vz,θθ + vz,zz ez .
ρ ρ

Finally, injecting Eqs. (5.10), (5.12) and (5.13) into Eq. (5.6)7 gives
   
1 1
curlv = vz,θ − vθ,z eρ + (vρ,z − vz,ρ )eθ + ((ρvθ ),ρ − vρ,θ ) ez . (5.15)
ρ ρ

5.7 Differential operators in spherical coordinates


The spherical coordinates r, φ, θ of a point p, whose Cartesian coordinates in the (fixed)
frame {o; e1 , e2 , e3 } are p = (x1 , x2 , x3 ), are shown in Fig. 5.3. They are related together
by
q
r = x21 + x22 + x23 ,
p
x21 + x22
φ = arctan ,
x3
x2
θ = arctan ,
x1

or conversely

x1 = r cos θ sin φ,
x2 = r sin θ sin φ,
x3 = r cos φ.

e3

𝝋
e𝜽
p
o er
r
e𝝋
𝜽
e1 e2

Figure 5.3: Spherical coordinates.

102
We note that r ≥ 0 and that the anomaly θ is bounded by 0 ≤ θ < 2π while the colatitude
φ by 0 ≤ φ ≤ π.
The procedure to determine the expression of the differential operators in spherical coordi-
nates, i.e. in the (moving) frame {p; er , eφ , eθ }, is identical to that used for the cylindrical
coordinates, but the analytical developments are even more complicated and long, so they
are omitted here and only the final formulae are given in the following:

1 1
gradf = f,r er + f,φ eφ + f,θ eθ ,
r r sin φ
 
1 1 1
gradv = vr,r er ⊗ er + (vr,φ − vφ )er ⊗ eφ + vr,θ − vθ er ⊗ eθ
r r sin φ
 
1 1 1
+ vφ,r eφ ⊗ er + (vφ,φ + vr )eφ ⊗ eφ + vφ,θ − vθ cot φ eφ ⊗ eθ
r r sin φ
 
1 1 1
+ vθ,r eθ ⊗ er + vθ,φ eθ ⊗ eφ + vθ,θ + vr + vφ cot φ eθ ⊗ eθ ,
r r sin φ
or, in matrix form,
   
1 1 1
 vr,r (vr,φ − vφ ) vr,θ − vθ
 r r sin φ 


 1 1 1 
 vφ,r
gradv =  (vφ,φ + vr ) vφ,θ − vθ cot φ ,
 r r sin φ 


 1 1 1 
vθ,r vθ,φ vθ,θ + vr + vφ cot φ
r r sin φ

1 2 1
divv = 2
(r vr ),r + ((vφ sin φ),φ + vθ,θ ),
r r sin φ
 
1 2 1 1 Lφφ + Lθθ cot φ
divL = (r Lrr ),r + Lrφ,φ + Lrθ,θ − + Lrφ er
r2 r r sin φ r r
 
1 2 1 1 1 cot φ
+ (r Lφr ),r + Lφφ,φ + Lφθ,θ + Lrφ + (Lφφ − Lθθ ) eφ
r2 r r sin φ r r
 
1 2 1 1 1 cot φ
+ (r Lθr ),r + Lθφ,φ + Lθθ,θ + Lrθ + (Lφθ + Lθφ ) eθ ,
r2 r r sin φ r r
 
1 2 1 f,θθ
∆f = 2 (r f,r ),r + 2 + (f,φ sin φ),φ ,
r r sin φ sin φ
   
2vr,r vr,φφ − 2vφ,φ vr,φ − 2vφ 1 vr,θθ 2vr
∆v= vr,rr + + + 2 + 2 − 2vθ,θ − 2 er
r r2 r tan φ r sin φ sin φ r
 
2vφ,r vφ,φφ+2vr,φ vφ,φ−vφ cot φ 1 vφ
+ vφ,rr + + + + 2 2 (vφ,θθ −2vθ,θ cos φ)− 2 eφ
r r2 r2 tan φ r sin φ r
     
2vθ,r vθ,φφ 2vφ,θ 1 1 vθ,θθ vθ
+ vθ,rr + + 2 + vθ,φ+ + +2vr,θ − 2 2 eθ ,
r r sin φ r2 tan φ r2 sin φ sin φ r sin φ

103
 
1
curlv = ((vθ sin φ),φ − vφ,θ ) er
r sin φ
 
1 1
+ vr,θ − (rvθ ),r eφ
r sin φ r
 
1
+ ((rvφ ),r − vr,φ ) eθ .
r

5.8 Exercises
1. Prove the relations of Eq. (5.2).
2. Prove the properties of the divergence in Eq. (5.3).
3. Prove the properties of the curl in Eq. (5.4).
4. Prove the identities in Eq. (5.5).
5. Prove that
dφ dv
= gradφ · n, = gradv n ∀n ∈ S.
dn dn
6. Prove the results of Eq. (5.6).
7. Consider a rigid body B, and a point p0 ∈ B. From the kinematics of rigid bodies,
we now that the velocity of another point p ∈ B is given by

v(p) = v(p0 ) + ω × (p − p0 ),

with ω the angular velocity. Prove that


1
ω = curlv, divv = 0.
2

8. In the infinitesimal theory of strain, a deformation is isochoric when divu = 0, u


being the corresponding displacement vector. Determine which, among the following
ones, are locally or globally isochoric deformations:
i. u = α(x1 , x2 , x3 ), α ∈ R, |α| ≪ 1;
ii. u = β(x2 + x3 , x1 + x3 , x1 + x2 ), β ∈ R, |β| ≪ 1;
iii. u = γ(x1 x2 , x2 x3 , x3 x1 ), γ ∈ R, |γ| ≪ 1;
iv. u = δ(sin x1 , − cos x2 , sin x3 ), δ ∈ R, |δ| ≪ 1.
9. In fluid mechanics, the condition divv = 0, with v being the velocity field, char-
acterizes incompressible flows. Verify that the following velocity fields, given in
cylindrical coordinates, correspond to incompressible flows (α ∈ R):
α
i. source or sink: v = eρ ;
ρ
α
ii. vortex: v = eθ ;
ρ

104
α
iii. doublet: v = (cos θeρ + sin θeθ ).
ρ2
10. A flow with curlv = o is said to be irrotational; check that the flows in the previous
exercise are irrotational.

105
106
Chapter 6

Curvilinear coordinates

6.1 Introduction
All the developments in the previous chapters are intended for the case where algebraic and
differential operators are expressed in a Cartesian frame, i.e. with rectangular coordinates.
The points of E are thus referred to a system of coordinates taken along straight lines
that are mutually orthogonal and with the same unit along each one of the directions
of the frame. Though this is a very important and common case, it is not the only
possibility and in many cases non rectangular coordinate frames are used or arise in the
mathematical developments (a typical example is that of the geometry of surfaces, see
Chapter 7). A non rectangular coordinate frame is a frame where coordinates can be
taken along non-orthogonal directions, or along some lines that intersect at right angles
but that are not straight lines, or even when both of these cases occur. This situation
is often denoted in the literature as that of curvilinear coordinates; the transformations
to be done to algebraic and differential operators in the case of curvilinear coordinates is
the topic of this chapter.

6.2 Curvilinear coordinates, metric tensor


Let us consider an arbitrary origin o of E and an orthonormal basis e = {e1 , e2 , e3 } of V;
we indicate the coordinates of a point p ∈ E with respect to the frame R = {o; e1 , e2 , e3 }
by xk : p = (x1 , x2 , x3 ). Then, we also consider another set of coordinate lines for E,
where the position of a point p ∈ E with respect to the same arbitrary origin o of E is now
determined by a set of three numbers z j : p = {z 1 , z 2 , z 3 }. Nothing is a-priori required of
coordinates z j , namely they do not need to be a set of Cartesian coordinates, i.e. referring
to an orthonormal basis of V. In principle, the coordinates z j can be taken along non
straight lines, that do not need to be mutually orthogonal at o and also with different
units along each line. That is why we call the z j s curvilinear coordinates, see Fig. 6.1.
Any point p ∈ E can be identified by either set of coordinates; mathematically, this means
that there must be an isomorphism between the xk s and the z j s, i.e. invertible relations

107
z3
x3

x2
o z2
x1
z1
Figure 6.1: Cartesian and curvilinear coordinates.

of the kind
z j = z j (x1 , x2 , x3 ) = z j (xk ),
∀j, k = 1, 2, 3 (6.1)
xk = xk (z 1 , z 2 , z 3 ) = xk (z j ),
exist between the two sets of coordinates. The distance between two points p, q ∈ E
is1 q
p
s = (p − q) · (p − q) = (xpk − xqk )(xpk − xqk )
but this is no longer true for curvilinear coordinates:
p
s ̸= (z j p − z j q )(z j p − z j q ).

However, if p → q, we can define


p q
dxk = xpk − xqk , dz j = z j − z j ,

so using Eq. (6.1)2


∂xk j
dxk = dz . (6.2)
∂z j

The (infinitesimal) distance between p and q will then be


r
∂xk ∂xk j l
p q
ds = dxk dxk = dz dz = gjl dz j dz l ,
∂z j ∂z l
where
∂xk ∂xk
gjl = glj = (6.3)
∂z j ∂z l
are the covariant2 components of the metric tensor3 g ∈ Sym(V). We note that, as g
defines a positive quadratic form (the square length of a vector), it is a positive definite
symmetric tensor, so
det g > 0. (6.4)
1
The distance between two points p and q is still defined as the Euclidean norm of (p − q), i.e., it is
independent of the set of coordinates.
2
The notion of co- and contra-variant components will be detailed in the next section.
3
As usually done in the literature, we indicate the metric tensor by g, i.e. using a lowercase letter,
though it is a 2nd-rank tensor, not a vector.

108
Coming back to the vector notation, from Eq. (6.2) we get4

∂xi k
dx = dxi ei = dz ei ;
∂z k
introducing the vector gk ,
∂xi
gk := ei , (6.5)
∂z k
we can write
dx = dz k gk .
We see hence that a vector dx can be expressed as a linear combination of the vectors gk ;
these form therefore a basis, called the local basis. Generally speaking, gk ∈/ S and it is
j
clearly tangent to the lines z = const. This can be seen in Fig. 6.2 for a two-dimensional
case:

z1=const.

g1 g2

𝛥x z2=const.
𝛥z 2)
, z 2+
x(z ,

x(z 1

e2
1 z2 )

o e1

Figure 6.2: Tangent vectors to the curvilinear coordinates lines.

xi (z 1 , z 2 + ∆z 2 ) − xi (z 1 , z 2 ) 2 ∂xi
dx = lim ∆x = lim 2
∆z e i = 2
ei dz 2 = g2 dz 2 .
∆x→0 ∆x→0 ∆z ∂z
Then
∂xi ∂xj ∂xi ∂xj
gk · gl =
k
ei · l ej = k l δij = gkl , (6.6)
∂z ∂z ∂z ∂z
i.e. the components of the metric tensor g are the scalar products of the tangent vectors
gk s. If the curvilinear coordinates are orthogonal, i.e. if gh · gk = 0 ∀h, k = 1, 2, 3, h ̸= k,
then g is diagonal. If, in addition, gk ∈ S ∀k = 1, 2, 3, then g = I: It is the case of
Cartesian coordinates. As an example, let us consider the case of polar coordinates,
( 1 p

x1 = r cos θ, z = r = x21 + x22 ,
x2
x2 = r sin θ, z 2 = θ = arctan .
x1
4
The differential dx is a vector because it is the difference of two infinitely close points; that is why it
is not needed to denote it in bold letters.

109
Hence, see Fig. 6.3,

∂x1 ∂x2
g1 = e1 + e2 = cos θe1 + sin θe2 = er ,
∂z 1 ∂z 1
∂x1 ∂x2
g2 = 2 e 1 + e2 = −r sin θe1 + r cos θe2 = reθ .
∂z ∂z 2

We remark that |g1 | = 1 but |g2 | =


̸ 1 and it is variable with the position.

re𝜃=g2

e𝜃 er=g1

e2 r

r=c
st.

ons
n
co

𝜃 t.
𝜃=

e1

Figure 6.3: Tangent vectors to the polar coordinates lines.

6.3 Co- and contra-variant components


A geometrical way to introduce the concept of covariant and contravariant components is
to consider how to represent a vector v in the z−system. There are basically two ways,
cf. Fig. 6.4, referred, for the sake of simplicity, to a planar case:
i. contravariant components: v is projected parallel to z 1 and z 2 ; they are indicated by
superscripts: v = (v 1 , v 2 , v 3 );
ii. covariant components: v is projected perpendicularly to z 1 and z 2 ; they are indicated
by subscripts: v = (v1 , v2 , v3 );

x2 x2
v2

vc2 vc2
z2 z2
v2 v v

𝛼2 z1 v1 𝛼2 z1 v1
𝛼1 𝛼1
vc1 x1 vc1 x1

Figure 6.4: Contravariant (left) and covariant (right) components of a vector in a plane.

110
Still referring to the planar case in Fig. 6.4, if the Cartesian components5 of v are
v = (v1c , v2c ), we get
 1 
v = h(v1c sin α2 − v2c cos α2 ), v1 = v1c cos α1 + v2c sin α1 ,
2 c c (6.7)
v = h(−v1 sin α1 + v2 cos α1 ), v2 = v1c cos α2 + v2c sin α2 ,

and conversely
 c 
v1 = v 1 cos α1 + v 2 cos α2 , v1c = h(v1 sin α2 − v2 sin α1 ),
(6.8)
v2c = v 1 sin α1 + v 2 sin α2 , v2c = h(−v1 cos α2 + v2 cos α1 ),

with
1
h= .
sin(α2 − α1 )
It is apparent that the Cartesian coordinates are at the same time co- and contra-variant.
Still on a planar scheme, we can see how to pass from a system of coordinates to another
one, cf. Fig. 6.5. For a point p the Cartesian coordinates (x1 , x2 ) are related to the

x2
z2

x2 p

z2

𝛼2
z1
𝛼1 z1
x1 x1

Figure 6.5: Relation between Cartesian and contravariant components.

contravariant ones by

x1 = z 1 cos α1 + z 2 cos α2 ,
x2 = z 1 sin α1 + z 2 sin α2 ,

and, conversely,

z 1 = h(x1 sin α2 − x2 cos α2 ),


z 2 = h(−x1 sin α1 + x2 cos α1 ).

So, differentiating, we get


∂x1 ∂x1
= cos α1 , = cos α2 ,
∂z 1 ∂z 2
∂x2 ∂x2
= sin α1 , = sin α2 ,
∂z 1 ∂z 2
5
In the following, we use the superscript c to indicate a Cartesian component: vic is the i-th Cartesian
component of v ∈ V and Lcij the ij-th Cartesian component of L ∈ Lin(V).

111
and
∂z 1 ∂z 1
= h sin α2 , = −h cos α2 ,
∂x1 ∂x2
∂z 2 ∂z 2
= −h sin α1 , = h cos α1 .
∂x1 ∂x2
Injecting these expressions into Eqs. (6.7) and (6.8) gives
∂x1 ∂x2
v1 = v1c + v2c 1 ,
∂z 1 ∂z → v = ∂xk v c (6.9)
i
∂x1 ∂x2 ∂z i k
v2 = v1c 2 + v2c 2 ,
∂z ∂z
and
∂z 1 ∂z 1
v 1 = v1c + v2c ,
∂x1 ∂x2 i ∂z i c
→ v = v . (6.10)
2 c ∂z
2
c ∂z
2
∂xk k
v = v1 + v2 ,
∂x1 ∂x2
Now, if we calculate
∂z i c
ghi v i = ghi v ,
∂xk k
from Eq. (6.3) and by the chain rule6 we get
∂xj ∂xj ∂z i c
ghi v i = v
∂z h ∂z i ∂xk k
∂xj ∂xj c ∂xj ∂xk
= h vk = h δjk vkc = h vkc = vh ,
∂z ∂xk ∂z ∂z
i.e. we obtain the rule of lowering of the indices for passing from contravariant to covariant
components:
vh = ghi v i .
Introducing the inverse7 to ghi as
∂z h ∂z i
g hi = , (6.11)
∂xk ∂xk
we get, still using the chain rule,
∂xk c ∂z h ∂z i ∂xk c
g hi vi = g hi vk = v
∂z i ∂xj ∂xj ∂z i k
∂z h ∂xk c ∂z h ∂z h c
= vk = δjk vkc = v = vh,
∂xj ∂xj ∂xj ∂xk k
6
The reader can easily see that, in practice, the chain rule allows us to handle the derivatives as
fractions.
7
To prove that the contravariant components g pq are the inverse of the covariant ones, gpq , is direct:
∂z p ∂z q ∂xj ∂xj
g pq gpq = = δjk δjk = 1.
∂xk ∂xk ∂z p ∂z q

112
which is the rule of raising of the indices for passing from covariant to contravariant
components:
v h = g hi vi .
Again applying the chain rule, by Eq. (6.9) we get

∂z i ∂z i ∂xk c ∂xk c
vi = vk = v = δkl vkc ,
∂xl ∂xl ∂z i ∂xl k
i.e.
∂z i
= vkcvi , (6.12)
∂xk
which is the converse of Eq. (6.9). In a similar way, we get the converse of Eq. (6.10):
∂xk i
vkc = v. (6.13)
∂z i
Let us now calculate the norm v of a vector v; starting from the Cartesian components
and using the last two results,
s s
√ p c c ∂z i ∂z j ∂z i ∂z j p
v = v · v = vk vk = vi vj = vi vj = g ij vi vj ,
∂xk ∂xk ∂xk ∂xk

or also
r r
√ p ∂xk i ∂xk j ∂xk ∂xk i j p
v= v·v = vkc vkc = v v = v v = gij v i v j
∂z i ∂z j ∂z i ∂z j
and even
s s
√ p ∂z i ∂xk ∂z i ∂xk j q i p
v = v · v = vkc vkc = vi j v j = vi v = δ j v i v j = vi v i .
∂xk ∂z ∂xk ∂z j

Through Eq. (6.13) and by the definition of the tangent vectors to the lines of curvilinear
coordinates, Eq. (6.5), for a vector v we get
∂xi
v = vic ei = v k e i = v k gk .
∂z k
We see hence that the contravariant components are actually the components of v in the
local basis, composed by vectors gk s, tangent to the lines of curvilinear coordinates. In a
similar manner, if we introduce the dual basis whose vectors gk are defined as

∂z k
gk := ei (6.14)
∂xi
and proceeding in the same way, we obtain that

∂z k
v = vic ei = vk ei = vk gk ,
∂xi

113
i.e. the covariant components are actually the components of v in the dual basis. Finally,
for a vector, we have, alternatively,
v = vic ei = v k gk = vk gk . (6.15)

Just as for the gk s, we have


 h   k 
h k ∂z ∂z ∂z h ∂z k ∂z h ∂z k
g ·g = ei · ej = δij = = g hk ;
∂xi ∂xj ∂xi ∂xj ∂xi ∂xi
moreover
 h  
∂z h ∂xj ∂z h ∂xi ∂z h

h ∂z ∂xj
g · gk = ei · k
e j = k
δij = k
= k
= δ hk ,
∂xi ∂z ∂xi ∂z ∂xi ∂z ∂z
and, by the symmetry of the scalar product,
δhk := gh · gk = gk · gh = δ kh .
The last equations defines the orthogonality conditions for the g-vectors. Using these
results and Eq. (6.15) we also have
v k = δ kh v h = gk · v h gh = gk · v = gk · vh gh = g kh vh ,
vk = δkh vh = gk · vh gh = gk · v = gk · v h gh = gkh v h ,
thus finding again the rules of raising and lowering of the indices.
What was done for vectors can be transposed, using a similar approach, to tensors. In
particular, for a second-rank tensor L, we get
∂z i ∂z j c
Lij = L ,
∂xh ∂xk hk (6.16)
∂xh ∂xk c
Lij = L ,
∂z i ∂z j hk
for the contravariant and covariant components, respectively, while we can also introduce
the mixed components
∂z i ∂xk c
Li j = L ,
∂xh ∂z j hk
(6.17)
∂xh ∂z j c
Li j = L .
∂z i ∂xk hk
Conversely,
∂xh ∂xk ij
Lchk = L ,
∂z i ∂z j
∂z i ∂z j
Lchk = Lij ,
∂xh ∂xk
(6.18)
∂xh ∂z j i
Lchk = L ,
∂z i ∂xk j
∂z i ∂xk j
Lchk = L .
∂xh ∂z j i

114
Also for L, the rule of lowering or raising the indices is valid:
Lij = g ih g jk Lhk , Lij = gih gjk Lhk . (6.19)
From Eq. (6.18) and by the same definitions of gij , eq.(6.3), and g ij , Eq. (6.11), we
get
∂xi ∂xj
L = Lcij ei ⊗ ej = h k Lhk ei ⊗ ej = Lhk gh ⊗ gk
∂z ∂z
and
∂z h ∂z k
L = Lcij ei ⊗ ej = Lhk ei ⊗ ej = Lhk gh ⊗ gk .
∂xi ∂xj
In a similar manner, the tensor mixed components are also found:
∂xi ∂z k h
L = Lcij ei ⊗ ej = L ei ⊗ ej = Lhk gh ⊗ gk
∂z h ∂xj k
and
∂z k ∂xi k
L = Lcij ei ⊗ ej = L ei ⊗ ej = Lhk gh ⊗ gk .
∂xj ∂z h h
We see hence that a second-rank tensor can be given with four different combinations
of coordinates; even more complex is the case of higher order tensors, which will not be
treated here.
∂z i
Still by eqs.(6.3) and (6.11) and applying the chain rule to δji = j , we get
∂z
∂xk ∂xk ∂xh ∂xk
gij = i j
= δhk ,
∂z ∂z ∂z i ∂z j
∂z i ∂z j ∂z i ∂z j
g ij = = δhk ,
∂xk ∂xk ∂xh ∂xk
(6.20)
i ∂z i ∂xk
δj= δhk ,
∂xh ∂z j
∂xh ∂z j
δi j = δhk .
∂z i ∂xk
So, applying Eq. (6.18) to the identity tensor, we get
∂xi ∂xj hk
I = δij ei ⊗ ej = h k
I ei ⊗ ej = I hk gh ⊗ gk ,
∂z ∂z
but by Eqs. (6.16) and (6.20),
∂z h ∂z k
I hk = δij = g hk
∂xi ∂xj
so, finally,
I = g hk gh ⊗ gk .
Proceeding in a similar manner, we can also get
I = ghk gh ⊗ gk = δ hk gh ⊗ gk = δhk gh ⊗ gk .
We see hence that the ghk s represent I in covariant coordinates, the g hk s in the contravari-
ant ones and the δhk s and δ hk s in mixed coordinates.

115
6.4 Spatial derivatives of fields in curvilinear coordi-
nates
Let φ a spatial8 scalar field, φ : E → R. Generally speaking,
φ = φ(z j (xi )),
or also
φ = φ(xj (z k )),
where the xj s, z k s are, respectively, Cartesian and curvilinear coordinates, related as in
Eq. (6.1). By the chain rule
∂φ ∂φ ∂z k
= k (6.21)
∂xj ∂z ∂xj
and inversely
∂φ ∂φ ∂xj
k
= .
∂z ∂xj ∂z k
We remark that the last quantity transforms like the components of a covariant vector,
cf. Eq. (6.9).
The gradient of φ is the vector that in the Cartesian basis, cf. Eq. (5.6)1 , is given by
∂φ
∇φ = ej ;
∂xj
so by Eqs. (6.14) and (6.21) we get that, in the dual basis,
∂φ ∂z k ∂φ k
∇φ = e j = g .
∂z k ∂xj ∂z k
We see hence that in curvilinear coordinates the nabla operator, Eq. (5.7), is defined
by
∂·
∇(·) = k gk . (6.22)
∂z
The contravariant components of the gradient can be obtained by the covariant ones
upon multiplication by the components of the inverse (contravariant) metric tensor, Eq.
(6.11):
∂φ ∂z h ∂z k ∂φ ∂xj ∂φ ∂z h ∂φ ∂z h ∂φ ∂z h
g hk = = δij = → ∇φ = gh .
∂z k ∂xi ∂xi ∂xj ∂z k ∂xj ∂xi ∂xj ∂xj ∂xj ∂xj

Let us now consider a vector field v : E → V; we want to calculate the spatial derivative
of its Cartesian components. By the chain rule and Eq. (6.13), we get
∂vic ∂vic ∂z k ∂z k ∂ ∂z k ∂xi ∂v h ∂ 2 xi l
   
∂xi h
= k = v = + v
∂xj ∂z ∂xj ∂xj ∂z k ∂z h ∂xj ∂z h ∂z k ∂z k ∂z l
∂z k ∂xi ∂v h ∂z h ∂ 2 xm l
 
= + v ,
∂xj ∂z h ∂z k ∂xm ∂z k ∂z l
8
The term spatial here refers to differentiation with respect to spatial coordinates, which can be
Cartesian or curvilinear.

116
whence
∂z h ∂xj ∂vic ∂v h ∂z h ∂ 2 xm l
= + v. (6.23)
∂xi ∂z k ∂xj ∂z k ∂xm ∂z k ∂z l
Comparing this result with Eq. (6.17)1 we see that the first member actually corresponds
to the components of a mixed tensor field, which is the gradient of the vector field v, that
we write as
∂v h
v h;k = k + Γhkl v l , (6.24)
∂z
where the functions
∂z h ∂ 2 xm
Γhkl = (6.25)
∂xm ∂z k ∂z l
are the Christoffel symbols. We immediately see that Γhkl = Γhlk . The quantity v h;k is the
covariant derivative of the contravariant components v h . The proof that the Christoffel
symbols can also be written as
 
h 1 hm ∂gmk ∂gml ∂gkl
Γkl = g + − m (6.26)
2 ∂z l ∂z k ∂z
is left to the reader as an exercise.
Proceeding in a similar way for the covariant components of v, but now using Eqs. (6.12)
and (6.17)1 , we get
∂vh
vh;k = k − Γlkh vl ,
∂z
which is the covariant derivative of the covariant components vh .
Using Eqs. (6.23) and (6.24), we conclude that, cf. Eq. (5.6)3 ,
∂vic
divv = = v h;h .
∂xi

Then, applying the operator divergence so defined to the gradient of the scalar field φ we
obtain, in curvilinear coordinates z k , the Laplacian ∆φ as
   
hk ∂φ ∂ hk ∂φ ∂φ
∆φ = g k
= h g k
+ Γhhj g jk k . (6.27)
∂z ;h ∂z ∂z ∂z

Using the definition of the nabla operator in curvilinear coordinates, Eq. (6.22), jointly
to the fact that, cf. Section 5.5,

∆f := div∇φ = ∇ · ∇φ,

we get the the following representation of the Laplace operator in curvilinear coordi-
nates:
∂ 2 (·) h k ∂gh ∂(·) k
  
∂ ∂(·) h k
∆(·) = ∇ · ∇(·) = g · g = g ·g + k h ·g
∂z k ∂z h ∂z k ∂z h ∂z ∂z
∂ 2 (·) ∂gh ∂(·)
= k h g hk + k · gk h .
∂z ∂z ∂z ∂z
117
Let us now calculate the spatial derivatives of the components of a 2nd -rank tensor L: By
Eqs. (6.18)1 and (6.25) we get

∂Lcij ∂z h ∂
 
∂xi ∂xj np
= L
∂xk ∂xk ∂z h ∂z n ∂z p
∂z h ∂xi ∂xj ∂Lnp
 
n rp p nr
= + Γhr L + Γhr L ,
∂xk ∂z n ∂z p ∂z h

which implies that

∂Lnp n rp p nr ∂xk ∂z n ∂z p ∂Lcij


+ Γhr L + Γhr L = h . (6.28)
∂z h ∂z ∂xi ∂xj ∂xk

So, using Eq. (6.16), we can conclude that the expression

∂Lnp
Lnp;h = h
+ Γnhr Lrp + Γphr Lnr (6.29)
∂z
represents the covariant derivative of the contravariant components of the second-rank
tensor L. In a similar manner, this time by Eq. (6.18)2 , we obtain the covariant derivatives
of the covariant components of L:

∂Lnp r r ∂xk ∂xi ∂xj ∂Lcij


− Γ L
nh rp − Γ L
ph nr = ,
∂z h ∂z h ∂z n ∂z p ∂xk

i.e.
∂Lnp
Lnp;h = − Γrph Lnr − Γrnh Lpr . (6.30)
∂z h
The same procedure with Eqs. (6.18)3,4 gives the covariant derivatives of the mixed
components9 of L:

∂Lnp
Lnp;h = h
+ Γnhr Lrp − Γrph Lnr ,
∂z (6.31)
∂Lpn
Lpn;h = − Γrph Lrn + Γnhr Lpr .
∂z h

If in Eqs. (6.28) and (6.29) we set p = h, we get

∂Lnh ∂xk ∂z n ∂z h ∂Lcij


Lnh;h = n rh h nr
+ Γhr L + Γhr L = h
∂z h ∂z ∂xi ∂xj ∂xk
n ∂Lc n ∂Lc
∂z ij ∂z ij
= δkj = ,
∂xi ∂xk ∂xi ∂xj

which are the components of the contravariant vector field divL.


9
Equations (6.29), (6.30) and (6.31) represent the different forms of the components of an operator
depending upon three indices, i.e. of a third-rank tensor: ∇L, the gradient of L.

118
6.5 Exercises
1. Write g and ds for cylindrical coordinates.
2. Write g and ds for spherical coordinates.
3. Find the length of a helix traced on a circular cylinder of radius R between the
angles θ and θ + 2π.
4. A curve is traced in a quarter circle of radius R, see Fig. 6.6, with ρ proportional to
θ. When the quarter of circle is rolled into a cone, the curve appears as indicated
in the figure. Determine the length ℓ of the curve, first using the polar coordinates
in the plane of the quarter circle, then the cylindrical ones for the case of the curve
on the cone (exercise given in the book by W. H. Müller, see the suggested texts).

R0

𝜌(z)
x2
𝜌 R h
x3
𝜃
x1 R
x2
x1
Figure 6.6: Curve in a plane and on a cone.

5. Calculate g for a planar system of coordinates composed of two axes z 1 and z 2


inclined, respectively, at α1 and α2 on the axis x1 . Then, find the vectors gk and
gk , k = 1, 2, check the orthogonality conditions gh · gk = δ hk , determine the norm
of these vectors and design them.
6. Calculate the gi s for a system of spherical coordinates.
7. In the plane, elliptical coordinates are defined by the relations

x1 = c cosh z 1 cos z 2 , x2 = c sinh z 1 sin z 2 , z 1 ∈ (0, ∞), z 2 ∈ [0, 2π);

show that the lines z 1 = const. and z 2 = const. are confocal ellipses and hyperbolae,
determine the axes of the ellipses in terms of the parameter c, discuss the limit case
of ellipses that degenerate into a crack and determine its length. Finally, find g, g1
and g2 .
8. Determine the co- and contravariant components of a tensor L in cylindrical coor-
dinates.
9. Determine the co- and contravariant components of a tensor L in spherical coordi-
nates.

119
10. Show that
trL = Lcii = gij Lij = g ij Lij = Li i = Lj j .

11. Prove Eq. (6.26).


12. Prove the Lemma of Ricci:
∂gjk
h
= Γijh gik + Γikh gji .
∂z

13. Using Eq. (6.26), find the Christoffel symbols for the cylindrical, spherical and
elliptical (in the plane) coordinates.
14. Write the Laplacian ∆f of a spatial scalar field f in cylindrical and spherical coor-
dinates.
15. Prove that
g np;h = gnp;h = 0.

120
Chapter 7

Surfaces in E

7.1 Surfaces in E, coordinate lines, tangent planes


A function f (u, v) : Ω ⊂ R2 → E of class ≥C1 and such that its Jacobian
 
∂f1 ∂f1
 ∂u ∂v 
 
 ∂f2 ∂f2 
J =  ∂u ∂v 

 
 ∂f ∂f 
3 3
∂u ∂v
has maximum rank (rank[J]=2) defines a surface in E, see Fig. 7.1. We say also that f is
an immersion of Ω into E and that the subset Σ ⊂ E image of f is the support or trace of
the surface f .

x3

f,v
f(u,v)
p=f(q) Tp𝛴
f,u
v
𝛺 v=
𝜸
v0
^
0

o
u=u

x2
𝛴
𝜸
v0 q

u0 x1
u
Figure 7.1: General scheme of a surface and of the tangent space at a point p.

As usual, we indicate the derivatives with respect to the variables u, v by, for example,

121
∂f
= f,u etc. The condition on the rank of J is equivalent to impose that
∂u
f,u (u, v) × f,v (u, v) ̸= o ∀(u, v) ∈ Ω. (7.1)

This allows us to introduce the normal to the surface f as the vector N ∈ S defined
by
f,u × f,v
N := . (7.2)
|f,u × f,v |
A regular point of Σ is a point where N is defined; if N is defined ∀p ∈ Σ then the surface
is said to be regular.
A function γ(t) : G ⊂ R → Ω whose parametric equation is γ(t) = (u(t), v(t)) describes
a curve in Ω whose image, through f , is the curve, see Fig. 7.1,

b (t) = f (u(t), v(t)) : G ⊂ R → Σ ⊂ E.


γ

As a special case of curve in Ω, let us consider the curves of the type v = v0 or u = u0 , with
u0 , v0 being some constants. Then, their image through f are two curves f (u, v0 ), f (u0 , v)
on Σ called coordinate lines, see again Fig. 7.1. The tangent vectors to the coordinate
lines are, respectively, the vectors f,u (u, v0 ) and f,v (u0 , v), while the tangent to a curve
γ
b (t) = f (u(t), v(t)) is the vector
du dv
b ′ (t) = f,u
γ + f,v , (7.3)
dt dt
i.e. the tangent vector to any curve on Σ is a linear combination of the tangent vectors
to the coordinate lines. We remark that the tangent vectors f,u (u, v0 ) and f,v (u0 , v) are
necessarily non-null and linear independent as consequence of the assumption on the rank
of J and hence of the existence of N, i.e. of the regularity of Σ. They determine a plane
that contains the tangents to all the curves on Σ passing by p = f (u0 , v0 ) and form a basis
on this plane, called the natural basis. Such a plane is the tangent plane to Σ in p and is
indicated by Tp Σ; this plane is actually the space spanned by f,u (u, v0 ) and f,v (u0 , v) and
is also called the tangent vector space.
Let us consider two open subsets Ω1 , Ω2 ⊂ R2 ; a diffeomorphism1 of class Ck between Ω1
and Ω2 is a bijective map ϑ : Ω1 → Ω2 of class Ck with also its inverse of class Ck ; the
diffeomorphism is smooth if k = ∞.
Let Ω1 , Ω2 be two open subsets of R2 , f : Ω2 → E a surface and ϑ : Ω1 → Ω2 a smooth
diffeomorphism. Then the surface F = f ◦ ϑ : Ω1 → E is a change of parameterization for
f . In practice, the function defining the surface changes, but not Σ, its trace in E. Let
(U, V ) be the coordinates in Ω1 and (u, v) those in Ω2 . Then, by the chain rule,
∂u ∂v
F,U = f,u + f,v ,
∂U ∂U
∂u ∂v
F,V = f,u + f,v ,
∂V ∂V
1
The definition of diffeomorphism, of course, can be given for subsets of Rn , n ≥ 1; here, we bound
the definition to the case of interest.

122
or, denoting by Jϑ the Jacobian of ϑ,
   
F,U ⊤ f,u
= [Jϑ ] ,
F,V f,v

whence, making the cross product, one gets immediately

F,U × F,V = det[Jϑ ] f,u × f,v .

This result shows that the regularity of the surface, condition (7.1), the tangent plane
and the tangent space vector do not depend upon the parameterization of Σ. From the
last equation, we get also

N(U, V ) = sgn(det[Jϑ ]) N(u, v);

we say that the change of parameterization preserves the orientation if det[Jϑ ] > 0, and
that it inverses the parameterization in the opposite case.

7.2 Surfaces of revolution


A surface of revolution is a surface whose trace is obtained by letting a plane curve, say
γ, rotate around an axis, say x3 . To be more specific and without loss of generality, let
γ : G ⊂ R → R2 be a regular curve of the plane x2 = 0, whose parametric equation
is 
x1 = φ(u),
γ(u) : φ(u) > 0 ∀u ∈ G. (7.4)
x3 = ψ(u),
Then, the subset Σγ ⊂ E defined by

Σγ := (x1 , x2 , x3 ) ∈ E|x21 + x22 = φ2 (u), x3 = ψ(u), u ∈ G




is the trace of a surface of revolution of the curve γ(u) around the axis x3 . A general
parameterization of such a surface is

 x1 = φ(u) cos v,
f (u, v) : G × (−π, π] → E| x2 = φ(u) sin v, (7.5)
x3 = ψ(u).

It is readily checked that this parameterization actually defines a regular surface:


 ′     
 φ (u) cos v   −φ(u) sin v   −φ(u)ψ ′ (u) cos v 
f,u = φ′ (u) sin v , f,v = φ(u) cos v → f,u × f,v = −φ(u)ψ ′ (u) sin v
ψ ′ (u) 0 φ(u)φ′ (u)
     

so that
|f,u × f,v | = φ2 (u)(φ′2 (u) + ψ ′2 (u)) ̸= 0 ∀u ∈ G

123
for being γ(u) a regular curve, i.e. with γ ′ (u) ̸= o ∀u ∈ G. A meridian is a curve in E
intersection of the trace of f , Σγ , with a plane containing the axis x3 ; the equation of a
meridian is obtained fixing the value of v, say v = v0 :

 x1 = φ(u) cos v0 ,
x2 = φ(u) sin v0 ,
x3 = ψ(u).

A parallel is a curve in E intersection of Σγ with a plane orthogonal to x3 ; the equation


of a parallel, which is a circle with center on the axis x3 , is obtained by fixing the value
of u, say u = u0 :

 x1 = φ(u0 ) cos v,
x2 = φ(u0 ) sin v,
x3 = ψ(u0 ),

or also 
x21 + x22 = φ(u0 )2 ,
x3 = ψ(u0 );

the radius of the circle is φ(u0 ). A loxodrome or rhumb line is a curve on Σγ crossing
all the meridians at the same angle2 . Some important examples of surfaces of revolution

Figure 7.2: Surfaces of revolution. From the left: sphere, catenoid, pseudo-sphere, hyper-
bolic hyperboloid.

are:
• the sphere:

h π πi  x1 = cos u cos v,
f (u, v) : − , × (−π, π] → E| x2 = cos u sin v,
2 2
x3 = sin v;

• the catenoid:

 x1 = cosh u cos v,
f (u, v) : [−a, a] × (−π, π] → E| x2 = cosh u sin v,
x3 = u;

2
This concept is important for marine and aerial navigation

124
• the pseudo-sphere:

 x1 = sin u cos v,

f (u, v) : [0, a] × (−π, π] → E| x2 = sin u sin v, (7.6)
 x3 = cos u + ln tan u ;


2

• the hyperbolic hyperboloid:



 x1 = cos u − v sin u,
f (u, v) : [−a, a] × (−π, π] → E| x2 = sin u + v cos u,
x3 = v.

7.3 Ruled surfaces


A ruled surface (also named a scroll) is a surface with the property that through every
one of its points, there is a straight line that lies on the surface. A ruled surface can be
seen as the set of points swept by a moving straight line. We say that a surface is doubly
ruled if through every one of its points, there are two distinct straight lines that lie on the
surface.
Any ruled surface can be represented by a parameterization of the form

f (u, v) = γ(u) + vλ(u), (7.7)

where γ(u) is a regular smooth curve, the directrix, and λ(u) is a smooth curve. Fixing
u = u0 gives a generator line f (u0 , v) of the surface; the vectors λ(u) ̸= o describe the
directions of the generators. Some important examples of ruled surfaces are:
• Cones: For these surfaces, all the straight lines pass through a point, the apex of
the cone; choosing the apex as the origin, then it must be λ(u) = kγ(u), k ∈ R →

f (u, v) = vγ(u);

• Cylinders: A ruled surface is a cylinder ⇐⇒ λ(u) = const. In this case, it is always


possible to choose λ(u) ∈ S and γ(u) a planar curve lying in a plane orthogonal to
λ(u); in fact, it is sufficient to choose the curve γ ∗ (u) = (I − λ(u) ⊗ λ(u))γ(u);
• Helicoids: A surface generated by rotating and simultaneously displacing a curve,
the profile curve, along an axis is a helicoid. Any point of the profile curve is the
starting point of a circular helix. Generally, we get a helicoid if

γ(u) = (0, 0, φ(u)), λ(u) = (cos u, sin u, 0), φ(u) : R → R.

• Möbius strip: It is a ruled surface with

γ(u) = (cos 2u, sin 2u, 0), λ(u) = (cos u cos 2u, cos u sin 2u, sin u).

125
Figure 7.3: Ruled surfaces: (from left) elliptical cone, elliptical cylinder, helicoid and
Möbius strip.

7.4 First fundamental form of a surface


Let us consider two vectors of Tp Σ, say w1 , w2 ; we want to calculate their scalar product
in terms of their components in the natural basis {f,u , f,v } of Tp Σ. If w1 = a1 f,u + b1 f,v
and w2 = a2 f,u + b2 f,v , then
2
w1 · w2 = a1 a2 f,u + (a1 b2 + a2 b1 )f,u · f,v + b1 b2 f,v2 ,
which can be rewritten as the form
I(w1 , w2 ) = w1 · gw2 ,
where3  
f,u · f,u f,u · f,v
g=
f,v · f,u f,v · f,v
is precisely the metric tensor g of Σ, cf. Eq. (6.6). In fact, f,u and f,v are the tangent
vectors to the coordinate lines on Σ, i.e. they coincide with the vectors gk s.
I(w1 , w2 ) is the first fundamental form (or simply the first form) of f (u, v). If w1 = w2 =
w = af,u + bf,v , then
I(w) = w2 = a2 f,u
2
+ 2abf,u · f,v + b2 f,v2 .
By the same definition of scalar product, I(w1 , w2 ) is a positive definite, bilinear, sym-
metric form ∀w ∈ Tp Σ.
Through I(·, ·) we can calculate some important quantities regarding the geometry of
Σ:
• Metric on Σ: ∀ds ∈ Σ,
ds2 = ds · ds = I(ds);
so, if
ds = f,u du + f,v dv,
then
ds2 = f,u
2
du2 + 2f,u · f,v du dv + f,v2 dv 2 ; (7.8)
3
Often, in texts on differential geometry, tensor g is indicated as
 
E F
g= ,
F G
where E := f,u · f,u , F := f,u · f,v , G := f,v · f,v .

126
• Length ℓ of a curve γ : [t1 , t2 ] ⊂ R → Σ: We know, see Eq. (4.4), that the length of
a curve is the integral of the tangent vector:
Z t2 Z t2 p

ℓ= |γ (t)|dt = γ ′ (t) · γ ′ (t)dt
t1 t1

and hence, see Eq. (7.3), if we call w = (u′ , v ′ ) the tangent vector to γ, expressed
by its components in the natural basis,
Z t2 q Z t2 p
ℓ= ′2 2 ′ ′ ′2 2
u f,u + 2u v f,u · f,v + v f,v dt = (u′ , v ′ ) · g (u′ , v ′ )dt
t t1
Z 1t2 p (7.9)
= I(w)dt;
t1

• Angle θ formed by two vectors w1 , w2 ∈ Tp Σ:

w1 · w2 I(w1 , w2 )
cos θ = =p p ;
|w1 ||w2 | I(w1 ) I(w2 )

• Area of a small surface on Σ: Let f,u du and f,v dv be two small vectors on Σ, forming
together the angle θ, that are the transformed, through4 f : Ω → Σ, of two small
orthogonal vectors du, dv ∈ Ω; then the area dA of the parallelogram determined
by them is
q
dA = |f,u du × f,v dv| = |f,u × f,v |du dv = f,u 2 f 2 sin2 θdu dv
,v
q q
= f,u 2 f 2 (1 − cos2 θ)du dv = 2 f 2 − f 2 f 2 cos2 θdu dv
f,u
,v ,v ,u ,v
q p
= f,u 2 f 2 − (f · f )2 du dv = det gdu dv.
,v ,u ,v


The term det g is hence the dilatation factor of the areas; recalling Eq. (6.4), we
see that the previous expression has a sense ∀f (u, v), i.e. for any parameterization
of the surface.

7.5 Second fundamental form of a surface


Be f : Ω → Σ a regular surface, {f,u , f,v } the natural basis for Tp Σ and N ∈ S the normal
to Σ defined as in (7.2). We call map of Gauss of Σ the map φΣ : Σ → S that associates
to each p ∈ Σ its N : φΣ (p) = N(p). To each subset σ ⊂ Σ, the map of Gauss associates
hence a subset σS ⊂ S, Fig. 7.4 (e.g. the Gauss map of a plane is just a point on S).

4
For the sake of conciseness, from now on we will indicate a surface as the function f : Ω → Σ, with
f = f (u, v), (u, v) ∈ Ω ⊂ R2 and Σ ⊂ E.

127
x3 x3

𝝋𝝨(p)
N(p)
𝝈p 𝝈S
N(p)
o x2 S o x2
𝛴

x1 x1

Figure 7.4: The map of Gauss.

We want to study how N(p) varies at the varying of p on Σ. The idea is that the change
of N(p) on Σ is related to the curvature of the surface5 . For this purpose, we calculate
the change in N per unit length of a curve γ(s) ∈ Σ, i.e. we study how N varies along
any curve of Σ per unit of length of the curve itself; that is why we parameterize the
curve with its arc-length s 6 . Let N = Ni (u, v)ei ; then, if τ ∈ S is the tangent to the
curve,
 
dN dNi (u(s), v(s)) ∂Ni du ∂Ni dv
= ei = + ei
ds ds ∂u ds ∂v ds
dN
= ∇Ni · τ ei = (ei ⊗ ∇Ni )τ = (∇N) τ = .

The change in N is hence related to the directional derivative of N along the tangent τ
to γ(s), which is a linear operator on Tp Σ. Moreover, as N ∈ S, then, cf. Eq. (4.1),
N · N,u = N · N,v = 0 ⇒ N,u , N,v ∈ Tp Σ.
We then call Weingarten operator LW : Tp Σ → Tp Σ the opposite of the directional
derivative of N:
dN
LW (τ ) := − .

Hence,
LW (f,u ) = −N,u , LW (f,v ) = −N,v . (7.10)
Because LW is linear, then there exists a tensor X on Tp Σ such that
LW (v) = Xv ∀v ∈ Tp Σ. (7.11)
For any two vectors w1 , w2 ∈ Tp Σ, we define second fundamental form of a surface,
denoted by II(w1 , w2 ) the bilinear form
II(w1 , w2 ) := I(LW (w1 ), w2 ).
5
For curves, the curvature is linked to the change of τ , but for surfaces this should not be meaningful,
as τ is not unique ∀p ∈ Σ while N is.
6
Actually, it is also possible to introduce the following concepts more generally, i.e. for any parame-
terization of the curve but, for the sake of simplicity, in the following we just use the parameter s.

128
Theorem 43. (Symmetry of the second fundamental form). ∀w1 , w2 ∈ Tp Σ, II(w1 , w2 ) =
II(w2 , w1 ).

Proof. Because I and LW are linear, it is sufficient to prove the thesis for the natural
basis {f,u , f,v } of Tp Σ, and, by the symmetry of I, it is sufficient to prove that

I(LW (f,u ), f,v ) = I(f,u , LW (f,v )),

i.e. that
I(−N,u , f,v ) = I(f,u , −N,v )
and in the end that
N,u · f,v = f,u · N,v .
To this purpose, we recall that

N · f,u = 0 = N · f,v .

So, differentiating the first equation by v and the second one by u, we get

N,v · f,u = −N · f,uv = N,u · f,v . (7.12)

The second fundamental form defines a quadratic, bilinear symmetric form:

II(w1 , w2 ) = I(LW (w1 ), w2 ) = I(w1 , LW (w2 ))


= I(w1 , Xw2 ) = w1 · gXw2 = w1 · Bw2 ,

where
B := gX. (7.13)
In the natural basis {f,u , f,v } of Tp Σ, by Eq. (7.12), it is7

Bij = II(f,i , f,j ) = I(LW (f,i ), f,j ) = −N,i · f,j = N · f,ij ; (7.14)

tensor X can then be calculated by Eq. (7.13):

X = g−1 B. (7.15)

By Eq. (7.14), because f,ij = f,ji or simply because II(·, ·) is symmetric, we get that

B = B⊤ .
7
In many texts on differential geometry, the following symbols are used:

L = f,uu · N = −f,u · N,u ,


M = f,uv · N = −f,u · N,v ,
N = f,vv · N = −f,v · N,v .

129
7.6 Curvatures of a surface
Let f : Ω → Σ be a regular surface and γ(s) : G ⊂ R → Σ be a regular curve on Σ
parameterized with the arc length s. We call curvature vector of γ(s) the vector κ(s)
defined as
κ(s) := c(s)ν(s) = γ ′′ (s),
where ν(s) is the principal normal to γ(s). By Eq. (4.11), it is also

κ(s) = γ ′′ (s).

Then, we call normal curvature κN (s) of γ(s) the projection of κ(s) onto N(s), the normal
to Σ:
κN (s) := κ(s) · N(s) = c(s) ν(s) · N(s) = γ ′′ (s) · N(s).

Theorem 44. The normal curvature κN (s) of γ(s) ∈ Σ depends uniquely on τ (s):

κN (s) = τ (s) · Bτ (s) = II(τ (s), τ (s)). (7.16)

Proof.
γ(s) = γ(u(s), v(s)) → τ (s) = γ ′ (s) = f,u u′ + f,v v ′ ,
therefore τ = (u′ , v ′ ) in the natural basis and

κ(s) = γ ′′ (s) = f,u u′′ + f,v v ′′ + f,uu u′2 + 2f,uv u′ v ′ + f,vv v ′2

and finally, by Eqs. (7.2) and (7.14),

κN (s) = γ ′′ (s) · N(s) = B11 u′2 + 2B12 u′ v ′ + B22 v ′2 = τ · Bτ = II(τ , τ ).

If now s = s(t) is a change of parameter for γ, then

γ ′ (t) = |γ ′ (t)|τ (t),

so by the linearity of II(·, ·), we get

II(γ ′ (t), γ ′ (t)) = |γ ′ (t)|2 II(τ (t), τ (t)) = |γ ′ (t)|2 κN (t)

and finally,
II(γ ′ (t), γ ′ (t))
κN (t) = .
I(γ ′ (t), γ ′ (t))
To each point p ∈ Σ, it corresponds uniquely (in the assumption of regularity of the
surface f : Ω → Σ) a tangent plane and a tangent space vector Tp Σ. In p, there are
infinite tangent vectors to Σ, all of them belonging to Tp Σ. We can associate a curvature
to each direction t ∈ Tp Σ, i.e. to each tangent direction, in the following way: Let us
consider the bundle H of planes whose support is the straight line through p and parallel

130
to N. Then, any plane H ∈ H is a normal plane to Σ in p; each normal plane is uniquely
determined by a tangent direction t and the (planar) curve γ N t := H ∩ Σ is called a
normal section of Σ. If ν and N are, respectively, the principal normal to γ N t and the
normal to Σ in p, then
ν = ±N
for each normal section. We have, in this way, defined a function that to each tangent
direction t ∈ Tp Σ associates the normal curvature κN of the normal section γ N t :

II(t, t)
κN : S ∩ Tp Σ → R| κN (t) = .
I(t, t)

By the bilinearity of the second fundamental form, κN (t) = κN (−t).


A point p ∈ Σ is said to be a umbilical point if κN (t) = const. ∀t, it is a planar point
if κN (t) = 0 ∀t. In all the other points, κN takes a minimum and a maximum value on
distinct directions t ∈ Tp Σ.
Because B = B⊤ , by the spectral theorem, there exists an orthonormal basis {u1 , u2 } of
Tp Σ such that
B = βj uj ⊗ uj ,
with βj the eigenvalues of B. In such a basis, by Eq. (7.13) we get

II(ui , ui ) ui · Bui ui · gXui


κN (ui ) = = = , i = 1, 2.
I(ui , ui ) ui · gui ui · gui

Then, because {u1 , u2 } is an orthonormal basis, g = I and

κN (ui ) = ui · Xui , i = 1, 2,

i.e. X and B share the same eigenvectors. Moreover, cf. Section 2.8, we know that the
two directions u1 and u2 are the directions whereupon the quadratic form in the previous
equation gets its maximum, κ1 , and minimum, κ2 , values, and in such a basis,

X = κi ui ⊗ ui .

We call κ1 and κ2 the principal curvatures of Σ in p and u1 , u2 the principal directions of


Σ in p, see Fig. 7.5.
We call Gaussian curvature K the product of the principal curvatures:

K := κ1 κ2 = det X.

By Eq. (7.15) and the theorem of Binet, it is also

det B
K= . (7.17)
det g

131
x3

N
u2 Tp𝛴

p u1

o x2
𝛴

x1
Figure 7.5: Principal curvatures.

We define mean curvature H of a surface8 f : Ω → Σ at a point p ∈ Σ the mean of the


principal curvatures at p:
κ1 + κ2 1
H := = trX.
2 2
Of course, a change in the parameterization of a surface can change the orientation, cf.
Section 7.1, i.e. it can transforms N into its opposite one and, by consequence, change the
sign of the second fundamental form and hence of the normal and principal curvatures.
These last are hence defined to less the sign, and the mean curvature too, while the
principal directions, umbilicality, flatness and Gaussian curvature are intrinsic to Σ, i.e.
they do not depend on its parameterization.

7.7 The theorem of Rodrigues


Then principal directions of curvature have a property which is specified by the

Theorem 45. (Theoreom of Rodrigues). Let f (u, v) be a surface of class at least C2


and λ = (λu , λv ) ∈ Tp Σ; then
dN(p)
= −κλ λ (7.18)

if and only if λ is a principal direction; κλ is the principal curvature relative to λ.

Proof. Let λ be a principal direction of Tp Σ. Because N ∈ S, then


dN
· N = 0; (7.19)

8
The concept of mean curvature of a surface was introduced for the first time by Sophie Germain in
her celebrated work on the elasticity of plates (1815).

132
moreover,
  
0 0 0  λu 
dN
= ∇N λ =  0 0 0  λv = N,u λu + N,v λv . (7.20)

N,u N,v 1 0
 

Let µ = (µu , µv ) be the other principal direction of Tp Σ, then


λ · µ = 0 → I(λ, µ) = II(λ, µ) = 0.
Moreover,
dN
· µ = −II(λ, µ) = 0,

which implies, together with Eq. (7.19),
dN
= αλ. (7.21)

Therefore,
dN
· λ = −II(λ) = αλ · λ = αI(λ)

and finally,
II(λ)
α=− = −κλ .
I(λ)
Contrarily, if we assume Eq. (7.21), as before we get α = −κλ and to end, we just need
to prove that λ is a principal direction. From Eqs. (7.20) and (7.21), we get
λu N,u + λv N,v = −κλ (λu f,u + λv f,v ).
Projecting this equation onto f,u and f,v gives the two equations
Lλu + M λv = κλ (Eλu + F λv ),
(7.22)
M λu + N λv = κλ (Eλu + Gλv ),
with the symbols E, F, G, L, M and N defined in Notes 3 and 7 and used here for the
sake of conciseness. Let w = (wu , wv ) ∈ Tp Σ and consider the function
ζ(w, κλ ) = II(w) − κλ I(w);
∂ζ ∂ζ
it is easy to check that ζ, and take zero value for w = λ0 , with λ0 the eigenvector
∂wu ∂wv
of the principal direction relative to κλ , which gives the system of equations
II(λ0 ) − κλ I(λ0 ) = 0,



 ∂II(λ0 ) ∂II(λ0 )


− κλ = 0,
 ∂wu ∂wu
 ∂II(λ0 ) − κλ ∂II(λ0 ) = 0.



∂wv ∂wv
Developing the derivatives and making some standard passages, Eq. (7.22) is found again,
which proves that λ is necessarily the principal direction relative to κλ .

This theorems hence states that the derivative of N along a given direction is a vector
parallel to such a direction only when this is a principal direction of curvature.

133
7.8 Classification of the points of a surface
Let f : Ω → Σ be a regular surface and p ∈ Σ a non-planar point. Then, we say that
• p is an elliptic point if K(p) > 0;
• p is a hyperbolic point if K(p) < 0;
• p is a parabolic point if K(p) = 0.
We remark that, by Eq. (7.17), because det g > 0, Eq. (6.4), the value of det B is
sufficient to determine the type of a point on Σ.

Theorem 46. If p is an elliptical point of σ, then there exists a neighbourhood U ∈ Σ of


p such that all the points q ∈ U belong to the same half-space into which E is divided by
the tangent plane Tp Σ.

Proof. For the sake of simplicity and without loss of generality, we can always chose a
parameterization f (u, v) of the surface such that p = f (0, 0). Expanding f (u, v) into
a Taylor’s series around (0, 0), we get the position of a point q = f (u, v) ∈ Σ in the
nighbourhood of p (though not indicated for the sake of brevity, all the derivatives are
intended to be calculated at (0, 0)):

1
f (u, v) = f,u u + f,v v + (f,uu u2 + 2f,uv uv + f,vv v 2 ) + o(u2 + v 2 ).
2
The distance with sign d(q) of q ∈ Σ from the tangent plane Tp Σ is the projection onto
N, i.e.:

1
d(q) = (f,uu u2 + 2f,uv uv + f,vv v 2 ) · N + o(u2 + v 2 )
2
1
= (B11 u2 + 2B12 uv + B22 v 2 ) + o(u2 + v 2 ),
2
or, equivalently, once we set w = uf,u + vf,v ,

1
d(q) = II(w, w) + o(u2 + v 2 ). (7.23)
2
If p is an elliptic point, the principal curvatures have the same sign because K = κ1 κ2 >
0 ⇒ the sign of II(w, w) does not depend upon w, i.e. upon the tangent vector. As a
consequence the sign of d(q) does not change with w ⇒ ∀q ∈ U, Σ is on the same side of
the tangent plane Tp Σ.

Theorem 47. If p is a hyperbolic point of Σ, then for each neighbourhood U ∈ Σ of p,


there are points q ∈ U that are in half-spaces on the opposite sides with respect to the
tangent plane Tp Σ.

134
Figure 7.6: Elliptic (left), hyperbolic (center), and parabolic (last two on the right),
points.

Proof. The proof is identical to that of the previous theorem until Eq. (7.23); now, if p is
a hyperbolic point, the principal curvatures have opposite signs and by consequence d(q)
changes of sign at least two times in any neighbourhood U of p ⇒ there are points q ∈ U
lying in half-spaces on the opposite sides with respect to the tangent plane Tp Σ.

In a parabolic point, there are different possibilities: Σ is on one side of the space with
respect to Tp Σ, like for the case of a cylinder, or not, like, as an example, for the points
(0, v) of the surface, see Fig. 7.6,

 x = (u3 + 2) cos v,
y = (u3 + 2) sin v,
z = −u.

This is also the case for planar points: e.g., the point (0, 0, 0) is a planar point for both
the surfaces
z = x4 + y 4 , z = x3 − 3xy 2 ,
but in the first case, all the surface is on one side of the tangent plane, while it is on both
sides for the second case (the so-called monkey’s saddle), see Fig.7.7.

Figure 7.7: Two different planar points.

7.9 Developable surfaces


Let us now consider a ruled surface f : Ω → Σ as in Eq. (7.7); then

f,u = γ ′ + vλ′ , f,v = λ, f,u × f,v = γ ′ × λ + vλ′ × λ, f,uv = λ′ , f,vv = o.

135
2
Consequently, B22 = N · f,vv = 0 ⇒ det B = −B12 : The points of Σ are hyperbolic or
parabolic. Namely, the parabolic points are those with
f,u × f,v
B12 = N · f,uv = · f,uv = 0 ⇐⇒ (γ ′ × λ + vλ′ × λ) · λ′ = γ ′ × λ · λ′ = 0.
|f,u × f,v |
We remark that the ruled surfaces made of parabolic points have null Gaussian curvature
everywhere: K = 0.
Let us consider ruled surfaces having only parabolic points; then, we have the follow-
ing

Theorem 48. For a ruled surface f (u, v) = γ(u) + vλ(u), the following are equivalents:
i. γ ′ , λ, λ′ are linearly dependent;
ii. N,v = o.

Proof. Condition ii implies that N does not change along a straight line lying on the
ruled surface ⇒ f,u × f,v = γ ′ × λ + vλ′ × λ does not depend on v as well. This is possible
⇐⇒ γ ′ × λ and λ′ × λ are linearly dependent, i.e. ⇐⇒

(γ ′ × λ) × (λ′ × λ) = (λ′ × λ · γ ′ )λ − (λ′ × λ · λ)γ ′ = (λ′ × λ · γ ′ )λ = o,

i.e. when λ, λ′ and γ ′ are coplanar, which proves the thesis.

We say that a ruled surface is developable if one of the conditions of Theorem 48 is satisfied.
A developable surface is a surface that can be flattened without distortion onto a plane,
i.e. it can be bent without stretching or shearing or, vice-versa, it can be obtained by
transforming a plane. We remark that only ruled surfaces are developable (but not all
the ruled surfaces are developable).
It is immediate to check that a cylinder or a cone are developable surfaces, while the
helicoid, the hyperbolic hyperboloid or the hyperbolic paraboloid are not. Another clas-
sical example of developable surface is the ruled surface of the tangents to a curve: Let
γ(t) : G ⊂ R → E be a regular smooth curve; then the ruled surface of the tangents to γ
is the surface f (u, v) : G × R → Σ defined by

f (u, v) = γ(u) + vγ ′ (u).

Fig. 7.8 shows the ruled surface of the tangents to a cylindrical helix.

Figure 7.8: The ruled surface of the tangents to a cylindrical helix.

136
7.10 Points of a surface of revolution
Let us now consider a surface of revolution f : Ω → Σγ as in Eq. (7.5) and, for the sake
of simplicity, let u be the natural parameter of the curve γ(u) in Eq. (7.4) generating the
surface. Then

φ′2 (u) + ψ ′2 (u) = 1, ψ ′′ (u)φ′ (u) − ψ ′ (u)φ′′ (u) = c(u).

We can then calculate:


• the vectors of the natural basis:
 ′   
 φ (u) cos v   −φ(u) sin v 
f,u = φ′ (u) sin v , f,v = φ(u) cos v ;

ψ (u) 0
   

• the normal to the surface


 
 −ψ ′ (u) cos v 
N= −ψ ′ (u) sin v ;

φ (u)
 

• the metric tensor (i.e. the first fundamental form):


 
1 0
g= ;
0 φ2 (u)

• the second derivatives of f :


 ′′     
 φ (u) cos v   −φ′ (u) sin v   −φ(u) cos v 
f,uu = φ′′ (u) sin v , f,uv = φ′ (u) cos v , f,vv = −φ(u) sin v ;
ψ ′′ (u) 0 0
     

• tensor B (i.e. the second fundamental form):


 
c(u) 0
B= ;
0 φ(u)ψ ′ (u)

• the Gaussian curvature K:


det B c(u)ψ ′ (u)
K = det X = = .
det g φ(u)

Therefore, points of Σγ where c(u) and ψ ′ (u) have the same sign are elliptic, but hyperbolic
otherwise9 . Parabolic points correspond to inflexion points of γ(u) if c(u) = 0, or to points
with horizontal tangent to γ(u) if ψ ′ (u) = 0.
9
Recall that in a revolution surface, φ(u) > 0 ∀u.

137
As an example, let us consider the case of the pseudo-sphere, Eq. (7.6). Then,
u
φ(u) = sin u, ψ(u) = cos u + ln tan .
2
Some simple calculations give (u is not the arc-length
p of γ(u), which actually is a tractrix,

see Exercise 9, Chapter 4; hence |γ (u)| = φ (u) + ψ ′2 (u) ̸= 1)
′2

1
ψ ′ (u) = − sin u + , c(u) = −| tan u|;
sin u
as a consequence

c(u)ψ ′ (u) (− sin u + sin1 u )| tan u|


K= p =− = −1.
φ(u) φ′2 (u) + ψ ′2 (u) sin u| cot u|

Finally, K = const. = −1, which is the reason for the name of this surface.

7.11 Lines of curvature, conjugated directions, asymp-


totic directions
A line of curvature is a curve on a surface with the property of being tangent, at each
point, to a principal direction.

Theorem 49. The lines of curvature of a surface are the solutions to the differential
equation
X21 u′2 + (X22 − X11 )u′ v ′ − X12 v ′2 = 0.

Proof. A curve γ(t) : G ⊂ R → Σ ⊂ E is a line of curvature ⇐⇒

γ ′ (t) = f,u u′ + f,v v ′

is an eigenvector of X(t) ∀t, i.e. ⇐⇒ there exists a function µ(t) such that

X(t)γ ′ (t) = µ(t)γ ′ (t) ∀t.

In the natural basis of Tp Σ, this condition reads as (we omit the dependence upon t for
the sake of conciseness)
  ′   ′ 
X11 X12 u u
=µ ,
X21 X22 v′ v′

which is satisfied ⇐⇒ the two vectors at the left- and right-hand sides are proportional,
i.e. ⇐⇒
X11 u′ + X12 v ′ u′
 
det = 0 → X21 u′2 + (X22 − X11 )u′ v ′ − X12 v ′2 = 0.
X21 u′ + X22 v ′ v ′

138
As a corollary, if X is diagonal, then the coordinate lines are at the same time, principal
directions and lines of curvature.

Theorem 50. A curve γ(u) : G ⊂ R → Σ is a line of curvature ⇐⇒ the surface

f (u, v) = γ(u) + vN(γ(u)), (7.24)

is developable.

Proof. From Theorem 48, f (u, v) is developable ⇐⇒ γ ′ · N × N′ = 0. Because γ ′ and


N′ ∈ Tp Σ, which is orthogonal to N, the surface will be developable ⇐⇒ γ ′ × N′ = o.
Moreover, writing
γ ′ = f,u u′ + f,v v ′
it is
N′ = N,u u′ + N,v v ′ = −LW (γ ′ ),
hence f (u, v) is developable ⇐⇒ LW (γ ′ ) × γ ′ = o, i.e. when γ ′ is a principal direction.

The curve in Eq. (7.24) is called the ruled surface of the normals.
Let p be a non-planar point of a surface f : Ω → Σ and v1 , v2 two vectors of Tp Σ. We
say that v1 and v2 are conjugated if II(v1 , v2 ) = 0. The directions corresponding to v1
and v2 are called conjugated directions. Hence, the principal directions at a point p are
conjugated; if p is an umbilical point, any two orthogonal directions are conjugated.
The direction of a vector v ∈ Tp Σ is said to be asymptotic if it is autoconjugated, i.e. if
II(v, v) = 0. An asymptotic direction is hence a direction where the normal curvature is
null. In a hyperbolic point, there are two asymptotic directions, in a parabolic point only
one and in an elliptic point, there are not asymptotic directions. An asymptotic line is
a curve on a surface with the property of being tangent at every point to an asymptotic
direction. The asymptotic lines are the solution of the differential equation

II(γ ′ , γ ′ ) = 0 → B11 u′2 + 2B12 u′ v ′ + B22 v ′2 = 0;

in particular, if B11 = B22 = 0 and B ̸= O, then the coordinate lines are asymptotic lines.
Asymptotic lines exist only in the regions where K ≤ 0.

7.12 The Dupin’s conical curves


The conical curves of Dupin are the real curves in Tp Σ whose equations are

II(v, v) = ±1, v ∈ S.

Let {u1 , u2 } be the basis of the principal directions. Using polar coordinates, we can
write
v = ρeρ , eρ = cos θu1 + sin θu2 .

139
Therefore,
II(v, v) = ρ2 II(eρ , eρ ) = ρ2 κN (eρ ),
and the conicals’ equations are

ρ2 (κ1 cos2 θ + κ2 sin2 θ) = ±1.

With the Cartesian coordinates ξ = ρ cos θ, η = ρ sin θ, we get

κ1 ξ 2 + κ2 η 2 = ±1.

The type of conical curves depend upon the kind of point on Σ:


• Elliptical points: The principal curvatures have the same sign → one of the conical
curves is an ellipse, the other one the null set (actually, it is not a real curve).
• Hyperbolic points: The principal curvatures have opposite signs → the conical curves
are conjugated hyperbolae whose asymptotes coincide with the asymptotic direc-
tions.
• Parabolic points: At least one of the principal curvatures is null → one of the
conical curves degenerates into a couple of parallel straight lines, corresponding to
the asymptotic direction, the other one is the null set.
The three possible cases are depicted in Fig. 7.9

Figure 7.9: The conical curves of Dupin; from the left: elliptic, hyperbolic and parabolic
points.

7.13 The Gauss-Weingarten equations


Let f : Ω → Σ be a surface; for any point p ∈ Σ, consider the basis {f,u , f,v , N}, also called
the Gauss’ basis. It is the equivalent of the Frenet-Serret basis for the surfaces. We want
to calculate the derivatives of the vectors of this basis, i.e. we want to obtain, for the
surfaces, something equivalent to the Frenet-Serret equations.
N ∈ S and N · f,u = N · f,v = 0, but, in general, f,u , f,v ∈
/ S and f,u · f,v ̸= 0. In other
words, we are dealing with a case of non-orthogonal (curvilinear) coordinates. So, if w is
the coordinate along the normal N, let us call, for the sake of convenience,

u = z1, v = z2,

140
while, for the vectors,
f,u = f,1 = g1 , f,v = f,2 = g2 ,
with g1 , g2 exactly the g-vectors of the coordinate lines on Σ. Then (no summation on i
in the following equations),
∂gi 1 ∂(gi · gi ) 1 ∂gii
j
· gi = j
= ,
∂z 2 ∂z 2 ∂z j i, j = 1, 2;
∂gi ∂(gi · gj ) ∂gi ∂gij 1 ∂gii
i
· gj = i
− j · gi = i
− ;
∂z ∂z ∂z ∂z 2 ∂z j
for the last equation we have used the identity
∂gj ∂gi
i
= f,ji = f,ij = j , i, j = 1, 2.
∂z ∂z
Using Eq. (6.26), it can be proved that it is also10
∂gi
· gh = Γhij i, j, h = 1, 2.
∂z j
Moreover, by Eq. (7.14),
∂gi
· N = f,ij · N = Bij i, j = 1, 2,
∂z j
and by Eqs. (7.10), (7.11),
∂N
· gj = −LW (gi ) · gj = −Xgi · gj = −Xji , i, j = 1, 2,
∂z i
while, because N ∈ S, then from Eq. (4.1),
∂N
· N = 0 ∀i = 1, 2.
∂z i
Finally, the decomposition of the derivatives of the vectors of the basis {f,u , f,v , N} onto
these same vectors gives the equations
∂gi
= Γhij gh + Bij N,
∂z j i, j = 1, 2; (7.25)
∂N
= −Xij gi ,
∂z j
these are the Gauss-Weingarten equations (the first one is due to Gauss and the second
to Weingarten).
Moreover, if we make the scalar product of the Gauss equations by g1 and g2 , i.e.
∂gi
gk · j
= gk · (Γhij gh + Bij N), i, j, k = 1, 2,
∂z
10
The proof is rather cumbersome and it is omitted here; in many texts on differential geometry, the
Christoffel symbols are just introduced in this way, as the projection of the derivatives of vectors gi s onto
the same vectors, i.e. as the coefficients of the Gauss equations.

141
we get the following three systems of equations:
1 ∂g11

Γ111 g11 + Γ211 g21 =
 ,
2 ∂z 1 (7.26)
Γ1 g12 + Γ2 g22 = ∂g12 − 1 ∂g11 ;

11 11
∂z 1 2 ∂z 2

1 ∂g11

Γ112 g11 + Γ212 g21 =
 ,
2 ∂z 2 (7.27)
Γ1 g12 + Γ2 g22 = 1 ∂g22 ;

12 12
2 ∂z 1
∂g 1 ∂g22

Γ122 g11 + Γ222 g21 = 12

2
− ,
∂z 2 ∂z 1 (7.28)
Γ1 g12 + Γ2 g22 = 1 ∂g22 .

22 22
2 ∂z 2
The determinant of each one of these systems is simply det g ̸= 0 → it is possible to
express the Christoffel symbols as functions of the gij s and of their derivatives, i.e. as
functions of the first fundamental form (hence, of the metric tensor).

7.14 The Theorema Egregium


The following theorem is a fundamental result due to Gauss:

Theorem 51. (Theorema Egregium). The Gaussian curvature K of a surface f (u, v) :


Ω → Σ depends only upon the first fundamental form of f .

Proof. Let us write the identity

∂ 2 g1 ∂ 2 g1
=
∂z 1 ∂z 2 ∂z 2 ∂z 1
using the Gauss equations (7.25)1 :

Γ111 g1,2 + Γ211 g2,2 + B11 N,2 + Γ111,2 g1 + Γ211,2 g2 + B11,2 N =


Γ112 g1,1 + Γ212 g2,1 + B12 N,1 + Γ112,1 g1 + Γ212,1 g2 + B12,1 N,

∂(·)
where, for the sake of shortness, we have abridged by (·),j . Then, we use again
∂z j
Eqs. (7.25) to express g1,1 , g1,2 , g2,2 , N,1 and N,2 ; after doing that and equating to 0 the
coefficient of g2 , we get

B11 X22 − B12 X21 = Γ111 Γ212 + Γ211 Γ222 + Γ211,2 − Γ112 Γ211 − Γ212 Γ212 − Γ212,1 ;

from Eq. (7.13), we get that

B11 = g11 X11 + g12 X21 , B12 = g11 X12 + g12 X22 ,

142
which, injected into the previous equation, gives

g11 det X = Γ111 Γ212 + Γ211 Γ222 + Γ211,2 − Γ112 Γ211 − Γ212 Γ212 − Γ212,1 . (7.29)

Equating to zero the coefficient of g1 , a similar expression can also be get for g12 . Because g
is positive definite, it is not possible that g11 = g12 = 0. So, remembering that K = det X
and the result of the previous section, we see that it is possible to express K through the
coefficients of the first fundamental form and of its derivatives.

7.15 Minimal surfaces


A minimal surface is a surface f : Ω → Σ having the mean curvature H = 0 ∀p ∈ Σ.
Typical minimal surfaces are the catenoid and the helicoid11 . Other minimal surfaces are
the Enneper’s surface
u3

x = u − + uv 2 ,

 1

33

v
x2 = v − + u2 v,
3



x3 = u2 − v 2 ,

the Costa’s and the Schwarz’s surfaces, Fig. 7.10.

Figure 7.10: From the left, the minimal surfaces of Enneper, Costa and Schwarz.

Theorem 52. The non-planar points of a minimal surface are hyperbolic.

Proof. This is a direct consequence of the definition of mean curvature H and of hyperbolic
points: H = 0 ⇐⇒ κ1 κ2 < 0.

Let f : Ω → Σ be a regular surface and Q a subset of Ω with its boundary ∂Q a closed


regular curve in Ω; then, R = f (Q) ⊂ Σ is a simple region of Σ. Let h : Q → R be
a smooth function. Then, we call normal variation of R the map φ : Q × (−ϵ, ϵ) → E
defined by
φ(u, v, t) = f (u, v) + t h(u, v)N(u, v).
11
Minimal surfaces have some interesting applications in the mechanics of tensile structures composed
of prestressed membranes. Also, it can be shown that a soap film, when not bounding a closed region,
takes the form of a minimal surface.

143
For each fixed t, φ(u, v, t) is a surface with

φ,u (u, v, t) = f,u (u, v) + t h(u, v)N,u (u, v) + t h,u (u, v)N(u, v),
φ,v (u, v, t) = f,v (u, v) + t h(u, v)N,v (u, v) + t h,v (u, v)N(u, v).

If the first fundamental form of f is represented by the metric tensor g, we look for the
metric tensor gt representing the first fundamental form of φ(u, v, t) ∀t:
t
g11 = φ,u · φ,u = g11 + 2t h f,u · N,u + t2 (h2 N2,u + h2,u ),
t
g12 = φ,u · φ,v = g12 + t h(f,u · N,v + f,v · N,u ) + t2 (h2 N,u · N,v + h,u h,v ),
t
g22 = φ,v · φ,v = g22 + 2t h f,v · N,v + t2 (h2 N2,v + h2,v ),

and by Eq. (7.14),


t
g11 = g11 − 2t h B11 + t2 (h2 N2,u + h2,u ),
t
g12 = g12 − 2t h B12 + t2 (h2 N,u · N,v + h,u h,v ),
t
g22 = g22 − 2t h B22 + t2 (h2 N2,v + h2,v ),

whence
det gt = det g − 2th(g11 B22 − 2g12 B12 + g22 B11 ) + o(t2 ).
Then, by Eq. (7.15), we get easily that

g11 B22 − 2g12 B12 + g22 B11 = 2H det g,

so that
det gt = det g(1 − 4thH) + o(t2 ).
We can now calculate the area A(t)of the simple region Rt = φ(u, v, t) corresponding to
the subset Q: Z p
t
A = det g(1 − 4thH) + o(t2 )dudv;
Q

For ϵ ≪ 1 At is differentiable and its derivative for t = 0 is


 t Z
dA p
= − 2hH det gdudv.
dt t=0 Q

dAt
 
Theorem 53. A surface f : Ω → Σ is minimal ⇐⇒ = 0 ∀R ⊂ Σ and for
dt t=0
each normal variation.

Proof. If f is minimal, the condition is clearly satisfied (H = 0). Conversely, let us


suppose that ∃p = f (u, v) ∈ Σ|H(p) ̸= 0. Consider r1 , r2 ∈ R such that |H| =
̸ 0 in the
1
circle D2 with center p and radius r2 and |H| > 2 |H(p)| in the circle D1 with center p
and radius r1 . Then, we chose a smooth function h(u, v) such that i) h = H inside D1 ,

144
ii) hH > 0 inside D2 and iii) h = 0 outside D2 . For the normal variation defined by such
h(u, v) we have
 t Z Z
dA p p
− = 2hH det gdu dv ≥ 2H 2 det gdu dv
dt t=0 D D1
Z 2 2p
H(p) H(p)2
≥ det gdu dv = A(f (D1 ))
D1 2 2
 t
dA
⇒ <0
dt t=0

which contradicts the hypothesis.

The meaning of this theorem justifies the name of minimal surfaces: These are the surfaces
that have the minimal area among all the surfaces that share the same boundary.

7.16 Geodesics
Let f : Ω → Σ be a surface and γ(t) : G ⊂ R → Σ a curve on Σ. A vector function
w(t) : G → Tγ(t) Σ is called a vector field12 along γ(t). We call covariant derivative of
w(t) along γ(t) the vector field Dγ w(t) : G → V defined as13

dw
Dγ w := (I − N ⊗ N) ,
dt
i.e. the projection of the derivative of w onto Tγ(t) Σ. It is always possible to decompose
w(t) into its components in the natural basis {f,u , f,v }:

w(t) = w1 (t)f,u (γ(t)) + w2 (t)f,v (γ(t)).

Differentiating, we get (a prime here denotes the derivative with respect to t)

w′ = w1′ f,u + w1 (f,uu u′ + f,uv v ′ ) + w2′ f,v + w2 (f,uv u′ + f,vv v ′ )

and using the Gauss equations, Eq. (7.25)1 , we obtain (summation on the dummy indexes
and u1 stands for u while u2 for v)

w′ = f,k wk′ + (Γkij f,k + Bij N)wi u′j , i, j, k = 1, 2,

so that the projection onto Tγ(t) Σ, i.e. Dγ w(t), is

Dγ w = (wk′ + Γkij wi u′j )f,k . (7.30)

A parallel vector field w along γ is a vector field having Dγ w = o ∀t. A regular curve γ
is a geodesic of Σ if the vector field γ ′ of the vectors tangent to γ is parallel along γ.
12
More correctly, w(t) is a curve of vectors; however, it is normally called a vector field along a curve.
13
The operator that gives the projection of w onto a vector orthogonal to N ∈ S, i.e. onto Tγ(t) Σ, is
I − N ⊗ N, cf. Exercise 2, Chapter 2.

145
Theorem 54. A curve γ is a geodesic of Σ ⇐⇒ ν × N = o.

Proof. If γ is a geodesic, then the derivative of its tangent γ ′ has a component only along
N, i.e. γ ′′ × N = o ⇒ γ ′ · γ ′′ = 0. The principal normal to γ, ν, is orthogonal to
γ ′ ⇒ ν × N = o. Vice versa, if ν × N = o, then γ ′′ is orthogonal to γ ′ ⇒ Dγ γ ′ = o, i.e.
γ is a geodesic.

Theorem 55. If γ is a geodesic, then |γ ′ | = const.

d(γ ′ · γ ′ )
Proof. In a geodesic, γ ′ · γ ′′ = 0 ⇒ = 0 ⇒ |γ ′ | = const.
dt
This result shows that, in a geodesic, the parameter is always the natural parameter
s.
Let γ(s) be a curve on Σ parameterized by the the arc-length s. We call geodesic curvature
of γ(s) the function
κg := Dγ τ · (N × τ ),
where τ = γ ′ ∈ S is the tangent vector to γ. Because N × τ ∈ S lies in Tγ Σ, the com-
ponent of τ ′ orthogonal to Tγ Σ gives a null contribution to κg , so we can also write

κg = τ ′ · (N × τ ).

Theorem 56. A regular curve γ(s) is a geodesic ⇐⇒ κg = 0 ∀s.

Proof. If γ is a geodesic, clearly κg = 0. Vice versa, if κg = 0, then τ , τ ′ and N are


linearly dependent, i.e. coplanar. Because τ ′ · τ = N · τ = 0 ⇒ τ ′ × N = o ⇒ by
Theorem 54, γ is a geodesic.

Let us now write Eq. (7.30) in the particular case of w = γ ′ , i.e. w1 = u′ , w2 = v ′ :

Dγ w = (u′′k + Γkij u′i u′j )f,k ;

therefore, the geodesics are the solutions to the system of differential equations
(
u′′ + Γ111 u′2 + 2Γ112 u′ v ′ + Γ122 v ′2 = 0,
(7.31)
v ′′ + Γ211 u′2 + 2Γ212 u′ v ′ + Γ222 v ′2 = 0.

It can be shown that ∀p ∈ Σ and ∀w(p) ∈ Tp Σ the geodesic is unique.


Let p be a point of a regular surface f : Ω → Σ and α(v) : G ⊂ R → Σ a smooth
regular curve on Σ, with v being the natural parameter and such that p = α(0). Consider
the geodesic γ v passing through q = α(v) and such that γ ′v (0) = N(α(v)) × τ (v), with
τ (v) the (unit) tangent vector to α(v). Consider the map f (u, v) : Ω → Σ defined by

146
posing f (u, v) = γ v (u); this is a surface whose coordinates (u, v) are called semigeodesic
coordinates.
Let us see the form that the first fundamental form (i.e. the metric tensor g), the
Christoffel symbols and the Gaussian curvature take in semigeodesic coordinates. Curves
f (u, v0 ) = γ v0 (u) are geodesics, and u is hence their natural parameter. Therefore,
f,u ∈ S ⇒ g11 = 1. Then, f,uu (u, v0 ) is the derivative of the tangent vector to a geodesic
f (u, v0 ) = γ v0 (u) ⇒ f,uu (u, v0 ) does not have a component along the tangent, hence,
Eq. (7.26)1 ⇒ Γ111 = Γ211 = 0. Then, by Eq. (7.26)2 , we get g12,u = 0 ⇒ g12 does
not depend upon u ⇒ g12 (u, v) = g12 (0, v) ∀u. Moreover, let θ be the angle between
the curve α, i.e. between the coordinate line f (0, v), whose tangent vector is f,v (0, v),
π
and the geodesic γ v (u), whose tangent vector at (0, v) is γ ′v (u). Then, θ = because

2
γ v (0) = N(α(v)) × τ (v). As a consequence, g12 (0, v) = 0 ⇒ g12 (u, v) = 0 ∀(u, v) ∈ Ω.
Finally, setting g22 = g,  
1 0
g= ,
0 g
with g > 0 because g is positive definite. Through systems (7.26)−(7.28), we obtain
g,u g,u g,v
Γ112 = 0, Γ212 = , Γ122 = − , Γ222 = ,
2g 2 2g

and using Eq. (7.29), we get


2
g,uu g,u
K = det X = − + 2.
2g 4g

Given two points p1 , p2 ∈ Σ, we define the distance d(p1 , p2 ) as the infimum of the lengths
of the curves on Σ joining the two points. We end with an important characterization of
geodesics:

Theorem 57. Geodesics are the curves of minimal distance between two points of a
surface.

Proof. Let γ : G ⊂ R → Σ be a geodesic on Σ, parameterized with the arc-length, and α


a smooth regular curve through p and orthogonal to γ. Through α, we set up a system of
semigeodesic coordinates in a neighbourhood U of p. With an opportune parameterization
α(t), in such coordinates we can get p = f (0, 0) and γ described by the equation v = 0.
Let q ∈ U be a point in γ, and consider a regular curve connecting p with q. The length
ℓ(p, q) of such a curve is
Z qp Z q
ℓ(p, q) = u′2 +g v ′2 dt ≥ u′ dt = |uq − up |.
p p

Observing that p = (up , 0), q = (uq , 0), we remark that |uq − up | is exactly the length of
γ between p and q, because γ is parameterized with its arc-length.

147
There is another, direct and beautiful way to show that geodesics are the shortest path
lines: the use of the methods of the calculus of variations14 . The length ℓ(p, q) of a curve
γ(t) ∈ Σ between two points p and q is given by the functional (7.9); it depends upon the
first fundamental form, i.e. upon the metric tensor g on Σ. For the sake of conciseness,
let w = (w,t1 , w,t2 ) be the tangent vector to the curve γ(w1 , w2 ) ∈ Σ. Then,
Z qp Z q

ℓ(p, q) = I(w)dt = w · gwdt.
p p

The curve γ(t) that minimizes ℓ(p, q) is the solution to the Euler-Lagrange equations
d ∂F ∂F d ∂F ∂F
− =o → k
− = 0, k = 1, 2,
dt ∂w,t ∂w dt ∂w,t ∂wk

where

q
F (w, w,t , t) = w · gw = gij w,ti w,tj .
It is more direct, and equivalent, to minimize J 2 (t), i.e. to write the Euler-Lagrange
equations for
Φ(w, w,t , t) := F 2 (w, w,t , t) = gij w,ti w,tj .
Therefore:
∂Φ
k
= 2gjk w,tj ,
∂w,t
∂Φ ∂ghj h j
= w w , j, h, k = 1, 2.
∂w k ∂wk ,t ,t   
d ∂Φ j dgjk j j ∂gjk l j
= 2 gjk w,tt + w = 2 gjk w,tt + w w ,
dt ∂w,tk dt ,t ∂wl ,t ,t

The Euler-Lagrange equations are hence

j ∂gjk h j 1 ∂ghj h j
gjk w,tt + w,t w,t − w w = 0, , j, h, k = 1, 2,
∂w h 2 ∂wk ,t ,t
which can be rewritten as
 
j 1 ∂gjk ∂ghk ∂ghj
gjk w,tt + h
+ j
− k
w,th w,tj = 0, j, h, k = 1, 2.
2 ∂w ∂w ∂w
14
The reader is addressed to texts on the calculus of variations for an insight into this matter, cf. the
suggested texts. Here, we just recall the fundamental fact to be used in the proof concerning geodesics:
Let Z b
J(t) = F (x, x′ , t)dt
a

be a functional to be minimized by a proper choice of the function x(t) (in the case of the geodesics,
J = ℓ(p, q)); then, such a minimizing function can be found as a solution to the Euler-Lagrange equations
d ∂F ∂F

− = o.
dt ∂x ∂x

148
Multiplying by g lk , we get
 
j 1 ∂gjk ∂ghk ∂ghj
g lk
gjk w,tt + g lk + − w,th w,tj = 0, j, h, k, l = 1, 2.
2 ∂wh ∂wj ∂wk

and finally, because


g lk gjk = δlj
and by Eq. (6.26), we get
l
w,tt + Γljh w,tj w,th = 0, j, h, l = 1, 2.

These are the differential equations whose solution is the curve of minimal length between
two points of Σ; comparing these equations with those of a geodesic of Σ, Eq. (7.31), we
see that they are the same: The geodesics of a surface are hence the curves of minimal
distance on the surface.
The Christoffel symbols of a plane are all null; as a consequence, the geodesic lines of a
plane are straight lines. In fact, only such lines have a constant derivative.
Through systems (7.26)−(7.28), we can calculate the Christoffel symbols for a revolution
surface, Eq. (7.5), which are all null excepted

φ′
Γ212 = , Γ122 = −φ φ′ ,
φ

so the system of differential equations (7.31) becomes



 u′′ − φ φ′ v ′2 = 0,
φ′ (7.32)
 v ′′ + 2 u′ v ′ = 0.
φ

It is direct to check that the meridians (u = t, v = v0 ) are geodesic lines, while the parallels
(u = u0 , v = t) are geodesics ⇐⇒ φ′ (u0 ) = 0.

7.17 The Gauss-Codazzi compatibility conditions


Let us consider a surface Σ whose points are determined by the vector function r : Ω ⊂
R2 → Σ ⊂ E, r(α1 , α2 ) = xi (α1 , α2 )ϵi , with ϵi , i = 1, 2, 3, the vectors of the orthonormal
basis of the reference frame R = {o; ϵ1 , ϵ2 , ϵ3 } and the parameters α1 , α2 chosen in such
a way that the lines α1 = const., α2 = const. are lines of curvature, i.e. tangent at each
point to the principal directions of curvature and hence mutually orthogonal15 . With such
a choice, cf. Eq. (7.8),
ds2 = A21 dα12 + A22 dα22 ,
15
Here, the symbol r is preferred to f , like α1 to u and α2 to v, to recall that we have made the
particular choice of coordinate lines that are lines of curvature. All the developments could be done in
a more general case, but this choice is made to obtain simpler relations, which preserves anyway the
generality.

149
with
r
q dxi dxi
A1 = r2,α1 = ,
dα1 dα1
r
q
2
dxi dxi
A2 = r,α2 =
dα2 dα2
the so-called Lamé’s parameters. We remark that along the lines of curvatures, i.e. the
lines αi = const., i = 1, 2, which in short, from now, on we call the lines αi , it is

ds1 = A1 dα1 ,
ds2 = A2 dα2 ,

and hence,
ds1
λ1 = = A1 e1 ,
dα1
(7.33)
ds2
λ2 = = A 2 e2
dα2
are the vectors tangent to the lines of curvature. Let
1 1
e1 = r,α1 , e2 = r,α , e3 = e1 × e2 (= N); (7.34)
A1 A2 2

these three vectors form the orthonormal (local) natural basis e = {e1 , e2 , e3 }. We always
make the choice of α1 , α2 such that e3 is always directed toward the convex side of Σ if
the point is elliptic or parabolic or toward the side of the centres of negative curvature, if
the point is hyperbolic.
We consider a vector v = v(p), p ∈ Σ,

v = v1 e1 + v2 e2 + v3 e3 ,

and we want to calculate how it transforms when p changes. To this end, we need to
calculate how e1 , e2 , e3 change with α1 , α2 . Let q ∈ Σ be a point in the neighborhood of
p on the line αi and let us first consider the change of e3 in passing from p to q. Because
p and q belong to the same line αi , by the theorem of Rodrigues, we get (no summation
on i in the following equations)

∂e3
= −κi λi , i = 1, 2,
∂λi

i.e., by Eq. (7.33),


∂e3 Ai
= ei ,
∂αi Ri
with
1
Ri = −
κi

150
de3

e3(p) e3(q)

p q lin
e 𝛼i
𝞶 Ri

Figure 7.11: Variation of N = e3 along a line of curvature.

the (principal) radius of curvature along the line αi . The minus sign in the previous
equation is due to the choice made above for orienting e3 = N, which gives always
N = −ν, with ν the principal normal to the line αi . This result can also be obtained
directly, see Fig. 7.11:
e3 (q) = e3 (p) + de3
and in the limit of q → p, de3 tends to be parallel to q − p and

lim(q − p) = λi = Ai ei .
q→p

By the similitude of the triangles, it is evident that


|de3 | |q − p|
= ;
|e3 | Ri
moreover,
∂e3
de3 = dαi ei .
∂αi
Finally, as |e3 | = 1, we get again
∂e3 Ai
= ei . (7.35)
∂αi Ri
Implicitly, in this last proof, we have used the theorem of Rodrigues because we have
assumed that de3 is parallel to λi , as it is, because line αi is a line of curvature.
We now move on to determine the changes in e1 and e2 ; for this purpose, we remark
that
∂r,α1 ∂ 2r ∂ 2r ∂r,α2
= = = ,
∂α2 ∂α2 ∂α1 ∂α1 ∂α2 ∂α1
so by Eq. (7.34) we get
∂(A1 e1 ) ∂(A2 e2 )
= . (7.36)
∂α2 ∂α1
∂ej
Let us study now ; as |ej | = 1, j = 1, 2,
∂αi
∂ej
· ej = 0 ∀i, j = 1, 2. (7.37)
∂αi

151
Because e1 · e2 = 0,
∂e1 ∂(e1 · e2 ) ∂e2 ∂e2
· e2 = − e1 · = −e1 · .
∂α1 ∂α1 ∂α1 ∂α1
By Eq. (7.36), we get
∂e2 1 ∂(A1 e1 ) 1 ∂A2
= − e2 ,
∂α1 A2 ∂α2 A2 ∂α1
which when inserted into the previous equation gives, by Eq. (7.37),
∂e1 1 ∂(A1 e1 ) 1 ∂A2 A1 ∂e1 1 ∂A1 1 ∂A1
· e2 = − · e1 + e2 · e1 = − · e1 − e1 · e1 = − .
∂α1 A2 ∂α2 A2 ∂α1 A2 ∂α2 A2 ∂α2 A2 ∂α2
Then, because e1 · e3 = 0,
∂e1 ∂(e1 · e3 ) ∂e3 ∂e3
· e3 = − e1 · = −e1 · ,
∂α1 ∂α1 ∂α1 ∂α1
and by Eq. (7.35),
∂e3 A1
= e1 ,
∂α1 R1
so finally,
∂e1 A1
· e3 = − .
∂α1 R1
Again, through Eqs. (7.36) and (7.37), we get
∂e1 1 ∂(A2 e2 ) 1 ∂A1 A2 ∂e2 1 ∂A2 1 ∂A2
· e2 = · e2 − e1 · e2 = · e2 + e2 · e2 =
∂α2 A1 ∂α1 A1 ∂α2 A1 ∂α1 A1 ∂α1 A1 ∂α1
and also, by Eq. (7.35),
∂e1 ∂(e1 · e3 ) ∂e3 ∂e3 A2
· e3 = − e1 · = −e1 · = − e1 · e2 = 0.
∂α2 ∂α2 ∂α2 ∂α2 R2
The derivatives of e2 can be found in a similar way, and resuming, we have
∂e1 1 ∂A1 A1
= − e2 − e3 ,
∂α1 A2 ∂α2 R1
∂e1 1 ∂A2
= e2 ,
∂α2 A1 ∂α1
∂e2 1 ∂A1
= e1 ,
∂α1 A2 ∂α2
(7.38)
∂e2 1 ∂A2 A2
= − e1 − e3 ,
∂α2 A1 ∂α1 R2
∂e3 A1
= e1 ,
∂α1 R1
∂e3 A2
= e2 .
∂α2 R2

152
Passing now to the second-order derivatives, imposing the equality of mixed derivatives,
gives some important differential relations between the Lamé’s parameters Ai and the
radii of curvature Ri . In fact, from the identity

∂ 2 e3 ∂ 2 e3
= ,
∂α1 ∂α2 ∂α2 ∂α1

and Eqs. (7.38)5,6 , we get


   
∂ A1 ∂ A2
e1 = e2 ,
∂α2 R1 ∂α1 R2

whence    
∂ A1 A1 ∂e1 ∂ A2 A2 ∂e2
e1 + = e2 + .
∂α2 R1 R1 ∂α2 ∂α1 R2 R2 ∂α1
Inserting now Eqs. (7.38)2,3 into the last result and rearranging the terms gives
       
∂ A1 1 ∂A1 ∂ A2 1 ∂A2
− e1 − − e2 = 0,
∂α2 R1 R2 ∂α2 ∂α1 R2 R1 ∂α1

that to be true needs that the two following conditions be identically satisfied:
 
∂ A1 1 ∂A1
− = 0,
∂α2 R1 R2 ∂α2
  (7.39)
∂ A2 1 ∂A2
− = 0.
∂α1 R2 R1 ∂α1

The above equations are known as the Codazzi conditions. Let us now consider the other
identity
∂ 2 e1 ∂ 2 e1
= ;
∂α1 ∂α2 ∂α2 ∂α1
again using Eq. (7.38), with some standard operations, this identity can be transformed
to
         
∂ 1 ∂A2 ∂ 1 ∂A1 A1 A2 ∂ A1 1 ∂A1
+ + e2 + − e3 = 0.
∂α1 A1 ∂α1 ∂α2 A2 ∂α2 R1 R2 ∂α2 R1 R2 ∂α2

Also in this case, for this equation to be identically satisfied, each of the expressions in
square brackets must vanish, which gives two more differential conditions, but of which
only the first one is new, as the second one corresponds to Eq. (7.39)1 . The new condition
is hence    
∂ 1 ∂A2 ∂ 1 ∂A1 A1 A2
+ + = 0, (7.40)
∂α1 A1 ∂α1 ∂α2 A2 ∂α2 R1 R2
which is known as the Gauss condition. The last identity

∂ 2 e2 ∂ 2 e2
=
∂α1 ∂α2 ∂α2 ∂α1

153
does not add any independent condition, which can be easily checked. The meaning of the
Gauss-Codazzi conditions, Eqs. (7.39) and (7.40), is that of compatibility conditions: Only
when these conditions are satisfied by functions A1 , A2 , R1 and R2 , then such functions
represent the Lamé’s parameters and the principal radii of curvature of a surface, i.e.
only in this case they define a surface, except for its position in space. The Gauss-
Codazzi conditions are important in establishing the equations of the classical theory of
shells.

7.18 Exercises
1. Prove that a function of the type x3 = f (x1 , x2 ), with f : Ω ⊂ R2 → R smooth,
defines a surface.
2. Show that the catenoid is the rotation surface of a catenary, then find its Gaussian
curvature.
3. Show that the pseudo-sphere is the rotation surface of a tractrice and explain why
the surface has this name (hint: look for its Gaussian curvature).
4. Prove that the regularity of a cone f (u, v) = vγ(u) is satisfied at each point except
at the apex and at the points on the straight lines tangent to γ(u).

Figure 7.12: A hyperbolic hyperboloid, left, and a hyperbolic paraboloid, right.

5. Prove that the hyperbolic hyperboloid is a doubly ruled surface and determine the
angle θ formed by two straight lines belonging to the two sets of lines on the surface,
see the left panel of Fig. 7.12.
6. Prove that the hyperbolic paraboloid whose Cartesian equation is x3 = x1 x2 , right
panel of Fig. 7.12, is a doubly ruled surface and determine the angle θ formed by
two straight lines belonging to the two sets of lines on the surface. Where does
π
θ= ?
2
7. Consider the parameterization

f (u, v) = (1 − v)γ(u) + vλ(u),

154
with

γ(u) = (cos(u − α), sin(u − α), −1), λ(u) = (cos(u + α), sin(u + α), 1).

Show that:
• for α = 0, one gets a cylinder with equation x21 + x22 = 1;
π
• for α = , one gets a cone with equation x21 + x22 = x23 ;
2
π
• for 0 < α < , one gets a hyperbolic hyperboloid with equation
2
x21 + x22 x23
− = 1.
cos2 α cot2 α

8. Calculate the metric tensor of a sphere of radius R, write its first fundamental form,
determine the area of a sector of surface between the longitudes θ1 and θ2 and the
length of the parallel at the latitude π/4 between these two longitudes.
9. Prove that the surface defined by
 
cos v sin v sinh u
f (u, v) : Ω = R × (−π, π] → E| f (u, v) = , ,
cosh u cosh u cosh u

is a sphere. Then, show that the image of any straight line on Ω is a loxodromic
line on the sphere.
10. Calculate the vectors of the natural basis, the tensors g, B, X and the first and
second fundamental forms for the catenoid.
11. Calculate the same for the helicoid of parametric equation

f (u, v) = γ(u) + vλ(u),

with
γ(u) = (0, 0, u), λ(u) = (cos u, sin u, 0).

12. Show that the catenoid and the helicoid are made of hyperbolic points.
13. Determine the geodesic lines of a circular cylinder.

155
156
Suggested texts

There are many textbooks on tensors, differential geometry and calculus of variations. The
style, content, language of such books greatly depend upon the scientific community the
authors belong to: pure or applied mathematicians, physicists, theoretical mechanicians
or engineers. It is hence difficult to suggest some readings in the domain, and ultimately,
it is mostly a matter of personal taste.
This textbook is greatly inspired by some classical methods, style and language that are
typical in the community of theoretical mechanics; the following few suggested readings,
among several possible others, belong to such a kind of scientific literature. They are
classical textbooks and though the list is far from being exhaustive, they constitute a
solid basis for the topics briefly developed in this manuscript, where the objective is to
present mathematics for mechanics.
A good introduction to tensor algebra and analysis, which greatly inspired the content of
this manuscript, are the two introductory chapters of the classical textbook
• M. E. Gurtin: An introduction to continuum mechanics. Academic Press, 1981,
or also, in a similar style, the long article
• P. Podio-Guidugli: A primer in elasticity. Journal of Elasticity, v. 58: 1-104, 2000.
A short, effective introduction to tensor algebra and differential geometry of curves can
be found in the following text of exercises on analytical mechanics:
• P. Biscari, C. Poggi, E. G. Virga: Mechanics notebook. Liguori Editore, 1999.
A classical textbook on linear algebra that is recommended is
• P. R. Halmos: Finite-dimensional vector spaces. Van Nostrand Reynold, 1958.
In the previous textbooks, tensor algebra in curvilinear coordinates is not developed; an
introduction to this topic, especially intended for physicists and engineers, can be found
in
• W. H. Müller: An expedition to continuum theory. Springer, 2014,
which has largely influenced Chapter 6.
Two modern and application-oriented textbooks on differential geometry of curves and
surfaces are

157
• V. A. Toponogov: Differential geometry of curves and surfaces - A concise guide.
Birkhäuser, 2006,
• A. Pressley: Elementary differential geometry. Springer, 2010.
A short introduction to the differential geometry of surfaces, oriented toward the mechan-
ics of shells, can be found in the classical book
• V. V. Novozhilov: Thin shell theory. Noordhoff LTD., 1964.
For what concerns the calculus of variations, a still valid textbook in the matter (but not
only) is
• R. Courant, D. Hilbert: Methods of mathematical physics. Interscience Publishers,
1953.
Two very good and classical textbooks with an introduction to the calculus of variations
for engineers are
• C. Lanczos: The variational principles of mechanics. University of Toronto Press,
1949,
• H. L. Langhaar: Energy methods in applied mechanics. Wiley, 1962.

158
Solutions to the exercises

Chapter 1
1. Suppose o1 ̸= o2 ; then, apply to a point p the definition of vector null for both of
them.
2. Use v + o = v and make the scalar product with a vector w.
3. Make the norm of v + o = v and use the above result.
4. |u − v| = |u + v| ⇐⇒ |u − v|2 = |u + v|2 ⇒ (u − v) · (u − v) = (u + v) · (u + v) ⇒
u · v = 0.
5. By linearity, ∀v = vi ei , ψ(v) = vi ψ(ei ); moreover, u · v = ui vi , so setting
u = ψ(ei )ei , ψ(v) = u · v. Uniqueness: Suppose ∃u1 ̸= u2 |ψ(v) = u1 · v = u2 · v ⇒
(u1 − u2 ) · v = 0 ∀v ⇐⇒ u1 − u2 = o ⇒ u1 = u2 .
6. Let θu , θv be the angles formed by w with u and v, respectively; then, w · u =
w · v ⇒ uw cos θu = vw cos θv ⇒ cos θu = cos θv , as u, v ∈ S;
θu = θv ⇒ cos θu = cos θv ⇒ uw cos θu = vw cos θv ⇒ u · w = v · w.
7. i) Coplanar vectors: Let p be a point of the plane of the vectors ⇒ Mrp · R = 0
because R ∈ to the plane, while Mrp = (pi − p) × vpi is of course orthogonal to it.
Then, let q be a point ∈ / to the plane of the vectors → Mrq = Mrp + (p − q) × R ⇒
Mrq · R = Mrp · R + (p − q) × R · R = R × R · (p − q) = 0.
Let e ∈ S| vpi = αi e ∀i = 1, ..., n ⇒ R = ni=1 αP
P
ii) Parallel vectors: ie ⇒
∀o ∈ E, Mro = ni=1 (pi − o) × αi e ⇒ Mro · R = ( ni=1 αi (pi − o)) × e · ( ni=1 αi )e = 0.
P P

8. It is a direct consequence of the theorem of reduction of the systems of applied


vectors for the case R = o, with o being any point.
applied to p, then the system can be reduced to Rp plus
9. If R isP
Mp = ni=1 (p − p) × vip = o.
r

10. i) By the theorem of reduction, if R is applied to o, the system is reduced to Ro


and Mro ⇒ to only R if Mro = o.
ii) Because for coplanar or parallel vectors ∀o ∈ A, Mro = o, then the system is
equivalent to R applied to any point of A.

159
11. If a system is equilibrated, then, by definition, any equivalent system is equilibrated.
Conversely, if it exists another system equilibrated and equivalent, then by the
relation of equivalence also the system in object is equilibrated, and this is true for
any other equivalent system, which hence must be equilibrated.
12. Let vp , vq = −(vp )q be two opposite vectors applied to p and q, respectively ⇒ R =
vp + vq = o; moreover, Mrp = (q − p) ⊗ vq ⇒ ∀o ̸= p, Mro = Mrp + (p − o) × R =
Mrp ⇒ Mro = o ∀o ∈ E ⇐⇒ q − p = o.
13. If all the vectors pass through a point p, then the system is equivalent to Rp (Exercise
9) and if R = o, then the system is equilibrated.

Chapter 2
1. ∀u, L, Lu = L(u + o) = Lu + Lo ⇐⇒ Lo = o
2. u ∈ S → (u ⊗ u)v = v cos θu; (I − u ⊗ u)v = v − v cos θu, which is orthogonal to
u : (I − u ⊗ u)v · u = v · u − v · u u · u = o, as u ∈ S.
3. i) ∀a, b ∈ V, a · (αA)b = (αA)⊤ a · b and by the linearity of the scalar product,
a · (αA)b = αa · Ab = αA⊤ a · b ⇒ (αA)⊤ = αA⊤ .
ii) ∀a, b ∈ V, a · (A + B)b = (A + B)⊤ a · b and by the linearity of the scalar product
and of tensors, a·(A+B)b = a·Ab+a·Bb = A⊤ a·b+B⊤ a·b = (A⊤ +B⊤ )a·b ⇒
(A + B)⊤ = A⊤ + B⊤ .
iii) ∀u, v ∈ V u·(a⊗b)Av = u·(a⊗b)(Av) = (a⊗b)⊤ u·(Av) = (b⊗a)u·(Av) =
A⊤ (b⊗a)u·v = A⊤ (a·u b)·v = a·u A⊤ b·v = ((A⊤ b)⊗a)u·v = u·((A⊤ b)⊗a)⊤ v =
u · (a ⊗ A⊤ b)v.
4. By linearity and the definition of O : ∀u ∈ V, (L + O)u = Lu + Ou = Lu ⇐⇒
L + O = L.
5. i) trI = tr(δij ei ⊗ ej ) = δij tr(ei ⊗ ej ) = δij ei · ej = δij δij = δii = 3.
ii) trL = tr(L + O) = trL + trO ⇐⇒ trO = 0.
6. tr(AB) = tr((Aij ei ⊗ ej )(Bhk eh ⊗ ek )) = Aij Bhk tr((ei ⊗ ej )(eh ⊗ ek ))
= Aij Bhk ej · eh tr(ei ⊗ ek ) = Aij Bhk ej · eh ei · ek = Aij Bhk δjh δik = Aij Bji ; in
a similar way, we prove that tr(BA) = Bij Aji ; because i, j are dummy indexes,
Aij Bji = Bij Aji ⇒ tr(AB) = tr(BA).
7. i) L⊤ · M⊤ = tr((L⊤ )⊤ M⊤ ) = tr(LM⊤ ) = tr(M⊤ L) = M · L = L · M.
ii) LM · N = tr((LM)⊤ N) = tr(M⊤ L⊤ N) = M · L⊤ N;
|
= tr(N(LM)⊤ ) = tr((NM⊤ )L⊤ ) = tr(L⊤ (NM⊤ ))
= L · NM⊤ .
8. i) (a ⊗ b)(c ⊗ d) = ((a ⊗ b)(c ⊗ d))ij ei ⊗ ej = (a ⊗ b)ik (c ⊗ d)kj ei ⊗ ej
= ai bk ck dj ei ⊗ ej = b · c ai dj ei ⊗ ej = b · c a ⊗ d.

160
ii) A(a ⊗ b) = A(a ⊗ b)ij ei ⊗ ej = Aik (a ⊗ b)kj ei ⊗ ej = Aik ak bj ei ⊗ ej
= (Aa)i bj ei ⊗ ej = (Aa) ⊗ b.
9. L · v ⊗ w = tr(L⊤ (v ⊗ w)) = tr((L⊤ v) ⊗ w) = L⊤ v · w = v · Lw.
10. A = A⊤ , B = −B⊤ ⇒ A · B = A⊤ · B⊤ = A · (−B) = −A · B ⇐⇒ A · B = 0.
11. i) A = A⊤ ⇒ A · L = A · (Ls + La ) = A · Ls + A · La = A · Ls .
ii) B = −B⊤ ⇒ B · L = B · (Ls + La ) = B · Ls + B · La = B · La .
12. i) A · (BCD) = tr(A⊤ BCD) = tr((B⊤ A)⊤ CD) = (B⊤ A) · (CD).
ii) A · (BCD) = (BCD) · A = tr((BCD)⊤ A) = tr(A(D⊤ C⊤ B⊤ ))
= tr((AD⊤ )(C⊤ B⊤ )) = tr((C⊤ B⊤ )(AD⊤ )) = tr((BC)⊤ (AD⊤ )) = BC · AD⊤
= AD⊤ · BC.
13. L ∈ Sym(V) ⇒ L · W = 0, as already proved. Now, if L · W = 0 ∀W ∈ Skw(V),
/ Sym(V) ⇒ L = Ls + La ⇒ L · W = Ls · W + La · W = La · W = 0; if
suppose L ∈
in Skw(V), we chose W = La , we get W · La = La · La = 0 ⇐⇒ La = O ⇒ L ∈
Sym(V).
1 1
14. I2 = (tr2 L − trL2 ) = (Lii Ljj − Lij Lji )
2 2
= L11 L22 + L11 L33 + L22 L33 − L12 L21 − L13 L31 − L23 L32 .
    
0 −a3 a2  b1   c 1 
15. a × b · c =  a3 0 −a1  b2 · c2 = a2 b 3 c 1 − a3 b 2 c 1 + a3 b 1 c 2 −
−a2 a1 0 b3 c3 
   
0 −a3 a2
a1 b3 c2 + a1 b2 c3 − a2 b1 c3 = det  a3 0 −a1  .
−a2 a1 0
16. Let L1 ̸= L2 be two distinct inverse tensors of L; then L1 L = I = L2 L ⇒
L1 L − L2 L = O ⇒ (L1 − L2 )L = O ∀L ⇐⇒ L1 − L2 = O ⇒ L1 = L2 .
17. (a ⊗ b)ij = ai bj ; it is then sufficient to write the matrix representing (a ⊗ b) and to
compute its determinant.
1 1
18. (αL)−1 (αL) = α(αL)−1 L = I ⇒ (αL)−1 L = I ⇒ (αL)−1 LL−1 = IL−1 ⇒
α α
−1 1 −1
(αL) = L .
α
19. |W2 | = W · W = tr(W⊤ W) = −tr(WW) = tr(I − w ⊗ w) = 3 − 1 = 2 ⇒
1
WW = − |W2 |(I − w ⊗ w).
2
20. Let w1 = (a1 , b1 , c1 ), w2 = (a2 , b2 c2 ); then, form W1 , W2 and compute the two
scalar products.
21. u × v = o ⇐⇒ v = ku, k ∈ R. So, u × v = o ⇒ u ⊗ v = ku ⊗ u ∈ Sym(V).
Conversely, if u⊗v ∈ Sym(V) then ∀w, w·v u = (u⊗v)w = (v⊗u)w = w·u v ⇒
w·v
v= u ⇒ u × v = o.
w·u
161
22. L = L⊤ ⇒ (RLR⊤ )⊤ = RL⊤ R⊤ = RLR⊤ ; moreover, u · Lu > 0 ∀u ⇒
u · (RLR⊤ )u = (R⊤ u) · L(R⊤ u) > 0.
   3
sph 1 sph 1 sph
23. i) det(L − λI) = det trLI − λ I = trLI − λ I det I
3 3
 3
1 1
= trLI − λsph I = 0 ⇒ λsph
i = trL, i = 1, 2, 3.
3 3
 
sph sph sph 1
ii) (L − λi I)v = o ∀i = 1, 2, 3 ⇒ L − trLI v = o ⇒
3
sph sph
(L − L )v = o ⇒ Ov = o, which is true ∀v.
   
dev dev dev dev 1 dev
24. det(L − λ I) = det(L − L − λ I) = det L − trL + λ I
  3
= det L − λsph + λdev I = 0 ⇒ λ = λsph + λdev is an eigenvalue of L ⇒
λdev = λ − λsph .

Chapter 3
1. ∀L ∈ Lin(V), (ei ⊗ ej ) ⊠ (ek ⊗ el )L = (ei ⊗ ej )L(ek ⊗ el )⊤ = (ei ⊗ ej )L(el ⊗ ek ) =
(ei ⊗ ej )((Lel ) ⊗ ek ) = ej · (Lel )ei ⊗ ek = Ljl ei ⊗ ek ; moreover (ei ⊗ ek ⊗ ej ⊗ el )L =
(ei ⊗ ek ) ⊗ (ej ⊗ el )L = ((ej ⊗ el ) · L)(ei ⊗ ek ) = ((ej ⊗ el ) · (Lpq ep ⊗ eq ))(ei ⊗ ek ) =
Lpq δjp δlq (ei ⊗ ek ) = Ljl ei ⊗ ek ⇒ (ei ⊗ ej ) ⊠ (ek ⊗ el ) = ei ⊗ ek ⊗ ej ⊗ el .
2. ∀L, M ∈ Lin(V), L · (AB)M = A⊤ L · BM = B⊤ A⊤ L · M ⇒ (AB)⊤ = B⊤ A⊤ .
3. ∀C ∈ Lin(V), (A ⊗ BL)C = (A ⊗ B)LC = B · LCA = L⊤ B · CA = (A ⊗ L⊤ B)C.
4. ∀L ∈ Lin(V), ((A ⊠ B)(C ⊠ D))L = A ⊠ BCLD⊤ = ACLD⊤ B⊤
= (AC) ⊠ (D⊤ B⊤ )⊤ L = (AC) ⊠ (BD)L.
5. Let A = Aijkl ei ⊗ej ⊗ek ⊗el = Aijkl (ei ⊗ek )⊠(ej ⊗el ) and B = Bpqrs ep ⊗eq ⊗er ⊗es =
Bpqrs (ep ⊗ er ) ⊠ (eq ⊗ es ).
Then, AB = Aijkl Bpqrs ((ei ⊗ ek ) ⊠ (ej ⊗ el ))((ep ⊗ er ) ⊠ (eq ⊗ es ))
= Aijkl Bpqrs ((ei ⊗ ek )(ep ⊗ er )) ⊠ ((ej ⊗ el )(eq ⊗ es ))
= Aijkl Bpqrs δkp δlq(ei ⊗ er ) ⊠ (ej ⊗ es ) = Aijkl Bklrs (ei ⊗ ej ⊗ er ⊗ es ).
6. ∀L ∈ Lin(V), (A⊗B)(C⊠D)L = (A⊗B)CLD⊤ = B·(CLD⊤ )A = BD·(CL)A =
C⊤ BD · L A = (C⊤ ⊠ D⊤ )B · L A = A ⊗ ((C⊤ ⊠ D⊤ )B)L.
7. ∀L ∈ Lin(V), (A⊠B)(C⊗D)L = D·L(A⊠B)C = (D·L)ACB⊤ = ACB⊤ (D·L) =
((A ⊠ B)C)(D · L) = (((A ⊠ B)C) ⊗ D)L.
8. (P ⊗ P)ijhk = Pij Phk = (p ⊗ p)ij (p ⊗ p)hk = pi pj ph pk
= pi ph pj pk = (p ⊗ p)ih (p ⊗ p)jk = Pih Pjk = (P ⊠ P)ijhk .
9. i) IA = (I⊠I)A = (I⊠I)Aijkl (ei ⊗ej )⊗(ek ⊗el ) = Aijkl ((I⊠I)(ei ⊗ej ))⊗(ek ⊗el ) =
Aijkl (I(ei ⊗ ej )I⊤ ) ⊗ (ek ⊗ el ) = Aijkl (ei ⊗ ej ⊗ ek ⊗ el ) = A.

162
ii) AI = Aijkl (ei ⊗ ej ) ⊗ (ek ⊗ el )(I ⊠ I) = Aijkl (ei ⊗ ej ) ⊗ ((I ⊠ I)ek ⊗ el ) =
Aijkl (ei ⊗ ej ) ⊗ (I(ek ⊗ el ))I⊤ ) = Aijkl (ei ⊗ ej ⊗ ek ⊗ el ) = A.
10. (A ⊗ B) · (C ⊗ D) = tr4 ((A ⊗ B)⊤ (C ⊗ D)) = tr4 ((B ⊗ A)(C ⊗ D))
= tr4 ((B ⊗ A)C) ⊗ D = tr4 (A · C B ⊗ D) = A · C tr4 (B ⊗ D) = A · C B · D.
I I I I 1
11. ⊗ = √ ⊗ √ = I ⊗ I = Ssph .
|I| |I| 3 3 3
1
12. ∀L ∈ Lin(V), L = Lsph + Ldev and Lsph = trL I → just one number is sufficient
3
to determine L ⇒ dim(Sph(V)) = 1. Then, Ldev = L − Lsph is determined by 5
sph

numbers : dim(Dev(V)) = dim(Lin(V) − Sph(V)) = 6 − 1 = 5.


  
1 1 1 1 1
13. i) Ssph Ssph = I⊗I I ⊗ I = (I ⊗ I) = I · I I ⊗ I = I ⊗ I = Ssph .
3 3 9 9 3
ii) Ddev Ddev = (Is − Ssph )(Is − Ssph ) = Is − 2Ssph + Ssph Ssph = Is − Ssph = Ddev .
iii) Ssph Ddev = Ssph (Is − Ssph ) = Ssph − Ssph Ssph = Ssph − Ssph = O.
sym ek ⊗ el + el ⊗ ek
14. i) Sijkl = (ei ⊗ ej ) · Ssym (ek ⊗ el ) = (ei ⊗ ej ) ·
2
1 skw skw ek ⊗ el − el ⊗ ek
= (δik δjl + δil δjk ), Wijkl = (ei ⊗ ej ) · W (ek ⊗ el ) = (ei ⊗ ej ) ·
2 2
1 sym skw sym dev
= (δik δjl − δil δjk ), → Sijkl + Wijkl = δik δjl = Iijkl ⇒ S + W = I.
2
sym skw trp
ii) Sijkl − Wijkl = δil δjk = Tijkl ⇒ Ssym − Wdev = Ttrp .
   
1 1 1 1
sph
15. i) S · S = sph
I⊗I · I ⊗ I = tr4 ((I ⊗ I)⊤ (I ⊗ I)) = tr4 ((I ⊗ I)(I ⊗ I))
3 3 9 9
1 1 1 1
= tr4 ((I ⊗ I)I) ⊗ I = tr4 (I · I I ⊗ I) = tr4 (I⊤ ⊗ I) = I · I = 1.
9 9 3 3
1
ii) Ddev · Ddev = (Is − Ssph ) · (Is − Ssph ) = Is · Is − 2Is · Ssph + Ssph · Ssph = (δih δjk +
4
δih δjk + δik δjh 1 1
δik δjh )(δih δjk + δik δjh ) − 2 δij δhk + 1 = (δih δih δjk δjk + 2δik δjh δih δjk +
2 3 4
1
δik δik δjh δjh ) − (δij δhk δih δjk + δij δhk δik δjh ) + 1. Now, one should check that δih δih =
3
δjk δjk = δik δik = δjh δjh = δik δjh δih δjk = δij δhk δih δjk = δij δhk δik δjh = 3 ⇒
Ddev · Ddev = 5.
δih δjk + δik δjh 1
iii) Ssph · Ddev = Ssph · (Is − Ssph ) = Is · Ssph − Ssph · Ssph = δij δhk − 1 =
2 3
1 1
(δih δjk δij δhk + δik δjh δij δhk ) − 1 = (3 + 3) − 1 = 0.
6 6
16. SR = I − 2n ⊗ n → SR = (I − 2n ⊗ n) ⊠ (I − 2n ⊗ n)
= I ⊠ I − 2(n ⊗ n ⊠ I + I ⊠ n ⊗ n + 4n ⊗ n ⊠ n ⊗ n)
= I − 2(n ⊗ n ⊠ ei ⊗ ei + ei ⊗ ei ⊠ n ⊗ n + 4n ⊗ n ⊗ n ⊗ n)
= I − 2(n ⊗ ei ⊗ n ⊗ ei + ei ⊗ n ⊗ ei ⊗ n + 4n ⊗ n ⊗ n ⊗ n).

163
By components, (SR )ijkl = (I − 2n ⊗ n)ik (I − 2n ⊗ n)jl = (δik − 2ni nk )(δjl − 2nj nl )
= δik δjl − 2(δik nj nl + δjl ni nk ) + 4ni nj nk nl .
17. i) R0 = 0: This is the case of the so-called R0 -orthotropy.
ii) R1 = 0: This is the case of the square symmetry (all the components depend
upon 4θ).
π
iii) Φ0 − Φ1 = k , k ∈ {0, 1}: This is the case of the two ordinary orthotropies.
4
iv) R0 = R1 = 0: This is the condition for isotropy. Nothing depends upon θ ⇒ all
the directions are equivalent and thus, at the same time, axes of elastic symmetry.

Chapter 4
1. i) (u(t) + v(t)) − (u(t0 ) + v(t0 )) = (t − t0 )(u + v)′ + o(t − t0 ) and also u(t) − u(t0 ) +
v(t) − v(t0 ) = (t − t0 )u′ + (t − t0 )v′ + o(t − t0 ) ⇒ (u + v)′ = u′ + v′ .
The proof that (L + M)′ = L′ + M′ and that (L + M)′ = L′ + M′ can be done in a
similar way.
ii) We indicate, in short, α(t) = α, v(t) = v, α(t0 ) = α0 , v(t0 ) = v0 , α′ (t0 ) =
α0′ , v′ (t0 ) = v0′ → αv − α0 v0 = (αv)′0 (t − t0 ) + o(t − t0 );
moreover α = α0 + α0′ (t − t0 ) + o(t − t0 ), v = v0 + v0′ (t − t0 ) + o(t − t0 ) ⇒ αv − α0 v0 =
(α0 + α0′ (t − t0 ) + o(t − t0 ))(v0 + v0′ (t − t0 ) + o(t − t0 )) − α0 v0 = α0′ v0 (t − t0 ) + v0 o(t −
t0 )+α0 v0′ (t−t0 )+α0′ v0′ (t−t0 )2 +v0′ (t−t0 )o(t−t0 )+α0 o(t−t0 )+α0′ (t−t0 )o(t−t0 )+
o(t − t0 )2 = (α0′ v0 + α0 v0′ )(t − t0 ) + o(t − t0 ), so by comparison, (αv)′ = α′ v + αv′ .
By the same technique, one can easily prove the differentiation rule for all the
product-like quantities: (u · v)′ , (u × v)′ , (u ⊗ v)′ , (αL)′ , (Lv)′ , (LM)′ , (L · M)′ ,
(L ⊗ M)′ , (L ⊠ M)′ , (αL)′ , (LL)′ , (LM)′ , (L · M)′ .
2. The proof is the same for all the cases, we just write that for v(t). Using the two
properties shown in the previous exercise, we get: v(t) = vi (t)ei ⇒
v′ (t) = (vi (t)ei )′ = vi′ (t)ei + vi (t)e′i = vi′ (t)ei , because ei does not depend on t.
2 + θ2
3. i) p(θ) = (aθ cos θ, aθ sin θ) → c = 3 .
a(1 + θ2 ) 2
a √ √
ii) ℓ = (θ 1 + θ2 + ln(θ + 1 + θ2 )).
2
iii) Let pi , pi+1 be two consecutive intersection points of the spiral (i denotes the
order of the intersection point) with a straight line passing through the origin and
inclined at θ; their distances from the origin are ri = a(θ + 2πi), ri+1 = a(θ + 2π(i +
1)) ⇒ |pi+1 − pi | = ri+1 − ri = 2πa that does not depend upon θ.
4. i) r = a ebθ ⇒ r = 0 ⇐⇒ θ → +∞, for b < 0, −∞ for b > 0.
1
ii) p(θ) = (a ebθ cos θ, a ebθ sin θ) → c = √ .
a ebθ 1 + b2

164
a√
iii) ℓ = 1 + b2 ebθ .
b
iv) With the same meaning as in the previous exercise, |pi+1 − pi | = ri+1 − ri =
a(e2πb − 1)eb(θ+2πi) , i.e. the distance depend upon the order of the intersection:
ri+1
= e2πb , which is a geometric progression.
ri
1
v) τ = √ (b cos θ − sin θ, b sin θ + cos θ) ⇒
1 + b2
ab ebθ (p − o) · τ b
(p − o) · τ = √ ⇒ cos φ = = √ ⇒ φ, the angle between τ
1 + b2 |p − o||τ | 1 + b2
and p − o, is constant.
1 1
vi) ν = √ (−b sin θ − cos θ, b cos θ − sin θ) ⇒ q = p + ν = ab ebθ (−sinθ, cos θ)
1 + b2 c
is a point of the evolute, whose polar equation is hence r = ab ebθ , which is still a
logarithmic spiral.
1
5. i) c = .

1
ii) ℓ = aθ2 .
2
1
iii) ν = (− sin θ, cos θ) ⇒ p + ρν = p + ν = a(cos θ, sin θ), which is the parametric
c
equation of a circle of center o and radius a.
1 b
6. i) τ = √ (−a sin ωθ, a cos ωθ, b) ⇒ cos φ = τ · e3 = √ , which is
a2
+b 2 a + b2
2
independent of θ.

ii) ℓ = ωθ a2 + b2 .
a
iii) c = 2 .
a + b2
b
iv) ϑ = − .
a2 + b2
v) Let pi , pi+1 be two points, intersection of a same generatrix of the cylinder with
2π 2π
the helix, i.e. for, say, θ + i and θ + (i + 1) ⇒ d = (pi+1 − pi ) · e3 = 2πb.
ω ω
vi) By definition, a curve is a helix ⇐⇒ τ · e3 = const; differentiating gives
τ ′ · e3 + τ · e3 = τ ′ · e3 = 0 ⇒ by the first equation of Frenet-Serret cν · e3 =
0 ⇒ ν · e3 = 0 ⇒ β is tangent to the cylinder and β · e3 = const. Moreover,
differentiating again, ν · e3 + ν · e′3 = ν ′ · e3 = 0 ⇒ by the third equation of Frenet-
c β · e3
Serret (−cτ − ϑβ) · e3 = 0 ⇒ −cτ · e3 = ϑβ · e3 ⇒ = − = const.
ϑ τ · e3
c
Conversely, if p(s) is a curve with = α = const. through the first and second
ϑ
1 1 c
equations of Frenet-Serret, we get ν = τ ′ = β ′ ⇒ τ ′ = β ′ = αβ ′ ⇒ (τ −
c ϑ ϑ
165
αβ)′ = 0 ⇒ v = τ − αβ = const. ⇒ τ · v = 1 ⇒ τ forms with v a constant angle,
and because v is a constant vector, τ · e3 = const. ⇒ the curve is a helix.
vii) p′ × p′′ = ω 3 (ab sin ωθ, −ab cos ωθ, a2 ) ⇒ A = abω 3 , B = a2 ω 3 .
7. i) p(θ) = R(θ − sin θ, 1 − cos θ).
 
θ
ii) ℓ(θ) = 4R 1 − cos ⇒ ℓ(2π) = 8R.
2
1
iii) c = p .
2R 2(1 − cos θ)
1 1
iv) ν = p (sin θ, cos θ − 1) ⇒ q(θ) = p(θ) + ν = R(θ + sin θ, cos θ − 1);
2(1 − cos θ) c
this curve is the evolute of the cycloid and it can be obtained also as
q(θ) = p(θ + π) − (π, 2)R, i.e., it is the same cycloid p(θ) translated by −(π, 2)R.
1
8. i) c = .
cosh2 t
 
sinh t 1 1
ii) ν = − , ⇒ q(t) = p(t) + ν = (t − sinh t cosh t, 2 cosh t) is the
cosh t cosh t c
evolute.
 
R ′ 1 sinh t
iii) s = |p (t)|dt = sinh t, τ = , ⇒ b(t) = p(t) + (a − s)τ =
 cosh
 t cosh t
a − sinh t a − sinh t
t+ , cosh t + sinh t is the equation of the involutes.
cosh t cosh t
9. 
i) τ = (cos t, sin t); 
tangentto the tractrix at p: z = p + wτ =
t
(1 + w) cos t + ln tan , (1 + w) sin t . Intersection point g with x1 axis for
  2  
t
w = −1 → g = ln tan , 0 ⇒ |p − g| = 1 ∀t.
2
sin t2
ii) ℓ = ln .
sin t1
iii) c = tan t.
   
1 t 1
iv) ν = (− sin t, cos t) ⇒ q = p + ν = ln tan , . Setting σ =
  c 2 sin t
t t 1
ln tan ⇒ tan = eσ , = cosh σ ⇒ q = (σ, cosh σ), which is the equa-
2 2 sin t
tion of a catenary.
10. i) p(θ) = (cos√θ, sin θ, sin θ) ⇒ p′ (= − sin θ, cos θ, cos θ), p′′ = (− cos θ, − sin θ, − sin θ)
2 √
⇒c= 3 ⇒ c max = 2.
(1 + cos2 θ) 2
ii) p′′′ = (sin θ, − cos θ, − cos θ) ⇒ p′ × p′′ · p′′′ = cos θ − cos θ = 0 ⇒ ϑ = 0 ⇒ the
curve is planar.

166

11. i) v = ṗ, τ = ⇒ ṗ = |ṗ|τ ⇒ v = vτ , v = |ṗ| is the scalar velocity.
|ṗ|
|τ̇ | v
ii) a = p̈ = v̇ = (vτ )· = v̇τ + v τ̇ ; τ̇ = |τ̇ |ν; c =
⇒ |τ̇ | = |ṗ|c = ⇒
|ṗ| ρ
2 2
v v
a = v̇τ + ν; v̇ is the tangential acceleration and the centripetal acceleration.
ρ ρ
mv 2
iii) f = ma = mv̇τ + ν = fτ τ + fν ν; fτ = mv̇ is the tangential force, which is
ρ
mv 2
responsible for the change of the scalar velocity; fν = is the centripetal force,
ρ
which is responsible for the path change.

Chapter 5
1. i) By Eq. (5.1)3 : grad(v ·w) = (gradw)⊤ v +(gradv)⊤ w = (gradw)⊤ v +(gradw)v −
(gradw)v+(gradv)⊤ w+(gradv)w−(gradv)w = (gradw)v+(gradv)w+((gradw)⊤ −
gradw))v + ((gradv)⊤ − gradv))w = (gradw)v + (gradv)w − (curlw) × v −
(curlv) × w = (gradw)v + (gradv)w + v × (curlw) + w × (curlv).
ii) By Eqs. (5.1)2,3 and Exercise 3 iii), Chapter 2, grad(u · v w) = u · v gradw +
w ⊗ grad(u · v) = u · v gradw + w ⊗ ((gradu)⊤ v + (gradv)⊤ u) = u · v gradw +
w ⊗ (gradu)⊤ v + w ⊗ (gradv)⊤ u = u · v gradw + (w ⊗ v)gradu + (w ⊗ u)gradv.
iii) By Theorem 30 i), div((gradv)v − (divv)v) + (divv)2 = div((gradv)v) −
div((divv)v) + (divv)2 = (gradv)⊤ · gradv + v · div(gradv)⊤ − (divv)2 −
v · grad(divv) + (divv)2 = (gradv)⊤ · gradv + v · div(gradv)⊤ − v · div(gradv)⊤ =
(gradv)⊤ · gradv.
2. i) By Theorem 30 i), div(gradv)⊤ = div(vj,i ei ⊗ ej ) = vj,i div(ei ⊗ ej ) +
ei ⊗ ej gradvj,i = ei ⊗ ej vj,ik ek = vj,ik δjk ei = vj,ij ei = vj,ji ei = (divv),i ei =
grad(divv).
ii) By iii) of the previous exercise and Theorem 30 i), div((gradv)v − (divv)v) =
gradv · (gradv)⊤ − (divv)2 and also div((gradv)v − (divv)v) = div(gradv)v) −
div((divv)v) = div((gradv)v) − (divv)2 − v · grad(divv), so comparing the two
results div((gradv)v) = gradv · (gradv)⊤ + v · grad(divv).
iii) By Theorem 30 i) and iv), div(φLv) = φdiv(Lv) + Lv · gradφ = Lv · gradφ +
φL⊤ · gradv + φv · divL⊤ .
3. i) By Theorem 28 ii) and Eq. (2.29), ∀a = const. ∈ V, (curl(φv))×a = (grad(φv)−
(grad(φv)⊤ )a = (φgradv + v ⊗ gradφ − (φgradv + v ⊗ gradφ)⊤ )a = (φgradv +
v ⊗ gradφ − φ(gradv)⊤ − gradφ ⊗ v)a = φ(curlv) × a + a · gradφ v − a · v gradφ =
φ(curlv) × a + a × (v × gradφ) = φ(curlv) × a − (v × gradφ) × a = (φcurlv −
v × gradφ) × a ⇒ curl(φv) = φcurlv + gradφ × v.
ii) Using the Ricci’s alternator for the cross product, v × w = ϵpqr vq wr ep and

167
curl(v × w) = ϵijk (v × w)k,j ei = ϵijk ϵkqr (vq wr ),j ei = ϵijk ϵkqr (vq,j wr + vq wr,j )ei =
ϵkij ϵkqr (vq,j wr +vq wr,j )ei . Then, because ϵkij ϵkqr = δiq δjr −δir δjq we get curl(v×w) =
(δiq δjr −δir δjq )(vq,j wr +vq wr,j )ei = δiq δjr (vq,j wr +vq wr,j )ei −δir δjq (vq,j wr +vq wr,j )ei =
(vi,j wj + vi wj,j )ei − (vj,j wi + vj wi,j )ei = gradv w − gradw v + vdivw − wdivv.
4. Ri) Through Theorem 30 iv) we get ∂Ω v·Ln dA = ∂Ω L⊤ v·n dA = Ω div(L⊤ v)dV =
R R R


L · gradv + v · divL dV.
R R
Rii) By Theorems 29 and R 30 i), ∀a = const. ∈ V, ∂Ω (Ln)⊗v a dA = ∂Ω a·v Ln dA =
div(a · v L)dV =R Ω a · v divL + Lgrad(a · v)dV = Ω a · v divL + L(grada)⊤ v +
R

L(gradv)⊤ a dV = Ω ((divL) ⊗ v + L(gradv)⊤ )a dV.
R R R
R By Theorem 30 ii), ∂Ω (w · n)v dA = ∂Ω (v ⊗ w)n dA = Ω div(v ⊗ w)dV =
iii)

v divw + (gradv)w dV.
5. i) Take u = αn, n ∈ S → φ(p + αn) = φ(p) + α gradφ · n + o(α) ⇒
dφ φ(p + αn) − φ(p)
:= limα→0 = gradφ · n.
dn α
ii) In a similar way, v(p + αn) = v(p) + α gradv n + o(α) ⇒
dv v(p + αn) − v(p)
:= limα→0 = gradv n.
dn α
6. i) Applying the first proof of the previous exercise to n = ei , i = 1, 2, 3, we get
df
immediately := f,i = gradf · ei := (gradf ),i ⇒ gradf = f,i ei .
dei
ii) Applying the second proof of the previous exercise to n = ej , j = 1, 2, 3, we obtain
dv
:= v,j = vi,j ei = gradv ej = (gradv)ik ei ⊗ ek ej = δkj (gradv)ik ei = (gradv)ij ei
dej
⇒ (gradv)ij = vi,j ⇒ gradv = vi,j ei ⊗ ej .
iii) divv := tr(gradv) = tr(vi,j ei ⊗ ej ) = vi,j tr(ei ⊗ ej ) = δij vi,j = vi,i .
iv) ∀u = const. ∈ V, (divL) · u := div(L⊤ u) ⇒ (divL)i ui = (L⊤ u)j,j = (Lij ui ),j =
Lij,j ui ⇒ (divL)i = Lij,j ⇒ divL = Lij,j ei .
v) ∆f := div(gradf ) = div(f,i ei ) = f,ii .
vi) ∆v := div(gradv) = div(vi,j ei ⊗ ej ) = vi,jj .
vii) For the sake of brevity, let w = curlv; ∀u ∈ V,  
0 −w3 w2  u1 
(curlv) × u := (gradv − gradv⊤ )u ⇒  w3 0 −w1  u2 =
−w w 0 u
 
   2  1  3
v1,1 v1,2 v1,3 v1,1 v2,1 v3,1  u1 
 v2,1 v2,2 v2,3  −  v1,2 v2,2 v3,2  u2 =
v v v v v v u
 
3,1 3,2 3,3 1,3 2,3 3,3  3
   
0 v1,2 − v2,1 v1,3 − v3,1  u1   v3,2 − v2,3 
 v2,1 − v1,2 0 v2,3 − v3,2  u2 ⇒ curlv = v1,3 − v3,1 .
v3,1 − v1,3 v3,2 − v2,3 0 u3 v2,1 − v1,2
   

168
7. i) v(p) = v(p0 ) + ω × (p − p0 ) ⇒ ∃Wω ∈ Skw(V)| v(p) = v(p0 ) + Wω (p − p0 ), with
Wω the axial tensor of ω. Moreover, by the definition of gradient, v(p) = v(p0 ) +
gradv + gradv⊤ gradv − gradv⊤
(gradv)(p − p0 ) ⇒ Wω = gradv ⇒ gradv = + =
2 2
gradv + gradv⊤ gradv − gradv⊤
Wω = −Wω⊤ ⇐⇒ = O ⇒ gradv = ⇒
2 2

gradv − gradv
v(p) = v(p0 ) + (p − p0 ) so, by the definition of curl,
2
1 1
v(p) = v(p0 ) + curlv × (p − p0 ), and comparing the two results we get ω = curl.
2 2
ii) By the definition of divergence and Eq. (2.8), divv = tr(gradv) = trWω = 0.
8. i) divu = 3α → nowhere isochoric.
ii) divu = 0 → globally isochoric.
iii) divu = γ(x1 + x2 + x3 ) = 0: isochoric on the points of the plane x1 + x2 + x3 = 0.
iv) divu = δ(cos x1 + sin x2 + cos x3 ) = 0: isochoric on the points of the surface
cos x1 + sin x2 + cos x3 = 0.
9. Using Eq. (5.14), we get:
α α α
i) vθ = vz = 0, vρ = ⇒ divv = − 2 + 2 = 0.
ρ ρ ρ
α
ii) vρ = vz = 0, vθ = ⇒ divv = 0.
ρ
α cos θ α sin θ 2α cos θ 2α cos θ
iii) vρ = 2
, vθ = 2
, vz = 0 ⇒ divu = − + = 0.
ρ ρ ρ3 ρ3
10. Using Eq. (5.15), we get:
i) vρ,θ = vρ,z = vθ,ρ = vθ,z = vz,ρ = vz,θ = 0 ⇒ curlv = o.
 
α α α
ii) vρ,θ = vρ,z = vθ,z = vz,ρ = vz,θ = 0, vθ,ρ = − 2 ⇒ curlv = 0, 0, 2 − 2 = o.
ρ ρ ρ
α cos θ 2α sin θ
iii) vρ,θ = − 2
, vρ,z = 0, vθ,ρ = − , vθ,z = 0, vz,ρ = vz, θ = 0 ⇒
 ρ ρ3 
α cos θ α cos θ α cos θ
curlv = 0, 0, 3
−2 + = o.
ρ ρ3 ρ3

Chapter 6
 
1 0 0
1. Setting ρ = z 1 , θ = z 2 , z = z 3 , by Eq. (6.3) we get g =  0 ρ2 0  ⇒
0 0 1
p p
h k 2
ds = ghk dz dz = dρ + ρ dθ + dz . 2 2 2

169
1 2 3
2. Setting
 r = z , θ = z , φ= z , proceeding in a similar way, we get
1 0 0 p p
2
g = 0 r sin φ 0  ⇒ ds = ghk dz h dz k = dr2 + r2 sin2 φdθ2 + r2 dφ2 .
 2

0 0 r2
p
3. For cylindrical coordinates, cf. Exercise 1, ds = dρ 2 2 2 2
√ + ρ dθ + dz , and for a curve
2 2 2
on a circular cylinder, ρ = R ⇒ dρ = 0 ⇒ ds = R dθ + dz ; if the equation of
dz
the helix is p(θ) = R cos θe1 + R sin θe2 + bθe3 , then = b ⇒ dz = b dθ ⇒
√ R θ+2π √ √ dθ
ds = R2 + b2 dθ ⇒ ℓ = θ R2 + b2 dθ = 2π R2 + b2 .
2R 2R √ 2R √ Rπ
4. i) r = θ ⇒ dr = dθ; ds = dr2 + r2 dθ2 = 1 + θ2 dθ ⇒ ℓ = 02 ds =
π π π
2R R π2 √ R √ 1√

2 2
π
2
1 π
1 + θ dθ = [θ 1 + θ + arcsinhθ]0 = 2
4 + π + arcsinh R ∼
π 0 π 4 π 2
1.324R.
π R p 15 R0 h
ii) R = 2πR0 ⇒ R0 = ⇒ h = R2 − R02 = R; ρ(z) = z, z = θ⇒
2 4 √ 4 h√ 2π
R0 R 15 R 15
ρ(θ) = θ ⇒ ρ(θ) = θ, z(θ) = Rθ ⇒ dρ = dθ, dz = R dθ; equa-
2π 8π 8π 8π 8π
√ R
tion of the conical helix: p(θ) = ρ(z) cos θe1 +ρ(z) sin θe2 + 15θe3 = (θ cos θe1 +

√ p R√ 2
θ sin θe2 + 15θe3 ) ⇒ ds = dρ2 + ρ2 dθ2 + dz 2 = dθ + θ2 dθ2 + 15dθ2 =

R√ R 2π R R 2π √
16 + θ2 dθ ⇒ ℓ = 0 ds = 0
16 + θ2 dθ =
8π  8π
2π
R 1 √ 2
θ
θ 16 + θ + 8arcsinh ∼ 1.324R.
8π 2 4 0
1 2 1
5. i) Referring to Fig.
 6.5 and by Eq. (6.3), x1 =z cos α1 + z cos α2 , x2 = z sin α1 +
1 cos(α2 − α1 )
z 2 sin α2 ⇒ g = .
cos(α2 − α1 ) 1
ii) By Eq. (6.5), g1 = cos α1 e1 + sin α1 e2 , g2 = cos α2 e1 + sin α2 e2 .
1
iii) z 1 = h(x1 sin α2 −x2 cos α2 ), z 2 = h(−x1 sin α1 +x2 cos α1 ), h = ⇒
sin(α2 − α1 )
sin α2 cos α2 sin α1
by Eq. (6.14), g1 = e1 − e2 , g 2 = − e1 +
sin(α2 − α1 ) sin(α2 − α1 ) sin(α2 − α1 )
cos α1
e2 .
sin(α2 − α1 )
cos α1 sin α2 − sin α1 cos α2
iv) g1 · g1 = = 1,
sin(α2 − α1 )
− sin α1 cos α2 + sin α2 cos α1
g2 · g 2 = = 1,
sin(α2 − α1 )
− cos α1 sin α1 + sin α1 cos α1
g1 · g 2 = = 0,
sin(α2 − α1 )
cos α2 sin α2 − sin α2 cos α2
g2 · g 1 = = 0.
sin(α2 − α1 )

170
s
sin2 α2 1
v) |g1 | = |g2 | = 1, |g1 | = sin2 α1 + 2 , |g2 | = .
sin (α2 − α1 ) | sin(α2 − α1 )|

vi)

g2
x2
g2

g1
𝛼2
𝛼1
x1
g1

6. Referring to Exercise 2 and by Eq. (6.5), we get g1 = cos θ sin φe1 + sin θ sin φe2 +
cos φe3 , g2 = −r sin θ sin φe1 + r cos θ sin φe2 , g3 = r cos θ cos φe1 + r sin θ cos φe2 −
r sin φe3 .

7. i) If z 1 = const. ⇒ x1 = a cos z 2 , x2 = b sin z 2 , with a = c cosh z 1 = const., b =


x2 x2
c sinh z 1 = const. ⇒ 21 + 22 = 1: family of ellipses all with the same focuses
√ a b
xe = ± a2 − b2 = ±c.

ii) If z 2 = const. ⇒ x1 = A cosh z 1 , x2 = B sinh z 1 , with A = c cos z 2 = const., B =


x2 x2
c sin z 2 = const. ⇒ 12 − 22 = 1: family of hyperbolae all with the same focuses
√ A B
xh = ± A2 + B 2 = ±c ⇒ xe = xh .

iii) The axes of the ellipses are 2a = 2c cosh z 1 and 2b = 2c sinh z 1

iv) A crack along the horizontal axis corresponds to b → 0, which happens ⇐⇒


z 1 → 0 ⇒ cosh z 1 → 1 and a → c ⇒ length of the crack: 2c.

c2
v) Applying Eq. (6.3), we get g = (cosh 2z 1 − cos 2z 2 )I.
2

vi) By Eq. (6.5), g1 = c sinh z 1 cos z 2 e1 +c cosh z 1 sin z 2 e2 , g2 = −c cosh z 1 sin z 2 e1 +


c sinh z 1 cos z 2 e2 . We note that g1 · g2 = 0.

8. i) By Eq. (6.16)1 , setting z 1 = ρ, z 2 = θ, z 3 = z, we get:

171
L11 = Lx11 cos2 θ + (Lx12 + Lx21 ) sin θ cos θ + Lx22 sin2 θ,
1
L12 = ((Lx22 − Lx11 ) sin θ cos θ + Lx12 cos2 θ − Lx21 sin2 θ),
ρ
L13 = Lx13 cos θ + Lx23 sin θ,
1
L21 = ((Lx22 − Lx11 ) sin θ cos θ − Lx12 sin2 θ + Lx21 cos2 θ),
ρ
1
L = 2 (Lx11 sin2 θ − (Lx12 + Lx21 ) sin θ cos θ + Lx22 cos2 θ),
22
ρ
L23 = −Lx13 sin θ + Lx23 cos θ,
L31 = Lx31 cos θ + Lx32 sin θ,
L32 = −Lx31 sin θ + Lx32 cos θ,
L33 = Lx33 .

ii) The covariant components can alternatively be found by Eq. (6.16)2 or, using
the results of the previous point, by Eq. (6.19)2 ; by this latter way, using the result
of Exercise 1, we get easily L11 = L11 , L12 = ρ2 L12 , L13 = L13 , L21 = ρ2 L21 ,
L22 = ρ4 L22 , L23 = ρ2 L23 , L31 = L31 , L32 = ρ2 L32 , L33 = L33 .

9. i) In this case, we first calculate the covariant components: By Eq. (6.16)2 , setting
z 1 = r, z 2 = θ, z 3 = φ, we get:

L11 = Lx11 cos2 θ sin2 φ + (Lx12 + Lx21 ) sin θ cos θ sin2 φ + (Lx13 + Lx31 ) cos θ sin φ cos φ
+Lx22 sin2 θ sin2 φ + (Lx23 + Lx32 ) sin θ sin φ cos φ + Lx33 cos2 φ,
L12 = −r cos θ sin θ sin2 φLx11 + r sin2 φ(Lx12 cos2 θ − Lx21 sin2 θ)
+r sin θ cos θ sin2 φLx22 + r sin φ cos φ(Lx32 cos θ − Lx31 sin θ),
L13 = r cos2 θ sin φ cos φLx11 + r sin θ cos θ sin φ cos φ(Lx12 + Lx21 )
+r sin2 θ cos φ sin φLx22 − r sin2 φ sin θLx23 + r cos θ(Lx31 cos2 φ − Lx13 sin2 φ)
−r cos φ sin φLx33 ,
L21 = −r sin θ cos θ sin2 φLx11 + r sin2 φ(Lx21 cos2 θ − Lx12 sin2 θ)
+r sin θ cos θ sin2 φLx22 + r sin φ cos φ(Lx23 cos θ − Lx13 sin θ),
L22 = r2 sin2 θ sin2 φLx11 − r2 sin θ cos θ sin2 φ(Lx12 + Lx21 ) + r2 cos2 θ sin2 φLx22 ,
L23 = −r2 sin θ cos θ sin φ cos φ(Lx22 − Lx11 ) + r2 sin φ cos φ(Lx21 cos2 θ − Lx12 sin2 θ)
+r2 sin2 φ(Lx13 sin θ − Lx23 cos θ),
L31 = r sin φ cos φ(Lx11 cos2 θ + Lx22 sin2 θ) + r sin θ cos θ sin φ cos φ(Lx12 + Lx21 )
+r cos θ(Lx13 cos2 φ − Lx31 sin2 φ) + r sin θ(Lx23 cos2 φ − Lx32 sin2 φ)
−r sin φ cos φLx33 ,
L32 = r2 sin θ cos θ sin φ cos φ(Lx22 − Lx11 ) + r2 sin φ cos φ(Lx12 cos2 θ − Lx21 sin2 θ)
+r2 sin2 φ(Lx31 sin θ − Lx32 cos θ),
L33 = r2 cos2 φ(Lx11 cos2 θ + Lx22 sin2 θ) + r2 sin θ cos θ cos2 φ(Lx12 + Lx21 )
−r2 cos θ sin φ cos φ(Lx13 + Lx31 ) − r2 sin θ sin φ cos φ(Lx23 + Lx32 ) + r2 sin2 φLx33 .

ii) For the contravariant components, we use Eq. (6.19)1 , after having calculated
gcont ; this can be done either using Eq. (6.11) or simply observing that gcov is

172
 
0 0 1
1  1 
diagonal (see Exercise 2) and that g pq = ⇒ gcont =  0 r2 sin2 φ 0 ⇒
 
gpq  1 
0 0
r2
L 12 L 13 L 21 L 22
L11 = L11 , L12 = 2 2 , L13 = 2 , L21 = 2 2 , L22 = 4 4 ,
r sin φ r r sin φ r sin φ
23 L23 31 L31 32 L32 33 L33
L = 4 2 , L = 2 , L = 4 2 , L = 4 .
r sin φ r r sin φ r
10. i) trL = Lxhh , Eq. (2.7).
∂xh ∂xk ij
ii) By Eq. (2.7)1 ⇒ Lxhh = L δhk = gij Lij .
∂z i ∂z j
∂z i ∂z j
iii) By Eq. (2.7)2 ⇒ Lxhh = Lij δhk = g ij Lij .
∂xh ∂xk
∂xh ∂z j i
iv) By Eq. (2.7)3 ⇒ Lxhh = L δhk = δi j Li j = Li i .
∂z i ∂xk j
∂z i ∂xk j
v) By Eq. (2.7)4 ⇒ Lxhh
= j
Li δhk = δ i j Li j = Lj j .
∂xh ∂z
 
1 hm ∂gmk ∂gml ∂gkl
11. g + − m =
2 ∂z l ∂z k ∂z
h m
 
1 ∂z ∂z ∂ ∂xp ∂xp ∂ ∂xp ∂xp ∂ ∂xp ∂xp
+ − m k l =
2 ∂xp ∂xp ∂z l ∂z m ∂z k ∂z k ∂z m ∂z l ∂z ∂z ∂z
1 ∂z h ∂z m ∂ 2 xp ∂xp ∂xp ∂ 2 xp ∂ 2 xp ∂xp ∂xp ∂ 2 xp

+ + +
2 ∂xp ∂xp ∂z l ∂z m ∂z k ∂z m ∂z l ∂z k ∂z k ∂z m ∂z l ∂z m ∂z k ∂z l
∂ 2 xp ∂xp ∂ 2 xp ∂xp ∂z h ∂z m ∂xp ∂ 2 xp ∂z h ∂ 2 xp

− m k l − m l k = = = Γhkl .
∂z ∂z ∂z ∂z ∂z ∂z ∂xp ∂xp ∂z m ∂z k ∂z l ∂xp ∂z k ∂z l
∂z i ∂z m ∂xq ∂xq ∂z m ∂xq ∂z m
12. First, we remark that g im gik = = δ pq = = δmk , and
∂xp ∂xp ∂z i ∂z k ∂xp ∂z
k ∂z k 
1 ∂g mj ∂g mh ∂g jh
similarly g im gij = δmj . Then, Γijh gik +Γikh gji = g im + − gik +
    2
 ∂z h ∂z j  ∂z m
∂gmk ∂gmh ∂gkh 1 ∂gmj ∂gmh ∂gjh
g im h
+ k
− m gji = g im gik + − m +
∂z ∂z ∂z  2  ∂z h ∂z j ∂z
∂g mk ∂g mh ∂g kh 1 ∂g mj ∂g mh ∂g jh
g im gij h
+ k
− m = δmk h
+ j
− m +
 ∂z ∂z ∂z 2 ∂z ∂z ∂z 
∂gmk ∂gmh ∂gkh 1 ∂gkj ∂gkh ∂gjh ∂gjk ∂gjh ∂gkh
δmj + − m = + − + + −
∂z h ∂z k ∂z 2 ∂z h ∂z j ∂z k ∂z h ∂z k ∂z j
∂gjk
= .
∂z h
13. i) gρρ = gzz = 1, gθθ = ρ2 and the other components are null ⇒ g ρρ = g zz = 1, g θθ =
1

ρ2

173
 
1 ρm ∂gmθ ∂gmθ ∂gθθ 1 ∂gθθ
Γρθθ = g + − m = − g ρρ = −ρ,
2  ∂θ ∂θ ∂z  2 ∂ρ
1 ∂gmρ ∂gmθ ∂gρθ 1 ∂gθθ 1
Γθρθ = g θm + − m = g θθ = , the other Γkij are null.
2 ∂θ ∂ρ ∂z 2 ∂ρ ρ
1 1
ii) grr = 1, gφφ = r2 sin2 φ, gθθ = r2 ⇒ g rr = 1, g φφ = 2 ,g
θθ
= 2 and the
r2 sin φ r
other components
 are null ⇒ 
φ 1 φm ∂g mφ ∂gmr ∂gφr 1 ∂gφφ 1
Γφr = g + − m = g φφ = ,
2  ∂r ∂φ ∂z  2 ∂r r
1 ∂gθm ∂grm ∂gθr 1 ∂gθθ 1
Γθθr = g θm + − m = g θθ = ,
2  ∂r ∂θ ∂z  2 ∂r r
1 ∂g φm ∂g φm ∂g φφ 1 ∂g φφ
Γrφφ = g rm + − m
= − g rr = −r,
2  ∂φ ∂φ ∂z  2 ∂r
1 ∂gθm ∂gθm ∂gθθ 1 ∂gθθ
Γrθθ = g rm + − m = − g rr = −r sin2 φ,
2  ∂θ ∂θ ∂z  2 ∂r
1 ∂g θm ∂g φm ∂g θφ 1 ∂g θθ
Γθθφ = g θm + − m = g θθ = cot φ,
2  ∂φ ∂θ ∂z  2 ∂φ
1 ∂gθm ∂gθm ∂gθθ 1 ∂gθθ
Γφθθ = g φm + − m = − g φφ = − sin φ cos φ,
2 ∂θ ∂θ ∂z 2 ∂φ
Γφrφ = Γφφr , Γθrθ = Γθθr , Γθφθ = Γθθφ , the other Γkij are null.

c2 2
iii) g11 = g22 = (cosh 2z 1 − cos 2z 2 ) ⇒ g 11 = g 22 = 2 and the
2 c (cosh 2z − cos 2z 2 )
1
other components are null ⇒
sinh 2z 1
 
1 1 1m ∂g 1m ∂g 1m ∂g 11 1 11 ∂g11
Γ11 = g 1
+ − m = g = ,
2  ∂z ∂z 1 ∂z  2 ∂z 1 cosh 2z 1 − cos 2z 2
1 ∂g1m ∂g2m ∂g12 1 11 ∂g11 sin 2z 2
Γ112 = g 1m + − = g = ,
2 ∂z 2 ∂z 1 ∂z m 2 ∂z 2 cosh 2z 1 − cos 2z 2
and Γ212 = Γ221 = Γ111 = −Γ122 , Γ222 = Γ112 = Γ121 = −Γ211 .

14. Applying Eq. (6.27), we get:


     
∂ ρk ∂f ρ jk ∂f ∂ θk ∂f θ jk ∂f ∂ zk ∂f
i) ∆f = g + Γρj g + g + Γθj g + g +
∂ρ ∂z k ∂z k ∂θ ∂z k ∂z k ∂z ∂z k
∂f ∂ ρρ ∂f ∂f ∂ ∂f ∂f ∂f ∂ ∂f
Γzzj g jk k = g + Γρρθ g θθ + g θθ + Γθθθ g θθ + Γθθρ g ρρ + g zz =
∂z ∂ρ ∂ρ ∂θ ∂θ ∂θ ∂θ ∂ρ ∂z ∂z
∂ 2f 1 ∂ 2f 1 ∂f ∂ 2f 1 1
+ + + = (ρ f ,ρ ) ,ρ + f,θθ + f,zz , which is the same already
∂ρ2 ρ2 ∂θ2 ρ ∂ρ ∂z 2 ρ ρ2
found in Section 5.6.
     
∂ rk ∂f r jk ∂f ∂ θk ∂f θ jk ∂f ∂ φk ∂f
ii) ∆f = g + Γrj g + g + Γθj g + g +
∂r ∂z k ∂z k ∂θ ∂z k ∂z k ∂φ ∂z k
∂f ∂ ∂f ∂f ∂ ∂f ∂f ∂ ∂f ∂f
Γφφj g jk k = g rr +Γθθr g rk k + g θθ +Γθθφ g φk k + g φφ +Γφφr g rk k =
∂z ∂r ∂r ∂z ∂θ ∂θ ∂z ∂φ ∂φ ∂z
∂ 2f 1 ∂ 2f 2 ∂f 1 ∂f 1 ∂ 2f
+ 2 2 + + cot φ 2 + =
∂r2 r sin φ ∂θ2 r ∂r r ∂φ r2 ∂φ2

174
 
1 2 1 f,θθ
2
(r f,r ),r + 2 + (f,φ sin φ),φ , which is to be compared to the one
r r sin φ sin φ
given in Section 5.7.

np ∂g np n rp p nr np ∂z n ∂z p
15. i) By Eqs. (6.11), (6.25) and (6.29), g;h = + Γhr g + Γhr g , g = ,
∂z h ∂xk ∂xk
∂z n ∂ 2 xm p ∂z p ∂ 2 xt np ∂ ∂z n ∂z p ∂z n ∂ ∂z p
Γnhr = , Γhr = ⇒ g ;h = + +
∂xm ∂z h ∂z r ∂xt ∂z h ∂z r ∂xk ∂xh ∂xk ∂xk ∂xk ∂z h
n r p p n r n p p n
∂z ∂ ∂xm ∂z ∂z ∂z ∂ ∂xt ∂z ∂z ∂z ∂ ∂xm ∂z ∂z ∂ ∂xt ∂z
+ = + =
∂xm ∂z r ∂z h ∂xq ∂xq ∂xt ∂z r ∂z h ∂xs ∂xs ∂xm ∂z h ∂xq ∂xq ∂xt ∂z h ∂xs ∂xs
∂z n ∂xm
0 because, e.g., h
= δnh , = δmq etc., so their derivatives are null.
∂z ∂xq

∂ 2 xk ∂xk ∂xk ∂ 2 xk
ii) By Eqs. (6.3), (6.25) and (6.30), gnp;h = h n p + n h p −
∂z ∂z ∂z ∂z ∂z ∂z
∂z r ∂ 2 xm ∂xq ∂xq ∂z r ∂ 2 xt ∂xs ∂xs ∂ 2 xk ∂xk ∂xk ∂ 2 xk
− = h n p + n h p−
∂xm ∂z p ∂z h ∂z n ∂z r ∂xt ∂z n ∂z h ∂z p ∂z r ∂z ∂z ∂z ∂z ∂z ∂z
∂ 2 xm ∂xq ∂ 2 xt ∂xs ∂ 2 xk ∂xk ∂xk ∂ 2 xk ∂ 2 xq ∂xq
δqm p h n − δst n h p = h n p + n h p − p h n −
∂z ∂z ∂z ∂z ∂z ∂z ∂z ∂z ∂z ∂z ∂z ∂z ∂z ∂z ∂z
∂ 2 xs ∂xs
= 0.
∂z n ∂z h ∂z p

Chapter 7
1. It is sufficient to pose x1 = u, x2 = v, x3 = f (u, v) ⇒ p(u, v) defines a surface
because as f (u, v)
 is smooth, p(u, v) is also smooth, and because the Jacobian is
1 0
[J] =  0 1 , then rank[J] = 2.
f,u f,v

 x1 = cosh u cos v,
2. Catenoid: f (u, v) : x2 = cosh u sin v, Meridians: v = const.; if, for example,
x3 = u;


 x1 = cosh u,
v=0⇒ x2 = 0, is a catenary in the plane (x1 , x3 ).
x = u,

 3   
 sinh u cos v   − cosh u sin v 
cosh2 u
 
0
f,u = sinh u sin v , f,v = cosh u cos v ⇒g= .
0 cosh2 u
1 0
   
   
 − cosh u cos v  1  − cos v 
f,u × f,v = − cosh u sin v , |f,u × f,v | = cosh2 u ⇒ N = − sin v ;
cosh u 
sinh u cosh u sinh u
  
   
 − sinh u sin v   cosh u cos v 
f,uv = f,vu = sinh u cos v , f,uu = cosh u sin v ,
0 0
   

175
 
 − cosh u cos v   
1
−1 0
f,vv = − cosh u sin v ⇒B= ⇒K=− 4 .
0 1 cosh u
0
 

 x1 = sin u cos v,

3. Pseudo-sphere: f (u, v) : x2 = sin u sin v,  Meridians: v = const.; if, for
 x3 = cos u + ln tan u .

 2
 x1 = sin u,

example, v = 0 ⇒ x2 = 0,  is a tractrix in the plane (x1 , x3 ).
 x3 = cos u + ln tan u ,


  2

 cos u cos v 


 − sin u sin v 
 
cos2 u

f,u = cos u sin v , f,v = sin u cos v ⇒ g =  sin2 u 0 .


1 0 sin2 u
 − sin u + 0
 
  
 sin u   
 − cos2 u cos v  1  − cos2 u cos v 
f,u ×f,v = − cos2 u sin v , |f,u ×f,v | = | cos u| ⇒ N = − cos2 u sin v ;
| cos u| 
sin u cos u  sin u cos u
  

 − sin u cos v 
 
   − cos u sin v 
f,uu = − sin u sin v , f,uv = f ,vu = cos u cos v ,
 − cos u − cos u 
  
0

sin2 u
2
 
  cos u
 − sin u cos v   − sin u| cos u| 0 
f,vv = − sin u sin v ⇒B= 2
 ⇒ K = −1.
 cos u sin u 
0 0
 
| cos u|
4. Cone: f (u, v) = vγ(u) ⇒ f,u = vγ ′ , f,v = γ → N ̸= o ⇐⇒ v ̸= 0 and γ ̸= αγ ′ , i.e.
everywhere except at the apex of the cone and on straight lines tangent to γ.
5. The most general equation of the hyperbolic hyperboloid is
 x1 = a(cos u − v sin u),
f (u, v) : x2 = b(sin u + v cos u), a, b, c ∈ R ⇒ f (u, v) = γ(u) + vλ(u), with
x3 = c v,

 = −a sin ue1 + b cos ue2 + ce3 ⇒ f (u, v) is a ruled


γ(u) = a cos ue1 + b sin ue2 , λ(u)
 x1 = a(cos u0 − v sin u0 ),
surface. Fixing u = u0 ⇒ x2 = b(sin u0 + v cos u0 ), equation of a bundle of
x3 = c v, 

 x1 = a(cos u0 − v sin u0 ),
straight lines belonging to f (u, v), as well as x2 = b(sin u0 + v cos u0 ), The angle
x3 = −c v.

a2 sin2 u0 + b2 cos2 u0 − c2
formed by two straight lines of the two sets is θ = arccos 2 2 .
a sin u0 + b2 cos2 u0 + c2

 x1 = u,
6. x3 = x1 x2 ; setting u = x1 , v = x2 ⇒ f (u, v) : x2 = v, is of the type
x3 = u v,

f (u, v) = γ(u) + vλ(v), with γ(u) = ue1 , λ(u) = e2 + ue3 .

176
 
 x1 = u0 ,  x1 = u,
The straight lines x2 = v, and x2 = v0 , belong of course to f (u, v);
x3 = u0 v, x3 = u v 0 ,
 
u0 v0 π
they form the angle θ = arccos p ; θ = ⇐⇒ 0 = v0 = 0, i.e. at
(1 + u20 )(1 + v02 ) 2
(0, 0, 0).
7. i) γ(u) = (cos u, sin u, −1), λ(u) = (cos u, sin u, 1) ⇒ f (u, v) = (cos u, sin u, 2v −
1), of the form f (u, v) = γ 1 (u) + vλ1 (u), with γ 1 (u) = (cos u, sin u, −1), λ1 (u) =
(0, 0, 2) = const. ⇒ f (u, v) is a cylinder whose Cartesian equation is x21 + x22 = 1.
ii) γ(u) = (sin u, − cos u, −1), λ(u) = (− sin u, cos u, 1) ⇒
f (u, v) = (2v − 1)(− sin u, cos u, 1), of the form f (u, v) = γ 2 (u) + vλ2 (u), with
γ 2 (u) = (sin u, − cos u, −1), λ1 (u) = (−2 sin u, 2 cos u, 2) = −2γ 2 (u) ⇒ f (u, v) is a
cone whose Cartesian equation is x21 + x22 = x23 .

x1 = (1 − v) cos(u − α) + v cos(u + α),
iii) ⇒
 x2 = (1 − v) sin(u − α) + v sin(u + α), x3 = −(1 − v) + v,
x1 = cos u cos α − sin u sin α(2v − 1),
x2 = sin u cos α + cos u sin α(2v − 1), x3 = 2v −1.
sin α  x1 = cos α(cos u − w sin u),

Change in parameter w = (2v − 1) ⇒ x2 = cos α(sin u + w cos u), ⇒
cos α  x3 = w cos α ,

 sin α
 x1 = a(cos u − w sin u), cos α
x2 = a(sin u + w cos u), a = cos α, c = , which is the parametric equa-
sin α
x3 = cw,

x2 x2 x2
tion of a hyperbolic hyperboloid with Cartesian equation 21 + 22 − 23 = 1 →
a a c
x21 + x22 x23
− = 1.
cos2 α cot2 α
8. i) Sphere of radius R : x21 + x22 + x23 = R2 ⇒ using the spherical coordinates
θ = u, φ = v for expressing the xi s, we get the parametric equation f (u, v) :
 x1 = R cos θ sin φ,
x2 = R sin θ sin φ, ⇒ f,u = R(− sin u sin v, cos u sin v, 0),
x3 = R cos φ,

 2 2 
R sin v 0
f,v = R(cos u cos v, sin u cos v, − sin v) ⇒ g = .
0 R2
ii) If w = af,u + bf,v ∈ Tp Σ, I(w) = w · gw = R2 (a2 sin2 v + b2 ).
Rθ Rπ√ Rθ Rπ√
iii) A = θ12 0 det gdu dv = θ12 0 R4 sin2 vdu dv = 2R2 (θ2 − θ1 ).
R


 x1 = √ cos t,
2



π  R
iv) Parallel: putting u = t, v = , γ(t) : x2 = √ sin t, ⇒
4  2
R


 x3 = √ ,


2

177
 R

 x1 = − √ sin t,

 2 du dv
γ ′ (t) : R ⇒ γ ′
(t) = f ,u + f,v = f,u ⇒ in the natural basis of
x2 = √ cos t, dt dt
2




x3 = 0,
Tp Σ, w = (1, 0) is the tangent vector to the parallel γ(t) ⇒
R2 Rθ p R
I(w) = w · gw = R2 sin2 v = ⇒ ℓ = θ12 I(w)dt = √ (θ2 − θ1 ).
2 2
cos2 v sin2 v sinh2 u 1 sinh2 u cosh2 u
9. i) x21 + x22 + x23 = + + = + = =1→
cosh2 u cosh2 u cosh2 u cosh2 u cosh2 u cosh2 u
Cartesian equation of a sphere of centre (0, 0, 0) and radius R = 1.

cos(v0 + bt)
x1 = ,






 cosh(u0 + at)
u = u0 + a t,
 sin(v0 + bt)
ii) Straight line in Ω : ⇒ curve on Σ : γ(t) : x2 = ,
v = v0 + b t; 
 cosh(u0 + at)

 sinh(u0 + at)
 x3 = ,


cosh(u0 + at)
du dv
or also γ(t) = f (u(t), v(t)) ⇒ γ ′ (t) = f,u + f,v = af,u + bf,v .
dt dt
x2
Meridians: setting v = const. = v̂ ⇒ µ(u) = f (u, v̂); in fact = tan v̂ = const. →
x1
equation of a vertical plane. Tangent to the meridian µ(u) : µ′ (u) = f,u ⇒ in the
I(γ ′ , µ′ )
natural basis {f,u , f,v }, γ ′ (t) = (a, b), µ′ (u) = (1, 0); cos θ = p .
I(γ ′ )I(µ′ )
   
sinh u sinh u 1 sin v cos v
f,u = − cosh v , − sin v , , f,v = − , ,0 ⇒
cosh2 u cosh2 u cosh2 u cosh u cosh u
1 ′ ′ ′ ′ a ′ ′ ′ a2 + b 2
g= I ⇒ I(γ , µ ) = γ · gµ = , I(γ ) = γ · gγ = ,
cosh2 u cosh2 u cosh2 u
1 a
I(µ′ ) = µ′ · gµ′ = 2 ⇒ cos θ = √ = const. ⇒ γ(t) is a loxodromic
cosh u a2 + b 2
line on the sphere.

 x1 = φ(u) cos v,
10. i) Catenoid f (u, v) : x2 = φ(u) sin v, with φ(u) = cosh u, ψ(u) = u ⇒
x3 = ψ(u).

φ (u) = sinh u, φ (u) = cosh u, ψ ′ (u) = 1, ψ ′′ (u) = 0 ⇒
′ ′′

f,u = (sinh u cos v, sinh u sin v, 1), f,v = (− cosh u sin v, cosh u cos v, 0).

ii) g = cosh2 uI.

iii) f,u× f,v = (− cos v, − sin v, − sinh


 u), |f,u × f,v | = cosh u ⇒
cos v sin v sinh u
N= − ,− ,− .
cosh u cosh u cosh u
f,uu = (cosh u cos v, cosh
 u sin v,  0), f,uv = f,vu = (− sinh u sin v, sinh u cos v, 0),
−1 0
f,vv = −f,uu ⇒ B = .
0 1

178
1 1
iv) g−1 = −1
2 I ⇒ X = g B = B.
cosh u cosh2 u
v) Let w = (a, b) ∈ Tp Σ ⇒ I(w) = w · gw = cosh2 u(a2 + b2 ).
vi) II(w) = w · Bw = b2 − a2 .

 x1 = v cos u,
11. i) Helicoid f (u, v) : x2 = v sin u, ⇒ f,u = (−v sin u, v cos u, 1), f,v = (cos u, sin u, 0).
x3 = u,

 
1 + v2 0
ii) g = .
0 1
√ 1
iii) f,u ×f,v = (− sin u, cos u, −v), |f,u ×f,v | = 1 + v2 ⇒ N = √ (− sin u, cos u, −v).
1 + v2
f,uu = (−v cos u, −v sin
 u, 0), f,uv = f,vu = (− sin u, cos u, 0), f,vv = (0, 0, 0) ⇒
1 0 1
B= √ .
1+v 2 1 0
1
 
" 1 # 0 3
−1 0  −1 (1 + v 2 ) 2 
iv) g = 1 + v2 .⇒X=g B= 1
.
0 1
 
1 0
(1 + v 2 ) 2
v) Let w = (a, b) ∈ Tp Σ ⇒ I(w) = w · gw = (1 + v 2 )a2 + b2 .
2ab
vi) II(w) = w · Bw = √ .
1 + v2
1
12. i) Catenoid (see Exercise 2): K = − < 0 ∀u ⇒ hyperbolic points.
cosh4 u
det B 1
ii) Helicoid (see Exercise 11): K = = − < 0 ∀v ⇒ hyperbolic
det g (1 + v 2 )2
points.

 x1 = R cos v,
13. Parametric equation of a circular cylinder of radius R → f (u, v) : x2 = R sin v,
x3 = u,

with u = z, v = θ of a system of cylindrical  coordinates. Referring to Eq. (7.7),
u′′ = 0, u(t) = αt + α1 ,
φ(u) = R, ψ(u) = u ⇒ Eq. (7.32) is ′′ ⇒ with
v = 0, v(t) = βt + β1 ,
α, α1 , β, β1 = const. If α1  = β1 = 0, we get thegeodesic γ(t) passing through
 x1 = R cos v(t),  x1 = R cos(βt),
(R, 0, 0) for t = 0 ⇒ γ(t) : x2 = R sin v(t), ⇒ x2 = R sin(βt), ⇒ equation
x3 = u(t), x3 = αt,
 
of a helix if α, β ̸= 0, of a circle (cross section) if α = 0, β ̸= 0 and of a straight line
on the cylinder (generatrix) if α ̸= 0, β = 0.

179

You might also like