Understanding Vectors and Scalars in Maths
Understanding Vectors and Scalars in Maths
Term 1
Vectors
Vectors are one of the fundamental objects in the development of any physical theory. Vectors allow us to
track the location of objects such as space craft, asteroids, fundamental particles and interacting chemical
species. But vectors also allow use to characterise non-localised physical phenomena such as electric
and magnetic fields, examples of so called vector fields which are, in essence, continuous sets of arrows
showing the direction the lines of force of the field (an example is shown in Figure 1.1). Thus we start by
covering the fundamentals of these important mathematical objects.
1.1.1 Scalar
We start by defining something which is not a vector, a scalar this will in part serve as a means of
highlighting the special properties of vectors. A scalar is a number that (completely) describes some
physical quantity. For example:
• I have 10 fingers
• Temperature, 𝑇 , in ◦ 𝐾, or ◦ 𝐶, or ◦ 𝐹
• Density, 𝜌, in 𝑘𝑔𝑚−3
• Time, 𝑡, in seconds 𝑠
3
Figure 1.1: A (magnetic) vector field from a simulation of a magnetic flux rope emerging into the Sun’s
atmosphere. The arrows represent the vector field and they are the solution to rather complex set of partial
differential equations.
A scalar comes with a set of units, relating the number to the real world quantity at hand. Crucially there
is no notion of direction associated with this definition.
Note: A scalar can be constant, or it can depend on position, time, etc. When our scalar depends on such
a quantity, we usually indicate the dependence via brackets. For example, if the temperature depends on
time, we may denote by 𝑇 (𝑡) the temperature at time 𝑡.
The values that scalars can take are often restricted to some pre-determined set. Typically, this set is
the real numbers ℝ, but sometimes it might be the complex numbers ℂ or even the rationals ℚ. The
complex numbers are in fact completely fundamental to most physical theories owing to the fact they can
be used to describe radiation (the complex exponential). The rationals are perhaps more the preserve of
methematicians seeking to make things difficult for themselves.....
1.1.2 Vectors
Some quantities cannot be described by a single number. For example, if you are an air traffic controller
then knowing the wind speed alone is not too helpful; you also need to know the direction of the wind.
From the perspective of this course, a vector is an object with both a direction and a (scalar) magnitude
4
(or ‘length’). For example:
• Acceleration, 𝐚, in 𝑚𝑠−2
• Force, 𝐅, in Newtons 𝑁
A vector also comes with a set of units which correspond to its (scalar) magnitude.
Note: In order to distinguish a vector from a scalar we often use bold type 𝐯 in printed text, and either
underline 𝑣 or place an arrow above 𝑣⃗ in handwritten text. As before, a vector can depend on position,
time, etc, and we use brackets to convey this dependence. For example, the force exerted on a charged
particle by a magnetic field may depend on the position of the particle, and we may denote by 𝐁(𝑥) the
force exerted when at position 𝑥.
There is no such thing as an absolute position. For example Durham is due south of Edinburgh but north
of London. Its not north and south or even exclusively either. Instead its position is described as relative
to some other location. So, If we want to describe the position of some body we first need we need frame
of reference. This frame of reference is called the “The origin" 𝑂 in space. Geometrically, we can then
think of every other point in space as being displaced from 𝑂 by some distance in some direction. We
may view the directed line from 𝐎 to a point 𝐴 as a position vector 𝐩(𝑎).
In order to make this interpretation more explicit, we will often denote this vector by 𝐩(𝑎) = 𝑂𝐴
⃖⃖⃖⃖⃖⃗ and refer
More generally, we will represent the directed line beginning at the point 𝐵 and ending at the point 𝐶 by
5
the vector 𝐵𝐶.
⃖⃖⃖⃖⃖⃗
The length of the directed line from 𝐵 to 𝐶 is precisely the magnitude of the vector 𝐵𝐶
⃖⃖⃖⃖⃖⃗ and is de-
noted |𝐵𝐶|.
⃖⃖⃖⃖⃖⃗ More generally, the magnitude of any vector 𝐯 is denoted by |𝐯| and is, of course, a scalar.
As a special case, if |𝐯| = 1 we shall call 𝐯 a unit vector and will sometimes convay this fact by giving
the vector a ‘hat’ 𝐯.
̂
vation is that the same vector can represent the displacement between two distinct pairs of vectors. For
example, in the following diagram the vectors 𝐚 and 𝐛 have the same direction and length.
that both the direction and the magnitude of a vector are invariant under translation, aka movement
of the origin.
Having said that, if we have 𝐚 = 𝐛 we can still say 𝐚 is not the same as 𝐛! This is of course the case if
we know 𝑂 and 𝐵 represent different points i.e. 100𝑘𝑚 north of London is not the same as as 100𝑘𝑚
north of Birmingham (although some folk I have met in the past would have told me everything north of
Watford is "the north" ;) ).
If the origin 𝑂 is chose and fixed and there is some vector 𝐩(𝑡) which is changing in time (i.e. a rocket
leaving earth) then if we differentiate 𝐩(𝑡) we will obtain a velocity vector 𝐯 = d𝐩
d𝑡
which would actually
not depend on the origin, its fixed and a constant so disappears when we differentiate (we will see this
6
more formally in chapter 2).
So velocities and indeed accelerations don’t depend on what specific location is chosen for the origin, this
is seen as a crucial point as reasonable physical theories should not depend on an entirely arbitrary choice
of origin.
A vector can also be rotated in space, and obviously this does change its direction. For example, if the
wind were blowing directly to our right at 5𝑚𝑠−1 and we turned 90◦ (or 𝜋∕2 radians) anticlockwise, then in
our new frame of reference the wind would then be blowing directly towards our face. However, notice the
very important property than the wind speed would still be 5𝑚𝑠−1 ; the magnitude of a vector is invariant
under rotation. We shall see later that this actually implies the body is accelerating even though its speed
is fixed! Notice also that any such rotation would not affect scalar measurements such as the temperature
where we were standing.
Comment (not covered in this course): There are other quantities of physical significance that are not
described by scalars or vectors, but require even more numbers to specify them. Examples of these
are, moments of inertia, stress, and strain, which each need six pieces of information to specify them
completely. These are called tensors and you may well meet them in the future.
From a mathematical viewpoint, vectors are defined through their algebra. There are two fundamental
operations of vector algebra; vector addition and scalar multiplication.
another vector 𝐛, then we will end up at some other point in space, say 𝐶. But, the point 𝐶 already has its
own position vector 𝐜 = 𝑂𝐶.
⃖⃖⃖⃖⃖⃗ We therefore define the sum of two vectors 𝐚 and 𝐛 to be the vector 𝐜 and
we write
𝐚+𝐛=𝐜 [vector addition inputs two vectors and outputs a third vector]
7
Pictorially, we add vectors by placing them “head-to-tail".
From the diagram we can deduce the first important property of vector addition:
Another important property is that when adding more than two vectors, the order of addition does not
matter. See if you can draw a diagram in order to verify the following:
Next, let’s consider the origin 𝑂 itself. The position vector of the origin must represent zero displacement
in every direction (since the origin is not displaced from itself). We call this vector the zero vector and
denote it by 𝟎. The zero vector has the unique property among all vectors that its magnitude |𝟎| = 0. It
also satisfies the following rule under addition:
8
The inverse vector: a.k.a going back on yourself
result, if 𝐚 = 𝑂𝐴
⃖⃖⃖⃖⃖⃗ then we define the vector −𝐚 by −𝐚 = 𝐴𝑂.
⃖⃖⃖⃖⃖⃗
The vector −𝐚 has the same magnitude as the vector 𝐚, but has the opposite direction.
Note: Given two vectors 𝐚 and 𝐛, for notational convenience we will usually drop the parentheses and
just write 𝐚 − 𝐛 for 𝐚 + (−𝐛).
Beginning at the origin 𝑂, let’s say we want to travel to a point 𝐵 which lies exactly half way between
the origin and some given point 𝐴. If 𝐴 and 𝐵 have position vectors 𝐚 and 𝐛 respectively, then 𝐛 has the
same direction as 𝐚, but is of only half the length, and we write 𝐛 = 12 𝐚. More generally, we can combine
any real scalar 𝜆 in ℝ and any vector 𝐚 to obtain a new vector
The vector 𝐛 is the vector with the same direction as 𝐚 if 𝜆 > 0 (and the opposite direction if 𝜆 < 0), but
with magnitude scaled by |𝜆|.
9
In the case that 𝜆 = 0, we output the zero vector; that is, 0𝐯 = 𝟎 for any vector 𝐯.
Note: We say that the vectors 𝜆𝐚 are parallel to 𝐚 for all 𝜆 ≠ 0 (the zero vector has no direction so is not
parallel to anything). Sometimes we may distinguish those with 𝜆 < 0 by saying they are anti-parallel to
𝐚.
(𝑖) We have 1𝐚 = 𝐚
In this course we will usually restrict our attention to the ‘real’ vector spaces ℝ2 (two-dimensional Eu-
clidean space) and ℝ3 (three-dimensional Euclidean space). Our scalars will be taken from the real num-
bers ℝ. [If you are doing physics, you may soon encounter spacetime, which is also a vector space; here
each vector has three Euclidean coordinates and a fourth component, time.]
For the space ℝ3 let us consider the Cartesian coordinates defined by the 𝑥−axis, 𝑦−axis and 𝑧−axis
respectively. The origin 𝑂 has coordinates (0, 0, 0).
10
If a point 𝑃 in ℝ3 has coordinates (𝑥, 𝑦, 𝑧) then to move from 𝑂 to 𝑃 we travel the distance 𝑥 along the
𝑥−axis, the distance 𝑦 along the 𝑦−axis, and then the distance 𝑧 along the 𝑧−axis. We write its position
vector as the column, consisting of these three scalars
⎛𝑥⎞
⎜ ⎟
⃖⃖⃖⃖⃖⃗ = 𝐩 = ⎜𝑦⎟ .
𝑂𝑃
⎜ ⎟
⎜𝑧⎟
⎝ ⎠
The size of the final displacement (i.e., the distance from 𝑂 to 𝑃 ) is precisely the magnitude of the vector
𝐩 and via trigonometry is given by
√
|𝐩| = 𝑥2 + 𝑦2 + 𝑧2 .
( 0
) ( 1
)
Example [Link] Suppose that 𝐚 = 3 and 𝐛 = −1 . Then
−4 2
( 1
) ( 0
) √ √
𝐚+𝐛= 2 ; −3𝐚 = −9 ; and |𝐚| = 02 + 32 + (−4)2 = 25 = 5.
−2 12
11
1.3.1 Basis vectors
(𝑥)
Notice that we can write any vector 𝐩 = 𝑦
𝑧
in ℝ3 in a unique way as the combination
⎛𝑥⎞ ⎛1 ⎞ ⎛0 ⎞ ⎛0 ⎞
⎜ ⎟ ⎜ ⎟ ⎜ ⎟ ⎜ ⎟
⎜𝑦⎟ = 𝑥 ⎜0⎟ + 𝑦 ⎜1⎟ + 𝑧 ⎜0⎟ = 𝑥𝐢 + 𝑦𝐣 + 𝑧𝐤.
⎜ ⎟ ⎜ ⎟ ⎜ ⎟ ⎜ ⎟
⎜𝑧⎟ ⎜0 ⎟ ⎜0 ⎟ ⎜1 ⎟
⎝ ⎠ ⎝ ⎠ ⎝ ⎠ ⎝ ⎠
The vectors in the set {𝐢, 𝐣, 𝐤} are called the Cartesian Basis vectors; they are unit vectors in the direction
of each Cartesian coordinate axis. We call the scalars in the set {𝑥, 𝑦, 𝑧} the components (or coordinates)
of the vector 𝐩 with respect to the basis {𝐢, 𝐣, 𝐤}. It is easy to see that given 𝐚 = 𝑎1 𝐢 + 𝑎2 𝐣 + 𝑎3 𝐤 and
𝐛 = 𝑏1 𝐢 + 𝑏2 𝐣 + 𝑏3 𝐤 we can rephrase the vector operations above as
There is actually nothing too special about the Cartesian basis vectors {𝐢, 𝐣, 𝐤}.
{ }
Definition [Link] In a vector space 𝑉 , a set of vectors 𝐞1 , 𝐞2 , … , , 𝐞𝐷 is a basis for 𝑉 if every
vector 𝐮 in 𝑉 can be expressed uniquely in the form
∑
𝐷
𝐮= 𝑢𝑖 𝐞𝑖 ,
𝑖=1
The scalars 𝑢1 , 𝑢2 , … , 𝑢𝐷 are the components (or the coordinates) of the vector 𝐮 with respect to the
{ }
basis 𝐞1 , 𝐞2 , … , , 𝐞𝐷 . Each vector space can have many different bases.
{( ) ( )}
Example [Link] (a) Let us consider the case 𝑉 = ℝ2 . It is clear that the set 10 , 01 is a basis
( ) ( ) ( ) ( )
for ℝ2 since every position vector 𝑥𝑦 can be written as 𝑥𝑦 = 𝑥 10 + 𝑦 01 . To see this is unique,
( ) ( ) ( )
let us suppose we could also write 𝑥𝑦 = 𝑤 10 + 𝑧 01 for some scalars 𝑤 and 𝑧. Then
12
By comparing 𝑥− and 𝑦− coefficients we see that we must have had 𝑤 = 𝑥 and 𝑧 = 𝑦.
{( 1 ) ( 1 )}
(b) The above basis is not particularly special. We can also show that the set 0 , 1 is also a
basis for ℝ2 . To see this, notice that
In fact, one can show that any two vectors in ℝ2 that are not scalar multiples of one another are a basis
for ℝ2 .
More generally, one can show that any two distinct bases for the same vector space must share the same
cardinality (that is, there is the same number of vectors in each set). The number of vectors in a basis
for a vector space is called the dimension of the vector space. The dimension can be interpreted as the
number of independent directions in the vector space. For instance, in ℝ3 any three vectors that do not lie
in the same plane form a basis. The Cartesian basis vectors {𝐢, 𝐣, 𝐤} are just one example.
Note: Formally, we say that a given set of directions (i.e., vectors, say 𝐚, 𝐛, 𝐜) are (linearly) independent
if they cannot be scaled and added to form a ‘complete loop’. That is, the only way we can have 𝛼𝐚 +
𝛽𝐛 + 𝛾𝐜 = 𝟎 for scalars 𝛼, 𝛽, 𝛾 is if 𝛼 = 𝛽 = 𝛾 = 0.
We have seen that we can add and subtract vectors, but can we multiply vectors? The answer is yes, but.
Unlike multiplying numbers where thee is only one possible definition, vectors have multiple components
and so there are multiple possible definitions. Do we produce a single number, i.e.a operation which takes
two vectors and returns a scalar (rather than another vector), or some multiplication which returns a new
vector. We will start with the first case.
13
We will approach the construction of our operation from a geometric viewpoint, and will construct it in
such a way that, like vector addition, it is invariant under rotation. Given that we are to produce a scalar
from two vectors, say 𝐚 and 𝐛, our candidate operation should be depend on intrinsic (and geometrically
invariant) scalar quantities of the vectors themselves. As we have observed, the magnitudes |𝐚| and |𝐛|
of the vectors are invariant under rotation. But there is another natural rotation-invariant scalar quantity
associated with any pair of vectors - the (planar) angle 𝜃 between them.
Notice that the angle is periodic (an angle of 2𝜋 radians - or 360◦ - simply corresponds to the vector being
parallel again). For this reason, it is more natural to consider the cosine or sine of the angle.
Definition [Link] Given two vectors 𝐚 and 𝐛 we define the scalar (or dot) product of the two vectors
to be
𝐚 ⋅ 𝐛 = |𝐚| |𝐛| cos(𝜃),
Clearly, the scalar product vanishes whenever one of the vectors is the zero vector. However, a curious
special case occurs when the cosine of the angle is zero (which occurs when 𝜃 = 𝜋∕2). Here, the scalar
product also vanishes. We say that two non-zero vectors 𝐚 and 𝐛 are orthogonal if 𝐚 ⋅ 𝐛 = 0.
More generally this angle means the product is largest in magnitude when the two vectors are aligned so
it quantifies a degree of alignment. Indeed if we choose unit vectors it literally quantifies that.
14
1.4.1 Properties of the scalar product
Commutivity
We can very quickly deduce some useful properties of the scalar product. For example, it is immediate
that for any vectors 𝐚 and 𝐛 we have
(𝐷2) 𝐚 ⋅ 𝐚 = |𝐚|2
Linearity
In search of further properties, we now make an important observation. Consider the following diagram:
As we can see, the scalar product 𝐚 ⋅ 𝐛 can be interpreted as “the component of 𝐛 in the direction of 𝐚”
multiplied by |𝐚|. We can use this observation to deduce a very important property of the scalar product.
Given three vectors 𝐚, 𝐛 and 𝐜 let 𝜃 be the angle between 𝐚 and 𝐜, let 𝜑 be the angle between 𝐛 and 𝐜,
and let 𝛾 be the angle between 𝐚 + 𝐛 and 𝐜. The component of 𝐚 + 𝐛 in the direction of 𝐜 is therefore
|𝐚 + 𝐛| cos(𝛾).
15
It is clear from the above diagram that the component of 𝐚 + 𝐛 in the direction of 𝐜 is also the sum of the
component of 𝐚 in the direction of 𝐜 and the component of 𝐛 in the direction of 𝐜. That is; we have
Multiplying through by |𝐜| gives the first of the following two linearity properties of the scalar product.
(𝐿1) (𝐚 + 𝐛) ⋅ 𝐜 = 𝐚 ⋅ 𝐜 + 𝐛 ⋅ 𝐜
We can use these properties to derive an algebraic definition of the dot product. For example, consider
the vector space ℝ3 with Cartesian basis {𝐢, 𝐣, 𝐤}. Since each basis vector is of unit length and cos(0) = 1
we have
By properties (𝐿1) and (𝐿2), it follows that for any two vectors 𝐚 = 𝑎1 𝐢 + 𝑎2 𝐣 + 𝑎3 𝐤 and 𝐛 = 𝑏1 𝐢 + 𝑏2 𝐣 + 𝑏3 𝐤
that
𝐚 ⋅ 𝐛 = (𝑎1 𝐢 + 𝑎2 𝐣 + 𝑎3 𝐤) ⋅ (𝑏1 𝐢 + 𝑏2 𝐣 + 𝑏3 𝐤)
= 𝑎1 𝐢 ⋅ (𝑏1 𝐢 + 𝑏2 𝐣 + 𝑏3 𝐤) + 𝑎2 𝐣 ⋅ (𝑏1 𝐢 + 𝑏2 𝐣 + 𝑏3 𝐤) + 𝑎3 𝐤 ⋅ (𝑏1 𝐢 + 𝑏2 𝐣 + 𝑏3 𝐤)
= 𝑎1 𝑏1 𝐢 ⋅ 𝐢 + 𝑎1 𝑏2 𝐢 ⋅ 𝐣 + 𝑎1 𝑏3 𝐢 ⋅ 𝐤 + 𝑎2 𝑏1 𝐣 ⋅ 𝐢 + 𝑎2 𝑏2 𝐣 ⋅ 𝐣 + 𝑎2 𝑏3 𝐣 ⋅ 𝐤 + 𝑎3 𝑏1 𝐤 ⋅ 𝐢 + 𝑎3 𝑏2 𝐤 ⋅ 𝐣 + 𝑎3 𝑏3 𝐤 ⋅ 𝐤
= 𝑎1 𝑏1 + 𝑎2 𝑏2 + 𝑎3 𝑏3 .
16
Thus, the dot product is simply the sum of the products of the components.
√
Example [Link] (a) Suppose the vector 𝐚 has length 1 and the vector 𝐛 has length 2, and that the
angle between the two vectors is 𝜋∕4 radians. Then
√ √
𝐚 ⋅ 𝐛 = |𝐚| |𝐛| cos(𝜋∕4) = (1)( 2)(1∕ 2) = 1.
( 0
) ( 1
)
(b) Suppose that 𝐚 = 3 and 𝐛 = −1 . Then
−4 2
⎛0⎞ ⎛1⎞
⎜ ⎟ ⎜ ⎟
𝐚 ⋅ 𝐛 = ⎜ 3 ⎟ ⋅ ⎜−1⎟ = (0)(1) + (3)(−1) + (−4)(2) = −11.
⎜ ⎟ ⎜ ⎟
⎜−4⎟ ⎜ 2 ⎟
⎝ ⎠ ⎝ ⎠
Application 1:
Consider a rhombus; a quadrilateral whose four sides all have the same length and opposite sides are
parallel. Suppose the sides are given by the two vectors 𝐚 and 𝐛. Then, the diagonals 𝐜 and 𝐝 of the
rhombus, as indicated in the diagram below, are given by 𝐜 = 𝐚 + 𝐛 and 𝐝 = 𝐚 − 𝐛.
Consider the dot product of the two diagonals. We have by linearity that
𝐜 ⋅ 𝐝 = (𝐚 + 𝐛) ⋅ (𝐚 − 𝐛) = 𝐚 ⋅ 𝐚 − 𝐚 ⋅ 𝐛 + 𝐛 ⋅ 𝐚 − 𝐛 ⋅ 𝐛 = 𝐚 ⋅ 𝐚 − 𝐛 ⋅ 𝐛 = |𝐚|2 − |𝐛|2 = 0.
Thus, the angle between the diagonals is 𝜋∕2 and we have shown the diagonals of a rhombus are perpen-
dicular!
17
Application 2:
Consider the triangle defined by three vectors 𝐚, 𝐛 and 𝐜 and let 𝜃 be the angle between 𝐚 and 𝐛.
|𝐜|2 = |𝐚 − 𝐛|2 = (𝐚 − 𝐛) ⋅ (𝐚 − 𝐛) = 𝐚 ⋅ 𝐚 − 𝐚 ⋅ 𝐛 − 𝐛 ⋅ 𝐚 + 𝐛 ⋅ 𝐛
= |𝐚|2 + |𝐛|2 − 2𝐚 ⋅ 𝐛
= |𝐚|2 + |𝐛|2 − 2|𝐚| |𝐛| cos(𝜃).
Thus, we have derived the Cosine rule.
Before we proceed, I will ask of you a question (which we will answer shortly) if the saclar product is:
A curious observation in ℝ3 is the following. Given any two (non-zero) points 𝐴 and 𝐵 in ℝ3 there exists
a plane Π passing through the origin 𝑂 and both 𝐴 and 𝐵. Clearly, the position vectors 𝐚 = 𝑂𝐴
⃖⃖⃖⃖⃖⃗ and
18
⃖⃖⃖⃖⃖⃗ lie on this plane, and so long as 𝐚 and 𝐛 are not parallel the plane Π is unique. Indeed, this is
𝐛 = 𝑂𝐵
the plane used in order to define the angle between the two vectors and we say that Π is spanned by the
vectors 𝐚 and 𝐛. In fact, one can show that position vector 𝐩 = 𝑂𝑃
⃖⃖⃖⃖⃖⃗ of any point 𝑃 in the plane Π can be
written uniquely as a linear combination 𝐩 = 𝜆𝐚 + 𝜇𝐛 for some scalars 𝜆 and 𝜇. (You should compare
this with a basis.)
However, we can, in fact, completely specify Π with just a single vector! Let 𝐧 be a vector orthogonal to
both 𝐚 and 𝐛. Then, by properties (L1) and (L2) we also have that
so 𝐩 is also orthogonal to 𝐧. Thus, the plane Π is precisely the set of points 𝑃 in ℝ3 whose position vector
⃖⃖⃖⃖⃖⃗ is orthogonal to 𝐧; that is,
𝐩 = 𝑂𝑃
{ }
Π = all points 𝑃 in ℝ3 such that 𝐩 ⋅ 𝐧 = 0 .
Crucially, as we rotate the vectors 𝐚 and 𝐛 within this plane, the vector 𝐧 does not change and we call it a
normal to the plane Π. We will use a normal as the direction for the output of our vector product.
19
But what about the length of our vector product? For the scalar product we considered the quantity
|𝐚| |𝐛| cos(𝜃). So as to provide a quantity which carries different information to the scalar product, we
shall this time use |𝐚| |𝐛| sin(𝜃) as the magnitude of our product. In order to guarantee this quantity is
always positive, we will assume that the angle 𝜃 is always measured in such a way that 0 ≤ 𝜃 ≤ 𝜋. There
is, however, a problem. Given a plane Π through the origin, a normal to Π is not unique - we could just
as well pick −𝐧. To make our choice consistent we use as convention the “right-hand rule”.
Definition [Link] Given two position vectors 𝐚 and 𝐛 in ℝ3 we define the vector (or cross) product
of the two vectors to be
𝐚 × 𝐛 = |𝐚| |𝐛| sin(𝜃)𝐧,
̂
where 𝜃 is the planar angle (in [0, 𝜋]) between them, and 𝐧̂ is a unit vector orthogonal to both 𝐚 and 𝐛
chosen to satisfy the “right-hand rule” below:
20
It can be shown that the vector product satisfies the same linearity properties as the scalar product.
(𝐿1) (𝐚 + 𝐛) × 𝐜 = 𝐚 × 𝐜 + 𝐛 × 𝐜
However, the basic properties (𝐷1) and (𝐷2) of the scalar product do not translate. Indeed, by the “right-
hand rule” we have
(𝑉 2) 𝐚 × 𝐚 = 𝟎
The second property follows from sin(0) = 0. Remarkably, the magnitude |𝐚| |𝐛| sin(𝜃) of the vector
product holds its own geometric meaning. Given two vectors 𝐚 and 𝐛 consider the parallelogram spanned
by them.
As with the scalar product, we can use its linearity properties to deduce an algebraic description of the
vector product. By properties (𝐿1) and (𝐿2), it follows that for any two vectors 𝐚 = 𝑎1 𝐢 + 𝑎2 𝐣 + 𝑎3 𝐤 and
𝐛 = 𝑏1 𝐢 + 𝑏2 𝐣 + 𝑏3 𝐤 we have
𝐚 × 𝐛 = (𝑎1 𝐢 + 𝑎2 𝐣 + 𝑎3 𝐤) × (𝑏1 𝐢 + 𝑏2 𝐣 + 𝑏3 𝐤)
= 𝑎1 𝐢 × (𝑏1 𝐢 + 𝑏2 𝐣 + 𝑏3 𝐤) + 𝑎2 𝐣 × (𝑏1 𝐢 + 𝑏2 𝐣 + 𝑏3 𝐤) + 𝑎3 𝐤 × (𝑏1 𝐢 + 𝑏2 𝐣 + 𝑏3 𝐤)
= 𝑎1 𝑏1 𝐢 × 𝐢 + 𝑎1 𝑏2 𝐢 × 𝐣 + 𝑎1 𝑏3 𝐢 × 𝐤 + 𝑎2 𝑏1 𝐣 × 𝐢 + 𝑎2 𝑏2 𝐣 × 𝐣 + 𝑎2 𝑏3 𝐣 × 𝐤
+𝑎3 𝑏1 𝐤 × 𝐢 + 𝑎3 𝑏2 𝐤 × 𝐣 + 𝑎3 𝑏3 𝐤 × 𝐤
= (𝑎2 𝑏3 − 𝑎3 𝑏2 )𝐢 − (𝑎1 𝑏3 − 𝑎3 𝑏1 )𝐣 + (𝑎1 𝑏2 − 𝑎2 𝑏1 )𝐤.
21
Thus, we have an expression for the vector product in Cartesian notation.
Note: You may recognise the quantities above as the determinants of certain 2 × 2 matrices. Indeed we
could write
| |
| 𝐢 𝐣 𝐤|
| | | | | | | |
|𝑎2 𝑎3 | |𝑎1 𝑎3 | |𝑎1 𝑎2 | | |
𝐚×𝐛 = | |𝐢 − | |𝐣 + | |𝐤 = |𝑎 𝑎 𝑎 | .
| | | | | | | 1 2 3|
| 𝑏2 𝑏3 | | 𝑏1 𝑏3 | | 𝑏1 𝑏2 | | |
| | | | | | | |
| 𝑏1 𝑏2 𝑏3 |
| |
The final quantity above (with a little abuse of notation) denotes the determinant of a 3 × 3 matrix. It is
not vital that you know how to calculate such a quantity now - see SMA term 2 - but it is a useful notation
to get used to nonetheless!
Of course, we can therefore also express the above product in coordinate notation via
⎛𝑎 ⎞ ⎛𝑏 ⎞ ⎛𝑎 𝑏 − 𝑎 𝑏 ⎞
⎜ 1⎟ ⎜ 1⎟ ⎜ 2 3 3 2⎟
⎜𝑎 ⎟ × ⎜𝑏 ⎟ = ⎜𝑎 𝑏 − 𝑎 𝑏 ⎟ .
⎜ 2⎟ ⎜ 2⎟ ⎜ 3 1 1 3⎟
⎜𝑎 ⎟ ⎜𝑏 ⎟ ⎜𝑎 𝑏 − 𝑎 𝑏 ⎟
⎝ 3⎠ ⎝ 3⎠ ⎝ 1 2 2 1⎠
Note that given two vectors we can very quickly find a unit vector normal to them via the formula
𝐚×𝐛 𝐚×𝐛
𝐧̂ = = .
|𝐚| |𝐛| sin(𝜃) |𝐚 × 𝐛|
22
( 0
) ( 1
)
Example [Link] (a) Suppose that 𝐚 = 3 and 𝐛 = −1 . Then
−4 2
⎛ 0 ⎞ ⎛ 1 ⎞ ⎛(3)(2) − (−4)(−1)⎞ ⎛ 2 ⎞
⎜ ⎟ ⎜ ⎟ ⎜ ⎟ ⎜ ⎟
𝐚 × 𝐛 = ⎜ 3 ⎟ × ⎜−1⎟ = ⎜ (−4)(1) − (0)(2) ⎟ = ⎜−4⎟ .
⎜ ⎟ ⎜ ⎟ ⎜ ⎟ ⎜ ⎟
⎜−4⎟ ⎜ 2 ⎟ ⎜ (0)(−1) − (3)(1) ⎟ ⎜−3⎟
⎝ ⎠ ⎝ ⎠ ⎝ ⎠ ⎝ ⎠
( 0
) (2) (1) (2)
Note 3 ⋅ −4 = 0 = −1 ⋅ −4 and the vector product is indeed orthogonal to its inputs.
−4 −3 2 −3
(c) Consider the triangle 𝑂𝐴𝑃 whose vertices are the origin 𝑂, the point 𝐴 with Cartesian coordinates
(1, 2, 1), and the point 𝑃 have Cartesian coordinates (−1, 1, 3). Find a unit normal to the plane Π in
ℝ3 containing 𝑂𝐴𝑃 .
(1) ( −1 )
⃖⃖⃖⃖⃖⃗ =
The position vectors 𝑂𝐴 2 ⃖⃖⃖⃖⃖⃗ =
and 𝑂𝑃 1 both lie on the plane Π, and so a normal to the
1 3
plane is given by
⎛1⎞ ⎛−1⎞ ⎛ (2)(3) − (1)(1) ⎞ ⎛ 5 ⎞
⎜ ⎟ ⎜ ⎟ ⎜ ⎟ ⎜ ⎟
⃖⃖⃖⃖⃖⃗ = ⎜2⎟ × ⎜ 1 ⎟ = ⎜(1)(−1) − (1)(3)⎟ = ⎜−4⎟ .
⃖⃖⃖⃖⃖⃗ × 𝑂𝑃
𝐧 = 𝑂𝐴
⎜ ⎟ ⎜ ⎟ ⎜ ⎟ ⎜ ⎟
⎜1⎟ ⎜ 3 ⎟ ⎜(1)(1) − (2)(−1)⎟ ⎜ 3 ⎟
⎝ ⎠ ⎝ ⎠ ⎝ ⎠ ⎝ ⎠
√
To find a unit normal we need to rescale 𝐧 to be of length 1. We have |𝐧| = 52 + (−4)2 + 32 =
√ √
50 = 5 2 and so a unit normal is given by
⎛5⎞ ⎛ √1 ⎞
⎜ ⎟ ⎜ 2 ⎟
𝐧 1 ⎜ ⎟ ⎜− √4 ⎟ ,
𝐧̂ = = √ −4 =
|𝐧| 5 2 ⎜⎜ ⎟⎟ ⎜ 5 2⎟
⎜ √3 ⎟
⎝3⎠ ⎝ 5 2 ⎠
1 4 3
or in Cartesian form 𝐧̂ = √ 𝐢 − √ 𝐣 + √ 𝐤.
2 5 2 5 2
23
1.6 The Scalar Triple Product
We have found an easy way of calculating the area of a parallelogram spanned by two vectors, but we can
go further. Consider the parallelepiped (or “3𝐷 parallelogram”) spanned by three vectors 𝐚, 𝐛 and 𝐜.
The volume of the parallelepiped is given by the area of the base parallelogram 𝑃 (say, the one spanned
by 𝐛 and 𝐜 that is shaded in the above diagram) multiplied by the height, ℎ. Note that by the “right-hand
rule” the angle between 𝐛 × 𝐜 and the vector 𝐚 is the angle 𝜃 indicated on the diagram. Finally, we have
that ℎ = |𝐚| cos 𝜃 and therefore
Definition [Link] The quantity 𝐚 ⋅ (𝐛 × 𝐜) is called the scalar triple product of the three vectors 𝐚, 𝐛
and 𝐜. We sometimes denote the scalar triple product by [𝐚, 𝐛, 𝐜].
Note that there was nothing special about which side of the parallelepiped we chose as the base paral-
lelogram in the above calculation. Reproducing the formula for different choices of base parallelogram,
and combining these with the properties of the scalar and vector products, we can deduce some useful
properties of the scalar triple product:
24
Warning: Note that the scalar triple product can be negative if the vectors are permuted. In general,
the volume of a parallelepiped spanned by the vectors 𝐚, 𝐛 and 𝐜 is given by |𝐚 ⋅ (𝐛 × 𝐜)|.
Note: We can expand the matrix notation of the vector product to the scalar triple product. We have
| |
| 𝐢 𝐣 𝐤|
| | ⎛||𝑏 𝑏 || | | | |
| | |𝑏1 𝑏3 | |𝑏1 𝑏2 | ⎞
| | ⎜
𝐚 ⋅ (𝐛 × 𝐜) = 𝐚 ⋅ |𝑏1 𝑏2 𝑏3 | = (𝑎1 𝐢 + 𝑎2 𝐣 + 𝑎3 𝐤) ⋅ | | 2 3|
𝐢 − | | |𝐣 + | | 𝐤⎟
| | ⎜|𝑐 𝑐 || |
|
|
| |
|𝑐1 𝑐2 | ⎟
| | ⎝| 2 3 | |𝑐1 𝑐3 | | | ⎠
|𝑐1 𝑐2 𝑐3 |
| |
| | | | | |
| 𝑏2 𝑏3 | |𝑏1 𝑏3 | | 𝑏1 𝑏2 |
= 𝑎1 || | − 𝑎2 |
| |
| + 𝑎3 |
| |
|
|
|𝑐2 𝑐3 | |𝑐1 𝑐3 | | 𝑐1 𝑐2 |
| | | | | |
| |
|𝑎1 𝑎2 𝑎3 |
| |
| |
= ||𝑏1 𝑏2 𝑏3 || .
| |
| |
| 𝑐1 𝑐2 𝑐3 |
| |
So, the scalar triple product is the determinant of the 3 × 3 matrix whose rows are the vectors 𝐚, 𝐛 and 𝐜.
(Again, if you are unfamiliar with determinants you can simply use the vector and scalar product formulae
successively.)
Example [Link] (a) Find the volume of the parallelepiped with edges given by 𝐚 = −3𝐢 + 2𝐣 + 2𝐤,
𝐛 = 𝐢 + 2𝐣 + 𝐤 and 𝐜 = −𝐢 + 𝐣 + 3𝐤.
( √ √ )
1+ 3 1+ 3
(b) Find the volume of the tetrahedron with vertices (1, 0, 0) and (0, 1, 0) and , 2 ,0 and
( √ √ ) 2
1+ 3 1+ 3 2
√ , √ , √ .
2 3 2 3 3
1
The volume of a tetrahedron is given by 𝑉 = 3
× area of base × height. Let us first find the vec-
tors defining the edges of the tetrahedron. Pick some “base” vertex, say (1, 0, 0), and consider the
following diagram.
25
The edges spanning the tetrahedron (possibly be commuting the labels) can be taken to be
( √ √ ) ( √ √ )
1+ 3 1+ 3 2 1− 3 1+ 3 2
𝐚 = √ , √ , √ − (1, 0, 0) = √ , √ ,√ ;
( 2 √3 2 √3 )3 ( 2 3√ 2 3√ 3)
1+ 3 1+ 3 −1+ 3 1+ 3
𝐛 = 2
, 2 ,0 − (1, 0, 0) = 2
, 2 ,0 ;
𝐜 = (0, 1, 0) − (1, 0, 0) = (−1, 1, 0).
We may take as the “base” of our tetrahedron the triangle spanned by the vectors 𝐛 and 𝐜. Note that
the area of this triangle is precisely 12 of the area of the parallelogram spanned by 𝐛 and 𝐜. Moreover,
the height of the tetrahedron is the same as the height of the parallelepiped spanned by 𝐚, 𝐛 and 𝐜.
Thus, the volume of the tetrahedron is
1 1
𝑉 = 3
× area of base × height = 6
× area of base of parallelepiped × height of parallelepiped
1
= 6
× volume of parallelepiped
1
= 6
|𝐚 ⋅ (𝐛 × 𝐜)|
| 1−√3 √ |
|⎛ √ ⎞ ⎡⎛ −1+ 3 ⎞ ⎛−1⎞⎤ |
| ⎜ 2 3 ⎟ ⎢⎜ 2 ⎟ ⎜ ⎟⎥ |
| √ √ |
1 |⎜ 1+ 3 ⎟ ⎢⎜ 1+ 3 ⎟ |
= 6| √ ⋅ × ⎜ 1 ⎟⎥ |
| ⎜ 2 3 ⎟ ⎢⎜ 2 ⎟ ⎜ ⎟⎥ |
| ⎜ 2 ⎟ ⎢⎜ ⎟ ⎜ ⎟⎥ ||
| √
| ⎝ 3 ⎠ ⎣⎝ 0 ⎠ ⎝ 0 ⎠⎦ |
| |
| 1−√3 |
|⎛ √ ⎞ ⎛ 0 ⎞ |
|⎜ 2 3 ⎟ ⎜ ⎟ |
| √ |
| |
= 61 |⎜ 1+√ 3 ⎟ ⋅ ⎜ 0 ⎟ |
| ⎜ 2 3 ⎟ ⎜√ ⎟ |
|⎜ 2 ⎟ ⎜ ⎟ |
| √ |
|⎝ 3 ⎠ ⎝ 3⎠ |
| |
( )√
1 2 1
= 6
√ 3 = 3
.
3
26
1.7 Vector Triple Product
We have seen how we can combine three vectors in order to find a scalar with some geometric meaning.
Can we do the same but produce a vector?
Notice that we can repeatedly take vector products to combine as many vectors as we wish. For example,
given three vectors 𝐚, 𝐛 and 𝐜 in ℝ3 we can consider the vector 𝐚 × (𝐛 × 𝐜). What special properties does
this product of vectors have?
First, recall the vector product is orthogonal to its inputs. In particular, the vector 𝐛 × 𝐜 is orthogonal to
both 𝐛 and 𝐜. Furthermore, for any non-zero vector 𝐚 we must also have that the vector 𝐚 × (𝐛 × 𝐜) is
orthogonal to 𝐛 × 𝐜. Hence, the vector 𝐚 × (𝐛 × 𝐜) must lie in the plane spanned by the vectors 𝐛 and 𝐜.
𝐚 × (𝐛 × 𝐜) = 𝜆𝐛 + 𝜇𝐜.
Now, notice that the vector 𝐚 × (𝐛 × 𝐜) must also be orthogonal to 𝐚, and so by the linearity of the scalar
product these scalars must satisfy we have
If 𝐚⋅𝐛 = 𝐚⋅𝐜 = 0 then 𝐚 is parallel to 𝐛×𝐜 and so 𝐚×(𝐛×𝐜) = 𝟎 = 0𝐛+0𝐜 (and we have immediately found
our unique 𝜆 and 𝜇). So, without loss of generality, assume that 𝐚 ⋅ 𝐛 ≠ 0 and let us find the solutions to
27
the above equation. In this case 𝜆 = −𝜇 𝐚⋅𝐛
𝐚⋅𝐜
is uniquely specified by 𝜇. For ease of notation, let us rescale
𝜇 and instead consider the scalar 𝜔 = − 𝐚⋅𝐛
𝜇
. We thus have
( )
𝐚⋅𝐜
𝐚 × (𝐛 × 𝐜) = 𝜆𝐛 + 𝜇𝐜 = −𝜇 𝐛 + 𝜇𝐜 = 𝜔 [(𝐚 ⋅ 𝐜)𝐛 − (𝐚 ⋅ 𝐛)𝐜] . (1.7.3)
𝐚⋅𝐛
Let us now proceed via example. Consider the case when 𝐚 = 𝐢, 𝐛 = 𝐢 and 𝐜 = 𝐣. Then the left-hand side
of equation (1.7.3) is
𝐚 × (𝐛 × 𝐜) = 𝐢 × (𝐢 × 𝐣) = 𝐢 × 𝐤 = −𝐣
Since the same scalar 𝜔 must be valid for every choice of our vectors, we must have that 𝜔 = 1.
Definition [Link] The vector triple product of three vectors 𝐚, 𝐛 and 𝐜 in ℝ3 is the vector
𝐚 × (𝐛 × 𝐜) = (𝐚 ⋅ 𝐜)𝐛 − (𝐚 ⋅ 𝐛)𝐜.
One can also check rigorously that both expressions are the same by using Cartesian coordinates.
Properties of the vector triple product: From the anti-symmetry of the vector product we immediately
have that
𝐚 × (𝐛 × 𝐜) = −(𝐛 × 𝐜) × 𝐚 = (𝐜 × 𝐛) × 𝐚 = −𝐚 × (𝐜 × 𝐛).
In other words, the order of the vectors really matters! But what about the order of the operations? Can
we just write 𝐚 × 𝐛 × 𝐜 in the same way we can write 1 × 2 × 3 for real multiplication?
Again, we proceed via example with 𝐚 = 𝐢, 𝐛 = 𝐢 and 𝐜 = 𝐣. We have already seen that 𝐚 × (𝐛 × 𝐜) = −𝐣,
but on the other hand we have
(𝐚 × 𝐛) × 𝐜 = (𝐢 × 𝐢) × 𝐣 = 𝟎 × 𝐣 = 𝟎 ≠ −𝐣.
Thus, the order in which we perform the operations also really matters!
28
Let’s look further into these quantities. In general, what is the difference between 𝐚×(𝐛×𝐜) and (𝐚×𝐛)×𝐜?
Well, we have
Example [Link] Find the vector triple product of the vectors 𝐚 = −3𝐢 + 2𝐣 + 2𝐤, 𝐛 = 𝐢 + 2𝐣 + 𝐤 and
𝐜 = −𝐢 + 𝐣 + 3𝐤.
⎛1 ⎞ ⎛−1⎞ ⎛14⎞
⎜ ⎟ ⎜ ⎟ ⎜ ⎟
= 11 2 − 3 ⎜ 1 ⎟
⎜ ⎟ = ⎜19⎟ .
⎜ ⎟ ⎜ ⎟ ⎜ ⎟
⎜1 ⎟ ⎜3⎟ ⎜2⎟
⎝ ⎠ ⎝ ⎠ ⎝ ⎠
29
Given a straight line 𝓁 in Euclidean space (be it either ℝ2 or ℝ3 ) we can specify it by its direction (via
some vector 𝐛) and, to distinguish it from other parallel lines, some reference point 𝐴 (with position vector
⃖⃖⃖⃖⃖⃗ on the line.
𝐚 = 𝑂𝐴)
where 𝐚 is the position vector of a reference point on the line and 𝐛 is a direction vector for the line.
This representation is called the parametric equation of the line 𝓁 (with parameter 𝜆).
Note: We can have many different looking parametric equations for the same line since we are free to
choose any point on the line as the reference point and any vector parallel to 𝐛 as the direction.
Example [Link] Find the parametric equation of the line in ℝ3 that is parallel to the 𝑦-axis, lies
entirely in the 𝑥𝑦-plane, and whose 𝑥-coordinate is always 1.
First we must choose a point on the line. The obvious point to choose is 𝐴 = (1, 0, 0), which has
position vector 𝐚 = 𝐢. As for the direction, we know the line is parallel to the 𝑦-axis so we may choose
any vector in the direction of the 𝑦-axis, say 𝐛 = 𝐣. The parametric equation of the line is therefore
given by 𝐫 = 𝐚 + 𝜆𝐛 = 𝐢 + 𝜆𝐣 (for 𝜆 in ℝ).
30
Note that one could choose a different point on the line as the reference, say 𝐚 = 𝐢 + 7𝐣, and for
the direction we could freely choose a different scaling, say 𝐛 = −5𝐣. Then the line could also be
expressed as 𝐫 = (𝐢 + 7𝐣) + 𝜇(−5𝐣) = 𝐢 + (7 − 5𝜇)𝐣 (for 𝜇 in ℝ).
Point-to-Line:
Given some line 𝓁 with parametric equation 𝐫 = 𝐚 + 𝜆𝐛, we may ask how close this line gets to a given
point 𝐶.
(with position vector 𝐫, say). By the properties of the scalar product, this quantity must satisfy
| ⃖⃖⃖⃖⃖⃗|2
|𝑂𝑃 | = |𝐫|2 = 𝐫 ⋅ 𝐫 = (𝐚 + 𝜆𝐛) ⋅ (𝐚 + 𝜆𝐛) = |𝐚|2 + 2𝜆𝐚 ⋅ 𝐛 + 𝜆2 |𝐛|2 .
| |
But, this is just a quadratic in the variable 𝜆, so we can find the minimum by differentiating. We have
( )
𝑑 | ⃖⃖⃖⃖⃖⃗|2
0 = |𝑂𝑃 | = 2𝐚 ⋅ 𝐛 + 2𝜆|𝐛|2
𝑑𝜆 | |
31
Notice that
( ( ) ) ( )
𝐚⋅𝐛 𝐚⋅𝐛
𝐫𝐌 ⋅ 𝐛 = 𝐚− 𝐛 ⋅𝐛 = 𝐚⋅𝐛− 𝐛 ⋅ 𝐛 = 𝐚 ⋅ 𝐛 − 𝐚 ⋅ 𝐛 = 0,
|𝐛|2 |𝐛|2
We can now calculate the minimum distance from the origin to the line by simply considering the mag-
nitude of the vector 𝐫𝐌 . Again using the properties of the scalar product, we have
( ( ) ) ( ( ) ) ( )2
𝐚⋅𝐛 𝐚⋅𝐛 (𝐚⋅𝐛)2 𝐚⋅𝐛
|𝐫𝐌 |2 = 𝐚 − |𝐛| 2
𝐛 ⋅ 𝐚 − |𝐛|2
𝐛 = |𝐚| 2
− 2 |𝐛|2
+ |𝐛|2
|𝐛|2
|𝐚|2 |𝐛|2 −(𝐚⋅𝐛)2
= |𝐛|2
|𝐚|2 |𝐛|2 −|𝐚|2 |𝐛|2 cos2 (𝜃)
= |𝐛|2
2
= |𝐚| sin (𝜃).
2
The above formula works in any Euclidean space, but if we are working in ℝ3 we can write
where ̂𝐛 = 𝐛
|𝐛|
is a unit vector in the direction of 𝐛. Moreover, we can translate the whole of the diagram
above by the position vector 𝐜 of some point 𝐶, in order to conclude the following.
32
Line-to-Line (Parallel):
Next, we seek to find a formula for the distance between two lines 𝓁1 and 𝓁2 . First, let us consider the
case that the lines in question are parallel. In this case, let 𝐚𝟏 be the position vector of a reference point 𝐴1
on the line 𝓁1 , and let 𝐚𝟐 be the position vector of a reference point 𝐴2 on the line 𝓁2 . Then, parametric
equations for the two lines can be given in the form
where 𝐛 is a direction vector for both lines. Moreover, the vector 𝐚𝟏 − 𝐚𝟐 takes us from the point 𝐴2 to the
point 𝐴1 . Consider the following diagram.
It is then clear that the minimum distance 𝑑 between the parallel lines is given by
|𝐛| |(𝐚𝟏 − 𝐚𝟐 ) × 𝐛|
|𝐚1 − 𝐚2 | sin(𝜃) = |𝐚𝟏 − 𝐚𝟐 | sin(𝜃) = .
|𝐛| |𝐛|
Thus:
|(𝐚𝟏 − 𝐚𝟐 ) × 𝐛|
𝑑 = = |(𝐚𝟏 − 𝐚𝟐 ) × ̂𝐛|,
|𝐛|
Note: Another way to reach this formula is to consider the parallelogram with sides 𝐛 and 𝐚𝟏 − 𝐚𝟐 . The
height of this parallelogram (when considering 𝐛 to be the base) is precisely the minimum distance we
need, and is given by the area of the parallelogram divided by the length of the base . By section 1.5 we
|(𝐚𝟏 −𝐚𝟐 )×𝐛|
know this to be ℎ = |𝐛|
, and we again reach the formula above.
33
1.8.2 Line-to-Line (non-parallel):
Next, consider the case of two non-parallel lines 𝓁1 and 𝓁2 . Again, each line must have a parametric
equation, say
𝓁1 ∶ 𝐫𝟏 = 𝐚𝟏 + 𝜆𝐛𝟏 ; and 𝓁2 ∶ 𝐫𝟐 = 𝐚𝟐 + 𝜇𝐛𝟐 , (𝜆, 𝜇 in ℝ).
The minimum distance 𝑑 between the lines occurs when we measure in a direction perpendicular to both 𝓁1
and 𝓁2 .
Lets argue more geometrically this time. Consider the parallelepiped pictured below, two of whose edges
lie on the lines.
Evidently, the minimum distance between the two lines is then given by the height of this parallelepiped.
From section 1.6 we know that this is precisely the volume (which is given by the magnitude of the scalar
triple product of three of its edges) divided by the area of the shaded base parallelogram (which is given
by the magnitude of vector product of its edges). Thus:
|(𝐚 − 𝐚 ) ⋅ (𝐛 × 𝐛 )|
𝑑 = |
𝟏 𝟐 𝟏 𝟐 |
.
|𝐛𝟏 × 𝐛𝟐 |
34
(0) (1)
Example [Link] (a) Find the shortest distance from the line 𝐫 = 𝐚 + 𝜆𝐛 = 1 +𝜆 1 to (i) the
0 0
origin; and (ii) the point 𝐜 = (2, −2, 0).
(i) By the formula derived above, the shortest distance from the line to the origin is
|⎛ ⎞ ⎛ ⎞| |⎛ ⎞|
| 0 | | 0 |
|⎜ ⎟ ⎜1⎟| |⎜ ⎟|
|(𝐚 − 𝟎) × 𝐛| |
1 |⎜ ⎟ ⎜ ⎟| | 1 ||⎜ ⎟|| 1
= √ | 1 × 1 | = √ | 0 | = √ .
|𝐛| ⎜ ⎟ ⎜ ⎟
2 ||⎜ ⎟ ⎜ ⎟|| ⎜ ⎟
2 ||⎜ ⎟|| 2
|⎝0⎠ ⎝0⎠| |⎝−1⎠|
| | | |
( 0 ) ( 2 ) ( −2 )
(ii) We have 𝐚 − 𝐜 = 1 − −2 = 3 , so the shortest distance from the line to 𝐜 is
0 0 0
| ⎛ ⎞ ⎛ ⎞| |⎛ ⎞|
| −2 1 | | 0 |
| ⎜ ⎟ ⎜ ⎟| |⎜ ⎟|
|(𝐚 − 𝐜) × 𝐛| |
1 | ⎜ ⎟ ⎜ ⎟| | 1 ||⎜ ⎟|| 5
= √ | 3 × 1 | = √ | 0 | = √ .
|𝐛| 2 ||⎜⎜ ⎟⎟ ⎜⎜ ⎟⎟|| 2 ||⎜⎜ ⎟⎟|| 2
|⎝ 0 ⎠ ⎝0⎠| |⎝−5⎠|
| | | |
Find (i) the distance between 𝓁1 and 𝓁2 , and (ii) the distance between 𝓁1 and 𝓁3 .
( −1 )
(i) We should first check whether the lines are parallel. By inspection the direction vectors 1
(3) 2
and 0 are not scalar multiples of one another, so the lines are not parallel. Indeed, we have
( −1 ) −2( 3 ) ( −2 )
1 × 0 = 4 ≠ 𝟎. Thus, by the formula derived above, the distance from 𝓁1 to 𝓁2 is
2 −2 −3
|(( 1 ) ( 1 )) ( −2 )|
| 2 − 1 ⋅ 4 || ( ) ( −2 )|
|(𝐚 − 𝐚 ) ⋅ (𝐛 × 𝐛 )| | 3 1 || 0
| 𝟏 𝟐 𝟏 𝟐 |
= | 1 −3 |
= √ | 1 ⋅ 4 || = √ .
2
|𝐛𝟏 × 𝐛𝟐 | | ( ) |
| −2 | 29 | 2 −3 | 29
| −34 |
| |
(ii) In this case, the lines 𝓁1 and 𝓁3 are parallel since the direction vectors are scalar multiples of one
(2) ( −1 )
another: −2 = −2 1 . By the formula derived above, the distance from 𝓁1 to 𝓁3 is therefore
−4 2
|(( 1 ) ( 0 )) ( −1 )| √
| 2 − 3 × 1 || ( 1 ) ( −1 )|
|(𝐚𝟏 − 𝐚𝟑 ) × 𝐛| | 3 |
= | | = | −1 × 1 | = √50 = √5 .
0 2 1
(
| −1 |) √ | |
|𝐛| | 1 | 6| 3 2 | 6 3
| 2 |
| |
35
1.9 Affine Planes: Equations of planes
We have already encountered planes a number of times, and we have also discussed how we can specify
a point on a plane using two vectors on the plane. But, all the planes we have seen so far have passed
through the origin. More generally, a plane is specified using two non-parallel direction vectors 𝐛 and 𝐜
and a reference point 𝐴 (with position vector 𝐚 = 𝑂𝐴,
⃖⃖⃖⃖⃖⃗ say) through which the plane passes.
Each point on the plane relates to some choice of the scalars 𝜆 and 𝜇. For example, to reach a point 𝑃 on
the plane with position vector 𝐫 = 𝐚 + 2𝐛 + 𝐜 (so with 𝜆 = 2 and 𝜇 = 1), we travel from the origin along
the vector 𝐚 to the point 𝐴, then travel across the plane a distance of 2|𝐛| in the direction of 𝐛, and finally
travel across the plane a distance of |𝐜| in the direction of 𝐜. In the case 𝐚 = 𝟎 the plane clearly passes
through the origin. An affine plane is a plane that does not pass through the origin.
You may be familiar with the Cartesian equation of a plane; so something in the form 𝑎𝑥 + 𝑏𝑦 + 𝑐𝑧 = 𝑑.
It is quite easy to find the Cartesian equation of a plane from its parametric equation. To see this, pick
some vector 𝐧 that is orthogonal to both direction vectors 𝐛 and 𝐜. Then, observe that for any point 𝑃 on
the plane, distinct from the reference point 𝐴, the vector 𝐴𝑃
⃖⃖⃖⃖⃖⃗ lies on the plane and so is also orthogonal
36
to 𝐧. So, if 𝑃 has position vector 𝐫 = 𝐚 + 𝜆𝐛 + 𝜇𝐜 then
⃖⃖⃖⃖⃖⃗ = 𝐧 ⋅ (𝐫 − 𝐚);
0 = 𝐧 ⋅ 𝐴𝑃 or, in other words, 𝐧 ⋅ 𝐫 = 𝐧 ⋅ 𝐚.
( 𝑛1 )
Furthermore, we know how to find a normal 𝐧 = 𝑛2
𝑛3
to the plane: one is simply given by the vector
product 𝐧 = 𝐛 × 𝐜 of the direction vectors. Denote by 𝑑 the constant 𝑑 = 𝐧 ⋅ 𝐚 = 𝑛1 𝑎1 + 𝑛2 𝑎2 + 𝑛3 𝑎3 . Then,
(𝑥)
assuming the point 𝑃 on the plane has Cartesian coordinates 𝐫 = 𝑧𝑦 we attain the Cartesian equation
of the plane:
𝑛1 𝑥 + 𝑛2 𝑦 + 𝑛3 𝑧 = 𝐧 ⋅ 𝐫 = 𝐧 ⋅ 𝐚 = 𝑑.
Given an affine plane Π with parametric equation 𝐫 = 𝐚 + 𝜆𝐛 + 𝜇𝐜 (for 𝜆, 𝜇 in ℝ), let 𝐫𝐌 be the position
vector of the point 𝑀 on Π closest to the origin. Evidently, in order to attain the minimum distance, the
vector 𝐫𝐌 will be orthogonal to the plane Π. In other words, if 𝑐 = |𝐫𝐌 | is the minimum distance from Π
to the origin, then we must have that 𝐫𝐌 = ± 𝑐 𝐧,
̂ where 𝐧̂ is a unit normal to the plane. In particular, by
property (D2) of the scalar product
Thus, the distance from a plane to the origin is simply give by |𝐧̂ ⋅ 𝐚|, where 𝐧̂ is a unit normal.
Example [Link] (a) Find the parametric and Cartesian equations of the plane containing the points
(1, 0, 0), (2, 1, 0) and (1, 1, 2).
To answer this, let’s arbitrarily take 𝐴 = (1, 0, 0) (with position vector 𝐚 = 𝐢) to be a reference point
(2) (1) (1)
on the plane and find two direction vectors. Clearly, we may take 𝐛 = 1 − 0 = 1 and
(1) (1) (0) 0 0 0
(b) Find the shortest distance from the plane 𝑥 + 2𝑦 + 3𝑧 = 4 to the origin.
37
(1)
A normal to the plane can be read off as 𝐧 = 2 . Moreover, by letting 𝑦 = 𝑧 = 0 it is clear the point
(4) 3
with position vector 𝐚 = 0 lies on the plane. Thus, the minimum distance is given by
0
( ) ( 4 )|
|𝐧 ⋅ 𝐚| 1 | 1 4
|𝐧̂ ⋅ 𝐚| = = √ || 2 ⋅ 0 || = √ .
|𝐧| 14 |
3 0 | 14
38
Chapter 2
Kinematics
Now we’ve become comfortable with vectors we will explore how the theory we have developed can help
us to model real world problems. It is fair to say that vectors are an integral part of the language of science,
particularly physics. In this section we will focus on a class of applications relating to the mechanics of
particle motion in ℝ3 , the study of which is called kinematics. In particular, we now allow our scalars
and vectors to depend on another variable, namely, time.
The trajectory of a particle 𝑃 is a path in space parametrised by time, 𝑡. We denote by 𝐫(𝑡) its position
vector at time 𝑡, and naturally call this the position of the particle. We assume some suitable origin has
been chosen as indicated in figure 2.1. It is often useful to think of a particle’s position in terms of its
Cartesian coordinates. We often denote the components
where 𝐢, 𝐣 and 𝐤 are the Cartesian basis vectors (see figure 2.1) .
Example [Link] (a) Consider a particle 𝑃 travelling at constant speed along a straight line in space.
If we denote by 𝐚 the position of the particle at time 𝑡 = 0, say, then the position 𝐫(𝑡) of the particle
39
Figure 2.1: Depictions of the vector description of particle motion.
𝐫(𝑡) = 𝐚 + 𝑡𝐛,
Here, the vector 𝐫(𝑡) traces out a circle in the 𝑥𝑦-plane centred at the origin and with radius 𝑅.
40
2𝜋 𝜔
The particle makes complete revolutions of the circle with period 𝜔
(so with frequency 2𝜋
). The
quantity 𝜔 is referred to as the angular frequency.
Given knowledge of the position of a particle, a natural question to ask is: in which direction is the particle
travelling at time 𝑡, and at what speed?
The velocity 𝐯(𝑡) of a particle is defined to be the rate of change of its position with respect to time; that
is, we have 𝐯(𝑡) = 𝑑
𝑑𝑡
𝐫(𝑡). We often use the “dot” notation 𝐫(𝑡)
̇ = 𝑑
𝑑𝑡
𝐫(𝑡) for derivatives with respect to
time. But, wait! We know how to differentiate scalar functions (and can hopefully recall the product and
chain rules), but how do we differentiate vectors? We will see how to do this in detail (by taking a certain
limit) in term 2, but for now we simply define:
Definition [Link] The derivative (with respect to the variable 𝑡) of a vector with Cartesian form
𝐫(𝑡) = 𝑥(𝑡)𝐢 + 𝑦(𝑡)𝐣 + 𝑧(𝑡)𝐤 is
[ ] [ ] [ ]
𝑑 𝑑 𝑑 𝑑
̇
𝐫(𝑡) = 𝐫(𝑡) = 𝑥(𝑡)𝐢
̇ + 𝑦(𝑡)𝐣
̇ + 𝑧(𝑡)𝐤
̇ = 𝑥(𝑡) 𝐢 + 𝑦(𝑡) 𝐣 + 𝑧(𝑡) 𝐤.
𝑑𝑡 𝑑𝑡 𝑑𝑡 𝑑𝑡
That is to say we just differentiate the individual components (this requires 𝐢, 𝐣, 𝐤 don’t change in time,
which they don’t).
Note: Formally, we derive the velocity at a given time 𝑡 by considering the displacement 𝜹𝐫(𝑡) = 𝐫(𝑡 +
𝛿𝑡) − 𝐫(𝑡) that occurs in a very small amount of time 𝛿𝑡 immediately after the time 𝑡.
41
𝜹𝐫(𝑡)
The velocity is precisely the limiting value (as 𝛿𝑡 → 0) of the rate of change 𝛿𝑡
of this quantity. For
details on the precise meaning of a limit of this kind, see the end of term 1 of SMA. However, the key thing
to take away is that the direction of the velocity is tangent to the path of the particle.
The speed of a particle at time 𝑡 is a scalar and is defined to be the magnitude 𝑣(𝑡) = |𝐯(𝑡)| of the velocity.
Note, that, as discussed in the previous chapter, this does not depend on the choice of origin.
Acceleration:
The velocity of a particle will usually depend on the time, and so we can ask how it changes in time. The
acceleration 𝐚(𝑡) of a particle is the rate of change in its velocity. To be precise, if a particle has position
𝐫(𝑡) = 𝑥(𝑡)𝐢 + 𝑦(𝑡)𝐣 + 𝑧(𝑡)𝐤 then
[ ] [ 2 ] [ 2 ]
𝑑 𝑑2 𝑑 𝑑
𝐚(𝑡) = ̇
𝐯(𝑡) = 𝐯(𝑡) = 𝐫̈ (𝑡) = 𝑥(𝑡)𝐢
̈ + 𝑦(𝑡)𝐣
̈ + 𝑧(𝑡)𝐤
̈ = 𝑥(𝑡) 𝐢 + 𝑦(𝑡) 𝐣 + 𝑧(𝑡) 𝐤.
𝑑𝑡 𝑑𝑡2 𝑑𝑡2 𝑑𝑡2
Example [Link] (a) As before, consider a particle 𝑃 travelling at constant speed along a straight
line in space; so
( ) ( )
𝑎1 +𝑡𝑏1 𝑏1
𝐫(𝑡) = 𝐚 + 𝑡𝐛 = 𝑎2 +𝑡𝑏2 ; ̇ =
𝐯(𝑡) = 𝐫(𝑡) 𝑏2 = 𝐛; and ̇ = 𝟎.
𝐚(𝑡) = 𝐯(𝑡)
𝑎3 +𝑡𝑏3 𝑏3
(b) Consider the particle 𝑃 whose position is given by 𝐫(𝑡) = 𝑅 cos(𝜔𝑡)𝐢 + 𝑅 sin(𝜔𝑡)𝐣. Then
̇
𝐯 = 𝐫(𝑡) = −𝑅𝜔 sin(𝜔𝑡)𝐢 + 𝑅𝜔 cos(𝜔𝑡)𝐣; and 𝐚 = 𝐫̈ (𝑡) = −𝑅𝜔2 cos(𝜔𝑡)𝐢 − 𝑅𝜔2 sin(𝜔𝑡)𝐣.
42
(i) We have 𝐫(𝑡) ⋅ 𝐯(𝑡) = 0, so the velocity is always perpendicular to the position;
√
(ii) The speed |𝐯(𝑡)| = (−𝑅𝜔 sin(𝜔𝑡))2 + (𝑅𝜔 cos(𝜔𝑡))2 = 𝑅𝜔 is always constant acceleration
is a vector and changing direction counts as acceleration ;
(iii) We have 𝐚(𝑡) = −𝜔2 𝐫(𝑡) so the particle is always accelerating directly towards the centre of the
circular path, and with constant magnitude |𝐚(𝑡)| = 𝑅𝜔2 .
43
2.2 Forces
In the previous circular motion example, Example [Link](b), it is clear that there must be some force
pulling the particle towards the centre of the circular path; be it the tension in a rope or gravity towards
a much heavier object. Of course, to us a force is simply a vector - it pulls in a certain direction with a
certain strength.
2.2.1 Gravity:
In general, the gravitational force acting on an object of mass 𝑚 (whose centre of mass has position vector
𝐫) due to a body of mass 𝑀 (whose centre of mass is at the origin) is given by
𝐺𝑀𝑚
𝐅𝑔 = − 𝐫̂ ,
|𝐫|2
where 𝐺 = 6.7 × 10−11 𝑘𝑔 −2 𝑚3 𝑠− 2 is a scalar known as Newton’s (universal) gravitational constant and
𝐫̂ is a unit vector in the direction of 𝐫.
Near the surface of the earth, with 𝑀 = mass of earth, |𝑟| = radius of earth, 𝐫̂ = 𝐤 we have 𝐺𝑀
|𝐫|2
=𝑔=
9.81𝑚𝑠−2 , the (local) gravitational constant of earth’s gravity. Thus, the force exerted upon the object
due to gravity is 𝐅𝑔 = − 𝐺𝑀𝑚
|𝐫|2
𝐫̂ = −𝑚𝑔𝐤.
Pedantic point: Strictly forces should not depend directly on a position vector, the full form of Newton’s
law has the vector 𝐫̂ as being the direction of the difference 𝐫2 − 𝐫1 between masses 𝑚1 , 𝑚2 :
𝐺𝑚1 𝑚2
𝐅𝑔 = − 𝐫̂ ,
|𝐫|2
The difference does not depend on the choice of origin as it is subtracted off. But for obvious reasons
when we are on earth we generally assume the origin to be the centre of the earth.
Often, several forces can act on a particle and we must add the force vectors to get the total force.
44
A very important property of the forces acting on an object is that the total force 𝐅 must obey Newton’s
Second Law of Motion:
𝐅 = 𝑚𝐯̇ = 𝑚𝐚; where 𝑚 is the mass of the object (and is assumed to be constant).
In particular, an object exhibits zero acceleration precisely when the sum of the forces acting upon it
satisfies 𝐅 = 𝟎, and the object does not then accelerate until it encounters a change in total force. We
encountered this situation in Example [Link](a), where the particle simply travelled at constant velocity
in a straight line for ever.
Example [Link] Consider a block on a slope at angle 𝜃 to the horizontal. As we increase 𝜃 we may
wish to ask when the block will start sliding down. Typically this problem would be modelled by
considering three forces acting on the block.
(1) Gravity, 𝐅𝑔 = −𝑚𝑔 𝐤, acting ‘downwards’ with magnitude 𝑚𝑔, where the scalar 𝑔 is the gravita-
tional constant;
(2) Reaction, 𝐅𝑅 = −𝑅 sin 𝜃 𝐢 + 𝑅 cos 𝜃 𝐤, acting normal to the surface with magnitude 𝑅. It stops
the object ‘falling through’ the surface;
(3) Friction, 𝐅𝑓 = 𝜇𝑅 cos 𝜃 𝐢 + 𝜇𝑅 sin 𝜃 𝐤, acting to oppose motion with magnitude 𝜇𝑅, where the
scalar 𝜇 is the coefficient of friction between the two surfaces. Friction acts orthogonally to the
reaction force.
Let us fix 𝜃 to be the critical angle where the block is still at rest, but would move if we were to
45
tilt the slope any further. Then by Newton’s Second Law
𝑅 sin 𝜃
In order for the first component of this equation to be satisfied we must have 𝜇 = 𝑅 cos 𝜃
= tan 𝜃.
Thus, the critical angle is 𝜃 = arctan 𝜇.
(Even though it doesn’t look that way, for the last component of the equation to be satisfied we
have the same constraint 𝜇 = tan 𝜃. This is because the magnitude of the reaction force actually
depends on the angle via the relationship 𝑅 = 𝑚𝑔 cos 𝜃.)
2.2.2 Electromagnetism:
If a particle is charged, it is affected by electromagnetic forces. In an electric field 𝐄 and a magnetic field
𝐁 a particle of charge 𝑞 feels a force
Example [Link] Consider an electron 𝐸 with charge −𝑒 in a constant magnetic field 𝐁 = 𝐵𝐤. If the
( 𝑣𝑥 )
velocity of the particle is given by 𝐯 = 𝑣𝑣𝑦 , then by Newton’s second law, we have
𝑧
⎛𝑣̇ ⎞ ⎛𝑣 ⎞ ⎛0⎞ ⎛𝑣 ⎞
⎜ 𝑥⎟ ⎜ 𝑥⎟ ⎜ ⎟ ⎜ 𝑦 ⎟
𝑚 𝑣̇𝑦 = 𝑚𝐯̇ = 𝑚𝐚 = 𝐅𝑒𝑚 = −𝑒 (𝟎 + 𝐯 × 𝐵𝐤) = −𝑒𝐵 𝑣𝑦 × 0 = −𝑒𝐵 ⎜−𝑣𝑥 ⎟ .
⎜ ⎟ ⎜ ⎟ ⎜ ⎟
⎜ ⎟ ⎜ ⎟ ⎜ ⎟ ⎜ ⎟
⎜𝑣̇ ⎟ ⎜𝑣 ⎟ ⎜1⎟ ⎜ 0 ⎟
⎝ 𝑧⎠ ⎝ 𝑧⎠ ⎝ ⎠ ⎝ ⎠
46
for some real scalar 𝑅. In this case, we have a helical trajectory since (up to some constants)
⎛𝑅 cos(𝜔𝑡)⎞
⎜ ⎟
𝐫 = 𝑅 sin(𝜔𝑡) ⎟ .
⎜
⎜ ⎟
⎜ 𝑐𝑡 ⎟
⎝ ⎠
We saw in the case of circular motion that the velocity of a particle was orthogonal to the acceleration;
ergo, by Newton’s second law (𝐅 = 𝑚𝐚 = 𝑚̈𝐫 ) the velocity was orthogonal to the force. In other words
we had 𝐅 ⋅ 𝐯 = 0. In general, by the product rule for the derivative of a dot product (see Problem Sheet
3), we have
( )
1 1 𝑑 1 𝑑 ( 2) 𝑑 1 2
𝐅 ⋅ 𝐯 = 𝑚(̈𝐫 ⋅ 𝐫)
̇ = 𝑚(̈𝐫 ⋅ 𝐫̇ + 𝐫̇ ⋅ 𝐫̈ ) = 𝑚 (𝐫̇ ⋅ 𝐫)
̇ = 𝑚 |𝐫|
̇ = 𝑚𝑣 ,
2 2 𝑑𝑡 2 𝑑𝑡 𝑑𝑡 2
𝑊 = 𝐅 ⋅ 𝐯 𝑑𝑡.
∫𝐶
We will see how to actually integrate a function over a trajectory (or ‘contour’) much later in the course.
Note, you might remember from A-Level mechanics the rule, work done is force times distance. The
47
point being here that the integral of 𝐯 is the opposite of differentiating 𝐫, it calculates a distance travelled,
but in more general its distance and direction, hence why it has a vector form.
Example [Link] (a) Magnetic fields: The force imparted on a charged particle (of charge 𝑞) in a
magnetic field 𝐁 satisfies
𝐅𝑒𝑚 ⋅ 𝐯 = 𝑞 (𝐯 × 𝐁) ⋅ 𝐯 = 𝑞 (𝐯 × 𝐯) ⋅ 𝐁 = 𝑞 (𝟎 ⋅ 𝐁) = 0,
by the cyclic property of the scalar triple product. Thus, there is no work done and the kinetic energy
is constant.
(b) Gravity: The gravitational force imparted upon an object of mass 𝑚 satisfies
48
2.3.1 Angular Momentum
The momentum of a particle is defined to be the vector 𝐩 = 𝑚𝐯; its mass times its velocity. By Newton’s
second law we have 𝐅 = 𝐩,
̇ so when the total force satisfies 𝐅 = 𝟎 the momentum is conserved (it does
not change with time). Indeed this was the case in Example [Link](a) when the particle continued in
a straight line at constant speed. But what about for circular motion as in Example [Link](b)? Here,
momentum was not conserved (since 𝐅 ≠ 𝟎 and so the velocity changed with time). However, it turns out
there is another quantity that was conserved. We define the angular momentum (about the origin) of a
particle to be the vector
𝐉 = 𝐫 × 𝐩.
Let us consider how this vector changes with time. First, note that by the product rule for the derivative
of a cross product (see Problem Sheet 3), we have
𝐉̇ = 𝐫̇ × 𝐩 + 𝐫 × 𝐩̇ = 𝑚(𝐫̇ × 𝐫)
̇ + 𝐫 × 𝐅 = 𝐫 × 𝐅.
Definition [Link] We call a force 𝐅 a central force if it acts entirely in the direction of the position
of the particle (the ‘radial direction’); that is if we can write 𝐅 = 𝐹 (𝐫) 𝐫̂ for some scalar 𝐹 (𝐫).
Clearly, for any central force the angular momentum is preserved; that is, we have 𝐫 × 𝐅 = 𝟎.
Finally, notice that 𝐉 is orthogonal to both 𝐫 and 𝐯 since, for example, 𝐫 ⋅ 𝐉 = 𝐫 ⋅ (𝐫 × 𝐩) = 0. Thus, in the
case of circular motion it actually provides the axis of rotation.
49
Pedant’s return
Angular momentum as presented depends ona position vector and hence the choice of origin, again what
we really want is a quantity like:
𝐉 = (𝐫 − 𝐫0 ) × 𝐩.
as in we are looking at a rotation around a fixed location 𝐫0 , but we can often identify the suitable location
𝐫0 and make it the origin to keep things simple.
50
2.4 2D Polar Coordinates
Coordinates are simply a way of labelling points, and so far we have exclusively used the Cartesian co-
ordinate system. But, when central forces are at play some natural symmetry is often exhibited by the
motion of the objects in the system. To take advantage of this fact, it is sometimes useful to adapt our
coordinate system to better suit the situation we are faced with.
Instead of specifying points in the plane relative to the two coordinate axes, one might instead specify a
point by the distance to and the direction (measured, say, in terms of the anticlockwise angle from the
𝑥-axis) from the origin.
The polar coordinates (𝑟, 𝜃) of a point 𝑃 are precisely these quantities. We can freely switch from polar
back to Cartesian coordinates (𝑥, 𝑦), and vice versa, through the following correspondences:
√
𝑥 = 𝑟 cos 𝜃; 𝑦 = 𝑟 sin 𝜃; 𝑟= 𝑥2 + 𝑦2 ; 𝜃 = arctan (𝑦∕𝑥).
In the case of the circular motion of Example [Link](b), we simply have that the first polar coordinate 𝑟 =
𝑅 is constant, whilst the second 𝜃 = 𝜔𝑡 is simply a constant multiple of time. Here, we can immediately
see that polar coordinate system can encode the trajectory of the particle in a far more efficient way than
Cartesians.
We can actually go even further [difficult section of the notes alert]. Let’s discuss how we can express
polar coordinates using the language of vectors. Imagine we’re in a 2D spacecraft in a circular orbit of
the 2D earth, looking out of the front window, and that we want to make a map of the 2D universe from
our perspective. Naturally we might attempt to specify the position of objects relative to us by specifying
a set of ‘Cartesian-like’ basis vectors emanating from our spacecraft. To do so we need two orthogonal
unit vectors, which we label 𝐞𝑟 (representing 𝐢 from our perspective) and 𝐞𝜃 (representing 𝐣 from our
51
perspective).
One might notice that in the radial direction we have 𝐞𝑟 = 𝐫̂ . Since the second vector is pointing in an
angular direction, the notation 𝐞𝜃 = 𝜽̂ is often used for the second vector.
Regardless of the future trajectory of an object (in this case, a spacecraft), we can always define two
vectors in this way, with 𝐞𝑟 = 𝐫̂ and 𝐞𝜃 orthogonal, basing their directions on the position of the object.
To an observer back on the 2D earth, two things are particularly interesting about the vectors 𝐞𝑟 and 𝐞𝜃 .
{ }
1. The pair 𝐞𝑟 , 𝐞𝜃 form a basis for the plane, which we call the polar basis (indeed, as we saw
earlier, any two nonparallel vectors form a basis). In fact, the first polar coordinate is precisely the
coordinate of the position of a particle with respect to this basis; that is, we have 𝐫 = 𝑟 𝐞𝑟 .
2. The polar basis vectors depend on the position of the object whose trajectory we are tracing. Im-
{ }
portantly, as the object continues on its trajectory the vectors 𝐞𝑟 , 𝐞𝜃 change with time! One might
find it helpful to label them 𝐞𝑟 (𝑡) and 𝐞𝜃 (𝑡) for this reason.
At two distinct times, say 𝑡1 and 𝑡2 , the polar basis vectors might look quite different.
52
{ }
At any particular point in time we can easily switch our coordinate system between the polar basis 𝐞𝑟 , 𝐞𝜃 ,
as defined by the current position of the object, and the fixed Cartesian basis {𝐢, 𝐣} via
Of course, we need to be careful because the angle 𝜃 changes according to the position of the object (i.e.,
𝜃 depends on time)! This makes sense, as in the particular case of circular motion we encountered we
have 𝜃 = 𝜔𝑡.
Since the polar basis vectors change with time, it will be useful to know the rate at which they change.
By simply applying the product and chain rules we have
𝑑 𝑑 𝑑
𝐞̇ 𝑟 = (cos 𝜃 𝐢 + sin 𝜃 𝐣) = (cos 𝜃) 𝐢 + (sin 𝜃) 𝐣 = − 𝜃̇ sin 𝜃 𝐢 + 𝜃̇ cos 𝜃 𝐣 = 𝜃̇ 𝐞𝜃 ,
𝑑𝑡 𝑑𝑡 𝑑𝑡
and
𝑑 𝑑 𝑑
𝐞̇ 𝜃 = (− sin 𝜃 𝐢 + cos 𝜃 𝐣.) = − (sin 𝜃) 𝐢 + (cos 𝜃) 𝐣 = − 𝜃̇ cos 𝜃 𝐢 − 𝜃̇ sin 𝜃 𝐣 = −𝜃̇ 𝐞𝑟 .
𝑑𝑡 𝑑𝑡 𝑑𝑡
basis.
Example [Link] (a) Circular motion: We have 𝐫 = 𝑟 𝐫̂ = 𝑅 𝐞𝑟 and 𝜃 = 𝜔𝑡. Thus, by the product
rule
𝐯 = 𝐫̇ = 𝑅 𝐞̇𝑟 = 𝑅 𝜃̇ 𝐞𝜃 = 𝑅𝜔 𝐞𝜃 .
This shouldn’t be a surprise since for circular motion we already knew that velocity is orthogonal to
position and that the speed is 𝑣 = 𝑅𝜔. Moreover,
as expected. We can see here how polar coordinates are perfectly suited to circular motion.
(b) General motion: The polar coordinates are far less useful for a general non-circular motion,
where we have, say, 𝐫 = 𝑟(𝑡)̂𝐫 = 𝑟(𝑡) 𝐞𝑟 . In this case, we know nothing of 𝜃(𝑡) and so by the product
rule
𝐯 = 𝐫̇ = 𝑟̇ 𝐞𝑟 + 𝑟 𝐞̇ 𝑟 = 𝑟̇ 𝐞𝑟 + 𝑟𝜃̇ 𝐞𝜃 .
53
2.4.1 Cylindrical polars
While 2𝐷 polars are useful in modelling circular motion, more sophisticated examples of particle motion
in ℝ3 are more suited to a more sophisticated system.
The cylindrical polar coordinates (𝜌, 𝜙, 𝑧) of a point with Cartesian coordinates (𝑥, 𝑦, 𝑧) are defined in
such a way that 𝑥 = 𝜌 cos 𝜙, 𝑦 = 𝜌 sin 𝜙, and the 𝑧-coordinate stays the same. The ranges of the first two
coordinates are 𝑟 ≥ 0 and 0 ≤ 𝜙 < 2𝜋, while the 𝑧 coordinate can obviously take any value in ℝ. It is
easiest to visualize these coordinates when drawn onto a cylinder.
Notice that these vectors are all orthogonal to each other and are of unit length (so form an orthonormal
basis for ℝ3 ). Also, when the 𝑧 coordinate is zero, we have that 𝐞𝜌 and 𝐞𝜙 are precisely the 2𝐷 polar basis
vectors for points the 𝑥𝑦-plane. However, for 𝑧 ≠ 0 the vector 𝐞𝜌 is not parallel to position. As before, the
angle 𝜙 (and in turn, the basis vectors 𝐞𝜌 and 𝐞𝜙 ) depend on the position of the particle, and so depend on
time. We can again differentiate the previous expressions to get
54
To see how this new coordinate system can be useful, let us reconsider Example [Link](b), and express
the motion of the particle in terms of the cylindrical basis vectors. As we will see, this system is perfectly
suited for motion of this kind. For the helical trajectory, the position of the particle at time 𝑡 was
𝑐 𝑐
𝐫 = 𝑅 cos(𝜔𝑡)𝐢 + 𝑅 sin(𝜔𝑡)𝐣 + 𝑐𝑡𝐤 = 𝑅(cos(𝜙)𝐢 + sin(𝜙)𝐣) + 𝜙𝐤 = 𝑅𝐞𝜌 + 𝜙𝐞 ,
𝜔 𝜔 𝑧
where 𝜙 = 𝜔𝑡. Since 𝜙̇ = 𝜔, it is now very easy to write down the velocity and acceleration vectors, and
we see that we can now encode the information about the motion of the particle extremely efficiently:
𝑐 ̇
𝐯 = 𝐫̇ = 𝑅𝐞̇𝜌 + 𝜙 𝐞 = 𝑅𝜙̇ 𝐞𝜙 + 𝑐 𝐞𝑧 = 𝑅𝜔 𝐞𝜙 + 𝑐 𝐞𝑧 ; 𝐚 = 𝐯̇ = 𝑅𝜔 𝐞̇𝜙 = −𝑅𝜔𝜙̇ 𝐞𝜌 = −𝑅𝜔2 𝐞𝜌 .
𝜔 𝑧
For motion that exhibits spherical symmetry, say a central force in ℝ3 , we can use the following system
of coordinates. Consider the following quantities concerning points in ℝ3 :
𝑟 = |𝐫| = distance from the origin; 𝜃 = angle from 𝑧-axis; 𝜙 = angle in the 𝑥𝑦-plane from 𝑥-axis,
These quantities are the spherical polar coordinates (𝑟, 𝜃, 𝜙) of a point in ℝ3 , and the Cartesian coordi-
nates (𝑥, 𝑦, 𝑧) of the point are related via
55
As with the other polar coordinate systems, we can define some spherical basis vectors.
In exact analogy with 2𝐷-polars, the basis vector 𝐞𝑟 = 𝐫̂ is a unit vector in the direction of position (so
that the position of a particle is always given by 𝐫 = |𝐫| 𝐫̂ = 𝑟 𝐞𝑟 ). On the other hand, the basis vector 𝐞𝜙
is precisely the angular basis vector from 2D polars and has no 𝐤 component (note that, confusingly, the
convention in 3D is to use 𝜙 for the angle in the 𝑥𝑦−plane, and not 𝜃, as we do for 2D polars). Once again
we can write these basis vectors in terms of the usual Cartesian basis via
̇ 𝜃 + 𝜙̇ sin(𝜃) 𝐞𝜙 ;
𝐞̇ 𝑟 = 𝜃𝐞 ̇ 𝑟 + 𝜙̇ cos(𝜃) 𝐞𝜙 ;
𝐞̇ 𝜃 = −𝜃𝐞 𝐞̇ 𝜙 = −𝜙̇ sin(𝜃) 𝐞𝑟 − 𝜙̇ cos(𝜃) 𝐞𝜃 .
It follows by the product rule that for general motion 𝐫 = 𝑟 𝐞𝑟 (with the distance 𝑟 depending on time) we
have
𝐯 = 𝐫̇ = ̇ 𝑟 + 𝑟𝐞̇ 𝑟
𝑟𝐞 = ̇ 𝜃 + 𝑟𝜙̇ sin(𝜃) 𝐞𝜙 ,
̇ 𝑟 + 𝑟𝜃𝐞
𝑟𝐞
56
and (with a bit of work!) we have
𝐚 = 𝐯̇ = (̈𝑟 − 𝑟𝜃̇ 2 − 𝑟𝜙̇ 2 sin2 𝜃) 𝐞𝑟 + (𝑟𝜃̈ + 2𝑟̇ 𝜃̇ − 𝑟𝜙̇ 2 sin 𝜃 cos 𝜃) 𝐞𝜃 + (𝑟𝜙̈ sin 𝜃 + 2𝑟̇ 𝜙̇ sin 𝜃 + 2𝑟𝜃̇ 𝜙̇ cos 𝜃) 𝐞𝜙 .
Of course, these formulae are not very interesting (or useful) most of the time, but in the special case of
central forces they are particularly well suited.
Recall that by definition, for a central force we have that 𝐅 = −𝐹 (𝐫) 𝐫̂ = −𝐹 (𝐫) 𝐞𝑟 for some scalar function
𝐹 (𝐫); that is, the force is always parallel to the position. If 𝜃̇ = 𝜙̇ = 0 then the velocity is also parallel to
the position and the particle remains on a line.
Otherwise, at any given time the position and the velocity are nonparallel and therefore define a plane
in ℝ3 passing through the origin. Since the force vector (and therefore the acceleration vector) must also
lie on this plane, the particle can never leave the plane. In this case we can always change our inertial
frame of reference so that this plane is precisely the equatorial 𝑥𝑦-plane corresponding to 𝜃 = 𝜋∕2. We
have in essence, reduced our 3𝐷 problem to a 2𝐷 one.
On the 𝑥𝑦-plane, the first two basis vectors 𝐞𝑟 and 𝐞𝜙 coincide with the 2𝐷 polar basis vectors, while
𝐞𝜃 = −𝐤. Since 𝜃 = 𝜋∕2, we also have 𝜃̇ = 0 and so the velocity simplifies to 𝐯 = 𝑟̇ 𝐞𝑟 + 𝑟𝜙̇ 𝐞𝜙 . Of
particular interest is the angular momentum 𝐉 = 𝐫 × 𝐩 = 𝐫 × (𝑚𝐯), which is
We have already seen that angular momentum is conserved for central forces, although one can show
directly (see Problem Sheet 3) that the third component of angular momentum alone is always preserved
for central forces. Thus the quantity
Furthermore, since the force 𝐅 = −𝐹 (𝐫)𝐞𝑟 is parallel to position, Newton’s second law implies that the
acceleration 𝐚 = (1∕𝑚)𝐅 = −(𝐹 (𝐫)∕𝑚)𝐞𝑟 must be also be parallel to position. It follows that the 𝐞𝜃 and
𝐞𝜙 components of acceleration must vanish in the expression given for general motion above. (In fact,
this observation alone leads to the conclusion that 𝑟2 𝜙̇ sin2 𝜃 is constant.) By comparing the remaining 𝐞𝑟
component of acceleration with that of the force, we must therefore have
( )
̇ 2 ℎ2
−𝐹 (𝐫) = 𝑚(̈𝑟 − 𝑟𝜙 sin 𝜃) = 𝑚 𝑟̈ − 3 .
𝑟
57
Application: (Gravity)
In the special case of a gravitational force, the above has remarkable consequences. Firstly, the gravita-
tional force 𝐅𝑔 = − 𝐺𝑀𝑚
𝑟2
𝐞𝑟 is a central force. Therefore, we immediately have an explanation as to why,
say, a planet orbits the sun within a fixed plane.
But what about the shape of the trajectory of the orbit of the astronomical object? We have
( )
𝐺𝑀𝑚 ℎ2 ̇ 2
− 𝑟2 = −𝐹 (𝐫) = 𝑚 𝑟̈ − 𝑟3 ⟺ − 𝑟𝐺𝑀𝑚
̇
̇ 𝑟 − 𝑟𝑚ℎ
= 𝑚𝑟̈
(𝑟2
) ( 𝑟 )
3
𝑑 𝐺𝑀𝑚 𝑑 1 2 1 ℎ2
⟺ 𝑑𝑡 𝑟
= 𝑑𝑡 2 𝑚𝑟̇ + 2 𝑚 𝑟2 .
The constant 𝑇 above is again the total energy, and we actually have exactly the same conservation of
energy equation we found in Example [Link](b); to see this simply recall that the speed satisfies 𝑣2 =
|𝐯|2 = |𝑟̇ 𝐞𝑟 + 𝑟𝜙̇ 𝐞𝜙 |2 = 𝑟̇ 2 + 𝑟2 𝜙̇ 2 = 𝑟̇ 2 + ℎ2 ∕𝑟2 . The difference this time is that our conservation of energy
equation is written entirely in terms of a scalar parameter 𝑟 and its derivative 𝑟.̇ Such an equation is called
an Ordinary Differential Equation, which will be the focus of the next section of the course. One can
solve this ODE for 𝑟 to find it has the solution
√
ℎ2 2𝑇 ℎ2
𝑟 = , where 𝑈= + 𝐺2 𝑀 2 is a constant .
𝐺𝑀 + 𝑈 cos 𝜙 𝑚
Of course, 𝑟 simply represents how far from the origin the astronomical object is. One might then recog-
nise the shape of the trajectories this solution reveals, depending on the total energy 𝑇 and the constant
ℎ:
Thus, we can explain all possible trajectories of astronomical objects exposed to a single source the gravity.
58
Given a circular or elliptical orbit, one can actually deduce Kepler’s second law of planetary motion
simply from the fact that ℎ = 𝑟2 𝜙̇ is a constant, but we’ll leave that fun fact for a physics module.
In conclusion, we’d already seen how important the language of vectors is in physics, but now we can
also see that it is going to be very important to be able so solve ODEs!
59
60
Chapter 3
3.1 Introduction
Differential equations are defined to be equations involving the derivatives of quantities. They play a
crucial role in modern science, enabling us to mathematically model many real world processes with
astonishing accuracy and predicitive power. Most of the fundamental laws of physics are written in terms
of differential equations. In esscence because they result from the balance of forces which, as we have
seen in the previous chapter result form deirvatives of position vectors.
• Newton’s Second Law. The equation 𝐅 = 𝑚𝐚 relates the force 𝐅 on a particle to its acceleration 𝐚,
which in turn is defined as the second time derivative of its position 𝐫. Thus Newton’s second law
can be written as the differential equation:
𝑑 2𝐫
𝐅=𝑚 .
𝑑𝑡2
This is in some sense the ‘grand-daddy’ of all differential equations, as Newton developed the
concept of differentiation and calculus to deal with this equation!
• Schrödinger’s Equation. Newton’s second law might have reached its ‘best before date’ when
Newtonian Mechanics was superseded by quantum mechanics around a hundred years ago, but
physics simply replaced one differential equation by a new one; Schrodinger’s equation, one version
of which is
ℏ2 𝜕 2 𝜓 𝜕𝜓
− + 𝑉 (𝑥)𝜓 = 𝑖ℏ .
2𝑚 𝜕𝑥2 𝜕𝜓
61
The derivatives in this equation are now ‘partial’ (as indicated by the ‘floppy d’ or partial sign 𝜕),
because the variable 𝜓 = 𝜓(𝑥, 𝑡) depends on two variables 𝑥 and 𝑡 so we have to worry about which
one we have to differentiate by.
𝜕2𝑢 2
2𝜕 𝑢
− 𝑐 =0
𝜕𝑡2 𝜕𝑥2
where you can think of 𝑢(𝑥, 𝑡) as the ‘height’ of the wave above the point 𝑥. The equation appears
in many contexts, in particular in describing electromagnetic fields. Again notice the appearance
of ‘partial derivatives’ as the height of the wave 𝑢 = 𝑢(𝑥, 𝑡) depends on more than one variable.
If you are not convinced yet of the ubiquity of differential equations, let us construct a simple example to
demonstrate the power of differential equations to model situations mathematically
Let 𝑃 (𝑡) be the population of a small kingdom. The current population (at time 𝑡 = 0) is 𝑃 (0) =
10, 000, 000 = 107 . We would like to predict the population in 42 years time. We assume that
each year around 1∕70 of the population die, i.e. 𝑃 ∕70, and the number of children born is 𝑃 ∕60.
Therefore we write that the rate of change of population is
𝑑𝑃 𝑃 𝑃 𝑃
= − = .
𝑑𝑡 60 70 420
We have constructed a differential equation to model the situation, and can use this to find the
population at later times. We have heard about exponential population growth, so guess that 𝑃 (𝑡) =
𝑎𝑒𝑏𝑡 for some constants 𝑎 and 𝑏. Substituting this into the differential equation we find
𝑑𝑃 𝑑 ( 𝑏𝑡 ) 𝑃 𝑎 𝑏𝑡
= 𝑎𝑒 = 𝑎𝑏𝑒𝑏𝑡 = = 𝑒 .
𝑑𝑡 𝑑𝑡 420 420
which is satisfied provided we take 𝑏 = 1∕420. Notice that our differential equation does not
determine the value of 𝑎, so that its solution has one arbitrary parameter 𝑎 in it. On the other hand
we can fix 𝑎 by using the initial value of 𝑃 ; that is
This type of extra information (in this case the initial value of the population) is referred to as an
‘inital’ or ’boundary condition’, and can be used to determine the arbitrary parameters in the general
62
solution to differential equations. Now that we know 𝑎 and 𝑏, we have that
so we can determine the population at all future times. In particular, after 42 years we have that
so the population has grown by around 10%. Not too bad really.
(Of course if you are really pedantic you might complain that 𝑃 (𝑡) should be an integer since people
come in integer numbers,and in the above we have assumed that not only that 𝑃 (𝑡) is a continuous
function but that it is a differentiable function of time, but for large populations this is a pretty
reasonable approximation!)
3.1.1 Definitions
In order to make some sense of the jungle of differential equations, it makes sense to introduce some
definitions which are used to classify them.
• Dependent variables: Quantities which depend on other variables are called dependent variables.
In the example above the population 𝑃 (𝑡) is a dependent variable since it is a function of the time
𝑡. In general these are the quantities we solve differential equations for.
• Independent variables: Quantities which do not depend on other variables are called independent
variables. In the example above the time 𝑡 is an independent variable, which is a parameter which
the solution will depend on.
• Ordinary differential equations: If there is only one independent variable then all the deriva-
tives in the differential equation will be ‘ordinary derivatives’ (i.e. 𝑑𝑃
𝑑𝑡
) and is called an ordinary
differential equation (often abbreviated to an ODE).
• Partial differential equations: If there are more than one independent variables, so that the de-
pendent variables depend on more than one thing, then the derivatives are necessarily ‘partial’,
and the resulting differential equation is a partial differential equation (often abbreviated to PDE).
Schrödinger’s Equation and the Wave equation are examples of PDEs.
63
• Equation order: Differential equations are classified according to their order. A differential equa-
tion is called an equation of order 𝑁 if the highest derivative appearing in the equation is an 𝑁-th
order derivative. So for example
𝑑𝑃
= 𝜇𝑃
𝑑𝑡
is a first order equation, as the highest derivative is the first derivative 𝑑𝑃
𝑑𝑡
.
( )5
𝑑 2𝑥 𝑑𝑥
𝑚 − =0
𝑑𝑡2 𝑑𝑡
𝑑2𝑥
is a second order equation, as the highest derivative is the second derivative 𝑑𝑡2
and
( )4
𝑑 2𝑢 𝑑 3𝑢
= ln(𝑢)
𝑑𝑥2 𝑑𝑢3
𝑑3𝑢
is a third order equation, as the highest derivative appearing is 𝑑𝑢3
.
• The general solution to a differential equation of order 𝑁 has 𝑁 arbitrary parameters in it. For ex-
ample, our Population Model was an equation of order 1, and we found a solution 𝑃 = 𝑎 exp(𝑡∕420),
which contains one arbitrary parameter 𝑎 in it. This is the most general solution to the equation.
In class an (inconsequential) mistake was made. The solution to the population equation:
d𝑃 1
= 𝑃 (3.1.1)
d𝑡 420
where 𝑎 was an arbitrary constant. In class a typo had this being written as:
𝑎 𝑡∕420
𝑃 = e . (3.1.3)
420
so it still satisfies the equation. If we then apply the initial condition 𝑃 (0) = 107 we either get
64
but in both cases we have in the end:
𝑃 (𝑡) = 107 e𝑡∕420 . (3.1.6)
The critical point is that a differential equation alone is not a fully defined mathematical system, it has
ambiguities, the number of such ambiguities, the constants that must be set is dictated by the order of the
equation. This means one must supply 𝑛 additional pieces of information often called boundary or initial
conditions (𝑃 (0) is an initial condition as its at 𝑡 = 0). These can really matter as the following example
shows:
d2 𝑢
2
= −𝑤2 𝑢. (3.1.7)
d𝑡
Whose general solution was shown to be:
𝑢(0) = 𝐴 = 0 (3.1.9)
(you will need to know the basic cos, sin values). Then
d𝑢
(0) = 𝐵𝑤 cos(𝑤 ∗ 0). (3.1.11)
d𝑡
So 𝐵𝑤 = 𝑎 and the solution is:
𝑎
𝑢(𝑡) = sin(𝑤𝑡). (3.1.12)
𝑤
If 𝑎 > 0 we have an oscillatory solution (a wave) if 𝑎 = 0 then the only possible solution is the so-called
trivial solution 𝑢(𝑡) = 0 oscillations are not possible.
Even more extreme, consider the conditions 𝑢(0) = 0 𝑢(𝜋∕𝑤) = 1. The first one we know gives.
so
𝑢(𝜋∕𝑤) = 𝐵 sin(𝜋) = 0 = 1. (3.1.14)
65
Which is impossible! So for certain boundary conditions we cannot actually solve this equation/system.
In fact this is because the frequency 𝑤 sets aside certain special values of 𝑡 which will prove crucial in
chapter 4...
So, in summary, a complete problem must have a differential equation of order 𝑛 and 𝑛 values/conditions
on the funciton in order to be fully determined.
First order ODEs are the simplest class of differential equations. Presented with a first order ODE you
have a fighting chance of being able to solve it exactly, although there are many that cannot. Our approach
will to give a catalogue or classification of some easily solved types of first order ODEs, and then give
the algorithm for solving each type. The four types we will deal with are
a) Separable
b) Homogeneous
c) Linear
d) Bernouilli
Note that these types are NOT mutually exclusive; that is it is perfectly possible for an equation to be, for
example, separable AND Bernouilli or homogeneous AND linear. I have ordered the types in what I think
is the ease of solving them, so separable is the easiest, homogeneous and linear are about the same and
Bernouilli is the hardest to solve. So if you find your equation is separable AND Bernouilli, you should
probably try to solve it as a separable equation because it will be quicker to do so. The routine for solving
first order differential equations will be
2. Apply the algorithm given to find the general solution. This should have one arbitrary parameter in
it (being a first order equation)
3. If a boundary condition is given (e.g. 𝑦(𝑥) = 2 when 𝑥 = 5), use it to determine the value of the
arbitrary parameter.
66
Separable Equations
Now to make a decent equation of this we simply put an integral sign in front of both sides; that is we
write
𝑑𝑦 𝑥2
= .
𝑑𝑥 𝑦
𝑦 𝑑𝑦 = 𝑥2 𝑑𝑥
𝑦 𝑑𝑦 = 𝑥2 𝑑𝑥
∫ ∫
𝑦2 𝑥3
⇒ = +𝐶
2 3√
2𝑥3
⇒𝑦 = ± + 2𝐶.
3
Note that to solve the first order ODE, we have integrated the equation once and this is responsible
for introducing the one arbitrary parameter 𝐶 which appears in the general solution.
67
• Example 2: Population Model No. 2 We might try to model population growth more accurately
by including a term in our equation to reflect that if the population grows too large, then more
people will die due to lack of sufficient food. We shall do this by adding a term which becomes
very negative as the population 𝑃 grows large;
𝑑𝑃
= 𝜇𝑃 − 𝜆𝑃 2 .
𝑑𝑡
If we plot 𝑑𝑃
𝑑𝑡
against 𝑃 we find the following graph: From the figure you see that the population
dP
dt
µ
λ P
⎛ ⎞
𝜆 ⎜ 𝑃 ⎟
⇒ ln(𝑃 ) − ln(1 − 𝑃 ) = ln = 𝜇𝑡 + 𝑐
𝜇 ⎜1 − 𝜆𝑃 ⎟
⎝ 𝜇 ⎠
𝑃
⇒ 𝜆
= 𝑒𝜇𝑡+𝑐 = 𝑒𝜇𝑡 𝑒𝑐 = 𝐷𝑒𝜇𝑡
1 − 𝜇𝑃
𝐷𝑒𝜇𝑡 1
⇒𝑃 = = .
1 + 𝜇 𝐷𝑒𝜇𝑡 𝐷−1 𝑒−𝜇𝑡 + 𝜇𝜆
𝜆
As 𝑡 → ∞, 𝑒−𝜇𝑡 → 0 so
1 𝜇
𝑃 (𝑡) → 𝜆
= ,
0+ 𝜆
𝜇
68
2.0
1.0
1.5
0.8
0.6
1.0
0.4
0.5
0.2
(a) (b)
Homogeneous Equations
𝑑𝑦 𝑑(𝑢𝑥)
= = 𝑓 (𝑢)
𝑑𝑥 𝑑𝑥
𝑑𝑢 𝑑𝑥 𝑑𝑢
⇒ 𝑥+𝑢 = 𝑥 + 𝑢 = 𝑓 (𝑢)
𝑑𝑥 𝑑𝑥 𝑑𝑥
𝑑𝑢 𝑓 (𝑢) − 𝑢
⇒ = .
𝑑𝑥 𝑥
Looking at the last equation for 𝑢, we recognise it as a separable ODE as dealt with in the last section, so
we can solve it by that algorithm; i.e. we write
𝑑𝑢 𝑑𝑥
= .
∫ 𝑓 (𝑢) − 𝑢 ∫ 𝑥
and proceed to calculate the integrals. Let us see how this works in practice.
69
• Example Find the solution to
𝑑𝑦
2𝑥𝑦 − 𝑦2 + 𝑥2 = 0
𝑑𝑥
The right-hand side is a function of the ratio 𝑦∕𝑥 so the equation is homogeneous. Making the
substitution 𝑦∕𝑥 = 𝑢 we have
𝑑𝑦 𝑑(𝑢𝑥) 𝑑𝑢 ( )
1 1
= = 𝑥+𝑢 = 𝑢−
𝑑𝑥 𝑑𝑥 𝑑𝑥 2 𝑢
𝑑𝑢 𝑢 1 1 ( 2 )
⇒ 𝑥 = − −𝑢=− 𝑢 +1
𝑑𝑥 2 2𝑢 2𝑢
2𝑢 1
⇒ 𝑑𝑢 = − 𝑑𝑥
∫ 𝑢2 + 1 ∫ 𝑥
⇒ ln(𝑢2 + 1) = − ln(𝑥) + 𝑐 = ln(1∕𝑥) + 𝑐
𝑒𝑐 𝐷
⇒ 𝑢2 + 1 = =
𝑥√ 𝑥
𝑦 𝐷
⇒ =𝑢 = ± −1
𝑥 𝑥
√
𝐷
⇒ 𝑦 = ±𝑥 − 1.
𝑥
In this case we have the boundary condition 𝑦 = 1 at 𝑥 = 1 which we can use to determine the
arbitrary constant 𝐷 in our general solution
√
𝑦(1) = 1 = ± 𝐷 − 1 ⇒ 𝐷 = 2
so finally
√
2
𝑦=𝑥 − 1.
𝑥
where we take the positive square root so that 𝑦 is positive when 𝑥 = 1. As indicated in panel (b)
√ √
of Fig. 3.1 this tends to zero when 𝑥 → 0 as 2∕𝑥 − 1 ≈ 2∕𝑥 in this limit so 𝑦 → 𝑥 𝑥2 = 2𝑥
in this limit. Then at 𝑥 = 2 the argument of the root hits 0 then becomes negative, hence not real.
This means the solution is only valid (real) for 𝑥 ∈ [0, 2].
70
6
1.0
4
2
0.5
2 4 6 8 10
0.5 1.0 1.5 2.0 2.5 3.0 -2
-0.5 -4
-6
-1.0
(a) (b)
Linear Equations
𝑑𝑦
+ 𝑓 (𝑥)𝑦 = 𝑔(𝑥).
𝑑𝑥
Note that it is the way that 𝑦 appears that makes the equation linear; only 𝑑𝑦
𝑑𝑥
and 𝑦 itself appear, rather
than any powers or other complicated functions of 𝑦 or its derivative. On the other hand 𝑥 can appear
pretty much as it likes; the functions 𝑓 (𝑥) and 𝑔(𝑥) are arbitrary functions!
𝑑(𝑦𝑥2 ) 𝑑𝑦 2
= 𝑥 + 𝑦.2𝑥.
𝑑𝑥 𝑑𝑥
Suppose now I have the ODE
𝑑𝑦 2 1
+ 𝑦 = 3.
𝑑𝑥 𝑥 𝑥
Comparing with the general form of linear equations given above, I recognise that this is a linear equation
with 𝑓 (𝑥) = 2∕𝑥 and 𝑔(𝑥) = 1∕𝑥3 . If I multiply through this equation by the factor 𝑥2 , then it becomes
𝑑𝑦 2 1
𝑥 + 𝑦.2𝑥 = ,
𝑑𝑥 𝑥
but I recognise that I can rewrite the left hand side as the derivative of a product as above; that is
𝑑𝑦 2 𝑑(𝑦𝑥2 ) 1
𝑥 + 𝑦.2𝑥 = =
𝑑𝑥 𝑑𝑥 𝑥
71
I can integrate both sides of the equation with respect to 𝑥 and find
𝑑(𝑦𝑥2 ) 1
𝑑𝑥 = 𝑦𝑥2 = 𝑑𝑥 = ln(𝑥) + 𝑐
∫ 𝑑𝑥 ∫ 𝑥
ln(𝑥) + 𝑐
⇒𝑦 = ,
𝑥2
so solving the equation. This function is shown in Figure ??(a) for 𝑐 = 1, it has a negative asymptote
at 𝑥 → 0 due to the logarithm then as 𝑥 gets large the 1∕𝑥2 decay takes over. It has an intermediate
maximum (can you find it?). The form is quite similar to a repulsive/attractive potential common in
molecular chemistry, where the maximum would be the ideal molecule separation.
It turns out we can always do this, that is multiply the linear equation by a function, the so called integrating
factor 𝑤(𝑥), such that the left hand side turns into the derivative of the product of 𝑦𝑤(𝑥). To see that this
works compare the linear equation multiplied through by 𝑤(𝑥)
𝑑𝑦
𝑤(𝑥) + 𝑤(𝑥)𝑓 (𝑥)𝑦 = 𝑤(𝑥)𝑔(𝑥)
𝑑𝑥
with the derivative of the product
𝑑 𝑑𝑦 𝑑𝑤
(𝑤(𝑥)𝑦) = 𝑤(𝑥) + 𝑦.
𝑑𝑥 𝑑𝑥 𝑑𝑥
The left hand sides of these two expressions match provided that
𝑑𝑤
= 𝑤(𝑥)𝑓 (𝑥)
𝑑𝑥
𝑑𝑤
⇒ = 𝑓 (𝑥)𝑑𝑥
∫ 𝑤 ∫
( )
⇒ ln(𝑤) = 𝑓 (𝑥)𝑑𝑥 ⇒ 𝑤(𝑥) = exp 𝑓 (𝑥)𝑑𝑥 .
∫ ∫
( )
Choosing the function 𝑤(𝑥) = exp ∫ 𝑓 (𝑥)𝑑𝑥 we can rewrite the linear equation
𝑑𝑦 𝑑 1
𝑤(𝑥) + 𝑤(𝑥)𝑓 (𝑥)𝑦 = (𝑤(𝑥)𝑦) = 𝑤(𝑥)𝑔(𝑥) ⇒ 𝑦 = 𝑤(𝑥)𝑔(𝑥)𝑑𝑥.
𝑑𝑥 𝑑𝑥 𝑤(𝑥) ∫
• Example Find 𝑦(𝑥) satisfying the differential equation
𝑑𝑦
+ cot(𝑥)𝑦 = cos(𝑥)
𝑑𝑥
This ODE is of the linear form with 𝑓 (𝑥) = cot(𝑥) and 𝑔(𝑥) = cos(𝑥). The integrating factor is
given by
( ) ( )
cos(𝑥)
𝑤(𝑥) = exp 𝑓 (𝑥)𝑑𝑥 = exp 𝑑𝑥 = exp(ln(sin(𝑥))) = sin(𝑥),
∫ ∫ sin(𝑥)
72
so multiplying the equation by the integrating factor 𝑤(𝑥) = sin(𝑥) we find
𝑑𝑦 cos(𝑥)
sin(𝑥) + sin(𝑥) 𝑦 = cos(𝑥) sin(𝑥)
𝑑𝑥 sin(𝑥)
𝑑𝑦 𝑑
⇒ sin(𝑥) + cos(𝑥)𝑦 = (sin(𝑥)𝑦) = cos(𝑥) sin(𝑥)
𝑑𝑥 𝑑𝑥
sin2 (𝑥)
⇒ sin(𝑥)𝑦 = cos(𝑥) sin(𝑥)𝑑𝑥 = +𝐶
∫ 2
sin(𝑥) 𝐶
⇒𝑦 = + .
2 sin(𝑥)
This solution is shown in 3.2(b) for various values of 𝐶. The 1∕ sin(𝑥) behaviour leads to periodic asymp-
totes at 𝑥 = 0, 𝜋, 2𝜋, … . Mostly this dominates the sin(𝑥)∕2 term, but if 𝐶 is quite small we see some
sign this function oscillating superimposed on the 𝐶∕ sin(𝑥) function.
Bernouilli Equations
𝑑𝑦
𝑦−𝑛 + 𝑓 (𝑥)𝑦1−𝑛 = 𝑔(𝑥).
𝑑𝑥
Now we make the substitution 𝑧 = 𝑦1−𝑛 . Noting that
𝑑𝑧 𝑑(𝑦1−𝑛 ) 𝑑𝑦 𝑑𝑦 1 𝑑𝑧
= = (1 − 𝑛)𝑦−𝑛 ⇒ 𝑦−𝑛 = ,
𝑑𝑥 𝑑𝑥 𝑑𝑥 𝑑𝑥 1 − 𝑛 𝑑𝑥
we see that our ODE can be written in terms of 𝑧 as
1 𝑑𝑧
+ 𝑓 (𝑥)𝑧 = 𝑔(𝑥)
1 − 𝑛 𝑑𝑥
𝑑𝑧
⇒ + (1 − 𝑛)𝑓 (𝑥)𝑧 = (1 − 𝑛)𝑔(𝑥).
𝑑𝑥
This is now a linear equation for 𝑧, so we can solve it with an integrating factor using the method described
in the last section.
73
• Example 1 Find the solution to the ODE
𝑑𝑦
+ 𝑦 = 𝑥𝑦3
𝑑𝑥
In this case 𝑛 = 3 so that 𝑧 = 𝑦1−𝑛 = 𝑦1−3 = 𝑦−2 , that is 𝑦 = 𝑧−1∕2 . In terms of 𝑧 the equation
becomes
𝑑(𝑧−1∕2 )
+ 𝑧−1∕2 = 𝑥𝑧−3∕2
𝑑𝑥
1 𝑑𝑧
⇒ − 𝑧−3∕2 + 𝑧−1∕2 = 𝑥𝑧−3∕2
2 𝑑𝑥
𝑑𝑧
⇒ − 2𝑧 = −2𝑥.
𝑑𝑥
This as promised is a linear equation for 𝑧, so we proceed to solve it by finding the integrating factor;
( )
𝑤(𝑥) = exp −2𝑑𝑥 = 𝑒−2𝑥 .
∫
𝑑𝑧
𝑒−2𝑥 − 2𝑒−2𝑥 𝑧 = −2𝑥𝑒−2𝑥
𝑑𝑥
𝑑(𝑒−2𝑥 𝑧)
⇒ = −2𝑥𝑒−2𝑥
𝑑𝑥
⇒ 𝑒−2𝑥 𝑧 = −2𝑥𝑒−2𝑥 𝑑𝑥
∫
= 𝑥𝑒−2𝑥 − 𝑒−2𝑥 𝑑𝑥
∫
1
= 𝑥𝑒−2𝑥 + 𝑒−2𝑥 + 𝐶
2
1
⇒ 𝑧 = 𝑥 + + 𝐶𝑒2𝑥
2
( )−1∕2
1
⇒ 𝑦 = 𝑧−1∕2 = 𝑥 + + 𝐶𝑒2𝑥 .
2
We also have that 𝑦(0) = 1 which implies that 1 = (1∕2 + 𝐶)−1∕2 so 𝐶 = 1∕2 and
( )−1∕2
1 1
𝑦 = 𝑥 + + 𝑒2𝑥 .
2 2
𝑑𝑃
= 𝜇𝑃 − 𝜆𝑃 2
𝑑𝑡
74
where 𝜇 corresponded to the difference in the rates of births and deaths, and 𝜆 represented the extra
deaths incurred when the population became large as a result of food shortages. To refine our model
further, we shall now take into account technical innovations which mean that the world’s ability
to produce food improves with time. We shall take 𝜆 to be exponentially decaying with time, i.e.
𝜆 = 𝜆0 𝑒−𝜉𝑡 for some constants 𝜆0 and 𝜉. This means that the negative term in the expression for
𝑃̇ becomes smaller as time advances, so that the population 𝑃 has to grow larger before this term
plays a role in decreasing the population growth.
75
If 𝜇 > 𝜉 then as 𝑡 → ∞ we will have 𝑒−𝜇𝑡 ≪ 𝑒−𝜉𝑡 and
( )−1
𝜆0 −𝜉𝑡 (𝜇 − 𝜉)𝑒𝜉𝑡 (𝜇 − 𝜉)
𝑃 ≈ 𝑒 = =
(𝜇 − 𝜉) 𝜆0 𝜆(𝑡)
In this case the population growth because of excess births of deaths parametrised by 𝜇 is faster
than technology innovation so the population reaches its maximum size when it runs out of food,
and growth is then determined by the (slower) improvement in food given by 𝑒𝜉𝑡 . If 𝜇 ≫ 𝜉 then we
recover the answer 𝑃 = 𝜇∕𝜆 which is what Population Model No. 2 gave.
If on the other hand 𝜉 > 𝜇 then technological advancement is fast enough such that the supply of
food always outstrips the natural exponential population growth. In this case as 𝑡 → ∞ we will
have 𝑒−𝜇𝑡 ≫ 𝑒−𝜉𝑡 and
( )−1 1
𝑃 ≈ 𝐶𝑒−𝜇𝑡 = 𝑒𝜇𝑡
𝐶
which is the exponential growth we had from Population model no. 1. In this case 𝜆 decays fast
enough that the negative term proportional to 𝑃 2 in the equation is always small and never plays a
significant rôle.
• Second order equations are generally much more difficult to solve than first order equations. In
particular non-linear second order differential equations are very difficult and can rarely be solved
exactly.
• The general solution to a second order differential equation has two arbitrary parameters in it (as
opposed to the one arbitrary constant we found for first order equations). This is because in some
sense we need to do two indefinite integrations to solve a second order equation, and each integration
introduces an arbitrary constant
• BUT second order equations are quite common in applications, for example Newton’s Second Law.
Fortunately the equations that appear are very often linear, which makes them easier to solve.
76
Note that ‘linear’ constrains how the dependent variable 𝑦(𝑥) enters the equation, not how 𝑥 appears, just
as in the first order case. When 𝐶(𝑥) = 0 we call the equation homogeneous; when 𝐶(𝑥) ≠ 0 the equation
is called inhomogeneous. Just to be confusing, note that this use of the word homogeneous in this case
has nothing whatsoever to do with the word homogeneous in the first-order case!
Homogeneous equations are relatively easy to solve because of the principle of superposition. This says
that if 𝑦1 (𝑥) and 𝑦2 (𝑥) are two specific solutions to the equation then any linear combination 𝑦 = 𝛼𝑦1 (𝑥) +
𝛽𝑦2 (𝑥) (where 𝛼 and 𝛽 are constants) is also a solution. This is easy to see by substituting 𝑦 into the
equation:
𝑑 2𝑦 𝑑𝑦
+ 𝐴(𝑥) + 𝐵(𝑥)𝑦 =
𝑑𝑥2 𝑑𝑥
𝑑2 ( ) 𝑑 ( ) ( )
𝛼𝑦 1 (𝑥) + 𝛽𝑦 2 (𝑥) + 𝐴(𝑥) 𝛼𝑦 (𝑥) + 𝛽𝑦 (𝑥) + 𝐵(𝑥) 𝛼𝑦 (𝑥) + 𝛽𝑦 (𝑥)
𝑑𝑥 2
( 2 𝑑𝑥 (1
)
2 1
)
2
2
𝑑 𝑦1 𝑑𝑦1 𝑑 𝑦2 𝑑𝑦2
= 𝛼 + 𝐴(𝑥) + 𝐵(𝑥)𝑦 1 + 𝛽 + 𝐴(𝑥) + 𝐵(𝑥)𝑦2
𝑑𝑥2 𝑑𝑥 𝑑𝑥2 𝑑𝑥
= 𝛼.0 + 𝛽.0 = 0
where in the last line we have used the fact that both 𝑦1 (𝑥) and 𝑦2 (𝑥) are solutions. If we can find two
solutions 𝑦1 (𝑥) and 𝑦2 (𝑥) which are independent (i.e. not proportional to one another) from these we
can construct a solution by linearly combining them which depends on two arbitrary parameters, and this
means it is the general solution. So knowing any two independent particular solutions we can find all the
solutions. This is what makes linear equations easier to solve than non-linear equations; in the latter case,
if you know two particular solutions, that is all you know!
The easiest class of second-order equations to solve are those of the form
𝑑 2𝑦 𝑑𝑦
𝑎 2
+𝑏 + 𝑐𝑦 = 0,
𝑑𝑥 𝑑𝑥
where 𝑎, 𝑏 and 𝑐 are all constants. Following the previous section we know that if we can find two
independent solutions then we can find all the solutions. We might guess that the equation might be
solved by some sort of exponential function, since the differentiating an exponential function gives back
some multiple of the original equation, so there is a good chance that the all the terms in the above equation
might cancel with each other. Let us a guess a function of the form 𝑦 = 𝑒𝜆𝑥 might satisfy our equation.
77
Substituting it in we find
( )
𝑎𝜆2 𝑒𝜆𝑥 + 𝑏𝜆𝑒𝜆𝑥 + 𝑐𝑒𝜆𝑥 = 0 ⇒ 𝑎𝜆2 + 𝑏𝜆 + 𝑐 𝑒𝜆𝑥 = 0,
𝑎𝜆2 + 𝑏𝜆 + 𝑐 = 0.
This latter equation is called the auxiliary equation and generally it has two solutions:
√
−𝑏 ± 𝑏2 − 4𝑎𝑐
𝜆± = .
2𝑎
Thus in general we have found two particular solutions 𝑦1 = 𝑒𝜆+ 𝑥 and 𝑦2 = 𝑒𝜆− 𝑥 and can therefore write
the general solution to the equation as
However the nature of these solutions will turn out to be different depending on whether the roots 𝜆± to
the auxiliary equation are real, complex, or repeated. We will deal with each of these three cases in turn
1. 𝑏2 > 4𝑎𝑐 ∶ In this case the auxiliary equation has two distinct, real roots.
𝜆2 − 3𝜆 + 1 = 0
√ √
3± 9−4 3 5
⇒ 𝜆± = = ±
2 2 2
2. 𝑏2 < 4𝑎𝑐 In this case the auxiliary equation has two distinct, complex roots.
78
We find the auxiliary equation
𝜆2 − 2𝜆 + 5 = 0
√ √
2 ± 4 − 20 −16
⇒ 𝜆± = =1± = 1 ± 2𝑖
2 2
Strictly this is the most general solution, but the most general complex solution, as it contains
explicit ‘𝑖’s. To find the most general real solution it helps to remember that 𝑒𝑖𝜃 = cos(𝜃) +
𝑖 sin(𝜃) so we can write
( )
𝑦 = 𝑒𝑥 𝛼𝑒2𝑖𝑥 + 𝛽𝑒−2𝑖𝑥
= 𝑒𝑥 (𝐶 cos(2𝑥) + 𝐷 sin(2𝑥))
3. 𝑏2 = 4𝑎𝑐: In this case there is only one real root to the auxiliary equation,
√
−𝑏 ± 𝑏2 − 4𝑎𝑐 𝑏
𝜆± = =− .
2𝑎 2𝑎
This means that 𝜆+ = 𝜆− and our method of guessing 𝑦 = 𝑒𝜆𝑥 produces only one particular solution
since 𝑦1 (𝑥) = 𝑒𝜆+ 𝑥 = 𝑦2 (𝑥). No longer can we generate the general solution containing two arbitrary
parameters by taking linear combinations of 𝑦1 and 𝑦2 , since 𝛼𝑦1 + 𝛽𝑦2 = (𝛼 + 𝛽)𝑦1 = 𝛾𝑦1 where
we take 𝛾 = 𝛼 + 𝛽.
To understand where the missing solution to our equation has gone and how to recover it, consider
an ODE where the roots of the auxiliary equation are very nearly but not exactly the same; let us
call them 𝜆 and 𝜆+𝜖 where 𝜖 is very small. As the equation still has two distinct, albeit very similar,
79
roots, the equation will still have the solutions 𝑦1 = 𝑒𝜆𝑥 and 𝑦2 = 𝑒(𝜆+𝜖)𝑥 . But note that if we let
𝜖 → 0 the combination 𝑦2 − 𝑦1 = 𝑒(𝜆+𝜖)𝑥 − 𝑒𝜆𝑥 will disappear too. This is how our solution goes
missing, but we can recover it if we scale it by the (large) number 1∕𝜖. Then we have
𝑦2 − 𝑦1 𝑒(𝜆+𝜖)𝑥 − 𝑒𝜆𝑥 𝑑(𝑒𝜆𝑥 )
lim = lim = = 𝑥𝑒𝜆𝑥 ,
𝜖→0 𝜖 𝜖→0 𝜖 𝑑𝜆
where we have used the limit definition of a derivative to calculate the limit. Thus our ‘missing’
solution is 𝑥𝑒𝜆𝑥 and we can recover the general solution to an ODE with repeated roots as
𝜆2 − 6𝜆 + 9 = (𝜆 − 3)2 = 0
which has a single real root 𝜆 = 3. By the above the general solution to the ODE is
𝑦 = (𝛼 + 𝛽𝑥)𝑒3𝑥 .
An inhomogeneous equation
𝑑 2𝑦 𝑑𝑦
𝑎 2
+𝑏 + 𝑐𝑦 = 𝑓 (𝑥)?
𝑑𝑥 𝑑𝑥
contains a term (the right-hand side above) which is not zero and has no 𝑦 in it. Let us begin by considering
an example.
• Example
𝑑 2𝑦 𝑑𝑦
− 3 + 2𝑦 = 𝑥.
𝑑𝑥2 𝑑𝑥
We start by looking for any particular solution 𝑦𝑝 to our equation. Since the right-hand side is
a polynomial (albeit a very simple polynomial) and the derivatives of polynomials are just other
polynomials we might guess that substituting a polynomial for 𝑦 might work. So guessing 𝑦𝑝 =
𝑟 + 𝑠𝑥 and substituting into the equation we find
80
we see that to balance both sides we should take 𝑠 = 1∕2 and 𝑟 = 3∕4, so that
3 𝑥
𝑦𝑝 = + .
4 2
It is no longer true that if 𝑦𝑝 solves the equation then a multiple 𝛼𝑦𝑝 solves the equation, since
substituting 𝛼𝑦𝑝 into the left-hand side will give 𝛼𝑥, not 𝑥 as would be required for the equation to
be satisfied. Instead suppose that we have found another solution 𝑦, and now consider 𝑦𝐻 = 𝑦 − 𝑦𝑝 .
Substituting 𝑦𝐻 into the left-hand side of the equation we find that
𝑑 2 𝑦𝐻 𝑑𝑦𝐻 𝑑 2 (𝑦 − 𝑦𝑝 ) 𝑑(𝑦 − 𝑦𝑝 )
2
−3 + 2𝑦𝐻 = 2
−3 + 2(𝑦 − 𝑦𝑝 )
𝑑𝑥 𝑑𝑥 𝑑𝑥 𝑑𝑥 ( )
( 2 ) 2
𝑑 𝑦 𝑑𝑦 𝑑 𝑦 𝑝 𝑑𝑦𝑝
= 2
−3 + 2𝑦 − 2
−3 + 2𝑦𝑝
𝑑𝑥 𝑑𝑥 𝑑𝑥 𝑑𝑥
= (𝑥) − (𝑥) = 0
so we see that 𝑦𝐻 , the difference between the two solutions 𝑦 and 𝑦𝑝 to the inhomogeneous equation,
𝑑 2 𝑦𝐻 𝑑𝑦𝐻
solves the homogeneous equation 𝑑𝑥2
−3 𝑑𝑥
+ 2𝑦𝐻 = 0. In this case the auxiliary equation is
𝜆2 − 3𝜆 + 2 = 0 ⇒ (𝜆 − 1)(𝜆 − 2) = 0 ⇒ 𝜆 = 1, 2
3 𝑥
𝑦 = 𝑦𝐻 + 𝑦𝑝 = 𝛼𝑒𝑥 + 𝛽𝑒2𝑥 + +
4 2
Note that since 𝑦 has inherited two arbitrary constants from 𝑦𝐻 , it must be the most general solution
to the inhomogeneous equation.
From this example we can see that the most general solution 𝑦 to an inhomogeneous equation can be
written in the form 𝑦 = 𝑦𝐻 + 𝑦𝑝 where
1. 𝑦𝑝 is any particular solution we can found which satisfies the inhomogeneous equation
2. 𝑦𝐻 is the general solution to the homogeneous equation, found by solving the auxiliary equation
etc..
Finding 𝑦𝐻 is entirely algorithmic, but in our example we simply guessed a form for 𝑦𝑝 . Guessing is
generally the fastest way to find 𝑦𝑝 , but there are some functions 𝑓 (𝑥) appearing commonly which it is
good to learn what the best guesses are. We shall go through these in turn:
81
• Case 1: 𝑓 (𝑥) = 𝑎𝑒𝑏𝑥
In this case a good guess is 𝑦𝑝 = 𝑟𝑒𝑏𝑥 since differentiating exponentials gives back exponentials.
𝑑 2𝑦
− 9𝑦 = 2𝑒5𝑥 .
𝑑𝑥2
𝑦𝐻 = 𝛼𝑒3𝑥 + 𝛽𝑒−3𝑥 .
Given that 𝑓 (𝑥) = 2𝑒5𝑥 , we guess that 𝑦𝑝 is of the form 𝑦𝑝 = 𝑟𝑒5𝑥 . Substituting this into the
equation we have
1 𝑒5𝑥
25𝑟𝑒5𝑥 − 9𝑟𝑒5𝑥 = 2𝑒5𝑥 ⇒ 𝑟 = ⇒ 𝑦𝑝 = .
8 8
𝑒5𝑥
𝑦 = 𝑦𝐻 + 𝑦𝑝 = 𝛼𝑒3𝑥 + 𝛽𝑒−3𝑥 + .
8
The guess 𝑦𝑝 = 𝑟𝑒𝑏𝑥 does not work if 𝜆 = 𝑏 is a solution to the auxiliary equation. In this case 𝑒𝑏𝑥
solves the homogeneous equation, so by definition, substituting in 𝑦𝑝 = 𝑟𝑒𝑏𝑥 into the left-hand side
of the ODE will give zero, not the right-hand side 𝑓 (𝑥) = 𝑎𝑒𝑏𝑥 . In this case the correct guess is
𝑦𝑝 = 𝑟𝑥𝑒𝑏𝑥 .
82
we find that
𝑑 2 𝑦𝑝 𝑑 2 (𝑟𝑥𝑒3𝑥 )
− 9𝑦𝑝 = − 9𝑟𝑥𝑒3𝑥
𝑑𝑥2 𝑑𝑥2
( )
= 6𝑟𝑒3𝑥 + 9𝑟𝑥𝑒3𝑥 − 9𝑟𝑥𝑒3𝑥 = 6𝑟𝑒3𝑥
⇒ 6𝑟𝑒3𝑥 = 𝑒3𝑥
1
⇒𝑟 =
6
1 3𝑥
⇒ 𝑦𝑝 = 𝑥𝑒 .
6
. Putting this together we have
1
𝑦 = 𝑦𝐻 + 𝑦𝑝 = 𝑥𝑒3𝑥 + 𝛼𝑒3𝑥 + 𝛽𝑒−3𝑥 .
6
• Case 1b: Exception to the exception to the rule for 𝑓 (𝑥) = 𝑎𝑒𝑏𝑥 .
If the auxiliary equation is 𝜆2 − 2𝑏𝜆 + 𝑏2 = (𝜆 − 𝑏)2 = 0, it has a repeated solution 𝜆 = 𝑏. In this case
we have seen that the solution to the homogeneous equation is 𝑦𝐻 = (𝛼 + 𝛽𝑥)𝑒𝑏𝑥 . This means that
substituting in either 𝛼𝑒𝑏𝑥 or 𝛽𝑥𝑒𝑏𝑥 into the left-hand side of the equation will give zero satisfying
the homogeneous equation, and so cannot solve the inhomogeneous equation. The correct guess is
found by multiplying by 𝑥 again, that is we take 𝑦𝑝 = 𝑟𝑥2 𝑒𝑏𝑥 .
𝜆2 − 6𝜆 + 9 = (𝜆 − 3)2 = 0
so the solution to the homogeneous equation is 𝑦𝐻 = (𝛼 + 𝛽𝑥)𝑒3𝑥 . We should try 𝑦𝑝 = 𝑟𝑥2 𝑒3𝑥 .
Substituting into the left-hand side of the equation we have
𝑑 2 𝑦𝑝 𝑑𝑦𝑝 𝑑 2 (𝑟𝑥2 𝑒3𝑥 ) 𝑑(𝑟𝑥2 𝑒3𝑥 )
−6 + 9𝑦𝑝 = − 6 + 9𝑟𝑥2 𝑒3𝑥
𝑑𝑥2 𝑑𝑥 𝑑𝑥2 𝑑𝑥
( ) ( )
= 2𝑟𝑒3𝑥 + 12𝑟𝑥𝑒3𝑥 + 9𝑟𝑥2 𝑒3𝑥 − 6 2𝑟𝑥𝑒3𝑥 + 3𝑟𝑥2 𝑒3𝑥 + 9𝑟𝑥2 𝑒3𝑥
= 2𝑟𝑒3𝑥
⇒ 2𝑟𝑒3𝑥 = 𝑒3𝑥
1
⇒𝑟 =
2
1 2 3𝑥
⇒ 𝑦𝑝 = 𝑥𝑒 .
2
83
. Putting this together we have
( )
1 𝑥2
𝑦 = 𝑦𝐻 + 𝑦𝑝 = 𝑥2 𝑒3𝑥 + (𝛼 + 𝛽𝑥)𝑒3𝑥 = 𝛼 + 𝛽𝑥 + 𝑒3𝑥 .
2 2
• Case 2: 𝑓 (𝑥) = 𝐴 cos(𝜇𝑥) + 𝐵 sin(𝜇𝑥) In this case the correct guess is to take 𝑦𝑝 = 𝑟 cos(𝜇𝑥) +
𝑠 sin(𝜇𝑥).
𝑑 2𝑦 𝑑𝑦
2
−4 + 3𝑦 = 2 sin(3𝑥).
𝑑𝑥 𝑑𝑥
The auxiliary equation in this case is
𝜆2 − 4𝜆 + 3 = (𝜆 − 3)(𝜆 − 1) = 0
𝑑 2 𝑦𝑝 𝑑𝑦𝑝
−4 + 3𝑦𝑝 = (−9𝑟 cos(3𝑥) − 9𝑠 sin(3𝑥)) − 4 (−3𝑟 sin(3𝑥) + 3𝑠 cos(3𝑥))
𝑑𝑥2 𝑑𝑥
+3 (𝑟 cos(3𝑥) + 𝑠 sin(3𝑥))
−12𝑠 − 6𝑟 = 0 ⇒ 𝑟 = −2𝑠
12𝑟 − 6𝑠 = 2
1 2
⇒ −24𝑠 − 6𝑠 = 2 ⇒ 𝑠 = − , 𝑟=
15 15
2 1
⇒ 𝑦𝑝 = cos(3𝑥) − sin(3𝑥).
15 15
Putting this together we have
2 1
𝑦 = 𝑦𝐻 + 𝑦𝑝 = 𝛼𝑒3𝑥 + 𝛽𝑒−𝑥 + cos(3𝑥) − sin(3𝑥).
15 15
84
𝑦𝑝 . Therefore substituting 𝑦𝑝 = 𝛼 cos(𝜇𝑥) + 𝛽 sin(𝜇𝑥) into the left-hand side of our equation will
give zero, not 𝑓 (𝑥) = 𝐴 cos(𝜇𝑥) + 𝐵 sin(𝜇𝑥). The solution as ever is to multiply by 𝑥; that is we
should guess 𝑦𝑝 = 𝑟𝑥 cos(𝜇𝑥) + 𝑠𝑥 sin(𝜇𝑥).
𝜆2 + 4 = 0
𝑑𝑦𝑝
= 𝑟 cos(2𝑥) + 𝑠 sin(2𝑥) − 2𝑟𝑥 sin(2𝑥) + 2𝑠𝑥 cos(2𝑥)
𝑑𝑥
𝑑 2 𝑦𝑝
= −4𝑟 sin(2𝑥) + 4𝑠 cos(2𝑥) − 4𝑟𝑥 cos(2𝑥) − 4𝑠𝑥 sin(2𝑥).
𝑑𝑥2
𝑑 2 𝑦𝑝 𝑑𝑦𝑝
−4 + 3𝑦𝑝 = (−4𝑟 sin(2𝑥) + 4𝑠 cos(2𝑥) − 4𝑟𝑥 cos(2𝑥) − 4𝑠𝑥 sin(2𝑥))
𝑑𝑥2 𝑑𝑥
+4 (𝑟𝑥 cos(2𝑥) + 𝑠𝑥 sin(2𝑥))
1
𝑦 = 𝑦𝐻 + 𝑦𝑝 = 𝛼 cos(2𝑥) + 𝛽 sin(2𝑥) − 𝑥 cos(2𝑥).
4
If 𝑓 (𝑥) is an 𝑛-th order polynomial then the correct guess is to also take 𝑦𝑝 to be an 𝑛-th order
polynomial; that is
𝑦𝑝 = 𝑏𝑛 𝑥𝑛 + 𝑏𝑛−1 𝑥𝑛−1 + ⋯ + 𝑏1 𝑥 + 𝑏0 .
85
– Example Solve the equation
𝑑 2𝑦
2
− 𝑦 = 𝑥2 .
𝑑𝑥
The auxiliary equation in this case is
𝜆2 − 1 = 0
𝑦 = 𝑦𝐻 + 𝑦𝑝 = 𝛼𝑒𝑥 + 𝛽𝑒−𝑥 − 𝑥2 − 2
Again if the auxiliary solution has one(two) solutions of the form 𝜆 = 0, then one needs to multiply
by 𝑥 once (twice). That is one should take 𝑦𝑝 to be an (𝑛+1)-th ((𝑛+2)-th) order polynomial instead
of an 𝑛-th order polynomial.
In this case the guess for 𝑦𝑝 should be the corresponding sum of guesses for cases 1-3.
Although usually the fastest way to solve an inhomogeneous equation is to make an educated guess for
the particular integral, sometimes the term on the right-hand side does not fall into one of the categories
listed in the preceding section. In such cases one can sometimes still solve the equation but turning the
86
second order linear equation into a pair of first order linear equations which can be solved in turn. We shall
illustrate the method through an example. There are two methods which can prove fruitful, and which we
will practice in te first instance on the same equation
We expect our solution to have both homogeneous exponential behaviour and some other unknown
behaviour and try for a solution which packages them together as:
Where we seek to get an equation for 𝑣(𝑥), the trailing function. We have, using the product rule:
d𝑦 d𝑣
= 𝜆e𝜆𝑥 𝑣 + e𝜆𝑥 ,
d𝑥 d𝑥
d2 𝑦 2 𝜆𝑥 𝜆𝑥 d𝑣
2
𝜆𝑥 d 𝑣
= 𝜆 e 𝑣 + 2𝜆e + e .
d𝑥2 d𝑥 d𝑥2
Substituting the second derivative into our equation and collecting terms we find
( )
𝜆𝑥
( 2 ) 𝜆𝑥 d2 𝑣 d𝑣
e 𝑣 𝜆 − 1 +e 2
+ 2𝜆 = 𝑥𝑒𝑥 .
⏟⏞⏞⏞⏞⏞⏟⏞⏞⏞⏞⏞⏟ d𝑥 d𝑥
Homog eqn
The first part, as indicated, is basically the homogeneous equation so 𝜆 = ±1. This leaves an
equation:
d2 𝑣 d𝑣
2
+ 2𝜆 = 𝑥𝑒(1−𝜆)𝑥 .
d𝑥 d𝑥
If we set 𝑢 = d𝑣
d𝑥
then we have a linear fist order equation:
d𝑢
+ 2𝜆𝑢 = 𝑥𝑒(1−𝜆)𝑥 .
d𝑥
If we choose 𝜆 = 1 then you should confirm using the integrating factor that we get:
𝑥2 𝑥
𝑢(𝑥) = 𝐶e2𝑥 + 𝑥∕2 − 1∕4, ⇒ 𝑣(𝑥) = 𝐶1 e−2𝑥 + − + 𝐶2
4 4
87
for constants 𝐶1 and 𝐶2 . Now, we have
𝑦(𝑥) = e𝑥 𝑣(𝑥),
𝑥2 𝑥 𝑥2 𝑥
= e𝑥 (𝐶1 e−2𝑥 + − + 𝐶2 ) = 𝐶1 e𝑥 + 𝐶2 e−𝑥 + e𝑥 − e𝑥 .
4 4 4 4
Note the first part is the Homogenous solution! You should confirm for yourself that if we had
instead used the 𝜆 = −1 case we would have got exactly the same answer.
Instead, we begin as normal by solving the auxiliary equation; in this case we have
𝜆2 − 1 = (𝜆 − 1)(𝜆 + 1) = 0,
so that 𝜆 = ±1. Rather than write down the homogeneous solution 𝑦𝐻 , instead we remark that the
differential operator in our ODE can be factorised in the same way as the auxiliary equation; that is
𝑑 2𝑦
2
− 𝑦 = 𝑥𝑒𝑥
( 2 𝑑𝑥 )
𝑑
⇒ − 1 𝑦 = 𝑥𝑒𝑥
𝑑𝑥2
( )( )
𝑑 𝑑
⇒ −1 + 1 𝑦 = 𝑥𝑒𝑥 .
𝑑𝑥 𝑑𝑥
To turn our second order linear equation into two first order equations, we define
( ) 𝑑𝑦
𝑑
𝑢= +1 𝑦= +𝑦
𝑑𝑥 𝑑𝑥
88
Thus we can now view our equation as equivalent to the two equations
𝑑𝑦
+𝑦 = 𝑢
𝑑𝑥
𝑑𝑢
− 𝑢 = 𝑥𝑒𝑥 .
𝑑𝑥
both of which are first order and linear. Solving the second one first, we can see that the integrating
factor is 𝑒−𝑥 so it can be written as
𝑑𝑢 −𝑥
𝑒 − 𝑢𝑒−𝑥 = 𝑥
𝑑𝑥
𝑑(𝑢𝑒−𝑥 )
⇒ = 𝑥
𝑑𝑥
𝑥2
⇒ 𝑢𝑒−𝑥 = +𝐶
2
𝑥2 𝑥
⇒𝑢 = 𝑒 + 𝐶𝑒𝑥 .
2
Having solved for 𝑢 we can substitute this result into the other first order equation so that
𝑑𝑦 𝑥2
+ 𝑦 = 𝑢 = 𝑒𝑥 + 𝐶𝑒𝑥 .
𝑑𝑥 2
𝑑𝑦 𝑥 𝑥2 2𝑥
𝑒 + 𝑦𝑒𝑥 = 𝑒 + 𝐶𝑒2𝑥
𝑑𝑥 2
𝑑(𝑦𝑒𝑥 ) 𝑥2 2𝑥
⇒ = 𝑒 + 𝐶𝑒2𝑥
𝑑𝑥 2 ( )
𝑥 𝑥2 2𝑥 2𝑥
⇒ 𝑦𝑒 = 𝑒 + 𝐶𝑒 𝑑𝑥
∫ 2
𝑥2 2𝑥 𝑥𝑒2𝑥 𝐶
= 𝑒 − 𝑑𝑥 + 𝑒2𝑥 + 𝐷
4 ∫ 2 2
𝑥2 2𝑥 𝑥𝑒2𝑥 𝑒2𝑥 𝐶 2𝑥
= 𝑒 − + + 𝑒 +𝐷
4 4 8 2
2 𝑥 𝑥
𝑥 𝑥 𝑥𝑒 𝑒 𝐶
⇒𝑦 = 𝑒 − + + 𝑒𝑥 + 𝐷𝑒−𝑥
4 4 8 2
𝑥2 𝑥 𝑥 𝑥
⇒𝑦 = 𝑒 − 𝑒 + 𝐸𝑒𝑥 + 𝐷𝑒−𝑥
4 4
The first two terms are what we would have called 𝑦𝑝 beforehand and the last two terms are 𝑦𝐻 .
Note that 𝑦𝑝 has been generated by this method without any recourse to guessing.
89
3.2.5 Simultaneous First Order Equations
In fact one way to solve such equations is very similar to the way we solve (non-differential) simultaneous
equations. One can solve such equations by using one equation to solve for one variable in terms of the
other and then substituting into the remaining equation. So for example if we have
3𝑦 + 𝑧 = 5
5𝑦 − 2𝑧 = 1
then solving the first equation for 𝑧 we have 𝑧 = 5 − 3𝑦 and substituting this into the second equation
gives
5𝑦 − 2𝑧 = 1
⇒ 5𝑦 − 2(5 − 3𝑦) = 1
⇒ 11𝑦 − 10 = 1
⇒ 𝑦 = 1, 𝑧 = 5 − 3.1 = 2.
We shall follow the same approach and illustrate the method with an example
The following equations are meant to model the populations of lions (𝐿) and deer (𝐷) on a small
island:
𝑑𝐷
= 𝐷 − 8𝐿
𝑑𝑡
𝑑𝐿
= (𝐷 − 300) − 5𝐿
𝑑𝑡
90
We solve the second equation for 𝐷,
𝑑𝐿
𝐷= + 5𝐿 + 300
𝑑𝑡
We have succeeded in turning our two coupled first order equations into a single second order
inhomogeneous equation for 𝐿 which we know how to solve. As usual we have that the general
solution is the sum of 𝐿𝐻 which solves the homogeneous equation and a particular solution 𝐿𝑝 .
The auxiliary equation is
𝜆2 + 4𝜆 + 3 = 0 ⇒ (𝜆 + 3)(𝜆 + 1) = 0
so 𝜆 = −3, −1 and 𝐿𝐻 = 𝛼𝑒−3𝑡 + 𝛽𝑒−𝑡 . As the right hand side is a constant (0-th order polynomial)
we should try 𝐿𝑝 = 𝑏0 , a constant. Substituting in gives 3𝑏0 = 300 so 𝐿𝑝 = 𝑏0 = 100. Therefore
we find
An interesting feature of this solution is that as 𝑡 → ∞, 𝐿 → 100 and 𝐷 → 800. Therefore these
are the stable populations of Lions and Deer that the island can sustain!
91
3.3 Physical applications of second order differential equations
Second order differential equations play an important role in many physical situations. As an example we
will consider a car suspension system. Below is a very simplified model of the rear wheel of a car. The
SPRING
SHOCK ABSORBER
HINGE
spring provides the force counteracting gravity which stops the wheel of the car being forced up into the
arch; the shock absorber takes excess energy out of the vertical motion of the car to stop it bouncing up
and down.
In this case the above picture can be simplified even further. Essentially the car body is like a
mass 𝑚 suspended on top of a spring. If 𝑥 measures the vertical displacement of the car from its
Equilibrium position
x Car Body mass m
Spring
Ground
equilibrium position than the extra force from the spring will be 𝜆𝑥 where 𝜆 is a negative constant,
taking into account that the spring force always acts in the opposite direction to the displacement.
Taking 𝜆 = −𝑚𝑘2 then Newton’s second law gives
𝑑 2𝑥
𝑚 2
= −𝑚𝑘2 𝑥 ⇒ 𝑥̈ + 𝑘2 𝑥 = 0.
𝑑𝑡
This important equation is called the simple harmonic oscillator (SHO). Since it is a second-order
linear we can solve it using techniques developed in the previous chapter. The auxiliary equation is
𝜆 + 𝑘2 = 0 which gives that 𝜆 = ±𝑖𝑘 so the general solution to the SHO is
92
Suppose initially the car is moving horizontally but in vertical equilibrium so that 𝑥 = 0, when it
hits a pot-hole which has the effect of jolting the car upwards with velocity 𝑣0 . Mathematically we
have 𝑥(0) = 0, and 𝑥(𝑡)
̇ = 𝑣0 . In this case we have
𝑣0
𝑥(0)
̇ = 𝐵𝑘 cos(0) = 𝐵𝑘 ⇒ 𝐵 =
𝑘
so
𝑣0
𝑥(𝑡) = sin(𝑘𝑡).
𝑘
𝑣0
A graph of 𝑥 against 𝑡 looks like which means the car oscillates endlessly with amplitude 𝑘
. Pre-
x(t)
sumably this is not very comfortable for the passengers, and why we will need a shock absorber to
take the energy out of the system.
The shock absorber in contrast to the spring does not produce a force based on its shape but rather
on the rate at which we are trying to change its shape. It produces a force proportional to 𝑥̇ so we
have
93
so now Newton’s second law reads
𝑑 2𝑥
𝑚 = 𝐹spring + 𝐹shock absorber
𝑑𝑡2
𝑑𝑥
= −𝑚𝑘2 𝑥 − 2𝑚𝜇
𝑑𝑡
𝑑 2𝑥 𝑑𝑥
⇒ + 2𝜇 + 𝑘2 𝑥 = 0
𝑑𝑡2 𝑑𝑡
Again this is an important equation, the damped simple harmonic oscillator, which is a second-
order linear ODE which we can solve. The auxiliary equation is 𝜆2 + 2𝜇𝜆 + 𝑘2 = 0 which has
solutions √
−2𝜇 ± (2𝜇)2 − 4𝑘2 √
𝜆± = = −𝜇 ± 𝜇 2 − 𝑘2 .
2
The nature of the solutions will depend on whether the argument of the square root is positive,
negative or zero, and we shall deal with each of these cases in turn.
First we consider the case where the shock absorber does not exert much force; this corresponds to
small 𝜇, or more precisely 𝜇 < 𝑘. If we define 𝜔 > 0 by 𝜇2 − 𝑘2 = −𝜔2 , then the solutions to the
auxiliary equation are
√
𝜆± == −𝜇 ± 𝜇 2 − 𝑘2 = −𝜇 ± 𝑖𝜔
⇒ 𝑥 = 𝐵𝑒−𝜇𝑡 sin(𝜔𝑡)
⇒ 𝑥(0)
̇ = −𝜇𝐵𝑒0 sin(0) + 𝜔𝐵𝑒0 cos(0) = 𝜔𝐵 = 𝑣0
𝑣
⇒𝐵 = 0
𝜔
𝑣0 −𝜇𝑡
⇒𝑥 = 𝑒 sin(𝜔𝑡).
𝜔
Below we show the graph of 𝑥. The sin(𝜔𝑡) factor ensures that it remains an oscillating function, but
now with an exponentially decaying amplitude thanks to the factor 𝑒−𝜇𝑡 . For a passenger in the car,
94
x(t)
this is probably a better experience to the car without a shock absorber, as the effects of the pot-hole
die away exponentially with time, but still they may be subject to several vertical oscillations before
things settle down. Perhaps because of this, this case is called the underdampled simple harmonic
oscillator
• Model 2b: Strong shock absorber: overdamping Let us now go to the opposite extreme and bolt
a very ‘strong’ shock absorber to our car. Physically, this means that the shock absorber strongly
resists any motion, more mathematically it means the parameter 𝜇 is large. More precisely we shall
assume now that 𝜇 > 𝑘 so that we can write 𝜇 2 − 𝑘2 = 𝜅 2 , where 0 < 𝜅 < 𝜇. From this is follows
that the solutions to the auxiliary equation are
√
𝜆± == −𝜇 ± 𝜇 2 − 𝑘2 = −𝜇 ± 𝜅
Note that since 0 < 𝜅 < 𝜇, both solutions to they auxiliary equation 𝜆± = −𝜇 ± 𝜅 are negative; thus
𝑥(𝑡) decays exponentially with time. Sticking with our boundary conditions 𝑥(0) = 0 and 𝑥(0)
̇ = 𝑣0
95
we have that
𝑥(0) = 𝐴𝑒0 + 𝐵 0 = 𝐴 + 𝐵 = 0
⇒ 𝐵 = −𝐴
( )
⇒ 𝑥(𝑡) = 𝐵𝑒−𝜇𝑡 𝑒𝜅𝑡 − 𝑒−𝜅𝑡 = 2𝐵𝑒−𝜇𝑡 sinh(𝜅𝑡)
⇒ 𝑥(𝑡)
̇ = −2𝜇𝐵𝑒−𝜇𝑡 sinh(𝜅𝑡) + 2𝜅𝐵𝑒−𝜇𝑡 cosh(𝜅𝑡)
⇒ 𝑥(0)
̇ = −2𝜇𝐵𝑒0 sinh(0) + 2𝜅𝐵𝑒0 cosh(0) = 2𝜅𝐵 = 𝑣0
𝑣0
⇒𝐵 =
2𝜅
𝑣0 −𝜇𝑡
⇒ 𝑥(𝑡) = 𝑒 sinh(𝜅𝑡).
𝜅
We show the graph of 𝑥(𝑡) below. The solution no longer oscillates and simply exponentially de-
x(t)
cays. Perhaps the most worrying feature is that the graph of 𝑥(𝑡) can have very sharp bends in it,
corresponding to large accelerations and therefore large jerky forces felt by the passengers. In the
extreme case that 𝜇 → ∞,s effectively the suspension cannot move and the wheel is bolted directly
to the chassis, with the result that the passengers feel every jolt from the road! As such, this case is
called the overdamped simple harmonic oscillator.
Having dealt with the cases where 𝜇 < 𝑘 and 𝜇 > 𝑘, it only remains to deal with the case where
𝜇 = 𝑘. In this case the auxiliary equation has the repeated root 𝜆± = −𝜇, and so we have the general
solution
𝑥(𝑡) = (𝐴𝑡 + 𝐵)𝑒−𝜇𝑡
𝑥(𝑡) = 𝑣0 𝑡𝑒−𝜇𝑡 .
96
The graph of 𝑥(𝑡) looks very much like the previous case, except perhaps a little smoother. This
x(t)
case, called the critically damped simple harmonic oscillator, represents the boundary between
oscillating solutions and non-oscillating solutions. If we were to make the shock absorber even a
little weaker, the car would bounce up and down more than once.
3.3.2 Resonance
In this section we shall ask what happens when our simple harmonic oscillator is subject to an ‘external’
oscillating force, and discover that something very interesting happens when the frequency of the external
force matches the natural frequency of the SHO.
For simplicity we will only deal with the undamped oscillator. For example a mass 𝑚 moving in a line is
attached to a spring. The distance from its equilibrium position is given by 𝑥 and the force exerted by the
spring is −𝑚𝑘2 𝑥. Newton’s Second Law reads
x
𝑑 2𝑥 2 𝑑 2𝑥
𝑚𝑎 = 𝑚 = 𝐹 = −𝑚𝑘 𝑥 ⇒ + 𝑘2 𝑥 = 0.
𝑑𝑡2 𝑑𝑡2
In addition the mass is subject to an external oscillating force 𝐹ext = 𝑚𝑋 sin(𝜈𝑡) with (angular) frequency
𝜈, then the above equation is modified so that
𝑑 2𝑥
+ 𝑘2 𝑥 = 𝑋 sin(𝜈𝑡).
𝑑𝑡2
97
This is an inhomogeneous second-order equation, and the right-hand side is of a form we have dealt
with (Case 2). The general solution to this equation can be written 𝑥 = 𝑥𝐻 + 𝑥𝑝 where 𝑥𝐻 solves the
homogeneous equation and 𝑥𝑝 is a particular solution. The auxiliary equation is 𝜆2 + 𝑘2 = 0 so 𝜆± = ±𝑖𝑘
and we recover the usual SHO solution 𝑥𝐻 = 𝐴 cos(𝑘𝑡) + 𝐵 sin(𝑘𝑡). We take 𝑥𝑝 = 𝑟 cos(𝜈𝑡) + 𝑠 sin(𝜈𝑡).
Substituting in we find that
𝑑 2 𝑥𝑝
+ 𝑘2 𝑥𝑝 = 𝑋 sin(𝜈𝑡)
𝑑𝑡2
( )
⇒ −𝑟𝜈 2 cos(𝜈𝑡) − 𝑠𝜈 2 sin(𝜈𝑡) + 𝑘2 (𝑟 cos(𝜈𝑡) + 𝑠 sin(𝜈𝑡)) = 𝑋 sin(𝜈𝑡)
The general solution is a sum of oscillation with (angular) frequency 𝑘, the natural frequency of the
SHO, and frequency 𝜈, the (angular) frequency of the external force. Let us assume that initially both the
displacement and velocity of the particle vanish, i.e. 𝑥(0) = 𝑥(0)
̇ = 0. In this case we find that
( )
𝑋 𝜈
𝑥(𝑡) = − sin(𝑘𝑡) + sin(𝜈𝑡) .
𝑘2 − 𝜈 2 𝑘
Now let us assume that the two frequencies 𝑘 and 𝜈 are almost equal. We will define constants 𝜔 and 𝜖
such that 𝑘 = 𝜔 + 𝜖 and 𝜈 = 𝜔 − 𝜖, where 𝜖 ≪ 1. We have that
𝑋
𝑥(𝑡) ≈ (− sin(𝑘𝑡) + sin(𝜈𝑡))
4𝜔𝜖
𝑋
= (− sin ((𝜔 + 𝜖)𝑡) + sin ((𝜔 − 𝜖)𝑡))
4𝜔𝜖
𝑋
= (−2 cos(𝜔𝑡) sin(𝜖𝑡))
4𝜔𝜖
𝑋
= − cos(𝜔𝑡) sin(𝜖𝑡).
2𝜔𝜖
98
2. the term cos(𝜔𝑡) is an oscillation with (angular) frequency 𝜔 ≈ 𝑘 ≈ 𝜈, i.e. at approximately the
frequency of the external force and the SHO.
3. the term sin(𝜖𝑡) has very very low frequency as 𝜖 ≪ 1, and represents a very slow change in the
amplitude of the oscillation.
Plotting the displacement 𝑥(𝑡) as a function of time gives the following: The phenomenon of a slowly
x(t)
π
X
2ω
changing amplitude when two signals with very close frequencies (in this case 𝑘 and 𝜈) are added together
is called beats.
We have considered what happens when the two frequencies 𝑘 and 𝜈 are close to one another. What
happens when they coincide. Unfortunately the solution we found appears to blow up in the limit that
𝜖 → 0 when the two frequencies become identical. It is easier in this case to start again. If 𝜈 = 𝑘 then the
ODE becomes
𝑑 2𝑥
+ 𝑘2 𝑥 = 𝑋 sin(𝑘𝑡).
𝑑𝑡2
It is still true that the general solution can be written in the form 𝑥 = 𝑥𝐻 + 𝑥𝑝 and that 𝑥𝐻 = 𝐴 cos(𝑘𝑡) +
𝐵 sin(𝑘𝑡). However in this case taking 𝑥𝑝 = 𝑟 cos(𝜈𝑡) + 𝑠 sin(𝜈𝑡) as a guess for the particular integral is a
bad idea as this is identical in form to our homogenous solution 𝑥𝐻 , and we know that since this solves the
homogeneous equation substituting it into the left hand side of the ODE will give zero not 𝑋 sin(𝑘𝑡). This
is therefore the ‘exception to the rule’ or case 2a, so we must instead try taking 𝑥𝑝 = 𝑟𝑡 cos(𝜈𝑡) + 𝑠𝑡 sin(𝜈𝑡).
Substituting this into our equation we find that a particular solution to the equation is given by
𝑋
𝑥𝑝 = − 𝑡 cos(𝑘𝑡)
2𝑘
99
so that the general solution is
𝑋
𝑥(𝑡) = 𝐴 cos(𝑘𝑡) + 𝐵 sin(𝑘𝑡) − 𝑡 cos(𝑘𝑡)
2𝑘
𝑋 𝑋
𝑥(𝑡) = sin(𝑘𝑡) − 𝑡 cos(𝑘𝑡).
2𝑘2 2𝑘
The first term has magnitude at most 𝑋∕(2𝑘2 ) whilst the second term has a magnitude which grows
linearly with time so this term dominates for large times. Plotting the function 𝑥(𝑡) we find the following:
The amplitude of the system keeps on growing and growing as the SHO endlessly absorbs energy from
x(t)
the external source. In this situation either the system or its mathematical description will most likely
break down at some stage. This phenomenon is called resonance.
100
Chapter 4
Fourier Series/analysis
Fourier analysis is concerned with the study of periodic functions. Periodic functions are functions which
repeat themselves; a graph of a typical periodic function is shown below. Mathematically we say that the
f (x)
Note that if 𝑓 (𝑥) is periodic with period 𝑎 then from the definition we have that
and by the same sort of proof it is easy to show that 𝑓 (𝑥 + 𝑛𝑎) = 𝑓 (𝑥) for any integer 𝑛.
The basic idea of the Fourier series is that any (with some technical caveats) periodic function can be
represented as a sum of cos and sin functions in the form:
∑
∞
∑
∞
𝑓 (𝑥) = 𝑎0 .1 + 𝑎𝑛 cos(𝑛𝑥) + 𝑏𝑛 sin(𝑛𝑥).
𝑛=1 𝑛=1
101
The aim of the game is to find the right amount of 𝑎𝑛 and 𝑏𝑛 :
Example
The function 𝑓 (𝑥) is defined to be periodic with period 2𝜋 so that 𝑓 (𝑥 + 2𝜋) = 𝑓 (𝑥). It is defined on
These two facts define the function 𝑓 (𝑥) for all 𝑥, and the graph of the function is shown below:
f (x)
−π π x
3π
We have seen that given a periodic function 𝑓 (𝑥) we would like to write it as a Fourier series
∑
∞
∑
∞
𝑓 (𝑥) = 𝑎0 .1 + 𝑎𝑛 cos(𝑛𝑥) + 𝑏𝑛 sin(𝑛𝑥).
𝑛=1 𝑛=1
Given a function 𝑓 (𝑥) we would like to find the Fourier coefficients 𝑎𝑛 , 𝑏𝑛 . To put it another way, we
would like to determine how much of a particularly frequency 𝑛 appears in the function 𝑓 (𝑥).
To see how to work the coefficients out, we shall proceed by analogy with vectors. Suppose we have three
mutually orthogonal vectors 𝐱1 , 𝐱2 and 𝐱3 . Mathematically, the statement that the vectors are orthogonal
102
means that
𝐱1 .𝐱2 = 𝐱1 .𝐱3 = 𝐱2 .𝐱3 = 0.
Since they point in different directions to one another, they form a basis for three-dimensional vectors,
which means that we can write any vector 𝐯 as a linear combination of 𝐱1 , 𝐱2 and 𝐱3 ;
∑
3
𝐯= 𝑎𝑖 𝐱𝑖 = 𝑎1 𝐱1 + 𝑎2 𝐱2 + 𝑎3 𝐱3 ,
𝑖=1
for some constants 𝑎1 , 𝑎2 , 𝑎3 . Now the fact that the vectors 𝐱1 , 𝐱2 and 𝐱3 are mutually orthogonal makes
finding the coefficients 𝑎𝑖 for a given vector 𝐯 very simple. Suppose we want to find 𝑎1 . Taking the dot
product of the above equation with 𝐱1 we see
( )
𝐯.𝐱1 = 𝑎1 𝐱1 + 𝑎2 𝐱2 + 𝑎3 𝐱3 .𝐱𝟏
= 𝑎1 𝐱1 .𝐱1 + 𝑎2 .0 + 𝑎3 .0
𝐯.𝐱1
= 𝑎1 𝐱1 .𝐱1 ⇒ 𝑎1 =
𝐱1 .𝐱1
We have used the mutual orthogonality of 𝐱1 , 𝐱2 and 𝐱3 in the third line above; cunningly taking the dot
product has removed both the terms involving 𝑎2 and 𝑎3 , leaving us with only the term involving 𝑎1 which
we can rearrange to solve for 𝑎1 . If we had wanted to calculate 𝑎2 instead, we would begin by dotting the
equation with 𝐱𝟐 as this would have similarly isolated the 𝑎2 term and we would instead find that
𝐯.𝐱2
𝑎2 = .
𝐱2 .𝐱2
If this is a bit abstract so let us see how it works in a particular example:
𝐱1 = 𝐢 + 𝐣
𝐱2 = 𝐢 − 𝐣
𝐱3 = 2𝐤
∑
3
𝐯= 𝑎𝑖 𝐱𝑖 = 𝑎1 𝐱1 + 𝑎2 𝐱2 + 𝑎3 𝐱3 ,
𝑖=1
103
One way to solve this problem would be to substitute in the explicit expressions for 𝐯, 𝐱1 , 𝐱2 and 𝐱3
into the above equation and equate coefficients of 𝐢, 𝐣 and 𝐤, which would give three simultaneous
equations for 𝑎1 , 𝑎2 and 𝑎3 . But instead we shall use the formulae above which give
𝐯.𝐱1 (6, −4, 6).(1, 1, 0) 2
𝑎1 = = = =1
𝐱1 .𝐱1 (1, 1, 0).(1, 1, 0) 2
𝐯.𝐱2 (6, −4, 6).(1, −1, 0) 10
𝑎2 = = = =5
𝐱2 .𝐱2 (1, −1, 0).(1, −1, 0) 2
𝐯.𝐱3 (6, −4, 6).(0, 0, 2) 12
𝑎3 = = = =3
𝐱3 .𝐱3 (0, 0, 2).(0, 0, 2) 4
and one can check that
𝐯 = 𝐱1 + 5𝐱2 + 3𝐱3
Although not obvious at first sight, the problem of finding Fourier coefficients 𝑎𝑛 , 𝑏𝑛 is actually very
similar to the problem of finding 𝑎1 , 𝑎2 and 𝑎3 in the above vector problem. Colour coding the equations,
we have
∑
3
𝐯 = 𝑎𝑖 𝐱𝑖 = 𝑎1 𝐱1 + 𝑎2 𝐱2 + 𝑎3 𝐱3
𝑖=1
∑
∞
∑
∞
𝑓 (𝑥) = 𝑎0 .1 + 𝑎𝑛 cos(𝑛𝑥) + 𝑏𝑛 sin(𝑛𝑥).
𝑛=1 𝑛=1
In both equations we are trying to express a general vector/periodic function (in green) in terms of some
linear combination of some chosen set of basis vectors/periodic functions (in blue) with some numerical
coefficients (in red). To reinforce the similarity consider the following table:
The first three rows of the table we have already discussed - the last two rows are for the moment myste-
rious!
104
The inner product
In order to find e.g. the coefficient 𝑎1 in our vector example, the dot product played a vital role in isolating
the term containing 𝑎1 from all the other terms on the right had side of the equation. We need the analogous
quantity in the case of for periodic functions; it is called the inner product.
We define the inner product between two periodic functions 𝑓 (𝑥), 𝑔(𝑥) (of period 2𝜋) as follows:
𝜋
(𝑓 (𝑥), 𝑔(𝑥)) = 𝑓 (𝑥)𝑔(𝑥)𝑑𝑥.
∫−𝜋
• Example From the definition, calculate the inner product of the periodic function 𝑓 (𝑥) = 𝑥 for
−𝜋 < 𝑥 ≤ 𝜋 and the function 𝑔(𝑥) = sin(𝑥). The definition tells us that
𝜋
(𝑓 (𝑥), 𝑔(𝑥)) = 𝑓 (𝑥)𝑔(𝑥)𝑑𝑥
∫−𝜋
𝜋
= 𝑥 sin(𝑥)𝑑𝑥
∫−𝜋
𝜋
= [−𝑥 cos(𝑥)]𝜋−𝜋 + cos(𝑥)𝑑𝑥
∫−𝜋
= 2𝜋 + [sin(𝑥)]𝜋−𝜋 = 2𝜋.
From its definition we can deduce that the inner product on periodic functions shares a large number of
properties in common with the dot product on vectors:
• The dot product 𝐚.𝐛 is a function of two vectors which gives you back a number. Similarly the inner
product is a function of two periodic functions 𝑓 (𝑥), 𝑔(𝑥) which returns a number.
• Both the dot product and inner product are symmetric, that is 𝐚.𝐛 = 𝐛.𝐚. Similarly
𝜋 𝜋
(𝑓 (𝑥), 𝑔(𝑥)) = 𝑓 (𝑥)𝑔(𝑥)𝑑𝑥 = 𝑔(𝑥)𝑓 (𝑥)𝑑𝑥 = (𝑔(𝑥), 𝑓 (𝑥)).
∫−𝜋 ∫−𝜋
• Both products are linear, that is to say (𝛼𝐚 + 𝛽𝐜).𝐛 = 𝛼𝐚.𝐛 + 𝛽𝐜.𝐛, and similarly for three periodic
function 𝑓 (𝑥),𝑔(𝑥) and ℎ(𝑥) we have
𝜋
(𝛼𝑓 (𝑥) + 𝛽ℎ(𝑥), 𝑔(𝑥)) = (𝛼𝑓 (𝑥) + 𝛽ℎ(𝑥))𝑔(𝑥)𝑑𝑥
∫−𝜋
𝜋 𝜋
= 𝛼 𝑓 (𝑥)𝑔(𝑥)𝑑𝑥 + 𝛽
ℎ(𝑥)𝑔(𝑥)𝑑𝑥
∫−𝜋 ∫−𝜋
= 𝛼(𝑓 (𝑥), 𝑔(𝑥)) + 𝛽(ℎ(𝑥), 𝑔(𝑥)).
105
• The length or norm squared of a vector 𝐚 is written using the dot product as |𝐚|2 = 𝐚.𝐚 ≥ 0.
Similarly from our inner product we can define the norm squared of a periodic function as
𝜋 𝜋
||𝑓 (𝑥)||2 = (𝑓 (𝑥), 𝑓 (𝑥)) = 𝑓 (𝑥)𝑓 (𝑥)𝑑𝑥 = (𝑓 (𝑥))2 𝑑𝑥 ≥ 0
∫−𝜋 ∫−𝜋
where the last inequality follows because the integrand (𝑓 (𝑥))2 is non-negative.
• Finally, and of great importance to us in the current context, two vectors 𝐚 and 𝐛 are orthogonal to
one another if 𝐚.𝐛 = 0. We define two periodic functions 𝑓 (𝑥) and 𝑔(𝑥) to be orthogonal to one
another if their inner product (𝑓 (𝑥), 𝑔(𝑥)) = 0.
• Example Two periodic functions are given as 𝑓 (𝑥) = 𝑥 for −𝜋 < 𝑥 ≤ 𝜋 and 𝑔(𝑥) = 1. Calculate
the norm squared of 𝑓 (𝑥) and 𝑔(𝑥) and show that they are orthogonal to one another.
𝜋 𝜋 [ 3 ]𝜋
𝑥 2𝜋 3
||𝑓 (𝑥)|| =
2 2
𝑓 (𝑥) 𝑑𝑥 = 2
𝑥 𝑑𝑥 = =
∫−𝜋 ∫−𝜋 3 −𝜋 3
𝜋 𝜋
||𝑔(𝑥)||2 = 12 𝑑𝑥 = [𝑥]𝜋−𝜋 = 2𝜋
𝑓 (𝑥)2 𝑑𝑥 =
∫−𝜋 ∫−𝜋
𝜋 [ 2 ]𝜋
𝑥
(𝑓 (𝑥), 𝑔(𝑥)) = (𝑥, 1) = 𝑥.1𝑑𝑥 = = 0.
∫−𝜋 2 −𝜋
Since the inner product of 𝑓 (𝑥) and 𝑔(𝑥) vanishes, by definition they are orthogonal functions.
Now we come across a remarkable fact. The blue ‘basis’ periodic functions in the set {1, cos(𝑛𝑥), sin(𝑛𝑥)}
are all mutually orthogonal to one another; that is to say if I pick any two different functions from this
set then their inner product vanishes. So for example (sin(3𝑥), cos(7𝑥)) = 0, or (1, sin(5𝑥)) = 0. We will
not prove this for every case here. Let us just prove that the sines are all orthogonal to one another, i.e.
(sin(𝑚𝑥), sin(𝑛𝑥)) = 0 for 𝑚, 𝑛 > 0 and 𝑚 ≠ 𝑛.
𝜋
(sin(𝑚𝑥), sin(𝑛𝑥)) = sin(𝑚𝑥) sin(𝑛𝑥)𝑑𝑥
∫−𝜋
𝜋
1
= (cos((𝑚 − 𝑛)𝑥) − cos((𝑚 + 𝑛)𝑥)) 𝑑𝑥
∫−𝜋 2
[ ( )]𝜋
1 sin((𝑚 − 𝑛)𝑥) sin((𝑚 + 𝑛)𝑥)
= −
2 𝑚−𝑛 𝑚+𝑛 −𝜋
= 0,
106
where we have used a trigonometric relation on the second line and that sin(𝑀𝜋) = 0 for any integer
multiple 𝑀 of 𝜋. Notice that this proof breaks down when 𝑚 = 𝑛 as the denominator 𝑚 − 𝑛 vanishes in
this case. In that case 𝑚 = 𝑛 we instead write
𝜋
|| sin(𝑚𝑥)||2 = (sin(𝑚𝑥), sin(𝑚𝑥)) = sin2 (𝑚𝑥)𝑑𝑥
∫−𝜋
𝜋
1
= (1 − cos(2𝑚𝑥)) 𝑑𝑥
∫−𝜋 2
[ ( )]𝜋
1 sin(2𝑚𝑥)
= 𝑥−
2 2𝑚 −𝜋
= 𝜋.
After this somewhat lengthy diversion, we now have all that we need to find Fourier coefficients. Recall
that in the vector analogy to find the coefficient 𝑎1 of 𝐱𝟏 , we dotted the equation for 𝐯 with 𝐱𝟏 ; that is
( ) 𝐯.𝐱1
𝐯.𝐱1 = 𝑎1 𝐱1 + 𝑎2 𝐱2 + 𝑎3 𝐱3 .𝐱𝟏 = 𝑎1 𝐱1 .𝐱1 ⇒ 𝑎1 = .
𝐱1 .𝐱1
Suppose we wish to calculate the coefficient 𝑏7 of sin(7𝑥) in the Fourier series
∑
∞
∑
∞
𝑓 (𝑥) = 𝑎0 .1 + 𝑎𝑛 cos(𝑛𝑥) + 𝑏𝑛 sin(𝑛𝑥).
𝑛=1 𝑛=1
Proceeding by analogy we should take the inner product of this equation with the thing that 𝑏7 is the
coefficient of, i.e. sin(7𝑥). This gives
( )
∑
∞
∑
∞
(𝑓 (𝑥), sin(7𝑥)) = 𝑎0 .1 + 𝑎𝑛 cos(𝑛𝑥) + 𝑏𝑛 sin(𝑛𝑥), sin(7𝑥)
𝑛=1 𝑛=1
∑
∞
∑
∞
= 𝑎0 (1, sin(7𝑥)) + 𝑎𝑛 (cos(𝑛𝑥), sin(7𝑥)) + + 𝑏𝑛 (sin(𝑛𝑥), sin(7𝑥))
𝑛=1 𝑛=1
∑
∞
∑∞
= 𝑎0 .0 + 𝑎𝑛 0 + 𝑏𝑛 (sin(𝑛𝑥), sin(7𝑥))
𝑛=1 𝑛=1
107
where we have used the orthogonality of 1 with sin(7𝑥) and cos(𝑛𝑥) with sin(7𝑥) to set the first two terms
to zero. But additionally (sin(𝑛𝑥), sin(7𝑥)) = 0 for 𝑛 ≠ 7, so the only term that is not zero in the infinite
sum over 𝑛 is the 𝑛 = 7 term so we have
where in the last line we have used the definition of the inner product. Using a similar technique we can
calculate all of the Fourier [Link] result, known as Euler’s formula is given below
𝜋
1
𝑎0 = 𝑓 (𝑥)𝑑𝑥 (4.0.1)
2𝜋 ∫−𝜋
𝜋
1
𝑎𝑛 = 𝑓 (𝑥) cos(𝑛𝑥)𝑑𝑥 (4.0.2)
𝜋 ∫−𝜋
𝜋
1
𝑏𝑛 = 𝑓 (𝑥) sin(𝑛𝑥)𝑑𝑥. (4.0.3)
𝜋 ∫−𝜋
∑
∞
∑
∞
𝑓 (𝑥) = 𝑎0 + 𝑎𝑛 cos(𝑛𝑥) + 𝑏𝑛 sin(𝑛𝑥).
𝑛=1 𝑛=1
108
Using Euler’s formula we have
𝜋 [ ]𝜋
1 1 𝑥2
𝑎0 = 𝑥𝑑𝑥 = =0
2𝜋 ∫−𝜋 2𝜋 2 −𝜋
𝜋
1
𝑎𝑛 = 𝑥 cos(𝑛𝑥)𝑑𝑥
𝜋 ∫−𝜋
[ ]𝜋 𝜋
1 sin(𝑛𝑥) 1 sin(𝑛𝑥)
= 𝑥 − 𝑑𝑥
𝜋 𝑛 −𝜋
∫
𝜋 −𝜋 𝑛
1
= 0 + 2 [cos(𝑛𝑥)]𝜋−𝜋
𝜋𝑛
= 0
𝜋
1
𝑏𝑛 = 𝑥 sin(𝑛𝑥)𝑑𝑥
𝜋 ∫−𝜋
[ ]𝜋 𝜋
1 cos(𝑛𝑥) 1 cos(𝑛𝑥)
= − 𝑥 + 𝑑𝑥
𝜋 𝑛 −𝜋 𝜋 ∫−𝜋 𝑛
2 1
= (−1)𝑛+1 + 2 [sin(𝑛𝑥)]𝜋−𝜋
𝑛 𝜋𝑛
2
= (−1)𝑛+1
𝑛
Substituting these values into the Fourier series we get
∑
∞
2(−1)𝑛+1
𝑓 (𝑥) = sin(𝑛𝑥)
𝑛=1
𝑛
2 1 2
= 2 sin(𝑥) − sin(2𝑥) + sin(3𝑥) − sin(4𝑥) + sin(5𝑥) + … ,
3 2 5
as we stated earlier. Note that in calculating these integrals we have used two results about trigono-
metrical functions which are terribly useful when calculating Fourier coefficients:
The calculation of Fourier coefficients can often be greatly simplified if the function 𝑓 (𝑥) we are analysing
is either an even or an odd function. As a reminder
⎧ ⎫ ⎧
⎪ even ⎪ ⎪ 𝑓 (𝑥) = 𝑓 (−𝑥)
a function 𝑓 (𝑥) is ⎨ ⎬ if ⎨
⎪ odd ⎪ ⎪ 𝑓 (𝑥) = −𝑓 (−𝑥).
⎩ ⎭ ⎩
109
Geometrically, the graph of a function which is even is symmetric about the ‘y-axis’, whilst the graph
of an odd function is symmetric if rotated through 180◦ about the origin. As it happens, all the ‘basis’
f (x) f (x)
x x
functions which we use in a Fourier series {1, sin(𝑛𝑥), cos(𝑛𝑥)} are either even or odd;
Let us have another look at the last example in the previous section. There we calculated the Fourier
coefficients for the periodic function defined to be 𝑓 (𝑥) = 𝑥 for −𝜋 < 𝑥 ≤ 𝜋. From the graph of this
function (below) it is clear that the function is ODD. When we calculated the Fourier coefficients we
f (x)
−π π x
3π
found that the 𝑎𝑛 vanished for all 𝑛 (including 𝑛 = 0), meaning that the Fourier series for 𝑓 (𝑥) did not
contain any cos(𝑛𝑥) or 1 pieces, but instead 𝑓 (𝑥) could be expanded solely in terms of the remaining
sin(𝑛𝑥) terms:
∑
∞
2(−1)𝑛+1
𝑓 (𝑥) = sin(𝑛𝑥)
𝑛=1
𝑛
2 1 2
= 2 sin(𝑥) − sin(2𝑥) + sin(3𝑥) − sin(4𝑥) + sin(5𝑥) + … .
3 2 5
In retrospective, this is obvious. Since 𝑓 (𝑥) is an odd function, it makes sense that it should have a
Fourier series in terms of only the odd sin(𝑛𝑥) functions rather than the even functions cos(𝑛𝑥). This
remark saves us having to work out the coefficients 𝑎𝑛 at all; we could have immediately set them to zero!
110
Correspondingly, if 𝑓 (𝑥) were an even function, we should intuitively expect its Fourier series to be an
expansion in terms of the even functions cos(𝑛𝑥) and 1, and contain no odd pieces with sin(𝑛𝑥), so we
could immediately set 𝑏𝑛 = 0.
These results can be proved from Euler’s formula quite easily. The proof relies on a simple result about
the integrals of even and odd functions. From the figure we see that
g(x) g(x)
A A B
0
−π π π x
0 x −π
−B
𝜋 0 𝜋 𝜋
𝑔(𝑥)𝑑𝑥 = 𝑔(𝑥)𝑑𝑥 + 𝑔(𝑥)𝑑𝑥 = 2𝐴 = 2 𝑔(𝑥)𝑑𝑥
∫−𝜋 ∫−𝜋 ∫0 ∫0
if 𝑔(𝑥) is ODD.
Let 𝑓 (𝑥) be an even function, that is 𝑓 (𝑥) = 𝑓 (−𝑥). Then it follows that since
111
By similar reasoning if 𝑓 (𝑥) is odd then
𝜋 𝜋
1 1
𝑎0 = 𝑓 (𝑥)𝑑𝑥 = (ODD)𝑑𝑥 = 0
∫
2𝜋 −𝜋 2𝜋 ∫−𝜋
𝜋 𝜋
1 1
𝑎𝑛 = 𝑓 (𝑥) cos(𝑛𝑥)𝑑𝑥 = (ODD)𝑑𝑥 = 0
𝜋 ∫−𝜋 𝜋 ∫−𝜋
𝜋 𝜋 𝜋
1 1 2
𝑏𝑛 = 𝑓 (𝑥) sin(𝑛𝑥)𝑑𝑥 = (EVEN)𝑑𝑥 = 𝑓 (𝑥) sin(𝑛𝑥)𝑑𝑥.
𝜋 ∫−𝜋 𝜋 ∫−𝜋 𝜋 ∫0
Let us do a couple of more examples which show how these ideas make it much quicker to find Fourier
coefficients.
⎧
⎪ −1 for − 𝜋 < 𝑥 < 0,
𝑓 (𝑥) = ⎨
⎪ 1 for 0 ≤ 𝑥 ≤ 𝜋.
⎩
∑
∞
∑
∞
𝑓 (𝑥) = 𝑎0 + 𝑎𝑛 cos(𝑛𝑥) + 𝑏𝑛 sin(𝑛𝑥).
𝑛=1 𝑛=1
From the graph of the function it is clear that 𝑓 (𝑥) is odd. Therefore we can immediately write that
f (x)
−π 0 π 2π x
112
𝑎0 = 𝑎𝑛 = 0. The remaining coefficients can be calculated from
𝜋
2
𝑏𝑛 = 𝑓 (𝑥) sin(𝑛𝑥)𝑑𝑥
𝜋 ∫0
𝜋
2
= sin(𝑛𝑥)𝑑𝑥
𝜋 ∫0
[ ]𝜋
2 cos(𝑛𝑥)
= −
𝜋 𝑛 0
2
= (cos(0) − cos(𝑛𝜋))
𝜋𝑛
2
= (1 − (−1)𝑛 ))
𝜋𝑛
⎧
⎪ 0 for 𝑛 even
= ⎨
⎪ 𝜋𝑛 for 𝑛 odd.
4
⎩
so we have that
( )
∑ 4 4 sin(3𝑥) sin(5𝑥)
𝑓 (𝑥) = sin(𝑛𝑥) = sin(𝑥) + + +… .
𝑛 odd
𝑛𝜋 𝜋 3 5
• Example 2 A periodic function of period 2𝜋 is given by 𝑓 (𝑥) = |𝑥| for −𝜋 < 𝑥 ≤ 𝜋. Find the
Fourier coefficients in the expansion
∑
∞
∑
∞
𝑓 (𝑥) = 𝑎0 + 𝑎𝑛 cos(𝑛𝑥) + 𝑏𝑛 sin(𝑛𝑥).
𝑛=1 𝑛=1
From the graph of the function it is clear that 𝑓 (𝑥) is even. Therefore we can immediately write
f (x)
−π 0 π 2π x
113
that 𝑏𝑛 = 0. The remaining coefficients can be calculated from
𝜋
1
𝑎0 = 𝑓 (𝑥)𝑑𝑥
𝜋 ∫0
𝜋
1
= 𝑥𝑑𝑥
𝜋 ∫0
[ ]𝜋
1 𝑥2
=
𝜋 2 0
𝜋
= .
2
𝜋
2
𝑎𝑛 = 𝑓 (𝑥) cos(𝑛𝑥)𝑑𝑥
𝜋 ∫0
𝜋
2
= 𝑥 cos(𝑛𝑥)𝑑𝑥
𝜋 ∫0
[ ]𝜋 𝜋
2 sin(𝑛𝑥) 2 sin(𝑛𝑥)
= 𝑥 − 𝑑𝑥
𝜋 𝑛 0 𝜋 ∫0 𝑛
2
= 0 + 2 [cos(𝑛𝑥)]𝜋0
𝜋𝑛
2
= ((−1)𝑛 ) − 1)
𝜋𝑛2
⎧
⎪ 0 for 𝑛 even
= ⎨
⎪ − 𝜋𝑛2 for 𝑛 odd.
4
⎩
so we have that
( )
𝜋 ∑ 4 𝜋 4 cos(3𝑥) cos(5𝑥)
𝑓 (𝑥) = − 2
cos(𝑛𝑥) = − cos(𝑥) + + +… .
2 𝑛 odd 𝑛 𝜋 2 𝜋 9 25
Odd/Even extensions
For even and odd functions the formulae for the Fourier coefficients can be written in terms of integrals
between 0 and 𝜋. We have seen that for odd functions we have
𝜋
2
𝑏𝑛 = 𝑓 (𝑥) sin(𝑛𝑥)𝑑𝑥.
𝜋 ∫0
For even functions we have
𝜋 𝜋
1 1
𝑎0 = 𝑓 (𝑥)𝑑𝑥 = = 𝑓 (𝑥)𝑑𝑥
∫
2𝜋 −𝜋 𝜋 ∫0
𝜋 𝜋
1 2
𝑎𝑛 = 𝑓 (𝑥) cos(𝑛𝑥)𝑑𝑥 = 𝑓 (𝑥) cos(𝑛𝑥)𝑑𝑥.
𝜋 ∫−𝜋 𝜋 ∫0
Thus if 𝑓 (𝑥) is even or odd, the Fourier coefficients only depend on the values of 𝑓 (𝑥) for 0 ≤ 𝑥 ≤ 𝜋.
This is because in these cases we can deduce the function for all values of 𝑥 from these values. Suppose
𝑓 (𝑥) is given in the range 0 ≤ 𝑥 ≤ 𝜋 as for example shown below.
114
f (x)
x
−2π −π 0 π 2π 3π
We define the odd extension 𝑓ODD (𝑥) of 𝑓 (𝑥) to be the odd, periodic function such that 𝑓ODD (𝑥) = 𝑓 (𝑥)
for 0 ≤ 𝑥 ≤ 𝜋. This information is pretty much enough to specify 𝑓ODD (𝑥), since we can use the oddness
property to define the function for −𝜋 ≤ 𝑥 ≤ 0
f (x)
x
−2π −π 0 π 2π 3π
and then use the periodicity property to extend the function to the whole line.
f (x)
x
−2π −π 0 π 2π 3π
Similarly we can define the even extension 𝑓EVEN (𝑥) of 𝑓 (𝑥) to be the even, periodic function such that
115
𝑓EVEN (𝑥) = 𝑓 (𝑥) for 0 ≤ 𝑥 ≤ 𝜋.
116
f (x)
x
−2π −π 0 π 2π 3π
• Example
The function is defined to be 𝑓 (𝑥) = 𝑥2 for 0 ≤ 𝑥 < 𝜋. Sketch the odd and even periodic extensions,
𝑓ODD and EVEN respectively, and find their Fourier series.
Even extension
Odd extension
For 𝑓ODD , the Fourier coefficients 𝑎0 and 𝑎𝑛 vanish, and the remaining coefficients are given by
𝜋
2
𝑏𝑛 = 𝑥2 sin(𝑛𝑥)𝑑𝑥
𝜋 ∫0
2 [ 2 ]𝜋 𝜋
4
= −𝑥 cos(𝑛𝑥) 0 + 𝑥 cos(𝑛𝑥)𝑑𝑥
𝜋𝑛 𝜋𝑛 ∫0
𝜋
2𝜋 4 4
= (−1)𝑛+1 + 2 [𝑥 sin(𝑛𝑥)]𝜋0 − 2 sin(𝑛𝑥)𝑑𝑥
𝑛 𝜋𝑛 𝜋𝑛 ∫0
2𝜋 4
= (−1)𝑛+1 + 0 + 3 [cos(𝑛𝑥)]𝜋0
𝑛 𝜋𝑛
2𝜋 4
= (−1)𝑛+1 + 3 [(−1)𝑛 − 1] .
𝑛 𝜋𝑛
117
For 𝑓EVEN , the Fourier coefficients 𝑏𝑛 vanish, and the remaining coefficients are given by
𝜋
1
𝑎0 = 𝑥2 𝑑𝑥
∫
𝜋 0
[ ]𝜋
1 𝑥3 𝜋2
= = .
𝜋 3 0 3
𝜋
2
𝑎𝑛 = 𝑥2 cos(𝑛𝑥)𝑑𝑥
𝜋 ∫0
2 [ 2 ]𝜋 𝜋
4
= 𝑥 sin(𝑛𝑥) 0 − 𝑥 sin(𝑛𝑥)𝑑𝑥
𝜋𝑛 𝜋𝑛 ∫0
𝜋
4 4
= 0 + 2 [𝑥 cos(𝑛𝑥)]𝜋0 − 2 cos(𝑛𝑥)𝑑𝑥
𝜋𝑛 𝜋𝑛 ∫0
4 4
= 2
(−1)𝑛 − 3 [sin(𝑛𝑥)]𝜋0
𝑛 𝜋𝑛
4
= (−1)𝑛 .
𝑛2
So far we have restricted our attention to periodic functions of period 2𝜋, since these admit a Fourier
expansion in terms of the ‘natural’ functions sin(𝑛𝑥), cos(𝑛𝑥) with integer angular frequency.
0
−3L −2L −L L 2L 3L
−2π 0
−π π 2π 3π 4π
To deal with functions whose periodicity is different to 2𝜋 our approach will be to scale the variable 𝑥 to
118
turn our function into a function whose period is 2𝜋, borrow the results from the above sections and then
scale back.
More mathematically, suppose that 𝑓 (𝑥) is a function with periodicity 2𝐿; that is 𝑓 (𝑥 + 2𝐿) = 𝑓 (𝑥). If
we define a new function 𝑔(𝑥) = 𝑓 (𝐿𝑥∕𝜋) then we see that
( ) ( ) ( )
𝐿(𝑥 + 2𝜋) 𝐿𝑥 𝐿𝑥
𝑔(𝑥 + 2𝜋) = 𝑓 =𝑓 + 2𝐿 = 𝑓 = 𝑔(𝑥),
𝜋 𝜋 𝜋
i.e. 𝑔(𝑥) is a function with periodicity 2𝜋. Thus we can write the expansion
( ) ∑∞
∑∞
𝐿𝑥
𝑔(𝑥) = 𝑓 = 𝑎0 + 𝑎𝑛 cos(𝑛𝑥) + 𝑏𝑛 sin(𝑛𝑥).
𝜋 𝑛=1 𝑛=1
∑
∞ ( ) ∑ ∞ ( )
𝜋𝑛𝑧 𝜋𝑛𝑧
𝑓 (𝑧) = 𝑎0 + 𝑎𝑛 cos + 𝑏𝑛 sin .
𝑛=1
𝐿 𝑛=1
𝐿
Using a similar argument for the other Fourier Coefficients, we arrive at the general Euler formulae for a
periodic functions 𝑓 (𝑥) of period 2𝐿 such that 𝑓 (𝑥) = 𝑓 (𝑥 + 2𝐿) (relabelling 𝑧 as 𝑥 again!):
∑
∞ ( ) ∑ ∞ ( )
𝜋𝑛𝑥 𝜋𝑛𝑥
𝑓 (𝑥) = 𝑎0 + 𝑎𝑛 cos + 𝑏𝑛 sin .
𝑛=1
𝐿 𝑛=1
𝐿
119
where
𝐿
1
𝑎0 = 𝑓 (𝑥)𝑑𝑧
2𝐿 ∫−𝐿
𝐿 ( )
1 𝜋𝑛𝑥
𝑎𝑛 = 𝑓 (𝑥) cos 𝑑𝑥
𝐿 ∫−𝐿 𝐿
𝐿 ( )
1 𝜋𝑛𝑥
𝑏𝑛 = 𝑓 (𝑥) sin 𝑑𝑥.
𝐿 ∫−𝐿 𝐿
Note that
{ ( ) ( )}
• The function 𝑓 (𝑥) has now been expanded in terms of the ‘basis’ functions 1, cos 𝜋𝑛𝑥
𝐿
, sin 𝜋𝑛𝑥
𝐿
.
These are orthogonal functions with periodicity 2𝐿 and (angular) frequencies 𝜋𝑛∕𝐿. The larger 𝐿
is, the less often the function repeats itself and more ‘fractional’ frequencies are needed to reproduce
the function!
• The formulae for the Fourier coefficients is quite easy to remember. Simply replace all the 𝜋’s at
( )
the front of the formulae with 𝐿’s. Since e.g. 𝑎𝑛 is now the coefficient of the function cos 𝜋𝑛𝑥
( ) 𝐿
rather than cos(𝑛𝑥), then it is cos 𝐿 which appears in the formula for 𝑎𝑛 rather than cos(𝑛𝑥).
𝜋𝑛𝑥
• Example Find the Fourier Series of the function 𝑓 (𝑥) where 𝑓 (𝑥) = 𝑓 (𝑥 + 2) and
⎧
⎪ 0 for − 1 ≤ 𝑥 < 0
𝑓 (𝑥) = ⎨
⎪ 1 for 0 ≤ 𝑥 < 1
⎩
∑
∞
∑
∞
𝑓 (𝑥) = 𝑎0 + 𝑎𝑛 cos (𝜋𝑛𝑥) + 𝑏𝑛 sin (𝜋𝑛𝑥) .
𝑛=1 𝑛=1
120
The Fourier coefficients are given by
1
1
𝑎0 = 𝑓 (𝑥)𝑑𝑥
2 ∫−1
1
1 1
= 𝑑𝑥 = .
2 ∫0 2
1
𝑎𝑛 = 𝑓 (𝑥) cos(𝑛𝜋𝑥)𝑑𝑥
∫−1
1 [ ]1
sin(𝑛𝜋𝑥)
= cos(𝑛𝜋𝑥) = =0
∫0 𝑛𝜋 0
1
𝑏𝑛 = 𝑓 (𝑥) sin(𝑛𝜋𝑥)𝑑𝑥
∫−1
1 [ ]1
cos(𝑛𝜋𝑥) 1
= sin(𝑛𝜋𝑥) = − = [1 − (−1)𝑛 ] .
∫0 𝑛𝜋 0 𝑛𝜋
so we can write
1 ∑ 2
𝑓 (𝑥) = + sin(𝑛𝜋𝑥).
2 𝑛 odd 𝑛𝜋
When we analysed second order differential equations, and found complex solutions to the auxiliary equa-
( )
tion 𝜆± = 𝑟 ± 𝑖𝑠, this led naturally to complex solutions 𝑦 = 𝑒𝑟𝑥 𝛼𝑒𝑖𝑠𝑥 + 𝛽𝑒−𝑖𝑠𝑥 . In that case we saw that
this could be reëxpressed in the form
This relied on the ability to write sines and cosines in terms of complex exponentials and vice-versa, via
the relations
1 ( 𝑖𝑠𝑥 ) 1 ( 𝑖𝑠𝑥 )
cos(𝑠𝑥) = 𝑒 + 𝑒−𝑖𝑠𝑥 , sin(𝑠𝑥) = 𝑒 − 𝑒−𝑖𝑠𝑥
2 2𝑖
or conversely
𝑒𝑖𝑛𝑥 = cos(𝑛𝑥) + 𝑖 sin(𝑛𝑥) , 𝑒−𝑖𝑛𝑥 = cos(𝑛𝑥) − 𝑖 sin(𝑛𝑥).
It should therefore come as no surprise that since one can write a periodic function 𝑓 (𝑥) (with period 2𝜋)
as an expansion in terms of sines and coses,
∑
∞
∑
∞
𝑓 (𝑥) = 𝑎0 + 𝑎𝑛 cos(𝑛𝑥) + 𝑏𝑛 sin(𝑛𝑥)
𝑛=1 𝑛=1
121
This is often a much neater and more powerful way to write Fourier series. Note that we have managed
to replace the two sums over sines and coses by a single sum, although now the sum for 𝑛 runs over both
positive and negative integer values. Moreover we shall now show that the three Euler formulae for 𝑎0 ,
𝑎𝑛 and 𝑏𝑛 can now be replaced by a single formula for 𝑑𝑛 .
To see what the formula for 𝑑𝑛 should be, let us take the standard form for the Fourier series in terms of
coses and sines and convert it into complex exponential form;
∑
∞
∑
∞
𝑓 (𝑥) = 𝑎0 + 𝑎𝑛 cos(𝑛𝑥) + 𝑏𝑛 sin(𝑛𝑥)
𝑛=1 𝑛=1
∑∞
𝑎𝑛 ( 𝑖𝑛𝑥 ) ∑ ∞
𝑏𝑛 ( 𝑖𝑛𝑥 )
−𝑖𝑛𝑥
= 𝑎0 + 𝑒 +𝑒 + 𝑒 − 𝑒−𝑖𝑛𝑥
𝑛=1
2 𝑛=1
2𝑖
( ) ( )
∑∞
𝑎 𝑛 𝑏𝑛 ∑ ∞
𝑎𝑛 𝑏𝑛
= 𝑎0 + + 𝑖𝑛𝑥
𝑒 + − 𝑒−𝑖𝑛𝑥 .
𝑛=1
2 2𝑖 𝑛=1
2 2𝑖
This expression has essentially three terms in it which we need to compare this equation with
∑
𝑓 (𝑥) = 𝑑𝑛 𝑒𝑖𝑛𝑥 .
𝑛∈ℤ
The first constant term 𝑎0 corresponds to the 𝑛 = 0 term 𝑑0 𝑒𝑖0𝑥 = 𝑑0 in the above expression so that we
have
𝜋
1
𝑑0 = 𝑎0 = 𝑓 (𝑥)𝑑𝑥.
2𝜋 ∫−𝜋
The second term contains all the terms 𝑒𝑖𝑛𝑥 for 𝑛 > 0. So for 𝑛 > 0 comparing the two equations we have
( )
𝑎𝑛 𝑏𝑛
𝑑𝑛 = +
2 2𝑖
( 𝜋 ) ( 𝜋 )
1 1 1 1
= 𝑓 (𝑥) cos(𝑛𝑥)𝑑𝑥 + 𝑓 (𝑥) sin(𝑛𝑥)𝑑𝑥
2 𝜋 ∫−𝜋 2𝑖 𝜋 ∫−𝜋
𝜋
1
= 𝑓 (𝑥) (cos(𝑛𝑥) − 𝑖 sin(𝑛𝑥)) 𝑑𝑥
2𝜋 ∫−𝜋
𝜋
1
= 𝑓 (𝑥)𝑒−𝑖𝑛𝑥 𝑑𝑥
2𝜋 ∫−𝜋
The third and last term contains all the terms of the form 𝑒−𝑖𝑛𝑥 where 𝑛 > 0, or to put it another way all
122
the term of the form 𝑒𝑖𝑛𝑥 where 𝑛 < 0. Comparing the two equations we see that for 𝑛 < 0 we have
( )
𝑎−𝑛 𝑏−𝑛
𝑑𝑛 = −
2 2𝑖
( 𝜋 ) ( 𝜋 )
1 1 1 1
= 𝑓 (𝑥) cos(−𝑛𝑥)𝑑𝑥 − 𝑓 (𝑥) sin(−𝑛𝑥)𝑑𝑥
2 𝜋 ∫−𝜋 2𝑖 𝜋 ∫−𝜋
𝜋
1
= 𝑓 (𝑥) (cos(𝑛𝑥) − 𝑖 sin(𝑛𝑥)) 𝑑𝑥
2𝜋 ∫−𝜋
𝜋
1
= 𝑓 (𝑥)𝑒−𝑖𝑛𝑥 𝑑𝑥.
2𝜋 ∫−𝜋
Remarkably the three expressions for 𝑑𝑛 are all consistent with the single formula
𝜋
1
𝑑𝑛 = 𝑓 (𝑥)𝑒−𝑖𝑛𝑥 𝑑𝑥.
2𝜋 ∫−𝜋
Note that it is 𝑒−𝑖𝑛𝑥 that appears in the integrand and not 𝑒𝑖𝑛𝑥 . This is perhaps unexpected because in the
case of Euler’s formulae for Fourier coefficients in terms of sines and coses, the function appearing in the
integrand was always tied to the coefficient we were trying to calculate, e.g. the formula for 𝑏7 involved
integrating 𝑓 (𝑥) sin(7𝑥) and 𝑏7 was the coefficient of sin(7𝑥). Here on the other hand the formula for 𝑑𝑛
involves integrating over 𝑓 (𝑥)𝑒−𝑖𝑛𝑥 , even though 𝑑𝑛 is the coefficient of 𝑒𝑖𝑛𝑥 .
Example
The function 𝑓 (𝑥) is periodic with period 2𝜋, and is defined to be 𝑓 (𝑥) = 𝑒2𝑥 for −𝜋 ≤ 𝑥 < 𝜋. Find the
Fourier expansion of 𝑓 (𝑥) in complex exponential form.
123
We should expect that this series will produce real values. To see that this is the case we use the fact the
positive 𝑛 coefficients must satisfy
𝑎𝑛 𝑏𝑛 (−1)𝑛 sinh(2𝜋) (−1)𝑛 sinh(2𝜋)(2 + 𝑖𝑛)
+ = = (4.0.5)
2 2𝑖 𝜋(2 − 𝑖𝑛) 𝜋(4 + 𝑛2 )
and the negative ones
𝑎𝑛 𝑏𝑛 (−1)𝑛 sinh(2𝜋) (−1)𝑛 sinh(2𝜋)(2 − 𝑖𝑛)
− = = (4.0.6)
2 2𝑖 𝜋(2 + 𝑖𝑛) 𝜋(4 + 𝑛2 )
From which we have
4(−1)𝑛 sinh(2𝜋)
𝑎𝑛 = , (4.0.7)
𝜋(4 + 𝑛2 )
and
𝑏𝑛 2𝑖𝑛(−1)𝑛 sinh(2𝜋)
= , (4.0.8)
𝑖 𝜋(4 + 𝑛2 )
2𝑛(−1)𝑛+1 sinh(2𝜋)
𝑏𝑛 = . (4.0.9)
𝜋(4 + 𝑛2 )
which is just read from the complex sum:
sinh(2𝜋)
𝑎0 = . (4.0.10)
2𝜋
Parseval’s Theorem is a powerful way to evaluate all sorts of infinite sums which would otherwise be
difficult to evaluate. For instance we will be able to show that
∑ 1 1 1 𝜋2
= 1 + + + ⋯ = ≈ 1.2237.
𝑛 odd
𝑛2 9 25 8
Parseval’s Theorem is basically Pythagoras’ Theorem applied to periodic functions. One way to ‘prove’
Pythagoras theorem is particularly straightforward in vector notation. If 𝐯 = 𝑥𝐢 + 𝑦𝐣 then the length of the
hypotenuse squared is
which is the familiar result. An important point in this calculation is that when the brackets are expanded,
the ‘cross-term’ 2𝑥𝑦𝐢.𝐣 = 0 because of the orthogonality of 𝐢 and 𝐣. Going back to the more general vector
analogy considered earlier, if 𝐱𝟏 , 𝐱𝟐 and 𝐱𝟑 are orthogonal (but not necessarily unit length) vectors and
∑
3
𝐯= 𝑎𝑖 𝐱𝑖 = 𝑎1 𝐱1 + 𝑎2 𝐱2 + 𝑎3 𝐱3 ,
𝑖=1
124
(x, y)
v = xi + yj
(0, 0) x
|𝐯|2 = 𝐯.𝐯
( ) ( )
= 𝑎1 𝐱1 + 𝑎2 𝐱2 + 𝑎3 𝐱3 . 𝑎1 𝐱1 + 𝑎2 𝐱2 + 𝑎3 𝐱3
where again in the third line in expanding out the brackets we have been able to throw away all the cross-
terms because of the orthogonality of the vectors 𝐱𝟏 , 𝐱𝟐 and 𝐱𝟑 .
∑
∞
∑
∞
𝑓 (𝑥) = 𝑎0 .1 + 𝑎𝑛 cos(𝑛𝑥) + 𝑏𝑛 sin(𝑛𝑥).
𝑛=1 𝑛=1
125
where in the third line again we have thrown away all the cross-terms in the expansion of the brackets
because of the orthogonality of the functions, and in the fourth line we have used our earlier result that
(1, 1) = 2𝜋, and that (cos(𝑛𝑥), cos(𝑛𝑥)) = (sin(𝑛𝑥), sin(𝑛𝑥)) = 𝜋. Putting these two expressions for the
length squared |𝑓 (𝑥)|2 together we arrive at Parseval’s Theorem:
𝜋 ∑
∞
( )
2 2
(𝑓 (𝑥)) 𝑑𝑥 = 2𝜋(𝑎0 ) + 𝜋 (𝑎𝑛 )2 + (𝑏𝑛 )2 .
∫−𝜋 𝑛=1
Let us see how this can be used to evaluate the sum given at the start of this section
• Example By considering the length squared of the periodic function 𝑓 (𝑥) of period 2𝜋 where
⎧
⎪ −1 for − 𝜋 < 𝑥 < 0,
𝑓 (𝑥) = ⎨
⎪ 1 for 0 ≤ 𝑥 ≤ 𝜋.
⎩
show that
∑ 1 1 1 𝜋2
2
= 1 + + + ⋯ = ≈ 1.2237.
𝑛 odd
𝑛 9 25 8
We have already calculated the Fourier series for this function (see Example 1 in the section of Even
and Odd functions), and found that
( )
∑ 4 4 sin(3𝑥) sin(5𝑥)
𝑓 (𝑥) = sin(𝑛𝑥) = sin(𝑥) + + +… .
𝑛 odd
𝑛𝜋 𝜋 3 5
or in other words that 𝑎0 = 𝑎𝑛 = 0 and
⎧
⎪ 0 for 𝑛 even
𝑏𝑛 = ⎨
⎪ 𝜋𝑛 for 𝑛 odd.
4
⎩
For our function 𝑓 (𝑥) = (±1) = 1 so the left hand side of Parseval’s theorem is given by
2 2
𝜋 𝜋
2
(𝑓 (𝑥)) 𝑑𝑥 = 𝑑𝑥 = 2𝜋,
∫−𝜋 ∫−𝜋
and substituting the known values of the Fourier coefficients into the right-hand side gives
∑ ( 4 )2
2𝜋 = 0 + 𝜋
𝑛 odd
𝜋𝑛
16 ∑ 1 ∑ 1 𝜋 𝜋2
= ⇒ = 2𝜋 = .
𝜋 𝑛 odd 𝑛2 𝑛 odd
𝑛 2 16 8
126
In the lecture I will briefly discuss the non-examinable topic about what the Parseval theorem tells us
about which functions can be represented as a Fourier series and which not.
127