0% found this document useful (0 votes)
59 views186 pages

Mathematical Methods for Integration and Analysis

The document outlines various mathematical methods, focusing on integration, multidimensional integration, vector calculus, Fourier series, Laplace transforms, and complex functions. It includes acknowledgments for adapted chapters and a detailed table of contents that lists topics covered in each section. The content is structured to provide foundational knowledge and techniques for solving mathematical problems across different areas of calculus and analysis.

Uploaded by

gw7qsfcrx9
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
59 views186 pages

Mathematical Methods for Integration and Analysis

The document outlines various mathematical methods, focusing on integration, multidimensional integration, vector calculus, Fourier series, Laplace transforms, and complex functions. It includes acknowledgments for adapted chapters and a detailed table of contents that lists topics covered in each section. The content is structured to provide foundational knowledge and techniques for solving mathematical problems across different areas of calculus and analysis.

Uploaded by

gw7qsfcrx9
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

A . B A S S O M , E . C R I P P S , A . D E V I L L E R S , L . J E N N I N G S , A . N I E M E Y E R , T.

S T E M L E R ,

L . S T O YA N O V

M AT H E M AT I C A L
METHODS 2
2 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

Acknowledgements: The chapter on one dimensional inte-


gration is adapted from notes originally by Luchezar Stoyanov,
adapted by Jenny Hopwood, and then Les Jennings.
The chapters on multidimensional integration and vector calcu-
lus are adapted from notes by Edward Green of the University of
Adelaide.
The chapter on Fourier series is adapted from notes by Des Hill.
The chapter on Laplace transforms is adapted from notes origi-
nally by Luchezar Stoyanov.
Contents

1 Integration and some applications 7


1.1 Inverse functions 7
1.2 Differentiation and anti-differentiation 9
1.3 Techniques of integration 9
1.4 Riemann sums and integrals 15
1.5 The Fundamental Theorem of Calculus (FTC) 18
1.6 Some applications of Riemann sums and integrals 20
1.7 Improper integrals 23
1.8 Quadrature 25

2 Double and triple integrals 27


2.1 Double integrals 27
2.2 Triple integrals 35
2.3 Centre of mass 40
2.4 More applications to physics 41

3 Change of coordinates in double and triple integrals 43


3.1 Change of coordinates in double integrals 43
3.2 Change of coordinates in triple integrals 48
3.3 Three important coordinate changes 50

4 Path and surface integrals 57


4.1 Revision of parametric forms 57
4.2 Length of curves 62
4.3 Path integrals of a function 65
4.4 Areas of surfaces 65
4.5 Surface integrals of a function 68
4 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

5 Vector fields 71
5.1 Path integrals of vector fields in two dimensions 72
5.2 Flux through a surface in R3 76
5.3 Work done by the gradient of a scalar field 77

6 Conservative vector fields 79


6.1 Which vector fields are conservative? 80
6.2 Finding potentials 84

7 Fourier series 91
7.1 Introduction 91
7.2 Calculation of the Fourier coefficients 94
7.3 Functions of an arbitrary period 96
7.4 Convergence of Fourier series 97
7.5 Functions defined over a finite interval 98
7.6 Even and odd functions 99
7.7 Fourier cosine series for even functions 100
7.8 Fourier sine series for odd functions 101
7.9 Half-range expansions 102
7.10 Parseval’s Theorem 103
7.11 Differentiation of Fourier series 105
7.12 Integration of Fourier series 107

8 Laplace Transforms 109


8.1 The Laplace transform and its inverse 109
8.2 Inverse Laplace transforms of rational functions 114
8.3 The Laplace transform of derivatives and integrals of f (t) 115
8.4 Solving differential equations 118
8.5 Shift theorems 120
8.6 Derivatives of transforms 125
8.7 Convolution 125
8.8 Appendix: Some applications of the Laplace transform 130
mathematical methods 2 5

9 Complex Functions — Derivatives 141


9.1 Complex numbers and some basic functions 141
9.2 The derivative of a function of a complex variable 144
9.3 Rules for derivatives of analytic functions 145
9.4 The Cauchy–Riemann equations 146
9.5 Solutions of Laplace’s equation 149

10 Probability and Statistics 153


10.1 A recap on probability models 154
10.2 Random variables 158
10.3 Statistical Inference 171

11 Index 185
1
Integration and some applications

1.1 Inverse functions

Recall that a function f ( x ) assigns to each value x in its domain


a unique value y = f ( x ) in its range. In applications we might
view a function as a machine which, for a given input x, generates
a unique output y = f ( x ). It is then natural to ask ourselves, given
an output value y, for which input value x is y = f ( x )? Or we
might ask: can we find another function g such that x = g(y)? In
other words, we are looking for a function g with the property that
x = g( f ( x )) for x in a certain domain. If such a function g exists,
it is called the inverse function of f and is usually denoted by x =
f −1 (y). It exists provided that f is one-to-one and that the domains
and ranges of the two functions are suitably restricted, so that both
functions are onto. It is usual to call x the variable in the domain
of f and to call y the variable in the range of f , so that y = f ( x )
and x = f −1 (y). Hence y = f ( x ) = x2 has an inverse function

x = f −1 (y) = y provided we choose x ∈ [0, ∞), y ∈ [0, ∞). Also,
y = f ( x ) = x3 has an inverse function x = f −1 (y) = y1/3 for all x
and y; y = f ( x ) = e x has an inverse x = f −1 (y) = ln y for all x but In the literature, the natural logarithm
with y ∈ (0, ∞). For suitable domains, we have ln x is also sometimes denoted by
log x.

f f −1 ( y ) = y f −1 f ( x ) = x.
 
and

Later on, we might discard the variable y and also use x as the
variable in the domain of f −1 ( x ) on the understanding that here x
is in the range of f .

1.1.1 Inverse trigonometric functions, derivatives

The inverse trigonometric functions exist for suitable domains and


ranges and are very important in many applications.
If y = sin x, then given y ∈ [−1, 1] we can find the angle x ∈
[−π/2, π/2] using x = asin(y). (Or arcsin(y), or less preferred,
 −1
sin−1 (y) as it can be confused with sin(y) .)
Similarly, if y = cos x, then given y ∈ [−1, 1] we can find the
angle x ∈ [0, π ] using x = acos(y) (or arccos(y)).
Finally, if y = tan x, we can find the angle x ∈ (−π/2, π/2)
8 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

x Figure 1.1: acos(y) and asin(y)


π

π/2

-1 0 1 y

−π/2

Figure 1.2: atan(y)


x

π/2

-10 0 10 y

−π/2

using x = atan(y) (or arctan(y)).


The derivative of the function y = asin x is obtained by differ-
dy
entiating x = sin y with respect to x, giving 1 = cos y . Since
√ dx
2 2
x = sin y, we have x + (cos y) = 1, and so cos y = 1 − x2 (always
positive as y ∈ [−π/2, π/2]). It follows that
d asin( x ) 1
= √ for x ∈ (−1, 1).
dx 1 − x2
d asin( x)
By similar methods we can show that, Notice that
dx
≥ 1,
d acos( x)
d acos( x ) 1 ≤ −1, and
dx
= −√ for x ∈ (−1, 1), d atan( x)
dx 1 − x2 0< ≤ 1.
dx
and Is asin( x ) + acos( x ) constant?
d atan( x ) 1
= .
dx 1 + x2
Exercise 1.1.1. For a function f and its inverse f −1 , where y = f ( x ) or
x = f −1 (y), show that
 0 1  0 1
f −1 f −1 ( y ) = 0 −1 

f (x) = 0 or
f (x) f f (y)
where, remember, that 0 means differentiation with respect to the argument
of the function.
mathematical methods 2 9

1.2 Differentiation and anti-differentiation

In high school you learnt anti-differentiation as the inverse process


(or undoing process) to differentiation, and that this is related to area
under a curve. In these notes we introduce finding areas under
curves, call the process integration and prove that it is related to
anti-differentiation by the Fundamental Theorem of Calculus.
If f is a function on some interval, then an anti-derivative of f is
any function F such that F 0 ( x ) = f ( x ) on the interval. The entries in See Definition 1.5.
Table 1.1 should be known by all.

derivative ← function → anti-derivative Table 1.1: Derivatives and anti-


derivatives. Note that a, c denote
f 0 (x) f (x) F(x) constants. * Care using this formula is
0 a ax + c needed, as the domain of integration
cannot contain a. ** Provided cos x 6= 0
x n +1 in the domain of integration, which is
nx n−1 x n (n 6= 0, −1) +c where tan( x ) is undefined.
n+1
−1 1
( x − a )2
(x 6= a) ln | x − a| + c *
x−a
−b 1 −1
( x − a ) b +1
(x 6= a, b ≥ 2) (b−1)( x − a)b−1
+c *
( x − a)b
ex exp( x ) = e x ex + c
1
ln x (x > 0) —
x
cos x sin x − cos x + c
− sin x cos x sin x + c
1
tan x — **
cos2 x
1
√ asin x (−1 < x < 1) —
1 − x2
1
−√ acos x (−1 < x < 1) —
1 − x2
1
atan x —
1 + x2
1
— √ (−1 < x < 1) asin x + c
1 − x2
1
— −√ (−1 < x < 1) acos x + c
1 − x2
1
— atan x + c
1 + x2

1.3 Techniques of integration

This section introduces some important techniques for finding an


anti-derivative of a given function. The first technique, integra-
tion using partial fractions, deals with rational functions. Later we
examine two rules of differentiating functions and consider their
implications for integration.
10 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

1.3.1 Partial fractions


A rational function is a function which consists of a polynomial
divided by a polynomial. For example, f ( x ) = 3x −1 is a rational
x 2 −1
function.
It is hard to find an anti-derivative of f ( x ) = 3x −1 directly.
x 2 −1
However, if we observe that
3x − 1 2 1
f (x) = = + ,
x2 − 1 x+1 x−1
it is easy to find an anti-derivative, as

F(x) = 2 ln(| x + 1|) + ln(| x − 1|).

In general, we would like to solve the following problem, called


decomposing into partial fractions:

P( x )
Given a rational function f ( x ) = Q( x) with deg( P( x )) <
deg( Q( x )) find polynomials gi ( x ) and Ai ( x ) such that

A1 ( x ) Ak ( x ) What should you do if the numerator


f (x) = +···+ , is of higher or equal degree to the
( g1 ( x ))n1 ( gk ( x ))nk denominator?
A degree 2 polynomial is irreducible
where the polynomials gi ( x ) have degree 1 or are irreducible if it has no zeroes.
of degree 2 and deg( Ai ( x )) < deg( gi ( x )) and n1 , . . . , nk are
positive integers.

We now discuss various methods of how we can obtain such a


decomposition.
Case 1: Denominator has distinct linear factors

P( x ) A1 Ak
f (x) = = +...+ ,
( x − a1 ) . . . ( x − a k ) x − a1 x − ak
where a1 , . . . , ak are pairwise distinct. Note that then

P ( x ) = A 1 ( x − a 2 ) · · · ( x − a k ) + · · · + A k ( x − a 1 ) · · · ( x − a k −1 ).

So choosing x = ai we see that

P ( a i ) = A i ( a i − a 1 ) · · · ( a i − a i −1 ) · ( a i − a i +1 ) · · · ( a i − a k )

This leads us to the following formula:


P ( ai )
Ai = .
( a i − a 1 ) · · · ( a i − a i −1 ) · ( a i − a i +1 ) · · · ( a i − a k )
3x − 1 3x − 1 A1 A2
Example 1.1. f ( x ) = = = + ,
x2 − 1 ( x + 1)( x − 1) x+1 x−1
where
3(−1) − 1
A1 = =2
−1 − 1
3(1) − 1
A2 = = 1.
2

2 1
Hence f ( x ) = x +1 + x −1 .
mathematical methods 2 11

Note that an anti-derivative for f ( x ) can be found in Table 1.1.


Case 2: Denominator has repeated linear factors

P( x ) B1 Bc
f (x) = = +...+ .
( x − a)c x−a ( x − a)c
P( x ) = B1 ( x − a)c−1 + · · · + Bc−1 ( x − a)1 + Bc .
So choosing x = a we see that P( a) = Bc . If we then evalu-
ate the expression for further (simple) values of x, we can deduce
equations for the remaining Bi s.

Example 1.2.
P( x ) 2x + 3 B1 B2 B3
f (x) = = 3
= + 2
+ .
Q( x ) ( x − 1) ( x − 1) ( x − 1) ( x − 1)3

Then B3 = P(1) = 5. We now evaluate in x = 0 and x = 2.

x=0 −3 = − B1 + B2 − 5 (1.1)
x=2 7 = B1 + B2 + 5 (1.2)

2 5
Hence B1 = 0 and B2 = 2. Thus f ( x ) = ( x −1)2
+ ( x −1)3
.

Note that an anti-derivative for f ( x ) can be found in Table 1.1.


Case 3: Denominator has an irreducible factor of degree 2

P( x ) A1 C x + C2
f (x) = = + 21 .
( x − a)( x2 + bx + c) x−a x + bx + c
P( x ) = A1 ( x2 + bx + c) + (C1 x + C2 )( x − a). Hence (as before)

P( a)
A1 = .
( a2 + ba + c)

If we then evaluate the expression for further (simple) values of x,


we can deduce equations for C1 and C2 .

Example 1.3.

P( x ) 2x + 3 A1 C x + C2
f (x) = = 2
= + 12 .
Q( x ) ( x − 1)( x + 1) ( x − 1) x +1

Then A1 = P(1)/2 = 5/2. We now evaluate in x = 0 and x = −1.

x=0 −3 = − 25 + C2 C2 = −1/2
x = −1 − 41 = − 54 − C21 + C2
2 C1 = −5/2

5/2 (5/2) x + 1/2


f (x) = − .
( x − 1) x2 + 1
C1 x + C2
We note that an anti-derivative of f ( x ) = is There is no need to memorise this
x2 + bx + c formula, as we will learn another
method: substitution.
(bC1 − 2C2 )
 
C1 2 (b + 2x )
ln( x + bx + c) − √ atan √ .
2 −b2 + 4c −b2 + 4c
12 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

A different Method (Method of Undetermined Coefficients):


We would like to find an expression
5x + 7 c c x + c3
2
= 1 + 22 . (1.3.1)
( x − 2)( x + 4x + 5) x − 2 x + 4x + 5
Note that x2 + 4x + 5 = ( x + 2)2 + 1 cannot be factored to real
roots. We need a method to find c1 , c2 and c3 . Putting the RHS over
a common denominator, that of the LHS of course, gives
c1 ( x2 + 4x + 5) + (c2 x + c3 )( x − 2)
( x − 2)( x2 + 4x + 5)
(c1 + c2 ) x2 + (4c1 − 2c2 + c3 ) x + (5c1 − 2c3 )
= ,
( x − 2)( x2 + 4x + 5)
Now equating the numerator with the numerator of the LHS of Putting x = 2 into both LHS and
(1.3.1) gives RHS of the numerators (when over a
common denominator) gives directly
17 = 17c1 , or c1 = 1.
5x + 7 = (c1 + c2 ) x2 + (4c1 − 2c2 + c3 ) x + (5c1 − 2c3 )

which implies, because this has to be true for all x,

x2 : c1 + c2 = 0, x1 : 4c1 − 2c2 + c3 = 5, 1 = x0 : 5c1 − 2c3 = 7,

a set of three simultaneous equations in three variables, which has


solution c1 = 1, c2 = −1 and c3 = −1. Hence
5x − 7 1 −x − 1
= + .
( x − 2)( x2 + 4x + 5) x − 2 x2 + 4x + 5

1.3.2 Integration by parts


Integration by parts is the rule for integrating the product rule of
differentiation. We have
d
f ( x ) g ( x ) = f 0 ( x ) g ( x ) + f ( x ) g 0 ( x ),

dx
hence on integrating
 b Z b
d 
f ( x ) g( x ) = f ( x ) g( x ) dx (by the FTC, see Section 1.5)
a a dx
Z b Z b
= f 0 ( x ) g( x ) + f ( x ) g0 ( x )dx.
a a
Re-arranging we get:

Z b  b Z b
0
f ( x ) g( x )dx = f ( x ) g( x ) − f ( x ) g0 ( x )dx,
a a This
a R 0 rule is sometimes remembered as
u v = [uv] − uv0 .
R
Z
an antiderivative for f 0 ( x ) g( x ) is f ( x ) g( x ) − f ( x ) g0 ( x )dx.

Renaming f ( x ) as F ( x ) with f 0 ( x ) becoming F 0 ( x ) = f ( x )


(derivative anti-derivative pair) gives the first line in Table 1.2.
Z b  b Z b
f ( x ) g( x )dx = F ( x ) g( x ) − F ( x ) g0 ( x )dx
a a a
Z
and an antiderivative for f ( x ) g( x ) is F ( x ) g( x ) − F ( x ) g0 ( x )dx.
mathematical methods 2 13

Z 4
Example 1.4. Find x ln xdx.
2
Solution: Notice from the above formula that the f function is integrated
while the g function is differentiated, so we need to choose which one is f
and which one is g to make the integral do-able using simple algebra. Here
we are going to choose f ( x ) = x and g( x ) = ln x (in this case choosing
the reverse is possible too but it makes the algebra a little bit harder). Thus
F ( x ) = x2 /2 and g0 ( x ) = 1/x.
4
x2
Z 4  Z 4 2
x 1
x ln xdx = ln x − dx
2 2 2 2 x 2
 2 4
x Recall that ln( ab) = ln( a) + ln(b) and
= 8 ln(4) − 2 ln(2) − = 14 ln(2) − 3
4 2 so ln( an ) = n ln( a).

Example 1.5. Find an antiderivative for x2 cos x.

Solution: Obviously x2 is good to differentiate to 2x and then to 2,


whereas integrating cos x twice gets back to − cos x. So let f ( x ) = cos x As an exercise, see what would
and g( x ) = x2 , then F ( x ) = sin x and g0 ( x ) = 2x giving, happen if you took f ( x ) = x2 and
g( x ) = cos( x ).
Z   Z
2 2
x cos xdx = x sin x − 2x sin xdx.

Repeating the procedure


Z   Z
2x sin xdx = 2x (− cos x ) − 2(− cos x )dx

= −2x cos x + 2 sin x,

Hence Z
x2 cos xdx = x2 sin x + 2x cos x − 2 sin x + c.

You can check this by differentiating the right-hand side.

1.3.3 Integration by substitution


Integration by Substitution makes the chain rule anti-differentiation
0
= f 0 g( x ) g0 ( x ) or

easier to see in Table 1.2. Recall that f g( x )
0
= f g( x ) g0 ( x ), so

equivalently F g( x )

 Z
f g( x ) g0 ( x )dx

F g( x ) =

In practice we do a variable substitution u = g( x ) so that the


integral “looks nicer”. Now consider u as the variable and we have
du = g0 ( x )dx. So
Z Z
f g( x ) g0 ( x )dx =
   
f u du = F u = F g( x ) .

Rr √
Example 1.6. Find −r r2 − x2 dx where r is a positive constant.
14 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

Solution: Using x = r sin θ it follows that dx = r cos θdθ. In addition


we need to take into account that x = ±r results in θ = ±π/2. Therefore
we get:
Z r p Z π/2
r2 − x2 dx = r cos θ r cos θ dθ
−r −π/2
r2
Z π/2 Z π/2
2
=r cos2 θ dθ = (1 + cos 2θ ) dθ
−π/2 2 −π/2
r2 πr2
 π/2
1
= θ + sin 2θ = .
2 2 −π/2 2

Example 1.7. Find


Z 1
x+1
dx.
0 x2 + 4x + 5
We notice that + 4x + 5 = ( x + 2)2 + 1 is irreducible (has no roots), so
x2
this is already in the partial fraction form.

Solution: Using the substitution u = x + 2, we get dx = du, when


x = 0, u = 2, and when x = 1, u = 3. Substituting x = u − 2 into the
integrand, changing limits to the new variable, gives
Z u =3 Z u =3 
u−1

u 1
du = − du
u =2 u 2 + 1 u =2 u2 + 1 u2 + 1
 3 We did another substitution here:
1 2
= 2
ln(u + 1) − atan u Rv =u u + 1, so dv = 2udu, to get
2 u2 +1
du = 21 ln(u2 + 1)
2
1 1
= ln 10 − atan 3 − ln 5 + atan 2
2 2
1
= ln 2 − atan 3 + atan 2.
2
x +1 1
An anti-derivative for x2 +4x +5
is 2 ln(( x + 2)2 + 1) − atan( x + 2) .

This method can be generalised to find an anti-derivative of any


+C2
fraction of the form xC21+xbx +c
where x2 + bx + c is irreducible.
Z
Example 1.8. Find the general formula for f 0 ( x )e f ( x) dx.

Solution: Use the substitution u( x ) = f ( x ) so that du = f 0 ( x )dx, now


substitute into the integral,
Z Z
f 0 ( x )e f ( x) dx = eu du = eu + c = e f ( x) + c.

Lines 5-9 of Table 1.2 can be shown similarly.


Substitution also has other uses in simplifying the integration
of functions where it is not clear what the substitution should be.
Sometimes trial and error are needed.
mathematical methods 2 15

derivative function anti-derivative Table 1.2: Differentiation and anti-


differentiation using the product
Z
0
f 0 ( x ) g( x ) + f ( x ) g0 ( x ) f ( x ) g( x ) F ( x ) g( x ) − F(x) g (x) + c rule/integration by part and the chain
rule/substitution method. Note that
f 0 g( x ) g0 ( x )
 
f g( x ) — c is a constant. * f ( x ) must not be
zero for any point in the interval of
f 0 g( x ) g0 ( x )
 
— f g( x ) + c integration.
f 0 ( x ) g( x ) − f ( x ) g0 ( x ) f (x)

g2 ( x ) g( x )
 n +1
n f (x)
— f 0 (x) f (x) +c
n+1
— f 0 ( x )e f ( x) e f (x) + c
f 0 ( x )/ f ( x )

— ln | f ( x )| + c *
f 0 ( x ) sin f ( x )
 
— − cos f ( x ) + c
f 0 ( x ) cos f ( x )
 
— sin f ( x ) + c

1.4 Riemann sums and integrals

Motivation : Area
Let f ( x ) ≥ 0 be a continuous function on some finite interval
[ a, b], and let A be the area of the region under the graph of f , i.e.
the area of {( x, y) ∈ R2 : x ∈ [ a, b], 0 ≤ y ≤ f ( x )}.
To calculate an approximation to A, choose n ∈ N and consider a
partition a b

P : a = x 0 < x 1 < x 2 < . . . < x n −1 < x n = b

of [ a, b] into n subintervals [ x j−1 , x j ]( j = 1, 2, ..., n) of (arbitrary)


lengths ∆x j = x j − x j−1 . In general the lengths ∆x j do not have to
be equal; however, we define the size || P|| = max ∆x j , 1 ≤ j ≤ n.
Next, for any j choose a point c j ∈ [ x j−1 , x j ], and consider the
rectangle in the plane with base the interval [ x j−1 , x j ] and height
f (c j ). If A j is the area of the region under the graph of f and above a b

the interval [ x j−1 , x j ], then A j ≈ Area of rectangle = f (c j )∆x j , so


x0 x1 x2 x3 x4 x5 x6 x7 x8 x9 x10 x11

A = A1 + A2 + ... + An
n
≈ f (c1 )∆x1 + f (c2 ) ∆x2 + ... + f (cn )∆xn = ∑ f (c j )∆x j .
j =1

The approximation in the above formula is likely to become bet-


ter as the size of the subintervals becomes smaller. This is demon-
strated by the pictures on the margin: The red rectangles use a
coarser partition of the interval [ a, b] than the green rectangles and
a b
overlaying both pictures we can see that the area of the green rect- x0 x1 x2 x3 x4 x5 x6 x7 x8 x9 x10 x11 x12 x13 x14 x15 x16 x17 x18 x19 x20 x21x22

angles is a better approximation to the area of the curve than the


area of the red rectangles.
Hence one would expect that it would converge to A, i.e. become
exact, as || P|| → 0. That is, it would be expected that for ‘most’
functions
n
∑ f (c j )∆x j ,
a b
A = lim x0 x1 x2 x3 x4 x5 x6 x7 x8 x9 x10 x11
|| P||→0 j =1
16 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

where the limit is taken over all possible partitions and all possible
choices of c j s. These considerations lie beneath the following defini-
tion of the Riemann integral (there are other types of integrals).

Definition 1.9. (Riemann Sums and Definite integrals)


We say that f is bounded when there
Let f be a bounded function on some finite interval [ a, b] . Let P be any exist constants N, M such that N ≤
partition of [ a, b] into n subintervals [ x j−1 , x j ], j = 1, 2, ..., n, and c j any f ( x ) ≤ M ∀ x ∈ [ a, b].
point in the jth interval. The sum
n
SP = ∑ f (c j )∆x j
j =1

is called a Riemann Sum for f determined by the partition P and the


points c j . The definite integral (Riemann) of f over [ a, b] is defined by a
limiting process
Z b n

a
f ( x )dx = lim
|| P||→0
∑ f (c j )∆x j ,
j =1

if this limit exists, in which case we say that f is (Riemann) integrable


over [ a, b].

Notes:

• The variable x in the left-hand-side of the definition is a dummy


variable, i.e. it can be replaced by any other variable.

• This means that in principle all possible partitions would have


to be taken into account when determining whether the limit
exists or not. Fortunately, it can be shown that it is enough to
consider only uniform partitions (partitions with subintervals
of equal length; i.e. with ∆x j = (b − a)/n ∀ j), in which case
|| P|| → 0 is the same as n → ∞. That is, if for some function f
the limit exists for this special kind of partition, then it exists for
general partitions as well and the limit is the same. It is also not
necessary to consider all possible choices of the points c j .

• Finally, note that the definition of the definite integral has not
used derivatives or anti-derivatives.

• For the interested student: the exact meaning of the limit in the
Rb
definition is the following: ∃ L ∈ R (L = a f ( x )dx) such that
∀e > 0, ∃δ > 0 s.t. ∀ partitions P of [ a, b] and for any choice of
the points c j ∈ [ x j−1 , x j ]( j = 1, 2, ..., n), | L − ∑nj=1 f (c j )∆x j | <
e ∀ P s.t. || P|| < δ.F

Example 1.10. Let f ( x ) = c = constant on some interval [ a, b]. Then for


any uniform partition and any choice of the points c j ,
!
n n n
∑ f (c j )∆x j = ∑ c(b − a)/n = c(b − a) ∑ 1/n = c ( b − a ).
j =1 j =1 j =1
mathematical methods 2 17

So Z b
cdx = lim c(b − a) = c(b − a).
a n→∞

This is just the area of the rectangle of base length b − a and height c, as
would be expected.

Example 1.11. Let f ( x ) = x on the interval [0, 1]; f ( x ) is bounded


since 0 ≤ f ( x ) ≤ 1. Consider the partition P of [0, 1] into n subintervals
of equal length ∆x j = 1/n; i.e. x j = j/n, j = 0, 1, ..., n. Choose c j =
j/n ∀ j. Then using the formula 1 + 2 + ... + n = n(n + 1)/2 gives This formula is easily proved using
induction.
n n
1 j
SP = ∑ f ( j/n) n = ∑ n2
j =1 j =1

n ( n +1)
1 + 2 + · · · + ( n − 1) + n 2 n+1
= = =
n2 n2 2n
1
which converges to 12 as n → ∞. Therefore 0 xdx = 1/2, which is
R

expected since this is the area of the triangle formed by the graph of f ( x ),
the x-axis and the line x = 1.

Definition 1.12. If f is integrable over [ a, b], we define


Z a Z b Z a
f ( x )dx ≡ − f ( x )dx, and f ( x )dx = 0.
b a a

1.4.1 Geometrical meaning of the Riemann integral


If f is integrable over [ a, b] and f ( x ) ≥ 0, ∀ x ∈ [ a, b], then the area
of the region under the graph of f and bounded by the x-axis is
Rb
given by a f ( x )dx, since this was our motivation for defining the
Riemann integral.
In general (when f is integrable over [ a, b]),
Z b
f ( x )dx = A+ − A−
a

is an oriented area, where

A+ = Area{( x, y) : x ∈ [ a, b], f ( x ) ≥ 0, 0 ≤ y ≤ f ( x )}

is the area of the region above the x-axis and under the graph of f ,
while

A− = Area{( x, y) : x ∈ [ a, b], f ( x ) ≤ 0, f ( x ) ≤ y ≤ 0}

is the (positive) area of the region under the x-axis and above the
graph of f .
It is clear from the definition of the integral that every Riemann
sum gives an approximation to the value of the integral. In general,
partitions of smaller size subintervals lead to better approximations.

Definition 1.13. A function f ( x ) is called piecewise continuous on a


given interval [ a, b] if f has only finitely many points of discontinuity in
[ a, b].
18 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

The following theorem shows that the family (set) of integrable


functions over a given interval is rather large.
Theorem 1.14. If f is a bounded and piecewise continuous function on a
finite interval [ a, b], then f is integrable over [ a, b].

Theorem 1.15. (Properties of Definite Integrals). Let f and g be


integrable functions over [ a, b]. Then:
Rb
(a) The functions f ± g are integrable and a ( f ± g)( x )dx =
Rb Rb
a f ( x )dx ± a g ( x )dx.

(b) For every constant c ∈ R the function c f is integrable and


Rb Rb
a ( c f )( x )dx = c a f ( x )dx.
Rb Rb
(c) If f ( x ) ≤ g( x ) ∀ x ∈ [ a, b], then a f ( x )dx ≤ a g( x )dx.

(d) If m, M are constants such that m ≤ f ( x ) ≤ M ∀ x ∈ [ a, b],


then Z b
m(b − a) ≤ f ( x )dx ≤ M (b − a),
a

(e) For any c ∈ ( a, b)


Z b Z c Z b
f ( x )dx = f ( x )dx + f ( x )dx.
a a c

Proof. (a) Consider an arbitrary partition P of [ a, b] into subintervals


[ x j−1 , x j ] and choose an arbitrary c j ∈ [ x j−1 , x j ] for any j. Then for
the corresponding Riemann Sums for f ± g, f and g the following
relationship holds:
SP ( f ± g) = ∑nj=1 ( f ± g)(c j )∆x j = ∑nj=1 f (c j )∆x j ± ∑nj=1 g(c j )∆x j =
S P ( f ) ± S P ( g ).
Since f and g are integrable, taking limits as || P|| → 0 gives the
result.
Properties (b), (c) and (e) can be derived in a similar way from
the definition of the integral, while (d) is a consequence of (c) and
Example 1.10.

1.5 The Fundamental Theorem of Calculus (FTC)

Let f be a continuous function on an interval [ a, b] (so it bounded


by the Extreme Value Theorem1 and hence is integrable by Theo- 1
See Section 8.1 of the notes Mathe-
rem 1.14). matical Methods 1.

Definition 1.16. (Antiderivative)


A function F ( x ) is called an antiderivative for f ( x ) on [ a, b] if and
only if F ( x ) is continuous on [ a, b], differentiable on ( a, b) and F 0 ( x ) =
f ( x ) ∀ x ∈ ( a, b).
mathematical methods 2 19

If F1 ( x ) and F2 ( x ) are both antiderivatives of f on [ a, b], then


F1 ( x ) − F2 ( x ) is a constant on [ a, b]; i.e. there exists a constant c
such that F1 ( x ) = F2 ( x ) + c ∀ x ∈ [ a, b].

Theorem 1.17. (The Fundamental Theorem of Calculus). Let


f ( x ) be a continuous function on a finite closed interval [ a, b].
(a) The function
Z x
A( x ) = f (t)dt , x ∈ [ a, b],
a

is continuous on [ a, b], differentiable on ( a, b) and

A0 ( x ) = f ( x )

for all x ∈ ( a, b). That is, A( x ) is an antiderivative for f on [ a, b],


such that A( a) = 0.
(b) If F is any antiderivative for f on [ a, b], then
Z b
f ( x )dx = F (b) − F ( a).
a

Note. The following notation is frequently used:


b  b
F (b) − F ( a) = F ( x ) or F ( x ) .
a a

Proof. Sketch of Proof.


A( x +h)− A( x )
(a) Let x ∈ ( a, b). First show that A0 ( x + ) = limh→0+ h =
f ( x ), as follows. Let h > 0 be so small that x + h ≤ b (this will al-
ways be possible). By Property (e) of Theorem 1.15
Z x+h Z x
A( x + h) − A( x ) = f (t)dt − f (t)dt
a a
Z x Z x+h Z x
= f (t)dt + f (t)dt − f (t)dt
a x a
Z x+h
= f (t)dt
x

Since f is continuous on [ a, b], and therefore on [ x, x + h], it


follows from the Extreme Value Theorem that f has a minimum
and a maximum in [ x, x + h]; i.e. ∃u, v ∈ [ x, x + h] s.t.

f (u) ≤ f (t) ≤ f (v) for t ∈ [ x, x + h].

Then property (d) of Theorem 1.15 implies


Z x+h
f (u)h ≤ A( x + h) − A( x ) = f (t)dt ≤ f (v)h.
x

and therefore
A( x + h) − A( x )
f (u) ≤ ≤ f ( v ).
h
20 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

Now u → x and v → x as h → 0, and hence the continuity of f


implies that f (u) → f ( x ) and f (v) → f ( x ) as h → 0. It follows from
A( x +h)− A( x )
the Squeeze Theorem that A0 ( x + ) = limh→0+ h = f ( x ).
0 −
In the same way it can be shown that A ( x ) = f ( x ), so A is
differentiable at x and A0 ( x ) = f ( x ).
The continuity of A at a and b can be proved using arguments
similar to those above. This is much easier than proving the differ-
entiability.
(b) Since both F and A are antiderivatives for f on [ a, b], it fol-
lows that F ( x ) = A( x ) + c ( x ∈ [ a, b]) for some constant c ∈ R.
Thus, F (b) − F ( a) = ( A(b) + c) − ( A( a) + c) = A(b) − A( a) =
Rb Ra Rb
a f ( t )dt − a f ( t )dt = a f ( t )dt, as required.

Note: The FTC says that (for a continuous function and for x in
an appropriate interval) d a f (t)dt = f ( x ) and so d x f (t)dt =
Rx Ra
dx dx
− f ( x ).

Definition 1.18. (Indefinite integrals)


R
The indefinite integral f ( x )dx of a continuous function f ( x ) on
an interval I is the family of all antiderivatives of f on I. Any two anti-
derivatives differ by at most a constant so f ( x )dx = F ( x ) + c, c ∈ R,
R

where F is any antiderivative of f on I.

Example 1.19. An antiderivative of x2 on R is x3 /3, so x2 dx =


R

x3 /3 + c, c ∈ R.
R x4
Example 1.20. Let g( x ) = 1 1+1ln t dt, for x > 1. Find g0 ( x ).
Rx
Solution: Consider the function A( x ) = 1 1+1ln t dt, so A0 ( x ) = 1+1ln x
by the FTC. Since g( x ) = A( x4 ), application of the Chain Rule gives
3 3
g0 ( x ) = A0 ( x4 )( x4 )0 = 1+4x
ln( x4 )
= 1+44xln(x) for any x > 1.
Generalising this example, we can prove the following theorem.

Theorem 1.21. Let h1 ( x ) and h2 ( x ) be differentiable functions, on some


R h (x)
interval J, taking values in ( a, b) and let g( x ) = h 2( x) f (t)dt then
1

0
h20 ( x ) − f h1 ( x ) h10 ( x ).
 
g ( x ) = f h2 ( x )

Remark 1.22. Continuing, if h1 ( x ) and h2 ( x ) are differentiable func-


R h (x)
tions, on some interval J, taking values in ( a, b) and g( x ) = h 2( x) f ( x, t)dt,
1
where f is a function of two variables, then Hard exercise: try proving this from
first principles, ie definition of deriva-
Z h2 ( x )
∂ f ( x, t) tive.
g0 ( x ) = dt + f x, h2 ( x ) h20 ( x ) − f x, h1 ( x ) h10 ( x ).
 
h1 ( x ) ∂x

1.6 Some applications of Riemann sums and integrals

1.6.1 Volumes by cross-sections


Let S be a solid in space. Suppose that the cross-sectional areas
of S relative to the x-axis are known. That is, let the orthogonal
mathematical methods 2 21

projection of S onto the x-axis be an interval [ a, b], and suppose that


for any x ∈ [ a, b] the cross-sectional area A( x ) of S perpendicular
to the x-axis is known. The problem is to find the volume, V, of S
and Riemann sums will be used to solve it: this technique is called
volume by cross-sections.
Consider an integer n ≥ 1 and a partition P : a = x0 < x1 < x2 <
... < xn = b of [ a, b] into n subintervals [ x j−1 , x j ]( j = 1, 2, . . . , n). For
any j choose an arbitrary point c j ∈ [ x j−1 , x j ]. Let Vj be the volume
of the slice of S between the cross-sections at x = x j−1 and x = x j .
Then Vj can be approximated by the volume of a cylinder with base
of area A(c j ) and height ∆x j = x j − x j−1 . Then Vj ≈ A(c j )∆x j and
therefore V = ∑nj=1 Vj ≈ ∑nj=1 A(c j )∆x j . This is the Riemann sum
for the area function A( x ) determined by the partition P and the
choice of the points c j . If the function A( x ) is sufficiently regular
(for instance, if it is continuous) taking the limit as || P|| → 0 gives
Rb
the exact result V = a A( x )dx. y
1
Example 1.23. Let D be the region in the xy-plane under the graph of
the function y = x1/3 , 0 ≤ x ≤ 1. Find the volume, V, of the solid, S,
obtained by rotating D about the x-axis.
Solution: Given x ∈ [0, 1], let A( x ) be the cross-sectional area of S at x.
0 1 x
Then A( x ) = πx2/3 . Therefore
Z 1 Z 1
V= A( x )dx = πx2/3 dx = 3πx5/3 /5|10 = 3π/5.
0 0

1.6.2 Work done by a force


It is known from elementary physics that if a body moves in a
straight line under a constant force F , then the work done by the
force is given by Work = Force × Distance = Fd where d is the
distance moved. In the case of variable forces this simple formula
must be adjusted, as shown below.
Suppose that an object moves along the x-axis from a to b under
a force f which is a function of x. Consider a partition P : a = x0 <
x1 < . . . < xn = b of the interval [ a, b] into subintervals [ x j−1 , x j ]
and choose an arbitrary c j ∈ [ x j−1 , x j ], j = 1, 2, . . . , n. Then the
work Wj done in moving the object from x j−1 to x j is approximately
f (c j )∆x j , and therefore the total work done is W = ∑nj=1 Wj ≈
∑nj=1 f (c j )∆x j .
This is a Riemann sum for the function f , so assuming that the
function f is sufficiently regular and letting || P|| → 0 gives
Z b
W= f ( x )dx.
a
Example 1.24. A bucket is lifted from the ground into the air by pulling
in 10m of rope at a constant speed. If the weight of the bucket is 5N and
that of the rope is 0.8N/m, how much work is done lifting the bucket and
the rope?
Solution: The work done on the bucket is W1 = 5 × 10 = 50 Joules.
Consider a vertical x-axis pointing upwards with 0 at ground level. When
22 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

the bucket is at level x, the length of the part of the rope that remains to be
lifted is 10 − x, so its weight is f ( x ) = (10 − x )0.8 N. Thus, the work
done to lift the rope is
Z 10
W2 = 0.8(10 − x )dx = 40 Joules.
0
Hence the total work done is W = W1 + W2 = 50 + 40 = 90 J.

Example 1.25. A tank has the shape of an inverted cone with height 10m
and base radius 4m. It is filled with water to a height of 8m. It is required
to pump all the water out of the tank by inserting a hose through the top
to which a pump is attached. Calculate the work done against gravity to
achieve this. (Assume that the gravity constant is g = 9.8m/s2 and the
density (mass per unit volume) of the water is 1000kg/m3 ). 4

Solution: Choose a vertical x-axis pointing upwards with 0 at the vertex


of the cone. Consider a partition P : 0 = x0 < x1 < . . . < xn = 8
of the interval [0, 8] into n subintervals and choose an arbitrary c j in the
jth subinterval. Let Vj be the volume of the layer of water in the tank
lying between levels x j−1 and x j . If r j is the radius of the horizontal cross-
section of the tank at level c j , then Vj ≈ πr2j ∆x j . Using similar triangles, 10

r j /4 = c j /10. 8

The force required to raise this volume against the force of gravity is
then given by (1000Vj ) g = 10(9.8)16πc2j ∆x j .
The work done in raising the jth layer to the top of the tank is then
Wj ≈ (1568πc2j ∆x j )(10 − c j ), and so the total work done is

Z 8
2048
W= 1568πx2 (10 − x )dx = 1568π ≈ 3.36 × 106 J.
0 3
Note: In some problems of this kind the density will be given as
weight per unit volume (i.e. the gravity constant has already been
taken into account); or SI2 units (Newtons, Joules) are not being 2
International System of Units (ab-
used. breviated SI from French: Système
international d’unités

1.6.3 Length of curves


f(xj) Lj
Let f be a continuously differentiable function on an interval [ a, b];
i.e. f 0 ( x ) is continuous on [ a, b]. The problem is to calculate the f(xj−1)
} f(xj) − f(xj−1)

length of a curve C = {( x, f ( x )) : a ≤ x ≤ b} in R2 . xj − xj−1


Consider a partition P of the interval as in previous examples.
Let L j be the length of the part of the curve above the subinterval xj−1 xj
[ x j−1 , x j ]. Then
2 2
q L j = (xj − xj−1) + ( f(xj) − f(xj−1)) 2
L j ≈ ( x j − x j−1 )2 + ( f ( x j ) − f ( x j−1 ))2 .

By the Mean Value Theorem there exists c j ∈ ( x j−1 , x j ) such that Recall the Mean Value Theorem:
f ( x j ) − f ( x j−1 ) = f 0 (c j )( x j − x j−1 ). Let f be continuous on an interval
q [ a, b] and differentiable on ( a, b),
Therefore L j ≈ 1 + ( f 0 (c j ))2 ∆x j . Hence for the total length of then there exists c ∈ ( a, b) such that
f (b)− f ( a)
q f 0 (c) = b− a .
the curve we obtain L ≈ ∑nj=1 (1 + ( f 0 (c j ))2 ∆x j .
p
This is a Riemann sum for the function 1 + ( f 0 ( x ))2 , so letting
Rbp
|| P|| → 0 gives L = a 1 + ( f 0 ( x ))2 dx.
mathematical methods 2 23

Example 1.26. Find the length of the curve y = ( x2 + 2)3/2 /3 for


0 ≤ x ≤ 1.
R1p
Solution: y0 ( x ) = x ( x2 + 2)1/2 so L = 0 1 + x2 ( x2 + 2) dx =
R1p R1
0 (1 + x2 )2 dx = 0 (1 + x2 ) dx = [ x + x3 /3]10 = 4/3 .

1.7 Improper integrals

In this section we describe two types of improper integrals.

1.7.1 Improper integrals over infinite intervals (Type I improper in-


tegrals)

Definition 1.27. (Type I improper integrals)


(a) Let the function f be defined on [ a, ∞) for some a ∈ R and integrable
over [ a, t] for any t > a. The improper integral of f over [ a, ∞) is defined
to be Z ∞ Z t
f ( x )dx = lim f ( x )dx.
a t→∞ a

If the limit exists then the improper integral is called convergent. If the
limit does not exist then the improper integral is called divergent.
(b) Similarly
Z b Z b
f ( x )dx = lim f ( x )dx.
−∞ t→−∞ t

(c) Finally, suppose f ( x ) is defined for all x ∈ R. Consider an arbitrary


c ∈ R and define
Z ∞ Z c Z ∞
f ( x )dx = f ( x )dx + f ( x )dx.
−∞ −∞ c

The improper integral is convergent if and only if for some c ∈ R both


integrals on the right-hand-side are convergent.

Remark 1.28. It can be shown that the choice of c is not important; i.e.
if for one particular choice of c the integrals on the right-hand-side are
convergent, then the same is true for any other choice of c and the sum of
the two integrals is always the same3 . 3
Try to prove this.
Z ∞
1
Example 1.29. Find the improper integral dx if it is convergent
1 x3
or show that it is divergent.
Solution: For any t > 1,
Z t  t  
1 1 1 1 1
3
dx = − 2 = − 2+ →
1 x 2x 1 2t 2 2
as t → ∞. The improper integral is convergent and its value is 1/2.
(This example shows that the area A under the graph of f ( x ) = 1/x3
and above the interval [1, ∞) is finite even though the ‘boundaries’ of the
area are infinitely long.)
24 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

Z ∞
1
Example 1.30. Find the improper integral dx if it is convergent or
1 x
show that it is divergent.
Solution: For any t > 1 we have
Z t
dx
= [ln x ]1t = ln t → ∞
1 x
R∞
as t → ∞. Hence the integral 1 1/x dx is divergent. (The area under
the graph of f ( x ) = 1/x and above the interval [1, ∞) is unbounded.)
Example 1.31. Find all constants p ∈ R such that the improper integral
Z ∞
dx
p is convergent.
1 x
Solution: For p = 1 the previous example shows that the integral is
divergent. Suppose p 6= 1. Then
 − p +1  t
t 1− p − 1
Z t
dx x
p = = .
1 x −p + 1 1 1− p
When 1 − p < 0, limt→∞ t1− p = 0 so the integral is convergent. When
1 − p > 0, limt→∞ t1− p = ∞ and therefore the integral is divergent.
Hence the integral is divergent for p ≤ 1 and otherwise convergent.

1.7.2 Improper integrals over finite intervals (Type II improper inte-


grals)
Sometimes we want to integrate a function over an interval, even
though the function is not defined at some points in the interval.

Definition 1.32. (Type II improper integrals)


(a) Assume that for some a < b the function f is defined and continuous
on [ a, b) however it has some kind of singularity at b, e.g. f ( x ) → ∞ or
−∞ as x → b− . Define the improper integral of f over [ a, b] by
Z b Z t
f ( x )dx = lim f ( x )dx.
a t→b− a

If the limit exists, then the improper integral is called convergent, other-
wise it is divergent.
(b) In a similar way, if f is continuous on ( a, b] however it has some
kind of singularity at a, e.g. f ( x ) → ∞ or −∞ as x → a+ . We define
Z b Z b
f ( x )dx = lim f ( x )dx.
a t→ a+ t

(c) If for some c ∈ ( a, b), f is continuous on each of the intervals [ a, c)


and (c, b], however it has some kind of singularity at c, define
Z b Z c Z b
f ( x )dx = f ( x )dx + f ( x )dx
a a c
Z t Z b
= lim f ( x )dx + lim f ( x )dx.
t→c− a t→c+ t

The improper integral on the left-hand side is convergent if and only if


both improper integrals on the right-hand-side are convergent.
mathematical methods 2 25

R1 √
Example 1.33. Consider the integral 0 1/ x dx. This is an improper
√ √
integral, since 1/ x is not defined at 0 and 1/ x → ∞ as x → 0.
For 0 < t < 1 we have
Z 1 √ √ √
1/ x dx = [2 x ]1t = 2 − 2 t → 2
t
as t → 0+ . So the improper integral is convergent and its value is 2.
Z 2
1
Example 1.34. Consider the integral dx. It is improper, since
0 x−1
f ( x ) = 1/( x − 1) is not defined at x = 1 and is unbounded near x = 1.
However, f ( x ) is continuous on [0, 1) ∪ (1, 2].
Therefore
Z 2 Z 1 Z 2
dx dx dx
= + ,
0 x−1 0 x−1 1 x−1
and the integral on the left-hand-side is convergent if and only if both
integrals on the right-hand-side are convergent.
Consider Z 1 Z t
1 1
dx = lim dx.
0 x − 1 t → 1 − 0 x − 1
Given 0 < t < 1, we have
Z t
1/( x − 1)dx = [ln | x − 1|]0t = ln |t − 1| − 0 = ln(1 − t) → −∞
0
as t → 1− .
Hence this integral is divergent and therefore the original integral is
also divergent4 . 4
Note that if one integral is shown
Z 1 to be divergent we do not need to
Exercise 1.7.1. For what values of p is x p dx improper, and in that check if the other one is convergent or
0 not, we already know the answer: the
case for what values of p is the integral divergent? Compare with Example original improper integral is divergent.
1.31.

1.8 Quadrature

There are some functions for which we do not know an anti-


derivative, so we cannot apply the FTC. Also sometimes, we only
know some values of f but not the whole function f precisely. In
these cases, we might want to approximate an integral using a
weighted sum of function values, as computers can do this much
faster than we can look up a table of integrals. Computing defi-
nite integrals by sums is called quadrature. Computing indefinite
integrals or solutions to ODEs is called numerical integration. Three
basic ways for quadrature are called the trapezoidal, mid-point and
Simpson’s rules. There are other more sophisticated and accurate
methods.
For a small interval ( a, a + h) say, the following table describes
R a+h
the three approximation methods for a f ( x ) dx and the corre-
sponding error made by considering this approximation.
R a+h
Rule approximation for a f ( x ) dx error, where c ∈ ( a, a + h)
h h3 00

trapezoidal rule 2 f ( a) + f ( a + h) − 12 f (c)
h3 00
mid-point rule h f ( a + h/2) 24 f ( c )
h h5 (4)

Simpson’s rule 6 f ( a) + 4 f ( a + h/2) + f ( a + h) − 2880 f (c)
26 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

If we have a large interval [ a, b], we can divide it into N subinter-


vals of length h = (b − a)/N and then sum the approximations on
the subintervals. This gives the composite quadrature rules, based on
the function values f ( xi ), where xi = a + ih, i = 0, 1, . . . , N, and/or
the mid-points, xi + h/2.
Rb
Rule approximation
 for a f ( x ) dx  error, where c ∈ ( a, b)
N − 3
1
composite trapezoidal rule 2h f ( a) + 2 ∑i=1 f ( xi ) + f (b) − Nh 00
12 f ( c )
 
Nh3 00
composite mid-point rule h ∑iN=−0 1 f ( xi + h/2) 24 f ( c )
 
h N −1 N −1 Nh5 (4)
composite Simpson’s rule 6 f ( a ) + 2 ∑ i =1 f ( x i ) + 4 ∑ i =0 f ( x i + h/2 ) + f ( b ) − 2880 f (c)

R1
Example 1.35. Approximate 0 1+1x2 dx using the composite trapezoidal
rule for N = 4. Deduce an interval for the true value of the integral.
1
Solution: Let f ( x ) = 1+ x 2
. We have h = 1/4. So
Z 1
1 1
dx ≈ ( f (0) + 2 f (1/4) + 2 f (1/2) + 2 f (3/4) + f (1))
0 1 + x2 8
1
= (1 + 2 · 16/17 + 2 · 4/5 + 2 · 16/25 + 1/2)
8
= 5323/6800 ≈ 0.78279 . . .
(6x2 −2)
To estimate the error, we need f 00 ( x ) = (1+ x 2 )3
, which is an increasing
function, so for c ∈ (0, 1), −2 < f 00 (c) < 1/2. Since the error is
3
− Nh 00
12 f ( c ) for some c ∈ (0, 1), we have

1 1 1
− · < error < − · (−2).
16 · 12 2 16 · 12
Hence the true value is in (0.78279 − 0.02605, 0.78279 + 0.01042) =
(0.75674, 0.79321). Exercise: evaluate the integral directly and check the
value is in this interval.

Example 1.36. Show that Simpson’s rule on [0, h] gives exact values of
the integral for all cubic polynomials.
Solution: Let an arbitrary cubic be p( x ) = c0 + c1 x + c2 x2 + c3 x3 . Then
Rh
the definite integral is 0 p( x )dx = c0 h + c1 h2 /2 + c2 h3 /3 + c3 h4 /4.
The Simpson rule approximation is 6h ( p(0) + 4p(h/2) + p(h))
 2  3 !  !
h h h h 2 3

= c0 + 4 c0 + c1 + c2 + c3 + c0 + c1 h + c2 h + c3 h
6 2 2 2
 
h 3
= 6c0 + 3c1 h + 2c2 h + c3 h3
2
= c0 h + c1 h2 /2 + c2 h3 /3 + c3 h4 /4.
6 2
The error term is actually proportional to the fourth derivative of the inte-
grand at some unknown point, so using the fact that the fourth derivative
of a cubic polynomial is zero everywhere, we also get the result.

Exercise 1.8.1. Use the first Taylor polynomial of f about h/2 to show
1 3 00
that the error for the mid-point rule on [0, h] is 24 h f (c), for some un-
known point c ∈ (0, h). Note that using the Taylor polynomial approxima-
tion assumes that certain derivatives exist everywhere in the interval.
2
Double and triple integrals

In this chapter we provide an introduction to the theory and ap-


plications of Vector Calculus. This generalises some familiar ideas
from calculus, such as finding gradients, integrals, etc. to higher
dimensions. This is hugely important for tackling a wide range of
problems in physics and engineering. Usually in the course, we will
concentrate on two- or three-dimensional settings, as these are most
appropriate from the point of view of applications.
Recall that in R2 the area of a region between the x-axis and the
graph of a one-variable funtion is given by the definite integral of
the function. In this chapter, we are looking at a similar concept for
functions of two or three variables.

2.1 Double integrals

2.1.1 Double integral on a rectangular region


In MATH1001 we introduced scalar valued functions of two vari- z
z = f(x,y)
ables (c.f. Chap. 6 in the notes). Let us assume we have such a
function z = f ( x, y), defined in some region R ∈ R2 . To consider a
concrete example, we assume that the function describes the height
of an object above the xy–plane. We would like to determine the
volume of the solid bounded above by the surface z = f ( x, y) and y
a
lying directly above some region R. Let us first consider the case
that this region R is a rectangular region in the xy-plane, that is, b
c d
R = {( x, y)| a ≤ x ≤ b, c ≤ y ≤ d}for some a, b, c, d ∈ R. Note that R x
z
is closed and bounded.
To approximate the volume V of the solid above R and bounded
above by the surface z = f ( x, y) we shall draw a connection be-
tween Riemann sums and integrals. Partition the region R as fol-
lows. Given an arbitrary integer N ≥ 1, consider a partition of the y
x–interval [ a, b] a

b
Px : a = x0 < x1 < x2 < . . . < x N −1 < x N = b. c d
x
Volume under z = f ( x, y).
Similarly, given another arbitrary integer M ≥ 1, consider a parti-
tion of the y–interval [c, d]

Py : c = y0 < y1 < y2 < . . . < y M−1 < y M = d.


28 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

Together, these two partitions provide a partition of R into rect-


angles Ri,j = {( x, y)| xi−1 ≤ x ≤ xi , y j−1 ≤ y ≤ y j } with sides of size y
∆ x i = x i − xi−1
∆xi = xi − xi−1 and ∆y j = y j − y j−1 . Inside each rectangle Ri,j we yM = d

pick a point ( xi∗ , y∗j ) (with xi−1 ≤ xi∗ ≤ xi and y j−1 ≤ y∗j ≤ y j ), and yj
∆ yj = yj − yj−1
the value of z = f ( xi∗ , y∗j ) describes approximately the height of the
surface above the xy-plane at the point ( xi∗ , y∗j ) and hence f ( xi∗ , y∗j ) y0 = c

multiplied by ∆xi · ∆y j gives us an approximation of the volume of x


x0= a xi xN = b
the column below the surface and above Ri,j . Obviously the sum Partitioning of the area R.
over all i, j approximates the volume V of the solid.

z Figure 2.1: Volume element at the


point ( xi , yi ) of height z = f ( xi , yi ).

f(xi ,yj )
y
a
xi
b
c yj d
x

Definition 2.1. (Double integral over a rectangular region)


A function f defined over a rectangular region R is integrable if

M N

M,N →∞
∑ ∑ f (xi∗ , y∗j )∆xi ∆y j
lim
∆x,∆y→0 j=1 i =1

exists (over all partitions of R and choices of ( xi∗ , y∗j )).


We call this limit the double integral of f over R, and denote it by
ZZ
f ( x, y)dA.
R

If f ( x, y) ≥ 0 for all (x,y) in R, then the double integral is the volume


of the solid bounded by z = f ( x, y) and the xy-plane over R.

M N
VPx ,Py = ∑ ∑ f (xi∗ , y∗j )(∆xi ∆y j ) height of column · base area
j =1 i =1
!
M N
= ∑ ∑ f ( xi∗ , y∗j )∆xi ∆y j Take limit as N → ∞
j =1 i =1
M Z b

The wall is the intersection of the solid
→ ∑ a
f ( x, y∗j ) dx ∆y j area of wall at y = y∗j · thickness with the plane y = y∗j
j =1
Z d Z b 
→V= f ( x, y) dx dy Taking limit as M → ∞
c a
mathematical methods 2 29

Notice that

∑ ∑ f (xi∗ , y∗j )∆xi ∆y j = ∑ ∑ f (xi∗ , y∗j )∆y j ∆xi ,


j i i j

and that in the calculation above we could have first taken the limit
M → ∞ and then N → ∞. We formalise this idea with the fol-
lowing theorem, which allows us to evaluate double integrals over
rectangular regions using well-known methods of integration with
respect to a single variable and shows that we can change the order
of integration:

Theorem 2.2. (Fubini’s Theorem for rectangular regions) Let f ( x, y)


be a continuous function on the rectangular region R defined by
R = {( x, y) | a ≤ x ≤ b and c ≤ y ≤ d}. Then
ZZ Z d Z b 
f ( x, y)dA = f ( x, y)dx dy
R c a
Z b Z d 
= f ( x, y)dy dx.
a c

Example 2.3. Let f ( x, y) = x2 + xy and

R = {( x, y) : 1 ≤ x ≤ 2, −1 ≤ y ≤ 1}.

Then
ZZ Z 1 Z 2
f ( x, y)dA = ( x2 + xy) dxdy
R y=−1 x =1
Z 1 Z 2 
= ( x2 + xy) dx dy
−1 1

1 3 1 2 2
Z 1  
= x + x y dy
−1 3 2 1
Z 1
7 3
= ( + y) dy
−1 3 2
 1
7 3
= y + y2
3 4 −1
   
7 3 2 7 3 2 14
= (1) + (1) − (−1) + (−1) = .
3 4 3 4 3
Note that attention needs to be paid to the order of the dx, dy
relative to the order of the integral symbols!

2.1.2 Double integral over a bounded region


By definition, a region in R2 is a subset R of the plane R2 whose
boundary ∂R is a finite union of continuously differentiable curves.
Having considered double integrals on rectangular regions
leads us to the general definition of double integrals over arbitrary
bounded regions in R2 . We achieve this by embedding R as a sub-
region into a rectangular region R̂. We use partitions Px , Py of R̂,
30 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

as in the previous section. The difference is that now we only sum


over the Ri,j s that are completely contained in R. Obviously, the
small rectangles that are outside R do not contribute, and the con-
tributions of those that overlap the boundary are negligible when
the subrectangles of the partition are small.

Definition 2.4. (Double integral)


Let f ( x, y) be a scalar function which is continuous on a bounded region
R ⊆ R2 . The double integral of f over the region R is
ZZ

R
f ( x, y)dA = lim
M,N →∞
∑ f ( xi∗ , y∗j )(∆xi ∆y j ),
∆x,∆y→0 Ri,j ⊂ R

provided that this limit exists.


If f ( x, y) ≥ 0 for all (x,y) in R, then the double integral is the volume
of the solid bounded by z = f ( x, y) and the xy-plane over R.

Remark about Definition 2.4:


Let Rb be an arbitrary rectangular region in R2 containing the
region R in its interior. Consider the extension fˆ of the function
f over the rectangular region R b defined by fˆ( x, y) = f ( x, y) if
( x, y) ∈ R and fˆ( x, y) = 0 if ( x, y) ∈
/ R.
It can be shown that
ZZ ZZ
f ( x, y)dA = fˆ( x, y) dA,
R R
b

where the right-hand-side is the double integral of fˆ over the rect-


angular region R
b defined as in the previous sub-section.

Properties of double integrals over bounded regions:


ZZ ZZ
1. c f dA = c f dA for any constant c ∈ R;
R R
ZZ ZZ ZZ
2. ( f + g) dA = f dA + g dA;
R R R
ZZ ZZ ZZ
3. f dA = f dA + f dA
R1 ∪ R2 R1 R2
for regions R1 and R2 with disjoint interiors;
ZZ
4. If f ( x, y) ≥ 0 on R, then f dA is the volume of the solid
R
under the graph of f and above the region R in the x, y-
plane;
ZZ
5. More generally, f dA = V+ − V− , where V+ is the
R
volume of the part of the solid under the graph of f and
above the x, y-plane, while V− is the volume of the part of
the solid above the graph of f and below the x, y-plane.
mathematical methods 2 31

Theorem 2.5. (Existence of Double integral over a bounded


region) Let f ( x, y) be a bounded real-valued function defined over
a bounded region R in R2 , and let f be continuous over R except
possibly over a finite union of continuously differentiable curves.
Then the double integral of f over R exists.

Notice that if the function we integrate is the constant function


f ( x, y) = 1, then the volume of the solid over R and under f is
equal to the area of the region R.

Theorem 2.6. ZZ
1 dA = Area( R).
R

Evaluations of double integrals over bounded regions

First consider regions between the graphs of two functions of x,


sometimes called regions of type I:

R = {( x, y) : a ≤ x ≤ b , g( x ) ≤ y ≤ h( x )},

where g( x ) and h( x ) are continuous functions of x ∈ [ a, b] and


g( x ) ≤ h( x ) for all x ∈ [ a, b].

y y = h( x )

y = g( x )
c

a b x

Let f ( x, y) be continuous on R.
Take an arbitrary rectangle R b = [ a, b] × [c, d] containing R in its
interior and consider the extension fˆ of f which is 0 outside R.
Then, using the Remark about Definition 2.4 made earlier and
Fubini’s Theorem for rectangular regions, we get
ZZ ZZ Z b Z d 
f ( x, y)dxdy = ˆ
f ( x, y)dxdy = ˆ
f ( x, y)dy dx
R R
b a c
Z b Z g( x ) Z h( x )
= [ fˆ( x, y) dy + fˆ( x, y)dy
a c g( x )
Z d
+ fˆ( x, y)dy ] dx
h( x )
Z b Z h( x ) 
= f ( x, y)dy dx.
a g( x )
32 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

This is so since for any given x ∈ [ a, b] we have fˆ( x, y) = 0 for


y < g( x ) and also for y > h( x ), while for y ∈ [ g( x ), h( x )] we have
fˆ( x, y) = f ( x, y).
Hence for regions R of type I we have
ZZ Z b Z h( x ) 
f ( x, y)dxdy = f ( x, y)dy dx.
R a g( x )

Next consider regions between the graphs of two functions of y,


sometimes called regions of type II:

R = {( x, y) : c ≤ y ≤ d , G (y) ≤ x ≤ H (y)},

where G (y) and Hyx ) are continuous functions of y ∈ [c, d] and


G (y) ≤ H (y) for all y ∈ [c, d].
As in the previous case one shows that if f ( x, y) is continuous on
R, then
ZZ  Z d Z H (y)

f ( x, y) dxdy = f ( x, y) dx dy .
R c G (y)

Example 2.7. Let R be the triangle bounded by the lines x = 0, y = 1 Since this triangle has half the area of a
and y = x. Compute the area of R using a double integral. unit–square, its area is A = 1/2.

y
Solution: The most important tool that we have is that we are able to y=1

sketch R. Let us assume that we want to integrate first with respect to


x. For a fixed y, we see that x is in between 0 and y. Moreover y is in x=0
between 0 and 1. Therefore we get

R = {( x, y) | 0 ≤ y ≤ 1, 0 ≤ x ≤ y}, y=x or x=y

and so the area is: x


ZZ Z 1 Z y Triangular region.
A= 1 dA = 1 dx dy
R y =0 x =0
1
y2
Z 1 Z 1 
y 1
= [ x ]0 dy = y dy = = .
0 0 2 0 2

Similarly by changing the order of integration we rewrite

R = {( x, y) | 0 ≤ x ≤ 1, x ≤ y ≤ 1}

and deduce:
ZZ Z 1 Z 1
A= 1 dA = 1 dy dx
R x =0 y = x
1
x2
Z 1 Z 1 
1
= [y]1x dx = (1 − x ) dx = x − = .
0 0 2 0 2

Note that y = x is the lower bound of the first integration.

The above discussion yields the following more general version


of Fubini’s Theorem.
mathematical methods 2 33

Theorem 2.8. (Fubini’s Theorem for regions bounded by functions) Let


f ( x, y) be a continuous function1 over a region R. 1
It is enough to assume f is bounded
on R and continuous over R except
1. If R is defined by possibly over a finite union of continu-
ously differentiable curves.
R = {( x, y) | a ≤ x ≤ b and g( x ) ≤ y ≤ h( x )},

where g( x ) and h( x ) are continuous on [ a, b], then


ZZ Z b  Z h( x ) 
f ( x, y)dA = f ( x, y)dy dx.
R a g( x )

2. If R is defined by

R = {( x, y) | c ≤ y ≤ d and H (y) ≤ x ≤ G (y)},

where G (y) and H (y) are continuous on [c, d], then


ZZ Z d  Z G (y) 
f ( x, y)dA = f ( x, y)dx dy.
R c H (y)

As seen in our example above, if g( x ) and h( x ) have inverse


functions on the interval [ a, b] and their range is the interval [c, d],
then we may choose G to be the inverse function of g and H to be
the inverse function of h.
We can summarise integration on non–rectangular regions as
follows:
• identify the region by drawing a sketch.
• identify where the functions intersect.
• identify which of the functions has a larger value for a given x
between the points of intersection.
• identify which of the inverse functions is larger for a given y
between the intersections.
• the inner (first) integral will have bounds that are functions of
the other variable.
• the inner integral can be understood as the length of the area
increment as a function of the other variable.
• the outer (second) integral will add up the area increments be-
tween the two intersections and therefore only have bounds that
are scalars!

Remark. Every region R in the plane can be represented as a union


of finitely many regions R1 , R2 , . . . , Rk with disjoint interiors such
that each Ri is either type I or type II region (or both). For any
bounded function f ( x, y) on R, Property 3 implies
ZZ ZZ ZZ ZZ
f dA = f dA + f dA + . . . + f dA.
R R1 R2 Rk

For each i, one can then evaluate the double integral of f over Ri
using Fubini’s Theorem.

Often, we might be given a region which is described by curves,


which are not functions. For example, consider a circle of radius
34 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

r around the origin. Then we might consider the region R of all


points lying in the circle, i.e. R = {( x, y) | x2 + y2 ≤ r2 }. The curve
x2 + y2 = r2 is clearly not a function. However, we can identify
two functions which describe the points in the region, namely we
√ √
consider g( x ) = − r2 − x2 and h( x ) = r2 − x2 and hence
R = {( x, y) | −r ≤ x ≤ r and g( x ) ≤ y ≤ h( x )}.
The following example also uses curves, not just functions, to
bound the region.
Example 2.9. Let R be the region bounded by y = x3 and x = y2 . Find
y y = x3
its area twice, using both orders of integration.
Solution: Obviously these curves intersect at (0, 0) and (1, 1). We re-
strict our attention to only some points on the curves to obtain functions,

namely define functions g( x ) = x3 and h( x ) = x. For 0 ≤ x ≤ 1 we
x
see that g( x ) ≤ h( x ). Therefore:
Z 1 Z h( x ) Z 1 Z √x
5
A= 1 dy dx = 1 dy dx = x=y2
x =0 y = g ( x ) x =0 y= x3 12

Let G (y) = g−1 (y) = 3 y and H (y) = h−1 (y) = y2 . Then we find:
Z 1 Z G (y) Z 1 Z y1/3
5
A= 1 dx dy == 1 dx dy = .
y =0 x = H (y) y =0 x = y2 12
Choosing f ( x, y) = 1 yields the area of the region R, but we can
also use double integrals to evaluate any function f ( x, y) on rect-
angular or non–rectangular regions. In many applications f ( x, y)
might be understood as a function describing the density per unit
area, and in this case the integral represents the mass of the area
(for instance).
Changing the order of integration might make calculations easier
for a given f ( x, y).
Example 2.10. Let R be the region defined in the previous example.
Integrate f ( x, y) = xy on R.
Solution:

Z 1 Z √x Z 1 2 x
y
xy dy dx = x dx =
x =0 y = x 3 0 2 y= x3
Z 1
1  5
= x2 − x7 dx = .
2 0 48
Or
Z 1 Z y1/3 Z 1  2 y1/3
x
xy dx dy = y dx =
y =0 x = y2 0 2 x = y2
Z 1
1  5
= y5/3 − y5 dx = .
2 0 48
While both calculations are easy, the first double integral does not contain
fractions in the exponents and is therefore simpler.
When the functions bounding the region do not have an inverse
function which is one-to-one, we may have to split the region up
into two or more sub–regions.
mathematical methods 2 35

Example 2.11. Let R be the region bounded by the curves y = x − 2 and


x = y2 . Write the integral of f over R using both orders of integration.
Solution: The two curves intersect at the points (1, −1) and (4, 2). As
we can see from the picture in the margin, we can calculate the integral of
f over R by: y
y = x−2
Z 2 Z y +2 (4, 2)
I= f ( x, y) dx dy.
y=−1 x = y2

To change the order of integration, we observe that the region R can be R2


R1
split into two regions, the red region R1 and the green region R2 , namely, x
(1, −1)
√ √
R1 = {( x, y) | 0 ≤ x ≤ 1 and − x ≤ y ≤ x } x = y2

and

R2 = {( x, y) | 1 ≤ x ≤ 4 and x − 2 ≤ y ≤ x }.
The definition of the double integral by Riemann sums allows us to split
the integral over R into two different regions. Therefore we get by Fubini’s
more general theorem:
ZZ ZZ ZZ
I= f ( x, y) dy dx = f ( x, y) dy dx + f ( x, y) dy dx
R R1 R2
Z 1 Z √x Z 4 Z √x
= √ f ( x, y) dy dx + f ( x, y) dy dx
x =0 y=− x x =1 y = x −2

This example illustrates the fact that we can avoid a lot of calcu-
lations by being smart when choosing our order of integration.

2.2 Triple integrals

With the help of double integrals we can calculate the volume of a


a solid above a surface g( x, y) and beneath another surface h( x, y).
For the integrals we considered in the previous section, we simply
chose the function g( x, y) = 0. Obviously double integrals do not
depend on this particular choice and we see that the volume of the z
solid thus described is:
ZZ ZZ
V= h( x, y) dA − g( x, y) dA
Z ZR R

= (h( x, y) − g( x, y)) dA.


R y
a
Suppose R = {( x, y)| a ≤ x ≤ b, c( x ) ≤ y ≤ d( x )}. Then
b
Z b Z d( x ) c d
x
V= (h( x, y) − g( x, y)) dy dx. Volume of solid bounded by two surfaces.
x=a y=c( x )

The last term is the difference between z = h( x, y) and z = g( x, y).


This term would also arise when evaluating the triple integral:
Z b Z d( x) Z h( x,y)
V= 1 dz dy dx
x=a y=c( x ) z= g( x,y)

We now consider bounded regions R ⊆ R3 , which we call solids.


36 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

More precisely, by a solid we mean a bounded subset R in R3


whose boundary ∂R is a finite union of continuously differentiable
surfaces. A continuously differentiable surface in R3 is by defini-
tion a level surface S : f ( x, y, z) = k of a continuously differentiable
function f with ∇ f 6= 0 on S.
The idea of dividing a region into small rectangular subre-
gions generalises in R3 into dividing solids into small right par-
allelepipeds. This idea leads us to the definition of a triple integral:

Definition 2.12. (Triple integral)


Let f ( x, y, z) be a scalar function which is continuous on a solid R ⊆ R3 .
The triple integral of f over R is
ZZZ L M N

R
f ( x, y, z)dV =
L,M,N →∞
lim ∑ ∑ ∑ f (xi∗ , y∗j , z∗k ) ∆xi ∆y j ∆zk ,
∆x,∆y,∆z→0 k =1 j=1 i =1

provided that this limit exists.

Again we used Riemann sums to define triple integrals. We


view triple integrals as the integral of f over some solid R that is
bounded by functions. By the similarity between the definition of
double and triple integrals we see immediately that:

ZZZ ZZZ
1. k f dV = k f dV for a constant, k
R R
ZZZ ZZZ ZZZ
2. ( f 1 + f 2 ) dV = f 1 dV + f 2 dV
R R R
ZZZ ZZZ ZZZ
3. f dV = f dV + f dV
R1 ∪ R2 R1 R2
for solids R1 and R2 with disjoint interiors.

Theorem 2.13. (Existence of triple integrals) Let f be a bounded


function on a solid R in R3 such that f is continuous on R, except
possibly on a finite union ofZ Zcontinuously
Z differentiable surfaces. Then
f is integrable over R, i.e. f dV exists.
R

Remark: Geometrical Meaning of triple integral. While this goes


beyond the scope of this unit, it should be noted that if f ( x, y, z) ≥
RRR
0 on some solid R and f is continuous on R, then R f dV repre-
sents the 4D volume of the 4D solid "below" the graph of f in the
space R4 . More precisely,
ZZZ
f dV = Vol4 (S),
R
mathematical methods 2 37

where

S = {( x, y, z, t) ∈ R4 : ( x, y, z) ∈ R , 0 ≤ t ≤ f ( x, y, z)}.
RRR
In general if f is not necessarily non-negative, then R f dV rep-
resents an oriented 4D volume (which may be positive, negative or
zero).

There is however an important special case when the triple inte-


gral represents a 3D volume. Namely, if the function we integrate is
the constant function f ( x, y, z) = 1, then the integral of f over R is
simply the volume of R.

ZZZ
Theorem 2.14. 1 dV = Vol( R).
R

What concerns evaluations of triple integrals, first we need to


know how to do this over rectangular boxes. The following is easy
to derive from the definition.

Theorem 2.15. (Fubini’s Theorem) Let f ( x, y, z) be a bounded


function on a rectangular box

B = [ a, b] × [c, d] × [ p, q]
= {( x, y, z) ∈ R3 : a ≤ x ≤ b, c ≤ y ≤ d, p ≤ z ≤ q}

such that f is continuous on B except possibly on a finite union of


continuously differentiable surfaces. Then
ZZZ Z b Z d Z q  
f dV = f ( x, y, z)dz dy dx
B a c p

Z b Z q Z d  
= f ( x, y, z)dy dz dx, etc.
a p c
ZZZ
That is, all 6 possible iterated integrals are equal to f dV.
B

Important Remark:
Z bZ dZ q Z b Z d Z q  
f ( x, y, z) dz dy dx = f ( x, y, z) dz dy dx.
a c p a c p

That is:

• the integration with respect to z is from p to q, then

• the integration with respect to y is from c to d, and finally

• the integration with respect to x is from a to b.


38 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

All other arrangements of the variables (as in Fubini’s Theorem)


are treated according to the same rule.
Generally speaking, evaluations of triple integrals are done either
by

• a reduction to a double integral (and then using the procedures


from Sects. 2.1.1 and 2.1.2 above), or

• by using a change of coordinates (see Ch. 3 below),

or, sometimes, by using both.

Reduction to a double integral

The reduction to a double integral is particularly simple when


R is the solid between the graphs of two functions of two of the
variables x, y, z. E.g. assume

R = {( x, y, z) : ( x, y) ∈ D , g( x, y) ≤ z ≤ h( x, y)},

where D is some region in the x, y-plane and g( x, y) ≤ h( x, y) are


continuous functions on D.
If f ( x, y, z) is a continuous function on R, then
ZZZ Z Z Z h( x,y) 
f dV = f ( x, y, z) dz dxdy.
R D g( x,y)

This can be derived using an argument similar to that for double in-
tegrals (see "Evaluations of double integrals over bounded regions"
above).
Similarly, if

R = {( x, y, z) : ( x, z) ∈ D , g( x, z) ≤ y ≤ h( x, z)},

where D is some region in the x, z-plane and g( x, z) ≤ h( x, z) are


continuous functions on D, then
ZZZ Z Z Z h( x,z) 
f dV = f ( x, y, z) dy dxdz,
R D g( x,z)

etc.

Sometimes (though not very often) it is possible to write down a


triple integral immediately as a repeated integral, i.e. a particular
sequence of single integrals of the form:
Z b Z d( x) Z h( x,y)
f ( x, y, z) dz dy dx.
x=a y=c( x ) z= g( x,y)
| {z }
integrate on line
| {z }
integrate over area
| {z }
integrate over solid

Finding the right bounds can be challenging or very simple. We


are going to discuss methods for this in the following two exam-
ples.
mathematical methods 2 39

Example 2.16. Let f ( x, y, z) be a continuous function on the region R,


bounded by x2 + y2 ≤ r2 and 0 ≤ z ≤ h (where r and h are positive
constants). Obviously the solid R is a cylinder with radius r and height h.
Moreover we see that the height (bounds for z) does not depend2 on x or 2
This means that we could integrate
y. Therefore we may want to integrate first over either x or y. The given with respect to z any time we want
√ since this integral will only introduce a
inequality yields bounds: y = ± r2 − x2 and x = ±r. Therefore we can multiplicative scalar (here h). But often
write the triple integral as: such integrals are done either first or
last.
ZZZ Z h Z r Z √r 2 − x 2
f ( x, y, z) dV = √ f ( x, y, z) dy dx dz
R z =0 x =−r y=− r2 − x2

In the following example we are dealing with a more complex


geometry of an object. Consequently the bounds of the integrals
will be quite different and can introduce some difficulties into the
calculation.

Example 2.17. A channel weir is a solid bounded by z = x2 − 1, z = 0,


z − 3y = 0, 2z + y = 0. Its form is a river bed whose sides slope up like
a parabola (z = x2 − 1 and z = 0) that is bounded by two planes with
different slopes (z − 3y = 0 and 2z + y = 0). If we want to calculate its
volume we choose f ( x, y, z) = 1. The bounds of the triple integral can be
found by projecting the three dimensional body along the x, y and z–axis.
If we choose our order of integration wisely, we can express the volume as
a single integral. The bounds for x, y, z can be expressed as follows:

−1 ≤ x ≤ 1
2
x −1 ≤ z ≤ 0
z/3 ≤ y ≤ −2z
The channel weir.

So we can calculate the volume with a single integral if we first inte- y


grate with respect to y:
ZZZ
V= 1 dV
R
Z 1 Z 0 Z −2z
= 1 dy dz dx
x =−1 z= x2 −1 y=z/3
Z 1 Z 0
= [y]− 2z
y=z/3 dz dx x
x =−1 z = x 2 −1
Z 1 Z 0
7
= (− z) dz dx
x =−1 z= x2 −1 3 projection along z−axis
Z 1  0
7 1 z
= − z2 dx
x =−1 3 2 z = x 2 −1
Z 1 2 x
7 2
= x − 1 dx
x =−1 6
 1
7 1 5 2 3 42 projection along y−axis
= x − x +x = .
6 5 3 x =−1 45 z
y

projection along x−axis


Three possible projections of the channel weir.
40 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

2.3 Centre of mass

Twice so far we have mentioned that the function f ( x, y) or f ( x, y, z)


can be a density per area or per volume. In common mechanical
and physics applications we can use double and triple integrals of
mass density functions to calculate the centre of mass as well as the
moment of inertia of the corresponding bodies.
Given a mass density function ρ( x, y, z) the total mass of a body If the body is very thin we might
can be determined by want to drop the
RR third dimension and
evaluate M = R ρ( x, y) dA.

ZZZ
M= ρ( x, y, z) dV.
R

After we evaluate the first moments 3 3


The first moments measure the
distribution of the area/volume of a
ZZZ shape in relationship to an axis.

Myz = xρ( x, y, z) dV
Z Z ZR
Mxz = yρ( x, y, z) dV
Z Z ZR
Mxy = zρ( x, y, z) dV
R

we can evaluate the centre of mass

Myz Mxz Mxy


(Cx , Cy , Cz ) = ( , , ).
M M M

Note that the first moments project the density distribution ρ( x, y, z)


onto the relevant axis.
When dealing with thin objects, we can assume we are working
in the xy-plane and solve double instead of triple integrals. In that
case the first moments and the centre of mass are:

ZZ
My = xρ( x, y) dV
Z ZR
Mx = yρ( x, y) dV
R
My M x
(Cx , Cy ) = ( , ).
M M

Example 2.18. Calculate the mass and the centre of mass for a half-disk
x2 + y2 ≤ r2 , y ≥ 0, where the mass density function is ρ( x, y) = 1. (so the mass is just the area)

ZZ Z r Z √r 2 − x 2
Mdisk = M = ρ( x, y) dA = dy dx
disk x =−r y =0

πr2
Z r Z r
2 2
[y]y=r 0− x dx =
p
= r2 − x2 dx = .
x =−r x =−r 2

See Example 1.6 for the last equality.


mathematical methods 2 41

The first moments are:


ZZ Z r Z √r 2 − x 2
My = xρ( x, y) dA = x dy dx
disk x =−r y =0
Z r  r
p 1
= x r2 − x2 dx = − (r2 − x2 )3/2 =0
x =−r 3 x =−r
ZZ Z r Z √r 2 − x 2
Mx = yρ( x, y) dA = y dy dx
disk x =−r y =0

r2 − x2
y2 r2 − x2
Z r   Z r
= dx = dx
x =−r 2 y =0 x =−r 2
r
r2 x x3 2r3

= − =
2 6 x =−r 3
4r
In conclusion the centre of mass is ( My /M, Mx /M ) = (0, 3π ).
Note that the first coordinate is easily deduced from the symmetry of the
problem.

Exercise 2.3.1. Calculate the centre of mass for a cylinder x2 + y2 ≤ r2


and 0 ≤ z ≤ h, where the mass density function is ρ( x, y, z) = 1.

2.4 More applications to physics

In the dynamics of rigid bodies there exist a nice similarity be-


tween the linear motion of a body and the rotation of a body. For
example, in one variable, the energy E and the momentum (linear
momentum p or angular momentum L) are:

1
E= Mv2 , p = Mv, if the body moves in a straight line
2
1
E = Iω 2 , L = Iω, if the body is rotating.
2
Notice that in the second equation, the role of the total mass M is
replaced by the moment of inertia I and the role of the velocity v
is replaced by the angular velocity ω. The symbol E represents the
kinetic energy in the first equation and the rotational kinetic energy
in the second one.
If A is the axis of rotation, it can be shown that the scalar I is
given by the expression
ZZZ
I= ( distance from ( x, y, z) to A)2 ρ( x, y, z) dV.
R

In practice we will generally consider the case where A is one of


the main axes. The moment of inertia about the main axes can be
determined by computing the second (principal) moments:
ZZZ
Ixx = (y2 + z2 )ρ( x, y, z) dV
R
ZZZ
Iyy = ( x2 + z2 )ρ( x, y, z) dV
R
ZZZ
Izz = ( x2 + y2 )ρ( x, y, z) dV,
R
42 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

where Ixx , Iyy , Izz give the moment of inertia about the x, y or z–axis
respectively. Intuitively, as a mass is moved away from the x-axis
Ixx increases.
For the interested student:
If we want to know the value of E and I for an arbitrary axis A,
RRR 2
the integral R ( distance from ( x, y, z ) to A) ρ ( x, y, z ) dV is often
not easy to compute. Instead we can compute further moments:
ZZZ
Iyz = yzρ( x, y, z) dV
Z Z ZR
Ixz = xzρ( x, y, z) dV
ZZZ R
Ixy = xyρ( x, y, z) dV,
R

which defines the last moments in the moment of inertia matrix (or
sometimes called inertia tensor):
 
Ixx − Ixy − Ixz
I = − Ixy Iyy − Iyz  .
 
− Ixz − Iyz Izz

Then E = 21 ω T I ω, where ω is a column vector4 , and I = 1


| ω |2
ω T I ω, 4
Recall that the angular velocity ω is a
vector in the direction of A.
or put another way I = uT I u where u is a unit vector in the direc-
tion of A.
3
Change of coordinates in double and triple integrals

We have seen several examples of how one-dimensional integrals


can sometimes be simplified by making an appropriate change
of variable (or substitution). In this chapter we will develop this
idea further to look at how the idea generalises to two and three
dimensions. First however, remember that the general form of
integration by substitution can be expressed in one dimension by Note that we switched the roles of u
and x from Subsection 1.3.3, but since
Z b Z d u and x are dummy variables, we are
f ( x ) dx = f ( g(u)) g0 (u) du, allowed to do that.
a c

with the substitution:

x = g ( u ),
dx = g0 (u) du,

where, supposing that g has an inverse function, the bounds or


limits change as follows:

x=a =⇒ u = g−1 ( a) = c,
x=b =⇒ u = g−1 (b) = d.

The first observation we can make is that the substitution is actually


a change of coordinates, also called a transformation. Instead of inte-
grating with variable x over [ a, b] we integrate with variable u over
[c, d]. We can change the coordinates similarly in double and triple
integrals.

3.1 Change of coordinates in double integrals

In many cases the evaluation of a double integral


ZZ
f ( x, y) dxdy
R

over a certain region R in R2 can be significantly simplified by a


change of coordinates. This is done by using some new coordinates
e.g. u, v and a certain transformation g(u, v) = ( x, y) which tells
us how the new coordinates u, v are related to the old ones x, y. We
44 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

write such a change of coordinates as:

( x, y) = g(u, v) = (φ(u, v), ψ(u, v)),


 
∂g(u, v)
dx dy = det du dv,
∂(u, v)
S = g −1 ( R ),

where S is the new region of integration, i.e. the domain of the new
coordinates u, 
v. 
∂g(u,v)
The matrix ∂(u,v) is the Jacobian matrix of the transformation
g(u, v) at (u, v) (recall the definition from Section 7.5 of the Mathe-
matical Methods 1 notes), that is,
!
  ∂φ(u,v) ∂φ(u,v)
∂g(u, v)
= ∂ψ(u,v) ∂ψ(u,v) ,
∂u ∂v
∂(u, v)
∂u ∂v

and its determinant J (u, v) is called the Jacobian of g(u, v) at (u, v).
Note that the change of coordinates requires a transformation
function g that maps a two-dimensional region onto another two-
dimensional region. Moreover g can be differentiated on S. This
mapping has to be one–to–one, so that for each point ( x, y) in R
there is a unique point (u, v) in S with g(u, v) = ( x, y).
More precisely, we have the following.

Theorem 3.1. (Change of Variables in Double Integrals)


Assume that g : S −→ R is a transformation mapping a region S
in R2 onto another region R = g(S) in R2 such that:

(i) g is continuously differentiable on S and J (u, v) 6= 0 on S;

(ii) g is one–to–one, i.e. each point ( x, y) ∈ R has a unique preimage


(u, v) ∈ S.

Then
 
∂g(u, v)
ZZ ZZ
f ( x, y) dx dy = f (g(u, v)) det du dv.
R S ∂(u, v)

The above formula remains true if the conditions (i) and (ii) are sat-
isfied on S except possibly on a finite union of continuously differen-
tiable curves.

In practice S is just a description of the domain of integration in


terms of u, v instead of x, y.
In what follows we will briefly try to explain why the Jacobian
J (u, v) appears in the formula for change of variables.
To understand how R and S are related we consider the transfor-
mation:      
u φ(u, v) x
7−→ g(u, v) = = .
v ψ(u, v) y
mathematical methods 2 45

For convenience we assume that S is a rectangular region with


boundaries parallel to the u and the v-axes. Consider in S a line
parallel to the u–axis (say v = v0 ). The mapping from S to R trans-
forms this line into an arc g(u, v0 ) in the xy–plane. Similarly, we
obtain an arc for any line parallel to the v–axis. Using lines that
are parallel to the u and v–axes we can partition our region S into
subrectangles.

Figure 3.1: Action of g


v y

S R

u x

The one–to–one property of the transformation g guarantees


that the arcs g(u, vk ) of the partitioning of R do not intersect each
other1 for different values of vk . Similarly for the arcs g(u j , v). If we 1
that means that they are in some
sense parallel even on R
consider the shaded region in S, we see that the corner points of the
rectangle are:
u j + ∆u j u j + ∆u j
       
uj uj
, , , ,
vk vk vk + ∆vk vk + ∆vk
and therefore the area of the shaded region is ∆u j ∆vk . Under the
action of the transformation the area may be changed. First we note
that the corner points are now:
φ(u j , vk ) φ(u j + ∆u j , vk )
   
,
ψ(u j , vk ) ψ(u j + ∆u j , vk )
φ(u j , vk + ∆vk ) φ(u j + ∆u j , vk + ∆vk )
   
, .
ψ(u j , vk + ∆vk ) ψ(u j + ∆u j , vk + ∆vk )
For a very fine partitioning of S the area in R delimited by these
corner points can be approximated by the area of a parallelogram.
In R2 the area of a parallelogram is given by the determinant of the
vectors defining two adjacent sides of the parallelogram.
Theorem 3.2. The area of the parallelogram defined by the vectors
! !
a1 b1
a= , b=
a2 b2
is !
a1 b1
det .
a2 b2 b H = |b| sin θ

Proof. We see from the picture that the area of the parallelogram is
θ
| a|.H = | a||b| sin θ, where θ is the angle between a and b. a
Now we are considering both vectors as part of the plane z = 0
in 3 dimensions. We set a0 = ( a1 , a2 , 0), b0 = (b1 , b2 , 0). Notice that
the angle between a0 and b0 is also θ, and that | a0 | = | a|, |b0 | = |b|.
46 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

The cross product a0 × b0 has length (modulus) | a0 ||b0 | sin θ (seen


in a Tutorial in Mathematical Methods 1). Since
 
e1 e2 e3 !
0 0 a1 a2
a × b = det  a1 a2 0  = det e3 ,
 
b1 b2
b1 b2 0

we get that
! !
0 0 a1 a2 a1 b1
| a||b| sin θ = | a ||b | sin θ = det = det
b1 b2 a2 b2

Since the shaded parallelogram in R is defined by the vectors

φ(u j + ∆u j , vk ) φ(u j , vk ) φ(u j , vk + ∆vk ) φ(u j , vk )


       
− and − ,
ψ(u j + ∆u j , vk ) ψ(u j , vk ) ψ(u j , vk + ∆vk ) ψ(u j , vk )

the area of the shaded region in R is given by the absolute value of


the determinant:
!
φ(u j + ∆u j , vk ) − φ(u j , vk ) φ(u j , vk + ∆vk ) − φ(u j , vk )
det .
ψ(u j + ∆u j , vk ) − ψ(u j , vk ) ψ(u j , vk + ∆vk ) − ψ(u j , vk )

For ∆u j small we can approximate the first difference by using a


linear Taylor expansion:

∂φ(u j , vk )
φ(u j + ∆u j , vk ) − φ(u j , vk ) ≈ ∆u j .
∂u
Using similar approximations for ∆vk and for ψ, we get an area
approximation in R that is:
∂φ(u j ,vk ) ∂φ(u j ,vk ) ∂φ(u j ,vk ) ∂φ(u j ,vk )
! !
∆u j ∆vk
det ∂ψ(u ,v )
∂u ∂v
∂ψ(u j ,vk ) = det ∂ψ(u j ,vk ) ∂ψ(u j ,vk ) ∆u j ∆vk
∂u ∂v

∂u
j k
∆u j ∂v ∆v k ∂u ∂v
∂g(u j , vk )
 
= det ∆u j ∆vk .
∂(u, v)
We can use Riemann sums over our partitioning to approximate the
double integral of f ( x, y) on R:
N M 
∂g(u j , vk )

∑∑ f (φ(u j , vk ), ψ(u j , vk )) det
∂(u, v)
∆u j ∆vk
j =1 k =1

The limit of the Riemann sum (as N, M → ∞ and the partitions


become finer) converges to an integral and we obtain:
 
∂g(u, v)
ZZ ZZ
f ( x, y) dx dy = f (φ(u, v), ψ(u, v)) det du dv.
R S ∂(u, v)
We state again that the change of coordinates requires that each
point in R has one corresponding point in S.
RR
Example 3.3. Evaluate the integral D (3x − 2y) dA over the parallelo-
gram D bounded by the lines 3x − 2y = 2, 3x − 2y = −1, 2x + y = 14
and 2x + y = 0.
mathematical methods 2 47

Solution: The difficulty with this problem is that the boundaries of the
region do not align conveniently with the coordinate axes. Guided by the
fact that a linear change of variables maps parallelograms in one coor-
dinate system to parallelograms in the other, it would seem that putting
u = 3x − 2y and v = 2x + y would be a good choice. Then the domain in
(u, v) space is defined by −1 ≤ u ≤ 2 and 0 ≤ v ≤ 14.
To calculate the Jacobian we first need to find x and y in terms of u
and v; it is easy to show that with u and v defined as above then x =
(u + 2v)/7 and y = (3v − 2u)/7. Then
! !
∂x ∂x 1 2
7 7 1
det ∂y ∂y = det
∂u ∂v
2 3 = .
∂u ∂v
−7 7 7

Then it follows that

1 2 2
Z 14 Z 2 Z 14 
3 14

1 1
ZZ Z
(3x − 2y) dA = (u) · dudv = u dv = dv = 3.
D 0 −1 7 7 0 2 −1 14 0

RR y4
Example 3.4. Evaluate I = D x dxdy over the region D contained
between the parabolas x = 1 − y2 and x = 4(1 − y2 ).
Solution: First we need to understand the geometry of the region D. It is
clear that the two parabolas are symmetric about the x-axis; they meet at
x = 0, y = ±1 and they cut the x-axis at x = 1 and x = 4 respectively.
If we therefore define new coordinates x = v(1 − u2 ) and y = u then
the two parabolas are given by v = 1 and v = 4 and they meet at u = ±1.
Notice that in these coordinates the appropriate region of integration is
again just a simple rectangle with sides parallel to the u and v axes; this is
much simpler than the corresponding region in ( x, y) space.
Once more we must compute the Jacobian
! !
∂x ∂x
−2uv 1 − u2
J = det ∂y ∂y = det
∂u ∂v = −1(1 − u2 ) = u2 − 1.
∂u ∂v
1 0

Notice that the region of integration satisfies −1 ≤ u ≤ 1. This determi-


nant is less or equal to 0 on the region and so we must not forget to take
its absolute value: as here J = u2 − 1 ≤ 0 we take | J | = 1 − u2 .
Transforming the integral we find that

y4 u4
ZZ Z 4Z 1
I = dxdy = | J |dudv
D x 1 −1 v (1 − u2 )
u4
Z 4Z 1 Z 4Z 1 4
u
= (1 − u2 )dudv = dudv
1 v (1 − u2 )
−1 1 −1 v
Z 4  5 1 Z 4
u 2 1 2
= dv = dv = ln 4.
1 5v −1 5 1 v 5
48 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

3.2 Change of coordinates in triple integrals

This is very similar to the case of double integrals. In an integral of


the form ZZZ
f ( x, y, z) dxdydz,
R
where R is a solid in R3 , we would to change the coordinates x, y, z
to new coordinates u, v, w using a transformation g:

( x, y, z) = g(u, v, w),
 
∂g(u, v, w)
dx dy dz = det du dv dw,
∂(u, v, w)
S = g −1 ( R ),

where S is the pre-image of R under g, that is, the domain of the


new coordinates u, v, w.  
∂g(u,v,w)
As before, the matrix ∂(u,v,w) is the Jacobian matrix of the
transformation g(u, v, w) at (u, v, w), that is, if

g(u, v, w) = ( g1 (u, v, w), g2 (u, v, w), g3 (u, v, w)),

then
1 ( u,v,w ) ∂g1 (u,v,w) ∂g1 (u,v,w) 
 ∂g
  ∂u ∂v ∂w
∂g(u, v, w)  ∂g2 (u,v,w) ∂g2 (u,v,w) ∂g2 (u,v,w) 
= .
∂(u, v, w)
 ∂u ∂v ∂w
∂g3 (u,v,w) ∂g3 (u,v,w) ∂g3 (u,v,w)
∂u ∂v ∂w

Its determinant J (u, v, w) is called the Jacobian of g(u, v, w) at


(u, v, w).

Theorem 3.5. (Change of Variables in Triple Integrals)


Assume that g : S −→ R is a transformation mapping a solid S in
R3 onto another solid R = g(S) in R3 such that:

(i) g is continuously differentiable on S and J (u, v, w) 6= 0 on S;

(ii) g is one–to–one, i.e. each point ( x, y, z) ∈ R has a unique pre-


image (u, v, w) ∈ S.

Then
ZZZ
f ( x, y, z) dx dy dz =
R
 
∂g(u, v, w)
ZZZ
f (g(u, v, w)) det du dv dw.
S ∂(u, v, w)

The above formula remains true if the conditions (i) and (ii) are sat-
isfied on S except possibly on a finite union of continuously differen-
tiable surfaces.

To understand the above formula we can easily extend the proce-


dure described in Sect. 3.1 to triple integrals. The main ideas can be
summarised as follows:
mathematical methods 2 49

1. we partition S using right parallelepipeds (instead of rectangles


in two dimensions);

2. the transformation g deforms these blocks, but for sufficiently


fine partitioning their volume can be approximated by assum-
ing they are parallelepipeds (instead of parallelograms in two
dimensions);

3. the volume of these parallelepipeds is given by the determinant


of the vectors that span this body. This determinant is the Jaco-
bian of the transformation g.

Example 3.6. A torus (doughnut) has major radius b and minor radius a
b
(0 < a < b). a

Use the transformation


   
x (b + r cos α) cos θ
 y  = g(θ, α, r ) =  (b + r cos α) sin θ 
   
z r sin α

where 0 ≤ θ ≤ 2π, 0 ≤ α ≤ 2π and 0 ≤ r ≤ a to determine the volume


of the doughnut.
Solution: Often the most tricky part of changing variables in three di-
mensions is the calculation of the appropriate Jacobian. Here we need to
determine
 ∂x ∂x ∂x   
∂θ ∂α ∂r −(b + r cos α) sin θ −r sin α cos θ cos α cos θ
J = det  ∂y ∂y ∂y 
= det (b + r cos α) cos θ −r sin α sin θ cos α sin θ  .
  
∂θ ∂α ∂r  
∂z
∂θ
∂z
∂α
∂z
∂r
0 r cos α sin α

To evaluate this determinant the zero in the third row indicates that the
simplest calculation will be using an expansion by this third row and then

J = r cos α(b + r cos α) cos α(sin2 θ + cos2 θ ) + sin α(b + r cos α)r sin α(sin2 θ + cos2 θ )

= r cos2 α(b + r cos α) + r sin2 α(b + r cos α) = r (b + r cos α).


Note that on the region of integration, J is non-negative, so | J | = J.
Given this, the volume of the doughnut can be found by evaluating
Z a Z 2π Z 2π Z a Z 2π
r (b + r cos α)dαdθdr = [r (bα + r sin α)]2π
0 dθdr
0 0 0 0 0
Z a Z 2π Z a
2
= 2πb rdθdr = 4π b rdr = 2π 2 a2 b.
0 0 0
Hence the volume of the torus is 2π 2 a2 b.

For the interested student:


While integration by substitution can be generalised to double
and triple integrals the underlying theory that proves this is far
from trivial. Here we only highlighted some concepts that explain
the method and work in the common and generic transformations
that arise in practice. In this section we just assumed that there
exists a suitable transformation g and moreover a partitioning of S.
The corresponding proof is beyond the level of this course.
50 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

3.3 Three important coordinate changes

The example of the torus above illustrates one of the main obsta-
cles to using coordinate transformations in three dimensions. Fre-
quently the calculation of the Jacobian is somewhat tedious and
lengthy but many practical problems involve circular, cylindrical
or spherical geometries. There are three sets of coordinates that
are specifically designed to assist with the solution of these classes
of problem and, rather that having to work from first principles
each time, it is convenient to be able to use the standard results for
these coordinates. Accordingly, in the remainder of this chapter,
we introduce these three special coordinate systems known as polar
coordinates, cylindrical coordinates and spherical coordinates.
The choice of substitution/new coordinate system for solving a
double or triple integral depends on the particular problem. Above
we discussed a problem with a geometry that was simple on S but
complex on R. In these cases it is useful to make the transformation
and evaluate the original problem in the new coordinate system. In
the applied sciences, like physics, the new coordinate systems are
sometimes called natural coordinate systems.

3.3.1 Polar coordinates


y
In Example 2.18 we learnt how to determine the area of a disc
x2 + y2 ≤ r2 of radius r using a double integral. While the first
integration of the double integral had an obvious solution, the
second integration needed to be solved by substitution. This disc– ρ sin θ ρ
example highights the power of coordinates changes: the geometry θ

given by the functional dependence of x and y is complex but in- x


ρ cos θ
stinctively we feel that it should not be. After all it is just a disc and
therefore everything we need is a radius 0 ≤ ρ ≤ r and some angle
0 ≤ θ < 2π. To determine the right kind of transformation from x, y
to ρ, θ we make use of some simple geometry that we learnt when
the polar coordinates transformation.
the trigonometric functions were introduced in high school. For a
point ( x, y) in the disc, we define its angle θ as the angle between
the vector ( x, y) and the half-axis x ≥ 0, and ρ as the modulus of
( x, y) (the length of the vector ( x, y)). We immediately see that:

x = ρ cos θ,
y = ρ sin θ.

While in our sketch we have chosen some special θ = θ1 and ρ = r


it is clear that each point on the line defined by θ = θ1 will only
lead to a variation of ρ in [0, r ]. Similarly, all the other points on the
disc can be uniquely2 described by a pair of polar coordinates ρ, θ. 2
except for the origin
Our new coordinates comply with common, historical defini-
tions:
• the radius ρ is always defined as a positive number ρ ≥ 0.
• θ is the angle between ( x, y) and x ≥ 0 and 0 ≤ θ < 2π.
While these are not the only possible choices for our polar coor-
mathematical methods 2 51

dinate system (for instance we could choose −π ≤ θ < π), it is


the one that is generally accepted for historical reasons. Just like
in other sciences such axiomatic choices3 are necessary even in the 3
for example the choice of an SI–unit
abstract mathematical sciences. They enable us to communicate and system

adapt results.
Using the proposed transformation:
   
ρ cos θ x
g(ρ, θ ) = = ,
ρ sin θ y

we see that:
!
−ρ sin θ
 
∂g(ρ, θ ) cos θ
det = det = ρ(cos2 θ + sin2 θ ) = ρ.
∂(ρ, θ ) sin θ ρ cos θ

This leads us to our the substitution theorem in the polar coordi-


nate system:

Theorem 3.7. Let R be an area on the R2 plane given by the Carte-


sian coordinates x and y. The double integral of a function f ( x, y)
RR
defined on R: R f ( x, y) dx dy can be transformed to the polar coor-
dinates ρ and θ by substitution of:
   
x ρ cos θ
= g(ρ, θ ) = ,
y ρ sin θ

where we take ρ ≥ 0, 0 ≤ θ < 2π. Then, for S = g−1 ( R),


ZZ ZZ
f ( x, y) dx dy = f (ρ cos θ, ρ sin θ )ρ dρ dθ.
R S

As in Theorem 3.1, the above change of coordinates can be used


for any bounded function f ( x, y) which is continuous on R except
possibly over a finite union of differentiable curves.
We will use this change of coordinate every time R is 2-dimensional
and has some sort of rotational symmetry.
We can use the transformation g to determine S. For a full disc
this might be trivial, since 0 ≤ ρ ≤ r and 0 ≤ θ ≤ 2π. But for
parts of the disc, x and y values and therefore the corresponding
polar coordinates are restricted by additional inequalities. This will
become clearer with some examples.

Example 3.8. The area of a disc x2 + y2 ≤ r2 can be calculated easily


using polar coordinates: S = {(ρ, θ )|0 ≤ ρ ≤ r, 0 ≤ θ < 2π } and so
ZZ Z 2π Z r
A = 1 dx dy = 1 · ρ dρ dθ
disc 0 0
Z 2π  2 r
ρ r2 2π
= dθ = [θ ]0 = πr2 .
0 2 0 2
52 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov y
y = bx

Example 3.9. Let R be the “pizza slice” given by x2 + y2 ≤ r2 , y ≥ ax y = ax


and y ≤ bx (see picture). Describe S = g−1 ( R) in polar coordinates.
θ0 θ1
Solution: Using our transformations we get: x

θ0 = arctan ( a) ,
θ1 = arctan (b) ,

and therefore two new inequalities in polar coordinates:

0≤ρ≤r
θ0 ≤ θ ≤ θ1 ,

that describe the same geometrical region, but in polar coordinates. So


S = g−1 ( R) = {(ρ, θ )|0 ≤ ρ ≤ r, θ0 ≤ θ ≤ θ1 }.

These examples give us some new interpretation for the deter-


minant of the Jacobian matrix. While in Cartesian coordinates the
geometry is complex (a disc) in polar coordinates it is simple (a
rectangle). The factor ρ corrects for the change in area, or density,
that the change of coordinates cause. Therefore it is often called a
density correction.
One more example:

Example 3.10. Compute the area of R, described by

4 ≤ x 2 + y2 ≤ 9
1
x ≤ y ≤ 2x
2

y θ

R S

r
x

Solution: The region is a ring segment. From the first inequality we get
bounds on the radius
2 ≤ ρ ≤ 3.
For the lower bound on θ note that we need to use y = x/2. Substitution
of g gives us θ0 = arctan 1/2 and similarly by using y = 2x we find
θ1 = arctan 2. Therefore the area is:
Z arctan 2 Z 3 Z arctan 2  2 3  
ρ 5 1
A= ρ dρ dθ = dθ = arctan 2 − arctan .
arctan 1/2 2 arctan 1/2 2 2 2 2

Using polar coordinates can also simplify double integrals of


functions. Especially in mechanics and electro–statics4 we often 4
e.g. forces in a gravitational field and
in a electrical field. We’ll come back to
them in Chapters 5 and 6.
mathematical methods 2 53

need to integrate functions of the form:


x y
f ( x, y) = p +p .
x2 + y2 x2 + y2

Using polar coordinates simplifies f ( x, y) dramatically and we get:

f (ρ, θ ) = cos θ + sin θ.

Therefore in double integrals of such functions the only ρ depen-


dence results from our density correction (Jacobian).
Moreover the following exercise shows that polar-like coordi-
nates are useful for regions other than circular geometries. Geome-
tries that are similar to discs often are easier to integrate in polar
coordinates too.

Exercise 3.3.1. Determine the area of an ellipse given by the inequality:

x2 y2
+ ≤ 1.
a2 b2
(The change of coordinates you need to make here is similar to polar coor-
dinates but is slightly different.)

3.3.2 Cylindrical coordinates


z
Everything we learnt about polar coordinates we can apply to
cylindrical coordinates, since cylinders – the simplest structure in ρ
this coordinate system – are discs with an added third dimension.
Therefore we can use this transformation function: ξ
   
ρ cos θ x
g(ρ, θ, ξ ) =  ρ sin θ  =  y  ,
   
ξ z y
θ
where we added a third Cartesian dimension ξ = z and kept the x
the cylindrical coordinate system.
known conventions for ρ and θ. The Jacobian of the transformation
g is:  
  cos θ −ρ sin θ 0
∂g(ρ, θ, ξ )
det = det  sin θ ρ cos θ 0 = ρ.
 
∂(ρ, θ, ξ )
0 0 1
Therefore we obtain the following substitution theorem in the cylin-
drical coordinate system:

Theorem 3.11. Let R be a region in R3 given by the Cartesian


coordinates x, y and z. The triple integral of a function f ( x, y, z)
RRR
defined on R: R f ( x, y, z ) dx dy dz can be transformed to the
cylindrical coordinates ρ, θ and ξ by substitution of:
   
x ρ cos θ
 y  = g(ρ, θ, ξ ) =  ρ sin θ  ,
   
z ξ
54 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

where we take ρ ≥ 0, 0 ≤ θ < 2π, ξ ∈ R. Then, for S = g−1 ( R),


ZZZ ZZZ
f ( x, y, z) dx dy dz = f (ρ cos θ, ρ sin θ, ξ )ρ dρ dθ dξ.
R S

As before, here we assume that f ( x, y, z) is a bounded function


on R which is continuous on R except possibly over a finite union
of continuously differentiable surfaces.
Since ξ = z often we just use z and distinguish between the
Cartesian coordinates and the cylindrical ones by context.
The following simple generic examples are left to the reader.
Exercise 3.3.2. Determine the volume of a cylinder given by: x2 + y2 ≤
9 and 0 ≤ z ≤ h.
Exercise 3.3.3. Determine the volume of the body given by: x2 + y2 ≤ 9,
x ≥ 0, y ≥ 0 and 0 ≤ z ≤ h.
Instead we want to focus on an example where we use the meth-
ods of integrating non–rectangular regions in the new coordinate
system.
Example 3.12. Determine the volume of a cone of height h, given in
Cartesian coordinates by the inequalities:
 z 2
x 2 + y2 ≤ r 2 1 − , 0 ≤ z ≤ h.
h
Solution: Using our transformation we can rewrite the inequalities in
cylindrical coordinates:
 z
0 ≤ ρ ≤ r 1− , 0 ≤ θ ≤ 2π, 0≤z≤h
h
Note that the transformation has reduced the complexity of the problem.
In Cartesian coordinates we had just 2 inequalities, one of which was
quadratic. Now we have the same number of inequalities as integrals to
solve. Therefore we get:
ZZZ ZZZ
V= dx dy dz = ρ dρ dz dθ =
R S
Z 2π Z h Z r (1− z ) Z 2π Z h  2 r (1− hz )
h ρ
= ρ dρ dz dθ = dz dθ =
0 0 0 0 0 2 0
r 2 Z 2π Z h 
2z z2

= 1− + 2 dz dθ
2 0 0 h h
h
r2 z2 z3
Z 2π  
= z− + 2 dθ
2 0 h 3h 0
r2
Z 2π  
h
= dθ
2 0 3
1
= πr2 h.
3
Note that the boundary of S can be found by visualising the three dimen-
sional body in its projections along the x or y–axis. The projections are
triangles bounded by z = 0 and lines that intersect the z–axis at h and the
other axis at ±r, and a circle of radius r.
mathematical methods 2 55

3.3.3 Spherical coordinates


While the cylindrical coordinate system is just the polar coordinate
system with an added third Cartesian dimension, the spherical
coordinate system requires transformation of all three Cartesian
coordinates. But again there are a several similarities to the polar
coordinate system. As we have seen a disc can be described using
one radius and one angle. To get the third dimension we just add
another angle:
   
r cos θ sin φ x
g(r, θ, φ) =  r sin θ sin φ  =  y  .
   
r cos φ z
This time r is the distance from the origin, and θ is still the angle
θ
from x-axis in a plane parallel to the xy-plane. Our third coordinate
φ is the angle down from the z-axis. Commonly the range of φ is φ
r

0 ≤ φ ≤ π. The reader might want to confirm that a range of size


π for φ is enough to describe every point in R3 as long as r ≥ 0
and the range of θ is 0 ≤ θ < 2π. We can think about θ being the
longitude and φ the latitude of a point. The Jacobian of g is:
 
  cos θ sin φ −r sin θ sin φ r cos θ cos φ the spherical coordinate system.
∂g(r, θ, φ)
det = det  sin θ sin φ r cos θ sin φ r sin θ cos φ 
 
∂(r, θ, φ)
cos φ 0 −r sin φ
= −r2 sin φ.
 
∂g(r,θ,φ)
Since 0 ≤ φ ≤ π, sin φ is non-negative, det ∂(r,θ,φ) = r2 sin φ,
and therefore the last substitution theorem follows as:

Theorem 3.13. Let R be a region in R3 given by the Cartesian


coordinates x, y and z. The triple integral of a function f ( x, y, z)
RRR
defined on R: R f ( x, y, z ) dx dy dz can be transformed to the
spherical coordinates r, θ and φ by substitution of:
   
x r cos θ sin φ
 y  = g(r, θ, φ) =  r sin θ sin φ  ,
   
z r cos φ

where we take r ≥ 0, 0 ≤ φ ≤ π, 0 ≤ θ < 2π.


Then, for S = g−1 ( R),
ZZZ
f ( x, y, z) dx dy dz
R
ZZZ
= f (r cos θ sin φ, r sin θ sin φ, r cos φ)r2 sin φ dr dθ dφ.
S

As in Theorem 3.5, here we assume that f ( x, y, z) is a bounded


function on R which is continuous on R except possibly over a
finite union of continuously differentiable surfaces.
56 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

Obviously – like in polar coordinates – inequalities of the Carte-


sian coordinates x, y, z that describe a body in R3 can be used to
determine the equivalent inequalities in the spherical coordinate
system. We illustrate the integration by substitution in spherical
coordinates using two examples:

Example 3.14. Evaluate the volume of an ice cream cone given by:
π
0 ≤ r ≤ R, 0 ≤ θ ≤ 2π, 0≤φ≤ .
4
Solution: Given these inequalities, we can determine the volume Vcone of
the cone by:
ZZZ ZZZ
Vcone = 1 dx dy dz = r2 sin φ dr dφ dθ
cone S
Z 2π Z π/4 Z R Z 2π Z π/4  3  R
2 r
= r sin φ dr dφ dθ = sin φ dφ dθ
0 0 0 0 03 0
√ !
R3
Z 2π
R3 2 2πR3  √ 
= [− cos φ]0π/4 dθ = 1− [θ ]2π
0 = 2− 2 .
3 0 3 2 6

A wedge and a cone.

Example 3.15. Evaluate the volume of a wedge given by:

π 2π
0 ≤ r ≤ R, ≤θ≤ , 0 ≤ φ ≤ π.
4 3
Solution:
Given these inequalities, we can determine the volume Vwedge of the
wedge by:
ZZZ ZZZ
Vwedge = 1 dx dy dz = r2 sin φ dr dφ dθ
wedge S
2π R
r3
Z R 
5π 3
Z πZ
3
= 2
r sin φ dr dθ dφ = [θ ]2π/3 π
π/4 [− cos φ ]0 = R .
0 π
4 0 3 0 18
4
Path and surface integrals

We have already seen in Section 1.6.3 how to determine the length


of a curve lying in the ( x, y) plane and given in the form y = f ( x ).
This is useful, but in applications there is often need to find lengths
of curves on more complicate surfaces; one example would be
the length of a path taken by a climber as they scaled the side of
a mountain. Unless the mountainside is particularly simple, the
techniques we know about so far are not sufficient to enable us to
calculate the length of the route taken. Moreover, can we evaluate
the surface area of the mountain itself? Such questions require
knowledge of the methods of path and surface integrals that are
explored below.

4.1 Revision of parametric forms

Before we begin to tackle these new type of integral it is important


to revise our knowledge of parametric representations of curves
and surfaces. It will turn out that determining suitable parametric
forms is the starting point of path and surface integrals and the
ability to find such parameterisations is one that needs to be to
mastered. Although much of the following might be familiar, and
some may be revision from high–school, we reintroduce parametric
forms first by exploring some simple examples that later will be
generalised.

4.1.1 Parametric forms of paths


Given some curve in R3 that an object travels along in the time t ∈ C
[ a, b], its location can be determined by finding the position vector y
r(t). Moreover, if the curve is differentiable, we can determine the
velocity v(t) of the object by finding the tangent to the curve. Then
we have the
“position” vector:
r(t) = ( x, y, z) r(t) v (t)
and the “velocity” vector:

dr
x
v(t) = (t) = ṙ(t) a path in the xy–plane.
dt
58 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

where we have used the usual notation that ṙ = dr/dt. By varying


t from a to b we obtain a function r(t) that is called the parametric
form of the curve. We denote by C the curve in its entirety, that is,
C = {r(t)|t ∈ [ a, b]}.

Example 4.1. Find a parametric form of the straight line from (1, 2) to
(3, 6).
Solution: Obviously we are trying to find functions x (t) and y(t) that
together form the components of the parametric form r(t) = ( x (t), y(t)).
Both of them will be linear and therefore of the form:

x ( t ) = x0 + m x t y(t) = y0 + my t.

We can choose any closed t interval over which our parameterisation is to


be applied – suppose we take 0 ≤ t ≤ 1 as the domain of the function.
With this choice we see that for t = 0 we need to be at the starting point
(1, 2). This obviously determines the values x (0) = x0 = 1 and y(0) =
y0 = 2. For t = 1 we want to be at the end point (3, 6) and therefore
x (1) = 3 and y(1) = 6. Substituting these values into our line equations
we get m x = 2 and my = 4. Therefore a parametric form of the line is:
   
x (t) 1 + 2t
r(t) = = , t ∈ [0, 1].
y(t) 2 + 4t
This simple example highlights several important points:

1. a parametric form of a curve is a vector valued function r(t).

2. we need to specify the domain on which r(t) is defined:


t ∈ [ a, b].

3. the range of r(t) includes both the starting and end points:
r( a) is the starting point and r(b) is the end point.

By specifying some interval for t we make a choice of the initial


and final values of t corresponding to the two ends of the line.
This selection is not unique and therefore parametric forms are not
unique representations of curves. We are going to explore this in
the next example but one might first check that a parameterisation
is not unique by tackling the following exercise.

Exercise 4.1.1. Find two parametric forms r(t) of the straight line from
y
(1, 2) to (3, 6) where in the first case t ∈ [0, 2] and in the second case
t ∈ [−1, 1].

Example 4.2. Find a parametric form of the half circle x2 + y2 = 1,


y ≥ 0 from (−1, 0) to (1, 0).
Solution: Given such an implicit curve, the first thing we want to do is x
rewrite it as an explicit curve (in other words, a function):
p
y = 1 − x2 .
mathematical methods 2 59

Note that the inequality y ≥ 0 implies that we only recover the positive
square root solution of y2 = 1 − x2 . The domain of our explicit curve is
x ∈ [−1, 1] and the range y ∈ [0, 1]. One possible parametric form can be
found by simply choosing t = x and then we get:
   
x (t) √ t
r(t) = = , t ∈ [−1, 1].
y(t) 1 − t2

From our discussion on natural coordinate systems we know that this


parametric form might not be the best way to describe a half circle. Us-
ing our knowledge of polar coordinates we can see that a second possible
parametric form is:

cos(π − t)
   
x (t)
r(t) = = , t ∈ [0, π ].
y(t) sin(π − t)

The advantage of this second parametric form becomes obvious


when we try to describe a complete circle:

Example 4.3. Find a parametric form of the whole circle x2 + y2 = 1.


Solution: In Cartesian coordinates we need to express the circle in two
different segments
p p
y = 1 − x2 , y = − 1 − x2 .

If we further require that a particle moves once around the circle starting
and ending at (1, 0) we can write it as the union of these rather messy
parametric forms:

1−t
 
r(t) = p , t ∈ [0, 2],
1 − (1 − t )2
p t−1
 
r(t) = , t ∈ [0, 2].
− 1 − ( t − 1)2

On the other hand in polar coordinates the circle can be expressed in a


much easier single form:
 
cos t
r(t) = , t ∈ [0, 2π ].
sin t

With the help of these examples we have now revised the meth-
ods for finding parametric forms of curves. We can summarise
these methods in a more general context in the following way.
Curve given as an explicit equation: Let C be the curve given by
y = f ( x ) on the domain x ∈ D . Let ( x0 , y0 ) be the starting point
and ( x1 , y1 ) the end point of the curve. A parametric form r(t) of C
can be found by choosing t = x, where t ∈ [ x0 , x1 ]. Then
 
t
r(t) = , t ∈ [ x0 , x1 ].
f (t)

Curve given as an implicit equation: Let C be the curve given by


g( x ) + f (y) = n, where n is some constant real number and g and
f are some invertible functions. Moreover let ( x0 , y0 ) and ( x1 , y1 ) be
the starting and end point of C. A parametric form r(t) of C can be
60 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

found by making the implicit equation explicit: y = f −1 (n − g( x ))


and following the method for explicit equations. Note that it is
sometimes necessary to split up the curve C into segments each
with their own parametric representation. C will then be the union
of these segments.
Sometimes it is easier to use a different parameterisation though
(see Example 4.2).

Exercise 4.1.2. Find parametric forms of C for the following:

1. The ellipse:
x2 y2
+ = 1.
a2 b2
Let ( x0 , y0 ) = ( x1 , y1 ) = ( a, 0).

2. The astroid:
x2/3 + y2/3 = 1.

Let ( x0 , y0 ) = ( x1 , y1 ) = (1, 0).

3. A spiral whose distance from the origin increases by 1 each time the
path completes one revolution.

Note that all these examples are in R2 , but of course we can


similarly descibe curves in R3 by
 
x (t)
r(t) =  y(t)  , t ∈ [ a, b].
 
z(t)

4.1.2 Parametric forms of surfaces


While we only need one parameter to describe a path in space, a
surface is specified by an equation involving two parameters. As a
familiar example we might want to think about the coordinates we
use to describe our location on earth:

( x, y, z) = f (u, v) = ( cos u sin v, sin u sin v, cos v ).

This function gives a point on a sphere of unit radius, where u


represents longitude, v represents latitude.

7−→

(u, v) ∈ R2 7−→ ( x, y, z) ∈ R3

Using the method we have learnt for paths, we can define a


mathematical methods 2 61

position vector and its associated tangent vectors:

( x, y, z) = S(u, v),
∂S
Su (u, v) = (u, v),
∂u
∂S
Sv (u, v) = (u, v).
∂v
Since the surface S depends on two parameters, we obviously have
two tangent vectors at each point (u, v). From MATH1001 we re-
member that it is often useful to work with the vector that is nor-
mal to the surface:
N = Su × Sv ;
in general this vector is a function of u and v.
As we have seen, it is not a difficult task to find a parametric
form when an explicit function is given for a path. The same holds
for surfaces:

Example 4.4. Find a parametric form of the surface S given by z =


x2 − y3 , 0 ≤ x ≤ 1 and 0 ≤ y ≤ 4.
Solution: Obviously z depends only on two variables (x and y) and
therefore we can immediately conclude:

x = u, y = v, z = u2 − v3 .

Therefore

S(u, v) = (u, v, u2 − v3 ), 0 ≤ u ≤ 1, 0 ≤ v ≤ 4.

Again, like in the previous section, the simplest way to find the
parametric form of an implicitly given surface is to transform it into
an explicit one.

Example 4.5. Find a parametric representation of a cone with height h,


basis radius R and apex at (0, 0, 0).
Solution: From the previous chapter we remember that this cone is given
by:
 z 2
x 2 + y2 = R2 1 − , 0 ≤ z ≤ h,
h
and in cylindrical coordinates (with x = ρ cos θ and y = ρ sin θ) we have
 z
ρ = R 1− , 0 ≤ θ ≤ 2π, 0 ≤ z ≤ h.
h
As we can see the radius ρ depends on z but θ and z do not depend on any
other variable. This makes z and θ our favoured choice for u and v. We
get:
u u
x = R(1 − ) cos v, y = R(1 − ) sin v, z=u
h h 
 u u
S(u, v) = R(1 − ) cos v, R(1 − ) sin v, u ,
h h
0 ≤ u ≤ h, 0 ≤ v ≤ 2π.

Exercise 4.1.3. Find a parametric form of the cone in Cartesian coordi-


nates.
62 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

4.2 Length of curves

We shall now explore how the parametric form can be used to


determine lengths of curves.
Consider a path C having a differentiable parametric form r(t)
over the range a ≤ t ≤ b. By this we mean that the limit
r(t + ∆t) − r(t)
lim exists for all t ∈ ( a, b)
∆t→0 ∆t
and we write
dr r(t + ∆t) − r(t)
ṙ(t) = (t) = lim .
dt ∆t→0 ∆t
Clearly, if all components of r(t) are differentiable then the vector
function itself is also differentiable.
To determine the length L(C ) of C we need to partition the in-
terval [ a, b] on which C is defined into subintervals [tk−1 , tk ]. Say
Z = {t0 , t1 , . . . , tn }, where t0 = a and tn = b. This partitioning
allows us to determine the distance between two consecutive points
r(tk−1 ), r(tk ) (k = 1, . . . , n), that is |r(tk ) − r(tk−1 )|. The sum of
these distances gives us an approximation to the length of the path
C = {r(t)|t ∈ [ a, b]}:
n
L(C, Z ) = ∑ |r(tk ) − r(tk−1 )|.
k =1

r(t 3) r(t 4)

r(t2) r(t5)
r(t1)

r(t0)
As the number of points in the partition increases (that is, as n
increases and the points on the curve become closer together) the
approximation to the length of the curve increases – a result that is
essentially a consequence of the triangle inequality. Indeed, in the
limit as n → ∞ the approximation length will converge to the actual
length of the curve C. That is,
n
L(C ) = lim
n→∞
∑ |r(tk ) − r(tk−1 )|.
k =1

Let ∆t = tk − tk−1 . Then we can also write L(C ) in the form


n
r ( t k ) − r ( t k −1 )
L(C ) = lim
n→∞
∑ ∆t
∆t.
k =1
mathematical methods 2 63

As the number of points in the partition n → ∞ their average


spacing ∆t must go to zero (because they all have to fit on the finite
curve defined by t ∈ [ a, b]). Informally we can write This argument can be extended so as
to be made completely rigorous but
n n as this is not required for this unit
r ( t k ) − r ( t k −1 ) dr
L(C ) = lim
∆t→0
∑ ∆t
∆t ≈ lim ∑
∆t→0
(t ) ∆t
dt k−1
the more abstract argument is not
k =1 k =1 included in these notes.

which is the limit of a Riemann sum. Written in terms of integrals


and derivatives we get the following theorem.

Theorem 4.6. If C is given in parametric form by {r(t)| a ≤ t ≤ b},


then Z b Z bq
dr
L(C ) = dt = r˙1 2 + . . . + r˙p 2 dt,
a dt a
dri
where r(t) = (r1 (t), · · · , r p (t)), ṙi = dt , and p = 2 or 3 according
to the ambient space being R2 or R3 .

Given this definition of the length of a path C as an integral over


an interval, it is now possible to consider the path–length function
s(t). In simple terms, the function s(t) is the length of the curve
from its beginning (that is, the point corresponding to t = a) to
the point corresponding to t. This path–length function s is clearly
monotone increasing and continuous and, assuming that the points
r(tk−1 ) and r(tk ) are close together then, approximately,

s(tk ) − s(tk−1 ) ≈ |r(tk ) − r(tk−1 )|.

(Informally this is saying that distance along the path between two
points that are near to each other is approximately the same as the
straight-line distance between them.)
If we divide both sides of this equation by ∆t = tk − tk−1 and
take the limit as ∆t → 0 then
s ( t k ) − s ( t k −1 ) |r(tk ) − r(tk−1 )|
lim = lim
∆t→0 ∆t ∆t→0 ∆t
so that
ds dr
= .
dt dt
Using this result we can conclude that it is possible to write the line
integral as
Z b Z b
dr ds
L(C ) = dt = dt,
a dt a dt

which is often abbreviated to just


Z
L(C ) = ds.
C

Notice that the details of the parameterisation are placed in the


background and not written within this notation. Indeed if we
Rb
wrote limits on the integral in the form of a , this would imply
64 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

something about the parameterisation used as we would be say-


ing that t ∈ [ a, b]. We have already seen that there are many ways
that a particular path can be parameterised and, intuitively, it is
reasonable to expect that the length of a curve should be a quantity
that is independent of the parameterisation chosen. Consequently
R
there is no ambiguity in simply writing the path integral as C ds
for its value is the same regardless as to how the path is parame-
terised. In this integral we can simply think of ds as an infinitesimal
segment of the curve in much the same manner as we considered
infinitesimally thin slices when calculating volumes of revolution
or infinitesimally thin vertical or horizontal strips when evaluating
double integrals. In the next section we shall extend these ideas to
look at infinitesimal patches that, when amalgamated, combine to
form a surface.

Example 4.7. Determine the length of the spiral S given by the parame-
terisation r : [0, 2π ] 7→ R3 , where r(t) = (cos t, sin t, t).
Solution: We compute

ṙ(t) = (− sin t, cos t, 1)


p √
and therefore |ṙ(t)| = (− sin t)2 + (cos t)2 + 12 = 2 for all t. Hence
the length of the spiral is
Z 2π √ √ √
L(S) = 2dt = 2[t]2π
0 = 2 2π.
0

Example 4.8. Find the length of the part of the astroid x2/3 + y2/3 = 1
which is contained in the second quadrant (x ≤ 0, y ≥ 0). y

Solution: First we look for a parameterisation of the astroid. We see that


x1/3 and y1/3 are on a circle of radius 1, so we have the usual parameteri-
sation  1/3   
x cos t
= .
y1/3 sin t
So we choose the parameterisation x

cos3 t
 
r(t) = .
sin3 t
Since our curve is in the second quadrant, it goes from (0, 1) to (−1, 0)
which corresponds to π/2 ≤ t ≤ π.
Now
−3 cos2 t sin t
 
ṙ(t) =
3 sin2 t cos t
so
q p
|ṙ| = 9 cos2 t sin2 t(cos2 t + sin2 t) = 9 cos2 t sin2 t = −3 cos t sin t.

Notice that the quantity |ṙ| must be positive and that for the range
π/2 ≤ t ≤ π it is the case that cos t ≤ 0 and sin t ≥ 0 so the prod-
uct cos t sin t ≤ 0. Hence |ṙ| = −3 cos t sin t. Then
3h 3 3
Z π iπ
L= −3 cos t sin t dt = − sin2 t = − (0 − 1) = .
π/2 2 π/2 2 2
mathematical methods 2 65

4.3 Path integrals of a function

Let f ( x, y) or f ( x, y, z) be a continuous function defined on a


smooth curve C. This means that f can be evaluated for all t ∈ [ a, b]
and we require the function to be continuous in t. Moreover we
require that ṙ exists and is continuous. We can now define path
integrals.

Definition 4.9. (Path integral)


We define the integral of a function f ( x, y) or f ( x, y, z) over a path C =
{r(t)| a ≤ t ≤ b} to be
Z Z
f ds := f (r(t)) |ṙ(t)| dt.
C C

We can think of this integral (sometimes also called a line inte-


gral) as an integration of f along a special path. We need to evalu-
ate f on the path, which we indicate by f ( x (t), y(t)) or f ( x (t), y(t), z(t)).
Imagine for example that f ( x, y) gives us the height of a mountain
range at the position ( x, y) on a map. While driving through this
range, we don’t care about the heights that we do not cross, but
might want to know the cumulative height along our route C. The
R
integral C f ds will determine this.
R
Exercise 4.3.1. Evaluate the path integral C ( x + y + z)ds where C has
the parametric equation

r(t) = (cos t, sin t, t), 0 ≤ t ≤ π.

First we calculate the tangent vector:



ṙ(t) = (− sin t, cos t, 1) ⇒ |ṙ(t)| = 2.

Now our function is f ( x, y, z) = ( x + y + z). To evaluate it on the path


we substitute the components of r:

f ( x (t), y(t), z(t)) = cos t + sin t + t.

and hence
Z Z π √
( x + y + z)ds = (cos t + sin t + t) 2 dt
C 0
π √ √ 
t2 π2
 
= sin t − cos t + 2 = 2 2+ .
2 0 2

4.4 Areas of surfaces

We will generalise the techniques of path integrals to show how


we can evaluate areas of surfaces and double integrals of functions
defined over surfaces which are embedded in R3 . Let S(u, v) be
66 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

a differentiable (and hence continuous) surface in R3 (defined on


some domain D : a ≤ u ≤ b, c ≤ v ≤ d). We denote
S = {S(u, v)| a ≤ u ≤ b, c ≤ v ≤ d}
and we would like to determine the surface area of S. At each point
(u, v), the derivatives Su and Sv exist and hence the normal vector
N(u, v) = Su × Sv

is well-defined.
Consider some infinitesimal rectangle with sides of length ∆u
and ∆v, based at (u, v) in the domain D. This rectangle has an area
∆u∆v in D. Moreover, as the surface S is parameterised in terms of
u and v, the rectangle in D will correspond to some infinitesimal
surface ∆S in S:
v y
v+ v
S
v
S

u u+ u u x
z
Note that in general the element ∆S in S will not be rectangular
but will have some other shape. Unless the parameterisation S is
extremely special it is not possible to find the area of ∆S precisely
but we can derive a very good approximation for its value. We use
now an argument similar to one we used in Chapter 3. Consider
the vertex S(u, v) of ∆S. The shape of ∆S is approximately that of
the parallelogram defined by the vectors
S(u + ∆u, v) − S(u, v) and S(u, v + ∆v) − S(u, v).
The area of the parallelogram is the length of the cross-product
of these vectors. That is
∆S ≈ |(S(u + ∆u, v) − S(u, v)) × (S(u, v + ∆v) − S(u, v))|
∂S ∂S
≈ (u, v)∆u × (u, v)∆v
∂u ∂v
≈ |Su (u, v) × Sv (u, v)| ∆u∆v
≈ |N(u, v)| ∆u∆v.
Hence the total surface area of S is approximately the Riemann sum
∑ N(ui , v j ) ∆ui ∆v j . Taking the limit for size of partitions going to
0, we get the following theorem.

Theorem 4.10. Let S be the surface given by the parameterisation


S = {S(u, v)| a ≤ u ≤ b, c ≤ v ≤ d}. Then the surface area of S,
RR
denoted by S dS, is equal to
ZZ Z dZ b
∂S ∂S
× du dv = |N(u, v)| du dv.
D ∂u ∂v c a
mathematical methods 2 67

This means that we can find the area of a quite complicated


surface S if it can be conveniently parameterised in terms of u and
v. Then, rather than trying to evaluate the area of S directly, we
Note that this theorem can easily be
instead calculate the integral of
generalised to general domains D (not
∂S ∂S necessarily rectangular).
×
∂u ∂v
over an appropriate region D.
Remark. In fact, the formula in Theorem 4.10 can be used as a
definition of surface area.

Example 4.11. Find the area of the part of the surface z = x + y2 that
lies above the triangle with vertices (0, 0), (1, 1) and (0, 1).
Solution: We use u = x and v = y to parametrize the surface:

S(u, v) = (u, v, u + v2 ) v
1
and D
D = {(u, v)|0 ≤ v ≤ 1, 0 ≤ u ≤ v}.
0 1 u
Then the tangent vectors are Su = (1, 0, 1) and Sv = (0, 1, 2v) and hence
 
e1 e2 e3 p
Su × Sv = det  1 0 1  = (−1, −2v, 1) ⇒ |Su × Sv | = 2 + 4v2
 
0 1 2v
Therefore we get,
Z 1 Z v p Z 1 p
A = 2 + 4v2 du dv = v 2 + 4v2 dv
v =0 u =0 v =0 We use a substitution w = 2 + 4v2 ,
i1 √ √ dw = 8vdv to solve the last integral.
1 h 6 2
= (2 + 4v2 )3/2 = − .
12 0 2 6
Note: Computing the integral in the other order of integration is much
harder.
Example 4.12. Find the surface area of the torus (doughnut)

S(α, θ ) = ([b + a cos α] cos θ, [b + a cos α] sin θ, a sin α) b


a

where 0 ≤ α ≤ 2π, 0 ≤ θ ≤ 2π. Note that b is the radius of the circular


centre of the torus and a is the radius of a vertical cross-section of the
torus. These two numbers are constant and a < b. Our two parameters
are θ, the angle on the torus around from x = 0, and α, the angle on this
around from z = 0.
Solution: We get

Sα = (− a sin α cos θ, − a sin α sin θ, a cos α) = a(− sin α cos θ, − sin α sin θ, cos α),

Sθ = (−[b + a cos α] sin θ, [b + a cos α] cos θ, 0) = [b + a cos α](− sin θ, cos θ, 0).
Then
 
e1 e2 e3
N = Sα × Sθ = a[b + a cos α] det  − sin α cos θ − sin α sin θ cos α 
 
− sin θ cos θ 0
= − a[b + a cos α](cos α cos θ, cos α sin θ, sin α),
68 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

and so
p
|N| = a|b + a cos α| cos2 θ cos2 α + sin2 θ cos2 α + sin2 α = a[b + a cos α]. Note that b − a ≤ b + a cos α ≤ b + a,
and so 0 < b + a cos α.

We can now compute the surface area:


Z 2π Z 2π
A= a[b + a cos α]dθdα = 2πa [bα + a sin α]2π 2
α=0 = 4π ab.
α =0 θ =0

Note: This is the product of the perimeters of two circles of radii a and b.

4.5 Surface integrals of a function

Let f ( x, y, z) be a continuous function defined on a smooth surface


S. This means that f can be evaluated for all u ∈ [ a, b] and v ∈
[c, d] and we require the function to be continuous in both u and v.
Moreover we require that N exists and is continuous. We can now
define surface integrals.

Definition 4.13. (Surface integral)


We define the integral of a function f ( x, y, z) over a surface S =
{S(u, v)|(u, v) ∈ D } to be
Z ZZ
f dS := f (S(u, v)) |N(u, v)| du dv,
S D

∂u ( u, v ) × ∂v ( u, v ).
∂S ∂S
where N(u, v) =

This integral can be regarded as an integration of f ( x, y, z) on a


special surface S. We need to evaluate f on the surface, which we
indicate by f ( x (u, v), y(u, v), z(u, v)).
RR
Example 4.14. Evaluate the surface integral S xy dS, where S is the
triangle with vertices (1, 0, 0), (0, 2, 0), (0, 0, 2).

Solution: The surface is contained in a plane so we can parameterise it as


follows:

S(u, v) = (1, 0, 0) + u(−1, 2, 0) + v(−1, 0, 2) = (1 − u − v, 2u, 2v).

The points (1, 0, 0), (0, 2, 0), (0, 0, 2) correspond respectively to (u, v) =
(0, 0), (1, 0), (0, 1), so the domain is

D = {(u, v)|0 ≤ u ≤ 1, 0 ≤ v ≤ 1 − u}.

We compute

∂S ∂S √
× = |(−1, 2, 0) × (−1, 0, 2)| = |(4, 2, 2)| = 2 6
∂u ∂v

and
f (S(u, v)) = (1 − u − v)2u.
mathematical methods 2 69

So the surface integral is


ZZ Z 1 Z 1− u √
xy dS = (1 − u − v)(2u)2 6dvdu
S u =0 v =0
√ Z1 h i 1− u
= 2 6 2u (1 − u)v − v2 /2 du
0 0
√ Z1
= 2 6 u(1 − u)2 du
0

√  u2 2u3 u4 1 6
= 2 6 − + =
2 3 4 0 6
5
Vector fields

In MATH1001 we discussed the importance of the gradient of a


scalar multi-variable function f ( x, y, z) defined to be the vector-
valued function
 
∂f ∂f ∂f
F( x, y, z) = ( x, y, z), ( x, y, z), ( x, y, z) .
∂x ∂y ∂z

Such a function is an example of a vector field.

Definition 5.1. (Vector fields)


A vector field in R p (p = 2 or 3) is a continuously differentiable map For the interested student: in some
from R p to R p . cases, a vector field can be defined on
We write it: an open subset of R p but not on the
whole of R p .

 
! p( x, y, z)
p( x, y)
F( x, y) = in R2 , F( x, y, z) =  q( x, y, z)  in R3 .
 
q( x, y)
r ( x, y, z)

We see that for each point v of R p the vector field gives us a vector
F(v), whose components are given by the corresponding scalar functions
p, q or r evaluated at v.

Vector fields arise in numerous situations. One simple example


occurs in ocean dynamics or meteorology. In order to describe the
motion of water or air fully, it is necessary to know the velocity of
each particle within the system and, of course, since these particles
can move in three dimensions, a complete specification of the ve-
locity requires knowledge of its three components. This then tells
us that the velocity distribution is a vector field. Or, as a second ex-
ample, think about the space probe NASA sent to Saturn. After its
launch the Cassini probe interacted five times with planets of our
solar system. During each of these "fly–bys" the gravitational field
of the planets accelerated the probe making it possible to reach Sat-
urn in less than seven years. As we know from high–school physics,
the gravitational force acting on a body is directed towards the cen-
tre of a planet and decreases with distance from it. Therefore the
72 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

pull acting on the space probe depends on the vector displacement


of the probe relative to the planet and is a vector field.
Other examples of vector fields that are important for applica-
tions include the electric field in a capacitor or the magnetic field
around a bar magnet (both made visible by using iron fragments).
Some examples of vector fields are illustrated below.

The electrical and magnetic field.

A vector field describing wind speed and direction.

In this chapter we are going to encounter further examples of


vector fields in the course of exploring their properties. Many prac-
tical problems such as predicting the motion of a body through a
gravitational field, or finding the path taken by a charged particle
sitting within an electric field or the transport of a parcel of wa-
ter advected by an ocean current, require understanding of vector
fields. It should be stressed that the ideas developed here consti-
tute very much a basic introduction to the substantial topic of the
calculus of vector fields. These ideas will be refined and extended
in Mathematical Methods 3 (MATH2501) where, in particular, the
powerful results of integration theorems will be developed and
discussed.

5.1 Path integrals of vector fields in two dimensions

Consider a particle that moves along some path C in R2 , where


C is parameterised by r(t) for t ∈ [ a, b]. At any position of the
particle we can easily find the tangent vector by simply differen-
tiating the parametric form of C and writing ṙ(t) = dr(t)/dt. In
our subsequent discussion it is going to be convenient to refer re-
peatedly to the unit tangent vector T(t). This is straightforwardly
deduced from knowledge of ṙ(t), for T is parallel to this vector
but has length one; hence the unit tangent vector T(t) is just ṙ(t)
mathematical methods 2 73

divided by its modulus: T = ṙ/|ṙ|.


We say that a vector v is normal to the curve at the point r(t) if
v is orthogonal to ṙ(t) (that is v · ṙ(t) = 0) and define n(t) to be a
unit vector normal to the curve at the point r(t). Note that there are
two such vectors. To avoid any ambiguity, we adopt the following
standard convention: we require that the determinant with columns
n(t) and T(t) (in that order) is equal to 1. This means that if we
rotate n(t) anticlockwise into 90◦ we get T(t). If C is a closed curve1 , 1
A closed curve is a curve whose
like a circle, and is parameterised in the anti-clockwise direction, terminal point is also its initial point,
that is r( a) = r(b).
then this standard convention yields the outwards direction.
We also define another vector: N = n |ṙ|, which has the same
direction as n but the same length as ṙ. This vector is actually easier
to determine than n so it is the one we use in practice. Note that
if ṙ = (c, d) then (−d, c) and (d, −c) are orthogonal to ṙ and have
√ d c
the same length ( c2 + d2 ) as ṙ. Since the determinant is
−c d
positive, we always have that N is equal to (d, −c). That will give
us N being in the same direction as n.
Now assume we have a vector field F : R2 7→ R2 . We can use
T and n to decompose F on C into two parts: one parallel to T and
one parallel to n. If we write

F = αT + βn

for some quantities α and β, then these can be found by taking the Note that α and β depend on t!
scalar product with T and n, respectively. If we take the first of
these then

F · T = α(T · T) + β(n · T) =⇒ α = F · T = F(r(t)) · T(t)

since T · T = 1 (unit vector) and T · n = 0. Similarly it follows that


β = F · n = F(r(t)) · n(t). We deduce that the scalar product F · T
gives us the magnitude of the vector field tangential to the path
while F · n represents the magnitude of F normal to the path.

C T(t)

r (t) n (t)

Imagine we have a bead threaded on a wire curve C and that a


force F acts on the bead. In general this force can be resolved into
two components: one is tangential to the wire and tends to push
the bead along it while the other is normal to the wire but does not
move the bead as the bead cannot come off the wire! If we want
to find the total work done by the force in moving the bead along
C then it is not the total force that is of importance, but rather the
74 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

component of the force tangential to the wire. We can see that in


this instance we would need to integrate the tangential component
of F along the curve C:

Z Z b Z b
F · T ds = F · T |ṙ| dt = F · ṙ dt.
C a a

Sometimes it is of interest to know the integral of the normal com-


ponent of F along the path:

Z Z b Z b
F · n ds = F · n |ṙ| dt = F · N dt.
C a a

Note that we used here the definition of path integral seen in Chap-
ter 4.

Definition 5.2. (Circulation)


Given a vector field F in R2 and some curve C with a parametric repre- Note that we can define circulation in
R3 exactly the same way.
sentation r(t) for t ∈ [ a, b], the circulation of F along C is
Z Z b
F · dr := F(r(t)) · ṙ(t) dt.
C a

If C is a closed curve, we make the convention of parameterising C in the


anti-clockwise direction.

Definition 5.3. (Flux across a curve in R2 )


Note that this definition is NOT valid
Given a vector field F in R2 and some curve C with a parametric repre-
in R3 .
sentation r(t) for t ∈ [ a, b], the flux of F across C is
Z Z b
F · dn := F(r(t)) · N(t) dt,
C a

where N(t) = (d(t), −c(t)) if ṙ(t) = (c(t), d(t)). If C is a closed


curve, we make the convention of parameterising C in the anti-clockwise
direction, so that N points outwards.

Note that the flux is also sometimes called the flow of F.

Example 5.4. Consider the closed curve given by the semi-circle x2 +


y2 = 1, y > 0 together with y = 0 for −1 ≤ x ≤ 1. Calculate the
circulation along C and the flux across C of the vector field
! !
x 2 x+y
F = .
y π −x
mathematical methods 2 75

1.2

0.8

0.6

0.4

0.2

−0.2
−1 −0.8 −0.6 −0.4 −0.2 0 0.2 0.4 0.6 0.8 1

Solution: Our first task is to parameterise the curve C. It is clearly made


up from two parts so our parameterisations will also need two forms.
Furthermore we need to traverse C in an anti-clockwise sense so let us
divide it into the semi-circle (call it A) and the part of the x-axis with
−1 ≤ x ≤ 1 (call this B). We also need to ensure that we move in an anti-
clockwise sense so our parameterisation must move along the semi-circle A
from the point (1, 0) to the point (−1, 0) and then back along the straight
line B from (−1, 0) to (1, 0). With these noted, it is clear that a suitable
parameterisation for the entire curve  C is 
cos t
Path A: r(t) = , 0≤t≤π
sin t
 
t
Path B: r(t) = , −1 ≤ t ≤ 1.
0
We now compute ṙ and N for each path.
− sin t
   
cos t
Path A: ṙ(t) = N(t) =
cos t sin t
   
1 0
Path B: ṙ(t) = N(t) = .
0 −1
For both paths, it is easy to check that ṙ has magnitude one (so is the unit
tangent T), that N is orthogonal to ṙ with the same magnitude and that
N(t) = (d(t), −c(t)) if ṙ(t) = (c(t), d(t)). Since C is closed, it is
simple to check that for both parts A and B the vector N is indeed outward
pointing.
We are now in a position to calculate both the flux across C and the
circulation around it. For the flux
Z Z π
Path A: F · dn = F(r(t)) · N(t) dt
A 0
Z    
2 π cos t + sin t cos t
= · dt
π 0 − cos t sin t
2 π
Z
= cos2 t dt
π 0
 π
1 π 1 sin(2t)
Z
= cos(2t) + 1 dt = + t = 1.
π 0 π 2 0
Z Z 1
Path B: F · dn = F(r(t)) · N(t) dt
B −1
2 1 t+0
Z     Z 1
0 2
= · dt = t dt = 0.
π −1 −t −1 π −1
For the total flux we have to sum up the two results and get
Z Z Z
F · dn = F · dn + F · dn = 1.
C A B
76 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

For the circulation along C:


Z Z π
Path A: F · dr = F(r(t)) · ṙ(t) dt
A 0
− sin t
Z    
2 π cos t + sin t
= · dt
π 0 − cos t cos t
2 π
Z 
= − cos t sin t − sin2 t − cos2 t dt
π 0
2 π
Z
=− (cos t sin t + 1) dt
π 0
" #π
2 sin2 t
=− + t = −2.
π 2
0
Z 1
2 1 t+0
Z    
1
Z
Path B: F · dr = F(r(t)) · ṙ(t) dt = · ds
B −1 π −1 − t 0
2 1
Z
= t dt = 0.
π −1
And therefore the total circulation is
Z Z Z
F · dr = F · dr + F · dr = −2.
C A B

5.2 Flux through a surface in R3

We will now generalise the above method to calculate the flux of


a vector field across a surface, in R3 . As an example one might
suppose that F describes the velocity of a particular flow of wa-
ter. In this case, the flux of F admits a particularly easy physical
interpretation for the flux across a given surface S is a measure of
how much water flows through the surface. If S is a closed surface,
like a sphere, then a positive flux means that the net flow of water
across S is outward – that is more water is leaving the volume con-
tained by S than is entering it. Similarly a negative flux means there
is a net flow into the interior of S.
In order to compute the flux through a surface S we must first
propose a suitable parameterisation for S – remember that it takes
two parameters to describe a surface. Consequently suppose that S
is given by a parameterisation S(u, v) for (u, v) ∈ D; at any point
on S there are two tangent vectors Su and Sv . We say that a vector
v is normal to the surface if it is orthogonal to both Su and Sv ; we
know from its very definition that N = Su × Sv satisfies these
requirements. We need a unit vector n normal to the surface but, of
course there are two possibilities depending on the orientation of
the vector. The standard convention is to simply take n = N/|N|.
If S is a closed2 surface and is appropriately parameterised, then n 2
A closed surface is a surface that is the
will point outwards (if it does not, we can just interchange the roles boundary of a solid region of R3 .

of u and v so that it does).


To compute the flux of a vector field F through a surface we need
to know its component normal to the surface which is simply F · n.
It follows that the total flux is
ZZ
F · n dS
S
mathematical methods 2 77

which is a surface integral (see previous chapter) equal to


ZZ ZZ
F · n |N| dudv = F · N dudv.
D D

Definition 5.5. (Flux across a surface in R3 )


Given a vector field F of R3 and some surface S with a parametric rep-
resentation S(u, v) for (u, v) ∈ D, the flux across S (in the standard
direction of the unit normal vector n) is
Z ZZ
F · dS := F · N dudv.
S D

If C is a closed surface, we make sure that N points outwards (by inter-


changing the roles of u and v if needed ).

1
Example 5.6. Calculate the flux of F( x, y, z) = 3 ( x , y , z) through the
surface
S(u, v) = (v, (1 − v2 ) cos u, (1 − v2 ) sin u) 0 ≤ u < 2π, −1 < v < 1 .
Solution: We can compute
 
e1 e2 e3
∂S ∂S
N= × = det  0 −(1 − v2 ) sin u (1 − v2 ) cos u 
 
∂u ∂v
1 −2v cos u −2v sin u
= (1 − v2 )(2v, cos u, sin u).
We can check that N is indeed pointing outwards. In this example it is
probably easiest to proceed by looking at the point ( x, y, z) = (0, 1, 0)
which corresponds to the parameter values u = 0 and v = 0 and for which
N = (0, 1, 0). It is also clear that the surface must lie in −1 ≤ y ≤ 1
so the outward normal ought to have a positive y-component, which is the
case.
Now we can calculate the flux
Z ZZ
F · dS = F · N dudv
S D
1 1 2π
Z Z
= (v, (1 − v2 ) cos u, (1 − v2 ) sin u) · (1 − v2 )(2v, cos u, sin u)dudv
3 v=−1 u=0
1 1
Z 2π
2π 1 16π
Z Z
= (1 + v2 )(1 − v2 )dudv = (1 − v4 )dv = .
3 v=−1 u=0 3 −1 15

5.3 Work done by the gradient of a scalar field

A frequent application for path integrals over vector fields is to


calculate the work done by a vector field. Here we focus on a very
special field: the gradient field of a scalar function.
Definition 5.7. The gradient of a scalar function f ( x, y, z) is a vector
field:   
  ∂ f ( x,y,z) 
x ∂x ∂ x f ( x, y, z)
 ∂ f ( x,y,z) 

∇ f  y  :=  ∂y  =  ∂y f ( x, y, z)  .
   
z ∂ f ( x,y,z) ∂z f ( x, y, z)
∂z
78 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

We should remember that work is force multiplied by distance.


In applications, force fields often arise as gradient fields of some
scalar functions called the potential.
From Section 5.1 we can see that path integrals of vector fields
can be used to calculate the work done by a vector field. If we have
a path C (given by a parametrisation r) through some vector field
F, that is a force field, only the component of F parallel to T will do
some work. Therefore we have

Theorem 5.8. Let r describe the path of a particle moving through a force
field F. The work done is the circulation of F along C:
Z
W= F · dr.
C

Example 5.9. Let a particle move from the origin A = (0, 0, 0) to


B = (1, 1, 1) on a straight line. A force field has potential f ( x, y, z) =
x3 y + xyz + 2x + y + z2 + c where c is a constant. Calculate the work
done by the force.
Solution: First we need to calculate the force field by calculating the
gradient:

∇ f = (∂ x f , ∂y f , ∂z f ) = (3x2 y + yz + 2, x3 + xz + 1, xy + 2z).

So our force field is F( x, y, z) = (3x2 y + yz + 2, x3 + xz + 1, xy + 2z).


The path can be parameterised by r(t) = (t, t, t) where 0 ≤ t ≤ 1.
Therefore the velocity is ṙ = (1, 1, 1) and we get for the work:
Z Z 1
W= F · dr = F(t, t, t) · (1, 1, 1) dt
C 0
Z 1
= (3t3 + t2 + 2, t3 + t2 + 1, t2 + 2t) · (1, 1, 1) dt
0
Z 1
= 4t3 + 3t2 + 2t + 3 dt = 6.
0

Note that this example used the definition of circulation in R3 .


6
Conservative vector fields

In the previous chapter we introduced vector fields and performed


some relatively simple calculations using them. However within
the family of vector fields there are some members which possess a
very particular and important property. These may be illustrated by
re-visiting the final example of Chapter 5 in which we saw that the
work done by the force

F( x, y, z) = (3x2 y + yz + 2, x3 + xz + 1, xy + 2z)

in moving a particle from the origin A = (0, 0, 0) to the point


B = (1, 1, 1) along the straight line AB is 6 units.
Now suppose we try to move the particle from A to B again,
and under the same force, but this time take a slightly less direct
path. For the sake of argument, let us progress from A to B by first
taking the straight line to A1 = (1, 0, 0), then the straight line to
A2 = (1, 1, 0) and then finally to B. Notice that, rather than moving
the particle along the diagonal of the unit cube in one move, this
more circuitous route takes the particle first along the x-axis, then
parallel to the y-axis and then lastly parallel to the z-axis.
We can evaluate the work done along each of the three legs of
the journey using, what should be by now, a fairly standard proce-
dure. To remind ourselves of the argument involved:
Step 1. The first path from A to A1 is parameterised by r(t) =
(t, 0, 0) where 0 ≤ t ≤ 1 so the velocity is ṙ = (1, 0, 0). The work
done is then
Z Z 1 Z 1
F · dr = F(t, 0, 0) · (1, 0, 0) dt = (2, t3 + 1, 0) · (1, 0, 0) dt = 2.
AA1 0 0

Step 2. The second path from A1 to A2 is parameterised by


r(t) = (1, t, 0) where 0 ≤ t ≤ 1 so the velocity is ṙ = (0, 1, 0).
The work done
Z Z 1 Z 1
F · dr = F(1, t, 0) · (0, 1, 0) dt = (3t + 2, 2, t) · (0, 1, 0) dt = 2.
A1 A2 0 0

Step 3. The third path from A2 to B is parameterised by r(t) =


(1, 1, t) where 0 ≤ t ≤ 1 so the velocity is ṙ = (0, 0, 1). The work
80 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

done
Z Z 1 Z 1
F · dr = F(1, 1, t) · (0, 0, 1) dt = (5 + t, 2 + t, 1 + 2t) · (0, 0, 1) dt
A2 B 0 0

= [t + t2 ]10 = 2.

Notice that the sum of these three parts is 6; exactly the same as
we calculated for the direct path from A to B. (Incidentally, note
that the fact that the three results for the individual steps are each
equal to two is pure coincidence; in general we would expect the
work done on the various parts of the path to be different.)
The fact that the results for the two routes taken from A to be B
are the same is not fortuitous for F( x, y, z) is an example of what is
known as a conservative field. As an exercise, design some other path
from A to B and then compute the total work done by F in moving
from A to B. You should get the same answer 6 irrespective of how
simple or complicated your choice of path is (though, of course, if
you choose a very intricate way of moving from A to B you will
have a lot of calculations to do and are more likely to make a mis-
take!). The conclusion we draw is that it does not matter how one
moves between A and B the work done is the same independent
of the path traversed. However, the work done does depend on the
endpoints A and B; if one (or both) of them moves then the work
done is likely to change.
We can encapsulate these ideas in the following definition:

Definition 6.1. (Conservative vector fields)


A vector field is called conservative if the path integral (circulation)
R
C F · dr from a starting point S to a finishing point F depends on the
locations of S and F but not on the particular path C taken between them.

Conservative fields occur frequently in physical and engineering


applications. One familiar example is the gravitational field; we
know that the potential energy of a particle is given by mgh, where
m is the mass, g is gravity and h is the height of the particle above
(or below) some reference level. If we lift an object from A to B the
change in the potential energy depends only on the difference in
heights of A and B and not on the details of the path taken by the
object to get from A to B.
In view of the physical significance of conservative fields it is
of some importance to recognise when a given vector field is (or is
not) conservative. This is the issue we tackle next.

6.1 Which vector fields are conservative?

In order to determine whether a vector field might be conservative


let us start with a special family which consists of fields that can be
mathematical methods 2 81

expressed as the gradient of some scalar function, i.e.

F = ∇ψ

path by r(t) for t ∈ [ a, b] so that t = a corresponds to the start point


S and t = b to the end point F then
Z Z b Z b
F · dr = F(r(t)) · ṙ(t) dt = ∇ψ(r(t)) · ṙ(t) dt
C a a
Z b 
∂ψ ∂ψ ∂ψ
= (r(t))ṙ1 (t) + (r(t))ṙ2 (t) + (r(t))ṙ3 (t) dt
a ∂x ∂y ∂z
Z b

= (r(t)) dt
a dt
by the chain rule. Now we see from the Fundamental theorem of
calculus that this last integral is simply

ψ (r(b)) − ψ (r( a)) = ψ( F ) − ψ(S).

We see therefore that the circulation is


Z
F · dr = ψ( F ) − ψ(S)
C

or, in other words, its value depends only on the locations of the
endpoints S and F. Note that in working out this integral we have
not had to say anything about the details of the path C, except
for the very weak assumption that we can parameterise it in an
appropriate way. Therefore the vector field F satisfies our earlier
definition of a conservative field that the circulation depends only
on the choice of endpoints and is independent of the path taken
between the two.
So what do we know now? Well, we can be assured that any
field that can be written as the gradient of a scalar function is con-
servative. What we do not yet know is whether there are any other
conservative fields; i.e. are there conservative fields that cannot be
expressed as the gradient of a suitable scalar function?
So suppose that some vector field F is known to be conservative
and suppose

F( x, y, z) = ( f 1 ( x, y, z), f 2 ( x, y, z), f 3 ( x, y, z)).

Then define a function


Z
ψ( x, y, z) = F · dr
C

where C is any path from the origin 0 to the point ( x, y, z); this is
well defined since F is conservative. Notice ψ is a scalar function
(from R3 to R) since the circulation of a field along a curve is a real
number.
Then, for a fixed point ( x, y, z) and a fixed (small) real number
∆x, Z
ψ( x + ∆x, y, z) − ψ( x, y, z) = F · dr
C0
82 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

where C 0 is a path that joins ( x, y, z) to ( x + ∆x, y, z). In order to


evaluate this integral it is simplest to parameterise C 0 so that on this
curve
r(t) = ( x + t∆x, y, z) , 0 ≤ t ≤ 1;
it is clear that t = 0 and t = 1 give the two end-points ( x, y, z) and
( x + ∆x, y, z) as required. Now ṙ(t) = (∆x, 0, 0) so that
Z 1
ψ( x + ∆x, y, z) − ψ( x, y, z) = F(r(t)) · (∆x, 0, 0) dt
0
Z 1
= ( f 1 (r(t)), f 2 (r(t)), f 3 (r(t))) · (∆x, 0, 0) dt
0
Z 1
= f 1 ( x + t∆x, y, z)∆x dt
0
so that
ψ( x + ∆x, y, z) − ψ( x, y, z)
Z 1
= f 1 ( x + t∆x, y, z) dt.
∆x 0
Now by definition of partial derivatives we know that
∂ψ ψ( x + ∆x, y, z) − ψ( x, y, z)
( x, y, z) = lim
∂x ∆x →0 ∆x
Z 1
= lim f 1 ( x + t∆x, y, z) dt
∆x →0 0
Z 1
= f 1 ( x, y, z) dt = f 1 ( x, y, z).
0
We can repeat this argument twice more, once with a curve
joining ( x, y, z) to ( x, y + ∆y, z) and again with one joining ( x, y, z)
to ( x, y, z + ∆z). We won’t go through all the details again because
it should be fairly clear how the conclusions are modified. We find
that
∂ψ ∂ψ
= f2 and = f3 .
∂y ∂z
So if we combine these results we see that
 
∂ψ ∂ψ ∂ψ
( f1 , f2 , f3 ) = , , = ∇ψ.
∂x ∂y ∂z
Hence we conclude that if F is a conservative vector field then it can
be written as the gradient of some scalar function ψ known as the
potential.
To summarise, we have proved the following theorem.

Theorem 6.2. Let F be a vector field in an open subset U of R3 or


R2 .

(a) If F = ∇Ψ for some scalar function Ψ, then F is conservative in


U. In this case, for any curve C in U we have
Z
F · dr = Ψ( E) − Ψ(S)
C

where S and E are the starting and ending points of C.

(b) If F conservative in U, then F = ∇Ψ for some scalar function Ψ.


mathematical methods 2 83

Ψ is called a potential for F.


Recall that a subset U of R3 (or R2 ) is called open if for every x ∈
U there exists a ball (reps. disk) with centre x contained entirely in
U.
Thus, according to the above theorem, a vector field F is con-
servative if and only if F = ∇ψ for some scalar function ψ (the
R
potential). Moreover, in this case, C F · dr = ψ( F ) − ψ(S) where S
and F are the starting and finishing points of C.

Corollary 6.3. If C is a closed curve and F is conservative, then


Z
F · dr = 0.
C

This follows immediately from the fact that S = F for a closed


curve.

Remark 6.4. Given that F = ∇ψ for a conservative field, and


that the gradient of a constant is zero, if we have one potential
then we can always add a constant to it and we will still have a
potential. The value to be assigned to the constant is sometimes
arbitrary and sometimes suggested by the context. For instance, if
we have the potential corresponding to the electric field around a
collection of charged particles, we might well choose the constant
so that the potential tends to zero far from the charges. Often when
measuring the potential energy of an object above the earth we pick
the constant so that the potential is zero at sea level.

Example 6.5. Find the work done in moving a particle of mass m from
(1, 1, 1) along a straight line C to (2, 2, 2) by the gravity field

( x, y, z)
F( x, y, z) = − GmM ,
( x2 + y2 + z2 )3/2
where G is the gravitational constant and M the mass of a second body
located at the origin.
Solution: First we need to check if the field is conservative. We note that
!
∂ 1 x
p =− 2 2 + z2 )3/2
∂x 2
x +y +z2 2 ( x + y

so that the first component of F,


!
x ∂ 1
− GmM 2 2 2 3/2
= GmM p .
(x + y + z ) ∂x x 2 + y2 + z2

The same pattern continues for the second and third components of F so
that we have
GmM
F = ∇ψ where ψ( x, y, z) = p .
x2 + y2 + z2

Now we have identified a potential we know by Theorem 6.2 that F is


conservative. Evaluating the work done is then straightforward for all we
84 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

need to do is find the difference in the values of ψ at the ends of the line C
so that the work done is just

GmM GmM GmM


ψ(2, 2, 2) − ψ(1, 1, 1) = √ − √ =− √ .
2 3 3 2 3

A key step in the last example was to identify the potential ψ


corresponding to the vector field F. Once we had ψ the remain-
der of the problem was really quite elementary. This then poses
the next question: if we are given a particular vector field is there
an easy way to determine whether it is conservative or not? Fur-
thermore, if it is conservative, is there a systematic way to find the
potential?

6.2 Finding potentials

In order to test whether a given vector field is conservative, it is


necessary to introduce the concept of the curl of a vector field.

Definition 6.6. (Curl of a vector field)


The curl of a vector field F = ( f 1 , f 2 , f 3 ) is the vector field

( ∂ y f 3 − ∂ z f 2 , ∂ z f 1 − ∂ x f 3 , ∂ x f 2 − ∂ y f 1 ).

We denote it by ∇ × F.

The notation ∇ × F suggests the notion of a vector cross-product.


In fact, treating the curl in this way as a ‘cross-product’ of the op-
erator ∇ = (∂ x , ∂y , ∂z ) and the vector F = ( f 1 , f 2 , f 3 ) is a con-
venient way of remembering the formula for the curl of a vector
field although ∇ is certainly not a vector in the conventional sense.
Indeed, using a (formal) expansion along the first row in the follow-
ing determinant, we get
 
e1 e2 e3
∇ × F = det ∂ x ∂ z  = ( ∂ y f 3 − ∂ z f 2 , ∂ z f 1 − ∂ x f 3 , ∂ x f 2 − ∂ y f 1 ).
 
∂y
f1 f2 f3

Here e1 = (1, 0, 0), e2 = (0, 1, 0) and e3 = (0, 0, 1).

Remark. If F( x, y) = ( P( x, y), Q( x, y)) is a vector field in the plane,


then F can be regarded as a vector field in R3 by setting

F( x, y, z) = ( P( x, y), Q( x, y), 0).

It then follows from above that

∇ × F = (0, 0, Q x − Py ).
mathematical methods 2 85

Example 6.7. Find the curl of the vector field F = ( xz, z2 , e xy ).


Solution: With f 1 = xz, f 2 = z2 , f 3 = e xy we have the three components
of the curl:

∂y f 3 − ∂z f 2 = ∂y (e xy ) − ∂z (z2 ) = xe xy − 2z,
∂z f 1 − ∂ x f 3 = ∂z ( xy) − ∂ x (e xy ) = 0 − ye xy = −ye xy ,
∂ x f 2 − ∂y f 1 = ∂ x (z2 ) − ∂y ( xz) = 0 − 0 = 0.

Hence ∇ × F = ( xe xy − 2z, −ye xy , 0).

Finding the curl of a vector field is the easiest way to determine


whether it is conservative. This is because we have the following
result:

Theorem 6.8. If a vector field F is conservative then ∇ × F = 0.

Proof. If F is a conservative vector field there exist a potential ψ


such that F = ∇ψ. The curl of F is therefore:
   
∂x ∂x ψ
∇ × F = ∇ × ∇ψ =  ∂y  ×  ∂y ψ  =
   
∂z ∂z ψ
 
∂y ∂z ψ − ∂z ∂y ψ
=  ∂z ∂ x ψ − ∂ x ∂z ψ  = 0.
 
∂ x ∂y ψ − ∂y ∂ x ψ

Note carefully what this statement says (and, perhaps more im-
portantly, what it does not). It tells us that the curl of a conservative
vector field is zero but not the converse; the argument above does
not allow us to conclude that if the curl of a vector field is zero then
the field is conservative. In fact this statement is true for a certain
kind of solids (regions in R2 ), but the proof is relatively long and
complicated so is not part of this unit.

Definition 6.9. (Simply-connected sets)


A subset U of R3 (or R2 ) is called simply-connected if U has no ‘holes’.
(Strictly speaking for sets U ⊂ R2 this means every closed curve in
U can be continuously contracted to a point without leaving U. For sets
U ⊂ R3 the condition is that every closed surface in U can be continu-
ously contracted to a point without leaving U.)

Theorem 6.10. Let F be a continuously differentiable vector field in


an open subset U of R3 .
86 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

1. If F is conservative, then ∇ × F = 0.

2. If U is simply-connected and ∇ × F = 0, then F is a conservative


vector field in U.

This then provides us a simple test as to whether a vector field


is conservative or not. All we have to do is to calculate its curl; if
the curl vanishes (that is, is equal to 0) and the domain is simply-
connected, then we have a conservative field. If the curl is non-zero
then the field is definitely not conservative.
Example 6.11. Show that F = (e2y , 1 + 2xe2y , 0) is a conservative vector
field.
Solution: F is defined in the whole of R3 which is simply-connected. We
need to show that ( g1 , g2 , g3 ) = ∇ × F = 0.
From the formulae above

g1 = ∂y (0) − ∂z (1 + 2xe2y ) = 0,
g2 = ∂z (e2y ) − ∂ x (0) = 0 − ye xy = 0
g3 = ∂ x (1 + 2xe2y ) − ∂y (e2y ) = 2e2y − 2e2y = 0.

Hence ∇ × F = 0 so F is indeed a conservative vector field.


We can now use Theorem 6.10 to derive a similar result for vec-
tor fields in the plane.

Theorem 6.12. Let F( x, y) = ( P( x, y), Q( x, y)) be a continuously


differentiable vector field in an open subset U of R2 .

1. If F is conservative, then Q x = Py in U.

2. If U is simply-connected and Q x = Py in U, then F is a conserva-


tive vector field in U.

This follows from Theorem 6.10 and the fact that ∇ × F =


(0, 0, Q x − Py ).
So now we have a criteria for determining whether a given vec-
tor field F is conservative, we next need to develop a method to
find the corresponding potential ψ. By Theorem 6.2 we know that
F = ∇ψ so that if F = ( f 1 , f 2 , f 3 ) then
∂ψ ∂ψ ∂ψ
f1 = , f2 = , and f3 = .
∂x ∂y ∂z
We now integrate these equations systematically one at a time. Let
us start with the first equation (although we can actually begin with
any of them). If we integrate with respect to x we will have that
Z
ψ= f 1 dx + g(y, z); (♣)
mathematical methods 2 87

in particular note the presence of the arbitrary function g(y, z). We


know that, when we do an indefinite integral expressed in terms of
a single variable, we should add an arbitrary constant (the familiar
‘+C’) because the derivative of this constant vanishes. The function
g(y, z) is the analogue when we integrate an expression containing
several variables. One way to look at this is to differentiate equation
(♣) with respect to x for we recover f 1 = ∂ψ/∂x whatever the form
of g(y, z).
We thus conclude that
Z
ψ= f 1 dx + g(y, z)

for some apparently arbitrary function g(y, z). However we are not
at liberty to choose g(y, z) as we wish, because we are yet to satisfy
the other two equations
∂ψ ∂ψ
f2 = and f3 = .
∂y ∂z
Let us take the first of these. If we substitute in (♣) then
Z 
f 2 = ∂y f 1 dx + ∂y g(y, z).

This tells us the form of ∂y g(y, z) so if we integrate this with respect


to y then we have g(y, z) = ĝ(y, z) + h(z) for some known function
ĝ(y, z) and arbitrary function h(z). Then we have
Z
ψ= f 1 dx + ĝ(y, z) + h(z).

To find h(z) we need to ensure that the third (and last) equation
f 3 = ∂ψ/∂z holds. Substituting in the form of ψ leads to an equa-
tion for h0 (z) which can be integrated to find h(z) within an arbi-
trary constant. Remember (Remark 6.4) that a potential can only be
tied down within a constant.
The formulae above are somewhat involved (and need not be
remembered). In practice the process of finding potentials is less
daunting as long as it is done in a systematic manner. Let us illus-
trate the method with a couple of examples.

Example 6.13. Show that the vector field F = ( x, ey sin z, ey cos z) is


conservative and find its potential.
Solution: We begin by calculating ( g1 , g2 , g3 ) = ∇ × F to ensure it is
zero. Now

g1 = ∂y (ey cos z) − ∂z (esin z ) = ey cos z − ey cos z = 0,


g2 = ∂z ( x ) − ∂ x (ey cos z) = 0
g3 = ∂ x (ey sin z) − ∂y ( x ) = 0.

Hence ∇ × F = 0 so F is indeed a conservative vector field.


To find the potential function ψ( x, y, z) such that F = ∇ψ we need to
solve
∂ψ ∂ψ ∂ψ
= x, = ey sin z, and = ey cos z.
∂x ∂y ∂z
88 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

The first of these implies that ψ( x, y, z) = x2 /2 + g(y, z) for some


function g. Substituting this in the second equation then implies that
Z
∂ψ ∂g
= = ey sin z ⇒ g(y, z) = ey sin z dy = ey sin z + h(z)
∂y ∂y

for some h(z). Hence

x2 x2
ψ( x, y, z) = + g(y, z) = + ey sin z + h(z). (♦)
2 2
Now we turn to the last equation ∂ψ/∂z = ey cos z. With ψ given by (♦)
it follows that

ey cos z + h0 (z) = ey cos z ⇒ h0 (z) = 0 ⇒ h(z) = C,

a constant. Given that this potential does not arise from any particular
situation we are free to choose the value of C and the conventional choice is
to take C = 0. Then h(z) = 0 and (♦) gives the final potential as

x2
ψ( x, y, z) = + ey sin z.
2
Example 6.14. Let F = (ye xy , xe xy + z, y + 2) be a force field. Show that
F is conservative and find the corresponding potential function φ( x, y, z).
Find the work done in moving a particle from A(0, 0, 1) to B(1, 1, 1).
Verify this result by calculating the work done by F in moving the particle
along the straight-line path between A and B.
Solution: First we need to check that ( g1 , g2 , g3 ) = ∇ × F is zero. As

g1 = ∂y (y + 2) − ∂z ( xe xy + z) = 1 − 1 = 0,
g2 = ∂z (ye xy ) − ∂ x (y + 2) = 0
g3 = ∂ x ( xe xy + z) − ∂y (ye xy ) = (1 + xy)e xy − (1 + xy)e xy = 0,

∇ × F = 0 and so F is indeed a conservative force field.


To find the potential φ we need to solve the three equations

∂φ ∂φ ∂φ
= ye xy = xe xy + z and = y + 2.
∂x ∂y ∂z

In developing the theory and in the previous example we integrated the


three equations by starting with the one for ∂ x φ, then the one for ∂y φ and
finally the one involving ∂z φ. There is actually no need to stick to this
order though, and in some cases it is actually easier to integrate the system
by proceeding in some other sequence. To illustrate an alternative order of
integration, let us start this time with ∂φ/∂z = y + 2. Integrating with
respect to z gives
Z
φ= (y + 2) dz ⇒ φ = (y + 2)z + g( x, y);

note that integrating with respect to z means that our arbitrary function
depends on x and y. If we substitute this form into ∂φ/∂y = xe xy + z
then
∂g ∂g
z+ = xe xy + z ⇒ = xe xy
∂y ∂y
mathematical methods 2 89

which integrates so that


Z
g( x, y) = xe xy dy = e xy + h( x )

for some function h( x ). Hence

φ = (y + 2)z + e xy + h( x ).

Now we turn to the last equation ∂φ/∂x = ye xy which becomes

∂φ dh dh
= ye xy + = ye xy ⇒ =0
∂x dx dx
so h( x ) is a constant that we choose to be zero. Hence

φ = (y + 2)z + e xy

is a suitable potential.
To find the work done in moving the particle from A(0, 0, 1) to B(1, 1, 1)
R
we can use the result that the circulation C F · dr of a conservative field is
just the difference in the values of the potential at the endpoints of C. Thus
in this case the work done is simply

φ(1, 1, 1) − φ(0, 0, 1) = (3 + e) − (2 + 1) = e.

To check this result by direct integration we need to parameterise the


straight-line path between A and B. This is given by r(t) = (t, t, 1) where
0 ≤ t ≤ 1 so the velocity is ṙ = (1, 1, 0). The work done is then
Z Z 1 Z 1
2 2
F · dr = F(t, t, 1) · (1, 1, 0) dt = (tet , tet + 1, t + 2) · (1, 1, 0) dt
C 0 0
Z 1 h i1
2 2
= (2tet + 1) dt = et + t = e + 1 − 1 = e.
0 0

We see that the result is the same, as it should be of course. No-


tice though that the direct calculation is somewhat tricky in this
example and that evaluating the work done using the properties of
the potential is far easier.
7
Fourier series

We now go back to studying functions from R to R. In this chapter,


we focus on periodic functions.

7.1 Introduction

Because waves and vibrations tend to have a periodic structure,


that is, they repeat their basic shape in time or space, it is often
convenient to approximate a periodic function by a linear combi-
nation of perfect waves (the sine and cosine functions, which arise
in simple harmonic motion). Decomposing a general periodic func-
tion into a sum of trigonometrical functions is sometimes called
harmonic analysis. This process is often necessary because most
physical waves and vibrations do not have a perfect single sine or
cosine form but rather they are made up of a number of different
harmonics of the underlying system.
In order to develop the ideas required we need to formalise
the process. For convenience we shall assume for now that the
underlying periodicity (whether it be in space or time) is of length
2π. This eases the subsequent manipulation somewhat but turns
out not to be unduly restrictive; if in practice our function has a
periodicity of length different to 2π it is relatively easy to scale
our results so as to account for this change. However, it is worth
quickly reminding ourselves of the definition of the period of a
function.

Definition 7.1. (Period of a function)


We say that a function f from R to R has period P if f (t + P) = f (t)
for all t in R. In this case, we also say that f is P-periodic.

Periodic function

Graphically it is generally not difficult to identify a periodic 1

function for its sketch takes the form of a curve that clearly repeats
itself after an interval of length P. The rather peculiar function on
the right possesses discontinuities at t = (2n + 1)π for n ∈ Z but 0

nevertheless has a period 2π.


Recall that both sin(nt) and cos(nt) are 2π-periodic functions for
−1

−3π −2π −π 0 π 2π 3π
92 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

positive integer values of n. Note that the smallest period of sin(nt)


and cos(nt) is actually 2π/n, n ≥ 1.

cos(0*t), cos(t), cos(2t), cos(3t) Figure 7.1: Graphs of cos(nt) and


sin(nt) for n = 0, 1, 2, 3.
1

−1

−π −π/2 0 π/2 π

sin(t), sin(2t), sin(3t)

−1

−π −π/2 0 π/2 π

Our hope then is to approximate a general 2π-periodic function


f (t) in the form
a0
S N f (t) = + a1 cos(t) + b1 sin(t) + a2 cos(2t) + b2 sin(2t) + . . . + a N cos( Nt) + b N sin( Nt)
2
N
a
= 0 + ∑ ( an cos(nt) + bn sin(nt));
2 n =1

that is we aim to approximate f (t) by a finite sum of sines and


cosines of the form sin(nt) and cos(nt) for integers 0 ≤ n ≤ N.
If our approximation is well-behaved we should expect that it im-
proves as N goes to infinity.
The infinite sum
∞ There is no need to consider n negative
a0
FS f (t) = lim S N f (t) = + ∑ ( an cos(nt) + bn sin(nt)) (7.1) because sin(−nt) = − sin(nt) and
N →∞ 2 n =1 cos(−nt) = cos(nt).

is called the Fourier series expansion of f where the constants a0 ,


an , bn (n = 1, 2, 3, . . .) are known as the Fourier coefficients. The
various S N f (t) functions are called partial sums of the Fourier series
expansion.
mathematical methods 2 93

We will explain in the next section how to compute the Fourier


coefficients. The following example illustrates how the approxima-
tions get better when N increases.

Example 7.2. Consider the piecewise 2π-periodic function



1, −π < t ≤ 0,
f (t) = and f (t + 2π ) = f (t).
−1, 0 < t ≤ π,
Step function

−1

−3π −2π −π 0 π 2π 3π

We will see in Example 7.12 that


N
4 sin(nt)
S N f (t) = −
π ∑ n
.
n=1,n odd

We draw the Fourier sums for different values of N.

partial sum N=11 partial sum N=21 partial sum N=101 Figure 7.2: Graph of S N f for different
values of N, on [−π, π ].
1 1 1

0 0 0

−1 −1 −1

−π 0 π −π 0 π −π 0 π

Notice how the approximation improves as we increase N (so we are


retaining more terms in the partial sum S N ( f )). We know that the true
function is equal to −1 for 0 < t ≤ π and the the approximation S N ( f )
oscillates about this value. When N = 11 this oscillation is quite notice-
able and relatively large; by the time N = 101 it is far less pronounced.
Notice also how the approximation S N f is relatively good away from dis-
continuities in the function f (t) but poorer as these points are approached.
As an example of this look at the N = 11 result; it is clear that oscillations
in the approximating function are small around t = π/2 but increase as
94 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

either t → 0 or t → π. This is a well-known behaviour known as Gibbs’


phenomena which tends to occur when the function f (t) has points of
discontinuity.

7.2 Calculation of the Fourier coefficients

Recall that the Fourier series representation of f (t) is


a0
FS f (t) = + a1 cos(t) + b1 sin(t) + a2 cos(2t) + b2 sin(2t) + · · ·
2
(7.1)
(Note at this juncture it seems a little strange to define the constant
term as a0 /2 rather than simply a0 . We will see why this is done
presently.)
We need a method of determining the values of the coefficients
a0 , an , bn , for n = 1, 2, 3, . . ., so that FS f (t) converges (if possible) to
f (t). To find a0 we simply integrate both sides from −π to π to get

a0
Z π Z π
FS f (t)dt = dt = πa0
−π −π 2

because
Z π Z π
cos(nt) dt = 0 and sin(nt) dt = 0.
−π −π

So we set
1
Z π
a0 = f (t)dt. (7.2a)
π −π

This says that a0 is twice the average value of f (t), or equiva-


lently, the zeroth order term a0 /2 is the average value of the func-
tion f (t).
We can show by direct evaluation of the integrals that
Z π
sin(mt) sin(nt)dt = 0, for any integers m, n, with m 6= n,
−π
Z π
sin(mt) cos(nt)dt = 0, for any integers m, n,
−π
Z π
cos(mt) cos(nt)dt = 0, for any integers m, n, with m 6= n,
−π
Z π Z π
sin2 (nt)dt = π and cos2 (nt)dt = π
−π −π
for any integer n. It is an exercise to verify these statements, using
trigonometric formulae such as

sin( A + B) + sin( A − B)
sin( A) cos( B) = ,
2
or using integration by parts (twice).
This allows us to calculate the other Fourier coefficients. To ob-
tain an we multiply (7.1) by cos(nt) and integrate from −π to π:
Z π
FS f (t) cos(nt)dt = πan
−π
mathematical methods 2 95

so we set
1
Z π
an = f (t) cos(nt)dt. (7.2b)
π −π
Note that only one term on the right-hand-side survives the in-
tegration. To obtain bn we multiply (7.1) by sin(nt) and integrate
from −π to π: Z π
FS f (t) sin(nt)dt = πbn ,
−π
so we set
1
Z π
bn = f (t) sin(nt)dt. (7.2c)
π −π
Again, only one term on the right-hand-side survives. The above
process is called expanding the function f as an infinite sum of
orthogonal1 functions. 1
We can think of functions as vectors,
R π define a dot product f 1 . f 2 =
and
−π f 1 ( t ) f 2 ( t ) dt, so that the functions
The expressions (7.2a, b, c) are called Euler’s formulae. in the Fourier series are mutually
orthogonal (dot product equal to 0).
1
Z π
a0 = f (t)dt Notice we need to assume here that
π −π f is sufficiently regular for all these
integrals to be defined (for instance
1
Z π
it is sufficient for f to be piecewise
an = f (t) cos(nt)dt
π −π continuous).

1
Z π
bn = f (t) sin(nt)dt
π −π

It is at this point that we can appreciate the reason we defined


the constant term in the Fourier series to be a0 /2 and not simply
a0 . If we put n = 0 in Euler’s formula for an this result collapses to
the expression defining a0 . Had the factor 1/2 not been inserted in
the definition of the Fourier series then the universal formula that
defines an for all values of n would not apply.

Example 7.3. Define the 2π-periodic function



0, −π < t ≤ 0,
f (t) = and f (t + 2π ) = f (t).
π − t, 0 < t ≤ π,

What is its Fourier series?

Solution: A simple calculation yields


Z 0
1 1 1
Z π Z π
π
a0 = 0dt + (π − t)dt = (π − t)dt =
π −π π 0 π 0 2

A more complicated calculation yields for n > 0, Useful anti-derivatives (integration by


parts):
1 1 − cos(nπ )
Z π
t 1
Z
an = (π − t) cos(nt)dt = t cos(nt)dt = sin(nt) + 2 cos(nt) + c
π 0 πn2 n n
and
Hence we can evaluate an as
t 1
Z
t sin(nt)dt = − cos(nt) + 2 sin(nt) + c
n n

0, n > 0 even,
an =
2/(πn2 ), n odd,
96 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

Figure 7.3: Graph of f for Example 7.3

−3π −2π −π 0 π 2π 3π

or, for k = 1, 2, . . .,

2/ π (2k − 1)2

a2k−1 =
a2k = 0.
A similarly complicated calculation yields
1 1
Z π
bn = (π − t) sin(nt)dt = .
π 0 n
Hence the Fourier series of the above function is
∞ ∞
π 2 cos(nt) sin(nt)
FS f (t) =
4
+
π ∑ n 2
+ ∑ n
n=1, n odd n =1
 
π 2 cos(3t) cos(5t)
= + cos(t) + + +···
4 π 9 25
 
sin(2t) sin(3t)
+ sin(t) + + +···
2 3

7.3 Functions of an arbitrary period

The above analysis extends to functions of arbitrary period 2L rather


than the special value 2π used above. In this case it turns out that

∞   ∞  
a0 nπt nπt
FS f (t) = + ∑ an cos + ∑ bn sin
2 n =1
L n =1
L

where
1 L
Z
a0 = f (t)dt,
L −L
Z L  
1 nπt
an = f (t) cos dt,
L −L L
Z L  
1 nπt
bn = f (t) sin dt.
L −L L
mathematical methods 2 97

Notice that these results revert to the expressions (7.2a, b, c)


when L = π, as indeed they should.

7.4 Convergence of Fourier series

It is of importance to consider the convergence properties of Fourier


series, in other words what happens to the partial sums S N ( f ) as
N → ∞. We would hope that in this limit the partial sums would
approach f (t) so that we could approximate the true value of f (t)
at any given point to a given accuracy just by taking enough terms
in the requisite partial sum. Unfortunately, this cannot be guaran-
teed. Rather, instead of being able to prove that the partial sums
converge to f (t) at every point, it is only possible to be assured that
the integral of the squared difference of the partial sums and the
function goes to zero.
RL
Theorem 7.4. Assume − L f (t)2 dt < ∞. Then
Z L
lim (S N f (t) − f (t))2 dt = 0.
N →∞ − L

In other words, while we cannot be sure of the behaviour of


the partial sums at any one single given point, we do know that
the integral of the square of the difference does approach zero as
N → ∞. The consequence of this is that while FS f (t) = f (t) at
almost all points, there could be a finite number of locations in the
interval [− L, L], where FS f (t) 6= f (t).
An additional issue arises for functions which possess a dis-
continuity. The function in Example 7.2 has a jump at t = 0; for
π < t ≤ 0 then f (t) = 1 but for 0 < t ≤ π we have f (t) = −1. What
does the Fourier series converge to at t = 0? This issue is settled by
the following theorem.

Theorem 7.5. Provided that f (t) and f 0 (t) are bounded and piece-
wise continuous on [− L, L], the Fourier series will converge to (be Recall Definition 1.13, of a piecewise
continuous function.
equal to) f (t) except at points of discontinuity, where it will con-
verge to the average of the right- and left-hand limits of f (t) at that
point, i.e.
f (t+ ) + f (t− )
2
where f (t+ ) is the right-hand limit and f (t− ) is the left-hand limit.

Example 7.6. Example 7.3 revisited. We see that f (t) and f 0 (t) are
bounded and piecewise continuous on [−π, π ] (with the only disconti-
nuity point being 0).2 Since for this function we have f (0− ) = 0 and 2
Note that f 0 (t) considered on its full
f (0+ ) = π, then by Theorem 7.5 the Fourier series FS f (t) converges domain also has discontinuity points
in the odd multiples of π.
to the average of these values, ie. π/2 at t = · · · , −2π, 0, 2π, · · · . The
graph of the Fourier series is then
98 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

Figure 7.4: Fourier series function of


Example 7.3

−3π −2π −π 0 π 2π 3π

Note that it is identical to the graph of f (t) except that it takes the
value π/2 at integer multiples of 2π whereas the function itself is 0 at
these points.

7.5 Functions defined over a finite interval

Writing a function in terms of a Fourier series is convenient for


many calculations. Up to now we have only discussed strictly peri-
odic functions but the ideas of Fourier series can be extended and
applied to many functions that are defined on a finite interval but
which appear to have no intrinsic periodic properties.
Suppose we have f (t) defined on some finite interval of length
2L given by − L < t ≤ L. (It might seem restrictive to assume
that the interval is centred on t = 0. However if the interval is not
centred on the origin it is straightforward to apply a translation and
consider the function in terms of a new co-ordinate t0 for which the
centre is at t0 = 0.)
We can now extend f (t) to all real values of t by defining the
periodic extension of f (t).

Definition 7.7. (Periodic extension)


Let f (t) be a function defined on the interval (− L, L]. The periodic
extension of f (t) is the function φ(t) defined by:

φ ( t ) = f ( t ), − L < t ≤ L, and φ(t + 2L) = φ(t) ∀t.

Now φ(t) is defined for all values of t and is naturally a periodic Exponential periodic extension
3
function of period 2L. Hence we are able to apply the theory of
Fourier series to φ(t).
2

Example 7.8. The graph of the Fourier series of the periodic extension of
et , −1 < t ≤ 1 is on the right. 1

Note that the periodic extension is discontinuous at t = · · · , −5, −3,


0
−3 −2 −1 0 1 2 3

Fourier series of the periodic extension of et .


mathematical methods 2 99

−1, 1, 3, 5,· · · .

7.6 Even and odd functions

Definition 7.9. (Even functions)


A function f (t) if even if and only if f (−t) = f (t) for all t. The graph of
an even function is symmetrical in the vertical axis.

Simple examples of even functions include


f (t) = 1, f ( t ) = t2 and f (t) = cos(t).
In particular all the even power functions f (t) = t2n are even.

Definition 7.10. (Odd functions)


A function f (t) if odd if and only if f (−t) = − f (t) for all t. The graph
of an odd function is 180◦ rotationally symmetric around the origin.

Elementary examples of odd functions include


f (t) = t, f ( t ) = t3 and f (t) = sin(t),

together with the function in Example 7.2. In particular all the odd
power functions f (t) = t2n+1 are odd.
Even and Odd functions
1

Graphs of the even functions cos(t)


0 and t2 , and the odd functions sin(t)
and t3 (non-dashed lines).

−1
−1 0 1

Properties of odd and even functions: Try proving them

(even) + (even) = (even), (odd) + (odd) = (odd)

(even) · (even) = (even), (odd) · (odd) = (even)


(odd) · (even) = (odd), (even) · (odd) = (odd)
Z L
(odd)dt = 0
−L
RL RL
If f (t) is even: −L f (t)dt = 2 0 f (t)dt.
100 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

In words, this tells us that the sum of two even (odd) functions is
itself even (odd). The product of two even or two odd functions is
even while the product of an odd and an even function is odd. The
integral results are particularly important for they facilitate some
great simplifications in the calculation of Fourier series.

7.7 Fourier cosine series for even functions

Even functions must have even Fourier series and hence bn = 0


for all n, giving a Fourier cosine series. There is nothing particularly
special about a Fourier cosine series; really it is little more than a
standard Fourier series with the property that all its sine terms are
absent because the coefficients bn all happen to vanish. We can use
the integration properties of odd and even functions to verify that
bn = 0 when f (t) is an even function. From its definition
1 L 1 L
 
nπt
Z Z
bn = f (t) sin dt = (even)(odd)dt
L −L L L −L
Z L
1
⇒ bn = (odd)dt = 0.
L −L

The Fourier series of an even function f (t) is the cosine


series
∞  
a0 nπt
+ ∑ an cos ,
2 n =1
L
where
Z L Z L  
2 2 nπt
a0 = f (t)dt and an = f (t) cos dt
L 0 L 0 L

Example 7.11. Determine the Fourier series of the even (‘Hats’3 ) function 3
Note that the name ‘hats’ function
 derives from the from of its graph
π + t, −π < t ≤ 0 sketched in the Figure.
f (t) = f (t + 2π ) = f (t).
π − t, 0 < t ≤ π

Hats

−3π −2π −π 0 π 2π 3π

Solution: Since f is even, its Fourier series is a cosine series. We can


compute:
2 π
Z
a0 = (π − t)dt = π,
π 0
mathematical methods 2 101


2 2(1 − (−1)n ) 0 if n even,
Z π
an = (π − t) cos(nt)dt = =
π 0 n2 π  24 if n odd,
n π

and hence the Fourier cosine series is


∞  
π 4 cos(nt) π 4 cos(3t) cos(5t)
FS f (t) = +
2 π ∑ n 2
= +
2 π
cos(t) +
9
+
25
+···
n=1, n odd

7.8 Fourier sine series for odd functions

Odd functions must have odd Fourier series and hence an = 0 for
all n leading to a Fourier sine series. Again it is relatively straightfor-
ward to check that the an = 0 because from the definition
1 L 1 L
 
nπt
Z Z
an = f (t) cos dt = (odd)(even)dt
L −L L L −L
Z L
1
⇒ an = (odd)dt = 0.
L −L

The Fourier series of an odd function f (t) is the sine series


∞  
nπt
∑ bn sin
L
,
n =1

where Z L  
2 nπt
bn = f (t) sin dt
L 0 L

Example 7.12. Determine the Fourier series of the function in Example


7.2.
Solution: Since f (t) is an odd function, its Fourier series is a sine series.
We compute

2
Z π
2(cos(nπ ) − 1) 0 if n even
bn = (−1) sin(nt) dt = = ,
π 0 nπ − 4 if n odd nπ

and hence the Fourier sine series is



4 sin(nt)
FS f (t) = −
π ∑ n
.
n=1,n odd

By Theorem 7.5 we know this converges to the following function:

−1

−3π −2π −π 0 π 2π 3π
102 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

7.9 Half-range expansions

Suppose that a function f (t) is only defined on [0, L]. We could


use the ideas described above to create a periodic function and
hence derive a Fourier series representation. However, with the
function defined on [0, L] it is possible to extend the function in
such a way that the resulting series contains only cosine terms or, if
the extension is made in another way, such that the series contains
only sine terms.
To see how to do this, we extend the domain of definition to
[− L, L], called a half-range expansion, in two ways. We accomplish
this by defining two new functions g(t) and h(t) according to the
following recipes.
Even expansion:

 f (t) if 0 ≤ t ≤ L
g(t) =
 f (−t) if − L ≤ t ≤ 0.

Now g(t) is an even function by construction; therefore the se-


ries for g(t) (or more precisely for its periodic extension) will be a
Fourier cosine series.
Odd expansion:

 f (t)

 if 0 < t ≤ L
h(t) = 0 if t = 0



− f (−t) if − L ≤ t < 0.

This time our function is an odd one so will be given by a


Fourier sine series.

Figure 7.5: Original function and even


Original function and odd expansions.
2
1

−1

−2
−2 −1 0 1 2

Even expansion Odd expansion


2 2
1 1

0 0

−1 −1

−2 −2
−2 −1 0 1 2 −2 −1 0 1 2

As an example look at the sketches in Figure 7.5. Here a func-


tion f (t) is defined for 0 < t < 2 (top panel). In the left lower
mathematical methods 2 103

diagram is shown the even expansion of f (t), that is the function


g(t) given above. This function is now defined on −2 < t < 2
and is clearly even (its graph is symmetric about the vertical axis).
On the other hand, the right lower diagram illustrates the odd ex-
pansion h(t). This time the graph possesses the characteristic 180o
rotational symmetry about the origin indicative of an odd function.
The two extended functions g(t) and h(t) clearly must be given by
Fourier cosine and Fourier sine series respectively. Notice that these
two (different) series converge to the same f (t) for 0 < t < 2 as
both g(t) and h(t) equal f (t) here but will naturally converge to
different values for t < 0.
Example 7.13. Find the Fourier series of the even and odd expansions of

f ( t ) = t2 , 0 ≤ t ≤ 1.

Solution:
Even expansion: The even expansion is just g(t) = t2 , −1 ≤ t ≤ 1.
We find the Fourier coefficients (as g(t) is even, bn = 0):
Z 1
2
a0 = 2 t2 dt = ,
0 3
4(−1)n
Z 1
an = 2 t2 cos(nπt) dt = .
0 π 2 n2
Then the Fourier cosine series is

1 4 (−1)n
+
3 π2 ∑ 2
cos(nπt).
n =1 n

This series converges to g(t) on (−1, 1) so in particular, it converges to


f (t) for 0 < t < 1. 
 t2 if 0 ≤ t ≤ 1
Odd expansion: The odd expansion is h(t) =
−t 2 if − 1 ≤ t < 0.
We find the Fourier coefficients (as h(t) is odd, a0 = an = 0):
−4 − 2(n2 π 2 − 2)(−1)n
Z 1
bn = 2 t2 sin(nπt) dt = .
0 π 3 n3
Hence the Fourier sine series is

2 2 + (n2 π 2 − 2)(−1)n

π3 ∑ n3
sin(nπt).
1

This series converges to h(t) on (−1, 1) so in particular, it converges to


f (t) for 0 < t < 1.
Notice these two series look very different, but they converge to the
same value f (t) = t2 for 0 < t < 1!

7.10 Parseval’s Theorem

There is a relationship between the sum of the squares of all of


the Fourier coefficients of a function and the integral of the square
of the function itself over one period. This relationship turns out
to be very useful in Engineering, Physics and other branches of
Mathematics.
104 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

Theorem 7.14 (Parseval’s Theorem). If a 2π-periodic, piecewise


continuous on [−π, π ], bounded function f (t) has a Fourier series
given by

a0
FS f (t) = + ∑ ( an cos(nt) + bn sin(nt)) ,
2 n =1

then
a20 ∞ 
1
Z π 
[ f (t)]2 dt = + ∑ a2n + bn2 .
π −π 2 n =1

The proof of this result is omitted here. What is more important


is to see how the theorem enables us to derive results concerning
the sums of infinite series of terms. We do this via a few examples.

Example 7.15. We shall apply Parseval’s theorem to Example 7.12. Recall


the function

1, −π < t ≤ 0,
f (t) = and f (t + 2π ) = f (t).
−1, 0 < t ≤ π,

We found that its Fourier series is


∞  
4
FS f (t) = ∑ −

sin(nt).
n=1,n odd

Parseval’s theorem says that


Z 0 ∞  2
1 1 4
Z π

π −π
12 dt +
π 0
(−1)2 dt = ∑ −

n=1, n odd

∞ ∞
16 1 1 π2
⇒ 1+1 =
π2 ∑ n2
⇒ ∑ n 2
=
8
n=1, n odd n=1, n odd

We can use the result of the previous example to find the value
of ∑∞ 1
n=1 n2 by noting that

∞ ∞ ∞
1 1 1
∑ 2
= ∑ 2
+ ∑ 2
n =1 n n=1, n odd
n n=1, n even n

and realising that



1 1 ∞ 1 1 ∞ 1
∑ (2k)2
= ∑
4 k =1 k 2
=
4 n∑ 2
k =1 =1 n

Making use of the result in the previous example gives



1 π2 1 ∞ 1 ∞
1 π2
∑ 2
=
8
+
4 n∑ 2
⇒ ∑ 2
=
6
.
n =1 n =1 n n =1 n

Leonhard Euler first proved that equality in 1741 (by a different


method though).
mathematical methods 2 105

Example 7.16. We shall apply Parseval’s theorem to Example 7.3. Recall


that the function is

0, −π < t ≤ 0,
f (t) = and f (t + 2π ) = f (t)
π − t, 0 < t ≤ π,

with Fourier series


∞ ∞
1 π 2 1
FS f (t) = · + ∑
2 2 n=1, n odd πn 2
cos ( nt ) + ∑ n
sin(nt).
n =1

Then Parseval’s theorem tells us that


∞ 2 ∞  2
π2

1 2 1
Z π

π 0
2
(π − t) dt =
8
+ ∑ πn2
+∑
n
n=1, n odd n =1

∞ ∞
π2 π2 4 1 1

3
=
8
+ 2
π ∑ n 4
+∑ 2
n
n=1, n odd n =1

5π 2 4 1 π2

24
=
π2 ∑ n4
+
6
.
n=1, n odd

1 π4
⇒ ∑ n4
=
96
.
n=1, n odd

An application of Parseval’s theorem to a suitably chosen func-


tion can often yield equivalent results for other infinite sums that
are often difficult to evaluate by other means.

7.11 Differentiation of Fourier series

If we wish to differentiate a function expressed as a Fourier series


it is tempting to simply differentiate each term in the infinite series
one by one. Very often in mathematics we need to be ultra-careful
when dealing with infinite sums because results that look as if they
ought to be reasonable and sensible are not always true! Therefore
it is not an obvious result that the differentiation of a Fourier se-
ries of a function is possible term by term but, fortunately, it can be
proved to yield something that is useful. In particular, for a 2π-
periodic function, if
∞ ∞
a0
FS f (t) = + ∑ an cos(nt) + ∑ bn sin(nt)
2 n =1 n =1

then
∞ ∞
(FS f )0 (t) = − ∑ nan sin(nt) + ∑ nbn cos(nt).
n =1 n =1
Moreover, the following theorem tells us what f 0 (t) is, if f is contin-
uous.

Theorem 7.17. If f is continuous, then (FS f )0 (t) = f 0 (t) at


points t where f (t) is differentiable.
106 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

In particular, note the requirement that the function f (t) be con-


tinuous; if it is not continuous then the result does not necessarily
follow.
Example 7.18. Recall the ’Hats’ function in Example 7.11; this function
is continuous. It is differentiable except at points an integer multiple of π.
Previously we showed that

π 4 cos(nt)
FS f (t) =
2
+
π ∑ n2
n=1, n odd

By Theorem 7.17, we have



4 sin(nt)
f 0 (t) = (FS f )0 (t) = −
π ∑ n
,
n=1, n odd

for all t not a multiple of π. Derivative of Hats FS

The graph of the original function (and of its Fourier series) appeared 1

in Example 7.11. We recognise from Example 7.12 that (FS f )0 (t) is the
Fourier series of the function from Example 7.2, so the graph of (FS f )0 (t)
0
is the one in the margin. Note that it has the value 0 at multiples of π,
the average of the left and right limits, but that f 0 (t) is not defined at
multiples of π. −1

Recall that a function cannot be differentiated at points of dis- −3π −2π −π 0 π 2π 3π

continuity. There are also problems with the convergence of the


differentiated series if f (t) is not continuous, as illustrated by the
following example.
Example 7.19. Let f (t) be the ‘Slopes’ function:
t
f (t) = , −π < t ≤ π, f (t + 2π ) = f (t).
2
The graph of its Fourier series is Figure 7.6. Notice it is discontinuous at
odd multiples of π.

FS of Slopes function Figure 7.6: Fourier series of the Slopes


2 function (Example 7.19)

−1

−2
−3π −2π −π 0 π 2π 3π

Since f is an odd function, it has a sine Fourier series and we can show
(exercise) that

(−1)n+1
FS f = ∑ sin(nt).
n =1
n
mathematical methods 2 107

If we naively find the derivative of the Fourier series term-by-term we


deduce that

(FS f )0 (t) = ∑ (−1)n+1 cos(nt) = cos(t) − cos(2t) + cos(3t) + · · · .
n =1
(♦)
We have an obvious problem here. We know4
that if an infinite series of 4
Recall the ‘Test for divergence’ from
terms is to converge then necessarily the nth term in the series must go to Mathematical Methods 1, Section 13.2.

zero as n → ∞ (although you should also remember that the fact that the
nth term → 0 is not in itself sufficient to guarantee the convergence of a
series). For example, if we try to evaluate (♦) at t = 0 we have
1−1+1−1+...
Actual derivative of slopes
and this series clearly does not converge. Moreover, note that t = 0 is 1

not a problem point for f and f 0 (0) = 1/2 so it is not the case that the
differentiated series only fails at points where f (t) is not differentiable or
has some other problem. If we evaluate (♦) at t = π we have −1 − 1 − 0

1 − · · · which is also clearly nonsense. The actual derivative function is


shown on the right and it is not defined at odd multiples of π.
The conclusion is that we must not differentiate the Fourier −1
−3π −2π −π 0 π 2π 3π

series of a non-continuous function and expect to obtain results


with any meaning.

7.12 Integration of Fourier series

It turns out that the Integration of Fourier series is more stable than
differentiation in the sense that fewer potential problems tend to
arise.

Theorem 7.20. Let f (t) be a 2π-periodic, piecewise continuous on


[−π, π ], bounded function. If a0 = 0 then
Z t ∞ ∞
an bn
−π
f (α)dα = ∑ n
sin(nt) − ∑
n
(cos(nt) − cos(nπ )) .
n =1 n =1

Recall that sin nπ = 0 for integer values of n so the first right-


hand side term is simpler than the second. We must have a0 = 0
because Z t
a0 a0
dα = (t + π ),
−π 2 2
which is not a Fourier series component.
Example 7.21. Recall the Slopes function from Example 7.19 and its
Fourier series

(−1)n+1
FS f = ∑ sin(nt).
n =1
n
Notice a0 = 0 and the function is bounded and piecewise continuous on
[−π, π ]. Thus by Theorem 7.20:
Z t ∞ ∞ n
cos(nt) − cos(nπ ) n cos( nt ) − (−1)
  
α
dα = ∑ (−1)n = ∑ (− 1 ) .
−π 2 n =1 n2 n =1 n2
108 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

From this we get


∞ ∞ ∞
t2 − π 2 (−1)n cos(nt) 1 π2 (−1)n cos(nt)
= ∑ − ∑ = − + ∑ .
4 n =1 n2 n =1 n
2 6 n =1 n2

t2 − π 2 2
Note that the average value of function 4 is − π6 , and that the
integral function is continuous.

Integral of FS Figure 7.7: The integral of the Fourier


1 series.

−1

−2

−3π −2π −π 0 π 2π 3π
8
Laplace Transforms

Laplace transforms represent a powerful method for tackling vari-


ous problems that arise in engineering and physical sciences. Most
often they are used for solving differential equations that cannot
be solved via standard methods. An introduction to the concepts
of and the language relating to Laplace transforms is our plan for
this chapter although some applications of Laplace transform meth-
ods will be outlined in an appendix (Section 8.8). More advanced
theory and uses of the transform are postponed until later units.

8.1 The Laplace transform and its inverse

We begin with a definition of the Laplace transform of a scalar


function f (t) defined for t ≥ 0.

Definition 8.1. (Laplace transform)


Given a function f (t) defined for all t ≥ 0, the Laplace transform (LT)
of f (t) is the function
Z ∞
F (s) = e−st f (t)dt
0

defined for all s ∈ R for which the above integral is convergent. We often
write F (s) as L( f ), or, more precisely L( f )(s).

It is worth remarking that here we are following traditional nota-


tion and denoting the variable of the initial function t (this is moti-
vated by regarding the function f (t) as defined for all ‘time’ t ≥ 0).
Performing the transformation will of course yield a function F (s)
and the usual designation of this Laplace transform variable is s
(although some texts might use p instead). Lastly, we point out
that the Laplace transforms of functions f (t), g(t), h(t), etc. are
normally denoted by their corresponding capital letters F (s), G (s),
H (s), etc.
If F = L( f ) is the Laplace transform of f (t), we say that f (t) is
the inverse Laplace transform (ILT) of F (s), written as f = L−1 ( F ). In
slightly cumbersome terms, this is saying that the inverse transform
of F (s) is that function f (t) whose Laplace transform is F (s).
110 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

We now determine the Laplace transforms of some simple func-


tions.

Example 8.2. If f (t) = 1 for t ≥ 0, then if s > 0


Throughout this chapter we use the
following notational convention. If for
Z ∞ 
e −st ∞ 1 1 1 aR function f (t) the improper integral
F (s) = e−st dt = − =− lim e−st + = . ∞
0 s 0 s t→∞ s s c f ( t ) dt exists and if g ( t ) is an anti-
derivative for f (t), then we write
[ g(t)]∞
c for limt→∞ ( g ( t ) − g ( c )) .
This is so since limt→∞ e−st = 0 for s > 0 (notice this limit does not exist
if s ≤ 0). Thus, the integral exists for s > 0, giving that

1
L(1) = .
s

Notice that here F (s) is not defined for all real values of s, just for s > 0.
The definition of the ILT now implies that
 
1
L −1 =1.
s

Example 8.3. For f (t) = tn for some integer n ≥ 0 then


Z ∞
L(tn ) = e−st tn dt .
0

Substituting u = ts gives
Z ∞  u n du Z ∞
1 n!
L(tn ) = e−u = n +1 un e−u du = n+1 ( for s > 0)
0 s s s 0 s

where the integral can be evaluated using induction and integration by


parts.

Example 8.4. Consider f (t) = e at for t ≥ 0, where a is a constant. Then


for s > a we have

Z ∞
" #∞
−st at e( a−s)t
F (s) = e e dt =
0 a−s
0

1 1 1
=− + lim e(a−s)t = .
a−s a − s t→∞ s−a
Thus, the integral exists for s > a (note F (s) does not exist for s ≤ a) and

1
L(e at ) = (s > a) .
s−a

Hence we can deduce that


 
1
L −1 = e at ( t ≥ 0) .
s−a

For a = 0 this result is consistent with Example 8.2.

Example 8.5. Let f (t) = sin( at) for some a 6= 0, and let F = L( f ).
Notice that for s > 0 we have lim e−st sin( at) = 0 (by the Squeeze Theo-
t→∞
rem); similarly, lim e−st cos( at) = 0. Using this and two integrations by
t→∞
mathematical methods 2 111

parts, we get
Z ∞ Z ∞
1
F (s) = e−st sin( at) dt = − (e−st )0 sin( at) dt
0 s 0
1 a ∞ −st
Z
= − [e−st sin( at)]0∞ + e cos( at) dt
s s 0
Z ∞
a
= 0− 2 (e−st )0 cos( at) dt
s 0
a a2 ∞ −st
Z
= − 2 [e−st cos( at)]0∞ − 2 e sin( at) dt
s s 0
a a2
= 2 − 2 F (s) .
s s
This gives an equation for F (s):

a a2
F (s) = 2
− 2 F (s) .
s s
It is a matter of simple algebra to rearrange to deduce that
a
L(sin( at)) = F (s) = ( s > 0).
s2 + a2
It is then immediately obvious that
 
a
L −1 = sin( at) .
s + a2
2

Exercise 8.1.1. Use similar methods to show that


 
s −1 s
L(cos( at)) = 2 and L = cos( at) for s > 0.
s + a2 s2 + a2

These are just a few of the more straightforward examples of


the Laplace transform. To obtain others we can use some of the
properties of the Laplace transform operation.

8.1.1 Linearity of the Laplace transform


The following property is an immediate consequence of the defini-
tion of Laplace transform.

Theorem 8.6. If the Laplace transforms L( f )(s) and L( g)(s) of two


functions f (t) and g(t) exist for s ≥ a for some a ∈ R, then for any
constants α ∈ R and β ∈ R we have

L (α f + βg) (s) = αL( f )(s) + βL( g)(s)

for s ≥ a.

Recall the hyperbolic trigonometric functions:

1 t 1
cosh(t) = (e + e−t ) and sinh(t) = (et − e−t ).
2 2
112 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

Example 8.7. Find L(cosh( at)) and L(sinh( at)).


Solution: Since cosh( at) = 12 (e at + e− at ), we have
1
L(e at ) + L(e−at )

L(cosh( at)) =
2 
1 1 1 s
= + = 2 ,
2 s−a s+a s − a2
provided both s > a and s > − a, i.e., s > | a|. Similarly,
1
L(e at ) − L(e−at )

L(sinh( at)) =
2 
1 1 1 a
= − = 2 ,
2 s−a s+a s − a2
provided s > | a|.

Example 8.8. Find L−1 ( s(s1−1) ).


1 1
Solution: We decompose in partial fractions: − 1s . Hence
s ( s −1)
= s −1
       
−1 1 −1 1 1 −1 1 −1 1
L =L − =L −L = et − 1.
s ( s − 1) s−1 s s−1 s
From the linearity of the Laplace transform follows the linearity
of the inverse Laplace transform.

Theorem 8.9. Let F (s) and G (s) be functions. If the inverse Laplace
transforms f (t) = L−1 ( F )(t) and g(t) = L−1 ( G )(t) exist, then for
any constants α ∈ R and β ∈ R we have

L−1 (αF + βG ) (t) = αL−1 ( F )(t) + βL−1 ( G )(t).

8.1.2 Existence of Laplace transforms


Recall Definition 1.13: A function f ( x ) is called piecewise contin-
uous on a given interval [ a, b] if f has only finitely many points of
discontinuity in [ a, b]. Piecewise continuous functions possess a
Laplace transform if they are of exponential order:

Definition 8.10. (Exponential order)


A function f (t), t ≥ 0, is of exponential order if f (t) is piecewise
continuous and bounded on every interval [0, T ] with T > 0 and there
exist constants M > 0 and γ ∈ R such that

| f (t)| ≤ Meγt for all t ≥ 0 .

When this holds we will say that the exponential order of f is ≤ γ.

Given a function is of exponential order ≤ γ we can then deduce


for what values of s its Laplace transform is defined:
mathematical methods 2 113

Theorem 8.11. If f (t) is of exponential order ≤ γ, then the Laplace


transform F (s) = L( f )(s) exists for all s > γ.

The Laplace transform of a given function is unique. Conversely,


if two functions have the same Laplace transform then they can
differ only at isolated points.

Example 8.12. For f (t) = e at , we saw in Example 8.4 that the transform
exists for s > a. This is consistent with Theorem 8.11 since f (t) is of
exponential order ≤ a: taking M = 1 and γ = a we see that | f (t)| ≤
Me at for all t ≥ 0.

2 2
Example 8.13. For f (t) = et there are no M and γ for which et ≤
2
Meγt for all t ≥ 0. In very informal terms, et grows more quickly than
2
eγt for any γ. The Laplace transform L(et ) does not exist in this case.
This example proves that not every well-defined function necessarily has a
Laplace transform.

8.1.3 Exercises

Based on the examples above you should be able to tackle the fol-
lowing problems:

Exercise 8.1.2. Use integration by parts to show that for any constants
a > 0 and ω ∈ R we have
ω
(a) L(e− at sin(ωt)) = , for s > 0
( s + a )2 + ω 2
s+a
(b) L(e− at cos(ωt)) = , for s > 0
( s + a )2 + ω 2
(Hint: Write down the definition of the Laplace transform in each case.
A suitable substitution will reduce the integrals to those in Example 8.5
thereby circumventing the need to do pages of laborious calculation.)

Exercise 8.1.3. Use the linearity of the Laplace transform and some of the
above examples to find the Laplace transforms of:
(a) f (t) = cos t − sin t,
(b) f (t) = t2 − 3t + 5,
(c) f (t) = 3e−t + sin(6t).

Exercise 8.1.4. Use the linearity of the inverse Laplace transform and
some of the above examples to find the inverse Laplace transforms of:
2
(a) F (s) = − , s > −16
s + 16
4s
(b) F (s) = 2 , s>3
s −9
3 1
(c) F (s) = + 2 , s > 7.
s−7 s
114 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

8.2 Inverse Laplace transforms of rational functions

If the Laplace transform F (s) = L( f ) of some function f (t) has the


special form
P(s)
F (s) = ,
Q(s)
where P(s) and Q(s) are polynomials, then we can find the inverse
Laplace transform f (t) = L−1 ( F ) using partial fractions. Remember
that partial fractions were discussed in Chapter 1 of these notes.
Notice that at this stage we know the inverse Laplace transforms
of the following basic rational functions (see Examples 8.4, 8.5 and
Exercise 8.1.1)
 
1
L −1 = e−at for s > − a,
s+a
   
1 1 s
L −1 = sin( at), L −1 = cos( at), for s > 0.
s2 + a2 a s2 + a2
Apart from this, we also have the following two formulae (see Ex-
ample 8.7) which are sometimes useful
   
1 1 s
L −1 2 = sinh ( at ) , L −1
= cosh( at), for s > | a|.
s − a2 a s2 − a2

To recall the method of partial fractions and demonstrate how it


applies to problems involving inverse Laplace transforms, we look
at two examples.
2s − 1
Example 8.14. Suppose F (s) = , s > 1, and we want to
(s2 − 1)(s + 3)
find f (t) = L−1 ( F ). First, using partial fractions, we write

2s − 1 A B C
F (s) = = + + .
(s − 1)(s + 1)(s + 3) s−1 s+1 s+3

Using the method explained in Chapter 1, we get A = 1/8, B = 3/4,


C = −7/8.
Thus,
1 3 7
F (s) = + −
8( s − 1) 4( s + 1) 8( s + 3)
and therefore

f (t) = L−1 ( F (s))


     
1 −1 1 3 1 7 1
= L + L −1 − L −1
8 s−1 4 s+1 8 s+3
1 t 3 −t 7 −3t
= e + e − e .
8 4 8

2s2 − s + 4
Example 8.15. Suppose F (s) = , for s ≥ 0. To find f (t) =
s3 + 4s
− 1
L ( F ), we first use partial fractions:

2s2 − s + 4 A Bs + C
F (s) = = + 2 .
s ( s2 + 4) s s +4
mathematical methods 2 115

Using the method explained in Chapter 1, we get A = 1. Taking value


s = 1, we get 1 + B+5 C = 55 or B + C = 0. Taking value s = −1, we
get −1 + − B5+C = −75 or C − B = −2. This system has solution B = 1,
C = −1.
Thus,
1 s−1 1 s 1
F (s) = + 2 = + 2 −
s s +4 s s + 4 s2 + 4
and therefore
     
−1 −1 1 −1 s −1 1
f (t) = L ( F ) = L +L −L
s s2 + 4 s2 + 4

1
= 1 + cos(2t) − sin(2t).
2

8.2.1 Exercise
Exercise 8.2.1. Use partial fractions to find the inverse Laplace trans-
forms of:
2s
(a) F (s) = −
(s + 1)(s2 + 1)
1
(b) F (s) = 4 .
s − 16

8.3 The Laplace transform of derivatives and integrals of f (t)

Later we shall show that some of the most important applications


of Laplace transforms are to the solutions of differential equa-
tions. To that end it is important to know the forms of the Laplace
transforms of the derivatives and integral of f (t). The form of the
Laplace transform of f 0 (t) is given in the following theorem:

Theorem 8.16. If f (t) is continuous and of exponential order ≤ γ


and if f 0 (t) exists and is piecewise continuous and bounded over
[0, T ] for all T ≥ 0, then the Laplace transform of f 0 (t) exists for
s > γ and
L( f 0 )(s) = sL( f )(s) − f (0) .

Proof. This is easy to verify using integration by parts. Indeed, if


we denote by F (s) and G (s) the Laplace transforms of f (t) and
f 0 (t), respectively, then
Z ∞ ∞ Z ∞
e−st f 0 (t)dt = e−st f (t) 0 + s e−st f (t)dt

G (s) =
0 0
= − f (0) + lim e−st f (t) + sF (s) = sF (s) − f (0) ,
t→∞

since for s > γ we have |e−st f (t)| ≤ Me(γ−s)t → 0 as t → ∞.

It is worth pausing at this juncture and spending a few moments


reflecting on this result. There are a few very important proper-
ties that need to be appreciated before we proceed. First, note how
116 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

G (s) = L( f 0 )(s) is a multiple of F (s) minus a constant. Before


working through Theorem 8.16 we might well have expected that
the Laplace transform of f 0 (t) would have involved the derivative
of the Laplace transform of f (t), in other words dF/ds. But that
clearly is not the case; G (s) is related to F (s) by simple algebraic
multiplication and this is the first clue as to the usefulness of the
Laplace transform in solving differential equations. Laplace trans-
forms of derivatives of y(t) are changed to algebraic multiples of
its Laplace transform Y (s) and, as we shall see, this means that ul-
timately the solution of the problem reduces to a task of solving an
algebraic equation which is normally a much easier prospect than
analysing the original differential equation.
The second aspect of note in the result of Theorem 8.16 is the
presence of the value f (0). This too is not expected – normally
when dealing with the derivative of a function we are not con-
cerned with the value of the function itself at any given point. But
that is not the case for the Laplace transform; in order to find the
Laplace transform of f 0 (t) completely some knowledge of f (0) is
required.
Given the technique used to deduce the form of G (s) we can re-
peat the process to obtain the Laplace transforms of higher deriva-
tives of f (t). For example

L( f 00 )(s) = sL( f 0 )(s) − f 0 (0) = s[sL( f )(s) − f (0)] − f 0 (0) ,

so that
L( f 00 )(s) = s2 L( f )(s) − s f (0) − f 0 (0) .

Similarly,

L( f 000 )(s) = s3 L( f )(s) − s2 f (0) − s f 0 (0) − f 00 (0) ,

and, more generally (this can be proved using a simple induction),

L( f (n) )(s) = sn L( f )(s) − sn−1 f (0) − . . . − s f (n−2) (0) − f (n−1) (0) .

Once again we remark that L( f (n) )(s) involves no derivatives of


F (s) at all; it is given by a multiple sn of F (s) plus a polynomial of
degree n − 1 in s. The coefficients of this polynomial involve the
values of the first n − 1 derivatives of f (t) at t = 0.

Example 8.17. The above can be used to find L(sin( at)) by an alterna-
tive route than that taken in Example 8.5. Let f (t) = sin( at). Then f (t)
is continuous and of exponential order ≤ 0 (since | f (t)| ≤ 1e0t ) and f 0 (t)
is continuous, so we can apply Theorem 8.16. We have f 00 (t) = − a2 f (t)
and so

− a2 L( f ) = L( f 00 ) = s2 L( f ) − s f (0) − f 0 (0), for s > 0.


mathematical methods 2 117

Collecting the two terms involving L( f ),

( a2 + s2 )L( f ) = s f (0) + f 0 (0)

and, using that f (0) = 0 and f 0 (0) = a, we obtain


a
L( f ) = , for s > 0.
s2 + a2
Using all of the techniques above (and some more to come
later) we can construct a table of Laplace transforms of frequently-
encountered functions. Such a table is provided on page 129.
Next we consider how one can derive a formula for the Laplace
transform of an integral.

Theorem 8.18. If f (t) is of exponential order (so that L( f ) exists)


Z t
then g(t) = f (u) du is of exponential order ≤ γ for some γ.
0
Moreover, for s > γ and s 6= 0 we have

1
L( g)(s) = L( f )(s) .
s
In other words, if the Laplace transform of f (t) is F (s), then
  Z t
F (s) F (s)
L( g)(s) = and L−1 = L−1 ( F )(u) du .
s s 0

Sketch of Proof. Denote the Laplace transform of g(t) by G (s). By


definition of g(t), it is continuous and it can be proved that it is of
exponential order, say ≤ γ. Since g0 (t) = f (t) by the Fundamental
theorem of Calculus and since g(0) = 0, Theorem 8.16 implies (for
s > γ)

F (s) = L( f )(s) = L( g0 )(s) = sL( g)(s) − g(0) = sG (s) .

Thus, for s > γ, s 6= 0, we have G (s) = F (s)/s. 


This can be particularly useful in helping to determine the in-
verse transform of functions which have a factor s appearing in the
denominator.

Example 8.19. Find the inverse Laplace transform g(t) of G (s) =


1
, using Theorem 8.18. We could also use the partial fraction
s ( s2 + ω 2 ) method.
F (s) 1
Solution: Notice that G (s) = , where for F (s) = 2 we know
s s + ω2
1
f ( t ) = L −1 ( F ) = sin(ωt).
ω
Then Theorem 8.18 yields
Z t Z t
1 1
g(t) = L−1 ( G (s)) = f (u) du = sin(ωu) du = [1 − cos(ωt)] .
0 ω 0 ω2
118 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

8.3.1 Exercises
Exercise 8.3.1. Use the formula for the Laplace transform of a double
derivative to show that
2ωs
(a) L(t sin(ωt)) = 2
( s + ω 2 )2
s2 − ω 2
(b) L(t cos(ωt)) = 2 .
( s + ω 2 )2
Exercise 8.3.2. Use the formula for the Laplace transform of a derivative
to find:
(a) L(te at )
(b) L(tn e at )

 Use the
Exercise 8.3.3.  formula for
 the Laplace  transform of an integral
−1 1 −1 1
to find L and L (no partial fractions
s ( s + 3) s2 ( s + 3)
required).

8.4 Solving differential equations

Laplace transforms can be applied to initial-value problems for


linear ordinary differential equations by reducing them to the task
of solving an algebraic equation.
However it should be realised that Laplace transform meth-
ods will only be able to detect solutions of ordinary differential
equations that have Laplace transforms. While most solutions will
satisfy this requirement, not all will. We saw in Example 8.13 that
2
f (t) = et does not have a Laplace transform but this is a solution
of the differential equation
dy
− 2ty = 0;
dt
we could not therefore expect to derive a meaningful solution of
this equation using the Laplace transform. We need to bear in mind
that although the Laplace transform will find most solutions of
differential equations there are isolated cases when it will fail.
Despite this caution, it can be shown that all solutions of con-
stant coefficient differential equations are of exponential order so
we can use Laplace transform methods to seek a solution y(t) of

y00 (t) + ay0 (t) + by(t) = r (t) t ≥ 0 ,

where a, b are constants and r (t) is a given function, such that y(t)
satisfies the initial conditions

y ( 0 ) = K0 , y 0 ( 0 ) = K1 .

To solve this, first transform the differential equation, writing

L(y00 ) + aL(y0 ) + bL(y) = R(s),

where R(s) = L(r )(s). In terms of Y (s) = L(y), this gives the
equation

[s2 Y (s) − sy(0) − y0 (0)] + a[sY (s) − y(0)] + bY (s) = R(s) .


mathematical methods 2 119

This can be written in the form

(s2 + as + b)Y (s) = R(s) + (s + a)y(0) + y0 (0) = R(s) + (s + a)K0 + K1 ,

so we have
R ( s ) + ( s + a ) K0 + K1
Y (s) = .
s2 + as + b
Therefore
 
−1 −1 R ( s ) + ( s + a ) K0 + K1
y(t) = L (Y ) = L .
s2 + as + b

This method will become clearer with some examples.

Example 8.20. Solve the initial value problem

y00 (t) − y(t) = t , y(0) = 1 , y0 (0) = 1.

Solution: Applying the Laplace transform and denoting Y (s) = L(y),


we get
1
[s2 Y (s) − sy(0) − y0 (0)] − Y (s) = L(t) = 2 .
s
Using the initial conditions y(0) = y0 (0) = 1, we write it in the form

1
( s 2 − 1 )Y = s + 1 + .
s2
Solving for Y (s) gives

s+1 1 1 1 1
Y (s) = + = + − ,
s2 − 1 s2 ( s2 − 1) s − 1 s2 − 1 s2

using partial fractions or by direct observation. So, from the table of


Laplace transforms,
     
−1 −1 1 −1 1 −1 1
y(t) = L (Y )(t) = L +L 2
−L
s−1 s −1 s2
3 t 1 −t
= et + sinh t − t = e − e − t.
2 2
Example 8.21. Solve the initial value problem

y(4) (t) − y(t) = 0 , y(0) = 0 , y0 (0) = 1, y00 (0) = y000 (0) = 0.

Solution: Let Y (s) = L(y)(s). Applying the Laplace transform to the


DE, we get

[s4 Y (s) − s3 y(0) − s2 y0 (0) − sy00 (0) − y000 (0)] − Y (s) = 0 .

Using the initial conditions, this gives s4 Y (s) − s2 − Y (s) = 0, and


therefore
s2
Y (s) = 4 .
s −1
To find y(t) = L−1 (Y ) we need to find a convenient partial fraction
expansion for Y (s). The following will be adequate:

s2 1/4 1/4 1/2


Y (s) = = − + .
(s − 1)(s + 1)(s2 + 1) s − 1 s + 1 s2 + 1
120 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

Hence
1 −1 1 1 1 1 1
y ( t ) = L − 1 (Y ) = L ( ) − L −1 ( ) + L −1 ( 2 )
4 s−1 4 s+1 2 s +1
1 t 1 −t 1
= e − e + sin t .
4 4 2

An advantage of this technique is that it is not necessary to solve


for the general solution of the homogeneous differential equation
and then determine the arbitrary constants in that solution. Apart
from this it can also be used for higher order differential equations,
as we saw in Example 8.21.

8.4.1 Exercises

Exercise 8.4.1. Solve the initial value problems using the Laplace trans-
form :
(a) y0 (t) − 9y(t) = t, y(0) = 5
(b) y00 (t) − 4y0 (t) + 4y(t) = cos t, y(0) = 1, y0 (0) = −1
(c) y00 (t) − 5y0 (t) + 6y(t) = e−t , y(0) = 0, y0 (0) = 2
(d) y(4) (t) − 4y(t) = 0, y(0) = 1, y0 (0) = 0, y00 (0) = −2,
y000 (0) = 0.

8.5 Shift theorems

We have now seen the general strategy for solving differential equa-
tions using Laplace transforms; we transform the differential prob-
lem to an algebraic one for Y (s) and then, given our knowledge of
inverse Laplace transforms, we attempt to reconstruct the form of
y(t). It is this last step that is potentially the tricky one for there is
always the possibility that Y (s) is of a form we do not recognise.
The situation gets worse. It is relatively straightforward to find
the Laplace transform of a function in as much that given an f (t)
we can, at least theoretically, compute F (s) using the definition of
a Laplace transform but, unfortunately, there is no easy equivalent
definition for going in the reverse direction (ie. given F (s), deduce
f (t)). Thus it is of importance to expand our repertoire of easily
identifiable inverse functions and this is facilitated using two so-
called shift theorems.

Theorem 8.22. If F (s) is the Laplace transform of f (t) for s > b,


then the Laplace transform of e at f (t) is Note that there is no restriction on a
in this theorem: a can be positive or
negative.
L(e at f (t)) = F (s − a)

for s − a > b. Equivalently,

L−1 ( F (s − a)) = e at f (t) .


mathematical methods 2 121

Proof. We have
 Z ∞ Z ∞
L e at f (t) = e−st e at f (t) dt = e−(s− a)t f (t)dt = F (s − a) ,

0 0
which proves the statement.

This is called s-shifting, as the graph of the function F (s − a) is


obtained from that of F (s) by shifting a units (to the right if a > 0
and to the left if a < 0) on the s-axis. Putting this result in words it
tells us that if the Laplace transform of f (t) is F (s), then the shifted
function F (s − a) is the transform of e at f (t).
Example 8.23. Find the Laplace transform of e at tn .
n!
Solution: Recall Example 8.3: L(tn )(s) = for s > 0. Using this
s n +1
and Theorem 8.22 we get
n!
L(e at tn )(s) = L(tn )(s − a) = (s > a) .
( s − a ) n +1
For example,
4!
L(e2t t4 )(s) = L(t4 )(s − 2) = ( s > 2) .
( s − 2)5
Example 8.24. Find the Laplace transform of e at cos(ωt).
s
Solution: The Laplace transform of f (t) = cos(ωt) is F (s) =
s2 + ω 2
(s > 0). Hence
s−a
L e at cos(ωt) = L(cos(ωt))(s − a) =

(s > a) .
( s − a )2 + ω 2
1
Example 8.25. Find the inverse Laplace transform of (s− a)n
.
1 1
Solution: Notice (s− a)n
= F (s − a) for F (s) = sn . We know that
n − 1
f (t) = L−1 ( F )(s) = (nt −1)! .
It follows that, for any integer n ≥ 1, we have
e at tn−1
 
−1 1 at
L ( t ) = e f ( t ) = .
(s − a)n ( n − 1) !
For example,
e−2t t2 t2 e−2t
 
1
L −1 (t) = = .
( s + 2)3 2! 2
Using s-shifting, we can find the inverse Laplace transform of
as + b
any function of the form G (s) = 2 , where ps2 + qs + r =
ps + qs + r
0 has no real roots.
Example 8.26. Find the inverse Laplace transform of
1 1
G (s) = = .
s2 − 4s + 7 ( s − 2)2 + 3
1
Solution: We have G (s) = F (s − 2), where F (s) = , so
s2 + 3
√ √
L−1 ( F )(t) = sin( 3t)/ 3.
√ √
By Theorem 8.22, L−1 ( G )(t) = e2t sin( 3t)/ 3.
122 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

Here is a more complicated example.

Example 8.27. Find the inverse Laplace transform of

2s 2s
G (s) = = .
s2 + 2s + 5 ( s + 1)2 + 4

Solution: We use a similar method to the previous example.

2( s + 1) − 2
 
L −1 ( G ) ( t ) = L −1
( s + 1)2 + 4
2s − 2
 
= e − t L −1
s2 + 4
 
s 1
= e − t L −1 2 2 − 2
s + 22 s2 + 22
= e−t [2 cos(2t) − sin(2t)] .

The second shifting theorem is related to the so-called Heaviside


function H (t) defined by

0 , t<0,
H (t) =
1 , t≥0.

This is also called the unit step function.


Notice that for any a ∈ R the graph of H (t − a) is obtained from
the graph of H (t) by shifting a units (to the right if a > 0 and to the
left if a < 0), that is:

0 , t < a ,
H (t − a) =
1 , t ≥ a .

We point out that multiplying a given function g(t) by H (t − a)


has the effect of turning the function off until time t = a and then
activating it. More precisely, we have

 0 , t<a,
g(t) H (t − a) =
 g(t) , t ≥ a .

Multiplication by the pulse function H (t − a) − H (t − b), where


a < b has the effect of a switch. This function has value one for
a ≤ t < b and is zero for times t < a and t ≥ b. Thus the application
of this function is equivalent to turning on a switch at t = a then
turning it off again at a later t = b.


 0

 , t<a,
g(t)[ H (t − a) − H (t − b)] = g(t) , a ≤ t < b ,


0 , t≥b.

mathematical methods 2 123

Because the Heaviside function is quite so important in real


problems it is helpful to note the result of the following theorem,
called t-shifting theorem.

Theorem 8.28. If the Laplace transform of f (t) is F (s) for s > b,


then for any a ≥ 0 we have
There is a restriction on a in this
theorem: these results are not valid if a
L[ f (t − a) H (t − a)] = e−as F (s) is negative.

for s > b. Consequently,

L−1 e−as F (s) = f (t − a) H (t − a) .




You can try to prove this result (it involves the definition of the
Laplace transform of H and one change of coordinate).

Example 8.29. Find L( H (t − a)) where a ≥ 0.


Solution: We take f (t) = 1 and apply the theorem. We saw in Example
8.2 that F (s) = L( f ) = 1/s for s > 0. Thus

e− as
L ( H (t − a)) = L ( f (t − a) H (t − a)) = e−as F (s) = , for s > 0.
s
Example 8.30. Find L−1 (e−4s /s3 ).
Solution: We apply the theorem with a = 4 and F (s) = 1/s3 , so that
L−1 (e−4s /s3 ) = f (t − 4) H (t − 4). All we have to do is determine f (t).
From the table we get that f (t) = t2 /2, so that,

  ( t − 4) 2 0 if t < 4 ,
L−1 e−4s /s3 = H (t − 4) =
2 (t − 4)2 /2 if t ≥ 4 .

Example 8.31. Find L( g(t)) where



t if 0 ≤ t < 3 ,
g(t) =
1 − 3t if t ≥ 3 .

Solution: We can express f (t) with Heaviside functions:

g(t) = t[ H (t) − H (t − 3)] + (1 − 3t) H (t − 3)


Recall we are only concerned with
= tH (t) + (1 − 4t) H (t − 3) functions defined on [0, ∞). On that
interval, H (t) is nothing else than the
= t + (1 − 4t) H (t − 3). constant function 1.

This is still not in the form required in order to use Theorem 8.28: in
the second term we have to write 1 − 4t as a function of t − 3. Since
t = (t − 3) + 3, we have 1 − 4t = 1 − 4(t − 3) − 12 = −4(t − 3) − 11.
Thus
g(t) = t − [4(t − 3) + 11] H (t − 3) .
We now apply Theorem 8.28 with f (t) = 4t + 11.
4 11
L([4(t − 3) + 11] H (t − 3)) = L( f (t − 3) H (t − 3))(s) = e−3s F (s) = e−3s ( + ).
s2 s
124 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

Thus
1 4 11
L( g)(s) = 2
− e−3s ( 2 + ).
s s s
We now solve a differential equation where the right-hand side
uses Heavside functions.

Example 8.32. Solve the initial value problem

y00 + y = H (t − 1) − H (t − 2)

with initial conditions y(0) = 0 and y0 (0) = 1.


Solution: Taking transforms, Y (s) = L(y) satisfies

e−s e−2s
s2 Y (s) − sy(0) − y0 (0) + Y (s) = −
s s
and solving for Y yields that

1 1
Y (s) = + (e−s − e−2s ) 2
s2 + 1 s ( s + 1)
 
1 −s −2s 1 s We used partial fractions here.
= + ( e − e ) −
s2 + 1 s s2 + 1
   
1 −s 1 s −2s 1 s
= + e − − e − .
s2 + 1 s s2 + 1 s s2 + 1

We now need to find the inverse Laplace transformof Y (s). For


 the last
1 s
two terms, we will apply Theorem 8.28 with F (s) = s − s2 +1 , so that
f (t) = 1 − cos t, and with a = 1 or 2. We get

y(t) = sin t + H (t − 1)[1 − cos(t − 1)] − H (t − 2)[1 − cos(t − 2)].

Hence

sin t,

 0 ≤ t < 1,
y(t) = sin t + 1 − cos(t − 1), 1≤t<2


sin t − cos(t − 1) + cos(t − 2), 2 ≤ t.

Note that in this solution both y and y0 are continuous at t = 1 and t = 2.

8.5.1 Exercises
Exercise 8.5.1. Find the Laplace transform of the functions
(a) (t3 − 3t + 2)e−2t
(b) e4t (t − cos t)
(c) f (t) = 2t + 1 for 0 ≤ t < 2 and f (t) = 2 − 3t for t > 2.

Exercise 8.5.2. Find the inverse Laplace transforms of


e−s
(a)
( s − 5)3
se−2s
(b) 2 .
s +9
Exercise 8.5.3. Solve the initial value problem y00 (t) − 2y0 (t) − 3y(t) =
f (t), y(0) = 1, y0 (0) = 0, where f (t) = 0 for 0 ≤ t < 4 and f (t) = 12
for t ≥ 4.
mathematical methods 2 125

8.6 Derivatives of transforms

If f (t) is of exponential order ≤ γ and piecewise continuous and


bounded on [0, T ] for any T > 0, then by Theorem 8.11,
Z ∞
F (s) = L( f )(s) = e−st f (t)dt
0

exists for s > γ. Moreover we have the following:

Theorem 8.33. Derivative of transform. Under the above assump-


tions, F 0 (s) exists for all s > γ, and

− F 0 (s) = L (t f (t)) . (8.1)

Consequently,

L −1 F 0 ( s ) = − t f ( t )

( t ≥ 0).

The proof of this result is omitted here as our focus is on how we


might use Theorem 8.33 to find more function-transform pairs.
Example 8.34. In Exercise 8.3.1, we found the transform of g(t) =
t sin(ωt) by differentiating twice. We now have an easier method since by
taking f (t) = sin(ωt) in (8.1), so that F (s) = L(sin(ωt)) = s2 +ωω2 , we
get:
 
0 d ω 2ωs
L(t sin(ωt)) = − F (s) = − 2 2
= 2 .
ds s + ω ( s + ω 2 )2
Exercise 8.6.1. Use Theorem 8.33 twice to find the Laplace transform of
f (t) = t2 cos(ωt).

8.7 Convolution

One last idea relevant to the theory of Laplace transforms is that


of the convolution of two functions. Very often the solution of the
transformed problem can be written in the form Y (s) = F (s) G (s);
it is extremely tempting to suppose that y(t) = f (t) g(t) which
would say that the Laplace transform of a product of two func-
tions is the product of the Laplace transforms of the two functions.
Unfortunately things are not that simple and instead requires the
introduction of a concept known as the convolution.

Definition 8.35. (Convolution)


Given two functions f (t) and g(t), both of them being piecewise continu-
ous and bounded on every finite interval [0, T ], the convolution f ∗ g of f
and g is defined by
Z t
( f ∗ g)(t) = f (u) g(t − u) du .
0
126 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

Main properties of the convolution:

f ∗g = g∗ f commutative
f ∗ ( g1 + g2 ) = f ∗ g1 + f ∗ g2 distributive
( f ∗ g) ∗ h = f ∗ ( g ∗ h) associative
f ∗0 = 0∗ f = 0

However, note that f ∗ 1 is not equal to f in general and that


f ∗ f can be negative.

Theorem 8.36. (The Convolution Theorem) Let f (t) and g(t) be as


above and let F (s) = L( f ) and G (s) = L( g) be defined for s > γ.
Then
L( f ∗ g)(s) = F (s) G (s) (s > γ) .
Equivalently, L−1 ( F (s) G (s)) = ( f ∗ g)(t).

This result tells us that if f (t) and g(t) have Laplace transforms
F (s) and G (s) respectively then the function ( f ∗ g)(t) has Laplace
transform F (s) G (s).

Example 8.37. Find the inverse Laplace transform of

1
K (s) =
( s − 1)2 ( s − 3)2
using the Convolution Theorem. Another method would be to use
partial fractions as in Section 8.2.
Solution: Notice that, by Example 8.25,
   
−1 1 t −1 1
L = te = f (t) , L = te3t = g(t) .
( s − 1)2 ( s − 3)2
Therefore by the Convolution Theorem,
 
−1 −1 1 1
L (K ) = L · = f (t) ∗ g(t)
( s − 1)2 ( s − 3)2
Z t Z t
= f (u) g(t − u) du = ueu (t − u)e3(t−u) du
0 0
Z t
= e3t (tu − u2 )e−2u du .
0

To evaluate the latter integral we have to use several integrations by parts


which shows that
t − 1 + (t + 1)e−2t
Z t
(tu − u2 )e−2u du = .
0 4

t − 1 + (t + 1)e−2t (t − 1)e3t + (t + 1)et


Thus, L−1 (K ) = e3t = .
4 4
1
Example 8.38. Find the inverse Laplace transform of F (s) =
( s2 + 1)2
using the Convolution Theorem.
mathematical methods 2 127

1 1
Solution: Since F (s) = · and L−1 ( s2 1+1 ) = sin(t), by
( s2 + 1) ( s2 + 1)
the Convolution Theorem,
Z t
f (t) =L−1 ( F ) = sin(t) ∗ sin(t) = sin(u) sin(t − u) du
0
1 t
Z
= [cos(u − (t − u)) − cos(u + (t − u))] du (using the cosine sum formula)
2 0
Z t
1
= [cos(2u − t) − cos t] du
2 0
 t
1 1
= sin(2u − t) − u cos t
2 2 0
1 1
= sin t − t cos t.
2 2

Exercise 8.7.1. In the same manner as the previous example, show that

s2
 
−1 1 1
L = sin t + t cos t .
( s2 + 1)2 2 2

We will use an example to demonstrate how the convolution


Theorem can be applied to solve differential equations.

Example 8.39. Solve the initial value problem



1 if 0 ≤ t < 1 ,
y00 (t) + y(t) = f (t) =
0 if t ≥ 1 .

with initial conditions y(0) = 0, y0 (0) = 1.

Solution: First we write the right-hand side using Heaviside functions:


f (t) = H (t) − H (t − 1). Applying the Laplace transform to the DE, we
get
1 − e−s
s2 Y ( s ) − 1 + Y ( s ) = ,
s
where Y (s) = L(y). This gives (using partial fractions)

1 1
Y (s) = + (1 − e − s ) 2
s2 + 1 s ( s + 1)
1 1 s e−s 1
= 2 + − 2 − ·
s +1 s s +1 s s2 + 1

Taking the inverse Laplace transform of this and using the Convolution
Theorem for the last term, we get

y(t) = L−1 (Y ) = sin t + 1 − cos t − H (t − 1) ∗ (sin t) .

To evaluate the convolution integral, notice that when 0 ≤ u ≤ t < 1 we


have H (u − 1) = 0. Thus,
Z t
H (t − 1) ∗ (sin t) = H (u − 1) sin(t − u) du = 0, for t < 1
0
128 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

while for t ≥ 1 we have:


Z t
H (t − 1) ∗ (sin t) = H (u − 1) sin(t − u) du
0
Z t
= sin(t − u) du
1
= [cos(t − u)]1t
= [1 − cos(t − 1)] .

Hence H (t − 1) ∗ (sin t) = [1 − cos(t − 1)] H (t − 1).


Finally we get the solution to the initial value problem:

y(t) = sin t + 1 − cos t − H (t − 1)[1 − cos(t − 1)] .

8.7.1 Exercises
Exercise 8.7.2. Find the inverse Laplace transform of the functions using
the Convolution Theorem:
1
(a) 2
(s + 4)(s2 − 4)
s
(b) 2
(s + a2 )(s2 + b2 )
e−2s
(c) 2 .
s + 16
Exercise 8.7.3. Solve the following initial value problems using Laplace
transforms:
(a) y00 (t) + 4y0 (t) + 13y(t) = f (t), y(0) = y0 (0) = 0, where f (t) = 1
for 0 ≤ t < π and f (t) = 0 for t ≥ π.
(b) y00 (t) + 2y0 (t) + 2y(t) = sin t, y(0) = y0 (0) = 0.
mathematical methods 2 129

Laplace Transform Table Table 8.1: Laplace Transforms and


Inverses
Z ∞
L ( f (t)) = F (s) = f (t)e−st dt
0

SPECIFIC FUNCTIONS GENERAL RULES

F (s) f (t) F (s) f (t)

1 e− as
1 H (t − a)
s s
1 t n −1
, n ∈ Z+ e− as F (s) f (t − a) H (t − a)
sn ( n − 1) !
1
e− at F (s − a) e at f (t)
s+a
1 t n −1
, n ∈ Z + e− at sF (s) − f (0) f 0 (t)
(s + a)n ( n − 1) !
1 sin(ωt)
s2 F ( s ) − s f (0) − f 0 (0) f 00 (t)
s2 + ω 2 ω
s
cos(ωt) F 0 (s) −t f (t)
s2 + ω 2
1 e−at sin(ωt)
F (n) ( s ) (−t)n f (t)
( s + a )2 + ω 2 ω
Z t
s+a F (s)
e− at cos(ωt) f (u) du
( s + a )2 + ω 2 s 0

1 sin(ωt) − ωt cos(ωt) 
F (s) G (s) f ∗ g (t)
( s2 + ω 2 )2 2ω 3
s t sin(ωt)
( s2 + ω 2 )2 2ω
Higher derivatives
 
L f ( n ) ( t ) = s n F ( s ) − s n −1 f (0 ) − s n −2 f 0 (0 ) − · · · − s f ( n −2) (0 ) − f ( n −1) (0 )

The Convolution Theorem:


 Z t
L ( f ∗ g) = L ( f ) L ( g) where f ∗ g (t) = f (u) g(t − u) du
0
130 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

8.8 Appendix: Some applications of the Laplace transform

We conclude our discussion of the Laplace transform by illustrating


a few applications of the theory to problems that arise in various
branches of engineering. Our examples are by no means exhaustive
but give a flavour of the fields in which solutions using Laplace
transforms are of importance.
Note that the material in this Appendix is intended to be for
further reading. It will not be explicitly tested by questions in the
examination.

8.8.1 A Mechanical Model – the vibrating spring


Consider the motion of an object of mass m attached at the end of
a (helical) spring. Let us assume that the spring is either vertical or
horizontal and the displacement of the object is measured along an
x-axis which is positioned so that x = 0 coincides with the equilib-
rium position of the object and directed so that x > 0 corresponds
to an extension of the spring.
If the displacement of the object at time t is denoted by x (t),

the governing equation for x (t) is derived under the following
0
assumptions:
m +
• Newton’s Second Law:
d
(m x 0 (t)) = sum of all forces acting on the object at time t.
dt
When the mass of the object is constant (i.e. does not depend on
the time t), as in the case here, we simply have

m x 00 (t) = sum of all forces acting on the object at time t. (8.2)

• Hooke’s Law: The elastic force exerted by the free end of the
spring is proportional to the displacement of that end. The factor
k > 0 of proportionality is called the spring constant.

• The sum of all damping (frictional) forces (e.g. air resistance,


viscosity if the object is moving in a fluid, etc.) acting on the
object at time t is proportional to the velocity x 0 (t) of the object
at time t and has direction opposite to the direction of motion,
i.e. it has the form −cx 0 (t), where c ≥ 0 is a constant (the so-
called damping constant).

Assuming that there is also some driving external force f (t)


acting on the object at time t, the sum of all forces acting at time t
is f (t) − kx (t) − cx 0 (t). Thus by (8.2),

mx 00 (t) = f (t) − kx (t) − cx 0 (t) ,

which is equivalent to

mx 00 (t) + cx 0 (t) + kx (t) = f (t). (8.3)

This is the differential equation governing the motion of the object


at the end of the spring.
mathematical methods 2 131

Example 8.40. An object of mass 1 kg is suspended from a spring with


a spring constant of 4 N/m. The mass starts at equilibrium but is set in
motion by applying a constant downward force of F0 N from time t = 0
until time t = 1 when the force is turned off and the mass allowed to
oscillate freely. If the air resistance is negligible, find the position of the
object at times t ≥ 0.

Solution: From the assumptions, x (0) = x 0 (0) = 0, c = 0 and k = 4.


The external force f (t) is given by f (t) = F0 for 0 ≤ t ≤ 1 and f (t) = 0
for t > 1. The differential equation (8.3) has the form

x 00 (t) + 4x (t) = f (t) ,

and on applying the Laplace transform to both sides we get

s2 X (s) + 4X (s) = F (s) ,

where F (s) = L( f )(s). Notice that we can write f (t) in terms of Heavi-
side functions f (t) = F0 [1 − H (t − 1)], so

1 e−s
 
F (s) = F0 [L(1) − L( H (t − 1))] = F0 − .
s s

Hence
e−s
 
F (s) 1
X (s) = 2 = F0 − .
s +4 s ( s2 + 4) s ( s2 + 4)

Using partial fractions we have


 
1 1 1 s
= − 2
s ( s2 + 4) 4 s s +4

and therefore
      
1 1 1 s
L −1 = L −1 − L −1 2
s ( s2 + 4) 4 s s +4
1
= (1 − cos 2t).
4

Next, since L−1 (s/(s2 + 4)) = cos 2t, using the (second) Shifting
Theorem, we get

e−s 1 −1 e − s se−s
   
−1
L = L − 2
s ( s2 + 4) 4 s s +4
1 1
= H (t − 1) − H (t − 1) cos 2(t − 1).
4 4

Hence

e−s
   
1
x (t) = L−1 ( X (s)) = F0 L−1 − F0 L −1
s ( s2 + 4) s ( s2 + 4)
 
1 F F
= F0 − cos 2t − 0 H (t − 1) + 0 H (t − 1) cos 2(t − 1) .
4 4 4
132 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

8.8.2 Simple Electric Circuits


A basic electric circuit contains an electromotive force E(t) (sup-
plied by a cell, battery or generator), a resistor, an inductor and a
capacitor, connected in a series. As usual, the resistance R of the
resistor is measured in ohms Ω, the coefficient of inductance L of
the inductor is measured in henries H and the capacitance C of the
capacitor is measured in farads F. We will assume that R, L and C
are constant, i.e. do not depend on the time t.

Figure 8.1: Series circuit.


E
R L
C

Let q(t) be the charge of the capacitor at time t, measured in


coulombs. The current i (t) in the circuit at time t (measured in
amperes) is defined as the rate of change of the charge q at time t,
i.e.
dq
i (t) = (t) .
dt
From physics it is known that:

• the voltage drop across the resistor (at time t) is Ri = Ri (t) ;

• the voltage drop across the inductor is Li0 (t) ;


1
• the voltage drop across the capacitor is q(t) .
C
To describe the DE modelling the change of the charge with the
time one also needs Kirchoff’s Laws:

• Kirchoff’s Voltage Law: the algebraic sum of the voltage rises and
drops in any closed loop in a circuit is equal to zero.

• Kirchoff’s Current Law: the algebraic sum of the currents at any


junction of a circuit is equal to zero. That is, the total current
entering the junction is equal to the total current leaving the
junction.

Hence E(t) = Ri (t) + Li0 (t) + C1 q(t) . Substituting i (t) = q0 (t), we


obtain the following differential equation:

d2 q dq 1
L 2
+ R + q = E(t) . (8.4)
dt dt C
On differentiating both sides of this equation, one derives a second
order equation for the current i (t):

d2 i di 1 dE
L 2
+R + i= (t) .
dt dt C dt
mathematical methods 2 133

Figure 8.2: Circuit 2.


E A
10V B R=250KΩ C=10−6F Eout

Example 8.41. Consider the circuit shown in Figure 8.2.


Assume that the capacitor has zero charge initially and that there is
zero initial current. At time t = 2 seconds the switch is thrown from
position B to position A, held there for 1 second, and then switched back
to position B. Find the output voltage Eout = C1 q(t) (the voltage on the
capacitor).
Solution: From the assumptions, it follows that q(0) = 0, q0 (0) =
i (0) = 0, and the electromotive force E(t) = 0 for 0 ≤ t < 2 and for t ≥ 3
and E(t) = 10 V for 2 ≤ t < 3. Thus,

E(t) = 10[ H (t − 2) − H (t − 3)] .

From Kirchoff’s Voltage Law, Ri (t) + C1 q(t) = E(t) at any time t ≥ 0,


which gives the differential equation 8.4 with L = 0:

250000 q0 (t) + 106 q(t) = E(t) .

Taking Laplace transforms, we get

250000[sQ(s) − q(0)] + 106 Q(s) = e(s) ,

where Q(s) = L(q) and

10e−2s 10e−3s
e(s) = L( E)(s) = 10[L( H (t − 2)) − L( H (t − 3))] = − .
s s
10e−2s 10e−3s
Thus, 106 (s/4 + 1) Q(s) = + , and therefore
s s
 
4 4
Q(s) = 10−5 e−2s − e−3s
s ( s + 4) s ( s + 4)
Using partial fraction (or a direct observation), one gets
4 1 1
= − .
s ( s + 4) s s+4
This and the (second) Shift Theorem imply
    
−1 −5 −1 4 −2s −1 4 −3s
q(t) = L ( Q(s)) = 10 L e −L e
s ( s + 4) s ( s + 4)
 −2s
e−2s e−3s e−3s
   
e
= 10−5 L−1 − − L −1 −
s s+4 s s+4
= 10−5 [ H (t − 2) − e−4(t−2) H (t − 2) − H (t − 3) + e−4(t−3) H (t − 3)] .

Thus, the output voltage is Eout = 106 q(t)

= 10[ H (t − 2) − e−4(t−2) H (t − 2) − H (t − 3) + e−4(t−3) H (t − 3)] .


134 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

Output Voltage
Figure 8.3: Output voltage for Example
10 8.41
9

0
0 1 2 3 4 5 6

8.8.3 Models that lead to systems of differential equations


First we deal with a mechanical model slightly more complicated
than the earlier one.
Consider two masses m1 and m2 resting on a horizontal table.
Assume that the masses are connected to each other and to two
fixed supports by three unstretched springs S1 , S2 and S3 with
respective spring constants k1 , k2 and k3 .
The masses can be displaced from their equilibrium positions,
for example, by holding m1 , moving m2 to the right, and then re-
leasing both. (Naturally, there are many other possibilities for set-
ting the system in motion.) Apart from this, some external forces
f 1 (t) and f 2 (t) can act on the masses m1 and m2 , respectively. We
want to find the equations of motion of the two masses.

Figure 8.4: Spring–mass system.


k1 k2 k3

m1 m2

f1→ f2→

Let O1 and O2 be the positions of m1 and m2 , respectively, when


the system is at rest. Denote by x1 (t) (resp. x2 (t)) the displacement
of m1 (resp. m2 ) from its equilibrium position O1 (resp. O2 ) at time
t, assuming that xi > 0, when the displacement is to the right of
Oi . Then at any given time t, the length of the spring S1 has been
changed by x1 (t) units, S2 by x2 (t) − x1 (t) units, and S3 by − x2 (t)
units.
To derive a differential equation for x1 (t) we essentially repeat
the earlier argument. By Newton’s Second Law, m1 x100 (t) equals the
sum of all forces acting on the mass m1 at time t. Apart from f 1 (t),
there are three other forces acting on m1 : the restoring (elastic)
forces of the springs S1 and S2 and the damping force (e.g. due to
mathematical methods 2 135

friction). For S1 the restoring force at time t is −k1 x1 (t), while for
S2 the restoring force acting on m1 at time t is k2 [ x2 (t) − x1 (t)]. The
damping force is −c1 x10 (t) for some (damping) constant c1 ≥ 0.
Hence

m1 x100 (t) = f 1 (t) − k1 x1 (t) + k2 [ x2 (t) − x1 (t)] − c1 x10 (t) .

A similar differential equation is obtained from applying Newton’s


Second Law to the mass m2 . This gives the following system that
describes the motion of the two masses:

m x 00 + c x 0 + k x + k [ x − x ] = f (t) ,
1 1 1 1 1 1 2 1 2 1
(8.5)
m2 x 00 + c2 x 0 + k3 x2 + k1 [ x2 − x1 ] = f 2 (t) .
2 2

Given some initial conditions for x1 and x2 (i.e. xi (0) and xi0 (0) for
i = 1, 2), the system can be solved for x1 (t) and x2 (t). One possible
way of doing this is by applying the Laplace transform to each
differential equation, thus obtaining a system of algebraic equations
for X1 (s) = L( x1 (t)) and X2 (s) = L( x2 (t)).

Example 8.42. In the model considered above, assume that the friction is
negligible and the system is at rest initially (i.e. both masses are at their
equilibrium positions). From time t = 0 until time t = 3 seconds a force
of magnitude 2 N acts on the mass m2 in direction to the right. Find the
positions x1 (t) and x2 (t) at any time t ≥ 0, provided m1 = m2 = 1 kg
and all spring constants are the same: k1 = k2 = k3 = k > 0.

Solution: From the assumptions we have x1 (0) = x10 (0) = x2 (0) =


x20 (0) = 0 and c1 = c2 = 0 (no damping forces involved). Also
f 1 (t) = 0 and f 2 (t) = 2 for 0 ≤ t < 3 and f 2 (t) = 0 for t ≥ 3, i.e.
f 2 (t) = 2 − 2H (t − 3).
The system (8.5) now has the form

 x 00 + 2kx − kx = 0
1 1 2
 x 00 + 2kx2 − kx1 = f 2 (t) .
2

If we denote

2 2e−3s
F2 (s) = L( f 2 (t)) = 2L(1) − 2L( H (t − 3)) = − ,
s s
then applying the Laplace transform to each of the differential equations in
the above system leads to the algebraic equations

s2 X (s) + 2kX (s) − kX (s) = 0
1 1 2
(8.6)
s2 X2 (s) + 2kX2 (s) − kX1 (s) = F2 (s) .

Solving the system with respect to X1 (s) and X2 (s) (e.g. by expressing
X2 (s) in terms of X1 (s) from the first equation and then substituting it
into the second), one finds

k s2 + 2k
X1 ( s ) = F2 ( s ) , X 2 ( s ) = F2 (s) .
(s2 + 2k)2 − k2 (s2 + 2k)2 − k2
136 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

We need to apply L−1 to these two formulae to retrieve x1 (t) and


x2 (t). Notice that
 
1 1 1 1 1
= 2 = − .
(s2 + 2k)2 − k2 (s + k)(s2 + 3k) 2k s2 + k s2 + 3k

Thus,

2 2e−3s
  
1 1 1
X1 ( s ) = − −
2 s2 + k s2 + 3k s s
1 e−3s 1 e−3s
= − − + .
s ( s2 2 2 2
+ k) s(s + k) s(s + 3k) s(s + 3k)
Using partial fractions, one gets
   
1 1 1 s 1 1 1 s
= − 2 , = − 2 .
s ( s2 + k ) k s s +k s(s2 + 3k ) 3k s s + 3k
Thus,

1 e−3s s e−3s
   
1 1 s
X1 ( s ) = − 2 − − 2
k s s +k k s s +k
 −3s
s e−3s
  
1 1 s 1 e
− − 2 + − 2 .
3k s s + 3k 3k s s + 3k
Applying the inverse Laplace transform and using the Shift Theorem, one
finds
2 1 √ 1 √
x1 ( t ) = [1 − H (t − 3)] − cos( kt) + H (t − 3) cos( k(t − 3))
3k k k
1 √ 1 √
+ cos( 3kt) − H (t − 3) cos( 3k(t − 3)) .
3k 3k
Similarly, for X2 (s) we have

2 2e−3s
  
1 1 1
X2 ( s ) = + −
2 s2 + k s2 + 3k s s
1 1
= (1 − e−3s ) + (1 − e−3s )
s ( s2 + k ) s(s2 + 3k )
   
1 1 s 1 1 s
= − 2 (1 − e−3s ) + − 2 (1 − e−3s ) ,
k s s +k 3k s s + 3k
and therefore
4 1 √ 1 √
x2 ( t ) = [1 − H (t − 3)] − cos( kt) + H (t − 3) cos( k(t − 3))
3k k k
1 √ 1 √
− cos( 3kt) + H (t − 3) cos( 3k(t − 3)) .
3k 3k
Example 8.43. Next, consider an electric circuit containing two closed
loops as in the figure.
Here R1 and R2 are the resistances of the corresponding resistors, q1 (t)
and q2 (t) the charges of the corresponding capacitors, i1 (t), i2 (t) and
i (t) are the currents in the corresponding parts of the circuit, L is the
inductance of the single inductor, and E(t) is the electromotive force.
Using Kirchoff’s current law for the first junction, one sees that i1 −
i2 − i = 0, i.e. i = i1 − i2 . (The same result follows if the current law is
applied at the second junction.)
mathematical methods 2 137

Figure 8.5: Double loop circuit.

R1 R2

E L
C1 C2

Next, applying Kirchoff’s voltage law to the left loop, it follows that

di 1
R1 i1 ( t ) + L ( t ) + q1 ( t ) = E ( t ) .
dt C1

Using i = i1 − i2 and i j (t) = q0j (t) for j = 1, 2, we get

1
L[q100 (t) − q200 (t)] + R1 q10 (t) + q (t) = E(t) . (8.7)
C1 1

Similarly, it follows from Kirchoff’s voltage law applied to the right loop
that
1
− L[q100 (t) − q200 (t)] + R2 q20 (t) + q2 (t) = 0 . (8.8)
C2
The equations (8.7) and (8.8) form a system of differential equations that
determine q1 (t) and q2 (t) provided some initial conditions for these are
given.

One can deal in a similar way with other electric circuits. Some-
times it is more convenient to consider differential equations in-
volving the currents (i1 and i2 in the case above).

Example 8.44. Consider the electric circuit in the accompanying figure.


Find the currents i1 (t) and i2 (t) at any time t ≥ 0, assuming that all
charges and current are zero when the switch is closed at t = 0.

Figure 8.6: Example 8.44.

i1→ i2→
↓ i1−i2

E R1=200Ω R2=300Ω

L=0.5H C=5×105F

Solution: It is convenient to derive differential equations for the currents


i1 (t) and i2 (t). Let q2 (t) be the charge on the capacitor at time t. Since
i20 (t) = q2 (t), we have
Z t Z t
q2 ( t ) = q2 (0) + i2 (u) du = i2 (u) du .
0 0
138 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

Using this and Kirchoff’s laws, we obtain the following system of equa-
tions:
1

 i10 (t) + 200[i1 (t) − i2 (t)] = 50

2 Z t
1
300i2 (t) + 200[i2 (t) − i1 (t)] +
 i2 (u) du = 0 .
50 × 10−6 0
Applying the Laplace transform to each equation, one finds the following
system for I1 (s) = L(i1 (t)) and I2 (s) = L(i2 (t)):

100
sI1 (s) + 400I1 (s) − 400I2 (s) =

s
5I2 (s) − 2I1 (s) + 200 I2 (s) = 0 .

s
Solving, one finds I1 (s) and I2 (s), and then i1 (t) = L−1 ( I1 (s)) and
i2 (t) = L−1 ( I2 (s)). We leave the details to the reader.

8.8.4 Exercises
Exercise 8.8.1. A body of mass m is fixed at the end of a vertical spring
with spring constant k > 0. Initially the mass is at the equilibrium po-
sition. At time t = 0 the upper end of the spring starts to be moved
vertically by means of a rotating eccentric or cam with the periodic dis-
placement y(t) = a sin ωt. Assume that the air resistance is negligible
and that y(t) and the displacement x (t) of the mass at time t are measured
from their initial positions and both x and y are positive for displacements
downward. Find the position x (t) of the mass at any time t ≥ 0.

Exercise 8.8.2. Consider the electric circuit in the figure below. Assume
that the capacitor has zero charge initially and that there is zero initial
current. At time t = 2 seconds the switch is moved from position B to
position A, held there for 1 second, and then switched back to position B.
Find the current i (t) at any time t ≥ 0.

Figure 8.7: Exercise 8.8.2.


E A
10V B R=250KΩ C=10−6F

Exercise 8.8.3. Consider the system of masses and (vertical) springs in


the figure. Assume that air resistance is negligible and the initial displace-
ments and velocities are zero.
(a) Find the displacements x1 (t) and x2 (t) if the external forces acting
on the masses are f 1 (t) = 2 and f 2 (t) = 0 for all t ≥ 0.
(b) Find the displacements x1 (t) and x2 (t) if the external forces acting
on the masses are f 1 (t) = 1 − H (t − 2) and f 2 (t) = 0 for all t ≥ 0.

Exercise 8.8.4. Consider the mechanical system in the figure below,


where the springs have spring constants k1 and k2 . The spring to the left
mathematical methods 2 139

g→ Figure 8.8: Exercise 8.8.3.


k1=6 k2=2 k3=3

m1=1 m2=1

f1→ f2→

of M has damping constant c1 > 0, while the right one is undamped.


Assume that a periodic driving force f (t) = A sin ωt is acting on M
starting from t = 0.

Figure 8.9: Exercise 8.8.4.


k1 k2

M m

f→
c1

(a) Derive and solve the system of equations for the displacement func-
tions of M and m.

(b) Show that if m and k2 are chosen so that ω = k2 /m, then the
mass m cancels the forced vibrations of M. (In this case m is called a
vibration absorber.)

Exercise 8.8.5. Find the currents i1 (t) and i2 (t) for the electric circuit
in the figure, assuming that E(t) = 1 − H (t − 4) sin[2(t − 4)] and the
initial charges and currents are zero.

Figure 8.10: Exercise 8.8.5.

i1→ i2→
R1=2Ω R2=1Ω

E L=5H
R3=3Ω R4=4Ω
9
Complex Functions — Derivatives

9.1 Complex numbers and some basic functions

Recall that a complex number z = x + iy, where x is the real part of


z and y is the imaginary part of z, can also be written as z
y

r cos θ + i sin θ = r cis θ = reiθ ,




p r
where r = |z| = x2 + y2 is called the modulus of z and θ = arg(z)
its argument. This θ is the angle of the radius vector to the point
( x, y) in the argand diagram and, by convention, is usually chosen θ
x
to be in the interval (−π, π ].
√ √
Example 9.1. i = 1 cis π/2 = eiπ/2 and 1 + i = 2 cis π/4 = 2eiπ/4 .

Theorem 9.2 ( De Moivre’s theorem). If z = r cis θ = reiθ and


n ∈ Z, then zn = r n cis(nθ ).

For any m ∈ N and any complex number z = r cis θ, there are m


complex solutions to the equation x m = z: they are

r1/m cis(θ/m + 2πk/m) for k = 0, . . . , m − 1.

So z1/m does not take a unique value but actually has up to m dif-
ferent values! Hence simply writing f (z) = z1/m does not immedi-
ately give us a well-defined function although this can be done by
exercising a little more care. We can make z1/m a genuine function
by specifying just one of the roots, say the first, k = 0. This might
seem a little strange but we should recognise this type of behaviour
from a simple discussion of the square root of a positive real num-
ber. Everyone knows that the square root of 4 is 2; but the roots
of the equation x2 = 4 are 2 and −2. Thus if z is a real positive
number then z1/2 has two possible values but we almost always
identify the square root of a positive number uniquely by specify-
ing the square root to be positive as well. This makes the square
root function well-defined and it precisely this sort of restriction
we are imposing when we take one particular root of the complex
142 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

equation x m = z so that z1/m obeys all the rules to be satisfied by a


complex function.
More generally, zq , for q ∈ Q \ Z, is multi-valued and hence
f (z) = zq is not a function unless further impositions are made.
Once more we do this in practice by focusing on just one of the
roots. Even more generally, we would like to find a way to ensure
that f (z) = zb , b ∈ R is a well-defined function.
Similar problems are encountered if we try to make ez , log z,
sin z and cos z functions from C to C and this is what we will look
at next. There are a number of possible ways to proceed, but one
important property we need to preserve is the definitions of these
functions for real values of z. We already know many of the prop-
erties of the exponential, logarithmic and trigonometrical functions
for real arguments, and we must ensure that however we define
these functions for a complex z they must revert to our standard
functions when z is real. If we fail to do this, we are liable to cause
ourselves considerable strife further down the line!
So let us start with the exponential function exp(z) := ez . We
have already used eiθ (Euler’s formula) to represent the complex
number cos θ + i sin θ. Hence, using the usual rules for exponents,

ez = e x+iy = e x eiy = e x cos y + i sin y




which is defined in terms of known real functions. Note that


ez maps C to C\{0} and is onto, but observe also that e x+iy =
e x+i(y+2π ) , so this function is not one-to-one.
The ln function is defined to be the inverse function of e x for re-
als. Both real functions are one-to-one and onto for their respective
domains and ranges. We would therefore like now to define log z as For complex numbers we denote the
the inverse function of exp(z), that is, log z = w if ew = z (note that natural logarithm by log z.

log 0 does not exist). But we have just seen that ez , while defined
for all z ∈ C is not one-to-one. This is a similar issue as when we
wanted to find inverse functions for sin and cos in Chapter 1. We
will use here the same trick as we did then: by restricting the range
of log z we can form a well-defined complex function. The standard
range of log z is taken to be

{ x + iy | x ∈ R, −π < y ≤ π }

that is the real part of log z is unrestricted but the imaginary part
is confined to lie in the interval (−π, π ]. Confining the range of
a complex function to something smaller than the whole of the
complex plane is called choosing a branch of the function.
Let w = x + iy and suppose ew = z for z ∈ C. We have

z = ew = e x cos y + i sin y ,


so |z| = e x and arg(z) = y (note that arg(z) ∈ (−π, π ], as required


by our choice of branch). We conclude that log z = w = x + iy is
defined by x = ln |z|, and y = arg(z), where arg(z) ∈ (−π, π ]. In
summary,
log z := ln |z| + i arg(z).
mathematical methods 2 143

Note that log z is defined on C \ {0}. You might like to check


that this definition satisfies the usual inverse function identities
of f f −1 (z) = z = f −1 f (z) for the appropriate domains and
 

ranges.
Given this definition, we are now in a position to define ab for
a, b ∈ C, a 6= 0. We have a = elog a from the above definition, so
ab = eb log a . Note that in general this is multi-valued, because we
can choose different branches of the log function.

Example 9.3. Compute ii+1 .


Solution: By the formula, ii+1 = e(i+1) log i , so we need to compute log i.

log i = ln |i | + i arg(i ) = ln 1 + iπ/2 = iπ/2.

Thus

ii+1 = e(i+1)iπ/2 = e−π/2+iπ/2 = e−π/2 cis π/2 = ie−π/2 .

In particular, we get the definition of the function f (z) = zb =


eb log z , which coincides with the usual function if both z and b are
real numbers.
From eiθ = cos θ + i sin θ and e−iθ = cos θ − i sin θ, for real θ, and
the fact that eiz is well defined, we have
eiz − e−iz eiz + e−iz
sin z = cos z = , z ∈ C.
2i 2
Note that if z ∈ R, then cos z and sin z are the usual trigonometric
functions. It is easy to check that for z, w ∈ C,

cos2 z + sin2 z = 1,

sin(z + w) = sin z cos w + cos z sin w,


cos(z + w) = cos z cos w − sin z sin w,
so that these results, which are familiar for real arguments, con-
tinue to hold in the complex plane. Last, we mention that the other
complex trigonometric functions are defined similarly to their real
counterparts in terms of sin z and cos z. Just to take one example
then, for complex z we have tan z = sin z/ cos z.
As a summary, we have:

Definition 9.4. (Complex functions)


Let z = x + iy and b be complex numbers. Then

• ez = e x+iy = e x cos y + i sin y




• log(z) = ln |z| + i arg(z)

• zb = eb log z , for z 6= 0

• sin z = 1 iz
2i ( e − e−iz )

• cos z = 12 (eiz + e−iz )


144 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

9.2 The derivative of a function of a complex variable

Consider a complex valued function of a complex variable

f : A ⊂ C → C, where f (z) = f ( x + iy) = u( x, y) + iv( x, y).

The rules for differentiation of f follow similarly to those of a real


valued function of a real variable. The analysis follows that of a real
function of two variables with two real components

f : R2 → R2 : ( x, y) 7→ (u, v),

but because of the extra arithmetic structure of complex numbers,


some amazing results arise.
We need the concept of a neighbourhood of z0 ∈ C, being a disc in
the complex plane around the point z0 of radius r > 0, defined as

D ( z0 , r ) = { z ∈ C : | z − z0 | < r }.

An open set1 A ⊂ C is defined as a set of points where each point 1


This definition is equivalent to the
has a neighbourhood contained in A. one seen in Mathematical Methods 1:
none of the boundary points are in the
subset.

Definition 9.5. (Limit of f at a point)


Let f be defined on D (z0 , r )\{z0 } for some r > 0. Then

lim f (z) = a
z → z0

means that we can make the values of f (z) arbitrarily close to a by taking More precisely, for the interested
student: for every e > 0, there is
z to be sufficiently close to z0 , but not equal to z0 . This means that f (z) a δ > 0 such that z ∈ D (z0 , r ),
is close to a when z is close to z0 in all directions. Note that the limit a is z 6= z0 , and |z − z0 | < δ imply that
| f (z) − a| < e.
unique.

The following equations hold (as for real functions of real vari-
ables). Let limz→z0 f (z) = a and limz→z0 g(z) = b, then

1. limz→z0 f (z) ± g(z) = a ± b.

2. limz→z0 f (z) g(z) = ab.

3. limz→z0 f (z)/g(z) = a/b, if b 6= 0.

4. If h is defined on the image of a neighbourhood of z0 under f



and limw→ a h(w) = c, then limz→z0 h f (z) = c.

Definition 9.6. (Continuity at a point, continuity on an open set)


Let A ⊂ C be an open set and let f : A → C be a function. We say f Once again z is allowed to approach
z0 from any direction, and the same
is continuous at z0 ∈ A if and only if limz→z0 f (z) = f (z0 ), and f is limit must be attained.
continuous on A if f is continuous at each z0 ∈ A.

We remark that the function log z is not continuous at z0 , when


z0 is a real negative number.
mathematical methods 2 145

Definition 9.7. (Complex Differentiability)


Let A ⊂ C be an open set and f : A → C. Then f is differentiable in
the complex sense at z0 ∈ A if

f ( z ) − f ( z0 )
lim
z → z0 z − z0

exists. This limit is denoted f 0 (z0 ) (or sometimes d f /dz (z0 )) and is a
complex number. When f is complex differentiable at every point in A,
we say f is analytic on A. Note that the phrase “analytic at z0 ” means
analytic in a neighbourhood of z0 .

Compare the above to the definition of the derivative of f : R →


R, where there are only two directions to consider (x → a+ and
x → a− ). Compare to the definition of the differentiability of a
function g : R2 → R2 .

Example 9.8. The exponential and trigonometric functions are analytic


on any open set of their domain. Moreover

d exp(z)
= exp(z),
dz
d sin(z) d cos(z)
= cos(z), = − sin(z).
dz dz
Example 9.9. The logarithmic function log z is analytic on any open set
not containing any real negative numbers, and

d log(z)
= 1/z.
dz
Theorem 9.10. Differentiability implies continuity. If f 0 (z0 ) exists then
f is continuous at z0 .

9.3 Rules for derivatives of analytic functions

Suppose that f and g are analytic on an open set A ⊂ C.

(1) Let a, b ∈ C. Then a f + b g is analytic on A, and

( a f + b g ) 0 ( z ) = a f 0 ( z ) + b g 0 ( z ).

(2) f g is analytic on A and

( f g ) 0 ( z ) = f 0 ( z ) g ( z ) + f ( z ) g 0 ( z ).
 −1
Similarly f /g = f (z) g(z) , where g(z) 6= 0 ∀z, is analytic on
A, and the quotient rule holds:
 −1  −2 0
( f /g)0 (z) = f 0 (z) g(z) − f (z) g(z) g (z)
f 0 (z) g(z) − f (z) g0 (z)
= 2 .
g(z)
146 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

(3) Any monomial is analytic on C, with (zn )0 = nzn−1 (n ≥ 1) and


hence any polynomial too, using (1).

(4) Any rational function

a0 + a1 z + · · · + a n z n
b0 + b1 z + · · · + bm zm
is analytic on C \ {zeroes of the denominator} (there are at most
m zeroes).

(5) Let B ⊂ C be an open set and suppose that f ( A) ⊂ B. Let


h : B → C be analytic on B. Then h ◦ f : A → C defined by

h f (z) is analytic and

d
( h ◦ f )(z) = h0 f (z) f 0 (z).

dz

The conclusion is that the complex derivatives of the exponential,


logarithm, sin and cos and polynomial functions are exactly what
we would expect from our knowledge of the real results. Moreover
the product, quotient and chain rules of differentiation can be ap-
plied to complex functions in precisely the same way we would use
them in the real case. The upshot is that many of our long estab-
lished results for derivatives of real functions transfer seamlessly to
complex functions. However we shall find that complex differentia-
tion introduces some extra structure and results and these are best
captured within the so–called Cauchy-Riemann equations.

9.4 The Cauchy–Riemann equations

This result connects the idea of f as a complex function, u + iv, of a


complex variable, x + iy, with f considered as a function from R2 to
R2 , that is, mapping ( x, y) to (u, v):

f ( x + iy) = u( x, y) + iv( x, y)

Example 9.11. Consider the exponential function

exp(z) = exp( x + iy) = e x cos y + ie x sin y.

Then u( x, y) = e x cos y and v( x, y) = e x sin y. Computing the partial


derivatives of u and v, we see that

∂u ∂v ∂u ∂v
= e x cos y = and = −e x sin y = − .
∂x ∂y ∂y ∂x

So for the exponential function we see that the partial derivatives


of the real part u( x, y) = e x cos y are related to the derivatives
of the imaginary part v( x, y) = e x sin y according to u x = vy
and uy = −v x . To see whether these results might hold for other
functions we recall the definition of complex differentiability of a
function that
f ( z ) − f ( z0 )
f 0 (z0 ) = lim
z → z0 z − z0
mathematical methods 2 147

as z → z0 from any direction. Let us evaluate this limit twice; once


when z → z0 so that z − z0 is real and again when this quantity is
purely imaginary.
For the first calculation, if z0 = x0 + iy0 then z = ( x0 + ∆x ) + iy0
where ∆x ∈ R and ∆x → 0. Then
f ( z ) − f ( z0 ) f (z0 + ∆x ) − f (z0 )
lim = lim
z → z0 z − z0 ∆x →0 ∆x
u( x0 + ∆x, y0 ) + iv( x0 + ∆x, y0 ) − u( x0 , y0 ) − iv( x0 , y0 )
= lim
∆x →0 ∆x
u( x0 + ∆x, y0 ) − u( x0 , y0 ) v( x0 + ∆x, y0 ) − v( x0 , y0 )
= lim + i lim
∆x →0 ∆x ∆x →0 ∆x
∂u ∂v
=
( x0 , y0 ) + i ( x0 , y0 )
∂x ∂x
using the definition of partial derivatives of real functions. Hence
we conclude that
∂u ∂v
f 0 ( z0 ) = ( x0 , y0 ) + i ( x0 , y0 ). (♥)
∂x ∂x
Now let us repeat the argument again, with the one difference
that z approaches z0 from the imaginary direction so that z =
x0 + i (y0 + ∆y), where ∆y ∈ R and ∆y → 0. Then we proceed so
that
f ( z ) − f ( z0 ) f (z0 + i∆y) − f (z0 )
lim = lim
z → z0 z − z0 ∆y→0 i∆y
u( x0 , y0 + ∆y) + iv( x0 , y0 + ∆y) − u( x0 , y0 ) − iv( x0 , y0 )
= lim
∆y→0 i∆y
u( x0 , y0 + ∆y) − u( x0 , y0 ) v( x0 , y0 + ∆y) − v( x0 , y0 )
= lim + i lim
∆y→0 i∆y ∆y→0 i∆y
1 ∂u ∂v ∂v ∂u
= ( x0 , y0 ) + ( x0 , y0 ) = ( x0 , y0 ) − i ( x0 , y0 ).
i ∂y ∂y ∂y ∂y
This time we have found that
∂v ∂u
f 0 ( z0 ) = ( x0 , y0 ) − i ( x0 , y0 ) (♠)
∂y ∂y

We now have the two expressions (♥) and (♠) for the value of
f 0 (z0 ) depending on how we allow z → z0 . But from the definition
of complex differentiability, if the function f (z) is differentiable at
z = z0 then the values of f 0 (z0 ) must be the same however we let
z → z0 . Hence the two forms (♥) and (♠) must be identical so
comparing the real and imaginary parts gives us

∂u ∂v ∂v ∂u
( x0 , y0 ) = ( x0 , y0 ) and ( x0 , y0 ) = − ( x0 , y0 ).
∂x ∂y ∂x ∂y

Hence our observation in Example 9.11 that these equations hold


for exp(z) is not peculiar to the exponential function; in fact they
have a much wider application in as much that they are satisfied for
any complex differentiable function. This result can be summarised
in an important theorem for complex differentiable functions:
148 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

Theorem 9.12. Let f : A ⊂ C → C be a function, with A an open


set and let us write f (z) = f ( x + iy) = u( x, y) + iv( x, y).

1. Let z0 = x0 + iy0 .
If f 0 (z0 ) exists then f is differentiable in the sense of real variables
at ( x0 , y0 ) and u, v satisfy

∂u ∂v ∂u ∂v
( x0 , y0 ) = ( x0 , y0 ) and ( x0 , y0 ) = − ( x0 , y0 ).
∂x ∂y ∂y ∂x

These two partial differential equations are called the Cauchy–


Riemann equations.

2. If f 0 (z0 ) does exist, then there are four expressions involving


partial derivatives of u and v to express it, (where each partial
derivative is evaluated at ( x0 , y0 ))

∂u ∂v ∂v ∂u ∂u ∂u ∂v ∂v
f 0 ( z0 ) = +i = −i = −i = +i .
∂x ∂x ∂y ∂y ∂x ∂y ∂y ∂x

3. The function f is analytic on A if and only if ∂u ∂u ∂v ∂v


∂x , ∂y , ∂x and ∂y
exist on A, are continuous on A, and satisfy the Cauchy–Riemann
equations on A.

It follows, by the calculations of Example 9.11, that the exponen-


tial function is analytic on any open set of C and that
d exp(z) ∂u ∂v
= +i = e x cos y + ie x sin y = exp(z),
dz ∂x ∂x
as claimed in Example 9.8.

Exercise 9.4.1. Show that f 1 (z) = c, f 2 (z) = z, f 3 (z) = z2 satisfy the


Cauchy–Riemann equations on C but that f 4 (z) = z = x − iy (complex
conjugate) does not. Deduce that f 4 (z) = z is not analytic.

The Cauchy-Riemann equations can be used to prove some very


powerful results for complex functions.

Example 9.13. Suppose that an analytic function f (z) = u( x, y) +


iv( x, y) is of a constant modulus; that is | f |2 = u2 + v2 is constant.
Differentiating this equation with respect to x and y in turn gives

2uu x + 2vv x = 0 and 2uuy + 2vvy = 0. (1, 2)

By the Cauchy-Riemann equations u x = vy and uy = −v x so (2) becomes

−uv x + vu x = 0. (3)

Now taking u × (1) + v × (3) and v × (1) − u × (3) give

( u2 + v2 ) u x = 0 and (u2 + v2 )v x = 0.

There are now two options. Either u2 + v2 = 0 in which case u = v = 0


and f (z) is the zero function. Alternatively we have u x = v x = 0 in
mathematical methods 2 149

which case the Cauchy-Riemann equations tell us that uy = vy = 0


as well. Now if u x = uy = 0 then u( x, y) is a constant and a similar
argument holds for v( x, y); hence f is a constant function.
Thus we conclude that if an analytic function is of constant modulus
then it must be a constant with zero derivative.

This is just one relatively simple example of some very surpris-


ing properties of analytic functions. We can use these ideas for all
manner of tasks that seem unrelated to complex functions - one
common application is in the evaluation of difficult integrals that
cannot be tackled with any of the standard methods. Some of the
more advanced uses of analytic functions will be discussed else-
where.

9.5 Solutions of Laplace’s equation

Laplace’s equation2 , which is a partial differential equation for a real 2


which has nothing to do with the
function of two real variables, φ( x, y), on a region of R2 , is given by Laplace transform, except for the
name!
∂2 φ ∂2 φ
+ 2 = 0.
∂x2 ∂y
In later units we will see that Laplace’s equation occurs in many
physical applications and has particular relevance to a suite of
problems arising in fluid mechanics and electromagnetism. Here
we are not concerned with these important uses of Laplace’s equa-
tion but rather want to explore its close connection to analytic com-
plex functions.
If we differentiate the first Cauchy–Riemann equation with re-
spect to x and the second with respect to y, then we have u xx = v xy
and v xy = −uyy . If we eliminate v xy it follows that

∂2 u ∂2 u
+ 2 = 0;
∂x2 ∂y
in other words the real part of an analytic function satisfies Laplace’s
equation. On the other hand, we can also differentiate the two
Cauchy–Riemann equations with respect to y and x in turn to
obtain u xy = vyy and vyy = −u xy . Now if we eliminate u xy we
conclude that
∂2 v ∂2 v
+ =0
∂x2 ∂y2
as well.
Arbitrary sums of these functions are also solutions, because the
Laplace equation is linear and homogeneous.

Exercise 9.5.1. Show that if φ1 ( x, y) and φ2 ( x, y) are solutions to


Laplace’s equation, then aφ1 + bφ2 (where a, b ∈ R) is also a solution.

Example 9.14. Show that the real and imaginary parts of z2 satisfy
Laplace’s equation.
Solution: As z2 = ( x + iy)2 = x2 − y2 + 2ixy we have that u( x, y) =
x2 − y2 and v( x, y) = 2xy. Simple differentiation gives that u xx = 2 and
150 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

uyy = −2 so u xx + uyy = 0 and Laplace’s equation is satisfied. Similarly


v xx = vyy = 0 so the imaginary part also satisfies Laplace’s equation.

Exercise 9.5.2. Show that the real and imaginary parts of the function
f (z) = z3 satisfy Laplace’s equation.

Exercise 9.5.3. Show that e ay cos( ax ), a ∈ R is a solution to Laplace’s


equation. What other multiplicative combinations of trigonometric and
exponential functions are solutions?

One interesting application of Laplace’s equation is the obser-


vation that given a solution u( x, y) we can construct a complex
analytic function f (z) whose real part is u( x, y).

Theorem 9.15. Both the real and imaginary parts of any analytic
function, considered as functions of x and y, are solutions to Laplace’s
equation:

∂2 u ∂2 u ∂2 v ∂2 v
+ 2 =0 and + 2 = 0.
∂x2 ∂y ∂x 2 ∂y

Moreover if φ( x, y) is a solution to Laplace’s equation, then there


exists a complex analytic function f (z) whose real part is φ( x, y) and
a complex analytic function g(z) whose imaginary part is φ( x, y).

x
Example 9.16. Show that u( x, y) = x 2 + y2
satisfies Laplace’s equation
for ( x, y) 6= (0, 0). Find the corresponding analytic f (z) having real part
u( x, y).
x
Solution: If u = x 2 + y2
then elementary differentiation gives

∂u y2 − x 2 ∂2 u 2x3 − 6xy2
= 2 ⇒ = 2
∂x ( x + y2 )2 ∂x 2 ( x + y2 )3
and
∂u 2xy ∂2 u −2x3 + 6xy2
=− 2 ⇒ = .
∂y ( x + y2 )2 ∂y2 ( x 2 + y2 )3
As u xx = −uyy it follows that u xx + uyy = 0 so u( x, y) satisfies Laplace’s
equation.
To find the complex function f = u( x, y) + iv( x, y), we attempt to find
v( x, y) so that the Cauchy-Riemann equations are satisfied. This requires

∂v ∂u y2 − x 2
= = 2 (♣)
∂y ∂x ( x + y2 )2
and
∂v ∂u 2xy
=− = 2 (♦).
∂x ∂y ( x + y2 )2
The method to solve this problem is similar to the method for finding
a potential, seen in Chapter 6. Notice that we will determine v up to a
constant, that we can then choose. If we integrate (♦) with respect to x we
have that
y
v( x, y) = − 2 + Q(y)
x + y2
mathematical methods 2 151

for some function Q(y) which is determined by satisfying (♣). Now

∂v y2 − x 2 dQ ∂u y2 − x 2
= 2 + = = .
∂y ( x + y2 )2 dy ∂x ( x 2 + y2 )2

Hence dQ/dy = 0 and Q is a constant, which we choose to be zero. Then

x − iy
f = u + iv =
x 2 + y2

and if z = x + iy then

z̄ z̄ 1
f (z) = = = .
| z |2 zz̄ z

Hence the differentiable function 1/z (for z 6= 0) is analytic with real part
the given u( x, y).

This last example provides us with some motivation that solu-


tions of Laplace’s equation can be associated with complex analytic
functions. In fact, solving Laplace’s equation for practical problems
is often difficult and is frequently best achieved by looking at re-
lated problems expressed in terms of complex variables. Exploring
this topic is beyond the scope of these notes but will be pursued in
related units in later years.
10
Probability and Statistics

This chapter describes probability theory and statistical inference


with an emphasis on reliability methods in engineering applica-
tions. Subsection 10.1 reviews basic probability and introduces
notation for concepts with which you should already be familiar.
The content is necessarily brief due to time constraints and this re-
vision of probability models considers only finite sample spaces.
Sections 10.2 - 10.3 contain the examinable material for this unit.
One of the most important aspects of probability theory is in
the description of random variables. For a long time most scientific
advancements were achieved by assuming a deterministic (non-
random) world. In this case, repeated experiments ought to give
rise to the same results. However, it was noticed that they did not.
The same experiment, repeated under identical conditions, could
give rise to different results. This variability of observed results
is described well by probability theory and random variables. We
can think of random variables as describing the results of an ex-
periment that is affected by some chance mechanism. Section 10.2
discusses probability models using the concept of random variables.
Subsection 10.2.1 examines two discrete random variables useful
for reliability engineering – the Bernoulli and binomial random vari-
ables. Subsection 10.2.2 discusses two continuous random variables
suitable for modelling reliability – the exponential and normal ran-
dom variables.
Section 10.3 introduces statistical inference. Probability theory is
useful at approximating many real world observations, if we know
the true probability model. Often we do not. If someone asks you to
bet on the toss of a coin, you may assume the chance is 50%, but
you do not know it. Perhaps the person who offers the bet is un-
savoury and uses a biased coin designed to weigh the odds against
you. Statistical inference is concerned with the use of observed
sample results (observed values of random variables) in making
statements about the unknown parameters of a probability model.
In effect there are a collection of candidate probability models
(P(heads) = 0.2, P(heads) = 0.5, P(heads) = 0.8 etc.), and statis-
tics is the process by which observed data can help identify the
most plausible of these probability models. Section 10.3.1 describes
one method of estimating probability models – obtaining point and
154 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

interval estimates of the model parameters. The task of the statis-


tician is often to assess whether or not a hypothesised value of the
model parameter is likely (e.g. P(heads) = 0.5? before making your
bet). Hypothesis testing is the process of deciding whether the data
provides evidence against statements and is discussed in Section
10.3.2.

10.1 A recap on probability models

A formal definition of a probability model begins with an experiment.


Definition 10.1. An experiment is any operation whose outcome is not
known with certainty.
The precise outcome of the experiment is not known with cer-
tainty but the possible outcomes are. This is the sample space of the
experiment.
Definition 10.2. The sample space S of an experiment is the set of all
possible outcomes from the experiment.
For example, if the experiment is coin tossing then S = { head, tails}.
If the experiment is die rolling then S = {1, 2, 3, 4, 5, 6}. The expe-
riment may be the weather tomorrow, in which case we must be
more careful how to construct the sample space. One possibility
may be S = {rain, snow, clear }. The sample space can be any set of
objects (e.g. numbers, figures, letters, points) and should include
those things which we are interested in. For example, in the case of
the weather tomorrow then rain, snow and clear are in S but tomor-
row’s opinion poll of the government is not.
A probability model also requires a collection of events that may
or may not occur:
Definition 10.3. An event is a subset of the sample space. An event
occurs if any one of its elements is the outcome of the experiment.
In the following two definitions, A and B are two events from
the same sample space S. For instance, consider the experiment
of die rolling and let A be the event that an even number is on
the upward facing side (A = {2, 4, 6}) and B be the event the die
shows 4 or more (B = {4, 5, 6}), of the common sample space
S = {1, 2, 3, 4, 5, 6}. Using these two events we will introduce some
notation you will already have seen.
Definition 10.4. The union of A and B (written A ∪ B) is the event
which consists of all the outcomes that belong to A or B or to both.
In the die rolling experiment above A ∪ B = {2, 4, 5, 6} and A ∪ B
occurs if the rolled die shows any one of these four outcomes.
Definition 10.5. The intersection of A and B (written A ∩ B) is the
event which consists of all the outcomes that belong to both A and B.
In the die rolling experiment A ∩ B = {4, 6} and A ∩ B occurs if
the rolled die shows a 4 or a 6.
mathematical methods 2 155

Definition 10.6. The complement of A with respect to S (written A) is


the event which consists of all the outcomes not belonging to A.
In the die rolling experiment A = {1, 3, 5} and A occurs if the
rolled die shows any one of these three outcomes.
Note also that S itself is an event as is the empty set ∅ = {}.
With possible outcomes described we assign probabilities to each
event. Formally, a probability function, P, is a real-valued function
defined on a set of subsets of the sample space S. Given S, and an
event A in S, we define P( A) to be the probability that A occurs. A
probability function satisfies the following three axioms:

Axiom 1: P(S) = 1. If we are tossing a coin the probability


This says it is certain something must happen. of the event S = {heads, tails} is 1 (that
event must occur).
Axiom 2: 0 ≤ P( A) ≤ 1, for each subset A of S.
This axiom says all probabilities are measured on a scale of 0
to 1, where 0 means impossible and 1 means certain.

Axiom 3: If A1 , A2 , . . . , An is a collection of disjoint subsets


of S then

P ( A1 ∪ A2 ∪ . . . ∪ A n ) = P ( A1 ) + P ( A2 ) + . . . + P ( A n ).

Axiom 3 is more subtle than the first two, and is known as the
additivity property of probability. It says we can calculate probabili-
ties of complicated events by adding up the probabilities of smaller
events provided the smaller events are disjoint and together make
up the entire complicated event. When we say disjoint we mean the
events do not intersect. This axiom can be extended to a countable
sequence of disjoint events A1 , A2 , . . . and is needed in general but
not for this unit.
Axioms 1-3 imply other basic properties or theorems that are
true for any probability function. We list four now.
Theorem 10.7. P(∅) = 0.
Proof. S ∪ ∅ = S and thus P(S ∪ ∅) = P(S) = 1. But S ∩ ∅ = ∅ so
that P(S ∪ ∅) = P(S) + P(∅) = 1 + P(∅) by Axioms 3 and 1. Thus
1 + P(∅) = 1 and P(∅) = 0.

Theorem 10.8. P( A) = 1 − P( A).


Proof. A ∪ A = S so P( A ∪ A) = P(S) = 1 by Axiom 1. But
A ∩ A = ∅ and therefore P( A ∪ A) = P( A) + P( A) by Axiom 3.
Thus P( A) + P( A) = 1 and the result follows immediately.

Theorem 10.9. P( A ∩ B) = P( B) − P( A ∩ B).


Proof. B = ( A ∩ B) ∪ ( A ∩ B) so P( B) = P(( A ∩ B) ∪ ( A ∩ B)). Since
( A ∩ B) ∩ ( A ∩ B) = ∅, P( B) = P( A ∩ B) + P( A ∩ B) and the result
follows immediately.
156 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

Theorem 10.10. P( A ∪ B) = P( A) + P( B) − P( A ∩ B).

Proof. A ∪ B = A ∪ ( A ∩ B) so P( A ∪ B) = P( A ∪ ( A ∩ B)). Since


A ∩ ( A ∩ B) = ∅ then

P( A ∪ ( A ∩ B)) = P( A) + P( A ∩ B)
= P( A) + P( B) − P( A ∩ B)

by Theorem 10.9. Thus

P ( A ∪ B ) = P ( A ) + P ( B ) − P ( A ∩ B ).

There are many applications where we know that an event B


has occurred and we are asked to calculate the probability that
another event A also occurs. This requires careful consideration
of the sample space. Implicit in our probability model above, the
probabilities are conditioned on S. To ask the probability of event
A of drawing an ace of clubs from a pack of cards is meaningless
unless we specify a suitable sample space. Perhaps we define S
to represent a canasta deck of cards (actually 2 standard 52 card
packs plus 4 jokers). We are then asking what is the probability
of drawing an ace of clubs, given the deck of cards is a canasta deck.
In this case, for well shuffled canasta deck, the probability will
equal 2/108 ≈ 0.0185. But what if we are told that the pack was
the standard pack of 52 cards (in other words, we keep from S just
one standard pack, and that is event B). The probability now (the
probability of A conditioned to event B) would be 1/52 ≈ 0.0192.
This is intuitively sensible and obvious but the concept is important
for conditional probability.

Definition 10.11. The conditional probability of A occurring, given


that B has occurred is P( A| B) = P( A ∩ B)/P( B) if P( B) > 0. If
P( B) = 0 we define P( A| B) = 0.

Thinking of probability as an area on a Venn diagram, Defini-


tion 10.11 says that the conditional probability of A given B is the
proportion of B inhabited by A ∩ B. In other words, S has been
reduced to the outcomes contained in B. If P( A| B) = P( A) then
knowledge of the occurrence of B has provided no additional in-
formation to our knowledge about the uncertainty associated with
the occurrence of A. This motivates the definition of independence in
probability.

Definition 10.12. If A and B are any two events in S we say A is


independent of B if P( A| B) = P( A).

This is equivalent to saying the proportion of A that inhabits S


is equal to the proportion of A ∩ B that inhabits B. If A and B are
independent then P( A ∩ B) = P( A) P( B).
Definition 10.11 leads directly to Bayes’ Theorem. Try to prove it.
mathematical methods 2 157

Theorem 10.13 (Bayes’ Theorem). Suppose we are given k events


A1 , A2 , . . . , Ak such that

1. A1 ∪ A2 ∪ . . . ∪ Ak = S

2. Ai ∩ A j = ∅ for all i 6= j

(that is the events form a partition of S) then for any event B

P( B| A j ) P( A j )
P( A j | B) = , j = 1, 2, . . . , k.
∑ik=1 P ( B | Ai ) P ( Ai )

We will illustrate this theorem with an example.

Example 10.14. We have two machines producing bolts. Machine 1


is quite defective and produces 30% damaged bolts, while Machine 2
produces only 2% damaged bolts. Machine 1 produces half as many bolts
per day as Machine 2. Given a damaged bolt what is the probability that it
has been made by Machine 1?
Solution: We take S to be all the bolts produced by both machines and Ai
the ones produced by Machine i. We also take B to be the set of damaged
bolts. By hypothesis we know that P( B| A1 ) = 30/100 and P( B| A2 ) =
2/100, and that P( A1 ) = 1/3, P( A2 ) = 2/3.
Knowing that the outcome of a specific experiment is a damaged bolt,
the probability that it comes from Machine 1 is

P ( B | A1 ) P ( A1 )
P ( A1 | B ) =
P ( B | A1 ) P ( A1 ) + P ( B | A2 ) P ( A2 )
30/100 · 1/3
=
30/100 · 1/3 + 2/100 · 2/3
= 30/34 ≈ 0.882.

We end this revision with a reminder on the different types of


sample spaces.

Definition 10.15. A discrete sample space is one which has a finite or


countable 1 number of elements. 1
A countable set is a set which has
the same size as the natural numbers,
The discussion up to now has only dealt with discrete sample more precisely, which can be put in
bijection with the natural numbers.
spaces and described experiments in terms of single element events
(outcomes) to which we can assign a probability. There are many
experiments however that identify a sample space consisting of
anywhere along a continuous line. For example, the experiment
may be the life time of a machine part and S has outcomes all of the
points on the non-negative real line.

Definition 10.16. A continuous sample space is one which has as


outcomes all of the points in some interval on the real line.

Although subsets of the continuous interval are events, the out-


comes would be the sets of single points in the interval. There are
an infinite number of points in the interval and the probability of
the occurrence of any one point must be zero. This is more easily
examined with the concept of random variables which we turn to
now.
158 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

10.2 Random variables

The content of the previous section was a reminder of what you


already know about probability models (for finite sample spaces). It
required defining a probability function (a real valued set function
that satisfies the axioms of probability) on S. It turns out there can
be simpler ways of working with probability assignments. Once we
have a probability model as described in Section 10.1 we may define
a random variable for that model. Intuitively, a random variable
assigns a numerical value to each possible outcome in the sample
space. Formally,

Definition 10.17. (Random Variable)


A random variable X is a real-valued function X : S → R with domain
being the sample space S.

Since X (s) is a real valued function the range of X (s) is always


a set of real numbers. The probability of X (s) = x (or X = x for
short) corresponds to the probability of outcome x.

Example 10.18. We throw a red die and a blue die. The outcome of the
experiment is a pair of numbers ( a, b), where a is the number on the red
die and b the number on the blue die. The sample space S is the set

{(1, 1), (1, 2), (1, 3), . . . , (6, 6)}

of all possible pairs of numbers that can occur.


We define the random variable X : S → R defined by X (( a, b)) =
a + b. Then X yields the sum of the values of the red and the blue dice.
The probability that X = 5 can be computed to be 4/36 = 1/9 (count
how many pairs give an outcome sum of 5).

In the remainder of this chapter we will use upper case letters


(for example X, Y, Z) to denote particular random variables and
lower case letters (for example x, y, z) to denote values in the range
of the random variable. This will become important later. Note that
often the outcome of an experiment is a number. In such situations
the introduction of random variables amounts to little more than a
new and convenient notation.
Obviously similar to sample spaces random variables can be
either discrete or continuous

Definition 10.19. A random variable X is discrete if its range forms


a discrete (finite or countable) set of real numbers. A random variable X
is continuous if its range forms a continuous (uncountable) set of real
numbers and the probability of X equalling any single value is 0.

We now discuss these two types of random variables in turn2 . 2


Actually there is a third type - a
mixture of discrete and continuous but
we will ignore these in this course.
mathematical methods 2 159

10.2.1 Discrete random variables


If X is a discrete random variable we can use the probability defined
on the subsets of S to define the probability mass function (p.m.f).

Definition 10.20. (Probability Mass Function)


The probability mass function for a discrete random variable X is a
function (denoted by p X ( x ) or P( X = x )) of a real variable x and is
defined to be

p X ( x ) = P( X = x ) = P({s ∈ S | X (s) = x }) for all real x.

If we have an experiment and sample space S defined for that


experiment and have defined the discrete random variable X on the
elements of S then we can find p X ( x ) as follows. Define the event

A( x ) = {s : X (s) = x } ⊆ S.

Then p X ( x ) = P( A( x )) for all real x. Note that A( x ) may equal


∅ for many x and therefore p X ( x ) is 0 for any such x. From the
axioms of probability models p X ( x ) must satisfy

p X ( x ) ≥ 0 and ∑ p X ( x ) = 1.
all x
From now on we shall write P( X = x ) for P( X (s) = x ) and
suppress the functional dependence of X on the elements S un-
less necessary. We also define the cumulative distribution function
(c.d.f.).

Definition 10.21. (Cumulative Distribution Function)


The cumulative distribution function for a discrete random variable X
is a function FX : R → R of a real variable t defined by

FX (t) = P( X ≤ t) = ∑ p X ( x ).
x ≤t

In other words, FX (t) gives the total accumulation of probability


for X equalling any number less than or equal to t and therefore it
can be proved to have the following properties:

Theorem 10.22. Let FX be the cumulative distribution function of a


discrete random variable, X. Then

1. 0 ≤ FX (t) ≤ 1

2. FX (t1 ) ≤ FX (t2 ) whenever t1 ≤ t2 (i.e. FX is monotone increasing)

3. limt→∞ FX (t) = 1

4. limt→−∞ FX (t) = 0
160 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

5. FX (t) is right-continuous, that is lim∆t→0+ FX (t + ∆t) = FX (t).

Either p X or FX can be used to evaluate probability statements


about X. Often we are interested not only in probabilities that X
lies in certain intervals but also a typical value of X – the expected
value of X. The concept of expectation is fundamental to all of
probability theory. Intuitively, the expected value of a random
variable is the average value that the random variable assumes.
More formally,

Definition 10.23. (Expected value)


For a discrete random variable X with probability mass function p X ( x ),
the expected value (written E[ X ] or µ X ) is defined by

E[ X ] = ∑ xp X ( x ).
range of X

Note the above definition can be generalised to include the aver-


age value of a function of a random variable E[ H ( X )] :

E[ H ( X )] = ∑ H ( x ) p X ( x ).
range of X

For instance E[ X 2 ] = ∑range of X x2 p X ( x ).


The expected value of X is often called the mean and is denoted
by µ X . So the mean of a random variable X is the average of the
values in the range of X, where the average is weighted by the
p.m.f.

Exercise 10.2.1. What is the expected value of the random variable X


from Example 10.18?

The operation of taking the expected value has many convenient


properties.

Theorem 10.24. If X and Y are random variables and c is a real number,

1. E[c] = c,

2. E[cX ] = cE[ X ],

3. E[ X + Y ] = E[ X ] + E[Y ].

Now that we understand expected value we can use it to define


other quantities that provide further information about a random
variable. Given a random variable X, µ X tells us the average value
of X but does not tell us how far X tends to be from µ X . For that
we define the variance of a discrete random variable.
mathematical methods 2 161

Definition 10.25. (Variance, standard deviation)


The variance of a discrete random variable X, denoted by Var( X ) or σX2 is
defined by

Var( X ) = E[( X − µ X )2 ] = ∑ ( x − µ X )2 p X ( x ),
range of X
p
and the standard deviation is defined by σX = Var( X ).

Intuitively, σX2 and σX are measures of how spread out the distri-
bution of X is or how much it varies. As a measure of variability
the variance is not so intuitive because it is measured in different
units than the random variable, but it is convenient mathemati-
cally. Therefore, the standard deviation σX is also defined and is the
square root of the variance, providing a measure of variability in
the same units as the random variable. The variance of a random
variable plays a very important role in probability theory and statis-
tical inference so we pause to present some properties of variances
(without proof).

Definition 10.26. (Independent variables)


Two discrete random variables X and Y are independent if

P( X = x and Y = y) = P( X = x ) P(Y = y)

for all x, y.

Theorem 10.27. Let X be any random variable with µ X = E[ X ] and


variance σX2 = Var( X ). Then the following hold true:

1. Var( X ) ≥ 0.

2. If a and b are real numbers then Var( aX + b) = a2 Var( X ).

3. Var( X ) = E( X 2 ) − µ2X = E[ X 2 ] − E[ X ]2 .

4. Var( X + Y ) = Var( X ) + Var(Y ) for Y a random variable indepen-


dent of X.

We are now equipped with enough information about random


variables to examine some important discrete random variables,
and their applicability to reliability engineering.

Bernoulli random variables


Many problems in reliability engineering deal with repeated trials.
For example, we may want to know the probability 9 out of 10
machines will be operating tomorrow or the probability that at least
1 nuclear power plant out of 15 experiences a serious violation. To
address these problems we first define a Bernoulli trial.
162 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

Definition 10.28. A Bernoulli trial is an experiment which has two


possible outcomes, generally called success and failure.

So S = {success, f ailure} and we define a random variable Y by


Y ( f ailure) = 0 and Y (success) = 1. Let p denote the probability of
success and let q denote the probability of failure. Then q = (1 − p).
The random variable Y is said to have the Bernoulli distribution
(written Y ∼ Bern( p)). In other words it has the p.m.f
(
p y q 1− y y = 0, 1.
pY ( y ) =
0 otherwise.

The expected value of Y is

E [Y ] = 1 · p + 0 · q = p

and the variance is calculated by

σY2 = E[Y 2 ] − E[Y ]2 = p(1 − p) = pq

because E[Y 2 ] = 12 · p + 02 · q = p and E[Y ]2 = p2 .

Binomial random variables


So a Bernoulli random variable can address questions such as
“What is the probability that one nuclear power plant fails?” Now,
consider a random variable that describes the outcomes of repeated
Bernoulli trials where we are interested in the probability of getting
x successes in n independent, identical Bernoulli trials. So setting
X = ∑in=1 Yi where Yi the represents the ith Bernoulli trial yields the
binomial random variable.

Definition 10.29. Let X be the total number of successes in n repeated,


independent Bernoulli trials with success probability p. X is called the
binomial random variable with parameters n and p, and written X ∼
Binom(n, p).

If X ∼ Binom(n, p) and q = 1 − p it can be proved that


(
(nx) p x qn− x x = 0, 1, 2, . . . , n.
pX (x) =
0 otherwise.

To understand the binomial p.m.f, if the sample space of each Yi ∼


Bern( p) is Si = {success, f ailure} then the sample space for X ∼
Binom(n, p) is the Cartesian product S = S1 × S2 × . . . × Sn . Since
the trials are independent we assign probabilities to the outcomes
in S by multiplying the values of the probabilities for the individual
Bernoulli trials. Now, each outcome in S must contain exactly x
successes, for some 0 ≤ x ≤ n, so the probability assigned to each
outcome is p x qn− x . So if we count out the number of elements in S
having exactly x success, the product of this number with p x qn− x
will yield p X ( x ) and it is known that the number of elements in S
having exactly x successes is (nx).
mathematical methods 2 163

It is clear that p X ( x ) ≥ 0 for x = (0, 1, 2, . . . , n) and


n n  
n
∑ X p ( x ) = ∑ x p x qn− x = ( p + q )n = ( p + 1 − p )n = 1
x =0 x =0

and therefore p X ( x ) is a p.m.f. Figure 10.1 contains plots of the


binomial [Link] and c.d.f.s.

Figure 10.1: Binomial [Link] (left


panel) and c.d.f.s (right panel) for
n = 20 and p = 0.5 (black) and p = 0.2
0.25

(clear).

1.0
● ● ● ● ● ● ●● ●● ● ● ● ● ●



0.20

0.8 ●


0.15

0.6


p

y
0.10

0.4

● ●


0.05

0.2



0.00


0.0


● ● ● ● ●

0 5 10 15 20 0 5 10 15 20

x x

From Theorem 10.24, Theorem 10.27 and the fact that the Bernoulli
trials are independent it can be shown that if X ∼ Binom(n, p) then
µ X = np and the variance is σX2 = npq by
" #
n n n
µ X = E[ X ] = E ∑ Yi = ∑ E[Yi ] = ∑ p = np.
i =1 i =1 i =1
" #
n n n
σX2 = Var ∑ Yi = ∑ Var[Yi ] = ∑ pq = npq.
i =1 i =1 i =1
To summarise:
Name p.m.f. ( Mean Variance
p y q 1− y y = 0, 1.
Y ∼ Bern( p) PY (y) = p pq
0 otherwise.
(
(nx) p x qn− x x = 0, 1, 2, . . . , n.
X ∼ Binom(n, p) PX ( x ) = np npq
0 otherwise.

We conclude this section with three examples.

Example 10.30. There are 10 machines in a factory. The machines operate


overnight while there are no workers in the factory. History suggests the
probability that each machine still works the next day is p = 0.95 inde-
pendent of the other machines’ operating status. If less than 9 machines
164 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

are working the next morning the factory can’t meet customer demand
and incur severe financial losses. How many machines do we expect to
be working the next day? What is the probability that at least 9 out of 10
machines will work? Provide a measure of variability (standard deviation)
for how many machines are still operating the next day.
Solution: Define X to be the number of machines working the next day
out of a total of 10 machines. Then X ∼ Binom(10, 0.95). We would
therefore expect E[ X ] = 10 × 0.95 = 9.5 machines to be operating the
next day. The probability that at least 9 machines out of 10 are working
the next day is
   
10 10
P( X ≥ 9) = p X (9) + p X (10) = 0.959 0.051 + 0.9510 0.050
9 10
≈ 0.3151 + 0.5987 = 0.9138.

A measure of the variability associated with working machines is σX =



10 · 0.95 · 0.05 ≈ 0.6892

Example 10.31. In the country Bigtrouble, the probability that a nu-


clear power plant experiences a serious violation every 10 years is 0.06.
Bigtrouble has 15 nuclear power plants that operate independently of each
other. How many nuclear power plants do you expect to experience a seri-
ous violation over the next decade? Even one serious violation may result
in untold catastrophe. What is the probability of experiencing at least one
serious violation?
Solution: Define X to be the number of nuclear power plants that experi-
ence a serious violation in a decade in Bigtrouble. X ∼ Binom(15, 0.06)
and we would expect E[ X ] = 15 × 0.06 = 0.9 power plants to experience
a serious violation. The probability of at least one power plant suffering a
serious violation is
 
15
P ( X ≥ 1) = 1 − p X (0) = 1 − 0.060 0.9415 ≈ 0.6047.
0
Residents should be nervous.

Example 10.32. As a final example consider the reliability of systems


built from several manufactured components of similar structure. System
1 has three components linked in such a way that the system is operating
if at least two of the three components are operating. In other words,
System 1 has one redundant component. Suppose there is another system,
System 2, that has two components of similar type to System 1. System
2 is operating only if both components are operating – System 2 has no
redundant components. What is the effect of redundancy on the reliability
of the system?
Solution: Suppose all the components in System 1 and System 2 operate
independently of one another and have the same probability of operating,
p. Denote X to be the number of operating components in System 1, so
X ∼ Bin(3, p). Then the probability System 1 operates is

P( X ≥ 2) = p X (2) + p X (3) = 3p2 q + p3 .

Now consider System 2. Denote W to be the number operating compo-


nents in System 2 then W ∼ Bin(2, p) and the probability the system is
mathematical methods 2 165

working is pW (2) = p2 . If we compare System 1 (which has a redundant


component) with System 2 we can assess the value of redundancy in the
reliability of System 1. This can be based on the ratio of the above two
probabilities , i.e.

( p X (2) + p X (3))/pW (2) = p2 ( p + 3q)/p2 = 3 − 2p.

Observe that the ratio is close to 3 for small p and close to 1 for large p
illustrating the general principle of reliability: that redundancy improves
system reliability when components are ‘unreliable’ but there is little
advantage in having redundancy when the components are highly reliable.

10.2.2 Continuous random variables


In the previous section we considered discrete random variables
for which P( X = x ) > 0 for some x. However, for some random
variables we also have P( X = x ) = 0 for all x ∈ R. Indeed, recall
the definition of a continuous random variable as stated in Definition
10.19. The immediate consequence is that we must focus on events
like X ∈ [ a, b] where [ a, b] is an interval of length b − a > 0. Note
first, if X is a continuous random variable then

P ( a ≤ X ≤ b ) = P ( a < X ≤ b ) = P ( a ≤ X < b ) = P ( a < X < b ).

Instead of considering the probability mass function p X ( x ) as for


discrete random variables we define the probability density func-
tion (p.d.f.).

Definition 10.33. (Probability Density Function)


A probability density function is a nonnegative function f such that its
integral from −∞ to ∞ is equal to 1. That is,
Z ∞
f ( x ) ≥ 0, for all x, and f ( x )dx = 1.
−∞

If X is a continuous random variable such that


Z b
P( a ≤ X ≤ b) = f ( x )dx,
a

then we say X has probability density function f and we denote f by f X .

In particular, if b = a + δ with δ a small positive number, and if f


is continuous at a, then we see that
Z a+δ
P( a ≤ X ≤ a + δ) = f X ( x )dx ≈ δ f ( a).
a

Definition 10.34. (Cumulative Distribution Function)


The cumulative distribution function for a continuous random variable
X with p.d.f. f X is the function FX : R → [0, 1] of a real variable t defined
166 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

by
Z t
FX (t) = P( X ≤ t) = f X ( x )dx.
−∞

Note that the properties of c.d.f.s for continuous random vari-


ables are the same as those listed for discrete random variables in
Theorem 10.22, with the stronger property that FX (t) is continu-
ous instead of right-continuous. Also recall the definitions for the
mean and variance of discrete random variables. We get similar
definitions for continuous random variables.

Definition 10.35. (Expected value, Variance)


If a continuous random variable X has p.d.f. f X , then its expected value
(written E[ X ]) is defined by
Z ∞
E[ X ] = µ X = x f X ( x )dx,
−∞

its variance is defined by


Z ∞
σX2 2
= Var( X ) = E[( X − µ X ) ] = ( x − µ X )2 f X ( x )dx,
−∞
p
and its standard deviation is defined by σX = Var( X ).

The following definition holds for both continuous and discrete


random variables.

Definition 10.36. (Independent Variables)


Let X and Y be two random variables. We say X and Y are independent
if
P( X ≤ x, Y ≤ y) = P( X ≤ x ) P(Y ≤ y) ∀ x, y.

Note that Theorem 10.24 and Theorem 10.27 hold also for contin-
uous random variables.

Exponential random variables


An important continuous random variable in reliability is one
with the exponential distribution. Let T be a random variable that
measures the length of time until a certain event occurs. Since time
is measured continuously T is a continuous random variable with
its range the positive numbers. The “rate” of occurrence of the
event is governed by a parameter λ > 0.

Definition 10.37. A random variable with the exponential distribu-


tion with rate parameter λ > 0 (written T ∼ Exp(λ)) has a probability
mathematical methods 2 167

density function (
λe−λx x>0
f T (x) = (10.1)
0 x≤0
Clearly, f T ( x ) ≥ 0 for all x and
Z ∞ Z ∞
f T ( x )dx = λe−λx dx = [−e−λx ]0∞ = 1
−∞ 0

so f T ( x ) is a p.d.f. The c.d.f. is given by


Z t Z t
FT (t) = f T ( x )dx = λe−λx dx = [−e−λx ]0t = 1 − e−λt for t > 0,
−∞ 0

and FT (t) = 0 for t ≤ 0.

Figure 10.2: Exponential p.d.f.s (left


panel) and c.d.f.s (right panel) for
λ = 0.5 (red), λ = 1 (blue) and λ = 1.5
(black)
1.0
1.5

0.8
0.6
1.0

F
f

0.4
0.5

0.2
0.0

0.0

0 1 2 3 4 5 0 1 2 3 4 5

t t

Inspection of Figure 10.2 shows that as λ gets larger the proba-


bility of lasting to a time t before the occurrence of the event tends
toward zero. Another way to see this is through the expected value
of T. Integration by parts yields the mean of an exponential random
variable as
Z ∞ Z ∞
µT = E[ T ] = x f T ( x )dx = xλe−λx dx
−∞ 0
Z ∞ ∞
i∞ e−λx

h 1
= − xe−λx + e−λx dx = 0 + = .
0 0 −λ 0 λ
So as λ increases our expected time to the event shrinks.
We now compute Var( T ) = E[ T 2 ] − E[ T ]2 . Using integration by
parts yields
Z ∞ Z ∞
E[ T ] 2
= 2
x f T ( x )dx = x2 λe−λx dx
−∞ 0
i∞ Z ∞
h 2 2
= − x2 e−λx +2 xe−λx dx = 0 + E[ T ] = 2 ,
0 0 λ λ
168 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

so
2 1 1
Var( T ) = E[ T 2 ] − E[ T ]2 = 2
− 2 = 2
λ λ λ
and
1
q
σT = Var( T ) = .
λ
To summarise:

For T ∼ Exp(λ):
(
λe−λx x>0
f T (x) =
0 x≤0
(
1 − e−λt t>0
FT (t) =
0 t≤0
1 1
E[ T ] = and Var( T ) = 2 .
λ λ

An exponential distribution can be used to model life lengths, as


in the following example.

Example 10.38. Past experience tells us the life time in hours of a certain
type of lightbulb produced by a manufacturer follows an Exp(λ) distribu-
tion, where λ = 0.001. How many hours do we expect a light bulb to last?
What is the standard deviation of the life time of a lightbulb? What is the
probability that a light bulb lasts more than 900 hours?
Solution: Let T denote the life time in hours of a light bulb. We know
T ∼ Exp(λ) and λ = 0.001. Therefore, the expected lifetime is 1/λ =
1/0.001 = 1000 hours. Also, the standard deviation is 1/λ = 1000
hours.
Finally, because T ∼ Exp(0.001)

P( T ≥ t) = 1 − FT (t) = e−λt for positive t.

Therefore the probability a randomly selected light bulb lasts longer than
900 hours is

P( T ≥ 900) = 1 − FT (900) = e−0.001∗900 = e−0.9 ≈ 0.4067

Normal random variables


The normal p.d.f. is the most commonly used of all p.d.f.s. This
is not only because it addresses a lot of practical problems but also
because of the Central Limit Theorem and its approximation to a
large number of other probability models. We will return later to
this in Section 10.3.
mathematical methods 2 169

Definition 10.39. (Normal distribution)


A random variable X has a normal distribution with parameters µ and
σ2 if and only if its probability density function is

1 2 /2σ2
f X (x) = √ e−( x−µ)
σ2 2π

for all real x. We write X ∼ N (µ, σ2 ). It can be computed that E[ X ] = µ


and Var( X ) = σ2 . N (0, 1) is called the standard normal distribution.

Inspection of the p.d.f. and Figure 10.3 shows a normal random


variable has a p.d.f. that is bell-shaped centred at µ and the spread
of the distribution is governed by σ.

Figure 10.3: Normal p.d.f.s (left panel)


and c.d.f.s (right panel) for µ = 0 and
σ = 0.5, (black) σ = 1 (blue) and σ = 2
(red).
0.8

1.0
0.8
0.6

0.6
0.4
f

0.4
0.2

0.2
0.0

0.0

-4 -2 0 2 4 -4 -2 0 2 4

x x

The normal c.d.f.


Z t
1 2 /2σ2
FX (t) = √ e−( x−µ) dx (10.2)
−∞ σ2 2π
has no closed form (approximations of examples are shown in
Figure 10.3). However, most text books and computer packages
provide approximations to (10.2) when µ = 0 and σ = 1. This is
because of the following.
Theorem 10.40. Suppose X ∼ N (µ, σ2 ). Then the random variable
X −µ
Z = σ has standard normal distribution N (0, 1).
Corollary 10.41. If X ∼ N (µ, σ2 ) then
P( a ≤ X ≤ b) = P ( Z ≤ (b − µ)/σ ) − P ( Z ≤ ( a − µ)/σ )
b−µ a−µ
   
= FZ − FZ .
σ σ
170 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

Theorem 10.40 and Corollary 10.41 say that we can write proba-
bility statements about any normally distributed random variable
in terms of the standard normal c.d.f. Section 10.3.3 at the end of
this chapter contains approximations of the P( Z < z) = FZ (z) for
z > 0 for a standard normal random variable. Some equalities will
be useful to remember: P( Z > z) = 1 − P( Z < z) and if −z < 0,
P ( Z < − z ) = P ( Z > z ) = 1 − P ( Z < z ).

Example 10.42. Let X be a random variable with a normal distribution


with mean 7 and variance 4. What is P( X < 2)?
Solution: We have µ = 7 and σ = 2 for X. Define the related random
variable Z = ( X − 7)/2, then P( X < 2) = P( Z < −2.5). By Theorem
10.40, Z has a standard normal distribution. We have P( Z < −2.5) =
P( Z > 2.5) = 1 − FZ (2.5). From the table we get FZ (2.5) = 0.9938, so
P( X < 2) = 1 − 0.9938 = 0.0062.

Exercise 10.2.2. Let X ∼ N (µ, σ2 ). What is the probability that (i)


X ∈ (µ − σ, µ + σ ), (ii) X ≤ µ + 2σ, (iii) X ≥ µ + 3σ?

In several problems in statistical inference the probability to a


standard normal is given, e.g. P( Z ≤ z) = 0.95, and it is asked what
is the corresponding value of z, e.g. z ≈ 1.64 in this case.

Definition 10.43. (Quantile)


For 0 ≤ α ≤ 1, let zα be the real number, called (1 − α)-quantile, such
that
1 − α = P( Z < zα ) = FZ (zα ),
where Z has the standard normal distribution. In other words

P( Z > zα ) = α.

Example 10.44. Determine the 90%-quantile of the standard normal


distribution.
Solution: The 90%-quantile corresponds to α = 0.1. So we are looking
for z0.1 such that 1 − 0.1 = 0.9 = P( Z < z0.1 ) = FZ (z0.1 ). Looking it up
in the table, we see that z0.1 = 1.28.

Exercise 10.2.3. Let Z ∼ N (0, 1). Deduce from the standard normal
table the (1 − α)-quantiles for α = 0.05, α = 0.025 and α = 0.0005.

We finish with an example where normal random variables are


used in reliability engineering.

Example 10.45. The length of screws produced by a machine are not


all the same but rather follow a normal distribution with mean 75 mm
and standard deviation 0.1 mm. If a screw is too long it is automatically
rejected. If 1% of screws are rejected, what is the length of the smallest
screw to be rejected?
Solution: Let X denote the length of a screw produced by the machine.
X ∼ N (µ, σ2 ) where µ = 75 and σ2 = 0.12 . The task is to find P( X >
mathematical methods 2 171

a−µ
a) = 0.01. We know P( X > a) = P( Z > σ ). We also know from the
table that P( Z > 2.33) ≈ 0.01 so
a − 75
≈ 2.33
0.1
and a ≈ 75.233. Thus if 1% of screws get rejected then the smallest to be
rejected would be approximately 75.233 mm.

10.3 Statistical Inference

In Section 10.1 and Section 10.2 we discussed probability theory.


The various calculations associated with the application of prob-
ability theory presented above rely on our knowledge of the true
model parameters. For instance, in Example 10.31 we are informed
the probability of failure for a nuclear power plant is 0.06 and in
Example 10.38 we know the rate parameter for the life time distri-
bution of light bulbs is 0.001. Knowledge of the parameter values
allows us to make probabilistic statements about the uncertainty of
the random variable.
We are now going to discuss statistical inference in which we are
faced with another type of uncertainty – the uncertainty about the
true parameter value in a probability model. In statistics we assume
a certain type of probability model generates observed data and we
use a collection of observations to estimate which particular model
is most plausible, to estimate intervals for plausible values of the
parameters and to test whether certain claims about parameters are
likely. The first two concepts will be examined in Section 10.3.1 and
testing is examined in Section 10.3.2.

10.3.1 Estimation
A very basic concept to statistical inference is that of a random sam-
ple. By sample we mean only a fraction or portion of the whole –
crash testing every car to inspect the effectiveness of airbags is not
a good idea. By random we mean essentially that the portion taken
is determined non-systematically. Heuristically, it is expected that a
haphazardly selected random sample will be representative of the
whole population because there is no selection bias. For the pur-
poses of this course (and often in practice), we will assume every
observation in a random sample is generated by the same probabil-
ity distribution and each observation is made independently of the
others. This motivates the following definition
Definition 10.46. A collection of random variables X1 , X2 , . . . , Xn is
independent and identically distributed (or i.i.d.) if the collection
is independent and if each of the n variables has the same distribution.
The i.i.d. collection X1 , X2 , . . . , Xn is called a random sample from the
common distribution.
Definition 10.47. Any function of the elements of a random sample,
which does not depend on unknown model parameters, is called a statistic.
It is itself a random variable.
172 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

Therefore, we can think of the numbers in our random sample


as being observed values of i.i.d. random variables. Denote X1 to
represent the first measurement we are to make, X2 to represent the
second measurement we are to make etc up to Xn to represent the
nth measurement we are to make. Thus, if X1 , X2 , . . . , Xn constitute
a random sample, the quantities
n
X1 + X2 , ∑ Xi , X n − X1
i =1

are statistics but


Xn − X1 + 2µ X
X1 + µ X , X2 /σX ,
σX
are not.
An estimator is a statistic that is used to provide an estimate to
a parameter in a probability model. We are now going to make a
very important distinction between an estimate and an estimator.
An estimator is a rule or method of obtaining our best guess of the
model parameters and is based on a random sample. An estimate
is the result of applying the estimator to a particular observed
random sample. For instance, you already know that for a random
sample X1 , X2 , . . . , Xn of normal random variables, one estimator of
µ is
X n = ( X1 + X2 + . . . + Xn )/n.
So X n is our estimator (rule for estimating) for µ from any normal
random sample of size n. Once we observe the random sample we
know that X1 = x1 , X2 = x2 , . . . , Xn = xn and then

x n = ( x1 + x2 + . . . + xn )/n

is our observed estimate for that particular data set. Of course, if


we take another random sample, X n will realise another value, x ∗n
say, different to x n . The estimator X n is itself a random variable.
It turns out that X n is a very important estimator (for reasons we
shall examine shortly) so to develop the ideas of estimation and
hypothesis testing in this section and the next we shall focus only
on X n .

Example 10.48. The manager of an automobile factory wants to know


on average how much labour time is required to produce an order of auto-
mobile mufflers using a heavy stamping machine. The manager decides to
collect a random sample of size n = 20 orders to estimate the true mean
time in hours taken to produce the mufflers. The data are

2.58 2.58 1.75 0.53 3.29 2.04 3.46 2.92 3.10 2.41
3.89 1.99 0.74 1.59 0.35 0.03 0.52 1.42 0.04 4.02

Give an estimate for µ, assuming that each time is a normal N (µ, σ2 ).


Solution: Assume that each Xi ∼ N (µ, σ2 ) (i.i.d. for i = 1, 2, . . . , 20)
represents hours worked for each order of muffler. Then x1 = 2.58, x2 =
2.58, . . . , x20 = 4.02 is the observed random sample and we estimate the
true mean µ as x20 = 1.9625.
mathematical methods 2 173

Example 10.48 provides a point estimate of the true mean, but if


we collected another set of muffler orders the new estimate would
be different. Obviously, we frequently do not collect multiple ran-
dom samples. With this in mind it is a good idea to provide an
interval of plausible values for the true mean, based on the one
sample observed. In statistics these are called confidence intervals.
We now discuss the rationale behind confidence intervals.
Since X n is a random variable we should ask

1. What is the expected value of X n ?

2. What is the variance or standard deviation of X n ?

3. What is the probability density of X n ?

If the answer to these questions is satisfactory then we can accept


X n is a satisfactory estimator and we can use the results to attach a
confidence interval to the true parameter value. We address these
questions first in the case of a normal random sample.
Suppose Xi ∼ N (µ, σ2 ) i.i.d. for i = 1, 2, . . . , n and our estimator
for µ is X n . Using the properties of expectations in Theorem 10.24
" #
n
1
E[ X n ] = E ∑ Xi
n i =1
1
= ( E[ X1 ] + E[ X2 ] + . . . + E[ Xn ])
n
= µ.

So the expected value of X n tells us if we took random samples re-


peatedly, then the average of the observed estimates would be cen-
tred around µ, the true value of the parameter we are estimating.
This is a nice result. Using the properties of variances in Theorem
10.27 and the fact that each Xi is independent we have that
" #
n
1
Var[ X n ] = Var ∑ Xi
n2 i =1
1
= (Var[ X1 ] + Var[ X2 ] + . . . + Var[ Xn ])
n2
= σ2 /n.

So the variance of X n tells us that as our information in the sam-


ple increases (that is as the number of observations increase) the
variance of the estimator X n shrinks. This property is also nice,
particularly since we know the variance will shrink around the true
mean because E[ X n ] = µ. Finally in the case of a normal random
sample we also have the following theorem.

Theorem 10.49. Suppose X1 , X2 , . . . , Xn are i.i.d. normal ran-


dom variables, each with parameters µ and σ2 . Then Y = X n =
1 n
n ( ∑i =1 Xi ) is a normal random variable with parameters µ and
2
σ /n.
174 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

Theorem 10.49 says for a normal random sample X n ∼ N (µ, σ2 /n).


The figure on the right shows the effect on the p.d.f. of X n as n in-

2.0
creases.
The above outlines some nice properties of the estimator of µ,

1.5
X n , if the random sample is generated from a normal distribution.
What if we wanted to construct an estimator for the mean from
a non-normal random sample? Now the Central Limit Theorem

f(x)

1.0
comes to the rescue. Before we can state the theorem, we need the
following definition.

0.5
Definition 10.50. We say that a sequence Y1 , Y2 , . . . of random variables
converges in distribution to the random variable Y if

0.0
−3 −2 −1 0 1 2 3
lim FYn ( x ) = FY ( x )
n→∞ x

Thick black line shows the p.d.f. of


for all x where FY ( x ) is continuous.
each Xi ∼ N (0, 1). The coloured thin
lines show the effect as n increases on
Omitting the proof, the Central Limit Theorem (CLT) states the the p.d.f. of X n ∼ N (0, 1/n). As n
following. increases the p.d.f.s tighten around the
true value of µ = 0.

Theorem 10.51. (Central Limit Theorem) Let X1 , X2 , X3 . . . be


a sequence of i.i.d. random variables, each with mean µ and variance
σ2 . Define the sequence Z1 , Z2 , . . . of random variables by

Xn − µ
Zn = √ , n = 1, 2, 3, ...
(σ/ n)

where
n
1
Xn =
n ∑ Xi .
i =1

Then Zn converges in distribution to N (0, 1).

Loosely, the CLT states that for any random sample as n gets
larger the distribution of the sample average X n , properly nor-
malised, approaches a standard normal distribution. Even more
˙ N (µ X , σX2 /n) where ∼
loosely, the CLT states, for large n, X n ∼ ˙
denotes approximately distributed. The CLT is one of the most
powerful results presented in probability theory. Obviously, know-
ing the distribution of the sample average is important because
it provides a measure of how precise we believe our estimate of
the model parameter to be. However, more powerfully, the CLT
means that for any random variable we can always perform statis-
tical inference on the mean or expected value. Figure 10.4 shows
an approximate p.d.f. of an exponential random variable and the
standardised estimator X n based on random samples of increasing
size from this random variable. The figure shows that although the
original density of each Xi is exponential, the p.d.f. of the standard-
ised estimator approaches the standard normal distribution as n
increases.
mathematical methods 2 175

Figure 10.4: The upper left hand panel


contains the approximate p.d.f. of
an exponential random variable. The

0.5
approximate p.d.f. of Zn when n = 2 is
shown in the upper right panel, when
n = 20 in the bottom left panel and
1.5

0.4
when n = 2000 in the bottom right
panel. The thick black line shows the
p.d.f. of the standard normal random

0.3
1.0

variable. As n increases on the p.d.f.

fZn
fX

of Zn converge to a standard normal

0.2
p.d.f.
0.5

0.1
0.0

0.0

0 1 2 3 4 5 6 −2 0 2 4 6 8 10 12

x zn
0.4
0.4

0.3
0.3

0.2
0.2
fZn

fZn

0.1
0.1
0.0

0.0

−2 0 2 4 6 −4 −2 0 2 4

zn zn

By providing the p.d.f. of the estimator X n , Theorem 10.49 and


Theorem 10.51 (in the case of large n) can be used to provide ranges
of plausible values of the mean parameter – a confidence interval.

Definition 10.52. Suppose we have a random sample whose probability


model depends on a parameter θ. The two statistics, L1 and L2 , form a
(1 − α) confidence interval if

P ( L1 < θ < L2 ) = 1 − α

no matter what the unknown value of θ.

Therefore, ( L1 , L2 ) form a random interval which covers the true


parameter θ with probability 1 − α. Consider a 95% confidence
interval for µ. In this case 1 − α = 0.95 so α = 0.05. From standard
normal tables we know P(−1.96 < Z < 1.96) ≈ 0.95 (i.e. zα/2 =
1.96) and we can write

σ σ
P( X n − 1.96 √ ≤ µ ≤ X n + 1.96 √ ) ≈ 0.95.
n n
176 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

For an observed random sample, we substitute the value x n for X n


to calculate the confidence interval. It is important to note that for
this observed data set, we cannot say the probability that µ lies in
this interval is 0.95. All we can say is that after repeated samples
we expect 95% of the confidence intervals constructed to contain
the true value of µ. We will now do an example using Bernoulli
random variables.

Example 10.53. A supplier of ammunition claims that of a certain type of


bullet they produce, less than 5% misfire. You decide to test this claim and
observe 1000 bullets fired. Of these 70 misfired. Provide a point estimate
and an interval estimate for the true probability of a given bullet misfiring
and assess whether you believe the suppliers claim.
Solution: Assuming each bullet is fired independently, we can treat each
observation as coming from a Bernoulli trial, with success to represent
a misfiring and failure to represent not misfire. Then Xi ∼ Bern( p)
(i = 1, 2, . . . , 1000) is a random variable representing whether the i-th
bullet misfires. Write
(
1 the bullet misfires
Xi =
0 the bullet does not misfire

We can estimate p by X n which in this case would provide the point


estimate x n = 70/1000 = 0.07. We can construct a 95% confidence
interval as
√ √
( x n − 1.96σ/ n, x n + 1.96σ/ n).

Naturally, we do not know σ but we do know for a random variable from


p
a Bernoulli trial that σ = p(1 − p) so we shall use our estimate of p
p
to calculate σ = 0.07 ∗ (1 − 0.07) ≈ 0.2551. So our 95% confidence
interval is

(0.07 − 1.96 ∗ 0.0081, 0.07 + 1.96 ∗ 0.0081) = (0.0541, 0.0859).

Notice the confidence interval does not extend below 0.05 and we should
probably doubt the producer’s claim that no more than 5% of their bullets
misfire.

10.3.2 Hypothesis testing


We commence this section on hypothesis testing with a motivating
example.

Example 10.54. Suppose that a processor of dairy products is packaging


cheese into blocks of 1kg. Variability exists of course in this process and
historical information suggests that the standard deviation of the weights
is σ = 10g. We weigh 25 blocks and compute that for these 25 blocks
the observed mean is 1006g. Assess whether the process is continuing to
operate satisfactorily. (We define an unsatisfactory process as a shift in the
mean, with no effect on the standard deviation.)
mathematical methods 2 177

The statistical procedure to assess whether the process is oper-


ating satisfactorily is to take a random sample of blocks of cheese.
Suppose we get the weights x1 , x2 , . . . , xn from the random vari-
ables X1 , X2 , . . . , Xn that describe this process, which are i.i.d. with
unknown mean µ and known standard deviation σ = 10g. The
question then is: “Is the observed sample average sufficiently close
to the desired mean of 1000g that we can regard the operation as
satisfactory, or is there some doubt?”
Consider x n = (∑in=1 xi ) /n as a realised value of X n and we
know by the CLT that X n ∼ ˙ N (µ, σ2 /n). When the process is oper-
ating to satisfaction X n ∼ ˙ N (1000, 102 /n). In our case we observe
n = 25 blocks and see that x n = 1006. Is the process operating
to satisfaction? How do we assess this? We do so by assessing
how unlikely or surprising the observed sample is if the true mean
were really 1000g. The distribution of X n implies regions of low
probability occur in the tails of the p.d.f. of X n and we begin by
examining how far out in the tail of the p.d.f. lies our observation
x n . Figure 10.5 shows the p.d.f. of X n if the process is operating sat-
isfactorily and shows the observed average, x n = 1006, lies far out
in the upper tail. This suggests the observed value is very unlikely
if the process were operating satisfactorily.

Figure 10.5: The p.d.f. of X n if the


process is working satisfactorily. The
observed estimate of x n = 1006 is
0.20

indicated with a black dot and lies far


out in the tail of the p.d.f.
0.15
0.10
fXn

0.05
0.00

990 995 1000 1005 1010

xn

As a first indication that the process is not operating satisfacto-


rily the above is fine but to make a hard decision we must define
more precisely what is “unlikely” or “surprising”. We measure this
by the probability of obtaining a value more extreme than what we
observed. That is

P( X n is more extreme than 1006|µ = 1000)


X n − 1000 1006 − 1000
 
= P √ > √
10/ 25 10/ 25
= P(| Z | > 3)
≈ 0.0026 (from the standard normal table).
Note that the left-hand side probability is read as the “proba-
178 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

bility of obtaining a value of X n more extreme than 1006 (in either


direction) given that µ = 1000.” The smallness of this value reflects
the extremeness of the observed data. On the basis of this probability
we could conclude either

1. the process operation is satisfactory and a rather surprising event


has occurred; or

2. the process operation is not satisfactory because if µ > 1000


then the probability could be larger and the event more likely to
occur.

If we ignore the possibility that a rather unlikely event has oc-


curred, i.e. getting an observed value x n at least as extreme as 1006,
then we would conclude that the process is unsatisfactory.
Now consider two other possible scenarios to assist us in under-
standing the above conclusion: we still test 25 cheese blocks but
obtain different observed means.

• For x n = 1012

P( X n more extreme than 1012|µ X = 1000) = P(| Z | > 6)

which is much smaller than P(| Z | > 4) so we would (perhaps


more confidently) conclude the process is unsatisfactory.

• For x n = 1002

P( X n more extreme than 1002|µ X = 1000) = P(| Z | > 1) = 0.3174

Since the latter probability is not very small there is no cause to


doubt the process operation is satisfactory. That is, if µ = 1000
we are not surprised by the data we have observed. The shaded
area in Figure 10.6 represents the probability that we observe
a value of X n more extreme than 1002 in either direction if the
process is working satisfactorily.

Figure 10.6: The p.d.f. of X n if the


process is working satisfactorily and
x n = 1002. The observed estimate is
0.20

indicated with a black dot and does


not lie far out in the tail of the p.d.f.
The probability of being more extreme
0.15

than 1002 in either direction is the


shaded area.
0.10
fXn

0.05
0.00

990 995 1000 1005 1010

xn
mathematical methods 2 179

The discussion to this point in the present section has been our
motivation for testing hypotheses; now we formalise some ideas.
By a statistical hypothesis we mean a statement about the value of a
specified parameter in a probability model. For example, based on
the probability model described by

X n ∼ N (µ, σ2 /n)

the hypothesis we considered earlier was H0 : µ = 1000. This


is called the null hypothesis. It represents a state of affairs that we
will believe in unless the sample data provides strong evidence
to the contrary. Often the null hypothesis represents the negation
of a claim for which we are seeking evidence from data. In such a
situation the claim itself is the alternative hypothesis, H1 .
The previously considered alternative hypothesis was H1 : µ 6=
1000. This is a two-sided hypothesis since no information was
given, apart from that in the sample data, regarding the direction of
the departure from the null hypothesis. In different circumstances
other possible alternative hypotheses are the one-side hypotheses
H1 : µ > 1000 or H1 : µ < 1000. Assessment of the null hypothe-
sis is then made using the observed value of a suitable test statistic
constructed from the random sample. For our purposes this is just
a standardised version of an estimator of the relevant parameter –

in our case Z = n( X n − µ)/σ. There are two methods used for
hypothesis testing: P-value and critical value.

P-value method
Based on the observed test statistic, z, we determine the P-value,
this being the probability of obtaining a value of the test statistic
at least as extreme as that observed. In determining the P-value we
use one or both tails of the distribution, depending on whether the
alternative hypothesis is one-sided (P( Z > z) or P( Z < z)) or
two-sided (P(| Z | > z)), respectively.

In summary, hypothesis testing with P-value requires

1. Identifying an appropriate probability model – e.g. normal,


exponential, Bernoulli, binomial.

2. Formulating an appropriate null hypothesis H0 and alterna-


tive hypothesis H1 taking into careful consideration whether
H1 should be one-sided (which way?) or two-sided.

3. Constructing the distribution of a suitable test statistic under


H0 .

4. Calculating the P-value based on the data.

5. Stating your conclusion. Reject H0 if the P-value is very small


as this suggests the data are highly unlikely if H0 is true. Do
180 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

not reject H0 if the P-value is not small because the observed


data are not surprising if H0 were true.

6. Interpreting your conclusion in the relevant context.

Often in a testing situation there is a specified value α satisfying


0 ≤ α ≤ 1 called the level of significance. Typically, α is small, con-
vential values being α = 0.05 or α = 0.01. Its role is as a yardstick
against which to judge the size of the P-value:

1. a P-value of less than or equal to α we interpret to mean that the


data cast doubt on H0 or there is sufficient evidence to suggest
H1 might hold;

2. a P-value of greater than or α we interpret to mean that the data


is consistent with H0 or provide not enough evidence to support
the claim expressed by H1 .

If the P-value is less than or equal to α the data is said to be statisti-


cally significant at the level α. Say α = 0.05 then in the cheese example
above if x n = 1006 the P-value is 0.0026 and we would reject the
null hypothesis but if x n = 1002 then the P-value is 0.3174 we
would not. We will now examine an example following the steps
above for hypothesis testing.

Example 10.55. The production manager of bullets claims less than 5%


misfire. You have strong doubts about this claim. You decide to collect
1000 bullets and test how many misfire. You observe 70 out of 1000 bul-
lets misfired. Test whether the manager is correct at the α = 0.05 level of
significance. Note, this example uses the same data and model as Example
10.53 for confidence intervals but we now carry out a hypothesis test.

1. Identify an appropriate probabilty model


We can model these data as Xi ∼ Bern( p) (i = 1, 2, . . . , 1000), rep-
resenting whether the i-th bullet misfires. A sensible estimator of
p is then just X n .

2. Formulate an appropriate null hypothesis H0 and alternative hypothesis


H1 .
We can test the claim by

H0 : p = 0.05
H1 : p > 0.05

The test is one sided because we would primarily be interested


in whether too many bullets misfired. We are not concerned if
the probability of a misfire is less than 0.05, only if it were more
unreliable than claimed and the probability is greater than 0.05.

3. Constructing the distribution of a suitable test statistic under H0 .


mathematical methods 2 181

The relevant test statistic is then

X −p
qn .
p (1− p )
n

Assuming H0 is true p = 0.05 and we have

X − 0.05
q n ∼ N (0, 1)
0.05(1−0.05)

60
1000

50
because the sample size is large enough for the CLT to apply.

40
4. Calculate the P-value based on the data
The observed value of the estimator is x n = 70/1000 = 0.07 and

30
n
fX
the observed test statistic is
0.07 − 0.05

20
p ≈ 2.9019.
0.05(1 − 0.05)/1000

10
Hence from tables at the end of this chapter the P-value is

0
P( X n > 0.07| p = 0.05) = P( Z > 2.9019) = 1 − P( Z <
0.02 0.03 0.04 0.05 0.06 0.07 0.08
2.9019) ≈ 0.0019, where Z ∼ N (0, 1). See the figure on the right
x/n
for a graphical representation of this result.
The p.d.f. of X n if the misfire rate of
5. State your conclusion the bullets were 0.05. The observed
The P-value ≈ 0.0019 is less than α = 0.05 and we reject H0 in estimate x n = 0.07 is indicated with
a black dot and lies far out in the
favour of H1 . tail of the p.d.f. The probability of
lying further out in the upper tail, the
6. Interpret your conclusion in the relevant context. P-value, is the shaded area.
There is sufficient evidence to reject the manager’s claim that
less than 5% of the bullets misfire.

Critical value method


When a significance level α has been specified (e.g. α = 0.05)
there is an alternative way of assessing the extremeness of the ob-
served value of the test statistic. From the appropriate tables we
obtain a so-called critical value which is zα or zα/2 respectively
depending on whether the alternative hypothesis is one-sided
or two sided. In other words, we determine the point zα so that
P( Z > zα ) = α (recall the definition of quantiles). The relevant tail
region, or regions, under the p.d.f. for the standardised test statistic
then constitute what is called the critical region. Then if the observed
value of the test statistic falls anywhere in the critical region we re-
ject the null hypothesis; otherwise we do not. The critical region is
what could also be described as the rejection region.

In summary, hypothesis testing with critical value requires

1. Identifying an appropriate probabilty model – e.g. normal,


exponential, Bernoulli, binomial.
182 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

2. Formulating an appropriate null hypothesis H0 and alterna-


tive hypothesis H1 taking into careful consideration whether
H1 should be one-sided (which way?) or two-sided.

3. Constructing the distribution of a suitable test statistic under


H0 .

4. Calculating the critical value based on the significance level


α.

5. Stating your conclusion based on the data. Reject H0 if the


observed statistic is bigger than the critical value. Do not
reject H0 if the observed statistic is less than the critical value.

6. Interpreting your conclusion in the relevant context.

Example 10.56. Aircrew escape systems are powered by a solid propel-


lant. The burning rate of this propellant is an important product char-
acteristic. Specifications require that the mean burning rate must be 50
centimeters per second. We know that the standard deviation of burning
rate is 4 centimeters per second. The experiementer decides to specify a
level of significance of α = 0.05 and selects a random sample of n = 25
and obtains an observed average burning rate of 51.3 centimeters per sec-
ond. Test whether the burning rate meets specifications versus alternative
that the burning rate does not meet specifications.

1. Identify an appropriate probabilty model


Assume the burning rate of each propellant constitutes a normal
random sample, Xi for i = 1, 2, . . . , n with unknown mean µ and
known standard deviation σ = 4. A sensible estimator of µ is
then X n .

2. Formulate an appropriate null hypothesis H0 and alternative hypothesis


H1 .
We can test the claim by

H0 : µ = 50
H1 : µ 6= 50

The test is two sided because we do not have any extra informa-
tion to specify which way we should test.

3. Constructing the distribution of a suitable test statistic under H0 .


The relevant test statistic is then
Xn − µ
q .
σ2
n

Assuming H0 is true µ = 50, we are told σ = 4 and we have

X n − 50
q ∼ N (0, 1).
16
25
mathematical methods 2 183

4. Calculate the critical value based on the significance level α.


Since the test is two-sided the critical value is zα/2 , Since P(| Z | >

0.5
zα/2 ) = α. From the table zα/2 = z0.025 = 1.96.

0.4
5. State your conclusion
The observed estimate is x n = 51.3 and the observed test

0.3
statistic is
51.3 − 50

fXn
= 1.625.
4/5

0.2
The observed statistic is less than the critical value, so we do not
reject H0 in favour of H1 .

0.1
6. Interpret your conclusion in the relevant context.

0.0
We do not have enough evidence to reject the claim that less ●

burning propellant operates to specification. 47 48 49 50 51 52 53

xn

The p.d.f. of X n if the burning of the


propellant is burning according to
specification. The observed estimate is
Final remarks on hypothesis testing indicated with a black dot and does
not lie far out in the tail of the p.d.f.
1. To reject H0 does not mean that the null hypothesis is false. Only The probability of lying further out in
that the data shows sufficient evidence to cast doubt on H0 . To the tail in either direction is the shaded
area.
not reject H0 does not mean it is true, only that the data shows
insufficient evidence against H0 .

2. As a consequence in any testing situation there are two possible


types of error

• Type I error which occurs when H0 is rejected even though it is


true or
• Type II error which occurs when H0 is not rejected even though
it is false.

3. It is easy to see that

P(Type I error) = P(Reject H0 | H0 is true) = α

However to give a specific value for the probability of a Type II


error requires consideration of a specified alternative value fo the
parameter of interest.

4. The P-value is NOT the probability that H0 is true. It is a proba-


bility calculated assuming the null hypothesis is true. Regard it as
a way of summarizing the extent of agreement between the data
and the model when H0 is true.
184 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

10.3.3 Statistical Tables

Cumulative Standard Normal Probabilities (for


z > 0). The table below gives FZ (z) where Z ∼ N (0, 1)
(see shaded area in the Figure). z

z 0.00 0.01 0.02 0.03 0.04 0.05 0.06 0.07 0.08 0.09
0.0 0.5000 0.5040 0.5080 0.5120 0.5160 0.5199 0.5239 0.5279 0.5319 0.5359
0.1 0.5398 0.5438 0.5478 0.5517 0.5557 0.5596 0.5636 0.5675 0.5714 0.5753
0.2 0.5793 0.5832 0.5871 0.5910 0.5948 0.5987 0.6026 0.6064 0.6103 0.6141
0.3 0.6179 0.6217 0.6255 0.6293 0.6331 0.6368 0.6406 0.6443 0.6480 0.6517
0.4 0.6554 0.6591 0.6628 0.6664 0.6700 0.6736 0.6772 0.6808 0.6844 0.6879
0.5 0.6915 0.6950 0.6985 0.7019 0.7054 0.7088 0.7123 0.7157 0.7190 0.7224
0.6 0.7257 0.7291 0.7324 0.7357 0.7389 0.7422 0.7454 0.7486 0.7517 0.7549
0.7 0.7580 0.7611 0.7642 0.7673 0.7704 0.7734 0.7764 0.7794 0.7823 0.7852
0.8 0.7881 0.7910 0.7939 0.7967 0.7995 0.8023 0.8051 0.8078 0.8106 0.8133
0.9 0.8159 0.8186 0.8212 0.8238 0.8264 0.8289 0.8315 0.8340 0.8365 0.8389
1.0 0.8413 0.8438 0.8461 0.8485 0.8508 0.8531 0.8554 0.8577 0.8599 0.8621
1.1 0.8643 0.8665 0.8686 0.8708 0.8729 0.8749 0.8770 0.8790 0.8810 0.8830
1.2 0.8849 0.8869 0.8888 0.8907 0.8925 0.8944 0.8962 0.8980 0.8997 0.9015
1.3 0.9032 0.9049 0.9066 0.9082 0.9099 0.9115 0.9131 0.9147 0.9162 0.9177
1.4 0.9192 0.9207 0.9222 0.9236 0.9251 0.9265 0.9279 0.9292 0.9306 0.9319
1.5 0.9332 0.9345 0.9357 0.9370 0.9382 0.9394 0.9406 0.9418 0.9429 0.9441
1.6 0.9452 0.9463 0.9474 0.9484 0.9495 0.9505 0.9515 0.9525 0.9535 0.9545
1.7 0.9554 0.9564 0.9573 0.9582 0.9591 0.9599 0.9608 0.9616 0.9625 0.9633
1.8 0.9641 0.9649 0.9656 0.9664 0.9671 0.9678 0.9686 0.9693 0.9699 0.9706
1.9 0.9713 0.9719 0.9726 0.9732 0.9738 0.9744 0.9750 0.9756 0.9761 0.9767
2.0 0.9772 0.9778 0.9783 0.9788 0.9793 0.9798 0.9803 0.9808 0.9812 0.9817
2.1 0.9821 0.9826 0.9830 0.9834 0.9838 0.9842 0.9846 0.9850 0.9854 0.9857
2.2 0.9861 0.9864 0.9868 0.9871 0.9875 0.9878 0.9881 0.9884 0.9887 0.9890
2.3 0.9893 0.9896 0.9898 0.9901 0.9904 0.9906 0.9909 0.9911 0.9913 0.9916
2.4 0.9918 0.9920 0.9922 0.9925 0.9927 0.9929 0.9931 0.9932 0.9934 0.9936
2.5 0.9938 0.9940 0.9941 0.9943 0.9945 0.9946 0.9948 0.9949 0.9951 0.9952
2.6 0.9953 0.9955 0.9956 0.9957 0.9959 0.9960 0.9961 0.9962 0.9963 0.9964
2.7 0.9965 0.9966 0.9967 0.9968 0.9969 0.9970 0.9971 0.9972 0.9973 0.9974
2.8 0.9974 0.9975 0.9976 0.9977 0.9977 0.9978 0.9979 0.9979 0.9980 0.9981
2.9 0.9981 0.9982 0.9982 0.9983 0.9984 0.9984 0.9985 0.9985 0.9986 0.9986
3.0 0.9987 0.9987 0.9987 0.9988 0.9988 0.9989 0.9989 0.9989 0.9990 0.9990
3.1 0.9990 0.9991 0.9991 0.9991 0.9992 0.9992 0.9992 0.9992 0.9993 0.9993
3.2 0.9993 0.9993 0.9994 0.9994 0.9994 0.9994 0.9994 0.9995 0.9995 0.9995
3.3 0.9995 0.9995 0.9995 0.9996 0.9996 0.9996 0.9996 0.9996 0.9996 0.9997
3.4 0.9997 0.9997 0.9997 0.9997 0.9997 0.9997 0.9997 0.9997 0.9997 0.9998
3.5 0.9998 0.9998 0.9998 0.9998 0.9998 0.9998 0.9998 0.9998 0.9998 0.9998
11
Index

alternative hypothesis, 179 estimator, 172 Jacobian matrix, 44


analytic, 145 Euler’s formula, 142
analytic at z0 , 145 Euler’s formulae, 95 Kirchoff’s Current Law, 132
anti-differentiation, 9 Even expansion, 102 Kirchoff’s Laws, 132
antiderivative, 18 Even functions, 99 Kirchoff’s Voltage Law, 132
areas of surfaces, 66 event, 154
expected value, 160, 166
Laplace transform, 109
experiment, 154
Bayes’ Theorem, 157 Laplace Transform Table, 129
exponential distribution, 166
Bernoulli distribution, 162 Laplace’s equation, 149
exponential order, 112
Bernoulli trial, 161 length of a curve, 22
binomial random variable, 162 lengths of curves, 62
first moments, 40
Flux across a curve in R2 , 74
Cauchy–Riemann equations, 148 Flux across a surface in R3 , 77 mean, 160
Central Limit Theorem, 174 Fourier coefficients, 92 mid-point rule, 25
chain rule, 13 Fourier cosine series, 100
change of coordinates, 44, 48 Fourier series expansion, 92 neighbourhood, 144
circulation, 74 Fourier sine series, 101 Newton’s Second Law, 130
closed curve, 73 Fubini’s Theorem, 29, 33, 37 normal distribution, 169
closed surface, 76 Fundamental Theorem of Calculus, null hypothesis, 179
complex Differentiability, 145 19 numerical integration, 25
complex functions, 143
composite quadrature rules, 26 gradient, 77 Odd expansion, 102
conditional probability, 156 Odd functions, 99
confidence intervals, 173 half-range expansion, 102 open set, 83, 144
conservative field, 80 Heaviside function, 122
continuity of complex functions, 144 Hooke’s Law, 130 P-value, 179
continuous random variable, 165 hypothesis testing, 176 partial fractions, 10
convolution, 125
partition, 15, 28
critical value, 181
improper integrals, 23 path integrals, 65
cumulative distribution function,
indefinite integral, 20 Period of a function, 91
159, 165
independence, 156 periodic extension, 98
curl, 84
independent and identically dis- piecewise continuous, 17
cylindrical coordinates, 53
tributed, 171 polar coordinates, 50
Independent variables, 161, 166 potential, 78, 82
De Moivre’s theorem, 141 Integration by parts, 12 probability density function, 165
definite integral, 16 Integration by Substitution, 13 probability function, 155
derivative of transform, 125 inverse function, 7 probability mass function, 159
differentiation, 9 inverse Laplace transform, 109, 114 probability model, 154
discrete random variable, 158 inverse trigonometric functions, 7 product rule, 12
disjoint sets, 155 Properties of Definite Integrals, 18
double integrals, 28 Jacobian, 44 pulse function, 122
186 a. bassom, e. cripps, a. devillers, l. jennings, a. niemeyer, t. stemler, l. stoyanov

quadrature, 25 Simpson’s rule, 25 trapezoidal rule, 25


quantile, 170 spherical coordinates, 55 triple integral, 36
standard deviation, 161, 166 Type I improper integrals, 23
random sample, 171 standard normal, 170 Type II improper integrals, 24
random variable, 158 standard normal distribution, 169
region, 29 statistic, 171
variance, 161, 166
Riemann Sum, 16 statistical hypothesis, 179
vector field, 71
statistical inference, 171
volume by cross-sections, 20
sample spaces, 157 surface integrals, 68
second (principal) moments, 41
simply-connected sets, 86 test statistic, 179 work done, 21

You might also like