0% found this document useful (0 votes)
4 views166 pages

Numerical Techniques

Uploaded by

v.rockwell1987
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views166 pages

Numerical Techniques

Uploaded by

v.rockwell1987
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Numerical Techniques for Chemists and

Chemical Engineers
Solving chemistry and engineering problems using Python

August 31, 2022


Copyright
This document is exclusively meant for students in the Molecular Science and Technology
BSc programme of Delft University of Technology and Leiden University. No part of this
document may be shared to others, electronically, on paper or in any other sneaky way
without permission of the author.

Publisher
First published in November 2020 by F.C. Grozema
Contents

Contents iii

Basic Introduction to Programming in Python 1


1 Introduction to Python 2
1.1 Getting Started with Python . . . . . . . . . . . . . . . . . . . . . . . . . . 2
Installing Python . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
Documentation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
1.2 Basic calculations in Python . . . . . . . . . . . . . . . . . . . . . . . . . . 3
Variables . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4
Comments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5
Printing and formatting text . . . . . . . . . . . . . . . . . . . . . . . . . . 5
1.3 Some remarks on syntax . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
Equality vs. assignments . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
Precedence of operators . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8
1.4 Importing modules . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8
1.5 Exercise: Height of ball . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9

2 Loops and lists 11


2.1 While-loops . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
2.2 Boolean expressions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13
2.3 Lists . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13
2.4 For-loops . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14
2.5 The range construction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15
2.6 Branching: if-else statements . . . . . . . . . . . . . . . . . . . . . . . . . . 16
2.7 Exercise: Money in the bank . . . . . . . . . . . . . . . . . . . . . . . . . . 17
2.8 Exercise: Cooking a perfect egg . . . . . . . . . . . . . . . . . . . . . . . . . 18

3 Structured programming and functions 19


3.1 Simple function . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19
3.2 Optional arguments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20
3.3 Multiple return values . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21
3.4 Exercise: Temperatures . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21
3.5 Exercise: A list as argument . . . . . . . . . . . . . . . . . . . . . . . . . . . 22
Mathematics and data analysis using Python 23
4 Mathematics using Numpy 24
4.1 Basics of Numpy . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24
4.2 Numpy arrays . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24
4.3 Arrays indexing and slicing . . . . . . . . . . . . . . . . . . . . . . . . . . . 26
4.4 Vectorization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27
4.5 The copy issue . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29
4.6 Exercise: Slicing and Copying . . . . . . . . . . . . . . . . . . . . . . . . . 29

5 Plotting data with Matplotlib 31


5.1 Plotting data with Matplotlib . . . . . . . . . . . . . . . . . . . . . . . . . . 31
5.2 Simple Curves . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31
5.3 Contour plots . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 32
5.4 Exercise: Plotting a function . . . . . . . . . . . . . . . . . . . . . . . . . . 33
5.5 Extra exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 34
Fourier series . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 34
Taylor expansion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 34

6 Curve Fitting 36
6.1 Linear regression . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 36
6.2 Polynomial regression . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 40
6.3 Fitting to any function . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 41
6.4 Some remarks on fitting . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42
6.5 Exercise: fitting to a Gaussian . . . . . . . . . . . . . . . . . . . . . . . . . . 43

7 Reading and analyzing data 44


7.1 Reading data from a file . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 44
7.2 Statistics and histograms . . . . . . . . . . . . . . . . . . . . . . . . . . . . 47
7.3 Exercise: fitting many curves . . . . . . . . . . . . . . . . . . . . . . . . . . 49
7.4 Extra exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 50
Process Data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 50

Vectors, Matrices and Systems of Linear Equations 52


8 Vectors and Matrices 53
8.1 Simple vectors and matrices . . . . . . . . . . . . . . . . . . . . . . . . . . 53
8.2 Slicing vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 54
8.3 Slicing Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 54
8.4 Complicated Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 56
8.5 Linear Algebra . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 57
8.6 Matrix-Vector multiplication . . . . . . . . . . . . . . . . . . . . . . . . . . 58
8.7 Determinant: Recursion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 59
8.8 Exercise: Matrix Multiplication . . . . . . . . . . . . . . . . . . . . . . . . . 61

9 Systems of linear equations 63


9.1 Example: Distillation Column . . . . . . . . . . . . . . . . . . . . . . . . . 63
9.2 Solution using numpy . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 64
9.3 The Gauss-Jordan algorithm . . . . . . . . . . . . . . . . . . . . . . . . . . 65
9.4 Gauss-Jordan implemented . . . . . . . . . . . . . . . . . . . . . . . . . . . 67
Computational cost . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 69
9.5 The Jacobi iterative method . . . . . . . . . . . . . . . . . . . . . . . . . . . 71
9.6 Exercise: Two-stage distillation column . . . . . . . . . . . . . . . . . . . . 72
9.7 Exercise: Jacobi Iteration . . . . . . . . . . . . . . . . . . . . . . . . . . . . 73
9.8 Extra exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 74
Lake contamination . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 74
The Gauss-Seidel method . . . . . . . . . . . . . . . . . . . . . . . . . . . . 74

Integration, Derivatives and Non-linear systems 76


10 Numerical integration 77
10.1 Equal intervals . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 77
Rectangle approximation . . . . . . . . . . . . . . . . . . . . . . . . . . . . 77
Trapezoid approximation . . . . . . . . . . . . . . . . . . . . . . . . . . . . 79
Simpson’s rule . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 80
Newton-Cotes formulas . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 81
10.2 Gaussian quadrature . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 81
10.3 Multiple Integrals . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 82
10.4 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 83
Composite methods . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 83
Simpson’s rule . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 85

11 Numerical differentiation 86
11.1 First derivative . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 86
Example . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 88
11.2 Second derivative . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 90
11.3 Partial derivatives in 2D . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 91
11.4 Summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 92
11.5 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 93
Second derivative . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 93

12 Non-linear equations 94
12.1 The bisection method . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 94
Recursion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 95
Recursive bisection method . . . . . . . . . . . . . . . . . . . . . . . . . . . 96
12.2 The Newton method . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 97
Two coupled equations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 99
Example . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 101
12.3 Reduced Newton Method . . . . . . . . . . . . . . . . . . . . . . . . . . . . 103
Multiple equations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 103
12.4 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 104
The bisection method . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 104
The Newton method . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 104
Newton for two equations . . . . . . . . . . . . . . . . . . . . . . . . . . . 106
12.5 Extra exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 106
Gradient descent in 2D . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 106

Ordinary differential equations 108


13 Classification of ODEs and linear differential equations 109
13.1 Classification of ODEs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 110
13.2 Linear ODEs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 112
13.3 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 116
Coupled first order reactions . . . . . . . . . . . . . . . . . . . . . . . . . . 116

14 Initial value problems 118


14.1 Forward Euler . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 119
14.2 Modified Euler . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 120
14.3 Runge-Kutta methods . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 120
14.4 Numerical VS Exact . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 122
14.5 ODE-solver in 𝑠𝑐𝑖𝑝 𝑦 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 123
Simultaneous ODEs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 124
The predator-prey model . . . . . . . . . . . . . . . . . . . . . . . . . . . . 124
14.6 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 126
Forward Euler . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 126
Runge-Kutta 4-th order . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 127
Coupled first order reactions . . . . . . . . . . . . . . . . . . . . . . . . . . 127
Epidemic propagation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 128

15 Boundary-value problems 130


15.1 The shooting method . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 130
Application . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 132
𝑁 equations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 133
15.2 The finite difference method . . . . . . . . . . . . . . . . . . . . . . . . . . 134
15.3 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 137
The shooting method . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 137
1D transport . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 138

Special topics 140


16 Monte Carlo schemes 141
16.1 Random numbers . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 141
16.2 Monte-Carlo Integration . . . . . . . . . . . . . . . . . . . . . . . . . . . . 143
Standard MC integration . . . . . . . . . . . . . . . . . . . . . . . . . . . . 144
Throwing darts . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 145
16.3 Random walks: Direct MC . . . . . . . . . . . . . . . . . . . . . . . . . . . 146
Random Walk in 1D . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 146
Diffusion: random walk . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 147
Example: stock prices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 148
Random walk in 2D . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 148
16.4 Metropolis Monte-Carlo . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 149
Lattice Metropolis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 150
Off-lattice Metropolis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 150
16.5 Exercises : Algorithm . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 152
Approximating 𝜋 by MC . . . . . . . . . . . . . . . . . . . . . . . . . . . . 152
Stock Price . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 152
Mixing of particles . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 152
Appendix 154
A Exponential of a matrix 155
Demonstration of eq. 13.32 . . . . . . . . . . . . . . . . . . . . . . . . . . . 155
Exponential of a matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 155

Index 157
Basic Introduction to Programming in
Python
Introduction to Python 1
Computers are exceptionally good at performing repetitive task at a 1.1 Getting Started with Python . 2
very high speed, and with virtually zero probability of mistakes. The Installing Python . . . . . . . . 3
word computer already implies that the main task of computers is to Documentation . . . . . . . . . 3
1.2 Basic calculations in Python . 3
do computations or calculations, which is the kind of task that occurs
Variables . . . . . . . . . . . . . 4
very often in engineering disciplines but also in basic sciences, including
Comments . . . . . . . . . . . . 5
chemistry and physics. Using computers, a wide variety of mathematical
Printing and formatting text . 5
problems can be solved, generally not by deriving analytical solution 1.3 Some remarks on syntax . . . 7
to a certain mathematical problem but by coming up with a numerical Equality vs. assignments . . . 7
solution. While analytical solutions are sometimes preferred, in a majority Precedence of operators . . . . 8
of cases it is not even possible to get to such a solution without making 1.4 Importing modules . . . . . . . 8
very severe approximations. In order to use a computer to help us solve 1.5 Exercise: Height of ball . . . . 9
out problems in science and engineering, we need a way of instructing
the computer what to do; i.e. we need a programming language. In
this course we use a language called Python, which is an easily usable
high-level programming language that is quickly becoming a standard
in scientific computing. One of the big advantages is that it is open
source (free) and can be used on virtually all platforms (Windows, Mac,
Linux). The language is easy to learn and programs written in Python are
generally easy to read and often much shorter than comparable programs
written in languages such as C or Fortran. A major difference between
Python and these languages is that C or Fortran codes have to be compiled
into an executable format before they can be run, while in Python the
code is interpreted while it is running. The disadvantage of this is that
computationally intensive programs run a lot slower than pre-compiled
ones written in C or Fortran. However, Python comes packed with an
enormous range of standard modules that contain many precompiled
functions and algorithms for many common (numerical) tasks. Examples
include opening and reading files, complicated mathematical functions,
plotting data in graphs and advanced numerical methods. The version
of Python that is currently used, is version 3.7 and that is what we
will use in this course (however most of the examples will also run
in Python 2.7). In this chapter we will give a general overview of the
programming language Python, covering most of the basic features that
all programming languages share and that will be used in the chapters
that follow.

1.1 Getting Started with Python

In order to use Python, we need to at least install an interpreter that


allows us to run Python codes. In addition, there are many useful tools,
for instance graphical user interfaces (GUIs) that make programming
in Python more convenient. A very extensive Python programming
environment is Anaconda, which contains many modules and libraries
by default, for example the ’Spyder’ editor that is specifically designed
for editing Python code.
1 Introduction to Python 3

Installing Python

There are many distributions of Python available, that sometimes come


with many different modules and help-programs. The standard distri-
bution that is used at TU-Delft is the Anaconda package, which can be
downloaded freely at:
[Link]
The anaconda package is very complete. Once start the ’Anaconda
Navigator’ you will see some of these programs. The part we will use
mostly is the Spyder development environment that will allow you to
write and execute code.

Documentation

There is extensive online documentation available for Python and there


are many online resources, including tutorials and examples for specific
applications. Python documentation can be easily accessed through
the Help function in the Spyder editor. The entire manual and other
documentation are also available on:
[Link]
This course does not even come close to covering all the possibilities
offered by Python. Therefore many concepts and tools will not be men-
tioned during the lectures. However several websites have great tutorials
and examples to learn how to code in Python. See for example:
[Link]

1.2 Basic calculations in Python

The first way in which Python can be used is as a basic calculator, where
you can just print the result of a simple equation, for instance to calculate
the height of a ball from the surface of the earth with an initial speed
𝑣0 :

1 2
𝑦(𝑡) = 𝑣0 𝑡 − 𝑔𝑡 (1.1)
2

Here 𝑦(𝑡) is the height of the ball, 𝑔 the gravitational constant, and 𝑡 the
time elapsed after the ball has been launched. We can simply compute
the height of the ball at 𝑡 = 0.6 in a Python program as:
1 print(5*0.6 - 0.5*9.81*0.6**2)

where we assumed an initial velocity of 𝑣 0 = 5 m s−1 . This line of code


already illustrates the use of the basic numerical operations adding (+),
subtracting (-), multiplication (*) and division (/). The square is indicated
by **2. This line of text is a complete Python program that can be saved
in a file, for instance [Link]. In the Spyder environment that we will use
in this course, we can run this program by selecting ’Run’ from the menu
(or by pressing the ’play’ button. The output of this calculation is 1.2342,
which will appear if the program is executed.
1 Introduction to Python 4

The line of code in the grey area also illustrates how pieces of example
code in this text can be recognised. The words in red are the so-called
reserved words in Python. They indicate a specific action that you want
to happen in the code, in this case printing the answer to the calculation
on the screen. These reserved words can not be used for anything else in
your code, for instance as the name of a variable.

Variables

While the piece of code above does something useful, it is not easy to
recognise the different parameters, let alone vary them easily. For instance
selecting a different time, t, or a different initial velocity requires us to
know the equation and what each number stands for. It is much easier
to change the different parameters if the equation would be written in
terms of the variables. In programming languages, and Python is no
exception, it is possible to define 𝑣 0 , 𝑔 , 𝑡 and 𝑦 as variables and combine
them into to the right hand side of the equation for our ball. The result
can be written as:
1 v0 = 5
2 g = 9.81
3 t = 0.6
4 y = v0*t - 0.5*g*t**2
5 print(y)

This illustrates how variables in Python are defined: by setting a name


equal to a numerical value or an expression containing variables that Listing 1.1: Python example
have been previously defined. The variable is automatically of the type
that is defined by the numerical value that is assigned to it; 𝑣 0 is an
integer, while 𝑔 and 𝑡 are floating point numbers. This assignment of
variables also marks a difference with more conventional programming
languages such as C and Fortran where variables have to be declared
(with the type of the variable indicated) before they are first used. The
reason for this is that the code in these languages is compiled before
running it, and all memory allocations are sorted out already at that time.
In Python, there is no pre-compiled code and all these issues are sorted
out at run time. When the code is written using variables, it is much
easier to read as in the example above since the original equation can be
recognised and the different variables have a specific meaning. In this
way it is also much easier to modify the numerical value of the variables
since they can be found more easily, but the result of the calculation is
exactly the same: 1.2342.
The names of the variables can be freely chosen and it is good practice
to use names that are easy to understand; i.e. some descriptive name.
This makes it easy to go back after you have written an code and actually
understand what you wrote before. The naming of the variables is fully
free in principle, as can be seen in the example below. The result of the
code is exactly the same as before.
1 initial_velocity = 5
2 acceleration_of_gravity = 9.81
3 TIME = 0.6
4 VerticalPositionOfBall = initial_velocity*TIME - \
5 0.5*acceleration_of_gravity*TIME**2
6 print(VerticalPositionOfBall)
1 Introduction to Python 5

This version illustrates two things. First of all, when an expression


becomes too long it may become convenient or necessary to extend it
over multiple lines. This can be done by putting a backslash at the very
end of the line (make sure that there are no blank spaces behind the
backslash). Secondly, the example shows that it may be a good idea to
keep the names of variables reasonably compact because most people
would agree that the readability of this last version is worse than the
previous one. The result of the code is still exactly the same.

Comments

So far we have only the discussed the actual program statements that
actually ’do’ something. For a readable code it is advisable to add
comments that do not actually do something but they describe the code.
This makes it a lot easier to follow what is going on in the code and will
make it easier to understand and modify a code at a later stage. In Python,
comments start with the hashtag character. All text after this character is
a comment and will not be considered when the code is executed. Our
code with comments looks like this:
1 # Program to calculate the height of a ball that moves in the vertical
direction
2 v0 = 5 # Initial velocity
3 g = 9.81 # acceleration of gravity
4 t = 0.6 # time
5 y = v0*t - 0.5*g*t**2 # calculation of the position of the ball
6 print(y)

This program again does exactly the same thing as before, but it is much
easier to understand for other people, they may actually figure out what
is going on in the code when they read it without explanation by the
programmer. Any code that consist of more than a few lines will greatly
benefit from the inclusion of comments and a smart choice of the names
of variables.

Printing and formatting text

In the programs we have looked at so far the output was just a single
number, which basically does the job we would like but it is in many cases
nice to have a somewhat more informative line of output. Ideally, we
would also like to have some control over the formatting of the number
in the output, for instance the number of decimal spaces. An example of
such output is given below:
At t=0.6 s, the height of the ball is 1.23 m.
Such kind of output can be achieved using the print statement that we
have seen before but the way it is used is a bit more complicated. The
formatting that is used here is the so-called ’printf’ formatting that has
its origin in the programming language C. A line of code that supplies
such output is the following:
1 print(’At t=%g s, the height of the ball is %.2f m.’ % (t, y))
1 Introduction to Python 6

This line is a bit complicated although the output as specified can be


recognized. The print statement prints everything that is enclosed in
quotes. The quotes can be single quotes as in ’text’ or double quotes as
in "text". The text between the quotes is called a ’string’. In the example
above, the string contains two slots where the numerical values of two
variables can be inserted. After the string that is contained in quotes, there
is a percentage sign that is followed by the variables that are to be inserted
in parentheses, separated by a comma. The slots in the string itself start
with a percentage sign: %g and %.2f in this case. The information that
follows behind the percentage sign defines the way in which the number
is formatted. Here ’g’ means that the real number should be written as
compactly as possible, either in a decimal or scientific notation. This is
useful for printing numbers on the screen, but it is nice to have some
control over how numbers are specified exactly. The second slot, %.2f
offers some more more flexibility, in this case a ’floating point number’ is
written in a decimal notation with two decimals.
Other options that are sometimes useful are scientific notations with
%e or %E in which a number is written in scientific format with either
’e’ or ’E’. In this notation is is also possible to exactly specify how the
number is to be written. For instance %14.6E means that a floating point
number is written in scientific notation in a field of 14 characters with 6
decimal places, with capital ’E’. An integer number can be inserted as
%d. A listing of all the possibilities of such formatting would fill up an
entire chapter. You can find more information on line as for example at:
[Link]
Our program, including the print statement now reads:
1 v0 = 5
2 g = 9.81
3 t = 0.6
4 y = v0*t - 0.5*g*t**2
5 print(’At t=%g s, the height of the ball is %.2f m.’ % (t, y))

Using this program, it is easy to experiment with some of the different


possibilities for printing the output of the numbers that generated. Python
also uses a more recent syntax to print the values of variables. Keeping
the same example as above we can also write:
1 " At t = {0:1.3f} sec., the height of the ball is {1:8f}".format(t, y)

This new syntax uses the method format() of the strings. Each slot where
we want to insert the value of a variable is now given by a {what:how}
where ’what’ is what we want to print and ’how’ specifies the format.
In the example above the whats are 0 and 1 and refers to the first and
second argument of the format() method. Alternatively we could have
written:
1 " At t = {a:1.3f} sec., the height of the ball is {b:8f}".format(a=t, b=y)

where names ( 𝑎 and 𝑏 ) were given to the variable to be printed. This


new syntax is more flexible and more powerful than the the old way
presented above. For example it is very simple to left- and right-justify
texts using the operator "<" and ">" as in
1 " {0:<20s} {price:6.2f} Euros".format(’Spam and Eggs’,price = 6.99)
2 " {0:>20s} {price:6.2f} Euros".format(’Spam and Eggs’,price = 6.99)
1 Introduction to Python 7

1.3 Some remarks on syntax

As we have seen above, computer programs consist of a collection of


statements that result in very specific actions. This can be calculations
or printing text or numbers on the screen. For the correct interpretation
of the statements it is important to follow the rule of a programming
language, or the ’syntax’ when writing lines of code. In Python, each line
in principle contains a single instruction, and we have already seen that a
line of code can be extended over multiple lines. It is in principle allowed
to write multiple instructions on a single line, separated by a semicolon.
We could have written our program for calculating the height of the ball
as:
1 v0=5;g=9.81;t=0.6;y=v0*t-0.5*g*t**2;print(y)

Clearly, it is not the best idea to collect all statements on a single line;
this line is almost impossible to read! In this line of code, the spaces, for
instance around ’=’ are omitted, which is also true for other mathematical
operations. A general convention is to use one blank space around
=, - and + and no spaces around *, / and **. It should be noted that
this is just a matter of readability, omitting the spaces is in principle
correct. For guidelines about syntax conventions you can refer to the
PEP8 documents at [Link]
This series of conventions are widely adopted in the Python community
and significantly improve the readability of codes.

Equality vs. assignments

When we first wrote the line of code that contained the equation for
calculating the height of a ball at time 𝑡 the result looked very much
like a mathematical equation. In mathematics the ’=’ sign means that
the expressions on both sides of this sign are in fact equal. In Python
(and most other programming languages) the ’=’ indicates what we call
an assignment. An assignment basically means the following: calculate
everything on the right-hand side of the ’=’ and store the result in the
variable on the left-hand side. Although this may seem very similar to
the mathematical meaning it is in fact very different. Consider the line
below:
1 y = y + 3

Mathematically this is clearly wrong; the value of a variable can not be


the same as that same variable plus 3. In a Python program the right
side is evaluated, meaning that the current value stored in 𝑦 is taken, 3
is added to it and the result is assigned to the variable 𝑦 . The old value
that was stored in 𝑦 is now lost since it was overwritten by the newly
calculated value. A short code that illustrates this is shown below:
1 y = 3
2 print(y)
3 y = y + 4
4 print(y)
5 y = y*y
6 print(y)

This program will result in the output of three numbers: 3, 7 and 49.
1 Introduction to Python 8

Precedence of operators

Formulas in Python are evaluated in the same way as they would normally
be evaluated in mathematics. This means that power operations, 𝑎 4 coded
as a**4 have precedence over multiplications and division, which in turn
have precedence over addition and subtraction. We can use parentheses
in our code to change the way an expression is evaluated.

1.4 Importing modules

The basic types of operations that are present in Python are very limited,
but it is already possible to perform many interesting calculations. As
we have seen, the basic arithmetic operations are present, but if we
need to calculate a square root for instance this is not included in the
standard operations. The good thing about Python is that there is a wide
range of modules that can be included that contain useful functions. For
instance the module ’math’ contains a square root function (sqrt) and
many other mathematical functions, including sin, cos, exp and log. In
order to use these functions it is necessary to ’import’ these functions or
a whole module into you Python program, which is done by the import
statement:
1 import math
2 v0 = 9.0
3 answer = [Link](v0)
4 print(answer)

As can be seen in this code, if the module is imported in this way, the
function name requires a prefix which is the name of the module where
it comes from. This prefix can be a bit annoying, especially if the function
is used very often. An alternative import syntax makes it possible to skip
the prefix, and only specific functions can be imported:
1 from math import sqrt
2 v0 = 9.0
3 answer = sqrt(v0)
4 print(answer)

After this, the function ’sqrt’ can be used without the prefix. It is also
possible to import more than one function at once:
1 from math import sqrt, exp, sin

or simply importing all functions from a certain module:


1 from math import *

In the last statement all functions in the math module are imported and
can be used in your code without the prefix. In general it is advised
to import only the functions are actually used in a program because
importing all results in a lot of names in the program that are not used,
but can also not be used for variables anymore. However it is sometimes
quite convenient, especially in the case of the math module, to just import
all functions.
Sometimes it can be convenient to import modules and functions and
give them new names in one go:
1 Introduction to Python 9

1 import math as m
2 #now the math module is called m in the program
3 v = [Link]([Link]) #pi is a constant in the math module that is now called
’m’

Again, all the red words are reserved words in Python. You can also
change the name of a given function that you are importing from a
module. This is done in a very similar way than for the changing the
name of the module itself
1 from math import sin as theSinus, cos as theCosinus, log as theLog
2 x = 5.0
3 v = theLog(x)
4 v = theSinus(x)*theCosinus(x) + theLog(x)

However this is generally not a very good idea as you can loose track of
where a given function is defined. Besides writing [Link](x) is generally
much faster much more readable as theSinus(x).
The math module is only one of the many modules that can be used in
Python, the virtually endless list contains modules related to designing
web interfaces, doing operations on all sorts of graphical formats (pic-
tures), interacting with the operating system of a computer, etc. In this
course we are dealing mostly with the use of Python in a scientific/engi-
neering environment and we focus on numerical mathematics techniques.
For the purpose of this course there are two additional modules that
we will encounter in the chapters that follow. One is a collection of a
wide variety of pre-programmed numerical methods to solve all sorts of
problems, including integration, matrix operations, solving differential
equations, etc. The library that contains these functions is called numpy.
The second module that is useful is called matplotlib, which contains
functions to make a graphical representation of data on screen or in a
picture. These modules will be introduced in later chapters but they are
imported in the same way as the math module here.

1.5 Exercise: Height of ball

For the first exercise you will reproduce the example presented in
section 1.2. You will therefore write a small program named ’[Link]’
that automatically computes the height of a ball that falls vertically
following:

1 2
𝑦 = 𝑣0 𝑡 − 𝑔𝑡 (1.2)
2

and print the answer. You’ll define the variable as shown in page 8. We
now ask the question : How long does it take the ball to reach a certain
height 𝑦 𝑐 ? Posing 𝑦 = 𝑦 𝑐 in the equation above and rearranging the
terms we find

1 2
𝑔𝑡 − 𝑣0 𝑡 + 𝑦 𝑐 = 0 (1.3)
2
1 Introduction to Python 10

You can solve this equation using the well known formula for a quadratic
equation and find the expression of the two roots

 q 
𝑡± = 𝑣 0 ± 𝑣02 − 2 𝑔 𝑦 𝑐 /𝑔 (1.4)

There are two times because the ball can reach 𝑦 𝑐 on its way up or down.
Update your code so that it also computes the values of 𝑡 𝑝𝑚 for 𝑦 𝑐 = 0.2.
What happens if you take 𝑦 𝑐 = 2.0 ? Change your code to intercept the
error message.
Note: the square root function can be loaded from the 𝑚𝑎𝑡 ℎ module.
1 import math as m
2 [Link](4.)
Loops and lists 2
As mentioned in Chapter 1, computers are exceptionally well-suited 2.1 While-loops . . . . . . . . . . . 11
to perform (boring) repetitive tasks in a automatic way. In computer 2.2 Boolean expressions . . . . . 13
programs such tasks can conveniently be performed in loops of which 2.3 Lists . . . . . . . . . . . . . . . 13
2.4 For-loops . . . . . . . . . . . . 14
different types exist. Additionally, what is often useful for large collections
2.5 The range construction . . . 15
of data (a row of numbers) to store them in a smart way, for instance
2.6 Branching: if-else statements 16
in a vectors or an array. In Python, such storage can be done in a ’list’.
2.7 Exercise: Money in the bank . 17
Loops and lists, together with functions and if-statements form the basic 2.8 Exercise: Cooking a perfect egg18
core of programming operations that are used in this course and a
thorough understanding of these concepts is essential for understanding
the following chapters.

2.1 While-loops

In the previous parts of this chapter we have evaluated a mathematical


formula by inserting a single value and printing the answer, for instance
the conversion of temperatures from Celcius to Fahrenheit. Our task
now is to print a table with a conversion list with two columns, the
temperatures in both scales. For instance a table running from -20 to +40
degrees C in steps of 5 degrees. A very naive solution to this problem is
shown in the following code:
1 C = -20; F = 9.0/5*C + 32; print(C, F)
2 C = -15; F = 9.0/5*C + 32; print(C, F)
3 C = -10; F = 9.0/5*C + 32; print(C, F)
4 C = -5; F = 9.0/5*C + 32; print(C, F)
5 C = 0; F = 9.0/5*C + 32; print(C, F)
6 C = 5; F = 9.0/5*C + 32; print(C, F)
7 C = 15; F = 9.0/5*C + 32; print(C, F)
8 C = 20; F = 9.0/5*C + 32; print(C, F)
9 C = 25; F = 9.0/5*C + 32; print(C, F)
10 C = 30; F = 9.0/5*C + 32; print(C, F)
11 C = 35; F = 9.0/5*C + 32; print(C, F)
12 C = 40; F = 9.0/5*C + 32; print(C, F)

This does the job perfectly, twelve almost identical lines of code (with
three statements on each line) that give the required result. Now imagine
wanting a table in steps 0.01 degrees? It is clear who is doing all the hard
work here, in terms of typing the lines of code. While it works perfectly
well, it isn’t using the capabilities of computers to the fullest. The main
issue is: all lines do the same thing, only the value in the variable C
changes. All programming languages are equipped with a construct to
deal with such problems in a very easy way, so-called loops. In Python
there are two kinds of loops, while-loops and for-loops.
A while loop is used to repeat a set of statements as long as a certain
condition is true. We will use this type of loop here as an example to
generate the list temperatures for our conversion table. A sequence of
operations, or an algorithm for this is listed below in text form.
2 Loops and lists 12

1. Print a line with dashes


2. C = -20
3. While C <= 40:
I F = 9C/5 + 32
I Print C and F
I Increment C by 5
4. Print line with dashes

This list of operations that we want to perform is a translation of our


task (making a conversion table) into what we call an algorithm. We have
already written it in such a way that we can directly translate it into a
piece of code in Python, in this case using a while loop. This illustrates
the main task of a programmer: translating a particular problem or task
(in our case a scientific or engineering problem) into a set of instructions
that tell a computer how to solve the problem or execute the task. The
implementation of this algorithm in Python code results in a very short
program:
1 print(’-’*20) # print table heading
2 C = -20 # initialize C
3 dC = 5 # set increment
4 while C <= 40: # loop heading with condition
5 F = (9.0/5)*C + 32 # first statement in loop
6 print(C, F) # second statement
7 C = C + dC # third statement
8 print(’-’*20) # bottom of table (not in loop)

This small code contains one of the most important features of program-
ming in Python: the block of statements that is to be executed in the while
loop has to be indented by the same amount. All operations under point
three in the algorithm above start in the code on the same position of the
line. In this way, the Python interpreter knows that these three statement
should be ’inside the loop’. At the top of the loop, line 4, Python checks
whether the condition is true (C <= 40), which is certainly true at the
beginning since the initial value of C is -20. Subsequently it executes
the three lines inside the loop (including printing the table entry) and
goes back to the top of the loop, line 4. Inside the loop the value of C has
changed so the check has to be performed again. This cycle of events is
repeated until, at some point, the condition is not true anymore: C has
the value of 45 after twelve cycles. At this point, the program does not go
into the loop anymore but skips to the first statement after the loop, line
8. The colon : behind the while statement on line 4 is essential because it
marks the start of the block of code on the next lines to be executed. It is
interesting to experiment with this code, for instance giving the last line
the same indentation as those inside the loop. In this case the lines in the
table will be separated by a line of dashes.
It is easy to see from this code that this way of solving this problem is
much more convenient and powerful than typing a separate line of code
for each temperature. On top of that, it is very easy now to make a much
longer list, incremented every 1 degree for instance by just changing the
values of dC. Note the use in the piece of code above of the printing
syntax:
1 print(’-’*20)

that prints exactly 20 dashes.


2 Loops and lists 13

2.2 Boolean expressions

An important aspect of while loops is that at the top of the loop there
is an expression of which it is evaluated whether it is true. This type
of expression is called a boolean expression that can either be true or
false. Other comparisons that can be useful are listed in the following
fragment:
1 C == 40 # C equals 40
2 C != 40 # C does not equal 40
3 C >= 40 # C is larger than or equal to 40
4 C > 40 # C is larger than 40
5 C < 40 # C is smaller than 40

The result of a boolean expression can be inverted (changed for True to


False or the other way around) by putting the keyword ’not’ in front of
it:
1 not C == 40 # if C equals 40 the result will be ’False’

Boolean expressions can also be combined with the keywords ’and’ and
’or’:
1 while x > 0 and y <= 1:
2 print(x, y)

If both separate conditions are true the overall result is ’True’, and in all
other cases the result will be ’False’.

2.3 Lists

Until now all the calculations we have made using variables contained
a single number in an isolated variable. In many cases, especially in
mathematics, it is more natural if these number are grouped together
in a list, or an array as it is called in many text about programming. An
illustration is the list of temperatures in the examples above. In Python a
list can be made just assigning a variable name to a list of numbers that
is contained in square brackets, separated by commas:
1 C = [-20, -15, -10, -5, 0, 5, 10, 15, 20, 25, 30, 35, 40]

The variable C now refers to a list object that contains 13 elements and
all these elements are integer numbers. Each element of the list can be
assessed separately using an index that refers to its position in the list.
The index runs from 0 to 12 in this case. To indicate the third element in
the list we can write C[2], which refers to an object of the type int, with
the value -10.
Some basic operations can be performed on lists for instance appending
elements at the end:
1 C = [ -10, -5, 0, 5, 10, 15, 20, 25, 30]
2 [Link](35) #add a new element at the end

Two lists can be added:


1 C = [ -10, -5, 0, 5, 10, 15, 20, 25, 30, 35]
2 C = C + [40, 45] #extend C at the end

or elements can be inserted at specific places in the list:


2 Loops and lists 14

1 C = [ -10, -5, 0, 5, 10, 15, 20, 25, 30, 35, 40, 45]
2 [Link](0, 15) # insert 15 a new element on index 0

Elements can be deleted:


1 del C[2]

and the length of a list can be examined by:


1 len(C)

There are many more operations that can be performed on lists that may
be handy in certain situations but the ones we will use in this course are
quite limited. The combination of a while (or for) loop and lists can be
very powerful since the elements can be addressed one by one in a loop
in a very short piece of code. For instance in the code below, a list is filled
with temperatures starting from -50 up to 200 degrees in steps of 2.5
degrees:
1 C =[]
2 C_value = -50
3 C_max = 200
4 while C_value <= C_max:
5 [Link](C_value)
6 C_value += 2.5

In the last line we have used the operator +=, which has the equivalent
effect of
1 C_value = C_value + 2.5

but is a lot shorter.

2.4 For-loops

For collections of data that are assembled in a list, we often want to


perform the same operation, walking through the elements of that list.
In Python (and other computer languages) there is a simple way to walk
through such a list in a for loop. For instance a loop can be used to print
all the elements in a list:
1 degrees = [0, 10, 20, 40, 100]
2 for C in degrees:
3 print(’list element: ’, C)
4 print(’The list has’, len(degrees), ’elements.’)

This for loop runs over all elements in the list and in every pass the
variable C takes the value of one of the elements of the list. Again, the for
specification ends with a colon : and the lines that follow below with the
same indentation are executed in each cycle of the loop. Using the for
loop we can now take a list of temperature in Celcius and convert them
to Fahrenheit:
1 Cdegrees = [-20, -15, -10, -5, 0, 5, 10, 15, 20, 25, 30, 35, 40]
2 for C in Cdegrees:
3 F = (9.0/5)*C + 32
4 print(C, F)

Executing this piece of code gives a rather ugly looking table, but we
have seen before that the way numbers are printed can be influenced in
the print statement very easily:
2 Loops and lists 15

1 Cdegrees = [-20, -15, -10, -5, 0, 5, 10, 15, 20, 25, 30, 35, 40]
2 for C in Cdegrees:
3 F = (9.0/5)*C + 32
4 print(’%5d %5.1f’ % (C, F))

Now the temperature are both printed in a fixed field-width of 5 characters.


The temperature in Celcius is an integer %5d, while the temperature in
Fahrenheit is a floating point number that is printed with one decimal,
%5.1f. We can use the same method to store the values of the temperature
in Fahrenheit using the append() function
1 Cdegrees = [-20, -15, -10, -5, 0, 5, 10, 15, 20, 25, 30, 35, 40]
2 Fdegrees = []
3 for C in Cdegrees:
4 F = (9.0/5)*C + 32
5 [Link](F)
6 print(’%5d %5.1f’ % (C, F))

Finally let’s mention that if the sole purpose of the for loop is to create
the Fdegrees list, this can be achieved in Python using only one line !
1 Cdegrees = [-20, -15, -10, -5, 0, 5, 10, 15, 20, 25, 30, 35, 40]
2 Fdegrees = [ (9.0/5)*C + 32 for C in Cdegrees ]

This is very Pythonic ! In the second line a new list is create where each
element depends is built up from the list Cdegrees.
We have now encountered two different kinds of loops, for-loops and
while-loops. These loops differ in the way they are implemented but in
principle, either of the two can be used to produce any loop. In fact, in
many programming languages only a single type of loop construct is
available, and this works perfectly fine. The for-loop just above can be
programmed using a while-loop with exactly the same result:
1 Cdegrees = [-20, -15, -10, -5, 0, 5, 10, 15, 20, 25, 30, 35, 40]
2 index = 0
3 while index < len(Cdegrees):
4 C = Cdegrees[index]
5 F = (9.0/5)*C + 32
6 print(’%5d %5.1f’ % (C, F))
7 index += 1

This code is slightly longer any some would argue that it is less clear than
the implementation with the for loop, but both give exactly the same
result. Therefore, it is up to the programmer to decide which construct
to use and it is advised to take the version that is the most intuitive for
the specific application.

2.5 The range construction

A useful construction in combination with for-loops is the ’range’ con-


struction. Instead of typing the entries in the list of temperatures in
Celcius we can generate this list using a for-loop. The range construction
can work in several ways:

I range(n) generates integers 0, 1, 2, ..., n-1


I range(start, stop, step) generates a list of integers start, start+step,
start+2*step etc., up to, but not including stop. For example range(2,
8, 3) generates 2 and 5, but not 8.
I range (start, stop) is the same as range(start, stop, 1)
2 Loops and lists 16

Using the range construct a for loop over a range of integers can be
written as:
1 for i in range(start, stop, step):
2 ...

We can use this construction to generate the list of temperatures as


before. However, the range construction only generates a sequence of
integers, which means that we can not use it directly to make a list of
temperature that increase in steps of 2.5 degrees. We can still use the
range construction but it requires an extra operation:
1 Cdegrees = []
2 for i in range(0, 21):
3 C = -10 + i*2.5
4 [Link](C)

Finally, it is perfectly fine to include a for- or while-loop (or more than


one) inside the list of statements that is executed inside another loop.
This leads to so-called nested loops that are encountered very often in
numerical methods:
1 for x in range(1, 11):
2 for y in range(1,11):
3 z = x**2 + y**2
4 print(x, y, z)

This code will will print the value of z, for every combination of x and y,
where x and y range from 1 to 10.

2.6 Branching: if-else statements

In a computer program there are often decisions to be made, on basis


of the results of calculations or the numbers that are read from a file
etc... Decisions can force the program to go into different ’branches’. If a
certain condition is met, we do one thing, if not, we do another thing. In
Python this can be done using an if-else statement which is very intuitive:
it basically does what the code says:
1 if C < -273.15:
2 print(’%g degrees Celcius is non-physical!’ % C)
3 print(’ The Fahreheit temperature will not be computed.’)
4 else:
5 F = 9.0/5*C + 32
6 print(F)
7 print(’end of program’)

The two print statement on lines 2 and 3 are only executed if C < 273.15 is
True, otherwise it is skipped and the code on lines 5 and 6 is executed. The
final printing statement on line 7 is always executed since the indentation
tells Python that it is not inside the block of statements in the if-else
statement. The else can also be skipped :
1 if C < -273.15:
2 print(’%g degrees Celcius is non-physical!’ % C)
3 F = 9.0/5*C + 32
4 print(F)
5 print(’end of program’)

In this case the conversion in Fahrenheit is not in the if-else block and is
therefore always executed, regardless of whether it makes sense or not.
2 Loops and lists 17

These statements have only given the choice between two options but it
is in fact possible using the elif keyword (short for else if) to ’branch’ the
program into as many different flows are needed:
1 if condition1:
2 <block of statements>
3 elif condition2:
4 <block of statements>
5 elif condition3:
6 <block of statements>
7 else:
8 <block of statements>
9 <next statement>

There are many other ways of using the if statement and some of these
will appear in the code that is used in further chapter in this course. In
most cases it is easy to understand how these constructs work because
as with a lot of statements in Python: they actually make some sense
in normal language too. It is also a good exercise to find your own
background information about the if-else statement since the online
documentation is virtually endless, including many examples. This is
true for all other Python constructs discussed in this chapter.

2.7 Exercise: Money in the bank

Let 𝑝 be a bank interest rate in percent per year. An initial amount 𝐴 has
then grow to

𝑝 𝑛
𝐴𝑛 = 𝐴0 (1 + ) (2.1)
100

after 𝑛 years. Make a program called ’[Link]’ that computes and print
on the screen how much money 𝐴0 = 1000 euros have grown to after
𝑛 = 1 , 2, 3, 4..., 𝑁 years with a 5% interest rate. To do that, make a for
loop over the value of 𝑛 and store the successive values of in a list
𝐴𝑛 = [1000 , ......]. Take 𝑁 = 50. If you don’t know how to start, try
something like the following:
1 A0 = ....
2 N = ...
3 p = ...
4 A = []
5 ....
6 for i in range(N):
7 [Link]( .... )
8 print(’year = %d \t money = %1.6f’ % (i, A[i]))

Of course the ... must be replaced by expressions if you want your


program to work. Notice that we declare 𝐴 as a list and then use the
𝑎𝑝𝑝𝑒𝑛𝑑() command to add values into the list. Notice as well the use of
the format %1.6 𝑓 that forces Python to print the money with 6 numbers
after the dot. Note as well the \t character that prints a tab.
Modify your code so that it stops when you became a millionaire !! You
can do that with a while loop. Careful with the while loop that can run
forever if the conditions is always true. It is usually a good idea to add
an additional condition that we know will force the while loop to exit at
some point. See the example below:
2 Loops and lists 18

1 n = 0 # counter
2 N_MAX = 1000 # maximum number of iterations
3 An = 0. # initial value
4

5 # while loop with two conditions


6 while (An < 1E6) and (n < N_MAX):
7

8 An = .... # compute the new value


9 [Link](An) # store it in A
10 n += 1 # don’t forget to increment n !!!
11

12 # print the result


13 print(’It took %d years to become a millionaire !!’ % ( ... ))

How long does it take to become a millionaire ? However what happen


in the code above if you set N_MAX = 100 ? This is because the 𝑝𝑟𝑖𝑛𝑡
statement is executed regardless of the fact that you re a millionaire or not.
Modify the code below by introducing 𝑖 𝑓 ... 𝑒 𝑙𝑠𝑒 branching statement to
print the number of year it took to become a millionaire if you became
one or to print that you never became a millionaire otherwise.
1 if(....):
2 print(’It took %d years to become a millionaire !!’ % ( ... ))
3 else:
4 print(’You never became a millionaire :(’)

2.8 Exercise: Cooking a perfect egg

As an egg cooks, the proteins first denature and then coagulate. When the
temperature exceeds a certain critical point, reactions begin and proceed
faster as the temperature increases. In the egg-white the proteins start to
coagulate at temperatures above 63 °C, while in the yolk the proteins start
to coagulate for temperatures above 70 °C. For a soft boiled egg, the white
needs to have been heated long enough to coagulate at a temperature
above 63 °C, but the yolk should not be heated above 70 °C. For a hard
boiled egg, the center of the yolk should be allowed to reach 70 °C. The
following formula expresses the time 𝑡 it takes (in seconds) for the center
of the yolk to reach the temperature 𝑇𝑦 (in Celsius degrees):

𝑀 2/3 𝑐𝜌1/3 𝑇𝑜 − 𝑇𝑤
 
𝑡= ln 0.76 (2.2)
𝐾𝜋2 (4𝜋/3)2/3 𝑇𝑦 − 𝑇𝑤

Here, 𝑀 , 𝜌, 𝑐 , and 𝐾 are properties of the egg: 𝑀 is the mass, 𝜌 is the


density, 𝑐 is the specific heat capacity, and 𝐾 is thermal conductivity.
Relevant values are 𝜌 = 1.038 g cm−3 , c = 3.7 Jg−1 K−1 , and K = 5.4 10−3 W
cm−1 K−1 . Furthermore, 𝑇𝑤 is the temperature (in °C) of the boiling water,
and 𝑇𝑜 is the starting temperature (in °C) of the egg before being put in
the water. Write a program called ’[Link]’ that calculates the cooking
time for different combination of the temperature and the mass of the
egg.
The eggs can be stored in the fridge (𝑇𝑜 = 4 °C) or at room temperature
(𝑇𝑜 = 20 °C). Add code to your program that compute the time it take to
cook a soft (𝑇𝑦 = 63 °C) and hard (𝑇𝑦 = 70 °C) boiled egg when 𝑇𝑜 = 4 °C
and when 𝑇𝑜 = 20 °C. and for a small egg (50 gram) and a large egg (65
gram). So in total you will have to calculate 8 results.
Structured programming and
functions 3
So far all our programs have been in the form of a single list of statements, 3.1 Simple function . . . . . . . . 19
which is perfectly fine for short programs. When programs get longer, it 3.2 Optional arguments . . . . . 20
becomes more complicated to follow the flow of the program, especially 3.3 Multiple return values . . . . 21
3.4 Exercise: Temperatures . . . 21
when there are loops and nested loops with many statement inside.
3.5 Exercise: A list as argument 22
Moreover, in many cases the same sequence of statements appears many
times in the same program. For these common tasks it is possible to
define our own functions, and they can be called just like the built in
functions or the functions that we import from a library.

3.1 Simple function

As an example consider the code below that calculates and prints the
factorial of the range of integer numbers from 1 to 10:
1 for n in range(1, 11):
2 result = 1
3 for m in range (1, n+1):
4 result = result*m
5 print(n, result)

While this is a very short piece of code it is already hard to imagine


what the code is doing without going through every line and trying to
reproduce the result. It would help greatly if we put some comments
behind every statement but we can also make the code a lot more readable
by introducing some ’structure’ in the code itself. For instance, we can
define a function that calculates the factorial, using the def statement:
1 def myfactorial(n):
2 result = 1
3 for m in range (1, n+1):
4 result = result*m
5 return(result)
6

7 for n in range(1, 11):


8 print(n, myfactorial(n))

In this code we have first defined a function called ’myfactorial’ which


uses a loop that calculates the factorial of the variable n. In the lines
below, the function is called with the argument n and the result is printed.
The last two lines of the code are perfectly clear now, without comments:
we make a loop over integers from 1 to 10 and for each number we print
a line with the number itself and the factorial of that number. The actual
calculation of the factorial is done in a function, in this case a function that
we have defined ourselves called ’myfactorial’. This function takes the
number that we supply to it (in the variable n) calculates the factorial and
returns the result to the main program. The introduction of a very general
function to perform an often occurring calculation has two effect:

I We have a separate function that we can call any time we need a


factorial. This happens to be a function that is also available in the
3 Structured programming and functions 20

module math: [Link], so in this case there was no need to


define ourselves, but we can define any function we want.
I The program now has a more clear structure: it is split up into
easily readable pieces, each performing their own subtask. The task
in a function stand on its own, it does not depend on the rest of
the program, only the result is used in the rest of the program. At
the same time the main program has become much more simple
to read because the complication of the actual calculation of the
factorial has been removed and put into a separate piece of code,
with a convenient name that actually describes what is happening:
the calculation of the factorial!

It should be noted that in large pieces of codes it is always advisable


to split up the code into manageable parts, functions. This keeps the
task for the programmer manageable since small part of code can be
implemented separately, even by different programmers. When writing
the main code of the program above (the last two lines) we just assume
that there is a function that calculates the factorial and use it. We can go
back later and actually implement this small subtask, or even better, get
someone else to do it.
The names of the variables that are used inside a function have nothing
to do with those that we use in the main program, and they also do not
have to be the same. The following is equivalent to the program before:
1 def myfactorial(n):
2 result = 1
3 for m in range (1, n+1):
4 result = result*m
5 return(result)
6

7 for number in range(1, 11):


8 print(number, myfactorial(number))

3.2 Optional arguments

In the previous subsection we have seen how to use one simple argument
as input of a function. Of course a function can take as many arguments as
you want. For example we can define a function that takes two numbers
as argument and return their product. This is simply done by the small
snippet of code below:
1 def myproduct(n, m):
2 return(n*m)

One very useful feature of Python is that some arguments can be defined
as optional. To do that you simply have to give a value to the optional
argument in the definition of the function.
1 def myfunc(w, a=0, b=2*[Link], n=101):
2 # the parameters a,b and n are here optional
3 # default values are used if these options are not specified
4 dx = (b-a)/(n-1)
5 for i in range(n):
6 x = a + i*dx
7 [Link]([Link](w*x))
8 return S
9

10 w = 0.1
3 Structured programming and functions 21

11 x = myfunc(0.1)
12 y = myfunc(0.1, b=4*[Link])

As you can see in the definition of the function the parameters 𝑎 ,𝑏 and 𝑛
are given a default value. Therefore we can omit them during the call of
the function as in line 8. In that case the default values of these parameter
are assumed. If on the other hand we want to use different values than
the default ones we just have to specify these values when calling the
function as on line 9. Note however that to specify a value we need to
explicitly write the name of the argument (here 𝑏 ) in the function call.
Such argument is generally called a keyword argument in Python

3.3 Multiple return values

It is also possible to return more than 1 value from a function. Suppose


we are interested in evaluating the height and velocity of a ball as given
by

1 2
𝑦(𝑡) = 𝑣0 𝑡 − 𝑔𝑡 (3.1)
2
𝑦 0(𝑡) = 𝑣0 − 𝑔𝑡 (3.2)

To return both 𝑦 and 𝑦 0 from a single function we can simply separate


their variables by a comma in the return statement
1 def yfunc(t, v0, g=9.81):
2 y = v0*t-0.5*g*t**2
3 yp = v0 - g*t
4 return y, yp
5

6 pos, vel = yfunc(0.6, 5.0, g=3.711)

Note that you also have to specify two values in the call of the function.
Note as well the use of the optional argument 𝑔 whose default values is
9.81 (like on Earth) but that we change to 3.711 (like on Mars) when we
call the function.

3.4 Exercise: Temperatures

The formula for converting Fahrenheit degrees to Celsius reads:

5
𝑇𝐶 = (𝑇𝐹 − 32) (3.3)
9

On the other hand converting Celsius to Fahrenheit can be done with

9
𝑇𝐹 = 𝑇𝐶 + 32 (3.4)
5
3 Structured programming and functions 22

1. Write a program called ’[Link]’ that contains two functions


𝑐𝑒 𝑙 2 𝑓 𝑎𝑟 and 𝑓 𝑎𝑟 2 𝑐𝑒 𝑙 that implement these formulas for one tem-
perature (each function takes one argument that is a floating point
number and returns the result of the conversion).
Reminder: a Celsius temperature cannot be lower than -273.15 and
a Fahrenheit cannot be lower than -459.67. If some temperatures
are not valid, your function should return ’None’ like this:
1 # Impossible calculation
2 return None

2. To verify the implementation, you can convert a Celsius temperature


to Fahrenheit and then back to Celsius again. That is, you can check
that a temperature 𝑐 equals far2cel(cel2far(c)).
3. Write a program that uses the function ’cel2far’ in a For-loop to
convert a series of temperature ranging from -273 to +227 °C in
steps of twenty degrees (make a list using the ’range’ function) and
print a table with temperature in C and F on the screen.

3.5 Exercise: A list as argument

The functions in the previous exercise take a single value as their argument
and return a single converted temperature. It is also possible to pass a list
of values to the functions, which can then also return a list of converted
temperatures. In this exercise you can rewrite one or both of the functions
written above so that they can convert a whole list of temperatures at
once. To achieve this the function should:
I establish the number of elements in the list that was passed to the
function: use the ’len()’ function.
I use a loop to convert all the elements in the list to other units and
store them all in a (second) list.
I return the list with converted values.
Mathematics and data analysis using
Python
Mathematics using Numpy 4
In the last chapters we have seen how to use computer programs to solve 4.1 Basics of Numpy . . . . . . . 24
simple problems. In this chapter we will turn to more dedicated methods 4.2 Numpy arrays . . . . . . . . . 24
that are specialized for scientific computing. Python is very well suited for 4.3 Arrays indexing and slicing 26
4.4 Vectorization . . . . . . . . . . 27
scientific use because specific modules have been written called 𝑛𝑢𝑚𝑝 𝑦
4.5 The copy issue . . . . . . . . . 29
and 𝑠𝑐𝑖𝑝 𝑦 . These modules contain a large number of functions and
4.6 Exercise: Slicing and Copying 29
methods to model and solve complex engineering problems. Together
with the visualization module 𝑚𝑎𝑡𝑝𝑙𝑜𝑡𝑙𝑖𝑏 these modules offer a very
powerful environment for scientific use.

4.1 Basics of Numpy

The module 𝑛𝑢𝑚𝑝 𝑦 provides Python with very data structures that
can be used in mathematical programming. In particular, it allows the
definition of vectors and matrices in an easy way. On top of that it contains
a multitude of linear algebra routines. In this chapter we present the
main features of the numpy module. Additional functionality will be
introduced throughout the next chapters when needed. You can also
find all the information related to the numpy module in the online
documentation [Link]
As all modules, 𝑛𝑢𝑚𝑝 𝑦 needs to be imported before it can be used.
Therefore to use the numpy module in your code the following line must
be present before any call to any numpy routine
1 import numpy as np

In this case, while importing, we have renamed numpy with a shorter


name, i.e. 𝑛𝑝 , since it will typically be used many times in a code. The
name that you choose is fully free and you may also just omit it and use
the full name ’numpy’ instead.

4.2 Numpy arrays

Numpy provides essentially all the tools to manipulate vectors, matrices,


etc .... The simplest way to create a vector is through the function
𝑛𝑢𝑚𝑝 𝑦.𝑎𝑟𝑟 𝑎 𝑦(). Using this function, a vector (array) can be created by
passing a Python list to the array() as an argument:
1 v = [Link]([1,2,3],dtype=np.float64)

This results in the creation of a one-dimensional array, or a vector:

𝑣 = 1.0 2.0 3.0



(4.1)

Note that in the code listing above we specifically specify the type of
numbers, i.e. 64-bit floating point numbers, even though the numbers in
4 Mathematics using Numpy 25

the list that was provided as an argument contains integer numbers. So


adding the ’dtype=np.float64’ forces the numbers to be stored as floating
point numbers. The same could have been achieved by just adding a
decimal point to each number.
The numpy-array created in this way looks very similar to the Lists
introduced in Chapter 2. The advantages of such arrays over Python lists
become clear if we examine the following piece of code:
1 import numpy as np
2

3 # define two Python Lists


4 a = [1., 2., 3.]
5 b = [4., 5., 6.]
6

7 # tranform these lists


8 # in arrays
9 A = [Link](a, dtype=np.float64)
10 B = [Link](b, dtype=np.float64)
11

12 # compute the sum


13 c = a + b
14 C = A + B
15

16 print(c)
17 print(C)
18

19 C *=10
20

21 print (C)

In this code two Python lists are defined and these are subsequently used
to create two numpy-arrays by passing them to the [Link] function.
The difference between the arrays and lists becomes clear if they are
added up. The sum of two lists stored in the variable 𝑐 leads to a list
containing six numbers, 𝑐 = [1 , 2 , 3 , 4 , 5 , 6], i.e, the concatenation and not
the mathematical sum of the two list. In general, when doing mathematics,
we prefer our arrays of numbers to behave as vectors and this is not what
happens if we use lists. This is solved in numpy arrays that have many of
the characteristics of vectors (and matrices for two-dimensional arrays).
In this case the sum becomes 𝐶 = [5 , 7 , 9] as expected mathematically.
This example nicely illustrates what numpy does: it provides numerical
and mathematical operations for Python.
A large number of methods can be used to create 1D arrays. A few of
them are shown in the snippet of code below
1 v = [Link](5) # [0 0 0 0 0]
2 v = [Link](5) # [1 1 1 1 1]
3 v = [Link](5) # [0 1 2 3 4]
4

5 # a vector containing 100 elements ranging between -pi and pi


6 v = [Link](-[Link],[Link],100)

The last example is particularly useful when we need to construct an


array to serve as the x-axis for evaluating and plotting functions and we
will use this function many time in the remainder of this book.
Apart from one-dimensional arrays or vector we can also create two-
dimensional ones, i.e. matrices. To achieve this we can pass a list that
contains two lists to the function [Link]().
1 mat = [Link]( [ [1,2,3] , [4,5,6] , [7,8,9] ])
4 Mathematics using Numpy 26

Note the extra set of brackets: the list contains three lists each contained
in their own brackets. This line of code creates the matrix

1 2 3
mat = ­4 5 6® (4.2)
© ª

«7 8 9¬

We will see in details how to handle matrices in the next chapter.

4.3 Arrays indexing and slicing

As for the lists each element of an array can be accessed individually


using its index. Reminder: indexes in Python start at 0 !. We can therefore
use loops to fill up arrays as in the example below
1 N = 100
2 X = [Link](-[Link],[Link],N)
3 Y = [Link](N)
4 for i in range(N):
5 Y[i] = [Link](X[i])

where we use the [Link]() function to create a linearly-spaced


x-axis. The np-zeros() function is called to create a numpy-array with
space for N numbers that are filled with zeros. This is very useful to
reserve space in the computer memory if it is not yet known which
numbers we want to store. This code uses a loop to fill up the array Y with
the sine values of the corresponding number in the array X: 𝑌 = sin(𝑋)
for 𝑋 = [−𝜋; 𝜋].

It is also possible to address multiple numbers of a numpy at once by


using so-called slicing. You can define such ’slices’ by giving a range of
indexes very similar as we have seen for the ’range’ function before:
1 Y = [Link](20)
2 Y[10:20] = [Link](10)

In the first line of code a vector 𝑌 = [0 , 1 , 2 , ..., 19] is created. Subsequently,


its last 10 elements are replaced by zeros. The general syntax to specify
such a range is
1 Y[start:end:increment]

where start/end define the first and last index in the range and increment
the distance between two elements of the range. For example
1 Y[0:10:2]

specifies the indexes [0,2,4,6,8], i.e. using an increment of 2. If you do not


specify the increment an implicit value of 1 is assumed. Similarly if you
do not specify the value of ’start’, a value of 0 is assumed. Similarly the
default value of end is the last element. Therefore different parts of the
array can be addressed easily:
1 Y[:10] # first 10 elements
2 Y[10:] # the elements from the 10th to the end
3 Y[:] # all the elements !
4 Mathematics using Numpy 27

In the last line no indexes are give, just the ’:’, which means that all
elements are addressed. Finally, we can also use negative indexes to start
counting from the end of the array, for instance the last, and second-last
elements are accessed by:
1 Y[-1] # last element
2 Y[-2] # second to last
3 ...

The true strength of slicing matrices becomes apparent when we apply


them to matrices. The indexing of matrix elements follow the same syntax
than the indexing of arrays. Therefore, the indication mat[i,j] gives the
elements of the matrix on the i-th line and j-th [Link] example for
the matrix given in eq. 4.2, the command x = mat[1,2] assigns a value of 6
to the variable 𝑥 . Remember that the index starts at 0, therefore, mat[0,0]
is the element on the first line and first column.
We can also access entire blocks of the matrix by defining range of indexes
as for the 1D vectors. This is particularly useful when we want to address
an entire column/line of a matrix with the instruction:
1 col_0 = mat[:,0] # first column
2 lin_0 = mat[1,:] # second line
3

The manipulation of numpy-arrays has many more approaches and in


the following chapters we will see more ways of handling different parts
of these arrays.

4.4 Vectorization

Arrays can be filled up with numbers using different methods. A straight-


forward approach is to use loops as we have done in the example at the
beginning of the chapter 4.3. However, the use of a user-defined loop is
generally quite slow. To solve this, we can exploit a numpy-feature called
vectorization. Instead of passing numbers to numpy functions one by one
as we did in the loop, we can pass a complete numpy array in one go:
1 N = 100
2 X = [Link](0,[Link],N)
3 Y = [Link](X)

In this code, the complete array is passed to the function [Link] and the
function returns a numpy-array with all the results.
While this leads to the exact same results as in the implementation using
a for-loop, this vector implementation is much faster. To illustrate this,
the time required to fill up the the array 𝑌 for different number of points
is listed in Table 4.1.

Table 4.1: Comparison of the time needed


𝑁 = 10 𝑁 = 1𝐸 3 𝑁 = 1𝐸 5 𝑁 = 1𝐸 7 to calculate the sine of a series of numbers,
either in a for-loop or by using Numpy-
𝑡loop 3.2 1E-5 1.7 1E-3 1.6 1E-1 15.5
vectorization.
𝑡vect 2.0 1E-6 1.4 1E-5 1.2 1E-3 1.2 1E-1
4 Mathematics using Numpy 28

It is obvious from this table that it becomes quite attractive to use this
vectorized approaches, especially when the number of elements in an
array is very large.
4 Mathematics using Numpy 29

4.5 The copy issue

It happens very often in programming that a copy of a certain array (or a


slice of it) is need in a separate array. This requires some specific attention
since just using an assignment statement is not enough to achieve this.
Consider the example in the code below
1 X = [Link](20).astype(’float64’)
2 Y = X[0:10]
3 Y += [Link](10)

This code creates an array 𝑋 and second one 𝑌 that contains the 10 first
elements of 𝑋 . Note here the use of the method 𝑎𝑠𝑡 𝑦𝑝𝑒() to force the
numbers to be floating-point numbers. The instruction + = is encountered
for the first time here and simply results in adding a certain number
(here 1) to all the elements of 𝑌 . The results of this code is that we have
indeed added one to all the elements of the array 𝑌 , however, this way of
creating 𝑌 does not result in a physical copy in the computer memory.
Instead 𝑌 simply points to the same physical positions in memory as the
first ten elements of 𝑋 . Therefore, if we now print 𝑋 we will see that its 10
first elements have also been incremented by 1! If we want to circumvent
this issue we have to make sure that we explicitly reserve new memory
positions to store an actual copy of the numbers. This can be achieved in
multiple ways as shown in the code below:
1 # these are example of safe copies
2

3 # create an array
4 A = [Link](20).astype(’float64’)
5

6 # copy with [Link]


7 B = [Link](A[:10])
8

9 # copy directly with [Link]


10 C = [Link](A[:10], copy=true)
11

12 # create a new array and set its elements


13 D = [Link](10)
14 D[:] = A[:10]

We have barely scratched the surface of what 𝑛𝑢𝑚𝑝 𝑦 is capable of. Many
more possibilities will be during in the next chapters to complete the
presentation of numpy. A companion module of 𝑛𝑢𝑚𝑝 𝑦 , called 𝑠𝑐𝑖𝑝 𝑦
will also be introduced. Scipy provides even more sophisticated methods
for scientific computing, as for example the manipulation of sparse
matrices, and is intensively used.

4.6 Exercise: Slicing and Copying

In this exercise we will do some manipulations on numpy-arrays to get


familiar with Numpy. To do this, we will write a code called ’[Link]’.
In this code, create a 1-dimensional array called 𝑇 containing N=251
elements ranging between 0 and 2𝜋 (use the ’[Link]’ function).
To start we simply wish to print selected elements of this array:

I print the second element of the array


I print the last element of the array
4 Mathematics using Numpy 30

I print the element between the indexes 50 and 150


I print the last 25 elements of the array

Now that we know how to address the different elements in an array


we want to compute functions using the values stored in T as x-values.
Compute the function 𝑋 = sin(𝑇). To do that you will define a new array
(X) as:
1 X = ....

Finally we want to illustrate a recurrent issue with the numpy array.


Make a copy of X and multiply this copy by 2 as below:
1 Y = X
2 Y *= 2

Use a for-loop to print a table containing the first 10 values in the arrays
𝑇 , 𝑋 and 𝑌 on the screen. If you inspect the numbers in the arrays X and
Y you will find something is wrong. How can you fix this problem?
Plotting data with Matplotlib 5
5.1 Plotting data with Matplotlib 31
5.1 Plotting data with Matplotlib 5.2 Simple Curves . . . . . . . . . . 31
5.3 Contour plots . . . . . . . . . 32
The plotting of numerical data is a very important aspect of science 5.4 Exercise: Plotting a function 33
general and in scientific computing in particular. Plotting data in graphs 5.5 Extra exercises . . . . . . . . . 34
is a very quick way to check and analyze the results of experiments or Fourier series . . . . . . . . . . 34
Taylor expansion . . . . . . . 34
simulations. The module Matplotlib is one of many libraries that adds
plotting capabilities to Python. The most general set of plotting functions
for numerical data is found in the submodule pyplot, which will we use
here. Like any other (sub-)module, pyplot library has to be imported in
the python code before we can use its functions:
1 import [Link] as plt

Matplotlib is a very large library, so large that entire books have been
written to document the different styles and capabilities it offers. A
lot of the information about the module can be found in its online
documentation [Link] The documentation
contains many coding examples and tutorials to help you mastering
Matplotlib. In this chapter a few basic examples of of making plots are
discussed. In the following chapters a few more of the pyplot capabilities
will become clear, but we will use only a very small subset of the many
possibilities that are available.

5.2 Simple Curves

The simplest example for plotting a graph is plotting the y-values of


a function against a range of x-values. Such an illustrative example is
shown in the the code snippet below where the most common ingredients
for such a graph are shown.
1 import numpy as np
2 import [Link] as plt
3

4 X = [Link](0, 2, 100)
5

6 # ploting the data


7 [Link](X, X**2, label=’X^2’, color=’black’, linewidth=3)
8 [Link](X, [Link](X), ’--’, label=’sin’, color=’blue’, linewidth=2)
9 [Link](X, [Link](X), ’o-’, label=’exp’, color=’red’, linewidth=2)
10

11 [Link](’Time’, fontsize=20)
12 [Link](’F(t)’, fontsize=20)
13

14 [Link](’My plot’)
15 [Link](loc=2)
16

17 [Link]()

This code first creates a numpy-array 𝑋 containing 100 elements ranging


from 0 to 2. Subsequently, it uses the function plot() to plot curves for three
5 Plotting data with Matplotlib 32

different functions; 𝑋 2 , sin(𝑋), and exp(𝑋). The required arguments of


plot() are an array of x-values and an array of y-values that have to have
the same length. The way in which the lines for the different functions are
plotted can be controlled by a series of optional arguments. By default
curves are plotted as plain lines (as for 𝑋 2 ). This can be changed with
the arguments ’–’ and ’o-’ that replace the plain line by a dashed line or a
circle-dash line. The color of the line can be set using the argument 𝑐𝑜𝑙𝑜𝑟
and its width by the argument 𝑙𝑖𝑛𝑒𝑤𝑖𝑑𝑡 ℎ . To include labels for the axes,
the functions [Link]() and [Link]() can be called with the text to be
printed along the axes as an argument. For each of the curves specified
in the plot(), a label was attached. These labels are used in a legend that
can be added to the graph by calling the function [Link](), where the
argument specifies the location of the legend on the graph (a value of 2
meaning that it will appear in the top left). The function [Link]() can be
used to add a title to the graph. Once all the ingredients of the plot have
been assembled, the function [Link]() is called that actually displays
the graph. The code above will generate the plot presented in Fig. 5.1

Figure 5.1: A plot showing a quadratic


and sinus and an exponential on the same
plot with different line style and legends.

5.3 Contour plots

The simple plot shown above is the simplest kind of plot that can be made
with Matplotlib and it will be sufficient for the most common situations.
However, Matplotlib contains a wealth of other plotting options. An
example is a so-called contour-plot that allows to plot a 3D function that
is a function of two arguments. An example of such a function is given
by:.

𝑥
𝑓 (𝑥, 𝑦) = (1 − + 𝑥 2 + 𝑦 3 ) exp(−𝑥 2 − 𝑦 2 ) (5.1)
2

For this function we would like to make a contour-plot, i.e. a series of


lines where the function-value is equal to a certain value. This can easily
be done using matplotlib as shown below.
5 Plotting data with Matplotlib 33

Figure 5.2: Contour plot of the function


𝑓 (𝑥, 𝑦) = (1 − 𝑥2 + 𝑥 2 + 𝑦 3 ) exp(−𝑥 2 − 𝑦 2 ).

1 import [Link] as plt


2 import numpy as np
3

5 def f(x, y):


6 f = (1.0 - 0.5*x + x**2 + y**3)*[Link](-x**2-y**2)
7 return f
8

9 n = 256
10 x = [Link](-2, 4, n)
11 y = [Link](-2, 4, n)
12

13 X, Y = [Link](x, y)
14

15 C = [Link](X, Y, f(X, Y), 8)


16 [Link](C, inline=1)
17 [Link]()

Note that the function [Link]() creates a grid of points encoded in


the matrices X,Y. These matrices are then used to compute the function
𝑓 (𝑥, 𝑦) using vectorization. A double ’nested’ loop over 𝑥 and 𝑦 can also
be used to compute all the values of 𝑓 but this would be considerably
slower. This code produces the plot shown in Fig. 5.2

5.4 Exercise: Plotting a function

In this exercise we will experiment with making some plots of a simple


function. Consider the so-called Gaussian function:

(𝑥 − 𝑏)2
 
𝑦(𝑥) = 𝑎 · exp − (5.2)
𝑐2

This type of bell-shaped function is often found in chemistry and physics,


for instance in optical absorption spectra. The three parameters a, b and c
define the height, the average value (position of the maximum) and the
width of the Gaussian function, respectively.
5 Plotting data with Matplotlib 34

1. Define a function that accepts three parameters: the x-value and


the three parameters a, b and c.
2. Use the function ’[Link]’ to generate a linear series of
x-values (100 numbers should give nice plots) from -2.0 to 2.0.
3. Use ’matplotlib’ to plot a Gaussian function for the following values
of c: 0.2, 0.5 and 0.8 (keep a = 2.5 and b = 0.0).
4. Make sure you make a nice plot with the three curves in different
colors, labels along the axes, a title and a legend.

5.5 Extra exercises

Fourier series

A step function is a piece-wise constant function given by:



 1 0 ≤ 𝑡 < 𝑇/2


𝑓 (𝑡) = 0 𝑡 = 𝑇/2 (5.3)

 −1 𝑇/2 < 𝑡 ≤ 𝑇

This function can be approximated by a Fourier-series that is a sum of


periodic terms. In principle we would need an infinitely long Fourier
series, but with a truncated series we could already have a reasonable
description. The Fourier-series for the step function is given by:

𝑛
2(2 𝑖 − 1)𝜋𝑡
 
4X 1
𝑆(𝑡, 𝑛) = sin (5.4)
𝜋 𝑖=1 2 𝑖 − 1 𝑇

It can be shown that 𝑆(𝑡, 𝑛) → 𝑓 (𝑡) when 𝑛 → ∞. In this exercise we will


explore ’visually’ good the approximation of the step-function is if we
only include a finite set of terms. Implement a code called ’[Link]’
that performs the following steps:

1. A function that calculates 𝑓 (𝑡). You will take 𝑇 = 2𝜋 and use 251
points between 𝑡 = 0 and 𝑡 = 𝑇 . This is the true step-function.
2. A function that calculates the fourier series 𝑆(𝑡, 𝑛) fan any value of
𝑛.
3. The calculation of 𝑆(𝑡, 𝑛) for n=2,5,10,30 and 100. Plot these different
approximations and compare to the true 𝑓 (𝑡) function in the same
graph.

Taylor expansion

In this exercise you will have to create a program called ’[Link]’ to


visualize the accuracy of the Taylor approximation of the sine function.
The sine function can be approximated by a polynomial according to the
following formula:

𝑛
𝑥 2 𝑗+1
(−1) 𝑗
X
sin(𝑥) ≈ 𝑆(𝑥 ; 𝑛) = (5.5)
𝑗=0 (2 𝑗 + 1)!
5 Plotting data with Matplotlib 35

where (2 𝑗 + 1)! is the factorial function that we have implemented in


section 3.1. The quality of the approximation increases with 𝑛 and that’s
what we want to visualize here. To do that you will write a function
that computes 𝑆(𝑥 ; 𝑛). You will then use this function to compute 𝑆(𝑥 ; 1),
𝑆(𝑥 ; 2), 𝑆(𝑥 ; 3), 𝑆(𝑥 ; 6) and 𝑆(𝑥 ; 12) on the interval 𝑥 ∈ [0 , 2𝜋]. Plot the
exact function sin(𝑥) as well as the approximations on the same graph.
Hint: To better visualize the approximation you might want to restrict
the y axis of your plot to [−1 , 1]. You can do that by using the function
[Link]. You can also specify the color and linewidth of each line plot
with the options color and linewidth of the plot() function. The little
snippet of code below illustrate how to do both of these things.
1 import [Link] as plt
2 ....
3 [Link](X,Y,color=’black’,linewidth=4
4 ...
5 [Link](-1,1)
6 [Link]()
Curve Fitting 6
A very common situation in chemistry and chemical engineering is that 6.1 Linear regression . . . . . . . 36
a certain quantity of a system is measured as a function of a certain 6.2 Polynomial regression . . . . 40
parameter. Examples include measurements of the vapour pressure of a 6.3 Fitting to any function . . . . . 41
6.4 Some remarks on fitting . . . 42
liquid as a function of temperature or the product yield as a function of
6.5 Exercise: fitting to a Gaussian 43
temperature in a chemical reactor. In some cases there is a linear relation
between the parameter that is varied and the quantity that is measured.
An example of such data is shown in Figure 6.1 where the yield of a
chemical reaction is plotted as a function of temperature. The data clearly
follow a linear relation. Trends in data, either linear or not, tend to be
more useful if they can be described by a mathematical expression that
approximately matches or ’fits’ the data. Such fits can be made with
software such as Excell or Origin, but an attractive approach is also to
use Python. In this chapter we discuss how fits to experimental data can
be made using Python.

We start of with a description of so-called linear regression, which is


treated in considerable detail in order to illustrate the general concept
and problems that may occur when we fit a function to a data set. The
method discussed can be extended in a straightforward way to include
fits to polynomials of any order. For these general regression fits we first
write our own code, but we also illustrate how to use a routine that is
supplied by Numpy. Finally, we describe a routine from the Scipy module
that allows fitting to any user-defined function.

6.1 Linear regression

The simplest example for making fits to experimental data is a linear


regression fit. This kind of fit occurs very often in engineering and it
illustrates the concepts very nicely. Therefore, we treat it in some detail.
The table below summarises some data from a chemical engineer who
has measured the yield of a certain reaction as a function of temperature. Table 6.1: Example of a linear relation
When the yield is plotted against temperature a clear linear trend emerges. between temperature and reaction yield
This means that we will attempt to fit this data using a linear relation,
T(degrees C) Yield(%)
𝑝(𝑥 𝑖 ) = 𝑐 1 𝑥 𝑖 + 𝑐0 .
100 45
The goal of the linear regression, and more generally of fitting data, is
110 51
to obtain an equation that describes the data as accurately as possible.
120 54
The most important step is the selection of a measure that tells us how
130 61
accurately the fit describes the data. One possibility is to minimise the
140 66
absolute value of the difference between the data points and the fitted
150 70
line. This turns out to be difficult in practice. A common and much more
160 74
convenient choice is to minimise the sum of the square of the differences
170 78
180 85
190 89
6 Curve Fitting 37

Figure 6.1: Scatter plot of the data in Table


6.1.

between the linear fit and the data points. In other words, we would like
to minimise the least squares error, 𝐸 𝑙𝑠 defined as:

𝑛
X
𝐸 𝑙𝑠 = (𝑦 𝑖 − 𝑝(𝑥 𝑖 ))2 (6.1)
𝑖=1

The function 𝑝(𝑦 𝑖 ) can be a polynomial of any order but as for the
purpose of this illustration we will restrict ourselves to a linear function.
We can therefore replace the function 𝑝(𝑦 𝑖 ) by the linear equation in
terms of 𝑐 0 and 𝑐 1 . The goal of linear regression is to minimise 𝐸 𝑙𝑠 , i.e. to
determine the values of the coefficients for which 𝐸 𝑙𝑠 is smallest. Since
eq. 6.1 is a quadratic equation we know that the minimum occurs where
the derivative of this function is zero. Therefore, we take the derivative
of eq. 6.1 with respect to the two unknown parameters:

𝑛
𝜕𝐸 𝑙𝑠 X
= −2 (𝑦 𝑖 − 𝑐 1 𝑥 𝑖 − 𝑐 0 ) = 0 , (6.2)
𝜕𝑐 0 𝑖=1

𝑛
𝜕𝐸 𝑙𝑠 X
= −2 ((𝑦 𝑖 − 𝑐 1 𝑥 𝑖 − 𝑐 0 ) · 𝑥 𝑖 ) = 0. (6.3)
𝜕𝑐 1 𝑖=1

These equations can be simplified noting that constants can be factored


out of the sum, which gives:

𝑛
X 𝑛
X
(𝑦 𝑖 ) = 𝑐 1 (𝑥 𝑖 ) + 𝑛 · 𝑐0 , (6.4)
𝑖=1 𝑖=1

𝑛
X 𝑛
X 𝑛
X
(𝑦 𝑖 𝑥 𝑖 ) = 𝑐1 (𝑥 2𝑖 ) + 𝑐 0 (𝑥 𝑖 ). (6.5)
𝑖=1 𝑖=1 𝑖=1
6 Curve Fitting 38

Figure 6.2: The same scatter plot with the


linear regression fit.

This results in two linear equations with two unknown variables, 𝑐 0 and
𝑐 1 that can be solved easily by hand, giving:

P𝑛
(𝑥 𝑖 ) 𝑛𝑖=1 (𝑦 𝑖 )
P
P𝑛
𝑖=1 (𝑦 𝑖 𝑥 𝑖 )
𝑖=1
− 𝑛
𝑐1 = (6.6)
( 𝑛𝑖=1 (𝑥 𝑖 ))2
P
P𝑛 2
𝑖=1 (𝑥 𝑖 ) − 𝑛

P𝑛 P𝑛
𝑖=1 (𝑦 𝑖 − 𝑐1 ) 𝑖=1 (𝑥 𝑖 )
𝑐0 = (6.7)
𝑛
Although these equations look rather fearsome, in a computer program
it is very easy to calculate all the summations that are involved. In the
code below the summations are evaluated using the numpy routine
[Link], while summations over the product of the elements in
the arrays are evaluated using [Link]. The sum() function simply
returns the sum of all the elements in an array, while dot() calculates
the dot-product of two vectors, which is the sum of the products of the
elements of two vectors. After sorting out the sums, the evaluation of
the two coefficients and hence the best possible linear fit is calculated in
two lines of code (lines 16 and 17). After calculation the best possible
linear fit the code prints the two coefficients to the screen and plots both
the experimental data (using [Link]) and the linear fit. The
output of the program is shown in Figure ?? (right) and it is clear that it
yields a nice linear fit. The resulting fit is the best possible in the sense
that it minimises the least square difference between the measured data
and the fitted line.
1 import numpy as np
2 import [Link] as plt
3

4 # creating numpy arrays, fill them with data


5 temp = [Link]([100., 110., 120., 130., 140., 150., 160., 170., 180.,
190.])
6 yld = [Link]([45., 51., 54., 61., 66., 70., 74., 78., 85., 89.])
7 N = [Link](temp)
8

9 # first we calculate all sums


6 Curve Fitting 39

10 Xsum = [Link](temp)
11 Ysum = [Link](yld)
12 XYsum = [Link](temp, yld)
13 X2sum = [Link](temp, temp)
14

15 # working out the least square linear fit


16 c1 = (XYsum - Xsum*Ysum/N)/(X2sum - Xsum**2/N)
17 c0 = (Ysum - c1*Xsum)/N
18

19 print(c0, c1)
20

21 regression = c0 + c1*temp
22

23 [Link](temp, regression, label=’LR-fit’) # plotting the fitted line


24 [Link](temp, yld, alpha=1.0, label=’experiment’) # plotting
experimental data
25 [Link]("Temperature", fontsize=20)
26 [Link]("Yield", fontsize=20)
27 [Link](loc=2)
28 [Link](’Yield of the reaction’)
29 [Link]()

This piece of code uses two internal numpy functions: [Link](X) and
[Link]. The first one computes the sum of all the elements contained in
the array X. The second one, i.e. [Link](X,Y) computes the dot product
of the two vectors passed in argument. In terms of equations we have

X
[Link](X) = 𝑋𝑛 (6.8)
𝑛
X
[Link](X,Y) = 𝑋𝑛 × 𝑌𝑛 (6.9)
𝑛

The approach used above here can be easily extended to polynomials


of any order and the evaluation is equally simple. In fact, as mentioned
above, regression fits of this type result in sets of linear equations that
can be combined into a matrix notation, in this case:

P𝑛     P𝑛
𝑛 𝑐0
 
𝑖=1
(𝑥 𝑖 ) 𝑖=1
(𝑦 𝑖 )
P𝑛 P𝑛 · = P𝑛 (6.10)
𝑖=1 (𝑥 𝑖 )
2
𝑖=1 (𝑥 𝑖 ) 𝑐1 𝑖=1 (𝑦 𝑖 𝑥 𝑖 )

The solution of such sets of linear equations is the subject of a later


chapter, where methods to solve them are treated in detail. The extension
to polynomials of higher order is very straightforward in this way and
numpy has very efficient routines for solving such systems. In the code
below the matrices 𝐴 and 𝑦 are set up to solve the coefficients 𝑐 0 and 𝑐 1 .
The results and the coefficients are exactly the same as for the code above
where we have explicitly written out the calculation of the coefficients.
1 import numpy as np
2 import [Link] as plt
3

4 # creating numpy arrays, fill them with data


5 temp = [Link]([100., 110., 120., 130., 140., 150., 160., 170., 180.,
190.])
6 yld = [Link]([45., 51., 54., 61., 66., 70., 74., 78., 85., 89.])
7

8 # first we calculate all sums


9 Xsum = [Link](temp)
10 Ysum = [Link](yld)
11 XYsum = [Link](temp, yld)
6 Curve Fitting 40

12 X2sum = [Link](temp, temp)


13

14 # Setting up the arrays


15 A = [Link]([[N, Xsum], [Xsum, X2sum]])
16 f = [Link]([Ysum, XYsum])
17

18 # This is where the set of linear equations is solved by numpy


19 c = [Link](A, f)
20

21 print(c)
22

23 regression = c[1]*temp + c[0]


24

25 [Link](temp, regression, label=’LR-fit’) # plotting the fitted line


26 [Link](temp, yld, alpha=1.0, label=’experiment’) # plotting
experimental data
27 [Link]("Temperature", fontsize=20)
28 [Link]("Yield", fontsize=20)
29 [Link](loc=2)
30 [Link](’Yield of the reaction’)
31 [Link]()

6.2 Polynomial regression

While it is very instructive to write the code for fitting a function to exper-
imental data, this is usually not necessary. In many software packages
(Origin, Igor, MATLAB) there are standard routines available for fitting
functions to any kind of data. Similarly, in the Python modules Numpy
and Scipy there is a variety of routines available for curve fitting. In the
section above we have written our own code for linear regression of any
data, no matter how large the data set is. For this type of fit there is a
simple routine available in numpy, called [Link]. The code below
uses this routine for the data that we have specified above and running
the code will result is exactly the same coefficients and plots as we have
seen above.
1 import numpy as np
2 import [Link] as plt
3

4 # creating numpy arrays, fill them with data


5 temp = [Link]([100., 110., 120., 130., 140., 150., 160., 170., 180.,
190.])
6 yld = [Link]([45., 51., 54., 61., 66., 70., 74., 78., 85., 89.])
7

8 regression = [Link](temp, yld, 1)


9 print(regression)
10

11 regressionvalues = regression[0]*temp + regression[1]


12

13 [Link](temp, regressionvalues, label=’LR-fit’) # plotting the fitted


line
14 [Link](temp, yld, alpha=1.0, label=’experiment’) # plotting
experimental data
15 [Link]("Temperature", fontsize=20)
16 [Link]("Yield", fontsize=20)
17 [Link](loc=2)
18 [Link](’Yield of the reaction’)
19 [Link]()
6 Curve Fitting 41

6.3 Fitting to any function

In most cases the data that you need to fit is not described by a polynomial.
The routine optimize.curve_fit from scipy can be used to fit any function.
This function can be as complicated as we like and it does not have to
be a simple analytical function. As shown in the example below, we can
define our own function (lines 5-6) and then just use that function, on
line 12, to fit it to a certain data set. In this example the data point are
generated using the same function and on line 10 a random gaussian
noise is added. The latter is just for making the data a bit noisy for the
purpose of the example. The resulting (noisy) data with the exponential
fit are shown in Figure 6.3.

Figure 6.3: Exponentially decaying trend


fitted with the general curve_fit routine
from scipy.

The optimize.curve_fit() routine return two things: the optimised coef-


ficients belonging to the function that are stored in an array ’popt’ and
the covariance matrix that is stored in ’pcov’ that contains information
about the accuracy of the fit. As said before, this routine will work for
any function that you define yourself and is therefore very useful for
fitting experimental data.
1 import numpy as np
2 import [Link] as plt
3 from [Link] import curve_fit
4

6 def func(x, a, b, c):


7 return a * [Link](-b * x) + c
8

9 x = [Link](0, 4, 50)
10 y = func(x, 2.5, 1.3, 0.5)
11 yn = y + 0.2*[Link](size=len(x))
12

13 popt, pcov = curve_fit(func, x, yn)


14

15 print(popt)
16

17 [Link](x, func(x, *popt), ’r-’, label="Fitted Curve") # plotting the


fitted line
6 Curve Fitting 42

18 [Link](x, yn, alpha=1.0, label=’Noisy data’) # plotting experimental


data
19 [Link]("X", fontsize=20)
20 [Link]("Y", fontsize=20)
21 [Link](loc=1)
22 [Link](’curve_fit example’)
23 [Link]()

6.4 Some remarks on fitting

The examples shown this section have consisted of quite nice data that
accurately follows a certain trend. This is not always the case and the
correct selection of data requires some attention. One of the very general
issues that are true for all curve fits is the fact that each point weighs
equally in the fit. This means that if there is one point out of 10 that is
very far off the trend, for instance a linear trend, the result can be a very
bad linear fit. An example is shown in Figure 6.4 where one of the points
in the linear regression example from above is moved way off the linear
trend. This results in a line that tries to include this point, resulting in a
line that clearly does not describe the data very well at all! A solution to
this is data selection before the fitting procedure is performed.
A related issue may occur in the fitting of nonlinear functions, for instance
an exponential function. Since in almost all curve fitting routines the
sum of the square of the difference between the data and the fitted
curve is minimised, it is the absolute difference that plays a role. For
an exponential function of which the values can span several orders of
magnitude this may result in a poor correspondence in the tail of the
exponential, where the absolute differences may still be small but the
relative difference can be very big. As a result, after fitting, the exponential
fit will follow the high points with reasonable accuracy but plotting it Figure 6.4: Illustration of the effect of out-
liers. One point that is far off the linear
on a log scale typically shows quite severe deviations in the tail of the trend can make the resulting linear trend
exponential. very bad.

Finally, fitting to a general function is something that can be done quite


easily with a wide variety of software packages, but it is usually a good
idea to think about it for a few seconds before diving into it. As an
example you may think about the exponential fit mentioned above, for
instance the rate of charge transfer in a donor-acceptor molecule as a
function of distance:

−𝛽
𝑘 𝐶𝑇 (𝑑) = 𝐴0 𝑒 𝑑 , (6.11)

which can easily be recast in the form of a linear equation, plotting the
natural logarithm of the rate against the inverse of the distance:

1
ln 𝑘 𝐶𝑇 = ln(𝐴0 ) − 𝛽 . (6.12)
𝑑

Fitting of this linear relation can easily be done, for instance by linear
regression as described in detail above. The result will in most case be a
more evenly weighted fit where the low points are better described than
when just fitting an exponential function directly. The latter is particularly
6 Curve Fitting 43

relevant when the correspondence to an exponential is not perfect, as is


usually the case.

6.5 Exercise: fitting to a Gaussian

In this exercise we will use the definition of the Gaussian function that
we did in Chapter 6 and make a fit to some experimental data. Some one
has done some measurements, obtaining a series of 10 x-values and a
corresponding series of 10 y-values:
1 import numpy as np
2

3 xn = [Link]([-1.99, -1.52, -1.17, -0.65, -0.19, 0.17, 0.61, 1.26,


1.55, 2.04])
4 yn = [Link]([-0.02, -0.08, -0.00, 0.32, 1.79, 2.08, 0.22, -0.15,
0.16, 0.07])

These code-lines can be downloaded from Brightspace: [Link].


We will now make a fit of a Gaussian function to these data points:

1. Use the lines above in your code to generate two numpy arrays.
2. Make a scatter-plot that visualizes the data.
3. Copy the function that you defined for the Gaussian in Chapter 6
and add it to your code.
4. Use the [Link] function 0 𝑐𝑢𝑟𝑣𝑒 _ 𝑓 𝑖𝑡 0 and your Gaussian
function to make the best fit to the data-points and print the values
of the parameters a, b and c to the screen.
5. Add a smooth Gaussian curve (100 points) to the scatter-plot that
you made above and confirm that the fitted curve nicely follows the
experimental data. Make sure you make a pretty plot with legend,
axis labels, a title and different colors for the scatter plot and the
fitted curve.
Reading and analyzing data 7
The rigorous statistical analysis of large amounts of data is an important 7.1 Reading data from a file . . . 44
aspect science and engineering problems. Experimental results often 7.2 Statistics and histograms . . 47
comes in the form of a large data set that needs further processing to 7.3 Exercise: fitting many curves 49
7.4 Extra exercises . . . . . . . . . 50
extract meaningful information. This further processing can be done in
Process Data . . . . . . . . . . 50
many ways, for instance simply in programs such as Excel or Origin,
or specialized software packages for statistical analysis such as SPSS.
Powerful tools have also been developed to perform statistical analysis
directly within Python. This is for example the case in the module
[Link]() that contains a range of methods to process complex data.
In this chapter we will discuss some very basic aspects of analyzing
large data sets. The subjects treated will be very useful and often directly
applicable to student research projects.

7.1 Reading data from a file

Before data can be processed in a Python code, we first need to read the
data from a file, and later maybe write the resulting processed data back
into another file. Several options are available to read/write data from
file depending on its format. The most common file formats are space-
separated values, where each line contains several numbers separated by
a blank space, and character-separated values (csv) where the numbers
are separated by a character, typically a comma or semicolon. This format
is very popular to share data between different programs. Excel can, for
example, export spreadsheets in .csv format.
Space-Separated Values Data files that contain column of data where
the different value on each line are separated by a space can be directly
read into your Python code by using the function 𝑛𝑢𝑚𝑝 𝑦.𝑙𝑜𝑎𝑑𝑡𝑥𝑡().
Similarly, you can write data to a file with the function 𝑛𝑢𝑚𝑝 𝑦.𝑠 𝑎𝑣𝑒𝑡𝑥𝑡().
As an example, we would like to import the data shown in Fig. 7.1, into a
Python code:

Figure 7.1: Example of a data file using


space separated values formatting.

This data file contains, as the first column, a list of numbers running
from 0 , 1 , 2...., 𝑁 . The other columns in the file are filled with a series of
random numbers. If the name of this file is ’[Link]’, this data can be
loaded into a numpy array using the function [Link]():
7 Reading and analyzing data 45

1 matrix_data = [Link](’[Link]’)

This function will read the data from the file and return a two-dimensional
numpy-array that is assigned to matrix_data. This matrix contains all the
numbers present in the file. We can now access the data in the file, for
instance to plot it in a graph by taking different ’slices’ of it. For example
the first column can be accessed using the indexing that indicates all
lines in the first index (:) while only selecting the column index 0:
1 first_col = matrix_data[:,0]

Note that in this case the variable name first_col points to the same
values in memory as the first column in the array column_data, it is not
a copy!
Similarly, the values contained in a two-dimensional numpy-array can be
written to a file using the function [Link]():
1 [Link](’my_file.dat’,mat_data)

This function will generates a file called my_file.dat that contains all the
elements of the matrix mat_data. The save and read commands have a lot
of other argument as for example the format used to print each number
(%f,%e,...), but the default options work very well in most cases.
Character-separated values. The space separated values may seem like
the most logical format, but character-separated values formatting is
used by many programs and experimental equipment. These files cannot
be imported using the [Link]() function as they contains commas. A
typical .csv file is shown in Fig. 7.2

Figure 7.2: Example of a data file using


comma separated values formatting.

This file contains a header, i.e. a first line containing a description of


the data contained in the file. On the lines below that columns of data
appear where the values on each line are separated by commas. The
values contained in this file can be loaded into a numpy-array using the
function [Link]():
1 matrix_data = [Link](’[Link],delimiter=’,’)

This function allows loading data in a generic format in a numpy array


by specifying the delimiter used in the file. If the Header line starts
with a "#" sign it will be ignored during the loading. If it does not
start with a ’#’ symbol, the loading of the data will most likely fail.
Therefore, one must always check this *.csv file before trying to load it
with [Link](). A dedicated module csv has been implemented for
Python [Link] This module offers
advanced reading and writing capabilities that are beyond the scope of
this text.
7 Reading and analyzing data 46

General file format. It is always possible to manually load all the values
contained in a file using the standard input/output functionality of
Python. This can be be done by reading the file line by line and storing
the values contained on each line in a List or a [Link]. This method
can be used to read any type of files with any formatting (or no formatting
at all). We illustrate this approach in the snippet of code below that reads
a .csv file with a header.
1 import numpy as np
2

4 def read_csv(filename):
5

6 # open the file


7 f = open(filename, ’r’)
8

9 # read all the lines


10 data_string = [Link]()
11 nLines = len(data_string)
12

13 # close the filename


14 [Link]()
15

16 # declare a list where we store the data


17 data_float = []
18

19 # for all the lines but the first


20 # since it s a header
21 for iL in range(1, nLines):
22

23 # split the line at each comma


24 line = data_string[iL].split(’,’)
25

26 # append the float to the list


27 data_float.append([float(x) for x in line])
28

29 # return the float


30 return [Link](data_float)

This examples requires much more explanation since the loading of the
numerical data in the file takes place in many steps. The function open()
opens the file called filename, where the option ’r’ indicates that the file
is for reading only. This is also the default option. Other options such
as ’w’ allows writing to a file. The function [Link]() reads all the
lines contained in the file one by one and stores them in the variable
data_string. This variable is a 1D array where each element correspond
to one line of characters. Hence, data_string[0] is the first line stored as a
string of characters and so on. The next important function that is used
in the code above is line = data_string[iL].split(’,’). This function takes
one individual line ( data_string[iL]) and splits it at each comma ,. For
example the string ’1.00, 2.00, 3.00’ will become a 1D vector containing
3 separate elements: [’1.00’, ’2.00’, ’3.00’ ]. All these elements are still
strings of characters and not numbers. To convert these text-strings into
numbers we use the function float(x) that converts a string into a floating
point number if possible.
The approach of reading a file line-by-line and processing the data in the
form we want is a bit more complicated than the much simpler numpy
functions used above but it does allow full control over the reading of
the data. In addition, it is sometimes impossible to use the [Link]() or
[Link]() and the method shown here is the only solution capable
7 Reading and analyzing data 47

of doing the job correctly.

7.2 Statistics and histograms

Once we have loaded the data into a proper numpy array we can start
processing its elements. In the file presented above, the first column
represent the time of the experiment and the each following column
contains the value of certain physical quantity measured at these times.
We can therefore plot the time evolution of this data using matplotlib
as shown in Fig. 7.3.
1 import numpy as np
2 import [Link] as plt
3

4 DATA = [Link](’[Link]’)
5 X = DATA[:, 0]
6 Y = DATA[:, 1:]
7

8 [Link](X, Y[:, 1], color=’#FF003F’, label=’1’)


9 [Link](X, Y[:, 0], color=’#00AFFF’, label=’1’)
10

11 [Link](’time’, fontsize=15)
12 [Link](’Measurements’, fontsize=15)
13 [Link]()

Figure 7.3: Time evolution of two mea-


surements contained in the data file.

The data looks very noisy and not much information can be extracted
from inspection of the plots. We can just observe that the mean value of
the blue curve seems lower than that of the red ones. To obtain a more
quantitative analysis of the data we can analyze the data in some more
detail. Different modules such as Numpy and Scipy offers a vast library
of statistical function that can be accessed easily. For example the mean
values of the elements of a vector 𝑋 defined as

1 𝑁−
X1
𝜇= 𝑋[𝑛] (7.1)
𝑁 𝑛=0

can be calculated via the two functions


1 m = [Link](X)
2 m = [Link]()

provided that 𝑋 is a Numpy-array. Similar syntax exist for the minimum


and maximum values contained in an array (min and max). The variance
7 Reading and analyzing data 48

Table 7.1: Output of calculations of the


mean median min. max. var. std. mean, median, minimum and maximum
values, variance and standard deviation
Data set 1 -0.89 -0.89 -1.17 0.19 0.075 0.27
of a series of numbers.
Data set 2 0.06 -0.01 -0.75 2.51 0.14 0.38

and standard deviation of a time series are defined by

s
𝑁−
1 X1 1 𝑁−
X1
var = (𝑋[𝑛] − 𝜇)2 𝜎= (𝑋[𝑛] − 𝜇)2 (7.2)
𝑛 𝑛=1 𝑛 𝑛=1
where 𝜇 is the mean value of the vector. These two quantities can be
calculated for the values in a numpy-array can be calculated using the
functions [Link]() and [Link]() or the equivalent syntax [Link](X) [Link](X).
These quantities define how the values are dispersed around the mean
values and this plays a key role in error analysis. Finally, the median
of a time series, i.e. the value that is superior to half of the number
contained in the series and inferior to the second half is returned by the
function [Link](X).The code below illustrates the basic usage of these
functions
1 nLine, nCol = [Link]
2 for iCol in range(nCol):
3 print ’\n === Data Column %02d’ %(iCol+1)
4 print("\tMean : %1.6f" %(Y[:,iCol].mean()))
5 print("\tMedian : %1.6f" %([Link](Y[:,iCol])))
6 print("\tMinimum : %1.6f" %(Y[:,iCol].min()))
7 print("\tMaximum : %1.6f" %(Y[:,iCol].max()))
8 print("\tVariance : %1.6f" %(Y[:,iCol].var()))
9 print("\tStd deviation : %1.6f" %(Y[:,iCol].std()))

This code will produce the values reported in the table below for the first
two columns of data.
The calculation of these quantities confirms our observations. The mean
values of the first two columns are -0.89 and 0.06, respectively. In addition,
the variance of the first time series is smaller than for the second indicating
a smaller spread of values around the mean. Finally, we can see that the
mean and median values are equal for the first time series but not for the
second, which indicates that the shapes of the distributions of random
values are quite different.
The same insights can also be obtained if we plot these data, but not as a
function of the time but as a histogram. Matplotlib offers an easy way to
construct such a histogram. Using the function [Link](X,bins=nbin,...)
will directly generate a histogram of the data contained in X using a
certain number of intervals (or bins) to compute the count. The use of
this function is illustrated in the code below.
1 import numpy as np
2 import [Link] as plt
3

4 # inport the data


5 DATA = [Link](’[Link]’)
6 X = DATA[:, 0]
7 Y = DATA[:, 1:]
8

9 # plot histograms
10 [Link](Y[:, 0], bins=50, facecolor=’#00AFFF’, alpha=0.5)
11 [Link](Y[:, 1], bins=50, facecolor=’#FF003F’, alpha=0.5)
12 [Link](’Count’, fontsize=15)
7 Reading and analyzing data 49

Figure 7.4: Histogram of the values con-


tained in the first two columns of the data
file.

13 [Link](’Measurements’, fontsize=15)
14 [Link]()

In this code the function [Link]() accepts as the most important argument
a numpy array containing the data to be plotted. In addition, different
optional arguments can be supplied to specify the number of ’bins’ to be
used or the color of the histogram and its level of transparency (alpha).
This generates the plot in Fig. 7.4.
Plotting the histogram gives a quick insight in the main differences
between the two data sets. The blue distribution seems quite symmetric
around its mean value, which explains why its mean is identical to its
median. In contrast, the red distributions is not symmetric, leading to the
difference between its mean and median value. These two distributions
consequently look like a normal (Gaussian) and a Poisson distribution,
respectively.

7.3 Exercise: fitting many curves

In a lab course, students have measured the so-called fluorescence


lifetime of a specific perylene bisimid (PDI) molecule. The data of all
these students is collected in a file [Link]. In this file, the
first column contains the time in nanoseconds and all other columns
describe the relative number of emitted photons at the corresponding
times recorded by the different students.

1. Write a program that reads the data in the file


2. Now make a ’scatter plot’ where you plot the first data column
against the time axis (the first column in the file).
3. Define a function for the exponential description of fluorescence
intensity I given below.

𝑡
 
𝐼(𝑡) = 𝐴 · exp − (7.3)
𝜏𝑓 𝑙
where t is time, is the fluorescence life time and A is a pre-
exponential factor.
4. Use the [Link] function 0 𝑐𝑢𝑟𝑣𝑒 _ 𝑓 𝑖𝑡 0 (see Chapter 6) to fit
this one data set to the exponential function that you defined above.
Make a plot where you combine the scatter plot you made above
with the optimal fit.
7 Reading and analyzing data 50

5. Now modify your code to automatically fit the fluorescence decay


curves for all experiments in the file and store the different values
for 𝜏 𝑓 𝑙 in an array. This requires a for loop that runs over the
number of data columns in the array that you read from the file.
6. Make a histogram of the fluorescence lifetimes, 𝜏 𝑓 𝑙 , obtained above
using 20 bins.

7.4 Extra exercises

Process Data

You’ve been hired by a new start-up to characterize the reliability of


their manufacturing process. The start up builds little devices that when
illuminated with a ultra-short laser pulse shines from a brief period of
time. To characterize their devices they record the light coming out of
the device and obtain for each device a curve showing the light intensity
as a function of time. Few of these curves are shown in Fig. 7.5. These
curves can be fitted with an exponential function.

𝐹(𝑡) = exp(−𝑘𝑡) (7.4)

where 𝑘 is the decay parameter. Of course each device will lead to a


different value of 𝑘 . The fit of few of curves are shown in Fig. 7.5.

Figure 7.5: Light intensity measured for 3


different devices. The dash line show the
result of the exponential fit performed to
evaluate the values of 𝑘 for each device.
You can already see that very different
𝑘 -values are expected for the different de-
vices.

To facilitate your analysis the company has built 250 devices on three
different days. They’ve measured the light coming out of these devices
and gave you 3 space separated files ’exp_data_1.dat’, ’exp_data_2.dat’
and ’exp_data_3.dat’ that you can download on the Blackboard. Each
file contains 251 columns where the first column is the time (i.e. the
x-axis) and the 250 others are the light intensity coming out of the 250
devices built and tested that day. The company wants you to estimate
the distribution of decay parameters obtained for all these devices. To
analyze the data create a script called ’[Link]’. In this script you
will do for each file:

I Load the data in your program


I Fit the 250 signals it contains with an exponential decay to evaluate
the corresponding value of 𝑘 . You will end up with a distribution
of k-values that characterize all the devices tested that day.
7 Reading and analyzing data 51

I Plot the histogram of the k-value. You should obtain a figure similar
to Fig. 7.6 (we show here the distributions obtained for the three
files.)
I Estimate the mean value and standard deviation of each of the
distribution.

The target of the company was to obtain a mean value of 𝑘 ranging


between 1.5 < 𝜇 < 2.5 with a standard deviation of 𝜎 < 2.0. Is this goal
achieved for the different days? The company also wants to know if the
devices created on different days are similar or not. Use a t-test on the
three distributions of k-values to answer that question.

Figure 7.6: Histograms of the k-values ob-


tained by fitting the data contained in the
three files.
Vectors, Matrices and Systems of
Linear Equations
Vectors and Matrices 8
The correct handling of vectors and matrices plays a role in many 8.1 Simple vectors and matrices 53
problems in chemistry and chemical engineering. Methods involving 8.2 Slicing vectors . . . . . . . . . 54
matrices and vectors are for instance used in solving sets of linear 8.3 Slicing Matrices . . . . . . . . 54
8.4 Complicated Matrices . . . . 56
equations, solving differential equations, and in deriving approximate
8.5 Linear Algebra . . . . . . . . . 57
solution of the Schrödinger equation. As we have already seen in Chapter
8.6 Matrix-Vector multiplication 58
4, the module Numpy contains a convenient data structure for dealing
8.7 Determinant: Recursion . . . 59
with vectors and matrices. In addition, Numpy contains a wide range 8.8 Exercise: Matrix Multiplica-
of function to manipulate vectors and matrices. In this chapter we will tion . . . . . . . . . . . . . . . . . . . 61
discuss several other approaches to creating and manipulating matrices
and we will see how we can do mathematics with matrices in Numpy.

8.1 Simple vectors and matrices

A vector is a one-dimensional array of numbers that is usually represented


as:

𝑉 = 𝑣0 𝑣1 ... 𝑣𝑁

(8.1)

where 𝑣 0 is the first element of the vector, 𝑣 1 the second .... As we have
already seen such a numpy-array can be constructed using the numpy-
function-array() that takes a Python-list of numbers as an argument:

1 v=[Link]([12.24, 23.10, 7.89, 1.03, 0.01])

This leads to a one-dimensional array where we have set 𝑣 0 = 12.24,


𝑣1 = 23.10, and so on. We have also already seen how we can create a
two-dimensional numpy-array or a matrix, for instance defined as:

𝑎 0 ,0 𝑎 0 ,1 ... 𝑎0,𝑁
­ 𝑎 1 ,0 𝑎 1 ,1 ... 𝑎1,𝑁 ®
© ª
𝐴 = ­­ . .. .. ®® (8.2)
­ .. . ... . ®
« 𝑎 𝑁 ,0 𝑎 𝑁 ,1 ... 𝑎 𝑁 ,𝑁 ¬

Where 𝐴 is the name of the matrix and 𝑎 𝑖,𝑗 its element on the 𝑖 -th row
and 𝑗 -th column. Matrices can be constructed using the numpy-function
array() that takes a list filled with lists as its argument in this case. For
example the code:
1 M = [Link]([[1, 2, 3], [4, 5, 6], [7, 8, 9]])

leads to the definition of the matrix:

1 2 3
M = ­4 5 6® (8.3)
© ª

«7 8 9¬
8 Vectors and Matrices 54

Note that the argument given to the array() function is a list of which the
elements are lists themselves. Therefore, there is an extra set of square
brackets.

8.2 Slicing vectors

It is often useful to extract part of a large matrix. We have seen examples of


that in Chapter 7 where we wanted to decompose a large two-dimensional
matrix into columns for plotting or manipulating experimental data sets.
If the indexes of the elements to be extracted are contiguous, we can use
the ":" operator to extract a certain part of a matrix. The operator ":" can
be used to generate a series of indexes of an array. Hence in the code
1 u = [Link](10) #u = [0 1 2 3 4 5 6 7 8 9]
2 v = u[1:3] #v = [1 2]

𝑣 contains the values of 𝑢 ranging from the first to the second index
(remember the index starts at 0). In general the statement:
1 v = u[a:b]

extracts the elements of 𝑢 ranging from 𝑎 upto (but not including) 𝑏 and
stores them in 𝑣 . If you omit the 𝑎 or 𝑏 , the default values 𝑎 = 0 and
𝑏 = 𝑁 will be used, where 𝑁 is the total length of the vector. We have
therefore:
1 u = [Link](10) #u = [0 1 2 3 4 5 6 7 8 9]
2 v = u[:4] #v = [0 1 2 3]
3 v = u[6:] #v = [6 7 8 9]
4 v = u[:] #v = [0 1 2 3 4 5 6 7 8 9]
5 v = u[1:-2] #v = [1 2 3 4 5 6 7]

On line 5 a negative index is used which means that the second element
from the end of the array is indicated. The index -1 corresponds to the
last element of the array, -2 the second last and so on.

8.3 Slicing Matrices

A similar approach can be used to extract elements from a matrix by


specifying indexes for the line and the column of the element. As for
one-dimensional array, the indexing of matrices in Python starts at 0.
Hence the indication:
1 mat[1,1]

refers to the element on the 2nd row and 2nd column of the matrix. To
extract a complete row of a complete column of a matrix we can use:
1 A = [Link](25)
2 mat = [Link](A,(5, 5))
3 col = mat[:, 2]
4 line = mat[3, :]

The first two lines of this code serve to generate a two-dimensional matrix.
First, we use the function arange() to create a linear array. This array is
transformed into a two-dimensional array using the function reshape()
that takes as its second argument a ’shape’ that consists of two numbers.
These two numbers indicating the shape have to be contained between
8 Vectors and Matrices 55

brackets because they, together, form a single argument to the function


reshape(). The statements on lines 3 and 4 indicate the column with index
2 and the row with index 3, respectively. It is also possible to address a
block of element, or a sub-matrix, by using ranges of elements:
1 sub = mat[:4, :4]

This statement extracts a 3-by3 sub-matrix containing only the rows and
column ranging between from the first (with index 0) to the third (with
index 3).
If the indexes of the elements to be extracted are not contiguous, the
numpy function 𝑖𝑥 can be used to extract them. This function constructs
a new matrix taking elements from an existing matrix where the ’mesh’
of elements has to be supplied as an argument. This mesh consist of two
Python lists that contain the row and column indexes to be extracted and
can be used as a ’mask’ to extract certain values to form a sub-matrix. An
example is shown in the code below.
1 mat = [Link](25).reshape(5, 5)
2 ind1 = [0, 2, 4]
3 ind2 = [0, 2, 4]
4 sub = mat[np.ix_(ind1, ind2)]

The first line of the code shows an alternative use of the reshape() function
in which is is directly appended to the arange function that generates a
linear array. The effect is exactly the same as in the example above where
the reshape function was called separately. On the second and third lines
of the code the row and column indexes are indicated in a Python list.
On line 4, these lists are passed as arguments to the [Link]() function
to generate the mesh of indexes to be assigned to the sub-matrix. The
codes eventually leads to the following matrices:

0 1 2 3 4
­5 6 7 8 9® 0 2 4
© ª
mat = ­10 11 12 13 14® −→ sub = ­10 12 14® (8.4)
­ ® © ª
­15 16 17 18 19® «20 22 24¬
­ ®

«20 21 22 23 24¬

The sub-matrix is generated by taking only the elements on the 0th, 2nd
and 4th row and 0th, 2nd and 4th [Link] this case the same indexes
have been chosen for the columns and rows, but they can be different
also, as long as they do not contain an index larger than the size of the
matrix.
It is finally possible to change the values of the elements contains in a
row, a column or a part of a matrix, for instance by a random number as
in the example below.
1 mat = [Link](25).reshape(5, 5).astype(float)
2

3 mat[:, 2] = [Link](5)
4 mat[1, :] = [Link](5)
5

6 ind2 = [0, 2, 4]
7 mat[np.ix_(ind1, ind2)] = [Link](3, 3)
8 Vectors and Matrices 56

On the first line we have used the function ’astype(float)’, which forces
Python to interpret the number in the matrix as floating point num-
bers rather integers, which is the default for [Link](). The function
[Link]() generates a series of random numbers between 0
and 1. The argument, in this case 5, indicates how many random numbers
are to be returned. The conversion of the elements in the matrix ’mat’
to floating point numbers is important here because otherwise, all the
random numbers would be truncated to make them an integer and hence
they would all be equal to 0.

8.4 Complicated Matrices

So far we have only considered square matrices, however, there are


many examples where matrices of other shapes are required. Several
of the functions in numpy cab generate matrices or arbitrary size. For
instance the functions that we used before for one-dimensional matrices,
[Link]() and [Link]() can also be called with an argument
that consists of two numbers between brackets, a so-called tuple, that
defines the shape of the matrix.
1 import numpy as np
2 N = 10 # Number of columns
3 M = 15 # Number of rows
4 Z = [Link]((N, M)) # a null matrix note the double bracket
5 O = [Link]((N, M)) # a matrix with ones everywhere
6 E = [Link](N) # a square diagonal matrix (identity matrix)

The two components of the ’shape’ are pass with an extra set of brackets,
(𝑁 , 𝑀), to indicate that they together form a single argument! On the
last line of the code above we use the function [Link]() to generate
an identity matrix, which is a square matrix by definition and hence no
information other than the size is needed.

Arbitrarily shaped matrices can also be generated using the reshape()


function that we have seen above. For example the statement:
1 X = [Link](N*M).reshape(N, M)

first generates a one-dimensional array ranging from 0 to (𝑁 ∗ 𝑀) − 1 and


subsequently reshapes this array into a matrix of 𝑁 rows and 𝑀 columns.
Calling the [Link]() function with two arguments leads to
the generation of a matrix filled with random numbers:
1 R = [Link](N, M)

The elements elements are random numbers, uniformly distributed


between 0 and 1. The generation of random numbers is important for
some types of simulations, for instance in Monte Carlo methods that are
discussed later in this text.
So-called nested loops can also be used to fill up matrices with numbers.
In order to create a 𝑁 × 𝑀 matrix with the value of each of the elements
given by 𝑀 𝑖,𝑗 = 𝑖 ∗ 𝑗 we can use the following code:
1 mat = [Link]((N, M))
2 for iL in range(N):
3 for iC in range(M):
4 mat[iL, iC] = iL*iC
8 Vectors and Matrices 57

This code contains two loops: one loop inside another loop, or nested
loops. The first loop, called outer loop, runs over all the rows, while
the second one (the inner loop) runs over the columns. This code works
perfectly fine and gives full control over what happens, but in many cases
there are faster ways of doing the same thing by exploiting vectorization
in numpy routines. As an example, the matrix created above can also be
obtained by calculating the outer product of two vectors. If we have two
vectors u and v that are 𝑀 × 1 and 𝑁 × 1 vectors, then the outer product
of these two vectors is given by:

𝑢1 
𝑇
 ©𝑢1 𝑣1 𝑢1 𝑣 2 𝑢1 𝑣 3
u ⊗ v = uv = 𝑢2  𝑣 1 𝑣2 𝑣 3 = ­𝑢2 𝑣1 𝑢2 𝑣 2 𝑢2 𝑣 3 ®
 
(8.5)
ª
𝑢3 
  «𝑢3 𝑣1 𝑢3 𝑣2 𝑢3 𝑣 3 ¬

where 𝑀 = 𝑁 = 3. In a compact notation, (u ⊗ v)𝑖𝑗 = 𝑢𝑖 𝑣 𝑗 . The final


matrix can be directly obtained via the numpy 𝑜𝑢𝑡𝑒𝑟 function:
1 u = [Link](10)
2 v = [Link](10)
3 mat = [Link](u, v)

Note: The inner product of two vectors, hu , vi = u𝑇 v = 𝑢𝑖 𝑣 𝑖 , can also


P
𝑖
be calculated via the 𝑖𝑛𝑛𝑒𝑟 function of numpy:
1 u = [Link](10)
2 v = [Link](10)
3 IP = [Link](u, v)

For two one dimensional vectors the function [Link]() has the same
effect as the function [Link]() that we have seen earlier.

8.5 Linear Algebra

Calculations with matrices are closely related to linear algebra. A large


number of routines for linear algebra have been implemented in Numpy.
An example is the multiplication of two matrices 𝐴 and 𝐵 defined by:

𝐴11 ... 𝐴1 𝑚 𝐵11 ... 𝐵1 𝑝


© . .. .. ª® © . .. .. ª®
𝐴 = ­­ .. . . ® and 𝐵 = ­­ .. . . ® (8.6)
«𝐴𝑛 1 ... 𝐴𝑛𝑚 ¬ «𝐵𝑚 1 ... 𝐵𝑚𝑝 ¬

The product of these two matrices is given by:

𝐶11 ... 𝐶1𝑝 𝑚


© . .. .. ª®
𝐶 = 𝐴𝐵 = ­­ ..
X
. . ® with 𝐶 𝑖𝑗 = 𝐴 𝑖𝑘 𝐵 𝑘 𝑗 (8.7)
𝑘=1
«𝐶 𝑛 1 ... 𝐶 𝑛𝑝 ¬

This complex operation can be performed with the function 𝑛𝑢𝑚𝑝 𝑦.𝑚𝑎𝑡𝑚𝑢𝑙()
or 𝑛𝑢𝑚𝑝 𝑦.𝑑𝑜𝑡():
8 Vectors and Matrices 58

1 m = 10
2 n = 15
3 p = 5
4 A = [Link](n,m)
5 B = [Link](m,p)
6 C = [Link](A,B)

In addition to such basic operations, Numpy contains a large number of


functions to perform more complicated operations on matrices. There are
for instance functions for calculating of the eigenvalues and eigenvectors,
the determinant and the inverse of a matrix. These operations can be
performed as shown in the code below in a single instruction:
1 import [Link] as npla
2 det = [Link](A) # determinant of the matrix A
3 w, v = [Link](A) # eigenvalues (w) and eigenvectors(v) of the matrix A
4 Am1 = [Link](A) # inverse of the matrix A

8.6 Matrix-Vector multiplication

The manipulation of lines and columns in a matrix is encountered in


many problems. An illustrative example is the multiplication of a matrix
and a vector. We will consider the matrix-vector multiplication: u = 𝐴 · v.
This multiplication is fully written as:

𝑢0 𝑎 0 ,0 𝑎 0 ,1 ... 𝑎0,𝑁 𝑣0
­ 𝑢1 ® ­ 𝑎1,0 𝑎 1 ,1 ... 𝑎1,𝑁 ® ­ 𝑣1 ®
© ª © ª © ª
­ . ®=­ . .. .. ®® · ­­ .. ®® (8.8)
­ . ® ­ . .
­ . ® ­ . ... . ® ­ . ®
«𝑢 𝑁 ¬ « 𝑎 𝑁 , 0 𝑎 𝑁 ,1 ... 𝑎 𝑁 ,𝑁 ¬ «𝑣 𝑁 ¬

where

𝑁
X
𝑢𝑖 = 𝑎 𝑖,𝑗 · 𝑣 𝑗 (8.9)
𝑗=0

In order to calculate the different elements of the vector u we need to


calculate, for each element, the inner product of a given row of 𝐴 with the
the vector v. This series of inner-products can be calculated using a loop
that runs over all the lines in the matrix. In the numpy-array indexing
the 𝑖 -th row of the matrix can be accessed by A[i,:]and the inner product
of this line with the vector v can be calculated by:
1 [Link](A[i, :], v)

Therefore, the main part of the calculation of this matrix-vector product


is the loop:
1 for i in range(nLine):
2 u[i] = [Link](A[i, :], v)

where the elements of the vector u are calculated one by one. For each
element, the inner product of the 𝑖 -th row of A with v. This for-loop can
be included in a Python program as shown below.
8 Vectors and Matrices 59

1 import numpy as np
2

3 def mat_vect(A, v):


4

5 nLine = [Link][0]
6 nCol = [Link][1]
7 sizeV = [Link][0]
8

10 if nCol != sizeV:
11 print("Error : Size inconsistent")
12 return
13

14 u = [Link](nLine)
15 for iL in range(nLine):
16 u[iL] = [Link](A[iL, :], v)
17 return u
18

19 A = [Link](4, 5)
20 v = [Link](5)
21

22 u = mat_vect(A, v)
23 ucheck = [Link](A, v)
24

25 print([Link](u-ucheck))

This program contains a function that performs the multiplication of a


matrix with a vector and return the resulting vector. Note that the function
first checks whether the size of the matrix and vector are compatible
or not using the command ’shape’. This is a very typical use of the
if-statement in functions that perform numerical operations since the
input of the function should be correct, otherwise it makes no sense to
actually do the required calculations.

8.7 Determinant: Recursion

The determinant of a square matrix is an important quantity that plays a


role in matrix transformations and linear algebra. For a 2x2 matrix it is
easily calculated as:

𝑎 𝑏
Det(M) = | M | = = 𝑎·𝑑−𝑐·𝑏 (8.10)
𝑐 𝑑

For square matrices of higher order than two it become somewhat


more complex. Determinants of higher order matrices can in general be
calculated by decomposing them in co-factors and ’minor’ matrices. For
a 3x3 matrix this results in:

𝑎 𝑏 𝑐
𝑒 𝑓 𝑑 𝑓 𝑑 𝑒
|M| = 𝑑 𝑒 𝑓 =𝑎· −𝑏· +𝑐· (8.11)
ℎ 𝑖 𝑔 𝑖 𝑔 ℎ
𝑔 ℎ 𝑖

This decomposition is done by taking one row or column of a matrix


and then making a series of combinations of the co-factor multiplied
by the determinant of the corresponding minor matrix. In the example
above we take the co-factors from the first line and the minor matrix is
constructed by taking out the line and the column of the co-factor, i.e. for
8 Vectors and Matrices 60

the first element ( 𝑎 ) the minor matrix is the right-bottom 2x2 matrix that
remains, etc. Not that the sign in front of the co-factor alternates. Now
the calculation of the determinant of this 3x3 matrix has been reduced to
the calculation of determinants of three 2x2 matrices, which was defined
in Eq.8.10. The same approach can be used for matrices of higher order,
decomposing them step by step until they are given only in the form of
2x2 matrices. The general approach for the decomposition of matrices of
any order is given by the Laplace-expansion, in this case along the 𝑖 -th
row:

𝑛
(−1)𝑖+𝑗 𝑚 𝑖𝑗 𝐴 𝑖,𝑗
X
|M| = (8.12)
𝑖=1

The 𝑚 𝑖 𝑗 are element of the matrix M that in this case run along one of the
lines of the matrix. The matrix A is the (𝑛 − 1) x (𝑛 − 1) ’minor’ matrix
that does not contain the line and column of the current element 𝑚 𝑖 𝑗 .
Special attention should be paid to the factor (−1)𝑖+𝑗 that defines the sign
of the different contributions. As seen in the evaluation for the 3x3 matrix
the sign alternates along the line that is used for the decomposition. Note
that we use the common convention in mathematics here, which implies
that the indexing of matrices starts at 1. This is different from the numpy
indexing of arrays that starts at 0!
The evaluation of determinants of higher order square matrices quickly
becomes tedious and prone to small mistakes. However, it is done by a
systematic step-wise approach that is very suitable to do in a Python
code. In fact, it is an excellent example that can be solved by a so-called
recursive methods. In a recursive method we use a function that actually
calls itself, generally multiple times. This is exactly what we do in the
code below: it decomposes the matrix for which we want to calculate
the determinant into its minors and then we call the same function to
calculate the determinant of those minors! This sequence continues until
we arrive at a 2x2 matrix for which the determinant is easily calculated.
1 import numpy as np
2

3 def my_determinant(M):
4 rows = [Link][0]
5 columns = [Link][1]
6 # Start by checking whether the matrix is square
7 if rows != columns:
8 print(’this is not a square matrix’)
9 return(0)
10

11 # If it is a two-by-two matrix, return the simple cross-solution


12 if rows == 2:
13 return (M[0,0]*M[1,1] - M[1,0]*M[0,1])
14

15 # In all other cases: create submatrix, and add the terms up with the
16 # appropriate prefactor
17 # the function my_determinant is called recursively until we end up at the
18 # 2x2 matrix
19

20 else:
21 sum = 0
22 for i in range(columns):
23 submatrix = [Link]((rows-1, columns-1))
24 submatrix[0:, :i] = M[1:, :i]
25 submatrix[0:, i:] = M[1:, i+1:]
8 Vectors and Matrices 61

26 sum += M[0, i] * (-1)**i * my_determinant(submatrix)


27 return (sum)
28

29

30

31 matrix = [Link]([[2, -3, 1], [1, 4, 5], [2, 0, -1]])


32 print(matrix)
33

34 determinant = my_determinant(matrix)
35 print(’My determinant:’, determinant)
36

37 det = [Link](matrix)
38 print(’linalg:’, det)

In this code we define the function my_determinant() and check in the


first few lines whether it is actually a square matrix. Subsequently, the
code check whether it is a 2x2 matrix and if this is the case it simply
returns a numerical value for the determinant. The remainder of the
code deals with higher order matrices in the ’else’ part of the if-else
construction. The loop runs over all the elements of the first line of
the matrix and for each of these elements the minor matrix (called
’submatrix’) is constructed in two steps, first all the elements to the left
of the current index are added (if there are any) and then those to the
right. After that, on line 26 we add up the different contributions of the
minor matrices in a step-wise way with the appropriate co-factor and
sign. Note that the sign takes the indexing in Python (starting at 0) into
account correctly. Line 26 is very general and it works for matrices of any
order because the determinant that is to be calculated is simply obtained
by a call to the function itself! This sequence is repeated until we end up
at 2x2 matrices. When the size of the matrix becomes larger, the function
my_determinant will be calling itself many times, which can easily be
checked by adding a print() statement inside the function.
This function to calculate the determinant of a matrix of any order shows
the power of recursive functions and such functions can be very elegant
solutions to many problems in mathematics.

8.8 Exercise: Matrix Multiplication

To practice the handling of matrices in numpy we will, in this exercise,


create a Python code called ’[Link]’. We first going to write a function
that calculates the product of two matrices. This function will take two
matrices as input arguments and returns the product of these matrices as
a result. The equation for the matrix product is given in eq. 8.7. To obtain
the number of lines and columns of a matrix use the function:
1 [Link][0] # number of rows
2 [Link][1] # number of columns

The implementation of the multiplication of a matrix with a vector is


discussed above. The matrix-matrix multiplication follows a similar
pattern. The main part of the code consists of two nested loops as in the
pseudo-code below.
1 def matmul(A,B):
2 .....
3 for iR in range(nRow):
4 for iC in range(nCol):
8 Vectors and Matrices 62

6 C[iR,iC] = ......
7

8 return C

To test the function, generate two random matrices with the statements:
1 N, M, P = 100, 50, 150
2 A = [Link](N, M)
3 B = [Link](M, P)

This leads to the creation of two random matrices. You can check that
your function works correctly by comparing the result it gives to the
result provided by the standard numpy function to multiply matrices:
C = [Link](A,B). In addition you can time the performance of the function
with two nested loops, by using the function clock() from the module
time.
1 import time
2 start = [Link]()
3 # put the instruction you want here
4 end = [Link]()
5 print(’Calculation done in %f sec’ % (end-start))

Calculate the time required to calculation the multiplication of square


matrices of size 𝑁 , 𝑀, 𝑃 = (10 , 15 , 30), (100 , 150 , 300), (500 , 600 , 1000)
filled with random numbers with your function. Compare this time with
the time required to perform the same operation using the function
𝑛𝑝.𝑑𝑜𝑡().
Systems of linear equations 9
Many problems in science and engineering involve systems of coupled 9.1 Example: Distillation Column63
linear equations. Such systems of linear equations have multiple unknown 9.2 Solution using numpy . . . . 64
variables ( 𝑥 1 , 𝑥 2 , ...𝑥 𝑁 ) that are found in the same number of linear 9.3 The Gauss-Jordan algorithm 65
9.4 Gauss-Jordan implemented 67
equations. Examples of such systems include electrical circuits, systems
Computational cost . . . . . . 69
of masses connected by springs and distillations columns with mass
9.5 The Jacobi iterative method . 71
balances. If the system is truly linear, there will be no squares etc. of the
9.6 Exercise: Two-stage distillation
variables ( 𝑥 2𝑘 ) or cross-products of them ( 𝑥 𝑖 𝑥 𝑗 ) and there are no non-linear column . . . . . . . . . . . . . . . 72
terms such as sin(𝑥 𝑖 ). In such cases we can write the system of equations 9.7 Exercise: Jacobi Iteration . . 73
as: 9.8 Extra exercises . . . . . . . . . 74
Lake contamination . . . . . 74
The Gauss-Seidel method . 74
𝑎 1 𝑥1 + 𝑎2 𝑥 2 + . . . 𝑎 𝑛 𝑥 𝑛 = 𝑓1 (9.1)
𝑏1 𝑥1 + 𝑏2 𝑥 2 + . . . 𝑛 𝑛 𝑥 𝑛 = 𝑓2 (9.2)
..
. (9.3)
𝑧 1 𝑥 1 + 𝑧 2 𝑥 2 + . . . 𝑧 𝑛 𝑥 𝑛 = 𝑓𝑛 (9.4)

In this system of equations, 𝑎 𝑖 , 𝑏 𝑖 ... and 𝑓𝑖 are the multiplication vari-


ables and the results-vector, respectively. This represents a system of
𝑛 -equations with 𝑛 unknowns (the values 𝑥1 , 𝑥2 , ...𝑥 𝑁 ). Such a system of
linear equations can be rewritten in the form of a matrix multiplication:

𝑎1 𝑎2 ... 𝑎𝑛 𝑥1 𝑓1
­𝑏 1 𝑏2 ... 𝑏 𝑛 ® ­ 𝑥 2 ® ­ 𝑓2 ®
© ª © ª © ª
­. .. .. ®® · ­­ .. ®® = ­­ .. ®® (9.5)
­. .
­. .® ­ . ® ­.®
«𝑧1 𝑧2 ... 𝑧 𝑛 ¬ « 𝑥 𝑛 ¬ « 𝑓𝑛 ¬

The first matrix contains all the constant multiplication factors. This
matrix is multiplied with the middle vector that holds the unknowns
and the right sides is a vector that holds the solution of each equation.
Such a problem is often written as A · x = f. Given a set of constants 𝑎 𝑖 ,
𝑏 𝑖 ... and 𝑓𝑖 , the goal is to find the values of all the different unknowns.

9.1 Example: Distillation Column

A typical chemical example of a system of linear equations is the mass bal-


ance of a distillation column. Distillations are used to separate mixtures
of compounds based on their different boiling temperatures. The mathe-
matical modeling of the distillation process can be very complicated and
involve non-linear equation. In this chapter we will explore a simplified
model that is illustrated in Fig. 9.1. In this example, the input mixture that
flows into the column consists of three compounds: 30 kg/s of methane
(M), 25 kg/s of ethane (E) and 10 kg/s of propane (P). The input mixture
9 Systems of linear equations 64

is separated into three output flows: the overhead-stream that is rich


in methane (90%), the middle stream that is rich in ethane (50%), and
the bottom-flow that is rich in propane (70%). Fig. 9.1 contains some
additional information about the mass fractions in each flow, indicated
as 𝑥 𝑖 and the total mass-flow rates of each output-stream are given in
kg/s and indicated by 𝑚 𝑖 .
Ideal distillation columns are operated in the steady state each every
kg of a given compound that enters the column must be matched by
the same amount of the same compound leaving the column. Using
the principle of conservation-of-mass 𝑚 𝑖𝑛 = 𝑚 𝑜𝑢𝑡 , we can equate the
total mass entering the column with the total mass leaving it. Such a
mass-balance for methane results in:

30 = 0.9𝑚1 + 0.3𝑚2 + 0.1𝑚3 (9.6)

Similar linear equations can be written down for ethane and propane
and the overall system of linear equations can be written in a matrix form
A · x = f. This results in the matrix equation:

0.9 0.3 0.1 𝑚1 30.0


­0.1 0.5 0.2® · ­𝑚2 ® = ­25.0® (9.7)
© ª © ª © ª

«0.0 0.2 0.7¬ «𝑚3 ¬ «10.0¬

Figure 9.1: Simplified model of a distilla-


tion column. A flow composed of methane,
ethane and propane enters the column.
The three compounds are then separated
in 3 streams each mainly composed of one
of the compounds.

9.2 Solution using numpy

The system of equations 9.7 can be solved directly using a numpy function.
The numpy has a library (or sub-module) that contains many common
linear algebra functions. The most versatile sub-module is [Link]
that can be imported in a Python program in the usual way. In the code
below we rename the [Link] sub-module to ’nl’ while loading to
9 Systems of linear equations 65

reduce the amount of typing we have to do. We can then use function
[Link]() to solve systems of linear equations:
1 import numpy as np
2 import [Link] as nl
3

4 # create the matrix A


5 A = np. array ([[0.9, 0.3, 0.1],
6 [0.1, 0.5, 0.2],
7 [0.0, 0.2, 0.7]])
8

9 # create the right-hand side


10 f = [Link]([30.0, 25.0, 10.0])
11

12 # solve the linear system with [Link]()


13 x = [Link](A,f)
14

15 # print the solution


16 print(x)
17

18 # check the solution


19 check_sol = [Link](A,x)
20 print(check_sol)
21 print(f)

The output of this program is given by:

𝑚1 = 17.88 𝑘 𝑔/𝑠 𝑚2 = 45.96 𝑘 𝑔/𝑠 𝑚3 = 1.15 𝑘 𝑔/𝑠 (9.8)

Using [Link]() function we have directly solved this system of linear


equations and determined the total mass flow in each of the streams of
the column. In the code, the validity of the solution is verified by directly
computing the product A · x and comparing this solution with the vector
𝑓 . The method implemented by [Link]() is said to be ’exact’, which
means that the error is determined by the hardware of the computer. In
other words the first 8-12 digits of the results will be exact, depending
on the computer and the software that is used. Methods that calculate
such exact solutions are called 𝑑𝑖𝑟𝑒 𝑐𝑡 . These methods are suitable for
relatively small linear systems (up to ∼ 1000 equations). For larger systems,
𝑖𝑡𝑒𝑟 𝑎𝑡𝑖𝑣𝑒 solutions are usually preferred. Such iterative algorithms can
be used to calculate the solution up to a previously determined accuracy
and are suitable for very large matrices.

9.3 The Gauss-Jordan algorithm

The most well-known 𝑑𝑖𝑟𝑒 𝑐𝑡 solution for solving sets of linear equations
is the Gauss-Jordan algorithm, also referred to as Gaussian elimination.
This approach relies on the use of an augmented or extended matrix,
made up of the matrix 𝐴 but to which the column vector 𝑓 is added
on the right side. The Gauss-Jordan algorithm transforms this matrix
in two steps to a matrix where all the elements of the original matrix
𝐴 are zero, except for the diagonals that become 1 after the procedure.
After this is achieved, the last column will contain the solution vector x.
This sequential transformation of the augmented matrix is schematically
indicated by:
9 Systems of linear equations 66

𝑎1 𝑎2 ... 𝑎𝑛 𝑓1 𝑎0 𝑎20 ... 𝑎 0𝑛 𝑓10


© 1
𝑏1 𝑏2 ... 𝑏𝑛 𝑓2 ­ 0 𝑏 20 ... 𝑏 0𝑛 𝑓20
© ª ª
­ ® ®
­ .. .. .. .. ® −→ ­ .
­ . .. .. .. ®
­
­ . . . .
®
® ­ . . . .
®
®
« 𝑧1 𝑧2 ... 𝑧𝑛 𝑓𝑛 ¬ « 0 0 ... 𝑧 0𝑛 𝑓𝑛0 ¬
1 0 ... 0 𝑥1
0 1 ... 0 𝑥2
© ª
­ ®
−→ ­­ .. .. .. .. ® (9.9)
­ . . . .
®
®
« 0 0 ... 1 𝑥𝑛 ¬

As seen in the above equation the algorithm takes place in two steps. First
we transform 𝐴 into an upper-triangular matrix, i.e. a matrix containing
non-zero elements only on and above its diagonal. In the second step,
we back-substitute the solution of the equation to obtain the solution
of the linear system. To perform these matrix transformations, only 3
operations, that do not change the solutions of the linear system are
allowed. These operations are:

I The interchange of two lines.


I Multiplication/division of a line by a non-zero number.
I Addition/subtraction of a multiple of another line.

This approach can be readily applied to the linear system in Eq. 9.7 for
the distillation column. The augmented matrix of this system reads:

0.9 0.3 0.1 30


­ 0.1 0.5 0.2 25 ® (9.10)
© ª

« 0.0 0.2 0.7 10 ¬

We will refer to the first second and third row (line) of this matrix by 𝐿1 ,
𝐿2 and 𝐿3 , respectively. Our first goal it to eliminate the term 𝑏 1 = 0.1 in
𝐿2 . To do so, we perform the following operation: 𝐿2 = 𝐿2 − 00..91 𝐿1 . The
result of this operation is shown in the second matrix of Eq. 9.11, where
𝑏1 = 0. We can then eliminate the term 𝑐 2 = 0.2 from this new matrix by
performing 𝐿3 = 𝐿3 − 00..462
𝐿2 , leading to the last matrix in Eq. 9.11. Note
that 𝑐 1 happened to be already equal to 0, otherwise we should have also
made that zero by adding a certain factor time 𝐿1 , in addition to adding
a factor time 𝐿2 .

0.9 0.3 0.1 30 0.90 0.30 0.10 30.00


­ 0.1 0.5 0.2 25 ® → ­ 0.00 0.46 0.18 21.66 ®
© ª © ª

« 0.0 0.2 0.7 10 ¬ « 0.00 0.20 0.70 10.00 ¬


0.90 0.30 0.10 30.00
→ ­ 0.00 0.46 0.18 21.66 ® (9.11)
© ª

« 0.00 0.00 0.61 0.71 ¬

In Eq. 9.11 all floating point numbers have here been truncated to two
decimals to save space. Using the procedure above we have transformed
the matrix 𝐴 into an upper triangular matrix. We can now use the
so-called back-substitution approach to arrive at an ’identity matrix’
9 Systems of linear equations 67

containing only ones on the diagonal and zeros everywhere else. This
can be done systematically by starting from the bottom. We therefore
start with 𝐿3 = 𝐿3/0.61, leading directly to 𝑚3 = 1.15 as seen in the
first matrix in Eq. 9.12. We can use that new matrix to solve for 𝑚2 by
performing 𝐿2 = 1/0.46(𝐿2 − 0.18 ∗ 𝐿3 ). Finally we solve for 𝑚1 by doing:
𝐿1 = 1/0.90(𝐿1 − 0.30𝐿2 − 0.10𝐿3 ).

0.90 0.30 0.10 30.00 0.90 0.30 0.10 30.00


­ 0.00 0.46 0.18 21.66 ® → ­ 0.00 1.00 0.0 45.96 ®
© ª © ª

« 0.00 0.00 1.00 1.15 ¬ « 0.00 0.00 1.00 1.15 ¬


1.00 0.00 0.00 17.88
→ ­ 0.00 1.00 0.0 45.96 ®
© ª

« 0.00 0.00 1.00 1.15 ¬

As expected, the last column now contains the solution of this linear
system. Although the method that is used by [Link]() is a bit more
complex than the Gauss-Jordan algorithm shown, here it relies on the
same general principle and of course leads to exactly the same solution.

9.4 Gauss-Jordan implemented

As explained above, the Gauss-Jordan algorithm consists of two main


operations:

1. The transformation of the matrix to an upper diagonal form.


2. The backward substitution to make it a diagonal matrix with ones
on the diagonal.

These two stages of the solution are somewhat different and therefore the
implementation requires that we write two functions that each perform
one of these two operations.
A simple implementation of the first stage of the Gauss-Jordan algorithm
(i.e. obtaining an upper diagonal matrix) is shown in the code below.
1 import numpy as np
2

3 ########################################################
4 # Function to decompose the matrix
5 # in an upper triangular form
6 ########################################################
7

8 def LU(A,f):
9

10 # get the size of the system


11 n = len(f)
12

13 # check the size


14 if ([Link][0] != n) or ([Link][1] != n):
15 print(" \t Inconsistent size in LU decomposition")
16 info = 0
17 return M, info
18

19 # create the augmented matrix


20 M = [Link]((n, n+1))
21 M[:, :-1] = A
22 M[:, -1] = f
9 Systems of linear equations 68

23

24 # loop through all the colum


25 # to get rid of the lower part
26 for iC in range(n-1):
27

28 # for each column loop over all the lines


29 # that are below the diagonal
30 # to set to 0 their elements
31 for iL in range(iC+1, n):
32

33 # check if the diagonal element


34 # is null
35 if M[iC, iC] == 0:
36

37 print(" \t Zero on the diagonal, LU failed")


38 info = 0
39 return M, info
40

41 # eliminate the element


42 M[iL, :] = M[iL, :] - M[iL, iC]/M[iC, iC] * M[iC, :]
43

44 # if we succed we return info = 1 and the upper augmetned matrix


45 info = 1
46 return M, info

The function LU() receives a matrix A and the right hand side vector ’f’ as
arguments. In the first few lines of code the dimensions of the matrix are
checked. Subsequently, a loop over all the columns is performed (with
exception of the last one) to eliminate all the terms below the diagonal
using the method discussed above. Note that we also make sure that
the diagonal element M[iC, iC] is not null as it would lead to an error
when we try to divide by M[iC, iC]. A more advanced version of the
Gauss-Jordan algorithm is required to treat such cases (interchanging
the different lines would solve this problem). Once all the terms are
eliminated we return the modified augmented or extended matrix M as
well as an integer value called ’info’, that equals 1 if the elimination was
successful and 0 otherwise.
The implementation of the second stage, the backward substitution, is
listed in the code below.
1

2 ########################################################
3 # Function to backsubstitute the results
4 # and get the final solution
5 ########################################################
6 def BS(M):
7

8 # get the size of the matrix


9 n = [Link][0]
10

11 # loop over all the lines


12 # starting by the end
13 for iL in range(n-1, -1, -1):
14

15 # check if we have diagonal elements on the diagonal


16 if M[iL, iL] == 0:
17 print(" \t Zero on the diagonal, LU failed")
18 info = 0
19 return M, info
20

21

22 # divide the line by the diagonal element of M


23 M[iL, :] /= M[iL, iL]
24
9 Systems of linear equations 69

25 # loop over all the lines that are above this one
26 for iLL in range(iL-1, -1, -1):
27 M[iLL, :] -= M[iLL, iL]*M[iL, :]
28

29 info = 1
30 return M, info
31

32

33 ########################################################
34 ########################################################

This function takes the upper diagonal augmented matrix, M, created


by the LU function as its input argument. The function starts from the
last row of the matrix to eliminate the upper diagonal part until a form
similar to last one in eq. 9.9 is obtained. The function returns the diagonal
form of the matrix M where the last column holds the solution of the
system.
The combination of these two functions performs a very common opera-
tion and we may wish to use them quite often in all sorts of programs.
To make it easier to use these functions in the future we can write our
own Python module. This is very easy as it is simply done by putting
the two functions presented above in the same file that we will call
’gauss_jordan.py’. This file can now be used as a module in any other file
located in the same folder. To use this module we just have to import it
(exactly as Numpy etc ...). If we write a file called ’[Link]’ we can use
the function written in ’gauss_jordan.py’ as illustrated below:
1 import numpy as np
2 from gauss_jordan import *
3

5 # create the matrix A


6 A = [Link](
7 [[0.9, 0.3, 0.1],
8 [0.1, 0.5, 0.2],
9 [0.0, 0.2, 0.7]])
10

11 # create the right-hand side


12 f = [Link]([30.0, 25.0, 10.0])
13

14 # decompose the matrix


15 M, info = LU(A, f)
16

17 # backsubstitute
18 M, info = BS(M)
19

20 print(M[:, -1])

Computational cost

The Gaussian elimination and backward substitution are relatively ex-


pensive computationally. During the Gaussian elimination we must loop
over over the the 𝑛 columns and 𝑛 rows of the matrix to eliminate all
the terms. For each entry we must calculate 𝑛 multiplications leading to
an overall cost of 𝑛 3 operations. To compare the time required to solve
linear systems of increasing size with our method compared to the time
required by the [Link]() method we can use the code below.
1 import numpy as np
2 import [Link] as npla
9 Systems of linear equations 70

3 import gauss_jordan as gj
4 import time
5 import [Link] as plt
6

7 # size of the matrix to be calculated


8 SIZE = [10, 50, 100, 500, 1000]
9

10 # create list
11 cpu_time_numpy = []
12 cpu_time_mycode = []
13

14 # loop over the size


15 for size in SIZE:
16

17 # create the system


18 A = [Link](size,size)
19 f = [Link](size)
20

21 # nummpy
22 t0 = [Link]()
23 [Link](A, f)
24 cpu_time_numpy.append([Link]()-t0)
25

26 # mycode
27 t0 = [Link]()
28 M,info = [Link](A, f)
29 M,info = [Link](M)
30 cpu_time_mycode.append([Link]()-t0)
31

32 # plot the results


33 [Link](SIZE, cpu_time_numpy, ’o-’, linewidth=3, color=’black’, label
=’Numpy’)
34 [Link](SIZE, cpu_time_mycode, ’o-’, linewidth=4, color=’#005FFF’,
label=’My code’)
35 [Link](’Size of the system’)
36 [Link](’Computation Time’)
37 [Link]([0, 1250])
38 [Link](loc=2)
39 [Link]()

The results of this code are shown in Fig. 9.2. As you can see finding the
solution of relatively big linear system is rather inexpensive. However,
the graph also shows clearly that [Link]()is considerably more
efficient than our simple Gauss-Jordan implementation. Numpy uses a
different algorithm to solve linear systems that is more complicated, but
the Gauss-Jordan algorithm is a very good illustration.

Figure 9.2: Computational time required


to solve linear systems of increasing size
with the Numpy method and with our
own algorithm.
9 Systems of linear equations 71

9.5 The Jacobi iterative method

The direct Gauss-Jordan approach to solve a system of linear equations


is the same as it would be done by hand on paper. This is generally
straightforward to implement and give a very systematic approach to
find the solution. However, such direct methods are usually not the
most efficient way to find a solution, and they can suffer from numerical
instabilities. An approach that is often more efficient is the use of iterative
methods. The general approach in an iterative method is to improve the
solution step-by-step, where in each iteration the solution becomes closer
to the exact solution, until a certain threshold is [Link] requires
an estimate of the error in the values that are optimized at each stage.
In the case of solving our system of linear equations this error is the
so-called residual res = f − 𝐴 · x0 , where x0 is a vector containing the
values that are being optimized during the algorithm to minimize res.
The algorithm continues until the residual is small enough, i.e. when its
norm is smaller than a predefined threshold-value. To illustrate iterative
approaches we discuss the Jacobi method and apply it again to the linear
system for the distillation column in Eq. 9.7. Iterative methods generally
start from a certain initial guess of the solution and this starting point is
used to improve on this guess in subsequent steps. The Jacobi therefore
(𝑘) (𝑘) (𝑘)
results in a series of solution vectors m(𝑘) = [𝑚1 , 𝑚2 , 𝑚3 ]. The case
𝑘 = 0 corresponds to the initial guess from which the second iteration
i.e. m 𝑘=1 is calculated, and so on. In a slightly changed notation we can
write our system of linear equations as:

𝑎11 𝑎12 𝑎13 𝑚1 𝑓1


­ 𝑎21 𝑎22 𝑎23 ® · ­𝑚2 ® = ­ 𝑓2 ® (9.12)
© ª © ª © ª

« 𝑎31 𝑎32 𝑎33 ¬ «𝑚3 ¬ « 𝑓3 ¬

From these equations we can write for each of the values of 𝑚 in terms of
all the other values in the vector m. Of course at the start of the algorithm
we do not know any of the values in this vector yet, but we will guess
them at first. Therefore, we can now write these equations for each 𝑚 𝑖 in
terms of the values of the values of all other values in m as they were in
the previous cycle (or at the starting point in terms of the initial guess:

!
(𝑘+1) 1 X (𝑘)
𝑚𝑖 = 𝑓𝑖 − 𝑎 𝑖𝑗 𝑚 𝑗 (9.13)
𝑎 𝑖𝑖 𝑗≠𝑖

As an example we apply this iterative approach to our problem of the


distillation column. For the first iteration we assume (guess) that the
(0) (0) (0)
values of the mass flow rates are 𝑚1 = 𝑚2 = 𝑚3 = 20. We will see,
that even with this bad guess the Jacobi method converges rapidly to a
correct solution. The first iteration of the method leads to:
9 Systems of linear equations 72

1
(1)
𝑚1 = (30.0 − 0.3 ∗ 20 − 0.1 ∗ 20) = 24.44 (9.14)
0.9
(1) 1
𝑚2 = (25.0 − 0.1 ∗ 20 − 0.2 ∗ 20) = 38.00 (9.15)
0.5
(1) 1
𝑚3 = (10.0 − 0.0 ∗ 20 − 0.2 ∗ 20) = 8.57 (9.16)
0.7

If we compare the the new value of the masses we can see that they are
collectively closer to the exact solution than our initial guess. We then
take these updated values and insert them as the new initial guess and
perform the same operation:

(2) 1
𝑚1 = (30.0 − 0.3 ∗ 38.00 − 0.1 ∗ 8.57) = 19.71 (9.17)
0.9
(2) 1
𝑚2 = (25.0 − 0.1 ∗ 24.44 − 0.2 ∗ 8.57) = 41.68 (9.18)
0.5
(2) 1
𝑚3 = (10.0 − 0.0 ∗ 24.44 − 0.2 ∗ 38.00) = 3.43 (9.19)
0.7

which gives updated values that are even closer to the exact solution. If
we keep iterating that way the solution will converge quickly to values
that are really close to the exact solution that we obtained from the
Gauss-Jordan method. In the exercise below we will write our own code
to implement the Jacobi algorithm and compare its performance to exact
solutions for different linear systems.

9.6 Exercise: Two-stage distillation column

The example of the distillation column used in this chapter is rather


simple. In this exercise we will consider the set of coupled distillation
columns shown in Fig. 9.3. We will write a Python code ’[Link]’
that computes the output mass flow rates 𝑚1 , 𝑚2 , 𝑚3 of the distillation
column shown in this figure. In this two-stage column three compound
contained in the input flow are separated in two-stages.
Hint: to solve this system of equations, write a system of linear equation
with 1: the balance of the total mass flow rate, 2: the mass flow rate of
methane and 3: the mass flow rate of ethane.
9 Systems of linear equations 73

Figure 9.3: Simplified model of a two-


stage distillation column. A flow com-
posed of methane, ethane and propane
enters the column. The three compounds
are then separated in 2 streams in the first
column and the bottom steam enters an
additional column that perform another
distillation.

9.7 Exercise: Jacobi Iteration

In this exercise we will write a python script ’[Link]’ that implements


the Jacobi algorithm for solving a system of linear equations and compare
its performance to the function [Link]().
We will write a function that takes 3 input arguments: the matrix 𝐴, the
right-hand vector f and an initial-guess vector for the solution of the
linear system x0 . The function will return the final value of the solution.
In the function we will first determine the residual of the initial guess. A
𝑤 ℎ𝑖𝑙𝑒 loop should be used iterate the solution vector m(𝑘) and update
the residual res = f − 𝐴 · x 𝑘 . The 𝑤 ℎ𝑖𝑙𝑒 loop will stop only if the norm of
the residual is smaller than a threshold that is defined before. Typically
this threshold should be lower than 10−5 to obtain accurate solutions.
The norm of a vector can be obtained using the function:
1 [Link](X)

To ensure that the while-loop stops even if the solution is not converging,
which can sometimes happen, we typically impose a maximum number
of iterations 𝐼𝑇𝐸𝑅𝑀𝐴𝑋 and combine several conditions to stop the
𝑤 ℎ𝑖𝑙𝑒 loop:
1 while (res>tol) and (niter<ITER_MAX):
2 # block of instruction
3 # ...
4 res = .....
5 niter += 1

The loop will continue as long as (𝑟𝑒 𝑠 > 𝑡𝑜𝑙) and 𝑛𝑖𝑡𝑒𝑟 < 𝐼𝑇𝐸𝑅𝑀𝐴𝑋).
Do not forget to update 𝑟𝑒 𝑠 and increment 𝑛𝑖𝑡𝑒𝑟 within the loop otherwise
it will never stop! Test your Jacobi function to solve the linear system in
Eq. 9.7.
9 Systems of linear equations 74

9.8 Extra exercises

Lake contamination

In this exercise we will estimate the concentration of polychlorinated


bipheny (PCB) in the Great Lakes between the U.S.A. and Canada. The
diagram shown in Fig. 9.4 shows the relevant flows between the lakes.
The concentration of PCB in each lake is constant and therefore all the
PCB that enters this ’lake-system’ must also leave. Use this information
to write a Python script ’[Link]’ that calculates the balance equation of
each lake. For example the balance of Lake Superior reads:

𝑘𝑔
180 = 𝑄 𝑆𝐻 𝐶𝑆 (9.20)
𝑦𝑟

Once you have written these equations rearrange them so that only term
without any unknown appears on the right-hand side. You can then solve
this system of equations with one of the methods seen in this chapter
and calculate the concentration of PCB contained in each lake.
An environmental organization is considering to build a bypass that
would go directly from lake Michigan to Lake Ontario with a flow rate of
20 𝑘𝑚 3 /𝑦𝑟 in order to reduce the concentration of PCB in lake Michigan.
Examine the effect of this bypass by rewriting the set of linear.

Figure 9.4: PCB contamination of the


Great Lakes. On this diagram you can
see all the input/output mass flow rates
in the different lakes. The output of each
lake is given by the product of its concen-
tration 𝐶 𝑖 and a constant 𝑄 𝑖𝑗 whose value
are given in the table.

The Gauss-Seidel method

An improvement of the Jacobi method is provided by the so-called Gauss-


Seidel method. In this approach the 𝑘 + 1-th iteration of the solutions is
obtained via the equation:

!
(𝑘+1) 1 X (𝑘+1)
X (𝑘)
𝑚𝑖 = 𝑓𝑖 − 𝑎 𝑖𝑗 𝑚 𝑗 − 𝑎 𝑖𝑗 𝑚 𝑗 (9.21)
𝑎 𝑖𝑖 𝑗<𝑖 𝑗>𝑖

The difference with the Jacobi method is that we always use the most
recent information we have of the solutions, even if they have not been
9 Systems of linear equations 75

(𝑘+1)
calculated for all values yet. For example, the calculation of 𝑚2 in
(𝑘)
the Jacobi method only depends on 𝑚1,3 . However, in the Gauss-Seidel
(𝑘+1) (𝑘+1) (𝑘)
method 𝑚2 is calculated from 𝑚1 (that we compute from 𝑚2,3 ) and
(𝑘)
𝑚3 . Implement your own version of the Gauss-Seidel algorithm and
compare its performance to the Jacobi-method. Note: the implementation
largely follows the Jacobi approach, but taking into account that the
solution-values from the current cycle are used where available.
Integration, Derivatives and
Non-linear systems
Numerical integration 10
The evaluation of the definite integral (or quadrature) of a function, as 10.1 Equal intervals . . . . . . . . . 77
shown in EQ. 10.1 is an operation that occurs in many problems, and Rectangle approximation . . 77
is a central issue in the integration of differential equations. Analytical Trapezoid approximation . . 79
Simpson’s rule . . . . . . . . . 80
solutions for the evaluation of such integral are sometimes available, but
Newton-Cotes formulas . . . . 81
in many cases they are difficult to derive, and often impossible. A specific
10.2 Gaussian quadrature . . . . . . 81
case where analytic integration is impossible is in cases where the data
10.3 Multiple Integrals . . . . . . 82
is obtained from experiments, i.e. only a series of tabulated values. The 10.4 Exercises . . . . . . . . . . . . . 83
numerical integration of functions is straightforward and in this chapter Composite methods . . . . . 83
we will discuss some common, efficient approaches for integration. Simpson’s rule . . . . . . . . . 85

∫ 𝑏
𝐼= 𝑓 (𝑥)𝑑𝑥 (10.1)
𝑎
In general, integrals can be calculated by summing up the areas of a
series of geometric shapes, such as rectangles or trapezes. The summation
is done in intervals and in principle an arbitrary accuracy can be reached
by making the size of the intervals smaller and smaller. In this chapter
we will consider two of such ’composite’ methods. We will first consider
approaches that use evenly spaced intervals, based on the so-called
Newton-Cotes formulas. Subsequently, we will consider more efficient
methods where the sizes of the intervals vary and are optimized to
give the best accuracy. Such methods, of which Gaussian quadrature is
an example, are more difficult to implement but several functions are
available in numpy to perform such integrations.

10.1 Equal intervals

The calculation of a definite integral in 1D is identical to calculating the


area below the curve. The most straightforward way to approximate the
definite integral of a given function numerically is by dividing the range
over which the integral is to be calculated into 𝑛 equal intervals of length
ℎ . Progressively decreasing the value of ℎ can lead to the approximation
of any integrals with arbitrary accuracy.

Rectangle approximation

The simplest method to compute a definite integral is the rectangle


approximation shown in Fig. 10.1. In this approximation the X-axis is
divided in a series of equally spaced intervals. In a first approximation
the height of the rectangle filling the interval between 𝑥 𝑖 and 𝑥 𝑖+1 is set
10 Numerical integration 78

to 𝑓 (𝑥 𝑖 ), i.e. the value of the function at the beginning of the domain. In


this case the approximated integral over this interval is given by:

∫ 𝑏 𝑁−
X1
𝑓 (𝑥)𝑑𝑥 ≈ ℎ 𝑓 (𝑥 𝑗 ) (10.2)
𝑎 𝑗=1

Figure 10.1: Illustration of simplest com-


posite method for the numerical integra-
tion of curves. The interval is divided in
small rectangles that spans the integra-
tion interval. The integral is obtained by
summing up the area of all rectangles.

This rectangle approximation to calculate the value of the integral in


Eq. 10.2 can easily be implemented in a Python code. In the code listing
below a function ’int_rect()’ is implemented that takes the name of a
function, the bounds of the interval ( 𝑎 and 𝑏 ) and the number of intervals
as arguments.
1 def int_rect(func, a, b, N):
2 res = 0
3 xpts = [Link](a, b, N)
4 h = xpts[1]-xpts[0]
5 for x in xpts[:-1]:
6 res += h*func(x)
7 return res

Three of the arguments of this function are optional: the ones that defines
the integration interval and the number of points in which the interval
in divided. For all the points in the interval, except for the last one, the
function adds up the area of all the rectangle, ℎ × 𝑓 (𝑥) and stores them in
the variable the variable 𝑟𝑒 𝑠 . An example of how this function is called
is given by:
1 A = int_rect([Link], a=0, b=[Link], N=1001)

The integration procedure above can in principle attain any accuracy


desired by making the width of the interval approach zero. This will also
increase the number of times the function value has to be evaluated and
hence the computational cost will continue to increase with increasing
accuracy. There are several methods that improve on the accuracy of
numerical integration, compared to this simple approach. The first of
these that we consider is the composite midpoint method, which is
schematically shown in Fig. 10.2a. In the midpoint-approach the area
under the curve is divided in a series of rectangles of width of ℎ just
as above. However, the difference is that the height of each rectangle
is defined by value the function at the midpoint of the width of the
rectangle. This way of calculating the integral is therefore given by:

∫ 𝑏 𝑁−
X1

𝑥 𝑗 + 𝑥 𝑗+1

𝑓 (𝑥)𝑑𝑥 ≈ ℎ 𝑓 (10.3)
𝑎 𝑗=1
2
10 Numerical integration 79

This approach of integrating is typically more accurate but it depends on


the shape of the function. For instance, for a linearly decreasing function,
the first rectangle approximation in Eq. 10.2 will always overestimate
the integral (albeit that the overestimation decreases with the number of
intervals), while the mid-point rule in Eq. 10.3 gives the exact solution
for any number of intervals.

Figure 10.2: Illustration of the the com-


posite midpoint method to evaluate the
∫𝑏
definite integral 𝑎
𝑓 (𝑥)𝑑𝑥 .

Trapezoid approximation

An alternative approach that has a similar accuracy and dependence on


the number of intervals is the trapezoid approximation. The method is
graphically shown in Fig. 10.3. The area under the curve is divided into a
series of trapezoids, rather than rectangles. We can also interpret this as
approximation of the function itself between two neighboring points by
a linear functions. As we will see below, this offers a systematic way to
improve the approximation of the integration. In this approximation the
integral is defined as:

" #
𝑏
ℎ 𝑁−
X1

𝑓 (𝑥)𝑑𝑥 ≈ 𝑓 (𝑥 𝑗 ) + 𝑓 (𝑥 𝑗+1 ) (10.4)
𝑎 2 𝑗=1

The trapezoidal approximation in Eq.10.4 is very similar to that for the


mid-point approximation in Eq.10.3 but is different. In the mid-point
approximation the actual function value in the middle of the interval is
taken, while for the trapezoid approximation the function value is the
average of two neighboring points. The computational cost of the two
approaches and their accuracy is similar but it does depend on the shape
of the function that is considered. The implementation of the trapezoid
rule is analogous to that for the rectangular approximations and is the
subject of one of the exercises at the end of this chapter.

Figure 10.3: Illustration of the the trape-


zoidal rule where trapezoids are used to
∫𝑏
evaluate the definite integral 𝑎
𝑓 (𝑥)𝑑𝑥 .
10 Numerical integration 80

Simpson’s rule

In the trapezoidal approximation the shape of the function between two


neighboring points is approximated by a linear function. This approach
can easily be extended to account for the curvature of the function in
between two points. While two neighboring points define a line, we can
also take three neighboring points and approximate the function over
this interval as a parabola. This parabola is uniquely defined by the three
points. This method is illustrated in Fig.10.4.

Figure 10.4: In Simpson’s rule the function


𝑓 (𝑥) is approximated over two neighbor-
ing intervals using a quadratic function,
𝑃(𝑥).

The integration can now be performed for the two intervals at once and
it is easily derived that this integral is defined as:

𝑖+2
ℎ 

𝑓 (𝑥)𝑑𝑥 ≈ 𝑓 (𝑥 𝑖 ) + 4 𝑓 (𝑥 𝑖+1 ) + 𝑓 (𝑥 𝑖+2 )

(10.5)
𝑖 3

The implementation of Simpson’s rule is very similar to for instance the


trapezoid approach but requires a loop over two intervals at a time to
prevent intervals from being counted twice. This also implies that the
number of points over which the summation is done should be an odd
number larger than or equal to 3.
Using Simpson’s rule gives errors that are generally orders of magnitude
smaller than for the trapezoid rule. To illustrate this we consider the
integral:

∫ 3
1
𝑑𝑥 = ln(3) (10.6)
1 𝑥

In Table 10.1 error for the numerical evaluation of this integral is given for
both the trapezoid rule and for Simpson’s rule. It is clear that Simpson’s
rule is by far superior and quickly approaches the numerical accuracy of
the computer itself. Importantly, the computational cost of Simpson’s rule
is similar as that for the trapezoidal approach, but for 𝑁 = 501 the error
is 5-6 orders of magnitude smaller! Therefore Simpson’s rule is generally
a high-accuracy method at a modest computational cost, making it a
method of great practical use.
10 Numerical integration 81

N Error trapezoid Error Simpson Table 10.1: Error in the evaluation of the
integral in Eq.10.6 for the trapezoidal ap-
5 1.8054e-02 1.3877e-03 proach and Simpson’s rule.
9 4.5984e-03 1.1306e-04
25 5.1401e-04 1.5619e-06
53 1.0956e-04 7.1788e-08
101 2.9628e-05 5.2624e-09
249 4.8175e-06 1.3923e-10
501 1.1852e-06 8.4284e-12

Table 10.2: Newton-Cotes formulas for different orders of approximation.

Name  Formula

𝑓 (𝑥 𝑖 ) + 𝑓 (𝑥 𝑖+1 )

Trapezoidal rule 2

𝑓 4 𝑓 (𝑥 𝑖+1 ) + 𝑓 (𝑥 𝑖+2 )
 
Simpson’s rule 3 (𝑥 𝑖 ) +
3ℎ
𝑓 (𝑥 𝑖 ) + 3 𝑓 (𝑥 𝑖+1 ) + 3 𝑓 (𝑥 𝑖+2 ) + 𝑓 (𝑥 𝑖+3 )
 
Simpson’s 3/8 rule 8
2ℎ
7 𝑓 (𝑥 𝑖 ) + 32 𝑓 (𝑥 𝑖+1 ) + 12 𝑓 (𝑥 𝑖+2 ) + 32 𝑓 (𝑥 𝑖+3 ) + 7 𝑓 (𝑥 𝑖+4 )
 
Boole’s rule 45

Newton-Cotes formulas

The trapezoidal approach and Simpson’s rule are closely related method.
In the first the function over an interval is approximated as a straight
line, while in the second a quadratic function is used. We can extend
this sequence of methods to higher order polynomials. In every step an
additional interval is used to define the method and a higher polynomial is
obtained that runs through these points. Following this general approach
we obtain the (closed) Newton-Cotes formulas summarized in Table
10.2. The higher order methods are not often used, especially because
in modern computers it is easy to use a large number of intervals in the
evaluation. Generally, Simpson’s rule gives a good trade-off for accuracy
vs. computational cost.

10.2 Gaussian quadrature

Until now we have only considered integration methods with equally


spaced intervals. While the Newton-Cotes formulas that we have dis-
cussed give a very good approximation of integrals at a modest cost,
there are more efficient methods, but these are considerably more com-
plicated also. One of the most used approaches to numerical integration
is so-called Gaussian quadrature. In this approach the grid that is used to
calculate the integral is not uniform but varies. The Gaussian quadrature
was constructed to yield exact results for polynomials of degree 2𝑛 − 1
or less by taking suitable values for the positions, 𝑥 𝑖 , and the weights, 𝑤 𝑖 ,
summarized as:

∫ 1 𝑛
X
𝑓 (𝑥)𝑑𝑥 = 𝑤 𝑖 𝑓 (𝑥 𝑖 ) (10.7)
−1 𝑖=1

In the most common implementation the function over the integration


interval is approximated by a Legendre polynomial, 𝑃𝑛 (𝑥) through a
very specific set of points 𝑥 𝑖 , (where 𝑥 𝑖 is the 𝑖 -th ’Gauss-node’). In
10 Numerical integration 82

order to apply this method to an arbitrary interval [𝑎 ; 𝑏] a change of the


integration interval to [−1 , 1] is necessary. This can be done by:

𝑏 1
𝑏−𝑎 𝑏−1 𝑎+𝑏
∫ ∫
𝑓 (𝑥)𝑑𝑥 = 𝑓( 𝑥+ )𝑑𝑥 (10.8)
𝑎 2 −1 2 2
𝑛
𝑏−𝑎 X 𝑏−1 1+𝑏
= 𝑤𝑖 𝑓 ( 𝑥𝑖 + ) (10.9)
2 𝑖=1
2 2

It can be shown that the the weights 𝑤 𝑖 are then given by:

2
𝑤𝑖 = (10.10)
(1 − 𝑥 𝑖 )2 [𝑃𝑛0 (𝑥 𝑖 )]2

Gaussian quadrature approaches are the methods of choice to numerically


evaluate integrals as they lead to great numerical accuracy at a very low
computational cost. They are conceptually very similar to the composite
methods with evenly spaced intervals that we have considered above but
they are substantially more complicated to implement. However, many
of such methods are already available in numpy/scipy libraries and we
can use them in our own codes. In the example below we show that the
Gaussian quadrature method implemented in 𝑠𝑐𝑖𝑝 𝑦 returns the value of
the integral of 𝑓 (𝑥) = 𝑥 sin(𝑥) with an error of 10−12 , which is the close
to machine precision, i.e. the highest possible accuracy of a numerical
method.
1 import numpy as np
2 import math
3 from [Link] import quadrature as quad
4

5 def func(x):
6 return x*[Link](x)
7

8 def exactIntegral(a, b):


9 Iab = -b*[Link](b)+[Link](b)+a*[Link](a)-[Link](a)
10 return Iab
11

12 a = 0.0
13 b = 2.0
14

15 exact = exactIntegral(a, b)
16 estimate = quad(func, a, b)
17

18 print("Exact %1.6f Numerical %1.6f" % (exact, estimate[0]))


19 print("Error %1.3e" % [Link](exact-estimate[0]))

10.3 Multiple Integrals

The methods discussed above all relate to the calculation of integrals


along one coordinate. Expansion to two or more coordinates is generally
straightforward, and the resulting equations are only slightly more com-
plicated that those for the one-dimensional case. For integrals dependent
on two variables, the integration can still be imagined as adding up
a series of volume elements such as cuboids. We can easily derive a
10 Numerical integration 83

two-dimensional analogue to the mid-point approximation used for one


variable for the general function:

∫ 𝑏 ∫ 𝑑 
𝐼= 𝑓 (𝑥, 𝑦)𝑑𝑦 𝑑𝑥 (10.11)
𝑎 𝑐

For the inner integral we substitute the midpoint rule, yielding:

∫ 𝑑 𝑁𝑦
X
𝐼1 = 𝑓 (𝑥, 𝑦)𝑑𝑦 = ℎ 𝑦 𝑓 (𝑥, 𝑦 𝑗 ) (10.12)
𝑐 𝑗=0

where 𝑦 𝑗 is the midpoint of the 𝑗 -th interval. Introducing this expression


in 𝐼 and using the midpoint method again leads to

𝑁𝑦
!
∫ 𝑏 X
𝐼 = ℎ𝑦 𝑓 (𝑥, 𝑦 𝑗 ) 𝑑𝑥
𝑎 𝑗=0
𝑁𝑦 ∫ 𝑏
X
= ℎ𝑦 𝑓 (𝑥, 𝑦 𝑗 )𝑑𝑥
𝑗=0 𝑎
𝑁𝑦
𝑁𝑥 X
X
= ℎ𝑥 ℎ𝑦 𝑓 (𝑥 𝑖 , 𝑦 𝑗 ) (10.13)
𝑖=0 𝑗=0

The resulting expression for integral is now the sum of the volume of the
rectangular columns that spanning the 𝑥 𝑦 plane, where the height of the
column is taken to be the function value at the average coordinate in both
the 𝑥 - and the 𝑦 -direction. Similar expressions can be obtained for any of
the composite methods that we have used above in one direction.

10.4 Exercises

Composite methods

In this exercise we will implement the midpoint and trapeze composite


methods for the numerical integration of simple functions illustrated in
Figures 10.2 and 10.3. We will first implement these two numerical approx-
imations and compare their accuracy with more advanced techniques
implemented in scipy. In the second part, we will use the implementation
to compute the amount of energy required to heat up a material.
Part I : Implementation As a reminder, the midpoint and trapeze method
approximate the definite integral

∫ 𝑏
𝐼= 𝑓 (𝑥)𝑑𝑥 (10.14)
𝑎

by the summing the area of a number of rectangles that spans the interval
to be integrated. The expression for the midpoint and trapeze methods
are given by
10 Numerical integration 84

𝑁−
X1

𝑥 𝑗 + 𝑥 𝑗+1

𝐼𝑚𝑖𝑑 = ℎ 𝑓 (10.15)
𝑗=1
2
" #
ℎ 𝑁−
X1
𝐼𝑡𝑟 𝑎𝑝 = 𝑓 (𝑥 𝑗 ) + 𝑓 (𝑥 𝑗+1 ) (10.16)
2 𝑗=1

where ℎ is the width of the rectangle. In the midpoint method the function
is evaluated in the middle of the 𝑗 -th rectangle (see Fig. 10.2). In contrast,
in the trapezoid method the average function values is taken of the two
points on either side of the interval (Fig. 10.3).
The two functions to be implemented should take as argument, the
beginning and end of the integration interval, 𝑎 and 𝑏 , the number of
integration points 𝑁 and the function to be integrated 𝑓 𝑢𝑛𝑐 . You will
compare the results provided by code with results given by the routine
[Link] presented in section 10.2.

To test the code, calculate the integral of

𝑥
𝑓 (𝑥) = 2𝜋𝑥 2 sin(𝜋𝑥) exp (− ) (10.17)
2𝜋

between 0 and 0.1. Plot the difference in the Δ𝑚𝑖𝑑 = 𝐼𝑚𝑖𝑑 − 𝐼 𝑠𝑝 and
Δ𝑡𝑟 𝑎𝑝 = 𝐼𝑡𝑟 𝑎𝑝 − 𝐼 𝑠𝑝 (where 𝐼 𝑠𝑝 is the value of the integral obtained with
the quadrature of scipy) as a function of 𝑁 . As Δ𝑡𝑟𝑎𝑝 and Δ𝑚𝑖𝑑 can reach
very small values, use a loglog plot to visualize them:
1 [Link](....)
2 [Link]

Part II: Enthalpy Calculation The calculation of the amount of energy


required to increase the temperature of a material is a common problem
in engineering. For example, if we have a cubic meter of nitrogen how
much energy is required to increase its energy by 1𝑜 𝐶 ? A refinery has
hired you to automate such calculation and compare different techniques
to do it.
The enthalpy change associated with a temperature change from 𝑇1 to 𝑇2
is given by:

∫ 𝑇2
Δ𝐻 = 𝐶 𝑝 (𝑇)𝑑𝑇 (10.18)
𝑇1

where 𝐶 𝑝 (𝑇) is the heat capacity of the materials at a constant pressure


𝑝 . This quantity is often given as a polynomial. For example the heat
capacity of nitrogen gas (given in 𝑘𝐽/(𝑚𝑜𝑙 𝑜 𝐶)) is:

𝐶 𝑝 (𝑇) = 0.0290 + 0.2199 × 10−5𝑇 + 0.5723 × 10−8𝑇 2 − 2.871 × 10−12𝑇 3


(10.19)
10 Numerical integration 85

where 𝑇 is in Celsius. On the other hand the heat capacity of carbon is


given by the polynomial expression:

𝐶 𝑝 (𝑇) = 0.1118 + 1.095 × 10−5𝑇 + 489.1/𝑇 2 (10.20)

where 𝑇 is in Kelvin. The refinery wants you to compute the enthalpy


change for both material starting at 𝑇 = 20 𝑜 C and raising the temperature
to 𝑇 = 100 𝑜 C. Report the enthalpy change in 𝑘𝐽/(𝑚𝑜𝑙). You should
numerically evaluate the integral using the [Link]
function.
The answers are 2.33 kJ/mol for nitrogen and 9.65 kJ/mol for carbon.

Simpson’s rule

Simpson’s rule is a composite method similar to the trapezoid method.


Instead of approximating each interval with a trapeze, Simpson’s rule
uses a quadratic polynomial as can be seen in the figure. The integral
between the points xi and xi+2 is given by

𝑖+2
ℎ 

𝑓 (𝑥)𝑑𝑥 ≈ 𝑓 (𝑥 𝑖 ) + 4 𝑓 (𝑥 𝑖+1 ) + 𝑓 (𝑥 𝑖+2 )

(10.21)
𝑖 3

Note that in the implementation the summation requires an even number


of intervals, and hence an odd number of points! In this exercise we
will implement a function that calculates an integral for any function
using Simpson’s rule and compare its performance to the trapezoid
approximation for the function:

∫ 𝜋
1
𝑒 𝑥 sin(𝑥) = (1 + 𝑒 𝜋 ) (10.22)
0 2

Take the following steps:


I Write a function that calculates and returns the function values for
the function 𝑓 (𝑥) = 𝑒 𝑥 sin(𝑥).
I Use the trapezoid function from the exercise above to calculate the
integral in Eq.10.22.
I Write an analogous function that calculates the integral using
Simpson’s method.
I Calculate the integral of the function above between 0 and 𝜋 using
the functions written above for the following values of N: 5, 11, 51,
255, 501, [Link] their error by calculation the difference
with the exact integral (see Eq.10.22).
Numerical differentiation 11
In any undergraduate course the analytical evaluation of derivative 11.1 First derivative . . . . . . . . . 86
of functions plays an important role. The derivative is useful in many Example . . . . . . . . . . . . . 88
problems in science and engineering, for instance when minimizing 11.2 Second derivative . . . . . . . 90
11.3 Partial derivatives in 2D . . . 91
or maximizing a certain quantity. An example is the optimization of
11.4 Summary . . . . . . . . . . . . 92
geometries of molecules in electronic structure calculations. In a more
11.5 Exercises . . . . . . . . . . . . . 93
general sense, many approaches to solve systems of non-linear equations
Second derivative . . . . . . . 93
use derivatives. Finally, in the solution of differential equations that will
be discussed in later chapters, the derivative takes a central point. While
derivatives can often be obtained analytically, in this chapter we will
focus on numerical methods to obtain derivatives. All approaches for the
numerical calculation of derivatives proceed from a Taylor expansion of
the function that is considered. In a Taylor expansion, the value of the
function in a point at a distance ℎ away from the point 𝑥 𝑖 is given by:

𝑑 𝑓 (𝑥 𝑖 ) ℎ 2 𝑑2 𝑓 (𝑥 𝑖 )
𝑓 (𝑥 𝑖 + ℎ) = 𝑓 (𝑥 𝑖 ) + ℎ + +... (11.1)
𝑑𝑥 2 𝑑𝑥 2
𝑑 𝑓 (𝑥 𝑖 )
= 𝑓 (𝑥 𝑖 ) + ℎ + 𝑂(ℎ 2 ) (11.2)
𝑑𝑥

In principle, the Taylor expansion goes on infinitely, with each term


adding to the accuracy. In this case we truncate the expansion after the
term with the first derivative, where the term 𝑂(ℎ 2 ) means that the error
is on the order of ℎ 2 .

11.1 First derivative

Following the Taylor expansion of the function around the point 𝑥 𝑖 , we


can rearrange the truncated expression in Eq. 11.2 to obtain an expression
for the first derivative of 𝑓 (𝑥) in terms of the function values 𝑓 (𝑥 𝑖 ) and
𝑓 (𝑥 𝑖+1 ):

𝑑 𝑓 (𝑥 𝑖 ) 𝑓 (𝑥 𝑖 + ℎ) − 𝑓 (𝑥 𝑖 )
= + 𝑂(ℎ 2 ) (11.3)
𝑑𝑥 ℎ
𝑓 (𝑥 𝑖 + ℎ) − 𝑓 (𝑥 𝑖 )
≈ (11.4)

This approximation is called the forward approximation of the first


derivative and the error is here on the order of ℎ . Numerical differentiation
are often performed on experimental data such where the values of 𝑓 (𝑥)
are a result of measurements and are available as a series of points
arranged in a 1D array, separated by a fixed distance on the X-axis. To
make this explicit we can change the notation to 𝑓 (𝑥 = 0) = 𝑓0 , 𝑓 (ℎ) = 𝑓1 ,
11 Numerical differentiation 87

𝑓 (2 ℎ) = 𝑓2 , etc ... Using this notation the forward approximation of the


first derivative reads:

𝑑 𝑓𝑖 𝑓𝑖+1 − 𝑓𝑖
= (11.5)
𝑑𝑥 ℎ

The use of this approximation is illustrated in Fig. 11.1a. It is easy to


see that this approximation becomes more and more accurate when the
distance between the points decreases, however, reducing the value of ℎ
to very small values can also introduce numerical inaccuracies.

Figure 11.1: Illustration of the different nu-


merical approximation for the first deriva-
tives of a function a) forward approxima-
tion, b) backward approximation.

Therefore, we would like to derive a more accurate approximation of


the derivative while keeping ℎ the same. To derive such an approach
we will first derive a similar expression, the so-called the backward
approximation of the derivative. This approximation is given by the
Taylor approximation of 𝑓 (𝑥 𝑖 − ℎ) in the same way as for the forward
approximation:

𝑑 𝑓 (𝑥 𝑖 )
𝑓 (𝑥 𝑖 − ℎ) = 𝑓 (𝑥 𝑖 ) − ℎ + 𝑂(ℎ 2 ) (11.6)
𝑑𝑥

The rearrangement in this case leads to:

𝑑 𝑓𝑖 𝑓𝑖 − 𝑓𝑖−1
= (11.7)
𝑑𝑥 ℎ

This backward approximation has the same 𝑂(ℎ) accuracy as the forward
approximation and its use is illustrated in Fig. 11.1b. Now we have two
equations for the derivative of the function in point 𝑥 𝑖 . We can combine
these two equations by adding them up:

𝑑 𝑓 (𝑥 𝑖 ) 𝑓𝑖+1 − 𝑓𝑖 𝑓𝑖 − 𝑓𝑖−1 𝑓𝑖+1 − 𝑓𝑖−1


2 = + = (11.8)
𝑑𝑥 ℎ ℎ ℎ

This leads to the disappearance of the term with 𝑓 (𝑥 𝑖 ), resulting in the


following equation for the first derivative at point 𝑥 𝑖 in terms of the
function values 𝑓 (𝑥 𝑖−1 ) and 𝑓 (𝑥 𝑖+1 ):

𝑑 𝑓 (𝑥 𝑖 ) 𝑓𝑖+1 − 𝑓𝑖−1
= (11.9)
𝑑𝑥 2ℎ
This approximation is called the centered difference approximation and
it can be shown that its accuracy is on the order 𝑂(ℎ 2 ) instead of 𝑂(ℎ)
for the forward and backward approximation. This means that for a
given step size ℎ , the centered difference approximation will give much
11 Numerical differentiation 88

more accurate results than either the backward or the forward forward
approximation. The improved accuracy can be easily imagined when
looking at the illustration of the centered approximation in Fig. 11.2. The
red line through the two points has a slope that is much closer to the
slope of the function itself, as compared to the forward and backward
approximations. This is even true for relatively large values of ℎ .

Figure 11.2: Illustration of the centered ap-


proximation for the first derivative defined
by the points before and after the point for
which the derivative is to be calculated.

Example

To illustrate the accuracy of the three differentiation methods discussed


above we consider the derivative of the function in Eq.11.10. We will
compare the results of the three different methods to the exact values of
the derivatives. The function that we consider is given by:

𝑓 (𝑥) = 𝑥 sin(𝑥) (11.10)

As a first step, we write a piece of code to calculate and plot the first
derivative of this function. The code below calculates the derivative using
the centered difference approximation.
1 import [Link] as plt
2 import numpy as np
3

4 # compute the derivative numerically


5 def df_dx(y, x):
6 h = x[1]-x[0]
7 df = 0.5/h*([Link](y, -1)-[Link](y, 1))
8 df[0] = 1./h*(y[1]-y[0])
9 df[-1] = 1./h*(y[-1]-y[-2])
10 return df
11

12 # compute the derivative analytically


13 def df_dx_exact(x):
14 return [Link](x)+x*[Link](x)
15

16 # function to derive
17 def func(x):
18 return x*[Link](x)
19

20 # main part of the code


21 X = [Link](0, 4*[Link], 1E5)
22 F = func(X)
23 dF = df_dx(F,X)
24 dF_exact = df_dx_exact(X)
11 Numerical differentiation 89

25

26 # plot the result


27 [Link](X, F, color=’black’, linewidth=3, label=’f’)
28 [Link](X, dF, color=’#006FFF’, linewidth=3, label=’df’)
29 [Link](X[0:-1:5000], dF_exact[0:-1:5000], ’o’, markersize=8, color=’#006
FFF’, label=’df_exact’)
30 [Link](’X’, fontsize=20)
31 [Link](’f(X)’, fontsize=20)
32 [Link](loc=2)
33 [Link]()

The central part of this code is the use of the [Link]() function:
1 [Link](y,1)

This functions returns an array identical to 𝑦 , but with all the elements
have been ’rolled’ along the vector. This means that all values are moved
to the right by one position. The last value of the original array will ’roll
over’ to the first position. For example:

𝑦 = [1 , 2 , 3, 4, 5] −→ 𝑦roll = [5 , 1 , 2 , 3 , 4] (11.11)

The trick with the [Link]() function works well for points that
have neighboring points on both sides. The first and last element of the
derivative do not have neighbors on both sides and therefore cannot be
evaluated using the central difference. The only solution is to use the
forward and backward difference approximation for these two points,
which is done in the code above on two separate lines. The comparison
between the numerical and exact values of the derivative for the function
in Eq.11.10 is shown in Fig. 11.3.

Figure 11.3: Example of numerical deriva-


tion of the function 𝑓 (𝑥) = 𝑥 sin(𝑥) and
comparison with its exact derivative.

We can further compare the accuracy of the forward, backward and


central method for the calculation of the derivative. In the Python code
below the value of the derivative of 𝑓 (𝑥) at 𝑥 = 1.0 is calculated for
several values of the interval ℎ .
1 import numpy as np
2

3 def fun(x):
4 return x*[Link](x)
5

6 def exactDeriv(x):
7 return [Link](x) + x*[Link](x)
8

9 x0 = 1.0
10 exact = exactDeriv(x0)
11

12 for i in range(1, 7):


13
11 Numerical differentiation 90

Figure 11.4: Illustration of the error ob-


tained for the different numerical approx-
imation of the derivative. As you can
see the centered approximation leads to
smaller error down to ℎ = 10−6 .

14 h = 10**(-i)
15

16 fa = (fun(x0+h)-fun(x0))/h
17 ba = (fun(x0)-fun(x0-h))/h
18 ca = (fun(x0+h)-fun(x0-h))/2/h
19

20 print(" h = %1.4e" % h)
21

22 print(" forward approximation error %1.4e " % [Link](exact-fa))


23 print(" bacward approximation error %1.4e " % [Link](exact-ba))
24 print(" centered approximation error %1.4e " % [Link](exact-ca))
25 print("")

As can be seen in Fig. 11.4, the error decreases much faster for the centered
approximation than for the backward and forward approximation that
both have a similar accuracy.

11.2 Second derivative

In all three approaches used above for the first derivative, we started from
the Taylor expansion. This is a very common approach in the derivation
of numerical methods, especially for derivatives and it can also be used
to derive an expression for the second derivative. Start once again from
the Taylor approximation of a function around 𝑥 𝑖 ± ℎ , but now truncating
it after the term with the second derivatives:

𝑑 𝑓 (𝑥 𝑖 ) ℎ 2 𝑑2 𝑓 (𝑥 𝑖 )
𝑓 (𝑥 𝑖 + ℎ) = 𝑓 (𝑥 𝑖 ) + ℎ + + 𝑂(ℎ 3 ) (11.12)
𝑑𝑥 2 𝑑𝑥 2
𝑑 𝑓 (𝑥 𝑖 ) ℎ 2 𝑑2 𝑓 (𝑥 𝑖 )
𝑓 (𝑥 𝑖 − ℎ) = 𝑓 (𝑥 𝑖 ) − ℎ + − 𝑂(ℎ 3 ) (11.13)
𝑑𝑥 2 𝑑𝑥 2

These two expressions both contain the second derivative, and upon
adding them together the terms with the first derivative disappear:

𝑑2 𝑓 (𝑥 𝑖 )
𝑓 (𝑥 𝑖 + ℎ) + 𝑓 (𝑥 𝑖 − ℎ) = 2 𝑓 (𝑥 𝑖 ) + ℎ 2 + 𝑂(ℎ 4 ) (11.14)
𝑑𝑥 2

This equation can be rearranged to obtain an approximation for the


second derivative. Using the simplified notation for 1D array we now
11 Numerical differentiation 91

obtain:

𝑑 2 𝑓𝑖 𝑓𝑖+1 − 2 𝑓𝑖 + 𝑓𝑖−1
≈ (11.15)
𝑑𝑥 2 ℎ2

which is 𝑂(ℎ 2 ) accurate. This is called the centered difference approxi-


mation for the second derivative. More accurate expressions have been
derived for the second derivative. However, the expression 11.15 is used
in most applications in science and engineering as it is accurate and fast.
The function 𝑛𝑝.𝑟𝑜𝑙𝑙() can also be used to rapidly calculate the values of
the second derivative:
1 d2f = 1./h/h * ([Link](f,-1)-2*f+[Link](f,1))

Again, the elements 𝑑 2 𝑓 [0] and 𝑑 2 𝑓 [−1] are not evaluated correctly.
To evaluate these elements we have to use the forward and backward
expression that in fact makes the second derivative of the first and last
points equal to their neighbours:

𝑑 2 𝑓𝑖 1
= 2 ( 𝑓𝑖+2 − 2 𝑓𝑖+1 + 𝑓𝑖 ) + 𝑂(ℎ) (11.16)
𝑑𝑥 2 ℎ
𝑑 2 𝑓𝑖 1
= 2 ( 𝑓𝑖 − 2 𝑓𝑖−1 + 𝑓𝑖−2 ) + 𝑂(ℎ) (11.17)
𝑑𝑥 2 ℎ

11.3 Partial derivatives in 2D

We have seen how to extract numerical expressions of the first and second
derivatives for one-dimensional functions. Similar techniques can be
used to derive the partial derivatives of a function of two variables 𝑓 (𝑥, 𝑦).
We will adopt the notation:

𝑓 (𝑥 𝑖 , 𝑦 𝑗 ) = 𝑓𝑖𝑗 (11.18)

For such a function the grid we use has to extend along both the variable
axes. We will take an evenly spaced grid with an increment ℎ in the 𝑥
direction and 𝑘 in the 𝑦 direction. For the first derivatives the central
difference formula directly leads to:

𝑑 1
𝑓 (𝑥 𝑖 , 𝑦 𝑗 ) = ( 𝑓𝑖+1,𝑗 − 𝑓𝑖−1,𝑗 ) (11.19)
𝑑𝑥 2ℎ
𝑑 1
𝑓 (𝑥 𝑖 , 𝑦 𝑗 ) = ( 𝑓𝑖,𝑗+1 − 𝑓𝑖,𝑗−1 ) (11.20)
𝑑𝑦 2𝑘

The second partial derivatives are also given by the previously derived
expressions:
11 Numerical differentiation 92

𝑑2 1
𝑓 (𝑥 𝑖 , 𝑦 𝑗 ) = ( 𝑓𝑖+1,𝑗 − 2 𝑓𝑖,𝑗 + 𝑓𝑖−1,𝑗 ) (11.21)
𝑑𝑥 2 ℎ2
𝑑2 1
𝑓 (𝑥 𝑖 , 𝑦 𝑗 ) = ( 𝑓𝑖,𝑗+1 − 2 𝑓𝑖,𝑗 + 𝑓𝑖,𝑗−1 ) (11.22)
𝑑𝑦 2 𝑘2

The mixed partial derivatives can then be obtained by combining first


derivatives expression

𝑑2 𝑑 𝑑
 
𝑓 (𝑥 𝑖 , 𝑦 𝑗 ) = 𝑓 (𝑥 𝑖 , 𝑦 𝑗 ) (11.23)
𝑑𝑥𝑑𝑦 𝑑𝑥 𝑑𝑦
𝑑 1
 
= ( 𝑓𝑖,𝑗+1 − 𝑓𝑖,𝑗−1 ) (11.24)
𝑑𝑥 2 𝑘
1
𝑓𝑖+1,𝑗+1 − 𝑓𝑖−1,𝑗+1 − 𝑓𝑖+1,𝑗−1 + 𝑓𝑖−1,𝑗−1(11.25)

=
4ℎ 𝑘

11.4 Summary

There are many ways of deriving finite difference expressions for the
derivatives of a given function depending on the method used (forward,
backward, central) and the desired accuracy. In the table below, some
expressions are summarized for future reference.
11 Numerical differentiation 93

Forward difference : Error of order ℎ Forward difference : Error of order ℎ 2

𝑑𝑦 𝑖 1 𝑑𝑦 𝑖 1
= (𝑦 𝑖+1 − 𝑦 𝑖 ) = (−𝑦 𝑖+2 + 4 𝑦 𝑖+1 − 3 𝑦 𝑖 )
𝑑𝑥 ℎ 𝑑𝑥 2ℎ
𝑑2 𝑦𝑖 1 𝑑 𝑦𝑖
2
1
= 2 (𝑦 𝑖+2 − 2 𝑦 𝑖+1 + 𝑦 𝑖 ) = 2 (−𝑦 𝑖+3 + 4 𝑦 𝑖+2 − 5 𝑦 𝑖+1 + 2 𝑦 𝑖 )
𝑑𝑥 2 ℎ 𝑑𝑥 2 ℎ

Backward difference : Error of order ℎ Backward difference : Error of order ℎ 2

𝑑𝑦 𝑖 1 𝑑𝑦 𝑖 1
= (𝑦 𝑖 − 𝑦 𝑖−1 ) = (3 𝑦 𝑖 − 4 𝑦 𝑖−1 + 𝑦 𝑖−2 )
𝑑𝑥 ℎ 𝑑𝑥 2ℎ
𝑑2 𝑦𝑖 1 𝑑2 𝑦𝑖 1
= 2 (𝑦 𝑖 − 2 𝑦 𝑖−1 + 𝑦 𝑖−2 ) = 2 (2 𝑦 𝑖 − 5 𝑦 𝑖−1 + 4 𝑦 𝑖−2 − 𝑦 𝑖−3 )
𝑑𝑥 2 ℎ 𝑑𝑥 2 ℎ

Central difference : Error of order ℎ 2

𝑑 2 𝑓𝑖 𝑓𝑖+1 − 2 𝑓𝑖 + 𝑓𝑖−1
=
𝑑𝑥 2 ℎ2
𝑑2 1
𝑓 (𝑥 𝑖 , 𝑦 𝑗 ) = 𝑓𝑖+1,𝑗+1 − 𝑓𝑖−1,𝑗+1 − 𝑓𝑖+1,𝑗−1 + 𝑓𝑖−1,𝑗−1

𝑑𝑥𝑑𝑦 4ℎ 𝑘

11.5 Exercises

Second derivative
I Write a function that takes as its argument an array of values and
numerically evaluates the second derivative in each point. Make
sure you treat the points at either end of the interval correctly. All
the equations needed are given in section 11.2.
I To test your script use it to calculate the second derivative of the
function 𝑓 (𝑥) = 𝑥𝑠𝑖𝑛(𝑥) in the interval 0 to 4𝜋.
I Compare your result to the analytical solution of the problem and
plot in a single image the function, your solution of the second
derivative and the analytical solution, similar to Fig.11.3.
Non-linear equations 12
In chapter 9 we have discussed how systems of linear equation can be 12.1 The bisection method . . . . 94
solved using techniques from linear algebra, i.e. using matrix operations Recursion . . . . . . . . . . . . 95
such as Gauss-Jordan elimination or the Jacobi algorithm to find the Recursive bisection method 96
12.2 The Newton method . . . . . 97
N solutions to a set of N linear equations. However many problems
Two coupled equations . . . 99
in science and engineering involve non-linear equations, rather than
Example . . . . . . . . . . . . . 101
linear ones. Such nonlinear equations can contain powers of the variables
12.3 Reduced Newton Method . . 103
( 𝑥 3𝑖 ), products of different variables ( 𝑥 𝑖 𝑥 𝑗 ) or even more complicated Multiple equations . . . . . . 103
terms (cos(𝑥 𝑖 ), 𝑒 𝑥 𝑘 ...). Such non-linear equations or systems of multiple 12.4 Exercises . . . . . . . . . . . . . 104
non-linear equations can not be solved in the same straightforward way The bisection method . . . . 104
and different approaches are needed. In this chapter we will discuss two The Newton method . . . . . 104
general methods to find the solutions of non-linear equations numerically; Newton for two equations . 106
the bisection method and Newton’s method. 12.5 Extra exercises . . . . . . . . . 106
Gradient descent in 2D . . . 106

12.1 The bisection method

Finding solutions for any equation is generally called root-finding, be-


cause all equations can be rearranged into a form that is equal to zero
( 𝑓 (𝑥) = 0). As an example we take the function:

𝑓 (𝑥) = 4 𝑥 3 − 2 𝑥 2 + 4 𝑥 − 4 = 0 (12.1)

Solving this equation requires finding the values of 𝑥 such that 𝑓 (𝑥) = 0.
An intuitive algorithm by which we can find the root of an arbitrary
function is the bisection method. The algorithm locates the root of 𝑓 (𝑥)
in a predefined interval [a, b] where 𝑓 (𝑎) and 𝑓 (𝑏) have opposite signs
(i.e. the value of 𝑥 for which the function value is zero is between 𝑎 and
𝑏 . The method requires that there is only a single root between a and b.
This is illustrated in Fig.12.1.
The first step in the algorithm is to determine the middle of the interval,
between 𝑎 and 𝑏 . By calculation the function value, 𝑓 (𝑐) and comparing
it to 𝑓 (𝑎) and 𝑓 (𝑏) we can determine whether it is on the same side of
the ’root’ as either 𝑎 or 𝑏 . In the case of Fig. 12.1, 𝑓 (𝑐) has the same sign
as 𝑓 (𝑏) and is therefore on the same side of the root (remember that 𝑓 (𝑎)
and 𝑓 (𝑏) are required to have opposite signs). Therefore, we replace 𝑏 by
𝑐 and start from the new values of 𝑎 and 𝑏 .
The algorithm for the bisection method can be summarized as follow:

1. compute the midpoint 𝑐 = (𝑎 + 𝑏)/2 and the value 𝑓 (𝑐)


2. if 𝑓 (𝑐) has the same sign than 𝑓 (𝑎) search the root in the interval
[c,b] by setting 𝑎 = 𝑐 if 𝑓 (𝑐) has opposite sign than 𝑓 (𝑎) search the
root in the interval [a,c] by setting 𝑏 = 𝑐
3. check the convergence by computing 𝜖 = |𝑏 − 𝑎| and iterate the
process (i.e. go back to 1) until we reach the desired accuracy (i.e.
𝜖 < 𝑀𝐼𝑁 _𝐸𝑃𝑆 )
12 Non-linear equations 95

Figure 12.1: The bisection methods: the in-


terval [a,b] is defined such that it contains
a single root of the function.

The method successively divides the search interval in two parts, select the
interval where the root is and repeat the process on the new interval until
the points of the interval are closer together than a predefined convergence
criterion. This algorithm can be implemented in a straightforward way
in a ’while’-loop. However, the structure of the algorithm is very suitable
for the use of a recursive approach that we have seen before when
determining the determinant of a matrix.

Recursion

Recursion is a very powerful approach that can lead to very elegant


functions. but it can also be somewhat complicated to grasp; a recursive
function is a function that calls itself! To illustrate the concept of recursive
functions we will write a function that calculates the factorial of a number
in a recursive way. The implementation is given in the code below:
1 def factorial(n):
2 if n==1: return 1
3 else: return n*factorial(n-1)

The function factorial() is called within its own definition. To understand


what is happening here we can introduce a few print statements to keep
track of the different calls of the function factorial(). The introduction
of print statements is generally a very good approach to follow what
happens in a codes and is very useful when debugging.
1 def factorial(n):
2 print(’Factorial function called with n=%d’ %(n))
3 if n==1:
4 return 1
5 else:
6 res = n*factorial(n-1)
7 print(’Intermediate result for %d * factorial(%d) : %d’ %(n,n-1,res))
8 return res
9

10 factorial(5)

This code produces the following output:


12 Non-linear equations 96

1 Factorial function called with n=5


2 Factorial function called with n=4
3 Factorial function called with n=3
4 Factorial function called with n=2
5 Factorial function called with n=1
6 Intermediate result for 2 * factorial(1) : 2
7 Intermediate result for 3 * factorial(2) : 6
8 Intermediate result for 4 * factorial(3) : 24
9 Intermediate result for 5 * factorial(4) : 120

The code is first called with 𝑛 = 5 that leads to a call of the function
with 𝑛 = 4, which leads to a call for 𝑛 = 3, etc., until finally we obtain
a call with 𝑛 = 1. This calls return a value of 1 that is multiplied by 2
and returned as the return value of the call factorial(2). This value of 2 is
then multiplied by 3 and returned as return value of the factorial(3) and
so on. Finally, we obtain the correct value of factorial(5). The recursive
approach fully omits the need for a for-loop that we have used in Chapter
3 to calculate the factorial and implements the calculation of the factorial
in a single if-else statement.

Recursive bisection method

As the example of the calculation of the factorial shows, recursive function


lead to very elegant codes, but we can always choose whether we use
a recursive implementation of not. In one of the exercises at the end of
the chapter we will implement the bisection methods using a while-loop.
In the code below a recursive implementation of the method is given.
The core part of the bisection methods is on lines 32-35 of the code. On
line 35 we first determine whether the absolute values of the difference
between the two bounds of the interval [a,b] is larger than the required
maximum error. If this is not the case, the method is converged. If not,
we replace one of the bounds of the interval by 𝑐 , depending on the sign
of the function value 𝑓 (𝑐) and call the function itself again. All other
lines in the code deal with some necessary ’bookkeeping’ calculating
the function values 𝑓 (𝑎), etc. and making sure that we keep track of the
number of iterations.
1 import numpy as np
2

3 # function we want to study


4 def def_func(x):
5 y = 4*x**3-2*x**2+4*x-4
6 return y
7

8 # recursive function for the bisection


9 def find_root(f,a,b,n=0,err=1E-3,NMAX=100):
10

11 print(’’)
12 print(’-’*20)
13 print(’Iteration %d/%d’ %(n, NMAX))
14

15 # determibe the midpoint and the values of f


16 c = (a+b)*0.5
17 fa, fb, fc = f(a), f(b), f(c)
18

19 print(’Domain a = %f c = %f b = %f’ % (a, c, b))


20 print(’Values f(a) = %f f(c) = %f f(b) = %f’ % (fa, fc, fb))
21

22 # increment the iteration counter


23 n += 1
12 Non-linear equations 97

24

25 # exit if we have reached the maximum iteration number


26 if n>NMAX:
27 ’Solution not found after %d iterations’ %(n)
28 return c
29

30 # if the desired accuracy has not been reached


31 # we change the interval and call the function again
32 if [Link](b-a)>err:
33 if f(a)*f(c) > 0: a = c
34 else: b = c
35 find_root(f,a,b,n=n)
36

37 # if we have reached the desired accuracy


38 # we exit the function and return the root
39 else:
40 print(’Root found at x=%f’ % c)
41 return c
42

43 # definition of the total search interval


44 minX, maxX = -10.0, 10.0
45

46 # call for the recursive function


47 c = find_root(def_func, minX,maxX)

It is important to note that the method of implementation will not affect


the outcome of the calculation. This will be the same, regardless of
whether we use recursion or not.

12.2 The Newton method

A second, very versatile methods to calculate the root of a non-linear


function is the Newton approach that requires both the function itself,
and its derivative. The methods always converges to the closest root of the
function and therefore, the initial ’guess’ value that is used determines
whether it finds the desired root. This is similar to the bisection method
that has very specific conditions for the choice of the interval. To illustrate
the Newton-approach we will consider the following function:

𝑓 (𝑥) = 𝑥 2 − 2𝑥 = 0 (12.2)
The starting point for deriving the Newton method is a Taylor expansion
as we have seen before already. Using the Taylor expansion we can
expand the function around a certain starting value 𝑥 0 in terms of the
derivatives of the function in 𝑥 0 :

(𝑥 − 𝑥0 )2 00
𝑓 (𝑥) = 𝑓 (𝑥0 ) + (𝑥 − 𝑥0 ) 𝑓 0(𝑥0 ) + 𝑓 (𝑥 0 ) + 𝑂(𝑥 − 𝑥0 )3 (12.3)
2!

where the last term tells us that the error is on the order of (𝑥 − 𝑥 0 )3 ).
We are interested in finding the value for which 𝑓 (𝑥) = 0 using the first
derivative of the function at 𝑥 0 . Therefore, we can truncate the Taylor
expansion to obtain a first order approximation:

0 = 𝑓 (𝑥 0 ) + (𝑥 − 𝑥 0 ) 𝑓 0(𝑥 0 ) (12.4)
12 Non-linear equations 98

This truncation means that for the time being, we assume that the
function is linear with a slope equal to the derivative of 𝑓 (𝑥). In this
linear approximation we can easily find the value 𝑥 for which it becomes
zero by rearranging it into:

𝑓 (𝑥0 )
𝑥 = 𝑥0 − (12.5)
𝑓 0(𝑥0 )

Starting from the initial guess 𝑥 0 we obtain a better approximation, 𝑥 1


that is closer to the solution. This is illustrated in Fig.12.2. The value 𝑥 1
can then be used as the new ’guess’ for which we can again use Eq.12.5 to
find 𝑥 2 , etc., until the error in the calculation is smaller than a predefined
threshold.

Figure 12.2: Illustration for the Newton’s


method used to find the solution of non-
linear equation.

in Eq.12.5, the value of 𝑥 𝑛+1 is calculated using the the function value
and the derivative at 𝑥 𝑛 : 𝑓 (𝑥 𝑛 ) and 𝑓 0(𝑥 𝑛 ). For a simple function such as
the quadratic one in Eq.12.2 the derivative is easily obtained analytically,
but in general we can insert in Eq.12.5 the numerical value for the first
derivative:

𝑓 (𝑥0 ) 2ℎ
𝑥 = 𝑥0 − = 𝑥0 − 𝑓 (𝑥 0 ) (12.6)
𝑓 (𝑥 0 )
0 𝑓 (𝑥0 + ℎ) − 𝑓 (𝑥0 − ℎ)
In this case we have used the centered approximation for the derivative,
allowing us to obtain an accurate value without the need for analytical
derivatives. The Newton method for a single dimension is implemented
in the code below:
1 import numpy as np
2

3 # non linear function


4 def func(x):
5 return x**2 - 2*x
6

8 # compute the derivative numerically


9 def fprime(func, **kwargs):
10 x = [Link](’x’)
12 Non-linear equations 99

11 h = [Link](’h’)
12 return (func(x+h)-func(x-h))/2./h
13

14 # initial guess
15 x0 = 5.0
16

17 # step for the derivatives


18 step = 1E-3
19

20 # initial convergence and tolerance


21 eps, tol = 1, 1E-6
22

23 # iterations and maximum iterations


24 niter, MAXITER = 1, 100
25

26 # loop until convergence is reached


27 while (eps>tol) and (niter<MAXITER):
28

29 # compute the function and its derivative at x=x0


30 f = func(x0)
31 fp = fprime(func, x=x0, h=step)
32

33 # compute the new guess with the newton method


34 x1 = x0 - f/fp
35

36 # compute the covergence criteria and


37 eps = [Link](x1-x0)
38

39 # if we have converged we print the solution


40 if eps<tol:
41 print(’\n \tConvergence reached’)
42 print(’\tSolution at x = %f\n’ % (x1))
43

44 # else we update the old guess


45 else:
46 x0 = x1

Two coupled equations

The Newton method presented here can be extended to solve systems


involving systems of multiple nonlinear equations. We will first present
a simple example with only two equations and then expand this to
the general solution for 𝑘 equations in the next section, which follows
in a straightforward way from the solution of a system of only two
equations.
We consider a system with unknown variables 𝑥 1 and 𝑥 2 that are related
to each other by two nonlinear equations that have the general form:

𝑓1 (𝑥1 , 𝑥2 ) = 0 (12.7)
𝑓2 (𝑥1 , 𝑥2 ) = 0 (12.8)

We can again start from a Taylor series to expand each of these functions
(0) (0)
around an initial estimate of the solution indicated by 𝑥 1 and 𝑥 2
12 Non-linear equations 100

(0) (0) 𝜕 𝑓1 (0) 𝜕 𝑓1 (0)


𝑓1 (𝑥1 , 𝑥2 ) = 𝑓1 (𝑥1 , 𝑥2 ) + | (0) · (𝑥1 − 𝑥1 ) + | (0) · (𝑥2 − 𝑥2 ) + ...
𝜕𝑥1 𝑥 𝜕𝑥 2 𝑥
(0) (0) 𝜕 𝑓2 (0) 𝜕 𝑓2 (0)
𝑓2 (𝑥1 , 𝑥2 ) = 𝑓2 (𝑥1 , 𝑥2 ) + | (0) · (𝑥1 − 𝑥1 ) + | (0) · (𝑥2 − 𝑥2 ) + ...
𝜕𝑥1 𝑥 𝜕𝑥 2 𝑥

Since we are dealing with functions that depend on two variables, the
Taylor expansion includes partial derivative with respect to both of these.
The left-hand side of these equation can be set equal to zero since we are
looking for the root. If we truncate the Taylor series to the first derivative
terms and rearrange the equations we obtain:

𝜕 𝑓1 (0) 𝜕 𝑓1 (0) (0) (0)


| (0) · (𝑥1 − 𝑥1 ) + | (0) · (𝑥2 − 𝑥2 ) = − 𝑓1 (𝑥1 , 𝑥2 ) (12.9)
𝜕𝑥1 𝑥 𝜕𝑥2 𝑥
𝜕 𝑓2 (0) 𝜕 𝑓2 (0) (0) (0)
| (0) · (𝑥1 − 𝑥1 ) + | (0) · (𝑥2 − 𝑥2 ) = − 𝑓2 (𝑥1 , 𝑥2 ) (12.10)
𝜕𝑥1 𝑥 𝜕𝑥2 𝑥

To simplify these equations we introduce a new variable 𝛿 :

(0) (0)
𝛿1 = 𝑥1 − 𝑥1 (12.11)
(0) (0)
𝛿2 = 𝑥2 − 𝑥2 (12.12)

leading to the following equations:

𝜕 𝑓1 (0) 𝜕 𝑓1 (0) (0) (0)


| (0 ) · 𝛿 1 + | (0) · 𝛿2 = − 𝑓1 (𝑥 1 , 𝑥2 ) (12.13)
𝜕𝑥1 𝑥 𝜕𝑥2 𝑥
𝜕 𝑓2 (0) 𝜕 𝑓2 (0) (0) (0)
| (0 ) · 𝛿 1 + | (0) · 𝛿2 = − 𝑓2 (𝑥 1 , 𝑥2 ) (12.14)
𝜕𝑥1 𝑥 𝜕𝑥2 𝑥

These two equation s can be rewritten in the form of a matrix-equation,


analogous to the approach taken for systems of linear equations as
considered in Chapter 9:

𝜕 𝑓1 𝜕 𝑓1
! ! !
(0) (0)
|
𝜕𝑥 1 𝑥 (0)
|
𝜕𝑥2 𝑥 (0) 𝛿 𝑓
𝜕 𝑓2 𝜕 𝑓2 · 1(0) = − 1(0) (12.15)
|
𝜕𝑥 1 𝑥 (0)
|
𝜕𝑥2 𝑥 (0)
𝛿2 𝑓2

or even

𝕁 · 𝜹 = −f (12.16)
The matrix 𝕁 is the Jacobian matrix or first derivative matrix of the system.
The elements of the Jacobian can either be calculated analytically or
numerically for any well-behaved function. The numerical calculation of
partial first derivatives was discussed in Chapter 11. Once we can set up
the Jacobian, the matrix equation in Eq. 12.15 can be solved using any of
the techniques discussed in Chapter 9 for linear systems. This solution
(0) (0)
provides values for 𝛿 1 and 𝛿 2 that can be added to our initial guess of
the solution to obtain an improved guess of the solution:
12 Non-linear equations 101

(1) (0) (0)


𝑥1 = 𝑥1 + 𝛿1 (12.17)
(1) (0) (0)
𝑥2 = 𝑥2 + 𝛿2 (12.18)

These new values can now be used to construct a Jacobian again, in


Eq. 12.15 to obtain better and better approximations in an iterative way.

Example

As an example suppose we consider the set of nonlinear equations;

𝑓1 (𝑥1 , 𝑥2 ) = 𝑥13 + 𝑥22 = 0 (12.19)


𝑓2 (𝑥1 , 𝑥2 ) = 𝑥12 − 𝑥23 =0 (12.20)

We can write the Jacobian of this system in terms of the analytical


derivatives:

𝜕 𝑓1 𝜕 𝑓1
!
3 𝑥 12 2𝑥2
 
𝜕𝑥1 𝜕𝑥 2
𝕁= 𝜕 𝑓2 𝜕 𝑓2 = (12.21)
2𝑥1 −3𝑥 22
𝜕𝑥1 𝜕𝑥 2

The Newton algorithm is initiated by taking initial guess values for the
solutions as:

(0)
𝑥1 = 1 (12.22)
(0)
𝑥2 =2 (12.23)

Inserting these values leads to the linear system of equations:

!
(0)
𝛿
   
3 4 5
· 1(0) = − (12.24)
2 −12 𝛿2 − 7

Solving this matrix equation with, for example, the Jacobi iterative
method leads to:

(0)
𝛿1 = −0.7272727 (12.25)
(0)
𝛿2 = −0.70454545 (12.26)

and therefore we can calculate the updated values of 𝑥 1 and 𝑥 2 :

(1) (0) (0)


𝑥1 = 𝑥1 + 𝛿1 = 1.0 − 0.7272727 = 0.272727272 (12.27)
(1) (0) (0)
𝑥2 = 𝑥2 + 𝛿2 = 2.0 − 0.70454545 = 1.29545455 (12.28)
12 Non-linear equations 102

We can now use these new values of the solution to recompute the
(1) (1)
Jacobian 𝕁 and the right-hand side vector f to obtain 𝛿 1 and 𝛿 2 and
then new a solution. The Newton method is implemented in the code
listed below.
1 import numpy as np
2

3 # non linear function


4 def func(x1,x2):
5 return [Link]([x1**3 + x2**2, x1**2 - x2**3])
6

7 # compute the Jacobian


8 def jacobian(func, **kwargs):
9 x1 = [Link](’x1’)
10 x2 = [Link](’x2’)
11 J = [Link]([[3*x1**2,2*x2 ], [2*x1,-3*x2**2]])
12 return J
13

14 # initial guess
15 x0 = [Link]([1.0, 2.0])
16

17 print(’initial guess’)
18 print("x1 = %f \t x2 = %f" % (x0[0], x0[1]))
19 print(’’)
20 # initial convergence and tolerance
21 eps, tol = 1, 1E-6
22

23 # iterations and maximum iterations


24 niter, MAXITER = 1, 1000
25

26 # loop until convergence is reached


27 while (eps>tol) and (niter<MAXITER):
28

29 # compute the function and the jacobian


30 f = func(x0[0], x0[1])
31 j = jacobian(func, x1=x0[0], x2=x0[1])
32

33 # solve the linear system


34 delta = [Link](j, -f)
35

36 # uptade the solution


37 x1 = x0 + delta
38

39 # check if we have converged


40 f1 = func(x1[0], x1[1])
41 eps = [Link](f1)
42

43 print("x1 = %f \t x2 = %f \t eps = %f" % (x1[0], x1[1], eps))


44

45

46 if (eps < tol):


47 print(’\n \tConvergence reached in %d iteration’ % niter)
48 print(’\tSolution at x = [%f %f]\n’ % (x1[0],x1[1]))
49

50 else:
51 x0 = x1
52 niter += 1

As we can see the code converges towards the correct solution, i.e. 𝑥 1 = −1
𝑥2 = 1 in 613 iterations. As can be seen by printing the convergence criteria,
the error increases quite dramatically in the first iterations of the method.
This explains why so many iterations are required to converge toward
the solution. This also illustrates that the choice of the initial guess is
very important: if we choose values close to the actual solution we have a
good chance of fast convergence, but sometimes some tweaking of the
12 Non-linear equations 103

initial guess is needed.

12.3 Reduced Newton Method

To limit the number of step taken during the optimization we can search
along the direction of the Newton step for a point where the error is less
than the error at the starting point. This can be done simply by adding
the following lines of code at line 32 of the code shown above
1 # uptade the solution
2 isearch = 0
3 max_search = 10
4 eps1 = 2*eps
5

6 while (eps1 > eps) and (isearch<max_search):


7

8 # new solution
9 x1 = x0 + 0.5**(isearch)*delta
10

11 # check if we have converged


12 f1 = func(x1[0],x1[1])
13 eps1 = [Link](f1)
14

15 # increment
16 isearch += 1
17

18 # update the convergence


19 eps = eps1

As we can see, in this code the new solution is recalculated as

𝑥 (𝑛+1) = 𝑥 (𝑛) + 0.5isearch 𝛿(𝑛) (12.29)

with isearch is incremented from 0 to 10 as long as the error is larger than


the one obtained with 𝑥 (𝑛) . Using this simple optimization of the Newton
step, the method converges to the correct solution in only 11 iterations!

Multiple equations

The illustration of Newton’s method as presented above for two coupled


nonlinear equations directly shows how it can be extended to deal with
a system of 𝑘 coupled nonlinear equations:

𝑓1 (𝑥1 , . . . 𝑥 𝑘 ) = 0 (12.30)
...... (12.31)
𝑓 𝑘 (𝑥1 , . . . 𝑥 𝑘 ) = 0 (12.32)

Using the same approach as above we can collect these 𝑘 equations and
12 Non-linear equations 104

write them as a matrix equation that can be solved as a linear system:

𝜕 𝑓1 𝜕 𝑓1 𝜕 𝑓1
© 𝜕𝑥1 𝜕𝑥 2
... 𝜕𝑥 𝑘 ª 𝛿1 𝑓1
­ 𝜕 𝑓2 𝜕 𝑓2 𝜕 𝑓2 ®
... ­ 𝛿2 ® ­ 𝑓2 ®
© ª © ª
­ 𝜕𝑥1 𝜕𝑥 2 𝜕𝑥 𝑘 ®
­ .. .. .. .. ® · ­ .. ® = − ­ .. ® (12.33)
­ ® ­ ® ­ ®
­ . . . . ® ­.® ­.®
𝛿 « 𝑓𝑘 ¬
­ 𝜕𝑓 𝜕 𝑓𝑘 𝜕 𝑓𝑘
®
𝑘
... « 𝑘¬
« 𝜕𝑥1 𝜕𝑥 2 𝜕𝑥 𝑘 ¬

The elements of the Jacobian can be either calculated via their analytic
expressions or via the numerical expression of the first derivatives given in
Eq. 11.9. The Newton approach to solving systems of non-linear equations
nicely illustrates how the different methods that we have considered in
earlier chapters come together. For solving systems of multiple equation,
we require both the numerical calculation of derivative and the solution
of a linear system of equations collected in a matrix equation.

12.4 Exercises

The bisection method

In this chapter we have described the bisection method that can be used
to find the root of any function. In this exercise we will implement the bi-
section method. We briefly summarize the method again for convenience.
To find the root of a function 𝑓 (𝑥) in the interval [a, b] the bisection
method works in 3 steps:

1. calculate the midpoint of the interval 𝑐 = (𝑎 + 𝑏)/2 and the value


of the function 𝑓 (𝑐)
2. if 𝑓 (𝑎) and 𝑓 (𝑐) have the same sign then the root must be in the
interval [c,b] and we set 𝑎 = 𝑐 . Otherwise we set 𝑏 = 𝑐 .
3. check whether the new interval is small enough to consider it
converged by comparing 𝜖 = |𝑏 − 𝑎| to a desired threshold. If the
desired accuracy is not reached, go back to step 1 until convergence
is reached.

Implement this method in a script called ’[Link]’ using a simple while


loop. Test the implemented method on the function:

𝑓 (𝑥) = 4𝑥 3 − 2 𝑥 2 + 4 𝑥 − 4 (12.34)

The Newton method

In this exercise we will implement the Newton method for the solution of
a non-linear equation and use the method to determine the cause of a gas
tank explosion. The Newton method has been presented in section 12.2
and a pseudo code is given on page 98. You can use this structure to
create your own implementation of the Newton’s method.
Part I : Implementation As a reminder, the Newton method allows to
find the solution of a nonlinear equation, such as 𝑓 (𝑥) = 𝑥 2 − 2 𝑥 = 0, in
an iterative approach. We therefore start with an initial guess 𝑥 0 . Most
12 Non-linear equations 105

likely, this guess does not satisfy 𝑓 (𝑥 0 ) = 0, i.e, it is not the solution we
are looking for. We therefore have to improve the quality of our solution.
To do that the Newton method relies of the iterative procedure

𝑓 (𝑥 𝑛 )
𝑥 𝑛+1 = 𝑥 𝑛 − (12.35)
𝑓 0(𝑥 𝑛 )

where 𝑓 0(𝑥 𝑛 ) is the derivative of 𝑓 (𝑥) evaluated at 𝑥 = 𝑥 𝑛 .This iterative


procedure converges toward the exact solution leading to 𝑓 (𝑥 𝑁 ) ≈ 0.
Implement the Newton method and test the implementation by finding
the solution of

𝑥 𝑥
cos(𝑥) exp(− )− =0 (12.36)
2𝜋 𝜋

Start with an initial guess of 𝑥 0 = 5. Your method should reach the


solution 𝑥 = 1.127117 in a few iterations. The Newton method applied
to this problem is illustrated in Fig. 12.3. What happen if you change the
initial guess to 𝑥 0 = −5 ?

Figure 12.3: Illustration of the functioning


of the Newton method applied to eq 12.36.
Starting from 𝑥 0 = 5 the method converges
to the solution of the equation, i.e. 𝑥 =
1.127117.

Part II : Tank Explosion A propane tank was installed in a backyard.


Unfortunately, on a particularly hot day the tank exploded. You have
been hired to calculate the number of mole 𝑛 of propane contained in the
tank just before the explosion. To do so you will use the van der Waals
equation of state

!
𝑎 𝑉

𝑃+ − 𝑏 − 𝑅𝑇 = 0 (12.37)
𝑉 2
 𝑛
𝑛

where

27 𝑅 𝑇𝐶
2 2
1 𝑅𝑇𝐶
𝑎= 𝑏= (12.38)
64 𝑃𝑐 8 𝑃𝑐
12 Non-linear equations 106

Your firm has provided you with the last reading of the tank 𝑇 = 384 K
and P = 4891.3 kPa. The volume of the tank is 𝑉 = 0.15 𝑚 3 . In addition
the critical temperature and pressure of propane are 𝑇𝑐 = 369.9 K and
𝑃𝑐 = 4254.6 kPa. Finally the gas constant is 𝑅 = 0.008314 𝑚 3 𝑘𝑃𝑎/(𝑚𝑜𝑙.𝐾).
Write a Python code called ’[Link]’ and solve the van der Waals equation
to obtain the number of mole 𝑛 in the tank at the time of the explosion.
To solve this equation, you can use the Newton method implemented
above. Check that there is only one solution of this equation by plotting
the left side of the van der Waals equation as a function of 𝑛 .
The correct answer is 𝑛 = 546.51.

Newton for two equations

Consider the following two nonlinear equations:

𝑓1 (𝑥1 , 𝑥2 ) = 3 𝑥12 + 4 𝑥22 − 1 = 0 (12.39)


𝑓2 (𝑥1 , 𝑥2 ) = 𝑥23 − 8 𝑥 13 −1=0 (12.40)

Implement a version of the Newton-code presented above on page 100 to


solve this set of equations.
The general approach to this is summarized below:
1. Write two functions that return the function values for both func-
tions
2. Write a function called ’Jacobian’ that sets up the Jacobian matrix
and takes as its argument the two functions to be solved and the
two initial guess values for 𝑥 1 and 𝑥 2 . The derivatives should be
calculated using numerical derivatives.
3. use the while-loop like in the code on page 100 to solve the system
of equations.
The solutions are: 𝑥 1 = −0.497251208 and 𝑥 2 = 0.254078592.

12.5 Extra exercises

Gradient descent in 2D

Let’s assume that we need to find the minimum of a function of two


variables: 𝑓 (𝑥, 𝑦). To do so we can start at a given point (𝑥 0 , 𝑦0 ), compute
the gradient of the function at that point ∇ 𝑓 (𝑥, 𝑦) and move ’a little bit’
in the direction of the gradient to reach a second point (𝑥 1 , 𝑦1 ). Hence
the relation to from one point ( 𝑝 ) to the next one is given by

𝑝 𝑛+1 = 𝑝 𝑛 − 𝛼∇ 𝑓 (𝑥, 𝑦)| 𝑝 𝑛 (12.41)

Implement this approach, called the gradient descent approach to find


the minimum of two functions
12 Non-linear equations 107

𝑓1 (𝑥, 𝑦) = 𝑥2 + 4 ∗ 𝑦2 𝑥, 𝑦 ∈ [−(12.42)
3; 3]
𝑥2 𝑦 2
𝑓2 (𝑥, 𝑦) = − sin( − + 3) cos(2 𝑥 + 1 − 𝑒 𝑦 ) 𝑥, 𝑦 ∈ [−3; 1]
(12.43)
2 4

In each case you will start the search for different initial guess of the
solution ( 𝑥 0 ) and report on the minimum found by the gradient de-
scent approach. You will obtain figure similar to the one represented in
Fig. 12.4.

Figure 12.4: Finding the minimum of a 2D


surface via the gradient descent approach
on two different surface a) 𝑓1 (𝑥, 𝑦) and b)
𝑓2 (𝑥, 𝑦) and using 3 different initial guess.
Ordinary differential equations
Classification of ODEs and linear
differential equations 13
It regularly happens in scientific problems that we want to evaluate how 13.1 Classification of ODEs . . . . 110
a given system, that is initially in a known state, evolves if we change one 13.2 Linear ODEs . . . . . . . . . . 112
of the parameters. A typical example is a chemical reaction where we are 13.3 Exercises . . . . . . . . . . . . . 116
Coupled first order reactions 116
usually interested in the evolution of the concentrations of the different
species in time. For example, we may want to know what happens to the
concentration of a chemical species if we suddenly introduce a certain
reactant in the batch. In such problems, we generally want to know how
fast and how important the changes are. As an example we consider the
following chemical reaction:

k1 k3
A+B↽ −⇀
−−−−− C + D −−−→ E
k2

To model this reaction we start from a mass balance : Input + Generation


= Output + Accumulation. The evolution of the five different species
in these two chemical reactions can be written in the form of a set of
differential equation, one for each species:

𝑑𝐶 𝐴
= −𝑘1 𝐶 𝐴 𝐶 𝐵 + 𝑘2 𝐶 𝐶 𝐶 𝐷 (13.1)
𝑑𝑡
𝑑𝐶 𝐵
= −𝑘1 𝐶 𝐴 𝐶 𝐵 + 𝑘2 𝐶 𝐶 𝐶 𝐷 (13.2)
𝑑𝑡
𝑑𝐶 𝐶
= 𝑘1 𝐶𝐴 𝐶𝐵 − 𝑘2 𝐶𝐶 𝐶𝐷 − 𝑘3 𝐶𝐶 𝐶𝐷 (13.3)
𝑑𝑡
𝑑𝐶 𝐷
= 𝑘1 𝐶𝐴 𝐶𝐵 + 𝑘2 𝐶𝐶 𝐶𝐷 − 𝑘3 𝐶𝐶 𝐶𝐷 (13.4)
𝑑𝑡
𝑑𝐶𝐸
= +𝑘3 𝐶 𝐶 𝐶 𝐷 (13.5)
𝑑𝑡

In this system of first-order non-linear differential equations 𝐶 𝑋 is the


concentration of 𝑋 in the reaction mixture. The solution of this system
of differential equations gives the concentrations of all the five species
as a function of time for a given set of initial conditions. These initial
conditions in this case are the concentrations 𝐶 𝐴 and 𝐶 𝐵 at the beginning
of the reaction, i.e. 𝐶 𝑋 (𝑡 = 0). Such a problem, where the initial conditions
are known is called an initial value problem. There are also situations where
the variables (𝐶 𝑋 ) are known for different values of the independent
variable (in this case 𝑡 ), for instance the desired concentrations after
a certain amount of time. Such (systems of) differential equations are
called boundary value problems. In this chapter we will first discuss the
classification of differential equations and then present a general solution
for systems of linear differential equations. The general numerical solution
of initial value problems is discussed in Chapter 14, while numerical
solutions of boundary value problems are discussed in Chapter 15.
13 Classification of ODEs and linear differential equations 110

13.1 Classification of ODEs

The system of differential equations presented above is in the ideal form


to be solved numerically. In this so-called canonical form the system of
equations contains only first derivatives and none of the terms is directly
dependent on the independent variable (i.e. 𝑡 ). Any equation where none
of the term depend on the explicit variable is called autonomous. The
system does contain terms such as 𝐶 𝑋 𝐶𝑌 and is therefore nonlinear. More
complicated ODE’s are often encountered in chemistry and chemical
engineering. Several examples of differential equations are given below.

𝑑𝑦
1st order = 𝑓 (𝑦) (13.6)
𝑑𝑡
𝑑2 𝑦 𝑑𝑦
2nd order +𝑦 = 𝑒 −𝑡 (13.7)
𝑑𝑡 2 𝑑𝑡
𝑑3 𝑦 𝑑2 𝑦 𝑑𝑦
3rd order + 𝑎 +𝑏 +𝑦=0 (13.8)
𝑑𝑡 3 𝑑𝑡 2 𝑑𝑡

The first of these equations is the general form of a first order differential
equation, containing only first derivative terms. Eq. 13.7 contains a
second derivative and is therefore of the 2nd order. In addition, one of the
terms 𝑒 −𝑡 depends explicitly on the independent variable. Any equation
where one term depends on the explicit variable is called non-autonomous.
Similarly, eq. 13.8 contains not only a second derivative but also a third
derivative and is therefore of the third order. Numerical integration of
ODEs is easier and computationally more efficient if the system is in a
canonical form: first-order only and autonomous. Usually, higher-order
and non-autonomous equation can be recast in a canonical form by
substitution, i.e. replacing the variables used by another.
Consider the third-order eq. 13.8. To transform this high-order ODE in a
series of coupled first-order ODEs we can introduce new variables as:

𝑑𝑦
𝑦1 ≡ (13.9)
𝑑𝑡
𝑑𝑦1 𝑑2 𝑦
𝑦2 ≡ = 2 (13.10)
𝑑𝑡 𝑑𝑡

𝑑𝑦2 𝑑3 𝑦
and therefore 𝑑𝑡 ≡ 𝑑𝑡 3
. Using these new variables Eq. 13.8 becomes:

𝑑𝑦
= 𝑦1 (13.11)
𝑑𝑡
𝑑𝑦1
= 𝑦2 (13.12)
𝑑𝑡
𝑑𝑦2
= −𝑎 𝑦2 − 𝑏 𝑦1 − 𝑦 (13.13)
𝑑𝑡

and we have succeeded in transforming a high-order ODE in a series of


coupled first order ODE’s.
13 Classification of ODEs and linear differential equations 111

A similar approach can be taken for non-autonomous equations. Consider


Eq. 13.7. This equation is of the second order and is non-autonomous.
To transform this equation in a canonical form we can introduce the
following variables:

𝑑𝑦
𝑦1 ≡ (13.14)
𝑑𝑡
𝑦2 ≡ 𝑒 −𝑡 (13.15)

With these substitutions we obtain the system of equations:

𝑑𝑦
= 𝑦1 (13.16)
𝑑𝑡
𝑑𝑦1 𝑑2 𝑦
≡ = −𝑦 𝑦1 + 𝑦2 (13.17)
𝑑𝑡 𝑑𝑡 2
𝑑𝑦2
≡ −𝑒 −𝑡 = −𝑦2 (13.18)
𝑑𝑡

Once again we have reduced a higher-order non-autonomous ODE in a


canonical form that is much easier to solve numerically. It is important
to note that these equations are still nonlinear due to the term 𝑦 𝑦1 in the
𝑑𝑦1
equation for 𝑑𝑡 .

The substitution of variables may seem like an artificial approach. There-


fore it is interesting to look at this using a practical example. In Eq13.20
Newton’s first equation for the movement of a mass under the influence
of gravity is given. This equation can be solved if two initial conditions
are known; the position, 𝑥(𝑡 = 0) and the initial velocity, 𝑣 0 .

𝑑2 𝑥
= −𝑔 (13.19)
𝑑𝑡 2
(13.20)

By explicitly introducing the velocity as a variable, we can rewrite Eq.13.20


into a system of two linear differential equations:

𝑑𝑥
= 𝑣 (13.21)
𝑑𝑡
𝑑𝑣
= −𝑔 (13.22)
𝑑𝑡
(13.23)

These two equations can be solved using the same two initial values
mentioned above.
13 Classification of ODEs and linear differential equations 112

13.2 Linear ODEs

Many problems in physical chemistry can be described by a coupled


set of linear ordinary equation with constant coefficients. A common
example can be found in the form of first-order chemical reactions where
one species converts into another with a certain rate. Consider as an
example the following reaction scheme:

k1 k3
A↽ −⇀
−−−−− B ↽ −⇀
−−−−− C
k2 k4

The evolution of the concentrations of the four species involved in this


series of first order reaction can be described by a set of linear first order
differential equations in terms of the rate constants, 𝑘 𝑛 :

𝑑𝐶 𝐴 (𝑡)
= −𝑘 1 𝐶 𝐴 + 𝑘2 𝐶 𝐵 (13.24)
𝑑𝑡
𝑑𝐶 𝐵 (𝑡)
= 𝑘 1 𝐶 𝐴 + 𝑘4 𝐶 𝐶 − (𝑘2 + 𝑘 3 )𝐶 𝐵 (13.25)
𝑑𝑡
𝑑𝐶 𝐶 (𝑡)
= −𝑘 4 𝐶 𝐶 + 𝑘 3 𝐶 𝐵 (13.26)
𝑑𝑡
(13.27)

Just as we have seen in Chapter 9 for systems of linear equations, we can


assemble these four differential equations into a matrix form:

𝑑𝐶 𝐴
𝑑𝑡 ª
© 𝑑𝐶 −𝑘 1 𝑘2 0 𝐶𝐴
­ 𝐵 ® = ©­ 𝑘1 −(𝑘2 + 𝑘3 ) 𝑘4 ª® · ©­ 𝐶 𝐵 ª® (13.28)
­ 𝑑𝑡 ®
𝑑𝐶 𝐶 0 𝑘3 −𝑘4 ¬ «𝐶 𝐶 ¬
« 𝑑𝑡 ¬ «
or using a more compact notation

𝑑
C=K·C (13.29)
𝑑𝑡

To obtain an idea on what the solution of this matrix equation may look
like we will first recall the solution of a single first order linear differential
equation: equation of the type

𝑑
𝑦 = 𝑘𝑦 (13.30)
𝑑𝑡

with the initial condition 𝑦(𝑡0 = 0) = 𝑦0 . The solution for this equation
can be derived by separating the variables and integrating both sides of
the equation:

𝑦 𝑡
𝑑𝑦 𝑦
∫ ∫
= 𝑘𝑑𝑡 −→ 𝑙𝑛 = 𝑘𝑡 −→ 𝑦 = 𝑒 𝑘𝑡 𝑦0 (13.31)
𝑦0 𝑦 𝑡0 𝑦0
The solution of the matrix equation 13.29 takes a similar form when we
introduce the exponential of a matrix. The exponential of a matrix is
13 Classification of ODEs and linear differential equations 113

defined in such a way that the solution of the set of differential equation
assembled in the matrix equation 13.29 takes the form:

C = 𝑒 K 𝑡 C0 (13.32)

In Appendix A it is demonstrated that a solution of this form satisfies


the matrix form of the differential equations in Eq. 13.29. The remaining
problem is calculation of the exponential of a matrix, 𝑒 K𝑡 . As demonstrated
in Appendix A, the exponential of a matrix is given by:

𝑒 K𝑡 = U𝑒 𝚲𝑡 U−1 (13.33)

where U is the matrix of the eigenvectors of K and 𝚲 is a diagonal matrix


that contains the 𝑒 𝑖 𝑔𝑒𝑛𝑣𝑎𝑙𝑢𝑒 𝑠 (𝜆1 , 𝜆2 , ...., 𝜆 𝑁 ) of K. The exponential of
such a diagonal matrix is then simply:

𝑒 𝜆1 𝑡 0 0 ... 0
0 𝑒 𝜆2 𝑡 0 ... 0
© ª
­ ®
0 0 𝑒 𝜆2 𝑡 ... 0
­ ®
𝑒 𝚲𝑡 = ­­ ® (13.34)
.. .. .. .. .. ®
­
­ . . . . .
®
®
« 0 0 0 ... 𝑒 𝜆𝑁 𝑡 ¬
It is important to note here that the element by element exponentiation
only gives the correct matrix exponential for a diagonal matrix. For
non-diagonal matrices this is not the case! The calculation of the matrix
exponential in Eq. 13.33 can easily be done in a Python program by
first determining the eigenvectors and eigenvalues of the matrix, and
subsequently calculating the relevant matrix products. However, as can
be expected the module scipy contains a function [Link]() that
returns the exponential of a matrix that it is given as an argument. These
two approaches to calculate the matrix exponential are compared in the
code below:
1 import numpy as np
2 import [Link] as scln
3 import time
4

5 # create a matrix
6 N = 500
7 A = 0.1*[Link](N,N)
8 t = 0.01
9

10 # compute the exponential with scipy


11 t0 = time.process_time()
12 expA_scipy = [Link](A)
13 t1 = time.process_time()
14

15 # compute the exponential directly


16 t2 = time.process_time()
17 Lambda, U = [Link](A)
18 expA_direct = [Link](U,[Link]([Link]([Link](Lambda)),[Link](U)))
19 t3 = time.process_time()
20

21 # compute the error and print the performance


22 error = [Link]([Link](expA_scipy-expA_direct))/N/N
23

24 if error < 1E-3:


25 print(’Scipy %1.6f sec. \t Direct %1.6f’ %((t1-t0),(t3-t2)))
13 Classification of ODEs and linear differential equations 114

26 else:
27 print(’error %1.6e’ %(error))

Not surprisingly, the use of the scipy function results in a much faster
evaluation of the matrix exponential; being almost ten times faster for a
4x4 matrix.
To see how this works in practice we will return to our system of first
order chemical reactions presented above. The solution Eq. 13.32 can be
used to calculate the evolution of the concentrations of 𝐴, 𝐵 and 𝐶 as a
function of time. As a starting point, we need to know the initial value of
each of the concentrations, collected in the vector C0 . We assume that
these concentrations are given by:

𝐶 𝐴 (0) = 1 𝐶 𝐵 (0) 𝐶 𝐶 (0) (13.35)

For this example the rate constants of all four equations are known:

𝑘 1 = 1 min−1 𝑘2 = 0 min−1 𝑘3 = 2 min−1 𝑘 4 = 3 min−1


(13.36)
If we collect all this information in the matrix Eq. 13.32, we obtain the
following expression for the concentration at time 𝑡 𝑝 in the vector C𝑡 𝑝 :

𝐶 𝐴 (𝑡 𝑝 )  −1 0 0  1
­ 𝐶 𝐵 (𝑡 𝑝 ) ® = exp ­ 1 −2 3 ® 𝑡 𝑝  · ­0®
© ª  © ª
(13.37)
© ª 
«𝐶 𝐶 (𝑡 𝑝 )¬ 2 −3¬  «0¬
 0
«

While it is interesting to calculate the concentration for a single time,


it is usually more interesting to plot the evolution of the different con-
centrations as a function of time. This is easily possible for the matrix
exponentiation approach as the matrix Eq. 13.32 hold for any value of 𝑡 .
Therefore, we can divide the total evolution time that we are interested in
in small intervals of same length (Δ𝑡 ): 𝑇 = [0 , Δ𝑡, 2Δ𝑡, 3Δ𝑡, ...., 𝑁𝑡𝑜𝑡 Δ𝑡 ],
and then calculate the solution of the equation for each time value. To
avoid recalculating the matrix exponential at each time-step, we can use
a convenient property of matrices. If we want the solution at 𝑡 𝑛 = 𝑛Δ𝑡 .
The solution at this time can be written as:

𝑛
C(𝑡 𝑛 = 𝑛Δ𝑡) = 𝑒 K𝑛Δ𝑡 C0 = 𝑒 KΔ𝑡 C0 (13.38)

Therefore, we only need to calculate 𝑒 KΔ𝑡 and all multiples of the time
step only requires and additional multiplication. For a whole series of
time steps between 𝑡0 = 0Δ𝑡 and 𝑡 𝑓 = 𝑁𝑡𝑜𝑡 Δ𝑡 , we can therefore use the
recurrence relationship:

C𝑛+1 = 𝑒 KΔ𝑡 C𝑛 (13.39)

where C𝑛 ≡ C(𝑡 𝑛 = 𝑛Δ𝑡).


In the Python code below we illustrate how to implement the problem
that we have just discussed and the resulting graph is shown in Fig. 13.1.
13 Classification of ODEs and linear differential equations 115

1 import numpy as np
2 import [Link] as plt
3 import [Link] as scln
4

5 # rate constants
6 k1 = 1
7 k2 = 0
8 k3 = 2
9 k4 = 3
10

11 # initial condition
12 C0 = [Link]([1,0,0])
13

14 # evolution time
15 tmax = 5
16 nT = 250
17 T = [Link](0,tmax,nT)
18 dT = T[1]-T[0]
19

20 # matrix of the rates


21 K = [Link]([[-1,0,0],[1,-2,3],[0,2,-3]])
22

23 # Calculate the exponential of a matrix using scipy


24 eK = K * dT
25 eKdt = [Link](eK)
26

27 # store the solution


28 C = []
29 [Link](C0)
30

31 for i in range(1,nT):
32 C0 = [Link](eKdt,C0)
33 [Link](C0)
34

35 #print C
36 C = [Link](C)
37 [Link](T,C[:,0],linewidth=2,label=’C_A’)
38 [Link](T,C[:,1],linewidth=2,label=’C_B’)
39 [Link](T,C[:,2],linewidth=2,label=’C_C’)
40

41 [Link](’time (min)’,fontsize=12)
42 [Link](’Concentration’,fontsize=12)
43 [Link](loc=1)
44 [Link]()

As we can see in this graph, the solution obeys the ’conversation of


mass’ principle with 𝐶 𝐴 (𝑡) + 𝐶 𝐵 (𝑡) + 𝐶 𝐶 (𝑡) = 𝑐𝑡𝑒 . In addition, the
asymptotic solution follows the equilibrium condition 𝐶 𝐵 /𝐶 𝐶 = 𝑘 4 /𝑘 3 .
After obtaining a solution for a set of reactions it is always important to
check whether it obeys common ’chemical’ sense by considering such
mass-balance and equilibrium arguments since it is easy to see when a
solution is totally wrong.
13 Classification of ODEs and linear differential equations 116

Figure 13.1: Concentration profiles of the


a chemical reaction model by the equation
13.37.

13.3 Exercises

Coupled first order reactions

In this exercise we will use the exponentiation method to calculate how


the concentrations of different species evolve as a function of time. The
reaction scheme below shows a series of coupled first order reactions.
The starting product, 𝐴 transforms into 𝐵. 𝐵 is in equilibrium with 𝐶 ,
which finally reacts to form 𝐷 .

k1 k2 k4

A −−−→ B −
−−−⇀
−− C −−−→ D
k3

Initially, we only add the starting material 𝐴, implying that the concen-
tration 𝐶 𝐴 = 1.0, while the concentration of the other four species are
zero. The reaction rate constants are given by:

𝑘1 = 2.5min−1
𝑘2 = 0.5min−1
𝑘3 = 0.3min−1
𝑘4 = 0.1min−1

The goal of this exercise is to make a plot of the concentrations of the four
species as a function of time. To achieve this take the following steps:
1. Set up the system of differential equations that describe the reac-
tions.
2. Collect the differential equations in a matrix form.
3. Calculate the exponential of the matrix using the [Link]()
function.
4. Calculate the concentrations of 𝐴, 𝐵, 𝐶 and 𝐷 during the first 10
minutes in 250 steps using a for-loop.
13 Classification of ODEs and linear differential equations 117

5. Plot the concentrations as a function of time.


You can play around with the end-time to see if eventually all of 𝐴 is
transformed into 𝐷 .
Initial value problems 14
In the last chapter we have classified differential equation and have 14.1 Forward Euler . . . . . . . . . 119
looked at a specific form, the linear first order differential equations that 14.2 Modified Euler . . . . . . . . . 120
can be solved exactly by using exponentiation of matrices. However, in 14.3 Runge-Kutta methods . . . . 120
14.4 Numerical VS Exact . . . . . 122
many cases it is not possible to write a system of coupled differential
14.5 ODE-solver in 𝑠𝑐𝑖𝑝𝑦 . . . . . 123
equations in a matrix form such as Eq. 13.37. This is for example the case
Simultaneous ODEs . . . . . 124
if the equations contain nonlinear terms, which regularly happens for
The predator-prey model . . 124
rate equations of chemical reactions. This is for instance obvious in Eq. 14.6 Exercises . . . . . . . . . . . . . 126
13.5 where several terms depend on products of different concentrations. Forward Euler . . . . . . . . . 126
Luckily, many numerical methods are available to solve such systems of Runge-Kutta 4-th order . . . 127
non-linear differential equations. We will consider a form of differential Coupled first order reactions 127
equations that can be written in a canonical form where the first derivative Epidemic propagation . . . . 128
of 𝑦 is a function 𝑓 (𝑡, 𝑦 :

𝑑y
= f(𝑡, y) (14.1)
𝑑𝑡

In this chapter we consider so-called initial value problem, which means


that we know the initial conditions:

y(𝑡0 ) = y0 (14.2)

As a starting point we first consider the case where only one differ-
ential equation defines the system to illustrate the numerical solution
of non-linear differential equations in general. These approaches can
easily be extended to multiple simultaneous nonlinear ODEs. In the
numerical solution of nonlinear ODEs the goal is to obtain for instance
the concentrations of different reaction species as a function of time. In
such cases, ’time’, 𝑡 is called the independent variable. A first step is the
discretization of the independent variable (here 𝑡 ) in a series of intervals:
𝑡 = [𝑡0 , 𝑡1 , ...𝑡 𝑖 , ...𝑡 𝑁 ]. The solution is then obtained as a series of values
𝑦 = [𝑦0 , 𝑦1 , ...𝑦 𝑖 , ...𝑦 𝑁 ] and if the interval becomes small enough we get a
continuous solution for the system of differential equations. The solution
of the general differential equation above can be obtained by rearranging
the terms in equation 14.1 and integrating on both sides:

∫ 𝑦 𝑖+1 ∫ 𝑡 𝑖+1
𝑑𝑦 = 𝑓 (𝑡, 𝑦)𝑑𝑡 (14.3)
𝑦𝑖 𝑡𝑖
∫ 𝑡 𝑖+1
𝑦 𝑖+1 − 𝑦 𝑖 = 𝑓 (𝑡, 𝑦)𝑑𝑡 (14.4)
𝑡𝑖
∫ 𝑡 𝑖+1
𝑦 𝑖+1 = 𝑦 𝑖 + 𝑓 (𝑡, 𝑦)𝑑𝑡 (14.5)
𝑡𝑖
14 Initial value problems 119

This result shows that as long as we know 𝑦 𝑖 , we can always calculate


the next y-value, 𝑦 𝑖+1 by integrating the function 𝑓 (𝑡, 𝑦) over the time
interval 𝑑𝑡 . We have already seen several approaches to numerically
integrate functions in Chapter 10 and we can in principle use any of
these approaches to solve differential equations. In the following sections
we will consider a few integration schemes that give approximations of
different quality.

14.1 Forward Euler

The simplest approximation to calculate the integral over a function


between two points is the rectangle approximation that we encountered
in Chapter 10. In this approximation it is assumed that, to a first approxi-
mation, the function value is constant over the interval and the integral
is just the function value 𝑦 𝑖 times the width of the interval, ℎ :

∫ 𝑡 𝑖+1
𝑓 (𝑡, 𝑦)𝑑𝑡 = ℎ 𝑓 (𝑡 𝑖 , 𝑦 𝑖 ) (14.6)
𝑡𝑖

Introducing this expression in 14.5 leads to the so-called forward Euler


method for the solution of ODEs

𝑦 𝑖+1 = 𝑦 𝑖 + ℎ 𝑓 (𝑡 𝑖 , 𝑦 𝑖 ) (14.7)

The forward Euler approach can be understood graphically as shown


Fig.14.1 where true value of 𝑦 is given as a function of 𝑡 . In order to obtain
𝑑𝑦
an estimate for 𝑦 𝑖+1 we take the derivative 𝑑𝑡 at the 𝑖 -th point (which is
given by 𝑓 (𝑦 𝑖 , 𝑡 𝑖 ) as defined by the differential equation itself) and make
a step of length ℎ in the direction of the derivative. In other words, we
know the function value in the 𝑖 -th point and we know its derivative,
so by assuming the function is linear over a small interval we obtain
𝑦 𝑖+1 . The accuracy of the Euler method is of the same order than the
rectangle approximation for the integral, i.e. 𝑂(ℎ). The Euler equation
is an explicit equation, meaning that 𝑦 𝑖+1 appears only on the left-hand
side of the equation.

Figure 14.1: Illustration of the forward-


Euler method to integrate a differential
equation.
14 Initial value problems 120

14.2 Modified Euler

Improved approaches to integrating functions that we have seen in


Chapter 10 are the mid-point and the trapezoidal approximation. Both
have improved accuracy and can be used to calculate the integral in
Eq. 14.5. In this section we will use the trapezoid rule as an example. As
a reminder, in the trapezoid approach the integral over an interval ℎ is
given by:

𝑡 𝑖+1


𝑓 (𝑡, 𝑦)𝑑𝑡 = 𝑓 (𝑡 𝑖 , 𝑦 𝑖 ) + 𝑓 (𝑡 𝑖+1 , 𝑦 𝑖+1 )

(14.8)
𝑡𝑖 2

Introducing this approximation in Eq. 14.5, we obtain the equation for


the so-called modified Euler method


𝑦 𝑖+1 = 𝑦 𝑖 + 𝑓 (𝑡 𝑖 , 𝑦 𝑖 ) + 𝑓 (𝑡 𝑖+1 , 𝑦 𝑖+1 )

(14.9)
2

This equation is implicit as 𝑦 𝑖+1 appears on both sides of the equation.


We can make this equation explicit by making a prediction for the term
𝑦 𝑖+1 that appears on the right hand side of Eq. 14.9. Such a prediction
can be obtained using the forward Euler approach given by Eq. 14.7.

(𝑦 𝑖+1 )𝑃𝑟 = 𝑦 𝑖 + ℎ 𝑓 (𝑡 𝑖 , 𝑦 𝑖 ) (14.10)

We can then use this prediction to obtain a corrected value for 𝑦 𝑖+1 :


(𝑦 𝑖+1 )𝐶𝑜𝑟 = 𝑦 𝑖 + 𝑓 (𝑡 𝑖 , 𝑦 𝑖 ) + 𝑓 𝑡 𝑖+1 , (𝑦 𝑖+1 )𝑃𝑟

(14.11)
2

If we consider what this modified Euler methods means in terms of


the graphical representation in Fig.14.9, we come to the conclusion
that in Eq.14.11 we use the average of the slopes of the function at the
start and the end of the integration interval. As we know from the
calculation of the derivative, such a centered approximation is a much
better description of the derivative over the interval, and hence give a
considerable improvement over the simple forward Euler method.

14.3 Runge-Kutta methods

The modified Euler method described in the previous section is the


first in a series of improvements that all trace back to a progressive
improvement of the estimate of the average slope of the curve over
the integration interval. Such methods form a wide class of so-called
Runge-Kutta methods to solve ODEs and these are said to be second,
third, fourth, etc. order. The RK methods are an extension of the modified
Euler method presented above. To illustrate this, we rewrite Eq. 14.11
as:
14 Initial value problems 121

ℎ ℎ
𝑦 𝑖+1 = 𝑦𝑖 + 𝑓 (𝑡 𝑖 , 𝑦 𝑖 ) + 𝑓 (𝑡 𝑖+1 , 𝑦 𝑖+1 ) (14.12)
2 2
= 𝑦𝑖 + 𝑤1 𝑘1 + 𝑤2 𝑘2 (14.13)

where in this case:

1
𝑤1 = 𝑘1 = ℎ 𝑓 (𝑡 𝑖 , 𝑦 𝑖 ) (14.14)
2
1
𝑤2 = 𝑘2 = ℎ 𝑓 (𝑡 𝑖 + ℎ, 𝑦 𝑖 + 𝑘 1 ) (14.15)
2

since 𝑦 𝑖+1 is approximated by 𝑦 𝑖 + ℎ 𝑓 (𝑡 𝑖 , 𝑦 𝑖 ) in the prediction stage of


the modified Euler method. The equations presented above define the
second-order Runge-Kutta methods. This approach can be generalized
to any order:

𝑚
X
𝑦 𝑖+1 = 𝑦 𝑖 + 𝑤𝑖 𝑘𝑖 (14.16)
𝑖=1

Similarly to the second-order RK methods the 𝑘 𝑖 are defined by :

𝑘1 = ℎ 𝑓 (𝑡 𝑖 , 𝑦 𝑖 ) (14.17)
𝑘2 = ℎ 𝑓 (𝑡 𝑖 + 𝑐 2 ℎ, 𝑦 𝑖 + 𝑎21 𝑘1 ) (14.18)
𝑘3 = ℎ 𝑓 (𝑡 𝑖 + 𝑐 3 ℎ, 𝑦 𝑖 + 𝑎31 𝑘1 + 𝑎32 𝑘2 ) (14.19)
... (14.20)
𝑘𝑚 = ℎ 𝑓 (𝑡 𝑖 + 𝑐 𝑚 ℎ, 𝑦 𝑖 + 𝑎 𝑚 1 𝑘1 + 𝑎 𝑚 2 𝑘2 + ... + 𝑎 𝑚,𝑚−1 𝑘 𝑚−1()14.21)

or in a more compact notation:

𝑝−1
X
𝑘 𝑝 = ℎ 𝑓 (𝑡 𝑖 + 𝑐 𝑝 ℎ, 𝑦 𝑖 + 𝑎 𝑝𝑙 𝑘 𝑙 ) (14.22)
𝑙=1

The number of terms, 𝑚 , in the expansion defines the degree of complexity


and the accuracy of the method. Each degree of complexity is entirely
defined by the series of coefficients 𝑐 𝑖 and 𝑎 𝑖𝑗 .
The derivation of these coefficients is no straightforward but has been
performed for many orders. For example the fourth-order RK method is
defined by the equations:

1
𝑦 𝑖+1 = 𝑦 𝑖 + (𝑘1 + 2 𝑘2 + 2 𝑘3 + 𝑘4 ) (14.23)
6

with:
14 Initial value problems 122

𝑘1 = ℎ 𝑓 (𝑡 𝑖 , 𝑦 𝑖 ) (14.24)
ℎ 𝑘1
 
𝑘2 = ℎ 𝑓 𝑡𝑖 + , 𝑦𝑖 + (14.25)
2 2
ℎ 𝑘2
 
𝑘3 = ℎ 𝑓 𝑡𝑖 + , 𝑦𝑖 + (14.26)
2 2
𝑘4 ℎ 𝑓 𝑡 𝑖 + ℎ, 𝑦 𝑖 + 𝑘3

= (14.27)

This method is probably the most used technique in practice to solve ODEs
and it can be programmed in a relatively simple way. The implementation
of the fourth-order RK method is the subject of one of the exercises in
this chapter.

14.4 Numerical VS Exact

To illustrate the performances of the forward and modified Euler methods,


we consider the simplest chemical reaction possible

k
A −−−→ B

The differential equations that describe the concentrations of 𝐴 and 𝐵


are then given by:

𝑑 𝑑
𝐶 𝐴 = −𝑘𝐶 𝐴 𝐶 𝐵 = 𝑘𝐶 𝐴 (14.28)
𝑑𝑡 𝑑𝑡

Using the initial conditions 𝐶 𝐴 (0) = 1 and 𝐶 𝐵 (0) = 0, the exact solutions
of these ODEs can be readily derived

𝐶 𝐴 (𝑡) = 𝑒 −𝑘𝑡 𝐶 𝐵 (𝑡) = (1 − 𝑒 −𝑘𝑡 ) (14.29)

We can therefore assess the accuracy of the Euler methods by compar-


ing their results to the analytic solution shown above. In the Python
code below the different methods to integrate differential equations are
implemented:
1 import numpy as np
2

3 # the euler method


4 def forward_euler(func, yi, ti, dt):
5 return yi + dt*func(yi,ti)
6

7 # the modified euler method


8 def modified_euler(func, yi, ti, dt):
9 y_pr = yi + dt*func(yi,ti)
10 return yi + 0.5*dT*(func(yi,ti)+func(y_pr,ti+dt))
11

12 # 4th order RK
13 def RK4(func, yi, ti, dt):
14 k1 = dt*func(yi,ti)
15 k2 = dt*func(yi+0.5*k1,ti+0.5*dt)
16 k3 = dt*func(yi+0.5*k2,ti+0.5*dt)
17 k4 = dt*func(yi+k3,ti+dt)
14 Initial value problems 123

18 return yi + 1./6*( k1+2*k2+2*k3+k4 )


19

20 # variation of CA
21 def df(ca,t):
22 return -k*ca

The quality of the numerical approximations are compared to the exact


solution in Fig.14.2

Figure 14.2: Comparison of the numerical


and exact solution for the concentration
of 𝐴 during the simple chemical reaction
𝐴 → 𝐵. The solution were obtained for 2
different number of points along the time
axis 𝑁 = 10 and 𝑁 = 500.

14.5 ODE-solver in 𝑠𝑐𝑖𝑝𝑦

For the understanding of approaches to integrate ODEs it is very in-


structive to program these relatively simple methods ourselves. However
many more methods have been implemented already in scipy, which
can be used to solve ODEs very efficiently. In this section we mention
one of the available functions; [Link](). In the Python code
presented below we illustrate how to use this function. It should be noted
that there are many more approaches in scipy and other modules and if
needed documentation of these and other methods can easily be found
online, together with examples on their use.
1 import numpy as np
2 from [Link] import solve_ivp
3 import [Link] as plt
4

5 # variation of CA
6 def df(t,y,k=0.2):
7 return -k*y # return the value
8

9 # time argument
10 t_start = 0.0 # min
11 t_end = 30.0 # min
12

13 tmax,nT = 30.0,500
14 T = [Link](0,tmax,nT)
15

16 # initial condition
17 C0, T0 = 1.0, 0.0
18

19 # storage
20

21 t_span = [Link]([t_start, t_end])


22

23 result = solve_ivp(df,t_span,[Link]([C0]), t_eval=T, method=’RK45’)


24

25 [Link](result.t, result.y[0],linewidth=2)
26 [Link](’time (min)’,fontsize=12)
27 [Link](’Concentration’,fontsize=12)
28 [Link]()
14 Initial value problems 124

In the code above we used the method referred to as ’RK45’ which is a


Runge-Kutta method but many other methods are available also.

Simultaneous ODEs

We have so far only discussed the different integration methods for ODEs
by applying them to one equation at a time. The same methods can be
extended in a straightforward way to deal with systems coupled ODEs.
To illustrate this, we consider the following set of 𝑛 ODEs

𝑑𝑦1
= 𝑓1 (𝑡, 𝑦1 , 𝑦2 , ...𝑦𝑛 ) (14.30)
𝑑𝑡
𝑑𝑦2
= 𝑓2 (𝑡, 𝑦1 , 𝑦2 , ...𝑦𝑛 ) (14.31)
𝑑𝑡
... (14.32)
𝑑𝑦𝑛
= 𝑓𝑛 (𝑡, 𝑦1 , 𝑦2 , ...𝑦𝑛 ) (14.33)
𝑑𝑡

This system of coupled differential equations can be written in a vector


form as 𝑑𝑡𝑑 Y = F(𝑡, Y). If we for instance want to apply the forward Euler
method to solve this system starting from the initial condition contained
in a vector Y0 , containing all the initial values of the 𝑦 𝑖 at t=0, we can
directly use Eq. 14.7 and write

Y𝑖+1 = Y𝑖 + ℎ F(𝑡, Y𝑖 ) (14.34)

Any other methods presented above for a single ODE can be used to
solve a set of simultaneous ODEs and we will explore this in one of the
exercises at the end of this chapter.

The predator-prey model

To illustrate the solution of simultaneous coupled ODEs, we consider one


of the most famous initial value problem: the predator-prey problem, also
known as the Lotka-Volterra equations. In this problem the population of
two species, a predator (𝐴) and a prey (𝐵), are studied. The population-
change of predator is given by:

𝑑𝐴
= 𝛼𝐴𝐵 − 𝛽𝐴 (14.35)
𝑑𝑡

where 𝛼 is the growth rate due to prey consumption and 𝛽 is the death
rate by overpopulation. The population change of prey-animals is given
by

𝑑𝐵
= 𝛾𝐵 − 𝜇𝐴𝐵 (14.36)
𝑑𝑡

where 𝛾 is the birth rate and 𝜇 the death rate due to the predator. We can
use any of the methods presented above to study the dynamics of these
populations. The code below shows how to do this using the fourth-order
14 Initial value problems 125

RK methods and comparing it to the scipy integration function. The


solutions found for the populations are shown in Fig. 14.3, showing an
oscillatory trend.
1 import numpy as np
2 import [Link] as plt
3 from [Link] import solve_ivp
4

5 # 4th order RK
6 def RK4(func, yi, ti, dt):
7 k1 = dt*func(ti,yi)
8 k2 = dt*func(ti+0.5*dt,yi+0.5*k1)
9 k3 = dt*func(ti+0.5*dt,yi+0.5*k2)
10 k4 = dt*func(ti+dt,yi+k3)
11 return yi + 1./6*( k1+2*k2+2*k3+k4 )
12

13 # population change
14 def df(t,Y):
15 Yp = [Link](2)
16 Yp[0] = alpha*Y[0]*Y[1] - beta*Y[0]
17 Yp[1] = gamma*Y[1] - mu*Y[0]*Y[1]
18 return Yp
19

20 # time
21 tmax, nT = 10.0,500
22 T = [Link](0,tmax,nT)
23 dT = T[1]-T[0]
24 # parameters
25 alpha, beta, gamma, mu = 1.5,2.0,2.0,1.0
26 # initial population
27 A0, B0 = 2.0, 0.2
28 T0 = 0.0
29

30 # storage
31 POP = [Link]((nT,2))
32 POP[0,:] = [A0,B0]
33

34 # compute the population dynamics


35 for iT in range(1,nT):
36 POP[iT,:] = RK4(df, POP[iT-1,:], T[iT-1], dT)
37

38

39

40 # scipy approach
41 t_span = [Link]([0., tmax])
42

43 result =solve_ivp(df, t_span, [Link]([A0, B0]), t_eval=T, method=’RK45’)


44

45 [Link](result.t, result.y[0],linewidth=2, label=’Predator’)


46 [Link](result.t, result.y[1],linewidth=2, label=’Prey’)
47 [Link](’time (min)’,fontsize=12)
48 [Link](’Population’,fontsize=12)
49 [Link](loc=2)
50 [Link]()
14 Initial value problems 126

Figure 14.3: Solution of the predator prey


model. The population of both species
present a cyclic evolution: when there are
too many predators they starve until the
population of prey is restored and so on.

14.6 Exercises

Forward Euler

In this first exercise you will implement the forward Euler method for
the solution of ODE’s. As a reminder, the forward Euler method allows
finding the solution of the differential equation

𝑑𝑦
= 𝑓 (𝑡, 𝑦) (14.37)
𝑑𝑡

with the initial solution 𝑦(𝑡0 ) = 𝑦0 by the recurrence

𝑦 𝑖+1 = 𝑦 𝑖 + ℎ 𝑓 (𝑡 𝑖 , 𝑦 𝑖 ) (14.38)

To implement this approach you will write a function, for example called
forwardEuler() that returns the value of 𝑦 𝑖+1 on taking the value of 𝑦 𝑖 as an
argument. Subsequently, this function can be called in a for loop running
over all the time steps desired. To test this program, calculate the solution
of the simple reaction

𝑘
𝐴→
− 𝐵 (14.39)

with the initial condition 𝐶 𝐴 (0) = 1 and 𝐶 𝐵 (0) = 0 (where 𝐶 𝑋 (𝑡) is the
concentration of 𝑋 at time 𝑡 ) and a value of 𝑘 = 0.2 min−1 . Simulate the
dynamics of this reaction over a period of 30 min.
The differential equations for the evolution of 𝐶 𝐴 and 𝐶 𝐵 can be easily
derived and solved analytically. You can therefore test your program by
comparing the solution it provides with the exact solution that you will
have derived analytically. Vary the number of time step you use in the
simulations of the reaction. Use between 10 and 500 and comment on the
accuracy of the method.
14 Initial value problems 127

Runge-Kutta 4-th order

In this follow up exercise we will update our program by implementing


the RK4 method. The equation defining the RK-4 methods are given by

1
𝑦 𝑖+1 = 𝑦 𝑖 + (𝑘1 + 2 𝑘2 + 2 𝑘3 + 𝑘4 ) (14.40)
6

with:

𝑘1 =
ℎ 𝑓 (𝑡 𝑖 , 𝑦 𝑖 )
𝑘2 =
ℎ 𝑘1
 
ℎ 𝑓 𝑡 𝑖 + , 𝑦𝑖 +
2 2
(14.41)
𝑘3 =
ℎ 𝑘2
 
ℎ 𝑓 𝑡 𝑖 + , 𝑦𝑖 +
2 2
𝑘4 =
ℎ 𝑓 𝑡 𝑖 + ℎ, 𝑦 𝑖 + 𝑘3


As previously we write a function that returns the value of 𝑦 𝑖+1 from 𝑦 𝑖


and uses this function in a for-loop to obtain the solution of the ODEs.
We will once again use our program to compute the concentrations 𝐶 𝐴 (𝑡)
𝑘
and 𝐶 𝐵 (𝑡) of the reactions 𝐴 →
− 𝐵 under the initial condition 𝐶 𝐴 (0) = 1
and 𝐶 𝐵 (0) = 0. Vary the number of time step you use in the simulations
of the reaction between 10 and 500 time steps and compare the accuracy
to the forward Euler method.

Coupled first order reactions

In this exercise we consider the same set of coupled linear reactions as in


section 13.3 for the reaction scheme below, but this time we will use the
Runge-Kutta function that we have programmed above to integrate the
set of equations.

k1 k2 k4

A −−−→ B −
−−−⇀
−− C −−−→ D
k3

As a reminder, we only add the starting material 𝐴, implying that the


concentration 𝐶 𝐴 = 1.0, while the concentration of the other four species
are zero. The reaction rate constants are given by:

𝑘1 = 2.5min−1
𝑘2 = 0.5min−1
𝑘3 = 0.3min−1
𝑘4 = 0.1min−1
14 Initial value problems 128

Just like in section 13.3, the goal of this exercise is to make a plot of the
concentrations of the four species as a function of time. To achieve this
take the following steps:
1. Write your Runge-Kutta routine or copy it from the exercise above.
2. Write a function that returns a vector containing the function values
calculated for the right-hand side of the differential equations (one
value for each differential equation)
3. Calculate the concentrations of 𝐴, 𝐵, 𝐶 and 𝐷 during the first 10
minutes in 250 steps using a for-loop.
4. Plot the concentrations as a function of time.
5. Compare the resulting graph with the one obtained from the
exponentiation solution in Chapter 3. It should be exactly the same!
You can again play around with the end-time to see if eventually all of
𝐴 is transformed into 𝐷 . Note: in this particular case we are dealing
with linear differential equations, which is particularly easy to solve
and the Runge-Kutta method will give an exact result, regardless of the
number of steps taken in the integration. This is in general not the case
for non-linear differential equations.

Epidemic propagation

You have been hired by the World Health Organization to estimate the
impact of a disease outbreak on a city of 50.000 people. The WHO has
developed a model, called the 𝑆𝐼𝑅 model to simulate the propagation of
the infection. Let 𝑆 be the number of healthy people that can contract the
virus, (𝑆 stands for susceptible here) 𝐼 is the number of infected people
and 𝑅 the number of people who have recovered from the infection and
are immune to it. The evolution of the populations 𝑆 , 𝐼 and 𝑅 is then
given by

𝑑𝑆
= −𝛾 · 𝑆 · 𝐼 (14.42)
𝑑𝑡
𝑑𝐼
= 𝛾·𝑆·𝐼−𝛼·𝐼 (14.43)
𝑑𝑡
𝑑𝑅
= 𝛼·𝐼 (14.44)
𝑑𝑡

where 𝛾 is the transmission rate between sick and healthy people and 𝛼 the
recovery rate. For the present case the initial conditions are 𝑆(0) = 50000,
𝐼(0) = 2 and 𝑅(0) = 0. The transmission rate (𝛾) has been estimated at
10−4 day−1 while the recovery rate is 0.5 day−1 . Simulate the dynamics
of the three different populations over a period of 30 days. Report the
maximum number of infected people that can be expected and when
this will happen.

To limit the effect of the outbreak, the WHO has proposed to quarantine
the population. They estimate that it will decreases the transmission rate
to 0.5 10−4 day−1 but will slow down recovery to 𝛼 = 0.1 day−1 . Before
implementing their plan they asked you to forecast the impact of the
14 Initial value problems 129

quarantine on the maximum number of infected people. Comment on


the efficiency of this this quarantine ?
HINT: To solve the system of differential equations you can either use the
routines you have written in the exercises above or use the [Link]
module.
Boundary-value problems 15
In the previous chapters we have dealt with (systems of) differential 15.1 The shooting method . . . . 130
equations for which the initial conditions are known. Therefore, we could Application . . . . . . . . . . . 132
use approaches such as the forward-Euler or Runde-Kutta methods to 𝑁 equations . . . . . . . . . . 133
15.2 The finite difference method 134
calculate at other values of the independent variable. Examples of this
15.3 Exercises . . . . . . . . . . . . . 137
include chemical reactions where the initial concentrations are known. In
The shooting method . . . . 137
so-called boundary value problems, the boundary conditions for one or
1D transport . . . . . . . . . . 138
more of the independent variables are not known at the beginning of the
integration interval. This is for instance the case in a chemical reaction
where we aim to optimize the initial concentrations of the reactants to
reach a certain fixed concentration of products, i.e. the final concentration
is given but not the initial one. Such a boundary value problem with two
unknowns can be written as:

𝑑
𝑦1 = 𝑓1 (𝑥, 𝑦1 , 𝑦2 ) (15.1)
𝑑𝑥
𝑑
𝑦2 = 𝑓2 (𝑥, 𝑦1 , 𝑦2 )
𝑑𝑥

where the independent variable 𝑥 ∈ [𝑥 0 , 𝑥 𝑓 ] There are two boundary


values that are given by:

𝑦1 (𝑥 = 𝑥0 ) = 𝑦1 , 0 (15.2)
𝑦2 (𝑥 = 𝑥 𝑓 ) = 𝑦2 , 𝑓

This problem is graphically shown in Fig. 15.1 where for one of the
functions the initial values is known, while for the other the y-value is
known at the end of the interval. Boundary values problem occur in
many branches of chemical engineering. Several methods of different
complexity have been developed to solve such problems. In this chap-
ter, we will consider two of these: the shooting method and the finite
difference methods.

15.1 The shooting method

In the shooting method, a boundary value problem is treated as an initial


value problem, allowing us to take advantage of the methods that we
have already seen for such problems. The starting point in the shooting
method is to make a guess for the missing initial conditions, i.e. the
values of the variables that are not know for 𝑥 = 𝑥 0 . Using this ’initial
guess’ we can then solve the differentials equations as an initial value
problem, and subsequently we find a way to optimize the guessed initial
conditions to satisfy the actual boundary conditions.
15 Boundary-value problems 131

Figure 15.1: Graphical illustration of the


boundary value problem in equations 15.1
and 15.2. The boundary values are shown
in the green and red circles.

To illustrate this approach, we consider the boundary value problem


defined in Eq. 15.1 with the condition 15.2. To solve this problem using
the shooting method we first make the initial guess that

𝑦2 (𝑥0 ) = 𝛾 (15.3)

We can then solve the equations 15.1 as an initial value problem with the
initial condition 𝑦1 (𝑥 0 ) = 𝑦1,0 and 𝑦2 (𝑥 0 ) = 𝛾 . This can be done using any
of the methods presented in previous chapters. Since the value of 𝑦2 (𝑥 0 )
was a guess, it will not generally be correct and therefore it misses the
target at 𝑥 = 𝑥 𝑓 and the solution does not satisfy the boundary condition
𝑦2 (𝑥 = 𝑥 𝑓 ) = 𝑦2, 𝑓 . Therefore, we have to optimize our guess-value until
we obtain the correct boundary condition for 𝑦2 (𝑥 𝑓 ). To optimize the
value of 𝛾 in a systematic way we define a so-called cost-function:

Φ(𝛾) = 𝑦2 (𝑥 𝑓 , 𝛾) − 𝑦2, 𝑓 (15.4)

The cost function quantifies how far the solution obtained with the
guessed value of 𝑦2 (𝑥 0 ) is from the boundary condition. The cost function
can be then be expanded in a Taylor series:

𝜕Φ
Φ(𝛾 + Δ𝛾) = Φ(𝛾) + Δ𝛾 + 𝑂(Δ𝛾 2 ) (15.5)
𝜕𝛾

When the value of 𝑦2 (𝑥 𝑓 , 𝛾) is equal to the boundary condition the cost-


function Φ(𝛾 + Δ𝛾) = 0. Therefore, we want to use the Taylor expansion
to optimize the initial guess. If we truncate the higher-order terms in the
expansion we therefore obtain:

𝜕Φ −Φ(𝛾)
0 = Φ(𝛾) + Δ𝛾 −→ Δ𝛾 = (15.6)
𝜕𝛾 𝜕Φ
𝜕𝛾

𝜕Φ 𝜕𝑦2 (𝑥 𝑓 ,𝛾)
Since 𝜕𝛾
= 𝜕𝛾
we finally obtain

 𝜕𝑦2 (𝑥 𝑓 , 𝛾) −1
 
Δ𝛾 = − 𝑦2 (𝑥 𝑓 , 𝛾) − 𝑦2, 𝑓 (15.7)
𝜕𝛾
15 Boundary-value problems 132

The value Δ𝛾 is a correction to the previous values of 𝛾 used to solve the


system. We can therefore obtain a new value of 𝛾 from the old ones until
convergence is reached. This leads to a series of values for 𝛾 , denoted 𝛾𝑛 ,
obtained via the recurrence relationship:

𝛾𝑛+1 = 𝛾𝑛 + 𝜌Δ𝛾 0<𝜌≤1 (15.8)

where the factor 𝜌 is introduced to facilitate the convergence. We can keep


updating the value of 𝛾 until |Δ𝛾| < 𝜖 . An important part in Eq. 15.7 that
we still have to consider is the partial derivative. This partial derivative
is unknown we can approximate it numerically:

𝜕𝑦2 (𝑥 𝑓 , 𝛾) 𝑦2 (𝑥 𝑓 , 𝛾𝑛+1 ) − 𝑦2 (𝑥 𝑓 , 𝛾𝑛 )
= (15.9)
𝜕𝛾 𝛾𝑛+1 − 𝛾𝑛

The numerical approximation requires that we know the cost-function


for two for different values of 𝛾 , i.e. we need to make a guess for both 𝛾0
and 𝛾1 .

Application

It is easier to understand the shooting method if we apply it to a simple


problem. Consider the following boundary problem that we would like
to solve:

𝑑2 𝑦
= 4(𝑦 − 𝑥) (15.10)
𝑑𝑥 2

in the domain 𝑥 ∈ [0 , 1] and with two boundary conditions 𝑦(0) = 0


and 𝑦(1) = 2. Note that we need two boundary conditions to solve this
second order differential equation since we have to integrate twice and
hence need to derive the values for two integration constants. In order
to solve this differential equation more easily we can rewrite it into its
𝑑𝑦1
canonical form with the substitutions 𝑦1 = 𝑦 and 𝑦2 = 𝑑𝑥 , which leads
to:

𝑑𝑦1
= 𝑦2 (15.11)
𝑑𝑥
𝑑𝑦2
= 4(𝑦1 − 𝑥) (15.12)
𝑑𝑥

We now have two first order differential equation but only for the first
equation the initial value is known ( 𝑦1 (0) = 0). For the second equation
we do not have an initial value, but we do know the final value for the
first equation, 𝑦1 (1) = 2. Therefore, we need to find a suitable initial
value for 𝑦2 that satisfies the extra boundary condition, 𝑦1 (1) = 2. As
a starting point we make a guess and say: 𝑦2 (0) = 𝛾0 = 1.0. Using this
initial value, we can now solve the system of equations above as an initial
value problem (for example with 𝑠𝑐𝑖𝑝 𝑦.𝑖𝑛𝑡𝑒 𝑔𝑟 𝑎𝑡𝑒.𝑠𝑜𝑙𝑣 𝑖 𝑣𝑝 ) and check
whether the boundary condition 𝑦1 (1) = 2 is respected or not. As seen
15 Boundary-value problems 133

in Fig. 15.2, our initial guess misses the target and hits 𝑦1 (1 , 𝛾0 ) = 1
instead.
We therefore have to improve our guess. However, at this stage we only
know 1 value 𝛾0 and we cannot improve the value of 𝛾 in a systematic
way using Eq. 15.9. In order to evaluate the partial derivative in Eq. 15.7
we need an additional value and we therefore make a second guess,
𝑦2 (0) = 𝛾1 = 0.0. As shown in Fig. 15.2 the solution obtained with
this initial condition also misses the target (even more) and leads to
𝑦1 (1, 𝛾1 ) = −0.813. However we now have two values for and we can
use both these values to estimate how to update our value of 𝛾 using Eq.
15.7 and write

 𝑦1 (𝑥 𝑓 , 𝛾1 ) − 𝑦1 (𝑥 𝑓 , 𝛾0 ) −1
 
Δ𝛾 = − 𝑦1 (𝑥 𝑓 , 𝛾1 ) − 𝑦1, 𝑓 (15.13)
𝛾1 − 𝛾0

and obtain a new value for 𝛾 as 𝛾2 = 𝛾1 + Δ𝛾 . We have set here 𝜌 = 1


for simplicity. The procedure can then be repeated until we obtain
𝑦1 (1, 𝛾𝑛 ) = 2 ± 𝜖 or equivalently |Δ𝛾| < 𝜖. The solution is illustrated in
Fig. 15.2 we see that after only 3 iterations the shooting method converges
toward the correct boundary condition. This is the case for any linear
boundary value problem. For non-linear problems more iterations are
generally required.

Figure 15.2: Resolution of the boundary


value problem (15.1) with the shooting
method. Only 3 iterations are necessary.

𝑁 equations

The shooting method presented here can be generalized to the case of 𝑁


simultaneous differential equations using an approach similar to that for
the Newton method for solving non-linear equations that we have seen
earlier. We consider a system of 𝑁 equations:

𝑑𝑦 𝑖
= 𝑓 (𝑥, 𝑦1 , 𝑦2 , ...𝑦 𝑁 ) 𝑥 ∈ [𝑥 , 𝑥 𝑓 ] (15.14)
𝑑𝑥
15 Boundary-value problems 134

The boundary conditions are split between initial conditions for 𝑟 equa-
tions and at least 𝑁 − 𝑟 final conditions:

𝑦 𝑗 (𝑥0 ) = 𝑦 𝑗,0 𝑗 = 1 , 2 , ...𝑟 (15.15)


𝑦 𝑗 (𝑥 𝑓 ) = 𝑦 𝑗, 𝑓 𝑗 = 𝑟 + 1 , ....𝑁 (15.16)

Following the principle we have outlined above for the shooting method
we will make an initial guess for the missing initial conditions:

(0)
𝑦 𝑗 (𝑥0 ) = 𝛾 𝑗 𝑗 = 𝑟 + 1 , ....𝑁 (15.17)

The system of equations 15.14 can then be solved as an initial value


problem again, using any method available for that. This will provide
some values for the 𝑦 𝑗 (𝑥 𝑓 ) with 𝑗 = 𝑟 + 1 , ....𝑁 that generally do not
respect the boundary conditions (15.16). To improve our guess-values we
therefore compute the derivative of these incorrect solutions with respect
of our guess for each of the guess values. We can do this for all of the
guess values simultaneously by the calculation of the 𝐽 𝑎𝑐𝑜𝑏𝑖𝑎𝑛 matrix,
exactly like we have seen in Newton’s method for non-linear equations:

𝜕𝑦𝑟+1 𝜕𝑦𝑟+1 𝜕𝑦𝑟+1


© 𝜕𝛾𝑟+1 𝜕𝛾𝑟+2
... 𝜕𝛾𝑁 ª
­ 𝜕𝑦𝑟+2 𝜕𝑦𝑟+2 𝜕𝑦𝑟+2 ®
...
J(𝑥 𝑓 , 𝜸) = ­ 𝜕𝛾𝑟+1 𝜕𝛾𝑟+2 𝜕𝛾𝑁 ® (15.18)
­
­ ... ... ... . . . ®®
®
­ 𝜕𝑦 𝜕𝑦 𝑁 𝜕𝑦 𝑁
𝑁
« 𝜕𝛾𝑟+1 𝜕𝛾𝑟+2
... 𝜕𝛾𝑁 ¬

We can then compute the correction to our value for 𝜸 by:

𝑦𝑟+1 (𝑥 𝑓 , 𝛾𝑟+1 )
 −1 ©­ 𝑦𝑟+1 (𝑥 𝑓 , 𝛾𝑟+2 )ª®
Δ𝜸 = J(𝑥 𝑓 , 𝜸)

·­ (15.19)
...
®
­ ®
« 𝑦 𝑁 (𝑥 𝑓 , 𝛾 𝑁 ) ¬

and finally update the values of 𝛾 in a step-wise approach.

𝜸(𝑛+1) = 𝜸(𝑛) + 𝜌Δ𝜸 (15.20)

15.2 The finite difference method

A second technique to solve boundary value problems is the finite


difference approach(FD). A boundary value problem is generally given
by:

𝑑2 𝑦 𝑑𝑦
+ 𝛼(𝑥) + 𝑦𝛽(𝑥) = 𝑓 (𝑥) (15.21)
𝑑𝑥 2 𝑑𝑥

on the domain 𝑎 ≤ 𝑥 ≤ 𝑏 and with the boundary conditions given at


𝑎 an 𝑏 , i.e. 𝑦(𝑎) = 𝑦 𝑎 and 𝑦(𝑏) = 𝑦𝑏 . The finite difference method is
based on the idea of replacing the first and the second derivatives in this
15 Boundary-value problems 135

equation by numerical approximations at different points on the 𝑥 axis.


The finite difference method starts by dividing the 𝑥 -axis in a series of
evenly spaced points, called nodes and marked as 𝑥 𝑖 , with the interval
between two points being given by 𝑥 𝑖+1 − 𝑥 𝑖 = ℎ . If we have 𝑁 + 1 nodes
numbered from 0 to 𝑁 we have 𝑁 intervals. In the [𝑎, 𝑏] interval the
location of each node is given by 𝑥 𝑖 = 𝑎 + 𝑖 ℎ .
As we have seen in chapter 11, the first and second derivatives in Eq. 11
can be approximated by the expressions:

𝑑𝑦 𝑦 𝑖+1 − 𝑦 𝑖−1
= (15.22)
𝑑𝑥 2ℎ
𝑑2 𝑦 𝑦 𝑖+1 − 2 𝑦 𝑖 + 𝑦 𝑖−1
= (15.23)
𝑑𝑥 2 ℎ2

Using these approximations we can rewrite Eq. 15.21 as


(𝑦 𝑖+1 − 2 𝑦 𝑖 + 𝑦 𝑖−1 ) + 𝛼(𝑥 𝑖 )(𝑦 𝑖+1 − 𝑦 𝑖−1 ) + ℎ 2 𝛽(𝑥 𝑖 )𝑦 𝑖 = ℎ 2 𝑓 (𝑥 𝑖 ) (15.24)
2

By writing this equation for each node and accounting for the boundary
conditions we obtain a system of linear equations given by:

𝑦0 = 𝑦 𝑎 (15.25)
ℎ ℎ
𝑦 𝑖+1 (1 + 𝛼(𝑥 𝑖 )) + 𝑦 𝑖 (ℎ 2 𝛽(𝑥 𝑖 ) − 2) + 𝑦 𝑖−1 (1 − 𝛼(𝑥 𝑖 )) = ℎ 2 (15.26)
𝑓 (𝑥 𝑖 )
2 2
𝑦 𝑁 = 𝑦𝑏 (15.27)

This linear system of equations can be written is a matrix form:

𝑎 00 𝑎01 ... 𝑎 0𝑁 𝑦0 𝑏0
­ 𝑎10 𝑎11 ... 𝑎 1 𝑁 ® ­ 𝑦1 ® ­ 𝑏 1 ®
© ª © ª © ª
­ . .. .. .. ®® · ­­ .. ®® = ­­ .. ®® (15.28)
­ . .
­ . . . ® ­ . ® ­ . ®
«𝑎𝑁 0 𝑎𝑁1 ... 𝑎 𝑁 𝑁 ¬ « 𝑦 𝑁 ¬ «𝑏 𝑁 ¬

or more compactly

A·Y=B (15.29)

The solution of this matrix equation is a straightforward solution for


a system of linear equation that can for instance be solved using
𝑛𝑝.𝑙𝑖𝑛𝑎𝑙 𝑔.𝑠𝑜𝑙𝑣𝑒(), resulting in all the values of 𝑦 along the 𝑥 -axis with the
appropriate boundary conditions. The finite difference method therefore
transforms a boundary value problem into a linear system of equations
that is easily solvable.
15 Boundary-value problems 136

As an example, the code below uses the finite difference method to solve
the boundary value problem:

𝑑2 𝑦
+𝑦=0 (15.30)
𝑑𝑥 2

on the domain 0 ≤ 𝑥 ≤ 𝜋/2 with the Dirichlet boundary condition


𝑦(0) = 1.0 and 𝑦(𝜋/2) = 1.0. The results of the program is plotted in
Fig. 15.3.
1 import numpy as np
2 import [Link] as plt
3 from [Link] import solve
4

5 # Number of nodes
6 N = 10
7 X = [Link](0,[Link]/2.,N+1)
8 h = X[1]-X[0]
9

10 # define the matrices


11 A = [Link]((N+1,N+1))
12 b = [Link](N+1)
13

14 # boundary condition at x=0


15 A[0,0] = 1.0
16 b[0] = 1.0
17

18 # boundary condition at x = pi/2


19 A[-1,-1] = 1.0
20 b[-1] = 1.0
21

22 # all the other nodes


23 for i in range(1,N):
24 A[i,i-1] = 1.0
25 A[i,i] = -2.0 + h**2
26 A[i,i+1] = 1.0
27

28 Y = solve(A,b)
29

30 [Link](X,Y,marker=’o’,color=’black’)
31 [Link](X,[Link](X)+[Link](X),linewidth=2,color=’blue’)
32 [Link](’X’)
33 [Link](’Y’)
34 [Link]()
15 Boundary-value problems 137

Figure 15.3: Example of numerical solu-


tion of a BVP using the FD method. The
figure shows the numerical solution of
𝑑2 𝑦
+ 𝑦 = 0 with the boundary conditions
𝑑𝑥 2
𝑥(0) = 1.0 and 𝑥(𝜋/2) = 1.0. (dots) The
plain line shows the analytical solution of
this equation.

15.3 Exercises

The shooting method

In this first exercise you will implement the shooting method for the
solution of boundary value problems as described in section 15.1. As a
reminder, the shooting method treats a boundary value problem as an
initial value problem. We therefore make a guess of the initial values that
are not provided, solve the equations and compare the solution obtained
with the boundary conditions provided. This allows us to make a better
guess for the missing initial condition. We can therefore converge toward
the correct solution by iterating the process. A detailed description of
the method can be found in section 15.1 and 15.1.
To test your implementation you will compute the solution of the bound-
ary value problem given by Eq. 15.10, i.e.

𝑑2 𝑦
= 4(𝑦 − 𝑥) (15.31)
𝑑𝑥 2

for 𝑥 ∈ [0 , 1] and with the boundary conditions 𝑦(0) = 0 and 𝑦(1) = 2.


Since this system contains a second derivative we need to transform it
first to a canonical form leading to

𝑑𝑦1
= 𝑦2 (15.32)
𝑑𝑥
𝑑𝑦2
= 4(𝑦1 − 𝑥) (15.33)
𝑑𝑥

We now know the initial value for the first equation ( 𝑦1 (0) = 0) but
not for the second. We therefore have to guess it and improve our
guess following the procedure introduced in section 15.1. As shown in
15 Boundary-value problems 138

section 15.1, the shooting method should converge rapidly toward the
solution as represented in Fig. 15.2.
HINT: To help you in the implementation we provide below a possible
pseudo code for the shooting method
1 from [Link] import solve_ivp
2

3 # function that retrurns


4 # the systems of equation
5 def df(Y,x):
6 dy1 = ...
7 dy2 = ...
8 return [Link]([dy1,dy2])
9

10 ....
11

12 # boundary condition
13 y1_0,y1_f = 0.0, 2.0
14

15 # guesses for gamma


16 # remember we need two
17 gamma0, gamma1 = 1.0, 0.0
18 ....
19

20 # solve the system for the first guess


21 Y = odeint(....)
22 ...
23

24 # error obtained with the first guess


25 error = ....
26

27 # loop until convergence


28 while error > EPS and niter < MAXITER:
29

30 # solve for the new value of gamma


31 Y = solve_ivp(....)
32 ...
33

34 # compute the error


35 error = ...
36

37 # compute Delta gamma


38 dG = ....
39

40 # update the gammaXs


41 gamma0 = gamma1
42 gamma1 += dG
43

1D transport

We consider a Newtonian fluid undergoing laminar pressure-driven flow


between two parallel and infinite plates. The plates are separated by a
distance 𝐵. The bottom plate is fixed while the top plate moves in the 𝑦
direction at a constant speed 𝑉𝑢𝑝 . This system is represented in Fig. 15.4a.
We know the value of the constant dynamic pressure gradient Δ𝑃 Δ𝑥 and
we want to compute the velocity profile 𝑣(𝑦) at different heights in the
layer between the plates. If we suppose that the transport occurs only in
15 Boundary-value problems 139

the 𝑥 direction the Navier-Stokes equation reduces to

𝑑2 𝑣 Δ𝑃
𝜇 = ≡𝐺 (15.34)
𝑑𝑦 2 Δ𝑥

with the boundary conditions 𝑣(𝑦 = 0) = 0 and 𝑣(𝑦 = 𝐵) = 𝑉𝑢𝑝 . The goal
of this exercise is to compute the profile 𝑣(𝑦) using the expression for the
numerical evaluation of the second derivative and the techniques used
to solve linear systems of equations. Using the finite difference approach
we can rewrite the eq. 15.34 as

𝑣 𝑖+1 − 2𝑣 𝑖 + 𝑣 𝑖−1
𝜇 =𝐺 (15.35)
ℎ2

Note that G You can therefore rewrite the problem as a system of linear
equations

1 −2 1 ... ... 𝑎 0𝑁 𝑣0 𝑠0
­ 0 1 −2 1 ... 𝑎 1𝑁 ® ­ 𝑣 1 ® ­ 𝑠 1 ®
© ª © ª © ª
­ . .. .. ®® ­­ .. ®® ­­ .. ®®
­ . .
­ . . ® ­ . ® ­ . ®
­ . .. .. ®® · ­­ .. ®® = ­­ .. ®® (15.36)
­ . .
­ . . ® ­ . ® ­ . ®
­ . .. .. ®® ­­ .. ®® ­­ .. ®®
­ . .
­ . . ® ­ . ® ­ . ®
«𝑎𝑁 0 𝑎𝑁2 ... 𝑎 𝑁 𝑁 ¬ «𝑣 𝑁 ¬ « 𝑠 𝑁 ¬

In this equation 𝑣 𝑖 is the velocity profile at the 𝑖 -th point. You must
determine the expression for the 𝑎 𝑖𝑖 and 𝑠 𝑖 . Once you have established
this system of linear equations, use one of the techniques presented in
chapter 9 (e.g. [Link]) to solve it and obtain the values of the
𝑣 𝑖 . Compare your numerical results to the analytic expression given by

𝑦 1 Δ𝑃 2
𝑣(𝑦) = 𝑉𝑢𝑝 + (𝑦 − 𝑦𝐵) (15.37)
𝐵 2𝜇 Δ𝑥

During the calculation you will take : 𝐵 = 5 · 10−3 m, 𝑉𝑢𝑝 = (1/6) · 10−4
m/s, 𝜇 = 10−3 . You will vary the dynamic pressure gradient from 0.0
to -0.06 Pa/m (take 5 points linearly spaced between these two values ).
Take about 50 to 100 points along the 𝑦 axis. You will then compare your
numerical results to the analytic solution. You should find variations of
the velocity profile similar to the ones represented in Fig. 15.4b.

Figure 15.4: Laminar flow between two


parallel plate. a) Representation of the
problem the top plate is moving at a speed
𝑉𝑢𝑝 while the bottom one is fixed. This cre-
ates a flow of the liquid present between
the two plates with a well defined profile
𝑣(𝑦). b) Velocity profile of the liquid ob-
tained via the exact solution (black line) or
a numerical approach (blue line and dots).
Special topics
Monte Carlo schemes 16
In most of the previous chapters we have considered the calculation of 16.1 Random numbers . . . . . . . 141
physical quantities (heat, velocities) that were assumed to be continuous 16.2 Monte-Carlo Integration . . 143
functions of some spatial or temporal variable. For example, the velocity Standard MC integration . . 144
Throwing darts . . . . . . . . 145
profile in a water channel can be assumed to be a function of the lateral
16.3 Random walks: Direct MC . 146
dimension of the pipe 𝑣(𝑧). In reality, this profile is defined by the
Random Walk in 1D . . . . . 146
individual velocities of the water molecules that compose the fluid
Diffusion: random walk . . . 147
moving through the pipe. However at the microscopic scale the heat Example: stock prices . . . . 148
of a material is directly given by the motion of the atoms composing Random walk in 2D . . . . . 148
the system. We have therefore in the previous chapters calculated these 16.4 Metropolis Monte-Carlo . . 149
various physical quantities over a large domain where the microscopic Lattice Metropolis . . . . . . 150
details are averaged. Off-lattice Metropolis . . . . 150
16.5 Exercises : Algorithm . . . . 152
In this chapter we will focus on the description of matter on the mi- Approximating 𝜋 by MC . . 152
croscopic scale and focus in particular on Monte-Carlo simulations. Stock Price . . . . . . . . . . . 152
Monte-Carlo schemes are based on random number and are therefore Mixing of particles . . . . . . 152
not so-called deterministic simulations but they do give insight into the
statistical distributions that play a key role in thermodynamics. In Monte
Carlo methods, a central position is reserved for random numbers, and
therefore we will first discuss generation of random numbers in Python
codes.

16.1 Random numbers

As their name indicates, a random-number-generator produces a se-


quence of numbers that are randomly distributed according to a well-
defined distribution. The sub-module 𝑛𝑢𝑚𝑝 𝑦.𝑟 𝑎𝑛𝑑𝑜𝑚 contains a vari-
ety of different random number generators. The most common one is
[Link]() that is used in the code below:
1 r = [Link]()

The functions returns one random number in the open interval [0 , 1[. The
generation of random numbers is done by deterministic algorithm and
therefore they are not really random. However, good random number
generators lack any patterns in the series of numbers generated and can,
for all practical purposes be considered as truly random. In fact, it is
sometimes useful when coding to obtain the exact same sequence of
random numbers when running a code multiple times, especially for
debugging purposes. This reproducibility of the random numbers is the
result of the algorithms used, which generally produce numbers based
on a certain ’seed-number’. By default, this seed value is taken for the
system-time on the computer that runs the code but it is also possible to
imposing a given seed to the generator with the function:
1 r = [Link](13)
16 Monte Carlo schemes 142

The seed (here equals to 13) is an integer that entirely determines the full
sequence of random number generated. For example in the following
code:
1 [Link](13) # seed the generator
2 r1 = [Link](10) # creates 10 random numbers
3 r2 = [Link](10) # creates another 10 random numbers
4

5 [Link](13) # re-seed the generator


6 r3 = [Link](10) # creates 10 random numbers

𝑟 1 and 𝑟 3 will be identical.


Different distributions of random numbers have been implemented
in numpy. In this chapter we will only use the uniform distribution
introduced above, and the normal (Gaussian) distribution:

1 𝑥2
𝑝(𝑥) = √ 𝑒 − 2 (16.1)
2𝜋

A series of random numbers distributed according to a normal dis-


tribution can be generated using the function [Link]():

1 X = [Link](100)

The arguments leads to the generation of 100 random numbers following


a normal distribution. We can visualize and compare the uniform and
normal distribution using the code below:
1 import numpy as np
2 import [Link] as plt
3

4 N = 1E4
5 XYu = [Link](N,2)
6 XYn = [Link](N,2)
7

8 [Link](XYu[:,0],XYu[:,1],’o’,color=’#007FFF’)
9 [Link](XYu[:,0],20,color=’#007FFF’)
10

11 [Link](XYn[:,0],XYn[:,1],’o’,color=’#007FFF’)
12 [Link](XYn[:,0],20,color=’#007FFF’)

This code generates two arrays (each with two columns representing the
𝑥 and 𝑦 values) containing the position of 𝑁 = 1𝐸4 points that are either
distributed either uniformly or according to a normal distribution in the
𝑥, 𝑦 plane. These points are visualized both by plotting them directly
and by computing the histogram of their 𝑥 coordinate. The results of this
simple code is show in Fig. 16.1
By default [Link]() generates random numbers between
0 and 1, however, other distributions can sometimes be useful. As an
example the command:
1 X = [Link](min,max,N)

generates 𝑁 points randomly distributed between 𝑚𝑖𝑛 and 𝑚𝑎𝑥 in-


stead of 0 and 1 as in 𝑟 𝑎𝑛𝑑𝑜𝑚.𝑟 𝑎𝑛𝑑(). Similarly, calling the function
[Link]() with additional arguments allows to obtain any
kind of normal distribution:
1 X = [Link](loc = mu, scale = sigma, N=100)
16 Monte Carlo schemes 143

Figure 16.1: Random number generators.


Top panels show the position of the points
in the 𝑥 − 𝑦 plane. A clear distinction can
be seen between the two distribution. The
bottom panels show the histogram of the
𝑥 coordinate of the points.

generates 𝑁 points that are distributed according to a normal distribution


centered around 𝑙𝑜𝑐 and with a standard deviation given by 𝑠𝑐𝑎𝑙𝑒 . The
function generates number according to the distribution:

1 (𝑥−𝜇)2

𝑝(𝑥) = √ 𝑒 2𝜎2 (16.2)
2𝜋𝜎 2

The random number generators above all generate floating point numbers.
For some purposes it is of interest to generate a random sequence of
integers, for which the function [Link].random_integers() is
available:
1 X = [Link].random_integers(low,high,N)

This command creates 𝑁 random integers in the interval [𝑙𝑜𝑤, ℎ𝑖 𝑔 ℎ]. A


similar command
1 X = [Link](low,high,N)

also creates 𝑁 random integers but in the interval [𝑙𝑜𝑤, ℎ𝑖 𝑔 ℎ[ (i.e. the
high value is never generated). In addition to the random generators
contained in 𝑛𝑢𝑚𝑝 𝑦 the module 𝑠𝑐𝑖𝑝 𝑦.𝑠𝑡 𝑎𝑡𝑠 gives access to a large
number of distributions that may be useful for specific physical systems.
For example the code:
1 from [Link] import maxwell
2 r = [Link](size=1000)

will generate 1000 random numbers following a Maxwell distribution


that can be used to sample the velocities of atoms in a gas.

16.2 Monte-Carlo Integration

One of the Monte Carlo schemes that we will consider is the use of random
numbers in the calculation of integrals. We have seen that integrals over
a certain integration interval can be calculated by dividing the interval in
16 Monte Carlo schemes 144

Figure 16.2: Illustration of the standard


Monte Carlo integration approach.

a series of equally spaced sub-interval and approximating the integral of


each by a rectangle, for instance according to the midpoint approximation.
A similar approach can be followed using random numbers.

Standard MC integration

In the standard Monte-Carlo integration approach the integral

∫ 𝑏
𝐼= 𝑓 (𝑥)𝑑𝑥 (16.3)
𝑎

is evaluated by the numerical summation:

𝑁
𝑏−𝑎X
𝐼= 𝑓 (𝑥 𝑖 ) (16.4)
𝑁 𝑖=1

where the points 𝑥 𝑖 are chosen randomly instead of using a regular grid.
This is graphically shown in Fig.16.2. This approach to integration may
seem somewhat surprising, but in fact it turns out to be a very efficient
way of evaluation high-dimensional integrals.
To illustrate the Monte-Carlo integration method we will calculate the
∫2
integral 1 (2 + 3 𝑥)𝑑𝑥 , of which the exact value is 6.5. The code below
evaluates the integral using the Monte Carlo method:
1 import numpy as np
2

3 def f(x):
4 return 2.+3.*x
5

6 def MCInt(a,b,f,N):
7 X = [Link](a,b,N)
8 I = (float(b-a)/N) * [Link](f(X))
9 return I
10

11 a,b,N = 1.,2.,1000
12 I = MCInt(a,b,f,N)
13 print(I)

By executing this code we find a value of ≈ 6.5 ± 0.05 with only random
100 points. Of course the exact value changes each time we execute the
program as the sampling points are chosen randomly. We can add more
points to obtain a better estimate but for simple 1-dimensional integrals
the Monte-Carlo integration converges slowly and 1E6 points are required
to obtain an error below 10−3 . Therefore, for simple integrals the methods
16 Monte Carlo schemes 145

Figure 16.3: Illustration of the hit or miss


Monte Carlo integration of the function
𝑓 (𝑥), represented by the solid line, be-
tween 𝑎 and 𝑏 . 𝑁 random points are gen-
erated and we count the ones that fall
below the curve.

introduce in Chapter 10 are much more efficient. However for functions


of many variables the Monte-Carlo integration clearly outperform any
other methods.

Throwing darts

A second approach of Monte-Carlo Integration is the so-called throwing


of darts, which illustrated in Fig.16.3. In this approach, 𝑁 points are
randomly generated in a square where the function 𝑓 (𝑥) is defined. 𝑁0
of these points will fall below the curve. Since we know the area of the
square 𝐴 = 𝑦0 ∗ (𝑏 − 𝑎) we can estimate the value of the integral by the
fraction of the points below the curve, multiplied by the total area:

𝑏
𝑁0

𝑦= 𝑓 (𝑥)𝑑𝑥 ' × [𝑦0 (𝑏 − 𝑎)] (16.5)
𝑎 𝑁

In the Python code below this method is used to calculate the integral:

4
𝑥

𝐼= 4 + sin(2 𝑥) exp( ) (16.6)
0 2

1 import numpy as np
2

3 def f(x):
4 return 4 + [Link](2*x)*[Link](0.5*x)
5

6 def MCInt(a,b,y0,f,N):
7 X = [Link](a,b,N)
8 Y = [Link](0,y0,N)
9 N0 = Y[Y<=f(X)].size
10 return float(N0)/N*y0*(b-a)
11

12 a,b,y0,N = 0.,4.,15.,10000
13 I = MCInt(a,b,y0,f,N)
14 print(I)

For 100 points this approach leads to 𝐼 ' 17.9 where the quadrature
method implemented in 𝑠𝑐𝑖𝑝 𝑦 leads to 𝐼 = 17.8365. The same method is
used in one of the exercises at the end of this chapter to determine the
value of 𝜋.
16 Monte Carlo schemes 146

16.3 Random walks: Direct MC

In many physical processes there is an element of randomness. A well-


known example is the Brownian motion of small particles in a gas or
liquid. Such processes are often described as random walks, in which, as
the name suggests, the motion of the particles is dictated by random forces.
Therefore, at each point in time the particle can change direction randomly
to continue its path. Such motion is very useful to describe systems whose
dynamics is not entirely deterministic and where randomness plays a
crucial role. This is for example the case for the motion of particles in
liquids and gases as mentioned above, but also for the evolution of stock
prices that exhibit a certain randomness.

Random Walk in 1D

The simulation of a random walk in 1D is then straightforward using


the uniform generator. Starting from 𝑥 = 0, we can generate a random
number 𝜖 ranging between 0 and 1. The walker then move to the left
if 𝜖 < 0.5 and to the right if 𝜖 > 0.5. We can let the walker walk for a
while to see where he’s going. Of course if we restart the walk we will
obtain a different walk as each step is chosen randomly. Hence it is very
important to average random walk (and in general stochastic processes)
to obtain meaningful information. The snippet of code below show how
to implement a 1D random walk and average the distance walked by
the walker to extract meaningful quantities. The results are shown in
Fig. 16.4.
1 import numpy as np
2 import [Link] as plt
3

4 # number of walker
5 M = 1000
6

7 # size and number of step


8 dX, NT = 2.,1000
9

10 # positions
11 X = [Link]((NT,M))
12

13 # loop over time


14 for iT in range(1,NT):
15

16 # generate random number


17 EPS = [Link](M)
18

19 # propagate the walkers


20 X[iT,EPS<=0.5] = X[iT-1,EPS<=0.5] - dX
21 X[iT,EPS>0.5] = X[iT-1,EPS>0.5] + dX
22

23 [Link]([Link](NT),[Link]([Link](X**2,1)),color=’blue’,linewidth=2)
24 [Link]([Link](NT),[Link](X**1,1),color=’black’,linewidth=2)
25 [Link]([Link](NT),[Link]([Link](0.,NT,NT))*dX,color=’#007FFF’,
linewidth=2)
26 [Link]([Link](NT),[Link](M),color=’#5F5F6F’,linewidth=2)
27 [Link]()

As we can see on the first panel of Fig. XX, every walker take a very
different trajectory than than the other ones. Some go left (negative
position), some right (positive values). However the mean value of the
16 Monte Carlo schemes 147

Figure 16.4: Evolution of random walker


in 1D. Left - 10 examples of random walk
Middle - Average value of the position
and displacement of 1000 walkers. Right
- Distribution of the final position of 1000
walkers.

position (< 𝑋 >) is null over time since the probability to go left equals
the one to go right. However it can be shown that the average of the
square distance, i.e. < 𝑋 2 > increase linearly with the number of steps
𝑁 . Taking the square root of < 𝑋 2 > we obtain the root-mean-squared
distance which represent the average positive distance away from the 0.
We easily see that this distance increases with the square-root of 𝑁 :

√ √
< 𝑋 2 >(𝑁) = 𝑁Δ𝑋 (16.7)

As we can see our numerical simulations here with 1000 walkers gives
agrees very well with this theoretical results. Running the simulation with
even more walkers would lead to an even better agreement. Finally we
can note from the last panel that the distribution of final positions follows
a Gaussian distribution centered around 0. That also agree perfectly with
analytical analysis that are however beyond the scope of this chapter.

Diffusion: random walk

The random walk presented here can be directly related to PDEs as


the ones studied in the previous sections. Let’s assume that the walker
arrives at the positions 𝑚 after 𝑁 + 1 steps. This implies that at time 𝑁
the walker either came from the position 𝑚 − 1 with a probability 𝑝 or
from 𝑚 + 1 with a probability 𝑞 . We can therefore write

𝑃(𝑚, 𝑁 + 1) = 𝑝 × 𝑃(𝑚 − 1 , 𝑁) + 𝑞 × 𝑃(𝑚 + 1 , 𝑁) (16.8)

In our case 𝑝 = 𝑞 = 0.5 as the walker is truly random. In the equation


above 𝑃(𝑚, 𝑁) means that at time 𝑁 the walker is at the position 𝑚 . We
can subtract 𝑃(𝑚, 𝑁) from both side, divide by Δ𝑡 and introduce a few
Δ𝑥 to obtain

𝑃(𝑚, 𝑁 + 1) − 𝑃(𝑚, 𝑁) Δ𝑥 2 𝑃(𝑚, 𝑁) − 𝑃(𝑚 − 1 , 𝑁) Δ𝑥 2 𝑃(𝑚 + 1 , 𝑁) − 𝑃(𝑚, 𝑁)


= −𝑝 +𝑞
Δ𝑡 Δ𝑡 Δ𝑥 2 Δ𝑡 Δ𝑥 2
(16.9)
If we now take 𝑝 = 𝑞 = 𝛼 and we gather all the terms together we
obtain

𝑃(𝑚, 𝑁 + 1) − 𝑃(𝑚, 𝑁) Δ𝑥 2 𝑃(𝑚 + 1, 𝑁) − 2𝑃(𝑚, 𝑁) + 𝑃(𝑚 − 1 , 𝑁)


=𝛼
Δ𝑡 Δ𝑡 Δ𝑥 2
(16.10)
16 Monte Carlo schemes 148

Figure 16.5: Evolution of the stock price


following eq. 16.12.

We recognize here the finite difference expression of the partial derivative


and can therefore write

𝜕 𝜕2
𝑃(𝑥, 𝑡) = 𝐷 2 𝑃(𝑥, 𝑡) (16.11)
𝜕𝑡 𝜕𝑥

We therefore arrive to the diffusion equation studied in chapter 6 for the


diffusion of drug through the body of a patient. The parameter 𝐷 is here
called the diffusion coefficient. The random walk techniques constitutes
therefore a statistical approach of solving the diffusion equation by
tracking the trajectory of a collection of ’particles’. The more particle we
follow the better the agreement with the continuous diffusion equation.

Example: stock prices

1D Random walks play a central role in many fields including finance. A


common mathematical mode for the evolution of stock prices is given by
the equation


𝑥 𝑛 = 𝑥 𝑛−1 + Δ𝑡𝜇𝑥 𝑛−1 + 𝜎𝑥 𝑛−1 𝑟𝑛−1 Δ𝑡 (16.12)

where 𝑥 𝑛 is the stock price at time 𝑡 𝑛 , Δ𝑡 the time interval, 𝜇 is the growth
rate of the stock price, 𝜎 its volatility and 𝑟0 , 𝑟1 , . . . 𝑟𝑛−1 are normally
distributed random numbers with means 0 and unit standard deviation.
A common technique to evaluate the expected price of the stock is to
simulate 𝑁 realization of the equation above providing 𝑥 0 , 𝜇, 𝜎 and Δ𝑡 .
A few realizations, simulated with 𝑥 0 = 100, Δ𝑇 = 0.1, 𝜇 = 𝜎 = 0.01, are
shown in Fig. XX
From these realizations alone it is rather hard to make any prediction.
Running a large number of these simulations allows to get a better idea
of the price evolution. Three cases are reported in Fig. XX. As you can
see there depending on the growth rate and volatility it is a more or less
good idea to invest in the stock.

Random walk in 2D

Random walks can be generalized to the 2D. The process is very similar to
the 1D cases except that each walker can not only go left or right but also
16 Monte Carlo schemes 149

Figure 16.6: Top Panels : Mean (blue) and


median (red) value of the stock price. Bot-
tom Panel : Distribution of the final stock
price. Three cases are shown: Left 𝜇 = 0.01,
𝜎 = 0.1; Middle 𝜇 = 0.1, 𝜎 = 0.1; Right
𝜇 = 0.1, 𝜎 = 0.01.

Figure 16.7: 2D random walks. a) Ten steps


for 5 walkers on a regular grid. b) Many
more small steps. c) Evolution of the dis-
tance walked by 5 walker. d) Average dis-
tance walked by 100 walkers.

up or down. These four directions are equiprobable with a probability


of 1/4. As an example one trajectory representing the random walk of 5
different particles is represented in Fig. XX. During one of the exercises
of the session you will implement yourself a 2D random walk process
and study the properties of this particular system.

16.4 Metropolis Monte-Carlo

Consider the the 2D random walk presented above. During the simulation
the motion of each particle is uncorrelated with those of the others
particles. The probability for one particle to move in a given direction is
independent of the position of the other ones. As a consequence all the
possible configurations of 𝑁 particles are equiprobable. This is of course
not the case when the particles interact with each other. In presence of
interactions the particle will tend to adopt preferential configurations
that minimize the potential energy of the entire system. The simulation
of such system can be done with the Metropolis Monte-Carlo algorithm
16 Monte Carlo schemes 150

presented here.

Lattice Metropolis

To illustrate the metropolis algorithm let’s consider two types of ’atoms’


labeled 𝐴 and 𝐵 and that can diffuse on a rigid 2D lattice as represented
in Fig. XX. To be clear each atoms can switch its position with one of
its neighbor. If the the atom switches position with a vacancy the move
is equivalent to a diffusion. Otherwise we have exchange between tow
atoms. We assume that each atom only interacts with its nearest neighbor.
Furthermore we assume that each 𝐴 atoms prefers to have 𝐵 neighbors
and inversely. We therefore write the potential interaction

𝑈𝐴𝐴 = 𝑈𝐵𝐵 = 𝐸rep 𝑈𝐴𝐵 = 𝐸att (16.13)

For a given configuration of the atoms the total potential energy of the
system can be obtained by summing up all the interactions between the
pairs of atoms. We assume that in their initial configuration the atoms 𝐴
and 𝐵 occupied each half of the domain as represented in Fig. XX. We
want to obtain the equilibrium configuration for a given temperature 𝑇 .
The metropolis algorithm then proceed as follows:

(1) Consider one particle with a randomly chosen position on the lattice
𝑖, 𝑗
(2) Randomly chose one move for the atom at 𝑖, 𝑗 and perform that move
(3) Compute the energy change Δ𝐸 that is induced by this trial movement
(4) if Δ𝐸 < 0 the move is accepted and we return to (2)
(5) if Δ𝐸 > 0, a random number 𝑟 is generated between 0 and 1.
(6) if 𝑟 < exp(−Δ𝐸/𝑘 𝐵 𝑇), the move is accepted. Otherwise the move is
rejected. In both case we return to (2).

This very simple algorithm ensure that the final configuration (providing
enough Monte-Carlo moves have been performed) respect a Boltzman
distribution, i.e. is in equilibrium with the temperature imposed to the
system. We have reported in Fig. XX two final configurations obtained
after 10000 moves for two different temperature. During the simulations
we have set 𝐸rep = −𝐸att = 1.0. The thermal energy was then 𝑘 𝐵 𝑇 = 0.5
(low temperature) and 𝑘 𝐵 𝑇 = 10.0 (high temperature). As you can see on
the corresponding configuration the a very ordered phase is obtained at
low temperature as a very small number of unfavorable moves (i.e. move
that increases the total potential energy) are performed. On the contrary
at high temperature many unfavorable move are performed resulting in
a much more disordered phase.

Off-lattice Metropolis

The Metropolis algorithm can be generalized to the case where the


particles are not fixed to a specific grid. To illustrate that method, let’s
consider an ensemble of identical particles that are allowed to move in a
16 Monte Carlo schemes 151

Figure 16.8: Lattice Monte-Carlo for


atomic diffusion with two types of atoms
(𝐴 red and 𝐵 blue). Top-Left: Illustration of
the possible moves. Top-Right: Initial con-
figuration of the atoms. Bottom-Left: Con-
figuration obtained for 10000 Monte-Carlo
moves and at low temperature. Bottom-
Right: Configuration obtained for 10000
Monte-Carlo moves and at high tempera-
ture.

two dimensional space. We assume that the particles interact with each
other via a Lennard-Jones potential

𝜎 12  𝜎 6
 
𝑈(𝑟) = 4𝜖 − (16.14)
𝑟 𝑟

where 𝑟 is the distance between two particles, 𝜖 the interaction strength


and 𝜎 the optimum distance between the particle. The variation of
the potential is represented in Fig. XX. We want to simulate how these
particles aggregates with each other. For two particles at zero temperature
the problem is trivial. The Metropolis Monte Carlo method allows finding
the solution of this problem for a large number of particles and at finite
temperature. The algorithm proceeds as follows:

(1) Generate an initial configuration of 𝑁 particles distributed in the


2D space. To avoid difficulty when particles are very close and
therefore 𝑈 is very large, the particles are usually distributed along
a regular mesh.
(2) Consider one particle with a randomly chosen index 𝑖 and randomly
change its position to r’𝑖 = r𝑖 + 𝛿
(3) Compute the energy change Δ𝐸 that is induced by this trial movement
(4) if Δ𝐸 < 0 the move is accepted and we return to (2)
(5) if Δ𝐸 > 0, a random number 𝑟 is generated between 0 and 1.
(6) if 𝑟 < exp(−Δ𝐸/𝑘 𝐵 𝑇), the move is accepted. Otherwise the move is
rejected. In both case we return to (2).

As you can see the algorithm is almost identical to the previous case. The
principal difficulties arises in the calculation of the total energy that can
be extremely time consuming if all the pair wise interaction are calculated
16 Monte Carlo schemes 152

at each Monte-Carlo moves. Efficient tricks have been developed to avoid


that computational load and are presented in the next section in the
framework of molecular dynamics simulations.

16.5 Exercises : Algorithm

Approximating 𝜋 by MC

In this first exercise, we propose to evaluate 𝜋 with the help of a Monte


Carlo technique. Consider the figure XX where a circle of radius 𝑟 is
embedded in a square of length 2𝑟 . The circle and the square have an
area of 𝜋𝑟 2 and 4𝑟 2 respectively. The ratio between the two is therefore
of 𝑝 = 𝜋/4. Let’s focus on the upper right quarter of the figure. To
evaluate 𝑝𝑖 , generate a series of points ( 𝑥 , 𝑦 ) randomly distributed (with
an uniform distribution) in the upper right quarter of the figure. For each
point decide if it belongs to the circle, i.e. 𝑥 2 + 𝑦 2 ≤ 𝑟 2 . The ratio of points
inside the circle to all the points generated during the procedure should
equal 𝜋/4. Compare the accuracy of this approach to the case where a
regular mesh of points is used. Make sure to use the same number of
points in both calculations.

Stock Price

Create a program that simulate the evolution of 𝑀 realizations of the


price of a stock following the equation


𝑥 𝑛 = 𝑥 𝑛−1 + Δ𝑡𝜇𝑥 𝑛−1 + 𝜎𝑥 𝑛−1 𝑟𝑛−1 Δ𝑡 (16.15)

Take an initial price of $50, and explore different values for the growth
rate and the volatility.

Mixing of particles

We here consider a box divided in two equal size part by a wall. One
half of the box contains 𝑁 particles that are uniformly distributed in a
random fashion. We now remove suddenly the wall and we want to study
how the molecule fills in the entire box. We can simulate this process
using a 2D random walk in a box. We set the dimension of the box to be
𝐿𝑋 = 1 ad 𝐿𝑌 = 1 for simplicity.
To initiate the simulations you will randomly place 𝑁 particles (N=100 is
a good way to start) in the region defined by the boundary [0; 12] × [0; 1].
Implement a 2D random walk dynamics to simulate the trajectory of
the particles. The particles cannot escape from the box. Hence a particle
at one edge of the domains cannot go in all 4 directions. Visualize the
results as an animation.
16 Monte Carlo schemes 153

Figure 16.9: Integration with the Monte


Carlo technique. One can determine the
value of 𝑝𝑖 using random numbers
Appendix
Exponential of a matrix A
Demonstration of eq. 13.32

The exponential of a matrix is given by

K2 𝑡 2 K3 𝑡 3
𝑒 K𝑡 = I + K 𝑡 + + + ... (A.1)
2! 3!

We can use this expression to demonstrate that eq. 13.32 is a valid solution
of eq. 13.29

𝑑 𝑑 K𝑡
C = 𝑒 C0 (A.2)
𝑑𝑡 𝑑𝑡 
𝑑 K2 𝑡 2 K3 𝑡 3

= I + K𝑡 + + + ... C0 (A.3)
𝑑𝑡 2! 3!
K3 𝑡 2
 
= K+K 𝑡+2
+ ... C0 (A.4)
2!
K2 𝑡 2
 
= K I + K𝑡 + + ... C0 (A.5)
2!
= K 𝑒 K𝑡 C0 = KC (A.6)

Exponential of a matrix

We have seen in chapter ?? that a nonsingular matrix of order 𝑛 has 𝑛


eigenvalues and 𝑛 eigenvectors. Hence such matrix respect

𝐴 = 𝑈Λ𝑈 −1 (A.7)

where the 𝑖 -th column of 𝑈 contains the 𝑖 -th eigenvector and where Λ
is a diagonal matrix containing the eigenvalues of 𝐴. The characteristic
equations 𝐴𝑋 = 𝑋Λ gives by right-multiplying by 𝑈 −1

𝐴𝑋𝑋 −1 = 𝐴 = 𝑈Λ𝑈 −1 (A.8)

Any power of the matrix 𝐴𝑛 can then be expressed as

𝐴𝑛 = [𝑈Λ𝑈 −1 ]𝑛 = 𝑈Λ𝑛 𝑈 −1 (A.9)

The exponential of the matrix then reads


A Exponential of a matrix 156

1
𝑒𝐴 = 𝐼 + 𝑈Λ𝑈 −1 + 𝑈Λ2 𝑈 −1 + ... (A.10)
 2! 
1 2
= 𝑈 𝐼+Λ+ Λ + ... 𝑈 −1 (A.11)
2!
= 𝑈 𝑒 Λ 𝑈 −1 (A.12)
Index

+=, 14 fitting curves, 36


:, 27, 54 float(), 46
%, 6 for loop, 14
4th order Runge-Kutta method, 121 format(), 6
forward approximation, 86
anaconda, 3 forward Euler method, 119
append(), 13, 15 function definition, 19
argument of function, 19 functions, 19
array, 53
arrays, 24 Gauss-Jordan elimination, 65
astype(), 29 Gauss-Jordan implementation, 67
augmented matrix, 65 Gauss-Seidel algorithm, 74
autonomous differential equation, 110 Gaussian distribution , 142
Gaussian quadrature, 77, 81
backward approximation, 87 GUI, 2
bisection method, 94, 104
boolean, 13 histograms, 47, 49
boundary conditions, 130
boundary value problem, 130, 134 if-else statement, 16
branching, 16 implicit method, 120
import statement, 8
canonical form, 110, 132 indentation, 12
centered difference approximation, 87, 98 indexing of arrays, 26
character-separated values, 44, 45 initial guess, 71, 130
classification of ODEs, 109 initial value problems, 118
comments, 5 insert(), 13
composite methods, 77, 83 integration, 77
contour plots, 32 inverse of matrix, 58
copying numpy arrays, 29 iterative algorithm, 71
cost-function, 131
coupled linear equations, 63 Jacobi algorithm, 71
coupled non-linear equations, 99 Jacobian matrix, 100
csv, 45
len(), 14
curve fitting, 41
linalg, 64
del, 14 [Link](), 40
determinant of matrix, 58, 59 linear algebra, 57
Dirichlet boundary condition, 136 linear differential equations, 109, 112
discretization, 118 linear equations, 63
distillation column, 63 linear regression, 36
list, 11, 13
eigenvalue, 58, 113 loops, 11
eigenvector, 58, 113 Lotka-Volterra equation, 124
elif, 17
explicit method, 120 math module, 8
exponential of matrix, 112 matplotlib module, 31
[Link], 38
factorial, 19, 95 matrices, 25, 53
finite difference method, 134 matrix exponentiation, 112
first derivative, 86 matrix indexing, 53
INDEX 158

midpoint approximation, 78, 143 plotting curves, 31


modified Euler method, 120 polynomial, 37
modules, 8 polynomial regression, 40
Monte Carlo integration, 143 precedence, 8
Monte Carlo method, 141 predator-prey model, 124
multiple integrals, 82 print(), 6
multiple non-linear equations, 103 pyplot, 31
multiple return values, 21 [Link](), 49

nested loops, 56 random number, 141


Newton method, 97, 104 random number generator, 141
Newton-Cotes formulas, 77, 81 range(), 15
non-linear equations, 94 reading data, 44
normal distribution, 142 readlines(), 46
numerical differentiation, 86 rectangle approximation, 77, 119
numerical integration, 77 recursion, 59, 95
numpy array, 53 recursive bisection, 96
numpy module, 24 reduced Newton method, 103
[Link](), 25 return(), 19
[Link](), 24, 53 Runge-Kutta methods, 120
[Link](), 29
scatter plot, 38
[Link], 38
scipy, 36
[Link](), 57
scipy module, 24
[Link](), 56
[Link] submodule, 82
[Link](), 45
[Link](), 82
[Link](), 57
[Link]𝑖 𝑣𝑝(), 123
[Link](), 55
[Link](), 113
[Link], 58, 64
[Link].curve_fit, 41
[Link](), 64, 135
second derivative, 90
[Link](), 25
seed-number, 141
[Link](), 44
shooting method, 130, 137
[Link](), 57
Simpson’s rule, 80, 85
[Link](), 48
slicing matrices, 54
[Link](), 33
slicing of arrays, 26
[Link](), 25, 56
slicing vectors, 54
[Link](), 40
space-separated values, 44
[Link](), 142
Spyder, 2
[Link](), 55, 141
spyder, 3
[Link](), 143
statistics, 47
[Link](), 142
string, 6
[Link].random_integers(), 143
structured programming, 19
[Link](), 141
substitution, 110
[Link](), 142
system of linear equations, 63, 100, 135
[Link](), 54, 55
[Link](), 89 Taylor expansion, 86, 90, 97, 99, 131
[Link](), 44 throwing darts, 145
[Link](), 48 trapezoid approximation, 79
[Link](), 38 trapezoidal approximation, 120
[Link](), 48 two-dimensional arrays, 25
[Link](), 25, 56
variables, 4
open(), 46 vectorization, 27, 57
optional arguments, 20 vectors, 53

partial derivatives, 91 while loop, 11

You might also like