0% found this document useful (0 votes)
3 views69 pages

Chapter 13

Chapter 13 of 'Advanced Engineering Mathematics' discusses differential calculus for functions of several variables, covering concepts such as n-tuples, distance functions, neighborhoods, and connected sets. It introduces limits, continuity, and partial derivatives, including the conditions under which they exist and theorems related to their continuity. The chapter also explores examples and applications of these concepts in mathematical analysis.

Uploaded by

hqx45qq9f8
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views69 pages

Chapter 13

Chapter 13 of 'Advanced Engineering Mathematics' discusses differential calculus for functions of several variables, covering concepts such as n-tuples, distance functions, neighborhoods, and connected sets. It introduces limits, continuity, and partial derivatives, including the conditions under which they exist and theorems related to their continuity. The chapter also explores examples and applications of these concepts in mathematical analysis.

Uploaded by

hqx45qq9f8
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Chapter 13: Differential Calculus

of
Functions of Several Variables

Advanced Engineering Mathematics, 2005, 2nd Ed., Prentice Hall


[M. D. Greenberg]

Dr. [Link], Bilkent University, UMRAM


Preliminaries (I)
➢ A set of 𝑛-tuples is:
ℝ𝑛 = 𝑥ҧ = 𝑥1, 𝑥2, ⋯ , 𝑥𝑛 , 𝑥𝑗 ∈ ℝ
➢ A function of 𝑛-variables is:
𝑓: ℝ𝑛 → ℝ: 𝑥ҧ ↦ 𝑓 𝑥ҧ = 𝑓 𝑥1, ⋯ , 𝑥𝑛
➢ The distance from 𝑥ҧ to a fixed “point “ 𝑥ഥ′ is such a function:
1
𝑑 𝑥,ҧ 𝑥ഥ′ ≔ 𝑥1′
− 𝑥1 2 ′
+ ⋯ + 𝑥𝑛 − 𝑥𝑛 2 2

where 𝑥ҧ = 𝑥1, 𝑥2, ⋯ , 𝑥𝑛 ഥ=


and 𝑥′ 𝑥1′, 𝑥2′, ⋯ , 𝑥𝑛′ .
Note that:
𝑛=1 𝑑 𝑥,ҧ 𝑥ഥ′ = |𝑥1′ − 𝑥1|

𝑛=2 𝑑 𝑥,ҧ 𝑥ഥ′ = 𝑥1′ − 𝑥1 2 + 𝑥2′ − 𝑥2 2


Dr. [Link], Bilkent University, UMRAM

𝑛 = 3 𝑑 𝑥,ҧ 𝑥ഥ′ = 𝑥1′ − 𝑥1 2 + 𝑥2′ − 𝑥2 2 + 𝑥3′ − 𝑥3 2

2
Preliminaries (II)
➢ The neighborhood of 𝑥ഥ′ with “radius” r is
𝑁 𝑥ഥ′, 𝑟 ∶= 𝑥ҧ ∈ ℝ𝑛 : 𝑑 𝑥,ҧ 𝑥ഥ′ < 𝑟
𝑛 = 1: 𝑛 = 2:

𝑛 = 3 : A sphere of radius 𝑟, centered at 𝑥ഥ′


➢ A set 𝑆 ⊆ ℝ𝑛 is connected if each pair of points in 𝑆 can be joined by a
finite number of line segments connected end-to-end.

Not connected
Dr. [Link], Bilkent University, UMRAM

connected
➢ A straight line in 𝑛-dimensional space has the parametric equation:
𝑥1 = 𝑎1 + 𝑏1𝑡, ⋯ , 𝑥𝑛 = 𝑎𝑛 + 𝑏𝑛 𝑡 ; 𝑎𝑗 , 𝑏𝑗 ∈ ℝ , 𝑡 ∈ ℝ 𝑜𝑟 𝛼, 𝛽
3
Preliminaries (III)
➢ A point 𝑥ҧ ∈ ℝ𝑛 is an interior point of 𝑆 if :
𝑥ҧ ∈ 𝑆 & 𝑁 𝑥,ҧ 𝑟 ⊆ 𝑆 for some 𝑟 > 0.
It is a boundary point of 𝑆 if 𝑁(𝑥,ҧ 𝑟), for any 𝑟 > 0, contains some points
from 𝑆 as well as outside of 𝑆.

➢ A connected set is an open set or domain if


it contains none of its boundary points.
It is a closed set if it contains all of its boundary points.

𝑥1
𝑆
Dr. [Link], Bilkent University, UMRAM

𝑥2

Fact: A connected set 𝑆 is an open iff every point of 𝑆 is an interior point.


4
Ex. 2, p.616

Examples
➢ 𝑁(𝑥ഥ′ , 𝑟) is an open set for every 𝑟 > 0.
➢ 𝑁(𝑥ഥ′ , 𝑟) together with its boundary (which is a circle in 𝑛 = 2,
spherical shell in 𝑛 = 3, etc.) is a closed set.
➢ Note that ℝ = {𝑥: −∞ < 𝑥 < +∞} is both open and closed !
It has no boundary points so that we can say “it contains none” as
well as “it contains all” equally legitimately.
➢ Neither open nor closed !

𝑆
Dr. [Link], Bilkent University, UMRAM

5
Limit and Continuity (I)
➢ The limit of 𝑓(𝑥)ҧ as 𝑥ҧ approaches 𝑥ഥ′ is said to be 𝐿 and written as
lim 𝑓 𝑥ҧ = 𝐿,
ҧ 𝑥ƴ ҧ
𝑥→
if whenever an 𝜖 > 0 is given, one can find a 𝛿 > 0 such that:
0 < 𝑑 𝑥,ҧ 𝑥ഥ′ < 𝛿 ⇒ 𝑓 𝑥ҧ − 𝐿 < 𝜖
for every 𝑥ҧ ∈ 𝑁(𝑥ഥ′, 𝛿) (and 𝑥ҧ ≠ 𝑥′) that is also in the domain of definition of
𝑓(𝑥).
ҧ
Thus, no matter how small an 𝜖 > 0 is given, there is a 𝛿-neighborhood of 𝑥ഥ′ such
that 𝑓(𝑥)ҧ is 𝜖-close to 𝐿 for every 𝑥ҧ in that neighborhood.
Note: “0 <“ part in 0 < 𝑑 𝑥,ҧ 𝑥ഥ′ < 𝛿 says that 𝑓(𝑥ഥ′) need not equal 𝐿!
𝑥2 0 ≤ 𝑥 ≤ 3, 𝑥 ≠ 2
𝑓 𝑥 =ቊ ⇒ lim 𝑓 𝑥 = 4
12 ,𝑥 = 2 𝑥→2
Dr. [Link], Bilkent University, UMRAM

➢ 𝑓 𝑥ҧ is continuous at 𝑥ഥ′ if :
lim 𝑓 𝑥ҧ = 𝑓(𝑥ഥ′) .
ҧ 𝑥ƴ ҧ
𝑥→

6
Ex. -, p.618

Examples
1 𝑥>0
➢ 𝐻 𝑥 =ቐ 0 𝑥 < 0 is continuous everywhere except at 𝑥 =
1/2 𝑥 = 0
0, where the lim 𝐻(𝑥) does not exist.
𝑥→0
1
➢ 𝑓 𝑥ҧ = is continuous everywhere except at 𝑥ҧ = 0
𝑥12 +𝑥22 +𝑥32
because neither lim 𝑓(𝑥)ҧ nor 𝑓(0) exists.
ҧ
𝑥→0
➢ 𝑓 𝑥1 , 𝑥2 = 𝑥12 + 𝑒 𝑥2 sin(𝑥1 𝑥22 )
is continuous in the whole 𝑥1 𝑥2 -
plane since it is the product, sum, composition of continuous
functions.
Dr. [Link], Bilkent University, UMRAM

7
Chapter 13: Differential Calculus
of
Functions of Several Variables

Advanced Engineering Mathematics, 2005, 2nd Ed., Prentice Hall


[M. D. Greenberg]

Dr. [Link], Bilkent University, UMRAM


Partial Derivatives (I)
Suppose the domain of definition of 𝑓(𝑥, 𝑦) includes a neighborhood of
(𝑥0, 𝑦0). Then:
𝜕𝑓 𝑓 𝑥0 + Δ𝑥, 𝑦0 − 𝑓 𝑥0, 𝑦0
𝑓𝑥 𝑥0, 𝑦0 = ቚ ≔ lim ,
𝜕𝑥 𝑥0,𝑦0 Δ𝑥→0 Δ𝑥
𝜕𝑓 𝑓 𝑥0, 𝑦0 + Δ𝑦 − 𝑓 𝑥0, 𝑦0
𝑓𝑦 𝑥0, 𝑦0 = ቚ ≔ lim ,
𝜕𝑦 𝑥0,𝑦0 Δ𝑦→0 Δ𝑦
are the partial derivatives of 𝑓(𝑥, 𝑦) with respect to 𝑥 and 𝑦 at (𝑥0, 𝑦0).
If 𝑓𝑥 , 𝑓𝑦 exist in a neighborhood of (𝑥0, 𝑦0), then they are functions of
(𝑥, 𝑦) and one can define second-order derivatives as:
𝜕 𝜕𝑓 𝜕2 𝑓 𝜕 𝜕𝑓 𝜕2 𝑓
𝑓𝑥𝑥 = = 2 , 𝑓𝑦𝑥 = = ,
𝜕𝑥 𝜕𝑥 𝜕𝑥 𝜕𝑥 𝜕𝑦 𝜕𝑥𝜕𝑦
Dr. [Link], Bilkent University, UMRAM

𝜕 𝜕𝑓 𝜕2 𝑓 𝜕 𝜕𝑓 𝜕2 𝑓
𝑓𝑦𝑦 = = 2 , 𝑓𝑥𝑦 = = .
𝜕𝑦 𝜕𝑦 𝜕𝑦 𝜕𝑦 𝜕𝑥 𝜕𝑦𝜕𝑥

9
Partial Derivatives : Example (I)
➢ Theorem 13.3.1: If 𝑓𝑥 , 𝑓𝑦 , 𝑓𝑥𝑦 , 𝑓𝑦𝑥 are all continuous in a neighborhood
of (𝑥0, 𝑦0), then 𝑓𝑥𝑦 = 𝑓𝑦𝑥 at (𝑥0, 𝑦0).
➢ Remark: Unlike functions of one variable, 𝑓𝑥 and 𝑓𝑦 may exist at
(𝑥0, 𝑦0) without 𝑓 being continuous at (𝑥0, 𝑦0) !
1 , 𝑥 = 0 or 𝑦 = 0
▪ Ex. 13.3.7: 𝑓 𝑥, 𝑦 = ቊ
0 , 𝑥 ≠ 0 and 𝑦 ≠ 0
𝑓 Δ𝑥, 0 − 𝑓 0,0 1−1
𝑓𝑥 0,0 = lim = lim ,
Δ𝑥→0 Δ𝑥 Δ𝑥→0 Δ𝑥
𝑓 0, Δ𝑦 − 𝑓 0,0 1−1
𝑓𝑦 0,0 = lim = lim ,
Δ𝑦→0 Δ𝑦 Δ𝑦→0 Δ𝑦
Dr. [Link], Bilkent University, UMRAM

but lim 𝑓(𝑥, 𝑦) does not exists since it depends on how we approach the
𝑥,𝑦→0
origin. It follows that 𝑓(𝑥, 𝑦) is not continuous at (0,0).

10
Partial Differential Equation
➢ If 𝑓 = 𝑥12 + ⋯ + 𝑥𝑛2 𝛼 and
𝑓𝑥1𝑥1 + ⋯ + 𝑓𝑥𝑛𝑥𝑛 = 0, (A)
then find 𝛼 ∈ ℝ. We compute first 𝑓𝑥𝑘𝑥𝑘 :
𝜕
𝑓𝑥𝑘𝑥𝑘 = 2𝛼𝑥𝑘 𝑥12 + ⋯ + 𝑥𝑛2 𝛼−1
𝜕𝑥𝑘
= 2𝛼 𝑥12 + ⋯ + 𝑥𝑛2 𝛼−1 + 4𝛼 𝛼 − 1 𝑥𝑘2 𝑥12 + ⋯ + 𝑥𝑛2 𝛼−2
= 2𝛼 𝑥12 + ⋯ + 𝑥𝑛2 𝛼−2[𝑥12 + ⋯ + 𝑥𝑛2 + 2 𝛼 − 1 𝑥𝑘2]
Now (A) holds if and only if
𝑛

෍[𝑥12 + ⋯ + 𝑥𝑛2 + 2 𝛼 − 1 𝑥𝑘2] = 𝑛 + 2 𝛼 − 1 𝑥12 + ⋯ + 𝑥𝑛2 = 0,


Dr. [Link], Bilkent University, UMRAM

𝑘=1
2−𝑛
2−𝑛
which in turn holds iff 𝛼 = or 𝑓 = 𝑥12 + ⋯ + 𝑥𝑛2 2 .
2

11
Chapter 13: Differential Calculus
of
Functions of Several Variables

Advanced Engineering Mathematics, 2005, 2nd Ed., Prentice Hall


[M. D. Greenberg]

Dr. [Link], Bilkent University, UMRAM


Composite Functions & Chain Rule (I)
Let us first consider a function of one-variable 𝑓(𝑥), which is differentiable
over 𝑋 = 𝑥: 𝑎 < 𝑥 < 𝑏 .
Suppose 𝑥 = 𝑥(𝑡) is in turn differentiable over 𝑇 = {𝑡: 𝛼 < 𝑡 < 𝛽}, such
that 𝑥 𝑡 ∈ X , ∀𝑡 ∈ 𝑇. Then, 𝐹 𝑡 ≔ 𝑓(𝑥 𝑡 ) is defined and is
differentiable over 𝑇 with
𝑑𝐹 𝑑𝑓 𝑑𝑥
= ⋅
𝑑𝑡 𝑑𝑥 𝑑𝑡
The 𝐹 is called the “composite function” of 𝑡 or “composition of 𝑓 and 𝑥”.
This 1-D result can be generalized.
x(t) f(x)
Dr. [Link], Bilkent University, UMRAM

α 𝛽 a b f(x(t)) ℝ

13
Ex. 2, p.626

Composite Functions & Chain Rule


➢ Theorem 13.4.1: Let 𝑓 𝑥, 𝑦 , 𝑓𝑥 𝑥, 𝑦 , and 𝑓𝑦 (𝑥, 𝑦) be continuous
over an open region (set) ℛ in 𝑥𝑦-plane. Let 𝑥 = 𝑥 𝑡 , 𝑦 = 𝑦(𝑡)
be differentiable over an interval 𝑇 ⊆ ℝ such that 𝑥 𝑡 , 𝑦 𝑡 ∈
ℛ for ∀𝑡 ∈ 𝑇.
Then, 𝐹 𝑡 ≔ 𝑓 𝑥 𝑡 , 𝑦 𝑡 is a differentiable function of t over 𝑇 and
𝑑𝐹 𝜕𝑓 𝑑𝑥 𝜕𝑓 𝑑𝑦
= ⋅ + ⋅ (24)
𝑑𝑡 𝜕𝑥 𝑑𝑡 𝜕𝑦 𝑑𝑡
Proof: See page 625 of book !
➢ Let 𝑟 𝑥, 𝑦 = 𝑥 2 𝑦 − 𝑒 2𝑦 , where 𝑥 = 3𝑡 2 , 𝑦 = sin(𝑡), 1 < 𝑡 < 4.
Now, 𝑅 𝑡 = 𝑟 𝑥 𝑡 , 𝑦 𝑡 = 9𝑡 4 sin 𝑡 − 𝑒 2 sin 𝑡 .
Dr. [Link], Bilkent University, UMRAM

Using (24) we have: 𝑅′ 𝑡 = 2𝑥𝑦 ⋅ 6𝑡 + 𝑥 2 − 2𝑒 2𝑦 ⋅ cos 𝑡


𝑅′ 𝑡 = 36𝑡 3 sin 𝑡 + 9𝑡 4 − 2𝑒 2 sin 𝑡 cos 𝑡
which
14
is equal to 𝑅′(𝑡) obtained by a direct derivation.
Chapter 13: Differential Calculus
of
Functions of Several Variables

Advanced Engineering Mathematics, 2005, 2nd Ed., Prentice Hall


[M. D. Greenberg]

Dr. [Link], Bilkent University, UMRAM


Taylor’s Formula in 1D (I)
Recall that in 1D case, by fundamental theorem of calculus, we can write
𝑥 𝑥
𝑓 𝑥 =𝑓 𝑎 +න 𝑓′ 𝑥 𝑑𝑥 𝑓′ 𝑥 = 𝑓′ 𝑎 + න 𝑓′′ 𝑥 𝑑𝑥
𝑟𝑒𝑝𝑙𝑎𝑐𝑒 𝑓 𝑏𝑦 𝑓′
𝑎 𝑎
for 𝑎, 𝑥 ∈ ℝ. Then, we can write
𝑥 𝑥
𝑓 𝑥 = 𝑓 𝑎 + න 𝑓′ 𝑎 + න 𝑓′′ 𝑥 𝑑𝑥 𝑑𝑥 ,
𝑎 𝑎
𝑥 𝑥
𝑓 𝑥 = 𝑓 𝑎 + 𝑓′ 𝑎 𝑥 − 𝑎 + න න 𝑓′′ 𝑥 𝑑𝑥𝑑𝑥 ,
𝑎 𝑎
This generalizes to (25):
𝑥 − 𝑎 𝑛−1
𝑓 𝑥 = 𝑓 𝑎 + 𝑓′ 𝑎 𝑥 − 𝑎 + ⋯ + 𝑓 𝑛−1 𝑎 + 𝑅𝑛 𝑥 ,
𝑛−1 !
Dr. [Link], Bilkent University, UMRAM

which is called the Taylor’s formula with 𝑛-th order error 𝑅𝑛(𝑥) given by
𝑛−𝑡𝑖𝑚𝑒𝑠
𝑥 𝑥 𝑛
𝑅𝑛 𝑥 ≔ ‫𝑓 𝑎׬ 𝑎׬‬
⋯ (𝑥) 𝑑𝑥 ⋯ 𝑑𝑥
16
Taylor’s Formula in 1D (II)
𝑛−𝑡𝑖𝑚𝑒𝑠
𝑥 𝑥 𝑛
𝑅𝑛 𝑥 ≔ ‫𝑓 𝑎׬ 𝑎׬‬
⋯ (𝑥) 𝑑𝑥 ⋯ 𝑑𝑥 (26)
provided that 𝑓(𝑥) is differentiable 𝑛 − 1-times in the interval [𝑎, 𝑥].
If its n-th derivative exists and is continuous in [𝑎, 𝑥] (or in 𝑥, 𝑎 ), then we
can show that, for some suitable 𝜉 ∈ [𝑎, 𝑥],
𝑓𝑛 𝜉
𝑅𝑛 𝑥 = 𝑥 − 𝑎 𝑛.
𝑛!
To see this, let 𝒎 and 𝑴 be the minimum and maximum of 𝑓 𝑛 (𝑡) for t ∈
[𝑎, 𝑥], both of which exist since 𝑓 𝑛 (𝑡) is continuous in that interval. Then,
(26) gives:
𝑥 𝑥 𝑥 𝑥
Dr. [Link], Bilkent University, UMRAM

𝒎 න ⋯ න 𝑑𝑥 ⋯ 𝑑𝑥 ≤ 𝑅𝑛 𝑥 ≤ 𝑴 න ⋯ න 𝑑𝑥 ⋯ 𝑑𝑥
𝑎 𝑎 𝑎 𝑎
𝑥−𝑎 𝑛 𝑥−𝑎 𝑛
𝒎 ≤ 𝑅𝑛 𝑥 ≤ 𝑴 .
𝑛! 𝑛!
17
Taylor’s Formula in 1D (III)
𝑥−𝑎 𝑛 𝑥−𝑎 𝑛
𝒎 ≤ 𝑅𝑛 𝑥 ≤ 𝑴
𝑛! 𝑛!
Now, since 𝑓 𝑛 (𝑡) must take on all values between its minimum 𝒎 and its
maximum 𝑴, by its continuity, it follows that for some 𝜉 ∈ [𝑎, 𝑥] :
𝑓𝑛 𝜉
𝑅𝑛 𝑥 = 𝑥−𝑎 𝑛
𝑛!
as claimed. This expression is called the Lagrange remainder of order 𝒏.
We thus have the result that, if one approximates 𝑓 𝑥 by (the
polynomial of order 𝑛 − 1 in 𝑥)
′ 𝑛−1
𝑥 − 𝑎 𝑛−1
𝑓 𝑥 ≅ 𝑓 𝑎 + 𝑓 𝑎 𝑥 − 𝑎 + ⋯+ 𝑓 𝑎 ,
𝑛−1 !
Dr. [Link], Bilkent University, UMRAM

using Taylor’s formula (25), then the “error of approximation” is given by


the absolute value
𝑅𝑛 𝑥 .
18
Ex. 1, p.632

Taylor Approximation
➢ Approximate 𝑒−𝑥 over 0.7 ≤ 𝑥 ≤ 1.3 by a second order polynomial of
𝑥 and obtain a bound on the error.
Let us choose 𝑎 = 1 so that (25) with 𝑛 = 3 gives:
−1
𝑒
𝑓 𝑥 = 𝑒−𝑥 = 𝑒−1 − 𝑒−1 𝑥 − 1 + 𝑥 − 1 2 + 𝑅3 𝑥 ,
2
with
𝑒−𝑧
𝑅3 𝑥 = − 𝑥 − 1 3, 𝑧 ∈ 1, 𝑥 or 𝑧 ∈ 𝑥, 1 ,
3!
whichever inequality 1 ≤ 𝑥 or 𝑥 ≤ 1 holds .
Note that:
𝑒−𝑧 𝑒 −0.7
Dr. [Link], Bilkent University, UMRAM

𝑅3 𝑥 = 𝑥−1 3 ≤ 0.3 3 ≅ 0.0022 ≔ 𝑒𝑟𝑟


6 6
This means that the error of approximation made in the 2nd order
polynomial above, is at most ±𝑒𝑟𝑟 as long as 0.7 ≤ 𝑥 ≤ 1.3.
19
Taylor Series 1D
If derivatives of all orders of 𝑓(𝑥) exist and if
lim 𝑅𝑛 𝑥 = 0 ,
𝑛→∞
for all values of 𝑥 in an interval 𝐼 including 𝑎, then

𝑘
𝑓 𝑎 𝑘
𝑓 𝑥 =෍ 𝑥−𝑎
𝑘!
𝑘=0
is the Taylor series of 𝑓(𝑥) about 𝑎 and it is convergent for all 𝑥 ∈ 𝐼.
Note:
▪ 𝑥 − 𝑎 0 ≡ 1 even if 𝑥 = 𝑎,
▪ 0! = 1.
Dr. [Link], Bilkent University, UMRAM

20
Ex. 3, p.634

Taylor Series Example


➢ Consider the Taylor series of 𝑓 𝑥 = 𝑒−𝑥 around 𝑎 = 0.

2 3 𝑘
𝑥 𝑥 𝑥
𝑒−𝑥 = 1 − 𝑥 + − + ⋯ = ෍ −1 𝑘 .
2! 3! 𝑘!
𝑘=0
Now the Lagrange remainder of order 𝑛 is
𝑓𝑛 𝜉 𝑒−𝜉 𝑛
𝑅𝑛 𝑥 = 𝑥 − 𝑎 𝑛 = −1 𝑛 𝑥
𝑛! 𝑛!
𝑒 −𝜉 𝑛
lim 𝑅𝑛 𝑥 = lim 𝑥 =0
𝑛→∞ 𝑛→∞ 𝑛!
for all 𝜉 ∈ [0, 𝑥] (or [𝑥, 0]) and for all 𝑥 ∈ ℝ.
It follows that the Taylor series of 𝑒−𝑥 exists for all finite values of 𝑥.
Dr. [Link], Bilkent University, UMRAM

21
Multi-Variable Taylor’s Formula 2D (I)
Given 𝑓(𝑥, 𝑦), let ℛ be an open set in 𝑥𝑦-plane and consider a line
extending from (𝑎, 𝑏) to 𝑥0, 𝑦0 ,
we can parameterize any point (𝑥, 𝑦) on the line by
𝑥 = 𝑎 + 𝑥0 − 𝑎 𝑡, 𝑦 = 𝑏 + 𝑦0 − 𝑏 𝑡,
where 𝑡 ∈ [0,1].
The composite function 𝐹(𝑡) given by
𝐹 𝑡 = 𝑓(𝑥 𝑡 , 𝑦 𝑡 ) has the expression:
𝐹 𝑡 = 𝑓 𝑎 + 𝑥0 − 𝑎 𝑡, 𝑏 + 𝑦0 − 𝑏 𝑡 ,
and it can be expanded about 𝑡 = 0 using 𝑛 −th order Taylor’s formula as
𝑡𝑛−1
𝐹 𝑡 =𝐹 0 + 𝐹′ 0 𝑡 + ⋯+ 𝐹 𝑛−1
0 + 𝑅𝑛 𝑡 , (28)
𝑛−1 !
Dr. [Link], Bilkent University, UMRAM

provided that 𝑛-th order derivative of 𝐹 exists and is continuous .


Here, we need to evaluate the coefficients 𝐹(0) , 𝐹′(0) and the remainder.
Let us first make some preparation.
22
Multi-Variable Taylor’s Formula 2D (II)
We now define an operator 𝐷 via chain differentiation rule
𝑑𝐹 𝜕 𝑑𝑥 𝜕 𝑑𝑦 𝜕 𝜕
= + 𝑓 ⇒ 𝑥0 − 𝑎 + 𝑦0 − 𝑏 =: 𝐷 (29)
𝑑𝑡 𝜕𝑥 𝑑𝑡 𝜕𝑦 𝑑𝑡 𝜕𝑥 𝜕𝑦
It follows that

𝑑
𝐹 0 = 𝐹 𝑡 ቚ = 𝐷𝑓 𝑥, 𝑦 ቚ ,
𝑑𝑡 𝑡=0 𝑎,𝑏
′′
𝑑 𝑑
𝐹 0 = 𝐹 𝑡 ቚ = 𝐷2𝑓 𝑥, 𝑦 ቚ ,
𝑑𝑡 𝑑𝑡 𝑡=0 𝑎,𝑏

(𝑘)
𝑑 𝑑
𝐹 0 = ⋯ 𝐹 𝑡 ቚ = 𝐷𝑘 𝑓 𝑥, 𝑦 ቚ ,
𝑑𝑡 𝑑𝑡 𝑡=0 𝑎,𝑏
Dr. [Link], Bilkent University, UMRAM

𝑘
and the remainder becomes, for some (𝜉,𝜂) that is on the line,
1
𝑅𝑛 𝑡 = 𝐷𝑛 𝑓 𝑥, 𝑦 ቚ .
𝑛! 𝜉,𝜂
23
Taylor’s Formula in 2D
Putting these expressions into
𝑡𝑛−1
𝐹 𝑡 =𝐹 0 + 𝐹′ 0 𝑡 + ⋯+ 𝐹 𝑛−1 0 + 𝑅𝑛 𝑡 , (28)
𝑛−1 !
and setting 𝑡 = 1, and noting that
𝐹 1 = 𝑓 𝑥0, 𝑦0 , 𝐹 0 = 𝑓 𝑎, 𝑏 ,
we obtain
1 1
𝑓 𝑥0,𝑦0 = 𝑓 𝑎,𝑏 + 𝐷𝑓|𝑎,𝑏 + ⋯+ 𝐷𝑛−1𝑓|𝑎,𝑏 + 𝑅𝑛 (30)
1! 𝑛−1 !
with
1 𝑛
𝑅𝑛 = 𝐷 𝑓ቚ
𝑛! 𝜉,𝜂
Dr. [Link], Bilkent University, UMRAM

for some values of 𝜉 ∈ 𝑎, 𝑥 or [𝑥, 𝑎] and 𝜂 ∈ 𝑏, 𝑦 or 𝑦, 𝑏 .


This is the “Taylor’s formula in two variables” with Lagrange remainder.

24
2D Mean Value Theorem
Let us replace (𝑥0, 𝑦0) with (𝑥, 𝑦) for notational simplicity and rewrite
the operator 𝐷 as
𝜕 𝜕
𝐷 ∶= 𝑥 − 𝑎 + 𝑦−𝑏 .
𝜕𝑥 𝜕𝑦
Then, Taylor’s formula for two variables will be:
1 1 𝑛−1
1 𝑛
𝑓 𝑥,𝑦 = 𝑓 𝑎,𝑏 + 𝐷𝑓 ቚ + ⋯+ 𝐷 𝑓ቚ + 𝐷 𝑓 ቚ
1! 𝑎,𝑏 𝑛−1 ! 𝑎,𝑏 𝑛! 𝜉,𝜂
For 𝑛 = 1, we get
𝑓 𝑥, 𝑦 = 𝑓 𝑎, 𝑏 + 𝐷𝑓 𝑥, 𝑦 |𝜉,𝜂

𝑓 𝑥, 𝑦 = 𝑓 𝑎, 𝑏 + 𝑓𝑥 𝜉, 𝜂 𝑥 − 𝑎 + 𝑓𝑦 𝜉, 𝜂 𝑦 − 𝑏
Dr. [Link], Bilkent University, UMRAM

This is the “Two-variable mean value theorem” ( Theorem 13.5.2).

25
Taylor’s Formula in 2D, 𝑛 = 2
Now, we set 𝑛 = 2 and obtain
1 1 2
𝑓 𝑥,𝑦 = 𝑓 𝑎,𝑏 + 𝐷𝑓 ቚ + 𝐷 𝑓 ቚ
1! 𝑎,𝑏 2! 𝜉,𝜂
Let us first calculate 𝐷2𝑓:
𝐷2𝑓 = 𝐷 𝐷𝑓
𝜕 𝜕 𝜕 𝜕
= 𝑥−𝑎 + 𝑦−𝑏 𝑥−𝑎 + 𝑦−𝑏 𝑓
𝜕𝑥 𝜕𝑦 𝜕𝑥 𝜕𝑦
= 𝑥 − 𝑎 2𝑓𝑥𝑥 + 2 𝑥 − 𝑎 𝑦 − 𝑏 𝑓𝑥𝑦 + 𝑦 − 𝑏 2𝑓𝑦𝑦 .
Substituting in above equation, we get
𝑓 𝑥, 𝑦 = 𝑓 𝑎, 𝑏 + 𝑓𝑥 𝑎, 𝑏 𝑥 − 𝑎 + 𝑓𝑦 𝑎, 𝑏 𝑦 − 𝑏
Dr. [Link], Bilkent University, UMRAM

1
+ [𝑓𝑥𝑥 𝜉, 𝜂 𝑥 − 𝑎 2 + 2𝑓𝑥𝑦 𝜉, 𝜂 𝑥 − 𝑎 𝑦 − 𝑏 + 𝑓𝑦𝑦 𝜉, 𝜂 𝑦 − 𝑏 2]
2
for some 𝜉, 𝜂.
26
Taylor Series, 2D
The Taylor series of 𝑓(𝑥, 𝑦) about (𝑎, 𝑏) is defined as
1 1 2
𝑓 𝑥,𝑦 = 𝑓 𝑎,𝑏 + 𝐷𝑓 𝑥,𝑦 ቚ + 𝐷 𝑓 𝑥,𝑦 ቚ + ⋯
1! 𝑎,𝑏 2! 𝑎,𝑏
Using
𝜕 𝜕
𝐷𝑓 = 𝑥 − 𝑎 + 𝑦−𝑏 𝑓 = 𝑥 − 𝑎 𝑓𝑥 + 𝑦 − 𝑏 𝑓𝑦 ,
𝜕𝑥 𝜕𝑦
𝐷2𝑓 = 𝐷 𝐷𝑓 = 𝑥 − 𝑎 2𝑓𝑥𝑥 + 2 𝑥 − 𝑎 𝑦 − 𝑏 𝑓𝑥𝑦 + 𝑦 − 𝑏 2𝑓𝑦𝑦 ,
one can write
𝑓 𝑥, 𝑦 = 𝑓 𝑎, 𝑏 + 𝑓𝑥 𝑎, 𝑏 𝑥 − 𝑎 + 𝑓𝑦 𝑎, 𝑏 𝑦 − 𝑏
1
+ 𝑓𝑥𝑥 𝑎, 𝑏 𝑥 − 𝑎 2 + 2𝑓𝑥𝑦 𝑎, 𝑏 𝑥 − 𝑎 𝑦 − 𝑏 + 𝑓𝑦𝑦 𝑎, 𝑏 𝑦 − 𝑏 2
2
Dr. [Link], Bilkent University, UMRAM

+⋯ถ ⋯
𝐻.𝑂.𝑇
Similar to 1D series, it exists provided the n-th remainder goes to zero as n
goes to infinity.
27
Ex. 5, p.639

2D Taylor Series Example


➢ Derive the Taylor series of 𝑓 𝑥, 𝑦 = 𝑒 𝑥𝑦 about point (1, 2) up to and
including second-order terms.
We compute 𝑓𝑥 = 𝑦𝑒 𝑥𝑦 , 𝑓𝑦 = 𝑥𝑒 𝑥𝑦 ,
𝑓𝑥𝑥 = 𝑦2𝑒 𝑥𝑦 , 𝑓𝑥𝑦 = 𝑓𝑦𝑥 = 1 + 𝑥𝑦 𝑒 𝑥𝑦 , 𝑓𝑦𝑦 = 𝑥2𝑒 𝑥𝑦 ,
By using the first two terms (linear and quadratic) of the Taylor series:
𝑓 𝑥, 𝑦 = 𝑓 𝑎, 𝑏 + 𝑓𝑥 𝑎, 𝑏 𝑥 − 𝑎 + 𝑓𝑦 𝑎, 𝑏 𝑦 − 𝑏 +
1 2 2
𝑓 𝑎, 𝑏 𝑥 − 𝑎 + 2𝑓𝑥𝑦 𝑎, 𝑏 𝑥 − 𝑎 𝑦 − 𝑏 + 𝑓𝑦𝑦 𝑎, 𝑏 𝑦 − 𝑏
2 𝑥𝑥
We can then write
𝑒 𝑥𝑦 = 𝑒2 + 2𝑒2 𝑥 − 1 + 𝑒2 𝑦 − 2 +
2
𝑒
2𝑒2 𝑥 − 1 2 + 3𝑒2 𝑥 − 1 𝑦 − 2 + 𝑦−2 2
+⋯
Dr. [Link], Bilkent University, UMRAM

2
Note that 𝑅1 = 𝑥 − 1 𝜂𝑒𝜉𝜂 + 𝑦 − 2 𝜉𝑒𝜉𝜂 , and
1
𝑅2 = 𝑥 − 1 2𝜂2 + 2 𝑥 − 1 𝑦 − 2 𝜉𝜂 + 𝑦 − 2 2𝜉2 𝑒𝜉𝜂.
2
28
Example: Linear Approximation of 𝑓(𝑥, 𝑦)
➢ Find a linear approximation to 𝑓 𝑥, 𝑦 = 𝑒 𝑥+𝑦 − 3𝑦 over 0.8 ≤ 𝑥 ≤ 1.2
and −0.1 ≤ 𝑦 ≤ 0.1 and give a bound on the error of approximation.
We have
𝑓𝑥 = 𝑒 𝑥+𝑦 , 𝑓𝑦 = 𝑒 𝑥+𝑦 − 3, 𝑓𝑥𝑥 = 𝑓𝑦𝑦 = 𝑓𝑥𝑦 = 𝑓𝑦𝑥 = 𝑒 𝑥+𝑦 .
Let 𝑎 = 1, 𝑏 = 0 (midpoints of the given intervals).
Then,
𝑓 𝑥, 𝑦 ≅ 𝑓 𝑎, 𝑏 + 𝑓𝑥 𝑎, 𝑏 𝑥 − 𝑎 + 𝑓𝑦 𝑎, 𝑏 𝑦 − 𝑏
⇒ 𝑓 𝑥, 𝑦 ≅ 𝑒 + 𝑒 𝑥 − 1 + 𝑒 − 3 𝑦 = 𝑒𝑥 + 𝑒 − 3 𝑦
is a linear approximation with error 𝑅2
1
𝑅2 = 𝑒𝜉+𝜂 𝑥 − 1 2 + 2 𝑥 − 1 𝑦 + 𝑦2 ,
2
Dr. [Link], Bilkent University, UMRAM

for some 𝜉 ∈ [0.8,1.2] and 𝜂 ∈ [−0.1,0.1].


Choosing 𝜉, 𝜂 such that they maximize 𝑅2, we can find a bound on the error as
1
𝑅2 ≤ 𝑒1.3 0.2 2 + 2 0.2 0.1 + 0.1 2 = 0.165.
2
29
Chapter 13: Differential Calculus
of
Functions of Several Variables

Advanced Engineering Mathematics, 2005, 2nd Ed., Prentice Hall


[M. D. Greenberg]

Dr. [Link], Bilkent University, UMRAM


Implicit Function
An equation 𝑓 𝑥, 𝑦 = 0 can be viewed as defining 𝑦 as a function of
𝑥 implicitly, so that it can be written as
𝑓 𝑥, 𝑦 𝑥 = 0.
Example: 𝑓 𝑥, 𝑦 = 𝑥 2 + 4𝑦 2 − 4 = 0
⇒ 𝑦 = ± 1 − 𝑥 2 /4.
We have two functions of 𝑥 defined by 𝑓 𝑥, 𝑦 = 0, which are upper
2
and lower parts of an ellipse 𝑥 ൗ4 + 𝑦 2 = 1.
The domain of definition of both are −2 ≤ 𝑥 ≤ 2. If we are
interested in another domain that is not a subset, e.g., 3 ≤ 𝑥 ≤ 8,
then 𝑓 𝑥, 𝑦 = 0 defines no real function.
Dr. [Link], Bilkent University, UMRAM

Almost always, it is not possible to solve for 𝑦 explicitly, as in 2𝑥𝑦 +


sin 𝑦 = 3, which is a transcendental equation.
31
Implicit Function Theorem
Definition: A function 𝑓 is of class 𝐶 1 , denoted 𝑓 ∈ 𝐶 1 , if 𝑓𝑥 and 𝑓𝑦
exist and are continuous in an open set (domain). It is of class 𝐶 2 ,
denoted 𝑓 ∈ 𝐶 2 , if 𝑓𝑥 , 𝑓𝑦 and 𝑓𝑥𝑥 , 𝑓𝑦𝑦 , 𝑓𝑥𝑦 , 𝑓𝑦𝑥 exist and are
continuous on an open set (domain).
➢ Theorem 13.6.1: Let 𝑓 𝑥0 , 𝑦0 = 0 and suppose 𝑓(𝑥, 𝑦) is 𝐶 1 in a
neighborhood of (𝑥0 , 𝑦0 ) with 𝑓𝑦 𝑥0 , 𝑦0 ≠ 0.
Then, 𝑓 𝑥, 𝑦 = 0 uniquely defines 𝑦(𝑥) in some neighborhood
𝑁 of 𝑥0 (i.e. an open interval about 𝑥0 ) such that 𝑦 𝑥0 = 𝑦0
and 𝑦′(𝑥) exists in 𝑁.
Note:
✓ This theorem assures us that there exists a unique differentiable
Dr. [Link], Bilkent University, UMRAM

function 𝑦(𝑥) in some neighborhood of 𝑥0 but it does not tell us


how large that neighborhood is.

32
Ex. 2, p.643

Example (I)
➢ 𝑓 𝑥, 𝑦 = 𝑥 2 + 4𝑦 2 − 4 = 0, and let 𝑥0 , 𝑦0 = 1, − 3/2
so that 𝑓 𝑥0 , 𝑦0 = 0.
We have 𝑓𝑥 𝑥, 𝑦 = 2𝑥, 𝑓𝑦 𝑥, 𝑦 = 8𝑦, which are continuous every
where, so that 𝑓 ∈ 𝐶 1 in a neighborhood of (1, − 3/2).
Furthermore,
𝑓𝑦 1, − 3/2 = −4 3 ≠ 0,
so that by the theorem, a unique differentiable function 𝑦(𝑥) exists
in a neighborhood of 𝑥0 = 1.
In fact, 𝑦(𝑥) = − 1 − 𝑥 2 /4 and is
the lower half of the ellipse over
Dr. [Link], Bilkent University, UMRAM

−2 < 𝑥 < 2 as shown in the picture.

33
Ex. 3, p.643

Example (II)
➢ 𝑓 𝑥, 𝑦 = 𝑦 − 2𝑥 𝑒 𝑦 − 𝑥 2 + 1 = 0, 𝑥0 , 𝑦0 = (1,2).
The given point is such that 𝑓 1,2 = 0.
The derivatives
𝑓𝑥 = −2𝑒 𝑦 − 2𝑥 and 𝑓𝑦 = 𝑦 − 2𝑥 + 1 𝑒 𝑦
are continuous everywhere in the 𝑥𝑦-plane.
Thus 𝑓 ∈ 𝐶 1 everywhere.
Also, 𝑓𝑦 1,2 = 𝑒 2 ≠ 0 so we know that 𝑦(𝑥) exists in a
neighborhood of 𝑥0 = 1.
However, an explicit expression for 𝑦(𝑥) is not clear. The previous
theorem only tells the existence of the function.
Dr. [Link], Bilkent University, UMRAM

34
How to Calculate the Function (I)
Let 𝑥0 , 𝑦0 be such that 𝑓 𝑥0 , 𝑦0 = 0. Let us see if we can calculate
the Taylor series of 𝑦(𝑥) about 𝑥0 :
1
𝑦 𝑥 = 𝑦 𝑥0 + 𝑦 ′ 𝑥0 𝑥 − 𝑥0 + 𝑦 ′′ 𝑥0 𝑥 − 𝑥0 2
+⋯ (A)
2
provided all derivatives exist about 𝑥0 .
We also have
𝑑
𝑓 𝑥, 𝑦 𝑥 = 𝑓𝑥 𝑥, 𝑦 + 𝑓𝑦 𝑥, 𝑦 𝑦 ′ = 0,
𝑑𝑥
which gives,

𝑓𝑥 𝑥, 𝑦
𝑦 =− .
𝑓𝑦 𝑥, 𝑦
Now, we can differentiate 𝑦′(𝑥) one more time to find 𝑦′′ as:
Dr. [Link], Bilkent University, UMRAM

′ ′
′′
𝑓𝑥𝑥 + 𝑓𝑥𝑦 𝑦 𝑓𝑦 − 𝑓𝑦𝑥 + 𝑓𝑦𝑦 𝑦 𝑓𝑥
𝑦 =− .
𝑓𝑦2
𝑑 𝑑 𝜕 𝑑𝑥 𝜕 𝑑𝑦
Hint: 𝑓 𝑥, 𝑦 𝑥 = 𝑓 = 𝑓 ⋅ + 𝑓 ⋅ = 𝑓𝑦𝑥 + 𝑓𝑦𝑦 𝑦 ′ .
35 𝑑𝑥 𝑦 𝑑𝑥 𝑦 𝜕𝑥 𝑦 𝑑𝑥 𝜕𝑦 𝑦 𝑑𝑥
How to Calculate the Function (II)
′ ′

𝑓𝑥 𝑥, 𝑦 ′′
𝑓𝑥𝑥 + 𝑓𝑥𝑦 𝑦 𝑓𝑦 − 𝑓𝑦𝑥 + 𝑓𝑦𝑦 𝑦 𝑓𝑥
𝑦 =− ,𝑦 = −
𝑓𝑦 𝑥, 𝑦 𝑓𝑦2
Substituting 𝑦′ from above, into 𝑦′′ we have
2 2
′′
2𝑓𝑥 𝑓𝑦 𝑓𝑥𝑦 − 𝑓𝑥 𝑓𝑦𝑦 − 𝑓𝑦 𝑓𝑥𝑥
𝑦 = 3 .
𝑓𝑦
Continuing this way, the terms of the Taylor series (A)
1 ′′
𝑦 𝑥 = 𝑦 𝑥0 + 𝑦 𝑥0 𝑥 − 𝑥0 + 𝑦 𝑥0 𝑥 − 𝑥0 2 + ⋯

2
can be computed from the partial derivatives of 𝑓. It remains merely
to evaluate the expressions for 𝑦 ′ , 𝑦 ′′ , ⋯ at (𝑥0 , 𝑦0 ).
Dr. [Link], Bilkent University, UMRAM

The condition given in Theorem 13.6.1, i.e., 𝑓𝑦 𝑥0 , 𝑦0 ≠ 0 is seen to


be crucial to make sure that above Taylor series exists.

36
Ex. 4, p.644

Example (II)
➢ 𝑓 𝑥, 𝑦 = 𝑦 − 2𝑥 𝑒 𝑦 − 𝑥 2 + 1 = 0, 𝑥0 , 𝑦0 = (1,2).
We can calculate 𝑦′ and 𝑦′′ as:
𝑦 −𝑦
𝑓𝑥 −2𝑒 − 2𝑥 2 + 2𝑥𝑒
𝑦′ = − = − 𝑦 𝑦
=
𝑓𝑦 𝑒 + 𝑦 − 2𝑥 𝑒 1 + 𝑦 − 2𝑥
2𝑒 −𝑦 − 2𝑥𝑒 −𝑦 ⋅ 𝑦 ′ 1 + 𝑦 − 2𝑥 − −2 + 𝑦′ 2 + 2𝑥𝑒 −𝑦
𝑦 ′′ =
1 + 𝑦 − 2𝑥 2
Evaluating these 𝑦′ and 𝑦′′ expressions at 𝑥0 , 𝑦0 = (1,2) gives
𝑦 ′ 1 = 2(1 + 𝑒 −2 ) and 𝑦 ′′ 1 = −6𝑒 −2 − 8𝑒 −4
Thus,
−2
1
𝑦 𝑥 =2+2 1+𝑒 𝑥−1 + −6𝑒 −2 − 8𝑒 −4 𝑥 − 1 2 + ⋯
Dr. [Link], Bilkent University, UMRAM

2!
is the desired Taylor series for 𝑦(𝑥).

37
Ex. 5, p.645

Example: Finding an Inverse Function


➢ Consider 𝑦 = sin(𝑥) and let us write
𝑓 𝑥,
ො 𝑦ො = 𝑥ො − sin 𝑦ො = 0,
and note that 𝑦(
ො 𝑥)
ො would be the inverse function of sin(𝑥).ො
Let us choose 𝑥ො = 0 and consider 𝑥ො0 , 𝑦ො0 = (0,0) that satisfies
𝑓 𝑥ො0 , 𝑦ො0 = 0.
As Implicit Fncn Thm requires 𝑓 ∈ 𝐶 1 since 𝑓𝑥ො = 1, 𝑓𝑦ො = −cos(𝑦).ො
Further, 𝑓𝑦ො 0,0 = − cos 0 = −1 ≠ 0.
By the theorem, 𝑦(
ො 𝑥)
ො exists for 𝑥ො in a neighborhood of 𝑥ො0 = 0.
We actually know that it exists in −1,1 ∈ ℝ:
𝑦( ො = sin−1 (𝑥)
ො 𝑥) ො
 𝑦(𝑥) = sin(𝑥)
Dr. [Link], Bilkent University, UMRAM

1 𝜋/2

−𝜋/2 −1
𝑥 𝑥ො
𝜋/2 1

38 −1 −𝜋/2
Multivariable Case: Jacobian (I)
Consider a system of 𝑛-equations
𝑓1 𝑥1 , ⋯ , 𝑥𝑛 , 𝑢1 , ⋯ , 𝑢𝑛 = 0
ቐ ⋮ (31)
𝑓𝑛 𝑥1 , ⋯ , 𝑥𝑛 , 𝑢1 , ⋯ , 𝑢𝑛 = 0
Does there exist 𝑢1 𝑥1 , ⋯ , 𝑥𝑛 , ⋯ , 𝑢𝑛 (𝑥1 , ⋯ , 𝑥𝑛 ), functions of 𝑛-
variables, satisfying these equations ?
Let 𝑃ത = (𝑥10 , ⋯ 𝑥𝑛0 , 𝑢10 , ⋯ , 𝑢𝑛0 ) be a point in 2𝑛-space such that
(31) holds at point 𝑃. ത
Fact: If 𝑓1 , ⋯ , 𝑓𝑛 ∈ 𝐶 1 in some neighborhood of 𝑃ത and if
𝜕𝑓1 𝜕𝑓1
… then there exists 𝐶 1
functions
𝜕𝑢1 𝜕𝑢𝑛
Dr. [Link], Bilkent University, UMRAM

det
⋮ ⋱ ⋮ ≠ 0 𝑢1 𝑥1 , ⋯ , 𝑥𝑛 , ⋯ , 𝑢𝑛 (𝑥1 , ⋯ , 𝑥𝑛 ) that
𝜕𝑓𝑛 𝜕𝑓𝑛 satisfy (31) in some neighborhood of
… (𝑥10 , ⋯ 𝑥𝑛0 ) in 𝑛-space.
𝜕𝑢1 𝜕𝑢𝑛 𝑃ത
39
Ex. 6, p.648

Example
➢ Consider the familiar change of variables from Cartesian
𝑦 𝑃
𝑥, 𝑦 coordinates to polar coordinates 𝑟, 𝜃 :
𝑟
𝑥 = 𝑟 cos(𝜃) 𝑓 𝑥, 𝑦, 𝑟, 𝜃 = 𝑥 − 𝑟 cos 𝜃 = 0
ቊ ⇒ ቊ1 𝜃
𝑦 = 𝑟 sin(𝜃) 𝑓2 𝑥, 𝑦, 𝑟, 𝜃 = 𝑦 − 𝑟 sin 𝜃 = 0 𝑥

Question: Do these relations define implicit functions 𝑟(𝑥, 𝑦) and


𝜃(𝑥, 𝑦) throughout some neighborhood of a given point 𝑃?
Answer: All partial derivatives of 𝑓1 and 𝑓2 exist and are continuous
in ℝ4 = 𝑥, 𝑦, 𝑟, 𝜃 . Further,
𝜕𝑓1 𝜕𝑓1
− cos 𝜃 𝑟 sin 𝜃
det 𝜕𝑟 𝜕𝜃 = det =𝑟 .
𝜕𝑓2 𝜕𝑓2 −sin 𝜃 −𝑟 cos 𝜃
Dr. [Link], Bilkent University, UMRAM

𝜕𝑟 𝜕𝜃
This is nonzero everywhere in ℝ4 except when 𝑟 = 0. It follows by
the theorem that, the functions 𝑟(𝑥, 𝑦) and 𝜃(𝑥, 𝑦) exist in a
neighborhood of any point (𝑥, 𝑦) that is nonzero. 40
Jacobian (I)
We denote the Jacobian determinant by

𝜕𝑓1 𝜕𝑓1

𝜕𝑢1 𝜕𝑢𝑛
𝜕 𝑓1 ,⋯,𝑓𝑛
𝐽 𝑢1 , ⋯ , 𝑢𝑛 = ≔ det ⋮ ⋱ ⋮ (32)
𝜕 𝑢1 ,⋯,𝑢𝑛
𝜕𝑓𝑛 𝜕𝑓𝑛

𝜕𝑢1 𝜕𝑢𝑛
It is called the Jacobian of 𝑓 ҧ with respect to 𝑢ത . The Jacobian comes
𝜕𝑢𝑗
up in calculating for 𝑖, 𝑗 = 1, ⋯ , 𝑛.
𝜕𝑥𝑖
➢ Case 𝑛 = 2: Let us assume that functions 𝑢, 𝑣 exist and they are
functions of 𝑥, 𝑦 defined implicitly by:
Dr. [Link], Bilkent University, UMRAM

𝑓 𝑥, 𝑦, 𝑢, 𝑣 = 0 𝐹 𝑥, 𝑦 ∶= 𝑓 𝑥, 𝑦, 𝑢(𝑥, 𝑦), 𝑣 𝑥, 𝑦 = 0,
ቋ⇒ ቐ
𝑔 𝑥, 𝑦, 𝑢, 𝑣 = 0 𝐺 𝑥, 𝑦 ∶= 𝑔 𝑥, 𝑦, 𝑢(𝑥, 𝑦), 𝑣 𝑥, 𝑦 = 0.
41
Jacobian (II)
Computing 𝑢 𝑥, 𝑦 , 𝑣(𝑥, 𝑦) requires some terms in the Taylor series
𝑢 𝑥, 𝑦 = 𝑢 𝑥0 , 𝑦0 + 𝑢𝑥 𝑥0 , 𝑦0 𝑥 − 𝑥0 + 𝑢𝑦 𝑥0 , 𝑦0 𝑦 − 𝑦0 + ⋯
𝑣 𝑥, 𝑦 = 𝑣 𝑥0 , 𝑦0 + 𝑣𝑥 𝑥0 , 𝑦0 𝑥 − 𝑥0 + 𝑣𝑦 𝑥0 , 𝑦0 𝑦 − 𝑦0 + ⋯
Now,
𝜕𝐹 𝜕𝐹
= 𝑓𝑥 + 𝑓𝑢 𝑢𝑥 + 𝑓𝑣 𝑣𝑥 = 0 , = 𝑓𝑦 + 𝑓𝑢 𝑢𝑦 + 𝑓𝑣 𝑣𝑦 = 0
𝜕𝑥 𝜕𝑦
𝜕𝐺 𝜕𝐺
= 𝑔𝑥 + 𝑔𝑢 𝑢𝑥 + 𝑔𝑣 𝑣𝑥 = 0 , = 𝑔𝑦 + 𝑔𝑢 𝑢𝑦 + 𝑔𝑣 𝑣𝑦 = 0
𝜕𝑥 𝜕𝑦
so that
𝑓𝑢 𝑓𝑣 𝑢𝑥 𝑓𝑥 𝑓𝑢 𝑓𝑣 𝑢𝑦 𝑓𝑦
= − , = −
𝑔𝑢 𝑔𝑣 𝑣𝑥 𝑔𝑥 𝑔𝑢 𝑔𝑣 𝑣𝑦 𝑔𝑦
Dr. [Link], Bilkent University, UMRAM

Rewriting, we have

42
Jacobian (III)
𝑓 𝑓 𝑓𝑣 𝑢𝑥 𝑓 𝑓 𝑓𝑣 𝑢𝑦
− 𝑥 = 𝑢 , − 𝑔𝑦 = 𝑢
𝑔𝑥 𝑔𝑢 𝑔𝑣 𝑣𝑥 𝑦 𝑔𝑢 𝑔𝑣 𝑣𝑦
Using Jacobian notation,
𝑓𝑥 𝑓𝑣 𝜕 𝑓, 𝑔 𝑓 𝑓𝑥 𝜕 𝑓, 𝑔
det det 𝑢
𝑔𝑥 𝑔𝑣 𝜕 𝑥, 𝑣 𝑔𝑢 𝑔𝑥 𝜕 𝑢, 𝑥
𝑢𝑥 = − =− , 𝑣𝑥 = − =−
𝑓 𝑓𝑣 𝜕 𝑓, 𝑔 𝑓 𝑓𝑣 𝜕 𝑓, 𝑔
det 𝑢 det 𝑢
𝑔𝑢 𝑔𝑣 𝜕 𝑢, 𝑣 𝑔𝑢 𝑔𝑣 𝜕 𝑢, 𝑣
Similarly,
𝑓 𝑓 𝜕 𝑓, 𝑔 𝑓 𝑓 𝜕 𝑓, 𝑔
det 𝑔𝑦 𝑔𝑣 det 𝑔𝑢 𝑔𝑦
𝑦 𝑣 𝜕 𝑦, 𝑣 𝑢 𝑦 𝜕 𝑢, 𝑦
𝑢𝑦 = − =− , 𝑣𝑦 = − = −
𝑓𝑢 𝑓𝑣 𝜕 𝑓, 𝑔 𝑓 𝑓𝑣 𝜕 𝑓, 𝑔
det det 𝑢
𝑔𝑢 𝑔𝑣 𝜕 𝑢, 𝑣 𝑔𝑢 𝑔𝑣 𝜕 𝑢, 𝑣
Dr. [Link], Bilkent University, UMRAM

Since 𝐽(𝑢, 𝑣) occurs in the denominator of every partial derivative of 𝑢, 𝑣, it


must be nonzero for the existence of every partial derivative 𝑢𝑥 , 𝑣𝑥 , 𝑢𝑦, 𝑣𝑦.
✓ You may refer to section 13.6.4 for an application to change of variables.
43
Jacobian (IV) Example
Ex. 1. Consider 𝑥 − 𝑢 𝑐𝑜𝑠𝑣 = 0, 𝑦 − 𝑢 𝑠𝑖𝑛𝑣 = 0. Do we have 𝑢 𝑥, 𝑦 , 𝑣 𝑥, 𝑦 defined about
𝜋
(𝑥, 𝑦) = (0,2)? The point 𝑃 = 𝑥, 𝑦, 𝑢, 𝑣 = 0,2,2, satisfies the equations. Denoting the
2
left-hand sides by 𝑓, 𝑔, we have
𝑓 𝑓𝑣 −𝑐𝑜𝑠𝑣 𝑢 𝑠𝑖𝑛𝑣
det 𝑢 = det = 𝑢|𝑃 = 2 ≠ 0
𝑔𝑢 𝑔𝑣 −𝑠𝑖𝑛𝑣 −𝑢 𝑐𝑜𝑠𝑣
so that 𝑢(𝑥, 𝑦), 𝑣(𝑥, 𝑦) exist about the point (0,2) by the previous Fact. Also,
𝑓 𝑓𝑣 1 𝑢 𝑠𝑖𝑛𝑣
det 𝑥 det
𝑔𝑥 𝑔𝑣 0 −𝑢 𝑐𝑜𝑠𝑣 = 𝑐𝑜𝑠𝑣 = 0,
𝑢𝑥 = − =−
𝑢 𝑢
𝑓𝑢 𝑓𝑥 −𝑐𝑜𝑠𝑣 1
det det
𝑣𝑥 = −
𝑔𝑢 𝑔𝑥
=− −𝑠𝑖𝑛𝑣 0 = −𝑠𝑖𝑛𝑣 = − 1
𝑢 𝑢 𝑢 2
Similarly,
𝑓 𝑓 0 𝑢 𝑠𝑖𝑛𝑣
det 𝑔𝑦 𝑔𝑣 det
𝑢𝑦 = −
𝑦 𝑣
=− 1 −𝑢 𝑐𝑜𝑠𝑣 = 𝑠𝑖𝑛𝑣 = 1,
𝑢 𝑢
Dr. [Link], Bilkent University, UMRAM

𝑓 𝑓 −𝑐𝑜𝑠𝑣 0
det 𝑔𝑢 𝑔𝑦 det
𝑣𝑦 = −
𝑢 𝑦
=− −𝑠𝑖𝑛𝑣 1 = 𝑐𝑜𝑠𝑣 = 0
𝑢 𝑢 𝑢
1
so that 𝑢 𝑥, 𝑦 ≅ 2 + 0 ∙ 𝑥 + 1 ∙ 𝑦 − 2 = 𝑦, 𝑣 𝑥, 𝑦 ≅ 𝜋Τ2 − 2 ∙ 𝑥 + 0 ∙ 𝑦 − 2 = 𝜋Τ2 − 𝑥Τ2

44
Jacobian (V) Example
Ex. 2. Let us consider
𝑓 = 𝑢𝑤3 − 𝑦, 𝑔 = 2𝑣 − 𝑤 + 𝑥, h = 𝑒𝑢𝑣 − 𝑧.
Then ,
𝜕(𝑓, 𝑔, ℎ) 𝑓𝑢 𝑓𝑣 𝑓𝑤 𝑤3 0 3𝑢𝑤2
= 𝑑𝑒𝑡 𝑔𝑢 𝑔𝑣 𝑔𝑤 = 𝑑𝑒𝑡 0 2 −1
𝜕(𝑢, 𝑣, 𝑤) ℎ𝑢 ℎ𝑣 ℎ𝑤 𝑣𝑒𝑢𝑣 𝑢𝑒𝑢𝑣 0
= 𝑤3∙ 𝑢𝑒𝑢𝑣 − 𝑣𝑒𝑢𝑣 ∙ 6𝑢𝑤2
= 𝑢𝑤2 𝑒𝑢𝑣 (𝑤 − 6𝑣)

𝜕(𝑓, 𝑔, ℎ) 𝑓𝑥 𝑓𝑣 𝑓𝑤 0 0 3𝑢𝑤2
= 𝑑𝑒𝑡 𝑔𝑥 𝑔𝑣 𝑔𝑤 = 𝑑𝑒𝑡 1 2 −1
Dr. [Link], Bilkent University, UMRAM

𝜕(𝑥, 𝑣, 𝑤) ℎ𝑥 ℎ𝑣 ℎ𝑤 0 𝑢𝑒𝑢𝑣 0
= 3𝑢2𝑤2𝑒𝑢𝑣

45
Chapter 13: Differential Calculus
of
Functions of Several Variables

Advanced Engineering Mathematics, 2005, 2nd Ed., Prentice Hall


[M. D. Greenberg]

Dr. [Link], Bilkent University, UMRAM


Definitions
A function 𝑓(𝑥) has an absolute maximum at 𝑥0 if 𝑓(𝑥) ≤ 𝑓(𝑥0 ) for
all 𝑥 in its domain of definition.
It has an absolute minimum at 𝑥0 if 𝑓 𝑥 ≥ 𝑓(𝑥0 ) for all 𝑥 in its
domain of definition.
It has a local maximum at 𝑥0 if 𝑓 𝑥 ≤ 𝑓(𝑥0 ) and/or local minimum
at 𝑥0 if 𝑓 𝑥 ≥ 𝑓(𝑥0 ) for all 𝑥 in some neighborhood of 𝑥0 .
We use the term extremum to mean either a maximum or a
minimum. In the figure bellow, 𝐴, 𝐵, 𝐶, 𝐷, 𝐹 are extremum points but
not 𝐸, which is called an inflection point.
𝑓 𝐷
𝐷: absolute maximum
Dr. [Link], Bilkent University, UMRAM

𝐹: absolute minimum 𝐵 𝐸
𝐴, 𝐶: local minimum
𝐵: local maximum 𝐴 𝐶 𝐹
𝐸: horizontal inflection point
47 𝑥
Vanishing Derivative for Local Extremum
➢ Theorem 13.7.1: If 𝑓(𝑥) has a local extremum at a point 𝑥0 , and
if 𝑓′(𝑥0 ) exists, then
𝑓 ′ 𝑥0 = 0.
Proof: Suppose 𝑥0 is a local minimum point so that 𝑓 𝑥0 ≤ 𝑓(𝑥) in
a neighborhood 𝑁(𝑥0 ). Since,

𝑓 𝑥 − 𝑓 𝑥0
𝑓 𝑥0 = lim ,
𝑥→𝑥0 𝑥 − 𝑥0
the left limit is negative and the right limit is positive as 𝑥 → 𝑥0 . This
implies that 𝑓 ′ 𝑥0 = 0 .
A similar argument applies if 𝑓 has a local maximum.
It follows that the condition 𝑓 ′ 𝑥0 = 0 is a necessary condition.
Dr. [Link], Bilkent University, UMRAM

It is not sufficient since the point 𝐸 in the previous figure also satisfies
this condition (zero slope), whereas 𝐸 is neither a local maximum nor a
local minimum, rather it is called a horizontal inflection point.
A sufficient condition is given in the next theorem.
48
Maximum, Minimum, Inflection Point (I)
➢ Theorem 13.7.2: Suppose that for some 𝑥0 in the domain of
definition of 𝑓(𝑥) and for 𝑛 ≥ 2, we have that
𝑓 ′ 𝑥0 = 0, 𝑓 ′′ 𝑥0 = 0, ⋯ , 𝑓 𝑛−1 𝑥0 = 0 but 𝑓 𝑛 𝑥0 ≠ 0,
𝑛
and that 𝑓 (𝑥) is continuous in some neighborhood of 𝑥0 .
Then, if:
▪ 𝑛 is even and 𝑓 𝑛 𝑥0 < 0 ⇒ 𝑥0 is a local maximum point,
▪ 𝑛 is even and 𝑓 𝑛 𝑥0 > 0 ⇒ 𝑥0 is a local minimum point,
▪ 𝑛 is odd ⇒ 𝑥0 is a horizontal inflection point.
Proof: Since by assumption 𝑓 𝑛 (𝑥) is continuous in some
neighborhood of 𝑥0 , there must exist a neighborhood 𝑁(𝑥0 )
Dr. [Link], Bilkent University, UMRAM

throughout which 𝑓 𝑛 (𝑥) has the same sign of 𝑓 𝑛 (𝑥0 ).


By Taylor’s formula with Lagrange remainder, i.e.:
𝑛
𝑓 𝜉
𝑓 𝑥 = 𝑓 𝑥0 + 𝑓′ 𝑥0 𝑥 − 𝑥0 + ⋯ + 𝑥 − 𝑥0 𝑛
49
𝑛!
Maximum, Minimum, Inflection Point (II)
𝑓𝑛 𝜉 𝑛
Then, 𝑓 𝑥 = 𝑓 𝑥0 + 0 + ⋯ + 0 + 𝑥 − 𝑥0
𝑛!
𝑓𝑛 𝜉
or , 𝑓 𝑥 =𝑓 𝑥0 + 𝑥 − 𝑥0 𝑛 , 𝑥 ∈ 𝑁 𝑥0 .
𝑛!
Since 𝑥 is in 𝑁(𝑥0 ), 𝜉 must be in 𝑁(𝑥0 ) as well, as 𝜉 is between 𝑥, 𝑥0 .
Thus, knowing that 𝑓 𝑛 𝜉 has the same sign as 𝑓 𝑛 (𝑥0 ) and that
𝑓𝑛 𝜉
𝑓 𝑥 − 𝑓 𝑥0 = 𝑥 − 𝑥0 𝑛 ,
𝑛!
It follows that
▪ If 𝑛 is even, 𝑓 𝑛 𝑥0 > 0, then 𝑓 𝑥 ≥ 𝑓 𝑥0 in 𝑁(𝑥0 ), so 𝑓 has a local
minimum at 𝑥0 ,
▪ If 𝑛 is even, 𝑓 𝑛 𝑥0 < 0, then 𝑓 𝑥 ≤ 𝑓 𝑥0 in 𝑁(𝑥0 ), so 𝑓 has a local
Dr. [Link], Bilkent University, UMRAM

maximum at 𝑥0 ,
▪ If 𝑛 is odd, then depending on whether 𝑥 < 𝑥0 or 𝑥 > 𝑥0 , we will have
opposite signs for 𝑓 𝑥 − 𝑓(𝑥0 ), which gives that 𝑥0 is an inflection
50 point.
Ex. 1,2, p.658

Examples
2 4
➢ Let 𝑓 𝑥 = 2 𝑥 − 𝑥 − 1 = 𝑥−1 over 0 < 𝑥 < ∞.
𝑓′ 𝑥 = 2 𝑥 − 1 3 / 𝑥.
Note that 𝑓 ′ 1 = 0; similarly, 𝑓 ′′ 1 = 𝑓 ′′′ 1 = 0 but
𝑓 4 1 = 3/2 ≠ 0.
Since 𝑛 is even and 𝑓 4 1 > 0, we have a local
minimum at 𝑥0 = 1. 

2
➢ 𝑓 𝑥 = 𝑥−1 ln 𝑥 ; 0 < 𝑥 < ∞.
𝑓 ′ 𝑥 = 2 𝑥 − 1 ln 𝑥 + 𝑥 − 1 2 /𝑥
= 𝑥 − 1 2 ln 𝑥 + 𝑥 − 1 /𝑥 ,
𝑓 ′′ 𝑥 = 2 ln 𝑥 + 4 𝑥 − 1 /𝑥 − 𝑥 − 1 2 /𝑥 2 ,
6
Dr. [Link], Bilkent University, UMRAM

𝑓 ′′′ 𝑥 = − 6(𝑥 − 1)/𝑥 2 + 2 𝑥 − 1 2 /𝑥 3 ,


𝑥
so that 𝑓 ′ 1 = 𝑓 ′′ 1 = 0 but 𝑓 ′′′ 1 = 6 ≠ 0.
This time 𝑛 is odd so 𝑓 has a horizontal inflection
point at 𝑥0 = 1. 51
Multivariable Case
Let us consider 𝑓 𝑥ҧ = 𝑓(𝑥1 , ⋯ , 𝑥𝑛 ) defined for 𝑥ҧ ∈ 𝒟 ⊆ ℝ𝑛 and a
point 𝑋1 , ⋯ , 𝑋𝑛 ∈ 𝒟. If there is a neighborhood 𝑁(𝑋) ത such that
𝑓 𝑥ҧ ≤ 𝑓 𝑋ത ∀𝑥ҧ ∈ 𝑁 𝑋ത ⇒ 𝑋ത is a local maximum point of 𝑓(𝑥), ҧ
𝑓 𝑥ҧ ≥ 𝑓 𝑋ത ∀𝑥ҧ ∈ 𝑁 𝑋ത ⇒ 𝑋ത is a local minimum point of 𝑓(𝑥). ҧ
➢ Theorem 13.7.3: For a function 𝑓 of 𝑛 real variables 𝑥1 , ⋯ , 𝑥𝑛 to
have a local extremum at a point 𝑋ത = (𝑋1 , ⋯ , 𝑋𝑛 ) in its domain
of definition, where 𝑓(𝑥)ҧ is 𝐶 1 in some neighborhood of 𝑋, ത it is
necessary that, at 𝑥ҧ = 𝑋ത , we have
𝜕𝑓 𝜕𝑓 𝜕𝑓
= 0, =0 , ⋯ , 𝜕𝑥𝑛
= 0. (33)
𝜕𝑥1 𝜕𝑥2
(PROOF?)
A point 𝑋ത satisfying (33) is called a critical point of 𝑓.
Dr. [Link], Bilkent University, UMRAM

As emphasized for single-variable case (in theorem 13.7.1), above


condition is necessary but not sufficient for the existence of an extremum

of 𝑓 at 𝑋.
 Not all critical points are local extremum.
52
Ex. -, p.660

Example
➢ Let 𝑓 𝑥ҧ = 𝑓 𝑥1 , 𝑥2 = 𝑥12 − 𝑥22 .
𝜕𝑓 𝜕𝑓
= 2𝑥1 , = −2𝑥2 , which are both zero at 𝑋ത = (0,0).
𝜕𝑥1 𝜕𝑥2
However, 𝑓 > 0 for |𝑥1 | > |𝑥2 | and 𝑓 < 0 for |𝑥1 | < |𝑥2 | and in any
neighborhood of (0,0) there are infinitely many such points.
Hence, 𝑋ത = (0,0) is neither a local maximum, nor a local minimum
point.
It is called a saddle point.
Dr. [Link], Bilkent University, UMRAM

53
Classification of Critical Points
➢ Theorem 13.7.4: Consider that
𝜕𝑓 𝜕𝑓

𝑓𝑥1 𝑋 = ത ത
𝑋 = 0, ⋯ , 𝑓𝑥𝑛 𝑋 = 𝑋ത = 0
𝜕𝑥1 𝜕𝑥𝑛
for some 𝑋ത in the domain of 𝑓.
ത and let det 𝐻
Let 𝑓 ∈ 𝐶 2 in some neighborhood 𝑁(𝑋) 𝑋ത ≠ 0 ,
𝑓𝑥1 𝑥1 𝑋ത … 𝑓𝑥1 𝑥𝑛 𝑋ത
where 𝐻≔ ⋮ ⋱ ⋮ (Hessian)
𝑓𝑥𝑛 𝑥1 𝑋ത … 𝑓𝑥𝑛 𝑥𝑛 𝑋ത
Then,
✓ If 𝐻 is positive definite ⇒ 𝑋ത is a local minimum point,
✓ If 𝐻 is negative definite ⇒ 𝑋ത is a local maximum point,
✓ If 𝐻 has at least one positive and one negative eigenvalue
Dr. [Link], Bilkent University, UMRAM

 𝑋ത is a saddle point,
✓ Otherwise,
Note: If det 𝐻 𝑋ത = 0, then the Hessian, by itself, does not give enough
information to classify 𝑋ത point.
54
Some Remarks
Eigenvalue:
A complex 𝜆 ∈ ℂ is an eigenvalue of a square matrix 𝑴, if it is a root
of the characteristic polynomial
𝑃 𝜆 = det(𝜆𝑰 − 𝑴) .
Positive/Negative definite:
If 𝑴 is symmetric, then all its eigenvalues are real numbers.
A symmetric matrix 𝑴 is
▪ positive- definite, if all its eigenvalues are positive.
▪ positive semi-definite if all its eigenvalues are nonnegative.
A symmetric matrix 𝑴 is:
Dr. [Link], Bilkent University, UMRAM

▪ negative-definite if −𝑴 is positive-definite.
▪ negative semi-definite if all eigenvalues of −𝑴 are nonnegative.

55
Ex. -, p.-

Examples
1 −1 𝜆−1 1
➢ 𝑀= , 𝜆𝐼 − 𝑀 =
−1 1 1 𝜆−1
2
det 𝜆𝐼 − 𝑀 = 𝜆 − 1 − 1 = 0 ⇒ 𝜆1 = 0, 𝜆2 = 2
So that 𝑀 is positive semi-definite.
1 −2 𝜆−1 2
➢ 𝑀= , 𝜆𝐼 − 𝑀 =
−2 1 2 𝜆−1
2
det 𝜆𝐼 − 𝑀 = 𝜆 − 1 − 4 = 0 ⇒ 𝜆1 = 3, 𝜆2 = −1
So that 𝑀 is neither positive nor negative (semi) definite.
1 −1/2 𝜆 − 1 1/2
➢ 𝑀= , 𝜆𝐼 − 𝑀 =
−1/2 1 1/2 𝜆 − 1
det 𝜆𝐼 − 𝑀 = 𝜆 − 1 2 − 1/4 = 0 ⇒ 𝜆1 = 1/2, 𝜆2 = 3/2
Dr. [Link], Bilkent University, UMRAM

So that 𝑀 is positive definite.

56
Ex. 3, p.662

Example
➢ Classify all local extrema or saddle points of the function, if any.
𝑓 𝑥, 𝑦 = ln 2𝑥 𝑦 − 1 + 1 .
2 𝑦−1 2𝑥
𝑓𝑥 = , 𝑓𝑦 = .
2𝑥 𝑦 − 1 + 1 2𝑥 𝑦 − 1 + 1
Setting 𝑓𝑥 = 0, 𝑓𝑦 = 0, gives the single critical point 𝑋ത = (0,1).
Now, 𝑓 ∈ 𝐶 2 in a neighborhood about this point 𝑋ത = (0,1), since
the curve 2𝑥 𝑦 − 1 + 1 =0 is safely away from that point.
By Calculating 𝑓𝑥𝑥 , 𝑓𝑦𝑦 , and 𝑓𝑥𝑦 = 𝑓𝑦𝑥 we have Hessian matrix as
4 𝑦−1 2 2 2𝑥 𝑦−1 +1 −4𝑥(𝑦−1)

2𝑥 𝑦−1 +1 2 2𝑥 𝑦−1 +1 2 0 2
𝐻 𝑋ത = =
2 2𝑥 𝑦−1 +1 −4𝑥(𝑦−1) 4𝑥 2 2 0

Dr. [Link], Bilkent University, UMRAM

2𝑥 𝑦−1 +1 2 2𝑥 𝑦−1 +1 2
𝑋ത
Eigenvalues of 𝐻(𝑋)ത are found from
𝜆2 − 4 = 0 as two points 2, −2 , so that the
point 𝑋ത = (0,1) is a saddle point of 𝑓(𝑥, 𝑦).
57
Relation to Multivariable Taylor Series
Multivariable Taylor’s formula can be written as
1
𝑓 𝑥ҧ = 𝑓 𝑋ത + 𝐽 𝑋ത 𝑥ҧ − 𝑋ത + 𝑥ҧ − 𝑋ത 𝑇 𝐻 𝑋ത 𝑥ҧ − 𝑋ത + ⋯
2
where 𝐽 is the Jacobian row vector [𝑓𝑥1 , ⋯ , 𝑓𝑥𝑛 ] and 𝐻 is the Hessian
matrix.
You may see how this expression can be used in a proof of the
classification Theorem 13.7.4.
If the necessary condition of Theorem 13.7.3 is satisfied, then
𝐽 𝑋ത =0 so that
1
𝑓 𝑥ҧ − 𝑓 𝑋 = 𝑥ҧ − 𝑋ത 𝑇 𝐻 𝑋ത 𝑥ҧ − 𝑋ത + ⋯

2
Dr. [Link], Bilkent University, UMRAM

The sign of 𝑓 𝑥ҧ − 𝑓 𝑋ത is determined by the dominant term on


the right, the sign of which, in turn, is determined by the matrix
𝐻 𝑋ത .
58
Constrained Extrema (I)
Consider the problem of finding the extrema for 𝑓 as
𝑓 𝑥1 , ⋯ , 𝑥𝑛 = minimum/maximum
subject to some constraints
𝑔1 𝑥1 , ⋯ , 𝑥𝑛 = 𝑐1 ,

𝑔𝑘 𝑥1 , ⋯ , 𝑥𝑛 = 𝑐𝑘 .
➢ Example:
Suppose we wish to find the point(s) on the plane 2𝑥 − 𝑦 + 4𝑧 = 3
that is closest to the origin.
Since the distance from the origin to (𝑥, 𝑦, 𝑧) is 𝑥 2 + 𝑦 2 + 𝑧 2 , we
Dr. [Link], Bilkent University, UMRAM

wish to minimize the function 𝑥 2 + 𝑦 2 + 𝑧 2 subject to the


constraint that 𝑥, 𝑦, 𝑧 are related according to 2𝑥 − 𝑦 + 4𝑧 = 3.
Note that if we minimize the distance we will also be minimizing the
square
59
of the distance, 𝑥 2 + 𝑦 2 + 𝑧 2 .
Constrained Extrema (II)
In this example, 𝑛 = 3 and 𝑘 = 1, and we can state the problem as
𝑓(𝑥, 𝑦, 𝑧) = 𝑥 2 + 𝑦 2 + 𝑧 2 = minimum,
subject to
𝑔 𝑥, 𝑦, 𝑧 = 2𝑥 − 𝑦 + 𝑧 = 3.
Note that we use simpler notation 𝑥, 𝑦, 𝑧 in place of 𝑥1 , 𝑥2 , 𝑥3 .
We also note that 𝑓 and 𝑔 are 𝐶 1 everywhere, and hence, in some
neighborhood of the point under search, at which 𝑓 has an
extremum.
We discuss two methods:
➢ Method of Elimination,
Dr. [Link], Bilkent University, UMRAM

➢ Method of Lagrange Multiplier.


Let us focus on the general problem, but with only one constraint.

60
Method of Elimination (I)
Let us use the chain rule and write
𝜕𝑓 𝜕𝑓 𝜕𝑓
𝑑𝑓 = ത
𝑃 𝑑𝑥 + ത
𝑃 𝑑𝑦 + 𝑃ത 𝑑𝑧 = 0
𝜕𝑥 𝜕𝑦 𝜕𝑧
at the extremum point 𝑃ത = (𝑥0 , 𝑦0 , 𝑧0 ).
Since 𝑥, 𝑦, 𝑧 are related to each other through the given constraint
equation 𝑔 𝑥, 𝑦, 𝑧 = 𝑐, the 𝑑𝑥, 𝑑𝑦, 𝑑𝑧 increments are not
independent and we can not infer that at the extremum point 𝑃: ത
𝜕𝑓 𝜕𝑓 𝜕𝑓
= 0, = 0, = 0.
𝜕𝑥 𝜕𝑦 𝜕𝑧
One way out, is to solve 𝑔 𝑥, 𝑦, 𝑧 = 𝑐 for 𝑧 as a function of 𝑥, 𝑦,
Dr. [Link], Bilkent University, UMRAM

using implicit function theorem, to obtain 𝑧 = 𝑍(𝑥, 𝑦), provided


ത ≠ 0 in a neighborhood of 𝑃,
𝑔𝑧 (𝑃) ത then
𝑓 𝑥, 𝑦, 𝑍 𝑥, 𝑦 =: 𝐹 𝑥, 𝑦 = extremum.
61
Method of Elimination (II)
𝑓 𝑥, 𝑦, 𝑍 𝑥, 𝑦 =: 𝐹 𝑥, 𝑦 = extremum.
Now, 𝐹 is a well-defined composite function in a neighborhood of
(𝑥0 , 𝑦0 ) with 𝑥 and 𝑦 as independent variables.
Thus,
𝜕𝐹 𝜕𝐹
𝑑𝐹 = 𝑑𝑥 + 𝑑𝑦 = 0
𝜕𝑥 𝜕𝑦
does imply that
𝜕𝐹 𝜕𝐹
= 0, =0
𝜕𝑥 𝜕𝑦
because 𝑑𝑥 and 𝑑𝑦 are arbitrary independent increments.
Dr. [Link], Bilkent University, UMRAM

Then, we can solve for 𝑥 and 𝑦, obtaining the critical point of


𝐹 𝑥, 𝑦 , and 𝑧 = 𝑍(𝑥, 𝑦) gives 𝑧.

62
Ex. 7, p.666

Example: Method of Elimination (I)


We use this method to solve the minimum distance problem:
minimize 𝑓 𝑥, 𝑦, 𝑧 = 𝑥 2 + 𝑦 2 + 𝑧 2
subject to
𝑔 𝑥, 𝑦, 𝑧 = 2𝑥 − 𝑦 + 𝑧 = 3.
Solving 𝑔 for, lets say 𝑦, gives 𝑦 = 𝑧 + 2𝑥 − 3, so
𝑓 𝑥, 𝑦, 𝑧 = 𝑥 2 + 𝑧 + 2𝑥 − 3 2 + 𝑧 2 ≔ 𝐹 𝑥, 𝑧 .
𝜕𝐹
= 2𝑥 + 4 𝑧 + 2𝑥 − 3 = 0 10 4 𝑥 12
𝜕𝑥
Then, ቑ⇒ =
𝜕𝐹
= 2 𝑧 + 2𝑥 − 3 + 2𝑧 = 0 4 4 𝑧 6
𝜕𝑧
which gives (𝑥0 , 𝑧0 ) = (1,1/2) and 𝑦0 = 𝑧0 + 2𝑥0 − 3 = −1/2.
Hence, 𝑃ത = (1, −1/2,1/2) is a critical point, which by
Dr. [Link], Bilkent University, UMRAM

𝐹 𝐹 10 4
𝐻 = [ 𝑥𝑥 𝑥𝑧 ] = , ⇒ 𝜆1 = 12, 𝜆2 = 2
𝐹𝑧𝑥 𝐹𝑧𝑧 4 4
is a minimum point since both eigenvalues are positive.
63
Method of Lagrange Multiplier (I)
Minimize 𝑓(𝑥, 𝑦, 𝑧) subject to 𝑔 𝑥, 𝑦, 𝑧 = 𝑐.
Note that 𝑑𝑓 vanishes at an extremum point, so does 𝑑𝑔, thus:
𝑑𝑓 = 𝑓𝑥 𝑑𝑥 + 𝑓𝑦 𝑑𝑦 + 𝑓𝑧 𝑑𝑧 = 0,
𝑑𝑔 = 𝑔𝑥 𝑑𝑥 + 𝑔𝑦 𝑑𝑦 + 𝑔𝑧 𝑑𝑧 = 0.
from which it follows that
𝑑𝑓 − 𝜆𝑑𝑔 = 𝑓𝑥 − 𝜆𝑔𝑥 𝑑𝑥 + 𝑓𝑦 − 𝜆𝑔𝑦 𝑑𝑦 + 𝑓𝑧 − 𝜆𝑔𝑧 𝑑𝑧 = 0,
for all 𝜆.
Since 𝑑𝑥, 𝑑𝑦, and 𝑑𝑧 are not independent increments, we can not
conclude that
𝑓𝑥 − 𝜆𝑔𝑥 = 0, 𝑓𝑦 − 𝜆𝑔𝑦 = 0, 𝑓𝑧 − 𝜆𝑔𝑧 = 0.
Dr. [Link], Bilkent University, UMRAM

However, note that if, e.g., 𝑔𝑧 ≠ 0 at the critical point 𝑃ത = (𝑥0, 𝑦0, 𝑧0),
𝑓𝑧 𝑃ത
then by choosing 𝜆 = ,
𝑔𝑧 𝑃ത
64
Method of Lagrange Multiplier (II)
𝜆 = 𝑓𝑧 𝑃ത /𝑔𝑧 𝑃ത ,
ത i.e.,
it remains only two terms to be vanished at 𝑃,
𝑑𝑓 − 𝜆𝑑𝑔 = 𝑓𝑥 − 𝜆𝑔𝑥 𝑑𝑥 + 𝑓𝑦 − 𝜆𝑔𝑦 𝑑𝑦 = 0,
where 𝑑𝑥, 𝑑𝑦 are now independent increments.
Consequently, we have four equations:
𝑓𝑥 − 𝜆𝑔𝑥 = 0, 𝑓𝑦 − 𝜆𝑔𝑦 = 0, 𝑓𝑧 − 𝜆𝑔𝑧 = 0,
and 𝑔 = 𝑐,
on four unknowns 𝑥, 𝑦, 𝑧, and 𝜆.
In effect, the Lagrange multiplier method forms a new function
𝑓 ∗ ≔ 𝑓 − 𝜆(𝑔 − 𝑐) and extremizing 𝑓 ∗ subject to no constraints.
Dr. [Link], Bilkent University, UMRAM

Note that we need to append 𝑔 = 𝑐 equation because we lost the


effect of constant 𝑐 after differentiation 𝑑𝑔 = 0.

65
Method of Lagrange Multiplier (III)
We can extend the procedure for multi-constraint case.
Consider the problem of finding the extrema for 𝑓(𝑥1 , ⋯ , 𝑥𝑛 )
subject to the constraints 𝑔𝑗 𝑥1 , ⋯ , 𝑥𝑛 = 0; 𝑗 = 1, ⋯ , 𝑘.
Such a problem can be solved by converting it to an unconstrained
problem of finding the extrema for
𝑓 ∗ 𝑥,ҧ 𝜆ҧ = 𝑓 𝑥ҧ − 𝜆1 𝑔1 𝑥ҧ − ⋯ − 𝜆𝑘 𝑔𝑘 𝑥ҧ ,
where 𝜆ҧ = 𝜆1 , ⋯ , 𝜆𝑘 .
Dr. [Link], Bilkent University, UMRAM

66
Ex. 7, p.668

Method of Lagrange: Example (II)


Consider the same problem of minimum distance of minimizing
𝑓 ∗ 𝑥, 𝑦, 𝑧 = 𝑥 2 + 𝑦 2 + 𝑧 2 − 𝜆 2𝑥 − 𝑦 + 𝑧 − 3 .
We have
𝜕𝑓 ∗ 𝜕𝑓 ∗
= 2𝑥 − 2𝜆 = 2𝑦 + 𝜆
𝜕𝑥 𝜕𝑦
𝜕𝑓 ∗ 𝜕𝑓 ∗
= 2𝑧 − 𝜆 = −(2𝑥 − 𝑦 + 𝑧 − 3)
𝜕𝑧 𝜕𝜆

Setting all these to zero, we get


2 0 0 −2 𝑥 0 𝑥 1
0 2 0 1 𝑦 0 𝑦 −0.5
= ⇒ = .
Dr. [Link], Bilkent University, UMRAM

0 0 2 −1 𝑧 0 𝑧 +0.5
2 −1 1 0 𝜆 3 𝜆 1
Hence 𝑋ത = (1, −0.5,0.5) is a critical point and from the geometry of
the problem, must be the only critical point on the plane.
67
Remarks
➢ In 2D, we can visualize “minimization” of 𝑓(𝑥, 𝑦) subject to
𝑔 𝑥, 𝑦 = 𝑐 as follows:
If (𝑥,Ƹ 𝑦)Ƹ is a solution point, then minimum is curve 𝐿2 in the figure,
where the level curves and the curve 𝑔 𝑥, 𝑦 = 𝑐 are shown.
Thus, at (𝑥,Ƹ 𝑦), Ƹ gradients of 𝑓 and 𝑔 are parallel, i.e., there exists a 𝜆
such that: 𝑦
𝑔 𝑥, 𝑦 = 𝑐
𝛻𝑓 = 𝜆𝛻𝑔
⇒ 𝑓𝑥 − 𝜆𝑔𝑥 = 0, 𝑓𝑦 − 𝜆𝑔𝑦 = 0.
𝐿1
𝑥
➢ The Hessian matrix sufficient 𝐿2
𝐿3
condition, i.e., det 𝐻 ≠ 0, 𝐿4
Dr. [Link], Bilkent University, UMRAM

does not apply to constrained


Level curves 𝑓 𝑥, 𝑦 = 𝐿𝑗 ; 𝑗 = 1, ⋯ , 4
problems via Lagrange multiplier
method.
68

You might also like