Multivariable Vector-Valued Functions Guide
Multivariable Vector-Valued Functions Guide
Vector-Valued
Functions
Contents
Contents 2
31.1 Multivariable Vector-Valued Functions . . . . . . . . . . . 3
31.2 Derivatives . . . . . . . . . . . . . . . . . . . . . . . . . . . 20
31.3 Linear Approximations & Differentials . . . . . . . . . . . 34
31.4 The Chain Rule . . . . . . . . . . . . . . . . . . . . . . . . 42
31.2 Optional: Limit Definition of the Derivative for Multivari-
able Vector-Valued Functions . . . . . . . . . . . . . . . . 55
MULTIVARIABLE VECTOR-VALUED FUNCTIONS 3
-2 -1 1 2
-1
-2
Figure 1: Graph of f ( x ) x 2
Figure 2: Graph of f ( x, y ) x 2 − y 2
1.0
0.5
0.0
-0.5
-1.0
20
10
1.0
0.5
0.0
-0.5
0
-1.0
Figure 3: Graph of s ( t )
f1 ( x1 ,x2 ,...,x n )
F ( x1 , x2 , . . ., x n )
f2 ( x1 ,x2 ,...,x n ) ,
..
f ( x ,x. ,...,x )
m 1 2 n
Figure 4: Graph of F
xy
R2 R3 given by the formula G x, y
image of the function G : → x2 .
y
Figure 5: Graph of G
-10 0 10
10
-10
10
-10
lim P ( x, y )
( x,y )→( a,b )
lim Q ( x, y )
lim G x, y ( x, y ) → ( a,b ) ,
( x,y ) → ( a,b )
( x,ylim R ( x, y )
) → ( a,b )
lim P ( x, y )
P ( x, y ) ( x, y )→( a,b )
lim Q ( x, y ) ( x, ylim
Q ( x, y ) ,
( x, y ) → ( a,b ) ) → ( a,b )
R ( x, y ) lim R ( x, y )
( x, y )→( a,b )
Example 1
3x+2y
R2 R3 be given by the formula F x, y
Let F : → x2 y3 . Find
2xe y
lim F x, y , if the limit exists.
( x,y ) → (1,0)
SOLUTION The component functions of F are the three multivariable
real-valued functions P, Q, R : R2 → R of F given by the formulas
P ( x, y ) 3x + 2y, and Q ( x, y ) x 2 y 3 and R ( x, y ) 2xe y . We
observing that each of these component functions is continuous, and
hence we see that
lim Q ( x, y ) lim x 2 y 3 12 · 03 0
( x, y ) → (1,0) ( x,y ) → (1,0)
We deduce that lim F x, y exists, and that
( x, y ) → (1,0)
lim (3x + 2y )
( x, y )→(1,0) 3
2 3
lim F x, y ( x,ylim x y 0 .
( x, y ) → (1,0) ) → ( 1,0 )
( x,ylim
2xe 2
y
) → (1,0)
f1 ( x1 ,x2 ,...,x n )
F ( x1 , x2 , . . ., x n )
f2 ( x1 ,x2 ,...,x n ) ,
..
f ( x ,x. ,...,x )
m 1 2 n
lim F ( x1 , x2 , . . ., x n )
( x1 ,x2 ,...,x n ) → ( a1 ,a2 ,...,a n )
lim f1 ( x 1 ,x 2 ,...,x n )
( x1 ,x2 ,...,x n )→( a1 ,a2 ,...,a n )
lim f2 ( x 1 ,x2 ,...,x n )
x 1 ,x2 ,...,x n ) → ( a 1 ,a 2 ,...,a n )
( ,
.
..
lim f m ( x 1 ,x2 ,...,x n )
( x1 ,x2 ,...,x n )→( a1 ,a2 ,...,a n )
2. The above limit can be written without using name of the function
MULTIVARIABLE VECTOR-VALUED FUNCTIONS 14
F by simply writing
f1 ( x1 ,x2 ,...,x n )
f2 ( x1 ,x2 ,...,x n )
lim ..
( x1 ,x2 ,...,x n ) → ( a1 ,a2 ,...,a n ) .
fm ( x1 ,x2 ,...,x n )
lim f1 ( x 1 ,x 2 ,...,x n )
( x1 ,x2 ,...,x n )→( a1 ,a2 ,...,a n )
lim f2 ( x 1 ,x 2 ,...,x n )
x 1 ,x 2 ,...,x n ) → ( a 1 ,a 2 ,...,a n )
( ,
.
..
lim f m ( x 1 ,x2 ,...,x n )
( x1 ,x2 ,...,x n )→( a1 ,a2 ,...,a n )
Example 2
sin(x2 +y2 )
Find lim x22 +y22 , if the limit exists.
( x, y ) → (0,0) x 2 −y2
x +y
sin ( x 2 +y 2 )
SOLUTION First, we consider the limit lim 2 2 . Using
( x,y ) → (0,0) x +y
polar coordinates, we know r 2 x 2 + y 2 . Clearly, as ( x, y ) gets closer
and closer to (0, 0) , then r gets closer and closer to 0. Hence, using
l’Hôpital’s Rule, we see that
sin ( x 2 + y 2 ) sin ( r 2 )
lim lim
( x,y ) → (0,0) x2 + y2 r→0 r2
2r cos ( r 2 )
lim
r→0 2r
lim cos ( r 2 ) cos (02 ) 1.
r→0
x 2 −y 2
Next, we consider the limit lim 2 2. Let us see what happens
( x,y ) → (0,0) x +y
as ( x, y ) approaches (0, 0) from different directions. First, suppose
that ( x, y ) approaches (0, 0) along the x-axis, which means that we set
y 0 and then take the limit, which yields
x2 − y2 x 2 − 02
lim lim
( x,0) → (0,0) x2 + y2 ( x,0) → (0,0) x 2 + 02
lim 1 1.
( x,0) → (0,0)
MULTIVARIABLE VECTOR-VALUED FUNCTIONS 15
x2 − y2 02 − y 2
lim lim
(0, y ) → (0,0) x2 + y2 (0, y ) → (0,0) 02 + y 2
lim − 1 −1.
( x,0) → (0,0)
Example 3
5x+y 2
R2 R3 be given by the formula F x, y
Let F : → sin ( x y ) . Is F
xe y
continuous?
SOLUTION To figure out if F is continuous, we need to figure
out if each of the component functions P ( x, y ) 5x + y 2 , and
Q ( x, y ) sin ( x y ) and R ( x, y ) xe y are continuous. We know that
polynomials in each of x and y are continuous, so each of 5x and y 2 are
continuous, and we we know that sums and products of continuous
functions are continuous, and hence P ( x, y ) 5x + y 2 and x y are both
continuous. We know that sin x and e y are continuous, and we know
that compositions of continuous functions are continuous, and hence
Q ( x, y ) sin ( x y ) and R ( x, y ) xe y are continuous. We conclude
that F is continuous.
MULTIVARIABLE VECTOR-VALUED FUNCTIONS 16
SUMMARY
f1 ( x1 ,x2 ,...,x n )
F ( x1 , x2 , . . ., x n )
f2 ( x1 ,x2 ,...,x n ) ,
..
f ( x ,x. ,...,x )
m 1 2 n
f1 ( x1 ,x2 ,...,x n )
F ( x1 , x2 , . . ., x n )
f2 ( x1 ,x2 ,...,x n ) ,
..
f ( x ,x. ,...,x )
m 1 2 n
( a1 , a2 , . . . , a n ) is given by
lim F ( x1 , x2 , . . ., x n )
( x1 ,x2 ,...,x n ) → ( a1 ,a2 ,...,a n )
lim f1 ( x 1 ,x 2 ,...,x n )
( x1 ,x2 ,...,x n )→( a1 ,a2 ,...,a n )
lim f2 ( x 1 ,x2 ,...,x n )
x 1 ,x2 ,...,x n ) → ( a 1 ,a 2 ,...,a n )
( ,
.
..
lim f m ( x 1 ,x2 ,...,x n )
( x1 ,x2 ,...,x n )→( a1 ,a2 ,...,a n )
2. The above limit can be written without using name of the function
F by simply writing
f1 ( x1 ,x2 ,...,x n )
f2 ( x1 ,x2 ,...,x n )
lim ..
( x1 ,x2 ,...,x n ) → ( a1 ,a2 ,...,a n ) .
fm ( x1 ,x2 ,...,x n )
lim f1 ( x 1 ,x 2 ,...,x n )
( x1 ,x2 ,...,x n )→( a1 ,a2 ,...,a n )
lim f2 ( x 1 ,x 2 ,...,x n )
x 1 ,x 2 ,...,x n ) → ( a 1 ,a 2 ,...,a n )
( ,
.
..
lim f m ( x 1 ,x2 ,...,x n )
( x1 ,x2 ,...,x n )→( a1 ,a2 ,...,a n )
EXAMPLES
MULTIVARIABLE VECTOR-VALUED FUNCTIONS 18
Example 4
√
9−x 2 −y 2
R2 R2 be given by the formula H x, y
Let H : → . Find
ln ( x−1)
and sketch the domain of H.
SOLUTION Thep domain of H is the set of all points ( x, y ) in R2
for which both 9 − x 2 − y 2 and ln ( x − 1) are defined. That is, the
domain is the set of all points ( x, y ) in R2 for which both inequalities
9 − x 2 − y 2 ≥ 0 and x − 1 > 0 hold, which is the same as x 2 + y 2 ≤ 32
and x > 1. The solution of x 2 + y 2 ≤ 32 is the interior of the circle of
radius 3 centered at the origin together with the circle, shown below
in blue; the solution of x > 1 is the half-plane that is to the right of
the vertical line x 1, not including the line, shown below in red. The
domain of H is the intersection of these two regions, which is shown
below in purple, where the part of the boundary of the region that is
solid is included in the region, and the part of the boundary that is
dashed is not included in the region.
EXERCISES
Basic Exercises
MULTIVARIABLE VECTOR-VALUED FUNCTIONS 19
1–3 Find and sketch the domain of each function. 2. Let G : R2 → R3 be given by the formula
" ln( x+y ) #
√
G x, y
x2 y .
1. Let F : R2
→ R2
be given by the formula sin ( x y )
√x−y 2
F x, y √ .
9−x
DERIVATIVES 20
31.2 Derivatives
that order. That will be the first row in our matrix. Next, for the function
Q ( x, y ) , which is in the second column of the output, we will find its two
∂Q ∂Q
partial derivatives, which are ∂x and ∂y , and write them horizontally, in
that order, forming the second row of our matrix. Finally, for the function
R ( x, y ) , which is in the third column of the output, we will find its two
partial derivatives, which are ∂R ∂x
and ∂R
∂y
, and write them horizontally, in
that order, forming the third row of our matrix. The matrix of partial
derivatives is then
∂P ∂P
∂Q∂x ∂y
∂Q
∂x
∂R .
∂y
∂R
∂x ∂y
This matrix is called the derivative of G, and is denoted DG; if we need
to specify the input variables, we would write DG ( x, y ) , which could
also be written as DG ( p ) , where p denotes a point in R2 .
Example 1
5x+y
R2 R3 be given by the formula G x, y
Let G : → 3x 2 . Find the
2x+e y
derivative of G.
SOLUTION Using the partial derivatives we computed in Equation (1)
of this section, we see that
∂P ∂P
5 1
∂x ∂y
∂Q ∂Q
DG x, y ∂x 6x 0 .
∂y
∂R ∂R
2 e y
∂x ∂y
f1 ( x1 ,x2 ,...,x n )
F ( x1 , x2 , . . ., x n )
f2 ( x1 ,x2 ,...,x n ) .
..
f ( x ,x. ,...,x )
m 1 2 n
DERIVATIVES 23
∂P ∂P
∂x ∂y
∂Q ∂Q
DF x, y ∂x .
∂y
∂R ∂R
∂x ∂y
∂f ∂f ∂f
D f ( x1 , x 2 , . . ., x n ) ... .
∂x1 ∂x2 ∂x n
f10 ( t )
column vector r 0 ( t ) ... . Once again, we can think of we can think
f 0 ( t )
n
1 n
of r : R → R as a multivariable vector-valued function. Viewed that
way, the function f has a derivative which is an n × 1 matrix, namely,
∂∂tf1
f 0 ( t )
1 . . Of course, an n × 1 matrix is the
the matrix Dr ( t ) ...
0 ..
∂ fn
fn (t )
∂t
same as a column vector, and so we see that viewing a r : R → Rn as
a multivariable vector-valued function just gives us the derivative we
already knew, except for thinking of it as a matrix rather than a column
vector.
All told, if F : Rn → Rm and n , 1 and m , 1, then the derivative
DF that we have currently defined is something genuinely new, and if
either n 1 or m 1, or both, then DF is just a slightly different way of
writing the derivatives we are already familiar with. Hence, what we are
considering at present incorporates as special cases everything we have
seen up till now regarding derivatives.
Finally, we note that for a function of the form F : Rn → Rm , the
columns of the derivative matrix of F are tangent vectors to parameter
curves of the function, as defined in Section 31.1.
1. D ( F + G )( p ) DF ( p ) + DG ( p ) ;
2. D ( F − G )( p ) DF ( p ) − DG ( p ) ;
3. D ( cF )( p ) cDF ( p ) .
Example 2
3x 2 y
Let F : R2 → R2 be given by the formula F x, y
5x+y 3
. Find the
Jacobian determinant of F.
6x y 3x 2
SOLUTION We compute DF x, y 5 3y 2
, and hence
det DF x, y 6x y · 3y 2 − 3x 2 · 5 18x y 3 − 15x 2 .
f P ( x,y ) g
a b . Specifically, suppose we have F x, y
fashioned notation c d Q ( x,y ) .
∂P ∂P
" #
∂x ∂y
Then DF x, y
∂Q ∂Q , and whereas we would then write the Jacobian
∂x ∂y
∂P ∂P
" #
∂x ∂y
determinant of F as det DF x, y det
∂Q ∂Q , the people who use
∂x ∂y
the alternative notation for determinants would write the Jacobian
∂P ∂P ∂P ∂P
∂x ∂y ∂x ∂y
determinant of F as ∂Q ∂Q . Because writing ∂Q ∂Q by hand can look
∂x ∂y ∂x ∂y
∂P ∂P ∂P ∂P
" # " #
∂x ∂y ∂x ∂y
very similar to ∂Q ∂Q , we recommend using the notation det ∂Q ∂Q ,
∂x ∂y ∂x ∂y
which is unambiguous. f P ( x,y ) g
Additionally, rather than writing F as Q ( x,y ) , where x and y are
the independent variables, it is sometimes convenient to write x and y
as the dependent variables, and u and v as the independent variables,
and one then writes xf g ( u,g v ) and y h ( u, v ) , which we would write
g ( u,v )
as T ( u, v ) xy h ( u,v ) . The Jacobian determinant of T is then
∂g ∂g
∂u ∂v
det ∂h ∂h . Sometimes, however, functions are not given names, so that
∂u ∂v
rather than writing x g ( u, v ) , we simply think of x as a function of
u and v, which we could write x x ( u, v ) , and similarly for y. The
Jacobian determinant of x and y as functions of u and v would then be
written, without function names and with the old fashioned notation for
∂x ∂x
∂u ∂v
the determinant, as ∂y ∂y . Moreover, instead of writing the Jacobian
∂u ∂v
determinant in this situation in either of the above notations, some people
∂ ( x,y )
use the old-fashioned notation ∂ ( u,v ) , which does not even need the name
of the function. Similar notation is used with more variables.
We summarize these ideas as follows.
det DF ( x1 , x2 , . . ., x n ) ,
f g
Observe that H ( e1 ) [ ac ] and H ( e2 ) db . As we saw in our
discussion of the cross product, the area of the area off thegparallelogram
formed by the two vectors H ( e1 ) and H ( e2 ) is | det ac db |. Hence, we
see that applying the linear map H to the original square results in
DERIVATIVES 29
f g f g
multiplying the area by | det ac db |. We note that the matrix ac db should
look familiar in this context, because it is precisely DH.
More generally, let F : Rn → Rn be a function, and let p be a point in
Rn . Of course, the function F need not be a linear map. However, as we
will see in more detail in Section 31.3, the matrix DF ( p ) is the matrix of
the linear map that best approximates F at the point p. Then, analogously
to what we saw above for the linear map H, it turns out more generally
that if we were to take a small region A in Rn that is near p, then the
area of F ( A ) is approximately equal to DF ( p ) times the area of A. The
smaller the region, and the closer to p, the better the approximation. The
proof of this fact is not simple, and it requires certain hypotheses on the
function F, but the inutive idea is simply that we use the matrix DF ( p ) ,
and that matrix multiplication is a linear map.
This geometric way of thinking about DF ( p ) will be used when we
discuss the change of variable formula for double and triple integrals.
SUMMARY
f1 ( x1 ,x2 ,...,x n )
F ( x1 , x2 , . . ., x n )
f2 ( x1 ,x2 ,...,x n ) .
..
f ( x ,x. ,...,x )
m 1 2 n
∂P ∂P
∂x ∂y
∂Q ∂Q
DF x, y ∂x .
∂y
∂R ∂R
∂x ∂y
1. D ( F + G )( p ) DF ( p ) + DG ( p ) ;
2. D ( F − G )( p ) DF ( p ) − DG ( p ) ;
3. D ( cF )( p ) cDF ( p ) .
det DF ( x1 , x2 , . . ., x n ) ,
f P ( u,v ) g
3. If F ( u, v ) Q ( u,v ) , then the Jacobian determinant of F is some-
∂x ∂x
∂u ∂v ∂ ( x,y )
times denoted ∂y ∂y or ∂ ( u,v )
.
∂u ∂v
EXAMPLES
Example 3
" 3x+2y−z #
R3 R3 be given by the formula H x, y, z
Let H : → x 2 +y 3 . Find
4xz 2
the Jacobian determinant of H.
3 2 −1
SOLUTION We compute DH x, y, z 2x 3y 2 0 . Hence, by
4z 2 0 8xz
expanding along the bottom row of this matrix, we deduce that
det DH x, y, z 4z 2 · (0 − (−3y 2 )) − 0 · (0 − (−2x )) + 8xz · (9y 2 − 4x )
12y 2 z 2 + 72x y 2 z − 32x 2 z.
Example 4
6x+y 2
R2 R2 be given by the formula F x, y
Let F : → 3x 2 +2y
. Find and
plot all points ( x, y ) in R2 for which det DF x, y 0.
f g
6 2y
We compute DF x, y 6x . Hence det DF x, y
SOLUTION 2
6·2−6x·2y 12−12x y. Therefore det DF x, y 0 yields 12−12x y 0,
which is the same as x y 1, which in turn is the equivalent to y x1 .
Hence, plotting all points for which det DF x, y 0 is the same as
drawing the graph of y x1 , which is shown below.
DERIVATIVES 32
COMMON MISTAKES
EXERCISES
Basic Exercises
DERIVATIVES 33
1–3 Find the derivative of each function. 4. Let F : R2 → R2 be given by the formula
x2 y
F x, y
3x−y 2
.
1. Let F : R3 R2
→ be given by the formula
3x 2 +yz
F x, y, z xz 3 +2y 5 . 5. Let G : R3 → R3 be given by the formula
3x+2y+z
G x, y, z
e 5x+y .
yz
2. Let G : R2 → R4 be given by the formula
e x +sin y
ln( x y )
G x, y √x .
6. Let H : R4 → R4 be given by the formula
x 3 y 4x−3y+2z−w
H x, y, z, w 3y+zw
.
4zw
w3
3. Let H : R4 → R3 be
given by the formula
x+2y+3z+4w
H x, y, z, w
xy .
6zw
7. Let K : R2 →f R2 be given
g by the formula
2 2
K x, y x −12x−y
xy . Find and plot all points
( x, y ) in R for which det DK x, y 0.
2
4–6 Find the Jacobian determinant of each function.
LINEAR APPROXIMATIONS & DIFFERENTIALS 34
In Calculus I, we saw the formula for the tangent line for a single-variable
real-valued function y f ( x ) at a point x a, which is
L ( x ) f ( a ) + f 0 ( a )( x − a ) . (1)
This formula was used to find approximate values for functions that are
difficult to compute. Similarly, in Calculus II, we saw the formula for the
tangent plane for a multivariable real-valued function z f ( x, y ) at a
point ( x, y ) ( a, b ) , which is
L ( x, y ) f ( a, b ) + f x ( a, b )( x − a ) + f y ( a, b )( y − b ) . (2)
Example 1
√
y+ x
R2 R2 be given by the formula F x, y
Let F : → √
x+ 3 y
. Find the
linear approximation of F at p (9, 8) .
1
1
" √
#
2 x
We compute DF x, y . Then F (9, 8)
11
SOLUTION 1 1
11
2
2y 3
1
1
and DF (9, 8) 6
1 1 , and we deduce that
12
" # " #!
x 9
L x, y F (9, 8) + DF (9, 8)
−
y 8
1 1
x + y + 23
" # " #" # " #
11 1 x−9
+ 6
1 6
1 .
11 1 12 y−8 x + 12 y + 43
and multivariable real-valued functions, here too our ability to make use
of Equation (3) to approximate F ( v ) depends upon finding a convenient
point p such that on the one hand p is very close to v (the closer p is to v
the better the approximation will be), and such that on the other hand it
is easy to compute both F ( p ) and DF ( p ) . Of course, in practice, it is not
always possible to find such a point p, but it is possible in enough cases
that this very easy to use method of approximation is of value.
Example 2
√
y+ x
R2 R2 be given by the formula F x, y
Let F : → √
x+ 3 y
. Use the
linear approximation to compute an approximate value for F (9.1, 7.8) .
SOLUTION We saw in Example 1 of this section that the linear ap-
proximation of F at p (9, 8) is
1
" # " #" #
11 1 x−9
L x, y + 6 .
1
11 1 12 y−8
Then
1
" # " #" # " #
11 1 9.1 − 9 1 649
F (9.1, 7.8) ≈ L (9.1, 7.8) + 6 1 .
11 1 12 7.8 − 8 60 665
∆y f ( a + ∆x ) − f ( a )
≈ L ( a + ∆x ) − f ( a ) f ( a ) + f 0 ( a )(( a + ∆x ) − a ) − f ( a ) (4)
f ( a ) ∆x.
0
LINEAR APPROXIMATIONS & DIFFERENTIALS 37
∆z F ( p + ∆v ) − F ( p )
≈ L ( p + ∆v ) − F ( p )
(5)
F ( p ) + DF ( p )(( p + ∆v ) − p ) − F ( p )
DF ( p ) ∆v.
dz DF ( p ) dv, (6)
dF ( p ) DF ( p ) dv.
LINEAR APPROXIMATIONS & DIFFERENTIALS 38
∆z ∆F ( p ) F ( p + ∆v ) − F ( p )
Example 3
√
y+ x
Let F : R2 → R2 be given by the formula F x, y
√
x+ 3 y
. Find dF ( p ) .
√1 1 " dx # 2√1 x dx + dy
2 x
dz dF ( p ) DF x, y dv 1 dx + 1 dy .
1
2 dy
2
2y 3 2y 3
SUMMARY
∆z ∆F ( p ) F ( p + ∆v ) − F ( p )
EXAMPLES
Example 4
Let F : R2 → R2 be a function. Suppose that F (1, 2)
5
7 and that
DF (1, 2) 34 15 . Estimate the value of F (1.02, 1.97) .
COMMON MISTAKES
that can be added on to the real numbers; the study of that approach,
which would take us very far afield, is called non-standard analysis.
EXERCISES
Basic Exercises
1–2 For each function, use the linear approximation 2. Let G : R3 → R2 be given by the formula
x+2y+3z
f g
to compute the requested approximation. G x, y z e √3 x yz . Compute an approximate
[ f ( g ( x )) ]0 f 0 ( g ( x )) g 0 ( x ) .
Composition of Functions
Rather than starting with a single function, for example h : R → R
given by the formula h ( x ) sin ( x 2 + 7) , and then decomposing it by
writing h ( x ) f ( g ( x )) , where f : R → R and g : R → R are given by the
formulas f ( x ) sin x and g ( x ) x 2 + 7, we now want to start with two
functions such as f and g and then combine them.
For example, let f : R → R and g : R → R be given by the formulas
f ( x ) x 2 and g ( x ) x + 3. We then want to construct a new function,
which will be given by the formula f ( g ( x )) ( x + 3) 2 . It would be very
convenient to have a name for this combined function. We could call this
function by a new letter, for example writing k : R → R as k ( x ) ( x + 3) 2 ,
but calling the function that results from combining f and g by an
arbitrary name such as k is confusing. It would be much better to give
the new function a name that reflects that it is made up of f and g, and
that is what we now state.
Let f : R → R and g : R → R be functions. The composition of f and
g is the function f ◦ g : R → R given by the formula
( f ◦ g )( x ) f ( g ( x )) . (1)
Composition of Functions
Let G : Rn → Rk and F : Rk → Rm be functions. The composition of F
and G is the function F ◦ G : Rn → Rm given by the formula
( F ◦ G )( p ) F ( G ( p )) . (2)
Example 1
let F : R3 → R4 and G : R2 → R3 be given by the formulas
" 5x+y # xy
F x, y, z 3z G x, y .
2xz
and x+2y
y−z x−y
Find the formula for ( F ◦ G ) x, y .
SOLUTION The function F ◦ G : R2 → R4 is given by the formula
[ ( f ◦ g )( x ) ]0 f 0 ( g ( x )) g 0 ( x ) . (3)
D ( F ◦ G )( p ) DF ( G ( p )) DG ( p ) , (4)
We will not prove Equation (4), but let us look at an example, to see
that it really works.
THE CHAIN RULE 46
Example 2
Let F : R3 → R2 and G : R2 → R3 be given by the formulas
f 2x y g x−4y
F x, y, z G x, y .
y−z and xy
x+y
Find each of DF ( G x, y ) and DG x, y and D ( F ◦ G ) x, y directly,
and verify the The Chain Rule via Matrix Multiplication for this
example.
SOLUTION The function F ◦ G : R2 → R2 is given by the formula
f 2( x−4y )( x y ) g f 2 y−8x y 2 g
( F ◦ G ) x, y ( x y )−( x+y ) 2xx y−x−y .
We then have
f g f g
2 ( x y ) 2 ( x−4y ) 0 2x y 2x−8y 0
DF ( G x, y ) .
0 1 −1 0 1 −1
Finally, we compute
f g1
−4
2x y 2x−8y 0
DF ( G x, y ) DG x, y
y x
0 1 −1 1 1
f 2x y·1+(2x−8y ) ·y+0·1 2x y· (−4) + (2x−8y ) ·x+0·1
g
0·1+1·y−1·1 0· (−4) +1·x−1·1
f 4x y−8y 2 2x 2 −16x y
g
y−1 x−1
D ( F ◦ G ) x, y .
Hence, we see that the The Chain Rule via Matrix Multiplication
works for this example.
it is also possible to state and use specific cases of this Chain Rule without
matrices. This non-matrix formulation of the Chain Rule will be written
entirely using Leibniz notation for derivatives. To start, let us recall the
statement of the Chain Rule for single-variable real-valued functions in
Leibniz Notation, which is as follows.
Let f : R → R and g : R → R be functions. We write y f ( u ) and
u g ( x ) . Suppose that f and g are differentiable. Then
dy dy du
. (5)
dx du dx
We make a few observations about this formulation of the Chain
Rule for single-variable real-valued functions. First, the formula in
Equation (5) is completely equivalent to the formula in Equation (3); only
the notation has changed. Second, observe that in Equation (5), we did
not make use of the names of the function f and g, but rather, all we
used was the names of the variables. Third, whereas it appears as if we
could prove Equation (5) by simply canceling the two appearances of
du in the right-hand side of the equation, such canceling, which useful
mnemonically, is not actually a valid thing to do. We will soon see why
it is not valid.
To see the non-matrix version of the Chain Rule for multivariable
vector-valued functions, we first look at various special cases, starting
with functions of the form f : R2 → R and r : R → R2 , and with the
composition f ◦ r. Let t be a real number. By Equation (4) of this section
we see that
D ( f ◦ r )( t ) D f ( r ( t )) Dr ( t ) . (6)
∂z ∂z
" dx #
dz
dt
dy .
dt ∂x ∂y dt
Observe that in the above version of the Chain Rule, we write single-
dy
variable derivatives dz dx
dt , dt and dt , whereas we write partial derivatives
∂z ∂z
∂x
and ∂y . It is important to keep track of which functions have one
variable, which then use the notation for single-variable derivatives,
and which functions have more than one variable, which then use the
notation for partial derivatives.
Example 3
√
Let z x 2 + y 3 , and x t and y sin t. Find dz
dt .
SOLUTION First, we compute the various derivatives (some partial
and some regular) that appear in the right-hand side of Equation (7)
∂z ∂z
of this section, obtaining ∂x 2x, and ∂y 3y 2 , and dx
dt
1
√ and
2 t
dy
dt cos t. Next, using Equation (7), we compute
dz ∂z dx ∂z dy
+
dt ∂x dt ∂y dt
1
2x · √ + 3y 2 · cos t
2 t
√ 1
2 t · √ + 3 (sin t ) 2 · cos t
2 t
1 + 3 sin2 t cos t.
dw ∂w dx ∂w dy ∂w dz
+ + .
dt ∂x dt ∂y dt ∂z dt
Clearly, the same idea holds for functions of more than three variables.
Next, we look at at functions of the form f : R2 → R and G : Rf2 → R2g ,
x ( s,t )
with the composition f ◦ G. We write z f ( x, y ) and G x, y y ( s,t ) .
THE CHAIN RULE 49
D ( f ◦ G ) x, y D f ( G x, y ) DG x, y ,
∂z ∂x ∂z ∂y ∂z ∂x ∂z ∂y
" #
∂z ∂z
+ + ,
∂s ∂t ∂x ∂s ∂y ∂s ∂x ∂t ∂y ∂t
∂z ∂z ∂x ∂z ∂y ∂z ∂z ∂x ∂z ∂y
+ and + . (8)
∂s ∂x ∂s ∂y ∂s ∂t ∂x ∂t ∂y ∂t
This last equation is the Chain Rule in the particular case of a function
of the form z f ( x, y ) , where each of x and y are functions of two
variables s and t.
We can now see that the wish of beginning calculus students to cancel
the two appearances of du in the right-hand side of Equation (5) is not
a valid thing to do. If such a cancellation were valid, then it would
plausibly also be valid to cancel the various instances of ∂x and ∂y in
either of the equations of Equation (8), but doing so in the left-hand
equation would yield ∂z ∂s
∂z + ∂z , which, upon cancelling, yields 1 1+1,
∂s ∂s
and that is clearly not possible. Hence, this type of canceling is not valid.
The same type of argument as above in the case of w f ( x, y, z ) ,
where each of x, y and z are functions of two variables s and t, yields
the equations
∂w ∂w ∂x ∂w ∂y ∂w ∂z ∂w ∂w ∂x ∂w ∂y ∂w ∂z
+ + and + + .
∂s ∂x ∂s ∂y ∂s ∂z ∂s ∂t ∂x ∂t ∂y ∂t ∂z ∂t
Clearly, the same idea holds for functions of more than three variables.
Example 4
∂z ∂z
Let z 5x 2 y, and x 3s + t and y sin ( st ) . Find ∂s
and ∂t
.
SOLUTION First, we compute the various derivatives (some partial
and some regular) that appear in the right-hand side of Equation (8) of
∂z ∂z
this section, obtaining ∂x 10x y, and ∂y 5x 2 , and ∂x
∂s
3, and ∂x
∂t
1,
THE CHAIN RULE 50
∂y dy
and ∂s t cos ( st ) and dt s cos ( st ) . Next, using Equation (8), we
compute
∂z ∂z ∂x ∂z ∂y
+
∂s ∂x ∂s ∂y ∂s
10x y · 3 + 5x 2 · t cos ( st )
10 (3s + t ) sin ( st ) · 3 + 5 (3s + t ) 2 · t cos ( st )
30 (3s + t ) sin ( st ) + 5t (3s + t ) 2 cos ( st )
and
∂z ∂z ∂x ∂z ∂y
+
∂t ∂x ∂t ∂y ∂t
10x y · 1 + 5x 2 · s cos ( st )
10 (3s + t ) sin ( st ) + 5s (3s + t ) 2 cos ( st ) .
∂z
Suppose we want to find the formula for . We then consider the
∂s
tree diagram in Figure 1, and we see that to go from z to s, there are two
routes, one of which is via x and the other via y. These two routes tell us
THE CHAIN RULE 51
that in the Chain Rule for this particular situation, we have two terms,
∂z ∂x ∂z ∂y
which are , representing the route via x, and , representing
∂x ∂s ∂y ∂s
the route via y. We then add these terms, and obtain the equation on
the left-hand side of Equation (8) of this section. Clearly, we can use the
analogous method in other cases.
Finally, we can write out the general version of the Chain Rule without
matrices as follows.
SUMMARY
Composition of Functions
Let G : Rn → Rk and F : Rk → Rm be functions. The composition of F
and G is the function F ◦ G : Rn → Rm given by the formula
( F ◦ G )( p ) F ( G ( p )) .
D ( F ◦ G )( p ) DF ( G ( p )) DG ( p ) ,
EXAMPLES
Example 5
The height and width of a rectangle are each changing as a function of
time. Suppose that at a certain instant of time, the height is 3 inches
and is increasing at a rate of 2 inches per second, and the width is 5
inches and is decreasing at a rate of 1 inch per second. Find the rate of
change of the area of the rectangle at that instant of time.
SOLUTION Let h denote the height of the rectangle and let w denote
the width of the rectangle. The area of the rectangle as a function
of height and width is the function A : R2 → R given by the formula
A ( h, w ) hw. Then ∂A
∂h
∂A
w and ∂w h, and hence at the given instant
∂A ∂A
of time we have ∂h 5 and ∂w 3.
Let t denote the time. We think of h and w as functions of t. The
information in the problem says that at the given instant of time, we
have dhdt 2 and dt −1.
dw
THE CHAIN RULE 53
By Equation (7) of this section we then see that at the given instant
of time, we have
dA ∂A dh ∂A dw
+ 5 · 2 + 3 · (−1) 7.
dt ∂h dt ∂w dt
COMMON MISTAKES
EXERCISES
Basic Exercises
y−x
e 2x+z
f g
6. F x, y, z and G x, y 13. The height and radius of a cylinder are each
y 2 +xz
4x .
xy
changing as a function of time. Suppose that at a
certain instant of time, the height is 3 inches and is
For each set of functions, find dz decreasing at a rate of 5 inches per second, and the
7–9 dt .
radius is 2 inches and is increasing at a rate of 2
inches per second. Find the rate of change of the
7. z x 4 + 5y, and x 3t + 1 and y t 2 . volume of the cylinder at that instant of time. Is the
volume increasing or decreasing at that instant of
8. z 3x y 2 , and x tan t and y t 3 .
time?
9. z sin ( x y ) , and x 2t 5 and y t + 2.
14. The length, width and height of a box are each
changing as a function of time. Suppose that at a
∂z ∂z certain instant of time, the length is 4 inches and is
10–12 For each set of functions, find ∂s
and ∂t
.
increasing at a rate of 2 inches per second, the
width is 3 inches and is decreasing at a rate of 6
10. z 3x + y 2 , and x s + 2t and y s − t.
inches per second, and the height is 2 inches and is
11. z x 2 y, and x s 2 t 3 and y 2t. increasing at a rate of 1 inch per second. Find the
rate of change of the volume of the box at that
12. z cos (2x + 3y ) , and x s 3 and y st. instant of time. Is the volume increasing or
decreasing at that instant of time?
Optional Section: Limit Definition of the Derivative for Multivariable Vector-Valued
Functions 55
f (a + h ) − f (a )
f 0 ( a ) lim , (1)
h→0 h
provided that the limit exists. Recall too that partial derivatives are
defined via similar limits.
The question arises, can the derivative of multivariable vector-
valued functions also be defined via some sort of limit? Of course,
it would be very troubling if such a limit definition existed but it
were not equivalent to the definition of the derivative of multivariable
vector-valued functions that we have already seen. We will now see
that a limit definition of the derivative of multivariable vector-valued
functions is indeed possible, and it yields the same matrix we are
already familiar with, though we will skip the rigorous details (which
are not trivial).
Let F : Rn → Rm be a multivariable vector-valued function, and let
p be a point in Rn . If we want to find the derivative of F at p via a limit,
the first thing one might attempt is an exact copy of the right-hand
side of Equation (1), but with the single-variable real-valued function
f replaced by the multivariable vector-valued function F, and with the
Optional Section: Limit Definition of the Derivative for Multivariable Vector-Valued
Functions 56
We will not go into further theoretical details, but let us look at the
relation of the limit definition and the matrix of partial derivatives via
an example.
x2
R2 R3 is given by the formula F x, y
Suppose that F : → 3y .
5x y
Then
2x 0
DF x, y 0 3 .
5y 5x
x
Let p y and k [ st ]. Then the left-hand side of Equation (3)
becomes
F ( p + k ) − [F ( p ) + DF ( p ) k]
lim
k→0 |k|
( x + s ) 2 x 2 2x 0 " #
3 ( y + t ) − *. 3y + 0 3 s +/
t
5 ( x + s )( y + t ) , 5x y 5y 5x
limf g
-
√
s → 0 s2 + t2
[t] 0
x 2 + 2xs + s 2 − x 2 − 2xs − 0
3y + 3t − 3y − 0 − 3t
5x y + 5xt + 5ys + 5st − 5x y − 5yt − 5x y
limf g √
[ st ]→ 00 s2 + t2
√ s 2 q s
s 2 +t 2 1+ t 22
s
lim 0 lim 0 . (4)
s→0 5st s→0 5t
t→0 √ t→0 q
s 2 +t 2
1+ st 22
To complete this calculation, we observe that no matter what non-zero
values of s and t we substitute in the expression q 1 2 , the denominator
1+ t 2
s
will always be at least 1, and therefore the whole fraction will be no
greater than 1. It then follows that
s 1
lim q lim s q 0,
s→0 t2 s→0 t2
t→0 1+ s2 t→0 1+ s2
and similarly
5t 1
lim q lim 5t q 0.
s→0 t2 s→0 t2
t→0 1+ s2 t→0 1+ s2
We deduce that the last expression in Equation (4) equals the zero
vector, which is what we wanted to verify.
Optional Section: Limit Definition of the Derivative for Multivariable Vector-Valued
Functions 59