Calculus Part1 Introduction
Calculus Part1 Introduction
I Introduction to Calculus 1
iii
CONTENTS
iv
Part I
Introduction to Calculus
1
Chapter 1
Learning Objectives
Describe the two central questions of calculus: the tangent problem and the area problem.
Understand why ordinary algebra is not enough to handle continuously changing quanti-
ties.
See, through a rst informal example, how zooming in forever turns a curve into a
straight line.
# Prerequisites
The idea of a graph: that a point (x, y) can be plotted on a pair of perpendicular axes.
Nothing else. If Part 0 (Mathematical Foundations) has not been studied yet, that is ne
for this chapter we will ag anything specic we need as we go.
Before calculus, essentially all of the mathematics you have ever done falls into one of two camps.
The rst camp is arithmetic and algebra of xed quantities. If a rectangle has width
4 and height 7, its area is 4 × 7 = 28. This is a xed, static fact. Nothing is changing. Two
numbers combine to produce a third number, once and for all.
The second camp is geometry of xed shapes. A circle of radius r has area πr2 . Again
for a specic value of r, we get a specic, unchanging area. Even though r is a variable in the
sense that it can stand for dierent circles, for any one circle it is just a xed number.
Calculus exists because the real world is not made of xed quantities. It is made of things
that are changing, continuously, from one instant to the next. A car's speed changes
continuously as the driver presses the gas pedal. The population of a country changes continu-
3
CHAPTER 1. WHAT IS CALCULUS, REALLY?
ously (well, one birth or death at a time, but for millions of people it behaves as if continuously).
The temperature in a room changes continuously as a window is opened. The amount of water
in a bathtub changes continuously as a tap runs.
Ordinary algebra is superb at describing a snapshot. Calculus is the mathematics of describing
continuous change itself and, remarkably, of using that description to answer very concrete,
very practical questions.
Every time you see a wall of algebra, rst ask: What is this trying to capture about the
real world? Calculus was not invented by mathematicians doodling for fun (though some
of it was!). Nearly every idea in this book was invented to answer a real, pressing question
often about motion, or about area, or about optimization. If you keep the real-world
question in mind, the algebra becomes much easier to follow, because you always know why
each step is being taken.
Let's make continuous change concrete with an example everyone has experienced: driving a
car, or riding in one.
Imagine you are on a highway. At every single moment, your car has a speed the number
that shows up on the speedometer. This speed is not usually constant. You speed up, you slow
down, you speed up again. The speedometer needle sweeps back and forth smoothly; it doesn't
jump from 50 km/h to 80 km/h in an instant, it passes through every value in between.
Here is the key question that motivates the rst half of this entire book:
If I know the exact position of the car at every moment in time, can I gure out its
exact speed at any single, particular instant?
This sounds like it should be easy after all, speed is just distance traveled divided by time
taken, right? Let's test that idea and see where it breaks down.
Suppose a car travels 120 kilometers in 2 hours. We are all taught that:
y(b) − y(a)
.
b−a
For motion, if y represents position and x represents time, this is exactly the average speed
formula above (with a sign to indicate direction, which we will discuss in Part 4).
4
1.2. THE CENTRAL PROBLEM: QUANTITIES THAT CHANGE CONTINUOUSLY
y(b) − y(a)
* Formula deep-dive:
b−a
Let's take this formula apart piece by piece, the way we will for every important formula in
this book.
What does it mean? The numerator y(b) − y(a) is the total change in the output quantity
how much y went up or down overall. The denominator b − a is the total change in the
input quantity how much time (or distance, or whatever x represents) elapsed. Dividing
total output change by total input change gives a single number: on average, how much did
y change per unit of x?
Why does it work? This is just the ordinary idea of rate from arithmetic, generalized.
60
If you earn $60 over 2 hours of work, your rate of pay is
2 = 30 dollars per hour even
though you may have been paid unevenly minute to minute. The formula doesn't care what
happens between a and b; it only compares the two endpoints. This is both its strength (it's
easy to compute you only need two data points) and its weakness (it hides everything
that happens in between, which is exactly the limitation this chapter is exposing).
Where does it come from? It comes directly from the denition of slope in coordinate
geometry: if you plot y against x, the two points (a, y(a)) and (b, y(b)) determine a straight
y(b)−y(a) rise
line (a secant line ), and is exactly the slope of that line, i.e.,
b−a run . Average rate
of change and slope of the secant line are the same number viewed two ways one as a
physical rate, one as a geometric steepness.
When should you use it, and when does it fall short? Use it whenever you only need
a summary rate over a whole interval average speed for a trip, average growth rate for a
year, average velocity for a swing of a pendulum. It falls short whenever you need to know
what is happening at one exact instant, which is exactly the gap that forces us toward the
shrinking-interval idea in the rest of this chapter.
So the tool we already have division gives us average rates over an interval of time.
What it fundamentally cannot give us, on its own, is the rate at a single instant, because an
instant has no duration. If we try to compute
y(a) − y(a) 0
= ,
a−a 0
0
we get the meaningless expression . This is not a minor technical annoyance it is the exact
0
obstacle that an entirely new branch of mathematics needed to be invented to overcome.
! The 00 trap
0
A very common instinct is to say well,
0 should just be 1, since anything divided by itself
is 1, or alternatively it should be 0, since the numerator is 0. Both instincts are wrong,
0
and more importantly, both miss the point. The expression
0 is undened it does not
0
correspond to a single number at all. The genius of calculus is not to nd a clever value for .
0
It is to realize that we never actually need to evaluate an expression at the problematic point;
instead we study what happens as we get arbitrarily close to it. This is the idea of a limit,
which we introduce informally in Section 1.3 of this chapter and treat fully and rigorously in
Part 3.
Here is the crucial idea, and if you understand this one idea deeply, you already understand the
conceptual heart of dierential calculus.
5
CHAPTER 1. WHAT IS CALCULUS, REALLY?
Instead of asking for the speed at the single instant t = 1.2 hours, let's ask for the average
speed over a very short interval that contains t = 1.2: say, from t = 1.2 to t = 1.2 + h, where h
is some small positive number of hours.
y(1.2 + h) − y(1.2)
average speed on [1.2, 1.2 + h] = .
h
Now here is the idea: what happens to this average speed as we make h smaller and smaller?
Ifh = 1 hour, we get the average speed over a whole hour probably not very representative of
the instantt = 1.2. If h = 0.1 hours (6 minutes), the average is a much better approximation of
the speed near t = 1.2. If h = 0.01 hours (36 seconds), better still. If h = 0.0001 hours, better
yet again.
y
e
position
rg
la
t,
n
h
a
c
ll
se
a
sm
t,
n
a
c
se
t = 1.2
time t
Figure 1.1: As h shrinks, the straight line joining the two points on the curve (called a secant
line ) swings around and hugs the curve more and more closely near t = 1.2. In the limit, it
becomes the tangent line and its slope is the instantaneous speed.
As h shrinks toward 0, two things happen simultaneously, and both are essential:
1. The numerator y(1.2 + h) − y(1.2) also shrinks toward 0 (the car doesn't move much in a
tiny sliver of time).
2. The ratio of these two shrinking quantities settles down and approaches a single, specic,
nite number.
That settling-down number is what we will call the instantaneous speed at t = 1.2. It is
0
not obtained by plugging in h=0 (which gives the meaningless ). It is obtained by watching
0
the trend as h approaches 0 without ever actually reaching it.
Try this yourself with a concrete function. Suppose position is given by y(t) = t2 (measured
in some convenient units). At t = 1.2:
6
1.3. A PREVIEW OF THE LIMIT IDEA
y(1.2 + h) − y(1.2)
h
h
(2.2)2 − (1.2)2
1 = 3.4
2
1 2
(1.3) − (1.2)
0.1 = 2.5
2
0.1 2
(1.21) − (1.2)
0.01 = 2.41
0.01
0.001 2.401
0.0001 2.4001
The pattern is unmistakable: as h → 0, the ratio is approaching 2.4, which happens to equal
2 × 1.2. This is not a coincidence it is our rst glimpse of the derivative of t2 , which we
will prove in general in Part 4 is 2t. For now, just notice: we found a denite, exact answer
to an instantaneous question, without ever dividing by zero.
What we just did watching a ratio settle down to a xed value as h approaches (but never
reaches) 0 is the seed of the single most important idea in all of calculus: the limit.
Informally:
We say that a quantity L is the limit of an expression as some variable approaches a particular
value, if we can make the expression as close as we like to L by making the variable close
enough (but not necessarily equal) to that particular value.
This is deliberately vague for now phrases like as close as we like and close enough are
placeholders for precise mathematical language that we will build carefully, piece by piece, in
Part 3. Trying to make this fully rigorous right now, before you have any intuition for it, would
be like trying to read the ne print of a contract in a language you haven't learned yet. So for
this entire rst Part of the book, we will lean on intuition, pictures, and numerical experiments
like the table above. Rigor comes later, once your intuition is strong enough to make the rigor
feel necessary rather than arbitrary.
Mathematicians did, in fact, try approaches like this historically (and we'll see some of that
0
struggle in Chapter 4). The trouble is that dierent problems, all involving a
0 -type ratio,
settle down to dierent nite values depending on the specic functions involved. There
0
is no single universal number that
0 could be assigned that would work for every problem.
0
The limit process sidesteps this entirely: instead of asking what number is
0 , it asks
what number does this specic ratio approach, and the answer is allowed to depend on the
problem. This exibility is exactly what makes it powerful.
Instantaneous rates of change (like instantaneous speed) form one half of calculus, called dier-
ential calculus. There is a second, seemingly unrelated question that turns out to be deeply
connected to the rst, forming the other half, called integral calculus.
7
CHAPTER 1. WHAT IS CALCULUS, REALLY?
Suppose that instead of knowing your position and wanting your speed, you know your speed
at every instant (say, from a very precise speedometer reading recorded continuously) and you
want to reconstruct how far you have traveled in total.
If your speed were constant, this would again be simple multiplication: distance = speed ×
time. But if speed is constantly changing, what do we do?
speed v(t)
time t
Figure 1.2: If we chop time into short intervals and pretend the speed is roughly constant on
each tiny interval, each thin rectangle's area (height = speed, width = time) approximates the
distance traveled in that interval. Adding up all the rectangles approximates the total distance.
As the rectangles get thinner and thinner, this approximation becomes exact this is the idea
behind the denite integral, developed fully in Part 5.
This is the second central question of calculus, sometimes called the area problem, because
geometrically, the total distance traveled corresponds exactly to the area trapped between the
speed graph v(t) and the horizontal axis. We will spend all of Part 5 making this idea completely
rigorous, but the core intuition slice into many thin pieces, approximate each piece simply,
add them all up, and let the pieces shrink to zero width is something you can already picture
clearly.
ÿ A stunning connection
For centuries, the tangent problem (instantaneous rate of change) and the area problem (total
accumulation) were studied by dierent people as seemingly unrelated puzzles. One of the
towering achievements of Newton and Leibniz in the late 1600s which we explore in depth
in the next chapter was realizing that these two problems are, in a precise sense, inverses
of each other. This single insight is called the Fundamental Theorem of Calculus, and it
is arguably the most important theorem in all of undergraduate mathematics. We will prove
it carefully in Part 5, but keep this preview in your mind as you read: dierentiation and
integration undo each other, much like squaring and taking a square root undo each other.
It is worth pausing to appreciate just how far-reaching these two questions instantaneous rate
of change, and total accumulation turn out to be. A few examples, all of which we will return
to with full mathematical detail later in the book:
Physics: Velocity is the rate of change of position; acceleration is the rate of change of
velocity. Newton's laws of motion are literally written in the language of derivatives.
8
1.6. COMMON MISCONCEPTIONS TO AVOID FROM THE START
Economics: Marginal cost (the cost of producing one more unit) is the derivative of the
total cost function. Total prot over a period can be found by integrating a rate of prot.
Medicine: The rate at which a drug concentration decays in the bloodstream is modeled
with derivatives; the total drug exposure over time is an integral.
Computer science and machine learning: Training a neural network is, at its core, an
optimization problem solved using derivatives (via an algorithm called gradient descent, which
we cover in Part 9).
Engineering: The bending of a beam, the ow of heat through a wall, the charge on a
capacitor all are described by dierential equations, the subject of Part 7.
Biology: Population growth models use derivatives to describe how a population's rate of
growth depends on its current size.
Car manufacturers need to know a vehicle's fuel consumption rate (liters per hour) as a
function of speed, in order to compute total fuel used (an integral) over a planned trip
with varying speeds say, city driving versus highway driving. Getting the instantaneous
rate right, and then correctly accumulating it, is literally why your car's estimated range
remaining display can update in real time as you drive.
Curves and graphs are extremely useful pictures of the ideas of calculus, but calculus is
fundamentally about relationships between changing quantities. You can do calculus with
no picture at all with tables of numbers, with formulas, with real-world data. The graph
is a visualization tool, not the subject itself.
It is tempting to imagine an innitely small number, like 0.000 . . . 0001 with innitely many
zeros. This is not how modern calculus is built (though it is close to how Leibniz originally
thought about it, and there are modern rigorous theories, like nonstandard analysis, that
make such innitesimals precise a topic we mention again in Part 3). In the standard
approach used throughout this book, we never use an actual innitely small number. We
use the limit idea: we consider ordinary, perfectly normal small numbers like h = 0.001, and
study the trend as h shrinks toward (but never touches) zero.
Every single calculation in calculus ultimately reduces to algebra at some point expanding
brackets, simplifying fractions, solving equations. Calculus extends what algebra can do, by
adding the idea of a limit on top of it. If your algebra is shaky, that will be the actual
bottleneck, not the new calculus ideas themselves. This is exactly why Part 0 of this book
exists: to make sure that bottleneck never appears.
9
CHAPTER 1. WHAT IS CALCULUS, REALLY?
Working mathematicians, physicists, and engineers rarely think in terms of long chains of algebra
when they reason about calculus day to day. Instead, they think in pictures and in trends:
When thinking about a derivative, an expert pictures zooming in on a graph until it looks like
a straight line, and thinks of the derivative as the slope of that line.
When thinking about an integral, an expert pictures slicing a region into thin strips and
adding them up, and thinks of the integral as the exact total once the strips are innitely thin.
When something changes fast, an expert thinks large derivative. When something changes
slowly or is momentarily at, an expert thinks derivative near zero.
Experts constantly sanity-check formulas using extreme or simple cases: What happens as
x → 0? As x → ∞? What if the function is constant? These sanity checks are a habit we will
build deliberately throughout this book, not just something to memorize as an afterthought.
This kind of visual, trend-based thinking is exactly what we are building toward the formal
notation and rigorous proofs are essential (and we will absolutely do them in full), but they are
the nal, precise expression of intuition you already have the capacity to build starting right
now.
A ball is dropped and its height above the ground, in meters, after t seconds, is given by
y(t) = 45 − 4.9t2 .
Estimate the instantaneous rate of change of height (i.e., the instantaneous vertical velocity)
at t=1 second, using the shrinking-interval method.
Step 1. Write the average rate of change formula over [1, 1 + h]:
y(1 + h) − y(1)
.
h
Step 2. Compute y(1) and y(1 + h) explicitly.
10
1.8. A FIRST WORKED EXAMPLE
Notice the crucial algebraic trick in Step 4: we needed to cancel a factor of h from top and
bottom before we could safely let h → 0. If we had tried to plug in h=0 before canceling,
0
we would have hit the meaningless
0 again. This simplify rst, then take the limit strategy
is the single most common technique in all of introductory dierential calculus, and we will
practice it extensively in Part 3 and Part 4.
y(t + h) − y(t)
* Formula deep-dive: the dierence quotient
h
What does it mean? It is the average rate of change of y over the short interval from t to
y(b)−y(a)
t+h identical in spirit to
b−a from earlier in this chapter, just written with b = t+h
anda = t, and with the interval width renamed h = b−a to emphasize that we intend to
shrink it.
Why does it work? Renaming the interval width as h turns two independent endpoints
a and b into a single base point t and a single step size h. This is purely a labeling
convenience, but it is an enormously useful one: it lets us treat t as xed and h as the only
thing changing, which is exactly what we need in order to ask what happens as h → 0?
Where does it come from? It comes from nothing more than substituting b = t + h into
the ordinary average-rate-of-change formula. There is no new mathematical content in the
renaming the power is entirely in what we do next : examine the trend as h → 0, which is
a genuinely new idea not present in ordinary algebra.
When should you use it? Use this exact pattern write the ratio in terms of a shrinking
h, expand algebraically, cancel the common factor of h, then let h→0 any time you are
asked to nd an instantaneous rate of change from rst principles (i.e., without using the
shortcut rules of Part 4). This is a skill examiners specically test, because it proves you
understand why the shortcut rules work, not just how to apply them.
C(10 + h) − C(10)
.
h
Step 2. Compute C(10):
= 250 + 6h + 0.1h2 .
11
CHAPTER 1. WHAT IS CALCULUS, REALLY?
C(10 + h) − C(10)
= 6 + 0.1h (h ̸= 0).
h
Step 5. Let h → 0:
6 + 0.1h −→ 6.
Conclusion. The marginal cost at q = 10 is $6 per unit meaning that, right around the
10th unit produced, each additional unit costs approximately $6 to make. Notice this is a
completely dierent real-world context (cost, not motion) governed by the exact same math-
ematics as the dropped-ball example. This is one of the most important lessons of calculus:
the same limiting process answers how fast questions in physics, economics, biology, and
beyond.
A quantity is given by y(t) = 3t. Using the same shrinking-interval method as in the worked
y(2 + h) − y(2)
example, compute and simplify. What do you notice about the result does
h
it even depend on h? Why might that make sense given that y(t) = 3t describes something
moving at a constant rate?
Hint: Substitute t=2 and t=2+h directly into y(t) = 3t and subtract.
Solution
y(2 + h) − y(2) 3h
= = 3.
h h
This does not depend on h at all the ratio is exactly 3 regardless of how large or small h is
(as long as h ̸= 0).This makes complete sense: y(t) = 3t describes a quantity changing at the
constant rate of 3 units per unit of time, so its average rate of change over any interval, long
or short, is always exactly 3. There is nothing to shrink toward the instantaneous rate
equals the average rate everywhere, because nothing is accelerating. This is a useful sanity
check: for straight-line (linear) relationships, calculus should simply recover the ordinary
slope you already know from algebra, and here it does exactly that.
Repeat the method of the worked example (dropped ball) for the function y(t) = t2 at the
general point t=a (instead of a specic number). That is, compute
y(a + h) − y(a)
h
fully simplied, and then state what value it approaches as h → 0.
12
1.8. A FIRST WORKED EXAMPLE
Solution
A tank is being lled with water. The volume of water in liters after t minutes is V (t) =
2t2 + 5t.
(a) Find the average rate at which the tank is lling between t = 3 and t = 3 + h, as a
simplied expression in h.
(b) Determine the instantaneous rate of lling at t=3 minutes (i.e., let h → 0).
(c) Explain, in one or two sentences, what this instantaneous rate physically represents.
Hint: This is structurally identical to the dropped-ball example just with a dierent
formula and two terms to expand instead of one.
Solution
(a)
V (3) = 2(9) + 15 = 18 + 15 = 33.
V (3+h) = 2(3+h)2 +5(3+h) = 2 9 + 6h + h2 +15+5h = 18+12h+2h2 +15+5h = 33+17h+2h2 .
For the marginal cost example above, C(q) = 200 + 4q + 0.1q 2 , nd the instantaneous rate of
change of cost at the general quantity q = a (not just a = 10). Then check that substituting
a = 10 into your general formula reproduces the answer $6 found in the worked example.
13
CHAPTER 1. WHAT IS CALCULUS, REALLY?
Hint: Follow exactly the same ve steps as the worked example, but keep a symbolic instead
of substituting 10.
Solution
Hint: You will need to expand (2 + h)3 and (2 + h)2 separately before combining. Use
(2 + h)3 = 8 + 12h + 6h2 + h3 .
Solution
P (2) = 23 − 2(2)2 + 5 = 8 − 8 + 5 = 5.
P (2 + h) = (2 + h)3 − 2(2 + h)2 + 5 = 8 + 12h + 6h2 + h3 − 2 4 + 4h + h2 + 5.
P (2 + h) − P (2)
= 4 + 4h + h2 −→ 4 as h → 0.
h
Conclusion: the instantaneous rate of growth at t = 2 hours is 4 thousand bacteria per
hour. Since this is positive, the population is increasing at that instant a negative result
would have indicated a shrinking population, and a result of exactly 0 would have indicated
a momentary pause (neither growing nor shrinking) at that instant.
14
1.8. A FIRST WORKED EXAMPLE
1 1
Hint: Start by writing f (a + h) − f (a) = − , then combine into a single fraction
a+h a
with common denominator a(a + h) before dividing by h.
Solution
1 1 a − (a + h) −h
f (a + h) − f (a) = − = = .
a+h a a(a + h) a(a + h)
Now divide by h:
f (a + h) − f (a) 1 −h −1
= · = (h ̸= 0).
h h a(a + h) a(a + h)
Notice the factor of h cancelled just as in every previous example it is simply hiding inside
the combined fraction rather than sitting outside a bracket. Now let h → 0:
−1 −1 −1
−→ = 2.
a(a + h) a·a a
1 1
Conclusion: the instantaneous rate of change of f (x) = x at x = a is − . Sanity
a2
1
check: this is always negative (for any a ̸= 0), which matches the graph of
x always sloping
downward on each branch. We will recognize this formula again in Part 4 as a special case
of the power rule applied to x−1 .
Quick FAQ
= Chapter Summary
Ordinary algebra excels at describing xed, unchanging quantities; calculus was invented
to describe quantities that change continuously.
The tangent problem asks: given a quantity's value at every instant, what is its in-
stantaneous rate of change at one particular instant? Attempting to answer this directly
0
leads to the undened expression
0.
The resolution is to compute an average rate of change over a shrinking interval of width
15
CHAPTER 1. WHAT IS CALCULUS, REALLY?
h, simplify algebraically (cancelling the common factor of h), and then examine the trend
as h → 0. This trend is called a limit, and it will be made fully rigorous in Part 3.
The area problem asks the reverse-avored question: given a rate at every instant,
what is the total accumulated quantity? This is answered by slicing into thin strips,
approximating, and summing leading to the denite integral, developed in Part 5.
The deep link between these two problems is the Fundamental Theorem of Calculus,
previewed here and proven in Part 5.
* Coming up next
In Chapter 2, we dig much deeper into the two questions you've just met the tangent
problem and the area problem working through several more worked examples of each
by hand, until the patterns become impossible to miss. Only after you've felt these ideas
work, in Chapter 3, do we pause to look carefully at innity and continuity, and nally,
in Chapter 4, tell the two-thousand-year story of where all of it actually came from
from ancient Greek attempts to compute areas, through the independent (and famously
contentious) development of calculus by Isaac Newton and Gottfried Wilhelm Leibniz, up to
how mathematicians nally made the limit idea fully rigorous almost 200 years later.
16
Chapter 2
Learning Objectives
Work through multiple concrete examples of the tangent problem using the shrinking-
interval (dierence quotient) method.
Work through multiple concrete examples of the area problem using shrinking rectangles.
Recognize the general pattern that will become the formal dierence quotient in Part 3
and Part 4.
Recognize the general pattern that will become the formal Riemann sum in Part 5.
Begin to see, through direct numerical evidence, why the two problems are connected.
# Prerequisites
Chapter 1. Comfort with expanding brackets such as (a + h)2 and (a + h)3 , and comfort
with basic summation of a short list of numbers.
In Chapter 1, we found the instantaneous rate of change of y(t) = t2 at a general point t=a by
computing
y(a + h) − y(a)
lim ,
h→0 h
y(a + h) − y(a)
and we found the answer 2a. The expression inside the limit, , is important
h
enough to deserve a name of its own.
f (a + h) − f (a)
h
is called the dierence quotient of f at x = a. It represents the average rate of change of
f over the interval from x = a to x = a + h. The derivative of f at x = a, once we dene
17
CHAPTER 2. THE TANGENT PROBLEM AND THE AREA PROBLEM, IN DEPTH
f (a + h) − f (a)
f ′ (a) = lim ,
h→0 h
whenever this limit exists.
We are previewing this formula now, informally, because working with it repeatedly by hand
on several dierent functions is the best possible preparation for the fully general theory in
Part 4. Let's build that experience.
f (a + h) − f (a)
* Formula deep-dive: f ′ (a) = lim
h→0 h
What does it mean? It packages the entire shrinking-interval idea from Chapter 1 into
a single symbol, f ′ (a), read f prime of a. Instead of describing the process in words every
time (compute the average rate over a shrinking interval and see what it approaches), we
now have one compact piece of notation for the resulting number.
Why does it work? It works because, for the well-behaved functions we meet in this book
(polynomials, and later trigonometric, exponential, and logarithmic functions), the dierence
quotient really does settle down to one specic number as h → 0, as we've now veried
directly in half a dozen examples. The limit notation limh→0 is simply honest bookkeeping:
it reminds us, every time we write f ′ (a), that a limiting process not a direct substitution
produced this number.
Where does it come from? It comes directly from combining two ideas we already have:
f (b)−f (a)
the average rate of change formula (this chapter and Chapter 1), rewritten with
b−a
b = a + h, together with the limit idea previewed in Chapter 1, Section 1.3.
When should you use it? This exact denition is what you fall back on whenever a problem
specically asks you to nd a derivative from rst principles or using the denition of the
derivative extremely common phrasing on exams. Whenever you see that phrasing, you
are being asked to write out these ve steps explicitly (set up the quotient, expand, simplify,
cancel h, take the limit), not to jump straight to a shortcut rule from Part 4, even once you
know those shortcuts.
(If this expansion is unfamiliar, Part 0, Section on Algebraic Expansions, derives it step by
step from (a + h)(a + h)(a + h); here we simply use the result.)
Step 2. Subtract f (a) = a3 :
18
2.1. REVISITING THE TANGENT PROBLEM, GENERALLY
f (a + h) − f (a)
= 3a2 + 3ah + h2 .
h
Step 5. Let h → 0. Both terms containing h vanish:
Conclusion: the instantaneous rate of change of f (x) = x3 at any point x=a is 3a2 .
x2 2x
x3 3x2
Do you notice a pattern relating the exponent to the result? Let's gather one more data
point before trusting the pattern.
(We derive this expansion in general, using Pascal's triangle and the binomial theorem, in
Part 0. For now, treat it as a given algebraic fact, veriable by direct multiplication of
(a + h)(a + h)(a + h)(a + h).)
Step 2. Subtract f (a) = a4 :
Step 4. Divide by h:
f (a + h) − f (a)
= 4a3 + 6a2 h + 4ah2 + h3 .
h
Step 5. Let h → 0. Every term except the rst contains at least one factor of h, so all
vanish:
4a3 + 6a2 h + 4ah2 + h3 −→ 4a3 .
Conclusion: the instantaneous rate of change of f (x) = x4 at x=a is 4a3 .
19
CHAPTER 2. THE TANGENT PROBLEM AND THE AREA PROBLEM, IN DEPTH
x2 2x
x3 3x2
x4 4x3
The pattern is now unmistakable: for f (x) = xn , the exponent n comes down in front as a
multiplying constant, and the new exponent is one less than before, n−1. This pattern is not
a coincidence, and in Part 4 we will prove completely rigorously, for any whole-number
exponent n (and later, for negative and fractional exponents too) that the instantaneous
rate of change of xn is nxn−1 . This single formula is called the power rule, and it is
one of the most-used results in all of calculus. Rather than telling you the rule up front
and asking you to trust it, we let you discover it yourself from direct evidence across three
separate cases which is close to how such patterns are typically discovered in mathematics
generally: notice a pattern in several worked cases, conjecture a general rule, then prove the
conjecture in general (which we do in Part 4).
Notice: the rate of change of 5x2 alone would be, by our earlier pattern (x
2 → 2x, scaled
by 5), equal to 5 × 2a = 10a. And the rate of change of 3x alone is just the constant 3
(matching Exercise 1.1, where y = 3t gave rate 3 everywhere). Adding these: 10a + 3
exactly what we computed for the combined function 5x2 + 3x. This suggests that the rate
of change of a sum of two functions is the sum of their individual rates of change, and that
constant multiples come along for the ride unchanged. Both of these observations will be
20
2.2. REVISITING THE AREA PROBLEM, GENERALLY
proven rigorously as the constant multiple rule and sum rule in Part 4 and they are
two of the most labor-saving results in the entire subject, because they mean we essentially
never need to redo a full dierence-quotient calculation for a complicated function built from
simpler pieces.
Now let's build the same kind of hands-on experience with the area problem. Recall from
Chapter 1: if we know a rate (like speed) at every instant, we can approximate total accumulation
(like distance) by slicing time into short intervals, treating the rate as constant on each tiny slice,
and adding up the resulting small pieces.
n
X
Total ≈ f (x∗i ) ∆x.
i=1
n
X
* Formula deep-dive: f (x∗i ) ∆x
i=1
Pn
What does it mean? The symbol i=1 (capital Greek sigma) means add up the following
expression as i runs through 1, 2, 3, . . . , n. So the whole formula means: for each strip i (from
∗
the rst strip to the n-th), evaluate the rate function at a chosen point xi inside that strip,
multiply by the strip's width ∆x to get that strip's rectangle area, and add all n of these
rectangle areas together.
Why does it work? Because on a very narrow strip, a continuously changing rate f (x)
doesn't have much room to change so treating it as constant (equal to its value at x∗i )
across that one narrow strip introduces only a small error. Multiplying a (nearly constant)
rate by a width to get an accumulated amount is just the elementary rate × time idea
from Chapter 1. Summing these small, nearly-correct pieces gives a total that is also nearly
correct and the total error shrinks as the strips get narrower, because each individual
strip's error shrinks and there being more strips doesn't outweigh that shrinkage (we make
this trade-o precise in Part 5).
Where does it come from? It comes from directly generalizing the single-rectangle idea
(rate × time = area) to many rectangles side by side, exactly as pictured in Chapter 1's
speed-graph diagram.
When should you use it? Use this rectangle-sum approach whenever you need a numer-
ical estimate of a total accumulation and either don't have, or don't yet know, an exact
antiderivative-based method (which we develop in Part 5). It's also exactly the method
used inside computers and calculators for numerical integration when an exact formula is
unavailable or impractical a topic we return to in Part 9.
21
CHAPTER 2. THE TANGENT PROBLEM AND THE AREA PROBLEM, IN DEPTH
Suppose a tap lls a bucket at a perfectly constant rate of f (x) = 4 liters per minute, for
x from 0 6 minutes. Use rectangles to nd the total volume, rst with n = 3 rectangles,
to
then with n = 6 rectangles, and compare to the answer you'd get from simple multiplication.
Simple multiplication check: rate × time = 4 × 6 = 24 liters.
With n = 3 rectangles: each rectangle has width ∆x = 6−0 3 = 2 minutes. Since the rate is
constant at 4 everywhere, every rectangle has height 4. Total area = 3 × (4 × 2) = 3 × 8 = 24
liters.
With n = 6 rectangles: each rectangle has width ∆x = 1 minute, height 4. Total area
= 6 × (4 × 1) = 24 liters.
Conclusion: for a constant rate, any number of rectangles gives exactly the right answer,
24 liters, with no approximation error at all which makes complete sense, since there is
nothing changing for the rectangles to approximate imperfectly.
Step 3. Sum the rectangle areas (height × width, and ∆x = 1 so this is just the sum of the
heights):
2 + 4 + 6 + 8 = 20.
Step 4. Compare to the exact geometric answer. The region under f (x) = 2x from 0 to 4 is
literally a triangle with base 4 and height f (4) = 8 (we can check this exactly using ordinary
geometry here, since f (x) = 2x is a straight line):
1 1
Area of triangle = × base × height = (4)(8) = 16.
2 2
Observation: Our rectangle approximation gave 20, but the true area is 16 an overesti-
mate of 4. This happened because using the right endpoint of an increasing function always
overshoots on each strip (each rectangle's top-right corner pokes above the actual line). Let's
see what happens with more, thinner rectangles.
36 × 0.5 = 18.
22
2.2. REVISITING THE AREA PROBLEM, GENERALLY
This is much closer to the true area of 16 than our earlier estimate of 20 was. The overestimate
has shrunk from 4 down to 2.
As we use more and more, thinner and thinner rectangles, our approximation (20 → 18 → · · · )
16. If we continued this process
is steadily closing in on the true area of n = 16, then
n = 100, then n = 10,000 the approximation would keep approaching 16 ever more closely,
and in the limit as n → ∞, it would equal 16 exactly. This is precisely the same kind of
limiting process we used for the tangent problem, just applied to a sum of many terms instead
of a single ratio. In Part 5, we will develop summation notation and limit techniques that let
us compute this exact limiting area without manually adding up thousands of rectangles
arriving at general formulas, just as the power rule gave us a general formula for derivatives
without needing to redo the dierence-quotient computation every time.
Using the same method as the cubic and quadratic worked examples, compute
f (a + h) − f (a)
lim for f (x) = 7x2 .
h→0 h
Hint: Use the pattern from the tip box: if x2 → 2x, what should 7x2 give you, based on
the constant-multiple observation? Then verify it directly with the full dierence-quotient
computation.
23
CHAPTER 2. THE TANGENT PROBLEM AND THE AREA PROBLEM, IN DEPTH
Solution
Direct computation:
Solution
Using the dierence-quotient method (rst principles), nd the instantaneous rate of change
of f (x) = x3 − 4x at the general pointx = a. Then use your general formula to nd the
instantaneous rate of change at a = −1, a = 0, and a = 2, and comment on the sign of each
answer.
Hint: Expand f (a + h) = (a + h)3 − 4(a + h) using the cubic expansion from earlier in this
chapter, then combine like terms before factoring out h.
Solution
f (a + h) − f (a)
= 3a2 + 3ah + h2 − 4 −→ 3a2 − 4 as h → 0.
h
General formula: instantaneous rate of change at x = a is 3a2 − 4.
At a = −1: 3(1) − 4 = −1 (decreasing). At a = 0: 3(0) − 4 = −4 (decreasing, and more
steeply so). At a = 2: 3(4) − 4 = 8 (increasing). The function switches from decreasing
24
2.2. REVISITING THE AREA PROBLEM, GENERALLY
Approximate the area under f (x) = x2 from x=0 to x = 3 (same function and interval
as Exercise 3.2) using the midpoint rule instead: use n = 3 rectangles, but evaluate f
at the midpoint of each strip rather than an endpoint. Compare your result to both the
left/right-endpoint estimates and to the true value of 9.
Hint: The three strips are [0, 1], [1, 2], [2, 3], with midpoints 0.5, 1.5, 2.5.
Solution
Midpoints: x = 0.5, 1.5, 2.5. Heights: f (0.5) = 0.25, f (1.5) = 2.25, f (2.5) = 6.25. Sum
= 8.75. Multiply by ∆x = 1: total = 8.75.
Compare: left/right endpoint methods (from a similar calculation to Exercise 3.2, adjusted
for n = 3) give 5 (left) and 14 (right), while the midpoint estimate gives 8.75 much closer
to the true value of 9 than either one-sided estimate. This illustrates a general and important
fact, proven carefully in Part 5: the midpoint rule is typically far more accurate than either
the left- or right-endpoint rule for the same number of rectangles, because overestimates on
one half of each strip tend to cancel with underestimates on the other half.
A signal processing circuit has an output voltage given by V (t) = 3t2 − 2t + 1 volts, where t
is measured in milliseconds.
(a) Using rst principles, nd a general formula for the instantaneous rate of change of
voltage with respect to time, at t = a.
(b) Find the specic instant(s) at which the voltage is momentarily not changing (i.e., where
the instantaneous rate of change is exactly zero).
(c) Explain what it would mean, physically, for the rate of change to be zero at an instant,
without it meaning the voltage itself is zero.
Hint: For part (b), set your answer from part (a) equal to 0 and solve the resulting linear
equation for a.
Solution
(a)
V (a + h) = 3(a + h)2 − 2(a + h) + 1 = 3a2 + 6ah + 3h2 − 2a − 2h + 1.
V (a) = 3a2 − 2a + 1.
V (a + h) − V (a) = 6ah + 3h2 − 2h = h(6a + 3h − 2).
V (a + h) − V (a)
= 6a + 3h − 2 −→ 6a − 2 as h → 0.
h
General formula: 6a − 2.
25
CHAPTER 2. THE TANGENT PROBLEM AND THE AREA PROBLEM, IN DEPTH
1
(b) Set 6a − 2 = 0, giving a= 3 ms.
(c) A zero instantaneous rate of change means that, at that exact instant, the voltage is
momentarily neither increasing nor decreasing it is at a turning point (in this case, since
the function is an upward-opening parabola, this instant is a momentary minimum voltage).
The voltage itself, V (1/3) = 3(1/9) − 2/3 + 1 = 1/3 − 2/3 + 1 = 2/3 volts, is not zero at all
rate of change is zero and the quantity itself is zero are entirely dierent statements, a
distinction worth keeping sharply in mind throughout Part 4's work on optimization.
Let f (x) = x2 on the interval [0, 3], as in Exercise 3.2. Using n rectangles of equal width
n
X n(n + 1)(2n + 1)
with right endpoints, and the summation formula i2 = (which we
6
i=1
prove by induction in Part 0 and use extensively in Part 5), nd a general expression for the
right-endpoint approximation as a function of n, and then determine what this expression
approaches as n → ∞. Conrm this matches the exact area of 9 mentioned earlier in this
chapter.
3i
Hint: With n strips over [0, 3], ∆x = 3/n and the i-th right endpoint is xi = n. The
Pn Pn 3i 2 3
Riemann sum is i=1 f (xi )∆x = i=1 n · n.
Solution
n 2 n n
X 3i 3 X 9i2 3 27 X 2 27 n(n + 1)(2n + 1)
· = 2
· = 3
i = 3· .
n n n n n n 6
i=1 i=1 i=1
Simplify:
27 n(n + 1)(2n + 1) 9(n + 1)(2n + 1)
= 3
= .
6n 2n2
Expand the numerator: (n + 1)(2n + 1) = 2n2 + 3n + 1, so the expression becomes
9 2n2 + 3n + 1
18n2 + 27n + 9 27 9
2
= 2
=9+ + 2.
2n 2n 2n 2n
27 9
As n → ∞, both
2n and 2n2 shrink toward 0, leaving exactly 9 conrming, via a fully
general symbolic computation (not just numerical approximation), that the exact area under
x2 from 0 to 3 is 9. This is precisely the kind of calculation the Fundamental Theorem of
Calculus (Part 5) will let us shortcut entirely, but seeing it done directly, once, is invaluable
for understanding what that theorem is actually saving us from.
= Chapter Summary
f (a + h) − f (a)
The dierence quotient generalizes the shrinking-interval method from
h
Chapter 1 to any function f .
Working through x2 → 2x, x3 → 3x2 , and 5x2 + 3x → 10a + 3 reveals two patterns a
power pattern and an additive/scaling pattern that will become the power rule and
the sum and constant multiple rules in Part 4.
The area problem is tackled by summing many thin rectangles; using more, thinner rect-
26
2.2. REVISITING THE AREA PROBLEM, GENERALLY
Both problems rely on the same underlying idea: compute something for a nite, nonzero
h (or ∆x), simplify algebraically, and then examine the limiting trend.
These worked patterns are strong motivation for building a fully rigorous theory of lim-
its next, in Part 3 which is exactly why Chapter 3 pauses to discuss innity and
innitesimals before we dive into that formal theory.
27
CHAPTER 2. THE TANGENT PROBLEM AND THE AREA PROBLEM, IN DEPTH
28
Chapter 3
Learning Objectives
Explain informally why there is no smallest positive real number, and what that has to
do with calculus.
Feel comfortable with the idea that a sum of innitely many things can add up to a nite
number.
# Prerequisites
y(1.2+h)−y(1.2)
Chapters 12. In particular, the numerical tables from Chapter 1 (values of
h
as h shrinks) and the rectangle approximations from Chapter 2.
The word innity gets used loosely in everyday speech I've told you a million times, there
are innite possibilities. In calculus, the symbol ∞ is used in two related but distinct ways, and
it is worth being crystal clear about both before going further, because confusing them is one of
the most common sources of error for beginners.
29
CHAPTER 3. INFINITY, INFINITESIMALS, AND CONTINUOUS CHANGE
(grows without any upper bound), never as a specic numerical value to be manipulated
directly with +, −, ×, ÷.
The rst, more common use of ∞ in this book is to describe a variable that grows larger and
larger, without any upper limit. When we write
x → ∞,
we mean: imagine x taking on larger and larger values 10, then 100, then 10,000, then
1,000,000, and so on, without ever stopping at some nal largest value. We are describing an
unending process, a direction of travel, not a destination that is ever actually reached.
1
This is exactly how we will use ∞ when we write things like lim =0 in Part 3: we mean
x→∞ x
1
that as x x gets closer and closer to 0, without x
takes on larger and larger values, the quantity
1
ever equaling some nal innite number, and indeed without
x ever actually equaling 0 either,
for any nite x.
The second use, seen already in Chapter 2's rectangle approximations, is to describe a process of
dividing something into more and more pieces, without bound: n→∞ meaning imagine using
more and more rectangles 10, then 100, then 10,000, and so on. Here again, ∞ describes
an unending process of renement, not a specic innite number of pieces that is ever actually
reached or held in your hand.
Notice that both uses of ∞ in this book describe a process without end, never a completed,
actual innite object. This distinction between a potential innity (an unending process)
and an actual innity (a completed innite totality) goes back to Aristotle, and it is
exactly the distinction that let Cauchy and Weierstrass (Chapter 4) make calculus fully
rigorous: the entire theory of limits is built using only potential innity (unending processes
we can describe with ordinary nite numbers at every stage), which sidesteps a great deal of
philosophical diculty.
Here is a simple but genuinely important question: what is the smallest positive real number?
Try to name one. Suppose you guess 0.0001. But 0.00001 is smaller and still positive. Suppose
you guess 0.00001 0.000001 is smaller still. No matter what positive number you
instead but
name, dividing it by 10 (or by 2, or by anything greater than 1) produces an even smaller positive
number.
There is no smallest positive real number. That is, for every positive real number ε, there
exists another positive real number smaller than ε (for instance, ε/2).
This might seem like a strange thing to dwell on, but it is precisely why the limit process
in Chapter 1 works the way it does. When we said let h shrink toward 0, we never mean h
reaches some ultimate, nal, smallest possible nonzero value. There is no such value. Instead,
30
3.3. DISCRETE CHANGE VERSUS CONTINUOUS CHANGE
we mean: for any tiny positive number you might name as a target closeness, we can nd values
of h that get the dierence quotient at least that close to the limiting value. This is the intuitive
seed of the formal epsilon-delta denition awaiting us in Part 3.
Suppose someone challenges you: I bet you can't make the dierence quotient of f (x) = x2
at a=1 get within 0.001 of the true limiting value 2. Let's meet that challenge directly, to
build a feel for what arbitrarily close really means in practice.
Step 1. Recall from Exercise 1.2 that the dierence quotient for f (x) = x2 at a = 1 simplies
to 2+h a = 1 in the general result 2a + h).
(using
Step 2. We want |(2 + h) − 2| < 0.001, i.e., |h| < 0.001.
Step 3. So any nonzero h with |h| < 0.001 for instance h = 0.0005 works: the
dierence quotient 2 + 0.0005 = 2.0005 is indeed within 0.001 of 2.
Step 4. Now suppose someone raises the bar: within 0.000001 instead. The same reasoning
gives: any h with |h| < 0.000001 works, e.g., h = 0.0000005.
Conclusion. No matter how small a target closeness we're challenged with, we can always
name a value of h (indeed, a whole range of values of h) that achieves it because there
is no smallest positive real number to run out of room against. This no matter how small
the challenge, we can always respond structure is exactly what the formal epsilon-delta
denition of a limit will state with complete precision in Part 3; we have just carried it out
by hand for one specic function.
Yes and this is an excellent, practical observation. Computers represent real numbers
using a nite number of binary digits (a scheme called oating-point arithmetic), so there
genuinely is a smallest positive number a given computer can represent (for standard double-
precision oating point, roughly 10−308 ). This is a limitation of digital representation, not a
mathematical fact about the real numbers themselves. It has real practical consequences for
numerical methods (which we cover in Part 9) for instance, some numerical dierentiation
and integration algorithms can actually become less accurate if you chooseh absurdly small,
because of rounding error, even though the mathematical theory says smaller h should always
be better. Knowing where the pure mathematics ends and the practical computing begins is
a genuinely useful skill for engineers and programmers.
Another distinction worth making crisp before we build formal limit theory: the dierence be-
tween something that changes in discrete jumps and something that changes continuously.
The number of students enrolled at a university changes discretely: it goes from say 20,000 to
20,001 when one more student enrolls; it does not pass through 20,000.5 students at some point
along the way that value simply has no meaning. By contrast, the height of a falling ball
(Chapter 1's example) changes continuously: as time increases smoothly, the height decreases
31
CHAPTER 3. INFINITY, INFINITESIMALS, AND CONTINUOUS CHANGE
smoothly, passing through every intermediate height value with no jumps or gaps.
students height
t t
discrete: step function continuous: smooth curve
One more idea worth planting now, ahead of Part 8 (Innite Series), because it will quietly
support your intuition throughout the whole book: it is entirely possible for innitely many
positive numbers to add up to a perfectly ordinary, nite total.
The ancient Greek philosopher Zeno of Elea posed a famous puzzle: to walk from your front
door to the end of your street, you must rst cross half the distance. Then you must cross
half of the remaining distance. Then half of what remains after that, and so on, forever.
Since this process never technically nishes listing steps, Zeno argued that motion itself
might be logically impossible, or at least paradoxical.
Modern calculus resolves this cleanly. If the street is, say, 10 meters long, the distances
covered at each stage are
The running total is creeping steadily toward 10, and never overshoots it. In fact, it can be
shown (we will do so properly with the theory of geometric series in Part 8) that this innite
sum equals exactly 10 not approximately 10, but precisely, exactly 10. Innitely many
positive terms, added together, produce a completely ordinary nite number. There is no
contradiction with walking down the street at all; Zeno's puzzle arises only if you assume,
incorrectly, that innitely many steps must take innitely long to enumerate one-by-one in
real time, or that an innite sum can never complete. Calculus shows precisely how and
why it does.
32
3.4. A SUM OF INFINITELY MANY THINGS CAN BE FINITE
What does it mean? If you start with a quantity d and keep adding half of whatever
remains, the running total after n steps is
1 1 1 1 1
d + + + ··· + n =d 1− n .
2 4 8 2 2
This compact formula (which we will prove carefully using the geometric series formula in
d
Part 8) says the total after n steps is always d minus a leftover piece
2n .
Why does it work? Each new term is exactly half the previous leftover distance, so the
d
still remaining distance after n steps is always
2n this is easy to check directly: after
d d
step 1, half of d remains ( ); after step 2, half of that remains ( ); and so on. The total
2 4
walked is simply d minus whatever still remains.
Where does it come from? It comes from recognizing the walked distance as d−
d
(remaining distance), and the remaining distance after n steps as
2n , by direct repeated
halving.
When should you use it? Whenever you see a quantity being repeatedly halved (or
repeatedly scaled by any xed fraction), this pattern total after n steps equals (starting
amount) minus (starting amount times the scaling factor to the n-th power) is the key to
nding both a formula for partial progress and, by letting n → ∞, the nal total. This is
precisely the geometric series pattern that reappears constantly in Part 8, in loan repayment
and compound interest calculations, and in computer science (e.g., analyzing algorithms that
repeatedly halve a problem size, like binary search).
Let's test the same idea with a dierent fraction. Suppose instead of halving, each new
1
distance is
3 of the remaining distance, starting from a street of length 9 meters. Find the
rst four terms, compute the running total at each stage, and predict the nal total.
Step 1: nd the terms. First term: 13 of 9 is 3. Remaining: 6. Second term:
1
3 of 6 is 2.
1 4 8 1 8 8
Remaining: 4. Third term: 4
3 of is . Remaining:
3 3 . Fourth term: 3 of 3 is 9 .
Step 2: running totals.
4 19 19 8 57+8 65
3, 3 + 2 = 5, 5+ 3 = 3 ≈ 6.33, 3 + 9 = 9 = 9 ≈ 7.22.
Step 3: predict the nal total. The running total is climbing toward 9 (the full street
length) exactly as in the halving case, just at a dierent pace. Indeed, this matches
the same reasoning as the deep-dive box: the remaining distance after each step is being
2 1 2 1
multiplied by
3 (since 3 is walked and 3 remains) rather than 2 , so the sum approaches 9
more slowly, but it still approaches 9 exactly, for the same underlying reason.
Conclusion. No matter what xed fraction of the remaining distance you choose to walk
at each step (as long as it's a xed positive fraction less than the whole), the innite process
of walk part of what remains, forever always totals up to exactly the original full distance.
This is strong evidence for a general pattern that Part 8 will pin down completely: a geometric
2 1
series with ratio r (here r= 3 or r= 2 ) satisfying −1 < r < 1 always converges to a specic
nite sum.
33
CHAPTER 3. INFINITY, INFINITESIMALS, AND CONTINUOUS CHANGE
This same phenomenon innitely many pieces adding up to something perfectly nite
is exactly what happens in the rectangle approximations of Chapter 2. As n → ∞, we
are conceptually summing innitely many innitely-thin rectangles, and yet the total area
is a completely ordinary, nite number (we saw 16 for the triangle example). Keep Zeno's
resolved paradox in your back pocket any time an innite process in this book feels intuitively
troubling chances are, a very similar resolution applies.
True or false, and explain briey: As h→0 in the dierence quotient, h eventually equals
some extremely tiny but nonzero xed number, and that is the number we plug in.
Hint: Revisit the no smallest positive real number theorem in this chapter.
Solution
False. There is no smallest positive real number, so there is no nal, tiniest h to plug in.
The phrase h → 0 describes an unending process (choosing smaller and smaller positive
values of h) together with a claim about the trend of the dierence quotient throughout that
process, not a single substitution of one particular tiny value of h.
Classify each of the following as an example of discrete change or continuous change, and
briey justify your answer: (a) the balance of a bank account that only posts interest once
per month; (b) the position of the second hand on an old-fashioned mechanical clock with a
smoothly sweeping (non-ticking) motion; (c) the number of goals scored in a football match,
as a function of elapsed match time.
Hint: Ask whether the quantity can meaningfully take on every value in between two nearby
moments, or only certain separated values.
Solution
(a) Discrete. The balance jumps once per month when interest posts; there is no meaningful
in-between balance representing partial interest at, say, day 15 of the month (assuming
interest is only credited monthly).
(b) Continuous. A smoothly sweeping second hand passes through every angular position
between two nearby moments in time, with no jumps.
(c) Discrete. The goal count jumps by whole numbers at discrete moments (when a goal is
scored) and is constant in between; there is no such thing as 1.5 goals at some intermediate
time.
Repeat the making arbitrarily close concrete worked example above, but for f (x) = x2 at
a=3 instead of a = 1, and a target closeness of 0.0002. That is: nd a bound on h that
guarantees the dierence quotient is within 0.0002 of the true limiting value.
Hint: First nd the general dierence-quotient formula at a=3 using 2a + h with a = 3,
then set up an inequality just as in the worked example.
34
3.4. A SUM OF INFINITELY MANY THINGS CAN BE FINITE
Solution
Hint: Multiply each height by 35 to get the next one; for the total, think of this as similar in
3
spirit to the Zeno-type sum, though here the terms shrink by a factor of
5 each time rather
1
than by
2.
Solution
3 3
Heights: 2, 2 × 5 = 1.2, 1.2 × 5 = 0.72, 0.72 × 35 = 0.432. Since each height is
3
5 of
the previous one (a xed ratio less than 1 in absolute value), the same reasoning as the
geometric-sum deep-dive box applies: the innite sum of all bounce heights approaches a
specic nite total rather than growing without bound we will compute this total exactly,
using the geometric series sum formula, in Part 8. (For reference, that formula will give total
2 2
= 1−3/5 = 2/5 =5 meters of total upward bounce height, but conrming this rigorously is a
Part 8 topic.)
Explain why the following statement is false, using ideas from this chapter: Since a com-
puter's smallest representable positive number is about 10−308 , mathematically there must
also be a smallest positive real number, just one that's even smaller than 10−308 .
Hint: Revisit the FAQ box distinguishing oating-point computer arithmetic from the math-
ematical real numbers.
Solution
This statement confuses a limitation of digital representation with a fact about the mathe-
matical real numbers. A computer's smallest representable positive number is a consequence
of using a xed, nite number of binary digits to store numbers it's an engineering con-
straint, not a mathematical one. The real numbers themselves, as a mathematical structure,
have no smallest positive element at all (proven in the Theorem box: for any positive ε, ε/2
is smaller and still positive). There is no even smaller true smallest number waiting to be
discovered the absence of a smallest positive real number is a proven mathematical fact,
not a measurement limitation waiting for better technology.
35
CHAPTER 3. INFINITY, INFINITESIMALS, AND CONTINUOUS CHANGE
to be approaching a xed nite total, or does it appear to keep growing without bound?
(You are not expected to prove your answer rigorously this is a preview of a subtlety fully
resolved in Part 8.)
Hint: Compute each sum as a decimal to make the growth pattern easy to see: n = 1 : 1.
n = 2: 1.5. n = 4: 1 + 0.5 + 0.333 + 0.25 ≈ 2.083. n = 8: add 15 , 61 , 17 , 18 to the n = 4 total.
Solution
= Chapter Summary
Discrete quantities jump between separated values; continuous quantities pass smoothly
through every intermediate value. Calculus, in its standard form, is built for continuous
quantities.
Innitely many positive quantities can sum to a perfectly ordinary nite total resolving
Zeno's paradox, and previewing both the rectangle-sum idea of Part 5 and the innite
series of Part 8.
* Coming up next
You've now built a strong intuitive and computational feel for calculus, entirely through
worked examples and direct numerical evidence the two central questions (the tangent
problem and the area problem), and the subtleties of innity and continuity that any rigorous
theory of limits will eventually need to handle carefully. Before we move on to build that
rigorous theory, Chapter 4 closes out this Part by telling the two-thousand-year story of
where all of these ideas actually came from a story that will make much more sense now
that you've felt the ideas rsthand, rather than only having been told about them.
36
Chapter 4
Learning Objectives
Describe the ancient Greek origins of the area problem through the method of exhaus-
tion.
Explain what Newton and Leibniz each contributed, and why a bitter priority dispute
broke out between their followers.
Understand why calculus was used successfully for over a century before it was made fully
logically rigorous.
Name the key gures who eventually placed calculus on rigorous foundations in the 1800s.
Appreciate why the notation we use today is a deliberate historical choice, not an arbitrary
one.
# Prerequisites
Long before anyone used the word calculus, ancient mathematicians were already wrestling
with one of its two central questions: the area problem. The area of a rectangle, a triangle, or
any shape made of straight edges was well understood in antiquity these are all computable
with pure geometry and algebra. But what about a shape with a curved boundary, like a circle,
or the region under a parabola?
The Greek mathematician Eudoxus of Cnidus (4th century BCE) developed an approach
called the method of exhaustion: approximate a curved region using a sequence of polygons
whose areas you can compute, and let the polygons t the curved region more and more snugly.
Archimedes (3rd century BCE) used this method with extraordinary skill for example,
calculating the area of a circle, and famously nding the area enclosed by a parabola and a
straight line, by inscribing and circumscribing sequences of triangles and carefully bounding the
true area between the two sequences.
37
CHAPTER 4. A BRIEF HISTORY OF CALCULUS
ÿ Archimedes' insight
Archimedes' method for the area under a parabola is, in hindsight, breathtakingly close to a
full-blown Riemann sum (the technique we build carefully in Part 5). He eectively summed
innitely many shrinking triangular pieces and showed the total converges to an exact value
nearly two thousand years before the formal theory of limits and series existed to justify the
technique rigorously. What Archimedes lacked was not the core idea, but the general-purpose
algebraic machinery to apply this technique quickly to any curve, rather than laboriously re-
deriving it by hand for each new shape.
Let's actually carry out a simplied version of Eudoxus and Archimedes's method ourselves,
to feel how it works. Consider a circle of radius r = 1, so its true area is πr2 = π ≈ 3.14159.
We will approximate this area using regular polygons inscribed inside the circle.
Step 1: a square (4 sides). A square inscribed in a unit circle has diagonal 2 (the circle's
2
√ √
diameter), so each side has length √
2
= 2. Its area is ( 2)2 /2 × 2 = 2 (a square with
d2 22
diagonal d has area
2 , so area = 2 = 2). This is a rather poor approximation: 2 versus the
true value 3.14159 . . ..
Step 2: a regular hexagon (6 sides). A regular hexagon inscribed in a unit circle can
be split into
√
6 equilateral triangles, each with side length 1 (equal to the radius) and area
3
4 ≈ 0.433. 6 × 0.433 ≈ 2.598. Already noticeably closer to π .
Total area:
Step 3: a regular 12-gon. Using the formula for the area of a regular n-gon inscribed in
1 360◦ 1 ◦
a unit circle, An = n sin , we get A12 = (12) sin(30 ) = 6 × 0.5 = 3.
2 n 2
Step 4: a regular 96-gon (Archimedes' actual choice). A96 = 12 (96) sin(3.75◦ ) ≈
48 × 0.06540 ≈ 3.1393.
Conclusion. As the number of sides n grows, the polygon hugs the circle more and more
closely, and its area approaches π exactly the same shrink the gap logic as the dierence
quotient in Chapter 1, just applied to area instead of instantaneous rate. Archimedes actually
carried this all the way to a 96-sided polygon by hand, using both inscribed and circumscribed
polygons to trap the true area of the circle between two increasingly tight bounds, ultimately
proving that 3 10 1
71 < π < 3 7 a genuinely remarkable feat of hand computation with no
algebraic notation, no trigonometric tables, and no calculators.
For close to two thousand years after Archimedes, progress on the area problem was slow.
Mathematicians in the Islamic Golden Age, in India, and later in Europe made further partial ad-
vances for instance, the Indian mathematician Madhava of Sangamagrama (14th century)
and the later Kerala school developed innite series for trigonometric functions, foreshadowing
ideas we will meet in Part 8. But a truly general method one that could handle the tangent
problem and the area problem for a huge variety of curves, quickly and mechanically did not
yet exist.
38
4.2. THE SEVENTEENTH CENTURY: SETTING THE STAGE
By the early-to-mid 1600s, several mathematicians had made important partial progress that set
the stage for calculus:
René Descartes and Pierre de Fermat developed analytic geometry the idea of de-
scribing curves with algebraic equations using coordinates (x, y). Without this, none of the
algebraic machinery of calculus would even have a language to operate in.
Fermat also developed an early method for nding the maximum or minimum points of a
curve, and for nding tangent lines startlingly close to the derivative techniques of Part 4,
though without a general theory behind them.
John Wallis extended these techniques and introduced notation and ideas that Newton
would directly build upon.
The pieces were all on the table. What was missing was a single, unied framework connecting
the tangent problem and the area problem, together with a compact, reusable notation and set
of rules. That unication is what earns Newton and Leibniz the title of inventors of calculus,
even though, as this section shows, they were building on a full century (or more) of accumulated
groundwork.
Isaac Newton (16421727), working largely in England, developed his version of calculus
which he called the method of uxions primarily during 16651666, while Cambridge
University was closed due to an outbreak of plague and Newton had retreated to his family
home. This period was so extraordinarily productive across multiple elds of Newton's work
that it is often called his annus mirabilis (year of wonders).
Newton's approach centered on the idea of a uxion the instantaneous rate of change of
a owing quantity (a uent ) which corresponds closely to what we now call a derivative.
Newton was deeply motivated by physics: he needed calculus as a tool to describe motion, and
in particular to derive and work with his laws of motion and law of universal gravitation. For
Newton, calculus was never an abstract exercise; it was inseparable from mechanics.
Remarkably, Newton did not publish his method of uxions for years after developing it. He
circulated some results privately among colleagues, and it was not until Philosophiæ Naturalis
Principia Mathematica (1687) his monumental work on mechanics and gravitation
that calculus-based reasoning appeared in print, though even there it was often translated
back into classical geometric arguments, which Newton considered more rigorous and more
acceptable to contemporary mathematicians. His fuller, more explicit treatments of uxions
were published even later (some posthumously). This delay would later fuel the bitter dispute
described below.
39
CHAPTER 4. A BRIEF HISTORY OF CALCULUS
Latin summa ), to represent the accumulation of innitely many innitesimally thin pieces
precisely the area-problem intuition from Chapter 1.
dy
Throughout this book you will see derivatives written as (Leibniz notation) and sometimes
′
dx
as y or f ′ (x)
Z (a notation later popularized by Joseph-Louis Lagrange). You will see integrals
written as f (x) dx. These are not arbitrary symbols to memorize they are historically
the result of Leibniz's innitesimal viewpoint, and they remain in universal use because they
make many calculations (especially the chain rule and substitution) almost automatic once
you get used to manipulating them. We introduce each piece of this notation carefully and
exactly once, in Parts 4 and 5, and use it consistently afterward.
Because Newton developed his ideas rst (mid-1660s) but Leibniz published rst (1684), and
because both men worked largely independently with only limited, ambiguous communication
between them, a ferocious dispute erupted fanned enthusiastically by supporters on both sides
over who deserved credit for inventing calculus, and whether Leibniz had seen unpublished
Newton material and used it without proper credit (an accusation of plagiarism).
The Royal Society of London (of which Newton was, tellingly, president at the time) con-
ducted an investigation in 1712 and concluded in Newton's favor a verdict now viewed by
historians as compromised by Newton's own inuence over the process. The dispute consumed
enormous energy on both sides for decades, and had a genuinely damaging side eect: British
mathematicians, loyal to Newton, continued using Newton's uxion notation and largely isolated
themselves from continental European mathematics for roughly a century, while European math-
ematicians adopted Leibniz's superior notation and made rapid progress. This is a large part of
why Leibniz's notation, not Newton's, is the notation used throughout the world today.
The modern historical consensus, well-supported by manuscript evidence, is that Newton and
Leibniz developed their versions of calculus independently of one another, at dierent times
(Newton earlier, in the 1660s; Leibniz later, in the 1670s, but published rst, in the 1680s),
40
4.6. A CENTURY OF POWERFUL BUT SHAKY FOUNDATIONS
using dierent core ideas (uxions and rates of ow for Newton; innitesimal dierences for
Leibniz) and dierent notation. Both deserve genuine, independent credit this is why
calculus is universally taught today as a joint achievement of Newton and Leibniz, rather
than crediting one man alone.
For roughly 150 years after Newton and Leibniz, calculus was used with enormous success to
describe planetary motion, uid ow, heat, elasticity, sound, and much more despite the fact
that its logical foundations were, frankly, shaky. Mathematicians (and physicists) manipulated
innitesimals like dx as though they were tiny nonzero numbers, sometimes treating them as zero
and sometimes as nonzero within the very same calculation, depending on what was convenient.
In 1734, the Irish philosopher and Bishop George Berkeley published a scathing critique of
calculus titled The Analyst, pointing out precisely this inconsistency. He memorably mocked
innitesimals as the ghosts of departed quantities quantities that were supposedly small
enough to be discarded in one step of a calculation, yet apparently large enough to divide by
in another step. This was a genuinely serious logical objection, and mathematicians of the
era did not have a fully satisfying rebuttal for almost another hundred years.
Despite this shaky foundation, calculus continued to produce spectacularly correct, veriable
results throughout the 1700s, in the hands of mathematicians like Leonhard Euler, Joseph-
Louis Lagrange, and Pierre-Simon Laplace. This is one of the strange and wonderful facts
about the history of mathematics: an extremely powerful and predictively accurate tool was in
widespread use well before anyone had rigorously justied why it worked.
The person most responsible for beginning to x these logical gaps was the French mathematician
Augustin-Louis Cauchy, in the 1820s. Cauchy was the rst to dene the derivative and
the integral carefully in terms of limits, rather than innitesimals precisely the approach
previewed informally in Chapter 1 of this book, and the approach we develop with full rigor in
Part 3.
The nal, fully rigorous version of the limit concept the famous epsilon-delta denition
that we will study in detail in Part 3 is credited primarily to the German mathematician
Karl Weierstrass, working in the mid-to-late 1800s. Weierstrass's denition nally banished
any need for vague language like arbitrarily small or approaches, replacing it with a precise,
checkable, purely algebraic statement involving two small positive numbers traditionally named
ε (epsilon) and δ (delta).
ÿ A 150-year gap
It is genuinely remarkable that roughly a century and a half separates Newton and Leibniz's
development of calculus (1660s1680s) from its fully rigorous logical foundation (Cauchy in
the 1820s, Weierstrass nishing the job by the 1870s1880s). For the entire intervening period,
physicists and mathematicians used calculus with complete practical condence, correctly
predicting the motions of planets, the behavior of uids, and the propagation of sound
and heat all without a fully airtight logical justication for the tools they were using.
41
CHAPTER 4. A BRIEF HISTORY OF CALCULUS
This history is a useful reminder, as you work through this book: intuition and correct
computation often arrive well before formal rigor, and both matter. We will develop your
intuition rst (Parts 12), and rigor immediately after (Part 3), rather than starting with
rigor alone, precisely because this mirrors how the subject itself actually developed and how
most people actually learn it best.
The story does not end in the 1800s. Later developments extended and enriched calculus consid-
erably:
Bernhard Riemann formalized the denite integral using the sums of thin rectangles pre-
viewed in Chapter 1 what we now call Riemann sums (Part 5).
Henri Lebesgue, around 1900, developed an even more general theory of integration (Lebesgue
integration), which extends Riemann's approach to handle a wider class of functions a topic
typically studied in advanced university analysis courses, briey mentioned in this book's ad-
vanced appendices.
In the 1960s, Abraham Robinson developed nonstandard analysis, a fully rigorous math-
ematical theory in which genuine innitesimals do exist as legitimate mathematical objects
vindicating, in a precise modern sense, the original intuitions of Leibniz (and even Newton)
that Berkeley had so eectively criticized two centuries earlier.
Calculus continues to expand today into multivariable calculus, dierential equations, and
the vector calculus that underlies electromagnetism and uid dynamics (Part 6), and into
the numerical and computational methods (Part 9) that make modern engineering, computer
graphics, and machine learning possible.
Put the following four milestones in correct chronological order: (i) Weierstrass's epsilon-
delta denition, (ii) Newton's method of uxions, (iii) Archimedes' method of exhaustion,
(iv) Leibniz's rst published papers on calculus.
Solution
Correct order: (iii) Archimedes (∼3rd century BCE) → (ii) Newton (1660s) → (iv) Leibniz
(1680s) → (i) Weierstrass (mid-to-late 1800s). Notice the enormous gap between Archimedes
and Newton (about 1900 years) compared to the much shorter gap between Newton/Leibniz
and Weierstrass (about 200 years) reecting how much faster mathematical progress
became once analytic geometry and algebraic notation were available.
In your own words (two to three sentences), explain why Bishop Berkeley's ghosts of departed
quantities criticism was a legitimate mathematical objection, not merely a philosophical
nitpick.
42
4.8. BEYOND WEIERSTRASS: CALCULUS KEEPS EVOLVING
Hint: Think back to the 00 trap from Chapter 1 Berkeley was pointing at exactly this
kind of inconsistency, decades before Cauchy and Weierstrass resolved it.
Solution
Using the same method-of-exhaustion approach as the worked example in this chapter (in-
scribed regular polygons in a unit circle), estimate the area using a regular 24-gon, and
conrm your answer is between the 12-gon estimate (3.000) and the 96-gon estimate (3.139).
360◦
An = 12 n sin 15◦ .
Hint: Use n with n = 24, so the interior angle in the sine is
Solution
1
A24 = (24) sin(15◦ ) = 12 × 0.2588 ≈ 3.106.
2
This indeed lies between the 12-gon value (3.000) and the 96-gon value (3.139), and it is
closer to π ≈ 3.14159 than the 12-gon estimate was conrming the pattern that more
sides (a ner approximation) means a closer estimate, exactly mirroring the smaller h gives
a better estimate pattern from the dierence quotient in Chapter 1.
Quick FAQ
Q: If calculus wasn't rigorous for 150 years, were results from that era simply
wrong? A: No almost all of the major results derived by Euler, Lagrange, the Bernoulli
family, and others during this period have since been fully veried using modern rigorous
methods. Their intuition for which manipulations were legitimate was excellent, even though
they lacked a formal justication for why. This is a common pattern in the history of
mathematics and physics: correct results often precede a fully rigorous justication.
Q: Why does this book teach limits (Part 3) before fully formalizing derivatives
and integrals, rather than the historical order (uxions/innitesimals rst)? A:
Pedagogically, starting with the intuitive shrinking-interval idea (as we did in Chapter 1),
then formalizing it with limits (Part 3), mirrors how most students actually build under-
standing fastest even though historically, rigor came 150 years after the original intuitive
methods. We get the best of both: the intuitive motivation appears rst, exactly as it did
historically, but you won't be left for 150 years without the rigorous version, unlike Newton
and Leibniz's original audience.
43
CHAPTER 4. A BRIEF HISTORY OF CALCULUS
= Chapter Summary
The area problem traces back to the ancient Greeks (Eudoxus, Archimedes), who de-
veloped the method of exhaustion a direct ancestor of the Riemann sum.
A bitter, largely nationalistic priority dispute between British and continental European
mathematicians followed, with lasting (and mostly unfortunate) consequences for British
mathematics.
R
Leibniz's notation (dy/dx, ) won out and remains standard worldwide, because of how
naturally it supports later techniques like the chain rule and substitution.
Calculus was used successfully for roughly 150 years before being placed on a fully
rigorous logical footing by Cauchy and Weierstrass in the 1800s, using the limit concept
previewed in Chapter 1.
The subject continued evolving through Riemann, Lebesgue, Robinson, and into the
modern computational and multivariable calculus covered later in this book.
This completes Part 1. You've now built a strong intuitive and computational feel for calculus
rst by feeling the tangent problem and area problem directly through worked examples
(Chapters 1 and 2), then by confronting the subtleties of innity and continuity that any
rigorous theory must handle (Chapter 3), and now, nally, by seeing the two-thousand-year
human story behind all of it. Part 2 turns to Functions the mathematical objects that
calculus operates on building (or reviewing, if you've seen some of this before) a complete
and rigorous toolkit for describing, graphing, transforming, and combining functions of every
major type, so that by Part 3, when limits are nally dened with full rigor, you will have
an enormous library of concrete functions to test that denition against.
44