0% found this document useful (0 votes)
3 views55 pages

IB Math AI HL - Complete Revision Guide

The document is a comprehensive revision guide for IB Mathematics Applications and Interpretation HL, structured in exam order covering topics 1-5. It includes foundational knowledge, algebra, geometry, and trigonometry, with clear explanations, formulas, worked examples, and common errors. Each chapter builds on prerequisite skills to ensure fluency in mathematical concepts essential for the exam.

Uploaded by

112aaravgupta
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views55 pages

IB Math AI HL - Complete Revision Guide

The document is a comprehensive revision guide for IB Mathematics Applications and Interpretation HL, structured in exam order covering topics 1-5. It includes foundational knowledge, algebra, geometry, and trigonometry, with clear explanations, formulas, worked examples, and common errors. Each chapter builds on prerequisite skills to ensure fluency in mathematical concepts essential for the exam.

Uploaded by

112aaravgupta
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

IB Mathematics AI HL - The Complete Revision Guide

Table of Contents

IB Mathematics: Applications and Interpretation HL


The Complete Revision Guide
Compiled from the IB Subject Guide (syllabus structure), the Haese Mathematics AI HL
textbook (theory and worked examples), and the Haese Background Knowledge book
(prerequisite skills).
This guide is built in exam order: Topics 1–5 as defined by the IB Guide, SL content before
AHL-only content within each topic. Every subtopic has (1) a plain-English theory
explanation, (2) the formulas you must know, (3) worked examples, and (4) common exam
errors / GDC notes. Read Chapter 0 first if any prerequisite algebra, geometry, or
arithmetic feels shaky — everything after it assumes you’re fluent in it.

Chapter 0: Foundations (Prerequisite Knowledge)


This chapter condenses the Background Knowledge book. It is not examined directly, but
every topic below assumes fluency in it. Skim it once; return to any section that trips you up.

0.1 Number
• Exponent notation: aⁿ means a multiplied by itself n times; a is the base, n the
exponent/power/index. a⁰ = 1 for any a ≠ 0; 0⁰ is undefined.
• Factors and primes: the factors of an integer are the whole numbers that divide it
exactly. A prime has exactly two factors (1 and itself); 1 is neither prime nor
composite. The Fundamental Theorem of Arithmetic: every composite number has a
unique prime factorisation.
• HCF / LCM: the highest common factor is found from the common prime factors
(lowest power of each shared prime); the lowest common multiple uses the highest
power of every prime appearing in either number.
• Roots: √a is the positive number which squares to a; if a is not a perfect square,
estimate by bracketing between consecutive perfect squares (e.g. √30 is between
√25=5 and √36=6).
• Order of operations: brackets, exponents, then multiplication/division left-to-
right, then addition/subtraction left-to-right (BEDMAS/PEMDAS).
• Absolute value |a|: the distance of a from 0, always ≥ 0. |a−b| = distance between a
and b on the number line.
• Rounding & approximation: round to a stated number of decimal places or
significant figures; the first significant figure is the first non-zero digit. Always carry
extra precision through a multi-step calculation and only round the final answer.

0.2 Algebra
• Like terms can be combined by adding/subtracting their coefficients only (2a + 3a
= 5a; a² and a are not like terms).
• Expansion: distribute multiplication over addition, e.g. a(b+c) = ab+ac; (a+b)(c+d) =
ac+ad+bc+bd.
• Factorisation reverses expansion: common factor, difference of squares
a²−b²=(a−b)(a+b), and quadratic trinomial factorisation (find two numbers
multiplying to give ac and summing to b, for ax²+bx+c).
• Solving linear equations: do the same operation to both sides to isolate the
variable; check by substitution.
• Rearranging formulas (changing the subject): treat the desired variable as the
unknown and isolate it using inverse operations, in reverse order of how the
expression was built.
• Substitution: replace variables with given numeric values, respecting order of
operations.

0.3 Measurement and Geometry


• Area formulas: rectangle = l×w; triangle = ½×base×height; parallelogram =
base×height; trapezium = ½(a+b)×h; circle = πr².
• Composite areas: split into rectangles/triangles/circle-parts and add or subtract.
• Volume formulas: prism = cross-sectional area × length; cylinder = πr²h; sphere =
(4/3)πr³; cone = (1/3)πr²h; pyramid = (1/3)×base area×height.
• Pythagoras’ theorem: in a right triangle, c² = a² + b², where c is the hypotenuse
(the side opposite the right angle). Used constantly in 3D geometry (Topic 3) —
always identify the right angle first.
• Similar figures: same shape, proportional sides, equal corresponding angles. If the
linear scale factor is k, areas scale by k² and volumes by k³.
• Transformation geometry: translations (shift by a vector), reflections (flip across a
line), rotations (turn about a point), enlargements (scale from a centre point) —
these underpin AHL 3.9 (matrix transformations) and Topic 2’s graph
transformations.
Worked example (algebraic fractions) Write (x−1)/2 + (x+3)/3 as a single fraction.
LCD=6: [3(x−1) + 2(x+3)] / 6 = (3x−3+2x+6)/6 = (5x+3)/6.

0.4 Measurement
• Perimeter: sum of all side lengths of a polygon; circumference of a circle = 2πr =
πd.
• Area: rectangle=lw; triangle=½bh; parallelogram=bh; trapezium=½(a+b)h;
circle=πr². Composite shapes: split into simple pieces and add/subtract.
• Volume: prism = cross-sectional area × length; cylinder=πr²h; sphere=(4/3)πr³;
cone=(1/3)πr²h; pyramid=(1/3)×base area×height.
• Unit conversion: always convert to consistent units before calculating (e.g. km/h to
m/s: multiply by 1000/3600 = 5/18).
Worked example Convert 36 km/h to m/s. 36 km/h = 36×1000 m / 3600 s = 10 m/s.

0.5 Pythagoras’ Theorem and Circle Geometry


Theory In a right-angled triangle, the square of the hypotenuse (the side opposite the right
angle) equals the sum of the squares of the other two sides. This single relationship
underlies most of Topic 3’s 3D geometry and right-angled trigonometry. Several circle facts
also rely on constructing a right angle: a tangent meets a radius at 90° at the point of
contact; the angle in a semicircle is always 90°; and the line from a circle’s centre
perpendicular to a chord bisects that chord — this last fact is the standard way to find the
distance from a centre to a chord, or a chord’s length given the radius and that distance.
Formula
c²=a²+b² (c = hypotenuse)
Worked example 1 A circle with radius 5 cm has a chord of length 4 cm. Find the shortest
(perpendicular) distance from the centre to the chord. The perpendicular from the centre
bisects the chord, giving a right triangle with hypotenuse 5 and one leg 2 (half the chord).
Distance = √(5²−2²) = √21 ≈ 4.58 cm.
Worked example 2 A rhombus has diagonals of length 10 cm and 24 cm. Find the side
length. Diagonals of a rhombus bisect each other at right angles, giving right triangles with
legs 5 and 12. Side = √(5²+12²) = √169 = 13 cm.

0.6 Coordinate Geometry


Theory Straight-line coordinate geometry (gradient, distance, midpoint) is the direct
prerequisite for SL 2.1 and SL 3.5 — the same formulas simply get reused throughout the
course with more context attached (perpendicular bisectors, tangent lines, vector
directions).
Formulas
m=(y₂-y₁)/(x₂-x₁) d=√((x₂-x₁)²+(y₂-y₁)²) midpoint=((x₁+x₂)/(2),(y₁+y₂)/(2))
Parallel: m₁=m₂. Perpendicular: m₁m₂=−1. Collinear points all share the same gradient
between any pair.
Worked example Find a, given that the line through P(a,−4) and Q(1,8) has gradient 3. (8−
(−4))/(1−a) = 3 ⇒ 12 = 3(1−a) ⇒ 12=3−3a ⇒ a=−3.
0.7 Transformation Geometry
Theory The four basic rigid/similarity transformations — translation (slide by a vector),
reflection (flip across a line), rotation (turn about a point by an angle), and enlargement
(scale from a centre point by a factor k) — are the geometric foundation for AHL 3.9’s
matrix transformations and AHL 2.8’s graph transformations. Translations and rotations
preserve size and shape exactly (congruence); enlargements preserve shape but scale size
by k (similarity, area scales by k²).
Worked example Rectangle ABCD has vertices A(1,4), B(4,4), C(4,2), D(1,2). Find the
image vertices under the translation vector (5,−1). Add the vector to each point: A’(6,3),
B’(9,3), C’(9,1), D’(6,1).

0.8 Similarity
Theory Two figures are similar if they are equiangular (all corresponding angles equal)
and their corresponding sides are all in the same ratio (the scale factor k). This is the
geometric basis for right-angled trigonometry itself (all right triangles with the same acute
angle are similar, which is why the ratio of sides depends only on the angle, not the
triangle’s size) and for the sine/cosine rules. When linear dimensions scale by k, areas scale
by k² and volumes by k³ — a frequently tested idea in applied/modelling contexts.
Worked example Triangles ABE and ACD share angle A, with AB=6, BC=4 (so AC=10), and
DE ∥ BC. Find x=DE given BE=3. Since the triangles are similar (equiangular, sharing angle
A and having corresponding parallel sides): AB/AC = BE/DE ⇒ 6/10 = 3/x ⇒ x = 5.

0.9 Basic Trigonometry (right-angled)


sinθ=(opposite)/(hypotenuse), cosθ=(adjacent)/(hypotenuse),
tanθ=(opposite)/(adjacent)
Memory aid: SOH CAH TOA. Always draw and label the triangle before applying a ratio;
identify which side you know and which you need.

Topic 1: Number and Algebra


SL 1.1 — Standard Form (Scientific Notation)
Theory Standard form writes any number as a × 10^k, where 1 ≤ a < 10 and k is an integer.
It exists to make extremely large or small numbers easy to read, compare, and compute
with — the exponent k immediately tells you the order of magnitude. To convert a number
into standard form, move the decimal point until only one non-zero digit remains before it;
k is positive if you moved the point left (large number) and negative if you moved it right
(small number). Calculators display standard form using an “E” or “^” notation (e.g. 5.2E30
meaning 5.2×10³⁰) — this is never acceptable as a final written answer; always convert
back to a×10^k notation.
When multiplying/dividing numbers in standard form, combine the “a” parts and
add/subtract the exponents; if the resulting “a” part falls outside [1,10), renormalise. When
adding/subtracting, first convert to the same power of 10.
Formulas / rules
• a×10^m × b×10^n = (a×b)×10^(m+n)
• (a×10^m) ÷ (b×10^n) = (a/b)×10^(m−n)
Worked example 1 Write 0.00000000000000000000000167 (the mass of a hydrogen
atom in grams) in standard form. Count digits from the first non-zero digit (1) to the
original decimal point: 24 places → 1.67×10⁻²⁴ g.
Worked example 2 Simplify (3.2×10⁸)×(5×10⁻³), giving your answer in standard form.
3.2×5 = 16, exponent: 8+(−3)=5, so 16×10⁵. Since 16 > 10, renormalise: 1.6×10¹×10⁵ =
1.6×10⁶.
Common errors: forgetting to renormalise when a≥10 after multiplying; leaving a
calculator’s E-notation in a final answer; sign errors on the exponent for small numbers.

SL 1.2 — Arithmetic Sequences and Series


Theory An arithmetic sequence increases (or decreases) by the same fixed amount, the
common difference d, from one term to the next. This models any situation with constant,
additive change: fixed pay rises, evenly-spaced seating, simple interest. Because the
difference is constant, the nth term formula is built by starting at u₁ and adding d exactly
(n−1) times (not n times — a very common error, since u₁ already accounts for zero
additions).
The sum of the first n terms, Sₙ, can be found either from the first term and the last term
(average of first and last, times the number of terms), or from the first term and the
common difference. Sigma notation Σ is just compact notation for “add up all these terms.”
Formulas
u_n = u₁ + (n-1)d S_n = (n)/(2)(u₁+u_n) = (n)/(2)(2u₁+(n-1)d)
Worked example 1 A theatre has 20 seats in row 1, increasing by 3 seats each subsequent
row, for 15 rows. Find the number of seats in the last row and the total number of seats. u₁₅
= 20 + 14(3) = 62 seats in the last row. S₁₅ = (15/2)(20+62) = 615 seats total.
Worked example 2 The 5th term of an arithmetic sequence is 17 and the 12th term is 45.
Find u₁ and d. u₅ = u₁+4d = 17; u₁₂ = u₁+11d = 45. Subtracting: 7d = 28 ⇒ d = 4, so u₁ =
17−4(4) = 1.
Common errors: using n instead of (n−1) in the nth-term formula; mixing up “number of
terms” with “number of differences” (there are always one fewer differences than terms).
SL 1.3 — Geometric Sequences and Series
Theory A geometric sequence multiplies by the same fixed ratio r from one term to the
next — this is the natural model for anything that changes by a constant percentage rather
than a constant amount: population growth, radioactive decay, compound interest, viral
spread. Because growth compounds, geometric sequences change much faster than
arithmetic ones over time — recognising which type a real-world scenario is (constant
addition vs constant multiplication/percentage) is the key modelling skill tested here.
The sum formula comes from a standard trick: write Sₙ, then write rSₙ (shifted by one
power), and subtract — nearly everything cancels, leaving Sₙ(1−r) = u₁(1−rⁿ).
Formulas
u_n = u₁ rⁿ⁻¹ S_n = (u₁(rⁿ-1))/(r-1), r≠1
Worked example 1 A bacteria population starts at 500 and triples every hour. Find the
population after 6 hours. This is the 7th term (u₁ = population at t=0): u₇ = 500(3)⁶ = 364
500.
Worked example 2 The 2nd term of a geometric sequence is 6 and the 5th term is 162.
Find r and u₁. u₅/u₂ = r³ = 162/6 = 27 ⇒ r = 3. Then u₁ = u₂/r = 6/3 = 2.
Common errors: confusing r as a “growth rate” (e.g. writing r = 0.05 for 5% growth
instead of r = 1.05); applying the arithmetic sum formula to a geometric sequence or vice
versa — always check whether consecutive terms add or multiply to the fixed amount first.

SL 1.4 — Financial Applications of Geometric Sequences


Theory Compound interest is a geometric sequence in disguise: each compounding period
multiplies the current value by a fixed growth factor. Interest can compound annually,
monthly, weekly, or daily — the more frequently it compounds, the faster the value grows
for the same annual rate, because interest starts earning interest sooner. Depreciation is
the mirror image — the value is multiplied by a factor less than 1 each period.
In the exam, this is nearly always solved directly on the GDC’s finance/TVM (time-value-of-
money) solver rather than by hand — know how to identify N (number of periods), I%
(annual rate), PV, PMT, and FV in your calculator’s finance app, and get the sign convention
right (money paid out is negative, money received is positive).
Formulas
FV = PV(1+(r)/(100k))^kn
where r = annual percentage rate, k = number of compounding periods per year, n =
number of years.
Worked example 1 €8000 is invested at 4.5% p.a. compounded monthly for 5 years. Find
the final value. FV = 8000(1 + 4.5/1200)⁶⁰ = 8000(1.00375)⁶⁰ ≈ €10 046.30.
Worked example 2 A car worth $30 000 depreciates by 12% per year. Find its value after
4 years. Value = 30000(0.88)⁴ ≈ $18 044.33.
Common errors: using the annual rate directly when compounding monthly (forgetting to
divide by k and multiply the exponent by k); sign errors on the GDC’s finance solver;
forgetting depreciation uses (1−rate), not (1+rate).

SL 1.5 — Laws of Exponents; Introduction to Logarithms


Theory The exponent laws let you manipulate powers algebraically without a calculator. A
logarithm answers the question “to what power must the base be raised to get this
number?” — it is literally defined as the inverse operation of exponentiation. This inverse
relationship, aˣ = b ⟺ logₐb = x, is the single most important idea to internalise here,
because it’s what lets you solve exponential equations (which you can’t isolate x from
algebraically otherwise).
Formulas
a^m aⁿ = a^m+n, (am)/(aⁿ)=am-n, (am)ⁿ=amn, (ab)ⁿ=aⁿbⁿ, a⁻ⁿ=frac1{aⁿ}, a⁰=1
a^x=b ⇔ log_a b = x (b>0)
Worked example 1 Solve 2ˣ = 50. Take log of both sides (any base): x = log₂50 = ln50/ln2
≈ 5.64.
Worked example 2 Simplify (2x³)² ÷ (4x⁵). (2x³)² = 4x⁶; dividing by 4x⁵ gives x.
Common errors: treating logₐb as “logₐ × b”; applying exponent laws to terms that aren’t
being multiplied/divided (e.g. incorrectly “simplifying” aⁿ+aᵐ); forgetting the base must be
positive and ≠1.

SL 1.6 — Approximation and Error


Theory Every measurement has a built-in uncertainty determined by its precision — a
length given as “4.1 cm” (to 1 d.p.) could really be anywhere from 4.05 to 4.15 cm, since
anything in that range rounds to 4.1. These are the upper and lower bounds. When a
measured (approximate) quantity is used in a further calculation (e.g. area from a radius),
the resulting bounds must be recalculated from the bounds of the input, not simply
bounded the same way — errors can compound or partially cancel depending on the
operation.
Percentage error quantifies how far an approximate/estimated value is from the true
(exact) value, as a proportion of the true value — always divide by the exact value, never
the approximate one.
Formulas
% error = |(v_A-v_E)/(v_E)|×100%
Lower bound = value − ½(unit of precision); upper bound = value + ½(unit of precision).
Worked example 1 A circle’s radius is measured as 2.5 cm (1 d.p.). Find the percentage
error in the area if the true radius could be anywhere in the bounds. Bounds on r: 2.45 ≤ r <
2.55. Area at r=2.5: A=19.63 cm². Max area (r=2.55): 20.43 cm². Max % error ≈
(20.43−19.63)/19.63×100% ≈ 4.05%.
Worked example 2 The exact value of a quantity is 80; it was estimated as 76. Find the
percentage error. |76−80|/80 × 100% = 5%.
Common errors: dividing by the estimated value instead of the exact value; forgetting the
absolute value (percentage error is never reported as negative); using the wrong half-unit
when finding bounds (e.g. using 0.1 instead of 0.05 for 1 d.p. data).

SL 1.7 — Amortization and Annuities


Theory A loan is repaid via amortization: each payment covers the interest accrued that
period plus a portion of the principal, so the principal balance shrinks over time and (with
fixed payments) more of each successive payment goes toward the principal. An annuity is
the mirror concept from the saver’s side: regular deposits that grow with compound
interest. Both are solved with the GDC’s finance solver in this course — the algebraic
annuity formula exists but is not required to be reproduced from memory.
Worked example 1 A €200 000 mortgage at 3.6% p.a. compounded monthly is repaid
over 25 years with equal monthly payments. Using a finance solver: N=300, I%=3.6,
PV=200000, FV=0 → monthly payment ≈ €1011.85; total interest paid over the loan ≈
€103 555.
Worked example 2 $200 is deposited at the end of each month into an account earning
3% p.a. compounded monthly, for 10 years. Find the final balance. N=120, I%=3,
PMT=−200, PV=0 → FV ≈ $27 940 (finance solver).
Common errors: mismatching the compounding frequency of the rate with the payment
frequency; wrong sign convention (outgoing payments should be negative in most
calculator conventions); confusing “N” as years instead of number of payment periods.

SL 1.8 — Systems of Equations and Polynomial Equations (Technology)


Theory Some systems and polynomial equations are messy or impossible to solve neatly
by hand — this course expects you to solve them with technology (GDC’s
equation/polynomial solver, or graphing to find intersections/zeros) and interpret the
output correctly, rather than reproduce algebraic techniques like elimination or
substitution by hand (though understanding those helps you sanity-check the technology’s
answer). In exams, a linear system is always guaranteed to have a unique solution.
Worked example 1 Solve: x+y+z=6, 2x−y+z=3, x+2y−z=4. Enter as a matrix/system on the
GDC: x=1, y=2, z=3.
Worked example 2 Find the real roots of x³−6x²+11x−6=0 using technology. GDC
polynomial solver: x = 1, 2, 3.
Common errors: entering coefficients in the wrong order/position in the GDC’s
simultaneous equation solver; forgetting a term with coefficient 0 (e.g. no z-term) must still
be entered as 0, not omitted.

AHL Content (Topic 1)


AHL 1.9 — Laws of Logarithms
Theory Since logarithms are inverse exponents, the log laws are direct translations of the
exponent laws: multiplying numbers corresponds to adding their logs (because multiplying
means adding exponents); raising to a power corresponds to multiplying the log by that
power. These laws let you combine/split log expressions to solve equations where the
variable appears inside multiple logarithms, by condensing everything into a single
logarithm on each side and then equating the arguments.
Formulas
log_a(xy)=log_ax+log_ay, log_a(frac xy)=log_ax-log_ay, log_a(x^m)=mlog_ax
Worked example 1 Solve 2ln x − ln(x−1) = ln4. ln(x²) − ln(x−1) = ln4 ⇒ ln(x²/(x−1)) = ln4
⇒ x²/(x−1)=4 ⇒ x²−4x+4=0 ⇒ (x−2)²=0 ⇒ x=2.
Worked example 2 Given log₂x = 3, find log₂(4x²). log₂4 + 2log₂x = 2 + 2(3) = 8.
Common errors: treating log(x+y) as log x + log y (this identity does not exist); losing
solutions by not checking that logarithm arguments must be positive (x=2 in example 1
must be checked against x−1>0 — it is valid).

AHL 1.10 — Rational Exponents


Theory A rational (fractional) exponent combines a root and a power: x^(m/n) means
“take the nth root, then raise to the mth power” (or equivalently, raise to the mth power
then take the nth root — order doesn’t matter for positive x). This unifies roots and powers
under a single set of exponent laws, so expressions with mixed radicals and powers can be
simplified using the same rules as integer exponents.
Formulas
x(m/n)=n√(xm)=(sqrt[n]x)^m, x^(-1/n)=(1)/(sqrt[n]x)
Worked example 1 Simplify 32^(1/5). 32 = 2⁵, so 32^(1/5) = 2.
Worked example 2 Simplify (x(3/2)·x(−1/4)) / x^(1/4). Exponent: 3/2 − 1/4 − 1/4 = 1, so
the answer is x.
Common errors: applying the root only to part of a product/quotient instead of the whole
expression; sign errors when the exponent’s numerator or denominator is negative.

AHL 1.11 — Sum of Infinite Geometric Series


Theory If |r| < 1, each successive term of a geometric sequence shrinks toward zero, so the
partial sums approach a finite limiting value even though there are infinitely many terms
— this only happens because the terms shrink fast enough. If |r| ≥ 1 the sum diverges
(grows without bound, or oscillates) and S∞ does not exist. This idea underlies fractal
geometry (repeatedly shrinking copies), the “bouncing ball” total-distance problem, and
the long-run behaviour of certain Markov chains.
Formula
S_∞=(u₁)/(1-r), |r|<1
Worked example 1 A ball dropped from 10 m rebounds to 60% of its previous height each
bounce. Find the total vertical distance travelled. Falls: 10, then repeated up+down trips of
10(0.6), 10(0.6)², … D = 10 + 2×[10(0.6)/(1−0.6)] = 10 + 2(15) = 40 m.
Worked example 2 Find S∞ for 8 + 4 + 2 + 1 + … r = 1/2, S∞ = 8/(1−0.5) = 16.
Common errors: applying S∞ when |r|≥1 (it simply doesn’t exist — check this first);
forgetting the “there and back” doubling in bounce-distance problems.

AHL 1.12 — Complex Numbers (Cartesian Form)


Theory Not every polynomial has real roots — x²+1=0 has no real solution, because no real
number squares to give −1. Complex numbers extend the real numbers by defining i =
√(−1), so every polynomial equation does have a solution (the Fundamental Theorem of
Algebra). A complex number z = a+bi has a real part a and imaginary part b; arithmetic with
complex numbers follows the ordinary rules of algebra, replacing i² with −1 wherever it
appears. The conjugate z̄ = a−bi is useful for dividing complex numbers (multiply top and
bottom by the conjugate of the denominator to make it real) and for the fact that non-real
roots of real-coefficient polynomials always come in conjugate pairs.
Formulas
z=a+bi, z̄ = a-bi, |z|=√(a²+b²)
Quadratic formula still applies when b²−4ac<0: x=(-b±√(b²-4ac))/(2a), giving a complex
conjugate pair.
Worked example 1 Solve x²−2x+5=0. x = (2±√(4−20))/2 = (2±√−16)/2 = (2±4i)/2 = 1±2i.
Worked example 2 Given z₁=3+2i, z₂=1−4i, find z₁×z₂ and z₁/z₂. z₁z₂ = 3−12i+2i−8i² =
3−10i+8 = 11−10i. z₁/z₂ = (3+2i)(1+4i) / (1²+4²) = (3+12i+2i+8i²)/17 = (−5+14i)/17 =
−5/17 + (14/17)i.
Common errors: forgetting i²=−1 when expanding brackets; dividing complex numbers
without multiplying by the conjugate; sign errors in the conjugate.

AHL 1.13 — Polar/Exponential Form of Complex Numbers


Theory Every complex number can also be located on the Argand diagram (a plane with
real axis horizontal, imaginary axis vertical) using its distance from the origin (modulus, r)
and angle from the positive real axis (argument, θ), instead of its Cartesian coordinates.
Multiplying complex numbers in polar/exponential form is far easier than in Cartesian
form: moduli multiply and arguments add, which geometrically means multiplication by a
complex number is a rotation-and-scale of the plane. This connects directly to AHL 3.9
(matrix transformations).
Formulas
z=r(cosθ+isinθ)=r cis θ = re^(iθ) r=|z|, θ=arg z
z₁z₂=r₁r₂ cis(θ₁+θ₂) (z₁)/(z₂)=(r₁)/(r₂)cis(θ₁-θ₂)
Worked example 1 Write z=1+i√3 in polar and exponential form. r=√(1+3)=2,
θ=arctan(√3/1)=π/3. z = 2 cis(π/3) = 2e^(iπ/3).
Worked example 2 Given z₁=4cis(π/6), z₂=2cis(π/4), find z₁z₂. r: 4×2=8; θ:
π/6+π/4=5π/12. z₁z₂ = 8cis(5π/12).
Common errors: using the wrong quadrant for θ (arctan alone doesn’t distinguish
quadrants — check the signs of a and b first); mixing degrees and radians.

AHL 1.14 — Matrices: Algebra, Determinants, Systems


Theory A matrix is a rectangular array of numbers; its order (m×n) tells you
rows×columns. Matrices add/subtract element-by-element (requires identical order) and
multiply by a genuinely different rule (row-by-column “dot products”) which is not
commutative — order matters. The determinant of a square matrix is a single number that
tells you whether the matrix is invertible (det ≠ 0) and, geometrically, the area/volume
scale factor of the linear transformation it represents. A linear system Ax=b can be solved
instantly once A⁻¹ is known: x = A⁻¹b.
Formulas (2×2 case, examinable by hand)
A=(a, b; c, d) ⇒ det A=ad-bc, A⁻¹=(1)/(ad-bc)(d, -b; -c, a)
Worked example 1 Solve using matrices: 2x+y=5, x−y=−1. A=(2,1;1,−1), det A=−2−1=−3.
A⁻¹ = (1/−3)(−1,−1;−1,2). (x;y) = A⁻¹(5;−1) = (1/−3)(−4;−7) = (4/3, 7/3).
Worked example 2 Given A=(1,2;3,4) and B=(0,1;1,0), find AB. AB = (1×0+2×1, 1×1+2×0;
3×0+4×1, 3×1+4×0) = (2,1; 4,3).
Common errors: multiplying matrices element-by-element instead of row-by-column;
assuming AB=BA (generally false); trying to invert a matrix with det=0 (it has no inverse).

AHL 1.15 — Eigenvalues and Eigenvectors


Theory An eigenvector of a matrix A is a special direction that the transformation A doesn’t
rotate — it only stretches (or shrinks/reverses) it, by a factor called the eigenvalue. Finding
these requires solving det(A−λI)=0 (the characteristic equation) for λ, then for each λ
solving (A−λI)v=0 for the direction v. If A has two distinct real eigenvalues, it can be
diagonalised as A=PDP⁻¹, which makes computing high powers of A (e.g. for long-run
population models) dramatically easier, since Aⁿ=PDⁿP⁻¹ and Dⁿ is trivial to compute (just
raise each diagonal entry to the nth power).
Formulas
det(A-λ I)=0, Aⁿ=PDⁿP⁻¹
Worked example 1 Find the eigenvalues/eigenvectors of A=(4,1; 2,3). det(A−λI) = (4−λ)
(3−λ)−2 = λ²−7λ+10=0 ⇒ λ=5, 2. λ=5: (A−5I)v=0 ⇒ −x+y=0 ⇒ v₁=(1,1). λ=2: 2x+y=0 ⇒
v₂=(1,−2).
Worked example 2 Using the above, find A³ applied to the vector (1,1). Since (1,1) is the
eigenvector for λ=5: A³(1,1) = 5³(1,1) = (125,125).
Common errors: forgetting to set up (A−λI), not (A−λ); sign slips solving the resulting
homogeneous 2×2 system; forgetting eigenvectors are only defined up to a scalar multiple
(any multiple of a found eigenvector is also valid).

Topic 2: Functions
SL 2.1 — Equations of a Straight Line
Theory A straight line’s defining property is constant gradient (rate of change) — the same
idea as a constant common difference in an arithmetic sequence, now applied continuously.
Gradient m = (change in y)/(change in x) between any two points on the line. Three
equivalent forms exist because different information is easiest to plug into each: gradient–
intercept form is best when you know the y-intercept directly; point–gradient form is best
when you know one point and the gradient; general form is a tidy way to state any line,
including vertical ones (which have no defined gradient).
Formulas
y=mx+c y-y₁=m(x-x₁) ax+by+d=0
m₁ parallel m₂ ⇔ m₁=m₂ m₁⊥ m₂ ⇔ m₁m₂=-1
Worked example 1 Find the line through (2,5) perpendicular to y=3x−1. m_perp = −1/3.
y−5 = −1/3(x−2) ⇒ y = −x/3 + 17/3.
Worked example 2 Find the equation of the line through (1,2) and (4,11). m =
(11−2)/(4−1) = 3. y−2=3(x−1) ⇒ y=3x−1.
Common errors: flipping the perpendicular-gradient rule (negative reciprocal, not just
negative); mixing up (x₁,y₁) order when computing gradient between two points.

SL 2.2 — Concept of a Function; Inverse Functions


Theory A function is a rule assigning exactly one output to each input in its domain; the
range is the resulting set of outputs. An inverse function reverses this rule — it asks “what
input produced this output?” — and only exists cleanly (as a function) if the original
function is one-to-one (each output comes from exactly one input), otherwise the “inverse”
would have to assign multiple outputs to some input, which breaks the definition of a
function. Graphically, f⁻¹ is the reflection of f in the line y=x, since swapping x and y swaps
the roles of input and output.
Formulas Domain of f⁻¹ = range of f; range of f⁻¹ = domain of f. To find f⁻¹ algebraically:
swap x and y in y=f(x), then solve for y.
Worked example 1 f(x)=√(2−x), domain x≤2. Find f⁻¹(10). Set √(2−x)=10 ⇒ 2−x=100 ⇒
x=−98, so f⁻¹(10)=−98.
Worked example 2 Find the inverse of f(x)=3x+2. Swap: x=3y+2 ⇒ y=(x−2)/3. f⁻¹(x) =
(x−2)/3.
Common errors: confusing f⁻¹(x) with 1/f(x) (they are completely different); forgetting to
restrict the domain when the original function isn’t one-to-one over its natural domain.

SL 2.3 — Graphing Functions


Theory “Sketch” and “draw” mean different things in IB exams: a sketch shows the general
shape and key features without needing graph paper or exact scale, while a draw
instruction requires an accurate, to-scale graph (typically on provided grid paper), with all
axes labelled and a suitable scale chosen. Regardless of which is asked, every graph needs
labelled axes, and any key features mentioned in the question (intercepts, asymptotes,
turning points) must be clearly shown/labelled.
Worked example Sketch y = x² − 4 and y = 2x on the same axes, and estimate their
intersection. Sketch shows the parabola crossing the x-axis at x=±2, y-intercept −4; the line
through the origin with gradient 2. They appear to cross near x≈−1 and x≈4 (confirm
algebraically or with GDC in SL2.4).
Common errors: unlabelled axes/scale losing marks even with a correct-shaped curve;
sketching a function outside a domain that was explicitly restricted in the question.

SL 2.4 — Key Features of Graphs


Theory Modern IB assessment expects you to read key features off a GDC-generated graph
rather than derive them all algebraically: maximum/minimum points, x- and y-intercepts,
symmetry, and asymptotic behaviour are typically found using the graphing app’s built-in
analysis tools (trace, minimum/maximum, zero, intersect). Vertical asymptotes occur
where a function is undefined (often division by zero); horizontal asymptotes describe the
function’s long-run behaviour as x→±∞.
Worked example 1 Find the point(s) of intersection of y=x²−1 and y=2x+2. GDC intersect
feature (or algebraically: x²−2x−3=0 ⇒ x=−1,3): (−1,0) and (3,8).
Worked example 2 For f(x) = 1/(x−2), state the equations of the asymptotes. Vertical
asymptote at x=2 (undefined there); horizontal asymptote at y=0 (as x→±∞, f(x)→0).
Common errors: reporting a turning point’s x-coordinate but forgetting the y-coordinate
(or vice versa); confusing “zero of a function” (x-intercept) with “y-intercept.”

SL 2.5 — Modelling with Functions


Theory Recognising which family of function fits a real-world situation is the core skill:
constant rate of change → linear; constant acceleration/area-type growth or a single
turning point → quadratic; constant percentage change with a long-run ceiling/floor →
exponential; two turning points or an S-shaped inflection → cubic; repeating, oscillating
behaviour → sinusoidal; a relationship where one quantity is proportional to a power of
another → direct/inverse variation. Each family has diagnostic graph features (listed
below) that let you match data or a description to the right model.
Key features by family

Model Form Key features


Linear f(x)=mx+c constant gradient m, y-
intercept c
Quadratic f(x)=ax²+bx+c axis of symmetry, vertex, up
to 2 zeros
Exponential f(x)=ka^x+c, ke^(rx)+c horizontal asymptote y=c
Model Form Key features
Direct/inverse variation f(x)=ax^n vertical asymptote x=0 if
n<0
Cubic f(x)=ax³+bx²+cx+d up to 2 turning points
Sinusoidal f(x)=a sin(bx)+d amplitude a, period 360°/b,
principal axis y=d

Worked example 1 The height of a ferris-wheel car is modelled by h(t)=15sin(30t)+17 (t


in minutes). Find the amplitude, period, and max/min height. Amplitude=15, principal axis
h=17 (max=32 m, min=2 m), period=360/30=12 minutes.
Worked example 2 A population is modelled by P(t) = 200(1.08)^t. State whether this
represents growth or decay, and find the initial population and annual growth rate. Since
base 1.08>1, this is growth; initial population P(0)=200; growth rate = 8% per year.
Common errors: picking a model family based only on “the data goes up” without
checking whether the increase is additive (linear) or multiplicative (exponential);
forgetting the vertical shift d changes the asymptote/principal axis, not the amplitude.

SL 2.6 — The Modelling Process


Theory This subtopic is examined as an explicit four-stage cycle and is a favourite source
of “state/comment/interpret” command-term questions: 1. Develop — choose an
appropriate function family based on the context and any given data shape; decide a
sensible domain (e.g. time can’t be negative). 2. Find parameters — use given data
points/initial conditions, solved simultaneously (by technology) to pin down the unknown
constants. 3. Test and reflect — check the model against data not used to build it;
comment on whether the shape/behaviour makes real-world sense. 4. Use — interpret
model outputs in context, and predict — while explicitly flagging that extrapolation
(predicting outside the range of the original data) is unreliable, since the underlying real-
world process may not continue behaving the same way.
Worked example A cup of coffee cools from 90°C in a 20°C room; after 10 minutes it is
60°C. Model: T(t)=20+70e^(−kt). Find k: 60=20+70e^(−10k) ⇒ e^(−10k)=4/7 ⇒ k =
−ln(4/7)/10 ≈ 0.0560. Predict T(30): T(30)=20+70e^(−0.0560×30) ≈ 33.4°C — but since
30 minutes is not far beyond the given data (10 min), this interpolation/mild extrapolation
is reasonably trustworthy; predicting T(500) would not be.
Common errors: substituting given data into the wrong parameter; failing to explicitly
discuss extrapolation risk when a question asks you to “comment on the validity” of a
prediction.
AHL Content (Topic 2)
AHL 2.7 — Composite and Inverse Functions
Theory A composite function (f∘g)(x)=f(g(x)) applies g first, then f, to the result — order
matters, and (f∘g) is generally different from (g∘f). Composing a function with its own
inverse, in either order, always returns the original input (they “undo” each other exactly),
which is in fact the defining property of an inverse function. Some functions only have an
inverse over a restricted domain (e.g. a parabola must be split at its vertex).
Worked example 1 f(x)=2x+1, g(x)=x². Find (f∘g)(3) and (g∘f)(3). (f∘g)(3)=f(9)=19. (g∘f)
(3)=g(7)=49.
Worked example 2 f(x)=(x−3)²−2, restricted to x≥3. Find f⁻¹(x). y=(x−3)²−2 ⇒
x−3=√(y+2) (positive root since x≥3) ⇒ x=3+√(y+2). f⁻¹(x)=3+√(x+2).
Common errors: computing (g∘f) when (f∘g) was asked for; forgetting the domain
restriction is required for the inverse of a non-one-to-one function to exist.

AHL 2.8 — Transformations of Graphs


Theory Every transformation of y=f(x) can be classified as either changing outputs (acting
on y, so it moves the graph vertically / stretches vertically) or changing inputs (acting on x,
so it moves the graph horizontally / stretches horizontally — and input-based
transformations behave “backwards” from what you’d intuitively expect, e.g. f(x+3) shifts
left). When multiple transformations are combined, the order in which they’re applied
generally matters and must match the order operations would be applied to a specific x-
value.
Formulas

Transformation Effect
y=f(x)+b vertical translation by b
y=f(x−a) horizontal translation by a
y=−f(x) reflection in x-axis
y=f(−x) reflection in y-axis
y=pf(x) vertical stretch, scale factor p
y=f(qx) horizontal stretch, scale factor 1/q

Worked example 1 Describe the transformation from y=sin x to y=4sin(2x). Vertical


stretch factor 4 (amplitude), then horizontal stretch factor 1/2 (period halves).
Worked example 2 The graph of y=f(x) has a minimum at (3,−2). Find the minimum point
of y=f(x−1)+5. Horizontal shift +1, vertical shift +5: new minimum at (4, 3).
Common errors: applying a horizontal shift in the wrong direction (f(x−a) shifts right by a,
not left); applying transformations in an order that doesn’t match how they’re written
algebraically.

AHL 2.9 — Further Modelling Functions


Theory Beyond the basic families, three further models appear regularly: natural-log
models (good for diminishing-returns relationships), phase-shifted sinusoidal models (for
oscillations that don’t start at their principal axis at t=0), and the logistic model, which
captures growth that starts exponential but levels off at a carrying capacity L as
resources/space become limited — the single most important “S-shaped curve” context in
the course (population in a confined habitat, spread of a disease/technology through a
finite population). Piecewise models stitch together different functions over different
domains — continuity at the join points is often the crux of the algebra required.
Formulas
f(x)=a+bln x f(x)=asin(b(x-c))+d f(x)=(L)/(1+Ce^-kx)
Logistic: horizontal asymptote f(x)=L (the carrying capacity) as x→∞.
Worked example 1 (logistic) Bacteria in a petri dish: P(t) = 500/(1+9e^(−0.4t)). Find the
carrying capacity and the initial population. Carrying capacity L=500. P(0) = 500/(1+9) =
50.
Worked example 2 (piecewise continuity) f(x) = 1+x for 0≤x<2, and f(x)=ax²+x for x≥2.
Find a for continuity at x=2. Left limit: 1+2=3. Right value: 4a+2. Set equal: 4a+2=3 ⇒
a=1/4.
Common errors: mistaking the logistic model’s initial value for its carrying capacity (they
are different constants — L is the asymptote, not P(0)); forgetting to check both pieces
agree exactly at the join when solving for a continuity parameter.

AHL 2.10 — Scaling Large/Small Numbers; Linearizing Data


Theory Raw data spanning many orders of magnitude (e.g. income, earthquake energy,
population) is easier to visualise and analyse on a logarithmic scale, because logs compress
multiplicative differences into additive ones. This same idea lets you test whether a dataset
follows an exponential or power relationship, and extract its parameters, by “linearizing”:
taking logs of one or both variables turns a curved relationship into a straight line, whose
gradient and intercept (found via linear regression — SL 4.4) reveal the original model’s
parameters.
Formulas If y=ka^x: ln y = ln k + x ln a (plot ln y vs x — a semi-log plot; gradient=ln a,
intercept=ln k). If y=ax^n: ln y = ln a + n ln x (plot ln y vs ln x — a log-log plot; gradient=n,
intercept=ln a).
Worked example 1 Data suspected to follow y=ax^n. A ln y vs ln x plot gives gradient 2.03,
intercept 1.10. Find the model. n≈2; ln a=1.10 ⇒ a≈3.00. Model: y≈3x².
Worked example 2 Data suspected to follow y=ka^x. A ln y vs x plot gives gradient 0.405,
intercept 2.30. Find k and a. ln a=0.405 ⇒ a=e^0.405≈1.50. ln k=2.30 ⇒ k=e^2.30≈9.97.
Common errors: taking logs of only one variable when a power model (both logs needed)
was intended, or vice versa for an exponential model; misreading which axis is which when
extracting gradient/intercept.

Topic 3: Geometry and Trigonometry


SL 3.1 — 3D Coordinate Geometry and Volumes
Theory Distances and midpoints in 3D extend directly from 2D by adding a third (z)
coordinate under the same square-root/averaging logic — this is simply Pythagoras’
theorem applied twice (once in a base plane, once vertically). Volume and surface area
formulas for pyramids, cones, and spheres are provided in the exam formula booklet, so the
real skill is identifying which right triangle inside a 3D solid lets you find an unknown length
or angle using Pythagoras/trigonometry — sketching a clear 2D cross-section of the solid is
usually the fastest route to the answer.
Formulas
d=√((x₂-x₁)²+(y₂-y₁)²+(z₂-z₁)²)
(Volume/surface-area formulas for pyramids, cones, spheres — given in the formula
booklet.)
Worked example 1 A cone has base radius 5 cm and slant height 13 cm. Find its height
and volume. Right triangle (radius, height, slant height): h=√(13²−5²)=12 cm.
V=(1/3)πr²h=(1/3)π(25)(12)=100π ≈ 314.2 cm³.
Worked example 2 Find the distance between A(1,2,3) and B(4,6,15).
d=√((4−1)²+(6−2)²+(15−3)²)=√(9+16+144)=√169=13.
Common errors: forgetting the third coordinate entirely (defaulting to the 2D distance
formula); picking the wrong right triangle within a complex solid — always sketch a clean
2D extract of just the relevant triangle first.

SL 3.2 — Right-Angled Trigonometry; Sine and Cosine Rules


Theory SOH-CAH-TOA only works in right-angled triangles. The sine and cosine rules
extend trigonometry to any triangle. The sine rule relates each side to the sine of its
opposite angle, and is used when you have an angle-side opposite pair plus one more piece
of information. The cosine rule generalises Pythagoras (notice it reduces to c²=a²+b² when
C=90°, since cos90°=0) and is used when you know two sides and the included angle (find
the third side), or all three sides (find any angle, using the rearranged form). At SL there’s
no ambiguous case to worry about.
Formulas
(a)/(sin A)=(b)/(sin B)=(c)/(sin C) c²=a²+b²-2abcos C Area=(1)/(2)absin C
Worked example 1 Triangle with a=8, b=10, C=55°. Find c and the area. c²=8²+10²−2(8)
(10)cos55°=164−91.7=72.3 ⇒ c≈8.50. Area=(1/2)(8)(10)sin55°≈32.8 units².
Worked example 2 Triangle with a=7, A=40°, B=65°. Find b. By sine rule: b/sin65° =
7/sin40° ⇒ b = 7sin65°/sin40° ≈ 9.87.
Common errors: using the cosine rule when the sine rule is more appropriate (and vice
versa) — check what information you actually have (SAS/SSS → cosine rule; AAS/ASA/SSA
→ sine rule); forgetting the “minus” in the cosine rule.

SL 3.3 — Applications of Trigonometry


Theory This subtopic doesn’t introduce new mathematics — it tests whether you can
translate a word problem into a labelled diagram and identify the correct triangle and trig
tool. Angles of elevation/depression are always measured from the horizontal. Bearings are
measured clockwise from north, always given as three digits (e.g. 065°). The key skill is
careful, systematic diagram construction: mark all known lengths/angles, identify right
angles, and decide whether right-angled trig, the sine rule, or the cosine rule applies.
Worked example 1 From a point 50 m from the base of a tower, the angle of elevation to
the top is 32°. Find the tower’s height. h = 50 tan32° ≈ 31.2 m.
Worked example 2 A ship sails 40 km on a bearing of 060°, then 30 km on a bearing of
150°. Find the distance from the start. The angle between the two legs (bearings 060° and
150°) is 90°, so Pythagoras applies directly: distance=√(40²+30²)=√2500=50 km.
Common errors: measuring elevation/depression from the vertical instead of the
horizontal; bearings measured anticlockwise or not starting from north; not converting a
compass bearing problem into an internal triangle angle correctly.

SL 3.4 — Arcs and Sectors


Theory An arc is a fraction of a circle’s circumference; a sector is the corresponding
fraction of its area — both scale by exactly the same fraction θ/360° (in degrees) of the full
circle, since the sector angle determines what proportion of the whole circle you have.
Formulas
Arc length=(θ)/(360)×2π r Sector area=(θ)/(360)×π r²
Worked example 1 Sector of radius 6 cm, angle 75°. Find the arc length and sector area.
Arc length = (75/360)(2π×6) ≈ 7.85 cm. Area = (75/360)(π×36) ≈ 23.6 cm².
Worked example 2 A sector has arc length 12 cm and radius 8 cm. Find the sector’s angle
in degrees. 12 = (θ/360)(2π×8) ⇒ θ = 12×360/(16π) ≈ 85.9°.
Common errors: using diameter instead of radius; forgetting sector area uses πr² (not the
arc-length formula’s 2πr).

SL 3.5 — Perpendicular Bisectors


Theory The perpendicular bisector of a segment AB is the set of all points equidistant from
A and B — this single geometric fact is what makes it the foundation of Voronoi diagrams
(SL 3.6): the boundary between two “closer to A” and “closer to B” regions must be exactly
this line. To find its equation, you need the midpoint of AB (a point the bisector passes
through) and the negative reciprocal of AB’s gradient (since the bisector is perpendicular
to AB).
Worked example 1 Find the perpendicular bisector of A(1,2) and B(5,8). Midpoint (3,5);
gradient AB = 6/4 = 3/2, so m_perp = −2/3. y−5 = −2/3(x−3).
Worked example 2 Find the perpendicular bisector of A(0,0) and B(6,0). Midpoint (3,0);
AB is horizontal so the perpendicular bisector is vertical: x=3.
Common errors: using the midpoint of AB as if it were on line AB extended rather than the
bisector; forgetting that a horizontal segment’s perpendicular bisector is a vertical line
(undefined gradient — can’t use point-gradient form directly).

SL 3.6 — Voronoi Diagrams


Theory Given a set of “sites” (points), a Voronoi diagram partitions the plane into cells, one
per site, where every point in a cell is closer to that site than to any other. Cell boundaries
(edges) are segments of perpendicular bisectors between neighbouring sites; the points
where three or more edges meet (vertices) are equidistant from three or more sites. In
exams, you are typically given (not asked to construct from scratch) some perpendicular
bisector equations and asked to find vertices, add a new site, or apply the diagram:
nearest-neighbour interpolation estimates an unknown value at a point using the value
recorded at that point’s cell site (e.g. estimating rainfall at an ungauged location using the
nearest weather station). The “toxic waste dump” problem asks for the location maximising
the minimum distance to every site — this is always found at a Voronoi vertex (or on the
edge of the diagram’s bounding region).
Worked example Given sites A(0,0), B(4,0), C(2,5), the perpendicular bisectors of AB and
AC intersect at the Voronoi vertex equidistant from all three sites. This point is the best
candidate for a facility that must be as far as possible from all three sites (or, in the toxic-
waste framing, the safest place to bury the waste, minimizing the worst-case risk to any
site).
Common errors: confusing which side of a perpendicular bisector belongs to which site
(the closer site is on the same side as the bisector “leans away from”); assuming the toxic-
waste-dump answer is always inside the triangle of sites — it can lie on a boundary edge of
the region under consideration.

AHL Content (Topic 3)


AHL 3.7 — Radian Measure
Theory A radian is defined so that an arc of length equal to the radius subtends an angle of
exactly 1 radian at the centre — this is why arc length and sector area formulas become
dramatically simpler in radians (no ×360 conversion factor needed). Radians are the
natural unit for calculus involving trigonometric functions (the derivative formulas in AHL
5.9 only hold in radians), so HL papers assume radians unless a degree symbol is explicitly
shown.
Formulas
2π rad=360°, arc length=rθ, sector area=(1)/(2)r²θ (θ in radians)
Worked example 1 Convert 5π/6 radians to degrees. (5π/6)×(180/π) = 150°.
Worked example 2 Sector radius 8 cm, angle π/3 rad. Find the arc length and sector area.
Arc length = 8×π/3 ≈ 8.38 cm. Area = (1/2)(64)(π/3) ≈ 33.5 cm².
Common errors: leaving the calculator in the wrong angle mode (degrees vs radians) —
always check before evaluating trig functions; forgetting the ×360 or π conversion factor
entirely disappears once working in radians.

AHL 3.8 — Unit Circle, Pythagorean Identity, Ambiguous Case


Theory The unit circle (radius 1, centred at the origin) defines cos θ and sin θ as the x- and
y-coordinates of the point reached by rotating θ from the positive x-axis — this extends
sine and cosine to any angle, not just those inside a right triangle (0°–90°). Because every
point on the unit circle satisfies x²+y²=1, we immediately get the Pythagorean identity. The
ambiguous case of the sine rule arises specifically when you’re given two sides and a non-
included angle (SSA): since sin θ = sin(180°−θ), there can be two different valid triangles
matching the given information, and you must check whether both are geometrically
possible (angle sum <180°).
Formulas
cos²θ+sin²θ=1, tanθ=(sinθ)/(cosθ)
Worked example 1 (ambiguous case) a=7, b=9, A=40°. Find the possible value(s) of B.
sinB/9 = sin40°/7 ⇒ sinB=0.8265 ⇒ B=55.8° or B=180−55.8=124.2°. Both give an angle
sum under 180° with A=40°, so both triangles are valid.
Worked example 2 Given sinθ=3/5 and θ is obtuse (in the second quadrant), find cosθ.
cos²θ=1−9/25=16/25 ⇒ cosθ=±4/5; since θ is obtuse, cosine is negative: cosθ=−4/5.
Common errors: automatically discarding the second (obtuse) solution in an SSA problem
without checking if it’s actually invalid; sign errors when using the Pythagorean identity
outside the first quadrant.

AHL 3.9 — Matrix Transformations


Theory Every linear transformation of the plane (reflection, rotation, stretch, enlargement)
can be represented by a 2×2 matrix acting on column-vector coordinates; adding a
translation vector afterwards extends this to any affine transformation. Composing two
transformations (apply one, then the other) corresponds to multiplying their matrices —
but because matrix multiplication isn’t commutative, the order of the multiplication must
match the order the transformations are actually applied in. The determinant’s absolute
value gives the area scale factor: |det|=1 for rotations/reflections (they preserve area), |
det|≠1 for stretches/enlargements. Repeatedly applying a transformation to a starting
shape (iteration) generates self-similar patterns — fractals.
Formulas
(x’; y’)=(a, b; c, d)(x; y)+(e; f)
Worked example 1 Find the image of (3,1) under a 90° anticlockwise rotation about the
origin. Rotation matrix (0,−1;1,0): (0×3−1×1, 1×3+0×1) = (−1,3). Area scale factor = |det|
=1.
Worked example 2 A stretch matrix (2,0; 0,1) is applied to the unit square. Find the area
scale factor and describe the transformation. det = 2×1−0×0=2, so the area doubles:
horizontal stretch, scale factor 2 (parallel to the x-axis).
Common errors: applying the translation vector before the matrix multiplication instead
of after; multiplying two transformation matrices in the wrong order for a composite
transformation.

AHL 3.10 — Vectors: Definitions and Operations


Theory A vector has both magnitude (size) and direction, unlike a scalar which is just a
number. Vectors are written as columns (or using i, j, k unit vectors along the axes) and
combine by adding/subtracting corresponding components — geometrically this
corresponds to “tip-to-tail” addition. Multiplying a vector by a scalar k stretches/shrinks it
(and reverses direction if k<0) but never changes the direction line it lies along — this is
the definition of parallel vectors. The magnitude is found via a 3D extension of Pythagoras;
dividing a vector by its own magnitude produces a unit vector — same direction, length
exactly 1 — useful for expressing “speed in a given direction” as a full velocity vector.
Formulas
v=(v₁,v₂,v₃), |v|=√(v₁²+v₂²+v₃²), v̂=(v)/(|v|)
Worked example 1 v=3i+4j represents a direction; an object travels at speed 7 m/s along
this direction. Find the velocity vector. |v|=5, unit vector = (0.6, 0.8). Velocity = 7×(0.6,0.8)
= (4.2, 5.6).
Worked example 2 Given a=(2,−1,3) and b=(−1,4,1), find 2a−b. 2a=(4,−2,6); 2a−b=(4−
(−1), −2−4, 6−1) = (5,−6,5).
Common errors: confusing a vector with its magnitude (a scalar) — a vector answer needs
direction/components, not just a single number; component-wise sign errors when
subtracting.

AHL 3.11 — Vector Equation of a Line


Theory A line in 2D or 3D is described by a single fixed point on it (the position vector a)
plus a direction vector b — every other point on the line is reached by travelling some
multiple λ of b from a. This single vector equation replaces the need for a Cartesian
equation in 3D (where “y=mx+c” doesn’t generalise cleanly), and λ acts as a parameter that
traces out every point on the line as it varies over all real numbers.
Formulas
r=a+λb x=x₀+λ l, y=y₀+λ m, z=z₀+λ n
Worked example 1 Find the vector equation of the line through (1,2,0) with direction
(2,−1,3). r = (1,2,0) + λ(2,−1,3), i.e. x=1+2λ, y=2−λ, z=3λ.

r=(1,2,0)+λ(2,−1,3). From x: 1+2λ=5 ⇒ λ=2. Check y: 2−2=0 ✓. Check z: 3(2)=6 ✓. Yes, the
Worked example 2 Determine whether the point (5,0,6) lies on the line

point lies on the line (λ=2 satisfies all three equations).


Common errors: using a direction vector that isn’t actually parallel to the line
(e.g. accidentally using two points’ sum instead of their difference); forgetting to check all
three component equations agree on the same λ value when testing a point.
AHL 3.12 — Vector Kinematics
Theory Position vectors extend 1D kinematics (displacement, velocity) into 2D/3D: an
object’s position at time t is r(t) = r₀+vt for constant velocity motion. To compare two
moving objects (e.g. two ships), the relative position vector B−A tells you the displacement
from one to the other at any time; setting its magnitude to a target value (like a collision-
avoidance distance) or minimising it (finding closest approach) are the two most common
question types. If both objects reach exactly the same position at the same time t, they
collide/meet; matching position without matching time only shows their paths cross, not
that they meet.
Worked example 1 Ship A: r_A=(2,0)+t(1,2). Ship B: r_B=(10,4)+t(−1,0). Determine if the
ships collide. Set x equal: 2+t=10−t ⇒ t=4. Check y: 0+2(4)=8, but B’s y=4 always (no t-
dependence) — 8≠4, so they do not meet at any common time.
Worked example 2 For the ships above, describe how you would find the time of closest
approach. Minimise |r_A(t)−r_B(t)| (a function of t) using calculus or the GDC’s minimum-
finding tool — this gives the time at which the distance between the ships is smallest.
Common errors: concluding two objects “meet” just because their paths cross
geometrically, without checking they’re at that crossing point at the same time; sign errors
setting up the relative position vector (must be consistent, e.g. always B−A).

AHL 3.13 — Scalar and Vector Products


Theory The scalar (dot) product combines two vectors into a single number that measures
how aligned they are — it’s zero exactly when the vectors are perpendicular, which makes
it the standard tool for angle-between-vectors and perpendicularity tests. The vector
(cross) product instead produces a new vector, perpendicular to both original vectors,
whose magnitude equals the area of the parallelogram they span — useful for area
calculations and finding a direction perpendicular to two given directions (e.g. a normal to
a plane, needed in some contexts). Resolving a vector into components along and
perpendicular to another vector (projections) uses both products together.
Formulas
v·w=|v||w|cosθ |v×w|=|v||w|sinθ
Component of a along b: (a·b)/|b|. Component of a perpendicular to b: |a×b|/|b|.
Worked example 1 Given a=(2,1,−1), b=(1,−1,2), find the angle between them.
a·b=2−1−2=−1. |a|=√6, |b|=√6. cosθ=−1/6 ⇒ θ≈99.6°.
Worked example 2 Find a×b for the vectors above. a×b = (1×2−(−1)(−1), (−1)(1)−2×2,
2×(−1)−1×1) = (1,−5,−3).
Common errors: confusing the dot product (gives a scalar) with the cross product (gives a
vector) — check what type of answer the question expects; sign/order errors in the cross-
product component formula (it is not symmetric — a×b=−(b×a)).

AHL 3.14 — Graph Theory Basics


Theory A graph (in this sense — a network, not a function plot) is a set of vertices (nodes)
connected by edges. The degree of a vertex counts how many edges touch it. A “simple”
graph has no loops (edge from a vertex to itself) and no multiple edges between the same
pair of vertices; a complete graph has every pair of vertices connected. Directed graphs
give edges a direction, splitting degree into in-degree and out-degree separately. A tree is a
connected graph with no cycles — exactly the structure needed for a minimum spanning
tree (AHL 3.16).
Worked example A graph has vertices A, B, C, D and edges AB, AC, BC, BD. Find the degree
of each vertex. deg(A)=2 (to B,C), deg(B)=3 (to A,C,D), deg(C)=2 (to A,B), deg(D)=1 (to B).
Sum of degrees = 8 = 2×(number of edges) — this always holds (handshake lemma), since
every edge contributes to exactly two vertices’ degree counts.
Common errors: forgetting a loop contributes 2 to a vertex’s degree, not 1; miscounting
degree in a directed graph by not separating in- and out-degree.

AHL 3.15 — Adjacency Matrices, Walks, Transition Matrices


Theory An adjacency matrix is a compact way to record a graph’s structure: entry (i,j)
records the number of edges directly connecting vertex i and vertex j. Its real power comes
from matrix powers: raising the adjacency matrix to the kth power gives, in each entry, the
number of distinct walks of exactly length k between the corresponding pair of vertices —
a genuinely useful computational shortcut for counting routes. Weighted adjacency tables
replace the “number of edges” entries with costs/distances/times for optimisation
problems. When the entries are converted to probabilities (each column summing to 1), the
matrix becomes a transition matrix — the foundation of Markov chains (AHL 4.19) and
search-ranking algorithms like PageRank.
Worked example For the triangle graph with adjacency matrix A=(0,1,1; 1,0,1; 1,1,0), find
A² and interpret one entry. A² = (2,1,1; 1,2,1; 1,1,2). The entry A²₁₁=2 means there are 2
walks of length 2 from vertex 1 back to itself (via vertex 2, and via vertex 3).
Common errors: confusing “number of walks” (which can revisit vertices/edges) with
“number of paths” (which cannot); building the adjacency matrix asymmetrically for an
undirected graph (it must be symmetric).
AHL 3.16 — Graph Algorithms
Theory This subtopic is about efficient real-world routing/network problems, each with a
precise vocabulary and matching algorithm: - A walk may repeat vertices/edges; a trail
repeats no edges; a path repeats no vertices; a circuit/cycle returns to the start. - An
Eulerian trail (traverses every edge exactly once) exists exactly when 0 or 2 vertices have
odd degree (0 → circuit possible; 2 → trail between those two vertices, but not a circuit). - A
Hamiltonian path/cycle visits every vertex exactly once — no simple degree test exists for
this, unlike the Eulerian case. - A minimum spanning tree connects all vertices with
minimum total edge weight and no cycles. Kruskal’s algorithm: repeatedly add the
cheapest remaining edge that doesn’t create a cycle. Prim’s algorithm: grow a tree from
one starting vertex, always adding the cheapest edge connecting the tree to a new vertex. -
The Chinese Postman Problem finds the minimum-weight route traversing every edge at
least once (for a delivery/inspection route): pair up the odd-degree vertices (there will be
an even number of them) to minimise the extra distance retraced, then duplicate those
shortest connecting paths. - The Travelling Salesman Problem finds the minimum-weight
Hamiltonian cycle. Since this is computationally very hard to solve exactly for large graphs,
two GDC-friendly technique give bounds: the nearest-neighbour algorithm (always move
to the closest unvisited vertex) gives an upper bound; the deleted-vertex algorithm
(delete a vertex, find the MST of what remains, add back the two cheapest edges from the
deleted vertex) gives a lower bound.
Worked example 1 (Kruskal’s algorithm) Edges (weight): AB(2), AC(3), BC(1), BD(4),

Add BC ✓ (weight 1). Add AB ✓ (weight 2). Skip AC (would form cycle A-B-C). Add BD ✓
CD(5). Find the minimum spanning tree. Sort by weight: BC(1), AB(2), AC(3), BD(4), CD(5).

(weight 4). MST = {BC, AB, BD}, total weight = 1+2+4 = 7.


Worked example 2 (Eulerian check) A graph has vertices with degrees 4, 3, 3, 2, 2. Does
it have an Eulerian trail, circuit, or neither? Exactly two vertices (the two degree-3 vertices)
have odd degree ⇒ an Eulerian trail exists (starting and ending at the two odd-degree
vertices), but not a circuit.
Common errors: running Kruskal’s without checking each candidate edge would create a
cycle; confusing the upper bound (nearest-neighbour) and lower bound (deleted-vertex)
algorithms for TSP — they are not interchangeable and give different types of estimate.

Topic 4: Statistics and Probability


SL 4.1 — Sampling and Data Concepts
Theory Statistics is only as trustworthy as the sample it’s built from. A population is
everyone/everything of interest; a sample is the subset actually measured, ideally chosen
so that it represents the population without bias. Different sampling techniques trade off
convenience against representativeness: simple random sampling (every member has
equal chance of selection) is the gold standard but often impractical; convenience sampling
(whoever’s easiest to reach) is fast but prone to serious bias; systematic sampling (every
kth member from a list) and quota sampling (fixed numbers from each subgroup, chosen
non-randomly) sit in between; stratified sampling divides the population into subgroups
(strata) and samples from each proportionally to its size — this guarantees fair
representation of every subgroup and is the technique most commonly tested numerically.
An outlier is a data point unusually far from the rest (formally, more than 1.5×IQR beyond
the nearer quartile) — it may be a genuine extreme value or a data-entry/measurement
error, and deciding which requires judgement, not just the formula.
Formulas Stratified sample size from stratum = (stratum size / population size) × total
sample size.
Worked example 1 A school of 800 students is stratified-sampled by year group to get a
total sample of 80. Year 12 has 250 students. Find the Year 12 sample size. (250/800)×80 =
25 students.
Worked example 2 A survey is conducted by phoning every 10th name on an alphabetical
class list. What sampling method is this, and what is one limitation? This is systematic
sampling. Limitation: if the list has a hidden pattern related to every 10th entry (unlikely
here, but possible in other contexts), the sample could be biased; more generally it excludes
anyone not on the original list.
Common errors: rounding a stratified sample size before it’s needed, causing totals not to
sum correctly; mislabelling convenience sampling as random sampling.

SL 4.2 — Presenting Data


Theory A histogram groups continuous data into equal-width intervals and displays
frequency as bar height — unlike a bar chart, bars touch because the underlying variable is
continuous. A cumulative frequency graph plots the running total of frequency against the
upper boundary of each class, and its shape (an S-curve) lets you read off the median (at
cumulative frequency n/2) and quartiles (at n/4 and 3n/4) graphically, without needing
the raw data. A box-and-whisker plot condenses the five-number summary (minimum, Q1,
median, Q3, maximum) into a compact visual, with outliers marked separately as individual
crosses/points beyond the whiskers — comparing box plots side-by-side is the standard
way to compare two distributions’ centre and spread at a glance, and a strongly
asymmetric box-and-whisker shape is a visual clue the underlying distribution is not close
to normal.
Formulas
IQR=Q₃-Q₁ Outlier if x<Q₁-1.5 IQR or x>Q₃+1.5 IQR
Worked example From a cumulative frequency graph of 200 data values, read off the
median and IQR. Median at cumulative frequency 100; Q1 at 50; Q3 at 150. Suppose these
read as 42, 35, 51 respectively: median=42, IQR=51−35=16.
Common errors: reading cumulative frequency graphs at n/4, n/2, 3n/4 of the frequency
axis rather than matching those values back to the correct x-axis value; forgetting outliers
are plotted individually, not included in the whisker length.

SL 4.3 — Central Tendency and Dispersion


Theory Mean, median, and mode each answer “what’s typical?” differently — the mean
uses every value (sensitive to outliers/skew), the median is the middle value when ordered
(robust to outliers), and the mode is the most frequent value (works even for non-
numeric/categorical data). For grouped (interval) data, the exact mean can’t be calculated
since individual values are unknown, so we estimate it using each interval’s midpoint as a
stand-in for every value in that interval. Standard deviation measures typical spread from
the mean; because it’s calculated from squared deviations, it’s always found via technology
in this course rather than by the raw formula. A crucial, frequently-tested idea: linear
transformations of data (adding a constant, or scaling by a constant) affect the mean and
standard deviation in predictable, different ways.
Formulas Adding k to every value: new mean = old mean + k; new σ = old σ (unchanged —
spread doesn’t change under a shift). Multiplying every value by k: new mean = k×old
mean; new σ = |k|×old σ (spread scales too).
Worked example 1 Data has mean 40, σ=5. Every value has 3 subtracted, then is doubled.
Find the new mean and σ. New mean = (40−3)×2 = 74. New σ = 5×2 = 10 (the subtraction
of 3 doesn’t affect σ; only the ×2 scaling does).
Worked example 2 Grouped data: intervals [0,10), [10,20), [20,30) with frequencies 5, 8,
7 respectively. Estimate the mean. Midpoints: 5, 15, 25. Estimated mean =
(5×5+8×15+7×25)/20 = (25+120+175)/20 = 16.
Common errors: applying the “×k” scaling rule to σ when only a shift (+k) was applied (σ
shouldn’t change under a pure shift); using interval endpoints instead of midpoints when
estimating a grouped mean.

SL 4.4 — Correlation and Linear Regression


Theory Pearson’s correlation coefficient r measures the strength and direction of a linear
relationship only — a strong curved relationship can have a low r even though the
variables are clearly related, which is a key conceptual trap. The regression line of y on x is
the specific “line of best fit” that minimises the sum of squared vertical distances from the
data points (least-squares) — critically, this is only appropriate for predicting y from x, not
the reverse, because it was built to minimise error in the y-direction specifically. Every
regression line passes through the mean point (x̄, ȳ). As always: correlation is not
causation (a strong r doesn’t prove one variable causes the other — both could be driven
by a third factor, or the association could be coincidental), and predicting well outside the
range of the original data (extrapolation) is unreliable since the linear relationship may not
continue to hold.
Formulas
-1≤ r≤1 y=ax+b (least-squares regression line, via technology)
Worked example 1 Hours studied vs test score gives r=0.87 and regression line
y=4.2x+52. Predict the score for 6 hours studied, and comment on predicting for 40 hours.
y=4.2(6)+52=77.2. Predicting for 40 hours would be unreliable extrapolation — well
beyond the range of the original study-hours data, where the linear trend may not continue
(e.g. diminishing returns, or a maximum possible score).
Worked example 2 A dataset has r=−0.92. Describe the relationship. A strong, negative
linear correlation — as one variable increases, the other tends to decrease markedly, in an
approximately straight-line pattern.
Common errors: using the y-on-x line to predict x from a given y (need a separate x-on-y
regression for that, not covered at SL, or simply note it’s inappropriate); interpreting r≈0 as
“no relationship” when a strong non-linear relationship may still exist (always look at the
scatter diagram too).

SL 4.5 — Probability Basics


Theory Probability quantifies how likely an event is, from 0 (impossible) to 1 (certain), as
the proportion of a sample space’s outcomes that satisfy the event (assuming equally likely
outcomes) — or, over many repeated trials, the long-run relative frequency. The
complement rule (P(A’)=1−P(A)) is often the fastest route to an answer when “at least one”
or “not” language appears, since it’s frequently far easier to compute the probability of the
opposite event.
Formulas
P(A)=(n(A))/(n(U)) P(A’)=1-P(A) Expected occurrences=n× P
Worked example 1 A class of 128 students has P(absent on a given day)=0.1. Find the
expected number of absentees. 128×0.1=12.8.
Worked example 2 A bag has 12 balls, 5 of which are red. Find P(not red). P(red)=5/12, so
P(not red)=1−5/12=7/12.
Common errors: forgetting expected value can be a non-integer (it’s a long-run average,
not a specific outcome); confusing “at least one” with “exactly one” when applying the
complement rule.
SL 4.6 — Combined and Conditional Probability
Theory Venn diagrams, tree diagrams, and sample-space (grid) diagrams are three
different visual tools for the same underlying combined-probability logic, and choosing the
right one for a given problem is half the battle. The addition rule accounts for double-
counting when two events can happen together (subtract the overlap); mutually exclusive
events (which can never happen together) simplify this since the overlap is zero.
Conditional probability P(A|B) restricts the sample space to only the outcomes where B has
already happened — this is the natural language of tree diagrams, where branch
probabilities are conditional probabilities, and multiplying along a branch gives a joint
(AND) probability. Independence is a special case where knowing B happened doesn’t
change the probability of A at all.
Formulas
P(A∪ B)=P(A)+P(B)-P(A∩ B) P(A|B)=(P(A∩ B))/(P(B))
Independent: P(A∩ B)=P(A)P(B)
Worked example 1 A bag has 5 red, 3 blue balls. Two are drawn without replacement.
Find P(both red). P(both red) = (5/8)×(4/7) = 5/14.
Worked example 2 P(A)=0.4, P(B)=0.5, P(A∩B)=0.2. Find P(A∪B) and determine if A, B
are independent. P(A∪B)=0.4+0.5−0.2=0.7. Independence check: P(A)×P(B)=0.2=P(A∩B),
so A and B are independent.
Common errors: using “with replacement” probabilities (unchanged denominators) for a
“without replacement” scenario, or vice versa; forgetting to subtract the overlap in the
addition rule for non-mutually-exclusive events.

SL 4.7 — Discrete Random Variables


Theory A discrete random variable’s full behaviour is captured by its probability
distribution — a table (or formula) listing every possible value and its probability, which
must sum to exactly 1 (a useful check, and often the key to finding an unknown parameter).
The expected value E(X) is a weighted average of the possible outcomes, weighted by how
likely each is — it represents the long-run average result over many repetitions, not
necessarily a value X can actually take. In a “fair game” context (X = a player’s net gain),
E(X)=0 is precisely the mathematical definition of fairness — neither side is expected to
profit long-run.
Formulas
Σ P(X=x)=1 E(X)=Σ xP(X=x)

distribution and find E(X). P(1)=5/18, P(2)=6/18, P(3)=7/18; sum=18/18=1 ✓ valid. E(X)
Worked example 1 X takes values 1, 2, 3 with P(X=x) = (4+x)/18. Verify this is a valid

= 1(5/18)+2(6/18)+3(7/18) = 38/18 ≈ 2.11.


Worked example 2 A game costs 2 to play; you win 10 with probability 0.1, otherwise
nothing. Find the expected net gain and state whether the game is fair. Net gain X: +8 with
p=0.1, −2 with p=0.9. E(X)=8(0.1)+(−2)(0.9)=0.8−1.8=−1 — not fair (the player is expected
to lose 1 per game on average).
Common errors: forgetting to check probabilities sum to 1 before trusting a
distribution/solving for an unknown constant within it; computing E(X) using outcome
values instead of net gain when a “cost to play” is involved.

SL 4.8 — Binomial Distribution


Theory The binomial distribution models the number of “successes” in a fixed number of
independent trials, each with the same two possible outcomes and the same success
probability — the classic scenario is repeated coin flips or any repeated yes/no experiment
(n fixed trials, constant p, independence between trials — all four conditions matter for the
model to be valid). Its mean and variance follow simple closed forms; individual
probabilities are calculated using the GDC’s binomial probability function rather than the
underlying combinatorial formula by hand.
Formulas
X~ B(n,p), E(X)=np, Var(X)=np(1-p)
Worked example 1 X~B(10, 0.3). Find E(X), Var(X), and P(X=4) using technology.
E(X)=10(0.3)=3. Var(X)=10(0.3)(0.7)=2.1. P(X=4)≈0.2001 (GDC).
Worked example 2 A factory has a 5% defect rate. In a sample of 60 items, find the
expected number of defects and P(at most 2 defective). X~B(60,0.05). E(X)=60(0.05)=3.
P(X≤2) via GDC (cumulative binomial) ≈ 0.4200.
Common errors: applying the binomial model when trials aren’t actually independent
(e.g. sampling without replacement from a small population — technically hypergeometric,
though the binomial is often used as an approximation when the population is large);
mixing up P(X=k) with P(X≤k) on the GDC (these are different calculator functions).

SL 4.9 — Normal Distribution


Theory The normal distribution is the symmetric, bell-shaped model that describes
countless natural and measured quantities clustering around a mean, with symmetric
tapering probability further away — it’s defined by just two parameters, the mean μ
(centre) and standard deviation σ (spread). The empirical rule gives quick sanity-check
benchmarks for roughly what proportion of data falls within 1, 2, or 3 standard deviations
of the mean. In exams, both “forward” problems (given x, find a probability) and “inverse”
problems (given a probability/percentile, find x) are solved directly with the GDC’s normal
distribution functions — no z-table lookups are required, though understanding the
standardised z-score conceptually helps interpret results.
Formulas
X~ N(μ,σ²)
Empirical rule: ≈68% within μ±σ, ≈95% within μ±2σ, ≈99.7% within μ±3σ.
Worked example 1 X~N(100,15²) (IQ scores). Find P(X>120) and the score
corresponding to the top 10%. P(X>120) ≈ 0.0912 (GDC normal CDF). Top 10%
(i.e. P(X<x)=0.90): inverse normal gives x≈119.2.
Worked example 2 Heights of adult men are N(175, 7²) cm. Find the proportion between
168 cm and 182 cm. This range is exactly μ±σ, so by the empirical rule this is
approximately 68% (GDC gives ≈68.3%).
Common errors: confusing “inverse normal” (given probability, find x) with “normal CDF”
(given x, find probability) on the GDC — always identify which direction the question
requires first; forgetting variance is σ² in the notation N(μ,σ²), so σ itself must be found by
taking a square root if variance is given.

SL 4.10 — Spearman’s Rank Correlation Coefficient


Theory Spearman’s rₛ measures the strength of a monotonic relationship (consistently
increasing or consistently decreasing, but not necessarily in a straight line) by first
converting both variables to ranks and then computing a correlation-type measure on
those ranks — this makes it robust to outliers and effective at detecting non-linear-but-
monotonic trends that Pearson’s r (which specifically targets linearity) would understate.
When two data values tie, their ranks are averaged (e.g. two values tied for 3rd/4th place
both get rank 3.5).
Worked example Two judges rank 6 contestants; computing rₛ from their ranks via
technology gives rₛ=0.83. Interpret this. A strong positive association in ranking — the
judges’ relative orderings largely agree, even though the underlying (unranked) scores
might not be linearly related.
Common errors: computing rₛ on the raw data instead of the ranks; forgetting to average
ranks for tied values, which shifts every subsequent rank incorrectly.

SL 4.11 — Hypothesis Testing


Theory Hypothesis testing is a formal procedure for deciding whether observed data
provides convincing evidence against a default assumption (H₀, the null hypothesis) in
favour of an alternative (H₁). The p-value is the probability of seeing data at least as
extreme as what was observed, if H₀ were actually true — a small p-value means the
observed data would be surprising under H₀, giving evidence against it. The decision rule is
simple and mechanical: reject H₀ if p < significance level (commonly 5%); otherwise, don’t
reject it (this is not the same as “proving H₀ true” — merely insufficient evidence against
it). - The χ² test for independence uses a contingency table to test whether two
categorical variables are associated; degrees of freedom = (rows−1)(columns−1). - The χ²
goodness-of-fit test checks whether observed data matches an assumed/expected
distribution; at SL, degrees of freedom = (number of categories − 1). - The t-test compares
the means of two populations from sample data when the population standard deviation
isn’t known (very common in practice) — this course uses the pooled two-sample t-test
assuming equal (but unknown) variances.
Formulas Reject H₀ if p-value < significance level, or if the test statistic exceeds the critical
value.
Worked example 1 (χ² independence) A 2×3 contingency table of gender vs preferred
subject. GDC gives χ²=8.45, df=2, p=0.0146. Test at the 5% level. Since p=0.0146<0.05,
reject H₀: there is significant evidence of an association between gender and subject
preference.
Worked example 2 (two-sample t-test) Two classes’ test scores are compared with a
pooled two-sample t-test; GDC gives p=0.032. Test at the 5% level. Since p=0.032<0.05,
reject H₀: there is significant evidence the two classes’ mean scores differ.
Common errors: stating a conclusion as “H₀ is true” rather than “insufficient evidence to
reject H₀” when p ≥ significance level; using the wrong degrees-of-freedom formula (mixing
up independence vs goodness-of-fit); forgetting expected frequencies below 5 can
invalidate a χ² test.

AHL Content (Topic 4)


AHL 4.12 — Design of Data Collection; Reliability and Validity
Theory Well-designed data collection avoids two related but distinct problems: bias (a
systematic tendency to over/under-represent certain outcomes, e.g. from leading questions
or an unrepresentative sample) and poor reliability/validity. Reliability asks whether a
measurement is consistent (would repeating it give the same result? — tested via test-
retest or parallel-forms methods); validity asks whether it’s actually measuring what it
claims to measure (content validity: does it cover the full concept; criterion validity: does it
correlate with an established outside measure). When continuous numerical data is
grouped into categories for a χ² test, the choice of category boundaries affects expected
frequencies — categories should be chosen so expected frequency in each is not too small
(typically ≥5), and degrees of freedom must be reduced by one for each parameter
estimated from the data itself (rather than assumed).
Common errors: treating “reliable” and “valid” as synonyms (a scale that’s consistently 2
kg wrong is reliable but not valid); forgetting to reduce degrees of freedom when
parameters were estimated from the sample for a goodness-of-fit test.

AHL 4.13 — Non-Linear Regression


Theory Beyond straight lines, technology can fit least-squares regression curves of several
standard families (quadratic, cubic, exponential, power, sine) directly to data — the fitting
principle is the same (minimise the sum of squared residuals, SS_res), just applied to a
curved model. R², the coefficient of determination, tells you what proportion of the
variability in y is explained by the chosen model; for a linear fit, R² is literally r² (Pearson’s
correlation squared), but for non-linear fits R² generalises that idea to any model shape. A
high R² shows the model fits the given data well but does not by itself guarantee the model
is appropriate for the underlying real-world process, especially outside the data’s range —
always sanity-check against the context and the shape of a residual plot.
Formulas
R²=1-(SS_(res))/(SS_(tot))
Worked example Fitting y=ax²+bx+c to data via GDC gives R²=0.982. Interpret this value,
and note one thing it does not tell you. 98.2% of the variability in y is explained by the
quadratic model — a very strong fit to this data. However, a high R² does not confirm the
quadratic form is the “true” underlying relationship (a cubic might fit even better, or the
good fit might not hold for x-values outside the sampled range).
Common errors: treating a high R² as proof of causation or of the “correct” model family;
comparing R² values across models fitted to different datasets (only valid for comparing
models on the same dataset).

AHL 4.14 — Linear Combinations of Random Variables


Theory When a random variable is transformed linearly (aX+b) or several independent
random variables are combined linearly, their expected values and variances follow
predictable rules — expectation is always linear (works even if the variables aren’t
independent), but variance only adds simply across independent variables (since any
correlation would introduce extra covariance terms, which are outside this syllabus). This
underlies the sample mean x̄ as an estimator: it is unbiased for μ (its expected value equals
the true population mean), and the sample variance formula with (n−1) in the denominator
is specifically constructed to be an unbiased estimator of σ².
Formulas
E(aX+b)=aE(X)+b Var(aX+b)=a²Var(X)
x̄ = (1)/(n)Σ x_i (unbiased for μ) s_(n-1)² (unbiased for σ²)
Worked example 1 X has mean 50, variance 16. Y=3X−4. Find E(Y) and Var(Y).
E(Y)=3(50)−4=146. Var(Y)=3²(16)=144.
Worked example 2 X and Y are independent, E(X)=10, E(Y)=6, Var(X)=4, Var(Y)=9. Find
E(X+Y) and Var(X+Y). E(X+Y)=10+6=16. Var(X+Y)=4+9=13 (variances add only because X,
Y are independent).
Common errors: applying Var(aX+b)=a²Var(X)+b (forgetting b vanishes entirely — a
constant shift doesn’t affect spread); adding variances of variables that are not independent
(invalid without a covariance correction, which is outside this course).

AHL 4.15 — Sampling Distributions and the Central Limit Theorem


Theory Every time you take a sample and compute its mean, you get a slightly different
value — the sample mean is itself a random variable, with its own distribution (the
sampling distribution of the mean), which is less spread out than the original population as
n grows (bigger samples give more reliable, less variable estimates of μ). If the original
population is normal, the sample mean’s distribution is exactly normal for any n. The
Central Limit Theorem is the more powerful and more frequently examined fact: even if
the original population is not normal, the sample mean’s distribution becomes
approximately normal once the sample size is reasonably large (n>30 is treated as
sufficient in this course) — this is why so many real-world averages behave approximately
normally regardless of the underlying data’s shape.
Formulas
X̄ ~ N(μ, (σ²)/(n)) (exactly if X is normal; approximately, by the CLT, for large n
regardless)
Worked example A population has mean 70, σ=12 (distribution shape unknown/non-
normal). For samples of size 50, describe the distribution of the sample mean. By the CLT
(n=50>30 is large enough): X̄ is approximately N(70, 144/50), i.e. N(70, 2.88).
Common errors: applying the CLT’s “n>30 is enough” rule as if it makes the sample mean’s
distribution exactly normal rather than approximately so; confusing σ (population SD) with
σ/√n (the standard deviation of the sample mean, sometimes called the standard error) —
these are frequently swapped by mistake in formulas.

AHL 4.16 — Confidence Intervals


Theory A confidence interval gives a plausible range for an unknown population
parameter (here, the mean), built from sample data, together with a stated confidence level
(typically 95%) describing the procedure’s long-run reliability — a 95% confidence interval
means that if you repeated the sampling-and-interval-construction process many times,
about 95% of the resulting intervals would contain the true population mean (it is a
common and important misconception to instead say “there’s a 95% probability μ is in this
particular interval” — μ is fixed, not random; it’s the interval that varies from sample to
sample). The choice between a normal distribution and a t-distribution for constructing the
interval depends only on whether σ (the population SD) is known — if it’s unknown (the
much more common real-world case), the t-distribution is used regardless of sample size,
because estimating σ from the sample itself introduces extra uncertainty that the t-
distribution’s slightly heavier tails account for.
Worked example A sample of n=25 has x̄=100, s=12 (σ unknown). Find a 95% confidence
interval for the population mean. Since σ is unknown, use the t-distribution (GDC t-
interval, df=24): (95.05, 104.95) approximately.
Common errors: using the normal distribution instead of the t-distribution whenever σ is
unknown, regardless of sample size (a very common and serious error); misinterpreting
the confidence level as a probability statement about the fixed population mean rather than
about the interval-construction procedure.

AHL 4.17 — Poisson Distribution


Theory The Poisson distribution models the count of rare, independent events occurring at
a constant average rate over a fixed interval of time or space — classic examples include
calls arriving at a call centre, radioactive decay counts, or typos per page. Unlike the
binomial, there’s no fixed “number of trials”; instead, a single parameter m (the mean rate
for the interval in question) determines the entire distribution, and remarkably, the mean
and variance are always exactly equal for a Poisson variable — this equality is itself a useful
diagnostic for whether Poisson is an appropriate model for given data. A convenient
additive property: summing two independent Poisson variables (e.g. combining two
separate time intervals, or two independent sources of events) gives another Poisson
variable, with parameters simply added.
Formulas
XsimPo(m), E(X)=Var(X)=m
Worked example 1 Calls arrive at a rate of 4 per hour, X~Po(4). Find P(X=6) using
technology, and the distribution over a 2-hour period. P(X=6)≈0.1042 (GDC). Over 2 hours:
Y~Po(8).
Worked example 2 Typing errors occur at a rate of 2 per page. Find the probability of at
most 1 error on a given page. X~Po(2). P(X≤1) via GDC (cumulative Poisson) ≈ 0.4060.
Common errors: using a Poisson model when events aren’t actually independent or the
rate isn’t actually constant (e.g. rush-hour call volume varying strongly by time of day
violates the constant-rate assumption); forgetting to rescale the parameter m when the
time/space interval in the question changes (e.g. from “per hour” to “per 2 hours”).
AHL 4.18 — Further Hypothesis Testing
Theory This extends SL 4.11’s hypothesis-testing framework across more scenarios:
testing a population mean against a known or unknown σ (normal/t-test), testing a
population proportion (using the binomial/normal approximation), testing the mean of a
Poisson-distributed population (always one-tailed, since Poisson means are naturally
bounded below by 0), and testing whether a population correlation coefficient ρ is
genuinely zero (i.e., whether an observed sample correlation is strong enough to be
statistically significant rather than a fluke of sampling). A paired t-test (e.g. before/after
measurements on the same subjects) is analysed as a single-sample test on the differences,
which is more powerful than treating the two sets as independent samples, because it
removes subject-to-subject variability from the comparison. Every hypothesis test carries
two possible error types: a Type I error (rejecting a true H₀ — a “false alarm,” with
probability exactly equal to the chosen significance level α) and a Type II error (failing to
reject a false H₀ — a “missed detection,” with probability β, related to the test’s statistical
power = 1−β).
Worked example Testing H₀: μ=50 vs H₁: μ>50 at 5% significance, with σ=10 known,
n=36. If the sample mean is 53.5, find the critical region and state the conclusion. Critical
region: x̄ > 50 + 1.645×(10/√36) = 50+2.74 = 52.74. Since 53.5>52.74, reject H₀:
significant evidence the mean exceeds 50.
Common errors: using a two-tailed critical value/p-value for a clearly one-tailed
hypothesis (or vice versa); treating a paired design as if the two samples were independent
(loses statistical power and can even reverse a conclusion); confusing Type I and Type II
errors’ definitions.

AHL 4.19 — Markov Chains and Transition Matrices


Theory A Markov chain models a system that moves between a finite set of “states” over
discrete time steps, where the probability of moving to each next state depends only on the
current state (not on the full history of how it got there) — the “memoryless” property. All
these transition probabilities are captured in a single transition matrix T; repeatedly
multiplying an initial state vector by T steps the system forward in time. For many such
systems (“regular” chains), the state vector converges to a steady state as time goes on,
regardless of the starting state — found either by repeatedly multiplying (numerically) or,
more elegantly, by recognising the steady state is the eigenvector of T corresponding to
eigenvalue 1 (connecting directly back to AHL 1.15).
Formulas
s_n=Tⁿs₀ Ts=s (steady state)
Worked example Weather model: if sunny today, P(sunny tomorrow)=0.7; if rainy,
P(rainy tomorrow)=0.6. Find the long-run proportion of sunny days. T=(0.7,0.4; 0.3,0.6).
Steady state: 0.7s₁+0.4s₂=s₁ and s₁+s₂=1 ⇒ 0.3s₁=0.4s₂ ⇒ s₁=4/7, s₂=3/7. Long-run,
≈57.1% of days are sunny.
Common errors: forgetting the state vector’s entries must sum to 1 at every stage (a
normalisation check); setting up the transition matrix with rows and columns transposed
relative to the convention being used (always double-check which index — row or column
— represents “from” vs “to”).

Topic 5: Calculus
SL 5.1 — Introduction to Limits and the Derivative
Theory A limit describes the value a function approaches as the input approaches some
point — even if the function isn’t actually defined right at that point. This course treats
limits informally, estimated numerically from a table of values getting closer and closer to
the target x, rather than through formal analytic limit techniques. The derivative is defined
as a limit of average rates of change over shrinking intervals — informally, it’s the
“instantaneous” gradient of a curve at a point, i.e. the gradient of the tangent line there.
This is the single most important conceptual leap in the whole course: gradient stops being
a single fixed number (as for a line) and becomes a function itself, f’(x), that can be
evaluated at any point.
Worked example Estimate lim(x→2) (x²−4)/(x−2) using a table of values. Values at x=1.9,
1.99, 2.01, 2.1 give function values 3.9, 3.99, 4.01, 4.1 — approaching 4. (Confirmed
algebraically: (x−2)(x+2)/(x−2) = x+2 → 4 as x→2.)
Common errors: trying to substitute x=2 directly into a 0/0 expression and concluding the
limit “doesn’t exist” without checking the actual approaching behaviour; confusing an
average rate of change (a slope between two points) with the instantaneous rate (a limit).

SL 5.2 — Increasing/Decreasing Functions


Theory The sign of the derivative directly tells you the shape of the original function:
positive f’(x) means the function is rising, negative means falling, and zero means
momentarily flat — a stationary point. This single idea (checking the sign of f’) is the engine
behind almost every curve-sketching and optimisation problem in this topic.
Worked example Given f’(x) = x²−4, determine the intervals where f is
increasing/decreasing. f’(x)=0 at x=±2. Testing signs: f’(x)>0 for x<−2 and x>2 (increasing
there); f’(x)<0 for −2<x<2 (decreasing there).
Common errors: analysing the sign of f(x) instead of f’(x) when determining
increasing/decreasing behaviour.
SL 5.3 — Derivative of Polynomials
Theory The power rule (bring the exponent down as a multiplying factor, then reduce the
exponent by 1) is derived from the limit definition but is used directly, term by term, for
any polynomial — this is the single most-used mechanical skill in the whole calculus topic
and needs to be completely automatic before moving on to more complex rules.
Formula
f(x)=axⁿ ⇒ f’(x)=anxⁿ⁻¹
Worked example 1 f(x)=3x⁴−2x²+5x−7. Find f’(x). f’(x) = 12x³−4x+5.
Worked example 2 Find the gradient of y=2x³−x at x=−1. y’=6x²−1; at x=−1: 6(1)−1=5.
Common errors: forgetting the derivative of a constant term is 0 (not the constant itself);
applying the power rule to a term that isn’t a pure power of x (e.g. inside a product or
composition — those need the rules in AHL 5.9).

SL 5.4 — Tangents and Normals


Theory At any point on a curve, the tangent line touches the curve and has gradient exactly
equal to the derivative evaluated there; the normal line is perpendicular to the tangent at
that same point, so its gradient is the negative reciprocal — this directly reuses the
perpendicular-line relationship from SL 2.1 and SL 3.5, now with the derivative supplying
the gradient instead of two given points.
Formulas Tangent at (x₁,y₁): y−y₁ = f’(x₁)(x−x₁). Normal at (x₁,y₁): y−y₁ = −1/f’(x₁) ×
(x−x₁).
Worked example f(x)=x²−3x. Find the equations of the tangent and normal at x=2.
f(2)=−2; f’(x)=2x−3 ⇒ f’(2)=1. Tangent: y+2=1(x−2) ⇒ y=x−4. Normal: y+2=−1(x−2) ⇒
y=−x.
Common errors: using f(x₁) as the gradient instead of f’(x₁); forgetting the normal’s
gradient is the negative reciprocal, not just the negative, of the tangent’s gradient.

SL 5.5 — Integration as Anti-Differentiation


Theory Integration reverses differentiation — given a derivative, find the original function.
Because differentiation destroys constant terms (their derivative is 0), reversing it can
never fully recover the original constant, so every indefinite integral carries an unknown
“+c”; a specific boundary condition (a known point on the original curve) is needed to pin
down c’s exact value. Definite integrals (with specific limits) instead compute a specific
numeric value — geometrically, the (signed) area between the curve and the x-axis over
that interval — and are evaluated using technology in this course.
Formulas
∫ axⁿ dx=(a)/(n+1)xⁿ⁺¹+c (n≠ -1)
Worked example 1 dy/dx = 3x²+x, and y=10 when x=1. Find y. y = x³ + x²/2 +
c. Substituting: 10 = 1+0.5+c ⇒ c=8.5. y = x³ + x²/2 + 8.5.
Worked example 2 Evaluate ∫₂⁶ (3x²+4) dx. [x³+4x]₂⁶ = (216+24)−(8+8) = 224.
Common errors: forgetting the “+c” on an indefinite integral (loses marks even if the rest
is correct); applying the power rule for integration to n=−1 (undefined by this formula —
that case needs ln|x|, covered in AHL 5.11).

SL 5.6 — Stationary Points


Theory Stationary (turning) points occur exactly where f’(x)=0 — geometrically, the
tangent is momentarily horizontal there. This course solves f’(x)=0 using technology rather
than always factorising by hand. A vital distinction: a local maximum or minimum is only
the highest/lowest point in its immediate neighbourhood — it is not necessarily the overall
(global) maximum/minimum of the function across its entire domain, especially for
functions without a restricted domain (a cubic, for instance, has no global max/min at all).
Worked example f(x)=x³−3x²−9x+5. Find and classify the stationary points.
f’(x)=3x²−6x−9=0 ⇒ x=−1, 3. f(−1)=10 (a local maximum — nearby values are lower on
both sides), f(3)=−22 (a local minimum).
Common errors: stating a local extremum is automatically the global extremum without
checking the function’s behaviour further out (or its domain restrictions).

SL 5.7 — Optimisation
Theory Optimisation problems translate a real-world “best” scenario (maximum volume,
minimum cost, maximum profit) into a function of a single variable, then find its stationary
point(s) using derivatives — the calculus itself is identical to SL 5.6, but the setup
(expressing the quantity to optimise in terms of one variable, using given constraints to
eliminate other variables) is usually the harder half of the problem. Always check the found
stationary point actually corresponds to a sensible answer in context (e.g. reject negative
lengths, or values making the constructed object physically impossible).
Worked example An open box is made from a 20 cm square of card, cutting corner
squares of side x and folding up the sides. Find x that maximises the box’s volume. V(x) =
x(20−2x)². Expand or use the product rule: V’(x) = (20−2x)² + x·2(20−2x)(−2) = (20−2x)
(20−6x) = 0 ⇒ x=10 or x=10/3. x=10 gives a box of zero volume (the sides fold flat —
rejected). x=10/3 ≈ 3.33 cm gives the maximum volume ≈ 592.6 cm³.
Common errors: forgetting to reduce the problem to a single variable before
differentiating (a two-variable expression can’t be optimised this way without first using a
constraint to eliminate one variable); failing to reject solutions that are mathematically
valid but physically meaningless in context.

SL 5.8 — Trapezoidal Rule


Theory When a function’s exact antiderivative is hard or impossible to find (or when only
data points are available, not a formula), a definite integral (area) can be approximated by
dividing the region into trapezia rather than rectangles, averaging the heights at each end
of each strip — this generally gives a much better approximation than rectangle-based
methods for the same number of strips, since trapezia follow the curve’s slope rather than
assuming it’s flat within each strip.
Formula
∫_a^b y dx ≈ (h)/(2)[y₀+y_n+2(y₁+…+y_(n-1))], h=(b-a)/(n)
Worked example Estimate ∫₀⁴ x² dx using the trapezoidal rule with 4 strips (h=1), and
compare to the exact value. y-values at x=0,1,2,3,4: 0, 1, 4, 9, 16. Estimate = (1/2)
[0+16+2(1+4+9)] = (1/2)(44) = 22. Exact value = 64/3 ≈ 21.33 — the trapezoidal rule
slightly overestimates here because y=x² is convex (curves upward, so the straight
trapezoid edges sit above the curve).
Common errors: using the wrong count of “middle” terms (all y-values except the first and
last are doubled — miscounting these is the most common slip); using an incorrect strip
width h (h = (b−a)/n, not (b−a)).

AHL Content (Topic 5)


AHL 5.9 — Derivatives of Standard Functions; Chain, Product, Quotient Rules
Theory Beyond polynomials, the derivatives of the standard transcendental functions (trig,
exponential, logarithmic) must be memorised — they’re the “atoms” that combine, via
three combination rules, to differentiate almost anything examinable. The chain rule
handles composite functions (a function of a function) — differentiate the “outer” function,
leaving the inner alone, then multiply by the derivative of the inner function; this is by far
the most-used rule and the most common source of errors when forgotten. The product
rule handles a product of two functions of x (neither constant): differentiate one, keep the
other unchanged, add the mirror-image term. The quotient rule does the analogous thing
for division, with an extra minus sign and a squared denominator. Related rates problems
chain two or more rates of change together via a shared variable using the chain rule
structurally — the two rates are connected because both variables depend on a common
third variable, often time.
Formulas
(d)/(dx)sin x=cos x, (d)/(dx)cos x=-sin x, (d)/(dx)tan
x=sec²x, (d)/(dx)ex=ex, (d)/(dx)ln x=frac1x
Chain rule: dy/dx = (dy/du)(du/dx). Product rule: (uv)’ = u’v+uv’. Quotient rule: (u/v)’ =
(u’v−uv’)/v².
Worked example 1 Differentiate y = sin(3x²). Chain rule: outer derivative cos(3x²), inner
derivative 6x. y’ = 6x cos(3x²).
Worked example 2 (related rates) A spherical balloon’s volume increases at 10 cm³/s.
Find the rate of change of its radius when r=5 cm. V=(4/3)πr³ ⇒ dV/dr=4πr²=100π (at
r=5). dr/dt = (dV/dt)/(dV/dr) = 10/100π ≈ 0.0318 cm/s.
Common errors: forgetting to multiply by the inner derivative when using the chain rule
(the single most common calculus error in the whole course); mixing up the order of
terms/sign in the quotient rule.

AHL 5.10 — The Second Derivative


Theory Differentiating the derivative itself gives the second derivative, f’‘(x), which
measures how the gradient is changing — i.e. the curve’s concavity. A positive second
derivative means the curve bends upward (concave up, like a smile — any stationary point
here is a local minimum); negative means it bends downward (concave down, a local
maximum). This gives a faster alternative to the “sign change either side” test from SL 5.6
for classifying stationary points, though it’s inconclusive when f’’=0 exactly. A point of
inflexion is where concavity itself changes sign — the curve switches from bending one
way to bending the other.
Worked example f(x)=x³−3x². Classify the stationary points and find any inflexion point.
f’(x)=3x²−6x=0 ⇒ x=0, 2. f’‘(x)=6x−6. f’‘(0)=−6<0 ⇒ local maximum at x=0. f’‘(2)=6>0 ⇒
local minimum at x=2. f’’(x)=0 at x=1, and concavity does change sign there ⇒ point of
inflexion at x=1.
Common errors: using the second derivative test when f’‘=0 (inconclusive — must fall
back to the sign-change method instead); forgetting a point of inflexion requires an actual
change in concavity, not just f’’=0.
AHL 5.11 — Further Integration; Integration by Inspection/Substitution
Theory The standard antiderivatives extend the standard derivatives list in reverse, with
one crucial addition: since d/dx(ln x)=1/x, integrating x⁻¹ gives ln|x| (the absolute value is
needed since ln is only defined for positive inputs, but x⁻¹ makes sense for negative x too).
“Integration by inspection” (a special, quick case of substitution) works when the integrand
is exactly of the form f(g(x))·g’(x) — recognising this pattern (an “inner function”
multiplied by its own derivative) lets you write down the antiderivative directly by
reversing the chain rule, without a full substitution write-up.
Formulas
∫ x⁻¹ dx=ln|x|+c ∫sin x dx=-cos x+c ∫cos x dx=sin x+c ∫ e^x dx=e^x+c
Worked example 1 Find ∫4x sin(x²) dx. Let u=x², du=2x dx, so 4x dx = 2 du. ∫2sin(u) du =
−2cos(u)+c = −2cos(x²)+c.
Worked example 2 Find ∫ sinx/cosx dx. Recognise this as −(d/dx[cos x])/cos x, i.e. of the
form −g’(x)/g(x): −ln|cos x|+c.
Common errors: attempting “inspection” when the derivative of the inner function is not
actually present as a factor (a full substitution or different technique is needed instead);
dropping the absolute value bars on ln|x| results.

AHL 5.12 — Areas and Volumes of Revolution


Theory A definite integral gives signed area — where the curve dips below the x-axis, the
integral there is negative, so finding a true (positive) total area between a curve and the
axis over a region that crosses the axis requires splitting the integral at the crossing
point(s) and taking the absolute value of any negative piece separately, rather than
integrating straight through. Rotating a curve fully around an axis sweeps out a 3D solid of
revolution; slicing this solid into thin circular discs perpendicular to the axis of rotation,
each with radius equal to the curve’s y-value (or x-value) at that point, and integrating their
areas (πy² or πx²) gives the total volume.
Formulas
V=∫_a^b π y² dx (about the x-axis) V=∫_a^b π x² dy (about the y-axis)
Worked example y=√x is rotated about the x-axis from x=0 to x=4. Find the resulting
volume. V=∫₀⁴ π(√x)² dx = π∫₀⁴ x dx = π[x²/2]₀⁴ = 8π.
Common errors: forgetting to square y before integrating for a volume of revolution (a
very common slip — the disc’s area is πr², i.e. π(y)², not πy); not splitting a signed-area
integral at an x-axis crossing when a total (unsigned) area is requested.
AHL 5.13 — Kinematics
Theory Displacement, velocity, and acceleration form a calculus chain: velocity is the
derivative of displacement, acceleration is the derivative of velocity (so the second
derivative of displacement) — and integrating reverses each step. A crucial distinction
tested repeatedly: displacement (net position change, from ∫v dt directly — can be
negative, cancelling out motion in opposite directions) versus total distance travelled
(always positive, requiring you to integrate |v(t)|, which in practice means splitting the
interval wherever v(t) changes sign and taking the absolute value of each piece separately).
Speed is simply |v| at an instant.
Formulas
v=(ds)/(dt), a=(dv)/(dt) Displacement=∫(t₁)^(t₂)v dt, Total distance=∫(t₁)^(t₂)|v|
dt
Worked example v(t)=t²−4t (m/s) for 0≤t≤5. Find the displacement and total distance
travelled. v(t)=0 at t=0, 4; v<0 on (0,4), v>0 on (4,5). Displacement = ∫₀⁵(t²−4t)dt =
[t³/3−2t²]₀⁵ = 125/3−50 = −25/3 ≈ −8.33 m. Total distance = |∫₀⁴v dt| + ∫₄⁵v dt = 32/3 +
7/3 = 13 m.
Common errors: computing total distance by simply integrating v(t) over the whole
interval without splitting at sign changes (this instead gives displacement, which can be
smaller in magnitude due to cancellation); confusing speed (always ≥0) with velocity (can
be negative).

AHL 5.14 — Setting Up and Solving Differential Equations (Separation of


Variables)
Theory A differential equation describes how a quantity’s rate of change relates to the
quantity itself (or to other variables) — the first skill is translating a verbal description
(“the rate of growth is proportional to the current amount”) directly into symbols (dG/dt =
kG). Separable differential equations are those where all the y-terms can be algebraically
gathered on one side and all the x-terms on the other, after which each side is integrated
independently — this only works when such a separation is actually possible (not every
differential equation is separable). The resulting “general solution” contains an arbitrary
constant, representing an entire family of curves satisfying the equation; a specific initial
condition picks out one particular member of that family.
Worked example Solve dy/dx = ky (the fundamental exponential-growth/decay equation)
by separation of variables. ∫(1/y) dy = ∫k dx ⇒ ln|y| = kx+C ⇒ y = Ae^(kx) (where A=e^C
absorbs the constant — this is exactly the exponential model family from SL 2.5/AHL 2.9).
Common errors: forgetting the constant of integration until after exponentiating (it must
be introduced during the integration step, before rearranging); attempting separation on
an equation where x and y terms cannot actually be fully isolated on opposite sides.
AHL 5.15 — Slope Fields
Theory A slope field visualises a differential equation without solving it: at a grid of points
across the plane, a short line segment is drawn with gradient equal to dy/dx evaluated at
that point (using the differential equation’s right-hand side) — a solution curve, if drawn,
must be tangent to the field’s direction everywhere it passes through, so the overall pattern
of segments lets you sketch the family of possible solution curves by eye, and see
qualitative long-run behaviour (e.g. curves converging to a particular value) even for
equations that are hard or impossible to solve exactly.
Worked example For dy/dx = x−y, describe how you’d sketch the solution curve through
(0,2) using a slope field. At each grid point, compute x−y and draw a short segment of that
gradient; starting at (0,2), follow the direction of the local segments continuously across
the grid to trace out the specific solution curve through that point.
Common errors: drawing a solution curve that crosses the local slope-field segments at an
angle instead of running tangent to them; forgetting the slope field is generated from the
differential equation’s right-hand side, not from any assumed solution formula.

AHL 5.16 — Euler’s Method; Coupled Systems


Theory When a differential equation can’t be solved exactly (or a numerical answer is all
that’s needed), Euler’s method approximates the solution curve step-by-step: starting from
a known point, take a small step of size h in the x-direction, using the current gradient
(from the differential equation) to estimate how far y moves over that step — then repeat
from the new point. Smaller step sizes give more accurate approximations but require
more steps; this is a first-order approximation, so error does accumulate over many steps.
The same idea extends to coupled systems, where two (or more) quantities’ rates of
change each depend on both current values (e.g. predator and prey populations, each
affecting the other’s growth rate) — both variables are stepped forward simultaneously at
each iteration.
Formula
y_(n+1)=y_n+h f(x_n,y_n), x_(n+1)=x_n+h
Worked example dy/dx = x+y, y(0)=1, step h=0.5. Perform one Euler step to estimate
y(0.5). y₁ = y₀ + h·f(x₀,y₀) = 1 + 0.5(0+1) = 1.5 (at x=0.5).
Common errors: using the gradient at the new point instead of the current point for each
step (this is a different, more advanced method); accumulating rounding errors across
many steps without carrying enough decimal precision.
AHL 5.17 — Phase Portraits for Coupled Linear Differential Equations
Theory For a coupled linear system dx/dt=ax+by, dy/dt=cx+dy, the long-run qualitative
behaviour of solution trajectories is entirely determined by the eigenvalues of the
coefficient matrix (a,b;c,d) — connecting this subtopic directly back to AHL 1.15’s
eigenvalue theory. Real eigenvalues of the same sign give trajectories that move directly
toward (both negative — stable) or away from (both positive — unstable) the origin;
opposite signs give a saddle (attracting along one direction, repelling along another);
complex eigenvalues give rotational/spiralling behaviour, with the sign of the real part
determining whether the spiral grows (unstable) or shrinks (stable) toward the origin, and
purely imaginary eigenvalues giving closed, non-decaying elliptical orbits (a centre) — this
last case describes undamped oscillation.
Behaviour by eigenvalue type

Eigenvalues Behaviour
Real, both positive Unstable node — trajectories move away
from origin
Real, both negative Stable node — trajectories move towards
origin
Real, opposite signs Saddle point
Complex, positive real part Unstable spiral
Complex, negative real part Stable spiral
Purely imaginary Centre — closed circular/elliptical orbits

Worked example dx/dt = x−y, dy/dt = x+y. Classify the phase portrait’s behaviour at the
origin. Matrix (1,−1;1,1) has eigenvalues 1±i (complex, positive real part) ⇒ trajectories
spiral outward from the origin (an unstable spiral).
Common errors: forgetting exact analytic solutions are only required for distinct real
eigenvalues (complex-eigenvalue cases are analysed qualitatively via the phase-portrait
table, not solved explicitly by hand); misreading the matrix from a written pair of coupled
equations (row order must match which equation is dx/dt vs dy/dt).

AHL 5.18 — Second-Order Differential Equations


Theory A second-order differential equation involves d²x/dt² and can always be converted
into an equivalent first-order coupled system by introducing a new variable for the first
derivative (y = dx/dt) — this lets Euler’s method (designed for first-order systems) handle
second-order equations too, and also lets the AHL 5.17 phase-portrait classification apply
directly to equations of the form d²x/dt² + a(dx/dt) + bx = 0, which describe damped or
undamped oscillatory systems (springs, circuits, and similar physical contexts).
Worked example d²x/dt² = −4x (simple harmonic motion). Rewrite as a first-order
coupled system and classify its phase portrait. Let y = dx/dt: dx/dt = y, dy/dt = −4x. Matrix
(0,1;−4,0) has eigenvalues ±2i (purely imaginary) ⇒ a centre — closed elliptical
trajectories, i.e. sustained oscillation, consistent with the known exact solution x(t) =
Acos(2t)+Bsin(2t).
Common errors: forgetting to actually introduce the substitution y=dx/dt before
attempting to treat a second-order equation with first-order methods; sign errors carrying
the original equation’s coefficients into the coupled system’s matrix.

Appendix A: IB Command Terms Glossary


Understanding exactly what each command term demands is worth easy marks. The IB
uses these consistently across all papers.
• Write down / State: give the answer with no working required (usually a direct
read-off or trivial calculation).
• Calculate / Find / Determine: obtain an answer showing relevant working — the
method matters for marks, not just the final number.
• Solve: obtain the answer(s) using algebraic and/or technological methods.
• Sketch: draw a graph showing general/key features, without requiring an accurate
scale or graph paper.
• Draw: produce an accurate, to-scale diagram/graph, usually on grid/graph paper,
with labelled axes.
• Show that: prove a given (stated) result using logical steps — since the answer is
given, working must be complete and convincing, not just asserted.
• Hence: use the immediately preceding result to reach the next answer (a different,
usually longer method may lose marks).
• Hence or otherwise: use the preceding result if helpful, but another valid method is
also acceptable.
• Justify / Explain: give a reason or set of reasons supported by evidence or
mathematical argument.
• Comment on: give a judgement based on a given statement or result, e.g. about the
validity or reasonableness of a model.
• Interpret: use knowledge/understanding to recognise trends and draw conclusions
from given information, in context.
• Deduce: reach a conclusion from the information given, usually via a short
logical/algebraic step.
• Verify: confirm that a given statement/result is correct by substitution or direct
check (not a full derivation).
• Suggest: propose a possible answer, solution, or hypothesis where a definitive
answer isn’t required/possible.
• Estimate: obtain an approximate value, often via rounding, a graph, or a model,
where exact calculation isn’t required or possible.
• Distinguish: give the difference(s) between two or more concepts/items.

Appendix B: Formula and GDC Quick-Reference by Topic


A compact lookup table for revision — not a substitute for understanding derivations covered
in the main chapters. Formulas marked (booklet) are given in the exam formula booklet;
those without are expected to be known.
Topic 1 — Number and Algebra
• Arithmetic: uₙ=u₁+(n−1)d; Sₙ=(n/2)(2u₁+(n−1)d) (booklet)
• Geometric: uₙ=u₁rⁿ⁻¹; Sₙ=u₁(rⁿ−1)/(r−1); S∞=u₁/(1−r), |r|<1 (booklet)
• Compound interest: FV=PV(1+r/100k)^(kn) (booklet)
• Percentage error: |vA−vE|/vE × 100%
• Log laws: log(xy)=logx+logy; log(x/y)=logx−logy; log(xᵐ)=m·logx (booklet)
• Complex numbers: z=a+bi; |z|=√(a²+b²); z=r·cis θ=re^(iθ) (booklet)
• 2×2 matrix inverse: A⁻¹ = (1/detA)(d,−b;−c,a)
• Eigenvalues: det(A−λI)=0
GDC skills to have automatic: solving simultaneous equations/polynomials numerically;
finance (TVM) solver; matrix operations (inverse, determinant, multiplication, powers);
complex number arithmetic and conversion between forms.
Topic 2 — Functions
• Line: y=mx+c; y−y₁=m(x−x₁); m₁⊥m₂ ⟺ m₁m₂=−1 (booklet)
• Exponential model: f(x)=ka^x+c or ke^(rx)+c (asymptote y=c)
• Sinusoidal: f(x)=a sin(b(x−c))+d; amplitude a; period 2π/b (rad) or 360/b (deg);
axis y=d (booklet)
• Logistic: f(x)=L/(1+Ce^(−kx)); carrying capacity L (booklet)
• Linearizing: ln y=ln k+x ln a (exponential); ln y=ln a+n ln x (power)
GDC skills: graphing and finding intersections/zeros/extrema; regression for all standard
model families.
Topic 3 — Geometry and Trigonometry
• Sine rule: a/sinA=b/sinB=c/sinC (booklet)
• Cosine rule: c²=a²+b²−2ab cosC (booklet)
• Area of triangle: ½ab sinC (booklet)
• Arc length: rθ (radians) or (θ/360)×2πr (degrees) (booklet)
• Sector area: ½r²θ (radians) or (θ/360)×πr² (degrees) (booklet)
• Vectors: |v|=√(v₁²+v₂²+v₃²); v·w=|v||w|cosθ; |v×w|=|v||w|sinθ (booklet)
• Line: r=a+λb (booklet)
GDC skills: 3D graphing; vector operations; matrix transformations; solving trig equations
graphically.
Topic 4 — Statistics and Probability
• IQR=Q₃−Q₁; outlier beyond Q₁−1.5IQR or Q₃+1.5IQR
• Binomial: X~B(n,p); E(X)=np; Var(X)=np(1−p) (booklet)
• Poisson: X~Po(m); E(X)=Var(X)=m (booklet)
• Normal: X~N(μ,σ²)
• Linear combinations: E(aX+b)=aE(X)+b; Var(aX+b)=a²Var(X) (booklet)
• CLT: X̄ ~N(μ,σ²/n) approx. for large n
• χ² independence df=(r−1)(c−1); goodness-of-fit df=(categories−1)
GDC skills: one-variable and regression statistics; binomial/Poisson/normal pdf & cdf and
their inverses; t-tests, z-tests, χ² tests directly from data or summary statistics; confidence
intervals.
Topic 5 — Calculus
• Power rule: d/dx(xⁿ)=nxⁿ⁻¹; ∫xⁿdx=xⁿ⁺¹/(n+1)+c (n≠−1) (booklet)
• Standard derivatives: sinx→cosx; cosx→−sinx; eˣ→eˣ; lnx→1/x (booklet)
• Chain/product/quotient rules (booklet)
• Trapezoidal rule: (h/2)[y₀+yₙ+2(y₁+…+yₙ₋₁)] (booklet)
• Kinematics: v=ds/dt, a=dv/dt; distance=∫|v|dt
• Euler’s method: yₙ₊₁=yₙ+h·f(xₙ,yₙ) (booklet)
GDC skills: numerical/graphical differentiation and integration; solving f’(x)=0; finding
definite integrals directly.

Appendix C: Practice Question Sets (with Full Worked Solutions)


Mixed SL/AHL questions per topic, in increasing difficulty, mirroring the style of Haese’s
“Review Set” exercises. Attempt each fully before reading the solution.

Topic 1 Practice Set


Q1. Write 0.0000456 in standard form. Solution: Move the decimal 5 places right to get 4.56
before the first non-zero digit’s position: 4.56×10⁻⁵.
Q2. The first term of an arithmetic sequence is 8 and the 10th term is 44. Find the common
difference and the sum of the first 20 terms. Solution: u₁₀=8+9d=44 ⇒ d=4. S₂₀=(20/2)
(2(8)+19(4))=10(16+76)=920.
Q3. €5000 is invested at 3.8% p.a. compounded quarterly for 8 years. Find the final value.
Solution: FV=5000(1+3.8/400)(32)=5000(1.0095)32 ≈ €6785.60.
Q4. Solve log₃(x+2)+log₃(x−2)=2. Solution: log₃((x+2)(x−2))=2 ⇒ x²−4=9 ⇒ x²=13 ⇒ x=√13
(reject −√13 since x−2>0 requires x>2, and √13≈3.6>2). x=√13.
Q5. Find the sum to infinity of 12 − 6 + 3 − 1.5 + … Solution: r=−0.5, S∞=12/(1−
(−0.5))=12/1.5=8.
Q6 (AHL). Solve z²+2z+5=0 for z ∈ ℂ. Solution: z=(−2±√(4−20))/2=(−2±4i)/2=−1±2i.
Q7 (AHL). Write z=−2+2i in polar form. Solution: r=√(4+4)=2√2; θ is in the second
quadrant, arctan(2/−2) adjusted: θ=3π/4. z=2√2 cis(3π/4).
Q8 (AHL). Find the eigenvalues of A=(5,2;2,2). Solution: det(A−λI)=(5−λ)(2−λ)
−4=λ²−7λ+6=0 ⇒ λ=1, 6.

Topic 2 Practice Set


Q1. Find the equation of the line through (−1,3) parallel to y=−2x+5. Solution: m=−2
(parallel). y−3=−2(x+1) ⇒ y=−2x+1.

f⁻¹(x)=2x−4. f⁻¹(6)=8; f(8)=(8+4)/2=6 ✓.


Q2. f(x)=(x+4)/2. Find f⁻¹(x) and verify f(f⁻¹(6))=6. Solution: Swap: x=(y+4)/2 ⇒ y=2x−4.

Q3. A population is modelled by P(t)=1200e^(0.03t). Find the initial population and the
population after 10 years. Solution: P(0)=1200. P(10)=1200e^0.3≈1620 (3 s.f.).
Q4. State the four stages of the modelling cycle. Solution: Develop a model; find
parameters; test and reflect; use the model (with awareness of extrapolation risk).
Q5 (AHL). f(x)=x²+1, g(x)=2x−3. Find (g∘f)(x) and (f∘g)(x). Solution: (g∘f)
(x)=2(x²+1)−3=2x²−1. (f∘g)(x)=(2x−3)²+1=4x²−12x+10.
Q6 (AHL). Describe the transformation taking y=x² to y=−2(x−3)²+1. Solution: Vertical
stretch factor 2, reflection in the x-axis, then translation 3 right and 1 up (order:
stretch/reflect before translating, matching the algebraic structure).
Q7 (AHL). A logistic model has L=800, and P(0)=50. If C=15, find k given P(5)=200.
Solution: 200=800/(1+15e^(−5k)) ⇒ 1+15e^(−5k)=4 ⇒ e^(−5k)=0.2 ⇒
k=−ln(0.2)/5≈0.322.
Q8 (AHL). A power-law dataset gives a log-log regression line ln y = 0.693+1.5 ln x. Find
the model y=ax^n. Solution: n=1.5; a=e^0.693≈2.0. y≈2x^1.5.

Topic 3 Practice Set


Q1. A triangle has sides a=9, b=12, c=15. Find its largest angle. Solution: Largest angle is
opposite the longest side (c=15): cosC=(9²+12²−15²)/(2×9×12)=0/216=0 ⇒ C=90°.
Q2. Find the arc length and sector area of a sector with radius 10 cm and angle 120°.
Solution: Arc=(120/360)(2π×10)≈20.9 cm. Area=(120/360)(π×100)≈104.7 cm².
Q3. Find the perpendicular bisector of A(2,−1) and B(6,3). Solution: Midpoint (4,1);
gradient AB=1, m_perp=−1. y−1=−(x−4) i.e. y=−x+5.
Q4. From a point 25 m from a building’s base, the angle of elevation to the top is 38°. Find
the building’s height. Solution: h=25tan38°≈19.5 m.
Q5 (AHL). Convert 210° to radians, and find the arc length of a sector with this angle and
radius 6 cm. Solution: 210×(π/180)=7π/6 rad. Arc length=6×7π/6=7π ≈ 22.0 cm.
Q6 (AHL). Find the angle between a=(1,2,2) and b=(2,−1,2). Solution: a·b=2−2+4=4; |a|=3, |
b|=3. cosθ=4/9 ⇒ θ≈63.6°.
Q7 (AHL). A graph has vertices with degrees 3,3,2,2,2. Does an Eulerian circuit or trail
exist? Solution: Two vertices (degree 3) are odd ⇒ an Eulerian trail exists (not a circuit).
Q8 (AHL). Use Kruskal’s algorithm to find the MST for edges AB(4), AC(2), BC(3), BD(5),

DE ✓, AC ✓, BC ✓, AB — creates cycle A-B-C-A, skip; BD ✓ connects D (already linked via


CD(6), CE(7), DE(1). Solution: Sorted: DE(1), AC(2), BC(3), AB(4), BD(5), CD(6), CE(7). Add

DE) to B ✓ (no cycle: check — B,C,A already connected; D,E already connected; BD joins the
two groups ✓). All 5 vertices connected with 4 edges: {DE, AC, BC, BD}, total weight =
1+2+3+5 = 11.

Topic 4 Practice Set


Q1. A school of 1200 students is stratified by year group. If Year 13 has 180 students, find
the Year 13 sample size for a total sample of 60. Solution: (180/1200)×60=9.
Q2. A dataset has Q1=24, Q3=38. Determine if a value of 60 is an outlier. Solution: IQR=14;
upper fence=38+1.5(14)=59. Since 60>59, 60 is an outlier.
Q3. X~B(15,0.4). Find E(X) and Var(X). Solution: E(X)=15(0.4)=6. Var(X)=15(0.4)(0.6)=3.6.
Q4. X~N(60,8²). Find P(50<X<70). Solution: This range is exactly μ±1.25σ (not exactly the
empirical-rule bracket) — GDC normal CDF: ≈0.7887.
Q5. Two events have P(A)=0.3, P(B)=0.6, and A,B are independent. Find P(A∪B). Solution:
P(A∩B)=0.3×0.6=0.18. P(A∪B)=0.3+0.6−0.18=0.72.
Q6 (AHL). X~Po(5). Find P(X≥3) using technology. Solution:
P(X≥3)=1−P(X≤2)≈1−0.1247=0.8753.
Q7 (AHL). A sample of n=40 has x̄=55, s=9. Construct a 90% confidence interval for μ.
Solution: σ unknown → t-distribution, df=39. GDC t-interval: ≈(52.6, 57.4).
Q8 (AHL). In a 3×2 contingency table, a χ² test gives χ²=6.2, df=2. Test H₀: independence at
the 5% level (critical value χ²₀.₀₅,₂=5.99). Solution: Since 6.2>5.99, reject H₀ — evidence of
association between the variables.

Topic 5 Practice Set


Q1. Differentiate f(x)=5x³−2x²+7. Solution: f’(x)=15x²−4x.
Q2. Find the equation of the tangent to y=x²−4x+1 at x=3. Solution: y(3)=9−12+1=−2;
y’=2x−4 ⇒ y’(3)=2. Tangent: y+2=2(x−3) ⇒ y=2x−8.
Q3. Find ∫(6x²−4x+1)dx. Solution: 2x³−2x²+x+c.
Q4. Find and classify the stationary points of f(x)=x³−12x. Solution: f’(x)=3x²−12=0 ⇒ x=±2.
f’‘(x)=6x. f’‘(2)=12>0 (local min); f’’(−2)=−12<0 (local max). Local max at x=−2 (f=16);
local min at x=2 (f=−16).
Q5. Use the trapezoidal rule with 4 strips to estimate ∫₀⁸ √x dx. Solution: h=2; x=0,2,4,6,8 →
y=0, 1.414, 2, 2.449, 2.828. Estimate=(2/2)
[0+2.828+2(1.414+2+2.449)]=1×[2.828+11.726]=14.55.
Q6 (AHL). Differentiate y=x²e^(3x) using the product rule. Solution: u=x², v=e^(3x); u’=2x,
v’=3e^(3x). y’=2xe(3x)+3x²e(3x)=xe^(3x)(2+3x).
Q7 (AHL). Find ∫6x²(x³+1)⁴ dx by inspection/substitution. Solution: Let u=x³+1, du=3x²dx,
so 6x²dx=2du. ∫2u⁴du = (2/5)u⁵+c = (2/5)(x³+1)⁵+c.
Q8 (AHL). A particle has velocity v(t)=3t²−12t (m/s) for 0≤t≤5. Find the total distance
travelled. Solution: v=0 at t=0,4. v<0 on (0,4), v>0 on (4,5). ∫₀⁴v dt=[t³−6t²]₀⁴=64−96=−32
(distance 32). ∫₄⁵v dt=[t³−6t²]₄⁵−[…]₄=(125−150)−(64−96)=−25−(−32)=7. Total
distance=32+7=39 m.

Appendix D: Practice Question Sets B (Further Practice)


Topic 1 — Set B
Q1. Write 3.2×10⁻⁴ as an ordinary decimal. Solution: 0.00032.
Q2. An arithmetic sequence has u₃=11 and u₈=31. Find u₁ and d, then find S₁₀. Solution:
5d=20 ⇒ d=4; u₁=11−2(4)=3. S₁₀=(10/2)(2(3)+9(4))=5(6+36)=210.
Q3. A car depreciates from $25000 by 15% per year. Find its value after 6 years. Solution:
25000(0.85)⁶ ≈ $9430.94.
Q4. Solve 5^(2x−1)=40 for x. Solution: (2x−1)ln5=ln40 ⇒ 2x−1=ln40/ln5≈2.292 ⇒
x≈1.646.
Q5. A length is measured as 12.5 cm, correct to 1 d.p. Find the upper and lower bounds.
Solution: 12.45 ≤ length < 12.55.
Q6 (AHL). Simplify log₂32 − log₂4 + log₂1. Solution: 5−2+0=3.
Q7 (AHL). Find z₁z₂ and z₁/z₂ given z₁=2cis(π/4), z₂=5cis(π/12). Solution:
z₁z₂=10cis(π/4+π/12)=10cis(π/3). z₁/z₂=(2/5)cis(π/4−π/12)=0.4cis(π/6).
Q8 (AHL). For A=(3,1;0,2), find A⁻¹. Solution: det=6, A⁻¹=(1/6)(2,−1;0,3)=(1/3,−1/6;
0,1/2).

Topic 2 — Set B
Q1. Find the equation of the line perpendicular to 4x+2y=6, passing through (3,−1).
Solution: Rearranged: y=−2x+3, m=−2, m_perp=1/2. y+1=(1/2)(x−3) ⇒ y=x/2−2.5.
Q2. f(x)=(2x−1)/(x+3), x≠−3. Find f⁻¹(x). Solution: Swap: x=(2y−1)/(y+3) ⇒ x(y+3)=2y−1
⇒ xy+3x=2y−1 ⇒ y(x−2)=−1−3x ⇒ f⁻¹(x)=(−1−3x)/(x−2).
Q3. State whether f(x)=200(0.85)^x models growth or decay, and find the percentage rate.
Solution: Base<1 ⇒ decay; rate = 15% per unit x.
Q4. Explain why extrapolating a linear model far beyond the data range is risky, using a
specific example. Solution: E.g. a linear model of a plant’s height over its first month would
predict impossible, unbounded growth if extended for years — real growth typically levels
off, which a model fitted only to early data cannot capture.
Q5 (AHL). f(x)=√x (x≥0), g(x)=x−4. Find (f∘g)(x) and its domain. Solution: (f∘g)(x)=√(x−4);
domain x≥4 (need x−4≥0 for the square root to be defined).
Q6 (AHL). Describe the single transformation taking y=cos x to y=cos(x)−3. Solution:
Vertical translation 3 units down.
Q7 (AHL). A logistic model has carrying capacity 1000 and P(0)=100, C=9. Find P(3) given
k=0.5. Solution: P(3)=1000/(1+9e^(−1.5))=1000/(1+9(0.2231))=1000/3.008≈332.
Q8 (AHL). Data follows y=ka^x. A semi-log plot (ln y vs x) has gradient −0.223 and
intercept 3.00. Find k and a. Solution: k=e³≈20.1; a=e^(−0.223)≈0.800.

Topic 3 — Set B
Q1. A triangle has a=6, b=8, C=60°. Find c and the area. Solution: c²=36+64−2(6)
(8)cos60°=100−48=52 ⇒ c≈7.21. Area=(1/2)(6)(8)sin60°≈20.8.
Q2. A cone has radius 4 cm and height 9 cm. Find its slant height and curved surface area
(πrl). Solution: l=√(16+81)=√97≈9.85. CSA=π(4)(9.85)≈123.8 cm².
Q3. Find the perpendicular bisector of A(−2,3) and B(4,−1). Solution: Midpoint (1,1);
gradient AB=−4/6=−2/3, m_perp=3/2. y−1=(3/2)(x−1).
Q4. A ship sails on a bearing of 245° for 60 km. Find how far south and west it has
travelled. Solution: Bearing 245° is 245−180=65° past south, measured toward west. South
component=60cos65°≈25.4 km; west component=60sin65°≈54.4 km.
Q5 (AHL). Convert 5π/4 rad to degrees, and find sin and cos of this angle. Solution: 225°.
sin(225°)=−√2/2≈−0.707; cos(225°)=−√2/2≈−0.707.
Q6 (AHL). For a=(3,0,4) and b=(0,5,0), find a×b and its magnitude. Solution:
a×b=(0×0−4×5, 4×0−3×0, 3×5−0×0)=(−20,0,15). |a×b|=√(400+225)=√625=25.
Q7 (AHL). A graph’s adjacency matrix is A=(0,1,0;1,0,1;0,1,0). Find A² and interpret A²₁₃.
Solution: A²=(1,0,1; 0,2,0; 1,0,1). A²₁₃=1: there is exactly 1 walk of length 2 from vertex 1
to vertex 3 (via vertex 2).
Q8 (AHL). Apply the nearest-neighbour algorithm starting from A on the weighted graph
with edges AB(3),AC(5),AD(9),BC(4),BD(6),CD(2), to estimate an upper bound for the TSP
tour. Solution: From A, nearest is B(3). From B, nearest unvisited is C(4). From C, nearest
unvisited is D(2). Return D→A(9). Total = 3+4+2+9=18 (upper bound).

Topic 4 — Set B
Q1. A population of 2000 is divided into 3 strata of sizes 800, 700, 500. Find the sample
size from each stratum for a total sample of 100. Solution: 800/2000×100=40;
700/2000×100=35; 500/2000×100=25.
Q2. A dataset has mean 25 and standard deviation 4. Every value is increased by 10%. Find
the new mean and standard deviation. Solution: Multiplying by 1.1: new
mean=25×1.1=27.5; new σ=4×1.1=4.4.
Q3. X~B(20,0.25). Find P(X=5) and P(X≤5) using technology. Solution: P(X=5)≈0.2023.
P(X≤5)≈0.6172.
Q4. A dataset gives r=0.15. Comment on what this suggests, bearing in mind the scatter
diagram is not shown. Solution: A very weak linear correlation — but this alone does not
rule out a strong non-linear relationship; the scatter diagram should always be checked
before concluding “no relationship.”
Q5 (AHL). X~Po(3.5). Find P(X=0) and interpret this value in context (e.g. calls per hour).
Solution: P(X=0)=e^(−3.5)≈0.0302 — about a 3% chance of zero events (e.g. zero calls)
occurring in the given interval.
Q6 (AHL). A sample of 50 gives a 95% CI for μ of (48.2, 55.8). State the point estimate for μ
and the margin of error. Solution: Point estimate (midpoint) = (48.2+55.8)/2=52. Margin of
error = (55.8−48.2)/2=3.8.
Q7 (AHL). Explain the difference between a Type I and Type II error in the context of a
medical test where H₀: “patient is healthy.” Solution: Type I (false positive): concluding a
healthy patient is sick. Type II (false negative): failing to detect that a sick patient is
actually sick.
Q8 (AHL). A weather Markov chain has T=(0.8,0.3; 0.2,0.7). Find the steady-state
distribution. Solution: 0.8s₁+0.3s₂=s₁ ⇒ 0.2s₁=0.3s₂ ⇒ s₁=1.5s₂; with s₁+s₂=1: s₂=0.4,
s₁=0.6. Steady state (0.6, 0.4).

Topic 5 — Set B
Q1. Differentiate f(x) = 4x^(1/2) − 3x⁻¹. Solution: f’(x) = 2x^(−1/2) + 3x⁻².
Q2. Find the normal to y=x²+2x at x=1. Solution: y(1)=3; y’=2x+2 ⇒ y’(1)=4. Normal
gradient=−1/4. y−3=−(1/4)(x−1).
Q3. Find ∫(4/x − 3sinx)dx. Solution: 4ln|x| + 3cosx + c.
Q4. A rectangle has perimeter 40 cm. Find the dimensions that maximise its area. Solution:
Let width=x, length=20−x. Area A(x)=x(20−x)=20x−x². A’(x)=20−2x=0 ⇒ x=10. Since the
second derivative is negative, this is a maximum: 10 cm × 10 cm (a square), area 100 cm².
Q5. Use the trapezoidal rule with 3 strips to estimate ∫₁⁴ (1/x) dx. Solution: h=1; x=1,2,3,4
→ y=1, 0.5, 0.333, 0.25. Estimate=(1/2)[1+0.25+2(0.5+0.333)]=(1/2)[1.25+1.667]=1.458
(exact value ln4≈1.386 — the trapezoidal rule overestimates here since 1/x is convex).
Q6 (AHL). Differentiate y = ln(x²+1) / x using the quotient rule. Solution: u=ln(x²+1), v=x;
u’=2x/(x²+1), v’=1. y’ = [2x²/(x²+1) − ln(x²+1)] / x².
Q7 (AHL). Solve the differential equation dy/dx = y/x (x>0) given y(1)=5. Solution:
Separate: ∫(1/y)dy=∫(1/x)dx ⇒ ln|y|=ln|x|+C ⇒ y=Ax. Using y(1)=5: A=5. y=5x.
Q8 (AHL). For dx/dt=−2x+y, dy/dt=x−2y, classify the phase portrait. Solution: Matrix
(−2,1;1,−2); eigenvalues from (−2−λ)²−1=0 ⇒ λ=−1,−3 (real, both negative). Stable node
— trajectories move toward the origin.

You might also like