The Holy Grail CONTENTS
2. einops . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13
3. jaxtyping . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14
1.10 The mental model . . . . . . . . . . . . . . . . . . . . . . . . . 14
Read it yourself . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15
Practice . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15
The forward pass: a neural network as a pure function 17
2.1 The linear layer is y = Wx + b . . . . . . . . . . . . . . . . . . . . 18
Batched linear . . . . . . . . . . . . . . . . . . . . . . . . . . 18
Counting parameters . . . . . . . . . . . . . . . . . . . . . . . 19
2.2 What “pure function” means . . . . . . . . . . . . . . . . . . . . 19
2.3 Why nonlinearity matters . . . . . . . . . . . . . . . . . . . . . 20
2.4 The activation zoo . . . . . . . . . . . . . . . . . . . . . . . . . 21
ReLU — max(0, x) . . . . . . . . . . . . . . . . . . . . . . . . 21
GELU — Gaussian Error Linear Unit . . . . . . . . . . . . . . 21
SiLU / Swish — x * sigmoid(x) . . . . . . . . . . . . . . . . . . 22
SwiGLU — the gated variant . . . . . . . . . . . . . . . . . . . 22
Sigmoid and Tanh . . . . . . . . . . . . . . . . . . . . . . . . 22
Softmax — the activation for the output . . . . . . . . . . . . . 23
2.5 Stacking layers as function composition: the MLP . . . . . . . . 23
2.6 Hand-tracing a tiny MLP . . . . . . . . . . . . . . . . . . . . . 24
2.7 Initialization: why bad init silently kills training . . . . . . . . . 25
2.8 Normalization: a preview . . . . . . . . . . . . . . . . . . . . . 26
2.9 Residual connections: a preview . . . . . . . . . . . . . . . . . 26
2.10 The forward pass mental model . . . . . . . . . . . . . . . . . 27
Read it yourself . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28
Practice . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28
The backward pass: autograd, gradients, the chain rule made
mechanical 29
3.1 Training is gradient computation . . . . . . . . . . . . . . . . . 30
3.2 The chain rule, mechanically . . . . . . . . . . . . . . . . . . . 30
Worked example . . . . . . . . . . . . . . . . . . . . . . . . . 31
3.3 Computation graphs as DAGs . . . . . . . . . . . . . . . . . . . 32
3.4 Reverse-mode autodiff . . . . . . . . . . . . . . . . . . . . . . . 32
Why reverse mode is sometimes called “backpropagation” . . . 33
3.5 What autograd actually stores . . . . . . . . . . . . . . . . . . . 33
3.6 The vector-Jacobian product, and why “gradient” is slightly the
wrong word . . . . . . . . . . . . . . . . . . . . . . . . . . . . 34
3.7 Common gradients to know cold . . . . . . . . . . . . . . . . . . 35
3.8 The memory cost of backward — and why it scales with batch size 36
3.9 The four toggles you must understand . . . . . . . . . . . . . . 37
3.10 Why mixed precision needs loss scaling . . . . . . . . . . . . . 38
iii