0% found this document useful (0 votes)
4 views1 page

page_3

The document outlines the structure and content of a text on neural networks, covering topics such as the forward and backward passes, activation functions, and training mechanics. It includes sections on linear layers, nonlinearity, initialization, normalization, and autograd. The document serves as a comprehensive guide for understanding neural network operations and training processes.

Uploaded by

Cường Mậm
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views1 page

page_3

The document outlines the structure and content of a text on neural networks, covering topics such as the forward and backward passes, activation functions, and training mechanics. It includes sections on linear layers, nonlinearity, initialization, normalization, and autograd. The document serves as a comprehensive guide for understanding neural network operations and training processes.

Uploaded by

Cường Mậm
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

The Holy Grail CONTENTS

2. einops . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13
3. jaxtyping . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14
1.10 The mental model . . . . . . . . . . . . . . . . . . . . . . . . . 14
Read it yourself . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15
Practice . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15

The forward pass: a neural network as a pure function 17


2.1 The linear layer is y = Wx + b . . . . . . . . . . . . . . . . . . . . 18
Batched linear . . . . . . . . . . . . . . . . . . . . . . . . . . 18
Counting parameters . . . . . . . . . . . . . . . . . . . . . . . 19
2.2 What “pure function” means . . . . . . . . . . . . . . . . . . . . 19
2.3 Why nonlinearity matters . . . . . . . . . . . . . . . . . . . . . 20
2.4 The activation zoo . . . . . . . . . . . . . . . . . . . . . . . . . 21
ReLU — max(0, x) . . . . . . . . . . . . . . . . . . . . . . . . 21
GELU — Gaussian Error Linear Unit . . . . . . . . . . . . . . 21
SiLU / Swish — x * sigmoid(x) . . . . . . . . . . . . . . . . . . 22
SwiGLU — the gated variant . . . . . . . . . . . . . . . . . . . 22
Sigmoid and Tanh . . . . . . . . . . . . . . . . . . . . . . . . 22
Softmax — the activation for the output . . . . . . . . . . . . . 23
2.5 Stacking layers as function composition: the MLP . . . . . . . . 23
2.6 Hand-tracing a tiny MLP . . . . . . . . . . . . . . . . . . . . . 24
2.7 Initialization: why bad init silently kills training . . . . . . . . . 25
2.8 Normalization: a preview . . . . . . . . . . . . . . . . . . . . . 26
2.9 Residual connections: a preview . . . . . . . . . . . . . . . . . 26
2.10 The forward pass mental model . . . . . . . . . . . . . . . . . 27
Read it yourself . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28
Practice . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28

The backward pass: autograd, gradients, the chain rule made


mechanical 29
3.1 Training is gradient computation . . . . . . . . . . . . . . . . . 30
3.2 The chain rule, mechanically . . . . . . . . . . . . . . . . . . . 30
Worked example . . . . . . . . . . . . . . . . . . . . . . . . . 31
3.3 Computation graphs as DAGs . . . . . . . . . . . . . . . . . . . 32
3.4 Reverse-mode autodiff . . . . . . . . . . . . . . . . . . . . . . . 32
Why reverse mode is sometimes called “backpropagation” . . . 33
3.5 What autograd actually stores . . . . . . . . . . . . . . . . . . . 33
3.6 The vector-Jacobian product, and why “gradient” is slightly the
wrong word . . . . . . . . . . . . . . . . . . . . . . . . . . . . 34
3.7 Common gradients to know cold . . . . . . . . . . . . . . . . . . 35
3.8 The memory cost of backward — and why it scales with batch size 36
3.9 The four toggles you must understand . . . . . . . . . . . . . . 37
3.10 Why mixed precision needs loss scaling . . . . . . . . . . . . . 38

iii

You might also like