DeepLearning Notes Sample
DeepLearning Notes Sample
al
Made by Anmol Kansal
Maulana Abul Kalam Azad University of Technology, West Bengal
s
May 26, 2026
an
Your Complete Concept Handbook
K
This is Book 1 of 2. It teaches you every concept in your syllabus
from rst principles, with diagrams, intuition, and memory hooks.
1 day Quick Revision Handbook (end of book) + Book 2 PYQs only. Skim
formula sheet and comparison tables.
al
3 days Read Units 2, 3, 5 thoroughly (these carry 40 of 70 marks). Then
Quick Revision + Book 2 PYQs.
7 days All 6 units from this book. Then Book 2 PYQs and expected ques-
s
tions.
Full prep Both books cover-to-cover, including Book 2's MCQs and numericals.
1 Introduction 3 5 Low
Subtotal: 36 70
A
Color Meaning
s al
an
K
ol
nm
A
Contents
al
1.2.1 Supervised Learning Two Flavours . . . . . . . . . . . . . . . . . . . . . . 10
1.2.2 Unsupervised Learning Three Flavours . . . . . . . . . . . . . . . . . . . . 10
1.2.3 Reinforcement Learning Key Terms . . . . . . . . . . . . . . . . . . . . . . 10
1.3 Why Deep Learning Works Now (and Didn't in 1990) . . . . . . . . . . . . . . . . 10
s
1.4 Issues / Challenges in Deep Learning . . . . . . . . . . . . . . . . . . . . . . . . . . 11
1.5 Quick Review of Fundamental Learning Techniques . . . . . . . . . . . . . . . . . . 11
an
1.6 Evaluation Metrics for Classication . . . . . . . . . . . . . . . . . . . . . . . . . . 11
1.7 Bias-Variance Tradeo . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12
1.8 Deep Dive: Worked Examples . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12
1.8.1 Worked Example 1 Confusion Matrix Metrics . . . . . . . . . . . . . . . . 12
1.8.2 Worked Example 2 Why F1 Beats Accuracy (Imbalanced Data) . . . . . . 13
K
1.8.3 Worked Example 3 Confusion Matrix with Zero O-Diagonals . . . . . . . 13
1.8.4 Worked Example 4 Entropy of a Coin . . . . . . . . . . . . . . . . . . . . 13
al
3.3.3 Tiny Worked Example (one weight) . . . . . . . . . . . . . . . . . . . . . . 27
3.4 Vanishing and Exploding Gradients . . . . . . . . . . . . . . . . . . . . . . . . . . . 27
3.4.1 How to Fix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27
s
3.5 Regularization Techniques . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28
3.5.1 L1 vs L2 Regularization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28
an
3.5.2 Dropout . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28
3.5.3 Other Regularization Techniques . . . . . . . . . . . . . . . . . . . . . . . . 28
3.6 Optimizers Beyond SGD . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28
3.6.1 Momentum . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29
3.6.2 RMSProp . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29
K
3.6.3 Adam (the default in 2024+) . . . . . . . . . . . . . . . . . . . . . . . . . . 29
3.6.4 Optimizer Comparison . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29
3.7 Model Selection . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29
3.7.1 Train / Validation / Test Split . . . . . . . . . . . . . . . . . . . . . . . . . 29
3.7.2 k-Fold Cross-Validation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30
ol
al
4.9.2 Worked Example 2 HMM Forward Algorithm (Numerical) . . . . . . . . . 40
4.9.3 Worked Example 3 Viterbi vs Forward (Same Recursion, Dierent Operator) 40
4.9.4 Worked Example 4 Entropy Calculations . . . . . . . . . . . . . . . . . . . 40
4.9.5 Worked Example 5 HMM Three Problems Mapped to Real Tasks . . . . . 41
s
4.10 Advanced Topics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 41
4.10.1 D-Separation in Bayes Networks . . . . . . . . . . . . . . . . . . . . . . . . 41
an
4.10.2 Conditional Independence in HMM (Markov Property) . . . . . . . . . . . . 41
4.10.3 The Forward-Backward Algorithm . . . . . . . . . . . . . . . . . . . . . . . 41
4.10.4 Baum-Welch (EM for HMM) . . . . . . . . . . . . . . . . . . . . . . . . . . 42
4.10.5 Linear-Chain CRF in Detail . . . . . . . . . . . . . . . . . . . . . . . . . . . 42
4.10.6 HMM vs CRF: When to Use Each . . . . . . . . . . . . . . . . . . . . . . . 43
K
4.10.7 Belief Propagation: The General Framework . . . . . . . . . . . . . . . . . . 43
4.10.8 Cross-Entropy = Entropy + KL Divergence . . . . . . . . . . . . . . . . . . 43
4.10.9 Mutual Information . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 44
al
5.9.7 Why LSTM Solves Vanishing Gradient . . . . . . . . . . . . . . . . . . . . . 55
5.9.8 GRU vs LSTM: Detailed Comparison . . . . . . . . . . . . . . . . . . . . . 55
5.9.9 Attention Mechanism: The Math . . . . . . . . . . . . . . . . . . . . . . . . 56
5.9.10 Transformer Encoder Block in Detail . . . . . . . . . . . . . . . . . . . . . . 56
s
5.9.11 Autoencoders: The Family . . . . . . . . . . . . . . . . . . . . . . . . . . . 57
5.9.12 Restricted Boltzmann Machine Energy and Training . . . . . . . . . . . . 57
an
5.9.13 Deep Belief Network: Layer-Wise Pre-training . . . . . . . . . . . . . . . . . 57
6.4.4 Attention . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 60
6.4.5 Transformer . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 61
6.4.6 BERT and GPT . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 61
6.5 Deep Dive: PCA (Was a 15-Mark Question in 2023!) . . . . . . . . . . . . . . . . . 61
6.5.1 The Goal . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 61
6.5.2 The Optimisation Problem . . . . . . . . . . . . . . . . . . . . . . . . . . . 61
6.5.3 Solve via Lagrangian . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 61
A
al
7.2.6 LSTM vs GRU . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 70
7.3 Critical Memory Hooks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 70
7.4 Architecture Quick Sketches . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 71
s
7.5 Top 10 Things You Must Know . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 72
7.6 The 24-Hour Final Strategy . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 72
an
K
ol
nm
A
al
Concept
Deep Learning is a sub-eld of Machine Learning that uses multi-layered neural networks
to automatically learn hierarchical representations from raw data without needing humans
s
to design features manually.
an
Plain English: Instead of telling the computer a face has two eyes, a nose, a mouth, look for
these, we just show it thousands of face photos. The network gures out by itself across many
layers that edges combine into shapes, shapes combine into features (eyes, noses), and features
combine into faces.
K
2-Mark Ready Answer
1.1.1 AI vs ML vs DL Hierarchy
nm
Articial Intelligence
Machine Learning
Deep
Learning
A
AI (1950s): Any system that mimics human intelligence. Includes rule-based expert systems.
ML (1980s): Systems that learn from data. Includes decision trees, SVMs, neural networks.
DL (2010s): ML with deep neural networks. Powered by big data + GPU + new algorithms.
Supervised Mapping x→y from labelled Email spam lter, digit recog-
data nition
al
Mnemonic SUR : Supervised = with teacher, Unsupervised = self-organize, Reinforcement
= trial-and-error.
s
1.2.1 Supervised Learning Two Flavours
an
Classication: output is a category. Binary (spam/not), multi-class (digit 0-9), multi-label.
Dimensionality reduction: compress data while keeping structure. PCA, t-SNE, autoen-
coders.
Intuition
RL feels like training a dog with treats. The dog (agent) takes actions (sit, jump). Good
actions get treats (reward); bad actions get nothing (or scolding = penalty). Over time the
dog learns the policy: which action in which situation maximises treats.
1. Big Data: ImageNet (14M images, 2009), Wikipedia, YouTube, Common Crawl. Networks
need millions of examples.
2. GPU Compute: NVIDIA's CUDA (2007) made parallel matrix multiplication 100× faster
than CPU.
3. Algorithms: ReLU (2010), Dropout (2012), Batch Norm (2015), Residual connections (2015),
Adam (2015), Attention/Transformer (2017).
al
Data-hungry • Overtting • Vanishing gradients • Expensive (compute) • Black-box •
Adversarial vulnerability • Catastrophic forgetting • Hyperparameter sensitivity
s
an
Method Type Key Idea
variance
Confusion Matrix:
A
Predicted + Predicted
Actual + TP FN
Actual FP TN
Key Formula
TP + TN
Accuracy =
TP + TN + FP + FN
TP
Precision = (of those predicted +, how many are really +?)
TP + FP
TP
Recall = (of all real +, how many did we catch?)
TP + FN
2·P ·R
F1 = (harmonic mean)
P +R
al
F1 vs Accuracy a recurring question.
Accuracy is misleading on imbalanced datasets. If 99% of emails are not spam, predicting
s
not spam always gives 99% accuracy but 0% recall on spam. F1 balances precision and recall,
catching this failure.
/! Examiner Trap
an
Trap: If all o-diagonal entries of the confusion matrix are zero, the classier is perfect on
this dataset (zero mistakes). But verify on a held-out test set it could be overtting!
K
1.7 Bias-Variance Tradeo
bias2
model complexity
Bias = how wrong on average. Variance = how much answers wobble across runs.
Need both low.
TP + TN 80 + 90
Acc = = = 0.85
Total 200
TP 80
Prec = = = 0.889
TP + FP 90
TP 80
Rec = = = 0.80
TP + FN 100
2·P ·R 2(0.889)(0.80)
F1 = = = 0.842
P +R 0.889 + 0.80
al
*** PYQ Favourite
s
Concrete example: 99% of emails are not spam. A trivial always not spam classier:
an
TP = 0 (never predicts spam)
TN = 99
FP = 0
ol
F1 = 0.
Lesson: Accuracy is misleading on imbalanced data. F1 punishes models that ignore the
nm
Problem (PYQ 2024 Q7g, 1M): If all o-diagonal entries of the confusion matrix are zero,
A
Answer: O-diagonal entries are the misclassications. If all are zero, there are no mistakes
⇒ the classier is perfect on this dataset. Accuracy = Precision = Recall = F1 = 1.0.
Caveat: If this is the training set, the model might be overtting. Verify on a held-out test
set before celebrating.
Problem: Compute the entropy (in bits) of (a) a fair coin, (b) a heavily biased coin with
p(H) = 0.9.
P
Formula: H=− x p(x) log2 p(x).
H = −[0.9 log2 0.9 + 0.1 log2 0.1] ≈ −[0.9(−0.152) + 0.1(−3.322)] = 0.469 bit .
Intuition
Less uncertainty ⇒ less entropy. A biased coin is more predictable, so its entropy is lower.
Maximum entropy is achieved by the uniform distribution.
al
3 paradigms: supervised, unsupervised, reinforcement.
DOVE-BACH.
s
Issues:
an
K
ol
nm
A
How examiners ask: Dene ANN. Linear neuron numerical. Sigmoid range and derivative.
Why ReLU sparse? Parameter count.
al
A real neuron has dendrites (receive signals), a soma (cell body that decides whether to re),
and an axon (sends signal out). Neurons connect at synapses with varying strengths.
Articial neuron mimics this minimally: weighted inputs → sum → activation → output.
s
x1
an
w1
x2 w2
Σ ϕ y
x3 w3
b
K
1 (bias)
weighted sum activation
Key Formula
n
!
X
y=ϕ wi x i + b
i=1
Order WIB-A : Weight × Input, sum, add Bias, apply Activation. Always.
2.2.1 Sigmoid
Key Formula
1
σ(z) = σ ′ (z) = σ(z)(1 − σ(z))
1 + e−z
Range: (0, 1) Max gradient: 0.25 at z = 0.
2.2.2 Tanh
Key Formula
ez − e−z
tanh(z) = tanh′ (z) = 1 − tanh2 (z)
ez + e−z
Range: (−1, 1) Max gradient: 1 at z = 0.
2.2.3 ReLU
Key Formula
al
(
′ 1 z>0
ReLU(z) = max(0, z) ReLU (z) =
0 z≤0
s
Use: hidden layers everywhere. Default choice for CNNs/MLPs.
an
Why ReLU is great:
Dead ReLU problem: if a neuron's input becomes < 0 for all training examples, its gradient
is permanently 0 and it never recovers. Fix: use Leaky ReLU (αz for z < 0) or ELU.
ol
Key Formula
nm
ezi
softmax(z)i = PK
zj
j=1 e
al
An MLP is a stack of fully-connected layers, each consisting of an ane transformation followed
by a non-linear activation: a[ℓ] = ϕ(W [ℓ] a[ℓ−1] + b[ℓ] ).
s
an
K
Input (3) Hidden (4) Output (2)
Key Formula
A
[ℓ]
params = n[ℓ] · (n[ℓ−1] + 1)
Layer 1: 8(3 + 1) = 32
Layer 2: 8(8 + 1) = 72
Layer 3: 3(8 + 1) = 27
Total: 131 parameters (112 weights + 19 biases)
A feed-forward network with one hidden layer and a non-polynomial activation can ap-
proximate any continuous function on a compact set, to arbitrary accuracy, given enough
neurons.
So why go deep? The theorem says one layer can work, but might need exponentially
many neurons. Deep networks achieve the same with exponentially fewer parameters, thanks to
hierarchical feature composition.
sal
Key Formula
an
where (the input).
/! Examiner Trap
A single perceptron can only solve linearly separable problems (Minsky-Papert critique,
1969). XOR requires at least one hidden layer. This is the historical motivation for MLPs.
Key insight: both produce a linear decision boundary wT x + b = 0, so they classify the same
way on separable data. The dierences are in training and output type.
∆wi = η · (t − y) · xi
1 ∂J ∂J
− y)2 , y =
P
Derivation: J = 2 (t wi xi . ∂wi = −(t − y)xi . Update: wi ← wi − η ∂w =
i
wi + η(t − y)xi .
al
The syllabus mentions cardinality, operations, and properties of fuzzy relations. This is rarely
the main focus but worth knowing:
s
Fuzzy set: elements have membership µ(x) ∈ [0, 1].
an
Cardinality: |A| =
P
x µA (x).
These are the kinds of problems MAKAUT examiners actually ask. Master them.
Problem (PYQ-style): A 4-input neuron has weights w = [1, 2, 3, 4]. Transfer function is
linear with proportionality constant k=2 (i.e., ϕ(z) = 2z ). Inputs are x = [4, 10, 5, 20]. No
bias. Find the output.
X
wi xi = 1(4) + 2(10) + 3(5) + 4(20) = 4 + 20 + 15 + 80 = 119.
i
/! Examiner Trap
Linear does NOT mean no activation. Linear with proportionality constant k means
ϕ(z) = kz . Identity is the special case k = 1. If you forget this, you'd write 119 (wrong).
Problem: A network has 3 input neurons, 2 hidden layers each with 8 neurons, and 3 output
neurons. Find:
(a) total biases, (b) total weights, (c) best loss + output activation.
btotal = 8 + 8 + 3 = 19 .
al
W [1] : 8 × 3 = 24
W [2] : 8 × 8 = 64
s
W [3] : 3 × 8 = 24
an
Total weights = 24 + 64 + 24 = 112 .
Step 3 output activation & loss: 3 outputs ⇒ multi-class classication ⇒ softmax +
categorical cross-entropy (CCE).
Problem (PYQ-style): A 4-layer network: x → a[1] → a[2] → a[3] → a[4] has layer sizes 3, 5,
3, 1.
nm
(d) Size of W
[2] , W [3] ?
(c) Because each layer has multiple input AND multiple output units, a 2-D structure (matrix)
is needed to encode all pairwise weights. A vector wouldn't suce. Biases are vectors because each
output unit has exactly one bias.
(d) W [2] ∈ R5×3 (3 → 5); W [3] ∈ R3×5 (5 → 3).
(e) If ϕ(z) = z everywhere except output, the entire hidden stack collapses:
Profound insight: a deep linear network is no more powerful than a single linear layer.
Non-linearity is what makes depth meaningful. This is a favourite examiner question.
Problem: Derive σ ′ (z). At what z is the gradient maximum? What is the max value?
e−z
σ ′ (z) = −(1 + e−z )−2 · (−e−z ) =
(1 + e−z )2
sal
1 e−z
= ·
1 + e−z 1 + e−z
1 1
= · 1−
1 + e−z 1 + e−z
Intuition
This is the mathematical origin of vanishing gradients. Through 10 sigmoid layers, gradient is
multiplied by at most (0.25)10 ≈ 9.5 × 10−7 . Eectively no learning reaches early layers.
mo
Beyond sigmoid/tanh/ReLU, modern DL uses many variants. Examiners sometimes ask comparison
An
al
SELU Scaled ELU Self-normalising activations
GELU z · Φ(z) (Gauss CDF) Used in BERT, GPT
Swish / SiLU z · σ(z) Found by NAS; smooth
Softplus Smooth ReLU approximation
s
ln(1 + ez )
Maxout maxi (wiT x + bi ) Generalises ReLU/Leaky-
an
ReLU
Softmax ezi / ezj Multi-class output
P
j
ezi e zk − e zi · e zi
P
∂ ŷi kP
= zk 2
∂zi ( ke )
= ŷi − ŷi2 = ŷi (1 − ŷi ).
ol
O-diagonal (i ̸= j ):
∂ ŷi 0 − ezi · ezj
= P z 2 = −ŷi ŷj .
∂zj ( k e k)
nm
Compactly: ∂ ŷi /∂zj = ŷi (δij − ŷj ) where δij is Kronecker delta.
Proof:
The clean cancellation is why softmax+CCE is the canonical pairing, just as sig-
moid+BCE gives ŷ − y . The complicated activation derivative cancels with the loss derivative.
All zeros: every neuron computes the same output, gets the same gradient, stays identical
forever. Symmetry never breaks.
Too small: activations shrink layer by layer (vanishing forward and backward).
sal
Too large: activations explode.
The principle: keep variance of activations and gradients approximately constant across layers.
Xavier / Glorot (for tanh / sigmoid):
2
W ∼ N 0,
nin + nout
an
He / Kaiming (for ReLU; accounts for half-activation):
2
W ∼ N 0,
nin
Why the factor of 2 in He? ReLU zeros out half the activations on average, so to maintain
√
variance, scale up by 2.
lK
Bias init: usually 0 (or small positive for ReLU to keep neurons active early).
Orthogonal init: good for RNN Whh avoids vanishing/exploding when iterated over time.
µB = xi (batch mean)
m
i=1
m
2 1 X
σB = (xi − µB )2 (batch variance)
m
i=1
x i − µB
x̂i = q (normalize)
2 +ϵ
σB
An
q γ, β are learnable, allowing the network to undo normalisation if useful (e.g., identity if γ =
2 + ϵ and β = µ ).
σB B
Test-time behaviour: use running averages of µ, σ 2 computed during training, not the test
batch's statistics.
Benets:
Layer Norm: normalize across feature dim for each example. Used in Transformers, RNNs.
s al
an
K
ol
nm
A
Official Store
[Link]