0% found this document useful (0 votes)
0 views26 pages

DeepLearning Notes Sample

Uploaded by

soumouchiha
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
0 views26 pages

DeepLearning Notes Sample

Uploaded by

soumouchiha
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

DEEPLEARNING

Semester Notes  Book 1 of 2


Paper Code: PCCAIML602 • [Link] AIML (Sem 6)

al
Made by Anmol Kansal
Maulana Abul Kalam Azad University of Technology, West Bengal

s
May 26, 2026

an
Your Complete Concept Handbook
K
This is Book 1 of 2. It teaches you every concept in your syllabus
from rst principles, with diagrams, intuition, and memory hooks.

For PYQ solutions, expected questions,


ol

MCQs, and numerical practice, refer to:

Book 2  PYQ & Practice Book


nm

How to Use This Book

Step 1: Pick a Fast-Track Path (next page).


Step 2: Read each unit's overview + diagrams rst.
Step 3: Read the text, then revisit highlighted boxes.
A

Step 4: Use the Quick Revision section at the end.


Step 5: Solve PYQs from Book 2.
Deep Learning Notes Made by Anmol Kansal

Fast-Track Study Paths


Pick the Path That Matches Your Time
Don't try to read everything if you don't have time. Pick a path based on how many days
you have left.

Time Left What to Study

1 day Quick Revision Handbook (end of book) + Book 2 PYQs only. Skim
formula sheet and comparison tables.

al
3 days Read Units 2, 3, 5 thoroughly (these carry 40 of 70 marks). Then
Quick Revision + Book 2 PYQs.

7 days All 6 units from this book. Then Book 2 PYQs and expected ques-

s
tions.

Full prep Both books cover-to-cover, including Book 2's MCQs and numericals.

Marks Weightage by Unit


an
K
Unit Topic Hours Marks Priority

1 Introduction 3 5 Low

2 Feed Forward NN 6 10 Medium


ol

3 Training Neural Networks 6 15 HIGH

4 Conditional Random 9 15 HIGH


Fields
nm

5 Deep Learning 6 15 HIGH


(CNN/RNN/DBN)

6 DL Research (Applications) 6 10 Medium

Subtotal: 36 70
A

!!! MUST MEMORIZE

Strategic insight: Units 3, 4, 5 together carry 45 of 70 marks (64%). If short on time,


focus here.

How to Read the Color-Coded Boxes

PCCAIML602 Anmol Kansal 2


Deep Learning Notes Made by Anmol Kansal

Color Meaning

Red Must memorize  formulas, denitions, key facts

Blue Concept explanation, intuition, theory

Green PYQ favourite  this exact topic appeared / will appear

Yellow Examiner trap  the tricky bit students miss

Purple Intuition / mental model

Gray Memory hook / mnemonic

s al
an
K
ol
nm
A

PCCAIML602 Anmol Kansal 3


Deep Learning Notes Made by Anmol Kansal

Contents

Fast-Track Study Paths 2

Marks Weightage by Unit 2

Color Coding Guide 3

1 Unit 1: Introduction to Deep Learning 9


1.1 What is Deep Learning? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9
1.1.1 AI vs ML vs DL Hierarchy . . . . . . . . . . . . . . . . . . . . . . . . . . . 9
1.2 Three Learning Paradigms . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10

al
1.2.1 Supervised Learning  Two Flavours . . . . . . . . . . . . . . . . . . . . . . 10
1.2.2 Unsupervised Learning  Three Flavours . . . . . . . . . . . . . . . . . . . . 10
1.2.3 Reinforcement Learning  Key Terms . . . . . . . . . . . . . . . . . . . . . . 10
1.3 Why Deep Learning Works Now (and Didn't in 1990) . . . . . . . . . . . . . . . . 10

s
1.4 Issues / Challenges in Deep Learning . . . . . . . . . . . . . . . . . . . . . . . . . . 11
1.5 Quick Review of Fundamental Learning Techniques . . . . . . . . . . . . . . . . . . 11

an
1.6 Evaluation Metrics for Classication . . . . . . . . . . . . . . . . . . . . . . . . . . 11
1.7 Bias-Variance Tradeo . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12
1.8 Deep Dive: Worked Examples . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12
1.8.1 Worked Example 1  Confusion Matrix Metrics . . . . . . . . . . . . . . . . 12
1.8.2 Worked Example 2  Why F1 Beats Accuracy (Imbalanced Data) . . . . . . 13
K
1.8.3 Worked Example 3  Confusion Matrix with Zero O-Diagonals . . . . . . . 13
1.8.4 Worked Example 4  Entropy of a Coin . . . . . . . . . . . . . . . . . . . . 13

2 Unit 2: Feed Forward Neural Network 15


2.1 The Biological Inspiration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15
ol

2.2 Activation Functions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15


2.2.1 Sigmoid . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15
2.2.2 Tanh . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16
2.2.3 ReLU . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16
nm

2.2.4 Softmax (output for multi-class) . . . . . . . . . . . . . . . . . . . . . . . . 16


2.2.5 Activation Comparison Table . . . . . . . . . . . . . . . . . . . . . . . . . . 16
2.3 Multi-Layer Perceptron (MLP) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17
2.3.1 Dimensions (memorize this!) . . . . . . . . . . . . . . . . . . . . . . . . . . . 17
2.3.2 Parameter Counting Formula . . . . . . . . . . . . . . . . . . . . . . . . . . 17
2.4 Universal Approximation Theorem . . . . . . . . . . . . . . . . . . . . . . . . . . . 18
2.5 Forward Propagation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18
A

2.6 Linear Separability and XOR . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18


2.7 Perceptron vs Logistic Regression . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18
2.8 Delta Rule (Widrow-Ho ) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19
2.9 Fuzzy Relations (Brief ) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19
2.10 Deep Dive: Worked Examples . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19
2.10.1 Worked Example 1  Linear Neuron Output . . . . . . . . . . . . . . . . . . 19
2.10.2 Worked Example 2  Parameter Counting (3-8-8-3 Network) . . . . . . . . . 20
2.10.3 Worked Example 3  Layer Dimensions & Forward Pass . . . . . . . . . . . 20
2.10.4 Worked Example 4  Sigmoid Derivative Derivation . . . . . . . . . . . . . 21
2.11 Advanced Topics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21
2.11.1 The Complete Activation Function Family . . . . . . . . . . . . . . . . . . . 21
2.11.2 Softmax Derivative (Often Asked) . . . . . . . . . . . . . . . . . . . . . . . 22
2.11.3 Softmax + CCE Gradient Simplies Beautifully . . . . . . . . . . . . . . . . 22

PCCAIML602 Anmol Kansal 4


Deep Learning Notes Made by Anmol Kansal

2.11.4 Weight Initialization Theory . . . . . . . . . . . . . . . . . . . . . . . . . . . 23


2.11.5 Batch Normalization: Full Forward Pass . . . . . . . . . . . . . . . . . . . . 23

3 Unit 3: Training Neural Networks 25


3.1 Risk Minimization and Loss Functions . . . . . . . . . . . . . . . . . . . . . . . . . 25
3.1.1 Common Loss Functions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25
3.1.2 Why Cross-Entropy is Used (Derivation) . . . . . . . . . . . . . . . . . . . . 25
3.2 Gradient Descent . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 26
3.2.1 Three Variants . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 26
3.2.2 Learning Rate Eects . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 26
3.3 Backpropagation  The Heart of Training . . . . . . . . . . . . . . . . . . . . . . . 26
3.3.1 The Two Passes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 26
3.3.2 The Recursion (Memorize!) . . . . . . . . . . . . . . . . . . . . . . . . . . . 27

al
3.3.3 Tiny Worked Example (one weight) . . . . . . . . . . . . . . . . . . . . . . 27
3.4 Vanishing and Exploding Gradients . . . . . . . . . . . . . . . . . . . . . . . . . . . 27
3.4.1 How to Fix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27

s
3.5 Regularization Techniques . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28
3.5.1 L1 vs L2 Regularization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28

an
3.5.2 Dropout . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28
3.5.3 Other Regularization Techniques . . . . . . . . . . . . . . . . . . . . . . . . 28
3.6 Optimizers Beyond SGD . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28
3.6.1 Momentum . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29
3.6.2 RMSProp . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29
K
3.6.3 Adam (the default in 2024+) . . . . . . . . . . . . . . . . . . . . . . . . . . 29
3.6.4 Optimizer Comparison . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29
3.7 Model Selection . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29
3.7.1 Train / Validation / Test Split . . . . . . . . . . . . . . . . . . . . . . . . . 29
3.7.2 k-Fold Cross-Validation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30
ol

3.8 Deep Dive: Worked Examples . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30


3.8.1 Worked Example 1  Log Loss from MLE (Full Derivation) . . . . . . . . . 30
3.8.2 Worked Example 2  Categorical Cross-Entropy Numerical . . . . . . . . . 30
3.8.3 Worked Example 3  Gradient Descent Update Derivation . . . . . . . . . . 31
nm

3.8.4 Worked Example 4  Backprop on 3-Input, 2-Hidden, 1-Output MLP . . . . 31


3.8.5 Worked Example 5  LR = 0, Vanishing Gradient Causes . . . . . . . . . . 32
3.8.6 Worked Example 6  Vanishing Gradient with Sigmoid (Numerical) . . . . . 32
3.8.7 Worked Example 7  Logistic Regression Weight Scaling Trick . . . . . . . . 33
3.9 Advanced Topics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 33
3.9.1 Momentum  Why It Helps . . . . . . . . . . . . . . . . . . . . . . . . . . . 33
3.9.2 AdaGrad  Per-Parameter Learning Rates . . . . . . . . . . . . . . . . . . . 33
A

3.9.3 RMSProp  Exponential Moving Average Fix . . . . . . . . . . . . . . . . . 34


3.9.4 Adam  The Default Modern Optimizer . . . . . . . . . . . . . . . . . . . . 34
3.9.5 Optimizer Comparison Table . . . . . . . . . . . . . . . . . . . . . . . . . . 34
3.9.6 The Optimization Landscape of Deep Nets . . . . . . . . . . . . . . . . . . . 34
3.9.7 Why Newton's Method is Rarely Used . . . . . . . . . . . . . . . . . . . . . 34
3.9.8 Learning Rate Schedules . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 35
3.9.9 Regularization  The Complete Catalogue . . . . . . . . . . . . . . . . . . . 35
3.9.10 Cross-Validation Strategies . . . . . . . . . . . . . . . . . . . . . . . . . . . 35

4 Unit 4: Conditional Random Fields & Probabilistic Models 36


4.1 Probabilistic Graphical Models  The Big Picture . . . . . . . . . . . . . . . . . . . 36
4.2 Bayesian Network . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 36
4.3 Markov Network (Markov Random Field) . . . . . . . . . . . . . . . . . . . . . . . 36
4.4 Linear-Chain Structure . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 37

PCCAIML602 Anmol Kansal 5


Deep Learning Notes Made by Anmol Kansal

4.5 Hidden Markov Model (HMM) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 37


4.5.1 The Three Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 37
4.5.2 Forward Algorithm . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 37
4.5.3 Viterbi Algorithm . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 38
4.6 Conditional Random Field (CRF) . . . . . . . . . . . . . . . . . . . . . . . . . . . . 38
4.6.1 HMM vs CRF . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 38
4.7 Belief Propagation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 38
4.8 Entropy and Information Theory . . . . . . . . . . . . . . . . . . . . . . . . . . . . 39
4.8.1 Properties . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 39
4.8.2 Other Quantities . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 39
4.9 Deep Dive: Worked Examples . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 39
4.9.1 Worked Example 1  Bayes Network Joint Probability . . . . . . . . . . . . 39

al
4.9.2 Worked Example 2  HMM Forward Algorithm (Numerical) . . . . . . . . . 40
4.9.3 Worked Example 3  Viterbi vs Forward (Same Recursion, Dierent Operator) 40
4.9.4 Worked Example 4  Entropy Calculations . . . . . . . . . . . . . . . . . . . 40
4.9.5 Worked Example 5  HMM Three Problems Mapped to Real Tasks . . . . . 41

s
4.10 Advanced Topics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 41
4.10.1 D-Separation in Bayes Networks . . . . . . . . . . . . . . . . . . . . . . . . 41

an
4.10.2 Conditional Independence in HMM (Markov Property) . . . . . . . . . . . . 41
4.10.3 The Forward-Backward Algorithm . . . . . . . . . . . . . . . . . . . . . . . 41
4.10.4 Baum-Welch (EM for HMM) . . . . . . . . . . . . . . . . . . . . . . . . . . 42
4.10.5 Linear-Chain CRF in Detail . . . . . . . . . . . . . . . . . . . . . . . . . . . 42
4.10.6 HMM vs CRF: When to Use Each . . . . . . . . . . . . . . . . . . . . . . . 43
K
4.10.7 Belief Propagation: The General Framework . . . . . . . . . . . . . . . . . . 43
4.10.8 Cross-Entropy = Entropy + KL Divergence . . . . . . . . . . . . . . . . . . 43
4.10.9 Mutual Information . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 44

5 Unit 5: Deep Learning Architectures 45


ol

5.1 Deep Feed Forward Network (DFNN) . . . . . . . . . . . . . . . . . . . . . . . . . 45


5.2 Convolutional Neural Network (CNN) . . . . . . . . . . . . . . . . . . . . . . . . . 45
5.2.1 The Convolution Operation . . . . . . . . . . . . . . . . . . . . . . . . . . . 45
5.2.2 Key Hyperparameters . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 46
nm

5.2.3 Pooling Layer . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 46


5.2.4 Why Weight Sharing Saves Parameters . . . . . . . . . . . . . . . . . . . . . 46
5.2.5 Typical CNN Architecture . . . . . . . . . . . . . . . . . . . . . . . . . . . . 47
5.2.6 Famous CNN Architectures . . . . . . . . . . . . . . . . . . . . . . . . . . . 47
5.2.7 ResNet's Skip Connection  Why It Works . . . . . . . . . . . . . . . . . . 47
5.3 Recurrent Neural Network (RNN) . . . . . . . . . . . . . . . . . . . . . . . . . . . . 47
5.3.1 Vanilla RNN . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 47
A

5.3.2 Why Standard NN Fails on Sequences . . . . . . . . . . . . . . . . . . . . . 48


5.3.3 RNN Architecture Variants . . . . . . . . . . . . . . . . . . . . . . . . . . . 48
5.3.4 Backpropagation Through Time (BPTT) . . . . . . . . . . . . . . . . . . . 48
5.4 Long Short-Term Memory (LSTM) . . . . . . . . . . . . . . . . . . . . . . . . . . . 48
5.5 Gated Recurrent Unit (GRU) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 49
5.5.1 LSTM vs GRU . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 49
5.6 Bidirectional RNN . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 49
5.7 Deep Belief Network (DBN) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 49
5.7.1 Restricted Boltzmann Machine . . . . . . . . . . . . . . . . . . . . . . . . . 50
5.8 Deep Dive: Worked Examples . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 50
5.8.1 Worked Example 1  CNN Output Size Calculation . . . . . . . . . . . . . . 50
5.8.2 Worked Example 2  Convolution Operation by Hand . . . . . . . . . . . . 50
5.8.3 Worked Example 3  Pooling with Two Examples . . . . . . . . . . . . . . . 51

PCCAIML602 Anmol Kansal 6


Deep Learning Notes Made by Anmol Kansal

5.8.4 Worked Example 4  Designing a CNN for MNIST . . . . . . . . . . . . . . 51


5.8.5 Worked Example 5  Weight Sharing Saves Parameters . . . . . . . . . . . . 52
5.8.6 Worked Example 6  One-hot Encoding for NER . . . . . . . . . . . . . . . 52
5.8.7 Worked Example 7  Why Standard NN Fails on Sequences . . . . . . . . . 53
5.8.8 Worked Example 8  LSTM Gate Trace . . . . . . . . . . . . . . . . . . . . 53
5.9 Advanced Topics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 53
5.9.1 Receptive Field Calculation . . . . . . . . . . . . . . . . . . . . . . . . . . . 53
5.9.2 Special Convolutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 54
5.9.3 Famous CNN Architectures: Deeper Comparison . . . . . . . . . . . . . . . 54
5.9.4 Inception Module Architecture . . . . . . . . . . . . . . . . . . . . . . . . . 54
5.9.5 ResNet Residual Block Mathematics . . . . . . . . . . . . . . . . . . . . . . 55
5.9.6 Backpropagation Through Time (BPTT)  Full Derivation . . . . . . . . . . 55

al
5.9.7 Why LSTM Solves Vanishing Gradient . . . . . . . . . . . . . . . . . . . . . 55
5.9.8 GRU vs LSTM: Detailed Comparison . . . . . . . . . . . . . . . . . . . . . 55
5.9.9 Attention Mechanism: The Math . . . . . . . . . . . . . . . . . . . . . . . . 56
5.9.10 Transformer Encoder Block in Detail . . . . . . . . . . . . . . . . . . . . . . 56

s
5.9.11 Autoencoders: The Family . . . . . . . . . . . . . . . . . . . . . . . . . . . 57
5.9.12 Restricted Boltzmann Machine  Energy and Training . . . . . . . . . . . . 57

an
5.9.13 Deep Belief Network: Layer-Wise Pre-training . . . . . . . . . . . . . . . . . 57

6 Unit 6: Deep Learning Research & Applications 59


6.1 Object Recognition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 59
6.1.1 Three Related Tasks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 59
K
6.1.2 Two-Stage Detectors (R-CNN Family) . . . . . . . . . . . . . . . . . . . . . 59
6.1.3 One-Stage Detectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 59
6.1.4 Evaluation Metrics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 59
6.2 Sparse Coding . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 59
6.3 Computer Vision Task Summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . 60
ol

6.4 Natural Language Processing (NLP) . . . . . . . . . . . . . . . . . . . . . . . . . . 60


6.4.1 Word Embeddings . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 60
6.4.2 Tokenization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 60
6.4.3 Encoder-Decoder for Machine Translation . . . . . . . . . . . . . . . . . . . 60
nm

6.4.4 Attention . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 60
6.4.5 Transformer . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 61
6.4.6 BERT and GPT . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 61
6.5 Deep Dive: PCA (Was a 15-Mark Question in 2023!) . . . . . . . . . . . . . . . . . 61
6.5.1 The Goal . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 61
6.5.2 The Optimisation Problem . . . . . . . . . . . . . . . . . . . . . . . . . . . 61
6.5.3 Solve via Lagrangian . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 61
A

6.5.4 The PCA Algorithm . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 62


6.5.5 Worked Numerical Example . . . . . . . . . . . . . . . . . . . . . . . . . . . 62
6.5.6 Explained Variance Ratio . . . . . . . . . . . . . . . . . . . . . . . . . . . . 62
6.5.7 Applications of PCA . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 62
6.5.8 Limitations of PCA . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 63
6.6 Advanced Topics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 63
6.6.1 Object Detection Evolution  Full Timeline . . . . . . . . . . . . . . . . . . 63
6.6.2 Detection Metrics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 63
6.6.3 Sparse Coding  Full Theory . . . . . . . . . . . . . . . . . . . . . . . . . . 64
6.6.4 Word Embeddings  Word2Vec Details . . . . . . . . . . . . . . . . . . . . . 64
6.6.5 Tokenization in Modern NLP . . . . . . . . . . . . . . . . . . . . . . . . . . 64
6.6.6 BERT vs GPT  Detailed Comparison . . . . . . . . . . . . . . . . . . . . . 65
6.6.7 T5 and BART: Encoder-Decoder Hybrids . . . . . . . . . . . . . . . . . . . 65

PCCAIML602 Anmol Kansal 7


Deep Learning Notes Made by Anmol Kansal

6.6.8 GAN  Generative Adversarial Networks . . . . . . . . . . . . . . . . . . . . 65


6.6.9 Diusion Models (Modern Generative SOTA) . . . . . . . . . . . . . . . . . 66
6.6.10 Transfer Learning Mechanisms . . . . . . . . . . . . . . . . . . . . . . . . . 66
6.6.11 Modern Topics Likely on Future Exams . . . . . . . . . . . . . . . . . . . . 66

7 Quick Revision Handbook 68


7.1 Formula Sheet . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 68
7.2 The Master Comparison Tables . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 69
7.2.1 Activations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 69
7.2.2 GD Variants . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 70
7.2.3 CNN vs RNN vs Transformer . . . . . . . . . . . . . . . . . . . . . . . . . . 70
7.2.4 L1 vs L2 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 70
7.2.5 HMM vs CRF . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 70

al
7.2.6 LSTM vs GRU . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 70
7.3 Critical Memory Hooks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 70
7.4 Architecture Quick Sketches . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 71

s
7.5 Top 10 Things You Must Know . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 72
7.6 The 24-Hour Final Strategy . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 72

an
K
ol
nm
A

PCCAIML602 Anmol Kansal 8


Deep Learning Notes Made by Anmol Kansal

1 Unit 1: Introduction to Deep Learning


Unit 1 Snapshot
Hours: 3 • Marks: 5 • Priority: Low
What you'll learn: What deep learning is, the three learning paradigms, why DL works
now, the big issues, and a quick review of classical ML.

Common exam questions: Dene deep learning. Dierences between super-


vised/unsupervised. F1 vs accuracy. Bias-variance.

1.1 What is Deep Learning?

al
 Concept

Deep Learning is a sub-eld of Machine Learning that uses multi-layered neural networks
to automatically learn hierarchical representations from raw data  without needing humans

s
to design features manually.

an
Plain English: Instead of telling the computer a face has two eyes, a nose, a mouth, look for
these, we just show it thousands of face photos. The network gures out by itself  across many
layers  that edges combine into shapes, shapes combine into features (eyes, noses), and features
combine into faces.
K
2-Mark Ready Answer

Q: What is Deep Learning?


Deep Learning is a branch of machine learning that uses neural networks with multiple hidden
layers to learn hierarchical feature representations directly from data, eliminating the need for
manual feature engineering.
ol

1.1.1 AI vs ML vs DL Hierarchy
nm

Articial Intelligence

Machine Learning

Deep
Learning
A

ˆ AI (1950s): Any system that mimics human intelligence. Includes rule-based expert systems.

ˆ ML (1980s): Systems that learn from data. Includes decision trees, SVMs, neural networks.

ˆ DL (2010s): ML with deep neural networks. Powered by big data + GPU + new algorithms.

PCCAIML602 Anmol Kansal 9


Deep Learning Notes Made by Anmol Kansal

1.2 Three Learning Paradigms


Paradigm What it learns Example

Supervised Mapping x→y from labelled Email spam lter, digit recog-
data nition

Unsupervised Hidden structure in unlabelled Customer clustering, dimen-


data sionality reduction

Reinforcement Optimal action from re- Chess engine, robot walking


ward/penalty signal

One-Line Memory Hook

al
Mnemonic SUR : Supervised = with teacher, Unsupervised = self-organize, Reinforcement
= trial-and-error.

s
1.2.1 Supervised Learning  Two Flavours

an
ˆ Classication: output is a category. Binary (spam/not), multi-class (digit 0-9), multi-label.

ˆ Regression: output is a number. House prices, temperature forecast, stock value.

1.2.2 Unsupervised Learning  Three Flavours


K
ˆ Clustering: group similar items. K-Means, DBSCAN.

ˆ Dimensionality reduction: compress data while keeping structure. PCA, t-SNE, autoen-
coders.

ˆ Density estimation: model p(x). GMM, VAE, GAN.


ol

1.2.3 Reinforcement Learning  Key Terms

ˆ Agent: the learner.


nm

ˆ Environment: the world the agent acts in.

ˆ State s: snapshot of the world.

ˆ Action a: what the agent does.

ˆ Reward r: feedback (positive or negative).

ˆ Policy π(a|s): agent's strategy.


A

Intuition

RL feels like training a dog with treats. The dog (agent) takes actions (sit, jump). Good
actions get treats (reward); bad actions get nothing (or scolding = penalty). Over time the
dog learns the policy: which action in which situation maximises treats.

1.3 Why Deep Learning Works Now (and Didn't in 1990)


Key Formula

The DL Revolution Equation:

Big Data + GPU Compute + Better Algorithms = Deep Learning

PCCAIML602 Anmol Kansal 10


Deep Learning Notes Made by Anmol Kansal

1. Big Data: ImageNet (14M images, 2009), Wikipedia, YouTube, Common Crawl. Networks
need millions of examples.

2. GPU Compute: NVIDIA's CUDA (2007) made parallel matrix multiplication 100× faster
than CPU.

3. Algorithms: ReLU (2010), Dropout (2012), Batch Norm (2015), Residual connections (2015),
Adam (2015), Attention/Transformer (2017).

1.4 Issues / Challenges in Deep Learning


!!! MUST MEMORIZE

Mnemonic DOVE-BACH  8 key issues:

al
Data-hungry • Overtting • Vanishing gradients • Expensive (compute) • Black-box •
Adversarial vulnerability • Catastrophic forgetting • Hyperparameter sensitivity

1.5 Quick Review of Fundamental Learning Techniques

s
an
Method Type Key Idea

Linear Regres- Supervised regr. y = wT x + b, minimise MSE


sion

Logistic Re- Supervised class. Sigmoid output, minimise log-


K
gression loss

k-NN Supervised class./regr. Predict based on k closest


training points

SVM Supervised class. Find max-margin hyperplane

Decision Tree Supervised Recursive split on best feature


ol

K-Means Unsupervised Iteratively assign points to k


centroids

PCA Unsupervised Project to directions of max


nm

variance

1.6 Evaluation Metrics for Classication


 Concept

Confusion Matrix:
A

Predicted + Predicted 

Actual + TP FN

Actual  FP TN

PCCAIML602 Anmol Kansal 11


Deep Learning Notes Made by Anmol Kansal

Key Formula

TP + TN
Accuracy =
TP + TN + FP + FN
TP
Precision = (of those predicted +, how many are really +?)
TP + FP
TP
Recall = (of all real +, how many did we catch?)
TP + FN
2·P ·R
F1 = (harmonic mean)
P +R

*** PYQ Favourite

al
F1 vs Accuracy  a recurring question.
Accuracy is misleading on imbalanced datasets. If 99% of emails are not spam, predicting

s
not spam always gives 99% accuracy but 0% recall on spam. F1 balances precision and recall,
catching this failure.

/! Examiner Trap

an
Trap: If all o-diagonal entries of the confusion matrix are zero, the classier is perfect on
this dataset (zero mistakes). But verify on a held-out test set  it could be overtting!
K
1.7 Bias-Variance Tradeo

error sweet spot variance


total
ol
nm

bias2
model complexity

ˆ High bias = undertting: model too simple, misses the pattern.

ˆ High variance = overtting: model too complex, memorises noise.


A

ˆ Goal: nd the sweet spot.

One-Line Memory Hook

Bias = how wrong on average. Variance = how much answers wobble across runs.
Need both low.

1.8 Deep Dive: Worked Examples


1.8.1 Worked Example 1  Confusion Matrix Metrics

*** PYQ Favourite

Problem: A classier produces TP = 80, TN = 90, FP = 10, FN = 20. Compute Accuracy,


Precision, Recall, F1.

PCCAIML602 Anmol Kansal 12


Deep Learning Notes Made by Anmol Kansal

Apply each formula:

TP + TN 80 + 90
Acc = = = 0.85
Total 200
TP 80
Prec = = = 0.889
TP + FP 90
TP 80
Rec = = = 0.80
TP + FN 100
2·P ·R 2(0.889)(0.80)
F1 = = = 0.842
P +R 0.889 + 0.80

1.8.2 Worked Example 2  Why F1 Beats Accuracy (Imbalanced Data)

al
*** PYQ Favourite

Problem (PYQ 2024 Q7f, 2M): Why is F1 better than accuracy?

s
Concrete example: 99% of emails are not spam. A trivial always not spam classier:

an
ˆ TP = 0 (never predicts spam)

ˆ TN = 99

ˆ FP = 0

ˆ FN = 1 (the real spam was missed)


K
Compute:

ˆ Accuracy = 99/100 = 99%  looks great!

ˆ Recall = 0/1 = 0%  catches NO spam.

ˆ
ol

Precision is undened (or 0).

ˆ F1 = 0.

Lesson: Accuracy is misleading on imbalanced data. F1 punishes models that ignore the
nm

minority class. Use F1 for spam, fraud detection, disease screening.

1.8.3 Worked Example 3  Confusion Matrix with Zero O-Diagonals

*** PYQ Favourite

Problem (PYQ 2024 Q7g, 1M): If all o-diagonal entries of the confusion matrix are zero,
A

what can you infer?

Answer: O-diagonal entries are the misclassications. If all are zero, there are no mistakes
⇒ the classier is perfect on this dataset. Accuracy = Precision = Recall = F1 = 1.0.
Caveat: If this is the training set, the model might be overtting. Verify on a held-out test
set before celebrating.

1.8.4 Worked Example 4  Entropy of a Coin

*** PYQ Favourite

Problem: Compute the entropy (in bits) of (a) a fair coin, (b) a heavily biased coin with
p(H) = 0.9.
P
Formula: H=− x p(x) log2 p(x).

PCCAIML602 Anmol Kansal 13


Deep Learning Notes Made by Anmol Kansal

(a) Fair coin:H = −2(0.5 log2 0.5) = −2(0.5)(−1) = 1.0 bit .


(b) Biased p = 0.9:

H = −[0.9 log2 0.9 + 0.1 log2 0.1] ≈ −[0.9(−0.152) + 0.1(−3.322)] = 0.469 bit .

Intuition

Less uncertainty ⇒ less entropy. A biased coin is more predictable, so its entropy is lower.
Maximum entropy is achieved by the uniform distribution.

Unit 1 Quick Recap


ˆ DL = ML with deep neural networks.

al
ˆ 3 paradigms: supervised, unsupervised, reinforcement.

ˆ DL boom = big data + GPU + algorithms.

ˆ DOVE-BACH.

s
Issues:

ˆ F1 > accuracy on imbalanced data.

ˆ Bias-variance tradeo = undert vs overt.

an
K
ol
nm
A

PCCAIML602 Anmol Kansal 14


Deep Learning Notes Made by Anmol Kansal

2 Unit 2: Feed Forward Neural Network


Unit 2 Snapshot
Hours: 6 • Marks: 10 • Priority: Medium
What you'll learn: The neuron model, activation functions, multi-layer perceptrons (MLPs),
forward propagation, fuzzy relations.

How examiners ask: Dene ANN. Linear neuron numerical. Sigmoid range and derivative.
Why ReLU sparse? Parameter count.

2.1 The Biological Inspiration

al
A real neuron has dendrites (receive signals), a soma (cell body that decides whether to re),
and an axon (sends signal out). Neurons connect at synapses with varying strengths.
Articial neuron mimics this minimally: weighted inputs → sum → activation → output.

s
x1

an
w1

x2 w2

Σ ϕ y
x3 w3

b
K
1 (bias)
weighted sum activation

Key Formula

The Articial Neuron:


ol

n
!
X
y=ϕ wi x i + b
i=1

where xi are inputs, wi are weights, b is bias, ϕ is the activation function.


nm

One-Line Memory Hook

Order WIB-A : Weight × Input, sum, add Bias, apply Activation. Always.

2.2 Activation Functions


A

2.2.1 Sigmoid

Key Formula

1
σ(z) = σ ′ (z) = σ(z)(1 − σ(z))
1 + e−z
Range: (0, 1) Max gradient: 0.25 at z = 0.

Use: binary classication output layer.


Problem: saturates at extremes → vanishing gradient.

PCCAIML602 Anmol Kansal 15


Deep Learning Notes Made by Anmol Kansal

2.2.2 Tanh

Key Formula

ez − e−z
tanh(z) = tanh′ (z) = 1 − tanh2 (z)
ez + e−z
Range: (−1, 1) Max gradient: 1 at z = 0.

Use: RNN hidden states (zero-centred is better than sigmoid).

2.2.3 ReLU

Key Formula

al
(
′ 1 z>0
ReLU(z) = max(0, z) ReLU (z) =
0 z≤0

s
Use: hidden layers everywhere. Default choice for CNNs/MLPs.

an
Why ReLU is great:

ˆ Gradient = 1 for z>0→ no vanishing.

ˆ Computationally cheap (just a max).

ˆ Produces sparse activations (about 50% zero)  biologically plausible.


K
/! Examiner Trap

Dead ReLU problem: if a neuron's input becomes < 0 for all training examples, its gradient
is permanently 0 and it never recovers. Fix: use Leaky ReLU (αz for z < 0) or ELU.
ol

2.2.4 Softmax (output for multi-class)

Key Formula
nm

ezi
softmax(z)i = PK
zj
j=1 e

Outputs are probabilities (sum to 1, each in (0, 1)).

2.2.5 Activation Comparison Table


A

Function Range Zero-centred? Max grad Vanish?

Sigmoid (0, 1) No 0.25 Severe

Tanh (−1, 1) Yes 1.0 Moderate

ReLU [0, ∞) No 1.0 No (for z > 0)


Leaky ReLU R Almost 1.0 No

Softmax (0, 1), sum=1   Used at output

PCCAIML602 Anmol Kansal 16


Deep Learning Notes Made by Anmol Kansal

!!! MUST MEMORIZE

Output activation ↔ task:

ˆ Regression → Linear (identity) + MSE loss

ˆ Binary classication → Sigmoid + BCE loss

ˆ Multi-class (one of K) → Softmax + CCE loss

ˆ Multi-label (any of K) → Sigmoid per class + BCE

2.3 Multi-Layer Perceptron (MLP)


 Concept

al
An MLP is a stack of fully-connected layers, each consisting of an ane transformation followed
by a non-linear activation: a[ℓ] = ϕ(W [ℓ] a[ℓ−1] + b[ℓ] ).

s
an
K
Input (3) Hidden (4) Output (2)

2.3.1 Dimensions (memorize this!)


ol

!!! MUST MEMORIZE

For a layer with n[ℓ−1] inputs and n[ℓ] outputs:


nm

[ℓ] ×n[ℓ−1] [ℓ]


W [ℓ] ∈ Rn , b[ℓ] ∈ Rn

2.3.2 Parameter Counting Formula

Key Formula
A

Total parameters in a layer:

[ℓ]
params = n[ℓ] · (n[ℓ−1] + 1)

The +1 accounts for the bias.

Quick example: for a 3→8→8→3 network:

ˆ Layer 1: 8(3 + 1) = 32
ˆ Layer 2: 8(8 + 1) = 72
ˆ Layer 3: 3(8 + 1) = 27
ˆ Total: 131 parameters (112 weights + 19 biases)

PCCAIML602 Anmol Kansal 17


Deep Learning Notes Made by Anmol Kansal

2.4 Universal Approximation Theorem


 Concept

A feed-forward network with one hidden layer and a non-polynomial activation can ap-
proximate any continuous function on a compact set, to arbitrary accuracy, given enough
neurons.

So why go deep? The theorem says one layer can work, but might need exponentially
many neurons. Deep networks achieve the same with exponentially fewer parameters, thanks to
hierarchical feature composition.

2.5 Forward Propagation

sal
Key Formula

For each layer ℓ:


z [ℓ] = W [ℓ] a[ℓ−1] + b[ℓ] a[ℓ] = ϕ(z [ℓ] )
a[0] = x

an
where (the input).

2.6 Linear Separability and XOR


 Concept
lK
A dataset is linearly separable if a single straight line (or hyperplane) can perfectly separate
the two classes.

(a) AND  separable (b) XOR  NOT separable


mo

/! Examiner Trap

A single perceptron can only solve linearly separable problems (Minsky-Papert critique,
1969). XOR requires at least one hidden layer. This is the historical motivation for MLPs.

2.7 Perceptron vs Logistic Regression


An

Aspect Perceptron Logistic Regression

Activation Step (hard threshold) Sigmoid (smooth)

Output type {0, 1} Probability in (0, 1)


Training rule Perceptron rule Gradient descent on log-
loss

Loss function None explicitly Log-loss / BCE (convex)

Convergence Only on separable data Always (convex problem)

Decision boundary Linear hyperplane Linear hyperplane

*** PYQ Favourite

Key insight: both produce a linear decision boundary wT x + b = 0, so they classify the same
way on separable data. The dierences are in training and output type.

PCCAIML602 Anmol Kansal 18


Deep Learning Notes Made by Anmol Kansal

2.8 Delta Rule (Widrow-Ho)


Key Formula

Delta rule: for a linear neuron with target t and output y:

∆wi = η · (t − y) · xi

Derived by minimising J = 12 (t − y)2 .

1 ∂J ∂J
− y)2 , y =
P
Derivation: J = 2 (t wi xi . ∂wi = −(t − y)xi . Update: wi ← wi − η ∂w =
i
wi + η(t − y)xi .

2.9 Fuzzy Relations (Brief)

al
The syllabus mentions cardinality, operations, and properties of fuzzy relations. This is rarely
the main focus but worth knowing:

s
ˆ Fuzzy set: elements have membership µ(x) ∈ [0, 1].

an
ˆ Cardinality: |A| =
P
x µA (x).

ˆ Union: µA∪B (x) = max(µA (x), µB (x)).


ˆ Intersection: µA∩B (x) = min(µA (x), µB (x)).
ˆ Complement: µĀ (x) = 1 − µA (x).
K
ˆ Properties: reexive (µ(x, x) = 1), symmetric (µ(x, y) = µ(y, x)), transitive (max-min composi-
tion).

2.10 Deep Dive: Worked Examples


ol

These are the kinds of problems MAKAUT examiners actually ask. Master them.

2.10.1 Worked Example 1  Linear Neuron Output


nm

*** PYQ Favourite

Problem (PYQ-style): A 4-input neuron has weights w = [1, 2, 3, 4]. Transfer function is
linear with proportionality constant k=2 (i.e., ϕ(z) = 2z ). Inputs are x = [4, 10, 5, 20]. No
bias. Find the output.

Step 1  compute weighted sum:


A

X
wi xi = 1(4) + 2(10) + 3(5) + 4(20) = 4 + 20 + 15 + 80 = 119.
i

Step 2  apply activation:

y = ϕ(119) = 2 · 119 = 238 .

/! Examiner Trap

Linear does NOT mean no activation. Linear with proportionality constant k means
ϕ(z) = kz . Identity is the special case k = 1. If you forget this, you'd write 119 (wrong).

PCCAIML602 Anmol Kansal 19


Deep Learning Notes Made by Anmol Kansal

2.10.2 Worked Example 2  Parameter Counting (3-8-8-3 Network)

*** PYQ Favourite

Problem: A network has 3 input neurons, 2 hidden layers each with 8 neurons, and 3 output
neurons. Find:
(a) total biases, (b) total weights, (c) best loss + output activation.

Architecture sketch: 3 → 8 → 8 → 3. Three layers with parameters.


Step 1  biases per layer = neurons in that layer:

btotal = 8 + 8 + 3 = 19 .

Step 2  weights between layers = (left) × (right):

al
W [1] : 8 × 3 = 24
W [2] : 8 × 8 = 64

s
W [3] : 3 × 8 = 24

an
Total weights = 24 + 64 + 24 = 112 .
Step 3  output activation & loss: 3 outputs ⇒ multi-class classication ⇒ softmax +
categorical cross-entropy (CCE).

!!! MUST MEMORIZE


K
Master formula: total parameters per layer = n[ℓ] (n[ℓ−1] + 1). The +1 counts the bias.
Verify: 8(4) + 8(9) + 3(9) = 32 + 72 + 27 = 131 = 112 + 19. ✓

2.10.3 Worked Example 3  Layer Dimensions & Forward Pass


ol

*** PYQ Favourite

Problem (PYQ-style): A 4-layer network: x → a[1] → a[2] → a[3] → a[4] has layer sizes 3, 5,
3, 1.
nm

(a) Size of W [1] ?


(b) Compute a[2] component-wise.

(c) Why are W 's matrices, not vectors?

(d) Size of W
[2] , W [3] ?

(e) If all hidden activations are linear, what happens?


A

(a) W [1] maps 3 inputs to 3 hidden units: W [1] ∈ R3×3 .


(b) a[2] = ϕ(W [2] a[1] + b[2] ) ∈ R5×1 . Each component:
 
3
[2] [2] [1] [2]
X
ai = ϕ Wij aj + bi  , i = 1, . . . , 5.
j=1

(c) Because each layer has multiple input AND multiple output units, a 2-D structure (matrix)
is needed to encode all pairwise weights. A vector wouldn't suce. Biases are vectors because each
output unit has exactly one bias.
(d) W [2] ∈ R5×3 (3 → 5); W [3] ∈ R3×5 (5 → 3).
(e) If ϕ(z) = z everywhere except output, the entire hidden stack collapses:

a[3] = W [3] W [2] W [1] x + const = We x + be .

PCCAIML602 Anmol Kansal 20


Deep Learning Notes Made by Anmol Kansal

!!! MUST MEMORIZE

Profound insight: a deep linear network is no more powerful than a single linear layer.
Non-linearity is what makes depth meaningful. This is a favourite examiner question.

2.10.4 Worked Example 4  Sigmoid Derivative Derivation

*** PYQ Favourite

Problem: Derive σ ′ (z). At what z is the gradient maximum? What is the max value?

Derivation: starting from σ(z) = (1 + e−z )−1 :

e−z
σ ′ (z) = −(1 + e−z )−2 · (−e−z ) =
(1 + e−z )2

sal
1 e−z
= ·
1 + e−z  1 + e−z 
1 1
= · 1−
1 + e−z 1 + e−z

Find the maximum: let


an
= σ(z) (1 − σ(z))

u = σ(z) ∈ (0, 1). Maximise f (u) = u(1 − u) = u − u2 :

f ′ (u) = 1 − 2u = 0 ⇒ u = 0.5 ⇒ σ(z) = 0.5 ⇒ z = 0.


lK
Max value = 0.5 · 0.5 = 0.25 at z = 0.

Intuition

This is the mathematical origin of vanishing gradients. Through 10 sigmoid layers, gradient is
multiplied by at most (0.25)10 ≈ 9.5 × 10−7 . Eectively no learning reaches early layers.
mo

2.11 Advanced Topics


2.11.1 The Complete Activation Function Family

Beyond sigmoid/tanh/ReLU, modern DL uses many variants. Examiners sometimes ask comparison
An

questions across the full family.

PCCAIML602 Anmol Kansal 21


Deep Learning Notes Made by Anmol Kansal

Activation Formula Key property


Identity f (z) = z Used in regression output
layer
Step / Heaviside 1[z > 0] Original perceptron; non-
dierentiable
Sigmoid 1/(1 + e−z ) Output (0, 1); vanishing grad
ez −e−z
Tanh ez +e−z Zero-centred (−1, 1)
ReLU max(0, z) Sparse, non-saturating (z >
0)
Leaky ReLU z if z > 0, else αz Avoids dead ReLU
PReLU Leaky with learnable α Adapts during training
ELU z if z > 0, else α(ez − 1) Smooth, mean closer to 0

al
SELU Scaled ELU Self-normalising activations
GELU z · Φ(z) (Gauss CDF) Used in BERT, GPT
Swish / SiLU z · σ(z) Found by NAS; smooth
Softplus Smooth ReLU approximation

s
ln(1 + ez )
Maxout maxi (wiT x + bi ) Generalises ReLU/Leaky-

an
ReLU
Softmax ezi / ezj Multi-class output
P
j

2.11.2 Softmax Derivative (Often Asked)

ŷi = ezi / k ezk ,


P
K
For softmax the Jacobian ∂ ŷi /∂zj has two cases:
Diagonal (i = j ):

ezi e zk − e zi · e zi
P
∂ ŷi kP
= zk 2
∂zi ( ke )
= ŷi − ŷi2 = ŷi (1 − ŷi ).
ol

O-diagonal (i ̸= j ):
∂ ŷi 0 − ezi · ezj
= P z 2 = −ŷi ŷj .
∂zj ( k e k)
nm

Compactly: ∂ ŷi /∂zj = ŷi (δij − ŷj ) where δij is Kronecker delta.

2.11.3 Softmax + CCE Gradient Simplies Beautifully


P
For one-hot label y and softmax output ŷ , CCE loss J =− k yk log ŷk .
∂J
Claim: ∂z = ŷi − yi .
i
A

Proof:

∂J X ∂ log ŷk X yk ∂ ŷk


=− yk =−
∂zi ∂zi ŷk ∂zi
k k
X yk
=− ŷk (δki − ŷi )
ŷk
k
X
=− yk (δki − ŷi )
k
X
= −yi + ŷi yk
k
X
= ŷi − yi (since yk = 1).
k

PCCAIML602 Anmol Kansal 22


Deep Learning Notes Made by Anmol Kansal

!!! MUST MEMORIZE

The clean cancellation is why softmax+CCE is the canonical pairing, just as sig-
moid+BCE gives ŷ − y . The complicated activation derivative cancels with the loss derivative.

2.11.4 Weight Initialization Theory

Why bad init kills networks:

ˆ All zeros: every neuron computes the same output, gets the same gradient, stays identical
forever. Symmetry never breaks.

ˆ Too small: activations shrink layer by layer (vanishing forward and backward).

sal
ˆ Too large: activations explode.

The principle: keep variance of activations and gradients approximately constant across layers.
Xavier / Glorot (for tanh / sigmoid):
 
2
W ∼ N 0,
nin + nout

an
He / Kaiming (for ReLU; accounts for half-activation):
 
2
W ∼ N 0,
nin
Why the factor of 2 in He? ReLU zeros out half the activations on average, so to maintain

variance, scale up by 2.
lK
Bias init: usually 0 (or small positive for ReLU to keep neurons active early).
Orthogonal init: good for RNN Whh  avoids vanishing/exploding when iterated over time.

2.11.5 Batch Normalization: Full Forward Pass

For a mini-batch {x1 , . . . , xm } at one layer:


m
1 X
mo

µB = xi (batch mean)
m
i=1
m
2 1 X
σB = (xi − µB )2 (batch variance)
m
i=1
x i − µB
x̂i = q (normalize)
2 +ϵ
σB
An

yi = γ x̂i + β (scale and shift)

q γ, β are learnable, allowing the network to undo normalisation if useful (e.g., identity if γ =
2 + ϵ and β = µ ).
σB B
Test-time behaviour: use running averages of µ, σ 2 computed during training, not the test
batch's statistics.
Benets:

ˆ Reduces internal covariate shift

ˆ Allows higher learning rates

ˆ Acts as regulariser (mini-batch noise)

ˆ Less sensitive to initialization

Layer Norm vs Batch Norm:

ˆ Batch Norm: normalize across batch dim for each feature.

ˆ Layer Norm: normalize across feature dim for each example. Used in Transformers, RNNs.

PCCAIML602 Anmol Kansal 23


Deep Learning Notes Made by Anmol Kansal

Unit 2 Quick Recap


ˆ Neuron: y = ϕ(wT x + b). Order: WIB-A.

ˆ Sigmoid max grad = 0.25 (causes vanishing).

ˆ Tanh range (−1, 1), zero-centred.

ˆ ReLU = max(0, z), produces sparse activations.

ˆ Output activation: linear/sigmoid/softmax depending on task.

ˆ Params per layer: n[ℓ] (n[ℓ−1] + 1).


ˆ Single perceptron ≡ logistic regression's decision boundary.

s al
an
K
ol
nm
A

PCCAIML602 Anmol Kansal 24


Thank You for Reading!
This was a freely shared demo version of the notes.

Want the Full Exam Preparation Bundle?

Complete Detailed Notes + PYQ Solutions + Question Bank + Revision


Material

Designed specially for Semester Exam Preparation

Purchase at a very affordable student-friendly price

Official Store
[Link]

Join My Telegram Group


[Link]

Contact via Telegram Bot


[Link]
YouTube Channel
[Link]

Study Smart. Revise Fast. Score Better.

Made by Anmol Kansal

You might also like