Lecture 6: Recurrent Neural Networks
Pranabendu Misra
Chennai Mathematical Institute
S.P. Jain Institute of Management & Research (SPJIMR)
13 June 2024
Dealing with sequences
Conventional neural networks map single
inputs to single outputs
Each input/output may be a vector of values
These are Feed Forward Networks.
S.P. Jain Institute of Management & Research (SPJIM
Pranabendu Misra (CMI) Lecture 6: Recurrent Neural Networks 2 / 17
Dealing with sequences
Conventional neural networks map single
inputs to single outputs
Each input/output may be a vector of values
These are Feed Forward Networks.
Some classification tasks require mapping a
sequence of inputs to an output
Identifying a music of video clip
S.P. Jain Institute of Management & Research (SPJIM
Pranabendu Misra (CMI) Lecture 6: Recurrent Neural Networks 2 / 17
Dealing with sequences
Conventional neural networks map single
inputs to single outputs
Each input/output may be a vector of values
These are Feed Forward Networks.
Some classification tasks require mapping a
sequence of inputs to an output
Identifying a music of video clip
Others require mapping a single input to a
sequence of outputs
Generating a caption for an image
S.P. Jain Institute of Management & Research (SPJIM
Pranabendu Misra (CMI) Lecture 6: Recurrent Neural Networks 2 / 17
Dealing with sequences
Mapping sequences to sequences
Language translation — read an entire input
sentence, then generate output
S.P. Jain Institute of Management & Research (SPJIM
Pranabendu Misra (CMI) Lecture 6: Recurrent Neural Networks 3 / 17
Dealing with sequences
Mapping sequences to sequences
Language translation — read an entire input
sentence, then generate output
Mapping sequences to sequences on the fly
Predict the next word in a sentence
S.P. Jain Institute of Management & Research (SPJIM
Pranabendu Misra (CMI) Lecture 6: Recurrent Neural Networks 3 / 17
Dealing with sequences
Mapping sequences to sequences
Language translation — read an entire input
sentence, then generate output
Mapping sequences to sequences on the fly
Predict the next word in a sentence
Context is important
The handwritten word is clearly defence
The n in isolation is illegible
S.P. Jain Institute of Management & Research (SPJIM
Pranabendu Misra (CMI) Lecture 6: Recurrent Neural Networks 3 / 17
Idea: Lets also remember the past!
S.P. Jain Institute of Management & Research (SPJIM
Pranabendu Misra (CMI) Lecture 6: Recurrent Neural Networks 4 / 17
Incorporating memory
Input sequence x (1) , x (2) , . . . , x (t)
Output sequence ŷ (1) , ŷ (2) , . . . , ŷ (t)
Allow ŷ (t) to also depend on previous inputs
x (1) , x (2) , . . . , x (t−1)
Hidden state : h(t)
h(t) depends on current input and previous
state
h(t) = f (W hx x (t) + W hh h(t−1) + bh )
Output is a function of the current state
y (t) = g (W yh h(t) + by )
S.P. Jain Institute of Management & Research (SPJIM
Pranabendu Misra (CMI) Lecture 6: Recurrent Neural Networks 5 / 17
Time unrolling
S.P. Jain Institute of Management & Research (SPJIM
Pranabendu Misra (CMI) Lecture 6: Recurrent Neural Networks 6 / 17
Time unrolling
S.P. Jain Institute of Management & Research (SPJIM
Pranabendu Misra (CMI) Lecture 6: Recurrent Neural Networks 6 / 17
Time unrolling
S.P. Jain Institute of Management & Research (SPJIM
Pranabendu Misra (CMI) Lecture 6: Recurrent Neural Networks 7 / 17
Time unrolling
S.P. Jain Institute of Management & Research (SPJIM
Pranabendu Misra (CMI) Lecture 6: Recurrent Neural Networks 7 / 17
Time Unrolling and Back Propagation Through Time
Time Unrolling makes it a (larger) Feed-Forward Network
but all copies share the parameters, so number of parameters doesn’t increase.
So we can do back-propagation to update the weights
S.P. Jain Institute of Management & Research (SPJIM
Pranabendu Misra (CMI) Lecture 6: Recurrent Neural Networks 8 / 17
Vanishing and Exploding Gradients gradients
Unfortunately, we end-up with a very deep network
Only the most recent parts of the input sequence are remembered; earlier parts are
forgotten
Back-Propagation suffers from vanishing or exploding gradients when unrolled over many
time steps.
S.P. Jain Institute of Management & Research (SPJIM
Pranabendu Misra (CMI) Lecture 6: Recurrent Neural Networks 9 / 17
Vanishing and Exploding Gradients gradients
Also, can’t unroll infinitely far into the future.
Truncated BPTT is what is done in practice, and also addresses these issues to an extent.
Unfortunately, long-term context is lost.
S.P. Jain Institute of Management & Research (SPJIM
Pranabendu Misra (CMI) Lecture 6: Recurrent Neural Networks 10 / 17
Problem: We are unable to remember important parts of information far
into the past.
S.P. Jain Institute of Management & Research (SPJIM
Pranabendu Misra (CMI) Lecture 6: Recurrent Neural Networks 11 / 17
Problem: We are unable to remember important parts of information far
into the past.
Why this happens:
Every piece xi of the input sequence is treated in the same manner by the RNN.
So important and non-important pieces both modify the hidden state, causing important
pieces of the input sequences further in the past to be forgotten.
S.P. Jain Institute of Management & Research (SPJIM
Pranabendu Misra (CMI) Lecture 6: Recurrent Neural Networks 11 / 17
Problem: We are unable to remember important parts of information far
into the past.
Idea: A mechanism that learns to distinguish between important and
not-important parts, and remembers the important parts.
S.P. Jain Institute of Management & Research (SPJIM
Pranabendu Misra (CMI) Lecture 6: Recurrent Neural Networks 12 / 17
Problem: We are unable to remember important parts of information far
into the past.
Idea: A mechanism that learns to distinguish between important and
not-important parts, and remembers the important parts.
Don’t update the hidden state all the time, but only when necessary.
Use gates to control information flow
A gate is typically a neuron or a simple neural network with sigmoid activation, that takes
the current input x (t) and the previous state h(t) and outputs a vector with values in [0, 1].
0 means gate is closed; 1 means gate is open
We thus arrive at Gated RNNs. Intuitively, gates learn to distinguish important and
non-important information. They only let important information update the internal-state.
S.P. Jain Institute of Management & Research (SPJIM
Pranabendu Misra (CMI) Lecture 6: Recurrent Neural Networks 12 / 17
LSTM
Long Short Term Memory (LSTM)
are a popular variant of gated RNNs.
They use multiple gates to decide
(t)
how the internal state / memory sc
is updated.
(t)
Here, xc is the input at time t,
which is supplied to all the gates
along with the previous hidden state
(t−1)
hc
Π represents point-wise product of
two vectors.
S.P. Jain Institute of Management & Research (SPJIM
Pranabendu Misra (CMI) Lecture 6: Recurrent Neural Networks 13 / 17
LSTM
(t)
fc is the Forget Gate. Decides what
(t−1) (t)
bits of sc is retained in sc and
what is forgotten.
(t)
ic is the Input Gate. Decides what
(t)
bits of the input xc are added to
(t)
sc .
g (t)c is the Input Node which
encodes the input x (t) . It’s output is
(t)
what is actually added to sc instead
of x (t) .
S.P. Jain Institute of Management & Research (SPJIM
Pranabendu Misra (CMI) Lecture 6: Recurrent Neural Networks 14 / 17
LSTM
We update the internal state as
(t) (t) (t) (t−1) (t)
sc = Π(gc , ic ) + Π(sc , fc )
(t)
The Output Gate oc decides what
bits of the internal state s (t) , is
output to the hidden state hc (t)
(t) (t) (t)
hc = Π(sc , oc )
S.P. Jain Institute of Management & Research (SPJIM
Pranabendu Misra (CMI) Lecture 6: Recurrent Neural Networks 15 / 17
LSTM unrolled in time
S.P. Jain Institute of Management & Research (SPJIM
Pranabendu Misra (CMI) Lecture 6: Recurrent Neural Networks 16 / 17
Bidirectional RNN
Useful when we need context
from both past and “future” e.g.
translation, handwriting
recognition etc.
The whole input must be
available; not online.
Training via BPTT
S.P. Jain Institute of Management & Research (SPJIM
Pranabendu Misra (CMI) Lecture 6: Recurrent Neural Networks 17 / 17