Chapter 2
Chapter 2
Architecture
of the Transformer Model
Siksha ‘O’ Anushandhan, ITER
Center for Data Science
Table of Contents
5 Positional Encoding
1 Introduction: The Transformer Revolution
6 Multi-Head Attention
2 Transformer Architecture Overview 7 Post-layer normalization
3 The Encoder Stack 8 Feedforward network
4 Input Embedding 9 The Decoder stack
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 2 / 33
The Language Revolution
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 3 / 33
The Birth of Transformers
December 2017
Key Achievement
The Transformer trained faster and obtained higher evaluation results than previous
architectures
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 4 / 33
From RNNs to Transformers
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 5 / 33
Modern Transformer Models
GPT-4
OpenAI
PaLM
Google Large Language
Original
Models (LLMs)
Transformer
The beginning
(2017)
LaMBDA of a new era!
Google
ChatGPT
OpenAI
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 6 / 33
Transformer Architecture
The Structure
6-layer encoder stack Output
6-layer decoder stack
Each layer contains
ENCODER
DECODER
sublayers
Output of layer ℓ = Input
of layer ℓ + 1
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 7 / 33
Attention mechanism
in a sequence
sat sat 0.15
Structural Consistency
All N=6 (can be 12, 24, etc) layers identical in
structure
Different learned weights per layer
Each layer explores different word associations
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 9 / 33
Encoder layer structure
Input Embedding
The input embedding sublayer converts the input tokens to vectors of
dimension dmodel = 512 using learned embeddings in the original
Transformer model.
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 10 / 33
Example of embedding
Text : The cat slept on the [Link] was too tired to get up.
Tokenized text : [1996, 4937, 7771, 2006, 1996, 6411, 1012, 2009, 2001, 2205, 5458,
2000, 2131, 2039, 1012]
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 11 / 33
Example of embedding
Text : The cat slept on the [Link] was too tired to get up.
Tokenized text : [1996, 4937, 7771, 2006, 1996, 6411, 1012, 2009, 2001, 2205, 5458,
2000, 2131, 2039, 1012]
There is not enough information in the tokenized text at this point to go further.
Need to apply some embedding method.
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 11 / 33
Word embedding
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 12 / 33
Word embedding
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 12 / 33
Word embedding
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 12 / 33
Word embedding
1 import nltk
2 from nltk . tokenize import sent_tokenize , word_tokenize
3 s = " The black cat sat on the couch and the brown dog slept on the
rug . "
4 data = [[ word . lower () for word in word_tokenize ( sentence ) ] for
sentence in sent_tokenize ( s ) ]
5 print ( data )
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 13 / 33
Skip Gram model
1 # Creating Skip Gram model
2 model = gensim . models . Word 2 Vec ( data , min_count = 1 , vector_size = 5
1 2 , window = 5 , sg = 1 )
Parameter Description
data This is the corpus (training data) for the Word2Vec model. It should be an iterable of
lists of words. Each inner list represents a sentence or a document, and each element in
the inner list is a word (string). In our case, data is [[’the’, ’black’, ’cat’, ...,
’rug’, ’.’]].
min count Ignores all words with total frequency lower than this.
vector size Dimensionality of the word vectors.
window Maximum distance between the current and predicted word within a sentence.
sg {0, 1}, Training algorithm: 1 for skip-gram; otherwis Continuous Bag of
Words(CBOW).
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 14 / 33
Skip Gram model
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 15 / 33
Skip Gram model
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 15 / 33
Positional Encoding
10000
dmodel
pos
PE (pos, 2i + 1) = cos 2i
10000 dmodel
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 16 / 33
Final Input
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 17 / 33
positional encoding
Define function to get positional encoding vector
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 18 / 33
positional encoding
Define function to get positional encoding vector
1 import numpy as np
2 def positional_encoding ( pos , d_model ) :
3 pe = np . zeros (( 1 , d_model ) )
4 for i in range ( 0 , d_model , 2 ) :
5 pe [ 0 ][ i ] = np . sin ( pos / ( 1 0 0 0 0 ** (( 2 * i ) / d_model ) ) )
6 if i + 1 < d_model :
7 pe [ 0 ][ i + 1 ] = np . cos ( pos / ( 1 0 0 0 0 ** (( 2 * i ) /
d_model ) ) )
8 return pe
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 18 / 33
positional encoding
Define function to get positional encoding vector
1 import numpy as np
2 def positional_encoding ( pos , d_model ) :
3 pe = np . zeros (( 1 , d_model ) )
4 for i in range ( 0 , d_model , 2 ) :
5 pe [ 0 ][ i ] = np . sin ( pos / ( 1 0 0 0 0 ** (( 2 * i ) / d_model ) ) )
6 if i + 1 < d_model :
7 pe [ 0 ][ i + 1 ] = np . cos ( pos / ( 1 0 0 0 0 ** (( 2 * i ) /
d_model ) ) )
8 return pe
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 18 / 33
positional encoding
Define function to get positional encoding vector
1 import numpy as np
2 def positional_encoding ( pos , d_model ) :
3 pe = np . zeros (( 1 , d_model ) )
4 for i in range ( 0 , d_model , 2 ) :
5 pe [ 0 ][ i ] = np . sin ( pos / ( 1 0 0 0 0 ** (( 2 * i ) / d_model ) ) )
6 if i + 1 < d_model :
7 pe [ 0 ][ i + 1 ] = np . cos ( pos / ( 1 0 0 0 0 ** (( 2 * i ) /
d_model ) ) )
8 return pe
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 18 / 33
Final Input
Find Cosine-similarity
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 19 / 33
Final Input
Find Cosine-similarity
1 cos_pos_enc_only = cosine_similarity ( pe_a , pe_b )
2 print ( f " Cosine similarity of positional encodings :
{ cos_pos_enc_only } " )
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 19 / 33
Final Input
Find Cosine-similarity
1 cos_pos_enc_only = cosine_similarity ( pe_a , pe_b )
2 print ( f " Cosine similarity of positional encodings :
{ cos_pos_enc_only } " )
Find final input for the model and the cosine similarity
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 19 / 33
Final Input
Find Cosine-similarity
1 cos_pos_enc_only = cosine_similarity ( pe_a , pe_b )
2 print ( f " Cosine similarity of positional encodings :
{ cos_pos_enc_only } " )
Find final input for the model and the cosine similarity
1 pos_embd_a = a . reshape ( 1 , -1 ) + pe_a
2 pos_embd_b = b . reshape ( 1 , -1 ) + pe_b
3
4 cos_sim_combined = cosine_similarity ( pos_embd_a , pos_embd_b )
5 print ( f " Cosine similarity of combined embeddings :
{ cos_sim_combined } " )
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 19 / 33
Multi-Head Attention
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 21 / 33
Core Idea of attention
Attention mechanism
Input: X ∈ Rn×dmodel
Compute:
Q = XW Q W Q , W K , W V are
K = XW K learnable matrices,
and W Q , W K , W V ∈
V = XW V Rdmodel ×dmodel
Q, K , V ∈ Rn×dmodel
Compute attention
T
Attention(Q, K , V ) = softmax QK
√ V
dk
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 22 / 33
Core Idea of attention
Multi-head Attention
Attention mechanism
Input: X ∈ Rn×dmodel
Input: X ∈ Rn×dmodel
Head i: (for i = 1, 2, . . . , m)
headi = Attention(Qi , Ki , Vi )
Q KT
Compute: = softmax √i di Vi
k
Q = XW Q W Q , W K , W V are Qi , Ki , Vi ∈ Rn×dk
K = XW K learnable matrices,
and W Q , W K , W V ∈
V = XW V Rdmodel ×dmodel
Q, K , V ∈ Rn×dmodel Compute attention
Heads = Concat(head1 , . . . , headm )
Compute attention
T
Attention(Q, K , V ) = softmax QK
√ V Final output
dk
output = Heads × W O
W O ∈ Rdmodel ×dmodel
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 22 / 33
Multi-head attention implementation
Define all parameters
1 import numpy as np
2 np . random . seed ( 4 2 )
3 n , d_model , num_heads = 5 , 8 , 2
4 d_k = d_model // num_heads
5 print ( " d_k per head : " , d_k ) # 4
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 23 / 33
Multi-head attention implementation
Define all parameters
1 import numpy as np
2 np . random . seed ( 4 2 )
3 n , d_model , num_heads = 5 , 8 , 2
4 d_k = d_model // num_heads
5 print ( " d_k per head : " , d_k ) # 4
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 23 / 33
Multi-head attention implementation
Define all parameters
1 import numpy as np
2 np . random . seed ( 4 2 )
3 n , d_model , num_heads = 5 , 8 , 2
4 d_k = d_model // num_heads
5 print ( " d_k per head : " , d_k ) # 4
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 23 / 33
Multi-head attention implementation
Define all parameters
1 import numpy as np
2 np . random . seed ( 4 2 )
3 n , d_model , num_heads = 5 , 8 , 2
4 d_k = d_model // num_heads
5 print ( " d_k per head : " , d_k ) # 4
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 23 / 33
Multi-head attention implementation
Compute Q, K , V
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 24 / 33
Multi-head attention implementation
Compute Q, K , V
1 Q = X @ W_Q
2 K = X @ W_K
3 V = X @ W_V
4 print ( " Q shape : " , Q . shape ) # (5 , 8)
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 24 / 33
Multi-head attention implementation
Compute Q, K , V
1 Q = X @ W_Q
2 K = X @ W_K
3 V = X @ W_V
4 print ( " Q shape : " , Q . shape ) # (5 , 8)
1 Q = Q . transpose ( 1 , 0 , 2 )
2 K = K . transpose ( 1 , 0 , 2 )
3 V = V . transpose ( 1 , 0 , 2 )
4 print ( " Q transposed shape : " ,
Q . shape ) # (2 , 5 , 4)
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 24 / 33
Multi-head attention implementation
Compute Q, K , V
1 Q = X @ W_Q
2 K = X @ W_K
3 V = X @ W_V
4 print ( " Q shape : " , Q . shape ) # (5 , 8)
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 25 / 33
Multi-head attention implementation
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 25 / 33
Multi-head attention implementation
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 26 / 33
Multi-head attention implementation
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 26 / 33
Post-layer normalization (Add and Norm)
Post-layer normalization
output = LayerNormalization(X + MHA(X ))
LayerNormalization
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 27 / 33
Post-layer normalization implementation
Compute X + final output
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 28 / 33
Post-layer normalization implementation
Compute X + final output
1 residual = X + final_output
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 28 / 33
Post-layer normalization implementation
Compute X + final output
1 residual = X + final_output
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 28 / 33
Post-layer normalization implementation
Compute X + final output
1 residual = X + final_output
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 28 / 33
Post-layer normalization implementation
Compute X + final output
1 residual = X + final_output
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 28 / 33
Post-layer normalization implementation
Compute X + final output
1 residual = X + final_output
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 28 / 33
Post-layer normalization implementation
Compute X + final output
1 residual = X + final_output
Exercise:
1 Define ReLU function.
2 Define Softmax function.
3 Define LayerNormalization function.
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 29 / 33
Feedforward network
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 30 / 33
Feedforward network implementation
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 31 / 33
Feedforward network implementation
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 31 / 33
Feedforward network implementation
Compute FNN(x)
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 31 / 33
Feedforward network implementation
Compute FNN(x)
1 fnn_hidden = np . maximum ( 0 , output @ W 1 + b 1 )
2 print ( fnn_hidden . shape )
3 fnn_output = fnn_hidden @ W 2 + b 2
4 print ( fnn_output . shape )
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 31 / 33
Feedforward network implementation
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 32 / 33
Feedforward network implementation
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 32 / 33
The decoder stack
Masked attention
When predicting token at position t, the model must NOT
see tokens at positions > t.
QK T + M
Attention(Q, K , V ) = softmax √ V
dk
where (
0 if j ≤ i
Mij =
−∞ if j > i
(ITER, Center for Data Science) Natural Language Processing with Transformer Dr. Debashis Bhowmik 33 / 33