Basic
Basic
Fall 2025
1
Acknowledgements
Many of the slides in this lecture are adapted from the sources below. Copyrights belong to the original authors.
Source: [Link]
For details, refer to: [Link]
Word2vec Overview
Source: [Link]
For details, refer to: [Link]
Word2vec Overview
Source: [Link]
For details, refer to: [Link]
Word2vec: Objective Function
Source: [Link]
For details, refer to: [Link]
Word2vec: Objective Function
Source: [Link]
For details, refer to: [Link]
Word2vec: Prediction Function = Prob[o|c]
Source: [Link]
For details, refer to: [Link]
Word2vec: To Train the Model
Source: [Link]
For details, refer to: [Link]
The problematic P[o|c]
exp 𝑢!" 𝑣𝑐
𝑃 𝑜 𝑐) = &
∑#$% exp(𝑢#" 𝑣𝑐 )
𝑈"
𝑉!
You first observe (actually sample) the following sentence from the training
corpus:
“word2vec Explained…”
Goldberg & Levy, arXiv
2014
Skip-Grams with Negative Sampling (SGNS)
Marco saw a furry little wampimuk hiding in the tree.
Take Log and the optimization problem becomes “similar” to the training of a binary logistic-regression classifier
( [Link] ):
Summary: How to learn Word2vec Embeddings via SGNS
21
NLP Tasks as Sequence to Sequence Modeling
22
NLP Tasks as Sequence to Sequence Modeling
► Example Scenarios
► Text à Text (e.g. Q/A, translation, text summarization)
► Image à Text (e.g. image captioning)
Output
𝑦! 𝑦" 𝑦# 𝑦$ sequence
magic?
Input
sequence 𝑥! 𝑥" 𝑥# 𝑥$
NLP Tasks as Sequence to Sequence Modeling
24
► Example Scenarios
► Text à Text (e.g. Q/A, translation, text summarization)
► Image à Text (e.g. image captioning)
state
ENCODER DECODER
context
vector
Input
sequence 𝑥! 𝑥" 𝑥# 𝑥$
Recurrent Neural Networks (RNN) – a Seq2Seq NN model
A simple RNN shown unrolled in time. Network layers are recalculated for
each time step, while weights U, V and W are shared across all time steps.
Better Capturing of Long-Range Dependence
using LSTM for Seq2Seq Modeling
► Encoder (LSTM) and decoder (LSTM)
► Fixed-length context vector
Initial state
ℎ! 𝑓 ℎ" 𝑓 ℎ# 𝑓 ℎ$ 𝑠% 𝑠! 𝑔 𝑠" 𝑔 𝑠# 𝑔 𝑠$
𝑥! 𝑥" 𝑥# 𝑥$ c 𝑦% 𝑦! 𝑦" 𝑦#
Context vector
ENCODER DECODER
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in Proceedings of the 27th International Conference on Neural Information
Processing Systems (NIPS), 2014, pp. 3104–3112.
LSTM still suffers from Information Bottleneck
► Encoder (LSTM) and decoder (LSTM)
► Fixed-length context vector (bottleneck)
Initial state
ℎ! 𝑓 ℎ" 𝑓 ℎ# 𝑓 ℎ$ 𝑠% 𝑠! 𝑔 𝑠" 𝑔 𝑠# 𝑔 𝑠$
BOTTLENECK
𝑥! 𝑥" 𝑥# 𝑥$ c 𝑦% 𝑦! 𝑦" 𝑦#
Context vector
ENCODER DECODER
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in Proceedings of the 27th International Conference on Neural Information
Processing Systems (NIPS), 2014, pp. 3104–3112.
1990s Prehistoric Era
Rule-based methods, parsing, RNNs, LSTMs
Attention Timeline
2014 Simple attention mechanisms
► Idea! Use a different context vector for each timestep in the decoder
𝑠. = 𝑔(𝑦./0 , 𝑠./0 , 𝒄𝒕 )
► Craft the context vector so that it “looks at” different parts of the input
sequence for each decoder timestep
D. Bahdanau, K. Cho, and Y. Bengio, “Neural Machine Translation by Jointly Learning to Align and Translate,” in 3rd International Conference on Learning Representations
(ICLR), 2015.
Sequence to Sequence with RNNs + Attention
38
Compute context vector
Find 𝑠!: 𝑡 = 1
× × × × 𝑐' = 4 𝛼',( ℎ(
+ (
Attention weights
𝛼!,! 𝛼!," 𝛼!,# 𝛼!,$ (normalize alignment
scores) 𝑠$ = 𝑔(𝑦$%& , 𝑠$%& , 𝒄𝒕 )
softmax
𝑦!
Alignment scores
𝑒!,! 𝑒!," 𝑒!,# 𝑒!,$ 𝛼',( represents the probability that
𝑒',( = 𝑓)'' (𝑠'*! , ℎ( )
the target word 𝑦' is aligned to, or
translated from, a source word 𝑥(
D. Bahdanau, K. Cho, and Y. Bengio, “Neural Machine Translation by Jointly Learning to Align and Translate,” in 3rd International Conference on Learning Representations
(ICLR), 2015.
Sequence to Sequence with RNNs + Attention
39
Compute context vector
Find 𝑠": 𝑡 = 2
× × × × 𝑐' = 4 𝛼',( ℎ(
+ (
Attention weights
𝛼",! 𝛼"," 𝛼",# 𝛼",$ (normalize alignment
scores) 𝑠$ = 𝑔(𝑦$%& , 𝑠$%& , 𝒄𝒕 )
softmax
𝑦! 𝑦"
Alignment scores
𝑒",! 𝑒"," 𝑒",# 𝑒",$
𝑒',( = 𝑓)'' (𝑠'*! , ℎ( )
ℎ! 𝑓 ℎ" 𝑓 ℎ# 𝑓 ℎ$ 𝑠% 𝑠! 𝑔 𝑠"
Initial state
𝑥! 𝑥" 𝑥# 𝑥$ 𝑐! 𝑦% 𝑐" 𝑦!
Context
ENCODER vector DECODER
D. Bahdanau, K. Cho, and Y. Bengio, “Neural Machine Translation by Jointly Learning to Align and Translate,” in 3rd International Conference on Learning Representations
(ICLR), 2015.
Sequence to Sequence with RNNs + Attention
40
Compute context vector
Find 𝑠' : 𝑡 = 𝑡
× × × × 𝑐' = 4 𝛼',( ℎ(
+ (
Attention weights
𝛼',! 𝛼'," 𝛼',# 𝛼',$ (normalized alignment
scores) 𝑠$ = 𝑔(𝑦$%& , 𝑠$%& , 𝒄𝒕 )
softmax
𝑦! 𝑦" 𝑦# 𝑦$
Alignment scores 𝑐'
𝑒',! 𝑒'," 𝑒',# 𝑒',$
𝑒',( = 𝑓)'' (𝑠'*! , ℎ( )
𝑠'*!
ℎ! 𝑓 ℎ" 𝑓 ℎ# 𝑓 ℎ$ 𝑠% 𝑠! 𝑔 𝑠" 𝑔 𝑠# 𝑔 𝑠$
Initial state
D. Bahdanau, K. Cho, and Y. Bengio, “Neural Machine Translation by Jointly Learning to Align and Translate,” in 3rd International Conference on Learning Representations
(ICLR), 2015.
Sequence to Sequence with RNNs + Attention
41
All steps are differentiable,
Compute context vector so we can backpropagate
Find 𝑠' → 𝑡 = 𝑡
through everything
× × × × 𝑐' = 4 𝛼',( ℎ(
+ (
Attention weights
𝛼',! 𝛼'," 𝛼',# 𝛼',$ (normalized alignment
scores) 𝑠$ = 𝑔(𝑦$%& , 𝑠$%& , 𝒄𝒕 )
softmax
𝑦! 𝑦" 𝑦# 𝑦$
Alignment scores 𝑐'
𝑒',! 𝑒'," 𝑒',# 𝑒',$
𝑒',( = 𝑓)'' (𝑠'*! , ℎ( )
𝑠'*!
ℎ! 𝑓 ℎ" 𝑓 ℎ# 𝑓 ℎ$ 𝑠% 𝑠! 𝑔 𝑠" 𝑔 𝑠# 𝑔 𝑠$
Initial state
Encoder is bi-directional: allows for the annotation of each
word to summarize both preceding and following words.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural Machine Translation by Jointly Learning to Align and Translate,” in 3rd International Conference on Learning Representations
(ICLR), 2015.
Sequence to Sequence with RNNs + Attention
Application: translation
09-10
2025-
Each pixel shows the weight 𝛼",$ of
the annotation of the 𝑖-th source
word for the 𝑡-th target word.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural Machine Translation by Jointly Learning to Align and Translate,” in 3rd International Conference on Learning Representations
(ICLR), 2015.
Sequence to Sequence with RNNs + Attention 43
Application:
09-10
2025-
UWaterloo
Slides created for CS886 at
text translation
RNN:
RNNenc
RNN + attention:
RNNsearch
D. Bahdanau, K. Cho, and Y. Bengio, “Neural Machine Translation by Jointly Learning to Align and Translate,” in 3rd International Conference on Learning Representations
(ICLR), 2015.
Motivating Transformers by
Understanding the Limitations of Recurrent Models
Caveating this discussion: while the challenges we’ll discuss originally motivated
Transformers, many have continued to make progress on RNNs over the years (S4, Mamba,
Linear Attention, GLA, Based, etc.)
Challenge 1: Long Interaction Distances
• Performance degrades
as the distance between
words increases due to
memory constraints:
Diluted impact of earlier The counselor situation
elements on output as
sequence progresses
“The” has gone through t layers
before getting to “situation”
While RNNs + “Attention” has made some progress, the
coupling of the “sequential” structure of RNN with Attention
still creates difficulties
53
Slides for video from:
Prof. Jia-Bin Huang
University of Maryland, College Park
[Link]
🤔 🤔 🤔
What? How? Why?
En ZH
How are you? 你好嗎?
Encoders Decoders
Encoders Decoders
Encoders Decoders
Encoders Decoders
Encoders Decoders
Encoders Decoders
Encoders Decoders
Encoders Decoders
Encoders Decoders
Encoders Decoders
𝑊!
2.7 0 0
𝑑 1.2 = 𝑑 0 0
⋮ 0 0
0.2 0 ⋮
0
Embedding Space Embedded Embedding Matrix
token
dog
0
# tokens
0.5 0 1
𝑊!
2.7 0 0
𝑑 1.2 = 𝑑 0 ⋯ 0
⋮ 0 0
0.2 0 ⋮
0
Embedding Space Embedded Embedding Matrix
token
Apple
dog
0
# tokens
0.5 0 1
𝑊!
2.7 0 0
𝑑 1.2 = 𝑑 0 ⋯ 0
⋮ 0 0
0.2 0 ⋮
0
Embedding Space Embedded Embedding Matrix
token
I bought an apple and an orange.
Apple
𝑊!
2.7 0 0
𝑑 1.2 = 𝑑 0 ⋯ 0
⋮ 0 0
0.2 0 ⋮
0
Embedding Space Embedded Embedding Matrix
token
Encoder
Embedded
Tokens
Token
𝑊< 𝑊< 𝑊< 𝑊< 𝑊< 𝑊< 𝑊<
Embedding
Tokens I bought an apple and an orange
Feed Feed Feed Feed Feed Feed Feed
Forward Forward Forward Forward Forward Forward Forward
Embedded
Tokens
Token
Embedding 𝑊< 𝑊< 𝑊< 𝑊< 𝑊< 𝑊< 𝑊<
Embedded
Tokens
Token
Embedding 𝑊< 𝑊< 𝑊< 𝑊< 𝑊< 𝑊< 𝑊<
Embedded
Tokens
Token
Embedding 𝑊< 𝑊< 𝑊< 𝑊< 𝑊< 𝑊< 𝑊<
Self-Attention
Embedded
Tokens
Token
Embedding 𝑊< 𝑊< 𝑊< 𝑊< 𝑊< 𝑊< 𝑊<
apple
Embedding Space
Embedded
Tokens
apple
Embedding Space
Embedded
Tokens
apple
Embedding Space
Embedded
Tokens
apple
Embedding Space
Embedded
Tokens
Embedding Space
Embedded
Tokens
Embedding Space
Embedded
Tokens
Embedded
Tokens 𝒙0∈ 𝑅8 𝒙4 ∈ 𝑅 8 𝒙5 ∈ 𝑅 8 𝒙6 ∈ 𝑅 8 𝒙7 ∈ 𝑅 8
Tokens I bought an apple watch
* * * * *
𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%
0.082 0.0495 0.0199 0.6034 0.2452
Sum normalization
Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
* * * * *
𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%
Softmax
𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%
= 𝒙&# 𝒙) = 𝒙&# 𝒙( = 𝒙&# 𝒙' = 𝒙&# 𝒙# = 𝒙&# 𝒙%
Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
* * * * *
𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%
Softmax
𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%
= 𝒙&# 𝒙) = 𝒙&# 𝒙( = 𝒙&# 𝒙' = 𝒙&# 𝒙# = 𝒙&# 𝒙%
Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
' 𝒙 ' 𝒙 ' 𝒙 ' 𝒙 ' 𝒙
Updated 𝒙'$ = 𝛼$,! ! + 𝛼$," " + 𝛼$,# # + 𝛼$,$ $ + 𝛼$,% %
feature
* * * * *
watch 𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%
Softmax
I apple 𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%
= 𝒙&# 𝒙) = 𝒙&# 𝒙( = 𝒙&# 𝒙' = 𝒙&# 𝒙# = 𝒙&# 𝒙%
an bought
Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
' 𝒙 ' 𝒙 ' 𝒙 ' 𝒙 ' 𝒙
Updated 𝒙'$ = 𝛼$,! ! + 𝛼$," " + 𝛼$,# # + 𝛼$,$ $ + 𝛼$,% %
feature
* * * * *
watch 𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%
apple
Softmax
I 𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%
= 𝒙&# 𝒙) = 𝒙&# 𝒙( = 𝒙&# 𝒙' = 𝒙&# 𝒙# = 𝒙&# 𝒙%
an bought
Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
' 𝒙 ' 𝒙 ' 𝒙 ' 𝒙 ' 𝒙
Updated 𝒙'$ = 𝛼$,! ! + 𝛼$," " + 𝛼$,# # + 𝛼$,$ $ + 𝛼$,% %
feature
* * * * *
𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%
delicious apple
Softmax
𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%
= 𝒙&# 𝒙) = 𝒙&# 𝒙( = 𝒙&# 𝒙' = 𝒙&# 𝒙# = 𝒙&# 𝒙%
Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
* * * * *
𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%
Softmax
= 𝑘𝒙)&&
𝛼#,) = #𝑞𝒙#) 𝛼#,(= 𝑘(& 𝑞# 𝛼#,' = 𝑘'& 𝑞# 𝛼#,#= 𝑘#& 𝑞# 𝛼#,% = 𝑘%& 𝑞#
𝑊 B ∈ 𝑅 %A ×%
𝑊 C ∈ 𝑅 %A ×%
𝑘& 𝑘@ 𝑘? 𝑞= 𝑘= 𝑘>
𝑊, 𝑊, 𝑊, 𝑊+ 𝑊, 𝑊,
Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
* * * * *
𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%
Softmax
= 𝑘𝒙)&&
𝛼#,) = #𝑞𝒙#) 𝛼#,( = 𝑘(& 𝑞# 𝛼#,' = 𝑘'& 𝑞# 𝛼#,# = 𝑘#& 𝑞# 𝛼#,% = 𝑘%& 𝑞#
𝑊 B ∈ 𝑅 %A ×% = 𝒙&# 𝒙( = 𝒙&# 𝒙' = 𝒙&# 𝒙# = 𝒙&# 𝒙%
𝑊 C ∈ 𝑅 %A ×%
𝑘& 𝑣& 𝑘@ 𝑣@ 𝑘? 𝑣? 𝑞= 𝑘= 𝑣= 𝑘> 𝑣>
𝑊 E
∈ 𝑅 %D×%
𝑊, 𝑊- 𝑊, 𝑊- 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊, 𝑊-
Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Updated 𝒙(' = *
𝛼#,) *
𝑣& + 𝛼#,( *
𝑣@ + 𝛼#,' *
𝑣? + 𝛼#,# *
𝑣= + 𝛼#,% 𝑣>
feature
* * * * *
𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%
Softmax
= 𝑘𝒙)&&
𝛼#,) = #𝑞𝒙#) 𝛼#,( = 𝑘(& 𝑞# 𝛼#,' = 𝑘'& 𝑞# 𝛼#,# = 𝑘#& 𝑞# 𝛼#,% = 𝑘%& 𝑞#
𝑊 B ∈ 𝑅 %A ×% = 𝒙&# 𝒙( = 𝒙&# 𝒙' = 𝒙&# 𝒙# = 𝒙&# 𝒙%
𝑊 C ∈ 𝑅 %A ×%
𝑘& 𝑣& 𝑘@ 𝑣@ 𝑘? 𝑣? 𝑞= 𝑘= 𝑣= 𝑘> 𝑣>
𝑊 E
∈ 𝑅 %D×%
𝑊, 𝑊- 𝑊, 𝑊- 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊, 𝑊-
Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Updated 𝒙(' = 𝑊" *
𝛼#,) *
𝑣& + 𝛼#,( *
𝑣@ + 𝛼#,' * *
0𝑣? + 𝛼#,#𝑣= + 𝛼#,% 𝑣>
feature
= *
= 0𝛼#,) 𝑊" 𝑊- 𝒙.
.
* * * * *
𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%
𝑊 F ∈ 𝑅 %×%D
Softmax
𝛼#,) = 𝑘)& 𝑞# 𝛼#,( = 𝑘(& 𝑞# 𝛼#,' = 𝑘'& 𝑞# 𝛼#,# = 𝑘#& 𝑞# 𝛼#,% = 𝑘%& 𝑞#
𝑊 B ∈ 𝑅 %A ×% = 𝒙&# 𝒙) = 𝒙&# 𝒙( = 𝒙&# 𝒙' = 𝒙&# 𝒙# = 𝒙&# 𝒙%
𝑊 C ∈ 𝑅 %A ×%
𝑘& 𝑣& 𝑘@ 𝑣@ 𝑘? 𝑣? 𝑞= 𝑘= 𝑣= 𝑘> 𝑣>
𝑊 E
∈ 𝑅 %D×%
𝑊, 𝑊- 𝑊, 𝑊- 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊, 𝑊-
Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Updated 𝒙(' = 𝑊" *
𝛼#,) *
𝑣& + 𝛼#,( *
𝑣@ + 𝛼#,' * *
0𝑣? + 𝛼#,#𝑣= + 𝛼#,% 𝑣>
feature
= *
= 0𝛼#,) 𝑊" 𝑊- 𝒙.
.
* * * * *
𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%
𝑊 F ∈ 𝑅 %×%D
Softmax
𝛼#,) = 𝑘)& 𝑞# 𝛼#,( = 𝑘(& 𝑞# 𝛼#,' = 𝑘'& 𝑞# 𝛼#,# = 𝑘#& 𝑞# 𝛼#,% = 𝑘%& 𝑞#
𝑊 B ∈ 𝑅 %A ×% = 𝒙&# 𝒙) = 𝒙&# 𝒙( = 𝒙&# 𝒙' = 𝒙&# 𝒙# = 𝒙&# 𝒙%
𝑊 C ∈ 𝑅 %A ×%
𝑘& 𝑣& 𝑘@ 𝑣@ 𝑘? 𝑣? 𝑞= 𝑘= 𝑣= 𝑘> 𝑣>
𝑊 E
∈ 𝑅 %D×%
𝑊, 𝑊- 𝑊, 𝑊- 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊, 𝑊-
Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
* * * * *
𝛼(,) 𝛼(,( 𝛼(,' 𝛼(,# 𝛼(,%
𝑊 F ∈ 𝑅 %×%D
Softmax
𝛼#,) = 𝑘)& 𝑞# 𝛼#,( = 𝑘(& 𝑞# 𝛼#,' = 𝑘'& 𝑞# 𝛼#,# = 𝑘#& 𝑞# 𝛼#,% = 𝑘%& 𝑞#
𝑊 B ∈ 𝑅 %A ×% = 𝒙&# 𝒙) = 𝒙&# 𝒙( = 𝒙&# 𝒙' = 𝒙&# 𝒙# = 𝒙&# 𝒙%
𝑊 C ∈ 𝑅 %A ×%
𝑘& 𝑣& 𝑞@ 𝑘@ 𝑣@ 𝑘? 𝑣? 𝑘= 𝑣= 𝑘> 𝑣>
𝑊 E
∈ 𝑅 %D×%
𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊, 𝑊- 𝑊, 𝑊- 𝑊, 𝑊-
Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
* * * * *
𝛼(,) 𝛼(,( 𝛼(,' 𝛼(,# 𝛼(,%
𝑊 F ∈ 𝑅 %×%D
Softmax
𝛼#,) = 𝑘)& 𝑞# 𝛼#,( = 𝑘(& 𝑞# 𝛼#,' = 𝑘'& 𝑞# 𝛼#,# = 𝑘#& 𝑞# 𝛼#,% = 𝑘%& 𝑞#
𝑊 B ∈ 𝑅 %A ×% = 𝒙&# 𝒙) = 𝒙&# 𝒙( = 𝒙&# 𝒙' = 𝒙&# 𝒙# = 𝒙&# 𝒙%
𝑊 C ∈ 𝑅 %A ×%
𝑘& 𝑣& 𝑞@ 𝑘@ 𝑣@ 𝑘? 𝑣? 𝑘= 𝑣= 𝑘> 𝑣>
𝑊 E
∈ 𝑅 %D×%
𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊, 𝑊- 𝑊, 𝑊- 𝑊, 𝑊-
Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Updated 𝒙() = 𝑊 F *
𝛼(,) *
𝑣& + 𝛼(,( *
𝑣@ + 𝛼(,' * *
0𝑣? + 𝛼(,#𝑣= + 𝛼(,% 𝑣>
feature
* * * * *
𝛼(,) 𝛼(,( 𝛼(,' 𝛼(,# 𝛼(,%
𝑊 F ∈ 𝑅 %×%D
Softmax
𝛼#,) = 𝑘)& 𝑞# 𝛼#,( = 𝑘(& 𝑞# 𝛼#,' = 𝑘'& 𝑞# 𝛼#,# = 𝑘#& 𝑞# 𝛼#,% = 𝑘%& 𝑞#
𝑊 B ∈ 𝑅 %A ×% = 𝒙&# 𝒙) = 𝒙&# 𝒙( = 𝒙&# 𝒙' = 𝒙&# 𝒙# = 𝒙&# 𝒙%
𝑊 C ∈ 𝑅 %A ×%
𝑘& 𝑣& 𝑞@ 𝑘@ 𝑣@ 𝑘? 𝑣? 𝑘= 𝑣= 𝑘> 𝑣>
𝑊 E
∈ 𝑅 %D×%
𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊, 𝑊- 𝑊, 𝑊- 𝑊, 𝑊-
Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Updated 𝒙() = 𝑊 F *
𝛼(,) *
𝑣& + 𝛼(,( *
𝑣@ + 𝛼(,' * *
0𝑣? + 𝛼(,#𝑣= + 𝛼(,% 𝑣>
feature
* * * * *
𝛼(,) 𝛼(,( 𝛼(,' 𝛼(,# 𝛼(,%
𝑊 F ∈ 𝑅 %×%D
Softmax
𝛼#,) = 𝑘)& 𝑞# 𝛼#,( = 𝑘(& 𝑞# 𝛼#,' = 𝑘'& 𝑞# 𝛼#,# = 𝑘#& 𝑞# 𝛼#,% = 𝑘%& 𝑞#
𝑊 B ∈ 𝑅 %A ×% = 𝒙&# 𝒙) = 𝒙&# 𝒙( = 𝒙&# 𝒙' = 𝒙&# 𝒙# = 𝒙&# 𝒙%
𝑊 C ∈ 𝑅 %A ×%
𝑘& 𝑣& 𝑞@ 𝑘@ 𝑣@ 𝑘? 𝑣? 𝑘= 𝑣= 𝑘> 𝑣>
𝑊 E
∈ 𝑅 %D×%
𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊, 𝑊- 𝑊, 𝑊- 𝑊, 𝑊-
Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
𝛼),) = 𝑘&G 𝑞&
𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊-
Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
1
𝛼),) 𝛼(,) 𝛼',) 𝛼#,) 𝛼%,) 𝑘1&G
𝛼),( 𝛼(,( 𝛼',( 𝛼#,( 𝛼%,( 𝑘1@G
𝛼),' 𝛼(,' 𝛼',' 𝛼#,' 𝛼%,' = 𝑘1?G 𝑞& 𝑞@ 𝑞? 𝑞= 𝑞>
𝛼),# 𝛼(,# 𝛼',# 𝛼#,# 𝛼%,# 𝑘1=G
𝛼),% 𝛼(,% 𝛼',% 𝛼#,% 𝛼%,% 𝑘1>G
1
𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊-
Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
1
1 1
𝑑/ 𝛼),) 𝛼(,) 𝛼1',) 𝛼#,) 𝛼%,) 𝑘1&G
=
𝛼),( 𝛼(,( 𝛼1',( 𝛼#,( 𝛼%,( 𝑘1@G
𝛼),' 𝛼(,' 𝛼1',' 𝛼#,' 𝛼%,' = 𝑘1?G 𝑞&1𝑞@ 𝑞? 𝑞= 𝑞>
1
𝛼),# 𝛼(,# 𝛼1',# 𝛼#,# 𝛼%,# 𝑘1=G 1
𝛼),% 𝛼(,% 𝛼1',% 𝛼#,% 𝛼%,% 𝑘1>G
1 1
Softmax
* * 1* * *
𝛼),) 𝛼(,) 𝛼1',) 𝛼#,) 𝛼%,)
* * * * *
𝛼),( 𝛼(,( 𝛼1',( 𝛼#,( 𝛼%,(
* * * * *
𝛼),' 𝛼(,' 𝛼1',' 𝛼#,' 𝛼%,'
* * * * *
𝛼),# 𝛼(,# 𝛼1',# 𝛼#,# 𝛼%,#
* * * * *
𝛼),% 𝛼(,% 𝛼1',% 𝛼#,% 𝛼%,%
1
𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊-
Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
1
1 1
𝑑/ 𝛼),) 𝛼(,) 𝛼1',) 𝛼#,) 𝛼%,) 𝑘1&G
=
𝛼),( 𝛼(,( 𝛼1',( 𝛼#,( 𝛼%,( 𝑘1@G
𝛼),' 𝛼(,' 𝛼1',' 𝛼#,' 𝛼%,' = 𝑘1?G 𝑞&1𝑞@ 𝑞? 𝑞= 𝑞>
1
𝛼),# 𝛼(,# 𝛼1',# 𝛼#,# 𝛼%,# 𝑘1=G 1
𝛼),% 𝛼(,% 𝛼1',% 𝛼#,% 𝛼%,% 𝑘1>G
1 1
Softmax
* * 1* * * 𝑘 = 𝑘 ! , 𝑘 " , ⋯ , 𝑘 (0 ) 𝐸 𝑘 * = 𝐸 𝑞* = 0
𝛼),) 𝛼(,) 𝛼1',) 𝛼#,) 𝛼%,)
*
𝛼),( *
𝛼(,( *
𝛼1',( *
𝛼#,( *
𝛼%,( 𝑞 = 𝑞! , 𝑞" , ⋯ , 𝑞(0 )
Var 𝑘 * = Var 𝑞* = 1
* * * * *
𝛼),' 𝛼(,' 𝛼1',' 𝛼#,' 𝛼%,' (0
* * * * *
𝛼),# 𝛼(,# 𝛼1',# 𝛼#,# 𝛼%,#
*
𝛼),% *
𝛼(,% *
𝛼1',% *
𝛼#,% *
𝛼%,% 𝑘 ) 𝑞 = / 𝑘 * 𝑞* Var 𝑘 ) 𝑞 = 𝑑+
1 *,!
𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊-
Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
1
1 1
𝑑/ 𝛼),) 𝛼(,) 𝛼1',) 𝛼#,) 𝛼%,) 𝑘1&G
=
𝛼),( 𝛼(,( 𝛼1',( 𝛼#,( 𝛼%,( 𝑘1@G
𝛼),'
𝛼),#
𝛼(,'
𝛼(,#
𝐴
𝛼1','
𝛼1',#
𝛼#,'
𝛼#,#
𝛼%,' =
𝛼%,#
𝐾 𝑘1?G! 𝑞&1𝑞@ 𝑞? 𝑞= 𝑞>
1
𝑄
𝑘1=G 1
𝛼),% 𝛼(,% 𝛼1',% 𝛼#,% 𝛼%,% 𝑘1>G
1 1
Softmax
* * 1* * *
𝛼),) 𝛼(,) 𝛼1',) 𝛼#,) 𝛼%,)
* * * * *
𝛼),( 𝛼(,( 𝛼1',( 𝛼#,( 𝛼%,(
*"
*
𝛼),'
*
𝛼),#
*
𝛼(,'
*
𝛼(,#
𝐴
𝛼1','
*
𝛼1',#
*
𝛼#,'
*
𝛼#,#
*
𝛼%,'
*
𝛼%,#
* * * * *
𝛼),% 𝛼(,% 𝛼1',% 𝛼#,% 𝛼%,%
1
𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊-
Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
1
1 1
𝑑/ 𝛼),) 𝛼(,) 𝛼1',) 𝛼#,) 𝛼%,) 𝑘1&G
=
𝛼),( 𝛼(,( 𝛼1',( 𝛼#,( 𝛼%,( 𝑘1@G
𝛼),'
𝛼),#
𝛼(,'
𝛼(,#
𝐴
𝛼1','
𝛼1',#
𝛼#,'
𝛼#,#
𝛼%,' =
𝛼%,#
𝐾 𝑘1?G! 𝑞&1𝑞@ 𝑞? 𝑞= 𝑞>
1
𝑄
𝑘1=G 1
𝛼),% 𝛼(,% 𝛼1',% 𝛼#,% 𝛼%,% 𝑘1>G
1 1
Softmax
* * 1* * * * * 1* * *
𝛼),) 𝛼(,) 𝛼1',) 𝛼#,) 𝛼%,) 𝛼),) 𝛼(,) 𝛼',) 𝛼#,) 𝛼%,)
* * * * * = = * * 1* * *
𝛼),( 𝛼(,( 𝛼1',( 𝛼#,( 𝛼%,( 𝛼),( 𝛼(,( 𝛼1',( 𝛼#,( 𝛼%,(
*" *"
*
𝛼),'
*
𝛼),#
*
𝛼(,'
*
𝛼(,#
𝐴
𝛼1','
*
𝛼1',#
*
𝛼#,'
*
𝛼#,#
*
𝛼%,'
*
𝛼%,#
1
𝒙&( 1𝒙(@ 𝒙(? 𝒙(= 𝒙(> = 𝑊 F 𝑞& 1𝑣@ 𝑣? 𝑣= 𝑣>
𝑣
1
*
𝛼),'
*
𝛼),#
*
*
𝐴
𝛼(,'
𝛼(,#
𝛼1','
*
𝛼1',#
*
𝛼#,'
*
𝛼#,#
*
𝛼%,'
*
𝛼%,#
* * * * * 1. 1 * * * * *
𝛼),% 𝛼(,% 𝛼1',% 𝛼#,% 𝛼%,% 𝛼),% 𝛼(,% 𝛼1',% 𝛼#,% 𝛼%,%
Output features
1 1
𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊-
Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
1
𝑑/ 𝛼),) 𝛼(,)
1
𝛼1',) 𝛼#,) 𝛼%,)
1
𝑘1&G
=
𝑄 = 𝑊+ 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%
𝛼),( 𝛼(,( 𝛼1',( 𝛼#,( 𝛼%,( 𝑘1@G
𝛼),'
𝛼),#
𝛼(,'
𝛼(,#
𝐴
𝛼1','
𝛼1',#
𝛼#,'
𝛼#,#
𝛼%,' =
𝛼%,#
𝐾 𝑘1?G! 𝑞&1𝑞@ 𝑞? 𝑞= 𝑞>
1
𝑄 𝐾 = 𝑊, 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%
𝑘1=G 1
1
𝛼),% 𝛼(,% 𝛼1',% 𝛼#,% 𝛼%,%
1
𝑘1>G 𝑉 = 𝑊- 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%
Softmax
* * 1* * * * * 1* * *
𝛼),) 𝛼(,) 𝛼1',) 𝛼#,) 𝛼%,) 𝛼),) 𝛼(,) 𝛼1',) 𝛼#,) 𝛼%,)
*
𝛼),( *
𝛼(,( *
𝛼1',( *
𝛼#,( *
𝛼%,( = = *
𝛼),( *
𝛼(,( *
𝛼1',( *
𝛼#,( *
𝛼%,(
*" *"
*
𝛼),'
*
𝛼),#
*
𝛼(,'
*
𝛼(,#
𝐴
𝛼1','
*
𝛼1',#
*
𝛼#,'
*
𝛼#,#
*
𝛼%,'
*
𝛼%,#
1
𝒙&( 1𝒙(@ 𝒙(? 𝒙(= 𝒙(> = 𝑊F 𝑞& 1𝑣@ 𝑣? 𝑣= 𝑣>
𝑣
1 𝑉
*
𝛼),'
*
𝛼),#
*
*
𝐴
𝛼(,'
𝛼(,#
𝛼1','
*
𝛼1',#
*
𝛼#,'
*
𝛼#,#
*
𝛼%,'
*
𝛼%,#
* * * * *
1. 1 * * * * *
𝛼),% 𝛼(,% 𝛼1',% 𝛼#,% 𝛼%,% 𝛼),% 𝛼(,% 𝛼1',% 𝛼#,% 𝛼%,%
Output features
1 1
𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊-
Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Single-head attention 𝑄 = 𝑊+ 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%
Attention 𝑄, 𝐾, 𝑉 = 𝑉 softmax
𝐾*𝑄 𝐾 = 𝑊, 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%
𝑑+
𝑉 = 𝑊- 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%
𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊-
Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Analogy for Q, K, V
► Library system
► Imagine you're looking for information on a specific topic (query)
► Each book in the library has a summary (key) that helps identify if it contains the information
you're looking for
► Once you find a match between your query and a summary, you access the book to get the
detailed information (value) you need
► Here, in Attention, we do a “soft match” across multiple values, e.g. get info from multiple
books (“book 1 is most relevant, then book 2, then book 3, etc.”)
𝐾*𝑄
Attention 𝑄, 𝐾, 𝑉 = 𝑉 softmax
𝑑+
Single-head attention 𝑄 = 𝑊+ 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%
Attention 𝑄, 𝐾, 𝑉 = 𝑉 softmax
𝐾*𝑄 𝐾 = 𝑊, 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%
𝑑+
𝑉 = 𝑊- 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%
𝑊H
B
∈ 𝑅 %A×%
𝑊HC ∈ 𝑅 %A ×%
𝑊HE ∈ 𝑅 %D ×%
Single-head attention 𝑄 = 𝑊+ 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%
Attention 𝑄, 𝐾, 𝑉 = 𝑉 softmax
𝐾*𝑄 𝐾 = 𝑊, 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%
𝑑+
𝑉 = 𝑊- 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%
𝑊H
B
∈ 𝑅 %A×% 𝑊&' 𝑄 𝑊&( 𝐾 𝑊&) 𝑉
𝑊HC ∈ 𝑅 %A ×% Attention
𝑊H
B
∈ 𝑅 %A×% 𝑊&' 𝑋 𝑊&( 𝑋 𝑊&) 𝑋
𝑊HC ∈ 𝑅 %A ×% Attention
Attention 𝑄, 𝐾, 𝑉 = 𝑉 softmax
𝐾*𝑄 𝐾 = 𝑊, 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%
𝑑+
𝑉 = 𝑊- 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%
𝑊H
B
∈ 𝑅 %A×% 𝑊&' 𝑄 𝑊&( 𝐾 𝑊&) 𝑉 𝑊!' 𝑄 𝑊!( 𝐾 𝑊!) 𝑉
Attention 𝑄, 𝐾, 𝑉 = 𝑉 softmax
𝐾*𝑄 𝐾 = 𝑊, 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%
𝑑+
𝑉 = 𝑊- 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%
𝑊H
B
∈ 𝑅 %A×% 𝑊&' 𝑄 𝑊&( 𝐾 𝑊&) 𝑉 𝑊!' 𝑄 𝑊!( 𝐾 𝑊!) 𝑉
Attention 𝑄, 𝐾, 𝑉 = 𝑉 softmax
𝐾*𝑄 𝐾 = 𝑊, 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%
𝑑+
Multi-head attention 𝑉 = 𝑊- 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%
𝑊H
B
∈ 𝑅 %A×% 𝑊&' 𝑄 𝑊&( 𝐾 𝑊&) 𝑉 𝑊!' 𝑄 𝑊!( 𝐾 𝑊!) 𝑉
Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Feed Forward Network (FFN)
ReLU
𝐹𝐹𝑁 𝒙
⋮
⋮
= 𝑾4 ReLU 𝑾0 𝒙 + 𝒃0 + 𝒃4
𝑾( 𝑾) 𝒙
Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Feed Feed Feed Feed Feed
Forward Forward Forward Forward Forward
Encoder #1
Multi-head Self-Attention
Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Feed Feed Feed Feed Feed
Forward Forward Forward Forward Forward
Encoder #1
Multi-head Self-Attention
Embedded
Tokens 𝒙7 𝒙5 𝒙6 𝒙0 𝒙4
Tokens watch an apple I bought
Positional encoding
Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Positional encoding
Position 𝑘 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15
2' 0 0 0 0 0 0 0 0 1 1 1 1 1 1 1 1 Slow oscillating
0 0 0 0 1 1 1 1 0 0 0 0 1 1 1 1
2(
Dimension
0 0 1 1 0 0 1 1 0 0 1 1 0 0 1 1
2)
0 1 0 1 0 1 0 1 0 1 0 1 0 1 0 1
21 Fast oscillating
Positional
𝑑
embedding
Embedded
Tokens 𝑑 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Positional encoding
sin 𝑤1 𝑘
Position 𝑘 cos(𝑤1 𝑘) Fast oscillating
sin 𝑤) 𝑘
Angular frequency cos 𝑤) 𝑘
𝑑 ⋮
𝑤H = 𝑁 %@H/J
⋮
sin 𝑤2 𝑘
( 6) Slow oscillating
𝑁 = 100,000 cos 𝑤2 𝑘
6)
( Image: [Link]
Positional
𝑑
embedding
Embedded
Tokens 𝑑 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Normalized Range
Positional encoding
sin 𝑤1 𝑘 Unique identifier, unlimited length
Position 𝑘 cos(𝑤1 𝑘) Relative positions as linear transform
sin 𝑤) 𝑘
Angular frequency cos 𝑤) 𝑘
𝑑 sin(𝑤. (𝑘 + ∆𝑘) sin 𝑤. 𝑘 cos 𝑤. ∆𝑘 + cos 𝑤. 𝑘 sin 𝑤. ∆𝑘
𝑤H = 𝑁 %@H/J ⋮ =
cos(𝑤. (𝑘 + ∆𝑘) cos 𝑤. 𝑘 cos 𝑤. ∆𝑘 − sin 𝑤. 𝑘 sin 𝑤. ∆𝑘
co𝑠
⋮
sin 𝑤2 𝑘
( 6)
𝑁 = 100,000 cos 𝑤2 𝑘
6)
(
Positional
𝑑
embedding
Embedded
Tokens 𝑑 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Normalized Range
Positional encoding
sin 𝑤1 𝑘 Unique identifier, unlimited length
Position 𝑘 cos(𝑤1 𝑘) Relative positions as linear transform
sin 𝑤) 𝑘
Angular frequency cos 𝑤) 𝑘
𝑑 sin(𝑤. (𝑘 + ∆𝑘) sin 𝑤. 𝑘 cos 𝑤. ∆𝑘 + cos 𝑤. 𝑘 sin 𝑤. ∆𝑘
𝑤H = 𝑁 %@H/J ⋮ =
⋮ cos(𝑤. (𝑘 + ∆𝑘) cos 𝑤. 𝑘 cos 𝑤. ∆𝑘 − sin 𝑤. 𝑘 sin 𝑤. ∆𝑘
sin 𝑤2 𝑘 cos 𝑤. ∆𝑘 sin 𝑤. ∆𝑘 sin 𝑤
sin 𝑤..𝑘𝑘
( 6) =
𝑁 = 100,000 cos 𝑤2 𝑘 − sin 𝑤. ∆𝑘 cos 𝑤. ∆𝑘 co𝑠
co𝑠 𝑤𝑤..𝑘𝑘
6)
(
Positional
𝑃IJ∆I = 𝑀𝑃I
embedding
𝑑 𝑷0 𝑷4 𝑷5 𝑷6 𝑷7
Embedded
Tokens 𝑑 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Embedded
Tokens 𝑑 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Position 𝑘=1 𝑘=2 𝑘=3 𝑘=4 𝑘=5
sin 𝑤1 𝑘
𝑑 𝑷0 𝑷4 𝑷5 𝑷6 𝑷7
cos(𝑤1 𝑘)
sin 𝑤) 𝑘
cos 𝑤) 𝑘
𝑑 ⋮ 𝒙9 𝒙9 𝒙9
⋮
sin 𝑤2 𝑘
6)
concat MLP +
(
cos 𝑤2
(
6)
𝑘 𝑷9 𝑷9 𝑷9
Feed Feed Feed Feed Feed
Forward Forward Forward Forward Forward
Encoder #1
Multi-head Self-Attention
Embedded
Tokens 𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝒙7 𝑷7
Tokens I bought an apple watch
Positional encoding
sin 𝑤1 𝑘
Position 𝑘 cos(𝑤1 𝑘)
Sinusoidal positional encoding
sin 𝑤) 𝑘
Angular frequency cos 𝑤) 𝑘 Relative positional encoding
𝑑 ⋮
𝑤H = 𝑁 %@H/J
⋮ KERPLE RoPE CoPE
sin 𝑤2 𝑘
( 6)
𝑁 = 100,000 cos 𝑤2 𝑘
6)
NoPE YaRN FIRE
(
Positional
embedding
𝑑 𝑷0 𝑷4 𝑷5 𝑷6 𝑷7
Embedded
Tokens 𝑑 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Feed Feed Feed Feed Feed
Forward Forward Forward Forward Forward
Encoder #2
Multi-head Self-Attention
Embedded
Tokens 𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝒙7 𝑷7
Tokens I bought an apple watch
Feed Feed Feed Feed Feed
Forward Forward Forward Forward Forward
Encoder #2
Multi-head Self-Attention
Embedded
Tokens 𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝒙7 𝑷7
Tokens I bought an apple watch
Residual connection
Multi-head Self-Attention
Embedded
Tokens 𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝒙7 𝑷7
Tokens I bought an apple watch
Residual connection
Multi-head Self-Attention
Embedded
Tokens 𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝒙7 𝑷7
Tokens I bought an apple watch
Residual connection
Layer normalization
LayerNorm LayerNorm LayerNorm LayerNorm LayerNorm
Multi-head Self-Attention
Embedded
Tokens 𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝒙7 𝑷7
Tokens I bought an apple watch
Residual connection
Layer normalization
LayerNorm LayerNorm LayerNorm LayerNorm LayerNorm
Embedded
Tokens 𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝒙7 𝑷7
Tokens I bought an apple watch
Residual connection
Layer normalization
Embedded
Tokens 𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝒙7 𝑷7
Multi-head Self-Attention
LayerNorm LayerNorm LayerNorm LayerNorm LayerNorm
Embedded
Tokens 𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝒙7 𝑷7
Tokens I bought an apple watch
Encoder #6
Encoder #5
Encoder #4
Encoder #3
Encoder #2
Encoder #1
Embedded
Tokens 𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝒙7 𝑷7
Tokens I bought an apple watch
Encoder #6
Encoder #5
Encoder #4
Encoder #3
Encoder #2
Encoder #1
Embedded
Tokens 𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6
Tokens How are you ?
Encoder #6
Encoder #1
Embedded
Tokens 𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6
Tokens How are you ?
你 好 吃 飽 嗎 ⋯ ? <end>
Softmax
𝒆0 𝒆4 𝒆5 𝒆6 Linear
Encoder #𝐿 Decoder #𝐿
⋮
⋮ ⋮
Encoder #1 Decoder #1
𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝑑 𝒛0 𝑷0
Softmax
𝒆0 𝒆4 𝒆5 𝒆6 Linear
Encoder #𝐿 Decoder #𝐿
⋮
⋮ ⋮
Encoder #1 Decoder #1
𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝑑 𝒛0 𝑷0 𝒛4 𝑷4
Softmax
𝒛0 𝒛4 𝒛5 𝒛6 Linear
Encoder #𝐿 Decoder #𝐿
⋮
⋮ ⋮
Encoder #1 Decoder #1
𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝑑 𝒛0 𝑷0 𝒛4 𝑷4
Softmax
𝒆0 𝒆4 𝒆5 𝒆6 Linear
Encoder #𝐿 Decoder #𝐿
⋮
⋮ ⋮
Encoder #1 Decoder #1
𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝑑 𝒛0 𝑷0 𝒛4 𝑷4 𝒛5 𝑷5
Softmax
𝒆0 𝒆4 𝒆5 𝒆6 Linear
Encoder #𝐿 Decoder #𝐿
⋮
⋮ ⋮
Encoder #1 Decoder #1
𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝑑 𝒛0 𝑷0 𝒛4 𝑷4 𝒛5 𝑷5
Softmax
𝒆0 𝒆4 𝒆5 𝒆6 Linear
Encoder #𝐿 Decoder #𝐿
⋮
⋮ ⋮
Encoder #1 Decoder #1
𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝑑 𝒛0 𝑷0 𝒛4 𝑷4 𝒛5 𝑷5 𝒛6 𝑷6
Encoder #𝐿
⋮
Encoder #1 Decoder #1
𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝑑 𝒛0 𝑷0 𝒛4 𝑷4 𝒛5 𝑷5 𝒛6 𝑷6
Encoder #𝐿
⋮
Multi-head Self-Attention
Encoder #1
Decoder #1
𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝑑 𝒛0 𝑷0 𝒛4 𝑷4 𝒛5 𝑷5 𝒛6 𝑷6
Encoder #𝐿
⋮
Multi-head Self-Attention
Encoder #1
Decoder #1
𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝑑 𝒛0 𝑷0 𝒛4 𝑷4 𝒛5 𝑷5 𝒛6 𝑷6
Encoder #𝐿
⋮
Multi-head Self-Attention
Encoder #1
Decoder #1
𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝑑 𝒛0 𝑷0 𝒛4 𝑷4 𝒛5 𝑷5 𝒛6 𝑷6
Encoder #𝐿 𝐾 G𝑄
Attention 𝑄, 𝐾, 𝑉 = 𝑉 softmax
𝑑K
⋮
𝐾 G𝑄
MaskedAttention Encoder
𝑄, 𝐾, 𝑉 =#1
𝑉 softmax +𝑀 Multi-head Self-Attention
𝑑K
0 0 0 0 0 Decoder #1
−∞ 0 0 0 0
𝒙0= 𝑷0
𝑑𝑀
−∞ −∞
𝒙4 𝑷04 0𝒙5 0𝑷5 𝒙6 𝑷6 𝑑 𝒛0 𝑷0 𝒛4 𝑷4 𝒛5 𝑷5 𝒛6 𝑷6
−∞ −∞ −∞ 0 0
are
−∞ −∞ −∞ −∞ you
0 ? <start> 你 好 嗎
𝒆0 𝒆4 𝒆5 𝒆6
Encoder #𝐿 𝐾 G𝑄
Attention 𝑄, 𝐾, 𝑉 = 𝑉 softmax
𝑑K
⋮
𝐾 G𝑄
MaskedAttention Encoder
𝑄, 𝐾, 𝑉 =#1
𝑉 softmax +𝑀 Masked Multi-head Self-Attention
𝑑K
0 0 0 0 0 Decoder #1
−∞ 0 0 0 0
𝒙0= 𝑷0
𝑑𝑀
−∞ −∞
𝒙4 𝑷04 0𝒙5 0𝑷5 𝒙6 𝑷6 𝑑 𝒛0 𝑷0 𝒛4 𝑷4 𝒛5 𝑷5 𝒛6 𝑷6
−∞ −∞ −∞ 0 0
are
−∞ −∞ −∞ −∞ you
0 ? <start> 你 好 嗎
𝒆0 𝒆4 𝒆5 𝒆6
你
Masked Multi-head Self-Attention
你 好 Training
examples
Decoder #1
你 好 嗎
𝑑 𝒛0 𝑷0 𝒛4 𝑷4 𝒛5 𝑷5 𝒛6 𝑷6
你 好 嗎 ?
<start> 你 好 嗎
LayerNorm LayerNorm LayerNorm LayerNorm
𝐾-./01-2
LayerNorm LayerNorm LayerNorm LayerNorm
𝑉-./01-2
𝒆0 𝒆4 𝒆5 𝒆6
Encoder-decoder Attention
⋮
Masked Multi-head Self-Attention
Encoder #1
Decoder #1
𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝑑 𝒛0 𝑷0 𝒛4 𝑷4 𝒛5 𝑷5 𝒛6 𝑷6
Softmax
𝑘) 𝑞( 𝑘(& 𝑞( 𝑘'& 𝑞( 𝑘#& 𝑞(
𝑞@
𝑊+
LayerNorm LayerNorm
𝑘& 𝑣& 𝑘@ 𝑣@ 𝑘? 𝑣? 𝑘= 𝑣=
𝒆0 𝒆4 𝒆5 𝒆6 Decoder #1
Encoder #𝐿 𝑑 𝒛0 𝑷0 𝒛4 𝑷4
⋮
<start> 你
* * * *
𝛼(,) 𝛼(,( 𝛼(,' 𝛼(,#
𝒛() = 𝑊" *
𝛼(,) *
𝑣& + 𝛼(,( *
𝑣@ + 𝛼(,' *
𝑣? + 𝛼(,#𝑣=
Softmax
𝑘) 𝑞( 𝑘(& 𝑞( 𝑘'& 𝑞( 𝑘#& 𝑞( Cross-attention
Encoder-decoder attention
𝑞@
𝑊+
LayerNorm LayerNorm
𝑘& 𝑣& 𝑘& 𝑣@ 𝑘& 𝑣? 𝑘& 𝑣=
𝒆0 𝒆4 𝒆5 𝒆6 Decoder #1
Encoder #𝐿 𝑑 𝒛0 𝑷0 𝒛4 𝑷4
⋮
(ignore the scaling 1/ 𝑑/ here for simplicity)
<start> 你
LayerNorm LayerNorm LayerNorm LayerNorm
𝐾-./01-2
LayerNorm LayerNorm LayerNorm LayerNorm
𝑉-./01-2
𝒆0 𝒆4 𝒆5 𝒆6
Encoder-decoder Attention
⋮
Masked Multi-head Self-Attention
Encoder #1
Decoder #1
𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝑑 𝒛0 𝑷0 𝒛4 𝑷4 𝒛5 𝑷5 𝒛6 𝑷6
Encoder #𝐿
⋮
Encoder #1 Decoder #1
𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝑑 𝒛0 𝑷0 𝒛4 𝑷4 𝒛5 𝑷5 𝒛6 𝑷6
Encoder #𝐿 Decoder #𝐿
⋮
⋮ ⋮
Encoder #1 Decoder #1
𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝑑 𝒛0 𝑷0 𝒛4 𝑷4 𝒛5 𝑷5 𝒛6 𝑷6
Good for:
Machine translation, summarization. QA
(when input/target are sufficiently different)
𝐾-./01-2 Softmax
𝑉-./01-2
𝒆0 𝒆4 𝒆5 𝒆6 Linear
Encoder #𝐿 Decoder #𝐿
⋮
⋮ ⋮
Encoder #1 Decoder #1
Examples:
Attention is all you need, T5, BART.
Good for:
Machine translation, summarization. QA
(when input/target are sufficiently different)
𝒆0 𝒆4 𝒆5 𝒆6
Encoder #𝐿
⋮
Encoder #1
Examples:
BERT, RoBERTa, DeBERTa, X-BERT
Good for:
Classification, sequence tagging, sentiment
analysis
(Understand text, but not generate them)
𝒆0
Encoder #𝐿
⋮
Encoder #1
Examples:
GPT-X (OpenAI), PaLM (Google), LLaMA (Meta)
BLOOM (BigScience)
Good for:
Text generation, multi-round conversation
Softmax
Linear
Decoder #𝐿
⋮
Decoder #1
Softmax Softmax
𝒆0 𝒆4 𝒆5 𝒆6 Linear Linear
2023
OPT-IML
BLOOMZ Galactica
Flan
T5
BLOOM
YaLM
Open-Source OPT
Tk
GPT-NeoX
ST-MoE
2022
GPT-J
GLM
GPT-Neo
Switch
2021
mT5 T0
DeBERTa
ELECTRA
2020
Distill T5
ALBERT BERT BART
XLNet
RoBERTa
ERNIE
GPT-2
Enc
ode
2019 r-D
eco
BERT der
Encoder-On GPT-1
ELMo ULMFiT ly Decoder-Only
2018
GloVe
FastText
Word2Vec
2023
OPT-IML
BLOOMZ Galactica
Flan
T5
Open-Source BLOOM
YaLM
OPT
Tk
GPT-NeoX
ST-MoE
2022
GPT-J
GLM
GPT-Neo
Switch
2021
mT5 T0
DeBERTa
ELECTRA
2020
Distill T5
ALBERT BERT BART
RoBERTa XLNet
ERNIE
GPT-2
Enc
ode
2019 r-D
BERT eco
der
Encoder-On GPT-1
ELMo ULMFiT ly Decoder-Only
2018
FastText GloVe
Word2Vec
16
8
Transformers vs. RNNs
► When sequence length (n) << representation dimension (d), the complexity per layer
is lower for a Transformer model compared to RNN models ; NOT true for real-world LLMs
Differences in Attention Mechanism of RNN vs. Transformer
Based on decoder hidden state and Based on dot-product of query and keys
Alignment
encoder hidden states (global attention)
Weighted sum of encoder hidden Attends to all encoder positions for every
Context
states at each step output
Transformer Advantages:
● # unparallelizable operations does not
RNN-Based Encoder-Decoder
increase with sequence length.
Model with Attention
● Each word interacts with each other, so
maximum interaction distance is O(1).
Transformer-Based
Encoder-Decoder Model