0% found this document useful (0 votes)
8 views170 pages

Basic

The document outlines the basic architecture of Transformers and the Word2vec model, focusing on the Skip-Gram with Negative Sampling (SGNS) approach for learning word embeddings. It discusses various NLP tasks modeled as sequence-to-sequence problems and critiques early neural network methods like RNNs and LSTMs for their limitations. The document also highlights the evolution of attention mechanisms leading to the rise of Transformers in NLP applications.

Uploaded by

howe108
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
8 views170 pages

Basic

The document outlines the basic architecture of Transformers and the Word2vec model, focusing on the Skip-Gram with Negative Sampling (SGNS) approach for learning word embeddings. It discusses various NLP tasks modeled as sequence-to-sequence problems and critiques early neural network methods like RNNs and LSTMs for their limitations. The document also highlights the evolution of attention mechanisms leading to the rise of Transformers in NLP applications.

Uploaded by

howe108
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

IERG5050 AI Foundation Models, Systems and Applications

Fall 2025

Transformers Part I: Basic Architecture

Prof. Wing C. Lau


wclau@[Link]
[Link]

1
Acknowledgements
Many of the slides in this lecture are adapted from the sources below. Copyrights belong to the original authors.

● Stanford CS25: Transformer United V4, Spring 2024, [Link]


Instructors: Div Garg, Steven Feng, Seonghee Lee, Emily Bunnapradist ;
Faculty Advisor: Prof. Chris Manning,
Overview Slides [Link]
● Stanford CS336: Language Modeling from Scratch, Spring 2025
○ by Profs. Tatsunori Hashimoto, Percy Liang, [Link]
● Stanford CS229S: Systems for Machine Learning, Fall 2023
by Profs. Azalia Mirhoseini, Simran Arora, [Link]
● Stanford CS224N: Natural Language Processing with Deep Learning, Winter 2021
by Prof. Chris Manning, [Link]
● Stanford CS231n: Deep Learning for Computer Vision, Spring 2023
by Prof. Fei-fei Li, [Link]
● CMU 11-667: Large Language Models: Methods and Applications, Fall 2024
by Profs. Chenyan Xiong and Daphne Ippolito, [Link]
● CMU 11-711: Advanced Natural Language Processing (ANLP), Spring 2024
by Prof. Graham Neubig, [Link]
● UPenn CIS7000: Large Language Models, Fall 2024
by Prof. Mayur Naik, [Link]
● Princeton COS597G: Understanding Large Language Models, Fall 2022
by Prof. Danqi Chen, [Link]
● UWaterloo CS886: Recent Advances on Foundation Models, Winter 2024
by Prof. Wenhu Chen, [Link]
● UMD CMSC848K: Multimodal Foundation Models, Fall 2024
by Prof. Jia-Bin Huang, [Link]
2
Natural Language Processing (NLP) & Language Modeling
Representing Words as Discrete Symbols
Problem with Words as Discrete Symbols
Representing Words by their Context
Word Vectors (aka Word Embeddings)
Word2vec: How to learn the Word Embedding

Source: [Link]
For details, refer to: [Link]
Word2vec Overview

Source: [Link]
For details, refer to: [Link]
Word2vec Overview

Source: [Link]
For details, refer to: [Link]
Word2vec: Objective Function

Source: [Link]
For details, refer to: [Link]
Word2vec: Objective Function

Source: [Link]
For details, refer to: [Link]
Word2vec: Prediction Function = Prob[o|c]

Source: [Link]
For details, refer to: [Link]
Word2vec: To Train the Model

Source: [Link]
For details, refer to: [Link]
The problematic P[o|c]

exp 𝑢!" 𝑣𝑐
𝑃 𝑜 𝑐) = &
∑#$% exp(𝑢#" 𝑣𝑐 )

The denominator is difficult to evaluate as


it involves the embedding of ALL words in the Universe (Vocabulary)

=> Use a trick to circumvent the problem via Negative Sampling


15
What is Word2vec ?
► word2vec is not a single algorithm
► It is a software package for representing words as vectors, containing:
► Two distinct models
► CBoW
► Skip-Gram (SG)
► Various training methods
► Negative Sampling (NS)
► Hierarchical Softmax
► A rich preprocessing pipeline
► Dynamic Context Windows
► Subsampling
► Deleting Rare Words
► We will focus on the Skip-Grams with Negative Sampling (SGNS) approach !

Slide by Omer Levy


SGNS starts with SAME basic setting
• SGNS finds a vector 𝑣𝑐 for each word 𝑐 in our vocabulary 𝑉!
• Each such vector has 𝑑 latent dimensions (e.g. 𝑑 = 100)
• Effectively, it learns a matrix 𝐶 whose rows represent the center vectors 𝑣𝑐
• Key point: it also learns a similar auxiliary matrix 𝑂 of outside context vectors
• In fact, each word has two embeddings: 𝑣𝑐 and 𝑢𝑜
𝑑 𝑑
𝑐:wampimuk =
(−3.1, 4.15, 9.2, −6.5, … )
𝐶 ≠ 𝑂

𝑈"
𝑉!

𝑜:wampimuk = “word2vec Explained…”


(−5.6, 2.95, 1.4, −1.3, … ) 17
Goldberg & Levy, arXiv
2014
Skip-Grams with Negative Sampling (SGNS)

You first observe (actually sample) the following sentence from the training
corpus:

Marco saw a furry little wampimuk hiding in the tree.

“word2vec Explained…”
Goldberg & Levy, arXiv
2014
Skip-Grams with Negative Sampling (SGNS)
Marco saw a furry little wampimuk hiding in the tree.

center word outside context word


wampimuk furry
wampimuk little 𝐷 (observed data)
wampimuk hiding
wampimuk in
… …
“word2vec Explained…” Goldberg & Levy, arXiv 2014 19
Skip-Grams with Negative Sampling (SGNS)
Maximize ∏ 𝑖 𝜎 𝑐⃗ ⋅ 𝑜𝑖 AND• Minimize ∏ 𝑖 𝜎 𝑐⃗ ⋅ 𝑜𝑖 ′
• 𝑜𝑖 was observed with 𝑐 ≡ Maximize ∏ 𝑖 [1- 𝜎 𝑐⃗ ⋅ 𝑜𝑖 ′ ]
where 𝜎 𝑧 = 1/ [1+exp(- 𝑧)] • 𝑜𝑖′ was NOT observed with 𝑐, they are
from the set of Negative Samples D’
randomly generated by the algorithm.
center word outside context
center word NOT outside context
wampimuk furry
wampimuk Australia
wampimuk little
wampimuk cyber
wampimuk hiding
wampimuk the
wampimuk in
wampimuk 1985

Take Log and the optimization problem becomes “similar” to the training of a binary logistic-regression classifier
( [Link] ):
Summary: How to learn Word2vec Embeddings via SGNS

21
NLP Tasks as Sequence to Sequence Modeling

22
NLP Tasks as Sequence to Sequence Modeling
► Example Scenarios
► Text à Text (e.g. Q/A, translation, text summarization)
► Image à Text (e.g. image captioning)

Output
𝑦! 𝑦" 𝑦# 𝑦$ sequence

magic?

Input
sequence 𝑥! 𝑥" 𝑥# 𝑥$
NLP Tasks as Sequence to Sequence Modeling
24

► Example Scenarios
► Text à Text (e.g. Q/A, translation, text summarization)
► Image à Text (e.g. image captioning)

► How? Usually Encoder-Decoder models


► e.g. RNNs, LSTMs, Transformers Output
𝑦! 𝑦" 𝑦# 𝑦$ sequence

state
ENCODER DECODER
context
vector

Input
sequence 𝑥! 𝑥" 𝑥# 𝑥$
Recurrent Neural Networks (RNN) – a Seq2Seq NN model

An RNN “Unrolled” along the Time axis


Recurrent Neural Networks (RNN)
(Vanilla) Recurrent Neural Networks (RNN)
Computational Graph for an RNN

Note the reusing of the SAME weight matrix fw at every time-step !


Example: Character-level Language Model
Example: Character-level Language Model
Example: Character-level Language Model
32
Weaknesses of early NN-based NLP approaches

► Short context length


► “Linear” reasoning - no attention mechanism to focus on other parts
► Earlier approaches (e.g. word2vec) do not adapt based on context.
Seq2Seq Models w/ Neural Nets: the Pre-Transformer Era
The inputs to each unit consists of
the current input xt, previous hidden
state ht-1, and previous context ct-1
● Recurrent Neural Networks (RNNs)
● Long Short-Term Memory Networks (LSTMs)
● Capture dependencies between input tokens
● Gates control the flow of information

A single LSTM unit displayed as a


computation graph.

The outputs are a new hidden state ht


and an updated context ct.

A simple RNN shown unrolled in time. Network layers are recalculated for
each time step, while weights U, V and W are shared across all time steps.
Better Capturing of Long-Range Dependence
using LSTM for Seq2Seq Modeling
► Encoder (LSTM) and decoder (LSTM)
► Fixed-length context vector

Input: sequence 𝑥& , … , 𝑥' Output: sequence 𝑦& , … , 𝑦'(


ℎ$ = 𝑓(𝑥$ , ℎ$%& ) 𝑠$ = 𝑔(𝑦$%& , 𝑠$%& , 𝑐)
𝑦! 𝑦" 𝑦# 𝑦$

Initial state

ℎ! 𝑓 ℎ" 𝑓 ℎ# 𝑓 ℎ$ 𝑠% 𝑠! 𝑔 𝑠" 𝑔 𝑠# 𝑔 𝑠$

𝑥! 𝑥" 𝑥# 𝑥$ c 𝑦% 𝑦! 𝑦" 𝑦#
Context vector
ENCODER DECODER
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in Proceedings of the 27th International Conference on Neural Information
Processing Systems (NIPS), 2014, pp. 3104–3112.
LSTM still suffers from Information Bottleneck
► Encoder (LSTM) and decoder (LSTM)
► Fixed-length context vector (bottleneck)

Input: sequence 𝑥& , … , 𝑥' Output: sequence 𝑦& , … , 𝑦'(


ℎ$ = 𝑓(𝑥$ , ℎ$%& ) 𝑠$ = 𝑔(𝑦$%& , 𝑠$%& , 𝑐)
𝑦! 𝑦" 𝑦# 𝑦$

Initial state

ℎ! 𝑓 ℎ" 𝑓 ℎ# 𝑓 ℎ$ 𝑠% 𝑠! 𝑔 𝑠" 𝑔 𝑠# 𝑔 𝑠$

BOTTLENECK

𝑥! 𝑥" 𝑥# 𝑥$ c 𝑦% 𝑦! 𝑦" 𝑦#
Context vector
ENCODER DECODER
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in Proceedings of the 27th International Conference on Neural Information
Processing Systems (NIPS), 2014, pp. 3104–3112.
1990s Prehistoric Era
Rule-based methods, parsing, RNNs, LSTMs

Attention Timeline
2014 Simple attention mechanisms

2017 Beginning of transformers


Attention is all you need

2018 Explosion of transformers in NLP


BERT, GPT-3

2018-2020 Explosion into other fields


Explosion into other fields: ViTs, Alphafold-2.

2021-2022 Start of Generative Era


Codex, Decision Transformers, GPT-X, DALL·E

2024 Present Day


Huge models, more applications: Chat-GPT, GPT-4, Gemini,
Llama and open-source LLMs, Whisper, Robotics Transformer,
Stable Diffusion, Sora, LLM Agents, Multimodal, and so much
more…!
Future (?!) 36
Sequence to Sequence with RNNs + Attention 37

► Idea! Use a different context vector for each timestep in the decoder
𝑠. = 𝑔(𝑦./0 , 𝑠./0 , 𝒄𝒕 )

► No more bottleneck through a single vector

► Craft the context vector so that it “looks at” different parts of the input
sequence for each decoder timestep

D. Bahdanau, K. Cho, and Y. Bengio, “Neural Machine Translation by Jointly Learning to Align and Translate,” in 3rd International Conference on Learning Representations
(ICLR), 2015.
Sequence to Sequence with RNNs + Attention
38
Compute context vector
Find 𝑠!: 𝑡 = 1
× × × × 𝑐' = 4 𝛼',( ℎ(
+ (
Attention weights
𝛼!,! 𝛼!," 𝛼!,# 𝛼!,$ (normalize alignment
scores) 𝑠$ = 𝑔(𝑦$%& , 𝑠$%& , 𝒄𝒕 )
softmax
𝑦!
Alignment scores
𝑒!,! 𝑒!," 𝑒!,# 𝑒!,$ 𝛼',( represents the probability that
𝑒',( = 𝑓)'' (𝑠'*! , ℎ( )
the target word 𝑦' is aligned to, or
translated from, a source word 𝑥(

ℎ! 𝑓 ℎ" 𝑓 ℎ# 𝑓 ℎ$ 𝑠% 𝑠! 𝛼',( reflects the importance of the


annotation ℎ( with respect to the
Initial state previous hidden state 𝑠'*! in
deciding the next state 𝑠' and
generating 𝑦'
𝑥! 𝑥" 𝑥# 𝑥$ 𝑐! 𝑦%
Context
ENCODER vector DECODER

D. Bahdanau, K. Cho, and Y. Bengio, “Neural Machine Translation by Jointly Learning to Align and Translate,” in 3rd International Conference on Learning Representations
(ICLR), 2015.
Sequence to Sequence with RNNs + Attention
39
Compute context vector
Find 𝑠": 𝑡 = 2
× × × × 𝑐' = 4 𝛼',( ℎ(
+ (
Attention weights
𝛼",! 𝛼"," 𝛼",# 𝛼",$ (normalize alignment
scores) 𝑠$ = 𝑔(𝑦$%& , 𝑠$%& , 𝒄𝒕 )
softmax
𝑦! 𝑦"
Alignment scores
𝑒",! 𝑒"," 𝑒",# 𝑒",$
𝑒',( = 𝑓)'' (𝑠'*! , ℎ( )

ℎ! 𝑓 ℎ" 𝑓 ℎ# 𝑓 ℎ$ 𝑠% 𝑠! 𝑔 𝑠"

Initial state

𝑥! 𝑥" 𝑥# 𝑥$ 𝑐! 𝑦% 𝑐" 𝑦!
Context
ENCODER vector DECODER

D. Bahdanau, K. Cho, and Y. Bengio, “Neural Machine Translation by Jointly Learning to Align and Translate,” in 3rd International Conference on Learning Representations
(ICLR), 2015.
Sequence to Sequence with RNNs + Attention
40
Compute context vector
Find 𝑠' : 𝑡 = 𝑡
× × × × 𝑐' = 4 𝛼',( ℎ(
+ (
Attention weights
𝛼',! 𝛼'," 𝛼',# 𝛼',$ (normalized alignment
scores) 𝑠$ = 𝑔(𝑦$%& , 𝑠$%& , 𝒄𝒕 )
softmax
𝑦! 𝑦" 𝑦# 𝑦$
Alignment scores 𝑐'
𝑒',! 𝑒'," 𝑒',# 𝑒',$
𝑒',( = 𝑓)'' (𝑠'*! , ℎ( )

𝑠'*!

ℎ! 𝑓 ℎ" 𝑓 ℎ# 𝑓 ℎ$ 𝑠% 𝑠! 𝑔 𝑠" 𝑔 𝑠# 𝑔 𝑠$

Initial state

𝑥! 𝑥" 𝑥# 𝑥$ 𝑐! 𝑦% 𝑐" 𝑦! 𝑐# 𝑦" 𝑐# 𝑦#


Context
ENCODER vector DECODER

D. Bahdanau, K. Cho, and Y. Bengio, “Neural Machine Translation by Jointly Learning to Align and Translate,” in 3rd International Conference on Learning Representations
(ICLR), 2015.
Sequence to Sequence with RNNs + Attention
41
All steps are differentiable,
Compute context vector so we can backpropagate
Find 𝑠' → 𝑡 = 𝑡
through everything
× × × × 𝑐' = 4 𝛼',( ℎ(
+ (
Attention weights
𝛼',! 𝛼'," 𝛼',# 𝛼',$ (normalized alignment
scores) 𝑠$ = 𝑔(𝑦$%& , 𝑠$%& , 𝒄𝒕 )
softmax
𝑦! 𝑦" 𝑦# 𝑦$
Alignment scores 𝑐'
𝑒',! 𝑒'," 𝑒',# 𝑒',$
𝑒',( = 𝑓)'' (𝑠'*! , ℎ( )

𝑠'*!

ℎ! 𝑓 ℎ" 𝑓 ℎ# 𝑓 ℎ$ 𝑠% 𝑠! 𝑔 𝑠" 𝑔 𝑠# 𝑔 𝑠$

Initial state
Encoder is bi-directional: allows for the annotation of each
word to summarize both preceding and following words.

𝑥! 𝑥" 𝑥# 𝑥$ 𝑐! 𝑦% 𝑐" 𝑦! 𝑐# 𝑦" 𝑐# 𝑦#


Context
ENCODER vector DECODER

D. Bahdanau, K. Cho, and Y. Bengio, “Neural Machine Translation by Jointly Learning to Align and Translate,” in 3rd International Conference on Learning Representations
(ICLR), 2015.
Sequence to Sequence with RNNs + Attention

Application: translation

09-10
2025-
Each pixel shows the weight 𝛼",$ of
the annotation of the 𝑖-th source
word for the 𝑡-th target word.

D. Bahdanau, K. Cho, and Y. Bengio, “Neural Machine Translation by Jointly Learning to Align and Translate,” in 3rd International Conference on Learning Representations
(ICLR), 2015.
Sequence to Sequence with RNNs + Attention 43

Application:

09-10
2025-
UWaterloo
Slides created for CS886 at
text translation

RNN:
RNNenc

RNN + attention:
RNNsearch

D. Bahdanau, K. Cho, and Y. Bengio, “Neural Machine Translation by Jointly Learning to Align and Translate,” in 3rd International Conference on Learning Representations
(ICLR), 2015.
Motivating Transformers by
Understanding the Limitations of Recurrent Models

Challenge 1: Modeling long-range dependencies


Challenge 2: Optimization due to vanishing and exploding gradients
Challenge 3: Slow (sequential, serial) bottleneck

Caveating this discussion: while the challenges we’ll discuss originally motivated
Transformers, many have continued to make progress on RNNs over the years (S4, Mamba,
Linear Attention, GLA, Based, etc.)
Challenge 1: Long Interaction Distances

E.g. “The counselor helped


frame the situation.”

• Performance degrades
as the distance between
words increases due to
memory constraints:
Diluted impact of earlier The counselor situation

elements on output as
sequence progresses
“The” has gone through t layers
before getting to “situation”
While RNNs + “Attention” has made some progress, the
coupling of the “sequential” structure of RNN with Attention
still creates difficulties

Key idea of the Decoders

RNN + Attention operation:


All tokens interact with all
other tokens’ representations
Encoders
Challenge 2: RNNs/ LSTMs are difficult to train!

Backpropagation through many


timesteps/”layers”…

Recall: backpropogation is about


updating the parameters in a way that
reduces the loss. We multiply with
respect to each set of parameters at
each timestep.

Figure Source: O’Reilly Media


Backpropagation through Time
Challenge 2: RNNs/LSTMs are difficult to train!

If the value we are multiplying is large,


our gradients will grow exponentially!
The model becomes unstable!

If the value we are multiplying is


small, our gradients will get smaller
each timestep, going to 0. The
network stops learning/learns too
slowly.

Figure Source: O’Reilly Media


Challenge 3: Parallelizability

Each time step needs to be processed before we can move


onto the next step.
Decoupling “Attention” from RNNs 51

► Recall: attention determines the importance of elements to be passed


forward in the model.
► These weights lets the model pay attention to the most
significant parts

► Objective: a more general attention mechanism not confined to RNNs


► We need a modified procedure to:
1. Determine weights based on context that indicate the
elements to attend to
2. Apply these weights to enhance attended features
Self-Attention and Transformers

● Allows to “focus attention” on particular aspects of


the input while generating the output.
● Done by using a set of parameters, called "weights,"
that determine how much attention should be paid to
each input token at each time step.
● These weights are computed using a combination of
the input and the current hidden state of the model.

In encoding the word "it", one attention head is


focusing most on "the animal", while another is
focusing on "tired". The model's representation
A. Vaswani et al. Attention Is All You Need. NeurIPS 2017. of the word "it" thus bakes in some of the
representation of both "animal" and "tired".
[Link]
Transformers – the current ”standard” for
building LLMs and Foundation Models

53
Slides for video from:
Prof. Jia-Bin Huang
University of Maryland, College Park
[Link]

🤔 🤔 🤔
What? How? Why?
En ZH
How are you? 你好嗎?

Sequence-to- Sequence model


你 好 吃 飽 嗎 ⋯ ? <end>
ZH

Encoders Decoders

En How are you? <start>


你 好 吃 飽 嗎 ⋯ ? <end>
ZH

Encoders Decoders

En How are you? <start> 你


你 好 吃 飽 嗎 ⋯ ? <end>
ZH

Encoders Decoders

En How are you? <start> 你


你 好 吃 飽 嗎 ⋯ ? <end>
ZH

Encoders Decoders

En How are you? <start> 你 好


你 好 吃 飽 嗎 ⋯ ? <end>
ZH

Encoders Decoders

En How are you? <start> 你 好


你 好 吃 飽 嗎 ⋯ ? <end>
ZH

Encoders Decoders

En How are you? <start> 你 好 嗎


你 好 吃 飽 嗎 ⋯ ? <end>
ZH

Encoders Decoders

En How are you? <start> 你 好 嗎


你 好 吃 飽 嗎 ⋯ ? <end>
ZH

Encoders Decoders

En How are you? <start> 你 好 嗎 ?


你 好 吃 飽 嗎 ⋯ ? <end>
ZH

Encoders Decoders

En How are you? <start> 你 好 嗎 ?


你 好 吃 飽 嗎 ⋯ ? <end>
ZH

Encoders Decoders

En How are you? <start> 你 好 嗎 ? <end>


Encoders

How are you?


8607 4339 2472 311 832 4037 11 719 1063 1541 956 25 3687 23936 13

cat dog bear cow indiv


1 0 0 0 0
0 1 0 0 0
0 0 1 0 0
One-hot encoding
0 0 0 1 0 Value 1 at
# tokens 0 0 0 0 0 3687th
⋮ ⋮ ⋮ ⋮ ⋮ entry
0 0 0 0 0
One-hot encoding

cat dog bear cow indiv


1 0 0 0 0
0 1 0 0 0
0 0 1 0 0
0 0 0 1 0 Value 1 at
0 0 0 0 0 3687th
⋮ ⋮ ⋮ ⋮ ⋮ entry
0 0 0 0 0
cat dog bear cow indiv
1 0 0 0 0
0 1 0 0 0
0 0 1 0 0
0 0 0 1 0 Value 1 at
0 0 0 0 0 3687th
⋮ ⋮ ⋮ ⋮ ⋮ entry
0 0 0 0 0
Embedding Space
dog
0
# tokens
0.5 0 1

𝑊!
2.7 0 0
𝑑 1.2 = 𝑑 0 0
⋮ 0 0
0.2 0 ⋮
0
Embedding Space Embedded Embedding Matrix
token
dog
0
# tokens
0.5 0 1

𝑊!
2.7 0 0
𝑑 1.2 = 𝑑 0 ⋯ 0
⋮ 0 0
0.2 0 ⋮
0
Embedding Space Embedded Embedding Matrix
token
Apple

dog
0
# tokens
0.5 0 1

𝑊!
2.7 0 0
𝑑 1.2 = 𝑑 0 ⋯ 0
⋮ 0 0
0.2 0 ⋮
0
Embedding Space Embedded Embedding Matrix
token
I bought an apple and an orange.
Apple

I bought an apple watch.


dog
0
# tokens
0.5 0 1

𝑊!
2.7 0 0
𝑑 1.2 = 𝑑 0 ⋯ 0
⋮ 0 0
0.2 0 ⋮
0
Embedding Space Embedded Embedding Matrix
token
Encoder

Embedded
Tokens

Token
𝑊< 𝑊< 𝑊< 𝑊< 𝑊< 𝑊< 𝑊<
Embedding
Tokens I bought an apple and an orange
Feed Feed Feed Feed Feed Feed Feed
Forward Forward Forward Forward Forward Forward Forward

Embedded
Tokens

Token
Embedding 𝑊< 𝑊< 𝑊< 𝑊< 𝑊< 𝑊< 𝑊<

Tokens I bought an apple and an orange


Feed Feed Feed Feed Feed Feed Feed
Forward Forward Forward Forward Forward Forward Forward

Embedded
Tokens

Token
Embedding 𝑊< 𝑊< 𝑊< 𝑊< 𝑊< 𝑊< 𝑊<

Tokens I bought an apple and an orange


Feed Feed Feed Feed Feed Feed Feed
Forward Forward Forward Forward Forward Forward Forward

Embedded
Tokens

Token
Embedding 𝑊< 𝑊< 𝑊< 𝑊< 𝑊< 𝑊< 𝑊<

Tokens I bought an apple and an orange


Feed Feed Feed Feed Feed Feed Feed
Forward Forward Forward Forward Forward Forward Forward

Self-Attention

Embedded
Tokens

Token
Embedding 𝑊< 𝑊< 𝑊< 𝑊< 𝑊< 𝑊< 𝑊<

Tokens I bought an apple and an orange


Self-Attention

apple

Embedding Space

Embedded
Tokens

Tokens I bought an apple and an orange


Self-Attention

apple
Embedding Space

Embedded
Tokens

Tokens I bought an apple and an orange


Self-Attention

apple

Embedding Space

Embedded
Tokens

Tokens I bought an apple watch


Self-Attention

apple

Embedding Space

Embedded
Tokens

Tokens I bought an apple watch


Self-Attention

Embedding Space

Embedded
Tokens

Tokens I bought an apple watch


Self-Attention

Embedding Space

Embedded
Tokens

Tokens I bought an apple watch


Self-Attention

𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%


= 𝒙&# 𝒙) = 𝒙&# 𝒙( = 𝒙&# 𝒙' = 𝒙&# 𝒙# = 𝒙&# 𝒙%

Embedded
Tokens 𝒙0∈ 𝑅8 𝒙4 ∈ 𝑅 8 𝒙5 ∈ 𝑅 8 𝒙6 ∈ 𝑅 8 𝒙7 ∈ 𝑅 8
Tokens I bought an apple watch
* * * * *
𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%
0.082 0.0495 0.0199 0.6034 0.2452

Sum normalization

1.1 0.67 Softmax


0.27 8.17 3.32

Exp Exp: exp


Exp 𝛼6,9 Exp Exp
𝛼6,9 =
∑; exp 𝛼6,;
0.1 −0.4 −1.3 2.1 1.2
𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%
= 𝒙&# 𝒙) = 𝒙&# 𝒙( = 𝒙&# 𝒙' = 𝒙&# 𝒙# = 𝒙&# 𝒙%

Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
* * * * *
𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%

Softmax
𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%
= 𝒙&# 𝒙) = 𝒙&# 𝒙( = 𝒙&# 𝒙' = 𝒙&# 𝒙# = 𝒙&# 𝒙%

Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
* * * * *
𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%

Softmax
𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%
= 𝒙&# 𝒙) = 𝒙&# 𝒙( = 𝒙&# 𝒙' = 𝒙&# 𝒙# = 𝒙&# 𝒙%

Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
' 𝒙 ' 𝒙 ' 𝒙 ' 𝒙 ' 𝒙
Updated 𝒙'$ = 𝛼$,! ! + 𝛼$," " + 𝛼$,# # + 𝛼$,$ $ + 𝛼$,% %
feature

* * * * *
watch 𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%

Softmax
I apple 𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%
= 𝒙&# 𝒙) = 𝒙&# 𝒙( = 𝒙&# 𝒙' = 𝒙&# 𝒙# = 𝒙&# 𝒙%

an bought

Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
' 𝒙 ' 𝒙 ' 𝒙 ' 𝒙 ' 𝒙
Updated 𝒙'$ = 𝛼$,! ! + 𝛼$," " + 𝛼$,# # + 𝛼$,$ $ + 𝛼$,% %
feature

* * * * *
watch 𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%

apple
Softmax
I 𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%
= 𝒙&# 𝒙) = 𝒙&# 𝒙( = 𝒙&# 𝒙' = 𝒙&# 𝒙# = 𝒙&# 𝒙%

an bought

Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
' 𝒙 ' 𝒙 ' 𝒙 ' 𝒙 ' 𝒙
Updated 𝒙'$ = 𝛼$,! ! + 𝛼$," " + 𝛼$,# # + 𝛼$,$ $ + 𝛼$,% %
feature

* * * * *
𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%
delicious apple
Softmax
𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%
= 𝒙&# 𝒙) = 𝒙&# 𝒙( = 𝒙&# 𝒙' = 𝒙&# 𝒙# = 𝒙&# 𝒙%

Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
* * * * *
𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%

Softmax
= 𝑘𝒙)&&
𝛼#,) = #𝑞𝒙#) 𝛼#,(= 𝑘(& 𝑞# 𝛼#,' = 𝑘'& 𝑞# 𝛼#,#= 𝑘#& 𝑞# 𝛼#,% = 𝑘%& 𝑞#
𝑊 B ∈ 𝑅 %A ×%

𝑊 C ∈ 𝑅 %A ×%
𝑘& 𝑘@ 𝑘? 𝑞= 𝑘= 𝑘>

𝑊, 𝑊, 𝑊, 𝑊+ 𝑊, 𝑊,

Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
* * * * *
𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%

Softmax
= 𝑘𝒙)&&
𝛼#,) = #𝑞𝒙#) 𝛼#,( = 𝑘(& 𝑞# 𝛼#,' = 𝑘'& 𝑞# 𝛼#,# = 𝑘#& 𝑞# 𝛼#,% = 𝑘%& 𝑞#
𝑊 B ∈ 𝑅 %A ×% = 𝒙&# 𝒙( = 𝒙&# 𝒙' = 𝒙&# 𝒙# = 𝒙&# 𝒙%

𝑊 C ∈ 𝑅 %A ×%
𝑘& 𝑣& 𝑘@ 𝑣@ 𝑘? 𝑣? 𝑞= 𝑘= 𝑣= 𝑘> 𝑣>
𝑊 E
∈ 𝑅 %D×%
𝑊, 𝑊- 𝑊, 𝑊- 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊, 𝑊-

Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Updated 𝒙(' = *
𝛼#,) *
𝑣& + 𝛼#,( *
𝑣@ + 𝛼#,' *
𝑣? + 𝛼#,# *
𝑣= + 𝛼#,% 𝑣>
feature

* * * * *
𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%

Softmax
= 𝑘𝒙)&&
𝛼#,) = #𝑞𝒙#) 𝛼#,( = 𝑘(& 𝑞# 𝛼#,' = 𝑘'& 𝑞# 𝛼#,# = 𝑘#& 𝑞# 𝛼#,% = 𝑘%& 𝑞#
𝑊 B ∈ 𝑅 %A ×% = 𝒙&# 𝒙( = 𝒙&# 𝒙' = 𝒙&# 𝒙# = 𝒙&# 𝒙%

𝑊 C ∈ 𝑅 %A ×%
𝑘& 𝑣& 𝑘@ 𝑣@ 𝑘? 𝑣? 𝑞= 𝑘= 𝑣= 𝑘> 𝑣>
𝑊 E
∈ 𝑅 %D×%
𝑊, 𝑊- 𝑊, 𝑊- 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊, 𝑊-

Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Updated 𝒙(' = 𝑊" *
𝛼#,) *
𝑣& + 𝛼#,( *
𝑣@ + 𝛼#,' * *
0𝑣? + 𝛼#,#𝑣= + 𝛼#,% 𝑣>
feature

= *
= 0𝛼#,) 𝑊" 𝑊- 𝒙.
.

* * * * *
𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%

𝑊 F ∈ 𝑅 %×%D
Softmax
𝛼#,) = 𝑘)& 𝑞# 𝛼#,( = 𝑘(& 𝑞# 𝛼#,' = 𝑘'& 𝑞# 𝛼#,# = 𝑘#& 𝑞# 𝛼#,% = 𝑘%& 𝑞#
𝑊 B ∈ 𝑅 %A ×% = 𝒙&# 𝒙) = 𝒙&# 𝒙( = 𝒙&# 𝒙' = 𝒙&# 𝒙# = 𝒙&# 𝒙%

𝑊 C ∈ 𝑅 %A ×%
𝑘& 𝑣& 𝑘@ 𝑣@ 𝑘? 𝑣? 𝑞= 𝑘= 𝑣= 𝑘> 𝑣>
𝑊 E
∈ 𝑅 %D×%
𝑊, 𝑊- 𝑊, 𝑊- 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊, 𝑊-

Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Updated 𝒙(' = 𝑊" *
𝛼#,) *
𝑣& + 𝛼#,( *
𝑣@ + 𝛼#,' * *
0𝑣? + 𝛼#,#𝑣= + 𝛼#,% 𝑣>
feature

= *
= 0𝛼#,) 𝑊" 𝑊- 𝒙.
.

* * * * *
𝛼#,) 𝛼#,( 𝛼#,' 𝛼#,# 𝛼#,%

𝑊 F ∈ 𝑅 %×%D
Softmax
𝛼#,) = 𝑘)& 𝑞# 𝛼#,( = 𝑘(& 𝑞# 𝛼#,' = 𝑘'& 𝑞# 𝛼#,# = 𝑘#& 𝑞# 𝛼#,% = 𝑘%& 𝑞#
𝑊 B ∈ 𝑅 %A ×% = 𝒙&# 𝒙) = 𝒙&# 𝒙( = 𝒙&# 𝒙' = 𝒙&# 𝒙# = 𝒙&# 𝒙%

𝑊 C ∈ 𝑅 %A ×%
𝑘& 𝑣& 𝑘@ 𝑣@ 𝑘? 𝑣? 𝑞= 𝑘= 𝑣= 𝑘> 𝑣>
𝑊 E
∈ 𝑅 %D×%
𝑊, 𝑊- 𝑊, 𝑊- 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊, 𝑊-

Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
* * * * *
𝛼(,) 𝛼(,( 𝛼(,' 𝛼(,# 𝛼(,%

𝑊 F ∈ 𝑅 %×%D
Softmax
𝛼#,) = 𝑘)& 𝑞# 𝛼#,( = 𝑘(& 𝑞# 𝛼#,' = 𝑘'& 𝑞# 𝛼#,# = 𝑘#& 𝑞# 𝛼#,% = 𝑘%& 𝑞#
𝑊 B ∈ 𝑅 %A ×% = 𝒙&# 𝒙) = 𝒙&# 𝒙( = 𝒙&# 𝒙' = 𝒙&# 𝒙# = 𝒙&# 𝒙%

𝑊 C ∈ 𝑅 %A ×%
𝑘& 𝑣& 𝑞@ 𝑘@ 𝑣@ 𝑘? 𝑣? 𝑘= 𝑣= 𝑘> 𝑣>
𝑊 E
∈ 𝑅 %D×%
𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊, 𝑊- 𝑊, 𝑊- 𝑊, 𝑊-

Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
* * * * *
𝛼(,) 𝛼(,( 𝛼(,' 𝛼(,# 𝛼(,%

𝑊 F ∈ 𝑅 %×%D
Softmax
𝛼#,) = 𝑘)& 𝑞# 𝛼#,( = 𝑘(& 𝑞# 𝛼#,' = 𝑘'& 𝑞# 𝛼#,# = 𝑘#& 𝑞# 𝛼#,% = 𝑘%& 𝑞#
𝑊 B ∈ 𝑅 %A ×% = 𝒙&# 𝒙) = 𝒙&# 𝒙( = 𝒙&# 𝒙' = 𝒙&# 𝒙# = 𝒙&# 𝒙%

𝑊 C ∈ 𝑅 %A ×%
𝑘& 𝑣& 𝑞@ 𝑘@ 𝑣@ 𝑘? 𝑣? 𝑘= 𝑣= 𝑘> 𝑣>
𝑊 E
∈ 𝑅 %D×%
𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊, 𝑊- 𝑊, 𝑊- 𝑊, 𝑊-

Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Updated 𝒙() = 𝑊 F *
𝛼(,) *
𝑣& + 𝛼(,( *
𝑣@ + 𝛼(,' * *
0𝑣? + 𝛼(,#𝑣= + 𝛼(,% 𝑣>
feature

* * * * *
𝛼(,) 𝛼(,( 𝛼(,' 𝛼(,# 𝛼(,%

𝑊 F ∈ 𝑅 %×%D
Softmax
𝛼#,) = 𝑘)& 𝑞# 𝛼#,( = 𝑘(& 𝑞# 𝛼#,' = 𝑘'& 𝑞# 𝛼#,# = 𝑘#& 𝑞# 𝛼#,% = 𝑘%& 𝑞#
𝑊 B ∈ 𝑅 %A ×% = 𝒙&# 𝒙) = 𝒙&# 𝒙( = 𝒙&# 𝒙' = 𝒙&# 𝒙# = 𝒙&# 𝒙%

𝑊 C ∈ 𝑅 %A ×%
𝑘& 𝑣& 𝑞@ 𝑘@ 𝑣@ 𝑘? 𝑣? 𝑘= 𝑣= 𝑘> 𝑣>
𝑊 E
∈ 𝑅 %D×%
𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊, 𝑊- 𝑊, 𝑊- 𝑊, 𝑊-

Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Updated 𝒙() = 𝑊 F *
𝛼(,) *
𝑣& + 𝛼(,( *
𝑣@ + 𝛼(,' * *
0𝑣? + 𝛼(,#𝑣= + 𝛼(,% 𝑣>
feature

* * * * *
𝛼(,) 𝛼(,( 𝛼(,' 𝛼(,# 𝛼(,%

𝑊 F ∈ 𝑅 %×%D
Softmax
𝛼#,) = 𝑘)& 𝑞# 𝛼#,( = 𝑘(& 𝑞# 𝛼#,' = 𝑘'& 𝑞# 𝛼#,# = 𝑘#& 𝑞# 𝛼#,% = 𝑘%& 𝑞#
𝑊 B ∈ 𝑅 %A ×% = 𝒙&# 𝒙) = 𝒙&# 𝒙( = 𝒙&# 𝒙' = 𝒙&# 𝒙# = 𝒙&# 𝒙%

𝑊 C ∈ 𝑅 %A ×%
𝑘& 𝑣& 𝑞@ 𝑘@ 𝑣@ 𝑘? 𝑣? 𝑘= 𝑣= 𝑘> 𝑣>
𝑊 E
∈ 𝑅 %D×%
𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊, 𝑊- 𝑊, 𝑊- 𝑊, 𝑊-

Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
𝛼),) = 𝑘&G 𝑞&

𝛼),( = 𝑘@G 𝑞&

𝛼),' = 𝑘?G 𝑞&

𝛼),# = 𝑘=G 𝑞&

𝛼),% = 𝑘>G 𝑞&

𝑞& 𝑘& 𝑣& 𝑞@ 𝑘@ 𝑣@ 𝑞? 𝑘? 𝑣? 𝑞= 𝑘= 𝑣= 𝑞> 𝑘> 𝑣>

𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊-

Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
1
𝛼),) 𝛼(,) 𝛼',) 𝛼#,) 𝛼%,) 𝑘1&G
𝛼),( 𝛼(,( 𝛼',( 𝛼#,( 𝛼%,( 𝑘1@G
𝛼),' 𝛼(,' 𝛼',' 𝛼#,' 𝛼%,' = 𝑘1?G 𝑞& 𝑞@ 𝑞? 𝑞= 𝑞>
𝛼),# 𝛼(,# 𝛼',# 𝛼#,# 𝛼%,# 𝑘1=G
𝛼),% 𝛼(,% 𝛼',% 𝛼#,% 𝛼%,% 𝑘1>G
1

𝑞& 𝑘& 𝑣& 𝑞@ 𝑘@ 𝑣@ 𝑞? 𝑘? 𝑣? 𝑞= 𝑘= 𝑣= 𝑞> 𝑘> 𝑣>

𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊-

Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
1
1 1
𝑑/ 𝛼),) 𝛼(,) 𝛼1',) 𝛼#,) 𝛼%,) 𝑘1&G
=
𝛼),( 𝛼(,( 𝛼1',( 𝛼#,( 𝛼%,( 𝑘1@G
𝛼),' 𝛼(,' 𝛼1',' 𝛼#,' 𝛼%,' = 𝑘1?G 𝑞&1𝑞@ 𝑞? 𝑞= 𝑞>
1
𝛼),# 𝛼(,# 𝛼1',# 𝛼#,# 𝛼%,# 𝑘1=G 1
𝛼),% 𝛼(,% 𝛼1',% 𝛼#,% 𝛼%,% 𝑘1>G
1 1
Softmax
* * 1* * *
𝛼),) 𝛼(,) 𝛼1',) 𝛼#,) 𝛼%,)
* * * * *
𝛼),( 𝛼(,( 𝛼1',( 𝛼#,( 𝛼%,(
* * * * *
𝛼),' 𝛼(,' 𝛼1',' 𝛼#,' 𝛼%,'
* * * * *
𝛼),# 𝛼(,# 𝛼1',# 𝛼#,# 𝛼%,#
* * * * *
𝛼),% 𝛼(,% 𝛼1',% 𝛼#,% 𝛼%,%
1

𝑞& 𝑘& 𝑣& 𝑞@ 𝑘@ 𝑣@ 𝑞? 𝑘? 𝑣? 𝑞= 𝑘= 𝑣= 𝑞> 𝑘> 𝑣>

𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊-

Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
1
1 1
𝑑/ 𝛼),) 𝛼(,) 𝛼1',) 𝛼#,) 𝛼%,) 𝑘1&G
=
𝛼),( 𝛼(,( 𝛼1',( 𝛼#,( 𝛼%,( 𝑘1@G
𝛼),' 𝛼(,' 𝛼1',' 𝛼#,' 𝛼%,' = 𝑘1?G 𝑞&1𝑞@ 𝑞? 𝑞= 𝑞>
1
𝛼),# 𝛼(,# 𝛼1',# 𝛼#,# 𝛼%,# 𝑘1=G 1
𝛼),% 𝛼(,% 𝛼1',% 𝛼#,% 𝛼%,% 𝑘1>G
1 1
Softmax
* * 1* * * 𝑘 = 𝑘 ! , 𝑘 " , ⋯ , 𝑘 (0 ) 𝐸 𝑘 * = 𝐸 𝑞* = 0
𝛼),) 𝛼(,) 𝛼1',) 𝛼#,) 𝛼%,)
*
𝛼),( *
𝛼(,( *
𝛼1',( *
𝛼#,( *
𝛼%,( 𝑞 = 𝑞! , 𝑞" , ⋯ , 𝑞(0 )
Var 𝑘 * = Var 𝑞* = 1
* * * * *
𝛼),' 𝛼(,' 𝛼1',' 𝛼#,' 𝛼%,' (0
* * * * *
𝛼),# 𝛼(,# 𝛼1',# 𝛼#,# 𝛼%,#
*
𝛼),% *
𝛼(,% *
𝛼1',% *
𝛼#,% *
𝛼%,% 𝑘 ) 𝑞 = / 𝑘 * 𝑞* Var 𝑘 ) 𝑞 = 𝑑+
1 *,!

𝑞& 𝑘& 𝑣& 𝑞@ 𝑘@ 𝑣@ 𝑞? 𝑘? 𝑣? 𝑞= 𝑘= 𝑣= 𝑞> 𝑘> 𝑣>

𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊-

Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
1
1 1
𝑑/ 𝛼),) 𝛼(,) 𝛼1',) 𝛼#,) 𝛼%,) 𝑘1&G
=
𝛼),( 𝛼(,( 𝛼1',( 𝛼#,( 𝛼%,( 𝑘1@G
𝛼),'
𝛼),#
𝛼(,'
𝛼(,#
𝐴
𝛼1','
𝛼1',#
𝛼#,'
𝛼#,#
𝛼%,' =
𝛼%,#
𝐾 𝑘1?G! 𝑞&1𝑞@ 𝑞? 𝑞= 𝑞>
1
𝑄
𝑘1=G 1
𝛼),% 𝛼(,% 𝛼1',% 𝛼#,% 𝛼%,% 𝑘1>G
1 1
Softmax
* * 1* * *
𝛼),) 𝛼(,) 𝛼1',) 𝛼#,) 𝛼%,)
* * * * *
𝛼),( 𝛼(,( 𝛼1',( 𝛼#,( 𝛼%,(
*"
*
𝛼),'
*
𝛼),#
*
𝛼(,'
*
𝛼(,#
𝐴
𝛼1','
*
𝛼1',#
*
𝛼#,'
*
𝛼#,#
*
𝛼%,'
*
𝛼%,#
* * * * *
𝛼),% 𝛼(,% 𝛼1',% 𝛼#,% 𝛼%,%
1

𝑞& 𝑘& 𝑣& 𝑞@ 𝑘@ 𝑣@ 𝑞? 𝑘? 𝑣? 𝑞= 𝑘= 𝑣= 𝑞> 𝑘> 𝑣>

𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊-

Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
1
1 1
𝑑/ 𝛼),) 𝛼(,) 𝛼1',) 𝛼#,) 𝛼%,) 𝑘1&G
=
𝛼),( 𝛼(,( 𝛼1',( 𝛼#,( 𝛼%,( 𝑘1@G
𝛼),'
𝛼),#
𝛼(,'
𝛼(,#
𝐴
𝛼1','
𝛼1',#
𝛼#,'
𝛼#,#
𝛼%,' =
𝛼%,#
𝐾 𝑘1?G! 𝑞&1𝑞@ 𝑞? 𝑞= 𝑞>
1
𝑄
𝑘1=G 1
𝛼),% 𝛼(,% 𝛼1',% 𝛼#,% 𝛼%,% 𝑘1>G
1 1
Softmax
* * 1* * * * * 1* * *
𝛼),) 𝛼(,) 𝛼1',) 𝛼#,) 𝛼%,) 𝛼),) 𝛼(,) 𝛼',) 𝛼#,) 𝛼%,)
* * * * * = = * * 1* * *
𝛼),( 𝛼(,( 𝛼1',( 𝛼#,( 𝛼%,( 𝛼),( 𝛼(,( 𝛼1',( 𝛼#,( 𝛼%,(
*" *"
*
𝛼),'
*
𝛼),#
*
𝛼(,'
*
𝛼(,#
𝐴
𝛼1','
*
𝛼1',#
*
𝛼#,'
*
𝛼#,#
*
𝛼%,'
*
𝛼%,#
1
𝒙&( 1𝒙(@ 𝒙(? 𝒙(= 𝒙(> = 𝑊 F 𝑞& 1𝑣@ 𝑣? 𝑣= 𝑣>
𝑣
1
*
𝛼),'
*
𝛼),#
*
*
𝐴
𝛼(,'
𝛼(,#
𝛼1','
*
𝛼1',#
*
𝛼#,'
*
𝛼#,#
*
𝛼%,'
*
𝛼%,#
* * * * * 1. 1 * * * * *
𝛼),% 𝛼(,% 𝛼1',% 𝛼#,% 𝛼%,% 𝛼),% 𝛼(,% 𝛼1',% 𝛼#,% 𝛼%,%
Output features
1 1

𝑞& 𝑘& 𝑣& 𝑞@ 𝑘@ 𝑣@ 𝑞? 𝑘? 𝑣? 𝑞= 𝑘= 𝑣= 𝑞> 𝑘> 𝑣>

𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊-

Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
1
𝑑/ 𝛼),) 𝛼(,)
1
𝛼1',) 𝛼#,) 𝛼%,)
1
𝑘1&G
=
𝑄 = 𝑊+ 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%
𝛼),( 𝛼(,( 𝛼1',( 𝛼#,( 𝛼%,( 𝑘1@G
𝛼),'
𝛼),#
𝛼(,'
𝛼(,#
𝐴
𝛼1','
𝛼1',#
𝛼#,'
𝛼#,#
𝛼%,' =
𝛼%,#
𝐾 𝑘1?G! 𝑞&1𝑞@ 𝑞? 𝑞= 𝑞>
1
𝑄 𝐾 = 𝑊, 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%
𝑘1=G 1
1
𝛼),% 𝛼(,% 𝛼1',% 𝛼#,% 𝛼%,%
1
𝑘1>G 𝑉 = 𝑊- 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%
Softmax
* * 1* * * * * 1* * *
𝛼),) 𝛼(,) 𝛼1',) 𝛼#,) 𝛼%,) 𝛼),) 𝛼(,) 𝛼1',) 𝛼#,) 𝛼%,)
*
𝛼),( *
𝛼(,( *
𝛼1',( *
𝛼#,( *
𝛼%,( = = *
𝛼),( *
𝛼(,( *
𝛼1',( *
𝛼#,( *
𝛼%,(
*" *"
*
𝛼),'
*
𝛼),#
*
𝛼(,'
*
𝛼(,#
𝐴
𝛼1','
*
𝛼1',#
*
𝛼#,'
*
𝛼#,#
*
𝛼%,'
*
𝛼%,#
1
𝒙&( 1𝒙(@ 𝒙(? 𝒙(= 𝒙(> = 𝑊F 𝑞& 1𝑣@ 𝑣? 𝑣= 𝑣>
𝑣
1 𝑉
*
𝛼),'
*
𝛼),#
*
*
𝐴
𝛼(,'
𝛼(,#
𝛼1','
*
𝛼1',#
*
𝛼#,'
*
𝛼#,#
*
𝛼%,'
*
𝛼%,#
* * * * *
1. 1 * * * * *
𝛼),% 𝛼(,% 𝛼1',% 𝛼#,% 𝛼%,% 𝛼),% 𝛼(,% 𝛼1',% 𝛼#,% 𝛼%,%
Output features
1 1

𝑞& 𝑘& 𝑣& 𝑞@ 𝑘@ 𝑣@ 𝑞? 𝑘? 𝑣? 𝑞= 𝑘= 𝑣= 𝑞> 𝑘> 𝑣>

𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊-

Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Single-head attention 𝑄 = 𝑊+ 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%

Attention 𝑄, 𝐾, 𝑉 = 𝑉 softmax
𝐾*𝑄 𝐾 = 𝑊, 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%
𝑑+
𝑉 = 𝑊- 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%

𝑞& 𝑘& 𝑣& 𝑞@ 𝑘@ 𝑣@ 𝑞? 𝑘? 𝑣? 𝑞= 𝑘= 𝑣= 𝑞> 𝑘> 𝑣>

𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊- 𝑊+ 𝑊, 𝑊-

Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Analogy for Q, K, V
► Library system
► Imagine you're looking for information on a specific topic (query)
► Each book in the library has a summary (key) that helps identify if it contains the information
you're looking for
► Once you find a match between your query and a summary, you access the book to get the
detailed information (value) you need
► Here, in Attention, we do a “soft match” across multiple values, e.g. get info from multiple
books (“book 1 is most relevant, then book 2, then book 3, etc.”)

𝐾*𝑄
Attention 𝑄, 𝐾, 𝑉 = 𝑉 softmax
𝑑+
Single-head attention 𝑄 = 𝑊+ 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%

Attention 𝑄, 𝐾, 𝑉 = 𝑉 softmax
𝐾*𝑄 𝐾 = 𝑊, 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%
𝑑+
𝑉 = 𝑊- 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%

𝑊&' 𝑊&( 𝑊&)


'
𝑊! 𝑊!( 𝑊!) ⋯
'
𝑊*+! (
𝑊*+! )
𝑊*+!

𝑊H
B
∈ 𝑅 %A×%

𝑊HC ∈ 𝑅 %A ×%

𝑊HE ∈ 𝑅 %D ×%
Single-head attention 𝑄 = 𝑊+ 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%

Attention 𝑄, 𝐾, 𝑉 = 𝑉 softmax
𝐾*𝑄 𝐾 = 𝑊, 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%
𝑑+
𝑉 = 𝑊- 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%

𝑊&' 𝑊&( 𝑊&)


'
𝑊! 𝑊!( 𝑊!) ⋯
'
𝑊*+! (
𝑊*+! )
𝑊*+!

𝑊H
B
∈ 𝑅 %A×% 𝑊&' 𝑄 𝑊&( 𝐾 𝑊&) 𝑉

𝑊HC ∈ 𝑅 %A ×% Attention

𝑊HE ∈ 𝑅 %D ×% Head1 ∈ 𝑅2+ ×4


Single-head attention
𝐾*𝑄
Attention 𝑄, 𝐾, 𝑉 = 𝑉 softmax
𝑑+
𝑋= 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%

𝑊&' 𝑊&( 𝑊&)


'
𝑊! 𝑊!( 𝑊!) ⋯
'
𝑊*+! (
𝑊*+! )
𝑊*+!

𝑊H
B
∈ 𝑅 %A×% 𝑊&' 𝑋 𝑊&( 𝑋 𝑊&) 𝑋

𝑊HC ∈ 𝑅 %A ×% Attention

𝑊HE ∈ 𝑅 %D ×% Head1 ∈ 𝑅2+ ×4


Single-head attention 𝑄 = 𝑊+ 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%

Attention 𝑄, 𝐾, 𝑉 = 𝑉 softmax
𝐾*𝑄 𝐾 = 𝑊, 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%
𝑑+
𝑉 = 𝑊- 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%

𝑊&' 𝑊&( 𝑊&)


'
𝑊! 𝑊!( 𝑊!) ⋯
'
𝑊*+! (
𝑊*+! )
𝑊*+!

𝑊H
B
∈ 𝑅 %A×% 𝑊&' 𝑄 𝑊&( 𝐾 𝑊&) 𝑉 𝑊!' 𝑄 𝑊!( 𝐾 𝑊!) 𝑉

𝑊HC ∈ 𝑅 %A ×% Attention Attention

Head1 ∈ 𝑅2+ ×4 Head) ∈ 𝑅2+ ×4 ⋯ Head56) ∈ 𝑅2+ ×4


𝑊HE ∈ 𝑅 %D ×%
Single-head attention 𝑄 = 𝑊+ 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%

Attention 𝑄, 𝐾, 𝑉 = 𝑉 softmax
𝐾*𝑄 𝐾 = 𝑊, 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%
𝑑+
𝑉 = 𝑊- 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%

𝑊&' 𝑊&( 𝑊&)


'
𝑊! 𝑊!( 𝑊!) ⋯
'
𝑊*+! (
𝑊*+! )
𝑊*+!

𝑊H
B
∈ 𝑅 %A×% 𝑊&' 𝑄 𝑊&( 𝐾 𝑊&) 𝑉 𝑊!' 𝑄 𝑊!( 𝐾 𝑊!) 𝑉

𝑊HC ∈ 𝑅 %A ×% Attention Attention

Head1 ∈ 𝑅2+ ×4 Head) ∈ 𝑅2+ ×4 ⋯ Head56) ∈ 𝑅2+ ×4


𝑊HE ∈ 𝑅 %D ×%
Single-head attention 𝑄 = 𝑊+ 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%

Attention 𝑄, 𝐾, 𝑉 = 𝑉 softmax
𝐾*𝑄 𝐾 = 𝑊, 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%
𝑑+
Multi-head attention 𝑉 = 𝑊- 𝒙! 𝒙" 𝒙# 𝒙$ 𝒙%

𝑊 F ∈ 𝑅 %×,%D 𝑊&' 𝑊&( 𝑊&)


'
𝑊! 𝑊!( 𝑊!) ⋯
'
𝑊*+! (
𝑊*+! )
𝑊*+!

𝑊H
B
∈ 𝑅 %A×% 𝑊&' 𝑄 𝑊&( 𝐾 𝑊&) 𝑉 𝑊!' 𝑄 𝑊!( 𝐾 𝑊!) 𝑉

𝑊HC ∈ 𝑅 %A ×% Attention Attention

Head1 ∈ 𝑅2+ ×4 Head) ∈ 𝑅2+ ×4 ⋯ Head56) ∈ 𝑅2+ ×4


𝑊HE ∈ 𝑅 %D ×%
Head1
Head)
MultiHeadedAttention 𝑄, 𝐾, 𝑉 = 𝑊9 ⋮
Head56)
Feed Forward Network (FFN)
𝐹𝐹𝑁 𝒙
= 𝑾4 ReLU 𝑾0 𝒙 + 𝒃0 + 𝒃4

Feed Feed Feed Feed Feed


Forward Forward Forward Forward Forward
Encoder #1
Multi-head Self-Attention

Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Feed Forward Network (FFN)
ReLU
𝐹𝐹𝑁 𝒙



= 𝑾4 ReLU 𝑾0 𝒙 + 𝒃0 + 𝒃4
𝑾( 𝑾) 𝒙

Feed Feed Feed Feed Feed


Forward Forward Forward Forward Forward
Encoder #1
Multi-head Self-Attention

Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Feed Feed Feed Feed Feed
Forward Forward Forward Forward Forward
Encoder #1
Multi-head Self-Attention

Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Feed Feed Feed Feed Feed
Forward Forward Forward Forward Forward
Encoder #1
Multi-head Self-Attention

Embedded
Tokens 𝒙7 𝒙5 𝒙6 𝒙0 𝒙4
Tokens watch an apple I bought
Positional encoding

Feed Feed Feed Feed Feed


Forward Forward Forward Forward Forward
Encoder #1
Multi-head Self-Attention

Embedded
Tokens 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Positional encoding
Position 𝑘 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15
2' 0 0 0 0 0 0 0 0 1 1 1 1 1 1 1 1 Slow oscillating
0 0 0 0 1 1 1 1 0 0 0 0 1 1 1 1
2(
Dimension
0 0 1 1 0 0 1 1 0 0 1 1 0 0 1 1
2)
0 1 0 1 0 1 0 1 0 1 0 1 0 1 0 1
21 Fast oscillating

Positional
𝑑
embedding

Embedded
Tokens 𝑑 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Positional encoding
sin 𝑤1 𝑘
Position 𝑘 cos(𝑤1 𝑘) Fast oscillating
sin 𝑤) 𝑘
Angular frequency cos 𝑤) 𝑘
𝑑 ⋮
𝑤H = 𝑁 %@H/J

sin 𝑤2 𝑘
( 6) Slow oscillating
𝑁 = 100,000 cos 𝑤2 𝑘
6)
( Image: [Link]

Positional
𝑑
embedding

Embedded
Tokens 𝑑 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Normalized Range
Positional encoding
sin 𝑤1 𝑘 Unique identifier, unlimited length
Position 𝑘 cos(𝑤1 𝑘) Relative positions as linear transform
sin 𝑤) 𝑘
Angular frequency cos 𝑤) 𝑘
𝑑 sin(𝑤. (𝑘 + ∆𝑘) sin 𝑤. 𝑘 cos 𝑤. ∆𝑘 + cos 𝑤. 𝑘 sin 𝑤. ∆𝑘
𝑤H = 𝑁 %@H/J ⋮ =
cos(𝑤. (𝑘 + ∆𝑘) cos 𝑤. 𝑘 cos 𝑤. ∆𝑘 − sin 𝑤. 𝑘 sin 𝑤. ∆𝑘
co𝑠

sin 𝑤2 𝑘
( 6)
𝑁 = 100,000 cos 𝑤2 𝑘
6)
(

Positional
𝑑
embedding

Embedded
Tokens 𝑑 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Normalized Range
Positional encoding
sin 𝑤1 𝑘 Unique identifier, unlimited length
Position 𝑘 cos(𝑤1 𝑘) Relative positions as linear transform
sin 𝑤) 𝑘
Angular frequency cos 𝑤) 𝑘
𝑑 sin(𝑤. (𝑘 + ∆𝑘) sin 𝑤. 𝑘 cos 𝑤. ∆𝑘 + cos 𝑤. 𝑘 sin 𝑤. ∆𝑘
𝑤H = 𝑁 %@H/J ⋮ =
⋮ cos(𝑤. (𝑘 + ∆𝑘) cos 𝑤. 𝑘 cos 𝑤. ∆𝑘 − sin 𝑤. 𝑘 sin 𝑤. ∆𝑘
sin 𝑤2 𝑘 cos 𝑤. ∆𝑘 sin 𝑤. ∆𝑘 sin 𝑤
sin 𝑤..𝑘𝑘
( 6) =
𝑁 = 100,000 cos 𝑤2 𝑘 − sin 𝑤. ∆𝑘 cos 𝑤. ∆𝑘 co𝑠
co𝑠 𝑤𝑤..𝑘𝑘
6)
(

Positional
𝑃IJ∆I = 𝑀𝑃I
embedding
𝑑 𝑷0 𝑷4 𝑷5 𝑷6 𝑷7

Embedded
Tokens 𝑑 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Embedded
Tokens 𝑑 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Position 𝑘=1 𝑘=2 𝑘=3 𝑘=4 𝑘=5

sin 𝑤1 𝑘
𝑑 𝑷0 𝑷4 𝑷5 𝑷6 𝑷7
cos(𝑤1 𝑘)
sin 𝑤) 𝑘
cos 𝑤) 𝑘
𝑑 ⋮ 𝒙9 𝒙9 𝒙9

sin 𝑤2 𝑘
6)
concat MLP +
(
cos 𝑤2
(
6)
𝑘 𝑷9 𝑷9 𝑷9
Feed Feed Feed Feed Feed
Forward Forward Forward Forward Forward
Encoder #1
Multi-head Self-Attention

Embedded
Tokens 𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝒙7 𝑷7
Tokens I bought an apple watch
Positional encoding
sin 𝑤1 𝑘
Position 𝑘 cos(𝑤1 𝑘)
Sinusoidal positional encoding
sin 𝑤) 𝑘
Angular frequency cos 𝑤) 𝑘 Relative positional encoding
𝑑 ⋮
𝑤H = 𝑁 %@H/J
⋮ KERPLE RoPE CoPE
sin 𝑤2 𝑘
( 6)
𝑁 = 100,000 cos 𝑤2 𝑘
6)
NoPE YaRN FIRE
(

Positional
embedding
𝑑 𝑷0 𝑷4 𝑷5 𝑷6 𝑷7

Embedded
Tokens 𝑑 𝒙0 𝒙4 𝒙5 𝒙6 𝒙7
Tokens I bought an apple watch
Feed Feed Feed Feed Feed
Forward Forward Forward Forward Forward
Encoder #2
Multi-head Self-Attention

Feed Feed Feed Feed Feed


Forward Forward Forward Forward Forward
Encoder #1
Multi-head Self-Attention

Embedded
Tokens 𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝒙7 𝑷7
Tokens I bought an apple watch
Feed Feed Feed Feed Feed
Forward Forward Forward Forward Forward
Encoder #2
Multi-head Self-Attention

Feed Feed Feed Feed Feed


Forward Forward Forward Forward Forward
Encoder #1
Multi-head Self-Attention

Embedded
Tokens 𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝒙7 𝑷7
Tokens I bought an apple watch
Residual connection

Feed Feed Feed Feed Feed


Forward Forward Forward Forward Forward

Multi-head Self-Attention

Embedded
Tokens 𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝒙7 𝑷7
Tokens I bought an apple watch
Residual connection

Feed Feed Feed Feed Feed


Forward Forward Forward Forward Forward

Multi-head Self-Attention

Embedded
Tokens 𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝒙7 𝑷7
Tokens I bought an apple watch
Residual connection
Layer normalization
LayerNorm LayerNorm LayerNorm LayerNorm LayerNorm

Feed Feed Feed Feed Feed


Forward Forward Forward Forward Forward

LayerNorm LayerNorm LayerNorm LayerNorm LayerNorm

Multi-head Self-Attention

Embedded
Tokens 𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝒙7 𝑷7
Tokens I bought an apple watch
Residual connection
Layer normalization
LayerNorm LayerNorm LayerNorm LayerNorm LayerNorm

LayerNorm 𝒙 = Feed Feed Feed Feed Feed


Forward Forward Forward Forward Forward
𝒙 − mean 𝒙
𝛾 +𝛽 LayerNorm LayerNorm LayerNorm LayerNorm LayerNorm
Variance 𝐱 + 𝜖
Multi-head Self-Attention
𝛾, 𝛽 ∈ 𝑅
Learnable parameters

Embedded
Tokens 𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝒙7 𝑷7
Tokens I bought an apple watch
Residual connection
Layer normalization

LayerNorm 𝒙 = Feed Feed Feed Feed Feed


Forward Forward Forward Forward Forward
𝒙 − mean 𝒙 LayerNorm LayerNorm LayerNorm LayerNorm LayerNorm
𝛾 +𝛽
Variance 𝐱 + 𝜖
Multi-head Self-Attention
𝛾, 𝛽 ∈ 𝑅
LayerNorm LayerNorm LayerNorm LayerNorm LayerNorm
Learnable parameters

Embedded
Tokens 𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝒙7 𝑷7

[Xiong et al. 2020]


Tokens I bought an apple watch
On Layer Normalization in the Transformer Architecture
Feed Feed Feed Feed Feed
Forward Forward Forward Forward Forward
LayerNorm LayerNorm LayerNorm LayerNorm LayerNorm

Multi-head Self-Attention
LayerNorm LayerNorm LayerNorm LayerNorm LayerNorm

Embedded
Tokens 𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝒙7 𝑷7
Tokens I bought an apple watch
Encoder #6

Encoder #5

Encoder #4

Encoder #3

Encoder #2

Encoder #1

Embedded
Tokens 𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝒙7 𝑷7
Tokens I bought an apple watch
Encoder #6

Encoder #5

Encoder #4

Encoder #3

Encoder #2

Encoder #1

Embedded
Tokens 𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6
Tokens How are you ?
Encoder #6

Encoder #1

Embedded
Tokens 𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6
Tokens How are you ?
你 好 吃 飽 嗎 ⋯ ? <end>

Softmax

𝒆0 𝒆4 𝒆5 𝒆6 Linear

Encoder #𝐿 Decoder #𝐿

⋮ ⋮

Encoder #1 Decoder #1

𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝑑 𝒛0 𝑷0

How are you ? <start>


你 好 吃 飽 嗎 ⋯ ? <end>

Softmax

𝒆0 𝒆4 𝒆5 𝒆6 Linear

Encoder #𝐿 Decoder #𝐿

⋮ ⋮

Encoder #1 Decoder #1

𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝑑 𝒛0 𝑷0 𝒛4 𝑷4

How are you ? <start> 你


你 好 吃 飽 嗎 ⋯ ? <end>

Softmax

𝒛0 𝒛4 𝒛5 𝒛6 Linear

Encoder #𝐿 Decoder #𝐿

⋮ ⋮

Encoder #1 Decoder #1

𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝑑 𝒛0 𝑷0 𝒛4 𝑷4

How are you ? <start> 你


你 好 吃 飽 嗎 ⋯ ? <end>

Softmax

𝒆0 𝒆4 𝒆5 𝒆6 Linear

Encoder #𝐿 Decoder #𝐿

⋮ ⋮

Encoder #1 Decoder #1

𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝑑 𝒛0 𝑷0 𝒛4 𝑷4 𝒛5 𝑷5

How are you ? <start> 你 好


你 好 吃 飽 嗎 ⋯ ? <end>

Softmax

𝒆0 𝒆4 𝒆5 𝒆6 Linear

Encoder #𝐿 Decoder #𝐿

⋮ ⋮

Encoder #1 Decoder #1

𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝑑 𝒛0 𝑷0 𝒛4 𝑷4 𝒛5 𝑷5

How are you ? <start> 你 好


你 好 吃 飽 嗎 ⋯ ? <end>

Softmax

𝒆0 𝒆4 𝒆5 𝒆6 Linear

Encoder #𝐿 Decoder #𝐿

⋮ ⋮

Encoder #1 Decoder #1

𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝑑 𝒛0 𝑷0 𝒛4 𝑷4 𝒛5 𝑷5 𝒛6 𝑷6

How are you ? <start> 你 好 嗎 ?<end>


𝒆0 𝒆4 𝒆5 𝒆6

Encoder #𝐿

Encoder #1 Decoder #1

𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝑑 𝒛0 𝑷0 𝒛4 𝑷4 𝒛5 𝑷5 𝒛6 𝑷6

How are you ? <start> 你 好 嗎 ?<end>


𝒆0 𝒆4 𝒆5 𝒆6

Encoder #𝐿

Multi-head Self-Attention
Encoder #1
Decoder #1

𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝑑 𝒛0 𝑷0 𝒛4 𝑷4 𝒛5 𝑷5 𝒛6 𝑷6

How are you ? <start> 你 好 嗎


𝒆0 𝒆4 𝒆5 𝒆6

Encoder #𝐿

Multi-head Self-Attention
Encoder #1
Decoder #1

𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝑑 𝒛0 𝑷0 𝒛4 𝑷4 𝒛5 𝑷5 𝒛6 𝑷6

are you ? <start> 你 好 嗎


𝒆0 𝒆4 𝒆5 𝒆6

Encoder #𝐿

Multi-head Self-Attention
Encoder #1
Decoder #1

𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝑑 𝒛0 𝑷0 𝒛4 𝑷4 𝒛5 𝑷5 𝒛6 𝑷6

are you ? <start> 你 好 嗎


𝒆0 𝒆4 𝒆5 𝒆6

Encoder #𝐿 𝐾 G𝑄
Attention 𝑄, 𝐾, 𝑉 = 𝑉 softmax
𝑑K

𝐾 G𝑄
MaskedAttention Encoder
𝑄, 𝐾, 𝑉 =#1
𝑉 softmax +𝑀 Multi-head Self-Attention
𝑑K
0 0 0 0 0 Decoder #1
−∞ 0 0 0 0

𝒙0= 𝑷0
𝑑𝑀
−∞ −∞
𝒙4 𝑷04 0𝒙5 0𝑷5 𝒙6 𝑷6 𝑑 𝒛0 𝑷0 𝒛4 𝑷4 𝒛5 𝑷5 𝒛6 𝑷6
−∞ −∞ −∞ 0 0

are
−∞ −∞ −∞ −∞ you
0 ? <start> 你 好 嗎
𝒆0 𝒆4 𝒆5 𝒆6

Encoder #𝐿 𝐾 G𝑄
Attention 𝑄, 𝐾, 𝑉 = 𝑉 softmax
𝑑K

𝐾 G𝑄
MaskedAttention Encoder
𝑄, 𝐾, 𝑉 =#1
𝑉 softmax +𝑀 Masked Multi-head Self-Attention
𝑑K
0 0 0 0 0 Decoder #1
−∞ 0 0 0 0

𝒙0= 𝑷0
𝑑𝑀
−∞ −∞
𝒙4 𝑷04 0𝒙5 0𝑷5 𝒙6 𝑷6 𝑑 𝒛0 𝑷0 𝒛4 𝑷4 𝒛5 𝑷5 𝒛6 𝑷6
−∞ −∞ −∞ 0 0

are
−∞ −∞ −∞ −∞ you
0 ? <start> 你 好 嗎
𝒆0 𝒆4 𝒆5 𝒆6


Masked Multi-head Self-Attention
你 好 Training
examples
Decoder #1
你 好 嗎
𝑑 𝒛0 𝑷0 𝒛4 𝑷4 𝒛5 𝑷5 𝒛6 𝑷6
你 好 嗎 ?
<start> 你 好 嗎
LayerNorm LayerNorm LayerNorm LayerNorm

Feed Feed Feed Feed


Forward Forward Forward Forward

𝐾-./01-2
LayerNorm LayerNorm LayerNorm LayerNorm
𝑉-./01-2
𝒆0 𝒆4 𝒆5 𝒆6
Encoder-decoder Attention

Encoder #𝐿 LayerNorm LayerNorm LayerNorm LayerNorm


Masked Multi-head Self-Attention
Encoder #1
Decoder #1

𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝑑 𝒛0 𝑷0 𝒛4 𝑷4 𝒛5 𝑷5 𝒛6 𝑷6

Ho are you ? <start> 你 好 嗎


* * * *
𝛼(,) 𝛼(,( 𝛼(,' 𝛼(,#

Softmax
𝑘) 𝑞( 𝑘(& 𝑞( 𝑘'& 𝑞( 𝑘#& 𝑞(

𝑞@

𝑊+
LayerNorm LayerNorm
𝑘& 𝑣& 𝑘@ 𝑣@ 𝑘? 𝑣? 𝑘= 𝑣=

𝑊, 𝑊- 𝑊, 𝑊- 𝑊, 𝑊- 𝑊, 𝑊- Masked Multi-head Self-Attention

𝒆0 𝒆4 𝒆5 𝒆6 Decoder #1

Encoder #𝐿 𝑑 𝒛0 𝑷0 𝒛4 𝑷4

<start> 你
* * * *
𝛼(,) 𝛼(,( 𝛼(,' 𝛼(,#
𝒛() = 𝑊" *
𝛼(,) *
𝑣& + 𝛼(,( *
𝑣@ + 𝛼(,' *
𝑣? + 𝛼(,#𝑣=
Softmax
𝑘) 𝑞( 𝑘(& 𝑞( 𝑘'& 𝑞( 𝑘#& 𝑞( Cross-attention
Encoder-decoder attention
𝑞@

𝑊+
LayerNorm LayerNorm
𝑘& 𝑣& 𝑘& 𝑣@ 𝑘& 𝑣? 𝑘& 𝑣=

𝑊, 𝑊- 𝑊, 𝑊- 𝑊, 𝑊- 𝑊, 𝑊- Masked Multi-head Self-Attention

𝒆0 𝒆4 𝒆5 𝒆6 Decoder #1

Encoder #𝐿 𝑑 𝒛0 𝑷0 𝒛4 𝑷4

(ignore the scaling 1/ 𝑑/ here for simplicity)
<start> 你
LayerNorm LayerNorm LayerNorm LayerNorm

Feed Feed Feed Feed


Forward Forward Forward Forward

𝐾-./01-2
LayerNorm LayerNorm LayerNorm LayerNorm
𝑉-./01-2
𝒆0 𝒆4 𝒆5 𝒆6
Encoder-decoder Attention

Encoder #𝐿 LayerNorm LayerNorm LayerNorm LayerNorm


Masked Multi-head Self-Attention
Encoder #1
Decoder #1

𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝑑 𝒛0 𝑷0 𝒛4 𝑷4 𝒛5 𝑷5 𝒛6 𝑷6

How are you ? <start> 你 好 嗎


𝐾-./01-2
𝑉-./01-2
𝒆0 𝒆4 𝒆5 𝒆6

Encoder #𝐿

Encoder #1 Decoder #1

𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝑑 𝒛0 𝑷0 𝒛4 𝑷4 𝒛5 𝑷5 𝒛6 𝑷6

How are you ? <start> 你 好 嗎


𝐾-./01-2 Softmax
𝑉-./01-2
𝒆0 𝒆4 𝒆5 𝒆6 Linear

Encoder #𝐿 Decoder #𝐿

⋮ ⋮

Encoder #1 Decoder #1

𝑑 𝒙0 𝑷0 𝒙4 𝑷4 𝒙5 𝑷5 𝒙6 𝑷6 𝑑 𝒛0 𝑷0 𝒛4 𝑷4 𝒛5 𝑷5 𝒛6 𝑷6

How are you ? <start> 你 好 嗎


Transformer & Multi-Head Attention

“Attention Is All You Need”


[Link]
Summary: Attention and Transformers
► Attention weights used to compute the context vector, which is a
weighted sum of the input at different positions
► Context vector is used to update the hidden state of the model,
which is used to generate the final output
► "Pay attention" to different parts of the input, depending on the task
at hand → more accurate and natural-sounding output, esp. when
working with longer inputs (e.g. paragraphs)
Ways Attention was used in the
*original* Transformer Architecture
• Encoder-decoder cross-attention
• Allow decoder layers to attend all parts of the latent representation produced by
the encoder
• Pull context from the encoder sequence over to the decoder

• Self-attention in the encoder


• Allow the model to attend to all positions in the previous encoder layer
• Embeds context about how elements in the sequence relate to one another

• Masked self-attention in the decoder


• Allow the model to attend to all positions in the previous decoder layer up to and
including the current position (during auto-regressive process)
• Prevent forward looking bias by stopping leftward information flow during training
• Also embed context about how elements in the sequence relate to one another
* A. Vaswani et al., “Attention is All you Need,” in Advances in Neural Information Processing Systems (NeurIPS), 2017.
Examples:
Attention is all you need, T5, BART.

Good for:
Machine translation, summarization. QA
(when input/target are sufficiently different)

𝐾-./01-2 Softmax
𝑉-./01-2
𝒆0 𝒆4 𝒆5 𝒆6 Linear

Encoder #𝐿 Decoder #𝐿

⋮ ⋮

Encoder #1 Decoder #1
Examples:
Attention is all you need, T5, BART.

Good for:
Machine translation, summarization. QA
(when input/target are sufficiently different)

𝒆0 𝒆4 𝒆5 𝒆6

Encoder #𝐿

Encoder #1
Examples:
BERT, RoBERTa, DeBERTa, X-BERT

Good for:
Classification, sequence tagging, sentiment
analysis
(Understand text, but not generate them)

𝒆0

Encoder #𝐿

Encoder #1
Examples:
GPT-X (OpenAI), PaLM (Google), LLaMA (Meta)
BLOOM (BigScience)

Good for:
Text generation, multi-round conversation

Softmax

Linear

Decoder #𝐿

Decoder #1
Softmax Softmax

𝒆0 𝒆4 𝒆5 𝒆6 Linear Linear

Encoder #𝐿 Decoder #𝐿 Decoder #𝐿



⋮ ⋮ ⋮

Encoder #1 Decoder #1 Decoder #1


<Input> <Target> <Input> <Target> <Input> ⋯

Prompting, in-context examples


Different parameters for encoder/decoder Shared parameters
LLaMA-2-Chat
Chat
GLM
LLaMA

2023
OPT-IML
BLOOMZ Galactica
Flan
T5

BLOOM
YaLM
Open-Source OPT
Tk

GPT-NeoX
ST-MoE
2022

GPT-J
GLM
GPT-Neo
Switch
2021

mT5 T0
DeBERTa
ELECTRA
2020
Distill T5
ALBERT BERT BART
XLNet
RoBERTa
ERNIE
GPT-2
Enc
ode
2019 r-D
eco
BERT der

Encoder-On GPT-1
ELMo ULMFiT ly Decoder-Only
2018

GloVe
FastText
Word2Vec

Image credit: [Link]


LLM LLaMA-2-Chat
Evolutionary Chat
GLM
Tree LLaMA

2023
OPT-IML
BLOOMZ Galactica
Flan
T5

Open-Source BLOOM
YaLM
OPT
Tk

GPT-NeoX
ST-MoE
2022

GPT-J
GLM
GPT-Neo
Switch
2021

mT5 T0
DeBERTa
ELECTRA
2020
Distill T5
ALBERT BERT BART
RoBERTa XLNet
ERNIE
GPT-2
Enc
ode
2019 r-D
BERT eco
der
Encoder-On GPT-1
ELMo ULMFiT ly Decoder-Only
2018

FastText GloVe
Word2Vec
16
8
Transformers vs. RNNs

Challenges with RNNs Transformers

● Long range dependencies ● Can model long-range


● Gradient vanishing and explosion dependencies
● Large # of training steps ● No gradient vanishing and explosion
● Sequential/recurrence → can’t parallelize ● Fewer training steps
● Complexity per layer: O(n*d2) ● Can parallelize computation!
● Complexity per layer: O(n2*d)

► When sequence length (n) << representation dimension (d), the complexity per layer
is lower for a Transformer model compared to RNN models ; NOT true for real-world LLMs
Differences in Attention Mechanism of RNN vs. Transformer

RNN with Attention


Feature Transformer
(Bahdanau et al. 2015)

Attention Type Additive (Bahdanau) Attention Scaled Dot Product Attention

Based on decoder hidden state and Based on dot-product of query and keys
Alignment
encoder hidden states (global attention)

Efficiency Processes sequences step-by-step Parallel processing of all positions

Weighted sum of encoder hidden Attends to all encoder positions for every
Context
states at each step output

Self-Attention Not used Self-attention in both encoder and decoder


Computational Dependencies for Recurrence vs. Attention

Transformer Advantages:
● # unparallelizable operations does not
RNN-Based Encoder-Decoder
increase with sequence length.
Model with Attention
● Each word interacts with each other, so
maximum interaction distance is O(1).

Transformer-Based
Encoder-Decoder Model

You might also like