0% found this document useful (0 votes)
9 views28 pages

Evolution of ChatGPT and GPT Models

ChatGPT is an AI language model developed by OpenAI, launched in November 2022, capable of understanding and generating human-like text. It is built on the Transformer architecture and has evolved through several iterations (GPT-1 to GPT-5), each improving in scale, capabilities, and performance. The model utilizes advanced techniques such as self-attention and parallel processing to handle complex natural language tasks effectively.

Uploaded by

sangrampanigrahi
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
9 views28 pages

Evolution of ChatGPT and GPT Models

ChatGPT is an AI language model developed by OpenAI, launched in November 2022, capable of understanding and generating human-like text. It is built on the Transformer architecture and has evolved through several iterations (GPT-1 to GPT-5), each improving in scale, capabilities, and performance. The model utilizes advanced techniques such as self-attention and parallel processing to handle complex natural language tasks effectively.

Uploaded by

sangrampanigrahi
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

ChatGPT

 ChatGPT is an AI language model or a chatbot released in November 2022 by OpenAI,


an artificial intelligence research company.
 You can think of it as a very advanced computer program that can understand text and
generate human-like responses.
 The "GPT" in ChatGPT stands for "Generative Pre-training Transformer" and refers to
the way it processes language.
 ChatGPT is an advanced artificial intelligence system or a very advanced computer
program that can understand text and generate human-like responses.
 ChatGPT is built on large language models that have been trained on vast amounts of
information, allowing it to recognize patterns in language, and engage in natural
conversation.
 ChatGPT works by predicting the next most likely word in a sentence based on the
context it has been given.
 This simple mechanism, combined with extensive training data and powerful neural
network architectures, enables it to perform a wide range of tasks such as explaining
complex concepts, generating creative content, summarizing long texts, and helping with
problem-solving.
 A free version of ChatGPT is available on OpenAI's website and the company also sells a
more advanced version on a subscription basis.

The Evolution of GPT Models


 The GPT (Generative Pre-trained Transformer) series represents a major progression in
natural language understanding and generation.
 Each version builds on the strengths of its predecessor, improving in scale, capabilities,
reasoning, speed, and alignment with human instructions.
 Below is a comprehensive overview of how GPT models have evolved over time.

1. GPT-1:

 GPT-1 is the original GPT model, introduced by OpenAI in 2018.


 GPT-1 introduced the idea of using the Transformer decoder architecture for
large-scale language modeling.
 This model was composed of 12 layers, each with 12 self-attention heads and a
total of 117 million parameters.
 It used unsupervised learning and was trained on the BookCorpus dataset, a
collection of 7,000 unpublished books.
 It demonstrated that pre-training on large text and then fine-tuning for specific
tasks could outperform earlier methods.
 GPT-1 proved that a single model could perform multiple NLP tasks with
minimal adjustments, laying the foundation for future generative AI.
2. GPT-2

 GPT-2 was released by OpenAI in 2019.


 GPT-2greatly expanded the scale and abilities of its predecessor, such as GPT-1.
 It was composed of 48 layers and a total of 1.5 billion parameters, such as 10
times bigger than GPT-1 parameters.
 This version was trained on a larger corpus of text data scraped from the Internet
(40GB of internet text or WebText), covering a more diverse range of topics and
styles.
 GPT-2’s text generation quality was so impressive that OpenAI initially withheld
the full model due to potential misuse. It brought global attention to large
language models.
 Notable Features
o Improved zero-shot learning
o Strong context understanding
o Ability to generate long, coherent paragraphs of text such as stories,
articles, and code-like patterns

3. GPT-3

 GPT-3 was introduced in 2020, marked another significant step up in scale and
performance.
 It was composed of multiple transformer layers and Trained on hundreds of
billions of words (a total of 175 billion parameters) from books, web data, and
programming sources.
 This model
 GPT-3 is the first GPT model widely adopted by the public (via APIs), which is
demonstrated an impressive ability to generate text that closely resembled human
language.
 The release of GPT-3 spurred widespread interest in the potential applications of
large language models, as well as discussions about the ethical implications and
challenges of such powerful models.
 GPT-3 sparked the global AI boom in conversational agents, prompting
widespread adoption in educational, professional, and commercial domains.
 Notable Abilities
o Few-shot learning: GPT-3 could learn a task from only a few examples
o Stronger reasoning and creativity
o Capable of producing human-like text, code, dialogues, and explanations.

4. GPT-3.5

 GPT-3.5 introduced in 2022, and made ChatGPT viable for everyday use.
 This model powered the first version of ChatGPT, making AI accessible to
millions worldwide.
 Notable Abilities
o Better instruction-following
o More stable responses
o Improved safety and alignment
o Optimized for chat interactions

5. GPT-4

 GPT-4 was introduced in 2023, and it transformed the GPT family by enabling
multimodal capabilities.
 GPT-4 trained on larger and more advanced parameters than GPT-3,
approximately estimated 1Trillion+ parameters via mixture-of-experts (Not
publicly disclosed).
 GPT-4 is a revolutionary multimodal language model with capabilities extending
to processing both text and image inputs, describing humor in images, and
summarizing text from screenshots.
 GPT-4’s interactions with external interfaces enable tasks beyond text prediction,
making it a transformative tool in natural language processing and various
domains.
 GPT-4 set a new industry standard for general-purpose AI models, that used in
ChatGPT+, enterprise, education, and coding tools.
 Notable Abilities
o More reliable long-context reasoning
o Multimodal input (text + image),
o Can process text and images
o Better understanding of images and diagrams
o Strong across reasoning, math, coding, and knowledge-intensive tasks,
such as law, medicine, etc.
o Support for safer, more grounded responses

6. GPT- 4.1, GPT- 4o, and Variants

 GPT-4.1, GPT-4o, and Variants were introduced in 2024 and these models
expanded GPT-4 into faster and more affordable variants.
 They enabled real-time AI assistants, integrated voice interactions, and broader
everyday use.
 GPT-4.1: More affordable and faster than GPT-4 with similar intelligence.
 GPT-4o (Omni):
o Truly multimodal (text, voice, image, audio)
o Real-time voice interaction
o Efficient and lightweight
o Good for mobile and edge devices
 GPT-4o Mini: Optimized for speed and lightweight applications

7. GPT-5 Series
 GPT-5 models were introduced in 2025, which presents the advanced reasoning
and autonomy.
 GPT-5 models represent the next generation with improved capability and deeper
understanding.
 GPT-5 models enable more autonomous workflows, advanced tutoring, smarter
assistants, and precise domain-specific problem solving.
 GPT-5:
o Major upgrades in reasoning, planning, and tool-use
o Enhanced long-context processing effortlessly
o Superior reasoning, including step-by-step logical and mathematical
capabilities
o More reliable hallucination control
 GPT-5.1 (Current model):
o Even stronger multi-step reasoning
o More stable at following instructions
o Better memory, safer outputs
o Improved multilingual capability
o Rich multimodal interaction across text, images, and audio
o More natural, human-like conversation flow
o Better tool use, enabling high-level automation

Summary of Evolution of GPT Models


Model Year Parameters Key Features
GPT-1 2018 117M First transformer LM
GPT-2 2019 1.5B Coherent long text
GPT-3 2020 175B Few-shot learning
GPT-3.5 2022 — Better instruction following
GPT-4 2023 — Multimodal, high reasoning
GPT-4o 2024 — Omni, real-time multimodality
GPT-5 2025 — Advanced reasoning, planning
GPT-5.1 2025 — Strongest reasoning, stable outputs

The Transformer Architecture: A Recap


 Transformer models are a type of model used in machine learning, particularly in the
field of natural language processing (NLP).
 Transformer model was introduced by Vaswani et al. in the paper “Attention is All You
Need.”
 The main advantage of Transformer models is that they process input data in parallel
rather than sequentially, allowing for more efficient computation and the ability to handle
longer sequences of data.
 Transformer models also introduced the concept of “attention,” enabling the model to
weigh the importance of different words in the input when generating an output.
 Key Pointers to Remember About the Transformer Architecture
1. Parallel Processing Instead of Recurrence

 Transformers eliminate recurrence (unlike RNNs, LSTMs, GRUs).


 All tokens in a sequence are processed simultaneously, enabling significant
computational speed-up during training on GPUs/TPUs.
 This parallelism is possible due to self-attention, which computes relationships
between tokens in a single step.

2. Self-Attention as the Core Mechanism

 Self-attention allows each token to attend to every other token in the sequence.
 It computes attention using Query (Q), Key (K), and Value (V) vectors derived
from embeddings.
 Self-attention helps the model capture:
o Long-range dependencies
o Contextual relationships
o Semantic similarity
 Scaled dot-product attention prevents extremely large scores during softmax by
dividing by √ d k , stabilizing gradients.

3. Multi-Head Attention (MHA)

 Instead of using a single attention head, the Transformer uses multiple heads.
 Each head learns different types of relationships (e.g., syntactic vs. semantic).
 Allows the model to gather multi-perspective contextual information.
 After processing, heads are concatenated and linearly projected.

4. Positional Encoding to Add Order

 Transformers have no inherent notion of sequence order.


 Positional encodings inject information about the position of each token.
 They can be:
o Sinusoidal (original paper)
o Learned embeddings (used in many modern variants)
 Provide information about relative and absolute positions.
5. Layer Normalization and Residual Connections

 Residual (skip) connections prevent vanishing gradients and allow deeper


networks.
 Layer normalization stabilizes training by normalizing inputs to sublayers.
 Typical pattern:
o LayerNorm( x + Sublayer(x) )

6. Feed-Forward Networks (FFN) After Attention

 Each layer includes a position-wise FFN applied independently to every token.


 FFN consists of two linear layers with a nonlinearity (ReLU/GELU).
 Increases the representational capacity of each layer.

7. Encoder–Decoder Architecture

Encoder

o Contains N identical layers with:


 Multi-head self-attention
 Feed-forward network
 It transforms input tokens into contextual representations.

Decoder

o Also contains N layers, but includes:


 Masked multi-head self-attention (prevents seeing future tokens)
 Encoder–decoder attention (attends to encoder outputs)
 Feed-forward network

8. Masking Mechanisms

 Padding mask: prevents attention to padded positions.


 Look-ahead mask (in decoder): prevents using future information during
training.

9. Tokenization and Embedding

 Uses subword tokenization like BPE, WordPiece, or SentencePiece.


 Embedding vectors represent tokens in a continuous vector space.

10. Transformer Advantages

 Captures long-range dependencies better than RNNs.


 Enables highly efficient parallel computation.
 Scales extremely well with model size (basis of GPT, BERT, T5, etc.).
 Achieves state-of-the-art results in NLP, vision, speech, and multimodal tasks.

11. Variants and Improvements

 Modern models build on the core Transformer architecture:


o BERT: encoder-only, for understanding tasks.
o GPT: decoder-only, for generative tasks.
o T5 / BART: full encoder–decoder for sequence transformation.
o Vision Transformers (ViT): apply Transformers to image patches.
o Longformer / BigBird: sparse attention for long sequences.

12. Computational Considerations

 Self-attention has O(n²) memory/time complexity → expensive for very long


sequences.
 Numerous efficient attention mechanisms (linear attention, sparse attention, Flash
Attention) have been developed to address this.
Architecture of ChatGPT
 ChatGPT is built on the Transformer architecture, specifically a decoder-only
Transformer.
 Its architecture can be divided down into several primary components:

1. Tokenizer
2. Embedding layer
3. Positional encoding
4. Multiple decoder blocks
5. Output Layer

1. Tokenization
 At first, the input sequence is broken down into a series of tokens or individual
sequence components.
 For instance, if the input is a sentence, the tokens are words.
 Before entering the model, text is broken into tokens using Byte-Pair Encoding (BPE)
or SentencePiece.
 Example: “ChatGPT is powerful” → ["ChatGPT", " is", " powerful"]

2. Embedding Layer

 Each token is converted into a dense vector (learned representation). Two embeddings
are added:
 Token embeddings
 This stage converts the input sequence into the mathematical domain that
software algorithms understand.
 Embedding then transforms the token sequence into a mathematical vector
sequence.
 Vectors are visualized as n-dimensional spaces with series of attributes
(coordinates) are represented with numbers.
 The vectors carry semantic and syntax information (such as, grammar, meaning,
and use in sentences) about any word.
 During the training process, Software can use the numbers to calculate the
relationships between words in mathematical terms and understand the human
language model.
 Embeddings provide a way to represent discrete tokens as continuous vectors that
the model can process and learn from.

3. Positional embeddings (to encode order of words)

 Positional encoding is a crucial component in the ChatGPT architecture because


the Transformer model processes tokens in parallel (not sequentially like RNNs),
it needs a way to know the order of tokens.
 With positional encoding, the model can preserve the order of the tokens and
understand the sequence context.
 Positional encoding adds information to each token's embedding to indicate its
position in the sequence.
 This is often done by using a set of functions that generate a unique positional
signal that is added to the embedding of each token.
PE( pos , 2 i)=sincos ¿ ¿
PE( pos , 2 i+1)=cos ¿ ¿

Where , pos correspond to the position of the word,

i is the current dimension index.


dmodel is the dimension of the model.

 The sine function is used for even indices of the embedding while the cosine
function is used for odd indices.
 These positional encodings (PE) is added to the token embeddings, so the model
can distinguish “dog chases cat” vs “cat chases dog.”

4. Stacked Transformer Decoder Blocks

 ChatGPT uses a decoder-only structure.


 Decoder-only meaning a stack of identical “decoder blocks” that handle both interpreting
input context and generating output.
 ChatGPT is built using a stack of many (or multiple) decoder blocks, each one
transforming the token embeddings step-by-step to progressively build deeper semantic
understanding.
 GPT uses only decoder blocks stacked, e.g., 12, 24, 48, 96 or 100+ depending on the
model size.
Model # Decoder Blocks

GPT-2 Small 12

GPT-2 Medium 24

GPT-3 (175B) 96

GPT-4-class models 100+

 Multiple decoder blocks allow hierarchical learning:


1. Lower Layers (1–12 layers) capture local patterns (e.g., "is", "are"); detects
patterns Word relationships and Local syntax.
2. Middle Layers (12–48 layers) capture grammar and structure; phrase meaning
and co-reference, (such as, linking pronouns to nouns)
3. Higher Layers (48+ layers) capture, world knowledge, deep semantics, Abstract
reasoning, Contextual decision-making, Task-specific patterns (like as,
summaries, answers, instructions)
 The deeper the stack, the better the model becomes at reasoning across long text
sequences.
 Stacking many blocks gives the model deeper capacity to learn:
1. Long-range dependencies
2. Hierarchical language patterns
3. Relationships between distant tokens
4. Semantic and syntactic structures
 These stacked blocks form the backbone of autoregressive language generation.
 Each token starts as an embedding vector with positional information and then it goes
through a sequence of transformations:
 Each Decoder block typically includes the following subcomponents:
 Multi-Head Self-Attention (MMHSA)
 Masked Multi-Head Self-Attention (MMHSA)
 Layer Normalization (LayerNorm)
 Residual Connections in ChatGPT
 Feed-Forward Network (MLP Layer)

1. Multi-Head Self-Attention (MMHSA)


 To understand the Multi-Head Self-Attention, first we need to study about
the following:
 Self-Attention
 The self-attention mechanism enables the model to weigh the
importance of different tokens within the sequence.
 Self-Attention allow each token in a sequence attend to (look at) other
tokens of the input to gather contextual information.
 The self-attention mechanism enables ChatGPT to capture long-range
dependencies in the conversation history.
 The self-attention mechanism allows the model to consider all tokens in
the conversation history regardless of their distance from the current
token.
 This capability is crucial for understanding the flow of the conversation
and generating responses that maintain coherence over extended
dialogues.
 A single self-attention mechanism is called a head.
 The head works as follows:
 First, the each input token is fed into three separate linear layers.
 Then each token is multiplied with the Queries (Q) and the keys
(K), then scaled, and turned into a probability distribution using a
softmax activation function.
 Think of this probability distribution as describing which indices
matter most for the output (i.e. which words in the prompt matter
for the next word to be predicted).
 Finally, the output is multiplied with Values (V).
 Attention score:

( )
T
QK
Attention ( K , Q ,V ) =softmax V
√ dk
Where, Where dk = dimension of the key vectors (scaling
factor).
 Multi-Head Attention
 Instead of one attention operation, the model uses multiple heads:
o Multi-head attention is nothing more than several individual
heads stacked on top of one another.
o Each head focuses on a different type of relation (syntax,
semantics, long-range context, etc.)
o The input to all heads is equivalent. However, each head has its
own weights and run attention in parallel
 After forwarding the input through all the heads, the output of the
heads is concatenated and passed through a linear layer which brings
the dimensionality back to the dimension of the initial input.
 The following figure gives an overview of the operations done in a
head and an overview of how multi-head attention works.
2. Masked Multi-Head Self-Attention (MMHSA)
 This mechanism allows the model to generate text one token at a
time while ensuring it cannot “see” future tokens.
 The masking mechanism involves applying a triangular mask to
the self-attention matrix, where all elements below the main
diagonal are set to negative infinity (or a very large negative
value).
 This effectively masks out the future tokens and allows the token
to only attend to its previous tokens and itself.
 Masking Is Needed, because ChatGPT generates text left-to-right.
So when predicting token t, it must not use information from future
tokens, such as t+1, t+2, …
 Example:
 Let’s consider an example of generating the sentence “I
love natural language processing” using ChatGPT.
 During the generation process, when the model is
predicting the word “language,” it should only attend to the
previous tokens “I,” “love,” “natural,” and “language”
itself.
 The attention to the word “processing” should be masked
out to maintain causality.
 If it could see future tokens, it would “cheat” during
training.
 Masking prevents attention to future tokens.
 Splits embedding into multiple “heads.”
 Each head learns different relations (syntax, meaning, long-range
dependencies).
 Output = weighted combination of all previous token embeddings.
 Example: For a sequence of 5 tokens: here, future tokens get 0
attention weight.

1 2 3 4 5
1 ✓ 0 0 0 0
2 ✓ ✓ 0 0 0
3 ✓ ✓ ✓ 0 0
4 ✓ ✓ ✓ ✓ 0
5 ✓ ✓ ✓ ✓ ✓

 Mathematical Formula (with Mask)

( )
T
QK
Attention ( K , Q ,V ) =softmax +M V
√ dk
Where mask M:

Mij = 0; if j ≤ i

Mij=−∞; if j > i

3. Layer Normalization (LayerNorm)


 Layer Normalization (LayerNorm) is a normalization technique applied
inside every Transformer block in ChatGPT.
 It stabilizes training, improves gradient flow, and makes token
representations more consistent.
 In ChatGPT models (, LayerNorm is applied before attention and
before the feed-forward network — a style called Pre-LayerNorm.
 For Example:
 Let the x is the each token’s hidden vector, such as x = [x1, x2, ... ,
xd]
 LayerNorm normalizes the features of that token by computing
Mean (μ) and Variance (σ2)
n
1
Mean(μ)= ∑ x i
n i=1
n
1
Variance(σ )= ∑ (xi −μ)2
2
n i=1

(xi −μ)
Normalization ( x^ i )=
√ σ 2+ ε
 Then scale and shift, hence

LayerNorm ( xi ) =γ x^ i + β

Where:
γ (gamma) = learnable scale
β (beta) = learnable shift
ε = small constant for numerical stability
4. Residual Connections in ChatGPT
 Residual connections (also called skip connections) are a critical
architectural component inside every Transformer block of ChatGPT.
 They allow information to bypass complex layers, making training deeper
models stable, faster, and more accurate.
 Residual connections were introduced in ResNet, but they are even more
important in large-scale language models like GPT.
 Residual connections, sometimes referred to as skip connections, allow the
model to retain important information from the previous layers and
provide a mechanism to “skip” some layers by zeroing the weights,
resulting in an over-specified model that can learn sparsity.
 A residual connection simply adds the input of a sub-layer directly to its
output:
Output = x + SubLayer(x)
Where,
x = input
SubLayer(x) = output of Self-Attention or Feed-Forward
Network
Output = x + F(x)
Where,
F(x) = SubLayer(x)
 This means the network learns a residual function (F(x))
 The model learns “just the difference” from the identity mapping.
 In Residual Connections in a GPT Decoder Block, Every GPT block has
two residual paths:
 Around Masked Multi-Head Self-Attention
x1 = x0 + Attention(LayerNorm(x0))
 Around the Feed-Forward Network
x2 = x1 + FFN(LayerNorm(x1))
 Residual Connections Are enable ChatGPT to:
 Train extremely deep networks
 Maintain stable gradients
 Preserve original information
 Learn meaningful refinements
 Scale to billions of parameters
 Achieve strong performance on language tasks
 Without residual connections, modern Transformers and ChatGPT would
not be trainable at their enormous scale.
5. Feed-Forward Network (FFN / MLP)
 In every Transformer decoder block of ChatGPT, after Masked Multi-
Head Self-Attention, the next major component is the Position-wise
Feed-Forward Network (FFN).
 GPT models use FFNs to expand and compress representation for deep
feature transformation.
 After attention, each token representation passes through a small neural
network (usually two linear layers + non-linear activation), which enabling
the model to learn richer representations.

FFN(x) = W2 (σ(W1x + b1)) + b2

Where:
x = input vector of size dmodel
W1 expands dimensionality → typically to 4× or 8× the model size.
 Example (GPT-3): dmodel =12288, then FFN intermediate
size = 4× 12288 = 49152.

 This expansion helps the model capture complex, high-


dimensional patterns.
W2 projects it back to dmodel
σ is a non-linear activation function in GPT, such as
GELU(Gaussian Error Linear Unit).
 GELU(x) = x ⋅ Φ(x)
 It smoothly gates input based on probability, improving
training stability and model quality.
 FFN processes each token independently
 Let a sequence is [x1, x2, x3,…] and The FFN
applies: FFN(x1), FFN(x2), x3), …
 This maintains the sequence structure while
enhancing token representations.

5. Output Layer in ChatGPT (GPT Models)

 In GPT-style Transformer models, the output layer is the final processing stage
responsible for converting the model’s internal hidden representations into actual text
tokens (words, subwords, or characters).
 It is the final step after all decoder blocks have processed the input.
 The output layer is extremely important because it transforms the model’s understanding
into predictions.
 The output layer in GPT has two main components:
 Linear Projection Layer (Fully Connected Layer)
 Softmax Layer

1. Linear Projection Layer (Fully Connected Layer)

 In ChatGPT’s Transformer architecture, the Linear Projection Layer


(often called a Fully Connected (FC) layer or dense layer) plays a crucial
role in transforming vector representations at various stages of the model.
 This Linear Projection Layer produces a final output projection to
vocabulary vector, called the logits vector.
 This layer maps the final hidden vector from the decoder to a vector of
size equal to the vocabulary size.
 Suppose:
Hidden state dimension = dmodel
Vocabulary size = V
 Then the linear transformation is:
logits = W ⋅ h + b
Where:
h = final hidden vector from the last decoder block
W = weight matrix size (V × dmodel)
b = bias vector of size V
 This produces a vector called the logits vector, which consists scores.
 Logits are unnormalized scores for each possible token.
 Example (simplified vocabulary of 4 tokens):

logits = [4.6, 1.2, -0.3, 0.8]

Some numbers may be negative.



These numbers are not probabilities.

These numbers do not sum to 1.

 Every score (or number) in this logits vector corresponds to one
token (e.g., "the", "cat", "model", etc.) in the vocabulary (e.g.,
50,000 scores).
2. Softmax Layer

 The Softmax function is one of the most important components in


ChatGPT’s architecture, especially in final token prediction.
 It converts raw scores (logits) into probabilities that sum to 1.
 Softmax is a mathematical function that transforms a vector of arbitrary
real numbers (positive, negative, or zero) into a probability distribution.
 let a logit vector: z = [z1, z2, ..., zn]
 Softmax converts these logits to probabilities over vocabulary tokens:
zi
e
P(z i )=Softmax (z i)= n

∑ ez j

j=1
Where,
P(z i ) is probability for each token.
Exponentiation each logit token such as, e z , is ensured
i

positivity.
n

∑ e z is the sum of all exponentiated logits


j

j=1

 Properties of softmax:
 Every output is between 0 and 1.
 Sum of all outputs is 1.
 Larger input values get exponentially higher probabilities.
 Softmax transforms unbounded scores into probabilities and ChatGPT
needs to make decisions based on probabilities.
 Softmax also creates contrast between tokens due to exponentiation.
 Smaller numbers become closer to zero, while higher values
become much more dominant.
 This helps the model strongly prefer certain relationships or next-
word predictions.
 Softmax supports gradient-based learning (backpropagation), so it
can be trained efficiently.
 Softmax determines which token ChatGPT outputs next.
 The highest probability token is selected (or sampled, depending
on decoding strategy).
 Example of Next Token Prediction using Softmax
 Suppose the logits for tokens are:

Token in Vocabulary “cat” “dog” “car”


Logit Vector 5.0 2.0 1.0
 Exponentiation each logit token such as, e z , such as i

For 1st token, e⁵ = 148.4


For 2nd token, e² = 7.39
For 3rd token, e¹ = 2.71
 The sum of all exponent logits, such as
n

∑ e z =148.4+7.39+ 2.71=158.5
j

j=1

 Probabilities ( P(z i )) for each token is computed using following


formula:
zi
e
P(z i )=Softmax (z i)= n

∑ ez j

j=1

So that,
148.4
P(cat)= ≈ 0.936
158.5

7.39
P(dog)= ≈ 0.046
158.5

2.71
P(car)= ≈ 0.017
158.5

Hence,
P = [0.936, 0.046, 0.017]
 The first token has the highest probability; hence ChatGPT is most
likely to generate it.
The output is “cat”.

ChatGPT’s Training Process

 The training process of ChatGPT is arguably the most crucial aspect responsible for its
remarkable ability to engage in human-like conversations.
 Roughly, we can divide the training process into two phases:
 Pre-training
 Fine-tuning

1. Pre-training

o Pre-training is the core foundation of how ChatGPT learns language, concepts,


world knowledge, and reasoning abilities.
o It is the largest and most computationally expensive stage in the entire training
pipeline.
o The effectiveness of pre-training stage determines everything ChatGPT does later
on, including following instructions, generating meaningful responses, solving
task.
o The first step in training ChatGPT is unsupervised learning, often referred to as
pre-training.
o Pre-training played a crucial role in developing a ChatGPT as it helped to learn
the basic rules of language and understand common word usage and phrases.
o Pre-training is a process where the model learns to predict the next token (next
word/subword) in a sentence using a massive text dataset.
o At this stage, the model was trained on a vast corpus of text data of approximately
570GB, consisting of books, articles, Wikipedia, and other internet text sources.
o It is an unsupervised learning stage—no labels, instructions, or examples from
humans are needed.
o The main objective is for the model to use a massive text dataset to learn the
statistical characteristics of the language, such as grammar and syntax, and to
predict the next token (word or subword) in a sentence.
o The training objective was to predict the next token in a given text based on the
context of preceding words.
 For example, given the sentence “The cat sat on the”, the model’s task is
to predict the next word. It might predict “mat”, “sofa”, “floor”, etc., based
on the patterns it has learned from the training data.
o By comparing its prediction with the actual next token in the sentence, the model
adjusts its weights during training to enhance the accuracy of future predictions.
o Now this model is known as the GPT Base model.
o The Pre-training knowledge is then used as a foundation for further
customization.
 Further adjustments are necessary to enhance its performance and make it
more conversational, like a chatbot.

2. Fine-tuning

o Fine-tuning is the stage that transforms a general language model (trained during
pre-training) into a useful, instruction-following, safe, and aligned
conversational assistant.
o It is performed after pre-training and plays a major role in making ChatGPT
respond helpfully.
o Fine-tuning mainly includes two major steps:
 Supervised Fine-Tuning (SFT)
 Reinforcement Learning from Human Feedback (RLHF)
o Together, this process shapes how ChatGPT behaves, answers questions, and
maintains safety.
a. Supervised Fine-Tuning (SFT)
 Supervised Fine-Tuning (SFT) is the first major stage in training
ChatGPT after pre-training.
 It teaches the model to follow instructions, produce useful responses, and
behave safely.
 Supervised Fine-Tuning (SFT) is the second major phase in the training
pipeline of ChatGPT, performed after pre-training but before
Reinforcement Learning from Human Feedback (RLHF)
 It transforms a raw language model—which simply predicts the next token
—into a model that can follow instructions, answer questions, generates
structured responses, and behaves like a helpful assistant.
 In this phase, the model is fine-tuned on labelled datasets tailored
to particular applications such as sentiment analysis, question
answering, or text summarization.
 For example, in a sentiment analysis task, the model might be
trained on a dataset of movie reviews labelled as positive or
negative. Given a new movie review, the model would learn to
predict the sentiment as either positive or negative based on the
examples it has seen during training.
 After pre-training GPT Base model is great at language, but it does not
generate conversational context-driven responses, Thus, the ChatGPT
Base model goes through a Supervised Fine-Tuning process to focus on
learning of prompt-and-response interactions.
 Supervised Fine-Tuning is a training stage where the model is fed
prompt–response pairs written by humans, and it learns to imitate these
high-quality outputs.
 During Supervised Fine-Tuning process the following steps are taken
place:
 Design Human-Created Dialogs:
 Human agent creates a set of dialogue scenarios by playing both
the requester and responder roles.
 In order to help the model learn how to respond appropriately, it
includes not only the query context (history) but also the expected
response for those queries.
 Some hypothetical examples are:
o Prompt: "What is the capital of France?"
o Response: "The capital of France is Paris."

o Prompt: "How do I make chocolate chip cookies?"


o Response: "To make chocolate chip cookies, you'll need
flour, butter, sugar, chocolate chips, and vanilla extract.
Start by creaming the butter and sugar together, and then
gradually add the dry ingredients. Finally, fold in the
chocolate chips and bake in the oven until golden brown."
 Creating the Training Dataset:
 The query context and the expected response for those queries
make up the Supervised Fine-Tuning Training Data Corpus.
 In the dataset the query context as input and the next response as
output.
 Training Process:
 The Base GPT model receives the dataset for training in a
supervised manner.
 The Base GPT model undergoes optimization by using Stochastic
Gradient Descent technique, where its outputs are adapted to match
the expected responses provided by human trainers.
 Now this model is known as the SFT ChatGPT Model, which can hold a
conversation more adeptly than a GPT Base model.

b. Reinforcement Learning from Human Feedback (RLHF)


 Finally, the Reinforcement Learning process is used to further train the
Supervised Fine-tuning model.
 Reinforcement Learning through Human Feedback (RLHF), transforms
the fine-tuned model into a sophisticated conversational assistant.
 While Supervised Fine-Tuning teaches the model to generate ideal
responses, RLHF polishes and optimizes its behavior to align with user
preferences, making interactions more natural, helpful, and safe.
 Reward Model (RM) Training
 This step uses human preferences to train a secondary model called
a Reward Model.
 The Supervised Fine-Tuning (SFT) model generates several
alternative outputs (several possible responses) for the same input
(or same user query)
 Human evaluators create a dataset by ranking these responses on a
scale from best to worst based on criteria like accuracy,
helpfulness, tone, etc.
 Following the creation of the dataset, the reward model is trained,
which assigns numerical scores to the model's responses. It serves
as a standard by which to measure the effectiveness of its future
outputs.
 Example of Reward Model Training Steps
 The following steps involved for reward Model training
o Step 1: Generate Multiple Responses
 For a given user prompt, “Explain overfitting in
machine learning.”, the Supervised Fine-Tuned (SFT)
model generates several possible responses.
 Response A: “Overfitting happens when a model
memorizes the training data and performs poorly on
new data.”
 Response B: “Overfitting is a condition where a
machine learning model captures noise, irrelevant
patterns, or random fluctuations from the training set.
As a result, the model performs extremely well on
training data but fails to generalize to unseen data.
Techniques like cross-validation, regularization
(L1/L2), early stopping, and dropout help reduce
overfitting.”
 Response C: “Overfitting is when a model does too
much learning and doesn’t work properly.”
o Step 2: Human Preference Data Set Creation
 In Step 2, these responses are examined and ranked by
human evaluators to create a dataset.
 The human evaluators are analyze that
 Response B is most detailed and accurate,
 Response A is correct but short
 Response C is unclear.
 Then rank these responses them as follows:
B>A>C
 This rank produces a set of preference pairs
like as follows:
o B>A
o A>C
o B>C
 These pairs help the RM learn how to assign higher
rewards to good answers.
o Step 3: Convert Preferences into Training Data
 For each pair:
 Chosen response = the preferred one (high rank)
 Rejected response = the less preferred one (high
rank)
 For Example:
o Pair 1: B > A
 Chosen = B
 Rejected = A
o Pair 2: A > C
 Chosen = A
 Rejected = C
o Pair 3: B > C
 Chosen = B
 Rejected = C
 These pairs are used to teach the RM that:
 Response B should get a higher reward than
Response A
 Response A should get higher reward than
Response Ã
 Response B should get higher reward than
Response C
o Step 4: Train the Reward Model
 The Reward Model reads each prompt–response pair
and outputs a scalar score:
r = RM (prompt, response)
Where,
r, a single scalar value, which represents how
good or preferred that response is according to
human preferences.
Reward Model (RM) is a neural network that
takes two inputs:
prompt – the user’s question or instruction
response – the model’s answer to that prompt
and it outputs:
 For Example:
 The Reward Model computes:
rA =RM(prompt, Response A) = 1.6
rB =RM(prompt, Response B) = 3.2
 Since 3.2 > 1.6, it concludes:
o Response B is better
o Response B aligns more with human
preferences
 The model is trained using pairwise ranking loss:
Loss = − log (σ(rchosen − rrejected))

Where,
rchosen = reward for the human-preferred
response
rrejected = reward for the less-preferred response
Loss is the pairwise ranking loss used to train
the Reward Model (RM) in ChatGPT’s RLHF
pipeline.
The loss encourages the model to make:
rchosen > rrejected
σ is the sigmoid function, which maps any value
to a probability between 0 and 1.
1
σ (x)= −x
1+e
 In this case, we compute:
σ(rchosen − rrejected)
If, rchosen ≫ rrejected, then σ(rchosen − rrejected) ≈ 1 → good
If, rchosen ≈ rrejected, σ(rchosen − rrejected) ≈ 0.5 → okay
If, rchosen < rrejected , σ(rchosen − rrejected) < 0.5 → bad
 The loss penalizes the model when it gives higher
reward to a bad answer.
 Interpretation of the Full Formula
Loss = − log (σ(rchosen − rrejected))

Example 1: If the chosen response scores higher


Given, rchosen = 4.0, rrejected = 1.0
rchosen − rrejected = 4.0 – 1.0 = 3.0
σ(rchosen − rrejected) = σ(3.0) ≈ 0.95
Loss = − log (σ(rchosen − rrejected))
= − log (0.95) ≈ 0.05
 Remark: Low loss → good → RM is
behaving correctly.
Example 2: If both responses have similar scores
Given, rchosen = 3.0, rrejected = 3.0
rchosen − rrejected = 3.0 – 3.0 = 0
σ(rchosen − rrejected) = σ(0) ≈ 0.5
Loss = − log (σ(rchosen − rrejected))
= − log (0.955) ≈ 0.693

 Remark: Medium loss → RM needs to


push scores farther apart.
Example 3: If the RM scores the rejected answer
higher scores
Given, rchosen = 1.0, rrejected = 3.0
rchosen − rrejected = 1.0 – 3.0 = - 2.0
σ(rchosen − rrejected) = σ(- 2.0) ≈ 0.12
Loss = − log (σ(rchosen − rrejected))
= − log (0.12) ≈ 2.12

 Remark: High loss → RM is wrong and must be


corrected.
How the Reward Model Is Used After Training
o Once trained, the Reward Model becomes the critic in the
reinforcement learning step (e.g., PPO or GRPO).
o The process is:
 The model generates several responses.
 The Reward Model assigns scores to each.
 RL optimizes the model so that it produces
responses with higher expected reward.
 Proximal Policy Optimization
 The tuned ChatGPT model is trained more using a reinforcement
learning algorithm, usually Proximal Policy Optimization (PPO).
 Responses are produced by the model and assessed using the
reward model. The reinforcement learning algorithm then tunes the
parameters of your model to maximize that reward score.

Pre-Training vs. SFT vs. RLHF

Feature Pre-Training SFT RLHF

Purpose Learn language patterns Teach instruction-following Optimize for human preference

Data Internet-scale Human-written examples Human rankings

Method Next-token prediction Supervised learning Reinforcement learning

Output “Raw GPT” Helpful assistant Human-preferred assistant

Advantages of ChatGPT:

1. Large Knowledge Base: ChatGPT has been trained on a massive dataset and has access
to a vast amount of information across various domains, which enables it to answer a
wide range of questions accurately.
2. 24/7 Availability: Unlike humans who require breaks and sleep, ChatGPT can operate
around the clock without any downtime. This makes it available to users at any time of
the day, including weekends and holidays.
3. Consistent Quality: ChatGPT can provide consistent and reliable answers to questions
without being influenced by emotions, fatigue, or personal biases. This ensures that users
get accurate and unbiased information every time.
4. Multilingual Support: ChatGPT can communicate in various languages, making it
accessible to a diverse range of users around the world.
5. Fast Response Time: ChatGPT can process and respond to queries quickly, which
makes it ideal for situations that require immediate responses.
6. Scalability: ChatGPT can handle an almost unlimited number of users simultaneously,
making it suitable for large-scale applications.
7. Personalized Experience: With its ability to learn and adapt to user preferences,
ChatGPT can provide a personalized experience, improving user engagement and
satisfaction.

Limitations of ChatGPT

1. Knowledge Cutoff: ChatGPT’s knowledge is limited to the information it was trained


on, which means that it may not have access to the latest information or updates in certain
domains.
2. Contextual Understanding: Although ChatGPT can generate responses based on the
input it receives, it may not always fully understand the context of a question or the
nuances of language, leading to inaccurate or irrelevant responses.
3. Biased Responses: ChatGPT may produce biased responses based on the biases present
in the data it was trained on, leading to inaccurate or discriminatory responses.
4. Lack of Emotional Intelligence: ChatGPT does not have emotions or emotional
intelligence, making it challenging to understand or respond to questions that require
empathy or sensitivity.
5. Security Concerns: As with any technology that interacts with users, there are security
concerns with ChatGPT, such as protecting user privacy, preventing malicious use, and
guarding against hacking attempts.
6. Need for Training: To improve its performance, ChatGPT requires continuous training
with relevant data and feedback, which can be time-consuming and resource-intensive.
7. Lack of Creativity: While ChatGPT can generate new text based on the input it receives,
it may not be able to produce creative or original responses.

Input Embeddings

Decoder Block 1

You might also like