0% found this document useful (0 votes)
13 views9 pages

Decoder-Only Transformer LLM Guide

Natural Language Processing
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
13 views9 pages

Decoder-Only Transformer LLM Guide

Natural Language Processing
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Decoder-Only Transformer (LLM) For Question

Asking
"FROM SCRATCH"

Notebook Structure

Data
Data source
Tokenization
Features and Target
Test data
Model Design
Positional encoding
Multi-head attention
Transformer Decoder
Final Architecture
Training script
Simplistic Inference Script
Issues and mistakes
Pre-training with a downstream task
Not masking Padding layers
Context window

In [1]: #necessary imports


import numpy as np
import pandas as pd
import torch
import [Link] as plot
import [Link] as F
from transformers import AutoTokenizer, AutoModel
import random

In [2]: if [Link].is_available():
device = [Link]("cuda")
print("GPU is available")
else:
device = [Link]("cpu")
print("GPU is not available, using CPU")

GPU is available

DATA

Data Source
The data I used for this project is the Stanford Question Ansewring Dataset (SQuAD).
SQuAD was prepared such that a question and a context would map to an Answer (Q+C -->
A). I modified this the data so that a Context would map to question (C --> Q).
Find my modified data here = link to dataset

In [3]: data = pd.read_json('{fill with path to your data}').to_dict(orient='list')

Tokenization
In [ ]: #bert tokenizer
tokenizer = AutoTokenizer.from_pretrained('bert-base-uncased')

In [5]: #Example
data['conversation'][0]

[{'from': 'human',
Out[5]:
'value': 'Beyoncé Giselle Knowles-Carter (/biːˈjɒnseɪ/ bee-YON-say) (born September 4, 1
981) is an American singer, songwriter, record producer and actress. Born and raised in
Houston, Texas, she performed in various singing and dancing competitions as a child, an
d rose to fame in the late 1990s as lead singer of R&B girl-group Destiny\'s Child. Mana
ged by her father, Mathew Knowles, the group became one of the world\'s best-selling gir
l groups of all time. Their hiatus saw the release of Beyoncé\'s debut album, Dangerousl
y in Love (2003), which established her as a solo artist worldwide, earned five Grammy A
wards and featured the Billboard Hot 100 number-one singles "Crazy in Love" and "Baby Bo
y".'},
{'from': 'gpt', 'value': 'When did Beyonce start becoming popular?'}]

In [6]: def tokenize_input(qa):


#1. tokenizing with a max seq length of 300 and padding layers
#2. Adding an <sos> and <eos> token to target values. In this case; [CLS] and [SEP]
seq_length = 300
q_tokens = tokenizer(qa[0]['value'],add_special_tokens=False)['input_ids']
a_tokens = tokenizer(qa[1]['value'],padding=True)['input_ids']

x_tokens = q_tokens + a_tokens[:-1]


y_tokens = q_tokens[1:] + a_tokens

x_pad = [0 for i in range(seq_length-len(x_tokens))]


y_pad = [0 for i in range(seq_length-len(x_tokens))]
final_x = x_tokens + x_pad
final_y = y_tokens + y_pad

return final_x, final_y

In [7]: #tokenizing all data


tokens = []
targets = []
for i in [Link](data['conversation'],len(data['conversation'])):
try:
x, y = tokenize_input(i)

if len(x) == 300:
[Link](x)
[Link](y)
except:
pass

X = [Link](tokens)
Y = [Link](targets)

Token indices sequence length is longer than the specified maximum sequence length for t
his model (718 > 512). Running this sequence through the model will result in indexing e
rrors
In [8]: [Link], [Link]

([Link]([124975, 300]), [Link]([124975, 300]))


Out[8]:

Test Data
Create your test data here

Model Design

Important Notes
Embedding layer: I used the emmbedding layer from the bert model.
Positional encoding: Sinusoidal encoding from Attention is all you need
Attention: Multihead (4 heads)
Linear projection: Projected input to 224 before passing it through decoder
Number of decoders: 8

Embedding layer
In [9]: class embed([Link]):
def __init__(self):
super().__init__()
[Link] = AutoModel.from_pretrained('bert-base-uncased')

def forward(self,x_tokens):
inputs = {'input_ids':x_tokens}
with torch.no_grad():
attention_mask = (inputs['input_ids'] != 0).int()
outputs = [Link](**inputs,attention_mask=attention_mask)
embeddings = outputs.last_hidden_state * attention_mask.unsqueeze(-1)
return embeddings
Positional Encoding
In [10]: ## sinusoidal positional encoding

class pos_enc([Link]):
def __init__(self) -> None:
super().__init__()

def forward(self,x):
batch_size, max_seq_length, dmodel = [Link]
pe = torch.zeros_like(x) #position encoding matrix

# Compute the positional encoding values


for pos in range(max_seq_length):
for i in range(0, dmodel):
if i % 2 == 0:
pe[:, pos, i] = [Link](pos / (10000 ** (2 * i / dmodel)))
else:
pe[:, pos, i] = [Link](pos / (10000 ** (2 * i / dmodel)))

x = x + pe
return x

Self-attention mechanisim
In [11]: class self_attention([Link]):
def __init__(self,no_of_heads: int ,shape: tuple, mask: bool=False, QKV: list=[]
'''
Initializes a Self Attention module as described in the "Attention is all you ne
This module splits the input into multiple heads to allow the model to jointly a
from different representation subspaces at different positions. After attention
on each head, the module concatenates and linearly transforms the results.

## Parameters:
* no_of_heads (int): Number of attention heads. To implement single head att

* shape (tuple): A tuple (seq_length, dmodel) where `seq_length` is the lengt


and `dmodel` is the dimensionality of the input feature space

* mask (bool, optional): If True, a mask will be applied to prevent attentio

* QKV (list, optional): A list containing pre-computed Query (Q), Key (K), a
The forward pass computes the multi-head attention for input `x` and returns the
'''
super().__init__()
self.h = no_of_heads
self.seq_length,[Link] = shape
[Link] = [Link]//self.h
[Link] = [Link](dim=-1)
[Link] = [Link]([[Link]([Link],[Link]) for
[Link] = [Link]([[Link]([Link],[Link]) for
[Link] = [Link]([[Link]([Link],[Link]) for
self.output_linear = [Link]([Link],[Link])
[Link] = mask
[Link] = QKV

def __add_mask(self,atten_values):
#masking attention values
mask_value = -1e9
mask = [Link]([Link](atten_values.shape) * mask_value, diagonal=1)
masked = atten_values + [Link](device)
return masked

def forward(self, x):


heads = []
for i in range(self.h):
# Apply linear projections in batch from dmodel => h x d_k
if [Link]:
q = [Link][i]([Link][0])
k = [Link][i]([Link][1])
v = [Link][i]([Link][2])
else:
q = [Link][i](x)
k = [Link][i](x)
v = [Link][i](x)

# Calculate attention using the projected vectors q, k, and v


[Link] = [Link](q, [Link](-1, -2)) / [Link]([Link]
if [Link]:
[Link] = self.__add_mask([Link])

attn = [Link]([Link])
head_i = [Link](attn, v)

[Link](head_i)

# Concatenate all the heads together


multi_head = [Link](heads, dim=-1)
# Final linear layer
output = self.output_linear(multi_head)

return output + x # Residual connection

Decoder
In [12]: class decoder_layer([Link]):
def __init__(self,shape: tuple,no_of_heads:int = 1):
'''
Implementation of Transformer Dencoder
Parameters:
shape (tuple): The shape (H, W) of the input tensor
no_of_heads (int): number of heads in the attention mechanism. set this to 1
Returns:
Tensor: The output of the encoder layer after applying attention, feedforwar
'''
super().__init__()

self.max_seq_length,[Link] = shape
def ff_weights():
layer1 = [Link]([Link],600)
layer2 = [Link](600,600)
layer3 = [Link](600,[Link])
return layer1,layer2,layer3

self.no_of_heads = no_of_heads

self.multi_head = self_attention(no_of_heads=no_of_heads, mask=True,


shape=(self.max_seq_length,[Link]

self.layer1,self.layer2,self.layer3 = ff_weights()
[Link] = [Link](dim=-1)
[Link] = [Link](shape)
self.relu1 = [Link]()
self.relu2 = [Link]()

def feed_forward(self,x):
f = self.layer1(x)
f = self.relu1(f)
f = self.layer2(f)
f = self.relu2(f)
f = self.layer3(f)

return [Link](f + x) #residual connection

def forward(self,x):
x = self.multi_head(x)
x = [Link](x)
x = self.feed_forward(x)
x = [Link](x)

return x

Full Model Architecture


In [13]: class architecture([Link]):
def __init__(self,n_classes,shape) -> None:
super().__init__()
self.max_seq_length,[Link] = shape
self.projected_dmodel = 224
self.embedding_layer = embed()
self.proj_to_224 = [Link]([Link], self.projected_dmodel)
[Link] = pos_enc()
self.decoder1 = decoder_layer(shape=(self.max_seq_length,self.projected_dmodel),
self.decoder2 = decoder_layer(shape=(self.max_seq_length,self.projected_dmodel),
self.decoder3 = decoder_layer(shape=(self.max_seq_length,self.projected_dmodel),
self.decoder4 = decoder_layer(shape=(self.max_seq_length,self.projected_dmodel),
self.decoder5 = decoder_layer(shape=(self.max_seq_length,self.projected_dmodel),
self.decoder6 = decoder_layer(shape=(self.max_seq_length,self.projected_dmodel),
self.decoder7 = decoder_layer(shape=(self.max_seq_length,self.projected_dmodel),
self.decoder8 = decoder_layer(shape=(self.max_seq_length,self.projected_dmodel),
# self.decoder5 = decoder_layer(shape=(self.max_seq_length,self.projected_dmodel
self.final_MLP = [Link](self.projected_dmodel,n_classes)
[Link] = [Link](dim=2)

def forward(self,x,temperature=1.0):
x = self.embedding_layer(x)
x = self.proj_to_224(x)
x = [Link](x)
x = self.decoder1(x)
x = self.decoder2(x)
x = self.decoder3(x)
x = self.decoder4(x)
x = self.decoder5(x)
x = self.decoder6(x)
x = self.decoder7(x)
x = self.decoder8(x)
x = self.final_MLP(x)
logits = x / temperature
x = [Link](logits)

return x

Training Script
In [ ]: !pip install torchmetrics

In [14]: from torchmetrics import Accuracy


from tqdm import tqdm

In [15]: dataset = [Link](X,Y)


loader = [Link](dataset,batch_size=20,num_workers=0,shuffle=False)

In [ ]: vocab_size = tokenizer.vocab_size
model = architecture(n_classes = vocab_size, shape = (300,768))
model = [Link](device)
model.load_state_dict([Link]("{fill with path to model weights if any}"))

In [76]: metric = Accuracy(num_classes=vocab_size,task='multiclass').to(device)


optimizer = [Link]([Link](),lr=0.0001)
criterion = [Link](ignore_index=0,label_smoothing=0.01)

In [ ]: from tqdm import tqdm

print('Training Started')
NUM_EPOCHS = 1

for epoch in range(NUM_EPOCHS):


[Link]() # Set the model to training mode
running_loss = 0.0
epoch_accuracy = 0.0
num_batches = len(loader)

# Initialize tqdm progress bar


with tqdm(total=num_batches, desc=f"Epoch {epoch + 1}", leave=True) as pbar:
for i, (x_batch, y_batch) in enumerate(loader):
x_batch, y_batch = x_batch.to(device), y_batch.to(device)

# Zero the parameter gradients


optimizer.zero_grad()

# Forward pass
outputs = model(x_batch)
# Flatten the outputs and y_batch tensors one dimension lower
outputs = [Link](-1, [Link][-1])
y_batch = y_batch.view(-1)

# Loss calculation
loss = criterion(outputs, y_batch).to(device)

# Backward pass and optimize


[Link]()
[Link]()

# Metrics
argmax_pred = [Link](axis=1)
[Link](argmax_pred, y_batch)

# Print statistics
running_loss += [Link]()
if i % 10 == 9: # update every 10 mini-batches
accuracy = [Link]().item()
epoch_accuracy += accuracy
pbar.set_postfix({'Loss': running_loss / (i + 1), 'Accuracy': accuracy})

# Update the progress bar


[Link](1)
# Save model weights periodically
if i % 10 == 9:
[Link](model.state_dict(), '/kaggle/working/model_weights.pth')

# Compute and print average loss and accuracy for the epoch
avg_loss = running_loss / num_batches
avg_accuracy = epoch_accuracy / (num_batches // 10) # since we're summing accuracy
print(f'Epoch {epoch + 1} - Loss: {avg_loss:.4f}, Accuracy: {avg_accuracy:.4f}')

print('Training Completed')

Training Started
Epoch 1: 66%|██████ | 2770/4166 [1:58:07<1:00:08, 2.59s/it, Loss=9.89, Accuracy=0.2
34]

In [ ]: #Link to download model


from [Link] import FileLink
FileLink(r'model_weights.pth')

Simplistic Inference Script


In [17]: text = '''Beyoncé Giselle Knowles-Carter (/biːˈjɒnseɪ/ bee-YON-say) (born September 4, 198

In [6]: def model_pred(tokens,temp):


[Link]()
with torch.no_grad():
pred = model(tokens,temp)
pred = [Link](-1, [Link][-1]).argmax(axis=1)
return pred

def tokenize_text(text):
seq_length = 300
q_tokens = tokenizer(text,add_special_tokens=False)['input_ids']
pad = [0 for i in range(seq_length-len(q_tokens))]
final_tokens = [q_tokens + pad]
last_index = len(q_tokens)-1

return [Link](final_tokens),last_index

def inference(text, starter='', temperature=1.0):


curr = 0
pred_list = []
t, last_token = tokenize_text(text + '[CLS]' + starter)
t = [Link](device)

while curr != 102:


print('\n',"Generating...")
all_pred = model_pred(t, temperature)
pred = all_pred[last_token].item()
pred_list.append(pred)
t[0][last_token + 1] = pred
last_token += 1
curr = pred

if curr > 10:


break
print("Question from the model: ".upper(), starter + ' ' + [Link](pred_lis

return starter + ' ' + [Link](pred_list)

inference(text, '')
QUESTION FROM THE MODEL: what is the name of the singer? [SEP]

Issues
Not Enough data: I trained on only 130K samples which is too small
Pre-trainig on a downstream task: Pretraining is supposed to be self supervised

Common questions

Powered by AI

The project employs several strategies to handle overfitting: 1) The dataset is randomized using random sampling to enhance data diversity. 2) Label smoothing is applied in the loss calculation to prevent the model from becoming overly confident in its predictions. 3) Regular monitoring of loss and accuracy per mini-batch helps to adjust training strategies dynamically. These methods collectively help in reducing overfitting by ensuring the model generalizes better and doesn't memorize training data .

The modified SQuAD dataset is used to train a model where the input is a context (C) and the output is a generated question (Q). This modification reverses the typical format where the SQuAD dataset usually maps questions and context to answers (Q+C --> A). By altering the data format to (C --> Q), the context now serves as the input from which the model generates a relevant question, as demonstrated in the Decoder-Only Transformer architecture .

Padding is utilized during input processing to ensure that all sequences are of uniform length, which is essential for batch processing in transformer models. In this particular model, sequences are padded to a length of 300 tokens, using zeroes where necessary. This uniformity allows the model to efficiently process data in batches, leveraging parallelism and maintaining the sequence alignment across different components of the model. Additionally, padding aids in the application of attention and ensures that the model treats sequences of varying lengths in a consistent manner .

A challenge that arises from sequence length limitations is that the model has a maximum sequence length it can handle (512 tokens in this specific case). Longer sequences can lead to indexing errors if not handled properly. In this project, the issue was addressed by setting a maximum sequence length to 300 tokens during tokenization, ensuring sequences fit within the model's capacity. Special tokens like [CLS] and [SEP] are used for demarcation, and zero padding is applied to ensure uniform input sizes, reducing indexing errors while maintaining model performance .

Multi-head attention is crucial because it allows the transformer model to simultaneously attend to information from different representation subspaces, enhancing its ability to focus on different parts of the input context at various scales. By using multiple heads, the model can capture a more comprehensive view of the input, which is necessary for generating diverse and contextually accurate questions. In this implementation, the model employs 8 decoder layers with multi-head attention, enhancing its capacity to process the complex interdependencies inherent in natural language tasks like question generation .

Checking for GPU availability is critical because transformer models, due to their complexity and massive computational demands, perform significantly faster on a GPU than on a CPU. The GPU accelerates model training by efficiently handling parallel computations, essential for large-scale models like transformers. This project implements a device check using PyTorch's torch.cuda.is_available() function, which ensures the model utilizes the GPU if available, falling back to the CPU otherwise, optimizing resource usage and training speed .

Residual connections play a crucial role in the transformer decoder architecture by allowing gradients to flow through the network without vanishing, thus improving model training. They provide shortcut paths for backpropagation, helping deeper networks converge more effectively. In this project, residual connections are used after multi-head attention and feed-forward layers, which help preserve input information and facilitate better learning dynamics, crucial for complex tasks like question generation .

Positional encodings are necessary because transformer models, unlike recurrent models, have no inherent understanding of the sequence order within the input context. The sinusoidal positional encoding method used in this project helps the model differentiate between different positions in a sequence, providing necessary information about the order of tokens. This enhances the model's ability to generate questions that appropriately relate to the sequential context of the given information, ensuring coherent and contextually relevant question formation .

Integrating multiple decoder layers in the model enhances its depth, thereby increasing its capacity to learn complex patterns from the input context. Each decoder layer, equipped with self-attention and feed-forward sub-layers, processes and refines the sequence representations, successively abstracting higher-level features. This cumulative process allows the model to generate questions that are more contextually accurate and semantically rich, better capturing the intricacies of the input text for effective question generation .

Pre-training on a downstream task can considerably enhance the performance of transformer models by providing them with a preliminary understanding of language that is refined on a more general dataset before fine-tuning on a specific task. In this project, the lack of sufficient pre-training was highlighted as an issue. Pre-training is supposed to be self-supervised, creating robust initial language representations that facilitate more efficient learning during task-specific training. Without extensive pre-training, the model trained with 130K samples may not fully capture the intricate language nuances needed for high-quality question generation, leading to compromised output .

You might also like