0% found this document useful (0 votes)
10 views30 pages

Distillation

The document discusses advancements in knowledge distillation (KD) for large language models (LLMs), highlighting various methods such as white-box, black-box, and meta KD. It emphasizes the importance of teacher-student dynamics in the KD process, including the need for task-aware teacher models and noise-free signals. Additionally, it explores recent innovations like MiniLLM, GKD, and MPDistil, which aim to enhance the efficiency and effectiveness of knowledge transfer between models.

Uploaded by

integration330
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
10 views30 pages

Distillation

The document discusses advancements in knowledge distillation (KD) for large language models (LLMs), highlighting various methods such as white-box, black-box, and meta KD. It emphasizes the importance of teacher-student dynamics in the KD process, including the need for task-aware teacher models and noise-free signals. Additionally, it explores recent innovations like MiniLLM, GKD, and MPDistil, which aim to enhance the efficiency and effectiveness of knowledge transfer between models.

Uploaded by

integration330
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Knowledge Distillation in LLMs

Tanmoy Chakraborty
Associate Professor, IIT Delhi
[Link]

Advances in Large Language Models


Qwen-3-Next-80B-A3B Announced on
September 12, 2025
The first in the series of next-generation foundation models that are optimized for
Qwen-3-Next-Blog
extreme context length and large-scale parameter efficiency

Qwen-3-Next-80B-
Qwen3-Next-80B-A3B uses
A3B introduces several
a highly sparse MoE design,
architectural innovations to
having a total of 80 billion
maximize performance
parameters with only 3
while minimizing
billion activated, making it
computational cost. It uses
highly efficient. A thinking
a combination of Gated
version is also released
DeltaNet and Gated
along with the base model.
Attention, enabling efficient
context modeling for ultra-
long sequences.
Model Compression

Model Pruning Knowledge Distillation


Knowledge Distillation (KD): Types
Hinton et al., 2015 Teacher model generate
soft labels (logits)

Teacher generated logits


used to fine-tune the
student
Advances in Large Language Models

Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty


Knowledge Distillation (KD): Types
Hinton et al., 2015 Teacher model generate Distance between
soft labels (logits) layer-wise
representations
used to enforce
student model to
imitate the teacher

Teacher generated logits


used to fine-tune the
student
Advances in Large Language Models

Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty


Sun et al., 2019
Divergence and Similarity Functions

Advances in Large Language Models

Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty


Categories of KD

● White-box KD: Full access to the teacher’s internal components (logits, hidden
states, attention maps)

● Meta KD: Teacher helps guide student training strategies (e.g., data selection,
curriculum)

● Black-box KD: Only the final output of the teacher is available, e.g., via API

Advances in Large Language Models

Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty


KD for Language Models
Kim and Rush 2016 extended the idea to word-level and sequence-level KD for language
models, which aligns the student model with the teacher’s output distributions

Advances in Large Language Models

Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty


KD for Language Models
• Applied in sequence generation tasks (e.g., machine translation)
• Student model is trained using the teacher’s best decoded sequence (e.g., via beam
search)

Advantages:
• Instead of label sequences, student mimics the teacher's generation process
• Better for long-form tasks like summarization or machine translation

Disadvantages:
• Beam search is computationally expensive
• Generated sequences may propagate teacher’s errors

Advances in Large Language Models

Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty


KD for LLMs – MiniLLM (Gu et al. 2023)
• MiniLLM (Gu et al. 2023) replace the forward Kullback-Leibler divergence (KLD) objective in the
standard KD approaches with reverse KLD, which is more suitable for KD on generative language
models.
• This prevents the student model from overestimating the low-probability regions of the teacher
distribution.

Advances in Large Language Models

Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty


KD for LLMs – GKD (Agarwal et al. 2024)
• Current KD methods for auto-regressive sequence models suffer from distribution mismatch between
output sequences seen during training and those generated by the student during inference.

• Instead of solely relying on a fixed set of output sequences, GKD trains the student on its self-
generated output (SGO) sequences by leveraging feedback from the teacher on such sequences.

Advances in Large Language Models

Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty


Adaptive SGO for KD – DistiLLM (Ko et al. 2024)
• Generating SGO for each step can increase distillation time significantly. Ko et al., suggested an
adaptive method with replay buffer to adaptively determine when to generate SGO vs. when to use
original ground truth texts for distilling knowledge.

• KD optimization stability depends on the smoothness of the distillation loss objective. Ko et al.,
suggested a skewed divergence loss, where a mixture probability of teacher and student logits is used.

Advances in Large Language Models

Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty


Confidence-Concentrated Loss for KD – ABKD
(Wang et al. 2025)
• Traditional distillation loss functions – forward KLD and reverse KLD tackles two different properties –
while FKLD makes student distribution overly smoothened (higher recall), RKLD captures prominent
modes of the teacher (higher precision).
• Wang et al., proposed a weighted scheme between FKLD and RKLD, capturing the confidence and
hardness of teacher-student output probabilities. ABKD is a generalized variation of the popularly
used divergence-based loss functions used in KD.

Advances in Large Language Models

Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty


Limitations of Vanilla KD

Knowledge sharing is unidirectional, i.e., teacher is not aware of student’s


capacity
Advances in Large Language Models

Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty


Image reference: Xu et al., 2024
KD with Meta Learning – Zhou et al., 2022
• Traditional KD approaches are uni-directional, i.e., teacher is mostly trained prior to the KD process;
therefore, teacher is unaware of the student’s capacity.
• The teacher pre-training procedure is not optimized for distillation purposes; good model may not be
always a good teacher
• To address these challenges, Zhou et al., proposed a meta-KD method where the teacher model is
also trained in a meta loop, enabling better knowledge dissipation in the subsequent KD step.

Advances in Large Language Models

Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty


MPDistil: Student-Aware Meta Distillation: Learning to teach

● A healthy competition between the teacher


and student can encourage both the models
to perform better.

● A better teacher can set a higher


benchmark for the student, enhancing
student’s performance.

● The student can devise better learning


strategy (curriculum) to perform better than
the teacher.

Sengupta, Dixit, Akhtar, Chakraborty. A Good Learner Can Teach Better: Teacher-Student Collaborative Knowledge Distillation. ICLR 2024.

Advances in Large Language Models

Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty


MPDistil: Step 1 -- Teacher Fine-tuning
1. Teacher Fine-tuning

Advances in Large Language Models

Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty


MPDistil: Step 2 -- Student Distillation
1. Teacher Fine-tuning 2. Student Distillation

Advances in Large Language Models

Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty


MPDistil: Step 3 -- Meta-teacher Learning
1. Teacher Fine-tuning 2. Student Distillation 3. Teacher Meta Learning
(on a quiz dataset)

Collaborative Loss Competitive Loss

Intuition: The meta-teacher obtains the hidden states from both teacher and student and creates a healthy
Advances in Large Language Models

competition
Advances inbetween
Large Languagethe models.
Models Tanmoy Chakraborty Tanmoy Chakraborty
MPDistil: Step 4 -- Student Curriculum Learning
1. Teacher Fine-tuning 2. Student Distillation 3. Teacher Meta Learning
(on a quiz dataset)

4. Student Curriculum
Learning

Why Curriculum Learning in KD?

In real world, a student might aim to improve her


understanding of Physics by studying selected concepts
from Mathematics.
Advances in Large Language Models

Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty


MPDistil: Step 4 -- Student Curriculum Learning
1. Teacher Fine-tuning 2. Student Distillation 3. Teacher Meta Learning
(on a quiz dataset)

4. Student Curriculum
Learning
Competing student tries to beat the teacher
A policy network selects
optimal curriculum to
fine-tune the student by
maximizing the reward
Advances in Large Language Models

Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty


ICLR’24
A “smart” student can beat a teach!!

Positive value
indicates the
student model is
better than the
teacher model
Advances in Large Language Models

Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty


Sengupta et al. ACL (Findings) 2025

Explaining Knowledge Distillation

Known: KD improves generalization abilities of student models.

Questions

(i) Post-KD, does student perfectly imitate a teacher?


(ii) What are the key drivers influencing the effectiveness of KD methods?

Advances in Large Language Models

Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty


Agreement b/w Teacher-Student Post-KD

Agreement: Overlap between the final output generated by teacher and students.
Teacher-student agreement improves post KD, mostly for smaller LMs (<7B).
Advances in Large Language Models

Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty


Fidelity b/w Teacher-Student Post-KD

•Fidelity: Ability of the student to imitate the teacher’s reasoning behaviors.


• Smaller LMs tend to have better fidelity post-KD.
• However, statistical tests show that fidelity does not necessarily improve the generalization
abilities of student models!!
Advances in Large Language Models

Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty


Fidelity vs Generalization Paradox of KD

•High teacher-student
fidelity, but wrong answer
predicted by student (poor
generalization)

Advances in Large Language Models

Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty


Fidelity vs Generalization Paradox of KD

•Low teacher-student
fidelity, but good
generalization

Therefore, the tradeoff between generalization vs fidelity-agreement remains prominent.

Advances in Large Language Models

Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty


Drivers behind Successful KD

1. Teacher model should be task-aware

Advances in Large Language Models

Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty


Drivers behind Successful KD

1. Teacher model should be task-aware 2. Teacher signals to student should be


noise-free.

Here 𝜎 is the amount of Gaussian noise added to the


teacher logits before distilling to student. For 𝜎, student
performance drops drastically.

Teacher model performance minimally affects student outcomes; however, the


teacher’s task-specific expertise isTanmoy
crucial
Advances in Large Language Models

Advances in Large Language Models Chakraborty Tanmoy Chakraborty


Drivers behind Successful KD

1. Teacher model should be task-aware 2. Teacher signals to student should be 3. Logit smoothing is important
noise-free. Here 𝜏 is the temperature used to
Here 𝜎 is the amount of Gaussian noise added to the smoothen the teacher logits. Too much
teacher logits before distilling to student. For 𝜎, student smoothing hurts student performance, but
performance drops drastically. moderate smoothing shows benefit.

Temperature (𝜏) in KD balances precision (𝜏 ↓) and recall (𝜏 ↑) of the student model.


Advances in Large Language Models

Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty

You might also like