Knowledge Distillation in LLMs
Tanmoy Chakraborty
Associate Professor, IIT Delhi
[Link]
Advances in Large Language Models
Qwen-3-Next-80B-A3B Announced on
September 12, 2025
The first in the series of next-generation foundation models that are optimized for
Qwen-3-Next-Blog
extreme context length and large-scale parameter efficiency
Qwen-3-Next-80B-
Qwen3-Next-80B-A3B uses
A3B introduces several
a highly sparse MoE design,
architectural innovations to
having a total of 80 billion
maximize performance
parameters with only 3
while minimizing
billion activated, making it
computational cost. It uses
highly efficient. A thinking
a combination of Gated
version is also released
DeltaNet and Gated
along with the base model.
Attention, enabling efficient
context modeling for ultra-
long sequences.
Model Compression
Model Pruning Knowledge Distillation
Knowledge Distillation (KD): Types
Hinton et al., 2015 Teacher model generate
soft labels (logits)
Teacher generated logits
used to fine-tune the
student
Advances in Large Language Models
Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty
Knowledge Distillation (KD): Types
Hinton et al., 2015 Teacher model generate Distance between
soft labels (logits) layer-wise
representations
used to enforce
student model to
imitate the teacher
Teacher generated logits
used to fine-tune the
student
Advances in Large Language Models
Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty
Sun et al., 2019
Divergence and Similarity Functions
Advances in Large Language Models
Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty
Categories of KD
● White-box KD: Full access to the teacher’s internal components (logits, hidden
states, attention maps)
● Meta KD: Teacher helps guide student training strategies (e.g., data selection,
curriculum)
● Black-box KD: Only the final output of the teacher is available, e.g., via API
Advances in Large Language Models
Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty
KD for Language Models
Kim and Rush 2016 extended the idea to word-level and sequence-level KD for language
models, which aligns the student model with the teacher’s output distributions
Advances in Large Language Models
Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty
KD for Language Models
• Applied in sequence generation tasks (e.g., machine translation)
• Student model is trained using the teacher’s best decoded sequence (e.g., via beam
search)
Advantages:
• Instead of label sequences, student mimics the teacher's generation process
• Better for long-form tasks like summarization or machine translation
Disadvantages:
• Beam search is computationally expensive
• Generated sequences may propagate teacher’s errors
Advances in Large Language Models
Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty
KD for LLMs – MiniLLM (Gu et al. 2023)
• MiniLLM (Gu et al. 2023) replace the forward Kullback-Leibler divergence (KLD) objective in the
standard KD approaches with reverse KLD, which is more suitable for KD on generative language
models.
• This prevents the student model from overestimating the low-probability regions of the teacher
distribution.
Advances in Large Language Models
Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty
KD for LLMs – GKD (Agarwal et al. 2024)
• Current KD methods for auto-regressive sequence models suffer from distribution mismatch between
output sequences seen during training and those generated by the student during inference.
• Instead of solely relying on a fixed set of output sequences, GKD trains the student on its self-
generated output (SGO) sequences by leveraging feedback from the teacher on such sequences.
Advances in Large Language Models
Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty
Adaptive SGO for KD – DistiLLM (Ko et al. 2024)
• Generating SGO for each step can increase distillation time significantly. Ko et al., suggested an
adaptive method with replay buffer to adaptively determine when to generate SGO vs. when to use
original ground truth texts for distilling knowledge.
• KD optimization stability depends on the smoothness of the distillation loss objective. Ko et al.,
suggested a skewed divergence loss, where a mixture probability of teacher and student logits is used.
Advances in Large Language Models
Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty
Confidence-Concentrated Loss for KD – ABKD
(Wang et al. 2025)
• Traditional distillation loss functions – forward KLD and reverse KLD tackles two different properties –
while FKLD makes student distribution overly smoothened (higher recall), RKLD captures prominent
modes of the teacher (higher precision).
• Wang et al., proposed a weighted scheme between FKLD and RKLD, capturing the confidence and
hardness of teacher-student output probabilities. ABKD is a generalized variation of the popularly
used divergence-based loss functions used in KD.
Advances in Large Language Models
Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty
Limitations of Vanilla KD
Knowledge sharing is unidirectional, i.e., teacher is not aware of student’s
capacity
Advances in Large Language Models
Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty
Image reference: Xu et al., 2024
KD with Meta Learning – Zhou et al., 2022
• Traditional KD approaches are uni-directional, i.e., teacher is mostly trained prior to the KD process;
therefore, teacher is unaware of the student’s capacity.
• The teacher pre-training procedure is not optimized for distillation purposes; good model may not be
always a good teacher
• To address these challenges, Zhou et al., proposed a meta-KD method where the teacher model is
also trained in a meta loop, enabling better knowledge dissipation in the subsequent KD step.
Advances in Large Language Models
Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty
MPDistil: Student-Aware Meta Distillation: Learning to teach
● A healthy competition between the teacher
and student can encourage both the models
to perform better.
● A better teacher can set a higher
benchmark for the student, enhancing
student’s performance.
● The student can devise better learning
strategy (curriculum) to perform better than
the teacher.
Sengupta, Dixit, Akhtar, Chakraborty. A Good Learner Can Teach Better: Teacher-Student Collaborative Knowledge Distillation. ICLR 2024.
Advances in Large Language Models
Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty
MPDistil: Step 1 -- Teacher Fine-tuning
1. Teacher Fine-tuning
Advances in Large Language Models
Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty
MPDistil: Step 2 -- Student Distillation
1. Teacher Fine-tuning 2. Student Distillation
Advances in Large Language Models
Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty
MPDistil: Step 3 -- Meta-teacher Learning
1. Teacher Fine-tuning 2. Student Distillation 3. Teacher Meta Learning
(on a quiz dataset)
Collaborative Loss Competitive Loss
Intuition: The meta-teacher obtains the hidden states from both teacher and student and creates a healthy
Advances in Large Language Models
competition
Advances inbetween
Large Languagethe models.
Models Tanmoy Chakraborty Tanmoy Chakraborty
MPDistil: Step 4 -- Student Curriculum Learning
1. Teacher Fine-tuning 2. Student Distillation 3. Teacher Meta Learning
(on a quiz dataset)
4. Student Curriculum
Learning
Why Curriculum Learning in KD?
In real world, a student might aim to improve her
understanding of Physics by studying selected concepts
from Mathematics.
Advances in Large Language Models
Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty
MPDistil: Step 4 -- Student Curriculum Learning
1. Teacher Fine-tuning 2. Student Distillation 3. Teacher Meta Learning
(on a quiz dataset)
4. Student Curriculum
Learning
Competing student tries to beat the teacher
A policy network selects
optimal curriculum to
fine-tune the student by
maximizing the reward
Advances in Large Language Models
Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty
ICLR’24
A “smart” student can beat a teach!!
Positive value
indicates the
student model is
better than the
teacher model
Advances in Large Language Models
Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty
Sengupta et al. ACL (Findings) 2025
Explaining Knowledge Distillation
Known: KD improves generalization abilities of student models.
Questions
(i) Post-KD, does student perfectly imitate a teacher?
(ii) What are the key drivers influencing the effectiveness of KD methods?
Advances in Large Language Models
Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty
Agreement b/w Teacher-Student Post-KD
Agreement: Overlap between the final output generated by teacher and students.
Teacher-student agreement improves post KD, mostly for smaller LMs (<7B).
Advances in Large Language Models
Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty
Fidelity b/w Teacher-Student Post-KD
•Fidelity: Ability of the student to imitate the teacher’s reasoning behaviors.
• Smaller LMs tend to have better fidelity post-KD.
• However, statistical tests show that fidelity does not necessarily improve the generalization
abilities of student models!!
Advances in Large Language Models
Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty
Fidelity vs Generalization Paradox of KD
•High teacher-student
fidelity, but wrong answer
predicted by student (poor
generalization)
Advances in Large Language Models
Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty
Fidelity vs Generalization Paradox of KD
•Low teacher-student
fidelity, but good
generalization
Therefore, the tradeoff between generalization vs fidelity-agreement remains prominent.
Advances in Large Language Models
Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty
Drivers behind Successful KD
1. Teacher model should be task-aware
Advances in Large Language Models
Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty
Drivers behind Successful KD
1. Teacher model should be task-aware 2. Teacher signals to student should be
noise-free.
Here 𝜎 is the amount of Gaussian noise added to the
teacher logits before distilling to student. For 𝜎, student
performance drops drastically.
Teacher model performance minimally affects student outcomes; however, the
teacher’s task-specific expertise isTanmoy
crucial
Advances in Large Language Models
Advances in Large Language Models Chakraborty Tanmoy Chakraborty
Drivers behind Successful KD
1. Teacher model should be task-aware 2. Teacher signals to student should be 3. Logit smoothing is important
noise-free. Here 𝜏 is the temperature used to
Here 𝜎 is the amount of Gaussian noise added to the smoothen the teacher logits. Too much
teacher logits before distilling to student. For 𝜎, student smoothing hurts student performance, but
performance drops drastically. moderate smoothing shows benefit.
Temperature (𝜏) in KD balances precision (𝜏 ↓) and recall (𝜏 ↑) of the student model.
Advances in Large Language Models
Advances in Large Language Models Tanmoy Chakraborty Tanmoy Chakraborty