0% found this document useful (0 votes)
7 views11 pages

Data-Free Knowledge Distillation for DMs

The paper presents a novel approach called Data-Free Knowledge Distillation for Diffusion Models (DKDM) that enables the training of new diffusion models without requiring access to large datasets. By leveraging existing pretrained diffusion models as data sources, DKDM aims to alleviate the challenges of high data acquisition costs and storage expenses associated with traditional training methods. Experimental results indicate that this data-free approach achieves competitive generative performance, and in some cases, outperforms models trained with full datasets.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views11 pages

Data-Free Knowledge Distillation for DMs

The paper presents a novel approach called Data-Free Knowledge Distillation for Diffusion Models (DKDM) that enables the training of new diffusion models without requiring access to large datasets. By leveraging existing pretrained diffusion models as data sources, DKDM aims to alleviate the challenges of high data acquisition costs and storage expenses associated with traditional training methods. Experimental results indicate that this data-free approach achieves competitive generative performance, and in some cases, outperforms models trained with full datasets.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

This CVPR paper is the Open Access version, provided by the Computer Vision Foundation.

Except for this watermark, it is identical to the accepted version;


the final published version of the proceedings is available on IEEE Xplore.

DKDM: Data-Free Knowledge Distillation for Diffusion Models with Any


Architecture

Qianlong Xiang1 , Miao Zhang1,†, Yuzhang Shang2 , Jianlong Wu1 , Yan Yan3 , Liqiang Nie1,†
1
Harbin Institute of Technology (Shenzhen) 2 Illinois Institute of Technology 3 University of Illinois Chicago
[Link]

Abstract Diffusion Models

Diffusion models (DMs) have demonstrated exceptional


generative capabilities across various domains, including
image, video, and so on. A key factor contributing to their
effectiveness is the high quantity and quality of data used Standard Highly
DKDM Data-Free
during training. However, mainstream DMs now consume Training Data-Dependent
increasingly large amounts of data. For example, train-
ing a Stable Diffusion model requires billions of image-text
pairs. This enormous data requirement poses significant
PAPHAEL
challenges for training large DMs due to high data acquisi- Pretrain DALL·E 2
tion costs and storage expenses. To alleviate this data bur- LDM
den, we propose a novel scenario: using existing DMs as SD 1.5 ADM
data sources to train new DMs with any architecture. We Imagen
refer to this scenario as Data-Free Knowledge Distillation ……
GLIDE
for Diffusion Models (DKDM), where the generative abil-
ity of DMs is transferred to new ones in a data-free man- Large-Scale Datasets Existing Models
ner. To tackle this challenge, we make two main contribu-
tions. First, we introduce a DKDM objective that enables Figure 1. Illustration of our DKDM concept: utilizing pretrained
diffusion models to train new ones, thus avoiding the high costs
the training of new DMs via distillation, without requiring
associated with increasingly large datasets.
access to the data. Second, we develop a dynamic itera-
tive distillation method that efficiently extracts time-domain
knowledge from existing DMs, enabling direct retrieval of and audio [4, 17, 61]. One reason for their superior perfor-
training data without the need for a prolonged generative mance is their training on large-scale, high-quality datasets.
process. To the best of our knowledge, we are the first to ex- However, this advantage also entails a drawback: train-
plore this scenario. Experimental results demonstrate that ing DMs requires substantial storage capacity, as shown in
our data-free approach not only achieves competitive gen- Tab. 1. For instance, training a Stable Diffusion model ne-
erative performance but also, in some instances, outper- cessitates the use of billions of image-text pairs [44].
forms models trained with the entire dataset. To alleviate this data burden, considering that numerous
pretrained DMs have been trained and released by various
organizations, we pose a novel question:
1. Introduction
Can we train new diffusion models by using existing
The advent of Diffusion Models (DMs) [16, 51, 53] heralds pretrained diffusion models as the data source, thereby
a new era in the generative domain, garnering widespread eliminating the need to access or store any dataset?
acclaim for their exceptional capability in producing sam-
Fig. 1 illustrates the concept of this scenario. Traditionally,
ples of remarkable quality [8, 36, 44]. These models have
training DMs requires access to large datasets. In contrast,
rapidly ascended to a pivotal role across a spectrum of gen-
in this paper, we explore how to utilize existing DMs to
erative applications, notably in the fields of image, video
train new models without any data. We formalize this sce-
† Corresponding authors nario as the Data-Free Knowledge Distillation for Diffusion

2955
Model #Param. #Images traditional DMs, as described by Ho et al. [16], is inappro-
priate due to the absence of the data. To address this, we
GLIDE [35] 5.0B 5.94B specially design a DKDM objective that aligns closely with
LDM [44] 1.5B 0.27B the original DM optimization objective, while the architec-
DALL·E 2 [40] 5.5B 5.63B ture of the model is no longer limited as in other distilla-
Imagen [45] 3.0B 15.36B tion methods. For the latter, we observe that compared
eDiff-I [2] 9.1B 11.47B to realistic samples, time-domain ones corrupted by certain
Stable Diffusion v1.5 [44] 0.9B 3.16B noise are more relevant to the optimization objective for
DMs. Therefore, we define the knowledge form in DKDM
Table 1. Comparison of prominent diffusion models on parameter
as these noisy samples, enabling direct learning from each
count and training dataset size, sourced from Kang et al. [19].
denoising step of pretrained DMs, without the need for a
time-consuming generative process to obtain realistic sam-
ples. In other words, the student model learns from the gen-
Models (DKDM) paradigm, which aims at transferring the
erative process of the pretrained DMs rather than from their
generative ability of the pretrained DMs towards new ones.
final generative outputs. Based on this definition, we pro-
Compared with previous work, our proposed DKDM
pose a dynamic iterative distillation method that generates
paradigm imposes strict requirements in three aspects.
substantial and diverse knowledge to enhance the training
∂ Data. Previous work usually requires access to datasets
of the student.
to train DMs for purposes such as model compression
To sum up, this paper introduces a novel method for
[62, 66] and reducing denoising steps [46, 54]. In con-
training DMs without the need for datasets, by leveraging
trast, DKDM mandates that the entire training process must
existing pretrained DMs as the data source. Experimen-
not access any datasets. This constraint eliminates the need
tal results indicate that models trained with our approach
to spend significant time downloading and storing datasets
demonstrate competitive generative performance. Further-
and helps circumvent data privacy issues, especially when
more, in some cases, our data-free method even outper-
training data is not released [69]. ∑ Architecture. We ob-
forms models trained with the entire dataset.
serve that previous work on knowledge distillation for DMs
often initializes student models with the architectures and
2. Preliminaries on Diffusion Models
weights of teacher models, limiting architectural flexibility.
One reason for this is to improve performance. For exam- In diffusion models [16], a Markov chain is defined to add
ple, Xie et al. [58] proposed a distillation method for DMs noises to data, and then diffusion models learn the reverse
and found that the performance will degrade when the stu- process to generate data from noises.
dent model is randomly initialized. On the contrary, DKDM Forward Process. Given a sample x0 ⇠ q x0 from
calls for training DMs with any architecture. ∏ Knowledge the data distribution, the forward process iteratively adds
Form. Leveraging deep generative models to synthesize Gaussian noise for T diffusion steps with the predefined
high-quality samples [1, 34, 57] for performance enhance- noise schedule ( 1 , . . . , T ):
ment on downstream tasks [25, 26, 42, 55, 56, 67, 68] is ⇣ p ⌘
a common practice. However, we argue that in DKDM, q xt |xt 1 = N xt ; 1 t x t 1
, t I , (1)
the knowledge form should not be realistic samples, be- T
cause generating and storing such samples requires enor- Y
q x1:T |x0 = q xt |xt 1
, (2)
mous space and time. For instance, to train a model like t=1
Stable Diffusion in this way, we would need to use the
teacher model to synthesize billions of image-text pairs in until a completely noise xT ⇠ N (0, I) is obtained. Ac-
advance and then use this massive synthetic dataset to train cording to Ho et al. [16], adding noise t times sequentially
the new model, which is impractical. Therefore, the knowl- to the original sample x0 to generate a noisy sample xt can
edge form in DKDM should be carefully designed. be simplified to a one-step calculation as follows:
Based on the above considerations, we summarize the p
requirements brought by the DKDM paradigm into two key q xt |x0 = N xt ; ↵ ¯ t x0 , (1 ↵ ¯t) I , (3)
challenges and solve them separately. The first challenge p p
xt = ↵ ¯ t x0 + 1 ↵ ¯ t ✏, (4)
involves training DMs with any architecture, while not ac- Qt
cessing the dataset. The second challenge involves effi- where ↵t := 1 t, ↵
¯ t := s=0 ↵s and ✏ ⇠ N (0, I).
ciently designing the knowledge form for distillation, pre- Reverse Process. The posterior q(xt 1 |xt ) depends on
venting it from becoming the main bottleneck in slowing the data distribution, which is tractable conditioned on x0 :
the training process, as the generation of DMs is inherently ⇣ ⌘
slow. For the former, the optimization objective used in q xt 1 |xt , x0 = N xt 1 ; µ̃ xt , x0 , ˜t I , (5)

2956
where µ̃t xt , x0 and ˜t can be calculated by: Student DMs with any architecture Pre-trained DMs

˜t := 1 ↵
¯t 1
t, (6)
1 ↵ ¯t
DDPM DDPM DKDM
p p Objective Objective Objective

¯t 1 t 0 ↵t (1 ↵¯t 1)
µ̃t xt , x0 := x + xt . (7) "!~! +!~! "
"
1 ↵ ¯t 1 ↵¯t +$ , $ ~ℬ#
"
$~ 1,1000 $~ 1,1000
Since x0 in the data is not accessible during generation, a (~) 0, * (~) 0, *
neural network parameterized by ✓ is used for approxima-
tion:
Dataset ! " Collection of
Dataset !
p✓ xt 1
|xt = N xt 1
; µ✓ xt , t , ⌃✓ xt , t I . (8) ( !" = ! ) Knowledge ℬ#

Optimization. To optimize this network, the variational


bound on negative log likelihood E[ log p✓ ] is estimated
by:
(a) Data-Based (b) Data-Free (c) Our DKDM
⇥ ⇤
Lvlb = Ex0 ,✏,t DKL (q(xt 1 |xt , x0 )||p✓ (xt 1 |xt ) .
(9) Figure 2. Illustration of our DKDM Paradigm. (a): standard data-
based training of DMs. (b): a straightforward data-free training
Ho et al. [16] found that predicting ✏ is a more efficient
approach. (c): our proposed framework for DKDM.
way when parameterizing µ✓ (xt , t) in practice, which can
be derived by Eqs. (4) and (7):
✓ ◆ (DKDM). Sec. 3.1 details the DKDM paradigm, focusing
1 t
µ✓ xt , t = p xt p ✏✓ xt , t . (10) on two principal challenges: the formulation of the opti-
↵t 1 ↵ ¯t
mization objective and the acquisition of knowledge for dis-
Thus, a reweighted loss function is designed as the ob- tillation. Sec. 3.2 describes our proposed optimization ob-
jective to optimize Lvlb : jective tailored for DKDM. Sec. 3.3 details our proposed
h i method for efficient retrieval of knowledge.
2
Lsimple = Ex0 ,✏,t ✏ ✏✓ xt , t . (11)
3.1. DKDM Paradigm
Improvement. In original DDPMs, Lsimple offers no The DKDM paradigm represents a novel scenario for train-
signal for learning ⌃✓ (xt , t) and Ho et al. [16] fixed it to t ing DMs. Unlike traditional methods, DKDM aims to lever-
or ˜t . Nichol and Dhariwal [36] found it to be sub-optimal age existing DMs as the data source to train new ones with
and proposed to parameterize ⌃✓ (xt , t) as a neural network any architecture, which eliminates the need for access to
whose output v is interpolated as: large or proprietary datasets.
⇣ ⌘ In standard data-based training of DMs, as depicted in
⌃✓ xt , t = exp v log t + (1 v) log ˜t . (12) Fig. 2a, a sample x0 ⇠ D is selected along with a timestep
t ⇠ [1, 1000] and random noise ✏ ⇠ N (0, I). The input
To optimize ⌃✓ (xt , t), Nichol and Dhariwal [36] use xt is computed using Eq. (4), and the denoising network is
Lvlb , in which a stop-gradient is applied to the µ✓ (xt , t) optimized according to Eq. (13) to generate outputs close
because it is optimized by Lsimple . The final hybrid objec- to ✏. However, without dataset access, DKDM cannot ob-
tive is defined as: tain training data (xt , t, ✏) to employ this standard method.
A straightforward data-free training approach, depicted in
Lhybrid = Lsimple + Lvlb , (13)
Fig. 2b, involves using DMs pretrained on D to generate a
where is used for balance between the two objectives. The synthetic dataset D0 , which is then used to train new DMs
process of training and sampling are guided by Eq. (13), cf . with varying architectures. Despite its simplicity, creating
Algorithms 2 and 3 in Sec. 7. D0 is time-intensive and impractical for large datasets.
While data-based training necessitates access to large-
3. Data-Free Knowledge Distillation for Diffu- scale datasets, data-free training incurs significant costs
sion Models in generating synthetic datasets. To address these chal-
lenges, we propose an effective and efficient framework for
In this section, we introduce a novel paradigm, termed DKDM, outlined in Fig. 2c, which incorporates a DKDM
Data-Free Knowledge Distillation for Diffusion Models Objective (described in Sec. 3.2) and a strategy for collect-

2957
⇥ ⇤
ing knowledge Bi (detailed in Sec. 3.3). This framework L0vlb = Ex̂t ,t DKL (p✓T (x̂t 1
|x̂t )kp✓S (x̂t 1
|x̂t ) . (18)
mitigates the challenges of distillation without datasets and By this formulation, the need for x0 in LDKDM is re-
reduces the costs associated with data-free training. moved by naturally leveraging the generative ability of the
3.2. DKDM Objective teacher. Optimized by the proposed LDKDM , the student
progressively learns the entire reverse diffusion process
Given a dataset D, the original optimization objective for a
from the teacher without reliance on the source datasets.
DM with parameters ✓ involves minimizing the KL diver-
However, the removal of the diffusion posterior and prior
gence Ex0 ,✏,t [DKL (q(xt 1 |xt , x0 )kp✓ (xt 1 |xt ))]. Our
in the DKDM objective introduces a significant bottleneck,
proposed DKDM objective comprises two primary goals:
resulting in notably slow learning rates. As depicted in
(1) eliminating the diffusion posterior q(xt 1 |xt , x0 ) and
Fig. 2a, standard training for DMs enables straightforward
(2) removing the diffusion prior xt ⇠ q(xt |x0 ) from the
acquisition of noisy samples xtii at an arbitrary diffusion
KL divergence, since they both are dependent on x0 ⇠ D.
step t ⇠ [1, T ] using Eq. (4). These samples are compiled
Eliminating the diffusion posterior q(xt 1 |xt , x0 ).
into a batch Bj = {xtii }, with j representing the training
In our framework, we introduce a teacher DM with param-
iteration. Conversely, our DKDM objective requires ob-
eters ✓T , trained on dataset D. This model can generate
taining a noisy sample x̂ti = G✓T (T ti ) through T ti
samples that conform to the learned distribution D0 . Opti-
denoising steps. Consequently, by considering the denois-
mized with the objective Eq. (13), the distribution D0 within
ing steps as the primary computational expense, the worst-
a well-learned teacher DM closely matches D. Our goal is
case time complexity of assembling a batch B̂j = {x̂tii } for
for a student, parameterized by ✓S , to replicate D0 instead
distillation is O(T b), where b denotes the batch size. This
of D, thereby obviating the need for q during optimization.
complexity significantly hinders the training process. To
Specifically, the pretrained teacher DM was optimized
address this issue, we introduce a method called dynamic
via the hybrid objective Eq. (13), which indicates that both
iterative distillation, detailed in Sec. 3.3.
the KL divergence DKL (q(xt 1 |xt , x0 )kp✓T (xt 1 |xt ))
and the mean squared error Ext ,✏,t [k✏ ✏✓T (xt , t)k2 ] 3.3. Efficient Collection of Knowledge
are minimized. Given the similarity in distribution be-
In this section, we present our efficient strategy for gather-
tween the teacher model and the dataset, we propose
ing knowledge for distillation, illustrated in Fig. 3. We be-
a DKDM objective that optimizes the student model
gin by introducing a basic iterative distillation method that
through minimizing DKL (p✓T (xt 1 |xt )kp✓S (xt 1 |xt ))
allows the student to learn from the teacher at each denois-
and Ext [k✏✓T (xt , t) ✏✓S (xt , t)k2 ]. This objective indi-
ing step, instead of requiring the teacher to denoise multi-
rectly minimizes DKL (q(xt 1 |xt , x0 )kp✓S (xt 1 |xt )) and
ple times within every training iteration to create a batch of
Ex0 ,✏,t [k✏ ✏✓S (xt , t)k2 ], despite the inaccessibility of the
noisy samples. Subsequently, to enhance the diversity of
posterior. The proposed DKDM objective is as follows:
noise levels within the batch samples, we develop an ad-
LDKDM = L0simple + L0vlb , (14) vanced method termed shuffled iterative distillation, which
allows the student to learn denoising patterns across varying
where L0simple guides the learning of µ✓S and L0vlb opti-
time steps. Lastly, we refine our approach to dynamic iter-
mizes ⌃✓S , as defined in following equations:
⇥ ⇤ ative distillation, significantly augmenting the diversity of
L0simple = Ex0 ,✏,t k✏✓T (xt , t) ✏✓S (xt , t)k2 , (15) data in the batch. This adaptation ensures that the student
⇥ ⇤ acquires knowledge from a broader array of samples over
L0vlb = Ex0 ,✏,t DKL (p✓T (xt 1 |xt )kp✓S (xt 1 |xt ) , time, avoiding repetitive learning from identical samples.
(16) Iterative Distillation. We introduce a method called it-
where q(xt 1 |xt , x0 ) is eliminated whereas the term xt ⇠ erative distillation, which closely aligns the optimization
q(xt |x0 ) remains to be removed. process with the generation procedure. In this approach,
Removing the diffusion prior q(xt |x0 ). Considering the teacher model consistently denoises, while the student
the generative ability of the teacher model, we utilize it to model continuously learns from this denoising. Each output
generate x̂t as a substitute for xt ⇠ q(xt |x0 ). We define from the teacher’s denoising step is incorporated into some
a reverse diffusion step x̂t 1 ⇠ p✓T (x̂t 1 |xt ) through the batch for optimization, ensuring the student model learns
equation x̂t 1 = g✓T (xt , t). Next, we represent a sequence from every output. Specifically, during each training itera-
of t reverse diffusion steps starting from T as G✓T (t). Note tion, the teacher performs g✓T (xt , t), which is a single-step
that G✓T (0) = ✏ where ✏ ⇠ N (0, I). For instance, G✓T (2) denoising, instead of G✓T (t), which would involve t-step
yields x̂T 2 = g✓T (g✓T (✏, T ), T 1). Consequently, x̂t denoising. Initially, a batch B̂1 = {x̂Ti } is formed from a
is obtained by x̂t = G✓T (T t) and the objectives L0simple set of sampled noises x̂Ti ⇠ N (0, I). After one step of
and L0vlb are reformulated as follows: distillation, the batch B̂2 = {x̂Ti 1 } is used for training.
⇥ ⇤
L0simple = Ex̂t ,t k✏✓T (x̂t , t) ✏✓S (x̂t , t)k2 , (17) This process is iterated until B̂T = {x̂1i } is reached, indi-

2958
!"% ~$(0, ()
" Batch set Batch set
Iterative Distillation
!!"
" !!!
"
& &
!!!
"
Batch Batch

!"#
" !#"
"
&
!#"
"
&
!#"
"
& '!
!#"
"
& '!

Shuffle Denoise Random Denoise Substitution


& &
!"$
" !$#
" !$#
" & '!
!$#
" !"$
"
Selection

!"$(!
" Teacher &#$!
!$(!
"
&
#$!
!$(!
" "
&#$!
!$(!
'! &#$!
!$(!
"
'!

Train
!")
" &
!)%
"
&
!)%
"
Student

Figure 3. Dynamic Iterative Distillation: An enlarged batch set is initially constructed by sampling from a Gaussian distribution. Next,
shuffle denoise is applied, wherein each sample is denoised random times. A batch is then randomly selected from this enlarged set for
training the student with the denoised results substituting for their counterparts in the batch set. This process is repeated iteratively.

cating that the batch has nearly become real samples with no Algorithm 1 Dynamic Iterative Distillation
noise. The cycle then restarts with the resampling of noise
Require: B̂0+ = {x̂Ti }
to form a new batch B̂T +1 = {x̂Ti }. This method allows + t
1: Get B̂1 = {x̂ii } with shuffle denoise, j = 0
the teacher model to provide an endless stream of data for
2: repeat
distillation. To further improve the diversity of the synthetic
3: j =j+1
batch B̂j = {x̂tii }, we investigate it from the perspectives
4: get B̂js from B̂j+ through random selection
of noise level ti and sample x̂i .
5: compute L?simple using Eq. (20)
Shuffled Iterative Distillation. Unlike the standard 6: compute L?vlb using Eqs. (12) and (21)
data-based training, the t values in an iterative distillation 7: take a gradient descent step on r✓ LDKDM
batch remain the same and do not follow a uniform distri- 8:
+
update B̂j+1
bution, resulting in significant instability during distillation. 9: until converged
To mitigate this issue, we integrate a method termed shuf-
fle denoise into our iterative distillation. Initially, a batch
B̂0s = {x̂Ti } is sampled from a Gaussian distribution. Sub- named dynamic iterative distillation. As shown in Fig. 3,
sequently, each sample undergoes random denoising steps, this method employs shuffle denoise to construct an en-
resulting in B̂1s = {x̂tii }, with ti following a uniform distri- larged batch set B̂1+ = {x̂tii }, where size |B̂j+ | = ⇢T |B̂js |,
bution. This batch, B̂1s , then initiates the iterative distillation
where ⇢ is a scaling factor. During distillation, a subset B̂js
process. By enhancing the diversity in the ti values within
the batch, this method balances the impact of different t val- is sampled from B̂j+ through random selection for optimiza-
ues during distillation. tion. The one-step denoised samples replace their counter-
+
parts in B̂j+1 . This method only has a time complexity of
Dynamic Iterative Distillation. There is a notable dis-
O(b) and significantly improves distillation performance.
tinction between standard training and iterative distillation
The final DKDM objective is defined as:
regarding the flexibility in batch composition. Consider two
samples, x̂1 and x̂2 , within a batch without differentiating L?DKDM = L?simple + L?vlb , (19)
their noise level. During standard training, the pairing of x̂1
⇥ ⇤
and x̂2 is entirely random. Conversely, in iterative distilla- L?simple = E(x̂t ,t)⇠B̂+ k✏✓T (x̂t , t) ✏✓S (x̂t , t)k2 ,
tion, batches containing x̂1 almost always include x̂2 . This (20)
⇥ ⇤
departure from the principle of independent and identically L?vlb = E(x̂t ,t)⇠B̂+ DKL (p✓T (x̂t 1 |x̂t )kp✓S (x̂t 1 |x̂t ) ,
distributed samples in a batch can potentially diminish the (21)
model’s generalization ability. where x̂t and t are produced by our proposed dynamic it-
To better align the distribution of the denoising data with erative distillation. The complete algorithm is detailed in
that of the standard training batch, we propose a method Algorithm 1.

2959
CIFAR10 32x32 CelebA 64x64 ImageNet 32x32
Method IS↑ FID↓ sFID↓ IS↑ FID↓ sFID↓ IS↑ FID↓ sFID↓
Teacher 9.52 4.45 7.09 3.08 4.43 6.10 13.63 4.67 4.03
Data-Based Training 8.73 7.84 7.38 3.04 5.39 7.23 9.99 10.56 5.24
Data-Limited Training (20%) 8.49 9.76 11.30 2.86 9.52 11.55 10.48 12.62 9.70
Data-Limited Training (15%) 8.44 11.07 12.47 2.84 9.60 11.48 10.43 13.50 10.76
Data-Limited Training (10%) 8.40 11.06 11.98 3.04 8.20 10.65 10.39 14.23 12.63
Data-Limited Training (5%) 8.39 10.91 11.99 2.86 9.64 11.27 10.48 13.63 10.62
Data-Free Training (0%) 8.28 12.06 13.23 2.87 10.66 12.71 10.47 13.20 9.56
Dynamic Iterative Distillation (Ours) 8.60 9.56 11.77 2.91 7.07 8.78 10.50 11.33 4.80

Table 2. Pixel-space performance comparison between data-limited training, data-free training and our dynamic iterative distillation on
CIFAR10 32 ⇥ 32 [22], CelebA 64 ⇥ 64 [27] and ImageNet 32 ⇥ 32 [7]. The term (P%) denotes the percentage of real data included in
the synthetic dataset. The best performance is indicated by boldface, while the second-best is denoted by underlining. Results from the
‘Teacher’ and ‘Data-Based Training’ are provided for reference only and are not included in the comparison.

4. Experiments tive distillation under different training conditions.


All the teacher models employ Convolutional Neural Net-
This section presents a series of experiments designed to works (CNNs). For the student models, we maintain
validate the efficacy of our proposed dynamic iterative dis- the same architecture but reduce the scale. Additionally,
tillation. In Sec. 4.1, we introduce our experimental set- we conduct cross-architecture experiments between CNN-
ting and establish relevant baselines for comparative analy- based and ViT-based (Vision Transformer [9, 39]) DMs on
sis from a data-centric perspective. Sec. 4.2 provides a com- CIFAR10. Details of the architecture are listed in Sec. 8.
parison between these baselines and our method, assessing
Metrics. The distance between the generated samples
performance separately in pixel and latent spaces. We also
and the reference samples can be estimated by the Fréchet
demonstrate the capability of our approach to train models
Inception Distance (FID) score [14]. In our experiments,
across different architectures. Lastly, Sec. 4.3 includes an
we utilize the FID score as the primary metric for eval-
ablation study to solidify the validation of our method.
uation. Additionally, we report sFID [32] and Inception
Score (IS) [47] as secondary metrics. Following previous
4.1. Experiment Setting
work [16, 36, 37], we generate 50K samples for DMs, and
Datasets, teachers and students. The training of high- we use the full training set in the corresponding dataset to
resolution diffusion models typically requires substantial compute the metrics. Without additional contextual states,
time, so these models are often developed in latent space all the samples are generated through 50 Improved DDPM
to expedite the process [44]. To assess our method, we con- sampling steps [36] in pixel space and 200 DDIM sampling
duct experiments in both pixel and latent spaces, focusing steps [52] in latent space. All of our metrics are calculated
on low and high-resolution generation, respectively. by ADM TensorFlow evaluation suite [8].
• Pixel space. We utilize three pretrained DMs as teacher Baselines. As DKDM is a new paradigm proposed in
models, following the configurations introduced by Ning this paper, traditional distillation methods are not suitable
et al. [37]. These models were trained separately on CI- to serve as baselines. Therefore, we establish two kinds of
FAR10 at a resolution of 32 ⇥ 32 [22], CelebA at 64 ⇥ 64 baselines from a data-centric perspective.
[27] and ImageNet at 32 ⇥ 32 [7]. • Data-Free Training, which is depicted in Fig. 2b, in-
• Latent space. We adopt two different DMs as teacher volves a teacher model generating a large quantity of
models, adhering to the configurations proposed by Rom- high-quality synthetic samples, matching the size of the
bach et al. [44]. These models were trained on CelebA- original dataset. These synthetic samples form the train-
HQ 256 ⇥ 256 [20] and FFHQ 256 ⇥ 256 [21]. It is im- ing set D0 for the student models, which are initialized
portant to note that the pre-trained models in the latent randomly and trained according to the standard proce-
space were typically trained using a simpler loss func- dure, cf . Algorithm 2 in Sec. 7. Details about our syn-
tion, denoted as Lsimple (11), without incorporating the thetic datasets can be found in Tab. 4.
KL divergence Lvlb (9). For our experiments, we adopt • Data-Limited Training integrates a fixed proportion
L?simple (20) as the DKDM objective. This approach al- (ranging from 5% to 20%) of the original dataset samples
lows us to investigate the effectiveness of dynamic itera- with the synthetic dataset D0 used in data-free training.

2960
CelebA-HQ 256 FFHQ 256 CNN T. ViT T.
Method
FID↓ sFID↓ FID↓ sFID↓ CNN S.
Teacher 5.69 10.02 5.93 7.52 Data-Free Training 9.64 44.62
Data-Based 9.09 12.10 8.91 8.75 Dynamic Iterative Distillation 6.85 13.17
Data-Limited (20%) 14.49 17.08 15.43 12.25 ViT S.
Data-Limited (15%) 14.89 16.98 16.02 12.48 Data-Free Training 17.11 63.15
Data-Limited (10%) 15.23 17.53 16.00 12.47 Dynamic Iterative Distillation 17.11 17.86
Data-Limited (5%) 15.07 17.64 15.86 12.56
Data-Free (0%) 15.36 17.56 16.32 12.75
Table 5. FID scores on CIFAR10 for cross-architecture distilla-
Ours 8.69 12.50 11.53 10.29
tion between CNN and ViT models. The FID score for the CNN
teacher model is 4.45, and that of the ViT teacher is 11.30. Abbre-
Table 3. Latent-space performance comparison between data- viations used: ‘T.’ stands for teacher, ‘S.’ stands for student.
limited training, data-free training and our dynamic iterative dis-
tillation on CelebA-HQ 256⇥256 [20] and FFHQ 256⇥256 [21].
The term (P%) denotes the percentage of real data included in the
fective strategy to mitigate the costs associated with large-
synthetic dataset. The best performance is indicated by boldface,
while the second-best is denoted by underlining. Results from the
scale datasets. Additionally, we observe instances where
‘Teacher’ and ‘Data-Based Training’ are provided for reference our method outperforms traditional data-based training ap-
only and are not included in the comparison. proaches, exemplified by the IS score on CIFAR10 and the
FID score on CelebA-HQ. This outcome indicates that, to
some extent, neural networks face challenges in learning the
Dataset #Images Method
complex reverse diffusion processes inherent in data-based
Pixel Space training, whereas the knowledge from pretrained teacher
CIFAR10 32 [22] 50,000 IDDPM-1000 [36] models is easier to learn. This insight further highlights an
CelebA 64 [27] 202,599 IDDPM-100 [36] additional benefit: our method not only reduces the reliance
ImageNet 32 [7] 1,281,167 IDDPM-100 [36] on extensive datasets but also potentially yields models with
superior performance. Moreover, we found our method
Latent Space
only consumes minor extra GPU memory while achieving
CelebA-HQ 256 [20] 25,000 DDIM-100 [52] faster training speed in latent space, cf . Sec. 10. Some gen-
FFHQ 256 [21] 60,000 DDIM-100 [52] erated results are visualized in Fig. 5.
Cross-Architecture Distillation. Our method tran-
Table 4. Image counts and generation methods for our baselines, scends specific model architectures, enabling distillation
which match the training set sizes of their respective teachers. The
from CNN-based DMs to ViT-based ones and vice versa.
notation ⇤ N indicates that the synthetic dataset is generated
As shown in Tab. 5, our method effectively facilitates
using N sampling steps with the ⇤ method.
cross-architecture distillation, yielding superior perfor-
mance compared to baselines. Additionally, our results sug-
It facilitates a comparative analysis between our purely gest that CNNs are more effective as compressed DMs.
data-free method and those able to partially access to the
original dataset. 4.3. Ablation Study
Additionally, we also report performance of data-based To validate our approach, we tested the FID score of our
training, illustrated in Fig. 2a, which serves as an upper per- progressively designed methods, including iterative, shuf-
formance limit for our analysis. fled iterative, and dynamic iterative distillation, over 200K
training iterations without early stop. The results, shown
4.2. Main Results in Fig. 4a, demonstrate that our dynamic iterative distilla-
Effectiveness. Tabs. 2 and 3 present the performance com- tion strategy not only converges more rapidly but also de-
parison between our dynamic iterative distillation method livers superior performance. The convergence curve for our
and baseline models in pixel and latent spaces, respec- method closely matches that of the baseline, which confirms
tively. Our trained students consistently outperform base- the effectiveness of the DKDM objective in alignment with
lines across various datasets and metrics, demonstrating the the standard optimization objective Eq. (13).
efficacy of our proposed DKDM objective and dynamic it- Further experiments explored the effects of varying ⇢ on
erative distillation approach. These results validate our ini- the performance of dynamic iterative distillation. As de-
tial hypothesis posited in Sec. 1 and confirm that leverag- picted in Fig. 4b, higher ⇢ values enhance the distillation
ing existing diffusion models to train new ones is an ef- process up to a point, beyond which performance gains di-

2961
Data-Free Iterative Distillation 0.010 0.020 0.030
Shuffled Iterative Distillation 0.050 0.100 0.200
Dynamic Iterative Distillation 0.300 0.400 Data-Free
300
50 CIFAR10 32
40
12
FID

FID
200 30

20
10
10
100 100 150 ImageNet 32
8

0 50 100 150 200 10 15 20


Training Iterations (k) Training Iterations (k)
(a) Ablation (b) Effect of ⇢ CelebA 64 CelebA-HQ 256 FFHQ 256

Figure 4. FID scores of analytical experiments on CIFAR10. (a): Ablation on Figure 5. Selected samples generated by our student
dynamic iterative distillation with ⇢ = 0.4. (b): Effect of different ⇢. models across five datasets.

minish. This outcome supports our hypothesis that dynamic Data-Free Knowledge Distillation. Traditional data-
iterative distillation enhances batch construction flexibility, free knowledge distillation typically transfers knowledge
thereby improving distillation efficiency. Beyond a certain from a slow teacher model to a lightweight student without
level of flexibility, further increases in ⇢ yield no significant needing access to the original training dataset, addressing
benefit to the distillation process. For information regarding privacy concerns. Early methods optimized randomly ini-
GPU memory consumption with varying ⇢, cf . Sec. 10. Ad- tialized noise to produce synthetic data [3, 33, 63] for distil-
ditional discussion and analytical experiments are available lation. Owing to the slow nature of this optimization, subse-
in Secs. 11 and 12. quent studies have employed generative models to synthe-
size training data [5, 6, 10, 11, 29, 31, 64, 65]. These ef-
5. Related Work forts primarily distilled knowledge for non-generative mod-
els, such as classification networks. In contrast, this paper
Knowledge Distillation for Diffusion Models. Knowledge
focuses on the distillation of generative diffusion models
Distillation (KD) [12, 15, 24, 59] is an effective method
themselves. We deeply dive into the generation mechanism
for transferring the capabilities from teacher models to stu-
of diffusion models and design an effective and efficient
dents for model compression [18, 23, 38, 41, 43, 48, 60].
method to produce synthetic data for distillation.
In the context of diffusion models, KD is usually adopted
to accelerate the inherently slow generation process, which 6. Conclusion
involves multiple sampling steps. Approaches in this do-
main generally fall into two categories: 1) reducing model In this paper, we aim at addressing rapidly increasing cost
size [62, 66] and 2) decreasing sampling steps [13, 28, 30, associated with the demand for large-scale datasets in train-
46, 49, 50, 54, 58]. The first strategy focuses on distilling ing diffusion models. To mitigate this data burden, we
smaller models to reduce inference time, while the second introduce Data-Free Knowledge Distillation for Diffusion
distills the multi-step sampling behavior of teacher models Models (DKDM), a novel scenario that utilizes pretrained
into fewer steps for the student, thereby accelerating genera- diffusion models to train new ones with any architecture,
tion. Distinct from these conventional acceleration-oriented while not requiring access to the original training dataset.
KD methods, our approach shifts focus towards the data To achieve this, we carefully design a DKDM objective
perspective, aiming to mitigate the extensive data require- and dynamic iterative distillation method, which separately
ments of training diffusion models by distilling knowledge guarantees effectiveness and efficiency in the training pro-
from teacher models to randomly initialized student models cess of the student model. To the best of our knowledge,
in a data-free manner. Among existing methods, the BOOT we are the first to explore this scenario and make initial ef-
method [13] is most closely related to our work and is also a forts. Our experiments show superior performance across
data-free KD approach. While its codebase remains entirely five datasets, including both pixel and latent spaces. Fur-
closed, the primary difference between BOOT and ours lies thermore, in some cases, our data-free method even outper-
in the architecture of the student model. The BOOT method forms models trained with the entire dataset. This offers a
retains both the structure and weights from the teacher, more efficient direction for training diffusion models from
thereby limiting the flexibility of the student. In contrast, a data perspective, providing a valuable insight for future
our method permits any student architecture. advancements.

2962
Acknowledgment [11] Gongfan Fang, Kanya Mo, Xinchao Wang, Jie Song, Shitao
Bei, Haofei Zhang, and Mingli Song. Up to 100x faster data-
This work was partially sponsored by the National Nat- free knowledge distillation. In Proceedings of the AAAI Con-
ural Science Foundation of China under Grant 62306084 ference on Artificial Intelligence, pages 6597–6604, 2022. 8
and U23B2051, and Shenzhen Science and Technol- [12] Jianping Gou, Baosheng Yu, Stephen J Maybank, and
ogy Program under Grant KQTD20240729102207002, Dacheng Tao. Knowledge distillation: A survey. Interna-
GXWD20231128102243003, KJZD20230923115113026, tional Journal of Computer Vision, 129(6):1789–1819, 2021.
ZDSYS20230626091203008. 8
[13] Jiatao Gu, Shuangfei Zhai, Yizhe Zhang, Lingjie Liu, and
References Josh Susskind. Boot: Data-free distillation of denois-
ing diffusion models with bootstrapping. arXiv preprint
[1] Shekoofeh Azizi, Simon Kornblith, Chitwan Saharia, Mo-
arXiv:2306.05544, 2023. 8
hammad Norouzi, and David J. Fleet. Synthetic data from
diffusion models improves imagenet classification. Transac- [14] Martin Heusel, Hubert Ramsauer, Thomas Unterthiner,
tions on Machine Learning Research, 2023. 2 Bernhard Nessler, and Sepp Hochreiter. Gans trained by a
two time-scale update rule converge to a local nash equilib-
[2] Yogesh Balaji, Seungjun Nah, Xun Huang, Arash Vahdat, Ji-
rium. Advances in Neural Information Processing Systems,
aming Song, Qinsheng Zhang, Karsten Kreis, Miika Aittala,
30, 2017. 6
Timo Aila, Samuli Laine, et al. ediff-i: Text-to-image dif-
fusion models with an ensemble of expert denoisers. arXiv [15] Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distill-
preprint arXiv:2211.01324, 2022. 2 ing the knowledge in a neural network. arXiv preprint
[3] Kuluhan Binici, Shivam Aggarwal, Nam Trung Pham, arXiv:1503.02531, 2015. 8
Karianto Leman, and Tulika Mitra. Robust and resource- [16] Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffu-
efficient data-free knowledge distillation by generative sion probabilistic models. Advances in Neural Information
pseudo replay. In Proceedings of the AAAI Conference on Processing Systems, 33:6840–6851, 2020. 1, 2, 3, 6
Artificial Intelligence, pages 6089–6096, 2022. 8 [17] Qingqing Huang, Daniel S Park, Tao Wang, Timo I Denk,
[4] Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dock- Andy Ly, Nanxin Chen, Zhengdong Zhang, Zhishuai Zhang,
horn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Jiahui Yu, Christian Frank, et al. Noise2music: Text-
Align your latents: High-resolution video synthesis with la- conditioned music generation with diffusion models. arXiv
tent diffusion models. In Proceedings of the IEEE/CVF Con- preprint arXiv:2302.03917, 2023. 1
ference on Computer Vision and Pattern Recognition, pages [18] Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao
22563–22575, 2023. 1 Chen, Linlin Li, Fang Wang, and Qun Liu. Tinybert: Distill-
[5] Hanting Chen, Yunhe Wang, Chang Xu, Zhaohui Yang, ing bert for natural language understanding. arXiv preprint
Chuanjian Liu, Boxin Shi, Chunjing Xu, Chao Xu, and Qi arXiv:1909.10351, 2019. 8
Tian. Data-free learning of student networks. In Proceed- [19] Minguk Kang, Jun-Yan Zhu, Richard Zhang, Jaesik Park,
ings of the IEEE/CVF international conference on computer Eli Shechtman, Sylvain Paris, and Taesung Park. Scal-
vision, pages 3514–3522, 2019. 8 ing up gans for text-to-image synthesis. In Proceedings of
[6] Yoojin Choi, Jihwan Choi, Mostafa El-Khamy, and Jungwon the IEEE/CVF Conference on Computer Vision and Pattern
Lee. Data-free network quantization with adversarial knowl- Recognition, pages 10124–10134, 2023. 2
edge distillation. In Proceedings of the IEEE/CVF Con-
[20] Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen.
ference on Computer Vision and Pattern Recognition Work-
Progressive growing of GANs for improved quality, stabil-
shops, pages 710–711, 2020. 8
ity, and variation. In International Conference on Learning
[7] Patryk Chrabaszcz, Ilya Loshchilov, and Frank Hutter. A
Representations, 2018. 6, 7
downsampled variant of imagenet as an alternative to the ci-
far datasets. arXiv preprint arXiv:1707.08819, 2017. 6, 7 [21] Tero Karras, Samuli Laine, and Timo Aila. A style-based
generator architecture for generative adversarial networks.
[8] Prafulla Dhariwal and Alexander Nichol. Diffusion models
In Proceedings of the IEEE/CVF conference on computer vi-
beat gans on image synthesis. Advances in Neural Informa-
sion and pattern recognition, pages 4401–4410, 2019. 6, 7
tion Processing Systems, 34:8780–8794, 2021. 1, 6
[9] Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, [22] Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple
Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, layers of features from tiny images. 2009. 6, 7
Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- [23] Xiaojie Li, Jianlong Wu, Hongyu Fang, Yue Liao, Fei Wang,
vain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is and Chen Qian. Local correlation consistency for knowl-
worth 16x16 words: Transformers for image recognition at edge distillation. In European conference on computer vi-
scale. In International Conference on Learning Representa- sion, pages 18–33. Springer, 2020. 8
tions, 2021. 6 [24] Xiaojie Li, Shaowei He, Jianlong Wu, Yue Yu, Liqiang Nie,
[10] Gongfan Fang, Jie Song, Xinchao Wang, Chengchao Shen, and Min Zhang. Mask again: Masked knowledge distilla-
Xingen Wang, and Mingli Song. Contrastive model inver- tion for masked video modeling. In Proceedings of the ACM
sion for data-free knowledge distillation. arXiv preprint International Conference on Multimedia, page 2221–2232.
arXiv:2105.08584, 2021. 8 ACM, 2023. 8

2963
[25] Xiaojie Li, Yibo Yang, Xiangtai Li, Jianlong Wu, Yue Yu, [39] William Peebles and Saining Xie. Scalable diffusion models
Bernard Ghanem, and Min Zhang. Genview: Enhanc- with transformers. In International Conference on Computer
ing view quality with pretrained generative model for self- Vision, pages 4172–4182, 2023. 6, 1
supervised learning. In Proceedings of the European Con- [40] Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu,
ference on Computer Vision. Springer, 2024. 2 and Mark Chen. Hierarchical text-conditional image gen-
[26] Zaijing Li, Yuquan Xie, Rui Shao, Gongwei Chen, Dong- eration with clip latents. arXiv preprint arXiv:2204.06125,
mei Jiang, and Liqiang Nie. Optimus-1: Hybrid multimodal 2022. 2
memory empowered agents excel in long-horizon tasks. In [41] Jun Rao, Liang Ding, Shuhan Qi, Meng Fang, Yang Liu, Li
The Thirty-eighth Annual Conference on Neural Information Shen, and Dacheng Tao. Dynamic contrastive distillation for
Processing Systems, 2024. 2 image-text retrieval. IEEE Transactions on Multimedia, 25:
[27] Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. 8383–8395, 2023. 8
Deep learning face attributes in the wild. In Proceedings of [42] Jun Rao, Xuebo Liu, Lian Lian, Shengjun Cheng, Yunjie
the IEEE international conference on computer vision, pages Liao, and Min Zhang. Commonit: Commonality-aware in-
3730–3738, 2015. 6, 7 struction tuning for large language models via data parti-
[28] Eric Luhman and Troy Luhman. Knowledge distillation in tions. In EMNLP, 2024. 2
iterative generative models for improved sampling speed. [43] Jun Rao, Xv Meng, Liang Ding, Shuhan Qi, Xuebo Liu, Min
arXiv preprint arXiv:2101.02388, 2021. 8 Zhang, and Dacheng Tao. Parameter-efficient and student-
[29] Liangchen Luo, Mark Sandler, Zi Lin, Andrey Zhmoginov, friendly knowledge distillation. IEEE Transactions on Mul-
and Andrew Howard. Large-scale generative data-free distil- timedia, 26:4230–4241, 2024. 8
lation. arXiv preprint arXiv:2012.05578, 2020. 8 [44] Robin Rombach, Andreas Blattmann, Dominik Lorenz,
[30] Chenlin Meng, Robin Rombach, Ruiqi Gao, Diederik Patrick Esser, and Björn Ommer. High-resolution image
Kingma, Stefano Ermon, Jonathan Ho, and Tim Salimans. synthesis with latent diffusion models. In Proceedings of
On distillation of guided diffusion models. In Proceedings the IEEE/CVF conference on computer vision and pattern
of the IEEE/CVF Conference on Computer Vision and Pat- recognition, pages 10684–10695, 2022. 1, 2, 6
tern Recognition, pages 14297–14306, 2023. 8 [45] Chitwan Saharia, William Chan, Saurabh Saxena, Lala
[31] Paul Micaelli and Amos J Storkey. Zero-shot knowledge Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour,
transfer via adversarial belief matching. Advances in Neu- Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans,
ral Information Processing Systems, 32, 2019. 8 et al. Photorealistic text-to-image diffusion models with deep
[32] Charlie Nash, Jacob Menick, Sander Dieleman, and Peter W. language understanding. Advances in Neural Information
Battaglia. Generating images with sparse representations. In Processing Systems, 35:36479–36494, 2022. 2
International Conference on Machine Learning, 2021. 6 [46] Tim Salimans and Jonathan Ho. Progressive distillation for
[33] Gaurav Kumar Nayak, Konda Reddy Mopuri, Vaisakh Shaj, fast sampling of diffusion models. In International Confer-
Venkatesh Babu Radhakrishnan, and Anirban Chakraborty. ence on Learning Representations, 2022. 2, 8
Zero-shot knowledge distillation in deep networks. In In- [47] Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki
ternational Conference on Machine Learning, pages 4743– Cheung, Alec Radford, Xi Chen, and Xi Chen. Improved
4751. PMLR, 2019. 8 techniques for training gans. In Advances in Neural Infor-
[34] Quang Nguyen, Truong Vu, Anh Tran, and Khoi Nguyen. mation Processing Systems, 2016. 6
Dataset diffusion: Diffusion-based synthetic data generation [48] V Sanh. Distilbert, a distilled version of bert: smaller, faster,
for pixel-level semantic segmentation. Advances in Neural cheaper and lighter. arXiv preprint arXiv:1910.01108, 2019.
Information Processing Systems, 36, 2024. 2 8
[35] Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav [49] Axel Sauer, Dominik Lorenz, Andreas Blattmann, and Robin
Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Rombach. Adversarial diffusion distillation. arXiv preprint
Mark Chen. Glide: Towards photorealistic image generation arXiv:2311.17042, 2023. 8
and editing with text-guided diffusion models. arXiv preprint [50] Axel Sauer, Frederic Boesel, Tim Dockhorn, Andreas
arXiv:2112.10741, 2021. 2 Blattmann, Patrick Esser, and Robin Rombach. Fast high-
[36] Alexander Quinn Nichol and Prafulla Dhariwal. Improved resolution image synthesis with latent adversarial diffusion
denoising diffusion probabilistic models. In International distillation. arXiv preprint arXiv:2403.12015, 2024. 8
Conference on Machine Learning, 2021. 1, 3, 6, 7 [51] Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan,
[37] Mang Ning, Enver Sangineto, Angelo Porrello, Simone and Surya Ganguli. Deep unsupervised learning using
Calderara, and Rita Cucchiara. Input perturbation reduces nonequilibrium thermodynamics. In International Confer-
exposure bias in diffusion models. In International Confer- ence on Machine Learning, 2015. 1
ence on Machine Learning, 2023. 6, 1 [52] Jiaming Song, Chenlin Meng, and Stefano Ermon. Denois-
[38] Wonpyo Park, Dongju Kim, Yan Lu, and Minsu Cho. ing diffusion implicit models. In International Conference
Relational knowledge distillation. In Proceedings of the on Learning Representations,, 2021. 6, 7
IEEE/CVF conference on computer vision and pattern [53] Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Ab-
recognition, pages 3967–3976, 2019. 8 hishek Kumar, Stefano Ermon, and Ben Poole. Score-based

2964
generative modeling through stochastic differential equa- [66] Dingkun Zhang, Sijia Li, Chen Chen, Qingsong Xie, and
tions. In International Conference on Learning Represen- Haonan Lu. Laptop-diff: Layer pruning and normalized dis-
tations, 2021. 1 tillation for compressing diffusion models. arXiv preprint
[54] Yang Song, Prafulla Dhariwal, Mark Chen, and Ilya arXiv:2404.11098, 2024. 2, 8
Sutskever. Consistency models. In International Conference [67] Haoyu Zhang, Meng Liu, Yuhong Li, Ming Yan, Zan Gao,
on Machine Learning, 2023. 2, 8 Xiaojun Chang, and Liqiang Nie. Attribute-guided collab-
[55] Yonglong Tian, Lijie Fan, Phillip Isola, Huiwen Chang, and orative learning for partial person re-identification. IEEE
Dilip Krishnan. Stablerep: Synthetic images from text-to- Transactions on Pattern Analysis and Machine Intelligence,
image models make strong visual representation learners. 45(12):14144–14160, 2023. 2
Advances in Neural Information Processing Systems, 36, [68] Haoyu Zhang, Meng Liu, Zixin Liu, Xuemeng Song, Yaowei
2024. 2 Wang, and Liqiang Nie. Multi-factor adaptive vision selec-
[56] Kun Wang, Hao Liu, Lirong Jie, Zixu Li, Yupeng Hu, and tion for egocentric video question answering. In Proceedings
Liqiang Nie. Explicit granularity and implicit scale corre- of the 41st International Conference on Machine Learning,
spondence learning for point-supervised video moment lo- pages 59310–59328. PMLR, 2024. 2
calization. In Proceedings of the 32nd ACM International [69] Yang Zhao, Yanwu Xu, Zhisheng Xiao, and Tingbo Hou.
Conference on Multimedia, pages 9214–9223, 2024. 2 Mobilediffusion: Subsecond text-to-image generation on
[57] Weijia Wu, Yuzhong Zhao, Hao Chen, Yuchao Gu, Rui Zhao, mobile devices. arXiv preprint arXiv:2311.16567, 2023. 2
Yefei He, Hong Zhou, Mike Zheng Shou, and Chunhua
Shen. Datasetdm: Synthesizing data with perception annota-
tions using diffusion models. Advances in Neural Informa-
tion Processing Systems, 36:54683–54695, 2023. 2
[58] Sirui Xie, Zhisheng Xiao, Diederik P Kingma, Tingbo Hou,
Ying Nian Wu, Kevin Patrick Murphy, Tim Salimans, Ben
Poole, and Ruiqi Gao. Em distillation for one-step diffusion
models. arXiv preprint arXiv:2405.16852, 2024. 2, 8
[59] Chuanguang Yang, Helong Zhou, Zhulin An, Xue Jiang,
Yongjun Xu, and Qian Zhang. Cross-image relational knowl-
edge distillation for semantic segmentation. In Proceedings
of the IEEE/CVF conference on computer vision and pattern
recognition, pages 12319–12328, 2022. 8
[60] Chuanguang Yang, Zhulin An, Libo Huang, Junyu Bi, Xin-
qiang Yu, Han Yang, Boyu Diao, and Yongjun Xu. Clip-kd:
An empirical study of clip model distillation. In Proceed-
ings of the IEEE/CVF Conference on Computer Vision and
Pattern Recognition, pages 15952–15962, 2024. 8
[61] Ling Yang, Zhilong Zhang, Yang Song, Shenda Hong, Run-
sheng Xu, Yue Zhao, Wentao Zhang, Bin Cui, and Ming-
Hsuan Yang. Diffusion models: A comprehensive survey of
methods and applications. ACM Computing Surveys, 56(4):
1–39, 2023. 1
[62] Xingyi Yang, Daquan Zhou, Jiashi Feng, and Xinchao Wang.
Diffusion probabilistic model made slim. In Proceedings of
the IEEE/CVF Conference on computer vision and pattern
recognition, pages 22552–22562, 2023. 2, 8
[63] Hongxu Yin, Pavlo Molchanov, Jose M Alvarez, Zhizhong
Li, Arun Mallya, Derek Hoiem, Niraj K Jha, and Jan Kautz.
Dreaming to distill: Data-free knowledge transfer via deep-
inversion. In Proceedings of the IEEE/CVF conference on
computer vision and pattern recognition, pages 8715–8724,
2020. 8
[64] Jaemin Yoo, Minyong Cho, Taebum Kim, and U Kang.
Knowledge extraction with no observable data. Advances
in Neural Information Processing Systems, 32, 2019. 8
[65] Shikang Yu, Jiachen Chen, Hu Han, and Shuqiang Jiang.
Data-free knowledge distillation via feature exchange and
activation region constraint. In Proceedings of the IEEE/CVF
Conference on Computer Vision and Pattern Recognition,
pages 24266–24275, 2023. 8

2965

You might also like