DiffColor: Text-Guided Image Colorization
DiffColor: Text-Guided Image Colorization
Abstract—Recent data-driven image colorization methods have inherent color ambiguity, such as clothes, houses, other man-
enabled automatic or reference-based colorization, while still made objects, and so on. To enhance the variety of produced
suffering from unsatisfactory and inaccurate object-level color colors, other researchers [15]–[17] suggested predicting per-
control. To address these issues, we propose a new method
arXiv:2308.01655v1 [[Link]] 3 Aug 2023
called DiffColor that leverages the power of pre-trained diffusion pixel color distributions instead of a single color. Additionally,
models to recover vivid colors conditioned on a prompt text, reference-based approaches [18]–[20] propose to transfer color
without any additional inputs. DiffColor mainly contains two statistics from a reference image to a grayscale image using
stages: colorization with generative color prior and in-context correspondences. However, their outputs often exhibit visual
controllable colorization. Specifically, we first fine-tune a pre- distortions such as irregular and inconsistent colors. Other
trained text-to-image model to generate colorized images using a
CLIP-based contrastive loss. Then we try to obtain an optimized works [14], [18], [21] also resort to Generative Adversarial
text embedding aligning the colorized image and the text prompt, Networks (GANs) [22] to encourage generated chroma distri-
and a fine-tuned diffusion model enabling high-quality image butions to be indistinguishable from real-life color images.
reconstruction. Our method can produce vivid and diverse colors Recently, diffusion-based models like denoising diffusion
with a few iterations, and keep the structure and background probabilistic models (DDPM) [23], [24] and score-based gen-
intact while having colors well-aligned with the target language
guidance. Moreover, our method allows for in-context coloriza- erative models [25], [26] have made significant progress in
tion, i.e., producing different colorization results by modifying image generation tasks [23], [26]. The latest studies [26],
prompt texts without any fine-tuning, and can achieve object- [27] have demonstrated even higher quality of image synthesis
level controllable colorization results. Extensive experiments and compared to other generative models such as variational au-
user studies demonstrate that DiffColor outperforms previous toencoders (VAEs) [28]–[30], flows [31]–[33], auto-regressive
works in terms of visual quality, color fidelity, and diversity of
colorization options. models [34], [35], and generative adversarial networks (GANs)
[22], [36], [37]. Moreover, a recent denoising diffusion im-
Index Terms—Diffusion Models, Image Colorization, Interac- plicit model (DDIM) [38] has further improved the sampling
tive Multimedia Artworks
efficiency and enabled almost perfect inversion [27].
In this work, we propose a novel approach, DiffColor, for
I. I NTRODUCTION high-fidelity text-guided image colorization using diffusion
models, which leverages the power of text descriptions to
Image colorization, the process of adding colors to black- provide accurate and diverse colorization results. Text-guided
and-white images, has gained significant attention in recent image colorization has the advantage of being able to incorpo-
years, with applications in various fields such as photography, rate prior knowledge about the object or scene being colorized,
advertising, and the film industry. However, colorization is making it particularly useful in domains where such prior
inherently an ill-posed problem due to the fact that it requires knowledge is available. DiffColor aims to recover vivid and
estimating missing color channels from only one grayscale diverse colors by using pre-trained diffusion models [39], [40],
value, making it a challenging task. Traditional colorization to incorporate text descriptions into the colorization process.
methods typically require user intervention, such as providing DiffColor mainly contains two stages: colorization with
color scribbles [1]–[3] or reference images [4]–[6] to obtain generative color prior and in-context controllable colorization.
satisfactory results. With the advent of deep learning, various Specifically, we first convert an input grayscale image into
techniques have been developed that leverage deep neural latent noises through forward diffusion. We fine-tune the score
networks and large-scale datasets such as ImageNet [7] or function in the reverse diffusion process, conditioned on the
COCO-Stuff [8] to learn colorization in an end-to-end fashion context description text embedding, to enrich the generated
[9]–[13]. image’s colors by using a contrastive loss. The contrastive
Numerous existing colorization methods [9], [11], [14] em- loss is built by taking a colorized image and context text as a
ploy pixel-level regression models to learn color distribution pair of positive samples and other anti-color texts as negative
priors from large-scale training data. However, these methods samples, incorporating more accurate generative color priors
tend to produce brownish average colors for objects with into colorization results. Then, we try to obtain an optimized
text embedding aligning the colorized image and the text
Jianxin Lin, Peng Xiao, Yijun Wang, Rongju Zhang, and Xiangxiang prompt, and a fine-tuned diffusion model enabling high-quality
Zeng are with the College of Computer Science and Electronic Engineering, image reconstruction. Finally, we can achieve two goals: 1) our
Hunan University, Changsha 410082, China (e-mail: linjianxin@[Link];
napping@[Link]; wyjun@[Link]; 1251553178@[Link]; method allows for in-context colorization, in which, unlike
xzeng@[Link]). existing text-driven image manipulation methods, DiffColor
2
Prompt: “A brown chihuahua sitting before the red “A brown/white dog sitting on a black/red bench, with
background” green/yellow grassland and blue sky”
Fig. 1. Text-Guided Image Colorization. Given a grayscale image and a target text prompt with color description, our proposed approach DiffColor can
produce high-quality colorization results and realize object-level controllable colorization.
can generate different colorization results by simply modifying further propose to refine the initial colorized image re-
prompt texts without any fine-tuning; 2) our method can construction and text embedding, resulting in in-context
produce object-level controllable colorization results. and object-level controllable image colorization.
We demonstrate the effectiveness of DiffColor through • Evaluation and comparison: The paper presents extensive
extensive experiments and user studies. The results show experiments and evaluations of the proposed framework,
that our method produces high-quality images that closely demonstrating its superiority over existing methods in
resemble the input image and align well with the target text, terms of colorization quality, diversity, and controllability.
and outperforms previous state-of-the-art methods in terms of These evaluations also highlight the effectiveness of the
colorization quality and diversity. Furthermore, we conducted text-guided approach and its potential for real-world
an ablation study to evaluate the effect of each element of our applications.
method.
Our contributions can be summarized as follows: II. R ELATED W ORKS
• Development of a new framework: The paper presents Learning-based colorization. Learning-based colorization
DiffColor which utilizes text-guided diffusion models methods automatically colorize grayscale images with large-
for the task of image colorization. This framework is scale training data and end-to-end learning models. This
designed to achieve high fidelity and controllable col- line of work mainly addresses the two key issues of color
orization results compared to existing methods. semantics and multi-modality in image colorization. PalGAN
• Incorporation of a novel color contrastive loss as guid- [21] predicts the pixel colors in a coarse-to-fine paradigm,
ance: The proposed framework incorporates a color con- decomposing colorization to palette estimation and pixel-wise
trastive loss by leveraging both context text and other assignment. Specifically, to build better semantic representa-
anti-color texts, bringing higher quality colorization re- tion, Iizuka et al. [11] and Zhao et al. [17] present a two-
sults. branch architecture that jointly learns and fuse local image
• Introducing a two-stage color refinement process: we features and global priors (e.g., semantic labels). Su et al.
3
[13] learn object-level semantics by training on the cropped of personalized images to create images of the same object in
object images. To address the challenge of multi-modality, a new environment. Imagic [59] accepts a single image and
several studies [15], [16] have suggested representing color a simple text prompt describing the desired edit, and aims to
prediction as pixel-level color classification. These approaches apply this edit while preserving a maximal amount of details
enable the assignment of multiple colors to each pixel based from the image. L-code [60] and Unicolor [61] also proposed
on the posterior probability. to colorize images with text prompts. However, unlike L-
To tackle the uncertainty and diversity problems of col- code and Unicolor require large-scale training, we are the
orization, reference-based methods attempt to transfer the first work proposing to leverage the powerful text-to-image
color statistics from a reference image to the input grayscale generation model, i.e., a pre-trained stable diffusion model,
image [5], [6], [18]–[20], [41], which also enabled automatic for text-guided image colorization.
colorization. The correspondences between the input and ref- In this work, we provide a novel approach for text-
erence images are computed based on low-level similarity guided image colorization which can achieve fidelity and text
measures at pixel [41], [42], semantic segments [4], [6], or alignment simultaneously, performing high-quality coloriza-
super-pixel levels [5], [43]. These methods highly rely on the tion globally and locally in one single image. The resulting
time-consuming procedure of reference image retrieving as image outputs align well with the target text while preserving
well as manual annotations of image regions [4], [43], and the overall background, structure, and composition of the
often suffer from unsatisfactory and inaccurate object-level original image.
color control. Then Wu et al. [18] investigate the integration
of a generative color prior learned from a pretrained BigGAN III. A PPROACH
[36] to improve the ability of a deep learning model to produce A. Overview architecture
diverse colored results.
These methods were insufficient in providing accurate se- Our method mainly consists of two stages: colorization
mantic guidance, such as natural language, for coloring. How with generative color prior and in-context controllable
to enhance the semantic modeling capability of context, and colorization. As shown in Figure 2, given an input grayscale
how to better incorporate semantic priors into the coloring image xgray ∈ R1×H×W and a context text prompt c which
process, are issues that need to be addressed. describes the content of the grayscale image, our goal of
Text-driven image synthesis. Text-to-image synthesis has the first stage is to produce accurate and vivid colorization
been an active research area in the field of generative models, results, preserving a maximal amount of details from xgray .
with numerous works focusing on generating realistic images To achieve this, we consider the task of text-guided image col-
from textual descriptions. Initial approaches relied on RNN- orization through a pre-trained text-to-image latent diffusion
based methods [44], which were later replaced by generative model. We begin by inputting xgray into a pre-trained encoder,
adversarial networks (GANs) [22], [36], [37] that yielded obtaining the corresponding latent code z. The denoising
improved results [45]. Subsequent enhancements to GAN- model ϵθ takes two inputs, namely the noisy latent zT and
based models included multi-stage architectures and attention a context text embedding which is the encoded context text
mechanisms [46]–[48]. DALL-E [49] introduced a two-stage prompt c of xgray by a CLIP text encoder. Then, we fine-
approach that used a discrete VAE and transformer to model tune the score function in the reverse diffusion process using
the joint distribution of text and image tokens without relying a contrastive loss that converts grayscale image to a primary
on GANs. New solutions to the problem of text-guided image colorized image xrgbpri . Next, to guarantee the reconstruction of
synthesis were introduced with the recent development of dif- the grayscale image’s details, a spatial alignment module is
fusion models [23]–[25], [38], [39], which have demonstrated utilized to align the content of the grayscale image xgray and
impressive results. However, these methods lack the ability the primary colorized image xrgb pri , obtaining colorized image
rgb
to edit parts of a real image while preserving the rest. As a x̂pri .
result, recent works have focused on taking advantage of these In the second stage, taking the colorized image x̂rgb pri and
models rather than training a large-scale text-to-image model a target text prompt ĉ which is constructed by adding color
from scratch. descriptions in front of the object words in the context text
Text-driven image manipulation. Recently, text-driven image prompt c as inputs, our goal is to achieve a fine-tuned
manipulation has achieved significant progress using GANs model that can edit the color of the image in a way that
[22], [36], [37], [50], [51], which are known for their high- satisfies the given text while achieving high fidelity. Inspired
quality generation, further with CLIP [52], which consists of by DreamBooth [58] for better image identity capture, we
a semantically rich joint image-text representation, pre-trained rewrite the context text prompt for fine-tuning like “A [*]
over millions of text-image pairs. Works that combined these dog sitting on a [*] wooden bench.”, where “[*]” is a unique
components [53]–[56] produced highly realistic manipulations identifier. Then we optimize the context text embedding to
using text only. To obtain more expressive generation capabil- find the one that maximizes the similarity with the colorized
ities, recent approaches often use pre-trained diffusion models image within the proximity of the target text embedding, as
for image editing. DiffusionCLIP [57] uses the CLIP model well as fine-tune the diffusion models to improve the alignment
to obtain gradients for image manipulation, and achieves between the generated image x̂rgb rec and the colorized image
remarkable results in style transfer. DreamBooth [58] fine- x̂rgb
pri . In inference time, we perform a linear interpolation
tune the complete diffusion model by utilizing a small set between the optimized context text embedding and the target
4
Decoder
Encoder
Pretrained
Forward Diffusion
Diffusion Model
𝑧 𝑧𝑇 𝑧
𝑥 𝑔𝑟𝑎𝑦
Extractor
𝑟𝑔𝑏
Feature
Context Prompt = “A dog sitting on a CLIP Text 𝑥𝑝𝑟𝑖
wooden bench, with grassland and sky” Encoder Language-guided LDM warp
𝑔𝑟𝑎𝑦
𝑥 Spatial 𝑟𝑔𝑏
Negative Prompt 1 = “A grayscale photograph” Correlation 𝑥ො𝑝𝑟𝑖
Negative Prompt 2 = “A picture with scratches” CLIP Text CLIP Image
Encoder Encoder Spatial Alignment
Negative Prompt 3 = “Brownish average colors for objects”
…… Contrastive CLIP Guidance
Context Prompt = “A [*] dog CLIP Text Context Optimized In-context Controllable Colorization
sitting on a [*] wooden bench, Encoder Emb Emb
with [*] grassland and [*] sky” Target Prompt = “A brown dog sitting on a black wooden bench, with green grassland and blue sky”
×T
Fine-tuned target Optimized
Interpolate
Diffusion Emb Emb
Model ×T
𝑟𝑔𝑏 𝑟𝑔𝑏 Fine-tuned
𝑥ො𝑝𝑟𝑖 Optimized 𝑥ො𝑟𝑒𝑐
Emb
Diffusion
Model
×T 𝑟𝑔𝑏 Spatial
𝑥𝑓𝑖𝑛 Alignment
Fine-tuned
Diffusion 𝑥 𝑔𝑟𝑎𝑦 𝑟𝑔𝑏
𝑥ො𝑓𝑖𝑛
Colorized Image Model
Reconstruction In-context Color Editing
Fig. 2. Framework of Diffcolor. DiffColor mainly contains two stages: 1) colorization with generative color prior that produces accurate and vivid
colorization results; 2) in-context controllable colorization that edits the color of the first-stage output in a way that satisfies the target text prompt.
text embedding to obtain a solution that achieves high fidelity to identify the correspondence between texts and images. In
to the input image while maintaining semantic alignment with this stage, we leverage a pre-trained CLIP model to extract
the target text, realizing in-context controllable colorization. knowledge effectively and build CLIP guidance through a
contrastive loss, which makes colorized image xrgb
pri (sampled
B. Colorization with Generative Color Prior from LDM) more similar to context text c and more dissimilar
In this stage, we aim to leverage rich color priors encapsu- to other anti-color texts c− s which are called negative text
lated in a pretrained generative model to guide the colorization prompts, as shown below:
process. This stage mainly consists of a pre-trained language-
exp e⊤ e+
guided latent diffusion model, a CLIP guidance, and a spatial
alignment module. LCST = − log PL , (2)
exp (e⊤ e+ ) + i=1 exp e⊤ e−
i
Language-guided latent diffusion model. Specifically, the
img rgb
base model for this stage is the latent diffusion model (LDM) where e = Eclip (xpri ), e+ = Eclip
text
(c), e− text −
i = Eclip (ci ),
[39], which consists of an auto-encoder trained on images and img
Eclip is the CLIP image encoder, and L is the number of
a diffusion model learned on the latent space. The encoder E anti-color texts. As shown in Figure 2, we can set several
of the auto-encoder first encodes the given grayscale image anti-color texts, such as “A grayscale photograph.”, “A picture
xgray to a latent representation z, i.e., z = E(xgray ). The with scratches.” and so on, to prevent colorized image from
diffusion model is trained to produce latent codes within the falling into such situations. The diffusion guidance loss is thus
pre-trained latent space. The language-guided LDM is learned set to the weighted sum LLDM + λLCST .
as follows: Spatial alignment. The sampling process of diffusion
text
model is inherently stochastic, meaning that images generated
LLDM = Et,ϵ [||ϵ − ϵθ (zt , t, Eclip (c))||22 ], (1) from the diffusion model can not keep the inherent image
where t is the time step, zt is the latent noised to time t, ϵ details, even in the case of a deterministic sampling process
is the unscaled noise sample, ϵθ is the denoising model, and like DDIM. Therefore, the reconstruction of the grayscale
text
Eclip is the CLIP text encoder which encodes the context text image’s details cannot be guaranteed. Therefore, we utilize
prompt c describing the content of the grayscale image to a a spatial alignment module as previous works [18], such as
context text embedding. CoCosNet [62], which is pre-trained by matching grayscale
CLIP guidance. CLIP [52] was proposed to efficiently learn images and their corresponding warped color images, to align
semantic representation from texts and images. It employs primary colorized image xrgb pri and grayscale image xgray ,
a text encoder and an image encoder that are pre-trained obtaining colorized image x̂rgbpri .
5
Similarity (LPIPS) [65] that measures the perceptual similarity In addition, Peak Signal-to-Noise Ratio (PSNR) and Struc-
between the colorized images and the ground-truth images. tural Similarity Index Measure (SSIM) [66] are presented
7
for unconditional colorization methods. For text-conditioned not be perfectly aligned with the original grayscale image,
colorization, the CLIP score [52] is utilized to measure the we also apply CoCosNet to spatially align the results for a
relevance between the colorized images and the text prompts, fair comparison. We evaluate our DiffColor method using the
evaluating how well the generated images match the provided same context prompts without color hints, which are designed
textual descriptions. by human priors, as Imagic.
3) Datasets: Since we aim to produce high-fidelity image Qualitative Comparison. The results of the comparison
colorization results and our model does not require large data between our DiffColor and other unconditional methods are
training, we follow Imagic [67] and SINE [63] to collect displayed in Figure 4. As we can see, our method can produce
high-resolution, free-to-use images from Unsplash and Flickr high-fidelity colorized images with vivid and diverse colors,
for our experiments. Similarly, the dataset we collect has 60 without color bleeding. For example, in the second row of
images with various categories, such as dog, cat, car, house, Figure 4, when the gray image contains a house with two sym-
tree, etc. metric roofs, unlike DiffColor, most of the existing methods
tend to colorize the two symmetric roofs with two different
B. Qualitative evaluation colors, lacking the capability of globally semantic perception.
In addition, most of the existing methods usually produce
We demonstrate the versatility of our proposed method,
images of brownish average colors in the third row. Although
DiffColor, by applying it to a range of real-world images
Imagic has a powerful ability to manipulate images’ content,
with diverse objects. We use simple text prompts that describe
it fails to transfer gray images to vivid color images, which
different colors of objects in the generated image. We collected
demonstrates the necessity of the proposed CLIP guidance
high-resolution, free-to-use images from Unsplash and Flickr
with contrastive loss.
for our experiments. After optimization, we randomly gener-
Quantitative Comparison. To better show our re-colored
ate eight results for each grayscale image using different η
images’ quality, we provide quantitative evaluations using
and select the best one. DiffColor can apply various image
FID, PSNR, SSIM and LPIPS metrics. As shown in Table
colorizations via text prompts, including single-object color
I, our method still achieves the best performance than other
editing and multiple-object color editing, as shown in Figure
methods in four different metrics, indicating that our method
1 and Figure 3. Moreover, we demonstrate that DiffColor
can recover more vivid and natural colors with diffusion model
enables post-text color editing by adding new text prompt at
priors.
the end of the original text prompt, e.g., adding “with yellow
2) Text-conditioned colorization: For text-conditioned col-
grassland” at the end of “A dog sitting on the bench”, as
orization, We conduct a comprehensive comparison of our
shown in the last row of Figure 3. The text prompts specify the
proposed DiffColor with three state-of-the-art methods, i.e.,
desired modifications to different objects in grayscale images,
UniColor [61], L-CoDe [60] and ControlNet [71]. Both L-
and we can observe that 1) our method can accurately change
CoDe and UniColor are object-level image colorization meth-
the colors of different objects without color bleeding while
ods and can realize colorization results with color text prompts.
preserving the remaining parts of the image colors; 2) except
ControlNet proposed a neural network structure to control
for specially given colors, our method can recover vivid and
pretrained large diffusion models to support additional input
natural colors to the original grayscale images due to the
conditions, for colorization task we use the grayscale images
powerful diffusion model priors.
as visual condition and text prompts as text condition. We use
results from the second stage of our method for comparison
C. Comparison to state-of-the-arts in this part.
1) Unconditional colorization: For unconditional coloriza- Qualitative Comparison. The results of the comparison
tion, we conduct a comparison with several state-of-the-art between our DiffColor and other text-conditioned methods
(SOTA) techniques capable of colorizing grayscale images are displayed in Figure 5. In this part, we present visual
without requiring any color hints: namely GCP [18], DeOldify quality comparisons involving four distinct samples, each
[68], ColTran [69], CT2 [70] and Imagic [67]. We use results associated with different language descriptions. These descrip-
from the first stage of our method for comparison in this part. tions include instances where color words and object words
DeOldify [68] uses GANs to realistically colorize images and are clearly aligned. From the results, we can find that our
is particularly effective for restoring old images and videos. proposed method emerges as the superior approach in terms
GCP [18] is a two-stage GAN-based framework that generates of synthesizing visually pleasing and description-consistent
diverse and vivid colorizations by incorporating multiple color colorization results.
suggestions with grayscale images. ColTran [69] is a self- Quantitative Comparison. To better demonstrate our
supervised transformer model that achieves state-of-the-art method’s ability in the task of text-guided image colorization,
performance without relying on handcrafted priors or paired we provide quantitative evaluations comparing with state-of-
training data. CT2 [70] employs a transformer architecture the-art text-guided image colorization methods. For methods
with color tokens for color representation, leveraging self- that employ text prompts to colorize grayscale images, we
attention to efficiently colorize images while preserving their place particular emphasis on the CLIP Score [52] that mea-
structure and details. In addition, we also compare our method sures the consistency between colorized objects and color
with Imagic [59], which also has the capability of recovering prompts. As shown in Table II, our text-based method obtains
image colors. However, since Imagic’s generated results may the best CLIP similarity, which indicates that our method can
8
Imagic
Gray Image GCP DeOldify ColTran CT2 (with CoCosNet) Ours
FID ↓ PSNR ↑ SSIM ↑ LPIPS ↓ GCP DeOldify ColTran CT2 Imagic Ours-1st
GCP [18] 78.56 26.51 0.8714 0.29 12.3% 9.7% 16.3% 10.7% 3.5% 47.5%
DeOldify [68] 56.63 27.97 0.8867 0.26
ColTran [69] 73.24 27.88 0.8833 0.27 TABLE IV
CT2 [70] 58.03 29.39 0.8966 0.25 U SER STUDY RESULTS . P REFERENCE RATES OF TEXT- CONDITIONED
Imagic [59] 121.33 22.41 0.8023 0.35 IMAGE COLORIZATION METHODS IN TERMS OF IMAGE QUALITY AND
SEMANTIC CONSISTENCY.
Ours-1st 48.67 29.83 0.9063 0.22
UniColor L-CoDe ControlNet Ours-2nd
Quality 25.6% 13.7% 19.5% 41.2%
TABLE II Consistency 27.3% 8.1% 26.5% 38.5%
Q UANTITATIVE COMPARISON OF TEXT- CONDITIONED IMAGE
COLORIZATION METHODS .
FID ↓ CLIP Score ↑ LPIPS ↓ 3) User Study.: To assess the subjective quality of our
UniColor [61] 62.87 28.20 0.24 method in terms of color vividness, colorfulness, and con-
L-CoDe [60] 87.80 25.42 0.28 sistency, we conducted a user study comparing it with other
ControlNet [72] 65.28 27.08 0.25 colorization methods.
Ours-2nd 51.30 28.62 0.22 For unconditional colorization, we recruited 25 participants
with normal or corrected-to-normal vision and without color
blindness. We selected 30 grayscale images from various
categories of image content. For each image, we displayed the
accurately locate the objects and colorize them with accurate grayscale version on the leftmost side, while five colorization
colors. In addition, it shows that our method can produce results were displayed randomly to avoid potential bias. The
higher image quality and fidelity than other methods with the participants were asked to select the best-colorized image
best FID and LPIPS scores. based on the quality of colorization in terms of vividness
9
and colorfulness. We present the results in Table III. Our D. Ablation study
DiffColor method was significantly preferred (47.5%) by the
users compared to other colorization methods, demonstrating Colorization with Generative Color Prior. We exhibit the
a distinct advantage in producing natural and vivid results. initial colorization results obtained from the grayscale image
after processing solely through the first stage as shown in
Figure 6 (a). By incorporating the contrastive loss derived from
the CLIP, we effectively confine the entire output of the first
stage within a color space. Furthermore, the reconstruction
For text-conditioned colorization, we recruited another 25 loss allows us to concurrently consider the quality of the
participants. In the evaluation process, except for selecting the image reconstruction. Consequently, through capitalizing on
best-colorized image in terms of vividness and colorfulness, the extensive prior knowledge embedded within the diffusion
participants are presented with a caption describing a color model, we successfully achieve grayscale image colorization
image, and they are asked to choose the colorized image based exclusively on content cues. It is important to note
that is most consistent with the caption. As shown in Table that our objective during the first stage does not entail the
IV, our approach consistently outperforms the other methods, diffusion model’s ability to perfectly reconstruct the grayscale
achieving the highest scores in both scenarios, i.e., image image’s details while simultaneously performing colorization.
quality and semantic consistency. Instead, we permit a certain degree of detail loss, which
10
First Stage
(without
First Stage Second Stage
Gray Image Contrastive Loss & Alignment) (without Alignment) First Stage (without Alignment) Second Stage
Prompt: “A dog sitting on a bench, with grassland and “A white dog sitting on a black bench,
sky” with green grassland and the red sky”
(a) (b)
Increasing 𝛈
Image Recolored Image
Fig. 7. Ablation study on interpolation parameter η. In this figure, we directly input real color images to the second stage of our method.
is subsequently mitigated through the utilization of spatial with the provided color prompt. By employing CoCosNet to
alignment techniques. achieve spatial alignment between the output and the grayscale
Controllable Colorization. We showcase the color editing image, we manage to extract the colors from the diffusion
outcomes of the second stage by applying it directly to real model’s direct output while maintaining the structure and
color images as shown in Figure 7. This demonstrates the details inherent in the original image. As shown in Table V,
second stage’s versatility, as it is not only capable of function- the alignment module significantly improves image structure
ing effectively within our proposed two-stage methodology but preservation in the first stage, and our method can produce
can also operate independently as a semantic color alteration results with a highly similar image structure in the second
technique for color images. stage.
Spatial Alignment. We compare the results with and Interpolation Parameter η. Figure 7 displays several
without the spatial alignment as shown in Figure 6. It turns different results with increasing η for the same image. This
out that, although the output of the diffusion model is not approach provides the user an option to colorize gray images
impeccable in terms of reconstruction, it is generally consistent when the color description is ambiguous or imprecise.
11
TABLE V [10] H. Bahng, S. Yoo, W. Cho, D. K. Park, Z. Wu, X. Ma, and J. Choo,
Q UANTITATIVE RESULTS FOR ABLATION STUDY. “Coloring with words: Guiding image colorization through text-based
palette generation,” in Proceedings of the european conference on
computer vision (eccv), 2018, pp. 431–447.
Stage First Second [11] S. Iizuka, E. Simo-Serra, and H. Ishikawa, “Let there be color! joint
W/O Align. W/ Align. W/O Align. W/ Align. end-to-end learning of global and local image priors for automatic
image colorization with simultaneous classification,” ACM Transactions
PSNR 21.13 29.83 25.56 29.16 on Graphics (ToG), vol. 35, no. 4, pp. 1–11, 2016.
SSIM 0.7643 0.9063 0.8476 0.8934 [12] C. Lei and Q. Chen, “Fully automatic video colorization with self-
regularization and diversity,” in Proceedings of the IEEE/CVF confer-
ence on computer vision and pattern recognition, 2019, pp. 3753–3761.
[13] J.-W. Su, H.-K. Chu, and J.-B. Huang, “Instance-aware image coloriza-
V. C ONCLUSION tion,” in Proceedings of the IEEE/CVF Conference on Computer Vision
and Pattern Recognition, 2020, pp. 7968–7977.
In this paper, we presented a novel method called Diff- [14] P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros, “Image-to-image translation
Color that addresses the limitations of classic and reference- with conditional adversarial networks,” in Proceedings of the IEEE
based image colorization methods. Our approach leverages conference on computer vision and pattern recognition, 2017, pp. 1125–
1134.
the power of pre-trained diffusion models to produce vivid [15] G. Larsson, M. Maire, and G. Shakhnarovich, “Learning representations
and diverse colors conditioned on a text prompt without for automatic colorization,” in Computer Vision–ECCV 2016: 14th
requiring any additional inputs. DiffColor consists of two main European Conference, Amsterdam, The Netherlands, October 11–14,
2016, Proceedings, Part IV 14. Springer, 2016, pp. 577–593.
stages: colorization with generative color prior and in-context [16] R. Zhang, P. Isola, and A. A. Efros, “Colorful image colorization,”
controllable colorization. We fine-tune a pre-trained text-to- in Computer Vision–ECCV 2016: 14th European Conference, Amster-
image model using a CLIP-based contrastive loss to generate dam, The Netherlands, October 11-14, 2016, Proceedings, Part III 14.
Springer, 2016, pp. 649–666.
colorized images and optimize the text embedding to align [17] J. Zhao, L. Liu, C. G. Snoek, J. Han, and L. Shao, “Pixel-level semantics
the colorized image with the target text prompt. Moreover, guided image colorization,” arXiv preprint arXiv:1808.01597, 2018.
we fine-tune a diffusion model to enable high-quality image [18] Y. Wu, X. Wang, Y. Li, H. Zhang, X. Zhao, and Y. Shan, “Towards
vivid and diverse image colorization with generative color prior,” in
reconstruction. Our method allows for in-context colorization Proceedings of the IEEE/CVF international conference on computer
and achieves object-level controllable colorization results. vision, 2021, pp. 14 377–14 386.
Extensive experiments and user studies demonstrate that [19] H. Camilo, M. Clément, and A. Bugeau, “Super-attention for exemplar-
based image colorization,” in Proceedings of the Asian Conference on
DiffColor outperforms previous works in terms of visual Computer Vision, 2022, pp. 4548–4564.
quality, color fidelity, and semantic consistency. We believe [20] M. He, D. Chen, J. Liao, P. V. Sander, and L. Yuan, “Deep exemplar-
that our approach represents a significant advancement in text- based colorization,” ACM Transactions on Graphics (TOG), vol. 37,
no. 4, pp. 1–16, 2018.
guided image colorization using diffusion models and provides [21] Y. Wang, M. Xia, L. Qi, J. Shao, and Y. Qiao, “Palgan: Image
a convenient and powerful tool for producing high-quality colorization with palette generative adversarial networks,” in Computer
colorized images based on textual descriptions. Our method Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October
23–27, 2022, Proceedings, Part XV. Springer, 2022, pp. 271–288.
has the potential to benefit a wide range of applications, [22] I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley,
including art, fashion, and product design, among others. S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial networks,”
Further research could explore the use of DiffColor for other Communications of the ACM, vol. 63, no. 11, pp. 139–144, 2020.
[23] J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,”
image editing tasks beyond colorization. Advances in Neural Information Processing Systems, vol. 33, pp. 6840–
6851, 2020.
R EFERENCES [24] J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli,
“Deep unsupervised learning using nonequilibrium thermodynamics,”
[1] Y.-C. Huang, Y.-S. Tung, J.-C. Chen, S.-W. Wang, and J.-L. Wu, “An in International Conference on Machine Learning. PMLR, 2015, pp.
adaptive edge detection based colorization algorithm and its applica- 2256–2265.
tions,” in Proceedings of the 13th annual ACM international conference [25] Y. Song and S. Ermon, “Generative modeling by estimating gradients
on Multimedia, 2005, pp. 351–354. of the data distribution,” Advances in neural information processing
[2] A. Levin, D. Lischinski, and Y. Weiss, “Colorization using optimization,” systems, vol. 32, 2019.
in ACM SIGGRAPH 2004 Papers, 2004, pp. 689–694. [26] Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and
[3] L. Yatziv and G. Sapiro, “Fast image and video colorization using B. Poole, “Score-based generative modeling through stochastic differ-
chrominance blending,” IEEE transactions on image processing, vol. 15, ential equations,” arXiv preprint arXiv:2011.13456, 2020.
no. 5, pp. 1120–1129, 2006. [27] P. Dhariwal and A. Nichol, “Diffusion models beat gans on image
[4] R. Ironi, D. Cohen-Or, and D. Lischinski, “Colorization by example.” synthesis,” Advances in Neural Information Processing Systems, vol. 34,
Rendering techniques, vol. 29, pp. 201–210, 2005. pp. 8780–8794, 2021.
[5] R. K. Gupta, A. Y.-S. Chia, D. Rajan, E. S. Ng, and H. Zhiyong, “Image [28] D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” arXiv
colorization using similar images,” in Proceedings of the 20th ACM preprint arXiv:1312.6114, 2013.
international conference on Multimedia, 2012, pp. 369–378. [29] A. Van Den Oord, O. Vinyals et al., “Neural discrete representation
[6] G. Charpiat, M. Hofmann, and B. Schölkopf, “Automatic image coloriza- learning,” Advances in neural information processing systems, vol. 30,
tion via multimodal predictions,” in Computer Vision–ECCV 2008: 10th 2017.
European Conference on Computer Vision, Marseille, France, October [30] A. Razavi, A. Van den Oord, and O. Vinyals, “Generating diverse
12-18, 2008, Proceedings, Part III 10. Springer, 2008, pp. 126–139. high-fidelity images with vq-vae-2,” Advances in neural information
[7] J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: processing systems, vol. 32, 2019.
A large-scale hierarchical image database,” in 2009 IEEE conference on [31] L. Dinh, J. Sohl-Dickstein, and S. Bengio, “Density estimation using
computer vision and pattern recognition. Ieee, 2009, pp. 248–255. real nvp,” arXiv preprint arXiv:1605.08803, 2016.
[8] H. Caesar, J. Uijlings, and V. Ferrari, “Coco-stuff: Thing and stuff classes [32] D. P. Kingma and P. Dhariwal, “Glow: Generative flow with invertible
in context,” in Proceedings of the IEEE conference on computer vision 1x1 convolutions,” Advances in neural information processing systems,
and pattern recognition, 2018, pp. 1209–1218. vol. 31, 2018.
[9] S. Guadarrama, R. Dahl, D. Bieber, M. Norouzi, J. Shlens, and [33] D. Rezende and S. Mohamed, “Variational inference with normalizing
K. Murphy, “Pixcolor: Pixel recursive colorization,” arXiv preprint flows,” in International conference on machine learning. PMLR, 2015,
arXiv:1705.07208, 2017. pp. 1530–1538.
12
[34] J. Menick and N. Kalchbrenner, “Generating high fidelity images with IEEE/CVF conference on computer vision and pattern recognition, 2021,
subscale pixel networks and multidimensional upscaling,” arXiv preprint pp. 2256–2265.
arXiv:1812.01608, 2018. [57] G. Kim, T. Kwon, and J. C. Ye, “Diffusionclip: Text-guided diffusion
[35] A. Van Den Oord, N. Kalchbrenner, and K. Kavukcuoglu, “Pixel models for robust image manipulation,” in Proceedings of the IEEE/CVF
recurrent neural networks,” in International conference on machine Conference on Computer Vision and Pattern Recognition, 2022, pp.
learning. PMLR, 2016, pp. 1747–1756. 2426–2435.
[36] A. Brock, J. Donahue, and K. Simonyan, “Large scale gan training for [58] N. Ruiz, Y. Li, V. Jampani, Y. Pritch, M. Rubinstein, and K. Aberman,
high fidelity natural image synthesis,” arXiv preprint arXiv:1809.11096, “Dreambooth: Fine tuning text-to-image diffusion models for subject-
2018. driven generation,” arXiv preprint arXiv:2208.12242, 2022.
[37] T. Karras, S. Laine, and T. Aila, “A style-based generator architecture [59] B. Kawar, S. Zada, O. Lang, O. Tov, H. Chang, T. Dekel, I. Mosseri, and
for generative adversarial networks,” in Proceedings of the IEEE/CVF M. Irani, “Imagic: Text-based real image editing with diffusion models,”
conference on computer vision and pattern recognition, 2019, pp. 4401– arXiv preprint arXiv:2210.09276, 2022.
4410. [60] S. Weng, H. Wu, Z. Chang, J. Tang, S. Li, and B. Shi, “L-code:
[38] J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit models,” Language-based colorization using color-object decoupled conditions,”
arXiv preprint arXiv:2010.02502, 2020. in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36,
[39] R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High- no. 3, 2022, pp. 2677–2684.
resolution image synthesis with latent diffusion models,” in Proceedings [61] Z. Huang, N. Zhao, and J. Liao, “Unicolor: A unified framework
of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- for multi-modal colorization with transformer,” ACM Transactions on
tion, 2022, pp. 10 684–10 695. Graphics (TOG), vol. 41, no. 6, pp. 1–16, 2022.
[40] C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, [62] X. Zhou, B. Zhang, T. Zhang, P. Zhang, J. Bao, D. Chen, Z. Zhang, and
K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans F. Wen, “Cocosnet v2: Full-resolution correspondence learning for image
et al., “Photorealistic text-to-image diffusion models with deep language translation,” in Proceedings of the IEEE/CVF Conference on Computer
understanding,” Advances in Neural Information Processing Systems, Vision and Pattern Recognition, 2021, pp. 11 465–11 475.
vol. 35, pp. 36 479–36 494, 2022. [63] Z. Zhang, L. Han, A. Ghosh, D. Metaxas, and J. Ren, “Sine: Single
[41] X. Liu, L. Wan, Y. Qu, T.-T. Wong, S. Lin, C.-S. Leung, and P.-A. Heng, image editing with text-to-image diffusion models,” arXiv preprint
“Intrinsic colorization,” in ACM SIGGRAPH Asia 2008 papers, 2008, arXiv:2212.04489, 2022.
pp. 1–9. [64] M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter,
[42] T. Welsh, M. Ashikhmin, and K. Mueller, “Transferring color to “Gans trained by a two time-scale update rule converge to a local
greyscale images,” in Proceedings of the 29th annual conference on nash equilibrium,” Advances in neural information processing systems,
Computer graphics and interactive techniques, 2002, pp. 277–280. vol. 30, 2017.
[43] A. Y.-S. Chia, S. Zhuo, R. K. Gupta, Y.-W. Tai, S.-Y. Cho, P. Tan, and [65] R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The
S. Lin, “Semantic colorization with internet images,” ACM Transactions unreasonable effectiveness of deep features as a perceptual metric,” in
on Graphics (ToG), vol. 30, no. 6, pp. 1–8, 2011. Proceedings of the IEEE conference on computer vision and pattern
[44] E. Mansimov, E. Parisotto, J. L. Ba, and R. Salakhutdinov, “Generating recognition, 2018, pp. 586–595.
images from captions with attention,” arXiv preprint arXiv:1511.02793, [66] W. Zhou, “Image quality assessment: from error measurement to struc-
2015. tural similarity,” IEEE transactions on image processing, vol. 13, pp.
[45] S. Reed, Z. Akata, X. Yan, L. Logeswaran, B. Schiele, and H. Lee, 600–613, 2004.
“Generative adversarial text to image synthesis,” in International con- [67] B. Kawar, S. Zada, O. Lang, O. Tov, H. Chang, T. Dekel, I. Mosseri, and
ference on machine learning. PMLR, 2016, pp. 1060–1069. M. Irani, “Imagic: Text-based real image editing with diffusion models,”
[46] T. Xu, P. Zhang, Q. Huang, H. Zhang, Z. Gan, X. Huang, and in Proceedings of the IEEE/CVF Conference on Computer Vision and
X. He, “Attngan: Fine-grained text to image generation with attentional Pattern Recognition, 2023, pp. 6007–6017.
generative adversarial networks,” in Proceedings of the IEEE conference [68] A. Salmona, L. Bouza, and J. Delon, “Deoldify: A review and imple-
on computer vision and pattern recognition, 2018, pp. 1316–1324. mentation of an automatic colorization method,” Image Processing On
[47] H. Zhang, T. Xu, H. Li, S. Zhang, X. Wang, X. Huang, and D. N. Line, vol. 12, pp. 347–368, 2022.
Metaxas, “Stackgan: Text to photo-realistic image synthesis with stacked [69] M. Kumar, D. Weissenborn, and N. Kalchbrenner, “Colorization trans-
generative adversarial networks,” in Proceedings of the IEEE interna- former,” arXiv preprint arXiv:2102.04432, 2021.
tional conference on computer vision, 2017, pp. 5907–5915. [70] S. Weng, J. Sun, Y. Li, S. Li, and B. Shi, “Ct 2: Colorization trans-
[48] ——, “Stackgan++: Realistic image synthesis with stacked generative former via color tokens,” in European Conference on Computer Vision.
adversarial networks,” IEEE transactions on pattern analysis and ma- Springer, 2022, pp. 1–16.
chine intelligence, vol. 41, no. 8, pp. 1947–1962, 2018. [71] jagilley. (2023) Controlnet. [Online]. Available: [Link]
[49] A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, jagilley/controlnet
and I. Sutskever, “Zero-shot text-to-image generation,” in International [72] M. A. Lvmin Zhang, “Adding conditional control to text-to-image
Conference on Machine Learning. PMLR, 2021, pp. 8821–8831. diffusion models,” arXiv preprint arXiv:2302.05543, 2023.
[50] T. Karras, M. Aittala, S. Laine, E. Härkönen, J. Hellsten, J. Lehtinen,
and T. Aila, “Alias-free generative adversarial networks,” Advances in
Neural Information Processing Systems, vol. 34, pp. 852–863, 2021.
[51] T. Karras, S. Laine, M. Aittala, J. Hellsten, J. Lehtinen, and T. Aila,
“Analyzing and improving the image quality of stylegan,” in Proceedings
of the IEEE/CVF conference on computer vision and pattern recognition,
2020, pp. 8110–8119.
[52] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal,
G. Sastry, A. Askell, P. Mishkin, J. Clark et al., “Learning transferable
visual models from natural language supervision,” in International
conference on machine learning. PMLR, 2021, pp. 8748–8763.
[53] R. Abdal, P. Zhu, J. Femiani, N. Mitra, and P. Wonka, “Clip2stylegan:
Unsupervised extraction of stylegan edit directions,” in ACM SIG-
GRAPH 2022 conference proceedings, 2022, pp. 1–9.
[54] R. Gal, O. Patashnik, H. Maron, A. H. Bermano, G. Chechik, and
D. Cohen-Or, “Stylegan-nada: Clip-guided domain adaptation of image
generators,” ACM Transactions on Graphics (TOG), vol. 41, no. 4, pp.
1–13, 2022.
[55] O. Patashnik, Z. Wu, E. Shechtman, D. Cohen-Or, and D. Lischinski,
“Styleclip: Text-driven manipulation of stylegan imagery,” in Proceed-
ings of the IEEE/CVF International Conference on Computer Vision,
2021, pp. 2085–2094.
[56] W. Xia, Y. Yang, J.-H. Xue, and B. Wu, “Tedigan: Text-guided di-
verse face image generation and manipulation,” in Proceedings of the