Dual Contrastive Attention for Image SR
Dual Contrastive Attention for Image SR
Keywords: With its powerful ability to model geometric transformations, the deformable convolutional network brings
Image super-resolution great improvements for single image super-resolution (SISR). Nevertheless, its location-variant sampling
Deformable convolution method leads to an escalation in spatial variance as the deformable convolutional layers are stacked,
Contrastive learning
consequently resulting in limited performance. Hence, we propose a novel and effective approach called dual
Attention mechanism
contrastive attention-guided deformable convolutional network (DCADCN) for SISR modeling. Specifically,
we propose an attention-guided deformable convolutional module with joint inner and external attention
mechanisms to fully exploit the correspondences between input and deformation features and preserve spatial
characteristics to the extent possible. Additionally, we propose a dual mixed feature extractor consisting of two
parallel sub-paths. This design allows for the learning of diverse and complementary spatial features. Further-
more, contrastive learning is applied to further amplify the role of key features and mitigate the interference
of noisy features. Extensive experimental results demonstrate that DCADCN is capable of effectively handling
classic SISR, SISR with blind noise, and real-world SISR tasks. Moreover, our method achieves comparable or
even better performance with lower computational cost compared to state-of-the-art methods.
✩ This document is the results of the research projects funded by the National Natural Science Foundation of China (No. 11571325) and the Fundamental
Research Funds for the Central Universities, China (No. CUC2019 A002).
✩✩ This paper has been recommended for acceptance by Zicheng Liu.
∗ Corresponding author.
E-mail addresses: qiao_fj@[Link] (F. Qiao), ygzhu@[Link] (Y. Zhu), gfli@[Link] (G. Li), ribbenlee@[Link] (B. Li).
[Link]
Received 30 July 2023; Received in revised form 14 February 2024; Accepted 24 February 2024
Available online 27 February 2024
1047-3203/© 2024 The Author(s). Published by Elsevier Inc. This is an open access article under the CC BY license ([Link]
F. Qiao et al. Journal of Visual Communication and Image Representation 100 (2024) 104097
the deformation features present in LR images. Although DCN-based persistent memory. In the second approach, Liang et al. [18] proposed
methods have demonstrated a performance enhancement in SISR [11, a baseline framework based on the Transformer architecture for im-
12], they still suffer from the following two limitations. Firstly, the age restoration. Qin et al. [19] utilized attentive residual mechanism
irregular sampling method employed by DCN introduces spatial vari- to enhance the feature extraction capacity and replaced regular up-
ance, resulting in the gradual loss of spatial information with increasing sampling operation with a multi-scale separable up-sampling module
network depth [13]. Secondly, both regular and irregular sampling for better performance. However, both types of methods inevitably
methods are crucial, yet DCN primarily focuses on irregular features face challenges in terms of computational requirements and memory
while neglecting regular features. consumption, limiting their practical applicability.
To practically remedy these limitations, we propose a dual con- To accelerate training efficiency and reduce memory consumption,
trastive attention-guided deformable convolutional network (DCADCN) researchers have focused on reducing the input size or compressing
for SISR. Specifically, to preserve spatial information, we employ an the SISR model. For reducing the input size, Dong et al. [20] adopted
attention-guided deformable convolutional module (ADCM) to adap- LR images as input instead of up-sampled HR images, performing an
tively weight features by modeling the correspondences between input upscale operation at the final layer of the network to predict HR
and deformation features. To obtain diverse features, we design a dual images. For model compression, Sun et al. [21] utilized a large kernel
mixed feature extractor (DMFE) comprising two parallel sub-paths to ConvNet and a feature fusion for mobile-friendly SISR. Liu et al. [22]
capture regular and irregular features, respectively. Furthermore, we proposed feature distillation connection and established a residual
incorporate contrastive learning as part of the loss function to train our feature distillation network for effective SISR while keep lightweight.
model. In summary, our main contributions are summarized as follow: Subsequently, several networks, such as RepRFN [23], FDIWN [24],
and FDSCSR [25] combined CNNs with information distillation to
• We present an attention-guided deformable convolution to adap- improve training efficiency.
tively weight features, which utilize joint inner and external at- For complex SISR scenarios, it is well known that digital images
tention mechanisms to fully explore the correspondences between are affected by various factors, such as unknown noise, shooting en-
input and deformation features. vironment, and optical performance of equipment. To address SISR
• We exploit a dual structure, called DMFE, to learn diverse and problems with blind noise, Tian et al. [8] proposed an ACNet that em-
complementary spatial characteristics by stacked convolutional ploys a asymmetric architecture and flexible up-sampling mechanism.
layers and ADCMs. Zhang et al. [26] designed a complex degradation model considering
• A contrastive loss item is designed and embedded into the ADCM blur kernel and noise, generating clean images and using a plug-and-
to further amplify the role of key features and mitigate the inter- play algorithm to obtain a noise prior and the final SR image. Later,
ference of redundant noisy information. To our best knowledge, Zhang et al. [27] proposed a more complex but practical degradation
we are the first to apply contrastive learning to DCN to constrain model for real-world SISR, which consists of randomly shuffled blur,
the correspondences between input and deformation features. downsampling and noise degradations.
• Experimental results show that the proposed method can handle Despite the significant advancements in SISR research, there is still
the tasks of classic SISR, SISR with blind noise, and real-world a urgent need for a network that can simultaneously enhance SISR
SISR, and achieves a better trade-off between performance and performance, improve training efficiency, and handle complex SISR
computational cost. scenarios.
The remaining sections of this paper are organized as follows. 2.2. Deformable Convolutional Network (DCN)
Section 2 provides a comprehensive review of the related work in the
areas of SISR, DCNs, attention mechanisms, and contrastive learning. DCN was proposed to model geometric transformation of objects [9,
In Section 3, we present our proposed DCADCN and provide a detailed 28]. Recent works [29,30] introduced DCN for frame alignment in
description of each component within the network. Section 4 outlines video processing area, thus making more efficient use of information
the experimental setup and presents the performance evaluation of from adjacent frames. Liu et al. [31] proposed an attention-guided
DCADCN in the context of classic SISR, SISR with blind noise, and real- deformable convolutional network (ADNet) for high dynamic range
world SISR. Section 5 discusses the differences between our work and imaging. ADNet was constructed with two branches, the one utilizing
existing methods and points out the strengths and limitations of the the spatial attention module for extracting attention features and the
present approach. Finally, Section 6 concludes this paper. other adopting DCN to align the gamma-corrected images. The DCN
can also be used for SISR. For instance, Zhang et al. [10] developed a
2. Related work DefRCN for SISR, which applied DCN to fully exploit the deformation
features present in LR images. Despite showing great potential over
2.1. Single image super-resolution CNN, DefRCN still suffers from two major drawbacks. Firstly, DefRCN
does not consider the correspondences between input and learned
Significant progress has been made in recent years in the field of features, resulting in the loss of spatial information. Second, it contains
SISR with the advent of deep learning techniques. Dong et al. [14] an excessive number of trainable parameters. Addressing these two
made groundbreaking contributions by being the first to develop a issues in our network could theoretically improve SISR performance.
CNN framework for SISR, which yielded improved performance. Subse-
quently, various deep learning methods have been developed to further 2.3. Attention mechanism
enhance SISR performance [15–19], improve training efficiency [20–
25] and address complex SISR scenarios [8,26,27], respectively. The attention mechanism is a unique architecture that automatically
To enhance SISR performance, Kim et al. [15] introduced a very determines the contributions of input to output [32]. Through such a
deep CNN, dubbed VDSR, which employed a residual learning strategy. mechanism, the efficiency of information utilization can be significantly
This research demonstrated that deeper CNNs lead to better reconstruc- improved. In recent years, attention mechanisms have gained popular-
tion results. However, CNN-based approaches that primarily extract ity in various domains, including image processing [33,34] and natural
local features, resulting in limited performance in modeling long-range language processing [35]. They have also been applied to enhance
dependencies. To tackle this issue, two main approaches have been the performance of SISR. For instance, Yang et al. [36] employed LR
explored: recurrent structures and non-local methods. In the first ap- and HR images as queries and keys in an attention mechanism to
proach, Tai et al. introduced memory blocks in MemNet [17] to capture explore the correspondences between them, resulting in more accurate
2
F. Qiao et al. Journal of Visual Communication and Image Representation 100 (2024) 104097
Fig. 1. Overall architecture of the proposed dual contrastive attention-guided deformable convolutional network (DCADCN) for SISR, which mainly consists of dual mixed feature
extractors (DMFEs) and an up-sampling block.
texture reconstruction. Nevertheless, the commonly used self-attention 3.2. Dual Mixed Feature Extractor (DMFE)
mechanism requires computing the relevance between each pixel in
images, leading to a substantial computational workload. Therefore, a Motivated by the dual-path network [44], we introduce the concept
simpler attention architecture needs to be explored for our network. of dual path to our network and propose DMFE as a basic component.
Differently, the dual-path network only uses Conv layers, while we use
not only Conv layers but also the proposed ADCMs in the two sub-
2.4. Contrastive learning paths. Moreover, our DMFE is designed with two feature fusion and
enhancement stages to extract richer features. More details of ADCM
will be given in Section 3.3.
Contrastive Learning is a machine learning method that trains mod-
The architecture of DMFE is illustrated in Fig. 1. Let 𝑋𝑐𝑘 denote the
els by comparing the similarities and differences between different
input of 𝑘th (1 ≤ 𝑘 ≤ 𝑑) DMFE with 𝑐 channels, the DMFE first splits 𝐹𝑐𝑘
samples. In recent years, contrastive learning has made significant 𝑘
into two parts with the same channels: 𝑋1∶𝑐∕2 (abbreviated as 𝑋1 ) and
advancements in computer vision and natural language processing [37– 𝑘
𝑋𝑐∕2+1∶𝑐 (abbreviated as 𝑋2 ). These two parts are fed into two parallel
39]. However, there has been limited research on applying contrastive
learning to low-level visual tasks like SISR. For instance, Wu et al. [40] sub-paths for two-stage feature extraction and enhancement. During
employed effective data enhancement strategies to generate positive the first stage, 3 × 3 Conv layers and ADCMs are utilized in two sub-
and negative samples for contrastive learning and achieved a perfor- paths respectively to extract different spatial features. These features
are concatenated in the channel dimension and then the number of
mance improvement on SISR issue. Leveraging the growing popularity
channels is recovered by a 1 × 1 Conv layer. In addition, ESA is utilized
of self-supervised learning, contrastive learning can also be used for
in two paths to enhance the feature information. In this way, the feature
blind SISR [41,42] to implicitly model different degradation processes.
extraction and enhancement procedure in stage 1 is formulated as:
The main difference between our work and the above methods is that
the positive and negative samples are generated in different ways. ( ( ( ( ))))
𝑋1 ′ = 𝑅𝐿 𝑓 3×3 𝑅𝐿 𝑓 3×3 𝑋1 ,
𝑐𝑜𝑛𝑣 𝑐𝑜𝑛𝑣
′
( ( ( ( ))))
𝑋2 = 𝑅𝐿 𝑓𝑎𝑑𝑐𝑚 𝑅𝐿 𝑓𝑎𝑑𝑐𝑚 𝑋2
3. Proposed method ( [ ]) (1)
𝑋1𝑆1 = 𝑓𝑒𝑠𝑎 𝑓 1×1 𝑋1 ′ , 𝑋2 ′
𝑐𝑜𝑛𝑣
( [ ])
3.1. Overview
𝑋2𝑆1 = 𝑓𝑒𝑠𝑎 𝑓 1×1 𝑋2 ′ , 𝑋1 ′ ,
𝑐𝑜𝑛𝑣
The overall architecture of the proposed DCADCN is illustrated in where 𝑋1𝑆1 and 𝑋2𝑆1 are the outputs of the two sub-paths in stage 1,
3×3 and 𝑓 1×1 represent the function of 3 × 3 and 1 × 1
respectively. 𝑓𝑐𝑜𝑛𝑣
Fig. 1. Given a LR image 𝐼𝐿𝑅 ∈ Rℎ×𝑤×𝑐 , where ℎ × 𝑤 represents the 𝑐𝑜𝑛𝑣
Conv layers, respectively. 𝑓𝑎𝑑𝑐𝑚 denotes the function of ADCM, 𝑅𝐿(⋅)
spatial resolution and 𝑐 is the number of channels, a Conv layer is first
is a ReLU activation, [⋅] is the channel-wise concatenation, and 𝑓𝑒𝑠𝑎
utilized for shallow features extraction. In the network backbone, we
denotes the function of ESA module, the details of which can be found
cascade 𝑑 dual mixed feature extractors (DMFEs) to obtain rich features
in [22].
for HR image reconstruction. To excavate diverse features, the DMFE During the second stage, 3 × 3 Conv layers and ADCMs are also
contains two parallel sub-paths that utilize Conv layers and ADCMs utilized in two sub-paths to extract spatial features. These features
respectively to cargo complementary information. These features from are residually summed with their corresponding inputs and then con-
sub-paths are cross-fertilized and enhanced by channel-wise concate- catenated in channel dimension. Finally, the feature information is
nation operation and the enhanced spatial attention (ESA) [22]. After enhanced by an ESA module. The output 𝑋 𝑘+1 of the 𝑘th DMFE is
several stacked DMFEs, a residual connection is used to aggregate fea- calculated by:
tures of different levels. Finally, an up-sampling block [43] is employed ( ( ( ( ))))
𝑋1 ′′ = 𝑅𝐿 𝑓 3×3 𝑅𝐿 𝑓 3×3 𝑋1𝑆1 + 𝑋1 ,
to obtain residual map, which is added to the bilinear up-sampling of (
𝑐𝑜𝑛𝑣
( (
𝑐𝑜𝑛𝑣
( 𝑆1 ))))
the input to attain the super-resolved image 𝐼𝑆𝑅 ∈ R𝐻×𝑊 ×𝑐 , where ′′
𝑋2 = 𝑅𝐿 𝑓𝑎𝑑𝑐𝑚 𝑅𝐿 𝑓𝑎𝑑𝑐𝑚 𝑋2 + 𝑋2 , (2)
𝑘+1
[ ′′ ′′
]
𝐻 > ℎ, 𝑊 > 𝑤. 𝑋 = 𝑓𝑒𝑠𝑎 𝑋1 , 𝑋2 ,
3
F. Qiao et al. Journal of Visual Communication and Image Representation 100 (2024) 104097
where 𝑝𝑙 and 𝛥𝑝𝑙 denote the pre-specified and the learned offsets for
the 𝑙–th location of input 𝑋, respectively. 𝛥𝑚𝑙 is a modulation scalar.
Next, we take 𝑋 and 𝐷 as the inputs to the inner attention mech-
anism, respectively. Their corresponding outputs are regarded as the
two basic components of external attention mechanism: query (𝑞) and
key (𝑘). Then the relevance of 𝑞 ∈ R𝑐 and 𝑘 ∈ R𝑐 is computed to obtain
Fig. 2. Calculation flow chart of attention-guided deformable convolutional module
the attention weights 𝐴. For each 𝑞𝑖 in 𝑞 and 𝑘𝑗 in 𝑘, the value of 𝐴 at
(ADCM). position (𝑖, 𝑗) is determined by normalized inner product and a softmax
activation:
⟨ ⟩
𝑞𝑖 𝑘𝑗
𝐴𝑖,𝑗 = sof tmax , , 𝐴𝑖,𝑗 ∈ 𝐴, 𝐴 ∈ R𝑐×𝑐 . (6)
3.3. Attention-guided deformable convolutional module ‖𝑞𝑖 ‖ ‖ ‖
‖ ‖ ‖𝑘𝑗 ‖
‖ ‖
Due to the irregular sampling locations, previous DCN leads to Another key component of the external attention mechanism is
increased spatial variance as the network deepens [28,45]. To retain value (𝑣), which is a linear mapping of deformation features 𝐷. Mul-
useful spatial information, we exploit attention mechanism to capture tiplying the attention weight matrix and the 𝑣, we finally obtain the
the correspondences between input and deformation features, resulting weighted features
in an attention-guided deformable convolutional module (ADCM).
𝑌̂ = 𝐴𝑣. (7)
As discussed in Section 2.3, it is not appropriate to apply atten-
tion mechanism directly to our network. How to design an effective Compared with the commonly used self-attention mechanism [32],
attention mechanism to guide the training procedure of DCN is a very our external attention mechanism mainly has two advantages. First,
important part in our research. Here, we mainly consider two points: the external attention is able to model the correspondences between
first, the attention mechanism must be computationally efficient as it the input and deformation features, while the self-attention mechanism
will be inserted into every deformable Conv layer. Second, it must can only calculate the relevance between each pixel of the input.
be capable of modeling the nonlinear relationship between input and Second, the external attention mechanism is more efficient because the
deformation features. relevance is calculated across channel descriptors rather than spatial
Based on the aforementioned analyses, there are four parts in the dimensions.
ADCM as shown in Fig. 2, which are channel-downscaling layer, in- Channel-upscaling layer. Corresponding to the channel-downscaling
ner attention mechanism, external attention mechanism, and channel- layer, the ADCM use channel-upscaling layer to recover the channel
upscaling layer. size. In addition, a residual connection is used to obtain the final
Channel-downscaling layer. To make the whole module lightweight output.
enough, the ADCM starts with a 1 × 1 Conv layer to reduce channel
size with ratio 𝑟. 3.4. Loss function
Inner attention mechanism. We employ an inner attention to fully
capture the interdependencies within the features. Unlike the previous To enhance the role of important features and mitigate the inter-
channel attention [6] that has a strong sensitivity to spatial characteris- ference of redundant noisy features, we increase the distance between
tics, we rescale channel-wise statistic instead of each pixel in the feature relevant and irrelevant features. Concretely, we introduce the concept
map. Given an input 𝑋 ∈ Rℎ×𝑤×𝑐 , the inner attention mechanism first of contrastive learning and propose a contrastive loss. Contrastive
utilize certain aggregation technique (e.g., global average pooling) to learning techniques have been proven to be useful in the field of
obtain a channel-wise statistic image SR [40–42]. In our case, we leverage contrastive loss to improve
1 ∑∑
ℎ 𝑤 the SR performance by pulling positive features closer and pushing
𝑋̂ = 𝐻𝐺𝑃 (𝑋) = 𝑋(𝑖, 𝑗), (3) negative features away in the feature space. Compared to the previous
ℎ × 𝑤 𝑖=1 𝑗=1
application of contrastive learning, we have three main differences.
where 𝑋̂ ∈ R𝑐 , 𝐻𝐺𝑃 (⋅) denotes the global average pooling function, The first difference lies in the application scenario. We have applied
𝑋(𝑖, 𝑗) is the pixel of 𝑋 at location (𝑖, 𝑗). contrastive learning to DCN for the first time, which helps to constrain
4
F. Qiao et al. Journal of Visual Communication and Image Representation 100 (2024) 104097
For classic SISR and SISR with blind noise, the DMFE number and
channel number of DCADCN are set to 5 and 64, respectively. For real-
world SISR, we train a larger model with the DMFE number of 10
and the channel number of 128. We use Adam optimizer to train our
DCADCN by setting 𝛽1 = 0.9, 𝛽2 = 0.999, and 𝜀 = 1e−8. There is a total
of 800 epochs in the training procedure. The mini-batch size is set to
Fig. 3. Illustration of the contrastive loss. For each ordered 𝐴𝑖 (𝐴𝑖 ∈ 𝐴, 1 ≤ 𝑖 ≤ 𝑐), the
top 𝑛 ⋅ 𝑐 features are viewed as positive parts and the last 𝑛 ⋅ 𝑐 features are viewed as
16. The learning rate is initialized to 5e−4 and reduced by half every
negative parts. 200 epochs. We warm up our DCADCN for the first 100 epochs with
only 𝑚𝑎𝑒 , and then train the network with the overall loss of . The
project is implemented by PyTorch 1.11 and trained on a Tesla V100
the correspondences between input and deformation features. The sec- GPU.
ond point concerns the difference in how similarity is calculated. While
the conventional methods calculate the similarity between different 4.3. Comparisons with state-of-the-arts
samples, our method focuses on calculating the similarity between
different channels of the same sample. The third point is the selection To comprehensively evaluate the SISR effectiveness, we compare
our approach with some state-of-the-art methods, including LESR-
of positive and negative parts. In our contrastive loss, we choose feature
CNN [3], ACNet [8], RepRFN [23], VLESR [52], FDIWN [24], Shuf-
channels with high relevance as positive parts and ones with low
fleMixer [21], ARRFN [19], FDSCSR [25], RLFN [53], DefRCN [10],
relevance as negative parts.
and RFDN [22]. For fairness, all compared methods use the same
As illustrated in Fig. 3, given the external attention weight matrix training set and down-sampling method. We use both quantitative and
𝐴 ∈ R𝑐×𝑐 of the ADCM, the contrastive loss 𝑐𝑜𝑛 can be formulated as: qualitative analysis to perform experiments.
Quantitative analysis: The PSNR and SSIM of different methods are
shown in Table 1. As we can see, our DCADCN performs better against
𝐴̂ 𝑖 = sort(𝐴𝑖 , Descending), 1 ≤ 𝑖 ≤ 𝑐, 𝐴𝑖 ∈ 𝐴, 𝐴̂ 𝑖 ∈ 𝐴,
̂ (8)
other compared methods on almost all datasets and scale factors. For
instance, compared to ACNet [8], our DCADCN achieves a gain both
∑𝑛⋅𝑐 ( )
1∑
𝑐
𝑗=1 exp 𝐴̂ 𝑖,𝑗 ∕(𝑛 ⋅ 𝑐) of PSNR of 0.39 dB and SSIM of 0.0019 for ×2 SISR on Set5 with
𝑐𝑜𝑛 = − log ∑𝑐 ( ) , (9) less computational cost (see Fig. 9). It is essential to point out that
𝑐 𝑖=1
𝑗=𝑐−𝑛⋅𝑐+1 exp 𝐴̂ 𝑖,𝑗 ∕(𝑛 ⋅ 𝑐) + 𝑚
our approach outperforms DefRCN [10] (also DCN-based method) on
where 𝐴̂ 𝑖 stands for the descending sort of 𝐴𝑖 , 𝐴𝑖 and 𝐴̂ 𝑖 are the 𝑖th all benchmark datasets for three scales. This benefits from the unique
row of 𝐴 and 𝐴, ̂ respectively. 𝑛 represents the percentage of positive network architecture (i.e., ADCM, DMFE, and contrastive loss) of our
and negative feature channels which is less than 50%. 𝑚 is a margin DCADCN. These observed results also show the effectiveness of our
constant and manually takes the value of 1e−6. DCADCN for reconstructing accurate SR images.
Qualitative analysis: Some visual information, such as texture, struc-
As a consequence, the overall loss function of DCADCN consists of
ture, clarity of definition, and edge information, can be used to in-
two parts: one is commonly used mean absolute error (MAE) loss 𝑚𝑎𝑒 ,
tuitively evaluate the reconstruction effect. To display the qualitative
the other is contrastive loss 𝑐𝑜𝑛 , which is formulated as
results, we use officially released codes of the compared methods to
= 𝑚𝑎𝑒 + 𝜆𝑐𝑜𝑛 , (10) predict the SR images. Fig. 4 shows visual comparisons of image details
on different datasets (×4) between our DCADCN and other competitors.
The original images ‘‘barbara’’, ‘‘253027’’, and ‘‘img030’’ are selected
𝑚𝑎𝑒 = ‖ ‖
‖𝐼𝐻𝑅 − 𝐼𝑆𝑅 ‖1 , (11) from Set14, B100, and U100, respectively. From the enlarged view,
we can observe that our method can reconstruct super-resolved (SR)
where ‖⋅‖1 is the 𝐿1 norm, 𝐼𝑆𝑅 and 𝐼𝐻𝑅 are the predicted SR image and images that are closer to the original HR images than the competitors.
the target HR image, respectively. 𝜆 is the penalty parameter which is For example, in the reconstructed images of ‘‘253027’’, our method
manually set up as 1e−3. restores clearer lines on the animal, while some other methods would
produce blurred lines with severe artifacts. These visual comparisons
can further demonstrate that our DCADCN significantly outperforms
4. Experiments compared SISR methods.
5
F. Qiao et al. Journal of Visual Communication and Image Representation 100 (2024) 104097
Table 1
Quantitative results on five benchmark datasets for different scale factors. The best and second best are highlighted in bold and underlined, respectively.
Method Params Scale Set5 Set14 B100 U100 M109
PSNR/SSIM PSNR/SSIM PSNR/SSIM PSNR/SSIM PSNR/SSIM
LESRCNN [3] – 37.65/0.9586 33.32/0.9148 31.95/0.8964 31.45/0.9206 –/–
ACNet [8] – 37.72/0.9588 33.41/0.9160 32.06/0.8978 31.79/0.9245 –/–
RepRFN [23] 386K 37.99/0.9609 33.57/0.9179 32.18/0.9004 31.95/0.9261 38.80/0.9774
VLESR [52] 311K 38.01/0.9605 33.58/0.9177 32.16/0.8993 32.14/0.9280 38.75/0.9770
FDIWN [24] – –/– –/– –/– –/– –/–
ShuffleMixer [21] 394K 38.01/0.9606 33.63/0.9180 32.17/0.8995 31.89/0.9257 38.83/0.9774
×2
ARRFN [19] 988K 38.01/0.9606 33.66/0.9179 32.20/0.8999 32.27/0.9295 –/–
FDSCSR [25] 466K 38.02/0.9606 33.51/0.9174 32.18/0.8996 32.24/0.9288 38.67/0.9771
DefRCN [10] – 38.02/0.9596 33.58/0.9151 32.21/0.8998 32.20/0.9286 –/–
RFDN [22] 534K 38.05/0.9606 33.68/0.9184 32.16/0.8994 32.12/0.9278 38.88/0.9773
RLFN [53] 527K 38.05/0.9607 33.68/0.9172 32.19/0.8997 32.17/0.9286 –/–
DCADCN [ours] 831K 38.11/0.9607 33.69/0.9191 32.21/0.8999 32.27/0.9297 38.85/0.9769
LESRCNN [3] – 33.93/0.9231 30.12/0.8380 28.91/0.8005 27.70/0.8415 –/–
ACNet [8] – 34.14/0.9247 30.19/0.8398 28.98/0.8023 27.97/0.8482 –/–
RepRFN [23] 392K 34.33/0.9272 30.30/0.8415 29.08/0.8058 27.95/0.8473 33.48/0.9434
VLESR [52] 319K 34.40/0.9272 30.34/0.8415 29.08/0.8043 28.16/0.8519 33.61/0.9445
FDIWN [24] 645K 34.46/0.9274 30.35/0.8423 29.10/0.8051 28.16/0.8528 –/–
ShuffleMixer [21] 415K 34.40/0.9272 30.37/0.8423 29.12/0.8051 28.08/0.8498 33.69/0.9448
×3
ARRFN [19] 996K 34.38/0.9272 30.36/0.8422 29.09/0.8050 28.22/0.8533 –/–
FDSCSR [25] 471K 34.42/0.9274 30.37/0.8429 29.10/0.8052 28.20/0.8532 33.55/0.9443
DefRCN [10] – 34.41/0.9263 30.34/0.8388 29.10/0.8044 28.16/0.8519 –/–
RFDN [22] 541K 34.41/0.9273 30.34/0.8420 29.09/0.8050 28.21 0.8525 33.67/0.9449
RLFN [53] – –/– –/– –/– –/– –/–
DCADCN [ours] 838K 34.46/0.9276 30.38/0.8427 29.11/0.8054 28.24/0.8534 33.76/0.9454
LESRCNN [3] 516K 31.88/0.8903 28.44/0.7772 27.45/0.7313 25.77/0.7732 –/–
ACNet [8] 1283K 31.83/0.8903 28.46/0.7788 27.48/0.7326 25.93/0.7798 –/–
RepRFN [23] 402K 32.15/0.8952 28.63/0.7824 27.60/0.7377 26.09/0.7834 30.52/0.9075
VLESR [52] 331K 32.17/0.8945 28.55/0.7802 27.55/0.7345 26.03/0.7830 30.48/0.9073
FDIWN [24] 664K 32.17/0.8941 28.55/0.7806 27.58/0.7364 26.02/0.7844 –/–
ShuffleMixer [21] 411K 32.21/0.8953 28.66/0.7827 27.61/0.7366 26.08/0.7835 30.65/0.9093
×4
ARRFN [19] 1008K 32.22/0.8952 28.60/0.7817 27.57/0.7355 26.09/0.7858 –/–
FDSCSR [25] 478K 32.25/0.8959 28.61/0.7821 27.58/0.7367 26.12/0.7866 30.51/0.9087
DefRCN [10] 1900K 32.21/0.8936 28.59/0.7810 27.57/0.7356 26.04/0.7841 –/–
RFDN [22] 550K 32.24/0.8952 28.61/0.7819 27.57/0.7360 26.11/0.7858 30.58/0.9089
RLFN [53] 543K 32.23/0.8961 28.61/0.7818 27.58/0.7359 26.15/0.7866 –/–
DCADCN [ours] 846K 32.27/0.8957 28.63/0.7819 27.58/0.7361 26.11/0.7858 30.58/0.9099
DCADCN-L [ours] 1031K 32.36/0.8962 28.70/0.7836 27.62/0.7373 26.16/0.7876 30.66/0.9109
Table 2 Table 3
Ablation results investigating the effects of DMFE. ✗: delete, &: replace. Ablation study investigating the effects of contrastive loss.
Case index DMFE Conv ADCM Set5 Set14 B100 Contrastive loss Set5 Set14 B100 U100 M109
1 ✗ ✗ ✗ 30.31 27.12 26.51 w 32.19 28.61 27.56 26.07 30.50
2 & ! ✗ 31.97 28.42 27.46 w/o 32.27 28.63 27.58 26.11 30.58
3 & ✗ ! 32.18 28.58 27.56
4 ! ! ! 32.27 28.63 27.58
6
F. Qiao et al. Journal of Visual Communication and Image Representation 100 (2024) 104097
Fig. 4. Visual comparisons of DCADCN with other competitors on different benchmark datasets (×4).
The FLOPs and inference time are effective metrics to evaluate the
computational cost. Here, we compare these metrics of our DCADCN-
Fig. 5. Ablation study on ADCM.
L with several SISR methods including LESRCNN [3], ACNet [8],
RepRFN [23], VLESR [52], FDIWN [24], ShuffleMixer [21], ARRFN
[19], FDSCSR [25], RLFN [53], and RFDN [22] on Set5 of ×4 SISR.
FLOPs are evaluated on a 1280 × 720 HR image, inference time are
counted on testing Set5 (×4) in a feed-forward process.
7
F. Qiao et al. Journal of Visual Communication and Image Representation 100 (2024) 104097
better model ACNet-M, our DCADCN needs less computational cost (see
Fig. 9).
Qualitative analysis: The visual comparisons of different methods for
SISR with different noise on different benchmark datasets (×4) are
shown in Fig. 10 (𝜎 = 25) and Fig. 11 (𝜎 = 35). As we can see, our
method is better at removing noise and recovering clearer and more
accurate image textures at different noise levels. In contrast, other
competitors are susceptible to noise, resulting in blurry SR images.
5. Discussion
8
F. Qiao et al. Journal of Visual Communication and Image Representation 100 (2024) 104097
Fig. 9. Average PSNR vs. FLOPs and inference time for SISR.
Fig. 10. Visual comparisons for SISR with noise level of 25.
Fig. 11. Visual comparisons for SISR with noise level of 35.
9
F. Qiao et al. Journal of Visual Communication and Image Representation 100 (2024) 104097
Table 4
Quantitative results for SISR with blind noise on different noise levels of 15, 25, 35, and 50. The best results are highlighted in bold.
Method Noise level Set5 Set14 B100 U100
PSNR/SSIM PSNR/SSIM PSNR/SSIM PSNR/SSIM
LESRCNN [3] 28.73/0.8177 26.56/0.7049 25.97/0.6557 24.29/0.7066
ACNet-M [8] 15 28.71/0.8177 26.59/0.7056 25.98/0.6563 24.33/0.7078
DCADCN [ours] 29.02/0.8268 26.72/0.7116 26.04/0.6609 24.51/0.7177
LESRCNN [3] 27.43/0.7849 25.65/0.6698 25.18/0.6203 23.55/0.6709
ACNet-M [8] 25 27.40/0.7848 25.67/0.6710 25.21/0.6221 23.63/0.6751
DCADCN [ours] 27.55/0.7911 25.77/0.6761 25.24/0.6262 23.78/0.6858
LESRCNN [3] 26.40/0.7580 24.88/0.6419 24.56/0.5954 22.99/0.6437
ACNet-M [8] 35 26.41/0.7587 24.93/0.6441 24.61/0.5970 23.05/0.6482
DCADCN [ours] 26.43/0.7630 24.98/0.6468 24.62/0.6005 23.15/0.6585
LESRCNN [3] 25.23/0.7246 24.00/0.6113 23.86/0.5681 22.25/0.6092
ACNet-M [8] 50 25.25/0.7254 24.03/0.6122 23.89/0.5693 22.33/0.6149
DCADCN [ours] 25.25/0.7307 24.07/0.6156 23.89/0.5731 22.40/0.6249
We propose a DCADCN for effective and efficient SISR. The main The authors declare that they have no known competing finan-
advantages of DCADCN are: (1) Improved image quality. The experi- cial interests or personal relationships that could have appeared to
mental results demonstrate that our method generates high-resolution influence the work reported in this paper.
images with improved visual quality compared to existing methods.
(2) Enhanced details. DCADCN can effectively recover fine details and Data availability
textures in SR images. The proposed ADCM enables to maintain the
overall structure and edge of the image during the SR process. (3) Data will be made available on request.
Wide range of application scenarios. Our method can effectively handle
several SISR tasks, such as classic SISR, SISR with blind noise, and real- Acknowledgments
world SISR. (4) Efficient computation. The proposed method has high
computational efficiency since it requires very low inference time to This document is the results of the research projects funded by the
process LR images. National Natural Science Foundation of China (No. 11571325) and the
While our method has made progress on classic image SR, SISR Fundamental Research Funds for the Central Universities, China (No.
with blind noise, and real-world SISR, there are still some limitations CUC2019 A002). This work is supported by Public Computing Cloud,
that need to be addressed: (1) Compared to some lightweight models, CUC.
our method requires more parameters and FLOPs. (2) When switching
tasks, our DCADCN needs to be retrained with newly synthesized References
data that has the same distribution, indicating it might not generalize
well to different types of images beyond those used for training. In [1] Y. Yoon, H. Jeon, D. Yoo, J. Lee, I.S. Kweon, Light-field image super-resolution
future research, we would like to explore more lightweight models and using convolutional neural network, IEEE Signal Process. Lett. 24 (6) (2017)
improve its ability to migrate across different data domains. 848–852.
[2] C. Ledig, L. Theis, F. Huszár, J. Caballero, A. Cunningham, A. Acosta, A. Aitken,
A. Tejani, J. Totz, Z. Wang, et al., Photo-realistic single image super-resolution
6. Conclusion using a generative adversarial network, in: Proceedings of the IEEE Conference
on Computer Vision and Pattern Recognition, 2017, pp. 4681–4690.
In this paper, we propose a dual contrastive attention-guided de- [3] C. Tian, R. Zhuge, Z. Wu, Y. Xu, W. Zuo, C. Chen, C.W. Lin, Lightweight image
formable convolutional network (DCADCN) for effective and efficient super-resolution with enhanced CNN, Knowl.-Based Syst. 205 (2020) 106235.
[4] B. Lim, S. Son, H. Kim, S. Nah, K. Mu Lee, Enhanced deep residual networks
SISR. Specifically, the joint inner and external attention mechanisms for single image super-resolution, in: Proceedings of the IEEE Conference on
enable the ADCM to model the correspondences between input and Computer Vision and Pattern Recognition Workshops, 2017, pp. 136–144.
deformation features and retain more spatial structural information. [5] Y. Tai, J. Yang, X. Liu, Image super-resolution via deep recursive residual
The dual mixed feature extractor (DMFE) utilizes both Conv layers network, in: Proceedings of the IEEE Conference on Computer Vision and Pattern
Recognition, 2017, pp. 3147–3155.
and ADCMs to obtain diverse and complementary features. Further-
[6] Y. Zhang, K. Li, K. Li, L. Wang, B. Zhong, Y. Fu, Image super-resolution using
more, DCADCN adopts contrastive loss to further amplify the role of very deep residual channel attention networks, in: Proceedings of the European
key features and mitigate the effect of redundant noisy information. Conference on Computer Vision, ECCV, 2018, pp. 286–301.
Experimental results show the superior performance of DCADCN on [7] Z. Zhang, X. Wang, C. Jung, DCSR: Dilated convolutions for single image
classic SISR, SISR with blind noise, and real-world SISR tasks, and super-resolution, IEEE Trans. Image Process. 28 (4) (2018) 1625–1635.
[8] C. Tian, Y. Xu, W. Zuo, C. Lin, D. Zhang, Asymmetric CNN for image
comprehensive ablation studies demonstrate the effectiveness of ADCM,
super-resolution, IEEE Trans. Syst. Man Cybern.: Syst. 52 (6) (2021) 3718–3730.
DMFE, and contrastive loss. Moreover, our DCADCN achieves a better [9] J. Dai, H. Qi, Y. Xiong, Y. Li, G. Zhang, H. Hu, Y. Wei, Deformable convolutional
trade-off between performance and computational cost. In the future, networks, in: Proceedings of the IEEE International Conference on Computer
we will explore the application of our model to other image restoration Vision, 2017, pp. 764–773.
tasks, such as deblurring and demosaicing. [10] Y. Zhang, Y. Sun, S. Liu, Deformable and residual convolutional network for
image super-resolution, Appl. Intell. 52 (1) (2022) 295–304.
[11] C. Zhao, W. Zhu, S. Feng, Superpixel guided deformable convolution network
CRediT authorship contribution statement for hyperspectral image classification, IEEE Trans. Image Process. 31 (2022)
3838–3851.
Fengjuan Qiao: Writing – original draft, Visualization, Validation, [12] Y. Huang, X. Hou, Y. Dun, J. Qin, L. Liu, X. Qian, L. Shao, Learning deformable
Methodology, Investigation, Formal analysis, Data curation, Conceptu- and attentive network for image restoration, Knowl.-Based Syst. 231 (2021)
107384.
alization. Yonggui Zhu: Writing – review & editing, Project adminis- [13] Y. Wang, J. Yang, L. Wang, X. Ying, T. Wu, W. An, Y. Guo, Light field image
tration, Funding acquisition. Guofang Li: Writing – review & editing, super-resolution using deformable convolution, IEEE Trans. Image Process. 30
Conceptualization. Bin Li: Software, Resources. (2020) 1057–1071.
10
F. Qiao et al. Journal of Visual Communication and Image Representation 100 (2024) 104097
[14] C. Dong, C.C. Loy, K. He, X. Tang, Image super-resolution using deep con- [36] F. Yang, H. Yang, J. Fu, H. Lu, B. Guo, Learning texture transformer network for
volutional networks, IEEE Trans. Pattern Anal. Mach. Intell. 38 (2) (2015) image super-resolution, in: Proceedings of the IEEE/CVF Conference on Computer
295–307. Vision and Pattern Recognition, 2020, pp. 5791–5800.
[15] J. Kim, J.K. Lee, K.M. Lee, Accurate image super-resolution using very deep [37] T. Chen, S. Kornblith, M. Norouzi, G. Hinton, A simple framework for contrastive
convolutional networks, in: Proceedings of the IEEE Conference on Computer learning of visual representations, in: International Conference on Machine
Vision and Pattern Recognition, 2016, pp. 1646–1654. Learning, PMLR, 2020, pp. 1597–1607.
[16] J. Kim, J.K. Lee, K.M. Lee, Deeply-recursive convolutional network for image [38] C.-H. Yeh, C.-Y. Hong, Y.-C. Hsu, T.-L. Liu, Y. Chen, Y. LeCun, Decoupled
super-resolution, in: Proceedings of the IEEE Conference on Computer Vision contrastive learning, in: European Conference on Computer Vision, Springer,
and Pattern Recognition, 2016, pp. 1637–1645. 2022, pp. 668–684.
[17] Y. Tai, J. Yang, X. Liu, C. Xu, Memnet: A persistent memory network for image [39] H. Wang, X. Guo, Z.-H. Deng, Y. Lu, Rethinking minimal sufficient representation
restoration, in: Proceedings of the IEEE International Conference on Computer in contrastive learning, in: Proceedings of the IEEE/CVF Conference on Computer
Vision, 2017, pp. 4539–4547. Vision and Pattern Recognition, 2022, pp. 16041–16050.
[18] J. Liang, J. Cao, G. Sun, K. Zhang, L. Van Gool, R. Timofte, Swinir: Image [40] G. Wu, J. Jiang, X. Liu, A practical contrastive learning framework for
restoration using swin transformer, in: Proceedings of the IEEE/CVF International single-image super-resolution, IEEE Trans. Neural Netw. Learn. Syst. (2023).
Conference on Computer Vision, 2021, pp. 1833–1844. [41] L. Wang, Y. Wang, X. Dong, Q. Xu, J. Yang, W. An, Y. Guo, Unsupervised
[19] J. Qin, R. Zhang, Lightweight single image super-resolution with attentive degradation representation learning for blind super-resolution, in: Proceedings
residual refinement network, Neurocomputing 500 (2022) 846–855. of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021,
[20] C. Dong, C.C. Loy, X. Tang, Accelerating the super-resolution convolutional pp. 10581–10590.
neural network, in: European Conference on Computer Vision, Springer, 2016, [42] Y. Zhang, L. Dong, H. Yang, L. Qing, X. He, H. Chen, Weakly-supervised
pp. 391–407. contrastive learning-based implicit degradation modeling for blind image
[21] L. Sun, J. Pan, J. Tang, Shufflemixer: An efficient convnet for image super-resolution, Knowl.-Based Syst. (2022) 108984.
super-resolution, Adv. Neural Inf. Process. Syst. 35 (2022) 17314–17326. [43] W. Shi, J. Caballero, F. Huszár, J. Totz, A.P. Aitken, R. Bishop, D. Rueckert,
[22] J. Liu, J. Tang, G. Wu, Residual feature distillation network for lightweight image Z. Wang, Real-time single image and video super-resolution using an efficient
super-resolution, in: European Conference on Computer Vision, Springer, 2020, sub-pixel convolutional neural network, in: Proceedings of the IEEE Conference
pp. 41–55. on Computer Vision and Pattern Recognition, 2016, pp. 1874–1883.
[23] W. Deng, H. Yuan, L. Deng, Z. Lu, Reparameterized residual feature network for [44] Y. Chen, J. Li, H. Xiao, X. Jin, S. Yan, J. Feng, Dual path networks, Adv. Neural
lightweight image super-resolution, in: Proceedings of the IEEE/CVF Conference Inf. Process. Syst. 30 (2017).
on Computer Vision and Pattern Recognition, 2023, pp. 1712–1721. [45] Z. Liu, B. Yang, G. Duan, J. Tan, Visual defect inspection of metal part surface
[24] G. Gao, W. Li, J. Li, F. Wu, H. Lu, Y. Yu, Feature distillation interaction weighting via deformable convolution and concatenate feature pyramid neural networks,
network for lightweight image super-resolution, in: Proceedings of the AAAI IEEE Trans. Instrum. Meas. 69 (12) (2020) 9681–9694.
Conference on Artificial Intelligence, Vol. 36, 2022, pp. 661–669, (1). [46] R. Timofte, E. Agustsson, L. Van Gool, M.-H. Yang, L. Zhang, Ntire 2017
[25] Z. Wang, G. Gao, J. Li, H. Yan, H. Zheng, H. Lu, Lightweight feature de- challenge on single image super-resolution: Methods and results, in: Proceedings
redundancy and self-calibration network for efficient image super-resolution, of the IEEE Conference on Computer Vision and Pattern Recognition Workshops,
ACM Trans. Multimedia Comput. Commun. Appl. 19 (3) (2023) 1–15. 2017, pp. 114–125.
[26] K. Zhang, W. Zuo, L. Zhang, Deep plug-and-play super-resolution for arbitrary [47] M. Bevilacqua, A. Roumy, C. Guillemot, M.L. AlberiMorel, Low-complexity
blur kernels, in: Proceedings of the IEEE/CVF Conference on Computer Vision single-image super-resolution based on nonnegative neighbor embedding, in:
and Pattern Recognition, 2019, pp. 1671–1681. Proceedings of the 23rd British Machine Vision Conference, BMVA Press, 2012,
[27] K. Zhang, J. Liang, L. Van Gool, R. Timofte, Designing a practical degradation pp. 1–10.
model for deep blind image super-resolution, in: Proceedings of the IEEE/CVF [48] R. Zeyde, M. Elad, M. Protter, On single image scale-up using sparse-
International Conference on Computer Vision, 2021, pp. 4791–4800. representations, in: International Conference on Curves and Surfaces, Springer,
[28] X. Zhu, H. Hu, S. Lin, J. Dai, Deformable convnets v2: More deformable, better 2010, pp. 711–730.
results, in: Proceedings of the IEEE/CVF Conference on Computer Vision and [49] D. Martin, C. Fowlkes, D. Tal, J. Malik, A database of human segmented natural
Pattern Recognition, 2019, pp. 9308–9316. images and its application to evaluating segmentation algorithms and measuring
[29] X. Wang, K.C. Chan, K. Yu, C. Dong, C. Change Loy, Edvr: Video restoration with ecological statistics, in: Proceedings Eighth IEEE International Conference on
enhanced deformable convolutional networks, in: Proceedings of the IEEE/CVF Computer Vision. ICCV 2001, Vol. 2, IEEE, 2001, pp. 416–423.
Conference on Computer Vision and Pattern Recognition Workshops, 2019. [50] J.B. Huang, A. Singh, N. Ahuja, Single image super-resolution from transformed
[30] Z. Luo, L. Yu, X. Mo, Y. Li, L. Jia, H. Fan, J. Sun, S. Liu, EBSR: Feature self-exemplars, in: Proceedings of the IEEE Conference on Computer Vision and
enhanced burst super-resolution with deformable alignment, in: Proceedings of Pattern Recognition, 2015, pp. 5197–5206.
the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, [51] Y. Matsui, K. Ito, Y. Aramaki, A. Fujimoto, T. Ogawa, T. Yamasaki, K. Aizawa,
pp. 471–478. Sketch-based manga retrieval using manga109 dataset, Multimedia Tools Appl.
[31] Z. Liu, W. Lin, X. Li, Q. Rao, T. Jiang, M. Han, H. Fan, J. Sun, S. Liu, 76 (20) (2017) 21811–21838.
Adnet: Attention-guided deformable convolutional network for high dynamic [52] D. Gao, D. Zhou, A very lightweight and efficient image super-resolution
range imaging, in: Proceedings of the IEEE/CVF Conference on Computer Vision network, Expert Syst. Appl. 213 (2023) 118898.
and Pattern Recognition, 2021, pp. 463–470. [53] F. Kong, M. Li, S. Liu, D. Liu, J. He, Y. Bai, F. Chen, L. Fu, Residual local
[32] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A.N. Gomez, Ł. Kaiser, feature network for efficient super-resolution, in: Proceedings of the IEEE/CVF
I. Polosukhin, Attention is all you need, Adv. Neural Inf. Process. Syst. 30 (2017). Conference on Computer Vision and Pattern Recognition, 2022, pp. 766–776.
[33] Y. Mei, Y. Fan, Y. Zhou, Image super-resolution with non-local sparse attention, [54] X. Wang, K. Yu, S. Wu, J. Gu, Y. Liu, C. Dong, Y. Qiao, C. Change Loy, Esrgan:
in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Enhanced super-resolution generative adversarial networks, in: Proceedings of
Recognition, 2021, pp. 3517–3526. the European Conference on Computer Vision (ECCV) Workshops, 2018.
[34] H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, H. Jégou, Train- [55] X. Ji, Y. Cao, Y. Tai, C. Wang, J. Li, F. Huang, Real-world super-resolution
ing data-efficient image transformers & distillation through attention, in: via kernel estimation and noise injection, in: Proceedings of the IEEE/CVF
International Conference on Machine Learning, PMLR, 2021, pp. 10347–10357. Conference on Computer Vision and Pattern Recognition Workshops, 2020, pp.
[35] M. Huang, Y. Liu, Z. Peng, C. Liu, D. Lin, S. Zhu, N. Yuan, K. Ding, L. Jin, 466–467.
Swintextspotter: Scene text spotting via better synergy between text detection [56] J. Liang, H. Zeng, L. Zhang, Efficient and degradation-adaptive network for
and text recognition, in: Proceedings of the IEEE/CVF Conference on Computer real-world image super-resolution, in: European Conference on Computer Vision,
Vision and Pattern Recognition, 2022, pp. 4593–4603. Springer, 2022, pp. 574–591.
11