0% found this document useful (0 votes)
6 views11 pages

Dual Contrastive Attention for Image SR

Uploaded by

Lý Phi Cường
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
6 views11 pages

Dual Contrastive Attention for Image SR

Uploaded by

Lý Phi Cường
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

J. Vis. Commun. Image R.

100 (2024) 104097

Contents lists available at ScienceDirect

J. Vis. Commun. Image R.


journal homepage: [Link]/locate/jvci

Full length article

Dual contrastive attention-guided deformable convolutional network for


single image super-resolution✩ , ✩✩
Fengjuan Qiao a , Yonggui Zhu b ,∗, Guofang Li a , Bin Li c
a
School of Information and Communication Engineering, Communication University of China, Beijing 100024, China
b
School of Data Science and Intelligent Media, Communication University of China, Beijing 100024, China
c
School of Mathematics and Statistics, Qilu University of Technology (Shandong Academy of Sciences), Jinan 250353, China

ARTICLE INFO ABSTRACT

Keywords: With its powerful ability to model geometric transformations, the deformable convolutional network brings
Image super-resolution great improvements for single image super-resolution (SISR). Nevertheless, its location-variant sampling
Deformable convolution method leads to an escalation in spatial variance as the deformable convolutional layers are stacked,
Contrastive learning
consequently resulting in limited performance. Hence, we propose a novel and effective approach called dual
Attention mechanism
contrastive attention-guided deformable convolutional network (DCADCN) for SISR modeling. Specifically,
we propose an attention-guided deformable convolutional module with joint inner and external attention
mechanisms to fully exploit the correspondences between input and deformation features and preserve spatial
characteristics to the extent possible. Additionally, we propose a dual mixed feature extractor consisting of two
parallel sub-paths. This design allows for the learning of diverse and complementary spatial features. Further-
more, contrastive learning is applied to further amplify the role of key features and mitigate the interference
of noisy features. Extensive experimental results demonstrate that DCADCN is capable of effectively handling
classic SISR, SISR with blind noise, and real-world SISR tasks. Moreover, our method achieves comparable or
even better performance with lower computational cost compared to state-of-the-art methods.

1. Introduction deformations of CNN, such as dilated convolution [7] and symmetric


convolution [8], have been proposed to expand the receptive field
Single image super-resolution (SISR) plays a crucial role in various without introducing additional parameters. Nevertheless, traditional
domains, including satellite imaging and surveillance monitoring. The CNN-based methods typically sample the LR image at fixed intervals,
primary objective of SISR is to estimate a high-resolution (HR) image and thus only local information contained in a small fixed receptive re-
from its low-resolution (LR) counterpart. However, since multiple HR gion is extracted. Consequently, such location-fixed sampling methods
images can degrade into the same LR image, SISR inherently becomes yield limited SISR performance, particularly in complex and irregular
an ill-posed task, posing a significant challenge in reconstructing high- scenarios.
quality HR images. To address this challenge, numerous learning-based In contrast to general CNNs, deformable convolutional networks
approaches have been proposed to establish a non-linear mapping (DCNs) [9] have several key advantages. Firstly, DCNs enable adaptive
between LR and HR image pairs.
sampling based on the input data, facilitating more precise feature
Convolutional neural networks (CNNs) have emerged as a funda-
extraction. Secondly, they exhibit high robustness to model complex
mental component in the field of SISR due to their powerful feature
and irregular geometric transformations. Lastly, DCNs reduce the grid
representation capabilities [1–3]. The feature extraction ability of CNN
effect by incorporating deformable offsets, thereby mitigating grid-
is deeply dependent on the effectiveness of convolutional receptive
like artifacts in the feature maps. Exploiting these advantages, Zhang
field [4]. Therefore, some studies have attempted to enlarge the re-
et al. [10] introduced DCN to the SISR domain and developed a de-
ceptive field by deepening or widening the CNN [5,6]. However, these
deeper or wider networks come at the cost of substantial computa- formable and residual convolutional network (DefRCN), which utilized
tional resources and memory consumption. To overcome this issue, residually connected deformation convolutional blocks to fully exploit

✩ This document is the results of the research projects funded by the National Natural Science Foundation of China (No. 11571325) and the Fundamental
Research Funds for the Central Universities, China (No. CUC2019 A002).
✩✩ This paper has been recommended for acceptance by Zicheng Liu.
∗ Corresponding author.
E-mail addresses: qiao_fj@[Link] (F. Qiao), ygzhu@[Link] (Y. Zhu), gfli@[Link] (G. Li), ribbenlee@[Link] (B. Li).

[Link]
Received 30 July 2023; Received in revised form 14 February 2024; Accepted 24 February 2024
Available online 27 February 2024
1047-3203/© 2024 The Author(s). Published by Elsevier Inc. This is an open access article under the CC BY license ([Link]
F. Qiao et al. Journal of Visual Communication and Image Representation 100 (2024) 104097

the deformation features present in LR images. Although DCN-based persistent memory. In the second approach, Liang et al. [18] proposed
methods have demonstrated a performance enhancement in SISR [11, a baseline framework based on the Transformer architecture for im-
12], they still suffer from the following two limitations. Firstly, the age restoration. Qin et al. [19] utilized attentive residual mechanism
irregular sampling method employed by DCN introduces spatial vari- to enhance the feature extraction capacity and replaced regular up-
ance, resulting in the gradual loss of spatial information with increasing sampling operation with a multi-scale separable up-sampling module
network depth [13]. Secondly, both regular and irregular sampling for better performance. However, both types of methods inevitably
methods are crucial, yet DCN primarily focuses on irregular features face challenges in terms of computational requirements and memory
while neglecting regular features. consumption, limiting their practical applicability.
To practically remedy these limitations, we propose a dual con- To accelerate training efficiency and reduce memory consumption,
trastive attention-guided deformable convolutional network (DCADCN) researchers have focused on reducing the input size or compressing
for SISR. Specifically, to preserve spatial information, we employ an the SISR model. For reducing the input size, Dong et al. [20] adopted
attention-guided deformable convolutional module (ADCM) to adap- LR images as input instead of up-sampled HR images, performing an
tively weight features by modeling the correspondences between input upscale operation at the final layer of the network to predict HR
and deformation features. To obtain diverse features, we design a dual images. For model compression, Sun et al. [21] utilized a large kernel
mixed feature extractor (DMFE) comprising two parallel sub-paths to ConvNet and a feature fusion for mobile-friendly SISR. Liu et al. [22]
capture regular and irregular features, respectively. Furthermore, we proposed feature distillation connection and established a residual
incorporate contrastive learning as part of the loss function to train our feature distillation network for effective SISR while keep lightweight.
model. In summary, our main contributions are summarized as follow: Subsequently, several networks, such as RepRFN [23], FDIWN [24],
and FDSCSR [25] combined CNNs with information distillation to
• We present an attention-guided deformable convolution to adap- improve training efficiency.
tively weight features, which utilize joint inner and external at- For complex SISR scenarios, it is well known that digital images
tention mechanisms to fully explore the correspondences between are affected by various factors, such as unknown noise, shooting en-
input and deformation features. vironment, and optical performance of equipment. To address SISR
• We exploit a dual structure, called DMFE, to learn diverse and problems with blind noise, Tian et al. [8] proposed an ACNet that em-
complementary spatial characteristics by stacked convolutional ploys a asymmetric architecture and flexible up-sampling mechanism.
layers and ADCMs. Zhang et al. [26] designed a complex degradation model considering
• A contrastive loss item is designed and embedded into the ADCM blur kernel and noise, generating clean images and using a plug-and-
to further amplify the role of key features and mitigate the inter- play algorithm to obtain a noise prior and the final SR image. Later,
ference of redundant noisy information. To our best knowledge, Zhang et al. [27] proposed a more complex but practical degradation
we are the first to apply contrastive learning to DCN to constrain model for real-world SISR, which consists of randomly shuffled blur,
the correspondences between input and deformation features. downsampling and noise degradations.
• Experimental results show that the proposed method can handle Despite the significant advancements in SISR research, there is still
the tasks of classic SISR, SISR with blind noise, and real-world a urgent need for a network that can simultaneously enhance SISR
SISR, and achieves a better trade-off between performance and performance, improve training efficiency, and handle complex SISR
computational cost. scenarios.

The remaining sections of this paper are organized as follows. 2.2. Deformable Convolutional Network (DCN)
Section 2 provides a comprehensive review of the related work in the
areas of SISR, DCNs, attention mechanisms, and contrastive learning. DCN was proposed to model geometric transformation of objects [9,
In Section 3, we present our proposed DCADCN and provide a detailed 28]. Recent works [29,30] introduced DCN for frame alignment in
description of each component within the network. Section 4 outlines video processing area, thus making more efficient use of information
the experimental setup and presents the performance evaluation of from adjacent frames. Liu et al. [31] proposed an attention-guided
DCADCN in the context of classic SISR, SISR with blind noise, and real- deformable convolutional network (ADNet) for high dynamic range
world SISR. Section 5 discusses the differences between our work and imaging. ADNet was constructed with two branches, the one utilizing
existing methods and points out the strengths and limitations of the the spatial attention module for extracting attention features and the
present approach. Finally, Section 6 concludes this paper. other adopting DCN to align the gamma-corrected images. The DCN
can also be used for SISR. For instance, Zhang et al. [10] developed a
2. Related work DefRCN for SISR, which applied DCN to fully exploit the deformation
features present in LR images. Despite showing great potential over
2.1. Single image super-resolution CNN, DefRCN still suffers from two major drawbacks. Firstly, DefRCN
does not consider the correspondences between input and learned
Significant progress has been made in recent years in the field of features, resulting in the loss of spatial information. Second, it contains
SISR with the advent of deep learning techniques. Dong et al. [14] an excessive number of trainable parameters. Addressing these two
made groundbreaking contributions by being the first to develop a issues in our network could theoretically improve SISR performance.
CNN framework for SISR, which yielded improved performance. Subse-
quently, various deep learning methods have been developed to further 2.3. Attention mechanism
enhance SISR performance [15–19], improve training efficiency [20–
25] and address complex SISR scenarios [8,26,27], respectively. The attention mechanism is a unique architecture that automatically
To enhance SISR performance, Kim et al. [15] introduced a very determines the contributions of input to output [32]. Through such a
deep CNN, dubbed VDSR, which employed a residual learning strategy. mechanism, the efficiency of information utilization can be significantly
This research demonstrated that deeper CNNs lead to better reconstruc- improved. In recent years, attention mechanisms have gained popular-
tion results. However, CNN-based approaches that primarily extract ity in various domains, including image processing [33,34] and natural
local features, resulting in limited performance in modeling long-range language processing [35]. They have also been applied to enhance
dependencies. To tackle this issue, two main approaches have been the performance of SISR. For instance, Yang et al. [36] employed LR
explored: recurrent structures and non-local methods. In the first ap- and HR images as queries and keys in an attention mechanism to
proach, Tai et al. introduced memory blocks in MemNet [17] to capture explore the correspondences between them, resulting in more accurate

2
F. Qiao et al. Journal of Visual Communication and Image Representation 100 (2024) 104097

Fig. 1. Overall architecture of the proposed dual contrastive attention-guided deformable convolutional network (DCADCN) for SISR, which mainly consists of dual mixed feature
extractors (DMFEs) and an up-sampling block.

texture reconstruction. Nevertheless, the commonly used self-attention 3.2. Dual Mixed Feature Extractor (DMFE)
mechanism requires computing the relevance between each pixel in
images, leading to a substantial computational workload. Therefore, a Motivated by the dual-path network [44], we introduce the concept
simpler attention architecture needs to be explored for our network. of dual path to our network and propose DMFE as a basic component.
Differently, the dual-path network only uses Conv layers, while we use
not only Conv layers but also the proposed ADCMs in the two sub-
2.4. Contrastive learning paths. Moreover, our DMFE is designed with two feature fusion and
enhancement stages to extract richer features. More details of ADCM
will be given in Section 3.3.
Contrastive Learning is a machine learning method that trains mod-
The architecture of DMFE is illustrated in Fig. 1. Let 𝑋𝑐𝑘 denote the
els by comparing the similarities and differences between different
input of 𝑘th (1 ≤ 𝑘 ≤ 𝑑) DMFE with 𝑐 channels, the DMFE first splits 𝐹𝑐𝑘
samples. In recent years, contrastive learning has made significant 𝑘
into two parts with the same channels: 𝑋1∶𝑐∕2 (abbreviated as 𝑋1 ) and
advancements in computer vision and natural language processing [37– 𝑘
𝑋𝑐∕2+1∶𝑐 (abbreviated as 𝑋2 ). These two parts are fed into two parallel
39]. However, there has been limited research on applying contrastive
learning to low-level visual tasks like SISR. For instance, Wu et al. [40] sub-paths for two-stage feature extraction and enhancement. During
employed effective data enhancement strategies to generate positive the first stage, 3 × 3 Conv layers and ADCMs are utilized in two sub-
and negative samples for contrastive learning and achieved a perfor- paths respectively to extract different spatial features. These features
are concatenated in the channel dimension and then the number of
mance improvement on SISR issue. Leveraging the growing popularity
channels is recovered by a 1 × 1 Conv layer. In addition, ESA is utilized
of self-supervised learning, contrastive learning can also be used for
in two paths to enhance the feature information. In this way, the feature
blind SISR [41,42] to implicitly model different degradation processes.
extraction and enhancement procedure in stage 1 is formulated as:
The main difference between our work and the above methods is that
the positive and negative samples are generated in different ways. ( ( ( ( ))))
𝑋1 ′ = 𝑅𝐿 𝑓 3×3 𝑅𝐿 𝑓 3×3 𝑋1 ,
𝑐𝑜𝑛𝑣 𝑐𝑜𝑛𝑣

( ( ( ( ))))
𝑋2 = 𝑅𝐿 𝑓𝑎𝑑𝑐𝑚 𝑅𝐿 𝑓𝑎𝑑𝑐𝑚 𝑋2
3. Proposed method ( [ ]) (1)
𝑋1𝑆1 = 𝑓𝑒𝑠𝑎 𝑓 1×1 𝑋1 ′ , 𝑋2 ′
𝑐𝑜𝑛𝑣
( [ ])
3.1. Overview
𝑋2𝑆1 = 𝑓𝑒𝑠𝑎 𝑓 1×1 𝑋2 ′ , 𝑋1 ′ ,
𝑐𝑜𝑛𝑣

The overall architecture of the proposed DCADCN is illustrated in where 𝑋1𝑆1 and 𝑋2𝑆1 are the outputs of the two sub-paths in stage 1,
3×3 and 𝑓 1×1 represent the function of 3 × 3 and 1 × 1
respectively. 𝑓𝑐𝑜𝑛𝑣
Fig. 1. Given a LR image 𝐼𝐿𝑅 ∈ Rℎ×𝑤×𝑐 , where ℎ × 𝑤 represents the 𝑐𝑜𝑛𝑣
Conv layers, respectively. 𝑓𝑎𝑑𝑐𝑚 denotes the function of ADCM, 𝑅𝐿(⋅)
spatial resolution and 𝑐 is the number of channels, a Conv layer is first
is a ReLU activation, [⋅] is the channel-wise concatenation, and 𝑓𝑒𝑠𝑎
utilized for shallow features extraction. In the network backbone, we
denotes the function of ESA module, the details of which can be found
cascade 𝑑 dual mixed feature extractors (DMFEs) to obtain rich features
in [22].
for HR image reconstruction. To excavate diverse features, the DMFE During the second stage, 3 × 3 Conv layers and ADCMs are also
contains two parallel sub-paths that utilize Conv layers and ADCMs utilized in two sub-paths to extract spatial features. These features
respectively to cargo complementary information. These features from are residually summed with their corresponding inputs and then con-
sub-paths are cross-fertilized and enhanced by channel-wise concate- catenated in channel dimension. Finally, the feature information is
nation operation and the enhanced spatial attention (ESA) [22]. After enhanced by an ESA module. The output 𝑋 𝑘+1 of the 𝑘th DMFE is
several stacked DMFEs, a residual connection is used to aggregate fea- calculated by:
tures of different levels. Finally, an up-sampling block [43] is employed ( ( ( ( ))))
𝑋1 ′′ = 𝑅𝐿 𝑓 3×3 𝑅𝐿 𝑓 3×3 𝑋1𝑆1 + 𝑋1 ,
to obtain residual map, which is added to the bilinear up-sampling of (
𝑐𝑜𝑛𝑣
( (
𝑐𝑜𝑛𝑣
( 𝑆1 ))))
the input to attain the super-resolved image 𝐼𝑆𝑅 ∈ R𝐻×𝑊 ×𝑐 , where ′′
𝑋2 = 𝑅𝐿 𝑓𝑎𝑑𝑐𝑚 𝑅𝐿 𝑓𝑎𝑑𝑐𝑚 𝑋2 + 𝑋2 , (2)
𝑘+1
[ ′′ ′′
]
𝐻 > ℎ, 𝑊 > 𝑤. 𝑋 = 𝑓𝑒𝑠𝑎 𝑋1 , 𝑋2 ,

3
F. Qiao et al. Journal of Visual Communication and Image Representation 100 (2024) 104097

Then, to fully capture channel-wise dependencies, we use a gating


mechanism to obtain the weight of each channel. Inspired by [6], the
gating mechanism contains two linear projection layers following a
ReLU and a sigmoid activation respectively. Finally, these weights are
utilized to rescale the channel-wise statistic:
( )
𝑦 = 𝑠 𝑋̂ ⋅ 𝑋,̂ (4)

where 𝑠(⋅) denotes the gating mechanism, represented as 𝑠(𝑋) =


𝑆𝑖(𝑊1 (𝑅𝐿(𝑊2 𝑋))), where 𝑊1 and 𝑊2 are the weights of linear projec-
tion functions, 𝑆𝑖(.) is the sigmoid activation.
External attention mechanism. We propose an external attention
mechanism to further capture the correspondence between channels of
input and deformation features.
First, a deformable Conv layer [28] is employed to obtain the
deformation features. Suppose 𝑤𝑙 represent the value at 𝑙th location
of Conv kernel, the deformation feature 𝐷 of input 𝑋 at location 𝑝 can
be obtained by

𝐿
( )
𝐷(𝑝) = 𝑤𝑙 ⋅ 𝑋 𝑝 + 𝑝𝑙 + 𝛥𝑝𝑙 ⋅ 𝛥𝑚𝑙 , (5)
𝑙=1

where 𝑝𝑙 and 𝛥𝑝𝑙 denote the pre-specified and the learned offsets for
the 𝑙–th location of input 𝑋, respectively. 𝛥𝑚𝑙 is a modulation scalar.
Next, we take 𝑋 and 𝐷 as the inputs to the inner attention mech-
anism, respectively. Their corresponding outputs are regarded as the
two basic components of external attention mechanism: query (𝑞) and
key (𝑘). Then the relevance of 𝑞 ∈ R𝑐 and 𝑘 ∈ R𝑐 is computed to obtain
Fig. 2. Calculation flow chart of attention-guided deformable convolutional module
the attention weights 𝐴. For each 𝑞𝑖 in 𝑞 and 𝑘𝑗 in 𝑘, the value of 𝐴 at
(ADCM). position (𝑖, 𝑗) is determined by normalized inner product and a softmax
activation:
⟨ ⟩
𝑞𝑖 𝑘𝑗
𝐴𝑖,𝑗 = sof tmax , , 𝐴𝑖,𝑗 ∈ 𝐴, 𝐴 ∈ R𝑐×𝑐 . (6)
3.3. Attention-guided deformable convolutional module ‖𝑞𝑖 ‖ ‖ ‖
‖ ‖ ‖𝑘𝑗 ‖
‖ ‖
Due to the irregular sampling locations, previous DCN leads to Another key component of the external attention mechanism is
increased spatial variance as the network deepens [28,45]. To retain value (𝑣), which is a linear mapping of deformation features 𝐷. Mul-
useful spatial information, we exploit attention mechanism to capture tiplying the attention weight matrix and the 𝑣, we finally obtain the
the correspondences between input and deformation features, resulting weighted features
in an attention-guided deformable convolutional module (ADCM).
𝑌̂ = 𝐴𝑣. (7)
As discussed in Section 2.3, it is not appropriate to apply atten-
tion mechanism directly to our network. How to design an effective Compared with the commonly used self-attention mechanism [32],
attention mechanism to guide the training procedure of DCN is a very our external attention mechanism mainly has two advantages. First,
important part in our research. Here, we mainly consider two points: the external attention is able to model the correspondences between
first, the attention mechanism must be computationally efficient as it the input and deformation features, while the self-attention mechanism
will be inserted into every deformable Conv layer. Second, it must can only calculate the relevance between each pixel of the input.
be capable of modeling the nonlinear relationship between input and Second, the external attention mechanism is more efficient because the
deformation features. relevance is calculated across channel descriptors rather than spatial
Based on the aforementioned analyses, there are four parts in the dimensions.
ADCM as shown in Fig. 2, which are channel-downscaling layer, in- Channel-upscaling layer. Corresponding to the channel-downscaling
ner attention mechanism, external attention mechanism, and channel- layer, the ADCM use channel-upscaling layer to recover the channel
upscaling layer. size. In addition, a residual connection is used to obtain the final
Channel-downscaling layer. To make the whole module lightweight output.
enough, the ADCM starts with a 1 × 1 Conv layer to reduce channel
size with ratio 𝑟. 3.4. Loss function
Inner attention mechanism. We employ an inner attention to fully
capture the interdependencies within the features. Unlike the previous To enhance the role of important features and mitigate the inter-
channel attention [6] that has a strong sensitivity to spatial characteris- ference of redundant noisy features, we increase the distance between
tics, we rescale channel-wise statistic instead of each pixel in the feature relevant and irrelevant features. Concretely, we introduce the concept
map. Given an input 𝑋 ∈ Rℎ×𝑤×𝑐 , the inner attention mechanism first of contrastive learning and propose a contrastive loss. Contrastive
utilize certain aggregation technique (e.g., global average pooling) to learning techniques have been proven to be useful in the field of
obtain a channel-wise statistic image SR [40–42]. In our case, we leverage contrastive loss to improve
1 ∑∑
ℎ 𝑤 the SR performance by pulling positive features closer and pushing
𝑋̂ = 𝐻𝐺𝑃 (𝑋) = 𝑋(𝑖, 𝑗), (3) negative features away in the feature space. Compared to the previous
ℎ × 𝑤 𝑖=1 𝑗=1
application of contrastive learning, we have three main differences.
where 𝑋̂ ∈ R𝑐 , 𝐻𝐺𝑃 (⋅) denotes the global average pooling function, The first difference lies in the application scenario. We have applied
𝑋(𝑖, 𝑗) is the pixel of 𝑋 at location (𝑖, 𝑗). contrastive learning to DCN for the first time, which helps to constrain

4
F. Qiao et al. Journal of Visual Communication and Image Representation 100 (2024) 104097

4.2. Implementation details

For classic SISR and SISR with blind noise, the DMFE number and
channel number of DCADCN are set to 5 and 64, respectively. For real-
world SISR, we train a larger model with the DMFE number of 10
and the channel number of 128. We use Adam optimizer to train our
DCADCN by setting 𝛽1 = 0.9, 𝛽2 = 0.999, and 𝜀 = 1e−8. There is a total
of 800 epochs in the training procedure. The mini-batch size is set to
Fig. 3. Illustration of the contrastive loss. For each ordered 𝐴𝑖 (𝐴𝑖 ∈ 𝐴, 1 ≤ 𝑖 ≤ 𝑐), the
top 𝑛 ⋅ 𝑐 features are viewed as positive parts and the last 𝑛 ⋅ 𝑐 features are viewed as
16. The learning rate is initialized to 5e−4 and reduced by half every
negative parts. 200 epochs. We warm up our DCADCN for the first 100 epochs with
only 𝑚𝑎𝑒 , and then train the network with the overall loss of . The
project is implemented by PyTorch 1.11 and trained on a Tesla V100
the correspondences between input and deformation features. The sec- GPU.
ond point concerns the difference in how similarity is calculated. While
the conventional methods calculate the similarity between different 4.3. Comparisons with state-of-the-arts
samples, our method focuses on calculating the similarity between
different channels of the same sample. The third point is the selection To comprehensively evaluate the SISR effectiveness, we compare
our approach with some state-of-the-art methods, including LESR-
of positive and negative parts. In our contrastive loss, we choose feature
CNN [3], ACNet [8], RepRFN [23], VLESR [52], FDIWN [24], Shuf-
channels with high relevance as positive parts and ones with low
fleMixer [21], ARRFN [19], FDSCSR [25], RLFN [53], DefRCN [10],
relevance as negative parts.
and RFDN [22]. For fairness, all compared methods use the same
As illustrated in Fig. 3, given the external attention weight matrix training set and down-sampling method. We use both quantitative and
𝐴 ∈ R𝑐×𝑐 of the ADCM, the contrastive loss 𝑐𝑜𝑛 can be formulated as: qualitative analysis to perform experiments.
Quantitative analysis: The PSNR and SSIM of different methods are
shown in Table 1. As we can see, our DCADCN performs better against
𝐴̂ 𝑖 = sort(𝐴𝑖 , Descending), 1 ≤ 𝑖 ≤ 𝑐, 𝐴𝑖 ∈ 𝐴, 𝐴̂ 𝑖 ∈ 𝐴,
̂ (8)
other compared methods on almost all datasets and scale factors. For
instance, compared to ACNet [8], our DCADCN achieves a gain both
∑𝑛⋅𝑐 ( )
1∑
𝑐
𝑗=1 exp 𝐴̂ 𝑖,𝑗 ∕(𝑛 ⋅ 𝑐) of PSNR of 0.39 dB and SSIM of 0.0019 for ×2 SISR on Set5 with
𝑐𝑜𝑛 = − log ∑𝑐 ( ) , (9) less computational cost (see Fig. 9). It is essential to point out that
𝑐 𝑖=1
𝑗=𝑐−𝑛⋅𝑐+1 exp 𝐴̂ 𝑖,𝑗 ∕(𝑛 ⋅ 𝑐) + 𝑚
our approach outperforms DefRCN [10] (also DCN-based method) on
where 𝐴̂ 𝑖 stands for the descending sort of 𝐴𝑖 , 𝐴𝑖 and 𝐴̂ 𝑖 are the 𝑖th all benchmark datasets for three scales. This benefits from the unique
row of 𝐴 and 𝐴, ̂ respectively. 𝑛 represents the percentage of positive network architecture (i.e., ADCM, DMFE, and contrastive loss) of our
and negative feature channels which is less than 50%. 𝑚 is a margin DCADCN. These observed results also show the effectiveness of our
constant and manually takes the value of 1e−6. DCADCN for reconstructing accurate SR images.
Qualitative analysis: Some visual information, such as texture, struc-
As a consequence, the overall loss function  of DCADCN consists of
ture, clarity of definition, and edge information, can be used to in-
two parts: one is commonly used mean absolute error (MAE) loss 𝑚𝑎𝑒 ,
tuitively evaluate the reconstruction effect. To display the qualitative
the other is contrastive loss 𝑐𝑜𝑛 , which is formulated as
results, we use officially released codes of the compared methods to
 = 𝑚𝑎𝑒 + 𝜆𝑐𝑜𝑛 , (10) predict the SR images. Fig. 4 shows visual comparisons of image details
on different datasets (×4) between our DCADCN and other competitors.
The original images ‘‘barbara’’, ‘‘253027’’, and ‘‘img030’’ are selected
𝑚𝑎𝑒 = ‖ ‖
‖𝐼𝐻𝑅 − 𝐼𝑆𝑅 ‖1 , (11) from Set14, B100, and U100, respectively. From the enlarged view,
we can observe that our method can reconstruct super-resolved (SR)
where ‖⋅‖1 is the 𝐿1 norm, 𝐼𝑆𝑅 and 𝐼𝐻𝑅 are the predicted SR image and images that are closer to the original HR images than the competitors.
the target HR image, respectively. 𝜆 is the penalty parameter which is For example, in the reconstructed images of ‘‘253027’’, our method
manually set up as 1e−3. restores clearer lines on the animal, while some other methods would
produce blurred lines with severe artifacts. These visual comparisons
can further demonstrate that our DCADCN significantly outperforms
4. Experiments compared SISR methods.

4.1. Datasets 4.4. Ablation study

We undertake extensive experiments to investigate the effects of


Following commonly used methods (e.g., [3,10]), we use 800 images
ADCM, DMFE, contrastive loss, the 𝑛 value in contrastive loss, and
from DIV2K dataset [46] for training. To synthesize LR images, we use
model depth.
bicubic down-sampling method with different scale factors (i.e., ×2, ×3,
Effect of ADCM: For evaluating the effectiveness of ADCM, we present
and ×4) for classic SISR, the same degradation model as ACNet [8] for
the SISR results of DCADCN with (w) and without (w/o) ADCM in
SISR with blind noise, the same degradation model as BSRGAN [27] for Fig. 5. For a fair comparison, all experiments are conducted under the
real-world SISR. Data augmentation is performed by horizontally flips same parameters and environment. As can be seen, the PSNR values of
and rotation of 90◦ , 180◦ , 270◦ . Besides, each HR image is cropped DCADCN w/o ADCM are lower than those by the method w ADCM.
into patches of 192 × 192. For testing, we use five standard bench- This is because the proposed ADCM can not only generate features
mark datasets: Set5 [47], Set14 [48], BSD100 (B100) [49], Urban100 in a flexible way but also adaptively weight deformation features by
(U100) [50], and Manga109 (M109) [51]. PSNR and SSIM metrics on calculating the relevance between input and deformation features. The
the Y channel in the transformed YCbCr space are used to evaluate SISR ADCM enables the deformation features to keep a certain degree of
performance. spatial relevance with the original input.

5
F. Qiao et al. Journal of Visual Communication and Image Representation 100 (2024) 104097

Table 1
Quantitative results on five benchmark datasets for different scale factors. The best and second best are highlighted in bold and underlined, respectively.
Method Params Scale Set5 Set14 B100 U100 M109
PSNR/SSIM PSNR/SSIM PSNR/SSIM PSNR/SSIM PSNR/SSIM
LESRCNN [3] – 37.65/0.9586 33.32/0.9148 31.95/0.8964 31.45/0.9206 –/–
ACNet [8] – 37.72/0.9588 33.41/0.9160 32.06/0.8978 31.79/0.9245 –/–
RepRFN [23] 386K 37.99/0.9609 33.57/0.9179 32.18/0.9004 31.95/0.9261 38.80/0.9774
VLESR [52] 311K 38.01/0.9605 33.58/0.9177 32.16/0.8993 32.14/0.9280 38.75/0.9770
FDIWN [24] – –/– –/– –/– –/– –/–
ShuffleMixer [21] 394K 38.01/0.9606 33.63/0.9180 32.17/0.8995 31.89/0.9257 38.83/0.9774
×2
ARRFN [19] 988K 38.01/0.9606 33.66/0.9179 32.20/0.8999 32.27/0.9295 –/–
FDSCSR [25] 466K 38.02/0.9606 33.51/0.9174 32.18/0.8996 32.24/0.9288 38.67/0.9771
DefRCN [10] – 38.02/0.9596 33.58/0.9151 32.21/0.8998 32.20/0.9286 –/–
RFDN [22] 534K 38.05/0.9606 33.68/0.9184 32.16/0.8994 32.12/0.9278 38.88/0.9773
RLFN [53] 527K 38.05/0.9607 33.68/0.9172 32.19/0.8997 32.17/0.9286 –/–
DCADCN [ours] 831K 38.11/0.9607 33.69/0.9191 32.21/0.8999 32.27/0.9297 38.85/0.9769
LESRCNN [3] – 33.93/0.9231 30.12/0.8380 28.91/0.8005 27.70/0.8415 –/–
ACNet [8] – 34.14/0.9247 30.19/0.8398 28.98/0.8023 27.97/0.8482 –/–
RepRFN [23] 392K 34.33/0.9272 30.30/0.8415 29.08/0.8058 27.95/0.8473 33.48/0.9434
VLESR [52] 319K 34.40/0.9272 30.34/0.8415 29.08/0.8043 28.16/0.8519 33.61/0.9445
FDIWN [24] 645K 34.46/0.9274 30.35/0.8423 29.10/0.8051 28.16/0.8528 –/–
ShuffleMixer [21] 415K 34.40/0.9272 30.37/0.8423 29.12/0.8051 28.08/0.8498 33.69/0.9448
×3
ARRFN [19] 996K 34.38/0.9272 30.36/0.8422 29.09/0.8050 28.22/0.8533 –/–
FDSCSR [25] 471K 34.42/0.9274 30.37/0.8429 29.10/0.8052 28.20/0.8532 33.55/0.9443
DefRCN [10] – 34.41/0.9263 30.34/0.8388 29.10/0.8044 28.16/0.8519 –/–
RFDN [22] 541K 34.41/0.9273 30.34/0.8420 29.09/0.8050 28.21 0.8525 33.67/0.9449
RLFN [53] – –/– –/– –/– –/– –/–
DCADCN [ours] 838K 34.46/0.9276 30.38/0.8427 29.11/0.8054 28.24/0.8534 33.76/0.9454
LESRCNN [3] 516K 31.88/0.8903 28.44/0.7772 27.45/0.7313 25.77/0.7732 –/–
ACNet [8] 1283K 31.83/0.8903 28.46/0.7788 27.48/0.7326 25.93/0.7798 –/–
RepRFN [23] 402K 32.15/0.8952 28.63/0.7824 27.60/0.7377 26.09/0.7834 30.52/0.9075
VLESR [52] 331K 32.17/0.8945 28.55/0.7802 27.55/0.7345 26.03/0.7830 30.48/0.9073
FDIWN [24] 664K 32.17/0.8941 28.55/0.7806 27.58/0.7364 26.02/0.7844 –/–
ShuffleMixer [21] 411K 32.21/0.8953 28.66/0.7827 27.61/0.7366 26.08/0.7835 30.65/0.9093
×4
ARRFN [19] 1008K 32.22/0.8952 28.60/0.7817 27.57/0.7355 26.09/0.7858 –/–
FDSCSR [25] 478K 32.25/0.8959 28.61/0.7821 27.58/0.7367 26.12/0.7866 30.51/0.9087
DefRCN [10] 1900K 32.21/0.8936 28.59/0.7810 27.57/0.7356 26.04/0.7841 –/–
RFDN [22] 550K 32.24/0.8952 28.61/0.7819 27.57/0.7360 26.11/0.7858 30.58/0.9089
RLFN [53] 543K 32.23/0.8961 28.61/0.7818 27.58/0.7359 26.15/0.7866 –/–
DCADCN [ours] 846K 32.27/0.8957 28.63/0.7819 27.58/0.7361 26.11/0.7858 30.58/0.9099
DCADCN-L [ours] 1031K 32.36/0.8962 28.70/0.7836 27.62/0.7373 26.16/0.7876 30.66/0.9109

Table 2 Table 3
Ablation results investigating the effects of DMFE. ✗: delete, &: replace. Ablation study investigating the effects of contrastive loss.
Case index DMFE Conv ADCM Set5 Set14 B100 Contrastive loss Set5 Set14 B100 U100 M109
1 ✗ ✗ ✗ 30.31 27.12 26.51 w 32.19 28.61 27.56 26.07 30.50
2 & ! ✗ 31.97 28.42 27.46 w/o 32.27 28.63 27.58 26.11 30.58
3 & ✗ ! 32.18 28.58 27.56
4 ! ! ! 32.27 28.63 27.58

strong evidence for the effectiveness of contrastive loss in enhancing


the overall performance of our method.
Effect of DMFE: Table 2 shows the ablation results that investigate the Effect of 𝑛 in contrastive loss: We reconstruct several models of
effects of DMFE. Four different cases are considered: case 1 removes DCADCN by taking different 𝑛 values of 0%, 5%, 10%, 15%, 20%, and
DMFE entirely, case 2 replace DMFE with four cascaded Conv layers,
25%. All experiments are trained on DIV2K and tested on Set5 (×4), and
case 3 replace it with four cascaded ADCMs, and case 4 is our complete
the results are presented in Fig. 7. When 𝑛 increases from 0% to 15%,
DCADCN. To ensure validity, the channel dimensions of both cases
both PSNR and SSIM increase for taking more informative features as
2 and 3 are equal to the sum of two sub-paths in DMFE. As one
positive features and uninformative features as negative features. The
can see, DCADCN outperforms the other cases across all benchmark
model attains the best performance when 𝑛 = 15%. Therefore, the final
datasets. These results indicate that the inclusion of DMFE contributes
value of 𝑛 is taken as 15%. However, the increase of 𝑛 from 15% to
to improved performance by enhancing the structural differences of
25% causes performance degradation as a result of taking redundant
sub-networks, thereby promoting the learning of diverse features.
Effect of contrastive loss: We further investigate the effect of con- noisy information as positive features.
trastive loss on SISR performance. Table 3 presents the results obtained Model depth analysis: We explore the effectiveness of model depth
from our method w and w/o the incorporation of contrastive loss. We on SISR performance, mainly in terms of the number of DMFEs. Exper-
can observe that the method w contrastive loss consistently outper- iments are conducted by using one to eight DMFEs in DCADCN for ×4
forms the method w/o contrastive loss across all benchmark datasets. SISR. As shown in Fig. 8, the PSNR increases with the increasing num-
Furthermore, we visualize the feature maps in the layer before up- ber of DMFEs on every benchmark dataset. This phenomenon highlights
sampling of DCADCN to verify the enhancement of feature extraction the potency of suggested DMFE and the potential for our DCADCN
by contrastive loss. As shown in Fig. 6, the contrastive loss brings larger to be expanded to deeper networks (e.g., DCADCN-L with 8 DMFEs)
weights to key features and smaller weights to irrelevant features. By to provide even better outcomes. Nevertheless, the increase in DMFEs
assigning different weights, contrastive loss amplifies the role of key also leads to an increase in model size. According to the experimental
features and ignores irrelevant information. This finding serves as a statistics, for each additional DMFE, the number of model parameters

6
F. Qiao et al. Journal of Visual Communication and Image Representation 100 (2024) 104097

Fig. 4. Visual comparisons of DCADCN with other competitors on different benchmark datasets (×4).

Fig. 6. Visualization of feature maps in the layer before up-sampling of DCADCN.

increases by 0.19M. To balance the model size and performance, we


choose the number of DMFEs to be 5 as the default setting.

4.5. Computational cost analysis

The FLOPs and inference time are effective metrics to evaluate the
computational cost. Here, we compare these metrics of our DCADCN-
Fig. 5. Ablation study on ADCM.
L with several SISR methods including LESRCNN [3], ACNet [8],
RepRFN [23], VLESR [52], FDIWN [24], ShuffleMixer [21], ARRFN
[19], FDSCSR [25], RLFN [53], and RFDN [22] on Set5 of ×4 SISR.
FLOPs are evaluated on a 1280 × 720 HR image, inference time are
counted on testing Set5 (×4) in a feed-forward process.

7
F. Qiao et al. Journal of Visual Communication and Image Representation 100 (2024) 104097

better model ACNet-M, our DCADCN needs less computational cost (see
Fig. 9).
Qualitative analysis: The visual comparisons of different methods for
SISR with different noise on different benchmark datasets (×4) are
shown in Fig. 10 (𝜎 = 25) and Fig. 11 (𝜎 = 35). As we can see, our
method is better at removing noise and recovering clearer and more
accurate image textures at different noise levels. In contrast, other
competitors are susceptible to noise, resulting in blurry SR images.

4.7. Extension for real-world SISR

We also extend our network for SISR in real-world scenarios. To ver-


ify the effectiveness, we retrain our DCADCN using the low-quality data
synthesized under the same degradation conditions as BSRGAN [27]
and test it on the real-world SISR benchmark dataset RealSRSet [27].
Fig. 7. Average PSNR and SSIM vs. 𝑛 values in contrastive loss.
Since the ground-truth of RealSRSet is not available, we only provide
the visual comparisons with several state-of-the-art real-world SISR
methods, including ESRGAN [54], RealSR [55], DASR [56], and BSR-
GAN [27]. As shown in Fig. 12, our DCADCN obtains competitive or
even better visual effect with clear and sharp edges, whereas some
competitors introduce undesired artifacts. For example, ESRGAN suffers
from blurry and RealSR generates over-sharpened images.

5. Discussion

5.1. Difference to ADNet

The proposed DCADCN shares a few similarities with ADNet [31]


as they both use deformable Conv and attention mechanisms. However,
the two are intrinsically different. They are designed for different tasks.
Fig. 8. Average PSNR vs. number of DMFEs on five benchmark datasets.
DCADCN is for low-resolution image reconstruction, while ADNet is
for high-dynamic range imaging. The main differences between ADNet
and DCADCN can be summarized as follows. The first one is the
purpose of deformable Conv operation. It is used to align feature
As depicted in Fig. 9, our DCADCN obtains the highest PSNR among
maps in ADNet while for accurate deformation feature extraction in
all competitors, which reflects the effectiveness of our method in im-
DCADCN. The second one is about the attention mechanism. AD-
proving SISR performance. DCADCN outperforms competitive methods
Net uses spatial attention to acquire attention feature maps, while
by a PSNR margin of up to 0.53 dB with similar or even lower total
DCADCN uses joint internal and external channel attention to calibrate
numbers of FLOPs. What is more, our method takes less inference time
deformation features. Our channel-by-channel approach requires less
compared to most methods of comparison. For example, compared to
computational resources than the pixel-by-pixel computation of spatial
RFDN which has an inference time of 0.078 s, the inference time of
attention. The third is the cohesion pattern of deformable Conv and at-
DCADCN is just 0.052 s. This result can be attributed to the unique
tention mechanism. In ADNet, the attention module and the deformable
network architecture and component design of our method, such as Conv alignment module are connected in parallel, and the two branches
DMFE, ADCM, etc. Therefore, our DCADCN obtains the highest PSNR operate independently. In our approach, deformable Conv layer and
with faster inference speed and decent FLOPs, having a better trade-off attention are connected in tandem, which enables the attention mech-
between performance and computational cost. anism to better constrain and calibrate the learning process to obtain
more accurate features. The last one lies in the optimized loss function.
4.6. Extension for SISR with blind noise ADNet only uses a single L1 loss, whereas our approach incorporates a
joint L1 loss and a contrastive loss to effectively constrain the learning
We further extend our network to handle SISR with blind noise. features.
Following [8], the degradation model is formulated as 𝐼𝐿𝑅 = 𝐼𝐻𝑅 ↓𝑠
+𝑔, where 𝑔 is additive white Gaussian noise (AWGN) with level 𝜎. 5.2. Difference to DefRCN
Consistent with [8], we retrain our DCADCN using the LR/HR pairs
on DIV2K synthesized by the above degradation model. The down- Although both our method and DefRCN use deformable Conv for
sampling factor is set to4 during training. It is worth noting that other feature extraction, we observe that our DCADCN is completely different
settings are the same as Section 4.2. In addition, we compare our from DefRCN. Firstly, the network structures are different. DefRCN is a
method with LESRCNN [3] and ACNet-M [8] for tackling SISR with residual stack of deformable Conv layers without additional prior infor-
noise levels of 15, 25, 35 and 50. Since the LESRCNN did not perform mation. Whereas our approach not only uses the attention mechanism
experiment under this degraded condition, we retrained the LESRCNN to guide the deformation feature extraction process, but also intro-
using newly synthesized data to ensure the fairness of the comparison. duces contrastive learning to further constrain the feature information.
Quantitative analysis: Table 4 shows the superiority of our DCADCN Secondly, the ways of feature extraction are different. DefRCN uses
over other competitors for SISR with blind noise. For instance, our single deformable Conv layers, while our approach builds a DMFE that
DCADCN outperforms ACNet-M by up to 0.31 dB, 0.13 dB, 0.06 dB, respectively utilizes deformable Conv layers and normal Conv layers
0.18 dB in PSNR and 0.0091, 0.006, 0.0046, 0.0099 in SSIM on four to obtain complementary feature information. More importantly, our
benchmark datasets when 𝜎 = 15. Besides, compared with the previous DCADCN achieves greater performance with less computational cost.

8
F. Qiao et al. Journal of Visual Communication and Image Representation 100 (2024) 104097

Fig. 9. Average PSNR vs. FLOPs and inference time for SISR.

Fig. 10. Visual comparisons for SISR with noise level of 25.

Fig. 11. Visual comparisons for SISR with noise level of 35.

Fig. 12. Visual comparisons for real-world SISR on RealSRSet.

9
F. Qiao et al. Journal of Visual Communication and Image Representation 100 (2024) 104097

Table 4
Quantitative results for SISR with blind noise on different noise levels of 15, 25, 35, and 50. The best results are highlighted in bold.
Method Noise level Set5 Set14 B100 U100
PSNR/SSIM PSNR/SSIM PSNR/SSIM PSNR/SSIM
LESRCNN [3] 28.73/0.8177 26.56/0.7049 25.97/0.6557 24.29/0.7066
ACNet-M [8] 15 28.71/0.8177 26.59/0.7056 25.98/0.6563 24.33/0.7078
DCADCN [ours] 29.02/0.8268 26.72/0.7116 26.04/0.6609 24.51/0.7177
LESRCNN [3] 27.43/0.7849 25.65/0.6698 25.18/0.6203 23.55/0.6709
ACNet-M [8] 25 27.40/0.7848 25.67/0.6710 25.21/0.6221 23.63/0.6751
DCADCN [ours] 27.55/0.7911 25.77/0.6761 25.24/0.6262 23.78/0.6858
LESRCNN [3] 26.40/0.7580 24.88/0.6419 24.56/0.5954 22.99/0.6437
ACNet-M [8] 35 26.41/0.7587 24.93/0.6441 24.61/0.5970 23.05/0.6482
DCADCN [ours] 26.43/0.7630 24.98/0.6468 24.62/0.6005 23.15/0.6585
LESRCNN [3] 25.23/0.7246 24.00/0.6113 23.86/0.5681 22.25/0.6092
ACNet-M [8] 50 25.25/0.7254 24.03/0.6122 23.89/0.5693 22.33/0.6149
DCADCN [ours] 25.25/0.7307 24.07/0.6156 23.89/0.5731 22.40/0.6249

5.3. Strengths and limitations Declaration of competing interest

We propose a DCADCN for effective and efficient SISR. The main The authors declare that they have no known competing finan-
advantages of DCADCN are: (1) Improved image quality. The experi- cial interests or personal relationships that could have appeared to
mental results demonstrate that our method generates high-resolution influence the work reported in this paper.
images with improved visual quality compared to existing methods.
(2) Enhanced details. DCADCN can effectively recover fine details and Data availability
textures in SR images. The proposed ADCM enables to maintain the
overall structure and edge of the image during the SR process. (3) Data will be made available on request.
Wide range of application scenarios. Our method can effectively handle
several SISR tasks, such as classic SISR, SISR with blind noise, and real- Acknowledgments
world SISR. (4) Efficient computation. The proposed method has high
computational efficiency since it requires very low inference time to This document is the results of the research projects funded by the
process LR images. National Natural Science Foundation of China (No. 11571325) and the
While our method has made progress on classic image SR, SISR Fundamental Research Funds for the Central Universities, China (No.
with blind noise, and real-world SISR, there are still some limitations CUC2019 A002). This work is supported by Public Computing Cloud,
that need to be addressed: (1) Compared to some lightweight models, CUC.
our method requires more parameters and FLOPs. (2) When switching
tasks, our DCADCN needs to be retrained with newly synthesized References
data that has the same distribution, indicating it might not generalize
well to different types of images beyond those used for training. In [1] Y. Yoon, H. Jeon, D. Yoo, J. Lee, I.S. Kweon, Light-field image super-resolution
future research, we would like to explore more lightweight models and using convolutional neural network, IEEE Signal Process. Lett. 24 (6) (2017)
improve its ability to migrate across different data domains. 848–852.
[2] C. Ledig, L. Theis, F. Huszár, J. Caballero, A. Cunningham, A. Acosta, A. Aitken,
A. Tejani, J. Totz, Z. Wang, et al., Photo-realistic single image super-resolution
6. Conclusion using a generative adversarial network, in: Proceedings of the IEEE Conference
on Computer Vision and Pattern Recognition, 2017, pp. 4681–4690.
In this paper, we propose a dual contrastive attention-guided de- [3] C. Tian, R. Zhuge, Z. Wu, Y. Xu, W. Zuo, C. Chen, C.W. Lin, Lightweight image
formable convolutional network (DCADCN) for effective and efficient super-resolution with enhanced CNN, Knowl.-Based Syst. 205 (2020) 106235.
[4] B. Lim, S. Son, H. Kim, S. Nah, K. Mu Lee, Enhanced deep residual networks
SISR. Specifically, the joint inner and external attention mechanisms for single image super-resolution, in: Proceedings of the IEEE Conference on
enable the ADCM to model the correspondences between input and Computer Vision and Pattern Recognition Workshops, 2017, pp. 136–144.
deformation features and retain more spatial structural information. [5] Y. Tai, J. Yang, X. Liu, Image super-resolution via deep recursive residual
The dual mixed feature extractor (DMFE) utilizes both Conv layers network, in: Proceedings of the IEEE Conference on Computer Vision and Pattern
Recognition, 2017, pp. 3147–3155.
and ADCMs to obtain diverse and complementary features. Further-
[6] Y. Zhang, K. Li, K. Li, L. Wang, B. Zhong, Y. Fu, Image super-resolution using
more, DCADCN adopts contrastive loss to further amplify the role of very deep residual channel attention networks, in: Proceedings of the European
key features and mitigate the effect of redundant noisy information. Conference on Computer Vision, ECCV, 2018, pp. 286–301.
Experimental results show the superior performance of DCADCN on [7] Z. Zhang, X. Wang, C. Jung, DCSR: Dilated convolutions for single image
classic SISR, SISR with blind noise, and real-world SISR tasks, and super-resolution, IEEE Trans. Image Process. 28 (4) (2018) 1625–1635.
[8] C. Tian, Y. Xu, W. Zuo, C. Lin, D. Zhang, Asymmetric CNN for image
comprehensive ablation studies demonstrate the effectiveness of ADCM,
super-resolution, IEEE Trans. Syst. Man Cybern.: Syst. 52 (6) (2021) 3718–3730.
DMFE, and contrastive loss. Moreover, our DCADCN achieves a better [9] J. Dai, H. Qi, Y. Xiong, Y. Li, G. Zhang, H. Hu, Y. Wei, Deformable convolutional
trade-off between performance and computational cost. In the future, networks, in: Proceedings of the IEEE International Conference on Computer
we will explore the application of our model to other image restoration Vision, 2017, pp. 764–773.
tasks, such as deblurring and demosaicing. [10] Y. Zhang, Y. Sun, S. Liu, Deformable and residual convolutional network for
image super-resolution, Appl. Intell. 52 (1) (2022) 295–304.
[11] C. Zhao, W. Zhu, S. Feng, Superpixel guided deformable convolution network
CRediT authorship contribution statement for hyperspectral image classification, IEEE Trans. Image Process. 31 (2022)
3838–3851.
Fengjuan Qiao: Writing – original draft, Visualization, Validation, [12] Y. Huang, X. Hou, Y. Dun, J. Qin, L. Liu, X. Qian, L. Shao, Learning deformable
Methodology, Investigation, Formal analysis, Data curation, Conceptu- and attentive network for image restoration, Knowl.-Based Syst. 231 (2021)
107384.
alization. Yonggui Zhu: Writing – review & editing, Project adminis- [13] Y. Wang, J. Yang, L. Wang, X. Ying, T. Wu, W. An, Y. Guo, Light field image
tration, Funding acquisition. Guofang Li: Writing – review & editing, super-resolution using deformable convolution, IEEE Trans. Image Process. 30
Conceptualization. Bin Li: Software, Resources. (2020) 1057–1071.

10
F. Qiao et al. Journal of Visual Communication and Image Representation 100 (2024) 104097

[14] C. Dong, C.C. Loy, K. He, X. Tang, Image super-resolution using deep con- [36] F. Yang, H. Yang, J. Fu, H. Lu, B. Guo, Learning texture transformer network for
volutional networks, IEEE Trans. Pattern Anal. Mach. Intell. 38 (2) (2015) image super-resolution, in: Proceedings of the IEEE/CVF Conference on Computer
295–307. Vision and Pattern Recognition, 2020, pp. 5791–5800.
[15] J. Kim, J.K. Lee, K.M. Lee, Accurate image super-resolution using very deep [37] T. Chen, S. Kornblith, M. Norouzi, G. Hinton, A simple framework for contrastive
convolutional networks, in: Proceedings of the IEEE Conference on Computer learning of visual representations, in: International Conference on Machine
Vision and Pattern Recognition, 2016, pp. 1646–1654. Learning, PMLR, 2020, pp. 1597–1607.
[16] J. Kim, J.K. Lee, K.M. Lee, Deeply-recursive convolutional network for image [38] C.-H. Yeh, C.-Y. Hong, Y.-C. Hsu, T.-L. Liu, Y. Chen, Y. LeCun, Decoupled
super-resolution, in: Proceedings of the IEEE Conference on Computer Vision contrastive learning, in: European Conference on Computer Vision, Springer,
and Pattern Recognition, 2016, pp. 1637–1645. 2022, pp. 668–684.
[17] Y. Tai, J. Yang, X. Liu, C. Xu, Memnet: A persistent memory network for image [39] H. Wang, X. Guo, Z.-H. Deng, Y. Lu, Rethinking minimal sufficient representation
restoration, in: Proceedings of the IEEE International Conference on Computer in contrastive learning, in: Proceedings of the IEEE/CVF Conference on Computer
Vision, 2017, pp. 4539–4547. Vision and Pattern Recognition, 2022, pp. 16041–16050.
[18] J. Liang, J. Cao, G. Sun, K. Zhang, L. Van Gool, R. Timofte, Swinir: Image [40] G. Wu, J. Jiang, X. Liu, A practical contrastive learning framework for
restoration using swin transformer, in: Proceedings of the IEEE/CVF International single-image super-resolution, IEEE Trans. Neural Netw. Learn. Syst. (2023).
Conference on Computer Vision, 2021, pp. 1833–1844. [41] L. Wang, Y. Wang, X. Dong, Q. Xu, J. Yang, W. An, Y. Guo, Unsupervised
[19] J. Qin, R. Zhang, Lightweight single image super-resolution with attentive degradation representation learning for blind super-resolution, in: Proceedings
residual refinement network, Neurocomputing 500 (2022) 846–855. of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021,
[20] C. Dong, C.C. Loy, X. Tang, Accelerating the super-resolution convolutional pp. 10581–10590.
neural network, in: European Conference on Computer Vision, Springer, 2016, [42] Y. Zhang, L. Dong, H. Yang, L. Qing, X. He, H. Chen, Weakly-supervised
pp. 391–407. contrastive learning-based implicit degradation modeling for blind image
[21] L. Sun, J. Pan, J. Tang, Shufflemixer: An efficient convnet for image super-resolution, Knowl.-Based Syst. (2022) 108984.
super-resolution, Adv. Neural Inf. Process. Syst. 35 (2022) 17314–17326. [43] W. Shi, J. Caballero, F. Huszár, J. Totz, A.P. Aitken, R. Bishop, D. Rueckert,
[22] J. Liu, J. Tang, G. Wu, Residual feature distillation network for lightweight image Z. Wang, Real-time single image and video super-resolution using an efficient
super-resolution, in: European Conference on Computer Vision, Springer, 2020, sub-pixel convolutional neural network, in: Proceedings of the IEEE Conference
pp. 41–55. on Computer Vision and Pattern Recognition, 2016, pp. 1874–1883.
[23] W. Deng, H. Yuan, L. Deng, Z. Lu, Reparameterized residual feature network for [44] Y. Chen, J. Li, H. Xiao, X. Jin, S. Yan, J. Feng, Dual path networks, Adv. Neural
lightweight image super-resolution, in: Proceedings of the IEEE/CVF Conference Inf. Process. Syst. 30 (2017).
on Computer Vision and Pattern Recognition, 2023, pp. 1712–1721. [45] Z. Liu, B. Yang, G. Duan, J. Tan, Visual defect inspection of metal part surface
[24] G. Gao, W. Li, J. Li, F. Wu, H. Lu, Y. Yu, Feature distillation interaction weighting via deformable convolution and concatenate feature pyramid neural networks,
network for lightweight image super-resolution, in: Proceedings of the AAAI IEEE Trans. Instrum. Meas. 69 (12) (2020) 9681–9694.
Conference on Artificial Intelligence, Vol. 36, 2022, pp. 661–669, (1). [46] R. Timofte, E. Agustsson, L. Van Gool, M.-H. Yang, L. Zhang, Ntire 2017
[25] Z. Wang, G. Gao, J. Li, H. Yan, H. Zheng, H. Lu, Lightweight feature de- challenge on single image super-resolution: Methods and results, in: Proceedings
redundancy and self-calibration network for efficient image super-resolution, of the IEEE Conference on Computer Vision and Pattern Recognition Workshops,
ACM Trans. Multimedia Comput. Commun. Appl. 19 (3) (2023) 1–15. 2017, pp. 114–125.
[26] K. Zhang, W. Zuo, L. Zhang, Deep plug-and-play super-resolution for arbitrary [47] M. Bevilacqua, A. Roumy, C. Guillemot, M.L. AlberiMorel, Low-complexity
blur kernels, in: Proceedings of the IEEE/CVF Conference on Computer Vision single-image super-resolution based on nonnegative neighbor embedding, in:
and Pattern Recognition, 2019, pp. 1671–1681. Proceedings of the 23rd British Machine Vision Conference, BMVA Press, 2012,
[27] K. Zhang, J. Liang, L. Van Gool, R. Timofte, Designing a practical degradation pp. 1–10.
model for deep blind image super-resolution, in: Proceedings of the IEEE/CVF [48] R. Zeyde, M. Elad, M. Protter, On single image scale-up using sparse-
International Conference on Computer Vision, 2021, pp. 4791–4800. representations, in: International Conference on Curves and Surfaces, Springer,
[28] X. Zhu, H. Hu, S. Lin, J. Dai, Deformable convnets v2: More deformable, better 2010, pp. 711–730.
results, in: Proceedings of the IEEE/CVF Conference on Computer Vision and [49] D. Martin, C. Fowlkes, D. Tal, J. Malik, A database of human segmented natural
Pattern Recognition, 2019, pp. 9308–9316. images and its application to evaluating segmentation algorithms and measuring
[29] X. Wang, K.C. Chan, K. Yu, C. Dong, C. Change Loy, Edvr: Video restoration with ecological statistics, in: Proceedings Eighth IEEE International Conference on
enhanced deformable convolutional networks, in: Proceedings of the IEEE/CVF Computer Vision. ICCV 2001, Vol. 2, IEEE, 2001, pp. 416–423.
Conference on Computer Vision and Pattern Recognition Workshops, 2019. [50] J.B. Huang, A. Singh, N. Ahuja, Single image super-resolution from transformed
[30] Z. Luo, L. Yu, X. Mo, Y. Li, L. Jia, H. Fan, J. Sun, S. Liu, EBSR: Feature self-exemplars, in: Proceedings of the IEEE Conference on Computer Vision and
enhanced burst super-resolution with deformable alignment, in: Proceedings of Pattern Recognition, 2015, pp. 5197–5206.
the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, [51] Y. Matsui, K. Ito, Y. Aramaki, A. Fujimoto, T. Ogawa, T. Yamasaki, K. Aizawa,
pp. 471–478. Sketch-based manga retrieval using manga109 dataset, Multimedia Tools Appl.
[31] Z. Liu, W. Lin, X. Li, Q. Rao, T. Jiang, M. Han, H. Fan, J. Sun, S. Liu, 76 (20) (2017) 21811–21838.
Adnet: Attention-guided deformable convolutional network for high dynamic [52] D. Gao, D. Zhou, A very lightweight and efficient image super-resolution
range imaging, in: Proceedings of the IEEE/CVF Conference on Computer Vision network, Expert Syst. Appl. 213 (2023) 118898.
and Pattern Recognition, 2021, pp. 463–470. [53] F. Kong, M. Li, S. Liu, D. Liu, J. He, Y. Bai, F. Chen, L. Fu, Residual local
[32] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A.N. Gomez, Ł. Kaiser, feature network for efficient super-resolution, in: Proceedings of the IEEE/CVF
I. Polosukhin, Attention is all you need, Adv. Neural Inf. Process. Syst. 30 (2017). Conference on Computer Vision and Pattern Recognition, 2022, pp. 766–776.
[33] Y. Mei, Y. Fan, Y. Zhou, Image super-resolution with non-local sparse attention, [54] X. Wang, K. Yu, S. Wu, J. Gu, Y. Liu, C. Dong, Y. Qiao, C. Change Loy, Esrgan:
in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Enhanced super-resolution generative adversarial networks, in: Proceedings of
Recognition, 2021, pp. 3517–3526. the European Conference on Computer Vision (ECCV) Workshops, 2018.
[34] H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, H. Jégou, Train- [55] X. Ji, Y. Cao, Y. Tai, C. Wang, J. Li, F. Huang, Real-world super-resolution
ing data-efficient image transformers & distillation through attention, in: via kernel estimation and noise injection, in: Proceedings of the IEEE/CVF
International Conference on Machine Learning, PMLR, 2021, pp. 10347–10357. Conference on Computer Vision and Pattern Recognition Workshops, 2020, pp.
[35] M. Huang, Y. Liu, Z. Peng, C. Liu, D. Lin, S. Zhu, N. Yuan, K. Ding, L. Jin, 466–467.
Swintextspotter: Scene text spotting via better synergy between text detection [56] J. Liang, H. Zeng, L. Zhang, Efficient and degradation-adaptive network for
and text recognition, in: Proceedings of the IEEE/CVF Conference on Computer real-world image super-resolution, in: European Conference on Computer Vision,
Vision and Pattern Recognition, 2022, pp. 4593–4603. Springer, 2022, pp. 574–591.

11

You might also like